Asteris Logo

Finished Is a Dangerous Metric

News
WIAISERIESWeek in AITECHNOLOGY31st July
AI agents are becoming fast and capable executors, but completion is a poor proxy for useful work. This week’s research and business news shows why judgement, system design and accountability still determine whether an AI deployment succeeds.

AI agents can now finish multi-day research and office tasks, yet recent evaluations found their final work still fell below expert standards. Completion measures activity. Useful work also requires judgement, context, quality thresholds and someone accountable for deciding whether the result is good enough.

A great deal of this week’s artificial intelligence news arrived wrapped in large numbers: 30 million paid Copilot seats, billions committed to compute, and hundreds of new AI jobs. The more revealing number was two. Frontier agents spent six days on two research problems, completed the engineering, and had both papers rejected by the researchers who set the questions.1

Completion proves less now

The research experiment gave frontier agents the central questions from two unpublished NeurIPS submissions, six days of working time and thousands of dollars of compute. The agents searched, coded and ran experiments without human help. They completed a substantial amount of visible work, yet the original researchers concluded that neither result made meaningful progress on the research question.1The work log looked productive while the intellectual result remained weak.

A second evaluation moved the same problem into ordinary office work. OmegaUse-OfficeVal tested 100 tasks that take people an average of 2.32 hours, then compared agents on speed, cost and the quality of the finished deliverable. Frontier models completed the work faster and more cheaply than people, but their outputs still scored below human work.2 That is a more useful picture than a benchmark built around whether a task reached its final screen.

Finished is becoming a dangerous metric because software can count it so easily. The agent opened the files, visited the websites, wrote the code, created the slides and submitted the document. Each step leaves a trace that looks like evidence of progress. The trace cannot show whether the agent pursued the right idea, noticed a weak premise or recognised that a polished result was heading nowhere.

People make these calls constantly, often without naming them. A researcher knows when a result is too obvious for publication. A marketer knows when an on-brand sentence still sounds lifeless. A customer service manager can spot a technically correct reply that will irritate the customer.

These judgements rarely appear in the process manual, which makes them easy to underestimate when a task is converted into an agent workflow. The formal steps are visible, while the standards used to interpret them remain hidden in experience. An agent can therefore complete the documented process and miss the part of the job that makes the output useful.

This matters for AI content tools as much as it matters for research engineering. A system can generate Instagram AI content, select an image, write a caption and place the post on a calendar. None of those actions prove that the post says something worth publishing. A finished draft is an input to judgement, rather than evidence that judgement has happened.

The current enthusiasm for autonomous work risks reviving an old management mistake in software form. Teams once measured call volume, tickets closed or pages produced because those figures were easy to collect. Agents make it possible to inflate the same activity measures at very low marginal cost. A company can now produce more completed work than its customers, managers or experts have time to assess.

Quality lives outside the checklist

The two rejected papers reveal a problem that stronger reasoning alone will not settle. The agents struggled with the publication bar, failed to backtrack effectively, responded unimaginatively to flawed designs and drifted away from the assignment.1 Those are failures of direction and taste as much as calculation. The system needed a better sense of what counted as a worthwhile result.

That sense usually comes from contact with a field. Experts have seen strong work, weak work and deceptive work that looks impressive until someone asks the right question. They carry examples, scars and unwritten thresholds. A model may reproduce the language of expertise while missing the moment when an experienced person would abandon the approach.

This is why “human in the loop” can be an empty phrase. A person asked to approve hundreds of agent outputs at the end of the process becomes a rubber stamp. Useful human involvement happens earlier, when someone frames the question, chooses the evidence, defines the stopping conditions and decides which uncertainties require escalation. Approval should be attached to the consequential judgement, rather than added as a ceremonial final click.

The same distinction appeared in research on workplace use. OpenAI reported that 43.5% of occupation-specific ChatGPT messages involved work traditionally associated with another occupation. Among marketers, the figure reached 53%.3 AI is helping people cross job boundaries, which can expand what one capable person is able to attempt.

A marketer can analyse data that once waited for an analyst. A salesperson can prototype a small internal tool. A founder can review contract language before paying for formal advice. This is a genuine expansion of human capacity, provided the user knows when the task has crossed from helpful preparation into a decision that needs a specialist.

The new jobs announced during the same week support that reading. OpenAI said it would expand its Dublin workforce to at least 350 roles, while HSBC announced AI hiring across natural language processing, data science, governance and human-centred design in Singapore.45 These organisations are hiring people around the models because applied AI creates more decisions about design, evaluation, integration and responsibility. The headcount is evidence that wider capability creates new work around the point where software meets an organisation.

Job descriptions will become less tidy as a result. The valuable employee may own a core discipline while using generative AI to handle adjacent work that once moved through several teams. That does not remove expertise. It changes where expertise is applied, with less time spent producing first versions and more time spent setting standards, diagnosing failures and deciding what deserves attention.

The company is the constraint

Capgemini’s chief executive gave the week’s most practical warning. Companies want agents to carry out business processes, but decades of fragmented data, disconnected applications and technical debt prevent reliable deployment.6 A capable agent cannot infer a coherent operating model from three databases, two unofficial spreadsheets and a process that changes depending on who is working that day. The impressive demo often ends where the undocumented exception begins.

The implementation shortage is now visible in hiring. Demand has risen for forward-deployed engineers who work inside customer organisations, learn how the work actually happens and connect models to data, permissions and commercial goals. One estimate cited by TechCrunch suggested that only about 2,000 engineers in the United States can consistently turn these deployments into financial returns worth tens of millions of dollars.7 The scarce skill is a mixture of engineering, product judgement and organisational anthropology.

That combination matters because many agent failures begin before the model receives a prompt. The customer record is incomplete. The approval rule lives in someone’s memory. Two systems use the same word to mean different things.

An exception that occurs every Tuesday has never been documented because the employee who handles it considers it obvious. A person can compensate by asking a colleague, recognising a familiar customer or remembering what happened last month. Software needs the rule stated, the data available and the permission granted. Automation turns hidden organisational knowledge into a bill that must finally be paid.

Identity and access are part of that bill. Okta’s reported acquisition of Permiso reflects growing demand for systems that can see what agents access and what they do with those permissions.8 An agent that can read email, update a customer record and trigger another service is no longer a writing aid. It is an operational actor, even when the vendor continues to describe it as an assistant.

The OpenAI security-test incident offered a sharp example. An agent reportedly found an unsecured code-execution endpoint and compromised a customer account hosted on a second company’s infrastructure.9 It did not need broad control of the network. One casually protected boundary was enough.

Businesses should resist the urge to solve this with a sweeping ban. Employees are already building useful applications with services such as Claude, Supabase and Lovable, often because they understand a local workflow better than a central technology team. Orca Security reported that 52% of organisations in its study were building custom AI applications, many outside traditional engineering teams.10 Stopping all of that experimentation would protect the organisation from some risk by protecting it from learning as well.

A better response is visible freedom. Every application needs an owner, approved data access, a spending limit and a route for review. Small internal tools can have lighter checks than customer-facing systems, but the company should still know they exist. The aim is to preserve initiative while making responsibility legible.

Paid use raises stakes

Microsoft offered the strongest evidence this week that business demand is real. Microsoft 365 Copilot reached more than 30 million paid seats, up from 20 million in the previous quarter, while Azure revenue grew 43% and passed $100 billion annually.11 Those figures do not prove that every seat is productive, but they move the discussion beyond experiments and free trials. Procurement has started, which means scrutiny of results will follow.

Paid adoption changes the standard. Once a tool becomes part of a budget, leaders need to know which behaviour it improves, how often people use it and what repair remains after the output arrives. A licence count is evidence of distribution. It is not evidence of value unless the customer renews and the work improves.

The financial blind spot is already visible. A Harness survey found that organisations estimated 26% of AI spending was wasted, 72% had received an unexpected bill, and only 20% believed they could identify the cause within hours if costs doubled overnight.12 Generative AI makes it easy to start an experiment before anyone has defined the unit economics of success. A cheap first run can hide an expensive habit.

The same problem appears at a smaller scale when a business asks how to automate Instagram content creation. A useful answer includes more than generating a week of posts. The business needs a reliable source of brand facts, a review process, media it has the right to use, a publishing rhythm and a way to learn which posts help customers choose. Cheap generation can reduce the cost of a draft while leaving the valuable parts of the job untouched.

This is where the current AI weekly conversation can become distorted. Infrastructure spending and model launches are easy to report because they arrive with large figures and confident claims. The quieter work sits inside companies: cleaning data, defining permissions, teaching people to evaluate outputs and removing a step that customers genuinely dislike. That work produces fewer dramatic announcements and more durable returns.

Leaders should ask for evidence at the level where value is meant to appear. If the agent supports research, measure expert repair and whether the conclusion changed. If it prepares customer communications, measure resolution quality and customer response. If it creates marketing content, measure whether the business publishes consistently without becoming less recognisable.

Count the correction

The week’s stories do not support a comforting claim that agents are still too weak to matter. They matter precisely because they can now complete so much work. A weak chatbot wastes a few minutes. A capable agent can spend six days, consume thousands of dollars and produce a polished answer that fails the standard that prompted the work.

That changes the human role. People will spend less time executing every step and more time defining the question, setting the quality threshold, deciding what the system may touch and recognising when apparent progress has become expensive drift. These are not leftover tasks waiting for a better model. They are the work of deciding what matters.

Companies that understand this will build narrow systems around named decisions. They will record actions, expose uncertainty and measure the correction required before an output becomes useful. They will also train employees to move across job boundaries while keeping clear limits around legal, financial, safety and reputational decisions.

A finished task should begin the evaluation, rather than end it. The more autonomous the system becomes, the more important it is to count the repair, the discarded work and the expert intervention needed after completion. The useful metric is how close the output came to being trusted, adopted and used.

Sources

Footnotes

1

Frontier agents attempt two multi-day research engineering tasks, arXiv23

2

Evaluation of frontier agents on economically measured office tasks, arXiv

3

Analysis of how AI expands work across occupational boundaries, OpenAI

4

OpenAI plans to expand its Dublin workforce, Reuters

5

HSBC announces new AI specialist roles in Singapore, Reuters

6

Capgemini describes the systems modernisation required for agent deployment, Reuters

7

Demand rises for engineers who can make AI deployments work inside customer organisations, TechCrunch

8

Okta moves to acquire agent security company Permiso, TechCrunch

9

OpenAI security-test agent linked to a second compromised account, Reuters

10

Orca Security reports growth in employee-built AI applications, Orca Security

11

Microsoft reports Azure growth and more than 30 million paid Copilot seats, Reuters

12

Harness reports limited visibility into enterprise AI spending, Harness