In spring 2026 the independent research institute METR faced an unusual problem. It was to evaluate a new AI model. Standard procedure: set tasks, measure performance, record the result.
The model had reached the ceiling of the test system. Its fifty-per-cent time horizon — the length of human working time at which a model still completes the task in half of all attempts — stood at a little over seventeen hours. The confidence interval, however, ran from eight and a half to fifty-five hours; at that spread the lower bound is the more honest number. And METR attached a note to the measurement: above sixteen hours, measurements with the current task suite are unreliable. The measured value therefore sits above the point at which the institute stops trusting its own instrument.
That is not marketing copy. That is a measurement institute having to rebuild its instrument because the technology outran the research.
The number you need to know
Anyone wanting to understand the development needs to internalise one figure: over the period 2019 to 2025, the length of tasks AI agents complete on their own doubled roughly every seven months — 196 days by METR's own calculation. For models from 2023 onwards the doubling time is four to five and a half months; for those from 2024 onwards, around three.
This is not a linear improvement but an exponential curve with a rising gradient. Extrapolate it over two years and you reach a result most annual plans do not price in: tasks that cost a human several weeks could fall into the autonomous range.
The paradox most people miss
Speed up one part of a process and the next-slowest part becomes visible. The term comes from computer science: Amdahl's law, formulated in 1967. The overall gain from speeding up a system is limited by the share that cannot be sped up.
Applied to organisations: When AI radically accelerates execution, the part requiring human judgement becomes the new bottleneck. Not because it gets worse — but because it is suddenly so much slower by comparison.
GitHub shows this at scale: in 2025 the platform recorded 1.12 billion contributions to public and open-source projects, thirteen per cent more than the year before, and just under a billion commits, up twenty-five per cent. How much of that is ultimately usable is nowhere measured reliably — which is itself part of the problem. More speed, more output, and the checking still lands on people.
Three observations from more than twenty years of transformation work
AI makes organisational failure visible — faster than any audit
When a company introduces AI-supported processes and the results disappoint, it is almost never the technology. It is whatever was not working before and now becomes visible faster. Unclear goals produce unclear results more quickly with AI. Decision paths that used to take weeks still take weeks — because the culture of alignment has not changed, only the tool.
McKinsey puts numbers on it. In its global AI survey of summer 2025, around six per cent of the organisations surveyed meet the definition of an AI high performer — they attribute more than five per cent of their EBIT to their use of AI. That is a narrow standard, and precisely for that reason a telling one. Two further figures, from an earlier study by the same firm, describe the gap: ninety-two per cent of companies planned to increase their AI spending over the following three years, while only one per cent of the executives surveyed described their own organisation as mature in its use of AI. The gap between willingness to invest and organisational maturity is not small. It is structural.
Responsibility does not delegate itself to algorithms
In a survey by the careers portal Resume Now of 968 US employees, seventy-two per cent said ChatGPT had given them better advice than their manager; forty-nine per cent found it more emotionally supportive during work-related stress. The provider is commercial and the survey is not certified as representative — the order of magnitude serves as an indication, not as evidence.
I read that not as a triumph of technology but as a finding about the quality of leadership — and as a warning. AI fills a vacuum. Where orientation is missing, where decisions are not communicated clearly, the machine steps in. Relieving in the short term, dangerous in the medium term: the capacity for human leadership atrophies when it is no longer needed.
Prioritisation is the new bottleneck skill
Modern systems can develop a solution starting from an unspecified problem. Humans supply the goal, no longer the method. The remaining performance gaps sit where someone must decide which goals should be pursued at all.
That is the new core skill: not execution, not even method — but the ability to ask the right questions. What are we trying to achieve? Which problems are worth solving? And, less comfortably: What do we stop doing? Without settling that, AI mainly serves to work faster on the wrong problem.
What leadership must deliver now
There is a temptation I observe in many companies: waiting for clarity. For better systems, for mature governance, for the right moment. It does not come. The curve does not wait.
Give orientation. Not as an instruction, but as a clear picture of how success will be recognised. AI systems cannot compensate for ambiguity in the goal. Leadership must supply that, more precisely than before.
Locate accountability. Especially when systems prepare decisions. Whoever implements a recommendation without understanding it still carries the consequence.
Enable learning loops. In an environment where capabilities double every four to five and a half months, an organisation's speed of learning is no longer a characteristic. It is the condition of survival.
Conclusion
Introducing tools without strengthening orientation, accountability and the capacity to learn merely accelerates the old system. An accelerated old system is not a transformation. It is a faster failure.
Self-improvement as a concept presupposes that the learning system is itself capable of learning. That holds for AI. It holds just as much for organisations.
Sources
- METR (29 Jan 2026): Time Horizon 1.1. Doubling time of the fifty-per-cent time horizon: 196 days across 2019–2025, around 89 days for models from 2024 onwards. Figures as of that blog entry; the continuously updated data page now shows different values. metr.org
- METR: Task-Completion Time Horizons of Frontier AI Models, updated continuously; entry of 8 May 2026, model Claude Mythos Preview (early), dated there 7 April 2026: fifty-per-cent time horizon 17.4 hours, 95-per-cent interval 8.5 to 55 hours, with the note “Measurements above 16 hrs are unreliable with our current task suite”. metr.org/time-horizons
- McKinsey & Company / QuantumBlack (November 2025): The state of AI in 2025 — Agents, innovation, and transformation. Survey of 25 June to 29 July 2025, 1,993 respondents in 105 countries, weighted by share of GDP. Source of the six-per-cent AI high performer figure (definition: more than five per cent of EBIT attributed to AI use). mckinsey.com (PDF)
- McKinsey & Company (28 Jan 2025): Superagency in the workplace. 3,613 employees and 238 C-level executives, surveyed October and November 2024, 81 per cent in the USA. Source of the ninety-two-per-cent and one-per-cent figures. mckinsey.com
- GitHub (28 Oct 2025): Octoverse 2025. 1.12 billion contributions to public and open-source projects, just under a billion commits, more than 180 million developers. github.blog
- Resume Now (4 Sep 2025): The AI Boss Effect. Own online survey of 18 June 2025 among 968 US employees; commercial provider, not certified as representative. resume-now.com
- Amdahl, Gene M. (1967): Validity of the Single Processor Approach to Achieving Large Scale Computing Capabilities. AFIPS Conference Proceedings, Bd. 30, Spring Joint Computer Conference, S. 483–485. The basis of Amdahl’s law.
Correction, 10 September 2026, extended 16 September 2026. Every figure has been traced back to its primary source; six were inaccurate and are corrected above: the confidence interval on the METR measurement given only at the top end, its date and point estimate, the doubling time assigned to the wrong window, two McKinsey figures from a different survey, a vendor poll presented as an analysis of several studies, and the GitHub number that was wrong twice over. One claim without a findable source has been removed. The article’s argument is unaffected.
This article offers a professional assessment and does not replace legal or management advice.