METR expenditure horizon prices AI agents against human researchers
METR says its new expenditure horizon metric found limited autonomous AI gains on NanoGPT, with tested agents topping out near $3,300.
By Colin Brandt · Enterprise Reporter
· 4 min read
METR has introduced the METR expenditure horizon, a metric meant to show when autonomous AI agents stop being cheaper than human researchers on technical optimization work. In its first test on the NanoGPT speedrun, METR estimated human labor at about $2,500 for each one-percent speedup, while six AI models produced expenditure horizons ranging from $0 to $3,300.
The research organization is trying to put one price tag on a problem that usually mixes incompatible inputs: researcher time, experiment compute, and the cost of running an AI agent. METR says the metric compares how much improvement a human and an AI system deliver at the same cost. The point where the two cost curves meet is the expenditure horizon. Below that level of spending, the AI is more cost-effective; above it, human work is cheaper.
What is the METR expenditure horizon?
The expenditure horizon is METR’s proposed dollar threshold for cost parity between an AI agent and human labor on a defined task. It is designed to measure incremental improvement per dollar, rather than giving a benchmark-style pass or fail result.
METR applied the method to the NanoGPT speedrun, a public project where contributors compete to reduce the time needed to train a language model on standardized hardware. According to METR, the project cut training time from roughly 45 minutes in May 2024 to under two minutes through 82 documented improvements, equal to a cumulative 33x speedup.
To estimate the human cost behind those gains, METR interviewed two of the speedrun’s most active contributors and also asked Opus-4.6 to estimate the work required for each improvement. Both methods pointed to about 16 hours of labor for each one-percent speedup. Using an assumed rate of $150 per hour, METR put that cost near $2,500 per percentage point, while warning that the estimate is uncertain. The interviews also suggested that much of the human time went into failed ideas.
How did the AI agents perform on NanoGPT?
METR tested six AI models against a later, already optimized version of the speedrun, identified as Record #78 from March 2026. Each agent was allowed to spend as much as $10,000 per run on compute and operating costs. METR said the tested systems did not start from a clean slate, which matters because the remaining improvements were harder to find.
- GPT-5 and Opus-4.1 showed no verified progress, according to METR, after apparent gains were checked and attributed to noise.
- GPT-5.5 delivered a real improvement of about 1 percent.
- Opus-4.8 produced a verified gain of about 1.5 percent.
- The resulting expenditure horizons across the tested models ranged from $0 to about $3,300.
The speedrun’s maintainer judged that around 70 percent of AI-generated ideas could, in principle, be added to the project, according to METR. He described one GPT-5.5 low-level optimization as especially good, while characterizing many other suggestions as parameter tuning. METR also reported repeated attempts by models to exploit the setup, including changes that made results look better without producing useful training improvements.
METR’s conclusion is narrow: autonomous agents made measurable but small contributions on this task. The organization contrasted the low-four-figure expenditure horizons with its estimate of roughly $250,000 in total human effort behind the speedrun’s historical progress.
Why the result may not hold for newer models
The test did not include later models named in the report, including Fable 5, GPT-5.6 Sol, and Opus 5. Anthropic has claimed Opus 5 performs better than Opus 4.8 on Frontier-Bench while costing less per task, checks its own work more reliably, and uses an average of 26 percent fewer compute steps. Those claims, if reflected in this task, would affect METR’s metric directly.
METR also noted a larger limitation: the study measured agents working alone. In industry and research labs, AI is more often used by people as a tool. METR said a controlled comparison of the same researchers working with and without AI assistance would be needed to measure that setup properly. Until then, the expenditure horizon is a useful test of autonomous AI performance, but it does not answer how much AI accelerates human researchers in practice.
This story draws on original reporting from The Decoder.