Andrew Ho training data startup targets AI’s generalization gap
Former OpenAI researcher Andrew Ho says AI labs may spend over $100 billion on targeted data as model scaling delivers uneven gains.
By Renata Fuchs · Policy Reporter
· 3 min read
Former OpenAI researcher Andrew Ho has left the company after eight months to start a training-data startup, arguing that bigger language models alone will not make AI systems generalize reliably. The Andrew Ho training data thesis is a direct bet on frontier AI economics: Ho says labs may need to spend more than $100 billion on targeted data collection in the years ahead, while funding, valuation and revenue for his new company were not disclosed.
Ho’s argument is that many valuable work tasks are too contextual to be captured well by current datasets. He has pointed to programming, one of the best-funded AI application areas, as an example where performance can still be uneven. In posts on X, Ho said many jobs do not fit neatly into environments where an AI system can be graded, and that even when a strong human workflow is visible, alternative paths can be hard to judge.
Why does Andrew Ho think AI labs need more training data?
Ho says existing training corpora do not represent enough of the practical skills that businesses would pay AI systems to perform. In his view, scaling models without collecting better task-specific data leaves models strong in benchmark-friendly domains and unreliable in work where feedback is ambiguous.
The startup’s first products are aimed at scientific work. One planned dataset covers complex bioinformatics analysis, an area Ho worked on at OpenAI through GeneBench Pro and where he says current systems such as GPT-5.6 Sol reach only about a 30% success rate. Another product targets routine lab work, including cases where researchers submit experiment photos to models for evaluation. Ho has said chemistry, materials science, healthcare and broader knowledge work are later targets.
Ho has also criticized the valuations of frontier AI labs including OpenAI and Anthropic, saying they face chronic unprofitability because they must keep spending on new models to stay ahead of lower-cost competitors such as Qwen and Kimi. That claim frames his startup less as an application-layer play and more as a bet that proprietary data pipelines become a major cost center for model companies.
What do other AI researchers say about model generalization?
Cambridge researcher Adam Hunt has made a similar case, saying his view of large language models has become more pessimistic. Hunt argues that newer models are becoming better in select areas, especially coding and difficult math, while language quality and basic logic have stalled or declined in some cases.
Hunt attributes part of that split to reinforcement learning. Code gives model builders clearer reward signals and more complete training examples than many other domains, making it easier to optimize. He argues that early signs of broad model generalization came from training over large text corpora, rather than evidence that the systems had developed general understanding. Hunt has put his confidence in that view at about 40% and has acknowledged that future technical advances, including combining specialized models, could change the result.
Google DeepMind researcher Tom Zahavy has offered another explanation in a position paper titled “LLMs can’t jump.” Zahavy argues that language models are useful for deduction and induction but struggle with creative abduction, the ability to propose a cause or explanation that has no prior linguistic pattern. He suggests action-controllable world models, which can run counterfactual experiments, as one possible path beyond that limit, with LLMs still playing a role inside broader AI systems.
The shared implication is that the next expensive input for AI may be less about generic web-scale text and more about data that captures real workflows, hard-to-grade decisions and domain-specific feedback. Ho’s startup is entering that market before it is clear whether the bottleneck is data, model architecture, or both.
This story draws on original reporting from The Decoder.