Andrew Ho is leaving OpenAI after just eight months, convinced that large language models generalize poorly. Even in well-funded areas like programming, he says, their performance is inconsistent.
The root cause, according to Ho, is a lack of training data. Most skills that matter economically are barely represented in existing datasets.
"Most work is highly contextual and not easily encoded into a gradable environment; even if we can observe a 'golden path' taken by a human which we believe to be good, it's challenging to understand whether alternate, counterfactual paths produce good or bad outcomes," Ho writes.
Scaling alone won't fix this, he argues, and expects AI labs will have to spend more than $100 billion on targeted data collection in the years ahead. Ho is also skeptical of the sky-high valuations at frontier labs like OpenAI or Anthropic, which he says are chronically unprofitable because they have to keep pouring growing sums into new models just to stay ahead of cheaper rivals like Qwen or Kimi.
His first products target two areas. One is datasets for complex scientific analyses in bioinformatics, where even current models like GPT-5.6 Sol hit only about a 30 percent success rate, a topic he worked on at OpenAI. The other is datasets for everyday lab work, such as when researchers submit photos of experiments to AI models for evaluation. Chemistry, materials science, healthcare, and broader knowledge work are planned to follow.







