Jul 29, 2026
AI

Waymo AI evals drive release readiness, engineering leader says

Waymo’s Manasi Joshi said AI projects are judged by test maturity, with continuous evals, curated data and human release reviews.

Colin Brandt

By Colin Brandt · Enterprise Reporter

· 4 min read

Waymo AI evals are central to how the Alphabet autonomous-vehicle company decides whether a project is ready for production, Manasi Joshi, Waymo’s director of engineering for systems intelligence and machine learning, said at VB Transform 2026. The point is material beyond robotaxis: Waymo is applying AI in a setting where model failures can affect people on public roads, not only workflows inside software.

Joshi said Waymo has organized development around what she described as eval-forced or eval-centric development. In that model, a project’s maturity is judged in part by the maturity of the evaluations built around it, rather than by model performance alone. For companies building AI agents for support, coding, finance or operations, the read-through is direct: a system that cannot be measured reliably is not ready for unsupervised use in a business process.

Waymo says it has logged more than 220 million fully autonomous, rider-only miles. The company also says its vehicles have produced 17 times fewer serious crash injuries than human drivers over the same distance. Those are company-provided figures, and Joshi tied the safety claims to Waymo’s evaluation approach, including the data used to test the systems.

How does Waymo test AI before launch?

Joshi said Waymo evaluates models during training, after training and through both open-loop and closed-loop simulations. Open-loop tests typically examine system behavior against recorded or simulated inputs without letting the model change the scenario, while closed-loop simulation lets the system’s actions affect what happens next.

Waymo’s evaluation system uses datasets, metrics and infrastructure designed to run at scale. Joshi said evaluation is not treated as a single pre-launch checkpoint. It continues across driving, simulation and validation as models, data and operating conditions change.

The company’s testing hierarchy starts with safety. Waymo uses its own driving logs, some third-party data and simulation to put systems through scenarios covering billions of synthetic miles, according to Joshi. Task owners select narrower datasets and measures for higher-risk or more complex cases, including vulnerable road users, railroad crossings, construction areas and other difficult road situations.

Human review remains part of the release process

Joshi said Waymo does not leave deployment calls entirely to automated tools. Production-readiness reviews include human oversight, and internal safety leaders approve software releases and expansions into new service areas. She said human lives are at stake, which is why the process is not fully automated.

That is the part many enterprise AI programs have been slow to formalize. If an agent can trigger refunds, change records, write production code or influence customer communications, the company needs named owners who can decide when the system is ready and when it should be held back.

Waymo is also using agents inside engineering

Joshi said Waymo uses AI agents as internal tools for engineers. The agents help examine data distributions, judge data efficiency and sort through issues found in vehicle telemetry, training runs and failed evaluation jobs. The stated goal is to reduce time spent on investigation and leave engineers more time for technical judgment.

Those internal agents are also evaluated, according to Joshi. The concern is that an inaccurate tool can send engineers toward the wrong diagnosis, even if it appears to save time.

Waymo is also trying to limit the infrastructure cost of its AI work. Joshi said demand for compute, storage, memory and networking is rising faster than available resources. The company is working on efficiency across data extraction, storage, distributed training, model distillation, simulation and evaluation, while selecting training examples for usefulness rather than volume alone.

Waymo began using transformers in 2017 and has since expanded its work into large language models, vision-language models and vision-language-action models, Joshi said. She said generative multimodal models now sit within the company’s foundation-model strategy. The broader lesson for AI operators is less about model branding than about controls: define the outcome, test against representative data, keep evaluating after launch and make accountable humans part of the release path.

This story draws on original reporting from VentureBeat.

More from AI

All AI →