AI agent evaluation shifts from clean traces to cohort baselines
LangChain, Conviva and CoreWeave executives said enterprises are rethinking how they test agents, with cheaper judge models and mo...
By Renata Fuchs 1 month agoSection
LangChain, Conviva and CoreWeave executives said enterprises are rethinking how they test agents, with cheaper judge models and mo...
By Renata Fuchs 1 month agoThe Information says Google’s Frozen v2 chip could improve inference efficiency by 6 to 10 times, with deployment planned for 2028...
By Renata Fuchs 1 month agoThe District 9 director used Seedance 2.0 for a 13-minute sci-fi horror short and says he plans a feature in the same format.
By Wei-Lin Zhao 1 month agoAt VB Transform 2026, Zillow and Glean executives said enterprise AI gains depend on persistent context, cost controls and pre-exi...
By Renata Fuchs 1 month agoMicrosoft plans to add AMD's Helios platform to Azure in 2026, while Anthropic is reportedly testing AMD hardware for Claude workl...
By Wei-Lin Zhao 1 month agoJumpCloud says self-rated AI maturity fell from 40% to 23% in six months as IT teams encounter governance gaps around production a...
By Renata Fuchs 1 month agoAxios says Commerce, the NSA and the White House have considered sanctions, warnings and liability rules that could push U.S. comp...
By Renata Fuchs 1 month agoThe AI platform says a malicious dataset led to credential theft, while its own AI tools compressed incident analysis from days to...
By Renata Fuchs 1 month agoMoonshot said demand for its new Kimi K3 model reached the limits of available compute within 48 hours, prompting a pause on new p...
By Wei-Lin Zhao 1 month agoNaveen Ayalla argues enterprise AI failures are often rooted in ingestion, validation and access-control gaps rather than LLM limi...
By Colin Brandt 1 month agoAlibaba says its 2.4 trillion-parameter Qwen 3.8 model will ship with open weights soon, targeting coding and productivity workloa...
By Colin Brandt 1 month agoGenCeption adapts Alibaba’s Wan2.1 video model for depth, segmentation and pose tasks, with strong benchmark results from limited...
By Wei-Lin Zhao 1 month ago