
From Perfect Traces to Broken Products: The New Paradigm for Evaluating AI Agents
A conversation with an AI agent can score perfectly in an individual evaluation yet signal a broken product. Leaders from LangChain, Conviva, and CoreWeave explained at VB Transform 2026 that the industry is moving from isolated traces to contrastive analysis across user cohorts, along with cheaper, more specific judge models.


