TheVortiq
Inteligencia Artificial

Confidence in AI Agent Evaluation Grows, but Failures Persist

A VentureBeat study reveals that companies with more experience in failures trust automated evaluation less, but advance faster toward autonomy.

August 13, 2026 · 5 min read

two hands touching each other in front of a pink background

TL;DR: Confidence in automated AI agent evaluation grew from 5% to 13% in July, but the production failure rate remained at 49%. Companies that have already suffered failures are the ones that advance most toward full autonomy.

What happened?

The second VentureBeat Pulse Research report on AI agent reliability, with data from July 2026, reveals a troubling paradox: while confidence in automated evaluation has soared, the production failure rate remains stubbornly stable. According to the study, based on 108 companies, 13% of organizations now fully trust automated evaluation, nearly tripling the 5% in June. However, 49% of companies deployed an agent or LLM feature that passed internal evaluations but failed in front of a customer, virtually unchanged from June's 50%. This is the first month with a methodology identical to the previous one, allowing trend analysis: what moved was confidence, not accuracy.

The data also shows that complaints about the lack of alignment between evaluations and real-world results fell from 29% to 19%, no longer being the main objection. But this optimism is not reflected in outcomes: the failure rate did not move. The report notes that the new confidence comes almost exclusively from companies that have not yet experienced a false-confidence failure: 24% of them fully trust automated evaluation, compared to only 4% of those that have been burned. This suggests that confidence is being driven by inexperience, not by better evaluations.

Why it matters

This phenomenon reflects a gap between confidence and evidence that has profound implications for AI adoption in enterprise environments. Historically, trust in automated systems has been a critical factor in technology adoption, but here we see a dangerous disconnect: companies trust tools whose effectiveness has not improved. This echoes the robotic process automation (RPA) bubble of the late 2010s, when companies adopted bots without properly evaluating their return on investment, leading to a wave of failures and a market reset.

The concentration of confidence in companies without failure experience is a classic pattern of optimism bias. In psychology, it is known as survival bias: those who have not suffered negative consequences tend to overestimate a system's reliability. This is especially dangerous in AI, where errors can be subtle and hard to detect until they impact the customer. Furthermore, the fact that companies that have suffered failures are the ones that advance most toward autonomy (85% already allow deployment without human intervention or are working on it, versus 61% of those that have not had failures) suggests that direct experience with failures does not slow automation but accelerates it. This could be because they seek more robust solutions, but also due to a kind of psychological 'sunk cost': they have already invested in AI and prefer to move forward rather than retreat.

The gap between evaluation and real-world results is not new; it was already documented in the first June report, which noted that evaluation coverage was insufficient. But the fact that confidence increases while the failure rate remains stable indicates that companies are prioritizing deployment speed over safety. This is reminiscent of the microservices crisis in the 2010s, when many companies adopted distributed architectures without adequate observability tools, leading to cascading failures in production.

Consequences for businesses

For technology leaders, these data imply that automated evaluation is not yet reliable for predicting real-world success. The gap between evaluation and real-world results remains a critical issue, and blindly trusting evaluations can lead to premature deployments that damage the company's reputation and customer trust. A recent example is the Air Canada case, whose AI chatbot provided incorrect information about baggage policies, resulting in a lawsuit and customer compensation. Although not mentioned in the report, this case illustrates the risks of insufficient evaluation.

Companies should adopt a hybrid approach: combine automated evaluations with human oversight at critical stages, especially in high-risk domains such as healthcare, finance, or customer service. Additionally, it is essential to implement production performance metrics and feed that data back into the evaluation cycle. Organizations that have suffered failures should view them as an opportunity to improve their evaluation pipelines, not as a justification to automate faster. Gradual automation, with canaries and progressive rollouts, can help detect problems before they affect all users.

Evaluation tools market

The evaluation tools market is maturing, but still has a long way to go. The proportion of companies without dedicated tools fell from 17% to 12%, indicating greater adoption. Platforms like Braintrust and DeepEval gained ground, according to the report, although no figures are specified. Additionally, the primary selection criterion shifted from cost to ease of integration (39% vs. 27% previously), indicating that companies are seeking solutions that fit their existing workflows. This shift is positive, as integration is a key factor for effective adoption, but it also suggests that companies are prioritizing convenience over effectiveness in predicting failures.

Historically, the software testing tools market has gone through similar phases: first focusing on cost, then on coverage, and finally on integration with CI/CD. AI evaluation tools are following a parallel path, but with the difference that AI is more complex and failures are less predictable. Companies should evaluate not only ease of integration, but also the tool's ability to align with real-world results, for example, through integration with production monitoring systems and the ability to learn from past failures.

Recommendations for readers

  • Do not fully trust automated evaluations; validate with production testing and human oversight. Implement a continuous monitoring system that detects deviations between what is evaluated and what is real.
  • If your company has already suffered a failure, do not assume that full automation is the answer; consider a gradual approach. Analyze the root causes of failures and adjust your evaluations accordingly. Experience with failures should be a driver for improving quality, not for accelerating deployment.
  • When choosing evaluation tools, prioritize integration over cost, but also assess their ability to align with real-world results. Look for tools that offer production testing and continuous feedback.
  • Foster a culture of learning and transparency: share failures within the organization so other teams can avoid similar mistakes. Confidence should be based on data, not on the absence of failures.
Confidence without evidence is dangerous; experience with failures should be a drive to improve, not to automate blindly.

Keep reading