Artificial intelligence is becoming a routine part of enterprise decision-making, helping organizations analyze contracts, negotiate with suppliers, evaluate data, and support strategic planning. But as businesses place greater trust in AI-generated recommendations, a familiar problem remains unresolved: Large language models can present incorrect information with remarkable confidence.

Rather than producing obvious hallucinations, today’s AI models often generate plausible, well-written responses that make errors more difficult to recognize. For organizations relying on AI to inform important business decisions, distinguishing confidence from accuracy is becoming an increasingly important challenge.

John Davie, founder and CEO of Buyers Edge Platform, encountered that problem while expanding AI use across his organization. His experience led to the development of CollectivIQ, a platform that compares responses from multiple leading AI models to help users identify areas of agreement, disagreement, and uncertainty before acting on the results.

TechNewsWorld spoke with Davie about why AI overconfidence concerns him more than traditional hallucinations, how enterprises can reduce the risks of AI-assisted decision-making, and whether consensus across multiple models can improve trust in AI-generated answers.

TechNewsWorld: As AI becomes more persuasive, what can organizations do to reduce the risk of employees acting on incorrect AI-generated information?

John Davie: To continue scaling AI effectively and responsibly, leaders should teach employees to interrogate and evaluate AI outputs by asking follow-up questions, challenging assumptions, and seeking alternative sources to verify responses. Teach employees that AI can lie and fabricate key data points.

If an answer is going to influence a supplier negotiation, a pricing decision, or a board presentation, don’t stop at the first response simply because it sounds convincing. Ultimately, though, I don’t think this can be solved through training alone.

If training alone isn’t enough, what should organizations do differently?

John Davie, Founder and CEO of Buyers Edge Platform

Davie: If we’re asking every employee to become an expert at detecting AI manipulation, something is bound to fall through the cracks. The technology itself has to provide more transparency into where answers come from and where uncertainty still exists.

That is what the CollectivIQ platform provides. Its result is transparent, consensus-driven intelligence instead of a single black-box answer.

It’s not enough to generate answers anymore. Instead, leaders need to strategically implement tools and processes that help employees assess how much confidence to place in those answers before acting on them.

 

Why are overconfident AI responses more dangerous than obvious hallucinations?

Davie: Early examples of AI errors were simple, like someone asking how many Rs are in the word “strawberry,” and the AI would say “two.” The same principle applies to glaring mathematical or logical fallacies.

However, a well-written fabricated answer that cites the right concepts and sounds genuinely thoughtful and confident is much more difficult to detect and more dangerous. And even when someone questions it, the model often doubles down instead of acknowledging uncertainty.

Researchers found that when they challenged AI outputs, the models became more persuasive in defending the wrong answer, trying to persuade the user that the output was correct.

Is this behavior becoming more common across today’s leading AI models?

Davie: Additional research suggests it is far from isolated, making it a potentially widespread issue for business leaders. One peer-reviewed study finds that AI models exhibit sycophantic behavior, or overconfident tendencies, in nearly 60% of queries. In nearly 15% of cases, the models abandon a correct answer in favor of an incorrect one after the human user expresses disagreement.

In other words, pushing back on the AI output doesn’t increase accuracy. It often results in a larger error. As AI becomes embedded across the enterprise, businesses are relying on it to support everything from financial analysis and contract reviews to supplier evaluations, strategic planning, and other major investments such as M&A due diligence.

That’s where the real productivity gains come from, but it only works if employees know when to trust and when to question the output. When AI confidently presents an incorrect answer, people are far more likely to accept it and build on it, allowing small mistakes to quietly influence important business decisions.

What led you to conclude that relying on a single AI model wasn’t enough?

Davie: Our employees were experimenting with AI and trying to embed it into their workflows, but everyone was taking a different approach. Each one liked a different AI model for different workflows. We had no visibility into which models people were using, at what cost, what data was being shared, or why one employee trusted one answer over another.

I realized this went beyond finding the perfect model to meet all employee needs. Instead, I recognized that the complex set of challenges we faced required a sophisticated solution housed under a single, governed platform. CollectivIQ allows us to leverage the strengths of multiple models while giving us visibility, security, transparency, and confidence in the outputs.

Users can query all the top LLMs at once, including ChatGPT, Claude, Gemini, and Grok. Comparing answers, understanding where models agree or disagree, and examining those differences helps inform important business decisions.

If leading AI models are trained on much of the same data, why does comparing their responses produce better results?

Davie: These models are trained on much of the same publicly available information. But they aren’t clones of one another. Each has a different architecture, weight, training method, and approaches to reasoning. They often arrive at the same conclusion for different reasons, or they may disagree altogether.

If several independent platforms reach the same conclusion through different reasoning paths, it provides a much stronger signal than relying on a single model’s opinion. But when those models disagree, that’s often the most valuable insight of all, because it tells you the answer may be more nuanced than it first appeared and deserves another look.

How should non-technical users interpret conflicting answers from multiple AI models?

Davie: No technical skills required. The user doesn’t have to choose between multiple competing models. CollectivIQ offers the non-technical employee the efficiency of one high-quality, consensus-backed answer without sacrificing transparency.

Think of it as assembling an advisory board rather than listening to a single expert. You’re getting the strongest collective thinking distilled into a single unified answer, the “Best of the Best” answer. We want users to see those disparities so they can conduct deeper research and ask more questions.

When AI models disagree, how can technology help users resolve those conflicts?

Davie: We took this a step further and created Argue Mode. So when two models disagree, they each get challenged to check their sources. Then they go back and forth behind the scenes. One model usually concedes to the other model. This is all transparent to the user and no special skills are required.

Share.
Exit mobile version