Hero imageMobile Hero image
  • Facebook
  • LinkedIn

September 11, 2026

Our author, Antoine Aymer, CTO for Quality Engineering & Testing at Sogeti, explains in Part 2 of this five-part blog series that the most damaging AI failures are often the quiet ones. Organizations should prioritize risks based on their likelihood and impact, rather than focusing only on obvious or high-profile failures.

Picture the AI disaster that keeps you awake. You almost certainly pictured a loud one: the chatbot that swears at a customer, the image generator that produces something grotesque, the screenshot that goes viral with your company’s name on it. That is the failure in every headline. And it will almost never be the one that ruins you.

The quiet failures are the deadly ones

Follow the logic and it turns cold. The loud, offensive, obviously wrong output is the one everyone anticipates, so it is the one you filtered, guardrailed, and red-teamed before launch. You watched for it precisely because you could imagine it. The failure that actually does the damage is the one you cannot picture, because it does not look like a failure at all. A language model that quietly agrees with a user’s mistaken assumption, because agreeing is pleasant and disagreeing is friction, is called sycophancy, and it ships a thousand wrong confirmations that each look like helpfulness. A model that states a fabricated regulation with total composure gets believed, acted on, and cited in a document a director signs unread. High likelihood, because these behaviors run on every ordinary request. Severe damage, because nothing about them looks alarming.
So your fear is a lagging indicator. The intensity of your worry about an AI failure is roughly proportional to how defended that failure already is, because worry is what provoked the defense. In 1953, the writer Theodore Sturgeon shrugged off critics who said ninety percent of science fiction was rubbish by replying that ninety percent of everything is rubbish. A language model obeys him: most of what it produces is fine, and the harm hides in the small remainder that looks exactly as fine as the rest. The dangerous region of the system is the calm region, the answers that read well and are wrong in ways no one thought to check. You are most exposed exactly where you feel most reassured.

You are most exposed exactly where you feel most reassured.

Grade every risk: likelihood times damage

Instinct cannot grade this for you, so you replace it with two cold questions asked of every risk in the catalogue. How likely is this failure in your system, given what it does and who uses it? And if it happens, how much damage does it do? A support bot that occasionally fabricates a store’s opening hours and a medical assistant that occasionally fabricates a dosage carry the same failure, hallucination, at wildly different costs. Likelihood, then damage, and you multiply them, because a risk earns your attention only when a real chance of happening meets a real cost of happening.

Grade all the risks that way and the flat catalogue resolves into a picture you can read at a glance: cool where a failure is rare or harmless, burning red where a probable failure meets a severe consequence. The value of the picture is that it sorts by real exposure and stays blind to how frightening a failure merely sounds. You will likely find that a handful of risks burn red, and the rest do not, and the reds sit where nobody had been looking, while the whole team was busy guarding against the swear word that was never going to slip through anyway.

But red means severe damage, and damage to whom, measured how? A leaked phone number, a biased hiring suggestion, and a wrong citation harm different people in different currencies. Before you trust a single red cell, you have to answer the question the picture quietly assumed: what does “good” even mean for this system, and whose good is it?

Turn AI risk into AI confidence. Explore and secure your AI Trust & Assurance Assessment today.

Read more articles

A confident answer is not a correct one, and your tests cann...

You know the launch you are afraid of. You are putting a language model in front of customers, or auditors, or your own …

The hardest problem in AI isn’t intelligence

AI agents like OpenClaw expose a shift from prompting to delegation, showing how autonomy and new interfaces reshape dig…

Transforming enterprise software delivery through agentic AI

Learn how Agentic Software Engineering transforms software delivery by combining AI agents, governance, and modern devel…