Hero imageMobile Hero image
  • Facebook
  • LinkedIn

September 23, 2026

Our author, Antoine Aymer, CTO for Quality Engineering & Testing at Sogeti, explains in Part 3 of this five-part blog series that a “good” AI system cannot be measured by a single score. Organizations must balance multiple dimensions, including functionality, reliability, security, privacy, sustainability, and humaneness, while considering the needs of different stakeholders. The key is to define what quality means for your specific use case and prioritize the risks that matter most.

Ask five people what a “good” AI system is and watch the answers scatter. The product lead means it gives correct, useful answers. The security officer means it cannot be tricked into leaking its instructions. The privacy officer means it never repeats one user’s data to another. The finance lead means it does not burn a fortune in compute per query. The ethicist means it treats every user fairly and never manipulates them. One phrase, “a good AI,” concealing at least six different systems, each certain theirs is the one that counts.

One score hides the trade-offs

Leaders crave a single grade, because one number can be reported and forgotten. Watch what that number hides. To assess an AI system honestly you have to look along separate axes at once: whether it is correct, whether it behaves consistently, whether it resists attack, whether it protects data, whether it is affordable to run, and whether it treats people decently. Call them functionality, reliability, security, privacy, sustainability, and humaneness. They do not rise together, and often they fight. A model tuned to be maximally helpful becomes easier to manipulate. One tuned to refuse anything risky becomes uselessly cautious and starts declining reasonable requests. A single score does not summarize those trades. It buries which one you chose to lose.

The point of splitting quality into six dimensions, with finer sub-qualities beneath each, is to drag those trades into the light. A model that scores brilliantly on giving correct answers and quietly fails on privacy will help a thousand users and expose one, and calling it an 8 out of 10 hides exactly that. The split forces you to see the exposure instead of averaging it away under the brilliance.

Whose ‘good’ are you grading?

The trap underneath is old, and anyone who has asked “who is my customer, truly?” has fallen in. A pharmaceutical company cannot answer whether its client is the patient, the doctor, the pharmacist, the insurer, or the regulator, because all five hold a different definition of a good outcome and the person who uses the drug is seldom the one who pays. An AI system has the same fracture. The user wants a helpful answer. The company wants no lawsuit. The regulator wants fairness it can audit. Serve one master’s definition of good and you can quietly betray another’s, and the model will do it fluently, sounding equally excellent to everyone in the room.

So quality is a question about whose harm you are counting. The method hands that question back to you on purpose: it asks you to weigh the six dimensions for this system, these users, this deployment, so that “damage” gets measured against the harm that actually matters here. When you assign those weights, the red map from the essay before recolors itself around your honest answer. The risks held still. Your eyes moved. A privacy failure you had shrugged at flares red once you admit your users handed the model their medical history.

Serve one master’s definition of good and you can quietly betray another’s.

Most teams discover they had been grading the dimension that dazzles a demo, the clever, fluent answer, and starving the one their users quietly depend on. Now you can see every failure mode, weighted by a definition of good you chose deliberately. Which forces the question that separates a professional from a coward: which of these risks do you dare to leave untested?

Discover how Sogeti helps organizations define, measure, and build trustworthy AI systems and secure your AI Trust & Assurance Assessment today

Antoine Aymer

Antoine Aymer

CTO for Quality Engineering & Testing, Sogeti

Read more articles

A confident answer is not a correct one, and your tests cann...

You know the launch you are afraid of. You are putting a language model in front of customers, or auditors, or your own …

The hardest problem in AI isn’t intelligence

Part 1 of a 4-part series on AI beyond intelligence. In this series, Kim Berg, Global CTO for Data & AI Engineering, exp…

Transforming enterprise software delivery through agentic AI

Learn how Agentic Software Engineering transforms software delivery by combining AI agents, governance, and modern devel…