{domain:"www.qualitydigest.com",server:"169.47.211.87"} Skip to main content

        
User account menu
Main navigation
  • Topics
    • Customer Care
    • Regulated Industries
    • Research & Tech
    • Quality Improvement Tools
    • People Management
    • Metrology
    • Manufacturing
    • Roadshow
    • QMS & Standards
    • Statistical Methods
    • Resource Management
  • Videos/Webinars
    • All videos
    • Product Demos
    • Webinars
  • Advertise
    • Advertise
    • Submit B2B Press Release
    • Write for us
  • Metrology Hub
  • Training
  • Subscribe
  • Log in
Mobile Menu
  • Home
  • Topics
    • Customer Care
    • Regulated Industries
    • Research & Tech
    • Quality Improvement Tools
    • People Management
    • Metrology
    • Manufacturing
    • Roadshow
    • QMS & Standards
    • Statistical Methods
    • Supply Chain
    • Resource Management
  • Login / Subscribe
  • More...
    • All Features
    • All News
    • All Videos
    • Training

The Inspection Passed. The Product Still Failed.

Too often, AI ‘human-in-the-loop’ describes a position in the workflow, not control over it

KULSUM/Adobe

Dan Leiva
Bio

CXAmplify

Wed, 08/05/2026 - 12:02
  • Comment
  • RSS

Social Sharing block

  • Print
Body

The vision system scanned the part and logged a pass. The SPC chart stayed inside its control limits. The line kept moving.

ADVERTISEMENT

Three weeks later, the customer complaint arrived. The defect had been there the whole time, sitting just inside the algorithm’s confidence threshold, close enough to spec that the model called it good.

Nobody did anything wrong. The camera worked. The model performed within its trained tolerance. The dashboard was green. The product still failed.

This is the accountability problem now facing quality organizations that have moved AI into inspection, supplier scoring, and CAPA workflows. The technology isn’t the issue. Most of these systems perform exactly as designed. The issue is that almost nobody redesigned who owns the decision once the machine started making it.

That isn’t a technology problem. It’s an operating model problem.

Presence is not authority

Most quality organizations already have a human somewhere in the AI-assisted process. A quality engineer reviews the exception report. A line supervisor gets a notification when the model flags a borderline call. On paper, that looks like a human in the loop.

In practice, “human in the loop” describes a position in the workflow, not control over it. Reviewing a dashboard isn’t the same as having the standing, the time, and the authority to stop what the dashboard is reporting.

I see this most clearly at the border between automated accept and automated reject. The model is confident on the easy calls and correctly should be. It’s the calls near the threshold, the ones that used to land on an experienced inspector’s bench, that now get resolved by a probability score instead of a person. The system executes because it’s built to execute, not because someone judged that execution was the right call in that instance.

Both the yield report and the customer complaint can be telling the truth at the same time.

I’ve watched this play out the same way within industries that look nothing alike on the surface. A supplier scorecard auto-approves a deviation because the statistical pattern resembles a thousand prior approvals. A claims system auto-clears a borderline case because the language pattern matches the approved template. A quality model auto-accepts a part because the defect sits one confidence point inside the threshold. In every case, the system did exactly what it was trained to do. In every case, the people who would have caught the exception were no longer positioned to see it in time.

The boundary your quality system doesn’t have

Every AI-assisted quality decision falls into one of three categories, and most organizations have never drawn a line between them.

Class 1: Some decisions are deterministic. A dimensional measurement against a hard tolerance. A barcode match. A torque value inside a fixed range. These are low-stakes, reversible, and rules-based. The system should handle them without a human in the way.

Class 2: Some decisions require judgment. A defect call near the confidence threshold. A supplier deviation request tied to a part with a thin performance history. A nonconformance that’s technically within spec but visually inconsistent with what the customer expects. The right answer depends on context the model doesn’t fully hold. A named person should own these, with the AI providing the recommendation.

Class 3: A smaller set of decisions are high stakes: anything touching a safety-critical characteristic, a regulatory requirement, or a potential field action. These are hard to undo and expensive to get wrong. A human owns the decision. The AI informs it; it doesn’t make it.

The failure pattern I see most in quality operations is the second category quietly getting treated like the first. Not through a deliberate decision. Through drift. The manual review step adds seconds to cycle time. Someone optimizes it away to protect throughput. Six months later, the model is auto-clearing calls nobody agreed it should be trusted to make alone, and the first sign of the problem is a customer complaint or an audit finding, not an internal alarm.

The fix starts with a document, not a policy statement. For every AI-assisted quality workflow, write down which class each decision belongs to, who the named owner is for Class 2 and Class 3 decisions, and what triggers a review of that classification. Model retraining, a new supplier, a spec change, or a complaint spike should all trigger a review automatically.

Nobody owns the full nonconformance

A defect rarely gets caught and resolved by one system or one team. Detection happens during inspection. Disposition gets decided by quality engineering. Root cause runs through manufacturing engineering. Supplier corrective action, if it applies, runs through supplier quality. The customer notification, if it reaches that point, runs through a completely different function.

Each team owns a station. Nobody owns what the stations produce together for the customer.

When a defect escapes, every station can show that its component worked. Inspection logged the reading. Disposition followed the documented criteria. The CAPA was opened and closed on time. The investigation becomes an audit of individual steps instead of an audit of the outcome, and the gap between the steps is where the failure actually lived.

The fix is what I call sequence ownership: one named person accountable for what the full chain produces for the customer, not just what each station logged. It should be someone who can explain what the customer experienced from first detection to final resolution, why the system was designed to produce that result, and what changes if it happens again.

When a decision travels through multiple systems, accountability must travel with it.

The line needs an andon cord for algorithms

Manufacturing already understands this problem better than most industries. The andon cord exists because someone on the floor needs the standing authority to stop the line the moment something looks wrong, without asking permission first.

AI-assisted quality systems need the same mechanism, extended to cover the calls the algorithm is making, not just the calls a person can see with their own eyes.

That mechanism needs four properties to be real instead of decorative:
• Authority, so the people closest to the part, not just plant management, can trigger it.
• Immediacy, so pulling it changes what happens now, not after a ticket gets reviewed next week.
• Traceability, so every activation is logged with a reason and an outcome, turning individual judgment calls into organizational learning.
• Protection, so using it never gets someone treated as the problem.

Miss any one of those properties and what you have built is an inbox, not an andon cord. It collects concerns without producing action, and the people who could see a problem coming learn to stay quiet instead.

Measure judgment, not just throughput

First-pass yield, automation coverage, inspection cycle time, and cost per unit inspected; these are legitimate metrics. They tell you whether the system is running. They don’t tell you whether the humans around it are still exercising judgment.

This is the efficiency trap, and it doesn’t announce itself with a bad quarter. It shows up as a good one. Yield improves. Variability tightens. Everything on the scorecard looks better, while the organization quietly loses the capability to catch what the model was never built to see.

Alongside the standard quality metrics, track a small set of key human indicators.

Override rate: How often are inspectors and quality engineers actually overriding the AI’s accept or reject call vs. simply approving it? Near zero isn’t a sign of a perfect model. It’s often a sign that people have stopped questioning the output.

Escalation quality: Are escalations catching real issues, or are they noise nobody trusts?

Explainability confidence: Can the person reviewing a flagged part actually answer why the model made that call, in plain language, without opening a data science ticket?

Andon activation rate and resolution time: Is the stop mechanism being used, and does using it actually change the outcome?

What you measure shapes what the system optimizes toward. If throughput is the only signal in the loop, you will get throughput, including in the quarters when throughput is the wrong outcome to optimize for.

Review these indicators on the same cadence you review SPC data, not once a year as part of a management review slide. Add an event-triggered review on top of the calendar: A model retrain, a new supplier, a spec change, or a complaint spike should each pull the decision boundaries back onto the table, the same way they would pull a process capability study back onto the table.

The question worth asking on your next walk of the floor

Quality organizations that hold up over the next decade will not be the ones that automate the most inspection points. They will be the ones that keep real authority over what that automation is allowed to decide alone.

That means a documented boundary between what the AI clears on its own and what a named person must sign off on. It means one owner accountable for the full nonconformance, not just each station along the way. It means a stop mechanism with teeth, available to the person standing closest to the part.

Walk your floor and ask one question: If a defect were flowing through your line right now because a model called it good, would anyone have the standing, the visibility, and the protection to stop it before it reaches the customer?

If the honest answer requires a meeting, the system is running your quality organization. You’re just watching it.

Add new comment

The content of this field is kept private and will not be shown publicly.
About text formats
Image CAPTCHA
Enter the characters shown in the image.

© 2026 Quality Digest. Copyright on content held by Quality Digest or by individual authors. Contact Quality Digest for reprint information.
“Quality Digest" is a trademark owned by Quality Circle Institute Inc.

footer
  • Home
  • Print QD: 1995-2008
  • Print QD: 2008-2009
  • Videos
  • Privacy Policy
  • Write for us
footer second menu
  • Subscribe to Quality Digest
  • About Us