{domain:"www.qualitydigest.com",server:"169.47.211.87"} Skip to main content

        
User account menu
Main navigation
  • Topics
    • Customer Care
    • Regulated Industries
    • Research & Tech
    • Quality Improvement Tools
    • People Management
    • Metrology
    • Manufacturing
    • Roadshow
    • QMS & Standards
    • Statistical Methods
    • Resource Management
  • Videos/Webinars
    • All videos
    • Product Demos
    • Webinars
  • Advertise
    • Advertise
    • Submit B2B Press Release
    • Write for us
  • Metrology Hub
  • Training
  • Subscribe
  • Log in
Mobile Menu
  • Home
  • Topics
    • Customer Care
    • Regulated Industries
    • Research & Tech
    • Quality Improvement Tools
    • People Management
    • Metrology
    • Manufacturing
    • Roadshow
    • QMS & Standards
    • Statistical Methods
    • Supply Chain
    • Resource Management
  • Login / Subscribe
  • More...
    • All Features
    • All News
    • All Videos
    • Contact
    • Training

AI May Be Making Your Quality Problem Worse

General-purpose AI does not eliminate human variation. It can amplify it and make weak reasoning harder to see

Dirk Dusharme/tiero/Adobe

Rick Heilshorn

MAD-Ai

Tue, 09/15/2026 - 12:03
  • Comment
  • RSS

Social Sharing block

  • Print
Body

Manufacturing has a people problem, but not in the way we usually talk about it.

ADVERTISEMENT

We talk about the skilled-labor shortage. We talk about retirements. We talk about doing more with fewer people. Usually those are treated as separate problems.

They aren’t. Rather, they are compounding.

Deloitte and The Manufacturing Institute estimate that 2.8 million manufacturing job openings through 2033 will come from retirements alone. Separate Manufacturing Institute research found that 97% of manufacturers expressed concern about the loss of institutional and technical knowledge as the workforce ages.

We know the knowledge is leaving. We’re just not very good at capturing it before it goes. APQC’s 2025 research found that only 8% of organizations consistently capture knowledge from retiring employees.

We can replace the person. We rarely replace everything that person knew.

Now look at the skills underneath the work. Recent national data show roughly one-third of the adult population performing at the lowest levels of numeracy and adaptive problem solving. Twenty-eight percent are now scoring at the lowest level of literacy. This isn’t a criticism of younger workers. It’s a workforce reality, and manufacturing is getting more complex at the same time.

Then add workload.

CADDi and SME’s “2026 Manufacturing Outlook Study” found that 79% of manufacturing leaders identified the skilled-labor shortage as their greatest challenge. A Quality Magazine “2024 State of the Profession” survey found that skilled labor was the most frequently cited job barrier among quality professionals, yet only 29% expected their quality-team head count to increase.

So the experienced people are leaving. Knowledge is walking out with them. Core skills are under pressure. Head count isn’t keeping pace. And the people who remain spend more of their day firefighting and trying to get the next thing off their desk.

What should we expect the output of that system to look like?

We’ve been living with variation for years

Put the same corrective action in front of three experienced quality engineers and ask them to grade it against the same standard. There’s a good chance you’ll get three different answers. How can they all be correct? More concerning, it’s possible that none of them are.

That’s the part the quality industry has been living with for a long time.

A 2019 automotive attribute-agreement study by Carmen Simion provides a useful example. Three trained inspectors evaluated the same 30 daytime running lights, twice, against the same known standard.

All three agreed with one another on only 36.67% of the parts. Their individual effectiveness against the known standard was 80.00%, 56.67%, and 70.00%.

The company improved its defect samples and operational definitions, retrained the inspectors, and achieved 95% agreement afterward. That follow-up was based on 19 of 20 parts from the original 30, so it was a small and previously exposed sample. But the improvement was still significant.

That’s exactly what manufacturing does. We find variation, improve the standard, train the process, and measure again.

We’ve spent decades doing this around physical work and inspection. Standardized work, control plans, poka-yoke, check fixtures, go/no-go gauges, automated inspection, and MSA all attack the same basic problem: Outcomes shouldn’t depend excessively on which person happens to be doing the work.

Yet when the activity moves from inspecting a part to grading an 8D, evaluating a corrective action, checking a PPAP, or deciding whether the evidence actually satisfies the requirement, we become much more tolerant of variation.

We call it judgment.

Then AI showed up.

General-purpose AI may quietly be making the problem worse

Most manufacturers didn’t begin their AI journey with purpose-built quality systems. AI arrived as general-purpose assistants. In a lot of companies, the rollout was effectively, “Here’s your access. Go use it.”

That sounds like a way to compensate for the knowledge and capacity problems we just described. But there’s a problem.

General-purpose AI doesn’t remove the human variable. It can amplify it.

A knowledgeable engineer gives the AI good context, recognizes the correct requirements, challenges bad assumptions, and catches weak output. A less experienced engineer brings a different set of assumptions, might not know which requirements matter, and might not recognize when something important is missing.

Both now have access to an extremely capable tool.

Your best engineer gets amplified. So does your weakest.

Research from Anthropic helps explain why this matters. In “Towards Understanding Sycophancy in Language Models,” researchers evaluated five leading AI assistants from four free-form tasks and found sycophantic behavior with all of them. Their work showed that assistants can favor responses aligned with a user’s stated beliefs over more truthful answers, and that human preference data can actually reward that behavior.

Think about what that looks like in a plant.

An engineer is overloaded and needs to close a corrective action. They already have an opinion about what happened and what should be required. They put that position into a general-purpose AI assistant. The AI returns something polished, detailed, confident, and professional looking.

The engineer skims it because there are 10 other things waiting.

The AI may have reinforced the assumption the engineer started with instead of challenging it. The result looks better. The reasoning might not be.

Variation hasn’t been reduced at all. It has been given better formatting, more confidence, and the appearance of authority.

We tested what happens when the user keeps pushing

We recently tested this behavior using an automotive quality scenario governed by IATF 16949, ISO 9001, AIAG Core Tools, and applicable customer-specific requirements.

A general-purpose AI system and a purpose-built quality workflow received the same scenario, and the same four increasingly aggressive follow-up questions.

The general-purpose system’s first answer was reasonably competent. Then the user pushed.

By the third turn, it conceded that the PFMEA and control plan could remain unchanged. By the fourth, it said a separate controlled record was unnecessary and described the proposed closure path as “auditor-proof.”

No new evidence had been introduced. Nothing about the underlying requirement changed.

The standard never moved. The AI did.

That’s what makes sycophancy particularly dangerous in quality work. The system doesn’t necessarily fail loudly. It can gradually negotiate its way toward the answer the user wanted while continuing to sound completely confident.

The first wave of enterprise AI reflects the same problem

This may help explain why enterprise AI deployment has struggled to create measurable value.

MIT’s Project NANDA reported in “The GenAI Divide: State of AI in Business 2025” that roughly 95% of enterprise gen AI initiatives in its dataset showed no measurable P&L effect.

That doesn’t mean 95% of AI technology doesn’t work. It means handing people powerful general-purpose tools hasn’t automatically translated into controlled, repeatable business processes.

A general-purpose assistant can write an email, summarize a meeting, explain statistics, generate software, build a presentation, and help plan a vacation. That’s extraordinary capability. It’s not the same thing as a qualified quality process.

A general-purpose assistant starts with the user. There may be no required inputs, no mandatory evidence, no fixed scoring rubric, and nothing forcing two engineers to apply the same method.

Two people can use the same AI and still get very different depth, different assumptions, and different conclusions.

The human variation is still there. AI can simply make it harder to see.

Purpose-built doesn’t mean automatically good

The answer isn’t to abandon AI. It’s to stop treating all AI as if it were the same thing. For repetitive analytical review work, we can take a different approach.

Build the method into the system.

If the job is grading an 8D, the workflow should know the grading criteria before the user arrives. It should know what evidence is required, what requirements apply, what constitutes an acceptable response, and what conditions should prevent a passing grade.

If the job is PPAP review, the same principle applies.

The system isn’t being asked to invent engineering truth from nothing. It’s evaluating submitted work against a known rubric, standard, or expert-adjudicated answer.

That distinction is important. Because now we have something we can test.

But purpose-built AI shouldn’t receive a free pass, either. Putting a workflow around an AI model doesn’t automatically make it accurate or repeatable.

If we claim it reduces variation, we should have to prove it.

Manufacturing already knows how to do this

We recently ran a repeatability and reproducibility study on one of our own 8D grading workflows. The study produced 2,370 individual criterion judgments. Repeatability was 93.2%. Reproducibility between two appraisers was 99%.

Those numbers mattered.

We didn’t turn model sampling down to force repeatability. The workflow ran at its normal configuration. The consistency came from the design of the method, not from suppressing model variation.

What the study found mattered even more.

Twelve criteria were unstable from run to run. When we investigated them, each traced back to ambiguity in our own written criteria.

We also discovered an input format that was silently failing to reach the model, even though the system continued producing a complete, professional-looking scorecard.

And yet, our customers were satisfied with the workflow. Normal use had exposed neither problem.

Measurement did.

That’s why I’ve started calling this application “Reasoning R&R.”

It’s not a new statistical method. Attribute agreement analysis already gives us much of the framework.

The questions are familiar.

Reasoning: Does the evaluation agree with the known standard?

Repeatability: Given the same input under the same controlled conditions, does the system reach the same substantive evaluation again?

Reproducibility: Do different appraisers applying the controlled method reach comparable conclusions?

The order matters.

A system that gives the same wrong answer every time is highly repeatable and completely useless.

AI shouldn’t get a lower standard

Manufacturing has a real problem to solve. Experienced knowledge is leaving. Core skills are under pressure. The people who remain are overloaded. Human judgment already varies, sometimes dramatically.

General-purpose AI can help with a tremendous number of things. But simply putting it in everyone’s hands doesn’t solve those problems. In some cases, it may quietly make them worse.

An overloaded engineer with an incorrect assumption no longer has only an incorrect assumption. They may now have an extremely articulate AI system reinforcing it while they skim the output and move to the next fire.

That’s not cognitive offloading. It’s variation amplification.

The answer isn’t better prompting.

For work that can be defined against known criteria, the answer is to encode the method, control the inputs, require the evidence, and then measure whether the resulting system actually performs against the standard.

That’s what manufacturing has always done when variation matters.

For years, we’ve asked whether AI is good. That’s no longer a good enough question.

If AI is going to help judge the quality of our work, we should be asking something much more familiar: Has the system been qualified?

Disclosure: My company builds AI workflows for quality functions. I have a commercial interest in this topic. The measurement methods discussed here, however, aren’t new. They come directly from measurement-system practices manufacturing has used for decades.

Top Stories
Beyond the Statistics: Implementing the New VDA AIAG SPC Manual
All Whys Are Not The Same
Flying Blind: What Your Quality Data Miss Every Shift
Book Preview
The Psychology of AI Adoption at Work
Momentum Technologies Achieves Record Purity Milestones for Rare Earths

Add new comment

The content of this field is kept private and will not be shown publicly.
About text formats

© 2026 Quality Digest. Copyright on content held by Quality Digest or by individual authors. Contact Quality Digest for reprint information.
“Quality Digest" is a trademark owned by Quality Circle Institute Inc.

footer
  • Home
  • Print QD: 1995-2008
  • Print QD: 2008-2009
  • Videos
  • Privacy Policy
  • Write for us
footer second menu
  • Subscribe to Quality Digest
  • About Us
  • Contact Us