Why Three Inspectors Can Disagree 63% of the Time
In 2019, three trained automotive inspectors evaluated the same 30 daytime running lights, twice, against the same known standard. They agreed with each other on 36.67% of the parts. Individually, their effectiveness against the standard ranged from 56.67% to 80%.
This Isn't a Training Failure
It's a measurement problem, and manufacturing has a whole discipline for that: attribute agreement analysis. When results depend heavily on who's doing the inspecting, you don't blame the inspector — you fix the standard, retrain against clearer defect samples, and measure again. In that study, the company did exactly that and got to 95% agreement.
We Never Ran the Same Test on Judgment
Physical inspection gets this rigor. Grading an 8D, evaluating a corrective action, deciding whether a PPAP submission satisfies the requirement — we call the disagreement there “experience” and move on. Nobody runs an attribute agreement study on whether two quality engineers would grade the same submission the same way.
What This Means for AI
If unaided human judgment tops out around 37% agreement, that's the real baseline any AI system grading quality work has to beat — not “sounds confident,” not “produces a professional-looking scorecard.” We built Reasoning R&R™ to run exactly this kind of study on our own 8D grading workflow, and the results — 93.2% repeatability, 99% reproducibility — are the subject of the next post.
Want to see how it holds up on your own documents? Book a demo: mad-ai.com/book-a-demo





Comments