In a 2019 study, three trained automotive inspectors evaluated the same 30 daytime running lights twice against a known standard, agreeing with each other only 36.67% of the time. This post explains why that's not a training failure but a measurement problem, and why manufacturing runs attribute agreement analysis on physical inspection but never on judgment work like grading an 8D. If unaided human judgment tops out near 37% agreement, that's the bar any AI grading quality w