# When two inspectors disagree: the measurement error hiding inside your tolerance

> A batch passes at the supplier and fails at the customer, and both sides have the readings to prove it. Before anyone argues about the parts, it is worth finding out how much of the tolerance the measurement itself is using up. On a surprising number of characteristics the answer is a third of it, or more.

Page: https://ganiindustrial.com/insights/when-two-inspectors-disagree · 2026-09-29 · Language: en

A customer once returned a batch of machined housings with a report showing a bore diameter out of tolerance on eleven parts out of fifty. We had measured the same fifty parts two days earlier and found all of them inside. Neither report was dishonest, and neither inspector was careless. Both had measured with a calibrated instrument, recorded what the instrument said, and signed. The parts were in the middle third of a 0.05 millimetre tolerance band, and the two measurement systems disagreed with each other by more than the band allowed.

The useful question in that situation is not whose number is right. It is how much of the tolerance each measurement system consumes before a part is even made. That is a question with a standard answer, and the study that produces it takes a morning.

## The tolerance is shared with the instrument

Think of a tolerance as a budget. A drawing that says a diameter must be 32 plus or minus 0.025 millimetres gives a total band of 0.05 mm. Everything that varies has to fit inside that band: the process, the fixture, the material, the temperature, and the measurement. If the measurement system by itself swings across a third of the band, only two thirds remain for actually making the part, and parts near the limits will be sorted essentially at random.

The usual rule, from the standard measurement system analysis practice that most automotive and aerospace customers require, puts a number on it. The spread of the measurement system, taken as six standard deviations, should be under 10 per cent of the tolerance to be comfortable; 10 to 30 per cent is acceptable with justification and depends on how critical the characteristic is; above 30 per cent the measurement is not fit for the decision it is being used to make. The same ratio can be taken against the actual process spread instead of the tolerance, and both are worth knowing: the first says whether you can judge a part, the second says whether you can see your own process moving.

## Repeatability, reproducibility, and the arithmetic between them

Two sources of variation are separated by the study. Repeatability is the same operator measuring the same part with the same instrument more than once, and it reflects the instrument and the way the part sits in it. Reproducibility is the difference between operators measuring the same parts, and it reflects technique and interpretation. The combined figure is the square root of the sum of their squares, which is worth remembering because it explains why fixing only the smaller one changes very little.

An example with real numbers. A micrometer study on a machined shoulder gives a repeatability standard deviation of 0.004 mm and a reproducibility standard deviation of 0.006 mm. The combined standard deviation is the square root of 0.004 squared plus 0.006 squared, which is 0.0072 mm. Six of those is 0.043 mm, against a tolerance band of 0.1 mm: 43 per cent. The instrument reads to a micrometre and the calibration certificate is current, and the measurement system is still not fit to sort these parts.

The same arithmetic shows where the effort belongs. Halving the repeatability alone, from 0.004 to 0.002, moves the combined figure only from 0.0072 to 0.0063, a four point improvement. Halving the reproducibility instead, from 0.006 to 0.003, moves it to 0.005, which is 30 per cent. The larger term dominates, and on hand measurements the larger term is nearly always the operator, which means the fix is usually a method and a fixture rather than a more expensive instrument.

## Resolution is not accuracy

A digital display promises nothing. Resolution only needs to be fine enough not to hide the variation: the working rule is that the smallest increment should be a tenth of the tolerance or better, so a 0.1 mm band wants a 0.01 mm resolution at worst. Beyond that, extra digits buy nothing. A caliper reading to 0.01 mm has a repeatability on a real workshop part closer to 0.02 or 0.03 mm, because the jaws can be rocked, the thumbwheel pressure varies and the part is held by hand. The display is not the measurement system; the person, the part, the fixture, the temperature and the instrument together are the measurement system.

## Where reproducibility actually comes from

Four causes cover most of it, and all four are cheap to remove once they are named.

The first is where on the part the measurement is taken. A bore that is slightly oval reads differently at 0 and at 90 degrees, and two inspectors who were never told where to measure will measure in different places, honestly, for years. A method sheet that says which feature, in which orientation, at which depth, usually halves reproducibility by itself.

The second is force. A micrometer has a ratchet because measuring force changes the reading, and that matters most on the parts people assume are rigid: a thin-walled tube, a plastic housing, a soft aluminium extrusion. Ten newtons applied by a confident hand against a thin wall is worth several micrometres of deflection.

The third is temperature, which is both the most ignored and the easiest to calculate. Steel expands about 11.7 micrometres per metre per degree. A 100 millimetre steel part measured at 28 degrees, straight off a machine, reads about 9 micrometres larger than the same part at the standard reference temperature of 20 degrees: nearly a tenth of a 0.1 mm tolerance, from heat alone. Aluminium moves twice as far. This is why a measurement taken at the machine and one taken in a measuring room an hour later are different measurements of the same part, and why a dispute about a part is often a dispute about two temperatures.

The fourth is how the part is held. A part laid on a surface plate by hand sits differently each time; the same part located in a simple fixture against three defined points sits the same way for everybody. The fixture is usually the cheapest improvement available, and on production characteristics we build one as a matter of course.

## Doing a study that is worth the morning

The classic study is ten parts, three operators, three trials. The parts are chosen to span the real range of production rather than being ten good ones, because the study also estimates process variation and ten identical parts make the measurement system look worse than it is. The parts are numbered where the operator cannot see the number while measuring, the order is randomised, and the operators measure in their normal way rather than in a careful way reserved for studies.

The output is three numbers worth reading. The percentage of tolerance consumed answers whether you can judge a part against a drawing. The percentage of total variation consumed answers whether you can see your process. And the number of distinct categories, which is roughly 1.41 times the process spread divided by the measurement spread, answers how many levels your gauge can actually distinguish; below five it cannot support a control chart, and a value of one or two means the gauge is telling you nothing except that the parts exist.

For go and no-go gauges and for anything judged by eye there are no standard deviations to compute, so the study changes shape: the same parts, including deliberately marginal ones, are judged several times by several people, and what is counted is agreement. Each inspector against themselves, each inspector against the others, and everybody against a reference decision made by metrology. Visual criteria for scratches, flash and finish routinely come out below 80 per cent agreement on the first study, which is usually the first honest evidence that a limit sample set is needed rather than a sentence in a specification.

## When a study fails, fix it in this order

The order matters because the cheap fixes are also the effective ones. First the fixture, so the part is located the same way every time. Then the method, written down with the location, the force and the temperature stated. Then training against that method, which is only worth doing once the method exists. Then the instrument, which is where everyone wants to start. Then, finally, a different measuring principle altogether, which is when a feature moves onto a coordinate measuring machine or an optical system.

There is a sixth option that engineers forget: ask whether the tolerance is real. A tolerance copied from a similar drawing, or tightened by habit, costs money in three places at once, since it narrows the process, demands a better measurement system and creates disputes about parts that function perfectly. When a measurement system consumes a third of a band, the question for the designer is whether the function needs that band, and the answer is often no.

## Calibration proves something else

A calibration certificate says that on one day, in a controlled room, an instrument agreed with a traceable standard within a stated uncertainty. It says nothing about what happens when your operator measures your part in your hall. Both are needed, and they answer different questions: calibration gives traceability and bias, the study gives capability in use. We have seen a plant with a perfect calibration system and no capability studies reject good material for a year, and we have never seen the reverse.

The uncertainty on the certificate also has a practical use at the limits. When a reading sits within the measurement uncertainty of a tolerance limit, the part has not been shown to be in or out of specification, and the decision rule has to be agreed in advance rather than invented during the argument. Agreeing with the customer at the start which side of the line gets the benefit of the doubt, and whether the limit is guard-banded by the uncertainty, costs one sentence in a quality agreement and saves a returned batch.

## What we hand over

For every characteristic we are asked to measure and report on production, our control plan names the instrument, its resolution, the fixture, the method, the date of the last capability study and the percentage it consumed. Test fixtures we build for electrical and functional measurements are validated the same way, with repeated measurements on known good and known bad units rather than on a single golden sample. Readings are stored against the serial number, so a disagreement years later is settled by data.

None of this is a quality department ritual. It is the difference between an inspection result that means something and a number that happens to be written on a report, and when two plants disagree about a batch it is the only thing that tells you which one of them is measuring.


---

GANI Industrial — engineering, manufacturing and lifecycle for European OEMs. Bursa, Türkiye.
Site index: https://ganiindustrial.com/llms.txt
