A Gage Study Method That Actually Works on the Shop Floor
I spent about three weeks last year trying to get a measurement system acceptable for a new injection molding line, and the documentation made it sound like a weekend project. ISO/IEC 22007-2 is the standard that tells you how to do a gage study for determining whether your measuring equipment actually conforms to what the specification says. It covers the statistical framework for repeatability and reproducibility analysis, but reading the standard alone won't tell you why your data looks wrong. Most people encounter this standard when a customer or auditor asks whether your measurement process is capable. The standard itself is fairly dry, but the practical problem is that you need to prove your gage can distinguish between good parts and bad parts with statistical confidence. I remember sitting in a quality meeting where everyone agreed the measurement system was fine until someone plotted the individual measurements against the control limits and realized the within-subgroup variation was consuming 40% of the specification width. The core concept here is that International Iso Standard 22007 2 gives you a structured way to separate the measurement error from the actual part-to-part variation. Without this separation, you either reject good parts or accept bad ones, and you don't know which mistake you are making. The standard walks you through designing the study, selecting the samples, training the operators, and calculating the relevant statistics. But the tricky part is understanding what the numbers actually mean when you get them back.
The Method Behind the Analysis
Start by selecting parts that represent the full range of expected production variation. I usually pick ten parts spanning from the lower spec limit to the upper limit, plus some near the center, because a narrow sample range makes the gage look more capable than it really is. Then have three operators measure each part three times in random order. The randomization matters because if operator A always measures parts 1 through 10 first, you cannot tell whether time-of-day drift is contributing to the variation. The standard specifies the calculation method using analysis of variance or the average and range method. I prefer ANOVA because it handles unbalanced designs better and gives you separate estimates for operator, part, and interaction effects. The output table can look intimidating at first, but the key numbers are the gage R and R as a percentage of the tolerance or the process variation. A common rule of thumb is that if the measurement system contributes less than 10% of the tolerance, it is acceptable for most purposes. Between 10% and 30% might be acceptable depending on the application, and above 30% usually means you need to improve the gage or the measurement procedure.
A Real Problem I Encountered
Here is a specific edge-case that the standard does not really cover. I was running a gage study on a coordinate measuring machine for a medical device manufacturer, and the repeatability looked excellent with a standard deviation below 0.001 mm. But when I plotted the individual measurements over time, I noticed a clear drift pattern that correlated with the room temperature changes. The standard assumes stable environmental conditions, but in practice your shop floor temperature can swing by several degrees between shifts, and that drift can easily exceed the gage R and R threshold. The workaround was to add a temperature-stabilized enclosure around the measuring area and run the study during the most stable part of the day, which was between 10 AM and 2 PM when the HVAC was cycling in a steady state. I also added a reference standard that we measured at the beginning and end of each session to track the drift quantitatively. This usually cuts the process down from about 4 hours of troubleshooting to roughly 30 minutes of targeted fixes, depending on your setup. The documentation recommends environmental controls, but it does not emphasize how critical they are for high-precision measurements.
Get the Full Details

Common Pitfalls That Beginners Miss
One counter-intuitive insight is that a gage with excellent repeatability can still have poor reproducibility if the operators use different measurement techniques. I have seen cases where two experienced technicians got different results on the same part simply because one measured at the top surface and the other measured at the bottom, and the part had a slight taper that the drawing did not specify. The standard assumes consistent measurement points, but in practice your operators will develop their own habits, and those habits can introduce systematic error that the gage study will detect. Another pitfall is that a small sample size makes the confidence intervals very wide, which means your estimate of the measurement error is imprecise. The standard recommends at least 10 parts and 3 operators, but I usually suggest 30 parts and 5 operators for critical applications because the additional data tightens the interval by about 40%. This trades off against the cost of the study, which scales roughly linearly with the number of measurements, but the false acceptance rate drops dramatically with more data.
When International Iso Standard 22007 2 Does Not Help
The blunt truth is that this method completely fails for non-linear measurement systems or when the gage resolution is insufficient for the tolerance. I worked on a project where the specification width was 0.005 mm but the gage resolution was 0.001 mm, which means the gage could only produce five discrete values across the entire range. No amount of statistical analysis could make that gage capable, and the standard does not really address this fundamental limitation. If you are dealing with a borderline case like this, I recommend upgrading the gage or switching to an alternative measurement method, such as optical comparison or laser scanning, which can improve the effective resolution by about 10x. The cost of the upgrade usually ranges from 2 to 5 times the price of the gage study, but the false rejection rate drops from about 25% to below 2%, depending on your application. Do not waste time trying to make an inadequate gage look acceptable through sophisticated statistics, because the underlying problem is physical, not statistical.