Measurement System Analysis – Attribute Agreement Analysis (AAA), Kappa and Gage R&R
Learning Objectives
By the end of this lesson, the learner will be able to:
- Explain the purpose and importance of Measurement System Analysis (MSA).
- Explain why measurement reliability must be established before analyzing process performance.
- Distinguish between variable and attribute measurement systems.
- Explain repeatability and reproducibility.
- Explain the purpose and methodology of Attribute Agreement Analysis (AAA).
- Distinguish within-appraiser, between-appraiser, and appraiser-to-standard agreement.
- Explain the purpose and interpretation of Kappa Analysis.
- Distinguish Kappa Analysis from Attribute Agreement Analysis.
- Understand when Kappa is appropriate for categorical data.
- Explain the purpose and methodology of Gage R&R.
- Distinguish repeatability from reproducibility in a Gage R&R study.
- Identify common sources of measurement-system variation.
- Interpret important MSA results and graphical outputs.
- Understand the appropriate use of Minitab for AAA, Kappa, and Gage R&R.
- Recognize common mistakes when evaluating measurement systems.
- Apply MSA findings appropriately during the Measure and Analyze phases of DMAIC.
1. Introduction
Measurement is one of the foundations of Six Sigma.
A Green Belt may collect hundreds or thousands of observations and perform sophisticated statistical analysis. However, if the measurement system itself is unreliable, the resulting conclusions may also be unreliable.
For example, suppose a manufacturing team reports:
Component diameter = 25.04 mm
Before using this measurement to study process variation, capability, or improvement, the team should ask:
Can the measurement system consistently and reliably measure the component diameter?
Similarly, consider an inspection system in which inspectors classify products as:
- Accept
- Reject
If inspectors frequently disagree when evaluating the same product, the classification system may not be sufficiently reliable.
Measurement System Analysis (MSA) provides a structured way to evaluate measurement systems.
In this lesson, three important analytical approaches are covered separately:
1. Attribute Agreement Analysis (AAA)
AAA evaluates the agreement and consistency of an attribute measurement system, including agreement:
- Within an appraiser
- Between appraisers
- Between appraisers and a known standard, where applicable
2. Kappa Analysis
Kappa is a statistical measure of categorical agreement beyond chance agreement.
Kappa analysis focuses specifically on the degree of agreement after accounting for the agreement that could be expected by chance.
3. Gage R&R
Gage R&R evaluates important sources of variation in a variable measurement system, particularly:
- Repeatability
- Reproducibility
These three topics are related because they all concern measurement-system reliability, but they are not the same analysis.
A key Six Sigma principle is:
Before analyzing process variation, establish that the measurement system is appropriate for the intended purpose.
2. Concept / Theory
2.1 What Is Measurement System Analysis?
Measurement System Analysis is the systematic evaluation of a measurement process to determine whether it is suitable for its intended use.
A measurement system includes more than the instrument.
It may include:
- Measuring instrument
- Operator or appraiser
- Measurement method
- Operating procedure
- Part or sample
- Measurement location
- Environment
- Calibration
- Data recording method
- Software
- Inspection standards
Therefore:
The measurement system is the complete process used to obtain and record a measurement.
2.2 Why MSA Is Important
Suppose a process is actually stable, but the measurement system produces inconsistent readings.
The resulting data may appear to show:
- Excessive variation
- Process instability
- Poor capability
- Increased defects
- Differences between operators
- False improvement or deterioration
Therefore, measurement-system variation must be understood before important process conclusions are made.
2.3 Two Broad Types of Measurement
Measurement systems can broadly be classified as:
Variable Measurement
The result is numerical.
Examples:
- Diameter = 25.03 mm
- Weight = 4.82 kg
- Temperature = 72.4°C
- Cycle time = 8.6 minutes
A variable measurement system is commonly evaluated using Gage R&R.
Attribute Measurement
The result is categorical.
Examples:
- Pass / Fail
- Accept / Reject
- Good / Defective
- Correct / Incorrect
- Type A / Type B / Type C
Attribute systems can be evaluated using Attribute Agreement Analysis and appropriate categorical agreement statistics such as Kappa.
2.4 Attribute Agreement Analysis (AAA)
Attribute Agreement Analysis is used to evaluate the consistency of an attribute measurement or inspection system.
An attribute measurement system does not produce a continuous numerical measurement. Instead, an appraiser assigns a category or classification.
For example:
Product inspection
- Accept
- Reject
or:
Defect classification
- Scratch
- Dent
- Crack
- No defect
AAA evaluates whether the appraisers classify the same items consistently.
2.5 What Does AAA Evaluate?
Attribute Agreement Analysis can evaluate several types of agreement.
Within-Appraiser Agreement
This evaluates whether the same appraiser gives the same classification when assessing the same item repeatedly.
Example:
Inspector A evaluates a component twice.
First assessment:
Reject
Second assessment:
Reject
This is evidence of within-appraiser agreement.
Between-Appraiser Agreement
This evaluates whether different appraisers give the same classification to the same item.
Example:
Inspector A:
Reject
Inspector B:
Reject
This contributes to between-appraiser agreement.
Appraiser-to-Standard Agreement
Where a known or established standard classification exists, AAA can evaluate whether an appraiser agrees with that standard.
Example:
Reference standard:
Reject
Inspector:
Reject
This represents agreement with the standard.
2.6 Why AAA Is Important
AAA provides a broader view of attribute measurement-system performance.
The Green Belt can determine whether problems are associated with:
- Individual appraiser consistency
- Differences between appraisers
- Difficulty of particular items
- Agreement with a known standard
This can help identify where the measurement system needs improvement.
2.7 Kappa Analysis
Kappa Analysis is a separate statistical concept from Attribute Agreement Analysis.
Kappa measures the degree of categorical agreement beyond what would be expected by chance.
The basic concept is:
Observed Agreement
compared with:
Expected Agreement by Chance
A commonly used formula for Cohen’s Kappa is:
κ = (Po − Pe) / (1 − Pe)
Where:
- Po = observed agreement
- Pe = expected agreement by chance
Kappa is therefore different from simply calculating the percentage of observations on which two classifications agree.
2.8 Why Percentage Agreement Alone May Be Misleading
Suppose two inspectors agree on 95% of classifications.
At first glance, this appears to be excellent agreement.
However, if almost every item belongs to one category, a high level of agreement may occur simply because that category is very common.
Kappa takes expected chance agreement into consideration.
Therefore:
Percentage agreement and Kappa are different measures and should not be treated as interchangeable.
2.9 Kappa Values
Kappa generally ranges from:
−1 to +1
A value approaching:
+1
indicates strong agreement beyond chance.
A value near:
0
indicates agreement approximately at the level expected by chance, under the assumptions of the statistic.
A negative value indicates agreement below the level expected by chance.
However, Kappa should not be interpreted using one universal set of numerical cutoffs without considering:
- Category prevalence
- Study design
- Number of categories
- Consequences of misclassification
- Intended use of the measurement system
2.10 Kappa Is Not the Same as AAA
This distinction is important for the Green Belt.
| Attribute Agreement Analysis | Kappa Analysis |
|---|---|
| Broader analysis of attribute measurement agreement | Specific statistical measure of categorical agreement |
| Can evaluate within-appraiser agreement | Quantifies agreement beyond chance in applicable categorical comparisons |
| Can evaluate between-appraiser agreement | Can quantify agreement between categorical ratings |
| Can evaluate agreement with a standard | Measures agreement after accounting for chance agreement |
| Provides a broader view of attribute-system performance | Focuses on a statistical agreement coefficient |
| May report several agreement measures | Produces Kappa statistics for applicable comparisons |
AAA and Kappa can therefore be used together, but they should not be described as identical.
2.11 Gage R&R
Gage R&R is primarily used for variable measurement systems.
R&R means:
Repeatability and Reproducibility
Gage R&R evaluates how much measurement variation comes from the measurement system compared with variation between parts.
2.12 Repeatability
Repeatability is the variation observed when:
- The same operator
- Measures the same part
- Using the same measurement system
- Under similar conditions
- Repeatedly
It is often associated with equipment or measurement-method variation.
Example:
An operator measures the same component five times.
Results:
25.02
25.03
25.02
25.08
25.03
If the readings vary more than expected, repeatability may be a concern.
2.13 Reproducibility
Reproducibility concerns variation between operators or appraisers.
Example:
Three operators measure the same components using the same measurement system.
If Operator A consistently obtains higher measurements than Operators B and C, the measurement system may have a reproducibility issue.
Potential causes include:
- Different measurement techniques
- Training differences
- Fixture use
- Reading interpretation
- Measurement location
- Operator interaction with the instrument
2.14 Part-to-Part Variation
Part-to-part variation represents the actual differences between the parts being measured.
For example:
Part A = 24.95 mm
Part B = 25.02 mm
Part C = 25.11 mm
Some variation exists because the physical parts themselves are different.
A useful Gage R&R study must allow the analysis to distinguish:
Measurement variation
from
Part-to-part variation.
2.15 Gage R&R Components
A typical Gage R&R analysis considers:
- Repeatability
- Reproducibility
- Gage R&R
- Part-to-part variation
- Total variation
The objective is to understand how much observed variation is attributable to the measurement system.
3. Key Topics
3.1 Attribute Agreement Analysis — Key Topics
AAA commonly examines:
- Within-appraiser agreement
- Between-appraiser agreement
- Appraiser-to-standard agreement
- Overall agreement
- Classification consistency
- Difficult or ambiguous items
The analysis helps identify whether the inspection process is consistent.
3.2 Kappa Analysis — Key Topics
Kappa analysis focuses on:
- Categorical data
- Agreement between classifications
- Observed agreement
- Expected chance agreement
- Agreement beyond chance
- Kappa coefficient
The Green Belt must understand that Kappa is not simply a percentage agreement calculation.
3.3 Gage R&R — Key Topics
Gage R&R focuses on:
- Repeatability
- Reproducibility
- Part-to-part variation
- Total measurement variation
- Percentage contribution or variation measures
- Number of distinct categories
3.4 Number of Distinct Categories
Gage R&R studies may report the:
Number of Distinct Categories (ndc)
This provides an indication of the measurement system’s ability to distinguish different levels of part-to-part variation.
If the measurement system cannot reliably distinguish meaningful differences between parts, it may not be suitable for the intended application.
3.5 Measurement-System Resolution
The measurement system must have sufficient resolution for the characteristic being studied.
For example, an instrument with very coarse resolution may not be suitable when the process differences of interest are very small.
The Green Belt should therefore consider:
- Instrument resolution
- Specification limits
- Process variation
- Measurement purpose
3.6 Bias
Bias is the systematic difference between the average observed measurement and an accepted reference value.
For example:
Reference:
50.00 mm
Average observed measurement:
50.12 mm
This may indicate positive bias.
Bias is different from:
- Repeatability
- Reproducibility
3.7 Stability
Stability refers to the consistency of the measurement system’s performance over time.
A measurement system may initially perform well but change over time because of:
- Instrument wear
- Environmental changes
- Calibration issues
- Method changes
Stability should therefore be considered where appropriate.
3.8 Linearity
Linearity concerns whether measurement bias remains reasonably consistent across the operating range of the measurement system.
A measurement system may perform differently at the low, middle, and high portions of its measurement range.
Linearity is therefore another important MSA concept for variable measurement systems.
3.9 Choosing the Appropriate MSA
| Data Type | Primary MSA Approach |
|---|---|
| Continuous / variable measurement | Gage R&R |
| Attribute / categorical inspection | Attribute Agreement Analysis |
| Categorical agreement beyond chance | Kappa Analysis |
| Systematic difference from reference | Bias |
| Measurement performance across range | Linearity |
| Measurement performance over time | Stability |
These analyses answer different questions and should not be unnecessarily substituted for one another.
4. Tools / Methodology
4.1 Attribute Agreement Analysis — Methodology
Step 1 — Define the Attribute
Clearly define the classification.
Example:
Defective = Any component containing a crack that meets the documented rejection criterion.
Avoid vague definitions such as:
Defective = Poor quality
Step 2 — Establish the Standard
Where appropriate, establish a reference classification using a reliable method.
The standard provides a basis for evaluating appraiser-to-standard agreement.
Step 3 — Select Representative Items
Select items representing the actual inspection situation.
Include appropriate examples of:
- Clearly acceptable items
- Clearly defective items
- Borderline items, when they are relevant to the actual inspection process
Step 4 — Select Appraisers
Use appraisers who normally perform the inspection.
Step 5 — Conduct Repeated Assessments
Each appraiser assesses the selected items according to the normal inspection procedure.
Repeated assessments allow within-appraiser agreement to be evaluated.
Step 6 — Analyze Agreement
Review:
- Within-appraiser agreement
- Between-appraiser agreement
- Appraiser-to-standard agreement
- Overall agreement
Step 7 — Investigate Disagreement
Where disagreement occurs, investigate:
- Inspection criteria
- Training
- Lighting
- Visual standards
- Part presentation
- Defect definitions
- Inspection procedure
4.2 Kappa Analysis — Methodology
Step 1 — Identify the Categorical Ratings
Determine the categories being compared.
Example:
Accept / Reject
or:
Defect A / Defect B / Defect C / No Defect
Step 2 — Identify the Comparison
Determine what is being compared.
Examples:
- Appraiser versus appraiser
- Rating versus reference classification
- Classification from one assessment versus another assessment
Step 3 — Determine Observed Agreement
Calculate how often the classifications agree.
Step 4 — Determine Expected Agreement
Estimate the agreement expected by chance under the Kappa framework.
Step 5 — Calculate Kappa
Use the appropriate Kappa statistic for the study design.
For two-category nominal classifications, Cohen’s Kappa is commonly used.
For more complex categorical situations, other forms of Kappa may be appropriate.
Step 6 — Interpret Kappa in Context
Consider:
- Kappa value
- Category distribution
- Sample size
- Study design
- Agreement percentage
- Consequences of classification errors
- Practical requirements of the inspection system
Do not interpret Kappa in isolation.
4.3 Gage R&R — Methodology
Step 1 — Define the Measurement
Specify:
- Characteristic
- Unit
- Instrument
- Measurement location
- Measurement method
- Acceptance requirements
Step 2 — Select Representative Parts
Select parts that represent an appropriate range of actual process variation.
Avoid selecting only nearly identical parts.
Step 3 — Select Operators
Use operators who normally perform the measurement.
Step 4 — Determine the Study Design
A typical study may use:
- Multiple parts
- Multiple operators
- Multiple trials
For example:
10 parts × 3 operators × 2 trials = 60 measurements
The exact design should be appropriate for the measurement system and intended analysis.
Step 5 — Randomize Measurements
Measurement order should be randomized appropriately to reduce potential order and recognition effects.
Step 6 — Conduct the Study
Each operator measures each selected part according to the defined procedure.
Step 7 — Analyze Gage R&R
Review:
- Repeatability
- Reproducibility
- Gage R&R
- Part-to-part variation
- Total variation
- Number of distinct categories
- Relevant graphical output
Step 8 — Investigate Problems
If measurement variation is excessive, investigate:
- Instrument
- Operator technique
- Method
- Fixture
- Environment
- Calibration
- Resolution
- Training
4.4 Minitab Application
Gage R&R
Minitab generally provides Gage R&R under:
Stat → Quality Tools → Gage Study → Gage R&R Study
The exact menu wording may vary by Minitab version.
Attribute Agreement Analysis
Minitab provides an Attribute Agreement Analysis procedure for evaluating attribute measurement systems.
The analysis can provide information about:
- Within-appraiser agreement
- Between-appraiser agreement
- Agreement with a standard
- Overall agreement
Kappa Analysis
Kappa statistics may be obtained through appropriate categorical agreement analyses depending on the study design and Minitab version.
The Green Belt should distinguish the Kappa statistic from the broader Attribute Agreement Analysis.
The exact Minitab menu and available options may vary by version.
5. Worked Example / Case Study
Case Study A — Attribute Agreement Analysis
A company inspects finished products using three inspectors.
Each product is classified as:
Accept or Reject.
The company selects 20 representative products.
Each inspector evaluates every product twice using the normal inspection procedure.
Step 1 — Define the Attribute
The company establishes:
Reject = Product fails one or more documented acceptance criteria.
Step 2 — Establish the Standard
A qualified review process establishes the reference classification for each study item.
Step 3 — Conduct the Study
Each inspector evaluates the 20 products twice.
The results are analyzed for:
- Within-inspector agreement
- Between-inspector agreement
- Inspector-to-standard agreement
Step 4 — Interpret the Findings
Suppose Inspector A agrees with their own previous classifications most of the time, but Inspectors B and C show considerable disagreement on borderline products.
This suggests that the inspection system may require investigation.
Possible causes include:
- Ambiguous inspection criteria
- Insufficient training
- Poor visual standards
- Difficulty distinguishing borderline conditions
Step 5 — Improvement
The company may:
- Clarify acceptance criteria.
- Create reference samples.
- Improve inspection lighting.
- Retrain inspectors.
- Standardize inspection technique.
The study should then be repeated to verify improvement.
Case Study B — Kappa Analysis
Two inspectors classify 100 products as:
Accept
or
Reject.
Suppose the observed agreement is high.
However, the Accept category is very common.
The Green Belt should not rely solely on percentage agreement.
Kappa analysis is performed to evaluate agreement beyond expected chance agreement.
Suppose the resulting Kappa is:
κ = 0.78
The result indicates substantial agreement beyond chance under the conditions of the study.
The Green Belt should still consider:
- The category distribution
- The consequences of incorrect classification
- The study design
- The intended use of the inspection system
Important Distinction
The analysis above demonstrates why the two concepts must remain separate:
AAA asks broadly:
How consistently are the appraisers classifying the items?
Kappa asks statistically:
How much categorical agreement exists beyond what would be expected by chance?
Both can contribute to understanding an attribute measurement system.
Case Study C — Gage R&R
A manufacturing company measures the diameter of a critical component.
Specification:
50.00 ± 0.20 mm
The study uses:
- 10 parts
- 3 operators
- 2 trials
Therefore:
10 × 3 × 2 = 60 measurements
Step 1 — Define the Measurement
Characteristic: Component diameter
Unit: mm
Instrument: Micrometer
Measurement location: Standardized location
Step 2 — Conduct the Study
Each operator measures all ten parts twice.
Measurements are randomized appropriately.
Step 3 — Analyze
The Gage R&R analysis separates measurement-system variation from part-to-part variation.
Suppose the results indicate:
- Repeatability is relatively small.
- Reproducibility is relatively small.
- Part-to-part variation is substantially larger.
- The measurement system has useful discrimination for the intended application.
The Green Belt can then evaluate whether the measurement system is suitable for the intended process decisions.
Step 4 — If Repeatability Is Excessive
Investigate:
- Instrument condition
- Resolution
- Measurement technique
- Contact force
- Fixture
- Measurement location
- Environmental conditions
Step 5 — If Reproducibility Is Excessive
Investigate:
- Operator training
- Standard work
- Measurement technique
- Instrument handling
- Reading interpretation
- Fixture use
- Measurement location
Important Learning Point
A measurement system can have different problems.
For example:
High repeatability variation
does not necessarily indicate the same problem as:
High reproducibility variation.
The Green Belt must identify the dominant source before selecting corrective action.
6. Practical Application
Exercise 1 — Select the Correct MSA
For each situation, identify the most appropriate primary analysis.
Situation A
A micrometer is used to measure component diameter.
Expected approach: Gage R&R
Situation B
Inspectors classify products as Accept or Reject.
Expected approach: Attribute Agreement Analysis
Situation C
Two categorical classifications are being compared and agreement beyond chance is required.
Expected approach: Kappa Analysis
The objective is to understand why each analysis is appropriate rather than simply memorizing the names.
Exercise 2 — Develop an AAA Study
Design an Attribute Agreement Analysis for a visual inspection process.
Define:
- Attribute
- Categories
- Standard
- Number of items
- Appraisers
- Repeated assessments
- Inspection conditions
- Agreement measures
- Method for investigating disagreement
Exercise 3 — Kappa Interpretation
Suppose two inspectors classify the same set of products.
The observed agreement is high, but one category is extremely common.
Ask the learner:
- Why might percentage agreement alone be insufficient?
- What additional statistic can account for expected chance agreement?
- What does a positive Kappa indicate?
- Why should Kappa still be interpreted in context?
Exercise 4 — Gage R&R Study Design
Design a Gage R&R study for a critical dimensional characteristic.
Specify:
- Number of parts
- Number of operators
- Number of trials
- Measurement instrument
- Measurement procedure
- Randomization
- Data collection format
- Analysis method
- Interpretation approach
Exercise 5 — Diagnose the Measurement Problem
Consider two scenarios.
Scenario 1
One operator obtains substantially different measurements when repeatedly measuring the same part.
Identify the likely primary concern:
Repeatability
Scenario 2
Different operators consistently obtain different measurements from the same parts.
Identify the likely primary concern:
Reproducibility
Then identify at least three potential causes for each situation.
Practical Checklist
Before relying on measurement data, ask:
General MSA
- Is the measurement clearly defined?
- Is the method standardized?
- Is the instrument appropriate?
- Is calibration appropriate and current where required?
- Is the measurement resolution sufficient?
- Is the measurement system suitable for its intended purpose?
Attribute Measurement
- Is the attribute clearly defined?
- Are the classification categories clear?
- Is there a reliable reference standard where appropriate?
- Have within-appraiser agreement and between-appraiser agreement been evaluated?
- Has agreement with the standard been evaluated where applicable?
- Has Attribute Agreement Analysis been performed?
Kappa
- Is the data categorical?
- What classifications are being compared?
- Has observed agreement been considered?
- Has chance agreement been considered?
- Is the appropriate Kappa statistic being used?
- Are category prevalence and study design being considered?
Variable Measurement
- Has Gage R&R been performed?
- Has repeatability been evaluated?
- Has reproducibility been evaluated?
- Has part-to-part variation been evaluated?
- Is the number of distinct categories adequate for the intended use?
- Have measurement-system problems been investigated?
7. Lesson Summary
Measurement System Analysis provides the foundation for trustworthy Six Sigma analysis.
This lesson distinguishes three important approaches.
Attribute Agreement Analysis (AAA)
AAA evaluates the performance of an attribute measurement system.
It can examine:
- Within-appraiser agreement
- Between-appraiser agreement
- Appraiser-to-standard agreement
- Overall classification consistency
Kappa Analysis
Kappa is a statistical measure of categorical agreement beyond chance agreement.
It should not be treated as simply another name for Attribute Agreement Analysis.
AAA provides a broader assessment of attribute measurement-system agreement, while Kappa provides a specific statistical measure of categorical agreement beyond chance.
Gage R&R
Gage R&R is primarily used for variable measurement systems.
It evaluates:
- Repeatability
- Reproducibility
- Gage R&R
- Part-to-part variation
- Measurement-system discrimination
The Green Belt should understand the fundamental distinction:
AAA evaluates attribute agreement.
Kappa quantifies categorical agreement beyond chance.
Gage R&R evaluates repeatability and reproducibility of variable measurements.
These analyses support different measurement-system questions and should be selected according to the type of data and the purpose of the measurement.
8. Lesson Learnt / Conclusion
The most important lesson is:
Before trusting process data, understand the measurement system that produced the data.
A Six Sigma team should not automatically assume that differences in measurements represent real differences in the process.
For attribute measurements, the Green Belt should understand both:
Attribute Agreement Analysis
and
Kappa Analysis
because they provide related but distinct information.
AAA helps evaluate the consistency of classifications across appraisers, repeated assessments, and standards where applicable.
Kappa provides a statistical assessment of categorical agreement beyond expected chance agreement.
For variable measurements, Gage R&R helps separate measurement-system variation into important components such as repeatability and reproducibility and compares them with part-to-part variation.
The appropriate MSA approach is therefore:
Identify the type of measurement → Define the measurement system → Select the appropriate MSA → Conduct the study → Analyze the results → Investigate sources of variation → Improve the measurement system → Reassess.
Once the measurement system is demonstrated to be suitable for its intended purpose, the Green Belt can proceed with greater confidence to:
- Process capability
- Hypothesis testing
- Control charts
- Root-cause analysis
- Regression analysis
- Improvement decisions
Reliable measurements create reliable data.
Reliable data create reliable Six Sigma decisions.