Skip to main content Skip to course navigation

Data Collection and Sampling

Learning Objective

By the end of this lesson, the learner will be able to:

  1. Explain the importance of systematic data collection in Six Sigma.
  2. Distinguish between a population and a sample.
  3. Explain the importance of representative data.
  4. Understand how sampling can reduce the effort required to study a process.
  5. Recognize common sources of sampling bias.
  6. Explain random, stratified, and systematic sampling.
  7. Use check sheets to support structured data collection.
  8. Understand how appropriate sampling supports reliable Six Sigma analysis.

1. Introduction

Data is the foundation of Six Sigma decision-making.

A Black Belt must collect data in a systematic and appropriate manner so that the information accurately represents the process being studied.

When it is not practical to collect information from every item or person in a population, sampling can be used.

The objective is to obtain useful information while reducing unnecessary data-collection effort.


2. Understanding the Population

A population is the complete group of items, transactions, people, or observations relevant to the study.

Examples include:

  • All products manufactured during a month
  • All customer transactions during a year
  • All patients visiting a hospital
  • All service calls received by a call centre

The population defines the complete group about which the Black Belt wants to understand process performance.


3. Understanding a Sample

A sample is a subset of observations selected from the population.

For example:

Population: 100,000 customer transactions

Sample: 1,000 selected transactions

Studying the sample can provide information about the population without examining every individual observation.

However, the sample should be selected appropriately so that it represents the population being studied.


4. Representative Data

Representative data reflects the important characteristics of the population or process being studied.

A Black Belt should consider whether the collected sample adequately represents:

  • Different time periods
  • Different shifts
  • Different locations
  • Different customer groups
  • Different operating conditions

If important groups are excluded, the analysis may provide a misleading view of process performance.


5. Sampling Bias

Sampling bias occurs when the method used to select observations systematically favours some observations or groups over others.

For example, suppose a retail organization wants to understand customer waiting time but collects data only during quiet periods.

The resulting sample may not represent the actual customer experience during busy periods.

Therefore, the Black Belt should consider when, where, and how the data is collected.


6. Random Sampling

Random sampling selects observations so that members of the population have an appropriate opportunity to be selected.

For example, a company may randomly select customer transactions from a complete transaction list.

Random sampling can help reduce selection bias when properly designed and implemented.


7. Stratified Sampling

Stratified sampling divides the population into relevant groups, or strata, and then samples from those groups.

Examples of strata may include:

  • Geographic regions
  • Customer categories
  • Product types
  • Production shifts

Application Example

A company operates in four regions:

  • North
  • South
  • East
  • West

Instead of collecting data only from one region, the Black Belt can structure the sample to include the relevant regions.

This can provide a more representative view of the process.


8. Systematic Sampling

Systematic sampling selects observations according to a defined interval from an ordered population or process.

For example, an organization may inspect every 10th transaction from an appropriate transaction sequence.

The sampling interval should be selected appropriately for the process and study design.


9. Check Sheets for Data Collection

A check sheet provides a structured method for recording observations.

It can help the Black Belt consistently capture information such as:

  • Defect types
  • Defect frequency
  • Process occurrences
  • Time-related events
  • Categories of problems

The important principle is to define what will be collected, how it will be recorded, and under what conditions before beginning data collection.


10. Tools & Techniques

The main tools and techniques associated with this lesson include:

  • Population definition
  • Sample selection
  • Random sampling
  • Stratified sampling
  • Systematic sampling
  • Check sheets
  • Data-collection planning
  • Bias identification

These techniques help the Black Belt obtain data that is useful for subsequent analysis.


11. Application Example — Customer Survey

A company wants to understand customer satisfaction across several geographic regions.

The customer population includes customers from:

North, South, East and West.

The Black Belt uses stratified sampling so that customers from the different regions are represented in the study.

The resulting data provides a broader view of customer experience than collecting responses from only one region.


12. Case Study — Retail Chain

A retail chain wants to understand customer waiting times.

Initially, the organization collects observations mainly from urban stores.

The Black Belt recognizes that the sample does not adequately represent the organization’s wider operations.

The study is expanded to include both urban and rural locations.

The analysis identifies longer waiting times in rural locations.

The organization responds by adjusting staffing arrangements.

The case illustrates the importance of selecting data that represents the process being studied.


13. From Data Collection to Analysis

A Black Belt should follow a logical sequence:

Define the Population

Determine the Required Data

Select an Appropriate Sampling Method

Collect the Data Systematically

Check for Bias and Representativeness

Analyse the Data

Poor sampling can affect the reliability of subsequent analysis, even when the statistical method itself is correctly applied.


14. Black Belt Perspective

A Black Belt should not simply ask:

“How much data do we have?”

The more important questions are:

“Where did the data come from?”

“How was it selected?”

“Does it represent the process?”

“Could there be sampling bias?”

Good statistical analysis begins with good data collection.


15. Lesson Practice

Scenario

A retail company wants to understand customer waiting time across its stores.

The company operates in:

  • Urban locations
  • Rural locations
  • Different regions
  • Different operating shifts

Questions

  1. What is the population in this study?
  2. What would constitute a sample?
  3. Why is representative sampling important?
  4. Which sampling method could help ensure different regions are represented?
  5. How could collecting data only during quiet periods introduce bias?
  6. How could a check sheet support systematic data collection?

16. Key Learning Points

After completing this lesson, the learner should be able to:

  1. Explain the purpose of systematic data collection.
  2. Define a population.
  3. Define a sample.
  4. Explain representative data.
  5. Recognize sampling bias.
  6. Explain random sampling.
  7. Explain stratified sampling.
  8. Explain systematic sampling.
  9. Understand the role of check sheets.
  10. Select an appropriate sampling approach for a process situation.
  11. Recognize the importance of data representativeness before analysis.

17. Lesson Conclusion

Effective Six Sigma analysis depends on reliable and representative data.

The Black Belt should understand:

Population → Sample → Sampling Method → Data Collection → Representativeness → Analysis

Appropriate sampling helps obtain useful information without necessarily studying the entire population.

The key principle is:

Good analysis begins with good data collection.

Lesson 7 — Data Collection & Sampling: COMPLETED

Thank you. Next lesson please

Lesson 8 — Gage R&R — Repeatability & Reproducibility

Learning Objective

By the end of this lesson, the learner should be able to:

  • Explain the purpose of Gage R&R.
  • Distinguish between repeatability and reproducibility.
  • Understand how measurement variation can affect process analysis.
  • Apply Gage R&R concepts to a measurement-system problem.
  • Recognize appropriate actions when measurement variation is identified.

1. Introduction

Before analysing process performance, a Black Belt must understand whether the measurement system itself is producing reliable data.

Gage R&R evaluates variation attributable to a measurement system.

Two important sources of measurement variation are:

  • Repeatability — variation when the same operator measures the same item repeatedly.
  • Reproducibility — variation associated with different operators using the measurement system.

If measurement variation is large, it can obscure the true process signal and weaken subsequent analysis.


2. Understanding Repeatability

Repeatability is the variation observed when:

  • the same operator,
  • measures the same item,
  • repeatedly,
  • using the same measurement system.

For example, an inspector measures the diameter of the same component several times.

If the readings vary substantially, the measurement system may not be sufficiently consistent under repeated measurement.


3. Understanding Reproducibility

Reproducibility is the variation associated with different operators using the measurement system.

For example, three inspectors measure the same product dimensions.

If their measurements differ significantly, the measurement process may be affected by differences between operators.

Possible areas requiring investigation may include:

  • Measurement technique
  • Interpretation of the measurement method
  • Standardization of inspection procedures
  • Operator training

The source specifically identifies different inspection techniques as a cause of reproducibility problems in its case study.


4. Measurement Variation and Process Analysis

A Black Belt must distinguish between:

Actual process variation

and

Measurement-system variation

If the measurement system introduces substantial variation, the observed data may not accurately represent the actual process.

Therefore, measurement-system understanding should precede detailed process analysis.

The purpose is not simply to obtain measurements, but to ensure that the measurements provide a sufficiently reliable basis for decisions.


5. Tools & Techniques

The source identifies the following tools:

  • Gage R&R Study
  • ANOVA Method
  • Minitab

These tools can be used to investigate variation associated with the measurement system and to understand whether repeatability and reproducibility require attention.


6. Application Example

Factory Product-Dimension Inspection

A factory has several inspectors measuring product dimensions.

The Black Belt wants to determine whether the inspectors are obtaining consistent measurements.

A Gage R&R study is conducted to investigate measurement variation.

The analysis considers:

  • Repeated measurements by the same operator.
  • Measurements made by different operators.
  • Whether the measurement system is introducing variation that could affect process analysis.

The results provide information that can guide further investigation of the measurement system.


7. Case Study

Automotive Plant — Reproducibility Problem

An automotive plant identified a reproducibility problem caused by different inspection techniques used by operators.

The organisation responded through:

  • Retraining
  • Standardization of inspection techniques

Following these actions, measurement variation was reduced by 40%.

Black Belt Learning

The case demonstrates that a measurement problem may originate not only from the measuring device but also from differences in how people use the measurement system.

The appropriate response is therefore to investigate the source of variation and address the identified cause.


8. Black Belt Perspective

A Black Belt should ask:

“Can I trust the measurement data before I use it to make process decisions?”

Gage R&R provides a structured way to investigate this question.

The Black Belt should understand:

Measurement System → Measurement Variation → Reliable Data → Process Analysis

If the measurement system is not adequately understood, conclusions about process variation may be misleading.


9. Lesson Practice

  1. Why must the measurement system be understood before process analysis?
  2. What is the difference between repeatability and reproducibility?
  3. What action might follow a reproducibility problem?
  4. A single inspector obtains different readings when repeatedly measuring the same item. Is this primarily a repeatability or reproducibility issue?
  5. Several inspectors obtain different readings for the same item. What type of measurement variation should be investigated?

10. Key Learning Points

  • Gage R&R evaluates variation attributable to a measurement system.
  • Repeatability concerns repeated measurements by the same operator.
  • Reproducibility concerns differences among operators.
  • Measurement variation can obscure the true process signal.
  • Measurement-system understanding is important before relying on process-analysis results.
  • Gage R&R studies, ANOVA and Minitab are relevant tools.
  • Retraining and standardization may be appropriate when reproducibility problems are identified.

11. Lesson Conclusion

Gage R&R helps the Black Belt understand whether measurement-system variation could affect the interpretation of process data.

The essential distinction is:

Repeatability = Same Operator

Reproducibility = Different Operators

A reliable measurement system provides a stronger foundation for the statistical analysis and improvement activities that follow.