Skip to main content Skip to course navigation

Data Types, Scales and Sampling Methods

Learning Objectives

  • Explain the role of data in Six Sigma improvement.
  • Distinguish between qualitative and quantitative data.
  • Distinguish between discrete and continuous data.
  • Understand the major measurement scales.
  • Explain nominal, ordinal, interval, and ratio scales.
  • Understand the difference between categorical and numerical data.
  • Identify appropriate data types for Six Sigma analysis.
  • Explain the purpose of sampling.
  • Understand common sampling methods.
  • Distinguish between a population and a sample.
  • Recognize sources of sampling bias.
  • Understand sample size considerations at a basic Green Belt level.
  • Apply appropriate sampling methods to practical Six Sigma situations.

1. Introduction

Six Sigma is a data-driven methodology.

Decisions made during DMAIC should be supported by reliable information rather than assumptions, opinions, or isolated observations.

The quality of analysis depends heavily on the quality and nature of the data being analyzed.

Before selecting a statistical tool, the Green Belt should understand:

  • What type of data is available?
  • How was the data measured?
  • What does each value represent?
  • What measurement scale is being used?
  • Does the data represent the entire population or a sample?
  • How was the sample selected?
  • Could the data contain bias?
  • Is the sample sufficiently representative?

Incorrectly identifying data can lead to inappropriate analysis and misleading conclusions.

Therefore:

Understand the data before selecting the statistical method.

2. Concept / Theory

2.1 What Is Data?

Data consists of recorded observations, measurements, classifications, or information collected for analysis and decision-making.

Examples include:

  • Number of defects.
  • Product dimensions.
  • Customer ratings.
  • Processing time.
  • Number of complaints.
  • Temperature.
  • Delivery status.
  • Defect type.
  • Customer category.

In Six Sigma, data helps the team understand process performance and variation.

2.2 Population and Sample

A population is the complete set of items, transactions, people, measurements, or events of interest.

A sample is a subset of the population selected for study.

For example:

If a company processes 100,000 invoices during a year, the 100,000 invoices may represent the population of interest.

If the Green Belt studies 1,000 invoices selected from that population, the 1,000 invoices constitute a sample.

The objective of sampling is to obtain useful information about the population without necessarily examining every member of it.

2.3 Why Sampling Is Used

Studying an entire population may be:

  • Expensive.
  • Time-consuming.
  • Impractical.
  • Destructive in some testing situations.
  • Unnecessary when an appropriately selected sample can provide useful information.

Sampling can therefore make data collection more practical.

However, the sample must be selected carefully.

A poor sample can produce misleading conclusions even when the statistical analysis itself is technically correct.

3. Key Topics

3.1 Qualitative Data

Qualitative data describes characteristics or categories rather than numerical measurements of magnitude.

Examples include:

  • Defect type.
  • Product color.
  • Customer complaint category.
  • Payment method.
  • Department.
  • Machine identification.
  • Delivery status.

For example:

CustomerComplaint Type
1Delivery Delay
2Billing Error
3Product Damage
4Delivery Delay

The categories provide useful information even though they are not measurements on a numerical scale.

3.2 Quantitative Data

Quantitative data represents numerical quantities.

Examples include:

  • Cycle time.
  • Weight.
  • Temperature.
  • Cost.
  • Length.
  • Number of defects.
  • Number of transactions.

Quantitative data can often be analyzed using statistical measures such as:

  • Mean.
  • Median.
  • Standard deviation.
  • Range.
  • Variance.
  • Percentiles.

3.3 Discrete Data

Discrete data consists of countable values.

Examples include:

  • Number of defects.
  • Number of customer complaints.
  • Number of rejected units.
  • Number of employees.
  • Number of defects per product.

A count generally takes whole-number values.

For example:

0, 1, 2, 3, 4…

A product may have 0 defects, 1 defect, or 2 defects.

It would generally not have 1.7 defects as a count of defects.

3.4 Continuous Data

Continuous data can take any value within a meaningful range of measurement.

Examples include:

  • Length.
  • Weight.
  • Temperature.
  • Time.
  • Pressure.
  • Diameter.
  • Voltage.

For example, a processing time may be:

10 seconds, 10.2 seconds, 10.25 seconds, 10.257 seconds, depending on measurement resolution.

Continuous data is particularly valuable when detailed information about process variation is required.

3.5 Categorical and Numerical Data

Another useful distinction is:

Categorical data identifies groups or classifications.

Examples:

  • Department.
  • Product type.
  • Defect category.
  • Shift.
  • Supplier.

Numerical data represents quantities or measurements.

Examples:

  • Cycle time.
  • Cost.
  • Temperature.
  • Number of defects.
  • Weight.

The nature of the data influences the statistical methods that can appropriately be applied.

3.6 Measurement Scales

Four commonly discussed measurement scales are:

  1. Nominal.
  2. Ordinal.
  3. Interval.
  4. Ratio.

These scales describe the mathematical meaning and structure of the values.


3.7 Nominal Scale

Nominal data represents categories without a meaningful numerical order.

Examples:

  • Department.
  • Product type.
  • Defect type.
  • Country.
  • Payment method.

For example:

CodeDepartment
1Production
2Quality
3Logistics

The numbers are merely labels.

Department 3 is not “greater” than Department 1 in a mathematical sense.

3.8 Ordinal Scale

Ordinal data represents categories that have a meaningful order or ranking.

Examples:

  • Poor / Fair / Good / Excellent.
  • Low / Medium / High.
  • Customer satisfaction ratings expressed as ordered categories.
  • Priority levels.

For example:

Low < Medium < High

The order is meaningful, but the numerical distance between categories is not necessarily equal.

3.9 Interval Scale

Interval data has ordered values with meaningful and equal intervals, but zero does not represent an absolute absence of the measured quantity.

A commonly used example is temperature measured in Celsius.

The difference between:

20°C and 30°C

is the same size interval as the difference between:

30°C and 40°C.

However, 40°C is not meaningfully described as having twice the temperature of 20°C.

3.10 Ratio Scale

Ratio data has:

  • Ordered values.
  • Equal intervals.
  • A meaningful zero point.

Examples include:

  • Weight.
  • Length.
  • Time duration.
  • Distance.
  • Cost.
  • Volume.

For example, 20 kg is twice the mass of 10 kg because zero represents the absence of mass on that measurement scale.

Ratio data supports a broad range of mathematical and statistical operations.

3.11 Summary of Measurement Scales

ScaleCategoriesOrderEqual IntervalsMeaningful Zero
NominalYesNoNoNo
OrdinalYesYesNot necessarilyNo
IntervalNot primarily categoricalYesYesNo
RatioNot primarily categoricalYesYesYes

Understanding the measurement scale helps the Green Belt determine which analytical methods are appropriate.

4. Tools / Methodology

4.1 Sampling

Sampling is the process of selecting a subset of observations from a population for analysis.

The basic sampling process is:

Define Population → Define Sampling Requirement → Select Sample → Collect Data → Analyze Data → Draw Conclusions

The sampling method should be appropriate for the objective of the study.

4.2 Simple Random Sampling

In simple random sampling, each member of the defined population has an appropriate opportunity to be selected according to the randomization procedure.

Example:

A company has 10,000 customer transactions and randomly selects 500 transactions for review.

Random selection helps reduce systematic selection based on convenience or personal preference.

4.3 Systematic Sampling

In systematic sampling, observations are selected at a defined interval after an appropriate starting point.

For example:

  • Select every 20th transaction.
  • Inspect every 50th unit.
  • Review every 10th record.

The sampling interval should be selected appropriately for the population and study objective.

Care should be taken if the process contains a pattern that coincides with the sampling interval.

4.4 Stratified Sampling

In stratified sampling, the population is divided into meaningful subgroups called strata, and samples are obtained from those groups.

Possible strata include:

  • Shift.
  • Region.
  • Product type.
  • Customer segment.
  • Department.
  • Supplier.

Example:

A manufacturing process operates on three shifts.

Instead of sampling only from the day shift, the Green Belt may sample from each shift to ensure that the different operating conditions are represented.

4.5 Cluster Sampling

In cluster sampling, the population is divided into groups or clusters, and selected clusters are sampled.

For example, a company may operate 100 branches and select a number of branches for detailed study.

Cluster sampling can be useful when the population is geographically dispersed or naturally grouped.

4.6 Convenience Sampling

Convenience sampling selects observations because they are easy to access.

For example:

A Green Belt asks the employees who happen to be available in the office that morning to complete a survey.

This approach may be easy, but it can introduce bias because the selected observations may not represent the entire population.

Convenience sampling should therefore be used cautiously when the objective requires representative conclusions.

4.7 Sampling Bias

Sampling bias occurs when the sampling method systematically produces a sample that does not adequately represent the population of interest.

Examples include:

  • Sampling only one shift when all shifts operate differently.
  • Sampling only customers who responded to a voluntary survey.
  • Selecting only recent transactions when older transactions may differ.
  • Selecting only easily accessible records.

A Green Belt should consider whether the sample could systematically differ from the population.

4.8 Sample Size

Sample size refers to the number of observations included in the sample.

The appropriate sample size depends on factors such as:

  • Purpose of the study.
  • Population characteristics.
  • Expected variation.
  • Desired precision.
  • Statistical method.
  • Confidence requirements.
  • Practical constraints.

There is no single sample size that is automatically correct for every Six Sigma project.

A Green Belt should avoid the assumption that “more data is always better.”

The objective is to obtain sufficient and appropriate data for the intended analysis.

4.9 Sampling Plan

A basic Sampling Plan should specify:

  • What will be measured?
  • From which population?
  • Where will the data be collected?
  • When will the data be collected?
  • How many observations are required?
  • How will observations be selected?
  • Who will collect the data?
  • What conditions should be recorded?

A well-defined Sampling Plan improves consistency and reduces unintended selection bias.

5. Worked Example / Case Study

Case Study — Manufacturing Defects Across Three Shifts

A manufacturing plant operates three shifts:

  • Shift A.
  • Shift B.
  • Shift C.

Management has observed an increase in product defects.

The Green Belt wants to investigate whether defect performance differs across operating conditions.

Step 1 — Define the Population

The population consists of production units manufactured during the defined study period.

Step 2 — Identify Relevant Groups

Because the process operates under three shifts, shift may be an important source of variation.

Step 3 — Select a Sampling Approach

The Green Belt may use stratified sampling so that observations from all three shifts are represented.

Step 4 — Collect Data

For each sampled unit, the team may record:

  • Shift.
  • Product type.
  • Date.
  • Defect status.
  • Defect category.
  • Relevant process measurements.

Step 5 — Analyze

The data can then be examined to determine whether defect patterns differ among the shifts.

The important principle is that the sampling plan should represent the process conditions relevant to the question being investigated.

Example of Poor Sampling

Suppose the Green Belt collects all observations from Shift A because that shift is easiest to access.

If Shift B and Shift C operate under different conditions, the resulting sample may not adequately represent the process.

The problem would not necessarily be the statistical analysis. The problem would be the sampling approach.

6. Practical Application

Exercise 1 — Customer Complaints

A company receives customer complaints from four regions:

  • North.
  • South.
  • East.
  • West.

The Green Belt wants to understand complaint types across the organization.

Questions:

  • What is the population?
  • What is the sample?
  • Should region be considered when sampling?
  • Which sampling method could ensure representation from each region?
  • What variables should be recorded?

Exercise 2 — Production Inspection

A production line manufactures 20,000 units per day.

The Green Belt cannot inspect every unit.

Questions:

  • What is the population?
  • What would constitute a sample?
  • Could systematic sampling be used?
  • What potential process patterns should be considered?
  • What information should be included in the Sampling Plan?

Exercise 3 — Customer Satisfaction

A company wants to measure customer satisfaction.

A team decides to ask only customers who voluntarily respond to an online survey.

Questions:

  • What potential sampling issue exists?
  • Could the respondents differ from customers who do not respond?
  • What steps could be considered to obtain more representative information?

6.1 Selecting the Appropriate Data Type

Consider these variables:

VariableLikely Data Type
Number of defectsDiscrete
Product weightContinuous
Defect categoryCategorical
Processing timeContinuous
Number of complaintsDiscrete
Customer satisfaction categoryOrdinal
Department nameNominal
Temperature in °CInterval

Correctly identifying the data type is an important preliminary step before statistical analysis.

6.2 Data Collection and Operational Definitions

A data value has little meaning if the team does not clearly define what it represents.

For example:

“Response Time”

could mean:

  • Time from customer email receipt to first response.
  • Time from customer email receipt to complete resolution.
  • Business hours only.
  • Calendar hours.

Therefore, the Green Belt should establish an appropriate operational definition.

An operational definition specifies exactly how a characteristic or measure will be identified and recorded.

7. Lesson Summary

Data is the foundation of data-driven Six Sigma improvement.

Important data classifications include:

  • Qualitative and quantitative.
  • Categorical and numerical.
  • Discrete and continuous.

The four common measurement scales are:

  • Nominal.
  • Ordinal.
  • Interval.
  • Ratio.

The Green Belt must understand the difference between a population and a sample.

Sampling provides a practical way to study a population, but the sampling method must be appropriate.

Common sampling methods include:

  • Simple random sampling.
  • Systematic sampling.
  • Stratified sampling.
  • Cluster sampling.

Convenience sampling is easy to perform but may introduce bias.

A sound Sampling Plan specifies what data will be collected, from whom or what, when, where, how, and by whom.

The quality of the statistical analysis depends substantially on the quality and appropriateness of the data collected.

8. Lesson Learnt / Conclusion

A Green Belt should never treat data as simply a collection of numbers.

Every data value has a context:

What was measured?
How was it measured?
From which population?
When was it measured?
How was the observation selected?
What does the value actually represent?

Understanding these questions helps prevent incorrect conclusions.

Appropriate data classification and sampling provide a strong foundation for the statistical analysis that follows in DMAIC.

Green Belt Takeaway:

Correct data classification and representative sampling are essential foundations for reliable Six Sigma analysis.