Basic Statistics in Process Improvement: Measures of Spread
What is covered in this Article
Understanding basic statistics in process improvement requires analyzing both process location (central tendency) and process spread (dispersion). While measures of central tendency identify where a process centers, measures of variation quantify the inconsistency, width, and predictability of operational outcomes. Two processes can share an identical average yet perform radically differently in quality due to variations in process spread.
Mastering basic statistics in process improvement empowers teams to evaluate baseline variation, identify extreme performance tails, and calculate precise process capability boundaries.
1. Understanding Process Spread
Process spread quantifies how far data points scatter around the center point.
┌──────────────────────────────────────────────┐
│ Measures of Variation │
└──────────────────────┬───────────────────────┘
│
┌─────────────────────────────────┼─────────────────────────────────┐
▼ ▼ ▼
┌──────────────────────────────┐ ┌──────────────────────────────┐ ┌──────────────────────────────┐
│ Range (R) │ │ Variance (σ² or s²) │ │ Std Deviation (σ or s) │
│ * Simplest spread metric │ │ * Average squared distance │ │ * Original unit variation │
│ * Uses extreme limits │ │ * Additive statistical tool │ │ * Key capability input │
└──────────────────────────────┘ └──────────────────────────────┘ └──────────────────────────────┘
2. Key Metrics of Process Dispersion
Range (
)
The mathematical difference between the maximum and minimum observed values in a sample dataset.
-
Limitation: Highly sensitive to extreme outliers because it relies exclusively on two boundary values.
Interquartile Range (IQR)
The range covering the middle of the data, calculated as the difference between the 3rd quartile (
or 75th percentile) and 1st quartile (
or 25th percentile).
-
Key Advantage: Highly resistant to extreme outliers, making it ideal for non-normal or heavily skewed datasets.
Variance (
or
)
The arithmetic average of squared deviations from the mean. It quantifies total variability across all data points.
-
Population Variance Formula:
-
Sample Variance Formula:
-
Bessel’s Correction (
): Dividing by
instead of
corrects sample estimation bias, ensuring an unbiased estimate of population variance.
Standard Deviation (
or
)
The square root of variance. It restores the measurement of spread back to the original operational units (e.g., millimeters, seconds, or grams).
-
Population Standard Deviation:
-
Sample Standard Deviation:
3. Mathematical Comparison of Spread Metrics
| Metric | Mathematical Formula | Sensitivity to Outliers | Primary Operational Application |
| Range ( |
Extreme | Small control chart subgroups ( |
|
| Interquartile Range (IQR) | Very Low | Skewed processes, cycle time non-parametric data. | |
| Sample Variance ( |
Moderate-High | Statistical modeling and ANOVA variance decomposition. | |
| Sample Std Dev ( |
Moderate-High | Process capability indices ( |
4. Worked Step-by-Step Calculation Example
Scenario
A beverage bottling line logs fill volumes (in milliliters) across 5 consecutive bottle samples:
Step 1: Calculate Sample Mean (
)
Step 2: Calculate Deviations and Squared Deviations
Sum of Squared Deviations ():
Step 3: Calculate Sample Variance (
)
Step 4: Calculate Sample Standard Deviation (
)
Step 5: Calculate Range (
)
Interpretation
While the average fill volume is , individual bottles vary with a standard deviation of
across a total span of
.
Frequently Asked Questions (FAQ)
Q1: Why do we square deviations when calculating variance?
Squaring deviations eliminates negative signs so that negative and positive deviations do not cancel each other out when summed. It also gives greater mathematical weight to larger deviations.
Q2: Why is standard deviation preferred over variance in process capability reporting?
Variance is expressed in squared units (e.g., ), which cannot be directly compared with operational tolerances. Standard deviation returns the metric to original units (e.g.,
).
Q3: When should the Interquartile Range (IQR) be used instead of standard deviation?
The Interquartile Range should be used when data is heavily skewed or contains extreme outliers, as it focuses strictly on the central of observations without distortion from extreme tails.
Q4: What is the relation between standard deviation and normal distribution z-scores?
The standard deviation serves as the distance metric along a normal distribution curve. Approximately of data lies within
,
within
, and
within
.
Q5: Why is sample variance calculated using instead of
?
Using (Bessel’s correction) compensates for the fact that sample data tends to underestimate true population variability, yielding an mathematically unbiased estimate of population variance.
Six Sigma Practice Exam Questions
1. An engineer measures 5 machined pins and obtains a sum of squared deviations from the sample mean equal to . What is the calculated sample variance (
)?
A)
B)
C)
D)
-
Correct Answer: B)
-
Explanation: Sample variance is calculated as
.
2. Which measure of spread is most appropriate for describing process variability in a heavily skewed non-parametric dataset containing extreme outliers?
A) Range
B) Population Standard Deviation
C) Interquartile Range (IQR)
D) Sample Variance
-
Correct Answer: C) Interquartile Range (IQR)
-
Explanation: IQR measures the spread of the middle
of data (
), providing an outlier-resistant metric for skewed data.
3. What is the sample standard deviation () of a process dataset with a calculated sample variance of
?
A)
B)
C)
D)
-
Correct Answer: A)
-
Explanation: Standard deviation is the square root of variance:
.
4. A quality technician calculates sample range across three consecutive subgroups. If maximum values are 12, 15, and 14, and minimum values are 8, 10, and 9 respectively, what is the average subgroup range ()?
A)
B)
C)
D)
-
Correct Answer: A)
-
Explanation: Ranges for each subgroup are:
,
, and
. Average Range
.
5. Why is the range metric considered less robust than standard deviation when analyzing sample sizes larger than ?
A) Range cannot be converted into metric units.
B) Range uses only the two extreme values and completely ignores sample points in between.
C) Range produces negative values for continuous data.
D) Range requires prior knowledge of population parameters.
-
Correct Answer: B) Range uses only the two extreme values and completely ignores sample points in between.
-
Explanation: Range depends strictly on
, ignoring intermediate values and becoming increasingly prone to extreme outlier distortion in larger datasets.
Written by Ravi Prakash—Quality Expert (38+ yrs exp). Connect on LinkedIn or Contact Us.
Support Our Website
Your support helps us continue delivering good content and useful data. If you’d like to contribute, here’s how:
- From anywhere in the world (PayPal): https://www.paypal.com/paypalme/rpbehara
- From India (UPI): speakingdata@ybl
If you would like to follow this Masterclass series on quality engineering and process statistics, check out our previous and upcoming lessons: