Menu Close

1021. Basic Statistics in Process Improvement-Measures of Spread

1021. Basic Statistics in Process Improvement

Basic Statistics in Process Improvement: Measures of Spread

Understanding basic statistics in process improvement requires analyzing both process location (central tendency) and process spread (dispersion). While measures of central tendency identify where a process centers, measures of variation quantify the inconsistency, width, and predictability of operational outcomes. Two processes can share an identical average yet perform radically differently in quality due to variations in process spread.

Mastering basic statistics in process improvement empowers teams to evaluate baseline variation, identify extreme performance tails, and calculate precise process capability boundaries.

1. Understanding Process Spread

Process spread quantifies how far data points scatter around the center point.

                  ┌──────────────────────────────────────────────┐
                  │             Measures of Variation            │
                  └──────────────────────┬───────────────────────┘
                                         │
       ┌─────────────────────────────────┼─────────────────────────────────┐
       ▼                                 ▼                                 ▼
┌──────────────────────────────┐ ┌──────────────────────────────┐ ┌──────────────────────────────┐
│          Range (R)           │ │  Variance (σ² or s²)         │ │  Std Deviation (σ or s)      │
│  * Simplest spread metric    │ │  * Average squared distance  │ │  * Original unit variation   │
│  * Uses extreme limits       │ │  * Additive statistical tool │ │  * Key capability input      │
└──────────────────────────────┘ └──────────────────────────────┘ └──────────────────────────────┘

2. Key Metrics of Process Dispersion

Range (R)

The mathematical difference between the maximum and minimum observed values in a sample dataset.

R = X_{\text{max}} - X_{\text{min}}

  • Limitation: Highly sensitive to extreme outliers because it relies exclusively on two boundary values.

Interquartile Range (IQR)

The range covering the middle 50% of the data, calculated as the difference between the 3rd quartile (Q_3 or 75th percentile) and 1st quartile (Q_1 or 25th percentile).

\text{IQR} = Q_3 - Q_1

  • Key Advantage: Highly resistant to extreme outliers, making it ideal for non-normal or heavily skewed datasets.

Variance (\sigma^2 or s^2)

The arithmetic average of squared deviations from the mean. It quantifies total variability across all data points.

  • Population Variance Formula: \sigma^2 = \frac{\sum_{i=1}^{N} (X_i - \mu)^2}{N}

  • Sample Variance Formula: s^2 = \frac{\sum_{i=1}^{n} (X_i - \bar{X})^2}{n - 1}

  • Bessel’s Correction (n-1): Dividing by n - 1 instead of n corrects sample estimation bias, ensuring an unbiased estimate of population variance.

Standard Deviation (\sigma or s)

The square root of variance. It restores the measurement of spread back to the original operational units (e.g., millimeters, seconds, or grams).

  • Population Standard Deviation: \sigma = \sqrt{\sigma^2}

  • Sample Standard Deviation: s = \sqrt{s^2}

3. Mathematical Comparison of Spread Metrics

Metric Mathematical Formula Sensitivity to Outliers Primary Operational Application
Range (R) X_{\text{max}} - X_{\text{min}} Extreme Small control chart subgroups (n \le 10).
Interquartile Range (IQR) Q_3 - Q_1 Very Low Skewed processes, cycle time non-parametric data.
Sample Variance (s^2) \frac{\sum (X_i - \bar{X})^2}{n - 1} Moderate-High Statistical modeling and ANOVA variance decomposition.
Sample Std Dev (s) \sqrt{s^2} Moderate-High Process capability indices (C_p, C_{pk}) and control charts.

4. Worked Step-by-Step Calculation Example

Scenario

A beverage bottling line logs fill volumes (in milliliters) across 5 consecutive bottle samples:

498, 500, 501, 503, 508

Step 1: Calculate Sample Mean (\bar{X})

\bar{X} = \frac{498 + 500 + 501 + 503 + 508}{5} = \frac{2510}{5} = 502.0\text{ mL}

Step 2: Calculate Deviations and Squared Deviations
  • (498 - 502)^2 = (-4)^2 = 16

  • (500 - 502)^2 = (-2)^2 = 4

  • (501 - 502)^2 = (-1)^2 = 1

  • (503 - 502)^2 = (1)^2 = 1

  • (508 - 502)^2 = (6)^2 = 36

Sum of Squared Deviations (SS):

SS = 16 + 4 + 1 + 1 + 36 = 58

Step 3: Calculate Sample Variance (s^2)

s^2 = \frac{SS}{n - 1} = \frac{58}{5 - 1} = \frac{58}{4} = 14.5\text{ mL}^2

Step 4: Calculate Sample Standard Deviation (s)

s = \sqrt{14.5} \approx 3.808\text{ mL}

Step 5: Calculate Range (R)

R = 508 - 498 = 10.0\text{ mL}

Interpretation

While the average fill volume is 502.0\text{ mL}, individual bottles vary with a standard deviation of 3.81\text{ mL} across a total span of 10.0\text{ mL}.

Frequently Asked Questions (FAQ)

Q1: Why do we square deviations when calculating variance?

Squaring deviations eliminates negative signs so that negative and positive deviations do not cancel each other out when summed. It also gives greater mathematical weight to larger deviations.

Q2: Why is standard deviation preferred over variance in process capability reporting?

Variance is expressed in squared units (e.g., \text{mL}^2), which cannot be directly compared with operational tolerances. Standard deviation returns the metric to original units (e.g., \text{mL}).

Q3: When should the Interquartile Range (IQR) be used instead of standard deviation?

The Interquartile Range should be used when data is heavily skewed or contains extreme outliers, as it focuses strictly on the central 50% of observations without distortion from extreme tails.

Q4: What is the relation between standard deviation and normal distribution z-scores?

The standard deviation serves as the distance metric along a normal distribution curve. Approximately 68.27% of data lies within \pm 1\sigma, 95.45% within \pm 2\sigma, and 99.73% within \pm 3\sigma.

Q5: Why is sample variance calculated using n - 1 instead of n?

Using n - 1 (Bessel’s correction) compensates for the fact that sample data tends to underestimate true population variability, yielding an mathematically unbiased estimate of population variance.

Six Sigma Practice Exam Questions

1. An engineer measures 5 machined pins and obtains a sum of squared deviations from the sample mean equal to 32.0\text{ mm}^2. What is the calculated sample variance (s^2)?

A) 6.4\text{ mm}^2

B) 8.0\text{ mm}^2

C) 2.83\text{ mm}^2

D) 16.0\text{ mm}^2

  • Correct Answer: B) 8.0\text{ mm}^2

  • Explanation: Sample variance is calculated as s^2 = \frac{SS}{n - 1} = \frac{32.0}{5 - 1} = \frac{32.0}{4} = 8.0\text{ mm}^2.

2. Which measure of spread is most appropriate for describing process variability in a heavily skewed non-parametric dataset containing extreme outliers?

A) Range

B) Population Standard Deviation

C) Interquartile Range (IQR)

D) Sample Variance

  • Correct Answer: C) Interquartile Range (IQR)

  • Explanation: IQR measures the spread of the middle 50% of data (Q_3 - Q_1), providing an outlier-resistant metric for skewed data.

3. What is the sample standard deviation (s) of a process dataset with a calculated sample variance of 0.0049\text{ grams}^2?

A) 0.07\text{ grams}

B) 0.0007\text{ grams}

C) 0.49\text{ grams}

D) 0.00245\text{ grams}

  • Correct Answer: A) 0.07\text{ grams}

  • Explanation: Standard deviation is the square root of variance: s = \sqrt{0.0049} = 0.07\text{ grams}.

4. A quality technician calculates sample range across three consecutive subgroups. If maximum values are 12, 15, and 14, and minimum values are 8, 10, and 9 respectively, what is the average subgroup range (\bar{R})?

A) 4.67

B) 5.00

C) 4.00

D) 6.00

  • Correct Answer: A) 4.67

  • Explanation: Ranges for each subgroup are: 12 - 8 = 4, 15 - 10 = 5, and 14 - 9 = 5. Average Range \bar{R} = \frac{4 + 5 + 5}{3} = \frac{14}{3} \approx 4.67.

5. Why is the range metric considered less robust than standard deviation when analyzing sample sizes larger than n = 10?

A) Range cannot be converted into metric units.

B) Range uses only the two extreme values and completely ignores sample points in between.

C) Range produces negative values for continuous data.

D) Range requires prior knowledge of population parameters.

  • Correct Answer: B) Range uses only the two extreme values and completely ignores sample points in between.

  • Explanation: Range depends strictly on X_{\text{max}} - X_{\text{min}}, ignoring intermediate values and becoming increasingly prone to extreme outlier distortion in larger datasets.

Written by Ravi Prakash—Quality Expert (38+ yrs exp). Connect on LinkedIn or Contact Us.

Support Our Website

Your support helps us continue delivering good content and useful data. If you’d like to contribute, here’s how:

If you would like to follow this Masterclass series on quality engineering and process statistics, check out our previous and upcoming lessons:

Posted in Continuous Improvement, Measure Phase, Process Improvement, Quality Tools, Six Sigma, Statistics