Menu Close

1020. Basic Statistics in Process Improvement-Measures of Central Tendency

1020. Basic Statistics in Process Improvement

Basic Statistics in Process Improvement: Measures of Central Tendency

Mastering basic statistics in process improvement is essential for evaluating where a process centers relative to customer specifications. Measures of central tendency summarize an entire dataset using a single representative value that identifies the midpoint or location of process performance.

Applying basic statistics in process improvement ensures that quality teams correctly interpret process centerlines, evaluate baseline capability, and avoid misinterpreting skewed data.

1. Understanding Process Location

Central tendency answers a fundamental operational question: “Where is our process operating on average?”

                  ┌──────────────────────────────────────────────┐
                  │         Measures of Central Tendency         │
                  └──────────────────────┬───────────────────────┘
                                         │
       ┌─────────────────────────────────┼─────────────────────────────────┐
       ▼                                 ▼                                 ▼
┌──────────────────────────────┐ ┌──────────────────────────────┐ ┌──────────────────────────────┐
│       Mean (Average)         │ │    Median (Middle Value)     │ │     Mode (Peak Frequency)    │
│  * Arithmetic center         │ │  * Outlier-resistant center │ │  * Most common value         │
│  * Best for normal data      │ │  * Best for skewed data     │ │  * Best for categorical data │
└──────────────────────────────┘ └──────────────────────────────┘ └──────────────────────────────┘

2. Statistical Foundations: Sample vs. Population

Before calculating central metrics, analysts must distinguish between a Population and a Sample. This distinction determines the mathematical symbols, formulas, and inferences used in data analysis.

  • Population: The complete collection of all elements, items, or outcomes of interest produced by a process. Quantities calculated from a population are called parameters and are represented by Greek letters (e.g., population mean \mu, population size N).

  • Sample: A subset of observations drawn from the population. Quantities calculated from sample data are called statistics and are represented by Roman letters (e.g., sample mean \bar{X}, sample size n).

In real-world process improvement, measuring an entire population is often impossible or cost-prohibitive. Analysts collect representative samples to infer the true parameters of the underlying population.

3. The Three Metrics of Central Tendency

Mean (Arithmetic Average)

The sum of all observed values divided by the total number of observations. It serves as the primary location metric for symmetrical distributions.

  • Sample Mean Formula: \bar{X} = \frac{\sum_{i=1}^{n} X_i}{n}

  • Population Mean Formula: \mu = \frac{\sum_{i=1}^{N} X_i}{N}

Median (Middle Point)

The exact center value when a dataset is ordered from smallest to largest. If the sample size n is even, the median is the arithmetic mean of the two central numbers.

  • Positional Index (Odd n): \text{Position} = \frac{n + 1}{2}

  • Key Advantage: It is robust against extreme outliers and skewed distributions.

Mode (Peak Frequency)

The value that occurs most frequently in a dataset.

  • A dataset with one peak is unimodal.

  • A dataset with two distinct peaks is bimodal (often signaling mixed data sources, such as two different shifts or operators).

  • A dataset with no repeating numbers has no mode.

4. Mathematical Comparison of Location Metrics

Metric Mathematical Definition Outlier Sensitivity Best Used For
Mean \bar{X} = \frac{\sum X}{n} High (Pulled toward extremes) Continuous, normally distributed quantitative data.
Median Central element of ordered set Low (Resistant to extremes) Skewed quantitative data, cycle times, ordinal ratings.
Mode \text{Max Frequency}(X) None Categorical, nominal, and discrete frequency distributions.

5. Worked Operational Examples

Example 1: Symmetrical Shaft Diameters (Normal Data)

A CNC machine produces 5 pin samples with the following measured diameters (in mm):

10.02, 10.04, 10.05, 10.06, 10.08

  • Mean Calculation:

    \bar{X} = \frac{10.02 + 10.04 + 10.05 + 10.06 + 10.08}{5} = \frac{50.25}{5} = 10.05\text{ mm}

  • Median Calculation:

    The dataset is odd (n=5) and already ordered. The 3rd value is 10.05\text{ mm}.

  • Interpretation: Because the data is symmetrical, the mean and median are identical (10.05\text{ mm}), providing a clear center point.

Example 2: Skewed Order Resolution Time (Outlier Impact)

An IT service desk logs 6 ticket resolution times (in minutes):

10, 12, 14, 15, 17, 130

  • Mean Calculation:

    \bar{X} = \frac{10 + 12 + 14 + 15 + 17 + 130}{6} = \frac{198}{6} = 33.0\text{ minutes}

  • Median Calculation:

    The dataset is even (n=6). The middle two terms are 14 and 15.

    \text{Median} = \frac{14 + 15}{2} = 14.5\text{ minutes}

  • Interpretation: The severe outlier of 130\text{ minutes} (a system outage ticket) pulls the mean up to 33.0\text{ minutes}. Reporting the mean gives the false impression that typical requests take over half an hour. The median (14.5\text{ minutes}) far more accurately represents typical customer experience.

6. Impact of Skewness on Central Tendency

The relationship between the mean, median, and mode visually reveals the shape of process data:

  • Symmetrical (Normal) Distribution: \text{Mean} = \text{Median} = \text{Mode}

  • Right-Skewed (Positive Skew): \text{Mean} > \text{Median} > \text{Mode} (Tail extends right toward high values).

  • Left-Skewed (Negative Skew): \text{Mean} < \text{Median} < \text{Mode} (Tail extends left toward low values).

Frequently Asked Questions (FAQ)

Q1: What is the main difference between a sample statistic and a population parameter?

A population parameter represents a true fixed value of an entire population (denoted by Greek letters like \mu), whereas a sample statistic is calculated from a subset of data (denoted by Roman letters like \bar{X}) to estimate the population parameter.

Q2: Why is the mean sensitive to extreme outliers?

The mean includes the exact numerical value of every observation in its summation calculation, meaning one extremely large or small value directly shifts the calculated average.

Q3: Can a dataset have more than one mode?

Yes. A dataset can be bimodal (two modes) or multimodal (multiple modes), which usually indicates that data from two or more distinct process streams were combined.

Q4: Which measure of central tendency is best for ordinal survey ratings?

The median is the preferred central metric for ordinal data (such as satisfaction scales from 1 to 5), as distances between ordinal categories are not mathematically equal.

Q5: What measure of central tendency should be used when data is heavily skewed?

The median should be used when data is heavily skewed, as it represents the 50th percentile of performance without being distorted by extreme tail values.

Six Sigma Practice Exam Questions

1. A quality lead measures an entire batch of 5,000 components produced in a closed lot to determine the true overall average length. The calculated mean of all 5,000 components is classified as a:

A) Sample Statistic

B) Population Parameter

C) Standard Error

D) Sample Median

  • Correct Answer: B) Population Parameter

  • Explanation: When data is gathered from the complete population rather than a subset, calculated summary values represent population parameters (denoted by \mu).

2. A process analyst logs 7 transaction processing times: 5, 5, 7, 8, 9, 12, 45 seconds. What is the median of this dataset?

A) 8\text{ seconds}

B) 13\text{ seconds}

C) 5\text{ seconds}

D) 9\text{ seconds}

  • Correct Answer: A) 8\text{ seconds}

  • Explanation: The dataset has n = 7 values arranged in order. The middle (4th) position value is 8\text{ seconds}.

3. In a right-skewed delivery time distribution, how do the mean and median typically compare?

A) The mean is smaller than the median.

B) The mean is equal to the median.

C) The mean is larger than the median.

D) The median equals the mode.

  • Correct Answer: C) The mean is larger than the median.

  • Explanation: In right-skewed data, long tail values pull the mean toward the right, making it larger than the median.

4. Which central tendency metric is appropriate for analyzing nominal categorical data, such as component defect types?

A) Mean

B) Median

C) Mode

D) Mid-range

  • Correct Answer: C) Mode

  • Explanation: Nominal data consists of un-ordered categories. The mode (most frequent category) is the only valid central metric.

5. A quality check records five dimensions: 1.01, 1.02, 1.02, 1.03, 1.07. What is the calculated sample mean (\bar{X})?

A) 1.020

B) 1.030

C) 1.025

D) 1.035

  • Correct Answer: B) 1.030

  • Explanation: Sum = 1.01 + 1.02 + 1.02 + 1.03 + 1.07 = 5.15. Mean = \frac{5.15}{5} = 1.030.

Written by Ravi Prakash—Quality Expert (38+ yrs exp). Connect on LinkedIn or Contact Us.

Support Our Website

Your support helps us continue delivering good content and useful data. If you’d like to contribute, here’s how:

If you would like to follow this Masterclass series on quality engineering and process statistics, check out our previous and upcoming lessons:

Posted in Continuous Improvement, Measure Phase, Process Improvement, Quality Tools, Six Sigma, Statistics