Services Combination Products Approach Industries Insights About Book a consultation

Prefer email? info@dpdynamicsol.com

Statistics

Process Capability (Cpk and Ppk): What Auditors Actually Expect

A capability index is a single number that compresses an entire process into a verdict. That is what makes it useful, and what makes it so easy to misuse. Most capability findings are not about the arithmetic — the arithmetic is trivial. They are about everything that should have been established before anyone computed it.

Key takeaways

  • Capability indices are meaningless on a process that is not in statistical control. Demonstrate stability first, or the number describes nothing.
  • Cpk uses within-subgroup variation; Ppk uses total variation. A large gap between them is a stability problem, not a rounding difference.
  • Three prerequisites decide whether the number is credible: stability, normality, and a qualified measurement system.
  • 1.33 is a convention, not a requirement. Tie your threshold to harm severity and say so.
  • A point estimate is not evidence. A Cpk of 1.33 from thirty units has a lower confidence bound near 1.0.

What the indices actually measure

All four indices answer variations of one question: how much room is there between the process and the specification limits, measured in standard deviations?

Cp compares the total specification width to the process spread. It assumes the process is centred, which is why it describes potential rather than reality:

Cp= USLLSL 6σ

Cpk drops the centring assumption by taking the worse of the two sides. This is why Cpk is the index that matters operationally — a process can have excellent spread and still sit dangerously close to one limit:

Cpk= min[ USLx¯ 3σ , x¯LSL 3σ ]

Readers of our sample size article will recognise the numerator: it is the same margin that drives variables sample size, just divided by three so that a value of 1.0 corresponds to three standard deviations of room.

Pp and Ppk use exactly the same formulas. The difference is entirely in which standard deviation you substitute — and that difference is the single most misunderstood point in the subject.

Cpk versus Ppk: the gap is the diagnosis

The distinction is not cosmetic and it is not a matter of preference.

  • Cpk uses the within-subgroup standard deviation, estimated from short-term variation inside rational subgroups. It describes what the process is capable of when nothing shifts — its inherent potential.
  • Ppk uses the overall standard deviation, computed from all the data as one set. It captures everything: within-subgroup noise plus any drift, shift, tool wear, or lot-to-lot movement across the study.

Because the overall standard deviation includes sources the within-subgroup estimate excludes, Ppk is normally the lower of the two. The informative quantity is the size of the gap:

ObservationWhat it means
Cpk ≈ PpkThe process is stable. Between-subgroup variation is negligible. Both numbers can be trusted.
Cpk > Ppk noticeablySomething is moving between subgroups — drift, shifts, setup differences. The process has potential it is not delivering.
Cpk ≫ PpkThe process is not in control. Neither index is valid yet. Fix the instability before quoting any capability number.
Report both, and explain the gap. A validation report showing only the higher of the two invites the obvious question, and the honest answer — "we reported Cpk because it looked better" — is not one you want to give. Reporting both, with a sentence interpreting the difference, demonstrates you understand your own process.

The three prerequisites nobody documents

1. Statistical control comes first

This is the finding that recurs most often, and it is conceptually fatal rather than merely procedural. Capability indices assume the process has a distribution — one mean, one standard deviation, stable over time. An out-of-control process does not have a single distribution; it has a moving one. Computing a capability index on it produces a number that describes no real state of the world.

Demonstrate control with an appropriate control chart before computing capability, and keep that chart in the record. If special-cause signals are present, investigate and resolve them — do not delete the points and recompute. Removing inconvenient data without a documented, investigated cause is a serious finding in its own right.

2. Normality is an assumption, not a formality

The familiar relationship between a capability index and a defect rate depends on the data being approximately normal. Many real device characteristics are not: flatness, concentricity, and particle counts are bounded at zero and often skewed; timing and force data can be bimodal when two tools or cavities are pooled.

Check it, record the check, and if the data is not normal either transform it with justification, fit an appropriate alternative distribution, or use a percentile-based capability method of the kind described in the ISO 22514 series. What is not acceptable is computing an index as though the data were normal and never mentioning the question.

3. The measurement system must be qualified first

Every observed value carries both product variation and measurement variation, and the observed spread combines them. A noisy gauge inflates your estimate of process spread, which deflates capability — so an adequate process can fail on the strength of a poor measurement system.

Run a measurement system analysis before the capability study, not after it disappoints. If gauge variation consumes a large share of the total, you are substantially measuring your instrument rather than your process, and no amount of additional data will fix it.

What the numbers actually correspond to

It helps to remember what an index implies about conformance, assuming a normal, centred, stable process:

CpkSigma to nearest limitApprox. nonconformingTypical use
1.00≈ 2,700 ppm (0.27%)Generally inadequate for a controlled characteristic
1.33≈ 63 ppmThe common default expectation
1.67≈ 0.6 ppmHigh-severity characteristics
2.00≈ 0.002 ppmCritical / life-sustaining functions

Two cautions. These figures assume a centred process; for an off-centre process the nonconforming fraction concentrates on the near side. And the widely quoted "3.4 parts per million" associated with six sigma is not the value above — it incorporates an assumed long-term mean shift of 1.5 standard deviations. Quoting 3.4 ppm alongside a short-term Cpk of 2.0 conflates two different conventions.

Choosing a threshold: 1.33 is a convention, not a law

Asked why the acceptance criterion is 1.33, most teams answer that it is standard. That answer has the same weakness as "we tested thirty units" — it describes a habit, not a rationale.

Tie the threshold to the severity of harm if the characteristic fails, using the same logic as your ISO 14971 analysis, and record the tiering in a procedure so it is applied consistently rather than negotiated per study:

Severity of harm if the characteristic failsTypical minimum capability
Negligible / minorCpk ≥ 1.00
SeriousCpk ≥ 1.33
Critical / life-threateningCpk ≥ 1.67, sometimes 2.00

This is a framework you own and must defend, not a rule handed down by any standard. Its value is that it converts an arbitrary-looking threshold into a traceable decision: this characteristic carries this severity, therefore this capability requirement.

The point estimate is not the evidence

Here is the part most capability reports omit entirely. A capability index computed from a sample is an estimate, and estimates carry uncertainty. For moderate sample sizes the approximate confidence interval is:

Cpk ± zα/2 19n + Cpk2 2(n1)

An approximation for reasonably large n. Use exact methods or qualified software for small samples, and cite the method in your protocol.

Apply it and the implications are uncomfortable. Suppose you observe a Cpk of exactly 1.33 — apparently a clean pass against the common threshold. The 95% lower confidence bound tells a different story:

Sample size (n)Observed Cpk95% lower confidence boundCan you claim ≥ 1.33?
301.33≈ 1.03No — barely clears 1.0
501.33≈ 1.10No
1001.33≈ 1.17No
2001.33≈ 1.21No

An observed Cpk equal to your threshold never demonstrates that the true capability meets it — it is a coin flip. To make the claim with confidence you need either a comfortably higher observed value or a much larger sample. This is why capability studies built on a single small batch are so fragile, and why an auditor who asks "what is the confidence interval on that?" is asking a fair and often unanswerable question.

Practical consequence: set your acceptance criterion on the lower confidence bound, not the point estimate — for example, "the 95% lower confidence bound on Cpk shall exceed 1.33." It is a stricter test, it is honest, and it is far more defensible than a bare number that happens to land above the line.

What auditors probe

  1. Was stability demonstrated first? Show the control chart. Without it, the index is unsupported.
  2. Cpk or Ppk — which did you report, and do you know why? The inability to explain the difference signals the analysis was run by software rather than understood.
  3. Was normality assessed? If the report does not show the check, assume it was not done.
  4. Is the measurement system qualified? Recent, relevant, and covering the same characteristic and range.
  5. Where is the threshold justified? "Industry standard" is not a justification. Severity-based tiering is.
  6. Is the sample representative? Multiple lots, operators, tools, and cavities — or a documented reason why not.
  7. What happens when it degrades? Capability at validation is a snapshot. What monitors it in production, and what triggers action?

Capability is not a one-time event

A capability index established during validation describes the process at that moment. Tools wear, suppliers change lots, operators turn over, environments drift. A process qualified at 1.5 two years ago may not be at 1.5 today, and the only way to know is to keep looking.

Build ongoing monitoring into the control plan: continued verification of the characteristic, periodic recalculation on production data, and a defined action threshold that triggers investigation before conformance is actually lost. This is also where capability connects back to the quality system — a downward capability trend is exactly the kind of signal that should feed CAPA and the risk file rather than sit unread in a report.

Common pitfalls

  • Computing capability on an unstable process. The most common and most fundamental error. Stability first, always.
  • Reporting whichever index looks better. Report both Cpk and Ppk and interpret the gap.
  • Silent normality assumptions. An unstated assumption is an undocumented one.
  • Quoting a point estimate as proof. Without a confidence bound you have an observation, not a demonstration.
  • One lot, one shift, one operator. Cheap to run, and it characterises that lot rather than the process.
  • Applying two-sided logic to a one-sided specification. If only one limit exists, only one ratio is meaningful.
  • Deleting outliers to rescue the number. Points may be excluded only with an investigated, documented cause.
  • Never recomputing after a process change. A validated capability is a claim about the process as it was, not as it is.

What to document

A capability study that will survive review states, in the protocol before execution: the characteristic and its specification limits, with an explicit note on whether they are one- or two-sided; its risk classification and the resulting capability threshold; the measurement system qualification status; the sampling plan across lots and sources of variation, with subgrouping rationale; the method for demonstrating stability; the normality assessment method and the plan if the data is not normal; which indices will be reported and which standard deviation each uses; and the acceptance criterion, stated on a confidence bound rather than a point estimate.

Written that way, the capability index stops being a number someone has to defend under pressure and becomes the visible conclusion of a documented chain of reasoning — which is the same standard every other statistical claim in your quality system should meet.

This article is general information, not statistical or regulatory advice. Statistical methods must be selected and applied by qualified personnel with reference to validated software, applicable standards, and the specific characteristics of your process and product.

David Plescia, Founder & Principal Consultant, DP Dynamic Solutions David Plescia Founder & Principal Consultant, DP Dynamic Solutions · 25+ years in medtech quality

DP Dynamic Solutions provides statistical support for medical device and combination product programmes — process capability and SPC, measurement system analysis, sample size justification, statistical analysis plans, and process validation support.

Statistics services
Let's talk

Is your capability data telling the truth?

We review capability studies the way an assessor would — stability, normality, measurement system, sampling, and whether the claim survives its own confidence interval.