"We tested thirty units" is a sample size. It is not a justification. The question a reviewer or auditor is really asking is: what claim are you making about the population, and with what confidence? Answer that, and the number falls out of the arithmetic instead of out of habit.
Key takeaways
- Sample size is an output of your claim, not an input you pick first.
- Attribute (pass/fail) testing is statistically expensive — 95%/95% with zero failures requires 59 units.
- Measuring a variable instead of scoring pass/fail typically cuts that sample size by half or more.
- Variables sample size is driven by margin — how many standard deviations separate your process from the limit. You need a capability estimate to size the study.
- Below roughly 0.55 Cpk, no sample size will pass a 95%/95% claim. That is a design problem, not a sampling problem.
- Confidence and reliability should be tiered by risk, and traceable to your ISO 14971 analysis.
- Write the rationale before the test. A justification produced after the data exists is worth very little.
Why "n=30" is not a justification
Thirty is a folk number. It comes from a rule of thumb about when a sampling distribution of means becomes roughly normal — a statement about estimating an average, which is almost never what a design verification test is about. Verification usually asserts something much stronger: that a specified proportion of the population conforms, with stated confidence.
The distinction matters in an audit. "Thirty is standard practice" invites the follow-up question that has no good answer: standard for what reliability, at what confidence, for a failure mode with what severity? A justification that cannot survive that question is a finding waiting to happen.
Start with the claim, not the number
Every defensible sample size begins with a sentence in this shape:
X is confidence — how sure you are about the estimate, driven by sample size. Y is reliability — the proportion of the population you are claiming conforms. Both are choices you must make deliberately and defend, and together with the data type they determine the sample size completely.
Confidence and reliability are frequently confused because both are expressed as percentages and both are commonly set to 95. They are not the same thing. Increasing confidence means being more certain about the same claim; increasing reliability means making a stronger claim. Reliability is by far the more expensive of the two.
The biggest lever: attribute versus variable data
Before any arithmetic, ask what kind of data the test actually produces.
- Attribute data is categorical — pass/fail, leak/no leak, present/absent. Each unit contributes one bit of information.
- Variable data is measured on a continuous scale — force in newtons, pressure in kPa, dimension in millimetres. Each unit contributes a value and its distance from the specification limit.
That extra information is worth a great deal. A unit that passes tells you only that it passed. A unit that measured 42 N against a 20 N minimum tells you it passed with enormous margin — evidence that the population is comfortably clear of the limit.
Attribute path: the success-run theorem
When the test yields pass/fail and you plan to accept on zero failures, the required sample size is:
where is confidence and is reliability, both expressed as decimals, and is rounded up to the next whole unit.
Some common results, all assuming zero failures observed:
| Confidence / Reliability | Required n (zero failures) | Typical use |
|---|---|---|
| 90% / 90% | 22 | Low-risk attributes, early development |
| 95% / 90% | 29 | Moderate risk |
| 90% / 95% | 45 | Moderate risk, stronger claim |
| 95% / 95% | 59 | The common default for verification |
| 95% / 99% | 299 | High-severity failure modes |
| 99% / 99% | 459 | Life-supporting / life-sustaining functions |
Two consequences worth internalising. First, the jump from 95% to 99% reliability is a fivefold increase in units — reliability is expensive, so choose it on the basis of risk rather than instinct. Second, this table assumes zero failures. A single failure does not simply cost you one unit; it invalidates the zero-failure basis, and re-establishing the same claim while allowing one failure requires a substantially larger sample. Plan for what happens if a unit fails before you start.
Variables path: tolerance intervals
With measured data you can use a one-sided normal tolerance interval, which asks whether the whole distribution — not just the units you tested — sits clear of the specification limit. For a lower specification limit, the test is:
where is the sample mean, the sample standard deviation, a factor determined by sample size, confidence, and reliability, and the lower specification limit.
The factor shrinks as grows — more data, less penalty for uncertainty. But notice the problem: this is an acceptance test. You cannot apply it until the data exists, so on its own it tells you nothing about how many units to build.
Closing the loop: how to choose n before you have data
Rearrange the acceptance condition and the answer appears. Dividing through by :
The left side is your margin — how many standard deviations separate the process from the specification limit. The right side depends only on sample size, confidence, and reliability.
That reframes the whole question. Sample size for variables data is driven by how much margin your process has. A capable process needs few units; a marginal process needs many; and a process without enough margin cannot pass at any sample size.
The margin term is directly related to process capability. For a one-sided lower limit, is that same distance expressed in units of , so the margin equals . That gives a table you can size a study from:
| Sample size (n) | k (95% / 95%, one-sided) | Margin required | ≈ Cpk required |
|---|---|---|---|
| 10 | ≈ 2.91 | 2.91 s | 0.97 |
| 15 | ≈ 2.57 | 2.57 s | 0.86 |
| 20 | ≈ 2.40 | 2.40 s | 0.80 |
| 30 | ≈ 2.22 | 2.22 s | 0.74 |
| 50 | ≈ 2.07 | 2.07 s | 0.69 |
| 100 | ≈ 1.93 | 1.93 s | 0.64 |
| ∞ (theoretical floor) | 1.645 | 1.65 s | 0.55 |
Values are illustrative — take exact factors from a validated table or qualified statistical software, and cite the source in your protocol.
The procedure is then three steps:
- Estimate your margin from prior data. Use pilot builds, development lots, a similar predicate process, or a design target. Compute the expected margin as . Be conservative: use the upper plausible value of , not the best case.
- Read off the smallest n whose k is below that margin. If your expected margin is 2.5 standard deviations, n = 20 (k ≈ 2.40) clears it; n = 15 (k ≈ 2.57) does not.
- Add headroom. Your estimate of is itself uncertain, and if the real spread comes in higher you fail the study after building all the units. Stepping up one row is cheap insurance; rebuilding the study is not.
The practical significance: the same 95%/95% claim that costs 59 units as an attribute test can often be made with 20 to 30 measured units — provided the data is approximately normal and the process has real margin. That is where the efficiency comes from, and it is also why the variables path demands an honest capability estimate up front rather than a habit.
The normality obligation
Tolerance intervals of this form assume normality. You must check it, not assume it — a normal probability plot plus a formal test such as Anderson-Darling, with the result recorded in the report. If the data is not normal, your options are to transform it and justify the transformation, to fit an appropriate alternative distribution, or to fall back on distribution-free methods. Note that non-parametric approaches give up the efficiency advantage almost entirely: a distribution-free 95%/95% one-sided claim lands you back around 59 units.
Choosing confidence and reliability by risk
The values should not be uniform across a project. Tie them to the severity of the harm associated with the failure mode in your ISO 14971 risk analysis, and state the tiering in a procedure so that it is applied consistently rather than negotiated test by test.
| Severity of associated harm | Typical confidence / reliability |
|---|---|
| Negligible / minor | 90% / 90% |
| Serious | 95% / 95% |
| Critical / life-threatening | 95% / 99% or higher |
This is a framework, not a rule handed down by any standard — you own the tiering and must defend it. Its value is that it converts an arbitrary-looking number into a traceable decision: this failure mode has this severity, therefore this reliability, therefore this sample size.
Sample size is not the whole sampling plan
A correct n drawn badly proves very little. Three considerations decide whether your sample actually represents the population you are claiming about:
- Lots and batches. Sixty units from a single lot characterise that lot. If lot-to-lot variation is a real source of variability — and for moulded, extruded, or coated components it usually is — draw across at least three lots.
- Sources of variation. Deliberately span the variation you expect in production: multiple operators, tools or cavities, machines, and where relevant environmental conditions. A sample that accidentally holds everything constant will understate true variability and overstate your capability.
- Worst case. Verification frequently needs to demonstrate performance at the boundaries — tolerance stack extremes, aged or sterilised product, conditioned samples. Define worst case explicitly and justify it.
Qualify the measurement system first
Every measured value contains both product variation and measurement variation. If the measurement system is noisy, part of what you record as product spread is instrument error — which inflates s, shrinks your tolerance interval margin, and can fail a perfectly capable design.
Run a measurement system analysis (Gage R&R) before the verification test, not after it fails. A common convention treats total gauge variation under 10% of the study variation as acceptable and 10–30% as conditionally acceptable depending on criticality. Discovering an inadequate gauge after building 60 units is an expensive way to learn this.
A worked example
Requirement: sterile barrier seal strength ≥ 1.5 N. Risk: breach compromises sterility; severity assessed as serious, so the procedure calls for 95% confidence and 95% reliability.
- Data type: the tensile fixture reports force in newtons — variable data. Record the values, not pass/fail.
- Size the study from prior margin. A 12-unit development build gave and (using the conservative upper estimate of spread). Expected margin standard deviations. Any clears that statistically.
- Let the binding constraint decide. The statistical minimum is roughly 6 units, but the sampling strategy needs 3 lots and 2 sealer machines, and the pilot estimate of is based on only 12 units. n = 30 — 10 units from each of three production lots, spanning both machines — satisfies both, with headroom if the real spread runs higher than the pilot suggested.
- Measurement system: Gage R&R completed on the tensile fixture in advance; result acceptable.
- Result: , . Normality confirmed by probability plot and Anderson-Darling.
- Analysis: with , . Lower tolerance limit 2.58 N.
- Conclusion: 2.58 N ≥ 1.5 N, so the requirement is met. We are 95% confident that at least 95% of the population exceeds 1.5 N.
Thirty units instead of fifty-nine, a stronger statement than "thirty units passed," and every step traceable to a decision recorded before the first sample was pulled. Note what the justification actually says: not "thirty is standard," but "six would satisfy the statistics given our margin; we used thirty because the sampling strategy and the uncertainty in our pilot estimate demanded it." That is a defensible answer to the only question that matters — why this number?
Common pitfalls
- Writing the justification afterwards. Pre-specify sample size, acceptance criteria, and analysis method in the protocol. Choosing the method once you have seen the data is not a defensible practice.
- Confusing a confidence interval with a tolerance interval. A confidence interval bounds the mean. Verification claims are about the population proportion. They are different calculations answering different questions.
- Discarding variable data. Recording pass/fail when the fixture produced a number is the single most common source of unnecessarily large studies.
- Sizing a variables study with no margin estimate. Picking n for measured data without any prior view of capability is guessing. If no pilot data exists, generate some — a small development build is far cheaper than a failed 30-unit verification.
- Using the optimistic estimate of spread. Sizing from the best-case you have ever seen produces a study that fails when normal variation shows up. Size from a conservative upper estimate.
- Silent normality assumptions. If the report does not show the normality check, an auditor is entitled to assume it was never done.
- One lot, one operator, one day. Cheap to execute, weak as evidence, and a predictable finding.
- Uniform 95/95 everywhere. Applying the same values to a cosmetic attribute and a life-supporting function shows the tiering was never actually risk-based.
What to document
A complete rationale — in the protocol, before execution — states the requirement being verified and its risk classification; the data type and why; the chosen confidence and reliability with justification; the statistical method and its source; for variables data, the prior margin or capability estimate used to size the study and where that estimate came from; the resulting sample size and the arithmetic; the sampling strategy across lots and sources of variation; the measurement system qualification status; and the pre-defined acceptance criteria and the plan if a failure occurs.
Where the sample size exceeds the statistical minimum — as it often will — say so and say why. "Statistically n = 6 suffices at our expected margin; we used n = 30 to span three lots and two machines" is a stronger rationale than a bare number, because it demonstrates you knew which constraint was binding.
Written that way, the sample size stops being a number someone has to defend under pressure and becomes the visible conclusion of a documented chain of reasoning.
This article is general information, not statistical or regulatory advice. Statistical methods must be selected and applied by qualified personnel with reference to validated tables, software, and applicable standards for your specific application.