Services Combination Products Approach Industries Insights About Book a consultation

Prefer email? info@dpdynamicsol.com

Statistics

Sample Size Justification for Design Verification

"We tested thirty units" is a sample size. It is not a justification. The question a reviewer or auditor is really asking is: what claim are you making about the population, and with what confidence? Answer that, and the number falls out of the arithmetic instead of out of habit.

Key takeaways

  • Sample size is an output of your claim, not an input you pick first.
  • Attribute (pass/fail) testing is statistically expensive — 95%/95% with zero failures requires 59 units.
  • Measuring a variable instead of scoring pass/fail typically cuts that sample size by half or more.
  • Variables sample size is driven by margin — how many standard deviations separate your process from the limit. You need a capability estimate to size the study.
  • Below roughly 0.55 Cpk, no sample size will pass a 95%/95% claim. That is a design problem, not a sampling problem.
  • Confidence and reliability should be tiered by risk, and traceable to your ISO 14971 analysis.
  • Write the rationale before the test. A justification produced after the data exists is worth very little.

Why "n=30" is not a justification

Thirty is a folk number. It comes from a rule of thumb about when a sampling distribution of means becomes roughly normal — a statement about estimating an average, which is almost never what a design verification test is about. Verification usually asserts something much stronger: that a specified proportion of the population conforms, with stated confidence.

The distinction matters in an audit. "Thirty is standard practice" invites the follow-up question that has no good answer: standard for what reliability, at what confidence, for a failure mode with what severity? A justification that cannot survive that question is a finding waiting to happen.

Start with the claim, not the number

Every defensible sample size begins with a sentence in this shape:

"We are [X]% confident that at least [Y]% of the population meets [the specified requirement]."

X is confidence — how sure you are about the estimate, driven by sample size. Y is reliability — the proportion of the population you are claiming conforms. Both are choices you must make deliberately and defend, and together with the data type they determine the sample size completely.

Confidence and reliability are frequently confused because both are expressed as percentages and both are commonly set to 95. They are not the same thing. Increasing confidence means being more certain about the same claim; increasing reliability means making a stronger claim. Reliability is by far the more expensive of the two.

The biggest lever: attribute versus variable data

Before any arithmetic, ask what kind of data the test actually produces.

  • Attribute data is categorical — pass/fail, leak/no leak, present/absent. Each unit contributes one bit of information.
  • Variable data is measured on a continuous scale — force in newtons, pressure in kPa, dimension in millimetres. Each unit contributes a value and its distance from the specification limit.

That extra information is worth a great deal. A unit that passes tells you only that it passed. A unit that measured 42 N against a 20 N minimum tells you it passed with enormous margin — evidence that the population is comfortably clear of the limit.

The most common expensive mistake in device verification is collecting variable data and then throwing the numbers away by recording only pass or fail. If the fixture reads a value, capture the value. It routinely turns a 59-unit test into a 20–30 unit test.

Attribute path: the success-run theorem

When the test yields pass/fail and you plan to accept on zero failures, the required sample size is:

n = ln (1C) ln (R)

where C is confidence and R is reliability, both expressed as decimals, and n is rounded up to the next whole unit.

Some common results, all assuming zero failures observed:

Confidence / ReliabilityRequired n (zero failures)Typical use
90% / 90%22Low-risk attributes, early development
95% / 90%29Moderate risk
90% / 95%45Moderate risk, stronger claim
95% / 95%59The common default for verification
95% / 99%299High-severity failure modes
99% / 99%459Life-supporting / life-sustaining functions

Two consequences worth internalising. First, the jump from 95% to 99% reliability is a fivefold increase in units — reliability is expensive, so choose it on the basis of risk rather than instinct. Second, this table assumes zero failures. A single failure does not simply cost you one unit; it invalidates the zero-failure basis, and re-establishing the same claim while allowing one failure requires a substantially larger sample. Plan for what happens if a unit fails before you start.

Variables path: tolerance intervals

With measured data you can use a one-sided normal tolerance interval, which asks whether the whole distribution — not just the units you tested — sits clear of the specification limit. For a lower specification limit, the test is:

x¯ ks LSL

where x¯ is the sample mean, s the sample standard deviation, k a factor determined by sample size, confidence, and reliability, and LSL the lower specification limit.

The k factor shrinks as n grows — more data, less penalty for uncertainty. But notice the problem: this is an acceptance test. You cannot apply it until the data exists, so on its own it tells you nothing about how many units to build.

Closing the loop: how to choose n before you have data

Rearrange the acceptance condition and the answer appears. Dividing through by s:

x¯LSL s k(n,C,R)

The left side is your margin — how many standard deviations separate the process from the specification limit. The right side depends only on sample size, confidence, and reliability.

That reframes the whole question. Sample size for variables data is driven by how much margin your process has. A capable process needs few units; a marginal process needs many; and a process without enough margin cannot pass at any sample size.

The margin term is directly related to process capability. For a one-sided lower limit, Cpk is that same distance expressed in units of 3s, so the margin equals 3Cpk. That gives a table you can size a study from:

Sample size (n)k (95% / 95%, one-sided)Margin required≈ Cpk required
10≈ 2.912.91 s0.97
15≈ 2.572.57 s0.86
20≈ 2.402.40 s0.80
30≈ 2.222.22 s0.74
50≈ 2.072.07 s0.69
100≈ 1.931.93 s0.64
∞ (theoretical floor)1.6451.65 s0.55

Values are illustrative — take exact factors from a validated table or qualified statistical software, and cite the source in your protocol.

The procedure is then three steps:

  1. Estimate your margin from prior data. Use pilot builds, development lots, a similar predicate process, or a design target. Compute the expected margin as (x¯LSL)/s. Be conservative: use the upper plausible value of s, not the best case.
  2. Read off the smallest n whose k is below that margin. If your expected margin is 2.5 standard deviations, n = 20 (k ≈ 2.40) clears it; n = 15 (k ≈ 2.57) does not.
  3. Add headroom. Your estimate of s is itself uncertain, and if the real spread comes in higher you fail the study after building all the units. Stepping up one row is cheap insurance; rebuilding the study is not.
Read the last row of that table carefully. Even with infinite units, a 95%/95% claim requires roughly 1.65 standard deviations of margin — about Cpk=0.55. Below that, no sample size will ever pass. If your sizing calculation demands hundreds of units, the statistics are not being difficult; they are telling you the design or the process lacks margin. Fix that instead of buying more units.

The practical significance: the same 95%/95% claim that costs 59 units as an attribute test can often be made with 20 to 30 measured units — provided the data is approximately normal and the process has real margin. That is where the efficiency comes from, and it is also why the variables path demands an honest capability estimate up front rather than a habit.

The normality obligation

Tolerance intervals of this form assume normality. You must check it, not assume it — a normal probability plot plus a formal test such as Anderson-Darling, with the result recorded in the report. If the data is not normal, your options are to transform it and justify the transformation, to fit an appropriate alternative distribution, or to fall back on distribution-free methods. Note that non-parametric approaches give up the efficiency advantage almost entirely: a distribution-free 95%/95% one-sided claim lands you back around 59 units.

Choosing confidence and reliability by risk

The values should not be uniform across a project. Tie them to the severity of the harm associated with the failure mode in your ISO 14971 risk analysis, and state the tiering in a procedure so that it is applied consistently rather than negotiated test by test.

Severity of associated harmTypical confidence / reliability
Negligible / minor90% / 90%
Serious95% / 95%
Critical / life-threatening95% / 99% or higher

This is a framework, not a rule handed down by any standard — you own the tiering and must defend it. Its value is that it converts an arbitrary-looking number into a traceable decision: this failure mode has this severity, therefore this reliability, therefore this sample size.

Sample size is not the whole sampling plan

A correct n drawn badly proves very little. Three considerations decide whether your sample actually represents the population you are claiming about:

  • Lots and batches. Sixty units from a single lot characterise that lot. If lot-to-lot variation is a real source of variability — and for moulded, extruded, or coated components it usually is — draw across at least three lots.
  • Sources of variation. Deliberately span the variation you expect in production: multiple operators, tools or cavities, machines, and where relevant environmental conditions. A sample that accidentally holds everything constant will understate true variability and overstate your capability.
  • Worst case. Verification frequently needs to demonstrate performance at the boundaries — tolerance stack extremes, aged or sterilised product, conditioned samples. Define worst case explicitly and justify it.

Qualify the measurement system first

Every measured value contains both product variation and measurement variation. If the measurement system is noisy, part of what you record as product spread is instrument error — which inflates s, shrinks your tolerance interval margin, and can fail a perfectly capable design.

Run a measurement system analysis (Gage R&R) before the verification test, not after it fails. A common convention treats total gauge variation under 10% of the study variation as acceptable and 10–30% as conditionally acceptable depending on criticality. Discovering an inadequate gauge after building 60 units is an expensive way to learn this.

A worked example

Requirement: sterile barrier seal strength ≥ 1.5 N. Risk: breach compromises sterility; severity assessed as serious, so the procedure calls for 95% confidence and 95% reliability.

  1. Data type: the tensile fixture reports force in newtons — variable data. Record the values, not pass/fail.
  2. Size the study from prior margin. A 12-unit development build gave x¯3.3 N and s0.45 N (using the conservative upper estimate of spread). Expected margin =(3.31.5)/0.454.0 standard deviations. Any n6 clears that statistically.
  3. Let the binding constraint decide. The statistical minimum is roughly 6 units, but the sampling strategy needs 3 lots and 2 sealer machines, and the pilot estimate of s is based on only 12 units. n = 30 — 10 units from each of three production lots, spanning both machines — satisfies both, with headroom if the real spread runs higher than the pilot suggested.
  4. Measurement system: Gage R&R completed on the tensile fixture in advance; result acceptable.
  5. Result: x¯=3.42 N, s=0.38 N. Normality confirmed by probability plot and Anderson-Darling.
  6. Analysis: with n=30, k2.22. Lower tolerance limit =3.42(2.22×0.38)= 2.58 N.
  7. Conclusion: 2.58 N ≥ 1.5 N, so the requirement is met. We are 95% confident that at least 95% of the population exceeds 1.5 N.

Thirty units instead of fifty-nine, a stronger statement than "thirty units passed," and every step traceable to a decision recorded before the first sample was pulled. Note what the justification actually says: not "thirty is standard," but "six would satisfy the statistics given our margin; we used thirty because the sampling strategy and the uncertainty in our pilot estimate demanded it." That is a defensible answer to the only question that matters — why this number?

Common pitfalls

  • Writing the justification afterwards. Pre-specify sample size, acceptance criteria, and analysis method in the protocol. Choosing the method once you have seen the data is not a defensible practice.
  • Confusing a confidence interval with a tolerance interval. A confidence interval bounds the mean. Verification claims are about the population proportion. They are different calculations answering different questions.
  • Discarding variable data. Recording pass/fail when the fixture produced a number is the single most common source of unnecessarily large studies.
  • Sizing a variables study with no margin estimate. Picking n for measured data without any prior view of capability is guessing. If no pilot data exists, generate some — a small development build is far cheaper than a failed 30-unit verification.
  • Using the optimistic estimate of spread. Sizing from the best-case s you have ever seen produces a study that fails when normal variation shows up. Size from a conservative upper estimate.
  • Silent normality assumptions. If the report does not show the normality check, an auditor is entitled to assume it was never done.
  • One lot, one operator, one day. Cheap to execute, weak as evidence, and a predictable finding.
  • Uniform 95/95 everywhere. Applying the same values to a cosmetic attribute and a life-supporting function shows the tiering was never actually risk-based.

What to document

A complete rationale — in the protocol, before execution — states the requirement being verified and its risk classification; the data type and why; the chosen confidence and reliability with justification; the statistical method and its source; for variables data, the prior margin or capability estimate used to size the study and where that estimate came from; the resulting sample size and the arithmetic; the sampling strategy across lots and sources of variation; the measurement system qualification status; and the pre-defined acceptance criteria and the plan if a failure occurs.

Where the sample size exceeds the statistical minimum — as it often will — say so and say why. "Statistically n = 6 suffices at our expected margin; we used n = 30 to span three lots and two machines" is a stronger rationale than a bare number, because it demonstrates you knew which constraint was binding.

Written that way, the sample size stops being a number someone has to defend under pressure and becomes the visible conclusion of a documented chain of reasoning.

This article is general information, not statistical or regulatory advice. Statistical methods must be selected and applied by qualified personnel with reference to validated tables, software, and applicable standards for your specific application.

David Plescia, Founder & Principal Consultant, DP Dynamic Solutions David Plescia Founder & Principal Consultant, DP Dynamic Solutions · 25+ years in medtech quality

DP Dynamic Solutions provides statistical support for medical device and combination product programmes — sample size justification, statistical analysis plans, process capability, measurement system analysis, and stability analysis.

Statistics services
Let's talk

Need a sample size you can actually defend?

We write statistical analysis plans and sample size rationales that hold up in audits and submissions — tied to your risk analysis, not to habit.