Quality control

Westgard rules, without the mnemonics.

Six rules that every laboratory system claims to implement, what each one is actually looking for, and the arithmetic that explains why switching all six on for every analyte makes your quality control worse rather than better.

Published28 Jul 2026
Corrected4 Aug 2026
Reading~11 min
AuthorThe LabFlow team
Quality controlWestgardLevey-Jennings

The article

Every laboratory information system sold in the last twenty years claims Westgard multi-rule quality control. Almost none of them will tell you which rules you should switch on, and the answer is emphatically not "all of them", for a reason that is arithmetic rather than opinion.

This is the article we wanted when we started implementing the rule evaluation in a laboratory system: what each rule is looking for, in plain terms, and what the consequences of running it are. Written by the LabFlow team, not by laboratory scientists, which matters for how you weigh the clinical parts and is why every rule below is traceable to a published source named at the foot of the page.

What quality control is actually checking#

A control is a specimen with a known value that you run alongside patient specimens. You already know roughly what it should say, so what it actually says tells you something about the measurement system rather than about the specimen.

That "roughly" is the whole subject. Repeat a measurement on identical material and you get a spread of answers, and that spread is characterised by two numbers your own laboratory establishes for your own instrument, method and control lot: a mean and a standard deviation. Both are yours. A mean printed on a manufacturer's insert is a starting point for a new lot and not a substitute for the twenty or so runs it takes to establish your own.

This is the first place implementations go wrong. If the mean and standard deviation come from the insert rather than from your own data, every rule below is being evaluated against somebody else's dispersion. The rules will still fire and the firing will mean nothing. Software should therefore record which lot a mean belongs to and where the figure came from, and refuse to evaluate against a lot that has no established statistics.

Given a mean and a standard deviation, a control result can be expressed as a distance from the mean measured in standard deviations. A result 2.3 standard deviations above the mean is written +2.3s, and every rule below is a statement about those distances. Nothing more complicated than that is happening.

The plot everyone draws#

The Levey-Jennings chart is the control values in run order, with the mean as a centre line and horizontal lines at one, two and three standard deviations either side. It is not a fancy visualisation. It exists because the two failure modes it separates look completely different on it, and look identical in a table of numbers.

Random error is one point a long way from the mean with its neighbours behaving normally. Systematic error is a run of points that are all slightly wrong in the same direction, none of which is dramatic on its own. Every rule below is designed to catch one or the other, and knowing which is which is what makes the rule names stop being noise.

Glucose, control level 2, twenty consecutive runs
mmol/L · mean 5.60 · SD 0.12
+3s+2s+1smean-1s-2s-3s
Mean±2s±3sIn controlRule violation
Hide the plotted values
Twenty consecutive control results with their distance from the mean in standard deviations, and any rule violated.
RunValueDistanceRule
15.58-0.17s—
25.63+0.25s—
35.55-0.42s—
45.61+0.08s—
55.69+0.75s—
65.52-0.67s—
75.66+0.50s—
85.57-0.25s—
95.86+2.17s2-2s
105.88+2.33s2-2s
115.60+0.00s—
125.54-0.50s—
135.64+0.33s—
145.59-0.08s—
155.71+0.92s—
165.53-0.58s—
175.99+3.25s1-3s
185.62+0.17s—
195.57-0.25s—
205.65+0.42s—

Two things are happening in that run. Runs 9 and 10 are both a little over two standard deviations high, neither dramatic, and together they are a systematic shift. Run 17 is on its own past three standard deviations, with its neighbours normal, and that is random error. A table of twenty numbers hides both; the plot shows them at a glance, which is why it has survived since 1950.

The six rules, one at a time#

The names encode the rule. The number before the dash is how many observations are involved, and what follows describes the limit. That is the entire mnemonic, and once you see it the names stop needing to be memorised.

1-2s warn1-3s reject2-2s rejectR-4s4-1s10x

1-2s — one observation past two standard deviations#

A single control result more than 2s from the mean. On its own this is a warning and not a rejection, and the next section is entirely about why. It exists to say "look at the other rules now", which is what makes the whole set a multi-rule procedure rather than six independent tests.

1-3s — one observation past three standard deviations#

A single control result more than 3s from the mean, which is a rejection. This is the classic random-error detector: a bubble, a short sample, a clot, a pipetting error on that one tube. Run 17 in the plot above is this rule firing.

2-2s — two in a row past two standard deviations, same side#

Two consecutive control results both past 2s on the same side of the mean. Rejection, and the classic systematic-error detector: a shift, a fresh calibration that landed slightly off, a reagent lot change, a drifting temperature. Runs 9 and 10 above are this rule firing.

The "same side" clause is load-bearing and is the most commonly mis-implemented part of the whole set. One control 2.1s high followed by one 2.1s low is not a shift, it is scatter, and treating it as a systematic error sends somebody looking for a cause that does not exist.

R-4s — a range of four standard deviations within a run#

Two controls in the same run, one at least 2s above the mean and one at least 2s below, so that the range between them exceeds 4s. Rejection, and it detects a loss of precision rather than a shift: the measurement has become imprecise enough that two specimens which should agree do not.

R-4s needs two control levels in the same run to be evaluable at all. A laboratory running a single level cannot use it, and a system that reports it as passing in that situation is reporting on a test it did not perform.

4-1s — four in a row past one standard deviation, same side#

Four consecutive results all more than 1s from the mean and all on the same side. Rejection. No individual point here looks remotely alarming, which is the point: this is a small systematic bias that a human eye scanning a chart will not see and that arithmetic finds immediately.

10x — ten in a row on the same side of the mean#

Ten consecutive results on one side of the mean, however close to it. Rejection, and the most sensitive systematic detector in the set. Ten coin tosses landing the same way is a one-in-512 event, so this is arithmetic about the mean having moved rather than about any individual result being wrong.

10x is also the rule most likely to be firing because the mean is stale. A control lot whose mean was established months ago against a since-recalibrated instrument produces exactly this signature. Before investigating the analyser, check when the statistics were last established.

Why 1-2s is a warning and not a rejection#

This is the single most consequential thing to understand about the set, and the thing most often got wrong in configuration.

If a measurement system is behaving perfectly, control results are distributed around the mean. In a normal distribution, about 95.45% of values fall within 2s, which leaves about 4.55% outside it. So roughly one control result in twenty-two will breach 2s while nothing whatsoever is wrong. That is not a failure of the system, it is what a 2s limit means.

Probability a single control breaches each limit, if nothing is wrong

  beyond 1s     31.7%       about 1 run in 3
  beyond 2s      4.55%      about 1 run in 22
  beyond 3s      0.27%      about 1 run in 370

With TWO control levels per run, at least one breaching:

  beyond 2s      8.9%       about 1 run in 11
  beyond 3s      0.54%      about 1 run in 185

Treat 1-2s as a rejection and you are rejecting nearly one run in eleven for no reason. That has three costs, and the third is the one that matters: repeated controls, delayed patient results, and a bench that learns within a fortnight that the rejection usually means nothing and starts repeating without investigating. A rule that cries wolf does not merely waste time, it trains people to ignore the rules that do not.

So 1-2s is a gate into the others. A 2s breach means evaluate 1-3s, 2-2s, R-4s, 4-1s and 10x now. If none of them fires, the run is in control and the 2s breach was the one in twenty-two.

The arithmetic of running everything#

Two numbers describe any quality control configuration. Error detection is the probability of rejecting a run in which something has genuinely gone wrong. False rejection is the probability of rejecting a run in which nothing has. Every rule you add improves the first and worsens the second, and the second is not free.

A laboratory running twenty analytes with two control levels, once a day, is performing forty control evaluations a day. At a per-run false rejection probability of even 1%, that is roughly one spurious rejection every two and a half days, each one costing a repeat, an investigation, and a delay to whatever patient results were riding on that run.

This is why the answer to "which rules should I run" is genuinely method-dependent. A method with plenty of analytical room between its own imprecision and the clinical tolerance for error needs very little: a single 1-3s rule can be sufficient. A method operating close to its clinical limit needs the full multi-rule set, and the false rejections are a price worth paying there. Choosing between those is a laboratory decision made against your own performance data, and no software vendor can make it for you.

What each rule is for, and when it is worth its false-rejection cost.
RuleDetectsObservationsVerdict
1-2sNothing on its own1Warning. A gate into the rest
1-3sRandom error1Reject
2-2sSystematic error, shift2 consecutive, same sideReject
R-4sRandom error, imprecision2 within a run, opposite sidesReject. Needs two levels
4-1sSystematic error, small bias4 consecutive, same sideReject
10xSystematic error, mean shift10 consecutive, same sideReject

Which rules to actually run#

The honest general answer is that this is your decision and it is made per method against your own imprecision and your own clinical tolerance. What can be said generally is narrower, and worth saying because it is where most configurations go wrong:

  • 1-2s should be configured as a warning. If your system only offers reject or ignore, configuring it as reject is the worse of the two mistakes.
  • R-4s cannot be evaluated on a single control level. A system that silently passes it in that case is reporting on a test it did not run.
  • 4-1s and 10x need history, so they are meaningless on a lot with only a handful of runs behind it. Software should say so rather than evaluate them against four points.
  • Rules configured per analyte beat rules configured globally, because analytical performance is a property of the method rather than of the laboratory.
  • Whatever you choose, the choice belongs in your own documented quality control procedure, which is what an assessor asks to see.

What happens when a rule fires#

A rejection is a statement about the run, not about the control. That distinction is the entire clinical point, and it is where quality control stops being a filing exercise: the patient results produced alongside that control are not releasable, because the evidence that the measurement system was working is the thing that just failed.

Which produces the obligation nobody enjoys: results already released on a run later found out of control have to be reviewed, and where the difference is clinically material, whoever received them has to be told. This is the same shape as an amended result, and for the same reason. Somebody may have acted on it.

Overrides are legitimate and must be recorded. There are real situations where a run is released over a violation, by a named person, with a reason. The problem is never the override, it is the override with no name and no reason attached, which turns the gate into a formality and is exactly what an assessment is looking for.

What a laboratory system should do with all this#

Four things, and most systems do the first two.

Record the control resulteveryone does thisLevel, lot, value, run, operator, and the mean and standard deviation in force at the time rather than the ones in force today.
Evaluate the configured rulesmost do thisAnd name the rule in the outcome. "Quality control failed" is not an outcome, it is a shrug. "1-3s, run 17, +3.25s" is something a person can act on.
Bind the verdict to the patient resultsthe one that mattersEvery patient result produced on that run carries the control verdict that was in force. Six months later, a question about one result can reach the control run behind it.
Refuse the releasethe one that is usually missingOut of control means release is refused, with the rule quoted. If the system records the violation and still lets the result go out with one click, the quality control was documentation.

The gap between the second and the fourth is the difference between a system that has a quality control screen and a system where quality control does something. It is worth asking a vendor to demonstrate the fourth rather than describe it, because both look identical in a slide.

Where LabFlow stops#

This is the section a vendor blog usually uses to reveal that everything above was setup. It is here for the opposite reason: this article would be dishonest without stating what the product it is published on actually does, including where it does nothing.

LabFlow evaluates the six rules above, names the rule that fired in the outcome, attaches the verdict to every patient result produced on that run, and refuses release when the run is out of control. An override is possible, is a recorded act with a person's name and a reason, and is in the audit log.

Choosing which rules apply to which analyteYours

LabFlow evaluates the set you configure. It does not select one for you, and a vendor that shipped a default set would be making your validation decisions on your behalf.

Calibration scheduling and remindersRecorded, not scheduled

Quality control is built and gates release. Calibration dates and results are recorded on the equipment register; nothing schedules a calibration and nothing sends a reminder, although the two are routinely described together.

Reading control results off the analyserNot built

LabFlow is not an instrument driver and ships no middleware. Control values arrive the same way patient results do, through whatever already talks to your analysers.

Peer-group or external quality assessment comparisonNot built

Internal quality control only. Comparing your performance against other laboratories using the same method is a different exercise with a different data source, and none of it is here.

Any of this validated in a live laboratoryNever

LabFlow has no laboratories in production. The rule evaluation is tested against constructed cases, which is not the same as having run for a year on a real bench, and the difference is not a detail.

Nothing above is machine learning and nothing above is predictive. It is arithmetic against numbers that were already recorded, which is precisely why it can be trusted in a clinical setting and precisely why calling it artificial intelligence would mislead a reviewer about how a clinical decision was reached.


Corrected 4 August 2026

The published version said R-4s compares two consecutive results. It compares two control results within the same run, which is a different test and cannot be evaluated by a laboratory running one control level. The original sentence read: "two consecutive results whose range exceeds 4s". Reported by a reader; corrected in place rather than quietly replaced, for the same reason a released result is versioned rather than overwritten.

Where to read the real thing.

This article is the LabFlow team's summary. None of the rules originate here and none of the clinical judgement about applying them does either, so the sources are named rather than absorbed into unattributed facts.

  • Westgard's own published description of the multirule procedure, which is where the rules and the error-detection argument come from. The definitions above follow it; any disagreement between the two should be resolved in its favour.
  • CLSI C24, the guideline document on statistical quality control for quantitative measurement procedures. This is the one your assessor will expect your written procedure to be built against, and it covers the parts this article deliberately does not, including how often to run controls and how to establish limits.
  • Your own accrediting body's requirements, which are the ones that actually bind you and which differ by country and by scope.
  • The normal distribution, for the percentages in the arithmetic section. Those are computed rather than cited, so they are checkable rather than quotable, which is the honest form for a number in an article like this.

Not clinical advice, and not a substitute for a validated procedure. No laboratory scientist has reviewed this article. It is software writing about a clinical domain, which is a genuinely different thing, and the distinction is kept explicit rather than blurred by confident phrasing.

Written by the team that implemented it.

The LabFlow team, who build LabFlow and wrote the rule evaluation this article describes. Not laboratory scientists: the domain here was read rather than practised, from standards documents, published rule sets and the source of the system being replaced.

That is enough to build software carefully and it is not the same as having run a bench for ten years. Where this article states something as fact it is traceable to a source above or to code; where it is inference, the sentence says so.

If something here is wrong, say so. It gets corrected in place with a dated note and the original quoted, as the correction above shows, and credited unless you ask otherwise.

This article

Published2026-07-28
Corrected2026-08-04
AuthorThe LabFlow team
Clinical reviewNone
SponsoredNo. There is nobody to sponsor it

More on who is behind LabFlow, and why there is no team page, is on the about page.

14 Jul 2026

HL7 v2 or FHIR: which one to build first

The honest answer is v2, and the reason is not technical merit. Includes the boundary between a message format and the network process that carries it.

About 9 minutes.

16 Jun 2026

Tenant isolation in a multi-tenant LIMS

Why a list query has to prove its own tenant from its own filters, and why a policy reading a field the query never filters on leaks with a clean 200.

About 12 minutes.

Ask a vendor to demonstrate the fourth thing.

Recording control results, evaluating rules and drawing the chart are table stakes and every product does them. Refusing a release because the run that produced the result was out of control is the one that changes what the software is for, and it is the one that looks identical to the others in a demonstration script.

Ask to see a release attempted and refused, with the rule quoted. Then ask what an override looks like and whose name ends up on it.

Found something wrong in this article? Say so. It is corrected in place with a dated note and the original quoted, exactly as the correction above was.