Every laboratory information system sold in the last twenty years claims Westgard multi-rule quality control. Almost none of them will tell you which rules you should switch on, and the answer is emphatically not "all of them", for a reason that is arithmetic rather than opinion.
This is the article we wanted when we started implementing the rule evaluation in a laboratory system: what each rule is looking for, in plain terms, and what the consequences of running it are. Written by the LabFlow team, not by laboratory scientists, which matters for how you weigh the clinical parts and is why every rule below is traceable to a published source named at the foot of the page.
What quality control is actually checking#
A control is a specimen with a known value that you run alongside patient specimens. You already know roughly what it should say, so what it actually says tells you something about the measurement system rather than about the specimen.
That "roughly" is the whole subject. Repeat a measurement on identical material and you get a spread of answers, and that spread is characterised by two numbers your own laboratory establishes for your own instrument, method and control lot: a mean and a standard deviation. Both are yours. A mean printed on a manufacturer's insert is a starting point for a new lot and not a substitute for the twenty or so runs it takes to establish your own.
This is the first place implementations go wrong. If the mean and standard deviation come from the insert rather than from your own data, every rule below is being evaluated against somebody else's dispersion. The rules will still fire and the firing will mean nothing. Software should therefore record which lot a mean belongs to and where the figure came from, and refuse to evaluate against a lot that has no established statistics.
Given a mean and a standard deviation, a control result can be expressed as a distance from the mean measured in standard deviations. A result 2.3 standard deviations above the mean is written +2.3s, and every rule below is a statement about those distances. Nothing more complicated than that is happening.
The plot everyone draws#
The Levey-Jennings chart is the control values in run order, with the mean as a centre line and horizontal lines at one, two and three standard deviations either side. It is not a fancy visualisation. It exists because the two failure modes it separates look completely different on it, and look identical in a table of numbers.
Random error is one point a long way from the mean with its neighbours behaving normally. Systematic error is a run of points that are all slightly wrong in the same direction, none of which is dramatic on its own. Every rule below is designed to catch one or the other, and knowing which is which is what makes the rule names stop being noise.
Hide the plotted values
| Run | Value | Distance | Rule |
|---|---|---|---|
| 1 | 5.58 | -0.17s | — |
| 2 | 5.63 | +0.25s | — |
| 3 | 5.55 | -0.42s | — |
| 4 | 5.61 | +0.08s | — |
| 5 | 5.69 | +0.75s | — |
| 6 | 5.52 | -0.67s | — |
| 7 | 5.66 | +0.50s | — |
| 8 | 5.57 | -0.25s | — |
| 9 | 5.86 | +2.17s | 2-2s |
| 10 | 5.88 | +2.33s | 2-2s |
| 11 | 5.60 | +0.00s | — |
| 12 | 5.54 | -0.50s | — |
| 13 | 5.64 | +0.33s | — |
| 14 | 5.59 | -0.08s | — |
| 15 | 5.71 | +0.92s | — |
| 16 | 5.53 | -0.58s | — |
| 17 | 5.99 | +3.25s | 1-3s |
| 18 | 5.62 | +0.17s | — |
| 19 | 5.57 | -0.25s | — |
| 20 | 5.65 | +0.42s | — |
Two things are happening in that run. Runs 9 and 10 are both a little over two standard deviations high, neither dramatic, and together they are a systematic shift. Run 17 is on its own past three standard deviations, with its neighbours normal, and that is random error. A table of twenty numbers hides both; the plot shows them at a glance, which is why it has survived since 1950.
The six rules, one at a time#
The names encode the rule. The number before the dash is how many observations are involved, and what follows describes the limit. That is the entire mnemonic, and once you see it the names stop needing to be memorised.
1-2s — one observation past two standard deviations#
A single control result more than 2s from the mean. On its own this is a warning and not a rejection, and the next section is entirely about why. It exists to say "look at the other rules now", which is what makes the whole set a multi-rule procedure rather than six independent tests.
1-3s — one observation past three standard deviations#
A single control result more than 3s from the mean, which is a rejection. This is the classic random-error detector: a bubble, a short sample, a clot, a pipetting error on that one tube. Run 17 in the plot above is this rule firing.
2-2s — two in a row past two standard deviations, same side#
Two consecutive control results both past 2s on the same side of the mean. Rejection, and the classic systematic-error detector: a shift, a fresh calibration that landed slightly off, a reagent lot change, a drifting temperature. Runs 9 and 10 above are this rule firing.
The "same side" clause is load-bearing and is the most commonly mis-implemented part of the whole set. One control 2.1s high followed by one 2.1s low is not a shift, it is scatter, and treating it as a systematic error sends somebody looking for a cause that does not exist.
R-4s — a range of four standard deviations within a run#
Two controls in the same run, one at least 2s above the mean and one at least 2s below, so that the range between them exceeds 4s. Rejection, and it detects a loss of precision rather than a shift: the measurement has become imprecise enough that two specimens which should agree do not.
R-4s needs two control levels in the same run to be evaluable at all. A laboratory running a single level cannot use it, and a system that reports it as passing in that situation is reporting on a test it did not perform.
4-1s — four in a row past one standard deviation, same side#
Four consecutive results all more than 1s from the mean and all on the same side. Rejection. No individual point here looks remotely alarming, which is the point: this is a small systematic bias that a human eye scanning a chart will not see and that arithmetic finds immediately.
10x — ten in a row on the same side of the mean#
Ten consecutive results on one side of the mean, however close to it. Rejection, and the most sensitive systematic detector in the set. Ten coin tosses landing the same way is a one-in-512 event, so this is arithmetic about the mean having moved rather than about any individual result being wrong.
10x is also the rule most likely to be firing because the mean is stale. A control lot whose mean was established months ago against a since-recalibrated instrument produces exactly this signature. Before investigating the analyser, check when the statistics were last established.
Why 1-2s is a warning and not a rejection#
This is the single most consequential thing to understand about the set, and the thing most often got wrong in configuration.
If a measurement system is behaving perfectly, control results are distributed around the mean. In a normal distribution, about 95.45% of values fall within 2s, which leaves about 4.55% outside it. So roughly one control result in twenty-two will breach 2s while nothing whatsoever is wrong. That is not a failure of the system, it is what a 2s limit means.
Probability a single control breaches each limit, if nothing is wrong beyond 1s 31.7% about 1 run in 3 beyond 2s 4.55% about 1 run in 22 beyond 3s 0.27% about 1 run in 370 With TWO control levels per run, at least one breaching: beyond 2s 8.9% about 1 run in 11 beyond 3s 0.54% about 1 run in 185
Treat 1-2s as a rejection and you are rejecting nearly one run in eleven for no reason. That has three costs, and the third is the one that matters: repeated controls, delayed patient results, and a bench that learns within a fortnight that the rejection usually means nothing and starts repeating without investigating. A rule that cries wolf does not merely waste time, it trains people to ignore the rules that do not.
So 1-2s is a gate into the others. A 2s breach means evaluate 1-3s, 2-2s, R-4s, 4-1s and 10x now. If none of them fires, the run is in control and the 2s breach was the one in twenty-two.
The arithmetic of running everything#
Two numbers describe any quality control configuration. Error detection is the probability of rejecting a run in which something has genuinely gone wrong. False rejection is the probability of rejecting a run in which nothing has. Every rule you add improves the first and worsens the second, and the second is not free.
A laboratory running twenty analytes with two control levels, once a day, is performing forty control evaluations a day. At a per-run false rejection probability of even 1%, that is roughly one spurious rejection every two and a half days, each one costing a repeat, an investigation, and a delay to whatever patient results were riding on that run.
This is why the answer to "which rules should I run" is genuinely method-dependent. A method with plenty of analytical room between its own imprecision and the clinical tolerance for error needs very little: a single 1-3s rule can be sufficient. A method operating close to its clinical limit needs the full multi-rule set, and the false rejections are a price worth paying there. Choosing between those is a laboratory decision made against your own performance data, and no software vendor can make it for you.
| Rule | Detects | Observations | Verdict |
|---|---|---|---|
| 1-2s | Nothing on its own | 1 | Warning. A gate into the rest |
| 1-3s | Random error | 1 | Reject |
| 2-2s | Systematic error, shift | 2 consecutive, same side | Reject |
| R-4s | Random error, imprecision | 2 within a run, opposite sides | Reject. Needs two levels |
| 4-1s | Systematic error, small bias | 4 consecutive, same side | Reject |
| 10x | Systematic error, mean shift | 10 consecutive, same side | Reject |
Which rules to actually run#
The honest general answer is that this is your decision and it is made per method against your own imprecision and your own clinical tolerance. What can be said generally is narrower, and worth saying because it is where most configurations go wrong:
- 1-2s should be configured as a warning. If your system only offers reject or ignore, configuring it as reject is the worse of the two mistakes.
- R-4s cannot be evaluated on a single control level. A system that silently passes it in that case is reporting on a test it did not run.
- 4-1s and 10x need history, so they are meaningless on a lot with only a handful of runs behind it. Software should say so rather than evaluate them against four points.
- Rules configured per analyte beat rules configured globally, because analytical performance is a property of the method rather than of the laboratory.
- Whatever you choose, the choice belongs in your own documented quality control procedure, which is what an assessor asks to see.
What happens when a rule fires#
A rejection is a statement about the run, not about the control. That distinction is the entire clinical point, and it is where quality control stops being a filing exercise: the patient results produced alongside that control are not releasable, because the evidence that the measurement system was working is the thing that just failed.
Which produces the obligation nobody enjoys: results already released on a run later found out of control have to be reviewed, and where the difference is clinically material, whoever received them has to be told. This is the same shape as an amended result, and for the same reason. Somebody may have acted on it.
Overrides are legitimate and must be recorded. There are real situations where a run is released over a violation, by a named person, with a reason. The problem is never the override, it is the override with no name and no reason attached, which turns the gate into a formality and is exactly what an assessment is looking for.
What a laboratory system should do with all this#
Four things, and most systems do the first two.
The gap between the second and the fourth is the difference between a system that has a quality control screen and a system where quality control does something. It is worth asking a vendor to demonstrate the fourth rather than describe it, because both look identical in a slide.
Where LabFlow stops#
This is the section a vendor blog usually uses to reveal that everything above was setup. It is here for the opposite reason: this article would be dishonest without stating what the product it is published on actually does, including where it does nothing.
LabFlow evaluates the six rules above, names the rule that fired in the outcome, attaches the verdict to every patient result produced on that run, and refuses release when the run is out of control. An override is possible, is a recorded act with a person's name and a reason, and is in the audit log.
LabFlow evaluates the set you configure. It does not select one for you, and a vendor that shipped a default set would be making your validation decisions on your behalf.
Quality control is built and gates release. Calibration dates and results are recorded on the equipment register; nothing schedules a calibration and nothing sends a reminder, although the two are routinely described together.
LabFlow is not an instrument driver and ships no middleware. Control values arrive the same way patient results do, through whatever already talks to your analysers.
Internal quality control only. Comparing your performance against other laboratories using the same method is a different exercise with a different data source, and none of it is here.
LabFlow has no laboratories in production. The rule evaluation is tested against constructed cases, which is not the same as having run for a year on a real bench, and the difference is not a detail.
Nothing above is machine learning and nothing above is predictive. It is arithmetic against numbers that were already recorded, which is precisely why it can be trusted in a clinical setting and precisely why calling it artificial intelligence would mislead a reviewer about how a clinical decision was reached.
The published version said R-4s compares two consecutive results. It compares two control results within the same run, which is a different test and cannot be evaluated by a laboratory running one control level. The original sentence read: "two consecutive results whose range exceeds 4s". Reported by a reader; corrected in place rather than quietly replaced, for the same reason a released result is versioned rather than overwritten.