The honest answer is v2, and the reason is not technical merit. It is that the systems you have to talk to this year speak v2, and the pipe that carries it is a separate problem from the message inside it.
Every laboratory integration conversation starts in the same place. Somebody asks whether to build HL7 v2 or FHIR, the room splits along roughly generational lines, and an hour later nobody has written anything down. The honest answer is v2 first, and the reason is not that v2 is better. It is that the systems you have to talk to this year speak it.
This is written from the implementer’s side rather than the standards side. We have built the message layer for a laboratory system and the parts that cost time were not the parts the specifications warn you about.
The question people think they are asking#
"HL7 v2 or FHIR" sounds like a choice between two ways of representing a laboratory result. It is not. It is a choice about which systems you can exchange results with on the day you go live, and that is a question about your customers’ estate rather than about either specification.
A hospital that bought its electronic health record in 2011 has a v2 interface engine, a team that knows it, and a change-control process for touching it. It may also have a FHIR endpoint, usually newer, often read-only, and frequently scoped to whatever a national programme required rather than to what you need. Asking that hospital to receive results over FHIR is asking for a project. Asking it to receive an ORU^R01 is asking for a configuration change.
The test is not which standard is better. It is which one lets a real laboratory send a real result to a real clinician in the first quarter. On today’s estate that is v2, and it will stay v2 for longer than anybody writing about interoperability finds comfortable.
What v2 actually is#
HL7 v2 is a pipe-delimited message format with segments, fields, components and subcomponents, and it is genuinely old. The parts that matter for a laboratory are few. An ORM or OML carries an order. An ORU carries a result. An ADT carries the patient demographics everything else hangs off. A result message is a patient identifier, an order identifier, one OBR describing the requested battery, and one OBX per analyte with its value, units, reference range, abnormal flag and result status.
That last list is the part worth noticing, because it is the same list a result row needs anyway. A system whose internal result model already carries the resolved reference interval and the abnormal flag has almost nothing to compute at message time. A system that stores a bare number and looks the range up on the way out has a harder job and a subtler bug: the range it looks up today is not necessarily the range that applied when the value was produced.
The awkwardness in v2 is not the format. It is the optionality. Almost every field is optional in the specification and mandatory in practice, and which ones are mandatory in practice differs per receiving system. This is what conformance profiles exist for, and it is why an integration is never finished when the parser is.
What FHIR is genuinely better at#
FHIR is a resource model with an HTTP API, and it is much better than v2 at three things a laboratory eventually needs.
- Being read. A clinician’s application asking "what were this patient’s last six haemoglobins" is a query, and v2 has no good answer to a query. FHIR’s Observation search is exactly that shape.
- Being described. A FHIR resource carries its own structure, so a receiver can validate what it was sent without a conformance document negotiated by email.
- Being extended honestly. Extensions are declared and typed rather than smuggled into a Z-segment that only two systems understand.
None of those three is the job on day one. On day one the job is to push a result into a system that is waiting for one, and push is where v2 has twenty-five years of installed advantage.
The pipe is not the message, and this is where estimates go wrong#
The single most expensive misunderstanding in laboratory integration is treating "we support HL7 v2" as one piece of work. It is two, and they have almost nothing in common.
Building and parsing the message is application work. It is a well-bounded problem with a testable answer: given this order and this result, emit this message; given this message, produce this order. It is the part a laboratory system should own, and it is the part that benefits from having a clean internal model.
Carrying the message is infrastructure. In practice that means MLLP, a minimal framing protocol over a raw TCP socket, usually inside a private network or a VPN, usually on a port a hospital network team has to open, usually terminating on a host somebody has to run, patch and monitor. It has nothing to do with laboratory logic and everything to do with somebody being on call when the socket drops at three in the morning.
These two are estimated together and delivered separately. A quote that says "HL7 v2 support" and means the first is honest work. A customer who hears it and means the second has bought an on-call rota they did not know about.
FHIR does not have this problem, because its transport is HTTP and everybody already runs HTTP. That is a real advantage and it is worth saying plainly: the reason FHIR feels easier is mostly that its pipe is one you already own.
Acknowledgements, and what "sent" means#
A v2 exchange has an acknowledgement message, and it is worth understanding what it does and does not tell you. An application acknowledgement says the receiving application accepted the message. A commit acknowledgement says it reached a queue. Neither says a clinician read the result, and a laboratory that reports "delivered" from an ACK is reporting a fact about a socket.
This matters for critical values in particular. A critical result is not discharged by a message being accepted. It is discharged by a person being told, and by that conversation being recorded with a time, a name and what was read back. Any laboratory system that lets an interface acknowledgement close a critical-value escalation has confused two different kinds of delivery.
Identifiers are where it actually breaks#
Parsing is not the hard part. Agreeing on who the patient is, is the hard part.
A v2 message identifies a patient in PID-3, which is a repeating field of identifier-plus-assigning-authority pairs. The receiving system has its own identifier and its own authority, and the two agree only if somebody made them agree. In practice you will meet a hospital number, a national number, a laboratory number and an accession number in the same conversation, and at least one of them will be reused across sites.
The same is true of the order. An order placed in the ordering system has a placer order number; the laboratory assigns a filler order number; a result must carry both, because the sender needs to recognise its own request coming back. A laboratory system that treats its accession number as the universal key will work perfectly in testing and fail the first time a hospital sends the same placer number for two different requests.
- Store every identifier you are given, with the authority that issued it, rather than collapsing them into one column.
- Never derive one identifier from another. A check digit is not an authority.
- Decide explicitly what happens when a message arrives for a patient you do not have. Creating one silently is how two records for the same person appear.
None of this is standard-specific. FHIR has the same problem with better names for it.
How to test an interface you cannot reach#
You will be asked to demonstrate an integration long before anybody gives you access to the system you are integrating with. The workable answer is to make the message layer testable without a network at all.
Build the message from a result the way the application will in production, then assert the output against a fixed expected message held beside the test. That catches the class of error that matters — a field that moved, a flag that stopped being set, a range that stopped being resolved — and it catches it in a second rather than in a scheduled window with a hospital integration team watching.
Then keep a small set of real messages, with identifiers replaced, as parser fixtures. Real messages contain things no specification prepares you for: trailing empty fields, unexpected repetitions, a segment in an order nobody documented. A parser that has only ever seen messages you generated yourself is a parser that has only been tested against your own assumptions.
The message layer is testable; the transport is not. That asymmetry is another reason to keep them apart. You can prove the first is correct on a laptop. The second is proved by running it.
What to build, in what order#
For a laboratory system starting from nothing, the order that has held up in practice is this.
- An internal result model that already carries what a message needs — resolved reference interval, abnormal flag, result status, specimen identity, ordering provider. Do this first and both standards get cheaper.
- Outbound ORU. It is the message that makes you useful to a customer, and it is the one they are already able to receive.
- Inbound ORM or OML, once a customer wants to place orders electronically rather than on paper or through your own interface.
- FHIR Observation and DiagnosticReport, as a read surface. This is where a patient application, a portal, or a modern clinical system will want to talk to you, and it is a much nicer thing to expose than a v2 query.
- ADT last, and only if you are consuming patient demographics rather than owning them.
The ordering is not a maturity ladder. It is a sequence of things a customer will pay for, arranged so that each one is useful before the next one starts.
Where LabFlow stops#
LabFlow exports a released report as an HL7 v2 message and a FHIR R4 bundle, and imports an HL7 v2 order as a draft a person confirms. All of it moves as files, with no live feed.
The MLLP listener is not built. Running a socket on your infrastructure, inside your network, with your monitoring and your on-call, is a decision about your operations rather than a feature of a laboratory system, and shipping a listener would mean pretending otherwise. That boundary is stated on the home page under the limits rather than discovered during an implementation.
If you are scoping this work for your own system, the sentence worth writing into the plan is the one this article is built around: the message format and the process that carries it are two projects, and only one of them is about laboratories.