The Hard Parts of Converting HL7 v2 to FHIR
I've spent years dealing with healthcare data interoperability, and the thing nobody tells you about converting between these standards is that the mapping looks simple until you hit the edge cases. The basic segments map almost perfectly, but that's where most projects quietly fall apart. You start with something like an ADT^A01 admission message and think, okay, Patient maps to Patient, Encounter maps to Encounter. That works for the first hundred records. Then you hit a patient who was admitted in 2019 using local codes, and their NPI is wrong, and the pharmacy system they transferred to uses a completely different identifier system. That's when the real work begins. The mapping itself isn't the hard part. It's the semantic gap between how HL7 v2 describes clinical information and how FHIR structures it. HL7 v2 is message-oriented. FHIR is resource-oriented. Those are fundamentally different ways of thinking about the same data, and it shows up everywhere once you start mapping for real.
The Hl7 To Fhir Mapping Process
Here's what actually happens when you sit down to do this. First, you pull a sample message from the source system and parse it by segment. PID, PV1, NK1, OBX, AL1. You go through each one and look at what FHIR resource it corresponds to. Then you start matching fields. MSH-7 becomes the resource.meta.lastUpdated timestamp. PV1-2 becomes the encounter.class code. That part is documented pretty well in the official implementation guides. The second step is handling the valuesets. HL7 v2 uses its own coding systems throughout. PV1-3 has the attending physician, which is usually an HD data type with an extension and an identifier. In FHIR, that needs to become a reference to a Practitioner resource, or sometimes a PractitionerRole depending on what information you have. If the source system only stores the provider's name and not their NPI, you're building a lookup table, not doing a straightforward mapping. For medication orders, the ORC and OBR segments get mapped to a ServiceRequest resource in FHIR. But here's the issue that catches most people off guard. In HL7 v2, the order and the observation are separate segments that share a common ORC field. In FHIR, they might end up as two different resources—a ServiceRequest and a diagnostic Report—that reference each other. The relationship has to be explicit, and if your source data doesn't preserve that linkage cleanly, you lose it in translation.
A Real Problem I Ran Into
Last year I was mapping lab results from an HL7 v2 ORM^O01 order followed by an ORU^R01 result pair. The lab system had a quirk where it would send the order confirmation with an empty OBX segment and then resubmit the full result in a follow-up message. Standard point-to-point mapping tools flattened both messages into the same FHIR Observation resource, which created duplicate entries with different timestamps but identical codes and values. The consuming application saw two results instead of one and flagged a data quality issue that didn't actually exist. The workaround was to add a deduplication step using the OBX-3-1 (observation ID) as the key, then merge the order metadata from the first message with the actual result data from the second. It added maybe twenty minutes of processing time per batch, but it prevented the downstream system from generating false alerts for every single lab result in the feed. Without that step, you'd get about eight percent duplicate observations in a typical quarterly data migration.
Get the Full Details

Where The Mapping Breaks Down
Address fields are the easiest place to lose data. In HL7 v2, an XAD segment has separate fields for street address line one, street address line two, city, state, zip code, and country. When you map this to FHIR's Address datatype, lines one and two collapse into a single line field. Most people just paste the full string into Line[0] and call it done. If your downstream system does any kind of geocoding or mail merge, that merged address line will fail validation or produce incorrect results. I've seen it happen repeatedly. The fix is to split the raw XAD-3 field on the carriage return delimiter and assign each line separately. It takes an extra six lines of code but prevents address validation failures downstream. Time zone handling is another silent data loss point. HL7 v2 uses the TQ data type for timing, which can include time zone offset information in MSH-27 or as part of individual timestamp fields. FHIR has a dedicated time zone element built into the DateTime datatype. Most mapping tools don't extract this and just drop it. If you're dealing with multi-site health systems across different time zones, your timestamps will appear to be off by hours after conversion. This showed up in our audit logs where a medication administered at 0200 central time appeared as 0200 eastern time in the FHIR store, which made the medication administration review look like a dosing error. The bigger structural issue is that HL7 v2 message types encode context that FHIR resources don't. An ADT^A03 discharge message tells you exactly what happened and when. In FHIR, you represent this as an Encounter with an status of "finished." But the clinical narrative of why the patient was discharged, who authorized it, and what the discharge instructions were often lives in unstructured OBX text segments in the original message. When you map those to FHIR Narrative or Diagnosis resources, you're either losing the detail or storing it as free text, which defeats the purpose of structured data exchange.
What Actually Works For Large-Scale Mapping
Building custom mappings from scratch works for small projects. For anything over a thousand message types across multiple source systems, you need a more systematic approach. The HL7 International publishing of the v2 to FHIR mapping guide is the starting point, but it's incomplete by design. It covers the common message types and leaves the rest as implementation decisions. The tools that have worked for my team are either commercial ETL platforms with built-in FHIR connectors, or open source frameworks like HAPI FHIR that let you write Java-based transformations. The HAPI approach is more work upfront but gives you version control over your mappings, which matters when you need to update the mapping logic after a FHIR specification change. The commercial tools are faster to deploy but you're locked into their mapping rules, and you can't customize edge case handling without waiting for vendor support. For terminology services, you need to plan for that part separately. Your mapped data will reference LOINC codes for labs, SNOMED CT for conditions, and RxNorm for medications. If the source system uses local codes, you need a code mapping table or an API that resolves them. This is often where projects stall because nobody accounted for the terminology service setup before starting the actual message mapping.
Validation is non-negotiable. After your mapping runs, you need to validate the output FHIR resources against the relevant FHIR profiles using a tool like the FHIR Validator. This catches structural errors that the mapping logic itself won't reveal. A resource that passes your custom validation checks might still fail the formal FHIR conformance tests if you missed a required element or used the wrong cardinality.

The Limits You Should Accept Up Front
HL7 v2 to FHIR mapping will never be lossless. Some information simply doesn't translate between the two models. Custom HL7 segments have no FHIR equivalent. V2 observation relationships that depend on segment ordering become explicit FHIR references that may not exist in the source data. Historical patient records that predate standardized coding systems require manual code assignment that automation can't handle reliably. If you're working with a single well-behaved source system and a limited set of message types, you can probably get a working mapping in two to four weeks. A multi-source environment with dozens of message types and custom segments will take three to six months for a production-ready implementation, depending on how much manual cleanup your data requires. Budget accordingly. The alternative some teams pursue is moving directly to FHIR at the source system level instead of maintaining a mapping layer. That's the right answer when you have control over the source system and enough time to implement it properly. It's not always possible, which is why this mapping work still exists and why people still need to know how to do it.