Working with Duplicates
EAP

 
Applies to: all deduplication strategies. The operations on this page repair duplicates that already exist in the repository, regardless of which ingestion strategy is used.

FHIR Repositories can end up with duplicate resources in them. This page details tools that Smile provides to work with duplicated data.

Cause of DuplicatesRecommended Approach
The same entities were imported from different source systems. E.g. Jane Smith's Patient record is imported from the Lab system as Patient/123 and from the Pharmacy system as Patient/456.Preserve the source data as is so that future data continues to be properly associated to the right patients. Use MDM to establish Golden Resources and MDM LINK records outside of the data and use MDM features like Observation?patient=Patient/123&_mdm=true to link the data at search time.
Data was accidentally duplicated on import. E.g. the same data was accidentally loaded twice (as POST new resources) or Conditional Create directives failed to match the intended target correctly. This kind of unintended duplication can also occur when translating HL7v2, CDA, or CSV data into FHIR resources.In the case where the data was accidentally duplicated, it may make sense to "clean up" the duplicates. See below for details on Smile CDR tools to merge such duplicates.

Deduplication Operations
EAP

 

Smile CDR provides several operations to deduplicate data:

  • $merge which is a backport of the FHIR R5 Patient/$merge specification to FHIR R4
  • $hapi.fhir.merge which extends the merge functionality to all FHIR resource types that have an identifier element
  • $hapi.fhir.undo-merge undoes the effects of a previous merge operation
  • $hapi.fhir.replace-references performs only the update references part of this $merge operation
  • $hapi.fhir.undo-replace-references undoes the effects of a $hapi.fhir.replace-references operation
  • $sdh.mdm-bundle-match processes FHIR Bundles to match resources using MDM rules and optionally merge them using survivorship
  • $sdh.mdm-deduplicate starts a batch job that will run in the background and use the mdm rules to deduplicate resources

$merge Operation
EAP

 

See the FHIR R5 Patient/$merge specification page for a description of this operation. See the bottom of this page for details on the current roadmap for enhancing Smile CDR deduplication functionality.

$merge Input Parameters

NameTypeDefaultNotes
source-patient-identifierIdentifier List of source patient identifiers
source-patientReference Source patient
target-patient-identifierIdentifier List of target patient identifiers
target-patientReference Target patient
result-patientPatient Optional merged patient resource
previewBooleanfalseIf true, no changes will be made and response will summarize what would happen were the merge to occur
delete-sourceBooleanfalseIf true, delete the source resource
resource-limitInteger512If the request is synchrononous and the number of resources to change exceeds this threshold, the operation will fail with 412 Precondition Failed. This parameter has no effect if the Prefer: respond-async header is set

See the FHIR R5 Patient/$merge specification for a detailed description of these input parameters.

resource-limit is a Smile CDR addition to protect users from accidentally changing too many resources at once. If resource-limit is larger than 10,000, the value 10,000 will be used.

If you request that the operation be performed asynchronously by providing the Prefer: respond-async HTTP header, then the resource-limit parameter is ignored.

When performed asynchronously, the operation is performed in batches of 1024 resource patches at a time, via PATCH transaction Bundles.

$merge Output Parameters

NameTypeNotes
inputParametersA copy of the input parameters used in the $merge operation
outcomeOperationOutcomeDetails about the result of the merge
resultPatientThe merged Patient resource
taskTaskIf the merge operation was performed asynchronously, this Task resource provides details about the status of the merge operation

Merge Provenance Resource

With the 2025.08 release, the $merge operation creates a Provenance resource upon successful completion. This Provenance resource contains, in its target element, the versioned references to the target patient, the source patient (if not deleted during the operation), and all other resources updated as part of the operation. The Provenance.activity is set to http://terminology.hl7.org/CodeSystem/iso-21089-lifecycle|merge, and the Provenance.agent.who is populated with a logical reference to the request user.

Merge Example 1: Merge Patient/2 into Patient/3 synchronously.

Input:

POST /Patient/$merge
Content-Type: application/fhir+json

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "source-patient",
      "valueReference": {
        "reference": "Patient/2"
      }
    },
    {
      "name": "target-patient",
      "valueReference": {
        "reference": "Patient/3"
      }
    }
  ]
}

Output:

HTTP/1.1 200 OK

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "input",
      "resource": {
        "resourceType": "Parameters",
        "parameter": [
          {
            "name": "source-patient",
            "valueReference": {
              "reference": "Patient/2"
            }
          },
          {
            "name": "target-patient",
            "valueReference": {
              "reference": "Patient/3"
            }
          }
        ]
      }
    },
    {
      "name": "outcome",
      "resource": {
        "resourceType": "OperationOutcome",
        "issue": [
          {
            "severity": "information",
            "details": {
              "text": "Merge operation completed successfully."
            }
          }
        ]
      }
    },
    {
      "name": "result",
      "resource": {
        "resourceType": "Patient",
        "id": "3",
        "identifier": [
          {
            "system": "SYS2A",
            "value": "VAL2A"
          },
          {
            "system": "SYS2B",
            "value": "VAL2B"
          },
          {
            "system": "SYSC",
            "value": "VALC"
          },
          {
            "use": "old",
            "system": "SYS1A",
            "value": "VAL1A"
          },
          {
            "use": "old",
            "system": "SYS1B",
            "value": "VAL1B"
          }
        ],
        "link": [
          {
            "other": {
              "reference": "Patient/2"
            },
            "type": "replaces"
          }
        ]
      }
    }
  ]
}

Merge Example 2: Preview merge Patient/2 into Patient/3 synchronously.

Input:

POST /Patient/$merge
Content-Type: application/fhir+json

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "source-patient",
      "valueReference": {
        "reference": "Patient/2"
      }
    },
    {
      "name": "target-patient",
      "valueReference": {
        "reference": "Patient/3"
      }
    },
    {
      "name": "preview",
      "valueBoolean": true
    }
  ]
}

Output:

HTTP/1.1 200 OK

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "input",
      "resource": {
        "resourceType": "Parameters",
        ...copy of input...
      }
    },
    {
      "name": "outcome",
      "resource": {
        "resourceType": "OperationOutcome",
        "issue": [
          {
            "severity": "information",
            "details": {
              "text": "Preview only merge operation - no issues detected"
            },
            "diagnostics": "Merge would update 25 resources"
          }
        ]
      }
    },
    {
      "name": "result",
      "resource": {
        "resourceType": "Patient",
        ...merged patient...
      }
    }
  ]
}

Merge Example 3: Merging Patients by identifier asynchronously.

Input:

POST /Patient/$merge
Content-Type: application/fhir+json
Prefer: respond-async

{
    "resourceType": "Parameters",
    "parameter": [ {
            "name": "source-patient-identifier",
            "valueIdentifier": {
                    "system" : "urn:oid:1.2.36.146.595.217.0.1",
                    "value" : "12345"
            }
        },
        {
            "name": "target-patient-identifier",
            "valueIdentifier": {
                    "system" : "urn:oid:1.2.36.146.595.217.0.1",
                    "value" : "12346"
            }
        }
    ]
}

Output:

HTTP/1.1 202 Accepted

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "input",
      "resource": {
        "resourceType": "Parameters",
        ...copy of input...
      }
    },
    {
      "name": "task",
      "resource": {
        "resourceType": "Task",
        "id": "352",
        "identifier": [
          {
            "system": "http://hapifhir.io/batch/jobId",
            "value": "26738f4d-266c-4ef6-934f-1d13b1b474b9"
          }
        ],
        "status": "in-progress"
      }
    }
  ]
}

You can poll the status of the returned Task resource to see when it completes. For more detailed status about the background job, you can view the status of the corresponding Smile CDR batch job either through the Web Admin Console, or through the Admin JSON API. The ID of the Smile CDR batch job is provided as an identifer on the returned Task.

Merge in Patient ID Partition Mode

In Patient ID Partition mode, a resource is stored in the partition of the Patient compartment it belongs to, and an update cannot change a resource's partition. A compartment resource therefore cannot simply be repointed at the target Patient; the $merge operation will:

  • Copy each resource in the source Patient's compartment into the target Patient's compartment. Each copy is assigned a new resource ID.
  • Rewrite every reference to an original compartment resource so it points to the new copy. This includes references between copied resources as well as references held by resources outside the compartment.
  • Redirect references to the source Patient so they point to the target Patient.
  • Delete the original compartment resources.

This is how every merge behaves in Patient ID Partition mode, including when the source and target Patients resolve to the same database partition.

In addition to the steps above, the source and target Patients themselves are updated as in an unpartitioned merge: copying identifiers from source to target, adding replaces / replaced-by links, optionally deleting the source via the delete-source parameter, etc.

Provenance

A merge in this mode records what it changed across several Provenance resources rather than a single one, grouped by the partition the change was made in and the kind of change. Each of them carries a provenance-group extension (http://hapifhir.io/fhir/StructureDefinition/provenance-group) whose value ties the group together:

  • One main Provenance for the operation as a whole. Its extension value is the bare group id, of the form merge|<source resource type>|<source id>|<target id>|<uuid>. It references the source and target Patients, contains the original input parameters of the operation, and does not list the individual changed resources.
  • One member Provenance for each combination of partition and change type that occurred. Its extension value is the group id qualified with both, for example merge|Patient|123|456|<uuid>;partition=0;changeType=update. The change types are create, update and delete, corresponding to the copies created in the target Patient's compartment, the resources whose references were rewritten, and the originals that were deleted. Each member lists the resources it accounts for in Provenance.target, in addition to the source and target Patients of the merge.

Together these capture every change the merge made, which is what allows the operation to be undone via $hapi.fhir.undo-merge.

Current Limitations

  • Synchronous execution only.
  • With MegaScale, a merge can span more than one database, and writes across databases cannot be committed as a single transaction. A failure partway through can therefore leave the merge partially applied. The returned error lists the resources that may have been left in their merged state; these must be reverted manually.

$hapi.fhir.merge Operation
EAP

 

The $hapi.fhir.merge operation extends the merge functionality to all FHIR resource types. While the $merge operation described above is specific to Patient resources, this generic merge operation can be invoked on any resource type that has an identifier element.

Key Differences from Patient/$merge

The $hapi.fhir.merge operation differs from the Patient-specific $merge operation in several ways:

Scope

The operation is available as a resource-level operation at {resourceType}/$hapi.fhir.merge for any FHIR resource type, not just Patient resources. For example:

  • Practitioner/$hapi.fhir.merge
  • Organization/$hapi.fhir.merge

Parameter Names

The operation uses generic parameter names instead of Patient-specific ones:

  • source-resource and target-resource (instead of source-patient and target-patient)
  • source-resource-identifier and target-resource-identifier (instead of source-patient-identifier and target-patient-identifier)
  • result-resource (instead of result-patient)

Resource Linking

The way merge relationships are tracked differs based on resource type:

  • Patient resources: Continue to use the native Patient.link field to track merge relationships (with type of replaces and replaced-by)

  • Non-Patient resources: Use FHIR extensions to track merge relationships:

    • The target resource receives an extension with URL http://hl7.org/fhir/StructureDefinition/replaces pointing to the source resource
    • The source resource receives an extension with URL http://hl7.org/fhir/StructureDefinition/replaced-by pointing to the target resource

    Example extension structure on target resource:

    "extension": [
      {
        "url": "http://hl7.org/fhir/StructureDefinition/replaces",
        "valueReference": {
          "reference": "Practitioner/100"
        }
      }
    ]
    

Provenance Activity Reason Codes

The operation automatically selects the appropriate FHIR activity reason code for the Provenance resource generated after the operation:

  • PATADMIN - Used for Patient resource merges
  • RECORDMGT - Used for all other resource type merges

$hapi.fhir.merge Input Parameters

NameTypeDefaultNotes
source-resource-identifierIdentifier List of source resource identifiers
source-resourceReference Source resource
target-resource-identifierIdentifier List of target resource identifiers
target-resourceReference Target resource
result-resourceResource Optional merged resource
previewBooleanfalseIf true, no changes will be made and response will summarize what would happen were the merge to occur
delete-sourceBooleanfalseIf true, delete the source resource
resource-limitInteger512If the request is synchronous and the number of resources to change exceeds this threshold, the operation will fail with 412 Precondition Failed.

These parameters mirror the Patient $merge operation parameters, but with generic naming. The resource-limit parameter works the same way: if larger than 10,000, the value 10,000 will be used. When performed asynchronously (via the Prefer: respond-async HTTP header), the resource-limit parameter is ignored and the operation is performed in batches of 1024 resource patches at a time.

$hapi.fhir.merge Output Parameters

NameTypeNotes
inputParametersA copy of the input parameters used in the $hapi.fhir.merge operation
outcomeOperationOutcomeDetails about the result of the merge
resultResourceThe merged resource (resource type matches the operation target)
taskTaskIf the merge operation was performed asynchronously, this Task resource provides details about the status of the merge operation

Merge Provenance Resource

Like the Patient $merge operation, the $hapi.fhir.merge operation creates a Provenance resource upon successful completion. This Provenance resource contains, in its target element, the versioned references to the target resource, the source resource, and all other resources updated as part of the operation. The Provenance.activity is set to http://terminology.hl7.org/CodeSystem/iso-21089-lifecycle|merge, and the Provenance.agent.who is populated with a logical reference to the request user.

The activity reason code in the Provenance resource is automatically selected based on resource type: PATADMIN for Patient resources and RECORDMGT for all other resource types.

Merge Example: Merge Practitioner/100 into Practitioner/200 synchronously

This example demonstrates merging two Practitioner resources. The operation uses the generic parameter names and creates extensions to track the merge relationship.

Input:

POST /Practitioner/$hapi.fhir.merge
Content-Type: application/fhir+json

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "source-resource",
      "valueReference": {
        "reference": "Practitioner/100"
      }
    },
    {
      "name": "target-resource",
      "valueReference": {
        "reference": "Practitioner/200"
      }
    }
  ]
}

Output:

HTTP/1.1 200 OK

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "input",
      "resource": {
        "resourceType": "Parameters",
        "parameter": [
          {
            "name": "source-resource",
            "valueReference": {
              "reference": "Practitioner/100"
            }
          },
          {
            "name": "target-resource",
            "valueReference": {
              "reference": "Practitioner/200"
            }
          }
        ]
      }
    },
    {
      "name": "outcome",
      "resource": {
        "resourceType": "OperationOutcome",
        "issue": [
          {
            "severity": "information",
            "details": {
              "text": "Merge operation completed successfully."
            }
          }
        ]
      }
    },
    {
      "name": "result",
      "resource": {
        "resourceType": "Practitioner",
        "id": "200",
        "extension": [
          {
            "url": "http://hl7.org/fhir/StructureDefinition/replaces",
            "valueReference": {
              "reference": "Practitioner/100"
            }
          }
        ],
        "identifier": [
          {
            "system": "http://example.org/practitioner-ids",
            "value": "PRAC200"
          },
          {
            "use": "old",
            "system": "http://example.org/practitioner-ids",
            "value": "PRAC100"
          }
        ],
        "name": [
          {
            "family": "Smith",
            "given": ["John"]
          }
        ]
      }
    }
  ]
}

Note that the merged Practitioner resource (Practitioner/200) includes:

  • An extension with URL http://hl7.org/fhir/StructureDefinition/replaces pointing to the source resource (Practitioner/100)
  • The identifiers from both the source and target resources, with the source identifiers marked with use: "old"
  • The merged content based on the merge rules

$hapi.fhir.undo-merge Operation
EAP

 

The $hapi.fhir.undo-merge operation undoes a previous merge (Patient/$merge or {resourceType}/$hapi.fhir.merge). It is called with the same source and target that were merged; see Undo Merge Examples for a full request.

It works from the Provenance resource or resources the merge created, which record every resource the merge changed, and restores each of them to its version from before the merge. The restore is done as an update, so it creates a new version of each resource. The operation runs as a transaction and restores either everything or nothing.

The operation is available as a resource-level operation at {resourceType}/$hapi.fhir.undo-merge for any FHIR resource type that supports merge operation. For example:

  • Patient/$hapi.fhir.undo-merge
  • Practitioner/$hapi.fhir.undo-merge
  • Organization/$hapi.fhir.undo-merge

Limitations

The hapi.fhir.undo-merge operation currently has the following limitations:

  • It fails if any resources to be restored have been subsequently changed since the merge operation was performed.
  • It can only run synchronously.
  • It fails if the number of resources to restore exceeds Internal Synchronous Search Size configuration. An external FHIR scripting approach can be used for larger undo operations.
  • There are additional limitations when using source identifiers instead of source resource references.
  • It requires resource history, since it works by restoring previous versions. It cannot be used when resource history is disabled.
  • With MegaScale, a restore that spans more than one database cannot be committed as a single transaction. If the undo fails partway through, the returned error identifies which Provenance resources' changes were restored and which were not.

$hapi.fhir.undo-merge Input Parameters

The $hapi.fhir.undo-merge input parameters are a subset of the $hapi.fhir.merge operation input parameters. They are used to identify the source and target resources from a merge operation that should be restored to their previous version. Resources can be identified either by reference or by identifiers.

NameTypeNotes
source-resource-identifierIdentifierList of source resource identifiers
source-resourceReferenceSource resource reference
target-resource-identifierIdentifierList of target resource identifiers
target-resourceReferenceTarget resource reference

Patient Backward Compatibility: For Patient resources, the Patient-specific parameter names (source-patient, target-patient, source-patient-identifier, target-patient-identifier) are still supported for backward compatibility with the Patient/$merge operation. Mixing parameter name styles will not work; use either all generic names or all Patient-specific names.

$hapi.fhir.undo-merge Output Parameters

NameTypeNotes
outcomeOperationOutcomeOutcome of the operation

Undo Merge Examples

Example: Undo merge of Practitioner/100 into Practitioner/200

Assuming that the $hapi.fhir.merge operation was previously performed to merge Practitioner/100 into Practitioner/200, you can undo that operation by calling $hapi.fhir.undo-merge with the same source and target parameters:

Input:

POST /Practitioner/$hapi.fhir.undo-merge
Content-Type: application/fhir+json

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "source-resource",
      "valueReference": {
        "reference": "Practitioner/100"
      }
    },
    {
      "name": "target-resource",
      "valueReference": {
        "reference": "Practitioner/200"
      }
    }
  ]
}

Output:

HTTP/1.1 200 OK

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "outcome",
      "resource": {
        "resourceType": "OperationOutcome",
        "issue": [
          {
            "severity": "information",
            "details": {
              "text": "Successfully restored 8 resources to their previous versions based on the Provenance resource: Provenance/1475/_history/1"
            }
          }
        ]
      }
    }
  ]
}

Example: Undo merge of Patient/2 into Patient/3 (using Patient-specific parameter names)

For Patient resources merged via Patient/$merge, you can use the Patient-specific parameter names for backward compatibility:

Input:

POST /Patient/$hapi.fhir.undo-merge
Content-Type: application/fhir+json

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "source-patient",
      "valueReference": {
        "reference": "Patient/2"
      }
    },
    {
      "name": "target-patient",
      "valueReference": {
        "reference": "Patient/3"
      }
    }
  ]
}

Output:

HTTP/1.1 200 OK

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "outcome",
      "resource": {
        "resourceType": "OperationOutcome",
        "issue": [
          {
            "severity": "information",
            "details": {
              "text": "Successfully restored 5 resources to their previous versions based on the Provenance resource: Provenance/2341/_history/1"
            }
          }
        ]
      }
    }
  ]
}

Limitations When Using Source Identifiers

When calling $hapi.fhir.undo-merge with source identifiers (instead of source resource reference), the following additional limitations apply:

  • Source identifier-based undo only works if the original merge also used source identifiers. The merge operation stores its input parameters as part of the Provenance resource it creates. When source identifiers are used in undo-merge, the undo-merge operation locates the correct Provenance resource by matching these stored input parameters. Use source-resource-identifier in undo-merge only if the merge operation was called with source-resource-identifier. In other words, if the original merge was called with a source reference (e.g., source-resource: Practitioner/123), then the undo-merge must also be called with a source reference rather than source identifiers.

  • All source identifiers provided to undo-merge must have been provided to the original merge. The matching logic requires that every source identifier provided to undo-merge exists in the Provenance's stored parameters. Providing a subset of the original source identifiers is allowed, but providing additional identifiers that were not part of the original merge will cause the operation to fail.

Best Practice: Either use the same source parameters as the original merge operation, or use source reference which always works.

$hapi.fhir.replace-references Operation
EAP

 

The $hapi.fhir.replace-references operation searches for all resources in the repository that have a reference to the source resource, and updates those references to point to the target resource. It is a simplified form of the $merge operation when all you want to do is update references. This operation creates a Transaction Bundle of Patch operations to update the references and returns the output of performing that transaction.

$hapi.fhir.replace-references Input Parameters

NameTypeDefaultNotes
source-reference-idString The id of the source resource reference to be replaced
target-reference-idString The id of the target resource reference that the references will be replaced with
resource-limitInteger512If the request is synchrononous and the number of resources to change exceeds this threshold, the operation will fail with 412 Precondition Failed. This parameter has no effect if the Prefer: respond-async header is set

The resource-limit parameter is available to control how many resources can be changed by this operation. If resource-limit is larger than 10,000, the value 10,000 will be used.

If you request that the operation be performed asynchronously by providing the Prefer: respond-async HTTP header, then the resource-limit parameter is ignored.

When performed asynchronously, the operation is performed in batches of 1024 resource patches at a time, via PATCH transaction Bundles.

$hapi.fhir.replace-references Output Parameters

NameTypeNotes
outcomeBundleThe result of the Bundle patch transaction
taskTaskIf the operation was performed asynchronously, this Task resource provides details about the status of the operation

See the $merge operation above for details about the returned Task resource in the case when the operation is performed asynchronously.

Replace References Provenance Resource

With the 2025.08 release, the $hapi.fhir.replace-references operation creates a Provenance resource upon successful completion. This Provenance resource contains, in its target element, the versioned references to the target resource, the source resource, and the resources updated as part of the operation. The Provenance.activity is set to http://terminology.hl7.org/CodeSystem/iso-21089-lifecycle|link, and the Provenance.agent.who is populated with a logical reference to the request user. Note that a provenance resource for this operation is not created if no resources were actually updated because the source resource is not referenced by any resources.

Replace References Example: Replace all references to Patient/2 with references to Patient/3 synchronously.

Input:

POST /$hapi.fhir.replace-references
Content-Type: application/fhir+json

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "source-reference-id",
      "valueString": "Patient/2"
    },
    {
      "name": "target-reference-id",
      "valueString": "Patient/3"
    }
  ]
}

Output:

HTTP/1.1 200 OK

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "outcome",
      "resource": {
        "resourceType": "Bundle",
        "id": "782add05-549c-4a7e-a687-38c22f2f12d0",
        "type": "transaction-response",
        "entry": [
          {
            "response": {
              "status": "200 OK",
              "location": "CarePlan/62/_history/2",
              "etag": "2",
              "outcome": {
                "resourceType": "OperationOutcome",
                "issue": [
                  {
                    "severity": "information",
                    "code": "informational",
                    "details": {
                      "coding": [
                        {
                          "system": "https://hapifhir.io/fhir/CodeSystem/hapi-fhir-storage-response-code",
                          "code": "SUCCESSFUL_PATCH",
                          "display": "Patch succeeded."
                        }
                      ]
                    },
                    "diagnostics": "Successfully patched resource \"CarePlan/62/_history/2\"."
                  }
                ]
              }
            }
          },
          ... etc outcome of the rest of the patch operations ...
        ]
      }
    }
  }
}

Limitations

The $hapi.fhir.replace-references operation currently has the following limitations:

  • It is not supported when the source and target resources reside in different partitions. Note that in Patient ID Partition mode a resource's partition is determined by the Patient compartment it belongs to, so replacing references from one Patient to another falls into this case.
  • It is not supported when MegaScale is in use.

$hapi.fhir.undo-replace-references Operation
EAP

 

The $hapi.fhir.undo-replace-references operation undoes the effects of the most recent $hapi.fhir. replace-references operation on the given source and target ids. This operation uses the Provenance resource that was created by the $hapi.fhir.replace-references operation, and restores all the resources that were updated as part of the operation back to their versions before the operation. This restore operation is done as an update, so it actually creates a newer version of each restored resource. The operation is performed as a transaction, so it either restores all or none.

Limitations

The hapi.fhir.undo-replace-references operation currently has the following limitations:

  • It fails if any resources to be restored have been subsequently changed since the $hapi.fhir.replace-references operation was performed.
  • It can only run synchronously.
  • It fails if the number of resources to restore exceeds Internal Synchronous Search Size configuration. An external FHIR scripting approach can be used for larger undo operations.
  • It requires resource history, since it works by restoring previous versions. It cannot be used when resource history is disabled.

$hapi.fhir.undo-replace-references Input Parameters

NameTypeDefaultNotes
source-reference-idString The id of the source resource, this must be the same as the source-reference-id that was used in the $hapi-fhir-replace-references operation being undone
target-reference-idString The id of the target resource, this must be the same as the target-reference-id that was used in the $hapi-fhir-replace-references operation being undone

$hapi.fhir.replace-references Output Parameters

NameTypeNotes
outcomeOperationOutcomeThe outcome of the operation

Undo Replace References Example:

Assuming that the $hapi.fhir.replace-references operation was previously performed to replace references from Patient/2 to Patient/3, you can undo that operation with the following request:

Input:

POST /$hapi.fhir.undo-replace-references
Content-Type: application/fhir+json

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "source-reference-id",
      "valueString": "Patient/2"
    },
    {
      "name": "target-reference-id",
      "valueString": "Patient/3"
    }
  ]
}

Output:

HTTP/1.1 200 OK

{
    "resourceType": "Parameters",
    "parameter": [
        {
            "name": "outcome",
            "resource": {
                "resourceType": "OperationOutcome",
                "issue": [
                    {
                        "severity": "information",
                        "diagnostics": "Successfully restored 8 resources to their previous versions based on the Provenance resource: Provenance/1234/_history/1"
                    }
                ]
            }
        }
    ]
}

$sdh.mdm-bundle-match Operation for Deduplication

 

The $sdh.mdm-bundle-match operation is particularly useful for preventing duplicate data when importing FHIR Bundles. This operation processes bundles to identify resources that match existing repository resources using MDM rules.

When to Use for Deduplication

Deduplication ScenarioRecommended ModeRationale
Prevent duplicate importsmerge=false (default)Removes matching resources from bundle, updates references
Data enrichment/correctionmerge=trueMerges new data with existing resources using survivorship rules
Bulk data migrationmerge=falseEfficiently deduplicates without changing existing resources
Incremental updatesmerge=trueApplies business rules to merge incremental changes

Relationship to Other Deduplication Operations

  • $merge: Use for post-import patient merging after duplicates are discovered
  • $hapi.fhir.replace-references: Use when you need only reference updates without resource merging
  • $sdh.mdm-bundle-match: Use for prevention during import, especially with transaction bundles

Complete Documentation

For complete technical documentation including parameters, examples, limitations, and the MDM Bundle Match Processor, see MDM Operations - $sdh.mdm-bundle-match.

$sdh.mdm-deduplicate Operation
EAP

 

The $sdh.mdm-deduplicate operation starts a batch job that runs in the background to find and merge duplicate resources using your configured MDM matching rules.

For each resource selected by the supplied search URL(s), the job runs MDM matching (the same rules used by $mdm-match) to find candidate duplicates. Unlike MATCH_AND_LINK MDM, it does not create MDM links or Golden Resources. Instead, each confirmed (exact MATCH) pair is merged using the same engine as the $merge operation described above: the source resource's data and references are moved onto the target, and by default the source resource is then deleted — reducing the total resource count.

To keep merges safe and reviewable, the job only merges unambiguous matches:

  • POSSIBLE_MATCH results are skipped and surfaced in the report for human review rather than merged automatically.
  • Transitive chains are skipped. If A matches B and B matches C, the chain is not resolved automatically.
  • Equivalent reverse pairs are de-duplicated. If both A→B and B→A are found, only the first is merged.
  • Resources with too many candidates are skipped. If the candidate search for a resource reaches the configured maximum threshold, MDM cannot determine its duplicates, so nothing is merged for it. The resource is named in the report as SKIPPED with a skipReason of TOO_MANY_CANDIDATES, rather than being silently reported as having no duplicates. See Resources Omitted From MDM Matching.

The operation (like all batch jobs) runs asynchronously.

It immediately returns 202 Accepted with an OperationOutcome and a Content-Location header pointing at the poll-for-status URL.

$sdh.mdm-deduplicate Input Parameters

NameTypeCardinalityDefaultNotes
urlstring1..* One or more FHIR search URLs selecting the resources to deduplicate, e.g. Patient?identifier=http://acme.org/mrn\|12345. At least one non-blank url is required.
previewboolean0..1falseIf true, the job reports what would be merged but persists no changes.
delete-sourceboolean0..1trueIf true (the default), the source resource of each merge is deleted. If false, the source is retained and marked inactive with a replaces/replaced-by link to the target.
Note that delete-source defaults to true for this operation, which is the opposite of the $merge operation default.

Example: Start a deduplication job

Submit one or more search URLs identifying the resources to deduplicate.

Input:

POST /$sdh.mdm-deduplicate
Content-Type: application/fhir+json

{
  "resourceType": "Parameters",
  "parameter": [
    {
      "name": "url",
      "valueString": "Patient?identifier=http://acme.org/mrn|12345"
    },
    {
      "name": "preview",
      "valueBoolean": false
    }
  ]
}

Output:

HTTP/1.1 202 Accepted
Content-Location: http://localhost:8000/$sdh.mdm-deduplicate.poll-for-status?_jobId=26738f4d-266c-4ef6-934f-1d13b1b474b9

{
  "resourceType": "OperationOutcome",
  "issue": [
    {
      "severity": "information",
      "code": "informational",
      "diagnostics": "$sdh.mdm-deduplicate job has been accepted. Poll for status at the following URL: http://localhost:8000/$sdh.mdm-deduplicate.poll-for-status?_jobId=26738f4d-266c-4ef6-934f-1d13b1b474b9"
    }
  ]
}

The Content-Location header (and the message in the returned OperationOutcome) contains the URL to poll for job status.

$sdh.mdm-deduplicate.poll-for-status Operation

Use the URL returned by $sdh.mdm-deduplicate to check on the job. The batch job instance id is supplied via the _jobId parameter.

$sdh.mdm-deduplicate.poll-for-status Input Parameters

NameTypeCardinalityNotes
_jobIdstring1..1The batch job instance id returned by $sdh.mdm-deduplicate.

Example: Poll a job that is still running

Input:

GET /$sdh.mdm-deduplicate.poll-for-status?_jobId=26738f4d-266c-4ef6-934f-1d13b1b474b9

Output:

HTTP/1.1 202 Accepted
X-Progress: $sdh.mdm-deduplicate job has started and is in progress Current step: find-matches. Overall progress: 40%.

{
  "resourceType": "OperationOutcome",
  "issue": [
    {
      "severity": "information",
      "code": "informational",
      "diagnostics": "$sdh.mdm-deduplicate job has started and is in progress Current step: find-matches. Overall progress: 40%."
    }
  ]
}

While the job is queued, in progress, or being finalized, the poll returns 202 Accepted. Keep polling until you receive 200 OK.

Example: Poll a completed job

When the job has completed, the poll returns 200 OK with a batch-response Bundle. The deduplication report is returned as the text of the contained OperationOutcome issue, summarizing how many pairs were merged (successTotal), skipped (skippedTotal), and failed (errorTotal), along with per-pair details.

Output:

HTTP/1.1 200 OK

{
  "resourceType": "Bundle",
  "type": "batch-response",
  "entry": [
    {
      "response": {
        "status": "200 OK",
        "outcome": {
          "resourceType": "OperationOutcome",
          "issue": [
            {
              "severity": "information",
              "code": "informational",
              "diagnostics": "{\"successTotal\":1,\"skippedTotal\":2,\"errorTotal\":0,\"results\":[{\"mergePair\":{\"targetId\":\"Patient/3\",\"sourceId\":\"Patient/2\",\"resourceType\":\"Patient\",\"exactMatch\":true},\"outcome\":\"SUCCESS\",\"message\":\"Successfully merged Patient/2 (source) into Patient/3 (target); preview=false.\"},{\"mergePair\":{\"targetId\":\"Patient/7\",\"sourceId\":\"Patient/6\",\"resourceType\":\"Patient\",\"exactMatch\":false},\"outcome\":\"SKIPPED\",\"message\":\"Possible MDM match: Patient/6 is a possible duplicate of Patient/7; surfacing for human review\",\"skipReason\":\"POSSIBLE_MATCH\"},{\"mergePair\":{\"targetId\":\"Patient/9\",\"resourceType\":\"Patient\",\"exactMatch\":true,\"skipReason\":\"TOO_MANY_CANDIDATES\"},\"outcome\":\"SKIPPED\",\"message\":\"Skipping merge for Patient/9 because too many match candidates were found; no merge source was selected.\",\"skipReason\":\"TOO_MANY_CANDIDATES\"}]}"
            }
          ]
        }
      }
    }
  ]
}

The report is a JSON document with the following fields:

FieldNotes
successTotalNumber of pairs successfully merged.
skippedTotalNumber of pairs skipped (possible matches, transitive chains, already-processed equivalent pairs, or resources with too many candidates).
errorTotalNumber of pairs where a merge was attempted but failed.
resultsPer-pair details: the mergePair (targetId, sourceId, resourceType, exactMatch), the outcome (SUCCESS, SKIPPED, or ERROR), a human-readable message, and — when the outcome is SKIPPED — a skipReason. Detailed results are capped at the first 1000 pairs; beyond that only the totals continue to increment.

When the outcome is SKIPPED, the skipReason field classifies why, so that consumers can act on the category without parsing the message text. It is absent for SUCCESS and ERROR results.

skipReasonNotes
TOO_MANY_CANDIDATESThe candidate search for the resource reached the configured maximum threshold, so MDM could not determine its duplicates. Nothing was merged for it, and the mergePair names only a targetId — there is no source. See Resources Omitted From MDM Matching.
POSSIBLE_MATCHThe pair matched as a POSSIBLE_MATCH rather than an exact MATCH. It is surfaced for human review rather than merged automatically.
EQUIVALENT_PAIRThe equivalent merge in the opposite direction was already processed in this job, so only the first of A→B and B→A is merged.
TRANSITIVE_PAIRMerging this pair would extend a transitive chain (A→B, B→C), which the job does not resolve automatically.

Deduplication Roadmap

 

The Deduplication features of Smile CDR are under active development. Here is a roadmap of new features we are planning to roll out.

Nov 2025 Release

  • Extend $merge to all Resource types
  • Create a third MDM mode: MATCH_AND_MERGE which is similar to MATCH_ONLY in that it uses MDM Rules but does not create any links or Golden Resources. When enabled, a MATCH_AND_MERGE MDM module will automatically perform a $merge operation on all inbound resources. Any resources that MATCH a single target will be merged into that resources following the MDM Survivorship rules and all references will be updated to point to the merged resource.
  • Provide a new $submit-for-deduplication operation that works like $mdm-submit and performs MATCH_AND_MERGE on all resources that match the criteria in the request. E.g. If Organization resources with _source=ABC were accidentally duplicated in your FHIR Repository, you could call $submit-for-deduplication with the criteria Organization?_source=ABC to submit all of those Organizations for deduplication. Ones that have existing matches would be deleted and all references updated to point to the remaining copy of that organization.