FHIR Repositories can end up with duplicate resources in them. This page details tools that Smile provides to work with duplicated data.
| Cause of Duplicates | Recommended Approach |
|---|---|
| The same entities were imported from different source systems. E.g. Jane Smith's Patient record is imported from the Lab system as Patient/123 and from the Pharmacy system as Patient/456. | Preserve the source data as is so that future data continues to be properly associated to the right patients. Use MDM to establish Golden Resources and MDM LINK records outside of the data and use MDM features like Observation?patient=Patient/123&_mdm=true to link the data at search time. |
| Data was accidentally duplicated on import. E.g. the same data was accidentally loaded twice (as POST new resources) or Conditional Create directives failed to match the intended target correctly. This kind of unintended duplication can also occur when translating HL7v2, CDA, or CSV data into FHIR resources. | In the case where the data was accidentally duplicated, it may make sense to "clean up" the duplicates. See below for details on Smile CDR tools to merge such duplicates. |
Smile CDR provides several operations to deduplicate data:
$merge which is a backport of the FHIR R5 Patient/$merge specification to FHIR R4$hapi.fhir.merge which extends the merge functionality to all FHIR resource types that have an identifier element$hapi.fhir.undo-merge undoes the effects of a previous merge operation$hapi.fhir.replace-references performs only the update references part of this $merge operation$hapi.fhir.undo-replace-references undoes the effects of a $hapi.fhir.replace-references operation$sdh.mdm-bundle-match processes FHIR Bundles to match resources using MDM rules and optionally merge them using survivorship$sdh.mdm-deduplicate starts a batch job that will run in the background and use the mdm rules to deduplicate resourcesSee the FHIR R5 Patient/$merge specification page for a description of this operation. See the bottom of this page for details on the current roadmap for enhancing Smile CDR deduplication functionality.
| Name | Type | Default | Notes |
|---|---|---|---|
| source-patient-identifier | Identifier | List of source patient identifiers | |
| source-patient | Reference | Source patient | |
| target-patient-identifier | Identifier | List of target patient identifiers | |
| target-patient | Reference | Target patient | |
| result-patient | Patient | Optional merged patient resource | |
| preview | Boolean | false | If true, no changes will be made and response will summarize what would happen were the merge to occur |
| delete-source | Boolean | false | If true, delete the source resource |
| resource-limit | Integer | 512 | If the request is synchrononous and the number of resources to change exceeds this threshold, the operation will fail with 412 Precondition Failed. This parameter has no effect if the Prefer: respond-async header is set |
See the FHIR R5 Patient/$merge specification for a detailed description of these input parameters.
resource-limit is a Smile CDR addition to protect users from accidentally changing too many resources at once. If resource-limit is larger than 10,000, the value 10,000 will be used.
If you request that the operation be performed asynchronously by providing the Prefer: respond-async HTTP header, then the resource-limit parameter is ignored.
When performed asynchronously, the operation is performed in batches of 1024 resource patches at a time, via PATCH transaction Bundles.
| Name | Type | Notes |
|---|---|---|
| input | Parameters | A copy of the input parameters used in the $merge operation |
| outcome | OperationOutcome | Details about the result of the merge |
| result | Patient | The merged Patient resource |
| task | Task | If the merge operation was performed asynchronously, this Task resource provides details about the status of the merge operation |
With the 2025.08 release, the $merge operation creates a Provenance resource upon successful completion.
This Provenance resource contains, in its target element, the versioned references to the target patient,
the source patient (if not deleted during the operation), and all other resources updated as part of the operation.
The Provenance.activity is set to http://terminology.hl7.org/CodeSystem/iso-21089-lifecycle|merge, and the Provenance.agent.who is populated with a logical reference to the request user.
Input:
POST /Patient/$merge
Content-Type: application/fhir+json
{
"resourceType": "Parameters",
"parameter": [
{
"name": "source-patient",
"valueReference": {
"reference": "Patient/2"
}
},
{
"name": "target-patient",
"valueReference": {
"reference": "Patient/3"
}
}
]
}
Output:
HTTP/1.1 200 OK
{
"resourceType": "Parameters",
"parameter": [
{
"name": "input",
"resource": {
"resourceType": "Parameters",
"parameter": [
{
"name": "source-patient",
"valueReference": {
"reference": "Patient/2"
}
},
{
"name": "target-patient",
"valueReference": {
"reference": "Patient/3"
}
}
]
}
},
{
"name": "outcome",
"resource": {
"resourceType": "OperationOutcome",
"issue": [
{
"severity": "information",
"details": {
"text": "Merge operation completed successfully."
}
}
]
}
},
{
"name": "result",
"resource": {
"resourceType": "Patient",
"id": "3",
"identifier": [
{
"system": "SYS2A",
"value": "VAL2A"
},
{
"system": "SYS2B",
"value": "VAL2B"
},
{
"system": "SYSC",
"value": "VALC"
},
{
"use": "old",
"system": "SYS1A",
"value": "VAL1A"
},
{
"use": "old",
"system": "SYS1B",
"value": "VAL1B"
}
],
"link": [
{
"other": {
"reference": "Patient/2"
},
"type": "replaces"
}
]
}
}
]
}
Input:
POST /Patient/$merge
Content-Type: application/fhir+json
{
"resourceType": "Parameters",
"parameter": [
{
"name": "source-patient",
"valueReference": {
"reference": "Patient/2"
}
},
{
"name": "target-patient",
"valueReference": {
"reference": "Patient/3"
}
},
{
"name": "preview",
"valueBoolean": true
}
]
}
Output:
HTTP/1.1 200 OK
{
"resourceType": "Parameters",
"parameter": [
{
"name": "input",
"resource": {
"resourceType": "Parameters",
...copy of input...
}
},
{
"name": "outcome",
"resource": {
"resourceType": "OperationOutcome",
"issue": [
{
"severity": "information",
"details": {
"text": "Preview only merge operation - no issues detected"
},
"diagnostics": "Merge would update 25 resources"
}
]
}
},
{
"name": "result",
"resource": {
"resourceType": "Patient",
...merged patient...
}
}
]
}
Input:
POST /Patient/$merge
Content-Type: application/fhir+json
Prefer: respond-async
{
"resourceType": "Parameters",
"parameter": [ {
"name": "source-patient-identifier",
"valueIdentifier": {
"system" : "urn:oid:1.2.36.146.595.217.0.1",
"value" : "12345"
}
},
{
"name": "target-patient-identifier",
"valueIdentifier": {
"system" : "urn:oid:1.2.36.146.595.217.0.1",
"value" : "12346"
}
}
]
}
Output:
HTTP/1.1 202 Accepted
{
"resourceType": "Parameters",
"parameter": [
{
"name": "input",
"resource": {
"resourceType": "Parameters",
...copy of input...
}
},
{
"name": "task",
"resource": {
"resourceType": "Task",
"id": "352",
"identifier": [
{
"system": "http://hapifhir.io/batch/jobId",
"value": "26738f4d-266c-4ef6-934f-1d13b1b474b9"
}
],
"status": "in-progress"
}
}
]
}
You can poll the status of the returned Task resource to see when it completes. For more detailed status about the background job, you can view the status of the corresponding Smile CDR batch job either through the Web Admin Console, or through the Admin JSON API. The ID of the Smile CDR batch job is provided as an identifer on the returned Task.
In Patient ID Partition mode, a resource is stored in the partition of the Patient compartment it belongs to, and an update cannot change a resource's partition. A compartment resource therefore cannot simply be repointed at the target Patient; the $merge operation will:
This is how every merge behaves in Patient ID Partition mode, including when the source and target Patients resolve to the same database partition.
In addition to the steps above, the source and target Patients themselves are updated as in an unpartitioned merge: copying identifiers from source to target, adding replaces / replaced-by links, optionally deleting the source via the delete-source parameter, etc.
A merge in this mode records what it changed across several Provenance resources rather than a single one, grouped by the partition the change was made in and the kind of change. Each of them carries a provenance-group extension (http://hapifhir.io/fhir/StructureDefinition/provenance-group) whose value ties the group together:
merge|<source resource type>|<source id>|<target id>|<uuid>. It references the source and target Patients, contains the original input parameters of the operation, and does not list the individual changed resources.merge|Patient|123|456|<uuid>;partition=0;changeType=update. The change types are create, update and delete, corresponding to the copies created in the target Patient's compartment, the resources whose references were rewritten, and the originals that were deleted. Each member lists the resources it accounts for in Provenance.target, in addition to the source and target Patients of the merge.Together these capture every change the merge made, which is what allows the operation to be undone via $hapi.fhir.undo-merge.
The $hapi.fhir.merge operation extends the merge functionality to all FHIR resource types. While the $merge operation described above is specific to Patient resources, this generic merge operation can be invoked on any resource type that has an identifier element.
The $hapi.fhir.merge operation differs from the Patient-specific $merge operation in several ways:
The operation is available as a resource-level operation at {resourceType}/$hapi.fhir.merge for any FHIR resource type, not just Patient resources. For example:
Practitioner/$hapi.fhir.mergeOrganization/$hapi.fhir.mergeThe operation uses generic parameter names instead of Patient-specific ones:
source-resource and target-resource (instead of source-patient and target-patient)source-resource-identifier and target-resource-identifier (instead of source-patient-identifier and target-patient-identifier)result-resource (instead of result-patient)The way merge relationships are tracked differs based on resource type:
Patient resources: Continue to use the native Patient.link field to track merge relationships (with type
of replaces and replaced-by)
Non-Patient resources: Use FHIR extensions to track merge relationships:
http://hl7.org/fhir/StructureDefinition/replaces pointing to the source resourcehttp://hl7.org/fhir/StructureDefinition/replaced-by pointing to the target resourceExample extension structure on target resource:
"extension": [
{
"url": "http://hl7.org/fhir/StructureDefinition/replaces",
"valueReference": {
"reference": "Practitioner/100"
}
}
]
The operation automatically selects the appropriate FHIR activity reason code for the Provenance resource generated after the operation:
PATADMIN - Used for Patient resource mergesRECORDMGT - Used for all other resource type merges| Name | Type | Default | Notes |
|---|---|---|---|
| source-resource-identifier | Identifier | List of source resource identifiers | |
| source-resource | Reference | Source resource | |
| target-resource-identifier | Identifier | List of target resource identifiers | |
| target-resource | Reference | Target resource | |
| result-resource | Resource | Optional merged resource | |
| preview | Boolean | false | If true, no changes will be made and response will summarize what would happen were the merge to occur |
| delete-source | Boolean | false | If true, delete the source resource |
| resource-limit | Integer | 512 | If the request is synchronous and the number of resources to change exceeds this threshold, the operation will fail with 412 Precondition Failed. |
These parameters mirror the Patient $merge operation parameters, but with generic naming. The resource-limit parameter works the same way: if larger than 10,000, the value 10,000 will be used. When performed asynchronously (via the Prefer: respond-async HTTP header), the resource-limit parameter is ignored and the operation is performed in batches of 1024 resource patches at a time.
| Name | Type | Notes |
|---|---|---|
| input | Parameters | A copy of the input parameters used in the $hapi.fhir.merge operation |
| outcome | OperationOutcome | Details about the result of the merge |
| result | Resource | The merged resource (resource type matches the operation target) |
| task | Task | If the merge operation was performed asynchronously, this Task resource provides details about the status of the merge operation |
Like the Patient $merge operation, the $hapi.fhir.merge operation creates a Provenance resource upon successful completion. This Provenance resource contains, in its target element, the versioned references to the target resource, the source resource, and all other resources updated as part of the operation. The Provenance.activity is set to http://terminology.hl7.org/CodeSystem/iso-21089-lifecycle|merge, and the Provenance.agent.who is populated with a logical reference to the request user.
The activity reason code in the Provenance resource is automatically selected based on resource type: PATADMIN for Patient resources and RECORDMGT for all other resource types.
This example demonstrates merging two Practitioner resources. The operation uses the generic parameter names and creates extensions to track the merge relationship.
Input:
POST /Practitioner/$hapi.fhir.merge
Content-Type: application/fhir+json
{
"resourceType": "Parameters",
"parameter": [
{
"name": "source-resource",
"valueReference": {
"reference": "Practitioner/100"
}
},
{
"name": "target-resource",
"valueReference": {
"reference": "Practitioner/200"
}
}
]
}
Output:
HTTP/1.1 200 OK
{
"resourceType": "Parameters",
"parameter": [
{
"name": "input",
"resource": {
"resourceType": "Parameters",
"parameter": [
{
"name": "source-resource",
"valueReference": {
"reference": "Practitioner/100"
}
},
{
"name": "target-resource",
"valueReference": {
"reference": "Practitioner/200"
}
}
]
}
},
{
"name": "outcome",
"resource": {
"resourceType": "OperationOutcome",
"issue": [
{
"severity": "information",
"details": {
"text": "Merge operation completed successfully."
}
}
]
}
},
{
"name": "result",
"resource": {
"resourceType": "Practitioner",
"id": "200",
"extension": [
{
"url": "http://hl7.org/fhir/StructureDefinition/replaces",
"valueReference": {
"reference": "Practitioner/100"
}
}
],
"identifier": [
{
"system": "http://example.org/practitioner-ids",
"value": "PRAC200"
},
{
"use": "old",
"system": "http://example.org/practitioner-ids",
"value": "PRAC100"
}
],
"name": [
{
"family": "Smith",
"given": ["John"]
}
]
}
}
]
}
Note that the merged Practitioner resource (Practitioner/200) includes:
http://hl7.org/fhir/StructureDefinition/replaces pointing to the source resource (Practitioner/100)use: "old"The $hapi.fhir.undo-merge operation undoes a previous merge (Patient/$merge or {resourceType}/$hapi.fhir.merge). It is called with the same source and target that were merged; see Undo Merge Examples for a full request.
It works from the Provenance resource or resources the merge created, which record every resource the merge changed, and restores each of them to its version from before the merge. The restore is done as an update, so it creates a new version of each resource. The operation runs as a transaction and restores either everything or nothing.
The operation is available as a resource-level operation at {resourceType}/$hapi.fhir.undo-merge for any FHIR resource type that supports merge operation. For example:
Patient/$hapi.fhir.undo-mergePractitioner/$hapi.fhir.undo-mergeOrganization/$hapi.fhir.undo-mergeThe hapi.fhir.undo-merge operation currently has the following limitations:
The $hapi.fhir.undo-merge input parameters are a subset of the $hapi.fhir.merge operation input parameters. They are used to identify the source and target resources from a merge operation that should be restored to their previous version. Resources can be identified either by reference or by identifiers.
| Name | Type | Notes |
|---|---|---|
| source-resource-identifier | Identifier | List of source resource identifiers |
| source-resource | Reference | Source resource reference |
| target-resource-identifier | Identifier | List of target resource identifiers |
| target-resource | Reference | Target resource reference |
Patient Backward Compatibility: For Patient resources, the Patient-specific parameter names (source-patient, target-patient, source-patient-identifier, target-patient-identifier) are still supported for backward compatibility with the Patient/$merge operation. Mixing parameter name styles will not work; use either all generic names or all Patient-specific names.
| Name | Type | Notes |
|---|---|---|
| outcome | OperationOutcome | Outcome of the operation |
Assuming that the $hapi.fhir.merge operation was previously performed to merge Practitioner/100 into Practitioner/200,
you can undo that operation by calling $hapi.fhir.undo-merge with the same source and target parameters:
Input:
POST /Practitioner/$hapi.fhir.undo-merge
Content-Type: application/fhir+json
{
"resourceType": "Parameters",
"parameter": [
{
"name": "source-resource",
"valueReference": {
"reference": "Practitioner/100"
}
},
{
"name": "target-resource",
"valueReference": {
"reference": "Practitioner/200"
}
}
]
}
Output:
HTTP/1.1 200 OK
{
"resourceType": "Parameters",
"parameter": [
{
"name": "outcome",
"resource": {
"resourceType": "OperationOutcome",
"issue": [
{
"severity": "information",
"details": {
"text": "Successfully restored 8 resources to their previous versions based on the Provenance resource: Provenance/1475/_history/1"
}
}
]
}
}
]
}
For Patient resources merged via Patient/$merge, you can use the Patient-specific parameter names for backward compatibility:
Input:
POST /Patient/$hapi.fhir.undo-merge
Content-Type: application/fhir+json
{
"resourceType": "Parameters",
"parameter": [
{
"name": "source-patient",
"valueReference": {
"reference": "Patient/2"
}
},
{
"name": "target-patient",
"valueReference": {
"reference": "Patient/3"
}
}
]
}
Output:
HTTP/1.1 200 OK
{
"resourceType": "Parameters",
"parameter": [
{
"name": "outcome",
"resource": {
"resourceType": "OperationOutcome",
"issue": [
{
"severity": "information",
"details": {
"text": "Successfully restored 5 resources to their previous versions based on the Provenance resource: Provenance/2341/_history/1"
}
}
]
}
}
]
}
When calling $hapi.fhir.undo-merge with source identifiers (instead of source resource reference), the following additional limitations apply:
Source identifier-based undo only works if the original merge also used source identifiers. The merge
operation stores its input parameters as part of the Provenance resource it creates. When source identifiers are
used in undo-merge, the undo-merge operation locates the correct Provenance resource by matching these stored input parameters. Use source-resource-identifier in undo-merge only if the merge operation was called with source-resource-identifier. In other words, if the original merge was called with a source reference (e.g., source-resource: Practitioner/123), then the undo-merge must also be called with a source reference rather than source identifiers.
All source identifiers provided to undo-merge must have been provided to the original merge. The matching logic requires that every source identifier provided to undo-merge exists in the Provenance's stored parameters. Providing a subset of the original source identifiers is allowed, but providing additional identifiers that were not part of the original merge will cause the operation to fail.
Best Practice: Either use the same source parameters as the original merge operation, or use source reference which always works.
The $hapi.fhir.replace-references operation searches for all resources in the repository that have a reference to the
source resource, and updates those references to point to the target resource. It is a simplified form of the $merge operation when all you want to do is update references.
This operation creates a Transaction Bundle of Patch operations to update the references and returns the output of performing that transaction.
| Name | Type | Default | Notes |
|---|---|---|---|
| source-reference-id | String | The id of the source resource reference to be replaced | |
| target-reference-id | String | The id of the target resource reference that the references will be replaced with | |
| resource-limit | Integer | 512 | If the request is synchrononous and the number of resources to change exceeds this threshold, the operation will fail with 412 Precondition Failed. This parameter has no effect if the Prefer: respond-async header is set |
The resource-limit parameter is available to control how many resources can be changed by this operation. If resource-limit is larger than 10,000, the value 10,000 will be used.
If you request that the operation be performed asynchronously by providing the Prefer: respond-async HTTP header, then the resource-limit parameter is ignored.
When performed asynchronously, the operation is performed in batches of 1024 resource patches at a time, via PATCH transaction Bundles.
| Name | Type | Notes |
|---|---|---|
| outcome | Bundle | The result of the Bundle patch transaction |
| task | Task | If the operation was performed asynchronously, this Task resource provides details about the status of the operation |
See the $merge operation above for details about the returned Task resource in the case when the operation is performed asynchronously.
With the 2025.08 release, the $hapi.fhir.replace-references operation creates a Provenance resource upon successful completion. This Provenance resource contains, in its target element, the versioned references to the target resource,
the source resource, and the resources updated as part of the operation.
The Provenance.activity is set to http://terminology.hl7.org/CodeSystem/iso-21089-lifecycle|link, and the Provenance.agent.who is populated with a logical reference to the request user. Note that a provenance resource for this operation is not created if no resources were actually updated because the source resource is not referenced by any resources.
Input:
POST /$hapi.fhir.replace-references
Content-Type: application/fhir+json
{
"resourceType": "Parameters",
"parameter": [
{
"name": "source-reference-id",
"valueString": "Patient/2"
},
{
"name": "target-reference-id",
"valueString": "Patient/3"
}
]
}
Output:
HTTP/1.1 200 OK
{
"resourceType": "Parameters",
"parameter": [
{
"name": "outcome",
"resource": {
"resourceType": "Bundle",
"id": "782add05-549c-4a7e-a687-38c22f2f12d0",
"type": "transaction-response",
"entry": [
{
"response": {
"status": "200 OK",
"location": "CarePlan/62/_history/2",
"etag": "2",
"outcome": {
"resourceType": "OperationOutcome",
"issue": [
{
"severity": "information",
"code": "informational",
"details": {
"coding": [
{
"system": "https://hapifhir.io/fhir/CodeSystem/hapi-fhir-storage-response-code",
"code": "SUCCESSFUL_PATCH",
"display": "Patch succeeded."
}
]
},
"diagnostics": "Successfully patched resource \"CarePlan/62/_history/2\"."
}
]
}
}
},
... etc outcome of the rest of the patch operations ...
]
}
}
}
}
The $hapi.fhir.replace-references operation currently has the following limitations:
The $hapi.fhir.undo-replace-references operation undoes the effects of the most recent $hapi.fhir. replace-references operation on the given source and target ids. This operation uses the Provenance resource that was created by the $hapi.fhir.replace-references operation, and restores all the resources that were updated as part of the operation back to their versions before the operation. This restore operation is done as an update, so it actually creates a newer version of each restored resource. The operation is performed as a transaction,
so it either restores all or none.
The hapi.fhir.undo-replace-references operation currently has the following limitations:
$hapi.fhir.replace-references operation was performed.| Name | Type | Default | Notes |
|---|---|---|---|
| source-reference-id | String | The id of the source resource, this must be the same as the source-reference-id that was used in the $hapi-fhir-replace-references operation being undone | |
| target-reference-id | String | The id of the target resource, this must be the same as the target-reference-id that was used in the $hapi-fhir-replace-references operation being undone |
| Name | Type | Notes |
|---|---|---|
| outcome | OperationOutcome | The outcome of the operation |
Assuming that the $hapi.fhir.replace-references operation was previously performed to replace references from Patient/2 to Patient/3,
you can undo that operation with the following request:
Input:
POST /$hapi.fhir.undo-replace-references
Content-Type: application/fhir+json
{
"resourceType": "Parameters",
"parameter": [
{
"name": "source-reference-id",
"valueString": "Patient/2"
},
{
"name": "target-reference-id",
"valueString": "Patient/3"
}
]
}
Output:
HTTP/1.1 200 OK
{
"resourceType": "Parameters",
"parameter": [
{
"name": "outcome",
"resource": {
"resourceType": "OperationOutcome",
"issue": [
{
"severity": "information",
"diagnostics": "Successfully restored 8 resources to their previous versions based on the Provenance resource: Provenance/1234/_history/1"
}
]
}
}
]
}
The $sdh.mdm-bundle-match operation is particularly useful for preventing duplicate data when importing FHIR Bundles. This operation processes bundles to identify resources that match existing repository resources using MDM rules.
| Deduplication Scenario | Recommended Mode | Rationale |
|---|---|---|
| Prevent duplicate imports | merge=false (default) | Removes matching resources from bundle, updates references |
| Data enrichment/correction | merge=true | Merges new data with existing resources using survivorship rules |
| Bulk data migration | merge=false | Efficiently deduplicates without changing existing resources |
| Incremental updates | merge=true | Applies business rules to merge incremental changes |
For complete technical documentation including parameters, examples, limitations, and the MDM Bundle Match Processor, see MDM Operations - $sdh.mdm-bundle-match.
The $sdh.mdm-deduplicate operation starts a batch job that runs in the background to find and merge duplicate resources using your configured MDM matching rules.
For each resource selected by the supplied search URL(s), the job runs MDM matching (the same rules used by $mdm-match) to find candidate duplicates. Unlike MATCH_AND_LINK MDM, it does not create MDM links or Golden Resources.
Instead, each confirmed (exact MATCH) pair is merged using the same engine as the $merge operation described above: the source resource's data and references are moved onto the target, and by default the source resource is then deleted — reducing the total resource count.
To keep merges safe and reviewable, the job only merges unambiguous matches:
POSSIBLE_MATCH results are skipped and surfaced in the report for human review rather than merged automatically.SKIPPED with a skipReason of TOO_MANY_CANDIDATES, rather than being silently reported as having no
duplicates. See
Resources Omitted From MDM Matching.The operation (like all batch jobs) runs asynchronously.
It immediately returns 202 Accepted with an OperationOutcome and a Content-Location header pointing at the poll-for-status URL.
| Name | Type | Cardinality | Default | Notes |
|---|---|---|---|---|
| url | string | 1..* | One or more FHIR search URLs selecting the resources to deduplicate, e.g. Patient?identifier=http://acme.org/mrn\|12345. At least one non-blank url is required. | |
| preview | boolean | 0..1 | false | If true, the job reports what would be merged but persists no changes. |
| delete-source | boolean | 0..1 | true | If true (the default), the source resource of each merge is deleted. If false, the source is retained and marked inactive with a replaces/replaced-by link to the target. |
delete-source defaults to true for this operation, which is the opposite of the $merge operation default.
Submit one or more search URLs identifying the resources to deduplicate.
Input:
POST /$sdh.mdm-deduplicate
Content-Type: application/fhir+json
{
"resourceType": "Parameters",
"parameter": [
{
"name": "url",
"valueString": "Patient?identifier=http://acme.org/mrn|12345"
},
{
"name": "preview",
"valueBoolean": false
}
]
}
Output:
HTTP/1.1 202 Accepted
Content-Location: http://localhost:8000/$sdh.mdm-deduplicate.poll-for-status?_jobId=26738f4d-266c-4ef6-934f-1d13b1b474b9
{
"resourceType": "OperationOutcome",
"issue": [
{
"severity": "information",
"code": "informational",
"diagnostics": "$sdh.mdm-deduplicate job has been accepted. Poll for status at the following URL: http://localhost:8000/$sdh.mdm-deduplicate.poll-for-status?_jobId=26738f4d-266c-4ef6-934f-1d13b1b474b9"
}
]
}
The Content-Location header (and the message in the returned OperationOutcome) contains the URL to poll for job status.
Use the URL returned by $sdh.mdm-deduplicate to check on the job. The batch job instance id is supplied via the _jobId parameter.
| Name | Type | Cardinality | Notes |
|---|---|---|---|
| _jobId | string | 1..1 | The batch job instance id returned by $sdh.mdm-deduplicate. |
Input:
GET /$sdh.mdm-deduplicate.poll-for-status?_jobId=26738f4d-266c-4ef6-934f-1d13b1b474b9
Output:
HTTP/1.1 202 Accepted
X-Progress: $sdh.mdm-deduplicate job has started and is in progress Current step: find-matches. Overall progress: 40%.
{
"resourceType": "OperationOutcome",
"issue": [
{
"severity": "information",
"code": "informational",
"diagnostics": "$sdh.mdm-deduplicate job has started and is in progress Current step: find-matches. Overall progress: 40%."
}
]
}
While the job is queued, in progress, or being finalized, the poll returns 202 Accepted. Keep polling until you receive 200 OK.
When the job has completed, the poll returns 200 OK with a batch-response Bundle. The deduplication report is returned as the text of the contained OperationOutcome issue, summarizing how many pairs were merged (successTotal), skipped (skippedTotal), and failed (errorTotal), along with per-pair details.
Output:
HTTP/1.1 200 OK
{
"resourceType": "Bundle",
"type": "batch-response",
"entry": [
{
"response": {
"status": "200 OK",
"outcome": {
"resourceType": "OperationOutcome",
"issue": [
{
"severity": "information",
"code": "informational",
"diagnostics": "{\"successTotal\":1,\"skippedTotal\":2,\"errorTotal\":0,\"results\":[{\"mergePair\":{\"targetId\":\"Patient/3\",\"sourceId\":\"Patient/2\",\"resourceType\":\"Patient\",\"exactMatch\":true},\"outcome\":\"SUCCESS\",\"message\":\"Successfully merged Patient/2 (source) into Patient/3 (target); preview=false.\"},{\"mergePair\":{\"targetId\":\"Patient/7\",\"sourceId\":\"Patient/6\",\"resourceType\":\"Patient\",\"exactMatch\":false},\"outcome\":\"SKIPPED\",\"message\":\"Possible MDM match: Patient/6 is a possible duplicate of Patient/7; surfacing for human review\",\"skipReason\":\"POSSIBLE_MATCH\"},{\"mergePair\":{\"targetId\":\"Patient/9\",\"resourceType\":\"Patient\",\"exactMatch\":true,\"skipReason\":\"TOO_MANY_CANDIDATES\"},\"outcome\":\"SKIPPED\",\"message\":\"Skipping merge for Patient/9 because too many match candidates were found; no merge source was selected.\",\"skipReason\":\"TOO_MANY_CANDIDATES\"}]}"
}
]
}
}
}
]
}
The report is a JSON document with the following fields:
| Field | Notes |
|---|---|
| successTotal | Number of pairs successfully merged. |
| skippedTotal | Number of pairs skipped (possible matches, transitive chains, already-processed equivalent pairs, or resources with too many candidates). |
| errorTotal | Number of pairs where a merge was attempted but failed. |
| results | Per-pair details: the mergePair (targetId, sourceId, resourceType, exactMatch), the outcome (SUCCESS, SKIPPED, or ERROR), a human-readable message, and — when the outcome is SKIPPED — a skipReason. Detailed results are capped at the first 1000 pairs; beyond that only the totals continue to increment. |
When the outcome is SKIPPED, the skipReason field classifies why, so that consumers can act on the category without parsing the message text. It is absent for SUCCESS and ERROR results.
| skipReason | Notes |
|---|---|
TOO_MANY_CANDIDATES | The candidate search for the resource reached the configured maximum threshold, so MDM could not determine its duplicates. Nothing was merged for it, and the mergePair names only a targetId — there is no source. See Resources Omitted From MDM Matching. |
POSSIBLE_MATCH | The pair matched as a POSSIBLE_MATCH rather than an exact MATCH. It is surfaced for human review rather than merged automatically. |
EQUIVALENT_PAIR | The equivalent merge in the opposite direction was already processed in this job, so only the first of A→B and B→A is merged. |
TRANSITIVE_PAIR | Merging this pair would extend a transitive chain (A→B, B→C), which the job does not resolve automatically. |
The Deduplication features of Smile CDR are under active development. Here is a roadmap of new features we are planning to roll out.
$merge to all Resource typesMATCH_AND_MERGE which is similar to MATCH_ONLY in that it uses MDM Rules but does not create any links or Golden Resources. When enabled, a MATCH_AND_MERGE MDM module will automatically perform a $merge operation on all inbound resources. Any resources that MATCH a single target will be merged into that resources following the MDM Survivorship rules and all references will be updated to point to the merged resource.$submit-for-deduplication operation that works like $mdm-submit and performs MATCH_AND_MERGE on all resources that match the criteria in the request. E.g. If Organization resources with _source=ABC were accidentally duplicated in your FHIR Repository, you could call $submit-for-deduplication with the criteria Organization?_source=ABC to submit all of those Organizations for deduplication. Ones that have existing matches would be deleted and all references updated to point to the remaining copy of that organization.