Supported Match Algorithms

 
Applies to: Deduplication on ingestion, MDM in EID mode, and Probabilistic MDM — the choice of algorithm is independent of the strategy you deploy.

Each field comparison rule in an MDM rule definition requires you to specify a matching algorithm. This page describes every built-in algorithm, the parameters that control its behaviour, and guidance to help you choose the right one for each field.

Matchers and Similarities

Algorithms fall into two categories based on what they return.

Matchers return a simple pass or fail. A matchField that uses a matcher reports true if the two field values satisfy the algorithm's matching condition, and false otherwise.

Similarities return a decimal score between 0.0 (completely different) and 1.0 (identical). A matchField that uses a similarity compares that score against a configurable matchThreshold. If the score meets or exceeds the threshold, the comparison passes; otherwise it fails.

Both categories accept an exact parameter, described below. Similarities also require a matchThreshold parameter.

Common Parameter: exact

All algorithms accept an optional boolean parameter named exact. When exact is false (the default), values are normalized before comparison: letters are converted to uppercase and all diacritical marks (accents, umlauts, etc.) are removed. When exact is true, values are compared exactly as stored, preserving original casing and accent characters.

For most administrative matching scenarios, the default exact: false is appropriate because it tolerates minor formatting inconsistencies. Use exact: true only when the raw stored form is meaningful — for instance, when matching a coded token value that is case-sensitive by specification.


Matcher Algorithms

The following algorithms are used in the matcher property of a matchField. Each returns a boolean pass/fail result.

Phonetic Matchers

Phonetic matchers encode each value as a phonetic key and then compare the keys. Two values match if their keys are identical. This approach tolerates alternate spellings of the same name that sound alike. Phonetic matchers are well-suited to fields containing person names and are commonly used for given name and family name comparisons.

Algorithm Description Match examples Non-match examples
SOUNDEX

The classic Soundex phonetic algorithm, originally designed for English-language names. Soundex encodes each name as a letter followed by three digits. It captures consonant sounds but discards repeated adjacent consonants and ignores vowels after the first character.

Soundex is widely recognized and easy to reason about, but it is coarse-grained: many distinct names share the same code, which can lead to false positives. It does not handle alternate transliterations or names that sound the same but begin with different consonants.

Use Soundex when you need a broadly tolerant match for common English-origin names and can accept occasional false positives.

Jon = John, Roberts = Robarts Thomas ≠ Tom, Smith ≠ Schmidt
REFINED_SOUNDEX

An improved variant of Soundex that assigns a unique digit to each consonant sound, retaining more phonetic information than the original. This produces finer-grained codes that reduce false positives while still tolerating common spelling variants.

Refined Soundex is a drop-in replacement for Soundex when you want better precision. Prefer it over SOUNDEX for new configurations.

Stephens = Stevens Sounds that differ by a single consonant position are more likely to be separated
METAPHONE

Metaphone encodes English-language names using rules for how consonants and vowel sequences are actually pronounced, producing a more phonetically accurate code than Soundex. It handles many common irregularities of English spelling.

Metaphone is more precise than Soundex and generates fewer false positives. It does not handle non-English names or alternate spellings that involve different starting consonants (for example, Smith and Schmidt have different Metaphone codes).

Use Metaphone for English-origin person names when you want better precision than Soundex.

Dury = Durie, Allsop = Alsop Allsop ≠ Allsob, Smith ≠ Schmidt
DOUBLE_METAPHONE

An enhancement of Metaphone that computes two phonetic codes for each value — a primary and an alternate — to account for ambiguous pronunciations and non-English name origins (Slavic, Germanic, Celtic, and others). A match is found if any combination of the primary or alternate codes aligns.

Double Metaphone handles a broader range of name origins and spelling variants than standard Metaphone. It does generate more potential matches than Metaphone, so verify its behaviour on your data before using it in a high-volume environment.

Use Double Metaphone when your data set includes names from multiple linguistic backgrounds.

Dury = Durie, Allsop = Allsob, Smith = Schmidt Names with sufficiently different pronunciation patterns
CAVERPHONE1

The first version of the Caverphone algorithm, developed in New Zealand for matching historical records. It converts names to a six-character code based on pronunciation rules tuned for names of British and European origin as they appear in New Zealand records.

Caverphone1 is more conservative than Caverphone2: it distinguishes between some spellings that Caverphone2 treats as equivalent.

Gail = Gael Gail ≠ Gale, Thomas ≠ Tom
CAVERPHONE2

The second version of Caverphone, which relaxes some of the encoding rules to tolerate a wider range of spelling variations. It uses a ten-character code, giving it more resolution than Caverphone1 while still being more permissive about variant forms.

Use Caverphone2 in preference to Caverphone1 when you expect varied historical spellings of the same name.

Gail = Gael, Gail = Gale Thomas ≠ Tom
NYSIIS

The New York State Identification and Intelligence System (NYSIIS) algorithm was designed for matching names in American public records. It applies a series of prefix and suffix transformations followed by a character-by-character encoding pass to produce a phonetic key.

NYSIIS performs well on Anglicized names and is often used in law enforcement and vital records contexts. It may not handle names of non-English origin as well as Double Metaphone.

Hansen = Hanson, Macintosh = Mackintosh Names with significantly different phonetic structure
COLOGNE

The Cologne phonetic algorithm (Kölner Phonetik) was developed in Germany and is optimized for German-language names. It applies encoding rules that reflect German pronunciation conventions. It assigns a digit to each phonetically significant character and merges adjacent identical digits.

Use Cologne when your data set contains names primarily of German or Central European origin.

Müller = Mueller, Schneider = Snyder Names with different consonant groupings in the German phonetic system
MATCH_RATING_APPROACH

The Match Rating Approach (MRA) is a phonetic algorithm that encodes names and then uses a comparison procedure based on the length of the encoded forms and the number of characters they share. It was developed for applications requiring a balance between recall and precision.

MRA is less commonly used than Soundex, Metaphone, or Double Metaphone. Consider it if you have an existing system that uses MRA and need to match against its output, or if your testing shows it performs better on your specific data set than the other available options.

Smith = Smythe Names that differ substantially in length or consonant structure

String Matchers

String matchers compare values as text strings rather than phonetic encodings. They are useful for fields where exact or near-exact textual agreement is the right criterion — for instance, coded tokens, numeric identifiers, birth dates, or fields where any phonetic substitution would be incorrect.

Algorithm Description Match examples Non-match examples
STRING

Compares two string values for equality. When exact is false (the default), values are first normalized by converting to uppercase and removing diacritical marks, then compared. When exact is true, the original stored values are compared without any normalization.

Use STRING for fields where you expect the values to agree character-for-character, such as a coded administrative gender, a birth date stored as a string, a phone number, or a postal code. It is also suitable for family name matching when phonetic tolerance is not wanted.

MCTAVISH = McTavish (exact: false), Müller = Muller (exact: false) MCTAVISH ≠ McTavish (exact: true), Smith ≠ Smyth
SUBSTRING

Returns a match when one value is a prefix of the other. The comparison is bidirectional: it passes if the left value starts with the right value, or if the right value starts with the left value. Normalization is applied according to the exact parameter before comparison.

Use SUBSTRING when abbreviated forms of a value are acceptable matches — for example, when "Bill" should match "Billy", or when a city name may be entered in truncated form. Be aware that very short values can produce many false positives.

Bill = Billy, 416 = 4169671111 Bert ≠ Egbert (Bert does not start with Egbert, and Egbert does not start with Bert)
NUMERIC

Strips all non-digit characters from both values and then compares the resulting digit-only strings for equality. Formatting characters such as spaces, hyphens, parentheses, and dots are all ignored.

Use NUMERIC for fields such as phone numbers or fax numbers where the same value can be stored in many different formats. Note that this algorithm discards decimal points, so it is not appropriate for matching decimal quantities or amounts.

4169671111 = (416) 967-1111 = 416-967-1111 Different digit sequences regardless of formatting

Date Matcher

Algorithm Description Match examples Non-match examples
DATE

Compares two date or date-time values in a precision-aware manner. It identifies whichever of the two values has the lower precision (for example, year-month versus full date), truncates the more precise value to that precision level, and then compares the two values as strings.

This means that 2019-12 (year-month precision) and 2019-12-19 (day precision) are considered a match, because the less precise value (2019-12) is consistent with the more precise one. Two values with identical precision are compared directly.

Use DATE for birth date fields when source systems may record dates at different levels of precision. Do not use it when you require a strictly identical date, since a year-only value such as 2019 would match any date within that year.

2019-12 = 2019-12-19, 1985 = 1985-07-14 2019-11 ≠ 2019-12, 1985-07-14 ≠ 1985-08-01

Name Matchers

Name matchers operate on FHIR HumanName fields. Unlike the string matchers, which work on a single primitive value, name matchers extract the given and family name components and compare them as a pair. They are designed specifically for the name path of Patient, Practitioner, and RelatedPerson resources.

Algorithm Description Match examples Non-match examples
NAME_ANY_ORDER

Compares a pair of given and family names as strings. Passes if one resource's given name matches the other's given name and both family names match, or if one resource's given name matches the other's family name and vice versa. This allows a match when given and family names have been entered in swapped order.

Use this matcher when your data sources may sometimes record given and family names in the wrong order — a common data quality issue with some legacy systems.

John Henry = Henry John (exact: false) John Henry ≠ John Smith
NAME_FIRST_AND_LAST

Compares a pair of given and family names as strings, requiring the given name to match the given name and the family name to match the family name. The names must be in the correct positions; reversed order does not produce a match.

Use this matcher when you are confident that names will always be recorded in the correct order and you want to avoid false positives that NAME_ANY_ORDER might introduce.

John Henry = John HENRY (exact: false) John Henry ≠ Henry John, John Henry ≠ John Smith
NICKNAME

Checks whether one given name is a recognized nickname or common variant of the other. The system maintains a built-in equivalence table of common English given name relationships (for example, Ken and Kenneth are equivalent, as are Bill and William).

The match is bidirectional: if Ken is a nickname for Kenneth, then both "Ken vs Kenneth" and "Kenneth vs Ken" will match. The comparison is case-insensitive regardless of the exact setting. Spelling variants of different names (such as Allen and Allan) are not treated as equivalent.

Use NICKNAME for given name fields when your data sources are likely to use informal short forms of names.

Ken = Kenneth, Bill = William, Bobby = Robert Allen ≠ Allan, Thomas ≠ Tom (no built-in equivalence)

Identifier Matcher

Algorithm Description Configuration parameters
IDENTIFIER

Compares two FHIR Identifier values. Passes when both the system URI and the value of the two identifiers are identical. An identifier with an empty or absent value never matches.

The optional identifierSystem parameter restricts matching to identifiers that belong to a specific system. When it is set, an identifier whose system does not equal the specified URI cannot match, even if its value happens to be identical to the other resource's identifier value. This is useful when a resource carries identifiers from multiple systems and you only want to match on one of them — for example, a Social Security Number system URI or a specific hospital's MRN namespace.

Use IDENTIFIER when an exact match on a structured identifier is the most reliable signal of identity, and use identifierSystem to avoid accidental matches between identifier values that belong to different namespaces.

  • identifierSystem (optional): URI of the identifier system to restrict matching to. If omitted, all identifier systems are considered.

Example — match only on SSN identifiers:

{
    "name": "identifier-ssn",
    "resourceType": "Patient",
    "resourcePath": "identifier",
    "matcher": {
        "algorithm": "IDENTIFIER",
        "identifierSystem": "http://hl7.org/fhir/sid/us-ssn"
    }
}

Extension Matcher

Algorithm Description
EXTENSION_ANY_ORDER

Matches two FHIR extension lists. Passes when at least one extension in the left resource's list has the same URL and the same value as at least one extension in the right resource's list. The order of extensions within each list does not affect the result.

Use this matcher when you need to compare resources on a FHIR extension field — for example, a custom race/ethnicity code extension, or a jurisdiction-specific identifier stored as an extension.

Empty Field Matcher

Algorithm Description
EMPTY_FIELD

Passes when both field values are absent or empty. Returns false if either field contains any non-empty value.

EMPTY_FIELD is the only matcher that considers empty fields at all; all other matchers silently skip comparisons when a field is absent. It is useful in matchResultMap rules that are intended to apply only when a particular field is not present — for example, to match records that share the same name and birth date but both lack an identifier.


Similarity Algorithms

The following algorithms are used in the similarity property of a matchField. Each returns a decimal score between 0.0 and 1.0, and a matchThreshold parameter is required to translate that score into a pass/fail result. A score greater than or equal to the threshold is a pass.

Similarity algorithms are useful when phonetic encoding is too coarse and exact string matching is too strict — for instance, when you expect minor typographic variations in a field value and want to tune precisely how much variation to tolerate.

String Similarity Algorithms

These algorithms compare the text representations of two values.

Algorithm Description Typical threshold range
JARO_WINKLER

Measures the similarity between two strings based on the number and order of matching characters, with a bonus applied when the strings share a common prefix. The prefix bonus makes it particularly sensitive to the beginnings of strings, which is a useful property for names and surnames because typographic errors are less common at the start of a word.

Jaro-Winkler is the most widely used similarity algorithm for person name matching. It performs well on short strings such as given names and family names.

0.80–0.92 for names; lower values increase recall, higher values increase precision
LEVENSCHTEIN

Computes the normalized Levenshtein edit distance: the minimum number of single-character insertions, deletions, and substitutions needed to transform one string into the other, divided by the length of the longer string. The result is then inverted so that 1.0 represents an identical pair and 0.0 represents the maximum possible edit distance.

Levenshtein treats all positions equally and does not give extra weight to prefixes. It is a good general-purpose choice when you expect scattered typographic errors throughout a string rather than concentrated near one end.

0.80–0.90; very sensitive to string length — shorter strings require a higher threshold to avoid false positives
COSINE

Splits each string into overlapping character n-grams (substrings of length k) and represents each string as a vector of n-gram frequencies. The similarity is the cosine of the angle between the two vectors. A score of 1.0 means the n-gram profiles are identical; 0.0 means they share no n-grams.

Cosine similarity handles transpositions and out-of-order fragments well, making it suitable for longer strings or addresses where word order may vary. It is less sensitive to very short strings.

0.75–0.90
JACCARD

Also operates on character n-grams. Computes the size of the intersection of the two n-gram sets divided by the size of their union. Unlike Cosine, Jaccard does not consider n-gram frequency — it simply asks whether each n-gram is present or absent in each string.

Jaccard is a reasonable alternative to Cosine when you prefer set membership over frequency weighting. It tends to score slightly lower than Cosine for the same pair of strings.

0.70–0.85
SORENSEN_DICE

A coefficient based on the same n-gram intersection concept as Jaccard, but calculated as twice the intersection size divided by the sum of both set sizes. This gives slightly more weight to shared n-grams than Jaccard does. The Sørensen-Dice coefficient is always greater than or equal to the Jaccard index for the same pair of strings.

Sørensen-Dice and Jaccard are closely related. Choose between them based on which threshold value produces better results on your test data.

0.75–0.90

Numeric Similarity Algorithms

These algorithms apply the same similarity measures as the string variants above, but first strip all non-digit characters from both values before computing the score. They are the scored counterparts to the NUMERIC matcher.

Algorithm Description
NUMERIC_JARO_WINKLER Strips non-digit characters from both values, then applies Jaro-Winkler similarity to the resulting digit strings. Useful for comparing phone numbers or other numeric identifiers where you want to tolerate a small number of digit errors or transpositions rather than requiring an exact match.
NUMERIC_LEVENSCHTEIN Strips non-digit characters from both values, then applies normalized Levenshtein distance to the resulting digit strings.
NUMERIC_COSINE Strips non-digit characters from both values, then applies Cosine n-gram similarity to the resulting digit strings.
NUMERIC_JACCARD Strips non-digit characters from both values, then applies Jaccard n-gram similarity to the resulting digit strings.
NUMERIC_SORENSEN_DICE Strips non-digit characters from both values, then applies Sørensen-Dice coefficient similarity to the resulting digit strings.

Choosing an Algorithm

The right algorithm depends on the field being compared, the quality of your data, and the degree of tolerance you need.

For person name fields: Use a phonetic matcher such as METAPHONE or DOUBLE_METAPHONE for family names and given names when you expect spelling variations caused by transcription or transliteration differences. Use DOUBLE_METAPHONE rather than METAPHONE if your population includes names from non-English-speaking backgrounds. Use NICKNAME for given names if informal short forms are common in your data. Use a similarity algorithm such as JARO_WINKLER when you want fine-grained control over how much spelling variation to tolerate, particularly for generating POSSIBLE_MATCH results that will be reviewed by a data steward.

For identifiers and coded fields: Use IDENTIFIER for FHIR Identifier fields, which carry both a system URI and a value. Use STRING with exact: false for coded token fields (such as administrative gender) when you want case-insensitive equality. Use STRING with exact: true when the field value is case-sensitive.

For phone numbers and other digit-heavy strings: Use NUMERIC when you require an exact match on the digit sequence regardless of formatting. Use a NUMERIC_* similarity algorithm when you want to tolerate transpositions or a small number of digit errors.

For birth dates: Use DATE when source systems may record dates at different precision levels. Use STRING when all systems record dates at the same precision and you want a strict equality check.

For address and free-text fields: Use a similarity algorithm such as JARO_WINKLER, COSINE, or LEVENSCHTEIN and tune the matchThreshold using representative test data. Address fields benefit from similarity algorithms because they are prone to abbreviations, transpositions, and inconsistent punctuation that phonetic matchers are not designed to handle.

If none of the built-in algorithms meet your requirements, you can implement and deploy your own — see Custom MDM Matching Algorithms.