The Delta Lake Export module extends Realtime Export to write FHIR resource changes to Delta Lake tables. This enables customers on data lake architectures (Databricks, Spark, cloud object stores) to receive FHIR data directly, without building custom ETL pipelines.
The module connects to a customer-managed Spark Connect server over gRPC and writes using version-guarded MERGE / UPDATE / DELETE: the table holds one row per resource, replaced in place on each update and physically removed on delete. Consumers query the table directly with a plain SELECT.
Delta Lake Export is configured as a separate module type (REALTIME_EXPORT_DELTA_LAKE) with its own configuration. It shares the same rules model as the JDBC-based export, but writes to Delta Lake instead of a database. Each module owns its own producer interceptor and publishes to a module-specific broker channel, so the two can operate independently with no shared state or coupled configuration.
The module connects to an external Spark Connect server (the customer's Spark cluster) via gRPC.
FHIR Storage Module
|
v
Interceptor (one per export module, CREATE/UPDATE/DELETE hooks)
| |
v v
Channel Channel
"realtime.export" "realtime.export.deltalake"
| |
v v
JDBC Realtime Export Delta Lake Export
|
v
Spark Connect (gRPC)
|
v
Customer-managed Spark cluster
(Docker, k8s, Databricks, EMR)
|
v
Delta Lake Tables
(Parquet + _delta_log)
The customer provides a Spark Connect server. Options include:
apache/spark:4.0.0 with the delta-connect-server plugin (see example compose file in the test resources)CDR connects via the deltalake.spark_connect_url configuration property (e.g. sc://spark-host:15002).
Each FHIR resource operation is written with a version-guarded Delta MERGE:
MERGE statement (shown below). Only updates whose source version is strictly greater than the target's replace the existing row, so out-of-order delivery is tolerated — stale messages are dropped.DELETE on the row's id. Deletes on non-existent tables are no-ops.MERGE INTO target t USING source s ON t.id = s.id
WHEN MATCHED AND t.version < s.version THEN UPDATE SET *
WHEN NOT MATCHED THEN INSERT *
Every parent row carries two system columns:
| Column | Type | Description |
|---|---|---|
id | STRING | Unqualified, versionless resource ID (e.g., Patient/123) |
version | INT | FHIR resource version |
Consumers can query current state with a plain SELECT:
SELECT * FROM patient_table WHERE id = 'Patient/123'
The trade-off is no built-in history on the primary table; consumers that need an audit trail should enable retainAllHistory (separate _history tables) or capture history upstream.
The Delta Lake Export module reuses the same rules JSON format as the JDBC Realtime Export module. See Realtime Export Rules Definition for details on configuring transformers, columns, and FHIRPath expressions.
module.realtime_export_delta_lake.config.deltalake.path=/path/to/delta/tables
module.realtime_export_delta_lake.config.deltalake.spark_connect_url=sc://spark-host:15002
module.realtime_export_delta_lake.config.script.text={"transformers": [...]}
module.realtime_export_delta_lake.config.script.file=/path/to/rules.json
module.realtime_export_delta_lake.config.channel.concurrent_consumers=1
module.realtime_export_delta_lake.config.channel.concurrent_retry_consumers=1
module.realtime_export_delta_lake.config.channel.prefix=
| Property | Description | Default |
|---|---|---|
deltalake.path | Path where Delta tables are created (local filesystem, s3a://, abfss://, gs:// — depends on Spark cluster configuration) | (required) |
deltalake.spark_connect_url | Spark Connect server URL (e.g. sc://spark-host:15002) | (required) |
script.text | Inline JSON rules configuration | Empty rules |
script.file | Path to external JSON rules file | (none) |
channel.concurrent_consumers | Number of concurrent message consumers | 1 |
channel.concurrent_retry_consumers | Number of concurrent retry consumers | 1 |
channel.prefix | Channel name prefix | (empty) |
Tables are created automatically on first write. The table schema is derived from the rules configuration, with system columns added automatically. New tables are created with auto-optimize TBLPROPERTIES enabled (delta.autoOptimize.optimizeWrite=true, delta.autoOptimize.autoCompact=true) so write batches are compacted automatically.
Repeating FHIR elements (e.g., Patient.name, Patient.address) are configured as child tables in the rules JSON. The Delta Lake export module supports two modes for handling child tables, controlled by the flatten property on the child table configuration:
flatten | Behavior |
|---|---|
true | The child elements are embedded in the parent row as a nested ARRAY<STRUCT<...>> Parquet column. Single Delta table per resource. |
false (default) | Each child element is written as a row in a separate Delta table with link columns. |
flatten: true)When flatten is true, the child table's tableName becomes the name of a nested array column in the parent table. The child transformer's columns become the fields of the struct inside the array.
{
"transformers": [
{
"resourceType": "Patient",
"tableName": "patient_table",
"columns": [
{"columnName": "is_active", "fhirPath": "active", "columnType": "BOOLEAN"}
],
"childTables": [
{
"fhirPath": "Patient.name",
"tableName": "names",
"flatten": true,
"childTransformer": {
"columns": [
{"columnName": "family", "fhirPath": "family", "columnType": "STRING"},
{"columnName": "use", "fhirPath": "use", "columnType": "STRING"}
]
}
}
]
}
]
}
The resulting patient_table schema includes a names column of type ARRAY<STRUCT<family: STRING, use: STRING>>. Spark consumers can query it via EXPLODE:
SELECT id, n.family, n.use
FROM patient_table
LATERAL VIEW EXPLODE(names) AS n
WHERE n.use = 'official'
Recursive nesting is supported — child tables can have their own flattened child tables.
flatten: false)When flatten is false (the default), each child element becomes a row in its own Delta table. These child tables include three link columns:
| Column | Description |
|---|---|
id | Auto-generated UUID for the child row |
parent_reference | The ID of the parent row this child belongs to |
source_resource_id | The unqualified ID of the top-level FHIR resource |
Non-flattened child tables use replace-children semantics: on every parent create or update, all existing child rows for the resource are deleted, then the new child rows are inserted. On parent delete, child rows are cascade-deleted before the parent row is removed.
Schema evolution is automatic. When the rules configuration adds a new column that is not present in the live Delta table, the writer runs ALTER TABLE ... ADD COLUMNS automatically. Existing rows get null for the new column.
| Schema change | Behavior |
|---|---|
| New column added to rules | Automatic ALTER TABLE ADD COLUMNS. Old rows get null. |
| Column removed from rules | Column persists in the table; new rows write null. No action needed. |
| Type change on an existing column | Fails fast with a clear error naming the column and both types. Operator must resolve manually. |
Schema comparison happens once on the first write to each table, then the resolved schema is cached for the lifetime of the module instance.
The following table shows how RTE column types map to Delta Lake (Parquet) types:
| RTE Column Type | Delta Lake Type | Notes |
|---|---|---|
STRING | StringType | |
INT | IntegerType | |
LONG | LongType | |
BOOLEAN | BooleanType | |
FLOAT | FloatType | |
DOUBLE | DoubleType | |
DATE_ONLY | DateType | Stored as epoch day |
DATE_TIMESTAMP | TimestampType | Stored as microseconds since epoch |
BLOB | BinaryType | |
CLOB | StringType | Mapped to string in Delta Lake |
| Aspect | JDBC Export | Delta Lake |
|---|---|---|
| Write pattern | INSERT/UPDATE/DELETE (mutable) | MERGE / UPDATE / DELETE (mutable, current state) |
| Current state | Always up to date in table | Always up to date in table |
| History | Optional (retainAllHistory) | Optional (retainAllHistory) — separate _history tables |
| History tables | Separate _history tables | Separate _history tables (when retainAllHistory: true) |
| Transaction atomicity | Database transaction | Per-row only |
| Schema creation | Manual (pre-create tables) | Automatic on first write |
| Schema evolution (add column) | Manual ALTER TABLE | Automatic ALTER TABLE ADD COLUMNS |
| Non-flattened child tables | Supported | Supported (replace-children) |
| Process footprint in CDR | JDBC driver only | Spark Connect client (~33 MB) |
| External dependency | Database | Spark Connect server |
| Target | JDBC databases | Delta Lake (Parquet + transaction log) |
You are about to leave the Smile Digital Health documentation and navigate to the Open Source HAPI-FHIR Documentation.