This page provides technical guidance for scaling Smile CDR’s relational storage to very large patient populations (50M+).
The primary driver of the database scaling discussion is estimated data volume.
If the estimated data volume will fit in a single database system, we recommend the simplicity of a single database.
If the data exceeds the practical limits of a single database, we then ask if the data represents several independent lines of business, or is a single homogeneous pool of patients.
REQUEST_TENANT mode.The following diagram summarizes the decision flow for choosing a scaling strategy.
flowchart LR
A[Estimate total db size] --> B{Will the data fit in a single db?}
B -- yes --> C[Deploy to a single database]
B -- no --> D{Can the data be segregated?}
D -- yes, multi-tenant --> E[Megascale Multi-Tenant]
D -- no, single-tenant --> F[Use Patient Id partitioning]
F --> G{Which Cloud?}
G -- AWS --> H[Use Patient Id partitioning with Aurora Limitless]
G -- Azure --> I[Use Patient Id partitioning with Azure/Citus]
G -- Other --> J[Talk to Us]
style A fill:#e3f2fd,stroke:#1565c0,color:#1565c0
style B fill:#fff9c4,stroke:#f9a825,color:#5d4037
style C fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32
style D fill:#fff9c4,stroke:#f9a825,color:#5d4037
style E fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32
style F fill:#e3f2fd,stroke:#1565c0,color:#1565c0
style G fill:#fff9c4,stroke:#f9a825,color:#5d4037
style H fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32
style I fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32
style J fill:#fce4ec,stroke:#c62828,color:#c62828
click C "/docs/fhir_storage_relational/fhir_storage_relational_module.html" "RDBMS FHIR persistence module" _blank
click E "/docs/fhir_storage_relational/megascale.html" "MegaScale (REQUEST_TENANT) multi-tenant scaling" _blank
click H "/docs/fhir_repository/patient_id_partitioning.html#pid-part-mode" "Patient ID partitioning on AWS Aurora Limitless" _blank
click I "/docs/fhir_repository/patient_id_partitioning.html#pid-part-mode" "Patient ID partitioning on Azure/Citus Elastic Clusters" _blank
click J "https://www.smiledigitalhealth.com/contact-us" "Contact Smile Digital Health" _blank
The primary challenge in scaling Smile CDR to 50 million patients and beyond is scaling the storage in the RDBMS. Cloud vendors have maximum sizes, the engines have maximum table sizes, and backup systems start to struggle once sizes cross 10s of TiB. We offer our MegaScale solution for combining multiple databases into a single FHIR persistence system. We also support database scaling technologies offered by the cloud vendors (AWS Aurora Limitless, and Azure/Citus Postgres).
| System | Limits |
|---|---|
| Postgres | Max 32 TB per table |
| AWS RDS for PostgreSQL | 64 TiB per database |
| AWS Aurora Serverless/Provisioned | 256 TiB per database |
| Azure PostgreSQL Flexible Server | 64 TB per database |
| Azure SQL DB Hyperscale | 128 TB per database |
| Azure PostgreSQL Elastic Cluster (Citus) | Unlimited (new in 2025.11) |
| AWS Aurora Limitless | Unlimited (new in 2025.11) |
| Oracle RAC (on-prem) | Unlimited |
The major database engines (Postgres, SQL Server, and Oracle) support partitioning technologies to support large scale systems. These allow partitioning tables, index structures, and storage. Smile CDR implements different application partitioning models to allocate FHIR data to table partitions, allowing the system to leverage this database technology.
Table partitioning technology has grown beyond single-database scale. Several vendors have developed clustering extensions based on table partitioning. In these technologies, the table sub-partitions can be distributed to separate database servers, allowing horizontal scaling of the database system while presenting a single SQL connection for applications. Oracle Database was a leader with RAC on version 10g. Then Citus added horizontal scaling extensions to Postgres. Citus was then acquired by Microsoft, and integrated with Azure. The Postgres Citus extensions have been offered as Azure Cosmos DB for Postgres, and the newer Elastic Clusters in Azure Database for PostgreSQL. We support all three variants: Citus open-source Postgres extensions, Azure Cosmos DB for Postgres, and Elastic Clusters in Azure Database for PostgreSQL. For simplicity, we will refer to these three implementations of Citus as Azure/Citus. AWS has developed their own horizontal Postgres extensions in AWS Aurora Limitless. We support all three vendor offerings.
Smile CDR implements two main partitioning models to leverage these database technologies: multi-tenant, and per-patient partitioning. In these modes, the system assigns every FHIR resource to a numeric partition. The database engine can then use these assignments to allocate rows to specific table partitions.
Multi-tenant partitioning (i.e. REQUEST_TENANT mode) is suited to systems holding data for several separate businesses, where the data within each tenant is completely segregated. This is well suited to organizations that may run several programs, each independent of the others.
Scaling Multi-Tenant
For systems with multiple tenants where each tenant can fit in a single database, our MegaScale solution works well. Writes and queries are routed to the host database for each tenant transparently. This model works well with standard databases, and is also compatible with partitioning within the database, and can be used with Oracle RAC, AWS Aurora Limitless, and Azure/Citus to scale to arbitrary sizes.
Patient ID partitioning allocates each patient to one of 15000 separate partitions. Resources in the FHIR Patient Compartment (e.g. Observation, ExplanationOfBenefit, etc.) associated with the patient will be stored with it in the patient partition. This system is appropriate for customers with a single, large patient population.
Scaling Patient ID Partitioning
The simplest way to scale patient ID partitioning beyond a single database is via a clustered vendor offering: AWS Aurora Limitless, Azure/Citus, or Oracle RAC. We have recently extended MegaScale to support patient ID partitioning. This support is new in 2025.11, and is still experimental. Performance testing of these solutions with patient ID partitioning is ongoing.
If the estimated data volume will fit in a single database system, we recommend the simplicity of a single database used with our RDBMS FHIR persistence module.
If the estimated data volume exceeds the platform limits, then we recommend either MegaScale, or one of the unlimited engine offerings (Oracle RAC, AWS Aurora Limitless, Azure/Citus). The choice between them is driven by the partitioning scheme: multi-tenant or per-patient.
You are about to leave the Smile Digital Health documentation and navigate to the Open Source HAPI-FHIR Documentation.