MegaScale is a mechanism for storing virtually unlimited amounts of data in a single FHIR server. It uses multiple database instances to create discrete pools of data which are logically separate, but are managed under a single Smile CDR FHIR Storage (RDBMS) module.
In its simplest terms, a MegaScale-enabled server can be thought of as a partitioned FHIR repository where individual partitions or groups of partitions are stored in separate database schemas, and potentially in separate physical database instances.
Using this strategy can be helpful in cases such as:
In MegaScale mode, one or more FHIR Endpoint modules are combined with a single FHIR Storage (RDBMS) module. Incoming FHIR requests include a tenant identifier which maps to a particular partition, which then specifies the target database. This architecture is shown in the diagram below.
In a MegaScale architecture, the dual concepts of Partitions and Shards are used. These two terms mean related but different things.
A Partition is a single grouping of resources. Any individual resource must be assigned to a single partition, and that partition will generally contain multiple resources.
One or more partitions are assigned to a given database schema. This grouping of Partitions to a single database schema is called a Shard.
The following diagram shows a potential mapping of the 15000 partitions defined in Patient ID Partition Mode to 3 shards. This is only one potential mapping however; it is possible to have fewer or more shards depending on anticipated storage and scaling requirements.
See MegaScale Patient ID Partition Selection Modes for details on how to use these partition selection modes with MegaScale.
MegaScale creates an architecture where different partitions are stored on different shards (see Partitions and Shards above). This has several implications to the semantics and operation of FHIR Transaction processing, but does not mean that FHIR transactions can not be used even if they span multiple shards.
When loading data using a FHIR Transaction Bundle, the system will automatically attempt to respect the semantics of the FHIR transaction as much as possible, but will make compromises where necessary if a transaction needs to span multiple shards.
When a transaction Bundle needs to write to multiple shards, it will be automatically split into multiple discrete FHIR Transaction Bundles and executed in sequence. The server will order these bundles according to resource dependencies within the Bundle, and will use the outcome of earlier bundles to inform the processing of later Bundles.
For example, suppose you are have configured your Partition Selection Mode to Patient ID Partition Selection Mode, with your Ancillary Resources on a separate MegaScale database from our Patient Resources. In this example, you might have Patient and Encounter resource referencing Organization and Location resources in the same FHIR Transaction Bundle.
In this scenario, the server will automatically process the Ancillary resources first. Any newly assigned resource IDs will be used in references from the subsequent Patient and Encounter resources.
As a result, it is not possible to have circular dependencies in FHIR Transaction Bundles executed on a MegaScale server where the cycle crosses shard boundaries. For example, if your Patient and Ancillary data are on separate shards, attempting to process a FHIR Transaction Bundle with a reference from a Patient to an Organization where the Organization also holds a reference to the Patient would result in an error.
FHIR searches may span multiple shards (see Partitions and Shards above). When a search is initiated and the chosen Partitioning Selection Mode indicates that multiple partitions should be searched, the system will begin a serial (one at a time) search across all selected partitions.
This is significantly less efficient than a search against a single partition or shard, and should be avoided wherever possible in favor of single-shard searches. The primary use case is for identification of individual resources (e.g. a conditional update on a resource by identifier where the partition is not known).
When searching multiple shards, the _sort parameter may not be used. Multi-shard searches attempting to use this parameter will result in an error.
This section lists the known limitations on this feature.
The following FHIR interactions have been tested:
/P1/$reindex), or all partitions using _ALL as the tenant name (e.g. POST /_ALL/$reindex).No other features, operations, or interactions have been tested or are expected to work with MegaScale.
You must ensure that all updates within a single Bundle target a single MegaScale database. This is true for REQUEST_TENANT partitioning mode, but may not be true for other partition modes like Patient-Id partitioning or custom partitioning solutions.
Search requests will only include results from a single database.
To enable MegaScale mode, the following settings must be set.
On the FHIR Storage (RDBMS) module:
true.DEFAULT partition.REQUEST_TENANT, REQUEST_HEADER, BUCKETED_PATIENT_ID, or PATIENT_ID.true.On the FHIR Endpoint module:
REQUEST_TENANT partition selection mode, Tenant Identification Strategy must be set to URL_BASED.When a resource is written that contains a reference whose target lives on a different shard (i.e. a different database), the write must open a nested transaction against the target shard to resolve the reference. Under high concurrent ingestion, a burst of such cross-shard resolutions can exhaust a target shard's connection pool while every borrower is also holding a connection in its source shard — a classic hold-and-wait deadlock across pools.
(Cross-partition references that resolve within the same shard are not affected by this — they reuse the source connection and do not contend for a second pool.)
To prevent the cross-shard cycle, MegaScale caps the number of cross-shard reference resolutions that may be in flight at any one time:
db.connectionpool.maxtotal - 4 if unset.The default works for most deployments. Tune it deliberately if the defaults do not match your topology.
db.connectionpool.maxtotal. Sizing the cap at or above pool capacity defeats its purpose — the cycle becomes possible again. Leave at least a few slots of headroom in each pool for the outer write that holds them.db_num_idle metric) before raising it.Two runtime metrics expose the gate's state. They are emitted by every FHIR Storage module instance that has MegaScale enabled, both through the Codahale MetricRegistry (visible on the cluster manager's metrics endpoints) and through OpenTelemetry.
| Metric | OpenTelemetry name | Type | What it means |
|---|---|---|---|
megascale_crossreference_queue_length | smilecdr.storage.megascale.crossreference.queue_length | Gauge | Number of ingest threads currently waiting for a permit. Zero in healthy operation. Sustained non-zero readings mean cross-shard demand exceeds the configured cap. |
megascale_crossreference_timeout_count | smilecdr.storage.megascale.crossreference.timeout_count | Cumulative counter | Total cross-shard resolutions that have failed because they could not acquire a permit before the acquire timeout. The rate of change of this counter (e.g. rate(megascale_crossreference_timeout_count[5m])) is the actionable signal — a non-zero rate means writes are failing right now. |
Recommended alerts and dashboards:
megascale_crossreference_timeout_count over the last few minutes — writes are being rejected.megascale_crossreference_queue_length sustained above zero for several minutes — saturation is imminent even if no writes have failed yet.db_num_active and db_num_idle for each shard's pool, so you can correlate gate saturation with actual pool exhaustion.MegaScale connection details are supplied using a Java Smile CDR Interceptor using the STORAGE_MEGASCALE_PROVIDE_DB_INFO pointcut.
See Example: MegaScale Connection Provider to see how this pointcut can be used. This example is also available in the Interceptor Starter Project.
You are about to leave the Smile Digital Health documentation and navigate to the Open Source HAPI-FHIR Documentation.