MegaScale
LMA

 
MegaScale has limitations, and will not be suitable for every use case. See Limitations below.

MegaScale is a mechanism for storing virtually unlimited amounts of data in a single FHIR server. It uses multiple database instances to create discrete pools of data which are logically separate, but are managed under a single Smile CDR FHIR Storage (RDBMS) module.

In its simplest terms, a MegaScale-enabled server can be thought of as a partitioned FHIR repository where individual partitions or groups of partitions are stored in separate database schemas, and potentially in separate physical database instances.

Using this strategy can be helpful in cases such as:

  • The total amount of data to be stored exceeds the natural limit of a single instance of the database technology being used
  • Data needs to be kept physically separate due to privacy or legal rules.

Architecture
LMA

 

In MegaScale mode, one or more FHIR Endpoint modules are combined with a single FHIR Storage (RDBMS) module. Incoming FHIR requests include a tenant identifier which maps to a particular partition, which then specifies the target database. This architecture is shown in the diagram below.

MegaScale Architecture

Partitions and Shards

 

In a MegaScale architecture, the dual concepts of Partitions and Shards are used. These two terms mean related but different things.

A Partition is a single grouping of resources. Any individual resource must be assigned to a single partition, and that partition will generally contain multiple resources.

One or more partitions are assigned to a given database schema. This grouping of Partitions to a single database schema is called a Shard.

The following diagram shows a potential mapping of the 15000 partitions defined in Patient ID Partition Mode to 3 shards. This is only one potential mapping however; it is possible to have fewer or more shards depending on anticipated storage and scaling requirements.

MegaScale Shards and Partitions

Patient ID Partition Selection Mode

 

See MegaScale Patient ID Partition Selection Modes for details on how to use these partition selection modes with MegaScale.

FHIR Transactions Spanning Multiple Shards

 

MegaScale creates an architecture where different partitions are stored on different shards (see Partitions and Shards above). This has several implications to the semantics and operation of FHIR Transaction processing, but does not mean that FHIR transactions can not be used even if they span multiple shards.

When loading data using a FHIR Transaction Bundle, the system will automatically attempt to respect the semantics of the FHIR transaction as much as possible, but will make compromises where necessary if a transaction needs to span multiple shards.

When a transaction Bundle needs to write to multiple shards, it will be automatically split into multiple discrete FHIR Transaction Bundles and executed in sequence. The server will order these bundles according to resource dependencies within the Bundle, and will use the outcome of earlier bundles to inform the processing of later Bundles.

For example, suppose you are have configured your Partition Selection Mode to Patient ID Partition Selection Mode, with your Ancillary Resources on a separate MegaScale database from our Patient Resources. In this example, you might have Patient and Encounter resource referencing Organization and Location resources in the same FHIR Transaction Bundle.

In this scenario, the server will automatically process the Ancillary resources first. Any newly assigned resource IDs will be used in references from the subsequent Patient and Encounter resources.

As a result, it is not possible to have circular dependencies in FHIR Transaction Bundles executed on a MegaScale server where the cycle crosses shard boundaries. For example, if your Patient and Ancillary data are on separate shards, attempting to process a FHIR Transaction Bundle with a reference from a Patient to an Organization where the Organization also holds a reference to the Patient would result in an error.

FHIR Searches Spanning Multiple Shards

 

FHIR searches may span multiple shards (see Partitions and Shards above). When a search is initiated and the chosen Partitioning Selection Mode indicates that multiple partitions should be searched, the system will begin a serial (one at a time) search across all selected partitions.

This is significantly less efficient than a search against a single partition or shard, and should be avoided wherever possible in favor of single-shard searches. The primary use case is for identification of individual resources (e.g. a conditional update on a resource by identifier where the partition is not known).

When searching multiple shards, the _sort parameter may not be used. Multi-shard searches attempting to use this parameter will result in an error.

Limitations
LMA

 

This section lists the known limitations on this feature.

FHIR Interactions

The following FHIR interactions have been tested:

  • Create/Update
  • Search (limitations)
  • $reindex – It is possible to reindex a single partition by providing the tenant name in the URL (e.g. /P1/$reindex), or all partitions using _ALL as the tenant name (e.g. POST /_ALL/$reindex).
  • $validate
  • $expunge – Expunge everything is verified to work, and will only expunge everything for a single MegaScale database at a time.
  • $delete-expunge – Delete expunge is verified to work, and will only delete and expunge for a single MegaScale database at a time.
  • FHIR Transactions (limitations)

No other features, operations, or interactions have been tested or are expected to work with MegaScale.

You must ensure that all updates within a single Bundle target a single MegaScale database. This is true for REQUEST_TENANT partitioning mode, but may not be true for other partition modes like Patient-Id partitioning or custom partitioning solutions.

Cross-Partition Searching

Search requests will only include results from a single database.

Configuration
LMA

 

To enable MegaScale mode, the following settings must be set.

On the FHIR Storage (RDBMS) module:

On the FHIR Endpoint module:

Cross-Partition Reference Concurrency

When a resource is written that contains a reference whose target lives on a different shard (i.e. a different database), the write must open a nested transaction against the target shard to resolve the reference. Under high concurrent ingestion, a burst of such cross-shard resolutions can exhaust a target shard's connection pool while every borrower is also holding a connection in its source shard — a classic hold-and-wait deadlock across pools.

(Cross-partition references that resolve within the same shard are not affected by this — they reuse the source connection and do not contend for a second pool.)

To prevent the cross-shard cycle, MegaScale caps the number of cross-shard reference resolutions that may be in flight at any one time:

Tuning Guidance

The default works for most deployments. Tune it deliberately if the defaults do not match your topology.

  • Pick a value strictly below the smallest target shard's db.connectionpool.maxtotal. Sizing the cap at or above pool capacity defeats its purpose — the cycle becomes possible again. Leave at least a few slots of headroom in each pool for the outer write that holds them.
  • Symmetric pools across shards keep tuning simple. If shards have different pool sizes, size the cap against the smallest pool you write into, not the average.
  • Raise it cautiously when throughput is the bottleneck. A higher cap permits more concurrent cross-shard work but consumes more target-pool connections per second. Verify the target shards have spare pool capacity (db_num_idle metric) before raising it.
  • Lower it (down to 1, if needed) for diagnosis. If you suspect cross-shard contention is degrading throughput, drop the cap aggressively and watch the queue length and timeout count. Persistent saturation at low caps points to ingest concurrency that exceeds your provisioned cross-shard capacity.
  • The acquire timeout should be generous in steady state. If a write waits longer than the timeout, it fails with HAPI-2701 and the upstream consumer (Kafka, Camel, etc.) is expected to redeliver. A short timeout will make transient bursts look like errors; a long one will mask real saturation. The 60-second default is intended to absorb normal bursts while still surfacing sustained overload.

Observing Saturation

Two runtime metrics expose the gate's state. They are emitted by every FHIR Storage module instance that has MegaScale enabled, both through the Codahale MetricRegistry (visible on the cluster manager's metrics endpoints) and through OpenTelemetry.

MetricOpenTelemetry nameTypeWhat it means
megascale_crossreference_queue_lengthsmilecdr.storage.megascale.crossreference.queue_lengthGaugeNumber of ingest threads currently waiting for a permit. Zero in healthy operation. Sustained non-zero readings mean cross-shard demand exceeds the configured cap.
megascale_crossreference_timeout_countsmilecdr.storage.megascale.crossreference.timeout_countCumulative counterTotal cross-shard resolutions that have failed because they could not acquire a permit before the acquire timeout. The rate of change of this counter (e.g. rate(megascale_crossreference_timeout_count[5m])) is the actionable signal — a non-zero rate means writes are failing right now.

Recommended alerts and dashboards:

  • Alert on a non-zero rate of megascale_crossreference_timeout_count over the last few minutes — writes are being rejected.
  • Alert on megascale_crossreference_queue_length sustained above zero for several minutes — saturation is imminent even if no writes have failed yet.
  • Dashboard showing queue length alongside db_num_active and db_num_idle for each shard's pool, so you can correlate gate saturation with actual pool exhaustion.

Connection Provider Interceptor
LMA

 

MegaScale connection details are supplied using a Java Smile CDR Interceptor using the STORAGE_MEGASCALE_PROVIDE_DB_INFO pointcut.

See Example: MegaScale Connection Provider to see how this pointcut can be used. This example is also available in the Interceptor Starter Project.