Monitoring Basics

 

This page outlines methods for monitoring Smile CDR.

Smile CDR provides several hooks that are suitable for integration with an external monitoring system.

HTTP Endpoint Health Check

 
The Endpoint Health Check is a legacy endpoint. For Kubernetes health probes and for monitoring a whole Smile CDR process, use the Process Status Endpoints below, which replace it.

All HTTP Servers supply an HTTP listener at the path /endpoint-health that can be used to query for the current availability of the endpoint. This endpoint will return an HTTP 200 with a simple JSON response payload as long at the endpoint is enabled and operational.

This is useful for systems requiring a simple way to test whether an endpoint is available, such as load balancers that perform automated failover based on a status test. The Endpoint Health Check is very lightweight and may be called frequently without causing extra load on the server.

The Endpoint Healh Check is available on any endpoint modules that supply an HTTP server (e.g. FHIR REST Endpoint or Web Admin Console. It will typically be served at /endpoint-health, but note that it will respect the context path if one is set. So for example if the context path is set to /admin, the Endpoint Health Check will be served at /admin/endpoint-health.

The following example shows a simple HTTP query against the Endpoint Health Check.

GET /endpoint-health

The server will respond with the following response:

200 OK
Content-type: application/json

{"status":"OPERATIONAL"}

If any health checks are failing on the given module, the server will return a response which includes the string NOT_OPERATIONAL as well as the IDs of any failing health checks. For example:

200 OK
Content-type: application/json

{
  "status" : "NOT_OPERATIONAL",
  "failingHealthChecks" : [ "connection_pool" ]
}

By default, the endpoint will return an HTTP 200 OK regardless of whether any health checks are failing. A different status code can be returned by adjusting the Unhealthy Status Code setting.

Process Status Endpoints

 

Every Smile CDR process can serve three endpoints on a dedicated status port, designed for Kubernetes startup, readiness and liveness probes. They are plain HTTP, so any monitoring tool can also poll them. They are off by default.

EndpointAnswers200503
/startupHas the process finished starting?StartedStill starting
/readinessCan the process serve traffic?Every health check is passingStarting, shutting down, a module failed to start, or a health check is failing
/livenessDoes the process need a restart?In every other caseA liveness check is failing

Unlike /endpoint-health, these endpoints:

  • cover the whole process, not one module's HTTP listener;
  • answer from early in startup, before any module has started;
  • need no login and have their own port, outside any module's HTTP server, so a module's context path does not apply.

Because they need no login, restrict the status port with a firewall rule, security group or network policy so that only the systems that probe it can reach it. This is also why the endpoints stay off until you configure them.

Configure them in the node configuration properties file, alongside node.id:

  • node.status.port turns the endpoints on and sets the port. As with every server port on the node, node.server_port_offset is added to it, and the result must be between 1 and 65535.
  • node.status.bind_address limits which network interface the endpoints listen on. The default, every interface, is what Kubernetes needs; use 127.0.0.1 if you only poll from the same host.
  • node.status.health_check_timeout_seconds is how long /readiness and /liveness wait for the health checks to finish: 1 to 600, default 5. It covers all the checks in the process together, because they run one after another. If a node reports healthCheckTimeout under normal load, raise it, and keep your probe's own timeout above it (Kubernetes defaults timeoutSeconds to 1 second); see Kubernetes Health Probes.
node.status.port=9300

Response payloads

Responses are JSON. Fields with no value are left out.

/startup returns STARTING (HTTP 503) while the process starts, then HTTP 200 and:

{
  "status" : "STARTED"
}

/readiness returns READY (HTTP 200) or NOT_READY (HTTP 503), with the node id (node.id from your configuration) and the process id (the name the process took when it joined the cluster, absent until then):

{
  "status" : "READY",
  "nodeId" : "Master",
  "processId" : "Algonquin"
}

When the process is not ready, failingHealthChecks lists each failing check by module (the module id from your configuration) and healthCheck name:

{
  "status" : "NOT_READY",
  "nodeId" : "Master",
  "processId" : "Algonquin",
  "failingHealthChecks" : [ {
    "module" : "persistence",
    "healthCheck" : "connection_pool"
  }, {
    "module" : "fhir_endpoint",
    "healthCheck" : "dependent_fhir_storage_health"
  } ]
}

A module that failed to start is listed under its own id as failedToStart. Conditions of the process itself use the module id _process: starting while it starts (no module checks run until then) and shuttingDown as soon as shutdown begins.

/liveness returns ALIVE (HTTP 200) or NOT_ALIVE (HTTP 503), with the same fields:

{
  "status" : "NOT_ALIVE",
  "nodeId" : "Master",
  "processId" : "Algonquin",
  "failingHealthChecks" : [ {
    "module" : "clustermgr",
    "healthCheck" : "memory_pressure"
  } ]
}

What /liveness reports

/liveness answers one question: does this process need to be restarted? It counts only liveness checks, the health checks that only a restart can clear, and ignores all others. An unreachable database therefore never restarts the process; /readiness takes it out of service instead.

A standard installation has one liveness check, memory_pressure, reported by the Cluster Manager. It fails when the process is so short of memory that it spends most of its time on garbage collection (reclaiming memory) instead of doing useful work. It measures once a minute, and fails after five measurements in a row where:

  • at least half of the minute was spent on garbage collection, and
  • more than 90% of the maximum memory (heap) was still in use afterwards.

Memory is measured after garbage collection because it is always high just before one, even on a healthy process. Each measurement that meets both conditions is logged as a WARN in the system log, and the failure as an ERROR.

Once memory_pressure fails, it does not clear until the process restarts. It also makes /readiness report NOT_READY, so the process stops receiving traffic until then. For a process that runs out of memory completely, see JVM Memory Settings.

In every other case /liveness reports ALIVE, including while the process is starting or shutting down, and when the liveness checks cannot finish within node.status.health_check_timeout_seconds or a module is restarting (the reason is logged). /readiness reports NOT_READY in those cases.

Reading a failing /readiness check

When a pod is not ready, or a monitor reports 503:

  1. Call /readiness and read failingHealthChecks.
  2. Look up each healthCheck in the table below.
  3. For the full error message, which the status port never shows, use the authenticated Runtime Status Health Checks operation. For a module that failed to start, see the startup log.

Results are reused for one second, so a brief failure such as a short database outage clears as soon as its cause does. During shutdown, /readiness always gives a fresh answer, so traffic stops at once.

Most names are a module's own checks. The process reports five of its own under _process, and healthCheckUnreadable is reported against the module it concerns:

healthCheckModuleWhat it meansFirst thing to check
connection_pool / connection_pool_writethe reporting moduleCannot get a (writable) database connectionIs the database up, reachable and writable?
http_responsesthe reporting moduleAt least 5 server errors (HTTP 5xx) in the last 5 minutes and no other responses. Clears on the next response that is not a server error, or at most about 6 minutes after the errors stopRecent errors, or a service this one depends on
dependent_fhir_storage_healththe reporting moduleThe FHIR storage this endpoint uses is unhealthyThe persistence module it points at
ft_indexing_statusthe reporting moduleFull-text indexing is behind or failingThe health of the search index (Elasticsearch or Lucene)
transaction_log_purge, stats_heartbeat, stats_cleanup_*the reporting moduleA background maintenance job is not runningThe scheduler and the database
memory_pressurethe Cluster Manager moduleThe process spends most of its time on garbage collection because it is short of memory. Does not clear until the process restarts, and also fails /livenessWhether the maximum memory (heap) suits the workload
hl7v2_mllp_listenerthe reporting moduleThe HL7v2 MLLP listener is downThe listener's port and configuration
sp_sync_analysis / sp_sync_syncthe reporting moduleSearch parameter synchronization is failingThe search parameter configuration
custom (or a module-specific name)the reporting moduleA check owned by a specific module or integrationThat module's own documentation
failedToStartthe module that failedThe module did not start, and will not recover by itselfThe module's configuration and the startup log
starting_processThe process is still startingNothing: expected during a start or restart
shuttingDown_processThe process is shutting downNothing: expected during a restart or planned stop
healthCheckTimeout_processThe health checks did not finish within node.status.health_check_timeout_secondsA check stuck on something slow, such as an unresponsive database. If the node simply has many checks, raise the timeout
healthCheckUnavailable_processThe checks could not run, because a check is stuck or too many requests arrived at onceEarlier healthCheckTimeout entries and the system log point to a stuck check; otherwise poll /readiness from fewer systems. Restart the process if a stuck check does not clear
healthCheckError_processThe checks could not run, usually because a module was restartingWhether a module is restarting; the system log has the error
healthCheckUnreadablethe module concernedThe module is running, but its health checks could not be read, usually because it is restartingWhether the module is restarting; this clears once it is back

One cause often produces two entries. In the example above, the unreachable database fails Persistence's connection_pool, and the FHIR endpoint that uses Persistence reports dependent_fhir_storage_health. Fix the database and both clear.

failedToStart

Unlike a failing health check, failedToStart does not clear on its own. Boot has finished, so /startup returns 200 while /readiness stays at 503 indefinitely. Every process running the same node configuration is affected, not just one.

Where you correct the configuration depends on node.propertysource, and the default is the less obvious of the two:

  • DATABASE (the default). Only the Cluster Manager is configured from the properties file. Every other module is configured from the Cluster Manager database after first startup, so editing the properties file and restarting will not change a failed module's configuration. Fix it in the Web Admin Console or through the Module Config endpoint, then restart the module or the process.
  • PROPERTIES. Module configuration is read from the properties file on every startup. Correct it there and restart the process.

Runtime Health Checks

 

As it is running, Smile CDR maintains a number of "health checks", which are simple status monitors that monitor whether a particular component of the system is operational.

For example, health checks exist to test database availability and health, HTTP server availability, scheduled job status, etc.

These health checks may be queried by a monitoring tool using the Runtime Status Health Checks operation on the JSON Admin API.