This page outlines methods for monitoring Smile CDR.
Smile CDR provides several hooks that are suitable for integration with an external monitoring system.
All HTTP Servers supply an HTTP listener at the path /endpoint-health that can be used to query for the current availability of the endpoint. This endpoint will return an HTTP 200 with a simple JSON response payload as long at the endpoint is enabled and operational.
This is useful for systems requiring a simple way to test whether an endpoint is available, such as load balancers that perform automated failover based on a status test. The Endpoint Health Check is very lightweight and may be called frequently without causing extra load on the server.
The Endpoint Healh Check is available on any endpoint modules that supply an HTTP server (e.g. FHIR REST Endpoint or Web Admin Console. It will typically be served at /endpoint-health, but note that it will respect the context path if one is set. So for example if the context path is set to /admin, the Endpoint Health Check will be served at /admin/endpoint-health.
The following example shows a simple HTTP query against the Endpoint Health Check.
GET /endpoint-health
The server will respond with the following response:
200 OK
Content-type: application/json
{"status":"OPERATIONAL"}
If any health checks are failing on the given module, the server will return a response which includes the string NOT_OPERATIONAL as well as the IDs of any failing health checks. For example:
200 OK
Content-type: application/json
{
"status" : "NOT_OPERATIONAL",
"failingHealthChecks" : [ "connection_pool" ]
}
By default, the endpoint will return an HTTP 200 OK regardless of whether any health checks are failing. A different status code can be returned by adjusting the Unhealthy Status Code setting.
Every Smile CDR process can serve three endpoints on a dedicated status port, designed for Kubernetes startup, readiness and liveness probes. They are plain HTTP, so any monitoring tool can also poll them. They are off by default.
| Endpoint | Answers | 200 | 503 |
|---|---|---|---|
/startup | Has the process finished starting? | Started | Still starting |
/readiness | Can the process serve traffic? | Every health check is passing | Starting, shutting down, a module failed to start, or a health check is failing |
/liveness | Does the process need a restart? | In every other case | A liveness check is failing |
Unlike /endpoint-health, these endpoints:
Because they need no login, restrict the status port with a firewall rule, security group or network policy so that only the systems that probe it can reach it. This is also why the endpoints stay off until you configure them.
Configure them in the node configuration properties file,
alongside node.id:
node.status.port turns the endpoints on and sets the port. As with every server port on the node,
node.server_port_offset is added to it, and the result must be between 1 and 65535.node.status.bind_address limits which network interface the endpoints listen on. The default,
every interface, is what Kubernetes needs; use 127.0.0.1 if you only poll from the same host.node.status.health_check_timeout_seconds is how long /readiness and /liveness wait for the
health checks to finish: 1 to 600, default 5. It covers all the checks in the process
together, because they run one after another. If a node reports healthCheckTimeout under normal
load, raise it, and keep your probe's own timeout above it (Kubernetes defaults timeoutSeconds
to 1 second); see
Kubernetes Health Probes.node.status.port=9300
Responses are JSON. Fields with no value are left out.
/startup returns STARTING (HTTP 503) while the process starts, then HTTP 200 and:
{
"status" : "STARTED"
}
/readiness returns READY (HTTP 200) or NOT_READY (HTTP 503), with the node id (node.id
from your configuration) and the process id (the name the process took when it joined the cluster,
absent until then):
{
"status" : "READY",
"nodeId" : "Master",
"processId" : "Algonquin"
}
When the process is not ready, failingHealthChecks lists each failing check by module (the module id from your configuration) and healthCheck name:
{
"status" : "NOT_READY",
"nodeId" : "Master",
"processId" : "Algonquin",
"failingHealthChecks" : [ {
"module" : "persistence",
"healthCheck" : "connection_pool"
}, {
"module" : "fhir_endpoint",
"healthCheck" : "dependent_fhir_storage_health"
} ]
}
A module that failed to start is listed under its own id as failedToStart. Conditions of the process itself use the module id _process: starting while it starts (no module checks run until then) and shuttingDown as soon as shutdown begins.
/liveness returns ALIVE (HTTP 200) or NOT_ALIVE (HTTP 503), with the same fields:
{
"status" : "NOT_ALIVE",
"nodeId" : "Master",
"processId" : "Algonquin",
"failingHealthChecks" : [ {
"module" : "clustermgr",
"healthCheck" : "memory_pressure"
} ]
}
/liveness reports/liveness answers one question: does this process need to be restarted? It counts only liveness checks, the health checks that only a restart can clear, and ignores all others. An unreachable database therefore never restarts the process; /readiness takes it out of service instead.
A standard installation has one liveness check, memory_pressure, reported by the Cluster Manager.
It fails when the process is so short of memory that it spends most of its time on garbage collection (reclaiming memory) instead of doing useful work. It measures once a minute, and fails after five measurements in a row where:
Memory is measured after garbage collection because it is always high just before one, even on a healthy process. Each measurement that meets both conditions is logged as a WARN in the system log, and the failure as an ERROR.
Once memory_pressure fails, it does not clear until the process restarts. It also makes
/readiness report NOT_READY, so the process stops receiving traffic until then. For a process that runs out of memory completely, see JVM Memory Settings.
In every other case /liveness reports ALIVE, including while the process is starting or shutting down, and when the liveness checks cannot finish within node.status.health_check_timeout_seconds or a module is restarting (the reason is logged). /readiness reports NOT_READY in those cases.
/readiness checkWhen a pod is not ready, or a monitor reports 503:
/readiness and read failingHealthChecks.healthCheck in the table below.Results are reused for one second, so a brief failure such as a short database outage clears as soon as its cause does. During shutdown, /readiness always gives a fresh answer, so traffic stops at once.
Most names are a module's own checks. The process reports five of its own under _process, and
healthCheckUnreadable is reported against the module it concerns:
healthCheck | Module | What it means | First thing to check |
|---|---|---|---|
connection_pool / connection_pool_write | the reporting module | Cannot get a (writable) database connection | Is the database up, reachable and writable? |
http_responses | the reporting module | At least 5 server errors (HTTP 5xx) in the last 5 minutes and no other responses. Clears on the next response that is not a server error, or at most about 6 minutes after the errors stop | Recent errors, or a service this one depends on |
dependent_fhir_storage_health | the reporting module | The FHIR storage this endpoint uses is unhealthy | The persistence module it points at |
ft_indexing_status | the reporting module | Full-text indexing is behind or failing | The health of the search index (Elasticsearch or Lucene) |
transaction_log_purge, stats_heartbeat, stats_cleanup_* | the reporting module | A background maintenance job is not running | The scheduler and the database |
memory_pressure | the Cluster Manager module | The process spends most of its time on garbage collection because it is short of memory. Does not clear until the process restarts, and also fails /liveness | Whether the maximum memory (heap) suits the workload |
hl7v2_mllp_listener | the reporting module | The HL7v2 MLLP listener is down | The listener's port and configuration |
sp_sync_analysis / sp_sync_sync | the reporting module | Search parameter synchronization is failing | The search parameter configuration |
custom (or a module-specific name) | the reporting module | A check owned by a specific module or integration | That module's own documentation |
failedToStart | the module that failed | The module did not start, and will not recover by itself | The module's configuration and the startup log |
starting | _process | The process is still starting | Nothing: expected during a start or restart |
shuttingDown | _process | The process is shutting down | Nothing: expected during a restart or planned stop |
healthCheckTimeout | _process | The health checks did not finish within node.status.health_check_timeout_seconds | A check stuck on something slow, such as an unresponsive database. If the node simply has many checks, raise the timeout |
healthCheckUnavailable | _process | The checks could not run, because a check is stuck or too many requests arrived at once | Earlier healthCheckTimeout entries and the system log point to a stuck check; otherwise poll /readiness from fewer systems. Restart the process if a stuck check does not clear |
healthCheckError | _process | The checks could not run, usually because a module was restarting | Whether a module is restarting; the system log has the error |
healthCheckUnreadable | the module concerned | The module is running, but its health checks could not be read, usually because it is restarting | Whether the module is restarting; this clears once it is back |
One cause often produces two entries. In the example above, the unreachable database fails Persistence's connection_pool, and the FHIR endpoint that uses Persistence reports
dependent_fhir_storage_health. Fix the database and both clear.
failedToStartUnlike a failing health check, failedToStart does not clear on its own. Boot has finished, so /startup
returns 200 while /readiness stays at 503 indefinitely. Every process running the same node configuration is affected, not just one.
Where you correct the configuration depends on node.propertysource,
and the default is the less obvious of the two:
DATABASE (the default). Only the Cluster Manager is configured from the properties
file. Every other module is configured from the Cluster Manager database after first startup, so
editing the properties file and restarting will not change a failed module's configuration.
Fix it in the Web Admin Console or through the
Module Config endpoint, then restart the
module or the process.PROPERTIES. Module configuration is read from the properties file on every startup.
Correct it there and restart the process.As it is running, Smile CDR maintains a number of "health checks", which are simple status monitors that monitor whether a particular component of the system is operational.
For example, health checks exist to test database availability and health, HTTP server availability, scheduled job status, etc.
These health checks may be queried by a monitoring tool using the Runtime Status Health Checks operation on the JSON Admin API.