Troubleshooting Storage Container Manager safemode
Use this page when the Storage Container Manager (SCM) stays in safemode longer than expected after startup or a restart, and client writes fail.
This page covers SCM safemode only. Ozone Manager (OM) and Recon have separate HA and warmup behavior. If OM metadata works but data writes fail, SCM safemode is a likely cause.
Under normal conditions, SCM exits safemode automatically once its exit rules are satisfied. This guide helps you find which rule is blocking exit and what to do about it.
For background on what safemode is, how it works, and the full list of exit rules, see Core Concepts → Architecture → Storage Container Manager (Safe Mode section).
When to use this page
- After a cluster restart or rolling restart, writes fail but the cluster otherwise looks up.
- After Datanode or SCM node failures, SCM restarted and has not left safemode.
- During initial cluster startup,
ozone admin safemode waittimes out. - Client or SCM logs mention SafeModePrecheck failures.
Symptoms
-
Client writes fail. OM or S3 may return errors that mention SCM safemode.
-
SCM logs or client stack traces include either:
SafeModePrecheck failed for allocateBlockor:
SafeModePrecheck failed for allocateContainerSCM blocks block allocation and container allocation while in safemode, so create and append operations that need new blocks or containers cannot proceed.
-
Reads may still work for existing data, depending on pipeline and container state.
Check safemode status
Run against the SCM leader. In an SCM HA cluster, identify the leader with:
ozone admin scm roles --service-id=<scm_service_id>
On a single-SCM cluster, run the commands below against that SCM host.
ozone admin safemode status --verbose
When SCM is in safemode, --verbose prints each exit rule and whether it has passed, for example:
SCM is in safe mode.
validated:false, DataNodeSafeModeRule, ...
validated:false, ContainerSafeModeRule, ...
ContainerSafeModeRule status text reports Ratis and EC container progress on separate lines.
You can also open the leader SCM web UI (default port 9876, see ozone.scm.http-address) and review the safemode rule status shown on the page.
To wait until safemode clears (for example during startup):
ozone admin safemode wait -t 240
If rules are close to passing, wait before forcing exit. SCM refreshes rule state every hdds.scm.safemode.rule.refresh.interval (default 5s). While in safemode, SCM also logs status periodically at hdds.scm.safemode.log.interval (default 1m) — useful when CLI output is inconclusive.
How automatic exit works
SCM enters safemode on startup and stays there until all configured exit rules pass.
DataNodeSafeModeRule is both one of the five exit rules and the pre-check gate: SCM will not evaluate container or pipeline rules until enough Datanodes have registered.
SCM evaluates five main exit rules before leaving safemode (documented in Core Concepts → Architecture → Storage Container Manager):
| Rule | What it checks | Default threshold (ozone-site.xml) |
|---|---|---|
DataNodeSafeModeRule | Minimum Datanodes registered with SCM | hdds.scm.safemode.min.datanode = 3 |
RatisContainerSafeModeRule | Percentage of Ratis containers with at least one reported replica | hdds.scm.safemode.threshold.pct = 0.99 |
ECContainerSafeModeRule | Percentage of EC containers with enough reported replicas (at least the data-node count for the EC layout) | hdds.scm.safemode.threshold.pct = 0.99 |
HealthyPipelineSafeModeRule | Percentage of pipelines where all Datanodes are reported healthy | hdds.scm.safemode.healthy.pipeline.pct = 0.10 |
OneReplicaPipelineSafeModeRule | Percentage of pipelines with at least one Datanode reported | hdds.scm.safemode.atleast.one.node.reported.pipeline.pct = 0.90 |
Failing rule (in --verbose output) | Start here |
|---|---|
| No progress; cannot reach SCM or no leader elected | §1 SCM HA quorum |
DataNodeSafeModeRule | §2 Datanodes |
ContainerSafeModeRule (Ratis or EC lines in status text) | §3 Containers and pipelines |
HealthyPipelineSafeModeRule or AtleastOneDatanodeReportedRule | §3 Containers and pipelines |
ozone admin safemode status --verbose groups the two container checks under a single ContainerSafeModeRule entry. Its status text reports Ratis and EC progress separately. The pipeline rule appears as AtleastOneDatanodeReportedRule (the implementation class is OneReplicaPipelineSafeModeRule).
Pipeline rules apply only when hdds.scm.safemode.pipeline.availability.check is true (the default).
If any rule shows validated:false in the verbose output, SCM remains in safemode.
Setting hdds.scm.safemode.enabled to false skips SCM safemode entirely. Use only in development or test environments, not in production.
Common causes and fixes
1. SCM HA quorum is not healthy
In an SCM HA cluster, only the leader SCM processes Datanode reports and drives safemode exit. Followers are expected to be running but inactive as leader.
If too many SCM processes are down, Raft cannot elect or keep a leader. Datanode reports may not be processed and safemode will not clear.
What to check
- All SCM processes configured in
ozone.scm.nodesare running. - A majority of SCM nodes are available (for example, at least 2 of 3 in a three-node SCM HA setup).
ozone admin scm rolesshows aLEADERand reachable followers.- The current leader is reachable from clients and Datanodes.
What to do
- Start any stopped SCM processes.
- Resolve leader election or network issues before investigating Datanode or container rules.
2. Not enough Datanodes are registered
DataNodeSafeModeRule must pass before SCM evaluates container or pipeline rules.
What to check
ozone admin datanode list
Look for Datanodes that are DEAD, STALE, or missing from the list.
Common reasons
- Datanode processes are not running.
- Datanodes are still starting up — for example, verifying containers and volumes on a large deployment can take time after restart.
- Network, certificate, or configuration problems prevent registration with SCM.
What to do
- Start missing Datanodes and wait for registration.
- Check Datanode logs for registration or TLS errors.
- Follow the startup order in Starting and stopping the cluster (SCM, then OM, then Datanodes).
- If you intentionally run a small cluster (for example, a single-node test), lower
hdds.scm.safemode.min.datanode— see the Flink integration guide for an example.
3. Container or pipeline thresholds are not met
Even with enough Datanodes online, SCM stays in safemode until enough Ratis containers, EC containers, and pipelines satisfy their thresholds.
- Containers: Ratis containers need at least one reported replica; EC containers need enough replicas for the EC data-node layout.
- Pipelines:
HealthyPipelineSafeModeRulerequires all Datanodes in a pipeline to be reported healthy.OneReplicaPipelineSafeModeRulerequires at least one Datanode reported per pipeline.
What to check
ozone admin container report
ozone admin pipeline list
The container report shows replication health (UNDER_REPLICATED, MISSING, UNHEALTHY, and related states). See Container Replication Report for how to interpret it.
Common reasons
- Many Datanodes were down and container reports have not caught up yet. Datanodes send container reports on
hdds.container.report.interval(default60m), so thresholds can take time to clear after a large restart. - Containers are under-replicated, missing, or corrupt after disk or node failures.
- Pipelines have not been created or reported yet (
hdds.scm.safemode.pipeline.creationcontrols background pipeline creation during safemode, defaulttrue). - On clusters whose default replication is EC, SCM may also require
RATIS/THREEpipelines during safemode whenozone.scm.pipeline.creation.ratis.threeistrue(the default).
What to do
- Restore failed Datanodes and disks, then wait for container reports and replication to progress.
- Investigate
MISSINGor persistentUNDER_REPLICATEDcontainers before forcing safemode exit. - For pipeline-related rules, confirm Datanodes are healthy and pipeline creation is enabled.
- On EC-default clusters, confirm
RATIS/THREEpipelines exist ifozone.scm.pipeline.creation.ratis.threeis enabled.
Force exit (last resort)
If you are confident the cluster is healthy and safemode is stuck due to a transient reporting delay, you can force SCM out of safemode:
ozone admin safemode exit
Forcing exit skips the automatic health checks. Do this only when you understand which rule is failing and have verified that the underlying problem is resolved or acceptable. Exiting safemode with missing replicas or too few Datanodes can cause write failures or data availability issues.
After a forced exit, SCM waits hdds.scm.wait.time.after.safemode.exit (default 5m) before starting some background recovery work.
Verify
After safemode clears (automatically or manually):
ozone admin safemode status
Expected output:
SCM is out of safe mode.
Then retry a write from your client application or:
ozone sh key put <volume>/<bucket>/safemode-test -f <local-file>
Still stuck?
- Confirm you checked safemode on the SCM leader, not a follower (see
ozone admin scm roles). - Review SCM logs around safemode status lines (emitted every
hdds.scm.safemode.log.interval). - If safemode is clear but writes still fail, the problem is likely elsewhere — for example write pipelines or client connectivity. See Troubleshooting client connectivity.
See also
- Core Concepts → Architecture → Storage Container Manager — safemode overview and exit rules
- Starting and stopping the cluster — includes
ozone admin safemode wait - Ozone Admin tools —
ozone admin safemodesubcommands - Configuration appendix — safemode-related settings