Node Health Monitoring
Fuzzball publishes facts about its compute nodes. It does not judge them.
Each node sends a periodic heartbeat, which is how the control plane tells a live node from a silent one, and exports hardware counters for a monitoring stack to scrape. Which counters matter, how long a fault must persist, and what severity it warrants are decided by Prometheus alerting rules – reviewable in git, tunable per site – not by thresholds inside the product.
When the monitoring stack concludes a node should stop taking work, it says so by placing a hold. That is the hook, and it is the only thing that takes a node out of service.
This split is new. Earlier releases scored each node 0-100, derived a health state from that score, and shipped an opt-in policy engine that could cordon, drain or replace a node on its own. The score, the health state, the condition vocabulary, the policy engine and the Fuzzball-native notification webhook are all retired – see What Was Retired.
Reading orchestrate service logs requires access to the Fuzzball control plane
namespace. Changing the heartbeat miss threshold requires edit access to the
FuzzballOrchestrate resource in the fuzzball-system namespace.
Three things, and only these:
| Signal | Where it appears |
|---|---|
Node status – unknown, not ready, ready, offline | fuzzball node list, fuzzball node show, the API |
| When the node last reported | fuzzball node show, fuzzball_scheduler_node_last_health_report_timestamp_seconds |
| Job failures attributed to the node | fuzzball_scheduler_node_attributed_job_failures |
Every change to a node’s status is recorded as a STATUS_CHANGED event
carrying both the previous and the new status, so a node’s history can be
rebuilt without parsing the event’s message. Read them with
fuzzball node events NODE --type status_changed.
Status reports what Fuzzball has observed about a node, and nothing else:
| Status | Means |
|---|---|
unknown | Fuzzball cannot see the node: no report yet, or reports have gone stale |
not ready | The node is reporting, but no provisioner definition has claimed it |
ready | Reporting, claimed, and able to accept work |
offline | The node is gone. Work placed there is released and the record is reaped |
A cordoned node is not a status. It reads ready and carries a cordon
hold, because “nothing is wrong with this machine” and “nobody wants new work
on it” are different statements. See
Node Cordoning and Uncordoning.
The attributed-failure counter is the one node signal Fuzzball can see that a monitoring stack cannot: it knows which node a job was running on when it stopped responding. It is exported raw, for rules to alert on.
The orchestrate extension needs a minimum version of Fuzzball Substrate on the node, currently v2.7.0, to pull and build container images. Each node’s extension reports both the version substrate states and the minimum that extension requires, so a control plane upgraded ahead of the fleet does not condemn nodes whose own extension is still serving image pulls. When a node runs an older substrate than its extension requires, the extension keeps running and reporting, and the control plane places a cordon hold on it:
| Hold id | Source | Effect | Placed when |
|---|---|---|---|
substrate-version | provisioner | cordon | The substrate on the node states a version older than the extension requires, or none at all |
The hold’s reason names both versions. While it is held the node refuses image pulls, so workflow stages that need an image fail on it with the same message, and the scheduler places no new work there. Work already running is left alone – the hold cordons, it does not evict: draining would move work that is running fine, and replacing a provisioned node would bring back the same substrate.
This is the one hold Fuzzball places on its own judgement rather than an operator’s or a monitoring rule’s, and it is the exception for a specific reason: whether the extension can serve a node’s image pulls is Fuzzball’s own knowledge, not something a rule reading metrics could determine.
The fix is to upgrade fuzzball-substrate on the node. The hold is released on the node’s next
health report – within 30 seconds by default – and it takes work again. Releasing it by hand with
fuzzball node hold release succeeds but achieves nothing: the next report places it again, because
the condition it describes is still true.
Each compute node reports every 30 seconds by default. When a node misses the
configured number of consecutive reports, three by default, the control plane
moves it to the unknown status, records a STATUS_CHANGED event, and logs a
warning naming the node and the time of its last report.
Heartbeat tracking does not depend on detecting a network disconnect. A node
whose hardware has hung, or whose substrate extension has stopped while the
host is still reachable, stops sending heartbeats and goes unknown on the
same timer as any other silent node.
A node that starts reporting again returns to ready when a provisioner
definition has already claimed it, and to not ready when none has – the
status it would have had if it had never gone silent. The health report itself
does not know whether the node is claimed; that is read from the node’s last
substrate resource report.
Only a health report brings a node out of unknown. A substrate resource
report for a node whose health reports have stopped does not return it to
service, because the two come from different places: substrate publishes the
resource, and the orchestrate extension reports the health. A node whose
extension has stopped while substrate keeps running would otherwise be handed
work again on its next capacity change, and taken back out on the next
liveness sweep.
A node that goes offline and re-registers is a different case: its last-report time is cleared with the status, so it comes back as a node that has never reported and is exempt from the sweep until it reports again. Without that, a rebooted node would be condemned on a timestamp from before it left.
A node whose health reporting is switched off after it has reported at least once staysunknown, and out of the placement set, until reporting resumes. Switch health reporting off before the node first reports, or leave it on: a node that has never reported is exempt from the liveness sweep, but one that stops is treated as unreachable.
On the API, that status is NODE_STATUS_UNKNOWN. It is a value of its own
rather than NODE_STATUS_UNSPECIFIED, so a client can tell “Fuzzball cannot
see this node” from “the server did not set this field”.
Set the number of consecutive reports a node can miss in the cluster’s
FuzzballOrchestrate custom resource:
apiVersion: deployment.ciq.com/v1alpha1
kind: FuzzballOrchestrate
metadata:
name: fuzzball-orchestrate
spec:
nodeHealth:
heartbeatMissThreshold: 3
See Node Health Configuration in the CRD reference for the full field description.
The threshold is a count, not a duration, and each node is measured against
its own reporting interval. Nodes that report at different intervals therefore
need no per-node tuning. With the defaults, a node goes unknown roughly 90
seconds after its last report.
A lower threshold detects a failed node sooner, at the cost of more false positives on a congested or high-latency network. A higher threshold reduces false positives but delays detection.
The reporting cadence is set per node, in the substrate orchestrate extension
configuration rather than on the FuzzballOrchestrate resource. This applies
to nodes whose extension configuration you manage directly; nodes provisioned
by the Fuzzball operator use the defaults below and do not expose these
settings.
binary: fuzzball-substrate-orchestrate
enabled: true
config:
health:
disabled: false
report-interval-seconds: 30
| Parameter | Type | Default | Description |
|---|---|---|---|
disabled | boolean | false | Turn off health reporting on this node entirely |
report-interval-seconds | integer | 30 | Seconds between reports. Values below 5 fall back to the default |
Turning health reporting off also removes the node from missed-heartbeat detection. A node that has never sent a health report never goesunknownfor missing one, because it has missed no heartbeat it promised – the same applies to a node running an extension too old to report health. Nodes that reported and then had reporting disabled are still detected.
The control plane puts a floor under the window it derives from these values, and a ceiling on the interval it derives it from. It waits at least 15 seconds after a node’s last report before recording the node as silent, however short the interval. It also treats any reported interval longer than 24 hours as 24 hours, so a node configured to report very rarely is still recorded as silent after the threshold multiplied by 24 hours – three days at the default threshold.
The same extension configuration carries the settings of the node’s image provider, which converts container images to EROFS and serves them to jobs:
binary: fuzzball-substrate-orchestrate
enabled: true
config:
image-provider:
mount-mode: auto
prune-interval: 6h
| Parameter | Type | Default | Description |
|---|---|---|---|
mount-mode | string | auto | How converted images are mounted for containers: fuse serves them through a FUSE daemon, loop through loop devices, and auto picks what the node’s kernel serves best. auto picks loop only where the kernel can also map file ownership on it (Linux 5.19 or newer), and serves any container whose mount the kernel still refuses to map through FUSE, so containers see the same file ownership on every node. With loop set explicitly, such a container fails to start instead |
prune-interval | duration | 6h | How often orphaned layers are removed from the node’s image store, as a Go duration such as 6h or 90m. Nodes sharing a store take a cluster-wide lock, so one prune runs at a time; the store records when it last ran and a restarting node prunes at once if that is older than the interval |
Status and holds appear in fuzzball node list:
fuzzball node listID CLUSTER HOSTNAME STATUS HOLDS JOBS
node-01 default compute-1 Ready - 3
node-02 default compute-2 Ready cordon 1
node-03 default compute-3 Unknown - 0
node-03 reading Unknown means Fuzzball cannot see it – it has not reported
yet, or its reports have stopped. That is not evidence its hardware is faulty.
node-02 reading Ready with a cordon hold is not a contradiction: the
hardware is fine and somebody has asked that no new work be placed there.
fuzzball node show adds when the node last reported, the failures attributed
to it, and each hold with its source and reason:
fuzzball node show node-02ID: 10.0.0.42/9100
Hostname: compute-2
CPU: x86_64
Status: Ready
Reported: 2026-09-17T14:05:00Z
Attributed job failures: 2 (last 2026-09-17T11:20:14Z)
Holds: 1
7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42 cordon monitoring "4 correctable ECC errors in 24h"
since 2026-09-17T11:22:03Z
The control plane publishes two per-node series:
| Metric | Type | Labels | Meaning |
|---|---|---|---|
fuzzball_scheduler_node_attributed_job_failures | gauge | node_id | Cumulative jobs that stopped responding on this node |
fuzzball_scheduler_node_last_health_report_timestamp_seconds | gauge | node_id | Unix timestamp of the node’s last health report |
The failure count is a gauge rather than a counter because it is restored from
the node record when the control plane restarts, rather than counted up from
zero. Alert on the level, not on increase() – a leader failover would read as
a burst of new failures.
The last-report timestamp is what a silent-node alert reads. It is a timestamp rather than an age so the value stays true without being republished every scrape; compute the age in the rule:
time() - max by (node_id) (fuzzball_scheduler_node_last_health_report_timestamp_seconds) > 600
It is absent until a node reports once, so pair a staleness rule with an
absent() check rather than treating a missing series as a node that never
reported. Publishing zero instead would read as “last reported in 1970” and
alert on every newly registered node.
Only the leader replica publishes node series, and it drops them when it stops
being leader, so every expression should aggregate with max by (node_id):
during a failover both the old and new leader can be scraped, and summing would
double count.
Hardware counters – ECC errors, machine checks, SMART attributes, GPU telemetry – are exported by Node Exporter and DCGM Exporter running on the node itself, not by the control plane.
Fuzzball can deploy Prometheus, Alertmanager and Grafana alongside a cluster,
preconfigured against the cluster’s metrics. The shipped rule pack keeps the
group name fuzzball-node-health, so a site that has already silenced or routed
that group keeps working.
The pack currently carries three rules, all written against the series above:
| Alert | Fires when |
|---|---|
FuzzballNodeUnknown | A node has not reported for 10 minutes |
FuzzballNodeAttributedJobFailures | Five or more jobs have stopped responding on one node |
FuzzballNodeHealthMetricsAbsent | No node series at all – no leader, or scraping is broken |
FuzzballNodeUnknown ships with the product rather than being left to a site’s
own rules, because a node going silent is only visible to Fuzzball. The bundled
stack scrapes Fuzzball’s own pods, so there is no Node Exporter up series that
would notice a host whose extension has died.
FuzzballNodeHealthMetricsAbsent exists because every other rule goes quiet
when the control plane publishes nothing, and an absent series must never read
as a healthy fleet.
The full rule pack and dashboards over the exporter series are a separate piece of work; the rules above are what can honestly be written against what Fuzzball publishes today.
When a node disappears – power loss, a killed agent, an instance terminated underneath Fuzzball – the work placed on it is already gone. Fuzzball releases it immediately rather than waiting for each lease to time out one at a time:
- Steps with
policy.retry.on: nodeare requeued and run again elsewhere. - Everything else fails, with the node named.
There is no judgement to make here, so this is not configurable: the node is gone and the work died with it. The only question is whether you are told now or after a timeout. A multinode job in particular no longer holds its remaining nodes for the length of that wait.
The failure names no cause. A node disappearing says nothing about its hardware – a rack losing power and an agent being killed look the same from the control plane – so Fuzzball reports what it knows and no more:
node 10.0.0.42/9100 stopped hosting this work; job failed. Set policy.retry.on to node for work that is safe to run again from the beginning
A step that declared policy.retry.on: node and has already been moved
maxEvictionRestarts times is failed with the count rather than that advice:
node 10.0.0.42/9100 stopped hosting this work; job failed after 3 node restart(s), the limit for this step
When a job ends because its node stopped responding, the error names the node:
job stopped responding on node 10.0.0.42/9100
If that node is held at the time, the error names the hold’s reason as well, so a user can tell immediately that the failure was not theirs:
node 10.0.0.42/9100 is held: GPU 2 uncorrectable ECC; job stopped responding on this node
The reason comes from the hold rather than from a diagnosis Fuzzball made: a person or an alerting rule wrote that sentence about this node deliberately. A node with nothing held against it is still named, because knowing which machine a job died on is useful on its own, and claims nothing.
The failure is also counted against the node’s attributed-failure total.
The control plane keeps the latest resource report for every attached node in a
JetStream key-value bucket named substrate-nodes. Bucket usage grows with the
number of nodes and with the size of each report, and a report grows with the
node’s core count, its devices and its annotations.
Size the bucket from the node count multiplied by the size of one report. A 192-core node with 8 GPUs and 10 annotations produces a report of roughly 4 KiB, so 3,000 such nodes need about 12 MiB. The bucket defaults to 64 MiB, which holds roughly 16,000 reports of that size.
A full bucket fails quietly – no error surfaces in the CLI, the API or a node’s status. JetStream stops accepting reports from nodes the bucket has not seen before, and stops accepting any report larger than the one it replaces. The cluster stops registering new nodes, and the scheduler keeps planning against the last resources it recorded for the nodes it already has. The only signals are the substrate bridge log, which reports a rejection at error level at most once every 30 seconds, and the metrics in Bucket Metrics.
Set the bucket’s size cap with nodesBucketMaxBytes in the cluster’s
FuzzballOrchestrate custom resource:
apiVersion: deployment.ciq.com/v1alpha1
kind: FuzzballOrchestrate
metadata:
name: fuzzball-orchestrate
spec:
fuzzball:
substrate:
nodesBucketMaxBytes: 1073741824
See Substrate Configuration in the CRD reference for the full field description.
When the substrate bridge restarts, it applies the new cap, resizing the existing bucket in place. Leaving the field unset changes nothing: the 64 MiB default is the size clusters ran with before the field existed.
Checkfuzzball_substrate_nodes_bucket_bytes_usedbefore lowering a cap. Any cap below what the bucket already stores produces the same failure as a full bucket. The field’s 64 MiB minimum only rejects caps below 64 MiB; it does not compare the new cap against what the bucket currently stores, so a bucket already holding more than 64 MiB can still be capped below its own contents.
The cap is a reservation, not a limit on actual use. JetStream subtracts it from the NATS server’s on-disk allowance of 10 GiB whether or not the bucket holds that much, and every other Fuzzball stream on that server draws on the same allowance.
That allowance is per server. A high-availability deployment replicates the bucket across three NATS servers, and each server reserves the full cap against its own 10 GiB. A 1 GiB cap therefore costs 1 GiB on every server, not 3 GiB on any one of them.
An oversized reservation stops the substrate bridge from starting rather than
degrading quietly, so the cap is bounded at 2 GiB. The Kubernetes API server
rejects a nodesBucketMaxBytes value below 64 MiB or above 2 GiB when the
custom resource is applied, so a mistyped size fails at apply time. A value that
reaches the orchestrate service configuration by another path is clamped into
that range rather than preventing the bridge from starting. The bridge logs the cap
in force at startup, and logs the configured value alongside it only when that
value had to be clamped. Raise nodesBucketMaxBytes as the node count grows; do
not set a large value up front.
Three series published by the substrate bridge cover the bucket:
| Metric | Type | Labels | Description |
|---|---|---|---|
fuzzball_substrate_nodes_bucket_bytes_used | gauge | - | Bytes stored in the substrate-nodes bucket |
fuzzball_substrate_nodes_bucket_bytes_max | gauge | - | Configured cap on the bucket |
fuzzball_substrate_nodes_bucket_put_failures_total | counter | - | Resource reports the bucket rejected |
The two gauges are sampled once a minute, regardless of how many nodes are
attached. Alert when fuzzball_substrate_nodes_bucket_bytes_used exceeds 80% of
fuzzball_substrate_nodes_bucket_bytes_max. That leaves room to raise the cap
and restart the substrate bridge before the bucket fills.
Both gauges are absent until the first successful sample, so absence means usage is unknown rather than zero. Pair a headroom alert with anabsent()check rather than treating a missing series as an empty bucket.
Events are retained for 90 days and can be read back per node through the
ListNodeEvents API, filtered by time and by event type and returned oldest
first in pages.
Read a node’s history with fuzzball node events, oldest first:
fuzzball node events 10.0.0.42/9100 --since 24hTIME | TYPE | DETAIL
2026-09-17T09:14:02Z | HOLD_PLACED | cordon by monitoring: 4 correctable ECC errors in 24h
2026-09-17T11:37:41Z | JOB_FAILURE_ATTRIBUTED | job alloc-93f2 failed on this node
2026-09-17T11:52:10Z | STATUS_CHANGED | status ready -> unknown: no health report for 3 consecutive intervals
Filter by type with --type, repeatable, and use -o json to pull the series
out for plotting.
STATUS_CHANGED is the event to watch for a node going silent. It replaces the
two events that used to describe it – a raised HEARTBEAT_MISSED condition and
a health-state transition into UNKNOWN – which said the same thing twice in a
vocabulary that no longer exists.
If you are upgrading, these are gone. Nothing needs to be done to remove them; they simply stop being produced.
| Retired | Replaced by |
|---|---|
| The 0-100 reliability score, its penalties, thresholds and explanation | Prometheus rules over the raw counters |
DEGRADED as a health state, and the closed condition-type list | Whatever your rules match on |
The policy engine and its observe / cordon / drain / replace modes | Which rules a site enables, and what action each carries |
evictOnFaultClasses | A rule with an evict action |
| The score-based placement gate | Holds |
| Automated node replacement and its churn ceiling | A deprovision hold; the provisioner supplies capacity again when work needs it |
| The manual-uncordon suppression window | Hold source – an administrator hold is one automation may never clear |
| The Fuzzball-native signed notification webhook | Alertmanager receivers |
| Score, condition, health-state and policy node events | HOLD_PLACED, HOLD_RELEASED, STATUS_CHANGED, JOB_FAILURE_ATTRIBUTED |
The nodeHealth central-config block keeps one field, maxEvictionRestarts.
The penalties, thresholds, failure-tally weights, policy mode and
evictOnFaultClasses it used to carry are ignored – a stored config that still
sets them keeps loading, so an upgrade does not need the file edited first.
On the API, the retired fields stay on the wire and are never populated:
reliability_score, health_state, conditions and score_explanation on a
node; the score, condition, threshold and policy fields on an event. A client
reading them sees the zero value rather than a stale judgement.
fuzzball node drainis deprecated in favour offuzzball node cordon, andfuzzball node evictin place of--active. It keeps working, but--activehas changed: it now honours what each step declared rather than requeueing everything on the node. A step whosepolicy.retry.onincludesnodeis requeued; anything else is failed. Earlier releases restarted steps whose authors had said they must never be restarted.