Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Node Health Monitoring

Fuzzball publishes facts about its compute nodes. It does not judge them.

Each node sends a periodic heartbeat, which is how the control plane tells a live node from a silent one, and exports hardware counters for a monitoring stack to scrape. Which counters matter, how long a fault must persist, and what severity it warrants are decided by Prometheus alerting rules – reviewable in git, tunable per site – not by thresholds inside the product.

When the monitoring stack concludes a node should stop taking work, it says so by placing a hold. That is the hook, and it is the only thing that takes a node out of service.

This split is new. Earlier releases scored each node 0-100, derived a health state from that score, and shipped an opt-in policy engine that could cordon, drain or replace a node on its own. The score, the health state, the condition vocabulary, the policy engine and the Fuzzball-native notification webhook are all retired – see What Was Retired.

Prerequisites

Reading orchestrate service logs requires access to the Fuzzball control plane namespace. Changing the heartbeat miss threshold requires edit access to the FuzzballOrchestrate resource in the fuzzball-system namespace.

What Fuzzball Observes

Three things, and only these:

SignalWhere it appears
Node status – unknown, not ready, ready, offlinefuzzball node list, fuzzball node show, the API
When the node last reportedfuzzball node show, fuzzball_scheduler_node_last_health_report_timestamp_seconds
Job failures attributed to the nodefuzzball_scheduler_node_attributed_job_failures

Every change to a node’s status is recorded as a STATUS_CHANGED event carrying both the previous and the new status, so a node’s history can be rebuilt without parsing the event’s message. Read them with fuzzball node events NODE --type status_changed.

Status reports what Fuzzball has observed about a node, and nothing else:

StatusMeans
unknownFuzzball cannot see the node: no report yet, or reports have gone stale
not readyThe node is reporting, but no provisioner definition has claimed it
readyReporting, claimed, and able to accept work
offlineThe node is gone. Work placed there is released and the record is reaped

A cordoned node is not a status. It reads ready and carries a cordon hold, because “nothing is wrong with this machine” and “nobody wants new work on it” are different statements. See Node Cordoning and Uncordoning.

The attributed-failure counter is the one node signal Fuzzball can see that a monitoring stack cannot: it knows which node a job was running on when it stopped responding. It is exported raw, for rules to alert on.

Substrate Version

The orchestrate extension needs a minimum version of Fuzzball Substrate on the node, currently v2.7.0, to pull and build container images. Each node’s extension reports both the version substrate states and the minimum that extension requires, so a control plane upgraded ahead of the fleet does not condemn nodes whose own extension is still serving image pulls. When a node runs an older substrate than its extension requires, the extension keeps running and reporting, and the control plane places a cordon hold on it:

Hold idSourceEffectPlaced when
substrate-versionprovisionercordonThe substrate on the node states a version older than the extension requires, or none at all

The hold’s reason names both versions. While it is held the node refuses image pulls, so workflow stages that need an image fail on it with the same message, and the scheduler places no new work there. Work already running is left alone – the hold cordons, it does not evict: draining would move work that is running fine, and replacing a provisioned node would bring back the same substrate.

This is the one hold Fuzzball places on its own judgement rather than an operator’s or a monitoring rule’s, and it is the exception for a specific reason: whether the extension can serve a node’s image pulls is Fuzzball’s own knowledge, not something a rule reading metrics could determine.

The fix is to upgrade fuzzball-substrate on the node. The hold is released on the node’s next health report – within 30 seconds by default – and it takes work again. Releasing it by hand with fuzzball node hold release succeeds but achieves nothing: the next report places it again, because the condition it describes is still true.

Heartbeat and Liveness

Each compute node reports every 30 seconds by default. When a node misses the configured number of consecutive reports, three by default, the control plane moves it to the unknown status, records a STATUS_CHANGED event, and logs a warning naming the node and the time of its last report.

Heartbeat tracking does not depend on detecting a network disconnect. A node whose hardware has hung, or whose substrate extension has stopped while the host is still reachable, stops sending heartbeats and goes unknown on the same timer as any other silent node.

A node that starts reporting again returns to ready when a provisioner definition has already claimed it, and to not ready when none has – the status it would have had if it had never gone silent. The health report itself does not know whether the node is claimed; that is read from the node’s last substrate resource report.

Only a health report brings a node out of unknown. A substrate resource report for a node whose health reports have stopped does not return it to service, because the two come from different places: substrate publishes the resource, and the orchestrate extension reports the health. A node whose extension has stopped while substrate keeps running would otherwise be handed work again on its next capacity change, and taken back out on the next liveness sweep.

A node that goes offline and re-registers is a different case: its last-report time is cleared with the status, so it comes back as a node that has never reported and is exempt from the sweep until it reports again. Without that, a rebooted node would be condemned on a timestamp from before it left.

A node whose health reporting is switched off after it has reported at least once stays unknown, and out of the placement set, until reporting resumes. Switch health reporting off before the node first reports, or leave it on: a node that has never reported is exempt from the liveness sweep, but one that stops is treated as unreachable.

On the API, that status is NODE_STATUS_UNKNOWN. It is a value of its own rather than NODE_STATUS_UNSPECIFIED, so a client can tell “Fuzzball cannot see this node” from “the server did not set this field”.

Heartbeat Miss Threshold

Set the number of consecutive reports a node can miss in the cluster’s FuzzballOrchestrate custom resource:

apiVersion: deployment.ciq.com/v1alpha1
kind: FuzzballOrchestrate
metadata:
  name: fuzzball-orchestrate
spec:
  nodeHealth:
    heartbeatMissThreshold: 3

See Node Health Configuration in the CRD reference for the full field description.

The threshold is a count, not a duration, and each node is measured against its own reporting interval. Nodes that report at different intervals therefore need no per-node tuning. With the defaults, a node goes unknown roughly 90 seconds after its last report.

A lower threshold detects a failed node sooner, at the cost of more false positives on a congested or high-latency network. A higher threshold reduces false positives but delays detection.

Node Reporting Interval

The reporting cadence is set per node, in the substrate orchestrate extension configuration rather than on the FuzzballOrchestrate resource. This applies to nodes whose extension configuration you manage directly; nodes provisioned by the Fuzzball operator use the defaults below and do not expose these settings.

binary: fuzzball-substrate-orchestrate
enabled: true
config:
  health:
    disabled: false
    report-interval-seconds: 30
ParameterTypeDefaultDescription
disabledbooleanfalseTurn off health reporting on this node entirely
report-interval-secondsinteger30Seconds between reports. Values below 5 fall back to the default
Turning health reporting off also removes the node from missed-heartbeat detection. A node that has never sent a health report never goes unknown for missing one, because it has missed no heartbeat it promised – the same applies to a node running an extension too old to report health. Nodes that reported and then had reporting disabled are still detected.

The control plane puts a floor under the window it derives from these values, and a ceiling on the interval it derives it from. It waits at least 15 seconds after a node’s last report before recording the node as silent, however short the interval. It also treats any reported interval longer than 24 hours as 24 hours, so a node configured to report very rarely is still recorded as silent after the threshold multiplied by 24 hours – three days at the default threshold.

Image Provider Settings

The same extension configuration carries the settings of the node’s image provider, which converts container images to EROFS and serves them to jobs:

binary: fuzzball-substrate-orchestrate
enabled: true
config:
  image-provider:
    mount-mode: auto
    prune-interval: 6h
ParameterTypeDefaultDescription
mount-modestringautoHow converted images are mounted for containers: fuse serves them through a FUSE daemon, loop through loop devices, and auto picks what the node’s kernel serves best. auto picks loop only where the kernel can also map file ownership on it (Linux 5.19 or newer), and serves any container whose mount the kernel still refuses to map through FUSE, so containers see the same file ownership on every node. With loop set explicitly, such a container fails to start instead
prune-intervalduration6hHow often orphaned layers are removed from the node’s image store, as a Go duration such as 6h or 90m. Nodes sharing a store take a cluster-wide lock, so one prune runs at a time; the store records when it last ran and a restarting node prunes at once if that is older than the interval

Viewing Node State

Status and holds appear in fuzzball node list:

fuzzball node list
ID        CLUSTER    HOSTNAME    STATUS     HOLDS    JOBS
node-01   default    compute-1   Ready      -        3
node-02   default    compute-2   Ready      cordon   1
node-03   default    compute-3   Unknown    -        0

node-03 reading Unknown means Fuzzball cannot see it – it has not reported yet, or its reports have stopped. That is not evidence its hardware is faulty.

node-02 reading Ready with a cordon hold is not a contradiction: the hardware is fine and somebody has asked that no new work be placed there.

fuzzball node show adds when the node last reported, the failures attributed to it, and each hold with its source and reason:

fuzzball node show node-02
ID:        10.0.0.42/9100
Hostname:  compute-2
CPU:       x86_64
Status:    Ready
Reported:  2026-09-17T14:05:00Z
Attributed job failures: 2 (last 2026-09-17T11:20:14Z)
Holds: 1
  7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42  cordon  monitoring  "4 correctable ECC errors in 24h"
    since 2026-09-17T11:22:03Z

Health Metrics

The control plane publishes two per-node series:

MetricTypeLabelsMeaning
fuzzball_scheduler_node_attributed_job_failuresgaugenode_idCumulative jobs that stopped responding on this node
fuzzball_scheduler_node_last_health_report_timestamp_secondsgaugenode_idUnix timestamp of the node’s last health report

The failure count is a gauge rather than a counter because it is restored from the node record when the control plane restarts, rather than counted up from zero. Alert on the level, not on increase() – a leader failover would read as a burst of new failures.

The last-report timestamp is what a silent-node alert reads. It is a timestamp rather than an age so the value stays true without being republished every scrape; compute the age in the rule:

time() - max by (node_id) (fuzzball_scheduler_node_last_health_report_timestamp_seconds) > 600

It is absent until a node reports once, so pair a staleness rule with an absent() check rather than treating a missing series as a node that never reported. Publishing zero instead would read as “last reported in 1970” and alert on every newly registered node.

Only the leader replica publishes node series, and it drops them when it stops being leader, so every expression should aggregate with max by (node_id): during a failover both the old and new leader can be scraped, and summing would double count.

Hardware counters – ECC errors, machine checks, SMART attributes, GPU telemetry – are exported by Node Exporter and DCGM Exporter running on the node itself, not by the control plane.

Bundled Monitoring Stack

Fuzzball can deploy Prometheus, Alertmanager and Grafana alongside a cluster, preconfigured against the cluster’s metrics. The shipped rule pack keeps the group name fuzzball-node-health, so a site that has already silenced or routed that group keeps working.

The pack currently carries three rules, all written against the series above:

AlertFires when
FuzzballNodeUnknownA node has not reported for 10 minutes
FuzzballNodeAttributedJobFailuresFive or more jobs have stopped responding on one node
FuzzballNodeHealthMetricsAbsentNo node series at all – no leader, or scraping is broken

FuzzballNodeUnknown ships with the product rather than being left to a site’s own rules, because a node going silent is only visible to Fuzzball. The bundled stack scrapes Fuzzball’s own pods, so there is no Node Exporter up series that would notice a host whose extension has died.

FuzzballNodeHealthMetricsAbsent exists because every other rule goes quiet when the control plane publishes nothing, and an absent series must never read as a healthy fleet.

The full rule pack and dashboards over the exporter series are a separate piece of work; the rules above are what can honestly be written against what Fuzzball publishes today.

Nodes That Go Offline

When a node disappears – power loss, a killed agent, an instance terminated underneath Fuzzball – the work placed on it is already gone. Fuzzball releases it immediately rather than waiting for each lease to time out one at a time:

  • Steps with policy.retry.on: node are requeued and run again elsewhere.
  • Everything else fails, with the node named.

There is no judgement to make here, so this is not configurable: the node is gone and the work died with it. The only question is whether you are told now or after a timeout. A multinode job in particular no longer holds its remaining nodes for the length of that wait.

The failure names no cause. A node disappearing says nothing about its hardware – a rack losing power and an agent being killed look the same from the control plane – so Fuzzball reports what it knows and no more:

node 10.0.0.42/9100 stopped hosting this work; job failed. Set policy.retry.on to node for work that is safe to run again from the beginning

A step that declared policy.retry.on: node and has already been moved maxEvictionRestarts times is failed with the count rather than that advice:

node 10.0.0.42/9100 stopped hosting this work; job failed after 3 node restart(s), the limit for this step

When a Job Fails on a Node

When a job ends because its node stopped responding, the error names the node:

job stopped responding on node 10.0.0.42/9100

If that node is held at the time, the error names the hold’s reason as well, so a user can tell immediately that the failure was not theirs:

node 10.0.0.42/9100 is held: GPU 2 uncorrectable ECC; job stopped responding on this node

The reason comes from the hold rather than from a diagnosis Fuzzball made: a person or an alerting rule wrote that sentence about this node deliberately. A node with nothing held against it is still named, because knowing which machine a job died on is useful on its own, and claims nothing.

The failure is also counted against the node’s attributed-failure total.

Node Report Storage

The control plane keeps the latest resource report for every attached node in a JetStream key-value bucket named substrate-nodes. Bucket usage grows with the number of nodes and with the size of each report, and a report grows with the node’s core count, its devices and its annotations.

Size the bucket from the node count multiplied by the size of one report. A 192-core node with 8 GPUs and 10 annotations produces a report of roughly 4 KiB, so 3,000 such nodes need about 12 MiB. The bucket defaults to 64 MiB, which holds roughly 16,000 reports of that size.

A full bucket fails quietly – no error surfaces in the CLI, the API or a node’s status. JetStream stops accepting reports from nodes the bucket has not seen before, and stops accepting any report larger than the one it replaces. The cluster stops registering new nodes, and the scheduler keeps planning against the last resources it recorded for the nodes it already has. The only signals are the substrate bridge log, which reports a rejection at error level at most once every 30 seconds, and the metrics in Bucket Metrics.

Set the bucket’s size cap with nodesBucketMaxBytes in the cluster’s FuzzballOrchestrate custom resource:

apiVersion: deployment.ciq.com/v1alpha1
kind: FuzzballOrchestrate
metadata:
  name: fuzzball-orchestrate
spec:
  fuzzball:
    substrate:
      nodesBucketMaxBytes: 1073741824

See Substrate Configuration in the CRD reference for the full field description.

When the substrate bridge restarts, it applies the new cap, resizing the existing bucket in place. Leaving the field unset changes nothing: the 64 MiB default is the size clusters ran with before the field existed.

Check fuzzball_substrate_nodes_bucket_bytes_used before lowering a cap. Any cap below what the bucket already stores produces the same failure as a full bucket. The field’s 64 MiB minimum only rejects caps below 64 MiB; it does not compare the new cap against what the bucket currently stores, so a bucket already holding more than 64 MiB can still be capped below its own contents.

The cap is a reservation, not a limit on actual use. JetStream subtracts it from the NATS server’s on-disk allowance of 10 GiB whether or not the bucket holds that much, and every other Fuzzball stream on that server draws on the same allowance.

That allowance is per server. A high-availability deployment replicates the bucket across three NATS servers, and each server reserves the full cap against its own 10 GiB. A 1 GiB cap therefore costs 1 GiB on every server, not 3 GiB on any one of them.

An oversized reservation stops the substrate bridge from starting rather than degrading quietly, so the cap is bounded at 2 GiB. The Kubernetes API server rejects a nodesBucketMaxBytes value below 64 MiB or above 2 GiB when the custom resource is applied, so a mistyped size fails at apply time. A value that reaches the orchestrate service configuration by another path is clamped into that range rather than preventing the bridge from starting. The bridge logs the cap in force at startup, and logs the configured value alongside it only when that value had to be clamped. Raise nodesBucketMaxBytes as the node count grows; do not set a large value up front.

Bucket Metrics

Three series published by the substrate bridge cover the bucket:

MetricTypeLabelsDescription
fuzzball_substrate_nodes_bucket_bytes_usedgauge-Bytes stored in the substrate-nodes bucket
fuzzball_substrate_nodes_bucket_bytes_maxgauge-Configured cap on the bucket
fuzzball_substrate_nodes_bucket_put_failures_totalcounter-Resource reports the bucket rejected

The two gauges are sampled once a minute, regardless of how many nodes are attached. Alert when fuzzball_substrate_nodes_bucket_bytes_used exceeds 80% of fuzzball_substrate_nodes_bucket_bytes_max. That leaves room to raise the cap and restart the substrate bridge before the bucket fills.

Both gauges are absent until the first successful sample, so absence means usage is unknown rather than zero. Pair a headroom alert with an absent() check rather than treating a missing series as an empty bucket.

Node Event History

Events are retained for 90 days and can be read back per node through the ListNodeEvents API, filtered by time and by event type and returned oldest first in pages.

Read a node’s history with fuzzball node events, oldest first:

fuzzball node events 10.0.0.42/9100 --since 24h
TIME                 | TYPE                    | DETAIL
2026-09-17T09:14:02Z | HOLD_PLACED             | cordon by monitoring: 4 correctable ECC errors in 24h
2026-09-17T11:37:41Z | JOB_FAILURE_ATTRIBUTED  | job alloc-93f2 failed on this node
2026-09-17T11:52:10Z | STATUS_CHANGED          | status ready -> unknown: no health report for 3 consecutive intervals

Filter by type with --type, repeatable, and use -o json to pull the series out for plotting.

STATUS_CHANGED is the event to watch for a node going silent. It replaces the two events that used to describe it – a raised HEARTBEAT_MISSED condition and a health-state transition into UNKNOWN – which said the same thing twice in a vocabulary that no longer exists.

What Was Retired

If you are upgrading, these are gone. Nothing needs to be done to remove them; they simply stop being produced.

RetiredReplaced by
The 0-100 reliability score, its penalties, thresholds and explanationPrometheus rules over the raw counters
DEGRADED as a health state, and the closed condition-type listWhatever your rules match on
The policy engine and its observe / cordon / drain / replace modesWhich rules a site enables, and what action each carries
evictOnFaultClassesA rule with an evict action
The score-based placement gateHolds
Automated node replacement and its churn ceilingA deprovision hold; the provisioner supplies capacity again when work needs it
The manual-uncordon suppression windowHold source – an administrator hold is one automation may never clear
The Fuzzball-native signed notification webhookAlertmanager receivers
Score, condition, health-state and policy node eventsHOLD_PLACED, HOLD_RELEASED, STATUS_CHANGED, JOB_FAILURE_ATTRIBUTED

The nodeHealth central-config block keeps one field, maxEvictionRestarts. The penalties, thresholds, failure-tally weights, policy mode and evictOnFaultClasses it used to carry are ignored – a stored config that still sets them keeps loading, so an upgrade does not need the file edited first.

On the API, the retired fields stay on the wire and are never populated: reliability_score, health_state, conditions and score_explanation on a node; the score, condition, threshold and policy fields on an event. A client reading them sees the zero value rather than a stale judgement.

fuzzball node drain is deprecated in favour of fuzzball node cordon, and fuzzball node evict in place of --active. It keeps working, but --active has changed: it now honours what each step declared rather than requeueing everything on the node. A step whose policy.retry.on includes node is requeued; anything else is failed. Earlier releases restarted steps whose authors had said they must never be restarted.