Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Node Cordoning and Uncordoning

Node cordoning marks compute nodes as unschedulable, preventing new workflows from being placed on them. This is commonly used during maintenance, node upgrades, or when preparing to remove nodes from the cluster.

A cordon is recorded as a hold on the node rather than as the node’s status. A node’s status reports what Fuzzball has observed about it – whether it is reporting, whether it is claimed, whether it has gone away – while a hold records what somebody has asked be done about it. The two answer different questions, so a cordoned node reports Ready and carries a cordon hold.

fuzzball node list shows holds in their own column, and fuzzball node show lists each hold with its source and reason. A node reading Ready with a cordon hold is not a contradiction: nothing is wrong with the hardware, and nobody wants new work placed on it.

Cordoning leaves running work alone. To move that work off a node, or to tear the instance down, see Evicting and Deprovisioning Nodes.

Prerequisites

Changes to compute node scheduling require that a fuzzball admin context is set up and authorized.

fuzzball context create <context_name> <api_url>
fuzzball context login -u <admin_username> -p '<admin_password>'

Cordon Node

The fuzzball node cordon command records a cordon hold on one or more compute nodes, so the scheduler places no new work on them. Work already running is left alone.

Cordon a Single Node

fuzzball node cordon cn-042 --reason "failing disk, replacing Friday"
cordon hold created for cn-042: 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42

--reason is optional but worth setting: it is what anyone asking why the node is held will read. The hold id is printed because it is what releases that specific hold later, and a generated id is not guessable.

The hold is recorded whatever the node’s current status. Cordoning a node that is Offline is not a no-op: the hold is on the node, so it still repels work when the node comes back. A bulk cordon reports each node separately rather than a total, so a node the server could not hold is named rather than lost in a count.

Cordon Several Nodes

fuzzball node cordon cn-042 cn-043 --reason "rack B maintenance"

Cordon All Nodes

fuzzball node cordon --all

Uncordon Node

The fuzzball node uncordon command releases the cordon holds an administrator placed, returning the node to service.

Uncordon a Single Node

fuzzball node uncordon cn-042
2 cordon hold(s) released for cn-042: 1 hold(s) remain (evict)

The output names the effect of each remaining hold, not just how many remain, because the question after a release is whether the node is back in service.

Uncordon Several Nodes

fuzzball node uncordon cn-042 cn-043

Uncordon All Nodes

fuzzball node uncordon --all

What Uncordon Does Not Release

Uncordoning releases administrator cordon holds. Several other kinds of hold survive it, and each exception is deliberate:

HoldWhy it survives
A hold with another effectAn evict or deprovision hold is not a cordon, and uncordon was not asked about it. Release it with fuzzball node hold release
A hold placed by monitoringIt is released when the alert that placed it resolves. Clearing evidence somebody else gathered should be a deliberate act, not a side effect
A hold placed by the provisionerThe provisioner placed it while taking the instance away, and clears it itself
substrate-versionFuzzball places it while the node’s substrate is too old to pull images, and clears it when the node reports a supported version. See Substrate Version

A Node Cordoned by Annotation

A node carrying the fuzzball.io/scheduler_cordoned_node resource annotation holds a cordon with the fixed id annotation-cordon. That hold is never released by uncordon or by hold release, and the command says so rather than reporting the node clear:

not releasing hold annotation-cordon on 10.0.0.42/9100: cordoned by the fuzzball.io/scheduler_cordoned_node node annotation; remove the annotation instead

Releasing it would not return the node to service in any case. The annotation is set out of band on the substrate resource, so Fuzzball re-derives the hold from it on the node’s next resource report – and for an idle node, whose resource events only fire when its capacity changes, that may be a long time. The node would read as clear and take work in the meantime. Remove the annotation instead.

Reserved Hold IDs

Two hold ids belong to Fuzzball:

Hold idPlaced by
annotation-cordonThe fuzzball.io/scheduler_cordoned_node resource annotation
substrate-versionA node whose substrate is older than its orchestrate extension requires

This only concerns callers of the API directly. PutNodeHold takes an optional hold_id, which makes the write idempotent, and rejects these two:

hold id "substrate-version" is reserved for holds Fuzzball places itself

Both are re-derived from an observation on every report and matched by id alone, so a hold you placed under one of those ids would have its reason overwritten by the next report, and would be released when the observation cleared. Any other id is yours to use.

fuzzball node hold add never sends a hold_id – the server generates one per node, so the same node can be held twice for two reasons and each released independently – so the CLI cannot hit this.

Seeing Why a Node Is Held

fuzzball node show reports every hold on a node, with who placed it and why:

fuzzball node show 10.0.0.42/9100
ID:        10.0.0.42/9100
Hostname:  cn-042
CPU:       x86_64
Status:    Ready
Health:    Healthy (last reported 2026-09-14T14:05:03Z)
Score:     100/100 (penalties: cluster)
Holds: 1
  7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42  cordon  admin  "failing disk, replacing Friday"
    since 2026-09-14T14:02:11Z

The hold history is also on the node event stream, so the question “who took this node out of service, and when” is answerable after the hold has been released and its record is gone.

Working With Holds Directly

fuzzball node cordon, uncordon, evict and deprovision are conveniences over the hold commands for the common cases. fuzzball node hold is the full surface.

Reading holds needs no special privilege: a hold is often the explanation for why somebody’s work is not running, and that is a diagnosis rather than a secret. Placing and releasing holds is restricted to cluster administrators.

Listing Holds

fuzzball node hold list lists every hold on every node in the active cluster:

fuzzball node hold list
NODE     ID                                    EFFECT  SOURCE      REASON
cn-042   7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42  cordon  admin       failing disk
cn-043   a3f9c210-55d4-41b8-8c07-6de0f1c94a77  evict   monitoring  GPU 2 uncorrectable ECC

Naming nodes narrows the listing to them, and --source and --effect narrow it further:

fuzzball node hold list cn-042 --source monitoring --order-by created

--order-by accepts node, id, source, effect and created. With no --order-by the server orders by node, which suits a fleet-wide listing. Sorting by created is newest first, so --order-by created and --order-by -created mean the same thing; use --order-by created:asc for oldest first.

Unlike other list commands, this one sorts by a single field. Comma-separated keys are rejected rather than partly honored:

fuzzball node hold list --order-by node,-created
invalid argument "node,-created" for "--order-by" flag: order by one field at a time; "node,-created" names 2

The table shows five columns; -o json and -o yaml carry the full record, including the node id and the created and refreshed timestamps.

Naming a single node that does not exist is an error. Naming several nodes filters the fleet-wide listing instead, so a mistyped name looks the same as a node that holds nothing. The command notes which names produced no rows, on standard error so -o json output stays clean:

no holds matched cn-042 -- the node may hold nothing, or the name may be wrong

Placing a Hold

fuzzball node hold add places a hold directly. The effect defaults to cordon, so naming one is how you place an evict or deprovision hold:

fuzzball node hold add cn-042 --effect evict --reason "GPU 2 uncorrectable ECC"
evict hold created for cn-042: 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42

The hold’s source is always admin: it is derived from the credential the request was made with rather than taken as a parameter, so automation cannot place a hold that looks like an operator’s. An administrator hold is the one automation may never clear.

Use --all-nodes to hold every node. It is named --all-nodes rather than --all because on hold release, --all already means every matching hold.

Any effect other than cordon prompts for confirmation first, because a fleet-wide cordon is recoverable by uncordoning and a fleet-wide eviction is not. --yes skips the prompt for scripts.

Releasing a Hold

Release one by id, or several with filters:

fuzzball node hold release 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42
fuzzball node hold release --all --node cn-042 --effect evict

An unscoped --all releases matching holds across the whole cluster, so it prompts for confirmation. --yes skips the prompt.

Releasing a Hold Somebody Else Placed

hold release releases administrator holds unless --source names another. That default is deliberate: releasing in bulk is the destructive direction, and a request that also cleared the monitoring stack’s holds would leave a node reading as clear while the fault that produced them is still there.

fuzzball node hold release --all --node cn-042 --source monitoring

A hold named by id but excluded by a filter is reported rather than silently skipped, so a release that did less than asked says so:

not releasing hold a3f9c210-55d4-41b8-8c07-6de0f1c94a77 on 10.0.0.43/9100: hold was placed by MONITORING, not ADMIN

When a Release Exits Non-Zero

A release that was declined fails, so a script chaining fuzzball node uncordon cn-042 && fuzzball workflow start ... stops when the node carries an annotation cordon. A node still held by an evict or deprovision hold is a different case: uncordon releases the cordon holds it was asked about, reports what remains, and exits zero. Check the remaining effects if a script must not submit work to a held node.

SituationExitReported as
A named hold matched but was declined – an annotation cordon, or a hold a filter excludednon-zero1 hold(s) not released
A named hold id matched nothing at allnon-zerono hold matched <id>
A named node was skipped because the write itself failednon-zero1 node(s) not changed
An unscoped --all released nothingzerono matching holds

The last row is the one exception, and it is deliberate: --all means “every matching hold”, so a hold that did not match is not a request that went unmet. Naming a node or a hold id is a request about that node or hold, and silence about it failing would be the wrong answer.

Spotting an Abandoned Hold

Holds do not expire. A monitoring integration that stops re-asserting its holds leaves capacity held with nothing watching it, and nothing else surfaces that. --stale filters to holds that have not been re-asserted within a window:

fuzzball node hold list --source monitoring --stale 1h

Command Behavior Worth Knowing

Two behaviors are deliberate and differ from what you might expect.

Naming no nodes is an error, not a no-op.

fuzzball node cordon
Error: name at least one NODE, or pass --all

A script whose node variable came back empty did not mean to ask for nothing to happen – it has a bug, and an error surfaces that rather than exiting quietly. fuzzball node drain already rejects the same shape, and the server rejects it too, so the two agree. Naming nodes and passing --all is rejected for the same reason – --all wins server-side, so accepting both would act on the whole cluster while the operator watched one node name scroll past.

Repeated cordon places another hold rather than doing nothing. Each call records its own hold with its own reason, so two cordons placed for two reasons stay two reasons. One uncordon releases all of them:

2 cordon hold(s) released for cn-042: 1 hold(s) remain (evict)

This differs from earlier releases, where cordoning an already-cordoned node did nothing.