Node Cordoning and Uncordoning
Node cordoning marks compute nodes as unschedulable, preventing new workflows from being placed on them. This is commonly used during maintenance, node upgrades, or when preparing to remove nodes from the cluster.
A cordon is recorded as a hold on the node rather than as the node’s
status. A node’s status reports what Fuzzball has observed about it – whether
it is reporting, whether it is claimed, whether it has gone away – while a
hold records what somebody has asked be done about it. The two answer different
questions, so a cordoned node reports Ready and carries a cordon hold.
fuzzball node listshows holds in their own column, andfuzzball node showlists each hold with its source and reason. A node readingReadywith a cordon hold is not a contradiction: nothing is wrong with the hardware, and nobody wants new work placed on it.Cordoning leaves running work alone. To move that work off a node, or to tear the instance down, see Evicting and Deprovisioning Nodes.
Changes to compute node scheduling require that a fuzzball admin context is set up and authorized.
fuzzball context create <context_name> <api_url>fuzzball context login -u <admin_username> -p '<admin_password>'The fuzzball node cordon command records a cordon hold on one or more compute
nodes, so the scheduler places no new work on them. Work already running is
left alone.
fuzzball node cordon cn-042 --reason "failing disk, replacing Friday"cordon hold created for cn-042: 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42
--reason is optional but worth setting: it is what anyone asking why the node
is held will read. The hold id is printed because it is what releases that
specific hold later, and a generated id is not guessable.
The hold is recorded whatever the node’s current status. Cordoning a node that
is Offline is not a no-op: the hold is on the node, so it still repels work
when the node comes back. A bulk cordon reports each node separately rather
than a total, so a node the server could not hold is named rather than lost in
a count.
fuzzball node cordon cn-042 cn-043 --reason "rack B maintenance"fuzzball node cordon --allThe fuzzball node uncordon command releases the cordon holds an administrator
placed, returning the node to service.
fuzzball node uncordon cn-0422 cordon hold(s) released for cn-042: 1 hold(s) remain (evict)
The output names the effect of each remaining hold, not just how many remain, because the question after a release is whether the node is back in service.
fuzzball node uncordon cn-042 cn-043fuzzball node uncordon --allUncordoning releases administrator cordon holds. Several other kinds of hold survive it, and each exception is deliberate:
| Hold | Why it survives |
|---|---|
| A hold with another effect | An evict or deprovision hold is not a cordon, and uncordon was not asked about it. Release it with fuzzball node hold release |
| A hold placed by monitoring | It is released when the alert that placed it resolves. Clearing evidence somebody else gathered should be a deliberate act, not a side effect |
| A hold placed by the provisioner | The provisioner placed it while taking the instance away, and clears it itself |
substrate-version | Fuzzball places it while the node’s substrate is too old to pull images, and clears it when the node reports a supported version. See Substrate Version |
A node carrying the fuzzball.io/scheduler_cordoned_node resource annotation
holds a cordon with the fixed id annotation-cordon. That hold is never
released by uncordon or by hold release, and the command says so rather
than reporting the node clear:
not releasing hold annotation-cordon on 10.0.0.42/9100: cordoned by the fuzzball.io/scheduler_cordoned_node node annotation; remove the annotation instead
Releasing it would not return the node to service in any case. The annotation is set out of band on the substrate resource, so Fuzzball re-derives the hold from it on the node’s next resource report – and for an idle node, whose resource events only fire when its capacity changes, that may be a long time. The node would read as clear and take work in the meantime. Remove the annotation instead.
Two hold ids belong to Fuzzball:
| Hold id | Placed by |
|---|---|
annotation-cordon | The fuzzball.io/scheduler_cordoned_node resource annotation |
substrate-version | A node whose substrate is older than its orchestrate extension requires |
This only concerns callers of the API directly. PutNodeHold takes an optional
hold_id, which makes the write idempotent, and rejects these two:
hold id "substrate-version" is reserved for holds Fuzzball places itself
Both are re-derived from an observation on every report and matched by id alone, so a hold you placed under one of those ids would have its reason overwritten by the next report, and would be released when the observation cleared. Any other id is yours to use.
fuzzball node hold add never sends a hold_id – the server generates one
per node, so the same node can be held twice for two reasons and each released
independently – so the CLI cannot hit this.
fuzzball node show reports every hold on a node, with who placed it and why:
fuzzball node show 10.0.0.42/9100ID: 10.0.0.42/9100
Hostname: cn-042
CPU: x86_64
Status: Ready
Health: Healthy (last reported 2026-09-14T14:05:03Z)
Score: 100/100 (penalties: cluster)
Holds: 1
7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42 cordon admin "failing disk, replacing Friday"
since 2026-09-14T14:02:11Z
The hold history is also on the node event stream, so the question “who took this node out of service, and when” is answerable after the hold has been released and its record is gone.
fuzzball node cordon, uncordon, evict and deprovision are conveniences
over the hold commands for the common cases. fuzzball node hold is the full
surface.
Reading holds needs no special privilege: a hold is often the explanation for why somebody’s work is not running, and that is a diagnosis rather than a secret. Placing and releasing holds is restricted to cluster administrators.
fuzzball node hold list lists every hold on every node in the active cluster:
fuzzball node hold listNODE ID EFFECT SOURCE REASON
cn-042 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42 cordon admin failing disk
cn-043 a3f9c210-55d4-41b8-8c07-6de0f1c94a77 evict monitoring GPU 2 uncorrectable ECC
Naming nodes narrows the listing to them, and --source and --effect narrow
it further:
fuzzball node hold list cn-042 --source monitoring --order-by created--order-by accepts node, id, source, effect and created. With no
--order-by the server orders by node, which suits a fleet-wide listing.
Sorting by created is newest first, so --order-by created and
--order-by -created mean the same thing; use --order-by created:asc for
oldest first.
Unlike other list commands, this one sorts by a single field. Comma-separated keys are rejected rather than partly honored:
fuzzball node hold list --order-by node,-createdinvalid argument "node,-created" for "--order-by" flag: order by one field at a time; "node,-created" names 2
The table shows five columns; -o json and -o yaml carry the full record,
including the node id and the created and refreshed timestamps.
Naming a single node that does not exist is an error. Naming several nodes
filters the fleet-wide listing instead, so a mistyped name looks the same as a
node that holds nothing. The command notes which names produced no rows, on
standard error so -o json output stays clean:
no holds matched cn-042 -- the node may hold nothing, or the name may be wrong
fuzzball node hold add places a hold directly. The effect defaults to
cordon, so naming one is how you place an evict or deprovision hold:
fuzzball node hold add cn-042 --effect evict --reason "GPU 2 uncorrectable ECC"evict hold created for cn-042: 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42
The hold’s source is always admin: it is derived from
the credential the request was made with rather than taken as a parameter, so
automation cannot place a hold that looks like an operator’s. An administrator
hold is the one automation may never clear.
Use --all-nodes to hold every node. It is named --all-nodes rather than
--all because on hold release, --all already means every matching hold.
Any effect other than cordon prompts for confirmation first, because a
fleet-wide cordon is recoverable by uncordoning and a fleet-wide eviction is
not. --yes skips the prompt for scripts.
Release one by id, or several with filters:
fuzzball node hold release 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42fuzzball node hold release --all --node cn-042 --effect evictAn unscoped --all releases matching holds across the whole cluster, so it
prompts for confirmation. --yes skips the prompt.
hold release releases administrator holds unless --source names
another. That default is deliberate: releasing in bulk is the destructive
direction, and a request that also cleared the monitoring stack’s holds would
leave a node reading as clear while the fault that produced them is still
there.
fuzzball node hold release --all --node cn-042 --source monitoringA hold named by id but excluded by a filter is reported rather than silently skipped, so a release that did less than asked says so:
not releasing hold a3f9c210-55d4-41b8-8c07-6de0f1c94a77 on 10.0.0.43/9100: hold was placed by MONITORING, not ADMIN
A release that was declined fails, so a script chaining
fuzzball node uncordon cn-042 && fuzzball workflow start ... stops when the
node carries an annotation cordon. A node still held by an evict or deprovision
hold is a different case: uncordon releases the cordon holds it was asked
about, reports what remains, and exits zero. Check the remaining effects if a
script must not submit work to a held node.
| Situation | Exit | Reported as |
|---|---|---|
| A named hold matched but was declined – an annotation cordon, or a hold a filter excluded | non-zero | 1 hold(s) not released |
| A named hold id matched nothing at all | non-zero | no hold matched <id> |
| A named node was skipped because the write itself failed | non-zero | 1 node(s) not changed |
An unscoped --all released nothing | zero | no matching holds |
The last row is the one exception, and it is deliberate: --all means “every
matching hold”, so a hold that did not match is not a request that went unmet.
Naming a node or a hold id is a request about that node or hold, and silence
about it failing would be the wrong answer.
Holds do not expire. A monitoring integration that stops re-asserting its holds
leaves capacity held with nothing watching it, and nothing else surfaces that.
--stale filters to holds that have not been re-asserted within a window:
fuzzball node hold list --source monitoring --stale 1hTwo behaviors are deliberate and differ from what you might expect.
Naming no nodes is an error, not a no-op.
fuzzball node cordonError: name at least one NODE, or pass --all
A script whose node variable came back empty did not mean to ask for nothing
to happen – it has a bug, and an error surfaces that rather than exiting
quietly. fuzzball node drain already rejects the same shape, and the server
rejects it too, so the two agree. Naming nodes and passing --all is
rejected for the same reason – --all wins server-side, so accepting both
would act on the whole cluster while the operator watched one node name scroll
past.
Repeated cordon places another hold rather than doing nothing. Each call
records its own hold with its own reason, so two cordons placed for two reasons
stay two reasons. One uncordon releases all of them:
2 cordon hold(s) released for cn-042: 1 hold(s) remain (evict)
This differs from earlier releases, where cordoning an already-cordoned node did nothing.