Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Evicting and Deprovisioning Nodes

Cordoning stops new work reaching a node. Two stronger actions deal with the work already on it, and with the instance underneath.

CommandWhat it does
fuzzball node cordonStops new placement; running work finishes
fuzzball node evictStops new placement and takes running work off
fuzzball node deprovisionStops new placement and tears the instance down once it is idle
fuzzball node drainDeprecated. Records a cordon hold, and with --active an evict hold as well – use cordon and evict instead

Each records a hold. See Node Cordoning and Uncordoning for how holds work and how to release them.

Prerequisites

These commands require a Fuzzball admin context.

fuzzball context create <context_name> <api_url>
fuzzball context login -u <admin_username> -p '<admin_password>'

Evicting a Node

fuzzball node evict cn-042 --reason "GPU 2 throwing uncorrectable ECC"
evict hold created for cn-042: 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42

What happens to the running work

What happens to each running step is what its workflow declared, not what the eviction wants. A step whose policy.retry.on includes node is requeued and runs again from the beginning elsewhere. Anything else is failed – as is a step that has already exhausted its retry budget, which is the lower of its own policy.retry.attempts and the cluster’s maxEvictionRestarts.

An evict hold is not permission to repeat work an author said must not be repeated. Fuzzball does not checkpoint, so re-running a step runs all of it again – including anything it has already written – and only the workflow’s author knows whether that is safe.

Either way the hold’s reason is recorded against the affected work, so a user learns why their job moved or ended:

Work that moved:

node 10.0.0.42/9100 was held: GPU 2 throwing uncorrectable ECC; work moved to another node

Work that was failed because it never declared policy.retry.on: node:

node 10.0.0.42/9100 was held: GPU 2 throwing uncorrectable ECC; job failed. Set
policy.retry.on to node for work that is safe to run again from the beginning

Work that declared it and has been moved as often as the cluster allows – see maxEvictionRestarts – reports the count instead, because repeating the advice would tell the author to set something they already set:

node 10.0.0.42/9100 was held: GPU 2 throwing uncorrectable ECC; job failed after 3 node restart(s), the limit for this step

The node is named by its id rather than its host name, because that is the identifier the scheduler holds when it moves the work.

The node stays held

An evict hold survives the eviction. The node stays out of service until the hold is released, even after the node is empty.

That is deliberate. A node emptied because something is wrong with it should not immediately be refilled – the scheduler would place work straight back onto the hardware that just destroyed the previous batch, and the next fault report would evict it again. Returning a node to service is an explicit act:

fuzzball node hold release --all --node cn-042 --effect evict

fuzzball node uncordon will not return an evicted node to service. It releases cordon holds, and an evict hold is not a cordon. The remaining holds are reported so the situation is visible rather than puzzling:

0 cordon hold(s) released for cn-042: 1 hold(s) remain (evict)

Deprovisioning a Node

fuzzball node deprovision cn-042 --reason "scaling down"
deprovision hold created for cn-042: 3f2a6c1e-8b47-4d19-9e5f-0a1b2c3d4e5f

The instance is torn down once it has no allocations left. Until then it keeps running the work it already had.

Waiting for the node to empty

Waiting is the whole difference between this and --force. An instance running a five-day job is held for five days: the hold asked for the node to be taken away when it was done, not for its work to be killed.

On a cloud pool that is usually not what you want, because an instance held for days is an instance billed for days. --evict takes the work off as well:

fuzzball node deprovision cn-042 --evict --reason "scaling down"
deprovision hold created for cn-042: 3f2a6c1e-8b47-4d19-9e5f-0a1b2c3d4e5f
evict hold created for cn-042: 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42

Two holds rather than one, because they say different things: releasing the eviction should not also cancel the teardown.

The deprovision hold is recorded first even though the eviction is what happens first. Fuzzball refuses to deprovision a node it did not provision, so placing the eviction first would end the work on a statically provisioned node and only then report that it cannot be torn down. In this order that request fails having changed nothing.

The eviction is then scoped to the nodes the deprovision hold actually landed on. That matters with --all: only a deprovision hold is refused on hardware Fuzzball did not provision, so evicting the whole fleet unconditionally would take the work off exactly the nodes that are not going to be torn down.

Immediate teardown

--force bypasses the hold system and tears the instance down at once, ending whatever is running on it. This is what fuzzball node deprovision did in earlier releases.

fuzzball node deprovision cn-042 --force
This is a behavior change. Without --force, deprovision now returns as soon as the hold is recorded, and the capacity is still there until the node empties. A script that deprovisions and then expects the capacity to be gone needs --force.

--force prints the provisioner’s own response per node, unchanged from earlier releases, so a script parsing that output keeps working.

--force cannot be combined with --all or --evict. Tearing down every instance in the cluster from one line, leaving no hold record to explain it, should not be easy. And --evict has nothing to do when the work ends with the instance either way.

Statically provisioned nodes

Fuzzball only deprovisions instances it created. A node the site registered itself is refused:

fuzzball node deprovision cn-042
Error: cannot deprovision static node cn-042

Destroying hardware somebody owns is not Fuzzball’s to do, and it is not recoverable. Use fuzzball node cordon or evict to take such a node out of service instead, and fuzzball node remove to drop its registration.

With --all, static nodes are skipped and reported rather than failing the whole call – otherwise the flag would be unusable on any cluster mixing owned and provisioned hardware:

skipped cn-042: statically provisioned; Fuzzball did not create this instance

A node skipped this way is not evicted either, even with --evict.

Naming a node that is skipped is different: the command exits non-zero, because on a named node the server only skips when the write genuinely failed. A script that chains fuzzball node evict cn-042 && reimage cn-042 therefore stops rather than reimaging a node that is still running work.

An eviction that fails on a node whose teardown hold was recorded also exits non-zero, even under --all. Only a deprovision hold is refused on hardware Fuzzball did not provision, so an eviction that did not land is always a real failure – the node has a teardown queued and its work was never taken off, which is the state worth stopping on.

A node that has not reported yet

Whether Fuzzball provisioned a node is recorded on the node’s substrate resource, which arrives shortly after the node registers. Ask before that report lands and the command tells you it cannot determine the answer yet, rather than guessing:

rpc error: code = Unavailable desc = node cn-042 has not reported its resources yet, so whether Fuzzball provisioned it is not known; retry

The transport prefix is kept because the condition is retryable: Unavailable is a code the CLI passes through verbatim, unlike the refusal above.

This is a retryable condition rather than a refusal, and it is reported separately from the static case above for that reason: a node mid-registration is not site-owned hardware, and saying it was would send you looking for a machine you do not have. Under --all such a node is skipped and reported the way a static one is.

Through a Federate cluster this message arrives wrapped, because a cluster that cannot be reached is reported the same way and the two are not distinguishable from the federate side. Look for the inner message rather than the outer one.

Fleet-Wide Actions

Both commands take --all, and both confirm first because neither is recoverable:

fuzzball node evict --all

The command asks Take running work off every node in the cluster? and waits for a Yes/No selection, defaulting to No.

Pass --yes to skip the prompt when running from a script. A prompt with no way past it makes the command unusable from automation.