Evicting and Deprovisioning Nodes
Cordoning stops new work reaching a node. Two stronger actions deal with the work already on it, and with the instance underneath.
| Command | What it does |
|---|---|
fuzzball node cordon | Stops new placement; running work finishes |
fuzzball node evict | Stops new placement and takes running work off |
fuzzball node deprovision | Stops new placement and tears the instance down once it is idle |
fuzzball node drain | Deprecated. Records a cordon hold, and with --active an evict hold as well – use cordon and evict instead |
Each records a hold. See Node Cordoning and Uncordoning for how holds work and how to release them.
These commands require a Fuzzball admin context.
fuzzball context create <context_name> <api_url>fuzzball context login -u <admin_username> -p '<admin_password>'fuzzball node evict cn-042 --reason "GPU 2 throwing uncorrectable ECC"evict hold created for cn-042: 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42
What happens to each running step is what its workflow declared, not what the
eviction wants. A step whose
policy.retry.on
includes node is requeued and runs again from the beginning elsewhere.
Anything else is failed – as is a step that has already exhausted its retry
budget, which is the lower of its own policy.retry.attempts and the cluster’s
maxEvictionRestarts.
An evict hold is not permission to repeat work an author said must not be repeated. Fuzzball does not checkpoint, so re-running a step runs all of it again – including anything it has already written – and only the workflow’s author knows whether that is safe.
Either way the hold’s reason is recorded against the affected work, so a user learns why their job moved or ended:
Work that moved:
node 10.0.0.42/9100 was held: GPU 2 throwing uncorrectable ECC; work moved to another node
Work that was failed because it never declared policy.retry.on: node:
node 10.0.0.42/9100 was held: GPU 2 throwing uncorrectable ECC; job failed. Set
policy.retry.on to node for work that is safe to run again from the beginning
Work that declared it and has been moved as often as the cluster allows – see
maxEvictionRestarts – reports the count instead, because repeating the
advice would tell the author to set something they already set:
node 10.0.0.42/9100 was held: GPU 2 throwing uncorrectable ECC; job failed after 3 node restart(s), the limit for this step
The node is named by its id rather than its host name, because that is the identifier the scheduler holds when it moves the work.
An evict hold survives the eviction. The node stays out of service until the hold is released, even after the node is empty.
That is deliberate. A node emptied because something is wrong with it should not immediately be refilled – the scheduler would place work straight back onto the hardware that just destroyed the previous batch, and the next fault report would evict it again. Returning a node to service is an explicit act:
fuzzball node hold release --all --node cn-042 --effect evict
fuzzball node uncordonwill not return an evicted node to service. It releases cordon holds, and an evict hold is not a cordon. The remaining holds are reported so the situation is visible rather than puzzling:0 cordon hold(s) released for cn-042: 1 hold(s) remain (evict)
fuzzball node deprovision cn-042 --reason "scaling down"deprovision hold created for cn-042: 3f2a6c1e-8b47-4d19-9e5f-0a1b2c3d4e5f
The instance is torn down once it has no allocations left. Until then it keeps running the work it already had.
Waiting is the whole difference between this and --force. An instance running
a five-day job is held for five days: the hold asked for the node to be taken
away when it was done, not for its work to be killed.
On a cloud pool that is usually not what you want, because an instance held for
days is an instance billed for days. --evict takes the work off as well:
fuzzball node deprovision cn-042 --evict --reason "scaling down"deprovision hold created for cn-042: 3f2a6c1e-8b47-4d19-9e5f-0a1b2c3d4e5f
evict hold created for cn-042: 7f3c1a90-2b11-4e6d-9a7f-2f0c5d8e1b42
Two holds rather than one, because they say different things: releasing the eviction should not also cancel the teardown.
The deprovision hold is recorded first even though the eviction is what happens first. Fuzzball refuses to deprovision a node it did not provision, so placing the eviction first would end the work on a statically provisioned node and only then report that it cannot be torn down. In this order that request fails having changed nothing.
The eviction is then scoped to the nodes the deprovision hold actually landed
on. That matters with --all: only a deprovision hold is refused on hardware
Fuzzball did not provision, so evicting the whole fleet unconditionally would
take the work off exactly the nodes that are not going to be torn down.
--force bypasses the hold system and tears the instance down at once, ending
whatever is running on it. This is what fuzzball node deprovision did in
earlier releases.
fuzzball node deprovision cn-042 --forceThis is a behavior change. Without--force,deprovisionnow returns as soon as the hold is recorded, and the capacity is still there until the node empties. A script that deprovisions and then expects the capacity to be gone needs--force.
--force prints the provisioner’s own response per node, unchanged from
earlier releases, so a script parsing that output keeps working.
--force cannot be combined with --all or --evict. Tearing down every
instance in the cluster from one line, leaving no hold record to explain it,
should not be easy. And --evict has nothing to do when the work ends with the
instance either way.
Fuzzball only deprovisions instances it created. A node the site registered itself is refused:
fuzzball node deprovision cn-042Error: cannot deprovision static node cn-042
Destroying hardware somebody owns is not Fuzzball’s to do, and it is not
recoverable. Use fuzzball node cordon or evict to take such a node out of
service instead, and
fuzzball node remove
to drop its registration.
With --all, static nodes are skipped and reported rather than failing the
whole call – otherwise the flag would be unusable on any cluster mixing owned
and provisioned hardware:
skipped cn-042: statically provisioned; Fuzzball did not create this instance
A node skipped this way is not evicted either, even with --evict.
Naming a node that is skipped is different: the command exits non-zero, because
on a named node the server only skips when the write genuinely failed. A script
that chains fuzzball node evict cn-042 && reimage cn-042 therefore stops
rather than reimaging a node that is still running work.
An eviction that fails on a node whose teardown hold was recorded also exits
non-zero, even under --all. Only a deprovision hold is refused on hardware
Fuzzball did not provision, so an eviction that did not land is always a real
failure – the node has a teardown queued and its work was never taken off,
which is the state worth stopping on.
Whether Fuzzball provisioned a node is recorded on the node’s substrate resource, which arrives shortly after the node registers. Ask before that report lands and the command tells you it cannot determine the answer yet, rather than guessing:
rpc error: code = Unavailable desc = node cn-042 has not reported its resources yet, so whether Fuzzball provisioned it is not known; retry
The transport prefix is kept because the condition is retryable: Unavailable
is a code the CLI passes through verbatim, unlike the refusal above.
This is a retryable condition rather than a refusal, and it is reported
separately from the static case above for that reason: a node mid-registration
is not site-owned hardware, and saying it was would send you looking for a
machine you do not have. Under --all such a node is skipped and reported the
way a static one is.
Through a Federate cluster this message arrives wrapped, because a cluster that cannot be reached is reported the same way and the two are not distinguishable from the federate side. Look for the inner message rather than the outer one.
Both commands take --all, and both confirm first because neither is
recoverable:
fuzzball node evict --allThe command asks Take running work off every node in the cluster? and waits
for a Yes/No selection, defaulting to No.
Pass --yes to skip the prompt when running from a script. A prompt with no
way past it makes the command unusable from automation.