Node Management
Cluster administrators can control the scheduling behavior of compute nodes in the Fuzzball cluster. You can mark nodes as unschedulable (cordoned) or schedulable, list all nodes, retrieve node details, and monitor node health.
All node management operations require administrative privileges. Please make sure you have created and logged into your admin context like so:
$ fuzzball context create <context_name> <api_url>
$ fuzzball context login -u <admin_username> -p '<admin_password>'List all compute nodes
$ fuzzball node listGet details about all compute nodes
$ fuzzball node show <node-id>Cordon a node (make unschedulable)
$ fuzzball node cordon <node-id>Uncordon a node (make schedulable)
$ fuzzball node uncordon <node-id>Remove the registration of a node that no longer exists
$ fuzzball node remove <node-id>Compute nodes can be in various states:
| Code | String | Description |
|---|---|---|
| 1 | Not Ready | Node is not ready |
| 2 | Ready | Node is active and operational |
| 3 | Offline | Node is unreachable |
| 4 | Cordoned | Node is unschedulable |
- Node Cordon - Mark nodes as unschedulable to prevent new workflows
- Node Uncordon - Mark nodes as schedulable to allow new workflows
- Node Get - Get details on a specific compute node
- Node List - Get details on all compute nodes
- Node Evict - Take running work off a node and keep it out of service
- Node Deprovision - Tear down a Fuzzball-provisioned instance once it is idle
- Node Hold - List, place and release holds directly
- Node Events - Read a node’s recorded event history
- Node Remove - Delete the registration of a node that no longer exists, releasing the capacity it advertised
Fuzzball publishes facts about its nodes and does not judge them. Nodes send a periodic heartbeat, which is how the control plane tells a live node from a silent one, and export hardware counters for a monitoring stack to scrape. Which counters matter and what they warrant is decided by Prometheus alerting rules, not by thresholds inside the product.
When the monitoring stack concludes a node should stop taking work, it places a hold. That is the only thing that takes a node out of service.
Two things happen with no configuration at all: the scheduler will not place work on a held node, and work stranded on a node that goes offline is released rather than left to time out.
Everything below is documented in Node Health Monitoring.
Seeing what a node is doing
- What Fuzzball observes – node status, when it last reported, and failures attributed to it
- Heartbeat and liveness – how a silent node is detected, and how it comes back
- Viewing node state – CLI and API
- Node event history – 90 days of what happened to a node, and why
Acting on it
- Cordoning a node – stop new work reaching it, leaving what is running alone
- Evicting and deprovisioning – take running work off, or tear the instance down once it is idle
- Nodes that go offline – work stranded on a node that disappeared is released, not left to time out
Being told about it
- Metrics – the cluster endpoint’s series, the node exporter on port 9101, and how they join
- Bundled monitoring stack – optional Prometheus, Alertmanager and Grafana, preconfigured
When something has already gone wrong
- Why a job failed on a node
- Support bundle – one command covering the control plane and the compute nodes
Upgrading
- What was retired – the reliability score, the policy engine, the notification webhook and what replaces each
fuzzball workflow list- View running & completed workflows