Fuzzball Documentation
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Node Management

Cluster administrators can control the scheduling behavior of compute nodes in the Fuzzball cluster. You can mark nodes as unschedulable (cordoned) or schedulable, list all nodes, retrieve node details, and monitor node health.

Prerequisites

All node management operations require administrative privileges. Please make sure you have created and logged into your admin context like so:

$ fuzzball context create <context_name> <api_url>

$ fuzzball context login -u <admin_username> -p '<admin_password>'

Basic Commands

List all compute nodes

$ fuzzball node list

Get details about all compute nodes

$ fuzzball node show <node-id>

Cordon a node (make unschedulable)

$ fuzzball node cordon <node-id>

Uncordon a node (make schedulable)

$ fuzzball node uncordon <node-id>

Remove the registration of a node that no longer exists

$ fuzzball node remove <node-id>

Node States

Compute nodes can be in various states:

CodeStringDescription
1Not ReadyNode is not ready
2ReadyNode is active and operational
3OfflineNode is unreachable
4CordonedNode is unschedulable

Node Management Commands

  • Node Cordon - Mark nodes as unschedulable to prevent new workflows
  • Node Uncordon - Mark nodes as schedulable to allow new workflows
  • Node Get - Get details on a specific compute node
  • Node List - Get details on all compute nodes
  • Node Evict - Take running work off a node and keep it out of service
  • Node Deprovision - Tear down a Fuzzball-provisioned instance once it is idle
  • Node Hold - List, place and release holds directly
  • Node Events - Read a node’s recorded event history
  • Node Remove - Delete the registration of a node that no longer exists, releasing the capacity it advertised

Node Health Monitoring

Fuzzball publishes facts about its nodes and does not judge them. Nodes send a periodic heartbeat, which is how the control plane tells a live node from a silent one, and export hardware counters for a monitoring stack to scrape. Which counters matter and what they warrant is decided by Prometheus alerting rules, not by thresholds inside the product.

When the monitoring stack concludes a node should stop taking work, it places a hold. That is the only thing that takes a node out of service.

Two things happen with no configuration at all: the scheduler will not place work on a held node, and work stranded on a node that goes offline is released rather than left to time out.

Everything below is documented in Node Health Monitoring.

Seeing what a node is doing

Acting on it

Being told about it

  • Metrics – the cluster endpoint’s series, the node exporter on port 9101, and how they join
  • Bundled monitoring stack – optional Prometheus, Alertmanager and Grafana, preconfigured

When something has already gone wrong

Upgrading

  • What was retired – the reliability score, the policy engine, the notification webhook and what replaces each
  • fuzzball workflow list - View running & completed workflows