Queue and Scheduling Observability
Fuzzball exposes the scheduler queue so you can see what is waiting, what is running, and — for any allocation that is not yet running — why. Three commands cover this:
fuzzball queue show— list allocations in the scheduler queue.fuzzball queue stats— aggregate scheduler counters and tick timings.fuzzball workflow why— the typed reason a workflow’s stages, or one specific allocation, are in their current state.
queue show (cluster-wide) and queue stats are cluster-admin commands. Any
authenticated user can list their own queued allocations with
queue show --mine, and workflow owners can run workflow why on their own
allocations without admin rights.
The cluster-wide views require an admin context:
$ fuzzball context create <context_name> <api_url>
$ fuzzball context login -u <admin_username> -p '<admin_password>'queue show --mine and workflow why on your own allocations work with an
ordinary user context.
The fuzzball queue show command lists allocations tracked by the scheduler,
ordered by effective scheduling priority.
$ fuzzball queue showID | WORKFLOW | JOB | QUEUE | STATUS | PRIORITY | WEIGHT | NODE | PREEMPTIBLE | QUEUE TIME | OWNER
217e5086-3f7f-5f03-a58e-c0f1b2156ff6 | 0d87ddcb-881c-4011-bea7-760bdf136447 | occ | running | running | 0 | 1 | 10.0.0.46/7331 | false | 2026-07-02 10:23:17AM | c30a8e57-cccb-448c-85c7-e49f197ceed4
cf043e9a-2cdf-59de-9ee3-8c41986ffd9d | 75ec995e-ce3f-4ab5-b946-5c789d7aeb70 | head | ready | - | 0 | 1 | | false | 2026-07-02 10:33:41AM | c30a8e57-cccb-448c-85c7-e49f197ceed4
Allocations in the pending queue are blocked on dependencies and not yet
scored, so their priority shows as -.
| Flag | Description |
|---|---|
--queue | Restrict to one queue: pending, ready, running, or finished. |
--priority-min | Only show allocations at or above the given effective priority. |
--owner | Restrict to allocations owned by the given identity id (admin only). |
--mine | Restrict to the calling user’s own allocations. |
--cluster | Query specific owning clusters by id (federate deployments); repeatable or comma-separated. Omit to query all. |
Non-admin users can see where their own allocations sit in the queue with
--mine:
$ fuzzball queue show --mineThe owner scope is enforced by the server, so --mine only ever returns the
caller’s own allocations. Because the scope is fixed to the caller, --mine
cannot be combined with --owner or --cluster.
To reorder your own queued work once you can see it, see Workflow Priority.
Without--mine,queue showreturns the cluster-wide queue and requires a cluster-admin context.--mineworks for any authenticated user.
The fuzzball queue stats command (cluster-admin) reports aggregate scheduler
counters, queue depths, tick timings, and the current backfill reservation.
$ fuzzball queue statsQueue depths
pending 0
ready 1
running 2
finished 3
Tick duration
total ticks 3732
p50 0.006s
p99 0.008s
Preemption / backfill
preemptions total 0
backfilled total 1
Reservation
active true
allocation id cf043e9a-2cdf-59de-9ee3-8c41986ffd9d
computed at 2026-07-02T17:34:17Z
In a federate deployment, pass --cluster to select owning clusters; queue
depths and counters are aggregated across clusters, while tick timings and
reservation state are reported per cluster.
When an allocation is not running, fuzzball workflow why returns a typed
reason and a human-readable detail string.
Give it a workflow id on its own and it explains every stage of that workflow
still waiting on the scheduler, in the same order fuzzball workflow describe
lists them:
$ fuzzball workflow why <workflow-id>Stages the workflow records as finished, failed, or canceled are left out,
because the scheduler discards an allocation once its stage ends. Add --all
to list them too; each one reports that the scheduler is not tracking it.
Workflow: 75ec995e-ce3f-4ab5-b946-5c789d7aeb70
Stage: fetch-inputs (File)
Allocation: 3b1f77a2-49c8-4f0e-bd6a-2c1b9e8d4a10
Reason: ready to schedule
Stage: train (Job)
Allocation: cf043e9a-2cdf-59de-9ee3-8c41986ffd9d
Reason: insufficient resources
Detail: insufficient resources: head-of-queue but no node can fit (est_start=2026-07-03T01:54:25Z)
Pass an allocation id as well to explain just that one:
$ fuzzball workflow why <workflow-id> <allocation-id>Both arguments are UUIDs. A stage’s id is its allocation id, so the ids come
from fuzzball workflow describe <workflow-id> -o json or from
queue show. The default workflow describe table does not include an id
column. --all chooses which stages to explain, so it is rejected in this
form rather than ignored.
Workflow: 75ec995e-ce3f-4ab5-b946-5c789d7aeb70
Allocation: cf043e9a-2cdf-59de-9ee3-8c41986ffd9d
Reason: insufficient resources
Detail: insufficient resources: head-of-queue but no node can fit (est_start=2026-07-03T01:54:25Z)
The output also carries a Received by cluster: <cluster-name> line after the
allocation id, naming the cluster answering the query — the cluster that owns
the allocation, which in a federate deployment is the orchestrate cluster the
workflow was routed to. The line is omitted only when the deployment has no
cluster ID configured (deployments managed by the Fuzzball operator always
have one).
Workflow owners can explain their own allocations; explaining another user’s
allocation requires the appropriate permission. workflow why operates on
allocations that are still in the queue — once an allocation finishes it is no
longer explainable. In the whole-workflow form that shows up per stage rather
than as an error, so one finished stage never hides the rest.
| Reason | Meaning |
|---|---|
READY_TO_SCHEDULE | Eligible; expected to place on an upcoming tick. |
WAITING_DEPENDENCY | Blocked on other stages it depends on (the detail lists them). |
AWAITING_PROVISIONING | A provisioning request is in flight for it. |
PROVISIONING_BLOCKED | Cannot provision a node right now (for example, a dynamic definition at its maxNodes cap). |
INSUFFICIENT_RESOURCES | Head of the queue but no node currently fits; the detail includes an estimated start time when the reservation has one. When no start can be estimated (for example, services with no walltime hold the nodes, occupancy is unattributable, or the allocation’s shape can never fit), the detail reports node availability as unestimable instead. Also reported for a multi-node job whose provisioner definition holds fewer nodes than the job needs: a gang is placed onto one definition, so it can never run there however long it waits, and the detail names the definition and both node counts. The detail also reports what the last placement attempt found – how many nodes were looked at, why each group of them was passed over, and the node with the most free memory and what fell short on it. |
BACKFILL_INELIGIBLE_NO_WALLTIME | Cannot backfill because the allocation has no walltime (for example, a service). |
BACKFILL_INELIGIBLE_WOULD_DELAY_HEAD | Its walltime would run past the reserved start of the blocked head, so backfilling it would delay that job. |
BACKFILL_INELIGIBLE_INTERNAL_NO_SURPLUS | An internal staging job (image or data fetch) whose walltime ceiling runs past the reserved start and whose resources do not fit within the spare capacity the blocked head leaves on the reserved node. |
BACKFILL_NO_FIT | Its walltime fits the gap, but its resource shape does not fit the free capacity before the reservation. Also reported for a candidate that is eligible behind an unestimable reservation (no estimated head start) but has not placed yet. The detail also reports what the last placement attempt found – how many nodes were looked at, why each group of them was passed over, and the node with the most free memory and what fell short on it. |
SERVICE_POOL_AT_CAPACITY | An autoscaled service replica whose provisioner definition cannot host its replica index, so it never places. The detail names the definition and the number of replicas the pool can hold; for a multi-node service it also gives the node cost per replica. Reported while the autoscaler is holding the replica back, and for an already-released replica whose pool shrank beneath it. When the pool can host no replica at all, the detail says why – no nodes registered for the definition, no node large enough for one rank, or simply no room right now. |
LEASE_GRANT_FAILING | The compute node keeps refusing the lease grant, so each placement is rolled back and the allocation returns to the queue. The detail names the node, how many consecutive attempts have failed, and the reason the node gave. See Blocking on a Refused Lease Grant. |
NODE_ANNOTATIONS_MISMATCH | Not enough nodes in the allocation’s provisioner definition both have capacity for the job and carry the node annotations it asks for. The scheduler does not provision a replacement while a node without the annotations still has capacity; the job places once a node carrying the annotations has room. The detail names the requested annotations and how many of the definition’s nodes had capacity but not the annotations. |
RUNNING | Currently running. |
FINISHED | Completed. |
CANCELLED | Cancelled. |
PREEMPTED | Evicted to free capacity for a higher-priority allocation. |