Workflow Endpoints
When a workflow service exposes network endpoints (configured via the network section in your
workflow definition), the fuzzball endpoint command (alias fuzzball endpoints) and its
subcommands let you list the endpoints you have access to, inspect a single endpoint, mint an
access token for a non-public endpoint, and manage persistent endpoints.
Endpoints normally live and die with their workflow. See Persistent endpoints for named endpoints whose URL stays the same across submissions.
The endpoint commands used to live underfuzzball workflow endpoints(orfuzzball workflow endpoint). Those forms still work but are deprecated; usefuzzball endpointinstead.
URL format change:pathendpoints are now served at/endpoints/{endpoint-id}. Earlier releases used/endpoints/accounts/{account-id}/workflows/{workflow-id}/{service}/{endpoint}, and URLs in that form no longer resolve – the proxy answers404. Re-read the URL fromfuzzball endpoint list, the Web UI, or theFB_ENDPOINT_URL_*environment variables, and update any bookmarks, scripts, or dashboards that stored the old form. A job that built the path of another service’s endpoint from its workflow and account IDs reads it fromFB_SERVICE_<SERVICE_NAME>_ENDPOINT_PATH_<ENDPOINT_NAME>instead.subdomainendpoints are unaffected: their URL was already keyed on the endpoint ID.
fuzzball endpoint list returns the workflow service endpoints currently alive on the cluster
that you have access to, and the persistent endpoints you can reach whether or not a workflow is
attached to them. Access is scoped to the groups (accounts) you belong to and to your
organization.
An endpoint is alive only while the workflow that declared it is still running. Once that workflow
reaches a terminal state — Finished, Errored, or Canceled (see
Workflow Status)
— its endpoints are removed, whether the service shut down normally or exited unexpectedly. A
persistent endpoint is not removed: it stays listed with a blank WORKFLOW column.
fuzzball endpoint get rechecks the owning workflow on every call and stops returning
the endpoint immediately. The list above does not recheck each row, so an endpoint whose workflow
has just ended can still appear there briefly; get is authoritative.
$ fuzzball endpoint list
ID NAME WORKFLOW SERVICE REPLICA TYPE SCOPE PERSISTENT UPDATE URL
endpoint-1a2b3c4d web 3f8c1d2e... jupyter path user No https://endpoints.example.com/endpoints/endpoint-1a2b3c4d.../
endpoint-5e6f7a8b inference 9a0b7c6d... inference subdomain group Yes ready https://endpoint-5e6f7a8b....endpoints.example.com/
endpoint-9c0d1e2f dashboard subdomain group Yes start https://endpoint-9c0d1e2f....endpoints.example.com/The REPLICA column is blank for an ordinary endpoint and for an autoscaled pool’s own endpoint.
It carries the replica number only on the per-replica endpoints described in
“Addressing individual replicas” below.
Flags:
| Flag | Description |
|---|---|
--persistent | Only persistent endpoints; --persistent=false lists only ephemeral ones. Omit to list both |
--name | Only endpoints with this name |
--cluster | Only endpoints owned by this orchestrate cluster ID |
-p, --page-size | Server batch size per request |
-m, --max-pages | Hard cap on the number of pages fetched (0 = no cap) |
-o, --output | Output format (table default, or json / yaml) |
Use -o json or -o yaml for scripting; the structured output includes additional fields such as
the workflow ID and the owning account and organization.
$ fuzzball endpoint list -o jsonfuzzball endpoint get shows the details of one endpoint by its ID, including its URL:
$ fuzzball endpoint get ENDPOINT_IDEndpoints whose scope is not public require a bearer token. fuzzball endpoint generate-token
mints a token bound to the endpoint and carrying your identity:
$ fuzzball endpoint generate-token ENDPOINT_IDPresent the returned token to the endpoint using the Authorization or FB-Authorization HTTP
header.
A token can only be minted while the owning workflow is still running. Once the workflow reaches a terminal state the request is refused, and tokens already issued for its endpoints stop working as soon as the endpoint is removed. Mint a fresh token against a running workflow instead of holding a long-lived one across workflow restarts.
A persistent endpoint is the exception: its token can be minted whether
or not a workflow is attached to it, and keeps working as workflows attach and detach, until the
token expires or the endpoint is deleted. Changing the endpoint’s scope also invalidates every
token minted before the change, so narrowing who can reach the endpoint takes effect at once; mint
a new token afterwards.
On a federated deployment, an endpoint belonging to a workflow the federate cluster has not yet learned about can still be read and minted; the cluster running the workflow remains authoritative and stops serving the endpoint once it is gone.
The --expiration flag sets the token lifetime. It accepts the units understood by Go durations
(s, m, h) plus day, week, month, and year units — 7d, 2w, 1mo, 1y (and long forms
like 7days). When omitted, the server default lifetime is used:
# Valid for 24 hours
$ fuzzball endpoint generate-token ENDPOINT_ID --expiration 24h
# Valid for 7 days
$ fuzzball endpoint generate-token ENDPOINT_ID --expiration 7d
# Valid for 1 month
$ fuzzball endpoint generate-token ENDPOINT_ID --expiration 1moIn addition to querying endpoint URLs with the CLI, the Fuzzball Web UI provides a Connect button on workflow detail pages and in the workflows list. When a workflow is running and has at least one named endpoint defined, clicking Connect opens that endpoint in a new browser tab. On the detail page, a workflow exposing several endpoints turns the button into a menu listing each one, so any of them can be opened directly.
This provides a convenient shortcut for accessing web-based services without needing to copy and paste URLs. For details on which endpoints are offered, and for information on client-script mode, see the Connecting to Workflow Services page.
The Web UI Endpoints page lists the endpoints you have access to, filtered by name, by persistence and, on a federate cluster, by cluster. It also creates, edits and deletes persistent endpoints, following the same rules as the CLI.
This form is deprecated. Usefuzzball endpoint list(filtered with--nameif needed) instead.
Invoked with a workflow ID (and optional service and endpoint names), fuzzball workflow endpoints prints the endpoint URL(s) derived from that workflow’s specification:
$ fuzzball workflow endpoints WORKFLOW_ID [SERVICE_NAME] [ENDPOINT_NAME]Arguments:
| Argument | Required | Description |
|---|---|---|
WORKFLOW_ID | Yes | The ID of a running workflow |
SERVICE_NAME | No | Filter results to a specific service |
ENDPOINT_NAME | No | Filter results to a specific endpoint within a service |
Get all endpoints for a workflow:
$ fuzzball workflow endpoints <workflow_id>
SERVICE ENDPOINT URL
jupyter web https://endpoints.example.com/endpoints/endpoint-1a2b3c4d.../
api-server api https://endpoints.example.com/endpoints/endpoint-5e6f7a8b.../Get all endpoints for a specific service, or a specific endpoint by name:
$ fuzzball workflow endpoints <workflow_id> jupyter
$ fuzzball workflow endpoints <workflow_id> jupyter webOutput the results as JSON (useful for scripting):
$ fuzzball workflow endpoints <workflow_id> -o jsonEndpoints are declared in the network section of a service definition. For example, a Jupyter
service that exposes a web endpoint:
services:
jupyter:
image:
uri: docker://jupyter/base-notebook:latest
script: |
#!/bin/bash
jupyter lab --ip=0.0.0.0 --no-browser
resource:
cpu:
cores: 2
memory:
size: 4GB
network:
ports:
- name: web
port: 8888
protocol: tcp
endpoints:
- name: web
port-name: web
protocol: https
type: path
scope: user
See the workflow syntax reference for full details on
the network configuration options.
The type field selects how an endpoint is addressed:
path(the default) serves the endpoint beneath a single cluster-wide hostname, ashttps://endpoints.<domain>/endpoints/<endpoint-id>/.subdomaingives the endpoint its own hostname,https://endpoint-<id>.endpoints.<domain>/.
Prefer subdomain for applications that assume they are served from the root of a domain. Many web
applications build absolute links or set cookies in ways that break beneath a path prefix, and a
dedicated hostname avoids the problem entirely.
The trade-off is a DNS requirement. Because a new hostname is generated for each endpoint,
subdomain endpoints require the cluster domain to resolve wildcards — every name matching
*.endpoints.<domain> must reach the cluster. Managed and cloud deployments handle this already.
For a single-host Docker Compose deployment, see
DNS resolution; note
that a /etc/hosts entry alone is not sufficient, since hosts files cannot express wildcards.
If asubdomainendpoint URL does not resolve whilepathendpoints on the same cluster work, the cluster domain is missing wildcard DNS rather than the endpoint being misconfigured.
An endpoint can carry annotations — arbitrary key-value pairs that Fuzzball stores with the
endpoint and returns from fuzzball endpoint list and get. Fuzzball does not interpret
them. They exist so a client that discovers endpoints through the API can tell which ones it is
meant to consume, and skip the rest:
services:
vllm:
image:
uri: docker://vllm/vllm-openai:latest
network:
ports:
- name: http
port: 8000
protocol: tcp
endpoints:
- name: openai
port-name: http
protocol: https
type: subdomain
scope: group
annotations:
example.com/api: openai
example.com/model: llama-3.1-8b
The annotations come back in the structured output:
$ fuzzball endpoint get ENDPOINT_ID
annotations:
example.com/api: openai
example.com/model: llama-3.1-8b
endpointName: openai
id: endpoint-5e6f7a8b
...They are not shown in the list table; use -o yaml or -o json to see them for a whole list.
Limits. Each endpoint may declare at most 32 annotations. A key may be up to 253 bytes and a
value up to 1024, with all keys and values together limited to 8192 bytes. A key is an optional
DNS subdomain prefix followed by / and an alphanumeric name of up to 63 characters, with dashes,
underscores and dots allowed inside the name – for example example.com/api or role. The
fuzzball.io prefix, and any subdomain of it, is reserved for the platform and rejected in your own
annotations. A workflow that exceeds any of these limits is rejected at submit time.
On a persistent endpoint, a workflow’s annotations apply only while that workflow is attached, and are removed when it ends.
An ordinary endpoint belongs to the workflow that declares it: its ID is derived from the workflow ID, so every submission produces a new URL, and the endpoint disappears when the workflow stops.
A persistent endpoint is a named endpoint that outlives the workflows serving it. You create it once, and its ID and URL stay the same for as long as it exists. Each workflow you submit with a matching endpoint name attaches to it, replacing the previous one. That makes it usable as a long-lived address for a service you redeploy: an inference server, a dashboard, an API you publish to colleagues.
While no workflow is attached, the endpoint keeps existing but has nothing behind it, and requests
to it return 503.
Create a persistent endpoint with fuzzball endpoint create, giving it a name:
$ fuzzball endpoint create inferenceThe name must be unique among your persistent endpoints on the cluster that owns the endpoint. On a federated deployment you can use the same name on several clusters; see Federated clusters.
Flags:
| Flag | Description |
|---|---|
--type | subdomain or path (default subdomain) |
--protocol | http or https (default http) |
--scope | Who can reach it: user, group, organization or public (default group) |
--strategy | Update strategy: start or ready (default ready) |
--cluster | The orchestrate cluster that will own the endpoint. Defaults to the local cluster; required on a federate cluster |
You can also create persistent endpoints from the Web UI Endpoints page, where a name is required.
A service endpoint attaches to a persistent endpoint when it is marked persistent: true: it then
serves your persistent endpoint whose name matches its own name. Its type, protocol and
scope come from the persistent endpoint. An endpoint without persistent: true belongs to its
workflow alone, even when its name matches one of your persistent endpoints.
services:
inference:
image:
uri: docker://vllm/vllm-openai:latest
persist: true
network:
ports:
- name: http
port: 8000
protocol: tcp
endpoints:
- name: inference
port-name: http
persistent: true
readiness-probe:
http-get:
path: /health
port: 8000
Resubmitting this definition – with a new image, new resources, whatever – attaches the new workflow to the same endpoint, and its URL does not move.
The following rules apply when attaching, and a submission that breaks one is rejected with an error naming the field to fix:
- The persistent endpoint has to exist: create it first with
fuzzball endpoint create. An endpoint markedpersistent: truewhose name matches none of yours is rejected rather than given a URL of its own. - Leave
type,protocolandscopeout to take the persistent endpoint’s. Any of them that is declared must match the persistent endpoint’s: a workflow declaringtype: pathcannot attach to asubdomainpersistent endpoint. - Only the user who created a persistent endpoint can attach workflows to it. Another user’s persistent endpoint of the same name is not a match.
- Two services of one workflow cannot attach to the same persistent endpoint.
- An autoscaled service cannot declare a persistent endpoint.
Workflow annotations on an attached endpoint apply only while that workflow is attached, and are
removed when it ends.
The id and update-strategy endpoint fields of earlier pre-releases are no longer part of the
workflow syntax: the endpoint is named by its name, and its update strategy is set on the
persistent endpoint itself.
When a new workflow attaches to a persistent endpoint that another workflow is serving, the
endpoint’s update strategy decides when the running workflow is cancelled. The strategy is a
property of the persistent endpoint, set with --strategy on create or update:
| Strategy | Behavior |
|---|---|
ready (default) | The running workflow keeps serving until the new service is ready. The endpoint then moves and the previous workflow is cancelled. No gap in service, but both workflows hold resources at the same time. |
start | The running workflow is cancelled once the new submission is accepted, freeing its resources for the new one. Use this when the cluster cannot run both at once. The endpoint returns 503 until the new workflow serves it. A submission rejected at validation leaves the running workflow serving. |
Under ready the handover waits for the service to report ready. A service that declares a
readiness-probe reports ready when that probe passes; a service without one reports ready as soon
as its container starts. Declaring a probe is what makes the handover meaningful – without one the
endpoint can move across before the service is able to answer requests – but it is not required.
Either way, the whole previous workflow is cancelled, not only the service that was attached.
A persistent endpoint belongs to the workflow submitted last that attaches to it. If you submit twice before the first submission is ready, the first one is cancelled when the second is accepted, whatever the strategy: it is not serving yet, and it could never take the endpoint over from the newer submission, even if it became ready first.
Because attaching cancels the workflow currently serving an endpoint, fuzzball workflow start,
fuzzball run and fuzzball workflow catalog start first list the running workflows a submission
would cancel and ask for confirmation:
$ fuzzball workflow start inference.yaml
Submitting this workflow takes over persistent endpoints and cancels the running workflows serving them:
PERSISTENT ENDPOINT SERVICE STRATEGY WORKFLOW TO CANCEL
inference (endpoint-5e6f...) inference ready (cancelled once the new service is ready) inference (9a0b7c6d...)
Continue? [y/N]Pass --yes (-y) to skip the prompt. Without a terminal to prompt on – in a script or CI job –
the command fails unless --yes is given.
The Web UI shows a confirmation dialog when starting, rerunning or running a template would cancel
such workflows. Over the MCP server, workflow_start and
catalog_start require confirm_endpoint_takeover=true in that case.
The check validates the workflow the way submitting it would, so a workflow it rejects – an
endpoint type that does not match the persistent endpoint’s, for example – is reported with the
reason and not submitted: fix the workflow and submit again. If the check cannot run at all, for
instance because the server is unreachable, nothing is submitted either, because there is no
telling which running workflows the submission would cancel. To submit anyway in that case, pass
--yes, choose Submit anyway in the Web UI dialog, or call the MCP tool again with
confirm_endpoint_takeover=true.
The server enforces the confirmation too. A submission that would cancel a running workflow is
refused, and nothing is cancelled, unless the StartWorkflow request sets
confirm_endpoint_takeover. The clients above set it once you confirm. A script calling the API
directly has to set it, and a takeover that appears between the check and the submission – another
workflow attaching to the endpoint meanwhile – is refused rather than let through: submit again to
see it and confirm.
fuzzball endpoint update changes one or more endpoints. Only the fields whose flags are given
change, and what can change depends on the endpoint:
| Endpoint | What can change |
|---|---|
| Ephemeral | It can only be made persistent with --persistent, optionally choosing --strategy (default ready) |
| Persistent, attached to a workflow, or claimed by one that is still starting | Only --strategy |
| Persistent, unattached | --type, --name, --scope and --strategy |
# Keep a running workflow's endpoint URL beyond the workflow
$ fuzzball endpoint update ENDPOINT_ID --persistent
# Cancel the serving workflow at submission from now on
$ fuzzball endpoint update ENDPOINT_ID --strategy start
# Rename an unattached persistent endpoint and open it to the organization
$ fuzzball endpoint update ENDPOINT_ID --name inference --scope organizationA persistent endpoint cannot be made ephemeral again; delete it instead. --name takes a single
ENDPOINT_ID, since names are unique. An endpoint made persistent keeps its name, which must not
already be taken by another of your persistent endpoints. Endpoints of an autoscaled service cannot
be made persistent.
The Web UI Endpoints page edits endpoints following the same rules.
A persistent endpoint belongs to the user who created it, so only they can attach workflows to it,
change it or delete it. Other users can still reach it according to its scope.
Deleting a persistent endpoint also cancels the workflow attached to it:
$ fuzzball endpoint delete ENDPOINT_ID
Deleted persistent endpoint endpoint-1a2b3c4d...
Canceled workflow 3f8c... that was attached to the endpointIt also cancels any workflow that attached to the endpoint and is still starting, which would otherwise run with no endpoint to serve. If one of those cancellations fails, the endpoint is kept; run the delete again.
Persistent endpoints are not cleaned up automatically. Delete the ones you no longer need.
A persistent endpoint lives on one orchestrate cluster for its whole life, chosen with --cluster
when it is created. A workflow that attaches to persistent endpoints is therefore always scheduled
on the cluster owning them, whatever the scoring would otherwise have chosen, and a workflow may not
attach to persistent endpoints spread across more than one cluster.
Names are unique per cluster, so you may have a persistent endpoint of the same name on several
clusters. When a workflow declares an endpoint with such a name, submit it to one of those clusters
with --cluster-id: the workflow attaches to that cluster’s endpoint. Submitted without a cluster,
the workflow is rejected with the list of clusters holding the name.
An autoscaled service runs a pool of interchangeable replicas rather than a single container, so its endpoint works differently: one endpoint addresses the whole pool. The endpoint exists for the life of the workflow, its URL never changes, and Fuzzball forwards each request to one of the replicas that is ready at that moment, cycling through them in turn.
services:
vllm:
image:
uri: docker://vllm/vllm-openai:latest
autoscaler:
replicas:
min: 1
max: 4
network:
ports:
- name: http
port: 8000
protocol: tcp
endpoints:
- name: api
port-name: http
protocol: http
type: subdomain
scope: group
Callers of https://<endpoint-id>.endpoints.<cluster-domain>/ reach the pool without knowing how
many replicas exist. Scaling up adds replicas to the rotation as they pass their readiness probe;
scaling down removes a replica from the rotation when its drain period begins, so requests already
in flight to it can finish while new ones go elsewhere.
Two rules apply to a pool’s endpoints, and a workflow that breaks either is rejected at submission:
- The endpoint must set
type: subdomain. Apathendpoint addresses one backend and has no way to select among a pool’s replicas. - The endpoint must reference a named port. Each replica serves on a randomly assigned host port, and Fuzzball tracks those ports per named port.
A pool with replicas.min: 0 keeps its endpoint while it is idle at zero replicas. A request
arriving then cannot be served, because there is no replica to serve it. Fuzzball answers
503 Service Unavailable with a Retry-After header and, at the same time, starts one replica.
The request is not held open: starting a replica of a large model takes minutes, far longer than
most clients wait. Retry after the cold start and the request is served normally.
This works every time the pool returns to zero, not only on the first request.
Two rules apply to waking a pool:
- The request must be authenticated. Starting a replica consumes cluster resources, GPUs in the
usual case, so only a caller Fuzzball can identify may trigger it. An endpoint with
scope: publicis served without authentication and therefore never wakes its pool. Give a pool you want woken on demand auser,groupororganizationscope, and reach it with a user token or an endpoint access token. - One replica per wake. A burst of requests against an idle pool starts a single replica rather
than one per request: further wakes for that endpoint are suppressed for a minute, or until the
replica that wake asked for has served a request. A pool that scales back to zero is therefore
wakeable again as soon as it goes idle, rather than waiting out a cooldown its own replica already
satisfied. Once a replica is running, the service’s own
scale-uptriggers grow the pool to match its load, up toreplicas.max.
Each wake is recorded on the workflow, so fuzzball workflow events <workflow-id> shows when a
request woke an idle pool.
Some callers do their own load balancing – an API gateway that tracks each backend’s health and
spreads requests according to its own policy, for example. Setting per-replica: true on a pool’s
endpoint gives every live replica its own endpoint as well, so such a caller sees one target per
replica instead of one target for the pool.
endpoints:
- name: api
port-name: http
protocol: http
type: subdomain
scope: group
per-replica: true
The pool’s own endpoint and the per-replica endpoints coexist; per-replica adds endpoints and
takes nothing away. Use the pool’s URL when you want Fuzzball to balance across the replicas, and
the per-replica URLs when the caller balances for itself.
A replica’s endpoint appears when the replica passes its readiness probe and disappears when the
autoscaler begins draining it, so a caller that refreshes its list of endpoints follows the pool as
it scales, and a retiring replica leaves the rotation before it is stopped. Both kinds of endpoint
come back from fuzzball endpoint list and the /endpoints API; a per-replica endpoint
reports replicaIndex (the replica it addresses, numbered from 1) and poolEndpointId (the pool’s
own endpoint), so a caller can group a pool’s replicas into one entry. replicaIndex is the
REPLICA column of the endpoint list table; poolEndpointId is available via -o json and
-o yaml.
per-replica: true requires type: subdomain and a service that defines autoscaler.replicas; a
workflow that sets it anywhere else is rejected at submission.
Every replica endpoint is a hostname of its own, created and removed as the pool scales, so a
cluster serving them needs the wildcard DNS that any subdomain endpoint needs – see
Choosing between path and subdomain.
Per-replica endpoints are local to the cluster running the workflow. In a federated deployment they are not published to the federate cluster, whose endpoint list carries the pool’s own endpoint instead – one address that always resolves to a ready replica.
On an endpoint whose scope is not public, Fuzzball authenticates every request before proxying it
and then removes the Authorization header, so a caller’s own token never reaches your service. To
let the service still know who it is serving, Fuzzball forwards a signed assertion of the caller’s
identity in the X-Fuzzball-Caller-Identity header.
This is what lets a service offer per-user behaviour – usage accounting, per-user access rules, audit trails – without asking the caller for a second credential of its own.
The header holds a JWT signed by the cluster. Its claims:
| Claim | Meaning |
|---|---|
aud | endpoint:<endpoint id> – what marks this as an assertion rather than a bearer token. Always check it against FB_ENDPOINT_ID |
sub | The calling user’s id |
email | The calling user’s email address |
account_id | For a user, the group the caller is acting in. For a caller presenting an endpoint token, the endpoint’s own account |
organization_id | The caller’s organization |
endpoint_id | The endpoint the assertion was minted for; the same id as in aud |
exp | Expiry, two minutes after minting |
iat | When it was minted |
iss | The cluster’s internal issuer, for information only. Never fetch from it – see below |
type | Always Bearer. An artefact of the shared minting path; it does not mean the assertion may be used as a credential |
cluster_id_origin | The cluster that issued it |
The assertion is signed with ES256. Fuzzball publishes the keys that verify it into the node trust store, the same directory it mounts the cluster CA into, so a service reads them from a local file rather than fetching them over the network:
/run/fuzzball-substrate/trusted-certs/signing-keys.json
The file is an ordinary JWKS document; match the assertion’s kid against it. Reject any assertion
whose signature does not verify, whose exp has passed, or whose aud is not your own endpoint.
Do not fetch keys from the URL in the assertion’s iss claim. Until the signature has been checked,
every claim is whatever the sender chose to put there, so a forged assertion would simply name a key
server the forger controls. The local file is the only source to trust.
Requireaud. An endpoint token – the credentialfuzzball endpoint generate-tokenhands out – is signed by the same cluster key, names the same issuer, and carries the same identity claims and the sameendpoint_id.audis the only claim that separates the two. A service that verifies the signature without checkingaudwill accept an endpoint token as an assertion naming whoever minted it, and endpoint tokens are meant to be shared and may be minted with any lifetime.
Your endpoint’s own id arrives in the environment, so the service does not have to derive it:
| Variable | Value |
|---|---|
FB_ENDPOINT_ID | The id of the service’s first endpoint |
FB_ENDPOINT_ID_<NAME> | The id of the endpoint named <NAME>, upper-cased with punctuation replaced by _ |
FB_ENDPOINT_REPLICA_ID, FB_ENDPOINT_REPLICA_ID_<NAME> | On a per-replica endpoint of a replica pool, this replica’s own endpoint id |
A pooled service is reachable two ways – through the pool’s endpoint and through its own
per-replica endpoint – and the assertion names whichever the request arrived on. Accept both
FB_ENDPOINT_ID and FB_ENDPOINT_REPLICA_ID when you set one up.
The example uses PyJWT with its cryptography extra
(pip install "pyjwt[crypto]" – note that the PyPI package jwt is a different library):
import json
import os
import jwt
KEYS_PATH = "/run/fuzzball-substrate/trusted-certs/signing-keys.json"
MY_AUDIENCES = [
f"endpoint:{os.environ[name]}"
for name in ("FB_ENDPOINT_ID", "FB_ENDPOINT_REPLICA_ID")
if os.environ.get(name)
]
raw = request.headers.get("X-Fuzzball-Caller-Identity")
if raw is None or not MY_AUDIENCES:
# No header, or nothing to bind it to: treat the caller as anonymous
# rather than verifying against an empty audience list.
return handle_anonymous_request()
with open(KEYS_PATH) as handle:
keys = jwt.PyJWKSet.from_dict(json.load(handle))
kid = jwt.get_unverified_header(raw)["kid"]
key = next(k for k in keys.keys if k.key_id == kid).key
claims = jwt.decode(
raw,
key,
algorithms=["ES256"],
audience=MY_AUDIENCES,
options={"require": ["exp", "sub", "aud"]},
)
user_id = claims["sub"]
iss is deliberately not pinned. The key file is what establishes the signer – only the cluster
can produce a signature that verifies against it – and aud is what binds the assertion to this
endpoint, so an issuer check adds nothing. There is also no issuer value to check against inside a
container: the claim names an internal address a workflow container cannot resolve.
A rotated signing key does not reach a running node. A node writes this file once, when its Fuzzball extension starts. The cluster serves the current key immediately, but an existing node keeps the set it was given, so after rotating the cluster signing key the extension has to be restarted on every node before services can verify again. Until then a service sees assertions signed with a
kidabsent from its file and must reject them.The same applies to withdrawing a key. A node served no keys removes the file, but only at its next extension start – a running node keeps what it already has, so a withdrawn key stays trusted there until the extension restarts.
The assertion is not a credential. Fuzzball refuses it as a bearer token, so a service cannot use one to call the Fuzzball API as the caller. Treat it as a statement about who called, nothing more.
A
publicendpoint authenticates nobody, so no assertion is sent. Your service must treat a missing header as an anonymous caller rather than failing closed on a value it expects to be present.Fuzzball always discards an
X-Fuzzball-Caller-Identityheader supplied by the client, at every scope, so a caller cannot choose the identity your service sees. Never trust the header without verifying its signature: a service also listens on its node’s port, where requests arrive without passing through the endpoint proxy at all.
Three other headers are added on every proxied request – X-Fuzzball-Workflow-ID,
X-Fuzzball-Account-ID and X-Fuzzball-Service. These describe the endpoint’s own workflow, not
the caller, and are unsigned; use them for tracing, not for authorization.
When a workflow service runs, Fuzzball automatically injects environment variables for each configured endpoint. These variables allow your service code to discover the public URLs assigned to its endpoints.
For each endpoint, two environment variables are provided:
FB_ENDPOINT_PATH_<ENDPOINT_NAME>— the full URL path to the endpointFB_ENDPOINT_URL_<ENDPOINT_NAME>— the complete URL including protocol and host
The <ENDPOINT_NAME> suffix is derived from the endpoint’s name field by:
- Converting the name to uppercase
- Replacing any character that is not a letter, digit, or underscore with an underscore (
_)
This transformation ensures that the generated environment variable names are valid POSIX shell identifiers and compatible with all container runtimes.
Because of it, distinct names can end up with the same suffix – lab-ui and lab_ui both become
LAB_UI. A workflow in which two endpoints of one service, or two services, would share a suffix is
rejected at submission; rename one of them.
Examples:
| Endpoint name | Environment variables |
|---|---|
web | FB_ENDPOINT_PATH_WEBFB_ENDPOINT_URL_WEB |
jupyter-one | FB_ENDPOINT_PATH_JUPYTER_ONEFB_ENDPOINT_URL_JUPYTER_ONE |
my.endpoint | FB_ENDPOINT_PATH_MY_ENDPOINTFB_ENDPOINT_URL_MY_ENDPOINT |
api_v2 | FB_ENDPOINT_PATH_API_V2FB_ENDPOINT_URL_API_V2 |
Simple alphanumeric endpoint names (containing only letters, digits, and underscores) are unaffected by this transformation beyond being converted to uppercase.
You can reference these environment variables in your service’s startup script to configure applications dynamically:
services:
web-app:
image:
uri: docker://myapp:latest
script: |
#!/bin/bash
echo "Service available at: $FB_ENDPOINT_URL_WEB"
./myapp --public-url="$FB_ENDPOINT_URL_WEB"
network:
ports:
- name: http
port: 8080
protocol: tcp
endpoints:
- name: web
port-name: http
protocol: https
type: path
scope: user
For an endpoint named jupyter-notebook, you would reference the sanitized variable names:
echo "Notebook URL: $FB_ENDPOINT_URL_JUPYTER_NOTEBOOK"
The variables above describe a service’s own endpoints. Every job and service in a workflow also gets the endpoints of all the workflow’s services, so a job can reach a sibling service through the same path the endpoint proxy uses:
FB_SERVICE_<SERVICE_NAME>_ENDPOINT_ID_<ENDPOINT_NAME>— the endpoint IDFB_SERVICE_<SERVICE_NAME>_ENDPOINT_PATH_<ENDPOINT_NAME>— the URL path,/endpoints/<endpoint-id>for apathendpoint and/for asubdomainoneFB_SERVICE_<SERVICE_NAME>_ENDPOINT_HOST_<ENDPOINT_NAME>— the hostFB_SERVICE_<SERVICE_NAME>_ENDPOINT_URL_<ENDPOINT_NAME>— the complete URL
<SERVICE_NAME> and <ENDPOINT_NAME> are sanitized as described above. For an endpoint attached
to a persistent endpoint, the ID is the persistent endpoint’s. The
per-replica endpoints of an autoscaled service are not included. If a job or service sets one of
these variables in its own env, its value is kept and the injected one is dropped.
A service that serves under its endpoint path, such as Jupyter started with
--NotebookApp.base_url=$FB_ENDPOINT_PATH, answers under that path on its internal address too, so
a job in the same workflow uses the path when calling it directly:
jobs:
check-notebook:
image:
uri: docker://alpine:latest
command:
- /bin/sh
- -c
- wget -qO- "http://jupyter.service:8888${FB_SERVICE_JUPYTER_ENDPOINT_PATH_JUPYTER}/lab"