@serve.zone/coreflow
Coreflow is the serve.zone cluster relay: the single session a cluster keeps open to Cloudly. A serve.zone cluster is outbound-only — nothing inside it is dialled from the outside — so every Cloudly-owned operation against cluster-side components travels through this one process.
Coreflow was the Docker Swarm reconciler until 32.0.0. That reconciler is gone: Pallet owns workload execution and Cloudly owns the decisions.
Issue Reporting and Security
For reporting bugs, issues, or security vulnerabilities, please visit community.foss.global/. This is the central community hub for all issue reporting. Developers who sign and comply with our contribution agreement and go through identification can also get a code.foss.global/ account to submit Pull Requests directly.
Runtime position
Cloudly control plane
<- one outbound TypedSocket session, authenticated with the cluster machine credential
Coreflow relay (this package)
-> Corestore control API (platform bindings, credential publication, legacy inventory,
backup, restore, archive)
<- Pallet nodes and the cluster ingress, on the relay's one TLS listener
The relay opens no other listener. Up to 32.x it also served an internal TypedSocket server on
port 3000 for Coretraffic; that server, COREFLOW_INTERNAL_HOSTNAME, COREFLOW_INTERNAL_PORT and
coreflowGetTrafficStats were removed in 33.0.0. The cluster ingress registers on the relay's TLS
listener, and traffic statistics travel as getClusterTrafficStatistics
(see Cluster ingress).
The relay holds the only Cloudly credential in the cluster: the cluster's own
relay credential, whose bearer it reads once at start from
SERVEZONE_CLUSTER_RELAY_AUTHORIZATION and keeps in memory. It authenticates by
exchanging that bearer for a cluster identity; @serve.zone/api then registers
that session server-side (registerCloudlyClientSession), which is what makes
the relay dispatchable. The relay authors no connection tags of its own —
Cloudly's TypedSocket server admits exactly one client tag, and it is not one of
Coreflow's.
Cloudly states on every accepted registerClusterRelay which generation of that
credential it authenticated the session as. The bearer itself is never logged,
never sent in a registration body and never written to disk; the generation is
what an operator reads this relay's authority by. A rotation has no overlap
window, so the bearer of the current generation reaches a relay only by being
placed in its environment and the process restarted — the relay never asks
Cloudly for one, and no request returns one. A relay that learns the generation
it holds is gone therefore parks instead of retrying, and says so once.
Protocol version
The protocol version is the installed @serve.zone/interfaces release. One
release of those interfaces is exactly one contract, so there is no protocol
number, version suffix or schemaVersion beside it anywhere in this package.
A relay stands on two hops and negotiates on both
(ts/coreflow.protocol.ts, ts/coreflow.classes.relaybound.ts):
- Cloudly hop. The relay is the client. It offers the release it was built
against on both carriers of that one session —
registerCloudlyClientSession, which@serve.zone/apiowns, andregisterClusterRelaybeside it — and reads back the offer Cloudly answered with. The offer is taken from@serve.zone/api'scloudlyClientProtocolOffer, so one session can never name two contracts. - Cluster hop. The relay is the server, with one offer per session kind it
serves:
palletRuntimefor the nodes andclusterIngressfor the ingress, each with its own oldest accepted peer release. It judges a peer's offer before the body is validated, before a bearer travels anywhere and before anything is forwarded — the contract validators are exact-key, so a peer of another major would otherwise be answered as malformed instead of being told which major it met. A refused peer is never carried to Cloudly.palletRuntimerequires 32.10.1 because that release binds every physical runtime session to the cluster Cloudly read from its owned node record;clusterIngressremains compatible with 32.0.0.
A refusal keeps the relay standing and makes it not ready. Both sides of a
refusal are a build somebody has to upgrade, so nothing is torn down: the
Cloudly session stays open, the carried nodes stay carried, the held DNS windows
keep being released, and the route table the ingress holds keeps being served.
What changes is what the relay says about itself — Coreflow.checkReadiness()
answers { ready: false, refusals } naming each hop that cannot serve and the
line that states why, until an accepted offer replaces it.
One statement is not a refusal to wait out: a superseded credential
generation. Everything else a relay hears about itself ends when somebody
changes the other side, and it keeps offering at the contract's cadence until
that happens. A rotation of this cluster's relay credential has no overlap
window, so the bearer this process holds authenticates nothing again and no
request returns another one: the relay parks. It stops offering
registrations, stops re-authenticating its session, says once which generation
it is bound to and that only a restart holding the bearer of the current one
serves the cluster again, and answers ready: false with that same line for the
rest of the process. It does not exit — a relay that crash-looped against an
environment only an operator changes would take its cluster's nodes with it.
Two independent things can stand on one hop, and an entry names which it is by
its refusal: the two offers, when the peer's build and this one disagree about
the contract, and null, when the peer's offer was one this relay serves and
only its answer was not the contract's — no version is what an operator changes
then, and an answer this build can act on is what ends it. They are held apart
because they end apart, and a hop serves again only once neither stands.
The refusal is retained and offered again every
protocol.refusedOfferRetryIntervalMs (5 minutes), the interval the contract
states rather than one this relay invented, so a Cloudly upgraded under a
refused relay is rejoined without a restart. On the cluster hop the peers
re-offer at their own cadence; the relay only has to notice, which it does the
moment an offer it can serve arrives.
What the relay does today
- Cloudly session (
ts/coreflow.connector.cloudlyconnector.ts): connects, authenticates with the cluster machine credential, renews the identity before it expires and re-authenticates after a reconnect with bounded backoff. Every forwarded node session travels on it, and the cluster config and certificates the relay needs are read over it. - Relay registration (
ts/coreflow.classes.relayregistration.ts): on every registration of its own Cloudly session, a relay that serves a listener tells Cloudly which build it is, which node it runs on and the endpoint it bound. See Registering this relay. - Runtime-session forwarding (
ts/coreflow.classes.sessionforwarder.ts): a Pallet node registers over the cluster-side listener and the relay carries that registration to Cloudly on its own session, byte for byte; Cloudly's controller pushes come back over the same session and are fanned out to the node theirsession.nodeIdnames. - DNS lease renewal escrow (
ts/coreflow.classes.leaseescrow.ts): holds the DNS lease renewals Cloudly signed ahead for each carried node, in memory, and hands a node the one whose window is current only while the relay has no Cloudly session. See DNS lease renewal escrow. - Corestore (
ts/coreflow.classes.corestoreclient.ts): the single client for the node-local control APIs of this cluster, with a bounded timeout per call and a named error carrying the method, path and status. Every call names the node it acts on, dials the endpoint that node published and authorizes itself with the credential Cloudly sealed to this relay. See Corestore through the relay. - Isolated restore (
ts/coreflow.classes.isolatedrestore.ts): carries Cloudly's five node-scoped isolated-restore controls to the node each one names, under the grant Cloudly signed for that one operation, and answers what that node stated about the restore. See Isolated restore. - Platform bindings (
ts/coreflow.classes.platformmanager.ts,ts/coreflow.classes.corestorepublication.ts): reconcilesdatabaseandobjectstoragebindings on the node Cloudly placed them on, under the publication grant Cloudly issued for each, and publishes each binding's credential material to Cloudly sealed to Cloudly's ingress recipient. Cloudly stores it as the service's Secrets and marks the binding ready; the relay then reports the binding's endpoint, which carries no credential. A binding without a grant is left untouched and unreported. A binding Corestore holds back for a legacy resource is reported with its named reason.pushnotificationbindings stay untouched; Cloudly owns them. Cloudly answers the desired state only to the relay whose registration it accepted, so nothing is reconciled while the relay starts: every acceptance reconciles against the cluster config read right after it, and every config push reconciles while that acceptance stands. A lost Cloudly session, or a registration refused by a name that states Cloudly does not dispatch to this relay, pauses reconciliation until the next acceptance; a refused read is logged in Cloudly's words and the relay keeps running. A cluster without a relay block registers nothing and so reconciles no bindings. - Image pull endpoint (
ts/coreflow.classes.registryproxy.ts): forwards the cluster's/v2/pulls to the registry Cloudly publishes, with the node's own credential, so a node fetches its images inside its cluster instead of dialling the control plane. See Image pulls through the relay. - Managed VPN hub (
ts/coreflow.classes.clustervpnhub.ts): runs this cluster's own smartvpn managed hub, so an inter-node packet crosses the cluster's network instead of the WAN twice. Cloudly keeps every authority — it compiles the network, issues the node credentials and owns the revisions — and pushes the complete network to the relay, which enforces it. See The cluster's managed VPN hub. - Cluster ingress (
ts/coreflow.classes.clusteringress.ts): holds the route table Cloudly computes for this cluster, hints the ingress that registered on the relay's own TLS listener — the one the cluster's Pallet nodes dial — at every newer revision, and hands it the table when it asks. Traffic statistics travel Cloudly → relay → that bound ingress, and are refused by name while none is bound. See Cluster ingress.
Pending relay functions
These are declared here because Cloudly may already ask for them. Wherever Cloudly can dispatch for one of them, the relay answers with a named refusal — never with silence or an empty result:
| Function | Status |
|---|---|
| Relay certificate renewal | not built; @push.rocks/smartserve fixes TLS material at construction and exposes no live rotation, and a restart-based renewal would drop every node socket and every in-flight request. It needs a smartserve API first |
| Desired-state cache and idempotent replay on reconnect | not built |
| Signing new network state during a Cloudly outage | not built, and not the relay's: DNS continuity comes from renewals Cloudly signs ahead and the relay only holds (see DNS lease renewal escrow). New workloads, membership and published ports still need Cloudly's compiler |
| Served-node roster and readiness publication | not built; the relay registers itself (registerClusterRelay) but that contract carries only its build, its node and its endpoint. Cloudly's client-tag policy admits only its UI-live tag, so publishing which nodes a relay serves needs a contract with a place for them |
Every method the Swarm reconciler used to answer is registered as a named
refusal (ts/coreflow.classes.retiredhandlers.ts), so a Cloudly dispatch reports
a retired capability and its new owner instead of timing out.
Cluster-side listener
ts/coreflow.classes.relaylistener.ts is the server Pallet nodes dial instead
of Cloudly, and the one server this relay runs: a SmartServe with its own
TypedSocket server and its own router, apart from the Cloudly-facing router of
the relay's outbound session, so the node-facing surface and the Cloudly-facing
one can never answer each other's methods. It speaks exactly what a Pallet
controller client sends — one TLS WebSocket upgrade at / — and it binds a
host address and port of its own, which are daemon-local facts and never
cluster state. Its HTTP surface beside that upgrade is the cluster's image pull
endpoint and nothing else (see
Image pulls through the relay); the two share
one port because a node dials one origin, and one certificate, for both.
All six methods a node sends its controller
(registerPalletRuntimeSession, getRuntimeManagedVpnCredential,
getRuntimeRegistryCredential, reportRuntimeNetworkProtection,
reportRuntimeAssignmentObservation,
reportRuntimeAssignmentTerminalReceipt) are forwarded to Cloudly on this
relay's own session, with their bodies untouched, and Cloudly's answers come
back unchanged. The relay is a transport, not a party: the bearer, the node
instance nonce, the credential generation and the session reference are
Cloudly's to read, and Cloudly recognises a relayed registration from the
identity tag it set on this relay's session, so nothing has to be added to one.
Image pulls through the relay
A cluster that is reached only through its relay must also fetch its images
through it. ts/coreflow.classes.registryproxy.ts is the endpoint a node's
container runtime dials for that: it is the same origin and the same
certificate the nodes already dial — the WebSocket upgrade at / and the pull
under /v2/ share one listener — and it forwards the exchange to the registry
Cloudly publishes, over a second outbound HTTPS connection of the relay's own.
Cloudly tells a node to use it by stating the relay's origin as the assignment's
workload.pullEndpoint; the digest-pinned reference is unchanged either way, so
the endpoint decides reachability and never content.
- The credential is the node's. Cloudly issues it over the forwarded
getRuntimeRegistryCredentialand authorizes the pull itself, exactly as it does a direct one. This relay holds no registry credential, decides nothing about who may pull what, and caches nothing: a cache would be a second copy of content Cloudly is the authority for. - One header is rewritten, and only its origin. The registry's
WWW-Authenticaterealm is an absolute URL naming Cloudly, which a node would otherwise follow straight out of its cluster, so the realm's origin becomes this relay's.service, which is the audience the issued token is checked against, is left exactly as the registry wrote it — rewriting it would invalidate every token this relay carried. A realm naming a third party is not rewritten: that challenge is that party's to answer. The relay states such a realm in its log once per origin, because the same challenge travels unchanged whenCLOUDLY_URLnames the same Cloudly under a different name than the public origin the realm is composed from — and then the cluster's nodes leave it for their credential with no other symptom. - Pulls only.
GETandHEADunder/v2/are forwarded; every other method is refused by name, because a registry write is Cloudly's own authorization path and no node performs one. Every path that is not/v2/is refused by name too: this is the listener's whole HTTP surface, so a node that reached it hears what it serves rather than a bare 404. - Nothing is held. The answer's body is streamed to the node as it arrives,
so a layer of any size crosses this relay without being collected in its
memory, and the listener does not re-encode it. The relay asks the registry for
identityand asks explicitly, because an omittedaccept-encodingis not an absent one — so the node reads the registry's own bytes under the length that describes them, and an answer the registry encoded anyway is passed on with no length at all rather than the encoded one. The forwarded request headers and the returned answer headers are both bounded, and only the headers a pull needs travel in either direction — no cookie, noHostof the node's, no transport header. - The registry is
CLOUDLY_URL. Cloudly composes both the realm and the token audience from its own public origin, which is the origin this relay already dials for its session, so the pull endpoint needs no variable of its own and cannot be pointed at a second registry by accident. - A failure is named. An unreachable registry is answered
502with the failure's class, never its message: that exchange carries the node's pull token.
The cluster's managed VPN hub
The relay is the one process every node of its cluster already dials, so it is
also the one that relays their packets. ts/coreflow.classes.clustervpnhub.ts
runs a @push.rocks/smartvpn server in forwardingMode: 'managed' — an
exclusive IPv4 node relay that creates no TUN, NAT, bridge or host route, so the
relay needs no network privilege for it. An inter-node packet crosses the
cluster's own network once instead of the WAN twice, and the plaintext a managed
hub sees by construction stays inside the cluster it belongs to.
- Cloudly keeps every authority. It compiles the network, issues each node's credential and owns the revisions. The relay receives public keys only, inside the network Cloudly pushes; no node private key ever reaches it.
- A node dials QUIC, and only QUIC. A managed hub terminates no TLS on its
WebSocket listener —
tlsCert/tlsKeyare read by the QUIC listener alone — so awss://hub cannot exist and the relay advertises a single UDP port. - The daemon binds lazily, once. A managed hub fixes its control subnet when
it binds (
IVpnServerConfig.subnet) and smartvpn refuses to configure a running daemon, so the relay starts it with the control prefix of the first network it admits. Nothing is lost by waiting: no node can dial the hub before Cloudly has composed its endpoint from the advertisement and compiled the network that names that node. applyClusterVpnNetworkis the whole authority, never a patch. A member or a grant the push leaves out is withdrawn. The relay admits it with the contract's own decision: a repeat of the held network is a replay that changes nothing, a later revision is taken over, and a network of another authority, one whose control prefix moved, or a revision this hub has passed is refused asvpn-network-authority-foreignorvpn-network-revision-stale— with nothing applied and the network it serves untouched.getClusterVpnNetworkis pulled on every accepted registration. A daemon restart starts empty and the relay's daemon lifetime is invisible to Cloudly, so the relay asks for its cluster's current network rather than waiting for the next push.nullmeans the cluster has none yet, and the relay binds nothing.- Nothing is persisted. The held network lives exactly as long as the daemon does. A relay restart re-keys the hub and drops its cluster's tunnels, which is a data plane and not an admission gate: workloads keep running and each node re-fetches its credential on its next reconcile. A daemon that exits on its own leaves the relay holding nothing, and the next push binds a fresh one.
- The daemon's own words stay at the boundary. smartvpn's
schemaVersion, its ISO-8601 expiry and itsenabledflag are composed when a snapshot is handed to the daemon and are never part of what the relay holds; the contract's Unix-millisecond expiry and presence-only membership are.
Registering this relay
Cloudly has to know which process is the relay of a cluster and where it
listens: it publishes that listen address as the A record of the relay's
certificate domain, and when two peers claim the role a registration on a live
verified socket supersedes the stored one. So on every registration of its own
Cloudly session — a reconnect and an identity renewal alike — a relay that
serves a listener sends registerClusterRelay with its build, the node it
runs on (COREFLOW_NODE_ID) and the endpoint it actually bound
(ts/coreflow.classes.relayregistration.ts). A cluster without a relay block
has no listener and registers nothing: there would be no endpoint to name.
- It carries this build's protocol offer.
relay.versionis the build, which is a different fact from the contract it speaks; Cloudly negotiates the offer before it reads anything else on the body, so a refused relay consumes no sequence. A refusal on this carrier — sent by Cloudly, or answered as an acceptance in a contract this build cannot serve — leaves the relay serving its nodes, marks the hop not ready and is offered again every five minutes. - Nothing in it is an authority. Which cluster this relay speaks for is the
verified machine identity on its socket, never a field; whether the node is
one of that cluster's is Cloudly's to decide against its own records; and
registrationSequencecounts registrations within this process and restarts at 0 when the process does, so it orders this peer's registrations and nothing between peers. One registration is in flight at a time, because two racing each other would have Cloudly refuse the older one as not advancing. - The endpoint is reported, not derived. The relay sends the address and port it bound, verbatim. A bind Cloudly cannot publish — a loopback address, or a hostname where the contract admits an IPv4 literal — is refused by name, and that refusal is the operator's signal to fix the bind. The relay never substitutes an address it thinks would work better.
- A relay that binds a hub advertises it here.
relay.vpnHubcarries the hub's Noise public key, the address it bound and the QUIC port it bound, reported exactly likelistenAddressandlistenPort. Cloudly composes the cluster's platform endpoint from it; a relay names no endpoint id of its own and claims no TLS identity. A relay that binds no hub sends no such key, and Cloudly statesrelay-vpn-hub-absentfor that cluster instead of selecting one. See The cluster's managed VPN hub. - A refusal changes nothing about what it serves.
relay-cluster-not-enabled,relay-node-foreign,relay-node-unknown,relay-identity-unverified,relay-sequence-not-advancingandrelay-credential-generation-staleare logged by name, and the nodes already connected keep being served: dropping them would turn a Cloudly opinion into a cluster outage. A transport failure is logged by class and code, like every other failure of a call this relay makes. relay-node-unknownmeans the node is not enrolled yet. Cloudly holds no node under this relay'sCOREFLOW_NODE_ID;relay-node-foreignis a node Cloudly holds for another cluster. Enrollment comes first: runspark enrollnodeon the node before the relay registers. The relay logs this step with the refusal and registers again on its next session renewal or reconnect.- The accepted origin is checked against the served one. Cloudly answers with the origin it published for the cluster, and this relay is the only party that sees both that origin and the one it terminates. A difference is a warn naming both, because the cluster's nodes would be dialling an origin this process does not serve. An accepted registration always carries an origin: a cluster nobody enabled a relay for is refused by name instead.
- The acceptance binds this process to one generation. Beside the origin and
the offer, an acceptance names the relay credential generation Cloudly
authenticated this socket as and the accepted registration record it just
wrote. Both are read with the contract's own exact-key reference reader, never
a shape check of this relay's own, and the relay states no generation of its
own on a registration — naming one would be asserting its own rank. The first
acceptance of a process pins the credential and logs it once
(
this relay is bound to cluster relay credential <id> generation <n>, public metadata: the bearer behind it is held by the session and never logged), and the accepted registration record is ordered: every later acceptance must name exactly the pinned credential and a strictly higher generation of that same record. - An answer that does not hold is not an acceptance. A Cloudly that states
neither reference, one that names another credential generation, and one whose
record does not advance all take the same branch as an unservable contract:
nothing is reported as accepted, the hop says it is not ready, the nodes this
relay already carries keep being served, and the offer is repeated at the
contract's cadence until an answer this build can act on arrives. The
credential case is the exception that also ends custody — rotation has no
overlap window, so a second generation on this socket means the cluster's
relay is another process: this relay states
relay-credential-generation-stalefor itself, drops every DNS lease renewal run it held, and parks. - A superseded generation parks this relay, and nothing else does. Both
carriers of the fact reach the same terminal state: Cloudly refusing a
registration with
relay-credential-generation-stale, and an acceptance that names a credential generation other than the pinned one. In that order, the relay gives up custody of what Cloudly dispatches, stops announcing its endpoint — no registration is offered again, at the cadence or on a reconnect — stops scheduling the identity refreshes and registration retries that would re-exchange a bearer Cloudly no longer accepts, and states once, as the one thing that blocks thecluster relay registrationhop, which generation it is bound to and that only a restart holding the bearer of the current generation serves this cluster again. Everything it already carries keeps running: the Cloudly socket stays open, the node sessions stay carried, the route table keeps being served and the process keeps running. There is nothing left for it to do, because no request returns a bearer — the current one reaches this relay only through its environment, and only a restart reads that again.
Forwarding a runtime session
ts/coreflow.classes.sessionforwarder.ts carries both directions and caches
nothing: a relay that answered from memory would be an authority, and Cloudly's
session store is the only one.
- Upstream. A registration waits up to five seconds for a Cloudly session that is momentarily reconnecting, then forwards within a nine-second budget — under the ten seconds a Pallet client gives its connection-restore callback, so a node hears a named refusal from the relay instead of running out its own clock. Every other forward is bounded at eight seconds, under the ten Cloudly allows itself. A refusal while the session is down is named and never silent, and it states the wait the relay granted — "no Cloudly session after waiting 5000 ms" — so a node operator reads how long the registration was held open instead of guessing. Each later session request must arrive on the exact node socket that registered the Cloudly-issued binding the request carries. A different socket, superseded session, cluster, or older relay registration generation is refused before forwarding. The cluster is part of Cloudly's issued binding; the relay never adds it to a registration or derives it from local configuration.
- Downstream.
applyRuntimeAssignment,applyRuntimeNetworkSigningAuthority,applyRuntimeNetworkProjectionandapplyRuntimeNetworkDnsLeaseRenewalare answered for as long as the listener runs and detached from the Cloudly router when it stops. Each is routed byrequest.session.nodeIdto the node's own socket, bounded at eight seconds, and the node's acknowledgement is returned verbatim. The relay never synthesises one. - When a node is gone. Its socket's close drops its registry entry, and a
later push for it is refused by name (
node <id> is not connected to this relay). There is no per-node retirement call to make: the runtime-session contract has none, and Cloudly fences the session itself. The relay invents nothing to fill that gap. - Secrets. The managed VPN credential, registry credential, recipient
possession challenge and assignment secret material carry sensitive bytes.
The relay forwards recipient-state, begin-enrollment, complete-enrollment and
material requests through the same current-session path. No secret-bearing
field or response body is read, stored or logged in either direction; every
forwarded request sets
skipHooks, and the relay registers no typedrequest hook and no raw-frame subscription. Its log lines name a method and a node, never a body. The value-free facts read on the way are the binding a registration answers with, each request's session binding for the current socket check, and the{ id, generation, digest }of a projection a node acknowledged, bound to the ACK withbindApplyRuntimeNetworkAuthorityResponse. The second one retires DNS lease renewals held for a projection the node has replaced.
A node still registers nowhere while this relay has no Cloudly session, and a
refused registration is not a retry: a Pallet controller client turns any
rejection from its connection-restore callback into a terminal
TypedSocketConnectionRestoreDeniedError, which fails that node's controller
lifetime until its process restarts. That is why the relay waits out a brief
reconnect before refusing.
Cluster ingress
A cluster's ingress terminates TLS for the hostnames its services publish and forwards to the host
ports Pallet published behind them. Cloudly computes that table — it is the only party that knows
every service, every ready endpoint, the node address each published port answers on and the
certificate each hostname needs — and this relay holds it for its cluster and hands it to the
ingress that asks (ts/coreflow.classes.clusteringress.ts).
- The table is memory only. It carries the private key of every hostname the ingress terminates, so it is never written, never journaled and never serialised into a log line or an error. What the relay says about it is its revision and how much it holds.
revisionis the only idempotency. A pushed table whose revision is not greater than the held one is ignored and the answer states the revision the relay actually holds, so a repeated push, a replayed push and a push that crossed a newer one all settle the same way. A table that failsvalidateClusterRouteTableis refused with the validator's first reason — which, by that contract, never quotes the material it refused — and the held table is untouched.- The relay re-reads on every registration Cloudly accepts, not only on a changed one: Cloudly
may have recomputed while the relay was away, and a table is idempotent by revision. Not on the
registration of its own Cloudly session, which comes first: Cloudly answers
getClusterRouteTableonly to the relay whose registration it accepted and refuses an earlier readrelay-not-connected. A failed read keeps the last good table, because an ingress serving yesterday's routes is a cluster that still answers. - No table is served as no table.
nullmeans Cloudly has computed none yet; the relay says so once per transition and answersnull, never an empty table. - A newer table is hinted, never delivered. Whenever the relay adopts a revision greater than
the one it held, it sends the bound ingress
pushClusterIngressRouteTableChanged { revision }. The hint carries the revision alone — no route, certificate or key — and the ingress answers it by pullinggetClusterIngressRouteTablewith the revision it applied. A hint that fails is logged and never retried: the table's revision is the idempotency, so the next pull catches up. A table adopted while no ingress is bound is hinted to nobody; the registration answers its revision.
The ingress registers on the same TLS listener the cluster's Pallet nodes dial, so isolation is
by authorization, not by routing: registerClusterIngress is forwarded to Cloudly with its
bearer untouched, only a registration Cloudly accepted binds that connection, a later accepted one
supersedes it, and the close hook releases only the connection currently bound.
getClusterIngressRouteTable is answered for that connection alone — any other caller, a node
included, is refused by name. getClusterTrafficStatistics travels the other way: Cloudly asks
over the relay's own session and the relay carries the question to the bound ingress and the page
back, or refuses by name when no ingress is bound. A named registration refusal from Cloudly
(cluster-ingress-registration-refused with its reason and retryAfterMs) travels back to the
ingress untouched. The relay refuses in the same shape by itself: not-ready while it has no
Cloudly session or cannot carry the registration to Cloudly, and registration-invalid for a body
validateClusterIngressRegistration refuses, both with retryAfterMs 5000.
Corestore through the relay
Corestore is node-local — a workload's database and bucket live on the node the workload runs on — and one relay serves a whole cluster. So every Corestore call this relay makes names the node it acts on, dials the endpoint that node published, and authorizes itself with a credential Cloudly sealed to this process. None of the three used to be true: the node was whichever coreflow had tagged itself with a hostname, the endpoint was a Swarm alias, and the credential was an environment variable.
- The endpoint comes from the cluster document
(
ts/coreflow.classes.corestoreendpoints.ts).IClusterNode.data.corestore.endpointis Cloudly's to publish and this relay's to read, adopted from the cluster config at boot, after every accepted registration and from every config push. A node without one is refused by name (CORESTORE_ENDPOINT_UNKNOWN) rather than dialled at a guessed hostname, and an endpoint Cloudly removed stops being dialled. The relay dials each node's Corestore directly: outbound-only constrains what leaves the cluster, not what moves inside it. - The credential comes from Cloudly, sealed
(
ts/coreflow.classes.corestorecredential.ts). The first Corestore-bound call generates an X25519 key pair, enrolls it as this cluster's secret recipient by proving possession of the private key, then readsgetCorestoreControlCredentialMaterialfor that capability and opens it. Enrolling is deliberately not part of starting: a cluster that runs no Corestore never needs a credential, and a boot that depended on one would fail for a subsystem it does not use. The plaintext token lives in one private field: never written, never journaled, never logged, never attached to an error. The private key never leaves the process — a restart enrolls a fresh one under a fresh recipient key id (relay-corestore-control:<uuid>), which Cloudly answers by retiring the previous recipient, so nothing is persisted. Key ids are unique across every recipient Cloudly retains for the cluster, so an id is never reused, and Cloudly keeps the set bounded by dropping the oldest revoked recipients. A failed enrollment isCorestoreCredentialEnrollmentError, logged with Cloudly's refusal code and named by every Corestore call it stops. A rejected call (401 or 403) drops the held token and reads it again, and a changed active recipient re-enrolls on the next registration of the relay's own session. - A binding is reconciled only under its grant
(
ts/coreflow.classes.platformmanager.ts). Cloudly publishes a Corestore publication grant per binding on the cluster config; the relay builds the exact value-free binding request, proves its digest is the one the grant covers, and reconciles through/control/credential-bindings/*. A binding without a grant is left exactly as it is and nothing is reported for it: Cloudly has either not issued the authority yet or already consumed it by accepting the binding's material, and in the second case the binding is ready by Cloudly's own word. A grant that names another service or capability than its binding is refused asbinding-not-granted. The generic provisioning route carries no authority, and a relay that used it would be creating a tenant's database on its own say-so. Object-storage retention evidence is verified by the shared validator with the grant's authority, so this relay holds no second opinion about an immutability promise. What the cluster no longer desires is retained, never deleted: interfaces defines publication grants and no deletion grant. - A failed binding is reconciled again with bounded backoff and reported once.
A 4xx answer to a binding's reconcile or material read — other than a bearer
rejection (401, 403) or a "come back later" status (408, 425, 429) — is a
Corestore refusal: the same request under the same grant gets the same answer.
Any other failure (a 5xx, an unreachable node, a publication Cloudly refused)
may pass on its own. Under the same grant the relay waits 5 s before trying a
failed binding again, doubling to at most 5 minutes after a refusal and 60 s
after any other failure, and wakes itself when the wait ends. It reports
failedand logs one named line once per distinct failure — a new grant, kind or class (Corestore's status and bounded reason, or the failure's name and code) — not once per attempt:corestore-binding-refused:for a refusal,corestore-binding-failed:for anything else, with the node, binding, grant operation and generation and no request value. A new grant is tried at once. When a binding that failed reconciles, the relay logscorestore-binding-failure-cleared:. - A database binding is addressed to its node's database endpoint. The
database request states the
corestoreDatabaseCloudly publishes on the node the binding is placed on (IPlatformDatabaseEndpoint { host, port }), so Corestore addresses the credential to where a workload there reaches the database instead of to its own public host. The relay checks the answered material withvalidateCorestoreDatabaseCredentialMaterialEndpointbefore it publishes anything, and refuses material addressed elsewhere. A node without a validcorestoreDatabasegets no request at all: the binding is reportedfailedwithdatabase-endpoint-missinguntil Cloudly states one, and a moved endpoint arrives as a new grant for the moved request. Database bindings therefore need peers that speak the endpoint: a Cloudly that publishescorestoreDatabaseand grants digests that include the endpoint (33.1.0 or later) and a Corestore that accepts it (33.0.0 requires it). Upgrade Cloudly first, then this relay, then Corestore; with Cloudly 33.0.0 every newly granted database binding is reportedfailedwithdatabase-endpoint-missing. - The credential reaches the workload through Cloudly, sealed
(
ts/coreflow.classes.corestorepublication.ts). After Corestore reconciled a binding and answered its material, the relay reads Cloudly's active ingress recipient (getSecretIngressRecipient), seals the material to it under the exact contextcreatePublishCorestoreCredentialMaterialEnvelopeContextstates, and sendspublishCorestoreCredentialMaterial. Cloudly stores the values as the service's Secrets, marks the binding ready and consumes the grant in one transaction; the relay verifies the receipt withvalidateCorestoreCredentialPublicationReceiptagainst the cluster its config names and the service's organization. Nothing is reported ready before that receipt. The mutation id is derived from the grant's operation, so every attempt under one grant is one mutation. A publication whose answer was lost is sent again as the identical sealed request, because Cloudly binds a receipt to the envelope digest; the relay therefore holds the sealed request (ciphertext only) per grant operation, and the verified receipt once it arrives, until the grant leaves the config push. A grant already published answers its kept receipt and seals nothing again. What the relay reports follows four rules:- A named refusal other than
INTERNAL_ERRORandREPLAY_CONFLICTdrops the held request and is reportedfailed; the next attempt, after the failed binding's wait, seals afresh, to a rotated recipient if Cloudly rotated one. - While a request sealed for the binding's current grant is held, answered
or not, no failure is reported
failed. That covers an open outcome — a lost answer, an answer without a named code, orINTERNAL_ERROR, which Cloudly also answers after a commit whose config push failed — and any later failure under that grant, a Corestore reconcile that fails before the held request goes out again included. It coversREPLAY_CONFLICTtoo, which proves Cloudly committed another envelope under the grant's mutation: the relay lost the request it held, as after a restart while the grant is still on the push, so it keeps the grant marked as committed and seals and sends nothing again under it. Cloudly may already have marked the binding ready and consumed the grant, and once the grant leaves the config push nothing would correct the report, so the relay logs a warning — naming the grant operation — on each attempt and the binding stays as Cloudly states it. - A receipt the shared validator refuses is never reported
failed: Cloudly committed the envelope, so the binding's readiness is its word. It is logged at error level with the grant operation, every reason the validator gave, and each field of the relay's own scope the receipt misstates (mutationId,organizationId,clusterId), expected against received; nothing of the material is logged. The request stays held without a receipt and goes out again unchanged, never sealed a second time. - A failure before any request is held for the grant — Corestore failing the
reconcile or answering unusable material, or a service without a cluster or
organization to bind the publication to — is reported
failed.
- A named refusal other than
- A binding endpoint carries no credential. The relay reports a database
binding as
{ name, capability, protocol: 'mongodb', networkAlias, port }— the host and port of the endpoint its request stated — and an object-storage binding as{ name, capability, protocol: 's3', networkAlias }with the node's Corestore host. NointernalUrlis reported: the database URI Corestore answers embeds the tenant's username and password, a binding is a Cloudly document an administrator reads, and a workload reads its connection from its Secrets. - A legacy resource is a named state, not a failure. Corestore refuses a
credential binding beside a service's legacy
/resources/provisionresource with a 409 whose body carries the reason. The relay reads it withreadCorestoreLegacyAdoptionRefusaland reports{ status: getPlatformBindingStatusOfReason(reason), statusReason: reason }:legacy-adoption-pendingasprovisioning,legacy-adoption-conflictandlegacy-adoption-mismatchasfailed. Every other refusal stays an ordinary failure without a reason, and a later report without one clears the stored reason. A report Cloudly already holds — status, reason and endpoints — is not sent again. - Placement comes from Cloudly through
getClusterPlatformDesiredState, which answers{ nodeId, binding }per binding. A binding without a node is a database nobody can place once one relay serves the whole cluster. Cloudly answers it only on the socket its relay election names — the one whoseregisterClusterRelayit accepted, and the only one of the cluster — and refuses any otherrelay-not-connectedorrelay-ambiguous, so the relay asks from an accepted registration on. - The inventory is an answer, not a guess
(
ts/coreflow.classes.corestoreinventory.ts).getClusterCorestoreInventory { nodeId }answersreachable: truewith an empty service list for a node whose Corestore answered and holds nothing — the proof a greenfield cluster is blank — andreachable: falsewithCORESTORE_ENDPOINT_UNKNOWNorCORESTORE_UNREACHABLEfor a node that could not be read, which proves nothing. The two are never the same answer. - The legacy inventory travels exactly as Corestore stated it.
getClusterCorestoreLegacyResources { nodeId }reads the node'sGET /control/legacy-resourceswith the control credential and answers{ nodeId, checkedAt, reachable: true, services }after bothvalidateCorestoreLegacyResourceInventoryandvalidateClusterCorestoreLegacyResourceInventoryaccept it — Cloudly checks an adoption against a digest of this answer, so nothing is reordered or renamed. A node without a published endpoint answersreachable: falsewithCORESTORE_ENDPOINT_UNKNOWN; a Corestore that refused, went away or answered an unusable inventory answersCORESTORE_UNREACHABLE. The unreachable answer lists nothing at all. - Backup, restore and pruning name their node
(
ts/coreflow.classes.clusterbackup.ts).backupClusterService,restoreClusterServiceandpruneClusterNodeArchiveact on the node the request names, through that node's endpoint, and refuse by name when the cluster publishes no endpoint for it, when the node cannot be reached, or when Cloudly has not issued the control credential. - A database closure streams; it is never base64 and never buffered. A
payload larger than a control body moves as a Corestore closure on the relay's
own session, chunk by chunk, with its digest computed while it crosses.
backupServiceDatabaseClosuresends: the relay opens Corestore's export, states the exact size, digest and receipt in the descriptor, streams the bytes and treats only the stream's acceptance receipt as storage — the response that came back first merely permitted the send.restoreServiceDatabaseClosurereceives: the relay opens Corestore's restore from the descriptor's exact size, because Corestore requires one exactContent-Length, identity transfer encoding, no content encoding and a body whose length matches its receipt, then accepts the stream only once Corestore answered. Progress, not duration, bounds both, and a cut transfer names the side that went quiet: the node for an export it stopped streaming, Cloudly for a body it was feeding into a node that was only waiting. A busy 256 MiB transfer is never interrupted. Neither direction buffers a closure: the receiving side is pulled by the Corestore request that writes it, so a Corestore slower than Cloudly slows the transfer down instead of deadlocking it, and every way out of a transfer — including the refusals nobody planned for — releases the body and the progress timer it acquired. replication.enabledis what carries a backup out of the cluster. A backup that asks for replication ends by sending this service's database closure to Cloudly, and a restore that states it takes its database back from Cloudly's closure instead of from the node's own archive, which is exactly what a replicated backup is allowed to have lost. A service without a database has no closure to carry and says so. Volume and object-storage snapshots stay node-local in both directions: the streamed contract covers the database closure and nothing else yet, and a streamed contract for volumes and object storage is a recorded follow-up. The relay's own logs say the same — a replicated backup carries the database closure, no more than that. A failed transfer does not fail the backup: the node-local snapshots happened, they are this backup's, and a relay that threw here would leave Cloudly without the ids of snapshots it now owns while a retry snapshotted the same backup id twice. The closure is refused with its reason and logged with the backup it belongs to. Cloudly authors the replication record — the relay's response carries noreplicationfield, because anIBackupReplicationResultdescribes Cloudly's own archive target, a path, a manifest and its digest — and derives that record from the closure it accepted for the backup: the closure is always sent before the response is, so an accepted closure is a replicated backup and a requested one that never arrived is a failed replication.- A restore is ordered, because it is not atomic. Every precondition is checked before the first step runs — the node's endpoint, its Corestore credential and, for a replicated backup, the Cloudly session the closure travels on — and then the steps run in one documented order, the remote one first: the database closure from Cloudly, then each volume (created before it is restored into, because a restore lands on nodes that never ran the service), then the node-local database and object-storage snapshots. A refusal names the step that ended, and that order says what had already run when it did.
- An isolated restore is carried, not driven. The five node-scoped isolated-restore controls are forwarded to the node each one names, and the node's own progress is what comes back. See Isolated restore.
- The node-tag routed Corestore methods are gone, not refused.
executeServiceBackup,executeServiceRestore,coreflowPruneNodeArchiveandcoreflowGetCorestoreInventoryare not registered at all: they were dispatched at the coreflow that had tagged itself with a node's hostname, no production Cloudly sends them to a cluster relay, and a cluster is flipped together with the Cloudly that talks to it. This relay answers exactly what its contract serves.
Isolated restore
An isolated restore replays a backup into a scratch namespace on one node,
under grants Cloudly signs for one operation at a time. The relay is a
carrier for it (ts/coreflow.classes.isolatedrestore.ts): it serves
prepareIsolatedRestoreOnNode, stageIsolatedRestoreArchive,
getIsolatedRestoreProgressOnNode, executeIsolatedRestoreOnNode and
cleanupIsolatedRestoreOnNode, each { nodeId, control } in and the node's
IIsolatedRestoreProgress out.
- The grant is the only authority, and this relay holds none of it. The compact grant inside each control names the operation, expires within minutes and binds the cluster, the node name, the canonical resource mappings and the archive manifest. A carrier can forward one but can neither widen nor mint one, so the control is posted to Corestore exactly as Cloudly wrote it. It is a bearer: the complete control is excluded from every log, every refusal and every answer on this hop.
- The control is judged before a bearer travels. The envelope and the
control body are read with
@serve.zone/interfaces' own readers — the same code Corestore runs on them — so a body Corestore would refuse is refused here, before the node is dialled. That is also what holds a staged chunk to its bound: at most 512 KiB decoded, with achunkSha256over exactly the bytes that travel, which Corestore verifies against the immutable manifest object before it stores them. - Cloudly cuts the chunks and threads the revision; this relay does neither.
Cloudly holds the source backup, so the bytes travel Cloudly → relay →
Corestore one chunk per call. Every answer carries the node's
revision, and the next control has to send it back as itsexpectedRevision; a control that does not is refused by the node as a conflict. The relay carries that number through unchanged and never retries a control: a retried chunk is a chunk written against a revision that has since moved. - Refusals are the contract's three names.
restore-node-not-carriedwhen this cluster publishes no Corestore endpoint for that node,restore-node-unreachablewhen the node did not answer — including while this relay holds no Corestore control credential, since nothing is dialled then — andrestore-node-refusedwith Corestore's own reason behind it for everything the node itself refused: a refused grant, a refused body, a chunk digest that does not match, a staleexpectedRevision, a restore it does not hold. Each is written<name>: <reason>, so a caller matches the name. - A refused grant costs this relay no credential. Corestore answers every
refused restore grant with 403 and a bearer it does not accept with 401, so on
these five routes only a 401 is a verdict on this relay's own Corestore
control credential. A grant that expired mid-staging is refused to the caller
and leaves the node's
databasecredential — the one binding reconciliation, backups and inventory are all made under — exactly where it was. - What comes back is what a node may state. Corestore answers its progress plus what the route it served did — a received-object receipt, the mappings it restored, the staging tree it deleted. Those members are dropped, the node this relay routed on is added, and the result is read through the contract's progress reader: an answer it cannot read is refused rather than repaired, and a node stating a node of its own is refused too.
- A control the node never answered is not a control that was applied. A
timeout is the one case where that cannot be proven from this side, so the
refusal states the timeout and the caller reads
getIsolatedRestoreProgressOnNode, whosestagedObject.nextOffsetsays the byte the node already holds.
Outage continuity
Cloudly keys a node's session by the carrier peer that registered it, so when this relay's own Cloudly session reconnects, Cloudly holds no session for any node behind it — and those nodes, whose own sockets never broke, have no reason to register again. Left alone, nothing would push, nothing would re-register, and every side would look healthy while the whole cluster was orphaned.
So the relay counts the connections its Cloudly session registers on — a
renewed identity on the unbroken socket keeps the same carrier peer at Cloudly
and changes nothing, a reconnect does not. When that count rises, and only while
its session is dispatchable, the relay closes every node session it carries that
was registered under an older count: 75 ms apart by default, compressed to fit a
5 s window however many nodes a cluster has, and said out loud per node:
node <id> asked to re-register: this relay's Cloudly session changed. Each node's client reconnects within a second and registers through
its own restore callback, with its own bearer and its own node instance nonce;
Cloudly issues a fresh binding and re-pushes what that node is owed. Workloads
keep running through the blip: a Pallet node keeps an ACTIVE generation whose
application and uplink are unchanged.
Closing a socket is the only lever a relay has over a Pallet client, and it is used carefully. While the relay has no session, nobody is closed — a node whose registration is refused is terminally failed, so it is better left connected and idle until the session returns.
The relay retains one entry per carried node — its connection, the binding Cloudly issued, and the connection count it was registered under — plus the count it is currently carrying. No bearer, no request or response body, never the managed-VPN credential, nothing on disk. It replays nothing: idempotency belongs to Cloudly, which answers an identical registration with the historical binding, and to the node, which re-admits a re-pushed assignment by reference.
What this cannot fix on its own: the DNS window lives inside the signed projection and is capped at 15 minutes, so a replayed projection carries a lease that has already expired. The DNS lease renewal escrow extends it through an outage with renewals Cloudly signed ahead. Managed tunnels stay unsolved: Cloudly withdraws a carried node's VPN grant when its carrier leaves, and only Cloudly can issue the next credential.
The origin the nodes dial is Cloudly's to decide, and the listener proves it can serve it: if that origin names a port, it must be the port this process binds, or construction fails by name rather than listening where nobody dials.
ts/coreflow.classes.relaycertificate.ts obtains the certificate for the
origin's domain through the Cloudly session, checks that it names that exact
domain, is complete and is not expired, and validates the material with
smartserve's own TLS validator before a listener is built from it. The private
key stays in memory: it is never logged, journaled or written, and a failure
names the domain and the reason, never the material.
The relay starts this listener only when Cloudly publishes a relay block for
the cluster (ICluster.data.relay, @serve.zone/interfaces 31.2.0). Without
one the relay logs that the cluster's nodes target Cloudly directly and serves
no listener. A block that fails validateClusterRelay is a startup failure
naming the reason, never a skipped listener: Cloudly believes a cluster with a
relay block is served.
A relay block that changes or disappears while the relay runs is reported by name and not adopted. The listener keeps serving the origin it started with, because restarting it would drop every node socket and every request in flight, and a node follows a moved relay only once Spark rewrites its identity anyway. Adopting a changed block is a daemon restart, which Spark owns.
DNS lease renewal escrow
A node's private DNS stops when its signed window ends, at most 15 minutes after
it was signed, and during a Cloudly outage nothing can sign the next one.
Cloudly therefore signs a run of IRuntimeNetworkDnsLeaseRenewals ahead of time
— 15-minute windows starting every 10 minutes, up to the contract's four-hour
horizon, each naming the node's exact projection — and hands the run to the
elected relay with holdRuntimeNetworkDnsLeaseRenewals.
ts/coreflow.classes.leaseescrow.ts holds it.
Custody. Memory only, one run per node, never on disk or in a database. The
relay holds no key and verifies no signature: every renewal is Cloudly's
statement, and the node verifies it against its own pinned authority. A hold is
answered { held } and accepted only:
- from Cloudly on this relay's own session. The handler lives on the Cloudly
router alone; the node listener answers the same method with the named refusal
escrow-cloudly-only. TypedSocket attaches no transport context to a request a client socket routes, so origin is enforced by which router serves the method, not by a per-request check; - when
validateHoldRuntimeNetworkDnsLeaseRenewalsRequestpasses, includingmaximumEscrowedRenewals(32) andmaximumEscrowHorizonMs(four hours), elseescrow-run-invalid; - for a node this relay carries, else
escrow-node-not-carried. The node's session is read the moment the hold arrives, before anything is awaited, and read again once the run is validated: a node that registered another session meanwhile is refusedescrow-run-foreign-session, including when only the cluster in its Cloudly-issued binding changed, and so is a run whose renewals are not in that session's scope bybindApplyRuntimeNetworkDnsLeaseRenewalRequest; - when it is not older than the run already held (below).
Holds apply one at a time, in the order they reach the relay.
Precedence. Projection generation decides first: a run for a later
generation of the held projection replaces it, and one for an earlier generation
is refused escrow-run-stale. For the same projection, the run whose last
renewal has the lower sequence is refused escrow-run-stale; an equal one is
the same top-up again and replaces. A run that holds nothing for the same
projection or a later one always replaces, because giving custody up only
shortens what a detached node keeps. Two projections that cannot be ordered — a
different projection id, or a different digest at the same generation — are
refused escrow-run-unordered rather than guessed.
Release. Only while the relay has no Cloudly session
(controlPlaneReady false), and only the renewal whose window has been open for
the admission margin and has not ended by the relay's clock
(notBefore + margin <= now < expiresAt, the latest such one when two overlap).
A node admits a window only once its own clock has reached notBefore, so an
offer exactly at the start is refused by any node whose clock is a moment behind
the relay's. The margin is 30 seconds (defaultLeaseEscrowTiming): Cloudly's
15-minute windows start every 10 minutes and overlap by five, so it covers
ordinary clock offset and leaves almost all of that overlap.
The relay offers the renewal the moment its session is lost, and then at every
window start plus the margin, as applyRuntimeNetworkDnsLeaseRenewal composed
from the node's current session binding and bound with the contract's delivery
binder. Only a node the forwarder carries is reached. The node's ACK is bound
with bindApplyRuntimeNetworkDnsLeaseRenewalResponse. A refused or failed
delivery is logged and offered again every 30 seconds until the window before
it in the run ends — the window a node that kept up still holds — or, for the
renewal that opens a run, until its own end, and then no more, so a refusal that
clears within seconds is retried in time and one that does not costs a bounded
number of offers. The moment the session is back, the timer is cleared and
nothing more is offered: Cloudly signs and delivers again. The one timer is
unref'd, owned by the escrow and cleared by stop(), which the listener calls
before its node registry empties.
What is due, not what is next, decides the first offer: a run handed over while the session is already lost — Cloudly's last top-up, or one that was still being validated when it went — is offered its current window at once rather than at the next window start, and a re-offer armed for one node never postpones another node's window that has just opened. Waiting for the next start would cost a node a whole grid step, ten minutes, of the five its windows overlap by.
Clock. The relay's clock only chooses which renewal to offer. The node admits a window only if it is current by its own qualified clock, so a wrong relay clock costs continuity and never grants a window early.
Drop rules. A node's held run is dropped when:
- the node registers a different session (a registration answered with the identical binding keeps it);
- the node acknowledges a projection other than the one the run renews or an earlier generation of it, because a node admits projections only forward and never admits a renewal of one it replaced;
- Cloudly refuses a registration of this relay with one of the refusals that
state this relay is not the one it dispatches to —
relay-cluster-not-enabled,relay-node-foreign,relay-node-unknown,relay-sequence-not-advancingorrelay-credential-generation-stale, the last of which says the cluster's relay credential has been rotated to a generation this socket does not hold, matched exactly asClusterRelayErrorwords them bynameClusterRelayRegistrationRefusaland read bystatesThisRelayIsNotDispatchedTo. That is the only election outcome a relay can observe. The first four are a state of Cloudly's that an operator can change there, so the relay is registered and accepted again after them and holds runs again; the fifth parks it for the rest of the process, so nothing is held for its nodes again until it is restarted with the bearer of the current generation. Every other answer is logged and keeps custody: Cloudly'sinvalid,conflictand generic rejections, its internal-error constant, a missing handler, a timeout — andrelay-identity-unverified, which Cloudly answers both for a credential it does not accept and when its own user and cluster reads fail, so it arrives exactly when Cloudly is failing, which is when the runs are needed. Keeping them is safe: every renewal is Cloudly's own signed statement, which the node verifies against its own pinned authority, so the relay gains no authority by holding it, and revocation latency up to the horizon is a limit this escrow already states below. Cloudly owes this one a fix of its own:getVerifiedCredentialContextandgetSingleClusterIdForUserIdmust distinguish "the credential is not valid" from "the records could not be read", so an infrastructure failure surfaces as a generic rejection instead of an authorization refusal. The refusals that drop custody aredata.clusterRelayUndispatchedRefusalsfrom@serve.zone/interfaces;relay-identity-unverifiedis read from Cloudly because the contract does not export it; - its last window has ended, which is checked at every release;
- the escrow stops.
What this does not solve.
- A full WAN cut. A node accepts a new window only with a fresh sample from its NTS-qualified clock, which needs the public internet. The escrow covers "Cloudly unreachable, internet reachable" only.
- Managed tunnels. A renewal extends DNS only. It does not extend an issued managed-VPN credential, and the VPN hub runs inside Cloudly, so cross-node traffic still stops with it.
- Restarts. A relay restart during an outage loses the held runs, so continuity ends with the window the node admitted last. A node restart cannot register, so there is no session to deliver to.
- Revocation latency. A detached cluster keeps DNS authority until its last held window ends. Cloudly must not reuse an address or name before then.
- Two live relays. Cloudly tells neither relay when a second one makes the
cluster
relay-ambiguous, so both may keep a run. They can only offer the same Cloudly-signed renewals, which a node answers as replays. - New network state during an outage, which needs Cloudly's compiler.
Configuration
Coreflow reads every variable below with @push.rocks/qenv, which looks in the
process environment, the .nogit overlays, the systemd credentials of the unit
($CREDENTIALS_DIRECTORY/<NAME>) and the Docker secrets of the service — one
reader for all of them, because a managed relay is configured through what its
service manager hands it and a read straight from process.env would be a
variable an operator can set and the relay never sees. A host release takes the
bearer as a systemd credential; see Host release.
qenv.yml lists the two unconditional ones under required:, and qenv resolves
them while the relay is constructed: a checkout started without one refuses before
any part starts, naming every missing one at once. The image carries dist_ts and
cli.js but not qenv.yml, so a managed relay is refused by the reader of the
variable it lacks, by that variable's name. Both refuse the same set, because every
entry under required: is one its reader treats as mandatory on every boot.
| Variable | Required | Purpose |
|---|---|---|
CLOUDLY_URL |
yes | Cloudly TypedSocket/HTTP endpoint. Its origin is also the registry the cluster's image pulls are forwarded to, because Cloudly publishes the registry and composes the pull realm and token audience from that same public origin. |
SERVEZONE_CLUSTER_RELAY_AUTHORIZATION |
yes | The bearer of this cluster's relay credential, exchanged for the relay identity. Read once at start and kept in memory only: never logged, never sent in a registration body, never written to disk. Boot fails without it, and a bearer whose generation Cloudly has rotated away is replaced by placing the current one here and restarting — nothing returns one over the wire. |
COREFLOW_NODE_ID |
only with a relay block | The cluster node this relay runs on, as Cloudly knows it. No default: Cloudly decides against its own records whether the node is one of the cluster's, so a guessed id would ask it to trust a guess. A node Cloudly does not hold is refused as relay-node-unknown: enroll it with spark enrollnode first. Spark does not pass it yet: a relayed cluster needs Spark to place COREFLOW_NODE_ID, COREFLOW_RELAY_BIND_HOSTNAME and COREFLOW_RELAY_BIND_PORT in the coreflow service environment. |
COREFLOW_RELAY_BIND_HOSTNAME |
only with a relay block | Address the cluster-side listener binds. No default: a daemon must not pick its own network exposure. |
COREFLOW_RELAY_BIND_PORT |
no | Port for that listener; defaults to 8443 and must be the port the relay origin names. |
COREFLOW_VPN_BIND_HOSTNAME |
only with a managed VPN hub | Address the cluster's managed VPN hub binds and advertises, as a unicast IPv4 literal. It must be one address for both: Cloudly composes the cluster's platform endpoint from what is advertised, so a hub that advertised an address it does not bind would be a cluster whose tunnels reach nothing. |
COREFLOW_VPN_QUIC_PORT |
only with a managed VPN hub | UDP port the hub binds for QUIC, the one transport a managed hub serves. |
COREFLOW_VPN_BIND_HOSTNAME and COREFLOW_VPN_QUIC_PORT are both or neither:
one without the other is refused at boot by the name of the missing variable. A
relay given neither serves no hub, advertises none and registers exactly as it
did before.
The pair is read beside COREFLOW_RELAY_BIND_HOSTNAME, on a boot whose cluster
has a relay block; a relay without one serves no listener and reads neither.
Whether the hub it advertises carries any traffic is decided outside this
process:
- The address must lie inside the runtime network plan's
vpn.hubPrefix. That prefix is Cloudly's (configureManagedVpn), never the relay's. Cloudly persists any advertisement the contract's validator admits, but it selects a hub only when the plan's hub prefix covers the advertised address: a hub outside it registers, and its cluster's nodes stay without VPN membership. - Every node of the cluster must reach the UDP port. A node dials
address:quicPortover QUIC directly, so the host firewall and the routing between the nodes' subnets must admit UDP to it, and a container that does not share the host's network must publish it as UDP. The address may be the one the listener binds: the listener is TCP, the hub UDP. - A bind the host refuses does not fail the boot. The daemon binds when
Cloudly states the cluster's first network, not at start. An address the host
does not hold, or a port another process holds, is logged as an error and
answered
state: 'failed'to that push; the relay keeps serving its nodes and the next push binds again. - Placing, changing or removing the pair takes a restart, and a restart re-keys the hub and drops its cluster's tunnels (see The cluster's managed VPN hub).
COREFLOW_INTERNAL_HOSTNAME and COREFLOW_INTERNAL_PORT were removed in 33.0.0
with the internal server they configured. Nothing reads them: a deployment that
still sets them boots unchanged and binds nothing on that port, and they can be
dropped from its environment.
Corestore needs no variable of its own any more. Where each node's Corestore
answers comes from the cluster document and the control credential from Cloudly,
so neither CORESTORE_CONTROL_URL nor CORESTORE_API_TOKEN is read.
Example .nogit/env.yml — qenv reads env.json, env.yml or env.yaml from that
directory, never a .env:
CLOUDLY_URL: https://cloudly.example.com
SERVEZONE_CLUSTER_RELAY_AUTHORIZATION: the-bearer-of-this-cluster-relay-credential
COREFLOW_NODE_ID: node-a
COREFLOW_RELAY_BIND_HOSTNAME: 192.0.2.15
Starting the relay
pnpm install
pnpm build
node cli.js
ts/index.ts exports runCli, runHostRelay (the host release's entry) and stop; Coreflow.start() starts the internal
server, opens the Cloudly session, registers every handler and only then opens
the relay for dispatch. A failure in any step rolls the started parts back in
reverse order. Registering with Cloudly and reconciling platform bindings are not
steps of the start: the relay registers once it is open, and a refused or late
acceptance leaves it running rather than failing the boot.
runCli stops the relay on SIGTERM or SIGINT and then ends the process: 0 when the
stop completed, 1 when it failed. A second signal while the relay stops ends the process at
once with 1. stop() is for a caller that owns the process; it also releases both signals.
A checkout or the image starts the relay without options: it reads qenv.yml from its
working directory, with the .nogit overlay beside it, and runs the smartvpn daemon of
its installed package. An entry that knows where its launch facts are passes
ICoreflowOptions to runCli (or to new Coreflow()):
| Option | Effect |
|---|---|
qenvDirectory |
The directory that holds qenv.yml. The relay then reads no .nogit overlay: a supervised relay is configured by its service manager only. |
vpnDaemonPath |
The exact smartvpn daemon executable the cluster's managed VPN hub runs. One that is missing or not executable fails the hub's bind by name; the package's own daemon is never used instead. |
Both must be absolute, normalized paths; anything else, and any other option, is refused by
name when the relay is constructed. The version the relay logs and registers with is the one
built into it (ts/00_commitinfo_data.ts), so it reads no package.json at run time.
Host release
The relay also ships as a host release: a self-contained Linux executable for linux-amd64 and
linux-arm64 that needs no Node.js, Deno or Docker on the host. Spark installs and supervises it
as a systemd unit. The Docker image stays as it is.
Release artifact
Every release tag publishes a sealed @git.zone/tspack set named coreflow-relay as assets of
the Gitea release v<version> of serve.zone/coreflow. The tag's CI job builds and publishes it
(.gitea/workflows/release.yml, managed by GitZone's tspackRelease assets).
| Asset | Content |
|---|---|
tspack-manifest.json |
Package @serve.zone/coreflow, name coreflow-relay, version, source commit, publishable: true, and the digest and mode of every file in each archive. |
SHA256SUMS.txt |
The digests of the manifest and both archives. |
coreflow-relay-<version>-linux-amd64.tar.gz, coreflow-relay-<version>-linux-arm64.tar.gz |
One release directory each. |
A consumer pins the version, the source commit and the manifest's SHA-256, downloads the manifest,
the checksums and its own archive, and verifies and extracts that archive with
TsPack.extractBundle (bundle IDs linux-amd64 and linux-arm64). That is the shape Spark
already consumes Pallet in.
Each archive holds one directory, and nothing else:
| File | Mode | Purpose |
|---|---|---|
coreflow |
0755 | The relay, compiled with Deno 2.9.7 from the built dist_ts (binary/coreflow-relay.js). The unit's ExecStart. |
qenv.yml |
0644 | The launch facts the relay refuses to start without. |
coreflow-vpn |
0755 | The smartvpn daemon the cluster's managed VPN hub runs, taken from the pinned @push.rocks/smartvpn release. |
coreflow-vpn.tsrust-build.json |
0644 | That daemon's tsrust build provenance: its project, version, commit, target and digest. |
coreflow.third-party-notices.txt |
0644 | The third-party notices of coreflow, written by @git.zone/tsdeno at compile time: the components of the Deno 2.9.7 runtime it contains (V8, ICU, Rust crates, Deno's own JavaScript) with their license expressions, where the runtime's complete corresponding source is obtained (and, for its LGPL components, how to rebuild the runtime and have deno compile use it), every npm package it embeds with the license and notice files that package ships, and every notice text in full. |
coreflow.third-party-notices.json |
0644 | The same notices, machine-readable (generator, binary, target, runtime, packages). |
license |
0644 | This project's license. |
notices/smartvpn/… |
0644 | The daemon's license and distributed notices, as its notice inventory lists them. |
The executable finds qenv.yml and coreflow-vpn beside itself (hostReleaseLayout,
createHostLaunchOptions) and reads nothing from its working directory and no .nogit overlay.
The relay writes into no file, so the directory can be read-only.
Unit contract
What a unit running the host release states, and what the relay answers:
| Aspect | Contract |
|---|---|
ExecStart |
<release directory>/coreflow, no arguments. |
Type |
exec. The relay sends no READY=1; readiness is the probe below. |
| User | Any. The relay needs no capability and no root: the managed VPN hub creates no TUN, route or NAT. DynamicUser=yes suits it. A port below 1024 for the listener or the hub would need AmbientCapabilities=CAP_NET_BIND_SERVICE. |
| Environment | CLOUDLY_URL, COREFLOW_NODE_ID, COREFLOW_RELAY_BIND_HOSTNAME, COREFLOW_RELAY_BIND_PORT (default 8443), and the pair COREFLOW_VPN_BIND_HOSTNAME and COREFLOW_VPN_QUIC_PORT when the cluster has a managed VPN hub; see Configuration. No other variable configures the relay. |
| Bearer | The credential SERVEZONE_CLUSTER_RELAY_AUTHORIZATION, by that exact name: LoadCredentialEncrypted=SERVEZONE_CLUSTER_RELAY_AUTHORIZATION (sealed with systemd-creds encrypt, found in /etc/credstore.encrypted/) or LoadCredential= for a plain file. Never Environment=. The value is read byte for byte, so the stored bearer carries no trailing newline. |
| Ports | TCP COREFLOW_RELAY_BIND_PORT on COREFLOW_RELAY_BIND_HOSTNAME (the relay origin, 8443); UDP COREFLOW_VPN_QUIC_PORT on COREFLOW_VPN_BIND_HOSTNAME while a managed network is held. The relay dials Cloudly at CLOUDLY_URL. |
| Readiness | The relay listener accepts TCP connections on COREFLOW_RELAY_BIND_HOSTNAME:COREFLOW_RELAY_BIND_PORT. It opens as the last step of the start, after the Cloudly session, the Corestore credential and the cluster config; the journal line the relay is open for Cloudly dispatch marks the same moment. A cluster without a relay block opens no listener. |
| Stop | SIGTERM (or SIGINT). The relay closes its Cloudly session and stops the VPN daemon itself, so the unit uses KillMode=mixed: with control-group the daemon would receive the signal beside the relay and the relay's own stop of it would race. A stop waits for Cloudly requests in flight; TimeoutStopSec= bounds it, after which systemd kills what is left. |
| Exit codes | 0: stopped on a signal and the stop completed, also when the signal came during the start. 1: the start failed (a missing variable or credential, a refused bearer, an unbindable listener; the cause is on stderr), the stop failed, or a second signal arrived while the relay stopped. |
| Restart | Restart=always with a short RestartSec=: every exit of a supervised relay is a relay that should run again, as --restart-condition any stated for the Swarm service. |
| Logs | stdout and stderr, one line per event, to the journal. The bearer is never written. |
Do not set MemoryDenyWriteExecute=: the executable embeds V8, which compiles code at run
time.
Building it locally
pnpm run build:relay
build:relay builds dist_ts and stages both targets under dist_relay/<target>
(scripts/build-relay.ts, scripts/relay-build.ts). It needs the pinned Deno, which is the
denoVersion of @git.zone/cli.assets.sealedRelease in .smartconfig.json. A staged directory
must hold exactly the files its @git.zone/tspack bundle declares, or the build fails naming the
difference; the smoke check starts the host target's executable once in an empty environment and
expects it to log its start and refuse the missing launch facts. Each compile writes the
executable's third-party notices beside it and fails instead when tsdeno cannot account for the
pinned Deno's runtime or for an embedded npm package without license text. The sealed set itself
is built only from the clean tagged source, by the release job (scripts/release-relay.ts).
test/test.hostrelease.node.ts builds both targets, seals them, checks the notices in each
extracted archive, and runs the relay from this host's archive the way its unit does.
Development
pnpm install
pnpm build
pnpm test
pnpm run watch
pnpm run build:relay
pnpm run build:docker
| Path | Purpose |
|---|---|
ts/index.ts |
CLI entry point exporting runCli, runHostRelay and stop. |
ts/coreflow.hostrelease.ts |
The layout of a host release and the launch options its executable derives from it. |
ts/coreflow.classes.coreflow.ts |
Composition root and lifecycle. |
ts/coreflow.connector.cloudlyconnector.ts |
The Cloudly session. |
ts/coreflow.classes.corestoreclient.ts |
Corestore control API client, one call per node. |
ts/coreflow.classes.corestorecredential.ts |
The sealed Corestore control credential and this relay's recipient. |
ts/coreflow.classes.corestoreendpoints.ts |
Where each node's Corestore answers. |
ts/coreflow.classes.corestoreinventory.ts |
What one node holds, its legacy resources, or why it could not be read. |
ts/coreflow.classes.corestorepublication.ts |
Sealing a binding's credential material to Cloudly's ingress recipient and verifying the receipt. |
ts/coreflow.classes.clusterbackup.ts |
Backup, restore and archive pruning, per node, and when a closure is carried. |
ts/coreflow.classes.closurecarrier.ts |
The database-closure transfer itself, both directions. |
ts/coreflow.classes.corestorerefusals.ts |
How one Corestore failure is named for Cloudly. |
ts/coreflow.classes.platformmanager.ts |
Platform binding reconciliation. |
ts/coreflow.classes.objectstorageretention.ts |
Compliance-retention intent capture and the expectation built from the grant. Verification itself is the shared validator's. |
ts/coreflow.classes.retiredhandlers.ts |
Named refusals for retired Swarm methods. |
ts/coreflow.classes.namedrefusals.ts |
The refusal registrar both refusal sets use. |
ts/coreflow.classes.relaylistener.ts |
The cluster-side TLS listener Pallet nodes dial. |
ts/coreflow.classes.relaycertificate.ts |
The listener's certificate, obtained from Cloudly. |
ts/coreflow.classes.relayregistration.ts |
Telling Cloudly which relay this is and where it listens. |
ts/coreflow.classes.sessionforwarder.ts |
Runtime-session forwarding, both directions, and outage continuity. |
ts/coreflow.classes.clusteringress.ts |
The cluster's route table and the ingress that serves it. |
test/test.relayboot.node.ts boots the relay against a loopback Cloudly and is
the fastest way to see the whole startup path.
Operational notes
Upgrading from 3.5.x. This release replaces the Swarm reconciler with the cluster relay. A running 3.5.x deployment is not upgraded in place: a cluster is flipped together with the Cloudly that talks to it, in the cutover sequence the platform design states (Cloudly 32 first, then the cluster's relay and its nodes, one cluster at a time). Starting this image against a Cloudly that still expects the reconciler leaves the relay standing but not ready, with the refused protocol retained on its status.
The released Spark daemon cannot start this relay. The Swarm-era spark asdaemon path
creates the relay's managed service with a Docker secret coreflowSecret that carries
SERVEZONE_ENVIRONMENT, CLOUDLY_URL and JUMPCODE only. This image never reads JUMPCODE
and boots only with SERVEZONE_CLUSTER_RELAY_AUTHORIZATION in its launch environment, so a
relay started through that path fails to boot naming the variable. A relay run from the image is
deployed out-of-band, by whoever holds the cluster's bearer, with that variable in the service's
secret. A host without Docker runs the host release as a unit instead, with the
bearer as that unit's systemd credential.
- One relay per cluster. It is the only holder of the cluster's Cloudly credential; node enrollment (Spark) still talks to Cloudly directly.
- A Corestore call fails closed when its credential is missing; it never falls back to an unauthenticated request, and it never reads one from the environment.
- Corestore data is node-local; snapshot before destructive changes or restores.
- The package is private and ships as part of the serve.zone platform deployment.
License and Legal Information
This repository contains open-source code licensed under the MIT License. A copy of the license can be found in the license file.
Please note: The MIT License does not grant permission to use the trade names, trademarks, service marks, or product names of the project, except as required for reasonable and customary use in describing the origin of the work and reproducing the content of the NOTICE file.
Trademarks
This project is owned and maintained by Task Venture Capital GmbH. The names and logos associated with Task Venture Capital GmbH and any related products or services are trademarks of Task Venture Capital GmbH or third parties, and are not included within the scope of the MIT license granted herein.
Use of these trademarks must comply with Task Venture Capital GmbH's Trademark Guidelines or the guidelines of the respective third-party owners, and any usage must be approved in writing. Third-party trademarks used herein are the property of their respective owners and used only in a descriptive manner, e.g. for an implementation of an API or similar.
Company Information
Task Venture Capital GmbH
Registered at District Court Bremen HRB 35230 HB, Germany
For any legal inquiries or further information, please contact us via email at hello@task.vc.
By using this repository, you acknowledge that you have read this section, agree to comply with its terms, and understand that the licensing of the code does not imply endorsement by Task Venture Capital GmbH of any derivative works.