2026-10-08 11:57:26 +00:00
…
2026-10-08 11:57:26 +00:00
2026-10-08 11:57:26 +00:00
…
2026-10-08 11:57:26 +00:00
…

@serve.zone/interfaces

@serve.zone/interfaces is the shared TypeScript contract package for the serve.zone ecosystem. It contains the public data shapes and TypedRequest interfaces used by Cloudly, Coreflow, Spark, Coretraffic, platform clients, SDKs, and external integrations to exchange infrastructure state without duplicating DTOs.

Issue Reporting and Security

For reporting bugs, issues, or security vulnerabilities, please visit community.foss.global/. This is the central community hub for all issue reporting. Developers who sign and comply with our contribution agreement and go through identification can also get a code.foss.global/ account to submit Pull Requests directly.

Install

pnpm add @serve.zone/interfaces

Public API

The root export exposes six namespaces:

import { appstore, data, platform, platformservice, protocol, requests } from '@serve.zone/interfaces';
Namespace Purpose
appstore App Store catalog, manifest, service requirement, and upgrade contracts.
data Durable platform object shapes such as clusters, services, deployments, images, domains, DNS entries, secrets, users, status, settings, backups, registries, BaseOS metadata, and task executions.
requests TypedRequest contracts for Cloudly and serve.zone control-plane RPC methods.
protocol The handshake two peers exchange when a session opens: what each side speaks, the oldest peer it accepts, and the named refusal when they cannot serve one session.
platform Current platform-service contracts for email, SMS, push notifications, letters, AI, databases, object storage, logging, backups, and SIP.
platformservice Legacy platform-service namespace kept for older consumers that still depend on the previous layout.

This package intentionally has no service implementation logic. It is a stable vocabulary for services that need to agree on payload shape, method names, and response types.

Identifier Vocabularies

Two vocabularies name things in these contracts, and which one a member uses is part of the contract.

Vocabulary Rule Members
Canonical identifier ^[A-Za-z0-9][A-Za-z0-9:._-]{0,199}$ — first character alphanumeric, up to 200 characters organizations, clusters, services, namespaces, sessions, assignments, attempts, authorities, pools, endpoints and every content id derived from them
Node id ^[A-Za-z0-9_-]{1,128}$ — the URL alphabet, a leading - or _ included, up to 128 characters every member that names a node of a cluster

A node may be named -abc or _abc: enrollment, the runtime session registration and the Spark wire admit a node id that begins with a separator, so every member that names the same node admits one too. data.isClusterNodeId(value) is that rule, and it is the rule at every node-bearing member: nodeId on a session binding, a registration, an enrollment, a Spark heartbeat, a relay block, an assignment, a workload lease, a handoff lease, a DNS-lease renewal run, a protection receipt, an egress authority, a cluster VPN node, an ingress registration, a traffic bucket, a Corestore inventory and an isolated-restore node control; id on a router incarnation and on a node retirement; cloudlyNodeId on a cluster runtime target and on a secret runtime target; and placement.nodeIds on a service runtime spec. A service that stores or forwards a node id should judge it with the same reader rather than with the canonical one, which would refuse a node these contracts themselves enrolled and registered.

A replica id is the one composed name: a producer names a replica slot <node id>.<index>, counted from zero. data.isRuntimeReplicaId(value) reads it by decomposition — the part before the last dot in the node vocabulary, the index on its own and never zero-padded — and reads a name that carries no such index, a zero-padded node-a.01 among them, as an opaque slot name in the canonical identifier vocabulary.

WorkloadInit release verifiers should use the dedicated Node-compatible subpath:

import {
  validateWorkloadInitReleaseAttestationStatement,
  validateWorkloadInitReleaseIdentity,
  workloadInitReleaseContract,
} from '@serve.zone/interfaces/runtime/workloadinit';

@serve.zone/interfaces/runtime/workloadinit exposes only the WorkloadInit release identity, attestation, approval, authority, digest, and validator contracts. Its runtime dependency graph is limited to the immutable-image digest validator; unlike @serve.zone/interfaces/runtime, it does not load the general runtime graph, plugins.js, SmartCrypto, TypedRequest, or Corestore runtime modules.

Directional Image Streams

Image transfer contracts use @api.global/typedrequest-interfaces 7.1.0 and the directional VirtualStream protocol used by TypedRequest 8 and TypedSocket 8. Directions in shared DTOs describe the requesting peer:

Contract field Requester endpoint TypedHandler endpoint
requests.image.IRequest_PushImageVersion.request.imageStream TVirtualStream<'send'> TVirtualStream<'receive'>
requests.image.IRequest_PullImageVersion.response.imageStream TVirtualStream<'receive'> TVirtualStream<'send'>

The method names remain pushImageVersion and pullImageVersion. Each stream carries ordered Uint8Array chunks. TypedRequest reverses stream directions for the handler automatically; do not reverse the shared declaration yourself.

Consumers must replace the removed undirected IVirtualStream API with explicit transport-created endpoints. Senders use send() or writable, then close; receivers use receive() or readable, then explicitly accept() after draining EOF and completing application storage or delivery. completion confirms the receiver's accepted receipt, while closed confirms transport cleanup. The upload response's allowed flag permits sending; it does not confirm storage. Receiver rejection, aborts and failed completion must reach the caller. Existing deployment log and shell push-message contracts retain their message shapes.

Human Credential Administration

requests.admin.getHumanCredential and mutateHumanCredential describe human credential inspection and mutation for recently authenticated administrators. Operations rotate a password, revoke password login, or revoke existing sessions. Each successful mutation advances the generation and invalidates older human UI/API and OCI registry sessions. Upstream OIDC identity bindings remain intact.

Send the expected generation and a stable mutation ID. After a lost response, authenticate again and repeat the same request. A replay returns the original metadata without applying the operation again; inspect current metadata separately. Requests containing passwords must never be logged. Responses contain no passwords or verifiers. Cloudly owns validation, atomic persistence, and session enforcement; these shared contracts alone do not implement an endpoint.

Node Credential Administration

requests.node defines getNodeCredential, rotateNodeCredential and revokeNodeCredential for verified platform-infrastructure administrators. getNodeConfig now also requires identity; callers using the previous unauthenticated request shape must update. Its response remains the public IClusterNode, never a persisted backend document.

data.INodeCredentialMetadata separates spark and pallet purposes and exposes lifecycle state, generation, rotation-required status and session epoch, but no bearer, hash, socket peer or controller process identity. Infrastructure credentials do not grant organization or workload permissions.

Mutations require the exact generation and session epoch plus a stable mutation ID. Rotation takes the SHA-256 hash of a node-generated CSPRNG credential that the node has already durably retained with that mutation ID. Lost-response retries repeat the complete request and receive the original metadata, never newly minted plaintext. A replay result is historical and cannot authorize a current session. Treat rotation request hashes as sensitive in transport hooks and logs.

These types do not implement authentication, storage, enrollment or endpoints. Cloudly must verify current administrator authority, validate complete inputs, atomically persist the credential transition and immutable actor-bound receipt, and fence subsequent control effects against current durable credential/session authority. Node identity storage and transport integration must be qualified before enabling rotation or Pallet enrollment.

Spark Host Reporting

data.sparkNodeHeartbeatContract defines the HTTPS POST endpoint /spark/nodes/heartbeat, transport byte limits and protocolRefusalStatus, the status a refused offer is answered with. Its exact request snapshot contains a Spark-purpose bearer, the sender's protocol offer and ISparkNodeReport: fractional host CPU, memory and disk observations, Linux host facts, and the verified bundle version, source commit and manifest digest. It has no Docker or container-count fields. Local runtime readiness remains distinct from unverified workload readiness.

The route reads the bounded body, reads the offer with protocol.readProtocolOffer, negotiates it against its own sparkNode offer and answers a refusal with protocolRefusalStatus and an IProtocolRefusal body — all before the node is authenticated, so a peer of another major is refused by name instead of by a credential verdict it cannot act on. The receipt carries the controller's own offer, so an accepted node judges the session it just reported into.

The Swarm-era Spark runtime keeps its own four routes in data.sparkSwarmNodeContracts, each stating the same three facts as the route above — heartbeat (/spark/swarm-nodes/heartbeat), metricsSample, actionResult and swarmObservation, every one of them an { endpoint, maxRequestBytes, maxResponseBytes } — so one bounded read serves every Spark route. The bounds differ because the bodies do: a heartbeat states the node's whole runtime description, a metrics-sample answer carries every action queued for that node, and one observation carries up to 1024 observed Swarm nodes. That is a second runtime, not an older version of this one: a worker in coreflow-node mode posts those bodies and a pallet node posts this one, and the family retires with coreflow rather than with this release. ISparkSwarmNodeHeartbeatRequest is now owned here, so Cloudly and Spark no longer carry private copies of it, and every Swarm-era body — request and answer alike — carries protocol and is judged in the same order. For fleet cutover, that same observation may carry a bounded local Docker snapshot from the reporting node and, only when control is available, a manager-visible service/task snapshot. The existing node-token, reporter-session, sequence, digest and accepted-receipt rules authenticate both; no separate fleet-report route exists.

Each of the four routes has its validator here, so a server never writes its own member checks: data.validateSparkSwarmNodeHeartbeatRequest, data.validateSparkMetricsSampleRequest, data.validateSparkActionResultRequest and data.validateSparkSwarmObservationRequest. All four answer the same way — one string per refused member, naming the member, and an empty list for an admissible body — so a route table can hold them side by side. They read the envelope first (exact key set, the sender's offer, then the reporter's identity), then the payload: a heartbeat's metrics and full runtime description including each described serve.zone service, a sample's four observations and optional network rates, and an action result whose status is one a node can actually report, never pending. Bodies are judged member by member rather than as canonical bytes, because only the observation is digested and the others carry fractional CPU, memory and rate observations; every level is copied away from the caller's object first, so an accessor is refused instead of run. None of them authenticates a node or acts on the body.

Cloudly must authenticate the bearer against current Spark authority and fence credential generation and session epoch transactionally before recording host liveness. ISparkNodeHeartbeatReceipt contains only server-derived acceptance evidence; the sender must match its node ID and credential generation to its exact active identity. Neither a receipt nor the sender's observation time grants current authority or workload readiness. Persist only the report and receipt in IClusterNode.data.sparkNodeReport, never the bearer. Keep legacy reporting separate. This endpoint does not deliver or acknowledge operator actions.

data.sparkNodeHostContract defines the node's desired-host route, /spark/nodes/host, with 1024-byte request and response bounds, authenticated and judged in the same order as the heartbeat. ISparkNodeHostRequest carries the Spark bearer and the sender's offer; ISparkNodeDesiredHost answers with the whole desired value of IClusterNode.data.hostnameIntent — the hostname and its generation — never a change, and hostname: null with generation 0 states that Cloudly holds no intent, so Spark leaves the host's hostname as it is. Spark applies the latest generation and states the outcome of its last attempt as the optional ISparkNodeReport.hostApply (applied, or failed with a bounded reason) on its next heartbeat; the node has converged when runtime.hostname states the intent's hostname. A Spark that sends hostApply raises its minimumPeerVersion so it reaches only a Cloudly that reads it.

The snapshot functions reject unknown keys, accessors, malformed identities and nonfinite/out-of-range numbers, then return detached frozen records. Transports own byte limits, timeouts and cancellation. These contracts do not implement a listener, credentials, persistence, freshness or shutdown.

Independent Node Enrollment

requests.node defines getNodeEnrollmentState for current Spark-authenticated adoption preconditions and enrollNode for atomic Spark/Pallet enrollment. These contracts do not implement endpoints, authentication, a database or a daemon.

The enrollment proposal binds the canonical HTTPS Cloudly origin, hostname, enrollment identity, original bootstrap proof hash and an exact ordered pair of independently prepared Spark/Pallet credential hashes and CAS counters. Fresh Jump starts both at zero; existing-node adoption rotates current Spark and requires Pallet authority to be absent. snapshotNodeEnrollment returns a detached, frozen snapshot; computeNodeEnrollmentDigest hashes strict canonical JSON under nodeEnrollmentContract.digestDomain, serve.zone/node-enrollment. Neither helper generates or persists credential material.

Transport proof is excluded from the digest to permit exact lost-response recovery. Bootstrap proof must hash to the original proof hash and is allowed only before a receipt exists. Pending-Spark proof must hash to the prepared Spark hash and is allowed only for an existing exact receipt. verifyNodeEnrollmentProof checks this binding and mode only: Cloudly must derive receipt state itself and fence live bootstrap/credential authority in the same owned transaction. Current Jump codes are 12 random bytes encoded as 16 canonical base64url characters.

bindNodeEnrollmentAcknowledgement checks the complete digest, both owners, exact next generations and unchanged adopted node ID, returning a detached frozen acknowledgement. Authenticate Cloudly before accepting it. The acknowledgement is immutable commit evidence, not ongoing authority or workload readiness. Server replay must still verify both active credential generations and hashes; reporter reconnect alone does not invalidate the credential. Spark and Pallet commit their local activation independently and converge after restart; there is no shared plaintext vault or cross-local-database ACID claim. Never log proof/request bodies or hashes, put them in URLs, or expose them through transport hooks.

requests.pallet defines the node-local preparePalletNodeEnrollment, bindPalletNodeEnrollment and activatePalletNodeEnrollment methods. They belong only on the protected root-owned Unix control socket, never a network router. Preparation requests initial Pallet authority (zero generation and session epoch) and returns only its independently retained credential hash. snapshotPalletNodeEnrollmentPreparation and bindPalletNodeEnrollmentPreparation validate that exact exchange.

Local confirmations explicitly distinguish a durable pending binding from the exact current active Pallet identity. bindPalletNodeEnrollmentBound verifies the whole enrollment digest and its pre-assignment node ID (null for fresh Jump). snapshotPalletNodeEnrollmentActivation captures the exact activation request, rejecting extra wrapper fields as well as malformed nested data. bindPalletNodeEnrollmentActive takes that complete activation request and verifies the full Cloudly acknowledgement, including both owners and the assigned node ID. Both confirmation binders return detached frozen snapshots captured before hashing. These are pure content checks: authenticate the local transport separately, and have Pallet prove its current owned state before issuing a confirmation. A historical receipt is insufficient, and Pallet activation does not imply Spark activation or workload readiness.

palletNodeEnrollmentContract specifies one four-byte unsigned big-endian length-prefixed UTF-8 JSON frame per direction, followed by write-half-close, at most 32768 request bytes and 16384 response bytes excluding the prefix. Implement the listener with half-open response support, strict EOF/trailing-byte rejection, bounded time/concurrency, root-only directory/socket permissions and drained cleanup. The package supplies no listener, permission checks or persistent state. No bearer, generic rotation operation or database access is part of this IPC API.

Node Runtime Routing Binding

requests.node.getNodeRuntimeBinding { nodeId, nodeToken } is a current Spark-node-authenticated read. Cloudly derives data.INodeRuntimeBindingRead from its current credential, node, cluster, durable runtime-controller, relay and cluster-runtime records. The answer contains the exact routing identity { cloudlyOrigin, nodeId, clusterId, controller, runtimeNamespace, relay }, the current { id, phase, generation } runtime reference and the Spark credential generation and session epoch that authenticated the read. snapshotNodeRuntimeBindingRead requires the runtime id to equal the routing cluster, and bindNodeRuntimeBindingRead binds the answer to Spark's exact immutable Cloudly origin, node and credential generation. Neither helper authenticates the server.

requests.pallet.bindPalletNodeRuntimeBinding and readPalletNodeRuntimeBinding belong on the existing root-owned Pallet enrollment control socket. Bind is absent-to-present only: the current active Pallet enrollment must name the same canonical Cloudly origin and node, an exact replay is idempotent and every different origin, node, cluster, controller, namespace or relay conflicts. Read requires the requested origin and node to equal both the active enrollment and the stored binding; an absent binding rejects. Spark applies and reads back the binding before starting runtime-serve. On an offline reboot it may read the exact persisted routing identity to select the relay, but that record grants no workload permission and no right to advance or run a cluster phase. Fresh authenticated runtime-session state remains the only workload authority.

The local operations carry no bearer, signing key, generic configuration field, phase mutation or caller-authored permission. nodeRuntimeBindingContract states the immutable local mutation and routing-only persistence rules.

A node can instead be driven by a controller on its own host, such as Onebox. data.INodeRuntimeLocalRoutingBinding is { nodeId, clusterId, controller, runtimeNamespace, registryHosts } with a controller of kind onebox; it names no Cloudly origin and no relay. nodeRuntimeLocalControllerContract fixes the transport: the controller is the TypedSocket client on Pallet's root-only Unix socket /run/serve.zone/pallet/controller.sock (mode 0600, server URL http://localhost), whose file permissions authenticate the peer. A node carries a Cloudly binding or a local one, never both, and the local binding, like the Cloudly one, grants no workload permission by itself.

The first request on a local connection is bindPalletLocalRuntimeController (see Authenticated Runtime Sessions). It is first bind or exact replay. Pallet answers data.INodeRuntimeLocalBindingConfirmation: the exact persisted binding plus the generation and SHA-256 hash of a bearer that Pallet generates and keeps. bindNodeRuntimeLocalBindingConfirmation binds that answer to the binding sent and to the credential the controller pinned in its own durable state at the first bind (null for that first bind). A different credential for the same binding rejects rather than being adopted, because it means the node lost the state the controller enrolled.

registryHosts lists, strictly ordered, the canonical host[:port] registries (isCanonicalRegistryHost) whose pull credential the local controller issues. isNodeRuntimeLocalRegistryWorkload is the local registry rule: a workload's registryHost, and its pullEndpoint when one is stated, must both be listed.

Service Machine Credentials

requests.admin.getServiceMachineCredential inspects redacted metadata; mutateServiceMachineCredential creates, rotates or revokes authority for one exact existing service and organization. ensure requires absent authority and a null expected generation. Rotation and revocation require the exact current generation. Retain the mutation ID and complete input for retry; a receipt replay returns historical metadata without restoring authority.

IServiceMachineGrant currently grants only platform:session. This is separate from deployment grants, human membership and infrastructure permissions. The backend must authenticate each JWT against the current active credential generation, expiry and exact service ownership. Structural snapshot helpers do not perform these checks. Credential changes, encrypted service SecretSet delivery and the replay receipt must commit in the same authorized transaction. The response contains only the immutable secret/version reference; no bearer or hash is returned. Revocation removes current delivery from the metadata.

Private Networks

requests.network defines organization-scoped network create/update/retire and service attachment operations. Mutations require a retained idempotency ID and an exact expected revision; null means no prior document. Cloudly must authorize each request against live canonical membership and policy, reserve aliases transactionally, reject cross-organization attachments, and retain historical receipts. These contracts do not implement the backend or node enforcement.

data.buildPrivateNetworkFqdn produces <alias>.n-<26 lowercase base32 characters>.o-<26 lowercase base32 characters>.internal. from server-owned immutable DNS keys. Aliases are lowercase ASCII labels unique per network. Display-name edits do not rename DNS. A sole attachment selects its search suffix; several attachments require an explicit default or no search suffix. Fully qualified names remain unambiguous across networks. Network policy permits member-to-member connectivity and explicitly configures or denies external DNS forwarding. Creating an alias does not publish public DNS, ingress or ports.

data.composeServicePrivateNetworkDnsPolicy(membership, networks, resolvers) derives the service's network suffixes, the single selected search suffix, and the intersection of external forwarding permissions. resolvers is the complete resolver list of the protected authority the service is projected under. A service attached to no private network forwards any public name to the first four of those resolvers in address and port order; when the authority declares no resolver it has no external resolution (externalDns: null, so its node answers REFUSED for every name). An attached service is decided by its networks alone: a network denying external DNS keeps denying its members whatever the authority declares. Supply every attached network exactly once, with active state and matching canonical organization/DNS keys. Missing, extra, inactive or inconsistent network definitions are rejected. Forwarding requires every attachment to permit both the queried public name and the same exact upstream IP/port; a denied or empty intersection returns externalDns: null. The default network changes only short-name search. Nested suffix unions are intersected at DNS label boundaries and canonicalized without dropping permitted branches; the composed list can contain up to 1,024 suffixes across 16 networks. This pure helper returns detached policy metadata. Its caller must authenticate and fence the included membership/network revisions before issuing node authority.

The snapshot helpers validate and detach content; they do not authenticate it, reserve names, prove membership, or report applied state. Packet authority remains until an applied denial or independent fence. DNS snapshots have a separate maximum 15-minute validity, with positive TTLs capped at five seconds and negative TTLs at one second. DNS expiry does not prove packet revocation or permit identity reuse. Retirement and attachment responses describe desired or historical metadata; consumers must separately track outstanding node revocations and readiness.

IRuntimeNetworkProtectedAuthority declares the complete controller-owned IPv4 protection set, disjoint private workload/transit pools, protected resolver and platform endpoints, and every affected egress authority, including offline nodes. Its snapshot and digest helpers bind the exact predecessor and all declaration content. They cannot prove that an operator's inventory is complete. Install the union of old and new protection on every affected egress owner, or independently fence that owner, before exposing a newly allocatable pool. Removing an entry from the declaration is not a revocation acknowledgement. Prefix arrays are sorted lexically and nonoverlapping; pools and endpoints are sorted by ID; egress owners are ordered by canonical JSON of [nodeId, runtimeNamespace].

continuesRuntimeNetworkProtection(previous, next) states which successor of a protected authority only adds. It answers true when next is previous itself, or exactly the next step of its chain (the same id and controller, next.previous naming previous) in which every egress owner, pool, resolver and platform endpoint of previous is still present with the same content and the prefixes still cover every address previous protected, however they are split. A cluster topology change is such a step: a node joins the egress, or a cluster's hub joins the platform endpoints. The step keeps its barrier: every egress owner still reports a receipt for next before an allocation binds to it. Any removal or change, a jump over a step, another chain or another controller answers false. Both authorities and their digests are validated, and either that does not validate is rejected.

IRuntimeNetworkProtectionReceipt reports a worker's durable, joined native protection journal separately from projection admission. It binds a complete protected-authority reference, node/controller scope, authenticated reporter, boot identity and exact native journal reference to a monotonic receipt chain. A null nativeBarrier explicitly removes allocation eligibility. A node sends one only in answer to the retirement request of a signed projection (see Projection Retirement Requests): its egress ownership is being retired with the withdrawn disposition, and it mints no positive receipt for the chain afterwards (runtimeNetworkProtectionContract.withdrawal). Nothing else withdraws. The node's allocation-pool guard is retained in the kernel when the node process ends, and its boot unit restores it before networking starts, so a clean stop, a crash, a reboot or being offline leaves the last receipt current. A reconnect binds a new reporter session, which the old receipt no longer matches, and the node's first positive receipt under that session restores eligibility. A boot or protection change requires new native journal evidence. These portable values do not inspect a kernel or authenticate their own producer.

reportRuntimeNetworkProtection uses the current authenticated physical peer. Its request binder accepts a historical outbox only against receiver-owned reporter history and identifies that use as historical. The response must match the exact sent receipt. bindRuntimeNetworkAllocationProtection requires the current persisted protection receipt and durable session for every declared egress owner, in canonical owner order. Missing, unavailable, stale-protection and superseded-session evidence reject. An offline owner cannot be omitted. The controller must fence all those records together with lease insertion, before reserving any handoff or workload address. There is no generic independent-fence flag, quorum or expiry: the one way an owner leaves is Egress Owner Retirement. Protection receipts neither prove workload readiness nor release packet/address/name quarantine.

IRuntimeNetworkHandoffLease binds an immutable node/router incarnation to its protected-authority reference, transit subnet and peers, protocol-qualified source ports, nonzero conntrack zone and 128-bit label. The label is exactly 32 lowercase hex characters. Port ranges are inclusive, sorted by protocol then first port, and disjoint within each protocol. bindRuntimeNetworkHandoffLease verifies both digests, scope and transit-pool containment and returns detached values. It does not authenticate a projection or attest to an actual namespace/interface.

snapshotRuntimeNetworkHandoffLedger verifies at most 256 retained leases for one node/controller epoch, including quarantined allocations. It rejects host handoff-port collisions even across router/runtime namespace changes, and zone or label reuse within one router incarnation. The caller must transactionally supply the complete retained set, serialize allocation and fence controller-epoch changes. These are collision checks, not an IPAM allocator or a native capability. No expiry, empty conntrack dump, process exit or table removal establishes flow drainage or permits handoff identity reuse. Signed complete projections, live interface proofs and the separate router/host apply journal remain responsibilities of Cloudly and Pallet; neither packet admission nor readiness follows from these helpers.

IRuntimeNetworkWorkloadLease is the immutable material referenced by assignment.workload.network. It binds one execution attempt and router incarnation to a workload pool, dedicated unbridged linkSubnet, workload address, exact /32 sourcePrefix, and router-side gateway. The gateway also supplies dnsServer on port 53, because CRI DNS settings cannot encode a port. Canonical private point-to-point subnets through /31 are accepted; the owning CNI implementation must qualify its chosen layout. The source prefix, link subnet, VPN control prefix and host/router transit subnet have different roles. bindRuntimeNetworkWorkloadLeaseToAssignment verifies immutable assignment identity without a circular assignment digest. Membership, aliases, readiness and projection revisions never change this address material.

IRuntimeNetworkProjection contains complete relevant network definitions, service memberships and endpoint leases, including remote ready endpoints and declared services with no replicas. It carries the current protected authority, exact historical allocation authorities, local handoff, a maximum 15-minute DNS window and explicit packet/lease/handoff withdrawals. All arrays use the documented canonical ordering: networks and endpoints by ID, memberships by canonical [organizationId, serviceId], references by canonical complete reference, and withdrawals by canonical complete grant. Projections are bounded to 896 KiB, 128 networks, 256 memberships/endpoints, 16 historical authorities and 4,096 directed grants. These transport bounds do not assert native capacity; the node must preflight the actual SmartVPN, Smartnftables and DNS limits before effectful application.

validateRuntimeNetworkProjection checks complete digest and scope bindings, pool containment, disjoint link subnets, network/organization keys, aliases, resolver authorization and ready-replica uniqueness. It does not authenticate controller inventory, prove observation provenance or authorize allocation reuse. Cloudly must derive readiness from the current authenticated assignment report, retain all affected offline authorities and complete its allocation/barrier transactions before issuance.

An endpoint may publish node ports through the optional IRuntimeNetworkEndpoint.publishedPorts. Each IRuntimeNetworkPublishedPort binds one hostPort (1..65535, unique per node and protocol across the projection) to the workload's targetPort, optionally on an explicit uplink hostIp (never the wildcard or loopback; binding it to the node's uplink is Pallet's enforcement, because the projection carries no uplink fact), and carries the Cloudly authorization reference that permitted it: a published port without that reference cannot exist, and the projection signature covers every byte of it. An endpoint that publishes nothing omits the member entirely and keeps its exact canonical bytes and digest; a present member is never an empty list, so "publishes nothing" has exactly one encoding.

Operators author that authorization through getServicePublishedPorts and setServicePublishedPorts, modelled on the private-network membership pair. IServicePublishedPorts returns the current authorized list with its policy revision, or null when a service never had a policy document. ISetServicePublishedPorts returns the accepted IServicePublishedPorts document with its new revision and a replayed flag, and carries expectedRevision (null for the first document), a retained mutationId for replay, and IRuntimeNetworkPublishedPortRequest entries of protocol, hostPort and targetPort, optionally hostPortEnd and symmetric: the request accepts no hostIp (an omitted address means the node uplink). An empty list is a valid policy meaning the service publishes nothing. Callers never supply an authorization reference: Cloudly derives each projection entry's reference from the accepted policy revision and digest, and node-level host-port placement stays Cloudly's decision.

Who authors a service's published ports follows the service. A service of an idp.global organization is that organization's: only a federated sign-in holding its freshly introspected owner role reads and writes them, and an infrastructure administrator does not substitute for it. A platform-owned service (data.IServicePlatformOwnership: organizationId, serviceId, the designating administrator's actorId, designatedAt) belongs to no idp.global organization, such as the cluster ingress or a platform call router; only a verified administrator reads and writes its published ports, as only an administrator writes its egress. An administrator designates such a service with designatePlatformOwnedService, withdraws it with undesignatePlatformOwnedService and reads it with getServicePlatformOwnership (requests.network.IReq_Admin_Cloudly_DesignatePlatformOwnedService, IReq_Admin_Cloudly_UndesignatePlatformOwnedService, IReq_Admin_Cloudly_GetServicePlatformOwnership). Only a platform organization's service is designated: data.isPlatformOrganizationId is true exactly for a canonical identifier that carries : (legacy:Service:<id>, Organization:<id>), and false for an idp.global nanoid, whichever character it starts with, and for any value that is not a canonical identifier, such as an empty or whitespace id. Cloudly refuses every other id as organization-not-platform, and further designation refusals are service-unknown, service-deleting and conflict (data.servicePlatformOwnershipRefusals). The designation grants no port, egress or placement by itself. Withdrawing it leaves the service's published ports as they are with no writer, because a platform organization has no idp.global owner, until an administrator designates the service again. A published-port write staged while the designation changes is refused as the policy write's conflict.

A published entry with hostPortEnd publishes every port from hostPort to hostPortEnd inclusive, each to the same port inside the workload, so its targetPort equals hostPort and a one-port range is written without the member. symmetric: true makes outbound flows the workload originates from a published target port leave the node from the matching host port, so a peer sees one address and port in both directions (SIP, RTP); it grants no egress of its own, and a projection endpoint carries a symmetric entry only while its publicEgress is true. Both members are stated only when present, so an entry without them keeps its bytes. A symmetric entry owns its inside ports: no other entry of the list and protocol, symmetric or not, publishes to its target port or any port of its range (findRuntimeNetworkPublishedPortInsideOverlap), which Cloudly refuses as published-port-symmetric-inside-overlap and the snapshots refuse outright. The entries of one list are sorted, never overlap within a protocol, number at most runtimeNetworkPublishedPortContract.maximumEntries (64) and cover at most its maximumPortsPerService (1024) ports; on one node, no two entries of a protocol share a port. The port space is split once for Cloudly and Pallet: handoff leases translate outbound flows to runtimeNetworkHandoffSnatPortRange (49152-65535), and Cloudly publishes only runtimeNetworkPublishedPortContract.publishablePorts (1-49151), refusing any other policy as published-port-reserved (isRuntimeNetworkPublishedPortReserved) and a symmetric entry without public egress as published-port-symmetric-without-egress. The snapshots do not enforce the band, so leases and policies written before it stay readable.

A node reads a projection as a whole, so one built against an interfaces release before these members refuses any projection that carries them, including entries of endpoints on other nodes. The session offer (IProtocolOffer.interfacesVersion) is the only statement a node makes about the contract it reads; there is no separate capability list, so a controller must not issue a ranged or symmetric entry into a projection for a node whose offer names an earlier release.

Every endpoint's publicEgress and platformEndpointIds come from the same service policy document. A verified administrator authors them through getServiceNetworkEgress and setServiceNetworkEgress (requests.network.IReq_Admin_Cloudly_GetServiceNetworkEgress / IReq_Admin_Cloudly_SetServiceNetworkEgress): platform endpoints are platform authority, and only an administrator writes the runtime spec whose attached network has to state the same values. IServiceNetworkEgress returns publicEgress and the selected platformEndpointIds (sorted, unique, at most runtimeNetworkAuthorityContract.maximumPlatformEndpoints) with the document's revision, or null when a service never had a policy document. ISetServiceNetworkEgress carries expectedRevision (null for the first document), a retained mutationId for replay and the exact desired values, and returns the accepted IServiceNetworkEgress with a replayed flag; Cloudly carries the published ports forward unchanged, and refuses an id that names no current platform endpoint. Published ports and egress are two halves of one document with one revision, so expectedRevision of either request is the revision the caller read from either answer, and an accepted write of either advances it.

One document per service holds its published ports and its egress under one revision. The first accepted write of either (expectedRevision: null) creates it: setServicePublishedPorts with egress publicEgress: false and platformEndpointIds: [], setServiceNetworkEgress with no published ports. From then on both reads return it, never null. It is an explicit policy: an attached spec stating publicEgress: false and platformEndpointIds: [] matches it and reaches no public address and no platform endpoint; network-policy-missing means the service has no document at all.

snapshotServiceNetworkEgress and snapshotSetServiceNetworkEgress check the shapes. An attached runtime spec is admitted only while it states the document's values (network-policy-mismatch), and nothing else grants a workload egress: a Corestore endpoint on the node's own host is delivered locally, but only to a workload that selects it.

getRuntimeNetworkProjectionPacketGrants derives directed member connectivity and explicit local public/platform egress independently of DNS readiness. composeRuntimeNetworkProjectionDnsViews builds one union view per local workload source address, with every authorized FQDN and complete ready-replica A set. An empty array declares NODATA; absent/unauthorized names remain NXDOMAIN, and IPv4-only service names have AAAA NODATA. The view carries the existing forwarding intersection and zero or one search suffix. Pallet must prove the actual sandbox/veth/source binding, block private/bare-name forwarding, clip TTLs and convert the effective window from getRuntimeNetworkEffectiveDnsWindow (see DNS Lease Renewals) using its qualified boot/time authority. These helpers neither cache queries nor renew a lease.

The DNS composer accepts an optional second argument of exact local workload lease references. Selection happens before view expansion while the complete validated projection still supplies every eligible local or remote target and declared empty name. Omit it for all local views or pass [] for none. Unknown, remote, stale and duplicate references reject. Selection is captured before asynchronous validation and output retains projection order. This lets Pallet compose only its currently attached sources without duplicating membership rules; the references themselves do not prove native attachments or DNS authority.

The optional IRuntimeNetworkProjection.router member carries explicit router-origin egress selections and withdrawals. A projection that omits it retains its exact canonical bytes, digest and signature and grants no router-origin flows. New issuers can select resolverIds and platformEndpointIds from the current protected inventory; an empty selection denies all router egress. DNS forwarding policy, workload egress and inventory presence alone supply no router grant. The whole member is signed with the projection; it is not an unsigned runtime option.

getRuntimeNetworkProjectionRouterPacketGrants validates the complete projection and returns up to 96 detached exact grants. Each binds the current local handoff reference and its router transit source address to a selected destination ID, IPv4 address, protocol and port. Resolver selection explicitly authorizes both TCP and UDP at the declared resolver port. Platform selection authorizes only its declared transport tuple, for example a VPN relay endpoint. Every selected protocol must have a source-port allocation in that handoff. No public wildcard, workload source or alternate source address can be supplied in this selection.

Router withdrawals retain these complete grant bodies, including the old handoff and source address. Removing a selection, changing an inventory endpoint or retiring the handoff must explicitly withdraw the old grants. Unknown denials and denials overlapping current grants reject. Up to 4,096 retained withdrawals carry forward until the caller supplies its exact joined application of the preceding projection, with the same receipt boundary as workload withdrawals. These values do not prove packet drainage, native enforcement or allocation reuse. Consumers must compose workload, router and host grants and their separate withdrawals; they must also fence the actual native source, routing and process lifetimes. DNS expiry or readiness changes do not revoke packet authority.

A projection whose controller kind is onebox names a controller that runs on the node itself, so it has no other node: it refuses every endpoint whose lease is not placed on the projection's own node and runtime namespace. Projections of a cloudly controller keep their remote endpoints.

The optional IRuntimeNetworkProjection.host member exists only on onebox projections; any other controller kind carrying it rejects. It lets the node's own host namespace, for example Onebox's reverse proxy, dial exact local workload ports directly, with no loopback publication, no loopback DNAT and no change to martian filtering. A projection that omits it retains its exact canonical bytes, digest and signature and grants no host-origin flows. Each ingress entry names the exact lease reference of a local endpoint of the same projection, a protocol, the workload's own port and an authorization reference. The authorization is signed with the projection but is not part of the grant identity, so re-authorizing an unchanged port neither creates nor withdraws a grant. Entries are unique and strictly ordered; at most 1,024 are allowed (runtimeNetworkHostContract).

getRuntimeNetworkProjectionHostPacketGrants validates the complete projection and returns detached exact grants. Each binds the current handoff reference and its transit host address as the source to the lease reference, lease IPv4 address, protocol and exact port as the destination. Unlike router grants, host grants need no per-protocol source-port allocation in the handoff. Host withdrawals follow the router rules: removing a selection or retiring the handoff or lease must explicitly withdraw the old complete grants, unknown denials and denials overlapping current grants reject, and up to 4,096 retained withdrawals carry forward until the caller supplies its exact joined application of the preceding projection.

The optional IRuntimeNetworkProjection.workloadIngress member exists only on cloudly projections; any other controller kind carrying it rejects. It is how the cluster ingress workload reaches the workloads it routes to without a published host port: each grants entry names the exact lease reference of the ingress workload (source), the exact lease reference of the target workload (destination), a protocol, the target's own listening port and an authorization reference. Unlike the full-mesh workload grants, it needs neither the same organization nor a shared private network, it opens exactly one port, and it is one-way: nothing flows back except the replies of a connection the ingress opened. Both leases must be endpoints of the projection, they must differ, and at least one of them must be placed on the projection's own node. Entries are unique and strictly ordered by [source.id, destination.id, protocol, port]; at most 1,024 are allowed (runtimeNetworkWorkloadIngressContract). The authorization is signed content, not grant identity. A projection that omits the member retains its exact canonical bytes, digest and signature. getRuntimeNetworkProjectionWorkloadIngressPacketGrants validates the complete projection and returns detached IRuntimeNetworkWorkloadIngressPacketGrants ({ source, destination: { lease, protocol, port } }). Their withdrawals follow the host rules: an omitted grant needs its exact withdrawal, unknown or current denials reject, and up to 4,096 retained withdrawals carry forward until the exact joined application of the preceding projection is supplied (Projection Application Receipts).

A signed grant is not yet a delivered flow. The receiving node enforces a grant only between two workloads attached behind its own router: the packet engine has no one-way grant across the inter-node tunnel, so a grant whose ingress lease and target lease run on different nodes is signed into both nodes' projections and admitted by both, yet neither builds a packet path for it and the flow stays denied. Until cross-node one-way grants exist, a controller may route a cluster ingress route to the target lease address and port only when the grant exists and both leases are placed on the same node.

The optional IRuntimeNetworkProjection.hostPlatformEndpointIds lists the ids of the protected authority's platform endpoints whose address the projection's node's own host carries, such as a relay listener or a Corestore beside the node on one machine. A node delivers workload flows to those endpoints locally instead of forwarding them. The list is sorted, unique, never empty when present, and every id must be a current platform endpoint; omitting it keeps the canonical bytes and states that no endpoint is local. It grants nothing: a workload still reaches only the endpoints it selects. The platform endpoint contract itself ({ id, address, protocol, port }) names no node, which is why this statement is per projection.

Projection Retirement Requests

The optional IRuntimeNetworkProjection.retirement (IRuntimeNetworkProjectionRetirement { disposition: 'withdrawn', protectedAuthority }) is Cloudly's request that a running node withdraw its allocation protection, because its egress ownership is being retired with the withdrawn disposition. It travels in the projection because that is the one authenticated, monotonic channel the node already admits, journals and reads on every pass, and whose acknowledged head Cloudly already tracks: the controller's signature covers it, so a node cannot invent it; the projection chain is bound to the controller epoch, node and runtime namespace, so it cannot be carried elsewhere; and a node's acknowledged head proves the node admitted it. A retired node never returns (runtimeNetworkEgressRetirementContract.reentry), so no later chain of the same node can meet a stale request.

protectedAuthority is the authority current when Cloudly first signed the request; its id names the receipt chain's authority chain. validateRuntimeNetworkProjection accepts it only on a cloudly projection and only naming the projection's own authority chain at a generation the projection has reached (the current authority itself at the same generation). admitRuntimeNetworkProjection latches it: it may first appear on any projection, including the first, and that first appearance names exactly the projection's current protectedAuthority (an older generation is refused, because the node's positive receipt under the current authority would then be newer than the request and the withdrawal could never complete). Afterwards every successor carries it unchanged, never dropped and never rewritten, so it stays fixed while the authority chain moves on. Omitted, it keeps the canonical bytes and asks for nothing; it grants and withdraws no packet flow. Consumers on interfaces earlier than 32.44.0 read projections with exact keys and refuse one that carries retirement, so Cloudly sends the request only to nodes that report a compatible interfaces version. snapshotRuntimeNetworkProjectionRetirement is the exact-key reader and runtimeNetworkProjectionRetirementContract states the rules.

The node answers with the null-barrier successor of its protection receipt chain and mints no positive receipt for that chain again. Its allocation-pool guard stays enforced: the answer withdraws eligibility, not protection.

bindRuntimeNetworkProtectionReceiptRetirement(receipt, previous, retirement) is the receiver's gate for every report from a node whose projections carry the request, beside admitRuntimeNetworkProtectionReceipt with the same previous (the receiver's current receipt for the node, or null). The receipt must belong to the retiring authority chain and previous, when given, must be the same chain's exact predecessor or the receipt itself. The null answer always passes. A positive receipt passes only while it can predate the node's admission of the request, so a receipt the node minted before it saw the request still drains from its outbox and the withdrawal behind it is never wedged: it names no authority newer than the request's (a newer one could only come from a projection already carrying the request), and its predecessor is not a null withdrawal (once the receiver holds the answer, no positive receipt follows). A positive receipt minted after the request against the same authority cannot be told apart from one minted before; the node's obligation is never to mint one.

admitRuntimeNetworkProjection checks exact predecessor/replay identity and rejects changed same-generation content, regressed membership/readiness evidence, changed immutable address/DNS keys, and omitted packet grants without explicit withdrawal. Pending withdrawals/tombstones must carry forward until the caller supplies the exact previous projection reference from its own durable joined application receipt (Projection Application Receipts). That reference is never taken from an incoming ACK. Keep the accepted history and apply journal until native effects and DNS/allocation quarantine are settled; a complete new snapshot does not erase that history.

IRuntimeNetworkSigningAuthority is separate from human JWT signing. Install its current revision through the existing authenticated physical connection and commit it with a durable CAS fence. Newly enrolled nodes may bootstrap the controller's current key generation; existing trust cannot be reset to do so. Exact successor revisions rotate the Ed25519 public key; publicKey: null revokes it. Rotation/revocation stops acceptance under the previous key, and requires newly signed authority for DNS. It does not deny outstanding packets or permit allocation reuse.

The Node-only @serve.zone/interfaces/runtime export supplies createRuntimeNetworkSigningKey, signRuntimeNetworkProjection and verifyRuntimeNetworkProjection. Signing produces ISignedRuntimeNetworkProjection { authority, projection, signature }, and the signature covers the canonical envelope { domain, authority, projection } under the dedicated serve.zone/runtime-network-projection-signature domain, so it binds the complete projection and the exact signing-authority revision. Verification requires receiver-owned trusted authority and node scope; an envelope contains no key to trust. Private KeyObjects stay process-local; Cloudly persists their exported material only through its encrypted internal secret store. The dedicated requests.runtimesession key/projection push contracts bind the current physical peer. Their response acknowledges durable admission only, never application, readiness, independent fencing or pool activation.

The getRuntimeManagedVpnCredential method contract in requests.runtimesession uses IGetRuntimeManagedVpnCredentialRequest and IGetRuntimeManagedVpnCredentialResponse. Its request carries the exact current session and signed projection reference. The secret response repeats those bindings and supplies the hub tuple the node dials, the transport it dials it over (quic, the one transport a managed hub serves), the managed authority ID, the Noise server public key, the client keypair and the expiry in Unix milliseconds. The response states no TLS name: QUIC is dialled on the address alone. The hub is the node's own cluster's relay (see Cluster VPN Hub And Network). Keep the response process-local; never persist, hash, log or embed it in a signed projection.

snapshotGetRuntimeManagedVpnCredentialRequest and its response counterpart capture closed detached JSON shapes and canonical 32-byte key encodings. bindGetRuntimeManagedVpnCredentialRequest(request, projection, trustedSession) validates the complete projection digest and binds node, controller, namespace, physical-session identity and projection reference. The response binder takes (response, request, projection, trustedSession, now, renewal = null) and additionally requires the hub's exact tuple to be selected in router.platformEndpointIds, carrying the protocol of the credential's transport — udp for quic, as runtimeNetworkAddressPlanContract.hubTransports states. A relay therefore cannot name itself: the endpoint has to be one the node's signed projection already selected. Expiry must be after the caller-qualified time and no later than the referenced projection's effective DNS window (see DNS Lease Renewals); acceptance before that window starts rejects.

A credential names one projection, but its key pair belongs to the node's managed VPN membership, which outlives it. Every workload add, removal, readiness change or policy change signs a new projection generation; none of them changes who the member is. continuesRuntimeManagedVpnMembership(previous, next) is that rule: next continues the membership when it is the same projection lineage (id), under a protected authority that continues the previous one (continuesRuntimeNetworkProtection: the same authority or its next step that only adds), with the same handoff lease (which names the node scope and carries the router incarnation), and is previous itself or a later generation. A node joining the egress or a cluster hub joining the platform endpoints therefore keeps every member's key pair and hub session. A protection step that removes or changes anything or skips a step, a fresh handoff (router reincarnation) or a new controller epoch is a new membership; an earlier generation or another projection of the same generation continues nothing. It validates both projections and rejects one that does not validate.

bindGetRuntimeManagedVpnCredentialSuccessor(response, previous, previousProjection, projection, trustedSession, now, renewal = null) binds the credential for the successor. response is bound exactly as the response binder binds an answer to { session: trustedSession, projection }, so the hub must still be selected and the expiry stays inside the successor's effective DNS window. It must also continue previous, the credential issued for previousProjection over the same session: the two projections are one membership and the transport, hub, authority, server key and client key pair are unchanged. Only projection and expiresAt move. The current effective window is the one limit: a membership whose current window does not cover now binds no successor, while the held credential's own expiry is not consulted, because a renewal or a successor may already have moved the window past it. The issuer binds the successor it states for a member with it, so the node keeps its key pair and its hub session across a projection change; a node binds a fresh answer with it to tell a successor, which keeps the live session, from a new membership, which replaces it. A previous credential for the same projection binds as its own successor, which is how a DNS lease renewal moves the expiry.

These helpers perform content validation. Callers authenticate and fence the current physical peer and projection before and after native work, qualify time independently, and retain ownership through revocation and shutdown. The native Noise implementation owns cryptographic validation and authentication. Credential delivery does not prove a connected VPN, TUN ownership, packet enforcement or workload readiness, and does not replace explicit packet withdrawals.

Projection Application Receipts

IRuntimeNetworkProjectionApplicationReceipt is a node's durable statement that its packet authority is bounded by one admitted projection. Its claim is a ceiling: every packet table the node holds was composed from projection, or the node holds no packet table at all, so the flows its kernel admits are a subset of that projection's grants. Only the node's joined native packet owner can supply it: the journaled application of the projection's policies on every role, or the release of every table (nativeJournal: null). An admission ACK, a single table inspection or a success flag is never such evidence (runtimeNetworkProjectionApplicationContract). A crash, a reboot or a restored guard only removes authority, so a receipt stays true until the node starts applying a successor; from then on only the successor's receipt counts, because only the exact predecessor of an admitted projection is ever selected.

A node's receipts form one monotonic chain bound to its node scope and reporter session; bootId records the boot that minted a receipt and is not compared. admitRuntimeNetworkProjectionApplicationReceipt(receipt, previousReceipt, scope, head) is the receiver's gate: the first receipt is generation 1, each successor names its exact predecessor, an identical receipt is a replay, and the named projection never goes back. The receipt must name the head's projection chain at a generation the head has reached: the head exactly, or an older generation. An older one was minted before the node admitted the head; it advances the chain and is never usable.

selectRuntimeNetworkProjectionAppliedPrevious(receipt, previous) returns the appliedPrevious argument of admitRuntimeNetworkProjection for the successor of previous: the receipt's projection reference when it is exactly previous, otherwise null. A controller passes the latest receipt it admitted for the node; the node passes its own durable receipt to its own admission. Only the exact predecessor counts. A receipt for an older projection proves nothing about the grants issued since, so it releases nothing.

This is why withdrawals now drop. A projection carries its predecessor's withdrawals, and may not grant a withdrawn flow again, until appliedPrevious names that predecessor. Once a receipt proves the node applied it, the node's packet authority is within the predecessor's grants, so the successor states only the withdrawals of grants the predecessor held and it drops. A flow withdrawn earlier can then be granted again. Without receipts, withdrawals only accumulate up to their bound, and a withdrawn flow stays withdrawn until a lease changes.

reportRuntimeNetworkProjectionApplication (IReq_Pallet_Controller_ReportRuntimeNetworkProjectionApplication) carries a receipt on the current authenticated physical session. bindReportRuntimeNetworkProjectionApplicationRequest binds it to the current session and the receiver-owned reporter history, as protection reports are bound: a receipt minted under an earlier session of the same node still drains from its outbox as historical. The response must match the exact sent receipt (bindReportRuntimeNetworkProjectionApplicationResponse). Trust rests on the authenticated session; the contract has no node signing key. A node's receipt only loosens that node's own enforcement, so a node that misreports it gains no authority over any other node. Each node's receipt chain is independent: a flow granted in the projections of two nodes is released on each by that node's own receipt. A receipt is not an admission acknowledgement, and an acknowledgement is never a receipt.

Egress Owner Retirement

An owner that is offline, revoked or gone can never acknowledge a successor, so while it stays declared no allocation is possible after the next protection change. IRuntimeNetworkEgressRetirement is the independent fence that lets it leave. Its owner is the controller, which records it from its own durable state against the exact protected authority the owner is about to leave, once all five hold:

  1. the node carries an INodeRuntimeRetirement row, whose disposition the fence copies;
  2. both node credentials are revoked (revokedCredentials, committed at revokedAt), so the node can no longer register, renew a DNS lease or fetch a managed VPN credential;
  3. every DNS and managed VPN window the node was ever issued has ended (packetWindowExpiresAt, or revokedAt when it was issued none), so no hub relays for it even if it was never told to stop, and fencedAt is not earlier;
  4. every handoff and workload lease the node was granted is quarantined (quarantinedHandoffs, at most 256, and quarantinedWorkloads, at most 512, both in canonical order) and is never handed out again;
  5. every remaining owner of that authority acknowledged a projection head that carries no endpoint of the node (clearedProjections, in canonical owner order, never the node itself). A remaining owner whose own lost fence is recorded against the same authority is left out: it never acknowledges anything again, so waiting for it would freeze every other owner's retirement, and it leaves behind its own fence. Every other remaining owner has exactly one entry.

lastReceipt is the node's last accepted protection receipt. A withdrawn node took part in its own retirement, so it always has one (its null withdrawal, the answer to its projection's retirement request); a lost node may have none. Neither disposition shortens the fence: a lost node leaves behind exactly the same controller-owned evidence as a withdrawn one. snapshotRuntimeNetworkEgressRetirement is the exact-key reader and computeRuntimeNetworkEgressRetirementDigest seals it under serve.zone/runtime-network-egress-retirement. Neither authenticates the evidence; the controller fences the rows it read together with the authority it stages.

bindRuntimeNetworkProtectedAuthorityRetirement(previous, next, fences, retiredOwners, exemptions = []) is the pure successor rule for the egress owners and the address space of any step, and answers { previous, next, fences, exemptions }. next must be the exact successor of previous under the same controller, with both digests verified. An owner of previous may be absent from next only when exactly one fence recorded against the exact previous reference, verified by its digest, names it; a fence for an owner next still names, or one previous never named, refuses the step.

exemptions are the lost fences of owners the step keeps, recorded against the same previous. They are read and returned, not consumed: each verifies by its digest, is recorded under the same controller against the exact previous reference, and names an owner both previous and next name that no fence and no other exemption names, in canonical owner order and at most 255 per step. A withdrawn or stale exemption refuses the step. Every consumed fence's clearedProjections must name exactly the owners the step keeps less the exempt owners, in both directions: a clearance that names an exempt owner or misses a kept one refuses the step. An exemption's own clearance is judged only in the step that consumes it as a fence. Lost owners can therefore be fenced in parallel against one authority and leave 16 per step: each step consumes up to 16 of their fences and holds the rest as exemptions, and every later step re-records the remaining fences against its own previous.

retiredOwners is the caller's durable history of fenced owners restricted to the owners previous or next names, in canonical order and at most 512 entries. Any entry refuses the step, so a fenced owner never returns, and the caller looks up only the owners of the two authorities instead of supplying a lifetime list. The caller must look up every owner that previous and next name: the binder cannot tell a missed lookup from an empty history, so an omitted entry silently drops the guarantee. Every pool of previous stays unchanged and the prefixes of next cover those of previous; a step that removes an owner also keeps every resolver. Owners may join in the same step, and a step without fences is the additive rule, which exemptions do not change. One step consumes at most 16 fences (runtimeNetworkEgressRetirementContract).

A retirement step is a new protected authority: every remaining owner reports a receipt citing it before anything new is allocated, while existing leases keep their historical eligibility. It also ends every managed VPN membership, as every protection change does, so each retirement costs the fleet one reconnect.

DNS Lease Renewals

A projection's dns window lasts at most fifteen minutes and is signed content, so a node's private DNS stops when the window ends unless a newer signed statement extends it. Re-issuing the projection would extend it too, but a node applies every new projection as new network state, which restarts its data plane. IRuntimeNetworkDnsLeaseRenewal extends the window and nothing else: it carries the node scope (controller, nodeId, runtimeNamespace), the exact projection reference it renews, a sequence, a fresh dns { issuedAt, notBefore, expiresAt } window and its digest. It names the projection by reference, so it never joins the projection chain and never changes what the node applied. The new shapes carry no schemaVersion and no version suffix; the installed package version is the protocol version.

issuedAt is the start of the window's slot, not the moment of signing: Cloudly signs renewals ahead of time on a grid, so the fifteen-minute rule (maximumDnsLeaseMs, the same limit a projection's window has) is measured from the slot. computeRuntimeNetworkDnsLeaseRenewalDigest hashes every field except digest under its own domain, and validateRuntimeNetworkDnsLeaseRenewal checks structure and digest only; a digest authenticates nothing.

The Node-only @serve.zone/interfaces/runtime export signs and verifies renewals with the same authority and key that sign the controller's projections, under a separate signature domain. signRuntimeNetworkDnsLeaseRenewal(renewal, authority, privateKey) produces ISignedRuntimeNetworkDnsLeaseRenewal { authority, renewal, signature }. verifyRuntimeNetworkDnsLeaseRenewal(signed, trustedAuthority, trustedScope, projectionAuthority) takes the authority and scope from the node's own stored trust, never from the envelope. projectionAuthority is the reference of the authority revision that signed the projection the node admitted, and a renewal verifies only when it was signed under exactly that revision. Rotation and revocation end renewals exactly as they end projections: nothing signed under an earlier revision verifies after, and a renewal signed under a rotated revision does not renew a projection signed under the one before. A renewal never closes a key gap; after rotation, DNS continues only with a newly signed projection.

admitRuntimeNetworkDnsLeaseRenewal(renewal, currentProjection, previousRenewal | null, trustedScope, now) is the pure admission decision after verification. The projection and the previous renewal are the node's own admitted state, and now is its caller-qualified clock, read before the first await like every other input. It admits a renewal only when:

  • the renewal and the projection are in the trusted scope, and the renewal names the current projection's exact reference. Once a successor projection is admitted, every renewal of its predecessor is refused, and a renewal never introduces a projection the node has not admitted;
  • its sequence is higher than the previous renewal's, with gaps allowed, or equal with byte-identical content, which is answered as replay whatever now is, so re-delivery is idempotent. The same sequence with any other content is refused;
  • its window is current: notBefore <= now < expiresAt. Admitting a window that has not opened would displace the coverage the node has now, so a sender delivers a window only once its slot has opened; a renewal refused for arriving early is simply delivered again once it has. An expired window is refused as well;
  • its window does not move issuedAt back and ends later than the previous window. Without a previous renewal, the projection's own window is the baseline. Windows need not touch, so a node whose lease lapsed can still admit the next renewal it is given.

The caller stores the admitted renewal in the same transaction as its trust fence, and deletes it in the transaction that admits a successor projection. Admission is not conversion: the node still turns the window into a deadline with its qualified clock, and must not use a window before its notBefore.

requests.runtimesession.IReq_Controller_Pallet_ApplyRuntimeNetworkDnsLeaseRenewal (applyRuntimeNetworkDnsLeaseRenewal) delivers one renewal as IApplyRuntimeNetworkDnsLeaseRenewalRequest { session, signed }. bindApplyRuntimeNetworkDnsLeaseRenewalRequest(request, trustedSession) requires the node's current physical-session binding and a renewal in that session's scope. It does not care who sent the request: a cluster relay pushes under the binding the node obtained itself. The answer, IApplyRuntimeNetworkDnsLeaseRenewalResponse { status, renewal }, repeats IRuntimeNetworkDnsLeaseRenewalReference { projection, sequence, digest }. bindApplyRuntimeNetworkDnsLeaseRenewalResponse(response, sentRenewal) compares it with the exact renewal that was sent. The ACK asserts durable admission only, not that the node serves DNS.

A cluster is outbound-only, so while its relay cannot reach Cloudly nobody can sign a new window. Cloudly therefore hands the elected relay a run of renewals it signed ahead of time, with requests.cluster.IReq_Cloudly_Relay_HoldRuntimeNetworkDnsLeaseRenewals (holdRuntimeNetworkDnsLeaseRenewals): IHoldRuntimeNetworkDnsLeaseRenewalsRequest { nodeId, projection, renewals } answered by { held }. Each request replaces what the relay held for that node, and an empty run holds nothing. The relay keeps the run in memory only, pushes only the renewal whose window is current and only while it has no Cloudly session (escrowRelease), and holds no key and no trust. snapshotHoldRuntimeNetworkDnsLeaseRenewalsRequest checks the rules a relay can evaluate without the projection itself; the node still measures the first window against its admitted projection, and every window against its clock, when it admits it. The run must have:

  • one node, one projection and one signing authority revision, in one scope;
  • strictly ascending sequences, each window extending the one before;
  • at most maximumEscrowedRenewals (32) renewals;
  • at most maximumEscrowHorizonMs (four hours) from the first window's issuedAt to the last window's expiresAt. On the intended grid of fifteen-minute windows every ten minutes, the horizon allows 23 windows, so it is the limit that binds.

validateHoldRuntimeNetworkDnsLeaseRenewalsRequest adds every renewal's digest check and verifies no signature. bindHoldRuntimeNetworkDnsLeaseRenewalsResponse(response, sentRequest) requires held to equal the number of renewals sent.

A hold is the one call whose outcome the controller cannot read back from its own records, so the refusal is the whole answer. data.holdRuntimeNetworkDnsLeaseRenewalsRefusals is the frozen list of names a relay leads its refusal with — escrow-run-invalid, escrow-node-not-carried, escrow-run-foreign-session, escrow-run-stale, escrow-run-unordered and escrow-cloudly-only — with data.THoldRuntimeNetworkDnsLeaseRenewalsRefusal as their union. The relay writes <name>: <reason> and the controller matches the name, so neither side keeps its own copy of the vocabulary.

The horizon is also the revocation latency: a detached cluster keeps DNS authority until its last held window ends. So a controller must not release an address or a DNS name before the latest expiresAt it has signed for that node. Renewals extend DNS only. They grant no packet authority and cannot introduce new network state, and they do not extend an already-issued VPN credential; a newly issued one may be bound to the renewed window.

getRuntimeNetworkEffectiveDnsWindow(projection, renewal | null) is the one rule for the window a node may use: the admitted renewal's if there is one, else the projection's own. renewal is the bare IRuntimeNetworkDnsLeaseRenewal that admission takes and returns, which is what a node stores; Cloudly passes signed.renewal. It checks digests and that the renewal names exactly that projection, but verifies no signature. Controller and node use it for every limit tied to the window, and bindGetRuntimeManagedVpnCredentialResponse takes the same bare admitted renewal as an optional last argument, so a newly issued managed-VPN credential may run to the end of the renewed window instead of the projection's original one.

import { data } from '@serve.zone/interfaces';
import { signRuntimeNetworkDnsLeaseRenewal, verifyRuntimeNetworkDnsLeaseRenewal } from '@serve.zone/interfaces/runtime';

const renewal: data.IRuntimeNetworkDnsLeaseRenewal = {
  ...scope,
  projection: { id: projection.id, generation: projection.generation, digest: projection.digest },
  sequence: 1,
  dns: { issuedAt: slotStart, notBefore: slotStart, expiresAt: slotStart + 900_000 },
  digest: `sha256:${'0'.repeat(64)}` as data.TSha256Digest, // well-formed placeholder, replaced below
};
renewal.digest = await data.computeRuntimeNetworkDnsLeaseRenewalDigest(renewal);
const signed = await signRuntimeNetworkDnsLeaseRenewal(renewal, signingAuthority, privateKey);

// On the node: stored trust, then the node's own admitted state and qualified clock.
const verified = await verifyRuntimeNetworkDnsLeaseRenewal(
  signed, storedAuthority, nodeScope, admittedProjectionAuthority);
const { renewal: admittedRenewal } = await data.admitRuntimeNetworkDnsLeaseRenewal(
  verified.renewal, admittedProjection, previousRenewal, nodeScope, qualifiedNow);
const window = await data.getRuntimeNetworkEffectiveDnsWindow(admittedProjection, admittedRenewal);

Runtime Network Address Plan

data.IRuntimeNetworkAddressPlan is the address space an administrator gives the runtime network, and the single input Cloudly composes the protected authority from: { id: 'address-plan', revision, prefixes, pools, resolvers, platformEndpoints, vpn: { controlPrefix, hubPrefix } }. Pools, resolvers and platform endpoints use the protected authority's own element shapes. controlPrefix is the private prefix VPN control addresses are allocated from, and hubPrefix is the protected prefix every cluster's VPN hub binds inside. The hub endpoints themselves are no administrator's input: each cluster's relay reports what it bound, and the composition turns that into platform endpoints, so platformEndpoints never lists one.

snapshotRuntimeNetworkAddressPlan refuses any plan that could describe an authority the authority contract refuses: it runs snapshotRuntimeNetworkProtectedAuthority over the plan's own sets, so overlapping pools, more than 16 resolvers and a resolver or endpoint outside prefixes or inside a pool are all refused by that one validator. On top it requires both VPN prefixes to sit inside prefixes, to overlap no pool, to hold no resolver or platform endpoint, and never to overlap each other — control addresses live inside the tunnel, hub endpoints on the underlay. controlPrefix is additionally a private prefix no longer than /30 (runtimeNetworkAddressPlanContract.maximumControlPrefixLength); no minimum length is stated, because the private-range rule already keeps it at /8 or longer. data.runtimeNetworkVpnControlPrefix(value) is that reader on its own, exported because a cluster's VPN network repeats the prefix its hub binds and has to read it exactly as the plan does.

Every list of a plan is stated in one canonical order, and a plan is never reordered for its writer: prefixes strictly ascending by the prefix text, and pools, resolvers and platformEndpoints strictly ascending by id. Both compare as strings, so the order is lexicographic, not numeric — 100.64.0.0/10 comes before 20.0.0.0/8, and resolver-10 before resolver-9 — and a value stated twice breaks it as well. A list out of order is the one plan refusal the contract names: snapshotRuntimeNetworkAddressPlan (and admitRuntimeNetworkAddressPlanChange through it) throws data.RuntimeNetworkAddressPlanStructureError, whose violation is { field: 'prefixes' | 'pools' | 'resolvers' | 'platformEndpoints', rule: 'order' } and whose message quotes nothing from the plan. Every other refusal stays the contract's generic rejection.

try {
  data.snapshotRuntimeNetworkAddressPlan(plan);
} catch (error) {
  if (error instanceof data.RuntimeNetworkAddressPlanStructureError) {
    error.violation; // { field: 'prefixes', rule: 'order' }
  }
}

data.IRuntimeNetworkVpnHub { clusterId, address, quicPort } is one cluster's hub as its relay bound it. runtimeNetworkVpnHubEndpoints(hub) is the one derivation of what it contributes: <clusterId>:quic on udp, one endpoint per row of runtimeNetworkAddressPlanContract.hubTransports. That table has one row, because a managed hub terminates TLS on its QUIC listener alone and no other transport of it can be dialled over a protected connection. Every reader calls the derivation — the composition that seals the authority and the router selection that names which endpoint a node may dial — so two readers can never derive two ids for one hub. composeRuntimeNetworkProtectedAuthority(plan, frame, hubs) is the composition itself: the producer supplies the frame (TRuntimeNetworkProtectedAuthorityFrame: authority id, generation, controller, previous reference and egress authorities) and the hubs of the clusters that registered one, and receives the complete authority with its digest, so the validator and the protection producer can never compose differently. Each hub's address must sit inside vpn.hubPrefix and one cluster is named once; an empty hubs composes the authority of a runtime network whose clusters have no hub yet. One hub costs one of the authority's 64 platform endpoints, so 64 clusters fit one authority.

A single host that carries both a cluster's hub and platform endpoints, for example the relay listener and a Corestore, holds more than one address. The hub address sits inside vpn.hubPrefix, as every hub's does. The plan's own platform endpoints may not: they sit on another address of the same host, inside prefixes, outside every pool and outside both VPN prefixes. Several endpoints may share that one address on distinct ports, because the authority requires only each address:port:protocol to be unique. For example, with hubPrefix 192.0.2.0/28, the hub binds 192.0.2.1 and the plan lists the relay listener and Corestore on 192.0.2.17 with two ports. The node's projection then names those endpoints in hostPlatformEndpointIds.

admitRuntimeNetworkAddressPlanChange(previous | null, next) is the pure decision for a plan write at the following revision; any other revision is a contract violation. Leases and the managed VPN outlive a plan revision, so it answers:

  • unchanged when the content equals the stored plan; Cloudly writes nothing;
  • refused with plan-vpn-immutable when either VPN prefix differs, plan-pool-removed when any pool is gone or changed its purpose or prefix, and plan-prefix-shrunk when the new prefixes, however split, no longer cover every address the old ones did;
  • accepted with the ids of the resolvers and platform endpoints the change removes or changes. An entry that keeps its id but differs in any field counts, so moving it goes through the same reference check as removing it.

TRuntimeNetworkAddressPlanRefusal adds plan-resolver-referenced and plan-endpoint-referenced, which only Cloudly can decide against live references, and plan-unsorted, which Cloudly answers for a RuntimeNetworkAddressPlanStructureError with the list as its reference, so every plan refusal has one type. requests.network carries setRuntimeNetworkAddressPlan { identity, expectedRevision, plan } (expectedRevision: null when no plan exists yet) and getRuntimeNetworkAddressPlan, which answers null until the first plan is set.

Node Network Readiness

requests.network.getRuntimeNetworkNodeReadiness { identity, nodeId } answers data.IRuntimeNetworkNodeReadiness: ready: true with nothing missing, or ready: false with at least one TRuntimeNetworkNodeReadinessGap. The first gap is liveness: node-observation-stale states that no accepted node report newer than data.clusterNodeObservationFreshnessMs (180 s, three of Spark's default 60 s reporting intervals) shows the node running, with lastObservedAt the time Cloudly accepted its last report, or null when the node's status is not online. The remaining gaps follow the node producer's order: the node joins the protected authority's egress (egress-pending), every egress owner reports current protection (protection-receipts-pending, naming those owners, at least one), a handoff is reserved (handoff-pending), the router selection carries it (selection-pending), the node's acknowledged projection carries it (projection-pending), and the node is a managed VPN member (vpn-member-pending).

data.IRuntimeRouterIncarnation { id, incarnation } is a node's router incarnation, from which its handoff and router incarnation ids derive; snapshotRuntimeRouterIncarnation checks the record. reincarnateRuntimeNetworkRouter { identity, nodeId, expectedIncarnation } advances it, so the node producer reserves the router a fresh handoff, and answers with the new incarnation.

Node-Bound Runtime Assignments

The data namespace exports IRuntimeAssignment, canonical digest helpers, evaluateRuntimeAssignment, and assignment observation validators/evaluators for Cloudly/Onebox controllers and Pallet nodes. These are shared contracts; they do not install a runtime, schedule containers or expose an RPC endpoint.

Each assignment binds a controller incarnation, organization, stable node and replica IDs, one immutable execution attempt, image/configuration and opaque network/storage/secret references. Group and site IDs are nullable topology provenance. Those workload references are also the authority vocabulary: data.runtimeWorkloadAuthorities is ['network', 'secrets', 'storage'], and data.requiredWorkloadAuthorities(workload) names the ones this workload cannot run without — filtered out of that vocabulary and sorted, so the answer is duplicate-free and compares byte for byte against the resolvedAuthorities a node stated when it registered. data.isResolvedWorkloadAuthorities(value) is the same rule read the other way, for a receiver checking a list it was sent. Canonical is strictly increasing in the code-unit order of the names themselves — a rule a peer in another language implements from the vocabulary alone, rather than from the order this package happens to declare it in — and both directions compare through that one rule, so neither can drift from the other.

workload.registryHost is the registry that published the image, which is the publication identity the image release is bound to, and it is never rewritten. workload.pullEndpoint is the optional host[:port] the node fetches the bytes from instead — a cluster whose relay forwards the registry states the relay's own origin, so a node pulls inside its cluster and nothing cluster-side dials the control plane. data.buildRuntimeAssignmentImageReference(workload) is the one place that choice is made: pullEndpoint ?? registryHost, digest-pinned either way, so the endpoint decides reachability and never content, and a controller and a node cannot disagree about the reference. Absence stays absence: an assignment that states no endpoint is different content, and so a different digest, from one that does.

All fields except the digest itself participate in a domain-separated SHA-256 digest. Digests provide integrity and compare-and-swap identity; authentication must separately establish the trusted controller/node scope.

An attempt starts at generation one with run, then advances through stop-preserve and remove-runtime-preserve-storage using exact predecessor generation/digest references. Exact retries return replay; conflicts, stale revisions, gaps, changed execution specifications and resurrection are rejected. Consumers must atomically persist the evaluator's detached snapshot before effects, retain removal tombstones and storage, and serialize ownership of each replica. Omission causes no change. A new attempt requires separate admission proving any older exclusive writer stopped or fenced. The supported offline policy is continue-node-bound; no heartbeat timeout or lease expiry proves a writer absent.

Observations bind the exact assignment, authenticated reporter session, durable per-assignment sequence and digest. Reconnect does not reset the counter. Observation history must reference the same assignment revision or its exact predecessor; receivers process each disposition's observations in order. Session authorization, clock/freshness policy and durable receipt persistence belong to the receiving service. Readiness requires matching running image evidence and run intent; assignment acceptance alone is never readiness. A valid observation can describe drift (such as a removed runtime while intent remains run), and does not authorize storage relocation or prove fencing.

A failed observation may state why the run failed in its optional cause (TRuntimeAssignmentFailureCause), a closed, value-free union keyed by kind (runtimeAssignmentFailureCauseKinds):

  • image-pull-refused with reason (runtimeImagePullRefusalReasons: unauthorized, not-found, tls, transport, deadline, unavailable, other) and the runtime's gRPC status grpcCode (1..16). Nothing of the run existed, so runtime is null;
  • exited with the container's exitCode (0..255), or null when the runtime stated none or one outside that range;
  • runtime-absent: the runtime holds nothing of a run that should be running;
  • network-epoch-ended: the node stopped a run whose network attachment belongs to an ended network epoch.

The snapshot refuses a cause on any other phase, an unknown kind, a missing or extra field and an out-of-range number, so no runtime message, registry or path can travel in it. The cause is sealed content: a causeless observation keeps exactly the bytes and digest it had before causes existed, and the same sequence stated with another cause is a conflict. A cause explains a failure and changes nothing about it: a controller replaces every failed run the same way. IServiceRuntimeAssignmentView.observation carries the cause, so an administrator reads why an attempt failed. A consumer on interfaces 32.39.0 or older rejects an observation that states a cause, so a node states one only to a controller that reads this contract.

Resolved Runtime Configuration

data.IRuntimeConfig seals the public environment, working directory, resource and identity policies, target ports, readiness timing and log limits together with the immutable image invocation. computeRuntimeConfigDigest hashes that exact content; bindRuntimeConfigToAssignment checks the config reference, service, organization, image and selected platform against one assignment. Secret material, registry credentials and host paths remain outside the configuration.

Both CPU and identity policies are required. cpuMillis: null explicitly means no CPU quota; a positive number sets a quota. runAsUser and runAsGroup must either both be numeric overrides or both be null. The null pair selects the pinned platform image's OCI user and group resolution against its immutable root filesystem. A consumer must not guess a group for a UID-only or named image user. Version 29 widens these numeric TypeScript fields to include those explicit null policies; consumers must handle them before accepting configurations from v29. Existing numeric configurations retain their exact meaning and digest.

Authenticated Runtime Sessions

requests.runtimesession defines registerPalletRuntimeSession, applyRuntimeAssignment, reportRuntimeAssignmentObservation, and getRuntimeRegistryCredential. Pallet initiates WSS through its configured Cloudly HTTPS origin, or, on a relayed cluster, through the relay origin its cluster publishes in ICluster.data.relay while Cloudly remains the credential authority. The receiver derives the physical peer, process instance and registration idempotency; callers cannot supply those authorities. A serialized IRuntimeSessionBinding provides content to compare with authenticated peer state and does not itself authenticate a connection. Its clusterId is issued from the controller-owned node record during registration; it is never selected by Pallet, and a current-session check must fence the node to that same cluster. Registry credentials are ephemeral, bounded and authorized against an exact current run assignment.

After reconnect or credential rotation, the current authenticated session may submit an unchanged historical observation from its durable outbox. Server-indexed history must prove the same node, runtime namespace and durable controller, an older session epoch and an eligible credential generation. A historical-recovery receipt can advance the durable counter and acknowledge the exact body; it cannot restore active readiness, routes or liveness. Those require a fresh observation from the current session. The contract helpers validate and classify content; authentication and atomic receipt persistence remain server-owned.

A registration is identified by its node, its credential generation and nodeInstanceNonce: a canonical lowercase UUID that Pallet generates once per process start and sends on every registration. From 32.0.0 the field is required, so a controller no longer has a nonce-less case to decide. One physical socket may carry the registrations of many nodes — a cluster relay forwards a whole cluster — so the transport alone can no longer tell a node restart from a replay. The nonce is not a secret, is not authentication, and grants nothing on its own: the bearer still authenticates every registration.

The registration also states resolvedAuthorities: which of data.runtimeWorkloadAuthorities (network, secrets, storage) this node build can resolve, sorted and duplicate-free so a replay compares byte for byte. It is a capability of the registering build, never a permission: a controller refuses an assignment whose data.requiredWorkloadAuthorities(workload) the list does not cover, instead of a node accepting work it cannot carry out. It is not part of the registration identity — a changed build is a new process and therefore a new nonce — so a replay that states different authorities is a conflict rather than an update.

Every registration carries protocol, and the controller judges the body in the one server order protocol.readProtocolOffer states: the bounded body, the offer, protocol.negotiateProtocol against this session kind's own offer, and only then the exact-key validator, the bearer and any mutation. For a node that order is what keeps a refused registration from touching its packet access. registerPalletRuntimeSession answers { session, protocol }, and data.bindRegisterPalletRuntimeSessionResponse binds that answer to the request the node sent and negotiates the two offers from the node's side.

A node bound to a local controller (Node Runtime Routing Binding) serves the same session family on its root-only Unix socket. The controller connects and first sends requests.runtimesession.IReq_Onebox_Pallet_BindLocalRuntimeController (bindPalletLocalRuntimeController); Pallet then fires the unchanged registerPalletRuntimeSession at that connection, and every other request keeps its direction. The controller checks the registration with data.bindRegisterPalletLocalRuntimeSessionRequest, which requires the node and credential generation of its pinned confirmation and the bearer whose SHA-256 that confirmation acknowledged. Pallet binds the answer with data.bindRegisterPalletLocalRuntimeSessionResponse: the session must name exactly its own local binding's controller, node, cluster and runtime namespace, the credential generation sent and a compatible offer. The Cloudly binder keeps refusing a onebox session.

A changed nonce, or a changed credential generation, is a new registration: the controller mints a new binding and a new session epoch and fences the node's previous session. An unchanged nonce with an unchanged generation is a replay and must return the exact historical binding. Bearer rotation therefore needs no transport reconnect: a node that rotates behind a relay registers again with its unchanged nonce and its new generation on the same live socket, and the other nodes that socket carries are untouched.

registerCloudlyClientSession separately authenticates a current human or machine JWT on a physical TypedSocket so the server can own its identity tag; it carries the same offer, judged in the same order before the JWT, and answers with the server's own offer. Both registration methods reject HTTP transport. Sensitive requests and registry responses must be excluded from hooks, logs, diagnostics and durable journals.

Assignment-Bound Runtime Secrets

@serve.zone/interfaces/runtime exports the Pallet assignment-secret contract. runtime.createRuntimeAssignmentSecretAuthority takes an already validated, immutable IResolvedSecretManifest and, for launcher-environment deliveries, the pinned WorkloadInit approval reference and platform artifact. It computes a value-free descriptor digest before Cloudly computes the generation-one run assignment digest. workload.secrets contains that descriptor's { id, generation, digest } reference. The manifest digest, descriptor digest and assignment digest form an ordered chain; none depends on a later digest. bindRuntimeAssignmentSecretAuthority checks that chain against the full manifest, run assignment and pinned runtime config, including image and invocation digest, organization, service, platform and required launcher mode. The producer must separately prove that the manifest and WorkloadInit approval are currently accepted by their owning stores.

The Pallet-session methods getRuntimeSecretRecipientState, beginRuntimeSecretRecipientEnrollment and completeRuntimeSecretRecipientEnrollment enroll one X25519 public recipient generation for the authenticated node. The challenge is sealed with SmartCrypto and binds the exact current session, node, controller epoch, key, generation and expiry. Cloudly must derive the node from the current bearer-authenticated node session and separately verify ownership of its physical peer or cluster relay, retain one challenge hash with its expiry, consume it once, advance recipient generation by durable CAS, retire the prior recipient and reject a revoked node or stale session. Pallet holds the corresponding private key only in process memory. After a daemon restart it must generate a new pair and re-enroll before requesting fresh material; no old envelope or private key is persisted for replay. Existing running mounts require separate native ownership recovery.

getRuntimeAssignmentSecretMaterial names the exact current generation-one run reference, its sealed secret authority reference and the active recipient key and generation. Its response carries the full value-free descriptor and manifest, exact sorted version coverage, an approved immutable WorkloadInit authority when launcher delivery requires it, and X25519 envelopes. Each envelope context binds the final assignment digest, node, organization, cluster, service, descriptor digest, image/invocation/secret rollouts, secret version and recipient generation. bindGetRuntimeAssignmentSecretMaterialRequest validates the request against server-owned current assignment/session and active recipient inputs. bindGetRuntimeAssignmentSecretMaterialResponse additionally validates the response against the pinned config and cluster bound into that authenticated session; neither helper decrypts. A serialized IRuntimeSessionBinding does not authenticate a peer. Cloudly still owns node-session authentication and peer or relay ownership, current slot/service authority, recipient CAS, accepted-version retention and issuance policy. Pallet owns recipient private-key custody, native memory mounts and exact terminal cleanup. Both request and response bodies are sensitive and must be excluded from hooks, logs, diagnostics and durable runtime configuration.

ISecretVersionRetentionReference uses runtime-assignment with the exact assignment ID as resourceId when a Pallet run still owns a secret version. This reference has no expiresAt: plan acceptance windows do not end a running assignment's hold. Cloudly releases it only after joining the exact generation-three removal receipt, which Pallet may emit only after its owned runtime and secret mounts are removed. Retention does not itself authorize fresh material issuance to a stale session or recipient.

Joined Runtime Terminal Receipts

data.IRuntimeAssignmentTerminalReceipt and requests.runtimesession.reportRuntimeAssignmentTerminalReceipt carry a separate receipt for the exact Pallet attempt's joined terminal execution. Sequence 1 binds the generation 2 stop assignment; sequence 2 binds generation 3 removal and the exact stop receipt digest. Both retain the original generation 1 run reference. The full immutable assignment chain and both receipt bodies remain durable after removal, permitting an exact earlier stop receipt to be acknowledged again.

Only a durably completed native operation can issue a terminal receipt. Native stop joins container and sandbox stop; native remove requires an already stopped sandbox and joins container and sandbox removal. The assignment owner separately requires the exact completed stop before admitting removal. Terminal evidence contains precise CRI identifiers, the kernel boot UUID, and stopped/absent state; this native path does not collect IP addresses. Pending, unknown and host-fenced journal entries never qualify. A proven boot change can permit a new joined terminal operation but cannot skip the stop/remove chain.

The synchronous snapshot helpers validate bounded exact data only. Digest, assignment-chain and receipt-chain helpers do not prove native execution or authenticate a reporter. bindRuntimeReporterSession and the terminal report binder require the current physical session and immutable server-indexed reporter history from their trusted owners. Reconnect recovery preserves every original receipt field and returns historical-recovery; it never restores active readiness, routes or liveness. The response binder rejects rewritten receipt bodies.

Cloudly must transact new acceptance against the exact global incumbent, stored assignment chain and current peer, credential and node authority. A first receipt after the incumbent changes is denied; only an already persisted exact receipt can be acknowledged again under current authentication. These receipts settle only the named Pallet attempt. They do not authorize storage relocation, prove whole-host or Docker/Coreflow absence, grant another controller or organization authority, or replace independently qualified readiness and routing evidence.

Runtime Workload Logs, Stats, Exec And Volume Archives

These requests ride on a node's authenticated runtime session (requests.runtimesession). Every request carries session, and the receiver compares it with its own current session and with the connection the request arrived on: a serialized session is never authentication. Exec and volume archives belong to a controller on the node's own host alone (controller.kind: 'onebox', which only a local controller binding carries), because the local socket's root-only permissions are what authenticate that controller.

Logs. IReq_Pallet_Controller_ReportRuntimeAssignmentLogs (reportRuntimeAssignmentLogs) pushes one IRuntimeAssignmentLogBatch { assignmentId, attemptId, captureId, sequence, entries, final } from the node to its controller. Logs are the controller's to keep (IRuntimeConfig.logs); the node's capture files are disposable transport artifacts, so the node pushes and the controller persists. The node discards a batch only after the controller acknowledged it, and sends the next one only then. Two rules below let it discard more: a not-kept answer, and its bound on ended captures awaiting acknowledgement.

  • entries are in capture order. A line entry { stream, timestampNs, text, partial } is one piece of one line: text holds no \n, invalid UTF-8 becomes U+FFFD, and a line longer than one entry (maximumEntryTextBytes, 64 KiB) travels as pieces that state partial: true up to the last. The joined line is at most the config's logs.maximumLineBytes.
  • A lost entry { reason, stream, timestampNs, bytes, lines } is evidence of output that did not arrive, placed where it was lost: buffer-overflow (the node's buffer reached logs.maximumBufferedBytes and it dropped the oldest bytes), line-cut (the rest of a line beyond maximumLineBytes, right after that line's last piece), capture-restart, not-captured (a container started without capture) and not-kept (output the node dropped while its controller kept no logs). capture-restart and not-captured cannot know how much was lost and state bytes and lines as null; the others count exactly.
  • captureId names one capture of the attempt. sequence counts that capture's batches from 1, and a number once sent is never sent again with other content (below). final states that the container is gone and its output drained; only a final batch may be empty. A report is at most 1 024 entries and one mebibyte of canonical JSON, and timestamps are nanoseconds since the Unix epoch as decimal digits.
  • discardedCaptures, optional, is node-level loss evidence carried by whichever batch the node creates next: how many ended captures it discarded whole, under its ended-capture bound or a not-kept answer (below), since it last stated a count. It is absent for none and at least 1 when stated.

The answer { status, assignmentId, captureId, sequence } (TRuntimeAssignmentLogsStatus) is accepted, replay for a batch the controller already holds, or not-kept from a controller that keeps no workload logs: it holds nothing, and the node sends no more batches on that session (a later session may be answered differently). A controller answers every batch with one of the three and never leaves the method unhandled, so a node tells a controller that keeps no logs from a failure worth retrying. A controller must therefore handle the method before a node that captures logs runs against it. Likewise, a receiver on interfaces 32.37.0 or earlier reads a batch by its exact schema and refuses one that carries discardedCaptures or a not-kept loss, on every retry: controllers must run a later interfaces release before any node sends either.

A sequence number, once sent, is never sent again with other content; a retry resends the batch unchanged, its discardedCaptures included. Reuse would lose data silently: a controller that stored the batch while its acknowledgement was lost answers replay and drops the new content with its loss evidence. So when the node drops a running capture's batch it sent that is still unacknowledged, which only not-kept permits, the next batch it sends of that capture takes a higher number it never sent and opens with a not-kept loss that counts what was dropped. A controller that holds the dropped batch sees no gap, and that loss also counts content it holds.

The controller never refuses a batch for a gap in sequence. Every gap the node leaves by discarding carries evidence: the batch after skipped numbers opens with a not-kept loss, and a capture discarded whole is counted in discardedCaptures. A gap before a batch that opens with a not-kept loss is therefore the node's discard, which that loss counts; a gap before any other batch is the controller's own loss, which it records. A capture whose final batch never arrives leaves no gap to read: it may still be running, and one discarded whole is counted, never named.

not-kept is the node's permission to discard for the session that answered it. The node may discard every ended capture completely, content and identity, and each capture that ends later while the answer stands, counting them in discardedCaptures. A running capture stays within logs.maximumBufferedBytes (overflow is a buffer-overflow loss), or the node drops its output, a sent and unacknowledged batch included. A later session that keeps logs receives each running capture from where it stands, after a not-kept loss that counts what was dropped.

While a controller is away or slow, ended captures wait for their final batch to be acknowledged, and a crash-looping workload ends one per restart. A node therefore keeps at most runtimeAssignmentLogContract.maximumEndedCaptures (64) ended captures awaiting acknowledgement, holding at most maximumEndedCaptureBytes (64 MiB, the largest logs.maximumBufferedBytes a config may set, runtimeConfigContract.maximumLogBufferedBytes, so one capture always fits) together. Past either bound it discards the oldest ended captures completely and counts them in the next batch's discardedCaptures. It does not name them: remembering them would grow without bound while the controller is away. The batch is the unit of retry for the count: a batch the node discards unacknowledged passes its count on to the next batch, so a count whose acknowledgement was lost may arrive again inside a later one.

The count belongs to the node, not to the assignment, attempt or capture of the batch that carries it. A controller that keeps logs records it against the node and the session the batch arrived on, once for each batch it accepts and never for a batch it answers replay. Counts are not idempotent: a lost acknowledgement can repeat a count, so their sum may exceed the captures discarded. A controller that keeps no logs does nothing with the count beyond answering not-kept; the node keeps it for a later session.

bindReportRuntimeAssignmentLogsRequest(request, assignment, logPolicy, trustedSession) checks the session, that the assignment is this node's and carries the batch's attempt, and every line's joined length against the policy. bindReportRuntimeAssignmentLogsResponse(response, request) binds the acknowledgement to the batch.

Reading kept logs. An operator reads what Cloudly keeps through two service requests, which take { identity, serviceId } and are authorized like every other read of that service. Cloudly keeps an attempt's logs from the moment it issues the attempt, so an attempt that failed before writing a line is listed with nothing held, and an attempt whose slot was replaced or released stays readable until its logs expire.

  • requests.service.getServiceRuntimeLogAttempts { identity, serviceId, cursor, limit } answers an IRuntimeAssignmentLogAttemptPage { attempts, nextCursor }: the service's attempts, newest issued first and ties in assignmentId order. An IRuntimeAssignmentLogAttemptView states serviceId, replicaId, nodeId, assignmentId, attemptId, issuedAt, lastAcceptedAt (null until a batch was accepted), expiresAt (when the attempt's logs are dropped whole: the latest of issuedAt, lastAcceptedAt and the time the attempt's slot released it, plus the config's logs.retentionMs; null while the slot still holds the attempt, which schedules no expiry however quiet it is), captures and capturesEnded (captures with a batch, and those whose final batch arrived), keptBytes (UTF-8 bytes of text held) and truncatedBatches. limit runs from 1 to runtimeAssignmentLogReadContract.maximumAttemptPageSize (256).
  • requests.service.getRuntimeAssignmentLogs { identity, serviceId, assignmentId, cursor, limit } answers an IRuntimeAssignmentLogPage { attempt, records, nextCursor }: one attempt's records in delivery order, captures in the order the first batch of each was accepted and each capture by sequence. limit counts records, from 1 to maximumPageRecords (256). A page also ends before its records would hold more than maximumPageEntries (4 096, four full batches) entries or its canonical JSON would cross maximumPageCanonicalBytes (2 MiB), so it may hold fewer than limit while nextCursor names more. One batch always fits: its record is smaller than the report that carried it.

A TRuntimeAssignmentLogRecord is one of three, so nothing is ever dropped silently:

  • batch { captureId, sequence, acceptedAt, final, entries }: a batch as the node sent it, its lines and loss entries unchanged;
  • gap { captureId, fromSequence, toSequence }: sequence numbers whose batches Cloudly did not receive or no longer holds, its own loss: numbers that never arrived and that the node did not account for, and accepted batches whose payload Cloudly has lost. Numbers the node skipped are not a gap: the batch after them opens with the not-kept loss that counts them;
  • truncated { captureId, throughSequence, reasons, batches, bytes, entries }: every batch Cloudly held of the capture up to throughSequence and has since dropped, oldest first, with exact counts. reasons (TRuntimeAssignmentLogTruncationReason, duplicate-free, in the canonical order batch-cap, byte-cap, retention) are batch-cap, Cloudly's bound on how many batches it keeps of one attempt, byte-cap, its bound on the text it keeps of one attempt, and retention, the config's logs.retentionMs. A node delivers about once a second, so a workload that writes now and then sends many small batches, which the batch bound holds where the text bound would not. Gaps inside the range are no longer stated.

snapshotRuntimeAssignmentLogPage enforces that order: a capture's records are contiguous and never return, a truncation only opens its capture, a gap starts at the number after the record before it, a gap never directly follows another gap (the controller states adjacent losses, received or not, as one gap, across pages too), a batch continues at the next number or, past numbers the node skipped, opens with a not-kept loss, and nothing follows a final batch. A page may resume inside a capture, so a capture's first record on a page is not judged against what came before it. A nextCursor is an opaque identifier, null on the last page, and never follows an empty page. snapshotRuntimeAssignmentLogAttemptView and snapshotRuntimeAssignmentLogAttemptPage read the attempt side. bindRuntimeAssignmentLogAttemptPage(page, serviceId, limit) and bindRuntimeAssignmentLogPage(page, serviceId, assignmentId, limit) match an answer to its read. A refused read carries IRuntimeAssignmentLogReadErrorData { code, retryable: false }, where TRuntimeAssignmentLogReadRefusal is attempt-not-found (no logs of that attempt for that service), cursor-invalid or limit-invalid.

Stats. IReq_Controller_Pallet_ReadRuntimeAssignmentStats (readRuntimeAssignmentStats) reads one sample of one attempt: { session, assignment } names the exact current revision, and the answer is { status: 'sampled', assignment, sample } or { status: 'not-running', assignment, sample: null }. An IRuntimeAssignmentStatsSample states timestampNs, the cumulative cpuUsageCoreNanoseconds (decimal digits: it outgrows a JSON number; utilisation comes from two samples), memoryWorkingSetBytes, memoryLimitBytes (the limit the node enforces, the config's memoryBytes), network { receivedBytes, transmittedBytes } of the sandbox, or null, and writableLayerBytes, or null. Any controller may read stats; a node answers one read per assignment at a time. bindReadRuntimeAssignmentStatsRequest(request, assignment, trustedSession) and bindReadRuntimeAssignmentStatsResponse(response, request) bind both sides.

Exec. IReq_Onebox_Pallet_OpenRuntimeAssignmentExec (openRuntimeAssignmentExec) runs one command in the container of a run assignment's running attempt. The request is { exec: { session, assignment, command, tty, terminalSize }, stdin }: command is the argument vector the runtime executes directly (at most 256 arguments and 64 KiB, no NUL, a non-empty first element), terminalSize { columns, rows } is stated exactly for tty, and stdin is a sending VirtualStream or null, whose end is the process's end of input. The answer { exec: { execId, assignment, tty, idleTimeoutMs }, stdout, stderr } returns receiving streams that outlive the request; a tty session's whole output is on stdout and its stderr is null. resizeRuntimeAssignmentExec changes a tty session's terminal size, and closeRuntimeAssignmentExec ends a session, releases its slot and answers how it ended: exited with exitCode (128 plus the signal number for a signalled process) or terminated with the termination that ended it first, closed, idle-timeout or container-stopped. runtimeAssignmentExecContract fixes the bounds: four sessions per node, counting ended ones not yet closed, and 15 minutes without a byte in either direction before the node ends one. A session also ends with the controller session it was opened under, and its execId is gone with it. bindOpenRuntimeAssignmentExecRequest, bindResizeRuntimeAssignmentExecRequest and bindCloseRuntimeAssignmentExecRequest take the node's trusted session and refuse any controller but a local one; the response binders match the answer to its request.

Volume archives. A local controller backs up and restores a node-local volume as one POSIX pax archive (runtimeLocalVolumeArchiveContract). The first entry is the volume root ./, and entries are directories, regular files, symbolic links (stored, never followed), hard links to earlier entries and FIFOs, with numeric uid/gid, exact mode bits and modification times. Sockets are omitted; a device node is refused, and extended attributes and ACLs are not carried. An archive is a whole number of 512-byte blocks and at least 1 536 bytes. A node moves one archive per volume at a time.

  • readRuntimeLocalVolumeArchive (IReq_Onebox_Pallet_ReadRuntimeLocalVolumeArchive): { session, claim } names the node's current generation of a bound directory claim. The node refuses while a running container holds the volume. It reads a volume nobody holds, or one whose holder is stopped, which a node never restarts, and grants the volume to no container meanwhile. The answer { volume: { claim, format: 'pax', holder }, archive } streams the archive. It carries no integrity, because the node learns the length and digest only while writing.
  • writeRuntimeLocalVolumeArchive (IReq_Onebox_Pallet_WriteRuntimeLocalVolumeArchive): { volume: { session, claim, format, byteLength, digest }, archive } writes an archive as the bytes of an import claim that waits in import-required, the online form of the offline import. The stream's integrity must be getRuntimeLocalVolumeArchiveIntegrity(volume). The node unpacks into a private staging tree and publishes it as the volume only after the whole archive verified. The tree keeps the archive's ownership, as every imported tree does. The answer { claim, status: 'receiving' } permits sending; the stream's acceptance receipt confirms the volume is ready, and a rejected or aborted stream publishes nothing.

bindReadRuntimeLocalVolumeArchiveRequest(request, claim, trustedSession) and bindWriteRuntimeLocalVolumeArchiveRequest require a local controller's current session and the exact reference of a bound directory claim of its controller, cluster, node and runtime namespace, an import claim for a write. Whether the volume is ready or still import-required, and who holds it, are the node's ledger checks.

Cluster Runtime Phase

data.IClusterRuntime { id, phase, generation, changedAt, changedBy } records which runtime owns a cluster's workloads: TClusterRuntimePhase is swarm, draining, switching or pallet. It is stored beside the cluster document, because a cluster update merges caller data and no caller may move a phase with it; generation is its compare-and-swap counter. snapshotClusterRuntime checks the record. Which phase may follow which belongs to the cutover, whose requests ship with it.

requests.cluster.getClusterRuntime { identity, clusterId } answers { runtime, blockers }. Each TClusterRuntimeBlocker names what it is about: cluster-wide blockers (relay-not-enabled, relay-not-registered, address-plan-missing, vpn-unconfigured) name nothing, node blockers (node-without-pallet-credential, node-arch-unsupported) name the nodeId, and service blockers (service-multi-cluster-scope, service-authority-unsupported, platform-binding-present) name the serviceId.

Fleet Runtime Cutover Authority

data.IFleetCutoverCensus is Cloudly's accepted, Spark-authenticated Docker census. It carries only immutable Docker service/spec identity, exact mode and replica count, hashed constraint identity, the ownership labels used by the cutover, exact tasks and value-free mounts. A manager reports only manager-visible /nodes, /services and /tasks state. Every node separately reports its own complete local physical-container list through the existing authenticated sparkSwarmNodeContracts.swarmObservation route; a manager never attests remote Docker-daemon contents. The local list is fenced by exact sorted container-ID set digests read before and after all bounded inspect calls, so a container appearing or disappearing during inspection refuses the snapshot. The accepted census references two consecutive manager observations with identical Swarm and inventory digests and one current accepted local observation for every target node. Every reference preserves the authenticating node, credential generation, reporter session, sequence, observation time and fixed freshUntil; assembling a later census cannot renew an old observation. It never transports environment, arbitrary labels, mount paths, volume names, registry credentials or secret values.

data.computeFleetCutoverDockerServiceSpecDigest hashes the complete open-schema Docker Service.Spec, including response properties unknown to this Interfaces version, under serve.zone/fleet-cutover-docker-service-spec. The helper returns only the digest: the full spec, including any environment values or secret references it contains, never belongs in the observation or logs. data.computeFleetCutoverDockerConstraintSetIdentity sorts exact constraint strings in UTF-16 code-unit order, preserves duplicates, and binds both the resulting list and its count. data.snapshotFleetCutoverDockerDesiredTaskCount admits only the non-negative safe integer from ServiceStatus.DesiredTasks; the producer must request service status and must not derive a replacement from task state. These helpers consume only the Docker response fields they name, so additional response properties do not make a newer daemon unreadable.

The only established historical service-identity ownership label is data.fleetCutoverDockerLabelNames.serviceId, whose exact value is serve.zone.serviceId. ownershipLabels.platformOwner: null means there is no known Docker-label evidence of a platform owner. It does not classify the service as an application, and neither a service name nor an unrecognized label supplies that classification. Cloudly's accepted administrator classification remains the authority for unlabeled legacy platform services.

Mount identity has three explicit canonical domains. serve.zone/fleet-cutover-mount-source-identity hashes { type, source }, where source is the exact UTF-8 Mount.Source returned by Docker, including an empty string. serve.zone/fleet-cutover-mount-target-identity hashes { target }, where target is the exact UTF-8 Mount.Destination. serve.zone/fleet-cutover-mount-identity then hashes { type, classification, sourceIdentityDigest, targetIdentityDigest, readOnly }. All three use strict canonical JSON and no case, slash, path, volume-name or alias normalization. Only the digests and classification cross the wire. unknown type/classification remains visible and prevents Cloudly accepting a complete scope; missing, unmanaged or conflicting service/container ownership also prevents completion rather than omitting the container. Structural validation does not establish freshness, authentication, manager consensus, ownership or Docker truth; Cloudly establishes those policy facts while dereferencing the accepted observations.

Docker exposes two different volume sources. A service-spec Mount.Source is the logical volume name, while a local container MountPoint.Source is its physical host path and MountPoint.Name is the logical name. data.createFleetCutoverDockerServiceMountDescriptor and data.createFleetCutoverDockerLocalMountDescriptor therefore keep their exact raw-source mount identities distinct. They never substitute Name for Source. The local-only data.createFleetCutoverDockerPhysicalMountEvidence adds a value-free correlation from the physical mountIdentityDigest to the service-spec identity derived from exact MountPoint.Name for volumes, or exact Source for bind and tmpfs mounts. Its independently domain-separated correlation digest binds that pair. When a physical container states mountCorrelations, the list must be sorted, unique, bounded, and a complete one-to-one mapping of its mounts. The property is optional only so observations created before this evidence existed still parse; absence proves no join, and a cutover consumer requiring a physical/service match must refuse it. Unsupported local mount types and volume mounts without a non-empty Name cannot produce correlation evidence.

The shared exact classifier recognizes only a writable /var/run/docker.sock to /var/run/docker.sock bind as docker-socket, a read-only /sys/fs/cgroup to /host/cgroup bind as cgroup, every volume as persistent-volume, other binds as bind, and tmpfs as tmpfs. Other Docker types remain unknown; no path normalization or service-name inference is performed.

data.IFleetCutoverScope is the server-owned authoritative service set. It starts incomplete; every complete entry binds an exact Docker service id to the current service authority and an accepted classification authority. unknown cannot appear in a complete structural snapshot. An authenticated administrator may submit IFleetCutoverScopeClassificationRequest against the exact current census, including for an unlabeled legacy Coretraffic or Corestore service. Cloudly still compares the declaration with current service, platform-service, credential, storage and census authority before accepting it. A classification is an operator decision, never a client proof boolean or a service-name inference.

data.TFleetCutoverPhaseAdvance permits only swarm -> draining -> switching -> pallet and carries opaque references to owner-accepted records. requests.fleetcutover.advanceFleetCutoverPhase dereferences and joins them inside Cloudly's phase CAS. Leaving draining requires the accepted fresh census/quiescence decision, the complete current scope and target set, every runtime spec held, and the accepted set of root-local volume fences. Entering pallet still requires every spec held. switching grants session and network preparation only; it runs no service. Structural snapshots enforce the closed forward shape but do not claim the referenced records exist or are current.

Before the first swarm -> draining CAS, Spark starts the replacement Coreflow32 relay as the exact Docker service/spec/image, delivers only its new generation-bearing credential and obtains the live accepted relay registration. Cloudly then revokes the old credential and accepts the current Spark bundle/unit/configuration and old-mode refusal evidence. This is a pre-switch Docker owner; it is never represented as a Pallet assignment. Downstream service holds and storage fences therefore depend on a live relay without requiring assignment authority that cannot exist yet.

The storage bootstrap separates desired identity, observed legacy identity and provider-proved physical lineage. IFleetCutoverPreparedVolumeHandoffIntent is Cloudly's accepted pre-switch statement. It carries the full sealed IRuntimeStorageClaimIntent and the exact persistent observedMount, including Docker source, target and combined identity digests. The Docker source digest commits only to Docker's exact source bytes; it is never physical lineage. The prepared handoff deliberately has no sourceLineageDigest. Corestore creates that independent commitment from its private device, inode and content proof while the native attachment fence is held.

An administrator changes prepared handoffs through prepareFleetCutoverVolumeHandoffs. Each compare-and-swap mutation contains 1..32 combined upserts and explicit removals. An upsert supplies only { nodeId, serviceId, mountIdentityDigest, request }, where request is the full canonical retained ReadWriteOnce filesystem request. Cloudly derives the volume, controller, organization, namespace, claim identity, evidence references and request digest from current owner rows. It re-derives the complete required set from the admitted manager task inventory and accepted local observations. The maximum possible physical set is maximumTasks * maximumMountsPerOwner, or 4096 * 32 = 131072; node placement is already carried by each admitted task, so a service-count bound would undercount replicated placement.

IFleetCutoverPreparedVolumeHandoffAggregate is a revisioned header with required/actual counts, state, page count and a complete set digest. It does not inline page references. Each immutable IFleetCutoverPreparedVolumeHandoffPage contains 1..32 sorted handoff references, with the owning nodeId committed beside every exact handoff revision, plus its content digest and the exact aggregate revision. Readers pin one aggregate, fetch every consecutive page and reject missing, duplicate, stale or mixed-revision pages. The set digest covers the ordered page-content digests. A 32-entry mutation is therefore a wire bound, not a fleet volume ceiling.

Spark reads one pinned page from /spark/swarm-nodes/fleet-cutover-prepared-volume-handoffs. The response carries the full immutable page and the authenticated node's complete full-handoff subset, including an explicit empty subset. Spark forwards held fences through the optional fleetCutoverHeldVolumeFences member available in every Swarm-observation state. The page's node-bearing references determine the exact subset; fetch responses and fence evidence must match every committed reference in page order, so omission is rejected and an empty report is valid only for a node with no entry on that page. The request binds the evidence node to the authenticated reporter. That evidence has a dedicated 128 KiB encoded budget and never displaces the existing census snapshots. One node's page report proves only that node's subset; Cloudly establishes global completeness across accepted reports for every page and node.

Pallet persists TFleetCutoverVolumeFence in held state while draining. A hold grants no claim, session, assignment, attachment, reservation, confirmation or execution authority. After the real current runtime reaches pallet, Pallet obtains IFleetCutoverStorageAdoptionGrant through its live authenticated session. The grant binds the exact prepared handoff, held fence, current pallet runtime, session generation/epoch and observed mount. Pallet keeps its per-volume lane while it invokes the isolated Corestore one-shot. Corestore performs physical-lineage proof and the initial ready-claim CAS under its native attachment fence, without calling back into Pallet, and returns the full claim with IFleetCutoverCorestoreAdoptionReceipt. Pallet revalidates the live session, runtime phase and grant after that potentially long operation before consuming the durable hold. A stale grant leaves the hold intact.

bindFleetCutoverAdoptionToStorageClaim joins the provider receipt and full ready claim to the prepared desired intent. bindFleetCutoverStorageAdoptionBundle additionally joins the current grant and consumed fence. Neither binder equates the provider physical-lineage digest with the observed Docker source digest or requires the provider's opaque resourceRef to equal Cloudly's volume ID. They are content checks; current owner dereferencing, peer authentication and native attachment fencing remain mandatory.

IFleetCutoverCoreflowRetirementEvidence binds the real pre-switch Docker Coreflow32 relay by exact service ID, container ID, spec version/spec digest and image digest to accepted manager and local node observations. It requires no Pallet assignment. Cloudly dereferences those observations to require a live exact container without Docker/cgroup mounts, then joins its generation-bearing relay credential and accepted live registration to the current Spark credential, bundle, unit and accepted observation. It separately binds invalidation of the legacy credential, refusal of the old callable mode, and Spark's configuration CAS. The legacy source is identified by SHA-256 and may be retained-inert or removed; deletion and secure erasure are not cutover requirements. Its exact owner generation is retained with that disposition. Retained recovery evidence is valid only while the released owner cannot consume it into live mode and the old credential remains revoked. relay.credential, relay.acceptedRegistration, legacy.credential and legacy.invalidation reference the records of Cluster Relay Credential. The old-mode refusal, the configuration CAS and the Spark bundle and unit identity are Spark-owned and come from Spark Legacy Runtime Retirement.

Before any post-pallet assignment exists, an administrator calls requests.fleetcutover.prepareFleetCutoverIngress with the exact server-created service ID, current pallet runtime, complete accepted scope and target set, held exact-cluster spec revision, and expectedDesignation. Null creates the first designation; an exact current designation reference permits only same-service credential rotation. Replacing the designated service is a different owner flow and is refused here. Cloudly's one owning transaction verifies the admin, same-organization live service, current references, held exact-cluster spec, zero configured and actual published/public ports, and absence of a conflicting designation. It generates the managed launcher-environment bearer privately, stores its active version and hash authority, advances the held spec if its secret input changed, and accepts IFleetCutoverIngressDesignation. The public designation exposes only credential ID/generation and the active ISecretVersionReference; neither it nor the response exposes the bearer, hash, private key, or any other secret value. Preparation requires no assignment, runtime session, ingress registration, legacy absence, publication, port-transfer, readback, or IFleetCutoverIngressAuthority.

The preparation request is a closed compare-and-swap body. A replay requires the same mutation ID and exactly the original normalized request payload and returns the recorded designation/spec result; it does not renew generation, acceptance time, freshness, or authority. Every downstream authorization must dereference the current designation, held spec, runtime, scope, and target set. A historical replay response never grants current authority. snapshotFleetCutoverIngressPreparationRequest and snapshotFleetCutoverIngressDesignation check exact structure, reference relationships and the designation's canonical digest. They cannot establish admin identity, currentness, organization, service liveness, held state, zero ports, or absence of a conflicting designation; Cloudly owns all of those acceptance checks.

The remaining post-pallet order is represented without a second route or publication model. First Cloudly releases only the designated new ingress spec under an accepted empty published-port policy. It then accepts the existing cluster-ingress registration and route-table revision, exact Spark evidence that the old ingress is absent, the current source and target publication records and owning port-transfer receipt, and a readback of the current route revision as IFleetCutoverIngressAuthority. Only an application release may cite that authority. The legacy Docker Coretraffic instance therefore remains through switching; the new ingress runs without public ports before exact port ownership moves, and applications remain held until the new authoritative route is read back.

The owner methods live under requests.fleetcutover: administrator scope classification, storage handoff preparation, ingress preparation and forward phase CAS; root-private prepared-fence hold, read and owner-triggered adoption; isolated Corestore prepared-volume adoption; ingress-authority read; and ingress/application spec release. Current-session adoption authorization, accepted bundle reporting and full claim readback live under requests.runtimesession. The published 32.x holdFleetCutoverVolumeFence, consumeFleetCutoverVolumeFence and adoptFleetCutoverVolume declarations remain for type compatibility only. Production owners do not register or adapt those incomplete operations; they are marked for removal at the next protocol major. Spark transports local and manager evidence through its existing authenticated Swarm-observation route, and Cloudly assembles census and Coreflow-retirement records from accepted owner state; there is no caller-supplied census or retirement proof endpoint. Accepted evidence references are opaque { id, generation, digest } records. They contain no secret bytes and gain meaning only when the named server owner dereferences them under current authorization and policy.

Service Runtime Specs And Status

data.IServiceRuntimeSpec is what Cloudly runs for one service on a Pallet cluster: { id, organizationId, clusterId, placement, replicasPerNode, hold, profile, network, revision }. It is stored complete and beside the service document, so a hold written by a cutover never races an operator's service edit.

  • placement is { kind: 'every-node' } or { kind: 'nodes', nodeIds }, with nodeIds sorted, unique, never empty and at most maximumPinnedNodes, the protected authority's egress limit.
  • replicasPerNode runs from 1 to 64; Cloudly refuses a service desiring more than maximumReplicas (64) as replica-limit.
  • profile { cpuMillis, memoryBytes, readiness, logs, readonlyRootfs, runAs } is the runtime config's execution policy ahead of any image: the same readiness and log shapes, and the same CPU, memory (at least 16 MiB), identity, log, readiness timing and HTTP request rules, which both validators share. runAs: null runs the image's own OCI user. Only the service's target ports decide whether a readiness probe's target exists: the runtime config requires the service to declare it, and the producer refuses a service that does not as readiness-target-undeclared.
  • network is { mode: 'isolated' } or { mode: 'attached', publicEgress, platformEndpointIds } with sorted unique endpoint ids.

snapshotServiceRuntimeSpec checks all of that; whether the cluster, the nodes and the endpoints exist is Cloudly's. One canonical service may have one runtime target on each of several clusters. Every public target keeps the same IServiceRuntimeSpec: id is the exact service id and clusterId names that target. Cloudly's private persistence joins those two values; no composite public service id is introduced.

requests.service.setServiceRuntimeSpec { identity, serviceId, expectedRevision, spec } writes TServiceRuntimeSpecWritableData, and Cloudly copies the organization from the service. Its compare-and-swap is for (serviceId, spec.clusterId): null means that exact target is absent, and each target advances its own revision. The request bytes and public spec shape are unchanged.

getServiceRuntimeSpecForCluster { identity, serviceId, clusterId } reads one exact target. getServiceRuntimeSpecs { identity, serviceId, cursor, limit } reads targets in cluster-id order, with an opaque cursor and a limit from 1 through serviceRuntimeContract.maximumSpecPageSize. Both reads answer each spec with its producer status, which is null until the producer first decides the spec. A status never comes without its spec.

The original getServiceRuntimeSpec { identity, serviceId } remains a scalar compatibility read: zero targets answer { spec: null, status: null }, and one target answers it. More than one target is refused with IServiceRuntimeSpecAmbiguityErrorData { code: 'service-runtime-spec-ambiguous', retryable: false } in TypedResponseError.errorData; Cloudly never selects a target implicitly.

removeServiceRuntimeSpecForCluster { identity, serviceId, clusterId, expectedRevision } removes one exact target only after Cloudly transactionally derives and fences all authority and safety facts: the revision still matches, the target is held, every target slot is fully settled and released, and the cluster's current phase still belongs to the spec producer. Callers provide no phase, settlement, release or authority booleans or evidence.

data.IServiceRuntimeStatus { id, specRevision, state, reason, changedAt } is the producer's verdict on one spec revision, and each reason belongs to exactly one state. converged and converging carry reason: null. waiting carries a TServiceRuntimeWaitingReason: node-network-pending, conflict, failure-backoff, storage-import-pending (an import local storage claim whose bytes have not arrived on its node), bindings-pending (the secret manifest does not cover every enabled binding of the service yet, or an object-storage binding is placed on no node), legacy-adoption-pending (a Corestore binding of the service waits for the offline adoption of its legacy resource, see Corestore Legacy Resource Adoption) or objectstorage-retention-pending (an object-storage binding declares compliance retention whose evidence has not arrived). refused carries a TServiceRuntimeRefusal: cluster-runtime-missing, organization-missing, image-plan-missing, platform-unsupported, secrets-unsupported, volumes-unsupported, port-protocol-unsupported, ports-too-many, readiness-target-undeclared, environment-too-large, no-eligible-node, replica-limit, node-authority-unsupported, network-policy-missing, network-policy-mismatch, network-workload-pool-missing, network-workload-pool-exhausted, volume-driver-unsupported, volume-options-unsupported, volume-shared-unsupported, volume-placement-unsupported, volume-initialization-unknown, volume-ownership-unknown, objectstorage-endpoint-missing, objectstorage-binding-invalid, objectstorage-environment-conflict, objectstorage-endpoint-unselected, database-endpoint-missing, database-endpoint-unselected, cluster-relay-missing or cluster-ingress-environment-conflict. The first four volume refusals name what a local storage claim cannot express: a volume driver other than a node-local directory, driver or mount options, one volume mounted by more than one replica or service, and a placement that does not pin each volume to one node. The other two name what Cloudly will not guess, see Service Volume Initialization. Of the four object-storage refusals, objectstorage-endpoint-missing, objectstorage-binding-invalid and objectstorage-environment-conflict name why a binding's environment cannot be composed, see Object-Storage Binding Environment, and objectstorage-endpoint-unselected, like database-endpoint-unselected, names an endpoint the workload's network does not select (below). The two cluster- refusals concern only a cluster's designated ingress service, whose environment composeClusterIngressEnvironment composes, see Cluster Ingress: cluster-relay-missing is a cluster that states no valid relay, which an administrator enables with enableClusterRelay, and cluster-ingress-environment-conflict an ingress service whose own environment already sets one of clusterIngressEnvironmentNames. ports-too-many is a service declaring more target ports than the runtimeConfigContract.maximumTargetPorts (64) a config carries, where port-protocol-unsupported is about one port's protocol; node-authority-unsupported is a node whose session resolved fewer workload authorities than the spec needs, refused at the admit step rather than handed to a node that would silently park it. An attached network requires an explicit service policy (IServiceNetworkEgress): its absence is network-policy-missing, and a disagreement with the requested egress or platform endpoints is network-policy-mismatch. A binding endpoint the workload's network does not select would leave its connection dark, so it is refused instead: database-endpoint-unselected for an enabled database binding whose node's valid corestoreDatabase is not selected, and objectstorage-endpoint-unselected for an enabled object-storage binding whose node's valid corestoreObjectStorage is not. An endpoint is selected when one of the spec's platformEndpointIds names a platform endpoint of the current protected authority with the endpoint's host as its address, tcp as its protocol and the endpoint's port as its port — for object storage the origin's port, or 80 for http: and 443 for https: when it states none. An isolated network selects nothing, and platform endpoints are IPv4, so a DNS-name or IPv6 host is never selected. Both are judged only for a valid endpoint: a node stating none, or an invalid one, stays database-endpoint-missing or objectstorage-endpoint-missing. platform.findUnselectedCorestoreEndpoints({ database, objectStorage, platformEndpoints, network }) is that rule: given one node's corestoreDatabase and corestoreObjectStorage (each null when no enabled binding of that kind is placed on the node), the current protected authority's platformEndpoints and the spec's network, it returns the TCorestoreEndpointSelectionRefusals that apply, sorted, or an empty list. A missing workload address pool in the current protected authority is network-workload-pool-missing; existing eligible pools without a free workload subnet are network-workload-pool-exhausted. Neither policy nor pool is silently defaulted. A refusal never retires running slots. snapshotServiceRuntimeStatus enforces the pairing.

Service Volume Initialization

A local storage claim has to say how its volume starts: empty with an owner and mode, or import of bytes that already exist. IServiceVolume.initialization (TServiceVolumeInitialization) is where the service states it, per volume, beside the mount path it belongs to:

  • { kind: 'empty', ownership } creates an empty root. ownership is { kind: 'explicit', uid, gid, mode }, or { kind: 'run-as', mode }, which takes the owner from the runtime spec's profile.runAs. The mode is always stated; it holds directory permission bits only (serviceVolumeInitializationContract.directoryModeBits, 0o3777, setuid refused).
  • { kind: 'import', source: { kind: 'docker-volume', volumeName } } waits for the bytes of an existing Docker volume on the claim node's host, named exactly as Docker lists it (a Corestore-driver volume included). The operator copies that tree into the node's import directory; the claim does not carry the source, and the node never reads it.

It lives on the volume and not on the runtime spec because it is a fact about the volume's data, which the service declares, while the runtime spec carries only execution policy and is stored exactly, so a new member there would refuse every stored spec. It is server-managed: setServiceVolumeInitialization { identity, serviceId, mountPath, expected, initialization } is its only write. The volume is named by its exact mount path; expected is the value the caller read, or null for none, and any other stored value refuses. createService and updateService take volumes as TServiceWritableVolume (no initialization) and Cloudly refuses one that states it; a stored initialization stays with the volume of the same mount path and goes with the volume. Refusals travel as IServiceVolumeInitializationErrorData { code, retryable: false }: volume-not-found, conflict, volume-initialization-invalid, and volume-initialization-bound, a change while a bound local storage claim of the volume exists.

validateServiceVolumeInitialization(value, path?) names every reason a value is refused. deriveServiceVolumeClaimInitialization(initialization, runAs) is the one rule for the claim fields, and nothing in it is a default:

Volume initialization runAs Claim initialization / initialOwnership
absent any refused volume-initialization-unknown
empty, explicit any empty / the stated { uid, gid, mode }
empty, run-as { user, group } empty / { uid: user, gid: group, mode }
empty, run-as null refused volume-ownership-unknown
import any import / null

An absent initialization is refused rather than created empty, because an empty volume in front of a service whose data sits in a legacy Docker volume would start that service without its data. A run-as ownership under runAs: null is refused because the image's own user is not a number Cloudly knows. The rule decides the first generation of a claim only: every later generation carries that claim's own initialization and ownership, whatever runAs says by then, since the node refuses a claim whose initial ownership changed.

Runtime Assignment Views

data.IServiceRuntimeAssignmentView is one replica slot's current assignment as an administrator reads it: serviceId, replicaId, nodeId, the current assignment reference, its disposition, the latest accepted observation { phase, ready, observedAt, sequence } (which may cite an earlier revision), terminalReceiptCount (0 to 2) and delivery. delivered means the node's own evidence, an observation or a terminal receipt, cites the current revision; until then it is undelivered.

requests.service.getServiceRuntimeAssignments { identity, serviceId } answers every slot of one service. requests.cluster.getClusterRuntimeAssignments { identity, clusterId, cursor, limit } pages through every slot on a cluster's nodes; limit runs from 1 to serviceRuntimeContract.maximumAssignmentPageSize (256), and Cloudly refuses any other. requests.service.pushServiceRuntimeChanged { serviceId } tells admin UIs that a service's spec, status or assignments changed; it carries no state, so a late push never overwrites a newer read. A slot's replaced and released attempts leave these views; their kept workload logs stay readable through getServiceRuntimeLogAttempts and getRuntimeAssignmentLogs (see Runtime Workload Logs, Stats, Exec And Volume Archives).

Node Retirement

data.INodeRuntimeRetirement { id, clusterId, state, disposition, actorId, since } records that one node leaves its cluster's Pallet runtime. TNodeRuntimeRetirementState is retiring while the producer still stops and removes the node's slots, the compiler quarantines its handoff and leases, both node credentials are revoked and the node's egress fence is not yet recorded, and retired once the protected authority no longer names the node (see Egress Owner Retirement). It is stored beside the node document, because a node update merges caller data and no caller may retire a node with it. The producers read the row's presence rather than its state: a node that has one is no longer an egress member and desires no slots, so the row is written before anything is taken away and stays as long as the node document. snapshotNodeRuntimeRetirement checks the record.

TNodeRuntimeRetirementDisposition (nodeRuntimeRetirementDispositions, snapshotNodeRuntimeRetirementDisposition) says how the node leaves. withdrawn: the node still runs and takes part — its slots are removed against its own removal receipts, Cloudly asks it through its projection's retirement request (see Projection Retirement Requests), and its Pallet answers by withdrawing its protection receipt before its credentials are revoked. lost: an administrator states the node will never report again, which authorizes abandoning its slots without removal receipts and nothing else; its egress fence is the same as a withdrawn node's.

requests.node.retireClusterNode { identity, nodeId, expectedGeneration, disposition? } answers { retirement }. expectedGeneration is the cluster runtime generation the caller read with getClusterRuntime, so a cutover that moved the phase meanwhile refuses the write. A request that states no disposition asks for withdrawn, the retirement this request always described; lost is taken only when stated, and asking again as lost escalates a withdrawn retirement that is still retiring. Every refusal is one of data.nodeRuntimeRetirementRefusals — node-missing, cluster-runtime-not-pallet, node-already-retired, conflict or node-still-reporting (a lost retirement of a node whose last accepted report is younger than clusterNodeObservationFreshnessMs) — and takes nothing away.

Corestore Database Export Packaging

The Node-only @serve.zone/interfaces/runtime export provides metadata-only Corestore database-export package contracts. A package request carries exact input length and SHA-256, approved source-backup and complete-output-receipt digests, a path-safe attempt ID, and an explicit absent or exact allocation expectation. It never carries the export bytes or Base64 content. Immutable completion binds the normal Corestore database backup receipt and exact closure identity; delivery wrappers may report exact replay without changing completion. Release is permitted only against the caller's exact fsynced closure and receipt identities. API-token authentication remains transport authority, and these Corestore contracts do not claim to verify Cloudly authority-B HMACs.

Exact input identities include one terminating LF and are limited to 96 MiB plus that LF. Package control metadata is limited to 16 KiB and keeps the existing 256 MiB Corestore closure limit.

One media type and one backup format are stated here for every process that handles a Corestore database backup: runtime.corestoreDatabaseBackupMediaType (application/vnd.serve-zone.corestore-database-backup) is the type an archive is served and stored under, and runtime.corestoreDatabaseBackupFormat (corestore.database.backup) is the format a receipt states. Corestore's control server, its export packager, Onebox and the cluster relay import them instead of writing the bytes again, so the side that writes a transfer and the side that reads it cannot disagree about its type.

Service creation carries canonical ownership separately from caller-writable service data. requests.service.IRequest_Any_Cloudly_CreateService requires a top-level organizationId, while data.TServiceWritableData excludes the server-managed ownership field. Consumers should use data.validateOrganizationId() to validate the identifier shape and verify the request's ownership authorization separately before persistence.

Protocol handshake

The protocol version is the installed @serve.zone/interfaces release. There is no protocol number beside it, no per-shape version suffix and no schemaVersion field: one release of this package is exactly one contract, so naming the release states everything about what a peer sends and what it can read.

Two peers exchange an offer when a session opens and refuse an incompatibility by name:

import { protocol } from '@serve.zone/interfaces';
import { TypedResponseError } from '@api.global/typedrequest';

// One offer per session kind a deployable serves, built once at start-up. A deployable that
// depends on nothing newer than its own major accepts that whole major.
const [installedMajor] = protocol.protocolVersion.split('.');
const ownOffer = protocol.createProtocolOffer(`${installedMajor}.0.0`);

// Server, on every carrier: the bounded body, then the offer, then everything else — the
// exact-key validator, the bearer and any mutation.
const peerOffer = protocol.readProtocolOffer(requestBody);
const negotiation = protocol.negotiateProtocol(ownOffer, peerOffer);
if (!negotiation.compatible) {
  throw new TypedResponseError(
    protocol.describeProtocolRefusal(negotiation.refusal),
    negotiation.refusal,
  );
}

// Client, on the payload of a failed hop: `TypedResponseError.errorData` on a TypedRequest, the
// parsed body on the Spark heartbeat's refusal status. Everything else is not a refusal.
const refusal = protocol.readProtocolRefusal(errorData);
if (refusal) {
  logger.warn(protocol.describeProtocolRefusal(refusal));
  // A refusal stands until an operator upgrades one of the two sides, so the client waits.
  setTimeout(offerAgain, protocol.refusedOfferRetryIntervalMs);
}
Export Purpose
protocolVersion The installed release, taken from the package's own commit info.
IProtocolOffer { interfacesVersion; minimumPeerVersion } — what a side speaks and the oldest peer it accepts.
createProtocolOffer(minimumPeerVersion) This build's offer. A minimum above the installed release, or in another major, throws here — at start-up, not at the first registration.
validateProtocolOffer(value, path?) Every reason an offer is refused, as an array of messages; empty for a valid offer.
validateProtocolOfferOn(body, path?) The same reasons for the protocol key of a body, read as a data property. The naming twin of readProtocolOffer, for validators that list reasons instead of throwing.
readProtocolOffer(body) The offer on a body, read from the protocol key alone and returned detached.
negotiateProtocol(local, remote) { compatible: true } or { compatible: false; refusal }. Pure; throws when either side's offer names something that is not a release version.
IProtocolRefusal { code: 'protocol-incompatible'; reason; refusing; refused }, carried as TypedResponseError.errorData or as a plain HTTP JSON body.
TProtocolIncompatibility 'major-mismatch', 'peer-below-minimum', 'self-below-peer-minimum'.
readProtocolRefusal(value) The refusal inside an error payload (TypedResponseError.errorData) or a response body, or null when the value is not one.
describeProtocolRefusal(refusal) One log line naming both sides, opening with the frozen protocol-incompatible code.
refusedOfferRetryIntervalMs 300_000 — the shared cadence at which a refused client offers again.

negotiateProtocol checks in a fixed order: different majors are major-mismatch even when a minimum would also fail, then a remote below the local minimum is peer-below-minimum, then a local below the remote minimum is self-below-peer-minimum. The last two are mirror images of one another — one holding on this side means the other holds on the peer — and they can never hold at once, which would require each version to be below the other. An offer that does not name two release versions is a local defect, not an incompatibility between two peers: negotiateProtocol throws on it, and a peer-sent body is read with readProtocolOffer first, which refuses it by name. Both offers and refusals come back frozen and detached from the values they were read from.

Every carrier reads the offer once, as data. A protocol key, or a member of the offer under it, that is an accessor rather than a plain value is refused as "must use its exact schema" and is never invoked — by readProtocolOffer, by validateProtocolOfferOn, by validateProtocolOffer and by readProtocolRefusal alike. No snapshot, validator or client on any carrier runs code a peer attached to a body, and none of them can observe one value while a session is served by another. This is stated here and nowhere else: the families that carry an offer go through these functions rather than restating the rule.

A carrier that negotiates puts the offer in a top-level protocol key on both its request and its response, and a server reads it before it validates the rest of the body, authenticates or changes any state. Validators in this package are exact-key, so a body from a later major is unreadable by an older build; reading the offer first is what turns that into a named refusal instead of a generic contract rejection, and it keeps a refused peer's credentials, sessions and packet access untouched.

Versions are MAJOR.MINOR.PATCH and nothing else — no prerelease, no build metadata, no v prefix, no leading zeros — and are compared per numeric component, so 32.10.0 is above 32.9.0. That grammar is public as data.IReleaseVersion, data.parseReleaseVersion, data.isReleaseVersion, data.compareReleaseVersions and data.maximumReleaseVersionLength: any policy that has to order two released versions — a minimum bundle a server believes, a supported floor — reads them here instead of splitting strings of its own, which is how a comparison ends up saying 32.9.0 is above 32.10.0. minimumPeerVersion is in the same major as interfacesVersion and never above it. Each deployable owns its minimums, one per session kind it serves, and raises one only in the commit that starts depending on a later interfaces minor, for that session kind alone.

Secrets v24 Pre-Cutover Contract

Version 24 retains the value-free v23 architecture while replacing the remaining pre-cutover lifecycle and runtime authority gaps. Secret values enter Cloudly only as strict SmartCrypto X25519 envelopes or an explicit bounded server-generation request. Coreflow receives a full digest-verified manifest plus one sealed envelope per pinned SecretVersion through the Node-only runtime export.

This is a contract release before consumer cutover. It does not claim that clean-v2 migration, backup verification, historical secret erasure, or storage cleanup has completed. Those operations remain separately gated and must be proven by their owning services before destructive cleanup.

Breaking removals include:

  • The SecretGroup and SecretBundle data/request modules and request namespaces.
  • getServiceSecretBundlesAsFlatObject and every service/preflight field tied to bundled flattening or aggregate runtime files.
  • Plaintext resolved runtime values, generic platform config/credential maps, credential-bearing Cloudly settings, and serialized object-storage credential references.
  • Hosted-app control-token identities and credential-bearing bootstrap actions.

The current secret contract includes:

  • TSecretValueInput, active-only IActiveSecretRecipientMetadata, the requests.secret.IReq_GetSecretIngressRecipient contract with method getSecretIngressRecipient, and fixed-order create/rotate/App Store context builders.
  • ISecretEnvelopeAdmissionBinding under the Node-only runtime export binds an exact envelope and request context to one recipient generation. The binding is reproducible and does not itself prove admission. Retiring-key retries require a trusted Cloudly mutation receipt created atomically while that recipient was active; caller-supplied admission or issuance timestamps are not part of the contract.
  • getSecretVersionPurgePreflight returns bounded advisory reference pages and complete blocker counts, including indefinite runtime-assignment holds until joined terminal removal proof. purgeSecretVersion identifies one exact version and fences the mutation by secret, version, and target revisions. The non-issuable purge-pending lifecycle and revisioned pending/erasing/failed/succeeded operation keep external erasure durable across retries without treating preflight as authorization.
  • IResolvedSecretManifest contracts, stable Docker resource naming, WorkloadInit map/wrapper helpers, and ISealedResolvedSecretMaterial under @serve.zone/interfaces/runtime.
  • Two-step Coreflow X25519 recipient enrollment and exact recipient lifecycle validators.
  • Mandatory IImmutableContainerInvocation evidence on immutable image deployment plans.

Secrets v24 Runtime Registration And Reporting

The Node-only @serve.zone/interfaces/runtime export adds the live contract that Cloudly must validate before publishing secret-bearing desired state to a Coreflow connection:

  • getSecretRecipientEnrollmentState returns either an exact generation-zero empty state or the complete valid recipient set for the cluster derived from the verified JWT.
  • getCoreflowSecretRuntimeRegistrationExpectation returns either an available expectation or an explicit unavailable reason. The available expectation binds the live reporter session, active recipient, fresh generation-fenced target authority, and active WorkloadInit approval through expectationDigest.
  • Spark sends authenticated, sequenced local Swarm membership observations and, on managers, complete manager snapshots. Cloudly derives scope from the Spark credential, reconciles manager consensus privately, and publishes one fresh single-Swarm target authority or a targetless unavailable state. Structural contract validation does not itself establish manager consensus.
  • TSparkSwarmObservation lets workers report only their local Swarm node ID, because Docker does not expose the Swarm cluster ID to workers. Cloudly may associate that node with a cluster only through authenticated node scope and accepted manager consensus. The transport contract designates sparkSwarmNodeContracts.swarmObservation, whose endpoint is /spark/swarm-nodes/swarm-observation and whose maxRequestBytes covers the largest snapshot this validator accepts; every body on it carries the sender's protocol offer. A conforming Spark sender must retain one exact request until it receives a request-bound acceptance receipt. A conforming Cloudly consumer must authenticate the node before inspecting replay state and derive cluster scope from that persisted identity.
  • computeSparkSwarmObservationDigest uses strict canonical JSON and the serve.zone/spark-swarm-observation domain. A reporter session starts at sequence one. The contract requires a Cloudly consumer to accept only the exact next sequence with a strictly advancing observedAt, while an exact same-sequence/same-digest retry returns the byte-equivalent persisted receipt before age checks. It must reject conflicting, stale, skipped, retired-session, old, future, or non-advancing reports. evidence-refused is the one rejection about content rather than sequencing: it names an evidence member this server refuses for this body and every resend of it, so the node must change the body instead of retrying it, while a server-side condition a later attempt could pass stays acceptance-race — telling those two apart while it accepts evidence is the consumer's own work. A reused session ID may be treated as new only after it leaves the bounded retired window, and its sequence-one observation must still advance the permanent timestamp high-water mark. The consumer must persist acceptance state and its receipt in one atomic boundary.
  • validateSparkSwarmObservation validates exact structure and attached snapshots, while validateSparkSwarmObservationRequest and validateSparkSwarmObservationResponse validate transport and request binding. None of these functions authenticates a node, persists replay state, establishes manager consensus, or creates runtime authority.
  • WorkloadInit approval binds a clean stable release identity, version-tagged OCI index, exact amd64 and arm64 platform/executable digests, policy generation, and a Cloudly summary of detached Cosign DSSE/SLSA verification. Public shape and digest validators do not verify the signature, public-key trust root, or private Cloudly policy. The provenance identifiers a verifier matches are workloadInitReleaseContract.buildType (https://serve.zone/buildtypes/workloadinit-gitzone-tsdocker) and builderId (https://serve.zone/builders/workloadinit-manual-release); the in-toto statement type, the SLSA predicate type, the Cosign bundle format and the OCI index media type keep the exact bytes their own specifications define. An approval carries no shape version, and IWorkloadInitApprovalAuthorityReference states only the counter authorityGeneration beside the approval digest it points at.
  • coreflowSecretRuntimeRegistrationTagId is the sole dedicated TypedSocket tag identifier. Its payload is exactly ICoreflowSecretRuntimeRegistration; cluster scope comes from the verified connection identity rather than the tag.
  • ICoreflowSecretRuntimeRegistration references the exact expectation, target generation/digest, and WorkloadInit authority. It must cover every Cloudly-node/Swarm-cluster/Swarm-node identity and approved per-platform manifest and installed-executable digest without self-asserting placement or approval.
  • capabilities is a sorted, duplicate-free list of data.TCoreflowRuntimeCapability names drawn from data.coreflowRuntimeCapabilities, judged by data.isCoreflowRuntimeCapabilityList. A build states what it can do by name; which shape it speaks is the interfaces version both peers exchanged. Every name in data.requiredCoreflowRuntimeCapabilities must be present. That list is the vocabulary minus data.optionalCoreflowRuntimeCapabilities, which holds corestore-inventory alone: it is the one capability a build may lack, because it depends on the node's Corestore deployment rather than on the contracts the build was compiled against.
  • validateCoreflowSecretRuntimeRegistration compares the registration with a trusted expectation built by Cloudly, including the live reporter session. Missing, extra, duplicate, reordered, or mismatched node evidence fails closed. Consumers must discard the registration on transport disconnect, tag removal or replacement, or any live session, placement, artifact, or recipient expectation change before publishing more secret-bearing state.
  • reportSecretDeploymentState reports only applying, applied, drifted, or failed for the manifest and plan revision selected by Cloudly. The validator receives the trusted cluster ID after JWT verification and the trusted live reporter session. Wire-provided manifest scope is checked against, and never replaces, that trusted authority.

The package validates report shape, digest, trusted cluster, and trusted live session; it does not persist replay state or mutate deployment plans. Deployment report consumers must persist an atomic receipt keyed by the verified cluster, reporter session, service, and positive sequence. The report digest uses fixed-order JSON and excludes only the JWT identity and reportDigest. The consumer accepts the exact next sequence once, returns the prior response for a same-digest replay, rejects a different-digest replay, and ensures a plan revision CAS failure consumes neither the sequence nor a receipt. Timestamps are informational and never replace live session, placement, artifact, recipient, sequence, or plan-revision fences.

createWorkloadInitEnvironmentMap now rejects manifests without any launcher-environment delivery. Consumers must bypass WorkloadInit for file-only manifests.

Portable Storage Contracts

App Store templates can declare logical, template-local storageClasses and stable named storageRequests. The same manifest is fulfilled by Onebox or Cloudly without exposing a physical provider:

const storageConfig: appstore.IAppStoreVersionConfig = {
  image: 'example/database:1.0.0',
  port: 5432,
  storageClasses: {
    databaseFast: {
      kind: 'filesystem',
      purpose: 'database',
      required: {
        performanceTier: 'highIops',
        durability: 'persistent',
        hardQuota: true,
        snapshots: 'native',
        encryptedInTransit: true,
      },
    },
    backupCapacity: {
      kind: 'objectStorage',
      purpose: 'backup',
      required: {
        performanceTier: 'capacity',
        durability: 'persistent',
        hardQuota: true,
        encryptedInTransit: true,
      },
    },
  },
  storageRequests: [
    {
      id: 'database-data',
      kind: 'filesystem',
      storageClass: 'databaseFast',
      mountPath: '/var/lib/example',
      accessMode: 'ReadWriteOnce',
      capacity: { request: '20GiB', limit: '40GiB' },
      reclaimPolicy: 'retain',
      protection: { backup: 'required', snapshots: 'native' },
    },
    {
      id: 'backup-archive',
      kind: 'objectStorage',
      storageClass: 'backupCapacity',
      accessMode: 'readWrite',
      capacity: { request: '100GiB', limit: '1TiB' },
      reclaimPolicy: 'retain',
      delivery: {
        type: 'file',
        targetPath: '/run/secrets/backup-archive.json',
        format: 'servezone-object-storage',
        uid: 1000,
        gid: 1000,
        mode: 0o400,
      },
      protection: { versioning: 'required', retentionDays: 30 },
    },
  ],
  requiresFeatures: [
    appstore.appStoreStorageFeatureIds.bindings,
    appstore.appStoreStorageFeatureIds.filesystem,
    appstore.appStoreStorageFeatureIds.objectStorage,
    appstore.appStoreStorageFeatureIds.objectStorageFile,
  ],
};

Capacity quantities are positive integers followed by KiB, MiB, GiB, or TiB. Storage request IDs survive upgrades and restores. Logical class keys express requirements and preferences only; Onebox and Cloudly map them to operator policy independently.

An app may declare multiple objectStorage requests. Each request resolves to its own endpoint, bucket, and value-free credential management scope. Launcher environment delivery uses an explicit key map and file delivery uses one managed JSON Secret at a unique target path, so two bindings cannot share credential destinations accidentally.

platform.storage contains separate capability advertisements and resolved binding/status contracts. Resolved object-storage bindings expose connection metadata plus a service-owned credential management scope and delivery policy, never credential references or values. Filesystem bindings expose the container mount and access mode, never a host path.

A binding's requestDigest is stated by one function, so the control plane that issues it and the consumer that re-derives it compute the same value:

const requestDigest = await platform.createStorageRequestSha256(storageRequest);

It normalizes the request to its contract shape, canonicalizes it and returns bare lowercase hex over the request alone — no prefix, and nothing wrapped around the request. A body that is not a storage request is refused with platform.StorageContractError instead of hashed, and platform.canonicalizeStorageRequest returns the exact bytes the digest is taken over. platform.storageRequestCanonicalDigestGoldenVectors states both for a filesystem request and for both object-storage deliveries.

Which requests exist at all is stated by the same contract:

const storageRequest = platform.normalizeStorageRequest(declaredRequest);

platform.normalizeStorageRequest reads a body against the exact request schema and returns the request in its contract shape — required members present, optional members carried only when declared, canonical environment keys, no empty protection block, no explicitly undefined member, no -0 or unsafe integer, and plain JSON data members only. It is the value the other two functions are derived from: canonicalizing and digesting a request normalize it first, so a consumer that admits a request through this function and a control plane that digests it cannot accept different sets of requests. The input is never mutated and the returned request is a fresh object. A body that is not a storage request is refused with platform.StorageContractError, whose message names the failing member so a consumer can map the refusal to its own status vocabulary instead of restating the rules.

Portable manifests and resolved bindings intentionally have no fields for Synology, NFS, Kerberos, Corestore, Kubernetes, Docker drivers, servers, exports, mount options, provider credential values, or local fallback paths. Runtimes must reject unknown manifest fields and unsupported required feature IDs before provisioning. Legacy volumes and platformRequirements.s3 remain deprecated inputs for strict resolver normalization only. platformRequirements.mail declares that the runtime provisions a CoreMail binding for the instance and injects the platform-owned MAIL_COREMAIL_*, MAIL_FROM and SMTP_* identities; a template never declares mail credentials or sender addresses itself.

Runtime filesystem storage claims

data.IRuntimeStorageClaimIntent is the non-ready Cloudly-owned precursor to a claim. It fixes the controller epoch, organization, cluster, node, runtime namespace, service, claim ID/generation and predecessor, plus IRuntimeStorageClaimRequestedFilesystemBinding. That requested binding fixes the service, full normalized retained ReadWriteOnce filesystem request and its canonical bare digest. It structurally cannot state a provider binding ID/generation, policy selection, status, observed generation, resource reference or granted capability. The intent is exact-schema and digest-bound under serve.zone/runtime-storage-claim-intent; it grants no assignment, session, execution or attachment authority.

data.IRuntimeStorageClaim is the portable, Cloudly-issued authority for one ready, retained ReadWriteOnce filesystem binding on one cluster node. It names the organization, service, cluster, node, runtime namespace, controller epoch, binding ID/generation, bare canonical storage-request digest, opaque resource reference, policy class/revision, and canonical container mount path. The binding must report a persistent, single-node capability and matching observed generation. The claim carries no host path, physical inode, Docker option, credential, execution attempt, or assignment digest.

data.snapshotRuntimeStorageClaim makes a detached exact-schema snapshot; data.computeRuntimeStorageClaimDigest hashes its full content except digest under the serve.zone/runtime-storage-claim domain; and data.validateRuntimeStorageClaim verifies the stated digest. Cloudly seals the claim's { id, generation, digest } in IRuntimeAssignment.workload.storage. At admission, data.bindRuntimeStorageClaimToAssignment(claim, assignment, expectedClusterId) checks that reference and the exact controller, organization, service, node, namespace and cluster. The cluster ID must come from the authenticated cluster scope, because an assignment does not contain one. These functions prove content consistency only: the caller must authenticate Cloudly, check current provider readiness and policy, durably admit the assignment, and obtain a fenced native attachment from the storage owner before a CRI mount. An absent claim or an unverified opaque resource reference never authorizes a mount.

data.bindRuntimeStorageClaimToIntent(claim, intent) verifies both canonical digests and requires the ready claim to preserve every Cloudly-owned identity and requested filesystem fact: service, request ID/digest, mount path, access mode and reclaim policy. Only the full claim may add the storage provider's binding identity/generation, selected policy, matching observed generation, ready status, opaque resource reference and persistent single-node capabilities. This is a content join: callers still authenticate Cloudly and the provider, check current readiness and retain the native fence.

Runtime local storage claims

data.IRuntimeLocalStorageClaim is a controller's authority for one retained volume of one service, mounted on one node. Cloudly and Onebox issue it alike (controller.kind is cloudly or onebox). It names the claim revision (id, generation, previous), the organization, cluster, node, runtime namespace, service and the service's own volumeId, and it states:

  • source (TRuntimeLocalStorageSource), one of:
    • { kind: 'directory' } (TRuntimeLocalStorageDirectorySource), a node-local directory the node derives from the claim's identity under a root only the node knows;
    • { kind: 'nfs', server, export, subPath, version } (TRuntimeLocalStorageNfsSource), the directory subPath below the NFS export server:export;
    • { kind: 'smb', server, share, subPath, version, credential } (TRuntimeLocalStorageSmbSource), the directory subPath below the SMB share //server/share. credential is an IRuntimeAuthorityReference { id, generation, digest } naming the credential the node authenticates with; the claim never carries a user name, password or domain, and the node resolves the material through the controller's secret delivery. NFS authenticates the node's host, so an nfs source names no credential;
  • initialization: empty creates an empty root with initialOwnership { uid, gid, mode }, which an empty claim must state, so a workload without runAs never meets a root-owned directory by accident; import waits for the volume's bytes to be imported and states initialOwnership: null, because an imported tree keeps its own ownership. mode holds directory permission bits only (runtimeLocalStorageClaimContract.directoryModeBits, 0o3777; setuid is refused). SMB carries no POSIX ownership, so the node presents an SMB volume with initialOwnership, and an smb claim is always empty;
  • mount { mountPath, readOnly }, judged by the storage request's own mount path rule: canonical, absolute, not /, and never overlapping /run/secrets;
  • accessMode: 'ReadWriteOnce', reclaimPolicy: 'retain', and lifecycle: bound keeps the volume, purge orders its bytes deleted and always follows a bound generation of the same claim.

A network source names the remote directory and nothing about how to mount it, so the node derives every mount parameter from these fields alone and a claim cannot smuggle an option in:

Field Rule
server a lowercase fully qualified hostname of two or more labels whose last label is not numeric (a single label would resolve through each node's own search domains), or a canonical IP literal (a dotted quad without leading zeros, or RFC 5952 IPv6 without brackets or zone). No scheme, user info or port: the node connects to the protocol's standard port
export /, or / and canonical segments without a trailing slash
subPath . for the export or share itself, or canonical relative segments
share one segment of at most 80 characters, without +, ;, [ or ]
version nfs: 3, 4.0, 4.1 or 4.2; smb: 3.0, 3.02 or 3.1.1

A segment is printable ASCII without a space, at most 255 characters, never . or .., and holds none of ", *, ,, /, :, <, =, >, ?, \ or |: a comma or = would read as a mount option, and the rest have meaning in a mount source or an SMB name. export and subPath are at most 1024 characters. ASCII keeps one spelling per segment, but one remote directory can still be named several ways: a hostname and an address of the same server, a different export/subPath split, or a nested subPath. The contract cannot detect these aliases, so the controller must refuse overlapping bound claims itself. The versions are the kernel's own vers= spellings, and each pins one protocol version: a bare NFS 4 or SMB 3 negotiates one, which a claim that must mean the same mount on every node cannot allow. NFS version 2 and the SMB 1 and 2 dialects are refused; SMB 3 is the first dialect that encrypts. The bounds and version lists are runtimeLocalStorageClaimContract.nfsVersions, smbVersions, maximumServerLength, maximumPathLength, maximumSegmentLength and maximumShareLength.

ReadWriteOnce holds on the node that mounts the claim. A network directory is reachable from more than one node, so the controller issues at most one bound claim per remote directory; a node cannot see another node's mounts. A purge of a network claim deletes everything below subPath, the whole export or share for ..

The claim carries no host path, device, mount option or credential value. snapshotRuntimeLocalStorageClaim makes a detached exact-schema snapshot, computeRuntimeLocalStorageClaimDigest hashes it under serve.zone/runtime-local-storage-claim, validateRuntimeLocalStorageClaim verifies the stated digest, and getRuntimeLocalStorageClaimReference is the { id, generation, digest } an assignment seals in workload.storage. bindRuntimeLocalStorageClaimToAssignment(claim, assignment, expectedClusterId) checks that reference and the exact controller, organization, node, runtime namespace, service and cluster; a purge claim binds to no assignment, because no workload may mount a volume being deleted.

requests.runtimesession.IReq_Controller_Pallet_ApplyRuntimeLocalStorageClaim (applyRuntimeLocalStorageClaim) pushes one claim over the node's authenticated session: { session, claim } → { status, state, claim }. status is accepted, replay or historical; state is ready, import-required, purged or uncertain (a volume the node can vouch for neither way, which it neither mounts nor deletes until an operator decides); claim is the sent claim's reference. bindApplyRuntimeLocalStorageClaimRequest(request, trustedSession) requires the request's session to equal the node's current one and the claim to name its controller, cluster, node and runtime namespace. bindApplyRuntimeLocalStorageClaimResponse(response, request) binds the answer to the sent claim and, for accepted and replay, to a state the claim can be in: a bound claim is ready, import-required (an import claim only) or uncertain, a purge is purged or uncertain. A historical answer describes the newer generation the node holds.

requests.runtimesession.IReq_Any_Cloudly_GetRuntimeLocalStorageClaims (getRuntimeLocalStorageClaims) lets an administrator read the claims Cloudly issued in one cluster, optionally on one node: { identity, clusterId, nodeId?, cursor?, limit? } → { claims, nextCursor? }. Each entry is an IRuntimeLocalStorageClaimSummary { claimId, generation, clusterId, nodeId, serviceId, volumeId, mountPath, sourceKind, initialization, lifecycle, state } of the current generation. claimId is the claim's id, the name the node keys the volume by and the import/<claimId> directory an operator places an import volume's bytes in; a reader takes it from here and never derives it, because how a controller composes its claim ids is the controller's own business. state is what the node answered about exactly that generation, or null until the node acknowledged it, and it obeys the same rule as an apply answer. A summary names the source by kind alone and carries no digest, controller epoch or SMB credential reference.

The read is paged: cursor is the opaque nextCursor of the previous page, limit is at most runtimeLocalStorageClaimContract.maximumSummaryPageSize (200), the page is in strictly ascending claimId order, and nextCursor is omitted on the final page and never stated on an empty one. summarizeRuntimeLocalStorageClaim(claim, state) is the one projection of a claim into a summary, snapshotRuntimeLocalStorageClaimSummary and snapshotGetRuntimeLocalStorageClaimsRequest / snapshotGetRuntimeLocalStorageClaimsResponse check exact schemas, and bindGetRuntimeLocalStorageClaimsResponse(response, request) binds a page to the read that asked for it: every summary names the requested cluster and node, and the page holds no more than the requested limit.

platform.storagemigration defines the provider-neutral cutover contract for a named object-storage binding. Corestore atomically owns and fences the source binding after validating a distinct, unfenced active-object snapshot. Onebox only receives a held candidate, stages that exact candidate, stops the matching workload generation, and attests the quiesced state. The candidate cannot start until destinationBindingStartAuthorized is true. Corestore continues returning consumerAction: 'startDestination' until Onebox submits the mutation-fenced consumer activation request and the acknowledgement becomes durable evidence. Candidate-issued abort tombstones authorize only startSource and never retain the staged candidate binding.

Every migration DTO and status has an exact, versioned runtime normalizer. Unknown fields, provider or pool identifiers, unbounded strings, unsafe integers, stale mutation revisions, identity drift, and lifecycle-inconsistent fields are rejected. Migration-created digests use strict canonical JSON and a bare lowercase 64-hex SHA-256 value; portable golden vectors cover the source snapshot, target request, prepare intent, and candidate binding. Pre-cutover failures may retry or abort; after the durable commit point, recovery can only retry or roll forward. Physical pool IDs, mount details, provider receipts, and publication capabilities remain private.

Primary exports include:

  • IObjectStorageMigrationPrepareRequest and TObjectStorageMigrationStatus for the immutable intent and status journal.
  • IObjectStorageMigrationConsumerQuiesceRequest with IStorageMigrationConsumerQuiesceEvidence for exact candidate staging and source-workload shutdown.
  • IObjectStorageMigrationConsumerActivationRequest with IStorageMigrationConsumerActivationEvidence for durable destination-start acknowledgement.
  • normalizeObjectStorageMigrationStatus, bindObjectStorageMigrationConsumerQuiesceRequest, and bindObjectStorageMigrationConsumerActivationRequest for strict ingress and current-revision mutation fencing.
  • The create*Sha256 helpers for source snapshots, prepare intent, target requests, candidate/active bindings, and persisted staging or activation evidence.

The phase and consumer-action progression is exact:

Phase consumerAction Destination start authorized
preparing, transferring wait No
awaitingConsumerQuiesce stageCandidateAndStop No
finalizing, committing wait No
readyToStart startDestination Yes
cleanupPending, complete none Yes; durable activation evidence is required
aborting wait No
aborted startSource No; only the active source binding may restart

Normalize every status before acting, and bind consumer mutations to that exact status revision:

const status =
  await platform.storagemigration.normalizeObjectStorageMigrationStatus(
    untrustedStatusPayload,
  );

if (status.phase === 'awaitingConsumerQuiesce') {
  const request =
    await platform.storagemigration.bindObjectStorageMigrationConsumerQuiesceRequest(
      untrustedQuiescePayload,
      status,
    );
  await submitQuiesceAcknowledgement(request);
}

if (status.phase === 'readyToStart') {
  if (
    !status.destinationBindingStartAuthorized ||
    status.consumerAction !== 'startDestination'
  ) {
    throw new Error('destination binding is not authorized to start');
  }
  await startWorkload(status.activeBinding);
  const request =
    await platform.storagemigration.bindObjectStorageMigrationConsumerActivationRequest(
      untrustedActivationPayload,
      status,
    );
  await submitActivationAcknowledgement(request);
}

if (status.phase === 'cleanupPending') {
  // Cleanup is reachable only after this durable acknowledgement was accepted.
  const durableActivation = status.consumerActivationEvidence;
}

if (status.phase === 'aborted') {
  if (
    status.destinationBindingStartAuthorized ||
    status.consumerAction !== 'startSource'
  ) {
    throw new Error('invalid aborted migration status');
  }
  await startWorkload(status.activeBinding);
}

Here startWorkload, submitQuiesceAcknowledgement, and submitActivationAcknowledgement are consumer-owned operations, not package exports. Canonical digests are produced from normalized payloads:

const snapshotSha256 =
  await platform.storagemigration.createUnfencedObjectStorageBindingControlSnapshotSha256(
    snapshotDigestPayload,
  );
const migrationSha256 =
  await platform.storagemigration.createObjectStorageMigrationSha256(
    prepareRequest,
  );
const candidateSha256 =
  await platform.storagemigration.createObjectStorageMigrationBindingSha256(
    candidateBinding,
  );
const activationRecordSha256 =
  await platform.storagemigration.createObjectStorageMigrationPersistedActivationSha256(
    activationDigestPayload,
  );

Data Contracts

Use data when you need object shapes that are persisted, exchanged between services, or exposed through the Cloudly API.

import { data } from '@serve.zone/interfaces';

const service: data.IService = {
  id: 'service-api',
  data: {
    name: 'api',
    description: 'Public API service',
    imageId: 'image-api',
    imageVersion: '1.0.0',
    environment: {
      NODE_ENV: 'production',
    },
    serviceCategory: 'workload',
    deploymentStrategy: 'limited-replicas',
    scaleFactor: 2,
    balancingStrategy: 'round-robin',
    targetPorts: [
      {
        name: 'web',
        port: 3000,
        protocol: 'http',
        default: true,
      },
      {
        name: 'ssh',
        port: 2222,
        protocol: 'ssh',
      },
    ],
    ports: {
      web: 3000, // legacy compatibility shorthand during migration
    },
    domains: [
      {
        name: 'api',
        protocol: 'https',
        targetPort: 'web',
      },
    ],
    publicPortMappings: [
      {
        name: 'ssh-public',
        publicPort: 2222,
        targetPort: 'ssh',
        protocol: 'tcp',
        exclusive: true,
      },
    ],
    deploymentIds: [],
  },
};

Common data contracts include:

  • ICluster and IClusterNode for cluster membership and provisioning state, including the optional ICluster.data.relay block described below.
  • clusternodehostname for the static hostname Cloudly wants a node to carry: IClusterNodeHostnameIntent on IClusterNode.data.hostnameIntent, the default scheme cloudly-worker-<name> (clusterNodeHostnameContract, deriveDefaultClusterNodeHostname, proposeClusterNodeName) and the named clusterNodeHostnameRefusals of setClusterNodeHostname and resetClusterNodeHostname.
  • IService, IDeployment, IImage, IRegistryTarget, and IExternalRegistry for workload delivery.
  • Service port contracts including IServiceTargetPort, IServiceDomainRoute, and IServicePublicPortMapping for canonical backend targets, domain target references, and edge/Coretraffic TCP/UDP public exposure.
  • IDomain, IDnsEntry, and traffic contracts for routing and DNS management.
  • Traffic and gateway route contracts including ICoretrafficPortRouteConfig, routing portRoutes, and IGatewayClientRoute client-owned route views. Gateway route intent supports optional match domains, transport, and remoteIngress, plus explicit route priority and managedRouteKind. Ownership can combine hostname with routeRef so a normal route and a path-specific managed route for the same hostname reconcile independently.
  • Value-free ISecretMetadata, ISecretVersionMetadata, and ISecretSetMetadata contracts for operator views, plus exact-version IResolvedSecretManifest contracts for cluster delivery. Manifest helpers bind immutable image rollout and invocation evidence, enforce canonical digests and globally unique launcher/file targets, and track per-cluster desired/applied/previous-accepted rollout state. Platform-provider and system owners are valid metadata owners but are rejected from workload manifests. listSecrets exposes the dedicated targetSecretsRevision CAS fence; every create, rotate, and lifecycle mutation consumes and returns that aggregate owner fence, while setServiceSecretSetAttachments returns the independent secretConfigurationRevision. Generic service writes own neither revision. Purge is a separate exact-version mutation with its own version revision and durable operation; it is not a logical-secret lifecycle action.
  • Mail gateway contracts for domain authorities, address bindings, WorkApp bindings, managed SMTP/API credentials, spool items, delivery journals, and inbound/outbound message payloads.
  • Service-level mail configuration through IService.data.mail, including per-address inbound smtpForward settings and outbound credential metadata. Cloudly settings include dcrouter gateway, SMTP submission, and inbound forward-target keys for reconciling those bindings.
  • Web Push contracts for environment-specific service bindings, public credential state, public VAPID key rotation metadata, privacy-minimal notification signals, and redacted delivery state. Subscription endpoints, browser key material, provider ciphertext, VAPID private keys, and credential secrets are intentionally absent from public binding and status DTOs.
  • Service-level Web Push declaration through IService.data.webPush. Immutable deployment declarations can require the pushnotification platform capability alongside database and object-storage capabilities; this does not turn Web Push into a Corestore resource or volume capability.
  • IUser, JWT-only IIdentityCredential, full IIdentity, and token-related contracts for authentication context. IIdentity extends IIdentityCredential with server-issued user metadata.
  • ICloudlyConfig, ICloudlySettings, status, server, bare-metal, BaseOS, backup, and task execution interfaces for control-plane state.

Cluster Relay Block

ICluster.data.relay is the optional IClusterRelay { origin: string } that states how cluster-side components reach this cluster's relay. It is absent on a cluster whose nodes still target Cloudly directly, so its presence is what marks a cluster as relayed.

origin is the exact value a Pallet node stores as its relayOrigin — the origin its runtime session registers over, as described in Authenticated Runtime Sessions. It must be an https: origin with no path, query, fragment, userinfo or trailing slash, so that new URL(origin).origin === origin; an explicit port is allowed, and because URL drops a default port and lowercases the host, https://relay.example:443 and https://Relay.example are not canonical origins. validateClusterRelay(value, path?) returns one named reason per violation and an empty array for a valid block.

The relay's certificate subject is not a second field: every consumer derives it from the origin with clusterRelayCertificateDomainName(relay), which returns the origin's hostname — the name Cloudly issues the certificate for and owns the A record for. An origin naming an address literal therefore has no issuable certificate.

data.clusterRelayUndispatchedRefusals names the five registration refusals that end a relay's custody — relay-cluster-not-enabled, relay-node-foreign, relay-node-unknown, relay-sequence-not-advancing and relay-credential-generation-stale — with data.TClusterRelayUndispatchedRefusal as their union. relay-node-foreign names a node of another cluster; relay-node-unknown names a node Cloudly does not hold at all, because no Spark enrollment has created it yet, so the node a relay's COREFLOW_NODE_ID names is enrolled before the relay registers. Each is Cloudly reading its own records and stating that it does not dispatch to this relay, so a relay that hears one drops what it holds for its nodes instead of waiting. A refusal that only says Cloudly could not verify the relay states nothing about dispatch and is deliberately not in the list.

const relay: data.IClusterRelay = {
  origin: 'https://relay.cluster-a.serve.zone:8443',
};

const reasons = data.validateClusterRelay(relay); // []
const certificateDomainName = data.clusterRelayCertificateDomainName(relay);
// 'relay.cluster-a.serve.zone'

Enabling a relay is one operator flip per cluster. enableClusterRelay { identity, clusterId } answers with the block Cloudly published, and disableClusterRelay removes it and answers with an explicit relay: null. The caller names only the cluster: Cloudly derives the origin itself from that cluster and its own public hostname, and refuses the request when no DNS zone it manages matches the derived name, so a relay name that nothing can publish is never written. Both requests push the cluster configuration to the connected relay afterwards, so a relay learns the change without reconnecting.

Once its socket is authenticated with registerCloudlyClientSession, the relay sends registerClusterRelay carrying IRegisterClusterRelayRequest: the relay build, the cluster node it runs on, the exact endpoint it bound for its cluster's Pallet nodes, its protocol offer, and a registrationSequence that counts up once per registration within one relay process and restarts at 0 when that process restarts. validateRegisterClusterRelayRequest(value, path?) returns one named reason per violation, like every validator here, and checks shape only; it validates the offer through the one grammar protocol owns. Cloudly reads and negotiates that offer before it takes the verified cluster connection or decides the sequence, so a refused registration consumes no sequence, and the accepted answer carries Cloudly's own offer beside the origin. relay.version stays the relay build — a different fact from protocol.interfacesVersion, the contract it speaks. listenAddress is an IPv4 literal because Cloudly publishes it verbatim as the A record of the relay's certificate domain name — a hostname would need a CNAME and an IPv6 address an AAAA record, so admitting either is an additive change once Cloudly publishes it. A loopback or link-local address is refused by its own name, because a relay that bound one can never serve the nodes of its cluster.

Authority stays with Cloudly, and none of it is a request field. The cluster is the one the verified machine identity on that socket names; whether relay.nodeId is a node of that cluster is decided against Cloudly's own records; and one cluster has one relay. A registration for a cluster whose relay nobody enabled is refused by name, because there is no origin to register into, so an accepted registration always answers with the origin Cloudly published: that is how a relay checks that the certificate it holds and the name its nodes are told to reach are the same one.

registrationSequence restarts with the relay process, so it orders registrations only within one live peer: there it must strictly increase, a refused registration does not consume it (the same body may be re-sent once the cause is gone), and a retry that repeats a sequence Cloudly already accepted is refused. Across peers it says nothing, because a restarted relay counts from zero again; there Cloudly's own acceptance time orders the registrations, and a registration arriving on a live verified socket supersedes the stored one.

An administrator reads the result with getClusterRelayRegistrations { identity, clusterIds? }, answered with IClusterRelayRegistration per cluster: the cluster, the node, the listen address and port, the relay version, the sequence that relay counted, Cloudly's acceptedAt — and live, whether a verified socket carries that relay at this instant. live is a liveness fact and grants nothing: a registration that is perfectly current reads live: false while its socket is gone, and a cluster whose relay never registered is absent from the answer rather than present with an empty one. validateClusterRelayRegistration(value, path?) returns one named reason per violation, as every validator here does.

const registration: data.IRegisterClusterRelayRequest = {
  protocol: protocol.createProtocolOffer('32.0.0'),
  relay: {
    version: '4.0.0',
    nodeId: 'node-a',
    listenAddress: '10.0.4.7',
    listenPort: 8443,
  },
  registrationSequence: 0,
};

const registrationReasons = data.validateRegisterClusterRelayRequest(registration); // []

How the relay reaches each node's Corestore is part of the same document: see Corestore Through The Relay.

Cluster Relay Credential

A cluster's relay authenticates with its own generation-bearing credential instead of the shared cluster authority every cluster-side component used before it. IClusterRelayCredential is public metadata: the cluster, a generation that advances by one per accepted mint or rotation, the machine userId it authenticates as, and delivery, the ISecretVersionReference of the immutable encrypted version the bearer was delivered as. No record, receipt, request or response of this contract carries a bearer, a hash or key material, except the one answer that exists to hand the bearer over (see the export below), and every reader is exact-keyed, so a peer that adds a secret-looking member is refused rather than copied.

The credential authenticates the cluster's whole control-plane session, exactly what the shared authority authenticated, so a relay changes one credential and keeps every call it makes. The gain is a generation Cloudly fences each socket against and an authority it can rotate and retire on its own; the scope is not narrower. One cluster has one relay and therefore one credential record, and it never expires by itself: rotation is its lifecycle.

mutateClusterRelayCredential { identity, mutation } takes TClusterRelayCredentialMutation. ensure requires expectedGeneration: null and an absent record; rotate requires the exact current generation, never a minimum. The caller retains the complete request and its mutationId across retries: the identical body answers with the original result and replayed: true without renewing delivery, generation or authority, and a changed body fails closed. The first mint also records the shared authority the cluster still carries as IClusterLegacyRelayCredential, so a replacement never exists without a named predecessor. Rotation has no overlap: the previous generation stops authenticating at commit, and a relay still holding it hears relay-credential-generation-stale, drops what it holds for its nodes, reads the delivered version and reconnects. There is no revocation of the relay credential, because a rotation already covers a compromised one without a state in which the cluster has no relay authority at all.

An accepted registerClusterRelay answers with two references beside the origin and the offer: credential, the generation Cloudly authenticated the socket as, and acceptedRegistration, the record that acceptance wrote. The relay states no generation of its own — the registration request is unchanged — because a relay naming its own generation would be asserting its own rank. IClusterRelayAcceptedRegistration is that record: the registration facts the served read shape also carries, its own id, a generation that advances once per acceptance, the credential reference and a digest. Liveness is not part of it. IClusterRelayRegistration keeps exactly its published members, the facts plus live.

retireLegacyClusterRelayCredential { identity, retirement } revokes the shared authority, and is the last step of replacing it. IClusterLegacyRelayCredentialRetirementRequest carries compare-and-swap inputs only: the exact expectedLegacyGeneration, the replacement credential reference and the acceptedRegistration reference the caller observed. Cloudly dereferences both against its own current records, requires the registration to have been made under that replacement, and requires a verified live socket carrying it; a caller can state no liveness or proof of its own. Acceptance writes the one-shot IClusterLegacyRelayCredentialInvalidation (generation: 1) and advances the legacy record to state: 'revoked' naming it. The invalidation names the legacy record at the generation before the revocation and the record names the invalidation after it, so neither digest covers the other's. data.clusterRelayCredentialRefusals names every other answer, with data.TClusterRelayCredentialRefusal as their union. getClusterRelayCredential { identity, clusterId } reads the chain: the credential, the legacy record and the current accepted registration, each null while no record exists.

Each digest-bearing record has a snapshot…, compute…Digest and validate… function under its own digest domain in data.clusterRelayCredentialContract and data.clusterRelayContract. bindClusterLegacyRelayCredentialInvalidation(invalidation, legacyCredential, acceptedRegistration) joins an invalidation to the legacy record it revoked — the active record it names or that record's revoked successor — and to a registration whose credential equals the invalidation's replacement. It is a content check; dereferencing each record as current and requiring the live socket remain Cloudly's.

exportClusterRelayAuthorization { identity, clusterId, expectedGeneration } answers the bearer itself, so an operator hands it to the relay without an exec into Cloudly's process. Cloudly mints nothing to answer it: the bearer is the encrypted version the credential's delivery names, read at exactly expectedGeneration (never a minimum) and proved against the hash that generation issued, so a rotation committing meanwhile is refused rather than answered. A new bearer is a rotate followed by an export at the new generation. It requires a freshly authenticated administrator, exactly as a credential mutation does. IClusterRelayAuthorizationExport carries the sealed credential at that generation, the authorization (the bearer, byte for byte what the relay presents: clusterRelayToken_<credential id>_<64 hex>, no line terminator) and the receipt. It is the only answer of this contract that carries a bearer; Cloudly never logs it, and its reader writes it where the relay reads it and nowhere else.

Every answered export is audited: Cloudly writes IClusterRelayAuthorizationExportReceipt (id, clusterId, the exact credential reference, the administrator's actorId, exportedAt) before it answers, so no bearer leaves without a record, and a retried export is a second receipt rather than a replay. A receipt is value-free: no bearer, no hash. getClusterRelayAuthorizationExports { identity, clusterId } answers a cluster's receipts, oldest first. Refusals are data.clusterRelayAuthorizationExportRefusals (data.TClusterRelayAuthorizationExportRefusal): relay-credential-absent, relay-credential-generation-mismatch, relay-authorization-missing and relay-authorization-ambiguous. A reader checks an answer with bindClusterRelayAuthorizationExport(answer, { clusterId, expectedGeneration }) before it writes a byte: the credential's digest holds, it is the requested cluster's at the requested generation, the bearer names that credential, and the receipt names that exact credential. Every refusal is the contract's one value-free message and never quotes the answer. isClusterRelayAuthorization states the bearer's shape, and snapshotClusterRelayAuthorizationExportRequest, snapshotClusterRelayAuthorizationExportReceipt and snapshotClusterRelayAuthorizationExport are the exact-keyed readers.

All five requests read only the verified IIdentityCredential, never an identity member a caller states about itself.

Spark Legacy Runtime Retirement

The Spark-owned half of IFleetCutoverCoreflowRetirementEvidence — the refused old mode, the configuration compare-and-swap, and the bundle and unit identity — is transported on the Swarm observation a node already posts. There is no route, request or endpoint for it. That is forced rather than preferred: the evidence requires relay.nodeObservation and spark.observation to be the identical accepted reference, so these facts must reach Cloudly on the very observation that carries the node's Docker census.

data.ISparkLegacyRuntimeRetirement is the optional observation member, available in every observation state because it describes the host and not its Docker daemon. Accepting it is state-independent; completing the evidence is not. IFleetCutoverCoreflowRetirementEvidence completes only from a member observation of the relay host, because relay.nodeObservation and spark.observation must be the identical accepted reference and that reference has to carry the relay host's local Docker snapshot, which no other observation state can state. It is { refusedMode, bundle, configuration } and nothing else. It names no node: Cloudly binds it to the reporter it already authenticated, so there is no node identity on it to disagree with. It carries no digest of its own either — computeSparkSwarmObservationDigest canonicalizes the whole observation, so a second domain here would only be a second thing to keep in step.

refusedMode is one of data.sparkLegacyRuntimeModes, the two runtimes a released Spark could be put into by its mode configuration key. data.sparkLegacyRuntimeModeServices is the record over that union naming the Docker services each mode owns, so a refusal states a mode and a reader derives the set. data.ISparkNodeBundleIdentity { version, sourceCommit, manifestSha256, unitName, unitDigest } is read back from the installed tree and the installed unit file, never from the running process's own build constants: a build can state its own version, it cannot state that the host will start that same code again. data.ISparkLegacyConfigurationState { ownerGeneration, sourceSha256, disposition } names the released key-value source by digest alone. No decoded member of that file is representable in any shape here, because it carries the node's released bearer and its jump code; sourceSha256 covers the whole file, whose smallest unguessable part is a 256-bit token. disposition is retained-inert or removed — the two of data.sparkLegacyConfigurationDispositions, with data.TSparkLegacyConfigurationDisposition as their union; erasure is not a cutover requirement.

Cloudly accepts an observation first and then advances two of its own records, which is what the evidence's two references point at. data.ISparkLegacyModeRefusal and data.ISparkLegacyConfigurationRetirement are revisions, not one-shot receipts: a node restates both on every observation, a bundle upgrade changes what the boot path is, and a disposition changes once. Both carry the accepted observation they were derived from, its reporter session and its sequence, and both obey the revision rule every record of this package obeys — a first generation has no predecessor and a later one names exactly the generation before it. The configuration record carries two counters on purpose: its own generation is contiguous because it is this record's revision, while configuration.ownerGeneration is the node's own and may skip, since an observation that never arrived must not look like a skipped revision.

advancesSparkLegacyConfigurationState(current, next) is the compare-and-swap rule, null for a node's first accepted state: the same sourceSha256, a strictly higher ownerGeneration, and no return from removed to retained-inert. Restating the accepted generation is a replay and therefore not an advance, which is what keeps acceptedAt and the named observation from aging forward on a node that has stopped changing. data.sparkLegacyRetirementRefusals names every other answer, with data.TSparkLegacyRetirementRefusal as their union. bindSparkLegacyRetirement(modeRefusal, configurationRetirement) joins the two receipts a retirement cites, requiring the same cluster, node, reporter session, sequence and accepted observation. It is a content check; whether each receipt is its node's current revision remains Cloudly's.

The answer travels on the observation's own accepted receipt, as ISparkSwarmObservationAcceptedReceipt.fleetCutoverLegacyRetirement, a data.TSparkLegacyRetirementOutcome: { state: 'accepted', generation } for the revision this observation advanced both records to, { state: 'restated', generation } for the replay that renewed nothing, and { state: 'refused', refusal, retryable: false } naming one of the refusals above. A refused statement does not reject the observation. The statement is judged against records rather than against the message — a malformed one is already refused by validateSparkLegacyRuntimeRetirement — so everything left is a fact about this node's history, while the observation's liveness and census are true either way and the cluster's target set depends on their freshness. A refusal is never retryable: it answers one exact statement, so resending it cannot change the answer, and a genuine server-side race is the observation's own acceptance-race one level up. The flag is carried anyway, so a node meeting a refusal name a later minor added still learns it must stop. validateSparkSwarmObservationResponse binds the answer to the statement in both directions: an accepted receipt that answers a statement the observation never made, or leaves one it did make unanswered, is refused. A rejected receipt carries no outcome at all, because an observation that was not accepted never reached the statement. data.snapshotSparkLegacyRetirementOutcome and data.validateSparkLegacyRetirementOutcome read it exact-key per state, like every other shape here.

An omitted member is not a refusal and advances nothing, so the existence of these records is not what makes a retirement honest — the freshness of the observation they name is. Assembly must require that observation to still be inside fleetCutoverContract.observationFreshnessMs, exactly as IFleetCutoverSparkObservationReference already does, or a node that stopped refusing would go on looking retired.

Cloudly reads this member before Spark writes it. Every observation validator in this package is exact-key, so a server built against an earlier minor refuses the whole observation rather than ignoring a member it does not know — the peers of a session kind are upgraded before the side that starts depending on the later minor, and the minimumPeerVersion of that session kind is raised in the same commit that starts depending on it. A Spark that begins stating this member therefore raises its sparkSwarmNode minimum alone, which requires one minimum per session kind rather than one shared across a deployable's offers. The duty runs both ways: a receipt is read exact-key too, so a Cloudly that begins answering with fleetCutoverLegacyRetirement raises its own sparkSwarmNode minimum in the same commit, or an older Spark refuses the receipt for a member it does not know.

Cluster VPN Hub And Network

A cluster's relay is also that cluster's managed VPN hub: every node of the cluster dials it, and an inter-node workload packet is relayed inside the cluster instead of crossing the control plane. Cloudly keeps every authority — it compiles the network and issues each node's credential — and carries no packet.

A relay that bound a hub says so in its registration, in relay.vpnHub: data.IClusterRelayVpnHub { publicKey, address, quicPort }. It reports what it bound, exactly like listenAddress and listenPort: address is an IPv4 literal inside the address plan's vpn.hubPrefix and quicPort is the port it bound for the one transport a managed hub serves. publicKey is the hub's Noise public key in the one canonical 32-byte base64 encoding every Noise key of this contract is read with — this advertisement, the managed VPN credential's keys and a cluster network's nodes all pass the same reader. A relay names no endpoint id of its own: Cloudly composes the endpoints with runtimeNetworkVpnHubEndpoints (see Runtime Network Address Plan). The member is optional — a relay that serves no hub registers exactly as before, and Cloudly selects no hub for that cluster, stating relay-vpn-hub-absent. validateClusterRelayVpnHub(value, path?) returns one named reason per violation.

data.IClusterVpnNetwork { authorityId, controlPrefix, revision, nodes, grants } is the whole managed VPN of one cluster. It is never a patch: a node or a grant the network does not carry is withdrawn, and membership is presence, so there is no enabled flag. controlPrefix is the subnet the cluster's hub serves its control plane on — the address plan's vpn.controlPrefix, read by the same runtimeNetworkVpnControlPrefix reader. A hub fixes that subnet when it binds and no other carrier hands a relay the address plan, so the network it is pushed states it; the plan contract never lets it change, which is why a network that moves the prefix under the same authority is refused as vpn-network-authority-foreign rather than by a name of its own. Each node states { nodeId, publicKey, controlAddress, expiresAt, workloadPrefixes }, where expiresAt is that node's credential expiry in Unix milliseconds and each prefix is { cidr, policyDomain }. Each grant is one directed { sourceDomain, destinationDomain } pair. clusterVpnControlPolicyDomain(address) and clusterVpnWorkloadPolicyDomain(leaseDigest) state the domain names once for both sides.

snapshotClusterVpnNetwork requires sorted, unique members, prefixes and grants, requires every control address to lie inside controlPrefix and every workload prefix to stay clear of it — the two rules the hub enforces against the subnet it bound — and requires every grant to join two workload domains the network itself carries, on two different nodes. That is what a per-cluster hub means: a grant whose other end sits on another cluster's hub cannot be expressed at all, which is why Cloudly refuses it while it compiles, by name, with grant-cross-cluster-unsupported from data.clusterVpnNetworkSelectionRefusals. clusterVpnNetworkContract bounds a network at 256 nodes, 128 workload prefixes per node, 4096 grants and 917 504 canonical bytes, and authorityId at clusterVpnNetworkContract.maximumAuthorityIdBytes — 128 UTF-8 bytes, the bound the hub daemon itself holds for the authority it binds. Bytes rather than code points, because that is what the daemon counts. The same bound is read wherever the id is admitted, snapshotClusterVpnNetwork and the managed VPN credential's authorityId alike, so an id no hub could take over is refused where the network is compiled instead of arriving back as a hub reporting state: 'failed'.

applyClusterVpnNetwork { network } is Cloudly's push over the relay's already registered session; the relay answers { appliedRevision, state }, where state is idle, applying or failed and idle means the pushed revision is the applied one — which is what bindApplyClusterVpnNetworkResponse(response, sentRequest) checks. admitClusterVpnNetwork(held | null, next) is the pure decision the relay makes against what its hub holds: accepted for a higher revision or an empty hub, replay for a repeat, and refused with vpn-network-revision-stale or vpn-network-authority-foreign from data.applyClusterVpnNetworkRefusals. getClusterVpnNetwork is the pull the other way, with no request body — the relay's socket names the cluster — answering { network } or { network: null } while the cluster has none. A hub daemon starts empty and its lifetime is the relay's own, so a relay that just bound one pulls instead of waiting for the next push.

Cluster Ingress

A cluster's ingress terminates TLS for the hostnames its services publish and forwards to the host ports Pallet published for the workloads behind them. Cloudly computes what it serves — it is the only party that knows every service, every ready endpoint, the node address each published port answers on and the certificate each hostname needs — and the cluster relay carries that table to the ingress and the ingress's counters back. These contracts carry no version suffix and no schema field: the protocol version is the installed @serve.zone/interfaces version.

IClusterRouteTable { revision, issuedAt, routes, portRoutes } is the whole table, never a delta. revision is its only idempotency: a relay and an ingress apply a table whose revision is greater than the one they hold and ignore anything else, so a repeated push, a replayed push and a push that crossed a newer one all settle the same way. IClusterHttpRoute names one hostname, its destinations ({ address, port }, unicast IPv4 and a real port), the certificate that terminates it and optional Basic authentication; IClusterPortRoute forwards a TCP or UDP listen port without terminating anything. validateClusterRouteTable(value, path?) returns one named reason per violation, refuses two routes for one hostname and two port routes for one transport and port, and — because a validation reason travels into logs — never quotes certificate material or a password in a reason.

The table is secret-bearing end to end: it carries the private key of every hostname the ingress serves. Exclude it from hooks, logs, journals and disk on every hop, exactly as a runtime session credential is excluded.

Method Direction Body
pushClusterRouteTable Cloudly → relay { table } → { appliedRevision }
getClusterRouteTable relay → Cloudly {} → { table | null }
registerClusterIngress ingress → relay → Cloudly { bearer, version, nodeId, capabilities?, protocol } → { accepted: true, revision | null, protocol }
getClusterIngressRouteTable ingress → relay { knownRevision | null } → { table | null }
getClusterTrafficStatistics Cloudly → relay → ingress { fromDayUtc, toDayUtc, cursor?, limit } → { statistics, nextCursor? }
pushClusterIngressRouteTableChanged relay → ingress { revision } → {}
designateClusterIngress admin → Cloudly { identity, designation } → { designation, spec, replayed }
getClusterIngressDesignation admin → Cloudly { identity, clusterId } → { designation | null }

null for a table means Cloudly has computed none yet — a fact, not an empty table, and an ingress must not serve it as one. An ingress announces itself with a method, not a connection tag: the relay learns what connected from what it registered through, which a server owns, rather than from a label a client authored.

Who may be the ingress is Cloudly's decision, not the relay's. The ingress shares the relay's node listener, so anything that reaches it could ask to register, and a relay cannot tell an ingress from anything else. registerClusterIngress therefore carries the bearer Cloudly minted into the ingress workload's assignment environment and travels both hops with one body: the ingress sends it to its relay, the relay forwards it untouched, and Cloudly verifies the bearer and derives the cluster from the identity tag on the relay's own session. The bearer is sensitive ephemeral input — excluded from hooks, logs, diagnostics and journals on every hop, and never quoted in a validation reason.

Isolation on that shared listener is therefore by authorization, not by routing: only the connection whose registration Cloudly accepted is the ingress, getClusterIngressRouteTable and getClusterTrafficStatistics are answered for that connection alone, and any other connection — including a Pallet node's — is refused by name. validateClusterIngressRegistration(value, path?) checks shape only: whether the bearer is the one Cloudly minted is Cloudly's to verify. It reads the registration's protocol offer through the one grammar protocol owns, and both hops judge the body in the one server order protocol.readProtocolOffer states — the offer is negotiated before this shape check and before the bearer — so a relay refuses a peer it cannot serve without asking Cloudly, and an accepted registration carries the accepting side's own offer back.

A refused registration is named. A protocol mismatch stays the IProtocolRefusal every hop already carries; every other refusal travels as IClusterIngressRegistrationRefusal { code: 'cluster-ingress-registration-refused', reason, retryAfterMs } in TypedResponseError.errorData, forwarded by the relay untouched. clusterIngressRegistrationRefusalReasons lists the reasons in the order Cloudly judges them: not-ready (Cloudly cannot judge registrations yet), registration-invalid (the body fails validateClusterIngressRegistration), relay-not-connected (the forwarding relay is not the connection Cloudly elected for its cluster), ingress-not-designated, node-mismatch (the node is not one the designation lets the ingress run on), assignment-not-running (no ready assignment of the ingress service runs on that node under its current session, or it does not carry the designation's current credential), bearer-refused and conflict (the designation, assignment or session changed while the registration was judged). retryAfterMs is how long the ingress waits before it registers again, 1 s to 300 s (clusterIngressRegistrationContract). validateClusterIngressRegistrationRefusal names every violation and readClusterIngressRegistrationRefusal(errorData) answers the refusal or null, reading plain data only, as protocol.readProtocolRefusal does. No refusal quotes the bearer.

pushClusterIngressRouteTableChanged is the relay's hint that it holds a newer table. It carries the revision alone (IClusterIngressRouteTableChange, validateClusterIngressRouteTableChange) and nothing secret; the ingress answers it by pulling getClusterIngressRouteTable with the revision it already applied and ignores a hint that is not newer, so a lost or repeated hint costs nothing.

clusterIngressEnvironmentNames spells the ingress workload's environment once for both sides: SERVEZONE_CLUSTER_INGRESS_BEARER (the registration bearer, delivered as a secret), SERVEZONE_CLUSTER_RELAY_ORIGIN (the exact IClusterRelay.origin) and SERVEZONE_CLUSTER_NODE_ID (the node the assignment runs on). Cloudly writes them; the ingress reads them.

composeClusterIngressEnvironment(serviceEnvironment, { relayOrigin, nodeId }) is how Cloudly writes the two non-secret names for one replica of the designated ingress service: the service's own environment with SERVEZONE_CLUSTER_RELAY_ORIGIN and SERVEZONE_CLUSTER_NODE_ID added (IClusterIngressLaunchFacts; relayOrigin is the cluster's IClusterRelay.origin or null, nodeId the node the replica's assignment runs on). The bearer is not added: it reaches the workload through its secret manifest. The answer is a TClusterIngressEnvironmentComposition:

status reason When
composed — environment is the service's own with both names added.
refused cluster-relay-missing relayOrigin is null or validateClusterRelay refuses it. Answered first.
refused cluster-ingress-environment-conflict The service's own environment already sets one of the three names; the message names every such key and no value.

refused carries the contract's message, and both reasons are TClusterIngressEnvironmentRefusal, members of TServiceRuntimeRefusal. A nodeId that is not a node id is the caller's error and is thrown. Both values are part of what the replica runs, so a moved relay or another node is another environment. The inputs are not changed.

A cluster born on Pallet has no cutover scope, census or target set, so its ingress is designated by IClusterIngressDesignation { id, generation, previous, clusterId, runtime, ingressServiceId, preparedSpecRevision, credential, actorId, acceptedAt, digest } instead of the fleet-cutover designation. runtime is the cluster's pallet phase reference and credential (IClusterIngressCredentialReference) the public identity of the managed bearer, never its value. snapshotClusterIngressDesignation, computeClusterIngressDesignationDigest (domain serve.zone/cluster-ingress-designation) and validateClusterIngressDesignation judge it, and snapshotClusterIngressDesignationRequest the administrator's compare-and-swap intent (IClusterIngressDesignationRequest { mutationId, clusterId, ingressServiceId, currentRuntime, expectedSpecRevision, expectedDesignation }). A cluster holds exactly one kind of designation: Cloudly refuses to designate a cluster that holds a fleet-cutover ingress designation, and the reverse. Which nodes the ingress runs on is the ingress service's runtime spec placement.

IClusterTrafficStatistics is one UTC day of one hostname's traffic on one node — requests, bytesIn, bytesOut, connections and the backend failures the proxy can observe (connect, handshake, request). Day buckets rather than rates: a rate answers "how busy is it now", which the live proxy already answers, while a bucket survives a restart, a poll that straddles midnight and a Cloudly outage. HTTP status classes are deliberately absent — the proxy exposes no status counters today, and a number this contract cannot source would be a guess. Readers page it exactly like CoreMail's mail statistics — the same fromDayUtc, toDayUtc, cursor and limit, answered as { statistics, nextCursor? } — re-reading yesterday and today on every pass because a day only stops changing once it has closed. dayUtc must name a day that exists: the shared UTC calendar-day rule round-trips through Date.UTC, so 2026-13-45 is refused rather than stored.

const reasons = data.validateClusterRouteTable(table); // []
const pageReasons = data.validateClusterTrafficStatisticsPage({ statistics: [] }); // []
const ingressReasons = data.validateClusterIngressRegistration({
  bearer: ingressBearerFromAssignmentEnvironment,
  version: '3.0.0',
  nodeId: 'node-a',
  protocol: protocol.createProtocolOffer('32.0.0'),
}); // []

A route to a node's published host ports needs the address that node answers on, and only the node can prove it: it holds the uplink observation the address belongs to. It therefore states it on the network report it already sends — IReportRuntimeNetworkProtectionRequest gains the optional uplinkAddress, a unicast IPv4 literal that is neither loopback nor link-local (runtimeNodeUplinkAddress). Absence stays absent: a node that states no address is told apart from one that states a value, and nothing downstream defaults or infers one. Cloudly stores what the node stated and resolves it into the route table; the relay never substitutes an address of its own.

The same report carries the optional hostAddresses: every IPv4 address the node's host carries on interfaces the node does not manage itself, sorted as strings, unique, never empty when present, at most runtimeNetworkProtectionContract.maximumHostAddresses (32), each one admitted by runtimeNodeUplinkAddress. Workload bridges and the managed VPN interface are the node's own and are left out. Cloudly names a plan platform endpoint in that node's hostPlatformEndpointIds when the endpoint's address is one of them. A relay registration already tells Cloudly which host carries the relay listener and hub; hostAddresses is how a host without a relay, such as a node with a Corestore beside it on one machine, states the same. Absent states nothing and is never defaulted.

Shared service port helpers are exported from data so Cloudly, App Store resolution, Coreflow, Coretraffic, Onebox, and dcrouter agree on the same normalization rules:

const normalizedPorts = data.normalizeServicePortConfig(service.data);
const defaultTarget = data.resolveDefaultServiceTargetPort(normalizedPorts.targetPorts);
const webTarget = data.resolveServiceTargetPort(normalizedPorts.targetPorts, 'web');

if (webTarget && data.isHttpServiceTargetProtocol(webTarget.protocol)) {
  // safe to use as a domain route target
}

normalizeServicePortData() returns service data with canonical targetPorts, domain targetPort refs, and publicPortMappings while removing legacy domain port fields from normalized writes.

Hosted-App Authorization and Platform OIDC

An App Store version can declare that it supports platform-managed OpenID Connect. This is a capability declaration only: Onebox and Cloudly keep OIDC disabled until an administrator explicitly enables it for that exact app instance, and they can disable it again without changing the template.

import { appstore } from '@serve.zone/interfaces';

const config: appstore.IAppStoreVersionConfig = {
  image: 'registry.example.com/example/app:1.0.0',
  platformOidc: {
    redirectPath: '/auth/oidc/callback',
    roles: [
      { id: 'admin', label: 'Administrator' },
      { id: 'user', label: 'User' },
    ],
    environmentVariables: {
      issuerUrl: 'SERVEZONE_PLATFORM_OIDC_ISSUER',
      clientId: 'SERVEZONE_PLATFORM_OIDC_CLIENT_ID',
      clientSecret: 'SERVEZONE_PLATFORM_OIDC_CLIENT_SECRET',
      redirectUri: 'SERVEZONE_PLATFORM_OIDC_REDIRECT_URI',
      audience: 'SERVEZONE_PLATFORM_OIDC_AUDIENCE',
    },
    clientAuthenticationMethod: 'client_secret_basic',
  },
};

redirectPath is a canonical callback path on the app's HTTPS origin. Registration validation also requires that canonical app origin explicitly and rejects cross-origin, normalized, query-bearing, fragment-bearing, or duplicate callback URLs. The five environment values are environment-key names, not credentials embedded in the manifest. A host injects the generated client secret through launcher environment delivery and injects the other registration values only while OIDC is enabled.

data.IHostedAppRoleAssignment binds a stable user subject to an immutable appInstanceId. That same app instance is the OIDC client_id, the ID-token aud, and the servezone_app_instance_id claim. Tokens include only the assigned roles for that audience; preferred_username is display metadata and must not be treated as identity authority.

requests.hostedapp exports the shared authorization RPCs used by a host dashboard. getHostedAppAccessConfiguration returns human user summaries, role assignments, and hosted-app summaries. setHostedAppRoleAssignment returns the assignment or null when it is removed, while setHostedAppPlatformOidc returns the current registration state. getHostedAppOidcAuthorization returns the app and role summary for a pending request; completeHostedAppOidcAuthorization and cancelHostedAppOidcAuthorization return the redirect URL. All six requests require a full data.IIdentity.

Validate untrusted manifests, registrations, and claims with:

  • data.validateHostedAppRoleDefinitions
  • data.validateHostedAppPlatformOidcRegistration
  • data.validateHostedAppPlatformOidcClaims
  • appstore.validateAppStorePlatformOidcCapability
  • appstore.validateAppStoreVersionPlatformOidc

Hosted lifecycle RPCs accept only IIdentityCredential. After JWT verification, handlers validate IHostedAppMachineClaims and derive the exact app instance and service from servezone_app_instance_id and servezone_service_id; callers cannot submit those selectors. Bootstrap actions are either a canonical same-origin setupRoute path or a nonsecret message. Control tokens, usernames, passwords, and token URLs are not part of the lifecycle contract. Each server action carries a CAS revision; completion requires its exact ID, current revision, and ready status so a delayed request cannot complete a replacement action.

App Store Catalog Contract

Onebox and Cloudly load the App Store at runtime as one signed catalog. The contract is what a consumer has to understand; everything else is data that arrives with the next catalog revision and needs no consumer release: new apps and versions, image digests, changelogs, descriptions.

import { appstore } from '@serve.zone/interfaces';

appstore.appStoreCatalogContract.id; // 'serve.zone/appstore-catalog'
appstore.appStoreCatalogContract.major; // 1
appstore.appStoreCatalogContract.minor; // 0

An IAppStoreCatalog carries contract, a strictly increasing revision, publishedAt, sourceRevision and every app (info, latestVersion, optional upgradeStrategy, versions). Each IAppStoreCatalogVersion holds the full config and its configDigest, and once withdrawn its withdrawn: { at, reason } record (IAppStoreCatalogWithdrawal). latestVersion names a listed version that is not withdrawn (under semver the highest of them); only an app whose every version is withdrawn names a withdrawn one. A catalog version config is an IAppStoreVersionConfig without the live-source members (source, resolvedSource, resolvedImageDigest), the per-version upgradeStrategy and appStoreVersion; its image is digest-pinned. A publisher ships an IAppStoreCatalogEnvelope: the catalog, its catalogDigest and one or more Ed25519 signatures by key id.

Compatibility rules:

  • A consumer reads a catalog of its own contract id and major whose minor is not newer than its own. Anything else is refused as contractUnsupported (appstore.judgeAppStoreCatalogContract).
  • Every level has exact keys; an unknown member is schemaInvalid (appstore.validateAppStoreCatalog).
  • requiresFeatures gates one version: a consumer passes the feature ids it fulfils (supportedFeatureIds, e.g. appStoreStorageFeatureIds values). A version naming any other id stays listed, is refused for install and upgrade there, and its config is not interpreted, so a new config capability ships under a feature id without refusing the rest of the catalog (appstore.getAppStoreCatalogMissingFeatureIds).
  • A published app@version stays in every later catalog and never changes (appstore.judgeAppStoreCatalogSuccession): a later catalog that no longer lists it is publishedVersionRemoved, one that lists it with another configDigest is publishedVersionMutated, and a revision that does not increase is revisionRegressed.
  • A version names the oldest host releases that run it (minOneboxVersion, minCloudlyVersion). A host passes itself to appstore.judgeAppStoreCatalogInstallTarget as { kind: 'onebox' | 'cloudly', version } (IAppStoreCatalogConsumer), which refuses a version needing a newer release of it as consumerVersionTooOld; the version stays listed and installs once the host is upgraded.
  • A publisher retires a version by withdrawing it, never by removing it. A withdrawn version keeps its entry, config and digest, so services installed from it keep reading it, and is refused as an install or upgrade target, like a version gated on a feature the consumer lacks (appstore.judgeAppStoreCatalogInstallTarget). A withdrawal is final: a later catalog that revokes or rewrites it is publishedWithdrawalChanged; a version withdrawn by mistake is republished under a new version.

Digests are SHA-256 over strict canonical JSON (keys in UTF-16 code-unit order, safe integers only) of a domain-separated value: computeAppStoreCatalogConfigDigest, computeAppStoreCatalogDigest and computeAppStoreCatalogSigningKeyId. A signature covers createAppStoreCatalogSignaturePayload(catalogDigest). These helpers use Web Crypto and are safe in the browser entry.

Signing and verification are Node-only:

import {
  createAppStoreCatalogSigningKey,
  signAppStoreCatalog,
  verifyAppStoreCatalogEnvelope,
} from '@serve.zone/interfaces/runtime';

const verification = await verifyAppStoreCatalogEnvelope(fetchedEnvelope, {
  trustedKeys, // IAppStoreCatalogTrustedKey[] shipped with the consumer release
  supportedFeatureIds,
});
if (!verification.accepted) {
  // verification.refusal.reason: schemaInvalid | digestMismatch | signatureInvalid | contractUnsupported
}

verifyAppStoreCatalogEnvelope authenticates before it interprets: envelope shape, catalogDigest, one valid signature by a trusted key (a failing signature by a trusted key refuses the envelope), then contract, schema and every configDigest. The consumer then judges succession against its active catalog.

A publisher keeps its signing key as a file: importAppStoreCatalogSigningKeyPem(text) reads one unencrypted PKCS#8 PEM Ed25519 private key (openssl genpkey -algorithm ed25519) into an IAppStoreCatalogSigningKey with its public key and key id, and exportAppStoreCatalogSigningKeyPem(key) writes one. getAppStoreCatalogTrustedKey(key) is the IAppStoreCatalogTrustedKey a consumer ships for it. A trusted key also converts from and to an SPKI PEM public key (openssl pkey -pubout) with importAppStoreCatalogTrustedKeyPem and exportAppStoreCatalogTrustedKeyPem. Each refuses another PEM label or key type and a key id that does not match its public key, and no error quotes its input.

A service records the App Store version it runs as IService.data.appStoreTemplate (IAppStoreServiceTemplateSnapshot: appId, version, configDigest, config, catalogRevision), taken from a verified catalog by appstore.createAppStoreServiceTemplateSnapshot. It is server-managed (TServiceWritableData excludes it): the host writes it at install, upgrade and catalog adoption, and runtime paths read it, never the live catalog. While no catalog is active, an unavailable status states its cause (loading, failed, refused) for programs next to the operator's reason, and App Store install, upgrade and upgrade preview are refused with the retryable code APPSTORE_CATALOG_UNAVAILABLE (IAppStoreCatalogUnavailableErrorData). IReq_Any_ImportAppStoreCatalog imports an operator's envelope text (origin operatorImport) and answers the status after it, so an air-gapped host needs no registry.

A host loads the catalog from the npm registries appstore.resolveAppStoreCatalogRegistries(list) answers: the configured list (Cloudly's ICloudlySettings.appStoreCatalogRegistries, an Onebox setting) in canonical form, or appStoreCatalogDefaultRegistries (npmjs) when it is empty. appstore.validateAppStoreCatalogRegistries is the one rule for a configured list: https: base URLs without credentials, query or fragment, each registry once.

A host's App Store views carry the listing as IAppStoreApp and IAppStoreAppMeta: versions names the install targets, and the optional versionStates (IAppStoreAppVersionState) states every listed version, withdrawn and feature-gated ones included, with its judgeAppStoreCatalogInstallTarget answer as installRefusal.

TAppStoreCatalogStatus is the operator's view, discriminated by availability: ready (IAppStoreCatalogReadyStatus) always carries the active catalog's revision, digest, key and origin, including a restored last-known-good catalog; unavailable (IAppStoreCatalogUnavailableStatus) always carries a reason and no active catalog. Both state supportedContract, supportedFeatureIds, lastAttempt, nextAttemptAt and the last refused catalog. Loading a catalog never gates the boot, readiness or health of anything else; while none is active only App Store install, upgrade and upgrade-preview are refused. requests.appstore.IReq_Any_GetAppStoreCatalogStatus reads it, IReq_Any_RefreshAppStoreCatalog runs one attempt now and answers the state after it, and IReq_Any_GetAppStoreTemplates may answer it as catalogStatus.

The catalog validator judges every supported version config in full; no consumer needs a second validator. Besides the catalog's own rules (member set, image pinning, the feature gate, minimum versions, upgrade flags, healthCheck) it applies appstore.validateAppStoreVersionConfig, the member rules a host applies before it installs a version, in this order:

  • image, then the port model (appstore.validateAppStoreVersionPortModel): port or targetPorts is required, and targetPorts, domains and publicPortMappings are read as exactly their types and resolved through data.normalizeServicePortConfig.
  • containerArgs and the environment declarations.
  • Storage (appstore.validateAppStoreVersionStorage): the exact config member set (appstore.appStoreVersionConfigKeys), platformRequirements, storageClasses and storageRequests (declared together, every class used), the storage feature ids requiresFeatures must name, portable legacy volumes, and every environment, secret-file and mount destination used once.
  • Platform OIDC, then Docker publishedPorts (appstore.validateAppStorePublishedPorts).

Each storage request is judged on its own by the storage contract behind platform.normalizeStorageRequest: members, capacity quantities and limits, protection, delivery and mount path, with that contract's messages and bounds, and a legacy volume's mount path as the request it becomes. The version config adds what only holds between requests: each request's class is declared, of its kind and backs a declared limit (hardQuota), backup or snapshots; ids, environment keys, secret files and mounts are used once; the legacy names legacyFilesystem, legacyObjectStorage and legacy-s3 stay reserved. Two App Store narrowings apply on top: a request id is a stable lowercase identifier (it shares a namespace with the derived legacy-volume-* ids), and a file delivery targets one file directly below /run/secrets.

Each of these validators stops at the first violation and answers it as a one-element list starting Invalid <label>, so a config is refused for the same reason everywhere.

TypedRequest Contracts

Use requests when registering handlers with @api.global/typedrequest or when creating typed requests through a TypedSocket client.

import { requests } from '@serve.zone/interfaces';

type GetClustersRequest = requests.cluster.IReq_Any_Cloudly_GetClusters;

const methodName: GetClustersRequest['method'] = 'getClusters';

Each request interface follows the same pattern:

interface IExampleRequest {
  method: 'methodName';
  request: Record<string, unknown>;
  response: Record<string, unknown>;
}

requests.config.IRequest_Any_Cloudly_GetClusterConfig accepts data.IIdentityCredential, which contains only the JWT needed for server-side identity resolution. Username/password login and machine-token exchange responses continue to return the full data.IIdentity.

Immutable Deployment Contracts

Cloudly deployment authority is expressed as an exact data.IServiceDeploymentGrant. A service grant applies to one owned existing service. An organization-service-slot grant reserves authority for one exact future service ID in an organization; it is not an organization-wide wildcard. configureServiceDeploymentMachineUser accepts this discriminated grant object instead of separate service and capability fields.

A deployer authenticates as a deployment machine user: an API machine user with a bearer token and the grants attached to it. An administrator manages these users with two requests:

  • requests.admin.getDeploymentMachineUsers lists every deployment machine user as data.IDeploymentMachineUserMetadata: its id, username, token expiries and grants, never a token value.
  • requests.admin.mutateDeploymentMachineUser takes one data.TDeploymentMachineUserMutation. create makes a user with one token under a username no other user holds; data.isDeploymentMachineUsername is the rule (lowercase, at most 64 characters, no : or /), because the username is also the registry login name. rotate-token replaces every token of the user with a new one, and revoke-token removes them all while the grants stay. An expiry must be in the future and at most deploymentMachineUserContract.maximumTokenLifetimeMs (one year) away (data.isDeploymentMachineTokenExpiryAdmissible). The token is answered exactly once, on the first answer to create or rotate-token; a replay of the same mutationId answers the historical metadata with token: null, so a lost token is replaced by a rotation.

Both refuse by the names in data.deploymentMachineUserRefusals, carried as data.IDeploymentMachineUserErrorData. data.snapshotDeploymentMachineUserMutation and data.snapshotDeploymentMachineUserMetadata are the structural checks; authorizing the administrator and the target user is the backend's. A reservation is made only by an API machine user: any other actor, an administrator included, is refused DEPLOYMENT_MACHINE_IDENTITY_REQUIRED, which is also a preflight blocker code, because only that user's registry push is bound to the operation it reserved.

The immutable deployment workflow is:

  1. Reserve the exact service, namespace, registry repository/tag, and route intent with reserveServiceDeployment.
  2. Push the OCI index to the returned exact tag with the authenticated deployer identity.
  3. Promote the authenticated release evidence with promoteServiceImageRelease.
  4. Observe exact rollout and runtime-digest evidence with getServiceDeploymentStatus.
  5. For a greenfield service, expose and verify its public route with promoteServiceDeploymentRoute only after the immutable rollout succeeds.

A legacy deployOnPush service without deployable source moves to a greenfield successor by a legacy import, which replaces step 2:

  • While the legacy replicas still run — before its runtime is frozen — an administrator reads the legacy Swarm service and every container of its tasks on the hosts that run them, and submits that as hostObservation (data.ILegacyServiceHostObservation: Swarm service name, spec image and mode with its desired replicas; per replica the node hostname, Swarm task id, container id, the engine's image store, the container's image id, the image's RepoDigests, the container state and its health, none when no healthcheck applies) to captureLegacyServiceRuntimeAttestation (requests.service.IRequest_Any_Cloudly_CaptureLegacyServiceRuntimeAttestation). data.validateLegacyServiceHostObservation judges its shape, data.selectLegacyServiceHostObservationRootDigest names the one root every replica runs from the service's own repository, and data.judgeLegacyServiceHostObservation judges it against that root as Cloudly reads it from its own registry: exactly the desired replicas of a replicated service, its spec image naming the repository by the source tag or the root digest, every replica running and healthy or none, each container's image id the root digest on the containerd image store or one of the root's platform image configuration digests on the classic store. The source tag need not still name the root. The stored data.ILegacyServiceRuntimeAttestation records that root, its platform manifests, the replicas, the evidence source administrator-host-observation and the observation itself; data.validateLegacyServiceRuntimeAttestation states its rules. Incomplete evidence is refused, never stored, and there is no substitute once the runtime is stopped. getLegacyServiceRuntimeAttestations lists what Cloudly holds for a legacy service; both requests are for administrators only.
  • The greenfield reservation, and every preflight of it, names legacyImport (data.ILegacyImageImportReference: legacy service, attestation, asserted source root digest) instead of sourceRevision, sourceDirty and release evidence. Refusals are the blockers LEGACY_IMPORT_ATTESTATION_MISSING, LEGACY_IMPORT_SOURCE_MISMATCH and LEGACY_IMPORT_PLATFORM_MISMATCH.
  • Instead of the registry push, the deployer calls importServiceDeploymentLegacyImage: Cloudly imports the attested image from its own registry into the reserved tag — an index root as it is, a single platform manifest under a one-descriptor index of the manifest's own family — and records release evidence of source cloudly-legacy-import naming the legacy source. Steps 3 to 5 follow unchanged; selectTrustedImageReleaseEvidence accepts that evidence only with evidenceSource: 'cloudly-legacy-import'.
  • The reservation and the import need the capability deployment:import-legacy, which only an organization-service-slot grant can carry.

data.IDeploymentRouteRequest.proxied carries provider-specific DNS proxy intent as an optional boolean. New route declarations should set it explicitly; omission remains valid for persisted historical operations and legacy callers.

data.IServiceDeploymentOperation is the durable revisioned compare-and-set fence for this workflow. Image promotion requires Cloudly-created trusted evidence that binds the operation, actor, repository, exact tag, root digest, and OCI index media type. The service request group also exposes deployment preflight, exact-digest rollback, retry, and cleanup contracts.

Two of these carriers accept an optional protocol offer (protocol.IProtocolOffer, see Protocol handshake): getDeploymentPreflight, through data.IDeploymentPreflightRequestData.protocol, and getServiceById. Those are the read-only requests that open a deployment and an adoption, and they are the only deployment carriers that take one. The mutating carriers stay byte-exact deliberately: Cloudly derives reserveServiceDeployment's requestDigest from the whole request body, so an offer inside one of them would bind a durable operation's identity to whichever @serve.zone/interfaces release the deployer had installed, and an in-flight deployment would stop being resumable across a deployer upgrade. Stating the offer on the opening read names developer-machine skew before anything is reserved and leaves nothing durable behind.

The member is optional — a peer of this release states nothing by offering and an older deployer sends none, so absence is a valid request, not a refusal. A stated offer is exact-key like every other offer in this package: validateDeploymentPreflightRequest judges it with protocol.validateProtocolOfferOn, the naming twin of the protocol.readProtocolOffer a server negotiates from, and answers a malformed or accessor-backed offer with a single INVALID_REQUEST blocker naming the offer and its reasons. Whether two peers can serve one session is never a preflight blocker: that is protocol.negotiateProtocol's decision, answered as an IProtocolRefusal.

Gateway request contracts include getGatewayClientRoutes (requests.gateway.IReq_GetGatewayClientRoutes) for listing owned IGatewayClientRoute[] route views, and syncGatewayClientRoute for idempotently syncing or deleting hostname-owned, routeRef-owned, and combined hostname-plus-routeRef routes. A client can label canonical intent with managedRouteKind: 'letsencrypt-http01-forward' and set a higher priority for a path-specific HTTP-01 route while retaining a separate normal route for the same hostname. Mail request contracts include syncMailAddressBinding, deleteMailAddressBinding, rotateMailCredential, and getMailDeliveryStatus. IReq_GetMailDeliveryStatus looks up a delivery spool item by spoolItemId, returns data.IMailDeliveryStatus, and accepts IMailSubmissionRequestAuth so service-mail credentials can query their own accepted, queued, deferred, delivered, or failed status. Typed outbound messages may set replyTo to one bare ASCII mailbox address; arbitrary Reply-To values do not belong in the custom header bag. Invalid values and typed-field/custom-header conflicts return stable TMailSubmissionErrorCode values. TMailAddressBindingSync.outboundEnabled explicitly controls whether a gateway should maintain a managed outbound SMTP credential for an address binding. Binding credential metadata is public; rotateMailCredential returns the new secret only in its one-time IMailCredentialOneTimeSecret response.

Web Push Contracts

New Web Push integrations use requests.webpush. Control-plane methods and application delivery methods deliberately use different, non-overlapping authentication types:

  • listWebPushBindings, syncWebPushBinding, deleteWebPushBinding, rotateWebPushCredential, and rotateWebPushVapidKey use control-plane identity or gateway API-token authentication.
  • getWebPushServiceStatus, enqueueWebPush, cancelWebPush, and getWebPushDeliveryStatus require a Web Push application credential. The gateway derives the owner exclusively from that credential; application requests cannot submit owner identity.

syncWebPushBinding may return the initial application credential secret once, and rotateWebPushCredential may return its replacement once. Binding and status DTOs contain only public credential and VAPID metadata. enqueueWebPush requires a credential-scoped idempotency key, an opaque application subscription ID, the browser Push API subscription, the VAPID key ID used for that browser subscription, and a privacy-minimal notificationAvailable signal.

import { requests } from '@serve.zone/interfaces';

type EnqueueWebPush = requests.webpush.IReq_EnqueueWebPush;
type WebPushStatus = requests.webpush.IReq_GetWebPushDeliveryStatus;

A delivery state of pushServiceAccepted means only that the remote push service accepted the encrypted request. It does not prove browser receipt, notification display, or user interaction.

Gateway Client Lifecycle and DNS

syncGatewayClientRoute accepts an optional dnsMode. Omission means skip for older clients. observe reports DNS without changing it. reconcile makes the gateway authoritative for the exact route hostname: it claims or replaces manual A, AAAA, and CNAME records, including already-correct manual values. Its optional dns result contains a closed status, retryability, the desired A/AAAA target, overwritten-record evidence, checkedAt, and authoritativeVerifiedAt once the provider or authoritative server confirms the state. Consumers can carry that evidence while retrying public propagation instead of treating an immediate recursive lookup miss as permanent.

Use requests.gateway.IReq_ProvisionGatewayClientCredential to replace an admin/bootstrap token with a first-class gateway-client credential. The admin-authenticated request idempotently upserts a data.IGatewayClient, durably creates a new bound credential, returns its raw value once, and then revokes older credentials bound to that client. It never revokes the bootstrap/admin credential. A successful response is discriminated with success: true and always includes the action, durable client, one-time credential, and revocation count.

import { requests } from '@serve.zone/interfaces';

const provisioning: requests.gateway.IReq_ProvisionGatewayClientCredential['request'] = {
  apiToken: 'admin-bootstrap-token',
  provisioning: {
    id: 'cloudly-main',
    type: 'cloudly',
    name: 'Cloudly main',
    hostnamePatterns: ['*'],
    allowedRouteTargets: [
      {
        host: 'coretraffic.internal',
        ports: [],
        allowAnyPort: true,
      },
    ],
    capabilities: {
      readDomains: true,
      readDnsRecords: true,
      readRoutes: true,
      syncRoutes: true,
      syncDnsRecords: true,
      readMail: true,
      manageMail: true,
      readCertificates: true,
      requestCertificates: true,
    },
  },
};

getGatewayClientContext returns effective live policy. A gatewayClient role necessarily includes the credential ID, bound client ID/type, and policyGeneration; consumers should reject admin/operator or mismatched contexts rather than falling back to a caller-supplied owner ID. getGatewayClientMailOverview provides an owner-scoped domain and recent-message summary. getGatewayClientMailDomainCount derives ownership exclusively from the authenticating gateway credential and returns { count: number } for its distinct configured mail domains. Cloudly therefore holds only its gateway-client credential, as the system secret DCROUTER_GATEWAY_API_TOKEN (see Cloudly Gateway Credential Bootstrap below); neither it nor the former dcrouterOpsApiToken is part of data.ICloudlySettings.

Cloudly Gateway Credential Bootstrap

Cloudly enrolls itself at dcrouter as one least-privilege gateway client. The only way a gateway credential enters Cloudly is a dcrouter admin token an administrator hands in once, for the first enrollment and for every re-bootstrap a stale policy asks for — a changed gateway setting, or a runtime network address plan whose DNS record allowance moved.

  • requests.settings.IReq_Any_Cloudly_BootstrapExternalGatewayCredential (bootstrapExternalGatewayCredential): { identity, expectedGatewayUrl, expectedGatewayClientId, bootstrapToken: { mode: 'sealed', envelope } } → { state }. The token is sealed to Cloudly's active secret ingress recipient (requests.secret.IReq_GetSecretIngressRecipient) under data.createGatewayBootstrapEnvelopeContext({ gatewayUrl, gatewayClientId }), so it opens only for a request naming the same dcrouter URL (HTTPS only, since an admin token is on the wire) and client id, and Cloudly compares both with its own settings before the token leaves it. Cloudly asks dcrouter whether the token is an admin token, provisions, finalizes and stores its own gateway-client credential with it, and answers the enrollment that results. The admin token is used for that one call and never stored; revoke it at dcrouter afterwards. A refusal leaves the held credential unchanged and is one of data.externalGatewayBootstrapRefusals: gateway-unconfigured, gateway-target-mismatch, bootstrap-token-invalid, bootstrap-token-rejected, bootstrap-token-not-admin, bootstrap-candidate-pending, gateway-unreachable and bootstrap-provisioning-refused.
  • requests.settings.IReq_Any_Cloudly_GetExternalGatewayEnrollmentState (getExternalGatewayEnrollmentState): { identity } → { state }.

data.IExternalGatewayEnrollmentState { status, reason?, lastCheckedAt?, nextRetryAt?, cachedValidationProof, finalizationPending } never carries a credential. status is one of data.externalGatewayEnrollmentStatuses: unconfigured, checking, ready, degraded (dcrouter unreachable; Cloudly retries) and admin-bootstrap-required. Both sides judge values with requests.settings.validateBootstrapExternalGatewayCredentialRequest(request, recipient) (schema, recipient, context digest and the ciphertext bound, without opening anything), requests.settings.validateGetExternalGatewayEnrollmentStateRequest, data.validateExternalGatewayEnrollmentState and data.isExternalGatewayBootstrapToken(bytes), which accepts 1 to 1024 bytes of visible ASCII (data.externalGatewayBootstrapContract).

import * as smartcrypto from '@push.rocks/smartcrypto';
import { data, requests } from '@serve.zone/interfaces';

const target = { gatewayUrl: 'https://gateway.example.com', gatewayClientId: 'cloudly.example.com' };
const request: requests.settings.IReq_Any_Cloudly_BootstrapExternalGatewayCredential['request'] = {
  identity: { jwt: adminJwt },
  expectedGatewayUrl: target.gatewayUrl,
  expectedGatewayClientId: target.gatewayClientId,
  bootstrapToken: {
    mode: 'sealed',
    envelope: await smartcrypto.sealX25519Envelope({
      plaintext: adminTokenBytes,
      recipientPublicKey: ingressPublicKey,
      recipientKeyId: ingressRecipient.recipientKeyId,
      context: data.createGatewayBootstrapEnvelopeContext(target),
    }),
  },
};

Gateway Client Hostname Records and Certificates

A gateway client can hold an exact hostname without a gateway route: an address record at a value it chooses, and an exact certificate for that one name. Cloudly uses this for its cluster relay names, whose A record must carry the relay node's own address and whose certificate the relay serves itself. All three requests accept only a gateway-client credential (identity is never). The gateway resolves the client from that credential, and the hostname must match the client's hostnamePatterns. data.IGatewayClientHostnameOwnership names only { appId, hostname }. One ownership holds at most one address record and one certificate. A hostname held by another ownership, by a gateway route claim, or by a record the gateway does not manage is a conflict. The gateway never overwrites or adopts it.

  • syncGatewayClientDnsRecord (requests.gateway.IReq_SyncGatewayClientDnsRecord) publishes, replaces or withdraws the ownership's A or AAAA record. It needs the token capability syncDnsRecords, and the gateway advertises it as capabilities.dns.clientRecords. The value must be a canonical unicast literal inside the client's policy field allowedDnsRecordAddresses, a list of canonical CIDR prefixes. An absent or empty list allows no record. The result reuses the route DNS vocabulary (data.TGatewayClientDnsRecordStatus): target-invalid for an address outside the allowance, record-conflict with conflictingRecords as evidence, and zone-unavailable or mutation-failed when retryable. Records are never proxied.
  • getGatewayClientCertificate (requests.gateway.IReq_GetGatewayClientCertificate) answers ready with data.IGatewayClientCertificate (PEM chain, PEM private key, validFrom, validUntil, renewAfter), pending with retryAfterMs while an issuance runs, or refused with a data.TGatewayClientCertificateRefusal. Asking again after renewAfter yields a renewed certificate. It needs requestCertificates and readCertificates, and the gateway advertises it as capabilities.certificates.exactIssuance. The ready answer carries a private key; keep it sealed and never log it.
  • releaseGatewayClientCertificate (requests.gateway.IReq_ReleaseGatewayClientCertificate) removes the ownership's certificate material and any issuance in flight. It is idempotent and answers released or unchanged.

Both sides judge values with the same helpers:

  • data.validateGatewayClientHostnameOwnership
  • data.validateGatewayClientDnsRecord
  • data.validateGatewayAllowedDnsRecordAddresses
  • data.isGatewayClientDnsRecordAllowed
  • data.isGatewayAddressPrefix
  • data.isGatewayClientHostname

Only one spelling of each address is accepted: the dotted quad, or RFC 5952 text for IPv6. data.gatewayClientHostnameContract states the bounds: TTL 60 to 86400 seconds with a default of 300, at most 64 allowance prefixes, an appId of at most 128 characters, and a hostname of at most 253 characters.

import { data, requests } from '@serve.zone/interfaces';

const record: data.IGatewayClientDnsRecord = { type: 'A', value: '198.51.100.85', ttl: 60 };
data.isGatewayClientDnsRecordAllowed(record, ['198.51.100.0/24']); // true

const request: requests.gateway.IReq_SyncGatewayClientDnsRecord['request'] = {
  apiToken: gatewayClientToken,
  ownership: { appId: 'cluster-relay:clusterexample', hostname: 'relay.clusterexample.example.com' },
  record,
};

CoreMail contracts

requests.coremail is the shared contract boundary for authenticated workload sessions, Coreflow reconciliation, and the CoreMail-to-dcrouter gateway session. Only the three authentication handshakes carry reusable peer credentials. Subsequent transfer operations may carry a scoped, short-lived, one-time bearer capability, while every subsequent workload request derives tenant, service, binding, capabilities, allowed senders, and the composite credential ID/version identity from the server-owned TypedSocket peer. Successful workload authentication also returns the effective binding state and the exact allowedOperations derived from data.coreMailWorkloadOperationPolicy. Disabled bindings never authenticate; draining bindings permit outbound status plus inbound list/fetch/ack only.

Gateway recipient resolution uses four strict outcomes. accept, defer, and reject are authoritative only for recipients owned by an active binding. unhandled means the CoreMail peer does not own that recipient, so the gateway may continue its next configured resolver. Consumers must apply data.normalizeCoreMailRecipientResolutions() against the exact requested recipient set before acting on a peer response.

CoreMail binds every issued routing handle to a digest of the exact normalized envelope and source sent at resolve time. A following coreMailGatewayPrepareInboundHandoff must resend byte-identical values, including the complete rcptTo set in the same order. Both peers use data.normalizeCoreMailGatewayMessage and data.normalizeCoreMailConnectionInfo so the digested bytes come from one shared implementation. Outbound status is monotonic per transportMessageId, and a permanently unknown transportMessageId is answered with a terminal failed status rather than an error, because CoreMail retries errors indefinitely.

Large content never travels inside TypedRequest JSON. Outbound body parts and attachments use prepare/upload/complete operations with short-lived one-time HTTP transfer grants. Inbound delivery uses bounded delivery listing followed by prepare/fetch/complete and an explicit acknowledgement after the workload has processed the exact byte count and SHA-256 digest. data.coreMailLimits defines the 64 KiB control-frame boundary, bounded structured content, the 30 MiB serialized MIME ceiling, a 56 KiB inbound page budget, bounded opaque cursors, transfer deadlines, and five-minute grant lifetime.

Coreflow applies data.ICoreMailDesiredState with a config-epoch compare-and-set fence. It stages the digest-fenced JSON snapshot through a bounded one-time HTTP upload, then applies it by reconciliation ID, so a large binding set never bypasses the 64 KiB control-frame boundary. Desired bindings contain password verifiers and secret references only, never plaintext workload or dcrouter credentials. Use data.normalizeCoreMailDesiredState, data.canonicalizeCoreMailDesiredState, and data.createCoreMailDesiredStateDigest at every producer and consumer boundary. Use data.verifyCoreMailDesiredStateDigest before applying a staged snapshot, data.normalizeCoreMailCredentialVerifier before accepting verifier metadata, data.normalizeCoreMailControlBootstrap for startup authority, and data.normalizeCoreMailGatewayPeerDesiredState for dcrouter peer state. The normalizers reject unknown fields, noncanonical mailboxes, ambiguous active recipient ownership, malformed SHA-256 values, and credential-lifecycle inconsistencies. CoreMail credential verifiers use the argon2id format and the exact policy exported as data.coreMailCredentialVerifierPolicy. The PHC parameter block is read by name, so any encoder's parameter order is accepted as long as m, t and p each appear once at the policy cost; the normalized verificationHash is always re-emitted as $argon2id$v=19$m=…,t=…,p=…$<salt>$<hash>, the canonical order the desired-state digest is computed over.

data.ICoreMailDesiredState.smtp is the optional CoreMail SMTP submission listener (MSA). It carries the listener port, the advertised EHLO hostname, and the runtime secret keys holding certificate and private-key PEM text; plaintext PEM material never appears in desired state. SMTP AUTH uses the binding identity directly: the username is the bindingId and the password is any current or retiring credential secret of that binding. The listener stays unavailable until the referenced PEM material resolves. Omitting smtp keeps the exact canonical bytes and digest of an API-only CoreMail. Accepted submissions report their entry point through the optional ICoreMailSubmission.source, either api or smtp.

data.ICoreMailServiceMailStatistics reports per-service outbound and inbound counters for one UTC calendar day. Control sessions read them through coreMailGetServiceMailStatistics using an inclusive day range, an optional service filter, and an opaque cursor. Apply data.normalizeCoreMailServiceMailStatistics before using reported counters.

data.ICoreMailControlBootstrap is installed before ordinary reconciliation. It contains the CoreMail service identity and verifier metadata only. The matching plaintext control credential is delivered exclusively to Coreflow through resolved runtime secrets. data.coreMailRuntimeKeys publishes the canonical environment keys for the verifier-only bootstrap payload and the separate control and gateway secret values; no consumer may derive or embed plaintext material in desired state.

The CoreMail contracts keep durable authority at the stable tenant, service, and binding identity while revisions, config epochs, and composite credentialId/version values fence sessions and new actions. Credential versions are authority-wide monotonic and unique even when rotation changes the credential ID. active bindings accept new mail, draining bindings permit existing status and inbound fetch/ack work without accepting new mail, and disabled bindings reject authentication. Cloudly retains a draining binding until pending inbound delivery reaches zero as reported by ICoreMailBindingReconciliationStatus.pendingInboundCount.

data.ICoreMailGatewayPeerDesiredState gives dcrouter the authoritative HTTPS CoreMail transfer origin associated with an authenticated CoreMail service. The same origin is carried in CoreMail desired state and returned by workload and gateway authentication. It must never be inferred from a socket, Host header, TypedSocket tag, or unrestricted peer input. CoreMail compares both authoritative views before handing over a path-only transfer grant. Cloudly and Onebox provision that peer state through listCoreMailGatewayPeers, syncCoreMailGatewayPeer, and deleteCoreMailGatewayPeer on the gateway-client mail surface, and the producer is the single owner of peer.transferOrigin. A mail inbound target may point at a CoreMail service using type: 'coreMail' with coreMail.coreMailServiceId, mirrored by the coreMail member of IServiceMailInboundConfig.targetType.

Inbound delivery pagination uses an opaque CoreMail-owned cursor, is bounded by data.coreMailLimits.inboundPageSize, and returns nextCursor only when another page may exist. Consumers must not construct or parse cursor contents. Cursor signing keys are value-free runtime references with one current and bounded retiring versions; plaintext remains in resolved runtime secrets.

Quota windows are fixed UTC minute/day buckets. A new outbound quota unit is consumed only by the first durable insertion of an idempotency identity, while replays consume none. Pending inbound includes every state except acknowledged. These semantics are exported as data.coreMailQuotaPolicy. Every binding carries finite messagesPerMinute, messagesPerDay, and maxPendingInbound values; omitted or unlimited quotas are not valid desired state. data.coreMailRetentionPolicy retains terminal outbound, acknowledged inbound, and idempotency receipts for 30 days and expired capabilities for 24 hours. Pending inbound is never age-purged.

HTTP transfers use canonical /transfers/<uuid> paths and one-time Bearer tokens. data.coreMailTransferTokenPolicy requires canonical 256-bit base64url token material, while issuedAt and expiresAt prove the exact five-minute lifetime. PUT succeeds with 204 and GET with 200. Content length and type must match the grant; digest integrity is bound by grant metadata and repeated in the completion RPC rather than an optional HTTP digest header. Apply the exported strict normalizers for outbound message descriptors, envelopes, method-specific upload/download grants, submissions, gateway outbound statuses, inbound deliveries, desired state, bootstrap state, and gateway-peer state at their corresponding untrusted request and response boundaries. Normalizer failures throw data.CoreMailContractError; its readonly code defaults to INVALID_REQUEST, while an otherwise valid outbound part that exceeds its kind-specific byte budget reports PAYLOAD_LIMIT_EXCEEDED.

Control and gateway credential rotation is ordered: provision the candidate plaintext through resolved runtime secrets, publish and activate the matching argon2id verifier, roll or reconnect every affected replica, observe per-task authentication and readiness, then mark the previous verifier retiring with acceptUntil. Remove the previous verifier and secret only after its acceptance window has elapsed and no session uses that composite credential identity. Reconciliation status carries the exact CoreMail task, rollout generation, and image digest so Coreflow can correlate every response with its authoritative current task roster. Apply data.normalizeCoreMailReconciliationStatus before using pending-inbound or composite active-session counts for drain and rotation decisions.

CoreMail control-plane settings and operator RPCs

data.ICloudlySettings carries the CoreMail control plane: coreMailEnabled, coreMailServiceId, the canonical coreMailControlEndpointUrl (https://host/socket), the path-free coreMailTransferOrigin, dcrouter's coreMailGatewayEndpointUrl (wss://…), the composite coreMailControlCredentialId/coreMailControlCredentialVersion, the submission listener's coreMailSmtpHost, coreMailSmtpPort and coreMailSmtpTlsMode, and the coreMailDefault* binding quota defaults. No credential or PEM value is part of settings: data.cloudlySystemSecretKeys publishes coreMailControlCredentialSecret, coreMailGatewayCredentialSecret, coreMailSmtpTlsCertificatePem, and coreMailSmtpTlsPrivateKeyPem as the stable keys of the removed values. The submission listener's PEM material is operator-supplied in v1; automatic issuance is a later workstream.

Secrets that Cloudly manages for CoreMail use the coremail management source and a coremail:<identifier> management scope, validated by the existing data.validateSecretManagementScope and secret-metadata validators.

IServiceMailConfig.coreMail publishes the service's CoreMail binding as public metadata only: binding and composite credential identity, the argon2id verification hash, the binding state, and any retiring predecessor with its acceptUntil. Plaintext never appears there, and the field is excluded from data.TServiceWritableData, so only the control plane writes it while the rest of mail stays caller-writable. IServiceMailConfig.allowedSenders makes the outbound sender allowlist explicit instead of inferring it from per-address outbound.enabled.

requests.coremailAdmin carries the identity-authenticated operator surface: getCoreMailServiceMailStatistics for per-service daily counters, getCoreMailControlStatus for the applied config epoch, desired-state digest and last reconciliation status, and listServiceMailTargets for the mail target inventory (data.IServiceMailTargetSummary with data.IServiceMailTargetAddressSummary rows).

Request groups are exported by product area:

  • requests.admin
  • requests.appstore
  • requests.baremetal
  • requests.baseos
  • requests.backup
  • requests.certificate
  • requests.cluster
  • requests.config
  • requests.coremail
  • requests.coremailAdmin
  • requests.corestore
  • requests.corestorecontrol
  • requests.deployment
  • requests.dns
  • requests.domain
  • requests.externalRegistry
  • requests.gateway
  • requests.hostedapp
  • requests.identity
  • requests.image
  • requests.inform
  • requests.log
  • requests.mail
  • requests.migration
  • requests.network
  • requests.node
  • requests.platform
  • requests.routing
  • requests.secret
  • requests.service
  • requests.settings
  • requests.status
  • requests.task
  • requests.version
  • requests.webpush

Secret material response contracts are not exported from the universal browser-facing entrypoint.

The root entrypoint does use @push.rocks/smartcrypto to parse and validate strict ingress envelopes. Browser consumers that import the root contract can therefore include SmartCrypto and its browser-compatible crypto dependencies in their bundle. This package never opens private keys; sealed runtime material contracts remain isolated to the Node-only /runtime subpath.

Node runtimes import the isolated subpath:

import {
  verifySealedResolvedSecretMaterial,
} from '@serve.zone/interfaces/runtime';
import type {
  IReq_GetResolvedSecretMaterial,
  ISecretMaterialExpectation,
  ISealedResolvedSecretMaterial,
} from '@serve.zone/interfaces/runtime';

async function acceptMaterial(
  material: ISealedResolvedSecretMaterial,
  expectation: ISecretMaterialExpectation,
) {
  if (!await verifySealedResolvedSecretMaterial(material, expectation)) {
    throw new Error('secret material does not match its immutable manifest');
  }
}

Runtime material is shaped as { manifest, entries }, where each entry contains only secretVersionId and a strict X25519 envelope. The helper verifies the canonical manifest digest, envelope context digests, active recipient key, sorted exact one-to-one version coverage, every request fence, and the trusted local organization/cluster expectation. It never decrypts. Organization and cluster are derived from the verified cluster JWT and cannot be selected by request fields.

Corestore Runtime Credentials

The Node-only /runtime export defines getCorestoreControlCredentialMaterial for recipient-bound sealed retrieval and publishCorestoreCredentialMaterial for ingress-sealed database or object storage credential publication. isCorestoreControlToken() validates the canonical token boundary, while createCorestoreControlCredentialPlaintextBytes() emits its exact JSON plaintext directly from validated UTF-8 bytes without first constructing a JavaScript token string. Publication grants carry the service, binding, reconciliation generation, binding-request digest, target Secrets revision, and bounded validity window. TCorestoreCredentialBindingRequest with validateCorestoreCredentialBindingRequest(), createCorestoreCredentialBindingRequestDigestInput(), and computeCorestoreCredentialBindingRequestSha256() produces that digest from the exact value-free provider, scope, environment, bucket, and optional retention request. Every object-storage request binds the exact durable bucketName; a database request may state the endpoint its connection values are addressed to, which then is part of the digest (see Database Binding Endpoint). Object-storage plaintext keeps the logical accessKeyId and secretAccessKey fields while createCorestoreCredentialSecretValues() expands them to the six canonical Corestore S3_* and AWS_* Secret aliases. Receipts bind that exact key coverage, created SecretVersion references, the admitted envelope/context digest, and the trusted organization and cluster scope. Cluster configuration DTOs add optional sorted corestoreCredentialPublicationGrants; existing consumers may omit them. When validateCorestoreObjectStorageCredentialMaterial() receives a trusted retention expectation, matching retention evidence is mandatory; missing or different evidence fails validation.

The runtime export also defines exact Corestore database backup receipts, allocation references, restore requests, and restore responses. Use normalizeCorestoreDatabaseAllocationReference(), encodeCorestoreDatabaseAllocationReference(), and computeCorestoreDatabaseAllocationReferenceSha256() for the allocation boundary; the receipt, restore-request, and restore-response APIs follow the same exact normalize, encode, and SHA-256 naming. They enforce bounded canonical JSON without performing a backup or restore. A restore request may carry expectedDatabaseAllocation so the restore implementation can fence the exact scratch allocation, and may carry the caller-owned restoreAttemptId to select a new fenced idempotency attempt. The attempt ID uses the database-backup identifier grammar and 256-character maximum. Omitting the snapshot databaseAllocation field together with request-level expectedDatabaseAllocation and restoreAttemptId preserves legacy request bytes and digests exactly. corestoreDatabaseBackupRuntimeLimits publishes the distinct 96 MiB snapshot payload, one-byte-larger framed snapshot-original, 128 MiB verified closure plaintext, and 256 MiB closure boundaries used by Corestore.

Corestore Control Token Administration

The elected relay of every cluster calls its nodes' Corestore control API with one control token (CORESTORE_API_TOKEN), which Cloudly stores for both Corestore provider configs and hands to the relay sealed (see Corestore Runtime Credentials). requests.corestorecontrol lets an administrator store that token without an exec into Cloudly's process.

setCorestoreControlToken { identity, token, expected } stores token under both provider configs in one transaction, so either both hold it or neither does. token must pass isCorestoreControlToken (32 to 4096 UTF-8 bytes, no whitespace) and is the only member of this contract that carries the token: it is never in an answer, a receipt, a log or a hook, and a sender never persists the request. expected is the data.TCorestoreControlTokenState the administrator last read: { state: 'absent' } or { state: 'present', revision }, never a minimum, so a set that committed in between is refused rather than overwritten. revision is 1 for the first token ever stored and advances by exactly one with every accepted set; no request removes a token. It requires a freshly authenticated administrator, exactly as a relay credential mutation does. The answer, data.ICorestoreControlTokenSetResult, is the state the set produced and its receipt. Refusals are data.corestoreControlTokenSetRefusals (data.TCorestoreControlTokenSetRefusal): token-invalid, token-present, token-absent, token-revision-mismatch and provider-unavailable.

Every accepted set is audited: Cloudly writes data.ICorestoreControlTokenSetReceipt (id, revision, via, actorId, setAt) in the transaction that stores the token. via is admin-request with the administrator's actorId, or container-cli with data.corestoreControlTokenContainerCliActorId (operator:secrets-cli) when Cloudly's in-container secrets set-corestore-control-token stored it; that command advances the same revision. A receipt is value-free: no token, no hash. getCorestoreControlTokenState { identity } answers data.ICorestoreControlTokenStateRead: the state and its receipts, oldest first, contiguous and ending at the state's revision; no token stored means no receipt. Only a token stored before receipts existed, which Cloudly's migration records as revision 1, has no receipt, so the receipts start at revision 1 or 2. A read answers at most data.corestoreControlTokenMaximumReceipts (4096) receipts: past that many sets it answers the latest 4096, whose oldest may name any revision.

A reader checks a set answer with bindCorestoreControlTokenSet(answer, request): the state is present at revision 1 after an absent expectation and at the expected revision plus one otherwise, and the receipt is an admin-request receipt for exactly that revision. snapshotCorestoreControlTokenSetRequest, snapshotCorestoreControlTokenSetResult, snapshotCorestoreControlTokenSetReceipt, snapshotCorestoreControlTokenState and snapshotCorestoreControlTokenStateRead are the exact-keyed readers. Every refusal is the contract's one value-free message and never quotes the token.

Corestore Through The Relay

A serve.zone cluster is outbound-only and Corestore is node-local, so Cloudly never reaches a Corestore itself: every Cloudly-owned operation against one travels through the cluster relay, and the relay is told which node it means. These contracts carry no version suffix and no schema field: the protocol version is the installed @serve.zone/interfaces version.

The node is a field, never a connection. One relay serves every node of its cluster, so the connection that answers proves the cluster and nothing more. backupClusterService, restoreClusterService, pruneClusterNodeArchive and getClusterCorestoreInventory each name nodeId, and they supersede executeServiceBackup, executeServiceRestore, coreflowPruneNodeArchive and coreflowGetCorestoreInventory, which were dispatched at whichever coreflow had tagged itself with a node hostname. The superseded four stay declared until the coordinated major so a mixed fleet keeps compiling; nothing new should use them.

IClusterCorestoreInventory { nodeId, checkedAt, reachable, errorCode?, services } is what one node's Corestore holds, and validateClusterCorestoreInventory(value, path?) returns one named reason per violation. A reachable inventory with an empty services list is the proof that a namespace is blank — the answer a greenfield cluster gives. An unreachable node proves nothing, so reachable: false is refused without an errorCode and refused with any service, and a reachable answer is refused with one: the two states can never be read as each other.

Method Direction Body
backupClusterService Cloudly → relay { nodeId, backupId, service, tags?, replication? } → { snapshots, replication? }
restoreClusterService Cloudly → relay { nodeId, backupId, service, snapshots, clear?, resourceTypes?, replication? } → { restored }
pruneClusterNodeArchive Cloudly → relay { nodeId, retention, dryRun? } → { found, result? }
getClusterCorestoreInventory Cloudly → relay { nodeId } → { inventory }
getClusterPlatformDesiredState relay → Cloudly {} → { capabilities, providerConfigs, bindings }
backupServiceDatabaseClosure relay → Cloudly { backupId, serviceId, nodeId, descriptor, closure } → { accepted }
restoreServiceDatabaseClosure relay → Cloudly { backupId, serviceId, nodeId } → { descriptor, closure }
prepareIsolatedRestoreOnNode Cloudly → relay { nodeId, control } → { progress }
stageIsolatedRestoreArchive Cloudly → relay { nodeId, control } → { progress }
getIsolatedRestoreProgressOnNode Cloudly → relay { nodeId, control } → { progress }
executeIsolatedRestoreOnNode Cloudly → relay { nodeId, control } → { progress }
cleanupIsolatedRestoreOnNode Cloudly → relay { nodeId, control } → { progress }

Placement is a field too. getClusterPlatformDesiredState answers with IClusterPlatformBindingPlacement { nodeId, binding } rather than a bare binding list: a database binding without a node is a database nobody can place once one relay serves the whole cluster. Cloudly derives the placement from where the service is assigned; the relay never guesses it.

Closures stream; they are never base64. A database backup larger than the portable handoff envelope moves as a Corestore closure, and the bytes travel as a directional VirtualStream on the relay's own session, exactly like pushImageVersion and pullImageVersion. Neither method carries an identity: the cluster is the one the verified machine identity on that session names, and a request-level identity would be a second, weaker authority for the same fact. The relay is the requesting peer in both directions, because only outbound requests exist: backupServiceDatabaseClosure sends (TVirtualStream<'send'> in its request) and restoreServiceDatabaseClosure receives (TVirtualStream<'receive'> in its response). accepted permits sending and never confirms storage; the stream's acceptance receipt does. The base64 uploadBackupArchiveObject and downloadBackupArchiveObject contracts are retiring with this path — Corestore already documents its repository-wide archive object routes as legacy — and are removed in the coordinated major.

ICorestoreDatabaseClosureDescriptor { size, sha256, receipt } from the Node-only /runtime export travels in the parent request, and validateCorestoreDatabaseClosureDescriptor(value, path?) refuses it by name. size is mandatory, and it is why the descriptor exists: Corestore's database backup restore requires one exact Content-Length, identity transfer encoding and no content encoding, and refuses a body whose byte length differs from its receipt, so the receiver must open its Corestore request before the first byte arrives. A stream that only learned its length at EOF could never be restored. size and sha256 are cross-checked against receipt.closure.archiveBytes and receipt.closure.sha256, so a descriptor can never describe a closure its own receipt contradicts.

Isolated restore carries its grant. Corestore accepts exactly one authority on /isolated-restore/*: a compact RS256 grant that names one operation, expires within minutes and binds the cluster, the node name, the canonical resource mappings and the archive manifest. The five methods above are therefore one shape — { nodeId, control } → { progress } — where control is the body Corestore itself reads (data.IIsolatedRestoreControlPrepareRequest, …WriteRequest, …StatusRequest, …ExecuteRequest, …CleanupRequest) with the grant inside it. The relay forwards the control untouched and can neither widen it nor mint one; it adds no identity and no protocol offer either, because the cluster is the one the verified machine identity on the session registerCloudlyClientSession opened names, and the protocol offer was exchanged by that handshake and again by registerClusterRelay's own protocol member. Every one of them is sensitive ephemeral input: the grant is a bearer credential, declared once on data.IIsolatedRestoreControlAuthority, so exclude the complete request from hooks, logs, diagnostics and durable journals on every hop.

The order an operator's restore runs in. Cloudly calls createIsolatedRestore with the node it dispatches to, then, on that node: prepareIsolatedRestoreOnNode with expectedRevision: 0, stageIsolatedRestoreArchive once per bounded chunk of every manifest object, executeIsolatedRestoreOnNode, and cleanupIsolatedRestoreOnNode when the rehearsal is over; getIsolatedRestoreProgressOnNode reads the same state without changing it. Staging is bounded by data.isolatedRestoreContractLimits: at most 512 KiB decoded per chunk, 4,096 chunks per object and 64 MiB per object, each chunk authenticated by its own chunkSha256 against the immutable object descriptor. Every mutation carries the revision the node last stated as its expectedRevision, so a replay settles instead of racing.

The node id routes, the node name is authorized. createIsolatedRestore takes targetNodeId — the cluster document's node id, which is what a relay routes on — while the signed grant keeps targetNodeName, which is what the operator authorized and what Corestore verifies. Cloudly is the only party holding both, so IIsolatedRestoreRecord states both and every control call is dispatched by the id. data.bindIsolatedRestoreProgressToNode(progress, { nodeId, restoreId }) is where an answer is tied back: the control body names neither the node nor the restore, so without it a relay could answer with another node's — or another restore's — perfectly canonical progress.

What a node can state, in Corestore's own words. data.IIsolatedRestoreProgress is Corestore's staging summary plus the node the control was carried to: { nodeId, restoreId, stagingArchiveId, status, revision, lastFence, expectedObjects, receivedObjects, completedMappings, totalMappings, preparedAt, updatedAt, executedAt?, stagedObject? }, with status one of data.isolatedRestoreProgressStatuses — prepared, executing, failed, executed, cleaned. stagedObject is present only on a staged chunk and says where the next chunk continues; it is named apart from the object receipt Corestore's own write answer carries, which is a different shape and is not projected. data.normalizeIsolatedRestoreProgress(value) reads it exactly, refusing an unknown status, a count larger than the plan it belongs to, a revision or lastFence below 1 — a committed state carries both — an executing or executed answer that does not hold the whole archive, and an executed answer that left a mapping behind or does not say when it finished. A node states nothing of the control plane's record — no source backup, no requester, no verification — so Cloudly composes its record from this and its own knowledge, exactly as it does with the Corestore inventory. A relay that cannot reach the node answers one of data.isolatedRestoreNodeControlRefusals instead: restore-node-not-carried, restore-node-unreachable, or restore-node-refused followed by Corestore's own reason.

IIsolatedRestoreDatabaseMapping.target.databaseAllocation carries Corestore's safe allocation reference, present exactly when the target database is allocation-managed. It lives inside the mapping rather than beside it because the canonical mapping digest is what a grant binds: an allocation reference outside it would be authority nobody signed. This is what an allocated isolated restore was missing — Corestore refuses one whose exact reference its authority never covered. The reference itself is declared once, as data.ICorestoreDatabaseAllocationReference, and re-exported from /runtime so both trees name one structure.

The endpoint comes from the cluster document. IClusterNode.data.corestore { endpoint } is where one node's Corestore control API answers, owned by Cloudly and read by the relay, stored in Cloudly's database like every other field of the cluster document — never an alias, a file or a relay environment variable. validateClusterNodeCorestore(value, path?) admits a canonical http: or https: origin with no userinfo, path, query, fragment or trailing slash. http: is admitted because Corestore terminates no TLS today and the hop is cluster-private by construction; https: is admitted so putting TLS in front of Corestore needs no contract change. Absent means the node runs no Corestore the relay may reach, which an inventory reports as errorCode: 'CORESTORE_ENDPOINT_UNKNOWN' rather than guessing a hostname.

An administrator sets it with setClusterNodeCorestoreEndpoint { identity; nodeId; endpoint }, answered by Cloudly with the node as it now stands. Cloudly validates the endpoint with validateClusterNodeCorestore before it writes, takes the node's own real-write fence inside the writing transaction so a concurrent deletion conflicts instead of leaving an endpoint behind for a node that is gone, and pushes the cluster configuration afterwards so every connected relay learns the change without reconnecting — the same push that the cluster relay block uses when an operator enables or disables a relay. endpoint: null clears it, which states the same fact as never having set one. There is no node-wide update request: one field with one validator cannot drift into a generic node mutation whose per-field authority nobody defined.

The endpoint says where, never with what. The bearer control token that authorizes every call is obtained from Cloudly through getCorestoreControlCredentialMaterial, sealed to the X25519 recipient the relay enrolled with beginSecretRecipientEnrollment and completeSecretRecipientEnrollment — see Corestore Runtime Credentials. No credential travels in the cluster document, and none is read from the relay's environment.

Corestore Legacy Resource Adoption

A Corestore root converted from a release before 32 still holds the databases and buckets /resources/provision created, one per service and capability, in its manifest. Corestore refuses a credential binding beside such a resource until an operator authorizes that exact binding to adopt it, offline, with a plan file; a fresh binding would otherwise start from an empty database or bucket while the data stays behind. These contracts carry that flow across Corestore, the relay, Cloudly and Coreflow, all exported from data and platform:

  1. Inventory. Corestore answers GET /control/legacy-resources (corestoreLegacyAdoptionContract.inventoryPath, bearer-authenticated like every control path) with ICorestoreLegacyResourceInventory { services }: every service, sorted by id, with one or two TCorestoreLegacyResource entries, database before objectstorage. A database is named by databaseName, a bucket by bucketName and retention (its IObjectStorageRetentionIntent, or null). Each carries its adoption: { state: 'unadopted', blocker } or { state: 'authorized', bindingId }. The blocker names what would stop the offline adoption, null when nothing would: not-adoptable (an allocation-managed database, another provider, a resource name that is not its database or bucket name, or a bucket name that is not a canonical S3 name), ordinary-credentials-rotated, owned-by-binding or fence-lineage-conflict (the bucket's object-storage fence lineage cannot move to the adopting binding: the bucket has more than one lineage, or its lineage does not validate, belongs to another authority or has an unfinished mutation). Only a bucket has a fence lineage, so a database never names fence-lineage-conflict. An adoptable bucket must have a canonical S3 name, because a binding request carries only such a name. A consumed adoption removes the legacy entry, so it never appears.

  2. Relay. getClusterCorestoreLegacyResources { nodeId } asks the relay for one node's inventory, answered as TClusterCorestoreLegacyResourceInventory: reachable: true with the services, or reachable: false with errorCode CORESTORE_ENDPOINT_UNKNOWN or CORESTORE_UNREACHABLE and no list. getCorestoreLegacyResources { identity, clusterId, nodeId } is the administrator's read through Cloudly, with inventorySha256 (computeClusterCorestoreLegacyResourceInventorySha256, over the node id and the services, not the read time; null for an unreachable node).

  3. Adoption. adoptCorestoreLegacyResources { identity, clusterId, nodeId, expectedInventorySha256, resources } is the only path that creates a binding carrying a legacy bucket name. resources is a sorted, unique { serviceId, capability } selection (validateCorestoreLegacyResourceSelections). Cloudly reads the inventory again, requires the expected digest, and creates each adopting binding in one transaction: enabled, placed on nodeId, with corestoreLegacyAdoption { nodeId, resourceName, inventorySha256 }, and for a bucket objectstorageBucketName set to the legacy bucket and objectstorageRetention.intent set to the legacy intent exactly, or absent when the legacy bucket has none. validatePlatformBindingCorestoreLegacyAdoption(binding, inventory, sha) is that rule. A resource whose adopting binding already exists is answered unchanged. upsertPlatformBinding keeps refusing objectstorageBucketName, and refuses corestoreLegacyAdoption and statusReason as well. The answer lists the bindings and the plan. Refusals travel as ICorestoreLegacyAdoptionErrorData { code, retryable }: legacy-inventory-unreachable (the only retryable one), legacy-inventory-changed, legacy-resource-absent, legacy-resource-blocked, legacy-resource-authorized-elsewhere, legacy-resource-adopted-elsewhere, service-absent, binding-exists, successor-organization-mismatch, successor-not-greenfield and successor-owns-legacy-resource.

    Successor. A selection entry may name successorServiceId, another service that replaces the legacy owner: a fresh immutable-deployment service of the same organization, created by its greenfield reserve, that takes over the legacy data. The adopting binding is then the successor's, and its corestoreLegacyAdoption.legacyServiceId names the legacy owner (present only when the owner is another service; never the binding's own). The successor must own no legacy resource of that capability and hold no other binding of it, and a legacy resource is adopted by one service only. The selection refuses two entries adopting into the same service capability.

  4. Plan. ICorestoreLegacyResourceAdoptionPlanEntry is exactly the { serviceId, capability, bindingId } object Corestore's legacy-resource-adoption migration reads, plus successorServiceId when the binding belongs to a successor; serviceId is always the legacy owner. validateCorestoreLegacyResourceAdoptionPlan applies that migration's rules: a non-empty array of at most 4096 exact entries, canonical identifiers, no legacy service capability, adopting service capability or binding id twice, any order. createCorestoreLegacyResourceAdoptionPlan(bindings) builds it from the adopting bindings, sorted by legacy service then capability. The plan restates no legacy name or retention intent: Corestore holds those itself, and the binding request that later reaches it carries the bucket and intent it compares.

  5. Status. Until the operator has run the migration, Corestore refuses the binding with a 409 whose body is { ok: false, error, code }. readCorestoreLegacyAdoptionRefusal(statusCode, body) reads code as a TPlatformBindingStatusReason, and Coreflow reports it through updatePlatformBindingStatus { statusReason }; validatePlatformBindingStatusReason pairs each reason with its one status (getPlatformBindingStatusOfReason):

statusReason Status Corestore refusal
legacy-adoption-pending provisioning the service owns a legacy resource of this capability that no adoption authorizes
legacy-adoption-conflict failed the adoption authorizes another binding id, or it was already consumed
legacy-adoption-mismatch failed the legacy identity changed since authorization, its database tenant is absent, or the bucket or retention intent differs

IPlatformBinding.statusReason holds the reported reason; a status report without one clears it. While a binding of a service reports legacy-adoption-pending, the service's runtime status waits as legacy-adoption-pending instead of bindings-pending.

Object-Storage Binding Environment

An object-storage binding gives its workload two things. The access key and the secret reach it only as Secrets: the six aliases createCorestoreCredentialSecretValues() expands (S3_ACCESS_KEY, S3_ACCESS_KEY_ID, S3_SECRET_KEY, S3_SECRET_ACCESS_KEY, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY), published through the sealed Secret set. Where the bucket is, and which retention it carries, is non-secret and reaches it through the runtime configuration: Cloudly composes it into IRuntimeConfig.environment, the complete public environment, so no further runtime spec or assignment field carries it. The names are the ones Corestore and the cluster relay handed workloads before, so an application that read them keeps working unchanged.

Name Value
S3_ENDPOINT, AWS_ENDPOINT_URL The endpoint origin, for example http://192.0.2.15:9000.
S3_ENDPOINT_HOST The origin's host; an IPv6 literal keeps its brackets.
S3_PORT The origin's port, 80 or 443 when the origin states none.
S3_USE_SSL true for https:, false for http:.
S3_REGION, AWS_REGION The signing region.
S3_BUCKET IPlatformBinding.objectstorageBucketName.
AWS_S3_FORCE_PATH_STYLE Always true: a node-local origin has no wildcard DNS.
S3_RETENTION_MODE intent.mode, always compliance.
S3_RETENTION_POLICY_ID intent.policyId.
S3_RETENTION_DURATION_SECONDS intent.retentionDurationSeconds.
S3_RETENTION_INTENT_SHA256 intentSha256.
S3_RETENTION_SENTINEL_KEY receipt.sentinelKey.
S3_RETENTION_CONFIGURED_AT receipt.configuredAt, Unix milliseconds.
S3_RETENTION_SENTINEL_CREATED_AT receipt.sentinel.createdAt, Unix milliseconds.
S3_RETENTION_SENTINEL_RETAIN_UNTIL receipt.sentinel.retainUntil, Unix milliseconds.
S3_RETENTION_SENTINEL_PAYLOAD_SHA256 receipt.sentinel.payloadSha256.
S3_RETENTION_SENTINEL_METADATA_SHA256 receipt.sentinel.metadataSha256.
S3_RETENTION_SENTINEL_ETAG receipt.sentinel.etag.

Every number is a canonical decimal. objectStorageEndpointEnvironmentKeys, objectStorageRetentionEnvironmentKeys and objectStorageBindingEnvironmentKeys publish the lists in this order, with the matching T…EnvironmentKey and T…Environment types.

Where the endpoint comes from. Corestore is node-local, so the endpoint is a fact about the node that holds the bucket: IClusterNode.data.corestoreObjectStorage is an optional IPlatformObjectStorageEndpoint { endpoint, region }, checked by validatePlatformObjectStorageEndpoint() (a canonical http: or https: origin and a lowercase region). It sits beside corestore, not inside it: the relay dials the control endpoint and a workload dials this one, and IClusterNodeCorestore keeps its exact schema, so a relay that validates it goes on accepting every node document. An administrator sets it with setClusterNodeCorestoreObjectStorageEndpoint { identity; nodeId; objectStorage }, answered by Cloudly with the node as it now stands; objectStorage: null clears it. The request is separate from setClusterNodeCorestoreEndpoint, so setting the control endpoint never touches it. A node without it cannot give a workload an object-storage environment; nothing derives a port, a scheme or a region.

Composing. createObjectStorageBindingEnvironment(binding, endpoint) returns the nine endpoint names for a binding that declares no retention, and all twenty when it declares retention with evidence. It refuses by name, and never guesses, when the binding is not objectstorage, carries no canonical bucket, or declares retention without evidence: a bucket whose immutability is not proven yet must not be handed to its workload as if it were ordinary storage, so the workload waits for the evidence. The evidence is validated with validateObjectStorageRetentionEvidence() against the binding's own service, binding id, bucket and intent, and the reconciliation generation and request digest it states. The split helpers are createObjectStorageEndpointEnvironment(endpoint, bucketName) and createObjectStorageRetentionEnvironment(evidence, expectation). mergeObjectStorageBindingEnvironment(environment, bindingEnvironment) adds one binding's names to a workload environment and refuses, naming every key, a name that is already set — by the service itself or by a second object-storage binding — instead of letting one value replace another.

Reporting. composeObjectStorageServiceEnvironment(serviceEnvironment, sources) composes a whole service and answers in the service runtime vocabulary, so a producer never classifies the errors above. Each IObjectStorageBindingEnvironmentSource { binding, endpoint } is an enabled object-storage binding and the corestoreObjectStorage of the node it is placed on, null when that node states none; a binding placed on no node is not a source, and its service waits as bindings-pending. The answer is a TObjectStorageServiceEnvironmentComposition:

status reason When
composed — environment is the service's own with every binding's names added.
waiting objectstorage-retention-pending A binding declares compliance retention whose evidence has not arrived.
refused objectstorage-endpoint-missing The binding's node states no valid corestoreObjectStorage.
refused objectstorage-binding-invalid createObjectStorageBindingEnvironment() refuses the binding: not objectstorage, no canonical bucket, no canonical retention intent, or evidence that does not prove this binding.
refused objectstorage-environment-conflict The service's own environment or another binding already sets one of the binding's names.

waiting and refused carry the bindingId and the contract's message. Bindings are taken in id order, so the verdict does not depend on the order of the sources, and a refusal outranks a wait: a binding waiting for its evidence already claims all twenty names, so a collision with it is refused now rather than once the evidence arrives. The reasons are TObjectStorageEnvironmentWaitingReason and TObjectStorageEnvironmentRefusal, members of TServiceRuntimeWaitingReason and TServiceRuntimeRefusal.

Validating. validateObjectStorageBindingEnvironment() accepts exactly the endpoint names and all or none of the retention names, each value the one its endpoint and evidence derive, and refuses any other name — so a credential alias in the public environment is refused by name. validateObjectStorageEndpointEnvironment() and validateObjectStorageRetentionEnvironment() judge each half on its own. The environment does not carry the evidence's authority, so validating it checks what a workload reads, not the proof behind it.

Database Binding Endpoint

A database binding's connection values embed its password, so they are Secrets as a whole and are addressed once, when the provider produces them — not composed into the public environment like an object-storage endpoint. What they are addressed to is a fact about the node, owned by Cloudly:

  • IPlatformDatabaseEndpoint { host, port } is where a workload on a node reaches the node's Corestore database. host is an IPv4 literal, an IPv6 literal without brackets or a lowercase DNS name, exactly as a URL states it; port is 1 to 65535. validatePlatformDatabaseEndpoint() names every reason it is refused, and formatPlatformDatabaseEndpointHost() writes the host as a URL authority states it.
  • IClusterNode.data.corestoreDatabase is optional and sits beside corestore and corestoreObjectStorage. An administrator sets it with setClusterNodeCorestoreDatabaseEndpoint { identity; nodeId; database }, answered with the node as it now stands; database: null clears it. Setting another Corestore endpoint never touches it.
  • The database variant of TCorestoreCredentialBindingRequest carries an optional endpoint. When present it is validated and included in the request digest; a request without it keeps the digest earlier releases computed. A moved endpoint is therefore a different request, which is granted and published again. Whether a request without an endpoint is admitted is the provider's decision; this contract supplies no default host.
  • createCorestoreDatabaseUri(endpoint, { databaseName, username, password }) is the one URI form: mongodb://<username>:<password>@<host>:<port>/<databaseName>?authSource=<databaseName>&directConnection=true, with URI-component encoding and an IPv6 host in brackets. createCorestoreDatabaseCredentialValues(endpoint, identity) derives all eleven corestoreDatabaseCredentialKeys from it (MONGODB_HOST is the host without brackets, MONGODB_PORT the decimal port).
  • validateCorestoreDatabaseCredentialMaterialEndpoint(material, endpoint) accepts database material only when every value is the one derived from the endpoint and the material's own database name, username and password, and refuses by key name — never by value — material addressed elsewhere, such as Corestore's own public host.
  • TServiceRuntimeRefusal gains database-endpoint-missing: an enabled database binding placed on a node that states no valid corestoreDatabase is neither granted nor published, and its workload is refused rather than started with a guessed host.
  • A workload reaches that endpoint only when its attached network selects it as a platform endpoint; an unselected valid endpoint is refused as database-endpoint-unselected (see Service Runtime Specs And Status for the matching rule), as an object-storage endpoint is as objectstorage-endpoint-unselected.

Launcher Environment Delivery

Launcher environment delivery uses stable szsv-<base32-sha256> Docker resource names and /run/serve.zone/secrets/<resource> source paths. The nonsecret map is written to /run/serve.zone/workloadinit-map.json with mode 0444. Runtime assets under /opt/serve.zone/runtime-assets are read-only, the wrapper itself lives at /opt/serve.zone/runtime-assets/workloadinit/workloadinit, and it executes workloadinit run --map /run/serve.zone/workloadinit-map.json -- <argv...> without a shell, symlink, or aggregate value file. A resource name is the lowercase base32 of a SHA-256 over serve.zone/docker-secret-resource, a NUL and the SecretVersion id, so every launcher secret of a service is republished under a new Docker resource name when a cluster moves to 32.0.0. Resources and maps remain retained while any desired, applied, or previous-accepted manifest references them and are deleted only after Docker confirms rollout or service removal.

Platform Contracts

Use platform for current platform-service capabilities and application-facing platform RPCs.

import { platform } from '@serve.zone/interfaces';

type SendEmailRequest = platform.email.IReq_SendEmail;
type PlatformBinding = platform.IPlatformBinding;

const sendEmailMethod: SendEmailRequest['method'] = 'sendEmail';

Available platform modules:

  • platform.email for transactional email, recipient registration, email status, and email stats.
  • platform.sms for SMS delivery and verification-code delivery.
  • platform.pushnotification is the deprecated legacy device-token push contract. New browser Web Push integrations use requests.webpush.
  • platform.letter for physical letter workflows.
  • platform.ai, platform.database, platform.objectstorage, platform.logging, platform.backup, and platform.sip for infrastructure and application capabilities.
  • platform.storage for provider-neutral storage classes, requests, capabilities, and resolved bindings.
  • platform.objectstorageretention for value-free immutable-retention intent, capability and sentinel evidence, authority binding, digest helpers, and strict validators.
  • platform.objectstorageenvironment for the non-secret workload environment of an object-storage binding; see Object-Storage Binding Environment.
  • platform.storagemigration for fenced object-storage migration intent, status, consumer acknowledgements, canonical digests, and strict normalizers.
  • platform.types provider and binding metadata is value-free. Provider-specific operational config stays adapter-internal, while public DTOs expose typed endpoints and optional credential management scopes only.

Optional IPlatformBinding.objectstorageBucketName is the durable value-free Cloudly authority for an exact Corestore bucket. It is required whenever a caller needs to establish trusted bucket authority for retention validation. Cluster runtimes report the exact provider-returned bucket through the optional requests.platform.IReq_Any_Cloudly_UpdatePlatformBindingStatus.request.objectstorageBucketName field when updating binding status. validatePlatformObjectStorageBucketName() applies the canonical ObjectStorage S3 bucket boundary without normalizing input, and platformObjectStorageBucketNameLimits publishes its 3-byte minimum and 63-byte maximum. IPlatformBinding.objectstorageRetention carries compliance-mode intent and optional provider evidence bound to the exact service, binding, reconciliation generation, request digest, and authoritative bucket. Use validatePlatformBindingObjectStorageRetention() for the full binding boundary or the narrower intent/evidence validators when the trusted authority context is already available; none of these contracts contains credential material.

Legacy Platformservice Contracts

platformservice is retained for older integrations that still consume the previous namespace layout.

import { platformservice } from '@serve.zone/interfaces';

type LegacySendEmailRequest = platformservice.mta.IRequest_SendEmail;

New code should prefer platform unless it must remain compatible with an active legacy consumer.

Contract Ownership

Only ecosystem-wide public contracts belong in this package. Cloudly-internal implementation details, service-private DTOs, and temporary migration helpers should stay in their owning service until they become real shared contracts.

Good candidates for this package:

  • Types persisted or exchanged across multiple serve.zone services.
  • TypedRequest contracts used by more than one project.
  • SDK-facing interfaces that external consumers should be able to rely on.

Poor candidates for this package:

  • Private implementation details of one service.
  • Runtime helpers or convenience wrappers, unless they are shared contract normalization or validation helpers used by multiple packages.
  • Compatibility aliases without an active consumer.

Development

pnpm install
pnpm run build
pnpm test
pnpm run buildDocs

The package is authored as ESM TypeScript and built with tsbuild tsfolders.

This repository contains open-source code licensed under the MIT License. A copy of the license can be found in the license.md file.

Please note: The MIT License does not grant permission to use the trade names, trademarks, service marks, or product names of the project, except as required for reasonable and customary use in describing the origin of the work and reproducing the content of the NOTICE file.

Trademarks

This project is owned and maintained by Task Venture Capital GmbH. The names and logos associated with Task Venture Capital GmbH and any related products or services are trademarks of Task Venture Capital GmbH or third parties, and are not included within the scope of the MIT license granted herein.

Use of these trademarks must comply with Task Venture Capital GmbH's Trademark Guidelines or the guidelines of the respective third-party owners, and any usage must be approved in writing. Third-party trademarks used herein are the property of their respective owners and used only in a descriptive manner, e.g. for an implementation of an API or similar.

Company Information

Task Venture Capital GmbH
Registered at District Court Bremen HRB 35230 HB, Germany

For any legal inquiries or further information, please contact us via email at hello@task.vc.

By using this repository, you acknowledge that you have read this section, agree to comply with its terms, and understand that the licensing of the code does not imply endorsement by Task Venture Capital GmbH of any derivative works.

S
Description
Shared TypeScript data and RPC contracts for the serve.zone ecosystem.
Readme
31 MiB
Languages
TypeScript 100%