@serve.zone/pallet

Pallet is the in-development node-local containerd execution component for serve.zone. Its public API provides a bounded, read-only native CRI v1 probe. Backend-private identity, enrollment and stateless execution mechanisms are under development; production assignment delivery and reconciliation remain unwired.

Issue Reporting and Security

For reporting bugs, issues, or security vulnerabilities, please visit community.foss.global/. This is the central community hub for all issue reporting. Developers who sign and comply with our contribution agreement and go through identification can also get a code.foss.global/ account to submit Pull Requests directly.

Development status and ownership

Cloudly will own cross-node desired state and placement. Pallet will own local execution, including the outbound Cloudly connection on workers. Onebox will control its own local workloads through authenticated Pallet IPC, without depending on Cloudly to bootstrap itself. Spark remains the host installer and supervisor. Coreflow's required behavior must be ported before its legacy Swarm runtime can be retired.

Those orchestration and authorization capabilities are not implemented by the public probe. It exposes no public listener and persists no application state. Do not replace a production runtime with this development foundation.

No Pallet name carries a version token: collections, singleton row ids, record-id hash domains, the native structs and their TypeScript twins, the dormant link aliases (pallet-handoff: / pallet-workload:), the host and workload evidence kinds and the process protocol names are named by what they hold. No Pallet shape carries a shape-version field either — not the native structs, not the identity, application, guard, attachment, ACTIVE, barrier, tunnel or DNS records, and not the build manifests this repository owns. Every one of those key sets is exact, so a value that still states schemaVersion is refused as an unknown key rather than read as an older shape. What contract a peer speaks is stated once, by the handshake below, and what a persisted shape means is stated by the installed build. No released Pallet version was ever deployed, so nothing is migrated and nothing is aliased — a store written by an older build fails its assertion on every read, a dormant link an older build left behind carries a foreign alias, interface name and MAC address because the seed domain changed with the name, so it is never recognized as this build's pair, and a node's stores under /var/lib/serve.zone/pallet are reset with the install.

Private CRI execution

ts/runtime/classes.executiondriver.ts owns a bounded native CRI client through the separate --executor-management process mode. It is absent from the public package facade and enrollment IPC. The eventual assignment owner must authenticate the controller, validate immutable material, durably admit the assignment and serialize its effects before invoking this mechanism.

The driver supports inspect, run, stop and explicit runtime removal against a dedicated containerd 2.3 release and CRI v1 socket. Run supplies a digest-pinned image, its platform image-config digest, exact argv, environment, working directory, CPU/memory policy, user/group policy and root filesystem write policy. Explicit cpuMillis: null leaves CPU quota unlimited. Explicit null for both runAsUser and runAsGroup selects the pinned image's User through containerd's rootfs resolution; a numeric pair overrides it. Missing fields and partial pairs are invalid.

Run also states the sandbox's resolver, dns: { servers, searches, options }, and the executor always sends it as the CRI sandbox's dns_config: containerd copies the host's /etc/resolv.conf only for a sandbox that states none, and a node whose host file names the systemd-resolved stub 127.0.0.53 would hand every workload a server it cannot reach (lab 2026-10-05 RUN 8, D10). An attached run's resolver is derived from its durable network preparation (PalletNetworkOwner.readExecutionResolver): its single server is the lease's DNS server — the node's per-workload resolver, which answers the private names of the run's networks and forwards public names as the contract's DNS view for the lease decides — and its search domain is the default network's suffix without the root dot, so containerd writes search <suffix> (only with a default network) and nameserver <lease DNS server>. An isolated run states three empty lists, for which containerd writes an empty file. The executor refuses any other shape as invalid input before any CRI call: an isolated run with a server, an attached run without exactly one canonical unicast IPv4 server, more than six search domains or a domain that is not a lowercase DNS name without the root dot. A run whose network owner answers no resolver is not started (unresolved-authority). Under @serve.zone/interfaces 32.48.0 a workload on no private network forwards any public name to the protected authority's resolvers (the address plan's), so a cluster declares the resolvers its nodes' workloads use there; with none declared such a workload resolves no name and its resolver answers REFUSED. Containers use private namespaces, runtime-default seccomp and no-new-privileges, without privileged mode or host networking; the only host paths a container ever receives are the read-only per-file binds of its own staged secret mount and one private bind per node-local volume it was granted ("Local storage claims"), which the executor derives from its fixed volume root and the volume's key — no host path crosses the IPC. Every sandbox names the absolute cgroup parent /pallet, so workload cgroups sit outside any service unit's cgroup (see Pallet's own containerd). Network policy and application readiness remain separate owners. Secret material delivery is this node's own, through the staged mount.

Every runtime object carries the immutable generation-one run digest and complete controller/node/assignment/replica/attempt ownership labels. Recovery requires one exact sandbox and container, including a full check for foreign sandbox children. Image identity uses the actual config digest and the requested repository digest; display tags are not identity. Running attempts can be recovered; exited or stopped attempts are never restarted. Stop preserves runtime objects, including containers created but never started. Removal requires confirmed stopped state and retains storage; images are left to the image collector below. Neither operation depends on a healthy image cache or network.

Input is captured before asynchronous work without invoking accessors. Native frames and CRI responses have fixed bounds; errors expose static codes rather than runtime messages or registry credentials. Cancellation, parent EOF and signals terminate the owned client. A timeout or client exit does not cancel a server-side CRI operation, prove rollback or fence a writer. The durable assignment owner must retain uncertain effects and resolve them before admitting replacement.

One failure is definitive instead: a PullImage the runtime refuses while the run has no sandbox and no container. Nothing of the attempt exists then, so the native owner answers IMAGE_PULL_FAILED with errorData { grpcCode, cause } — the runtime's gRPC status code and one word of a closed set, unauthorized, not-found, tls, transport, deadline, unavailable or other — and never the runtime's message, which names the registry, the reference and the token endpoint (PalletImagePullError). containerd 2.3 reports a registry's refusal as Unknown with the HTTP status line in its message, so the code decides first and the message's fixed markers after it. With a sandbox already there the same refusal stays a runtime failure. The execution owner ends that run's operation failed (no evidence), frees the native lane at once, observes the assignment failed once — with cause { kind: 'image-pull-refused', reason, grpcCode } (see "Failure causes" below); its controller times a replacement from the observation it accepted — and writes Pallet image pull refused: assignment <id> generation <n>: <cause> (gRPC <code>). to stderr once per assignment and cause. The attempt is never pulled again: its network preparation belongs to that one operation, and a new attempt — which Cloudly admits after its failure backoff — is the controller's. Before this, the refused pull stayed pending, held the one native lane for every other assignment of the node until a reboot fenced it, and its assignment stayed unknown for good.

A sandbox the runtime refuses to create is the same kind of fact when nothing of the attempt is left afterwards: containerd tears a sandbox down when its network setup fails, and the native owner reads the runtime again before it answers. With neither a sandbox nor a container of the attempt left it answers SANDBOX_REFUSED with errorData { grpcCode } (PalletSandboxRefusedError), never the runtime's message; with anything left the refusal stays a runtime failure. What the execution owner does next depends on whether a CNI ADD of the run bound its attachment to the refused sandbox (PalletNetworkOwner.readExecutionSandbox):

  • No ADD bound it — the network owner refused the ADD itself, not ready yet or its epoch not proven yet, which its next pass settles. The execution owner drives the same run operation again — nothing of it exists, so no effect is repeated — after 1 s and then 2 s, at most palletSandboxRefusalAttempts (3) times, and writes Pallet run sandbox refused: assignment <id> generation <n>: attempt <k> of 3 (gRPC <code>). to stderr for each refusal. Before each unbound ADD, Pallet refreshes the run's preparation in a transaction against the currently completed network configuration. A configuration still settling refuses by attachment.reprepareExecution.applicationPending; a lease no longer admitted refuses by its endpoint check. No native descriptor is created in either case. Preparation can change only while no ADD journal exists; refreshing and ADD admission fence the same application lane, including an ADD whose commit acknowledgement was lost. An attached run still unbound after three refusals stays pending, observed uncertain, and writes Pallet run waiting: assignment <id> generation <n>: network-attachment-pending. Its next reconcile, including after a process restart in the same boot, first proves both attachment and runtime absence before driving the same operation again. An unreadable runtime or attachment keeps it uncertain (network-run-unjoined with the owner's failure label).
  • An ADD bound it — the attachment owner journaled the pair, and answered it DOWN because the node's barrier still did not hold after the ADD's own bounded wait, or its raise failed. That binding is the run's for good: the attachment journal is immutable history the DNS and ACTIVE owners verify, so no other sandbox of the run is ever attached, and a re-drive could only be refused (attachment.reconcile.containerMismatch). The run is not driven again. While its operation is still pending, the bound pair is removed as that operation's rollback, as a DEL removes it, so a DEL that comes later finds it absent; the binding stays. The runtime's own DEL of the refused sandbox may have been answered before the ADD journaled, and a DEL after the run's operation ended was refused uncertain for good (Grasberg 2026-10-10). The node writes Pallet run sandbox refused: assignment <id> generation <n>: attempt 1 of 3 (gRPC <code>): attachment-bound, rolled back, not driven again., or attachment-bound, rollback refused < <label>, not driven again. with the removal's owner failure label, and the CNI broker has already written the attachment's own reason (Pallet workload attachment down: run <run digest> ADD: <label>.). An isolated run's binding is not rolled back.
  • Whether an ADD bound it cannot be read. A second sandbox is never driven on a guess: the run is not driven again, and the node writes Pallet run sandbox refused: assignment <id> generation <n>: attempt <k> of 3 (gRPC <code>): attachment unreadable < <label>, not driven again., where <label> is the read's owner failure label (nothing on the read path reports it otherwise).

A refused isolated run, or an attached run whose ADD bound a sandbox or whose binding cannot be read, ends exactly like a refused pull: its operation failed (no evidence) with cause { kind: 'runtime-absent' } — the runtime holds no sandbox or container of a run that should be running — the lane freed, and the assignment observed failed once. Pallet never runs that assignment revision again; the controller replaces the failed attempt with a new one, whose new run digest gets a fresh sandbox and attachment. Before this, the pending run held the lane until a reboot fenced it, and every other run of the node waited behind it as lane-busy. Through Pallet 36.3.0 a barrier lost for seconds after the gate failed the run this way: the ADD answered its pair DOWN at once, the runtime tore the sandbox down, and all three re-drives were refused attachment.reconcile.containerMismatch (production 2026-10-07).

A run whose own sandbox exists and is not ready — its network setup failed and its teardown did not complete, so containerd keeps it NOTREADY, or it stopped — is refused by the native run before any effect as SANDBOX_STOPPED (PalletSandboxStoppedError). A sandbox never becomes ready again and an attempt is never recreated after its first effect, so the refusal is definitive: the run ends as a pending run whose sandbox stopped does in its join — recorded failed (PalletAssignmentStore.failEndedRun), an attached run whose network epoch the network owner reads as ended stating network-epoch-ended, any other no cause, written as Pallet run failed: assignment <id> generation <n>: sandbox-stopped., the lane released, observed failed once — and the controller's stop and removal take the sandbox away. Pallet 36.2.0 answered it RUNTIME_STATE_CONFLICT; the run waited run-uncertain, held the node's one lane until a reboot, and every other run and the removal of its own service waited lane-busy behind it (lab 2026-10-06 RUN 9, D11).

A run starts only while its network can take it. Before its first effect — the registry credential, the volumes' transaction, the pull — an attached run asks the network owner whether the workload datapath is up, by the admission gate's own reading (PalletNetworkOwner.executionBarrier), and an isolated run whether the network owner is ready. A refusal is a wait, never a failure: nothing of the run is begun, the lane stays free, and the node writes Pallet run waiting: assignment <id> generation <n>: network-barrier:<site>. with the barrier's own admission.barrier.* site. Pallet 36.0.2 drove a run admitted before a process restart while its new network owner was still raising the generation; the CNI plugin answered Pallet network owner unavailable, the three re-drives were spent in seven seconds, and a run with nothing wrong with it failed runtime-absent. A barrier lost between the gate and the sandbox is waited out by the sandbox's own CNI ADD, bounded (see the workload attachment section).

A running attached run whose pair belongs to a network epoch that ended fails for good. The node process restarted and the router namespace went with it, taking the sandbox's link, while CRI still reports the address the sandbox was given, so nothing on the runtime side shows the loss. The execution owner asks the network owner for the run's epoch on every reconcile of a running attached run, and gets a closed answer: current, ended, none (nothing of the run is attached on this node, which leaves nothing to fence) or unknown (not provable yet: the network owner has not started, or the run's pair is not journaled, its add is not complete, or it was removed). An unknown epoch is waited on, never taken as current: the reconcile stays uncertain and writes Pallet run waiting: assignment <id> generation <n>: network-epoch-unknown. An ended one is recorded failed with cause { kind: 'network-epoch-ended' } (the native lane must be free; it is fenced, not taken), written to stderr as Pallet run failed: assignment <id> generation <n>: network-epoch-ended., and its sandbox is stopped — retried from the durable failure until the runtime shows it not running. The run is observed failed once, stating that cause; a new attempt, the controller's, attaches in the current epoch. The controller's stop and removal of the failed attempt run as for any other attempt, each with its terminal receipt, so its slot advances.

A run whose sandbox stopped while its controller still asks it to run fails for good as well. Nothing of this node stops a run under disposition run without failing it first, so a stopped sandbox there — every sandbox after a reboot, or one whose pause process the runtime lost — ended without a stop intent; it never runs again, and an attempt is never restarted. Its completed run operation, of this boot or an earlier one, is recorded failed (the native lane must be free; it is fenced, not taken), the node writes Pallet run failed: assignment <id> generation <n>: <cause>., and the run is observed failed once; a new attempt is the controller's. An attached run whose pair was attached in an earlier boot (a reboot ends every network epoch) or whose epoch the network owner reads as ended states network-epoch-ended; any other states no cause, written sandbox-stopped on stderr, because the controller's vocabulary has none for it. The failure is certain, so an epoch the network owner cannot prove yet only leaves the cause unnamed. Pallet 36.1.0 observed such a run stopped under disposition run for good, which Cloudly never replaces: after a reboot its attached workloads stayed down until their desire changed.

A run a reboot fenced before it completed ends the same way. The first start of a new boot fences the run operation the earlier boot left pending (host-fenced): that boot stopped every sandbox and ended every network epoch, and an attempt is never run again. The run is recorded failed (the native lane must be free; it is fenced, not taken), written to stderr and observed failed once — an attached run stating network-epoch-ended, with its sandbox stopped should one still run, any other stating no cause (sandbox-stopped on stderr) — and the controller's stop, removal and next attempt follow. Pallet 36.1.2 and earlier left such a run waiting operation-pending for good.

A launcher's WorkloadInit image. A launcher plan mounts the approved WorkloadInit image as a CRI image volume at /opt/serve.zone/runtime-assets/workloadinit, and containerd resolves an image volume only from its own store; it never pulls one, and refuses the container while it does not hold the image. The run pulls it by the plan's digest-pinned reference — code.foss.global/serve.zone/workloadinit@<platform manifest digest>, the release contract's registry, which serves it anonymously — after the workload image and before anything of the attempt is created, without a credential, and requires the stored image to list that reference. An image the store already holds under that exact reference is used as stored and not pulled again: the digest names its content and the pull proves no credential, so asking the registry again would only let a registry that cannot answer fail a run whose image is on the node (lab 2026-10-05 RUN 6, O2). The workload image is pulled on every run, with the registry credential the controller grants for it. Its registry refusing it while nothing of the attempt exists ends the run like a refused workload pull, image-pull-refused; with a sandbox already there, as when a run is resumed, it stays a runtime failure. The node needs egress to that registry: the cluster relay forwards only the controller's own registry.

A container's mount list. Before it starts a created container, and on every later inspection, the native owner demands that the container's mount list, as ContainerStatus answers it, is exactly the attempt's: the plan's secret binds, the launcher's map and WorkloadInit image volume, then one bind per granted volume, in that order. containerd 2.3 answers the list it stored at CreateContainer, in the request's order, after resolving each image volume in place: it sets the volume's host_path to its own mount point (<state>/io.containerd.grpc.v1.cri/image-volumes/<sandbox id>/<manifest digest>) and clears its uid and gid mappings, and echoes every other field as requested. That host path is the runtime's answer, never Pallet's, so an image volume is compared without it; every field Pallet sets — container path, image reference and the rest of its image spec, sub path, readonly, recursive_read_only, propagation, SELinux relabel, id mappings — and every bind's host path compare exactly, so a container that differs in any of them, or carries a mount more or less, or in another order, is refused. Pallet 36.1.4 compared the lists exactly, host path included, and refused every launcher run's own container before its start as foreign (OWNERSHIP_MISMATCH, lab 2026-10-04, N6). A refused mount list answers MOUNT_MISMATCH (PalletMountMismatchError). Nothing changes a created container's mounts, so the refusal is definitive for the attempt: the run, from the drive that created the container or from the join of a drive whose reply was lost, is recorded failed with no cause (the controller's vocabulary has none for it), written as Pallet run failed: assignment <id> generation <n>: mount-mismatch., its lane released, and observed failed once; the container is never started, and the controller's stop and removal take it away and release the mount as for any failed attempt. Before, the run stayed pending and held the node's one native lane, so every other run of the node waited lane-busy behind it.

Native failures by name. Every other failure of a native run, stop or removal answers one of the native owner's closed codes, and a failed CRI call adds errorData { call, grpcCode, cause }: the call (create-container, pull-image, run-pod-sandbox, …), the gRPC status code and one cause word — image-volume-unresolved (containerd does not hold an image an image volume names), name-reserved (it still reserves the sandbox or container name for an earlier request), deadline or other — never the runtime's message. The execution owner names it as PalletExecutionRuntimeError:<code>:cri.<call>.<cause>.<gRPC code> (executionRuntimeErrorOf) in the run's run-uncertain line.

A secret run's drive that ended short of its container. A run with secret material publishes its mount under its pending operation before its drive, so a drive that fails after that — the runtime created the sandbox and refused the container, or a reply was lost — leaves the operation pending with the mount published. Its next reconcile, in the same process or a later one of the same boot, joins it: a fresh helper must find exactly the recorded mount, or the run stays uncertain (secret-run-unjoined, or secret-run-unjoined < <owner failure label> when the join failed, a runtime that could not be read for example) with every record kept. A container that runs, or ran and exited, settles the operation with that evidence. A sandbox that stopped never runs again, and the native run refuses to take one up, so the run fails once as a completed run whose sandbox stopped does: recorded failed with its lane released, written to stderr and observed failed, an attached run whose epoch the network owner reads as ended stating network-epoch-ended, any other no cause; Pallet 36.1.2 and earlier resumed it and it waited run-uncertain until a reboot. Anything else is driven again as the same operation, without opening the material again: the native run takes it up from exactly what the runtime holds — the ready sandbox, a container created and never started, or nothing yet — so no effect that exists is repeated, and a refused pull or bound sandbox ends it failed as on a first drive. An attached run with no ADD binding stays pending after its bounded sandbox retries, as above. Every failed drive of a run is named, Pallet run waiting: assignment <id> generation <n>: run-uncertain < <owner failure label>. Pallet 36.1.2 and earlier settled only a running container there, and a run whose container the runtime refused once waited secret-run-unjoined until a reboot (lab 2026-10-04, N3), and then operation-pending.

A stop of a run this node never received. The controller can settle an attempt whose run was sent but never delivered: its stop (generation 2) is then the first revision of that attempt this node sees, and the shared decision (evaluateRuntimeAssignment) admits no generation 2 without its generation 1. A settlement revision repeats its run's immutable attempt and changes only its generation, previous, disposition and digest, so the node rebuilds the run from the stop and admits both revisions in one transaction only when the rebuilt run's digest is the one the stop's previous names — each through the same decision as any admission, and through the same slot, predecessor and session fences. Nothing of the run ever began here: its stop finds nothing (absent), is observed stopped with its terminal receipt, and the controller's removal follows as for any other stop. A first-seen removal is never rebuilt, since its stop receipt is the node's own. Pallet 36.1.0 refused such a stop as a siteless PalletAssignmentError:conflict on every delivery, and the controller sent it again every 6.5 s for good.

Every conflict of the assignment store names its site, assignment.<method>.<check> (TPalletAssignmentStoreSite), beside the execution admission barrier's admission.barrier.*; PalletAssignmentError requires one for conflict at the type level. An admission the shared decision rejects names its reason: assignment.admit.predecessorMismatch, .stale, .revisionConflict, .attemptMismatch, .dispositionRegression, .scopeMismatch or .invalid.

Failure causes. A failed observation states why the run failed in cause (@serve.zone/interfaces TRuntimeAssignmentFailureCause, value-free) whenever the node can name it: image-pull-refused with the classification and gRPC status of a refused pull, network-epoch-ended for a run stopped because its network epoch ended or an attached run whose sandbox a reboot stopped, exited with the container's exit status (0..255, or null for any other) for a container that ended on its own, and runtime-absent for a run the runtime no longer holds. image-pull-refused, network-epoch-ended and runtime-absent are kept on the run's operation record (cause, the same shape), so the observation states them whenever it is published, after a restart too; a failed row 35.x wrote carries none, and its observation states none. A cause explains a failure and changes nothing about it. A controller must read @serve.zone/interfaces 32.40.0 or later (Cloudly 33.9.0 or later) before this Pallet runs against it: a receiver on 32.39.0 or earlier refuses an observation that states a cause.

A run that waits before its pull says what it waits on: Pallet run waiting: assignment <id> generation <n>: <step>., where the step is lane-absent, lane-other-boot, lane-busy (another operation holds the native lane), operation-in-flight (this run's own unfinished intent holds it), operation-pending, operation-not-begun, secret-run-unjoined, secret-run-unjoined < <owner failure label>, secret-mount-uncertain, secret-mount-uncertain < <owner failure label>, storage-pending:<site>, registry-credential < <owner failure label> or run-uncertain < <owner failure label> (the run's drive failed; a native failure reads PalletExecutionRuntimeError:<code>[:cri.<call>.<cause>.<gRPC code>], see below). A new step is written at once and a repeated one again with its count, under the same repeat rule as refused controller requests. A staging failure names its chain: a controller that refused the secret-material request appears as PalletControllerRefusalError:<reason>:controller.refusal.getRuntimeAssignmentSecretMaterial, where <reason> is the check the controller named in its refusal. A failure before the run's operation began leaves nothing behind, so the next reconcile requests the material again; one after it leaves the pending operation to the next boot, which fails it, or — once the mount is published — to the join that drives it again.

A reconcile that refuses outright — an isolated run with a TCP or HTTP readiness probe (unsupported-readiness), an authority the node cannot resolve (unresolved-authority), an assignment it cannot read (invalid), or any other failure outside a named wait — is written Pallet run refused: assignment <id>: <owner failure label>., once per assignment and refusal: again only when the refusal changes or after a reconcile of the assignment resolved. A refusal its wait line already names (storage-pending:<site>, registry-credential < …) is not written twice. The controller learns none of these from an observation: TRuntimeAssignmentFailureCause has no cause for a refusal before any runtime effect, and a failed observation without one would only have the controller replace the attempt — for unsupported-readiness with one the node refuses the same way. Pallet 36.1.2 and earlier wrote nothing for them; the node runtime only counted them in its pass status (rejected, unavailable, unresolved-authority).

The pull's credential reaches containerd as AuthConfig.server_address https://<host[:port]>, the registry host of the reference it pulls. containerd's CRI hands a credential only to the registry whose host equals url.Parse(server_address).Host; Go reads a bare host:port as a scheme and an opaque part with an empty host, so a bare address matched no registry and every pull ran anonymously.

pnpm exec tstest test/test.execution.node.ts --verbose --logfile --timeout 60

The Unix fixture covers exact execution and recovery, lost mutation replies, foreign/ambiguous resources, image mismatch, degraded shutdown, never-started containers, malformed IPC, parent death and cancellation. These mechanism tests do not establish durable assignment admission or authorize a production cutover.

Image collection

Pallet's containerd holds only the images Pallet pulled for its own runs and the bundled sandbox image. PalletExecutionOwner.collectImages removes the images no durable assignment still needs. A collection is due when a lifetime starts and after every removal the owner proves. The node runtime runs it at the end of a full assignment sweep, on the same native lane as a reconcile, so nothing pulls or creates a container meanwhile.

The owner states what to keep. Every assignment whose runtime removal is not proven keeps its run's image config (platformEvidence.imageConfigDigest), whatever its disposition: a stopped attempt is still inspected and removed with its image, and an admitted one still pulls it. If any retained assignment's image cannot be named, nothing is collected. The native pass (collectExecutionImages, rust/src/execution.images.rs) also keeps:

  • every image a CRI container still uses, in any state and matched by any name the image carries;
  • every image the runtime reports pinned;
  • the sandbox image by its bundled name.

It removes the rest in id order with CRI RemoveImage and requires ImageStatus absence afterwards; an acknowledged removal that left the image in place fails the pass as RUNTIME_STATE_CONFLICT. It reports the images listed, removed with their size, kept by reason, left for the next pass, and the bytes the image filesystem reports in use.

Bound Value
retained image configs per pass 1,024
images the runtime may hold when a pass begins 1,024 (refused as RUNTIME_STATE_CONFLICT beyond)
images removed per pass 64; the rest are remaining and keep the collection due

The node runtime runs the collector at the end of a full assignment sweep, and only while the node holds an authenticated controller session: before the first one a node admits no native effect, so a due collection makes no CRI call and runs on the first sweep after the session is established. The collector runs only while the durable execution lane is idle on this boot, because an operation still pending there may be a run whose pull the runtime is finishing. A pass that cannot run stays due and is tried again after at most one minute (palletImageCollectionRetryMs). It is recorded on the node status images as one of:

  • deferred with lane-pending;
  • failed with retain-unresolved, retain-bound, or runtime plus the native code.

A lane or assignment store that cannot be read is recorded the same way: an unreadable lane is not proven idle and is deferred with lane-pending, an unreadable retained assignment is failed with retain-unresolved. Besides its own refusals (no controller session, a busy lane, an owner not ready), the collector rejects only when the identity runtime refuses its assignment store or its own code fails, a defect the status cannot name. The node writes Pallet image collection failed unexpectedly: <class> <code>[:<site>]. to stderr, which Spark forwards to the node journal, once per distinct failure until a collection runs again.

It never fails the node: an image left behind costs disk, not correctness. No journal is kept: removing an unreferenced image is idempotent, and a later run pulls by digest.

pnpm exec tstest test/test.imagecollection.node.ts --verbose --logfile --timeout 120

Workload stats

A controller reads one resource sample of a running attempt with readRuntimeAssignmentStats (@serve.zone/interfaces 32.36.0). The request names the exact current revision of an assignment on this node, and the controller client binds it to the session and that revision before any native call; like every inbound request, it acts only on the connection that holds the session.

PalletExecutionOwner.readStats reads outside the native lane, so a read never waits for or blocks a reconcile. Each read runs its own native child (sampleExecutionStats, rust/src/execution.stats.rs): it proves the attempt's exact CRI identity, calls CRI ContainerStats and PodSandboxStats, and proves the same identity again. A change between the two proofs is a conflict, never a sample of whatever runs now. An attempt without a running container is answered not-running, without a runtime stats call.

A sample states the cumulative CPU time across every core (decimal digits, because it outgrows a JSON number; utilisation comes from two samples), the memory working set, the memory limit the node enforces (the config's memoryBytes), the sandbox's default interface counters (or null) and the writable-layer usage (or null). A counter a JSON number cannot carry exactly is refused as RUNTIME_STATE_CONFLICT.

Bound Value
reads per assignment at a time 1 (another is refused as busy)
reads across the node at a time 8 (palletObservationBounds)
native deadline per read 10 s
pnpm exec tstest test/test.executionstats.node.ts --verbose --logfile --timeout 60

Workload logs

Every attempt's stdout and stderr are captured by the runtime and delivered to the controller, which keeps them (IRuntimeConfig.logs: SmartData owns log metadata, SmartBucket the payloads). The node keeps no log history of its own: its capture files are disposable transport artifacts on tmpfs.

Capture. The sandbox names the attempt's capture directory /run/serve.zone/pallet/logs/<run digest hex> and the container its file workload.log below it; containerd writes the CRI log format there (<RFC 3339 time> <stdout|stderr> <P|F> <content>). An attempt whose container was created without a log path (a container from an earlier Pallet) is stated once per process as a not-captured loss.

Delivery. PalletLogCapture reads the file and sends reportRuntimeAssignmentLogs batches (runtimeAssignmentLogContract, @serve.zone/interfaces 32.36.0) on the session, one at a time: the next batch leaves only after the controller acknowledged the previous one, and an unacknowledged batch is sent again unchanged. A CRI piece longer than one entry (64 KiB) travels as partial entries; a line longer than the config's logs.maximumLineBytes is cut there and followed by a line-cut loss of its stream, counting the cut bytes. Every delivery pass (palletLogDeliveryBounds) sends at most eight batches per capture before the next capture's turn.

Controllers that keep no logs. A controller answers every batch accepted, replay or not-kept. not-kept states that it keeps no workload logs and holds nothing: the node sends that session no more batches (runtimeAssignmentLogContract.notKept, @serve.zone/interfaces 32.37.0). The answer is also the node's permission to discard for that session (@serve.zone/interfaces 32.38.0), and Pallet uses it while that session stays the current, live one: every flush and every hold on the delivery cadence discards each ended capture whole — its directory, its cursor and its waiting batch — and each capture that ends later, and has each running capture drop its output instead of holding it. A running capture drops every unread whole CRI line and a batch it sent that holds output; it records the dropped batch's number with its cursor, so that number is never sent again, not even by a restarted process, and its next batch opens with one not-kept loss counting the dropped bytes and whole lines (loss evidence the dropped batch carried leads it too). A sent batch that holds only loss evidence waits unchanged. The next session is asked again, and a controller that keeps logs receives each running capture from where it stands, after that loss. Any other failure of a batch keeps it the capture's next one and sends it again unchanged on the next pass.

Refused batches. Each capture is its own delivery unit. A batch the controller answers with a refusal stays its capture's next one and only that capture waits: the pass goes on with every other capture, the refused capture waits out its retry delay on the delivery cadence (the delivery interval doubled with each refusal in a row, at most 30 s, as every refused report does) and then sends the same batch again, unchanged; flush() sends it at once. Nothing is dropped for a refusal: the waiting capture stays within its buffer bound, so output it cannot hold is a counted buffer-overflow loss, and an ended one counts toward the ended-capture bound (Ended captures below), past which the oldest are discarded whole and counted. The refusal is written to stderr, which Spark forwards to the node journal, as Pallet log batch refused: assignment <assignmentId> capture <captureId>: <label>., once and then with its count while it repeats, and it stands in the pass's rejection until the capture delivers, ends or is discarded, or a new controller connection is asked afresh. The contract names no refusal of a batch (TRuntimeAssignmentLogsStatus is accepted, replay or not-kept), so Pallet cannot tell a refusal that is final from one that is not, and treats every refusal as this capture's to retry. A failure that is not an answered refusal — a timeout, a lost connection, a session this client no longer holds, an answer that does not bind to its batch — still ends the pass's log delivery (Cloudly log survey 2026-10-05: one refused capture ended every pass at itself and starved the node's other captures).

Ended captures. A capture whose container ended waits for its final batch to be acknowledged, and a crash-looping workload ends one per restart, so the node keeps at most runtimeAssignmentLogContract.maximumEndedCaptures (64) ended captures awaiting acknowledgement, holding at most maximumEndedCaptureBytes (64 MiB) together, each measured as the buffer bound measures it: in the file bytes it holds for delivery (Bounds below). Each pass holds every ended capture within its own buffer bound and rewrites its files to what it still needs before it is measured, so a capture at the largest buffer (64 MiB) always fits whole. Past either bound — a controller that is down, fails every batch or leaves the method unhandled — it discards the oldest ended captures whole, oldest by when this process saw them end. Every capture discarded whole, under this bound or a not-kept answer, is counted in the discardedCaptures of the next batch the node creates, of any capture. The count is persisted (pallet_log_discards): it survives a restart, travels unchanged with the batch that carries it on every retry, a restarted process creating that batch again with the same number and count, passes on to the next batch when that batch is discarded unacknowledged (dropped under not-kept, discarded with its capture, or lost to a restart that cannot continue its capture), and is kept while a not-kept answer stands. The owner keeps in memory only running captures and ended ones within the bound: a capture that finished or was discarded leaves no directory, cursor or entry behind.

Rollout order. A controller must answer reportRuntimeAssignmentLogs before a Pallet with log capture runs against it. A controller that leaves the method unhandled fails every batch: the node stays within its bounds, but it retries the same batch on every flush and drops the output beyond the buffer bound, so no workload output reaches such a controller. Cloudly answers not-kept (since 33.8.0). Onebox has no Pallet log backend yet (Onebox 33.1.1): run this Pallet against Onebox only once one ships.

A controller must also run @serve.zone/interfaces 32.38.0 or later before this Pallet runs against it. A receiver on 32.37.0 or earlier reads a batch by its exact schema and refuses one that carries discardedCaptures or a not-kept loss, on every retry: against it the node stays within its bounds as against any failing controller, but that capture's output no longer arrives. Cloudly 33.8.1 is the first Cloudly on 32.38.0 (33.8.0 is on 32.37.0); an Onebox Pallet log backend must be on 32.38.0 when it ships.

Bounds. What the controller has not acknowledged is the capture's buffer, at most the config's logs.maximumBufferedBytes. Pallet measures it in the file bytes the capture holds for delivery, which is what sits on the tmpfs: the file bytes the batch waiting for its acknowledgement spans, plus the file bytes not yet read into a batch, CRI headers and the cut rests of lines beyond logs.maximumLineBytes included. The waiting batch counts with its whole span whether or not a reclaiming pass (below) already removed those bytes from disk. A batch never spans more file bytes than the bound (its first CRI line is always taken), so the waiting batch is always kept and sent again unchanged; beyond the bound the oldest unread whole CRI lines are dropped and stated where they were dropped as a buffer-overflow loss, which counts their exact output bytes and whole lines and leads the next batch. Output of short or cut lines therefore reaches the bound sooner than its output bytes alone would: a line cut to a few bytes still counts every byte the runtime wrote for it. Drops while a batch waits extend that one loss, which keeps the time the loss began, so however long the controller stays away the evidence is a single entry. Once the node has read past the buffer bound (never below 1 MiB) it renames the file aside and has the runtime reopen the log path (CRI ReopenContainerLog through the native reopenExecutionLog, which proves the attempt's exact identity before and after); the renamed file is drained first and removed once the cursor moved past it. On disk an attempt therefore holds at most the read part below the rotation bound plus its unacknowledged buffer. The runtime reopens only a running container's log, and its answer ends a capture only on definitive evidence that nothing writes the file any more: the container exited, its sandbox stopped, or it was removed. A container created but not yet started, or one in an unknown state, may still write: the reopen is refused as indeterminate like a runtime that could not answer, so the renamed file gets its name back, the capture keeps running and no rewrite touches its files, and a later pass asks again.

A rotation survives a Pallet process that stops inside it. Between the rename and the runtime's reopen, the renamed file is still the one the runtime writes, and on disk that window is a renamed file (workload.log.1) without a fresh workload.log. Every pass that finds it so asks the runtime to reopen before it reads on, unless the container ended, and every step that rewrites or removes the renamed file completes the rotation first, so none ever replaces or unlinks a file the runtime still writes: output written there would otherwise be neither delivered nor counted, and would grow unseen on the tmpfs. When the runtime cannot answer a reopen, the renamed file gets its name back by a hard link, never a rename over the name, so a fresh file the runtime created meanwhile (a reopen whose answer was lost) is never replaced; a process that stops between the link and the removal of the renamed name leaves one file under both names, and the next pass removes the renamed name. A reclaiming rewrite that stopped before its rename leaves a copy (workload.log.1.compact) that the capture removes when it opens.

A pass that had to drop also reclaims the read part, which the waiting batch no longer needs on disk: it rotates the file at once and rewrites the renamed file from the first unread byte, so after that pass the attempt holds on disk only its unread bytes, within the bound, however long a batch waits. The bound holds whatever the controller does: every delivery pass keeps the captures it reaches within it, a pass whose batch fails holds every capture before it ends, and the controller client holds every capture on the delivery cadence (deliveryIntervalMs) whether or not a controller is connected, so a controller that is down, fails every batch or leaves the method unhandled never lets a capture grow. A hold that fails writes Pallet log capture hold failed: <class> <code>[:<site>]. to stderr, which Spark forwards to the node journal, once per distinct failure until a hold succeeds. The capture's delivery, hold and reconcile run one at a time. A capture's own failure — its files, or a runtime that cannot answer its reopen or cannot yet tell whether its container will write — is that capture's alone: the delivery pass or hold goes on with every other capture, the failed one is tried again on the next pass, and the pass reports the first such failure once it is done (a delivery pass rejects, a hold writes the line above). A controller's failure that is not a refusal of one batch still ends the pass's log delivery, after every capture is held (Refused batches above).

An ended capture is never written again, so it keeps on disk only what it still needs: the file bytes of its waiting batch and the file bytes not yet read, CRI headers and the cut rests of long lines included. Once its files hold more beyond those bytes than those bytes and 64 KiB (palletLogCaptureBounds.settleSlackBytes), it rewrites them into one file (workload.log.settled, renamed over workload.log), checked when it ends and after every acknowledgement. The cursor moves to the rewritten file in one transaction before the rename, so a restart at any point continues from exactly the recorded byte. An ended capture therefore holds on the tmpfs at most twice the file bytes it still needs plus 64 KiB, instead of every byte the running container left behind. The bound is amortized on purpose: rewriting after every acknowledgement would hold only the needed bytes, but would copy what is left once per batch, quadratic in the capture's size (about 2.3 GiB of copying to drain one 64 MiB capture in 900 KB batches). Rewriting only once the files are at least twice what is needed copies each byte a bounded number of times, never more than was acknowledged or dropped before, so draining a capture copies at most its own size. Since the ended captures measure at most maximumEndedCaptureBytes together in the file bytes they need, after every pass they hold on the tmpfs at most 2 × 64 MiB + 64 × 64 KiB = 132 MiB (138 412 032 bytes) together.

Durable cursor. Each capture keeps one record in the node's own database (pallet_log_cursors): the capture, its last recorded sequence, the position after it (the kernel identity — device, inode, birth time — of the file it points into, the byte offset and the line state of both streams), the evidence that leads the next batch, and the waiting batch. A batch is recorded when it is created, before it is sent: where its file bytes end, and every entry its creation generated with its timestamp — the losses it opens with and the empty entries a final batch closes open lines with — and the discard count it states. The cursor moves on only after an acknowledgement is recorded. A restarted Pallet process continues the same capture from exactly the recorded byte, mid-line included, and creates the waiting batch again byte for byte from those file bytes and the recorded entries, so a sequence number once sent is never sent with other content. Evidence that leads a batch holds one loss per reason, however many passes drop before it is sent. Only a cursor whose file is gone — a reboot emptied the tmpfs, or a reclaiming pass rewrote the file since the last acknowledgement — starts a new capture with a capture-restart loss of unknown extent. When the container stopped, failed or was removed, the capture drains the rest, closes a line left open, sends one final batch, and once that batch is acknowledged records the capture finished, then removes its directory and then its cursor; a restart between those steps completes the removal and never counts the delivered capture as discarded. A capture discarded whole first, under a not-kept answer or past the ended-capture bound, loses both as well. Its record goes first, so a restart before its directory is gone leaves a directory without a record; the attempt's assignment stays listed (the assignment store keeps every attempt's state as a tombstone), so the next process opens that directory as a new capture of the ended attempt, delivers what it holds and removes it. The first pass never removes a directory merely for lacking a record: a running capture has none until its first batch is created, and a directory whose attempt no state names proves nothing about its writer (a database restored from an older copy while the container runs), so such a directory stays until the reboot that empties the tmpfs. A batch dropped under not-kept records the cursor too, at the dropped number and past the dropped output, together with the evidence the next batch opens with, so a restarted process skips the same number and states the same loss. The first pass of a process forgets every cursor whose capture directory is gone (a reboot emptied the tmpfs) and passes on the count of every carrying batch no open capture creates again; a cursor recorded finished goes uncounted.

pnpm exec tstest test/test.logcapture.node.ts --verbose --logfile --timeout 60
pnpm exec tstest test/test.logowner.node.ts --verbose --logfile --timeout 120

Private workload readiness

PalletExecutionOwner binds the published runtime config policy to the exact admitted assignment before invoking a native process, TCP or HTTP readiness probe. Probes inspect the complete CRI identity before and after network IO and connect only to the evidenced container IP addresses, from inside the sandbox's own network namespace: the host has no route to a workload, by design. The execution owner reads that namespace from the run's attached network journal (the pair's sandbox namespace), only while the journal is live and names exactly the CRI sandbox the evidence shows; a running run without one is refused as unresolved-authority. An isolated run has no such namespace: its readiness must be the process's, and a TCP or HTTP policy for it is refused (unsupported-readiness). The native opens the journaled path once, requires a network namespace of exactly the journaled device and inode, and enters it on a dedicated thread that runs the probe on its own runtime and ends with it; another namespace at the path fails the sample namespace-mismatch, an unopenable one namespace-unavailable. An attached run's readiness also requires its workload pair raised, whatever its target: inside the sandbox namespace the native reads the pair's sandbox interface (named by the attachment receipt) and requires it up and running and the namespace's default route through it, before any target probe; otherwise the sample fails link-down. A listener on the workload's own address answers over a link that is down, so a probe alone never saw a lost attachment: Pallet 36.0.2 reported a workload ready for 51 minutes while its pair was DOWN after a crash restart. An isolated run's process target needs no namespace. HTTP uses the configured Host and request target, checks final response headers without reading a body, and never follows redirects. HTTPS verifies system trust and the configured hostname/SNI. There is no insecure TLS or DNS-target fallback.

SmartData stores probe evidence, exact container start nanoseconds, boot ID, threshold counters and the next due time. Linux CLOCK_BOOTTIME provides ordering across native children and daemon restarts. Initial delay and startup grace use a conservative anchor from the first completed native observation of that exact runtime identity. An already running container therefore waits the configured delay when first discovered; a forward wall-clock step cannot shorten it. Thresholds and intervals come from the immutable config. Changed runtime identity, boot, config or clock regression resets readiness; pending native effects supply no readiness proof. Probes never overlap. A reconcile before the next due time returns readiness-pending with no new observation.

A ready observation must match the durable sample's assignment, timestamp, image, sandbox and container. Its outbox transaction fences the current native lane. Restart and lost acknowledgements retain the exact existing observation; they cannot reuse a cached probe to create a new ready report. This remains a private composition. Authenticated controller transport and production route promotion are separate integration requirements.

Qualification uses the actual native probe, real HTTP/HTTPS listeners, a controlled CRI server and disposable NoSQLDB stores. The focused tests are test.readiness.node.ts, test.readinessstore.node.ts and test.executionowner.node.ts; they do not establish new real-containerd or host power-loss qualification.

Private packet policy composition

The backend-private composePalletPacketPolicies maps a validated network projection and complete current attachment/handoff receipts into a combined Smartnftables router policy and its host-transit policy. It captures inputs before asynchronous validation. Attached local workloads receive exact veth sources; selected remote workload grants require an explicit caller-owned TUN. Reserved but unattached local workloads receive no packet paths. Private DNS uses only each workload's gateway and explicit TCP/UDP port 53 rules.

Workload egress comes from the published projection grant helper. Router-origin DNS and platform traffic comes only from explicit signed router selections, with the current handoff's transit source and exact selected destination tuple. Public workload grants grant no router authority. Both policies carry the complete protected union and use only the current handoff's leased transport ranges. Withdrawals grant no traffic; DNS readiness does not change packet authority.

Published host ports come only from the verified projection's endpoints placed on this node. Every entry keeps its Cloudly authorization reference; hostIp is the node uplink address or absent; the workload must be attached with a current receipt; a port held by a platform endpoint or resolver on the uplink address, a port inside a leased SNAT range, or an entry beyond the bound refuses the endpoint's whole set by a bounded publication.* site and blocks nothing else. An entry may publish a contiguous range (hostPortEnd, each host port to the same port inside the workload) and may be symmetric: outbound flows the workload opens from its published port(s) leave from the same port(s) on both hops, so a SIP or RTP peer sees one address and port in both directions (@serve.zone/interfaces 32.31.0; pinned at 32.32.0, which also refuses a symmetric entry whose inside port another entry of its protocol shares). Every port of a range counts: a range that reaches a leased SNAT range or a platform endpoint's port anywhere is refused like a single port. The contract admits a symmetric entry only on an endpoint with public egress and refuses the whole projection otherwise, before Pallet composes anything. The bound is the contract's 64 entries per generation, a range counting as one, taken over the node's signed publication list in lease id order: an endpoint counts against it once the projection itself admits its set, whatever the node later refuses of it (uplink, attachment or compiled capacity), so the bound follows from the projection alone and binds the allocation-pool guard as well (see below). An endpoint that would pass it is refused publication.capExceeded. The accepted set is handed to the host-transit policy as hostIp:hostPort[-hostPortEnd] -> transitAddress:hostPort[-hostPortEnd], journaled with the ACTIVE generation and retired with it. Leased outbound translation keeps using the handoff lease's own source-port ranges; Cloudly allocates new leases in 49152-65535 (runtimeNetworkHandoffSnatPortRange), so publishable ports never meet them, while a lease allocated earlier keeps its range and still refuses any publication inside it. Pallet's session registration offers the installed interfaces release, so Cloudly sends range and symmetric entries only to a node that reads them. The native compiler decides what fits its atomic budget: while it refuses either policy as EXHAUSTED on a bound that publications spend (their count, or the rule bytes, target bytes or operations of the batch), Pallet refuses publishing endpoints by name (publication.capExceeded), last first, and recompiles both. A policy the compiler refuses as INVALID input, or as EXHAUSTED on any other bound, is a composition defect that no refused publication mends: the generation fails with PalletPacketPolicyRejectedError (invalid:packet.policyInvalid or exhausted:packet.capacityExhausted), which carries the engine's bounded reason and, for EXHAUSTED, the bound with its limit and actual count. The router policy carries the second hop of the same accepted set, transitAddress:hostPort[-transitPortEnd] -> workloadAddress:targetPort with the same symmetric flag; it is composed from the same verified entries, cross-checked against the journal and applied by the native compiler, and the raise owns the sandbox default route the workload answers through: CNI ADD installs default via <router address> (static, metric 100) in every pod namespace and advertises exactly that route in its result. Proven end to end in the DHCP-hub packet qualification: an external client reaches the workload over TCP and UDP through both translations, the workload sees the client's own address, the client sees the published uplink address back, a refused entry never opens, and a successor generation without publications leaves both ports dark.

The host hop translates a publication to the router's transit address, which lies in the transit pool the node's allocation-pool guard denies by current destination in every hook, and by current source on the reply: without an exception the guard drops every published flow in FORWARD, on a docker-shared and an exclusive host alike (lab 2026-10-06 RUN 9, D12). The guard therefore carries the node's publications as its exceptions (allocationPoolGuard.publishedPorts, @push.rocks/smartnftables 4.4.0). Both hops take their entries from one signed publication list, signedPublications in ts/network/packet/publication.schema.ts: every endpoint the verified projection places on this node that publishes ports, with the site the projection itself refuses its whole set at (publication.unauthorized, publication.snatOverlap). composePalletPublications applies what only the node knows on top of it for the host hop, and composePalletGuardPublications takes every entry of the list the projection does not refuse, at most the 64-entry bound and so far within the guard compiler's 1024, in the host hop's translated form (one transitTarget): to the handoff's router transit address with the host hop's port. The guard admits each only translated from outside every pool, with its replies; it checks no source, link or uplink address, so the host hop's translation, bound to the uplink, is what lets a flow in. For one projection the guard holds every entry the host hop translates, and beyond them only the entries of an endpoint refused for the node's uplink, attachment or capacity, to which nothing is translated (test.guardpublications asserts the equality). A new projection's publications reach the guard with its next transition, like its host grants. The docker-shared guest (test/native/dockershared.ts, run by qualify-active-registry.py --docker-node-binary … --docker-bundle …) composes one signed projection's publications for both the host hop and the guard and proves, on the docker-shared host under Docker's FORWARD DROP with the contribution and on the exclusive host: TCP, UDP and both ends of a UDP range are delivered with no guard (docker-shared), dark under the guard as Pallet 36.2.0 composed it, and delivered once the guard's transition adds them, the workload seeing the external client's own address and the client the uplink address.

Host grants let the node's own host namespace — Onebox's reverse proxy on a single host — dial an exact workload port directly, with no loopback publication, no route_localnet, no loopback DNAT and no loosened martian filtering. They come only from the contract's getRuntimeNetworkProjectionHostPacketGrants over the verified projection, so only a onebox projection's signed host selections grant any; Pallet proves each again against that projection — the source is the current handoff's transit host address, the destination the exact lease address and port of an endpoint on this node — and refuses the composition otherwise. The same exact tuples are compiled into all three tables the flow crosses: the host-transit and router policies of the ACTIVE generation, for attached workloads only, and the allocation-pool guard, where they are its only exceptions. A new projection composes a new generation and guard target, so a withdrawn selection is an ordinary atomic replacement; an absent or empty set leaves every policy byte-identical.

A granted flow needs the host to route the lease address through the selected handoff, and the ACTIVE generation owns that route. A generation whose projection selects host flows plans one host route per workload pool that holds a selected lease (hostRoutePrefixes, derived again from the journaled projection on every read): the pool via the router's transit address, sourced from the host's transit address, so the flow leaves with exactly the source its grant names. It is added through a command socket in the host namespace inside the registered host transition that raises the link, deleted inside the one that lowers it, checked in every inventory of the raised link, and removed by recovery when it outlives a fence. Any other source's change to a route touching the host link (through it, or over its transit prefix or pool routes) still retires the uplink observation, so it is never accepted as the generation's own. The route widens nothing: the guard and both hops still admit only the exact granted tuples, and a generation without host selections — every Cloudly generation — plans none and is byte-identical.

Workload ingress grants let the cluster ingress workload reach the port a target workload listens on directly, instead of a hairpin through a published uplink port. They come only from the contract's getRuntimeNetworkProjectionWorkloadIngressPacketGrants over the verified projection, so only a cloudly projection's signed workloadIngress selections grant any; Pallet proves each again against that projection — two different endpoints, at least one placed on this node, the exact lease address of each and the target's listening port — and refuses the composition otherwise. They compile into the router policy's workloadGrants (@push.rocks/smartnftables 4.0.0): a stateful one-way flow of the exact protocol and port, answered only by the replies of a connection the ingress workload opened, never translated, and never open in the other direction. A grant compiles only when both of its workloads are attached behind this router in the generation; an unattached end, or a target on another node, receives no packet authority here, because the compiler has no one-way grant across the tunnel. Large port ranges such as RTP media stay uplink publications. An absent or empty set leaves the router policy byte-identical.

On an exclusive host (see Host forwarding modes) every host-transit policy a new generation composes carries exclusiveForwarding (@push.rocks/smartnftables 4.1.0): the host's IPv4 forwarding is on only while that generation is UP, and the table's forward chain ends in one unconditional drop after the handoff, publication and symmetric flows, so a LAN peer, a hairpin through the uplink or traffic between two other links is never forwarded. On a docker-shared host the policy is shared, and Docker's FORWARD policy DROP with the node's DOCKER-USER contribution does that work. The member is the composition's explicit exclusiveForwarding binding; a generation is recomposed with the value its journaled host policy carries, so one an earlier release journaled without it is recomposed byte for byte, fenced and released, never raised again.

A projection's hostPlatformEndpointIds names the protected platform endpoints this node's own host serves — on a single host the cluster hub, the relay listener and Corestore. They compile into the host-transit policy's localPlatformEndpoints: the host delivers leased flows to them in its INPUT instead of forwarding them, from the exact handoff, with the lease's transit source address and a source port of its range, and only the replies go back. The workload and router selections that reach them are unchanged: a workload still reaches only the endpoints it selects and the router only the hub its block selects. Every address a declared endpoint holds is added to the uplink binding, so the compiler verifies at apply, recovery and inspection that the host holds it; each must be a permanent /32 of the uplink, which the uplink observation admits beside the DHCP lease (see Retained uplink observation). An undeclared platform endpoint or a resolver on one of those addresses refuses the composition as conflict. Without the member the uplink binding and the host policy are byte-identical.

This function performs no native apply, link activation or durable journal write. Its caller must authenticate the projection, retain and fence actual namespace, attachment, TUN and uplink generations, and own routes, DHCP changes and SNAT address lifetime. Native preparation must confirm graph capacity before a complete transition is journalled and applied. Composition alone provides no enforcement, workload readiness, allocation reuse or packet-drain evidence.

Private DNS lease arithmetic

The backend-private ts/network/lease.ts converts an authenticated lease of at most 15 minutes to a conservative native CLOCK_BOOTTIME deadline. Its input is a verified UTC interval anchored to the same boot clock, with an independently qualified rate-error bound. It accounts for uncertainty and the fastest permitted UTC progression. Process recovery retains the exact deadline; persisted boot-time and UTC lower bounds reject clock regression.

A clock this node cannot read lapses that DNS acquisition and never fails the lifetime that owns the clock, and which lapse it is follows from what the clock states. A sample it cannot qualify leaves the acquisition without fresh time evidence: it continues on the window already proved, keeping the exact deadline this boot converted, and lapses only where there is no such deadline — another boot, or a window this boot has not converted. A clock that states no boot reading at all — unqualified, or lost with its native owner — is a clock-unavailable acquisition outright: that pass withdraws every name and the next pass binds them again. No reader is refused for asking while another holds the clock; losing the owned process still fails the node through the clock's own failure signal.

Reboot recovery requires newly verified independent time evidence for that boot. The original absolute expiry stays unchanged, and a saved Cloudly timestamp cannot renew it. The returned anchor, deadline and high-water state must be persisted through SmartData inside the complete authenticated projection before applying a newer native DNS revision. Expiry does not revoke packet grants or prove that an old writer is absent.

The private PalletIndependentClock owns an authenticated NTS measurement process through the native executor. Its pinned Chrony 4.8-servezone2 variant returns exact current Unix time and tracking metadata from one observation. Legacy Chrony tracking offsets cannot supply this precision when the RTC is years wrong. The private process never adjusts the host clock or loads host Chrony/GnuTLS settings. It uses bundled trust certificates, fixed Ubuntu NTS peers, no drift/cookie files, and a root-private Unix socket. One native operation runs at a time, and no caller is refused for asking second: a caller that asks for a reading already in flight is answered with that reading — a reading is of an instant, so sharing it is exact and the owner is asked once — and a caller that asks for the other reading waits for the operation in flight and then takes its turn, in arrival order. Loss of that owned process fails the node lifetime; unqualified or offline samples supply no time authority.

First qualification must not inherit long outage backoff from attempts made before startup connectivity is ready. The owned Chrony variant caps unresolved-source retry at 28 seconds until the first source resolves, and both connection- and TLS-class NTS-KE retry at 16 seconds until the instance first receives authenticated exchange data. Upstream exponential backoff resumes after those successes, including after cookies are later exhausted or the instance resets. Sampling continues to require every qualification check below; an unqualified reading grants no lease extension, and an expired native or VPN lease is never revived. The next network pass retries qualification and renewal without changing the live lease.

The owner checks coherent tracking, selection and NTS authentication reports, root error, freshness and kernel clock brackets. Supported KVM Linux clock profiles bound BOOTTIME progression by 400,000 ppm; unsupported clock sources, PPS/custom tick settings, VM pause, snapshot restore and live migration do not provide a qualified clock. The bootstrap certificate's validity interval is 1970–2100. A wall-clock reading or an NTP-synchronised flag alone is insufficient.

Clock ownership does not persist projections, supervise DNS or admit workloads. The DNS persistence and lifecycle owner remains unfinished. Lease arithmetic checks cover uncertainty, deadline boundaries, process recovery, reboot proof requirements, rollback and malformed stored state. The isolated native clock harness additionally checks actual NTS measurement, wrong-year RTCs, offline denial and joined Node/Deno cleanup; only recorded successful runs qualify the specified source and kernel.

Private network reservations

The backend-private ts/network/allocation/ store retains protected-authority revisions and immutable handoff leases through the identity runtime's existing SmartData/NoSQLDB connection. runNetworks() exposes callback-scoped operations; shutdown joins admitted operations before closing the engine. Every write, including replay, requires an owning-code identity fence inside its transaction. The caller authenticates controller authority before entering this interface.

A fresh node can stage the complete current authority. An initialized node must advance through exact consecutive references in the same controller epoch and node scope. Reservations reference the current authority and validate against all retained leases; a common revision fence serializes competing allocations. Historical leases continue to bind their exact authority. Quarantine preserves the lease and forbids reuse of its source-port range, conntrack zone and label within the shared contract's respective scopes. Replay retains quarantine. listHandoffs() returns the complete bounded retained history; there is no expiry, pruning, release or reuse operation.

This stores allocation intent. It does not authenticate signed projections, activate pools, acknowledge host/router barriers, declare flow drainage, allocate sandbox IPs or admit packet traffic. Those owners must bind this durable history to exact native receipts before activation. Tests use actual NoSQLDB to check concurrent conflicts, transactional identity fencing, stored digest corruption, lost commit acknowledgements, complete restart recovery and joined shutdown.

Private signed network admission

PalletControllerClient receives the published Interfaces 30.9.0 signing-authority and complete network-projection RPCs on its private outbound TLS connection. The current physical session, active credential and controller/node/namespace scope fence each admission transaction. Socket input is inert and bounded; the transport accommodates the full 896 KiB projection contract.

ts/network/projection/ persists public signing trust, immutable public revision history, one complete signed projection and bounded immutable workload leases through SmartData on the existing NoSQLDB owner. Key rotation and projection admission write a common revision fence. Rotation or revocation prevents the old envelope from being read as current authority; historical keys retain verification of recovery material. Re-signing an identical projection under the new key replaces its envelope before acknowledging replay. Replays never renew the DNS deadline.

Projection admission retains historical protected authorities and handoffs in the same allocation transaction. Full local and remote workload leases survive withdrawal and restart, including the addresses needed for later denial. Retired leases stay quarantined; their subnets and execution attempts cannot be reused. The store retains at most 512 workload leases and rejects capacity exhaustion. Admission carries pending withdrawals and tombstones until the node's own durable application receipt names exactly the previous projection (Projection application receipts); then it admits a successor that drops them. An ACK establishes durable intent only.

The callback-scoped runNetworkProjections() facade expires with its callback; shutdown joins admitted work. inspectRetainedProjection() is recovery material, while readCurrentProjection() requires the current signing key. Neither method establishes an application, time, or current-identity fence for a caller. Pool allocation, CNI, native realization receipts, network activation, qualified DNS time and live worker rollout remain separate owners.

The private application-store composition reads current verified signing intent and the complete retained handoff/workload history in one SmartData transaction. Before recording native intent, it rechecks that source fingerprint and writes both common trust and allocation fences in the caller's transaction. Concurrent key rotation, projection admission, reservation or quarantine therefore conflicts with stale intent; a failed identity guard rolls back both writes. This interface does not issue a native receipt or expose an application capability over RPC.

Tests test.networkapplicationsource.node.ts, test.networkprojectionstore.node.ts and test.controllernetwork.node.ts exercise actual file-backed NoSQLDB and TLS sockets, including concurrent rotation, identity rollback, lost ACKs, physical reconnect, shutdown and a complete namespace with 12,288 DNS names. They establish admission behavior, not live packet policy.

Private DNS lease renewal

A renewal carries a fresh DNS window for one exact admitted projection and no applied network state. PalletControllerClient receives applyRuntimeNetworkDnsLeaseRenewal on the same outbound connection as the two admissions, and the envelope's session must be the binding this node obtained itself. That comparison is what makes a cluster relay's push admissible: the relay holds signed statements for a node it carries and adds no trust of its own.

Admission is one transaction in ts/network/projection/, in this order: the stored signing trust, this node's current projection, the signature under that trust and under the revision that signed the projection, the renewal this node already holds for that projection, the published admission decision against the qualified clock's lower bound, the write, and the common trust fence every network admission takes. accepted replaces the single stored row, replay changes no renewal and repeats the same bound acknowledgement, and every other outcome is the refusal the other two admissions give; the trust fence is taken on both outcomes, so a concurrent rotation conflicts rather than interleaves. What persists is the complete signed envelope, so every later read re-proves it against the retained signer rather than trusting the node's own table. An envelope this node can no longer prove is a lapsed renewal, not a broken node: the DNS lane discards it and converts the projection's own window, while a credential bind refuses. The persisted DNS lease and DNS intent each state the renewal their window was converted from; no released version persisted either shape, so there is no migration, and a store carried over from an older build fails its assertion on every read and must be started from clean pallet_dns_lease and pallet_dns_intents collections.

A renewal belongs to one projection envelope: that projection's exact reference and the revision that signed it. The transaction that admits any envelope the stored renewal does not belong to discards it — a successor projection, and equally the same body re-signed after a key rotation, which keeps the reference and changes the signer. A renewal of a projection this node does not hold, at a sequence it has already passed, with different bytes at a sequence it holds, or whose slot has not provably opened at the qualified reading is refused; the sender simply delivers it again once it has. A rotated signing key renews nothing: the re-signed projection discards the renewal held and the node lives on the window that projection itself signs. While that window still runs nothing else changes; once it has ended the node withdraws DNS and stays dark until the controller delivers a renewal signed by the new revision. Pallet asks for none — a renewal arrives on the connection this node already holds — and converts the one it admits on its next DNS pass, which runs once a second.

Everything window-bounded then follows the effective window — the admitted renewal's when one renews this projection, the projection's own otherwise. The DNS lease converts that window in the same transaction that writes the lease and records the renewal it converted, the composed DNS snapshot proves its validUntilBoottimeMs against it, and the managed-VPN credential may expire inside it and never beyond it. Nothing else moves: the application, its fingerprint, the native generation and the protection receipt are untouched, and a node that loses qualified time or current signing authority still withdraws DNS exactly as before.

test.networkdnsleaserenewal.node.ts, test.dnsrenewalview.node.ts, test.dnsleasestore.node.ts, test.independentclock.node.ts and test.controllertunnel.node.ts state admission, supersession, replay, each refusal, the boundary at an equal expiry, the moved deadline, restart recovery, a stored renewal this node can no longer prove on either lane, a reading two overlapping callers share, a clock this node cannot read, and the bound the credential keeps.

Private identity store

The separate backend-private ts/identity/ implementation owns Pallet-only credential state in an explicitly supplied SmartData database. It is not imported by the public probe facade and has no Cloudly connection. Its protected lifecycle and node-local enrollment control owners remain private and unwired from the production daemon.

Preparation commits a 32-byte cryptographic palletToken before returning its SHA-256 proposal. Binding checks this owner's durable hash and the complete enrollment digest. Activation requires the exact bound acknowledgement and writes its immutable receipt in the same transaction. Initial adoption may name an existing Cloudly node while Pallet is still generation zero. Rotations retain the old active bearer until acknowledgement; historical receipts cannot reactivate it or replace pending material. Inspection, proposals and receipts contain no bearer.

The enrollment-specific helpers use published Interfaces 28.2 snapshots, digest and acknowledgement binding. Those helpers prove content binding, not remote authentication: the eventual authenticated transport/coordinator owns that trust boundary. Spark credentials never enter Pallet's private store.

Tests use disposable file-backed NoSQLDB 10.5.1 instances and prove exact replay across engine restarts, concurrent preparation, acknowledgement rollback, independent database bindings and rejection of changed or malformed input:

pnpm exec tstest test/test.identitystore.node.ts --verbose --logfile --timeout 60

Private identity runtime

ts/identity/classes.identityruntime.ts owns a Linux file-backed NoSQLDB 10.5.1 engine and its SmartData connection. Production defaults to UID 0 and /var/lib/serve.zone/pallet; trusted owning code supplies the installed engine's absolute path and SHA-256. It rejects symlinks, unsafe ancestors, wrong ownership, engine hardlinks, nonexecutable or writable engine files, and non-0700 data roots. It never repairs permissions or discovers an alternative executable. Ordinary startup requires existing storage; only explicit first provisioning may create it. The installer must exclude concurrent changes to the selected engine and paths.

NoSQLDB's fileStorageStartup: 'current-format-only' rejects legacy/mixed storage and orphan migration staging without modifying that tree. The native file-root lease is the only cross-process database lock. Model preparation happens after native ownership and readiness, with no duplicated format detector or lock helper. Each attempt uses its own mode-0700 /tmp/pallet-identity-* directory for the Unix socket; it contains no persisted application data. Normal cleanup removes only that owned socket and empty directory. Parent death can leave an empty disposable directory for OS temporary-file cleanup; startup never scans or deletes others.

run() admits a guarded identity facade, not the database or store object. Its methods expire when the callback settles, and admitted calls are tracked even if the callback does not await them. Stop immediately blocks new callbacks, drains admitted work, closes SmartData, and then confirms native exit. A failed database close retains the running engine's lease until cleanup is retried. Start/stop deadlines bound caller waiting, not resource ownership; pending or failed cleanup blocks restart. Native exit invalidates readiness and requires stop before restart.

pnpm exec tstest test/test.identityruntime.node.ts --verbose --logfile --timeout 60

The Linux tests qualify competing Node owners, successor receipt recovery, startup/stop races, cleanup failures, exact pending recovery after native SIGKILL, and active receipt recovery after an idle Node parent's SIGKILL. These are not hardware power-loss or ARM64 hardware tests, and do not prove bounded native exit when a management command is stuck. Production process supervision and coherent installer/engine packaging remain integration gates. When packaging that engine, retain NoSQLDB's own license and complete third-party notices alongside the binary.

Private enrollment control

ts/control/classes.enrollmentcontrol.ts owns its identity runtime and a root-only Unix listener at /run/serve.zone/pallet/control.sock. It acquires the native database lease before touching that socket. The installer provisions the trusted platform parent; the control owner creates its mode-0700 IPC leaf when absent, creates a mode-0600 socket and leaves the directory in place after shutdown. Existing directories are checked, never repaired. Paths must be canonical, symlink-free, protected from untrusted writers and fit Linux's socket path limit.

Only the five published requests.pallet enrollment and runtime-binding methods are registered. Prepare returns a current, independently retained hash-only proposal; bind proves the full pending enrollment; activate checks the complete authenticated Cloudly acknowledgement and then proves the exact current local identity. Historical receipts cannot masquerade as current activation. Spark must authenticate Cloudly before forwarding its acknowledgement or runtime routing binding over this trusted local channel. bindPalletNodeRuntimeBinding accepts the first exact routing tuple or an exact replay only after its origin and node match the current active Pallet identity. readPalletNodeRuntimeBinding requires that same current identity and rejects an absent binding. The binding is a separate singleton: absence is the unbound state, and the released identity document's relayOrigin field remains readable but inert rather than being expanded into unauthenticated routing data. The IPC does not expose a bearer, database, generic rotation or workload execution method.

TypedRequest 8.0.3 owns routing and envelope identity. Hooks and incoming response routing are disabled; wire-supplied local authority and unknown methods are denied. Each connection carries one four-byte big-endian length-prefixed JSON frame per direction followed by write-half-close. The server explicitly keeps the response half open until its reply is flushed. Requests are limited to 32768 bytes and responses to 16384, with a default 10-second total connection deadline and 16 admitted operations. Invalid UTF-8, incomplete frames, trailing bytes and missing EOF are rejected. No TCP listener or automatic application retry is introduced.

Disconnects and deadlines close transport but do not abandon an admitted database operation or free its admission slot. Shutdown stops admission, cancels socket I/O, drains handlers, closes the owned listener and then stops the identity runtime. Caller deadlines retain the actual cleanup owner and block restart. A live pre-existing socket is never removed; stale recovery requires the native lease, a completed connection-refused probe, and unchanged owner/inode checks. Because Node/libuv unlinks on listener close, shutdown checks its recorded socket and parent before calling close. A replaced name retains failed-cleanup ownership until the installer/operator restores the original owned path. These checks assume the trusted installer excludes concurrent root-owned path changes; they do not claim protection from hostile root. No arbitrary file or recursive cleanup occurs.

pnpm exec tstest test/test.enrollmentcontrol.node.ts --verbose --logfile --timeout 60

Qualification uses real Unix sockets and disposable native NoSQLDB stores for half-close delivery, independent restart, competing ownership, stale/live/replaced paths, malformed frames, hook isolation, cancellation, admission limits, historical receipt rejection and lost activation results. Spark's concrete client, coordinated daemon integration and production packaging remain separate integration gates. Before distributing a bundled control process, include its JavaScript dependency MIT/Apache-2.0 notices in addition to the existing Rust and NoSQLDB notice material.

Enrollment process lifecycle

The private PalletEnrollmentProcess owns one foreground control lifetime. Readiness follows the native database lease, model preparation and protected Unix listener. Cancellation during startup prevents readiness. Parent stdin EOF, SIGTERM and SIGINT stop admission and drain the existing control owner. A 250 ms watch of the public readiness getter terminates a failed lifetime; it never restarts the engine or retries a database operation. Each control request already checks native readiness independently of that watch.

The enrollment CLI accepts exactly enrollment-serve or enrollment-provision. Normal startup requires existing state; only explicit provisioning may create it. Trusted build code supplies the exact engine, digest and protected paths. No runtime path, UID, digest, bearer or configuration arrives through argv, stdin or environment. Stdin carries only parent lifetime: content is rejected without an echo. PALLET_CONTROL lines use pallet.enrollment.process and contain only ready, stopped, or a static failure reason. stopped follows confirmed cleanup; a failure does not authorize a successor until the supervisor confirms process exit.

pnpm exec tstest test/test.enrollmentprocess.node.ts --verbose --logfile --timeout 60

The lifecycle and source-process tests cover cancellation, EOF/signals, missing state, invalid input, native failure and retained cleanup ownership. A standalone Deno (the pinned 2.9.7) Linux amd64 fixture also qualifies real Unix half-close responses, durable replay, engine and parent SIGKILL, and stale-socket recovery from an unrelated working directory with an empty environment. That fixture supplies trusted test paths and UID. The production artifact, complete notice bundle, installer/supervisor integration and target-platform qualification remain gates; this source adapter is not a production node installation.

Foreground node process

The same control executable accepts runtime-serve for one private PalletNodeProcess lifetime. It reopens existing enrolled state and verifies the protected sibling pallet-runtime executor, pallet-guard, pallet-dns and the managed VPN pallet-vpn against their compiled-in SHA-256 identities before starting local owners. The released build activates the node network: it passes activation with pallet-vpn and its digest, so a node bound to its cluster relay raises its ACTIVE generation and managed VPN tunnel (see "Private ACTIVE tunnel transitions"), and a node bound to a local controller forwards as a single host. It never initializes missing storage or accepts paths, credentials or settings from argv, stdin or environment. The containerd CRI socket is /run/pallet/containerd/containerd.sock: Pallet's own containerd instance, never the host's shared /run/containerd/containerd.sock, which on a Docker CE host is Docker's containerd.io daemon with CRI disabled. The native runtime refuses that shared socket and Docker's sockets by path and by file identity: a socket with the device and inode of the shared containerd socket, Docker's API socket or Docker's embedded containerd socket is refused, so a symlink, hard link or bind mount of one of them cannot pass under another name. Pallet and Docker can therefore run side by side on one host without sharing images, sandboxes or CRI configuration.

The same lifetime runs in two modes that differ only in what ends it, and both take one argument, required: --forwarding-mode <docker-shared|exclusive>, who owns the host's IPv4 forwarding (see Host forwarding modes); anything else is refused as invalid_arguments. runtime-serve is a supervised child: its parent holds its standard input, and EOF there stops it (Spark's runnode). runtime-unit is the main process of the node service unit (see "Service unit contract"): its standard input is null, EOF is not a stop request, and SIGTERM or SIGINT stops it. Both refuse data on standard input as invalid_stdin.

A start fences what the previous lifetime retained before it reports ready, so a start after a crash does real work first. A supervisor must allow a start at least 600 s before it gives up on ready, in either mode (palletNodeStartTimeoutSeconds in @serve.zone/pallet-bundle). Pallet bounds each native request of a start, not the start as a whole: a request that misses its deadline ends the start with failed. The longest path whose length does not grow with the node's workloads is a start after a crash, on a host whose guard unit already ran in this boot, with an ACTIVE generation retained and a guard expansion pending. Its request deadlines add up to 478 s:

  • identity database 10 s (its startup deadline, migrations included), router namespace 18 s (6 s spawn including the 5 s descriptor-store barrier, two 6 s requests), independent clock 11 s (3 s spawn, 8 s start);
  • guard pass 274 s at smartnftables' 30 s per request: owner open 33 s (with its 3 s spawn), replay of the committed transition 90 s (prepare, reconcile, inspect), the pending expansion 120 s (prepare, then its replay), detach 31 s;
  • handoff set open 11 s (namespace read 3 s, spawn 3 s, open 5 s);
  • fence of the ended ACTIVE generation 131 s: the native's epoch proof 5 s, the retained host packet table's owner 36 s (namespace read, spawn, open), its recompilation 30 s and its release 60 s (release, inspect);
  • application recovery 15 s (epoch proof, prepare and reconcile at 5 s each);
  • execution host identity 8 s (3 s spawn, 4 s request, 1 s termination grace of a failed child).

Each retained workload attachment of the ended epoch adds its 5 s epoch proof (at most 512, twice the projection's 256 endpoints). The DNS pass, the database transactions, the executable digests and the storage recovery have no deadline of their own. 600 s covers the fixed path and leaves 131 s for those; a supervisor of a node that retains many workloads allows more, and PalletControlLifetime admits up to an hour (469 s + 512 × 5 s = 3029 s). Under the node unit (Type=exec) the service manager does not wait for ready, so TimeoutStartSec= does not bound a start; an installer that waits for the unit's ready line allows the same limit.

Runtime lifecycle lines keep the PALLET_CONTROL prefix with the distinct pallet.node.process protocol. ready means local identity, namespace and execution owners are ready; it does not assert controller connectivity, workload readiness or DNS availability. A start the packet engine refuses because of the running kernel ends with failed reason unsupported_kernel instead of owner_failed (see "Host kernel requirements"). A node that fails because its coordinator refused the stated forwarding mode for a fact of the host that persists until an operator changes it ends with failed reason forwarding_mode_refused, before or after ready, its refusal named on stderr: under exclusive persisted forwarding (forwarding.exclusive.persisted) or a running Docker (forwarding.exclusive.docker_present); under docker-shared no iptables-nft frontend (forwarding.docker_shared.iptables_nft) or a host that does not forward when the uplink opens (uplink.forwarding.open). Spark ends the node on it rather than restart it into the same refusal; everything transient stays owner_failed, and the guard modes never send it. Initial Cloudly connectivity runs independently. The runtime session requires the immutable node runtime routing binding and registers at its cluster relay. An absent binding fails startup; there is no direct-to-Cloudly fallback or second configured relay authority. Every registration response must match the persisted node, cluster, Cloudly controller and runtime namespace before it becomes execution authority, and reconnect rechecks the current identity and the exact persisted binding. The persisted tuple grants routing only; phase and workload authority always come from the fresh authenticated runtime session. Enrollment, the node credential and the registry host stay with Cloudly whichever transport carries the session: workload.registryHost is the publication identity the registry credential is fenced to, and it is never rewritten. Where the image bytes are fetched is a separate statement, workload.pullEndpoint — the relay's own origin for a cluster whose relay forwards the registry, so the node pulls inside its cluster and nothing cluster-side dials the control plane for bytes. The reference stays digest-pinned either way, so the endpoint decides reachability and never content, and the credential this node asks Cloudly for is addressed to whoever serves them. The contract this node relies on from the relay: it is expected to forward the node bearer verbatim and to serve the unchanged typed request contracts, so nothing in the session binding changes for Pallet. New workloads still require the exact authenticated session and registry grant, while persisted node-bound effects retain their existing offline rules. Network and secret references resolve through their owning resolvers on this same session before any native effect; storage references resolve to the local storage claims this node admitted on the same session ("Local storage claims").

Every inbound request a node refuses answers its sender the one refusal message of its kind, so a sender learns that it was refused and never why. The node writes why to stderr, which Spark forwards to the node journal: Pallet controller request refused: <method>: <chain>., the owner_failed labelling of class names, codes and sites: a new line at once, and a repeated one again with its count (<line> (repeated N times since <ISO time>)) one minute after it was last written, then at doubling intervals up to ten minutes. A network pass that fails is written the same way, as Pallet network pass failed: <chain>., until a pass completes: a failing pass renews nothing, so the node's tunnel lease runs out behind it. An assignment scan whose page the store refuses to read is written the same way, as Pallet assignment scan failed: <chain>., until a page is read: the scan keeps its cursor and reads the page again on its next interval, while the workloads, the network passes and the controller session keep running. The node fails only when an owner is no longer live, a database server that stopped included; Pallet 36.3.6 and earlier failed it on any such refusal, and the restart that followed took every site down (Grasberg 2026-10-10, one MongoServerError on this read). A database server's error is named by the server's name for its code and the number (MongoServerError:IllegalOperation(20)), never by its message, which can carry a document's key value or a path. The owner_failed labelling is bounded to 448 characters; a longer chain is cut in the middle (...), keeping its outermost owners and its last 192 characters, which name the innermost cause — what actually refused.

A controller may push on a session as soon as it answered its registration, but the node takes the session up only after it read that answer and its active credential again (on the controller socket, once the socket reports itself connected). From the registration on, until then, an inbound request on that connection waits for the session, at most palletSessionAdoptionWaitMs (5 s), and then acts under it or is refused as any request without one; one that outlasts the wait is refused with the site controller.session.adoptionPending.

Protocol handshake

The contract this node speaks is the installed @serve.zone/interfaces release, and it is stated once: ts/controller/protocol.ts builds one offer for the one session kind Pallet is (palletRuntime) from protocol.createProtocolOffer, with this build's own minimum (32.0.0, raised only by the commit that starts depending on a later minor). Every registration carries that offer, and the answer is the wrapped { session, protocol } the controller returns, negotiated against the offer that was sent before the binding is bound.

Two peers that cannot serve one session are named rather than collapsed into a transport failure. A controller that refuses this build answers an IProtocolRefusal; an accepting controller whose own offer this node cannot serve is refused by this node. Either way the node reports the status protocol-incompatible and retains the refusal in getStatus(). It is not an owned failure: the workloads keep running, the network stays up, the packet policies stay applied and the node's owners are not torn down — and the node process is not ended either, because protocol-incompatible is a live node with no controller it can speak to.

Recovery needs no operator action on the node. The refused offer is made again every protocol.refusedOfferRetryIntervalMs (five minutes, the one interval @serve.zone/interfaces states so that no client invents a second one) until a controller accepts it, through @api.global/typedsocket's TypedSocketRestoreDeferral: the refused attempt is closed, the transport waits, and it offers again without consuming one of its reconnect retries, so a node refused for days keeps them for real connection failures. An accepted offer drops the retained refusal, returns the client and the node to ready and resumes the pass, so upgrading the controller under a running fleet brings the fleet back by itself. The node's own poll loop keeps its interval while refused — the pass stays empty and admits nothing — because that loop is what carries the accepted offer back into the node's state.

A restoration the transport refuses outright is the one terminal outcome. TypedSocket releases itself, because a retry could only repeat the peer's verdict, and that leaves the node with no controller at all — the condition an exhausted transport leaves it in — so the node fails owned: the process ends and the supervisor's restart is the next offer, which lands on the same five-minute cadence if the controller it meets has still not moved. The denial is named on the controller status as restoreDenial — the transport's own reason and, when the refusal was this client's, the cause behind it — beside any refusal the node was already standing on, so the status the process ends on says which verdict ended it.

The registration also states resolvedAuthorities, the sorted, duplicate-free list of workload authorities this build can resolve. It is the same constant the execution owner enforces (resolvableWorkloadAuthorities), so what the node promises and what it refuses can never drift: an assignment whose requiredWorkloadAuthorities the list does not cover is refused as unresolved-authority. This build resolves network, secrets and storage.

Local controller

A node can be driven by a controller on its own host — Onebox — instead of Cloudly (nodeRuntimeLocalControllerContract in @serve.zone/interfaces). A node carries a Cloudly identity or a local controller, never both: a first local bind is refused while any Cloudly enrollment state exists, and a Cloudly enrollment is refused once a local controller is bound. Both claim one shared pallet_node_authority singleton in the transaction that writes their identity, and the first claim inserts it, so of a concurrent Cloudly preparation and first local bind exactly one commits. The other's commit conflicts with it, the transaction is retried, and the retry finds the committed authority and refuses as conflict. An offline read that finds both kinds anyway refuses instead of choosing one. The persisted identity selects the transport when the controller client starts; there is no configured choice between them.

The local controller is authenticated by the socket's permissions and nothing else: the socket is owned by root with mode 0600 in a root-only directory, so any process running as root on the host is trusted as the controller. The bearer Pallet generates and the controller pins at its first bind detect a Pallet that lost its state (a later bind answering another credential is refused by the controller); they do not keep out an intruder with root.

Pallet serves the local controller on the root-only Unix socket /run/serve.zone/pallet/controller.sock through SmartServe's unixSocket listener, which applies mode 0600 before any peer can connect and refuses a path another process serves. The controller is the TypedSocket client (http://localhost over the socket), but the session keeps its directions. Its first request on a connection is bindPalletLocalRuntimeController: the first bind persists the binding and a bearer this node generates, and every later bind must be the identical binding and answers the same credential generation and SHA-256 hash. The bearer itself never leaves this node except in its own registration. After answering, the node fires the unchanged registerPalletRuntimeSession at that exact connection with that bearer, and it binds the answered session to its own persisted binding (bindRegisterPalletLocalRuntimeSessionResponse) before it becomes execution authority. Admission, reports, registry and secret requests then run as on the outbound transport, but only on the connection that holds the session: another connection on the socket, bound or not, never acts under it. A connection whose registration is accepted supersedes the previous session and closes its connection; a failed registration closes only its own. A controller that refuses this build's offer leaves the node protocol-incompatible until a later connection's registration is accepted.

A local controller has no Cloudly origin to stand for its registry. It issues a pull credential only for an image whose registryHost, and pullEndpoint when one is stated, are both listed in the binding's registryHosts (isNodeRuntimeLocalRegistryWorkload); any other image is refused before a credential is requested.

A node bound to a local controller always serves it from its node lifetime (runtime-unit under its service unit, or runtime-serve). Only network-acquire-local admits the first bind of an unbound node (see Initial projection acquisition). Networked workloads of a local controller need the attach barrier like every other workload; on a node without a managed VPN it holds as a single host (see Private ACTIVE tunnel transitions).

Stdin EOF, SIGTERM and SIGINT close admission and join startup, controller and native operations. The CNI broker remains available through CRI rollback, then joins its admitted handlers before the network and database owners close. The private clock stops and joins its Chrony/query children before the retained router namespace closes and the database owner is released. Terminal node failure ends the lifetime with a static failure reason. The 250 ms liveness watch never restarts the node or retries an operation. A stopped event follows successful cleanup; Spark must also confirm child exit before admitting a successor.

This process composition does not start containerd, provide private DNS, activate Spark's daemon or establish production readiness. Spark must consume the released Pallet bundle and commission its guard before starting this runtime, and Pallet's own containerd (containerd-serve, below) serves the CRI socket /run/pallet/containerd/containerd.sock this runtime uses.

Pallet's own containerd

The sealed control bundle carries Pallet's container runtime under pallet-containerd/: containerd 2.3.6 and its runc v2 shim (the unmodified upstream release executables), runc 1.5.1+servezone1 (built from the signed upstream sources, statically against musl), the pause 3.10.2-servezone1 sandbox image as an OCI archive (pause.tar) and manifest.json, which records every file's size, mode and SHA-256. The control executable compiles in that manifest's digest. No upstream CNI plugin ships: the only network plugin is Pallet's own pallet-cni, which is the verified pallet-runtime executable, and containerd's internal loopback serves lo.

pallet-control containerd-serve is the main process of the containerd service unit. It verifies its sibling pallet-runtime against its compiled-in digest and starts it in its --containerd-management mode, which:

  • verifies pallet-containerd/ against the manifest: exact names, modes, sizes and digests, root-owned and not group- or world-writable up to /;
  • refuses to start a second containerd while one answers on /run/pallet/containerd/containerd.sock;
  • writes the configuration below to /run/pallet/config/config.toml (mode 0600), the CNI network pallet-workloads to /run/pallet/config/cni/conf/10-pallet.conf and a link /run/pallet/config/cni/bin/pallet-cni to the verified pallet-runtime, all recreated on every start;
  • starts containerd with an empty environment except a system PATH (for host helpers containerd resolves by name, such as apparmor_parser), bound to its own lifetime (PR_SET_PDEATHSIG), and waits up to 60 s until CRI reports version v2.3.6 and RuntimeReady;
  • imports the bundled pause archive through containerd's own Transfer service over a Streaming session, as ctr image import does, when containerd does not already hold it with the pinned config digest, unpacks it for the overlayfs snapshotter under pallet.local/pause:3.10.2-servezone1, and requires CRI to report that image id. Nothing is pulled from a registry, and no ctr executable ships.

Then the process reports PALLET_CONTROL {"protocol":"pallet.containerd.process","event":"ready"} on stdout and, when started by a Type=notify unit, READY=1 over systemd's notification socket. containerd's own log lines pass through to the process's standard error. containerd's exit fails the process (failed, owner_failed, exit status 1); it is never restarted inside the process — the service manager restarts the unit. Once containerd is ready, SIGTERM or SIGINT stops it with SIGTERM, waits up to 30 s, removes /run/pallet/config and reports stopped. A SIGTERM during startup kills the starting containerd at once and fails the process; the left-over /run/pallet/config is recreated by the next start. Stopping containerd never stops a workload: every shim, and every container beneath it, keeps running and is reattached by the next containerd. Standard input is not read (a unit's null stdin is expected; data on it is refused).

The generated configuration is complete and is Pallet's own:

Setting Value
version 4
root, state /var/lib/pallet/containerd, /run/pallet/containerd
imports [] — the host's /etc/containerd/conf.d never applies
required_plugins io.containerd.grpc.v1.cri
disabled_plugins io.containerd.internal.v1.opt, io.containerd.image-verifier.v1.bindir, io.containerd.nri.v1.nri
gRPC / TTRPC /run/pallet/containerd/containerd.sock / .sock.ttrpc, uid 0, gid 0
CRI stream server 127.0.0.1:10010, stream_idle_timeout = "15m", no TLS streaming, no CRI TCP service
images overlayfs snapshotter, sandbox image pallet.local/pause:3.10.2-servezone1, no registry host directory
runtime runc via pallet-containerd/containerd-shim-runc-v2 and pallet-containerd/runc, runc state /run/pallet/runc, SystemdCgroup = false; containerd's PATH is pallet-containerd first, then /usr/sbin:/usr/bin:/sbin:/bin, so the runtime info containerd asks of a shim configured by path (-info, without the runtime's options, which looks up runc by name) finds Pallet's own runc, never a host's
CDI enable_cdi = false, no specification directories
CNI /run/pallet/config/cni/{bin,conf}, one configuration, internal loopback
other image-defined volumes ignored, network namespaces under the state directory

Workloads get cgroups under the absolute parent /pallet (the executor names it on every sandbox), never under the unit that runs containerd: runc can enable the CPU, memory and PID controllers there because the cgroup holds no processes, and a stop of the containerd unit reaches no workload.

The CRI stream server (exec, attach and port forwarding) listens on loopback only, and only root may dial it: every allocation-pool guard policy carries the Smartnftables loopback port owner { address: '127.0.0.1', port: 10010, uid: 0 } (see Offline allocation-pool guard journal). Another user's connection is reset, and the port is dropped on every interface but loopback. The containerd unit requires the guard unit, and a guard pass that did not apply the owner fails, so containerd never serves without the rule.

Service unit contract

Spark on fleet nodes and Onebox on its own host install the unit; Pallet never writes a unit file. The unit must:

  • run <bundle>/pallet-control containerd-serve with no further arguments as its main process, as root, StandardInput=null;
  • use Type=notify with NotifyAccess=all (the ready notification comes from the native owner, a child of the main process);
  • use KillMode=process, so stopping the unit signals only containerd-serve, which stops containerd itself, and every shim and workload survives;
  • use Delegate=yes, as containerd's own unit does, so systemd leaves the cgroups of the processes it starts to them;
  • set TimeoutStopSec= of at least 45 s (30 s containerd grace plus the owner's joins), Restart=always with a short RestartSec=, LimitNOFILE=infinity, TasksMax=infinity and OOMScoreAdjust=-999;
  • order After= and Requires= the retained guard unit, and be ordered Before= the node unit, which Wants= it: the node runtime is the only CRI client.

A node unit of Pallet's own (Onebox) runs the node lifetime. Spark does not use it, nor createPalletNodeServiceDefinition or palletNodeServiceSettingsMatch: its node unit runs spark runnode, which supervises runtime-serve (standard input held, --router-namespace-descriptor 3 with the namespace its own unit's store kept; see "Router namespace lifetime"). Pallet's node unit must:

  • run <bundle>/pallet-control runtime-unit --forwarding-mode <mode> as its main process, as root, StandardInput=null, Type=exec, with the host's forwarding mode (see Host forwarding modes) and no other argument;
  • use KillMode=mixed: SIGTERM reaches only the main process, which joins every owner and child it started, and whatever remains of the unit's cgroup when the stop timeout ends is killed before the service manager starts a successor;
  • set TimeoutStopSec= of at least 360 s (a five-minute native operation, then the database joins) and Restart=on-failure with a short RestartSec=: a failed lifetime is restarted by the service manager, never inside the process;
  • order After= both the guard and the containerd unit, Requires= the guard unit and Wants= the containerd unit;
  • keep the router namespace: NotifyAccess=all (the namespace keeper, a child of the main process, stores it), FileDescriptorStoreMax=1 and FileDescriptorStorePreserve=yes (systemd 254 or later), so the manager hands it back through LISTEN_FDS to the next main process; an installer releases the store (releasePalletNodeDescriptorStore) when it changes the forwarding mode and on uninstall.

A start of the node may take up to 600 s before its ready line (a start fences what the previous lifetime retained first; the derivation is in "Foreground node process"). Type=exec does not wait for it; an installer that does allows at least that long.

The node unit is started only after the node's first projection is acquired (network-acquire-local or network-acquire); an unbound node fails its start.

The guard unit is a oneshot (Type=oneshot, RemainAfterExit=yes) that runs guard-recover with null stdin, without default dependencies, ordered Before= network-pre.target, systemd-networkd.service and shutdown.target, which it conflicts with, and required by network-pre.target and systemd-networkd.service. Its executable is pallet-control or an installer's own executable that verifies the installed bundle before it starts that mode. The installer writes it only after the first guard-commission.

@serve.zone/pallet-bundle (from this repository's ts_bundle/) states these three definitions for an installer and reads the loaded containerd and node units back against this contract.

The QEMU CNI guest's containerd-serve scenario runs this process from a build:control directory without systemd (see Isolation and qualification); the unit's own KillMode, Delegate and notification behaviour is qualified with the installer.

Router namespace lifetime

The verified pallet-runtime executable also owns one private router namespace for each node process lifetime. Its exact --network-namespace-management mode creates an unnamed Linux network namespace on the initial native thread, before Tokio, sockets or readiness. Before it exposes the namespace it switches IPv4 forwarding off in all and default and disables IPv6 in both, and verifies each (with lo's forwarding): a new namespace copies the host's IPv4 forwarding, and the router must forward nothing unless an UP generation's packet policies filter it. This requires permission to create network namespaces; failure prevents node readiness and controller admission.

The node opens the keeper's live namespace descriptor, compares its device and inode, rejects its own namespace, and confirms the native nonce and identity over the original stdio connection. A reused PID or saved metadata cannot establish authority. Nothing mounts or names the namespace.

Kept across restarts. Under a service manager that passed NOTIFY_SOCKET, the keeper removes any previous FDNAME=pallet-router from the manager's descriptor store with FDSTOREREMOVE=1 and waits for BARRIER=1 before reporting ready, whether it creates or adopts a namespace. The store stays empty while the lifetime can mutate the network. Only a clean stop, after every namespace consumer has joined and the handoff set was confirmed, sends keepNetworkNamespace: remove any kept descriptor, then FDSTORE=1 with FDPOLL=0 and SCM_RIGHTS, then a barrier. A lost coordinator, lapsed lease, crash or failed cleanup cannot hand an unsafe namespace to every later restart. Only that keeper is given the socket: the node removes NOTIFY_SOCKET from its own environment before any child starts, and nothing ever sends READY=1, MAINPID= or STOPPING= — the main process is the supervisor's. An unconfirmed removal (namespace.release) refuses startup; an unconfirmed store (namespace.store) refuses the clean handover.

The next node process receives the kept namespace as descriptor 3 — runtime-serve --forwarding-mode <mode> --router-namespace-descriptor 3 from a supervisor such as Spark's runnode, or through LISTEN_FDS (exactly one descriptor named pallet-router, for this process) as runtime-unit. It reopens it as a close-on-exec handle and closes the inherited number before any child starts (compiled Deno marks the number close-on-exec first, then closes its own duplicate of it through fs and the number itself through libc, in that order; a number that is inherited but cannot be closed fails the node before any child starts), so the handle is the process's only reference and no child holds the namespace. It adopts it only as the epoch the application lane last journaled, with no unfinished ACTIVE generation. A keeper started with --adopt proves the descriptor a network namespace (NS_GET_NSTYPE) of the journaled device and inode, in the journaled boot (a soft reboot keeps the store, a reboot empties it), enters it and reads its own namespace back. The adopted namespace keeps the journal's identity — its nonce and creator PID are the epoch's tokens — so the attachment, handoff-set and ACTIVE journals see the same epoch: attached pairs are recovered in place and the handoff set is completed again. A handover from an earlier version that still has an unfinished ACTIVE generation is rejected even when its descriptor identity matches: its lease may have lapsed or its coordinator failed. Its epoch ends rather than replaying the same conflict at every restart. A handover that cannot be proven — no journal, another boot, not a network namespace, another device or inode — is closed, named on stderr (Pallet router namespace not adopted: namespace.adopt.<check>.), and a fresh namespace is created; the runs of the ended epoch then fail as network-epoch-ended. A reboot empties the store, and so does a stop of a unit systemd then unloads: FileDescriptorStorePreserve=yes keeps the store only while the unit stays loaded, which an enabled unit does, and systemd unloads an unreferenced inactive unit and closes its store. An installer releases the store when it changes the forwarding mode and on uninstall (releasePalletNodeDescriptorStore).

A stop hands the set over. The next process of the epoch completes the handoff set it adopts against the receipt the application lane journaled, and a member that receipt holds is never created again: a pair that vanished under its receipt is a conflict (handoffset.reconcile.member_vanished), not a creation. So a stop of a node whose router namespace remains eligible for a clean handover (PalletNamespaceOwner.kept: the manager's socket is available and no owning failure invalidated it) releases nothing in it: the handoff coordinator's handOverHandoffSet withdraws and discards a generation it still holds and closes the uplink observation, exactly as the release does, then proves the set it leaves — no transition pending, and the complete inventory, the dormant pairs and every workload pair attached through them, equal to its receipt — and ends its protocol with every link in place; the owner's closure reads released and handedOver. A set that changed under its receipt is refused by name (handoffset.hand_over.<check>) and the stop is unconfirmed, as a refused release is. Pallets 36.0.0 to 36.1.1 released the set at every stop without an attached workload, and every successor in the same boot failed handoff_conflict:handoffset.command.reconcile until a reboot emptied the store (lab 2026-10-04); a stop with an attached workload was refused instead, since a release would remove a pair a running sandbox still uses. A namespace nobody keeps ends with the process, so its set is released as before. A failed lifetime or crash leaves the descriptor store empty and the namespace ends with its last live descriptor, every link in it with it. Once the store is released (releasePalletNodeDescriptorStore) or emptied by a reboot, the next process creates a fresh namespace and proves the previous epoch's host attachments absent ("Previous-epoch host attachment audit"); kernel namespace teardown is asynchronous, so that audit may refuse the first start right after a release while the kernel is still deleting the pairs.

Attachments outlive their generation. An attachment created under an UP ACTIVE generation is journaled bound to it (activePreparedDigest), and a DEL names the generation held at that moment, so a DEL of a pair bound to a generation that has since ended would never be admitted. Whenever a generation of the node's own epoch is DOWN — withdrawn in process, or recovered DOWN by a process that adopted the kept namespace — when a node process of that epoch starts, and again before any successor is prepared, the attachment owner carries the epoch's live attachments past it: the handoff coordinator re-reads each pair in the kernel (inspectWorkload, the links by name, address, marker and index in their own namespaces, as CHECK does), and a pair that is complete or raised with its journaled receipt is released from the ended generation in one journaled, idempotent step. It then reads as a pair prepared before ACTIVE: the successor is prepared over it, its own recovery accounts for it, and its DEL names whichever generation holds it — the successor, or none in between. A pair that cannot be proven is fenced instead: its run reads as of an ended epoch, so the execution owner stops its sandbox and fails it as network-epoch-ended; its DEL releases and removes it; and no successor is prepared while one is fenced (active.activate.unprovenAttachment). Each carry proves every pair again, so a restarted process fences the same pairs. A fenced run is no DNS source: no name is served for a pair that is not proven.

Carried pairs are raised again. A recovery fences every pair of the generation it takes DOWN — a crash, or a stop that left the generation UP in the journal — and the successor is prepared over those pairs, but only an ADD ever raised a pair, and no ADD comes for a running sandbox: Pallet 36.0.2 left such a workload running with its link DOWN, and reported it ready. Once the barrier of a newly raised generation holds, the attachment owner re-reads every live, carried pair of the epoch in the kernel and raises each one that is not raised (PalletWorkloadAttachmentOwner.raiseCarried), writing Pallet workload attachment raised again: run <digest>.; one it cannot prove or raise is fenced — its run reads as of an ended epoch and is stopped and failed by name — with Pallet workload attachment down: run <digest> CARRY: <label>.. A pair a clean withdrawal left raised is only read.

A sandbox that vanished while the node was down. A process that adopted the kept namespace restores each attachment of its epoch by capturing its sandbox namespace at the journaled path. When that path is gone and nothing of the pair is left in the router namespace — no link by its router name, alias or either MAC address, no address or route in its subnet, read twice — the pair went with its sandbox (a pair's two ends live and die together), and the coordinator says so (recoverWorkloadAttachment answers restored: false) instead of ending. The pair is fenced from the start: its run fails as network-epoch-ended, no ADD or CHECK is answered for it, an ACTIVE recovery takes it as retired, it reports absent, and its DEL proves its absence again and completes it. A path that still exists — the sandbox, or something else at its name — or a pair whose router side is still there still refuses, and the node fails as before.

A DOWN the coordinator does not hold. Every packet-cleanup step reads the generation's DOWN from the coordinator: the generation restored by its own withdrawal, or the recovery it accepted. The coordinator of a later process of the same epoch holds neither, so a generation an earlier process took DOWN without finishing its packet cleanup is taken DOWN again as a recovery (PalletHandoffSetOwner.holdsDown), which a fresh coordinator admits before it adopts its set and which changes nothing over a generation already DOWN, and the cleanup is finished under that DOWN.

A lost coordinator costs the generation, not the node. When the handoff coordinator ends — its process exits, for instance after it fenced a kernel event the UP generation cannot tolerate, or a command finds it gone — while this keeper still holds the router namespace, the node joins the application owner composed over that coordinator and starts a new one over the same namespace. Its start recovers the attachments and fences the retained generation DOWN exactly as a restarted process of the same epoch does, and the next admitted pass raises a new one; passes are refused as unavailable meanwhile. A replacement must be followed by a pass that raises a generation (barrier up or forwarding) before another loss is replaced: a second loss with no raise between them fails the node, so a conflict that recurs on every raise never becomes a rebuild loop. The bound is kept by the node process; a restarted process admits one replacement again. A lost keeper, a native refusal, any other failure, or a replacement that does not start still fails the node.

The private PalletNamespaceOwner.withDescriptor() API retains the source descriptor through an admitted consumer's asynchronous spawn. Callbacks must not close it or keep its number after returning. Shutdown fences new borrowers and joins admitted callbacks. The enclosing node must also join all DNS, VPN and packet-policy processes before closing this owner. Unexpected keeper exit fails the node, stops admission and retains the parent descriptor until explicit joined cleanup; it never silently recreates the old namespace.

The isolated Linux 6.18.35 amd64 fixture runs both Node and compiled Deno (the pinned 2.9.7) against the real namespace keeper and published Smartnftables 1.4.0 binary. It verifies descriptor inheritance, actual namespace identity, keeper-loss fencing, joined borrowing and unchanged parent links and routes, and — on a host that forwards IPv4 — that the namespace is exposed with forwarding off in all, default and lo and IPv6 disabled in all and default. The full node composition is separately checked through its compiled control bundle. Privileged ARM, containerd/CNI attachment, DNS and complete network policy remain separate gates.

Staged secret material mounts

The same verified executable owns one staged mount per admitted assignment that carries a secret reference. Its exact --secret-mount-management mode stages that attempt's material on its initial native thread, before Tokio, sockets or any other work: it unshares a private mount namespace, mounts a noswap,nodev,nosuid,noexec tmpfs there, writes one root-owned file per entry with the delivered mode, places the ownership marker inside the mount and detaches it with open_tree. This requires permission to create mount namespaces and mounts, and it must run in the host mount namespace — the helper refuses its own work when it does not.

The material never crosses the JSON IPC. PalletSecretMountOwner spawns one child per operation and writes the opened values as one published Smartrust sensitive frame on descriptor 3, bounded by this attempt's own declared total rather than by a shared ceiling. Prepare and attach share that one child, because the detached mount dies with the process that holds it; inspect and release run in fresh children, because recovery may never depend on a process that is already gone. No value is ever a return value, a persisted field, an environment value, a command line or an IPC member, and the caller's buffer is cleared after the write.

Publication follows durable state, never the other way round: the mount plan and the pending run commit in one transaction before the first native step, the prepared receipt — secured parent mount, base and staged device/inode, detached mount id, ownership nonce — is persisted before move_mount publishes the mount, and the published mount id is verified equal to the id the detached clone already allocated. The executor then demands exactly that plan's read-only private per-file binds from the container, never a bind of the mount directory that holds the marker, and refuses a live container whose mount list is not that plan.

Recovery reads the persisted receipt alone. A fresh helper verifies the recorded mount by descriptor, unmounts strictly, proves that exact mount id gone and the base directory empty before removing it, and answers uncertain with an exact reason on any mismatch — which retains the mount, the directory and the record and deletes nothing. Stop retains the mount; release is an explicit owner step after proven CRI absence, never a side effect of removal. That order is the only thing that protects a running workload: a container's own bind of a file in the mount lives in its own mount namespace and does not make the host mount busy, so the kernel would not refuse a release taken out of order. A record from an earlier boot names nothing this kernel can find, because /run is a fresh tmpfs after every boot, so it is retired without any native step.

The chain /run/serve.zone/pallet/secrets is walked through no-follow descriptors. /run/serve.zone is the parent the installer supplies for every serve.zone product on the node, so it is required to be exactly what the CNI plugin requires of it — a root-owned directory no one else can write — while the two components below it belong to Pallet alone and must be root-private.

The offline QEMU qualification drives the shipped binary itself, over the same protocol, through every stage boundary on 6.18.35-0-virt, 6.8.0-124-generic and 7.0.0-22-generic, and the containerd guest runs one admitted assignment end to end: sealed material opened, mount published, the container created with exactly the plan's bind, the marker invisible inside it, stop retaining the mount and removal releasing it after proven absence.

Local storage claims

A controller — Cloudly or Onebox — gives a service a retained, node-local volume with applyRuntimeLocalStorageClaim (IRuntimeLocalStorageClaim, @serve.zone/interfaces 32.24.0). A claim names the volume by identity only: controller, organization, cluster, node, runtime namespace, service and the service's volumeId. It carries no host path, no device, no mount option and no credential; this node resolves the identity to its own storage.

Layout. Every volume lives under /var/lib/serve.zone/pallet-storage: volumes/<volumeKey> is a volume, staging/ holds one being created, trash/ one being purged and import/ the bytes an operator places for an import claim. The volume key is the SHA-256 of the claim's identity and id, so it is stable across every generation and a re-epoched controller keeps its volumes. The root and its four directories are root-owned 0700 below trusted ancestors; a volume is reachable only as the bind a granted container receives. The root is created only while no volume was ever materialised: a root that later goes missing — an unmounted disk — is refused by name (storage.root.missing) and never replaced by an empty one.

Ledger. The claim ledger (pallet_storage_claims, pallet_storage_names, pallet_storage_grants) lives in the node's own NoSQLDB through @lossless.org/client/nosqldb. Admission is one transaction fenced to the session the claim arrived on: a new claim, the next generation (whose previous must name the one this node holds), a replay of the same generation or a historical older one. A same-generation claim with another digest, a next generation that changes the volume's identity, source, initialization or initial ownership, a generation after a purge and a second live claim of the same service volume are refused by name and write nothing. The answer is the volume's state: ready, import-required, purged or uncertain.

Crash safety. Every filesystem effect runs between a ledger intent and the completion that records its evidence, one at a time. An empty claim commits provisioning, then creates the directory in staging/ with the claim's initialOwnership and mode and publishes it by one atomic rename; the ledger records the directory's device, inode and birth time. A purge commits purging, renames the recorded directory into trash/ and removes it there, never following a link the workload left inside. Before the ledger records a step done, every directory it created, renamed or removed is fsynced with its parents (a new layout directory with its parent; the volume with staging/ and volumes/ before ready; volumes/ and trash/ before purged), and a failed sync keeps the intent pending, so a power loss can only undo a step whose intent the next start still finishes. On start, and on every replay, an interrupted intent finishes from what is on disk: a leftover staging tree is this intent's own and is removed, a published directory is adopted only while it is still exactly the fresh, empty directory the intent creates, and a purge whose volume already sits in the trash completes. Evidence that contradicts the ledger — a foreign directory at the volume's name, a volume whose identity changed — leaves the claim uncertain, which this node neither mounts nor deletes until an operator decides. A pending intent that cannot finish at start is listed in the execution owner's storageRecoveryRefusals and stays durable.

ReadWriteOnce. A run that seals storage references is planned before the registry, secret material or any native step: every reference must name the claim's current generation, the claim must bind to the run (bindRuntimeLocalStorageClaimToAssignment), its volume must be ready, still be the recorded directory and be held by nobody else, and no two volumes of the run may nest or cover a path Pallet or the runtime binds itself (/run/secrets, /run/serve.zone, /opt/serve.zone/runtime-assets, /proc, /sys, /dev, /etc/hosts, /etc/hostname, /etc/resolv.conf). A run that cannot have its volumes waits as storage-pending, with the step named (storage.claim.held, storage.claim.importPending, storage.claim.generation, …). The grant is taken in the transaction that begins the run, so a concurrent purge or a second holder conflicts instead of both proceeding. Stop keeps the volume held; the proven CRI absence of the removed container releases it. Release never deletes bytes: only a purge generation does, and a held volume's purge is refused (storage.purge.held) until its holder is removed. The container's mount list is part of the attempt's identity, so every later run or inspect reconcile demands exactly the grant's binds.

Offline import. An import claim waits as import-required until its bytes arrive. With runtime-serve stopped, the operator places each volume's tree at /var/lib/serve.zone/pallet-storage/import/<claimId> on the same filesystem and runs pallet-control storage-import. The pass (in ts_migration/, like every data adoption step) records the placed directory's identity, publishes it as the volume by one atomic rename, fsyncs import/ and volumes/ and records it ready; the tree keeps its ownership, modes and every other attribute, and nothing is copied. It reports the adopted, waiting and refused claims on its ready event (pallet.storage.import). A tree on another filesystem, a file instead of a directory or a volume name already in use is refused by name and moved nowhere; an interrupted pass finishes on the next run, and a recorded tree found in neither place leaves the claim uncertain. Keep a copy of the source first when a rollback may need it.

Network filesystems are not part of this build. The contract (32.24.0) also states nfs and smb sources, but this node materialises node-local directories only: such a claim is refused before the ledger is touched (storage.source.unsupported), so neither the claim nor its service volume name is recorded.

Dormant host/router handoff

The private ts/network/classes.handoffowner.ts mechanism consumes an already reserved, digest-verified immutable handoff lease and a live PalletNamespaceOwner. start() prepares an inert intent bound to the current boot and both namespaces. The caller must persist that complete intent before reconcile(previousReceipt) can create the exact veth pair, with its peer created directly in the router namespace. Both endpoints remain administratively DOWN. Only the leased IPv4 addresses and their kernel local /32 routes are present; connected prefix routes belong to eventual activation.

The separate --network-handoff-management native mode opens one netlink socket in each retained namespace on its initial thread and restores the parent namespace before polling either socket. It uses bounded native operations without shell commands, named mounts or filesystem state. Linux ignores aliases in this veth creation request, so the owner first verifies the newly created DOWN endpoints, sets both ownership markers through link updates, and verifies them before adding addresses. An interrupted partial creation never qualifies for adoption or deletion.

Recovery requires the exact boot, lease, namespace, link indices, peer indices, names, MACs, markers, addresses and local routes. Extra or changed state is rejected. Deletion uses the verified endpoint index and destroys only that exact pair. A release states the administrative state its own contract guarantees: every handoff path releases a dormant pair, so an endpoint raised behind the owner's back is refused as handoff_conflict and the pair stays, and only the raise measurement releases a raised pair. Removing a raised workload pair in production is the stage-aware progress rollback, never this release. Privileged exclusivity belongs to the enclosing node; these observations are not a compare-and-swap guarantee against concurrent privileged network writers.

The TypeScript owner serializes operations and fences namespace or native-process failure. getPrepared() and getReceipt() retain detached immutable evidence after failure. close() joins the native lifetime; a successful return with released: false means cleanup remains unconfirmed and the lease must remain quarantined. A rejected close has not proved that join; joined exposes its state. Join every namespace consumer before closing the keeper. Kernel namespace teardown is asynchronous, so keeper exit alone cannot establish link deletion or pool reuse.

Both native handoff management modes use bounded nonblocking pipe or stream-socket stdio on the same thread that owns their namespace sockets. Host, router and retained sandbox netlink connections remain polled while input is idle or partial and while a response waits for the parent to read it. Partial input survives cancelled reads; malformed or oversized frames terminate the protocol. Output has the same three-second terminal bound as native commands. A partial response is never retried on the same stream. No reader thread can retain the mutation lock. This establishes protocol-owner liveness; the dormant sockets still have no topology multicast subscriptions and provide no continuous route or link fence.

This mechanism is not yet composed into signed node network admission. Durable realization receipts, router/host/Docker policy barriers, CNI, DNS, link activation and production qualification remain separate prerequisites. The offline handoff fixture is test/native/handoff.ts, driven by test/native/qualify-handoff.py in a disposable amd64 VM without a network device, host disk or host mount. It runs the actual native binary through Node and compiled Deno, including partial-state and foreign-state faults, restart recovery, keeper loss and joined cleanup. ARM64 is built but this privileged fixture does not establish ARM64 execution qualification.

Complete dormant handoff set

ts/network/classes.handoffsetowner.ts supplies the private node-wide mechanism for complete retained membership. Start it with the live namespace, persist the empty getPrepared() intent, then reconcile(target, null) to inspect and establish an empty native receipt. prepare([{ lease, presence }]) accepts all 256 complete, digest-verified leases permitted by the allocation contract, including absent/quarantined history, with at most 32 desired present handoffs. Persist the returned target and the last applied receipt before calling reconcile(target, previous). Preparation does not change the kernel. Receipts contain an exact DOWN pair or explicit absence for every member; absence never frees an allocation.

The --network-handoff-set-management process retains both namespace descriptors and owns one host-network-namespace abstract Unix socket lock. The single-pair mechanism uses the same lock, so no two Pallet handoff mutators can coexist on that host namespace, even with different router namespaces. This IPC lock is concurrent process exclusion, not authentication or durable ownership. The node must still exclude other privileged network writers and namespace capabilities.

Before any transition effect, bounded complete link, address and all-family/table route dumps inspect the whole expected set. The private router namespace must contain only its empty, DOWN loopback and exact dormant handoffs. Unknown router interfaces, addresses or routes fail ownership. Unrelated host interfaces and routes remain permitted; unexpected Pallet markers/names, misplaced known MACs, transit address conflicts and surviving/reused receipt indices are rejected. Stable bound facts are re-read after inspection and after the complete transition. Typed requests retain the terminal NLMSG_DONE; every inventory requires a successful completion and rejects interrupted, filtered, malformed, unexpected or oversized responses. Both namespace connections also reject unsolicited messages and receive-buffer loss. The pinned maintained netlink-proto source corrects decoding at its owning layer so malformed packets cannot disappear before a later successful completion.

Transitions retain all prior members and cannot revive an absent member. Exact present-to-absent deletion keeps the lease in the set. Separate pair operations are not atomic: after a lost response, persistently prepared intent and the prior receipt can adopt exact completed members and finish the remaining transition in the same live namespace. Changed or half-created pairs stay failed-owned. EOF or process termination preserves DOWN effects for that recovery. Full node death does not make an unnamed namespace recoverable. close() deletes only exact known members and joins the native lifetime; uncertainty preserves receipts and returns released:false, with no guessed cleanup or allocation reuse. Over a router namespace a service manager keeps, close() hands the set over instead (handOverHandoffSet, "A stop hands the set over" under Router namespace lifetime): it deletes nothing, proves the set equal to its receipt and returns released: true, handedOver: true.

The native transport permits 1,048,576-byte frames; TypeScript captures inert complete-set data within the corresponding bounded budget. The serialization test measures a 701,951-byte replay with maximum-length lease and namespace identities. Native operations retain their three-second deadline and the bridge its five-second deadline: the 256-retained/32-present guest operations completed within the native deadline in the isolated fixture. Complete dumps establish absent members without issuing redundant per-member link dumps. The offline handoff fixture additionally covers empty/nonempty exhaustive inventory, unrelated host links, competing single/set owners, a competing router namespace, quarantine, lost-response and partial-transition recovery, 256 retained members with 32 present paths and joined wrapper shutdown under Node and compiled Deno. This stage exposes no UP method and issues no nativeBarrier: the authenticated SmartData application journal, positive protection evidence and eventual firewall/DNS/VPN/CNI lifecycle remain required.

After reconciling the complete dormant topology, the private handoff-set owner can openUplink() and return a boot-, host-namespace- and random-generation-bound observation. inspectUplink(generation) performs fresh bounded kernel and networkd reads; copied facts cannot recreate the observer. A wrong generation is rejected without retiring the current observation. closeUplink(generation) joins the read-only observer and does not release topology or change DHCP.

A generation routes its transit out through the uplink, so the host must forward IPv4 while it is UP. Who owns that switch is a fact of the host, which its installer states on every start of the node (--forwarding-mode, see "Host forwarding modes"); the observation follows it:

  • exclusive: the coordinator owns the switch. The observation opens only on a host that does not forward (all, default and the uplink link off; otherwise uplink.forwarding.open), records the binding with those switches off, and from then on compares every read against it with the three switches at the value the coordinator holds. A sysctl write notifies with the kernel's own port, so the coordinator arms its switch before writing: the armed switch admits the forwarding-only NETCONF changes to the value it writes, once per index — the uplink, all, default, the held host link and every dormant link alike — and, switching on, the feature change of a link whose LRO the kernel turns off for forwarding; the read that follows drains every notification the write queued, proves the binding unchanged but for the three switches, and requires the kernel to have notified each of them.
  • docker-shared: Docker owns the switch. The observation opens only on a host that forwards on all three (uplink.forwarding.open otherwise: Docker's switch is gone), and the coordinator never writes it.

Any other forwarding change of the binding, armed or not, retires the observation, as does a NETCONF change of any other setting. An operation states the forwarding it requires — exclusive: on exactly while the generation is UP; docker-shared: on throughout — and is refused by name otherwise (uplink.forwarding.on, uplink.forwarding.off).

The initial supported host has one physical Ethernet uplink, a finite bound systemd-networkd DHCPv4 lease, one main-table IPv4 DHCP default route and the three canonical unselected IPv4 policy rules. The observer records exact link, address, gateway, source, metric and IPv4/IPv6 all/default/interface NETCONF settings. Beside the lease the uplink may carry permanent /32 IPv4 host addresses — the addresses a platform service on this host binds, whether networkd configured them (Address=) or found them on the link. Each is a universe-scope address with an infinite lifetime, so it adds no prefix route and never becomes the default route's source; any other IPv4 address on the link still refuses the observation, and adding or removing one retires it like every other address change on the uplink. The observation states them (hostAddresses, sorted as strings); an observation a 33.0.0 node journaled carries none, which states nothing. Existing host IPv6 addresses remain possible; this observation grants no IPv6 workload authority. Networkd remains the only DHCP mutator.

The native lifetime subscribes before capture and drives both RTNL notifications and the fixed system D-Bus connection during inspection and quiet management I/O. Notifications that arrive while the opening read runs, before any binding is known, are held and judged against the binding that read finds, exactly like later ones: the subscription predates the read, so the binding reflects every change before them, a change that can affect it still fails the open closed, and another link's churn during the read (a container starting beside Pallet at boot) does not. The observation judges notifications ahead of its own work, so it bounds their rate, not their number: more than 4096 within one one-second window of BOOTTIME is a flood that would starve a read or the wait for the lease's renewal, and it retires the observation; notifications it ignores never accumulate toward the bound however long it lives, so continuous churn of other links, Pallet's own workload veths included, never ends it. At most 4096 notifications are held while the opening read runs. Its unique networkd owner and exact returned link object path remain bound to the lease. The observation follows the lease networkd renews, so a node keeps forwarding through every DHCP renewal: at the lease's renewal time (T1), on a property event of the link, and when networkd rewrites the leased address with new lifetimes, the native reads the uplink and its lease again. While networkd renews or rebinds, the lease in force holds until its valid lifetime ends, never longer. A fresh acquisition from the same networkd owner for exactly the same binding — link, address, prefix, gateway, source, metric, NETCONF settings and host addresses — whose lifetimes the kernel address already carries becomes the lease in force; the observation keeps its generation and its receipt, which states the lease it opened on. The lease's expiry, service replacement, protocol loss, a read that finds the binding changed in any fact or the lease moved backwards, and every other relevant kernel change retire the coordinator. Current values cannot erase an observed kernel change and restoration: only the leased address rewritten in place asks for a read, and any other IPv4 address, link or NETCONF change of the uplink, a NETCONF change of all or default, and every IPv4 policy rule change still retire. Unrelated link notifications are ignored.

A route change retires the observation when it can change the binding's path. The binding is IPv4, and its snapshot proves the three canonical rules, so an IPv4 lookup consults exactly the local, main and default tables (255, 254, 253) until a rule change retires the observation. A route retires it when it is IPv4, sits in one of those tables, and leaves through the uplink, has a destination that overlaps the leased subnet (which holds the leased address and the gateway; every default route overlaps it) or covers a host address, or has its gateway in the leased subnet. A route that cannot be judged exactly — a repeated or contradictory table or attribute, multipath, a nexthop object, an encapsulation, an unknown attribute or another family — retires it as well. IPv6 routes, IPv6 rules and the uplink's IPv6 addresses cannot change the IPv4 binding and are ignored, as are IPv4 routes of other links outside those prefixes and routes in tables no canonical rule selects. So another tool's links coming and going leave the observation alone: Docker starting a container or stopping, with the kernel's IPv6 link-local, multicast and local routes of its veths and bridges, or a bridge's IPv4 connected, local and broadcast routes. Pallet's own host links are judged apart. A route in a consulted table that leaves through the held (ACTIVE) host link or a dormant one, or whose destination or gateway lies in that link's transit prefix or in the pool routes it carries, and any IPv6 route naming it, belongs to that link: outside a pending host transition it retires the observation, and during one only the kernel's own such change reaches the transition, which accepts exactly the change its link state makes. Every other route, during a transition too, is judged against the uplink binding as above.

The kernel can stamp a change's notification with the requesting socket's port and sequence (an address or route notification does, networkd's lease rewrite included), so a sequence number never makes a notification Pallet's: only the port of the command socket a pending host transition bound does, and only for that transition's pool routes. Every other notification — the kernel's own or another requester's — is judged on what it changed; another requester's change to a host link in transition or to a route touching it is never the transition's. Close the observer before further dormant topology changes.

The ACTIVE observers of the router and sandbox namespaces tolerate one kind of notification no matter the phase: a link change whose kernel change mask names only IFF_PROMISC or IFF_ALLMULTI — what a capture on the link (tcpdump) or an administrator's promisc/allmulti setting makes the kernel send. Such a change forwards and drops nothing; the generation's inventory reads past a link's promiscuity and all-multicast counts and flags. An idle coordinator reads the generation's complete inventory again at once, exactly as an inspection reads it (handoffset.drive.active_reverify names a failure of that read); an owned step notes the change, and its own inventory reads the links. Every other link change — up, carrier, MTU, name, address, a change mask beyond those flags or none — still ends the observer. The dormant handoff-set audit stays strict. The host's uplink observation is unchanged.

A coordinator that ends while it waits — its uplink observation retired, an ACTIVE observer failed, a namespace channel lost — lowers every ACTIVE link on its way out, as it always did. It names the watch that ended it in a bounded PALLET_FAULT line (handoffset.drive.*, uplink.watch.*, uplink.follow.*); an ACTIVE observer that refused an event names the event's kind in its own line (active.observe.unexpected_event.<kind>: new_link_flags, new_link_state — any link message without a change mask, such as a carrier, operational-state, MTU or master change — del_link, new_address, del_address, new_route, del_route, new_rule, del_rule, netconf, control or other), and the handoff-set owner writes its failure chain once, as it retires, to stderr, which Spark carries into the node journal: Pallet network handoff set failed: <chain>. The stop that follows names the same chain as the reason its release could not be confirmed (noderuntime.close.networkRetained < …). A command the coordinator refuses names itself the same way (handoffset.command.<command>, after any line the refusing check wrote), so a refusal answered over IPC with only its code still reaches the chain as PalletNativeFault:<code>:<site>.

This mechanism exposes no UP or packet-readiness result. Active routing, firewall composition, durable active journals and joined withdrawal still belong to the activation coordinator. The ignored Rust test network_handoff::kernel::uplink::tests::qualified_networkd_host_observation provides an explicit read-only check on a real networkd-managed host; it changes no interfaces, addresses, routes, sysctls or services.

Host forwarding modes

A host forwards IPv4 between all of its links once forwarding is on. Docker switched it on when it started and set the iptables FORWARD policy to DROP, so a Docker host forwards only what a rule admits; a host without Docker has neither, and switching forwarding on alone would let a LAN peer use the node as its gateway. Who owns the switch is a fact of the host, which the node never infers: pallet-control runtime-serve and runtime-unit take --forwarding-mode <docker-shared|exclusive>, required, with no default (palletForwardingModes in @serve.zone/pallet-bundle; Spark states it on fleet nodes). The handoff coordinator receives the same mode (--network-handoff-set-management --forwarding-mode <mode>) and proves that mode's host conditions as it starts, before anything else, refusing by name:

  • exclusive — the node is the host's only forwarding owner:

    • no persisted IPv4 forwarding: any net.ipv4.ip_forward or net.ipv4.conf.<scope>.forwarding other than 0 in /etc/sysctl.conf or a *.conf of /etc/sysctl.d, /run/sysctl.d, /usr/local/lib/sysctl.d, /usr/lib/sysctl.d or /lib/sysctl.d refuses (forwarding.exclusive.persisted): it makes the host forward unfiltered from boot until the coordinator starts;
    • no Docker under the service manager: docker.service or docker.socket loaded and neither inactive nor failed refuses (forwarding.exclusive.docker_present); the exclusive drop and the switched-off forwarding would stop every container's forwarding, and a restarting Docker would switch forwarding on again behind the observation. A system bus with no service manager on it has no such unit.

    Then forwarding is switched off, and on again only inside activateActive once the generation is otherwise UP behind its exclusive host-transit table (see Retained uplink observation and Private packet policy composition). Right before that, the packet engine reads every owner on the forward hook (inspectForwardHooks, Smartnftables 4.2): each forward base chain must be the generation's own host table or the node's allocation-pool guard, by exact table name and kernel handle (active.activate.foreignForwarder), Docker's DOCKER-USER and DOCKER-FORWARD chains must be absent (active.activate.dockerPresent) and no legacy xtables table may be registered (active.activate.legacyForwarder). The generation is refused before it is raised, and forwarding stays off.

  • docker-shared — Docker owns the switch:

    • the host provides the iptables-nft frontend the contribution runs: /usr/sbin/iptables-nft, iptables-nft-save and iptables-nft-restore, root-owned (the symlinks and their target), at 1.8.10 or newer on nf_tables (forwarding.docker_shared.iptables_nft);
    • forwarding is never written, at start, activation, withdrawal or exit; the observation binds only a host that forwards.

    Every generation's host-transit policy is shared (no forward drop), and its flows pass Docker's FORWARD policy through the node's DOCKER-USER contribution (@push.rocks/smartnftables ManagedDockerForwarding), which the engine applies only over Docker's policy DROP with DOCKER-USER and DOCKER-FORWARD as the first two jumps (its link fence follows below). Per generation, one owner (the node's owner name, a fresh instance kept in the journal, pallet_active_docker_forwarding): the first step over the generation's host table once it is applied and journaled, before activateActive; a step over every later host table of the generation (tunnel phase, amendment, lowering), so the engine's reading of its barrier stays exact; each step admitted on the journal before the engine sees it and acknowledged after. The withdrawal releases it after the handoff is down and before the host table it admitted. A stop withdraws the generation it holds instead of leaving it UP, so no contribution rule outlives the node process; only a crash leaves one, for the next process to release. A fresh owner recovering a generation reopens the contribution under its journaled identity and releases it: an applied contribution through release, and a last step that was admitted and never acknowledged without applying it (Smartnftables 4.6 releaseTransition). Re-applying such a step would need its barrier — the host table it was prepared over — which the engine requires in this boot and may already be gone. The release request, the step's exact transition under the owner's engine identity of this boot, is journaled first, so a retry replays it exactly. The engine deletes exactly the rules of that step's target and previous contribution it still finds, and the end records the rule indices it deleted ({ kind: 'transition', request, removed }). A rule under the owner's prefix that is neither refuses the whole release by name before any deletion (docker.releaseTransition:FOREIGN_OWNER, a journal/table disagreement). The refusal settles nothing, so no successor begins (active.begin.previousContribution). Nothing retries it within the process; each restart replays the journaled request and meets the same refusal, failing closed until the table changes. The offending rule is one in table ip filter whose comment carries the owner's prefix, snftd1:<ownerId>: (the owner name is pallet_docker_ and 40 hex digits, journaled as the contribution's identity.ownerId; the comment continues with the instance, the contribution digest and the rule index), but that is not exactly one of the step's target or previous contribution rules in DOCKER-USER: another instance's or another step's rule, or one altered, moved to another chain or duplicated. An operator lists the owner's rules with iptables-nft -t filter -S | grep 'snftd1:pallet_docker_' and deletes the stray ones by their listed specification (iptables-nft -t filter -D <chain> <spec>). Deleting all of them is also safe: the release runs only once the generation is DOWN or being withdrawn with its handoff down, no successor begins before it settles, and it deletes only the step's rules it still finds, so the next start releases the step with whatever is left and records it. A contribution whose steps name an earlier boot ended with that boot, which took every DOCKER-USER rule with it: it is recorded as ended with its boot ({ kind: 'boot', bootId }) and no engine is asked, since the engine refuses any step or release of another boot. That holds for whatever DOWN the generation carries and whether its cleanup finished, and also for a release begun and never confirmed (production 2026-10-08: a docker-shared node rebooted into 36.3.1 over a DOWN generation whose first step was never confirmed under 36.2.1 failed every start at docker.reconcile:CONFLICT). test.activedockerboot.node.ts covers that journal, left recovered and discarded, across a reboot (no engine call, a successor raised) and a release an earlier boot began; test.activedockerrelease.node.ts the same journal in this boot (released through releaseTransition, never reconciled, before its host table), a lost release result replayed exactly, a FOREIGN_OWNER refusal, and a DOWN whose cleanup never began. The orchestrator refuses to raise a generation whose policy does not match the mode (active.activate.sharedForwarding, active.activate.exclusiveForwarding), and a refused contribution step fails the activation by site (docker.prepare, docker.reconcile, docker.release, docker.releaseTransition).

    The contribution (Smartnftables 4.2) admits exactly the flow kinds the shared host table admits, each with its replies: the leased flows (the router's own egress, such as its VPN client dialling the hub, included), the inbound publications and the symmetric publications. Its link fence: the first step applies only with the generation's handoff down; a later step, rules changed or not, applies while the generation is UP when the handoff link both steps bind stays exactly the same — every step of a generation binds its one selected handoff, so a publication amendment, the tunnel phase and the lowering follow in place — and a link a step adds or drops must be down. The release needs the bound handoff down or gone: a link whose index no link holds any more counts as down, so a successor in the same boot — after a crash, its router namespace gone — releases the previous contribution before it applies the next. An owner that holds no rule of its own closes on any host (Smartnftables 4.3): one whose first step the engine refused — a host whose FORWARD policy accepts — and one of a journaled identity on a host whose Docker and filter table are gone both close as released. A close that cannot confirm its cleanup fails CLEANUP_UNCONFIRMED with the engine's native code as its cause, which the node's failure chain names (PalletDockerForwardingError::docker.close < ManagedNftablesError:CLEANUP_UNCONFIRMED < ManagedNftablesError:<code>).

Moving a host from docker-shared to exclusive (the D9 step) is done with the node stopped: state exclusive, stop the node (its stop withdraws the generation and releases the contribution and tables; forwarding stays Docker's), remove Docker and every persisted forwarding setting, and reboot. The rebooted host has no table and forwards nothing; the coordinator's guards pass, and the first UP generation switches forwarding on behind its exclusive table. Stating the wrong mode fails closed and by name: exclusive on a Docker host refuses forwarding.exclusive.docker_present, docker-shared on a host without Docker refuses at the observation (uplink.forwarding.open) or in the contribution (Docker's chains or policy absent).

The D9 safeguard is that order, and only that order: stop the node, remove Docker, remove every persisted forwarding setting, reboot, then start the node as exclusive. Never remove Docker from a running docker-shared host and carry on without the reboot: Docker's switch stays on, and nothing in the node switches it off (a docker-shared coordinator never writes it), while Docker's FORWARD policy DROP — the only thing that kept the host from forwarding between all its links — has no owner left to restore it: once it is gone with Docker's chains (a firewall reload, a flush, the package's cleanup), the host forwards everything until the reboot. An exclusive coordinator a rebooted host refuses (persisted, docker_present) writes nothing, its exit included: the host's forwarding is left exactly as the refusal found it.

The start guards have limits; they narrow the window, they do not close it:

  • forwarding.exclusive.docker_present sees Docker only as docker.service or docker.socket under systemd on the system bus. A Docker daemon started outside systemd, or under another unit name (a snap's snap.docker.dockerd.service), is not seen at start. The activation-time read of the forward hook still refuses Docker's DOCKER-USER and DOCKER-FORWARD chains, but it is a point in time.
  • forwarding.exclusive.persisted reads the sysctl configuration files only. Forwarding switched on from the kernel command line (sysctl.net.ipv4.ip_forward=1), by a networkd IPForward= / IPv4Forwarding= setting, or by any other tool at boot is not seen: such a host forwards unfiltered from boot until the coordinator starts and switches it off. Once the coordinator runs, the observation refuses a host switched on behind it (uplink.forwarding.open) and fences a switch changed under an UP generation.

Downgrading a node from 36.1 to 36.0.3 or earlier is not supported once a generation has detached an attachment in place: the transition row that records it (detached, below) is refused by the exact reader of those releases, so their process fails at start. A node that never detached one reads its journal under 36.0.x unchanged.

Downgrading a node from 36 to 35.x is not supported in place: 35.x cannot read what 36 journals. It recomposes an exclusive generation's host policy without exclusiveForwarding, so the journaled policy no longer matches, and it does not know the contribution journal (pallet_active_docker_forwarding), so it never releases a contributed DOCKER-USER rule. The contributed rules live only in the kernel's ruleset, which a reboot empties; 36 records a contribution of an earlier boot as ended with that boot and never releases it through an engine. A contribution journal ended by a transition release (kind: 'transition') is refused by 36.3.1 and earlier, so a node that wrote one cannot be downgraded in place until two later generations were raised. Independently of the journal, once 36.3.2 reconciled a table on a node, a downgrade to 36.3.1 or earlier works only after a reboot: 36.3.2's Smartnftables 4.6 writes its managed sets with their key and data types, which the Smartnftables 4.4.x of those releases refuses to adopt, and a reboot empties the kernel's ruleset.

Private ACTIVE tunnel transitions

The retained PalletHandoffSetOwner accepts beginActiveTunnel(effect), finishActiveTunnel(effect, expectedReceipt) and inspectActiveTunnel(effect) while its exact ACTIVE generation is UP. The effect binds the complete ACTIVE preparation and a creation plan or previous tunnel receipt to a durable intent digest. It includes the authenticated SmartVPN generation, complete plan digest, router namespace, exclusive interface name, /32 address, MTU, split routes and same-boot BOOTTIME deadline. Tunnel routes cannot capture owned transit or local workload prefixes, the control address or the hub endpoint.

Persist the intent before begin. Retain the actual SmartVPN client in the router namespace, prepare its tunnel, freshly inspect that generation, and derive the expected receipt from its actual interface index before finish. Native code observes the bounded ordered kernel transition, checks the full namespace twice and returns to its strict observer. It does not create the tunnel or infer the continued existence of the VPN owner from a saved receipt. Packet policy must be installed and verified before enabling VPN forwarding or raising workloads.

For deletion, stop forwarding while retaining the UP tunnel, lower the packet policies back to their base shape while the device still exists, persist and begin the deletion intent, then disconnect and join the VPN device owner before finish with null, and release both packet policies only after the generation's withdrawal. Each generation therefore carries a chain of policy pairs per role, all journaled and each exactly one revision above the pair it replaced: the base pair at the generation number (composed with no tunnel, so remote peers are deferred rather than permitted), the tunnel pair (the same projection composed with the retained TUN bound as the pallet_vpn endpoint), the lowered pair (the base shape again) — a tunnel and a lowered pair for every device the generation raised, each device's above the lowering of the one before — and between them one amended pair for every workload attached while the generation is UP (below). A generation raises at most 32 devices (maximumActiveTunnelCreations): every read of the chain walks the tunnel lane back to its first command and revalidates each row, so the lane bounds the cost of every tunnel-phase, amendment and cleanup step (64 rows at most on a lane this release admits). The journal refuses the next creation by name (exhausted, active.tunnel.exhausted) before it admits a command. The orchestrator refuses it before it asks for a credential to raise a device or dials, and refuses a replacement under another credential before it lowers the live device or dials; only the renewal's own credential request, which names the other identity, comes first. The native's own bound is larger (256 session generations per generation, maximumNativeTunnelGenerations, remembered so none is reused), and the journal's must never exceed it. The pass fails with that name, its owner withdraws the generation in order, and the next pass raises a fresh generation whose lane starts empty. The bound governs admission only: releases up to 35.2.0 bounded a lane by the native alone, so a lane they wrote holds up to 512 rows (maximumStoredTunnelCommands). It stays readable — inspected, recovered, lowered, cleaned up and withdrawn as any other, each chain read costing what it cost under that release — and only its next creation is refused. The compiler releases only the exact graph a transition names, so no revision is derived from a phase: every pair states its revision and the acknowledged pair it replaced, and activePacketChain proves the whole chain on every path that admits or acknowledges a pair above the base pair, and before a recovery re-applies or a cleanup releases the live pair. The tunnel pair is applied only after finishTunnel, because the compiler binds the device by the interface index the receipt carries, and always before forwarding is enabled; the lowered pair is applied before the delete command, so no enforced graph ever names an absent link. Linux omits deletion notifications for the static split routes; full inventory must prove their absence. ACTIVE withdrawal is rejected while a tunnel or unfinished transition remains. Unexpected notifications, owner loss or the absolute lease deadline retire ACTIVE and fence known interfaces DOWN. Recovery requires joining the VPN owner first and proving the whole namespace through a fresh ACTIVE owner. No receipt grants allocation reuse or proves packet drain.

The private protocol allows 2 MiB frames, including complete retained history and up to 1,024 split routes.

Both the durable tunnel journal and the application lifecycle composition now exist. PalletActiveStore persists every managed TUN command as its own row — beginTunnel admits one create or delete before the device owner is asked to act, acknowledgeTunnel records the native begin, and finishTunnel records the terminal fact — chained per ACTIVE intent so a delete consumes exactly the receipt its predecessor left behind, and a normal withdrawal is refused while a device is still retained or a transition unacknowledged. A recovery records DOWN only; it never claims a deletion, and the retained row stays as history. The creation row also carries the tunnel pair applied over the device (packets) and, once the tunnel is being taken down, the lowered pair (lowered); a cleanup row names the live pair — the last one the chain admitted — by its exact transitions and policy digests, and the packet owner's release is fenced on exactly those digests.

A workload may be attached while the generation is UP. The native admits the pure prepareWorkloadAttachment under a held generation, and an addWorkloadAttachment for a pair the generation does not hold yet is created by the generation itself, through its own observers, while it is UP with no tunnel transition pending, holds fewer than 64 pairs (prepared and attached together) and has no tunnel reaching into the new subnet. 64 is what the packet engine is measured to hold: Smartnftables 3.0 compiles a router of 64 workloads, each with local DNS, public TCP and UDP egress, four platform endpoints and up to three publications, inside its atomic batch with room to 89. The native registry, its router-namespace inventory and configuration (one link per present handoff and per pair, 96) and its admission fence (192 targets) are bounded to match; the compiler's prepare() stays the authority on what a given graph fits. The bound above 32 is verified by compilation and unit tests only; the root QEMU guest qualification has not yet run at 64. The attachment owner asks the held generation before it journals the ADD intent (PalletHandoffSetOwner.admitAttachment, handoffset.admitAttachment.notUp/capacity/tunnelOverlap); a refused ADD releases its preparation and fails by that name with no journal, because a pending one would fence every amendment of the generation while it is held. addWorkload judges the same rules again before the native is asked (handoffset.addWorkload.*), because a native refusal retires the owner. In the other order a tunnel is refused where it would reach an attached pair's subnet, by the store before the command is journaled and by the owner before the native is asked. The pair is DOWN and outside every policy at that point. Before it is raised the orchestrator amends the generation (PalletActiveOrchestrator.amend): beginAmendment compiles the live pair again with the attachment as a further source, in the live shape, one revision up, and appends the row { attachment, shape, revision, host, router } to the journal's ordered amendments log in the same commit that fences the current source as an extension of the one the generation carries — the same settled application, every carried workload exact, the new attachment complete; finishAmendment appends each role's acknowledgement after its apply. The intent, the native preparation and the UP record never change: activeSources(journal) is the intent's source extended by the log, and the tunnel pair, the lowered pair, every recompilation and the cleanup compose over it. Every read of the journal verifies each carried workload again against the completed application it was admitted under (activeSourceHistory): the application the generation was raised over for the workloads it was prepared over and every attachment amended in before its first in-place transition, and transition N's application for the attachments amended in after transition N. Amendments and transitions share the generation's one revision sequence, so that order is the journal's own and nothing is recorded beside it. Pallet 36.0.1 verified every amendment against the raise application, so a node that attached a workload after an in-place transition could no longer read its own journal: every pass failed, its tunnel lease ran out, and the node failed on every start. The same journal reads again under this release; such a node recovers without any manual step. A generation that already carries the attachment has nothing to amend, a lowered generation takes no attachment, and a node without a held UP generation amends nothing: the pair is dormant and the next generation is prepared over it. DNS needs nothing new — the completed attachment is a source of the next resolver revision exactly as after a dormant ADD. A recovery sends the native attachedWorkloads, the complete receipts of the restored pairs the preparation does not cover (required, [] when none); the native restores DOWN only when its registry is exactly the preparation's workloads plus those. A restored attachment it cannot list — an ADD whose receipt never became durable, or one whose removal is already intended — is removed first, in the fresh process and before its native adopts anything, and journaled exactly as a DEL journals it (PalletWorkloadAttachmentOwner.recoveryRemovals). A failed journal step, a foreign intent or a failed native removal refuses the recovery by its step (handoffset.recoverActive.beginRemoval/removalIntent/incompleteAttachment/finishRemoval).

A pair raised under a held generation stays UP through its withdrawal, and the withdrawal restores the router namespace's baseline configuration. The raise is therefore refused by name (active.raise.ipv6_baseline) unless that baseline keeps IPv6 disabled on the pair's router side — all/disable_ipv6 must be 1 before the generation is prepared — because IPv6 enabled again on an UP link brings addresses and routes no owned transition admits. The same rule applies to pairs raised before the generation is prepared (active.prepare.ipv6_baseline). The native router owner establishes and verifies all and default IPv6-disabled baselines immediately after creating its private namespace, before exposing it.

The router namespace is fail-closed between generations for the same reason: the withdrawal releases both packet policies while the raised pairs stay UP, so the IPv4 forwarding it restores must be off. The native router owner switches all and default forwarding off at creation, and a generation is prepared only when every forwarding switch it captures — all, default, lo and each router link — is off; otherwise the preparation is refused by name (active.prepare.forwarding_baseline). Activation turns forwarding on only after both packet policies are applied, and the withdrawal turns it off again before the policies are released, so workload traffic through the router stops from a generation's withdrawal until its successor is UP (the withdrawal window of a projection change the generation cannot move to in place, below) and is never forwarded unfiltered. For a containerd sandbox, the admitted native ADD disables IPv6 only on its newly created, identity-proven DOWN eth0 before addressing it; the exact pair DEL removes that link. Containerd runs CNI before creating the sandbox container, so CRI sandbox sysctls cannot provide this pre-CNI guarantee. The sandbox's namespace-wide defaults and the host's sysctls are never changed. A node with the mpls_router module loaded is not supported: every link registration emits an MPLS netconf record, which the attach vocabulary refuses.

The tunnel credential is the controller's. PalletControllerClient.resolveTunnel issues getRuntimeManagedVpnCredential { session, projection } over whatever transport the runtime session uses — the relay when the node is bound to one — for the exact projection the generation carries — the one it was raised from, or the one it last moved to in place (below) — and a reading of the qualified clock, and turns the answer into the session request the device owner authenticates with: the cluster hub's address and port, key material, authority and node ids, and a same-boot BOOTTIME deadline bounded to fifteen minutes or the credential's own expiry, whichever is sooner. The node dials its own cluster's hub, never a platform-wide one: the credential names the hub endpoint this node's signed projection selected, and the contract binds that id, its transport's protocol, its address and its port to that selection. The transport decides the dial form — a managed QUIC hub is a bare host:port, the one transport the contract states — so a credential naming any other is refused rather than dialled as if it were QUIC. Every refusal is named (controller.tunnel.notLive, disconnected, projectionAbsent, projectionChanged, boot, credential, transport); the key material is returned to the caller only and is never journaled or logged. The orchestrator asks for it once the generation is UP and DNS is serving; a refused credential or a refused authentication leaves the generation UP without a tunnel, reports the refusal on the barrier as tunnelFailure (code and site, nothing else) and on the application pass as activation.tunnel:<label>, and the next kept pass asks again. Once the command is admitted every later step is a journaled transition and fails the activation as such.

A tunnel lease lasts at most fifteen minutes and is moved in place, never by tearing the tunnel down: every admitted DNS lease renewal moves the window the credential lives in, so the next kept pass asks the controller for a fresh credential, and a pass that finds the lease deadline within three minutes of the qualified clock does the same (renewalMarginMs). Under the identity the session authenticated with (hub address and key, client key, authority and node) both bounds on the device move in place: first Pallet's own native authority deadline (PalletHandoffSetOwner.renewActiveTunnel, native renewActiveTunnelLease), which the native arms at the lease the device was created under and at which it fails the generation closed, then the session's lease with SmartVPN's renewManagedSession (PalletManagedVpnOwner.renewLease). The device, its packet policies and forwarding are untouched, and nothing is journaled: the deadline bounds this process's custody, and a successor fences the generation without it. A credential under another identity cannot renew a session, so the tunnel is lowered and a new one raised with it. A refused renewal leaves the lease in force, names the refusal on the barrier's tunnelFailure and is asked again on the next pass; a refused native renewal never asks the session, so the session's lease never outlives the native's.

The hub may move a live tunnel's split routes in place. A SmartVPN 2.5.0 hub that offers assignment updates sends the node its committed assignment, and the client moves the device's routes through its own netlink socket — every withdrawn route deleted before any new one is added — before Pallet learns the new assignment, so the native observer sees the route events first. It accepts them provisionally, and only on the held device: an exact split route (main table, static, link scope, unicast, no gateway, source or metric) to a private prefix of eight to thirty-two bits that is no default route, overlaps neither the device's address nor its remote, no other route and nothing the generation owns (transit, workload and attached subnets), and, for a deletion, one the device carries. Any other route event — another device, a default or public route, a foreign shape — fences the generation exactly as before. From the first change the routes are pending: they differ from the journaled receipt, every inventory read expects the provisional set — a read that raced a move is read again once the events the kernel already reported are taken, at most four times — and they must be named by a journaled revision within sixty seconds (active.observe.tunnel_routes_deadline otherwise; the bound is twice the thirty-second request budget one queued native command may hold the revision behind, and the BOOTTIME start never moves, so nothing extends it). The orchestrator follows the session's managed-session-assignment events (IPalletTunnelSessionOwner.watchRoutes): it journals the revision first — its number and the exact routes, address-ordered, in its own collection (pallet_active_tunnel_routes, one row per device, the adopted revision and at most one pending, each later than the last) — and then hands it to the native (reviseActiveTunnelRoutes, PalletHandoffSetOwner.reviseActiveTunnel), which re-dumps the namespace and adopts the revision only when the device carries exactly its routes with everything else unchanged; the adoption, with the BOOTTIME instant the routes first differed, is journaled after. A revision that differs, a stale one or a deadline that passes is a conflict, named on the barrier's tunnelFailure, and fences the generation as any other. Device steps — raise, lowering, renewal and revision — run one at a time. A replayed revision is adopted once; a revision journaled before a restart stays history, because the device ends with the process and the next process fences the generation. A device's deletion consumes the receipt of its last adopted revision. The router policy keeps reaching the tunnel's peers by the prefixes its projection states: routes ahead of the projection carry nothing until the generation moves to that projection.

A projection change the generation can follow moves it in place (PalletActiveOrchestrator.transition). The application owner decides it before any dormant effect: the next application intent is the exact successor of the one the held generation carries, within its epoch, over the same uplink observation; it keeps the dormant handoff set and the receipt its predecessor completed with, the selected handoff and the host routes the generation installed; and its projection continues the node's managed VPN membership (continuesRuntimeManagedVpnMembership, @serve.zone/interfaces 32.42.0: the same handoff and projection lineage, a later generation, and a protected authority that is the same or exactly its next step that only adds — a node joining the egress, a hub joining the platform endpoints; workload endpoints, egress, private networks and the router selection may move). Such an intent is completed without the native — its receipt is its predecessor's — and the held generation is not withdrawn. beginTransition then admits, by the same rule again, the pair that moves it: the live pair compiled again over the successor application, in the live shape, with every attachment the live pair carries that the successor's source still admits, one revision up, publications refused for capacity by name exactly as a fresh generation refuses them, in the commit that fences the current source exactly; finishTransition journals each role's acknowledgement after its apply. Each transition is one row of its own collection (pallet_active_transitions, numbered per generation, at most 64 per generation), and the chain proves it like an amendment. The intent, its native preparation, the UP record, the device, its session and its lease stay as they were; the packet policies' protected authority, workload endpoints, egress and the router's tunnel peers move. The application a generation carries is its last transition's once both roles acknowledged it, its own before any; while a transition is half applied it carries neither, and only a fence settles it. Renewals ask for the credential of the projection it carries. A positive protection receipt for an application is appended only while a held UP generation carries exactly that application — both roles of its last transition acknowledged — and its row is fenced in the receipt's commit, so the receipt of a new protected authority follows the policies that enforce it; the receipt's native barrier stays the guard intent, which a new authority always moves. Anything else — a successor that ends the membership (a protection step that removes or changes anything, a skipped step, another handoff), another target or receipt, other host routes, a rebound uplink, the transition bound, a refused or failed transition — withdraws the generation and raises a fresh one, as every projection change did before. A process that ends between a transition's roles leaves the live pair to the cleanup the next process runs after it fences the generation.

A transition moves past an attachment the successor no longer carries instead of refusing. The current source (PalletWorkloadAttachmentStore.fenceActiveSources) names every retained attachment of the epoch it does not carry, and why: removing once its DEL is intended, absent once its absence is proven, and unadmitted while it is live but the application no longer reserves and projects its lease. A carried attachment that is released any of these ways is left out of the successor pair, and the transition row names it in detached (the sorted attachment ids; omitted when the transition detaches nothing, so such a row keeps the bytes an earlier release wrote); the chain applies it, so every later pair composes over the attachments the live pair still carries, and an amendment extends the live pair rather than everything the generation ever carried. A live, unadmitted attachment is then detached (PalletWorkloadAttachmentOwner. detachAttachments): its run reads as of an ended epoch, so its execution owner stops it — and fails it network-epoch-ended while its controller still asks it to run — no name is served for it, and its DEL takes the pair out of the generation; the policy that no longer carries it already denies its flows. A fresh generation is prepared only over the whole epoch. Before carrying pairs into it, activation reconciles released attachments while no generation is held, at most eight per pass. An existing removal keeps its original operation and ACTIVE binding. For an unadmitted complete pair, the current settled application's lease withdrawal authorizes a durable released-attachment network removal under the application fence. Its digest binds that completed application's reference, whose retained signed source proves withdrawal on every replay. It releases an ended creation binding without needing another CRI DEL or an assignment lane a failed run already freed. Runtime and storage authority do not change. The native removes only the exact journaled pair and proves absence, including when its sandbox vanished. Absence is persisted before activation; neither elapsed time nor attempts erase a row. The release logs removing, absence proven, or a bounded failing-step label with activation waits. Unfinished rows retain their intent for the next network pass or owner restart, and the run continues to read as of an ended epoch after this removal. While any released attachment lacks durable absence, activation refuses by name (active.activate.releasedAttachment) until its DEL proved it absent. Pallet 36.0.2 refused the whole transition instead (attachment.activeSource.unadmitted): every stop of an attached workload — a spec change, a deletion, a rollout — withdrew the generation, and with it the node's tunnel, for about two and a half minutes, and an attachment removed under a held generation broke its next amendment or transition the same way. Cloudly 33.13.2 keeps a stopping run's lease reserved until the node's stop receipt, so a stop normally reaches this node as a DEL first and a retired lease second; either order moves in place.

A DEL takes a carried attachment out of the live pair before its pair is deleted (PalletActiveOrchestrator.detach, called by PalletWorkloadAttachmentOwner once the removal is intended and before the native removal): beginDetachment compiles the live pair again without it, in the live shape and over the live application, one revision up, and appends it to the amendments log as a detaching row { attachment, shape, revision, detaches: true, host, router } — attachment exactly as the pair carried it — in a commit that fences the attachment's own removal intent (attachment.fenceRemoval.notIntended otherwise); finishAmendment acknowledges each role. A detachment is admitted after the generation's last device was lowered too, in the base shape, so a DEL never waits for the next device; no other amendment and no transition follows a lowering. activeSources leaves a detached attachment out, activeHistoricalWorkloads keeps every workload the generation ever held for its recovery, and activeSourceHistory verifies a detached attachment where it was admitted. The packet engine binds every link a policy names and refuses one that names a link the kernel no longer holds (ENODEV, which Smartnftables 4.3.0 reports as UNAVAILABLE), so a pair must never outlive the links it binds. Pallet 36.1.5 and earlier deleted the pair under the live pair, and every pair compiled over it later was refused: a node stop that followed a DEL failed active.stop.loweredPackets < Error < ManagedNftablesError:UNAVAILABLE, joined the device owner with the device's deletion never begun, and the handoff set retired on that unarmed deletion (active.observe.unexpected_event.new_link_flags: a link closed while UP is reported with IFF_UP in its change mask before its RTM_DELLINK, the first event of an armed deletion); a pass's withdrawal and an amendment after such a DEL failed the same way (lab 2026-10-04, N7 and run 3's N4b). A journal that carries a detaching row is refused by Pallet 36.1.5 and earlier: once one exists, downgrading the node in place to 36.1.5 or earlier is not supported.

A lease that lapses anyway ends the native authority at the same instant: the native fails closed, the generation is fenced and the node fails owned, and its supervisor starts a fresh process, which fences the generation by the proof that its epoch ended and raises a new one with a fresh credential. A session whose transport to the hub is lost stops instead: SmartVPN ends its packet I/O and reports managed-session-state stopped, keeping the device for an ordered teardown. The barrier stops holding at that moment — holds requires the session's packet I/O to be live — so attached workload links demote on their next CHECK and admission refuses, and the next node scan runs a network pass at once. That pass lowers the dead tunnel through the journal exactly as a withdrawal does (the lowered packet pair, the delete command, the device released) and closes its owner, then asks for a fresh credential and raises a new tunnel over the same generation; a refused credential leaves the generation UP without a tunnel, named, until a later pass succeeds. A retired session — its device lost behind the owner — is a loss of the owner, not a stop, and fails the activation as before.

PalletNetworkApplicationOwner takes an optional activation option and, when it is present, drives one ACTIVE generation after each completed dormant application: preparation, both packet policies applied and journaled, native UP, local DNS proven to be serving, then the tunnel commands above and finally forwarding. A native preparation whose intent commit is rejected (the source superseded while the policies compiled, active.source.fence) is discarded before the rejection is reported: the handoff set admits nothing else under a held preparation, and no journal row would name it. Without that option the owner stays exactly as dormant as before and the DOWN wire contract is unchanged. PalletNetworkOwner.barrier() exposes the result as one of absent, prepared, up or forwarding, with the retained device's journal reference when there is one. Only forwarding sets holds, and only holds may gate workload UP; a node with no managed VPN configured settles at up and its barrier deliberately does not hold, with one exception: the single host. A node bound to a local controller (Onebox) never dials a tunnel, whether or not a managed VPN is configured — the released bundle configures one on every node, and the persisted binding decides — and a generation reaches forwarding without a tunnel when three facts hold: its scope is that local onebox controller, every endpoint of the signed projection it was raised from is placed on this node (the contract already refuses a remote endpoint in a onebox projection; Pallet proves it again from the journal row and fails closed), and this process raised the generation itself. There is no remote peer to reach, so the base packet policies already carry every rule the generation needs. A barrier that does not hold names the missing fact as singleHostGap (controllerNotLocal or remoteEndpoint). Execution admission then rests on the single-host evidence instead of a tunnel receipt: the evidence names the projection, the admitting transaction re-proves the fact from the journal row and fences it, and a refusal is admission.barrier.singleHostUnproven. A tunnel command on such a generation is a corruption and refuses. Every other node — bound to Cloudly directly or through the relay — keeps the mandatory tunnel barrier unchanged. The barrier is always recomputed from the journal and the live owners, never read back from a stored pass. On loss this composition fences and withdraws first and then reports a barrier that does not hold; it never re-activates the lost generation by itself, because re-activation requires explicit newer state from the control plane.

Store reads under load

Every store read hands each row to its collection's validator, and the same rows — application intents and results carrying the complete signed projection, signing revisions, transitions, retained workloads — are read on every network pass, DNS refresh and admission. With a 64 KiB projection Pallet 36.0.2 spent a whole core re-validating them (lab 2026-10-03). Each validator now runs once per exact document per process (memoizedDocumentAssertion, ts/storage/validation.ts): the verdict is remembered by a SHA-256 over the document's JSON, which states every value it admits exactly — a non-finite number, negative zero, undefined, a Date or other class, a proxy makes a document unkeyed, and an unkeyed document is validated every time. Validators are pure functions of the document, only acceptance is remembered, and a document that differs in one byte is validated in full. The asynchronous proofs of immutable rows — the retained projection's signature and validation, signing revision, workload and handoff lease digests, guard host grants — run once per exact input the same way (verifyOnce); an application intent's whole retained lease set is proven once per exact set. The retained workload rows are read once per change: the projection store keeps the rows an application-source read last found, keyed by the stored revision of the projection row that read took in the same transaction. Only an admission writes the workload rows, and the same commit always writes the projection row and so moves its revision; rows and key therefore come from one snapshot, every application-source read of every pass shares them, and a transaction whose snapshot predates an admission keys what it read by the old revision, which no later reader holds (test.projectionworkloadcache.node.ts). The memo holds at most 8,192 entries and lives in the process only. An identical projection push, which a controller repeats on every delivery round, is a replay: the store compares the exact envelope it admitted under the signer it still trusts, writes and fences nothing, and the controller client schedules no network pass for it. test.storebudget.node.ts bounds all of it over a 64 KiB projection with 122 retained workloads: after the first pass, a pass validates at most 16 documents, reads no workload row and makes at most 300 row reads and proof lookups in all (measured 249; 36.0.2 made about 2,700, most of them the retained workloads read a dozen times); eight identical pushes validate none, read no workload row and schedule no pass; and after a changed projection the next pass reads the retained rows once and the pass after it not at all.

An idle node does constant work per second, and a controller's re-delivery writes nothing (lab 2026-10-03, Pallet 36.1.0: about half a core with two idle workloads, and a new attached assignment never admitted inside the relay's 8 s timeout while three ran). The execution driver keeps its --executor-management child between calls and ends it only after a call that failed and on close, so the scan's runtime inspections and clock readings are IPC round trips instead of a process spawn each. A run whose reconcile found it waiting only for its next readiness probe is settled until that probe, and at most palletSettledRecheckMs (5 s): while its durable state is exactly the one it was settled on — and, for an attached run, its pair is still of the current network epoch — the scan answers it readiness-pending from those reads and one clock reading, without inspecting the runtime or opening a transaction. A run whose network preparation was refused (IPalletExecutionNetworkOwner.prepareExecution, for example attachment.preparation.endpoint.check) is answered from that refusal — the same thrown refusal, or the same wait line for a secret run — while its durable state is unchanged and no network pass ended since (networkPasses), and at most palletRefusedPreparationRecheckMs (5 s): the preparation is fenced against the completed application the last pass selected, so nothing else a scan sees can change it, and the scan asks the controller for no registry credential or secret material meanwhile (lab 2026-10-04 RUN 4, P2: two refused runs held an exclusive node with no workload at about 40 % of a core). Assignment histories and configs are proved once per exact content (verifyOnce), and the delivery pass reads no terminal outbox of a first generation, which has none. An assignment re-delivery that is a replay or a historical answer is neither fenced nor gated: it writes nothing, so it no longer competes with every other fenced commit of the node for the identity row, which is what made a new admission lose its commit on nearly every attempt. The execution gate composes its barrier evidence once per admission and every retry of the admitting transaction only fences it again. test.admissionbudget.node.ts (eight re-deliveries take no identity fence and validate nothing; a new admission commits on its first attempt while every running assignment is re-delivered back to back), test.admissionguardretry.node.ts (a retried attempt composes nothing) and test.idlebudget.node.ts (one child for seven driver calls; ten scans of a settled run spawn nothing, open no transaction and make no runtime call; a delivery pass reads no first-generation outbox; ten scans of a run refused at its preparation ask for no credential and open no transaction) bound it.

Nor does an idle node's work grow with its history (lab 2026-10-05 RUN 6, P2: an exclusive node with no workload at all held 30–35 % of a core, against 4.8 % on a node that had run nothing yet). Every second the scan, the delivery pass and the log reconcile each read, decoded and proved every assignment the node ever held, and the database read each of those rows for them, though a retired assignment — removed, its terminal receipts and removed observation acknowledged, of which the controller keeps no record — never changes again. The assignment store keeps the ids of the assignments it has not proved retired, built from the whole inventory by the first scan of a process; listCurrent and listReportable read only those rows, by id, and an id proved retired leaves for good. An admission that was accepted, or failed and may have committed all the same, adds its id once its transaction settled, and an id with no row is dropped by a later scan that no admission overlapped. list still reads the whole history. The log owner reads the whole inventory on its first pass, which reopens what an earlier process left of a retired assignment, and afterwards lists only what is not retired plus the state of each capture it still holds. test.idlehistory.node.ts (an idle tick over three and over six retired assignments reads no assignment, outbox or slot row; a new admission is live at once) and test.logowner.node.ts (an ended capture whose assignment retired while it waited is still delivered) bound it.

A network pass whose application source is superseded while it runs — a projection, signer or allocation admitted between its read and its commit — is refused by name (projection.fenceApplicationSource.superseded) and changed nothing; the next scan runs a pass at once. The guard's side of a pass loses the same way: a projection admitted between the guard pass's read of its bootstrap source and its proposal of the expansion it prepared against it is refused as guard.propose.superseded; native preparation is inert and the proposal did not commit, so the pass ends without retiring the network owner, and the next one expands the guard to the newer source. Any other refusal of a guard pass still retires the owner. The node writes Pallet network pass deferred: <site>. once and then only with its count, instead of a failure line (test.networkpassrace.node.ts, test.networkguardrace.node.ts).

Durable application intent and results

The private identity runtime exposes runNetworkApplications() over its existing SmartData/NoSQLDB owner. readSource(scope) returns the current verified signed projection and complete retained lease history. Native preparation happens outside the database. begin(scope, fingerprint, prepared, guard) then persists that exact whole-set target, its previous receipt and the authenticated source, writing the common trust/allocation fences and the real identity guard in the same transaction. Only the current reserved handoff is desired present; every other retained lease remains explicitly absent. Omitted history and revived absent members are rejected. Ordinary successors remain within the same boot and native namespace.

The journal keeps one pending transition and immutable intent/result records. Exact retries join existing intent; a competing proposal cannot replace it. finish(scope, intentReference, nativeReceipt) records only the result matching that persisted target. It does not require new admission authority: signing-key, projection and credential changes must not discard an already-owned native result. Historical signatures remain verifiable through retained signing revisions. Lost completion acknowledgments replay exactly without advancing or rolling back the current head. No journal method deletes history or frees quarantined allocations.

Keep begin, native reconciliation and finish inside one admitted runtime callback so database shutdown joins the entire operation. Retained facade methods expire when their callback returns. The persistence capture budget covers both transition sides and complete signed source; native frames retain their independent 2 MiB limit. This store does not start a native owner, issue nativeBarrier, report protection, clear DNS withdrawals or authorize workloads. The private application owner supplies node composition and a separate durable unavailable-report outbox.

beginEpoch(scope, fingerprint, prepared, evidence, guard, workloadProofs) records a qualified same-boot absence or different-boot fence into a fresh namespace. The caller must hold the fresh private native owner continuously through its audit, journal commit and reconciliation. The store resolves the evidence's opaque journal reference to the actual persisted pending intent, or applied intent when none is pending, and binds its complete target and exact known receipt. A single transaction retains the old applied/pending references in an immutable epoch record, fences current signed authority and identity, and publishes the next intent. Evidence from a different target, previous receipt, host, boot or namespace is rejected.

The first intent in an epoch has previousEpoch and no native predecessor: reconcile it with previous = null. Its generation follows the old pending or applied head. The still-reserved logical handoff may be realized DOWN again; every retained lease and quarantine survives. Old unfinished intents stay in history without a fabricated result and reject late completion after the epoch fence. Previously recorded results still permit exact acknowledgment replay. An intent has one shape: it names the epoch it opens or null, and a record that does not is refused rather than read. Repeated failures before a fresh intent completes continue the same history. Different-boot recovery requires the distinct native previousBootTerminated proof, persisted as new-boot-host-fence; it cannot use a same-boot absence receipt. Reused numeric namespace identities do not make two kernels the same epoch.

Dormant workload attachment journals

The private ts/network/attachment/ composition binds one admitted pending Run to its immutable workload lease, completed application and exact container ID. Native preparation opens the sandbox namespace once and retains its descriptor. The ADD journal commits before any veth mutation; its immutable identity cannot be rebound after cleanup. Ambiguous admission acknowledgements are resolved on the same application transaction lane before an unadmitted descriptor is discarded. Partial creation uses the recorded removal intent and exact native absence check. The sandbox accepts containerd's internal loopback baseline. Workload links stay DOWN until this node's activation barrier holds; an ADD waits for it, bounded.

Application and attachment operations share one queue. Handoff retirement must fence every unfinished attachment in its transaction. A replacement router waits for distinct historical workload evidence: stable absence in the exact retained sandbox on the same boot, or termination of a different actual kernel boot. A missing or substituted sandbox path is not same-boot absence. Previous interface indices belong to their original namespaces and are never host capabilities.

An application epoch record has one shape and always binds a complete sorted workload proof set. Immutable coverage records commit with that epoch and freeze the old journals. Coverage permits historical DEL to acknowledge the proved absence, while preserving the original unfinished journal, all allocations and quarantine. It does not create an ordinary DEL receipt or authorize another ADD. Quiescence stops preparations; already admitted ADD/DEL operations drain before the shared native owner closes.

Production execution admits network: 'attached' only after the assignment's exact projected lease is prepared with the pending Run operation. The CNI owner then proves and journals its attachment before returning a usable result. An isolated assignment prepares a separate durable CNI intent with no lease or application reference. A secret reference is resolved by this node: its material is staged and published as this attempt's own mount before the container that receives it, as "Staged secret material mounts" states. A storage reference is granted from this node's claim ledger in the transaction that begins the run, as "Local storage claims" states.

Private CNI transport

The node runtime owns a root-private CNI broker at /run/serve.zone/pallet/cni.sock. The installer supplies the trusted /run/serve.zone parent; the broker verifies its ancestry and provisions the 0700 leaf and 0600 socket. The native executable enters CNI mode through --cni or the installed basename pallet-cni. The client checks path ownership, socket identity and the kernel-reported server UID before sending one bounded frame. It never stores journals, allocates addresses or opens sandbox namespaces.

Transport cancellation does not cancel an admitted attachment operation. Detached handlers retain broker capacity until their native effect and journal write finish. An invocation the plugin stopped waiting for while the owner still held it — its connection closed, by the plugin or by a broker that is closing, or the broker's own 30 s ran out — is written as Pallet workload attachment abandoned: run <run digest> <command>: <transport.closed|broker.deadline>; the owner continues., and what the owner answered after as Pallet workload attachment answered after abandonment: run <run digest> <command>: <status or label>., each once, then with its count while it repeats. Pallet 36.3.6 and earlier wrote neither, so an ADD that bound its run after the runtime had given up on it left nothing in the node's journal (Grasberg 2026-10-10). Node shutdown joins CRI, including synchronous rollback DEL, before closing the broker; the broker then joins all admitted handlers before its borrowed owners can close. Replaced socket paths are never unlinked during cleanup.

The fixed profile uses CNI 1.1.0, the single pallet-workloads configuration and containerd's use_internal_loopback=true. VERSION reports protocol support. DEL can acknowledge exact recorded absence with empty stdout.

A refusal of the network owner names its check. The attachment owner refuses by site (attachment.live.<state>, attachment.cleanup.<state>, attachment.reconcile.* such as attachment.reconcile.unprepared, attachment.isolated.*, attachment.epochEnded); the broker answers { error: 'rejected', refusal: 'refused.<code>:<site>[:<native site>]' } and writes Pallet workload attachment refused: run <run digest> <command>: <label>. to the node's stderr, once, then with its count while it repeats; the plugin answers the CNI error Pallet network owner refused the attachment; cleanup may still be required with the label as its details, which containerd states after the message, and writes it on its diagnostic line (PALLET_FAULT site=cni.refused code=<label>). A transport failure, or a broker answer without a label, is still Pallet network owner unavailable; cleanup may still be required. An ADD that arrives while the attachment owner waits for its network epoch to be proven (pending-epoch: it started over an attachment of an earlier epoch that is not yet proven gone, which the application's next pass settles) waits for the owner to become ready, within palletAttachmentWaitMs (20 s from its arrival, shared with its wait for the barrier, inside the broker's 30 s request timeout), and is then judged as any ADD; a close while it waits refuses it at once (attachment.reconcile.closing). An ADD journals its pair only within that same wait: one that reaches the application's lane after it, behind a slow lane or an earlier ADD of its run, or would commit its journal after it, is refused attachment.reconcile.expired, journals nothing and discards a native preparation it already made. Its caller has given up on it by then, and a pair journaled later bound the run for good to a sandbox the runtime had already torn down (Grasberg 2026-10-10). Every other refusal — a CHECK in pending-epoch, a closed or failed owner, 64 operations in flight (attachment.reconcile.busy), a missing preparation — is answered at once (lab 2026-10-05 RUN 6, O1: one run's sandbox was refused three times in six seconds as "network owner unavailable", with nothing to say what had refused it). The broker answers its own refusals the same way, before the attachment owner sees the invocation: refused.unavailable:broker.busy while 32 earlier invocations, abandoned ones included, are still held, refused.unavailable:broker.notListening while it closes, and refused.unavailable:broker.identityStopped once the node's identity runtime stopped; a STATUS request is then closed unanswered. Pallet 36.3.6 and earlier closed such a connection without an answer, which the plugin reports as "network owner unavailable". A connection beyond 32 open at once is still closed by the socket server before the broker sees it. A quiescing owner still admits ADD and DEL for runs already admitted, as before.

For an isolated assignment, the pending Run transaction durably binds its assignment, operation, boot and fixed CNI configuration. ADD opens the exact containerd sandbox namespace and proves it is distinct from the host and router namespaces. Containerd 2.3 requires a real eth0 with an address even for an isolated pod, so the native owner creates a namespace-local Linux dummy device with 192.0.2.1/32. It has no peer, uplink, default route or host grant; IPv6 address generation is disabled and every address, neighbour and route in both families is checked. The kernel's own ff00::/8 local-table multicast route is allowed only when it is bound to this peerless dummy, with no IPv6 address, gateway, RA or unicast route. The first ADD commits the exact container ID, namespace identity and double-proven initial absence in SmartData before NEWLINK. The completed ADD then records the dummy index, deterministic MAC and double native proof. If a process stops between NEWLINK, alias, address and UP, replay may remove only the exact DOWN partial dummy under that durable intent, prove its absence twice and recreate it. Ambiguous or foreign state remains quarantined. CHECK requires a completed journal and reproves its exact endpoint. The CNI result reports the real dummy and address, with no routes or DNS grants. DEL verifies and removes that exact device before retiring the journal; replayed DEL is idempotent. A DEL for a sandbox whose ADD never journaled its intent is answered absent without a native step: the dummy is created only under that intent, so nothing of Pallet's is in the sandbox, and the runtime's teardown of a refused sandbox completes (Pallet 36.2.0 refused it as attachment.isolated.delUnjournaled, and containerd kept the sandbox NOTREADY). No workload lease, veth or activation grant is made for this mode.

The isolated reconcile runs on the handoff-set coordinator, which owns no state in the sandbox, and is admitted while an ACTIVE generation is held: Pallet 36.2.0 refused it there (the coordinator's ACTIVE command filter did not list it), so every isolated run of a node whose generation was up had its sandbox refused, and the owner over the coordinator retired on that refusal, failing the node runtime (owner_failed; lab 2026-10-06 RUN 9, D11). A command refused under a held generation now writes handoffset.active.withdrawal_required and the command's site as fault lines. What the native refuses of an isolated sandbox itself — its path gone, a namespace that is not a fresh sandbox, or anything its inventory or its dummy's lifecycle finds — is answered handoff_sandbox_refused: the set is not marked failed, the coordinator and any generation stay, and PalletHandoffSetOwner.reconcileIsolatedSandbox rejects that call alone with conflict at handoffset.isolatedSandbox.refused, carrying the native fault lines; the run's sandbox setup fails and the execution owner handles it as any refused sandbox. An attached workload is answered the same way when its sandbox's path is gone as its preparation captures it — the runtime removed the sandbox while its ADD waited (workload.prepare.sandbox_gone): PalletHandoffSetOwner.prepareWorkload rejects that ADD alone with conflict at handoffset.prepareWorkload.sandboxRefused, no ADD bound the run's attachment, and the run is driven again as for any refused sandbox. Pallet 36.3.6 and earlier answered it unavailable, which ended the coordinator and every attachment with it (Grasberg 2026-10-10). A failure of the coordinator itself (its watches, entering or leaving the namespace, the command's deadline) still retires the owner. Native CRI evidence must report only 192.0.2.1 for an isolated running sandbox. Host-side TCP and HTTP readiness probes are refused before effects because they would target host services.

An ADD whose pair is admitted, journaled and complete while the activation barrier does not hold waits for the barrier: it leaves the application's lane — the generation that raises the barrier needs it — and is woken by the ACTIVE orchestrator's barrier signal (PalletActiveOrchestrator.watchBarrier, the one source; the attachment owner reads the barrier again on each wake), by its own owner's state changing, or by its close. Its two waits, for a pending network epoch and for the barrier, share one bound, palletAttachmentWaitMs (20 s) from the ADD's arrival, inside the CNI broker's request timeout and the plugin's own 30 s deadline; every exit clears the wait's timer and listeners. On a node without activation there is no barrier to wait for. Once the barrier holds, the held generation's policies are amended to carry the pair when they do not yet, and the pair is raised and proven in the same step; a barrier still not holding at the deadline answers the pair DOWN with an activation error and the label barrier.unheld:attachment.reconcile.barrierWait, and the runtime tears the sandbox down. That sandbox was the run's one binding, so the run then fails instead of being driven again (see "Private CRI execution"). A refused amendment or raise leaves the attachment admitted, journaled and DOWN and names its step on the pass (amend.<code>:<site>, raise.<code>:<site>); it never fails the ADD. The plugin writes that label only to its own stderr, which containerd discards, so the CNI broker writes it to the node's stderr as well: Pallet workload attachment down: run <run digest> <ADD|CHECK>: <label>., once, then with its count while it repeats. A successful ADD returns a CNI result carrying the interface, its address, the node resolver and the one route the raise installed: 0.0.0.0/0 via the router side of the pair (static, priority 100). Only state this owner installed is reported, and the advertised route is proven equal to the route found in the pod namespace. CHECK reports the raised attachment while the barrier still holds and demotes to the same DOWN activation error the moment it is lost, without touching kernel state. A repeated ADD on an already raised attachment is idempotent and never destroys it. STATUS reports authenticated broker and network owner availability to containerd, independently of any attached activation barrier; it cannot grant a workload network. An attached ADD or CHECK without the exact live barrier still reports DOWN, while an isolated ADD uses only its own pending Run preparation. GC cannot infer cleanup authority from a runtime-provided attachment list and returns an error until an admitted removal or historical proof can be resolved.

Application owner and node lifecycle

PalletNetworkApplicationOwner holds one private handoff-set owner for the node's router namespace. Each bounded pass keeps journal inspection, completion of a retained current-epoch intent, current source selection, epoch qualification, admission, native effects and completion inside one identity-runtime callback. It derives complete membership from the authenticated store; RPC callers cannot provide a prepared target, native receipt or absence proof. A stale source or identity rejects new admission. Completion of an already admitted effect survives credential changes, controller disconnect and quiescence.

The private PalletNetworkOwner composes this application owner with the guard and protection outbox. It restores the existing offline guard before starting the application or controller. While connected, the node reconciles the network with the actual physical session's private identity guard before scanning assignments whenever a network pass is due: the first scan, every scan after new network input from the controller (a new connection, an admitted signing authority, projection or DNS lease renewal), after a rejected pass, as soon as a managed VPN tunnel stopped on its own, and otherwise at most networkRefreshMs apart (default 30 s; the assignment scan itself runs every pollIntervalMs, default 1 s). A full pass re-proves the guard with a fresh engine process and re-reads and re-verifies the application and its signed source, so it does not run on every scan. The standing DNS refresh runs every 30 s for the same reason; every change DNS depends on (an application pass, a workload ADD or DEL) refreshes it at once, and the resolver child enforces its lease's BOOTTIME expiry itself. A refresh whose DNS store read or intent the database refuses changed nothing native: that pass ends alone, with the resolver child joined, and the owner stays ready. DNS is down until the next refresh serves again, within 30 s; the node writes Pallet DNS refresh refused: <chain>., once, then with its count while it repeats. A DEL whose refresh is refused so still removes its pair, and an ADD's refusal answers that ADD alone. Pallet 36.3.6 and earlier retired DNS and with it the node on any such refusal, and a DEL that met one closed DNS for good (Grasberg 2026-10-10). The protection delivery loop decides that nothing is pending from its lane row alone. Without current authority it waits; there is no offline network admission path. The status snapshot reports the application reference and admission rejection without exposing source records. Native ownership loss or an unrecordable native result fails the node with ownership retained for cleanup.

Shutdown stops admission, joins any pending begin/effect/finish callback, then joins the native handoff owner before closing the namespace and database. An unconfirmed native join keeps those resources owned; failed cleanup cannot report a clean stop.

A normal stop (SIGTERM or SIGINT, or stdin EOF for runtime-serve) withdraws an ACTIVE generation and journals it DOWN in either forwarding mode before a clean router namespace handover. The native refuses to withdraw a generation that still holds a device, and a device owner joined without a begun deletion takes its device away behind the native, which retires on that unarmed deletion; either way the handoff set's release stays unconfirmed. So the orchestrator's close refuses every new command and joins the ones it admitted before; then it lowers a held tunnel exactly as a withdrawal lowers it — forwarding stopped, an amendment or in-place transition pair whose command ended short of its acknowledgement applied again on its unacknowledged roles and acknowledged, the lowered packet pair applied, the deletion journaled and begun, the device owner released, the deletion finished — and only then joins the device owner. It then begins and confirms the native generation withdrawal, journals DOWN, releases the packet policies and any Docker contribution, and discards the preparation. The native's closeHandoffSet (handOverHandoffSet over a kept router namespace, which deletes no pair: "A stop hands the set over") then closes the dormant set or hands its retained pairs to the successor without deleting them. The process ends with stopped and exit code 0. The journal records the generation DOWN with its tunnel lane settled on the deletion and its policy cleanup complete. The successor can adopt the clean namespace and raise a fresh generation; only an unfinished failed or crashed lifetime needs the ended-epoch fence (the retained-generation paragraph under "Standalone control build"). The application owner joins every workload ADD and DEL before that close, since the lowering runs outside the lane an ADD's amendment and raise run on; so no other journal or packet work interleaves with the lowering. A stop over a generation with no tunnel withdraws the generation without any device command. A lowering or withdrawal that fails is an owner failure: the stop still joins every owner, fails with the failed step as its cause and ends owner_failed, and the successor fences the generation as after a crash. The step is named active.stop.<step> (settleTunnel, settleLive, stopForwarding, loweredPackets, beginDeletion, nativeBeginDeletion, acknowledgeDeletion, releaseDevice, nativeFinishDeletion, finishDeletion, and for a withdrawal beginWithdrawal, nativeWithdraw, finishWithdrawal, releaseDocker, releasePackets, discard, closeSession), with the step's own failure below it, and the network owner's close carries what refused it, so the node's Pallet runtime-serve failed: … line reads down to the refusing step.

The lowered pair is compiled only over an acknowledged live pair. A stop completes a pair left short of that (settleLive) without checking it against the current source — the controller may have removed an attachment it carries since the admitting commit fenced it — so it does so only once forwarding stopped, moves no DOCKER-USER contribution (Docker's FORWARD policy keeps admitting only the host table the contribution last followed), and the lowered pair replaces it at once. A pass's lowering refuses such a pair (active.prepareSuccessor.unapplied). While the attachment the pair extends is still in the source, the amendment or transition command that admitted it replays it against the current source. Once the controller has removed that attachment, no pass can complete the pair: passes stay refused until the node stops — whose settleLive completes it — or the next process fences it, so the recovery is a node restart. This needs an amendment whose apply was refused on one role and whose attachment was then removed.

A normal stop of a docker-shared node instead withdraws the generation it holds — handoff down, the Docker forwarding contribution released, then the host table — hands the dormant set and every attached workload pair over intact to the next process over the kept router namespace ("A stop hands the set over"), and never writes the host's IPv4 forwarding, at stop or at exit; the process ends with stopped and exit code 0 and no contribution rule outlives it ("Host forwarding modes"). The stop carries no attachment past the generation it withdrew: the application owner has joined and closed its attachment owner before the orchestrator's close, and the next process of the epoch carries the attachments at its start, before any pass ("Attachments outlive their generation").

Without the activation option all handoffs remain DOWN. With it the application owner opens the physical DHCP uplink observation before the orchestrator and raises one ACTIVE generation over it on the first pass that has an application; a generation that is UP for the same application over the held observation is kept on later passes, an observation that no longer proves itself withdraws the generation bound to it and is reopened on the next pass, and a failed activation is fenced and named on the pass (activationFailure). The same label is written once to stderr, which Spark carries into the node journal, as Pallet network activation refused: <label>; a packet engine refusal adds the engine request and hop (preparePolicy, host or router) and the engine's code and bounded refusal text, and any other refusal that has a cause adds that cause's chain (< <chain>, class names, codes and sites, cut in the middle so the line stays within the 512 characters Spark forwards): a packet operation the label can only call unavailable still names the engine request's code beneath it. A refusal repeated on every pass is written again with its count one minute after it was last written, then at doubling intervals up to ten minutes, and a refusal that returns after activation held is written at once. Every network store conflict names the check that refused — the ACTIVE family as active.<owner>.<check>, every other store as <store>.<method>.<check> — and the type of requireNetwork and of PalletNetworkStoreError makes the site mandatory for conflict, so the label always says which one. Without a managed VPN the generation settles at UP and the barrier deliberately does not hold, so attached workload links stay DOWN — except on a single host bound to a local controller, whose barrier holds as described above whether or not a managed VPN is configured. Lease release and quarantine remain with Cloudly's allocation owner; Pallet resolves only an exact assignment-bound projected lease and cannot create or release allocation authority.

Durable protection and withdrawal reports

Every connected network pass freshly recovers the actual allocation-pool guard, completes the dormant application journal, then attempts a positive receipt. Only a real guard owner that has reconciled, inspected enforcement, detached its PERSIST policy and joined its native child can supply that private capability. The returned bootstrap JSON and a durable guard reference alone cannot mint it. The capability expires when the owner closes; composition keeps it through the outbox transaction. RPC never accepts a native proof or guard selector.

Fresh positive minting fences the settled current guard and application lanes, current signed projection and signing key, retained allocation ledger and actual physical reporter/identity in one SmartData transaction. Guard and application must agree on source, boot, host namespace and protected authority. The report's nativeBarrier derives from that exact completed guard intent. It establishes allocation-pool denial, without claiming workload connectivity, DNS readiness, packet drain or lease release. Every handoff remains DOWN.

Every positive report also states the node's own uplink address — the address a route to its published ports uses, read from the retained uplink observation — on the request envelope (uplinkAddress, @serve.zone/interfaces 31.4+). The address is a fact of the pass that minted the receipt: a changed or absent address is a new receipt generation, so the pass after a rebind reports it absent (the stale observation is closed) and the next pass, over a fresh observation, reports the new one. Nothing defaults or infers it, and a withdrawal states none.

Beside it the report states hostAddresses (@serve.zone/interfaces 32.23.0): every address the same observation binds on the uplink — the lease address and the permanent /32 host addresses beside it — sorted as strings. Cloudly names a plan platform endpoint in this node's hostPlatformEndpointIds when its address is one of them. They are read together with uplinkAddress from one observation and are part of the receipt's identity in the same way: another set is a new receipt, and a report without an uplink address states none. Only the uplink's addresses are stated, because the packet compiler serves a host-local platform endpoint only on an address the uplink binding carries. An outbox entry a 33.0.0 node wrote carries none and is read as stating none.

The SmartData outbox retains immutable receipt revisions, original reporter bindings and completed-application references. Positive reads verify the historical completed guard too. One pending receipt supplies backpressure; a newer application cannot replace it. Historical replay uses those immutable proofs without requiring old guard/application heads to remain current. Exact ACKs fence the current physical connection and identity. Reconnect sends the original receipt body, then a later fresh native pass may queue a current-session successor. Replaying historical positive evidence does not re-establish current-session eligibility.

A stop withdraws nothing (@serve.zone/interfaces 32.44.0, runtimeNetworkProtectionContract.withdrawal). Quiescence stops admission and joins the native pass and the execution owner; the last receipt stays current and the node stays eligible for allocation, because its allocation-pool guard stays in the kernel when the process ends and its boot unit restores it before networking. A clean stop, a restart, a crash, a reboot and being offline are all the same to Cloudly: a pending receipt stays in the outbox for the next session's historical replay, and a reconnect binds a new reporter session whose first positive receipt restores eligibility. Physical disconnect alone does not remove Cloudly's durable allocation eligibility.

Only Cloudly's retirement request withdraws. It travels in the signed projection (IRuntimeNetworkProjection.retirement, { disposition: 'withdrawn', protectedAuthority }): Cloudly signs it into a node whose egress ownership it retires with the withdrawn disposition, its first appearance names exactly the projection's current protected authority, and every successor carries it unchanged (runtimeNetworkProjectionRetirementContract). Cloudly sends it only to a node that reports a compatible interfaces version: Pallet's registration offer states the release it carries (protocol.interfacesVersion, 32.50.0, beside minimumPeerVersion 32.0.0). The network pass that completes an application whose admitted projection carries the request latches it for the node's receipt chain, once, in the same commit as its admission guard (pallet_network_protection_retirement), and mints no positive receipt; from then on no pass does, and the store refuses one (protection.enqueueOwnedProtected.retiring), across restarts. The controller client answers before the pass resolves: it drains and ACKs any pending predecessor (a positive receipt minted before the request still drains), appends the chain's exact null successor and requires its ACK. The null retains its predecessor's historical application, authority and boot, so the answer works during key and projection gaps; its reporter is the current physical session. A withdrawal without a latched request is refused (protection.enqueueWithdrawal.notRetiring). A failed answer leaves the immutable pending history intact and fails the pass, naming the step it failed on (controller.withdrawal.*) and that step's own error; the next pass reads the latched request again and finishes the answer. The node's allocation-pool guard is never released for it: the answer withdraws eligibility, not protection.

A report the controller, or the relay carrying the session, refuses fails as PalletControllerRefusalError:<reason>:controller.refusal.<method> with the peer's TypedResponseError as its cause, in every chain it ends up in — a retirement answer's controller.withdrawal.predecessor included. <reason> is the most specific word-only token the peer stated: its error payload's reason, else its payload's code (Cloudly answers a refused report with { code: 'runtime-session-report-refused', reason }), else its refusal text when that is plain prose (words and spaces, at most 96 characters), turned into hyphen-joined lowercase words, else unstated. A token is lowercase letters in hyphen-joined words, at most 64 characters: no digit, dot, colon, slash or other punctuation passes, so no address, path, digest, session or node id reaches a label. A word made only of letters does pass as the peer stated it, so prose that names, say, a letters-only host name carries that name into the token; Cloudly's payload reasons and the relay's refusal text are words from their own fixed vocabularies. A delivery pass that fails while the client still admits work writes Pallet controller delivery failed: <chain>. (the owner_failed labelling: class names, codes, sites) to stderr, which Spark forwards to the node journal, once per distinct failure until a pass succeeds again: a receipt the controller keeps refusing is named when it is refused, not first at the stop it then fails.

A refusal holds back only what the contracts order behind the refused report. A delivery pass sends four kinds of unit, in this order: the protection chain, the application receipt chain, each assignment of one bounded page, and each capture's workload log batches. Within a unit the order is the contract's: protection receipts and application receipts are each one monotonic chain with one pending receipt, so a refused receipt holds every later one of its chain; an assignment's terminal receipts go in sequence and before its observation (the controller accepts a removed observation only after the joined removal receipt), so a refused one holds that assignment's later reports; a capture's batches go in sequence. Nothing orders one unit behind another — a protection receipt neither gates nor proves workload readiness, and assignments are independent of each other and of protection — so a refused unit ends only itself. The pass goes on with every other unit and, once it is done, rejects with every refusal still standing on the connection (one, or an AggregateError of all of them), those of units still waiting out their retry delay included, which the stderr line above names. A refused unit then waits before the delivery cadence sends it again: the delivery interval, doubled with each refusal in a row, at most 30 s (refusedReportMaximumRetryMs), while every other unit keeps the cadence. A unit that delivers, or a new controller connection, starts again from the interval. flush() sends every unit at once, whatever it waits out. Any failure that is not an answered refusal — a timeout, a lost connection, an answer that does not bind to its report, a store or capture failure — still ends the pass; when refusals stand, the pass rejects with an AggregateError of that failure first and every standing refusal after it, so a failure that recurs on every pass never hides a refusal from the stderr line.

Projection application receipts

A projection carries its predecessor's packet-grant withdrawals, and may not grant a withdrawn flow again, until the receiver supplies the exact previous projection as applied (admitRuntimeNetworkProjection's appliedPrevious). Pallet states that fact as IRuntimeNetworkProjectionApplicationReceipt (@serve.zone/interfaces 32.50.0): its durable claim that every packet table it holds was composed from one admitted projection, or that it holds none, so the flows its kernel admits are within that projection's grants.

Only the network pass mints one, after it completed an application (PalletNetworkOwner.reconcile, beside the protection receipt), and only from the ACTIVE journal's own evidence, fenced in the minting commit (PalletActiveStore.fenceApplicationReceipt):

  • the UP generation, neither withdrawing nor DOWN, carries exactly that application, its last in-place transition acknowledged on both roles (activeCarriedApplication), so every pair the kernel may hold — that transition's and every amendment or tunnel pair after it — was composed over it; on a docker-shared node the generation's DOCKER-USER contribution, a rule set of its own that follows each host table after it is acknowledged, must have its last step acknowledged on a host table of that same application. The receipt's nativeJournal names the generation;
  • or the node holds no packet table at all (nativeJournal: null): no generation was ever journaled, or the head generation is DOWN with its packet cleanup recorded and complete — both tables released in its own epoch, the host table that survives a same-boot epoch proof released through the engine, none left after a reboot — and its contribution released or ended with its boot.

Anything else mints nothing: an activation not yet UP, a transition awaiting an acknowledgement, a withdrawal in progress, a DOWN whose tables are still held (an epoch-proven generation's surviving host table included). The head is the only generation that can hold a table, because a successor is journaled only once its predecessor's tables are released or are the very tables its policies replace, and only once its predecessor's DOCKER-USER contribution is released or ended with its boot (active.begin.previousContribution). A DOWN whose contribution release failed or never ran is released by the next pass's fence (fenceRetained) before it raises a successor. With no generation journaled, the receipt and a first generation's intent both write the application lane, so they serialise. An admission ACK, a table inspection or a success flag never mints a receipt, and no transport facade can mint one (runNetworkApplicationReceipts reads and acknowledges only).

The receipts form one chain (applied:<scope digest>), kept in the identity database beside the projection chain they name (pallet_network_application_receipt_lane, …_entries, no migration). A new receipt is minted only for a newer projection than the latest one names; a pass whose application names an older one fails (applicationReceipt.append.regressed), and projection admission refuses a projection older than the latest receipt's (projection.admitProjection.behindReceipt), so the node never applies packet authority from a projection older than its latest receipt. The chain continues across restarts and reboots and never restarts at generation 1 while the database holds it; losing the database loses the node's enrolled identity with it, and a restored older copy's next receipt is refused by the controller, which holds a later one, exactly as that copy's older projection chain is.

The node's own projection admission reads its latest receipt in the admitting transaction and passes selectRuntimeNetworkProjectionAppliedPrevious(receipt, previous): the previous projection when the receipt names exactly it, else null. With it, a successor that drops the predecessor's withdrawals, or grants a flow withdrawn earlier again, is admitted; without it, both are refused as before.

One receipt is pending at a time. The controller client sends it on the current authenticated session (reportRuntimeNetworkProjectionApplication), requires the answer to name the exact receipt sent, and its authority use to be current for a receipt minted under this session and historical for one minted under an earlier one, then acknowledges it in the outbox. It sends receipts only to a controller whose accepted registration offer states interfacesVersion 32.50.0 or later (controllerReadsApplicationReceipts); an older controller keeps the session and the receipts wait in the outbox. Pallet's own offer states 32.50.0, which is how a controller learns that this node mints receipts and passes its own receipt to its admission.

test.activeapplicationreceipt.node.ts covers the evidence: none while a generation is prepared but not UP, a transition is pending, a withdrawal is in progress, a DOWN's tables are held or an epoch-proven host table survives. test.activeapplicationreceiptdocker.node.ts covers a docker-shared contribution that follows or lags its host table, a predecessor contribution that survived its DOWN, and a first generation begun while a no-table receipt commits. test.networkapplicationreceipt.node.ts covers minting with no table, one pending receipt, admission with and without a receipt, the chain across database restarts, historical replay and offer gating.

Offline allocation-pool guard journal

PalletNetworkGuardOwner composes the private identity database with published Smartnftables 2.1. commission() explicitly creates the first guard intent; recover() requires an existing journal. The caller supplies a trusted native binary path and must join this owner before closing the database. The node runtime composes this private owner with application completion and positive reporting; Spark installs the separately commissioned boot dependency.

The fixed receiver-owned trust record supplies the node scope, checked against the active local identity. Bootstrap verifies the latest admitted projection with its retained signer, including a crash between key rotation and the next signed projection. That historical material supplies denials only; DNS and new projection admission still require the current key. The policy denies exactly the declared allocation pools, without treating management LANs or resolvers as allocation pools. Its only exceptions are the exact host grants of the source's signed onebox projection and the publications its signed projection places on this node (see Published ports). They are derived from that verified projection whenever a target is proposed and proved again from the retained source on every read of an intent, so a stored row never widens the guard by itself. A guard Pallet 36.2.0 or earlier committed carries no publications; it is read as applied, narrower than its source, and the next guard pass appends the same-source transition that adds them, as it does for the stream port owner below. Every policy also restricts the CRI stream server of Pallet's own containerd, 127.0.0.1:10010, to uid 0 (localTcpPortOwners, Smartnftables 4.0.0). A pass refuses unless the policy it applied carries that owner; a guard committed before the owner existed gains it through one appended transition of the same source.

SmartData persists each complete native identity and original previous-Applied/null to prepared-target transition before reconciliation. Immutable result records bind the full native receipt to that intent. One pending lane prevents competing targets. Same-boot recovery repeats the original transition, including after a lost reply or completed result. A different actual boot records the old applied/pending heads in an immutable epoch and restores the prior denial with a fresh process instance and previous: null. A changed namespace in the same boot is rejected.

Recovery restores the committed denial before applying a newer inert projection. Every previously guarded pool must remain with the exact same id, purpose and prefix. Native preparation must reproduce the durable target, and exact enforced inspection and verified detachment must complete before success. An ambiguous failure uses closeRetaining() to join the child without deleting its policy or inventing an Applied receipt. No journal result alone grants allocation eligibility.

test/native/qualify-guard-boot.py runs real NoSQLDB and nftables in four isolated offline root boots under Node and compiled Deno. It covers same-boot completed and pending recovery, lost/undelivered native calls, new-boot epochs, signing-key gaps, newer staged pool expansion, and actual pool denial with unrelated traffic controls. On both boots a root TCP connection to 127.0.0.1:10010 succeeds and the same connection as uid 65534 is refused, before and after the new-boot recovery. It also exercises PalletGuardProcess and reopens the same database after its oneshot completes. Supplying --control-directory dist_control/linux-amd64-<digest> also runs the built production pallet-control guard-recover with null stdin and verifies its completed events, database release and continued pool denial. Production still requires the commissioned Spark boot dependency, and live activation the qualification of the complete workload networking, DNS and VPN path on a real host.

Host kernel requirements

The packet engine (Smartnftables 4.6) requires Linux 6.9 or newer, with the nftables filter and NAT modules (nf_tables, nft_nat, nft_chain_nat, nf_nat and the reject modules the port owner rule uses) available. A kernel below that floor, or one lacking a feature a policy needs, is refused by the engine as UNSUPPORTED_KERNEL; Pallet names that refusal instead of hiding it. The guard owner and PalletGuardProcess fail with code unsupported_kernel, and guard-recover, guard-commission and runtime-serve end with failed reason unsupported_kernel on their process protocol, exit code 1. Every other engine failure stays owner_failed. A host must boot a supported kernel before Spark commissions the guard.

PalletGuardProcess verifies the guard executable against trusted bundle metadata, opens only existing enrolled storage and runs one offline recovery or commission. It joins the native owner before stopping NoSQLDB. A native cleanup failure keeps the database lease owned for an explicit close() retry. Cancellation joins admitted work and fails the oneshot; success follows verified policy retention, native child exit and database shutdown.

The compiled control accepts guard-recover for ordinary boot and guard-commission for explicit first commissioning. Both accept exactly one mode argument and no runtime configuration. Null stdin is valid for these oneshots; unexpected input or SIGTERM/SIGINT fails without readiness. Successful completion emits ready and stopped under pallet.guard.process after cleanup. These events provide no positive allocation or workload protection receipt. The installer must acquire its first authenticated projection and commission before enabling recovery as a boot prerequisite.

test/native/qualify-protection-boot.py separately qualifies the complete private network owner in four offline Linux6.18.35 boots, under Node and compiled Deno. It verifies actual nftables denial with independent UDP controls, real dormant handoffs, completed application/guard agreement, historical positive replay after reboot, a fresh new-boot positive, the null withdrawal that answers a retirement request carried by a projection a newly advanced signing key signs, and PalletNodeProcess cleanup that reopens the same database. The packet client uses the pinned Node executable for both runtimes so a Deno UDP compatibility error cannot be interpreted as enforcement. No kernel, native-owner or database result is stubbed in these guests. They have no network device or host mount. Supplying --control-directory dist_control/linux-amd64-<digest> also runs the built production pallet-control runtime-serve on each runtime's second boot. It verifies readiness, EOF-driven shutdown, clean child exit, database reopening and continued pool denial after the entire node owner has stopped.

Initial projection acquisition

pallet-control network-acquire opens only existing enrolled storage and a projection-only controller connection. It waits for signing authority and a projection delivered over that actual physical session; a retained older projection cannot satisfy acquisition. It rejects assignments and does not fetch registry credentials or deliver terminal, observation or protection reports. After admission stops, it joins every controller write, verifies the delivered projection under the final current key, and closes controller/storage before ready and stopped under pallet.network.acquisition. Null stdin is valid; input, signals or the bounded operation timeout fail the oneshot after joining its owners.

pallet-control network-acquire-local is the same acquisition for a node driven by a controller on its own host (Local controller). It serves the local socket instead of connecting to a relay, and it additionally admits the first bind of a node no controller has bound yet; a node with any Cloudly enrollment state refuses to start it. The store must already be provisioned.

pallet-control store-provision provisions it: a null-stdin oneshot on the pallet.store.provision protocol that creates the data root and the database, prepares every collection, runs the store's migrations and releases the database again. It writes no enrollment, identity, binding or authority record and serves no socket, so the node stays unbound until network-acquire-local admits its first bind. An existing store is accepted as it is; a store that carries a Cloudly identity is refused (PalletStoreProvisionError:cloudly_enrolled on the owner_failed diagnostic line). enrollment-provision also creates the store and leaves no Cloudly state behind by itself, but it serves the Cloudly enrollment socket for its lifetime, where a prepare request writes that state; a node for a local controller is therefore provisioned with store-provision.

The installer runs acquisition while management networking is available and before first guard-commission and boot-unit installation. Existing guard recovery remains offline and requires an existing journal; it never silently commissions a new one.

Previous-epoch host attachment audit

A fresh PalletHandoffSetOwner can call verifyPreviousAbsence({ journal, target, previous }) before it applies a current set. Supply the authenticated persisted application reference, exact old prepared target and its prior applied receipt, including an unfinished transition's complete retained membership. Native code recomputes both old identities and their transition relationship, requires the same boot and host namespace, and retains the shared host mutator lock. The journal reference is an opaque binding; native code does not authenticate it.

Two complete host inventory passes reject old host or misplaced router names and MACs, Pallet markers, retained host indices and transit address/route conflicts. Old router indices remain scoped to the old namespace. The fresh router must independently remain empty and DOWN. Any candidate or inventory loss rejects the audit; it performs no cleanup, repair, deletion or adoption. Persist the returned old/current namespace-bound evidence before permitting a new epoch's effects.

The evidence proves old host attachments absent under exclusive privileged ownership. It does not prove physical peer destruction, packet drain, allocation reuse or a current nativeBarrier. Kernel peer teardown can be asynchronous, and another privileged writer can invalidate negative observations. Full node recovery uses the application owner's journal and complete authenticated history; process death or a saved namespace identity cannot replace this audit. The isolated Node/Deno fixture covers surviving and renamed old attachments, reused host indices, unrelated host indices matching old router indices, transit conflicts, fresh-router state, old request tampering and wrapper lifecycle.

verifyPreviousBootFence(request) is the separate fresh-owner capability for a different kernel boot. It validates the exact historical target and prior receipt, requires different valid kernel boot UUIDs, and reads the current boot from the native owner. Two complete audits still reject current old identities, Pallet markers, transit conflicts and nonempty fresh-router state. Old interface indices are deliberately ignored: those numbers can legitimately belong to unrelated interfaces after reboot. The returned evidence binds both epochs and adds previousBootTerminated: true. It supplies no allocation release, packet drain or nativeBarrier; the dormant-only history and privileged ownership requirements still apply.

The offline test/native/qualify-handoff-boot.py fixture runs two boots each under Node and compiled Deno with independent disposable ext4 NoSQL roots. The actual application owner persists real signed authority and applied/pending history. A fixture-only interrupted completion retains the acknowledged DOWN target, then the fixture flushes the database then signals init through a FIFO while the native owner and DOWN pair remain live. Init forces poweroff. The next boot reloads that journal, qualifies reused host indices and conflict rejection, then the application owner persists the boot fence and realizes the same still-reserved lease DOWN. All application state uses SmartData/NoSQLDB; no journal sidecar or host network/disk attachment participates in the qualification. It also retains the exact unavailable report across poweroff, replays its original reporter binding, and queues the next boot's report only after acknowledging history.

Standalone control build

The control builder requires Deno 2.9.7 and the exact NoSQLDB 10.5.1 engine identities in binary/control-build.json. After the normal dependency install, build either Linux target or omit the target to build both:

pnpm run build:control linux-amd64
pnpm run build:control linux-arm64

@git.zone/tsdeno 2.0.3 derives a temporary frozen Deno lock from the production tree of pnpm-lock.yaml under its managed runtime-only manifest, and per target leaves out the npm native binaries built for the other architecture or for macOS (their ELF/Mach-O headers prove them foreign); the control never loads them, because it runs its own sibling engine, guard and DNS executables. TsDeno refuses any Deno other than the pinned denoVersion, fails a pallet-control smaller than the target's minSize in binary/control-build.json (200 MiB; a compile that lost its npm payload still exits 0 with a far smaller binary), and on the host's own architecture runs it once without arguments in an empty environment, where it must print its invalid_arguments refusal and exit 1; the other architecture's smoke check is skipped. The build verifies the selected NoSQLDB engine's bytes and clean owning-build provenance. It also builds and verifies the Pallet static-musl executor against this exact source, version and architecture. The Smartnftables 4.6.0 guard is verified against its pinned bytes and clean provenance. Each completed dist_control/<target>-<digest> directory contains pallet-control, its fixed siblings pallet-smartdb, pallet-runtime, pallet-guard, pallet-dns and pallet-vpn, the pallet-containerd/ release directory (containerd, its shim, runc, the pause archive and their manifest), their native provenance records and a control-build.json recording artifact hashes, source state, compiler versions and runtime lock hash. The compiled control keeps UID 0 and its protected production data/socket paths; it derives only the sibling engine path from its own executable. The installer owns protected platform-parent directories. Packaging includes every native guard notice from the pinned upstream manifest (the native-notices/manifest.json that tsrust notices generates, with the digest of every file) under notices/smartnftables/ and rejects changed or missing material. The archive also includes the exact pallet-clock/ loader, programs, shared libraries and trust files, plus full clock source archives, patches, recipes and notices under sources/clock/ and notices/clock/. Consumers must verify the complete inventory.

binary/vpn-engine.json pins the published SmartVPN 2.5.0 musl executable, provenance, package license and native notice inventory for each architecture. The builder and packer verify the complete distributed VPN notice set under notices/smartvpn/, including the upstream Rust and Cargo license texts. A changed binary, symlink, missing notice or mismatched build is rejected. The bundled pallet-vpn is the executable runtime-serve verifies and runs as the managed VPN device owner (see "Foreground node process").

The two engines Pallet compiles itself, Chrony (chronyd, chronyc) and runc with the pause executable, are pinned by their outputs as well as their inputs: binary/clock-engine.json and binary/runc-engine.json record each output's SHA-256 and size per architecture. The rule is that released engine bytes always equal those pins. A control build reuses a verified earlier output under .nogit/clock-engine/ or .nogit/runc-engine/, or else reads the pinned outputs from the build-input store (see "Build input store") and verifies them against the pins. Only when the store holds none of a build's outputs does it build them from the pinned sources, and that build also fails unless every output reproduces its pin byte for byte. A store that holds only some of one build's outputs, a transport failure that outlasts the retries, or a byte that differs from its pin is refused by name; none of them falls back to a source build.

The source build is proven on every pin change: a changed recipe, SDK or source pin changes the output pins, and the maintainer runs the qualification below, which builds the engines from source and fails unless every output reproduces its new pin, before seeding the store with those outputs (see "Build input store"). Until the store holds them, a control build finds nothing under the new addresses and compiles from source itself. The qualification is also run on demand:

pnpm run engines:qualify               # both architectures
pnpm run engines:qualify linux-arm64   # one architecture

engines:qualify ignores the store and any earlier output, builds Chrony and runc/pause from their pinned sources for each selected architecture, fails unless every output matches its pin, and prints one JSON report of the reproduced outputs with the time each build took. The linux-arm64 builds run under emulation and take minutes. The verified output it leaves in .nogit/, beside the sha256.txt the build recorded inside its SDK, is the only kind of engine output the store seed uploads.

A source build additionally requires a Docker client/Buildx and a Linux Docker daemon able to execute the selected pinned Alpine architecture. Its build context and result archive use the daemon API; it requires no host workspace bind mounts. Compiler execution is offline, unprivileged and capability-free. The build checks pinned source, SDK and output hashes; workers receive only the finished programs and distribution materials and perform no package installation.

pnpm run release:control runs build:control and reads only its result lines from stdout. Every other stdout line (tsbuild, tsrust) reaches the job log on stderr as it arrives, prefixed with its ISO 8601 arrival time, so the log shows where a build's time goes; the engines report on stderr whether they came from an earlier output, the store or a source build, and how long that took.

A release build is reproducible or it is not released. release:control runs only on the exact clean release tag and sets SOURCE_DATE_EPOCH to the tagged commit's committer time (a job that sets a different value is refused), so tsrust stamps that time into the executor's provenance and tsdeno compiles the Deno binaries against a private cache with every embedded file at that time. It then runs the whole build:control a second time and refuses to release unless both builds name the same output directories and produce an identical file tree (paths, bytes and modes); the failure names every differing path. The second build reuses Cargo's warm target directory, so it proves the TypeScript, provenance and Deno outputs rather than a Rust rebuild on a fresh machine. Ordinary build:control runs leave SOURCE_DATE_EPOCH unset.

Every runtime dependency change must update pnpm-lock.yaml with pnpm install. TsDeno 2 derives .tsdeno.deno.lock, pins versions and dependency edges to pnpm's production tree, verifies the compiled graph against that tree and removes the temporary lock after compilation. Do not pass --lock or --frozen to tsdeno; there is no committed deno.lock for this pnpm-owned graph.

Carry noticeLockSha256, the control notices and their pinned asset hash in binary/control-build.json onto the new pnpm lock. test/test.controllock.node.ts checks the frozen pnpm manifest/lock pairing in a private copy, the identity and asset digests, registry origins and notice coverage for every production package. The build record's lockSha256 also names pnpm-lock.yaml. The tracked .npmrc sets registry=https://registry.npmjs.org/; scripts/control-lock.mjs rejects another NPM_CONFIG_REGISTRY origin and production tarballs from another registry.

The earlier Linux amd64 artifact with SmartDB 5.8.0 was qualified as root in an isolated QEMU guest with a disposable local ext4 disk and no network or host mounts. Checks cover the fixed production paths, engine digest and permission guards, explicit provisioning, durable enrollment replay after process restart, EOF/signals, native and parent SIGKILL, and stale-socket recovery. NoSQLDB rejects volatile filesystems; a tmpfs data root cannot substitute for supported local storage. This is process recovery qualification, not a power-loss or ARM64 runtime claim.

The NoSQLDB 10.5.1 engine passes the local native identity persistence and lifecycle suite with SmartData 11.14.2. Pallet now reaches it through the nosqldb family of @lossless.org/client 1.5.1, which continues SmartData 11.14.2 with the same persisted format. Both engine targets are the published static musl executables; the Deno control executable still targets GNU Linux. Both control targets compile and package with their verified owner provenance and complete upstream notices. The updated amd64 control bundle also passes the root guest checks above on an isolated local ext4 disk, including durable enrollment replay after native and parent process crashes. ARM64 runtime and power-loss qualification remain open.

test/native/qualify-node-process.py also qualifies a forward upgrade when given both --previous-control-directory and --previous-probe. It boots that earlier bundle first, retains a read-only copy of the stopped ext4 disk, then boots the new bundle twice against the working disk. Each boot verifies its exact binaries and probe; the final result compares the node identity and signed network projection and verifies the backup is unchanged. The NoSQLDB 9.0.0 to 10.2.0 amd64 run passes all three offline root boots, including retained workload leases, signature replay, namespace keeper loss and joined process shutdown. This qualifies existing current-format Pallet state; it does not qualify containerd, private DNS, physical power loss or foreign legacy database roots.

The guest stages the build's complete runtime inventory (executor, guard, resolver, clock loader and assets) and the nftables and veth modules, enrolls with a routing binding and commissions the allocation-pool guard before runtime-serve, which now owns the node's network; the guest provisions the root-only runtime directory on every boot, as the installer does. A runtime-serve whose namespace keeper is killed reports failed and exits: when its close cannot confirm the network release, the node process joins the clock, the namespace and the database anyway (PalletNodeRuntime.abandon) instead of holding them for a retry nobody in the exiting process makes, and the successor recovers the unconfirmed state by evidence, as after a crash.

The database package is @lossless.org/nosqldb and the controller uses its NoSqlDbServer API. Existing smartdb storage-directory and bundle field/file identities remain stable; they do not select the deprecated npm package. Changing the package does not convert an incompatible database format. NoSQLDB 10 reads existing version-9 current-format stores directly; once it writes version-10 records, older engines cannot read them. Retain a stopped pre-upgrade backup before activating the new engine. Native current-format admission still rejects foreign legacy and mixed roots before application startup.

Pallet admits the released 32.0 attachment-journal shape before any network owner starts. Terminal absence and an already durable pending removal retain their original bytes, digests and ACTIVE source or amendment history; the latter may finish with its released removal formula. The admission writes only the pallet_migrations completion row and is replayable after interruption. A released running or pending-ADD row has no durable removal transition that this version can complete without rewriting its history, so startup fails with a pallet-attachment-owner-recovery-* error and writes no completion row. Stop the upgrade and retain the prior Pallet owner to finish that lifecycle before trying again. There is no collection-reset or cleanup command in this recovery path.

Every node process owns its own router namespace, so a restarted process runs in a new epoch, and so does every process after a reboot. A generation the previous process retained — UP, or DOWN with a packet cleanup it never finished — is held by no native of the new epoch, and the native admits RecoverActive only for a generation of its own. The restarted process therefore fences it before its native adopts anything, with the audits that fence the application journal's own epoch, over the generation's application target and receipt: in the same boot, verifyPreviousHandoffAbsence proves the previous epoch's host attachments absent (the router namespace, and every pair, device and route under it, ended with that process); after a reboot, verifyPreviousHandoffBootFence proves the boot over. The proof is journaled as a DOWN of kind epoch (same-boot-host-absence or new-boot-host-fence), taken from the fencing native's epoch (PalletHandoffSetOwner.readEpochFence), and its packet cleanup releases what outlived the epoch: in the same boot the host table, through the running engine with the generation's retained host identity; the router table ended with its namespace, and after a reboot no table is left. The successor is raised on a pair keyed anew. A generation of the process's own epoch is still withdrawn in order while the process holds it — a tunnel transition left pending by a failed step is finished first — and recovered by RecoverActive otherwise. One the process journaled but never raised — an activation that failed between its intent and UP, such as a beginActivation that lost its source fence — is held by its native only as a preparation, and that native, having adopted its set, admits no RecoverActive: the process discards the preparation and journals the native's discard as a DOWN of kind discarded (PalletHandoffSetOwner.readActiveDiscard), then releases both tables as after any DOWN. A later process that finds such a DOWN with its cleanup unfinished takes it again as a recovery. A DOWN with no cleanup recorded at all — its owner ended after the DOWN and before the cleanup began, for example while releasing the DOCKER-USER contribution — is fenced the same way (fenceRetained): its tables are complete only for a successor that replaces them in place in that epoch, and the successor of a fenced generation is always keyed anew. In its own epoch it is recovered and its tables released; in a later one it is proven ended and its cleanup recorded, with the host table released within the boot and none left after a reboot. A journal row that carries discarded is refused by Pallet 36.1.2 and earlier, so a node that wrote one cannot be downgraded in place until two later generations were raised.

Every ACTIVE generation records the packet engine that prepared it, the SHA-256 of the Smartnftables executable the node process verified before it started (packetEngine; generations journaled before 35.0.0 carry none, so their engine is unknown). The kernel keeps a generation's host table across a Pallet restart in the same boot, and the restarted process asks the running engine to release it once the generation is fenced. An engine whose compiled graph differs from the one that applied it cannot adopt it and answers Conflict; Smartnftables 3.0 changed the graph of every routerEgress table and of every scope with publications. When that Conflict comes over a table another or an unrecorded engine applied, the release fails by name (PalletPacketEngineChangedError, conflict:packet.retainedByPreviousEngine) and the network owner stays failed and fenced, because no retry of the running engine can change the answer. The remedy is a reboot, which takes the kernel tables with it and lets the new-boot proof fence the generation, or finishing the generation on the previous Pallet before upgrading again. An upgrade that changes the packet engine across a same-boot Pallet restart therefore needs a reboot whenever an ACTIVE generation is retained. A Conflict under the engine that applied the tables is not this refusal and fails as before.

Local build outputs are qualification assets; a dirty source marker is recorded explicitly and cannot identify a release. The tag-triggered workflow publishes clean-source control bundles as inputs for Spark integration. The bundle's runtime-serve activates the node network with the bundled pallet-vpn (see "Foreground node process"); running it still requires Spark's whole-bundle installation and supervision. The complete activated path — ACTIVE generation, managed VPN tunnel to a real cluster hub, lease renewal and workload traffic — is qualified in parts (the guests below and unit specs), not yet end to end on a real host. ARM64 runtime and workload execution are not yet qualified by this component release.

Build input store

Every input the control build downloads is read from Pallet's build-input store, never from where it was first published: the Chrony runtime and compiler-image APKs, the clock and containerd corresponding sources, the runc compiler image and Go toolchain, the runc sources and the containerd release archives. The store is the public Gitea generic package pallet-build-inputs of the serve.zone organisation. Each input is one file, addressed by its SHA-256 and the last path segment of its pin's url:

https://code.foss.global/api/packages/serve.zone/generic/pallet-build-inputs/<sha256>/<name>

The store also holds the pinned engine outputs (see "Standalone control build"), so a control build needs no emulated compile while the engine pins are unchanged. An engine output has no provenance URL; its file name states what it is, <engine>-<version>-<target>-<output>, for example:

https://code.foss.global/api/packages/serve.zone/generic/pallet-build-inputs/<sha256>/chrony-4.8-servezone2-linux-arm64-chronyd
https://code.foss.global/api/packages/serve.zone/generic/pallet-build-inputs/<sha256>/runc-1.5.1-servezone1-linux-arm64-pause

A pin's url stays its provenance record and is never read by a build, test or release step. Upstream locations do not last: Alpine's stable repositories keep only the newest revision of a package, so a revision-pinned dl-cdn URL disappears with the next security update, and Pallet 36.0.0's sealed release could not be built for that reason. scripts/control-clock-inputs.mjs reads one store address per pin, verifies its size and SHA-256, and caches it in .nogit/clock-inputs/<sha256>. An attempt must receive the response headers within 10 s and a body byte at least every 30 s, and ends after 60 s plus the pinned size at 256 KiB/s, so a slow but steady transfer of a large input completes and a stalled one does not hang the build. Transport failures and these bounds are retried four times; a store without the input fails at once with ClockInputStoreMissError, naming the address and the seed command.

The store is public, so it holds no binary without its source. binary/sdk-sources.json lists every package of both compiler images (binary/clock-sdk.json, binary/runc-sdk.json) with the licenses its binaries declare. For every package whose license expression names a GPL-family license it pins the exact APKBUILD at the package's aports commit and every source that APKBUILD checksums: aports inputs from https://github.com/alpinelinux/aports, archives from Alpine's v3.23 distfiles, each matched against its APKBUILD SHA-512. The runtime packages' sources are in binary/clock-sources.json and containerd's in binary/containerd-sources.json; the store holds all of them. test/test.buildinputs.node.ts refuses an SDK pin the inventory does not cover. The qualification-only APKs in binary/clock-qualification.json are not build inputs and are not in the store.

To refresh a pin:

  1. Change the pin (url, sha256, size, and every digest that pins its file). For a compiler-image package, update its binary/sdk-sources.json entry and sources too. A change that moves an engine's outputs also changes their pins in binary/clock-engine.json or binary/runc-engine.json; run pnpm run engines:qualify before seeding: it builds them from source and proves the new pins.
  2. Seed the store. pnpm run inputs:store --dry-run reports, without uploading, every input and engine output the store lacks with its size and every one it cannot obtain. A maintainer with package write access to the serve.zone organisation then runs pnpm run inputs:store with the token in GITEA_PACKAGES_TOKEN. The tool takes input bytes from .nogit/clock-inputs and every --cache <directory>, which it only reads, or else from the provenance URL. Engine outputs have no provenance URL and are never taken from a cache: it takes them only from a source build under .nogit/clock-engine/ and .nogit/runc-engine/ of this checkout or of a --checkout <directory> (another checkout of the same pins, whose input cache it also reads), whose sha256.txt, written by the build inside its SDK, records the pinned digest. An output read from the store carries no such record; an engine output without a source build is reported unobtainable and the tool exits 1. It verifies every byte against its pin, uploads only what the store lacks and never replaces a store file.
  3. Commit. A tagged release job reads only the store, and builds an engine from source only when the store holds none of its outputs.

TLS trust of the control executable

pallet-control verifies every TLS peer it dials — the controller or Cloudly origin of the enrollment, runtime and network-acquire sessions — against the Mozilla root bundle compiled into Deno and then the operating system trust store. A node whose controller, relay or Cloudly certificate is issued by a private or internal certificate authority trusts it the way every other service on the host does: install the CA certificate into the OS store (update-ca-certificates, update-ca-trust). Pallet has no trust setting of its own and never disables verification.

A compiled Deno executable carries no store selection and otherwise trusts the Mozilla bundle alone, so runPalletControlCli sets DENO_TLS_CA_STORE=mozilla,system for every mode, after the argument check and before any mode starts; Deno reads the variable once, when the process opens its first TLS client. An operator who sets DENO_TLS_CA_STORE explicitly keeps that choice, provided it lists only mozilla and system; any other value, including an empty one, stops the process with one failed line of reason invalid_ca_store on the mode's protocol and exit code 1, rather than failing every later connection. Spark starts pallet-control with a cleared environment, so its processes always run with the default. The managed QUIC tunnel of pallet-vpn, a native SmartVPN executable, authenticates the hub by the public key its credential names rather than by a certificate authority, so no root store applies to it.

pnpm exec tstest test/test.tlstrust.node.ts --verbose --logfile --timeout 120

The test pins the default and the refusal, runs the CLI with an invalid store in three modes, and, under the pinned Deno, dials a real HTTPS and WebSocket server whose private CA is present only in the OS store (SSL_CERT_FILE): it is refused as UnknownIssuer without the selection or with an explicit mozilla, and trusted with it.

Regenerating third-party notices

binary/control-third-party-notices.txt names the frozen npm graph, the Deno GNU Linux runtime and the NoSQLDB engine. When the Deno pin moves, move every pin first — denoVersion in binary/control-build.json, the sealed release's denoVersion in .smartconfig.json and both Deno lines of .gitea/workflows/release.yml — and add the new version's release commit and the SHA-256 of its deno_src.tar.gz release asset (GitHub states it as the asset digest) to binary/control-deno-sources.json. Then, with that Deno and cargo-about 0.9.1 on PATH:

pnpm run notices:control            # rewrite the notices and their pins
pnpm run notices:control --check    # regenerate, compare, write nothing

scripts/notices-control.mjs refuses unless every Deno pin, the running Deno and the pinned source agree. It downloads the pinned source of the Deno the notices name and of the pinned Deno, checks both against their digests, and takes the runtime's crate inventory with cargo-about over cli/rt for the control targets (--locked, so Cargo fetches the locked crates). It needs tar and network access to GitHub and crates.io, and works in a private temporary directory it removes afterwards.

The regeneration carries the Deno runtime part and leaves everything else byte-identical. A registry crate the notices already name keeps its entry; Deno's workspace crates and every new crate are written from the crate's own legal files, the workspace crates with Deno's LICENSE.md. A new crate without a legal file is refused, because it needs a reviewed entry. The native V8, Chromium Rust and Rust standard library notices are carried only while the pinned Deno runs the V8 and the Rust toolchain they name; otherwise the tool refuses until they are refreshed. The v8 crate cites rusty_v8's license at the crate's own commit. The embedded JavaScript keeps its file list: each listed file's legal comments are read again at the new release, each excerpt is found again at its new lines, and a changed or new source file under ext/, runtime/js/ or libs/core/ whose legal comments are not listed is refused for review. The asset hash and noticeBuildIdentity.denoVersion in binary/control-build.json follow the written file. --input <file> carries another notices file instead of the committed one: from the Pallet 34.0.0 notices (Deno 2.9.4) the tool reproduces the committed Deno 2.9.7 notices byte for byte. test/test.controldenonotices.node.ts covers the pin refusals and a dry run over a small fixture without cargo-about.

Control bundle packaging

@git.zone/tspack owns archive assembly, file hashes, executable modes, sealed manifests and complete archive verification. Pallet's adapter selects explicit control builds and checks their source, version, compiler, runtime lock, fixed paths and both native owners' provenance before passing inputs to TsPack. The selected license notices are pinned in binary/control-build.json to the runtime lock, Deno version and NoSQLDB owner build. The pinned native notice inventory binds the Cargo/compiler inputs and every reviewed notice; the archive preserves the native notice index, complete inventory and CRI provenance.

After sealing, the adapter extracts every bundle again from the sealed set and runs the verifier of @serve.zone/pallet-bundle over it with the build's own source identity, so a bundle that disagrees with what that package verifies for its consumers is never produced. Its paths and control record keys come from the same package (ts_bundle/inventory.ts). pack:control first compiles that package (tsbuild custom ts_bundle to dist_ts_bundle/, output on stderr) from the checked-out source, so it runs on a clean checkout and never loads a stale verifier; release:control loads the adapter only after build:control has compiled it.

# Use the exact project-relative directories printed by build:control.
pnpm run pack:control dist_control/linux-amd64-DIGESTPREFIX dist_control/linux-arm64-DIGESTPREFIX

# A release requires both architectures built from the clean vVERSION tag.
pnpm run pack:control --release dist_control/linux-amd64-DIGESTPREFIX dist_control/linux-arm64-DIGESTPREFIX

Ordinary packaging produces an explicitly unpublishable qualification set. A release produces the sealed set under dist_control_release/ and an inputs-VERSION-COMMITPREFIX.json description that records its exact packaging configuration and manifest digest. The output JSON identifies both paths.

A publishing workflow must retain the entire dist_control_release/ directory, including that input description and sealed set, before publishing any asset. Restore those original files on retry, then run:

pnpm run pack:control --release --reuse

Reuse verifies the retained source, configuration, manifest and every archive. It works without the original compiler outputs and rejects missing, corrupt or changed retained inputs. It never recompiles a control bundle: the only thing it compiles is the TypeScript of the bundle verifier above. Fresh Deno compilation may produce different bytes, so rebuilding cannot substitute for retaining a published set. TsPack and this adapter do not provide remote CI artifact storage or publication.

GitZone 6.6.2 or later manages .gitea/workflows/release.yml and the scripts/gitzone-*.mjs release scripts through the committed tspackRelease asset configuration. The tag-triggered workflow invokes the configured scripts/prepare-control-release.sh command in its disposable job container. It requires a root Linux Gitea Actions environment, installs Pallet's clang and musl-tools compiler prerequisites through apt, then executes the existing scripts/release-control.mjs exact-tag adapter to build and package its two explicit output directories. Compiler installation writes only to stderr so the adapter retains its single JSON result on stdout. The runner must be Gitea Runner 3.3.2 or later with its cache/results service reachable from job containers and runner.patch_actions enabled. The pinned stock artifact actions use the v4 protocol required by Gitea's REST retention inventory. It retains the complete output root in Gitea Actions for 90 days, downloads that retained copy, and verifies component reuse and every archive before creating a release draft. The publisher checks existing attachment bytes, adds only missing files, and publishes after complete readback. It never replaces release assets.

After a CI interruption, rerun the original Gitea run so it restores the original bytes. Missing or expired retention and conflicting remote attachments stop publication. The generic release receipt gitzone-release.json and Pallet's input description must remain beside the sealed set in the retained artifact. Use gitzone format --only assets --write --yes to update these managed scripts from a published GitZone version, then commit the result before releasing.

Build and test from source

This development package is marked private and is never published. The npm release target publishes one module of this repository, ts_bundle/, as @serve.zone/pallet-bundle with Pallet's release version (see its own ts_bundle/readme.md). Its tspublish.json takes the @git.zone/tspack and @push.rocks/smartdaemon ranges from the root devDependencies (devDependencyVersions), so neither enters the root dependencies that pnpm-lock.yaml supplies to the control build, and @serve.zone/interfaces from the root dependencies. From the repository:

pnpm install
pnpm test
pnpm build

pnpm test is a plain tstest --verbose --logfile; the test set lives in the @git.zone/tstest section of .smartconfig.json. Its prepare.once builds the native debug binary (cargo build --locked) and warms the control-pack inputs, then the files under test/ run six at a time (concurrency), longest first by the durations in tstest's run ledger, each in its own process with its own 180 s timeout, and the Rust tests run as the suite rust (cargo test --locked) beside them. Runs are recorded in the ledger; the repository never reuses a passed run (reusePassedRun: false). A test file is split before it needs more than two-thirds of its timeout at that concurrency, so contention cannot push it over.

Tests exercise the native debug binary against a disposable Unix-socket HTTP/2 fake CRI server. They do not connect to Docker or a production daemon. The native build uses Rust 1.95.0, locked Cargo dependencies and pinned vendored protoc to produce static musl Linux binaries through @git.zone/tsrust, named pallet_linux_amd64_musl and pallet_linux_arm64_musl. pnpm build builds only the host architecture's executor; pnpm run build:control builds both with tsrust --configured-targets, so a control build and the release always carry the executors of both targets built from the same source. ARM64 cross-compilation uses aarch64-linux-gnu-gcc as the linker driver with Rust's self-contained musl target libraries. The ring TLS provider also requires musl-gcc for amd64 C/assembly and Clang for its supported freestanding arm64-musl C build. These compiler choices are declared in rust/.cargo/config.toml; provision them on developer and release builders before running the build. Other packaged platforms are rejected explicitly.

After pnpm build, run PALLET_TEST_PACKAGED=1 pnpm exec tstest test/ --verbose --logfile --timeout 60 to exercise the host architecture's packaged executable and its default lookup instead of the native debug executable. Cross-compiling an ARM64 artifact does not establish ARM64 runtime qualification. For explicit emulator qualification, the same suite accepts an absolute PALLET_TEST_BINARY_PATH pointing to a test-only emulator launcher. Do not combine it with PALLET_TEST_PACKAGED; an emulated run does not qualify real node hardware.

The third-party notice index points to the native notices tsrust notices generates in native-notices/ for the locked Cargo graph, the Rust toolchain and the static runtime, with full license texts and the reviewed once_cell licenses and source attributions of ring under native-notices/crates/ring-0.17.14/extra/, and lists the material kept beside them. npm publication and distribution of the standalone Rust runtime-probe binaries remain disabled. Gitea releases distribute the control bundles described above, whose runtime-serve activates the node network as described in "Foreground node process".

Read-only probe

After a build, from the repository:

import { Pallet } from './dist_ts/index.js';

const pallet = new Pallet({
  socketPath: '/run/pallet/containerd/containerd.sock',
});
const evidence = await pallet.probeRuntime({ timeoutMs: 5000 });
console.log(evidence);
// {
//   runtimeName: 'containerd',
//   runtimeVersion: '<actual daemon version>',
//   runtimeApiVersion: 'v1',
//   runtimeReady: true | false,
//   networkReady: true | false,
// }

The socket must already exist and belong to a separately configured, authorized containerd instance. Pallet never creates one, chooses a default socket, searches PATH for its binary, or falls back to Docker. An explicit absolute binaryPath can select a verified installed executable or the development debug binary. A missing or nonexecutable selected binary fails with SmartRust 2's RustBinaryLocatorError (ERR_RUST_BINARY_EXPLICIT_PATH_INVALID). Selection never changes executable permissions or falls back to another packaged binary or a stale GNU/native build. After explicitly provisioning the selected executable, the same Pallet instance can retry.

Each probe owns a separate native child. Only one probe per Pallet instance is admitted at a time. The method confirms that child's exit before returning, including on failure. If cleanup cannot be confirmed, the method rejects and retains ownership of the child. Further probes are blocked until await pallet.close() successfully retries cleanup. An optional AbortSignal cancels the local request and terminates the owned child; it never stops containerd itself.

The native deadline covers connection and both RPCs (50–30000 ms, default 5000). The bridge readiness handshake, after executable discovery, is bounded to 3 seconds, with 1 second for graceful termination before forced child shutdown. Executable filesystem discovery in the shared bridge is not currently timed or abortable; this is not an end-to-end startup deadline. IPC and decoded gRPC responses are limited to 16 KiB. The probe accepts only containerd with CRI API v1.

RuntimeReady and NetworkReady must each occur exactly once; a missing or duplicate condition is an error, while explicit false remains false. Runtime diagnostic messages and verbose configuration are never returned. Runtime reported readiness is not workload health, version support qualification, network reachability, storage fencing, or permission to perform a migration.

Native failures carry a RustBridgeRequestError.responseErrorCode of INVALID_INPUT, SOCKET_UNAVAILABLE, RUNTIME_UNAVAILABLE, DEADLINE_EXCEEDED, UNSUPPORTED_RUNTIME, or INVALID_EVIDENCE. Bridge transport, cancellation, and startup failures remain distinct errors.

Isolation and qualification

Known Docker-private path components and aliases into them are rejected as accident prevention. This is not a security boundary against a hostile local user who can replace socket paths. Run only with the local permissions needed to inspect the intended daemon; containerd's socket is root-equivalent.

Fake-server tests establish protocol and lifecycle behavior only. Before runtime adoption, qualify a dedicated containerd 2.3 instance with explicitly owned root, state, socket and configuration paths, then verify actual container, network, storage and recovery behavior. No production cutover is implied.

Every QEMU qualification guest under test/native/ runs one bundled ES module. scripts/guest-bundle.mjs builds it from the guest's own TypeScript source with esbuild — the whole graph inlined, the NodeNext .js specifiers resolved to the .ts sources in this checkout, bare Node builtins rewritten to their node: form (which is what the deno compile --no-config --node-modules-dir=none half of each guest needs) and a createRequire banner for the CommonJS dependencies:

node scripts/guest-bundle.mjs --entry test/native/namespace.ts --out .nogit/debug/guest/namespace.mjs

It prints the bundle's path, size and SHA-256, which is what a qualification run records beside its other inputs. The runners take that file as --bundle (and qualify-active-registry.py as --wrapper-bundle / --tunnel-bundle).

qualify-cni-up.py --scenario single-host runs test/native/singlehost.ts, a node bound to a local Onebox controller, in the cni-up guest with no substituted seam. network-acquire-local serves the contract's socket (a root-owned 0600 socket in a root-only 0700 directory; a peer running as another user gets EACCES); the first bind persists the binding, an exact replay answers the same credential, a different binding refuses, a replayed bind's session supersedes the earlier connection's, which can then no longer apply anything, and the signed onebox projection is acquired over the current session. The node's runtime then raises one generation without a managed VPN: the real barrier reaches forwarding on the single-host fact with no tunnel command, execution admission is accepted on that evidence, two workloads attached through the real CNI path are raised under the generation, TCP and UDP flow between them with the client's own address preserved, and each resolves the other's name to its lease address through the node-local resolver. A third workload, on a network only the first shares, answers the first and is dark from the second. The router namespace reads IPv4 forwarding off in all and default before any generation exists, although the guest host forwards. From the host namespace it probes the granted TCP and UDP tuples, another port, another protocol, another lease, another source address and a granted lease that is not attached. The Cloudly counterpart — a generation without a tunnel receipt settles at up and its barrier does not hold — stays proven by the default cloudly scenario. With the generation's pool route in place the granted TCP and UDP tuples reach the workload from the transit host address, and another port, another protocol, another lease and another source address stay dark; a granted lease that is not attached stays dark as well, stopped by the host-transit table before the router receives a packet (both hops carry attached leases only). That last check captures every frame on both ends of the transit link while it probes (the guest stages tcpdump, without promiscuous mode, and parses its pcap) and counts only IPv4 packets to the unattached lease: the link also carries the ARP exchange both ends run on their own schedule — the router re-probes its entry for the host about 5 s after the granted flows — which a frame counter cannot tell from a leak. The same captures must see a granted flow on both ends, drop nothing and account for every frame the link counters counted; with the host-transit policy deliberately given the unattached lease's grant, the probe's SYN reaches the router and the check fails. A successor projection then withdraws every host selection under the held generation: the pass withdraws that generation before it completes the changed dormant set, raises the next generation on a re-keyed packet pair (forwarding, no host route planned), and the formerly granted tuple goes dark. Probers dial every 100 ms through the change: the allowed pair between the first two workloads, and six tuples the policies deny — the pair that shares no network over TCP and UDP, and workload to the host's transit and uplink addresses over TCP and UDP, each with a live listener behind it. East-west traffic stops for the withdrawal window, because the withdrawn generation restores a router that forwards nothing, and resumes when the successor is UP (about 27 s of a 45 s change); every denied tuple is dialled inside that window and never answers before, during or after it, and no listener logs its source. Without the forwarding pin the same guest measures the unfiltered router: the pair that shares no network answered 61 of 201 dials during the change. A further projection that only adds a workload, F, then moves the successor in place (the same generation, the successor's application carried); F attaches through the real CNI path and is raised under it, answers A, every read of the journal stands with F's amendment, and a fresh network owner over the same router namespace fences the generation at its start and raises the next one on its first pass. All of it passes on 6.18.35-0-virt. The packet hub guest (packet.ts) plans the same pool route and proves it present only while its generation is UP, gone after a withdrawal, and gone after a fresh owner's recovery of a lost native. It also proves the host's forwarding: off with both exclusive tables applied and before the generation is UP, on while it is UP, off after the withdrawal. A LAN peer behind its own link on the guest host, routing through the host, dials a live TCP and UDP listener on the external host while the generation is UP and the host forwards: dark under the exclusive host table, answered (with the LAN source preserved) once the host table is replaced one revision up by the same scope without exclusiveForwarding, dark again when the member is restored, and dark with no generation and no table at all; the published two-hop flows pass under the exclusive tables.

qualify-cni.py --scenarios containerd-serve --control-directory dist_control/linux-amd64-… runs Pallet's own containerd from a build:control directory the way its service unit does, with a foreign /etc/containerd/conf.d drop-in planted in the guest. It proves the generated configuration governs (the drop-in's stream port stays closed, 127.0.0.1:10010 serves), the bundled sandbox image is imported and unpacked on overlayfs, a workload raised by the production execution owner through the real CNI path runs under /pallet with runc state under /run/pallet/runc, SIGTERM stops containerd and leaves the workload running, the next containerd-serve reattaches the exact sandbox and container, and killing containerd fails its supervisor (owner_failed) while the workload keeps running. The image collector keeps the running workload's image and the bundled sandbox image, and the proven removal frees the workload's image while the sandbox image stays.

qualify-cni.py --scenarios registry-auth --bundled-containerd-directory <…>/pallet-containerd pulls through Pallet's bundled containerd release (the pallet-containerd directory a control build or node test/helpers/controlpackinputs.mjs <root> <directory> assembles, verified against binary/containerd-release.json) from a token-auth registry fixture on guest loopback, named registry.pallet.test:5000 as a relay's host:port origin is: the workload repository answers 401 with a Bearer challenge, and its token endpoint issues a token only for the node's Basic credential. A refused credential reaches the token endpoint, the run ends failed with cause unauthorized, the lane is free and the refusal is written once; the node's credential is issued a token, the workload runs and is stopped and removed, and no token request is ever anonymous. With the former bare host:port server address the same guest sees containerd fetch an anonymous token and fail 401.

qualify-cni.py --scenarios owner-storage --control-directory dist_control/linux-amd64-… runs one local storage claim under the same bundled containerd and runc, on node and Deno. The claim materialises its volume root-private with the claim's ownership; containerd echoes exactly one private read-write bind of it and the bundled runc's workload carries it at the claimed path; the workload writes as its own user; a second run of the service and a purge are refused while the first container lives (storage.claim.held, storage.purge.held); a restarted owner rejoins the live container against the grant's mount list; the second run finds the first one's bytes after the first is removed; and the purge deletes the volume once both are gone.

qualify-cni.py --scenarios owner-readiness owner-epoch runs test/native/cniepoch.ts instead of cni.ts: bundle it with node scripts/guest-bundle.mjs --entry test/native/cniepoch.ts --out <guest>/cniepoch.mjs and deno compile --no-config --node-modules-dir=none -A, and pass them as --bundle and --deno-driver (the other inputs are those of owner-attached). Both scenarios boot the attached guest — the pinned official containerd, real CNI, the production network and execution owners, the QEMU user-mode NAT for the NTS clock only, and the owner-attached scenario's test-only activation barrier seam — once on node and once on Deno.

  • owner-readiness runs an attached workload whose readiness is TCP on a listener inside the sandbox. The host, a Cloudly-style node, has no route to the workload pool (its route to the workload's address leaves through the uplink) and its own connect times out, yet the workload becomes ready: the native probe connects from inside the sandbox's journaled network namespace, and no sample is a namespace refusal.
  • owner-epoch starts the node process in its own control group, runs an attached workload, and then kills the whole group (cgroup.kill), as a crashed service unit's restart does. The router namespace goes with it, and with it the sandbox's eth0, while CRI still reports the container running. A fresh node process on the same store and the same containerd reads the run's network epoch as ended, records the run failed with cause network-epoch-ended, writes Pallet run failed: …: network-epoch-ended., stops the sandbox (CRI then reports it exited), observes the failure once, and answers the controller's stop and removal with their chained terminal receipts; the removal's CNI DEL answers from the restart's epoch coverage of the old pair.

qualify-active-registry.py --paths-node-binary … --paths-bundle … (with the packet and SmartVPN binaries) runs test/native/singlehostpaths.ts: the three single-host paths of a Cloudly node whose platform services run on its own host. The guest adds the hub's and the relay's addresses to the uplink as permanent /32s before the uplink is observed, attaches and raises three workloads that share no private network, and composes one projection with the production composer twice: without hostPlatformEndpointIds and workloadIngress, and with them. Under the first pair a workload's dial of the relay it selects, the router namespace's dial of the hub and the ingress workload's TCP and UDP dials of its target are all dark; the composed pair replaces it one revision up under the same UP generation and all four are delivered — the relay and the hub see the transit source address with a port of the leased range, the target sees the ingress workload's own address — and the real SmartVPN hub on the host authenticates the router namespace's real managed QUIC client. Another workload to the relay, the Corestore port the workload did not select, an undeclared hub port, the host opening toward the router, the target opening back toward the ingress workload, another workload to the target and another target port stay dark under both pairs, each against a live listener that logs no peer. Then the protected authority takes its next step that only adds (a platform endpoint joins), and the pair composed over the successor replaces the live one, one revision up, while one TCP connection from the ingress workload to the relay keeps exchanging numbered lines every 20 ms: every line sent across the replacement comes back on the same connection, with no reset and no stall over 500 ms; both policies then carry the successor's authority digest, fresh dials to the relay and the target are delivered and the negatives stay dark.

qualify-active-registry.py --tunnel-node-binary … --tunnel-bundle … runs test/native/activetunnel.ts, whose moves case raises the device under the real SmartVPN 2.5.0 hub and lets the hub move the node's split route in place (reconcileManagedNetwork at the next revision): the client moves the device's route, the native keeps the generation UP with the moved route pending, the revision the device owner reports is adopted (a replay answers the same adoption), and both packet policies stay enforced. A route on another device of the router namespace then fences the generation, and a fresh owner recovers it DOWN and releases the live tunnel pair through a packet owner of the same identities.

The pinned upstream CRI protocol and the containerd Transfer and Streaming protocols, with their Apache-2.0 attributions, are under rust/proto/.

This repository contains open-source code licensed under the MIT License. A copy of the license can be found in the repository license file. The vendored Kubernetes protocol is separately licensed under Apache-2.0.

Please note: The MIT License does not grant permission to use the trade names, trademarks, service marks, or product names of the project, except as required for reasonable and customary use in describing the origin of the work and reproducing the content of the NOTICE file.

Trademarks

This project is owned and maintained by Task Venture Capital GmbH. The names and logos associated with Task Venture Capital GmbH and any related products or services are trademarks of Task Venture Capital GmbH or third parties, and are not included within the scope of the MIT license granted herein.

Use of these trademarks must comply with Task Venture Capital GmbH's Trademark Guidelines or the guidelines of the respective third-party owners, and any usage must be approved in writing. Third-party trademarks used herein are the property of their respective owners and used only in a descriptive manner, e.g. for an implementation of an API or similar.

Company Information

Task Venture Capital GmbH
Registered at District Court Bremen HRB 35230 HB, Germany

For any legal inquiries or further information, please contact us via email at hello@task.vc.

By using this repository, you acknowledge that you have read this section, agree to comply with its terms, and understand that the licensing of the code does not imply endorsement by Task Venture Capital GmbH of any derivative works.

S
Description
In-development node-local containerd execution foundation for Cloudly and Onebox; currently a read-only native CRI probe.
Readme
120 MiB
v36.3.7
Latest
2026-10-10 16:25:58 +00:00
Languages
TypeScript 69.5%
Rust 18.7%
HTML 5.8%
JavaScript 2.9%
Python 2.5%
Other 0.5%