@serve.zone/pallet
Pallet is the in-development node-local containerd execution component for serve.zone. Its public API provides a bounded, read-only native CRI v1 probe. Backend-private identity, enrollment and stateless execution mechanisms are under development; production assignment delivery and reconciliation remain unwired.
Issue Reporting and Security
For reporting bugs, issues, or security vulnerabilities, please visit community.foss.global/. This is the central community hub for all issue reporting. Developers who sign and comply with our contribution agreement and go through identification can also get a code.foss.global/ account to submit Pull Requests directly.
Development status and ownership
Cloudly will own cross-node desired state and placement. Pallet will own local execution, including the outbound Cloudly connection on workers. Onebox will control its own local workloads through authenticated Pallet IPC, without depending on Cloudly to bootstrap itself. Spark remains the host installer and supervisor. Coreflow's required behavior must be ported before its legacy Swarm runtime can be retired.
Those orchestration and authorization capabilities are not implemented by the public probe. It exposes no public listener and persists no application state. Do not replace a production runtime with this development foundation.
No Pallet name carries a version token: collections, singleton row ids,
record-id hash domains, the native structs and their TypeScript twins, the
dormant link aliases (pallet-handoff: / pallet-workload:), the host and
workload evidence kinds and the process protocol names are named by what they
hold. No Pallet shape carries a shape-version field either — not the native
structs, not the identity, application, guard, attachment, ACTIVE, barrier,
tunnel or DNS records, and not the build manifests this repository owns. Every
one of those key sets is exact, so a value that still states schemaVersion is
refused as an unknown key rather than read as an older shape. What contract a
peer speaks is stated once, by the handshake below, and what a persisted shape
means is stated by the installed build. No released Pallet version was ever
deployed, so nothing is migrated and nothing is aliased — a store written by an older build fails its assertion on
every read, a dormant link an older build left behind carries a foreign alias,
interface name and MAC address because the seed domain changed with the name, so
it is never recognized as this build's pair, and a node's stores under
/var/lib/serve.zone/pallet are reset with the install.
Private CRI execution
ts/runtime/classes.executiondriver.ts owns a bounded native CRI client through
the separate --executor-management process mode. It is absent from the public
package facade and enrollment IPC. The eventual assignment owner must authenticate
the controller, validate immutable material, durably admit the assignment and
serialize its effects before invoking this mechanism.
The driver supports inspect, run, stop and explicit runtime removal against a
dedicated containerd 2.3 release and CRI v1 socket. Run supplies a digest-pinned image,
its platform image-config digest, exact argv, environment, working directory,
CPU/memory policy, user/group policy and root filesystem write policy. Explicit
cpuMillis: null leaves CPU quota unlimited. Explicit null for both runAsUser
and runAsGroup selects the pinned image's User through containerd's rootfs
resolution; a numeric pair overrides it. Missing fields and partial pairs are
invalid.
Run also states the sandbox's resolver, dns: { servers, searches, options }, and the
executor always sends it as the CRI sandbox's dns_config: containerd copies the host's
/etc/resolv.conf only for a sandbox that states none, and a node whose host file names the
systemd-resolved stub 127.0.0.53 would hand every workload a server it cannot reach (lab
2026-10-05 RUN 8, D10). An attached run's resolver is derived from its durable network
preparation (PalletNetworkOwner.readExecutionResolver): its single server is the lease's DNS
server — the node's per-workload resolver, which answers the private names of the run's networks
and forwards public names as the contract's DNS view for the lease decides — and its search
domain is the default network's suffix without the root dot, so containerd writes
search <suffix> (only with a default network) and nameserver <lease DNS server>. An isolated
run states three empty lists, for which containerd writes an empty file. The executor refuses
any other shape as invalid input before any CRI call: an isolated run with a server, an attached
run without exactly one canonical unicast IPv4 server, more than six search domains or a domain
that is not a lowercase DNS name without the root dot. A run whose network owner answers no
resolver is not started (unresolved-authority). Under @serve.zone/interfaces 32.48.0 a
workload on no private network forwards any public name to the protected authority's
resolvers (the address plan's), so a cluster declares the resolvers its nodes' workloads use
there; with none declared such a workload resolves no name and its resolver answers REFUSED.
Containers
use private namespaces, runtime-default seccomp and no-new-privileges, without
privileged mode or host networking; the only host paths a container ever receives
are the read-only per-file binds of its own staged secret mount and one private
bind per node-local volume it was granted ("Local storage claims"), which the
executor derives from its fixed volume root and the volume's key — no host path
crosses the IPC. Every sandbox names the absolute cgroup parent /pallet, so
workload cgroups sit outside any service unit's cgroup (see Pallet's own
containerd). Network policy and application readiness
remain separate owners. Secret material delivery is this node's own, through the
staged mount.
Every runtime object carries the immutable generation-one run digest and complete controller/node/assignment/replica/attempt ownership labels. Recovery requires one exact sandbox and container, including a full check for foreign sandbox children. Image identity uses the actual config digest and the requested repository digest; display tags are not identity. Running attempts can be recovered; exited or stopped attempts are never restarted. Stop preserves runtime objects, including containers created but never started. Removal requires confirmed stopped state and retains storage; images are left to the image collector below. Neither operation depends on a healthy image cache or network.
Input is captured before asynchronous work without invoking accessors. Native frames and CRI responses have fixed bounds; errors expose static codes rather than runtime messages or registry credentials. Cancellation, parent EOF and signals terminate the owned client. A timeout or client exit does not cancel a server-side CRI operation, prove rollback or fence a writer. The durable assignment owner must retain uncertain effects and resolve them before admitting replacement.
One failure is definitive instead: a PullImage the runtime refuses while the run
has no sandbox and no container. Nothing of the attempt exists then, so the native
owner answers IMAGE_PULL_FAILED with errorData { grpcCode, cause } — the
runtime's gRPC status code and one word of a closed set, unauthorized,
not-found, tls, transport, deadline, unavailable or other — and never
the runtime's message, which names the registry, the reference and the token
endpoint (PalletImagePullError). containerd 2.3 reports a registry's refusal
as Unknown with the HTTP status line in its message, so the code decides first
and the message's fixed markers after it. With a sandbox already there the same
refusal stays a runtime failure. The execution owner ends that run's operation
failed (no evidence), frees the native lane at once, observes the assignment
failed once — with cause { kind: 'image-pull-refused', reason, grpcCode }
(see "Failure causes" below); its controller times a replacement from the
observation it accepted — and writes
Pallet image pull refused: assignment <id> generation <n>: <cause> (gRPC <code>).
to stderr once per assignment and cause. The attempt is never pulled again: its
network preparation belongs to that one operation, and a new attempt — which
Cloudly admits after its failure backoff — is the controller's. Before this,
the refused pull stayed pending, held the one native lane for every other
assignment of the node until a reboot fenced it, and its assignment stayed
unknown for good.
A sandbox the runtime refuses to create is the same kind of fact when nothing of
the attempt is left afterwards: containerd tears a sandbox down when its network
setup fails, and the native owner reads the runtime again before it answers. With
neither a sandbox nor a container of the attempt left it answers
SANDBOX_REFUSED with errorData { grpcCode } (PalletSandboxRefusedError),
never the runtime's message; with anything left the refusal stays a runtime
failure. What the execution owner does next depends on whether a CNI ADD of the
run bound its attachment to the refused sandbox
(PalletNetworkOwner.readExecutionSandbox):
- No ADD bound it — the network owner refused the ADD itself, not ready yet or its
epoch not proven yet, which its next pass settles. The execution owner drives the
same run operation again — nothing of it exists, so no effect is repeated — after
1 s and then 2 s, at most
palletSandboxRefusalAttempts(3) times, and writesPallet run sandbox refused: assignment <id> generation <n>: attempt <k> of 3 (gRPC <code>).to stderr for each refusal. Before each unbound ADD, Pallet refreshes the run's preparation in a transaction against the currently completed network configuration. A configuration still settling refuses byattachment.reprepareExecution.applicationPending; a lease no longer admitted refuses by its endpoint check. No native descriptor is created in either case. Preparation can change only while no ADD journal exists; refreshing and ADD admission fence the same application lane, including an ADD whose commit acknowledgement was lost. An attached run still unbound after three refusals stayspending, observeduncertain, and writesPallet run waiting: assignment <id> generation <n>: network-attachment-pending.Its next reconcile, including after a process restart in the same boot, first proves both attachment and runtime absence before driving the same operation again. An unreadable runtime or attachment keeps it uncertain (network-run-unjoinedwith the owner's failure label). - An ADD bound it — the attachment owner journaled the pair, and answered it DOWN
because the node's barrier still did not hold after the ADD's own bounded wait,
or its raise failed. That binding is the run's for good: the attachment journal
is immutable history the DNS and ACTIVE owners verify, so no other sandbox of the
run is ever attached, and a re-drive could only be refused
(
attachment.reconcile.containerMismatch). The run is not driven again. While its operation is still pending, the bound pair is removed as that operation's rollback, as a DEL removes it, so a DEL that comes later finds it absent; the binding stays. The runtime's own DEL of the refused sandbox may have been answered before the ADD journaled, and a DEL after the run's operation ended was refuseduncertainfor good (Grasberg 2026-10-10). The node writesPallet run sandbox refused: assignment <id> generation <n>: attempt 1 of 3 (gRPC <code>): attachment-bound, rolled back, not driven again., orattachment-bound, rollback refused < <label>, not driven again.with the removal's owner failure label, and the CNI broker has already written the attachment's own reason (Pallet workload attachment down: run <run digest> ADD: <label>.). An isolated run's binding is not rolled back. - Whether an ADD bound it cannot be read. A second sandbox is never driven on a
guess: the run is not driven again, and the node writes
Pallet run sandbox refused: assignment <id> generation <n>: attempt <k> of 3 (gRPC <code>): attachment unreadable < <label>, not driven again., where<label>is the read's owner failure label (nothing on the read path reports it otherwise).
A refused isolated run, or an attached run whose ADD bound a sandbox or whose binding cannot be
read, ends exactly like a refused pull: its operation failed
(no evidence) with cause { kind: 'runtime-absent' } — the runtime holds no
sandbox or container of a run that should be running — the lane freed, and the
assignment observed failed once. Pallet never runs that assignment revision
again; the controller replaces the failed attempt with a new one, whose new run
digest gets a fresh sandbox and attachment. Before
this, the pending run held the lane until a reboot fenced it, and every other run
of the node waited behind it as lane-busy. Through Pallet 36.3.0 a barrier lost
for seconds after the gate failed the run this way: the ADD answered its pair DOWN
at once, the runtime tore the sandbox down, and all three re-drives were refused
attachment.reconcile.containerMismatch (production 2026-10-07).
A run whose own sandbox exists and is not ready — its network setup failed and its
teardown did not complete, so containerd keeps it NOTREADY, or it stopped — is
refused by the native run before any effect as SANDBOX_STOPPED
(PalletSandboxStoppedError). A sandbox never becomes ready again and an attempt
is never recreated after its first effect, so the refusal is definitive: the run
ends as a pending run whose sandbox stopped does in its join — recorded failed
(PalletAssignmentStore.failEndedRun), an attached run whose network epoch the
network owner reads as ended stating network-epoch-ended, any other no cause,
written as Pallet run failed: assignment <id> generation <n>: sandbox-stopped.,
the lane released, observed failed once — and the controller's stop and removal
take the sandbox away. Pallet 36.2.0 answered it RUNTIME_STATE_CONFLICT; the run
waited run-uncertain, held the node's one lane until a reboot, and every other
run and the removal of its own service waited lane-busy behind it (lab
2026-10-06 RUN 9, D11).
A run starts only while its network can take it. Before its first effect — the
registry credential, the volumes' transaction, the pull — an attached run asks the
network owner whether the workload datapath is up, by the admission gate's own
reading (PalletNetworkOwner.executionBarrier), and an isolated run whether the
network owner is ready. A refusal is a wait, never a failure: nothing of the run is
begun, the lane stays free, and the node writes
Pallet run waiting: assignment <id> generation <n>: network-barrier:<site>. with the
barrier's own admission.barrier.* site. Pallet 36.0.2 drove a run admitted before
a process restart while its new network owner was still raising the generation; the
CNI plugin answered Pallet network owner unavailable, the three re-drives were
spent in seven seconds, and a run with nothing wrong with it failed
runtime-absent. A barrier lost between the gate and the sandbox is waited out by
the sandbox's own CNI ADD, bounded (see the workload attachment section).
A running attached run whose pair belongs to a network epoch that ended fails for
good. The node process restarted and the router namespace went with it, taking the
sandbox's link, while CRI still reports the address the sandbox was given, so
nothing on the runtime side shows the loss. The execution owner asks the network
owner for the run's epoch on every reconcile of a running attached run, and gets a
closed answer: current, ended, none (nothing of the run is attached on this
node, which leaves nothing to fence) or unknown (not provable yet: the network
owner has not started, or the run's pair is not journaled, its add is not complete,
or it was removed). An unknown epoch is waited on, never taken as current: the
reconcile stays uncertain and writes
Pallet run waiting: assignment <id> generation <n>: network-epoch-unknown. An
ended one is recorded failed with cause { kind: 'network-epoch-ended' }
(the native lane must be free; it is fenced, not taken), written to stderr as
Pallet run failed: assignment <id> generation <n>: network-epoch-ended., and its
sandbox is stopped — retried from the durable failure until the runtime shows it
not running. The run is observed failed once, stating that cause; a new
attempt, the controller's, attaches in the current epoch. The controller's stop
and removal of the failed attempt run as for any other attempt, each with its
terminal receipt, so its slot advances.
A run whose sandbox stopped while its controller still asks it to run fails for
good as well. Nothing of this node stops a run under disposition run without
failing it first, so a stopped sandbox there — every sandbox after a reboot, or one
whose pause process the runtime lost — ended without a stop intent; it never runs
again, and an attempt is never restarted. Its completed run operation, of this boot
or an earlier one, is recorded failed (the native lane must be free; it is
fenced, not taken), the node writes
Pallet run failed: assignment <id> generation <n>: <cause>., and the run is
observed failed once; a new attempt is the controller's. An attached run whose
pair was attached in an earlier boot (a reboot ends every network epoch) or whose
epoch the network owner reads as ended states network-epoch-ended; any other
states no cause, written sandbox-stopped on stderr, because the controller's
vocabulary has none for it. The failure is certain, so an epoch the network owner
cannot prove yet only leaves the cause unnamed. Pallet 36.1.0 observed such a run
stopped under disposition run for good, which Cloudly never replaces: after a
reboot its attached workloads stayed down until their desire changed.
A run a reboot fenced before it completed ends the same way. The first start of a
new boot fences the run operation the earlier boot left pending (host-fenced):
that boot stopped every sandbox and ended every network epoch, and an attempt is
never run again. The run is recorded failed (the native lane must be free; it is
fenced, not taken), written to stderr and observed failed once — an attached run
stating network-epoch-ended, with its sandbox stopped should one still run, any
other stating no cause (sandbox-stopped on stderr) — and the controller's stop,
removal and next attempt follow. Pallet 36.1.2 and earlier left such a run waiting
operation-pending for good.
A launcher's WorkloadInit image. A launcher plan mounts the approved
WorkloadInit image as a CRI image volume at /opt/serve.zone/runtime-assets/workloadinit,
and containerd resolves an image volume only from its own store; it never pulls one,
and refuses the container while it does not hold the image. The run pulls it by the
plan's digest-pinned reference — code.foss.global/serve.zone/workloadinit@<platform manifest digest>, the release contract's registry, which serves it anonymously —
after the workload image and before anything of the attempt is created, without a
credential, and requires the stored image to list that reference. An image the store
already holds under that exact reference is used as stored and not pulled again: the
digest names its content and the pull proves no credential, so asking the registry again
would only let a registry that cannot answer fail a run whose image is on the node (lab
2026-10-05 RUN 6, O2). The workload image is pulled on every run, with the registry
credential the controller grants for it. Its registry
refusing it while nothing of the attempt exists ends the run like a refused workload
pull, image-pull-refused; with a sandbox already there, as when a run is resumed,
it stays a runtime failure. The node needs egress to that registry: the cluster
relay forwards only the controller's own registry.
A container's mount list. Before it starts a created container, and on every
later inspection, the native owner demands that the container's mount list, as
ContainerStatus answers it, is exactly the attempt's: the plan's secret binds,
the launcher's map and WorkloadInit image volume, then one bind per granted
volume, in that order. containerd 2.3 answers the list it stored at
CreateContainer, in the request's order, after resolving each image volume in
place: it sets the volume's host_path to its own mount point
(<state>/io.containerd.grpc.v1.cri/image-volumes/<sandbox id>/<manifest digest>)
and clears its uid and gid mappings, and echoes every other field as requested.
That host path is the runtime's answer, never Pallet's, so an image volume is
compared without it; every field Pallet sets — container path, image reference
and the rest of its image spec, sub path, readonly, recursive_read_only,
propagation, SELinux relabel, id mappings — and every bind's host path compare
exactly, so a container that differs in any of them, or carries a mount more or
less, or in another order, is refused. Pallet 36.1.4 compared the lists exactly,
host path included, and refused every launcher run's own container before its
start as foreign (OWNERSHIP_MISMATCH, lab 2026-10-04, N6). A refused mount list
answers MOUNT_MISMATCH (PalletMountMismatchError). Nothing changes a created
container's mounts, so the refusal is definitive for the attempt: the run, from
the drive that created the container or from the join of a drive whose reply was
lost, is recorded failed with no cause (the controller's vocabulary has none for
it), written as Pallet run failed: assignment <id> generation <n>: mount-mismatch.,
its lane released, and observed failed once; the container is never started,
and the controller's stop and removal take it away and release the mount as for
any failed attempt. Before, the run stayed pending and held the node's one native
lane, so every other run of the node waited lane-busy behind it.
Native failures by name. Every other failure of a native run, stop or removal
answers one of the native owner's closed codes, and a failed CRI call adds
errorData { call, grpcCode, cause }: the call (create-container,
pull-image, run-pod-sandbox, …), the gRPC status code and one cause word —
image-volume-unresolved (containerd does not hold an image an image volume
names), name-reserved (it still reserves the sandbox or container name for an
earlier request), deadline or other — never the runtime's message. The
execution owner names it as PalletExecutionRuntimeError:<code>:cri.<call>.<cause>.<gRPC code>
(executionRuntimeErrorOf) in the run's run-uncertain line.
A secret run's drive that ended short of its container. A run with secret
material publishes its mount under its pending operation before its drive, so a
drive that fails after that — the runtime created the sandbox and refused the
container, or a reply was lost — leaves the operation pending with the mount
published. Its next reconcile, in the same process or a later one of the same boot,
joins it: a fresh helper must find exactly the recorded mount, or the run stays
uncertain (secret-run-unjoined, or secret-run-unjoined < <owner failure label>
when the join failed, a runtime that could not be read for example) with every
record kept. A container that runs, or ran and exited, settles the operation with
that evidence. A sandbox that stopped never runs again, and the native run refuses
to take one up, so the run fails once as a completed run whose sandbox stopped does:
recorded failed with its lane released, written to stderr and observed failed,
an attached run whose epoch the network owner reads as ended stating
network-epoch-ended, any other no cause; Pallet 36.1.2 and earlier resumed it and
it waited run-uncertain until a reboot. Anything else is driven
again as the same operation, without opening the material again: the native run
takes it up from exactly what the runtime holds — the ready sandbox, a container
created and never started, or nothing yet — so no effect that exists is repeated,
and a refused pull or bound sandbox ends it failed as on a first drive. An attached run with no
ADD binding stays pending after its bounded sandbox retries, as above. Every failed
drive of a run is named, Pallet run waiting: assignment <id> generation <n>: run-uncertain < <owner failure label>. Pallet 36.1.2 and earlier settled only a
running container there, and a run whose container the runtime refused once waited
secret-run-unjoined until a reboot (lab 2026-10-04, N3), and then operation-pending.
A stop of a run this node never received. The controller can settle an attempt
whose run was sent but never delivered: its stop (generation 2) is then the first
revision of that attempt this node sees, and the shared decision
(evaluateRuntimeAssignment) admits no generation 2 without its generation 1. A
settlement revision repeats its run's immutable attempt and changes only its
generation, previous, disposition and digest, so the node rebuilds the run from
the stop and admits both revisions in one transaction only when the rebuilt run's
digest is the one the stop's previous names — each through the same decision as
any admission, and through the same slot, predecessor and session fences. Nothing of
the run ever began here: its stop finds nothing (absent), is observed stopped
with its terminal receipt, and the controller's removal follows as for any other
stop. A first-seen removal is never rebuilt, since its stop receipt is the node's
own. Pallet 36.1.0 refused such a stop as a siteless PalletAssignmentError:conflict
on every delivery, and the controller sent it again every 6.5 s for good.
Every conflict of the assignment store names its site,
assignment.<method>.<check> (TPalletAssignmentStoreSite), beside the execution
admission barrier's admission.barrier.*; PalletAssignmentError requires one for
conflict at the type level. An admission the shared decision rejects names its
reason: assignment.admit.predecessorMismatch, .stale, .revisionConflict,
.attemptMismatch, .dispositionRegression, .scopeMismatch or .invalid.
Failure causes. A failed observation states why the run failed in
cause (@serve.zone/interfaces TRuntimeAssignmentFailureCause, value-free)
whenever the node can name it: image-pull-refused with the classification and
gRPC status of a refused pull, network-epoch-ended for a run stopped because its
network epoch ended or an attached run whose sandbox a reboot stopped, exited with the container's exit status (0..255, or null
for any other) for a container that ended on its own, and runtime-absent for a
run the runtime no longer holds. image-pull-refused, network-epoch-ended and
runtime-absent are kept on the run's operation record (cause, the same
shape), so the observation states them whenever it is published, after a restart
too; a failed row 35.x wrote carries none, and its observation states none. A
cause explains a failure and changes nothing about it. A controller must read
@serve.zone/interfaces 32.40.0 or later (Cloudly
33.9.0 or later) before this Pallet runs against it: a receiver on 32.39.0 or
earlier refuses an observation that states a cause.
A run that waits before its pull says what it waits on:
Pallet run waiting: assignment <id> generation <n>: <step>., where the step is
lane-absent, lane-other-boot, lane-busy (another operation holds the native
lane), operation-in-flight (this run's own unfinished intent holds it),
operation-pending, operation-not-begun, secret-run-unjoined,
secret-run-unjoined < <owner failure label>, secret-mount-uncertain, secret-mount-uncertain < <owner failure label>,
storage-pending:<site>, registry-credential < <owner failure label> or
run-uncertain < <owner failure label> (the run's drive failed; a native failure
reads PalletExecutionRuntimeError:<code>[:cri.<call>.<cause>.<gRPC code>], see below). A new
step is written at once and a repeated one again with its count, under the same
repeat rule as refused controller requests. A staging failure names its chain: a
controller that refused the secret-material request appears as
PalletControllerRefusalError:<reason>:controller.refusal.getRuntimeAssignmentSecretMaterial,
where <reason> is the check the controller named in its refusal. A failure
before the run's operation began leaves nothing behind, so the next reconcile
requests the material again; one after it leaves the pending operation to the
next boot, which fails it, or — once the mount is published — to the join that
drives it again.
A reconcile that refuses outright — an isolated run with a TCP or HTTP readiness
probe (unsupported-readiness), an authority the node cannot resolve
(unresolved-authority), an assignment it cannot read (invalid), or any other
failure outside a named wait — is written
Pallet run refused: assignment <id>: <owner failure label>., once per assignment
and refusal: again only when the refusal changes or after a reconcile of the
assignment resolved. A refusal its wait line already names (storage-pending:<site>,
registry-credential < …) is not written twice. The controller learns none of
these from an observation: TRuntimeAssignmentFailureCause has no cause for a
refusal before any runtime effect, and a failed observation without one would only
have the controller replace the attempt — for unsupported-readiness with one the
node refuses the same way. Pallet
36.1.2 and earlier wrote nothing for them; the node runtime only counted them in its
pass status (rejected, unavailable, unresolved-authority).
The pull's credential reaches containerd as AuthConfig.server_address
https://<host[:port]>, the registry host of the reference it pulls.
containerd's CRI hands a credential only to the registry whose host equals
url.Parse(server_address).Host; Go reads a bare host:port as a scheme and
an opaque part with an empty host, so a bare address matched no registry and
every pull ran anonymously.
pnpm exec tstest test/test.execution.node.ts --verbose --logfile --timeout 60
The Unix fixture covers exact execution and recovery, lost mutation replies, foreign/ambiguous resources, image mismatch, degraded shutdown, never-started containers, malformed IPC, parent death and cancellation. These mechanism tests do not establish durable assignment admission or authorize a production cutover.
Image collection
Pallet's containerd holds only the images Pallet pulled for its own runs and the bundled
sandbox image. PalletExecutionOwner.collectImages removes the images no durable assignment
still needs. A collection is due when a lifetime starts and after every removal the owner proves.
The node runtime runs it at the end of a full assignment sweep, on the same native lane as a
reconcile, so nothing pulls or creates a container meanwhile.
The owner states what to keep. Every assignment whose runtime removal is not proven keeps its
run's image config (platformEvidence.imageConfigDigest), whatever its disposition: a stopped
attempt is still inspected and removed with its image, and an admitted one still pulls it. If
any retained assignment's image cannot be named, nothing is collected. The native pass
(collectExecutionImages, rust/src/execution.images.rs) also keeps:
- every image a CRI container still uses, in any state and matched by any name the image carries;
- every image the runtime reports pinned;
- the sandbox image by its bundled name.
It removes the rest in id order with CRI RemoveImage and requires ImageStatus absence
afterwards; an acknowledged removal that left the image in place fails the pass as
RUNTIME_STATE_CONFLICT. It reports the images listed, removed with their size, kept by reason,
left for the next pass, and the bytes the image filesystem reports in use.
| Bound | Value |
|---|---|
| retained image configs per pass | 1,024 |
| images the runtime may hold when a pass begins | 1,024 (refused as RUNTIME_STATE_CONFLICT beyond) |
| images removed per pass | 64; the rest are remaining and keep the collection due |
The node runtime runs the collector at the end of a full assignment sweep, and only while the
node holds an authenticated controller session: before the first one a node admits no native
effect, so a due collection makes no CRI call and runs on the first sweep after the session is
established. The collector runs only while the durable execution lane is idle on this boot, because an
operation still pending there may be a run whose pull the runtime is finishing. A pass that
cannot run stays due and is tried again after at most one minute
(palletImageCollectionRetryMs). It is recorded on the node status images as one of:
deferredwithlane-pending;failedwithretain-unresolved,retain-bound, orruntimeplus the native code.
A lane or assignment store that cannot be read is recorded the same way: an unreadable lane is
not proven idle and is deferred with lane-pending, an unreadable retained assignment is
failed with retain-unresolved. Besides its own refusals (no controller session, a busy lane,
an owner not ready), the collector rejects only when the identity runtime refuses its assignment
store or its own code fails, a defect the status cannot name. The node writes
Pallet image collection failed unexpectedly: <class> <code>[:<site>]. to stderr, which Spark
forwards to the node journal, once per distinct failure until a collection runs again.
It never fails the node: an image left behind costs disk, not correctness. No journal is kept: removing an unreferenced image is idempotent, and a later run pulls by digest.
pnpm exec tstest test/test.imagecollection.node.ts --verbose --logfile --timeout 120
Workload stats
A controller reads one resource sample of a running attempt with readRuntimeAssignmentStats
(@serve.zone/interfaces 32.36.0). The request names the exact current revision of an assignment
on this node, and the controller client binds it to the session and that revision before any native
call; like every inbound request, it acts only on the connection that holds the session.
PalletExecutionOwner.readStats reads outside the native lane, so a read never waits for or blocks
a reconcile. Each read runs its own native child (sampleExecutionStats,
rust/src/execution.stats.rs): it proves the attempt's exact CRI identity, calls CRI
ContainerStats and PodSandboxStats, and proves the same identity again. A change between the two
proofs is a conflict, never a sample of whatever runs now. An attempt without a running container is
answered not-running, without a runtime stats call.
A sample states the cumulative CPU time across every core (decimal digits, because it outgrows a
JSON number; utilisation comes from two samples), the memory working set, the memory limit the node
enforces (the config's memoryBytes), the sandbox's default interface counters (or null) and the
writable-layer usage (or null). A counter a JSON number cannot carry exactly is refused as
RUNTIME_STATE_CONFLICT.
| Bound | Value |
|---|---|
| reads per assignment at a time | 1 (another is refused as busy) |
| reads across the node at a time | 8 (palletObservationBounds) |
| native deadline per read | 10 s |
pnpm exec tstest test/test.executionstats.node.ts --verbose --logfile --timeout 60
Workload logs
Every attempt's stdout and stderr are captured by the runtime and delivered to the controller, which
keeps them (IRuntimeConfig.logs: SmartData owns log metadata, SmartBucket the payloads). The node
keeps no log history of its own: its capture files are disposable transport artifacts on tmpfs.
Capture. The sandbox names the attempt's capture directory
/run/serve.zone/pallet/logs/<run digest hex> and the container its file workload.log below it;
containerd writes the CRI log format there (<RFC 3339 time> <stdout|stderr> <P|F> <content>). An
attempt whose container was created without a log path (a container from an earlier Pallet) is
stated once per process as a not-captured loss.
Delivery. PalletLogCapture reads the file and sends reportRuntimeAssignmentLogs batches
(runtimeAssignmentLogContract, @serve.zone/interfaces 32.36.0) on the session, one at a time:
the next batch leaves only after the controller acknowledged the previous one, and an
unacknowledged batch is sent again unchanged. A CRI piece longer than one entry (64 KiB) travels as
partial entries; a line longer than the config's logs.maximumLineBytes is cut there and followed
by a line-cut loss of its stream, counting the cut bytes. Every delivery pass
(palletLogDeliveryBounds) sends at most eight batches per capture before the next capture's turn.
Controllers that keep no logs. A controller answers every batch accepted, replay or
not-kept. not-kept states that it keeps no workload logs and holds nothing: the node sends that
session no more batches (runtimeAssignmentLogContract.notKept, @serve.zone/interfaces 32.37.0).
The answer is also the node's permission to discard for that session (@serve.zone/interfaces
32.38.0), and Pallet uses it while that session stays the current, live one: every flush and every
hold on the delivery cadence discards each ended capture whole — its directory, its cursor and its
waiting batch — and each capture that ends later, and has each running capture drop its output
instead of holding it. A running capture drops every unread whole CRI line and a batch it sent that
holds output; it records the dropped batch's number with its cursor, so that number is never sent
again, not even by a restarted process, and its next batch opens with one not-kept loss counting
the dropped bytes and whole lines (loss evidence the dropped batch carried leads it too). A sent
batch that holds only loss evidence waits unchanged. The next session is asked again, and a
controller that keeps logs receives each running capture from where it stands, after that loss. Any
other failure of a batch keeps it the capture's next one and sends it again unchanged on the next
pass.
Refused batches. Each capture is its own delivery unit. A batch the controller answers with a
refusal stays its capture's next one and only that capture waits: the pass goes on with every other
capture, the refused capture waits out its retry delay on the delivery cadence (the delivery interval
doubled with each refusal in a row, at most 30 s, as every refused report does) and then sends the
same batch again, unchanged; flush() sends it at once. Nothing is dropped for a refusal: the
waiting capture stays within its buffer bound, so output it cannot hold is a counted
buffer-overflow loss, and an ended one counts toward the ended-capture bound (Ended captures
below), past which the oldest are discarded whole and counted. The refusal is written to
stderr, which Spark forwards to the node journal, as Pallet log batch refused: assignment <assignmentId> capture <captureId>: <label>., once and then with its count while it repeats, and it
stands in the pass's rejection until the capture delivers, ends or is discarded, or a new
controller connection is asked afresh. The contract names no refusal of a batch
(TRuntimeAssignmentLogsStatus is accepted, replay or not-kept), so Pallet cannot tell a
refusal that is final from one that is not, and treats every refusal as this capture's to retry. A
failure that is not an answered refusal — a timeout, a lost connection, a session this client no
longer holds, an answer that does not bind to its batch — still ends the pass's log delivery
(Cloudly log survey 2026-10-05: one refused capture ended every pass at itself and starved the
node's other captures).
Ended captures. A capture whose container ended waits for its final batch to be acknowledged,
and a crash-looping workload ends one per restart, so the node keeps at most
runtimeAssignmentLogContract.maximumEndedCaptures (64) ended captures awaiting acknowledgement,
holding at most maximumEndedCaptureBytes (64 MiB) together, each measured as the buffer bound
measures it: in the file bytes it holds for delivery (Bounds below). Each pass holds every ended
capture within its own buffer bound and rewrites its files to what it still needs before it is
measured, so a capture at the largest buffer (64 MiB) always fits whole. Past either bound — a controller that is down, fails every batch or leaves the method
unhandled — it discards the oldest ended captures whole, oldest by when this process saw them end.
Every capture discarded whole, under this bound or a not-kept answer, is counted in the
discardedCaptures of the next batch the node creates, of any capture. The count is persisted
(pallet_log_discards): it survives a restart, travels unchanged with the batch that carries it on
every retry, a restarted process creating that batch again with the same number and count, passes
on to the next batch when that batch is discarded unacknowledged (dropped under not-kept, discarded
with its capture, or lost to a restart that cannot continue its capture), and is kept while a
not-kept answer stands. The owner keeps in memory only running captures and ended ones within the
bound: a capture that finished or was discarded leaves no directory, cursor or entry behind.
Rollout order. A controller must answer reportRuntimeAssignmentLogs before a Pallet with log
capture runs against it. A controller that leaves the method unhandled fails every batch: the node
stays within its bounds, but it retries the same batch on every flush and drops the output beyond
the buffer bound, so no workload output reaches such a controller. Cloudly answers not-kept
(since 33.8.0). Onebox has no Pallet log backend yet (Onebox 33.1.1): run this Pallet against
Onebox only once one ships.
A controller must also run @serve.zone/interfaces 32.38.0 or later before this Pallet runs against
it. A receiver on 32.37.0 or earlier reads a batch by its exact schema and refuses one that carries
discardedCaptures or a not-kept loss, on every retry: against it the node stays within its bounds
as against any failing controller, but that capture's output no longer arrives. Cloudly 33.8.1 is
the first Cloudly on 32.38.0 (33.8.0 is on 32.37.0); an Onebox Pallet log backend must be on
32.38.0 when it ships.
Bounds. What the controller has not acknowledged is the capture's buffer, at most the config's
logs.maximumBufferedBytes. Pallet measures it in the file bytes the capture holds for delivery,
which is what sits on the tmpfs: the file bytes the batch waiting for its acknowledgement spans, plus
the file bytes not yet read into a batch, CRI headers and the cut rests of lines beyond
logs.maximumLineBytes included. The waiting batch counts with its whole span whether or not a
reclaiming pass (below) already removed those bytes from disk. A batch never spans more file bytes
than the bound (its first CRI line is always taken), so the waiting batch is always kept and sent
again unchanged; beyond the bound the oldest unread whole CRI lines are dropped and stated where they
were dropped as a buffer-overflow loss, which counts their exact output bytes and whole lines and
leads the next batch. Output of short or cut lines therefore reaches the bound sooner than its
output bytes alone would: a line cut to a few bytes still counts every byte the runtime wrote for
it. Drops while a batch waits extend that one loss, which keeps the
time the loss began, so however long the controller stays away the evidence is a single entry.
Once the node has read past the buffer bound (never below 1 MiB)
it renames the file aside and has the runtime reopen the log path (CRI ReopenContainerLog through
the native reopenExecutionLog, which proves the attempt's exact identity before and after); the
renamed file is drained first and removed once the cursor moved past it. On disk an attempt
therefore holds at most the read part below the rotation bound plus its unacknowledged buffer.
The runtime reopens only a running container's log, and its answer ends a capture only on definitive
evidence that nothing writes the file any more: the container exited, its sandbox stopped, or it was
removed. A container created but not yet started, or one in an unknown state, may still write:
the reopen is refused as indeterminate like a runtime that could not answer, so the renamed file
gets its name back, the capture keeps running and no rewrite touches its files, and a later pass
asks again.
A rotation survives a Pallet process that stops inside it. Between the rename and the runtime's
reopen, the renamed file is still the one the runtime writes, and on disk that window is a renamed
file (workload.log.1) without a fresh workload.log. Every pass that finds it so asks the runtime
to reopen before it reads on, unless the container ended, and every step that rewrites or removes the
renamed file completes the rotation first, so none ever replaces or unlinks a file the runtime still
writes: output written there would otherwise be neither delivered nor counted, and would grow unseen
on the tmpfs. When the runtime cannot answer a reopen, the renamed file gets its name back by a hard
link, never a rename over the name, so a fresh file the runtime created meanwhile (a reopen whose
answer was lost) is never replaced; a process that stops between the link and the removal of the
renamed name leaves one file under both names, and the next pass removes the renamed name. A
reclaiming rewrite that stopped before its rename leaves a copy (workload.log.1.compact) that the
capture removes when it opens.
A pass that had to drop also reclaims the read part, which the waiting batch no longer needs on
disk: it rotates the file at once and rewrites the renamed file from the first unread byte, so after
that pass the attempt holds on disk only its unread bytes, within the bound, however long a batch
waits. The bound holds whatever the controller does: every delivery pass keeps the captures it
reaches within it, a pass whose batch fails holds every capture before it ends, and the controller
client holds every capture on the delivery cadence (deliveryIntervalMs) whether or not a
controller is connected, so a controller that is down, fails every batch or leaves the method
unhandled never lets a capture grow. A hold that fails writes Pallet log capture hold failed: <class> <code>[:<site>]. to stderr, which Spark forwards to the node journal, once per distinct
failure until a hold succeeds. The capture's delivery, hold and reconcile run one at a time.
A capture's own failure — its files, or a runtime that cannot answer its reopen or cannot yet tell
whether its container will write — is that capture's alone: the delivery pass or hold goes on with
every other capture, the failed one is tried again on the next pass, and the pass reports the first
such failure once it is done (a delivery pass rejects, a hold writes the line above). A controller's
failure that is not a refusal of one batch still ends the pass's log delivery, after every capture
is held (Refused batches above).
An ended capture is never written again, so it keeps on disk only what it still needs: the file
bytes of its waiting batch and the file bytes not yet read, CRI headers and the cut rests of long
lines included. Once its files hold more beyond those bytes than those bytes and 64 KiB
(palletLogCaptureBounds.settleSlackBytes), it rewrites them into one file (workload.log.settled,
renamed over workload.log), checked when it ends and after every acknowledgement. The cursor moves
to the rewritten file in one transaction before the rename, so a restart at any point continues from
exactly the recorded byte. An ended capture therefore holds on the tmpfs at most twice the file bytes
it still needs plus 64 KiB, instead of every byte the running container left behind. The bound is
amortized on purpose: rewriting after every acknowledgement would hold only the needed bytes, but
would copy what is left once per batch, quadratic in the capture's size (about 2.3 GiB of copying to
drain one 64 MiB capture in 900 KB batches). Rewriting only once the files are at least twice what
is needed copies each byte a bounded number of times, never more than was acknowledged or dropped
before, so draining a capture copies at most its own size. Since the ended captures measure at most
maximumEndedCaptureBytes together in the file bytes they need, after every pass they hold on the
tmpfs at most 2 × 64 MiB + 64 × 64 KiB = 132 MiB (138 412 032 bytes) together.
Durable cursor. Each capture keeps one record in the node's own database
(pallet_log_cursors): the capture, its last recorded sequence, the position after it (the kernel
identity — device, inode, birth time — of the file it points into, the byte offset and the line state
of both streams), the evidence that leads the next batch, and the waiting batch. A batch is recorded
when it is created, before it is sent: where its file bytes end, and every entry its creation
generated with its timestamp — the losses it opens with and the empty entries a final batch closes
open lines with — and the discard count it states. The cursor moves on only after an acknowledgement
is recorded. A restarted Pallet process continues the same capture from exactly the recorded byte,
mid-line included, and creates the waiting batch again byte for byte from those file bytes and the
recorded entries, so a sequence number once sent is never sent with other content. Evidence that
leads a batch holds one loss per reason, however many passes drop before it is sent. Only a
cursor whose file is gone — a reboot emptied the tmpfs, or a reclaiming pass rewrote the file
since the last acknowledgement — starts a new capture with a capture-restart loss of unknown
extent. When the container stopped, failed or was removed, the
capture drains the rest, closes a line left open, sends one final batch, and once that batch is
acknowledged records the capture finished, then removes its directory and then its cursor; a
restart between those steps completes the removal and never counts the delivered capture as
discarded. A capture discarded whole first, under a not-kept answer or past the ended-capture
bound, loses both as well. Its record goes first, so a restart before its directory is gone leaves
a directory without a record; the attempt's assignment stays listed (the assignment store keeps
every attempt's state as a tombstone), so the next process opens that directory as a new capture of
the ended attempt, delivers what it holds and removes it. The first pass never removes a directory
merely for lacking a record: a running capture has none until its first batch is created, and a
directory whose attempt no state names proves nothing about its writer (a database restored from an
older copy while the container runs), so such a directory stays until the reboot that empties the
tmpfs. A batch dropped under
not-kept records the cursor too, at the dropped number and past the dropped output, together with
the evidence the next batch opens with, so a restarted process skips the same number and states the
same loss. The first pass of a process forgets every cursor whose capture directory is gone (a
reboot emptied the tmpfs) and passes on the count of every carrying batch no open capture creates
again; a cursor recorded finished goes uncounted.
pnpm exec tstest test/test.logcapture.node.ts --verbose --logfile --timeout 60
pnpm exec tstest test/test.logowner.node.ts --verbose --logfile --timeout 120
Private workload readiness
PalletExecutionOwner binds the published runtime config policy to the exact
admitted assignment before invoking a native process, TCP or HTTP readiness probe.
Probes inspect the complete CRI identity before and after network IO and connect
only to the evidenced container IP addresses, from inside the sandbox's own
network namespace: the host has no route to a workload, by design. The execution
owner reads that namespace from the run's attached network journal (the pair's
sandbox namespace), only while the journal is live and names exactly the CRI
sandbox the evidence shows; a running run without one is refused as
unresolved-authority. An isolated run has no such namespace: its readiness must
be the process's, and a TCP or HTTP policy for it is refused
(unsupported-readiness). The native opens the journaled path once,
requires a network namespace of exactly the journaled device and inode, and
enters it on a dedicated thread that runs the probe on its own runtime and ends
with it; another namespace at the path fails the sample namespace-mismatch, an
unopenable one namespace-unavailable. An attached run's readiness also requires its
workload pair raised, whatever its target: inside the sandbox namespace the native
reads the pair's sandbox interface (named by the attachment receipt) and requires it
up and running and the namespace's default route through it, before any target
probe; otherwise the sample fails link-down. A listener on the workload's own
address answers over a link that is down, so a probe alone never saw a lost
attachment: Pallet 36.0.2 reported a workload ready for 51 minutes while its pair
was DOWN after a crash restart. An isolated run's process target needs no namespace. HTTP uses the configured Host and
request target, checks final response headers without reading a body, and never
follows redirects. HTTPS verifies system trust and the configured hostname/SNI.
There is no insecure TLS or DNS-target fallback.
SmartData stores probe evidence, exact container start nanoseconds, boot ID, threshold counters and the next
due time. Linux CLOCK_BOOTTIME provides ordering across native children and daemon
restarts. Initial delay and startup grace use a conservative anchor from the first
completed native observation of that exact runtime identity. An already running
container therefore waits the configured delay when first discovered; a forward
wall-clock step cannot shorten it. Thresholds and intervals come from the immutable
config. Changed runtime identity, boot, config or clock
regression resets readiness; pending native effects supply no readiness proof.
Probes never overlap. A reconcile before the next due time returns
readiness-pending with no new observation.
A ready observation must match the durable sample's assignment, timestamp, image, sandbox and container. Its outbox transaction fences the current native lane. Restart and lost acknowledgements retain the exact existing observation; they cannot reuse a cached probe to create a new ready report. This remains a private composition. Authenticated controller transport and production route promotion are separate integration requirements.
Qualification uses the actual native probe, real HTTP/HTTPS listeners, a controlled
CRI server and disposable NoSQLDB stores. The focused tests are
test.readiness.node.ts, test.readinessstore.node.ts and
test.executionowner.node.ts; they do not establish new real-containerd or host
power-loss qualification.
Private packet policy composition
The backend-private composePalletPacketPolicies maps a validated network
projection and complete current attachment/handoff receipts into a combined
Smartnftables router policy and its host-transit policy. It captures inputs before
asynchronous validation. Attached local workloads receive exact veth sources;
selected remote workload grants require an explicit caller-owned TUN. Reserved
but unattached local workloads receive no packet paths. Private DNS uses only each
workload's gateway and explicit TCP/UDP port 53 rules.
Workload egress comes from the published projection grant helper. Router-origin DNS and platform traffic comes only from explicit signed router selections, with the current handoff's transit source and exact selected destination tuple. Public workload grants grant no router authority. Both policies carry the complete protected union and use only the current handoff's leased transport ranges. Withdrawals grant no traffic; DNS readiness does not change packet authority.
Published host ports come only from the verified projection's endpoints placed
on this node. Every entry keeps its Cloudly authorization reference; hostIp
is the node uplink address or absent; the workload must be attached with a
current receipt; a port held by a platform endpoint or resolver on the uplink
address, a port inside a leased SNAT range, or an entry beyond the bound refuses
the endpoint's whole set by a bounded publication.* site and blocks nothing
else. An entry may publish a contiguous range (hostPortEnd, each host port to
the same port inside the workload) and may be symmetric: outbound flows the
workload opens from its published port(s) leave from the same port(s) on both
hops, so a SIP or RTP peer sees one address and port in both directions
(@serve.zone/interfaces 32.31.0; pinned at 32.32.0, which also refuses a symmetric entry whose inside port another entry of its protocol shares). Every port of a range counts: a range that
reaches a leased SNAT range or a platform endpoint's port anywhere is refused like
a single port. The contract admits a symmetric entry only on an endpoint with
public egress and refuses the whole projection otherwise, before Pallet composes
anything. The bound is the contract's 64 entries per generation, a range counting
as one, taken over the node's signed publication list in lease id order: an
endpoint counts against it once the projection itself admits its set, whatever the
node later refuses of it (uplink, attachment or compiled capacity), so the bound
follows from the projection alone and binds the allocation-pool guard as well (see
below). An endpoint that would pass it is refused publication.capExceeded. The accepted set is handed to the host-transit policy as
hostIp:hostPort[-hostPortEnd] -> transitAddress:hostPort[-hostPortEnd],
journaled with the ACTIVE generation and retired with it. Leased outbound
translation keeps using the handoff lease's own source-port ranges; Cloudly
allocates new leases in 49152-65535 (runtimeNetworkHandoffSnatPortRange), so
publishable ports never meet them, while a lease allocated earlier keeps its
range and still refuses any publication inside it. Pallet's session registration
offers the installed interfaces release, so Cloudly sends range and symmetric
entries only to a node that reads them. The native compiler decides what fits its atomic
budget: while it refuses either policy as EXHAUSTED on a bound that
publications spend (their count, or the rule bytes, target bytes or operations of
the batch), Pallet refuses publishing endpoints by name
(publication.capExceeded), last first, and recompiles both. A policy the
compiler refuses as INVALID input, or as EXHAUSTED on any other bound, is a
composition defect that no refused publication mends: the generation fails with
PalletPacketPolicyRejectedError (invalid:packet.policyInvalid or
exhausted:packet.capacityExhausted), which carries the engine's bounded reason
and, for EXHAUSTED, the bound with its limit and actual count. The router policy carries the second hop of
the same accepted set, transitAddress:hostPort[-transitPortEnd] -> workloadAddress:targetPort with the same symmetric flag;
it is composed from the same verified entries, cross-checked against the journal
and applied by the native compiler, and the raise owns the sandbox default route
the workload answers through: CNI ADD installs default via <router address>
(static, metric 100) in every pod namespace and advertises exactly that route in
its result. Proven end to end in the DHCP-hub packet
qualification: an external client reaches the workload over TCP and UDP through
both translations, the workload sees the client's own address, the client sees
the published uplink address back, a refused entry never opens, and a successor
generation without publications leaves both ports dark.
The host hop translates a publication to the router's transit address, which lies
in the transit pool the node's allocation-pool guard denies by current destination
in every hook, and by current source on the reply: without an exception the guard
drops every published flow in FORWARD, on a docker-shared and an exclusive
host alike (lab 2026-10-06 RUN 9, D12). The guard therefore carries the node's
publications as its exceptions (allocationPoolGuard.publishedPorts,
@push.rocks/smartnftables 4.4.0). Both hops take their entries from one signed
publication list, signedPublications in ts/network/packet/publication.schema.ts:
every endpoint the verified projection places on this node that publishes ports,
with the site the projection itself refuses its whole set at
(publication.unauthorized, publication.snatOverlap). composePalletPublications
applies what only the node knows on top of it for the host hop, and
composePalletGuardPublications takes every entry of the list the projection does
not refuse, at most the 64-entry bound and so far within the guard compiler's 1024, in the host hop's translated form (one transitTarget): to the
handoff's router transit address with the host hop's port. The guard admits each
only translated from outside every pool, with its replies; it checks no source, link
or uplink address, so the host hop's translation, bound to the uplink, is what lets
a flow in. For one projection the guard holds every entry the host hop translates,
and beyond them only the entries of an endpoint refused for the node's uplink,
attachment or capacity, to which nothing is translated (test.guardpublications
asserts the equality). A new projection's publications reach
the guard with its next transition, like its host grants. The docker-shared guest
(test/native/dockershared.ts, run by qualify-active-registry.py --docker-node-binary … --docker-bundle …) composes one signed projection's publications for both the host
hop and the guard and proves, on the docker-shared host under Docker's FORWARD
DROP with the contribution and on the exclusive host: TCP, UDP and both ends of a
UDP range are delivered with no guard (docker-shared), dark under the guard as Pallet
36.2.0 composed it, and delivered once the guard's transition adds them, the workload
seeing the external client's own address and the client the uplink address.
Host grants let the node's own host namespace — Onebox's reverse proxy on a
single host — dial an exact workload port directly, with no loopback publication,
no route_localnet, no loopback DNAT and no loosened martian filtering. They come
only from the contract's getRuntimeNetworkProjectionHostPacketGrants over the
verified projection, so only a onebox projection's signed host selections
grant any; Pallet proves each again against that projection — the source is the
current handoff's transit host address, the destination the exact lease address
and port of an endpoint on this node — and refuses the composition otherwise.
The same exact tuples are compiled into all three tables the flow crosses: the
host-transit and router policies of the ACTIVE generation, for attached workloads
only, and the allocation-pool guard, where they are its only exceptions. A new
projection composes a new generation and guard target, so a withdrawn selection
is an ordinary atomic replacement; an absent or empty set leaves every policy
byte-identical.
A granted flow needs the host to route the lease address through the selected
handoff, and the ACTIVE generation owns that route. A generation whose projection
selects host flows plans one host route per workload pool that holds a selected
lease (hostRoutePrefixes, derived again from the journaled projection on every
read): the pool via the router's transit address, sourced from the host's transit
address, so the flow leaves with exactly the source its grant names. It is added
through a command socket in the host namespace inside the registered host
transition that raises the link, deleted inside the one that lowers it, checked in
every inventory of the raised link, and removed by recovery when it outlives a
fence. Any other source's change to a route touching the host link (through it,
or over its transit prefix or pool routes) still retires the uplink observation, so
it is never accepted as the generation's own. The route widens nothing: the guard
and both hops still admit only the exact granted tuples,
and a generation without host selections — every Cloudly generation — plans none
and is byte-identical.
Workload ingress grants let the cluster ingress workload reach the port a target
workload listens on directly, instead of a hairpin through a published uplink port.
They come only from the contract's
getRuntimeNetworkProjectionWorkloadIngressPacketGrants over the verified
projection, so only a cloudly projection's signed workloadIngress selections
grant any; Pallet proves each again against that projection — two different
endpoints, at least one placed on this node, the exact lease address of each and
the target's listening port — and refuses the composition otherwise. They compile
into the router policy's workloadGrants (@push.rocks/smartnftables 4.0.0): a
stateful one-way flow of the exact protocol and port, answered only by the
replies of a connection the ingress workload opened, never translated, and never
open in the other direction. A grant compiles only when both of its workloads are
attached behind this router in the generation; an unattached end, or a target on
another node, receives no packet authority here, because the compiler has no
one-way grant across the tunnel. Large port ranges such as RTP media stay uplink
publications. An absent or empty set leaves the router policy byte-identical.
On an exclusive host (see Host forwarding modes) every host-transit policy a new
generation composes carries exclusiveForwarding (@push.rocks/smartnftables
4.1.0): the host's IPv4 forwarding is on only while that generation is UP, and the
table's forward chain ends in one unconditional drop after the handoff,
publication and symmetric flows, so a LAN peer, a hairpin through the uplink or
traffic between two other links is never forwarded. On a docker-shared host the
policy is shared, and Docker's FORWARD policy DROP with the node's DOCKER-USER
contribution does that work. The member is the composition's explicit
exclusiveForwarding binding; a generation is recomposed with the value its
journaled host policy carries, so one an earlier release journaled without it is
recomposed byte for byte, fenced and released, never raised again.
A projection's hostPlatformEndpointIds names the protected platform endpoints
this node's own host serves — on a single host the cluster hub, the relay
listener and Corestore. They compile into the host-transit policy's
localPlatformEndpoints: the host delivers leased flows to them in its INPUT
instead of forwarding them, from the exact handoff, with the lease's transit
source address and a source port of its range, and only the replies go back.
The workload and router selections that reach them are unchanged: a workload
still reaches only the endpoints it selects and the router only the hub its
block selects. Every address a declared endpoint holds is added to the uplink
binding, so the compiler verifies at apply, recovery and inspection that the host
holds it; each must be a permanent /32 of the uplink, which the uplink
observation admits beside the DHCP lease (see Retained uplink observation). An
undeclared platform endpoint or a resolver on one of those addresses refuses the
composition as conflict. Without the member the uplink binding and the host
policy are byte-identical.
This function performs no native apply, link activation or durable journal write. Its caller must authenticate the projection, retain and fence actual namespace, attachment, TUN and uplink generations, and own routes, DHCP changes and SNAT address lifetime. Native preparation must confirm graph capacity before a complete transition is journalled and applied. Composition alone provides no enforcement, workload readiness, allocation reuse or packet-drain evidence.
Private DNS lease arithmetic
The backend-private ts/network/lease.ts converts an authenticated lease of at
most 15 minutes to a conservative native CLOCK_BOOTTIME deadline. Its input is
a verified UTC interval anchored to the same boot clock, with an independently
qualified rate-error bound. It accounts for uncertainty and the fastest permitted
UTC progression. Process recovery retains the exact deadline; persisted boot-time
and UTC lower bounds reject clock regression.
A clock this node cannot read lapses that DNS acquisition and never fails the
lifetime that owns the clock, and which lapse it is follows from what the clock
states. A sample it cannot qualify leaves the acquisition without fresh time
evidence: it continues on the window already proved, keeping the exact deadline
this boot converted, and lapses only where there is no such deadline — another
boot, or a window this boot has not converted. A clock that states no boot
reading at all — unqualified, or lost with its native owner — is a
clock-unavailable acquisition outright: that pass withdraws every name and the
next pass binds them again. No reader is refused for asking while another holds
the clock; losing the owned process still fails the node through the clock's own
failure signal.
Reboot recovery requires newly verified independent time evidence for that boot. The original absolute expiry stays unchanged, and a saved Cloudly timestamp cannot renew it. The returned anchor, deadline and high-water state must be persisted through SmartData inside the complete authenticated projection before applying a newer native DNS revision. Expiry does not revoke packet grants or prove that an old writer is absent.
The private PalletIndependentClock owns an authenticated NTS measurement process
through the native executor. Its pinned Chrony 4.8-servezone2 variant returns exact
current Unix time and tracking metadata from one observation. Legacy Chrony
tracking offsets cannot supply this precision when the RTC is years wrong. The
private process never adjusts the host clock or loads host Chrony/GnuTLS settings.
It uses bundled trust certificates, fixed Ubuntu NTS peers, no drift/cookie files,
and a root-private Unix socket. One native operation runs at a time, and no caller
is refused for asking second: a caller that asks for a reading already in flight
is answered with that reading — a reading is of an instant, so sharing it is exact
and the owner is asked once — and a caller that asks for the other reading waits
for the operation in flight and then takes its turn, in arrival order. Loss of that
owned process fails the node lifetime; unqualified or offline samples supply no
time authority.
First qualification must not inherit long outage backoff from attempts made before startup connectivity is ready. The owned Chrony variant caps unresolved-source retry at 28 seconds until the first source resolves, and both connection- and TLS-class NTS-KE retry at 16 seconds until the instance first receives authenticated exchange data. Upstream exponential backoff resumes after those successes, including after cookies are later exhausted or the instance resets. Sampling continues to require every qualification check below; an unqualified reading grants no lease extension, and an expired native or VPN lease is never revived. The next network pass retries qualification and renewal without changing the live lease.
The owner checks coherent tracking, selection and NTS authentication reports, root error, freshness and kernel clock brackets. Supported KVM Linux clock profiles bound BOOTTIME progression by 400,000 ppm; unsupported clock sources, PPS/custom tick settings, VM pause, snapshot restore and live migration do not provide a qualified clock. The bootstrap certificate's validity interval is 1970–2100. A wall-clock reading or an NTP-synchronised flag alone is insufficient.
Clock ownership does not persist projections, supervise DNS or admit workloads. The DNS persistence and lifecycle owner remains unfinished. Lease arithmetic checks cover uncertainty, deadline boundaries, process recovery, reboot proof requirements, rollback and malformed stored state. The isolated native clock harness additionally checks actual NTS measurement, wrong-year RTCs, offline denial and joined Node/Deno cleanup; only recorded successful runs qualify the specified source and kernel.
Private network reservations
The backend-private ts/network/allocation/ store retains protected-authority
revisions and immutable handoff leases through the identity runtime's existing
SmartData/NoSQLDB connection. runNetworks() exposes callback-scoped operations;
shutdown joins admitted operations before closing the engine. Every write,
including replay, requires an owning-code identity fence inside its transaction.
The caller authenticates controller authority before entering this interface.
A fresh node can stage the complete current authority. An initialized node must
advance through exact consecutive references in the same controller epoch and
node scope. Reservations reference the current authority and validate against
all retained leases; a common revision fence serializes competing allocations.
Historical leases continue to bind their exact authority. Quarantine preserves
the lease and forbids reuse of its source-port range, conntrack zone and label
within the shared contract's respective scopes. Replay retains quarantine.
listHandoffs() returns the complete bounded retained history; there is no
expiry, pruning, release or reuse operation.
This stores allocation intent. It does not authenticate signed projections, activate pools, acknowledge host/router barriers, declare flow drainage, allocate sandbox IPs or admit packet traffic. Those owners must bind this durable history to exact native receipts before activation. Tests use actual NoSQLDB to check concurrent conflicts, transactional identity fencing, stored digest corruption, lost commit acknowledgements, complete restart recovery and joined shutdown.
Private signed network admission
PalletControllerClient receives the published Interfaces 30.9.0 signing-authority
and complete network-projection RPCs on its private outbound TLS connection.
The current physical session, active credential and controller/node/namespace
scope fence each admission transaction. Socket input is inert and bounded; the
transport accommodates the full 896 KiB projection contract.
ts/network/projection/ persists public signing trust, immutable public revision
history, one complete signed projection and bounded immutable workload leases
through SmartData on the existing NoSQLDB owner. Key rotation and projection
admission write a common revision fence. Rotation or revocation prevents the old
envelope from being read as current authority; historical keys retain verification
of recovery material. Re-signing an identical projection under the new key replaces
its envelope before acknowledging replay. Replays never renew the DNS deadline.
Projection admission retains historical protected authorities and handoffs in the same allocation transaction. Full local and remote workload leases survive withdrawal and restart, including the addresses needed for later denial. Retired leases stay quarantined; their subnets and execution attempts cannot be reused. The store retains at most 512 workload leases and rejects capacity exhaustion. Admission carries pending withdrawals and tombstones until the node's own durable application receipt names exactly the previous projection (Projection application receipts); then it admits a successor that drops them. An ACK establishes durable intent only.
The callback-scoped runNetworkProjections() facade expires with its callback;
shutdown joins admitted work. inspectRetainedProjection() is recovery material,
while readCurrentProjection() requires the current signing key. Neither method
establishes an application, time, or current-identity fence for a caller. Pool
allocation, CNI, native realization receipts, network activation, qualified DNS
time and live worker rollout remain separate owners.
The private application-store composition reads current verified signing intent and the complete retained handoff/workload history in one SmartData transaction. Before recording native intent, it rechecks that source fingerprint and writes both common trust and allocation fences in the caller's transaction. Concurrent key rotation, projection admission, reservation or quarantine therefore conflicts with stale intent; a failed identity guard rolls back both writes. This interface does not issue a native receipt or expose an application capability over RPC.
Tests test.networkapplicationsource.node.ts, test.networkprojectionstore.node.ts and test.controllernetwork.node.ts
exercise actual file-backed NoSQLDB and TLS sockets, including concurrent rotation,
identity rollback, lost ACKs, physical reconnect, shutdown and a complete namespace
with 12,288 DNS names. They establish admission behavior, not live packet policy.
Private DNS lease renewal
A renewal carries a fresh DNS window for one exact admitted projection and no
applied network state. PalletControllerClient receives
applyRuntimeNetworkDnsLeaseRenewal on the same outbound connection as the two
admissions, and the envelope's session must be the binding this node obtained
itself. That comparison is what makes a cluster relay's push admissible: the
relay holds signed statements for a node it carries and adds no trust of its own.
Admission is one transaction in ts/network/projection/, in this order: the
stored signing trust, this node's current projection, the signature under that
trust and under the revision that signed the projection, the renewal this node
already holds for that projection, the published admission decision against the
qualified clock's lower bound, the write, and the common trust fence every
network admission takes. accepted replaces the single stored row, replay
changes no renewal and repeats the same bound acknowledgement, and every other
outcome is the refusal the other two admissions give; the trust fence is taken on
both outcomes, so a concurrent rotation conflicts rather than interleaves. What
persists is the complete signed envelope, so every later read re-proves it against
the retained signer rather than trusting the node's own table. An envelope this
node can no longer prove is a lapsed renewal, not a broken node: the DNS lane
discards it and converts the projection's own window, while a credential bind
refuses. The persisted DNS lease and DNS intent each state the renewal their
window was converted from; no released version persisted either shape, so there is
no migration, and a store carried over from an older build fails its assertion on
every read and must be started from clean pallet_dns_lease and
pallet_dns_intents collections.
A renewal belongs to one projection envelope: that projection's exact reference and the revision that signed it. The transaction that admits any envelope the stored renewal does not belong to discards it — a successor projection, and equally the same body re-signed after a key rotation, which keeps the reference and changes the signer. A renewal of a projection this node does not hold, at a sequence it has already passed, with different bytes at a sequence it holds, or whose slot has not provably opened at the qualified reading is refused; the sender simply delivers it again once it has. A rotated signing key renews nothing: the re-signed projection discards the renewal held and the node lives on the window that projection itself signs. While that window still runs nothing else changes; once it has ended the node withdraws DNS and stays dark until the controller delivers a renewal signed by the new revision. Pallet asks for none — a renewal arrives on the connection this node already holds — and converts the one it admits on its next DNS pass, which runs once a second.
Everything window-bounded then follows the effective window — the admitted
renewal's when one renews this projection, the projection's own otherwise. The
DNS lease converts that window in the same transaction that writes the lease and
records the renewal it converted, the composed DNS snapshot proves its
validUntilBoottimeMs against it, and the managed-VPN credential may expire
inside it and never beyond it. Nothing else moves: the application, its
fingerprint, the native generation and the protection receipt are untouched, and
a node that loses qualified time or current signing authority still withdraws
DNS exactly as before.
test.networkdnsleaserenewal.node.ts, test.dnsrenewalview.node.ts,
test.dnsleasestore.node.ts, test.independentclock.node.ts and
test.controllertunnel.node.ts state admission, supersession, replay, each
refusal, the boundary at an equal expiry, the moved deadline, restart recovery, a
stored renewal this node can no longer prove on either lane, a reading two
overlapping callers share, a clock this node cannot read, and the bound the
credential keeps.
Private identity store
The separate backend-private ts/identity/ implementation owns Pallet-only
credential state in an explicitly supplied SmartData database. It is not imported
by the public probe facade and has no Cloudly connection. Its protected lifecycle
and node-local enrollment control owners remain private and unwired from the
production daemon.
Preparation commits a 32-byte cryptographic palletToken before returning its
SHA-256 proposal. Binding checks this owner's durable hash and the complete
enrollment digest. Activation requires the exact bound acknowledgement and writes
its immutable receipt in the same transaction. Initial adoption may name an
existing Cloudly node while Pallet is still generation zero. Rotations retain the
old active bearer until acknowledgement; historical receipts cannot reactivate it
or replace pending material. Inspection, proposals and receipts contain no bearer.
The enrollment-specific helpers use published Interfaces 28.2 snapshots, digest and acknowledgement binding. Those helpers prove content binding, not remote authentication: the eventual authenticated transport/coordinator owns that trust boundary. Spark credentials never enter Pallet's private store.
Tests use disposable file-backed NoSQLDB 10.5.1 instances and prove exact replay across engine restarts, concurrent preparation, acknowledgement rollback, independent database bindings and rejection of changed or malformed input:
pnpm exec tstest test/test.identitystore.node.ts --verbose --logfile --timeout 60
Private identity runtime
ts/identity/classes.identityruntime.ts owns a Linux file-backed NoSQLDB 10.5.1
engine and its SmartData connection. Production defaults to UID 0 and
/var/lib/serve.zone/pallet; trusted owning code supplies the installed engine's
absolute path and SHA-256. It rejects symlinks, unsafe ancestors, wrong ownership,
engine hardlinks, nonexecutable or writable engine files, and non-0700 data roots.
It never repairs permissions or discovers an alternative executable. Ordinary
startup requires existing storage; only explicit first provisioning may create it.
The installer must exclude concurrent changes to the selected engine and paths.
NoSQLDB's fileStorageStartup: 'current-format-only' rejects legacy/mixed storage
and orphan migration staging without modifying that tree. The native file-root
lease is the only cross-process database lock. Model preparation happens after
native ownership and readiness, with no duplicated format detector or lock helper.
Each attempt uses its own mode-0700 /tmp/pallet-identity-* directory for the Unix
socket; it contains no persisted application data. Normal cleanup removes only
that owned socket and empty directory. Parent death can leave an empty disposable
directory for OS temporary-file cleanup; startup never scans or deletes others.
run() admits a guarded identity facade, not the database or store object. Its
methods expire when the callback settles, and admitted calls are tracked even if
the callback does not await them. Stop immediately blocks new callbacks, drains
admitted work, closes SmartData, and then confirms native exit. A failed database
close retains the running engine's lease until cleanup is retried. Start/stop
deadlines bound caller waiting, not resource ownership; pending or failed cleanup
blocks restart. Native exit invalidates readiness and requires stop before restart.
pnpm exec tstest test/test.identityruntime.node.ts --verbose --logfile --timeout 60
The Linux tests qualify competing Node owners, successor receipt recovery, startup/stop races, cleanup failures, exact pending recovery after native SIGKILL, and active receipt recovery after an idle Node parent's SIGKILL. These are not hardware power-loss or ARM64 hardware tests, and do not prove bounded native exit when a management command is stuck. Production process supervision and coherent installer/engine packaging remain integration gates. When packaging that engine, retain NoSQLDB's own license and complete third-party notices alongside the binary.
Private enrollment control
ts/control/classes.enrollmentcontrol.ts owns its identity runtime and a root-only
Unix listener at /run/serve.zone/pallet/control.sock. It acquires the native
database lease before touching that socket. The installer provisions the trusted
platform parent; the control owner creates its mode-0700 IPC leaf when absent,
creates a mode-0600 socket and leaves the directory in place after shutdown.
Existing directories are checked, never repaired. Paths must be canonical,
symlink-free, protected from untrusted writers and fit Linux's socket path limit.
Only the five published requests.pallet enrollment and runtime-binding methods are registered.
Prepare returns a current, independently retained hash-only proposal; bind proves
the full pending enrollment; activate checks the complete authenticated Cloudly
acknowledgement and then proves the exact current local identity. Historical
receipts cannot masquerade as current activation. Spark must authenticate Cloudly
before forwarding its acknowledgement or runtime routing binding over this trusted
local channel. bindPalletNodeRuntimeBinding accepts the first exact routing tuple
or an exact replay only after its origin and node match the current active Pallet
identity. readPalletNodeRuntimeBinding requires that same current identity and
rejects an absent binding. The binding is a separate singleton: absence is the
unbound state, and the released identity document's relayOrigin field remains
readable but inert rather than being expanded into unauthenticated routing data. The IPC
does not expose a bearer, database, generic rotation or workload execution method.
TypedRequest 8.0.3 owns routing and envelope identity. Hooks and incoming response routing are disabled; wire-supplied local authority and unknown methods are denied. Each connection carries one four-byte big-endian length-prefixed JSON frame per direction followed by write-half-close. The server explicitly keeps the response half open until its reply is flushed. Requests are limited to 32768 bytes and responses to 16384, with a default 10-second total connection deadline and 16 admitted operations. Invalid UTF-8, incomplete frames, trailing bytes and missing EOF are rejected. No TCP listener or automatic application retry is introduced.
Disconnects and deadlines close transport but do not abandon an admitted database operation or free its admission slot. Shutdown stops admission, cancels socket I/O, drains handlers, closes the owned listener and then stops the identity runtime. Caller deadlines retain the actual cleanup owner and block restart. A live pre-existing socket is never removed; stale recovery requires the native lease, a completed connection-refused probe, and unchanged owner/inode checks. Because Node/libuv unlinks on listener close, shutdown checks its recorded socket and parent before calling close. A replaced name retains failed-cleanup ownership until the installer/operator restores the original owned path. These checks assume the trusted installer excludes concurrent root-owned path changes; they do not claim protection from hostile root. No arbitrary file or recursive cleanup occurs.
pnpm exec tstest test/test.enrollmentcontrol.node.ts --verbose --logfile --timeout 60
Qualification uses real Unix sockets and disposable native NoSQLDB stores for half-close delivery, independent restart, competing ownership, stale/live/replaced paths, malformed frames, hook isolation, cancellation, admission limits, historical receipt rejection and lost activation results. Spark's concrete client, coordinated daemon integration and production packaging remain separate integration gates. Before distributing a bundled control process, include its JavaScript dependency MIT/Apache-2.0 notices in addition to the existing Rust and NoSQLDB notice material.
Enrollment process lifecycle
The private PalletEnrollmentProcess owns one foreground control lifetime.
Readiness follows the native database lease, model preparation and protected Unix
listener. Cancellation during startup prevents readiness. Parent stdin EOF,
SIGTERM and SIGINT stop admission and drain the existing control owner. A 250 ms
watch of the public readiness getter terminates a failed lifetime; it never
restarts the engine or retries a database operation. Each control request already
checks native readiness independently of that watch.
The enrollment CLI accepts exactly enrollment-serve or enrollment-provision.
Normal startup requires existing state; only explicit provisioning may create it.
Trusted build code supplies the exact engine, digest and protected paths. No
runtime path, UID, digest, bearer or configuration arrives through argv, stdin or
environment. Stdin carries only parent lifetime: content is rejected without an
echo. PALLET_CONTROL lines use pallet.enrollment.process and contain only
ready, stopped, or a static failure reason. stopped follows confirmed cleanup;
a failure does not authorize a successor until the supervisor confirms process exit.
pnpm exec tstest test/test.enrollmentprocess.node.ts --verbose --logfile --timeout 60
The lifecycle and source-process tests cover cancellation, EOF/signals, missing state, invalid input, native failure and retained cleanup ownership. A standalone Deno (the pinned 2.9.7) Linux amd64 fixture also qualifies real Unix half-close responses, durable replay, engine and parent SIGKILL, and stale-socket recovery from an unrelated working directory with an empty environment. That fixture supplies trusted test paths and UID. The production artifact, complete notice bundle, installer/supervisor integration and target-platform qualification remain gates; this source adapter is not a production node installation.
Foreground node process
The same control executable accepts runtime-serve for one private
PalletNodeProcess lifetime. It reopens existing enrolled state and verifies the
protected sibling pallet-runtime executor, pallet-guard, pallet-dns and the
managed VPN pallet-vpn against their compiled-in SHA-256 identities before
starting local owners. The released build activates the node network: it passes
activation with pallet-vpn and its digest, so a node bound to its cluster
relay raises its ACTIVE generation and managed VPN tunnel (see "Private ACTIVE
tunnel transitions"), and a node bound to a local controller forwards as a single
host. It never initializes missing storage or accepts paths,
credentials or settings from argv, stdin or environment. The containerd CRI
socket is /run/pallet/containerd/containerd.sock: Pallet's own containerd
instance, never the host's shared /run/containerd/containerd.sock, which on a
Docker CE host is Docker's containerd.io daemon with CRI disabled. The native
runtime refuses that shared socket and Docker's sockets by path and by file
identity: a socket with the device and inode of the shared containerd socket,
Docker's API socket or Docker's embedded containerd socket is refused, so a
symlink, hard link or bind mount of one of them cannot pass under another name.
Pallet and Docker can therefore run side by side on one host without sharing
images, sandboxes or CRI configuration.
The same lifetime runs in two modes that differ only in what ends it, and both
take one argument, required: --forwarding-mode <docker-shared|exclusive>, who
owns the host's IPv4 forwarding (see Host forwarding modes); anything else is
refused as invalid_arguments. runtime-serve is a supervised child: its parent holds its standard input, and
EOF there stops it (Spark's runnode). runtime-unit is the main process of the
node service unit (see "Service unit contract"): its standard input is null, EOF
is not a stop request, and SIGTERM or SIGINT stops it. Both refuse data on
standard input as invalid_stdin.
A start fences what the previous lifetime retained before it reports ready, so
a start after a crash does real work first. A supervisor must allow a start at
least 600 s before it gives up on ready, in either mode
(palletNodeStartTimeoutSeconds in @serve.zone/pallet-bundle). Pallet bounds
each native request of a start, not the start as a whole: a request that misses
its deadline ends the start with failed. The longest path whose length does not
grow with the node's workloads is a start after a crash, on a host whose guard
unit already ran in this boot, with an ACTIVE generation retained and a guard
expansion pending. Its request deadlines add up to 478 s:
- identity database 10 s (its startup deadline, migrations included), router namespace 18 s (6 s spawn including the 5 s descriptor-store barrier, two 6 s requests), independent clock 11 s (3 s spawn, 8 s start);
- guard pass 274 s at smartnftables' 30 s per request: owner open 33 s (with its 3 s spawn), replay of the committed transition 90 s (prepare, reconcile, inspect), the pending expansion 120 s (prepare, then its replay), detach 31 s;
- handoff set open 11 s (namespace read 3 s, spawn 3 s, open 5 s);
- fence of the ended ACTIVE generation 131 s: the native's epoch proof 5 s, the retained host packet table's owner 36 s (namespace read, spawn, open), its recompilation 30 s and its release 60 s (release, inspect);
- application recovery 15 s (epoch proof, prepare and reconcile at 5 s each);
- execution host identity 8 s (3 s spawn, 4 s request, 1 s termination grace of a failed child).
Each retained workload attachment of the ended epoch adds its 5 s epoch proof (at
most 512, twice the projection's 256 endpoints). The DNS pass, the database
transactions, the executable digests and the storage recovery have no deadline of
their own. 600 s covers the fixed path and leaves 131 s for those; a supervisor of
a node that retains many workloads allows more, and PalletControlLifetime
admits up to an hour (469 s + 512 × 5 s = 3029 s). Under the node unit
(Type=exec) the service manager does not wait for ready, so TimeoutStartSec=
does not bound a start; an installer that waits for the unit's ready line
allows the same limit.
Runtime lifecycle lines keep the PALLET_CONTROL prefix with the distinct
pallet.node.process protocol. ready means local identity, namespace and execution owners
are ready; it does not assert controller connectivity, workload readiness or DNS
availability. A start the packet engine refuses because of the running kernel ends
with failed reason unsupported_kernel instead of owner_failed (see "Host
kernel requirements"). A node that fails because its coordinator refused the stated
forwarding mode for a fact of the host that persists until an operator changes it
ends with failed reason forwarding_mode_refused, before or after ready, its
refusal named on stderr: under exclusive persisted forwarding
(forwarding.exclusive.persisted) or a running Docker
(forwarding.exclusive.docker_present); under docker-shared no iptables-nft
frontend (forwarding.docker_shared.iptables_nft) or a host that does not forward
when the uplink opens (uplink.forwarding.open). Spark ends the node on it rather
than restart it into the same refusal; everything transient stays owner_failed, and
the guard modes never send it. Initial Cloudly connectivity runs independently. The runtime
session requires the immutable node runtime routing binding and registers at its
cluster relay. An absent binding fails startup; there is no direct-to-Cloudly
fallback or second configured relay authority. Every registration response must
match the persisted node, cluster, Cloudly controller and runtime namespace before
it becomes execution authority, and reconnect rechecks the current identity and
the exact persisted binding. The persisted tuple grants routing only; phase and
workload authority always come from the fresh authenticated runtime session.
Enrollment, the node credential and the registry host stay with Cloudly
whichever transport carries the session: workload.registryHost is the
publication identity the registry credential is fenced to, and it is never
rewritten. Where the image bytes are fetched is a separate statement,
workload.pullEndpoint — the relay's own origin for a cluster whose relay
forwards the registry, so the node pulls inside its cluster and nothing
cluster-side dials the control plane for bytes. The reference stays
digest-pinned either way, so the endpoint decides reachability and never
content, and the credential this node asks Cloudly for is addressed to whoever
serves them. The contract this node relies on from
the relay: it is expected to forward the node bearer verbatim and to serve the
unchanged typed request contracts, so nothing in the session binding changes
for Pallet. New workloads
still require the exact authenticated session and registry grant, while persisted
node-bound effects retain their existing offline rules. Network and secret
references resolve through their owning resolvers on this same session before any
native effect; storage references resolve to the local storage claims this node
admitted on the same session ("Local storage claims").
Every inbound request a node refuses answers its sender the one refusal message
of its kind, so a sender learns that it was refused and never why. The node
writes why to stderr, which Spark forwards to the node journal:
Pallet controller request refused: <method>: <chain>., the owner_failed
labelling of class names, codes and sites: a new line at once, and a repeated one
again with its count (<line> (repeated N times since <ISO time>)) one minute
after it was last written, then at doubling intervals up to ten minutes. A
network pass that fails is written the same way, as
Pallet network pass failed: <chain>., until a pass completes: a failing pass
renews nothing, so the node's tunnel lease runs out behind it. An assignment scan
whose page the store refuses to read is written the same way, as Pallet assignment scan failed: <chain>., until a page is read: the scan keeps its cursor and reads the
page again on its next interval, while the workloads, the network passes and the
controller session keep running. The node fails only when an owner is no longer
live, a database server that stopped included; Pallet 36.3.6 and earlier failed it
on any such refusal, and the restart that followed took every site down (Grasberg
2026-10-10, one MongoServerError on this read). A database server's error is named
by the server's name for its code and the number (MongoServerError:IllegalOperation(20)),
never by its message, which can carry a document's key value or a path. The owner_failed
labelling is bounded to 448 characters; a longer chain is cut in the middle
(...), keeping its outermost owners and its last 192 characters, which name
the innermost cause — what actually refused.
A controller may push on a session as soon as it answered its registration, but the
node takes the session up only after it read that answer and its active credential
again (on the controller socket, once the socket reports itself connected). From
the registration on, until then, an inbound request on that connection waits for
the session, at most palletSessionAdoptionWaitMs (5 s), and then acts under it or
is refused as any request without one; one that outlasts the wait is refused with
the site controller.session.adoptionPending.
Protocol handshake
The contract this node speaks is the installed @serve.zone/interfaces release,
and it is stated once: ts/controller/protocol.ts builds one offer for the one
session kind Pallet is (palletRuntime) from protocol.createProtocolOffer,
with this build's own minimum (32.0.0, raised only by the commit that starts
depending on a later minor). Every registration carries that offer, and the
answer is the wrapped { session, protocol } the controller returns, negotiated
against the offer that was sent before the binding is bound.
Two peers that cannot serve one session are named rather than collapsed into a
transport failure. A controller that refuses this build answers an
IProtocolRefusal; an accepting controller whose own offer this node cannot
serve is refused by this node. Either way the node reports the status
protocol-incompatible and retains the refusal in getStatus(). It is not an
owned failure: the workloads keep running, the network stays up, the packet
policies stay applied and the node's owners are not torn down — and the node
process is not ended either, because protocol-incompatible is a live node with
no controller it can speak to.
Recovery needs no operator action on the node. The refused offer is made again
every protocol.refusedOfferRetryIntervalMs (five minutes, the one interval
@serve.zone/interfaces states so that no client invents a second one) until a
controller accepts it, through @api.global/typedsocket's
TypedSocketRestoreDeferral: the refused attempt is closed, the transport waits,
and it offers again without consuming one of its reconnect retries, so a node
refused for days keeps them for real connection failures. An accepted offer drops
the retained refusal, returns the client and the node to ready and resumes the
pass, so upgrading the controller under a running fleet brings the fleet back by
itself. The node's own poll loop keeps its interval while refused — the pass stays
empty and admits nothing — because that loop is what carries the accepted offer
back into the node's state.
A restoration the transport refuses outright is the one terminal outcome.
TypedSocket releases itself, because a retry could only repeat the peer's
verdict, and that leaves the node with no controller at all — the condition an
exhausted transport leaves it in — so the node fails owned: the process ends and
the supervisor's restart is the next offer, which lands on the same five-minute
cadence if the controller it meets has still not moved. The denial is named on
the controller status as restoreDenial — the transport's own reason and, when
the refusal was this client's, the cause behind it — beside any refusal the node
was already standing on, so the status the process ends on says which verdict
ended it.
The registration also states resolvedAuthorities, the sorted, duplicate-free
list of workload authorities this build can resolve. It is the same constant the
execution owner enforces (resolvableWorkloadAuthorities), so what the node
promises and what it refuses can never drift: an assignment whose
requiredWorkloadAuthorities the list does not cover is refused as
unresolved-authority. This build resolves network, secrets and storage.
Local controller
A node can be driven by a controller on its own host — Onebox — instead of
Cloudly (nodeRuntimeLocalControllerContract in @serve.zone/interfaces). A
node carries a Cloudly identity or a local controller, never both: a first local
bind is refused while any Cloudly enrollment state exists, and a Cloudly
enrollment is refused once a local controller is bound. Both claim one shared
pallet_node_authority singleton in the transaction that writes their identity,
and the first claim inserts it, so of a concurrent Cloudly preparation and first
local bind exactly one commits. The other's commit conflicts with it, the
transaction is retried, and the retry finds the committed authority and refuses
as conflict. An offline read that finds
both kinds anyway refuses instead of choosing one. The persisted identity
selects the transport when the controller client starts; there is no configured
choice between them.
The local controller is authenticated by the socket's permissions and nothing
else: the socket is owned by root with mode 0600 in a root-only directory, so
any process running as root on the host is trusted as the controller. The
bearer Pallet generates and the controller pins at its first bind detect a
Pallet that lost its state (a later bind answering another credential is
refused by the controller); they do not keep out an intruder with root.
Pallet serves the local controller on the root-only Unix socket
/run/serve.zone/pallet/controller.sock through SmartServe's unixSocket
listener, which applies mode 0600 before any peer can connect and refuses a
path another process serves. The controller is the TypedSocket client
(http://localhost over the socket), but the session keeps its directions.
Its first request on a connection is bindPalletLocalRuntimeController: the
first bind persists the binding and a bearer this node generates, and every
later bind must be the identical binding and answers the same credential
generation and SHA-256 hash. The bearer itself never leaves this node except in
its own registration. After answering, the node fires the unchanged
registerPalletRuntimeSession at that exact connection with that bearer, and it
binds the answered session to its own persisted binding
(bindRegisterPalletLocalRuntimeSessionResponse) before it becomes execution
authority. Admission, reports, registry and secret requests then run as on the
outbound transport, but only on the connection that holds the session: another
connection on the socket, bound or not, never acts under it. A connection whose
registration is accepted supersedes the previous session and closes its
connection; a failed registration closes only its own. A controller that
refuses this build's offer leaves the node protocol-incompatible until a later
connection's registration is accepted.
A local controller has no Cloudly origin to stand for its registry. It issues a
pull credential only for an image whose registryHost, and pullEndpoint when
one is stated, are both listed in the binding's registryHosts
(isNodeRuntimeLocalRegistryWorkload); any other image is refused before a
credential is requested.
A node bound to a local controller always serves it from its node lifetime
(runtime-unit under its service unit, or runtime-serve). Only
network-acquire-local admits the first bind of an unbound node (see
Initial projection acquisition). Networked
workloads of a local controller need the attach barrier like every other
workload; on a node without a managed VPN it holds as a single host (see
Private ACTIVE tunnel transitions).
Stdin EOF, SIGTERM and SIGINT close admission and join startup, controller and
native operations. The CNI broker remains available through CRI rollback, then
joins its admitted handlers before the network and database owners close.
The private clock stops and joins its Chrony/query children before the retained
router namespace closes and the database owner is released. Terminal node failure ends
the lifetime with a static failure reason. The 250 ms liveness watch never restarts
the node or retries an operation. A stopped event follows successful cleanup;
Spark must also confirm child exit before admitting a successor.
This process composition does not start containerd, provide private DNS,
activate Spark's daemon or establish production readiness. Spark must consume the
released Pallet bundle and commission its guard before starting this runtime, and
Pallet's own containerd (containerd-serve, below) serves the CRI socket
/run/pallet/containerd/containerd.sock this runtime uses.
Pallet's own containerd
The sealed control bundle carries Pallet's container runtime under
pallet-containerd/: containerd 2.3.6 and its runc v2 shim (the unmodified
upstream release executables), runc 1.5.1+servezone1 (built from the signed
upstream sources, statically against musl), the pause 3.10.2-servezone1 sandbox
image as an OCI archive (pause.tar) and manifest.json, which records every
file's size, mode and SHA-256. The control executable compiles in that manifest's
digest. No upstream CNI plugin ships: the only network plugin is Pallet's own
pallet-cni, which is the verified pallet-runtime executable, and containerd's
internal loopback serves lo.
pallet-control containerd-serve is the main process of the containerd service
unit. It verifies its sibling pallet-runtime against its compiled-in digest and
starts it in its --containerd-management mode, which:
- verifies
pallet-containerd/against the manifest: exact names, modes, sizes and digests, root-owned and not group- or world-writable up to/; - refuses to start a second containerd while one answers on
/run/pallet/containerd/containerd.sock; - writes the configuration below to
/run/pallet/config/config.toml(mode0600), the CNI networkpallet-workloadsto/run/pallet/config/cni/conf/10-pallet.confand a link/run/pallet/config/cni/bin/pallet-cnito the verifiedpallet-runtime, all recreated on every start; - starts containerd with an empty environment except a system
PATH(for host helpers containerd resolves by name, such asapparmor_parser), bound to its own lifetime (PR_SET_PDEATHSIG), and waits up to 60 s until CRI reports versionv2.3.6andRuntimeReady; - imports the bundled pause archive through containerd's own Transfer service over
a Streaming session, as
ctr image importdoes, when containerd does not already hold it with the pinned config digest, unpacks it for theoverlayfssnapshotter underpallet.local/pause:3.10.2-servezone1, and requires CRI to report that image id. Nothing is pulled from a registry, and noctrexecutable ships.
Then the process reports PALLET_CONTROL {"protocol":"pallet.containerd.process","event":"ready"}
on stdout and, when started by a Type=notify unit, READY=1 over systemd's
notification socket. containerd's own log lines pass through to the process's
standard error. containerd's exit fails the process (failed, owner_failed,
exit status 1); it is never restarted inside the process — the service manager
restarts the unit. Once containerd is ready, SIGTERM or SIGINT stops it with SIGTERM,
waits up to 30 s, removes /run/pallet/config and reports stopped. A SIGTERM during
startup kills the starting containerd at once and fails the process; the left-over
/run/pallet/config is recreated by the next start. Stopping containerd
never stops a workload: every shim, and every container beneath it, keeps running
and is reattached by the next containerd. Standard input is not read (a unit's
null stdin is expected; data on it is refused).
The generated configuration is complete and is Pallet's own:
| Setting | Value |
|---|---|
version |
4 |
root, state |
/var/lib/pallet/containerd, /run/pallet/containerd |
imports |
[] — the host's /etc/containerd/conf.d never applies |
required_plugins |
io.containerd.grpc.v1.cri |
disabled_plugins |
io.containerd.internal.v1.opt, io.containerd.image-verifier.v1.bindir, io.containerd.nri.v1.nri |
| gRPC / TTRPC | /run/pallet/containerd/containerd.sock / .sock.ttrpc, uid 0, gid 0 |
| CRI stream server | 127.0.0.1:10010, stream_idle_timeout = "15m", no TLS streaming, no CRI TCP service |
| images | overlayfs snapshotter, sandbox image pallet.local/pause:3.10.2-servezone1, no registry host directory |
| runtime | runc via pallet-containerd/containerd-shim-runc-v2 and pallet-containerd/runc, runc state /run/pallet/runc, SystemdCgroup = false; containerd's PATH is pallet-containerd first, then /usr/sbin:/usr/bin:/sbin:/bin, so the runtime info containerd asks of a shim configured by path (-info, without the runtime's options, which looks up runc by name) finds Pallet's own runc, never a host's |
| CDI | enable_cdi = false, no specification directories |
| CNI | /run/pallet/config/cni/{bin,conf}, one configuration, internal loopback |
| other | image-defined volumes ignored, network namespaces under the state directory |
Workloads get cgroups under the absolute parent /pallet (the executor names it
on every sandbox), never under the unit that runs containerd: runc can enable the
CPU, memory and PID controllers there because the cgroup holds no processes, and
a stop of the containerd unit reaches no workload.
The CRI stream server (exec, attach and port forwarding) listens on loopback only,
and only root may dial it: every allocation-pool guard policy carries the
Smartnftables loopback port owner { address: '127.0.0.1', port: 10010, uid: 0 }
(see Offline allocation-pool guard journal).
Another user's connection is reset, and the port is dropped on every interface but
loopback. The containerd unit requires the guard unit, and a guard pass that did not
apply the owner fails, so containerd never serves without the rule.
Service unit contract
Spark on fleet nodes and Onebox on its own host install the unit; Pallet never writes a unit file. The unit must:
- run
<bundle>/pallet-control containerd-servewith no further arguments as its main process, as root,StandardInput=null; - use
Type=notifywithNotifyAccess=all(the ready notification comes from the native owner, a child of the main process); - use
KillMode=process, so stopping the unit signals onlycontainerd-serve, which stops containerd itself, and every shim and workload survives; - use
Delegate=yes, as containerd's own unit does, so systemd leaves the cgroups of the processes it starts to them; - set
TimeoutStopSec=of at least 45 s (30 s containerd grace plus the owner's joins),Restart=alwayswith a shortRestartSec=,LimitNOFILE=infinity,TasksMax=infinityandOOMScoreAdjust=-999; - order
After=andRequires=the retained guard unit, and be orderedBefore=the node unit, whichWants=it: the node runtime is the only CRI client.
A node unit of Pallet's own (Onebox) runs the node lifetime. Spark does not use it,
nor createPalletNodeServiceDefinition or palletNodeServiceSettingsMatch: its node
unit runs spark runnode, which supervises runtime-serve (standard input held,
--router-namespace-descriptor 3 with the namespace its own unit's store kept;
see "Router namespace lifetime"). Pallet's node unit must:
- run
<bundle>/pallet-control runtime-unit --forwarding-mode <mode>as its main process, as root,StandardInput=null,Type=exec, with the host's forwarding mode (see Host forwarding modes) and no other argument; - use
KillMode=mixed: SIGTERM reaches only the main process, which joins every owner and child it started, and whatever remains of the unit's cgroup when the stop timeout ends is killed before the service manager starts a successor; - set
TimeoutStopSec=of at least 360 s (a five-minute native operation, then the database joins) andRestart=on-failurewith a shortRestartSec=: a failed lifetime is restarted by the service manager, never inside the process; - order
After=both the guard and the containerd unit,Requires=the guard unit andWants=the containerd unit; - keep the router namespace:
NotifyAccess=all(the namespace keeper, a child of the main process, stores it),FileDescriptorStoreMax=1andFileDescriptorStorePreserve=yes(systemd 254 or later), so the manager hands it back throughLISTEN_FDSto the next main process; an installer releases the store (releasePalletNodeDescriptorStore) when it changes the forwarding mode and on uninstall.
A start of the node may take up to 600 s before its ready line (a start fences
what the previous lifetime retained first; the derivation is in "Foreground node
process"). Type=exec does not wait for it; an installer that does allows at
least that long.
The node unit is started only after the node's first projection is acquired
(network-acquire-local or network-acquire); an unbound node fails its start.
The guard unit is a oneshot (Type=oneshot, RemainAfterExit=yes) that runs
guard-recover with null stdin, without default dependencies, ordered Before=
network-pre.target, systemd-networkd.service and shutdown.target, which it
conflicts with, and required by network-pre.target and systemd-networkd.service.
Its executable is pallet-control or an installer's own executable that verifies
the installed bundle before it starts that mode. The installer writes it only after
the first guard-commission.
@serve.zone/pallet-bundle (from this repository's ts_bundle/) states these three
definitions for an installer and reads the loaded containerd and node units back
against this contract.
The QEMU CNI guest's containerd-serve scenario runs this process from a
build:control directory without systemd (see
Isolation and qualification); the unit's own
KillMode, Delegate and notification behaviour is qualified with the installer.
Router namespace lifetime
The verified pallet-runtime executable also owns one private router namespace
for each node process lifetime. Its exact --network-namespace-management mode
creates an unnamed Linux network namespace on the initial native thread, before
Tokio, sockets or readiness. Before it exposes the namespace it switches IPv4
forwarding off in all and default and disables IPv6 in both, and verifies
each (with lo's forwarding): a new namespace copies the host's IPv4
forwarding, and the router must forward nothing unless an UP generation's packet
policies filter it. This requires permission to create network namespaces;
failure prevents node readiness and controller admission.
The node opens the keeper's live namespace descriptor, compares its device and inode, rejects its own namespace, and confirms the native nonce and identity over the original stdio connection. A reused PID or saved metadata cannot establish authority. Nothing mounts or names the namespace.
Kept across restarts. Under a service manager that passed NOTIFY_SOCKET, the
keeper removes any previous FDNAME=pallet-router from the manager's descriptor
store with FDSTOREREMOVE=1 and waits for BARRIER=1 before reporting ready,
whether it creates or adopts a namespace. The store stays empty while the lifetime
can mutate the network. Only a clean stop, after every namespace consumer has joined
and the handoff set was confirmed, sends keepNetworkNamespace: remove any kept
descriptor, then FDSTORE=1 with FDPOLL=0 and SCM_RIGHTS, then a barrier.
A lost coordinator, lapsed lease, crash or failed cleanup cannot hand an unsafe
namespace to every later restart. Only that keeper is given the
socket: the node removes NOTIFY_SOCKET from its own environment before any child
starts, and nothing ever sends READY=1, MAINPID= or STOPPING= — the main
process is the supervisor's. An unconfirmed removal (namespace.release) refuses
startup; an unconfirmed store (namespace.store) refuses the clean handover.
The next node process receives the kept namespace as descriptor 3 —
runtime-serve --forwarding-mode <mode> --router-namespace-descriptor 3 from a
supervisor such as Spark's runnode, or through LISTEN_FDS (exactly one
descriptor named pallet-router, for this process) as runtime-unit. It reopens it
as a close-on-exec handle and closes the inherited number before any child starts
(compiled Deno marks the number close-on-exec first, then closes its own duplicate
of it through fs and the number itself through libc, in that order; a number that
is inherited but cannot be closed fails the node before any child starts), so the
handle is the process's only reference and no child holds the namespace. It adopts it only as the epoch the
application lane last journaled, with no unfinished ACTIVE generation. A keeper started with --adopt proves the
descriptor a network namespace (NS_GET_NSTYPE) of the journaled device and inode,
in the journaled boot (a soft reboot keeps the store, a reboot empties it), enters
it and reads its own namespace back. The adopted namespace keeps the journal's
identity — its nonce and creator PID are the epoch's tokens — so the attachment,
handoff-set and ACTIVE journals see the same epoch: attached pairs are recovered
in place and the handoff set is completed again. A handover from an earlier version
that still has an unfinished ACTIVE generation is rejected even when its descriptor
identity matches: its lease may have lapsed or its coordinator failed. Its epoch
ends rather than replaying the same conflict at every restart. A handover that cannot be proven — no journal,
another boot, not a network namespace, another device or inode — is closed, named
on stderr (Pallet router namespace not adopted: namespace.adopt.<check>.), and a
fresh namespace is created; the runs of the ended epoch
then fail as network-epoch-ended. A reboot empties the store, and so does a stop
of a unit systemd then unloads: FileDescriptorStorePreserve=yes keeps the store
only while the unit stays loaded, which an enabled unit does, and systemd unloads
an unreferenced inactive unit and closes its store. An installer releases the
store when it changes the forwarding mode and on uninstall
(releasePalletNodeDescriptorStore).
A stop hands the set over. The next process of the epoch completes the handoff set
it adopts against the receipt the application lane journaled, and a member that
receipt holds is never created again: a pair that vanished under its receipt is a
conflict (handoffset.reconcile.member_vanished), not a creation. So a stop of a node
whose router namespace remains eligible for a clean handover (PalletNamespaceOwner.kept:
the manager's socket is available and no owning failure invalidated it) releases
nothing in it: the handoff coordinator's handOverHandoffSet withdraws and discards a
generation it still holds and closes the uplink observation, exactly as the release
does, then proves the set it leaves — no transition pending, and the complete inventory,
the dormant pairs and every workload pair attached through them, equal to its receipt —
and ends its protocol with every link in place; the owner's closure reads released
and handedOver. A set that changed under its receipt is refused by name
(handoffset.hand_over.<check>) and the stop is unconfirmed, as a refused release is.
Pallets 36.0.0 to 36.1.1 released the set at every stop without an attached workload,
and every successor in the same boot failed handoff_conflict:handoffset.command.reconcile
until a reboot emptied the store (lab 2026-10-04); a stop with an attached workload was
refused instead, since a release would remove a pair a running sandbox still uses. A namespace nobody keeps ends
with the process, so its set is released as before. A failed lifetime or crash leaves
the descriptor store empty and the namespace ends with its last live descriptor,
every link in it with it. Once the store is released (releasePalletNodeDescriptorStore) or
emptied by a reboot, the next process creates a fresh namespace and proves the previous
epoch's host attachments absent ("Previous-epoch host attachment audit"); kernel
namespace teardown is asynchronous, so that audit may refuse the first start right after
a release while the kernel is still deleting the pairs.
Attachments outlive their generation. An attachment created under an UP
ACTIVE generation is journaled bound to it (activePreparedDigest), and a DEL
names the generation held at that moment, so a DEL of a pair bound to a
generation that has since ended would never be admitted. Whenever a generation of
the node's own epoch is DOWN — withdrawn in process, or recovered DOWN by a
process that adopted the kept namespace — when a node process of that epoch
starts, and again before any successor is prepared, the attachment owner carries the epoch's live attachments past it: the
handoff coordinator re-reads each pair in the kernel (inspectWorkload, the
links by name, address, marker and index in their own namespaces, as CHECK does),
and a pair that is complete or raised with its journaled receipt is released from
the ended generation in one journaled, idempotent step. It then reads as a pair
prepared before ACTIVE: the successor is prepared over it, its own recovery
accounts for it, and its DEL names whichever generation holds it — the successor,
or none in between. A pair that cannot be proven is fenced instead: its run reads
as of an ended epoch, so the execution owner stops its sandbox and fails it as
network-epoch-ended; its DEL releases and removes it; and no successor is
prepared while one is fenced (active.activate.unprovenAttachment). Each carry
proves every pair again, so a restarted process fences the same pairs. A fenced
run is no DNS source: no name is served for a pair that is not proven.
Carried pairs are raised again. A recovery fences every pair of the generation it
takes DOWN — a crash, or a stop that left the generation UP in the journal — and the
successor is prepared over those pairs, but only an ADD ever raised a pair, and no
ADD comes for a running sandbox: Pallet 36.0.2 left such a workload running with its
link DOWN, and reported it ready. Once the barrier of a newly raised generation holds,
the attachment owner re-reads every live, carried pair of the epoch in the kernel and
raises each one that is not raised (PalletWorkloadAttachmentOwner.raiseCarried),
writing Pallet workload attachment raised again: run <digest>.; one it cannot prove
or raise is fenced — its run reads as of an ended epoch and is stopped and failed by
name — with Pallet workload attachment down: run <digest> CARRY: <label>.. A pair a
clean withdrawal left raised is only read.
A sandbox that vanished while the node was down. A process that adopted the
kept namespace restores each attachment of its epoch by capturing its sandbox
namespace at the journaled path. When that path is gone and nothing of the pair
is left in the router namespace — no link by its router name, alias or either
MAC address, no address or route in its subnet, read twice — the pair went with
its sandbox (a pair's two ends live and die together), and the coordinator says
so (recoverWorkloadAttachment answers restored: false) instead of ending.
The pair is fenced from the start: its run fails as network-epoch-ended, no
ADD or CHECK is answered for it, an ACTIVE recovery takes it as retired, it
reports absent, and its DEL proves its absence again and completes it. A path
that still exists — the sandbox, or something else at its name — or a pair whose
router side is still there still refuses, and the node fails as before.
A DOWN the coordinator does not hold. Every packet-cleanup step reads the
generation's DOWN from the coordinator: the generation restored by its own
withdrawal, or the recovery it accepted. The coordinator of a later process of
the same epoch holds neither, so a generation an earlier process took DOWN
without finishing its packet cleanup is taken DOWN again as a recovery
(PalletHandoffSetOwner.holdsDown), which a fresh coordinator admits before it
adopts its set and which changes nothing over a generation already DOWN, and the
cleanup is finished under that DOWN.
A lost coordinator costs the generation, not the node. When the handoff
coordinator ends — its process exits, for instance after it fenced a kernel
event the UP generation cannot tolerate, or a command finds it gone — while this
keeper still holds the router namespace, the node joins the application owner
composed over that coordinator and starts a new one over the same namespace. Its
start recovers the attachments and fences the retained generation DOWN exactly as
a restarted process of the same epoch does, and the next admitted pass raises a
new one; passes are refused as unavailable meanwhile. A replacement must be
followed by a pass that raises a generation (barrier up or forwarding) before
another loss is replaced: a second loss with no raise between them fails the node,
so a conflict that recurs on every raise never becomes a rebuild loop. The bound
is kept by the node process; a restarted process admits one replacement again. A
lost keeper, a native refusal, any other failure, or a replacement that does not
start still fails the node.
The private PalletNamespaceOwner.withDescriptor() API retains the source
descriptor through an admitted consumer's asynchronous spawn. Callbacks must not
close it or keep its number after returning. Shutdown fences new borrowers and
joins admitted callbacks. The enclosing node must also join all DNS, VPN and
packet-policy processes before closing this owner. Unexpected keeper exit fails
the node, stops admission and retains the parent descriptor until explicit joined
cleanup; it never silently recreates the old namespace.
The isolated Linux 6.18.35 amd64 fixture runs both Node and compiled Deno (the pinned 2.9.7)
against the real namespace keeper and published Smartnftables 1.4.0 binary. It
verifies descriptor inheritance, actual namespace identity, keeper-loss fencing,
joined borrowing and unchanged parent links and routes, and — on a host that
forwards IPv4 — that the namespace is exposed with forwarding off in all,
default and lo and IPv6 disabled in all and default. The full node composition
is separately checked through its compiled control bundle. Privileged ARM,
containerd/CNI attachment, DNS and complete network policy remain separate gates.
Staged secret material mounts
The same verified executable owns one staged mount per admitted assignment that
carries a secret reference. Its exact --secret-mount-management mode stages that
attempt's material on its initial native thread, before Tokio, sockets or any
other work: it unshares a private mount namespace, mounts a
noswap,nodev,nosuid,noexec tmpfs there, writes one root-owned file per entry
with the delivered mode, places the ownership marker inside the mount and
detaches it with open_tree.
This requires permission to create mount namespaces and mounts, and it must run in
the host mount namespace — the helper refuses its own work when it does not.
The material never crosses the JSON IPC. PalletSecretMountOwner spawns one child
per operation and writes the opened values as one published Smartrust sensitive
frame on descriptor 3, bounded by this attempt's own declared total rather than by
a shared ceiling. Prepare and attach share that one child, because the detached
mount dies with the process that holds it; inspect and release run in fresh
children, because recovery may never depend on a process that is already gone. No
value is ever a return value, a persisted field, an environment value, a command
line or an IPC member, and the caller's buffer is cleared after the write.
Publication follows durable state, never the other way round: the mount plan and
the pending run commit in one transaction before the first native step, the
prepared receipt — secured parent mount, base and staged device/inode, detached
mount id, ownership nonce — is persisted before move_mount publishes the mount,
and the published mount id is verified equal to the id the detached clone already
allocated. The executor then demands exactly that plan's read-only private
per-file binds from the container, never a bind of the mount directory that holds
the marker, and refuses a live container whose mount list is not that plan.
Recovery reads the persisted receipt alone. A fresh helper verifies the recorded
mount by descriptor, unmounts strictly, proves that exact mount id gone and the
base directory empty before removing it, and answers uncertain with an exact
reason on any mismatch — which retains the mount, the directory and the record and
deletes nothing. Stop retains the mount; release is an explicit owner step after
proven CRI absence, never a side effect of removal. That order is the only thing
that protects a running workload: a container's own bind of a file in the mount
lives in its own mount namespace and does not make the host mount busy, so the
kernel would not refuse a release taken out of order. A record from an earlier
boot names nothing this kernel can find, because /run is a fresh tmpfs after
every boot, so it is retired without any native step.
The chain /run/serve.zone/pallet/secrets is walked through no-follow
descriptors. /run/serve.zone is the parent the installer supplies for every
serve.zone product on the node, so it is required to be exactly what the CNI
plugin requires of it — a root-owned directory no one else can write — while the
two components below it belong to Pallet alone and must be root-private.
The offline QEMU qualification drives the shipped binary itself, over the same
protocol, through every stage boundary on 6.18.35-0-virt, 6.8.0-124-generic
and 7.0.0-22-generic, and the containerd guest runs one admitted assignment
end to end: sealed material opened, mount published, the container created with
exactly the plan's bind, the marker invisible inside it, stop retaining the
mount and removal releasing it after proven absence.
Local storage claims
A controller — Cloudly or Onebox — gives a service a retained, node-local volume
with applyRuntimeLocalStorageClaim (IRuntimeLocalStorageClaim,
@serve.zone/interfaces 32.24.0). A claim names the volume by identity only:
controller, organization, cluster, node, runtime namespace, service and the
service's volumeId. It carries no host path, no device, no mount option and no
credential; this node resolves the identity to its own storage.
Layout. Every volume lives under /var/lib/serve.zone/pallet-storage:
volumes/<volumeKey> is a volume, staging/ holds one being created, trash/
one being purged and import/ the bytes an operator places for an import
claim. The volume key is the SHA-256 of the claim's identity and id, so it is
stable across every generation and a re-epoched controller keeps its volumes.
The root and its four directories are root-owned 0700 below trusted ancestors;
a volume is reachable only as the bind a granted container receives. The root is
created only while no volume was ever materialised: a root that later goes
missing — an unmounted disk — is refused by name (storage.root.missing) and
never replaced by an empty one.
Ledger. The claim ledger (pallet_storage_claims, pallet_storage_names,
pallet_storage_grants) lives in the node's own NoSQLDB through
@lossless.org/client/nosqldb. Admission is one transaction fenced to the
session the claim arrived on: a new claim, the next generation (whose previous
must name the one this node holds), a replay of the same generation or a
historical older one. A same-generation claim with another digest, a next
generation that changes the volume's identity, source, initialization or initial
ownership, a generation after a purge and a second live claim of the same service
volume are refused by name and write nothing. The answer is the volume's state:
ready, import-required, purged or uncertain.
Crash safety. Every filesystem effect runs between a ledger intent and the
completion that records its evidence, one at a time. An empty claim commits
provisioning, then creates the directory in staging/ with the claim's
initialOwnership and mode and publishes it by one atomic rename; the ledger
records the directory's device, inode and birth time. A purge commits purging,
renames the recorded directory into trash/ and removes it there, never
following a link the workload left inside. Before the ledger records a step
done, every directory it created, renamed or removed is fsynced with its parents
(a new layout directory with its parent; the volume with staging/ and
volumes/ before ready; volumes/ and trash/ before purged), and a failed
sync keeps the intent pending, so a power loss can only undo a step whose intent
the next start still finishes. On start, and on every replay, an
interrupted intent finishes from what is on disk: a leftover staging tree is this
intent's own and is removed, a published directory is adopted only while it is
still exactly the fresh, empty directory the intent creates, and a purge whose
volume already sits in the trash completes. Evidence that contradicts the ledger
— a foreign directory at the volume's name, a volume whose identity changed —
leaves the claim uncertain, which this node neither mounts nor deletes until an
operator decides. A pending intent that cannot finish at start is listed in the
execution owner's storageRecoveryRefusals and stays durable.
ReadWriteOnce. A run that seals storage references is planned before the
registry, secret material or any native step: every reference must name the
claim's current generation, the claim must bind to the run
(bindRuntimeLocalStorageClaimToAssignment), its volume must be ready, still be
the recorded directory and be held by nobody else, and no two volumes of the run
may nest or cover a path Pallet or the runtime binds itself (/run/secrets,
/run/serve.zone, /opt/serve.zone/runtime-assets, /proc, /sys, /dev,
/etc/hosts, /etc/hostname, /etc/resolv.conf). A run that cannot have its
volumes waits as storage-pending, with the step named (storage.claim.held,
storage.claim.importPending, storage.claim.generation, …). The grant is taken
in the transaction that begins the run, so a concurrent purge or a second holder
conflicts instead of both proceeding. Stop keeps the volume held; the proven CRI
absence of the removed container releases it. Release never deletes bytes: only a
purge generation does, and a held volume's purge is refused
(storage.purge.held) until its holder is removed. The container's mount list is
part of the attempt's identity, so every later run or inspect reconcile demands exactly the
grant's binds.
Offline import. An import claim waits as import-required until its bytes
arrive. With runtime-serve stopped, the operator places each volume's tree at
/var/lib/serve.zone/pallet-storage/import/<claimId> on the same filesystem and
runs pallet-control storage-import. The pass (in ts_migration/, like every
data adoption step) records the placed directory's identity, publishes it as the
volume by one atomic rename, fsyncs import/ and volumes/ and records it
ready; the tree keeps its ownership,
modes and every other attribute, and nothing is copied. It reports the adopted,
waiting and refused claims on its ready event (pallet.storage.import). A tree
on another filesystem, a file instead of a directory or a volume name already in
use is refused by name and moved nowhere; an interrupted pass finishes on the
next run, and a recorded tree found in neither place leaves the claim
uncertain. Keep a copy of the source first when a rollback may need it.
Network filesystems are not part of this build. The contract (32.24.0) also
states nfs and smb sources, but this node materialises node-local
directories only: such a claim is refused before the ledger is touched
(storage.source.unsupported), so neither the claim nor its service volume name
is recorded.
Dormant host/router handoff
The private ts/network/classes.handoffowner.ts mechanism consumes an already
reserved, digest-verified immutable handoff lease and a live PalletNamespaceOwner.
start() prepares an inert intent bound to the current boot and both namespaces.
The caller must persist that complete intent before reconcile(previousReceipt)
can create the exact veth pair, with its peer created directly in the router
namespace. Both endpoints remain administratively DOWN. Only the leased IPv4
addresses and their kernel local /32 routes are present; connected prefix routes
belong to eventual activation.
The separate --network-handoff-management native mode opens one netlink socket
in each retained namespace on its initial thread and restores the parent namespace
before polling either socket. It uses bounded native operations without shell
commands, named mounts or filesystem state. Linux ignores aliases in this veth
creation request, so the owner first verifies the newly created DOWN endpoints,
sets both ownership markers through link updates, and verifies them before adding
addresses. An interrupted partial creation never qualifies for adoption or deletion.
Recovery requires the exact boot, lease, namespace, link indices, peer indices,
names, MACs, markers, addresses and local routes. Extra or changed state is rejected.
Deletion uses the verified endpoint index and destroys only that exact pair. A
release states the administrative state its own contract guarantees: every
handoff path releases a dormant pair, so an endpoint raised behind the owner's
back is refused as handoff_conflict and the pair stays, and only the raise
measurement releases a raised pair. Removing a raised workload pair in
production is the stage-aware progress rollback, never this release.
Privileged exclusivity belongs to the enclosing node; these observations are not
a compare-and-swap guarantee against concurrent privileged network writers.
The TypeScript owner serializes operations and fences namespace or native-process
failure. getPrepared() and getReceipt() retain detached immutable evidence after
failure. close() joins the native lifetime; a successful return with
released: false means cleanup remains unconfirmed and the lease must remain
quarantined. A rejected close has not proved that join; joined exposes its state.
Join every namespace consumer before closing the keeper. Kernel namespace teardown
is asynchronous, so keeper exit alone cannot establish link deletion or pool reuse.
Both native handoff management modes use bounded nonblocking pipe or stream-socket stdio on the same thread that owns their namespace sockets. Host, router and retained sandbox netlink connections remain polled while input is idle or partial and while a response waits for the parent to read it. Partial input survives cancelled reads; malformed or oversized frames terminate the protocol. Output has the same three-second terminal bound as native commands. A partial response is never retried on the same stream. No reader thread can retain the mutation lock. This establishes protocol-owner liveness; the dormant sockets still have no topology multicast subscriptions and provide no continuous route or link fence.
This mechanism is not yet composed into signed node network admission. Durable
realization receipts, router/host/Docker policy barriers, CNI,
DNS, link activation and production qualification remain separate prerequisites.
The offline handoff fixture is test/native/handoff.ts, driven by
test/native/qualify-handoff.py in a disposable amd64 VM without a network device,
host disk or host mount. It runs the actual native binary through Node and compiled
Deno, including partial-state and foreign-state faults, restart recovery, keeper
loss and joined cleanup. ARM64 is built but this privileged fixture does not
establish ARM64 execution qualification.
Complete dormant handoff set
ts/network/classes.handoffsetowner.ts supplies the private node-wide mechanism
for complete retained membership. Start it with the live namespace, persist the
empty getPrepared() intent, then reconcile(target, null) to inspect and establish
an empty native receipt. prepare([{ lease, presence }]) accepts all 256 complete,
digest-verified leases permitted by the allocation contract, including
absent/quarantined history, with at most 32 desired present handoffs. Persist the returned
target and the last applied receipt before calling reconcile(target, previous).
Preparation does not change the kernel. Receipts contain an exact DOWN pair or
explicit absence for every member; absence never frees an allocation.
The --network-handoff-set-management process retains both namespace descriptors
and owns one host-network-namespace abstract Unix socket lock. The single-pair
mechanism uses the same lock, so no two Pallet handoff mutators can coexist on
that host namespace, even with different router namespaces. This IPC lock is
concurrent process exclusion, not authentication or durable ownership. The node
must still exclude other privileged network writers and namespace capabilities.
Before any transition effect, bounded complete link, address and all-family/table
route dumps inspect the whole expected set. The private router namespace must
contain only its empty, DOWN loopback and exact dormant handoffs. Unknown router
interfaces, addresses or routes fail ownership. Unrelated host interfaces and
routes remain permitted; unexpected Pallet markers/names, misplaced known MACs,
transit address conflicts and surviving/reused receipt indices are rejected.
Stable bound facts are re-read after inspection and after the complete transition.
Typed requests retain the terminal NLMSG_DONE; every inventory requires a
successful completion and rejects interrupted, filtered, malformed, unexpected
or oversized responses. Both namespace connections also reject unsolicited
messages and receive-buffer loss. The pinned maintained netlink-proto source
corrects decoding at its owning layer so malformed packets cannot disappear
before a later successful completion.
Transitions retain all prior members and cannot revive an absent member. Exact
present-to-absent deletion keeps the lease in the set. Separate pair operations
are not atomic: after a lost response, persistently prepared intent and the prior
receipt can adopt exact completed members and finish the remaining transition in
the same live namespace. Changed or half-created pairs stay failed-owned. EOF or
process termination preserves DOWN effects for that recovery. Full node death
does not make an unnamed namespace recoverable. close() deletes only exact
known members and joins the native lifetime; uncertainty preserves receipts and
returns released:false, with no guessed cleanup or allocation reuse. Over a router
namespace a service manager keeps, close() hands the set over instead
(handOverHandoffSet, "A stop hands the set over" under Router namespace lifetime): it
deletes nothing, proves the set equal to its receipt and returns released: true, handedOver: true.
The native transport permits 1,048,576-byte frames; TypeScript captures inert
complete-set data within the corresponding bounded budget. The serialization
test measures a 701,951-byte replay with maximum-length lease and
namespace identities. Native operations retain their three-second deadline and
the bridge its five-second deadline: the 256-retained/32-present guest operations
completed within the native deadline in the isolated fixture. Complete dumps establish absent members without issuing
redundant per-member link dumps. The offline handoff
fixture additionally covers empty/nonempty exhaustive inventory, unrelated host
links, competing single/set owners, a competing router namespace, quarantine,
lost-response and partial-transition recovery, 256 retained members with 32 present paths and joined
wrapper shutdown under Node and compiled Deno. This stage exposes no UP method
and issues no nativeBarrier: the authenticated SmartData application journal,
positive protection evidence and eventual firewall/DNS/VPN/CNI lifecycle remain required.
Retained uplink observation
After reconciling the complete dormant topology, the private handoff-set owner
can openUplink() and return a boot-, host-namespace- and random-generation-bound
observation. inspectUplink(generation) performs fresh bounded kernel and
networkd reads; copied facts cannot recreate the observer. A wrong generation
is rejected without retiring the current observation. closeUplink(generation)
joins the read-only observer and does not release topology or change DHCP.
A generation routes its transit out through the uplink, so the host must forward
IPv4 while it is UP. Who owns that switch is a fact of the host, which its
installer states on every start of the node (--forwarding-mode, see "Host
forwarding modes"); the observation follows it:
exclusive: the coordinator owns the switch. The observation opens only on a host that does not forward (all,defaultand the uplink link off; otherwiseuplink.forwarding.open), records the binding with those switches off, and from then on compares every read against it with the three switches at the value the coordinator holds. A sysctl write notifies with the kernel's own port, so the coordinator arms its switch before writing: the armed switch admits the forwarding-only NETCONF changes to the value it writes, once per index — the uplink,all,default, the held host link and every dormant link alike — and, switching on, the feature change of a link whose LRO the kernel turns off for forwarding; the read that follows drains every notification the write queued, proves the binding unchanged but for the three switches, and requires the kernel to have notified each of them.docker-shared: Docker owns the switch. The observation opens only on a host that forwards on all three (uplink.forwarding.openotherwise: Docker's switch is gone), and the coordinator never writes it.
Any other forwarding change of the binding, armed or not, retires the
observation, as does a NETCONF change of any other setting. An operation states
the forwarding it requires — exclusive: on exactly while the generation is UP;
docker-shared: on throughout — and is refused by name otherwise
(uplink.forwarding.on, uplink.forwarding.off).
The initial supported host has one physical Ethernet uplink, a finite bound
systemd-networkd DHCPv4 lease, one main-table IPv4 DHCP default route and the
three canonical unselected IPv4 policy rules. The observer records exact link,
address, gateway, source, metric and IPv4/IPv6 all/default/interface NETCONF
settings. Beside the lease the uplink may carry permanent /32 IPv4 host
addresses — the addresses a platform service on this host binds, whether
networkd configured them (Address=) or found them on the link. Each is a
universe-scope address with an infinite lifetime, so it adds no prefix route and
never becomes the default route's source; any other IPv4 address on the link
still refuses the observation, and adding or removing one retires it like every
other address change on the uplink. The observation states them
(hostAddresses, sorted as strings); an observation a 33.0.0 node journaled
carries none, which states nothing. Existing host IPv6 addresses remain possible; this observation grants
no IPv6 workload authority. Networkd remains the only DHCP mutator.
The native lifetime subscribes before capture and drives both RTNL notifications
and the fixed system D-Bus connection during inspection and quiet management I/O.
Notifications that arrive while the opening read runs, before any binding is
known, are held and judged against the binding that read finds, exactly like
later ones: the subscription predates the read, so the binding reflects every
change before them, a change that can affect it still fails the open closed, and
another link's churn during the read (a container starting beside Pallet at boot)
does not. The observation judges notifications ahead of its own work, so it
bounds their rate, not their number: more than 4096 within one one-second window of
BOOTTIME is a flood that would starve a read or the wait for the lease's renewal,
and it retires the observation; notifications it ignores never accumulate toward
the bound however long it lives, so continuous churn of other links, Pallet's own
workload veths included, never ends it. At most 4096 notifications are held while
the opening read runs.
Its unique networkd owner and exact returned link object path remain bound to
the lease. The observation follows the lease networkd renews, so a node keeps
forwarding through every DHCP renewal: at the lease's renewal time (T1), on a
property event of the link, and when networkd rewrites the leased address with
new lifetimes, the native reads the uplink and its lease again. While networkd
renews or rebinds, the lease in force holds until its valid lifetime ends, never
longer. A fresh acquisition from the same networkd owner for exactly the same
binding — link, address, prefix, gateway, source, metric, NETCONF settings and
host addresses — whose lifetimes the kernel address already carries becomes the
lease in force; the observation keeps its generation and its receipt, which
states the lease it opened on. The lease's expiry, service replacement, protocol
loss, a read that finds the binding changed in any fact or the lease moved
backwards, and every other relevant kernel change retire the coordinator.
Current values cannot erase an observed kernel change and restoration: only the
leased address rewritten in place asks for a read, and any other IPv4 address,
link or NETCONF change of the uplink, a NETCONF change of all or default, and
every IPv4 policy rule change still retire. Unrelated link notifications are
ignored.
A route change retires the observation when it can change the binding's path. The binding is IPv4, and its snapshot proves the three canonical rules, so an IPv4 lookup consults exactly the local, main and default tables (255, 254, 253) until a rule change retires the observation. A route retires it when it is IPv4, sits in one of those tables, and leaves through the uplink, has a destination that overlaps the leased subnet (which holds the leased address and the gateway; every default route overlaps it) or covers a host address, or has its gateway in the leased subnet. A route that cannot be judged exactly — a repeated or contradictory table or attribute, multipath, a nexthop object, an encapsulation, an unknown attribute or another family — retires it as well. IPv6 routes, IPv6 rules and the uplink's IPv6 addresses cannot change the IPv4 binding and are ignored, as are IPv4 routes of other links outside those prefixes and routes in tables no canonical rule selects. So another tool's links coming and going leave the observation alone: Docker starting a container or stopping, with the kernel's IPv6 link-local, multicast and local routes of its veths and bridges, or a bridge's IPv4 connected, local and broadcast routes. Pallet's own host links are judged apart. A route in a consulted table that leaves through the held (ACTIVE) host link or a dormant one, or whose destination or gateway lies in that link's transit prefix or in the pool routes it carries, and any IPv6 route naming it, belongs to that link: outside a pending host transition it retires the observation, and during one only the kernel's own such change reaches the transition, which accepts exactly the change its link state makes. Every other route, during a transition too, is judged against the uplink binding as above.
The kernel can stamp a change's notification with the requesting socket's port and sequence (an address or route notification does, networkd's lease rewrite included), so a sequence number never makes a notification Pallet's: only the port of the command socket a pending host transition bound does, and only for that transition's pool routes. Every other notification — the kernel's own or another requester's — is judged on what it changed; another requester's change to a host link in transition or to a route touching it is never the transition's. Close the observer before further dormant topology changes.
The ACTIVE observers of the router and sandbox namespaces tolerate one kind of
notification no matter the phase: a link change whose kernel change mask names
only IFF_PROMISC or IFF_ALLMULTI — what a capture on the link (tcpdump) or
an administrator's promisc/allmulti setting makes the kernel send. Such a
change forwards and drops nothing; the generation's inventory reads past a
link's promiscuity and all-multicast counts and flags. An idle coordinator reads
the generation's complete inventory again at once, exactly as an inspection
reads it (handoffset.drive.active_reverify names a failure of that read); an
owned step notes the change, and its own inventory reads the links. Every other
link change — up, carrier, MTU, name, address, a change mask beyond those flags
or none — still ends the observer. The dormant handoff-set audit stays strict.
The host's uplink observation is unchanged.
A coordinator that ends while it waits — its uplink observation retired, an
ACTIVE observer failed, a namespace channel lost — lowers every ACTIVE link on
its way out, as it always did. It names the watch that ended it in a bounded
PALLET_FAULT line (handoffset.drive.*, uplink.watch.*, uplink.follow.*);
an ACTIVE observer that refused an event names the event's kind in its own line
(active.observe.unexpected_event.<kind>: new_link_flags, new_link_state — any
link message without a change mask, such as a carrier, operational-state, MTU or master change — del_link,
new_address, del_address, new_route, del_route, new_rule, del_rule,
netconf, control or other), and the handoff-set owner writes its failure chain once, as it retires, to
stderr, which Spark carries into the node journal:
Pallet network handoff set failed: <chain>. The stop that follows names the
same chain as the reason its release could not be confirmed
(noderuntime.close.networkRetained < …). A command the coordinator refuses
names itself the same way (handoffset.command.<command>, after any line the
refusing check wrote), so a refusal answered over IPC with only its code still
reaches the chain as PalletNativeFault:<code>:<site>.
This mechanism exposes no UP or packet-readiness result. Active routing,
firewall composition, durable active journals and joined withdrawal still belong
to the activation coordinator. The ignored Rust test
network_handoff::kernel::uplink::tests::qualified_networkd_host_observation
provides an explicit read-only check on a real networkd-managed host; it changes
no interfaces, addresses, routes, sysctls or services.
Host forwarding modes
A host forwards IPv4 between all of its links once forwarding is on. Docker
switched it on when it started and set the iptables FORWARD policy to DROP, so a
Docker host forwards only what a rule admits; a host without Docker has neither,
and switching forwarding on alone would let a LAN peer use the node as its
gateway. Who owns the switch is a fact of the host, which the node never infers:
pallet-control runtime-serve and runtime-unit take --forwarding-mode <docker-shared|exclusive>, required, with no default (palletForwardingModes in
@serve.zone/pallet-bundle; Spark states it on fleet nodes). The handoff
coordinator receives the same mode (--network-handoff-set-management --forwarding-mode <mode>) and proves that mode's host conditions as it starts,
before anything else, refusing by name:
-
exclusive— the node is the host's only forwarding owner:- no persisted IPv4 forwarding: any
net.ipv4.ip_forwardornet.ipv4.conf.<scope>.forwardingother than 0 in/etc/sysctl.confor a*.confof/etc/sysctl.d,/run/sysctl.d,/usr/local/lib/sysctl.d,/usr/lib/sysctl.dor/lib/sysctl.drefuses (forwarding.exclusive.persisted): it makes the host forward unfiltered from boot until the coordinator starts; - no Docker under the service manager:
docker.serviceordocker.socketloaded and neither inactive nor failed refuses (forwarding.exclusive.docker_present); the exclusive drop and the switched-off forwarding would stop every container's forwarding, and a restarting Docker would switch forwarding on again behind the observation. A system bus with no service manager on it has no such unit.
Then forwarding is switched off, and on again only inside
activateActiveonce the generation is otherwise UP behind its exclusive host-transit table (see Retained uplink observation and Private packet policy composition). Right before that, the packet engine reads every owner on the forward hook (inspectForwardHooks, Smartnftables 4.2): each forward base chain must be the generation's own host table or the node's allocation-pool guard, by exact table name and kernel handle (active.activate.foreignForwarder), Docker'sDOCKER-USERandDOCKER-FORWARDchains must be absent (active.activate.dockerPresent) and no legacy xtables table may be registered (active.activate.legacyForwarder). The generation is refused before it is raised, and forwarding stays off. - no persisted IPv4 forwarding: any
-
docker-shared— Docker owns the switch:- the host provides the iptables-nft frontend the contribution runs:
/usr/sbin/iptables-nft,iptables-nft-saveandiptables-nft-restore, root-owned (the symlinks and their target), at 1.8.10 or newer on nf_tables (forwarding.docker_shared.iptables_nft); - forwarding is never written, at start, activation, withdrawal or exit; the observation binds only a host that forwards.
Every generation's host-transit policy is shared (no forward drop), and its flows pass Docker's
FORWARDpolicy through the node'sDOCKER-USERcontribution (@push.rocks/smartnftablesManagedDockerForwarding), which the engine applies only over Docker's policy DROP withDOCKER-USERandDOCKER-FORWARDas the first two jumps (its link fence follows below). Per generation, one owner (the node's owner name, a fresh instance kept in the journal,pallet_active_docker_forwarding): the first step over the generation's host table once it is applied and journaled, beforeactivateActive; a step over every later host table of the generation (tunnel phase, amendment, lowering), so the engine's reading of its barrier stays exact; each step admitted on the journal before the engine sees it and acknowledged after. The withdrawal releases it after the handoff is down and before the host table it admitted. A stop withdraws the generation it holds instead of leaving it UP, so no contribution rule outlives the node process; only a crash leaves one, for the next process to release. A fresh owner recovering a generation reopens the contribution under its journaled identity and releases it: an applied contribution throughrelease, and a last step that was admitted and never acknowledged without applying it (Smartnftables 4.6releaseTransition). Re-applying such a step would need its barrier — the host table it was prepared over — which the engine requires in this boot and may already be gone. The release request, the step's exact transition under the owner's engine identity of this boot, is journaled first, so a retry replays it exactly. The engine deletes exactly the rules of that step's target and previous contribution it still finds, and the end records the rule indices it deleted ({ kind: 'transition', request, removed }). A rule under the owner's prefix that is neither refuses the whole release by name before any deletion (docker.releaseTransition:FOREIGN_OWNER, a journal/table disagreement). The refusal settles nothing, so no successor begins (active.begin.previousContribution). Nothing retries it within the process; each restart replays the journaled request and meets the same refusal, failing closed until the table changes. The offending rule is one in tableip filterwhose comment carries the owner's prefix,snftd1:<ownerId>:(the owner name ispallet_docker_and 40 hex digits, journaled as the contribution'sidentity.ownerId; the comment continues with the instance, the contribution digest and the rule index), but that is not exactly one of the step's target or previous contribution rules inDOCKER-USER: another instance's or another step's rule, or one altered, moved to another chain or duplicated. An operator lists the owner's rules withiptables-nft -t filter -S | grep 'snftd1:pallet_docker_'and deletes the stray ones by their listed specification (iptables-nft -t filter -D <chain> <spec>). Deleting all of them is also safe: the release runs only once the generation is DOWN or being withdrawn with its handoff down, no successor begins before it settles, and it deletes only the step's rules it still finds, so the next start releases the step with whatever is left and records it. A contribution whose steps name an earlier boot ended with that boot, which took everyDOCKER-USERrule with it: it is recorded as ended with its boot ({ kind: 'boot', bootId }) and no engine is asked, since the engine refuses any step or release of another boot. That holds for whatever DOWN the generation carries and whether its cleanup finished, and also for a release begun and never confirmed (production 2026-10-08: a docker-shared node rebooted into 36.3.1 over a DOWN generation whose first step was never confirmed under 36.2.1 failed every start atdocker.reconcile:CONFLICT).test.activedockerboot.node.tscovers that journal, leftrecoveredanddiscarded, across a reboot (no engine call, a successor raised) and a release an earlier boot began;test.activedockerrelease.node.tsthe same journal in this boot (released throughreleaseTransition, never reconciled, before its host table), a lost release result replayed exactly, aFOREIGN_OWNERrefusal, and a DOWN whose cleanup never began. The orchestrator refuses to raise a generation whose policy does not match the mode (active.activate.sharedForwarding,active.activate.exclusiveForwarding), and a refused contribution step fails the activation by site (docker.prepare,docker.reconcile,docker.release,docker.releaseTransition).The contribution (Smartnftables 4.2) admits exactly the flow kinds the shared host table admits, each with its replies: the leased flows (the router's own egress, such as its VPN client dialling the hub, included), the inbound publications and the symmetric publications. Its link fence: the first step applies only with the generation's handoff down; a later step, rules changed or not, applies while the generation is UP when the handoff link both steps bind stays exactly the same — every step of a generation binds its one selected handoff, so a publication amendment, the tunnel phase and the lowering follow in place — and a link a step adds or drops must be down. The release needs the bound handoff down or gone: a link whose index no link holds any more counts as down, so a successor in the same boot — after a crash, its router namespace gone — releases the previous contribution before it applies the next. An owner that holds no rule of its own closes on any host (Smartnftables 4.3): one whose first step the engine refused — a host whose
FORWARDpolicy accepts — and one of a journaled identity on a host whose Docker and filter table are gone both close as released. A close that cannot confirm its cleanup failsCLEANUP_UNCONFIRMEDwith the engine's native code as its cause, which the node's failure chain names (PalletDockerForwardingError::docker.close < ManagedNftablesError:CLEANUP_UNCONFIRMED < ManagedNftablesError:<code>). - the host provides the iptables-nft frontend the contribution runs:
Moving a host from docker-shared to exclusive (the D9 step) is done with the
node stopped: state exclusive, stop the node (its stop withdraws the generation
and releases the contribution and tables; forwarding stays Docker's), remove
Docker and every persisted forwarding setting, and reboot. The rebooted host has
no table and forwards nothing; the coordinator's guards pass, and the first UP
generation switches forwarding on behind its exclusive table. Stating the wrong
mode fails closed and by name: exclusive on a Docker host refuses
forwarding.exclusive.docker_present, docker-shared on a host without Docker
refuses at the observation (uplink.forwarding.open) or in the contribution
(Docker's chains or policy absent).
The D9 safeguard is that order, and only that order: stop the node, remove
Docker, remove every persisted forwarding setting, reboot, then start the node as
exclusive. Never remove Docker from a running docker-shared host and carry on
without the reboot: Docker's switch stays on, and nothing in the node switches it
off (a docker-shared coordinator never writes it), while Docker's FORWARD
policy DROP — the only thing that kept the host from forwarding between all its
links — has no owner left to restore it: once it is gone with Docker's chains (a
firewall reload, a flush, the package's cleanup), the host forwards everything
until the reboot. An exclusive coordinator a rebooted host refuses (persisted,
docker_present) writes nothing, its exit included: the host's forwarding is
left exactly as the refusal found it.
The start guards have limits; they narrow the window, they do not close it:
forwarding.exclusive.docker_presentsees Docker only asdocker.serviceordocker.socketunder systemd on the system bus. A Docker daemon started outside systemd, or under another unit name (a snap'ssnap.docker.dockerd.service), is not seen at start. The activation-time read of the forward hook still refuses Docker'sDOCKER-USERandDOCKER-FORWARDchains, but it is a point in time.forwarding.exclusive.persistedreads the sysctl configuration files only. Forwarding switched on from the kernel command line (sysctl.net.ipv4.ip_forward=1), by a networkdIPForward=/IPv4Forwarding=setting, or by any other tool at boot is not seen: such a host forwards unfiltered from boot until the coordinator starts and switches it off. Once the coordinator runs, the observation refuses a host switched on behind it (uplink.forwarding.open) and fences a switch changed under an UP generation.
Downgrading a node from 36.1 to 36.0.3 or earlier is not supported once a
generation has detached an attachment in place: the transition row that records it
(detached, below) is refused by the exact reader of those releases, so their
process fails at start. A node that never detached one reads its journal under
36.0.x unchanged.
Downgrading a node from 36 to 35.x is not supported in place: 35.x cannot read
what 36 journals. It recomposes an exclusive generation's host policy without
exclusiveForwarding, so the journaled policy no longer matches, and it does not
know the contribution journal (pallet_active_docker_forwarding), so it never
releases a contributed DOCKER-USER rule. The contributed rules live only in the
kernel's ruleset, which a reboot empties; 36 records a contribution of an earlier
boot as ended with that boot and never releases it through an engine. A
contribution journal ended by a transition release (kind: 'transition')
is refused by 36.3.1 and earlier, so a node that wrote one cannot be downgraded
in place until two later generations were raised. Independently of the
journal, once 36.3.2 reconciled a table on a node, a downgrade to 36.3.1 or
earlier works only after a reboot: 36.3.2's Smartnftables 4.6 writes its
managed sets with their key and data types, which the Smartnftables 4.4.x of
those releases refuses to adopt, and a reboot empties the kernel's ruleset.
Private ACTIVE tunnel transitions
The retained PalletHandoffSetOwner accepts beginActiveTunnel(effect),
finishActiveTunnel(effect, expectedReceipt) and inspectActiveTunnel(effect)
while its exact ACTIVE generation is UP. The effect binds the complete ACTIVE
preparation and a creation plan or previous tunnel receipt to a durable intent
digest. It includes the authenticated SmartVPN generation, complete plan digest,
router namespace, exclusive interface name, /32 address, MTU, split routes and
same-boot BOOTTIME deadline. Tunnel routes cannot capture owned transit or local
workload prefixes, the control address or the hub endpoint.
Persist the intent before begin. Retain the actual SmartVPN client in the router namespace, prepare its tunnel, freshly inspect that generation, and derive the expected receipt from its actual interface index before finish. Native code observes the bounded ordered kernel transition, checks the full namespace twice and returns to its strict observer. It does not create the tunnel or infer the continued existence of the VPN owner from a saved receipt. Packet policy must be installed and verified before enabling VPN forwarding or raising workloads.
For deletion, stop forwarding while retaining the UP tunnel, lower the packet
policies back to their base shape while the device still exists, persist and
begin the deletion intent, then disconnect and join the VPN device owner before
finish with null, and release both packet policies only after the generation's
withdrawal. Each generation therefore carries a chain of policy pairs per role,
all journaled and each exactly one revision above the pair it replaced: the base
pair at the generation number (composed with no tunnel, so remote peers are
deferred rather than permitted), the tunnel pair (the same projection composed
with the retained TUN bound as the pallet_vpn endpoint), the lowered pair (the
base shape again) — a tunnel and a lowered pair for every device the generation
raised, each device's above the lowering of the one before — and between them one
amended pair for every workload attached while the generation is UP (below). A
generation raises at most 32 devices (maximumActiveTunnelCreations): every read of
the chain walks the tunnel lane back to its first command and revalidates each row,
so the lane bounds the cost of every tunnel-phase, amendment and cleanup step (64
rows at most on a lane this release admits). The journal refuses the next creation by name (exhausted,
active.tunnel.exhausted) before it admits a command. The orchestrator refuses it
before it asks for a credential to raise a device or dials, and refuses a
replacement under another credential before it lowers the live device or dials;
only the renewal's own credential request, which names the other identity, comes
first. The native's own bound is larger (256 session generations per generation,
maximumNativeTunnelGenerations, remembered so none is reused), and the journal's
must never exceed it. The pass fails with that name, its owner withdraws the
generation in order, and the next pass raises a fresh generation whose lane starts
empty. The bound governs admission only: releases up to 35.2.0 bounded a lane by the
native alone, so a lane they wrote holds up to 512 rows
(maximumStoredTunnelCommands). It stays readable — inspected, recovered, lowered,
cleaned up and withdrawn as any other, each chain read costing what it cost under
that release — and only its next creation is refused. The compiler releases only the exact graph
a transition names, so no revision is derived from a phase: every pair states
its revision and the acknowledged pair it replaced, and activePacketChain
proves the whole chain on every path that admits or acknowledges a pair above
the base pair, and before a recovery re-applies or a cleanup releases the live
pair. The tunnel pair is applied only after finishTunnel, because the
compiler binds the device by the interface index the receipt carries, and always
before forwarding is enabled; the lowered pair is applied before the delete
command, so no enforced graph ever names an absent link. Linux omits deletion
notifications for the static split routes; full inventory must prove their
absence. ACTIVE
withdrawal is rejected while a tunnel or unfinished transition remains. Unexpected
notifications, owner loss or the absolute lease deadline retire ACTIVE and fence
known interfaces DOWN. Recovery requires joining the VPN owner first and proving
the whole namespace through a fresh ACTIVE owner. No receipt grants allocation
reuse or proves packet drain.
The private protocol allows 2 MiB frames, including complete retained history and up to 1,024 split routes.
Both the durable tunnel journal and the application lifecycle composition now
exist. PalletActiveStore persists every managed TUN command as its own row —
beginTunnel admits one create or delete before the device owner is asked to
act, acknowledgeTunnel records the native begin, and finishTunnel records the
terminal fact — chained per ACTIVE intent so a delete consumes exactly the
receipt its predecessor left behind, and a normal withdrawal is refused while a
device is still retained or a transition unacknowledged. A recovery records DOWN
only; it never claims a deletion, and the retained row stays as history. The
creation row also carries the tunnel pair applied over the device (packets)
and, once the tunnel is being taken down, the lowered pair (lowered); a
cleanup row names the live pair — the last one the chain admitted — by its exact
transitions and policy digests, and the packet owner's release is fenced on
exactly those digests.
A workload may be attached while the generation is UP. The native admits the
pure prepareWorkloadAttachment under a held generation, and an
addWorkloadAttachment for a pair the generation does not hold yet is created
by the generation itself, through its own observers, while it is UP with no
tunnel transition pending, holds fewer than 64 pairs (prepared and attached
together) and has no tunnel reaching into the new subnet. 64 is what the packet
engine is measured to hold: Smartnftables 3.0 compiles a router of 64 workloads,
each with local DNS, public TCP and UDP egress, four platform endpoints and up to
three publications, inside its atomic batch with room to 89. The native registry,
its router-namespace inventory and configuration (one link per present handoff
and per pair, 96) and its admission fence (192 targets) are bounded to match; the
compiler's prepare() stays the authority on what a given graph fits. The bound above 32 is verified by compilation and unit tests only; the root QEMU guest qualification has not yet run at 64. The attachment owner
asks the held generation before it journals the ADD intent
(PalletHandoffSetOwner.admitAttachment,
handoffset.admitAttachment.notUp/capacity/tunnelOverlap); a refused ADD
releases its preparation and fails by that name with no journal, because a
pending one would fence every amendment of the generation while it is held.
addWorkload judges the same rules again before the native is asked
(handoffset.addWorkload.*), because a native refusal retires the owner. In the
other order a tunnel is refused where it would reach an attached pair's subnet,
by the store before the command is journaled and by the owner before the native
is asked. The pair is DOWN and outside every policy at that point. Before it is raised the
orchestrator amends the generation (PalletActiveOrchestrator.amend):
beginAmendment compiles the live pair again with the attachment as a further
source, in the live shape, one revision up, and appends the row
{ attachment, shape, revision, host, router } to the journal's ordered
amendments log in the same commit that fences the current source as an
extension of the one the generation carries — the same settled application,
every carried workload exact, the new attachment complete; finishAmendment
appends each role's acknowledgement after its apply. The intent, the native
preparation and the UP record never change: activeSources(journal) is the
intent's source extended by the log, and the tunnel pair, the lowered pair,
every recompilation and the cleanup compose over it. Every read of the journal
verifies each carried workload again against the completed application it was
admitted under (activeSourceHistory): the application the generation was raised
over for the workloads it was prepared over and every attachment amended in before
its first in-place transition, and transition N's application for the attachments
amended in after transition N. Amendments and transitions share the generation's
one revision sequence, so that order is the journal's own and nothing is recorded
beside it. Pallet 36.0.1 verified every amendment against the raise application,
so a node that attached a workload after an in-place transition could no longer
read its own journal: every pass failed, its tunnel lease ran out, and the node
failed on every start. The same journal reads again under this
release; such a node recovers without any manual step. A generation that already
carries the attachment has nothing to amend, a lowered generation takes no
attachment, and a node without a held UP generation amends nothing: the pair
is dormant and the next generation is prepared over it. DNS needs nothing new —
the completed attachment is a source of the next resolver revision exactly as
after a dormant ADD. A recovery sends the native attachedWorkloads, the
complete receipts of the restored pairs the preparation does not cover
(required, [] when none); the native restores DOWN only when its registry is
exactly the preparation's workloads plus those. A restored attachment it cannot
list — an ADD whose receipt never became durable, or one whose removal is
already intended — is removed first, in the fresh process and before its native
adopts anything, and journaled exactly as a DEL journals it
(PalletWorkloadAttachmentOwner.recoveryRemovals). A failed journal step, a
foreign intent or a failed native removal refuses the recovery by its step
(handoffset.recoverActive.beginRemoval/removalIntent/incompleteAttachment/finishRemoval).
A pair raised under a held generation stays UP through its withdrawal, and the
withdrawal restores the router namespace's baseline configuration. The raise is
therefore refused by name (active.raise.ipv6_baseline) unless that baseline
keeps IPv6 disabled on the pair's router side — all/disable_ipv6 must be 1
before the generation is prepared — because IPv6 enabled again on an UP link
brings addresses and routes no owned transition admits. The same rule applies
to pairs raised before the generation is prepared (active.prepare.ipv6_baseline).
The native router owner establishes and verifies all and default IPv6-disabled
baselines immediately after creating its private namespace, before exposing it.
The router namespace is fail-closed between generations for the same reason: the
withdrawal releases both packet policies while the raised pairs stay UP, so the
IPv4 forwarding it restores must be off. The native router owner switches all
and default forwarding off at creation, and a generation is prepared only when
every forwarding switch it captures — all, default, lo and each router link
— is off; otherwise the preparation is refused by name
(active.prepare.forwarding_baseline). Activation turns forwarding on only after
both packet policies are applied, and the withdrawal turns it off again before
the policies are released, so workload traffic through the router stops from a
generation's withdrawal until its successor is UP (the withdrawal window of a
projection change the generation cannot move to in place, below) and is never
forwarded unfiltered.
For a containerd sandbox, the admitted native ADD disables IPv6 only on its
newly created, identity-proven DOWN eth0 before addressing it; the exact pair
DEL removes that link. Containerd runs CNI before creating the sandbox container,
so CRI sandbox sysctls cannot provide this pre-CNI guarantee. The sandbox's
namespace-wide defaults and the host's sysctls are never changed.
A node with the mpls_router module loaded is not supported: every link
registration emits an MPLS netconf record, which the attach vocabulary refuses.
The tunnel credential is the controller's. PalletControllerClient.resolveTunnel
issues getRuntimeManagedVpnCredential { session, projection } over whatever
transport the runtime session uses — the relay when the node is bound to one —
for the exact projection the generation carries — the one it was raised from, or
the one it last moved to in place (below) — and a reading of the qualified clock, and turns the answer into the session request the device
owner authenticates with: the cluster hub's address and port, key material,
authority and node ids, and a same-boot BOOTTIME deadline bounded to fifteen
minutes or the credential's own expiry, whichever is sooner. The node dials its
own cluster's hub, never a platform-wide one: the credential names the hub
endpoint this node's signed projection selected, and the contract binds that id,
its transport's protocol, its address and its port to that selection. The
transport decides the dial form — a managed QUIC hub is a bare host:port, the
one transport the contract states — so a credential naming any other is refused
rather than dialled as if it were QUIC. Every refusal is named
(controller.tunnel.notLive, disconnected, projectionAbsent,
projectionChanged, boot, credential, transport); the key material is returned to the
caller only and is never journaled or logged. The orchestrator asks for it once
the generation is UP and DNS is serving; a refused credential or a refused
authentication leaves the generation UP without a tunnel, reports the refusal
on the barrier as tunnelFailure (code and site, nothing else) and on the
application pass as activation.tunnel:<label>, and the next kept pass asks
again. Once the command is admitted every later step is a journaled transition
and fails the activation as such.
A tunnel lease lasts at most fifteen minutes and is moved in place, never by
tearing the tunnel down: every admitted DNS lease renewal moves the window the
credential lives in, so the next kept pass asks the controller for a fresh
credential, and a pass that finds the lease deadline within three minutes of the
qualified clock does the same (renewalMarginMs). Under the identity the session
authenticated with (hub address and key, client key, authority and node) both
bounds on the device move in place: first Pallet's own native authority deadline
(PalletHandoffSetOwner.renewActiveTunnel, native renewActiveTunnelLease), which
the native arms at the lease the device was created under and at which it fails
the generation closed, then the session's lease with SmartVPN's
renewManagedSession (PalletManagedVpnOwner.renewLease). The device, its packet
policies and forwarding are untouched, and nothing is journaled: the deadline
bounds this process's custody, and a successor fences the generation without
it. A credential under another identity cannot renew a session, so the tunnel is
lowered and a new one raised with it. A refused renewal leaves the lease in force,
names the refusal on the barrier's tunnelFailure and is asked again on the next
pass; a refused native renewal never asks the session, so the session's lease
never outlives the native's.
The hub may move a live tunnel's split routes in place. A SmartVPN 2.5.0 hub that
offers assignment updates sends the node its committed assignment, and the client
moves the device's routes through its own netlink socket — every withdrawn route
deleted before any new one is added — before Pallet learns the new assignment, so
the native observer sees the route events first. It accepts them provisionally,
and only on the held device: an exact split route (main table, static, link
scope, unicast, no gateway, source or metric) to a private prefix of eight to
thirty-two bits that is no default route, overlaps neither the device's address
nor its remote, no other route and nothing the generation owns (transit,
workload and attached subnets), and, for a deletion, one the device carries. Any
other route event — another device, a default or public route, a foreign shape —
fences the generation exactly as before. From the first change the routes are
pending: they differ from the journaled receipt, every inventory read expects the
provisional set — a read that raced a move is read again once the events the kernel
already reported are taken, at most four times — and they must be named by a journaled revision within sixty
seconds (active.observe.tunnel_routes_deadline otherwise; the bound is twice the
thirty-second request budget one queued native command may hold the revision
behind, and the BOOTTIME start never moves, so nothing extends it). The
orchestrator follows the session's managed-session-assignment events
(IPalletTunnelSessionOwner.watchRoutes): it journals the revision first — its
number and the exact routes, address-ordered, in its own collection
(pallet_active_tunnel_routes, one row per device, the adopted revision and at
most one pending, each later than the last) — and then hands it to the native
(reviseActiveTunnelRoutes, PalletHandoffSetOwner.reviseActiveTunnel), which
re-dumps the namespace and adopts the revision only when the device carries
exactly its routes with everything else unchanged; the adoption, with the BOOTTIME
instant the routes first differed, is journaled after. A revision that differs, a
stale one or a deadline that passes is a conflict, named on the barrier's
tunnelFailure, and fences the generation as any other. Device steps — raise,
lowering, renewal and revision — run one at a time. A replayed revision is
adopted once; a revision journaled before a restart stays history, because the
device ends with the process and the next process fences the generation. A
device's deletion consumes the receipt of its last adopted revision. The router
policy keeps reaching the tunnel's peers by the prefixes its projection states:
routes ahead of the projection carry nothing until the generation moves to that
projection.
A projection change the generation can follow moves it in place
(PalletActiveOrchestrator.transition). The application owner decides it before
any dormant effect: the next application intent is the exact successor of the one
the held generation carries, within its epoch, over the same uplink observation;
it keeps the dormant handoff set and the receipt its predecessor completed with,
the selected handoff and the host routes the generation installed; and its
projection continues the node's managed VPN membership
(continuesRuntimeManagedVpnMembership, @serve.zone/interfaces 32.42.0: the same
handoff and projection lineage, a later generation, and a protected authority that
is the same or exactly its next step that only adds — a node joining the egress, a
hub joining the platform endpoints; workload endpoints, egress, private networks
and the router selection may move). Such an intent is completed without the native
— its receipt is its predecessor's — and the held generation is not withdrawn.
beginTransition then admits, by the same rule again, the pair that moves it: the
live pair compiled again over the successor application, in the live shape, with
every attachment the live pair carries that the successor's source still admits,
one revision up, publications refused for capacity by name exactly as a
fresh generation refuses them, in the commit that fences the current source
exactly; finishTransition journals each role's acknowledgement after its apply.
Each transition is one row of its own collection (pallet_active_transitions,
numbered per generation, at most 64 per generation), and the chain proves it like
an amendment. The intent, its native preparation, the UP record, the device, its
session and its lease stay as they were; the packet policies' protected authority,
workload endpoints, egress and the router's tunnel peers move. The application a
generation carries is its last transition's once both roles acknowledged it, its
own before any; while a transition is half applied it carries neither, and only a
fence settles it. Renewals ask for the credential of the projection it carries. A
positive protection receipt for an application is appended only while a held UP
generation carries exactly that application — both roles of its last transition
acknowledged — and its row is fenced in the receipt's commit, so the receipt of a
new protected authority follows the policies that enforce it; the receipt's native
barrier stays the guard intent, which a new authority always moves. Anything else
— a successor that ends the membership (a protection step that removes or changes
anything, a skipped step, another handoff), another target or receipt, other host
routes, a rebound uplink, the transition bound, a refused or failed transition —
withdraws the generation and raises a fresh one, as every projection change did
before. A process that ends between a transition's roles leaves the live pair to
the cleanup the next process runs after it fences the generation.
A transition moves past an attachment the successor no longer carries instead of
refusing. The current source (PalletWorkloadAttachmentStore.fenceActiveSources)
names every retained attachment of the epoch it does not carry, and why:
removing once its DEL is intended, absent once its absence is proven, and
unadmitted while it is live but the application no longer reserves and projects
its lease. A carried attachment that is released any of these ways is left out of
the successor pair, and the transition row names it in detached (the sorted
attachment ids; omitted when the transition detaches nothing, so such a row keeps
the bytes an earlier release wrote); the chain applies it, so every later pair
composes over the attachments the live pair still carries, and an amendment extends
the live pair rather than everything the generation ever carried. A live,
unadmitted attachment is then detached (PalletWorkloadAttachmentOwner. detachAttachments): its run reads as of an ended epoch, so its execution owner
stops it — and fails it network-epoch-ended while its controller still asks it to
run — no name is served for it, and its DEL takes the pair out of the generation;
the policy that no longer carries it already denies its flows. A fresh generation
is prepared only over the whole epoch. Before carrying pairs into it, activation
reconciles released attachments while no generation is held, at most eight per pass.
An existing removal keeps its original operation and ACTIVE binding. For an
unadmitted complete pair, the current settled application's lease withdrawal
authorizes a durable released-attachment network removal under the application
fence. Its digest binds that completed application's reference, whose retained
signed source proves withdrawal on every replay. It releases an ended creation binding without needing another CRI DEL or
an assignment lane a failed run already freed. Runtime and storage authority do
not change. The native removes only the exact journaled pair and proves absence,
including when its sandbox vanished. Absence is persisted before activation;
neither elapsed time nor attempts erase a row. The release logs removing,
absence proven, or a bounded failing-step label with activation waits.
Unfinished rows retain their intent for the next network pass or owner restart,
and the run continues to read as of an ended epoch after this removal. While any
released attachment lacks durable absence, activation refuses by name
(active.activate.releasedAttachment) until its DEL proved it absent. Pallet 36.0.2
refused the whole transition instead (attachment.activeSource.unadmitted): every
stop of an attached workload — a spec change, a deletion, a rollout — withdrew the
generation, and with it the node's tunnel, for about two and a half minutes, and an
attachment removed under a held generation broke its next amendment or transition
the same way. Cloudly 33.13.2 keeps a stopping run's lease reserved until the node's
stop receipt, so a stop normally reaches this node as a DEL first and a retired lease
second; either order moves in place.
A DEL takes a carried attachment out of the live pair before its pair is deleted
(PalletActiveOrchestrator.detach, called by PalletWorkloadAttachmentOwner once the
removal is intended and before the native removal): beginDetachment compiles the
live pair again without it, in the live shape and over the live application, one
revision up, and appends it to the amendments log as a detaching row
{ attachment, shape, revision, detaches: true, host, router } — attachment exactly
as the pair carried it — in a commit that fences the attachment's own removal intent
(attachment.fenceRemoval.notIntended otherwise); finishAmendment acknowledges each
role. A detachment is admitted after the generation's last device was lowered too, in
the base shape, so a DEL never waits for the next device; no other amendment and no
transition follows a lowering. activeSources leaves a detached attachment out, activeHistoricalWorkloads
keeps every workload the generation ever held for its recovery, and
activeSourceHistory verifies a detached attachment where it was admitted. The packet
engine binds every link a policy names and refuses one that names a link the kernel no
longer holds (ENODEV, which Smartnftables 4.3.0 reports as UNAVAILABLE), so a pair
must never outlive the links it binds. Pallet 36.1.5 and earlier deleted the pair under
the live pair, and every pair compiled over it later was refused: a node stop that
followed a DEL failed active.stop.loweredPackets < Error < ManagedNftablesError:UNAVAILABLE,
joined the device owner with the device's deletion never begun, and the handoff set
retired on that unarmed deletion (active.observe.unexpected_event.new_link_flags: a
link closed while UP is reported with IFF_UP in its change mask before its
RTM_DELLINK, the first event of an armed deletion); a pass's withdrawal and an
amendment after such a DEL failed the same way (lab 2026-10-04, N7 and run 3's N4b). A
journal that carries a detaching row is refused by Pallet 36.1.5 and earlier: once one
exists, downgrading the node in place to 36.1.5 or earlier is not supported.
A lease that lapses anyway ends the native authority at the same instant: the
native fails closed, the generation is fenced and the node fails owned, and its
supervisor starts a fresh process, which fences the generation by the proof that
its epoch ended and raises a new one with a fresh credential. A session whose transport to the hub is lost stops
instead: SmartVPN ends its packet I/O and reports managed-session-state
stopped, keeping the device for an ordered teardown. The barrier stops holding at that moment —
holds requires the session's packet I/O to be live — so attached workload links
demote on their next CHECK and admission refuses, and the next node scan runs a
network pass at once. That pass lowers the dead tunnel through the journal exactly
as a withdrawal does (the lowered packet pair, the delete command, the device
released) and closes its owner, then asks for a fresh credential and raises a new
tunnel over the same generation; a refused credential leaves the generation UP
without a tunnel, named, until a later pass succeeds. A retired session — its
device lost behind the owner — is a loss of the owner, not a stop, and fails the
activation as before.
PalletNetworkApplicationOwner takes an optional activation option and, when
it is present, drives one ACTIVE generation after each completed dormant
application: preparation, both packet policies applied and journaled, native UP,
local DNS proven to be serving, then the tunnel commands above and finally
forwarding. A native preparation whose intent commit is rejected (the source
superseded while the policies compiled, active.source.fence) is discarded before
the rejection is reported: the handoff set admits nothing else under a held
preparation, and no journal row would name it. Without that option the owner stays exactly as dormant as before and
the DOWN wire contract is unchanged. PalletNetworkOwner.barrier() exposes the
result as one of absent, prepared, up or forwarding, with the retained
device's journal reference when there is one. Only forwarding sets holds, and
only holds may gate workload UP; a node with no managed VPN configured settles
at up and its barrier deliberately does not hold, with one exception: the
single host. A node bound to a local controller (Onebox) never dials a tunnel,
whether or not a managed VPN is configured — the released bundle configures one on
every node, and the persisted binding decides — and a generation reaches
forwarding without a tunnel when three facts hold: its scope is that local onebox controller, every endpoint of the signed
projection it was raised from is placed on this node (the contract already
refuses a remote endpoint in a onebox projection; Pallet proves it again from
the journal row and fails closed), and this process raised the generation itself.
There is no remote peer to reach, so the base packet policies already carry every
rule the generation needs. A barrier that does not hold names the missing fact as
singleHostGap (controllerNotLocal or remoteEndpoint). Execution admission
then rests on the single-host evidence instead of a tunnel receipt: the evidence
names the projection, the admitting transaction re-proves the fact from the
journal row and fences it, and a refusal is admission.barrier.singleHostUnproven.
A tunnel command on such a generation is a corruption and refuses. Every other
node — bound to Cloudly directly or through the relay — keeps the mandatory tunnel
barrier unchanged. The barrier is always
recomputed from the journal and the live owners, never read back from a stored
pass. On loss this composition fences and withdraws first and then reports a
barrier that does not hold; it never re-activates the lost generation by itself,
because re-activation requires explicit newer state from the control plane.
Store reads under load
Every store read hands each row to its collection's validator, and the same rows —
application intents and results carrying the complete signed projection, signing
revisions, transitions, retained workloads — are read on every network pass, DNS
refresh and admission. With a 64 KiB projection Pallet 36.0.2 spent a whole core
re-validating them (lab 2026-10-03). Each validator now runs once per exact document
per process (memoizedDocumentAssertion, ts/storage/validation.ts): the verdict is
remembered by a SHA-256 over the document's JSON, which states every value it admits
exactly — a non-finite number, negative zero, undefined, a Date or other class, a
proxy makes a document unkeyed, and an unkeyed document is validated every time.
Validators are pure functions of the document, only acceptance is remembered, and a
document that differs in one byte is validated in full. The asynchronous proofs of
immutable rows — the retained projection's signature and validation, signing
revision, workload and handoff lease digests, guard host grants — run once per exact
input the same way (verifyOnce); an application intent's whole retained lease set is
proven once per exact set. The retained workload rows are read once per change: the
projection store keeps the rows an application-source read last found, keyed by the
stored revision of the projection row that read took in the same transaction. Only an
admission writes the workload rows, and the same commit always writes the projection
row and so moves its revision; rows and key therefore come from one snapshot, every
application-source read of every pass shares them, and a transaction whose snapshot
predates an admission keys what it read by the old revision, which no later reader
holds (test.projectionworkloadcache.node.ts). The memo holds at most 8,192 entries and
lives in the process only. An identical projection push, which a controller repeats on every
delivery round, is a replay: the store compares the exact envelope it admitted under
the signer it still trusts, writes and fences nothing, and the controller client
schedules no network pass for it. test.storebudget.node.ts bounds all of it over a
64 KiB projection with 122 retained workloads: after the first pass, a pass validates at
most 16 documents, reads no workload row and makes at most 300 row reads and proof
lookups in all (measured 249; 36.0.2 made about 2,700, most of them the retained
workloads read a dozen times); eight identical pushes validate none, read no workload
row and schedule no pass; and after a changed projection the next pass reads the
retained rows once and the pass after it not at all.
An idle node does constant work per second, and a controller's re-delivery writes nothing
(lab 2026-10-03, Pallet 36.1.0: about half a core with two idle workloads, and a new
attached assignment never admitted inside the relay's 8 s timeout while three ran). The
execution driver keeps its --executor-management child between calls and ends it only
after a call that failed and on close, so the scan's runtime inspections and clock
readings are IPC round trips instead of a process spawn each. A run whose reconcile found
it waiting only for its next readiness probe is settled until that probe, and at most
palletSettledRecheckMs (5 s): while its durable state is exactly the one it was settled
on — and, for an attached run, its pair is still of the current network epoch — the scan
answers it readiness-pending from those reads and one clock reading, without inspecting the
runtime or opening a transaction. A run whose network preparation was refused
(IPalletExecutionNetworkOwner.prepareExecution, for example
attachment.preparation.endpoint.check) is answered from that refusal — the same thrown
refusal, or the same wait line for a secret run — while its durable state is unchanged and no
network pass ended since (networkPasses), and at most palletRefusedPreparationRecheckMs
(5 s): the preparation is fenced against the completed application the last pass selected, so
nothing else a scan sees can change it, and the scan asks the controller for no registry
credential or secret material meanwhile (lab 2026-10-04 RUN 4, P2: two refused runs held an
exclusive node with no workload at about 40 % of a core). Assignment histories and configs are
proved once per exact content (verifyOnce), and the delivery pass reads no terminal
outbox of a first generation, which has none. An assignment re-delivery that is a replay
or a historical answer is neither fenced nor gated: it writes nothing, so it no longer
competes with every other fenced commit of the node for the identity row, which is what
made a new admission lose its commit on nearly every attempt. The execution gate composes
its barrier evidence once per admission and every retry of the admitting transaction only
fences it again. test.admissionbudget.node.ts (eight re-deliveries take no identity
fence and validate nothing; a new admission commits on its first attempt while every
running assignment is re-delivered back to back), test.admissionguardretry.node.ts (a
retried attempt composes nothing) and test.idlebudget.node.ts (one child for seven
driver calls; ten scans of a settled run spawn nothing, open no transaction and make no
runtime call; a delivery pass reads no first-generation outbox; ten scans of a run refused at
its preparation ask for no credential and open no transaction) bound it.
Nor does an idle node's work grow with its history (lab 2026-10-05 RUN 6, P2: an exclusive
node with no workload at all held 30–35 % of a core, against 4.8 % on a node that had run
nothing yet). Every second the scan, the delivery pass and the log reconcile each read,
decoded and proved every assignment the node ever held, and the database read each of
those rows for them, though a retired assignment — removed, its terminal receipts and
removed observation acknowledged, of which the controller keeps no record — never changes
again. The assignment store keeps the ids of the assignments it has not proved retired,
built from the whole inventory by the first scan of a process; listCurrent and
listReportable read only those rows, by id, and an id proved retired leaves for good. An
admission that was accepted, or failed and may have committed all the same, adds its id
once its transaction settled, and an id with no row is dropped by a later scan that no
admission overlapped. list still reads
the whole history. The log owner reads the whole inventory on its first pass, which reopens
what an earlier process left of a retired assignment, and afterwards lists only what is not
retired plus the state of each capture it still holds. test.idlehistory.node.ts (an idle
tick over three and over six retired assignments reads no assignment, outbox or slot row;
a new admission is live at once) and test.logowner.node.ts (an ended capture whose
assignment retired while it waited is still delivered) bound it.
A network pass whose application source is superseded while it runs — a projection,
signer or allocation admitted between its read and its commit — is refused by name
(projection.fenceApplicationSource.superseded) and changed nothing; the next scan
runs a pass at once. The guard's side of a pass loses the same way: a projection
admitted between the guard pass's read of its bootstrap source and its proposal of the
expansion it prepared against it is refused as guard.propose.superseded; native
preparation is inert and the proposal did not commit, so the pass ends without retiring
the network owner, and the next one expands the guard to the newer source. Any other
refusal of a guard pass still retires the owner. The node writes Pallet network pass deferred: <site>. once and then only with its count, instead of a failure line
(test.networkpassrace.node.ts, test.networkguardrace.node.ts).
Durable application intent and results
The private identity runtime exposes runNetworkApplications() over its existing
SmartData/NoSQLDB owner. readSource(scope) returns the current verified signed
projection and complete retained lease history. Native preparation happens outside
the database. begin(scope, fingerprint, prepared, guard) then persists that exact
whole-set target, its previous receipt and the authenticated source, writing the
common trust/allocation fences and the real identity guard in the same transaction.
Only the current reserved handoff is desired present; every other retained lease
remains explicitly absent. Omitted history and revived absent members are rejected.
Ordinary successors remain within the same boot and native namespace.
The journal keeps one pending transition and immutable intent/result records.
Exact retries join existing intent; a competing proposal cannot replace it.
finish(scope, intentReference, nativeReceipt) records only the result matching
that persisted target. It does not require new admission authority: signing-key,
projection and credential changes must not discard an already-owned native result.
Historical signatures remain verifiable through retained signing revisions. Lost
completion acknowledgments replay exactly without advancing or rolling back the
current head. No journal method deletes history or frees quarantined allocations.
Keep begin, native reconciliation and finish inside one admitted runtime callback
so database shutdown joins the entire operation. Retained facade methods expire
when their callback returns. The persistence capture budget covers both transition
sides and complete signed source; native frames retain their independent 2 MiB
limit. This store does not start a native owner, issue nativeBarrier, report
protection, clear DNS withdrawals or authorize workloads. The private application
owner supplies node composition and a separate durable unavailable-report outbox.
beginEpoch(scope, fingerprint, prepared, evidence, guard, workloadProofs) records a qualified
same-boot absence or different-boot fence into a fresh namespace. The caller must hold the
fresh private native owner continuously through its audit, journal commit and
reconciliation. The store resolves the evidence's opaque journal reference to
the actual persisted pending intent, or applied intent when none is pending,
and binds its complete target and exact known receipt. A single transaction
retains the old applied/pending references in an immutable epoch record, fences
current signed authority and identity, and publishes the next intent. Evidence
from a different target, previous receipt, host, boot or namespace is rejected.
The first intent in an epoch has previousEpoch and no native predecessor:
reconcile it with previous = null. Its generation follows the old pending or
applied head. The still-reserved logical handoff may be realized DOWN again;
every retained lease and quarantine survives. Old unfinished intents stay in
history without a fabricated result and reject late completion after the epoch
fence. Previously recorded results still permit exact acknowledgment replay.
An intent has one shape: it names the epoch it opens or null, and a record
that does not is refused rather than read. Repeated failures before a fresh
intent completes continue the same history.
Different-boot recovery requires the distinct native previousBootTerminated
proof, persisted as new-boot-host-fence; it cannot use a same-boot absence
receipt. Reused numeric namespace identities do not make two kernels the same epoch.
Dormant workload attachment journals
The private ts/network/attachment/ composition binds one admitted pending Run
to its immutable workload lease, completed application and exact container ID.
Native preparation opens the sandbox namespace once and retains its descriptor.
The ADD journal commits before any veth mutation; its immutable identity cannot
be rebound after cleanup. Ambiguous admission acknowledgements are resolved on
the same application transaction lane before an unadmitted descriptor is discarded.
Partial creation uses the recorded removal intent and exact native absence check.
The sandbox accepts containerd's internal loopback baseline. Workload links stay
DOWN until this node's activation barrier holds; an ADD waits for it, bounded.
Application and attachment operations share one queue. Handoff retirement must fence every unfinished attachment in its transaction. A replacement router waits for distinct historical workload evidence: stable absence in the exact retained sandbox on the same boot, or termination of a different actual kernel boot. A missing or substituted sandbox path is not same-boot absence. Previous interface indices belong to their original namespaces and are never host capabilities.
An application epoch record has one shape and always binds a complete sorted workload proof set. Immutable coverage records commit with that epoch and freeze the old journals. Coverage permits historical DEL to acknowledge the proved absence, while preserving the original unfinished journal, all allocations and quarantine. It does not create an ordinary DEL receipt or authorize another ADD. Quiescence stops preparations; already admitted ADD/DEL operations drain before the shared native owner closes.
Production execution admits network: 'attached' only after the assignment's
exact projected lease is prepared with the pending Run operation. The CNI owner
then proves and journals its attachment before returning a usable result. An
isolated assignment prepares a separate durable CNI intent with no lease or
application reference. A secret reference is resolved by this node: its material
is staged and published as this attempt's own mount before the container that
receives it, as "Staged secret material mounts" states. A storage reference is
granted from this node's claim ledger in the transaction that begins the run, as
"Local storage claims" states.
Private CNI transport
The node runtime owns a root-private CNI broker at
/run/serve.zone/pallet/cni.sock. The installer supplies the trusted
/run/serve.zone parent; the broker verifies its ancestry and provisions the
0700 leaf and 0600 socket. The native executable enters CNI mode through
--cni or the installed basename pallet-cni. The client checks path ownership,
socket identity and the kernel-reported server UID before sending one bounded
frame. It never stores journals, allocates addresses or opens sandbox namespaces.
Transport cancellation does not cancel an admitted attachment operation.
Detached handlers retain broker capacity until their native effect and journal
write finish. An invocation the plugin stopped waiting for while the owner still
held it — its connection closed, by the plugin or by a broker that is closing, or
the broker's own 30 s ran out — is written as
Pallet workload attachment abandoned: run <run digest> <command>: <transport.closed|broker.deadline>; the owner continues.,
and what the owner answered after as
Pallet workload attachment answered after abandonment: run <run digest> <command>: <status or label>.,
each once, then with its count while it repeats. Pallet 36.3.6 and earlier wrote
neither, so an ADD that bound its run after the runtime had given up on it left
nothing in the node's journal (Grasberg 2026-10-10). Node shutdown joins CRI, including synchronous rollback DEL, before
closing the broker; the broker then joins all admitted handlers before its
borrowed owners can close. Replaced socket paths are never unlinked during cleanup.
The fixed profile uses CNI 1.1.0, the single pallet-workloads configuration and
containerd's use_internal_loopback=true. VERSION reports protocol support. DEL
can acknowledge exact recorded absence with empty stdout.
A refusal of the network owner names its check. The attachment owner refuses by site
(attachment.live.<state>, attachment.cleanup.<state>, attachment.reconcile.* such as
attachment.reconcile.unprepared, attachment.isolated.*, attachment.epochEnded); the
broker answers { error: 'rejected', refusal: 'refused.<code>:<site>[:<native site>]' }
and writes Pallet workload attachment refused: run <run digest> <command>: <label>. to
the node's stderr, once, then with its count while it repeats; the plugin answers the CNI
error Pallet network owner refused the attachment; cleanup may still be required with
the label as its details, which containerd states after the message, and writes it on
its diagnostic line (PALLET_FAULT site=cni.refused code=<label>). A transport failure,
or a broker answer without a label, is still Pallet network owner unavailable; cleanup may still be required. An ADD that arrives while the attachment owner waits for its
network epoch to be proven (pending-epoch: it started over an attachment of an earlier
epoch that is not yet proven gone, which the application's next pass settles) waits for
the owner to become ready, within palletAttachmentWaitMs (20 s from its arrival, shared
with its wait for the barrier, inside the broker's 30 s request timeout), and is then
judged as any ADD; a close while it waits
refuses it at once (attachment.reconcile.closing). An ADD journals its pair only within
that same wait: one that reaches the application's lane after it, behind a slow lane or
an earlier ADD of its run, or would commit its journal after it, is refused
attachment.reconcile.expired, journals nothing and discards a native preparation it
already made. Its caller has given up on it by then, and a pair journaled later bound the
run for good to a sandbox the runtime had already torn down (Grasberg 2026-10-10). Every
other refusal — a CHECK in
pending-epoch, a closed or failed owner, 64 operations in flight
(attachment.reconcile.busy), a missing preparation — is answered at once (lab
2026-10-05 RUN 6, O1: one run's sandbox was refused three times in six seconds as
"network owner unavailable", with nothing to say what had refused it). The broker answers
its own refusals the same way, before the attachment owner sees the invocation:
refused.unavailable:broker.busy while 32 earlier invocations, abandoned ones included,
are still held, refused.unavailable:broker.notListening while it closes, and
refused.unavailable:broker.identityStopped once the node's identity runtime stopped; a
STATUS request is then closed unanswered. Pallet 36.3.6 and earlier closed such a
connection without an answer, which the plugin reports as "network owner unavailable".
A connection beyond 32 open at once is still closed by the socket server before the
broker sees it.
A quiescing owner still admits ADD and DEL for runs already admitted, as before.
For an isolated assignment, the pending Run transaction durably binds its
assignment, operation, boot and fixed CNI configuration. ADD opens the exact
containerd sandbox namespace and proves it is distinct from the host and router
namespaces. Containerd 2.3 requires a real eth0 with an address even for an
isolated pod, so the native owner creates a namespace-local Linux dummy device
with 192.0.2.1/32. It has no peer, uplink, default route or host grant;
IPv6 address generation is disabled and every address, neighbour and route in
both families is checked. The kernel's own ff00::/8 local-table multicast
route is allowed only when it is bound to this peerless dummy, with no IPv6
address, gateway, RA or unicast route. The first ADD commits the exact
container ID, namespace identity and double-proven initial absence in SmartData
before NEWLINK. The completed ADD then records the dummy index, deterministic
MAC and double native proof. If a process stops between NEWLINK, alias, address
and UP, replay may remove only the exact DOWN partial dummy under that durable
intent, prove its absence twice and recreate it. Ambiguous or foreign state
remains quarantined. CHECK requires a completed journal and reproves its exact
endpoint. The CNI result reports the real dummy and address,
with no routes or DNS grants. DEL verifies and removes that exact device before
retiring the journal; replayed DEL is idempotent. A DEL for a sandbox whose ADD
never journaled its intent is answered absent without a native step: the dummy is
created only under that intent, so nothing of Pallet's is in the sandbox, and the
runtime's teardown of a refused sandbox completes (Pallet 36.2.0 refused it as
attachment.isolated.delUnjournaled, and containerd kept the sandbox NOTREADY).
No workload lease, veth or activation grant is made for this mode.
The isolated reconcile runs on the handoff-set coordinator, which owns no state in
the sandbox, and is admitted while an ACTIVE generation is held: Pallet 36.2.0
refused it there (the coordinator's ACTIVE command filter did not list it), so
every isolated run of a node whose generation was up had its sandbox refused, and
the owner over the coordinator retired on that refusal, failing the node runtime
(owner_failed; lab 2026-10-06 RUN 9, D11). A command refused under a held
generation now writes handoffset.active.withdrawal_required and the command's
site as fault lines. What the native refuses of an isolated sandbox itself — its
path gone, a namespace that is not a fresh sandbox, or anything its inventory or
its dummy's lifecycle finds — is answered handoff_sandbox_refused: the set is not
marked failed, the coordinator and any generation stay, and
PalletHandoffSetOwner.reconcileIsolatedSandbox rejects that call alone with
conflict at handoffset.isolatedSandbox.refused, carrying the native fault lines;
the run's sandbox setup fails and the execution owner handles it as any refused
sandbox. An attached workload is answered the same way when its sandbox's path is
gone as its preparation captures it — the runtime removed the sandbox while its ADD
waited (workload.prepare.sandbox_gone): PalletHandoffSetOwner.prepareWorkload
rejects that ADD alone with conflict at handoffset.prepareWorkload.sandboxRefused,
no ADD bound the run's attachment, and the run is driven again as for any refused
sandbox. Pallet 36.3.6 and earlier answered it unavailable, which ended the
coordinator and every attachment with it (Grasberg 2026-10-10). A failure of the
coordinator itself (its watches, entering or leaving the namespace, the command's
deadline) still retires the owner. Native CRI evidence must report only
192.0.2.1 for an isolated running sandbox. Host-side TCP and HTTP readiness
probes are refused before effects because they would target host services.
An ADD whose pair is admitted, journaled and complete while the activation barrier
does not hold waits for the barrier: it leaves the application's lane — the
generation that raises the barrier needs it — and is woken by the ACTIVE
orchestrator's barrier signal (PalletActiveOrchestrator.watchBarrier, the one
source; the attachment owner reads the barrier again on each wake), by its own
owner's state changing, or by its close. Its two waits, for a pending network epoch
and for the barrier, share one bound, palletAttachmentWaitMs (20 s) from the
ADD's arrival, inside the CNI broker's request timeout and the plugin's own 30 s
deadline; every exit clears the wait's timer and listeners. On a node without
activation there is no barrier to wait for. Once the barrier holds, the held
generation's policies are amended to carry the pair when they do not yet, and the
pair is raised and proven in the same step; a barrier still not holding at the
deadline answers the pair DOWN with an activation error and the label
barrier.unheld:attachment.reconcile.barrierWait, and the runtime tears the
sandbox down. That sandbox was the run's one binding, so the run then fails
instead of being driven again (see "Private CRI execution"). A refused amendment or raise leaves the attachment admitted,
journaled and DOWN and names its step on the pass (amend.<code>:<site>,
raise.<code>:<site>); it never fails the ADD. The plugin writes that label only
to its own stderr, which containerd discards, so the CNI broker writes it to the
node's stderr as well:
Pallet workload attachment down: run <run digest> <ADD|CHECK>: <label>., once,
then with its count while it repeats. A successful ADD returns
a CNI result carrying the interface, its address, the node resolver and the one
route the raise installed: 0.0.0.0/0 via the router side of the pair
(static, priority 100). Only state this owner installed is reported, and the
advertised route is proven equal to the route found in the pod namespace.
CHECK reports the raised attachment while the barrier still holds and
demotes to the same DOWN activation error the moment it is lost, without
touching kernel state. A repeated ADD on an already raised attachment is
idempotent and never destroys it. STATUS reports authenticated broker and network
owner availability to containerd, independently of any attached activation
barrier; it cannot grant a workload network. An attached ADD or CHECK without the
exact live barrier still reports DOWN, while an isolated ADD uses only its own
pending Run preparation. GC cannot infer
cleanup authority from a runtime-provided attachment list and returns an error
until an admitted removal or historical proof can be resolved.
Application owner and node lifecycle
PalletNetworkApplicationOwner holds one private handoff-set owner for the node's
router namespace. Each bounded pass keeps journal inspection, completion of a
retained current-epoch intent, current source selection, epoch qualification,
admission, native effects and completion inside one identity-runtime callback.
It derives complete membership from the authenticated store; RPC callers cannot
provide a prepared target, native receipt or absence proof. A stale source or
identity rejects new admission. Completion of an already admitted effect survives
credential changes, controller disconnect and quiescence.
The private PalletNetworkOwner composes this application owner with the guard
and protection outbox. It restores the existing offline guard before starting
the application or controller. While connected, the node reconciles the network
with the actual physical session's private identity guard before scanning
assignments whenever a network pass is due: the first scan, every scan after new
network input from the controller (a new connection, an admitted signing
authority, projection or DNS lease renewal), after a rejected pass, as soon as a
managed VPN tunnel stopped on its own, and otherwise at most networkRefreshMs
apart (default 30 s; the assignment scan itself runs every pollIntervalMs,
default 1 s). A full pass re-proves the guard with a fresh engine process and
re-reads and re-verifies the application and its signed source, so it does not
run on every scan. The standing DNS refresh runs every 30 s for the same reason;
every change DNS depends on (an application pass, a workload ADD or DEL)
refreshes it at once, and the resolver child enforces its lease's BOOTTIME expiry
itself. A refresh whose DNS store read or intent the database refuses changed nothing
native: that pass ends alone, with the resolver child joined, and the owner stays ready.
DNS is down until the next refresh serves again, within 30 s; the node writes
Pallet DNS refresh refused: <chain>., once, then with its count while it repeats. A
DEL whose refresh is refused so still removes its pair, and an ADD's refusal answers
that ADD alone. Pallet 36.3.6 and earlier retired DNS and with it the node on any such
refusal, and a DEL that met one closed DNS for good (Grasberg 2026-10-10). The
protection delivery loop decides that nothing is pending from its
lane row alone. Without current
authority it waits; there is no offline network admission path. The status snapshot
reports the application reference and admission rejection without exposing source
records. Native ownership loss or an unrecordable native result fails the node
with ownership retained for cleanup.
Shutdown stops admission, joins any pending begin/effect/finish callback, then joins the native handoff owner before closing the namespace and database. An unconfirmed native join keeps those resources owned; failed cleanup cannot report a clean stop.
A normal stop (SIGTERM or SIGINT, or stdin EOF for runtime-serve) withdraws an
ACTIVE generation and journals it DOWN in either forwarding mode before a clean
router namespace handover. The native refuses to withdraw a generation that
still holds a device, and a device owner joined without a begun deletion takes
its device away behind the native, which retires on that unarmed deletion;
either way the handoff set's release stays unconfirmed. So the orchestrator's
close refuses every new command and joins the ones it admitted before; then it
lowers a held tunnel exactly as a withdrawal lowers it — forwarding stopped, an
amendment or in-place transition pair whose command ended short of its acknowledgement
applied again on its unacknowledged roles and acknowledged, the lowered packet pair applied, the deletion journaled and begun, the
device owner released, the deletion finished — and only then joins the device
owner. It then begins and confirms the native generation withdrawal, journals
DOWN, releases the packet policies and any Docker contribution, and discards
the preparation. The native's closeHandoffSet (handOverHandoffSet over a kept router namespace, which
deletes no pair: "A stop hands the set over") then closes the dormant set or hands
its retained pairs to the successor without deleting them. The process ends with
stopped and exit code 0. The journal records the
generation DOWN with its tunnel lane settled on the deletion and its policy
cleanup complete. The successor can adopt the clean namespace and raise a fresh
generation; only an unfinished failed or crashed lifetime needs the ended-epoch
fence (the retained-generation paragraph under "Standalone control build"). The application
owner joins every workload ADD and DEL before that close, since the lowering
runs outside the lane an ADD's amendment and raise run on; so no other journal
or packet work interleaves with the lowering. A stop over a generation with no
tunnel withdraws the generation without any device command. A lowering or withdrawal
that fails is an owner failure: the stop still
joins every owner, fails with the failed step as its cause and ends
owner_failed, and the successor fences the generation as after a crash. The
step is named active.stop.<step> (settleTunnel, settleLive, stopForwarding,
loweredPackets, beginDeletion, nativeBeginDeletion, acknowledgeDeletion,
releaseDevice, nativeFinishDeletion, finishDeletion, and for a withdrawal
beginWithdrawal, nativeWithdraw, finishWithdrawal, releaseDocker,
releasePackets, discard, closeSession), with the step's own failure below it,
and the network owner's close carries what refused it, so the node's
Pallet runtime-serve failed: … line reads down to the refusing step.
The lowered pair is compiled only over an acknowledged live pair. A stop completes a
pair left short of that (settleLive) without checking it against the current
source — the controller may have removed an attachment it carries since the admitting
commit fenced it — so it does so only once forwarding stopped, moves no DOCKER-USER
contribution (Docker's FORWARD policy keeps admitting only the host table the
contribution last followed), and the lowered pair replaces it at once. A pass's
lowering refuses such a pair (active.prepareSuccessor.unapplied). While the
attachment the pair extends is still in the source, the amendment or transition
command that admitted it replays it against the current source. Once the controller
has removed that attachment, no pass can complete the pair: passes stay refused until
the node stops — whose settleLive completes it — or the next process fences it, so
the recovery is a node restart. This needs an amendment whose apply was refused on
one role and whose attachment was then removed.
A normal stop of a docker-shared node instead withdraws the generation it
holds — handoff down, the Docker forwarding contribution released, then the host
table — hands the dormant set and every attached workload pair over intact to the next
process over the kept router namespace ("A stop hands the set over"), and never writes
the host's IPv4 forwarding, at stop or at exit; the
process ends with stopped and exit code 0 and no contribution rule outlives it
("Host forwarding modes"). The stop carries no attachment past the generation it
withdrew: the application owner has joined and closed its attachment owner before
the orchestrator's close, and the next process of the epoch carries the
attachments at its start, before any pass ("Attachments outlive their
generation").
Without the activation option all handoffs remain DOWN. With it
the application owner opens the physical DHCP uplink observation before the
orchestrator and raises one ACTIVE generation over it on the first pass that has
an application; a generation that is UP for the same application over the held
observation is kept on later passes, an observation that no longer proves itself
withdraws the generation bound to it and is reopened on the next pass, and a
failed activation is fenced and named on the pass (activationFailure). The same
label is written once to stderr, which Spark carries into the node journal, as
Pallet network activation refused: <label>; a packet engine refusal adds the
engine request and hop (preparePolicy, host or router) and the engine's code
and bounded refusal text, and any other refusal that has a cause adds that cause's
chain (< <chain>, class names, codes and sites, cut in the middle so the line
stays within the 512 characters Spark forwards): a packet operation the label can
only call unavailable still names the engine request's code beneath it. A refusal repeated on every pass is written again with
its count one minute after it was last written, then at doubling intervals up to
ten minutes, and a refusal that returns after activation held is written at once.
Every network store conflict names the check that refused — the ACTIVE family as
active.<owner>.<check>, every other store as <store>.<method>.<check> — and the
type of requireNetwork and of PalletNetworkStoreError makes the site mandatory for
conflict, so the label always says which one. Without
a managed VPN the generation settles at UP and the barrier deliberately does not
hold, so attached workload links stay DOWN — except on a single host bound to a
local controller, whose barrier holds as described above whether or not a managed
VPN is configured. Lease release and quarantine remain
with Cloudly's allocation owner; Pallet resolves only an exact assignment-bound
projected lease and cannot create or release allocation authority.
Durable protection and withdrawal reports
Every connected network pass freshly recovers the actual allocation-pool guard, completes the dormant application journal, then attempts a positive receipt. Only a real guard owner that has reconciled, inspected enforcement, detached its PERSIST policy and joined its native child can supply that private capability. The returned bootstrap JSON and a durable guard reference alone cannot mint it. The capability expires when the owner closes; composition keeps it through the outbox transaction. RPC never accepts a native proof or guard selector.
Fresh positive minting fences the settled current guard and application lanes,
current signed projection and signing key, retained allocation ledger and actual
physical reporter/identity in one SmartData transaction. Guard and application
must agree on source, boot, host namespace and protected authority. The report's
nativeBarrier derives from that exact completed guard intent. It establishes
allocation-pool denial, without claiming workload connectivity, DNS readiness,
packet drain or lease release. Every handoff remains DOWN.
Every positive report also states the node's own uplink address — the address
a route to its published ports uses, read from the retained uplink observation
— on the request envelope (uplinkAddress, @serve.zone/interfaces 31.4+).
The address is a fact of the pass that minted the receipt: a changed or absent
address is a new receipt generation, so the pass after a rebind reports it
absent (the stale observation is closed) and the next pass, over a fresh
observation, reports the new one. Nothing defaults or infers it, and a
withdrawal states none.
Beside it the report states hostAddresses (@serve.zone/interfaces 32.23.0):
every address the same observation binds on the uplink — the lease address and
the permanent /32 host addresses beside it — sorted as strings. Cloudly names a
plan platform endpoint in this node's hostPlatformEndpointIds when its address
is one of them. They are read together with uplinkAddress from one observation
and are part of the receipt's identity in the same way: another set is a new
receipt, and a report without an uplink address states none. Only the uplink's
addresses are stated, because the packet compiler serves a host-local platform
endpoint only on an address the uplink binding carries. An outbox entry a 33.0.0
node wrote carries none and is read as stating none.
The SmartData outbox retains immutable receipt revisions, original reporter bindings and completed-application references. Positive reads verify the historical completed guard too. One pending receipt supplies backpressure; a newer application cannot replace it. Historical replay uses those immutable proofs without requiring old guard/application heads to remain current. Exact ACKs fence the current physical connection and identity. Reconnect sends the original receipt body, then a later fresh native pass may queue a current-session successor. Replaying historical positive evidence does not re-establish current-session eligibility.
A stop withdraws nothing (@serve.zone/interfaces 32.44.0,
runtimeNetworkProtectionContract.withdrawal). Quiescence stops admission and
joins the native pass and the execution owner; the last receipt stays current and
the node stays eligible for allocation, because its allocation-pool guard stays in
the kernel when the process ends and its boot unit restores it before networking.
A clean stop, a restart, a crash, a reboot and being offline are all the same to
Cloudly: a pending receipt stays in the outbox for the next session's historical
replay, and a reconnect binds a new reporter session whose first positive receipt
restores eligibility. Physical disconnect alone does not remove Cloudly's durable
allocation eligibility.
Only Cloudly's retirement request withdraws. It travels in the signed projection
(IRuntimeNetworkProjection.retirement, { disposition: 'withdrawn', protectedAuthority }): Cloudly signs it into a node whose egress ownership it
retires with the withdrawn disposition, its first appearance names exactly the
projection's current protected authority, and every successor carries it
unchanged (runtimeNetworkProjectionRetirementContract). Cloudly sends it only to
a node that reports a compatible interfaces version: Pallet's registration offer
states the release it carries (protocol.interfacesVersion, 32.50.0, beside
minimumPeerVersion 32.0.0). The network pass that completes an application whose
admitted projection carries the request latches it for the node's receipt chain,
once, in the same commit as its admission guard (pallet_network_protection_retirement),
and mints no positive receipt; from then on no pass does, and the store refuses one
(protection.enqueueOwnedProtected.retiring), across restarts. The controller
client answers before the pass resolves: it drains and ACKs any pending predecessor
(a positive receipt minted before the request still drains), appends the chain's
exact null successor and requires its ACK. The null retains its predecessor's
historical application, authority and boot, so the answer works during key and
projection gaps; its reporter is the current physical session. A withdrawal without
a latched request is refused (protection.enqueueWithdrawal.notRetiring). A failed
answer leaves the immutable pending history intact and fails the pass, naming the
step it failed on (controller.withdrawal.*) and that step's own error; the next
pass reads the latched request again and finishes the answer. The node's
allocation-pool guard is never released for it: the answer withdraws eligibility,
not protection.
A report the controller, or the relay carrying the session, refuses fails as
PalletControllerRefusalError:<reason>:controller.refusal.<method> with the
peer's TypedResponseError as its cause, in every chain it ends up in — a
retirement answer's controller.withdrawal.predecessor included. <reason> is the most
specific word-only token the peer stated: its error payload's reason, else
its payload's code (Cloudly answers a refused report with
{ code: 'runtime-session-report-refused', reason }), else its refusal text
when that is plain prose (words and spaces, at most 96 characters), turned
into hyphen-joined lowercase words, else unstated. A token is lowercase
letters in hyphen-joined words, at most 64 characters: no digit, dot, colon,
slash or other punctuation passes, so no address, path, digest, session or
node id reaches a label. A word made only of letters does pass as the peer
stated it, so prose that names, say, a letters-only host name carries that
name into the token; Cloudly's payload reasons and the relay's refusal text are
words from their own fixed vocabularies. A
delivery pass that fails while the client still admits work writes
Pallet controller delivery failed: <chain>. (the owner_failed labelling:
class names, codes, sites) to stderr, which Spark forwards to the node journal,
once per distinct failure until a pass succeeds again: a receipt the controller
keeps refusing is named when it is refused, not first at the stop it then
fails.
A refusal holds back only what the contracts order behind the refused report.
A delivery pass sends four kinds of unit, in this order: the protection chain,
the application receipt chain, each assignment of one bounded page, and each
capture's workload log batches. Within a unit the order is the contract's:
protection receipts and application receipts are each one monotonic chain with
one pending receipt, so a refused receipt holds every later one of its chain; an
assignment's terminal receipts go in sequence and before its observation (the
controller accepts a removed observation only after the joined removal
receipt), so a refused one holds that assignment's later reports; a capture's
batches go in sequence. Nothing orders one unit behind another — a protection
receipt neither gates nor proves workload readiness, and assignments are
independent of each other and of protection — so a refused unit ends only
itself. The pass goes on with every other unit and, once it is done, rejects
with every refusal still standing on the connection (one, or an
AggregateError of all of them), those of units still waiting out their retry
delay included, which the stderr line above names. A refused unit then waits before the
delivery cadence sends it again: the delivery interval, doubled with each
refusal in a row, at most 30 s (refusedReportMaximumRetryMs), while every
other unit keeps the cadence. A unit that delivers, or a new controller
connection, starts again from the interval. flush() sends every unit at once,
whatever it waits out. Any failure that is
not an answered refusal — a timeout, a lost connection, an answer that does not
bind to its report, a store or capture failure — still ends the pass; when
refusals stand, the pass rejects with an AggregateError of that failure first
and every standing refusal after it, so a failure that recurs on every pass
never hides a refusal from the stderr line.
Projection application receipts
A projection carries its predecessor's packet-grant withdrawals, and may not
grant a withdrawn flow again, until the receiver supplies the exact previous
projection as applied (admitRuntimeNetworkProjection's appliedPrevious).
Pallet states that fact as IRuntimeNetworkProjectionApplicationReceipt
(@serve.zone/interfaces 32.50.0): its durable claim that every packet table it
holds was composed from one admitted projection, or that it holds none, so the
flows its kernel admits are within that projection's grants.
Only the network pass mints one, after it completed an application
(PalletNetworkOwner.reconcile, beside the protection receipt), and only from
the ACTIVE journal's own evidence, fenced in the minting commit
(PalletActiveStore.fenceApplicationReceipt):
- the UP generation, neither withdrawing nor DOWN, carries exactly that
application, its last in-place transition acknowledged on both roles
(
activeCarriedApplication), so every pair the kernel may hold — that transition's and every amendment or tunnel pair after it — was composed over it; on a docker-shared node the generation'sDOCKER-USERcontribution, a rule set of its own that follows each host table after it is acknowledged, must have its last step acknowledged on a host table of that same application. The receipt'snativeJournalnames the generation; - or the node holds no packet table at all (
nativeJournal: null): no generation was ever journaled, or the head generation is DOWN with its packet cleanup recorded and complete — both tables released in its own epoch, the host table that survives a same-boot epoch proof released through the engine, none left after a reboot — and its contribution released or ended with its boot.
Anything else mints nothing: an activation not yet UP, a transition awaiting an
acknowledgement, a withdrawal in progress, a DOWN whose tables are still held
(an epoch-proven generation's surviving host table included). The head is the
only generation that can hold a table, because a successor is journaled only
once its predecessor's tables are released or are the very tables its policies
replace, and only once its predecessor's DOCKER-USER contribution is released
or ended with its boot (active.begin.previousContribution). A DOWN whose
contribution release failed or never ran is released by the next pass's fence
(fenceRetained) before it raises a successor. With no generation journaled, the
receipt and a first generation's intent both write the application lane, so they
serialise. An admission ACK, a table inspection or a success flag never mints a
receipt, and no transport facade can mint one
(runNetworkApplicationReceipts reads and acknowledges only).
The receipts form one chain (applied:<scope digest>), kept in the identity
database beside the projection chain they name
(pallet_network_application_receipt_lane, …_entries, no migration). A new
receipt is minted only for a newer projection than the latest one names; a pass
whose application names an older one fails (applicationReceipt.append.regressed),
and projection admission refuses a projection older than the latest receipt's
(projection.admitProjection.behindReceipt), so the node never applies packet
authority from a projection older than its latest receipt. The chain continues
across restarts and reboots and never restarts at generation 1 while the
database holds it; losing the database loses the node's enrolled identity with
it, and a restored older copy's next receipt is refused by the controller, which
holds a later one, exactly as that copy's older projection chain is.
The node's own projection admission reads its latest receipt in the admitting
transaction and passes selectRuntimeNetworkProjectionAppliedPrevious(receipt, previous): the previous projection when the receipt names exactly it, else null.
With it, a successor that drops the predecessor's withdrawals, or grants a flow
withdrawn earlier again, is admitted; without it, both are refused as before.
One receipt is pending at a time. The controller client sends it on the current
authenticated session (reportRuntimeNetworkProjectionApplication), requires
the answer to name the exact receipt sent, and its authority use to be current
for a receipt minted under this session and historical for one minted under an
earlier one, then acknowledges it in the outbox. It sends receipts only to a
controller whose accepted registration offer states interfacesVersion 32.50.0
or later (controllerReadsApplicationReceipts); an older controller keeps the
session and the receipts wait in the outbox. Pallet's own offer states 32.50.0,
which is how a controller learns that this node mints receipts and passes its own
receipt to its admission.
test.activeapplicationreceipt.node.ts covers the evidence: none while a
generation is prepared but not UP, a transition is pending, a withdrawal is in
progress, a DOWN's tables are held or an epoch-proven host table survives.
test.activeapplicationreceiptdocker.node.ts covers a docker-shared contribution
that follows or lags its host table, a predecessor contribution that survived its
DOWN, and a first generation begun while a no-table receipt commits. test.networkapplicationreceipt.node.ts
covers minting with no table, one pending receipt, admission with and without a
receipt, the chain across database restarts, historical replay and offer gating.
Offline allocation-pool guard journal
PalletNetworkGuardOwner composes the private identity database with published
Smartnftables 2.1. commission() explicitly creates the first guard intent;
recover() requires an existing journal. The caller supplies a trusted native
binary path and must join this owner before closing the database. The node runtime
composes this private owner with application completion and positive reporting;
Spark installs the separately commissioned boot dependency.
The fixed receiver-owned trust record supplies the node scope, checked against
the active local identity. Bootstrap verifies the latest admitted projection with
its retained signer, including a crash between key rotation and the next signed
projection. That historical material supplies denials only; DNS and new projection
admission still require the current key. The policy denies exactly the declared
allocation pools, without treating management LANs or resolvers as allocation pools.
Its only exceptions are the exact host grants of the source's signed onebox
projection and the publications its signed projection places on this node (see
Published ports). They are derived from that verified projection whenever a target
is proposed and proved again from the retained source on every read of an intent, so
a stored row never widens the guard by itself. A guard Pallet 36.2.0 or earlier
committed carries no publications; it is read as applied, narrower than its source,
and the next guard pass appends the same-source transition that adds them, as it does
for the stream port owner below. Every policy also restricts the CRI
stream server of Pallet's own containerd,
127.0.0.1:10010, to uid 0 (localTcpPortOwners, Smartnftables 4.0.0). A pass
refuses unless the policy it applied carries that owner; a guard committed before
the owner existed gains it through one appended transition of the same source.
SmartData persists each complete native identity and original previous-Applied/null
to prepared-target transition before reconciliation. Immutable result records bind
the full native receipt to that intent. One pending lane prevents competing targets.
Same-boot recovery repeats the original transition, including after a lost reply or
completed result. A different actual boot records the old applied/pending heads in
an immutable epoch and restores the prior denial with a fresh process instance and
previous: null. A changed namespace in the same boot is rejected.
Recovery restores the committed denial before applying a newer inert projection.
Every previously guarded pool must remain with the exact same id, purpose and
prefix. Native preparation must reproduce the durable target, and exact enforced
inspection and verified detachment must complete before success. An ambiguous
failure uses closeRetaining() to join the child without deleting its policy or
inventing an Applied receipt. No journal result alone grants allocation eligibility.
test/native/qualify-guard-boot.py runs real NoSQLDB and nftables in four isolated
offline root boots under Node and compiled Deno. It covers same-boot completed and
pending recovery, lost/undelivered native calls, new-boot epochs, signing-key gaps,
newer staged pool expansion, and actual pool denial with unrelated traffic controls.
On both boots a root TCP connection to 127.0.0.1:10010 succeeds and the same
connection as uid 65534 is refused, before and after the new-boot recovery.
It also exercises PalletGuardProcess and reopens the same database after its
oneshot completes. Supplying --control-directory dist_control/linux-amd64-<digest>
also runs the built production pallet-control guard-recover with null stdin and
verifies its completed events, database release and continued pool denial.
Production still requires the commissioned Spark boot dependency, and live
activation the qualification of the complete workload networking, DNS and VPN path
on a real host.
Host kernel requirements
The packet engine (Smartnftables 4.6) requires Linux 6.9 or newer, with the
nftables filter and NAT modules (nf_tables, nft_nat, nft_chain_nat, nf_nat
and the reject modules the port owner rule uses) available. A kernel below that
floor, or one lacking a feature a policy needs, is refused by the engine as
UNSUPPORTED_KERNEL; Pallet names that refusal instead of hiding it. The guard
owner and PalletGuardProcess fail with code unsupported_kernel, and
guard-recover, guard-commission and runtime-serve end with failed reason
unsupported_kernel on their process protocol, exit code 1. Every other engine
failure stays owner_failed. A host must boot a supported kernel before Spark
commissions the guard.
PalletGuardProcess verifies the guard executable against trusted bundle metadata,
opens only existing enrolled storage and runs one offline recovery or commission.
It joins the native owner before stopping NoSQLDB. A native cleanup failure keeps
the database lease owned for an explicit close() retry. Cancellation joins
admitted work and fails the oneshot; success follows verified policy retention,
native child exit and database shutdown.
The compiled control accepts guard-recover for ordinary boot and
guard-commission for explicit first commissioning. Both accept exactly one mode
argument and no runtime configuration. Null stdin is valid for these oneshots;
unexpected input or SIGTERM/SIGINT fails without readiness. Successful completion
emits ready and stopped under pallet.guard.process after cleanup. These
events provide no positive allocation or workload protection receipt. The installer
must acquire its first authenticated projection and commission before enabling
recovery as a boot prerequisite.
test/native/qualify-protection-boot.py separately qualifies the complete private
network owner in four offline Linux6.18.35 boots, under Node and compiled Deno.
It verifies actual nftables denial with independent UDP controls, real dormant
handoffs, completed application/guard agreement, historical positive replay after
reboot, a fresh new-boot positive, the null withdrawal that answers a retirement
request carried by a projection a newly advanced signing key signs,
and PalletNodeProcess cleanup that reopens the same database. The packet client
uses the pinned Node executable for both runtimes so a Deno UDP compatibility
error cannot be interpreted as enforcement. No kernel, native-owner or database
result is stubbed in these guests. They have no network device or host mount.
Supplying --control-directory dist_control/linux-amd64-<digest> also runs the
built production pallet-control runtime-serve on each runtime's second boot.
It verifies readiness, EOF-driven shutdown, clean child exit, database reopening
and continued pool denial after the entire node owner has stopped.
Initial projection acquisition
pallet-control network-acquire opens only existing enrolled storage and a
projection-only controller connection. It waits for signing authority and a
projection delivered over that actual physical session; a retained older projection
cannot satisfy acquisition. It rejects assignments and does not fetch registry
credentials or deliver terminal, observation or protection reports. After admission
stops, it joins every controller write, verifies the delivered projection under
the final current key, and closes controller/storage before ready and stopped
under pallet.network.acquisition. Null stdin is valid; input, signals or the
bounded operation timeout fail the oneshot after joining its owners.
pallet-control network-acquire-local is the same acquisition for a node driven
by a controller on its own host (Local controller). It serves
the local socket instead of connecting to a relay, and it additionally admits the
first bind of a node no controller has bound yet; a node with any Cloudly
enrollment state refuses to start it. The store must already be provisioned.
pallet-control store-provision provisions it: a null-stdin oneshot on the
pallet.store.provision protocol that creates the data root and the database,
prepares every collection, runs the store's migrations and releases the database
again. It writes no enrollment, identity, binding or authority record and serves
no socket, so the node stays unbound until network-acquire-local admits its first
bind. An existing store is accepted as it is; a store that carries a Cloudly
identity is refused (PalletStoreProvisionError:cloudly_enrolled on the
owner_failed diagnostic line). enrollment-provision also creates the store and
leaves no Cloudly state behind by itself, but it serves the Cloudly enrollment
socket for its lifetime, where a prepare request writes that state; a node for a
local controller is therefore provisioned with store-provision.
The installer runs acquisition while management networking is available and before
first guard-commission and boot-unit installation. Existing guard recovery remains
offline and requires an existing journal; it never silently commissions a new one.
Previous-epoch host attachment audit
A fresh PalletHandoffSetOwner can call verifyPreviousAbsence({ journal, target, previous }) before it applies a current set. Supply the authenticated persisted
application reference, exact old prepared target and its prior applied receipt,
including an unfinished transition's complete retained membership. Native code
recomputes both old identities and their transition relationship, requires the
same boot and host namespace, and retains the shared host mutator lock. The
journal reference is an opaque binding; native code does not authenticate it.
Two complete host inventory passes reject old host or misplaced router names and MACs, Pallet markers, retained host indices and transit address/route conflicts. Old router indices remain scoped to the old namespace. The fresh router must independently remain empty and DOWN. Any candidate or inventory loss rejects the audit; it performs no cleanup, repair, deletion or adoption. Persist the returned old/current namespace-bound evidence before permitting a new epoch's effects.
The evidence proves old host attachments absent under exclusive privileged
ownership. It does not prove physical peer destruction, packet drain, allocation
reuse or a current nativeBarrier. Kernel peer teardown can be asynchronous, and
another privileged writer can invalidate negative observations. Full node
recovery uses the application owner's journal and complete authenticated history;
process death or a saved namespace identity cannot replace this audit.
The isolated Node/Deno fixture covers surviving and renamed old attachments,
reused host indices, unrelated host indices matching old router indices, transit
conflicts, fresh-router state, old request tampering and wrapper lifecycle.
verifyPreviousBootFence(request) is the separate fresh-owner capability for a
different kernel boot. It validates the exact historical target and prior receipt,
requires different valid kernel boot UUIDs, and reads the current boot from the
native owner. Two complete audits still reject current old identities, Pallet
markers, transit conflicts and nonempty fresh-router state. Old interface indices
are deliberately ignored: those numbers can legitimately belong to unrelated
interfaces after reboot. The returned evidence binds both epochs and adds
previousBootTerminated: true. It supplies no allocation release, packet drain
or nativeBarrier; the dormant-only history and privileged ownership requirements
still apply.
The offline test/native/qualify-handoff-boot.py fixture runs two boots each under
Node and compiled Deno with independent disposable ext4 NoSQL roots. The actual
application owner persists real signed authority and applied/pending history. A
fixture-only interrupted completion retains the acknowledged DOWN target, then
the fixture flushes the database
then signals init through a FIFO while the native owner and DOWN pair remain live.
Init forces poweroff. The next boot reloads that journal, qualifies reused host
indices and conflict rejection, then the application owner persists the boot fence
and realizes the same still-reserved lease DOWN. All application state uses SmartData/NoSQLDB; no journal
sidecar or host network/disk attachment participates in the qualification.
It also retains the exact unavailable report across poweroff, replays its original
reporter binding, and queues the next boot's report only after acknowledging history.
Standalone control build
The control builder requires Deno 2.9.7 and the exact NoSQLDB 10.5.1 engine identities
in binary/control-build.json. After the normal dependency install, build either
Linux target or omit the target to build both:
pnpm run build:control linux-amd64
pnpm run build:control linux-arm64
@git.zone/tsdeno 2.0.3 derives a temporary frozen Deno lock from the production
tree of pnpm-lock.yaml under its managed runtime-only manifest, and per target leaves out the npm native binaries built for the other
architecture or for macOS (their ELF/Mach-O headers prove them foreign); the control never loads them,
because it runs its own sibling engine, guard and DNS executables. TsDeno refuses any Deno other than the
pinned denoVersion, fails a pallet-control smaller than the target's minSize in
binary/control-build.json (200 MiB; a compile that lost its npm payload still exits 0 with a far
smaller binary), and on the host's own architecture runs it once without arguments in an empty
environment, where it must print its invalid_arguments refusal and exit 1; the other
architecture's smoke check is skipped. The build verifies the selected NoSQLDB engine's
bytes and clean owning-build provenance. It also builds and verifies the Pallet
static-musl executor against this exact source, version and architecture. The
Smartnftables 4.6.0 guard is verified against its pinned bytes and clean provenance.
Each completed dist_control/<target>-<digest> directory contains pallet-control,
its fixed siblings pallet-smartdb, pallet-runtime, pallet-guard, pallet-dns
and pallet-vpn, the pallet-containerd/ release directory (containerd, its shim,
runc, the pause archive and their manifest), their native provenance records and a
control-build.json recording artifact hashes, source state,
compiler versions and runtime lock hash. The compiled control keeps UID 0 and its
protected production data/socket paths; it derives only the sibling engine path
from its own executable. The installer owns protected platform-parent directories.
Packaging includes every native guard notice from the pinned upstream manifest
(the native-notices/manifest.json that tsrust notices generates, with the digest of every file)
under notices/smartnftables/ and rejects changed or missing material. The archive
also includes the exact pallet-clock/ loader, programs, shared libraries and
trust files, plus full clock source archives, patches, recipes and notices under
sources/clock/ and notices/clock/. Consumers must verify the complete inventory.
binary/vpn-engine.json pins the published SmartVPN 2.5.0 musl executable,
provenance, package license and native notice inventory for each architecture.
The builder and packer verify the complete distributed VPN notice set under
notices/smartvpn/, including the upstream Rust and Cargo license texts. A changed
binary, symlink, missing notice or mismatched build is rejected. The bundled
pallet-vpn is the executable runtime-serve verifies and runs as the managed VPN
device owner (see "Foreground node process").
The two engines Pallet compiles itself, Chrony (chronyd, chronyc) and runc with
the pause executable, are pinned by their outputs as well as their inputs:
binary/clock-engine.json and binary/runc-engine.json record each output's SHA-256
and size per architecture. The rule is that released engine bytes always equal those
pins. A control build reuses a verified earlier output under .nogit/clock-engine/
or .nogit/runc-engine/, or else reads the pinned outputs from the build-input store
(see "Build input store") and verifies them against the pins. Only when the store
holds none of a build's outputs does it build them from the pinned sources, and that
build also fails unless every output reproduces its pin byte for byte. A store that
holds only some of one build's outputs, a transport failure that outlasts the
retries, or a byte that differs from its pin is refused by name; none of them falls
back to a source build.
The source build is proven on every pin change: a changed recipe, SDK or source pin changes the output pins, and the maintainer runs the qualification below, which builds the engines from source and fails unless every output reproduces its new pin, before seeding the store with those outputs (see "Build input store"). Until the store holds them, a control build finds nothing under the new addresses and compiles from source itself. The qualification is also run on demand:
pnpm run engines:qualify # both architectures
pnpm run engines:qualify linux-arm64 # one architecture
engines:qualify ignores the store and any earlier output, builds Chrony and runc/pause
from their pinned sources for each selected architecture, fails unless every output
matches its pin, and prints one JSON report of the reproduced outputs with the time
each build took. The linux-arm64 builds run under emulation and take minutes. The
verified output it leaves in .nogit/, beside the sha256.txt the build recorded inside
its SDK, is the only kind of engine output the store seed uploads.
A source build additionally requires a Docker client/Buildx and a Linux Docker daemon able to execute the selected pinned Alpine architecture. Its build context and result archive use the daemon API; it requires no host workspace bind mounts. Compiler execution is offline, unprivileged and capability-free. The build checks pinned source, SDK and output hashes; workers receive only the finished programs and distribution materials and perform no package installation.
pnpm run release:control runs build:control and reads only its result lines from
stdout. Every other stdout line (tsbuild, tsrust) reaches the job log on stderr as it
arrives, prefixed with its ISO 8601 arrival time, so the log shows where a build's time
goes; the engines report on stderr whether they came from an earlier output, the store
or a source build, and how long that took.
A release build is reproducible or it is not released. release:control runs only on the
exact clean release tag and sets SOURCE_DATE_EPOCH to the tagged commit's committer time
(a job that sets a different value is refused), so tsrust stamps that time into the
executor's provenance and tsdeno compiles the Deno binaries against a private cache with
every embedded file at that time. It then runs the whole build:control a second time and
refuses to release unless both builds name the same output directories and produce an
identical file tree (paths, bytes and modes); the failure names every differing path. The
second build reuses Cargo's warm target directory, so it proves the TypeScript, provenance
and Deno outputs rather than a Rust rebuild on a fresh machine. Ordinary build:control
runs leave SOURCE_DATE_EPOCH unset.
Every runtime dependency change must update pnpm-lock.yaml with pnpm install.
TsDeno 2 derives .tsdeno.deno.lock, pins versions and dependency edges to pnpm's
production tree, verifies the compiled graph against that tree and removes the
temporary lock after compilation. Do not pass --lock or --frozen to tsdeno;
there is no committed deno.lock for this pnpm-owned graph.
Carry noticeLockSha256, the control notices and their pinned asset hash in
binary/control-build.json onto the new pnpm lock. test/test.controllock.node.ts
checks the frozen pnpm manifest/lock pairing in a private copy, the identity and
asset digests, registry origins and notice coverage for every production package.
The build record's lockSha256 also names pnpm-lock.yaml. The tracked .npmrc
sets registry=https://registry.npmjs.org/; scripts/control-lock.mjs rejects
another NPM_CONFIG_REGISTRY origin and production tarballs from another registry.
The earlier Linux amd64 artifact with SmartDB 5.8.0 was qualified as root in an isolated QEMU guest with a disposable local ext4 disk and no network or host mounts. Checks cover the fixed production paths, engine digest and permission guards, explicit provisioning, durable enrollment replay after process restart, EOF/signals, native and parent SIGKILL, and stale-socket recovery. NoSQLDB rejects volatile filesystems; a tmpfs data root cannot substitute for supported local storage. This is process recovery qualification, not a power-loss or ARM64 runtime claim.
The NoSQLDB 10.5.1 engine passes the local native identity persistence and
lifecycle suite with SmartData 11.14.2. Pallet now reaches it through the nosqldb family of
@lossless.org/client 1.5.1, which continues SmartData 11.14.2 with the same persisted format. Both engine targets are the published static musl executables;
the Deno control executable still targets GNU Linux. Both control targets compile
and package with their verified owner provenance and complete upstream notices.
The updated amd64 control bundle also passes the root guest checks above on an
isolated local ext4 disk, including durable enrollment replay after native and
parent process crashes. ARM64 runtime and power-loss qualification remain open.
test/native/qualify-node-process.py also qualifies a forward upgrade when given
both --previous-control-directory and --previous-probe. It boots that earlier
bundle first, retains a read-only copy of the stopped ext4 disk, then boots the
new bundle twice against the working disk. Each boot verifies its exact binaries
and probe; the final result compares the node identity and signed network
projection and verifies the backup is unchanged. The NoSQLDB 9.0.0 to 10.2.0
amd64 run passes all three offline root boots, including retained workload leases,
signature replay, namespace keeper loss and joined process shutdown. This
qualifies existing current-format Pallet state; it does not qualify containerd,
private DNS, physical power loss or foreign legacy database roots.
The guest stages the build's complete runtime inventory (executor, guard, resolver,
clock loader and assets) and the nftables and veth modules, enrolls with a routing
binding and commissions the allocation-pool guard before runtime-serve, which
now owns the node's network; the guest provisions the root-only runtime directory
on every boot, as the installer does. A runtime-serve whose namespace keeper is
killed reports failed and exits: when its close cannot confirm the network
release, the node process joins the clock, the namespace and the database anyway
(PalletNodeRuntime.abandon) instead of holding them for a retry nobody in the
exiting process makes, and the successor recovers the unconfirmed state by
evidence, as after a crash.
The database package is @lossless.org/nosqldb and the controller uses its
NoSqlDbServer API. Existing smartdb storage-directory and bundle field/file
identities remain stable; they do not select the deprecated npm package. Changing
the package does not convert an incompatible database format. NoSQLDB 10 reads
existing version-9 current-format stores directly; once it writes version-10
records, older engines cannot read them. Retain a stopped pre-upgrade backup
before activating the new engine. Native current-format admission still rejects
foreign legacy and mixed roots before application startup.
Pallet admits the released 32.0 attachment-journal shape before any network
owner starts. Terminal absence and an already durable pending removal retain
their original bytes, digests and ACTIVE source or amendment history; the latter
may finish with its released removal formula. The admission writes only the
pallet_migrations completion row and is replayable after interruption. A
released running or pending-ADD row has no durable removal transition that this
version can complete without rewriting its history, so startup fails with a
pallet-attachment-owner-recovery-* error and writes no completion row. Stop the
upgrade and retain the prior Pallet owner to finish that lifecycle before trying
again. There is no collection-reset or cleanup command in this recovery path.
Every node process owns its own router namespace, so a restarted process runs in
a new epoch, and so does every process after a reboot. A generation the previous
process retained — UP, or DOWN with a packet cleanup it never finished — is held
by no native of the new epoch, and the native admits RecoverActive only for a
generation of its own. The restarted process therefore fences it before its
native adopts anything, with the audits that fence the application journal's own
epoch, over the generation's application target and receipt: in the same boot,
verifyPreviousHandoffAbsence proves the previous epoch's host attachments
absent (the router namespace, and every pair, device and route under it, ended
with that process); after a reboot, verifyPreviousHandoffBootFence proves the
boot over. The proof is journaled as a DOWN of kind epoch (same-boot-host-absence
or new-boot-host-fence), taken from the fencing native's epoch
(PalletHandoffSetOwner.readEpochFence), and its packet cleanup releases what
outlived the epoch: in the same boot the host table, through the running engine
with the generation's retained host identity; the router table ended with its
namespace, and after a reboot no table is left. The successor is raised on a pair
keyed anew. A generation of the process's own epoch is still withdrawn in order
while the process holds it — a tunnel transition left pending by a failed step is
finished first — and recovered by RecoverActive otherwise. One the process
journaled but never raised — an activation that failed between its intent and UP,
such as a beginActivation that lost its source fence — is held by its native only
as a preparation, and that native, having adopted its set, admits no
RecoverActive: the process discards the preparation and journals the native's
discard as a DOWN of kind discarded (PalletHandoffSetOwner.readActiveDiscard),
then releases both tables as after any DOWN. A later process that finds such a DOWN
with its cleanup unfinished takes it again as a recovery. A DOWN with no cleanup
recorded at all — its owner ended after the DOWN and before the cleanup began, for
example while releasing the DOCKER-USER contribution — is fenced the same way
(fenceRetained): its tables are complete only for a successor that replaces
them in place in that epoch, and the successor of a fenced generation is always
keyed anew. In its own epoch it is recovered and its tables released; in a later
one it is proven ended and its cleanup recorded, with the host table released
within the boot and none left after a reboot. A journal row that
carries discarded is refused by Pallet 36.1.2 and earlier, so a node that wrote
one cannot be downgraded in place until two later generations were raised.
Every ACTIVE generation records the packet engine that prepared it, the SHA-256
of the Smartnftables executable the node process verified before it started
(packetEngine; generations journaled before 35.0.0 carry none, so their engine
is unknown). The kernel keeps a generation's host table across a Pallet restart
in the same boot, and the restarted process asks the running engine to release it
once the generation is fenced. An engine whose compiled graph differs from the
one that applied it cannot adopt it and answers Conflict; Smartnftables 3.0
changed the graph of every routerEgress table and of every scope with
publications. When that Conflict comes over a table another or an unrecorded
engine applied, the release fails by name (PalletPacketEngineChangedError,
conflict:packet.retainedByPreviousEngine) and the network owner stays failed
and fenced, because no retry of the running engine can change the answer. The
remedy is a reboot, which takes the kernel tables with it and lets the new-boot
proof fence the generation, or finishing the generation on the previous Pallet
before upgrading again. An upgrade that changes the packet engine across a
same-boot Pallet restart therefore needs a reboot whenever an ACTIVE generation is
retained. A Conflict under the engine that applied the tables is not this refusal
and fails as before.
Local build outputs are qualification assets; a dirty source marker is recorded
explicitly and cannot identify a release. The tag-triggered workflow publishes
clean-source control bundles as inputs for Spark integration.
The bundle's runtime-serve activates the node network with the bundled
pallet-vpn (see "Foreground node process"); running it still requires Spark's
whole-bundle installation and supervision. The complete activated path — ACTIVE
generation, managed VPN tunnel to a real cluster hub, lease renewal and workload
traffic — is qualified in parts (the guests below and unit specs), not yet end to
end on a real host. ARM64 runtime and workload execution are not yet qualified by
this component release.
Build input store
Every input the control build downloads is read from Pallet's build-input store, never
from where it was first published: the Chrony runtime and compiler-image APKs, the clock
and containerd corresponding sources, the runc compiler image and Go toolchain, the runc
sources and the containerd release archives. The store is the public Gitea generic
package pallet-build-inputs of the serve.zone organisation. Each input is one file,
addressed by its SHA-256 and the last path segment of its pin's url:
https://code.foss.global/api/packages/serve.zone/generic/pallet-build-inputs/<sha256>/<name>
The store also holds the pinned engine outputs (see "Standalone control build"), so a
control build needs no emulated compile while the engine pins are unchanged. An engine
output has no provenance URL; its file name states what it is,
<engine>-<version>-<target>-<output>, for example:
https://code.foss.global/api/packages/serve.zone/generic/pallet-build-inputs/<sha256>/chrony-4.8-servezone2-linux-arm64-chronyd
https://code.foss.global/api/packages/serve.zone/generic/pallet-build-inputs/<sha256>/runc-1.5.1-servezone1-linux-arm64-pause
A pin's url stays its provenance record and is never read by a build, test or release
step. Upstream locations do not last: Alpine's stable repositories keep only the newest
revision of a package, so a revision-pinned dl-cdn URL disappears with the next
security update, and Pallet 36.0.0's sealed release could not be built for that reason.
scripts/control-clock-inputs.mjs reads one store address per pin, verifies its size and
SHA-256, and caches it in .nogit/clock-inputs/<sha256>. An attempt must receive the
response headers within 10 s and a body byte at least every 30 s, and ends after 60 s plus
the pinned size at 256 KiB/s, so a slow but steady transfer of a large input completes and a
stalled one does not hang the build. Transport failures and these bounds are retried four
times; a store without the input fails at once with ClockInputStoreMissError, naming the
address and the seed command.
The store is public, so it holds no binary without its source.
binary/sdk-sources.json lists every package of both compiler images
(binary/clock-sdk.json, binary/runc-sdk.json) with the licenses its binaries
declare. For every package whose license expression names a GPL-family license it pins
the exact APKBUILD at the package's aports commit and every source that APKBUILD
checksums: aports inputs from https://github.com/alpinelinux/aports, archives from
Alpine's v3.23 distfiles, each matched against its APKBUILD SHA-512. The runtime
packages' sources are in binary/clock-sources.json and containerd's in
binary/containerd-sources.json; the store holds all of them.
test/test.buildinputs.node.ts refuses an SDK pin the inventory does not cover. The
qualification-only APKs in binary/clock-qualification.json are not build inputs and
are not in the store.
To refresh a pin:
- Change the pin (
url,sha256,size, and every digest that pins its file). For a compiler-image package, update itsbinary/sdk-sources.jsonentry and sources too. A change that moves an engine's outputs also changes their pins inbinary/clock-engine.jsonorbinary/runc-engine.json; runpnpm run engines:qualifybefore seeding: it builds them from source and proves the new pins. - Seed the store.
pnpm run inputs:store --dry-runreports, without uploading, every input and engine output the store lacks with its size and every one it cannot obtain. A maintainer with package write access to the serve.zone organisation then runspnpm run inputs:storewith the token inGITEA_PACKAGES_TOKEN. The tool takes input bytes from.nogit/clock-inputsand every--cache <directory>, which it only reads, or else from the provenance URL. Engine outputs have no provenance URL and are never taken from a cache: it takes them only from a source build under.nogit/clock-engine/and.nogit/runc-engine/of this checkout or of a--checkout <directory>(another checkout of the same pins, whose input cache it also reads), whosesha256.txt, written by the build inside its SDK, records the pinned digest. An output read from the store carries no such record; an engine output without a source build is reported unobtainable and the tool exits 1. It verifies every byte against its pin, uploads only what the store lacks and never replaces a store file. - Commit. A tagged release job reads only the store, and builds an engine from source only when the store holds none of its outputs.
TLS trust of the control executable
pallet-control verifies every TLS peer it dials — the controller or Cloudly origin
of the enrollment, runtime and network-acquire sessions — against the Mozilla root
bundle compiled into Deno and then the operating system trust store. A node whose
controller, relay or Cloudly certificate is issued by a private or internal certificate
authority trusts it the way every other service on the host does: install the CA
certificate into the OS store (update-ca-certificates, update-ca-trust). Pallet has
no trust setting of its own and never disables verification.
A compiled Deno executable carries no store selection and otherwise trusts the Mozilla
bundle alone, so runPalletControlCli sets DENO_TLS_CA_STORE=mozilla,system for
every mode, after the argument check and before any mode starts; Deno reads the variable
once, when the process opens its first TLS client. An operator who sets
DENO_TLS_CA_STORE explicitly keeps that choice, provided it lists only mozilla and
system; any other value, including an empty one, stops the process with one
failed line of reason invalid_ca_store on the mode's protocol and exit code 1,
rather than failing every later connection. Spark starts pallet-control with a cleared
environment, so its processes always run with the default.
The managed QUIC tunnel of pallet-vpn, a native SmartVPN executable, authenticates
the hub by the public key its credential names rather than by a certificate authority,
so no root store applies to it.
pnpm exec tstest test/test.tlstrust.node.ts --verbose --logfile --timeout 120
The test pins the default and the refusal, runs the CLI with an invalid store in three
modes, and, under the pinned Deno, dials a real HTTPS and WebSocket server whose
private CA is present only in the OS store (SSL_CERT_FILE): it is refused as
UnknownIssuer without the selection or with an explicit mozilla, and trusted with it.
Regenerating third-party notices
binary/control-third-party-notices.txt names the frozen npm graph, the Deno GNU
Linux runtime and the NoSQLDB engine. When the Deno pin moves, move every pin
first — denoVersion in binary/control-build.json, the sealed release's
denoVersion in .smartconfig.json and both Deno lines of
.gitea/workflows/release.yml — and add the new version's release commit and
the SHA-256 of its deno_src.tar.gz release asset (GitHub states it as the
asset digest) to binary/control-deno-sources.json. Then, with that Deno and
cargo-about 0.9.1 on PATH:
pnpm run notices:control # rewrite the notices and their pins
pnpm run notices:control --check # regenerate, compare, write nothing
scripts/notices-control.mjs refuses unless every Deno pin, the running Deno
and the pinned source agree. It downloads the pinned source of the Deno the
notices name and of the pinned Deno, checks both against their digests, and
takes the runtime's crate inventory with cargo-about over cli/rt for the
control targets (--locked, so Cargo fetches the locked crates). It needs
tar and network access to GitHub and crates.io, and works in a private
temporary directory it removes afterwards.
The regeneration carries the Deno runtime part and leaves everything else
byte-identical. A registry crate the notices already name keeps its entry;
Deno's workspace crates and every new crate are written from the crate's own
legal files, the workspace crates with Deno's LICENSE.md. A new crate without
a legal file is refused, because it needs a reviewed entry. The native V8,
Chromium Rust and Rust standard library notices are carried only while the
pinned Deno runs the V8 and the Rust toolchain they name; otherwise the tool
refuses until they are refreshed. The v8 crate cites rusty_v8's license at
the crate's own commit. The embedded JavaScript keeps its file list: each
listed file's legal comments are read again at the new release, each excerpt
is found again at its new lines, and a changed or new source file under ext/,
runtime/js/ or libs/core/ whose legal comments are not listed is refused
for review. The asset hash and noticeBuildIdentity.denoVersion in
binary/control-build.json follow the written file. --input <file> carries
another notices file instead of the committed one: from the Pallet 34.0.0
notices (Deno 2.9.4) the tool reproduces the committed Deno 2.9.7 notices byte
for byte. test/test.controldenonotices.node.ts covers the pin refusals and a
dry run over a small fixture without cargo-about.
Control bundle packaging
@git.zone/tspack owns archive assembly, file hashes, executable modes, sealed
manifests and complete archive verification. Pallet's adapter selects explicit
control builds and checks their source, version, compiler, runtime lock, fixed
paths and both native owners' provenance before passing inputs to TsPack. The selected
license notices are pinned in binary/control-build.json to the runtime lock,
Deno version and NoSQLDB owner build. The pinned native notice inventory binds
the Cargo/compiler inputs and every reviewed notice; the archive preserves the
native notice index, complete inventory and CRI provenance.
After sealing, the adapter extracts every bundle again from the sealed set and runs
the verifier of @serve.zone/pallet-bundle over it with the build's own source
identity, so a bundle that disagrees with what that package verifies for its
consumers is never produced. Its paths and control record keys come from the same
package (ts_bundle/inventory.ts). pack:control first compiles that package
(tsbuild custom ts_bundle to dist_ts_bundle/, output on stderr) from the
checked-out source, so it runs on a clean checkout and never loads a stale verifier;
release:control loads the adapter only after build:control has compiled it.
# Use the exact project-relative directories printed by build:control.
pnpm run pack:control dist_control/linux-amd64-DIGESTPREFIX dist_control/linux-arm64-DIGESTPREFIX
# A release requires both architectures built from the clean vVERSION tag.
pnpm run pack:control --release dist_control/linux-amd64-DIGESTPREFIX dist_control/linux-arm64-DIGESTPREFIX
Ordinary packaging produces an explicitly unpublishable qualification set. A
release produces the sealed set under dist_control_release/ and an
inputs-VERSION-COMMITPREFIX.json description that records its exact packaging
configuration and manifest digest. The output JSON identifies both paths.
A publishing workflow must retain the entire dist_control_release/ directory,
including that input description and sealed set, before publishing any asset.
Restore those original files on retry, then run:
pnpm run pack:control --release --reuse
Reuse verifies the retained source, configuration, manifest and every archive. It works without the original compiler outputs and rejects missing, corrupt or changed retained inputs. It never recompiles a control bundle: the only thing it compiles is the TypeScript of the bundle verifier above. Fresh Deno compilation may produce different bytes, so rebuilding cannot substitute for retaining a published set. TsPack and this adapter do not provide remote CI artifact storage or publication.
GitZone 6.6.2 or later manages .gitea/workflows/release.yml and the
scripts/gitzone-*.mjs release scripts through the committed tspackRelease
asset configuration. The tag-triggered workflow invokes the configured
scripts/prepare-control-release.sh command in its disposable job container.
It requires a root Linux Gitea Actions environment, installs Pallet's clang and
musl-tools compiler prerequisites through apt, then executes the existing
scripts/release-control.mjs exact-tag adapter to build and package its two
explicit output directories. Compiler installation writes only to stderr so the
adapter retains its single JSON result on stdout. The runner must be Gitea Runner 3.3.2 or
later with its cache/results service reachable from job containers and
runner.patch_actions enabled. The pinned stock artifact actions use the v4
protocol required by Gitea's REST retention inventory.
It retains the complete output root in Gitea Actions for 90 days, downloads that
retained copy, and verifies component reuse and every archive before creating a
release draft. The publisher checks existing attachment bytes, adds only missing
files, and publishes after complete readback. It never replaces release assets.
After a CI interruption, rerun the original Gitea run so it restores the
original bytes. Missing or expired retention and conflicting remote attachments
stop publication. The generic release receipt gitzone-release.json and Pallet's
input description must remain beside the sealed set in the retained artifact.
Use gitzone format --only assets --write --yes to update these managed scripts
from a published GitZone version, then commit the result before releasing.
Build and test from source
This development package is marked private and is never published. The npm
release target publishes one module of this repository, ts_bundle/, as
@serve.zone/pallet-bundle with Pallet's release version (see its own
ts_bundle/readme.md). Its tspublish.json takes the @git.zone/tspack and
@push.rocks/smartdaemon ranges from the root devDependencies
(devDependencyVersions), so neither enters the root dependencies that
pnpm-lock.yaml supplies to the control build, and
@serve.zone/interfaces from the root dependencies. From the repository:
pnpm install
pnpm test
pnpm build
pnpm test is a plain tstest --verbose --logfile; the test set lives in the
@git.zone/tstest section of .smartconfig.json. Its prepare.once builds the native
debug binary (cargo build --locked) and warms the control-pack inputs, then the files
under test/ run six at a time (concurrency), longest first by the durations in
tstest's run ledger, each in its own process with its own 180 s timeout, and the
Rust tests run as the suite rust (cargo test --locked) beside them. Runs are recorded
in the ledger; the repository never reuses a passed run (reusePassedRun: false). A
test file is split before it needs more than two-thirds of its timeout at that
concurrency, so contention cannot push it over.
Tests exercise the native debug binary against a disposable Unix-socket HTTP/2 fake
CRI server. They do not connect to Docker or a production daemon.
The native build uses Rust 1.95.0, locked Cargo dependencies and pinned vendored protoc to
produce static musl Linux binaries through @git.zone/tsrust, named
pallet_linux_amd64_musl and pallet_linux_arm64_musl. pnpm build builds only the
host architecture's executor; pnpm run build:control builds both with
tsrust --configured-targets, so a control build and the release always carry the
executors of both targets built from the same source. ARM64 cross-compilation
uses aarch64-linux-gnu-gcc as the linker driver with Rust's self-contained musl
target libraries. The ring TLS provider also requires musl-gcc for amd64 C/assembly
and Clang for its supported freestanding arm64-musl C build. These compiler choices
are declared in rust/.cargo/config.toml; provision them on developer and release
builders before running the build. Other packaged platforms are rejected explicitly.
After pnpm build, run PALLET_TEST_PACKAGED=1 pnpm exec tstest test/ --verbose --logfile --timeout 60 to exercise the host architecture's packaged executable and its
default lookup instead of the native debug executable. Cross-compiling an ARM64
artifact does not establish ARM64 runtime qualification.
For explicit emulator qualification, the same suite accepts an absolute
PALLET_TEST_BINARY_PATH pointing to a test-only emulator launcher. Do not combine
it with PALLET_TEST_PACKAGED; an emulated run does not qualify real node hardware.
The third-party notice index points to the native
notices tsrust notices generates in native-notices/ for the locked Cargo graph,
the Rust toolchain and the static runtime, with full license texts and the reviewed
once_cell licenses and source attributions of ring under
native-notices/crates/ring-0.17.14/extra/, and lists the material kept beside them. npm publication and
distribution of the standalone Rust runtime-probe binaries remain disabled.
Gitea releases distribute the control bundles described above, whose
runtime-serve activates the node network as described in "Foreground node
process".
Read-only probe
After a build, from the repository:
import { Pallet } from './dist_ts/index.js';
const pallet = new Pallet({
socketPath: '/run/pallet/containerd/containerd.sock',
});
const evidence = await pallet.probeRuntime({ timeoutMs: 5000 });
console.log(evidence);
// {
// runtimeName: 'containerd',
// runtimeVersion: '<actual daemon version>',
// runtimeApiVersion: 'v1',
// runtimeReady: true | false,
// networkReady: true | false,
// }
The socket must already exist and belong to a separately configured, authorized
containerd instance. Pallet never creates one, chooses a default socket, searches
PATH for its binary, or falls back to Docker. An explicit absolute binaryPath
can select a verified installed executable or the development debug binary. A
missing or nonexecutable selected binary fails with SmartRust 2's
RustBinaryLocatorError (ERR_RUST_BINARY_EXPLICIT_PATH_INVALID). Selection never
changes executable permissions or falls back to another packaged binary or a
stale GNU/native build. After explicitly provisioning the selected executable,
the same Pallet instance can retry.
Each probe owns a separate native child. Only one probe per Pallet instance is
admitted at a time. The method confirms that child's exit before returning,
including on failure. If cleanup cannot be confirmed, the method rejects and
retains ownership of the child. Further probes are blocked until await pallet.close()
successfully retries cleanup. An optional AbortSignal cancels the local request and
terminates the owned child; it never stops containerd itself.
The native deadline covers connection and both RPCs (50–30000 ms, default 5000).
The bridge readiness handshake, after executable discovery, is bounded to
3 seconds, with 1 second for graceful termination before forced child shutdown.
Executable filesystem discovery in the shared bridge is not currently timed or
abortable; this is not an end-to-end startup deadline. IPC and decoded gRPC responses are
limited to 16 KiB. The probe accepts only containerd with CRI API v1.
RuntimeReady and NetworkReady must each occur exactly once; a missing or
duplicate condition is an error, while explicit false remains false. Runtime
diagnostic messages and verbose configuration are never returned. Runtime
reported readiness is not workload health, version support qualification,
network reachability, storage fencing, or permission to perform a migration.
Native failures carry a RustBridgeRequestError.responseErrorCode of
INVALID_INPUT, SOCKET_UNAVAILABLE, RUNTIME_UNAVAILABLE,
DEADLINE_EXCEEDED, UNSUPPORTED_RUNTIME, or INVALID_EVIDENCE.
Bridge transport, cancellation, and startup failures remain distinct errors.
Isolation and qualification
Known Docker-private path components and aliases into them are rejected as accident prevention. This is not a security boundary against a hostile local user who can replace socket paths. Run only with the local permissions needed to inspect the intended daemon; containerd's socket is root-equivalent.
Fake-server tests establish protocol and lifecycle behavior only. Before runtime adoption, qualify a dedicated containerd 2.3 instance with explicitly owned root, state, socket and configuration paths, then verify actual container, network, storage and recovery behavior. No production cutover is implied.
Every QEMU qualification guest under test/native/ runs one bundled ES module.
scripts/guest-bundle.mjs builds it from the guest's own TypeScript source with
esbuild — the whole graph inlined, the NodeNext .js specifiers resolved to the
.ts sources in this checkout, bare Node builtins rewritten to their node:
form (which is what the deno compile --no-config --node-modules-dir=none half
of each guest needs) and a createRequire banner for the CommonJS dependencies:
node scripts/guest-bundle.mjs --entry test/native/namespace.ts --out .nogit/debug/guest/namespace.mjs
It prints the bundle's path, size and SHA-256, which is what a qualification run
records beside its other inputs. The runners take that file as --bundle (and
qualify-active-registry.py as --wrapper-bundle / --tunnel-bundle).
qualify-cni-up.py --scenario single-host runs test/native/singlehost.ts, a node
bound to a local Onebox controller, in the cni-up guest with no substituted seam.
network-acquire-local serves the contract's socket (a root-owned 0600 socket in
a root-only 0700 directory; a peer running as another user gets EACCES); the
first bind persists the binding, an exact replay answers the same credential, a
different binding refuses, a replayed bind's session supersedes the earlier
connection's, which can then no longer apply anything, and the signed onebox
projection is acquired over the current session. The node's runtime then raises one
generation without a managed VPN: the real barrier reaches forwarding on the
single-host fact with no tunnel command, execution admission is accepted on that
evidence, two workloads attached through the real CNI path are raised under the
generation, TCP and UDP flow between them with the client's own address preserved,
and each resolves the other's name to its lease address through the node-local
resolver. A third workload, on a network only the first shares, answers the first
and is dark from the second. The router namespace reads IPv4 forwarding off in
all and default before any generation exists, although the guest host forwards. From the host namespace it probes the granted TCP and UDP tuples, another
port, another protocol, another lease, another source address and a granted lease
that is not attached. The Cloudly counterpart — a generation without a tunnel
receipt settles at up and its barrier does not hold — stays proven by the default
cloudly scenario. With the generation's pool route in place the granted TCP and
UDP tuples reach the workload from the transit host address, and another port,
another protocol, another lease and another source address stay dark; a granted
lease that is not attached stays dark as well, stopped by the host-transit table
before the router receives a packet (both hops carry attached leases only). That
last check captures every frame on both ends of the transit link while it probes
(the guest stages tcpdump, without promiscuous mode, and parses its pcap) and
counts only IPv4 packets to the unattached lease: the link also carries the ARP
exchange both ends run on their own schedule — the router re-probes its entry for
the host about 5 s after the granted flows — which a frame counter cannot tell
from a leak. The same captures must see a granted flow on both ends, drop nothing
and account for every frame the link counters counted; with the host-transit
policy deliberately given the unattached lease's grant, the probe's SYN reaches
the router and the check fails. A
successor projection then withdraws every host selection under the held generation:
the pass withdraws that generation before it completes the changed dormant set,
raises the next generation on a re-keyed packet pair (forwarding, no host route
planned), and the formerly granted tuple goes dark. Probers dial every 100 ms
through the change: the allowed pair between the first two workloads, and six
tuples the policies deny — the pair that shares no network over TCP and UDP, and
workload to the host's transit and uplink addresses over TCP and UDP, each with a
live listener behind it. East-west traffic stops for the withdrawal window, because
the withdrawn generation restores a router that forwards nothing, and resumes when
the successor is UP (about 27 s of a 45 s change); every denied tuple is
dialled inside that window and never answers before, during or after it, and no
listener logs its source. Without the forwarding pin the same guest measures the
unfiltered router: the pair that shares no network answered 61 of 201 dials during
the change. A further projection that only adds a workload, F, then moves the
successor in place (the same generation, the successor's application carried); F
attaches through the real CNI path and is raised under it, answers A, every read of
the journal stands with F's amendment, and a fresh network owner over the same router
namespace fences the generation at its start and raises the next one on its first
pass. All of it passes on 6.18.35-0-virt. The packet hub guest (packet.ts) plans the same pool
route and proves it present only while its generation is UP, gone after a
withdrawal, and gone after a fresh owner's recovery of a lost native. It also proves
the host's forwarding: off with both exclusive tables applied and before the
generation is UP, on while it is UP, off after the withdrawal. A LAN peer behind
its own link on the guest host, routing through the host, dials a live TCP and
UDP listener on the external host while the generation is UP and the host
forwards: dark under the exclusive host table, answered (with the LAN source
preserved) once the host table is replaced one revision up by the same scope
without exclusiveForwarding, dark again when the member is restored, and dark
with no generation and no table at all; the published two-hop flows pass under
the exclusive tables.
qualify-cni.py --scenarios containerd-serve --control-directory dist_control/linux-amd64-…
runs Pallet's own containerd from a build:control directory the way its service
unit does, with a foreign /etc/containerd/conf.d drop-in planted in the guest. It
proves the generated configuration governs (the drop-in's stream port stays closed,
127.0.0.1:10010 serves), the bundled sandbox image is imported and unpacked on
overlayfs, a workload raised by the production execution owner through the real
CNI path runs under /pallet with runc state under /run/pallet/runc, SIGTERM stops
containerd and leaves the workload running, the next containerd-serve reattaches
the exact sandbox and container, and killing containerd fails its supervisor
(owner_failed) while the workload keeps running. The image collector keeps the
running workload's image and the bundled sandbox image, and the proven removal
frees the workload's image while the sandbox image stays.
qualify-cni.py --scenarios registry-auth --bundled-containerd-directory <…>/pallet-containerd
pulls through Pallet's bundled containerd release (the pallet-containerd directory a
control build or node test/helpers/controlpackinputs.mjs <root> <directory> assembles,
verified against binary/containerd-release.json) from a token-auth registry fixture on
guest loopback, named registry.pallet.test:5000 as a relay's host:port origin is: the
workload repository answers 401 with a Bearer challenge, and its token endpoint issues a
token only for the node's Basic credential. A refused credential reaches the token
endpoint, the run ends failed with cause unauthorized, the lane is free and the refusal
is written once; the node's credential is issued a token, the workload runs and is stopped
and removed, and no token request is ever anonymous. With the former bare host:port
server address the same guest sees containerd fetch an anonymous token and fail 401.
qualify-cni.py --scenarios owner-storage --control-directory dist_control/linux-amd64-…
runs one local storage claim under the same bundled containerd and runc, on node
and Deno. The claim materialises its volume root-private with the claim's
ownership; containerd echoes exactly one private read-write bind of it and the
bundled runc's workload carries it at the claimed path; the workload writes as
its own user; a second run of the service and a purge are refused while the
first container lives (storage.claim.held, storage.purge.held); a restarted
owner rejoins the live container against the grant's mount list; the second run
finds the first one's bytes after the first is removed; and the purge deletes the
volume once both are gone.
qualify-cni.py --scenarios owner-readiness owner-epoch runs test/native/cniepoch.ts
instead of cni.ts: bundle it with node scripts/guest-bundle.mjs --entry test/native/cniepoch.ts --out <guest>/cniepoch.mjs and deno compile --no-config --node-modules-dir=none -A, and pass them as --bundle and --deno-driver (the other
inputs are those of owner-attached). Both scenarios boot the attached guest — the pinned
official containerd, real CNI, the production network and execution owners, the QEMU
user-mode NAT for the NTS clock only, and the owner-attached scenario's test-only
activation barrier seam — once on node and once on Deno.
owner-readinessruns an attached workload whose readiness is TCP on a listener inside the sandbox. The host, a Cloudly-style node, has no route to the workload pool (its route to the workload's address leaves through the uplink) and its own connect times out, yet the workload becomesready: the native probe connects from inside the sandbox's journaled network namespace, and no sample is a namespace refusal.owner-epochstarts the node process in its own control group, runs an attached workload, and then kills the whole group (cgroup.kill), as a crashed service unit's restart does. The router namespace goes with it, and with it the sandbox'seth0, while CRI still reports the container running. A fresh node process on the same store and the same containerd reads the run's network epoch asended, records the runfailedwith causenetwork-epoch-ended, writesPallet run failed: …: network-epoch-ended., stops the sandbox (CRI then reports it exited), observes the failure once, and answers the controller's stop and removal with their chained terminal receipts; the removal's CNI DEL answers from the restart's epoch coverage of the old pair.
qualify-active-registry.py --paths-node-binary … --paths-bundle … (with the
packet and SmartVPN binaries) runs test/native/singlehostpaths.ts: the three
single-host paths of a Cloudly node whose platform services run on its own host.
The guest adds the hub's and the relay's addresses to the uplink as permanent
/32s before the uplink is observed, attaches and raises three workloads that
share no private network, and composes one projection with the production
composer twice: without hostPlatformEndpointIds and workloadIngress, and with
them. Under the first pair a workload's dial of the relay it selects, the router
namespace's dial of the hub and the ingress workload's TCP and UDP dials of its
target are all dark; the composed pair replaces it one revision up under the same
UP generation and all four are delivered — the relay and the hub see the transit
source address with a port of the leased range, the target sees the ingress
workload's own address — and the real SmartVPN hub on the host authenticates the
router namespace's real managed QUIC client. Another workload to the relay, the
Corestore port the workload did not select, an undeclared hub port, the host
opening toward the router, the target opening back toward the ingress workload,
another workload to the target and another target port stay dark under both
pairs, each against a live listener that logs no peer. Then the protected
authority takes its next step that only adds (a platform endpoint joins), and the
pair composed over the successor replaces the live one, one revision up, while one
TCP connection from the ingress workload to the relay keeps exchanging numbered
lines every 20 ms: every line sent across the replacement comes back on the same
connection, with no reset and no stall over 500 ms; both policies then carry the
successor's authority digest, fresh dials to the relay and the target are
delivered and the negatives stay dark.
qualify-active-registry.py --tunnel-node-binary … --tunnel-bundle … runs
test/native/activetunnel.ts, whose moves case raises the device under the real
SmartVPN 2.5.0 hub and lets the hub move the node's split route in place
(reconcileManagedNetwork at the next revision): the client moves the device's
route, the native keeps the generation UP with the moved route pending, the
revision the device owner reports is adopted (a replay answers the same adoption),
and both packet policies stay enforced. A route on another device of the router
namespace then fences the generation, and a fresh owner recovers it DOWN and
releases the live tunnel pair through a packet owner of the same identities.
The pinned upstream CRI protocol and the containerd Transfer and Streaming protocols,
with their Apache-2.0 attributions, are under rust/proto/.
License and Legal Information
This repository contains open-source code licensed under the MIT License. A copy of the license can be found in the repository license file. The vendored Kubernetes protocol is separately licensed under Apache-2.0.
Please note: The MIT License does not grant permission to use the trade names, trademarks, service marks, or product names of the project, except as required for reasonable and customary use in describing the origin of the work and reproducing the content of the NOTICE file.
Trademarks
This project is owned and maintained by Task Venture Capital GmbH. The names and logos associated with Task Venture Capital GmbH and any related products or services are trademarks of Task Venture Capital GmbH or third parties, and are not included within the scope of the MIT license granted herein.
Use of these trademarks must comply with Task Venture Capital GmbH's Trademark Guidelines or the guidelines of the respective third-party owners, and any usage must be approved in writing. Third-party trademarks used herein are the property of their respective owners and used only in a descriptive manner, e.g. for an implementation of an API or similar.
Company Information
Task Venture Capital GmbH
Registered at District Court Bremen HRB 35230 HB, Germany
For any legal inquiries or further information, please contact us via email at hello@task.vc.
By using this repository, you acknowledge that you have read this section, agree to comply with its terms, and understand that the licensing of the code does not imply endorsement by Task Venture Capital GmbH of any derivative works.