jkunz cf961a5028
Sealed release / build-and-release (push) Successful in 28m33s
v35.4.0
2026-09-30 20:45:44 +00:00
2026-09-30 20:45:44 +00:00
2026-09-30 20:45:44 +00:00
2026-09-30 20:45:44 +00:00
2026-09-30 20:45:44 +00:00

@serve.zone/pallet

Pallet is the in-development node-local containerd execution component for serve.zone. Its public API provides a bounded, read-only native CRI v1 probe. Backend-private identity, enrollment and stateless execution mechanisms are under development; production assignment delivery and reconciliation remain unwired.

Issue Reporting and Security

For reporting bugs, issues, or security vulnerabilities, please visit community.foss.global/. This is the central community hub for all issue reporting. Developers who sign and comply with our contribution agreement and go through identification can also get a code.foss.global/ account to submit Pull Requests directly.

Development status and ownership

Cloudly will own cross-node desired state and placement. Pallet will own local execution, including the outbound Cloudly connection on workers. Onebox will control its own local workloads through authenticated Pallet IPC, without depending on Cloudly to bootstrap itself. Spark remains the host installer and supervisor. Coreflow's required behavior must be ported before its legacy Swarm runtime can be retired.

Those orchestration and authorization capabilities are not implemented by the public probe. It exposes no public listener and persists no application state. Do not replace a production runtime with this development foundation.

No Pallet name carries a version token: collections, singleton row ids, record-id hash domains, the native structs and their TypeScript twins, the dormant link aliases (pallet-handoff: / pallet-workload:), the host and workload evidence kinds and the process protocol names are named by what they hold. No Pallet shape carries a shape-version field either — not the native structs, not the identity, application, guard, attachment, ACTIVE, barrier, tunnel or DNS records, and not the build manifests this repository owns. Every one of those key sets is exact, so a value that still states schemaVersion is refused as an unknown key rather than read as an older shape. What contract a peer speaks is stated once, by the handshake below, and what a persisted shape means is stated by the installed build. No released Pallet version was ever deployed, so nothing is migrated and nothing is aliased — a store written by an older build fails its assertion on every read, a dormant link an older build left behind carries a foreign alias, interface name and MAC address because the seed domain changed with the name, so it is never recognized as this build's pair, and a node's stores under /var/lib/serve.zone/pallet are reset with the install.

Private CRI execution

ts/runtime/classes.executiondriver.ts owns a bounded native CRI client through the separate --executor-management process mode. It is absent from the public package facade and enrollment IPC. The eventual assignment owner must authenticate the controller, validate immutable material, durably admit the assignment and serialize its effects before invoking this mechanism.

The driver supports inspect, run, stop and explicit runtime removal against a dedicated containerd 2.3 release and CRI v1 socket. Run supplies a digest-pinned image, its platform image-config digest, exact argv, environment, working directory, CPU/memory policy, user/group policy and root filesystem write policy. Explicit cpuMillis: null leaves CPU quota unlimited. Explicit null for both runAsUser and runAsGroup selects the pinned image's User through containerd's rootfs resolution; a numeric pair overrides it. Missing fields and partial pairs are invalid. Containers use private namespaces, runtime-default seccomp and no-new-privileges, without privileged mode or host networking; the only host paths a container ever receives are the read-only per-file binds of its own staged secret mount and one private bind per node-local volume it was granted ("Local storage claims"), which the executor derives from its fixed volume root and the volume's key — no host path crosses the IPC. Every sandbox names the absolute cgroup parent /pallet, so workload cgroups sit outside any service unit's cgroup (see Pallet's own containerd). Network policy and application readiness remain separate owners. Secret material delivery is this node's own, through the staged mount.

Every runtime object carries the immutable generation-one run digest and complete controller/node/assignment/replica/attempt ownership labels. Recovery requires one exact sandbox and container, including a full check for foreign sandbox children. Image identity uses the actual config digest and the requested repository digest; display tags are not identity. Running attempts can be recovered; exited or stopped attempts are never restarted. Stop preserves runtime objects, including containers created but never started. Removal requires confirmed stopped state and retains storage; images are left to the image collector below. Neither operation depends on a healthy image cache or network.

Input is captured before asynchronous work without invoking accessors. Native frames and CRI responses have fixed bounds; errors expose static codes rather than runtime messages or registry credentials. Cancellation, parent EOF and signals terminate the owned client. A timeout or client exit does not cancel a server-side CRI operation, prove rollback or fence a writer. The durable assignment owner must retain uncertain effects and resolve them before admitting replacement.

pnpm exec tstest test/test.execution.node.ts --verbose --logfile --timeout 60

The Unix fixture covers exact execution and recovery, lost mutation replies, foreign/ambiguous resources, image mismatch, degraded shutdown, never-started containers, malformed IPC, parent death and cancellation. These mechanism tests do not establish durable assignment admission or authorize a production cutover.

Image collection

Pallet's containerd holds only the images Pallet pulled for its own runs and the bundled sandbox image. PalletExecutionOwner.collectImages removes the images no durable assignment still needs. A collection is due when a lifetime starts and after every removal the owner proves. The node runtime runs it at the end of a full assignment sweep, on the same native lane as a reconcile, so nothing pulls or creates a container meanwhile.

The owner states what to keep. Every assignment whose runtime removal is not proven keeps its run's image config (platformEvidence.imageConfigDigest), whatever its disposition: a stopped attempt is still inspected and removed with its image, and an admitted one still pulls it. If any retained assignment's image cannot be named, nothing is collected. The native pass (collectExecutionImages, rust/src/execution.images.rs) also keeps:

  • every image a CRI container still uses, in any state and matched by any name the image carries;
  • every image the runtime reports pinned;
  • the sandbox image by its bundled name.

It removes the rest in id order with CRI RemoveImage and requires ImageStatus absence afterwards; an acknowledged removal that left the image in place fails the pass as RUNTIME_STATE_CONFLICT. It reports the images listed, removed with their size, kept by reason, left for the next pass, and the bytes the image filesystem reports in use.

Bound Value
retained image configs per pass 1,024
images the runtime may hold when a pass begins 1,024 (refused as RUNTIME_STATE_CONFLICT beyond)
images removed per pass 64; the rest are remaining and keep the collection due

The node runtime runs the collector at the end of a full assignment sweep, and only while the node holds an authenticated controller session: before the first one a node admits no native effect, so a due collection makes no CRI call and runs on the first sweep after the session is established. The collector runs only while the durable execution lane is idle on this boot, because an operation still pending there may be a run whose pull the runtime is finishing. A pass that cannot run stays due and is tried again after at most one minute (palletImageCollectionRetryMs). It is recorded on the node status images as one of:

  • deferred with lane-pending;
  • failed with retain-unresolved, retain-bound, or runtime plus the native code.

A lane or assignment store that cannot be read is recorded the same way: an unreadable lane is not proven idle and is deferred with lane-pending, an unreadable retained assignment is failed with retain-unresolved. Besides its own refusals (no controller session, a busy lane, an owner not ready), the collector rejects only when the identity runtime refuses its assignment store or its own code fails, a defect the status cannot name. The node writes Pallet image collection failed unexpectedly: <class> <code>[:<site>]. to stderr, which Spark forwards to the node journal, once per distinct failure until a collection runs again.

It never fails the node: an image left behind costs disk, not correctness. No journal is kept: removing an unreferenced image is idempotent, and a later run pulls by digest.

pnpm exec tstest test/test.imagecollection.node.ts --verbose --logfile --timeout 120

Workload stats

A controller reads one resource sample of a running attempt with readRuntimeAssignmentStats (@serve.zone/interfaces 32.36.0). The request names the exact current revision of an assignment on this node, and the controller client binds it to the session and that revision before any native call; like every inbound request, it acts only on the connection that holds the session.

PalletExecutionOwner.readStats reads outside the native lane, so a read never waits for or blocks a reconcile. Each read runs its own native child (sampleExecutionStats, rust/src/execution.stats.rs): it proves the attempt's exact CRI identity, calls CRI ContainerStats and PodSandboxStats, and proves the same identity again. A change between the two proofs is a conflict, never a sample of whatever runs now. An attempt without a running container is answered not-running, without a runtime stats call.

A sample states the cumulative CPU time across every core (decimal digits, because it outgrows a JSON number; utilisation comes from two samples), the memory working set, the memory limit the node enforces (the config's memoryBytes), the sandbox's default interface counters (or null) and the writable-layer usage (or null). A counter a JSON number cannot carry exactly is refused as RUNTIME_STATE_CONFLICT.

Bound Value
reads per assignment at a time 1 (another is refused as busy)
reads across the node at a time 8 (palletObservationBounds)
native deadline per read 10 s
pnpm exec tstest test/test.executionstats.node.ts --verbose --logfile --timeout 60

Workload logs

Every attempt's stdout and stderr are captured by the runtime and delivered to the controller, which keeps them (IRuntimeConfig.logs: SmartData owns log metadata, SmartBucket the payloads). The node keeps no log history of its own: its capture files are disposable transport artifacts on tmpfs.

Capture. The sandbox names the attempt's capture directory /run/serve.zone/pallet/logs/<run digest hex> and the container its file workload.log below it; containerd writes the CRI log format there (<RFC 3339 time> <stdout|stderr> <P|F> <content>). An attempt whose container was created without a log path (a container from an earlier Pallet) is stated once per process as a not-captured loss.

Delivery. PalletLogCapture reads the file and sends reportRuntimeAssignmentLogs batches (runtimeAssignmentLogContract, @serve.zone/interfaces 32.36.0) on the session, one at a time: the next batch leaves only after the controller acknowledged the previous one, and an unacknowledged batch is sent again unchanged. A CRI piece longer than one entry (64 KiB) travels as partial entries; a line longer than the config's logs.maximumLineBytes is cut there and followed by a line-cut loss of its stream, counting the cut bytes. Every delivery pass (palletLogDeliveryBounds) sends at most eight batches per capture before the next capture's turn.

Controllers that keep no logs. A controller answers every batch accepted, replay or not-kept. not-kept states that it keeps no workload logs and holds nothing: the node sends that session no more batches (runtimeAssignmentLogContract.notKept, @serve.zone/interfaces 32.37.0). The answer is also the node's permission to discard for that session (@serve.zone/interfaces 32.38.0), and Pallet uses it while that session stays the current, live one: every flush and every hold on the delivery cadence discards each ended capture whole — its directory, its cursor and its waiting batch — and each capture that ends later, and has each running capture drop its output instead of holding it. A running capture drops every unread whole CRI line and a batch it sent that holds output; it records the dropped batch's number with its cursor, so that number is never sent again, not even by a restarted process, and its next batch opens with one not-kept loss counting the dropped bytes and whole lines (loss evidence the dropped batch carried leads it too). A sent batch that holds only loss evidence waits unchanged. The next session is asked again, and a controller that keeps logs receives each running capture from where it stands, after that loss. Any other failure of a batch keeps it the capture's next one and sends it again unchanged on the next pass.

Ended captures. A capture whose container ended waits for its final batch to be acknowledged, and a crash-looping workload ends one per restart, so the node keeps at most runtimeAssignmentLogContract.maximumEndedCaptures (64) ended captures awaiting acknowledgement, holding at most maximumEndedCaptureBytes (64 MiB) together, each measured as the buffer bound measures it: in the file bytes it holds for delivery (Bounds below). Each pass holds every ended capture within its own buffer bound and rewrites its files to what it still needs before it is measured, so a capture at the largest buffer (64 MiB) always fits whole. Past either bound — a controller that is down, fails every batch or leaves the method unhandled — it discards the oldest ended captures whole, oldest by when this process saw them end. Every capture discarded whole, under this bound or a not-kept answer, is counted in the discardedCaptures of the next batch the node creates, of any capture. The count is persisted (pallet_log_discards): it survives a restart, travels unchanged with the batch that carries it on every retry, a restarted process creating that batch again with the same number and count, passes on to the next batch when that batch is discarded unacknowledged (dropped under not-kept, discarded with its capture, or lost to a restart that cannot continue its capture), and is kept while a not-kept answer stands. The owner keeps in memory only running captures and ended ones within the bound: a capture that finished or was discarded leaves no directory, cursor or entry behind.

Rollout order. A controller must answer reportRuntimeAssignmentLogs before a Pallet with log capture runs against it. A controller that leaves the method unhandled fails every batch: the node stays within its bounds, but it retries the same batch on every flush and drops the output beyond the buffer bound, so no workload output reaches such a controller. Cloudly answers not-kept (since 33.8.0). Onebox has no Pallet log backend yet (Onebox 33.1.1): run this Pallet against Onebox only once one ships.

A controller must also run @serve.zone/interfaces 32.38.0 or later before this Pallet runs against it. A receiver on 32.37.0 or earlier reads a batch by its exact schema and refuses one that carries discardedCaptures or a not-kept loss, on every retry: against it the node stays within its bounds as against any failing controller, but that capture's output no longer arrives. Cloudly 33.8.1 is the first Cloudly on 32.38.0 (33.8.0 is on 32.37.0); an Onebox Pallet log backend must be on 32.38.0 when it ships.

Bounds. What the controller has not acknowledged is the capture's buffer, at most the config's logs.maximumBufferedBytes. Pallet measures it in the file bytes the capture holds for delivery, which is what sits on the tmpfs: the file bytes the batch waiting for its acknowledgement spans, plus the file bytes not yet read into a batch, CRI headers and the cut rests of lines beyond logs.maximumLineBytes included. The waiting batch counts with its whole span whether or not a reclaiming pass (below) already removed those bytes from disk. A batch never spans more file bytes than the bound (its first CRI line is always taken), so the waiting batch is always kept and sent again unchanged; beyond the bound the oldest unread whole CRI lines are dropped and stated where they were dropped as a buffer-overflow loss, which counts their exact output bytes and whole lines and leads the next batch. Output of short or cut lines therefore reaches the bound sooner than its output bytes alone would: a line cut to a few bytes still counts every byte the runtime wrote for it. Drops while a batch waits extend that one loss, which keeps the time the loss began, so however long the controller stays away the evidence is a single entry. Once the node has read past the buffer bound (never below 1 MiB) it renames the file aside and has the runtime reopen the log path (CRI ReopenContainerLog through the native reopenExecutionLog, which proves the attempt's exact identity before and after); the renamed file is drained first and removed once the cursor moved past it. On disk an attempt therefore holds at most the read part below the rotation bound plus its unacknowledged buffer. The runtime reopens only a running container's log, and its answer ends a capture only on definitive evidence that nothing writes the file any more: the container exited, its sandbox stopped, or it was removed. A container created but not yet started, or one in an unknown state, may still write: the reopen is refused as indeterminate like a runtime that could not answer, so the renamed file gets its name back, the capture keeps running and no rewrite touches its files, and a later pass asks again.

A rotation survives a Pallet process that stops inside it. Between the rename and the runtime's reopen, the renamed file is still the one the runtime writes, and on disk that window is a renamed file (workload.log.1) without a fresh workload.log. Every pass that finds it so asks the runtime to reopen before it reads on, unless the container ended, and every step that rewrites or removes the renamed file completes the rotation first, so none ever replaces or unlinks a file the runtime still writes: output written there would otherwise be neither delivered nor counted, and would grow unseen on the tmpfs. When the runtime cannot answer a reopen, the renamed file gets its name back by a hard link, never a rename over the name, so a fresh file the runtime created meanwhile (a reopen whose answer was lost) is never replaced; a process that stops between the link and the removal of the renamed name leaves one file under both names, and the next pass removes the renamed name. A reclaiming rewrite that stopped before its rename leaves a copy (workload.log.1.compact) that the capture removes when it opens.

A pass that had to drop also reclaims the read part, which the waiting batch no longer needs on disk: it rotates the file at once and rewrites the renamed file from the first unread byte, so after that pass the attempt holds on disk only its unread bytes, within the bound, however long a batch waits. The bound holds whatever the controller does: every delivery pass keeps the captures it reaches within it, a pass whose batch fails holds every capture before it ends, and the controller client holds every capture on the delivery cadence (deliveryIntervalMs) whether or not a controller is connected, so a controller that is down, fails every batch or leaves the method unhandled never lets a capture grow. A hold that fails writes Pallet log capture hold failed: <class> <code>[:<site>]. to stderr, which Spark forwards to the node journal, once per distinct failure until a hold succeeds. The capture's delivery, hold and reconcile run one at a time. A capture's own failure — its files, or a runtime that cannot answer its reopen or cannot yet tell whether its container will write — is that capture's alone: the delivery pass or hold goes on with every other capture, the failed one is tried again on the next pass, and the pass reports the first such failure once it is done (a delivery pass rejects, a hold writes the line above). A controller's failure still ends the delivery pass, after every capture is held.

An ended capture is never written again, so it keeps on disk only what it still needs: the file bytes of its waiting batch and the file bytes not yet read, CRI headers and the cut rests of long lines included. Once its files hold more beyond those bytes than those bytes and 64 KiB (palletLogCaptureBounds.settleSlackBytes), it rewrites them into one file (workload.log.settled, renamed over workload.log), checked when it ends and after every acknowledgement. The cursor moves to the rewritten file in one transaction before the rename, so a restart at any point continues from exactly the recorded byte. An ended capture therefore holds on the tmpfs at most twice the file bytes it still needs plus 64 KiB, instead of every byte the running container left behind. The bound is amortized on purpose: rewriting after every acknowledgement would hold only the needed bytes, but would copy what is left once per batch, quadratic in the capture's size (about 2.3 GiB of copying to drain one 64 MiB capture in 900 KB batches). Rewriting only once the files are at least twice what is needed copies each byte a bounded number of times, never more than was acknowledged or dropped before, so draining a capture copies at most its own size. Since the ended captures measure at most maximumEndedCaptureBytes together in the file bytes they need, after every pass they hold on the tmpfs at most 2 × 64 MiB + 64 × 64 KiB = 132 MiB (138 412 032 bytes) together.

Durable cursor. Each capture keeps one record in the node's own database (pallet_log_cursors): the capture, its last recorded sequence, the position after it (the kernel identity — device, inode, birth time — of the file it points into, the byte offset and the line state of both streams), the evidence that leads the next batch, and the waiting batch. A batch is recorded when it is created, before it is sent: where its file bytes end, and every entry its creation generated with its timestamp — the losses it opens with and the empty entries a final batch closes open lines with — and the discard count it states. The cursor moves on only after an acknowledgement is recorded. A restarted Pallet process continues the same capture from exactly the recorded byte, mid-line included, and creates the waiting batch again byte for byte from those file bytes and the recorded entries, so a sequence number once sent is never sent with other content. Evidence that leads a batch holds one loss per reason, however many passes drop before it is sent. Only a cursor whose file is gone — a reboot emptied the tmpfs, or a reclaiming pass rewrote the file since the last acknowledgement — starts a new capture with a capture-restart loss of unknown extent. When the container stopped, failed or was removed, the capture drains the rest, closes a line left open, sends one final batch, and once that batch is acknowledged records the capture finished, then removes its directory and then its cursor; a restart between those steps completes the removal and never counts the delivered capture as discarded. A capture discarded whole first, under a not-kept answer or past the ended-capture bound, loses both as well. Its record goes first, so a restart before its directory is gone leaves a directory without a record; the attempt's assignment stays listed (the assignment store keeps every attempt's state as a tombstone), so the next process opens that directory as a new capture of the ended attempt, delivers what it holds and removes it. The first pass never removes a directory merely for lacking a record: a running capture has none until its first batch is created, and a directory whose attempt no state names proves nothing about its writer (a database restored from an older copy while the container runs), so such a directory stays until the reboot that empties the tmpfs. A batch dropped under not-kept records the cursor too, at the dropped number and past the dropped output, together with the evidence the next batch opens with, so a restarted process skips the same number and states the same loss. The first pass of a process forgets every cursor whose capture directory is gone (a reboot emptied the tmpfs) and passes on the count of every carrying batch no open capture creates again; a cursor recorded finished goes uncounted.

pnpm exec tstest test/test.logcapture.node.ts --verbose --logfile --timeout 60
pnpm exec tstest test/test.logowner.node.ts --verbose --logfile --timeout 120

Private workload readiness

PalletExecutionOwner binds the published runtime config policy to the exact admitted assignment before invoking a native process, TCP or HTTP readiness probe. Probes inspect the complete CRI identity before and after network IO and connect only to the evidenced container IP addresses. HTTP uses the configured Host and request target, checks final response headers without reading a body, and never follows redirects. HTTPS verifies system trust and the configured hostname/SNI. There is no insecure TLS or DNS-target fallback.

SmartData stores probe evidence, exact container start nanoseconds, boot ID, threshold counters and the next due time. Linux CLOCK_BOOTTIME provides ordering across native children and daemon restarts. Initial delay and startup grace use a conservative anchor from the first completed native observation of that exact runtime identity. An already running container therefore waits the configured delay when first discovered; a forward wall-clock step cannot shorten it. Thresholds and intervals come from the immutable config. Changed runtime identity, boot, config or clock regression resets readiness; pending native effects supply no readiness proof. Probes never overlap. A reconcile before the next due time returns readiness-pending with no new observation.

A ready observation must match the durable sample's assignment, timestamp, image, sandbox and container. Its outbox transaction fences the current native lane. Restart and lost acknowledgements retain the exact existing observation; they cannot reuse a cached probe to create a new ready report. This remains a private composition. Authenticated controller transport and production route promotion are separate integration requirements.

Qualification uses the actual native probe, real HTTP/HTTPS listeners, a controlled CRI server and disposable NoSQLDB stores. The focused tests are test.readiness.node.ts, test.readinessstore.node.ts and test.executionowner.node.ts; they do not establish new real-containerd or host power-loss qualification.

Private packet policy composition

The backend-private composePalletPacketPolicies maps a validated network projection and complete current attachment/handoff receipts into a combined Smartnftables router policy and its host-transit policy. It captures inputs before asynchronous validation. Attached local workloads receive exact veth sources; selected remote workload grants require an explicit caller-owned TUN. Reserved but unattached local workloads receive no packet paths. Private DNS uses only each workload's gateway and explicit TCP/UDP port 53 rules.

Workload egress comes from the published projection grant helper. Router-origin DNS and platform traffic comes only from explicit signed router selections, with the current handoff's transit source and exact selected destination tuple. Public workload grants grant no router authority. Both policies carry the complete protected union and use only the current handoff's leased transport ranges. Withdrawals grant no traffic; DNS readiness does not change packet authority.

Published host ports come only from the verified projection's endpoints placed on this node. Every entry keeps its Cloudly authorization reference; hostIp is the node uplink address or absent; the workload must be attached with a current receipt; a port held by a platform endpoint or resolver on the uplink address, a port inside a leased SNAT range, or an entry beyond the bound refuses the endpoint's whole set by a bounded publication.* site and blocks nothing else. An entry may publish a contiguous range (hostPortEnd, each host port to the same port inside the workload) and may be symmetric: outbound flows the workload opens from its published port(s) leave from the same port(s) on both hops, so a SIP or RTP peer sees one address and port in both directions (@serve.zone/interfaces 32.31.0; pinned at 32.32.0, which also refuses a symmetric entry whose inside port another entry of its protocol shares). Every port of a range counts: a range that reaches a leased SNAT range or a platform endpoint's port anywhere is refused like a single port. The contract admits a symmetric entry only on an endpoint with public egress and refuses the whole projection otherwise, before Pallet composes anything. The bound is the contract's 64 entries per generation, a range counting as one. The accepted set is handed to the host-transit policy as hostIp:hostPort[-hostPortEnd] -> transitAddress:hostPort[-hostPortEnd], journaled with the ACTIVE generation and retired with it. Leased outbound translation keeps using the handoff lease's own source-port ranges; Cloudly allocates new leases in 49152-65535 (runtimeNetworkHandoffSnatPortRange), so publishable ports never meet them, while a lease allocated earlier keeps its range and still refuses any publication inside it. Pallet's session registration offers the installed interfaces release, so Cloudly sends range and symmetric entries only to a node that reads them. The native compiler decides what fits its atomic budget: while it refuses either policy as EXHAUSTED on a bound that publications spend (their count, or the rule bytes, target bytes or operations of the batch), Pallet refuses publishing endpoints by name (publication.capExceeded), last first, and recompiles both. A policy the compiler refuses as INVALID input, or as EXHAUSTED on any other bound, is a composition defect that no refused publication mends: the generation fails with PalletPacketPolicyRejectedError (invalid:packet.policyInvalid or exhausted:packet.capacityExhausted), which carries the engine's bounded reason and, for EXHAUSTED, the bound with its limit and actual count. The router policy carries the second hop of the same accepted set, transitAddress:hostPort[-transitPortEnd] -> workloadAddress:targetPort with the same symmetric flag; it is composed from the same verified entries, cross-checked against the journal and applied by the native compiler, and the raise owns the sandbox default route the workload answers through: CNI ADD installs default via <router address> (static, metric 100) in every pod namespace and advertises exactly that route in its result. Proven end to end in the DHCP-hub packet qualification: an external client reaches the workload over TCP and UDP through both translations, the workload sees the client's own address, the client sees the published uplink address back, a refused entry never opens, and a successor generation without publications leaves both ports dark.

Host grants let the node's own host namespace — Onebox's reverse proxy on a single host — dial an exact workload port directly, with no loopback publication, no route_localnet, no loopback DNAT and no loosened martian filtering. They come only from the contract's getRuntimeNetworkProjectionHostPacketGrants over the verified projection, so only a onebox projection's signed host selections grant any; Pallet proves each again against that projection — the source is the current handoff's transit host address, the destination the exact lease address and port of an endpoint on this node — and refuses the composition otherwise. The same exact tuples are compiled into all three tables the flow crosses: the host-transit and router policies of the ACTIVE generation, for attached workloads only, and the allocation-pool guard, where they are its only exceptions. A new projection composes a new generation and guard target, so a withdrawn selection is an ordinary atomic replacement; an absent or empty set leaves every policy byte-identical.

A granted flow needs the host to route the lease address through the selected handoff, and the ACTIVE generation owns that route. A generation whose projection selects host flows plans one host route per workload pool that holds a selected lease (hostRoutePrefixes, derived again from the journaled projection on every read): the pool via the router's transit address, sourced from the host's transit address, so the flow leaves with exactly the source its grant names. It is added through a command socket in the host namespace inside the registered host transition that raises the link, deleted inside the one that lowers it, checked in every inventory of the raised link, and removed by recovery when it outlives a fence. Any other host route change still retires the uplink observation, so a route change from any other source is never accepted as the generation's own. The route widens nothing: the guard and both hops still admit only the exact granted tuples, and a generation without host selections — every Cloudly generation — plans none and is byte-identical.

Workload ingress grants let the cluster ingress workload reach the port a target workload listens on directly, instead of a hairpin through a published uplink port. They come only from the contract's getRuntimeNetworkProjectionWorkloadIngressPacketGrants over the verified projection, so only a cloudly projection's signed workloadIngress selections grant any; Pallet proves each again against that projection — two different endpoints, at least one placed on this node, the exact lease address of each and the target's listening port — and refuses the composition otherwise. They compile into the router policy's workloadGrants (@push.rocks/smartnftables 4.0.0): a stateful one-way flow of the exact protocol and port, answered only by the replies of a connection the ingress workload opened, never translated, and never open in the other direction. A grant compiles only when both of its workloads are attached behind this router in the generation; an unattached end, or a target on another node, receives no packet authority here, because the compiler has no one-way grant across the tunnel. Large port ranges such as RTP media stay uplink publications. An absent or empty set leaves the router policy byte-identical.

A projection's hostPlatformEndpointIds names the protected platform endpoints this node's own host serves — on a single host the cluster hub, the relay listener and Corestore. They compile into the host-transit policy's localPlatformEndpoints: the host delivers leased flows to them in its INPUT instead of forwarding them, from the exact handoff, with the lease's transit source address and a source port of its range, and only the replies go back. The workload and router selections that reach them are unchanged: a workload still reaches only the endpoints it selects and the router only the hub its block selects. Every address a declared endpoint holds is added to the uplink binding, so the compiler verifies at apply, recovery and inspection that the host holds it; each must be a permanent /32 of the uplink, which the uplink observation admits beside the DHCP lease (see Retained uplink observation). An undeclared platform endpoint or a resolver on one of those addresses refuses the composition as conflict. Without the member the uplink binding and the host policy are byte-identical.

This function performs no native apply, link activation or durable journal write. Its caller must authenticate the projection, retain and fence actual namespace, attachment, TUN and uplink generations, and own routes, DHCP changes and SNAT address lifetime. Native preparation must confirm graph capacity before a complete transition is journalled and applied. Composition alone provides no enforcement, workload readiness, allocation reuse or packet-drain evidence.

Private DNS lease arithmetic

The backend-private ts/network/lease.ts converts an authenticated lease of at most 15 minutes to a conservative native CLOCK_BOOTTIME deadline. Its input is a verified UTC interval anchored to the same boot clock, with an independently qualified rate-error bound. It accounts for uncertainty and the fastest permitted UTC progression. Process recovery retains the exact deadline; persisted boot-time and UTC lower bounds reject clock regression.

A clock this node cannot read lapses that DNS acquisition and never fails the lifetime that owns the clock, and which lapse it is follows from what the clock states. A sample it cannot qualify leaves the acquisition without fresh time evidence: it continues on the window already proved, keeping the exact deadline this boot converted, and lapses only where there is no such deadline — another boot, or a window this boot has not converted. A clock that states no boot reading at all — unqualified, or lost with its native owner — is a clock-unavailable acquisition outright: that pass withdraws every name and the next pass binds them again. No reader is refused for asking while another holds the clock; losing the owned process still fails the node through the clock's own failure signal.

Reboot recovery requires newly verified independent time evidence for that boot. The original absolute expiry stays unchanged, and a saved Cloudly timestamp cannot renew it. The returned anchor, deadline and high-water state must be persisted through SmartData inside the complete authenticated projection before applying a newer native DNS revision. Expiry does not revoke packet grants or prove that an old writer is absent.

The private PalletIndependentClock owns an authenticated NTS measurement process through the native executor. Its pinned Chrony 4.8-servezone1 variant returns exact current Unix time and tracking metadata from one observation. Legacy Chrony tracking offsets cannot supply this precision when the RTC is years wrong. The private process never adjusts the host clock or loads host Chrony/GnuTLS settings. It uses bundled trust certificates, fixed Ubuntu NTS peers, no drift/cookie files, and a root-private Unix socket. One native operation runs at a time, and no caller is refused for asking second: a caller that asks for a reading already in flight is answered with that reading — a reading is of an instant, so sharing it is exact and the owner is asked once — and a caller that asks for the other reading waits for the operation in flight and then takes its turn, in arrival order. Loss of that owned process fails the node lifetime; unqualified or offline samples supply no time authority.

The owner checks coherent tracking, selection and NTS authentication reports, root error, freshness and kernel clock brackets. Supported KVM Linux clock profiles bound BOOTTIME progression by 400,000 ppm; unsupported clock sources, PPS/custom tick settings, VM pause, snapshot restore and live migration do not provide a qualified clock. The bootstrap certificate's validity interval is 1970–2100. A wall-clock reading or an NTP-synchronised flag alone is insufficient.

Clock ownership does not persist projections, supervise DNS or admit workloads. The DNS persistence and lifecycle owner remains unfinished. Lease arithmetic checks cover uncertainty, deadline boundaries, process recovery, reboot proof requirements, rollback and malformed stored state. The isolated native clock harness additionally checks actual NTS measurement, wrong-year RTCs, offline denial and joined Node/Deno cleanup; only recorded successful runs qualify the specified source and kernel.

Private network reservations

The backend-private ts/network/allocation/ store retains protected-authority revisions and immutable handoff leases through the identity runtime's existing SmartData/NoSQLDB connection. runNetworks() exposes callback-scoped operations; shutdown joins admitted operations before closing the engine. Every write, including replay, requires an owning-code identity fence inside its transaction. The caller authenticates controller authority before entering this interface.

A fresh node can stage the complete current authority. An initialized node must advance through exact consecutive references in the same controller epoch and node scope. Reservations reference the current authority and validate against all retained leases; a common revision fence serializes competing allocations. Historical leases continue to bind their exact authority. Quarantine preserves the lease and forbids reuse of its source-port range, conntrack zone and label within the shared contract's respective scopes. Replay retains quarantine. listHandoffs() returns the complete bounded retained history; there is no expiry, pruning, release or reuse operation.

This stores allocation intent. It does not authenticate signed projections, activate pools, acknowledge host/router barriers, declare flow drainage, allocate sandbox IPs or admit packet traffic. Those owners must bind this durable history to exact native receipts before activation. Tests use actual NoSQLDB to check concurrent conflicts, transactional identity fencing, stored digest corruption, lost commit acknowledgements, complete restart recovery and joined shutdown.

Private signed network admission

PalletControllerClient receives the published Interfaces 30.9.0 signing-authority and complete network-projection RPCs on its private outbound TLS connection. The current physical session, active credential and controller/node/namespace scope fence each admission transaction. Socket input is inert and bounded; the transport accommodates the full 896 KiB projection contract.

ts/network/projection/ persists public signing trust, immutable public revision history, one complete signed projection and bounded immutable workload leases through SmartData on the existing NoSQLDB owner. Key rotation and projection admission write a common revision fence. Rotation or revocation prevents the old envelope from being read as current authority; historical keys retain verification of recovery material. Re-signing an identical projection under the new key replaces its envelope before acknowledging replay. Replays never renew the DNS deadline.

Projection admission retains historical protected authorities and handoffs in the same allocation transaction. Full local and remote workload leases survive withdrawal and restart, including the addresses needed for later denial. Retired leases stay quarantined; their subnets and execution attempts cannot be reused. The store retains at most 512 workload leases and rejects capacity exhaustion. Admission always carries pending withdrawals and tombstones because it has no joined native application receipt. An ACK establishes durable intent only.

The callback-scoped runNetworkProjections() facade expires with its callback; shutdown joins admitted work. inspectRetainedProjection() is recovery material, while readCurrentProjection() requires the current signing key. Neither method establishes an application, time, or current-identity fence for a caller. Pool allocation, CNI, native realization receipts, network activation, qualified DNS time and live worker rollout remain separate owners.

The private application-store composition reads current verified signing intent and the complete retained handoff/workload history in one SmartData transaction. Before recording native intent, it rechecks that source fingerprint and writes both common trust and allocation fences in the caller's transaction. Concurrent key rotation, projection admission, reservation or quarantine therefore conflicts with stale intent; a failed identity guard rolls back both writes. This interface does not issue a native receipt or expose an application capability over RPC.

Tests test.networkapplicationsource.node.ts, test.networkprojectionstore.node.ts and test.controllernetwork.node.ts exercise actual file-backed NoSQLDB and TLS sockets, including concurrent rotation, identity rollback, lost ACKs, physical reconnect, shutdown and a complete namespace with 12,288 DNS names. They establish admission behavior, not live packet policy.

Private DNS lease renewal

A renewal carries a fresh DNS window for one exact admitted projection and no applied network state. PalletControllerClient receives applyRuntimeNetworkDnsLeaseRenewal on the same outbound connection as the two admissions, and the envelope's session must be the binding this node obtained itself. That comparison is what makes a cluster relay's push admissible: the relay holds signed statements for a node it carries and adds no trust of its own.

Admission is one transaction in ts/network/projection/, in this order: the stored signing trust, this node's current projection, the signature under that trust and under the revision that signed the projection, the renewal this node already holds for that projection, the published admission decision against the qualified clock's lower bound, the write, and the common trust fence every network admission takes. accepted replaces the single stored row, replay changes no renewal and repeats the same bound acknowledgement, and every other outcome is the refusal the other two admissions give; the trust fence is taken on both outcomes, so a concurrent rotation conflicts rather than interleaves. What persists is the complete signed envelope, so every later read re-proves it against the retained signer rather than trusting the node's own table. An envelope this node can no longer prove is a lapsed renewal, not a broken node: the DNS lane discards it and converts the projection's own window, while a credential bind refuses. The persisted DNS lease and DNS intent each state the renewal their window was converted from; no released version persisted either shape, so there is no migration, and a store carried over from an older build fails its assertion on every read and must be started from clean pallet_dns_lease and pallet_dns_intents collections.

A renewal belongs to one projection envelope: that projection's exact reference and the revision that signed it. The transaction that admits any envelope the stored renewal does not belong to discards it — a successor projection, and equally the same body re-signed after a key rotation, which keeps the reference and changes the signer. A renewal of a projection this node does not hold, at a sequence it has already passed, with different bytes at a sequence it holds, or whose slot has not provably opened at the qualified reading is refused; the sender simply delivers it again once it has. A rotated signing key renews nothing: the re-signed projection discards the renewal held and the node lives on the window that projection itself signs. While that window still runs nothing else changes; once it has ended the node withdraws DNS and stays dark until the controller delivers a renewal signed by the new revision. Pallet asks for none — a renewal arrives on the connection this node already holds — and converts the one it admits on its next DNS pass, which runs once a second.

Everything window-bounded then follows the effective window — the admitted renewal's when one renews this projection, the projection's own otherwise. The DNS lease converts that window in the same transaction that writes the lease and records the renewal it converted, the composed DNS snapshot proves its validUntilBoottimeMs against it, and the managed-VPN credential may expire inside it and never beyond it. Nothing else moves: the application, its fingerprint, the native generation and the protection receipt are untouched, and a node that loses qualified time or current signing authority still withdraws DNS exactly as before.

test.networkdnsleaserenewal.node.ts, test.dnsrenewalview.node.ts, test.dnsleasestore.node.ts, test.independentclock.node.ts and test.controllertunnel.node.ts state admission, supersession, replay, each refusal, the boundary at an equal expiry, the moved deadline, restart recovery, a stored renewal this node can no longer prove on either lane, a reading two overlapping callers share, a clock this node cannot read, and the bound the credential keeps.

Private identity store

The separate backend-private ts/identity/ implementation owns Pallet-only credential state in an explicitly supplied SmartData database. It is not imported by the public probe facade and has no Cloudly connection. Its protected lifecycle and node-local enrollment control owners remain private and unwired from the production daemon.

Preparation commits a 32-byte cryptographic palletToken before returning its SHA-256 proposal. Binding checks this owner's durable hash and the complete enrollment digest. Activation requires the exact bound acknowledgement and writes its immutable receipt in the same transaction. Initial adoption may name an existing Cloudly node while Pallet is still generation zero. Rotations retain the old active bearer until acknowledgement; historical receipts cannot reactivate it or replace pending material. Inspection, proposals and receipts contain no bearer.

The enrollment-specific helpers use published Interfaces 28.2 snapshots, digest and acknowledgement binding. Those helpers prove content binding, not remote authentication: the eventual authenticated transport/coordinator owns that trust boundary. Spark credentials never enter Pallet's private store.

Tests use disposable file-backed NoSQLDB 10.5.1 instances and prove exact replay across engine restarts, concurrent preparation, acknowledgement rollback, independent database bindings and rejection of changed or malformed input:

pnpm exec tstest test/test.identitystore.node.ts --verbose --logfile --timeout 60

Private identity runtime

ts/identity/classes.identityruntime.ts owns a Linux file-backed NoSQLDB 10.5.1 engine and its SmartData connection. Production defaults to UID 0 and /var/lib/serve.zone/pallet; trusted owning code supplies the installed engine's absolute path and SHA-256. It rejects symlinks, unsafe ancestors, wrong ownership, engine hardlinks, nonexecutable or writable engine files, and non-0700 data roots. It never repairs permissions or discovers an alternative executable. Ordinary startup requires existing storage; only explicit first provisioning may create it. The installer must exclude concurrent changes to the selected engine and paths.

NoSQLDB's fileStorageStartup: 'current-format-only' rejects legacy/mixed storage and orphan migration staging without modifying that tree. The native file-root lease is the only cross-process database lock. Model preparation happens after native ownership and readiness, with no duplicated format detector or lock helper. Each attempt uses its own mode-0700 /tmp/pallet-identity-* directory for the Unix socket; it contains no persisted application data. Normal cleanup removes only that owned socket and empty directory. Parent death can leave an empty disposable directory for OS temporary-file cleanup; startup never scans or deletes others.

run() admits a guarded identity facade, not the database or store object. Its methods expire when the callback settles, and admitted calls are tracked even if the callback does not await them. Stop immediately blocks new callbacks, drains admitted work, closes SmartData, and then confirms native exit. A failed database close retains the running engine's lease until cleanup is retried. Start/stop deadlines bound caller waiting, not resource ownership; pending or failed cleanup blocks restart. Native exit invalidates readiness and requires stop before restart.

pnpm exec tstest test/test.identityruntime.node.ts --verbose --logfile --timeout 60

The Linux tests qualify competing Node owners, successor receipt recovery, startup/stop races, cleanup failures, exact pending recovery after native SIGKILL, and active receipt recovery after an idle Node parent's SIGKILL. These are not hardware power-loss or ARM64 hardware tests, and do not prove bounded native exit when a management command is stuck. Production process supervision and coherent installer/engine packaging remain integration gates. When packaging that engine, retain NoSQLDB's own license and complete third-party notices alongside the binary.

Private enrollment control

ts/control/classes.enrollmentcontrol.ts owns its identity runtime and a root-only Unix listener at /run/serve.zone/pallet/control.sock. It acquires the native database lease before touching that socket. The installer provisions the trusted platform parent; the control owner creates its mode-0700 IPC leaf when absent, creates a mode-0600 socket and leaves the directory in place after shutdown. Existing directories are checked, never repaired. Paths must be canonical, symlink-free, protected from untrusted writers and fit Linux's socket path limit.

Only the five published requests.pallet enrollment and runtime-binding methods are registered. Prepare returns a current, independently retained hash-only proposal; bind proves the full pending enrollment; activate checks the complete authenticated Cloudly acknowledgement and then proves the exact current local identity. Historical receipts cannot masquerade as current activation. Spark must authenticate Cloudly before forwarding its acknowledgement or runtime routing binding over this trusted local channel. bindPalletNodeRuntimeBinding accepts the first exact routing tuple or an exact replay only after its origin and node match the current active Pallet identity. readPalletNodeRuntimeBinding requires that same current identity and rejects an absent binding. The binding is a separate singleton: absence is the unbound state, and the released identity document's relayOrigin field remains readable but inert rather than being expanded into unauthenticated routing data. The IPC does not expose a bearer, database, generic rotation or workload execution method.

TypedRequest 8.0.3 owns routing and envelope identity. Hooks and incoming response routing are disabled; wire-supplied local authority and unknown methods are denied. Each connection carries one four-byte big-endian length-prefixed JSON frame per direction followed by write-half-close. The server explicitly keeps the response half open until its reply is flushed. Requests are limited to 32768 bytes and responses to 16384, with a default 10-second total connection deadline and 16 admitted operations. Invalid UTF-8, incomplete frames, trailing bytes and missing EOF are rejected. No TCP listener or automatic application retry is introduced.

Disconnects and deadlines close transport but do not abandon an admitted database operation or free its admission slot. Shutdown stops admission, cancels socket I/O, drains handlers, closes the owned listener and then stops the identity runtime. Caller deadlines retain the actual cleanup owner and block restart. A live pre-existing socket is never removed; stale recovery requires the native lease, a completed connection-refused probe, and unchanged owner/inode checks. Because Node/libuv unlinks on listener close, shutdown checks its recorded socket and parent before calling close. A replaced name retains failed-cleanup ownership until the installer/operator restores the original owned path. These checks assume the trusted installer excludes concurrent root-owned path changes; they do not claim protection from hostile root. No arbitrary file or recursive cleanup occurs.

pnpm exec tstest test/test.enrollmentcontrol.node.ts --verbose --logfile --timeout 60

Qualification uses real Unix sockets and disposable native NoSQLDB stores for half-close delivery, independent restart, competing ownership, stale/live/replaced paths, malformed frames, hook isolation, cancellation, admission limits, historical receipt rejection and lost activation results. Spark's concrete client, coordinated daemon integration and production packaging remain separate integration gates. Before distributing a bundled control process, include its JavaScript dependency MIT/Apache-2.0 notices in addition to the existing Rust and NoSQLDB notice material.

Enrollment process lifecycle

The private PalletEnrollmentProcess owns one foreground control lifetime. Readiness follows the native database lease, model preparation and protected Unix listener. Cancellation during startup prevents readiness. Parent stdin EOF, SIGTERM and SIGINT stop admission and drain the existing control owner. A 250 ms watch of the public readiness getter terminates a failed lifetime; it never restarts the engine or retries a database operation. Each control request already checks native readiness independently of that watch.

The enrollment CLI accepts exactly enrollment-serve or enrollment-provision. Normal startup requires existing state; only explicit provisioning may create it. Trusted build code supplies the exact engine, digest and protected paths. No runtime path, UID, digest, bearer or configuration arrives through argv, stdin or environment. Stdin carries only parent lifetime: content is rejected without an echo. PALLET_CONTROL lines use pallet.enrollment.process and contain only ready, stopped, or a static failure reason. stopped follows confirmed cleanup; a failure does not authorize a successor until the supervisor confirms process exit.

pnpm exec tstest test/test.enrollmentprocess.node.ts --verbose --logfile --timeout 60

The lifecycle and source-process tests cover cancellation, EOF/signals, missing state, invalid input, native failure and retained cleanup ownership. A standalone Deno (the pinned 2.9.7) Linux amd64 fixture also qualifies real Unix half-close responses, durable replay, engine and parent SIGKILL, and stale-socket recovery from an unrelated working directory with an empty environment. That fixture supplies trusted test paths and UID. The production artifact, complete notice bundle, installer/supervisor integration and target-platform qualification remain gates; this source adapter is not a production node installation.

Foreground node process

The same control executable accepts runtime-serve for one private PalletNodeProcess lifetime. It reopens existing enrolled state and verifies the protected sibling pallet-runtime executor, pallet-guard, pallet-dns and the managed VPN pallet-vpn against their compiled-in SHA-256 identities before starting local owners. The released build activates the node network: it passes activation with pallet-vpn and its digest, so a node bound to its cluster relay raises its ACTIVE generation and managed VPN tunnel (see "Private ACTIVE tunnel transitions"), and a node bound to a local controller forwards as a single host. It never initializes missing storage or accepts paths, credentials or settings from argv, stdin or environment. The containerd CRI socket is /run/pallet/containerd/containerd.sock: Pallet's own containerd instance, never the host's shared /run/containerd/containerd.sock, which on a Docker CE host is Docker's containerd.io daemon with CRI disabled. The native runtime refuses that shared socket and Docker's sockets by path and by file identity: a socket with the device and inode of the shared containerd socket, Docker's API socket or Docker's embedded containerd socket is refused, so a symlink, hard link or bind mount of one of them cannot pass under another name. Pallet and Docker can therefore run side by side on one host without sharing images, sandboxes or CRI configuration.

The same lifetime runs in two modes that differ only in what ends it. runtime-serve is a supervised child: its parent holds its standard input, and EOF there stops it (Spark's runnode). runtime-unit is the main process of the node service unit (see "Service unit contract"): its standard input is null, EOF is not a stop request, and SIGTERM or SIGINT stops it. Both refuse data on standard input as invalid_stdin.

A start fences what the previous lifetime retained before it reports ready, so a start after a crash does real work first. A supervisor must allow a start at least 600 s before it gives up on ready, in either mode (palletNodeStartTimeoutSeconds in @serve.zone/pallet-bundle). Pallet bounds each native request of a start, not the start as a whole: a request that misses its deadline ends the start with failed. The longest path whose length does not grow with the node's workloads is a start after a crash, on a host whose guard unit already ran in this boot, with an ACTIVE generation retained and a guard expansion pending. Its request deadlines add up to 469 s:

  • identity database 10 s (its startup deadline, migrations included), router namespace 9 s (3 s spawn, two 3 s requests), independent clock 11 s (3 s spawn, 8 s start);
  • guard pass 274 s at smartnftables' 30 s per request: owner open 33 s (with its 3 s spawn), replay of the committed transition 90 s (prepare, reconcile, inspect), the pending expansion 120 s (prepare, then its replay), detach 31 s;
  • handoff set open 11 s (namespace read 3 s, spawn 3 s, open 5 s);
  • fence of the ended ACTIVE generation 131 s: the native's epoch proof 5 s, the retained host packet table's owner 36 s (namespace read, spawn, open), its recompilation 30 s and its release 60 s (release, inspect);
  • application recovery 15 s (epoch proof, prepare and reconcile at 5 s each);
  • execution host identity 8 s (3 s spawn, 4 s request, 1 s termination grace).

Each retained workload attachment of the ended epoch adds its 5 s epoch proof (at most 512, twice the projection's 256 endpoints). The DNS pass, the database transactions, the executable digests and the storage recovery have no deadline of their own. 600 s covers the fixed path and leaves 131 s for those; a supervisor of a node that retains many workloads allows more, and PalletControlLifetime admits up to an hour (469 s + 512 × 5 s = 3029 s). Under the node unit (Type=exec) the service manager does not wait for ready, so TimeoutStartSec= does not bound a start; an installer that waits for the unit's ready line allows the same limit.

Runtime lifecycle lines keep the PALLET_CONTROL prefix with the distinct pallet.node.process protocol. ready means local identity, namespace and execution owners are ready; it does not assert controller connectivity, workload readiness or DNS availability. A start the packet engine refuses because of the running kernel ends with failed reason unsupported_kernel instead of owner_failed (see "Host kernel requirements"). Initial Cloudly connectivity runs independently. The runtime session requires the immutable node runtime routing binding and registers at its cluster relay. An absent binding fails startup; there is no direct-to-Cloudly fallback or second configured relay authority. Every registration response must match the persisted node, cluster, Cloudly controller and runtime namespace before it becomes execution authority, and reconnect rechecks the current identity and the exact persisted binding. The persisted tuple grants routing only; phase and workload authority always come from the fresh authenticated runtime session. Enrollment, the node credential and the registry host stay with Cloudly whichever transport carries the session: workload.registryHost is the publication identity the registry credential is fenced to, and it is never rewritten. Where the image bytes are fetched is a separate statement, workload.pullEndpoint — the relay's own origin for a cluster whose relay forwards the registry, so the node pulls inside its cluster and nothing cluster-side dials the control plane for bytes. The reference stays digest-pinned either way, so the endpoint decides reachability and never content, and the credential this node asks Cloudly for is addressed to whoever serves them. The contract this node relies on from the relay: it is expected to forward the node bearer verbatim and to serve the unchanged typed request contracts, so nothing in the session binding changes for Pallet. New workloads still require the exact authenticated session and registry grant, while persisted node-bound effects retain their existing offline rules. Network and secret references resolve through their owning resolvers on this same session before any native effect; storage references resolve to the local storage claims this node admitted on the same session ("Local storage claims").

Protocol handshake

The contract this node speaks is the installed @serve.zone/interfaces release, and it is stated once: ts/controller/protocol.ts builds one offer for the one session kind Pallet is (palletRuntime) from protocol.createProtocolOffer, with this build's own minimum (32.0.0, raised only by the commit that starts depending on a later minor). Every registration carries that offer, and the answer is the wrapped { session, protocol } the controller returns, negotiated against the offer that was sent before the binding is bound.

Two peers that cannot serve one session are named rather than collapsed into a transport failure. A controller that refuses this build answers an IProtocolRefusal; an accepting controller whose own offer this node cannot serve is refused by this node. Either way the node reports the status protocol-incompatible and retains the refusal in getStatus(). It is not an owned failure: the workloads keep running, the network stays up, the packet policies stay applied and the node's owners are not torn down — and the node process is not ended either, because protocol-incompatible is a live node with no controller it can speak to.

Recovery needs no operator action on the node. The refused offer is made again every protocol.refusedOfferRetryIntervalMs (five minutes, the one interval @serve.zone/interfaces states so that no client invents a second one) until a controller accepts it, through @api.global/typedsocket's TypedSocketRestoreDeferral: the refused attempt is closed, the transport waits, and it offers again without consuming one of its reconnect retries, so a node refused for days keeps them for real connection failures. An accepted offer drops the retained refusal, returns the client and the node to ready and resumes the pass, so upgrading the controller under a running fleet brings the fleet back by itself. The node's own poll loop keeps its interval while refused — the pass stays empty and admits nothing — because that loop is what carries the accepted offer back into the node's state.

A restoration the transport refuses outright is the one terminal outcome. TypedSocket releases itself, because a retry could only repeat the peer's verdict, and that leaves the node with no controller at all — the condition an exhausted transport leaves it in — so the node fails owned: the process ends and the supervisor's restart is the next offer, which lands on the same five-minute cadence if the controller it meets has still not moved. The denial is named on the controller status as restoreDenial — the transport's own reason and, when the refusal was this client's, the cause behind it — beside any refusal the node was already standing on, so the status the process ends on says which verdict ended it.

The registration also states resolvedAuthorities, the sorted, duplicate-free list of workload authorities this build can resolve. It is the same constant the execution owner enforces (resolvableWorkloadAuthorities), so what the node promises and what it refuses can never drift: an assignment whose requiredWorkloadAuthorities the list does not cover is refused as unresolved-authority. This build resolves network, secrets and storage.

Local controller

A node can be driven by a controller on its own host — Onebox — instead of Cloudly (nodeRuntimeLocalControllerContract in @serve.zone/interfaces). A node carries a Cloudly identity or a local controller, never both: a first local bind is refused while any Cloudly enrollment state exists, and a Cloudly enrollment is refused once a local controller is bound. Both claim one shared pallet_node_authority singleton in the transaction that writes their identity, and the first claim inserts it, so of a concurrent Cloudly preparation and first local bind exactly one commits. The other's commit conflicts with it, the transaction is retried, and the retry finds the committed authority and refuses as conflict. An offline read that finds both kinds anyway refuses instead of choosing one. The persisted identity selects the transport when the controller client starts; there is no configured choice between them.

The local controller is authenticated by the socket's permissions and nothing else: the socket is owned by root with mode 0600 in a root-only directory, so any process running as root on the host is trusted as the controller. The bearer Pallet generates and the controller pins at its first bind detect a Pallet that lost its state (a later bind answering another credential is refused by the controller); they do not keep out an intruder with root.

Pallet serves the local controller on the root-only Unix socket /run/serve.zone/pallet/controller.sock through SmartServe's unixSocket listener, which applies mode 0600 before any peer can connect and refuses a path another process serves. The controller is the TypedSocket client (http://localhost over the socket), but the session keeps its directions. Its first request on a connection is bindPalletLocalRuntimeController: the first bind persists the binding and a bearer this node generates, and every later bind must be the identical binding and answers the same credential generation and SHA-256 hash. The bearer itself never leaves this node except in its own registration. After answering, the node fires the unchanged registerPalletRuntimeSession at that exact connection with that bearer, and it binds the answered session to its own persisted binding (bindRegisterPalletLocalRuntimeSessionResponse) before it becomes execution authority. Admission, reports, registry and secret requests then run as on the outbound transport, but only on the connection that holds the session: another connection on the socket, bound or not, never acts under it. A connection whose registration is accepted supersedes the previous session and closes its connection; a failed registration closes only its own. A controller that refuses this build's offer leaves the node protocol-incompatible until a later connection's registration is accepted.

A local controller has no Cloudly origin to stand for its registry. It issues a pull credential only for an image whose registryHost, and pullEndpoint when one is stated, are both listed in the binding's registryHosts (isNodeRuntimeLocalRegistryWorkload); any other image is refused before a credential is requested.

A node bound to a local controller always serves it from its node lifetime (runtime-unit under its service unit, or runtime-serve). Only network-acquire-local admits the first bind of an unbound node (see Initial projection acquisition). Networked workloads of a local controller need the attach barrier like every other workload; on a node without a managed VPN it holds as a single host (see Private ACTIVE tunnel transitions).

Stdin EOF, SIGTERM and SIGINT close admission and join startup, controller and native operations. The CNI broker remains available through CRI rollback, then joins its admitted handlers before the network and database owners close. The private clock stops and joins its Chrony/query children before the retained router namespace closes and the database owner is released. Terminal node failure ends the lifetime with a static failure reason. The 250 ms liveness watch never restarts the node or retries an operation. A stopped event follows successful cleanup; Spark must also confirm child exit before admitting a successor.

This process composition does not start containerd, provide private DNS, activate Spark's daemon or establish production readiness. Spark must consume the released Pallet bundle and commission its guard before starting this runtime, and Pallet's own containerd (containerd-serve, below) serves the CRI socket /run/pallet/containerd/containerd.sock this runtime uses.

Pallet's own containerd

The sealed control bundle carries Pallet's container runtime under pallet-containerd/: containerd 2.3.6 and its runc v2 shim (the unmodified upstream release executables), runc 1.5.1+servezone1 (built from the signed upstream sources, statically against musl), the pause 3.10.2-servezone1 sandbox image as an OCI archive (pause.tar) and manifest.json, which records every file's size, mode and SHA-256. The control executable compiles in that manifest's digest. No upstream CNI plugin ships: the only network plugin is Pallet's own pallet-cni, which is the verified pallet-runtime executable, and containerd's internal loopback serves lo.

pallet-control containerd-serve is the main process of the containerd service unit. It verifies its sibling pallet-runtime against its compiled-in digest and starts it in its --containerd-management mode, which:

  • verifies pallet-containerd/ against the manifest: exact names, modes, sizes and digests, root-owned and not group- or world-writable up to /;
  • refuses to start a second containerd while one answers on /run/pallet/containerd/containerd.sock;
  • writes the configuration below to /run/pallet/config/config.toml (mode 0600), the CNI network pallet-workloads to /run/pallet/config/cni/conf/10-pallet.conf and a link /run/pallet/config/cni/bin/pallet-cni to the verified pallet-runtime, all recreated on every start;
  • starts containerd with an empty environment except a system PATH (for host helpers containerd resolves by name, such as apparmor_parser), bound to its own lifetime (PR_SET_PDEATHSIG), and waits up to 60 s until CRI reports version v2.3.6 and RuntimeReady;
  • imports the bundled pause archive through containerd's own Transfer service over a Streaming session, as ctr image import does, when containerd does not already hold it with the pinned config digest, unpacks it for the overlayfs snapshotter under pallet.local/pause:3.10.2-servezone1, and requires CRI to report that image id. Nothing is pulled from a registry, and no ctr executable ships.

Then the process reports PALLET_CONTROL {"protocol":"pallet.containerd.process","event":"ready"} on stdout and, when started by a Type=notify unit, READY=1 over systemd's notification socket. containerd's own log lines pass through to the process's standard error. containerd's exit fails the process (failed, owner_failed, exit status 1); it is never restarted inside the process — the service manager restarts the unit. Once containerd is ready, SIGTERM or SIGINT stops it with SIGTERM, waits up to 30 s, removes /run/pallet/config and reports stopped. A SIGTERM during startup kills the starting containerd at once and fails the process; the left-over /run/pallet/config is recreated by the next start. Stopping containerd never stops a workload: every shim, and every container beneath it, keeps running and is reattached by the next containerd. Standard input is not read (a unit's null stdin is expected; data on it is refused).

The generated configuration is complete and is Pallet's own:

Setting Value
version 4
root, state /var/lib/pallet/containerd, /run/pallet/containerd
imports [] — the host's /etc/containerd/conf.d never applies
required_plugins io.containerd.grpc.v1.cri
disabled_plugins io.containerd.internal.v1.opt, io.containerd.image-verifier.v1.bindir, io.containerd.nri.v1.nri
gRPC / TTRPC /run/pallet/containerd/containerd.sock / .sock.ttrpc, uid 0, gid 0
CRI stream server 127.0.0.1:10010, stream_idle_timeout = "15m", no TLS streaming, no CRI TCP service
images overlayfs snapshotter, sandbox image pallet.local/pause:3.10.2-servezone1, no registry host directory
runtime runc via pallet-containerd/containerd-shim-runc-v2 and pallet-containerd/runc, runc state /run/pallet/runc, SystemdCgroup = false
CDI enable_cdi = false, no specification directories
CNI /run/pallet/config/cni/{bin,conf}, one configuration, internal loopback
other image-defined volumes ignored, network namespaces under the state directory

Workloads get cgroups under the absolute parent /pallet (the executor names it on every sandbox), never under the unit that runs containerd: runc can enable the CPU, memory and PID controllers there because the cgroup holds no processes, and a stop of the containerd unit reaches no workload.

The CRI stream server (exec, attach and port forwarding) listens on loopback only, and only root may dial it: every allocation-pool guard policy carries the Smartnftables loopback port owner { address: '127.0.0.1', port: 10010, uid: 0 } (see Offline allocation-pool guard journal). Another user's connection is reset, and the port is dropped on every interface but loopback. The containerd unit requires the guard unit, and a guard pass that did not apply the owner fails, so containerd never serves without the rule.

Service unit contract

Spark on fleet nodes and Onebox on its own host install the unit; Pallet never writes a unit file. The unit must:

  • run <bundle>/pallet-control containerd-serve with no further arguments as its main process, as root, StandardInput=null;
  • use Type=notify with NotifyAccess=all (the ready notification comes from the native owner, a child of the main process);
  • use KillMode=process, so stopping the unit signals only containerd-serve, which stops containerd itself, and every shim and workload survives;
  • use Delegate=yes, as containerd's own unit does, so systemd leaves the cgroups of the processes it starts to them;
  • set TimeoutStopSec= of at least 45 s (30 s containerd grace plus the owner's joins), Restart=always with a short RestartSec=, LimitNOFILE=infinity, TasksMax=infinity and OOMScoreAdjust=-999;
  • order After= and Requires= the retained guard unit, and be ordered Before= the node unit, which Wants= it: the node runtime is the only CRI client.

The node unit runs the node lifetime:

  • run <bundle>/pallet-control runtime-unit with no further arguments as its main process, as root, StandardInput=null, Type=exec;
  • use KillMode=mixed: SIGTERM reaches only the main process, which joins every owner and child it started, and whatever remains of the unit's cgroup when the stop timeout ends is killed before the service manager starts a successor;
  • set TimeoutStopSec= of at least 360 s (a five-minute native operation, then the database joins) and Restart=on-failure with a short RestartSec=: a failed lifetime is restarted by the service manager, never inside the process;
  • order After= both the guard and the containerd unit, Requires= the guard unit and Wants= the containerd unit.

A start of the node may take up to 600 s before its ready line (a start fences what the previous lifetime retained first; the derivation is in "Foreground node process"). Type=exec does not wait for it; an installer that does allows at least that long.

The node unit is started only after the node's first projection is acquired (network-acquire-local or network-acquire); an unbound node fails its start.

The guard unit is a oneshot (Type=oneshot, RemainAfterExit=yes) that runs guard-recover with null stdin, without default dependencies, ordered Before= network-pre.target, systemd-networkd.service and shutdown.target, which it conflicts with, and required by network-pre.target and systemd-networkd.service. Its executable is pallet-control or an installer's own executable that verifies the installed bundle before it starts that mode. The installer writes it only after the first guard-commission.

@serve.zone/pallet-bundle (from this repository's ts_bundle/) states these three definitions for an installer and reads the loaded containerd and node units back against this contract.

The QEMU CNI guest's containerd-serve scenario runs this process from a build:control directory without systemd (see Isolation and qualification); the unit's own KillMode, Delegate and notification behaviour is qualified with the installer.

Router namespace lifetime

The verified pallet-runtime executable also owns one private router namespace for each node process lifetime. Its exact --network-namespace-management mode creates an unnamed Linux network namespace on the initial native thread, before Tokio, sockets or readiness. Before it exposes the namespace it switches IPv4 forwarding off in all and default and disables IPv6 in both, and verifies each (with lo's forwarding): a new namespace copies the host's IPv4 forwarding, and the router must forward nothing unless an UP generation's packet policies filter it. This requires permission to create network namespaces; failure prevents node readiness and controller admission.

The node opens the keeper's live namespace descriptor, compares its device and inode, rejects its own namespace, and confirms the native nonce and identity over the original stdio connection. A reused PID or saved metadata cannot establish authority. Nothing mounts or names the namespace, and no namespace capability is persisted. A new node lifetime creates a fresh namespace.

The private PalletNamespaceOwner.withDescriptor() API retains the source descriptor through an admitted consumer's asynchronous spawn. Callbacks must not close it or keep its number after returning. Shutdown fences new borrowers and joins admitted callbacks. The enclosing node must also join all DNS, VPN and packet-policy processes before closing this owner. Unexpected keeper exit fails the node, stops admission and retains the parent descriptor until explicit joined cleanup; it never silently recreates the old namespace.

The isolated Linux 6.18.35 amd64 fixture runs both Node and compiled Deno (the pinned 2.9.7) against the real namespace keeper and published Smartnftables 1.4.0 binary. It verifies descriptor inheritance, actual namespace identity, keeper-loss fencing, joined borrowing and unchanged parent links and routes, and — on a host that forwards IPv4 — that the namespace is exposed with forwarding off in all, default and lo and IPv6 disabled in all and default. The full node composition is separately checked through its compiled control bundle. Privileged ARM, containerd/CNI attachment, DNS and complete network policy remain separate gates.

Staged secret material mounts

The same verified executable owns one staged mount per admitted assignment that carries a secret reference. Its exact --secret-mount-management mode stages that attempt's material on its initial native thread, before Tokio, sockets or any other work: it unshares a private mount namespace, mounts a noswap,nodev,nosuid,noexec tmpfs there, writes one root-owned file per entry with the delivered mode, places the ownership marker inside the mount and detaches it with open_tree. This requires permission to create mount namespaces and mounts, and it must run in the host mount namespace — the helper refuses its own work when it does not.

The material never crosses the JSON IPC. PalletSecretMountOwner spawns one child per operation and writes the opened values as one published Smartrust sensitive frame on descriptor 3, bounded by this attempt's own declared total rather than by a shared ceiling. Prepare and attach share that one child, because the detached mount dies with the process that holds it; inspect and release run in fresh children, because recovery may never depend on a process that is already gone. No value is ever a return value, a persisted field, an environment value, a command line or an IPC member, and the caller's buffer is cleared after the write.

Publication follows durable state, never the other way round: the mount plan and the pending run commit in one transaction before the first native step, the prepared receipt — secured parent mount, base and staged device/inode, detached mount id, ownership nonce — is persisted before move_mount publishes the mount, and the published mount id is verified equal to the id the detached clone already allocated. The executor then demands exactly that plan's read-only private per-file binds from the container, never a bind of the mount directory that holds the marker, and refuses a live container whose mount list is not that plan.

Recovery reads the persisted receipt alone. A fresh helper verifies the recorded mount by descriptor, unmounts strictly, proves that exact mount id gone and the base directory empty before removing it, and answers uncertain with an exact reason on any mismatch — which retains the mount, the directory and the record and deletes nothing. Stop retains the mount; release is an explicit owner step after proven CRI absence, never a side effect of removal. That order is the only thing that protects a running workload: a container's own bind of a file in the mount lives in its own mount namespace and does not make the host mount busy, so the kernel would not refuse a release taken out of order. A record from an earlier boot names nothing this kernel can find, because /run is a fresh tmpfs after every boot, so it is retired without any native step.

The chain /run/serve.zone/pallet/secrets is walked through no-follow descriptors. /run/serve.zone is the parent the installer supplies for every serve.zone product on the node, so it is required to be exactly what the CNI plugin requires of it — a root-owned directory no one else can write — while the two components below it belong to Pallet alone and must be root-private.

The offline QEMU qualification drives the shipped binary itself, over the same protocol, through every stage boundary on 6.18.35-0-virt, 6.8.0-124-generic and 7.0.0-22-generic, and the containerd guest runs one admitted assignment end to end: sealed material opened, mount published, the container created with exactly the plan's bind, the marker invisible inside it, stop retaining the mount and removal releasing it after proven absence.

Local storage claims

A controller — Cloudly or Onebox — gives a service a retained, node-local volume with applyRuntimeLocalStorageClaim (IRuntimeLocalStorageClaim, @serve.zone/interfaces 32.24.0). A claim names the volume by identity only: controller, organization, cluster, node, runtime namespace, service and the service's volumeId. It carries no host path, no device, no mount option and no credential; this node resolves the identity to its own storage.

Layout. Every volume lives under /var/lib/serve.zone/pallet-storage: volumes/<volumeKey> is a volume, staging/ holds one being created, trash/ one being purged and import/ the bytes an operator places for an import claim. The volume key is the SHA-256 of the claim's identity and id, so it is stable across every generation and a re-epoched controller keeps its volumes. The root and its four directories are root-owned 0700 below trusted ancestors; a volume is reachable only as the bind a granted container receives. The root is created only while no volume was ever materialised: a root that later goes missing — an unmounted disk — is refused by name (storage.root.missing) and never replaced by an empty one.

Ledger. The claim ledger (pallet_storage_claims, pallet_storage_names, pallet_storage_grants) lives in the node's own NoSQLDB through @lossless.org/client/nosqldb. Admission is one transaction fenced to the session the claim arrived on: a new claim, the next generation (whose previous must name the one this node holds), a replay of the same generation or a historical older one. A same-generation claim with another digest, a next generation that changes the volume's identity, source, initialization or initial ownership, a generation after a purge and a second live claim of the same service volume are refused by name and write nothing. The answer is the volume's state: ready, import-required, purged or uncertain.

Crash safety. Every filesystem effect runs between a ledger intent and the completion that records its evidence, one at a time. An empty claim commits provisioning, then creates the directory in staging/ with the claim's initialOwnership and mode and publishes it by one atomic rename; the ledger records the directory's device, inode and birth time. A purge commits purging, renames the recorded directory into trash/ and removes it there, never following a link the workload left inside. Before the ledger records a step done, every directory it created, renamed or removed is fsynced with its parents (a new layout directory with its parent; the volume with staging/ and volumes/ before ready; volumes/ and trash/ before purged), and a failed sync keeps the intent pending, so a power loss can only undo a step whose intent the next start still finishes. On start, and on every replay, an interrupted intent finishes from what is on disk: a leftover staging tree is this intent's own and is removed, a published directory is adopted only while it is still exactly the fresh, empty directory the intent creates, and a purge whose volume already sits in the trash completes. Evidence that contradicts the ledger — a foreign directory at the volume's name, a volume whose identity changed — leaves the claim uncertain, which this node neither mounts nor deletes until an operator decides. A pending intent that cannot finish at start is listed in the execution owner's storageRecoveryRefusals and stays durable.

ReadWriteOnce. A run that seals storage references is planned before the registry, secret material or any native step: every reference must name the claim's current generation, the claim must bind to the run (bindRuntimeLocalStorageClaimToAssignment), its volume must be ready, still be the recorded directory and be held by nobody else, and no two volumes of the run may nest or cover a path Pallet or the runtime binds itself (/run/secrets, /run/serve.zone, /opt/serve.zone/runtime-assets, /proc, /sys, /dev, /etc/hosts, /etc/hostname, /etc/resolv.conf). A run that cannot have its volumes waits as storage-pending, with the step named (storage.claim.held, storage.claim.importPending, storage.claim.generation, …). The grant is taken in the transaction that begins the run, so a concurrent purge or a second holder conflicts instead of both proceeding. Stop keeps the volume held; the proven CRI absence of the removed container releases it. Release never deletes bytes: only a purge generation does, and a held volume's purge is refused (storage.purge.held) until its holder is removed. The container's mount list is part of the attempt's identity, so every later inspection demands exactly the grant's binds.

Offline import. An import claim waits as import-required until its bytes arrive. With runtime-serve stopped, the operator places each volume's tree at /var/lib/serve.zone/pallet-storage/import/<claimId> on the same filesystem and runs pallet-control storage-import. The pass (in ts_migration/, like every data adoption step) records the placed directory's identity, publishes it as the volume by one atomic rename, fsyncs import/ and volumes/ and records it ready; the tree keeps its ownership, modes and every other attribute, and nothing is copied. It reports the adopted, waiting and refused claims on its ready event (pallet.storage.import). A tree on another filesystem, a file instead of a directory or a volume name already in use is refused by name and moved nowhere; an interrupted pass finishes on the next run, and a recorded tree found in neither place leaves the claim uncertain. Keep a copy of the source first when a rollback may need it.

Network filesystems are not part of this build. The contract (32.24.0) also states nfs and smb sources, but this node materialises node-local directories only: such a claim is refused before the ledger is touched (storage.source.unsupported), so neither the claim nor its service volume name is recorded.

Dormant host/router handoff

The private ts/network/classes.handoffowner.ts mechanism consumes an already reserved, digest-verified immutable handoff lease and a live PalletNamespaceOwner. start() prepares an inert intent bound to the current boot and both namespaces. The caller must persist that complete intent before reconcile(previousReceipt) can create the exact veth pair, with its peer created directly in the router namespace. Both endpoints remain administratively DOWN. Only the leased IPv4 addresses and their kernel local /32 routes are present; connected prefix routes belong to eventual activation.

The separate --network-handoff-management native mode opens one netlink socket in each retained namespace on its initial thread and restores the parent namespace before polling either socket. It uses bounded native operations without shell commands, named mounts or filesystem state. Linux ignores aliases in this veth creation request, so the owner first verifies the newly created DOWN endpoints, sets both ownership markers through link updates, and verifies them before adding addresses. An interrupted partial creation never qualifies for adoption or deletion.

Recovery requires the exact boot, lease, namespace, link indices, peer indices, names, MACs, markers, addresses and local routes. Extra or changed state is rejected. Deletion uses the verified endpoint index and destroys only that exact pair. A release states the administrative state its own contract guarantees: every handoff path releases a dormant pair, so an endpoint raised behind the owner's back is refused as handoff_conflict and the pair stays, and only the raise measurement releases a raised pair. Removing a raised workload pair in production is the stage-aware progress rollback, never this release. Privileged exclusivity belongs to the enclosing node; these observations are not a compare-and-swap guarantee against concurrent privileged network writers.

The TypeScript owner serializes operations and fences namespace or native-process failure. getPrepared() and getReceipt() retain detached immutable evidence after failure. close() joins the native lifetime; a successful return with released: false means cleanup remains unconfirmed and the lease must remain quarantined. A rejected close has not proved that join; joined exposes its state. Join every namespace consumer before closing the keeper. Kernel namespace teardown is asynchronous, so keeper exit alone cannot establish link deletion or pool reuse.

Both native handoff management modes use bounded nonblocking pipe or stream-socket stdio on the same thread that owns their namespace sockets. Host, router and retained sandbox netlink connections remain polled while input is idle or partial and while a response waits for the parent to read it. Partial input survives cancelled reads; malformed or oversized frames terminate the protocol. Output has the same three-second terminal bound as native commands. A partial response is never retried on the same stream. No reader thread can retain the mutation lock. This establishes protocol-owner liveness; the dormant sockets still have no topology multicast subscriptions and provide no continuous route or link fence.

This mechanism is not yet composed into signed node network admission. Durable realization receipts, router/host/Docker policy barriers, CNI, DNS, link activation and production qualification remain separate prerequisites. The offline handoff fixture is test/native/handoff.ts, driven by test/native/qualify-handoff.py in a disposable amd64 VM without a network device, host disk or host mount. It runs the actual native binary through Node and compiled Deno, including partial-state and foreign-state faults, restart recovery, keeper loss and joined cleanup. ARM64 is built but this privileged fixture does not establish ARM64 execution qualification.

Complete dormant handoff set

ts/network/classes.handoffsetowner.ts supplies the private node-wide mechanism for complete retained membership. Start it with the live namespace, persist the empty getPrepared() intent, then reconcile(target, null) to inspect and establish an empty native receipt. prepare([{ lease, presence }]) accepts all 256 complete, digest-verified leases permitted by the allocation contract, including absent/quarantined history, with at most 32 desired present handoffs. Persist the returned target and the last applied receipt before calling reconcile(target, previous). Preparation does not change the kernel. Receipts contain an exact DOWN pair or explicit absence for every member; absence never frees an allocation.

The --network-handoff-set-management process retains both namespace descriptors and owns one host-network-namespace abstract Unix socket lock. The single-pair mechanism uses the same lock, so no two Pallet handoff mutators can coexist on that host namespace, even with different router namespaces. This IPC lock is concurrent process exclusion, not authentication or durable ownership. The node must still exclude other privileged network writers and namespace capabilities.

Before any transition effect, bounded complete link, address and all-family/table route dumps inspect the whole expected set. The private router namespace must contain only its empty, DOWN loopback and exact dormant handoffs. Unknown router interfaces, addresses or routes fail ownership. Unrelated host interfaces and routes remain permitted; unexpected Pallet markers/names, misplaced known MACs, transit address conflicts and surviving/reused receipt indices are rejected. Stable bound facts are re-read after inspection and after the complete transition. Typed requests retain the terminal NLMSG_DONE; every inventory requires a successful completion and rejects interrupted, filtered, malformed, unexpected or oversized responses. Both namespace connections also reject unsolicited messages and receive-buffer loss. The pinned maintained netlink-proto source corrects decoding at its owning layer so malformed packets cannot disappear before a later successful completion.

Transitions retain all prior members and cannot revive an absent member. Exact present-to-absent deletion keeps the lease in the set. Separate pair operations are not atomic: after a lost response, persistently prepared intent and the prior receipt can adopt exact completed members and finish the remaining transition in the same live namespace. Changed or half-created pairs stay failed-owned. EOF or process termination preserves DOWN effects for that recovery. Full node death does not make an unnamed namespace recoverable. close() deletes only exact known members and joins the native lifetime; uncertainty preserves receipts and returns released:false, with no guessed cleanup or allocation reuse.

The native transport permits 1,048,576-byte frames; TypeScript captures inert complete-set data within the corresponding bounded budget. The serialization test measures a 701,951-byte replay with maximum-length lease and namespace identities. Native operations retain their three-second deadline and the bridge its five-second deadline: the 256-retained/32-present guest operations completed within the native deadline in the isolated fixture. Complete dumps establish absent members without issuing redundant per-member link dumps. The offline handoff fixture additionally covers empty/nonempty exhaustive inventory, unrelated host links, competing single/set owners, a competing router namespace, quarantine, lost-response and partial-transition recovery, 256 retained members with 32 present paths and joined wrapper shutdown under Node and compiled Deno. This stage exposes no UP method and issues no nativeBarrier: the authenticated SmartData application journal, positive protection evidence and eventual firewall/DNS/VPN/CNI lifecycle remain required.

After reconciling the complete dormant topology, the private handoff-set owner can openUplink() and return a boot-, host-namespace- and random-generation-bound observation. inspectUplink(generation) performs fresh bounded kernel and networkd reads; copied facts cannot recreate the observer. A wrong generation is rejected without retiring the current observation. closeUplink(generation) joins the read-only observer and does not release topology or change DHCP.

The initial supported host has one physical Ethernet uplink, a finite bound systemd-networkd DHCPv4 lease, one main-table IPv4 DHCP default route and the three canonical unselected IPv4 policy rules. The observer records exact link, address, gateway, source, metric and IPv4/IPv6 all/default/interface NETCONF settings. Beside the lease the uplink may carry permanent /32 IPv4 host addresses — the addresses a platform service on this host binds, whether networkd configured them (Address=) or found them on the link. Each is a universe-scope address with an infinite lifetime, so it adds no prefix route and never becomes the default route's source; any other IPv4 address on the link still refuses the observation, and adding or removing one retires it like every other address change on the uplink. The observation states them (hostAddresses, sorted as strings); an observation a 33.0.0 node journaled carries none, which states nothing. Existing host IPv6 addresses remain possible; this observation grants no IPv6 workload authority. Networkd remains the only DHCP mutator.

The native lifetime subscribes before capture and drives both RTNL notifications and the fixed system D-Bus connection during inspection and quiet management I/O. Its unique networkd owner and exact returned link object path remain bound to the lease. The observation follows the lease networkd renews, so a node keeps forwarding through every DHCP renewal: at the lease's renewal time (T1), on a property event of the link, and when networkd rewrites the leased address with new lifetimes, the native reads the uplink and its lease again. While networkd renews or rebinds, the lease in force holds until its valid lifetime ends, never longer. A fresh acquisition from the same networkd owner for exactly the same binding — link, address, prefix, gateway, source, metric, NETCONF settings and host addresses — whose lifetimes the kernel address already carries becomes the lease in force; the observation keeps its generation and its receipt, which states the lease it opened on. The lease's expiry, service replacement, protocol loss, a read that finds the binding changed in any fact or the lease moved backwards, and every other relevant kernel change retire the coordinator. Current values cannot erase an observed kernel change and restoration: only the leased address rewritten in place asks for a read, and any other address, link, route, rule or NETCONF change of the uplink still retires. Unrelated link notifications are ignored; route/rule changes conservatively retire this observation until exact coordinator-owned transitions are implemented. Close the observer before further dormant topology changes.

A coordinator that ends while it waits — its uplink observation retired, an ACTIVE observer failed, a namespace channel lost — lowers every ACTIVE link on its way out, as it always did. It names the watch that ended it in a bounded PALLET_FAULT line (handoffset.drive.*, uplink.watch.*, uplink.follow.*), and the handoff-set owner writes its failure chain once, as it retires, to stderr, which Spark carries into the node journal: Pallet network handoff set failed: <chain>. The stop that follows names the same chain as the reason its release could not be confirmed (noderuntime.close.networkRetained < …).

This mechanism exposes no UP or packet-readiness result. Active routing, firewall composition, durable active journals and joined withdrawal still belong to the activation coordinator. The ignored Rust test network_handoff::kernel::uplink::tests::qualified_networkd_host_observation provides an explicit read-only check on a real networkd-managed host; it changes no interfaces, addresses, routes, sysctls or services.

Private ACTIVE tunnel transitions

The retained PalletHandoffSetOwner accepts beginActiveTunnel(effect), finishActiveTunnel(effect, expectedReceipt) and inspectActiveTunnel(effect) while its exact ACTIVE generation is UP. The effect binds the complete ACTIVE preparation and a creation plan or previous tunnel receipt to a durable intent digest. It includes the authenticated SmartVPN generation, complete plan digest, router namespace, exclusive interface name, /32 address, MTU, split routes and same-boot BOOTTIME deadline. Tunnel routes cannot capture owned transit or local workload prefixes, the control address or the hub endpoint.

Persist the intent before begin. Retain the actual SmartVPN client in the router namespace, prepare its tunnel, freshly inspect that generation, and derive the expected receipt from its actual interface index before finish. Native code observes the bounded ordered kernel transition, checks the full namespace twice and returns to its strict observer. It does not create the tunnel or infer the continued existence of the VPN owner from a saved receipt. Packet policy must be installed and verified before enabling VPN forwarding or raising workloads.

For deletion, stop forwarding while retaining the UP tunnel, lower the packet policies back to their base shape while the device still exists, persist and begin the deletion intent, then disconnect and join the VPN device owner before finish with null, and release both packet policies only after the generation's withdrawal. Each generation therefore carries a chain of policy pairs per role, all journaled and each exactly one revision above the pair it replaced: the base pair at the generation number (composed with no tunnel, so remote peers are deferred rather than permitted), the tunnel pair (the same projection composed with the retained TUN bound as the pallet_vpn endpoint), the lowered pair (the base shape again) — a tunnel and a lowered pair for every device the generation raised, each device's above the lowering of the one before — and between them one amended pair for every workload attached while the generation is UP (below). A generation raises at most 32 devices (maximumActiveTunnelCreations): every read of the chain walks the tunnel lane back to its first command and revalidates each row, so the lane bounds the cost of every tunnel-phase, amendment and cleanup step (64 rows at most on a lane this release admits). The journal refuses the next creation by name (exhausted, active.tunnel.exhausted) before it admits a command. The orchestrator refuses it before it asks for a credential to raise a device or dials, and refuses a replacement under another credential before it lowers the live device or dials; only the renewal's own credential request, which names the other identity, comes first. The native's own bound is larger (256 session generations per generation, maximumNativeTunnelGenerations, remembered so none is reused), and the journal's must never exceed it. The pass fails with that name, its owner withdraws the generation in order, and the next pass raises a fresh generation whose lane starts empty. The bound governs admission only: releases up to 35.2.0 bounded a lane by the native alone, so a lane they wrote holds up to 512 rows (maximumStoredTunnelCommands). It stays readable — inspected, recovered, lowered, cleaned up and withdrawn as any other, each chain read costing what it cost under that release — and only its next creation is refused. The compiler releases only the exact graph a transition names, so no revision is derived from a phase: every pair states its revision and the acknowledged pair it replaced, and activePacketChain proves the whole chain on every path that admits or acknowledges a pair above the base pair, and before a recovery re-applies or a cleanup releases the live pair. The tunnel pair is applied only after finishTunnel, because the compiler binds the device by the interface index the receipt carries, and always before forwarding is enabled; the lowered pair is applied before the delete command, so no enforced graph ever names an absent link. Linux omits deletion notifications for the static split routes; full inventory must prove their absence. ACTIVE withdrawal is rejected while a tunnel or unfinished transition remains. Unexpected notifications, owner loss or the absolute lease deadline retire ACTIVE and fence known interfaces DOWN. Recovery requires joining the VPN owner first and proving the whole namespace through a fresh ACTIVE owner. No receipt grants allocation reuse or proves packet drain.

The private protocol allows 2 MiB frames, including complete retained history and up to 1,024 split routes.

Both the durable tunnel journal and the application lifecycle composition now exist. PalletActiveStore persists every managed TUN command as its own row — beginTunnel admits one create or delete before the device owner is asked to act, acknowledgeTunnel records the native begin, and finishTunnel records the terminal fact — chained per ACTIVE intent so a delete consumes exactly the receipt its predecessor left behind, and a normal withdrawal is refused while a device is still retained or a transition unacknowledged. A recovery records DOWN only; it never claims a deletion, and the retained row stays as history. The creation row also carries the tunnel pair applied over the device (packets) and, once the tunnel is being taken down, the lowered pair (lowered); a cleanup row names the live pair — the last one the chain admitted — by its exact transitions and policy digests, and the packet owner's release is fenced on exactly those digests.

A workload may be attached while the generation is UP. The native admits the pure prepareWorkloadAttachment under a held generation, and an addWorkloadAttachment for a pair the generation does not hold yet is created by the generation itself, through its own observers, while it is UP with no tunnel transition pending, holds fewer than 64 pairs (prepared and attached together) and has no tunnel reaching into the new subnet. 64 is what the packet engine is measured to hold: Smartnftables 3.0 compiles a router of 64 workloads, each with local DNS, public TCP and UDP egress, four platform endpoints and up to three publications, inside its atomic batch with room to 89. The native registry, its router-namespace inventory and configuration (one link per present handoff and per pair, 96) and its admission fence (192 targets) are bounded to match; the compiler's prepare() stays the authority on what a given graph fits. The bound above 32 is verified by compilation and unit tests only; the root QEMU guest qualification has not yet run at 64. The attachment owner asks the held generation before it journals the ADD intent (PalletHandoffSetOwner.admitAttachment, handoffset.admitAttachment.notUp/capacity/tunnelOverlap); a refused ADD releases its preparation and fails by that name with no journal, because a pending one would fence every amendment of the generation while it is held. addWorkload judges the same rules again before the native is asked (handoffset.addWorkload.*), because a native refusal retires the owner. In the other order a tunnel is refused where it would reach an attached pair's subnet, by the store before the command is journaled and by the owner before the native is asked. The pair is DOWN and outside every policy at that point. Before it is raised the orchestrator amends the generation (PalletActiveOrchestrator.amend): beginAmendment compiles the live pair again with the attachment as a further source, in the live shape, one revision up, and appends the row { attachment, shape, revision, host, router } to the journal's ordered amendments log in the same commit that fences the current source as an extension of the one the generation carries — the same settled application, every carried workload exact, the new attachment complete; finishAmendment appends each role's acknowledgement after its apply. The intent, the native preparation and the UP record never change: activeSources(journal) is the intent's source extended by the log, and the tunnel pair, the lowered pair, every recompilation and the cleanup compose over it. A generation that already carries the attachment has nothing to amend, a lowered generation takes no attachment, and a node without a held UP generation amends nothing: the pair is dormant and the next generation is prepared over it. DNS needs nothing new — the completed attachment is a source of the next resolver revision exactly as after a dormant ADD. A recovery sends the native attachedWorkloads, the complete receipts of the restored pairs the preparation does not cover (required, [] when none); the native restores DOWN only when its registry is exactly the preparation's workloads plus those. A restored attachment it cannot list — an ADD whose receipt never became durable, or one whose removal is already intended — is removed first, in the fresh process and before its native adopts anything, and journaled exactly as a DEL journals it (PalletWorkloadAttachmentOwner.recoveryRemovals). A failed journal step, a foreign intent or a failed native removal refuses the recovery by its step (handoffset.recoverActive.beginRemoval/removalIntent/incompleteAttachment/finishRemoval).

A pair raised under a held generation stays UP through its withdrawal, and the withdrawal restores the router namespace's baseline configuration. The raise is therefore refused by name (active.raise.ipv6_baseline) unless that baseline keeps IPv6 disabled on the pair's router side — all/disable_ipv6 must be 1 before the generation is prepared — because IPv6 enabled again on an UP link brings addresses and routes no owned transition admits. The same rule applies to pairs raised before the generation is prepared (active.prepare.ipv6_baseline). The native router owner establishes and verifies all and default IPv6-disabled baselines immediately after creating its private namespace, before exposing it.

The router namespace is fail-closed between generations for the same reason: the withdrawal releases both packet policies while the raised pairs stay UP, so the IPv4 forwarding it restores must be off. The native router owner switches all and default forwarding off at creation, and a generation is prepared only when every forwarding switch it captures — all, default, lo and each router link — is off; otherwise the preparation is refused by name (active.prepare.forwarding_baseline). Activation turns forwarding on only after both packet policies are applied, and the withdrawal turns it off again before the policies are released, so workload traffic through the router stops from a generation's withdrawal until its successor is UP (a projection change's withdrawal window) and is never forwarded unfiltered. For a containerd sandbox, the admitted native ADD disables IPv6 only on its newly created, identity-proven DOWN eth0 before addressing it; the exact pair DEL removes that link. Containerd runs CNI before creating the sandbox container, so CRI sandbox sysctls cannot provide this pre-CNI guarantee. The sandbox's namespace-wide defaults and the host's sysctls are never changed. A node with the mpls_router module loaded is not supported: every link registration emits an MPLS netconf record, which the attach vocabulary refuses.

The tunnel credential is the controller's. PalletControllerClient.resolveTunnel issues getRuntimeManagedVpnCredential { session, projection } over whatever transport the runtime session uses — the relay when the node is bound to one — for the exact projection the generation was raised from and a reading of the qualified clock, and turns the answer into the session request the device owner authenticates with: the cluster hub's address and port, key material, authority and node ids, and a same-boot BOOTTIME deadline bounded to fifteen minutes or the credential's own expiry, whichever is sooner. The node dials its own cluster's hub, never a platform-wide one: the credential names the hub endpoint this node's signed projection selected, and the contract binds that id, its transport's protocol, its address and its port to that selection. The transport decides the dial form — a managed QUIC hub is a bare host:port, the one transport the contract states — so a credential naming any other is refused rather than dialled as if it were QUIC. Every refusal is named (controller.tunnel.notLive, disconnected, projectionAbsent, projectionChanged, boot, credential, transport); the key material is returned to the caller only and is never journaled or logged. The orchestrator asks for it once the generation is UP and DNS is serving; a refused credential or a refused authentication leaves the generation UP without a tunnel, reports the refusal on the barrier as tunnelFailure (code and site, nothing else) and on the application pass as activation.tunnel:<label>, and the next kept pass asks again. Once the command is admitted every later step is a journaled transition and fails the activation as such.

A tunnel lease lasts at most fifteen minutes and is moved in place, never by tearing the tunnel down: every admitted DNS lease renewal moves the window the credential lives in, so the next kept pass asks the controller for a fresh credential, and a pass that finds the lease deadline within three minutes of the qualified clock does the same (renewalMarginMs). Under the identity the session authenticated with (hub address and key, client key, authority and node) both bounds on the device move in place: first Pallet's own native authority deadline (PalletHandoffSetOwner.renewActiveTunnel, native renewActiveTunnelLease), which the native arms at the lease the device was created under and at which it fails the generation closed, then the session's lease with SmartVPN's renewManagedSession (PalletManagedVpnOwner.renewLease). The device, its packet policies and forwarding are untouched, and nothing is journaled: the deadline bounds this process's custody, and a successor fences the generation without it. A credential under another identity cannot renew a session, so the tunnel is lowered and a new one raised with it. A refused renewal leaves the lease in force, names the refusal on the barrier's tunnelFailure and is asked again on the next pass; a refused native renewal never asks the session, so the session's lease never outlives the native's.

A lease that lapses anyway ends the native authority at the same instant: the native fails closed, the generation is fenced and the node fails owned, and its supervisor starts a fresh process, which fences the generation by the proof that its epoch ended and raises a new one with a fresh credential. A session whose transport to the hub is lost stops instead: SmartVPN ends its packet I/O and reports managed-session-state stopped, keeping the device for an ordered teardown. The barrier stops holding at that moment — holds requires the session's packet I/O to be live — so attached workload links demote on their next CHECK and admission refuses, and the next node scan runs a network pass at once. That pass lowers the dead tunnel through the journal exactly as a withdrawal does (the lowered packet pair, the delete command, the device released) and closes its owner, then asks for a fresh credential and raises a new tunnel over the same generation; a refused credential leaves the generation UP without a tunnel, named, until a later pass succeeds. A retired session — its device lost behind the owner — is a loss of the owner, not a stop, and fails the activation as before.

PalletNetworkApplicationOwner takes an optional activation option and, when it is present, drives one ACTIVE generation after each completed dormant application: preparation, both packet policies applied and journaled, native UP, local DNS proven to be serving, then the tunnel commands above and finally forwarding. Without that option the owner stays exactly as dormant as before and the DOWN wire contract is unchanged. PalletNetworkOwner.barrier() exposes the result as one of absent, prepared, up or forwarding, with the retained device's journal reference when there is one. Only forwarding sets holds, and only holds may gate workload UP; a node with no managed VPN configured settles at up and its barrier deliberately does not hold, with one exception: the single host. A node bound to a local controller (Onebox) never dials a tunnel, whether or not a managed VPN is configured — the released bundle configures one on every node, and the persisted binding decides — and a generation reaches forwarding without a tunnel when three facts hold: its scope is that local onebox controller, every endpoint of the signed projection it was raised from is placed on this node (the contract already refuses a remote endpoint in a onebox projection; Pallet proves it again from the journal row and fails closed), and this process raised the generation itself. There is no remote peer to reach, so the base packet policies already carry every rule the generation needs. A barrier that does not hold names the missing fact as singleHostGap (controllerNotLocal or remoteEndpoint). Execution admission then rests on the single-host evidence instead of a tunnel receipt: the evidence names the projection, the admitting transaction re-proves the fact from the journal row and fences it, and a refusal is admission.barrier.singleHostUnproven. A tunnel command on such a generation is a corruption and refuses. Every other node — bound to Cloudly directly or through the relay — keeps the mandatory tunnel barrier unchanged. The barrier is always recomputed from the journal and the live owners, never read back from a stored pass. On loss this composition fences and withdraws first and then reports a barrier that does not hold; it never re-activates the lost generation by itself, because re-activation requires explicit newer state from the control plane.

Durable application intent and results

The private identity runtime exposes runNetworkApplications() over its existing SmartData/NoSQLDB owner. readSource(scope) returns the current verified signed projection and complete retained lease history. Native preparation happens outside the database. begin(scope, fingerprint, prepared, guard) then persists that exact whole-set target, its previous receipt and the authenticated source, writing the common trust/allocation fences and the real identity guard in the same transaction. Only the current reserved handoff is desired present; every other retained lease remains explicitly absent. Omitted history and revived absent members are rejected. Ordinary successors remain within the same boot and native namespace.

The journal keeps one pending transition and immutable intent/result records. Exact retries join existing intent; a competing proposal cannot replace it. finish(scope, intentReference, nativeReceipt) records only the result matching that persisted target. It does not require new admission authority: signing-key, projection and credential changes must not discard an already-owned native result. Historical signatures remain verifiable through retained signing revisions. Lost completion acknowledgments replay exactly without advancing or rolling back the current head. No journal method deletes history or frees quarantined allocations.

Keep begin, native reconciliation and finish inside one admitted runtime callback so database shutdown joins the entire operation. Retained facade methods expire when their callback returns. The persistence capture budget covers both transition sides and complete signed source; native frames retain their independent 2 MiB limit. This store does not start a native owner, issue nativeBarrier, report protection, clear DNS withdrawals or authorize workloads. The private application owner supplies node composition and a separate durable unavailable-report outbox.

beginEpoch(scope, fingerprint, prepared, evidence, guard, workloadProofs) records a qualified same-boot absence or different-boot fence into a fresh namespace. The caller must hold the fresh private native owner continuously through its audit, journal commit and reconciliation. The store resolves the evidence's opaque journal reference to the actual persisted pending intent, or applied intent when none is pending, and binds its complete target and exact known receipt. A single transaction retains the old applied/pending references in an immutable epoch record, fences current signed authority and identity, and publishes the next intent. Evidence from a different target, previous receipt, host, boot or namespace is rejected.

The first intent in an epoch has previousEpoch and no native predecessor: reconcile it with previous = null. Its generation follows the old pending or applied head. The still-reserved logical handoff may be realized DOWN again; every retained lease and quarantine survives. Old unfinished intents stay in history without a fabricated result and reject late completion after the epoch fence. Previously recorded results still permit exact acknowledgment replay. An intent has one shape: it names the epoch it opens or null, and a record that does not is refused rather than read. Repeated failures before a fresh intent completes continue the same history. Different-boot recovery requires the distinct native previousBootTerminated proof, persisted as new-boot-host-fence; it cannot use a same-boot absence receipt. Reused numeric namespace identities do not make two kernels the same epoch.

Dormant workload attachment journals

The private ts/network/attachment/ composition binds one admitted pending Run to its immutable workload lease, completed application and exact container ID. Native preparation opens the sandbox namespace once and retains its descriptor. The ADD journal commits before any veth mutation; its immutable identity cannot be rebound after cleanup. Ambiguous admission acknowledgements are resolved on the same application transaction lane before an unadmitted descriptor is discarded. Partial creation uses the recorded removal intent and exact native absence check. The sandbox accepts containerd's internal loopback baseline. Workload links stay DOWN until this node's activation barrier holds; an ADD never waits for it.

Application and attachment operations share one queue. Handoff retirement must fence every unfinished attachment in its transaction. A replacement router waits for distinct historical workload evidence: stable absence in the exact retained sandbox on the same boot, or termination of a different actual kernel boot. A missing or substituted sandbox path is not same-boot absence. Previous interface indices belong to their original namespaces and are never host capabilities.

An application epoch record has one shape and always binds a complete sorted workload proof set. Immutable coverage records commit with that epoch and freeze the old journals. Coverage permits historical DEL to acknowledge the proved absence, while preserving the original unfinished journal, all allocations and quarantine. It does not create an ordinary DEL receipt or authorize another ADD. Quiescence stops preparations; already admitted ADD/DEL operations drain before the shared native owner closes.

Production execution admits network: 'attached' only after the assignment's exact projected lease is prepared with the pending Run operation. The CNI owner then proves and journals its attachment before returning a usable result. An isolated assignment prepares a separate durable CNI intent with no lease or application reference. A secret reference is resolved by this node: its material is staged and published as this attempt's own mount before the container that receives it, as "Staged secret material mounts" states. A storage reference is granted from this node's claim ledger in the transaction that begins the run, as "Local storage claims" states.

Private CNI transport

The node runtime owns a root-private CNI broker at /run/serve.zone/pallet/cni.sock. The installer supplies the trusted /run/serve.zone parent; the broker verifies its ancestry and provisions the 0700 leaf and 0600 socket. The native executable enters CNI mode through --cni or the installed basename pallet-cni. The client checks path ownership, socket identity and the kernel-reported server UID before sending one bounded frame. It never stores journals, allocates addresses or opens sandbox namespaces.

Transport cancellation does not cancel an admitted attachment operation. Detached handlers retain broker capacity until their native effect and journal write finish. Node shutdown joins CRI, including synchronous rollback DEL, before closing the broker; the broker then joins all admitted handlers before its borrowed owners can close. Replaced socket paths are never unlinked during cleanup.

The fixed profile uses CNI 1.1.0, the single pallet-workloads configuration and containerd's use_internal_loopback=true. VERSION reports protocol support. DEL can acknowledge exact recorded absence with empty stdout.

For an isolated assignment, the pending Run transaction durably binds its assignment, operation, boot and fixed CNI configuration. ADD opens the exact containerd sandbox namespace and proves it is distinct from the host and router namespaces. Containerd 2.3 requires a real eth0 with an address even for an isolated pod, so the native owner creates a namespace-local Linux dummy device with 192.0.2.1/32. It has no peer, uplink, default route or host grant; IPv6 address generation is disabled and every address, neighbour and route in both families is checked. The kernel's own ff00::/8 local-table multicast route is allowed only when it is bound to this peerless dummy, with no IPv6 address, gateway, RA or unicast route. The first ADD commits the exact container ID, namespace identity and double-proven initial absence in SmartData before NEWLINK. The completed ADD then records the dummy index, deterministic MAC and double native proof. If a process stops between NEWLINK, alias, address and UP, replay may remove only the exact DOWN partial dummy under that durable intent, prove its absence twice and recreate it. Ambiguous or foreign state remains quarantined. CHECK requires a completed journal and reproves its exact endpoint. The CNI result reports the real dummy and address, with no routes or DNS grants. DEL verifies and removes that exact device before retiring the journal; replayed DEL is idempotent. No workload lease, veth or activation grant is made for this mode. Native CRI evidence must report only 192.0.2.1 for an isolated running sandbox. Host-side TCP and HTTP readiness probes are refused before effects because they would target host services.

ADD never blocks on the activation barrier. Without it the attachment is still admitted, journaled and reported DOWN with an activation error, and the runtime retries; with it the held generation's policies are amended to carry the pair when they do not yet, and the pair is raised and proven in the same step. A refused amendment or raise leaves the attachment admitted, journaled and DOWN and names its step on the pass (amend.<code>:<site>, raise.<code>:<site>); it never fails the ADD. A successful ADD returns a CNI result carrying the interface, its address, the node resolver and the one route the raise installed: 0.0.0.0/0 via the router side of the pair (static, priority 100). Only state this owner installed is reported, and the advertised route is proven equal to the route found in the pod namespace. CHECK reports the raised attachment while the barrier still holds and demotes to the same DOWN activation error the moment it is lost, without touching kernel state. A repeated ADD on an already raised attachment is idempotent and never destroys it. STATUS reports authenticated broker and network owner availability to containerd, independently of any attached activation barrier; it cannot grant a workload network. An attached ADD or CHECK without the exact live barrier still reports DOWN, while an isolated ADD uses only its own pending Run preparation. GC cannot infer cleanup authority from a runtime-provided attachment list and returns an error until an admitted removal or historical proof can be resolved.

Application owner and node lifecycle

PalletNetworkApplicationOwner holds one private handoff-set owner for the node's router namespace. Each bounded pass keeps journal inspection, completion of a retained current-epoch intent, current source selection, epoch qualification, admission, native effects and completion inside one identity-runtime callback. It derives complete membership from the authenticated store; RPC callers cannot provide a prepared target, native receipt or absence proof. A stale source or identity rejects new admission. Completion of an already admitted effect survives credential changes, controller disconnect and quiescence.

The private PalletNetworkOwner composes this application owner with the guard and protection outbox. It restores the existing offline guard before starting the application or controller. While connected, the node reconciles the network with the actual physical session's private identity guard before scanning assignments whenever a network pass is due: the first scan, every scan after new network input from the controller (a new connection, an admitted signing authority, projection or DNS lease renewal), after a rejected pass, as soon as a managed VPN tunnel stopped on its own, and otherwise at most networkRefreshMs apart (default 30 s; the assignment scan itself runs every pollIntervalMs, default 1 s). A full pass re-proves the guard with a fresh engine process and re-reads and re-verifies the application and its signed source, so it does not run on every scan. The standing DNS refresh runs every 30 s for the same reason; every change DNS depends on (an application pass, a workload ADD or DEL) refreshes it at once, and the resolver child enforces its lease's BOOTTIME expiry itself. The protection delivery loop decides that nothing is pending from its lane row alone. Without current authority it waits; there is no offline network admission path. The status snapshot reports the application reference and admission rejection without exposing source records. Native ownership loss or an unrecordable native result fails the node with ownership retained for cleanup.

Shutdown stops admission, joins any pending begin/effect/finish callback, then joins the native handoff owner before closing the namespace and database. An unconfirmed native join keeps those resources owned; failed cleanup cannot report a clean stop.

A normal stop (SIGTERM or SIGINT, or stdin EOF for runtime-serve) keeps an ACTIVE generation UP in the journal for the next process to fence, and lowers only its managed VPN tunnel. The native refuses to withdraw a generation that still holds a device, and a device owner joined without a begun deletion takes its device away behind the native, which retires on that unarmed deletion; either way the handoff set's release stays unconfirmed. So the orchestrator's close refuses every new command and joins the ones it admitted before; then it lowers a held tunnel exactly as a withdrawal lowers it — forwarding stopped, the lowered packet pair applied, the deletion journaled and begun, the device owner released, the deletion finished — and only then joins the device owner. The native's closeHandoffSet then withdraws the generation inside its own router namespace and releases the set, and the process ends with stopped and exit code 0. The journal keeps the generation UP with its tunnel lane settled on the deletion, and the kernel keeps the host table at the lowered pair: exactly what the next process's epoch fence proves ended and releases (the retained-generation paragraph under "Standalone control build"). The application owner joins every workload ADD and DEL before that close, since the lowering runs outside the lane an ADD's amendment and raise run on; so no other journal or packet work interleaves with the lowering. A stop over a generation with no tunnel lowers nothing. A lowering that fails is an owner failure: the stop still joins every owner, fails with the failed step as its cause and ends owner_failed, and the successor fences the generation as after a crash.

Without the activation option all handoffs remain DOWN. With it the application owner opens the physical DHCP uplink observation before the orchestrator and raises one ACTIVE generation over it on the first pass that has an application; a generation that is UP for the same application over the held observation is kept on later passes, an observation that no longer proves itself withdraws the generation bound to it and is reopened on the next pass, and a failed activation is fenced and named on the pass (activationFailure). The same label is written once to stderr, which Spark carries into the node journal, as Pallet network activation refused: <label>; a packet engine refusal adds the engine request and hop (preparePolicy, host or router) and the engine's code and bounded refusal text. A refusal repeated on every pass is written once, and a refusal that returns after activation held is written again. Without a managed VPN the generation settles at UP and the barrier deliberately does not hold, so attached workload links stay DOWN — except on a single host bound to a local controller, whose barrier holds as described above whether or not a managed VPN is configured. Lease release and quarantine remain with Cloudly's allocation owner; Pallet resolves only an exact assignment-bound projected lease and cannot create or release allocation authority.

Durable protection and withdrawal reports

Every connected network pass freshly recovers the actual allocation-pool guard, completes the dormant application journal, then attempts a positive receipt. Only a real guard owner that has reconciled, inspected enforcement, detached its PERSIST policy and joined its native child can supply that private capability. The returned bootstrap JSON and a durable guard reference alone cannot mint it. The capability expires when the owner closes; composition keeps it through the outbox transaction. RPC never accepts a native proof or guard selector.

Fresh positive minting fences the settled current guard and application lanes, current signed projection and signing key, retained allocation ledger and actual physical reporter/identity in one SmartData transaction. Guard and application must agree on source, boot, host namespace and protected authority. The report's nativeBarrier derives from that exact completed guard intent. It establishes allocation-pool denial, without claiming workload connectivity, DNS readiness, packet drain or lease release. Every handoff remains DOWN.

Every positive report also states the node's own uplink address — the address a route to its published ports uses, read from the retained uplink observation — on the request envelope (uplinkAddress, @serve.zone/interfaces 31.4+). The address is a fact of the pass that minted the receipt: a changed or absent address is a new receipt generation, so the pass after a rebind reports it absent (the stale observation is closed) and the next pass, over a fresh observation, reports the new one. Nothing defaults or infers it, and a withdrawal states none.

Beside it the report states hostAddresses (@serve.zone/interfaces 32.23.0): every address the same observation binds on the uplink — the lease address and the permanent /32 host addresses beside it — sorted as strings. Cloudly names a plan platform endpoint in this node's hostPlatformEndpointIds when its address is one of them. They are read together with uplinkAddress from one observation and are part of the receipt's identity in the same way: another set is a new receipt, and a report without an uplink address states none. Only the uplink's addresses are stated, because the packet compiler serves a host-local platform endpoint only on an address the uplink binding carries. An outbox entry a 33.0.0 node wrote carries none and is read as stating none.

The SmartData outbox retains immutable receipt revisions, original reporter bindings and completed-application references. Positive reads verify the historical completed guard too. One pending receipt supplies backpressure; a newer application cannot replace it. Historical replay uses those immutable proofs without requiring old guard/application heads to remain current. Exact ACKs fence the current physical connection and identity. Reconnect sends the original receipt body, then a later fresh native pass may queue a current-session successor. Replaying historical positive evidence does not re-establish current-session eligibility.

Quiescence stops admission and joins the native pass, drains and ACKs any pending predecessor, then withdraws an acknowledged positive with its exact null successor. The null retains its predecessor's historical application, authority and boot, so withdrawal works during key/projection gaps. Its reporter is the current physical session. The exact null ACK must be durably recorded before clean node shutdown. A failed ACK leaves the immutable pending history and enforced PERSIST policy intact, joins the child owners, and fails the process shutdown result. Physical disconnect alone does not remove Cloudly's durable allocation eligibility.

Offline allocation-pool guard journal

PalletNetworkGuardOwner composes the private identity database with published Smartnftables 2.1. commission() explicitly creates the first guard intent; recover() requires an existing journal. The caller supplies a trusted native binary path and must join this owner before closing the database. The node runtime composes this private owner with application completion and positive reporting; Spark installs the separately commissioned boot dependency.

The fixed receiver-owned trust record supplies the node scope, checked against the active local identity. Bootstrap verifies the latest admitted projection with its retained signer, including a crash between key rotation and the next signed projection. That historical material supplies denials only; DNS and new projection admission still require the current key. The policy denies exactly the declared allocation pools, without treating management LANs or resolvers as allocation pools. Its only exceptions are the exact host grants of the source's signed onebox projection. They are derived from that verified projection whenever a target is proposed and proved again from the retained source on every read of an intent, so a stored row never widens the guard by itself. Every policy also restricts the CRI stream server of Pallet's own containerd, 127.0.0.1:10010, to uid 0 (localTcpPortOwners, Smartnftables 4.0.0). A pass refuses unless the policy it applied carries that owner; a guard committed before the owner existed gains it through one appended transition of the same source.

SmartData persists each complete native identity and original previous-Applied/null to prepared-target transition before reconciliation. Immutable result records bind the full native receipt to that intent. One pending lane prevents competing targets. Same-boot recovery repeats the original transition, including after a lost reply or completed result. A different actual boot records the old applied/pending heads in an immutable epoch and restores the prior denial with a fresh process instance and previous: null. A changed namespace in the same boot is rejected.

Recovery restores the committed denial before applying a newer inert projection. Every previously guarded pool must remain with the exact same id, purpose and prefix. Native preparation must reproduce the durable target, and exact enforced inspection and verified detachment must complete before success. An ambiguous failure uses closeRetaining() to join the child without deleting its policy or inventing an Applied receipt. No journal result alone grants allocation eligibility.

test/native/qualify-guard-boot.py runs real NoSQLDB and nftables in four isolated offline root boots under Node and compiled Deno. It covers same-boot completed and pending recovery, lost/undelivered native calls, new-boot epochs, signing-key gaps, newer staged pool expansion, and actual pool denial with unrelated traffic controls. On both boots a root TCP connection to 127.0.0.1:10010 succeeds and the same connection as uid 65534 is refused, before and after the new-boot recovery. It also exercises PalletGuardProcess and reopens the same database after its oneshot completes. Supplying --control-directory dist_control/linux-amd64-<digest> also runs the built production pallet-control guard-recover with null stdin and verifies its completed events, database release and continued pool denial. Production still requires the commissioned Spark boot dependency, and live activation the qualification of the complete workload networking, DNS and VPN path on a real host.

Host kernel requirements

The packet engine (Smartnftables 4.0) requires Linux 6.9 or newer, with the nftables filter and NAT modules (nf_tables, nft_nat, nft_chain_nat, nf_nat and the reject modules the port owner rule uses) available. A kernel below that floor, or one lacking a feature a policy needs, is refused by the engine as UNSUPPORTED_KERNEL; Pallet names that refusal instead of hiding it. The guard owner and PalletGuardProcess fail with code unsupported_kernel, and guard-recover, guard-commission and runtime-serve end with failed reason unsupported_kernel on their process protocol, exit code 1. Every other engine failure stays owner_failed. A host must boot a supported kernel before Spark commissions the guard.

PalletGuardProcess verifies the guard executable against trusted bundle metadata, opens only existing enrolled storage and runs one offline recovery or commission. It joins the native owner before stopping NoSQLDB. A native cleanup failure keeps the database lease owned for an explicit close() retry. Cancellation joins admitted work and fails the oneshot; success follows verified policy retention, native child exit and database shutdown.

The compiled control accepts guard-recover for ordinary boot and guard-commission for explicit first commissioning. Both accept exactly one mode argument and no runtime configuration. Null stdin is valid for these oneshots; unexpected input or SIGTERM/SIGINT fails without readiness. Successful completion emits ready and stopped under pallet.guard.process after cleanup. These events provide no positive allocation or workload protection receipt. The installer must acquire its first authenticated projection and commission before enabling recovery as a boot prerequisite.

test/native/qualify-protection-boot.py separately qualifies the complete private network owner in four offline Linux6.18.35 boots, under Node and compiled Deno. It verifies actual nftables denial with independent UDP controls, real dormant handoffs, completed application/guard agreement, historical positive replay after reboot, a fresh new-boot positive, null withdrawal through signing-key advancement, and PalletNodeProcess cleanup that reopens the same database. The packet client uses the pinned Node executable for both runtimes so a Deno UDP compatibility error cannot be interpreted as enforcement. No kernel, native-owner or database result is stubbed in these guests. They have no network device or host mount. Supplying --control-directory dist_control/linux-amd64-<digest> also runs the built production pallet-control runtime-serve on each runtime's second boot. It verifies readiness, EOF-driven shutdown, clean child exit, database reopening and continued pool denial after the entire node owner has stopped.

Initial projection acquisition

pallet-control network-acquire opens only existing enrolled storage and a projection-only controller connection. It waits for signing authority and a projection delivered over that actual physical session; a retained older projection cannot satisfy acquisition. It rejects assignments and does not fetch registry credentials or deliver terminal, observation or protection reports. After admission stops, it joins every controller write, verifies the delivered projection under the final current key, and closes controller/storage before ready and stopped under pallet.network.acquisition. Null stdin is valid; input, signals or the bounded operation timeout fail the oneshot after joining its owners.

pallet-control network-acquire-local is the same acquisition for a node driven by a controller on its own host (Local controller). It serves the local socket instead of connecting to a relay, and it additionally admits the first bind of a node no controller has bound yet; a node with any Cloudly enrollment state refuses to start it. The store must already be provisioned.

pallet-control store-provision provisions it: a null-stdin oneshot on the pallet.store.provision protocol that creates the data root and the database, prepares every collection, runs the store's migrations and releases the database again. It writes no enrollment, identity, binding or authority record and serves no socket, so the node stays unbound until network-acquire-local admits its first bind. An existing store is accepted as it is; a store that carries a Cloudly identity is refused (PalletStoreProvisionError:cloudly_enrolled on the owner_failed diagnostic line). enrollment-provision also creates the store and leaves no Cloudly state behind by itself, but it serves the Cloudly enrollment socket for its lifetime, where a prepare request writes that state; a node for a local controller is therefore provisioned with store-provision.

The installer runs acquisition while management networking is available and before first guard-commission and boot-unit installation. Existing guard recovery remains offline and requires an existing journal; it never silently commissions a new one.

Previous-epoch host attachment audit

A fresh PalletHandoffSetOwner can call verifyPreviousAbsence({ journal, target, previous }) before it applies a current set. Supply the authenticated persisted application reference, exact old prepared target and its prior applied receipt, including an unfinished transition's complete retained membership. Native code recomputes both old identities and their transition relationship, requires the same boot and host namespace, and retains the shared host mutator lock. The journal reference is an opaque binding; native code does not authenticate it.

Two complete host inventory passes reject old host or misplaced router names and MACs, Pallet markers, retained host indices and transit address/route conflicts. Old router indices remain scoped to the old namespace. The fresh router must independently remain empty and DOWN. Any candidate or inventory loss rejects the audit; it performs no cleanup, repair, deletion or adoption. Persist the returned old/current namespace-bound evidence before permitting a new epoch's effects.

The evidence proves old host attachments absent under exclusive privileged ownership. It does not prove physical peer destruction, packet drain, allocation reuse or a current nativeBarrier. Kernel peer teardown can be asynchronous, and another privileged writer can invalidate negative observations. Full node recovery uses the application owner's journal and complete authenticated history; process death or a saved namespace identity cannot replace this audit. The isolated Node/Deno fixture covers surviving and renamed old attachments, reused host indices, unrelated host indices matching old router indices, transit conflicts, fresh-router state, old request tampering and wrapper lifecycle.

verifyPreviousBootFence(request) is the separate fresh-owner capability for a different kernel boot. It validates the exact historical target and prior receipt, requires different valid kernel boot UUIDs, and reads the current boot from the native owner. Two complete audits still reject current old identities, Pallet markers, transit conflicts and nonempty fresh-router state. Old interface indices are deliberately ignored: those numbers can legitimately belong to unrelated interfaces after reboot. The returned evidence binds both epochs and adds previousBootTerminated: true. It supplies no allocation release, packet drain or nativeBarrier; the dormant-only history and privileged ownership requirements still apply.

The offline test/native/qualify-handoff-boot.py fixture runs two boots each under Node and compiled Deno with independent disposable ext4 NoSQL roots. The actual application owner persists real signed authority and applied/pending history. A fixture-only interrupted completion retains the acknowledged DOWN target, then the fixture flushes the database then signals init through a FIFO while the native owner and DOWN pair remain live. Init forces poweroff. The next boot reloads that journal, qualifies reused host indices and conflict rejection, then the application owner persists the boot fence and realizes the same still-reserved lease DOWN. All application state uses SmartData/NoSQLDB; no journal sidecar or host network/disk attachment participates in the qualification. It also retains the exact unavailable report across poweroff, replays its original reporter binding, and queues the next boot's report only after acknowledging history.

Standalone control build

The control builder requires Deno 2.9.7 and the exact NoSQLDB 10.5.1 engine identities in binary/control-build.json. After the normal dependency install, build either Linux target or omit the target to build both:

pnpm run build:control linux-amd64
pnpm run build:control linux-arm64

@git.zone/tsdeno 1.14.0 compiles with the committed, frozen deno.lock under its managed runtime-only manifest, and per target leaves out the npm native binaries built for the other architecture or for macOS (their ELF/Mach-O headers prove them foreign); the control never loads them, because it runs its own sibling engine, guard and DNS executables. TsDeno refuses any Deno other than the pinned denoVersion, fails a pallet-control smaller than the target's minSize in binary/control-build.json (200 MiB; a compile that lost its npm payload still exits 0 with a far smaller binary), and on the host's own architecture runs it once without arguments in an empty environment, where it must print its invalid_arguments refusal and exit 1; the other architecture's smoke check is skipped. The build verifies the selected NoSQLDB engine's bytes and clean owning-build provenance. It also builds and verifies the Pallet static-musl executor against this exact source, version and architecture. The Smartnftables 4.0.1 guard is verified against its pinned bytes and clean provenance. Each completed dist_control/<target>-<digest> directory contains pallet-control, its fixed siblings pallet-smartdb, pallet-runtime, pallet-guard, pallet-dns and pallet-vpn, the pallet-containerd/ release directory (containerd, its shim, runc, the pause archive and their manifest), their native provenance records and a control-build.json recording artifact hashes, source state, compiler versions and runtime lock hash. The compiled control keeps UID 0 and its protected production data/socket paths; it derives only the sibling engine path from its own executable. The installer owns protected platform-parent directories. Packaging includes every native guard notice from the pinned upstream manifest (the native-notices/manifest.json that tsrust notices generates, with the digest of every file) under notices/smartnftables/ and rejects changed or missing material. The archive also includes the exact pallet-clock/ loader, programs, shared libraries and trust files, plus full clock source archives, patches, recipes and notices under sources/clock/ and notices/clock/. Consumers must verify the complete inventory.

binary/vpn-engine.json pins the published SmartVPN 2.4.1 musl executable, provenance, package license and native notice inventory for each architecture. The builder and packer verify the complete distributed VPN notice set under notices/smartvpn/, including the upstream Rust and Cargo license texts. A changed binary, symlink, missing notice or mismatched build is rejected. The bundled pallet-vpn is the executable runtime-serve verifies and runs as the managed VPN device owner (see "Foreground node process").

The Chrony source build additionally requires a Docker client/Buildx and a Linux Docker daemon able to execute the selected pinned Alpine architecture. Its build context and result archive use the daemon API; it requires no host workspace bind mounts. Compiler execution is offline, unprivileged and capability-free. The build checks pinned source, SDK and output hashes; workers receive only the finished programs and distribution materials and perform no package installation.

Every change to the dependencies in package.json moves the control graph, so it also moves the lock. Rebuild dist_ts, generate the lock with pnpm exec tsdeno install --entrypoint --lockfile-only --frozen=false --lock=deno.lock dist_ts/control/process-cli.js, and review the lock diff. Then carry noticeLockSha256, the control notices and their pinned asset hash in binary/control-build.json onto the new lock. test/test.controllock.node.ts runs the frozen resolution the control build uses, through scripts/control-lock.mjs, and fails while any of them lags. Ordinary control builds leave the lock frozen.

The lock is registry-neutral: its npm entries carry no tarball URL. Deno writes one only when the registry it resolved a package from is not npmjs, and from Deno 2.9.7 it refuses a lock whose tarball origin is neither the configured registry nor npmjs. The tracked .npmrc therefore sets registry=https://registry.npmjs.org/. Deno and pnpm both read it from the project root and prefer its registry over the one in ~/.npmrc, so a lock generated on a machine whose home configuration names a mirror still resolves from npmjs. Deno's NPM_CONFIG_REGISTRY environment variable overrides the project .npmrc; scripts/control-lock.mjs refuses to run while it names another registry, and refuses a lock that carries any tarball field.

The earlier Linux amd64 artifact with SmartDB 5.8.0 was qualified as root in an isolated QEMU guest with a disposable local ext4 disk and no network or host mounts. Checks cover the fixed production paths, engine digest and permission guards, explicit provisioning, durable enrollment replay after process restart, EOF/signals, native and parent SIGKILL, and stale-socket recovery. NoSQLDB rejects volatile filesystems; a tmpfs data root cannot substitute for supported local storage. This is process recovery qualification, not a power-loss or ARM64 runtime claim.

The NoSQLDB 10.5.1 engine passes the local native identity persistence and lifecycle suite with SmartData 11.14.2. Pallet now reaches it through the nosqldb family of @lossless.org/client 1.5.1, which continues SmartData 11.14.2 with the same persisted format. Both engine targets are the published static musl executables; the Deno control executable still targets GNU Linux. Both control targets compile and package with their verified owner provenance and complete upstream notices. The updated amd64 control bundle also passes the root guest checks above on an isolated local ext4 disk, including durable enrollment replay after native and parent process crashes. ARM64 runtime and power-loss qualification remain open.

test/native/qualify-node-process.py also qualifies a forward upgrade when given both --previous-control-directory and --previous-probe. It boots that earlier bundle first, retains a read-only copy of the stopped ext4 disk, then boots the new bundle twice against the working disk. Each boot verifies its exact binaries and probe; the final result compares the node identity and signed network projection and verifies the backup is unchanged. The NoSQLDB 9.0.0 to 10.2.0 amd64 run passes all three offline root boots, including retained workload leases, signature replay, namespace keeper loss and joined process shutdown. This qualifies existing current-format Pallet state; it does not qualify containerd, private DNS, physical power loss or foreign legacy database roots.

The guest stages the build's complete runtime inventory (executor, guard, resolver, clock loader and assets) and the nftables and veth modules, enrolls with a routing binding and commissions the allocation-pool guard before runtime-serve, which now owns the node's network; the guest provisions the root-only runtime directory on every boot, as the installer does. A runtime-serve whose namespace keeper is killed reports failed and exits: when its close cannot confirm the network release, the node process joins the clock, the namespace and the database anyway (PalletNodeRuntime.abandon) instead of holding them for a retry nobody in the exiting process makes, and the successor recovers the unconfirmed state by evidence, as after a crash.

The database package is @lossless.org/nosqldb and the controller uses its NoSqlDbServer API. Existing smartdb storage-directory and bundle field/file identities remain stable; they do not select the deprecated npm package. Changing the package does not convert an incompatible database format. NoSQLDB 10 reads existing version-9 current-format stores directly; once it writes version-10 records, older engines cannot read them. Retain a stopped pre-upgrade backup before activating the new engine. Native current-format admission still rejects foreign legacy and mixed roots before application startup.

Pallet admits the released 32.0 attachment-journal shape before any network owner starts. Terminal absence and an already durable pending removal retain their original bytes, digests and ACTIVE source or amendment history; the latter may finish with its released removal formula. The admission writes only the pallet_migrations completion row and is replayable after interruption. A released running or pending-ADD row has no durable removal transition that this version can complete without rewriting its history, so startup fails with a pallet-attachment-owner-recovery-* error and writes no completion row. Stop the upgrade and retain the prior Pallet owner to finish that lifecycle before trying again. There is no collection-reset or cleanup command in this recovery path.

Every node process owns its own router namespace, so a restarted process runs in a new epoch, and so does every process after a reboot. A generation the previous process retained — UP, or DOWN with a packet cleanup it never finished — is held by no native of the new epoch, and the native admits RecoverActive only for a generation of its own. The restarted process therefore fences it before its native adopts anything, with the audits that fence the application journal's own epoch, over the generation's application target and receipt: in the same boot, verifyPreviousHandoffAbsence proves the previous epoch's host attachments absent (the router namespace, and every pair, device and route under it, ended with that process); after a reboot, verifyPreviousHandoffBootFence proves the boot over. The proof is journaled as a DOWN of kind epoch (same-boot-host-absence or new-boot-host-fence), taken from the fencing native's epoch (PalletHandoffSetOwner.readEpochFence), and its packet cleanup releases what outlived the epoch: in the same boot the host table, through the running engine with the generation's retained host identity; the router table ended with its namespace, and after a reboot no table is left. The successor is raised on a pair keyed anew. A generation of the process's own epoch is still withdrawn in order while the process holds it — a tunnel transition left pending by a failed step is finished first — and recovered by RecoverActive otherwise.

Every ACTIVE generation records the packet engine that prepared it, the SHA-256 of the Smartnftables executable the node process verified before it started (packetEngine; generations journaled before 35.0.0 carry none, so their engine is unknown). The kernel keeps a generation's host table across a Pallet restart in the same boot, and the restarted process asks the running engine to release it once the generation is fenced. An engine whose compiled graph differs from the one that applied it cannot adopt it and answers Conflict; Smartnftables 3.0 changed the graph of every routerEgress table and of every scope with publications. When that Conflict comes over a table another or an unrecorded engine applied, the release fails by name (PalletPacketEngineChangedError, conflict:packet.retainedByPreviousEngine) and the network owner stays failed and fenced, because no retry of the running engine can change the answer. The remedy is a reboot, which takes the kernel tables with it and lets the new-boot proof fence the generation, or finishing the generation on the previous Pallet before upgrading again. An upgrade that changes the packet engine across a same-boot Pallet restart therefore needs a reboot whenever an ACTIVE generation is retained. A Conflict under the engine that applied the tables is not this refusal and fails as before.

Local build outputs are qualification assets; a dirty source marker is recorded explicitly and cannot identify a release. The tag-triggered workflow publishes clean-source control bundles as inputs for Spark integration. The bundle's runtime-serve activates the node network with the bundled pallet-vpn (see "Foreground node process"); running it still requires Spark's whole-bundle installation and supervision. The complete activated path — ACTIVE generation, managed VPN tunnel to a real cluster hub, lease renewal and workload traffic — is qualified in parts (the guests below and unit specs), not yet end to end on a real host. ARM64 runtime and workload execution are not yet qualified by this component release.

TLS trust of the control executable

pallet-control verifies every TLS peer it dials — the controller or Cloudly origin of the enrollment, runtime and network-acquire sessions — against the Mozilla root bundle compiled into Deno and then the operating system trust store. A node whose controller, relay or Cloudly certificate is issued by a private or internal certificate authority trusts it the way every other service on the host does: install the CA certificate into the OS store (update-ca-certificates, update-ca-trust). Pallet has no trust setting of its own and never disables verification.

A compiled Deno executable carries no store selection and otherwise trusts the Mozilla bundle alone, so runPalletControlCli sets DENO_TLS_CA_STORE=mozilla,system for every mode, after the argument check and before any mode starts; Deno reads the variable once, when the process opens its first TLS client. An operator who sets DENO_TLS_CA_STORE explicitly keeps that choice, provided it lists only mozilla and system; any other value, including an empty one, stops the process with one failed line of reason invalid_ca_store on the mode's protocol and exit code 1, rather than failing every later connection. Spark starts pallet-control with a cleared environment, so its processes always run with the default. The managed QUIC tunnel of pallet-vpn, a native SmartVPN executable, authenticates the hub by the public key its credential names rather than by a certificate authority, so no root store applies to it.

pnpm exec tstest test/test.tlstrust.node.ts --verbose --logfile --timeout 120

The test pins the default and the refusal, runs the CLI with an invalid store in three modes, and, under the pinned Deno, dials a real HTTPS and WebSocket server whose private CA is present only in the OS store (SSL_CERT_FILE): it is refused as UnknownIssuer without the selection or with an explicit mozilla, and trusted with it.

Regenerating third-party notices

binary/control-third-party-notices.txt names the frozen npm graph, the Deno GNU Linux runtime and the NoSQLDB engine. When the Deno pin moves, move every pin first — denoVersion in binary/control-build.json, the sealed release's denoVersion in .smartconfig.json and both Deno lines of .gitea/workflows/release.yml — and add the new version's release commit and the SHA-256 of its deno_src.tar.gz release asset (GitHub states it as the asset digest) to binary/control-deno-sources.json. Then, with that Deno and cargo-about 0.9.1 on PATH:

pnpm run notices:control            # rewrite the notices and their pins
pnpm run notices:control --check    # regenerate, compare, write nothing

scripts/notices-control.mjs refuses unless every Deno pin, the running Deno and the pinned source agree. It downloads the pinned source of the Deno the notices name and of the pinned Deno, checks both against their digests, and takes the runtime's crate inventory with cargo-about over cli/rt for the control targets (--locked, so Cargo fetches the locked crates). It needs tar and network access to GitHub and crates.io, and works in a private temporary directory it removes afterwards.

The regeneration carries the Deno runtime part and leaves everything else byte-identical. A registry crate the notices already name keeps its entry; Deno's workspace crates and every new crate are written from the crate's own legal files, the workspace crates with Deno's LICENSE.md. A new crate without a legal file is refused, because it needs a reviewed entry. The native V8, Chromium Rust and Rust standard library notices are carried only while the pinned Deno runs the V8 and the Rust toolchain they name; otherwise the tool refuses until they are refreshed. The v8 crate cites rusty_v8's license at the crate's own commit. The embedded JavaScript keeps its file list: each listed file's legal comments are read again at the new release, each excerpt is found again at its new lines, and a changed or new source file under ext/, runtime/js/ or libs/core/ whose legal comments are not listed is refused for review. The asset hash and noticeBuildIdentity.denoVersion in binary/control-build.json follow the written file. --input <file> carries another notices file instead of the committed one: from the Pallet 34.0.0 notices (Deno 2.9.4) the tool reproduces the committed Deno 2.9.7 notices byte for byte. test/test.controldenonotices.node.ts covers the pin refusals and a dry run over a small fixture without cargo-about.

Control bundle packaging

@git.zone/tspack owns archive assembly, file hashes, executable modes, sealed manifests and complete archive verification. Pallet's adapter selects explicit control builds and checks their source, version, compiler, runtime lock, fixed paths and both native owners' provenance before passing inputs to TsPack. The selected license notices are pinned in binary/control-build.json to the runtime lock, Deno version and NoSQLDB owner build. The pinned native notice inventory binds the Cargo/compiler inputs and every reviewed notice; the archive preserves the native notice index, complete inventory and CRI provenance.

After sealing, the adapter extracts every bundle again from the sealed set and runs the verifier of @serve.zone/pallet-bundle over it with the build's own source identity, so a bundle that disagrees with what that package verifies for its consumers is never produced. Its paths and control record keys come from the same package (ts_bundle/inventory.ts). pack:control first compiles that package (tsbuild custom ts_bundle to dist_ts_bundle/, output on stderr) from the checked-out source, so it runs on a clean checkout and never loads a stale verifier; release:control loads the adapter only after build:control has compiled it.

# Use the exact project-relative directories printed by build:control.
pnpm run pack:control dist_control/linux-amd64-DIGESTPREFIX dist_control/linux-arm64-DIGESTPREFIX

# A release requires both architectures built from the clean vVERSION tag.
pnpm run pack:control --release dist_control/linux-amd64-DIGESTPREFIX dist_control/linux-arm64-DIGESTPREFIX

Ordinary packaging produces an explicitly unpublishable qualification set. A release produces the sealed set under dist_control_release/ and an inputs-VERSION-COMMITPREFIX.json description that records its exact packaging configuration and manifest digest. The output JSON identifies both paths.

A publishing workflow must retain the entire dist_control_release/ directory, including that input description and sealed set, before publishing any asset. Restore those original files on retry, then run:

pnpm run pack:control --release --reuse

Reuse verifies the retained source, configuration, manifest and every archive. It works without the original compiler outputs and rejects missing, corrupt or changed retained inputs. It never recompiles a control bundle: the only thing it compiles is the TypeScript of the bundle verifier above. Fresh Deno compilation may produce different bytes, so rebuilding cannot substitute for retaining a published set. TsPack and this adapter do not provide remote CI artifact storage or publication.

GitZone 6.6.2 or later manages .gitea/workflows/release.yml and the scripts/gitzone-*.mjs release scripts through the committed tspackRelease asset configuration. The tag-triggered workflow invokes the configured scripts/prepare-control-release.sh command in its disposable job container. It requires a root Linux Gitea Actions environment, installs Pallet's clang and musl-tools compiler prerequisites through apt, then executes the existing scripts/release-control.mjs exact-tag adapter to build and package its two explicit output directories. Compiler installation writes only to stderr so the adapter retains its single JSON result on stdout. The runner must be Gitea Runner 3.3.2 or later with its cache/results service reachable from job containers and runner.patch_actions enabled. The pinned stock artifact actions use the v4 protocol required by Gitea's REST retention inventory. It retains the complete output root in Gitea Actions for 90 days, downloads that retained copy, and verifies component reuse and every archive before creating a release draft. The publisher checks existing attachment bytes, adds only missing files, and publishes after complete readback. It never replaces release assets.

After a CI interruption, rerun the original Gitea run so it restores the original bytes. Missing or expired retention and conflicting remote attachments stop publication. The generic release receipt gitzone-release.json and Pallet's input description must remain beside the sealed set in the retained artifact. Use gitzone format --only assets --write --yes to update these managed scripts from a published GitZone version, then commit the result before releasing.

Build and test from source

This development package is marked private and is never published. The npm release target publishes one module of this repository, ts_bundle/, as @serve.zone/pallet-bundle with Pallet's release version (see its own ts_bundle/readme.md). Its tspublish.json takes the @git.zone/tspack and @push.rocks/smartdaemon ranges from the root devDependencies (devDependencyVersions), so neither enters the root dependencies that deno.lock locks for the control build, and @serve.zone/interfaces from the root dependencies. From the repository:

pnpm install
pnpm test
pnpm build

Tests build a native debug binary and exercise it against a disposable Unix-socket HTTP/2 fake CRI server. They do not connect to Docker or a production daemon. The native build uses Rust 1.95.0, locked Cargo dependencies and pinned vendored protoc to produce static musl Linux binaries through @git.zone/tsrust, named pallet_linux_amd64_musl and pallet_linux_arm64_musl. pnpm build builds only the host architecture's executor; pnpm run build:control builds both with tsrust --configured-targets, so a control build and the release always carry the executors of both targets built from the same source. ARM64 cross-compilation uses aarch64-linux-gnu-gcc as the linker driver with Rust's self-contained musl target libraries. The ring TLS provider also requires musl-gcc for amd64 C/assembly and Clang for its supported freestanding arm64-musl C build. These compiler choices are declared in rust/.cargo/config.toml; provision them on developer and release builders before running the build. Other packaged platforms are rejected explicitly.

After pnpm build, run PALLET_TEST_PACKAGED=1 pnpm exec tstest test/ --verbose --logfile --timeout 60 to exercise the host architecture's packaged executable and its default lookup instead of the native debug executable. Cross-compiling an ARM64 artifact does not establish ARM64 runtime qualification. For explicit emulator qualification, the same suite accepts an absolute PALLET_TEST_BINARY_PATH pointing to a test-only emulator launcher. Do not combine it with PALLET_TEST_PACKAGED; an emulated run does not qualify real node hardware.

The third-party notice index points to the native notices tsrust notices generates in native-notices/ for the locked Cargo graph, the Rust toolchain and the static runtime, with full license texts and the reviewed once_cell licenses and source attributions of ring under native-notices/crates/ring-0.17.14/extra/, and lists the material kept beside them. npm publication and distribution of the standalone Rust runtime-probe binaries remain disabled. Gitea releases distribute the control bundles described above, whose runtime-serve activates the node network as described in "Foreground node process".

Read-only probe

After a build, from the repository:

import { Pallet } from './dist_ts/index.js';

const pallet = new Pallet({
  socketPath: '/run/pallet/containerd/containerd.sock',
});
const evidence = await pallet.probeRuntime({ timeoutMs: 5000 });
console.log(evidence);
// {
//   runtimeName: 'containerd',
//   runtimeVersion: '<actual daemon version>',
//   runtimeApiVersion: 'v1',
//   runtimeReady: true | false,
//   networkReady: true | false,
// }

The socket must already exist and belong to a separately configured, authorized containerd instance. Pallet never creates one, chooses a default socket, searches PATH for its binary, or falls back to Docker. An explicit absolute binaryPath can select a verified installed executable or the development debug binary. A missing or nonexecutable selected binary fails with SmartRust 2's RustBinaryLocatorError (ERR_RUST_BINARY_EXPLICIT_PATH_INVALID). Selection never changes executable permissions or falls back to another packaged binary or a stale GNU/native build. After explicitly provisioning the selected executable, the same Pallet instance can retry.

Each probe owns a separate native child. Only one probe per Pallet instance is admitted at a time. The method confirms that child's exit before returning, including on failure. If cleanup cannot be confirmed, the method rejects and retains ownership of the child. Further probes are blocked until await pallet.close() successfully retries cleanup. An optional AbortSignal cancels the local request and terminates the owned child; it never stops containerd itself.

The native deadline covers connection and both RPCs (50–30000 ms, default 5000). The bridge readiness handshake, after executable discovery, is bounded to 3 seconds, with 1 second for graceful termination before forced child shutdown. Executable filesystem discovery in the shared bridge is not currently timed or abortable; this is not an end-to-end startup deadline. IPC and decoded gRPC responses are limited to 16 KiB. The probe accepts only containerd with CRI API v1.

RuntimeReady and NetworkReady must each occur exactly once; a missing or duplicate condition is an error, while explicit false remains false. Runtime diagnostic messages and verbose configuration are never returned. Runtime reported readiness is not workload health, version support qualification, network reachability, storage fencing, or permission to perform a migration.

Native failures carry a RustBridgeRequestError.responseErrorCode of INVALID_INPUT, SOCKET_UNAVAILABLE, RUNTIME_UNAVAILABLE, DEADLINE_EXCEEDED, UNSUPPORTED_RUNTIME, or INVALID_EVIDENCE. Bridge transport, cancellation, and startup failures remain distinct errors.

Isolation and qualification

Known Docker-private path components and aliases into them are rejected as accident prevention. This is not a security boundary against a hostile local user who can replace socket paths. Run only with the local permissions needed to inspect the intended daemon; containerd's socket is root-equivalent.

Fake-server tests establish protocol and lifecycle behavior only. Before runtime adoption, qualify a dedicated containerd 2.3 instance with explicitly owned root, state, socket and configuration paths, then verify actual container, network, storage and recovery behavior. No production cutover is implied.

Every QEMU qualification guest under test/native/ runs one bundled ES module. scripts/guest-bundle.mjs builds it from the guest's own TypeScript source with esbuild — the whole graph inlined, the NodeNext .js specifiers resolved to the .ts sources in this checkout, bare Node builtins rewritten to their node: form (which is what the deno compile --no-config --node-modules-dir=none half of each guest needs) and a createRequire banner for the CommonJS dependencies:

node scripts/guest-bundle.mjs --entry test/native/namespace.ts --out .nogit/debug/guest/namespace.mjs

It prints the bundle's path, size and SHA-256, which is what a qualification run records beside its other inputs. The runners take that file as --bundle (and qualify-active-registry.py as --wrapper-bundle / --tunnel-bundle).

qualify-cni-up.py --scenario single-host runs test/native/singlehost.ts, a node bound to a local Onebox controller, in the cni-up guest with no substituted seam. network-acquire-local serves the contract's socket (a root-owned 0600 socket in a root-only 0700 directory; a peer running as another user gets EACCES); the first bind persists the binding, an exact replay answers the same credential, a different binding refuses, a replayed bind's session supersedes the earlier connection's, which can then no longer apply anything, and the signed onebox projection is acquired over the current session. The node's runtime then raises one generation without a managed VPN: the real barrier reaches forwarding on the single-host fact with no tunnel command, execution admission is accepted on that evidence, two workloads attached through the real CNI path are raised under the generation, TCP and UDP flow between them with the client's own address preserved, and each resolves the other's name to its lease address through the node-local resolver. A third workload, on a network only the first shares, answers the first and is dark from the second. The router namespace reads IPv4 forwarding off in all and default before any generation exists, although the guest host forwards. From the host namespace it probes the granted TCP and UDP tuples, another port, another protocol, another lease, another source address and a granted lease that is not attached. The Cloudly counterpart — a generation without a tunnel receipt settles at up and its barrier does not hold — stays proven by the default cloudly scenario. With the generation's pool route in place the granted TCP and UDP tuples reach the workload from the transit host address, and another port, another protocol, another lease and another source address stay dark; a granted lease that is not attached stays dark as well, stopped by the host-transit table before the router receives a packet (both hops carry attached leases only). That last check captures every frame on both ends of the transit link while it probes (the guest stages tcpdump, without promiscuous mode, and parses its pcap) and counts only IPv4 packets to the unattached lease: the link also carries the ARP exchange both ends run on their own schedule — the router re-probes its entry for the host about 5 s after the granted flows — which a frame counter cannot tell from a leak. The same captures must see a granted flow on both ends, drop nothing and account for every frame the link counters counted; with the host-transit policy deliberately given the unattached lease's grant, the probe's SYN reaches the router and the check fails. A successor projection then withdraws every host selection under the held generation: the pass withdraws that generation before it completes the changed dormant set, raises the next generation on a re-keyed packet pair (forwarding, no host route planned), and the formerly granted tuple goes dark. Probers dial every 100 ms through the change: the allowed pair between the first two workloads, and six tuples the policies deny — the pair that shares no network over TCP and UDP, and workload to the host's transit and uplink addresses over TCP and UDP, each with a live listener behind it. East-west traffic stops for the withdrawal window, because the withdrawn generation restores a router that forwards nothing, and resumes when the successor is UP (about 27 s of a 45 s change); every denied tuple is dialled inside that window and never answers before, during or after it, and no listener logs its source. Without the forwarding pin the same guest measures the unfiltered router: the pair that shares no network answered 61 of 201 dials during the change. All of it passes on 6.18.35-0-virt. The packet hub guest (packet.ts) plans the same pool route and proves it present only while its generation is UP, gone after a withdrawal, and gone after a fresh owner's recovery of a lost native.

qualify-cni.py --scenarios containerd-serve --control-directory dist_control/linux-amd64-… runs Pallet's own containerd from a build:control directory the way its service unit does, with a foreign /etc/containerd/conf.d drop-in planted in the guest. It proves the generated configuration governs (the drop-in's stream port stays closed, 127.0.0.1:10010 serves), the bundled sandbox image is imported and unpacked on overlayfs, a workload raised by the production execution owner through the real CNI path runs under /pallet with runc state under /run/pallet/runc, SIGTERM stops containerd and leaves the workload running, the next containerd-serve reattaches the exact sandbox and container, and killing containerd fails its supervisor (owner_failed) while the workload keeps running. The image collector keeps the running workload's image and the bundled sandbox image, and the proven removal frees the workload's image while the sandbox image stays.

qualify-cni.py --scenarios owner-storage --control-directory dist_control/linux-amd64-… runs one local storage claim under the same bundled containerd and runc, on node and Deno. The claim materialises its volume root-private with the claim's ownership; containerd echoes exactly one private read-write bind of it and the bundled runc's workload carries it at the claimed path; the workload writes as its own user; a second run of the service and a purge are refused while the first container lives (storage.claim.held, storage.purge.held); a restarted owner rejoins the live container against the grant's mount list; the second run finds the first one's bytes after the first is removed; and the purge deletes the volume once both are gone.

qualify-active-registry.py --paths-node-binary … --paths-bundle … (with the packet and SmartVPN binaries) runs test/native/singlehostpaths.ts: the three single-host paths of a Cloudly node whose platform services run on its own host. The guest adds the hub's and the relay's addresses to the uplink as permanent /32s before the uplink is observed, attaches and raises three workloads that share no private network, and composes one projection with the production composer twice: without hostPlatformEndpointIds and workloadIngress, and with them. Under the first pair a workload's dial of the relay it selects, the router namespace's dial of the hub and the ingress workload's TCP and UDP dials of its target are all dark; the composed pair replaces it one revision up under the same UP generation and all four are delivered — the relay and the hub see the transit source address with a port of the leased range, the target sees the ingress workload's own address — and the real SmartVPN hub on the host authenticates the router namespace's real managed QUIC client. Another workload to the relay, the Corestore port the workload did not select, an undeclared hub port, the host opening toward the router, the target opening back toward the ingress workload, another workload to the target and another target port stay dark under both pairs, each against a live listener that logs no peer.

The pinned upstream CRI protocol and the containerd Transfer and Streaming protocols, with their Apache-2.0 attributions, are under rust/proto/.

This repository contains open-source code licensed under the MIT License. A copy of the license can be found in the repository license file. The vendored Kubernetes protocol is separately licensed under Apache-2.0.

Please note: The MIT License does not grant permission to use the trade names, trademarks, service marks, or product names of the project, except as required for reasonable and customary use in describing the origin of the work and reproducing the content of the NOTICE file.

Trademarks

This project is owned and maintained by Task Venture Capital GmbH. The names and logos associated with Task Venture Capital GmbH and any related products or services are trademarks of Task Venture Capital GmbH or third parties, and are not included within the scope of the MIT license granted herein.

Use of these trademarks must comply with Task Venture Capital GmbH's Trademark Guidelines or the guidelines of the respective third-party owners, and any usage must be approved in writing. Third-party trademarks used herein are the property of their respective owners and used only in a descriptive manner, e.g. for an implementation of an API or similar.

Company Information

Task Venture Capital GmbH
Registered at District Court Bremen HRB 35230 HB, Germany

For any legal inquiries or further information, please contact us via email at hello@task.vc.

By using this repository, you acknowledge that you have read this section, agree to comply with its terms, and understand that the licensing of the code does not imply endorsement by Task Venture Capital GmbH of any derivative works.

S
Description
In-development node-local containerd execution foundation for Cloudly and Onebox; currently a read-only native CRI probe.
Readme
71 MiB
v35.4.0
Latest
2026-09-30 21:11:50 +00:00
Languages
TypeScript 67.3%
Rust 19.2%
HTML 7.4%
JavaScript 2.7%
Python 2.7%
Other 0.6%