@serve.zone/pallet
Pallet is the in-development node-local containerd execution component for serve.zone. Its public API provides a bounded, read-only native CRI v1 probe. Backend-private identity, enrollment and stateless execution mechanisms are under development; production assignment delivery and reconciliation remain unwired.
Issue Reporting and Security
For reporting bugs, issues, or security vulnerabilities, please visit community.foss.global/. This is the central community hub for all issue reporting. Developers who sign and comply with our contribution agreement and go through identification can also get a code.foss.global/ account to submit Pull Requests directly.
Development status and ownership
Cloudly will own cross-node desired state and placement. Pallet will own local execution, including the outbound Cloudly connection on workers. Onebox will control its own local workloads through authenticated Pallet IPC, without depending on Cloudly to bootstrap itself. Spark remains the host installer and supervisor. Coreflow's required behavior must be ported before its legacy Swarm runtime can be retired.
Those orchestration and authorization capabilities are not implemented by the public probe. It exposes no public listener and persists no application state. Do not replace a production runtime with this development foundation.
No Pallet name carries a version token: collections, singleton row ids,
record-id hash domains, the native structs and their TypeScript twins, the
dormant link aliases (pallet-handoff: / pallet-workload:), the host and
workload evidence kinds and the process protocol names are named by what they
hold. No Pallet shape carries a shape-version field either — not the native
structs, not the identity, application, guard, attachment, ACTIVE, barrier,
tunnel or DNS records, and not the build manifests this repository owns. Every
one of those key sets is exact, so a value that still states schemaVersion is
refused as an unknown key rather than read as an older shape. What contract a
peer speaks is stated once, by the handshake below, and what a persisted shape
means is stated by the installed build. No released Pallet version was ever
deployed, so nothing is migrated and nothing is aliased — a store written by an older build fails its assertion on
every read, a dormant link an older build left behind carries a foreign alias,
interface name and MAC address because the seed domain changed with the name, so
it is never recognized as this build's pair, and a node's stores under
/var/lib/serve.zone/pallet are reset with the install.
Private CRI execution
ts/runtime/classes.executiondriver.ts owns a bounded native CRI client through
the separate --executor-management process mode. It is absent from the public
package facade and enrollment IPC. The eventual assignment owner must authenticate
the controller, validate immutable material, durably admit the assignment and
serialize its effects before invoking this mechanism.
The driver supports inspect, run, stop and explicit runtime removal against a
dedicated containerd 2.3 release and CRI v1 socket. Run supplies a digest-pinned image,
its platform image-config digest, exact argv, environment, working directory,
CPU/memory policy, user/group policy and root filesystem write policy. Explicit
cpuMillis: null leaves CPU quota unlimited. Explicit null for both runAsUser
and runAsGroup selects the pinned image's User through containerd's rootfs
resolution; a numeric pair overrides it. Missing fields and partial pairs are
invalid. Containers
use private namespaces, runtime-default seccomp and no-new-privileges, without
privileged mode or host networking; the only host paths a container ever receives
are the read-only per-file binds of its own staged secret mount and one private
bind per node-local volume it was granted ("Local storage claims"), which the
executor derives from its fixed volume root and the volume's key — no host path
crosses the IPC. Every sandbox names the absolute cgroup parent /pallet, so
workload cgroups sit outside any service unit's cgroup (see Pallet's own
containerd). Network policy and application readiness
remain separate owners. Secret material delivery is this node's own, through the
staged mount.
Every runtime object carries the immutable generation-one run digest and complete controller/node/assignment/replica/attempt ownership labels. Recovery requires one exact sandbox and container, including a full check for foreign sandbox children. Image identity uses the actual config digest and the requested repository digest; display tags are not identity. Running attempts can be recovered; exited or stopped attempts are never restarted. Stop preserves runtime objects, including containers created but never started. Removal requires confirmed stopped state and retains storage; images are left to the image collector below. Neither operation depends on a healthy image cache or network.
Input is captured before asynchronous work without invoking accessors. Native frames and CRI responses have fixed bounds; errors expose static codes rather than runtime messages or registry credentials. Cancellation, parent EOF and signals terminate the owned client. A timeout or client exit does not cancel a server-side CRI operation, prove rollback or fence a writer. The durable assignment owner must retain uncertain effects and resolve them before admitting replacement.
pnpm exec tstest test/test.execution.node.ts --verbose --logfile --timeout 60
The Unix fixture covers exact execution and recovery, lost mutation replies, foreign/ambiguous resources, image mismatch, degraded shutdown, never-started containers, malformed IPC, parent death and cancellation. These mechanism tests do not establish durable assignment admission or authorize a production cutover.
Image collection
Pallet's containerd holds only the images Pallet pulled for its own runs and the bundled
sandbox image. PalletExecutionOwner.collectImages removes the images no durable assignment
still needs. A collection is due when a lifetime starts and after every removal the owner proves.
The node runtime runs it at the end of a full assignment sweep, on the same native lane as a
reconcile, so nothing pulls or creates a container meanwhile.
The owner states what to keep. Every assignment whose runtime removal is not proven keeps its
run's image config (platformEvidence.imageConfigDigest), whatever its disposition: a stopped
attempt is still inspected and removed with its image, and an admitted one still pulls it. If
any retained assignment's image cannot be named, nothing is collected. The native pass
(collectExecutionImages, rust/src/execution.images.rs) also keeps:
- every image a CRI container still uses, in any state and matched by any name the image carries;
- every image the runtime reports pinned;
- the sandbox image by its bundled name.
It removes the rest in id order with CRI RemoveImage and requires ImageStatus absence
afterwards; an acknowledged removal that left the image in place fails the pass as
RUNTIME_STATE_CONFLICT. It reports the images listed, removed with their size, kept by reason,
left for the next pass, and the bytes the image filesystem reports in use.
| Bound | Value |
|---|---|
| retained image configs per pass | 1,024 |
| images the runtime may hold when a pass begins | 1,024 (refused as RUNTIME_STATE_CONFLICT beyond) |
| images removed per pass | 64; the rest are remaining and keep the collection due |
The node runtime runs the collector at the end of a full assignment sweep, and only while the
node holds an authenticated controller session: before the first one a node admits no native
effect, so a due collection makes no CRI call and runs on the first sweep after the session is
established. The collector runs only while the durable execution lane is idle on this boot, because an
operation still pending there may be a run whose pull the runtime is finishing. A pass that
cannot run stays due and is tried again after at most one minute
(palletImageCollectionRetryMs). It is recorded on the node status images as one of:
deferredwithlane-pending;failedwithretain-unresolved,retain-bound, orruntimeplus the native code.
A lane or assignment store that cannot be read is recorded the same way: an unreadable lane is
not proven idle and is deferred with lane-pending, an unreadable retained assignment is
failed with retain-unresolved. Besides its own refusals (no controller session, a busy lane,
an owner not ready), the collector rejects only when the identity runtime refuses its assignment
store or its own code fails, a defect the status cannot name. The node writes
Pallet image collection failed unexpectedly: <class> <code>[:<site>]. to stderr, which Spark
forwards to the node journal, once per distinct failure until a collection runs again.
It never fails the node: an image left behind costs disk, not correctness. No journal is kept: removing an unreferenced image is idempotent, and a later run pulls by digest.
pnpm exec tstest test/test.imagecollection.node.ts --verbose --logfile --timeout 120
Workload stats
A controller reads one resource sample of a running attempt with readRuntimeAssignmentStats
(@serve.zone/interfaces 32.36.0). The request names the exact current revision of an assignment
on this node, and the controller client binds it to the session and that revision before any native
call; like every inbound request, it acts only on the connection that holds the session.
PalletExecutionOwner.readStats reads outside the native lane, so a read never waits for or blocks
a reconcile. Each read runs its own native child (sampleExecutionStats,
rust/src/execution.stats.rs): it proves the attempt's exact CRI identity, calls CRI
ContainerStats and PodSandboxStats, and proves the same identity again. A change between the two
proofs is a conflict, never a sample of whatever runs now. An attempt without a running container is
answered not-running, without a runtime stats call.
A sample states the cumulative CPU time across every core (decimal digits, because it outgrows a
JSON number; utilisation comes from two samples), the memory working set, the memory limit the node
enforces (the config's memoryBytes), the sandbox's default interface counters (or null) and the
writable-layer usage (or null). A counter a JSON number cannot carry exactly is refused as
RUNTIME_STATE_CONFLICT.
| Bound | Value |
|---|---|
| reads per assignment at a time | 1 (another is refused as busy) |
| reads across the node at a time | 8 (palletObservationBounds) |
| native deadline per read | 10 s |
pnpm exec tstest test/test.executionstats.node.ts --verbose --logfile --timeout 60
Workload logs
Every attempt's stdout and stderr are captured by the runtime and delivered to the controller, which
keeps them (IRuntimeConfig.logs: SmartData owns log metadata, SmartBucket the payloads). The node
keeps no log history of its own: its capture files are disposable transport artifacts on tmpfs.
Capture. The sandbox names the attempt's capture directory
/run/serve.zone/pallet/logs/<run digest hex> and the container its file workload.log below it;
containerd writes the CRI log format there (<RFC 3339 time> <stdout|stderr> <P|F> <content>). An
attempt whose container was created without a log path (a container from an earlier Pallet) is
stated once per process as a not-captured loss.
Delivery. PalletLogCapture reads the file and sends reportRuntimeAssignmentLogs batches
(runtimeAssignmentLogContract, @serve.zone/interfaces 32.36.0) on the session, one at a time:
the next batch leaves only after the controller acknowledged the previous one, and an
unacknowledged batch is sent again unchanged. A CRI piece longer than one entry (64 KiB) travels as
partial entries; a line longer than the config's logs.maximumLineBytes is cut there and followed
by a line-cut loss of its stream, counting the cut bytes. Every delivery pass
(palletLogDeliveryBounds) sends at most eight batches per capture before the next capture's turn.
Controllers that keep no logs. A controller answers every batch accepted, replay or
not-kept. not-kept states that it keeps no workload logs and holds nothing: the node sends that
session no more batches (runtimeAssignmentLogContract.notKept, @serve.zone/interfaces 32.37.0).
The answer is also the node's permission to discard for that session (@serve.zone/interfaces
32.38.0), and Pallet uses it while that session stays the current, live one: every flush and every
hold on the delivery cadence discards each ended capture whole — its directory, its cursor and its
waiting batch — and each capture that ends later, and has each running capture drop its output
instead of holding it. A running capture drops every unread whole CRI line and a batch it sent that
holds output; it records the dropped batch's number with its cursor, so that number is never sent
again, not even by a restarted process, and its next batch opens with one not-kept loss counting
the dropped bytes and whole lines (loss evidence the dropped batch carried leads it too). A sent
batch that holds only loss evidence waits unchanged. The next session is asked again, and a
controller that keeps logs receives each running capture from where it stands, after that loss. Any
other failure of a batch keeps it the capture's next one and sends it again unchanged on the next
pass.
Ended captures. A capture whose container ended waits for its final batch to be acknowledged,
and a crash-looping workload ends one per restart, so the node keeps at most
runtimeAssignmentLogContract.maximumEndedCaptures (64) ended captures awaiting acknowledgement,
holding at most maximumEndedCaptureBytes (64 MiB) together, each measured as the buffer bound
measures it: in the file bytes it holds for delivery (Bounds below). Each pass holds every ended
capture within its own buffer bound and rewrites its files to what it still needs before it is
measured, so a capture at the largest buffer (64 MiB) always fits whole. Past either bound — a controller that is down, fails every batch or leaves the method
unhandled — it discards the oldest ended captures whole, oldest by when this process saw them end.
Every capture discarded whole, under this bound or a not-kept answer, is counted in the
discardedCaptures of the next batch the node creates, of any capture. The count is persisted
(pallet_log_discards): it survives a restart, travels unchanged with the batch that carries it on
every retry, a restarted process creating that batch again with the same number and count, passes
on to the next batch when that batch is discarded unacknowledged (dropped under not-kept, discarded
with its capture, or lost to a restart that cannot continue its capture), and is kept while a
not-kept answer stands. The owner keeps in memory only running captures and ended ones within the
bound: a capture that finished or was discarded leaves no directory, cursor or entry behind.
Rollout order. A controller must answer reportRuntimeAssignmentLogs before a Pallet with log
capture runs against it. A controller that leaves the method unhandled fails every batch: the node
stays within its bounds, but it retries the same batch on every flush and drops the output beyond
the buffer bound, so no workload output reaches such a controller. Cloudly answers not-kept
(since 33.8.0). Onebox has no Pallet log backend yet (Onebox 33.1.1): run this Pallet against
Onebox only once one ships.
A controller must also run @serve.zone/interfaces 32.38.0 or later before this Pallet runs against
it. A receiver on 32.37.0 or earlier reads a batch by its exact schema and refuses one that carries
discardedCaptures or a not-kept loss, on every retry: against it the node stays within its bounds
as against any failing controller, but that capture's output no longer arrives. Cloudly 33.8.1 is
the first Cloudly on 32.38.0 (33.8.0 is on 32.37.0); an Onebox Pallet log backend must be on
32.38.0 when it ships.
Bounds. What the controller has not acknowledged is the capture's buffer, at most the config's
logs.maximumBufferedBytes. Pallet measures it in the file bytes the capture holds for delivery,
which is what sits on the tmpfs: the file bytes the batch waiting for its acknowledgement spans, plus
the file bytes not yet read into a batch, CRI headers and the cut rests of lines beyond
logs.maximumLineBytes included. The waiting batch counts with its whole span whether or not a
reclaiming pass (below) already removed those bytes from disk. A batch never spans more file bytes
than the bound (its first CRI line is always taken), so the waiting batch is always kept and sent
again unchanged; beyond the bound the oldest unread whole CRI lines are dropped and stated where they
were dropped as a buffer-overflow loss, which counts their exact output bytes and whole lines and
leads the next batch. Output of short or cut lines therefore reaches the bound sooner than its
output bytes alone would: a line cut to a few bytes still counts every byte the runtime wrote for
it. Drops while a batch waits extend that one loss, which keeps the
time the loss began, so however long the controller stays away the evidence is a single entry.
Once the node has read past the buffer bound (never below 1 MiB)
it renames the file aside and has the runtime reopen the log path (CRI ReopenContainerLog through
the native reopenExecutionLog, which proves the attempt's exact identity before and after); the
renamed file is drained first and removed once the cursor moved past it. On disk an attempt
therefore holds at most the read part below the rotation bound plus its unacknowledged buffer.
The runtime reopens only a running container's log, and its answer ends a capture only on definitive
evidence that nothing writes the file any more: the container exited, its sandbox stopped, or it was
removed. A container created but not yet started, or one in an unknown state, may still write:
the reopen is refused as indeterminate like a runtime that could not answer, so the renamed file
gets its name back, the capture keeps running and no rewrite touches its files, and a later pass
asks again.
A rotation survives a Pallet process that stops inside it. Between the rename and the runtime's
reopen, the renamed file is still the one the runtime writes, and on disk that window is a renamed
file (workload.log.1) without a fresh workload.log. Every pass that finds it so asks the runtime
to reopen before it reads on, unless the container ended, and every step that rewrites or removes the
renamed file completes the rotation first, so none ever replaces or unlinks a file the runtime still
writes: output written there would otherwise be neither delivered nor counted, and would grow unseen
on the tmpfs. When the runtime cannot answer a reopen, the renamed file gets its name back by a hard
link, never a rename over the name, so a fresh file the runtime created meanwhile (a reopen whose
answer was lost) is never replaced; a process that stops between the link and the removal of the
renamed name leaves one file under both names, and the next pass removes the renamed name. A
reclaiming rewrite that stopped before its rename leaves a copy (workload.log.1.compact) that the
capture removes when it opens.
A pass that had to drop also reclaims the read part, which the waiting batch no longer needs on
disk: it rotates the file at once and rewrites the renamed file from the first unread byte, so after
that pass the attempt holds on disk only its unread bytes, within the bound, however long a batch
waits. The bound holds whatever the controller does: every delivery pass keeps the captures it
reaches within it, a pass whose batch fails holds every capture before it ends, and the controller
client holds every capture on the delivery cadence (deliveryIntervalMs) whether or not a
controller is connected, so a controller that is down, fails every batch or leaves the method
unhandled never lets a capture grow. A hold that fails writes Pallet log capture hold failed: <class> <code>[:<site>]. to stderr, which Spark forwards to the node journal, once per distinct
failure until a hold succeeds. The capture's delivery, hold and reconcile run one at a time.
A capture's own failure — its files, or a runtime that cannot answer its reopen or cannot yet tell
whether its container will write — is that capture's alone: the delivery pass or hold goes on with
every other capture, the failed one is tried again on the next pass, and the pass reports the first
such failure once it is done (a delivery pass rejects, a hold writes the line above). A controller's
failure still ends the delivery pass, after every capture is held.
An ended capture is never written again, so it keeps on disk only what it still needs: the file
bytes of its waiting batch and the file bytes not yet read, CRI headers and the cut rests of long
lines included. Once its files hold more beyond those bytes than those bytes and 64 KiB
(palletLogCaptureBounds.settleSlackBytes), it rewrites them into one file (workload.log.settled,
renamed over workload.log), checked when it ends and after every acknowledgement. The cursor moves
to the rewritten file in one transaction before the rename, so a restart at any point continues from
exactly the recorded byte. An ended capture therefore holds on the tmpfs at most twice the file bytes
it still needs plus 64 KiB, instead of every byte the running container left behind. The bound is
amortized on purpose: rewriting after every acknowledgement would hold only the needed bytes, but
would copy what is left once per batch, quadratic in the capture's size (about 2.3 GiB of copying to
drain one 64 MiB capture in 900 KB batches). Rewriting only once the files are at least twice what
is needed copies each byte a bounded number of times, never more than was acknowledged or dropped
before, so draining a capture copies at most its own size. Since the ended captures measure at most
maximumEndedCaptureBytes together in the file bytes they need, after every pass they hold on the
tmpfs at most 2 × 64 MiB + 64 × 64 KiB = 132 MiB (138 412 032 bytes) together.
Durable cursor. Each capture keeps one record in the node's own database
(pallet_log_cursors): the capture, its last recorded sequence, the position after it (the kernel
identity — device, inode, birth time — of the file it points into, the byte offset and the line state
of both streams), the evidence that leads the next batch, and the waiting batch. A batch is recorded
when it is created, before it is sent: where its file bytes end, and every entry its creation
generated with its timestamp — the losses it opens with and the empty entries a final batch closes
open lines with — and the discard count it states. The cursor moves on only after an acknowledgement
is recorded. A restarted Pallet process continues the same capture from exactly the recorded byte,
mid-line included, and creates the waiting batch again byte for byte from those file bytes and the
recorded entries, so a sequence number once sent is never sent with other content. Evidence that
leads a batch holds one loss per reason, however many passes drop before it is sent. Only a
cursor whose file is gone — a reboot emptied the tmpfs, or a reclaiming pass rewrote the file
since the last acknowledgement — starts a new capture with a capture-restart loss of unknown
extent. When the container stopped, failed or was removed, the
capture drains the rest, closes a line left open, sends one final batch, and once that batch is
acknowledged records the capture finished, then removes its directory and then its cursor; a
restart between those steps completes the removal and never counts the delivered capture as
discarded. A capture discarded whole first, under a not-kept answer or past the ended-capture
bound, loses both as well. Its record goes first, so a restart before its directory is gone leaves
a directory without a record; the attempt's assignment stays listed (the assignment store keeps
every attempt's state as a tombstone), so the next process opens that directory as a new capture of
the ended attempt, delivers what it holds and removes it. The first pass never removes a directory
merely for lacking a record: a running capture has none until its first batch is created, and a
directory whose attempt no state names proves nothing about its writer (a database restored from an
older copy while the container runs), so such a directory stays until the reboot that empties the
tmpfs. A batch dropped under
not-kept records the cursor too, at the dropped number and past the dropped output, together with
the evidence the next batch opens with, so a restarted process skips the same number and states the
same loss. The first pass of a process forgets every cursor whose capture directory is gone (a
reboot emptied the tmpfs) and passes on the count of every carrying batch no open capture creates
again; a cursor recorded finished goes uncounted.
pnpm exec tstest test/test.logcapture.node.ts --verbose --logfile --timeout 60
pnpm exec tstest test/test.logowner.node.ts --verbose --logfile --timeout 120
Private workload readiness
PalletExecutionOwner binds the published runtime config policy to the exact
admitted assignment before invoking a native process, TCP or HTTP readiness probe.
Probes inspect the complete CRI identity before and after network IO and connect
only to the evidenced container IP addresses. HTTP uses the configured Host and
request target, checks final response headers without reading a body, and never
follows redirects. HTTPS verifies system trust and the configured hostname/SNI.
There is no insecure TLS or DNS-target fallback.
SmartData stores probe evidence, exact container start nanoseconds, boot ID, threshold counters and the next
due time. Linux CLOCK_BOOTTIME provides ordering across native children and daemon
restarts. Initial delay and startup grace use a conservative anchor from the first
completed native observation of that exact runtime identity. An already running
container therefore waits the configured delay when first discovered; a forward
wall-clock step cannot shorten it. Thresholds and intervals come from the immutable
config. Changed runtime identity, boot, config or clock
regression resets readiness; pending native effects supply no readiness proof.
Probes never overlap. A reconcile before the next due time returns
readiness-pending with no new observation.
A ready observation must match the durable sample's assignment, timestamp, image, sandbox and container. Its outbox transaction fences the current native lane. Restart and lost acknowledgements retain the exact existing observation; they cannot reuse a cached probe to create a new ready report. This remains a private composition. Authenticated controller transport and production route promotion are separate integration requirements.
Qualification uses the actual native probe, real HTTP/HTTPS listeners, a controlled
CRI server and disposable NoSQLDB stores. The focused tests are
test.readiness.node.ts, test.readinessstore.node.ts and
test.executionowner.node.ts; they do not establish new real-containerd or host
power-loss qualification.
Private packet policy composition
The backend-private composePalletPacketPolicies maps a validated network
projection and complete current attachment/handoff receipts into a combined
Smartnftables router policy and its host-transit policy. It captures inputs before
asynchronous validation. Attached local workloads receive exact veth sources;
selected remote workload grants require an explicit caller-owned TUN. Reserved
but unattached local workloads receive no packet paths. Private DNS uses only each
workload's gateway and explicit TCP/UDP port 53 rules.
Workload egress comes from the published projection grant helper. Router-origin DNS and platform traffic comes only from explicit signed router selections, with the current handoff's transit source and exact selected destination tuple. Public workload grants grant no router authority. Both policies carry the complete protected union and use only the current handoff's leased transport ranges. Withdrawals grant no traffic; DNS readiness does not change packet authority.
Published host ports come only from the verified projection's endpoints placed
on this node. Every entry keeps its Cloudly authorization reference; hostIp
is the node uplink address or absent; the workload must be attached with a
current receipt; a port held by a platform endpoint or resolver on the uplink
address, a port inside a leased SNAT range, or an entry beyond the bound refuses
the endpoint's whole set by a bounded publication.* site and blocks nothing
else. An entry may publish a contiguous range (hostPortEnd, each host port to
the same port inside the workload) and may be symmetric: outbound flows the
workload opens from its published port(s) leave from the same port(s) on both
hops, so a SIP or RTP peer sees one address and port in both directions
(@serve.zone/interfaces 32.31.0; pinned at 32.32.0, which also refuses a symmetric entry whose inside port another entry of its protocol shares). Every port of a range counts: a range that
reaches a leased SNAT range or a platform endpoint's port anywhere is refused like
a single port. The contract admits a symmetric entry only on an endpoint with
public egress and refuses the whole projection otherwise, before Pallet composes
anything. The bound is the contract's 64 entries per generation, a range counting
as one. The accepted set is handed to the host-transit policy as
hostIp:hostPort[-hostPortEnd] -> transitAddress:hostPort[-hostPortEnd],
journaled with the ACTIVE generation and retired with it. Leased outbound
translation keeps using the handoff lease's own source-port ranges; Cloudly
allocates new leases in 49152-65535 (runtimeNetworkHandoffSnatPortRange), so
publishable ports never meet them, while a lease allocated earlier keeps its
range and still refuses any publication inside it. Pallet's session registration
offers the installed interfaces release, so Cloudly sends range and symmetric
entries only to a node that reads them. The native compiler decides what fits its atomic
budget: while it refuses either policy as EXHAUSTED on a bound that
publications spend (their count, or the rule bytes, target bytes or operations of
the batch), Pallet refuses publishing endpoints by name
(publication.capExceeded), last first, and recompiles both. A policy the
compiler refuses as INVALID input, or as EXHAUSTED on any other bound, is a
composition defect that no refused publication mends: the generation fails with
PalletPacketPolicyRejectedError (invalid:packet.policyInvalid or
exhausted:packet.capacityExhausted), which carries the engine's bounded reason
and, for EXHAUSTED, the bound with its limit and actual count. The router policy carries the second hop of
the same accepted set, transitAddress:hostPort[-transitPortEnd] -> workloadAddress:targetPort with the same symmetric flag;
it is composed from the same verified entries, cross-checked against the journal
and applied by the native compiler, and the raise owns the sandbox default route
the workload answers through: CNI ADD installs default via <router address>
(static, metric 100) in every pod namespace and advertises exactly that route in
its result. Proven end to end in the DHCP-hub packet
qualification: an external client reaches the workload over TCP and UDP through
both translations, the workload sees the client's own address, the client sees
the published uplink address back, a refused entry never opens, and a successor
generation without publications leaves both ports dark.
Host grants let the node's own host namespace — Onebox's reverse proxy on a
single host — dial an exact workload port directly, with no loopback publication,
no route_localnet, no loopback DNAT and no loosened martian filtering. They come
only from the contract's getRuntimeNetworkProjectionHostPacketGrants over the
verified projection, so only a onebox projection's signed host selections
grant any; Pallet proves each again against that projection — the source is the
current handoff's transit host address, the destination the exact lease address
and port of an endpoint on this node — and refuses the composition otherwise.
The same exact tuples are compiled into all three tables the flow crosses: the
host-transit and router policies of the ACTIVE generation, for attached workloads
only, and the allocation-pool guard, where they are its only exceptions. A new
projection composes a new generation and guard target, so a withdrawn selection
is an ordinary atomic replacement; an absent or empty set leaves every policy
byte-identical.
A granted flow needs the host to route the lease address through the selected
handoff, and the ACTIVE generation owns that route. A generation whose projection
selects host flows plans one host route per workload pool that holds a selected
lease (hostRoutePrefixes, derived again from the journaled projection on every
read): the pool via the router's transit address, sourced from the host's transit
address, so the flow leaves with exactly the source its grant names. It is added
through a command socket in the host namespace inside the registered host
transition that raises the link, deleted inside the one that lowers it, checked in
every inventory of the raised link, and removed by recovery when it outlives a
fence. Any other host route change still retires the uplink observation, so a route
change from any other source is never accepted as the generation's own. The route
widens nothing: the guard and both hops still admit only the exact granted tuples,
and a generation without host selections — every Cloudly generation — plans none
and is byte-identical.
Workload ingress grants let the cluster ingress workload reach the port a target
workload listens on directly, instead of a hairpin through a published uplink port.
They come only from the contract's
getRuntimeNetworkProjectionWorkloadIngressPacketGrants over the verified
projection, so only a cloudly projection's signed workloadIngress selections
grant any; Pallet proves each again against that projection — two different
endpoints, at least one placed on this node, the exact lease address of each and
the target's listening port — and refuses the composition otherwise. They compile
into the router policy's workloadGrants (@push.rocks/smartnftables 4.0.0): a
stateful one-way flow of the exact protocol and port, answered only by the
replies of a connection the ingress workload opened, never translated, and never
open in the other direction. A grant compiles only when both of its workloads are
attached behind this router in the generation; an unattached end, or a target on
another node, receives no packet authority here, because the compiler has no
one-way grant across the tunnel. Large port ranges such as RTP media stay uplink
publications. An absent or empty set leaves the router policy byte-identical.
A projection's hostPlatformEndpointIds names the protected platform endpoints
this node's own host serves — on a single host the cluster hub, the relay
listener and Corestore. They compile into the host-transit policy's
localPlatformEndpoints: the host delivers leased flows to them in its INPUT
instead of forwarding them, from the exact handoff, with the lease's transit
source address and a source port of its range, and only the replies go back.
The workload and router selections that reach them are unchanged: a workload
still reaches only the endpoints it selects and the router only the hub its
block selects. Every address a declared endpoint holds is added to the uplink
binding, so the compiler verifies at apply, recovery and inspection that the host
holds it; each must be a permanent /32 of the uplink, which the uplink
observation admits beside the DHCP lease (see Retained uplink observation). An
undeclared platform endpoint or a resolver on one of those addresses refuses the
composition as conflict. Without the member the uplink binding and the host
policy are byte-identical.
This function performs no native apply, link activation or durable journal write. Its caller must authenticate the projection, retain and fence actual namespace, attachment, TUN and uplink generations, and own routes, DHCP changes and SNAT address lifetime. Native preparation must confirm graph capacity before a complete transition is journalled and applied. Composition alone provides no enforcement, workload readiness, allocation reuse or packet-drain evidence.
Private DNS lease arithmetic
The backend-private ts/network/lease.ts converts an authenticated lease of at
most 15 minutes to a conservative native CLOCK_BOOTTIME deadline. Its input is
a verified UTC interval anchored to the same boot clock, with an independently
qualified rate-error bound. It accounts for uncertainty and the fastest permitted
UTC progression. Process recovery retains the exact deadline; persisted boot-time
and UTC lower bounds reject clock regression.
A clock this node cannot read lapses that DNS acquisition and never fails the
lifetime that owns the clock, and which lapse it is follows from what the clock
states. A sample it cannot qualify leaves the acquisition without fresh time
evidence: it continues on the window already proved, keeping the exact deadline
this boot converted, and lapses only where there is no such deadline — another
boot, or a window this boot has not converted. A clock that states no boot
reading at all — unqualified, or lost with its native owner — is a
clock-unavailable acquisition outright: that pass withdraws every name and the
next pass binds them again. No reader is refused for asking while another holds
the clock; losing the owned process still fails the node through the clock's own
failure signal.
Reboot recovery requires newly verified independent time evidence for that boot. The original absolute expiry stays unchanged, and a saved Cloudly timestamp cannot renew it. The returned anchor, deadline and high-water state must be persisted through SmartData inside the complete authenticated projection before applying a newer native DNS revision. Expiry does not revoke packet grants or prove that an old writer is absent.
The private PalletIndependentClock owns an authenticated NTS measurement process
through the native executor. Its pinned Chrony 4.8-servezone1 variant returns exact
current Unix time and tracking metadata from one observation. Legacy Chrony
tracking offsets cannot supply this precision when the RTC is years wrong. The
private process never adjusts the host clock or loads host Chrony/GnuTLS settings.
It uses bundled trust certificates, fixed Ubuntu NTS peers, no drift/cookie files,
and a root-private Unix socket. One native operation runs at a time, and no caller
is refused for asking second: a caller that asks for a reading already in flight
is answered with that reading — a reading is of an instant, so sharing it is exact
and the owner is asked once — and a caller that asks for the other reading waits
for the operation in flight and then takes its turn, in arrival order. Loss of that
owned process fails the node lifetime; unqualified or offline samples supply no
time authority.
The owner checks coherent tracking, selection and NTS authentication reports, root error, freshness and kernel clock brackets. Supported KVM Linux clock profiles bound BOOTTIME progression by 400,000 ppm; unsupported clock sources, PPS/custom tick settings, VM pause, snapshot restore and live migration do not provide a qualified clock. The bootstrap certificate's validity interval is 1970–2100. A wall-clock reading or an NTP-synchronised flag alone is insufficient.
Clock ownership does not persist projections, supervise DNS or admit workloads. The DNS persistence and lifecycle owner remains unfinished. Lease arithmetic checks cover uncertainty, deadline boundaries, process recovery, reboot proof requirements, rollback and malformed stored state. The isolated native clock harness additionally checks actual NTS measurement, wrong-year RTCs, offline denial and joined Node/Deno cleanup; only recorded successful runs qualify the specified source and kernel.
Private network reservations
The backend-private ts/network/allocation/ store retains protected-authority
revisions and immutable handoff leases through the identity runtime's existing
SmartData/NoSQLDB connection. runNetworks() exposes callback-scoped operations;
shutdown joins admitted operations before closing the engine. Every write,
including replay, requires an owning-code identity fence inside its transaction.
The caller authenticates controller authority before entering this interface.
A fresh node can stage the complete current authority. An initialized node must
advance through exact consecutive references in the same controller epoch and
node scope. Reservations reference the current authority and validate against
all retained leases; a common revision fence serializes competing allocations.
Historical leases continue to bind their exact authority. Quarantine preserves
the lease and forbids reuse of its source-port range, conntrack zone and label
within the shared contract's respective scopes. Replay retains quarantine.
listHandoffs() returns the complete bounded retained history; there is no
expiry, pruning, release or reuse operation.
This stores allocation intent. It does not authenticate signed projections, activate pools, acknowledge host/router barriers, declare flow drainage, allocate sandbox IPs or admit packet traffic. Those owners must bind this durable history to exact native receipts before activation. Tests use actual NoSQLDB to check concurrent conflicts, transactional identity fencing, stored digest corruption, lost commit acknowledgements, complete restart recovery and joined shutdown.
Private signed network admission
PalletControllerClient receives the published Interfaces 30.9.0 signing-authority
and complete network-projection RPCs on its private outbound TLS connection.
The current physical session, active credential and controller/node/namespace
scope fence each admission transaction. Socket input is inert and bounded; the
transport accommodates the full 896 KiB projection contract.
ts/network/projection/ persists public signing trust, immutable public revision
history, one complete signed projection and bounded immutable workload leases
through SmartData on the existing NoSQLDB owner. Key rotation and projection
admission write a common revision fence. Rotation or revocation prevents the old
envelope from being read as current authority; historical keys retain verification
of recovery material. Re-signing an identical projection under the new key replaces
its envelope before acknowledging replay. Replays never renew the DNS deadline.
Projection admission retains historical protected authorities and handoffs in the same allocation transaction. Full local and remote workload leases survive withdrawal and restart, including the addresses needed for later denial. Retired leases stay quarantined; their subnets and execution attempts cannot be reused. The store retains at most 512 workload leases and rejects capacity exhaustion. Admission always carries pending withdrawals and tombstones because it has no joined native application receipt. An ACK establishes durable intent only.
The callback-scoped runNetworkProjections() facade expires with its callback;
shutdown joins admitted work. inspectRetainedProjection() is recovery material,
while readCurrentProjection() requires the current signing key. Neither method
establishes an application, time, or current-identity fence for a caller. Pool
allocation, CNI, native realization receipts, network activation, qualified DNS
time and live worker rollout remain separate owners.
The private application-store composition reads current verified signing intent and the complete retained handoff/workload history in one SmartData transaction. Before recording native intent, it rechecks that source fingerprint and writes both common trust and allocation fences in the caller's transaction. Concurrent key rotation, projection admission, reservation or quarantine therefore conflicts with stale intent; a failed identity guard rolls back both writes. This interface does not issue a native receipt or expose an application capability over RPC.
Tests test.networkapplicationsource.node.ts, test.networkprojectionstore.node.ts and test.controllernetwork.node.ts
exercise actual file-backed NoSQLDB and TLS sockets, including concurrent rotation,
identity rollback, lost ACKs, physical reconnect, shutdown and a complete namespace
with 12,288 DNS names. They establish admission behavior, not live packet policy.
Private DNS lease renewal
A renewal carries a fresh DNS window for one exact admitted projection and no
applied network state. PalletControllerClient receives
applyRuntimeNetworkDnsLeaseRenewal on the same outbound connection as the two
admissions, and the envelope's session must be the binding this node obtained
itself. That comparison is what makes a cluster relay's push admissible: the
relay holds signed statements for a node it carries and adds no trust of its own.
Admission is one transaction in ts/network/projection/, in this order: the
stored signing trust, this node's current projection, the signature under that
trust and under the revision that signed the projection, the renewal this node
already holds for that projection, the published admission decision against the
qualified clock's lower bound, the write, and the common trust fence every
network admission takes. accepted replaces the single stored row, replay
changes no renewal and repeats the same bound acknowledgement, and every other
outcome is the refusal the other two admissions give; the trust fence is taken on
both outcomes, so a concurrent rotation conflicts rather than interleaves. What
persists is the complete signed envelope, so every later read re-proves it against
the retained signer rather than trusting the node's own table. An envelope this
node can no longer prove is a lapsed renewal, not a broken node: the DNS lane
discards it and converts the projection's own window, while a credential bind
refuses. The persisted DNS lease and DNS intent each state the renewal their
window was converted from; no released version persisted either shape, so there is
no migration, and a store carried over from an older build fails its assertion on
every read and must be started from clean pallet_dns_lease and
pallet_dns_intents collections.
A renewal belongs to one projection envelope: that projection's exact reference and the revision that signed it. The transaction that admits any envelope the stored renewal does not belong to discards it — a successor projection, and equally the same body re-signed after a key rotation, which keeps the reference and changes the signer. A renewal of a projection this node does not hold, at a sequence it has already passed, with different bytes at a sequence it holds, or whose slot has not provably opened at the qualified reading is refused; the sender simply delivers it again once it has. A rotated signing key renews nothing: the re-signed projection discards the renewal held and the node lives on the window that projection itself signs. While that window still runs nothing else changes; once it has ended the node withdraws DNS and stays dark until the controller delivers a renewal signed by the new revision. Pallet asks for none — a renewal arrives on the connection this node already holds — and converts the one it admits on its next DNS pass, which runs once a second.
Everything window-bounded then follows the effective window — the admitted
renewal's when one renews this projection, the projection's own otherwise. The
DNS lease converts that window in the same transaction that writes the lease and
records the renewal it converted, the composed DNS snapshot proves its
validUntilBoottimeMs against it, and the managed-VPN credential may expire
inside it and never beyond it. Nothing else moves: the application, its
fingerprint, the native generation and the protection receipt are untouched, and
a node that loses qualified time or current signing authority still withdraws
DNS exactly as before.
test.networkdnsleaserenewal.node.ts, test.dnsrenewalview.node.ts,
test.dnsleasestore.node.ts, test.independentclock.node.ts and
test.controllertunnel.node.ts state admission, supersession, replay, each
refusal, the boundary at an equal expiry, the moved deadline, restart recovery, a
stored renewal this node can no longer prove on either lane, a reading two
overlapping callers share, a clock this node cannot read, and the bound the
credential keeps.
Private identity store
The separate backend-private ts/identity/ implementation owns Pallet-only
credential state in an explicitly supplied SmartData database. It is not imported
by the public probe facade and has no Cloudly connection. Its protected lifecycle
and node-local enrollment control owners remain private and unwired from the
production daemon.
Preparation commits a 32-byte cryptographic palletToken before returning its
SHA-256 proposal. Binding checks this owner's durable hash and the complete
enrollment digest. Activation requires the exact bound acknowledgement and writes
its immutable receipt in the same transaction. Initial adoption may name an
existing Cloudly node while Pallet is still generation zero. Rotations retain the
old active bearer until acknowledgement; historical receipts cannot reactivate it
or replace pending material. Inspection, proposals and receipts contain no bearer.
The enrollment-specific helpers use published Interfaces 28.2 snapshots, digest and acknowledgement binding. Those helpers prove content binding, not remote authentication: the eventual authenticated transport/coordinator owns that trust boundary. Spark credentials never enter Pallet's private store.
Tests use disposable file-backed NoSQLDB 10.5.1 instances and prove exact replay across engine restarts, concurrent preparation, acknowledgement rollback, independent database bindings and rejection of changed or malformed input:
pnpm exec tstest test/test.identitystore.node.ts --verbose --logfile --timeout 60
Private identity runtime
ts/identity/classes.identityruntime.ts owns a Linux file-backed NoSQLDB 10.5.1
engine and its SmartData connection. Production defaults to UID 0 and
/var/lib/serve.zone/pallet; trusted owning code supplies the installed engine's
absolute path and SHA-256. It rejects symlinks, unsafe ancestors, wrong ownership,
engine hardlinks, nonexecutable or writable engine files, and non-0700 data roots.
It never repairs permissions or discovers an alternative executable. Ordinary
startup requires existing storage; only explicit first provisioning may create it.
The installer must exclude concurrent changes to the selected engine and paths.
NoSQLDB's fileStorageStartup: 'current-format-only' rejects legacy/mixed storage
and orphan migration staging without modifying that tree. The native file-root
lease is the only cross-process database lock. Model preparation happens after
native ownership and readiness, with no duplicated format detector or lock helper.
Each attempt uses its own mode-0700 /tmp/pallet-identity-* directory for the Unix
socket; it contains no persisted application data. Normal cleanup removes only
that owned socket and empty directory. Parent death can leave an empty disposable
directory for OS temporary-file cleanup; startup never scans or deletes others.
run() admits a guarded identity facade, not the database or store object. Its
methods expire when the callback settles, and admitted calls are tracked even if
the callback does not await them. Stop immediately blocks new callbacks, drains
admitted work, closes SmartData, and then confirms native exit. A failed database
close retains the running engine's lease until cleanup is retried. Start/stop
deadlines bound caller waiting, not resource ownership; pending or failed cleanup
blocks restart. Native exit invalidates readiness and requires stop before restart.
pnpm exec tstest test/test.identityruntime.node.ts --verbose --logfile --timeout 60
The Linux tests qualify competing Node owners, successor receipt recovery, startup/stop races, cleanup failures, exact pending recovery after native SIGKILL, and active receipt recovery after an idle Node parent's SIGKILL. These are not hardware power-loss or ARM64 hardware tests, and do not prove bounded native exit when a management command is stuck. Production process supervision and coherent installer/engine packaging remain integration gates. When packaging that engine, retain NoSQLDB's own license and complete third-party notices alongside the binary.
Private enrollment control
ts/control/classes.enrollmentcontrol.ts owns its identity runtime and a root-only
Unix listener at /run/serve.zone/pallet/control.sock. It acquires the native
database lease before touching that socket. The installer provisions the trusted
platform parent; the control owner creates its mode-0700 IPC leaf when absent,
creates a mode-0600 socket and leaves the directory in place after shutdown.
Existing directories are checked, never repaired. Paths must be canonical,
symlink-free, protected from untrusted writers and fit Linux's socket path limit.
Only the five published requests.pallet enrollment and runtime-binding methods are registered.
Prepare returns a current, independently retained hash-only proposal; bind proves
the full pending enrollment; activate checks the complete authenticated Cloudly
acknowledgement and then proves the exact current local identity. Historical
receipts cannot masquerade as current activation. Spark must authenticate Cloudly
before forwarding its acknowledgement or runtime routing binding over this trusted
local channel. bindPalletNodeRuntimeBinding accepts the first exact routing tuple
or an exact replay only after its origin and node match the current active Pallet
identity. readPalletNodeRuntimeBinding requires that same current identity and
rejects an absent binding. The binding is a separate singleton: absence is the
unbound state, and the released identity document's relayOrigin field remains
readable but inert rather than being expanded into unauthenticated routing data. The IPC
does not expose a bearer, database, generic rotation or workload execution method.
TypedRequest 8.0.3 owns routing and envelope identity. Hooks and incoming response routing are disabled; wire-supplied local authority and unknown methods are denied. Each connection carries one four-byte big-endian length-prefixed JSON frame per direction followed by write-half-close. The server explicitly keeps the response half open until its reply is flushed. Requests are limited to 32768 bytes and responses to 16384, with a default 10-second total connection deadline and 16 admitted operations. Invalid UTF-8, incomplete frames, trailing bytes and missing EOF are rejected. No TCP listener or automatic application retry is introduced.
Disconnects and deadlines close transport but do not abandon an admitted database operation or free its admission slot. Shutdown stops admission, cancels socket I/O, drains handlers, closes the owned listener and then stops the identity runtime. Caller deadlines retain the actual cleanup owner and block restart. A live pre-existing socket is never removed; stale recovery requires the native lease, a completed connection-refused probe, and unchanged owner/inode checks. Because Node/libuv unlinks on listener close, shutdown checks its recorded socket and parent before calling close. A replaced name retains failed-cleanup ownership until the installer/operator restores the original owned path. These checks assume the trusted installer excludes concurrent root-owned path changes; they do not claim protection from hostile root. No arbitrary file or recursive cleanup occurs.
pnpm exec tstest test/test.enrollmentcontrol.node.ts --verbose --logfile --timeout 60
Qualification uses real Unix sockets and disposable native NoSQLDB stores for half-close delivery, independent restart, competing ownership, stale/live/replaced paths, malformed frames, hook isolation, cancellation, admission limits, historical receipt rejection and lost activation results. Spark's concrete client, coordinated daemon integration and production packaging remain separate integration gates. Before distributing a bundled control process, include its JavaScript dependency MIT/Apache-2.0 notices in addition to the existing Rust and NoSQLDB notice material.
Enrollment process lifecycle
The private PalletEnrollmentProcess owns one foreground control lifetime.
Readiness follows the native database lease, model preparation and protected Unix
listener. Cancellation during startup prevents readiness. Parent stdin EOF,
SIGTERM and SIGINT stop admission and drain the existing control owner. A 250 ms
watch of the public readiness getter terminates a failed lifetime; it never
restarts the engine or retries a database operation. Each control request already
checks native readiness independently of that watch.
The enrollment CLI accepts exactly enrollment-serve or enrollment-provision.
Normal startup requires existing state; only explicit provisioning may create it.
Trusted build code supplies the exact engine, digest and protected paths. No
runtime path, UID, digest, bearer or configuration arrives through argv, stdin or
environment. Stdin carries only parent lifetime: content is rejected without an
echo. PALLET_CONTROL lines use pallet.enrollment.process and contain only
ready, stopped, or a static failure reason. stopped follows confirmed cleanup;
a failure does not authorize a successor until the supervisor confirms process exit.
pnpm exec tstest test/test.enrollmentprocess.node.ts --verbose --logfile --timeout 60
The lifecycle and source-process tests cover cancellation, EOF/signals, missing state, invalid input, native failure and retained cleanup ownership. A standalone Deno (the pinned 2.9.7) Linux amd64 fixture also qualifies real Unix half-close responses, durable replay, engine and parent SIGKILL, and stale-socket recovery from an unrelated working directory with an empty environment. That fixture supplies trusted test paths and UID. The production artifact, complete notice bundle, installer/supervisor integration and target-platform qualification remain gates; this source adapter is not a production node installation.
Foreground node process
The same control executable accepts runtime-serve for one private
PalletNodeProcess lifetime. It reopens existing enrolled state and verifies the
protected sibling pallet-runtime executor, pallet-guard, pallet-dns and the
managed VPN pallet-vpn against their compiled-in SHA-256 identities before
starting local owners. The released build activates the node network: it passes
activation with pallet-vpn and its digest, so a node bound to its cluster
relay raises its ACTIVE generation and managed VPN tunnel (see "Private ACTIVE
tunnel transitions"), and a node bound to a local controller forwards as a single
host. It never initializes missing storage or accepts paths,
credentials or settings from argv, stdin or environment. The containerd CRI
socket is /run/pallet/containerd/containerd.sock: Pallet's own containerd
instance, never the host's shared /run/containerd/containerd.sock, which on a
Docker CE host is Docker's containerd.io daemon with CRI disabled. The native
runtime refuses that shared socket and Docker's sockets by path and by file
identity: a socket with the device and inode of the shared containerd socket,
Docker's API socket or Docker's embedded containerd socket is refused, so a
symlink, hard link or bind mount of one of them cannot pass under another name.
Pallet and Docker can therefore run side by side on one host without sharing
images, sandboxes or CRI configuration.
The same lifetime runs in two modes that differ only in what ends it.
runtime-serve is a supervised child: its parent holds its standard input, and
EOF there stops it (Spark's runnode). runtime-unit is the main process of the
node service unit (see "Service unit contract"): its standard input is null, EOF
is not a stop request, and SIGTERM or SIGINT stops it. Both refuse data on
standard input as invalid_stdin.
A start fences what the previous lifetime retained before it reports ready, so
a start after a crash does real work first. A supervisor must allow a start at
least 600 s before it gives up on ready, in either mode
(palletNodeStartTimeoutSeconds in @serve.zone/pallet-bundle). Pallet bounds
each native request of a start, not the start as a whole: a request that misses
its deadline ends the start with failed. The longest path whose length does not
grow with the node's workloads is a start after a crash, on a host whose guard
unit already ran in this boot, with an ACTIVE generation retained and a guard
expansion pending. Its request deadlines add up to 469 s:
- identity database 10 s (its startup deadline, migrations included), router namespace 9 s (3 s spawn, two 3 s requests), independent clock 11 s (3 s spawn, 8 s start);
- guard pass 274 s at smartnftables' 30 s per request: owner open 33 s (with its 3 s spawn), replay of the committed transition 90 s (prepare, reconcile, inspect), the pending expansion 120 s (prepare, then its replay), detach 31 s;
- handoff set open 11 s (namespace read 3 s, spawn 3 s, open 5 s);
- fence of the ended ACTIVE generation 131 s: the native's epoch proof 5 s, the retained host packet table's owner 36 s (namespace read, spawn, open), its recompilation 30 s and its release 60 s (release, inspect);
- application recovery 15 s (epoch proof, prepare and reconcile at 5 s each);
- execution host identity 8 s (3 s spawn, 4 s request, 1 s termination grace).
Each retained workload attachment of the ended epoch adds its 5 s epoch proof (at
most 512, twice the projection's 256 endpoints). The DNS pass, the database
transactions, the executable digests and the storage recovery have no deadline of
their own. 600 s covers the fixed path and leaves 131 s for those; a supervisor of
a node that retains many workloads allows more, and PalletControlLifetime
admits up to an hour (469 s + 512 × 5 s = 3029 s). Under the node unit
(Type=exec) the service manager does not wait for ready, so TimeoutStartSec=
does not bound a start; an installer that waits for the unit's ready line
allows the same limit.
Runtime lifecycle lines keep the PALLET_CONTROL prefix with the distinct
pallet.node.process protocol. ready means local identity, namespace and execution owners
are ready; it does not assert controller connectivity, workload readiness or DNS
availability. A start the packet engine refuses because of the running kernel ends
with failed reason unsupported_kernel instead of owner_failed (see "Host
kernel requirements"). Initial Cloudly connectivity runs independently. The runtime
session requires the immutable node runtime routing binding and registers at its
cluster relay. An absent binding fails startup; there is no direct-to-Cloudly
fallback or second configured relay authority. Every registration response must
match the persisted node, cluster, Cloudly controller and runtime namespace before
it becomes execution authority, and reconnect rechecks the current identity and
the exact persisted binding. The persisted tuple grants routing only; phase and
workload authority always come from the fresh authenticated runtime session.
Enrollment, the node credential and the registry host stay with Cloudly
whichever transport carries the session: workload.registryHost is the
publication identity the registry credential is fenced to, and it is never
rewritten. Where the image bytes are fetched is a separate statement,
workload.pullEndpoint — the relay's own origin for a cluster whose relay
forwards the registry, so the node pulls inside its cluster and nothing
cluster-side dials the control plane for bytes. The reference stays
digest-pinned either way, so the endpoint decides reachability and never
content, and the credential this node asks Cloudly for is addressed to whoever
serves them. The contract this node relies on from
the relay: it is expected to forward the node bearer verbatim and to serve the
unchanged typed request contracts, so nothing in the session binding changes
for Pallet. New workloads
still require the exact authenticated session and registry grant, while persisted
node-bound effects retain their existing offline rules. Network and secret
references resolve through their owning resolvers on this same session before any
native effect; storage references resolve to the local storage claims this node
admitted on the same session ("Local storage claims").
Protocol handshake
The contract this node speaks is the installed @serve.zone/interfaces release,
and it is stated once: ts/controller/protocol.ts builds one offer for the one
session kind Pallet is (palletRuntime) from protocol.createProtocolOffer,
with this build's own minimum (32.0.0, raised only by the commit that starts
depending on a later minor). Every registration carries that offer, and the
answer is the wrapped { session, protocol } the controller returns, negotiated
against the offer that was sent before the binding is bound.
Two peers that cannot serve one session are named rather than collapsed into a
transport failure. A controller that refuses this build answers an
IProtocolRefusal; an accepting controller whose own offer this node cannot
serve is refused by this node. Either way the node reports the status
protocol-incompatible and retains the refusal in getStatus(). It is not an
owned failure: the workloads keep running, the network stays up, the packet
policies stay applied and the node's owners are not torn down — and the node
process is not ended either, because protocol-incompatible is a live node with
no controller it can speak to.
Recovery needs no operator action on the node. The refused offer is made again
every protocol.refusedOfferRetryIntervalMs (five minutes, the one interval
@serve.zone/interfaces states so that no client invents a second one) until a
controller accepts it, through @api.global/typedsocket's
TypedSocketRestoreDeferral: the refused attempt is closed, the transport waits,
and it offers again without consuming one of its reconnect retries, so a node
refused for days keeps them for real connection failures. An accepted offer drops
the retained refusal, returns the client and the node to ready and resumes the
pass, so upgrading the controller under a running fleet brings the fleet back by
itself. The node's own poll loop keeps its interval while refused — the pass stays
empty and admits nothing — because that loop is what carries the accepted offer
back into the node's state.
A restoration the transport refuses outright is the one terminal outcome.
TypedSocket releases itself, because a retry could only repeat the peer's
verdict, and that leaves the node with no controller at all — the condition an
exhausted transport leaves it in — so the node fails owned: the process ends and
the supervisor's restart is the next offer, which lands on the same five-minute
cadence if the controller it meets has still not moved. The denial is named on
the controller status as restoreDenial — the transport's own reason and, when
the refusal was this client's, the cause behind it — beside any refusal the node
was already standing on, so the status the process ends on says which verdict
ended it.
The registration also states resolvedAuthorities, the sorted, duplicate-free
list of workload authorities this build can resolve. It is the same constant the
execution owner enforces (resolvableWorkloadAuthorities), so what the node
promises and what it refuses can never drift: an assignment whose
requiredWorkloadAuthorities the list does not cover is refused as
unresolved-authority. This build resolves network, secrets and storage.
Local controller
A node can be driven by a controller on its own host — Onebox — instead of
Cloudly (nodeRuntimeLocalControllerContract in @serve.zone/interfaces). A
node carries a Cloudly identity or a local controller, never both: a first local
bind is refused while any Cloudly enrollment state exists, and a Cloudly
enrollment is refused once a local controller is bound. Both claim one shared
pallet_node_authority singleton in the transaction that writes their identity,
and the first claim inserts it, so of a concurrent Cloudly preparation and first
local bind exactly one commits. The other's commit conflicts with it, the
transaction is retried, and the retry finds the committed authority and refuses
as conflict. An offline read that finds
both kinds anyway refuses instead of choosing one. The persisted identity
selects the transport when the controller client starts; there is no configured
choice between them.
The local controller is authenticated by the socket's permissions and nothing
else: the socket is owned by root with mode 0600 in a root-only directory, so
any process running as root on the host is trusted as the controller. The
bearer Pallet generates and the controller pins at its first bind detect a
Pallet that lost its state (a later bind answering another credential is
refused by the controller); they do not keep out an intruder with root.
Pallet serves the local controller on the root-only Unix socket
/run/serve.zone/pallet/controller.sock through SmartServe's unixSocket
listener, which applies mode 0600 before any peer can connect and refuses a
path another process serves. The controller is the TypedSocket client
(http://localhost over the socket), but the session keeps its directions.
Its first request on a connection is bindPalletLocalRuntimeController: the
first bind persists the binding and a bearer this node generates, and every
later bind must be the identical binding and answers the same credential
generation and SHA-256 hash. The bearer itself never leaves this node except in
its own registration. After answering, the node fires the unchanged
registerPalletRuntimeSession at that exact connection with that bearer, and it
binds the answered session to its own persisted binding
(bindRegisterPalletLocalRuntimeSessionResponse) before it becomes execution
authority. Admission, reports, registry and secret requests then run as on the
outbound transport, but only on the connection that holds the session: another
connection on the socket, bound or not, never acts under it. A connection whose
registration is accepted supersedes the previous session and closes its
connection; a failed registration closes only its own. A controller that
refuses this build's offer leaves the node protocol-incompatible until a later
connection's registration is accepted.
A local controller has no Cloudly origin to stand for its registry. It issues a
pull credential only for an image whose registryHost, and pullEndpoint when
one is stated, are both listed in the binding's registryHosts
(isNodeRuntimeLocalRegistryWorkload); any other image is refused before a
credential is requested.
A node bound to a local controller always serves it from its node lifetime
(runtime-unit under its service unit, or runtime-serve). Only
network-acquire-local admits the first bind of an unbound node (see
Initial projection acquisition). Networked
workloads of a local controller need the attach barrier like every other
workload; on a node without a managed VPN it holds as a single host (see
Private ACTIVE tunnel transitions).
Stdin EOF, SIGTERM and SIGINT close admission and join startup, controller and
native operations. The CNI broker remains available through CRI rollback, then
joins its admitted handlers before the network and database owners close.
The private clock stops and joins its Chrony/query children before the retained
router namespace closes and the database owner is released. Terminal node failure ends
the lifetime with a static failure reason. The 250 ms liveness watch never restarts
the node or retries an operation. A stopped event follows successful cleanup;
Spark must also confirm child exit before admitting a successor.
This process composition does not start containerd, provide private DNS,
activate Spark's daemon or establish production readiness. Spark must consume the
released Pallet bundle and commission its guard before starting this runtime, and
Pallet's own containerd (containerd-serve, below) serves the CRI socket
/run/pallet/containerd/containerd.sock this runtime uses.
Pallet's own containerd
The sealed control bundle carries Pallet's container runtime under
pallet-containerd/: containerd 2.3.6 and its runc v2 shim (the unmodified
upstream release executables), runc 1.5.1+servezone1 (built from the signed
upstream sources, statically against musl), the pause 3.10.2-servezone1 sandbox
image as an OCI archive (pause.tar) and manifest.json, which records every
file's size, mode and SHA-256. The control executable compiles in that manifest's
digest. No upstream CNI plugin ships: the only network plugin is Pallet's own
pallet-cni, which is the verified pallet-runtime executable, and containerd's
internal loopback serves lo.
pallet-control containerd-serve is the main process of the containerd service
unit. It verifies its sibling pallet-runtime against its compiled-in digest and
starts it in its --containerd-management mode, which:
- verifies
pallet-containerd/against the manifest: exact names, modes, sizes and digests, root-owned and not group- or world-writable up to/; - refuses to start a second containerd while one answers on
/run/pallet/containerd/containerd.sock; - writes the configuration below to
/run/pallet/config/config.toml(mode0600), the CNI networkpallet-workloadsto/run/pallet/config/cni/conf/10-pallet.confand a link/run/pallet/config/cni/bin/pallet-cnito the verifiedpallet-runtime, all recreated on every start; - starts containerd with an empty environment except a system
PATH(for host helpers containerd resolves by name, such asapparmor_parser), bound to its own lifetime (PR_SET_PDEATHSIG), and waits up to 60 s until CRI reports versionv2.3.6andRuntimeReady; - imports the bundled pause archive through containerd's own Transfer service over
a Streaming session, as
ctr image importdoes, when containerd does not already hold it with the pinned config digest, unpacks it for theoverlayfssnapshotter underpallet.local/pause:3.10.2-servezone1, and requires CRI to report that image id. Nothing is pulled from a registry, and noctrexecutable ships.
Then the process reports PALLET_CONTROL {"protocol":"pallet.containerd.process","event":"ready"}
on stdout and, when started by a Type=notify unit, READY=1 over systemd's
notification socket. containerd's own log lines pass through to the process's
standard error. containerd's exit fails the process (failed, owner_failed,
exit status 1); it is never restarted inside the process — the service manager
restarts the unit. Once containerd is ready, SIGTERM or SIGINT stops it with SIGTERM,
waits up to 30 s, removes /run/pallet/config and reports stopped. A SIGTERM during
startup kills the starting containerd at once and fails the process; the left-over
/run/pallet/config is recreated by the next start. Stopping containerd
never stops a workload: every shim, and every container beneath it, keeps running
and is reattached by the next containerd. Standard input is not read (a unit's
null stdin is expected; data on it is refused).
The generated configuration is complete and is Pallet's own:
| Setting | Value |
|---|---|
version |
4 |
root, state |
/var/lib/pallet/containerd, /run/pallet/containerd |
imports |
[] — the host's /etc/containerd/conf.d never applies |
required_plugins |
io.containerd.grpc.v1.cri |
disabled_plugins |
io.containerd.internal.v1.opt, io.containerd.image-verifier.v1.bindir, io.containerd.nri.v1.nri |
| gRPC / TTRPC | /run/pallet/containerd/containerd.sock / .sock.ttrpc, uid 0, gid 0 |
| CRI stream server | 127.0.0.1:10010, stream_idle_timeout = "15m", no TLS streaming, no CRI TCP service |
| images | overlayfs snapshotter, sandbox image pallet.local/pause:3.10.2-servezone1, no registry host directory |
| runtime | runc via pallet-containerd/containerd-shim-runc-v2 and pallet-containerd/runc, runc state /run/pallet/runc, SystemdCgroup = false |
| CDI | enable_cdi = false, no specification directories |
| CNI | /run/pallet/config/cni/{bin,conf}, one configuration, internal loopback |
| other | image-defined volumes ignored, network namespaces under the state directory |
Workloads get cgroups under the absolute parent /pallet (the executor names it
on every sandbox), never under the unit that runs containerd: runc can enable the
CPU, memory and PID controllers there because the cgroup holds no processes, and
a stop of the containerd unit reaches no workload.
The CRI stream server (exec, attach and port forwarding) listens on loopback only,
and only root may dial it: every allocation-pool guard policy carries the
Smartnftables loopback port owner { address: '127.0.0.1', port: 10010, uid: 0 }
(see Offline allocation-pool guard journal).
Another user's connection is reset, and the port is dropped on every interface but
loopback. The containerd unit requires the guard unit, and a guard pass that did not
apply the owner fails, so containerd never serves without the rule.
Service unit contract
Spark on fleet nodes and Onebox on its own host install the unit; Pallet never writes a unit file. The unit must:
- run
<bundle>/pallet-control containerd-servewith no further arguments as its main process, as root,StandardInput=null; - use
Type=notifywithNotifyAccess=all(the ready notification comes from the native owner, a child of the main process); - use
KillMode=process, so stopping the unit signals onlycontainerd-serve, which stops containerd itself, and every shim and workload survives; - use
Delegate=yes, as containerd's own unit does, so systemd leaves the cgroups of the processes it starts to them; - set
TimeoutStopSec=of at least 45 s (30 s containerd grace plus the owner's joins),Restart=alwayswith a shortRestartSec=,LimitNOFILE=infinity,TasksMax=infinityandOOMScoreAdjust=-999; - order
After=andRequires=the retained guard unit, and be orderedBefore=the node unit, whichWants=it: the node runtime is the only CRI client.
The node unit runs the node lifetime:
- run
<bundle>/pallet-control runtime-unitwith no further arguments as its main process, as root,StandardInput=null,Type=exec; - use
KillMode=mixed: SIGTERM reaches only the main process, which joins every owner and child it started, and whatever remains of the unit's cgroup when the stop timeout ends is killed before the service manager starts a successor; - set
TimeoutStopSec=of at least 360 s (a five-minute native operation, then the database joins) andRestart=on-failurewith a shortRestartSec=: a failed lifetime is restarted by the service manager, never inside the process; - order
After=both the guard and the containerd unit,Requires=the guard unit andWants=the containerd unit.
A start of the node may take up to 600 s before its ready line (a start fences
what the previous lifetime retained first; the derivation is in "Foreground node
process"). Type=exec does not wait for it; an installer that does allows at
least that long.
The node unit is started only after the node's first projection is acquired
(network-acquire-local or network-acquire); an unbound node fails its start.
The guard unit is a oneshot (Type=oneshot, RemainAfterExit=yes) that runs
guard-recover with null stdin, without default dependencies, ordered Before=
network-pre.target, systemd-networkd.service and shutdown.target, which it
conflicts with, and required by network-pre.target and systemd-networkd.service.
Its executable is pallet-control or an installer's own executable that verifies
the installed bundle before it starts that mode. The installer writes it only after
the first guard-commission.
@serve.zone/pallet-bundle (from this repository's ts_bundle/) states these three
definitions for an installer and reads the loaded containerd and node units back
against this contract.
The QEMU CNI guest's containerd-serve scenario runs this process from a
build:control directory without systemd (see
Isolation and qualification); the unit's own
KillMode, Delegate and notification behaviour is qualified with the installer.
Router namespace lifetime
The verified pallet-runtime executable also owns one private router namespace
for each node process lifetime. Its exact --network-namespace-management mode
creates an unnamed Linux network namespace on the initial native thread, before
Tokio, sockets or readiness. Before it exposes the namespace it switches IPv4
forwarding off in all and default and disables IPv6 in both, and verifies
each (with lo's forwarding): a new namespace copies the host's IPv4
forwarding, and the router must forward nothing unless an UP generation's packet
policies filter it. This requires permission to create network namespaces;
failure prevents node readiness and controller admission.
The node opens the keeper's live namespace descriptor, compares its device and inode, rejects its own namespace, and confirms the native nonce and identity over the original stdio connection. A reused PID or saved metadata cannot establish authority. Nothing mounts or names the namespace, and no namespace capability is persisted. A new node lifetime creates a fresh namespace.
The private PalletNamespaceOwner.withDescriptor() API retains the source
descriptor through an admitted consumer's asynchronous spawn. Callbacks must not
close it or keep its number after returning. Shutdown fences new borrowers and
joins admitted callbacks. The enclosing node must also join all DNS, VPN and
packet-policy processes before closing this owner. Unexpected keeper exit fails
the node, stops admission and retains the parent descriptor until explicit joined
cleanup; it never silently recreates the old namespace.
The isolated Linux 6.18.35 amd64 fixture runs both Node and compiled Deno (the pinned 2.9.7)
against the real namespace keeper and published Smartnftables 1.4.0 binary. It
verifies descriptor inheritance, actual namespace identity, keeper-loss fencing,
joined borrowing and unchanged parent links and routes, and — on a host that
forwards IPv4 — that the namespace is exposed with forwarding off in all,
default and lo and IPv6 disabled in all and default. The full node composition
is separately checked through its compiled control bundle. Privileged ARM,
containerd/CNI attachment, DNS and complete network policy remain separate gates.
Staged secret material mounts
The same verified executable owns one staged mount per admitted assignment that
carries a secret reference. Its exact --secret-mount-management mode stages that
attempt's material on its initial native thread, before Tokio, sockets or any
other work: it unshares a private mount namespace, mounts a
noswap,nodev,nosuid,noexec tmpfs there, writes one root-owned file per entry
with the delivered mode, places the ownership marker inside the mount and
detaches it with open_tree.
This requires permission to create mount namespaces and mounts, and it must run in
the host mount namespace — the helper refuses its own work when it does not.
The material never crosses the JSON IPC. PalletSecretMountOwner spawns one child
per operation and writes the opened values as one published Smartrust sensitive
frame on descriptor 3, bounded by this attempt's own declared total rather than by
a shared ceiling. Prepare and attach share that one child, because the detached
mount dies with the process that holds it; inspect and release run in fresh
children, because recovery may never depend on a process that is already gone. No
value is ever a return value, a persisted field, an environment value, a command
line or an IPC member, and the caller's buffer is cleared after the write.
Publication follows durable state, never the other way round: the mount plan and
the pending run commit in one transaction before the first native step, the
prepared receipt — secured parent mount, base and staged device/inode, detached
mount id, ownership nonce — is persisted before move_mount publishes the mount,
and the published mount id is verified equal to the id the detached clone already
allocated. The executor then demands exactly that plan's read-only private
per-file binds from the container, never a bind of the mount directory that holds
the marker, and refuses a live container whose mount list is not that plan.
Recovery reads the persisted receipt alone. A fresh helper verifies the recorded
mount by descriptor, unmounts strictly, proves that exact mount id gone and the
base directory empty before removing it, and answers uncertain with an exact
reason on any mismatch — which retains the mount, the directory and the record and
deletes nothing. Stop retains the mount; release is an explicit owner step after
proven CRI absence, never a side effect of removal. That order is the only thing
that protects a running workload: a container's own bind of a file in the mount
lives in its own mount namespace and does not make the host mount busy, so the
kernel would not refuse a release taken out of order. A record from an earlier
boot names nothing this kernel can find, because /run is a fresh tmpfs after
every boot, so it is retired without any native step.
The chain /run/serve.zone/pallet/secrets is walked through no-follow
descriptors. /run/serve.zone is the parent the installer supplies for every
serve.zone product on the node, so it is required to be exactly what the CNI
plugin requires of it — a root-owned directory no one else can write — while the
two components below it belong to Pallet alone and must be root-private.
The offline QEMU qualification drives the shipped binary itself, over the same
protocol, through every stage boundary on 6.18.35-0-virt, 6.8.0-124-generic
and 7.0.0-22-generic, and the containerd guest runs one admitted assignment
end to end: sealed material opened, mount published, the container created with
exactly the plan's bind, the marker invisible inside it, stop retaining the
mount and removal releasing it after proven absence.
Local storage claims
A controller — Cloudly or Onebox — gives a service a retained, node-local volume
with applyRuntimeLocalStorageClaim (IRuntimeLocalStorageClaim,
@serve.zone/interfaces 32.24.0). A claim names the volume by identity only:
controller, organization, cluster, node, runtime namespace, service and the
service's volumeId. It carries no host path, no device, no mount option and no
credential; this node resolves the identity to its own storage.
Layout. Every volume lives under /var/lib/serve.zone/pallet-storage:
volumes/<volumeKey> is a volume, staging/ holds one being created, trash/
one being purged and import/ the bytes an operator places for an import
claim. The volume key is the SHA-256 of the claim's identity and id, so it is
stable across every generation and a re-epoched controller keeps its volumes.
The root and its four directories are root-owned 0700 below trusted ancestors;
a volume is reachable only as the bind a granted container receives. The root is
created only while no volume was ever materialised: a root that later goes
missing — an unmounted disk — is refused by name (storage.root.missing) and
never replaced by an empty one.
Ledger. The claim ledger (pallet_storage_claims, pallet_storage_names,
pallet_storage_grants) lives in the node's own NoSQLDB through
@lossless.org/client/nosqldb. Admission is one transaction fenced to the
session the claim arrived on: a new claim, the next generation (whose previous
must name the one this node holds), a replay of the same generation or a
historical older one. A same-generation claim with another digest, a next
generation that changes the volume's identity, source, initialization or initial
ownership, a generation after a purge and a second live claim of the same service
volume are refused by name and write nothing. The answer is the volume's state:
ready, import-required, purged or uncertain.
Crash safety. Every filesystem effect runs between a ledger intent and the
completion that records its evidence, one at a time. An empty claim commits
provisioning, then creates the directory in staging/ with the claim's
initialOwnership and mode and publishes it by one atomic rename; the ledger
records the directory's device, inode and birth time. A purge commits purging,
renames the recorded directory into trash/ and removes it there, never
following a link the workload left inside. Before the ledger records a step
done, every directory it created, renamed or removed is fsynced with its parents
(a new layout directory with its parent; the volume with staging/ and
volumes/ before ready; volumes/ and trash/ before purged), and a failed
sync keeps the intent pending, so a power loss can only undo a step whose intent
the next start still finishes. On start, and on every replay, an
interrupted intent finishes from what is on disk: a leftover staging tree is this
intent's own and is removed, a published directory is adopted only while it is
still exactly the fresh, empty directory the intent creates, and a purge whose
volume already sits in the trash completes. Evidence that contradicts the ledger
— a foreign directory at the volume's name, a volume whose identity changed —
leaves the claim uncertain, which this node neither mounts nor deletes until an
operator decides. A pending intent that cannot finish at start is listed in the
execution owner's storageRecoveryRefusals and stays durable.
ReadWriteOnce. A run that seals storage references is planned before the
registry, secret material or any native step: every reference must name the
claim's current generation, the claim must bind to the run
(bindRuntimeLocalStorageClaimToAssignment), its volume must be ready, still be
the recorded directory and be held by nobody else, and no two volumes of the run
may nest or cover a path Pallet or the runtime binds itself (/run/secrets,
/run/serve.zone, /opt/serve.zone/runtime-assets, /proc, /sys, /dev,
/etc/hosts, /etc/hostname, /etc/resolv.conf). A run that cannot have its
volumes waits as storage-pending, with the step named (storage.claim.held,
storage.claim.importPending, storage.claim.generation, …). The grant is taken
in the transaction that begins the run, so a concurrent purge or a second holder
conflicts instead of both proceeding. Stop keeps the volume held; the proven CRI
absence of the removed container releases it. Release never deletes bytes: only a
purge generation does, and a held volume's purge is refused
(storage.purge.held) until its holder is removed. The container's mount list is
part of the attempt's identity, so every later inspection demands exactly the
grant's binds.
Offline import. An import claim waits as import-required until its bytes
arrive. With runtime-serve stopped, the operator places each volume's tree at
/var/lib/serve.zone/pallet-storage/import/<claimId> on the same filesystem and
runs pallet-control storage-import. The pass (in ts_migration/, like every
data adoption step) records the placed directory's identity, publishes it as the
volume by one atomic rename, fsyncs import/ and volumes/ and records it
ready; the tree keeps its ownership,
modes and every other attribute, and nothing is copied. It reports the adopted,
waiting and refused claims on its ready event (pallet.storage.import). A tree
on another filesystem, a file instead of a directory or a volume name already in
use is refused by name and moved nowhere; an interrupted pass finishes on the
next run, and a recorded tree found in neither place leaves the claim
uncertain. Keep a copy of the source first when a rollback may need it.
Network filesystems are not part of this build. The contract (32.24.0) also
states nfs and smb sources, but this node materialises node-local
directories only: such a claim is refused before the ledger is touched
(storage.source.unsupported), so neither the claim nor its service volume name
is recorded.
Dormant host/router handoff
The private ts/network/classes.handoffowner.ts mechanism consumes an already
reserved, digest-verified immutable handoff lease and a live PalletNamespaceOwner.
start() prepares an inert intent bound to the current boot and both namespaces.
The caller must persist that complete intent before reconcile(previousReceipt)
can create the exact veth pair, with its peer created directly in the router
namespace. Both endpoints remain administratively DOWN. Only the leased IPv4
addresses and their kernel local /32 routes are present; connected prefix routes
belong to eventual activation.
The separate --network-handoff-management native mode opens one netlink socket
in each retained namespace on its initial thread and restores the parent namespace
before polling either socket. It uses bounded native operations without shell
commands, named mounts or filesystem state. Linux ignores aliases in this veth
creation request, so the owner first verifies the newly created DOWN endpoints,
sets both ownership markers through link updates, and verifies them before adding
addresses. An interrupted partial creation never qualifies for adoption or deletion.
Recovery requires the exact boot, lease, namespace, link indices, peer indices,
names, MACs, markers, addresses and local routes. Extra or changed state is rejected.
Deletion uses the verified endpoint index and destroys only that exact pair. A
release states the administrative state its own contract guarantees: every
handoff path releases a dormant pair, so an endpoint raised behind the owner's
back is refused as handoff_conflict and the pair stays, and only the raise
measurement releases a raised pair. Removing a raised workload pair in
production is the stage-aware progress rollback, never this release.
Privileged exclusivity belongs to the enclosing node; these observations are not
a compare-and-swap guarantee against concurrent privileged network writers.
The TypeScript owner serializes operations and fences namespace or native-process
failure. getPrepared() and getReceipt() retain detached immutable evidence after
failure. close() joins the native lifetime; a successful return with
released: false means cleanup remains unconfirmed and the lease must remain
quarantined. A rejected close has not proved that join; joined exposes its state.
Join every namespace consumer before closing the keeper. Kernel namespace teardown
is asynchronous, so keeper exit alone cannot establish link deletion or pool reuse.
Both native handoff management modes use bounded nonblocking pipe or stream-socket stdio on the same thread that owns their namespace sockets. Host, router and retained sandbox netlink connections remain polled while input is idle or partial and while a response waits for the parent to read it. Partial input survives cancelled reads; malformed or oversized frames terminate the protocol. Output has the same three-second terminal bound as native commands. A partial response is never retried on the same stream. No reader thread can retain the mutation lock. This establishes protocol-owner liveness; the dormant sockets still have no topology multicast subscriptions and provide no continuous route or link fence.
This mechanism is not yet composed into signed node network admission. Durable
realization receipts, router/host/Docker policy barriers, CNI,
DNS, link activation and production qualification remain separate prerequisites.
The offline handoff fixture is test/native/handoff.ts, driven by
test/native/qualify-handoff.py in a disposable amd64 VM without a network device,
host disk or host mount. It runs the actual native binary through Node and compiled
Deno, including partial-state and foreign-state faults, restart recovery, keeper
loss and joined cleanup. ARM64 is built but this privileged fixture does not
establish ARM64 execution qualification.
Complete dormant handoff set
ts/network/classes.handoffsetowner.ts supplies the private node-wide mechanism
for complete retained membership. Start it with the live namespace, persist the
empty getPrepared() intent, then reconcile(target, null) to inspect and establish
an empty native receipt. prepare([{ lease, presence }]) accepts all 256 complete,
digest-verified leases permitted by the allocation contract, including
absent/quarantined history, with at most 32 desired present handoffs. Persist the returned
target and the last applied receipt before calling reconcile(target, previous).
Preparation does not change the kernel. Receipts contain an exact DOWN pair or
explicit absence for every member; absence never frees an allocation.
The --network-handoff-set-management process retains both namespace descriptors
and owns one host-network-namespace abstract Unix socket lock. The single-pair
mechanism uses the same lock, so no two Pallet handoff mutators can coexist on
that host namespace, even with different router namespaces. This IPC lock is
concurrent process exclusion, not authentication or durable ownership. The node
must still exclude other privileged network writers and namespace capabilities.
Before any transition effect, bounded complete link, address and all-family/table
route dumps inspect the whole expected set. The private router namespace must
contain only its empty, DOWN loopback and exact dormant handoffs. Unknown router
interfaces, addresses or routes fail ownership. Unrelated host interfaces and
routes remain permitted; unexpected Pallet markers/names, misplaced known MACs,
transit address conflicts and surviving/reused receipt indices are rejected.
Stable bound facts are re-read after inspection and after the complete transition.
Typed requests retain the terminal NLMSG_DONE; every inventory requires a
successful completion and rejects interrupted, filtered, malformed, unexpected
or oversized responses. Both namespace connections also reject unsolicited
messages and receive-buffer loss. The pinned maintained netlink-proto source
corrects decoding at its owning layer so malformed packets cannot disappear
before a later successful completion.
Transitions retain all prior members and cannot revive an absent member. Exact
present-to-absent deletion keeps the lease in the set. Separate pair operations
are not atomic: after a lost response, persistently prepared intent and the prior
receipt can adopt exact completed members and finish the remaining transition in
the same live namespace. Changed or half-created pairs stay failed-owned. EOF or
process termination preserves DOWN effects for that recovery. Full node death
does not make an unnamed namespace recoverable. close() deletes only exact
known members and joins the native lifetime; uncertainty preserves receipts and
returns released:false, with no guessed cleanup or allocation reuse.
The native transport permits 1,048,576-byte frames; TypeScript captures inert
complete-set data within the corresponding bounded budget. The serialization
test measures a 701,951-byte replay with maximum-length lease and
namespace identities. Native operations retain their three-second deadline and
the bridge its five-second deadline: the 256-retained/32-present guest operations
completed within the native deadline in the isolated fixture. Complete dumps establish absent members without issuing
redundant per-member link dumps. The offline handoff
fixture additionally covers empty/nonempty exhaustive inventory, unrelated host
links, competing single/set owners, a competing router namespace, quarantine,
lost-response and partial-transition recovery, 256 retained members with 32 present paths and joined
wrapper shutdown under Node and compiled Deno. This stage exposes no UP method
and issues no nativeBarrier: the authenticated SmartData application journal,
positive protection evidence and eventual firewall/DNS/VPN/CNI lifecycle remain required.
Retained uplink observation
After reconciling the complete dormant topology, the private handoff-set owner
can openUplink() and return a boot-, host-namespace- and random-generation-bound
observation. inspectUplink(generation) performs fresh bounded kernel and
networkd reads; copied facts cannot recreate the observer. A wrong generation
is rejected without retiring the current observation. closeUplink(generation)
joins the read-only observer and does not release topology or change DHCP.
The initial supported host has one physical Ethernet uplink, a finite bound
systemd-networkd DHCPv4 lease, one main-table IPv4 DHCP default route and the
three canonical unselected IPv4 policy rules. The observer records exact link,
address, gateway, source, metric and IPv4/IPv6 all/default/interface NETCONF
settings. Beside the lease the uplink may carry permanent /32 IPv4 host
addresses — the addresses a platform service on this host binds, whether
networkd configured them (Address=) or found them on the link. Each is a
universe-scope address with an infinite lifetime, so it adds no prefix route and
never becomes the default route's source; any other IPv4 address on the link
still refuses the observation, and adding or removing one retires it like every
other address change on the uplink. The observation states them
(hostAddresses, sorted as strings); an observation a 33.0.0 node journaled
carries none, which states nothing. Existing host IPv6 addresses remain possible; this observation grants
no IPv6 workload authority. Networkd remains the only DHCP mutator.
The native lifetime subscribes before capture and drives both RTNL notifications and the fixed system D-Bus connection during inspection and quiet management I/O. Its unique networkd owner and exact returned link object path remain bound to the lease. The observation follows the lease networkd renews, so a node keeps forwarding through every DHCP renewal: at the lease's renewal time (T1), on a property event of the link, and when networkd rewrites the leased address with new lifetimes, the native reads the uplink and its lease again. While networkd renews or rebinds, the lease in force holds until its valid lifetime ends, never longer. A fresh acquisition from the same networkd owner for exactly the same binding — link, address, prefix, gateway, source, metric, NETCONF settings and host addresses — whose lifetimes the kernel address already carries becomes the lease in force; the observation keeps its generation and its receipt, which states the lease it opened on. The lease's expiry, service replacement, protocol loss, a read that finds the binding changed in any fact or the lease moved backwards, and every other relevant kernel change retire the coordinator. Current values cannot erase an observed kernel change and restoration: only the leased address rewritten in place asks for a read, and any other address, link, route, rule or NETCONF change of the uplink still retires. Unrelated link notifications are ignored; route/rule changes conservatively retire this observation until exact coordinator-owned transitions are implemented. Close the observer before further dormant topology changes.
A coordinator that ends while it waits — its uplink observation retired, an
ACTIVE observer failed, a namespace channel lost — lowers every ACTIVE link on
its way out, as it always did. It names the watch that ended it in a bounded
PALLET_FAULT line (handoffset.drive.*, uplink.watch.*, uplink.follow.*),
and the handoff-set owner writes its failure chain once, as it retires, to
stderr, which Spark carries into the node journal:
Pallet network handoff set failed: <chain>. The stop that follows names the
same chain as the reason its release could not be confirmed
(noderuntime.close.networkRetained < …).
This mechanism exposes no UP or packet-readiness result. Active routing,
firewall composition, durable active journals and joined withdrawal still belong
to the activation coordinator. The ignored Rust test
network_handoff::kernel::uplink::tests::qualified_networkd_host_observation
provides an explicit read-only check on a real networkd-managed host; it changes
no interfaces, addresses, routes, sysctls or services.
Private ACTIVE tunnel transitions
The retained PalletHandoffSetOwner accepts beginActiveTunnel(effect),
finishActiveTunnel(effect, expectedReceipt) and inspectActiveTunnel(effect)
while its exact ACTIVE generation is UP. The effect binds the complete ACTIVE
preparation and a creation plan or previous tunnel receipt to a durable intent
digest. It includes the authenticated SmartVPN generation, complete plan digest,
router namespace, exclusive interface name, /32 address, MTU, split routes and
same-boot BOOTTIME deadline. Tunnel routes cannot capture owned transit or local
workload prefixes, the control address or the hub endpoint.
Persist the intent before begin. Retain the actual SmartVPN client in the router namespace, prepare its tunnel, freshly inspect that generation, and derive the expected receipt from its actual interface index before finish. Native code observes the bounded ordered kernel transition, checks the full namespace twice and returns to its strict observer. It does not create the tunnel or infer the continued existence of the VPN owner from a saved receipt. Packet policy must be installed and verified before enabling VPN forwarding or raising workloads.
For deletion, stop forwarding while retaining the UP tunnel, lower the packet
policies back to their base shape while the device still exists, persist and
begin the deletion intent, then disconnect and join the VPN device owner before
finish with null, and release both packet policies only after the generation's
withdrawal. Each generation therefore carries a chain of policy pairs per role,
all journaled and each exactly one revision above the pair it replaced: the base
pair at the generation number (composed with no tunnel, so remote peers are
deferred rather than permitted), the tunnel pair (the same projection composed
with the retained TUN bound as the pallet_vpn endpoint), the lowered pair (the
base shape again) — a tunnel and a lowered pair for every device the generation
raised, each device's above the lowering of the one before — and between them one
amended pair for every workload attached while the generation is UP (below). A
generation raises at most 32 devices (maximumActiveTunnelCreations): every read of
the chain walks the tunnel lane back to its first command and revalidates each row,
so the lane bounds the cost of every tunnel-phase, amendment and cleanup step (64
rows at most on a lane this release admits). The journal refuses the next creation by name (exhausted,
active.tunnel.exhausted) before it admits a command. The orchestrator refuses it
before it asks for a credential to raise a device or dials, and refuses a
replacement under another credential before it lowers the live device or dials;
only the renewal's own credential request, which names the other identity, comes
first. The native's own bound is larger (256 session generations per generation,
maximumNativeTunnelGenerations, remembered so none is reused), and the journal's
must never exceed it. The pass fails with that name, its owner withdraws the
generation in order, and the next pass raises a fresh generation whose lane starts
empty. The bound governs admission only: releases up to 35.2.0 bounded a lane by the
native alone, so a lane they wrote holds up to 512 rows
(maximumStoredTunnelCommands). It stays readable — inspected, recovered, lowered,
cleaned up and withdrawn as any other, each chain read costing what it cost under
that release — and only its next creation is refused. The compiler releases only the exact graph
a transition names, so no revision is derived from a phase: every pair states
its revision and the acknowledged pair it replaced, and activePacketChain
proves the whole chain on every path that admits or acknowledges a pair above
the base pair, and before a recovery re-applies or a cleanup releases the live
pair. The tunnel pair is applied only after finishTunnel, because the
compiler binds the device by the interface index the receipt carries, and always
before forwarding is enabled; the lowered pair is applied before the delete
command, so no enforced graph ever names an absent link. Linux omits deletion
notifications for the static split routes; full inventory must prove their
absence. ACTIVE
withdrawal is rejected while a tunnel or unfinished transition remains. Unexpected
notifications, owner loss or the absolute lease deadline retire ACTIVE and fence
known interfaces DOWN. Recovery requires joining the VPN owner first and proving
the whole namespace through a fresh ACTIVE owner. No receipt grants allocation
reuse or proves packet drain.
The private protocol allows 2 MiB frames, including complete retained history and up to 1,024 split routes.
Both the durable tunnel journal and the application lifecycle composition now
exist. PalletActiveStore persists every managed TUN command as its own row —
beginTunnel admits one create or delete before the device owner is asked to
act, acknowledgeTunnel records the native begin, and finishTunnel records the
terminal fact — chained per ACTIVE intent so a delete consumes exactly the
receipt its predecessor left behind, and a normal withdrawal is refused while a
device is still retained or a transition unacknowledged. A recovery records DOWN
only; it never claims a deletion, and the retained row stays as history. The
creation row also carries the tunnel pair applied over the device (packets)
and, once the tunnel is being taken down, the lowered pair (lowered); a
cleanup row names the live pair — the last one the chain admitted — by its exact
transitions and policy digests, and the packet owner's release is fenced on
exactly those digests.
A workload may be attached while the generation is UP. The native admits the
pure prepareWorkloadAttachment under a held generation, and an
addWorkloadAttachment for a pair the generation does not hold yet is created
by the generation itself, through its own observers, while it is UP with no
tunnel transition pending, holds fewer than 64 pairs (prepared and attached
together) and has no tunnel reaching into the new subnet. 64 is what the packet
engine is measured to hold: Smartnftables 3.0 compiles a router of 64 workloads,
each with local DNS, public TCP and UDP egress, four platform endpoints and up to
three publications, inside its atomic batch with room to 89. The native registry,
its router-namespace inventory and configuration (one link per present handoff
and per pair, 96) and its admission fence (192 targets) are bounded to match; the
compiler's prepare() stays the authority on what a given graph fits. The bound above 32 is verified by compilation and unit tests only; the root QEMU guest qualification has not yet run at 64. The attachment owner
asks the held generation before it journals the ADD intent
(PalletHandoffSetOwner.admitAttachment,
handoffset.admitAttachment.notUp/capacity/tunnelOverlap); a refused ADD
releases its preparation and fails by that name with no journal, because a
pending one would fence every amendment of the generation while it is held.
addWorkload judges the same rules again before the native is asked
(handoffset.addWorkload.*), because a native refusal retires the owner. In the
other order a tunnel is refused where it would reach an attached pair's subnet,
by the store before the command is journaled and by the owner before the native
is asked. The pair is DOWN and outside every policy at that point. Before it is raised the
orchestrator amends the generation (PalletActiveOrchestrator.amend):
beginAmendment compiles the live pair again with the attachment as a further
source, in the live shape, one revision up, and appends the row
{ attachment, shape, revision, host, router } to the journal's ordered
amendments log in the same commit that fences the current source as an
extension of the one the generation carries — the same settled application,
every carried workload exact, the new attachment complete; finishAmendment
appends each role's acknowledgement after its apply. The intent, the native
preparation and the UP record never change: activeSources(journal) is the
intent's source extended by the log, and the tunnel pair, the lowered pair,
every recompilation and the cleanup compose over it. A generation that already
carries the attachment has nothing to amend, a lowered generation takes no
attachment, and a node without a held UP generation amends nothing: the pair
is dormant and the next generation is prepared over it. DNS needs nothing new —
the completed attachment is a source of the next resolver revision exactly as
after a dormant ADD. A recovery sends the native attachedWorkloads, the
complete receipts of the restored pairs the preparation does not cover
(required, [] when none); the native restores DOWN only when its registry is
exactly the preparation's workloads plus those. A restored attachment it cannot
list — an ADD whose receipt never became durable, or one whose removal is
already intended — is removed first, in the fresh process and before its native
adopts anything, and journaled exactly as a DEL journals it
(PalletWorkloadAttachmentOwner.recoveryRemovals). A failed journal step, a
foreign intent or a failed native removal refuses the recovery by its step
(handoffset.recoverActive.beginRemoval/removalIntent/incompleteAttachment/finishRemoval).
A pair raised under a held generation stays UP through its withdrawal, and the
withdrawal restores the router namespace's baseline configuration. The raise is
therefore refused by name (active.raise.ipv6_baseline) unless that baseline
keeps IPv6 disabled on the pair's router side — all/disable_ipv6 must be 1
before the generation is prepared — because IPv6 enabled again on an UP link
brings addresses and routes no owned transition admits. The same rule applies
to pairs raised before the generation is prepared (active.prepare.ipv6_baseline).
The native router owner establishes and verifies all and default IPv6-disabled
baselines immediately after creating its private namespace, before exposing it.
The router namespace is fail-closed between generations for the same reason: the
withdrawal releases both packet policies while the raised pairs stay UP, so the
IPv4 forwarding it restores must be off. The native router owner switches all
and default forwarding off at creation, and a generation is prepared only when
every forwarding switch it captures — all, default, lo and each router link
— is off; otherwise the preparation is refused by name
(active.prepare.forwarding_baseline). Activation turns forwarding on only after
both packet policies are applied, and the withdrawal turns it off again before
the policies are released, so workload traffic through the router stops from a
generation's withdrawal until its successor is UP (a projection change's
withdrawal window) and is never forwarded unfiltered.
For a containerd sandbox, the admitted native ADD disables IPv6 only on its
newly created, identity-proven DOWN eth0 before addressing it; the exact pair
DEL removes that link. Containerd runs CNI before creating the sandbox container,
so CRI sandbox sysctls cannot provide this pre-CNI guarantee. The sandbox's
namespace-wide defaults and the host's sysctls are never changed.
A node with the mpls_router module loaded is not supported: every link
registration emits an MPLS netconf record, which the attach vocabulary refuses.
The tunnel credential is the controller's. PalletControllerClient.resolveTunnel
issues getRuntimeManagedVpnCredential { session, projection } over whatever
transport the runtime session uses — the relay when the node is bound to one —
for the exact projection the generation was raised from and a reading of the
qualified clock, and turns the answer into the session request the device
owner authenticates with: the cluster hub's address and port, key material,
authority and node ids, and a same-boot BOOTTIME deadline bounded to fifteen
minutes or the credential's own expiry, whichever is sooner. The node dials its
own cluster's hub, never a platform-wide one: the credential names the hub
endpoint this node's signed projection selected, and the contract binds that id,
its transport's protocol, its address and its port to that selection. The
transport decides the dial form — a managed QUIC hub is a bare host:port, the
one transport the contract states — so a credential naming any other is refused
rather than dialled as if it were QUIC. Every refusal is named
(controller.tunnel.notLive, disconnected, projectionAbsent,
projectionChanged, boot, credential, transport); the key material is returned to the
caller only and is never journaled or logged. The orchestrator asks for it once
the generation is UP and DNS is serving; a refused credential or a refused
authentication leaves the generation UP without a tunnel, reports the refusal
on the barrier as tunnelFailure (code and site, nothing else) and on the
application pass as activation.tunnel:<label>, and the next kept pass asks
again. Once the command is admitted every later step is a journaled transition
and fails the activation as such.
A tunnel lease lasts at most fifteen minutes and is moved in place, never by
tearing the tunnel down: every admitted DNS lease renewal moves the window the
credential lives in, so the next kept pass asks the controller for a fresh
credential, and a pass that finds the lease deadline within three minutes of the
qualified clock does the same (renewalMarginMs). Under the identity the session
authenticated with (hub address and key, client key, authority and node) both
bounds on the device move in place: first Pallet's own native authority deadline
(PalletHandoffSetOwner.renewActiveTunnel, native renewActiveTunnelLease), which
the native arms at the lease the device was created under and at which it fails
the generation closed, then the session's lease with SmartVPN's
renewManagedSession (PalletManagedVpnOwner.renewLease). The device, its packet
policies and forwarding are untouched, and nothing is journaled: the deadline
bounds this process's custody, and a successor fences the generation without
it. A credential under another identity cannot renew a session, so the tunnel is
lowered and a new one raised with it. A refused renewal leaves the lease in force,
names the refusal on the barrier's tunnelFailure and is asked again on the next
pass; a refused native renewal never asks the session, so the session's lease
never outlives the native's.
A lease that lapses anyway ends the native authority at the same instant: the
native fails closed, the generation is fenced and the node fails owned, and its
supervisor starts a fresh process, which fences the generation by the proof that
its epoch ended and raises a new one with a fresh credential. A session whose transport to the hub is lost stops
instead: SmartVPN ends its packet I/O and reports managed-session-state
stopped, keeping the device for an ordered teardown. The barrier stops holding at that moment —
holds requires the session's packet I/O to be live — so attached workload links
demote on their next CHECK and admission refuses, and the next node scan runs a
network pass at once. That pass lowers the dead tunnel through the journal exactly
as a withdrawal does (the lowered packet pair, the delete command, the device
released) and closes its owner, then asks for a fresh credential and raises a new
tunnel over the same generation; a refused credential leaves the generation UP
without a tunnel, named, until a later pass succeeds. A retired session — its
device lost behind the owner — is a loss of the owner, not a stop, and fails the
activation as before.
PalletNetworkApplicationOwner takes an optional activation option and, when
it is present, drives one ACTIVE generation after each completed dormant
application: preparation, both packet policies applied and journaled, native UP,
local DNS proven to be serving, then the tunnel commands above and finally
forwarding. Without that option the owner stays exactly as dormant as before and
the DOWN wire contract is unchanged. PalletNetworkOwner.barrier() exposes the
result as one of absent, prepared, up or forwarding, with the retained
device's journal reference when there is one. Only forwarding sets holds, and
only holds may gate workload UP; a node with no managed VPN configured settles
at up and its barrier deliberately does not hold, with one exception: the
single host. A node bound to a local controller (Onebox) never dials a tunnel,
whether or not a managed VPN is configured — the released bundle configures one on
every node, and the persisted binding decides — and a generation reaches
forwarding without a tunnel when three facts hold: its scope is that local onebox controller, every endpoint of the signed
projection it was raised from is placed on this node (the contract already
refuses a remote endpoint in a onebox projection; Pallet proves it again from
the journal row and fails closed), and this process raised the generation itself.
There is no remote peer to reach, so the base packet policies already carry every
rule the generation needs. A barrier that does not hold names the missing fact as
singleHostGap (controllerNotLocal or remoteEndpoint). Execution admission
then rests on the single-host evidence instead of a tunnel receipt: the evidence
names the projection, the admitting transaction re-proves the fact from the
journal row and fences it, and a refusal is admission.barrier.singleHostUnproven.
A tunnel command on such a generation is a corruption and refuses. Every other
node — bound to Cloudly directly or through the relay — keeps the mandatory tunnel
barrier unchanged. The barrier is always
recomputed from the journal and the live owners, never read back from a stored
pass. On loss this composition fences and withdraws first and then reports a
barrier that does not hold; it never re-activates the lost generation by itself,
because re-activation requires explicit newer state from the control plane.
Durable application intent and results
The private identity runtime exposes runNetworkApplications() over its existing
SmartData/NoSQLDB owner. readSource(scope) returns the current verified signed
projection and complete retained lease history. Native preparation happens outside
the database. begin(scope, fingerprint, prepared, guard) then persists that exact
whole-set target, its previous receipt and the authenticated source, writing the
common trust/allocation fences and the real identity guard in the same transaction.
Only the current reserved handoff is desired present; every other retained lease
remains explicitly absent. Omitted history and revived absent members are rejected.
Ordinary successors remain within the same boot and native namespace.
The journal keeps one pending transition and immutable intent/result records.
Exact retries join existing intent; a competing proposal cannot replace it.
finish(scope, intentReference, nativeReceipt) records only the result matching
that persisted target. It does not require new admission authority: signing-key,
projection and credential changes must not discard an already-owned native result.
Historical signatures remain verifiable through retained signing revisions. Lost
completion acknowledgments replay exactly without advancing or rolling back the
current head. No journal method deletes history or frees quarantined allocations.
Keep begin, native reconciliation and finish inside one admitted runtime callback
so database shutdown joins the entire operation. Retained facade methods expire
when their callback returns. The persistence capture budget covers both transition
sides and complete signed source; native frames retain their independent 2 MiB
limit. This store does not start a native owner, issue nativeBarrier, report
protection, clear DNS withdrawals or authorize workloads. The private application
owner supplies node composition and a separate durable unavailable-report outbox.
beginEpoch(scope, fingerprint, prepared, evidence, guard, workloadProofs) records a qualified
same-boot absence or different-boot fence into a fresh namespace. The caller must hold the
fresh private native owner continuously through its audit, journal commit and
reconciliation. The store resolves the evidence's opaque journal reference to
the actual persisted pending intent, or applied intent when none is pending,
and binds its complete target and exact known receipt. A single transaction
retains the old applied/pending references in an immutable epoch record, fences
current signed authority and identity, and publishes the next intent. Evidence
from a different target, previous receipt, host, boot or namespace is rejected.
The first intent in an epoch has previousEpoch and no native predecessor:
reconcile it with previous = null. Its generation follows the old pending or
applied head. The still-reserved logical handoff may be realized DOWN again;
every retained lease and quarantine survives. Old unfinished intents stay in
history without a fabricated result and reject late completion after the epoch
fence. Previously recorded results still permit exact acknowledgment replay.
An intent has one shape: it names the epoch it opens or null, and a record
that does not is refused rather than read. Repeated failures before a fresh
intent completes continue the same history.
Different-boot recovery requires the distinct native previousBootTerminated
proof, persisted as new-boot-host-fence; it cannot use a same-boot absence
receipt. Reused numeric namespace identities do not make two kernels the same epoch.
Dormant workload attachment journals
The private ts/network/attachment/ composition binds one admitted pending Run
to its immutable workload lease, completed application and exact container ID.
Native preparation opens the sandbox namespace once and retains its descriptor.
The ADD journal commits before any veth mutation; its immutable identity cannot
be rebound after cleanup. Ambiguous admission acknowledgements are resolved on
the same application transaction lane before an unadmitted descriptor is discarded.
Partial creation uses the recorded removal intent and exact native absence check.
The sandbox accepts containerd's internal loopback baseline. Workload links stay
DOWN until this node's activation barrier holds; an ADD never waits for it.
Application and attachment operations share one queue. Handoff retirement must fence every unfinished attachment in its transaction. A replacement router waits for distinct historical workload evidence: stable absence in the exact retained sandbox on the same boot, or termination of a different actual kernel boot. A missing or substituted sandbox path is not same-boot absence. Previous interface indices belong to their original namespaces and are never host capabilities.
An application epoch record has one shape and always binds a complete sorted workload proof set. Immutable coverage records commit with that epoch and freeze the old journals. Coverage permits historical DEL to acknowledge the proved absence, while preserving the original unfinished journal, all allocations and quarantine. It does not create an ordinary DEL receipt or authorize another ADD. Quiescence stops preparations; already admitted ADD/DEL operations drain before the shared native owner closes.
Production execution admits network: 'attached' only after the assignment's
exact projected lease is prepared with the pending Run operation. The CNI owner
then proves and journals its attachment before returning a usable result. An
isolated assignment prepares a separate durable CNI intent with no lease or
application reference. A secret reference is resolved by this node: its material
is staged and published as this attempt's own mount before the container that
receives it, as "Staged secret material mounts" states. A storage reference is
granted from this node's claim ledger in the transaction that begins the run, as
"Local storage claims" states.
Private CNI transport
The node runtime owns a root-private CNI broker at
/run/serve.zone/pallet/cni.sock. The installer supplies the trusted
/run/serve.zone parent; the broker verifies its ancestry and provisions the
0700 leaf and 0600 socket. The native executable enters CNI mode through
--cni or the installed basename pallet-cni. The client checks path ownership,
socket identity and the kernel-reported server UID before sending one bounded
frame. It never stores journals, allocates addresses or opens sandbox namespaces.
Transport cancellation does not cancel an admitted attachment operation. Detached handlers retain broker capacity until their native effect and journal write finish. Node shutdown joins CRI, including synchronous rollback DEL, before closing the broker; the broker then joins all admitted handlers before its borrowed owners can close. Replaced socket paths are never unlinked during cleanup.
The fixed profile uses CNI 1.1.0, the single pallet-workloads configuration and
containerd's use_internal_loopback=true. VERSION reports protocol support. DEL
can acknowledge exact recorded absence with empty stdout.
For an isolated assignment, the pending Run transaction durably binds its
assignment, operation, boot and fixed CNI configuration. ADD opens the exact
containerd sandbox namespace and proves it is distinct from the host and router
namespaces. Containerd 2.3 requires a real eth0 with an address even for an
isolated pod, so the native owner creates a namespace-local Linux dummy device
with 192.0.2.1/32. It has no peer, uplink, default route or host grant;
IPv6 address generation is disabled and every address, neighbour and route in
both families is checked. The kernel's own ff00::/8 local-table multicast
route is allowed only when it is bound to this peerless dummy, with no IPv6
address, gateway, RA or unicast route. The first ADD commits the exact
container ID, namespace identity and double-proven initial absence in SmartData
before NEWLINK. The completed ADD then records the dummy index, deterministic
MAC and double native proof. If a process stops between NEWLINK, alias, address
and UP, replay may remove only the exact DOWN partial dummy under that durable
intent, prove its absence twice and recreate it. Ambiguous or foreign state
remains quarantined. CHECK requires a completed journal and reproves its exact
endpoint. The CNI result reports the real dummy and address,
with no routes or DNS grants. DEL verifies and removes that exact device before
retiring the journal; replayed DEL is idempotent. No workload lease, veth or
activation grant is made for this mode. Native CRI evidence must report only
192.0.2.1 for an isolated running sandbox. Host-side TCP and HTTP readiness
probes are refused before effects because they would target host services.
ADD never blocks on the activation barrier. Without it the attachment is still
admitted, journaled and reported DOWN with an activation error, and the runtime
retries; with it the held generation's policies are amended to carry the pair
when they do not yet, and the pair is raised and proven in the same step. A
refused amendment or raise leaves the attachment admitted, journaled and DOWN
and names its step on the pass (amend.<code>:<site>, raise.<code>:<site>);
it never fails the ADD. A successful ADD returns
a CNI result carrying the interface, its address, the node resolver and the one
route the raise installed: 0.0.0.0/0 via the router side of the pair
(static, priority 100). Only state this owner installed is reported, and the
advertised route is proven equal to the route found in the pod namespace.
CHECK reports the raised attachment while the barrier still holds and
demotes to the same DOWN activation error the moment it is lost, without
touching kernel state. A repeated ADD on an already raised attachment is
idempotent and never destroys it. STATUS reports authenticated broker and network
owner availability to containerd, independently of any attached activation
barrier; it cannot grant a workload network. An attached ADD or CHECK without the
exact live barrier still reports DOWN, while an isolated ADD uses only its own
pending Run preparation. GC cannot infer
cleanup authority from a runtime-provided attachment list and returns an error
until an admitted removal or historical proof can be resolved.
Application owner and node lifecycle
PalletNetworkApplicationOwner holds one private handoff-set owner for the node's
router namespace. Each bounded pass keeps journal inspection, completion of a
retained current-epoch intent, current source selection, epoch qualification,
admission, native effects and completion inside one identity-runtime callback.
It derives complete membership from the authenticated store; RPC callers cannot
provide a prepared target, native receipt or absence proof. A stale source or
identity rejects new admission. Completion of an already admitted effect survives
credential changes, controller disconnect and quiescence.
The private PalletNetworkOwner composes this application owner with the guard
and protection outbox. It restores the existing offline guard before starting
the application or controller. While connected, the node reconciles the network
with the actual physical session's private identity guard before scanning
assignments whenever a network pass is due: the first scan, every scan after new
network input from the controller (a new connection, an admitted signing
authority, projection or DNS lease renewal), after a rejected pass, as soon as a
managed VPN tunnel stopped on its own, and otherwise at most networkRefreshMs
apart (default 30 s; the assignment scan itself runs every pollIntervalMs,
default 1 s). A full pass re-proves the guard with a fresh engine process and
re-reads and re-verifies the application and its signed source, so it does not
run on every scan. The standing DNS refresh runs every 30 s for the same reason;
every change DNS depends on (an application pass, a workload ADD or DEL)
refreshes it at once, and the resolver child enforces its lease's BOOTTIME expiry
itself. The protection delivery loop decides that nothing is pending from its
lane row alone. Without current
authority it waits; there is no offline network admission path. The status snapshot
reports the application reference and admission rejection without exposing source
records. Native ownership loss or an unrecordable native result fails the node
with ownership retained for cleanup.
Shutdown stops admission, joins any pending begin/effect/finish callback, then joins the native handoff owner before closing the namespace and database. An unconfirmed native join keeps those resources owned; failed cleanup cannot report a clean stop.
A normal stop (SIGTERM or SIGINT, or stdin EOF for runtime-serve) keeps an
ACTIVE generation UP in the journal for the next process to fence, and lowers
only its managed VPN tunnel. The native refuses to withdraw a generation that
still holds a device, and a device owner joined without a begun deletion takes
its device away behind the native, which retires on that unarmed deletion;
either way the handoff set's release stays unconfirmed. So the orchestrator's
close refuses every new command and joins the ones it admitted before; then it
lowers a held tunnel exactly as a withdrawal lowers it — forwarding
stopped, the lowered packet pair applied, the deletion journaled and begun, the
device owner released, the deletion finished — and only then joins the device
owner. The native's closeHandoffSet then withdraws the generation inside its
own router namespace and releases the set, and the process ends with stopped
and exit code 0. The journal keeps the generation UP with its tunnel lane
settled on the deletion, and the kernel keeps the host table at the lowered
pair: exactly what the next process's epoch fence proves ended and releases (the
retained-generation paragraph under "Standalone control build"). The application
owner joins every workload ADD and DEL before that close, since the lowering
runs outside the lane an ADD's amendment and raise run on; so no other journal
or packet work interleaves with the lowering. A stop over a generation with no
tunnel lowers nothing. A lowering that fails is an owner failure: the stop still
joins every owner, fails with the failed step as its cause and ends
owner_failed, and the successor fences the generation as after a crash.
Without the activation option all handoffs remain DOWN. With it
the application owner opens the physical DHCP uplink observation before the
orchestrator and raises one ACTIVE generation over it on the first pass that has
an application; a generation that is UP for the same application over the held
observation is kept on later passes, an observation that no longer proves itself
withdraws the generation bound to it and is reopened on the next pass, and a
failed activation is fenced and named on the pass (activationFailure). The same
label is written once to stderr, which Spark carries into the node journal, as
Pallet network activation refused: <label>; a packet engine refusal adds the
engine request and hop (preparePolicy, host or router) and the engine's code
and bounded refusal text. A refusal repeated on every pass is written once, and a
refusal that returns after activation held is written again. Without
a managed VPN the generation settles at UP and the barrier deliberately does not
hold, so attached workload links stay DOWN — except on a single host bound to a
local controller, whose barrier holds as described above whether or not a managed
VPN is configured. Lease release and quarantine remain
with Cloudly's allocation owner; Pallet resolves only an exact assignment-bound
projected lease and cannot create or release allocation authority.
Durable protection and withdrawal reports
Every connected network pass freshly recovers the actual allocation-pool guard, completes the dormant application journal, then attempts a positive receipt. Only a real guard owner that has reconciled, inspected enforcement, detached its PERSIST policy and joined its native child can supply that private capability. The returned bootstrap JSON and a durable guard reference alone cannot mint it. The capability expires when the owner closes; composition keeps it through the outbox transaction. RPC never accepts a native proof or guard selector.
Fresh positive minting fences the settled current guard and application lanes,
current signed projection and signing key, retained allocation ledger and actual
physical reporter/identity in one SmartData transaction. Guard and application
must agree on source, boot, host namespace and protected authority. The report's
nativeBarrier derives from that exact completed guard intent. It establishes
allocation-pool denial, without claiming workload connectivity, DNS readiness,
packet drain or lease release. Every handoff remains DOWN.
Every positive report also states the node's own uplink address — the address
a route to its published ports uses, read from the retained uplink observation
— on the request envelope (uplinkAddress, @serve.zone/interfaces 31.4+).
The address is a fact of the pass that minted the receipt: a changed or absent
address is a new receipt generation, so the pass after a rebind reports it
absent (the stale observation is closed) and the next pass, over a fresh
observation, reports the new one. Nothing defaults or infers it, and a
withdrawal states none.
Beside it the report states hostAddresses (@serve.zone/interfaces 32.23.0):
every address the same observation binds on the uplink — the lease address and
the permanent /32 host addresses beside it — sorted as strings. Cloudly names a
plan platform endpoint in this node's hostPlatformEndpointIds when its address
is one of them. They are read together with uplinkAddress from one observation
and are part of the receipt's identity in the same way: another set is a new
receipt, and a report without an uplink address states none. Only the uplink's
addresses are stated, because the packet compiler serves a host-local platform
endpoint only on an address the uplink binding carries. An outbox entry a 33.0.0
node wrote carries none and is read as stating none.
The SmartData outbox retains immutable receipt revisions, original reporter bindings and completed-application references. Positive reads verify the historical completed guard too. One pending receipt supplies backpressure; a newer application cannot replace it. Historical replay uses those immutable proofs without requiring old guard/application heads to remain current. Exact ACKs fence the current physical connection and identity. Reconnect sends the original receipt body, then a later fresh native pass may queue a current-session successor. Replaying historical positive evidence does not re-establish current-session eligibility.
Quiescence stops admission and joins the native pass, drains and ACKs any pending predecessor, then withdraws an acknowledged positive with its exact null successor. The null retains its predecessor's historical application, authority and boot, so withdrawal works during key/projection gaps. Its reporter is the current physical session. The exact null ACK must be durably recorded before clean node shutdown. A failed ACK leaves the immutable pending history and enforced PERSIST policy intact, joins the child owners, and fails the process shutdown result. Physical disconnect alone does not remove Cloudly's durable allocation eligibility.
Offline allocation-pool guard journal
PalletNetworkGuardOwner composes the private identity database with published
Smartnftables 2.1. commission() explicitly creates the first guard intent;
recover() requires an existing journal. The caller supplies a trusted native
binary path and must join this owner before closing the database. The node runtime
composes this private owner with application completion and positive reporting;
Spark installs the separately commissioned boot dependency.
The fixed receiver-owned trust record supplies the node scope, checked against
the active local identity. Bootstrap verifies the latest admitted projection with
its retained signer, including a crash between key rotation and the next signed
projection. That historical material supplies denials only; DNS and new projection
admission still require the current key. The policy denies exactly the declared
allocation pools, without treating management LANs or resolvers as allocation pools.
Its only exceptions are the exact host grants of the source's signed onebox
projection. They are derived from that verified projection whenever a target is
proposed and proved again from the retained source on every read of an intent, so
a stored row never widens the guard by itself. Every policy also restricts the CRI
stream server of Pallet's own containerd,
127.0.0.1:10010, to uid 0 (localTcpPortOwners, Smartnftables 4.0.0). A pass
refuses unless the policy it applied carries that owner; a guard committed before
the owner existed gains it through one appended transition of the same source.
SmartData persists each complete native identity and original previous-Applied/null
to prepared-target transition before reconciliation. Immutable result records bind
the full native receipt to that intent. One pending lane prevents competing targets.
Same-boot recovery repeats the original transition, including after a lost reply or
completed result. A different actual boot records the old applied/pending heads in
an immutable epoch and restores the prior denial with a fresh process instance and
previous: null. A changed namespace in the same boot is rejected.
Recovery restores the committed denial before applying a newer inert projection.
Every previously guarded pool must remain with the exact same id, purpose and
prefix. Native preparation must reproduce the durable target, and exact enforced
inspection and verified detachment must complete before success. An ambiguous
failure uses closeRetaining() to join the child without deleting its policy or
inventing an Applied receipt. No journal result alone grants allocation eligibility.
test/native/qualify-guard-boot.py runs real NoSQLDB and nftables in four isolated
offline root boots under Node and compiled Deno. It covers same-boot completed and
pending recovery, lost/undelivered native calls, new-boot epochs, signing-key gaps,
newer staged pool expansion, and actual pool denial with unrelated traffic controls.
On both boots a root TCP connection to 127.0.0.1:10010 succeeds and the same
connection as uid 65534 is refused, before and after the new-boot recovery.
It also exercises PalletGuardProcess and reopens the same database after its
oneshot completes. Supplying --control-directory dist_control/linux-amd64-<digest>
also runs the built production pallet-control guard-recover with null stdin and
verifies its completed events, database release and continued pool denial.
Production still requires the commissioned Spark boot dependency, and live
activation the qualification of the complete workload networking, DNS and VPN path
on a real host.
Host kernel requirements
The packet engine (Smartnftables 4.0) requires Linux 6.9 or newer, with the
nftables filter and NAT modules (nf_tables, nft_nat, nft_chain_nat, nf_nat
and the reject modules the port owner rule uses) available. A kernel below that
floor, or one lacking a feature a policy needs, is refused by the engine as
UNSUPPORTED_KERNEL; Pallet names that refusal instead of hiding it. The guard
owner and PalletGuardProcess fail with code unsupported_kernel, and
guard-recover, guard-commission and runtime-serve end with failed reason
unsupported_kernel on their process protocol, exit code 1. Every other engine
failure stays owner_failed. A host must boot a supported kernel before Spark
commissions the guard.
PalletGuardProcess verifies the guard executable against trusted bundle metadata,
opens only existing enrolled storage and runs one offline recovery or commission.
It joins the native owner before stopping NoSQLDB. A native cleanup failure keeps
the database lease owned for an explicit close() retry. Cancellation joins
admitted work and fails the oneshot; success follows verified policy retention,
native child exit and database shutdown.
The compiled control accepts guard-recover for ordinary boot and
guard-commission for explicit first commissioning. Both accept exactly one mode
argument and no runtime configuration. Null stdin is valid for these oneshots;
unexpected input or SIGTERM/SIGINT fails without readiness. Successful completion
emits ready and stopped under pallet.guard.process after cleanup. These
events provide no positive allocation or workload protection receipt. The installer
must acquire its first authenticated projection and commission before enabling
recovery as a boot prerequisite.
test/native/qualify-protection-boot.py separately qualifies the complete private
network owner in four offline Linux6.18.35 boots, under Node and compiled Deno.
It verifies actual nftables denial with independent UDP controls, real dormant
handoffs, completed application/guard agreement, historical positive replay after
reboot, a fresh new-boot positive, null withdrawal through signing-key advancement,
and PalletNodeProcess cleanup that reopens the same database. The packet client
uses the pinned Node executable for both runtimes so a Deno UDP compatibility
error cannot be interpreted as enforcement. No kernel, native-owner or database
result is stubbed in these guests. They have no network device or host mount.
Supplying --control-directory dist_control/linux-amd64-<digest> also runs the
built production pallet-control runtime-serve on each runtime's second boot.
It verifies readiness, EOF-driven shutdown, clean child exit, database reopening
and continued pool denial after the entire node owner has stopped.
Initial projection acquisition
pallet-control network-acquire opens only existing enrolled storage and a
projection-only controller connection. It waits for signing authority and a
projection delivered over that actual physical session; a retained older projection
cannot satisfy acquisition. It rejects assignments and does not fetch registry
credentials or deliver terminal, observation or protection reports. After admission
stops, it joins every controller write, verifies the delivered projection under
the final current key, and closes controller/storage before ready and stopped
under pallet.network.acquisition. Null stdin is valid; input, signals or the
bounded operation timeout fail the oneshot after joining its owners.
pallet-control network-acquire-local is the same acquisition for a node driven
by a controller on its own host (Local controller). It serves
the local socket instead of connecting to a relay, and it additionally admits the
first bind of a node no controller has bound yet; a node with any Cloudly
enrollment state refuses to start it. The store must already be provisioned.
pallet-control store-provision provisions it: a null-stdin oneshot on the
pallet.store.provision protocol that creates the data root and the database,
prepares every collection, runs the store's migrations and releases the database
again. It writes no enrollment, identity, binding or authority record and serves
no socket, so the node stays unbound until network-acquire-local admits its first
bind. An existing store is accepted as it is; a store that carries a Cloudly
identity is refused (PalletStoreProvisionError:cloudly_enrolled on the
owner_failed diagnostic line). enrollment-provision also creates the store and
leaves no Cloudly state behind by itself, but it serves the Cloudly enrollment
socket for its lifetime, where a prepare request writes that state; a node for a
local controller is therefore provisioned with store-provision.
The installer runs acquisition while management networking is available and before
first guard-commission and boot-unit installation. Existing guard recovery remains
offline and requires an existing journal; it never silently commissions a new one.
Previous-epoch host attachment audit
A fresh PalletHandoffSetOwner can call verifyPreviousAbsence({ journal, target, previous }) before it applies a current set. Supply the authenticated persisted
application reference, exact old prepared target and its prior applied receipt,
including an unfinished transition's complete retained membership. Native code
recomputes both old identities and their transition relationship, requires the
same boot and host namespace, and retains the shared host mutator lock. The
journal reference is an opaque binding; native code does not authenticate it.
Two complete host inventory passes reject old host or misplaced router names and MACs, Pallet markers, retained host indices and transit address/route conflicts. Old router indices remain scoped to the old namespace. The fresh router must independently remain empty and DOWN. Any candidate or inventory loss rejects the audit; it performs no cleanup, repair, deletion or adoption. Persist the returned old/current namespace-bound evidence before permitting a new epoch's effects.
The evidence proves old host attachments absent under exclusive privileged
ownership. It does not prove physical peer destruction, packet drain, allocation
reuse or a current nativeBarrier. Kernel peer teardown can be asynchronous, and
another privileged writer can invalidate negative observations. Full node
recovery uses the application owner's journal and complete authenticated history;
process death or a saved namespace identity cannot replace this audit.
The isolated Node/Deno fixture covers surviving and renamed old attachments,
reused host indices, unrelated host indices matching old router indices, transit
conflicts, fresh-router state, old request tampering and wrapper lifecycle.
verifyPreviousBootFence(request) is the separate fresh-owner capability for a
different kernel boot. It validates the exact historical target and prior receipt,
requires different valid kernel boot UUIDs, and reads the current boot from the
native owner. Two complete audits still reject current old identities, Pallet
markers, transit conflicts and nonempty fresh-router state. Old interface indices
are deliberately ignored: those numbers can legitimately belong to unrelated
interfaces after reboot. The returned evidence binds both epochs and adds
previousBootTerminated: true. It supplies no allocation release, packet drain
or nativeBarrier; the dormant-only history and privileged ownership requirements
still apply.
The offline test/native/qualify-handoff-boot.py fixture runs two boots each under
Node and compiled Deno with independent disposable ext4 NoSQL roots. The actual
application owner persists real signed authority and applied/pending history. A
fixture-only interrupted completion retains the acknowledged DOWN target, then
the fixture flushes the database
then signals init through a FIFO while the native owner and DOWN pair remain live.
Init forces poweroff. The next boot reloads that journal, qualifies reused host
indices and conflict rejection, then the application owner persists the boot fence
and realizes the same still-reserved lease DOWN. All application state uses SmartData/NoSQLDB; no journal
sidecar or host network/disk attachment participates in the qualification.
It also retains the exact unavailable report across poweroff, replays its original
reporter binding, and queues the next boot's report only after acknowledging history.
Standalone control build
The control builder requires Deno 2.9.7 and the exact NoSQLDB 10.5.1 engine identities
in binary/control-build.json. After the normal dependency install, build either
Linux target or omit the target to build both:
pnpm run build:control linux-amd64
pnpm run build:control linux-arm64
@git.zone/tsdeno 1.14.0 compiles with the committed, frozen deno.lock under its
managed runtime-only manifest, and per target leaves out the npm native binaries built for the other
architecture or for macOS (their ELF/Mach-O headers prove them foreign); the control never loads them,
because it runs its own sibling engine, guard and DNS executables. TsDeno refuses any Deno other than the
pinned denoVersion, fails a pallet-control smaller than the target's minSize in
binary/control-build.json (200 MiB; a compile that lost its npm payload still exits 0 with a far
smaller binary), and on the host's own architecture runs it once without arguments in an empty
environment, where it must print its invalid_arguments refusal and exit 1; the other
architecture's smoke check is skipped. The build verifies the selected NoSQLDB engine's
bytes and clean owning-build provenance. It also builds and verifies the Pallet
static-musl executor against this exact source, version and architecture. The
Smartnftables 4.0.1 guard is verified against its pinned bytes and clean provenance.
Each completed dist_control/<target>-<digest> directory contains pallet-control,
its fixed siblings pallet-smartdb, pallet-runtime, pallet-guard, pallet-dns
and pallet-vpn, the pallet-containerd/ release directory (containerd, its shim,
runc, the pause archive and their manifest), their native provenance records and a
control-build.json recording artifact hashes, source state,
compiler versions and runtime lock hash. The compiled control keeps UID 0 and its
protected production data/socket paths; it derives only the sibling engine path
from its own executable. The installer owns protected platform-parent directories.
Packaging includes every native guard notice from the pinned upstream manifest
(the native-notices/manifest.json that tsrust notices generates, with the digest of every file)
under notices/smartnftables/ and rejects changed or missing material. The archive
also includes the exact pallet-clock/ loader, programs, shared libraries and
trust files, plus full clock source archives, patches, recipes and notices under
sources/clock/ and notices/clock/. Consumers must verify the complete inventory.
binary/vpn-engine.json pins the published SmartVPN 2.4.1 musl executable,
provenance, package license and native notice inventory for each architecture.
The builder and packer verify the complete distributed VPN notice set under
notices/smartvpn/, including the upstream Rust and Cargo license texts. A changed
binary, symlink, missing notice or mismatched build is rejected. The bundled
pallet-vpn is the executable runtime-serve verifies and runs as the managed VPN
device owner (see "Foreground node process").
The Chrony source build additionally requires a Docker client/Buildx and a Linux Docker daemon able to execute the selected pinned Alpine architecture. Its build context and result archive use the daemon API; it requires no host workspace bind mounts. Compiler execution is offline, unprivileged and capability-free. The build checks pinned source, SDK and output hashes; workers receive only the finished programs and distribution materials and perform no package installation.
Every change to the dependencies in package.json moves the control graph, so it
also moves the lock. Rebuild dist_ts, generate the lock with pnpm exec tsdeno install --entrypoint --lockfile-only --frozen=false --lock=deno.lock dist_ts/control/process-cli.js, and review the lock diff. Then carry
noticeLockSha256, the control notices and their pinned asset hash in
binary/control-build.json onto the new lock. test/test.controllock.node.ts runs
the frozen resolution the control build uses, through scripts/control-lock.mjs,
and fails while any of them lags. Ordinary control builds leave the lock frozen.
The lock is registry-neutral: its npm entries carry no tarball URL. Deno writes
one only when the registry it resolved a package from is not npmjs, and from Deno
2.9.7 it refuses a lock whose tarball origin is neither the configured registry nor
npmjs. The tracked .npmrc therefore sets registry=https://registry.npmjs.org/.
Deno and pnpm both read it from the project root and prefer its registry over the
one in ~/.npmrc, so a lock generated on a machine whose home configuration names
a mirror still resolves from npmjs. Deno's NPM_CONFIG_REGISTRY environment
variable overrides the project .npmrc; scripts/control-lock.mjs refuses to run
while it names another registry, and refuses a lock that carries any tarball
field.
The earlier Linux amd64 artifact with SmartDB 5.8.0 was qualified as root in an isolated QEMU guest with a disposable local ext4 disk and no network or host mounts. Checks cover the fixed production paths, engine digest and permission guards, explicit provisioning, durable enrollment replay after process restart, EOF/signals, native and parent SIGKILL, and stale-socket recovery. NoSQLDB rejects volatile filesystems; a tmpfs data root cannot substitute for supported local storage. This is process recovery qualification, not a power-loss or ARM64 runtime claim.
The NoSQLDB 10.5.1 engine passes the local native identity persistence and
lifecycle suite with SmartData 11.14.2. Pallet now reaches it through the nosqldb family of
@lossless.org/client 1.5.1, which continues SmartData 11.14.2 with the same persisted format. Both engine targets are the published static musl executables;
the Deno control executable still targets GNU Linux. Both control targets compile
and package with their verified owner provenance and complete upstream notices.
The updated amd64 control bundle also passes the root guest checks above on an
isolated local ext4 disk, including durable enrollment replay after native and
parent process crashes. ARM64 runtime and power-loss qualification remain open.
test/native/qualify-node-process.py also qualifies a forward upgrade when given
both --previous-control-directory and --previous-probe. It boots that earlier
bundle first, retains a read-only copy of the stopped ext4 disk, then boots the
new bundle twice against the working disk. Each boot verifies its exact binaries
and probe; the final result compares the node identity and signed network
projection and verifies the backup is unchanged. The NoSQLDB 9.0.0 to 10.2.0
amd64 run passes all three offline root boots, including retained workload leases,
signature replay, namespace keeper loss and joined process shutdown. This
qualifies existing current-format Pallet state; it does not qualify containerd,
private DNS, physical power loss or foreign legacy database roots.
The guest stages the build's complete runtime inventory (executor, guard, resolver,
clock loader and assets) and the nftables and veth modules, enrolls with a routing
binding and commissions the allocation-pool guard before runtime-serve, which
now owns the node's network; the guest provisions the root-only runtime directory
on every boot, as the installer does. A runtime-serve whose namespace keeper is
killed reports failed and exits: when its close cannot confirm the network
release, the node process joins the clock, the namespace and the database anyway
(PalletNodeRuntime.abandon) instead of holding them for a retry nobody in the
exiting process makes, and the successor recovers the unconfirmed state by
evidence, as after a crash.
The database package is @lossless.org/nosqldb and the controller uses its
NoSqlDbServer API. Existing smartdb storage-directory and bundle field/file
identities remain stable; they do not select the deprecated npm package. Changing
the package does not convert an incompatible database format. NoSQLDB 10 reads
existing version-9 current-format stores directly; once it writes version-10
records, older engines cannot read them. Retain a stopped pre-upgrade backup
before activating the new engine. Native current-format admission still rejects
foreign legacy and mixed roots before application startup.
Pallet admits the released 32.0 attachment-journal shape before any network
owner starts. Terminal absence and an already durable pending removal retain
their original bytes, digests and ACTIVE source or amendment history; the latter
may finish with its released removal formula. The admission writes only the
pallet_migrations completion row and is replayable after interruption. A
released running or pending-ADD row has no durable removal transition that this
version can complete without rewriting its history, so startup fails with a
pallet-attachment-owner-recovery-* error and writes no completion row. Stop the
upgrade and retain the prior Pallet owner to finish that lifecycle before trying
again. There is no collection-reset or cleanup command in this recovery path.
Every node process owns its own router namespace, so a restarted process runs in
a new epoch, and so does every process after a reboot. A generation the previous
process retained — UP, or DOWN with a packet cleanup it never finished — is held
by no native of the new epoch, and the native admits RecoverActive only for a
generation of its own. The restarted process therefore fences it before its
native adopts anything, with the audits that fence the application journal's own
epoch, over the generation's application target and receipt: in the same boot,
verifyPreviousHandoffAbsence proves the previous epoch's host attachments
absent (the router namespace, and every pair, device and route under it, ended
with that process); after a reboot, verifyPreviousHandoffBootFence proves the
boot over. The proof is journaled as a DOWN of kind epoch (same-boot-host-absence
or new-boot-host-fence), taken from the fencing native's epoch
(PalletHandoffSetOwner.readEpochFence), and its packet cleanup releases what
outlived the epoch: in the same boot the host table, through the running engine
with the generation's retained host identity; the router table ended with its
namespace, and after a reboot no table is left. The successor is raised on a pair
keyed anew. A generation of the process's own epoch is still withdrawn in order
while the process holds it — a tunnel transition left pending by a failed step is
finished first — and recovered by RecoverActive otherwise.
Every ACTIVE generation records the packet engine that prepared it, the SHA-256
of the Smartnftables executable the node process verified before it started
(packetEngine; generations journaled before 35.0.0 carry none, so their engine
is unknown). The kernel keeps a generation's host table across a Pallet restart
in the same boot, and the restarted process asks the running engine to release it
once the generation is fenced. An engine whose compiled graph differs from the
one that applied it cannot adopt it and answers Conflict; Smartnftables 3.0
changed the graph of every routerEgress table and of every scope with
publications. When that Conflict comes over a table another or an unrecorded
engine applied, the release fails by name (PalletPacketEngineChangedError,
conflict:packet.retainedByPreviousEngine) and the network owner stays failed
and fenced, because no retry of the running engine can change the answer. The
remedy is a reboot, which takes the kernel tables with it and lets the new-boot
proof fence the generation, or finishing the generation on the previous Pallet
before upgrading again. An upgrade that changes the packet engine across a
same-boot Pallet restart therefore needs a reboot whenever an ACTIVE generation is
retained. A Conflict under the engine that applied the tables is not this refusal
and fails as before.
Local build outputs are qualification assets; a dirty source marker is recorded
explicitly and cannot identify a release. The tag-triggered workflow publishes
clean-source control bundles as inputs for Spark integration.
The bundle's runtime-serve activates the node network with the bundled
pallet-vpn (see "Foreground node process"); running it still requires Spark's
whole-bundle installation and supervision. The complete activated path — ACTIVE
generation, managed VPN tunnel to a real cluster hub, lease renewal and workload
traffic — is qualified in parts (the guests below and unit specs), not yet end to
end on a real host. ARM64 runtime and workload execution are not yet qualified by
this component release.
TLS trust of the control executable
pallet-control verifies every TLS peer it dials — the controller or Cloudly origin
of the enrollment, runtime and network-acquire sessions — against the Mozilla root
bundle compiled into Deno and then the operating system trust store. A node whose
controller, relay or Cloudly certificate is issued by a private or internal certificate
authority trusts it the way every other service on the host does: install the CA
certificate into the OS store (update-ca-certificates, update-ca-trust). Pallet has
no trust setting of its own and never disables verification.
A compiled Deno executable carries no store selection and otherwise trusts the Mozilla
bundle alone, so runPalletControlCli sets DENO_TLS_CA_STORE=mozilla,system for
every mode, after the argument check and before any mode starts; Deno reads the variable
once, when the process opens its first TLS client. An operator who sets
DENO_TLS_CA_STORE explicitly keeps that choice, provided it lists only mozilla and
system; any other value, including an empty one, stops the process with one
failed line of reason invalid_ca_store on the mode's protocol and exit code 1,
rather than failing every later connection. Spark starts pallet-control with a cleared
environment, so its processes always run with the default.
The managed QUIC tunnel of pallet-vpn, a native SmartVPN executable, authenticates
the hub by the public key its credential names rather than by a certificate authority,
so no root store applies to it.
pnpm exec tstest test/test.tlstrust.node.ts --verbose --logfile --timeout 120
The test pins the default and the refusal, runs the CLI with an invalid store in three
modes, and, under the pinned Deno, dials a real HTTPS and WebSocket server whose
private CA is present only in the OS store (SSL_CERT_FILE): it is refused as
UnknownIssuer without the selection or with an explicit mozilla, and trusted with it.
Regenerating third-party notices
binary/control-third-party-notices.txt names the frozen npm graph, the Deno GNU
Linux runtime and the NoSQLDB engine. When the Deno pin moves, move every pin
first — denoVersion in binary/control-build.json, the sealed release's
denoVersion in .smartconfig.json and both Deno lines of
.gitea/workflows/release.yml — and add the new version's release commit and
the SHA-256 of its deno_src.tar.gz release asset (GitHub states it as the
asset digest) to binary/control-deno-sources.json. Then, with that Deno and
cargo-about 0.9.1 on PATH:
pnpm run notices:control # rewrite the notices and their pins
pnpm run notices:control --check # regenerate, compare, write nothing
scripts/notices-control.mjs refuses unless every Deno pin, the running Deno
and the pinned source agree. It downloads the pinned source of the Deno the
notices name and of the pinned Deno, checks both against their digests, and
takes the runtime's crate inventory with cargo-about over cli/rt for the
control targets (--locked, so Cargo fetches the locked crates). It needs
tar and network access to GitHub and crates.io, and works in a private
temporary directory it removes afterwards.
The regeneration carries the Deno runtime part and leaves everything else
byte-identical. A registry crate the notices already name keeps its entry;
Deno's workspace crates and every new crate are written from the crate's own
legal files, the workspace crates with Deno's LICENSE.md. A new crate without
a legal file is refused, because it needs a reviewed entry. The native V8,
Chromium Rust and Rust standard library notices are carried only while the
pinned Deno runs the V8 and the Rust toolchain they name; otherwise the tool
refuses until they are refreshed. The v8 crate cites rusty_v8's license at
the crate's own commit. The embedded JavaScript keeps its file list: each
listed file's legal comments are read again at the new release, each excerpt
is found again at its new lines, and a changed or new source file under ext/,
runtime/js/ or libs/core/ whose legal comments are not listed is refused
for review. The asset hash and noticeBuildIdentity.denoVersion in
binary/control-build.json follow the written file. --input <file> carries
another notices file instead of the committed one: from the Pallet 34.0.0
notices (Deno 2.9.4) the tool reproduces the committed Deno 2.9.7 notices byte
for byte. test/test.controldenonotices.node.ts covers the pin refusals and a
dry run over a small fixture without cargo-about.
Control bundle packaging
@git.zone/tspack owns archive assembly, file hashes, executable modes, sealed
manifests and complete archive verification. Pallet's adapter selects explicit
control builds and checks their source, version, compiler, runtime lock, fixed
paths and both native owners' provenance before passing inputs to TsPack. The selected
license notices are pinned in binary/control-build.json to the runtime lock,
Deno version and NoSQLDB owner build. The pinned native notice inventory binds
the Cargo/compiler inputs and every reviewed notice; the archive preserves the
native notice index, complete inventory and CRI provenance.
After sealing, the adapter extracts every bundle again from the sealed set and runs
the verifier of @serve.zone/pallet-bundle over it with the build's own source
identity, so a bundle that disagrees with what that package verifies for its
consumers is never produced. Its paths and control record keys come from the same
package (ts_bundle/inventory.ts). pack:control first compiles that package
(tsbuild custom ts_bundle to dist_ts_bundle/, output on stderr) from the
checked-out source, so it runs on a clean checkout and never loads a stale verifier;
release:control loads the adapter only after build:control has compiled it.
# Use the exact project-relative directories printed by build:control.
pnpm run pack:control dist_control/linux-amd64-DIGESTPREFIX dist_control/linux-arm64-DIGESTPREFIX
# A release requires both architectures built from the clean vVERSION tag.
pnpm run pack:control --release dist_control/linux-amd64-DIGESTPREFIX dist_control/linux-arm64-DIGESTPREFIX
Ordinary packaging produces an explicitly unpublishable qualification set. A
release produces the sealed set under dist_control_release/ and an
inputs-VERSION-COMMITPREFIX.json description that records its exact packaging
configuration and manifest digest. The output JSON identifies both paths.
A publishing workflow must retain the entire dist_control_release/ directory,
including that input description and sealed set, before publishing any asset.
Restore those original files on retry, then run:
pnpm run pack:control --release --reuse
Reuse verifies the retained source, configuration, manifest and every archive. It works without the original compiler outputs and rejects missing, corrupt or changed retained inputs. It never recompiles a control bundle: the only thing it compiles is the TypeScript of the bundle verifier above. Fresh Deno compilation may produce different bytes, so rebuilding cannot substitute for retaining a published set. TsPack and this adapter do not provide remote CI artifact storage or publication.
GitZone 6.6.2 or later manages .gitea/workflows/release.yml and the
scripts/gitzone-*.mjs release scripts through the committed tspackRelease
asset configuration. The tag-triggered workflow invokes the configured
scripts/prepare-control-release.sh command in its disposable job container.
It requires a root Linux Gitea Actions environment, installs Pallet's clang and
musl-tools compiler prerequisites through apt, then executes the existing
scripts/release-control.mjs exact-tag adapter to build and package its two
explicit output directories. Compiler installation writes only to stderr so the
adapter retains its single JSON result on stdout. The runner must be Gitea Runner 3.3.2 or
later with its cache/results service reachable from job containers and
runner.patch_actions enabled. The pinned stock artifact actions use the v4
protocol required by Gitea's REST retention inventory.
It retains the complete output root in Gitea Actions for 90 days, downloads that
retained copy, and verifies component reuse and every archive before creating a
release draft. The publisher checks existing attachment bytes, adds only missing
files, and publishes after complete readback. It never replaces release assets.
After a CI interruption, rerun the original Gitea run so it restores the
original bytes. Missing or expired retention and conflicting remote attachments
stop publication. The generic release receipt gitzone-release.json and Pallet's
input description must remain beside the sealed set in the retained artifact.
Use gitzone format --only assets --write --yes to update these managed scripts
from a published GitZone version, then commit the result before releasing.
Build and test from source
This development package is marked private and is never published. The npm
release target publishes one module of this repository, ts_bundle/, as
@serve.zone/pallet-bundle with Pallet's release version (see its own
ts_bundle/readme.md). Its tspublish.json takes the @git.zone/tspack and
@push.rocks/smartdaemon ranges from the root devDependencies
(devDependencyVersions), so neither enters the root dependencies that
deno.lock locks for the control build, and
@serve.zone/interfaces from the root dependencies. From the repository:
pnpm install
pnpm test
pnpm build
Tests build a native debug binary and exercise it against a disposable Unix-socket
HTTP/2 fake CRI server. They do not connect to Docker or a production daemon.
The native build uses Rust 1.95.0, locked Cargo dependencies and pinned vendored protoc to
produce static musl Linux binaries through @git.zone/tsrust, named
pallet_linux_amd64_musl and pallet_linux_arm64_musl. pnpm build builds only the
host architecture's executor; pnpm run build:control builds both with
tsrust --configured-targets, so a control build and the release always carry the
executors of both targets built from the same source. ARM64 cross-compilation
uses aarch64-linux-gnu-gcc as the linker driver with Rust's self-contained musl
target libraries. The ring TLS provider also requires musl-gcc for amd64 C/assembly
and Clang for its supported freestanding arm64-musl C build. These compiler choices
are declared in rust/.cargo/config.toml; provision them on developer and release
builders before running the build. Other packaged platforms are rejected explicitly.
After pnpm build, run PALLET_TEST_PACKAGED=1 pnpm exec tstest test/ --verbose --logfile --timeout 60 to exercise the host architecture's packaged executable and its
default lookup instead of the native debug executable. Cross-compiling an ARM64
artifact does not establish ARM64 runtime qualification.
For explicit emulator qualification, the same suite accepts an absolute
PALLET_TEST_BINARY_PATH pointing to a test-only emulator launcher. Do not combine
it with PALLET_TEST_PACKAGED; an emulated run does not qualify real node hardware.
The third-party notice index points to the native
notices tsrust notices generates in native-notices/ for the locked Cargo graph,
the Rust toolchain and the static runtime, with full license texts and the reviewed
once_cell licenses and source attributions of ring under
native-notices/crates/ring-0.17.14/extra/, and lists the material kept beside them. npm publication and
distribution of the standalone Rust runtime-probe binaries remain disabled.
Gitea releases distribute the control bundles described above, whose
runtime-serve activates the node network as described in "Foreground node
process".
Read-only probe
After a build, from the repository:
import { Pallet } from './dist_ts/index.js';
const pallet = new Pallet({
socketPath: '/run/pallet/containerd/containerd.sock',
});
const evidence = await pallet.probeRuntime({ timeoutMs: 5000 });
console.log(evidence);
// {
// runtimeName: 'containerd',
// runtimeVersion: '<actual daemon version>',
// runtimeApiVersion: 'v1',
// runtimeReady: true | false,
// networkReady: true | false,
// }
The socket must already exist and belong to a separately configured, authorized
containerd instance. Pallet never creates one, chooses a default socket, searches
PATH for its binary, or falls back to Docker. An explicit absolute binaryPath
can select a verified installed executable or the development debug binary. A
missing or nonexecutable selected binary fails with SmartRust 2's
RustBinaryLocatorError (ERR_RUST_BINARY_EXPLICIT_PATH_INVALID). Selection never
changes executable permissions or falls back to another packaged binary or a
stale GNU/native build. After explicitly provisioning the selected executable,
the same Pallet instance can retry.
Each probe owns a separate native child. Only one probe per Pallet instance is
admitted at a time. The method confirms that child's exit before returning,
including on failure. If cleanup cannot be confirmed, the method rejects and
retains ownership of the child. Further probes are blocked until await pallet.close()
successfully retries cleanup. An optional AbortSignal cancels the local request and
terminates the owned child; it never stops containerd itself.
The native deadline covers connection and both RPCs (50–30000 ms, default 5000).
The bridge readiness handshake, after executable discovery, is bounded to
3 seconds, with 1 second for graceful termination before forced child shutdown.
Executable filesystem discovery in the shared bridge is not currently timed or
abortable; this is not an end-to-end startup deadline. IPC and decoded gRPC responses are
limited to 16 KiB. The probe accepts only containerd with CRI API v1.
RuntimeReady and NetworkReady must each occur exactly once; a missing or
duplicate condition is an error, while explicit false remains false. Runtime
diagnostic messages and verbose configuration are never returned. Runtime
reported readiness is not workload health, version support qualification,
network reachability, storage fencing, or permission to perform a migration.
Native failures carry a RustBridgeRequestError.responseErrorCode of
INVALID_INPUT, SOCKET_UNAVAILABLE, RUNTIME_UNAVAILABLE,
DEADLINE_EXCEEDED, UNSUPPORTED_RUNTIME, or INVALID_EVIDENCE.
Bridge transport, cancellation, and startup failures remain distinct errors.
Isolation and qualification
Known Docker-private path components and aliases into them are rejected as accident prevention. This is not a security boundary against a hostile local user who can replace socket paths. Run only with the local permissions needed to inspect the intended daemon; containerd's socket is root-equivalent.
Fake-server tests establish protocol and lifecycle behavior only. Before runtime adoption, qualify a dedicated containerd 2.3 instance with explicitly owned root, state, socket and configuration paths, then verify actual container, network, storage and recovery behavior. No production cutover is implied.
Every QEMU qualification guest under test/native/ runs one bundled ES module.
scripts/guest-bundle.mjs builds it from the guest's own TypeScript source with
esbuild — the whole graph inlined, the NodeNext .js specifiers resolved to the
.ts sources in this checkout, bare Node builtins rewritten to their node:
form (which is what the deno compile --no-config --node-modules-dir=none half
of each guest needs) and a createRequire banner for the CommonJS dependencies:
node scripts/guest-bundle.mjs --entry test/native/namespace.ts --out .nogit/debug/guest/namespace.mjs
It prints the bundle's path, size and SHA-256, which is what a qualification run
records beside its other inputs. The runners take that file as --bundle (and
qualify-active-registry.py as --wrapper-bundle / --tunnel-bundle).
qualify-cni-up.py --scenario single-host runs test/native/singlehost.ts, a node
bound to a local Onebox controller, in the cni-up guest with no substituted seam.
network-acquire-local serves the contract's socket (a root-owned 0600 socket in
a root-only 0700 directory; a peer running as another user gets EACCES); the
first bind persists the binding, an exact replay answers the same credential, a
different binding refuses, a replayed bind's session supersedes the earlier
connection's, which can then no longer apply anything, and the signed onebox
projection is acquired over the current session. The node's runtime then raises one
generation without a managed VPN: the real barrier reaches forwarding on the
single-host fact with no tunnel command, execution admission is accepted on that
evidence, two workloads attached through the real CNI path are raised under the
generation, TCP and UDP flow between them with the client's own address preserved,
and each resolves the other's name to its lease address through the node-local
resolver. A third workload, on a network only the first shares, answers the first
and is dark from the second. The router namespace reads IPv4 forwarding off in
all and default before any generation exists, although the guest host forwards. From the host namespace it probes the granted TCP and UDP tuples, another
port, another protocol, another lease, another source address and a granted lease
that is not attached. The Cloudly counterpart — a generation without a tunnel
receipt settles at up and its barrier does not hold — stays proven by the default
cloudly scenario. With the generation's pool route in place the granted TCP and
UDP tuples reach the workload from the transit host address, and another port,
another protocol, another lease and another source address stay dark; a granted
lease that is not attached stays dark as well, stopped by the host-transit table
before the router receives a packet (both hops carry attached leases only). That
last check captures every frame on both ends of the transit link while it probes
(the guest stages tcpdump, without promiscuous mode, and parses its pcap) and
counts only IPv4 packets to the unattached lease: the link also carries the ARP
exchange both ends run on their own schedule — the router re-probes its entry for
the host about 5 s after the granted flows — which a frame counter cannot tell
from a leak. The same captures must see a granted flow on both ends, drop nothing
and account for every frame the link counters counted; with the host-transit
policy deliberately given the unattached lease's grant, the probe's SYN reaches
the router and the check fails. A
successor projection then withdraws every host selection under the held generation:
the pass withdraws that generation before it completes the changed dormant set,
raises the next generation on a re-keyed packet pair (forwarding, no host route
planned), and the formerly granted tuple goes dark. Probers dial every 100 ms
through the change: the allowed pair between the first two workloads, and six
tuples the policies deny — the pair that shares no network over TCP and UDP, and
workload to the host's transit and uplink addresses over TCP and UDP, each with a
live listener behind it. East-west traffic stops for the withdrawal window, because
the withdrawn generation restores a router that forwards nothing, and resumes when
the successor is UP (about 27 s of a 45 s change); every denied tuple is
dialled inside that window and never answers before, during or after it, and no
listener logs its source. Without the forwarding pin the same guest measures the
unfiltered router: the pair that shares no network answered 61 of 201 dials during
the change. All of it passes on 6.18.35-0-virt. The packet hub guest (packet.ts) plans the same pool
route and proves it present only while its generation is UP, gone after a
withdrawal, and gone after a fresh owner's recovery of a lost native.
qualify-cni.py --scenarios containerd-serve --control-directory dist_control/linux-amd64-…
runs Pallet's own containerd from a build:control directory the way its service
unit does, with a foreign /etc/containerd/conf.d drop-in planted in the guest. It
proves the generated configuration governs (the drop-in's stream port stays closed,
127.0.0.1:10010 serves), the bundled sandbox image is imported and unpacked on
overlayfs, a workload raised by the production execution owner through the real
CNI path runs under /pallet with runc state under /run/pallet/runc, SIGTERM stops
containerd and leaves the workload running, the next containerd-serve reattaches
the exact sandbox and container, and killing containerd fails its supervisor
(owner_failed) while the workload keeps running. The image collector keeps the
running workload's image and the bundled sandbox image, and the proven removal
frees the workload's image while the sandbox image stays.
qualify-cni.py --scenarios owner-storage --control-directory dist_control/linux-amd64-…
runs one local storage claim under the same bundled containerd and runc, on node
and Deno. The claim materialises its volume root-private with the claim's
ownership; containerd echoes exactly one private read-write bind of it and the
bundled runc's workload carries it at the claimed path; the workload writes as
its own user; a second run of the service and a purge are refused while the
first container lives (storage.claim.held, storage.purge.held); a restarted
owner rejoins the live container against the grant's mount list; the second run
finds the first one's bytes after the first is removed; and the purge deletes the
volume once both are gone.
qualify-active-registry.py --paths-node-binary … --paths-bundle … (with the
packet and SmartVPN binaries) runs test/native/singlehostpaths.ts: the three
single-host paths of a Cloudly node whose platform services run on its own host.
The guest adds the hub's and the relay's addresses to the uplink as permanent
/32s before the uplink is observed, attaches and raises three workloads that
share no private network, and composes one projection with the production
composer twice: without hostPlatformEndpointIds and workloadIngress, and with
them. Under the first pair a workload's dial of the relay it selects, the router
namespace's dial of the hub and the ingress workload's TCP and UDP dials of its
target are all dark; the composed pair replaces it one revision up under the same
UP generation and all four are delivered — the relay and the hub see the transit
source address with a port of the leased range, the target sees the ingress
workload's own address — and the real SmartVPN hub on the host authenticates the
router namespace's real managed QUIC client. Another workload to the relay, the
Corestore port the workload did not select, an undeclared hub port, the host
opening toward the router, the target opening back toward the ingress workload,
another workload to the target and another target port stay dark under both
pairs, each against a live listener that logs no peer.
The pinned upstream CRI protocol and the containerd Transfer and Streaming protocols,
with their Apache-2.0 attributions, are under rust/proto/.
License and Legal Information
This repository contains open-source code licensed under the MIT License. A copy of the license can be found in the repository license file. The vendored Kubernetes protocol is separately licensed under Apache-2.0.
Please note: The MIT License does not grant permission to use the trade names, trademarks, service marks, or product names of the project, except as required for reasonable and customary use in describing the origin of the work and reproducing the content of the NOTICE file.
Trademarks
This project is owned and maintained by Task Venture Capital GmbH. The names and logos associated with Task Venture Capital GmbH and any related products or services are trademarks of Task Venture Capital GmbH or third parties, and are not included within the scope of the MIT license granted herein.
Use of these trademarks must comply with Task Venture Capital GmbH's Trademark Guidelines or the guidelines of the respective third-party owners, and any usage must be approved in writing. Third-party trademarks used herein are the property of their respective owners and used only in a descriptive manner, e.g. for an implementation of an API or similar.
Company Information
Task Venture Capital GmbH
Registered at District Court Bremen HRB 35230 HB, Germany
For any legal inquiries or further information, please contact us via email at hello@task.vc.
By using this repository, you acknowledge that you have read this section, agree to comply with its terms, and understand that the licensing of the code does not imply endorsement by Task Venture Capital GmbH of any derivative works.