2026-10-02 22:57:09 +00:00
…
…
2026-10-02 22:57:09 +00:00
2026-10-02 22:57:09 +00:00
…
2026-10-02 22:57:09 +00:00
…

@push.rocks/smartnftables

A TypeScript module for managing Linux nftables rules with a high-level, type-safe API. Handles NAT (DNAT/SNAT/masquerade), firewall rules, IP sets, and rate limiting — all from clean, declarative TypeScript.

Issue Reporting and Security

For reporting bugs, issues, or security vulnerabilities, please visit community.foss.global/. This is the central community hub for all issue reporting. Developers who sign and comply with our contribution agreement and go through identification can also get a code.foss.global/ account to submit Pull Requests directly.

Install

pnpm install @push.rocks/smartnftables
# or
npm install @push.rocks/smartnftables

The legacy SmartNftables helper tracks rules without applying them when it lacks root privileges. ManagedNftables always rejects unsupported enforcement; it has no dry-run or memory-only enforcement mode.

Managed workload policy

ManagedNftables provides a separate native owner for complete, interface-bound IPv4 policy. It requires Linux 6.9 or newer, CAP_NET_ADMIN in its network namespace, and working nftables OWNER/PERSIST support. The published native matrix is Linux x86_64 and arm64, statically linked with musl. Actual kernel support is checked when applying policy; starting the compiler does not establish enforcement.

An optional caller-owned networkNamespaceFd selects an already-created Linux network namespace. Keep that descriptor valid through start(). The native process validates and enters its inherited copy before readiness, owner identity capture or netlink activity, then closes the copy. Invalid descriptors and denied entry fail before readiness. The option is process configuration and never enters openOwner, policy commands or durable receipts. Native callers use --management --network-namespace-fd 3 with an inherited descriptor at FD 3.

The caller owns the namespace lifetime, links and routes. Retain a namespace owner across process loss and recovery: PERSIST retains policy only while its namespace exists. An entered process keeps its namespace membership if the source descriptor is closed, but a replacement process requires a valid descriptor. Receipts bind the actual target namespace device and inode; recovery in another namespace rejects. Await complete policy and process cleanup before releasing the caller's final namespace owner. Namespace entry does not create uplink authority, drain conntrack flows or authorize address and port reuse.

The Linux x86_64 musl candidate passes isolated Linux 6.18.35 kernel tests, including inherited namespace entry, retained policy after process loss, recovery in the original namespace, rejection in another namespace, and denied entry before readiness. ARM binaries are built; privileged ARM and complete Pallet workload integration remain separate qualification requirements.

The caller owns workload links and durable state. Keep links disabled until their policy is confirmed, serialize link changes with policy operations, and disable them before cleanup. Persist the full prepared transition before calling reconcile(), then retain the full returned receipt in the caller's SmartData store. This library does not persist runtime state in files.

import { ManagedNftables } from '@push.rocks/smartnftables';

const owner = new ManagedNftables({
  ownerId: 'node-policy',
  instanceId: 'caller_retained_unique_instance',
  tableName: 'snft_workloads',
});
await owner.start(); // Compiler/IPC readiness only; no kernel policy yet.

const target = await owner.prepare({
  schemaVersion: 1,
  revision: 1,
  endpoints: [
    { id: 'workload-a', interfaceIndex: 12, interfaceName: 'worka', interfaceKind: 'veth', sourcePrefixes: ['10.81.1.2/32'] },
    { id: 'workload-b', interfaceIndex: 13, interfaceName: 'workb', interfaceKind: 'veth', sourcePrefixes: ['10.81.2.2/32'] },
  ],
  rules: [
    {
      sourceEndpoint: 'workload-a', destinationEndpoint: 'workload-b',
      sourcePrefix: '10.81.1.2/32', destinationPrefix: '10.81.2.2/32',
      protocol: 'tcp', sourcePort: null, destinationPort: 443,
    },
    {
      sourceEndpoint: 'workload-b', destinationEndpoint: 'workload-a',
      sourcePrefix: '10.81.2.2/32', destinationPrefix: '10.81.1.2/32',
      protocol: 'tcp', sourcePort: 443, destinationPort: null,
    },
  ],
});
const transition = { previous: null, target };
// Persist this complete transition before the next operation.
const applied = await owner.reconcile(transition);
// Persist the complete applied result before enabling the caller-owned links.
const status = await owner.inspect();
if (!status.enforced) throw new Error('Workload policy is not confirmed.');

// After disabling/quiescing these workload links:
await owner.release(applied);
await owner.close();

Interface numbers and names in this example must be replaced with the caller's actual, already-created links. Native RTM_GETLINK checks the name/index pair, the declared veth or L3 tun kind and absence of a bridge/master attachment. The caller must prove and own each veth peer and workload network namespace. Both selectors must match for a grant. A stale name/index pair rejects application; renaming either selector retains denial. Do not reuse interface identities or workload addresses while their previous grants remain outstanding.

Method Behavior
start() Starts one private native process and returns idle status.
failureSignal Stable per-instance AbortSignal for lost owner confidence, including an unexpected idle process exit. It never resets or confirms cleanup.
prepare(policy) Validates, canonicalizes and hashes an inert complete policy without kernel writes.
reconcile({ previous, target }) Creates or replaces the entire policy in one kernel transaction. previous is null initially, otherwise the full prior applied result.
inspect() Reads the current owned graph and reports enforced, pending intent and a bounded error code.
release(applied) Verifies the exact retained graph and confirms deletion; repeats are idempotent.
releaseTransition({ identity, transition }) Selects terminal cleanup using the original full identity and transition when the applied result is unavailable. Accepts absence or deletes the exact previous/target graph without requiring surviving interfaces.
detach(applied) Selects terminal retention, verifies the same graph with PERSIST-only flags after dropping socket ownership, then joins the native process. Returns { detached: true }; retain the full applied result.
closeRetaining() Stops command admission, joins admitted operations and terminates the native process without requesting deletion or confirming retention. Use when the exact applied result is unavailable.
close() Stops command admission, joins admitted policy operations, confirms owned-table deletion unless retention was selected, then joins the native process. Failure retains the owner for explicit recovery.

Schema-v1 policies are directed. Return traffic requires its own grant; there is no broad connection-tracking bypass. A null endpoint denotes the actual local host, with an explicit prefix that cannot overlap any endpoint. Such a grant applies only to input/output. Forwarding requires two explicit endpoint references; a relay TUN is an endpoint with interfaceKind: 'tun' and its authenticated remote source prefixes. The relay owner must authenticate that remote source authority. Managed interfaces reject other IPv4 traffic and IPv6 traffic. Ethernet/ARP is outside these inet hooks; workload links must use the qualified routed layout. Unmatched interfaces retain their existing behavior. The compiler creates only inet input, forward and output filter chains; it does not configure links, routes, NAT, DNS, host defaults, forwarding sysctls, or relay sessions. An ungranted destination is denied independently of DNS answers.

The schema permits at most 32 interfaces, 16 source prefixes per interface and 128 directed rules, additionally bounded to 768 compiled operations and 100,000 encoded bytes. IPv4 prefixes must be canonical, non-overlapping local allocations; loopback, unspecified and multicast allocations reject. Ports require explicit TCP/UDP and all nullable fields must be present. The facade rejects getters, proxies, live objects and concurrent commands rather than queuing unbounded work.

One NETLINK_NETFILTER socket exclusively owns the table with OWNER/PERSIST flags. An acknowledged batch commits the complete graph atomically. Lost acknowledgements retain the exact pending transition; retry only that same body. Reconnect/restart uses the same caller-retained owner, instance, table and boot/namespace receipt. Recovery checks the complete graph before and after orphan adoption. The kernel's wrapping generation counter only fences concurrent transactions; it is never a durable revision or sufficient ownership proof. Foreign owners, unexpected tables, chains, rules, sets or objects inside the owned table reject without deletion. Other tables coexist independently. Linux's family-wide chain dump is explicitly filtered by table identity before graph comparison; foreign chains are never included in replacement or cleanup.

PERSIST deliberately retains accepted static grants if the native process crashes. It supplies no wall-clock lease or timed revocation. The caller must keep revocation pending until policy replacement or workload fencing is confirmed. Host reboot invalidates the old boot receipt; the caller must fence its old link/workload lifecycle before starting a new owner instance. Cleanup failures and uncertain operations must not be reported as successful release or used to transfer authority.

Subscribe to failureSignal before starting or admitting policy work, and check failureSignal.aborted when attaching to an existing instance. Unexpected child exit, ambiguous operation failures, native inspection reporting failed ownership, and unconfirmed cleanup abort the signal once with ManagedNftablesError code OWNER_LOST. Local input rejection, effect-free native compiler validation and native INVALID and EXHAUSTED responses do not abort it. Other native errors can occur after kernel effects or ownership loss and conservatively abort it. Successful orderly close, including a cancelled startup, does not signal an unexpected failure; ambiguous mutations still signal loss while close is draining them.

The caller must fence packet admission on loss. Recovery operations and exact retries remain available as documented, but successful recovery never resets the signal. It does not monitor foreign kernel changes while idle or prove policy withdrawal: PERSIST policy can remain after process death. Still await the chosen joined cleanup method and retain unresolved policy/allocation authority.

For an orderly restart that must preserve policy, persist the complete applied result and call await owner.detach(applied) before calling close(). The native operation validates the exact receipt, graph, interfaces, boot and namespace. It accepts this socket's OWNER/PERSIST table or an exact orphan; it never adopts a foreign owner. After closing the owner socket it reads the graph again and requires the same handle and compiled body with PERSIST-only flags. Native inspection then reports state: 'retained'. The facade returns success only after child termination.

Calling detach() irreversibly selects retention cleanup, including invalid input, transport failure, timeout, native rejection or malformed acknowledgement. Subsequent close() only joins and terminates; it never sends a table-deletion command. It cannot undo a close() already admitted earlier. Other policy operations reject. Resolve any pending policy transition by exact replay before requesting detachment. If a detach ACK is lost while the process remains available, only the same captured body may be retried. After process loss, start a fresh owner with identical options and retry detach(applied) against the exact orphan. Absent graphs, foreign active owners, changed interfaces, handles, bodies or boot/namespace identities reject. Native release/reconcile/close commands reject after a valid detach intent, even when post-drop verification fails; explicit recovery requires the retained body.

If reconciliation remains ambiguous and no complete applied result is available, call await owner.closeRetaining(). This permanently stops command admission, joins admitted work, and terminates and joins the native child without sending detachPolicy or closeOwner. It does not inspect the graph or acknowledge retention, enforcement, release, or namespace survival. An undelivered reconcile may have left no graph; a lost reply may have left the original or target graph. Keep the complete original transition durably and recover it through a fresh owner with identical options. Never fabricate an applied receipt to request detachment.

When cleanup is required instead of recovering enforcement, persist a IManagedNftTransitionRelease containing the identity returned by the original start() and the full original transition. After fencing workload packet admission and joining the old owner, start a fresh owner with the same options:

// Both values come from the caller's durable journal, before the lost apply ACK.
const cleanup = { identity: originalIdentity, transition: originalTransition };
await recoveryOwner.releaseTransition(cleanup);
await recoveryOwner.close();

releaseTransition() validates the original boot and namespace, canonical target, complete previous result when present, table identity and entire rule graph. It accepts an absent table, the previous graph, or the target graph; it never creates or replaces policy. Deleted or changed interfaces do not prevent this cleanup. Orphan adoption checks the graph before and after claiming OWNER, binds the prior or already-observed handle, and rejects another live owner. Deletion uses the exact handle and a generation-fenced batch, followed by an absence check.

Once admitted, only that captured cleanup body may be retried. Other policy writes, preparation and detachment reject; inspection never claims enforcement. An ordinary close() retries the selected cleanup before joining the child. closeRetaining() only joins, allowing a fresh process to retry the durably retained cleanup after an ambiguous acknowledgement. Local inert-input or busy rejection does not select cleanup. A successful release acknowledges that table operation at observation time, including an exact retry after a lost delete ACK. It does not prove packet drain, clear conntrack, release allocations, or cross a changed boot/namespace.

After closeRetaining(), all policy commands, including detach(), reject. Concurrent cleanup calls share the same join. A failed join keeps the native owner and retention choice for retry through closeRetaining() or close(). An ordinary close() admitted first keeps its deletion choice across failures; a later closeRetaining() rejects instead of reversing that choice.

Detachment confirms retention at observation time. A completed facade retry returns its existing acknowledgement; it is not a fresh kernel inspection. Retention does not guarantee future enforcement, namespace survival or reboot ordering, and does not withdraw external authority, drain packets, clear conntrack or release leases. A separately created owner can later recover the full policy and perform an explicit release after the caller has fenced that authority and packet lifecycle.

Apply and release receipts describe this owner's table and policy operations. They do not certify that packets admitted under an earlier graph have drained: foreign NFQUEUE or other deferred packet owners may still hold such packets. The caller must qualify the complete packet path or independently fence the workload before treating revocation as complete or reusing its addresses and interfaces. Table deletion does not release IP allocation authority or clean up conntrack/NAT state. This restriction also applies to private veth/TUN forwarding.

Refusals: INVALID and EXHAUSTED

Every failure is a ManagedNftablesError with a stable code. Two codes are refusals of the input itself, raised by the facade's capture or by the native owner before any kernel work:

  • INVALID: the input is outside the contract (schema, identifiers, addresses, overlaps, authority, or a bound of the topology's shape). A smaller policy does not help; the input must change.
  • EXHAUSTED: the policy is too big for one atomic replacement. It exceeds the compiled budget described under the router below, or an input count that exists to hold that budget. A smaller policy may fit.

Both carry the native refusal text in reason (bounded, printable ASCII, never policy content) and in the message, Managed nftables EXHAUSTED: <reason>. EXHAUSTED also names its bound in details, { bound, limit, actual }:

bound limit Counts
ruleBytes 100,000 encoded rules, chains and sets of the target
targetBytes 212,000 the complete target with its set elements
operations 768 messages of the target
endpoints 32 (v1), 128 (router) private endpoints
links 128 router links
rules 128 (v1), 1024 (router) private directed rules
grants 1024 router egress grants across every generation
hostGrants 1024 host grants of one scope
workloadGrants 1024 router workload grants
publishedPorts 1024 publications of one scope
localTcpPortOwners 8 loopback TCP port owners of the pool guard
restoreRules 192 Docker forwarding contribution rules
restoreBytes 100,000 Docker forwarding contribution command bytes

actual is always above limit. It is exact for input counts and for the schema-v1 budget, which measures its complete program. For the other compiled budgets (ruleBytes, targetBytes, operations, restoreRules, restoreBytes) compilation stops at the first message or rule past the limit, so the complete policy needs at least actual. Input counts are checked before their entries are validated, so a smaller policy can still be INVALID. Every other bound (identifier and interface-name lengths, source prefixes per endpoint, addresses per link, protected prefixes, platform endpoints, active generations, handoffs and allocations, ranges per allocation, guarded pools, the IPC and capture limits) describes the shape of the topology and stays INVALID. Other codes, and the facade's own local rejections, carry neither reason nor details. A native refusal outside this shape is PROTOCOL. A failed close() rejects CLEANUP_UNCONFIRMED, whose cause is the ManagedNftablesError that left cleanup unconfirmed, with the native code when the native owner refused.

try {
  prepared = await owner.prepare(policy);
} catch (error) {
  if (error instanceof ManagedNftablesError && error.code === 'EXHAUSTED' && error.details?.bound === 'publishedPorts') {
    // Publish fewer ports and prepare again.
  } else throw error;
}

Combined router egress and host transit

ManagedNftables<IManagedNftPolicyV2> accepts the exported schema-v2 policy. The prepared/applied/transition/status interfaces accept the same policy type parameter; existing callers default to the unchanged schema-v1 contract. V1 canonical digests and compiled bytes are preserved. V2 uses its own hash domain. An owner cannot transition between v1 private, v2 router, v2 host transit, and v2 allocation-pool guard policy kinds.

V2 scope Required authority and behavior
routerEgress Private endpoints and rules, one exact links binding per endpoint, a separate veth handoff, protection, and active generations. Private veth/TUN/local DNS and egress share one table so terminal private denial cannot override a separate egress table. Optional publishedPorts add the inbound second hop from the handoff to a workload endpoint, one port or a port range; both directions are classified into the default conntrack zone ahead of every leased classifier, so a published endpoint port is dedicated to its publication and never becomes leased egress. A symmetric publication also lets the workload open flows from its published ports. Optional hostGrants forward exact host-origin flows from the handoff to a workload endpoint. Optional workloadGrants let one workload open one exact port of another, one way.
hostTransit Exact handoff link/allocations pairs, complete protection, an explicit veth or Ethernet uplink, and its current snatAddress. It checks each handoff's leased source address and protocol/port range, default conntrack zone, direction, uplink, and protected destinations before outer SNAT. Optional publishedPorts add inbound uplink destination NAT inside the same generation, one port or a port range, and a symmetric publication also carries the workload's own flows from its published ports out through the uplink. Optional hostGrants let the host's own address on a handoff dial exact workload ports. Optional localPlatformEndpoints serve platform endpoints on the host's own addresses to leased flows. Optional exclusiveForwarding makes the table the host's only forwarding owner: every other forwarded packet drops.
allocationPoolGuard An authenticated authorityDigest and complete current allocation-pool prefixes. Installs host-wide IPv4 destination denial before any handoff exists, without link, uplink or SNAT dependencies. Optional hostGrants are its only exceptions. Optional localTcpPortOwners restrict loopback TCP ports to one local user each.

allocationPoolGuard accepts 1–64 canonical, disjoint RFC1918 prefixes. Supply the actual allocation pools, not the broader protected union containing management LANs, resolvers and platform endpoints. It applies to IPv4 INPUT, FORWARD and OUTPUT at filter priority 0, after normal destination NAT. It drops an original conntrack destination in a pool. Reply-direction packets additionally check their current source, then return; all remaining packets check their current destination. This denies DNAT into or away from a pool and untracked current-destination traffic, while permitting reverse-SNAT replies to authorized transit sources. It grants no egress permissions; exact hostTransit and router policies remain separate owners. Unrelated host traffic and IPv6 retain their existing behavior; the guard never denies forwarding between other links. On a host that forwards only for the handoffs, see Exclusive forwarding.

For example, a guard policy is { schemaVersion: 2, revision: 1, scope: { kind: 'allocationPoolGuard', authorityDigest, prefixes: ['10.240.0.0/16', '10.241.0.0/16'] } }. Hashing binds the supplied body, but does not authenticate the authority or prove its completeness. An empty hostTransit.handoffs array provides no host-wide denial and is not a substitute for this scope.

The caller must retain the authenticated authority, complete transition and applied receipt, verify inspect().enforced, and qualify the complete packet path. This guard alone proves neither early-boot/late-shutdown ordering nor protection after kernel reboot. It does not fence foreign packet queues, later packet rewrites, offloads or proxies, and supplies no quarantine release, conntrack drainage or allocation-reuse evidence. Historical leases remain the caller's durable ownership responsibility. PERSIST retains this table only while its network namespace exists; release() and ordinary close() delete it, so revoke any external protection claim before release, or use the verified retention operation described above.

Local bindings include name, index, kind, MAC (null for L3 TUN), interface-link index, and required IPv4 addresses. RTM_GETLINK/RTM_GETADDR verify those local facts at apply, recovery and inspection. Ethernet must be unbridged driver-backed Ethernet, including VirtIO; virtual VLAN/bond/bridge/dummy kinds are not inferred uplinks. The caller retains actual peer namespaces and link-generation ownership. These serialized facts are not native lifetime capabilities or a continuous link-change monitor. Address, DHCP and route changes require caller fencing.

protection carries an authority digest, non-overlapping protected IPv4 prefixes and exact platform endpoint IDs/address/protocol/port tuples. The caller must authenticate and supply complete authority. Hashing does not prove completeness. A public grant means its declared IPv4 prefix excluding the protected union; only an explicit platform endpoint grant admits a protected destination. Host checks cover both the current packet destination and original conntrack destination, so foreign DNAT cannot turn a protected destination into a public exception or redirect public traffic into protected space. INPUT diversion, local OUTPUT into handoffs, unmatched handoff traffic and IPv6 forwarding are denied.

Every handoff address and platform endpoint lies in the protected prefixes. The hostTransit uplink addresses need not: the authority is the platform's address space, and a node's uplink lease and snatAddress is usually a public address outside it. The host barrier protects the uplink addresses itself, adding each one no protected prefix covers as a /32 after the authority's prefixes, so a forwarded flow whose current or original destination is the host's uplink address, or one sourced from it toward a handoff, is denied like any protected destination. A scope whose uplink the prefixes already cover compiles exactly as before.

Published host ports

hostTransit.publishedPorts is optional. Absent and empty are the same canonical policy, so an unpublished scope keeps its exact previous digest and compiled bytes. Each entry is { protocol, hostPort, targetPort, targetAddress, hostIp }, with an optional hostPortEnd and an optional symmetric. hostIp must be an exact current uplink address, which rtnetlink verifies at apply, recovery and inspection; there is no wildcard, secondary-link or default-route inference. targetAddress must be a current leased transitSourceAddress of one allocation, and that binding is what selects the handoff link of the admission; an address that no active lease holds, a local host address or a bare workload prefix rejects. Ports are 1–65535, one protocol/hostPort is published exactly once across the scope including across uplink addresses, and a port published on the snatAddress may not fall inside any leased outbound source-port range of the same protocol, because outer SNAT translates to that same address.

hostPortEnd publishes the range hostPort..hostPortEnd. It must be greater than hostPort, and targetPort must equal hostPort: a range keeps every port, so port p of the uplink reaches port p of targetAddress. One port has exactly one canonical form, without hostPortEnd. Ranges follow the same rules as single ports: no two publications of one protocol overlap anywhere in the scope, and no published port on the snatAddress falls inside a leased range. A range translates the address only, which never changes a port.

Publications are set elements, not rules. The generation compiles CONSTANT named sets and maps of concatenated fields in which the port fields are inclusive intervals (the kernel's pipapo backend), so one element holds one port or a whole range: published_port maps a one-port publication's uplink address, protocols and port to its target address and port, published_address maps a range to its target address alone, and published_forward holds every publication's inbound tuple. A constant number of rules looks them up whatever the number of publications: one destination NAT rule per map in this table's own prerouting nat chain at priority -100, and four forward admissions ahead of the capture barrier — the ESTABLISHED original direction, a NEW opening per protocol with opening TCP restricted to SYN with FIN/RST/ACK clear, and the ESTABLISHED reply. Every rule checks the uplink, the default zone and the direction, and every key carries the exact handoff (index and name), the packet and conntrack protocols, the translated current address and port and the original conntrack destination, so foreign NAT cannot borrow the exception for another flow. Published traffic is not source-translated, so the workload observes the actual client address. The chain, its sets, its translations and its admissions are created, replaced and deleted with the generation: a target without the entry, a failed batch, or release removes them, and a caller never needs a separate teardown. Owner verification reads every set back with its description, every element with its interval bounds and map data, and every lookup with its registers, so apply, inspection, adoption and recovery reject a changed element exactly as they reject a changed rule.

Without symmetric, a publication carries inbound flows only: a flow the workload opens from its target address and port toward the uplink is denied. With symmetric: true, a flow from exactly targetAddress and the published target port(s) on the exact handoff may leave through the uplink, NEW (a TCP opening only with SYN) or ESTABLISHED, with its ESTABLISHED replies, and outer source NAT translates it to hostIp and the published uplink port. These flows are elements of one more map, published_symmetric, keyed on the handoff, protocols and the live and original source and holding the uplink address and port span; four admissions look it up as a set and one postrouting rule translates through it. The workload's outbound flows therefore leave from the same address and port its clients reach, as Docker host-mode NAT kept them; a range keeps each port unless the host itself already holds that exact tuple on hostIp. The admissions follow the capture barrier, so every protected current or original destination, the platform endpoints included, stays denied. A symmetric target port may not fall inside a leased range of targetAddress, and may not overlap the target ports of another publication of the same address and protocol. Absent and false are the same canonical policy, digest and compiled graph.

Input bounds are 1024 entries. The rule graph does not grow with them: with every kind of publication present the host compiles at most eleven rules and four sets, about 13,400 bytes of the 100,000-byte rule reserve, and each publication costs about 150 to 300 bytes of set elements, which share the target budget described under host grants. 64 symmetric ranges take about 19,600 element bytes. The caller still owns the route to targetAddress through that handoff, the workload listener, and every other packet owner on the host. As with leased egress, an ACCEPT here cannot override Docker's independent FORWARD DROP; the ManagedDockerForwarding contribution below admits every inbound and symmetric publication of the barrier through it, with the leased flows.

Each router generation binds an immutable lease reference, transit source address, leased TCP/UDP source-port ranges, a nonzero conntrack zone and a 16-byte label. Each directed grant selects one exact leased protocol range for SNAT. A grant's source is a workload veth or the actual router-local host with an exact assigned IPv4 source address; TUN endpoints retain private routing only. Router-local DNS and relay transports therefore need explicit local-origin grants. Raw PREROUTING and OUTPUT classify before conntrack; filter rules validate current and original tuples, links, direction, zone, state and generation label on every packet. Opening TCP requires SYN with FIN/RST/ACK clear and permits ECN negotiation. Replies must match the labelled original flow and current directed grant. There is no broad ESTABLISHED/RELATED bypass.

Published workload ports

routerEgress.publishedPorts is optional and is the second hop of a published host port. Absent and empty are the same canonical policy, so an unpublished scope keeps its exact previous digest and compiled bytes. Each entry is { protocol, transitPort, endpointPort, endpointAddress, transitSourceAddress }, with an optional transitPortEnd and an optional symmetric. transitSourceAddress must be a current leased generation address that is present on handoff, which rtnetlink verifies at apply, recovery and inspection; the handoff can carry several current leased addresses, so the arrival address is always explicit and never inferred. transitPort is the targetPort of the matching hostTransit publication. endpointAddress must be an address of one current veth workload endpoint of this scope: a router-local link or gateway address, the transit address, a TUN endpoint's private space, a platform endpoint, loopback and IPv6 all reject. Ports are 1–65535, one protocol/transitPort is published exactly once across the scope, and a published port may not fall inside any leased outbound source-port range of the same protocol on that transit address, because source NAT translates to that same address.

transitPortEnd publishes the range transitPort..transitPortEnd, the second hop of a host range. It must be greater than transitPort, and endpointPort must equal transitPort, so every port keeps its number on both hops; one port has exactly one canonical form, without transitPortEnd. No two publications of one protocol overlap anywhere in the scope, and no published port falls inside a leased range of its transit address. The destination NAT of a range translates the address only.

Publications are set elements here too: published_arrival (protocol, transit address and port), published_endpoint (workload link, protocol, address and port, the union of every publication's workload ports as disjoint canonical intervals), the published_port and published_address translation maps and published_forward, with the port fields as inclusive intervals. A constant number of rules looks them up whatever the number of publications: one destination NAT rule per map in this table's own prerouting nat chain at priority -100, two raw classifications, and four forward admissions placed ahead of the capture barrier — the ESTABLISHED original direction, a NEW opening per protocol with opening TCP restricted to SYN with FIN/RST/ACK clear, and the ESTABLISHED reply. The handoff ingress classification precedes the barrier's protected-source denial, because the uplink or management network that reaches a published port can itself be protected space (an uplink address outside the protected prefixes is still guarded as a /32). The endpoint ingress classification matches the exact workload link, address and published endpoint port, and precedes every leased classifier of every generation: raw runs before conntrack, so a client whose source port equals a leased grant's destination port would otherwise carry the workload's published packets into that generation, where the publication's flow does not exist. Both classifications match the arriving tuple only, set no generation zone and leave the flow in the default conntrack zone. The private anti-spoof guards still run first, and every admission checks the handoff, the default zone and the direction, and looks up the exact workload link (index and name), both protocols, the translated current address/port and the original conntrack destination.

A published endpoint port is therefore dedicated to its publication: a new outbound flow that the workload opens from that exact address and port on that exact link stays in the default zone too and is never translated by the leased outbound source NAT. Without symmetric such a flow is denied. Inbound published traffic is not source-translated on either hop, so the workload observes the actual client address. The chain, its sets, its translations, its classifications and its admissions are created, replaced and deleted with the generation: a target without the entry, a failed batch, or release removes them.

With symmetric: true the workload may open flows from exactly endpointAddress and the published endpoint port(s) on its exact link toward the handoff, NEW (a TCP opening only with SYN) or ESTABLISHED, with their ESTABLISHED replies, and source NAT translates them to transitSourceAddress and the published transit port. With the matching symmetric hostTransit publication the flow reaches the uplink from hostIp and the published uplink port, the same address and port its clients reach, as Docker host-mode NAT kept it; a range keeps each port unless that exact tuple is already taken on the translated address. These flows are elements of the published_symmetric map, keyed on the workload link, protocols and the live and original source and holding the transit address and port span, behind four admissions and one postrouting translation. Both directions are already in the default zone through the publication's endpoint and handoff classifications, so no leased classifier can capture them. The admissions follow the capture barrier, so every protected current or original destination, the platform endpoints included, stays denied; outbound flows anywhere else still need their leased grants. The endpoint ports of a symmetric publication may not overlap the endpoint ports of another publication of the same address and protocol, since each flow has exactly one published source. Absent and false are the same canonical policy, digest and compiled graph.

Input bounds are 1024 entries. The rule graph does not grow with them: with every kind of publication present the router compiles at most thirteen rules and six sets, about 15,100 bytes of the 100,000-byte rule reserve, and each publication costs about 250 to 450 bytes of set elements, which share the target budget with the router's endpoints, rules and grants. 64 symmetric ranges take about 28,000 element bytes. The caller still owns both hops: the exact matching hostTransit publication, the route to the workload through this router, the workload listener and every other packet owner.

Host grants

A host grant lets the node's own host network namespace dial one exact port of one local workload at the workload's address, from the host's own address on the current handoff (the transit host address). It is how a controller that runs on the node, such as a reverse proxy, reaches a workload without loopback publication, route_localnet, loopback DNAT or loosened martian filtering. Each entry is { protocol, sourceAddress, destinationAddress, destinationPort }: tcp or udp, two exact unicast IPv4 addresses and one port 1–65535. Prefixes, wildcards, ranges, lists, loopback, broadcast, IPv6 and duplicate tuples reject. There is no source-port selection and no translation.

The flow crosses three tables, so the caller compiles the same grants into each scope, and each scope binds them to the authority it proves:

  • allocationPoolGuard.hostGrants: destinationAddress lies in a guarded pool. The admissions sit in OUTPUT (original direction) and INPUT (reply) ahead of the guard. Everything else still meets the unconditional guard, including forwarded traffic and a translated flow that merely arrives at a granted tuple.
  • hostTransit.hostGrants: sourceAddress is an exact current address of exactly one handoff link and not an uplink address; destinationAddress is protected space that is neither host-local, a platform endpoint nor a leased transit source address. The admissions bind the exact handoff that holds the source address and precede its terminal OUTPUT and INPUT denials; forwarding and the uplink are unchanged.
  • routerEgress.hostGrants: destinationAddress is an exact current veth workload endpoint address of this scope, and sourceAddress is protected space that is neither router-local, a platform endpoint nor any workload's address. One raw handoff ingress classification precedes the barrier's protected-source denial, sets no generation zone and leaves the flow in the default conntrack zone; the forward admissions bind the handoff and the exact workload link, follow the private anti-spoof guards, and sit beside the publications ahead of the capture barrier. The workload's replies need no classifier: their destination is protected, so they return to the default zone ahead of every leased classifier, even when the host's source port equals a leased grant's destination port.

The grants of a scope live in one CONSTANT named set, host_grant, and the rule graph does not grow with them: four admissions per scope — ESTABLISHED original direction, a NEW UDP opening, a NEW TCP opening restricted to SYN with FIN/RST/ACK clear, and the ESTABLISHED reply — each check the default conntrack zone and the direction, then look up one concatenated key. The key carries every remaining operand: the variable interface's index and name (the handoff in host transit, the workload link in the router; the guard has none), the packet protocol and the conntrack protocol, the live source, destination and workload port, and the originally tracked source, destination and port. An element is one grant whose live tuple equals its original tuple, so a translated flow never matches and foreign NAT cannot borrow an admission for another flow. The router's raw classification looks up the exact arriving protocol, addresses and port in a second set, host_grant_arrival. Owner verification reads every set back with all of its elements, so apply, inspection, adoption and recovery reject a changed or foreign set exactly as they reject a changed rule; a bound CONSTANT set also refuses element changes from any other socket.

Absent and empty are the same canonical policy, so a scope without grants keeps its exact previous digest and compiled bytes and has no host grant set. Withdrawing a grant is an ordinary atomic replacement: the batch deletes the previous rules, chains and sets by handle and creates the complete target, so a target without the entry, a failed batch or release removes it, and an established flow stops with it. Apply grants hop by hop from the workload side (router, host transit, guard) and withdraw them in the reverse order; each table is atomic on its own, and a partially applied grant only stays dark.

Host-local platform endpoints

A platform endpoint normally lives elsewhere and host transit forwards leased flows to it. When the host serves an endpoint itself — a relay listener, a data port or a VPN hub on an address the host holds — the leased flow ends in the host's INPUT instead, where every handoff arrival is denied. The optional hostTransit.localPlatformEndpoints lists the ids of protection.platformEndpoints the host serves:

protection: { authorityDigest, prefixes: ['10.0.0.0/8', '192.0.2.0/24'], platformEndpoints: [
  { id: 'relay', address: '192.0.2.2', protocol: 'tcp', port: 8443 }] },
localPlatformEndpoints: ['relay'],

Each id names an existing endpoint whose address is an exact requiredIpv4Addresses entry of the uplink or of a handoff link, so rtnetlink verifies that the host holds it at apply, recovery and inspection. The refusal of a platform endpoint on a host address is lifted only for declared ids; unknown, duplicate or unheld ids reject. The router scope grants its workloads the endpoint like any platform endpoint (destination: { kind: 'platform', endpointId }).

INPUT admits the original direction only from the exact handoff that carries a lease, with the lease's transit source address and a source port in one of its ranges for the endpoint's protocol, in the default conntrack zone, to the exact endpoint address and port in both the live and the originally tracked tuple, NEW (a TCP opening only with SYN) or ESTABLISHED. OUTPUT admits only the ESTABLISHED reply back through that handoff. A flow translated in transit, another source address or port, another endpoint port or protocol and any host-origin opening toward the transit address keep meeting the handoff denials. The source port is checked against the leased range in both tuples, not for equality between them. Each leased range adds one INPUT and one OUTPUT jump when an endpoint of its protocol is declared, and each endpoint adds three rules. Absent and empty are the same canonical policy, digest and compiled bytes. The caller keeps a declared endpoint's address outside every guarded allocation pool and owns the listener.

Exclusive forwarding

A host that forwards IPv4 (net.ipv4.conf.all.forwarding=1) routes between all of its links. On a Docker host, Docker set that switch and also the iptables FORWARD policy DROP, and ManagedDockerForwarding admits the handoff flows through it. On a host without Docker nothing denies forwarding that does not touch a handoff: host transit governs only handoff traffic, and the allocation-pool guard only pool destinations, so a LAN peer could use the host as its gateway, the uplink could hairpin, and two other links could exchange traffic. hostTransit.exclusiveForwarding: true closes that path inside the barrier's own table:

const transit: IManagedNftPolicyV2 = { schemaVersion: 2, revision, scope: {
  kind: 'hostTransit', protection, handoffs, uplink, snatAddress, exclusiveForwarding: true } };

The forward chain ends in one unconditional drop, after every rule of the scope. The host then forwards exactly what the scope admits: the leased flows of each handoff to the uplink with their ESTABLISHED replies, the inbound publications and their replies, and the symmetric publications' flows. Everything else that reaches the forward hook drops, whatever its links: IPv4 between the uplink, a LAN or any other link, a hairpin through the uplink, flows that were established before the member was applied, and forwarded IPv6. There is no ESTABLISHED/RELATED bypass, as everywhere in this scope, so ICMP errors for leased flows stay denied as before. Allocation pools need no exception: pool addresses are routed inside the router namespace and leave it translated to the transit source, so on the host only transit addresses cross the forward hook, and the pool guard denies pool destinations there anyway. IPv6 is dropped because the scope's packet model is IPv4 only; a host whose IPv6 forwarding is off never presents IPv6 to the hook. Input, output and NAT are unchanged, and so are the host's own flows: host grants and host-local platform endpoints are INPUT and OUTPUT traffic. With no handoffs the scope forwards nothing at all.

nftables runs the forward base chain of every table, and a drop in any one of them is final, while an accept only ends its own chain. The member therefore drops forwarding that another owner (Docker, a container or VM bridge, a VPN router, a second routing daemon) accepts, and accepting in this table still cannot override another owner's drop. Use it only on a host where this table is the sole forwarding owner; on a Docker host keep it absent and use the Docker contribution below, which refuses an exclusive barrier. inspectForwardHooks(), described below, reports the other forward owners of the namespace for that decision. The member does not fence flowtable offload, packet queues, proxies or anything outside the forward hook.

The drop is part of the scope's graph: it is created, verified, adopted, recovered, replaced and released with the rest of the table in one batch, and a replacement without the member or release() removes it, so forwarding resumes for every link the moment the table goes. This package sets no sysctl. The denial holds only while an exclusive table is applied, so the caller enables host forwarding only after an exclusive scope is enforced and disables it before releasing that scope, or keeps an exclusive scope applied between generations (with no handoffs it is a pure host-wide forwarding denial) and replaces it in place. Persistent forwarding in sysctl.d would make the host forward unfiltered after every boot until the scope is applied again, so leave the boot default off. Absent and false are the same canonical policy, digest and compiled bytes; the member adds one rule.

Workload grants

routerEgress.workloadGrants is optional: a stateful one-way flow between two workloads behind the same router, such as an ingress workload reaching a target workload's service port across organizations. Each entry is { protocol, sourceAddress, destinationAddress, destinationPort }: tcp or udp, the exact current addresses of two different veth workload endpoints of the scope and one port 1–65535. Prefixes, wildcards, ranges, router-local and platform-endpoint addresses, TUN endpoints, both ends on one endpoint and duplicates reject. The source opens and the destination only answers; the reverse direction is a second grant.

Four FORWARD admissions follow the private anti-spoof guards whatever the number of grants: ESTABLISHED original direction, a NEW UDP opening, a NEW TCP opening restricted to SYN with FIN/RST/ACK clear, and the ESTABLISHED reply, each in the default conntrack zone. Every admission looks up the incoming link with the source address and the outgoing link with the destination address (index and name) in the CONSTANT set workload_link, and the protocols with the live and originally tracked tuple in the CONSTANT set workload_grant, so a spoofed, renamed or translated flow never matches and the destination can never open toward the source. Endpoint prefixes are protected, so both directions stay in the default zone ahead of every leased classifier. Owner verification reads both sets back element by element. Up to 1024 grants fit; the sets share the scope's target budget with hostGrants and every other set, and 1024 of each fit together beside the reference router; a scope beyond the budget is refused before any kernel work. Absent and empty are the same canonical policy, digest and compiled bytes, and withdrawal is an ordinary atomic replacement that also stops established flows.

These flows need no private rules. A private rule is stateless: answering through private rules needs a reverse rule, which would also let the destination open toward the source.

Loopback TCP port owners

allocationPoolGuard.localTcpPortOwners is optional and restricts a loopback TCP service to one local user, such as a container runtime's CRI stream server on 127.0.0.1:10010 that only root may reach. Each entry is { address, port, uid }: one exact IPv4 loopback host address in 127.0.0.0/8 (not the network or broadcast address), one port 1–65535 and one uid 0–4294967294. Prefixes, wildcards, IPv6, ranges, uid lists and a second entry for the same address and port reject; the limit is 8 entries. Absent and empty are the same canonical policy, so a guard without owners keeps its exact previous digest and compiled bytes.

const guard: IManagedNftPolicyV2 = { schemaVersion: 2, revision: 1, scope: {
  kind: 'allocationPoolGuard', authorityDigest, prefixes: ['10.240.0.0/16'],
  localTcpPortOwners: [{ address: '127.0.0.1', port: 10010, uid: 0 }] } };

Each entry compiles two rules at the head of the guard's base chains, ahead of the host grant admissions and the pool guard, so nothing earlier in the table accepts past them:

  • OUTPUT (filter priority 0, after ordinary output destination NAT): IPv4 TCP to the exact address and port with meta skuid != uid is rejected with a TCP reset, so another user's connect() fails at once with ECONNREFUSED. This covers every packet, not only the opening, and a flow a foreign output DNAT redirects to the port. meta skuid is the socket file's file-system uid as the user namespace owning the network namespace sees it, the initial one on a host: a process in a user namespace that shares the host network namespace is matched by its host uid, and a user namespace's root is not uid 0. A packet without a socket file carries no uid and does not match: a kernel reply, the teardown of an orphaned socket, or an in-kernel socket such as an NFS or CIFS client, which only a privileged mount creates.
  • INPUT: the exact address and port arriving on any interface but loopback (ifindex 1 in every network namespace) are dropped, so route_localnet or a prerouting DNAT cannot expose the service to another host.

The rules match the exact destination address. The service must bind exactly that address; a wildcard listener stays reachable through the host's other addresses. IPv6 is not covered, so the service must not listen on ::1 or ::. A socket keeps the uid it was created with, so a root process that hands its connected socket to another user delegates that connection. The reject expression needs the kernel's nft_reject_inet module (autoloaded on stock kernels). Owner verification reads back every rule, including the comparison operator, the uid and the reject type, so apply, inspection, adoption and recovery reject a changed rule. The entries are applied, retained, recovered and released with the rest of the guard's table; after a reboot the caller applies its retained intent again, like every other member.

Input bounds are 1024 grants per scope, the maximum a network projection carries. Elements are sent 256 to a message. At 1024 grants the reference router needs about 98 KB of elements (two sets), host transit about 66 KB and the guard about 45 KB.

Every replacement is one nfnetlink batch in one sendmsg, and Linux refuses a netlink message larger than the socket send buffer: the requested 1 MiB SO_SNDBUF is clamped to net.core.wmem_max and doubled, 425,984 bytes with the Linux default of 212,992. That is the hard ceiling. This package bounds a batch at 320,000 bytes, which fits every wmem_max of at least 160,016, and splits it: at most 64 bytes of batch framing, at most 107,520 bytes for the handle-only deletion of the previous graph (fewer than 768 rules, chains and sets of at most 140 bytes each), and 212,000 bytes for the complete target, of which rules, chains and sets may take at most 100,000 and set elements the rest. Scopes that grow with their inputs keep those inputs in set elements, so the element share grows where the rule share shrinks; prepare() refuses a target beyond either bound with EXHAUSTED before any kernel work. The private facade IPC carries up to three complete policies in one status and is bounded at 1,048,576 bytes per line; the facade captures at most 1,000,000 JSON bytes and 65,536 values per request or result. The caller owns the host route to the workload through the handoff, selecting the transit host address as the socket's source, the workload listener and every other packet owner.

Active overlapping classifiers, duplicate zones/labels and conflicting handoff allocations reject. Empty router generations retain private routing while denying egress. Router input bounds are 128 private endpoints and links, 1024 private rules, 32 active generations and 1024 total egress grants; other bounds include 128 protected prefixes, 96 platform endpoints, 16 ranges per allocation, and 32 host handoffs/active allocations. The complete target must fit the budget above; the router's rules grow only with its protected prefixes, so its elements bind first. Exhausted source-port ranges drop new flows rather than allocate outside the lease.

The router compiles its private endpoints, directed rules and leased grants into named sets and maps behind a fixed set of rules, whatever the number of workloads:

  • private_link holds every endpoint link (index and name) and the absent interface of router-local traffic; private_index and private_name hold every endpoint index and name for the terminal denials, either alone enough to deny; private_source holds each link index with its source prefixes. The anti-spoof guard drops a packet on an endpoint link whose address is not one of its prefixes, and anything but IPv4 there.
  • private_any and private_port hold the directed rules: incoming and outgoing link index (0 for the router itself), addresses, and for a transport rule the protocol and ports.
  • Per kind (platform ahead of the protected barrier, public behind it) and per shape (whether the grant names a source port), leased_*_zone maps a workload or router flow's link, protocol and tuple to its generation zone before conntrack, leased_*_return maps a translated return to its zone, and leased_*_flow maps the link, both protocols, the live and originally tracked tuple and the zone to the grant's source NAT; the filter admissions look it up as a set. leased_label holds each generation's zone and label, and leased_generation maps a zone to the label a fresh flow receives. A key's link index names its link because the rule also finds the interface's index and name in private_link, whose indices are unique. Every element is checked, like every rule, on apply, inspection, adoption and recovery. A decision corpus of about 133,000 packets across five router fixtures, frozen from the per-rule compiler of 2.6.0, is reproduced exactly by the set-backed compiler.

A router of 64 workloads, each with local DNS, public TCP and UDP egress, grants to four platform endpoints and one to three publications, compiles about 66,000 rule bytes and 105,000 element bytes; 89 such workloads fit the target budget:

Workloads Router rules Router elements Host rules Host elements
1 65,820 4,968 42,944 1,128
8 65,820 16,056 42,944 2,528
32 65,820 54,072 42,944 7,328
64 65,820 104,760 42,944 13,728

Upgrading from 2.x: version 3 changes the compiled identity of every schema-v2 routerEgress table and of every scope with publishedPorts, while their policy digests stay unchanged. Reconcile or release against such a table applied by 2.x returns Conflict: the new engine cannot adopt the old graph. Recover and release each such table with the exact previous engine and its retained intent and receipt, under the same traffic fence as below, before upgrading; tables of other scopes and without publications are unchanged.

Upgrading from 1.x: version 2 changes the compiled identity of schema-v2 routerEgress tables. Retained router tables and their receipts cannot be adopted by the new engine, even when the policy digest is unchanged. Fence workload and router-local traffic, then use the exact previous engine and retained intent/receipt to recover and release each router table. Join that owner before upgrading. Prepare and apply the policy with the new engine and persist its complete receipt before lifting the traffic fence. Keep lease quarantine and host/pool protection throughout; this migration provides no proof of conntrack drainage or safe allocation reuse. Do not discard an unresolved retained table or its old engine/receipt. Schema-v1, host-transit and allocation-pool-guard compiled identities are unchanged.

The caller must coordinate router and host apply order, retain complete intents and receipts, authenticate protected authority, prevent reuse of quarantined leases, and qualify other packet owners. An ACCEPT in this table cannot override Docker's independent FORWARD DROP. ManagedDockerForwarding, described below, owns the separate DOCKER-USER contribution. The caller must also control deferred packets, proxy/BPF/ offload paths and changing network authority. No table receipt proves flow drainage, conntrack cleanup, elapsed-time expiry, or safe address/port/zone reuse.

Docker forwarding contribution

ManagedDockerForwarding admits exactly the forwarded flows of an exact applied schema-v2 hostTransit receipt through Docker's existing IPv4 forwarding path. It reads and verifies that dedicated barrier's complete table and local links, without adopting its ownership. The contribution uses Docker's supported DOCKER-USER extension point. Docker's native nftables backend has no equivalent user chain and is unsupported.

The host must provide root-owned /usr/sbin/iptables-nft, /usr/sbin/iptables-nft-save and /usr/sbin/iptables-nft-restore, using the same 1.8.10-or-newer 1.8-series nf_tables frontend. prepare() captures the exact frontend version in its digest. Apply, inspection and recovery require that same version. Subprocesses use fixed argv, a clean environment, bounded input/output, one shared eight-second frontend deadline per request, and joined termination. No shell or raw rule API is exposed.

One node-level contribution owns a contiguous leading block in DOCKER-USER. FORWARD must already have policy DROP and its first two unconditional jumps must be DOCKER-USER and DOCKER-FORWARD. Their exact ordered raw graphs are checked without claiming their ownership. Other managed contribution owners, altered or duplicated owned rules, missing chains and changed placement reject. Unrelated rules following the owned block and Docker's own chains remain separate owners. Saved text verifies placement, count and normalized policy. Native netlink reads verify every contributed rule's complete ordered expression graph, including register flow and all match bytes; only counter values and attribute encoding order/flags are normalized. An exact xtables -C check also verifies each compiled deletion specification. These reads must share an unchanged ruleset generation; concurrent Docker reconciliation can therefore make inspection inconclusive. Conntrack source addresses use the canonical bare IPv4 form because an explicit /32 has different hidden match bytes in these frontends. A semantically equal rewrite with a different exact representation rejects, as does an early ACCEPT that the frontend reconstructs as the same command. The contribution never creates, flushes, adopts or deletes a Docker table/chain, and never changes a host forwarding policy. prepare() refuses a barrier with exclusiveForwarding as INVALID: its drop would also deny Docker's own container forwarding.

The contribution and the barrier's forward chain derive from one enumeration of the scope's forwarded flows, so the contribution admits exactly the flow kinds the non-exclusive barrier admits, each with its replies:

Flow kind Original direction Live tuple Conntrack original tuple
Leased range handoff → uplink transit source address and leased port range same source
Inbound publication uplink → handoff, after the barrier's DNAT inside target address and port(s) as destination published hostIp and port(s) as destination
Symmetric publication handoff → uplink, before the barrier's SNAT target address and port(s) as source same source

Every flow kind has three rules: the NEW opening (TCP only with SYN and FIN/RST/ACK clear), the ESTABLISHED original direction and the ESTABLISHED reply with swapped interfaces and tuple sides. A router's own egress, such as a VPN client in the router namespace dialling its hub, reaches the host as a leased flow: the router translates it to the transit address and a leased port range, so it is admitted by the leased rules exactly when the barrier admits it. Host grants and host-local platform endpoints use INPUT/OUTPUT and never cross FORWARD. Each rule binds the handoff and uplink names. The exact v2 host barrier independently checks interface indices, MAC/iflink facts, the default host conntrack zone, protected original/current destinations and explicit outer SNAT; it governs every forwarded packet on its handoffs, so the host forwards the intersection, which is the barrier's admitted set. A scope without publications keeps its exact 4.1.0 rules, comments and graphs. A complete contribution is bounded to 192 rules and 100,000 generated command bytes: three rules per leased range and per publication, and three more per symmetric publication. A larger scope is refused as EXHAUSTED (restoreRules/restoreBytes) before any subprocess.

Create the class with ownerId, a caller-retained instanceId of at least 16 identifier characters, and optional binaryPath/networkNamespaceFd. After start(), call prepare({ schemaVersion: 1, revision, barrier }), where barrier is the exact applied host policy. Persist the complete { previous, target } transition before reconcile(). The native owner fences the handoff links before and after its single no-flush restore transaction, which deletes the previous block and inserts the target block in one atomic nf_tables commit. A handoff link that previous and target both bind exactly may stay up, so a publication or lease amendment of a running generation applies in place: every flow both contributions admit passes throughout, a flow only the target admits passes from the commit, and the referenced barrier must already be the target. Every handoff link the transition adds or drops, and every link of a first apply, must be down; retain those exact links until cleanup finishes. Exact target replay can recover a lost response without rewriting live rules.

inspect().present confirms the contribution and its referenced barrier, not the complete host packet path or durable controller authority. Errors retain failed-owned state and pending intent. release(applied) and close() delete only exact contributed rules. close() first reads the filter table through the same generation-consistent frontend and netlink read, without requiring Docker's forwarding path: when no saved rule in any chain and no native DOCKER-USER or FORWARD rule carries this owner's comment prefix, cleanup is confirmed. So an owner whose first apply was refused, for example on a host with FORWARD policy ACCEPT or without Docker, closes on that host. A read that fails rejects CLEANUP_UNCONFIRMED whose cause carries the native code (UNAVAILABLE, UNSUPPORTED_BACKEND, PERMISSION, CONFLICT); it is never reported as cleanup. The fence of release(applied) and close() requires every handoff link down or gone: a link whose index no link holds any more, such as a veth that left with its router namespace when the node process ended, counts as down, so a restarted process in the same boot can release the previous contribution before applying the next one. A different link at a bound index still rejects. A cold boot needs a new barrier receipt and caller-authorized replay; old boot receipts cannot authorize native operations. The caller owns restart ordering, disjoint retained leases, exclusive privileged mutation authority and activation. The frontend transaction does not provide compare-and-swap against other privileged processes.

Forward hook owners

inspectForwardHooks({ binaryPath?, networkNamespaceFd? }) reads every packet-filter owner on the routed forward path of one network namespace and changes nothing. One native process takes the read and exits:

  • chains: every nf_tables base chain on the forward hook of the ip, ip6 and inet families, with family, table, tableHandle, chain, type, priority and policy, sorted by family, table and chain. Managed tables are included; compare table and tableHandle with a receipt's identity.tableName and tableHandle to find your own.
  • docker: whether Docker's iptables-nft DOCKER-USER and DOCKER-FORWARD chains exist in table ip filter.
  • legacyTables: the registered legacy xtables ipv4 and ipv6 tables. Their hooks, including FORWARD, are not nf_tables objects and never appear in chains.

The table and chain dumps share one ruleset generation, or the read is CONFLICT. Without CAP_NET_ADMIN in the namespace it is PERMISSION. Bridge and netdev hooks do not see routed packets and are not reported. Guard exclusiveForwarding with it: refuse while any chain other than your own barrier's, either Docker chain or any legacy table is present. The read is a point in time; the caller still owns exclusive authority over who may add a forward owner later.

Native qualification

The separate Docker fixture exercises the native contribution with pinned Docker packages on offline Ubuntu 26.04 and Ubuntu 24.04.4/HWE guests, including retained-state cold boots and published workloads. Its checked-in input manifest and per-run receipts identify the tested artifacts.

cargo test --manifest-path rust/Cargo.toml --locked runs unprivileged compiler, snapshot and bounded-subprocess tests. The ignored native cases require the explicitly marked disposable guest from test/native/qualify.py; requesting them on an ordinary host fails its scope check. Qualification uses an offline Linux 6.18.35 x86_64 guest with no host disks, mounts or external network backend. One VirtIO device connects only to a singleton QEMU-internal hub; packet-path tests use guest-owned namespaces, veth pairs and TUN. It covers UDP/TCP grants, direct-IP denial, source spoofing, IPv6 denial, renamed interfaces, foreign ownership, atomic rollback, generation conflicts and lost-ACK recovery. V2 tests exercise complete graph persistence/recovery, private and local DNS, local TCP, TUN isolation, observed UDP/TCP handoff ranges, range exhaustion, fragment reassembly, ECN SYN and invalid opening flags, protected current/original DNAT barriers, nonzero foreign zones, local diversion, IPv6 denial, revocation and restoration with disjoint generation ranges while old flows remain retained. Published host ports are qualified end to end: translated TCP and UDP delivery to the leased transit address with the actual client address preserved, reverse translation back to the published address, SYN-only opening, denial of non-SYN openings, denial of unpublished workload and uplink ports against a pre-policy positive control, and the port going dark with the withdrawn and released generation. Published workload ports are qualified across both hops in the same four-namespace topology: TCP and UDP reach a listener inside the workload namespace with the client address preserved and both translations reversed, only an exact SYN opens the publication against a positive control on the same injector, an unpublished transit port and the unpublished workload path stay behind the router barrier, and withdrawing only the router publication or releasing it takes the port dark while the host publication remains. A client whose source port equals a leased grant's destination port on that grant's public address completes the published TCP handshake and receives its UDP reply from the published address, so the leased classifiers of the same generation cannot capture the workload's published packets. Published ranges and symmetric publications are qualified across both hops in the same topology, with every publication a set element whose kernel dump matches the compiled sets and maps: the first, middle and last port of a UDP range reach their listeners with the client address preserved and answer from the published address, the ports just outside the range stay dark against pre-policy positive controls, symmetric UDP from a range port and from the signalling port and symmetric TCP from the signalling port reach the uplink peer from the published address with the same port, a later request from that peer reaches the workload on the same flow, the plain publication's workload port opens nothing, symmetric flows into the protected union stay denied against a pre-policy positive control, leased egress keeps its leased range, and the range goes dark with release. A 64-workload router is qualified in one batch, every workload in its own namespace with local DNS, public TCP and UDP egress, four platform endpoints and a publication, the first with the SIP shape: the kernel dumps every private, leased and published element and the owner verifies it, and on the first, middle and last workload local DNS answers, public UDP egress leaves from the leased range, UDP and TCP platform endpoints answer through both hops, the publication reaches its listener with the client address preserved, and another workload, a protected address outside the platform endpoints and an ungranted platform port stay denied against pre-policy positive controls, as does another workload's source address. The router compiles 82 rules for the 64 workloads. Host grants are qualified across all three tables in the same four-namespace topology with leased public egress on every port of both protocols and 1024 grants in every scope: a member far from the probed tuples passes, the tuple just outside the set stays dark, and against pre-policy positive controls, the host stays dark without grants and with router and transit grants alone (the guard still denies), the exact TCP and UDP tuples reach the workload from the transit host address once the guard exception is applied, another port, protocol, host source address and the workload-origin direction stay denied, withdrawing only the router grant closes the path and re-granting reopens it, and withdrawal in every scope stops new flows and the established one. A full guard set persists across owner loss, is re-verified element by element before adoption, survives lost-ACK replay and is replaced and released like the rest of the graph, and another socket cannot add or delete one of its elements. Host-local platform endpoints are qualified from a router namespace against pre-policy positive controls: without the member the handoff denials keep both a TCP endpoint on the uplink address and a UDP endpoint on the handoff address dark; with it the exact leased flows reach them with the transit source preserved, while a source port outside the lease or in the other protocol's range, an unleased source address, another endpoint port, a flow a foreign prerouting DNAT translated onto the endpoint and the host's own opening toward the transit address stay dark; withdrawal stops new flows and the established one, and release reopens the path. A graph with served endpoints persists across owner loss and survives lost-ACK replay. Workload grants are qualified between three workload namespaces behind the router, each with leased public egress on every port of both protocols: against pre-policy positive controls, the grant opens exactly the source workload's TCP and UDP flows to the granted ports with the source preserved; another port, another source workload, another destination workload, the destination's TCP and UDP openings toward the source stay dark; withdrawal stops new flows and the established one; release reopens the path. A set of 1024 grants persists across owner loss, is re-verified element by element and survives lost-ACK replay. Exclusive forwarding is qualified on a host namespace with a handoff, an uplink and a LAN link: against positive controls with no policy and with the scope without the member, the member denies the LAN's IPv4 to the uplink peer, the peer's way back into the LAN, the LAN's IPv6 and a LAN flow established before it, while leased egress still leaves translated and an unleased port stays denied; without handoffs leased egress drops too; replacing it with a scope without the member and releasing an exclusive scope both reopen the LAN path. The same probe fails against a compiler without the drop. Loopback TCP port owners are qualified against pre-policy positive controls on the same listeners: the owning uid (root, and uid 1000 for a second port) connects, every other uid is reset with ECONNREFUSED, an unowned port and another loopback address stay open to every user, a foreign output DNAT to the owned port stays closed to another user, and an uplink arrival translated to loopback with route_localnet enabled is dropped. The policy keeps enforcing after the owner detaches and its process is gone, a fresh owner verifies and adopts the exact graph, and release reopens every probe. The maximum of eight owners beside the guard persists across owner loss and survives lost-ACK replay. The arm64 binary is cross-built; native packet qualification is currently x86_64. Kernel 6.8 is unsupported. No production activation is implied by these tests.

Native dependency, standard-library and C runtime attribution is included in native-notices/, generated by tsrust notices from rust/Cargo.lock and the Rust toolchain pinned in rust-toolchain.toml. Every build verifies the committed notices first and refuses when they no longer match. Distribution must retain these notices alongside the binaries. The statically linked musl 1.2.5 comes from the Rust 1.95.0 toolchain pinned in rust-toolchain.toml, whose musl build includes the two CVE-2025-26519 patches.

Quick Start

import { SmartNftables } from '@push.rocks/smartnftables';

const nft = new SmartNftables();
await nft.initialize();

// Port forward 8080 → 192.168.1.100:80
await nft.nat.addPortForwarding('web', {
  sourcePort: 8080,
  targetHost: '192.168.1.100',
  targetPort: 80,
});

// Block a suspicious IP
await nft.firewall.blockIP('10.0.0.99');

// Rate limit HTTP to 100 req/s per IP
await nft.rateLimit.addRateLimit('http-limit', {
  port: 80,
  protocol: 'tcp',
  rate: '100/second',
  perSourceIP: true,
});

// Clean up everything when done
await nft.cleanup();

Architecture 🏗️

The library is organized around a facade pattern with specialized sub-managers:

SmartNftables (main facade)
├── nat          → NatManager       (DNAT, SNAT, masquerade)
├── firewall     → FirewallManager  (filter rules, IP sets, stateful tracking)
└── rateLimit    → RateLimitManager (packet/connection rate limiting)

All rules are tracked in rule groups identified by string IDs, so you can add, inspect, and remove them programmatically.

API Reference

SmartNftables — Main Facade

const nft = new SmartNftables({
  tableName: 'smartnftables', // nftables table name (default: 'smartnftables')
  family: 'ip',               // 'ip' | 'ip6' | 'inet' (default: 'ip')
  dryRun: false,               // generate commands without executing (default: false)
});
Method Description
initialize() Create the nftables table and NAT chains. Idempotent.
cleanup() Delete the entire table and clear all tracking.
status() Get an INftStatus report of the current managed state.
applyRuleGroup(id, commands) Apply and track a group of raw nft commands.
removeRuleGroup(id) Remove a tracked rule group.
getRuleGroup(id) Retrieve a tracked rule group by ID.

🌐 NAT — nft.nat

Port Forwarding (DNAT)

await nft.nat.addPortForwarding('my-service', {
  sourcePort: 443,
  targetHost: '10.0.0.5',
  targetPort: 8443,
  protocol: 'tcp',           // 'tcp' | 'udp' | 'both' (default: 'tcp')
  preserveSourceIP: false,    // skip masquerade if true (default: false)
});

await nft.nat.removePortForwarding('my-service');

Port Range Forwarding

Map a range of ports to another host:

// Forward ports 3000-3010 → 10.0.0.5:3000-3010
await nft.nat.addPortRange('dev-ports', 3000, 3010, '10.0.0.5', 3000, 'tcp');
await nft.nat.removePortRange('dev-ports');

SNAT (Source NAT)

await nft.nat.addSnat('egress', {
  sourceAddress: '203.0.113.1',
  targetPort: 80,
  protocol: 'tcp',
});

Masquerade

await nft.nat.addMasquerade('outbound', {
  targetPort: 443,
  protocol: 'tcp',
});

🛡️ Firewall — nft.firewall

Basic Rules

await nft.firewall.addRule('allow-ssh', {
  direction: 'input',         // 'input' | 'output' | 'forward'
  action: 'accept',           // 'accept' | 'drop' | 'reject'
  sourceIP: '10.0.0.0/24',
  destPort: 22,
  protocol: 'tcp',
  comment: 'Allow SSH from trusted network',
});

await nft.firewall.removeRule('allow-ssh');

When sourcePort or destPort is provided without protocol, TCP is used by default. Protocol-only rules match the specified layer-4 protocol.

Block an IP

await nft.firewall.blockIP('10.0.0.99');
await nft.firewall.blockIP('192.168.0.0/16', { direction: 'forward' });

Allow Only Specific IPs on a Port

// Only these IPs can reach port 3306 — everything else is dropped
await nft.firewall.allowOnlyIPs('db-access', ['10.0.0.1', '10.0.0.2'], 3306, 'tcp');

Stateful Connection Tracking

// Allow established/related, drop invalid — on the input chain
await nft.firewall.enableStatefulTracking('input');

IP Sets

Create named sets and match against them:

// Create a set of blocked IPs
await nft.firewall.createIPSet({
  name: 'blocklist',
  type: 'ipv4_addr',
  elements: ['10.0.0.50', '10.0.0.51'],
});

// Dynamically add/remove elements
await nft.firewall.addToIPSet('blocklist', ['10.0.0.52']);
await nft.firewall.removeFromIPSet('blocklist', ['10.0.0.50']);

// Clean up
await nft.firewall.deleteIPSet('blocklist');

You can also build set-matching rules directly with the low-level builder:

import { buildIPSetMatchRule } from '@push.rocks/smartnftables';

const rule = buildIPSetMatchRule('smartnftables', 'ip', {
  setName: 'blocklist',
  direction: 'input',
  matchField: 'saddr',
  action: 'drop',
});

⏱️ Rate Limiting — nft.rateLimit

Packet Rate Limiting

// Global: drop packets over 1000/second on port 80
await nft.rateLimit.addRateLimit('http-global', {
  port: 80,
  protocol: 'tcp',
  rate: '1000/second',
  burst: 50,
  action: 'drop',
});

// Per-IP: each source IP gets its own 100/second limit
await nft.rateLimit.addRateLimit('http-per-ip', {
  port: 80,
  protocol: 'tcp',
  rate: '100/second',
  perSourceIP: true,
});

await nft.rateLimit.removeRateLimit('http-per-ip');

Connection Rate Limiting

Limit the rate of new connections (uses ct state new):

await nft.rateLimit.addConnectionRateLimit('ssh-connrate', {
  port: 22,
  protocol: 'tcp',
  rate: '5/second',
  perSourceIP: true,
});

await nft.rateLimit.removeConnectionRateLimit('ssh-connrate');

🔧 Low-Level Rule Builders

For advanced use cases, you can generate raw nft command strings without applying them:

import {
  buildDnatRules,
  buildSnatRule,
  buildMasqueradeRule,
  buildFirewallRule,
  buildRateLimitRule,
  buildPerIpRateLimitRule,
  buildConnectionRateRule,
  buildIPSetCreate,
  buildIPSetAddElements,
  buildIPSetRemoveElements,
  buildIPSetDelete,
  buildIPSetMatchRule,
  buildTableSetup,
  buildFilterChains,
  buildTableCleanup,
} from '@push.rocks/smartnftables';

const commands = buildDnatRules('mytable', 'ip', {
  sourcePort: 8080,
  targetHost: '10.0.0.5',
  targetPort: 80,
});
// → ['nft add rule ip mytable prerouting tcp dport 8080 dnat to 10.0.0.5:80',
//    'nft add rule ip mytable postrouting tcp dport 80 masquerade']

Dry Run Mode 🧪

Generate commands without touching the kernel — perfect for testing, debugging, or CI:

const nft = new SmartNftables({ dryRun: true });
await nft.initialize();
await nft.nat.addPortForwarding('test', {
  sourcePort: 80,
  targetHost: '10.0.0.1',
  targetPort: 8080,
});

console.log(nft.status());
// Rules tracked in memory, nothing executed

Status Reporting 📊

const status = nft.status();
// {
//   initialized: true,
//   tableName: 'smartnftables',
//   family: 'ip',
//   isRoot: true,
//   activeGroups: 3,
//   groups: {
//     'nat:web': { ruleCount: 2, createdAt: 1711411200000 },
//     'fw:block-10_0_0_99': { ruleCount: 1, createdAt: 1711411200100 },
//     'ratelimit:http-limit': { ruleCount: 1, createdAt: 1711411200200 },
//   }
// }

Types

All interfaces and types are fully exported for use in your own code:

Type Description
INftDnatRule DNAT port forwarding rule config
INftSnatRule Source NAT rule config
INftMasqueradeRule Masquerade rule config
INftFirewallRule Firewall filter rule config
INftIPSetConfig IP set creation config
INftRateLimitRule Rate limiting rule config
INftConnectionRateRule New-connection rate limit config
ISmartNftablesOptions Constructor options
INftStatus Status report shape
TNftProtocol 'tcp' | 'udp' | 'both'
TNftFamily 'ip' | 'ip6' | 'inet'
TFirewallAction 'accept' | 'drop' | 'reject'
TCtState 'new' | 'established' | 'related' | 'invalid'

This repository contains open-source code licensed under the MIT License. A copy of the license can be found in the repository license file.

Please note: The MIT License does not grant permission to use the trade names, trademarks, service marks, or product names of the project, except as required for reasonable and customary use in describing the origin of the work and reproducing the content of the NOTICE file.

Trademarks

This project is owned and maintained by Task Venture Capital GmbH. The names and logos associated with Task Venture Capital GmbH and any related products or services are trademarks of Task Venture Capital GmbH or third parties, and are not included within the scope of the MIT license granted herein.

Use of these trademarks must comply with Task Venture Capital GmbH's Trademark Guidelines or the guidelines of the respective third-party owners, and any usage must be approved in writing. Third-party trademarks used herein are the property of their respective owners and used only in a descriptive manner, e.g. for an implementation of an API or similar.

Company Information

Task Venture Capital GmbH
Registered at District Court Bremen HRB 35230 HB, Germany

For any legal inquiries or further information, please contact us via email at hello@task.vc.

By using this repository, you acknowledge that you have read this section, agree to comply with its terms, and understand that the licensing of the code does not imply endorsement by Task Venture Capital GmbH of any derivative works.

S
Description
No description provided
Readme
4 MiB
Languages
Rust 53.1%
HTML 29.3%
TypeScript 14.3%
Python 3.3%