@push.rocks/smartnftables
A TypeScript module for managing Linux nftables rules with a high-level, type-safe API. Handles NAT (DNAT/SNAT/masquerade), firewall rules, IP sets, and rate limiting — all from clean, declarative TypeScript.
Issue Reporting and Security
For reporting bugs, issues, or security vulnerabilities, please visit community.foss.global/. This is the central community hub for all issue reporting. Developers who sign and comply with our contribution agreement and go through identification can also get a code.foss.global/ account to submit Pull Requests directly.
Install
pnpm install @push.rocks/smartnftables
# or
npm install @push.rocks/smartnftables
The legacy SmartNftables helper tracks rules without applying them when it lacks
root privileges. ManagedNftables always rejects unsupported enforcement; it has
no dry-run or memory-only enforcement mode.
Managed workload policy
ManagedNftables provides a separate native owner for complete, interface-bound
IPv4 policy. It requires Linux 6.9 or newer, CAP_NET_ADMIN in its network
namespace, and working nftables OWNER/PERSIST support. The published native matrix
is Linux x86_64 and arm64, statically linked with musl. Actual kernel support is
checked when applying policy; starting the compiler does not establish enforcement.
An optional caller-owned networkNamespaceFd selects an already-created Linux
network namespace. Keep that descriptor valid through start(). The native
process validates and enters its inherited copy before readiness, owner identity
capture or netlink activity, then closes the copy. Invalid descriptors and denied
entry fail before readiness. The option is process configuration and never enters
openOwner, policy commands or durable receipts. Native callers use
--management --network-namespace-fd 3 with an inherited descriptor at FD 3.
The caller owns the namespace lifetime, links and routes. Retain a namespace owner across process loss and recovery: PERSIST retains policy only while its namespace exists. An entered process keeps its namespace membership if the source descriptor is closed, but a replacement process requires a valid descriptor. Receipts bind the actual target namespace device and inode; recovery in another namespace rejects. Await complete policy and process cleanup before releasing the caller's final namespace owner. Namespace entry does not create uplink authority, drain conntrack flows or authorize address and port reuse.
The Linux x86_64 musl candidate passes isolated Linux 6.18.35 kernel tests, including inherited namespace entry, retained policy after process loss, recovery in the original namespace, rejection in another namespace, and denied entry before readiness. ARM binaries are built; privileged ARM and complete Pallet workload integration remain separate qualification requirements.
The caller owns workload links and durable state. Keep links disabled until their
policy is confirmed, serialize link changes with policy operations, and disable
them before cleanup. Persist the full prepared transition before calling
reconcile(), then retain the full returned receipt in the caller's SmartData
store. This library does not persist runtime state in files.
import { ManagedNftables } from '@push.rocks/smartnftables';
const owner = new ManagedNftables({
ownerId: 'node-policy',
instanceId: 'caller_retained_unique_instance',
tableName: 'snft_workloads',
});
await owner.start(); // Compiler/IPC readiness only; no kernel policy yet.
const target = await owner.prepare({
schemaVersion: 1,
revision: 1,
endpoints: [
{ id: 'workload-a', interfaceIndex: 12, interfaceName: 'worka', interfaceKind: 'veth', sourcePrefixes: ['10.81.1.2/32'] },
{ id: 'workload-b', interfaceIndex: 13, interfaceName: 'workb', interfaceKind: 'veth', sourcePrefixes: ['10.81.2.2/32'] },
],
rules: [
{
sourceEndpoint: 'workload-a', destinationEndpoint: 'workload-b',
sourcePrefix: '10.81.1.2/32', destinationPrefix: '10.81.2.2/32',
protocol: 'tcp', sourcePort: null, destinationPort: 443,
},
{
sourceEndpoint: 'workload-b', destinationEndpoint: 'workload-a',
sourcePrefix: '10.81.2.2/32', destinationPrefix: '10.81.1.2/32',
protocol: 'tcp', sourcePort: 443, destinationPort: null,
},
],
});
const transition = { previous: null, target };
// Persist this complete transition before the next operation.
const applied = await owner.reconcile(transition);
// Persist the complete applied result before enabling the caller-owned links.
const status = await owner.inspect();
if (!status.enforced) throw new Error('Workload policy is not confirmed.');
// After disabling/quiescing these workload links:
await owner.release(applied);
await owner.close();
Interface numbers and names in this example must be replaced with the caller's
actual, already-created links. Native RTM_GETLINK checks the name/index pair,
the declared veth or L3 tun kind and absence of a bridge/master attachment.
The caller must prove and own each veth peer and workload network namespace.
Both selectors must match for a grant. A stale
name/index pair rejects application; renaming either selector retains denial.
Do not reuse interface identities or workload addresses while their previous
grants remain outstanding.
| Method | Behavior |
|---|---|
start() |
Starts one private native process and returns idle status. |
failureSignal |
Stable per-instance AbortSignal for lost owner confidence, including an unexpected idle process exit. It never resets or confirms cleanup. |
prepare(policy) |
Validates, canonicalizes and hashes an inert complete policy without kernel writes. |
reconcile({ previous, target }) |
Creates or replaces the entire policy in one kernel transaction. previous is null initially, otherwise the full prior applied result. |
inspect() |
Reads the current owned graph and reports enforced, pending intent and a bounded error code. |
release(applied) |
Verifies the exact retained graph and confirms deletion; repeats are idempotent. |
releaseTransition({ identity, transition }) |
Selects terminal cleanup using the original full identity and transition when the applied result is unavailable. Accepts absence or deletes the exact previous/target graph without requiring surviving interfaces. |
detach(applied) |
Selects terminal retention, verifies the same graph with PERSIST-only flags after dropping socket ownership, then joins the native process. Returns { detached: true }; retain the full applied result. |
closeRetaining() |
Stops command admission, joins admitted operations and terminates the native process without requesting deletion or confirming retention. Use when the exact applied result is unavailable. |
close() |
Stops command admission, joins admitted policy operations, confirms owned-table deletion unless retention was selected, then joins the native process. Failure retains the owner for explicit recovery. |
Schema-v1 policies are directed. Return traffic requires its own grant; there is no broad
connection-tracking bypass. A null endpoint denotes the actual local host, with an
explicit prefix that cannot overlap any endpoint. Such a grant applies only to
input/output. Forwarding requires two explicit endpoint references; a relay TUN
is an endpoint with interfaceKind: 'tun' and its authenticated remote source
prefixes. The relay owner must authenticate that remote source authority.
Managed interfaces reject other IPv4 traffic and IPv6 traffic. Ethernet/ARP
is outside these inet hooks; workload links must use the qualified routed layout.
Unmatched interfaces retain their existing behavior. The compiler creates
only inet input, forward and output filter chains; it does not configure links,
routes, NAT, DNS, host defaults, forwarding sysctls, or relay sessions. An ungranted
destination is denied independently of DNS answers.
The schema permits at most 32 interfaces, 16 source prefixes per interface and 128 directed rules, additionally bounded to 768 compiled operations and 100,000 encoded bytes. IPv4 prefixes must be canonical, non-overlapping local allocations; loopback, unspecified and multicast allocations reject. Ports require explicit TCP/UDP and all nullable fields must be present. The facade rejects getters, proxies, live objects and concurrent commands rather than queuing unbounded work.
One NETLINK_NETFILTER socket exclusively owns the table with OWNER/PERSIST flags. An acknowledged batch commits the complete graph atomically. Lost acknowledgements retain the exact pending transition; retry only that same body. Reconnect/restart uses the same caller-retained owner, instance, table and boot/namespace receipt. Recovery checks the complete graph before and after orphan adoption. The kernel's wrapping generation counter only fences concurrent transactions; it is never a durable revision or sufficient ownership proof. Foreign owners, unexpected tables, chains, rules, sets or objects inside the owned table reject without deletion. Other tables coexist independently. Linux's family-wide chain dump is explicitly filtered by table identity before graph comparison; foreign chains are never included in replacement or cleanup.
PERSIST deliberately retains accepted static grants if the native process crashes. It supplies no wall-clock lease or timed revocation. The caller must keep revocation pending until policy replacement or workload fencing is confirmed. Host reboot invalidates the old boot receipt; the caller must fence its old link/workload lifecycle before starting a new owner instance. Cleanup failures and uncertain operations must not be reported as successful release or used to transfer authority.
Subscribe to failureSignal before starting or admitting policy work, and check
failureSignal.aborted when attaching to an existing instance. Unexpected child
exit, ambiguous operation failures, native inspection reporting failed ownership,
and unconfirmed cleanup abort the signal once with ManagedNftablesError code
OWNER_LOST. Local input rejection, effect-free native compiler validation and
native INVALID and EXHAUSTED responses do not abort it. Other native errors can occur after
kernel effects or ownership loss and conservatively abort it. Successful orderly
close, including a cancelled startup, does not signal an unexpected failure;
ambiguous mutations still signal loss while close is draining them.
The caller must fence packet admission on loss. Recovery operations and exact retries remain available as documented, but successful recovery never resets the signal. It does not monitor foreign kernel changes while idle or prove policy withdrawal: PERSIST policy can remain after process death. Still await the chosen joined cleanup method and retain unresolved policy/allocation authority.
For an orderly restart that must preserve policy, persist the complete applied
result and call await owner.detach(applied) before calling close(). The native
operation validates the exact receipt, graph, interfaces, boot and namespace. It
accepts this socket's OWNER/PERSIST table or an exact orphan; it never adopts a
foreign owner. After closing the owner socket it reads the graph again and requires
the same handle and compiled body with PERSIST-only flags. Native inspection then
reports state: 'retained'. The facade returns success only after child termination.
Calling detach() irreversibly selects retention cleanup, including invalid input,
transport failure, timeout, native rejection or malformed acknowledgement. Subsequent
close() only joins and terminates; it never sends a table-deletion command. It
cannot undo a close() already admitted earlier. Other policy operations reject.
Resolve any pending policy transition by exact replay before requesting detachment.
If a detach ACK is lost while the process remains available, only the same captured
body may be retried. After process loss, start a fresh owner with identical options
and retry detach(applied) against the exact orphan. Absent graphs, foreign active
owners, changed interfaces, handles, bodies or boot/namespace identities reject.
Native release/reconcile/close commands reject after a valid detach intent, even
when post-drop verification fails; explicit recovery requires the retained body.
If reconciliation remains ambiguous and no complete applied result is available,
call await owner.closeRetaining(). This permanently stops command admission,
joins admitted work, and terminates and joins the native child without sending
detachPolicy or closeOwner. It does not inspect the graph or acknowledge
retention, enforcement, release, or namespace survival. An undelivered reconcile
may have left no graph; a lost reply may have left the original or target graph.
Keep the complete original transition durably and recover it through a fresh owner
with identical options. Never fabricate an applied receipt to request detachment.
When cleanup is required instead of recovering enforcement, persist a
IManagedNftTransitionRelease containing the identity returned by the original
start() and the full original transition. After fencing workload packet
admission and joining the old owner, start a fresh owner with the same options:
// Both values come from the caller's durable journal, before the lost apply ACK.
const cleanup = { identity: originalIdentity, transition: originalTransition };
await recoveryOwner.releaseTransition(cleanup);
await recoveryOwner.close();
releaseTransition() validates the original boot and namespace, canonical target,
complete previous result when present, table identity and entire rule graph. It
accepts an absent table, the previous graph, or the target graph; it never creates
or replaces policy. Deleted or changed interfaces do not prevent this cleanup.
Orphan adoption checks the graph before and after claiming OWNER, binds the prior
or already-observed handle, and rejects another live owner. Deletion uses the exact
handle and a generation-fenced batch, followed by an absence check.
Once admitted, only that captured cleanup body may be retried. Other policy writes,
preparation and detachment reject; inspection never claims enforcement. An ordinary
close() retries the selected cleanup before joining the child. closeRetaining()
only joins, allowing a fresh process to retry the durably retained cleanup after
an ambiguous acknowledgement. Local inert-input or busy rejection does not select
cleanup. A successful release acknowledges that table operation at observation
time, including an exact retry after a lost delete ACK. It does not prove packet
drain, clear conntrack, release allocations, or cross a changed boot/namespace.
After closeRetaining(), all policy commands, including detach(), reject.
Concurrent cleanup calls share the same join. A failed join keeps the native
owner and retention choice for retry through closeRetaining() or close().
An ordinary close() admitted first keeps its deletion choice across failures;
a later closeRetaining() rejects instead of reversing that choice.
Detachment confirms retention at observation time. A completed facade retry returns its existing acknowledgement; it is not a fresh kernel inspection. Retention does not guarantee future enforcement, namespace survival or reboot ordering, and does not withdraw external authority, drain packets, clear conntrack or release leases. A separately created owner can later recover the full policy and perform an explicit release after the caller has fenced that authority and packet lifecycle.
Apply and release receipts describe this owner's table and policy operations. They do not certify that packets admitted under an earlier graph have drained: foreign NFQUEUE or other deferred packet owners may still hold such packets. The caller must qualify the complete packet path or independently fence the workload before treating revocation as complete or reusing its addresses and interfaces. Table deletion does not release IP allocation authority or clean up conntrack/NAT state. This restriction also applies to private veth/TUN forwarding.
Refusals: INVALID and EXHAUSTED
Every failure is a ManagedNftablesError with a stable code. Two codes are
refusals of the input itself, raised by the facade's capture or by the native owner
before any kernel work:
INVALID: the input is outside the contract (schema, identifiers, addresses, overlaps, authority, or a bound of the topology's shape). A smaller policy does not help; the input must change.EXHAUSTED: the policy is too big for one atomic replacement. It exceeds the compiled budget described under the router below, or an input count that exists to hold that budget. A smaller policy may fit.
Both carry the native refusal text in reason (bounded, printable ASCII, never
policy content) and in the message, Managed nftables EXHAUSTED: <reason>.
EXHAUSTED also names its bound in details, { bound, limit, actual }:
bound |
limit |
Counts |
|---|---|---|
ruleBytes |
100,000 | encoded rules, chains and sets of the target |
targetBytes |
212,000 | the complete target with its set elements |
operations |
768 | messages of the target |
endpoints |
32 (v1), 128 (router) | private endpoints |
links |
128 | router links |
rules |
128 (v1), 1024 (router) | private directed rules |
grants |
1024 | router egress grants across every generation |
hostGrants |
1024 | host grants of one scope |
workloadGrants |
1024 | router workload grants |
publishedPorts |
1024 | publications of one scope |
localTcpPortOwners |
8 | loopback TCP port owners of the pool guard |
restoreRules |
192 | Docker forwarding contribution rules |
restoreBytes |
100,000 | Docker forwarding contribution command bytes |
actual is always above limit. It is exact for input counts and for the
schema-v1 budget, which measures its complete program. For the other compiled
budgets (ruleBytes, targetBytes, operations, restoreRules, restoreBytes)
compilation stops at the first message or rule past the limit, so the complete
policy needs at least actual. Input counts are checked
before their entries are validated, so a smaller policy can still be INVALID.
Every other bound (identifier and interface-name lengths, source prefixes per
endpoint, addresses per link, protected prefixes, platform endpoints, active
generations, handoffs and allocations, ranges per allocation, guarded pools, the
IPC and capture limits) describes the shape of the topology and stays INVALID.
Other codes, and the facade's own local rejections, carry neither reason nor
details. A native refusal outside this shape is PROTOCOL. A failed close()
rejects CLEANUP_UNCONFIRMED, whose cause is the ManagedNftablesError that left
cleanup unconfirmed, with the native code when the native owner refused.
try {
prepared = await owner.prepare(policy);
} catch (error) {
if (error instanceof ManagedNftablesError && error.code === 'EXHAUSTED' && error.details?.bound === 'publishedPorts') {
// Publish fewer ports and prepare again.
} else throw error;
}
Combined router egress and host transit
ManagedNftables<IManagedNftPolicyV2> accepts the exported schema-v2 policy.
The prepared/applied/transition/status interfaces accept the same policy type
parameter; existing callers default to the unchanged schema-v1 contract. V1
canonical digests and compiled bytes are preserved. V2 uses its own hash domain.
An owner cannot transition between v1 private, v2 router, v2 host transit, and v2
allocation-pool guard policy kinds.
| V2 scope | Required authority and behavior |
|---|---|
routerEgress |
Private endpoints and rules, one exact links binding per endpoint, a separate veth handoff, protection, and active generations. Private veth/TUN/local DNS and egress share one table so terminal private denial cannot override a separate egress table. Optional publishedPorts add the inbound second hop from the handoff to a workload endpoint, one port or a port range; both directions are classified into the default conntrack zone ahead of every leased classifier, so a published endpoint port is dedicated to its publication and never becomes leased egress. A symmetric publication also lets the workload open flows from its published ports. Optional hostGrants forward exact host-origin flows from the handoff to a workload endpoint. Optional workloadGrants let one workload open one exact port of another, one way. |
hostTransit |
Exact handoff link/allocations pairs, complete protection, an explicit veth or Ethernet uplink, and its current snatAddress. It checks each handoff's leased source address and protocol/port range, default conntrack zone, direction, uplink, and protected destinations before outer SNAT. Optional publishedPorts add inbound uplink destination NAT inside the same generation, one port or a port range, and a symmetric publication also carries the workload's own flows from its published ports out through the uplink. Optional hostGrants let the host's own address on a handoff dial exact workload ports. Optional localPlatformEndpoints serve platform endpoints on the host's own addresses to leased flows. Optional exclusiveForwarding makes the table the host's only forwarding owner: every other forwarded packet drops. |
allocationPoolGuard |
An authenticated authorityDigest and complete current allocation-pool prefixes. Installs host-wide IPv4 destination denial before any handoff exists, without link, uplink or SNAT dependencies. Optional hostGrants are its only exceptions. Optional localTcpPortOwners restrict loopback TCP ports to one local user each. |
allocationPoolGuard accepts 1–64 canonical, disjoint RFC1918 prefixes. Supply
the actual allocation pools, not the broader protected union containing management
LANs, resolvers and platform endpoints. It applies to IPv4 INPUT, FORWARD and OUTPUT
at filter priority 0, after normal destination NAT. It drops an original conntrack
destination in a pool. Reply-direction packets additionally check their current
source, then return; all remaining packets check their current destination. This
denies DNAT into or away from a pool and untracked current-destination traffic,
while permitting reverse-SNAT replies to authorized transit sources. It grants no
egress permissions; exact hostTransit and router policies remain separate owners.
Unrelated host traffic and IPv6 retain their existing behavior; the guard never denies
forwarding between other links. On a host that forwards only for the handoffs, see
Exclusive forwarding.
For example, a guard policy is { schemaVersion: 2, revision: 1, scope: { kind: 'allocationPoolGuard', authorityDigest, prefixes: ['10.240.0.0/16', '10.241.0.0/16'] } }. Hashing binds the supplied body, but does not authenticate
the authority or prove its completeness. An empty hostTransit.handoffs array
provides no host-wide denial and is not a substitute for this scope.
The caller must retain the authenticated authority, complete transition and applied
receipt, verify inspect().enforced, and qualify the complete packet path. This
guard alone proves neither early-boot/late-shutdown ordering nor protection after
kernel reboot. It does not fence foreign packet queues, later packet rewrites,
offloads or proxies, and supplies no quarantine release, conntrack drainage or
allocation-reuse evidence. Historical leases remain the caller's durable ownership
responsibility. PERSIST retains this table only while its network namespace exists;
release() and ordinary close() delete it, so revoke any external protection
claim before release, or use the verified retention operation described above.
Local bindings include name, index, kind, MAC (null for L3 TUN), interface-link index, and required IPv4 addresses. RTM_GETLINK/RTM_GETADDR verify those local facts at apply, recovery and inspection. Ethernet must be unbridged driver-backed Ethernet, including VirtIO; virtual VLAN/bond/bridge/dummy kinds are not inferred uplinks. The caller retains actual peer namespaces and link-generation ownership. These serialized facts are not native lifetime capabilities or a continuous link-change monitor. Address, DHCP and route changes require caller fencing.
protection carries an authority digest, non-overlapping protected IPv4 prefixes
and exact platform endpoint IDs/address/protocol/port tuples. The caller must
authenticate and supply complete authority. Hashing does not prove completeness.
A public grant means its declared IPv4 prefix excluding the protected union;
only an explicit platform endpoint grant admits a protected destination. Host
checks cover both the current packet destination and original conntrack destination,
so foreign DNAT cannot turn a protected destination into a public exception or
redirect public traffic into protected space. INPUT diversion, local OUTPUT into
handoffs, unmatched handoff traffic and IPv6 forwarding are denied.
Every handoff address and platform endpoint lies in the protected prefixes. The
hostTransit uplink addresses need not: the authority is the platform's address
space, and a node's uplink lease and snatAddress is usually a public address
outside it. The host barrier protects the uplink addresses itself, adding each one
no protected prefix covers as a /32 after the authority's prefixes, so a
forwarded flow whose current or original destination is the host's uplink address,
or one sourced from it toward a handoff, is denied like any protected destination.
A scope whose uplink the prefixes already cover compiles exactly as before.
Published host ports
hostTransit.publishedPorts is optional. Absent and empty are the same canonical
policy, so an unpublished scope keeps its exact previous digest and compiled bytes.
Each entry is { protocol, hostPort, targetPort, targetAddress, hostIp }, with an
optional hostPortEnd and an optional symmetric. hostIp
must be an exact current uplink address, which rtnetlink verifies at apply, recovery
and inspection; there is no wildcard, secondary-link or default-route inference.
targetAddress must be a current leased transitSourceAddress of one allocation,
and that binding is what selects the handoff link of the admission; an address that
no active lease holds, a local host address or a bare workload prefix rejects. Ports
are 1–65535, one protocol/hostPort is published exactly once across the scope
including across uplink addresses, and a port published on the snatAddress may not
fall inside any leased outbound source-port range of the same protocol, because outer
SNAT translates to that same address.
hostPortEnd publishes the range hostPort..hostPortEnd. It must be greater than
hostPort, and targetPort must equal hostPort: a range keeps every port, so port
p of the uplink reaches port p of targetAddress. One port has exactly one
canonical form, without hostPortEnd. Ranges follow the same rules as single ports:
no two publications of one protocol overlap anywhere in the scope, and no published
port on the snatAddress falls inside a leased range. A range translates the
address only, which never changes a port.
Publications are set elements, not rules. The generation compiles CONSTANT named
sets and maps of concatenated fields in which the port fields are inclusive
intervals (the kernel's pipapo backend), so one element holds one port or a whole
range: published_port maps a one-port publication's uplink address, protocols and
port to its target address and port, published_address maps a range to its target
address alone, and published_forward holds every publication's inbound tuple. A
constant number of rules looks them up whatever the number of publications: one
destination NAT rule per map in this table's own prerouting nat chain at priority
-100, and four forward admissions ahead of the capture barrier — the ESTABLISHED
original direction, a NEW opening per protocol with opening TCP restricted to SYN
with FIN/RST/ACK clear, and the ESTABLISHED reply. Every rule checks the uplink, the
default zone and the direction, and every key carries the exact handoff (index and
name), the packet and conntrack protocols, the translated current address and port
and the original conntrack destination, so foreign NAT cannot borrow the exception
for another flow. Published traffic is not source-translated, so the workload
observes the actual client address. The chain, its sets, its translations and its
admissions are created, replaced and deleted with the generation: a target without
the entry, a failed batch, or release removes them, and a caller never needs a
separate teardown. Owner verification reads every set back with its description,
every element with its interval bounds and map data, and every lookup with its
registers, so apply, inspection, adoption and recovery reject a changed element
exactly as they reject a changed rule.
Without symmetric, a publication carries inbound flows only: a flow the workload
opens from its target address and port toward the uplink is denied. With symmetric: true, a flow from exactly targetAddress and the published target port(s) on the
exact handoff may leave through the uplink, NEW (a TCP opening only with SYN) or
ESTABLISHED, with its ESTABLISHED replies, and outer source NAT translates it to
hostIp and the published uplink port. These flows are elements of one more map,
published_symmetric, keyed on the handoff, protocols and the live and original
source and holding the uplink address and port span; four admissions look it up as a
set and one postrouting rule translates through it. The workload's outbound flows therefore leave
from the same address and port its clients reach, as Docker host-mode NAT kept them;
a range keeps each port unless the host itself already holds that exact tuple on
hostIp. The admissions follow the capture barrier, so every protected current or
original destination, the platform endpoints included, stays denied. A symmetric
target port may not fall inside a leased range of targetAddress, and may not overlap
the target ports of another publication of the same address and protocol. Absent and
false are the same canonical policy, digest and compiled graph.
Input bounds are 1024 entries. The rule graph does not grow with them: with every
kind of publication present the host compiles at most eleven rules and four sets,
about 13,400 bytes of the 100,000-byte rule reserve, and each publication costs
about 150 to 300 bytes of set elements, which share the target budget described
under host grants. 64 symmetric ranges take about 19,600 element bytes.
The caller still owns the route to targetAddress through that handoff, the
workload listener, and every other packet owner on the host. As with leased egress,
an ACCEPT here cannot override Docker's independent FORWARD DROP; the
ManagedDockerForwarding contribution below admits every inbound and symmetric
publication of the barrier through it, with the leased flows.
Each router generation binds an immutable lease reference, transit source address, leased TCP/UDP source-port ranges, a nonzero conntrack zone and a 16-byte label. Each directed grant selects one exact leased protocol range for SNAT. A grant's source is a workload veth or the actual router-local host with an exact assigned IPv4 source address; TUN endpoints retain private routing only. Router-local DNS and relay transports therefore need explicit local-origin grants. Raw PREROUTING and OUTPUT classify before conntrack; filter rules validate current and original tuples, links, direction, zone, state and generation label on every packet. Opening TCP requires SYN with FIN/RST/ACK clear and permits ECN negotiation. Replies must match the labelled original flow and current directed grant. There is no broad ESTABLISHED/RELATED bypass.
Published workload ports
routerEgress.publishedPorts is optional and is the second hop of a published
host port. Absent and empty are the same canonical policy, so an unpublished scope
keeps its exact previous digest and compiled bytes. Each entry is { protocol, transitPort, endpointPort, endpointAddress, transitSourceAddress }, with an
optional transitPortEnd and an optional symmetric.
transitSourceAddress must be a current leased generation address that is present
on handoff, which rtnetlink verifies at apply, recovery and inspection; the
handoff can carry several current leased addresses, so the arrival address is
always explicit and never inferred. transitPort is the targetPort of the
matching hostTransit publication. endpointAddress must be an address of one
current veth workload endpoint of this scope: a router-local link or gateway
address, the transit address, a TUN endpoint's private space, a platform endpoint,
loopback and IPv6 all reject. Ports are 1–65535, one protocol/transitPort is
published exactly once across the scope, and a published port may not fall inside
any leased outbound source-port range of the same protocol on that transit address,
because source NAT translates to that same address.
transitPortEnd publishes the range transitPort..transitPortEnd, the second hop of
a host range. It must be greater than transitPort, and endpointPort must equal
transitPort, so every port keeps its number on both hops; one port has exactly one
canonical form, without transitPortEnd. No two publications of one protocol overlap
anywhere in the scope, and no published port falls inside a leased range of its
transit address. The destination NAT of a range translates the address only.
Publications are set elements here too: published_arrival (protocol, transit
address and port), published_endpoint (workload link, protocol, address and port,
the union of every publication's workload ports as disjoint canonical intervals),
the published_port and published_address translation maps and
published_forward, with the port fields as inclusive intervals. A constant number
of rules looks them up whatever the number of publications: one destination NAT
rule per map in this table's own prerouting nat chain at priority -100, two raw
classifications, and four forward admissions placed ahead of the capture barrier —
the ESTABLISHED original direction, a NEW opening per protocol with opening TCP
restricted to SYN with FIN/RST/ACK clear, and the ESTABLISHED reply. The handoff
ingress classification precedes the
barrier's protected-source denial, because the uplink or management network that
reaches a published port can itself be protected space (an uplink address outside
the protected prefixes is still guarded as a /32). The endpoint ingress
classification matches the exact workload link, address and published endpoint
port, and precedes every leased classifier of every generation: raw runs before
conntrack, so a client whose source port equals a leased grant's destination port
would otherwise carry the workload's published packets into that generation, where
the publication's flow does not exist. Both classifications match the arriving
tuple only, set no generation zone and leave the flow in the default conntrack
zone. The private anti-spoof guards still run first, and every admission checks
the handoff, the default zone and the direction, and looks up the exact workload
link (index and name), both protocols, the translated current address/port and the
original conntrack destination.
A published endpoint port is therefore dedicated to its publication: a new
outbound flow that the workload opens from that exact address and port on that
exact link stays in the default zone too and is never translated by the leased
outbound source NAT. Without symmetric such a flow is denied. Inbound published
traffic is not source-translated on either hop, so the workload observes the actual
client address. The chain, its sets, its translations, its classifications and its
admissions are created, replaced and deleted with the generation: a target without
the entry, a failed batch, or release removes them.
With symmetric: true the workload may open flows from exactly endpointAddress
and the published endpoint port(s) on its exact link toward the handoff, NEW (a TCP
opening only with SYN) or ESTABLISHED, with their ESTABLISHED replies, and source
NAT translates them to transitSourceAddress and the published transit port. With
the matching symmetric hostTransit publication the flow reaches the uplink from
hostIp and the published uplink port, the same address and port its clients
reach, as Docker host-mode NAT kept it; a range keeps each port unless that exact
tuple is already taken on the translated address. These flows are elements of the
published_symmetric map, keyed on the workload link, protocols and the live and
original source and holding the transit address and port span, behind four
admissions and one postrouting translation. Both directions are already in the
default zone through the publication's endpoint and handoff classifications, so no
leased classifier can capture them. The admissions follow the capture barrier, so
every protected current or original destination, the platform endpoints included,
stays denied; outbound flows anywhere else still need their leased grants. The
endpoint ports of a symmetric publication may not overlap the endpoint ports of
another publication of the same address and protocol, since each flow has exactly
one published source. Absent and false are the same canonical policy, digest and
compiled graph.
Input bounds are 1024 entries. The rule graph does not grow with them: with every
kind of publication present the router compiles at most thirteen rules and six sets,
about 15,100 bytes of the 100,000-byte rule reserve, and each publication costs
about 250 to 450 bytes of set elements, which share the target budget with the
router's endpoints, rules and grants. 64 symmetric ranges take about 28,000 element
bytes. The caller still owns both hops: the exact matching hostTransit
publication, the route to the workload through this router, the workload listener
and every other packet owner.
Host grants
A host grant lets the node's own host network namespace dial one exact port of
one local workload at the workload's address, from the host's own address on the
current handoff (the transit host address). It is how a controller that runs on
the node, such as a reverse proxy, reaches a workload without loopback
publication, route_localnet, loopback DNAT or loosened martian filtering. Each
entry is { protocol, sourceAddress, destinationAddress, destinationPort }:
tcp or udp, two exact unicast IPv4 addresses and one port 1–65535. Prefixes,
wildcards, ranges, lists, loopback, broadcast, IPv6 and duplicate tuples reject.
There is no source-port selection and no translation.
The flow crosses three tables, so the caller compiles the same grants into each scope, and each scope binds them to the authority it proves:
allocationPoolGuard.hostGrants:destinationAddresslies in a guarded pool. The admissions sit in OUTPUT (original direction) and INPUT (reply) ahead of the guard. Everything else still meets the unconditional guard, including forwarded traffic and a translated flow that merely arrives at a granted tuple.hostTransit.hostGrants:sourceAddressis an exact current address of exactly one handoff link and not an uplink address;destinationAddressis protected space that is neither host-local, a platform endpoint nor a leased transit source address. The admissions bind the exact handoff that holds the source address and precede its terminal OUTPUT and INPUT denials; forwarding and the uplink are unchanged.routerEgress.hostGrants:destinationAddressis an exact current veth workload endpoint address of this scope, andsourceAddressis protected space that is neither router-local, a platform endpoint nor any workload's address. One raw handoff ingress classification precedes the barrier's protected-source denial, sets no generation zone and leaves the flow in the default conntrack zone; the forward admissions bind the handoff and the exact workload link, follow the private anti-spoof guards, and sit beside the publications ahead of the capture barrier. The workload's replies need no classifier: their destination is protected, so they return to the default zone ahead of every leased classifier, even when the host's source port equals a leased grant's destination port.
The grants of a scope live in one CONSTANT named set, host_grant, and the rule
graph does not grow with them: four admissions per scope — ESTABLISHED original
direction, a NEW UDP opening, a NEW TCP opening restricted to SYN with FIN/RST/ACK
clear, and the ESTABLISHED reply — each check the default conntrack zone and the
direction, then look up one concatenated key. The key carries every remaining
operand: the variable interface's index and name (the handoff in host transit, the
workload link in the router; the guard has none), the packet protocol and the
conntrack protocol, the live source, destination and workload port, and the
originally tracked source, destination and port. An element is one grant whose
live tuple equals its original tuple, so a translated flow never matches and
foreign NAT cannot borrow an admission for another flow. The router's raw
classification looks up the exact arriving protocol, addresses and port in a
second set, host_grant_arrival. Owner verification reads every set back with all
of its elements, so apply, inspection, adoption and recovery reject a changed or
foreign set exactly as they reject a changed rule; a bound CONSTANT set also
refuses element changes from any other socket.
Absent and empty are the same canonical policy, so a scope without grants keeps its exact previous digest and compiled bytes and has no host grant set. Withdrawing a grant is an ordinary atomic replacement: the batch deletes the previous rules, chains and sets by handle and creates the complete target, so a target without the entry, a failed batch or release removes it, and an established flow stops with it. Apply grants hop by hop from the workload side (router, host transit, guard) and withdraw them in the reverse order; each table is atomic on its own, and a partially applied grant only stays dark.
Host-local platform endpoints
A platform endpoint normally lives elsewhere and host transit forwards leased
flows to it. When the host serves an endpoint itself — a relay listener, a data
port or a VPN hub on an address the host holds — the leased flow ends in the
host's INPUT instead, where every handoff arrival is denied. The optional
hostTransit.localPlatformEndpoints lists the ids of protection.platformEndpoints
the host serves:
protection: { authorityDigest, prefixes: ['10.0.0.0/8', '192.0.2.0/24'], platformEndpoints: [
{ id: 'relay', address: '192.0.2.2', protocol: 'tcp', port: 8443 }] },
localPlatformEndpoints: ['relay'],
Each id names an existing endpoint whose address is an exact requiredIpv4Addresses
entry of the uplink or of a handoff link, so rtnetlink verifies that the host holds it
at apply, recovery and inspection. The refusal of a platform endpoint on a host
address is lifted only for declared ids; unknown, duplicate or unheld ids reject.
The router scope grants its workloads the endpoint like any platform endpoint
(destination: { kind: 'platform', endpointId }).
INPUT admits the original direction only from the exact handoff that carries a lease, with the lease's transit source address and a source port in one of its ranges for the endpoint's protocol, in the default conntrack zone, to the exact endpoint address and port in both the live and the originally tracked tuple, NEW (a TCP opening only with SYN) or ESTABLISHED. OUTPUT admits only the ESTABLISHED reply back through that handoff. A flow translated in transit, another source address or port, another endpoint port or protocol and any host-origin opening toward the transit address keep meeting the handoff denials. The source port is checked against the leased range in both tuples, not for equality between them. Each leased range adds one INPUT and one OUTPUT jump when an endpoint of its protocol is declared, and each endpoint adds three rules. Absent and empty are the same canonical policy, digest and compiled bytes. The caller keeps a declared endpoint's address outside every guarded allocation pool and owns the listener.
Exclusive forwarding
A host that forwards IPv4 (net.ipv4.conf.all.forwarding=1) routes between all of its
links. On a Docker host, Docker set that switch and also the iptables FORWARD policy
DROP, and ManagedDockerForwarding admits the handoff flows through it. On a host
without Docker nothing denies forwarding that does not touch a handoff: host transit
governs only handoff traffic, and the allocation-pool guard only pool destinations,
so a LAN peer could use the host as its gateway, the uplink could hairpin, and two
other links could exchange traffic. hostTransit.exclusiveForwarding: true closes
that path inside the barrier's own table:
const transit: IManagedNftPolicyV2 = { schemaVersion: 2, revision, scope: {
kind: 'hostTransit', protection, handoffs, uplink, snatAddress, exclusiveForwarding: true } };
The forward chain ends in one unconditional drop, after every rule of the scope.
The host then forwards exactly what the scope admits: the leased flows of each
handoff to the uplink with their ESTABLISHED replies, the inbound publications and
their replies, and the symmetric publications' flows. Everything else that reaches
the forward hook drops, whatever its links: IPv4 between the uplink, a LAN or any
other link, a hairpin through the uplink, flows that were established before the
member was applied, and forwarded IPv6. There is no ESTABLISHED/RELATED bypass, as
everywhere in this scope, so ICMP errors for leased flows stay denied as before.
Allocation pools need no exception: pool addresses are routed inside the router
namespace and leave it translated to the transit source, so on the host only
transit addresses cross the forward hook, and the pool guard denies pool
destinations there anyway. IPv6 is dropped because the scope's packet model is
IPv4 only; a host whose IPv6 forwarding is off never presents IPv6 to the hook.
Input, output and NAT are unchanged, and so are the host's own flows: host grants
and host-local platform endpoints are INPUT and OUTPUT traffic. With no handoffs
the scope forwards nothing at all.
nftables runs the forward base chain of every table, and a drop in any one of them
is final, while an accept only ends its own chain. The member therefore drops
forwarding that another owner (Docker, a container or VM bridge, a VPN router, a
second routing daemon) accepts, and accepting in this table still cannot override
another owner's drop. Use it only on a host where this table is the sole forwarding
owner; on a Docker host keep it absent and use the Docker contribution below, which
refuses an exclusive barrier. inspectForwardHooks(), described below, reports
the other forward owners of the namespace for that decision. The
member does not fence flowtable offload, packet queues, proxies or anything outside
the forward hook.
The drop is part of the scope's graph: it is created, verified, adopted, recovered,
replaced and released with the rest of the table in one batch, and a replacement
without the member or release() removes it, so forwarding resumes for every link
the moment the table goes. This package sets no sysctl. The denial holds only while
an exclusive table is applied, so the caller enables host forwarding only after an
exclusive scope is enforced and disables it before releasing that scope, or keeps
an exclusive scope applied between generations (with no handoffs it is a pure
host-wide forwarding denial) and replaces it in place. Persistent forwarding in
sysctl.d would make the host forward unfiltered after every boot until the scope
is applied again, so leave the boot default off. Absent and false are the same
canonical policy, digest and compiled bytes; the member adds one rule.
Workload grants
routerEgress.workloadGrants is optional: a stateful one-way flow between two
workloads behind the same router, such as an ingress workload reaching a target
workload's service port across organizations. Each entry is { protocol, sourceAddress, destinationAddress, destinationPort }: tcp or udp, the exact
current addresses of two different veth workload endpoints of the scope and one port
1–65535. Prefixes, wildcards, ranges, router-local and platform-endpoint addresses,
TUN endpoints, both ends on one endpoint and duplicates reject. The source opens and
the destination only answers; the reverse direction is a second grant.
Four FORWARD admissions follow the private anti-spoof guards whatever the number of
grants: ESTABLISHED original direction, a NEW UDP opening, a NEW TCP opening
restricted to SYN with FIN/RST/ACK clear, and the ESTABLISHED reply, each in the
default conntrack zone. Every admission looks up the incoming link with the source
address and the outgoing link with the destination address (index and name) in the
CONSTANT set workload_link, and the protocols with the live and originally tracked
tuple in the CONSTANT set workload_grant, so a spoofed, renamed or translated flow
never matches and the destination can never open toward the source. Endpoint prefixes
are protected, so both directions stay in the default zone ahead of every leased
classifier. Owner verification reads both sets back element by element. Up to 1024
grants fit; the sets share the scope's target budget with hostGrants and every
other set, and 1024 of each fit together beside the reference router; a scope
beyond the budget is refused before any kernel work. Absent and empty are the same canonical policy, digest and compiled bytes, and
withdrawal is an ordinary atomic replacement that also stops established flows.
These flows need no private rules. A private rule is stateless: answering through
private rules needs a reverse rule, which would also let the destination open toward
the source.
Loopback TCP port owners
allocationPoolGuard.localTcpPortOwners is optional and restricts a loopback TCP
service to one local user, such as a container runtime's CRI stream server on
127.0.0.1:10010 that only root may reach. Each entry is { address, port, uid }:
one exact IPv4 loopback host address in 127.0.0.0/8 (not the network or broadcast
address), one port 1–65535 and one uid 0–4294967294. Prefixes, wildcards, IPv6,
ranges, uid lists and a second entry for the same address and port reject; the
limit is 8 entries. Absent and empty are the same canonical policy, so a guard
without owners keeps its exact previous digest and compiled bytes.
const guard: IManagedNftPolicyV2 = { schemaVersion: 2, revision: 1, scope: {
kind: 'allocationPoolGuard', authorityDigest, prefixes: ['10.240.0.0/16'],
localTcpPortOwners: [{ address: '127.0.0.1', port: 10010, uid: 0 }] } };
Each entry compiles two rules at the head of the guard's base chains, ahead of the host grant admissions and the pool guard, so nothing earlier in the table accepts past them:
- OUTPUT (filter priority 0, after ordinary output destination NAT): IPv4 TCP to the
exact address and port with
meta skuid != uidis rejected with a TCP reset, so another user'sconnect()fails at once withECONNREFUSED. This covers every packet, not only the opening, and a flow a foreign output DNAT redirects to the port.meta skuidis the socket file's file-system uid as the user namespace owning the network namespace sees it, the initial one on a host: a process in a user namespace that shares the host network namespace is matched by its host uid, and a user namespace's root is not uid 0. A packet without a socket file carries no uid and does not match: a kernel reply, the teardown of an orphaned socket, or an in-kernel socket such as an NFS or CIFS client, which only a privileged mount creates. - INPUT: the exact address and port arriving on any interface but loopback (ifindex
1 in every network namespace) are dropped, so
route_localnetor a prerouting DNAT cannot expose the service to another host.
The rules match the exact destination address. The service must bind exactly that
address; a wildcard listener stays reachable through the host's other addresses.
IPv6 is not covered, so the service must not listen on ::1 or ::. A socket keeps
the uid it was created with, so a root process that hands its connected socket to
another user delegates that connection. The reject expression needs the kernel's
nft_reject_inet module (autoloaded on stock kernels). Owner verification reads
back every rule, including the comparison operator, the uid and the reject type, so
apply, inspection, adoption and recovery reject a changed rule. The entries are
applied, retained, recovered and released with the rest of the guard's table; after
a reboot the caller applies its retained intent again, like every other member.
Input bounds are 1024 grants per scope, the maximum a network projection carries. Elements are sent 256 to a message. At 1024 grants the reference router needs about 98 KB of elements (two sets), host transit about 66 KB and the guard about 45 KB.
Every replacement is one nfnetlink batch in one sendmsg, and Linux refuses a
netlink message larger than the socket send buffer: the requested 1 MiB SO_SNDBUF
is clamped to net.core.wmem_max and doubled, 425,984 bytes with the Linux default
of 212,992. That is the hard ceiling. This package bounds a batch at 320,000 bytes,
which fits every wmem_max of at least 160,016, and splits it: at most 64 bytes of
batch framing, at most 107,520 bytes for the handle-only deletion of the previous
graph (fewer than 768 rules, chains and sets of at most 140 bytes each), and
212,000 bytes for the complete target, of which rules, chains and sets may take at
most 100,000 and set elements the rest. Scopes that grow with their inputs keep
those inputs in set elements, so the element share grows where the rule share
shrinks; prepare() refuses a target beyond either bound with EXHAUSTED before any kernel work. The private facade IPC carries up to three complete policies in one
status and is bounded at 1,048,576 bytes per line; the facade captures at most
1,000,000 JSON bytes and 65,536 values per request or result. The caller owns the host route
to the workload through the handoff, selecting the transit host address as the
socket's source, the workload listener and every other packet owner.
Active overlapping classifiers, duplicate zones/labels and conflicting handoff allocations reject. Empty router generations retain private routing while denying egress. Router input bounds are 128 private endpoints and links, 1024 private rules, 32 active generations and 1024 total egress grants; other bounds include 128 protected prefixes, 96 platform endpoints, 16 ranges per allocation, and 32 host handoffs/active allocations. The complete target must fit the budget above; the router's rules grow only with its protected prefixes, so its elements bind first. Exhausted source-port ranges drop new flows rather than allocate outside the lease.
The router compiles its private endpoints, directed rules and leased grants into named sets and maps behind a fixed set of rules, whatever the number of workloads:
private_linkholds every endpoint link (index and name) and the absent interface of router-local traffic;private_indexandprivate_namehold every endpoint index and name for the terminal denials, either alone enough to deny;private_sourceholds each link index with its source prefixes. The anti-spoof guard drops a packet on an endpoint link whose address is not one of its prefixes, and anything but IPv4 there.private_anyandprivate_porthold the directed rules: incoming and outgoing link index (0 for the router itself), addresses, and for a transport rule the protocol and ports.- Per kind (platform ahead of the protected barrier, public behind it) and per
shape (whether the grant names a source port),
leased_*_zonemaps a workload or router flow's link, protocol and tuple to its generation zone before conntrack,leased_*_returnmaps a translated return to its zone, andleased_*_flowmaps the link, both protocols, the live and originally tracked tuple and the zone to the grant's source NAT; the filter admissions look it up as a set.leased_labelholds each generation's zone and label, andleased_generationmaps a zone to the label a fresh flow receives. A key's link index names its link because the rule also finds the interface's index and name inprivate_link, whose indices are unique. Every element is checked, like every rule, on apply, inspection, adoption and recovery. A decision corpus of about 133,000 packets across five router fixtures, frozen from the per-rule compiler of 2.6.0, is reproduced exactly by the set-backed compiler.
A router of 64 workloads, each with local DNS, public TCP and UDP egress, grants to four platform endpoints and one to three publications, compiles about 66,000 rule bytes and 105,000 element bytes; 89 such workloads fit the target budget:
| Workloads | Router rules | Router elements | Host rules | Host elements |
|---|---|---|---|---|
| 1 | 65,820 | 4,968 | 42,944 | 1,128 |
| 8 | 65,820 | 16,056 | 42,944 | 2,528 |
| 32 | 65,820 | 54,072 | 42,944 | 7,328 |
| 64 | 65,820 | 104,760 | 42,944 | 13,728 |
Upgrading from 2.x: version 3 changes the compiled identity of every schema-v2
routerEgress table and of every scope with publishedPorts, while their policy
digests stay unchanged. Reconcile or release against such a table applied by 2.x
returns Conflict: the new engine cannot adopt the old graph. Recover and release
each such table with the exact previous engine and its retained intent and
receipt, under the same traffic fence as below, before upgrading; tables of other
scopes and without publications are unchanged.
Upgrading from 1.x: version 2 changes the compiled identity of schema-v2
routerEgress tables. Retained router tables and their receipts cannot be adopted
by the new engine, even when the policy digest is unchanged. Fence workload and
router-local traffic, then use the exact previous engine and retained intent/receipt
to recover and release each router table. Join that owner before upgrading.
Prepare and apply the policy with the new engine and persist its complete receipt
before lifting the traffic fence. Keep lease quarantine and host/pool protection
throughout; this migration provides no proof of conntrack drainage or safe allocation reuse.
Do not discard an unresolved retained table or its old engine/receipt. Schema-v1,
host-transit and allocation-pool-guard compiled identities are unchanged.
The caller must coordinate router and host apply order, retain complete intents
and receipts, authenticate protected authority, prevent reuse of quarantined
leases, and qualify other packet owners. An ACCEPT in this table cannot override
Docker's independent FORWARD DROP. ManagedDockerForwarding, described below,
owns the separate DOCKER-USER contribution. The caller must also control deferred packets, proxy/BPF/
offload paths and changing network authority. No table receipt proves flow
drainage, conntrack cleanup, elapsed-time expiry, or safe address/port/zone reuse.
Docker forwarding contribution
ManagedDockerForwarding admits exactly the forwarded flows of an exact applied
schema-v2 hostTransit receipt through Docker's existing IPv4 forwarding path.
It reads and verifies that dedicated barrier's complete table and local links,
without adopting its ownership. The contribution uses Docker's supported
DOCKER-USER extension point.
Docker's native nftables backend has no equivalent user chain and is unsupported.
The host must provide root-owned /usr/sbin/iptables-nft,
/usr/sbin/iptables-nft-save and /usr/sbin/iptables-nft-restore, using the same
1.8.10-or-newer 1.8-series nf_tables frontend. prepare() captures the exact
frontend version in its digest. Apply, inspection and recovery require that same
version. Subprocesses use fixed argv, a clean environment, bounded input/output,
one shared eight-second frontend deadline per request, and joined termination.
No shell or raw rule API is exposed.
One node-level contribution owns a contiguous leading block in DOCKER-USER.
FORWARD must already have policy DROP and its first two unconditional jumps must
be DOCKER-USER and DOCKER-FORWARD. Their exact ordered raw graphs are checked
without claiming their ownership. Other managed contribution owners, altered or
duplicated owned rules, missing chains and changed placement reject. Unrelated
rules following the owned block and Docker's own chains remain separate owners.
Saved text verifies placement, count and normalized policy. Native netlink reads
verify every contributed rule's complete ordered expression graph, including
register flow and all match bytes; only counter values and attribute encoding
order/flags are normalized. An exact xtables -C check also verifies each compiled
deletion specification. These reads must share an unchanged ruleset generation;
concurrent Docker reconciliation can therefore make inspection inconclusive.
Conntrack source addresses use the canonical bare IPv4 form because
an explicit /32 has different hidden match bytes in these frontends. A
semantically equal rewrite with a different exact representation rejects, as does
an early ACCEPT that the frontend reconstructs as the same command.
The contribution never creates, flushes, adopts or deletes a Docker table/chain,
and never changes a host forwarding policy. prepare() refuses a barrier with
exclusiveForwarding as INVALID: its drop would also deny Docker's own container
forwarding.
The contribution and the barrier's forward chain derive from one enumeration of the scope's forwarded flows, so the contribution admits exactly the flow kinds the non-exclusive barrier admits, each with its replies:
| Flow kind | Original direction | Live tuple | Conntrack original tuple |
|---|---|---|---|
| Leased range | handoff → uplink | transit source address and leased port range | same source |
| Inbound publication | uplink → handoff, after the barrier's DNAT | inside target address and port(s) as destination | published hostIp and port(s) as destination |
| Symmetric publication | handoff → uplink, before the barrier's SNAT | target address and port(s) as source | same source |
Every flow kind has three rules: the NEW opening (TCP only with SYN and
FIN/RST/ACK clear), the ESTABLISHED original direction and the ESTABLISHED reply
with swapped interfaces and tuple sides. A router's own egress, such as a VPN client
in the router namespace dialling its hub, reaches the host as a leased flow: the
router translates it to the transit address and a leased port range, so it is
admitted by the leased rules exactly when the barrier admits it. Host grants and
host-local platform endpoints use INPUT/OUTPUT and never cross FORWARD. Each rule
binds the handoff and uplink names. The exact v2 host barrier independently checks
interface indices, MAC/iflink facts, the default host conntrack zone, protected
original/current destinations and explicit outer SNAT; it governs every forwarded
packet on its handoffs, so the host forwards the intersection, which is the
barrier's admitted set. A scope without publications keeps its exact 4.1.0 rules,
comments and graphs. A complete contribution is bounded to 192 rules and 100,000
generated command bytes: three rules per leased range and per publication, and
three more per symmetric publication. A larger scope is refused as EXHAUSTED
(restoreRules/restoreBytes) before any subprocess.
Create the class with ownerId, a caller-retained instanceId of at least 16
identifier characters, and optional binaryPath/networkNamespaceFd. After
start(), call prepare({ schemaVersion: 1, revision, barrier }), where barrier
is the exact applied host policy. Persist the complete { previous, target }
transition before reconcile(). The native owner fences the handoff links before
and after its single no-flush restore transaction, which deletes the previous block
and inserts the target block in one atomic nf_tables commit. A handoff link that
previous and target both bind exactly may stay up, so a publication or lease
amendment of a running generation applies in place: every flow both contributions
admit passes throughout, a flow only the target admits passes from the commit, and
the referenced barrier must already be the target. Every handoff link the
transition adds or drops, and every link of a first apply, must be down; retain
those exact links until cleanup finishes. Exact target replay can recover a lost
response without rewriting live rules.
inspect().present confirms the contribution and its referenced barrier, not
the complete host packet path or durable controller authority. Errors retain
failed-owned state and pending intent. release(applied) and close() delete
only exact contributed rules. close() first reads the filter table through the
same generation-consistent frontend and netlink read, without requiring Docker's
forwarding path: when no saved rule in any chain and no native DOCKER-USER or
FORWARD rule carries this owner's comment prefix, cleanup is confirmed. So an owner
whose first apply was refused, for example on a host with FORWARD policy ACCEPT or
without Docker, closes on that host. A read that fails rejects
CLEANUP_UNCONFIRMED whose cause carries the native code (UNAVAILABLE,
UNSUPPORTED_BACKEND, PERMISSION, CONFLICT); it is never reported as cleanup.
The fence of release(applied) and close() requires every handoff link down or gone:
a link whose index no link holds any more, such as a veth that left with its router
namespace when the node process ended, counts as down, so a restarted process in
the same boot can release the previous contribution before applying the next one.
A different link at a bound index still rejects. A cold boot needs a new
barrier receipt and caller-authorized replay; old boot receipts cannot authorize
native operations. The caller owns restart ordering, disjoint retained leases,
exclusive privileged mutation authority and activation. The frontend transaction
does not provide compare-and-swap against other privileged processes.
Forward hook owners
inspectForwardHooks({ binaryPath?, networkNamespaceFd? }) reads every
packet-filter owner on the routed forward path of one network namespace and
changes nothing. One native process takes the read and exits:
chains: every nf_tables base chain on the forward hook of theip,ip6andinetfamilies, withfamily,table,tableHandle,chain,type,priorityandpolicy, sorted by family, table and chain. Managed tables are included; comparetableandtableHandlewith a receipt'sidentity.tableNameandtableHandleto find your own.docker: whether Docker's iptables-nftDOCKER-USERandDOCKER-FORWARDchains exist in tableip filter.legacyTables: the registered legacy xtablesipv4andipv6tables. Their hooks, including FORWARD, are not nf_tables objects and never appear inchains.
The table and chain dumps share one ruleset generation, or the read is
CONFLICT. Without CAP_NET_ADMIN in the namespace it is PERMISSION. Bridge
and netdev hooks do not see routed packets and are not reported. Guard
exclusiveForwarding with it: refuse while any chain other than your own barrier's,
either Docker chain or any legacy table is present. The read is a point in time;
the caller still owns exclusive authority over who may add a forward owner later.
Native qualification
The separate Docker fixture exercises the native contribution with pinned Docker packages on offline Ubuntu 26.04 and Ubuntu 24.04.4/HWE guests, including retained-state cold boots and published workloads. Its checked-in input manifest and per-run receipts identify the tested artifacts.
cargo test --manifest-path rust/Cargo.toml --locked runs unprivileged compiler,
snapshot and bounded-subprocess tests. The ignored native cases require the explicitly marked disposable guest
from test/native/qualify.py; requesting them on an ordinary host fails its scope
check. Qualification uses an offline Linux 6.18.35 x86_64 guest with no host disks,
mounts or external network backend. One VirtIO device connects only to a singleton
QEMU-internal hub; packet-path tests use guest-owned namespaces, veth pairs and TUN.
It covers
UDP/TCP grants, direct-IP denial, source spoofing, IPv6 denial, renamed interfaces,
foreign ownership, atomic rollback, generation conflicts and lost-ACK recovery.
V2 tests exercise complete graph persistence/recovery, private and local DNS,
local TCP, TUN isolation, observed UDP/TCP handoff ranges, range exhaustion,
fragment reassembly, ECN SYN and invalid opening flags, protected current/original
DNAT barriers, nonzero foreign zones, local diversion, IPv6 denial, revocation
and restoration with disjoint generation ranges while old flows remain retained.
Published host ports are qualified end to end: translated TCP and UDP delivery to the
leased transit address with the actual client address preserved, reverse translation
back to the published address, SYN-only opening, denial of non-SYN openings, denial of
unpublished workload and uplink ports against a pre-policy positive control, and the
port going dark with the withdrawn and released generation. Published workload ports
are qualified across both hops in the same four-namespace topology: TCP and UDP reach
a listener inside the workload namespace with the client address preserved and both
translations reversed, only an exact SYN opens the publication against a positive
control on the same injector, an unpublished transit port and the unpublished
workload path stay behind the router barrier, and withdrawing only the router
publication or releasing it takes the port dark while the host publication remains.
A client whose source port equals a leased grant's destination port on that grant's
public address completes the published TCP handshake and receives its UDP reply
from the published address, so the leased classifiers of the same generation cannot
capture the workload's published packets.
Published ranges and symmetric publications are qualified across both hops in the
same topology, with every publication a set element whose kernel dump matches the
compiled sets and maps: the first, middle and last port of a UDP range reach their listeners
with the client address preserved and answer from the published address, the ports
just outside the range stay dark against pre-policy positive controls, symmetric UDP
from a range port and from the signalling port and symmetric TCP from the signalling
port reach the uplink peer from the published address with the same port, a later
request from that peer reaches the workload on the same flow, the plain publication's
workload port opens nothing, symmetric flows into the protected union stay denied
against a pre-policy positive control, leased egress keeps its leased range, and the
range goes dark with release.
A 64-workload router is qualified in one batch, every workload in its own namespace
with local DNS, public TCP and UDP egress, four platform endpoints and a
publication, the first with the SIP shape: the kernel dumps every private, leased
and published element and the owner verifies it, and on the first, middle and last
workload local DNS answers, public UDP egress leaves from the leased range, UDP and
TCP platform endpoints answer through both hops, the publication reaches its
listener with the client address preserved, and another workload, a protected
address outside the platform endpoints and an ungranted platform port stay denied
against pre-policy positive controls, as does another workload's source address.
The router compiles 82 rules for the 64 workloads.
Host grants are qualified across all three tables in the same four-namespace
topology with leased public egress on every port of both protocols and 1024 grants
in every scope: a member far from the probed tuples passes, the tuple just outside
the set stays dark, and against
pre-policy positive controls, the host stays dark without grants and with router and
transit grants alone (the guard still denies), the exact TCP and UDP tuples reach the
workload from the transit host address once the guard exception is applied, another
port, protocol, host source address and the workload-origin direction stay denied,
withdrawing only the router grant closes the path and re-granting reopens it, and
withdrawal in every scope stops new flows and the established one. A full guard
set persists across owner loss, is re-verified element by element before adoption,
survives lost-ACK replay and is replaced and released like the rest of the graph,
and another socket cannot add or delete one of its elements.
Host-local platform endpoints are qualified from a router namespace against
pre-policy positive controls: without the member the handoff denials keep both a TCP
endpoint on the uplink address and a UDP endpoint on the handoff address dark; with
it the exact leased flows reach them with the transit source preserved, while a
source port outside the lease or in the other protocol's range, an unleased source
address, another endpoint port, a flow a foreign prerouting DNAT translated onto the
endpoint and the host's own opening toward the transit address stay dark; withdrawal
stops new flows and the established one, and release reopens the path. A graph with
served endpoints persists across owner loss and survives lost-ACK replay.
Workload grants are qualified between three workload namespaces behind the router,
each with leased public egress on every port of both protocols: against pre-policy
positive controls, the grant opens exactly the source workload's TCP and UDP flows to
the granted ports with the source preserved; another port, another source workload,
another destination workload, the destination's TCP and UDP openings toward the
source stay dark; withdrawal stops new flows and the established one; release reopens
the path. A set of 1024 grants persists across owner loss, is re-verified element by
element and survives lost-ACK replay.
Exclusive forwarding is qualified on a host namespace with a handoff, an uplink and
a LAN link: against positive controls with no policy and with the scope without the
member, the member denies the LAN's IPv4 to the uplink peer, the peer's way back into
the LAN, the LAN's IPv6 and a LAN flow established before it, while leased egress
still leaves translated and an unleased port stays denied; without handoffs leased
egress drops too; replacing it with a scope without the member and releasing an
exclusive scope both reopen the LAN path. The same probe fails against a compiler
without the drop.
Loopback TCP port owners are qualified against pre-policy positive controls on the
same listeners: the owning uid (root, and uid 1000 for a second port) connects, every
other uid is reset with ECONNREFUSED, an unowned port and another loopback address
stay open to every user, a foreign output DNAT to the owned port stays closed to
another user, and an uplink arrival translated to loopback with route_localnet
enabled is dropped. The policy keeps enforcing after the owner detaches and its
process is gone, a fresh owner verifies and adopts the exact graph, and release
reopens every probe. The maximum of eight owners beside the guard persists across
owner loss and survives lost-ACK replay.
The arm64 binary is cross-built; native packet qualification is currently x86_64.
Kernel 6.8 is unsupported. No production activation is implied by these tests.
Native dependency, standard-library and C runtime attribution is included in
native-notices/, generated by tsrust notices from rust/Cargo.lock and the Rust
toolchain pinned in rust-toolchain.toml. Every build verifies the committed notices
first and refuses when they no longer match. Distribution must retain these notices
alongside the binaries. The statically linked musl 1.2.5 comes from the Rust 1.95.0
toolchain pinned in rust-toolchain.toml, whose musl build includes the two
CVE-2025-26519 patches.
Quick Start
import { SmartNftables } from '@push.rocks/smartnftables';
const nft = new SmartNftables();
await nft.initialize();
// Port forward 8080 → 192.168.1.100:80
await nft.nat.addPortForwarding('web', {
sourcePort: 8080,
targetHost: '192.168.1.100',
targetPort: 80,
});
// Block a suspicious IP
await nft.firewall.blockIP('10.0.0.99');
// Rate limit HTTP to 100 req/s per IP
await nft.rateLimit.addRateLimit('http-limit', {
port: 80,
protocol: 'tcp',
rate: '100/second',
perSourceIP: true,
});
// Clean up everything when done
await nft.cleanup();
Architecture 🏗️
The library is organized around a facade pattern with specialized sub-managers:
SmartNftables (main facade)
├── nat → NatManager (DNAT, SNAT, masquerade)
├── firewall → FirewallManager (filter rules, IP sets, stateful tracking)
└── rateLimit → RateLimitManager (packet/connection rate limiting)
All rules are tracked in rule groups identified by string IDs, so you can add, inspect, and remove them programmatically.
API Reference
SmartNftables — Main Facade
const nft = new SmartNftables({
tableName: 'smartnftables', // nftables table name (default: 'smartnftables')
family: 'ip', // 'ip' | 'ip6' | 'inet' (default: 'ip')
dryRun: false, // generate commands without executing (default: false)
});
| Method | Description |
|---|---|
initialize() |
Create the nftables table and NAT chains. Idempotent. |
cleanup() |
Delete the entire table and clear all tracking. |
status() |
Get an INftStatus report of the current managed state. |
applyRuleGroup(id, commands) |
Apply and track a group of raw nft commands. |
removeRuleGroup(id) |
Remove a tracked rule group. |
getRuleGroup(id) |
Retrieve a tracked rule group by ID. |
🌐 NAT — nft.nat
Port Forwarding (DNAT)
await nft.nat.addPortForwarding('my-service', {
sourcePort: 443,
targetHost: '10.0.0.5',
targetPort: 8443,
protocol: 'tcp', // 'tcp' | 'udp' | 'both' (default: 'tcp')
preserveSourceIP: false, // skip masquerade if true (default: false)
});
await nft.nat.removePortForwarding('my-service');
Port Range Forwarding
Map a range of ports to another host:
// Forward ports 3000-3010 → 10.0.0.5:3000-3010
await nft.nat.addPortRange('dev-ports', 3000, 3010, '10.0.0.5', 3000, 'tcp');
await nft.nat.removePortRange('dev-ports');
SNAT (Source NAT)
await nft.nat.addSnat('egress', {
sourceAddress: '203.0.113.1',
targetPort: 80,
protocol: 'tcp',
});
Masquerade
await nft.nat.addMasquerade('outbound', {
targetPort: 443,
protocol: 'tcp',
});
🛡️ Firewall — nft.firewall
Basic Rules
await nft.firewall.addRule('allow-ssh', {
direction: 'input', // 'input' | 'output' | 'forward'
action: 'accept', // 'accept' | 'drop' | 'reject'
sourceIP: '10.0.0.0/24',
destPort: 22,
protocol: 'tcp',
comment: 'Allow SSH from trusted network',
});
await nft.firewall.removeRule('allow-ssh');
When sourcePort or destPort is provided without protocol, TCP is used by default. Protocol-only rules match the specified layer-4 protocol.
Block an IP
await nft.firewall.blockIP('10.0.0.99');
await nft.firewall.blockIP('192.168.0.0/16', { direction: 'forward' });
Allow Only Specific IPs on a Port
// Only these IPs can reach port 3306 — everything else is dropped
await nft.firewall.allowOnlyIPs('db-access', ['10.0.0.1', '10.0.0.2'], 3306, 'tcp');
Stateful Connection Tracking
// Allow established/related, drop invalid — on the input chain
await nft.firewall.enableStatefulTracking('input');
IP Sets
Create named sets and match against them:
// Create a set of blocked IPs
await nft.firewall.createIPSet({
name: 'blocklist',
type: 'ipv4_addr',
elements: ['10.0.0.50', '10.0.0.51'],
});
// Dynamically add/remove elements
await nft.firewall.addToIPSet('blocklist', ['10.0.0.52']);
await nft.firewall.removeFromIPSet('blocklist', ['10.0.0.50']);
// Clean up
await nft.firewall.deleteIPSet('blocklist');
You can also build set-matching rules directly with the low-level builder:
import { buildIPSetMatchRule } from '@push.rocks/smartnftables';
const rule = buildIPSetMatchRule('smartnftables', 'ip', {
setName: 'blocklist',
direction: 'input',
matchField: 'saddr',
action: 'drop',
});
⏱️ Rate Limiting — nft.rateLimit
Packet Rate Limiting
// Global: drop packets over 1000/second on port 80
await nft.rateLimit.addRateLimit('http-global', {
port: 80,
protocol: 'tcp',
rate: '1000/second',
burst: 50,
action: 'drop',
});
// Per-IP: each source IP gets its own 100/second limit
await nft.rateLimit.addRateLimit('http-per-ip', {
port: 80,
protocol: 'tcp',
rate: '100/second',
perSourceIP: true,
});
await nft.rateLimit.removeRateLimit('http-per-ip');
Connection Rate Limiting
Limit the rate of new connections (uses ct state new):
await nft.rateLimit.addConnectionRateLimit('ssh-connrate', {
port: 22,
protocol: 'tcp',
rate: '5/second',
perSourceIP: true,
});
await nft.rateLimit.removeConnectionRateLimit('ssh-connrate');
🔧 Low-Level Rule Builders
For advanced use cases, you can generate raw nft command strings without applying them:
import {
buildDnatRules,
buildSnatRule,
buildMasqueradeRule,
buildFirewallRule,
buildRateLimitRule,
buildPerIpRateLimitRule,
buildConnectionRateRule,
buildIPSetCreate,
buildIPSetAddElements,
buildIPSetRemoveElements,
buildIPSetDelete,
buildIPSetMatchRule,
buildTableSetup,
buildFilterChains,
buildTableCleanup,
} from '@push.rocks/smartnftables';
const commands = buildDnatRules('mytable', 'ip', {
sourcePort: 8080,
targetHost: '10.0.0.5',
targetPort: 80,
});
// → ['nft add rule ip mytable prerouting tcp dport 8080 dnat to 10.0.0.5:80',
// 'nft add rule ip mytable postrouting tcp dport 80 masquerade']
Dry Run Mode 🧪
Generate commands without touching the kernel — perfect for testing, debugging, or CI:
const nft = new SmartNftables({ dryRun: true });
await nft.initialize();
await nft.nat.addPortForwarding('test', {
sourcePort: 80,
targetHost: '10.0.0.1',
targetPort: 8080,
});
console.log(nft.status());
// Rules tracked in memory, nothing executed
Status Reporting 📊
const status = nft.status();
// {
// initialized: true,
// tableName: 'smartnftables',
// family: 'ip',
// isRoot: true,
// activeGroups: 3,
// groups: {
// 'nat:web': { ruleCount: 2, createdAt: 1711411200000 },
// 'fw:block-10_0_0_99': { ruleCount: 1, createdAt: 1711411200100 },
// 'ratelimit:http-limit': { ruleCount: 1, createdAt: 1711411200200 },
// }
// }
Types
All interfaces and types are fully exported for use in your own code:
| Type | Description |
|---|---|
INftDnatRule |
DNAT port forwarding rule config |
INftSnatRule |
Source NAT rule config |
INftMasqueradeRule |
Masquerade rule config |
INftFirewallRule |
Firewall filter rule config |
INftIPSetConfig |
IP set creation config |
INftRateLimitRule |
Rate limiting rule config |
INftConnectionRateRule |
New-connection rate limit config |
ISmartNftablesOptions |
Constructor options |
INftStatus |
Status report shape |
TNftProtocol |
'tcp' | 'udp' | 'both' |
TNftFamily |
'ip' | 'ip6' | 'inet' |
TFirewallAction |
'accept' | 'drop' | 'reject' |
TCtState |
'new' | 'established' | 'related' | 'invalid' |
License and Legal Information
This repository contains open-source code licensed under the MIT License. A copy of the license can be found in the repository license file.
Please note: The MIT License does not grant permission to use the trade names, trademarks, service marks, or product names of the project, except as required for reasonable and customary use in describing the origin of the work and reproducing the content of the NOTICE file.
Trademarks
This project is owned and maintained by Task Venture Capital GmbH. The names and logos associated with Task Venture Capital GmbH and any related products or services are trademarks of Task Venture Capital GmbH or third parties, and are not included within the scope of the MIT license granted herein.
Use of these trademarks must comply with Task Venture Capital GmbH's Trademark Guidelines or the guidelines of the respective third-party owners, and any usage must be approved in writing. Third-party trademarks used herein are the property of their respective owners and used only in a descriptive manner, e.g. for an implementation of an API or similar.
Company Information
Task Venture Capital GmbH
Registered at District Court Bremen HRB 35230 HB, Germany
For any legal inquiries or further information, please contact us via email at hello@task.vc.
By using this repository, you acknowledge that you have read this section, agree to comply with its terms, and understand that the licensing of the code does not imply endorsement by Task Venture Capital GmbH of any derivative works.