NixOS edge router network design
The edge router in this homelab runs NixOS on an x86_64 appliance, terminating public WAN traffic and providing inter-VLAN routing, stateful firewalling, DHCP, and edge proxying for the fleet. This reference documents the router’s network architecture: tagged 802.1Q trunking, default-drop nftables zones with container bridge forwarding, Kea DHCP socket isolation, hairpin access without split-horizon DNS lies, loopback service-plane addressing, and the 18 post-deployment runtime health checks enforced by eaves doctor.
The router operates alongside the NAS and home-automation appliances described in Three NixOS hosts, one deploy interface, isolates smart-home peripherals as detailed in Isolating a smart-home fleet on an IoT VLAN, advertises internal routes to the tailnet in A Tailscale tailnet across a dual-site homelab, and aligns with the interface tuning covered in Tuning a 10GbE link end to end.
- Tagged trunk parent carries no address: The primary LAN trunk interface carries 802.1Q tagged child interfaces exclusively. The parent link runs with
LinkLocalAddressing = "no"and no IP address, eliminating raw-socket packet cross-capture in the DHCP daemon. - Admin segment is tagged VLAN 69: The management and workstation segment was migrated from an untagged native network to tagged VLAN 69 (
10.0.69.0/24) to eliminate switch port drift and prevent Kea from cross-allocating addresses across segments. - Scoped nftables table flushes: NixOS manages firewall rules with
flushRuleset = falseand opens with scopedflush table inet filterandflush table inet natcommands. This allows idempotent rule reloads without destroying Docker’s separateip natchains. - Docker bridge forward traversal: Linux kernel
br_netfilterdirects bridged container traffic through the host’s default-drop forward chain. Explicit intra-bridge and DNAT-conntrack rules permit container communication and testcontainer dialing while blocking lateral container-to-host hops. - Kea carrier re-detection: Kea is configured with
re-detect = trueso its interface manager re-binds raw sockets as switch ports complete their two-minute boot cycle, avoiding deaf DHCP listeners at startup. - Loopback service-plane addressing: Internal services and DNS listeners bind dedicated
/32loopback aliases onlo(including Knotea on10.0.10.5), bypassing Linux macvlan parent-child communication barriers. - Hairpin without split-horizon DNS: Public hostnames resolve to public IPs or loopback aliases across all networks; LAN packets hitting the router’s WAN IP terminate directly at the local reverse proxy, preserving authentic client source IPs in logs without NAT hairpins.
- 18-check doctor gate: Every deployment and automated run asserts 18 runtime conditions via
eaves doctor, testing kernel parameters, descriptor rings, default-drop firewall policies, orphan VLAN interfaces, and port ownership.
Architecture
Section titled “Architecture”The edge router bridges external WAN traffic, local VLANs, and Tailscale overlay routing through a unified nftables ruleset and host-networked edge services.
The diagram outlines the router’s packet paths:
- External traffic enters through the WAN interface or the Tailscale subnet router (
tailscale0). - The packet traverses the
inet filterinput chain for host services or the forward chain for inter-segment and egress routing. - Local management services and DNS queries terminate on loopback aliases (
lo), while DHCP is served by Kea directly to tagged VLAN children. - Filtered packets cross the physical SFP+ trunk link into the managed switch, which routes them to their destination VLANs based on 802.1Q tags.
VLAN segmentation and trunk architecture
Section titled “VLAN segmentation and trunk architecture”The physical link between the edge router and the core managed switch is a 10GbE direct-attach copper (DAC) cable on an Intel X710-DA2 controller. A secondary Intel i226-V 2.5GbE interface handles the WAN connection to the optical network terminal.
The bare trunk parent doctrine
Section titled “The bare trunk parent doctrine”The physical trunk interface does not hold an IPv4 address. Under systemd-networkd, the interface configuration explicitly drops link-local addressing and disables online requirements:
# 15-trunk.network[Match]Name=enp2s0f0np0
[Network]LinkLocalAddressing=noConfigureWithoutCarrier=yesVLAN=adminVLAN=wirelessVLAN=serversVLAN=turingVLAN=iotVLAN=guest
[Link]RequiredForOnline=noMTUBytes=1500Every routable subnet is bound to an explicit netdev of type vlan. In systemd-networkd, omitting MTUBytes causes the daemon to leave whatever value was previously set on the interface. The MTU is declared explicitly as 1500 across the trunk and all child interfaces to ensure consistent packet sizing across rebuilds.
Why admin moved to tagged VLAN 69
Section titled “Why admin moved to tagged VLAN 69”In the early router architecture, the administrative network operated as an untagged native network on the trunk parent link. This was abandoned for two operational reasons:
- Switch configuration drift: Port membership for untagged VLAN 1 on managed switches is frequently maintained as an implicit default table. Routine administrative edits to other ports caused switch firmware to silently alter untagged port memberships.
- Kea raw-socket cross-capture: Kea binds raw
AF_PACKETsockets to interfaces to process DHCP broadcast frames from unconfigured clients.1 When Kea listened on a bare parent interface alongside tagged VLAN interfaces, its raw socket captured tagged broadcast frames from child VLANs before the kernel stripped the 802.1Q headers. Kea loggedDHCPSRV_MULTIPLE_RAW_SOCKETS_PER_IFACEand issued wrong-subnet lease offers to newly connected clients on the untagged network, preventing devices from acquiring valid IP addresses.
Switching Kea to standard UDP sockets would resolve the cross-capture issue but would break initial DHCP discovery, as cold clients without an assigned IP address communicate via raw L2 broadcast. Moving the administrative network to tagged VLAN 69 solved this permanently. The kernel strips 802.1Q tags before passing packets to the VLAN netdevs, allowing Kea to bind cleanly to distinct virtual interfaces without cross-capturing frames. Switch access ports for workstations and out-of-band management assign PVID 69, keeping client configurations untagged.
Segment layout
Section titled “Segment layout”Each VLAN serves a distinct trust tier and administrative role:
| VLAN Name | 802.1Q Tag | Subnet / Range | Primary Purpose | DHCP Mode |
|---|---|---|---|---|
admin | 69 | 10.0.69.0/24 | Workstation (erfi1), PiKVM, home hub (erfipie / 10.0.69.7), switch management | Dynamic pool (.100 - .200) + reservations |
wireless | 100 | 10.0.72.0/24 | Wi-Fi clients and access point management on GL.iNet Flint 2 | Dynamic pool (.100 - .200) + reservations |
servers | 200 | 10.0.71.0/24 | Storage and compute host (servarr), Docker macvlan containers (servarr_lan) | Reservations-only (no dynamic pool) |
turing | 300 | 10.0.74.0/24 | Turing Pi 2 cluster: BMC controller and RK1 compute nodes | Reservations-only (no dynamic pool) |
iot | 400 | The IoT VLAN | Home automation plugs, sensors, and relays on dedicated SSID | Dynamic pool (.100 - .200) + reservations |
guest | 500 | The guest VLAN | Guest Wi-Fi devices with direct internet egress only | Dynamic pool (.100 - .200) |
The servers segment operates without a dynamic pool. IP addresses on 10.0.71.0/24 outside the router’s gateway are hand-allocated to Docker macvlan containers (such as the Kanban application at 10.0.71.66) and server infrastructure. The turing segment operates identically: the BMC controller holds a static reservation, cluster nodes assign static addresses, and a reserved block is assigned to the MetalLB load-balancer pool.
Trunk link buffer tuning
Section titled “Trunk link buffer tuning”At 9.4 Gbps link saturation on MTU 1500, a 10GbE network processes roughly 810,000 packets per second. The default ring buffer depth in the Intel i40e driver is 512 descriptors.2 Under CPU scheduling jitter or RSS core interrupts, a 512-descriptor buffer fills in less than a millisecond, causing the network controller to drop frames that have already arrived cleanly over the wire (tracked in the kernel counter rx_missed_errors).
Tuning the trunk link requires two mechanisms:
- MAC-matched systemd.link file: Descriptor rings must be set to 4096 via a
.linkfile. Matching must be performed againstmatchConfig.MACAddress, notmatchConfig.Name. Udev evaluates.linkfiles when the network interface first appears, at which point the kernel name is stilleth0. Matching against predictable interface names fails during boot, causing the system to fall back to default ring sizes. - Activation oneshot unit: Because udev does not re-evaluate
.linkfiles duringnixos-rebuild switch, an idempotent systemd oneshot unit executesethtool -G <trunk> rx 4096 tx 4096during activation if the current descriptor depth is below 4096.
On the WAN interface, ethtool -K <wan> rx-udp-gro-forwarding on is applied via a oneshot unit. Generic Receive Offload (GRO) forwarding prevents Tailscale WireGuard packets from being fragmented and re-segmented when routed through the edge.3
The service-plane address
Section titled “The service-plane address”Internal services and local DNS resolvers on the router do not attach to a physical VLAN interface. Instead, they bind to /32 loopback aliases configured on lo:
# 05-lo.network[Match]Name=lo
[Network]Address=10.0.10.5/32Loopback aliases are used for three reasons:
- Macvlan interface isolation: Linux macvlan sub-interfaces in bridge mode cannot communicate directly with the parent host interface. If the DNS resolver ran as a container on a macvlan network on the router, it would be unable to reach the host router as its recursive gateway. Binding services to
loon the host network avoids this architectural limitation. - Decoupled network identities: Services bound to loopback addresses hold stable identities that are independent of physical interface numbering. If a core service moves to a different physical machine, the router can route the
/32address to the new host without reconfiguring DHCP option 6 or modifying firewall rules across the fleet. - Universal local routing: Loopback addresses are inherently routable from every internal VLAN through the router’s local routing table, subject only to nftables input filtering.
The Knotea resolver binds directly to 10.0.10.5:53, providing recursive and authoritative DNS resolution across the homelab as detailed in knotea: one binary that is both your recursive resolver and your authoritative DNS. The edge reverse proxy binds to the service-plane loopback address for internal HTTP and SSH forwarding, documented in Edge Caddy as a native NixOS service.
nftables zone model and policies
Section titled “nftables zone model and policies”The router disables the default NixOS firewall module (networking.firewall.enable = false) and manages rules directly through networking.nftables.
Scoped table flushing
Section titled “Scoped table flushing”NixOS defaults to flushing the entire firewall ruleset upon activation. This causes severe regressions when running Docker on the host: Docker manages its own NAT and bridge isolation rules in the ip nat and ip filter tables using iptables-nft compatibility layers. A global ruleset flush deletes Docker’s masquerade rules, severing container network connectivity until dockerd is restarted.
Setting flushRuleset = false prevents NixOS from clearing Docker’s tables. However, a rebuild with flushRuleset = false only appends new rules, leaving obsolete chains active in memory. To achieve idempotent rebuilds without disturbing Docker, the configuration ruleset executes scoped table flushes:
table inet filterflush table inet filtertable inet natflush table inet natNixOS places all router-managed filtering in table inet filter and all router-managed NAT in table inet nat. Docker’s rules reside in table ip nat and table ip filter, allowing both systems to reload independently without collision.4
Input chain policies
Section titled “Input chain policies”The input chain drops all incoming traffic by default (policy drop). Conntrack accepts established and related packets while dropping invalid states. Inbound traffic is accepted based on source interface and service role:
table inet filter { chain input { type filter hook input priority 0; policy drop; iifname "lo" accept ct state vmap { established : accept, related : accept, invalid : drop }
# Universal LAN services iifname { "admin", "wireless", "servers", "turing", "iot", "guest" } icmp type echo-request accept iifname { "admin", "wireless", "servers", "turing", "iot", "guest" } udp dport { 67, 68 } accept
# Core LAN services (IoT excluded) iifname { "admin", "wireless", "servers", "turing", "guest" } udp dport 53 accept iifname { "admin", "wireless", "servers", "turing", "guest" } tcp dport 53 accept iifname { "admin", "wireless", "servers", "turing" } udp dport 5353 accept iifname { "admin", "wireless", "servers", "turing" } tcp dport { 80, 443, 2222, 2223 } accept iifname { "admin", "wireless", "servers", "turing" } udp dport 443 accept
# Administrative plane iifname "admin" tcp dport { 22, 8080 } accept iifname "tailscale0" tcp dport { 22, 8080 } accept iifname "tailscale0" icmp type echo-request accept
# Pinned ingress rules iifname "servers" ip saddr 10.0.71.66 tcp dport 8080 accept iifname "admin" ip saddr <switch-ip> udp dport 514 accept
# Docker bridge input to local services iifname { "docker0", "forgejo0" } udp dport 53 accept iifname { "docker0", "forgejo0" } tcp dport 53 accept iifname { "docker0", "forgejo0" } ip daddr <service-plane-ip> tcp dport { 443, 2222 } accept
# WAN ingress iifname "enp1s0" tcp dport { 80, 443, 2222, 2223 } accept iifname "enp1s0" udp dport 443 accept iifname "enp1s0" icmp type echo-request limit rate 5/second accept
# IoT audit logging iifname "iot" ip daddr != 224.0.0.0/4 limit rate 10/minute log prefix "iot-in-drop: " }}The input chain enforces key isolation boundaries:
- DNS isolation: The Knotea resolver on
10.0.10.5:53is accessible to all LAN segments except the IoT VLAN. IoT devices communicate with the Home Assistant hub exclusively by IP address, removing DNS lookups as an exfiltration channel. - Administrative control: SSH (port 22) and the Composer management API (port 8080) are reachable only from the administrative interface (
admin) and the Tailscale interface (tailscale0). The Kanban dashboard container on the servers VLAN is granted a single pinned exception to communicate with port 8080. - IoT surveillance: Any unpermitted packet from the IoT VLAN destined for the router is rate-limited to 10 logs per minute and recorded in the system journal with the prefix
iot-in-drop:.
Forward chain policies
Section titled “Forward chain policies”The forward chain enforces strict default-drop inter-VLAN routing:
chain forward { type filter hook forward priority 0; policy drop; ct state vmap { established : accept, related : accept, invalid : drop }
# Tailscale subnet routing: local VLANs only, no WAN exit iifname "tailscale0" ip saddr 100.64.0.0/10 ip daddr 10.0.0.0/8 accept
# Workstation full lateral reach iifname "admin" ip saddr 10.0.69.3 accept
# Hub to IoT management plane iifname "admin" ip saddr 10.0.69.7 oifname "iot" accept
# IoT to Hub: NTP only iifname "iot" ip daddr 10.0.69.7 udp dport 123 oifname "admin" accept iifname "iot" limit rate 10/minute log prefix "iot-fwd-drop: "
# Monitoring: Prometheus to Hub metrics iifname "servers" ip daddr 10.0.69.7 tcp dport 8123 oifname "admin" accept
# Internet egress (IoT excluded) iifname { "admin", "wireless", "servers", "turing", "guest" } ip daddr != 10.0.0.0/8 oifname "enp1s0" accept
# Docker bridge forwarding iifname "docker0" oifname "docker0" accept iifname "forgejo0" oifname "forgejo0" accept iifname { "docker0", "forgejo0" } oifname { "docker0", "forgejo0" } ct status dnat accept iifname { "docker0", "forgejo0" } ip daddr != 10.0.0.0/8 oifname "enp1s0" accept}The forward chain contains five critical routing controls:
- Tailscale subnet boundary: Inbound traffic from
tailscale0is restricted to RFC1918 internal space (10.0.0.0/8). The router does not forward tailnet traffic to the WAN interface, preventing the machine from operating as an unintentional exit node. - Administrative workstation: The primary workstation (
erfi1at10.0.69.3) is granted unrestricted lateral forwarding to manage infrastructure across all segments. - Smart-home confinement: The home automation hub (
erfipieat10.0.69.7) can initiate connections into the IoT VLAN for Home Assistant native API polling (port 6053), ESPHome over-the-air updates (port 3232), and web configuration (port 80). The IoT VLAN cannot initiate connections back into the network, with a single pinhole for NTP time synchronisation (UDP port 123) to Chrony on the hub. - Internet egress: All segments except the IoT VLAN may forward traffic out the WAN interface to non-internal addresses (
ip daddr != 10.0.0.0/8). The IoT VLAN has no egress rule; unpermitted forwarding attempts are logged at 10 per minute withiot-fwd-drop:before being dropped.
Docker bridge forward traversal
Section titled “Docker bridge forward traversal”When Docker creates bridge networks on a Linux host, the kernel’s br_netfilter module passes bridged Ethernet frames through the host’s iptables and nftables forward chains.5 In a default-drop firewall architecture, this causes container-initiated traffic to fail:
- Same-bridge communication: Packets between two containers on
docker0traverse the forward chain. Without explicit rules, container-initiated communication (such as an application container reaching a database container) is dropped by the default policy. The router addresses this with explicit intra-bridge acceptance:iifname <b> oifname <b> accept. - Cross-bridge testcontainer access: In continuous-integration workflows on the router’s Forgejo instance, a runner container on
forgejo0executes integration tests that spin up ephemeral testcontainers publishing ports on other Docker bridges. The runner dials the container via the host gateway port. Docker DNATs this traffic inPREROUTING. Because the destination IP is transformed to a container IP on another bridge, the packet enters the forward chain across two distinct interfaces. - The DNAT conntrack pinhole: Directly accepting all traffic between bridges would break container isolation. The forward chain inspects conntrack status:
iifname { "docker0", "forgejo0" } oifname { "docker0", "forgejo0" } ct status dnat accept. This restricts cross-bridge routing strictly to packets that underwent legitimate Docker port-publishing DNAT, while dropping attempts to route directly to raw container IPs.
NAT rules
Section titled “NAT rules”NAT configuration resides in table inet nat. The router handles masquerading and selective DNS steering:
table inet nat { chain prerouting { type nat hook prerouting priority dstnat; policy accept; # Media device DNS steering: redirect hardcoded queries to local Knotea resolver iifname "wireless" ip saddr <media-device-ip> ip daddr != 10.0.0.0/8 udp dport 53 dnat to 10.0.10.5 iifname "wireless" ip saddr <media-device-ip> ip daddr != 10.0.0.0/8 tcp dport 53 dnat to 10.0.10.5 } chain postrouting { type nat hook postrouting priority 100; policy accept; ip saddr 10.0.0.0/8 oifname "enp1s0" masquerade }}The prerouting chain intercepts DNS queries from devices with hardcoded upstream servers (such as smart TVs directing queries to public resolvers) and redirects them to the local Knotea resolver at 10.0.10.5:53. The postrouting chain masquerades all internal RFC1918 traffic exiting the WAN interface.
DHCP architecture with Kea
Section titled “DHCP architecture with Kea”The router runs ISC Kea DHCPv4 (kea-dhcp4-server.service) to manage IP leases across all local segments.
Switch boot race and socket recovery
Section titled “Switch boot race and socket recovery”In an edge router deployment, the host boots in seconds while managed switches require up to two minutes to load firmware, initialise ASIC matrices, and raise link carrier. When the router activates Kea at boot, the physical trunk interface and its 802.1Q child VLANs exist as netdevs in the kernel but lack carrier flags (LOWER_UP).
By default, Kea attempts to bind raw sockets to configured interfaces at startup. If an interface is not operational, socket allocation fails, Kea reports DHCPSRV_NO_SOCKETS_OPEN, and the daemon halts. If systemd attempts restarts while the switch is still booting, Kea trips systemd’s restart rate limits and enters a permanent failed state.
The configuration hardens Kea against this race through three options:
systemd.services.kea-dhcp4-server = { after = [ "sys-subsystem-net-devices-trunk.device" ] ++ map (i: "sys-subsystem-net-devices-${i}.device") vlanIfaces; unitConfig.StartLimitIntervalSec = lib.mkForce 0; serviceConfig = { Restart = lib.mkForce "always"; RestartSec = 5; };};
services.kea.dhcp4.settings = { interfaces-config = { interfaces = lanIfaces; re-detect = true; }; lease-database = { type = "memfile"; persist = true; name = "/var/lib/kea/leases4.csv"; }; valid-lifetime = 300;};Setting interfaces-config.re-detect = true instructs Kea’s interface manager (IfaceMgr) to poll the kernel link state periodically and re-bind raw sockets as interfaces gain carrier. Disabling systemd’s restart rate limits ensures Kea continues retrying until the switch completes its boot sequence.
Allocation and reservation policy
Section titled “Allocation and reservation policy”Kea generates a subnet4 declaration for each configured VLAN. Address pools are differentiated by trust tier:
- Dynamic pools (
.100 - .200): Configured onadmin,wireless,iot, andguest. Standard clients receive temporary leases with a 300-second valid lifetime, allowing rapid DHCP configuration changes to propagate across the fleet. - Reservations-only: Configured on
serversandturing. Thesubnet4blocks declarepool = [], disabling dynamic lease offers entirely. Devices on these segments must hold explicit MAC-to-IP reservations in Kea or configure static IP addresses. This prevents unknown hardware from leasing addresses on sensitive infrastructure segments.
Across all scopes, Kea distributes the Knotea loopback resolver address (10.0.10.5) as the primary DNS server.
Hairpin routing without DNS lies
Section titled “Hairpin routing without DNS lies”Split-horizon DNS - answering with private IP addresses for internal clients and public IP addresses for external clients - creates subtle operational failure modes. Modern web browsers and operating systems increasingly utilise DNS-over-HTTPS (DoH) or DNS-over-TLS (DoT) by default, bypassing local LAN resolvers and receiving public IP addresses. When split DNS is used, off-site laptops returning to the office cache public records, while on-site devices fail when accessing services via alternative resolvers.
The fleet eliminates split-horizon DNS lies for public hostnames. A single set of authoritative DNS records exists on public nameservers, pointing all service hostnames to the edge router’s public WAN IP.
Local edge termination
Section titled “Local edge termination”Because the public IP address terminates directly on the router’s WAN interface, hairpinned traffic requires no NAT translation:
- A LAN client (e.g.
10.0.69.3) resolves a public service hostname. The resolver returns the router’s public WAN IP address. - The client transmits packets destined for the public IP. Because the router is the client’s default gateway, the packet arrives on the router’s VLAN interface.
- The Linux routing engine inspects the destination IP, recognises it as an address assigned to a local interface, and directs the packet to the
inputfirewall chain. - The input rule
iifname { ... } tcp dport { 80, 443 } acceptpermits the connection. Caddy accepts the TCP connection directly on the host network.
Because the packet terminates locally on the router, no source NAT (masquerade) is required. Conntrack tracks the connection state, and return packets travel directly from the router back to the client’s LAN IP. Caddy’s access logs record the authentic client IP (10.0.69.3) rather than a masqueraded router gateway address.
Post-deploy verification: the eaves doctor gate
Section titled “Post-deploy verification: the eaves doctor gate”To ensure that changes to NixOS configurations do not introduce silent network regressions, the edge router deploys through a post-switch verification gate: eaves doctor.
The doctor is a standalone Go diagnostic binary that executes 18 automated runtime checks against the running system. It is invoked immediately following nixos-rebuild switch and runs periodically on an hourly timer. Any failure exits non-zero and triggers alerting.
The 18 checks
Section titled “The 18 checks”The 18 checks are implemented in internal/show/doctor.go within the eaves toolchain:
| # | Check Name | Target / Source | Assertion | Regression / Failure Mode Guarded |
|---|---|---|---|---|
| 1 | kernel-cmdline-i226 | /proc/cmdline | Contains pcie_port_pm=off and igc.eee_enable=0 | Intel i226-V PCIe power management and Energy Efficient Ethernet cause link flapping |
| 2 | ip-forwarding | /proc/sys/net/ipv4/ip_forward, .../ipv6/conf/all/forwarding | Both values equal 1 | Kernel forwarding disabled, causing the router to stop routing packets |
| 3 | wan-link | Kernel routing table & interfaces | Default route exists on an interface holding a scope-global IP | WAN interface loss, upstream DHCP failure, or missing default gateway |
| 4 | nft-default-policies | inet filter ruleset | input policy is drop and forward policy is drop | Inadvertent rebuild regression reverting firewall policies to open accept |
| 5 | nft-conntrack-accept | inet filter forward chain | Rule with ct state established accept exists | Missing conntrack state rule, dropping all established return traffic |
| 6 | nat-masquerade | inet nat postrouting chain | Rule with masquerade action exists | Missing outbound NAT, cutting off LAN client internet access |
| 7 | kea-running | Systemd unit state | kea-dhcp4-server.service is active | DHCP daemon crash or systemd start-limit failure |
| 8 | kea-topology | Kea config vs kernel links | Every Kea-served interface exists, is UP, has kind vlan, and holds an IP | Kea listening on bare trunk parent instead of tagged VLAN child |
| 9 | vlan-no-orphans | Kernel network interfaces | Every live 802.1Q VLAN link on the system has an IPv4 address | Leaked netdevs from rolled-back NixOS generations (networkd never deletes netdevs) |
| 10 | kea-no-raw-socket-warn | Systemd journal (Kea) | No MULTIPLE_RAW_SOCKETS_PER_IFACE in last 300 log entries | Raw socket packet cross-capture on bare parent interfaces |
| 11 | conntrack-headroom | /proc/sys/net/netfilter/ | Active conntrack count is below 60% (warn) and 80% (fail) of max | State table exhaustion under network load or SYN flooding |
| 12 | resolved-running | Systemd unit state | systemd-resolved.service is active | Local host name resolution failure on the router itself |
| 13 | docker-nat-intact | ip nat ruleset | If Docker is active, table ip nat contains chain DOCKER | Global nft flush ruleset wiped Docker’s iptables-nft rules |
| 14 | nixos-checkout | /etc/nixos git repository | Working tree is clean (--porcelain) and HEAD matches origin/main | Uncommitted on-box edits or detached HEAD diverging from git repository |
| 15 | trunk-link | Sysfs net speed/duplex | SFP+ trunk parent link is UP, has carrier, and operates at 10000/full | Silent link degradation or auto-negotiation downshift to 1GbE |
| 16 | trunk-errors | Sysfs statistics | 13 kernel error counters (missed, crc, fifo, frame) are zero | Physical cable, DAC, or switch port interface errors |
| 17 | trunk-rings | ethtool -g runtime status | Current RX and TX descriptor ring depths equal declared 4096 | MAC .link file failed to apply at boot, reverting buffers to driver default 512 |
| 18 | edge-services | Systemd units & ss -tlnp | Caddy, edgectl, and composer active; :8080 owned by composerd, :8082 by edgectl | Port collision between edge control plane and container management plane |
Three structural defects caught by the doctor
Section titled “Three structural defects caught by the doctor”The doctor suite was written to catch specific historical failures:
- Orphan VLAN interface leakage (
vlan-no-orphans): Systemd-networkd creates netdevs it is instructed to manage, but it never deletes virtual interfaces that are removed from subsequent configurations. Activating an older generation for testing permanently leaves its VLAN interfaces active in the kernel. In August 2026, an old generation was tested for twenty seconds and left four unmanaged links behind. These orphaned interfaces inherited the trunk’s MTU and sat active on tags that the current configuration no longer recognised. Thevlan-no-orphanscheck asserts that every 802.1Q child interface possesses a valid gateway IP address, flagging any unmanaged netdev leakage. - Descriptor ring regression (
trunk-rings): In September 2026, a power outage rebooted the router. The trunk.linkfile had been written matching onName=enp2s0f0np0. Because udev evaluates link files when the device is still namedeth0, the rule failed to match, and the driver initialised the rings at the default depth of 512. Within four hours, scheduling latency on the RSS core caused 239,000 dropped packets (rx_missed_errors). Thetrunk-ringscheck asserts the applied ring depth directly withethtool -g, ensuring buffer regressions fail the deployment gate immediately. - Edge management port collisions (
edge-services): When Caddy was migrated to run natively on the host network,edgectland the Docker container forcomposerboth attempted to bind host port 8080. Depending on which service started first,composerfailed to start orcomposer.erfi.iorouted traffic into theedgectlmetrics dashboard. Theedge-servicescheck parsesss -tlnpoutput to verify that port 8080 is owned exclusively bycomposerdand port 8082 is owned byedgectl.
Decision guide
Section titled “Decision guide”Use this guide when adding new network interfaces, container bridges, or routing rules to the edge router.
The decision sequence operates as follows:
- New LAN VLAN: Define the subnet under
vlansinconfiguration.nix. Ensure the trunk parent interface remains unnumbered, declareMTUBytes = 1500explicitly on the child netdev, assign a Kea subnet pool, and configure switch ports with the appropriate 802.1Q PVID. - New edge service: Assign a
/32loopback alias onlo. Do not use a container macvlan interface. Add explicit input chain acceptance rules in nftables for authorised LAN segments. - New Docker bridge: Assign an explicit bridge interface name using
com.docker.network.bridge.namein the application’s Compose file. Add the bridge name todockerBridgesinconfiguration.nixso nftables generates intra-bridge and DNAT forwarding rules. - Inter-VLAN pinhole: Add a targeted rule to the
forwardchain. Scope the rule strictly to the sending interface, the source IP, the destination IP, and the required destination port. Never grant broad inter-subnet forwarding.
Operational practices
Section titled “Operational practices”| Practice | Mechanism / Command | Operational Rationale |
|---|---|---|
| Verify post-deploy runtime health | eaves doctor | Asserts all 18 runtime conditions; verifies rulesets, descriptor rings, and socket bindings |
| Inspect dropped IoT traffic | journalctl -k | grep iot-in-drop | Audits unpermitted connection attempts from isolated smart-home peripherals |
| Inspect dropped forwarding attempts | journalctl -k | grep iot-fwd-drop | Verifies that isolated devices are not attempting lateral movement across segments |
| Check trunk interface descriptor rings | ethtool -g <trunk-interface> | Confirms RX and TX descriptor ring depths remain set to 4096 |
| Check physical interface link errors | ethtool -S <trunk-interface> | grep -E 'missed|errors' | Detects physical link degradation, DAC cable faults, or buffer overruns |
| Verify Kea socket bindings | journalctl -u kea-dhcp4-server | grep -i socket | Ensures Kea bound raw sockets to all tagged child VLANs without cross-capture |
Related docs
Section titled “Related docs”- Three NixOS hosts, one deploy interface - The fleet deployment workflow, git force-reset mechanics, and the role of
eaves doctorin deployment automation. - Edge Caddy as a native NixOS service - Native reverse proxy architecture, custom Nix derivations, and port allocations on the edge router.
- knotea: one binary that is both your recursive resolver and your authoritative DNS - The recursive and authoritative DNS server bound to loopback address
10.0.10.5:53. - Isolating a smart-home fleet on an IoT VLAN - Smart home isolation, Home Assistant integration, and NTP time distribution from the hub.
- A Tailscale tailnet across a dual-site homelab - Subnet router configuration and WireGuard routing over
tailscale0. - Tuning a 10GbE link end to end - Buffer ring tuning, interrupt mitigation, and MTU configuration on 10GbE hardware.
References
Section titled “References”-
Internet Systems Consortium, “DHCPv4 Server Configuration: Interface Selection,” Kea Administrator Reference Manual. https://kea.readthedocs.io/en/latest/arm/dhcp4-srv.html#interface-selection ↩
-
The Linux Kernel documentation, “Intel(R) Ethernet Controller X710/XL710 Family Linux Driver,” kernel.org. https://docs.kernel.org/networking/device_drivers/ethernet/intel/i40e.html ↩
-
Tailscale, “ethtool configuration for UDP GRO forwarding,” Tailscale Documentation. https://tailscale.com/kb/1320/performance-best-practices#ethtool-configuration ↩
-
Netfilter Project, “nftables HOWTO documentation,” Netfilter. https://wiki.nftables.org/wiki-nftables/index.php/Main_Page ↩
-
The Linux Kernel documentation, “Bridge Netfilter,” kernel.org. https://docs.kernel.org/networking/netfilter-sysctl.html ↩