diff --git a/README.md b/README.md index 26401ae..ff731d1 100644 --- a/README.md +++ b/README.md @@ -10,7 +10,7 @@ | Host | IP (mgmt) | Role | |---|---|---| -| `xlab-gateway` | 10.253.254.1 | Lab gateway / WAN router. Bond + VLANs (lan254, wan99, mgmt) + WireGuard tunnels with policy routing, NAT/masquerade. Kea DHCP4/6, radvd, fail2ban. | +| `xlab-gateway` | 10.253.254.1 | Lab gateway / WAN router. Bond + VLANs (lan254, wan99, mgmt) + WireGuard tunnels with policy routing, NAT/masquerade. Kea DHCP4/DDNS, IPv6 SLAAC, radvd, fail2ban. | | `skydick` | 10.0.1.1 | Storage server. ZFS data pool with hot spares; Samba (LDAP-backed passdb) + NFS + iSCSI. Jumbo frames (MTU 9200) over bonded 2×40G ConnectX-3 (LACP 802.3ad, `bond40g`). | ## Layout @@ -21,7 +21,7 @@ xlab-gateway/ default.nix # host config (boot, users, packages, smartd) networking.nix # bond/VLAN/WG/nftables/services.resolved - dhcp.nix # Kea DHCP4/6 + DDNS + dhcp.nix # Kea DHCP4/DDNS + IPv6 SLAAC advertisements disko.nix # ZFS root layout hardware-configuration.nix skydick/ @@ -60,10 +60,34 @@ ## Common gotchas -- **DNS**: both hosts route DNS via `10.0.0.1` (mosdns) with a fallback set - in `services.resolved.fallbackDns`. Don't add a co-primary nameserver to - `networking.nameservers` — systemd-resolved load-balances and bypasses - the analytics filter on 10.0.0.1. +- **DNS**: MosDNS at `10.0.0.1` remains the first normal upstream. Do not use + `FallbackDNS` as an availability mechanism: systemd-resolved ignores it + while any global or per-link DNS server is configured, even if that server + is unreachable. Xlab clients query the gateway-local stubs at + `10.253.254.1` / `fd99:23eb:1682:1::1`. The only upstreams are the main + MosDNS addresses `10.0.0.1` and `fd99:23eb:1682::1`; do not add a mainland + resolver as a same-priority secondary because resolved load-balances global + servers and can leak polluted foreign-domain answers to clients. + +- **Skyworks WLAN authentication**: the C9800 live configuration is WPA3-SAE + with PMF required. Keep OKC and band-select disabled on this FlexConnect + WLAN: live traces showed SAE clients being rejected with an invalid PMKID + after a 2.4/5 GHz reassociation. The WLC configuration is not currently + declared in this repository, so preserve the controller backup and audit + live state after upgrades. + +- **Xlab WireGuard control**: networkd owns `wg-to-wgnet` and + `wg-to-skyworks`, so do not use `wg-quick` to remove them. Use + `sudo xlab-wg status all`, `sudo xlab-wg disable `, + `sudo xlab-wg enable `, or + `sudo xlab-wg restart `. Disable state persists across the + health timer and reboot until explicitly enabled. + +- **Xlab campus admission**: WireGuard recovery starts only after the campus + WAN is admitted. For unattended cold boots, register the server address as + `open` with the unit network administrator or provision a dedicated Tunet + credential through agenix. Never embed the account password in a Nix value + or a systemd command line. - **IPv6 RA**: `skydick` runs IPv6 *enabled* (SLAAC, stable token `::d1c0` on `bond40g`). RA-provided DNS is kept out of systemd-resolved with `UseDNS = diff --git a/certs/skyw-ldap-ca.crt b/certs/skyw-ldap-ca.crt new file mode 100644 index 0000000..3d9d034 --- /dev/null +++ b/certs/skyw-ldap-ca.crt @@ -0,0 +1,20 @@ +-----BEGIN CERTIFICATE----- +MIIDTTCCAjWgAwIBAgIUE4FjjtBoCQup0iOUkpZrqTPMLiIwDQYJKoZIhvcNAQEL +BQAwLjEVMBMGA1UEAwwMU2t5dyBMREFQIENBMRUwEwYDVQQKDAxTa3l3IE5ldHdv +cmswHhcNMjYwNzA4MDkwMDQ2WhcNMzYwNzA1MDkwMDQ2WjAuMRUwEwYDVQQDDAxT +a3l3IExEQVAgQ0ExFTATBgNVBAoMDFNreXcgTmV0d29yazCCASIwDQYJKoZIhvcN +AQEBBQADggEPADCCAQoCggEBANcP/i49W8445J/qB30En1Y3bWG2Snn2ahJ6b0hZ +y9T/SHWGt8fpLKRWdsktlCr6F4rdD5zHH8EzgzldPm4nq6VwTR1iOrusOaDn0rIy +huyyuXJYFmvBdXjYdA+u2tufobeX3DoX9MSP1GEtzY02JbzpPUIti/lfS2jbAAKM +Y6r1vBeaYi8A2r7rcf5tuiR06ZPZTLU+xj2+Ge++KrgNsQ+d7WGUTIv4FMxxvHFq +6ZRRfzgZ5vYIV8W9SJUB7XDwWwDxrKAZAjqlu38JO0BkUKnh5R5SXysrctdpNdPl +n/66+qAUTE13GbORT9ZnoKEDsUhPwF+KRtxrsO1O/utpV50CAwEAAaNjMGEwHQYD +VR0OBBYEFE6Yq8HoIemagS5GicjpZ5u0uWezMB8GA1UdIwQYMBaAFE6Yq8HoIema +gS5GicjpZ5u0uWezMA8GA1UdEwEB/wQFMAMBAf8wDgYDVR0PAQH/BAQDAgEGMA0G +CSqGSIb3DQEBCwUAA4IBAQCEgaV1vbYJzc2vM375VffvEZUjSQFufwPBjr6FCcrK +zpDdj2vpHHh3v9GWOnOysKMhFVNhgZQnqhv/CSRrxP0RMuIM5PaBoi6VSIjy9lQG +UN+yxm25MBHJ9yV0V1HSbtWayMuPBcJjrkzsdxbaKQVkZfAa+J2HRtwQm4nFSv5l +zTjuCvmtIZSDGuDDmWKBIph3mP6cZyQ+L9kg+OdWuNpsaaUMBuII0FZ1aWkMT6g6 +DSUPyf5ZMaBNJX/RAc0+MxSs6Bz8SstBhSwMY2rEMYnHqyqsx5Q4pAhUURY4E46H +91897KiOfxEXvmUOnJ2uKDKvqeJu6XqQZrhQRIgGRQaT +-----END CERTIFICATE----- diff --git a/docs/2026-07-21-network-infrastructure-audit.md b/docs/2026-07-21-network-infrastructure-audit.md index 7f639bf..07ed34f 100644 --- a/docs/2026-07-21-network-infrastructure-audit.md +++ b/docs/2026-07-21-network-infrastructure-audit.md @@ -1,52 +1,45 @@ -# Network infrastructure audit - 2026-07-21 +# Network infrastructure audit - 2026-07-21/22 -Scope: the `skyworks` NixOS configurations for `xlab-gateway` and `skydick`, -the live main gateway (`door1` / `10.0.0.1`) and its separate -`skynet-server-gateway` repository, and the source policy in `nix-infra`. +Scope: the `skyworks` NixOS configurations for `xlab-gateway` and +`skydick`, the live main gateway (`door1`, `10.0.0.1:2222`) and its +`skynet-server-gateway` repository, the LDAP deployment in +`skynet-server-web`, and the source policy in `nix-infra`. + +## Operational cautions > [!CAUTION] -> **DO NOT rebuild, deploy, or otherwise activate a new NixOS generation on -> `skydick` until the iSCSI persistence problem below is fixed.** The current -> live LIO target can survive in the kernel while an activation replaces its -> saved configuration with `{}`. A later reboot or target restart can then -> remove the target and its client-visible LUN. +> Do not activate the staged SkyDick generation outside a coordinated storage +> window. The host currently has 33 established NFS/TCP channels from node1, +> node2, and main, including open MySQL data files, plus a logged-in 8 TiB +> iSCSI client at `10.0.200.11`. Dry activation would restart ZFS mount/share, +> NFS, Samba, InfluxDB, and NSS services. -## Evidence labels +## Status labels -- **LIVE + COMMITTED**: observed on the running host and represented by a - commit in that host's canonical repository. -- **REPO + EVALUATED**: present in this `skyworks` worktree and accepted by - Nix evaluation, but not claimed to be deployed. -- **LIVE FINDING**: observed on a running host during this audit. -- **RECOMMENDATION**: proposed remediation; not yet implemented unless stated. +- **LIVE + COMMITTED**: observed on the running host and committed in that + host's canonical repository. +- **LIVE + PROFILED**: active on the running host and selected for the next + boot, but still uncommitted in this worktree. +- **STAGED + BUILT**: represented by a complete NixOS closure, but deliberately + not activated. +- **LIVE FINDING**: observed on a running host and still requires work. -`nix flake check --offline --all-systems --no-build` passed on 2026-07-21 for -both NixOS configurations. The xlab changes below were activated with -deploy-rs, committed as `688ac05`, and passed live route, set, lifecycle, and -counter checks. +## Current state -## Completed policy work +| System | State | Result | +|---|---|---| +| Main gateway | **LIVE + COMMITTED** | Region policy, PotPlayer DNS exception, IPv6 DNS, TSIG rotation, SSH/input hardening, and tunnel CAKE lifecycle are deployed. Repository is clean. | +| Main LDAP stack | **LIVE + COMMITTED** | LDAP and Luminary credentials use Docker secrets; both containers are healthy. Commit `d507c33`. | +| Xlab | **LIVE + COMMITTED, AUTH PENDING** | Both WireGuard paths, automatic recovery, persistent manual disable, CN-direct routing, trusted DNS, DHCP/RA, and restricted public recovery access are live. Polluted AliDNS fallback was removed. Cold-boot campus admission still needs registered-IP status or a dedicated credential. | +| C9800 WLAN | **LIVE + SAVED** | `Skyworks` remains WPA3-SAE with PMF required; OKC and band-select are disabled to avoid the observed FlexConnect SAE PMKID reassociation failure. | +| SkyDick | **STAGED + BUILT** | NixOS 25.11 remains live and healthy. The final committed 26.05 closure is built, not activated because storage clients are active. | -### Alibaba, Taobao, and Tmall region-sensitive routing +## Region-sensitive routing -The source was `nix-infra` commit +The source is `nix-infra` commit `8692bebcfaf59be065d84d064c367cb01dddc691` (`mesh/proxy: force Alibaba -overseas CDN direct for intl-pinned clients`). Packet capture had shown about -21,000 packets per affected Taobao/Tmall session going to four AS45102 Alibaba -Singapore ranges which are not in APNIC-CN. Sending them through a foreign -exit caused Alibaba geo/risk controls to see a foreign source. DNS tagging -cannot reliably catch this path because the apps use HTTPDNS/GSLB. - -Source references: - -- `nix-infra/hosts/door-pek/networking.nix:630-648` documents the capture and - declares the four ranges; `:813-819` places the direct decision before the - per-client international pin. -- `nix-infra/hosts/door-sha/firewall.nix:218-230` declares the same set; - `:272-276` applies it in a comment that mentions Alibaba/JD. The commit adds - no JD-specific range or packet evidence beyond the same Alibaba set. - -The migrated policy is: +overseas CDN direct for intl-pinned clients`). It contains packet-capture +evidence for Taobao/Tmall traffic to four Alibaba AS45102 Singapore ranges: ```text 47.246.0.0/16 @@ -55,219 +48,281 @@ 139.95.0.0/16 ``` -- **LIVE + COMMITTED, main gateway**: commit `5b3cfc5` adds the four ranges to - `data/infra/nftables.netif`. `nft -c`, live set membership, an unmarked - `br-wan` route, and container health were verified. -- **LIVE + COMMITTED, xlab**: throw routes in routing table 1002 make those - non-APNIC-CN destinations fall through to the local campus WAN - (`hosts/xlab-gateway/networking.nix:393-402`). -- **LIVE + COMMITTED, xlab**: APNIC-delegated CN IPv4 and IPv6 destinations - are loaded into nft interval sets, marked `0x10`, and looked up in the main - table before the existing WireGuard rule - (`hosts/xlab-gateway/networking.nix:206-218`, `:412-427`, `:7-143`, - `:544-600`). The updater strictly validates downloaded fields, retains the - last validated file, rejects undersized data, checks the nft transaction - before applying it, and leaves the old `.1` path intact if no set is - available. -- **LIVE + COMMITTED, xlab**: forwarded packets claiming a source outside the - xlab IPv4 `/24` or IPv6 `/64` are dropped before the trusted-LAN accept - (`hosts/xlab-gateway/networking.nix:170-175`, `:197-203`). - -The live xlab service loaded 8,786 IPv4 and 2,039 IPv6 prefixes. Marked CN -route probes selected `wan99.0`; unmarked foreign probes selected -`wg-to-wgnet`; all four Alibaba ranges selected the campus WAN. The nft CN -counter increased on client traffic, the source guards were present, and the -host reported no failed units after activation. A deliberate nftables reload -repopulated both sets from the validated cache in under one second; during -that transient empty-set interval, packets safely follow the old `.1` path. - -These four `/16`s are the observed Alibaba exception set, not a complete or -future-proof inventory of every JD or Alibaba CDN. Keep packet counters and -capture evidence as the criterion for adding another static exception. - -### PotPlayer resolution - -- **LIVE + COMMITTED, main gateway**: commit `5e1d72f` adds `daumcdn.net` to - `data/infra/mosdns/force_proxy.txt`. EasyPrivacy's `tracker_domain.txt` - blanket match had turned the PotPlayer download domain into `NXDOMAIN`. - After restarting mosdns, A and AAAA resolution and an HTTP/2 200 response - for the exact `PotPlayerSetup64.exe` URL were verified. - -### Main-gateway IPv6 DNS - -- **LIVE + COMMITTED, main gateway**: commit - `2898736f8f42afd7bfd17ce25ef2047c27b6fd58` adds MosDNS UDP and TCP listeners - on `[fd99:23eb:1682::1]:53`. The address and both sockets, IPv6 UDP local and - Taobao resolution, IPv6 TCP DNSSEC with the AD bit, IPv4 regression, and - container health were verified. Only MosDNS was restarted. - -## P0 - fix before routine operations - -### 1. Skydick iSCSI state is erased by the declarative default - -**Verified configuration:** `hosts/skydick/datapool.nix:712-713` only enables -`services.target`; it does not set `services.target.config`. In the pinned -NixOS module, that option defaults to `{}` and manages -`/etc/target/saveconfig.json` as mode `0600`. The documented imperative -`targetcli ... saveconfig` workflow (`hosts/skydick/DATAPOOL.md:516-558`) -therefore writes a file that the next NixOS activation can replace with `{}`. - -**Impact:** an already loaded kernel target can make this look healthy until a -reboot, `iscsi-target` restart, or LIO clear. The restored target will then be -empty. - -**Required remediation:** first back up the live JSON and record `targetcli -ls`; then make restoration declarative or copy a persistent, encrypted runtime -file into place before `iscsi-target` starts. Do not put CHAP secrets directly -in a Nix expression because they would enter the world-readable Nix store. Add -an activation assertion that refuses to replace a non-empty live target with -an empty definition, then test restore and client login before allowing a -rebuild. - -### 2. The NFS pseudo-root can expose unlisted mounted datasets - -**Verified configuration:** `/srv` is exported to all of `10.0.0.0/16` with -`crossmnt` (`hosts/skydick/datapool.nix:479-480`). `crossmnt` implicitly exports -mounted child filesystems with the parent's options. This can expose unlisted -children such as monitoring, Time Machine, private, or future ZFS datasets and -can undermine narrower child exports. - -The surrounding ACLs are also broader than their prose suggests: - -- all LAN hosts can write `/srv/media` as UID 900 (`:482-485`); -- all LAN hosts map to the owner of `ye-lw21` and `zhuyz24` (`:487-500`); -- all LAN hosts receive `no_root_squash` on backup and VM trees (`:502-505`). - -This contradicts "Only you can access your tree" in -`hosts/skydick/DATAPOOL.md:293-304`. Remove `crossmnt` from the `fsid=0` -pseudo-root, export every intended child explicitly, and reduce client matches -to exact IPv4 `/32` and IPv6 `/128` identities. Use SMB or NFS Kerberos -(`sec=krb5p`) where user authentication is required. - -## P1 - high-priority reliability and security +The commit contains no JD-specific prefix or packet evidence. JD appears only +in a comment, so no unsupported JD range was invented during migration. ### Main gateway -1. **CAKE/IFB is not attached after WireGuard recreation (LIVE FINDING).** The - related units report healthy while the qdisc/filter is absent, so tunnel - fairness can silently disappear. Bind QoS lifecycle to the WireGuard device - recreation path, make setup idempotent, and alert on `tc` state rather than - unit state alone. -2. **Fifteen legacy `df99` peers bypass IPv6 policy; several peers appear - stale (LIVE FINDING).** Migrate active peers to classified source prefixes, - remove expired peers, and default-deny unclassified IPv6 sources. -3. **Public SSH permits passwords and forwarding, and its limiter is global - (LIVE FINDING).** A remote actor can consume the shared allowance and lock - out operators. Prefer VPN/source-restricted key-only access, disable - forwarding unless explicitly required, and use per-source limits. -4. **WAN IPv6 lacks equivalent source anti-spoofing (LIVE FINDING).** Add the - IPv6 reserved, ULA, link-local, multicast, and owned-prefix source checks at - the raw/prerouting boundary. +- **LIVE + COMMITTED**: `5b3cfc5` adds all four ranges to the direct-routing + nftables set before foreign-exit selection. +- **LIVE + COMMITTED**: `5e1d72f` forces `daumcdn.net` through the proxy path. + EasyPrivacy had classified the PotPlayer download CDN as a tracker, causing + MosDNS to return `NXDOMAIN`. A/AAAA lookup and the exact installer URL were + verified after deployment. +- **LIVE + COMMITTED**: `2898736` serves MosDNS over UDP/TCP on the advertised + gateway ULA as well as IPv4. +- **LIVE + COMMITTED**: `3b957cd` rotates the DDNS TSIG secret, retires the old + key, hardens gateway SSH/input exposure, scopes Avahi, and repairs the + CAKE/IFB lifecycle. `wg-sgp-qos.service` and + `wg-sgp-autorate.service` are currently active, and CAKE is attached to + `ifb-sgp` and `wg-outbound-sgp`. -### Xlab gateway +### Xlab client policy -1. **The TSIG DDNS secret is committed in plaintext.** It is embedded at - `hosts/xlab-gateway/dhcp.nix:172-178`. Treat it as compromised: rotate the - key on xlab and the authoritative server, move it to agenix, and render a - root-only runtime Kea fragment. -2. **The new source-spoof guards cover forwarding but not router input.** The - input chain accepts every service from the entire LAN, management VLAN, and - both WireGuard interfaces (`hosts/xlab-gateway/networking.nix:178-194`). Add - the source validation before those accepts and replace full-interface trust - with explicit DHCP, DNS/DDNS, SSH, ICMP, and WireGuard service rules. -3. **CN-direct routing needs monitoring and a controlled failover test.** The - normal-path route and counter checks passed after deployment. Add alerts on - APNIC cache age, set size, and the presence of both policy rules. In a - maintenance window, confirm the intended table-1002 fallback when the local - WAN default is unavailable. Registry allocation is a routing heuristic, - not exact geolocation. +The deployed xlab policy uses two nft interval-set layers: -### Skydick +- APNIC-delegated CN IPv4 and IPv6 prefixes are downloaded over HTTPS, + schema/size/count checked, nft syntax checked, and transactionally loaded. + The last validated cache remains usable if refresh fails. The current live + sets contain 8,786 IPv4 and 2,039 IPv6 prefixes. +- Static regional sets contain the four Alibaba ranges, six Tsinghua/CERNET + IPv4 ranges, and three CERNET IPv6 ranges. -1. **Storage services fail open onto placeholder directories.** Tmpfiles - creates the `/srv` tree (`hosts/skydick/datapool.nix:298-338`), while - `dick-zfs-properties` exits successfully when pool `dick` is missing and - only orders itself before NFS/Samba (`:340-368`). Evaluated NFS and Samba - units have no hard mount requirement. Gate NFS, Samba, InfluxDB, and iSCSI - behind a unit that asserts the expected ZFS datasets and mount identities. -2. **LDAP binds and credentials cross the LAN in plaintext.** NSS uses - `ldap://` with TLS disabled (`hosts/skydick/default.nix:326-340`); Samba's - LDAP passdb also disables SSL and strong authentication - (`hosts/skydick/datapool.nix:524-540`). Move both to validated StartTLS or - LDAPS before rotating the bind and admin credentials. -3. **Service exposure is host-wide rather than source/interface scoped.** NFS, - SMB, iSCSI, InfluxDB, and node-exporter ports are globally allowed - (`hosts/skydick/datapool.nix:715-731`, `modules/influxdb.nix:33-62`, - `modules/monitoring.nix:166-173`, `:296-297`). Restrict storage protocols to - exact storage clients/VLANs and monitoring ports to the monitoring host. - Enable TLS for InfluxDB and replace the shared fleet-wide operator token - documented in `hosts/skydick/README.md:20-42` with per-writer tokens and a - read-only Grafana token. -4. **NFS/RDMA is only partially wired.** Live inspection found the mlx4 RDMA - devices active and `/proc/fs/nfsd/portlist` containing `rdma 20049` twice, - but no active client RDMA mount was confirmed. The listener is declared at - `hosts/skydick/datapool.nix:441-469`; port 20049 is absent from the firewall - at `:715-731`, and NFS/RDMA clients require the export's `insecure` option - because they do not use a reserved source port. Scope TCP/UDP 20049 and that - export option to exact clients, reconcile the duplicate writer, and prove - `proto=rdma` from a client mount rather than inferring it from the server - listener. -5. **The pinned NixOS release is out of security support.** `flake.nix:5` - selects `nixos-25.11`; both hosts correctly keep `system.stateVersion = - "25.11"` (`hosts/xlab-gateway/default.nix:107`, - `hosts/skydick/default.nix:401`). NixOS 25.11 reached EOL after 2026-06-30; - [NixOS 26.05](https://nixos.org/blog/announcements/2026/nixos-2605/) is - supported through 2026-12-31. Upgrade the input to 26.05 without changing - `system.stateVersion`, and specifically rebuild/test the exact-source Samba - overlay at `hosts/skydick/datapool.nix:252-279`. +Matching LAN packets receive mark `0x10`. Rule priority 90 first looks in the +main table, so they leave through `wan99.0` while its default route exists. If +that default disappears, lookup continues to the existing source rule for +table 1002 and exits through `wg-to-wgnet`. The earlier static `throw` design +did not have this fallback and has been removed. -## P2 - cleanup and optimization +An isolated network-namespace test loaded the rendered nft rules and proved: -- **Main gateway:** bind Avahi only to intended LAN interfaces and remove stale - discovery ports; deduplicate policy rules; pin container images by digest; - enable Docker live-restore where compatible; and add a second, health-checked - foreign relay instead of retaining a single exit failure domain. -- **Xlab bonding:** `balance-xor` has no explicit transmit hash policy - (`hosts/xlab-gateway/networking.nix:266-275`). Routed traffic can collapse - onto one member with a layer-2 hash. Validate the switch LAG and use a - layer3+4 hash (or 802.3ad with matching switch configuration), then measure - per-member distribution. -- **Xlab CAPWAP workaround:** `capwap-df-clear` deletes and recreates the whole - ingress qdisc (`hosts/xlab-gateway/default.nix:87-104`). Move to stable - `clsact`/owned filter handles so restarting the unit cannot erase unrelated - ingress policy. -- **Skydick NIC tuning:** the active bond uses `enp130s0` and `enp130s0d1` - (`hosts/skydick/default.nix:79-100`), but ring tuning still targets the - retired `enp4s0f0np0` and `enp4s0f1np1` (`:191-194`). Retarget it and record - before/after drop, latency, and IRQ-affinity metrics. -- **Skydick discovery:** `samba-wsdd` is enabled with `openFirewall = false` - (`hosts/skydick/datapool.nix:707-710`), while UDP 3702/TCP 5357 are absent - from the firewall. Either open them only on the client LAN or disable WSD. -- **Skydick RoCE QoS:** configuration explicitly provides DSCP marking without - end-to-end PFC (`hosts/skydick/datapool.nix:404-410`). Complete and test PFC, - ECN/congestion control, MTU, and watchdog configuration at every hop before - calling the fabric lossless. -- **Privilege boundary:** `ye-lw21` is a Nix `trusted-user` - (`modules/common.nix:7-10`) even though the account is documented as - password-sudo only (`modules/users.nix:23-27`). Nix trusted users are - effectively root-capable; remove the entry unless that is intentional. -- **Secret rotation:** add restart/reload triggers for the Samba LDAP seed and - Telegraf token (`hosts/skydick/datapool.nix:371-389`, - `modules/monitoring.nix:88-99`) so rotating an agenix secret changes the - running credential without manual intervention. +```text +regional_direct_v4=10 regional_direct_v6=3 +healthy: 47.246.1.1 -> wan99.0 +failed WAN: 47.246.1.1 -> wg-to-wgnet table 1002 +``` -## Execution order +## Xlab deployed closure -1. Back up and make skydick iSCSI state activation-safe. Only then allow a - skydick rebuild. -2. Remove the NFS pseudo-root `crossmnt` exposure and narrow NFS client ACLs. -3. Add a client-path monitor for the repaired main-gateway IPv6 DNS listener, - then repair CAKE/IFB lifecycle and close the `df99` bypasses. -4. Rotate the exposed xlab TSIG key and harden router input/source validation. -5. Monitor the deployed xlab CN-direct policy and run its WAN-failure test in - a maintenance window; retain packet capture as the basis for any future - Alibaba or JD exception ranges. -6. Restrict skydick service exposure, enable LDAP/InfluxDB TLS, and add ZFS - mount guards. -7. Upgrade the NixOS input to 26.05 and run host-specific build and service - regression tests before deployment. +Active system and persistent system profile as of 2026-07-22: + +```text +/nix/store/1b50bcwxk7dcs5rhgny4p596shc5aabi-nixos-system-xlab-gateway-26.05.20260719.fd14620 +``` + +Included changes: + +- NixOS 26.05 while preserving `system.stateVersion = "25.11"`. +- APNIC-CN plus static regional direct routing with WAN-loss fallback. +- Gateway-local DNS listeners at `10.253.254.1` and + `fd99:23eb:1682:1::1`; DHCP and RA advertise those addresses. Their only + upstreams are main MosDNS at `10.0.0.1` and `fd99:23eb:1682::1`. + `FallbackDNS` is empty so resolved cannot select an unfiltered mainland + resolver for foreign domains. Both WireGuard endpoints are numeric, so + tunnel bootstrap has no DNS dependency. +- Explicit input/source ACLs, fixed WireGuard listen ports, and WAN SSH only + from main (`166.111.17.108`) or server3 (`166.111.17.81`). +- Explicit `wireguard` module loading and a handshake-aware recovery timer. + Networkd configures the netdevs with manual activation; `xlab-wg` waits for + the campus route, recreates missing/incomplete interfaces, raises enabled + tunnels, and cycles only peers whose latest handshake is stale. +- Persistent administrative disable markers let `xlab-wg disable` keep either + tunnel down across health checks and reboots. This replaces `wg-quick` for + interfaces owned by systemd-networkd. +- A numeric server3 endpoint removes the old boot-time DNS dependency. Explicit + WAN UDP rules allow either peer to initiate recovery, and WAN SSH is restricted + to main and server3 for break-glass access. +- Rotated agenix-backed Kea TSIG key, DHCPv4 plus SLAAC-only IPv6, symmetric + MSS clamping, safer CAPWAP filters, and layer3+4 bond hashing. +- Periodic route/set/cache health checks. + +Rendered `resolved.conf`, Kea DHCP4, radvd, networkd, and nftables files were +inspected. `radvd --configtest`, nftables `--check`, both generated health +scripts with `bash -n`, and the complete NixOS build passed. Kea parsed the +configuration through socket selection; its only test-host error was the +expected absence of xlab's `bond.lan254` on the remote build machine. + +### Recovered outage (2026-07-22) + +The failed boot journal disproved the initial missing-netdev hypothesis. Both +WireGuard netdevs and their agenix secrets existed, but networkd raised the +tunnels before the WAN route, DNS, and campus admission were usable. The +server3 endpoint hostname then failed resolution continuously for roughly 12 +hours. The old firewall did not admit inbound UDP 46961/51998, and the old +generation had no handshake-aware recovery task, so peer-initiated traffic +could not break the bootstrap failure. + +The campus authentication endpoint reports that the current admission session +started at 15:28:44 CST. `wg-to-wgnet` was manually raised at 15:29:37 and +handshook immediately. Xlab has no Tunet/GoAuthing unit or stored campus +credential, so the missing cold-boot admission was a necessary part of the +outage, not merely a WireGuard race. The new readiness service retries and +recovers automatically once admission exists, but full cold-boot automation +still requires either registered/open IP status or an agenix-backed dedicated +campus credential. + +The preserved evidence remains on xlab at: + +```text +/var/tmp/xlab-previous-boot-full-20260722.log +/var/tmp/xlab-previous-boot-network-20260722.log +/var/tmp/xlab-network-boot-20260722.log +``` + +The new closure was activated without rebooting and tested from both the normal +route and restricted public break-glass path. Activation restarted networkd; +the readiness service detected stale handshakes and recovered both tunnels in +about three seconds each. Further fault injection proved: + +- disabling `wg-to-skyworks` persists while the health service runs, then an + explicit enable restores a current handshake; +- deleting the complete `wg-to-skyworks` netdev causes it to be recreated, + configured, raised, and handshaken; +- forcing `wg-to-wgnet` down from the public recovery session causes automatic + recovery and restores `10.253.254.1` reachability. + +The deployed CN-direct health check passed with 8,786 IPv4 and 2,039 IPv6 APNIC +prefixes plus 10/3 static regional prefixes. Marked `223.5.5.5` and +`47.246.1.1` route through `wan99.0`; unmarked `1.1.1.1` routes through +`wg-to-wgnet`. No systemd units are failed. + +### Recovered DNS pollution (2026-07-22) + +Xlab's resolved configuration listed main MosDNS and AliDNS as peer global +servers. systemd-resolved does not treat that list as ordered failover; it had +selected `223.5.5.5`, which returned non-Google addresses for +`www.google.com` and similarly polluted YouTube names. This explained the +reported split symptom: `google.com` resolved and pinged, but its HTTPS redirect +to `www.google.com` failed. Path-MTU probes, symmetric MSS clamps, and real +client TCP flows were healthy. + +The live and profiled configuration now uses only the two main MosDNS tunnel +addresses. After removing the temporary runtime override and restarting +resolved, repeated `www.google.com` queries returned Google A/AAAA records and +an interface-bound HTTPS request returned HTTP 200. No systemd unit is failed. + +## Skyworks WLAN authentication recovery + +The controller is a C9800-CL running IOS XE 17.15.5; the xlab AP is an +AIR-AP3802I-H in FlexConnect mode. Authoritative post-change `show wlan id 1` +output confirms that `Skyworks` is WPA3-SAE-only with CCMP, PMF required, +H2E plus HNP, and FT disabled. The authentication mode itself was not changed +during this recovery. + +Always-on WLC trace provided a concrete interoperability failure for a +MacBook randomized address. It completed SAE on 5 GHz, received DHCP, moved to +2.4 GHz, then reassociated to 5 GHz and was deleted with: + +```text +SAE PMKID matching failed in roam case +Invalid PMKID +IE_VALIDATION_FAILURE type 53 +``` + +Another failing randomized client associated on 5 GHz but never completed the +SAE exchange and timed out in L2AUTH. This matches the FlexConnect/WPA3-SAE/OKC +failure family in +[Cisco CSCwp20425](https://www.cisco.com/c/en/us/td/docs/wireless/controller/9800/17-15/release-notes/rn-17-15-9800.html); +Cisco's documented workaround is to disable OKC. Enabled band-select increased +the chance of cross-band reassociation. The controller is already on Cisco's +[currently recommended 17.15.5 release](https://www.cisco.com/c/en/us/support/docs/wireless/catalyst-9800-series-wireless-controllers/214749-tac-recommended-ios-xe-builds-for-wirele.html), +so there is no justified in-train firmware upgrade to substitute for the +verified workaround. + +The WLC running configuration was backed up to +`bootflash:pre-skyworks-auth-fix-20260722.cfg`. OKC and band-select were then +disabled on WLAN 1, the WLAN was re-enabled, all three existing Skyworks SAE +clients returned to `Run`, and `write memory` completed. The previously +failing MacBook then rejoined on 5 GHz using SAE, remained associated well past +its former five-second failure point, obtained `10.253.254.103`, and exchanged +traffic without decrypt or policy errors. Six Skyworks SAE clients were in +`Run` concurrently with no excluded clients. WPA2 transition mode remains an +available compatibility fallback, but was not needed and would reduce the +current security level. + +## SkyDick staged closure + +Final build: + +```text +/nix/store/x3b5rnxpp37icqjr0ja91l2mhhm73igz-nixos-system-skydick-26.05.20260719.fd14620 +``` + +Included changes: + +- NixOS 26.05 with the exact-source Samba overlay rebuilt; state version stays + at 25.11. +- A fail-closed ZFS mount-identity gate for NFS, Samba, InfluxDB, and iSCSI. +- Persistent root-only LIO state at `/var/lib/target`, with `/etc/target` as a + directory symlink so `targetcli` atomic saves remain persistent. +- An iSCSI guard requiring block backstores, existing devices, explicit + non-wildcard portals, IQN ACLs, CHAP credentials, and LUN mappings. It also + prevents `iscsi-target.service` from stopping during a live switch. +- Explicit NFS children instead of pseudo-root `crossmnt`; exact media/backup + clients where verified; NFSv4 pseudo-root coverage for the declared IPv6 + clients. +- Source-restricted nftables rules for SSH, NFS, SMB, iSCSI, InfluxDB, and + node exporter. NetBIOS and unused WSD discovery are disabled. +- Certificate-verified StartTLS for NSS and Samba LDAP, secret restart + triggers, corrected 40 GbE ring tuning, and removal of the unnecessary Nix + trusted-user grant. +- Only main MosDNS at `10.0.0.1` and `fd99:23eb:1682::1` are global resolver + candidates. AliDNS and direct public resolvers were removed so the staged + closure cannot leak polluted answers or bypass main's split-DNS policy. + +The live iSCSI configuration was backed up again at +`/root/backups/target/20260721T204151Z`. Serializing the live kernel state was +byte-identical before and after (`babc4f7a...59650c246`), and the CHAP-aware +guard passes it. `targetcli sessions detail` saying `NOT AUTHENTICATED` refers +to optional mutual target authentication; initiator CHAP fields are present. + +Do not activate this closure until node1/node2 MySQL and all other NFS users +are stopped or unmounted and the iSCSI initiator is logged out or explicitly +accepted as maintenance risk. + +## LDAP runtime hardening + +- Root-only `slapcat` backups are under + `/root/backups/ldap/20260721T200920Z` on main. +- Commit `d507c33` moves LDAP admin/config/readonly and Luminary bind + credentials from container environment variables to Docker secrets. +- LDAP and Luminary were recreated and are healthy. Container inspection + confirmed no direct password variables and the expected secret mounts. +- StartTLS with hostname and the staged SkyDick CA bundle succeeds; NSS group + lookup and Samba `pdbedit` both pass. + +The live SkyDick generation still uses plaintext LDAP until its maintenance +deployment. This is a deployment-state issue, not a missing source fix. + +## Remaining findings + +1. **P0, xlab campus admission:** the active session was established manually + after boot. Prefer registering this server IP as `open` with the unit network + administrator. Otherwise provision a dedicated Tunet credential through + agenix and run a maintained authentication client; do not put the password + in the Nix store or service command line. +2. **P0, SkyDick maintenance:** deploy only after coordinated NFS and iSCSI + quiescence, then validate mounts, exports, LDAP TLS, Samba, InfluxDB, and + target restore before clients reconnect. +3. **P1, NFS identity:** `ye-lw21`, `zhuyz24`, and `/srv/system/vm` retain + broad subnet ACLs because the exact client contract is not yet known. + Replace them with `/32` and `/128` entries after inventory. Use Kerberos or + SMB when per-user authentication, rather than host identity, is required. +4. **P1, NFS/RDMA:** the live port list contains `rdma 20049` twice, but all + observed sessions are TCP/2049 and no RDMA client was proven. Port 20049 is + intentionally not opened by the staged firewall. Remove the unused listener + or validate a real `proto=rdma` client and then scope the firewall/export. +5. **P1, monitoring secrets:** InfluxDB remains HTTP and the documented token + model is shared more broadly than necessary. Move to TLS and separate + writer/read-only tokens. +6. **P2, LDAP indexes:** live slapd logs report unindexed searches for `cn`, + `sambaDomainName`, and `sambaSID`. The database is currently small; add + indexes with a backed-up, scheduled slapindex/restart rather than an online + ad hoc change. +7. **P2, main peers:** several legacy `df99` IPv6 peers and peers without + recent handshakes remain. Inventory ownership before removing or migrating + them; do not infer abandonment from one snapshot. +8. **P2, WLC source of truth:** the C9800 WLAN policy is live and saved, but it + is not declared in this repository. Export a sanitized controller baseline + or add an approved configuration-management path so OKC/band-select cannot + silently return after replacement or template changes. + +## Remaining execution order + +1. Finish xlab cold-boot campus admission through registered-IP status or an + agenix-backed dedicated credential. Retain the preserved journal and use + `xlab-wg` rather than `wg-quick` for its networkd-owned tunnels. +2. Schedule SkyDick storage downtime, stop/unmount NFS clients, log out iSCSI, + activate the staged closure, and perform the storage/LDAP/firewall checks. +3. Narrow the remaining NFS ACLs and address RDMA, InfluxDB TLS/tokens, LDAP + indexes, and legacy main-gateway peers as separately reversible changes. diff --git a/flake.lock b/flake.lock index 4789329..45eac8b 100644 --- a/flake.lock +++ b/flake.lock @@ -126,16 +126,16 @@ }, "nixpkgs": { "locked": { - "lastModified": 1773524153, - "narHash": "sha256-Jms57zzlFf64ayKzzBWSE2SGvJmK+NGt8Gli71d9kmY=", + "lastModified": 1784432872, + "narHash": "sha256-n3gKTBIV4ZA5VQpUakffBe3KGu4+mhPoA34rrqS0GkA=", "owner": "NixOS", "repo": "nixpkgs", - "rev": "e9f278faa1d0c2fc835bd331d4666b59b505a410", + "rev": "fd1462031fdee08f65fd0b4c6b64e22239a77870", "type": "github" }, "original": { "owner": "NixOS", - "ref": "nixos-25.11", + "ref": "nixos-26.05", "repo": "nixpkgs", "type": "github" } diff --git a/flake.nix b/flake.nix index b9b3f61..3616d18 100644 --- a/flake.nix +++ b/flake.nix @@ -2,7 +2,7 @@ description = "Skyworks infrastructure"; inputs = { - nixpkgs.url = "github:NixOS/nixpkgs/nixos-25.11"; + nixpkgs.url = "github:NixOS/nixpkgs/nixos-26.05"; disko = { url = "github:nix-community/disko"; diff --git a/hosts/skydick/DATAPOOL.md b/hosts/skydick/DATAPOOL.md index 0b59e19..2d4932c 100644 --- a/hosts/skydick/DATAPOOL.md +++ b/hosts/skydick/DATAPOOL.md @@ -45,8 +45,8 @@ ## Identity and authentication -- `skydick` resolves POSIX users and groups from LDAP at `ldap://10.0.0.1/`, base - `dc=skyw,dc=top` +- `skydick` resolves POSIX users and groups using certificate-verified StartTLS at + `ldap://ldap.skyw.top/`, base `dc=skyw,dc=top` - SMB now uses Samba's LDAP passdb (`ldapsam`) against the same directory tree - On this standalone server, the Samba account-domain object is expected to be `sambaDomainName=SKYDICK`, matching the NetBIOS name, not the browse workgroup `WORKGROUP` @@ -166,7 +166,9 @@ ## Connecting via NFS (Linux) -NFS uses NFSv4 with a pseudo-root at `/srv`. Mount paths omit `/srv`. +NFS uses NFSv4 with a read-only pseudo-root at `/srv`. Mount paths omit `/srv`. +Every intended child dataset is exported explicitly; the pseudo-root does not use `crossmnt`, so a +new ZFS dataset is not reachable until its export is reviewed and added. ### One-off mount @@ -409,13 +411,18 @@ "d /srv/users//vm/files 0750 -" ``` -Add to `services.nfs.server.exports`: +Add the parent and each mounted child to `services.nfs.server.exports`. Use exact client addresses; +an owner-mapped `all_squash` export to the whole LAN lets any LAN host act as that user. ``` -/srv/users/ 10.0.0.0/16(rw,sync,no_subtree_check,all_squash,anonuid=,anongid=) +/srv/users/ (rw,sync,no_subtree_check,all_squash,anonuid=,anongid=) +/srv/users//files (rw,sync,nohide,no_subtree_check,all_squash,anonuid=,anongid=) +/srv/users//bt-state (rw,sync,nohide,no_subtree_check,all_squash,anonuid=,anongid=) +/srv/users//vm (rw,sync,nohide,no_subtree_check,all_squash,anonuid=,anongid=) ``` -Replace `` and `` with the LDAP-backed numeric IDs from `getent passwd`. +Replace ``, ``, and `` with the assigned host address and LDAP-backed numeric +IDs. Add a second client term only when another stable address is an intentional owner endpoint. Example: the user previously called `ylw` in local NixOS config is now canonicalized to `ye-lw21` everywhere, so the per-user share path is `/srv/users/ye-lw21`. @@ -516,10 +523,12 @@ ### Provisioning an iSCSI block volume (admin) iSCSI hands a client a raw block device (a ZFS zvol) instead of a shared filesystem. Unlike SMB/NFS, -**iSCSI targets are imperative state**: `services.target.enable = true` in the Nix config only installs -the restore unit `iscsi-target.service` (note: *not* `target.service`), which runs `targetctl restore` -from `/etc/target/saveconfig.json` at boot. You build the target with `targetcli` and persist it with -`saveconfig`; it is **not** declared in the repo. +**iSCSI targets are imperative, secret-bearing state**. The root-only state lives in +`/var/lib/target`; `/etc/target` is a directory symlink so targetcli's atomic `saveconfig` rename stays +on persistent storage. `iscsi-target.service` (note: *not* `target.service`) restores it at boot. +The Nix configuration validates that the root-only JSON has a target, block backstore, non-wildcard +portal, explicit IQN ACL, CHAP credentials, and LUN mapping, and that every backing device exists. It +never copies CHAP material into the world-readable Nix store. An iSCSI "account" = a target LUN + a per-initiator **ACL** (allowlists the client's initiator IQN) + **CHAP** (username/password). Both together gate access. @@ -536,10 +545,16 @@ Steps (run on skydick as root): ```bash -# 1. Create the backing zvol. -s = thin/sparse. 16K volblocksize per pool convention. +# 1. Back up the existing secret-bearing state before every target change. +install -d -m 0700 /root/backups/target +stamp=$(date -u +%Y%m%dT%H%M%SZ) +cp --preserve=all /etc/target/saveconfig.json /root/backups/target/saveconfig.json.$stamp +targetcli ls > /root/backups/target/targetcli-ls.$stamp.txt + +# 2. Create the backing zvol. -s = thin/sparse. 16K volblocksize per pool convention. zfs create -s -V 8T -o volblocksize=16K dick/system/vm/ -# 2. Build the LIO target. Generate a CHAP secret (openssl is not on the box): +# 3. Build the LIO target. Generate a CHAP secret (openssl is not on the box): CHAP=$(tr -dc 'A-Za-z0-9' @@ -556,6 +571,9 @@ /iscsi/$TGT/tpg1/portals create 10.0.1.1 3260 saveconfig EOF + +# 4. Re-run the non-mutating persistence/backstore guard. +systemctl restart iscsi-target-config-ready.service ``` Notes / gotchas: @@ -565,7 +583,13 @@ `::0 3260` *before* creating a specific `10.0.1.1 3260` portal, or the specific bind fails with `EINVAL` (the wildcard already holds the port). - `generate_node_acls=0` + `authentication=1` means only the listed IQN, with the correct CHAP, can - attach. Firewall port 3260 is open with no source restriction, so CHAP + the IQN ACL are the guard. + attach. The host firewall additionally admits TCP 3260 only from the declared initiator at + `10.0.200.11`; update that source ACL deliberately when provisioning another initiator. +- In `targetcli sessions detail`, the parenthetical `NOT AUTHENTICATED` label refers to optional + *mutual* CHAP (`authenticate_target`), not whether the initiator supplied its CHAP credential. The + persistence guard checks the initiator-side `chap_userid` and `chap_password` fields directly. +- `nixos-rebuild switch` never restarts `iscsi-target.service`; perform a target restart only in an + explicit maintenance window because its stop action clears the live kernel target. - Verify: `targetcli ls /iscsi` and `ss -ltnp | grep 3260`. Client side (open-iscsi, run as root on the client): diff --git a/hosts/skydick/datapool.nix b/hosts/skydick/datapool.nix index cdc06c9..602622e 100644 --- a/hosts/skydick/datapool.nix +++ b/hosts/skydick/datapool.nix @@ -230,8 +230,107 @@ # final Unix authorization step. For stronger NFS isolation: use sec=krb5 or # tighter per-client IP restrictions. -{ config, pkgs, ... }: +{ config, lib, pkgs, ... }: +let + storageMounts = [ + { dataset = "dick/public"; mountPoint = "/srv/public"; } + { dataset = "dick/public/datasets"; mountPoint = "/srv/public/datasets"; } + { dataset = "dick/public/models"; mountPoint = "/srv/public/models"; } + { dataset = "dick/public/sdk"; mountPoint = "/srv/public/sdk"; } + { dataset = "dick/media"; mountPoint = "/srv/media"; } + { dataset = "dick/system/backup"; mountPoint = "/srv/system/backup"; } + { dataset = "dick/system/influxdb"; mountPoint = "/srv/system/influxdb"; } + { dataset = "dick/system/vm"; mountPoint = "/srv/system/vm"; } + { dataset = "dick/templates/vm"; mountPoint = "/srv/templates/vm"; } + { dataset = "dick/users/ldx"; mountPoint = "/srv/users/ldx"; } + { dataset = "dick/users/ldx/files"; mountPoint = "/srv/users/ldx/files"; } + { dataset = "dick/users/ldx/bt-state"; mountPoint = "/srv/users/ldx/bt-state"; } + { dataset = "dick/users/ldx/vm"; mountPoint = "/srv/users/ldx/vm"; } + { dataset = "dick/users/ldx/timemachine"; mountPoint = "/srv/users/ldx/timemachine"; } + { dataset = "dick/users/ye-lw21"; mountPoint = "/srv/users/ye-lw21"; } + { dataset = "dick/users/ye-lw21/files"; mountPoint = "/srv/users/ye-lw21/files"; } + { dataset = "dick/users/ye-lw21/bt-state"; mountPoint = "/srv/users/ye-lw21/bt-state"; } + { dataset = "dick/users/ye-lw21/vm"; mountPoint = "/srv/users/ye-lw21/vm"; } + { dataset = "dick/users/zhuyz24"; mountPoint = "/srv/users/zhuyz24"; } + { dataset = "dick/users/zhuyz24/files"; mountPoint = "/srv/users/zhuyz24/files"; } + { dataset = "dick/users/zhuyz24/bt-state"; mountPoint = "/srv/users/zhuyz24/bt-state"; } + { dataset = "dick/users/zhuyz24/vm"; mountPoint = "/srv/users/zhuyz24/vm"; } + ]; + + storageMountPoints = map (mount: mount.mountPoint) storageMounts; + storageServices = [ + "nfs-server.service" + "samba-smbd.service" + "influxdb2.service" + "iscsi-target.service" + ]; + + iscsiConfigGuard = pkgs.writeShellApplication { + name = "skydick-iscsi-config-guard"; + runtimeInputs = [ pkgs.coreutils pkgs.jq ]; + text = '' + set -euo pipefail + check_devices=true + if [[ "''${1:-}" == "--no-devices" ]]; then + check_devices=false + shift + fi + config="''${1:-/etc/target/saveconfig.json}" + + fail() { + echo "iSCSI configuration guard: $*" >&2 + exit 1 + } + + [[ -f "$config" ]] || fail "$config is missing or is not a regular file" + [[ -s "$config" ]] || fail "$config is empty" + [[ "$(stat -c '%u:%g' "$config")" == "0:0" ]] || fail "$config must be owned by root:root" + [[ "$(stat -c '%a' "$config")" == "600" ]] || fail "$config must have mode 0600" + + jq -e ' + type == "object" and + (.storage_objects | type == "array" and length > 0) and + all(.storage_objects[]; + .plugin == "block" and + (.dev | type == "string" and length > 0) + ) and + (.targets | type == "array" and length > 0) and + all(.targets[]; + (.tpgs | type == "array" and length > 0) and + all(.tpgs[]; + .attributes.authentication == 1 and + .attributes.generate_node_acls == 0 and + .attributes.cache_dynamic_acls == 0 and + (.portals | type == "array" and length > 0) and + all(.portals[]; + .port == 3260 and + .ip_address != "0.0.0.0" and + .ip_address != "::0" + ) and + (.luns | type == "array" and length > 0) and + (.node_acls | type == "array" and length > 0) and + all(.node_acls[]; + (.node_wwn | type == "string" and length > 0) and + (.chap_userid | type == "string" and length > 0) and + (.chap_password | type == "string" and length >= 12) and + (.mapped_luns | type == "array" and length > 0) + ) + ) + ) + ' "$config" >/dev/null || fail "$config lacks a secured block target, portal, ACL, CHAP credential, or LUN mapping" + + if $check_devices; then + device_count=0 + while IFS= read -r device; do + [[ -b "$device" ]] || fail "configured backstore $device is not a block device" + ((device_count += 1)) + done < <(jq -r '.. | objects | select(.plugin? == "block") | .dev? // empty' "$config") + ((device_count > 0)) || fail "$config has no usable block backstore" + fi + ''; + }; +in { # Build sambaFull with Spotlight/tracker support. # Fixes for tinysparql 3.x (tracker-sparql-3.0) compatibility: @@ -337,12 +436,55 @@ ]; + # Fail closed if ZFS did not mount the real datasets. tmpfiles deliberately + # creates these paths, so path existence alone cannot distinguish a dataset + # from an empty placeholder on the root filesystem. + systemd.services.dick-storage-ready = { + description = "Verify skydick datapool datasets are mounted"; + wants = [ "zfs-mount.service" ]; + after = [ "zfs-mount.service" ]; + before = storageServices; + requiredBy = storageServices; + unitConfig.RequiresMountsFor = storageMountPoints; + + serviceConfig = { + Type = "oneshot"; + RemainAfterExit = true; + }; + + script = '' + set -euo pipefail + + fail() { + echo "Datapool readiness check: $*" >&2 + exit 1 + } + + ${pkgs.zfs}/bin/zpool list -H -o name dick >/dev/null 2>&1 || fail "pool dick is not imported" + + check_mount() { + local dataset="$1" + local mountpoint="$2" + local actual + + actual="$(${pkgs.util-linux}/bin/findmnt --noheadings --raw --mountpoint "$mountpoint" --output SOURCE,FSTYPE 2>/dev/null || true)" + [[ "$actual" == "$dataset zfs" ]] || fail "$mountpoint is '$actual', expected '$dataset zfs'" + [[ "$(${pkgs.zfs}/bin/zfs get -H -o value mounted "$dataset")" == "yes" ]] || fail "$dataset is not mounted" + } + + ${lib.concatMapStringsSep "\n" (mount: '' + check_mount ${lib.escapeShellArg mount.dataset} ${lib.escapeShellArg mount.mountPoint} + '') storageMounts} + ''; + }; + # Keep the manually created dick pool aligned with the documented policy. systemd.services.dick-zfs-properties = { description = "Apply ZFS properties for the dick datapool"; - wants = [ "zfs-mount.service" ]; - after = [ "zfs-mount.service" ]; - before = [ "nfs-server.service" "samba-smbd.service" ]; + requires = [ "dick-storage-ready.service" ]; + after = [ "dick-storage-ready.service" ]; + before = storageServices; + requiredBy = storageServices; wantedBy = [ "multi-user.target" ]; serviceConfig = { @@ -352,10 +494,6 @@ script = '' set -euo pipefail - if ! ${pkgs.zfs}/bin/zpool list -H -o name dick >/dev/null 2>&1; then - exit 0 - fi - # Child datasets on the datapool currently inherit the pool default. # Setting the root keeps read-heavy SMB/NFS trees from generating atime # metadata writes without having to stamp every descendant explicitly. @@ -368,13 +506,39 @@ ''; }; + systemd.services.nfs-server = { + requires = [ "dick-storage-ready.service" "dick-zfs-properties.service" ]; + after = [ "dick-storage-ready.service" "dick-zfs-properties.service" ]; + unitConfig.RequiresMountsFor = storageMountPoints; + }; + + systemd.services.samba-smbd = { + requires = [ "dick-storage-ready.service" "dick-zfs-properties.service" ]; + after = [ "dick-storage-ready.service" "dick-zfs-properties.service" ]; + unitConfig.RequiresMountsFor = storageMountPoints; + restartTriggers = [ config.age.secrets.skydick-samba-ldap-admin.file ]; + }; + + systemd.services.samba-winbindd.restartTriggers = [ + config.age.secrets.skydick-samba-ldap-admin.file + ]; + + systemd.services.influxdb2 = { + requires = [ "dick-storage-ready.service" "dick-zfs-properties.service" ]; + after = [ "dick-storage-ready.service" "dick-zfs-properties.service" ]; + unitConfig.RequiresMountsFor = [ "/srv/system/influxdb" ]; + }; + systemd.services.samba-ldap-admin-password = { description = "Seed Samba LDAP admin password into secrets.tdb"; wants = [ "network-online.target" ]; after = [ "network-online.target" ]; - before = [ "samba-nmbd.service" "samba-winbindd.service" "samba-smbd.service" ]; - requiredBy = [ "samba-nmbd.service" "samba-winbindd.service" "samba-smbd.service" ]; - restartTriggers = [ config.environment.etc."samba/smb.conf".source ]; + before = [ "samba-winbindd.service" "samba-smbd.service" ]; + requiredBy = [ "samba-winbindd.service" "samba-smbd.service" ]; + restartTriggers = [ + config.environment.etc."samba/smb.conf".source + config.age.secrets.skydick-samba-ldap-admin.file + ]; serviceConfig = { Type = "oneshot"; @@ -477,30 +641,46 @@ mountdPort = 20003; exports = '' - /srv 10.0.0.0/16(rw,sync,fsid=0,crossmnt,no_subtree_check,root_squash) + # Traversal-only NFSv4 pseudo-root. Do not use crossmnt here: it would + # implicitly publish every present and future ZFS child below /srv. + /srv 10.0.0.0/16(ro,sync,fsid=0,no_subtree_check,root_squash) fd99:23eb:1682::75:15(ro,sync,fsid=0,no_subtree_check,root_squash) 2a0c:b641:69c:ada0::/64(ro,sync,fsid=0,no_subtree_check,root_squash) # Shared /srv/public 10.0.0.0/16(rw,sync,no_subtree_check,root_squash) - /srv/media 10.0.0.0/16(rw,async,no_subtree_check,all_squash,anonuid=900,anongid=997) + /srv/public/datasets 10.0.0.0/16(rw,sync,nohide,no_subtree_check,root_squash) + /srv/public/models 10.0.0.0/16(rw,sync,nohide,no_subtree_check,root_squash) + /srv/public/sdk 10.0.0.0/16(rw,sync,nohide,no_subtree_check,root_squash) + # door-pek is the sole qBittorrent/*arr writer observed on this export. + /srv/media 10.0.75.15(rw,async,no_subtree_check,all_squash,anonuid=900,anongid=997) fd99:23eb:1682::75:15(rw,async,no_subtree_check,all_squash,anonuid=900,anongid=997) /srv/media/library 10.0.0.0/16(ro,sync,no_subtree_check,root_squash) - # Per-user — explicit exports; all_squash maps every client UID to the owner. - # crossmnt lets NFS clients traverse into the child datasets - # (.../files, .../bt-state, .../timemachine, .../vm) without - # separate export entries. Equivalent to setting `nohide` on each - # child. Without it, the children appear as empty mountpoints over - # NFSv3 (this bit us 2026-05-14: door-pek's baidunetdisk bind - # silently wrote to the parent dataset's placeholder instead of - # the SMB-exposed `files` child). - # ldx: locked to a single client (10.0.75.15 / the matching v6 /128), not the LAN. - /srv/users/ldx 10.0.75.15(rw,sync,no_subtree_check,crossmnt,all_squash,anonuid=1000,anongid=100) - /srv/users/ldx fd99:23eb:1682::75:15(rw,sync,no_subtree_check,crossmnt,all_squash,anonuid=1000,anongid=100) - /srv/users/ye-lw21 10.0.0.0/16(rw,sync,no_subtree_check,crossmnt,all_squash,anonuid=1002,anongid=100) - /srv/users/zhuyz24 10.0.0.0/16(rw,sync,no_subtree_check,crossmnt,all_squash,anonuid=2200000020,anongid=100) - /srv/users/zhuyz24 2a0c:b641:69c:ada0::/64(rw,sync,no_subtree_check,crossmnt,all_squash,anonuid=2200000020,anongid=100) + # Per-user exports. Each mounted child is explicit, so an unrelated new + # dataset cannot inherit owner-mapped access through crossmnt. + /srv/users/ldx 10.0.75.15(rw,sync,no_subtree_check,all_squash,anonuid=1000,anongid=100) fd99:23eb:1682::75:15(rw,sync,no_subtree_check,all_squash,anonuid=1000,anongid=100) + /srv/users/ldx/files 10.0.75.15(rw,sync,nohide,no_subtree_check,all_squash,anonuid=1000,anongid=100) fd99:23eb:1682::75:15(rw,sync,nohide,no_subtree_check,all_squash,anonuid=1000,anongid=100) + /srv/users/ldx/bt-state 10.0.75.15(rw,sync,nohide,no_subtree_check,all_squash,anonuid=1000,anongid=100) fd99:23eb:1682::75:15(rw,sync,nohide,no_subtree_check,all_squash,anonuid=1000,anongid=100) + /srv/users/ldx/vm 10.0.75.15(rw,sync,nohide,no_subtree_check,all_squash,anonuid=1000,anongid=100) fd99:23eb:1682::75:15(rw,sync,nohide,no_subtree_check,all_squash,anonuid=1000,anongid=100) + + /srv/users/ye-lw21 10.0.0.0/16(rw,sync,no_subtree_check,all_squash,anonuid=1002,anongid=100) + /srv/users/ye-lw21/files 10.0.0.0/16(rw,sync,nohide,no_subtree_check,all_squash,anonuid=1002,anongid=100) + /srv/users/ye-lw21/bt-state 10.0.0.0/16(rw,sync,nohide,no_subtree_check,all_squash,anonuid=1002,anongid=100) + /srv/users/ye-lw21/vm 10.0.0.0/16(rw,sync,nohide,no_subtree_check,all_squash,anonuid=1002,anongid=100) + + # node1 (.10) and node2 (.20) currently hold open files in this tree, but + # that snapshot does not prove they are the user's only valid clients. + /srv/users/zhuyz24 10.0.0.0/16(rw,sync,no_subtree_check,all_squash,anonuid=2200000020,anongid=100) + /srv/users/zhuyz24/files 10.0.0.0/16(rw,sync,nohide,no_subtree_check,all_squash,anonuid=2200000020,anongid=100) + /srv/users/zhuyz24/bt-state 10.0.0.0/16(rw,sync,nohide,no_subtree_check,all_squash,anonuid=2200000020,anongid=100) + /srv/users/zhuyz24/vm 10.0.0.0/16(rw,sync,nohide,no_subtree_check,all_squash,anonuid=2200000020,anongid=100) + # Exact global-v6 client addresses are not yet inventoried; retain this + # existing ACL until they can replace the /64 without breaking clients. + /srv/users/zhuyz24 2a0c:b641:69c:ada0::/64(rw,sync,no_subtree_check,all_squash,anonuid=2200000020,anongid=100) + /srv/users/zhuyz24/files 2a0c:b641:69c:ada0::/64(rw,sync,nohide,no_subtree_check,all_squash,anonuid=2200000020,anongid=100) + /srv/users/zhuyz24/bt-state 2a0c:b641:69c:ada0::/64(rw,sync,nohide,no_subtree_check,all_squash,anonuid=2200000020,anongid=100) + /srv/users/zhuyz24/vm 2a0c:b641:69c:ada0::/64(rw,sync,nohide,no_subtree_check,all_squash,anonuid=2200000020,anongid=100) # System - /srv/system/backup 10.0.0.0/16(rw,sync,no_subtree_check,no_root_squash) + /srv/system/backup 10.0.0.1(rw,sync,no_subtree_check,no_root_squash) /srv/system/vm 10.0.0.0/16(rw,sync,no_subtree_check,no_root_squash) /srv/templates/vm 10.0.0.0/16(ro,sync,no_subtree_check,root_squash) @@ -520,6 +700,7 @@ enable = true; package = pkgs.sambaFull; openFirewall = false; + nmbd.enable = false; settings = { global = { @@ -527,7 +708,7 @@ "server string" = "Skydick Storage"; "netbios name" = "SKYDICK"; security = "user"; - "passdb backend" = "ldapsam:ldap://10.0.0.1"; + "passdb backend" = "ldapsam:ldap://ldap.skyw.top"; "ldap admin dn" = "cn=admin,dc=skyw,dc=top"; "ldap suffix" = "dc=skyw,dc=top"; "ldap user suffix" = "ou=people"; @@ -535,8 +716,8 @@ "ldap machine suffix" = "ou=machines"; "ldap delete dn" = "no"; "ldap passwd sync" = "only"; - "ldap ssl" = "off"; - "ldap server require strong auth" = "no"; + "ldap ssl" = "start tls"; + "ldap server require strong auth" = "yes"; "ldap connection timeout" = "5"; "ldap timeout" = "5"; "hosts allow" = "10.0. 127. 10.253."; @@ -705,29 +886,93 @@ }; services.samba-wsdd = { - enable = true; + # Discovery was firewall-blocked and had no usable network surface. + enable = false; openFirewall = false; }; - # iSCSI — vm zvols only - services.target.enable = true; - - # Firewall: storage service ports - networking.firewall = { - allowedTCPPorts = [ - 111 # RPC (NFS) - 2049 # NFS - 445 # SMB - 139 # NetBIOS (SMB) - 3260 # iSCSI - ]; - allowedUDPPorts = [ - 111 # RPC (NFS) - 2049 # NFS (NFSv4.1+) - 137 # NetBIOS Name Service - 138 # NetBIOS Datagram - ]; - allowedTCPPortRanges = [{ from = 20000; to = 20005; }]; - allowedUDPPortRanges = [{ from = 20000; to = 20005; }]; + # iSCSI state is intentionally imperative because targetcli can contain CHAP + # material. Keep the entire directory outside the Nix store: targetcli uses + # an atomic rename, so a directory symlink is required instead of a file + # symlink. The activation dependency migrates the current /etc/target before + # NixOS replaces the formerly generated saveconfig.json. + environment.etc."target/saveconfig.json".enable = lib.mkForce false; + environment.etc.target = { + source = "/var/lib/target"; + mode = "direct-symlink"; }; + + system.activationScripts.iscsiTargetState = { + deps = [ "users" "groups" "specialfs" ]; + text = '' + state_dir=/var/lib/target + legacy_dir=/etc/target + + if [[ -L "$legacy_dir" ]]; then + if [[ "$(${pkgs.coreutils}/bin/readlink -f "$legacy_dir")" != "$state_dir" ]]; then + echo "Refusing to replace unexpected $legacy_dir symlink" >&2 + exit 1 + fi + elif [[ -d "$legacy_dir" ]]; then + if [[ -e "$state_dir" ]]; then + echo "Both $legacy_dir and $state_dir exist; refusing an ambiguous iSCSI state merge" >&2 + exit 1 + fi + ${iscsiConfigGuard}/bin/skydick-iscsi-config-guard --no-devices "$legacy_dir/saveconfig.json" + ${pkgs.coreutils}/bin/install -d -m 0755 /var/lib + ${pkgs.coreutils}/bin/mv "$legacy_dir" "$state_dir" + ${pkgs.coreutils}/bin/chmod 0700 "$state_dir" + elif [[ ! -d "$state_dir" ]]; then + echo "No persistent iSCSI configuration exists at $legacy_dir or $state_dir" >&2 + exit 1 + fi + + ${iscsiConfigGuard}/bin/skydick-iscsi-config-guard --no-devices "$state_dir/saveconfig.json" + + # The previous NixOS generation tracked this copied file in /etc/.clean. + # Remove only its exact bookkeeping record before setup-etc follows the + # new directory symlink and mistakes the persistent state for an obsolete + # generated file. + if [[ -f /etc/.clean ]]; then + ${pkgs.gnused}/bin/sed -i '\|^target/saveconfig.json$|d' /etc/.clean + fi + ''; + }; + system.activationScripts.etc.deps = [ "iscsiTargetState" ]; + + systemd.services.iscsi-target-config-ready = { + description = "Verify persistent LIO target configuration"; + requires = [ "dick-storage-ready.service" ]; + after = [ "dick-storage-ready.service" ]; + before = [ "iscsi-target.service" ]; + requiredBy = [ "iscsi-target.service" ]; + unitConfig.RequiresMountsFor = [ "/srv/system/vm" ]; + serviceConfig = { + Type = "oneshot"; + RemainAfterExit = true; + }; + script = '' + ${iscsiConfigGuard}/bin/skydick-iscsi-config-guard /etc/target/saveconfig.json + ''; + }; + + # iSCSI — vm zvols only. A switch must never clear an active target/session; + # changed units take effect at the next controlled restart or reboot. + services.target.enable = true; + systemd.services.iscsi-target = { + requires = [ + "dick-storage-ready.service" + "dick-zfs-properties.service" + "iscsi-target-config-ready.service" + ]; + after = [ + "dick-storage-ready.service" + "dick-zfs-properties.service" + "iscsi-target-config-ready.service" + ]; + unitConfig.RequiresMountsFor = [ "/srv/system/vm" ]; + restartIfChanged = false; + stopIfChanged = false; + }; + } diff --git a/hosts/skydick/default.nix b/hosts/skydick/default.nix index c510c33..711d9a5 100644 --- a/hosts/skydick/default.nix +++ b/hosts/skydick/default.nix @@ -109,20 +109,36 @@ address = "10.0.0.1"; interface = "bond40g"; }; - # Single primary so systemd-resolved doesn't load-balance us off to a - # resolver that has no analytics-blocking. Fallback handled below. - nameservers = [ "10.0.0.1" ]; - - # IPv6 is enabled for SLAAC on bond40g. Keep DNS pinned to 10.0.0.1: + # IPv6 is enabled for SLAAC on bond40g. Keep DNS pinned to main MosDNS: # RA/DHCPv6-provided DNS is disabled in the networkd link config below. enableIPv6 = true; firewall = { enable = true; - allowedTCPPorts = [ 22 ]; + backend = "nftables"; + + # This host has a globally routed SLAAC address. Keep storage and + # monitoring services scoped to their actual clients instead of opening + # them on every IPv4 and IPv6 address. + extraInputRules = '' + iifname "bond40g" ip saddr 10.0.0.0/8 tcp dport 22 accept comment "SSH from routed private networks" + iifname "bond40g" ip6 saddr { fd99:23eb:1682::/48, 2a0c:b641:69c:ada0::/64 } tcp dport 22 accept comment "SSH from Skyworks IPv6 networks" + + iifname "bond40g" ip saddr 10.0.0.0/16 meta l4proto { tcp, udp } th dport { 111, 2049, 20001-20003 } accept comment "NFS IPv4 clients" + iifname "bond40g" ip6 saddr fd99:23eb:1682::75:15 meta l4proto { tcp, udp } th dport { 111, 2049, 20001-20003 } accept comment "door-pek NFS over ULA" + iifname "bond40g" ip6 saddr 2a0c:b641:69c:ada0::/64 meta l4proto { tcp, udp } th dport { 111, 2049, 20001-20003 } accept comment "declared NFS global IPv6 clients" + + iifname "bond40g" ip saddr { 10.0.0.0/16, 10.253.0.0/16 } tcp dport 445 accept comment "SMB client networks" + iifname "bond40g" ip saddr 10.0.200.11 tcp dport 3260 accept comment "rh1288v3 iSCSI initiator" + iifname "bond40g" ip saddr 10.0.75.15 tcp dport 8086 accept comment "door1 InfluxDB and Grafana" + iifname "bond40g" ip saddr { 10.0.0.1, 10.0.75.15 } tcp dport 9100 accept comment "node exporter scrapers" + ''; }; }; + networking.nftables.enable = true; + services.openssh.openFirewall = false; + # systemd-networkd: accept RA on bond40g for SLAAC, but keep DNS/search # domains and the default route under the explicit IPv4 config above. systemd.network.networks."40-bond40g" = { @@ -143,14 +159,23 @@ }; }; - # DNS routed through the network's mosdns at 10.0.0.1 so this host inherits - # CN-aware split routing and analytics blocking. AliDNS is the first - # fallback (close, clean, no GFW games), Cloudflare second. + # Use only main's filtered MosDNS endpoints. Global DNS servers are peers in + # resolved, not ordered failovers; adding AliDNS here can intermittently + # return polluted foreign-domain answers, while adding a public resolver + # bypasses the network's split-DNS and blocking policy. services.resolved = { enable = true; - fallbackDns = [ "223.5.5.5" "1.1.1.1" ]; + settings.Resolve = { + DNS = [ "10.0.0.1" "fd99:23eb:1682::1" ]; + FallbackDNS = ""; + LLMNR = false; + MulticastDNS = false; + }; }; + # Trust the private CA used by the LDAP service on the main gateway. + security.pki.certificateFiles = [ ../../certs/skyw-ldap-ca.crt ]; + # Wait only for bond40g, not individual member ports — a disconnected port # (cable maintenance) should not stall boot by 2 minutes. systemd.network.wait-online.anyInterface = true; @@ -188,9 +213,9 @@ fi done - # Increase Mellanox ring buffers for 10GbE burst absorption - for nic in enp4s0f0np0 enp4s0f1np1; do - ethtool -G "$nic" rx 4096 tx 4096 2>/dev/null || true + # Keep both active 40GbE bond members at their verified hardware maximum. + for nic in enp130s0 enp130s0d1; do + ethtool -G "$nic" rx 8192 tx 8192 2>/dev/null || true done ''; }; @@ -328,9 +353,9 @@ loginPam = false; nsswitch = true; daemon.enable = true; - server = "ldap://10.0.0.1/"; + server = "ldap://ldap.skyw.top/"; base = "dc=skyw,dc=top"; - useTLS = false; + useTLS = true; timeLimit = 5; bind = { @@ -341,10 +366,19 @@ }; daemon.extraConfig = '' + ssl start_tls + tls_reqcert demand + tls_cacertfile /etc/ssl/certs/ca-certificates.crt nss_initgroups_ignoreusers ALLLOCAL ''; }; + # The LDAP module only sees the stable /run/agenix path. Track the encrypted + # source so rotating the bind secret restarts nslcd and rewrites its config. + systemd.services.nslcd.restartTriggers = [ + config.age.secrets.skydick-ldap-bind.file + ]; + # ========================================================================== # MONITORING # ========================================================================== @@ -388,14 +422,20 @@ # ========================================================================== # INFLUXDB + TELEGRAF MONITORING # ========================================================================== - skyworks.influxdb.enable = true; + skyworks.influxdb = { + enable = true; + openFirewall = false; + }; skyworks.monitoring = { enable = true; influxUrl = "http://127.0.0.1:8086"; bucket = "skydick"; netInterfaces = [ "bond40g" ]; - nodeExporter.enable = true; + nodeExporter = { + enable = true; + openFirewall = false; + }; }; system.stateVersion = "25.11"; diff --git a/hosts/xlab-gateway/default.nix b/hosts/xlab-gateway/default.nix index eef4344..7bd28a8 100644 --- a/hosts/xlab-gateway/default.nix +++ b/hosts/xlab-gateway/default.nix @@ -1,6 +1,5 @@ # xlab-gateway - Lab Gateway / Router -# TODO: Migrate from Debian 12 to NixOS -# Current services: Kea DHCP4/6, DDNS, radvd, WireGuard, NAT, policy routing +# Current services: Kea DHCP4/DDNS, radvd, WireGuard, NAT, policy routing { config, pkgs, lib, ... }: { @@ -16,6 +15,32 @@ nixpkgs.hostPlatform = lib.mkDefault "x86_64-linux"; + # This router has no interactive virtual console. On this hardware, + # systemd-vconsole-setup can block indefinitely writing its UTF-8 escape to + # /dev/tty0, which also stalls switch-to-configuration. Disable the real + # setup (including in initrd), but retain no-op units so a live switch never + # hits a remove-then-reload ordering failure. + console.enable = false; + systemd.services.systemd-vconsole-setup = { + description = "No-op virtual console setup on headless xlab gateway"; + serviceConfig = { + Type = "oneshot"; + ExecStart = "${pkgs.coreutils}/bin/true"; + RemainAfterExit = true; + }; + }; + systemd.services.reload-systemd-vconsole-setup = { + description = "No-op virtual console reload on headless xlab gateway"; + wantedBy = [ "multi-user.target" ]; + reloadIfChanged = true; + serviceConfig = { + Type = "oneshot"; + ExecStart = "${pkgs.coreutils}/bin/true"; + ExecReload = "${pkgs.coreutils}/bin/true"; + RemainAfterExit = true; + }; + }; + boot = { loader = { systemd-boot.enable = true; @@ -41,7 +66,7 @@ }; # tcp_bbr is a module on this kernel (reno cubic only built-in) — load it # so tcp_congestion_control = bbr above takes effect. - kernelModules = [ "tcp_bbr" ]; + kernelModules = [ "tcp_bbr" "wireguard" ]; }; # WireGuard-crypto + routing appliance: keep cores pinned so WG throughput @@ -92,15 +117,35 @@ Type = "oneshot"; RemainAfterExit = true; ExecStart = pkgs.writeShellScript "capwap-df-clear-up" '' - set -e - for i in $(seq 1 30); do ${pkgs.iproute2}/bin/ip link show bond.lan254 >/dev/null 2>&1 && break; sleep 1; done - ${pkgs.iproute2}/bin/tc qdisc del dev bond.lan254 handle ffff: ingress 2>/dev/null || true - ${pkgs.iproute2}/bin/tc qdisc add dev bond.lan254 handle ffff: ingress - ${pkgs.iproute2}/bin/tc filter add dev bond.lan254 parent ffff: protocol ip prio 1 u32 \ - match ip protocol 17 0xff match ip dst 10.0.10.10/32 \ - action pedit ex munge ip df set 0 pipe action csum ip + set -eu + for _ in $(${pkgs.coreutils}/bin/seq 1 30); do + ${pkgs.iproute2}/bin/ip link show bond.lan254 >/dev/null 2>&1 && break + ${pkgs.coreutils}/bin/sleep 1 + done + ${pkgs.iproute2}/bin/ip link show bond.lan254 >/dev/null + + # clsact is shared infrastructure. Never replace or remove it, and own + # only the two CAPWAP priorities below. + if ! ${pkgs.iproute2}/bin/tc qdisc show dev bond.lan254 | ${pkgs.gnugrep}/bin/grep -q '^qdisc clsact '; then + ${pkgs.iproute2}/bin/tc qdisc add dev bond.lan254 clsact + fi + for port in 5246 5247; do + # `replace` without an explicit flower handle can still return + # EEXIST. Delete only our owned preference, then recreate it. + ${pkgs.iproute2}/bin/tc filter del dev bond.lan254 ingress \ + protocol ip pref "$port" 2>/dev/null || true + ${pkgs.iproute2}/bin/tc filter add dev bond.lan254 ingress \ + protocol ip pref "$port" flower \ + ip_proto udp dst_ip 10.0.10.10 dst_port "$port" \ + action pedit ex munge ip df set 0 pipe action csum ip + done ''; - ExecStop = "-${pkgs.iproute2}/bin/tc qdisc del dev bond.lan254 handle ffff: ingress"; + ExecStop = pkgs.writeShellScript "capwap-df-clear-down" '' + for port in 5246 5247; do + ${pkgs.iproute2}/bin/tc filter del dev bond.lan254 ingress \ + protocol ip pref "$port" 2>/dev/null || true + done + ''; }; }; diff --git a/hosts/xlab-gateway/dhcp.nix b/hosts/xlab-gateway/dhcp.nix index 32124f0..9426c1e 100644 --- a/hosts/xlab-gateway/dhcp.nix +++ b/hosts/xlab-gateway/dhcp.nix @@ -1,8 +1,34 @@ -# xlab-gateway DHCP + DDNS + radvd -# Kea DHCPv4/v6 on bond.lan254, DDNS forwarding to BIND9 at 10.0.0.1:5353 +# xlab-gateway DHCP + DDNS + IPv6 SLAAC +# Kea DHCPv4 on bond.lan254, DDNS forwarding to BIND9 at 10.0.0.1:5353 { config, pkgs, ... }: +let + lanDeviceUnit = "sys-subsystem-net-devices-bond.lan254.device"; + waitForLanAddresses = pkgs.writeShellScript "wait-for-xlab-lan-addresses" '' + set -eu + + for _ in $(${pkgs.coreutils}/bin/seq 1 30); do + if ${pkgs.iproute2}/bin/ip -4 -brief address show dev bond.lan254 \ + | ${pkgs.gnugrep}/bin/grep -Fq "10.253.254.1/24" \ + && ${pkgs.iproute2}/bin/ip -6 -brief address show dev bond.lan254 \ + | ${pkgs.gnugrep}/bin/grep -Fq "fd99:23eb:1682:1::1/64"; then + exit 0 + fi + ${pkgs.coreutils}/bin/sleep 1 + done + + echo "bond.lan254 did not acquire its configured addresses within 30 seconds" >&2 + exit 1 + ''; +in { + age.secrets.xlab-ddns-tsig = { + file = ../../secrets/xlab-ddns-tsig.age; + owner = "kea"; + group = "kea"; + mode = "0400"; + }; + # =========================================================================== # Kea DHCPv4 # =========================================================================== @@ -32,7 +58,7 @@ option-data = [ { name = "routers"; data = "10.253.254.1"; } - { name = "domain-name-servers"; data = "10.0.0.1"; } + { name = "domain-name-servers"; data = "10.253.254.1"; } { name = "domain-name"; data = "dev.skyw.top"; } { name = "domain-search"; data = "dev.skyw.top"; } ]; @@ -64,7 +90,7 @@ option-data = [ { name = "subnet-mask"; data = "255.255.255.0"; } { name = "routers"; data = "10.253.254.1"; } - { name = "domain-name-servers"; data = "10.0.0.1"; } + { name = "domain-name-servers"; data = "10.253.254.1"; } # Classless static routes: 10.0.0.0/16 via 10.253.254.1, default via 10.253.254.1 { code = 121; csv-format = false; data = "100A000AFDFE01000AFDFE01"; } # MS classless static routes (same) @@ -91,73 +117,21 @@ }; - systemd.services.kea-dhcp4 = { - after = [ "network-online.target" ]; - wants = [ "network-online.target" ]; - }; - - # =========================================================================== - # Kea DHCPv6 - # =========================================================================== - services.kea.dhcp6 = { - enable = true; - settings = { - interfaces-config.interfaces = [ "bond.lan254" ]; - - lease-database = { - type = "memfile"; - name = "/var/lib/kea/kea-leases6.csv"; - persist = true; - lfc-interval = 3600; - }; - - expired-leases-processing = { - reclaim-timer-wait-time = 10; - flush-reclaimed-timer-wait-time = 25; - hold-reclaimed-time = 3600; - max-reclaim-leases = 100; - max-reclaim-time = 250; - }; - - valid-lifetime = 86400; - preferred-lifetime = 72000; - renew-timer = 21600; - rebind-timer = 43200; - - subnet6 = [ - { - id = 1; - subnet = "fd99:23eb:1682:1::/64"; - pools = [ - { pool = "fd99:23eb:1682:1::100 - fd99:23eb:1682:1::ffff"; } - ]; - option-data = [ - { name = "dns-servers"; data = "fd99:23eb:1682::1"; } - { name = "domain-search"; data = "dev.skyw.top"; } - ]; - } - ]; - - ddns-send-updates = true; - ddns-qualifying-suffix = "dev.skyw.top."; - - dhcp-ddns = { - enable-updates = true; - max-queue-size = 1024; - ncr-protocol = "UDP"; - ncr-format = "JSON"; - sender-ip = "::1"; - sender-port = 0; - server-ip = "::1"; - server-port = 53001; - }; + # The global network-online target is intentionally disabled on this router, + # so wait for the actual LAN device and addresses before Kea opens sockets. + systemd.services.kea-dhcp4-server = { + requires = [ lanDeviceUnit ]; + wants = [ "kea-dhcp-ddns-server.service" ]; + after = [ + lanDeviceUnit + "systemd-networkd.service" + "kea-dhcp-ddns-server.service" + ]; + serviceConfig = { + ExecStartPre = [ waitForLanAddresses ]; + RestartSec = "2s"; }; }; - - systemd.services.kea-dhcp6 = { - after = [ "network-online.target" ]; - wants = [ "network-online.target" ]; - }; # =========================================================================== @@ -171,17 +145,18 @@ tsig-keys = [ { - name = "edge-ddns-key"; + # The main gateway temporarily accepts both names during rotation; + # xlab uses only the replacement key. + name = "edge-ddns-key-next"; algorithm = "HMAC-SHA256"; - # TODO: Move TSIG secret to agenix - secret = "qq+zsTGsWG4ENW9mazyE3/JFKhsUiUR1ex4geYv8OIo="; + secret-file = config.age.secrets.xlab-ddns-tsig.path; } ]; forward-ddns.ddns-domains = [ { name = "dev.skyw.top."; - key-name = "edge-ddns-key"; + key-name = "edge-ddns-key-next"; dns-servers = [{ ip-address = "10.0.0.1"; port = 5353; }]; } ]; @@ -189,20 +164,24 @@ reverse-ddns.ddns-domains = [ { name = "10.in-addr.arpa."; - key-name = "edge-ddns-key"; + key-name = "edge-ddns-key-next"; dns-servers = [{ ip-address = "10.0.0.1"; port = 5353; }]; } { name = "2.8.6.1.b.e.3.2.9.9.d.f.ip6.arpa."; - key-name = "edge-ddns-key"; + key-name = "edge-ddns-key-next"; dns-servers = [{ ip-address = "10.0.0.1"; port = 5353; }]; } ]; }; }; + systemd.services.kea-dhcp-ddns-server.restartTriggers = [ + config.age.secrets.xlab-ddns-tsig.file + ]; + # =========================================================================== - # radvd - IPv6 Router Advertisements + # radvd - IPv6 SLAAC, DNS, and search-domain advertisements # =========================================================================== services.radvd = { enable = true; @@ -210,7 +189,7 @@ interface bond.lan254 { AdvSendAdvert on; AdvManagedFlag off; - AdvOtherConfigFlag on; + AdvOtherConfigFlag off; MinRtrAdvInterval 30; MaxRtrAdvInterval 100; prefix fd99:23eb:1682:1::/64 { @@ -218,10 +197,19 @@ AdvAutonomous on; AdvRouterAddr on; }; - RDNSS fd99:23eb:1682::1 { + RDNSS fd99:23eb:1682:1::1 { AdvRDNSSLifetime 3600; }; + DNSSL dev.skyw.top { + AdvDNSSLLifetime 3600; + }; }; ''; }; + + systemd.services.radvd = { + requires = [ lanDeviceUnit ]; + after = [ lanDeviceUnit "systemd-networkd.service" ]; + serviceConfig.ExecStartPre = [ waitForLanAddresses ]; + }; } diff --git a/hosts/xlab-gateway/networking.nix b/hosts/xlab-gateway/networking.nix index b4b3a8e..f4ae56c 100644 --- a/hosts/xlab-gateway/networking.nix +++ b/hosts/xlab-gateway/networking.nix @@ -5,6 +5,245 @@ { config, pkgs, ... }: let + xlabWg = pkgs.writeShellApplication { + name = "xlab-wg"; + runtimeInputs = with pkgs; [ coreutils gnugrep gawk iproute2 systemd wireguard-tools ]; + text = '' + set -euo pipefail + + state=/var/lib/xlab-wireguard + command=status + target=all + if (( $# >= 1 )); then command=$1; shift; fi + if (( $# >= 1 )); then target=$1; shift; fi + if (( $# != 0 )); then + echo "usage: sudo xlab-wg {status|enable|disable|restart|health} [wgnet|skyworks|all]" >&2 + exit 2 + fi + if (( EUID != 0 )); then + echo "xlab-wg must run as root" >&2 + exit 1 + fi + + install -d -m 0750 "$state" + + configured() { + local iface=$1 port peer + case "$iface" in + wg-to-wgnet) + port=51998 + peer='H+PAPw+1MsE50Of4VMMPzMbzGG731CkNrIgaXbxcFwk=' + ;; + wg-to-skyworks) + port=46961 + peer='yyHQ8fg1riI9BJPUjydh/C2MTiA/p0Pb1f9Hc88BuCk=' + ;; + *) return 2 ;; + esac + + ip link show dev "$iface" >/dev/null 2>&1 \ + && [[ "$(wg show "$iface" listen-port 2>/dev/null)" == "$port" ]] \ + && [[ "$(wg show "$iface" peers 2>/dev/null)" == "$peer" ]] + } + + ensure_configured() { + local iface=$1 + configured "$iface" && return 0 + + echo "$iface is missing or incomplete; reloading networkd configuration" >&2 + networkctl reload + sleep 2 + networkctl reconfigure "$iface" 2>/dev/null || true + sleep 2 + configured "$iface" + } + + trigger_handshake() { + case "$1" in + wg-to-wgnet) + ping -q -c 1 -W 1 -I wg-to-wgnet 10.0.0.1 >/dev/null 2>&1 || true + ;; + wg-to-skyworks) + ping -q -c 1 -W 1 -I wg-to-skyworks 10.239.0.1 >/dev/null 2>&1 || true + ;; + esac + } + + handshake_age() { + local iface=$1 now latest + now=$(date +%s) + latest=$(wg show "$iface" latest-handshakes 2>/dev/null | gawk 'NR == 1 { print $2 }') + [[ "$latest" =~ ^[0-9]+$ ]] || latest=0 + if (( latest == 0 )); then + echo 2147483647 + else + echo $((now - latest)) + fi + } + + wait_for_handshake() { + local iface=$1 age + for _ in $(seq 1 10); do + trigger_handshake "$iface" + sleep 2 + age=$(handshake_age "$iface") + if (( age <= 180 )); then + return 0 + fi + done + return 1 + } + + disabled() { + [[ -e "$state/disabled-$1" ]] + } + + health_one() { + local iface=$1 now last=0 age + if disabled "$iface"; then + ip link set dev "$iface" down 2>/dev/null || true + echo "$iface is administratively disabled" + return 0 + fi + + ensure_configured "$iface" || { + echo "$iface configuration is incomplete after networkd reload" >&2 + return 1 + } + ip link set dev "$iface" up + + age=$(handshake_age "$iface") + if (( age <= 180 )); then + echo "$iface is healthy; latest handshake is ''${age}s old" + return 0 + fi + + now=$(date +%s) + if [[ -s "$state/last-recovery-$iface" ]]; then + read -r last < "$state/last-recovery-$iface" || last=0 + fi + if (( now - last < 300 )); then + echo "$iface handshake is stale and recovery is cooling down" >&2 + return 1 + fi + + echo "$now" > "$state/last-recovery-$iface" + echo "$iface handshake is stale; cycling the link after WAN readiness" >&2 + ip link set dev "$iface" down + sleep 1 + ip link set dev "$iface" up + + if wait_for_handshake "$iface"; then + age=$(handshake_age "$iface") + echo "$iface recovered; latest handshake is ''${age}s old" + return 0 + fi + + echo "$iface did not handshake after recovery" >&2 + return 1 + } + + status_one() { + local iface=$1 desired=enabled age=unavailable + disabled "$iface" && desired=disabled + if ip link show dev "$iface" >/dev/null 2>&1; then + age=$(handshake_age "$iface") + fi + printf '%s: desired=%s handshake_age=%ss\n' "$iface" "$desired" "$age" + ip -brief link show dev "$iface" 2>/dev/null || true + wg show "$iface" 2>/dev/null || true + } + + disable_one() { + local iface=$1 + touch "$state/disabled-$iface" + ip link set dev "$iface" down 2>/dev/null || true + echo "$iface disabled; automatic recovery will leave it down" + } + + enable_one() { + local iface=$1 + rm -f "$state/disabled-$iface" "$state/last-recovery-$iface" + ensure_configured "$iface" + ip link set dev "$iface" up + trigger_handshake "$iface" + echo "$iface enabled" + } + + restart_one() { + local iface=$1 + if disabled "$iface"; then + echo "$iface is disabled; use 'xlab-wg enable' instead" >&2 + return 1 + fi + ensure_configured "$iface" + ip link set dev "$iface" down + sleep 1 + ip link set dev "$iface" up + trigger_handshake "$iface" + rm -f "$state/last-recovery-$iface" + echo "$iface restarted" + } + + run_for_target() { + local action=$1 rc=0 + case "$target" in + wgnet|wg-to-wgnet) "$action" wg-to-wgnet || rc=1 ;; + skyworks|wg-to-skyworks) "$action" wg-to-skyworks || rc=1 ;; + all) + "$action" wg-to-wgnet || rc=1 + "$action" wg-to-skyworks || rc=1 + ;; + *) + echo "unknown WireGuard target: $target" >&2 + return 2 + ;; + esac + return "$rc" + } + + wait_for_wan() { + for _ in $(seq 1 60); do + if ip -4 route get 166.111.17.108 2>/dev/null | grep -q 'dev wan99.0'; then + return 0 + fi + sleep 2 + done + echo "campus WAN route did not become ready" >&2 + return 1 + } + + case "$command" in + status) + run_for_target status_one + ;; + disable|down) + systemctl stop xlab-wireguard-readiness.service 2>/dev/null || true + run_for_target disable_one + ;; + enable|up) + wait_for_wan + run_for_target enable_one + systemctl restart xlab-wireguard-readiness.service + ;; + restart) + wait_for_wan + systemctl stop xlab-wireguard-readiness.service 2>/dev/null || true + run_for_target restart_one + systemctl restart xlab-wireguard-readiness.service + ;; + health) + wait_for_wan + run_for_target health_one + ;; + *) + echo "usage: sudo xlab-wg {status|enable|disable|restart|health} [wgnet|skyworks|all]" >&2 + exit 2 + ;; + esac + ''; + }; + cnDirectSetScript = '' set -euo pipefail umask 027 @@ -170,23 +409,41 @@ chain forward { type filter hook forward priority filter; policy drop; iifname "bond.lan254" ip saddr != 10.253.254.0/24 counter drop comment "Drop spoofed LAN source" - iifname { "bond.lan254", "wg-to-wgnet" } accept ct state established,related accept + iifname "bond.lan254" accept + iifname "wg-to-wgnet" ip saddr 10.0.0.0/8 ip daddr 10.253.254.0/24 accept comment "Internal WG clients to xlab LAN" } } table inet input_filter { chain input { type filter hook input priority filter; policy drop; + ct state invalid drop + + # Reject forged LAN/management sources before service ACLs. DHCPv4 + # legitimately starts at 0.0.0.0. IPv6 DAD probes use the + # unspecified source, while normal control traffic is link-local. + iifname "bond.lan254" ip saddr != 10.253.254.0/24 ip saddr != 0.0.0.0 drop + iifname "bond.lan254" ip6 saddr :: ip6 nexthdr ipv6-icmp accept comment "Allow IPv6 duplicate-address detection" + iifname "bond.lan254" ip6 saddr != fd99:23eb:1682:1::/64 ip6 saddr != fe80::/10 drop + iifname "bond.mgmt" ip saddr != 192.168.1.0/24 drop + ct state established,related accept iif lo accept - # LAN, management, and WireGuard — trust fully - iifname { "bond.lan254", "bond.mgmt", "wg-to-wgnet", "wg-to-skyworks" } accept + # The router only exposes DHCP and SSH on trusted access networks. + iifname "bond.lan254" udp dport 67 accept + iifname "bond.lan254" meta l4proto { tcp, udp } th dport 53 accept comment "LAN DNS resolver" + iifname { "bond.lan254", "bond.mgmt" } tcp dport 22 accept + iifname "wg-to-wgnet" ip saddr 10.0.0.0/8 tcp dport 22 accept + iifname "wg-to-wgnet" ip6 saddr fd99:23eb:1682::/48 tcp dport 22 accept + iifname "wg-to-skyworks" tcp dport 22 accept # WAN — DHCP client replies (v4 + v6) iifname "wan99.0" udp sport 67 udp dport 68 accept iifname "wan99.0" udp sport 547 udp dport 546 accept + iifname "wan99.0" udp dport { 46961, 51998 } accept comment "WireGuard handshakes" + iifname "wan99.0" ip saddr { 166.111.17.81, 166.111.17.108 } tcp dport 22 accept comment "Break-glass SSH from server3 or main" # ICMP/ICMPv6 for path MTU discovery and diagnostics ip protocol icmp accept @@ -198,8 +455,9 @@ chain forward { type filter hook forward priority filter; policy drop; iifname "bond.lan254" ip6 saddr != fd99:23eb:1682:1::/64 counter drop comment "Drop spoofed LAN v6 source" - iifname { "bond.lan254", "wg-to-wgnet" } accept ct state established,related accept + iifname "bond.lan254" accept + iifname "wg-to-wgnet" ip6 saddr fd99:23eb:1682::/48 ip6 daddr fd99:23eb:1682:1::/64 accept comment "Internal WG clients to xlab LAN v6" } } @@ -210,10 +468,34 @@ set cn_direct_v4 { type ipv4_addr; flags interval; } set cn_direct_v6 { type ipv6_addr; flags interval; } + # These campus and Alibaba AS45102 ranges need the same failover + # semantics as the APNIC-CN data. Marking them selects main while a + # WAN default exists, then falls through to table 1002 when it does + # not. A throw route would skip table 1002 and blackhole an outage. + set regional_direct_v4 { + type ipv4_addr + flags interval + elements = { + 166.111.0.0/16, 101.5.0.0/16, 101.6.0.0/16, + 59.66.0.0/16, 183.172.0.0/16, 183.173.0.0/16, + 47.246.0.0/16, 43.109.0.0/16, 163.181.0.0/16, + 139.95.0.0/16 + } + } + set regional_direct_v6 { + type ipv6_addr + flags interval + elements = { + 2402:f000::/32, 2001:da8::/32, 2001:250::/35 + } + } + chain prerouting { type filter hook prerouting priority mangle; policy accept; iifname "bond.lan254" ip daddr @cn_direct_v4 counter meta mark set 0x10 comment "APNIC-CN direct via xlab WAN" iifname "bond.lan254" ip6 daddr @cn_direct_v6 counter meta mark set 0x10 comment "APNIC-CN v6 direct via xlab WAN" + iifname "bond.lan254" ip daddr @regional_direct_v4 counter meta mark set 0x10 comment "Campus and Alibaba direct via xlab WAN" + iifname "bond.lan254" ip6 daddr @regional_direct_v6 counter meta mark set 0x10 comment "CERNET v6 direct via xlab WAN" } } @@ -236,10 +518,12 @@ # truncated. Use `& (syn|rst) == syn` so SYN-ACK still matches # (it has SYN+ACK), only RST packets are excluded. # - # `set rt mtu` picks MSS from the egress route's MTU - # (wg-to-wgnet=1420 → 1380 v4 / 1360 v6), instead of the - # one-size-fits-all 1280 that was too low AND one-directional. - tcp flags & (syn|rst) == syn tcp option maxseg size set rt mtu + # Outbound SYN/SYN-ACK uses the tunnel route MTU. For packets + # arriving from WireGuard the egress LAN route is 1500, so clamp + # those explicitly to the 1420-byte tunnel ceilings as well. + oifname { "wg-to-wgnet", "wg-to-skyworks" } tcp flags & (syn|rst) == syn tcp option maxseg size set rt mtu + iifname { "wg-to-wgnet", "wg-to-skyworks" } meta nfproto ipv4 tcp flags & (syn|rst) == syn tcp option maxseg size set 1380 + iifname { "wg-to-wgnet", "wg-to-skyworks" } meta nfproto ipv6 tcp flags & (syn|rst) == syn tcp option maxseg size set 1360 } chain postrouting { type filter hook postrouting priority filter; policy accept; @@ -271,6 +555,8 @@ }; bondConfig = { Mode = "balance-xor"; + MIIMonitorSec = "100ms"; + TransmitHashPolicy = "layer3+4"; }; }; @@ -314,6 +600,7 @@ }; wireguardConfig = { PrivateKeyFile = config.age.secrets.xlab-wg-wgnet.path; + ListenPort = 51998; }; wireguardPeers = [ { @@ -333,12 +620,20 @@ }; wireguardConfig = { PrivateKeyFile = config.age.secrets.xlab-wg-skyworks.path; + ListenPort = 46961; }; wireguardPeers = [ { PublicKey = "yyHQ8fg1riI9BJPUjydh/C2MTiA/p0Pb1f9Hc88BuCk="; - AllowedIPs = [ "0.0.0.0/0" "::/0" ]; - Endpoint = "wgep.thu-skyworks.org:16777"; + AllowedIPs = [ + "10.0.2.144/32" + "10.74.0.0/16" + "10.239.0.0/16" + "fd5b:8249:afe:cb9a::/64" + ]; + # Avoid a boot-time dependency cycle: this tunnel previously waited + # indefinitely for DNS that was itself routed through wg-to-wgnet. + Endpoint = "166.111.17.81:16777"; PersistentKeepalive = 20; } ]; @@ -382,32 +677,8 @@ IPv6AcceptRA = false; }; routes = [ - #{ Destination = "10.0.0.0/16"; Gateway = "10.253.0.1"; } - # Throw routes for Tsinghua ranges → punches holes in WireGuard table - { Destination = "166.111.0.0/16"; Type = "throw"; Table = 1002; } - { Destination = "101.5.0.0/16"; Type = "throw"; Table = 1002; } - { Destination = "101.6.0.0/16"; Type = "throw"; Table = 1002; } - { Destination = "59.66.0.0/16"; Type = "throw"; Table = 1002; } - { Destination = "183.172.0.0/16"; Type = "throw"; Table = 1002; } - { Destination = "183.173.0.0/16"; Type = "throw"; Table = 1002; } - # Alibaba overseas CDN ranges used by Taobao/Tmall HTTPDNS. These - # addresses are outside chnroute, but sending them through the abroad - # tunnel makes Alibaba's geo/risk controls see a foreign source. A - # throw in the WG table falls through to main, so they leave directly - # through the campus WAN. Ported from nix-infra commit - # 8692bebcfaf59be065d84d064c367cb01dddc691. - { Destination = "47.246.0.0/16"; Type = "throw"; Table = 1002; } - { Destination = "43.109.0.0/16"; Type = "throw"; Table = 1002; } - { Destination = "163.181.0.0/16"; Type = "throw"; Table = 1002; } - { Destination = "139.95.0.0/16"; Type = "throw"; Table = 1002; } - # ========= 新增下面这一行 ========= - # 告诉策略路由表:如果是去往管理网段,不要走隧道,跳回 main 表处理 + # Keep the local management VLAN outside the foreign tunnel. { Destination = "192.168.1.0/24"; Type = "throw"; Table = 1002; } - # ================================== - # v6 CERNET throw routes → CERNET/Tsinghua v6 (incl. tuna 2402:f000::/32) direct via wan99.0, not the wg tunnel - { Destination = "2402:f000::/32"; Type = "throw"; Table = 1002; } - { Destination = "2001:da8::/32"; Type = "throw"; Table = 1002; } - { Destination = "2001:250::/35"; Type = "throw"; Table = 1002; } ]; routingPolicyRules = [ # nft table xlab_pbr marks APNIC-CN destinations for the local WAN. @@ -486,8 +757,8 @@ # Don't set link-specific DNS on the WAN. Tsinghua's resolvers # (166.111.8.28/29) are subject to GFW DNS poisoning, and link-DNS # would override the global services.resolved policy. Global DNS - # (10.0.0.1 → mosdns) handles CN routing internally; resolved's - # fallbackDns covers the case when 10.0.0.1 is unreachable. + # uses main's IPv4 and IPv6 MosDNS listeners. Numeric WireGuard + # endpoints keep tunnel bootstrap independent of DNS. DHCP = "yes"; }; dhcpV4Config = { @@ -508,22 +779,25 @@ # WireGuard: wg-to-wgnet (main tunnel for "freedom" routing) "50-wg-to-wgnet" = { matchConfig.Name = "wg-to-wgnet"; + linkConfig = { + ActivationPolicy = "manual"; + RequiredForOnline = "no"; + }; addresses = [ { Address = "10.253.1.37/32"; } { Address = "fd99:23eb:1682:fd::128/128"; } ]; routes = [ { Destination = "0.0.0.0/0"; Scope = "link"; Table = 1002; } - { Destination = "::/0"; Scope = "link"; Table = 1002; } + { Destination = "::/0"; Table = 1002; } # Internal ULA via tunnel (main table) so gateway + LAN can reach 10.0.0.1's IPv6 - { Destination = "fd99:23eb:1682::/48"; Scope = "link"; } + { Destination = "fd99:23eb:1682::/48"; } - # ========= 新增下面这两行 ========= + # Internal networks are reached through the main gateway tunnel. { Destination = "10.0.0.0/16"; Scope = "link"; } { Destination = "10.5.0.0/24"; Scope = "link"; } { Destination = "10.253.0.0/24"; Scope = "link"; } { Destination = "10.254.0.0/24"; Scope = "link"; } - # ================================== ]; }; @@ -531,6 +805,10 @@ # WireGuard: wg-to-skyworks "50-wg-to-skyworks" = { matchConfig.Name = "wg-to-skyworks"; + linkConfig = { + ActivationPolicy = "manual"; + RequiredForOnline = "no"; + }; addresses = [ { Address = "10.239.0.9/16"; } { Address = "fd5b:8249:afe:cb9a::9/64"; } @@ -542,6 +820,45 @@ }; }; + environment.systemPackages = [ xlabWg ]; + + # Networkd creates and configures the interfaces but leaves their + # administrative state alone. This unit waits for the campus route, raises + # enabled tunnels, validates real handshakes, and cycles only stale peers. + # Persistent disable markers let operators intentionally keep either path + # down without racing an automatic recovery timer. + systemd.services.xlab-wireguard-readiness = { + description = "Recover and verify xlab WireGuard handshakes"; + wants = [ "systemd-networkd.service" "nftables.service" ]; + after = [ "systemd-networkd.service" "nftables.service" ]; + wantedBy = [ "multi-user.target" ]; + serviceConfig = { + Type = "oneshot"; + ExecStart = "${xlabWg}/bin/xlab-wg health all"; + StateDirectory = "xlab-wireguard"; + StateDirectoryMode = "0750"; + CapabilityBoundingSet = [ "CAP_NET_ADMIN" "CAP_NET_RAW" ]; + NoNewPrivileges = true; + PrivateTmp = true; + ProtectHome = true; + ProtectKernelModules = true; + ProtectSystem = "strict"; + RestrictAddressFamilies = [ "AF_INET" "AF_INET6" "AF_NETLINK" "AF_UNIX" ]; + }; + }; + + systemd.timers.xlab-wireguard-readiness = { + description = "Periodically verify xlab WireGuard netdevs"; + wantedBy = [ "timers.target" ]; + timerConfig = { + OnBootSec = "30s"; + OnUnitInactiveSec = "2min"; + AccuracySec = "15s"; + RandomizedDelaySec = "10s"; + Persistent = true; + }; + }; + # Keep a local APNIC-CN destination set so xlab clients use xlab's own # campus uplink for domestic traffic instead of hairpinning through .1. # The cache is loaded before any network refresh, then refreshed daily. @@ -599,19 +916,103 @@ }; }; + # Verify the complete client path independently of the loader. This catches + # stale registry data and policy loss after either nftables or networkd is + # reloaded, even when the population unit itself still reports success. + systemd.services.xlab-cn-direct-health = { + description = "Verify xlab CN-direct policy and foreign-tunnel fallback"; + wants = [ "xlab-cn-direct-sets.service" ]; + after = [ + "nftables.service" + "systemd-networkd.service" + "xlab-cn-direct-sets.service" + ]; + path = with pkgs; [ coreutils findutils gnugrep iproute2 jq nftables ]; + serviceConfig = { + Type = "oneshot"; + CapabilityBoundingSet = [ "CAP_NET_ADMIN" ]; + NoNewPrivileges = true; + PrivateTmp = true; + ProtectHome = true; + ProtectSystem = "strict"; + RestrictAddressFamilies = [ "AF_INET" "AF_INET6" "AF_NETLINK" "AF_UNIX" ]; + }; + script = '' + set -euo pipefail + cache=/var/lib/xlab-cn-direct/delegated-apnic-extended-latest + + fail() { + echo "CN-direct health check: $*" >&2 + exit 1 + } + + [[ -s "$cache" ]] || fail "validated APNIC cache is missing" + [[ -n "$(find "$cache" -mmin -2880 -print -quit)" ]] || fail "APNIC cache is older than 48 hours" + + set_count() { + nft -j list set inet xlab_pbr "$1" \ + | jq '[.nftables[] | .set?.elem? // empty | .[]] | length' + } + + v4_count=$(set_count cn_direct_v4) + v6_count=$(set_count cn_direct_v6) + regional_v4_count=$(set_count regional_direct_v4) + regional_v6_count=$(set_count regional_direct_v6) + ((v4_count >= 8000)) || fail "IPv4 set is undersized: $v4_count" + ((v6_count >= 1500)) || fail "IPv6 set is undersized: $v6_count" + ((regional_v4_count >= 10)) || fail "regional IPv4 set is undersized: $regional_v4_count" + ((regional_v6_count >= 3)) || fail "regional IPv6 set is undersized: $regional_v6_count" + + ip -4 rule show | grep -Eq '^90:.*fwmark 0x10.*lookup main' \ + || fail "IPv4 mark rule is missing" + ip -6 rule show | grep -Eq '^90:.*fwmark 0x10.*lookup main' \ + || fail "IPv6 mark rule is missing" + + ip -4 route get 223.5.5.5 from 10.253.254.2 iif bond.lan254 mark 0x10 \ + | grep -q 'dev wan99.0' || fail "marked CN IPv4 does not use the campus WAN" + ip -6 route get 2400:3200::1 from fd99:23eb:1682:1::2 iif bond.lan254 mark 0x10 \ + | grep -q 'dev wan99.0' || fail "marked CN IPv6 does not use the campus WAN" + ip -4 route get 47.246.1.1 from 10.253.254.2 iif bond.lan254 mark 0x10 \ + | grep -q 'dev wan99.0' || fail "Alibaba exception does not use the campus WAN" + ip -4 route get 1.1.1.1 from 10.253.254.2 iif bond.lan254 \ + | grep -q 'dev wg-to-wgnet' || fail "foreign IPv4 does not use WireGuard" + ip -6 route get 2606:4700:4700::1111 from fd99:23eb:1682:1::2 iif bond.lan254 \ + | grep -q 'dev wg-to-wgnet' || fail "foreign IPv6 does not use WireGuard" + + echo "CN-direct healthy: APNIC v4=$v4_count v6=$v6_count; regional v4=$regional_v4_count v6=$regional_v6_count" + ''; + }; + + systemd.timers.xlab-cn-direct-health = { + description = "Periodically verify xlab client routing policy"; + wantedBy = [ "timers.target" ]; + timerConfig = { + OnBootSec = "10min"; + OnUnitInactiveSec = "15min"; + RandomizedDelaySec = "2min"; + Persistent = true; + }; + }; + # =========================================================================== # DNS RESOLUTION # =========================================================================== # Route DNS through the network's local mosdns at 10.0.0.1 so this host # inherits CN-aware split routing and the analytics blocking policy. - # Cloudflare is retained as fallback in case 10.0.0.1 is unreachable - # (bootstrap, maintenance, or partial outage). + # Do not add a mainland resolver as a same-priority "secondary": resolved + # load-balances global servers rather than treating their order as failover, + # so polluted foreign answers intermittently reach clients. Both addresses + # below terminate on main's filtered MosDNS over the primary WireGuard path. + # WireGuard endpoints are numeric, so tunnel recovery has no DNS dependency. services.resolved = { enable = true; - fallbackDns = [ "1.1.1.1" "2606:4700:4700::1111" ]; - extraConfig = '' - DNS=10.0.0.1 - ''; + settings.Resolve = { + DNS = [ "10.0.0.1" "fd99:23eb:1682::1" ]; + FallbackDNS = ""; + DNSStubListenerExtra = [ "10.253.254.1" "fd99:23eb:1682:1::1" ]; + LLMNR = false; + MulticastDNS = false; + }; }; # =========================================================================== diff --git a/modules/common.nix b/modules/common.nix index d03850a..7058779 100644 --- a/modules/common.nix +++ b/modules/common.nix @@ -7,7 +7,9 @@ nix.settings = { experimental-features = [ "nix-command" "flakes" ]; auto-optimise-store = true; - trusted-users = [ "root" "ldx" "ye-lw21" ]; + # Nix trusted users can import/store arbitrary derivations as root. + # Interactive sudo access does not justify a second root-equivalent path. + trusted-users = [ "root" "ldx" ]; substituters = [ "https://mirrors.tuna.tsinghua.edu.cn/nix-channels/store" "https://cache.nixos.org/" diff --git a/modules/influxdb.nix b/modules/influxdb.nix index be01679..4089797 100644 --- a/modules/influxdb.nix +++ b/modules/influxdb.nix @@ -28,6 +28,12 @@ default = 8086; description = "InfluxDB HTTP API listen port"; }; + + openFirewall = lib.mkOption { + type = lib.types.bool; + default = true; + description = "Whether to open the InfluxDB port without a source restriction"; + }; }; config = lib.mkIf cfg.enable { @@ -58,7 +64,6 @@ wants = [ "zfs-mount.service" ]; }; - # Open firewall for internal network only - networking.firewall.allowedTCPPorts = [ cfg.port ]; + networking.firewall.allowedTCPPorts = lib.optional cfg.openFirewall cfg.port; }; } diff --git a/modules/monitoring.nix b/modules/monitoring.nix index 2a7f7b5..40abe6f 100644 --- a/modules/monitoring.nix +++ b/modules/monitoring.nix @@ -82,6 +82,11 @@ type = lib.types.port; default = 9100; }; + openFirewall = lib.mkOption { + type = lib.types.bool; + default = true; + description = "Whether to open the exporter port without a source restriction"; + }; }; }; @@ -96,6 +101,12 @@ systemd.services.telegraf.serviceConfig.EnvironmentFile = config.age.secrets.influxdb-token.path; + # A changed age payload updates the runtime file in place; make the unit + # transition consume the new token without requiring an operator restart. + systemd.services.telegraf.restartTriggers = [ + config.age.secrets.influxdb-token.file + ]; + systemd.services.telegraf.path = [ "/run/wrappers" pkgs.lm_sensors pkgs.smartmontools pkgs.nvme-cli ]; services.telegraf = { @@ -293,7 +304,7 @@ }; }; - networking.firewall.allowedTCPPorts = lib.mkIf cfg.nodeExporter.enable + networking.firewall.allowedTCPPorts = lib.mkIf (cfg.nodeExporter.enable && cfg.nodeExporter.openFirewall) [ cfg.nodeExporter.port ]; }; } diff --git a/secrets/secrets.nix b/secrets/secrets.nix index 4ec6c6f..516cbd7 100644 --- a/secrets/secrets.nix +++ b/secrets/secrets.nix @@ -15,6 +15,7 @@ "xlab-wg-wgnet.age".publicKeys = admins ++ [ xlab-gateway ]; "xlab-wg-wgnet-psk.age".publicKeys = admins ++ [ xlab-gateway ]; "xlab-wg-warp.age".publicKeys = admins ++ [ xlab-gateway ]; + "xlab-ddns-tsig.age".publicKeys = admins ++ [ xlab-gateway ]; "influxdb-token.age".publicKeys = admins ++ [ skydick ]; "skydick-ldap-bind.age".publicKeys = admins ++ [ skydick ]; "skydick-samba-ldap-admin.age".publicKeys = admins ++ [ skydick ]; diff --git a/secrets/xlab-ddns-tsig.age b/secrets/xlab-ddns-tsig.age new file mode 100644 index 0000000..f955fba --- /dev/null +++ b/secrets/xlab-ddns-tsig.age Binary files differ