Scope: the skyworks NixOS configurations for xlab-gateway andskydick, the live main gateway (door1, 10.0.0.1:2222) and itsskynet-server-gateway repository, the LDAP deployment inskynet-server-web, and the source policy in nix-infra.
[!CAUTION]
Do not activate the staged SkyDick generation outside a coordinated storage
window. The host currently has 33 established NFS/TCP channels from node1,
node2, and main, including open MySQL data files, plus a logged-in 8 TiB
iSCSI client at10.0.200.11. Dry activation would restart ZFS mount/share,
NFS, Samba, InfluxDB, and NSS services.
| System | State | Result |
|---|---|---|
| Main gateway | LIVE + COMMITTED | Region policy, PotPlayer DNS exception, IPv6 DNS, TSIG rotation, SSH/input hardening, and tunnel CAKE lifecycle are deployed. Repository is clean. |
| Main LDAP stack | LIVE + COMMITTED | LDAP and Luminary credentials use Docker secrets; both containers are healthy. Commit d507c33. |
| Xlab | LIVE + COMMITTED, AUTH PENDING | Both WireGuard paths, automatic recovery, persistent manual disable, CN-direct routing, trusted DNS, DHCP/RA, and restricted public recovery access are live. Polluted AliDNS fallback was removed. Cold-boot campus admission still needs registered-IP status or a dedicated credential. |
| C9800 WLAN | LIVE + SAVED | Skyworks remains WPA3-SAE with PMF required; OKC and band-select are disabled to avoid the observed FlexConnect SAE PMKID reassociation failure. |
| SkyDick | STAGED + BUILT | NixOS 25.11 remains live and healthy. The final committed 26.05 closure is built, not activated because storage clients are active. |
The source is nix-infra commit8692bebcfaf59be065d84d064c367cb01dddc691 (mesh/proxy: force Alibaba overseas CDN direct for intl-pinned clients). It contains packet-capture
evidence for Taobao/Tmall traffic to four Alibaba AS45102 Singapore ranges:
47.246.0.0/16 43.109.0.0/16 163.181.0.0/16 139.95.0.0/16
The commit contains no JD-specific prefix or packet evidence. JD appears only
in a comment, so no unsupported JD range was invented during migration.
5b3cfc5 adds all four ranges to the direct-routing5e1d72f forces daumcdn.net through the proxy path.NXDOMAIN. A/AAAA lookup and the exact installer URL were2898736 serves MosDNS over UDP/TCP on the advertised3b957cd rotates the DDNS TSIG secret, retires the oldwg-sgp-qos.service andwg-sgp-autorate.service are currently active, and CAKE is attached toifb-sgp and wg-outbound-sgp.The deployed xlab policy uses two nft interval-set layers:
Matching LAN packets receive mark 0x10. Rule priority 90 first looks in the
main table, so they leave through wan99.0 while its default route exists. If
that default disappears, lookup continues to the existing source rule for
table 1002 and exits through wg-to-wgnet. The earlier static throw design
did not have this fallback and has been removed.
An isolated network-namespace test loaded the rendered nft rules and proved:
regional_direct_v4=10 regional_direct_v6=3 healthy: 47.246.1.1 -> wan99.0 failed WAN: 47.246.1.1 -> wg-to-wgnet table 1002
Active system and persistent system profile as of 2026-07-22:
/nix/store/1b50bcwxk7dcs5rhgny4p596shc5aabi-nixos-system-xlab-gateway-26.05.20260719.fd14620
Included changes:
system.stateVersion = "25.11".10.253.254.1 andfd99:23eb:1682:1::1; DHCP and RA advertise those addresses. Their only10.0.0.1 and fd99:23eb:1682::1.FallbackDNS is empty so resolved cannot select an unfiltered mainland166.111.17.108) or server3 (166.111.17.81).wireguard module loading and a handshake-aware recovery timer.xlab-wg waits forxlab-wg disable keep eitherwg-quick forRendered resolved.conf, Kea DHCP4, radvd, networkd, and nftables files were
inspected. radvd --configtest, nftables --check, both generated health
scripts with bash -n, and the complete NixOS build passed. Kea parsed the
configuration through socket selection; its only test-host error was the
expected absence of xlab's bond.lan254 on the remote build machine.
The failed boot journal disproved the initial missing-netdev hypothesis. Both
WireGuard netdevs and their agenix secrets existed, but networkd raised the
tunnels before the WAN route, DNS, and campus admission were usable. The
server3 endpoint hostname then failed resolution continuously for roughly 12
hours. The old firewall did not admit inbound UDP 46961/51998, and the old
generation had no handshake-aware recovery task, so peer-initiated traffic
could not break the bootstrap failure.
The campus authentication endpoint reports that the current admission session
started at 15:28:44 CST. wg-to-wgnet was manually raised at 15:29:37 and
handshook immediately. Xlab has no Tunet/GoAuthing unit or stored campus
credential, so the missing cold-boot admission was a necessary part of the
outage, not merely a WireGuard race. The new readiness service retries and
recovers automatically once admission exists, but full cold-boot automation
still requires either registered/open IP status or an agenix-backed dedicated
campus credential.
The preserved evidence remains on xlab at:
/var/tmp/xlab-previous-boot-full-20260722.log /var/tmp/xlab-previous-boot-network-20260722.log /var/tmp/xlab-network-boot-20260722.log
The new closure was activated without rebooting and tested from both the normal
route and restricted public break-glass path. Activation restarted networkd;
the readiness service detected stale handshakes and recovered both tunnels in
about three seconds each. Further fault injection proved:
wg-to-skyworks persists while the health service runs, then anwg-to-skyworks netdev causes it to be recreated,wg-to-wgnet down from the public recovery session causes automatic10.253.254.1 reachability.The deployed CN-direct health check passed with 8,786 IPv4 and 2,039 IPv6 APNIC
prefixes plus 10/3 static regional prefixes. Marked 223.5.5.5 and47.246.1.1 route through wan99.0; unmarked 1.1.1.1 routes throughwg-to-wgnet. No systemd units are failed.
Xlab's resolved configuration listed main MosDNS and AliDNS as peer global
servers. systemd-resolved does not treat that list as ordered failover; it had
selected 223.5.5.5, which returned non-Google addresses forwww.google.com and similarly polluted YouTube names. This explained the
reported split symptom: google.com resolved and pinged, but its HTTPS redirect
to www.google.com failed. Path-MTU probes, symmetric MSS clamps, and real
client TCP flows were healthy.
The live and profiled configuration now uses only the two main MosDNS tunnel
addresses. After removing the temporary runtime override and restarting
resolved, repeated www.google.com queries returned Google A/AAAA records and
an interface-bound HTTPS request returned HTTP 200. No systemd unit is failed.
The controller is a C9800-CL running IOS XE 17.15.5; the xlab AP is an
AIR-AP3802I-H in FlexConnect mode. Authoritative post-change show wlan id 1
output confirms that Skyworks is WPA3-SAE-only with CCMP, PMF required,
H2E plus HNP, and FT disabled. The authentication mode itself was not changed
during this recovery.
Always-on WLC trace provided a concrete interoperability failure for a
MacBook randomized address. It completed SAE on 5 GHz, received DHCP, moved to
2.4 GHz, then reassociated to 5 GHz and was deleted with:
SAE PMKID matching failed in roam case Invalid PMKID IE_VALIDATION_FAILURE type 53
Another failing randomized client associated on 5 GHz but never completed the
SAE exchange and timed out in L2AUTH. This matches the FlexConnect/WPA3-SAE/OKC
failure family in
Cisco CSCwp20425;
Cisco's documented workaround is to disable OKC. Enabled band-select increased
the chance of cross-band reassociation. The controller is already on Cisco's
currently recommended 17.15.5 release,
so there is no justified in-train firmware upgrade to substitute for the
verified workaround.
The WLC running configuration was backed up tobootflash:pre-skyworks-auth-fix-20260722.cfg. OKC and band-select were then
disabled on WLAN 1, the WLAN was re-enabled, all three existing Skyworks SAE
clients returned to Run, and write memory completed. The previously
failing MacBook then rejoined on 5 GHz using SAE, remained associated well past
its former five-second failure point, obtained 10.253.254.103, and exchanged
traffic without decrypt or policy errors. Six Skyworks SAE clients were inRun concurrently with no excluded clients. WPA2 transition mode remains an
available compatibility fallback, but was not needed and would reduce the
current security level.
Final build:
/nix/store/x3b5rnxpp37icqjr0ja91l2mhhm73igz-nixos-system-skydick-26.05.20260719.fd14620
Included changes:
/var/lib/target, with /etc/target as atargetcli atomic saves remain persistent.iscsi-target.service from stopping during a live switch.crossmnt; exact media/backup10.0.0.1 and fd99:23eb:1682::1 are global resolverThe live iSCSI configuration was backed up again at/root/backups/target/20260721T204151Z. Serializing the live kernel state was
byte-identical before and after (babc4f7a...59650c246), and the CHAP-aware
guard passes it. targetcli sessions detail saying NOT AUTHENTICATED refers
to optional mutual target authentication; initiator CHAP fields are present.
Do not activate this closure until node1/node2 MySQL and all other NFS users
are stopped or unmounted and the iSCSI initiator is logged out or explicitly
accepted as maintenance risk.
slapcat backups are under/root/backups/ldap/20260721T200920Z on main.d507c33 moves LDAP admin/config/readonly and Luminary bindpdbedit both pass.The live SkyDick generation still uses plaintext LDAP until its maintenance
deployment. This is a deployment-state issue, not a missing source fix.
open with the unit networkye-lw21, zhuyz24, and /srv/system/vm retain/32 and /128 entries after inventory. Use Kerberos orrdma 20049 twice, but allproto=rdma client and then scope the firewall/export.cn,sambaDomainName, and sambaSID. The database is currently small; adddf99 IPv6 peers and peers withoutxlab-wg rather than wg-quick for its networkd-owned tunnels.