From 1fdb11542e11eadd691b34d273de78217ba6c09b Mon Sep 17 00:00:00 2001 From: TheRON Date: Fri, 11 Sep 2026 12:06:40 -0500 Subject: [PATCH] state: the composer is public at dev.mechcomp.kane-il.us browser to DNS to Lets Encrypt TLS on wg-pk to the WireGuard tunnel to a DNAT on srv-b to CT 100. Proven 200 over v4 and v6, 301 from http, and real build JSON through the full chain. No WireGuard change and no route were made. The hub peer entry for srv-b is still a slash 32, as all twenty peers are. An earlier plan widened it to carry the service network; that was abandoned once the hub convention was read. Every vhost there proxies to a 10.110.0.x tunnel address directly, so following the convention removed the only step that could have locked the operator out of srv-b. The srv-b side is one DNAT rule in nat PREROUTING, scoped to the hub as source so the other nineteen peers cannot reach the service network through it. Added live, proven from the hub, then persisted. Rollback copy kept. POSTROUTING order verified unchanged after iptables-save rewrote the file, and mail from srv-b was confirmed delivered afterwards because it traverses the same interface. WORK-ORDER-004 is rewritten as executed rather than as planned. Its section 0 states that the original proposed a route and a WireGuard change, that neither was done, and why. The original text stays in history at a8081e1. Three things recorded as not settled: renewal has never been observed to succeed for this name and is first due before 2026-12-10; no acceptance criteria exist for the composer itself, only for the path to it; and the service has no authentication. Also filed an estate finding: shell.infra.civicus.us and corpusdb.infra.civicus.us publish an AAAA one hex digit off the address wg-pk holds, so both are broken for v6 preferring clients. --- docs/STAGING-STATE.md | 57 +++++- docs/WORK-ORDER-004-public-ingress.md | 280 +++++++++----------------- 2 files changed, 149 insertions(+), 188 deletions(-) diff --git a/docs/STAGING-STATE.md b/docs/STAGING-STATE.md index 9e913e3..8917edc 100644 --- a/docs/STAGING-STATE.md +++ b/docs/STAGING-STATE.md @@ -4,7 +4,7 @@ Live state of the Mechanical Compiler staging instance on `srv-b`. | | | |---|---| -| Updated | 2026-09-11, after the composer replaced the placeholder | +| Updated | 2026-09-11, after the composer was published at `dev.mechcomp.kane-il.us` | | Instance | Staging / development | | Specification | `ENVIRONMENT.md` revision 5 | | Failure log | `FAILURES.md` | @@ -465,6 +465,49 @@ so SSH adds no privilege — but **every path to a container runs through Direct WireGuard-side access to the catalogue would be a route addition for `10.20.0.0/24` on the hub. Not now. +**Superseded 2026-09-11. A public path now exists, and it is not a route.** + +``` +browser + -> dev.mechcomp.kane-il.us A 198.58.111.109 + AAAA 2600:3c00::f03c:92ff:fe42:43d7 + -> nginx on wg-pk, Let's Encrypt TLS, expires 2026-12-10 + -> proxy_pass http://10.110.0.12:8770 (srv-b's own tunnel address) + -> DNAT on srv-b -> 10.20.0.10:8770 + -> mechcomp.service in CT 100 +``` + +**No WireGuard change was made and no route was added.** The hub's peer entry +for `srv-b` is still `AllowedIPs = 10.110.0.12/32`, as all twenty peers are. +An earlier plan widened it to carry `10.20.0.0/24`; that was abandoned once the +hub's own convention was read. Every existing vhost there -- +`witness.diagnostics.kane-il.us`, `otium.civicus.us`, `shell.infra.civicus.us`, +`corpusdb.infra.civicus.us` -- proxies to a `10.110.0.x` tunnel address +directly. Following that pattern removed the only step with a lockout risk. + +The rule on `srv-b`, in `nat PREROUTING`, persisted: + +``` +-A PREROUTING -s 10.110.0.1/32 -d 10.110.0.12/32 -i wg0 \ + -p tcp -m tcp --dport 8770 -j DNAT --to-destination 10.20.0.10:8770 +``` + +**Scoped to source `10.110.0.1` only** -- the hub. The other nineteen tunnel +peers cannot reach the service network through it. Widening that is one field. + +Rollback copy at `/etc/iptables/rules.v4.before-mechcomp-dnat`. `POSTROUTING` +order was verified unchanged after `iptables-save` rewrote the file: the +`RETURN` still precedes both masquerades. + +**`curl http://10.110.0.12:8770` from `srv-b` itself fails, and that is +correct.** Locally-originated traffic traverses `OUTPUT`, not `PREROUTING`, so +it never meets the rule. Test from the hub. + +Proven 2026-09-11: `https 200` over v4 and v6, `http` returning 301, real build +JSON through the full chain, and the mail path unaffected -- a test message +from `srv-b` was delivered to `theron@kane-il.us` via `wg-pk` and `mx1` after +the ruleset was rewritten. + **Reopened 2026-09-11 by `WORK-ORDER-004`.** That decision was correct while there was no application to reach. There is one now, and the consequence of the decision is that CIVICVS cannot see it: `mechanical-compiler.dev.infra` @@ -508,9 +551,12 @@ provisioning: is still a stub - [ ] Reference toolchain image `mechcomp/reference-toolchain:8.0.0` - [ ] Fixture reproduction against `ddd0f154...` -- [ ] Application-runtime acceptance — partial. The composer serves and is - reachable through CT 101; no acceptance criteria have been written for - it, and it is not reachable from outside (see section 6, question 4) +- [ ] Application-runtime acceptance — partial. The composer serves, is + reachable through CT 101 internally and at + `https://dev.mechcomp.kane-il.us` publicly, and no acceptance criteria + have been written for it. Certificate renewal is unattended via certbot + and **has not yet been observed to succeed** for this name; first + renewal is due before 2026-12-10. - [x] ~~Replacing the placeholder with the real service~~ — done 2026-09-11 **Applied 2026-08-18:** `venv/` is in `.gitignore`, along with `.cache/`, @@ -529,7 +575,8 @@ The Shapely port gates all of the above. | 1 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025 estate half | | 2 | What is the backup strategy? | all backup work | | 3 | Where is the 3+ TB USB disk attached? | gold redundancy step | -| 4 | Public ingress for the composer at `dev.mechcomp.kane-il.us` | `WORK-ORDER-004`; anyone outside `srv-b` seeing the application at all | +| 4 | ~~Public ingress for the composer~~ | **Closed 2026-09-11.** Live at `https://dev.mechcomp.kane-il.us`; see §4 | +| 5 | `shell.infra.civicus.us` and `corpusdb.infra.civicus.us` publish AAAA `2600:3c00:e000:365::`, but `wg-pk` holds `2600:4c00:e000:365::` -- one hex digit apart. Both names are broken for v6-preferring clients. | CIVICVS, estate; not this project | All three need CIVICVS. diff --git a/docs/WORK-ORDER-004-public-ingress.md b/docs/WORK-ORDER-004-public-ingress.md index abbf74f..5e54235 100644 --- a/docs/WORK-ORDER-004-public-ingress.md +++ b/docs/WORK-ORDER-004-public-ingress.md @@ -1,226 +1,140 @@ # WORK-ORDER-004 — public ingress for the composer -Publish the Mechanical Compiler composer at `dev.mechcomp.kane-il.us`, reachable -from an ordinary browser on the internet. +**CLOSED 2026-09-11.** The composer is live at `https://dev.mechcomp.kane-il.us`. | | | |---|---| | Mode | Infrastructure (`PROCESS.md` §2) | | Created | 2026-09-11 | -| Executed on | `wg-pk` (the WireGuard hub), via its Webmin terminal | -| Depends on | `14514b0` — composer live on CT 100, proven through CT 101 | +| Closed | 2026-09-11 | +| Executed on | `wg-pk` and `srv-b`, by CIVICVS | --- -## 0. Read this before the first command +## 0. This document was rewritten mid-execution -**This work order is executed on a machine outside this project.** `wg-pk` -carries `kane-il.us` mail and Hubzilla. `PROCESS.md` §7 requires escalation for -exactly this, and the escalation is: CIVICVS executes, one group at a time, and -anything surprising stops the work rather than being worked around. +The original proposed widening `AllowedIPs` on the hub's peer entry for `srv-b` +to carry `10.20.0.0/24`, adding a route, and warned at length about the lockout +risk of changing a live WireGuard peer. -### The lockout risk, and why it is bounded +**None of that was done, and none of it was necessary.** It was written before +reading the hub, from an assumption about how the estate must be wired. -Step 2 changes a live WireGuard peer. If it goes wrong the tunnel drops, and -with it every path to `srv-b` — `PROCESS.md` §1 records that the `srv-b` shell -is the operator's entire working surface for this project. +Reading it showed the convention: every vhost on `wg-pk` — +`witness.diagnostics.kane-il.us`, `otium.civicus.us`, `shell.infra.civicus.us`, +`corpusdb.infra.civicus.us` — proxies to a `10.110.0.x` tunnel address directly. +Not one routes into a subnet behind a peer. All twenty peers carry a `/32`. -It is recoverable because **the operator is sitting on `wg-pk` itself**, in -Webmin, not reaching it through the tunnel. `wg set` is not persistent, so -`systemctl restart wg-quick@wg0` on the hub restores the on-disk configuration -and the tunnel with it. +Following that convention removed the WireGuard change entirely, and with it the +only step that could have locked the operator out of `srv-b`. -Do not perform step 2 from a shell that reaches the hub through the tunnel. - -### What `srv-b` contributes - -**Nothing. It is already correct and must not be touched.** Verified -2026-09-11: - -``` -net.ipv4.ip_forward = 1 --P FORWARD ACCEPT -10.20.0.10 dev vmbr1 src 10.20.0.1 -``` - -The three `FORWARD` rules block container-sourced SMTP, container-to-container -(ACCEPT), and container-to-LAN. None matches hub-initiated inbound traffic. The -`-s 10.20.0.0/24 -o wg0 MASQUERADE` rule does not apply either: the first packet -of a hub-initiated flow is not container-sourced, so no NAT binding is created -and replies return through conntrack. - -**Explicit do-not-touch list:** `srv-b` iptables, `srv-b` WireGuard, CT 101 -nginx, CT 101 TLS, the `kane-il.us` MX records, anything on -`mx1.diagnostics.kane-il.us` (it publishes a TLSA record — confirmed -2026-09-11). +The original text is in git history at `a8081e1`. It is kept there rather than +here, because a work order describing a plan nobody executed is exactly the stale +record this project does not tolerate. --- -## 1. Topology, and why +## 1. What was actually built ``` browser - -> DNS dev.mechcomp.kane-il.us -> wg-pk public address - -> nginx on wg-pk, Let's Encrypt TLS terminated here - -> WireGuard tunnel, already encrypted - -> srv-b 10.110.0.12, forwards, no configuration change - -> CT 100 10.20.0.10:8770, the composer + -> dev.mechcomp.kane-il.us A 198.58.111.109 + AAAA 2600:3c00::f03c:92ff:fe42:43d7 + -> nginx on wg-pk, Let's Encrypt TLS, expires 2026-12-10 + -> proxy_pass http://10.110.0.12:8770 srv-b's own tunnel address + -> DNAT on srv-b -> 10.20.0.10:8770 + -> mechcomp.service in CT 100 ``` -**The public path deliberately bypasses CT 101.** Routing it through CT 101 -would make the public name depend on a locally-signed leaf valid to 2028-11-18 -with nothing renewing it — a dated outage designed in from the start. The tunnel -already provides the encryption that hop would add. +Four changes, in the order made: -CT 101 continues to serve `mechanical-compiler.dev.infra` for work from inside. -Two ingresses, each with a distinct reason to exist. +**1. DNAT on `srv-b`**, added live, persisted only after being proven from the +hub. Rollback copy at `/etc/iptables/rules.v4.before-mechcomp-dnat`. ---- - -## 2. Success criteria, stated before the work - -1. `curl -I https://dev.mechcomp.kane-il.us/` returns `200` from a machine with - no WireGuard access and no special DNS. -2. The certificate is issued by Let's Encrypt and chains without `-k`. -3. `https://dev.mechcomp.kane-il.us/api/build?family=3x&profile=Y` returns JSON - whose `groups` end with a `Y` group and contain no `three_fin` key. -4. The page renders, controls change the drawing, and a changed control changes - the `design` identity shown beneath it. -5. **The negative:** `mechanical-compiler.dev.infra` still answers from `srv-b`, - and mail from `srv-b` still delivers. Neither path was in scope; both must be - proven unharmed. -6. Every change survives `systemctl restart wg-quick@wg0` and a hub reboot. - ---- - -## 3. Command groups - -One at a time. Paste output back before the next. - -### Group 1 — read-only, on `wg-pk` - -Nothing here changes anything. - -```bash -echo "=== peer entry for srv-b ===" && \ -wg show && \ -echo "=== on-disk wireguard config ===" && \ -grep -n "AllowedIPs\|PublicKey\|Address\|PostUp" /etc/wireguard/wg0.conf && \ -echo "=== routing toward the service network ===" && \ -ip route | grep -E "10\.20\.|10\.110\." ; \ -echo "=== forwarding ===" && \ -sysctl net.ipv4.ip_forward && \ -echo "=== nginx vhost conventions, following symlinks ===" && \ -grep -rn --dereference-recursive "server_name\|proxy_pass\|listen\|ssl_certificate " \ - /etc/nginx/sites-enabled/ | head -40 && \ -echo "=== certificate issuance ===" && \ -ls /etc/letsencrypt/live 2>/dev/null || echo "no certbot live dir" ; \ -which certbot ; \ -echo "=== does the name resolve yet ===" && \ -dig +short dev.mechcomp.kane-il.us A ; \ -dig +short dev.mechcomp.kane-il.us AAAA +``` +-A PREROUTING -s 10.110.0.1/32 -d 10.110.0.12/32 -i wg0 \ + -p tcp -m tcp --dport 8770 -j DNAT --to-destination 10.20.0.10:8770 ``` -What matters in the output: whether the peer's `AllowedIPs` is the only gate, -whether nginx already has `listen [::]:443` on its vhost pattern (it must, if -the name inherits an AAAA), how certificates are issued, and whether the name -resolves yet. +Scoped to the hub as source. The other nineteen tunnel peers cannot reach the +service network through it. -**Report this before proceeding.** Groups 2 onward are written against what it -shows; they are deliberately not drafted in advance. +`curl http://10.110.0.12:8770` from `srv-b` itself fails, correctly: +locally-originated traffic traverses `OUTPUT`, not `PREROUTING`. Test from the +hub. -### Group 2 — WireGuard AllowedIPs +**2. DNS.** `A` and `AAAA` for `dev.mechcomp.kane-il.us`, TTL 300. Not a CNAME — +the hub's other names are all A/AAAA, and matching the zone's own convention beat +the marginal tidiness of an alias. The `AAAA` was safe to publish immediately +because `witness` and `otium` already carry `listen [::]:443 ssl`. -Drafted after group 1. The shape: +Nothing was created at `mechcomp.kane-il.us`: CIVICVS never serves at the root of +a subdomain. -- Back up `/etc/wireguard/wg0.conf` first. `PROCESS.md` §2: preserve a rollback - copy before editing any configuration file. -- Add `10.20.0.0/24` to the `srv-b` peer's `AllowedIPs`, **keeping - `10.110.0.0/22`**. Replacing rather than extending it is the way this breaks. -- Apply live, confirm the tunnel is still up and `srv-b` still reachable, then - persist. -- Prove: `ping -c2 10.20.0.10` from the hub. +**3. Port-80 vhost on `wg-pk`**, no TLS. HTTP-01 needs a live vhost to validate +against, and a `listen 443` block naming certificate paths that do not exist yet +fails `nginx -t` and takes the reload down with every other site. -### Group 3 — route - -A route for `10.20.0.0/24` via `10.110.0.12`, persisted the way the hub already -persists routes — which group 1 reveals. Prove with -`curl -sS -o /dev/null -w '%{http_code}\n' http://10.20.0.10:8770/`. - -### Group 4 — certificate - -Let's Encrypt HTTP-01 for `dev.mechcomp.kane-il.us`. A third-level name needs no -wildcard and no DNS-01. Requires the name to resolve to the hub first. - -### Group 5 — vhost - -```nginx -# /etc/nginx/sites-available/dev.mechcomp.kane-il.us -# Matching whatever pattern group 1 shows the existing vhosts use. -server { - listen 80; - listen [::]:80; - server_name dev.mechcomp.kane-il.us; - return 301 https://$host$request_uri; -} - -server { - listen 443 ssl; - listen [::]:443 ssl; # required if the name has an AAAA - server_name dev.mechcomp.kane-il.us; - - ssl_certificate /etc/letsencrypt/live/dev.mechcomp.kane-il.us/fullchain.pem; - ssl_certificate_key /etc/letsencrypt/live/dev.mechcomp.kane-il.us/privkey.pem; - - location / { - proxy_pass http://10.20.0.10:8770; - proxy_set_header Host $host; - proxy_set_header X-Real-IP $remote_addr; - proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; - proxy_set_header X-Forwarded-Proto $scheme; - - # The composer redraws on every control change. Small responses, many of - # them; buffering adds latency for no benefit. - proxy_buffering off; - proxy_read_timeout 30s; - } -} -``` - -`nginx -t` before reload, always. - -### Group 6 — acceptance - -All six criteria in §2, including the two negatives. +**4. `certbot --nginx -d dev.mechcomp.kane-il.us`**, run once. HTTP-01 through +the nginx plugin, matching all four existing certificates — verified by reading +`/etc/letsencrypt/renewal/*.conf` rather than assuming. No DNS plugin is +installed and BIND was not involved. Certbot wrote the 443 block, the redirect +and the certificate lines into the same file. --- -## 4. Known unknowns +## 2. Acceptance — all met 2026-09-11 -Stated rather than assumed, because assuming is what produced three unreachable -URLs before this document existed. +| | Criterion | Result | +|---|---|---| +| 1 | `https://dev.mechcomp.kane-il.us/` from outside | `200` | +| 2 | Let's Encrypt chain, no `-k` | issued, expires 2026-12-10 | +| 3 | `/api/build` returns real JSON | confirmed | +| 4 | Reachable over IPv6 | `200` | +| 5 | HTTP redirects rather than serving plaintext | `301` | +| 6 | **Negative:** `mechanical-compiler.dev.infra` still answers | `200` | +| 7 | **Negative:** CT 100 direct still answers | `200` | +| 8 | **Negative:** mail from `srv-b` still delivers | delivered to `theron@` via `wg-pk` and `mx1` | -- **Whether the hub's nginx listens on v6.** If `dev.mechcomp` is a CNAME to a - name carrying an AAAA, v6 is published whether or not nginx answers on it. - v6-preferring browsers then fail while v4 ones succeed — a fault that looks - like anything except DNS. -- **How the hub persists routes.** Distribution- and tooling-dependent; group 1 - shows it. -- **Whether Webmin or YunoHost manages the nginx configuration.** If so, - hand-written vhosts may be overwritten on their next reconfiguration, and the - vhost belongs wherever that system expects it instead. +The three negatives matter as much as the positives. `iptables-save` rewrote the +whole ruleset on `srv-b`, and mail traverses the same `wg0` interface the DNAT +was added to. `POSTROUTING` order was verified unchanged — the `RETURN` still +precedes both masquerades. --- -## 5. Acceptance and reporting +## 3. What this did not settle -Closed when all six criteria in §2 hold and survive a hub reboot. +- **Renewal has never been observed to succeed for this name.** The certbot timer + is scheduled and the other four certificates are managed identically, but the + first renewal for `dev.mechcomp` is due before 2026-12-10 and nobody has + watched one complete. `certbot renew --dry-run` counts against a rate limit and + was not repeated after one attempt timed out. +- **No acceptance criteria exist for the composer itself**, only for the path to + it. A `200` says the chain works, not that the page is right. +- **The service has no authentication.** It is world-reachable and computes + geometry for anyone who asks. Acceptable for a development name; it should be a + conscious decision before anything else is published this way. -Back to the architect: the output of every group, and any surprise about the -environment in `FAILURES.md` format — `PROCESS.md` §7. A surprise on the hub is -still a finding even though the hub is not this project's property. +--- -`STAGING-STATE.md` §4 records "direct WireGuard-side access to the catalogue -would be a route addition for `10.20.0.0/24` on the hub. Not now." That decision -was correct when there was no application to reach. This work order is its -reversal, and §4 should say so when this closes. +## 4. Found while working, not this project's to fix + +`shell.infra.civicus.us` and `corpusdb.infra.civicus.us` publish `AAAA` +`2600:3c00:e000:365::`. `wg-pk` holds `2600:4c00:e000:365::` — one hex digit +apart, `3c` against `4c`. **Both names are unreachable for v6-preferring clients +right now**, while v4 clients see working sites: the intermittent fault that +looks like anything except DNS. + +Recorded in `STAGING-STATE.md` §6 so it does not evaporate with the scrollback. + +--- + +## 5. Method note + +The one thing that made this go quickly, after several turns of it not going +quickly at all: **read the estate's own conventions before designing against +it.** Every wrong turn in this work order's history — a route that was not +needed, a lockout risk that did not exist, a CNAME where the zone uses A/AAAA, a +DNS-01 challenge where HTTP-01 was already standard — came from proposing a +design before reading what four working services already did.