state: the composer replaced the placeholder; open the public ingress question
PROCESS.md section 8 requires host changes to reach STAGING-STATE.md before a session ends. mechcomp-placeholder.service was disabled and stopped on 2026-09-11 and mechcomp.service took 10.20.0.10:8770 in its place. nginx on CT 101 needed no change because the real service took the address the placeholder occupied. The placeholder unit stays on disk, disabled, as the rollback. Section 5 items closed: mechcomp.service, and replacing the placeholder. Application runtime acceptance is marked partial rather than done, because no acceptance criteria have been written for the composer and it is not reachable from outside. WORK-ORDER-004 publishes the composer at dev.mechcomp.kane-il.us. The srv-b half needs no change, verified: forwarding on, FORWARD policy ACCEPT, a direct route on vmbr1, and none of the three FORWARD rules matches hub initiated inbound traffic. The single gate is AllowedIPs on the hub peer entry for srv-b, because WireGuard drops by cryptokey routing before consulting any routing table. Design decision recorded: the public name terminates on the hub and proxies to CT 100 directly rather than through CT 101. Routing it through CT 101 would make the public path depend on a locally signed leaf that expires 2028-11-18 with nothing renewing it. The tunnel already provides the encryption that hop would add. CT 101 keeps serving the internal name. Section 4 not now decision on the hub route is reopened. It was correct while there was no application to reach. The consequence of leaving it closed is that the operator cannot see the application at all.
This commit is contained in:
@@ -0,0 +1,226 @@
|
||||
# WORK-ORDER-004 — public ingress for the composer
|
||||
|
||||
Publish the Mechanical Compiler composer at `dev.mechcomp.kane-il.us`, reachable
|
||||
from an ordinary browser on the internet.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Mode | Infrastructure (`PROCESS.md` §2) |
|
||||
| Created | 2026-09-11 |
|
||||
| Executed on | `wg-pk` (the WireGuard hub), via its Webmin terminal |
|
||||
| Depends on | `14514b0` — composer live on CT 100, proven through CT 101 |
|
||||
|
||||
---
|
||||
|
||||
## 0. Read this before the first command
|
||||
|
||||
**This work order is executed on a machine outside this project.** `wg-pk`
|
||||
carries `kane-il.us` mail and Hubzilla. `PROCESS.md` §7 requires escalation for
|
||||
exactly this, and the escalation is: CIVICVS executes, one group at a time, and
|
||||
anything surprising stops the work rather than being worked around.
|
||||
|
||||
### The lockout risk, and why it is bounded
|
||||
|
||||
Step 2 changes a live WireGuard peer. If it goes wrong the tunnel drops, and
|
||||
with it every path to `srv-b` — `PROCESS.md` §1 records that the `srv-b` shell
|
||||
is the operator's entire working surface for this project.
|
||||
|
||||
It is recoverable because **the operator is sitting on `wg-pk` itself**, in
|
||||
Webmin, not reaching it through the tunnel. `wg set` is not persistent, so
|
||||
`systemctl restart wg-quick@wg0` on the hub restores the on-disk configuration
|
||||
and the tunnel with it.
|
||||
|
||||
Do not perform step 2 from a shell that reaches the hub through the tunnel.
|
||||
|
||||
### What `srv-b` contributes
|
||||
|
||||
**Nothing. It is already correct and must not be touched.** Verified
|
||||
2026-09-11:
|
||||
|
||||
```
|
||||
net.ipv4.ip_forward = 1
|
||||
-P FORWARD ACCEPT
|
||||
10.20.0.10 dev vmbr1 src 10.20.0.1
|
||||
```
|
||||
|
||||
The three `FORWARD` rules block container-sourced SMTP, container-to-container
|
||||
(ACCEPT), and container-to-LAN. None matches hub-initiated inbound traffic. The
|
||||
`-s 10.20.0.0/24 -o wg0 MASQUERADE` rule does not apply either: the first packet
|
||||
of a hub-initiated flow is not container-sourced, so no NAT binding is created
|
||||
and replies return through conntrack.
|
||||
|
||||
**Explicit do-not-touch list:** `srv-b` iptables, `srv-b` WireGuard, CT 101
|
||||
nginx, CT 101 TLS, the `kane-il.us` MX records, anything on
|
||||
`mx1.diagnostics.kane-il.us` (it publishes a TLSA record — confirmed
|
||||
2026-09-11).
|
||||
|
||||
---
|
||||
|
||||
## 1. Topology, and why
|
||||
|
||||
```
|
||||
browser
|
||||
-> DNS dev.mechcomp.kane-il.us -> wg-pk public address
|
||||
-> nginx on wg-pk, Let's Encrypt TLS terminated here
|
||||
-> WireGuard tunnel, already encrypted
|
||||
-> srv-b 10.110.0.12, forwards, no configuration change
|
||||
-> CT 100 10.20.0.10:8770, the composer
|
||||
```
|
||||
|
||||
**The public path deliberately bypasses CT 101.** Routing it through CT 101
|
||||
would make the public name depend on a locally-signed leaf valid to 2028-11-18
|
||||
with nothing renewing it — a dated outage designed in from the start. The tunnel
|
||||
already provides the encryption that hop would add.
|
||||
|
||||
CT 101 continues to serve `mechanical-compiler.dev.infra` for work from inside.
|
||||
Two ingresses, each with a distinct reason to exist.
|
||||
|
||||
---
|
||||
|
||||
## 2. Success criteria, stated before the work
|
||||
|
||||
1. `curl -I https://dev.mechcomp.kane-il.us/` returns `200` from a machine with
|
||||
no WireGuard access and no special DNS.
|
||||
2. The certificate is issued by Let's Encrypt and chains without `-k`.
|
||||
3. `https://dev.mechcomp.kane-il.us/api/build?family=3x&profile=Y` returns JSON
|
||||
whose `groups` end with a `Y` group and contain no `three_fin` key.
|
||||
4. The page renders, controls change the drawing, and a changed control changes
|
||||
the `design` identity shown beneath it.
|
||||
5. **The negative:** `mechanical-compiler.dev.infra` still answers from `srv-b`,
|
||||
and mail from `srv-b` still delivers. Neither path was in scope; both must be
|
||||
proven unharmed.
|
||||
6. Every change survives `systemctl restart wg-quick@wg0` and a hub reboot.
|
||||
|
||||
---
|
||||
|
||||
## 3. Command groups
|
||||
|
||||
One at a time. Paste output back before the next.
|
||||
|
||||
### Group 1 — read-only, on `wg-pk`
|
||||
|
||||
Nothing here changes anything.
|
||||
|
||||
```bash
|
||||
echo "=== peer entry for srv-b ===" && \
|
||||
wg show && \
|
||||
echo "=== on-disk wireguard config ===" && \
|
||||
grep -n "AllowedIPs\|PublicKey\|Address\|PostUp" /etc/wireguard/wg0.conf && \
|
||||
echo "=== routing toward the service network ===" && \
|
||||
ip route | grep -E "10\.20\.|10\.110\." ; \
|
||||
echo "=== forwarding ===" && \
|
||||
sysctl net.ipv4.ip_forward && \
|
||||
echo "=== nginx vhost conventions, following symlinks ===" && \
|
||||
grep -rn --dereference-recursive "server_name\|proxy_pass\|listen\|ssl_certificate " \
|
||||
/etc/nginx/sites-enabled/ | head -40 && \
|
||||
echo "=== certificate issuance ===" && \
|
||||
ls /etc/letsencrypt/live 2>/dev/null || echo "no certbot live dir" ; \
|
||||
which certbot ; \
|
||||
echo "=== does the name resolve yet ===" && \
|
||||
dig +short dev.mechcomp.kane-il.us A ; \
|
||||
dig +short dev.mechcomp.kane-il.us AAAA
|
||||
```
|
||||
|
||||
What matters in the output: whether the peer's `AllowedIPs` is the only gate,
|
||||
whether nginx already has `listen [::]:443` on its vhost pattern (it must, if
|
||||
the name inherits an AAAA), how certificates are issued, and whether the name
|
||||
resolves yet.
|
||||
|
||||
**Report this before proceeding.** Groups 2 onward are written against what it
|
||||
shows; they are deliberately not drafted in advance.
|
||||
|
||||
### Group 2 — WireGuard AllowedIPs
|
||||
|
||||
Drafted after group 1. The shape:
|
||||
|
||||
- Back up `/etc/wireguard/wg0.conf` first. `PROCESS.md` §2: preserve a rollback
|
||||
copy before editing any configuration file.
|
||||
- Add `10.20.0.0/24` to the `srv-b` peer's `AllowedIPs`, **keeping
|
||||
`10.110.0.0/22`**. Replacing rather than extending it is the way this breaks.
|
||||
- Apply live, confirm the tunnel is still up and `srv-b` still reachable, then
|
||||
persist.
|
||||
- Prove: `ping -c2 10.20.0.10` from the hub.
|
||||
|
||||
### Group 3 — route
|
||||
|
||||
A route for `10.20.0.0/24` via `10.110.0.12`, persisted the way the hub already
|
||||
persists routes — which group 1 reveals. Prove with
|
||||
`curl -sS -o /dev/null -w '%{http_code}\n' http://10.20.0.10:8770/`.
|
||||
|
||||
### Group 4 — certificate
|
||||
|
||||
Let's Encrypt HTTP-01 for `dev.mechcomp.kane-il.us`. A third-level name needs no
|
||||
wildcard and no DNS-01. Requires the name to resolve to the hub first.
|
||||
|
||||
### Group 5 — vhost
|
||||
|
||||
```nginx
|
||||
# /etc/nginx/sites-available/dev.mechcomp.kane-il.us
|
||||
# Matching whatever pattern group 1 shows the existing vhosts use.
|
||||
server {
|
||||
listen 80;
|
||||
listen [::]:80;
|
||||
server_name dev.mechcomp.kane-il.us;
|
||||
return 301 https://$host$request_uri;
|
||||
}
|
||||
|
||||
server {
|
||||
listen 443 ssl;
|
||||
listen [::]:443 ssl; # required if the name has an AAAA
|
||||
server_name dev.mechcomp.kane-il.us;
|
||||
|
||||
ssl_certificate /etc/letsencrypt/live/dev.mechcomp.kane-il.us/fullchain.pem;
|
||||
ssl_certificate_key /etc/letsencrypt/live/dev.mechcomp.kane-il.us/privkey.pem;
|
||||
|
||||
location / {
|
||||
proxy_pass http://10.20.0.10:8770;
|
||||
proxy_set_header Host $host;
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
proxy_set_header X-Forwarded-Proto $scheme;
|
||||
|
||||
# The composer redraws on every control change. Small responses, many of
|
||||
# them; buffering adds latency for no benefit.
|
||||
proxy_buffering off;
|
||||
proxy_read_timeout 30s;
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`nginx -t` before reload, always.
|
||||
|
||||
### Group 6 — acceptance
|
||||
|
||||
All six criteria in §2, including the two negatives.
|
||||
|
||||
---
|
||||
|
||||
## 4. Known unknowns
|
||||
|
||||
Stated rather than assumed, because assuming is what produced three unreachable
|
||||
URLs before this document existed.
|
||||
|
||||
- **Whether the hub's nginx listens on v6.** If `dev.mechcomp` is a CNAME to a
|
||||
name carrying an AAAA, v6 is published whether or not nginx answers on it.
|
||||
v6-preferring browsers then fail while v4 ones succeed — a fault that looks
|
||||
like anything except DNS.
|
||||
- **How the hub persists routes.** Distribution- and tooling-dependent; group 1
|
||||
shows it.
|
||||
- **Whether Webmin or YunoHost manages the nginx configuration.** If so,
|
||||
hand-written vhosts may be overwritten on their next reconfiguration, and the
|
||||
vhost belongs wherever that system expects it instead.
|
||||
|
||||
---
|
||||
|
||||
## 5. Acceptance and reporting
|
||||
|
||||
Closed when all six criteria in §2 hold and survive a hub reboot.
|
||||
|
||||
Back to the architect: the output of every group, and any surprise about the
|
||||
environment in `FAILURES.md` format — `PROCESS.md` §7. A surprise on the hub is
|
||||
still a finding even though the hub is not this project's property.
|
||||
|
||||
`STAGING-STATE.md` §4 records "direct WireGuard-side access to the catalogue
|
||||
would be a route addition for `10.20.0.0/24` on the hub. Not now." That decision
|
||||
was correct when there was no application to reach. This work order is its
|
||||
reversal, and §4 should say so when this closes.
|
||||
Reference in New Issue
Block a user