state: the composer replaced the placeholder; open the public ingress question

PROCESS.md section 8 requires host changes to reach STAGING-STATE.md before a session ends. mechcomp-placeholder.service was disabled and stopped on 2026-09-11 and mechcomp.service took 10.20.0.10:8770 in its place. nginx on CT 101 needed no change because the real service took the address the placeholder occupied. The placeholder unit stays on disk, disabled, as the rollback.

Section 5 items closed: mechcomp.service, and replacing the placeholder. Application runtime acceptance is marked partial rather than done, because no acceptance criteria have been written for the composer and it is not reachable from outside.

WORK-ORDER-004 publishes the composer at dev.mechcomp.kane-il.us. The srv-b half needs no change, verified: forwarding on, FORWARD policy ACCEPT, a direct route on vmbr1, and none of the three FORWARD rules matches hub initiated inbound traffic. The single gate is AllowedIPs on the hub peer entry for srv-b, because WireGuard drops by cryptokey routing before consulting any routing table.

Design decision recorded: the public name terminates on the hub and proxies to CT 100 directly rather than through CT 101. Routing it through CT 101 would make the public path depend on a locally signed leaf that expires 2028-11-18 with nothing renewing it. The tunnel already provides the encryption that hop would add. CT 101 keeps serving the internal name.

Section 4 not now decision on the hub route is reopened. It was correct while there was no application to reach. The consequence of leaving it closed is that the operator cannot see the application at all.
This commit is contained in:
2026-09-11 06:59:53 -05:00
parent 14514b0bac
commit a8081e17dc
2 changed files with 261 additions and 4 deletions
+35 -4
View File
@@ -4,7 +4,7 @@ Live state of the Mechanical Compiler staging instance on `srv-b`.
| | | | | |
|---|---| |---|---|
| Updated | 2026-08-18, after repository seeding and container standardisation | | Updated | 2026-09-11, after the composer replaced the placeholder |
| Instance | Staging / development | | Instance | Staging / development |
| Specification | `ENVIRONMENT.md` revision 5 | | Specification | `ENVIRONMENT.md` revision 5 |
| Failure log | `FAILURES.md` | | Failure log | `FAILURES.md` |
@@ -238,6 +238,19 @@ Restart=on-failure, hardening set applied, enabled at multi-user.target
Exists solely to prove the proxy chain independently of the application. It is Exists solely to prove the proxy chain independently of the application. It is
removed when the real service arrives. removed when the real service arrives.
**Retired 2026-09-11.** `mechcomp-placeholder.service` is disabled and stopped.
`mechcomp.service` holds `10.20.0.10:8770` in its place, running
`/var/www/mechcomp/venv/bin/python -m mechcomp.web` as `mechcomp`, reading
`/etc/mechcomp/mechcomp.env` for its binding. nginx on CT 101 required no
change: the real service took the address the placeholder occupied.
The placeholder unit file is left on disk, disabled. It is the rollback:
`systemctl disable --now mechcomp.service` then
`systemctl enable --now mechcomp-placeholder.service`.
Proven end to end from `srv-b` at commit `14514b0`: `HTTPS 200` for the page
and a real JSON build payload through the TLS chain.
### Rollback copies on disk ### Rollback copies on disk
``` ```
@@ -452,6 +465,19 @@ so SSH adds no privilege — but **every path to a container runs through
Direct WireGuard-side access to the catalogue would be a route addition for Direct WireGuard-side access to the catalogue would be a route addition for
`10.20.0.0/24` on the hub. Not now. `10.20.0.0/24` on the hub. Not now.
**Reopened 2026-09-11 by `WORK-ORDER-004`.** That decision was correct while
there was no application to reach. There is one now, and the consequence of
the decision is that CIVICVS cannot see it: `mechanical-compiler.dev.infra`
resolves only on `srv-b`, CT 100 and CT 101, and `10.20.0.0/24` sits behind a
portless bridge with LAN traffic dropped.
The `srv-b` half needs no change, verified 2026-09-11: `ip_forward=1`,
`FORWARD` policy `ACCEPT`, a direct route on `vmbr1`, and none of the three
`FORWARD` rules matches hub-initiated inbound traffic. The single gate is
`AllowedIPs` on the hub's peer entry for `srv-b`, currently `10.110.0.0/22` --
WireGuard drops by cryptokey routing before consulting any routing table, so a
route without that entry does nothing.
No DHCP anywhere. Two containers with fixed addresses on a portless bridge is the No DHCP anywhere. Two containers with fixed addresses on a portless bridge is the
entire address space. entire address space.
@@ -477,11 +503,15 @@ None of the following can be honestly completed, and none may be fabricated by
provisioning: provisioning:
- [x] ~~Dependency install from committed manifests~~ — done 2026-08-18 - [x] ~~Dependency install from committed manifests~~ — done 2026-08-18
- [ ] `mechcomp.service`, `mechcomp-worker.service` - [x] ~~`mechcomp.service`~~ — done 2026-09-11, `deploy/mechcomp.service`
- [ ] `mechcomp-worker.service` — no worker exists yet; `src/mechcomp/worker/`
is still a stub
- [ ] Reference toolchain image `mechcomp/reference-toolchain:8.0.0` - [ ] Reference toolchain image `mechcomp/reference-toolchain:8.0.0`
- [ ] Fixture reproduction against `ddd0f154...` - [ ] Fixture reproduction against `ddd0f154...`
- [ ] Application-runtime acceptance - [ ] Application-runtime acceptance — partial. The composer serves and is
- [ ] Replacing the placeholder with the real service reachable through CT 101; no acceptance criteria have been written for
it, and it is not reachable from outside (see section 6, question 4)
- [x] ~~Replacing the placeholder with the real service~~ — done 2026-09-11
**Applied 2026-08-18:** `venv/` is in `.gitignore`, along with `.cache/`, **Applied 2026-08-18:** `venv/` is in `.gitignore`, along with `.cache/`,
`.local/`, `.ssh/`, `.gitconfig` and `.lesshst` — the service user's home is `.local/`, `.ssh/`, `.gitconfig` and `.lesshst` — the service user's home is
@@ -499,6 +529,7 @@ The Shapely port gates all of the above.
| 1 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025 estate half | | 1 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025 estate half |
| 2 | What is the backup strategy? | all backup work | | 2 | What is the backup strategy? | all backup work |
| 3 | Where is the 3+ TB USB disk attached? | gold redundancy step | | 3 | Where is the 3+ TB USB disk attached? | gold redundancy step |
| 4 | Public ingress for the composer at `dev.mechcomp.kane-il.us` | `WORK-ORDER-004`; anyone outside `srv-b` seeing the application at all |
All three need CIVICVS. All three need CIVICVS.
+226
View File
@@ -0,0 +1,226 @@
# WORK-ORDER-004 — public ingress for the composer
Publish the Mechanical Compiler composer at `dev.mechcomp.kane-il.us`, reachable
from an ordinary browser on the internet.
| | |
|---|---|
| Mode | Infrastructure (`PROCESS.md` §2) |
| Created | 2026-09-11 |
| Executed on | `wg-pk` (the WireGuard hub), via its Webmin terminal |
| Depends on | `14514b0` — composer live on CT 100, proven through CT 101 |
---
## 0. Read this before the first command
**This work order is executed on a machine outside this project.** `wg-pk`
carries `kane-il.us` mail and Hubzilla. `PROCESS.md` §7 requires escalation for
exactly this, and the escalation is: CIVICVS executes, one group at a time, and
anything surprising stops the work rather than being worked around.
### The lockout risk, and why it is bounded
Step 2 changes a live WireGuard peer. If it goes wrong the tunnel drops, and
with it every path to `srv-b` — `PROCESS.md` §1 records that the `srv-b` shell
is the operator's entire working surface for this project.
It is recoverable because **the operator is sitting on `wg-pk` itself**, in
Webmin, not reaching it through the tunnel. `wg set` is not persistent, so
`systemctl restart wg-quick@wg0` on the hub restores the on-disk configuration
and the tunnel with it.
Do not perform step 2 from a shell that reaches the hub through the tunnel.
### What `srv-b` contributes
**Nothing. It is already correct and must not be touched.** Verified
2026-09-11:
```
net.ipv4.ip_forward = 1
-P FORWARD ACCEPT
10.20.0.10 dev vmbr1 src 10.20.0.1
```
The three `FORWARD` rules block container-sourced SMTP, container-to-container
(ACCEPT), and container-to-LAN. None matches hub-initiated inbound traffic. The
`-s 10.20.0.0/24 -o wg0 MASQUERADE` rule does not apply either: the first packet
of a hub-initiated flow is not container-sourced, so no NAT binding is created
and replies return through conntrack.
**Explicit do-not-touch list:** `srv-b` iptables, `srv-b` WireGuard, CT 101
nginx, CT 101 TLS, the `kane-il.us` MX records, anything on
`mx1.diagnostics.kane-il.us` (it publishes a TLSA record — confirmed
2026-09-11).
---
## 1. Topology, and why
```
browser
-> DNS dev.mechcomp.kane-il.us -> wg-pk public address
-> nginx on wg-pk, Let's Encrypt TLS terminated here
-> WireGuard tunnel, already encrypted
-> srv-b 10.110.0.12, forwards, no configuration change
-> CT 100 10.20.0.10:8770, the composer
```
**The public path deliberately bypasses CT 101.** Routing it through CT 101
would make the public name depend on a locally-signed leaf valid to 2028-11-18
with nothing renewing it — a dated outage designed in from the start. The tunnel
already provides the encryption that hop would add.
CT 101 continues to serve `mechanical-compiler.dev.infra` for work from inside.
Two ingresses, each with a distinct reason to exist.
---
## 2. Success criteria, stated before the work
1. `curl -I https://dev.mechcomp.kane-il.us/` returns `200` from a machine with
no WireGuard access and no special DNS.
2. The certificate is issued by Let's Encrypt and chains without `-k`.
3. `https://dev.mechcomp.kane-il.us/api/build?family=3x&profile=Y` returns JSON
whose `groups` end with a `Y` group and contain no `three_fin` key.
4. The page renders, controls change the drawing, and a changed control changes
the `design` identity shown beneath it.
5. **The negative:** `mechanical-compiler.dev.infra` still answers from `srv-b`,
and mail from `srv-b` still delivers. Neither path was in scope; both must be
proven unharmed.
6. Every change survives `systemctl restart wg-quick@wg0` and a hub reboot.
---
## 3. Command groups
One at a time. Paste output back before the next.
### Group 1 — read-only, on `wg-pk`
Nothing here changes anything.
```bash
echo "=== peer entry for srv-b ===" && \
wg show && \
echo "=== on-disk wireguard config ===" && \
grep -n "AllowedIPs\|PublicKey\|Address\|PostUp" /etc/wireguard/wg0.conf && \
echo "=== routing toward the service network ===" && \
ip route | grep -E "10\.20\.|10\.110\." ; \
echo "=== forwarding ===" && \
sysctl net.ipv4.ip_forward && \
echo "=== nginx vhost conventions, following symlinks ===" && \
grep -rn --dereference-recursive "server_name\|proxy_pass\|listen\|ssl_certificate " \
/etc/nginx/sites-enabled/ | head -40 && \
echo "=== certificate issuance ===" && \
ls /etc/letsencrypt/live 2>/dev/null || echo "no certbot live dir" ; \
which certbot ; \
echo "=== does the name resolve yet ===" && \
dig +short dev.mechcomp.kane-il.us A ; \
dig +short dev.mechcomp.kane-il.us AAAA
```
What matters in the output: whether the peer's `AllowedIPs` is the only gate,
whether nginx already has `listen [::]:443` on its vhost pattern (it must, if
the name inherits an AAAA), how certificates are issued, and whether the name
resolves yet.
**Report this before proceeding.** Groups 2 onward are written against what it
shows; they are deliberately not drafted in advance.
### Group 2 — WireGuard AllowedIPs
Drafted after group 1. The shape:
- Back up `/etc/wireguard/wg0.conf` first. `PROCESS.md` §2: preserve a rollback
copy before editing any configuration file.
- Add `10.20.0.0/24` to the `srv-b` peer's `AllowedIPs`, **keeping
`10.110.0.0/22`**. Replacing rather than extending it is the way this breaks.
- Apply live, confirm the tunnel is still up and `srv-b` still reachable, then
persist.
- Prove: `ping -c2 10.20.0.10` from the hub.
### Group 3 — route
A route for `10.20.0.0/24` via `10.110.0.12`, persisted the way the hub already
persists routes — which group 1 reveals. Prove with
`curl -sS -o /dev/null -w '%{http_code}\n' http://10.20.0.10:8770/`.
### Group 4 — certificate
Let's Encrypt HTTP-01 for `dev.mechcomp.kane-il.us`. A third-level name needs no
wildcard and no DNS-01. Requires the name to resolve to the hub first.
### Group 5 — vhost
```nginx
# /etc/nginx/sites-available/dev.mechcomp.kane-il.us
# Matching whatever pattern group 1 shows the existing vhosts use.
server {
listen 80;
listen [::]:80;
server_name dev.mechcomp.kane-il.us;
return 301 https://$host$request_uri;
}
server {
listen 443 ssl;
listen [::]:443 ssl; # required if the name has an AAAA
server_name dev.mechcomp.kane-il.us;
ssl_certificate /etc/letsencrypt/live/dev.mechcomp.kane-il.us/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/dev.mechcomp.kane-il.us/privkey.pem;
location / {
proxy_pass http://10.20.0.10:8770;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
# The composer redraws on every control change. Small responses, many of
# them; buffering adds latency for no benefit.
proxy_buffering off;
proxy_read_timeout 30s;
}
}
```
`nginx -t` before reload, always.
### Group 6 — acceptance
All six criteria in §2, including the two negatives.
---
## 4. Known unknowns
Stated rather than assumed, because assuming is what produced three unreachable
URLs before this document existed.
- **Whether the hub's nginx listens on v6.** If `dev.mechcomp` is a CNAME to a
name carrying an AAAA, v6 is published whether or not nginx answers on it.
v6-preferring browsers then fail while v4 ones succeed — a fault that looks
like anything except DNS.
- **How the hub persists routes.** Distribution- and tooling-dependent; group 1
shows it.
- **Whether Webmin or YunoHost manages the nginx configuration.** If so,
hand-written vhosts may be overwritten on their next reconfiguration, and the
vhost belongs wherever that system expects it instead.
---
## 5. Acceptance and reporting
Closed when all six criteria in §2 hold and survive a hub reboot.
Back to the architect: the output of every group, and any surprise about the
environment in `FAILURES.md` format — `PROCESS.md` §7. A surprise on the hub is
still a finding even though the hub is not this project's property.
`STAGING-STATE.md` §4 records "direct WireGuard-side access to the catalogue
would be a route addition for `10.20.0.0/24` on the hub. Not now." That decision
was correct when there was no application to reach. This work order is its
reversal, and §4 should say so when this closes.