Files
mechanical-compiler/docs/PROCESS.md
T
TheRON 6ce8ece709 docs: F-037, the ingress corrections, and the priority order reversed
HANDOFF section 4 said WORK-ORDER-004 was pending and that CIVICVS could not
see any of this work. It was rewritten thirty-two minutes after 1fdb115 closed
the ingress, in the same session, against the work order's old state rather
than its new one -- so HEAD described a world where the thing you were already
looking at did not exist. Section 11 row 7 carried the same staleness. Both
corrected, and the correction says how it happened, because section 0 was added
to prevent exactly this and did not.

Three things 1fdb115 recorded as unsettled were never promoted into section 11
and are now rows 8 through 10: TLS renewal has never been observed to succeed
for this name and is first due before 2026-12-10; the composer has no
acceptance criteria of its own, only the path to it does; and the service is
world-reachable and unauthenticated, which section 4 item 5 called a conscious
decision to be made before publishing and which publishing has now made due.

Priority items 1 and 2 reversed. The original case for authorship first was
that records made before the field exists can never be attributed. That was too
strong, and too strong because the field is excluded from both hashes: a record
regenerated later keeps its input_id. What it does not keep is build_id, which
covers Shapely and GEOS, and boolean results on near-degenerate geometry shift
between GEOS releases -- the F-034 mechanism. Attribution is recoverable, not
free. The residual argument stands and the field landed first at 48d5665.

Recorded under item 2, because it is the first thing STL export runs into:
length_view is not in COMMON_GROUPS, so the Full Length branch is unreachable
from the composer and every export would silently be a 100 mm preview. The
sweep must equal model_length_mm(p), already published as LENGTH_MM, or the
record's VOLUME_MM3 and MASS_G describe a different object than the file beside
them. Also split the roadmap's STL/STEP bullet: STEP needs the kernel, STL does
not, and one line implying both was left behind by the section 5 correction.

F-037: a tar stream rooted at "." re-owned the repository root. tar x ran as
root and applied the "." entry's ownership to /var/www/mechcomp; the chown that
followed named only src and tests. Everything below the root was correct, so
547 tests passed and only git noticed. Corrected by chown on one path, owner
only, no -R. safe.directory was not added -- that is F-008 and would have
masked this and every later instance.

Two of F-037's three consequences are process defects of mine rather than facts
about the environment. The verification step ran before the landing that
destroyed it, so a git status ahead of the breaking step reads as a pass -- the
F-027 pattern in a new place. And section 7 says record a failure before
correcting it; I corrected first. Both recorded rather than quietly fixed.

PROCESS section 3 gains two REQs, because the delivery path as documented
produces F-037 every time: the chown must name the directory the files land in,
and a tar transport must name its top-level directories rather than root at
".". pct push is simpler for a single file and cannot reproduce it at all.

Section 8 gains the bytecode rule: PYTHONDONTWRITEBYTECODE=1 and clear
__pycache__ between mutations. A stale .pyc masked a real defect once and every
mutation result reported before that was optimistic by an unknown amount. A
mutation surviving on stale bytecode is indistinguishable from one surviving on
a weak test.

Open question 11 is new and is CIVICVS's: ct-baseline.sh exits 0 with the F-037
condition present, so by section 9a the ownership of a service working tree is
not part of the container standard. Whether it should be is a decision about
host property covering three projects.
2026-09-12 04:53:24 -05:00

15 KiB

PROCESS.md

How work gets done on the Mechanical Compiler.

Created 2026-08-17
Scope All work on a Mechanical Compiler instance
Companions ENVIRONMENT.md, STAGING-STATE.md, FAILURES.md, ROADMAP.md

0. Why this document exists

The four companion documents describe what the environment is. None of them describes how work gets done in it — where files land, who runs what, how a change travels from a conversation into a running container.

Until now that gap was filled by restating the mechanics inside each work order. That worked while the orders were written consecutively by one author. It does not survive a handoff: an assistant reading ENVIRONMENT.md learns how the host is built and nothing about how to operate on it.

Read this before your first command.


1. Who does what

Three parties, and confusing them is the most common way to waste a session.

Party Has Does
Operator (CIVICVS) A shell on srv-b and a file manager Runs every command. Uploads files. Decides.
Assistant This conversation Writes commands, reads output, writes code and documents
Architect This conversation, in a different mode Sets specification, accepts work, maintains the documents

The constraint that shapes everything

The operator has a shell on srv-b and a browser-based file manager. Nothing else.

No workstation git client. No direct container shell. No IDE. No ability to scp. Every file arrives by upload to a folder on srv-b; every command runs in that one shell.

An assistant that assumes otherwise produces instructions the operator cannot execute. This has already happened once.

What follows from it

  • Files reach a container as: upload to srv-b → pct push or expand and copy → pct exec to act on them.
  • srv-b is the only place a human types. Containers are reached through pct exec, never by SSH from elsewhere.
  • This is a legitimate deployment path, not a workaround. Do not design around a git client that does not exist.

2. Two modes of work, and they are not the same

Applying the wrong one is slow in one direction and dangerous in the other.

Infrastructure mode

For anything that changes the host, the containers, the network, or a service that other work depends on.

  • One command group at a time
  • Read-only before write
  • Record a failure before correcting it, in FAILURES.md format
  • Smallest corrective experiment — not the one that also fixes three adjacent worries
  • Preserve a rollback copy before editing any configuration file
  • Prove the negative: assert the forbidden path fails, not that the interface is gone
  • Assert health after a host reboot, never before

This is deliberate ceremony. Each step is irreversible or expensive to undo, and the failure log is a deliverable in its own right — it is the primary input to the eventual production automation.

Work orders 001 through 003 were all infrastructure mode.

Development mode

For anything inside the repository: code, tests, documents, fixtures.

  • Iterate freely. Edit, run the tests, edit again.
  • No command-group ceremony. A test-fix-rerun loop is not a provisioning step.
  • Failures during a normal red-green loop are not FAILURES.md entries.
  • Commit when something works, not when something is finished.

The distinction is reversibility. A wrong iptables rule can strand the operator. A wrong line of Python fails a test and gets deleted. Ceremony appropriate to the first is obstruction applied to the second.

Which mode am I in?

Ask: if this is wrong, what does it cost?

Cost Mode
A rebooted host, a lost shell, a broken service Infrastructure
A red test Development

If genuinely unsure, use infrastructure mode. Being slow is recoverable.


3. Getting files into the environment

The path

assistant produces a tarball
  -> operator downloads it
  -> operator uploads it to /root/incoming on srv-b   (file manager)
  -> operator expands it on srv-b
  -> pct push, or expand-then-copy, into CT 100
  -> pct exec to act on it

REQ — /root/incoming on srv-b is the landing area. One directory, always the same, so nothing has to be remembered between sessions.

REQ — Files landing in a container must end up owned by mechcomp. pct push writes as root; fix ownership immediately afterwards. See F-008 — repository operations run as the service user, and a root-owned file inside the tree causes exactly that failure later.

REQ (F-037) — The chown must name the directory the files landed in, not only the files. A tar stream rooted at . carries an entry for the destination directory itself, and tar x as root rewrites that directory's ownership. Naming only the payload subdirectories leaves the repository root root:root, and git then refuses the worktree with the F-008 message — which invites the F-008 mistake as its own remedy. Fix the owner. Never add safe.directory.

REQ (F-037) — Verification that depends on the repository being intact must run after the landing, and a landing step must not be able to destroy the check that would have caught it. A git diff placed before the step that broke git reports nothing wrong and looks like a pass.

REQ — An assistant delivering files states, in this order: what the archive contains, where it expands, what it overwrites, and how to verify it landed correctly. "Overwrites nothing" is a claim that must be checked, not assumed.

Delivering a whole tree

Prefer one tarball that expands over the existing tree. Individual file paths are error-prone to transcribe through a file manager.

Name the payload's top-level directories explicitly, when creating the archive and when extracting it: tar cf - -C <dir> src tests, never -C <dir> .. The second form is what produced F-037. For a single file, pct push is simpler and cannot reproduce it at all.

Delivering a single small file

For anything that fits comfortably on screen, a heredoc in the srv-b shell is faster than an upload, and leaves the content visible in the session transcript where it can be checked.


4. The repository

Gitea is the source of truth. https://gitea.barternetwork.us/TheRON/mechanical-compiler

CT 100 holds a clone at /var/www/mechcomp, owned by mechcomp.

Two directions, both valid

Gitea → container. git pull inside CT 100, as mechcomp. Use this whenever the change already exists in Gitea.

Upload → container → Gitea. Land the files, verify they work, then commit and push from inside CT 100. Use this when the assistant produced something new.

The second direction is the normal one for assistant-produced work, because the operator has no git client outside the environment.

REQ (F-008) — Every git command runs as mechcomp. Never as root, and never add a root safe.directory exception — it would mask every future instance of the same mistake.

What must be true before a push

  • The oracle verifies: python3 fixtures/strap-beam-8.0.0/make_fixtures.py --verify
  • The test suite runs, with a result the assistant has actually seen
  • git status shows nothing unexpected — particularly not venv/

Recording provenance

Artifacts are attributed to the commit that produced them, not to a timestamp or a file copy. When an artifact is generated, the commit SHA of the tree that produced it is part of its record. This is why the transport does not matter but the SHA does.


5. The development loop

Inside CT 100, as mechcomp, from /var/www/mechcomp:

make deps        once, and after any requirements change
make test        the loop
make verify-oracle   before any commit that touches fixtures

The oracle is the gate

fixtures/strap-beam-8.0.0/ holds 123 frozen cases. It is the acceptance criterion for the port, and it is not negotiable.

The ten rejected cases are part of the contract. A port that accepts them is wrong however good its numbers are elsewhere. Reproducing geometry is straightforward; keeping the constraint that made the geometry trustworthy is the actual work.

The oracle is never edited to make a test pass. If the port disagrees with it, the port is wrong until proven otherwise. If the oracle is genuinely wrong, that is a finding, it is recorded, and it is regenerated inside the pinned toolchain — never hand-edited.

Test discipline

A test must distinguish failure of the tool from failure of the thing under test (F-027). Every negative assertion first establishes that it could have observed the positive case. A test whose failure mode is indistinguishable from success manufactures confidence.

Prove a new harness by making it fail deliberately before trusting it green.


6. Work orders

A work order is an infrastructure-mode instruction set. Development work does not need one.

When to write one

  • The work changes the host, containers, network, or a shared service
  • The work has an acceptance criterion that can be stated in advance
  • The work will be executed by someone other than the author

Required sections

Section Purpose
Scope In and out. Explicit "do not touch" list.
Success criteria Strict, testable, stated before the work
Known state So the operator does not re-derive what is already recorded
Command groups One at a time, read-only first
Acceptance What must be true to close
Reporting What comes back to the architect

Executing one

  • One group at a time. Paste output back before the next.
  • Record a failure before correcting it.
  • A defect in the work order is a finding, not something to work around silently. F-024 and F-028 are work-order defects, logged as such.
  • "The problem is elsewhere and here is the evidence" is a complete outcome. Do not manufacture a local workaround for a remote problem.

7. When something goes wrong

Is it a FAILURES.md entry?

Yes if it is a surprise about the environment: something behaved differently from what the specification or a work order said, and a future automation writer would get it wrong the same way.

No if it is an ordinary development failure: a red test, a syntax error, a wrong algorithm. Those are the loop working.

Entry format

### F-nnn — one-line summary
Host or container. Phase.
**Observed:** verbatim where possible
**Cause:** proven, or explicitly "unproven"
**Correction:** the smallest change that fixed it
**Consequence:** what the specification or automation must do differently

An entry with no Consequence is either not understood or not worth recording.

"Unproven" is a respectable result. F-006, F-012 and F-022 are carried open and the log is better for it. Never invent a cause to close an entry, and never close one on a change that merely coincided with the symptom disappearing.

Escalating

Stop and ask the operator when:

  • A decision is needed that is not the assistant's to make
  • The work requires changing infrastructure outside this project
  • A REQ in the specification cannot be satisfied
  • The specification and observed reality disagree

The last one is not a blocker — reality wins, and the specification is corrected. But it is always worth saying out loud.


8. Session handoff

Sessions end abruptly. Assume the next assistant has these documents and nothing else.

Before a session ends

  • Anything working is committed and pushed
  • Anything learned is in FAILURES.md
  • Anything changed on a host is in STAGING-STATE.md
  • Any open decision is in STAGING-STATE.md section 6

Starting a session

Read in this order:

  1. PROCESS.md — this document
  2. STAGING-STATE.md — what is true right now
  3. FAILURES.md — what has already gone wrong
  4. ENVIRONMENT.md — the specification
  5. ROADMAP.md — where it is going

If writing provisioning automation, read FAILURES.md before the specification. Every entry is something a script written from the specification alone would have got wrong.

Authority when documents disagree

  1. STAGING-STATE.md — factual, wins on what is true
  2. FAILURES.md — evidence from contact with real hosts
  3. ENVIRONMENT.md — the specification, corrected when proven wrong
  4. ROADMAP.md — sequence

A permanent deviation is not a deviation. It is the specification. Promote it and delete the exception.


9. Instructions the operator can actually run

This section exists because it has already gone wrong.

REQ — One command group per message. Wait for output.

REQ — Commands run in the srv-b shell, or via pct exec from it. Nothing assumes a tool the operator does not have.

REQ — A group is copy-pasteable as a block, with echo markers separating sections so the output can be read back.

REQ — State what the expected output is. An operator who cannot tell success from failure cannot report usefully.

Do not deliver a numbered multi-stage procedure spanning several systems and expect it to be executed. It will not be, and the parts that are will not be separable in the output.

Do not assume the operator will improvise the missing step. If a command needs a directory to exist, create it in the same group.


9a. The baseline check

ct-baseline.sh is the executable definition of the container standard on this host. Read-only, runnable at any time, exits non-zero on divergence.

Run it:

  • before starting work on a container;
  • after any change to a container or to host firewall rules;
  • after any host reboot;
  • before considering backup work (see ENVIRONMENT.md §15 gate 4).

A property not checked by it is not part of the standard. If something should be uniform across containers, add it to the script. If it should not, leave it out and stop worrying about the difference. That is the whole point: it converts an endless comparison into a pass or a fail.

It covers containers belonging to more than one project, and encodes host requirements neither project owns alone. Treat it as host property.


10. Current mode

Infrastructure: complete and accepted. Host, containers, network isolation, bastion access, TLS, reverse proxy, mail alerting, disk monitoring. See STAGING-STATE.md.

Backup: deliberately postponed. Nothing exists yet whose loss would cost more than an afternoon; the documents are in Gitea. This changes when artifacts/ stops being empty.

Repository seeded at c7e32d8: reference implementation, frozen oracle, pinned toolchain, test harness. Dependencies installed in CT 100, oracle verified in-container, 3 passed, 236 skipped.

Container baseline established. All three containers on srv-b conform.

Development: starting. First work item is the Shapely port — pytest -n auto green against the 123 frozen cases.

From here the mode is development unless the work touches the host, the containers, or a shared service.