Skip to content

Open issues

UNVERIFIED: NOW Web Direct needs Classilla and MacWeb acceptance (2026-08-10, codex/web-proxy)

The Direct implementation, host supervision, PowerPC Workshop page, semantic rewriter, Reader, Wikipedia/Reddit handlers and optional local-model adapter are built and covered by local gates. None has crossed an actual classic browser yet. The first required rows are mac99/Classilla, PB1400c/Classilla, q800/MacWeb and PB180c/MacWeb. Each must record the request form the browser actually sends, navigation through rewritten links, page byte/chunk behavior, and cooperative liveness while NOW's dialogs and controls are active.

The address boundary is intentionally unresolved rather than guessed. 10.0.2.2 is the modern host as seen by this repository's QEMU user network; physical machines use the host's LAN address. No code calls 127.0.0.1 a classic-Mac endpoint. Open Transport and MacTCP listener/connect behavior for loopback and the guest's own address still require separate target probes, so there is no guest-local relay or relay wire contract.

Other open rows: forms, logins, cookies, uploads, CONNECT tunneling, a complete image transcoder, DNS-rebinding-resistant destination pinning, a persistent model service rather than per-request cold load, and model distribution. The inspected model artifact lacks a complete base-license and training-data redistribution record; local use is enabled, copying its weights is not.

METAL-VERIFIED: host project, classic identity, MPW build/run/dialog (2026-08-09)

The host Projects store, guest Development module, private import/candidate lanes, declarative ToolServer runtime, promotion guard, exact-product run and optional CodeKitten handoff are locally tested and both guests build. One bounded PPC emulator rung now passes: on mac99/OS 9.1, guest build 15a1c3087007 qualified mpw-ffff-00000cf0@structural-1, measured the active three-file project, completed MrC/PPCLink/Rez, reported an APPL/MMTR product at 1,700 data and 568 resource bytes, launched the exact product with matching process identity, cancelled another job and completed a later build.

That run did not exercise import, workspace restart, inactive candidate stage, divergent promote, successful promote or CodeKitten handoff. Those remain the emulator acceptance gap. The PowerBook metal rung still owes the exact MPW toolchain/version, both product fork sizes/digest, process identity and an odoc edit imported as a new revision.

The first PowerBook 1400c candidate-publication attempt on 2026-08-09 found a reproducible blocker before MPW ran. Guest build 15a1c3087007 qualified mpw-ffff-00007b37@structural-1; a NOW-owned host-home Hello World project was committed at revision 2 and staged twice. On both attempts the guest log reported Project.ckp, Hello.r and Main.c received with checksums OK, then candidate finalization refused candidate-unavailable with "The candidate request is malformed or no longer accepting files." A fresh build status was Job none, State idle, Actions 0 of 0, so this is a candidate acceptance/sealing defect, not an MPW result. Preserve that boundary when diagnosing it: transfer completion is proven; candidate re-find/seal is not. The refused stage result also omits the minted candidate ID, so MCP cannot address stage-status or stage-discard; the host's finalize-refusal path does not automatically discard the guest candidate. The two attempts may therefore have left inactive residue with no agent-visible recovery handle.

Updated 2026-08-10: the candidate was accepting. The host encoded expectedFiles as the JSON string "3", while the contract and guest parser require an integer; the parser therefore saw zero and the old compound guard collapsed that type error into candidate-unavailable. Project import had the same latent defect for its numeric cursor. Both fields now cross as typed JSON integers. The guest reports malformed digest, count, candidate root/folder, seal, measurement and marker failures separately and logs the failed predicate. The host logs each publication phase and, after a finalize refusal, either discards both candidates or retains the host receipt and returns its candidate ID for stage-status / stage-discard. A mutation-checked wire test guards both number fields. On a private mac99/OS 9.1 guest running build 995a13285c98, MCP staged the same revision-2 project, transferred its three files, sealed candidate candidate-fb94b4e9b84e4fd0 at the exact f3401211847f396634542276c44243a1262871510bdadcc8ca7238a36739c730 digest, observed it with stage-status, then discarded it. Candidate staging is therefore emulator-verified. The PowerBook has not yet rerun this patched stack, so the original metal failure is diagnosed and fixed but not metal-reverified.

The onboarding server can now place a separately supplied CodeKitten MacBinary at the setup-volume root, select it by default, advertise it as a standalone optional IDE, and serve it directly at /now/codekitten.bin. The host-mounted combined image preserved CodeKitten's 628,086-byte data fork, 6,888-byte resource fork and APPL/O9ID identity. On the Mac OS 9.1 emulator, the exact payload transferred and LaunchApplication returned success, but CodeKitten left the process table before the five-second observation. This is not a passed CodeKitten handoff or runtime gate. The combined image's Disk Copy mount was separately blocked by the pristine snapshot's unaccepted Apple license; the earlier image-carrier mount result below remains valid for the carrier, not for this new payload set.

Updated 2026-08-10, classic file preservation and complete host-home rung: The PowerBook rerun with the numeric candidate fix sealed the project, then MPW refused Sources:Main.c as “not a TEXT file.” Generic upload had reconstructed the bytes as BINA/ttxt; the development contract had treated a classic file as one data blob. The current branch makes data/resource forks and Finder type, creator and flags first-class through now_projects, project digests, candidate receipts, MacBinary guest transport/import and the host Git recovery archive. Legacy projects remain readable with identity reported as unknown, but staging refuses to invent metadata.

A fresh mac99/OS 9.1 acceptance used only the MCP Development/Projects surface after fixture-only registration of the emulator's existing MPW. Guest build 44a214ae1141 qualified mpw-ffff-00000cf0@structural-1; project 95ceb07504374a568705336a5728d19a carried a nonempty source resource fork and TEXT/MPS identity, candidate candidate-1c3ca2817afe4c24 sealed at digest 2d692e6239cb90862bb1a15c7d6f42ce63ea9026ac35663e9c41913725e417d0, and MrC/PPCLink/Rez produced APPL/H14E product product-5c96932bd4b1cd44 with 1,832 data and 578 resource bytes. Exact-product run matched the process; retained semantic UI observed the live frontmost alert and enabled OK item. After semantic dismissal the process disappeared and candidate discard passed.

That acceptance also found two authoring constraints rather than transport defects: the first sample omitted its application-owned QDGlobals qd, and a 26-byte debug product name left no room for PPCLink's .xcoff sidecar inside HFS's 31-byte component limit. The project was revised through MCP. Host and guest parsers now refuse the derived .xcoff overflow before staging; both guards were mutation-tested. The current fork/identity stack has not run on the PowerBook, so it remains emulator-verified rather than metal-verified.

Updated 2026-08-10, PowerBook acceptance complete: the signed host at source 3de370bf connected to PowerBook guest build 33bfe4e3e211 with Full Access. The guest qualified its selected mpw-ffff-00007b37@structural-1 ToolServer/MrC installation. MCP changed only the emulator project's toolchain pin, producing host revision 5, then sealed candidate candidate-d4752f9e46d84f9c at digest 6c3e96f8e4e368fc5004e2817ab9bcc087bd1eacba5bdd5cb0437c3b77a13669. MrC, PPCLink and Rez completed 3 of 3 actions. Product product-cf2b162d6ea8648a measured APPL/H14E, 1,832 data-fork bytes and 578 resource-fork bytes, digest cf2b162d6ea8648a7108b10b3732994ebc647f39c40ba74a1986ada1f6e564a8. Exact-product run reported a matched process identity. Retained semantic state and the person at the machine observed the frontmost Hello World alert and its enabled OK item; after the person dismissed it, a fresh process census showed HelloForks absent. Candidate discard then completed. The fork/Finder identity repair and complete host-home build/run/dialog lane are therefore metal-verified on the PowerBook 1400c. The test-receipt, guest-home promotion and CodeKitten handoff gaps below remain open and are not covered by this claim.

The run also exposed one agent-loop settlement defect that does not invalidate the build/run receipt. The first retained MCP snapshot contained the correct frontmost HelloForks process, dialog surface and enabled OK item, but carried baseComplete=false; the immediately following dialogItem act refused now-mirror-snapshot-unavailable before sending. Source inspection names two authorities at that seam: MirrorStateProjectionService reads MirrorStateEngineRegistry.snapshot, while MirrorDriveService resolves the act against NOWMirrorSource.scene. The former had the scene and the latter was nil. A separate process-list request transiently returned now-host-communication-failed, then succeeded on retry. An attended human can reconcile both cases, as this run did; an unattended agent cannot yet treat a published snapshot as proof that its named entities are actable. Reconcile the read and drive authorities and mutation-test that exact split before calling semantic runtime settlement autonomous-loop reliable.

Two MCP contract edges made diagnosis harder than it needed to be. The local server does not expose an early host-build/protocol compatibility verdict, so an older running host reported the generic now-host-invalid-response for each new Projects/Development call instead of naming that its response shape predated those tools. Separately, now_projects publishes one broad input schema for all operations, while project apply and workspace apply accept mutually exclusive revision/commit guards; supplying both reached only the generic invalid-arguments sentence. A version handshake and discriminated per-operation schema are required before treating ordinary agent mistakes or a stale deployment as product-domain failures.

Two implementation gaps block a full development/preservation/acceptance claim today:

  • now_development has separate build and exact-product run outcomes, but no test operation or ckproject.test-receipt/1. A test cannot be reported as a build configuration or a successful launch: its closed declarative action vocabulary, expected product identity and terminal receipt still need to be added to Project.ckp, the ToolServer service and both human/agent faces.

  • development-open proves CodeKitten launch, odoc dispatch and foregrounding without blocking the guest loop, but kAENoReply cannot prove the handler accepted the document. CodeKitten now has the handler; the remaining seam is an asynchronous acceptance receipt, not another IDE dependency.

The host Development import sheet also requires an opaque project ID. A bounded guest project catalog is still needed before that is a discoverable human workflow.

Updated 2026-08-10, autonomous-loop hardening: the three gaps above are closed in NOW and emulator-verified. Snapshot and act planning now resolve through one published Mirror state engine; every admitted direct action enters the operation journal, and wait_for_settlement waits by its UUID. A direct act with no declared postcondition ends unconfirmed, never confirmed from dispatch alone. Projects and Development publish discriminated operation schemas, mutation attempt IDs are mandatory, terminal local responses survive host restart, and compatibility names the host build, protocol, catalog digest and supported schema revisions before a domain request is decoded. The guest catalog makes import discoverable. Project.ckp now carries an all-or-none closed test plan and the PPC guest returns ckproject.test-receipt/1 only after the unchanged product's process identity matches.

A session-private mac99/OS 9.1 acceptance with guest build b1de53f2bfe9 and qualified mpw-ffff-00000cf0@structural-1 exercised three levels of loop: simple Hello World; a five-file, resource-fork-bearing Memory Meter with real MrC failure, repair, cancellation, restage and success; and a guest-only project imported into host scratch, edited, built, tested and promoted. A second built guest-home candidate was refused as guest-diverged after an out-of-band guest edit. Download matched the active edit byte-for-byte and the losing candidate remained inspectable until explicit discard. A deliberately lost stage response was retried with the same attempt ID and returned the one original candidate. The VM's staged resident reported source manifest 28ef6c07ee6d and fingerprint 085c4ebf8457; QEMU eventually exited after the repository shutdown helper and qemu-img check passed, but the fixture base was already marked HFS-dirty, so this run does not assert volume-clean fixture provenance.

Correction, later 2026-08-10: the base was not dirty. Direct tools/volclean.py inspection reports the source MPW image CLEAN and the completed session clone DIRTY. The shutdown applet went quiet, QEMU exited, and the qcow2 container passed, but HFS remained mounted. The helper then printed "already-unmounted machine" and returned success without asking HFS, contradicting its own recorded evidence that applet quiet is not an unmount. shutdown-guest.py now releases QEMU and makes the volume verdict the final return code; dirty and unknown both fail. volclean.py is directly executable as documented. The exact old post-_graceful return 0 mutation fails the new guard. This corrects cleanup observability; it does not invalidate the build, test, promotion, divergence or semantic receipts gathered before shutdown.

Second correction, later 2026-08-10: enforcing that verdict exposed the pre-INIT restart in scripts/spin-up-ppc: it launched only the direct ShutDwnPower fallback and therefore could no longer produce a clean receipt. The restart now launches NOW against the resident already active in the base and drives Finder Special > Shut Down over the wire before cold-booting the newly staged resident. The exact missing---wire mutation fails both rig guards. The first updater bake then caught a separate identity encoding error: NWid contained two full SHA-256 values while its fixed ABI reader consumed two 160-bit fields. NWid now carries exactly the same 40-hex prefixes as the resident table, and the old 128-hex payload fails the component identity gate.

The corrected shared bake from 8fe5baff2a4f is emulator-verified: both Finder shutdowns powered QEMU off in six seconds and left HFS clean; resident 1.2 reported source manifest fae73d4d5c7150b7105e0f0350446cf5111ee248, build fingerprint c725b32b77634fb8d1702b0a54f17fc2ce70e48f, all 511 declared capabilities, a passing act ABI self-test, and survival of all 14 census probes. The installed shared image is SHA-256 ccbff00f4ed1c18eed879b94798df191ea8dae15a83f2ffb0e465cd53126add4; qemu-img check and the HFS unmounted-bit check both passed. This is emulator-verified, not metal-verified.

The run also found one sharp authoring boundary: Project.ckp participates in the project digest as TEXT/NOWD. Uploading identical bytes as TEXT/MPS let the guest catalog parse the document but made import end in the generic coherent-snapshot mismatch. Re-uploading it with the canonical identity made the same snapshot import. The contract now documents the identity and import refuses the noncanonical type, creator, flags or resource fork explicitly before normalizing a host copy. The exact creator mutation was watched to fail.

The remaining open boundaries are narrower and separately owned:

  • MCP transport, historical state: at this checkpoint only the stdio bridge existed. The later update below records the HTTP addition and its subsequent ownership correction; this paragraph is retained to make the sequence and original scope error explicit rather than presenting HTTP as planned work.
  • CodeKitten acceptance: NOW still observes launch, asynchronous odoc dispatch and foregrounding rather than a returned handler receipt. Shared fixtures and any neutral receipt vocabulary must be completed in the sibling CodeKitten repository; NOW must not absorb the IDE or executor.
  • Starter payload: the relocatable, versioned Development starter-pack manifest and onboarding input are implemented and tested, including license and provenance fields. No redistributable MPW bytes are committed. A combined NOW + CodeKitten + MPW image cannot be accepted until payload licensing and provenance are settled.
  • Metal: the earlier PowerBook result remains the fork-aware host-home build/run/dialog proof. Typed test, retry/restart recovery, guest-home promotion and the new semantic settlement receipts remain emulator-only.

Updated later 2026-08-10, HTTP/receipt and varied-loop completion: the first two boundaries above moved. NOW now has an explicitly configured authenticated loopback HTTP MCP listener over the same dispatcher as stdio. Spawned tests compare exact initialization, notification lifecycle, ping, resource, prompt, complete tool descriptor/schema, real result and error behavior, and the same 46-tool no-host recipe reaches both transports. HTTP's own tests cover bearer, loopback Host, Origin, bounded session creation/deletion/expiry, incremental bodies, ambiguous framing and actual listener liveness. On the isolated VM, HTTP served 31 tools, returned 14 typed refusals, left the one person-approved transfer gated, and had zero failed or uncovered rows. This transport was an unapproved expansion of the hardening slice; completing its security and parity gates is a repair obligation, not retrospective approval of that scope choice.

NOW's CodeKitten handoff now waits for and validates ckproject.open-receipt/1; dispatch-only and malformed replies are refused. CodeKitten remains optional and no executor path depends on it. A neutral cross-repository project/receipt module remains future coordinated work rather than a reason for NOW to import the IDE.

The final mac99 run used guest build 27e37aeeaa0a, qualified mpw-ffff-00000cf0@structural-1, and base image SHA-256 be32b70a7fe546b144be76627bf4f20a1777a6fa2fb3e202ef1cd4f059ffe8e2. Every development and guest-file action used HTTP MCP. Four loops covered simple build/test/dismiss/cleanup; source resource-fork preservation; a six-file MrC failure, repair, cancellation, restage and success; and a project created only on the guest, imported to host scratch, edited, built, tested, promoted, deliberately diverged and later recovered by exact typed re-upload. The rebuilt host and HTTP MCP transport also proved a reused mutation UUID crosses as typed attempt-collision; the old missing/current-mismatched response ID is an exact failing mutation fixture.

That run added three open hardening findings:

  • loop-status retains some candidate receipts from an ended guest session, while stage-status against the current session says candidate-unavailable. Candidate receipts do not currently bind a guest session, and no host-only abandon operation has been product-approved. The recovery instruction is therefore not composable across reconnect yet.
  • Trash followed by restore reproduced fork sizes and type/creator but not the guest-project digest. The Files stat result does not expose Finder flags, so the caller could not distinguish a flag change from another hidden identity mismatch. Exact typed re-upload recovered the project; observability did not.
  • ckproject.test-receipt/1 intentionally leaves the product running for semantic assertions, so cleanup requires a separate dismissal/quit and process read. A cancelled build is terminal for its candidate and requires discard/restage. Both are deterministic, but an autonomous recipe must model them explicitly rather than retrying the last verb.

Things known to be wrong, unfinished, or unverified, with enough detail to pick any one of them up cold. Nothing here is being worked on right now; each is parked deliberately.

The distinction that matters in this list is broken (it does the wrong thing) versus unverified (it may well be right, but no one has watched it work on the PowerBook). Unverified is not a lesser problem — several of tonight's bugs lived in code that looked obviously correct.

Nothing on this page is corrected by editing it. A claim that has stopped being true gets a dated line saying so, under the entry that made it. The history is the point: several entries here are worth more for the shape of the mistake than for the fix.

Paths beginning mirror/ in entries dated before 2026-08-09 name the tree as it existed when the evidence was recorded. The standalone project now lives under archive/mirror-standalone-2026-08-09/; production MirrorKit and MirrorKitUI live under now-host/Packages/MirrorKit/. Historical entries retain their original path spelling so the ledger remains an honest receipt.

EMULATOR-VERIFIED, NOT METAL-VERIFIED: host-owned PowerPC application and NOW Extension updater (2026-08-10, codex/host-owned-updater)

The host now treats its validated canonical PPC application and NOW Extension artifacts as an update catalog. It advertises release version, exact build identity, byte count, SHA-256, channel and explicit signature state after rechecking the artifact against its generated sidecar. The guest Connection page distinguishes a different scratch build of the same version, requests the exact offered build, receives it over the existing fork-preserving MacBinary lane, verifies SHA-256 and Finder identity, and uses FSpExchangeFiles to keep rollback bytes. Application replacement exits through normal teardown before a Process Manager relaunch; Extension replacement reports restart-required and makes the retained old copy non-INIT. The guest also warns when the active resident release differs from the version the application expects.

The flow is still explicitly unsigned. SHA-256 detects corruption against the connected host's manifest; it does not authenticate a publisher. Unsigned installation therefore requires a local modal confirmation, and the shared console/wire command cannot bypass it. Release signing remains open until a pinned trust root, key rotation, revocation and recovery policy are chosen.

The native suites cover SHA-256, exact-build comparison, unsigned-consent ordering, host catalog validation, contract conformance and version-copy drift; the host suites and Debug/Release app builds pass, as do the PowerPC, 68K, Extension and rig cross-builds. Mutation runs watched stale artifact acceptance, bare signed-flag acceptance and remote unsigned-consent bypass fail by name.

One private mac99/OS 9.1 acceptance exercised the actual host and guest UI from an updater-capable 0.1 application with resident 1.1. Connections displayed the host's different 0.2 scratch build and unsigned 1.2 Extension, plus the 1.1/1.2 mismatch warning. Local confirmation served the exact SHA-addressed application artifact; the guest exchanged it, completed normal teardown, relaunched, and reconnected as 0.2. A second local confirmation exchanged the Extension and returned restart-required. After a guest-clean shutdown and cold boot of the same disk, the guest reported resident 1.2 active with all 511 capabilities and fingerprint 085c4ebf8457.

That acceptance exposed a terminal-state UI defect: after successful Extension installation the footer could retain Downloading... and leave Install enabled. The guest now retains restart-required state until restart, renders Extension installed. Restart this Mac to activate it., and disables the button. The exact source guard was mutation-tested and the corrected guest cross-builds, but the corrected wording was not re-driven through the live UI: the post-relaunch act plane refused generic Control Manager actions even while the updater itself remained connected. Checksum refusal, low-disk refusal and rollback-file recovery also remain suite-tested rather than emulator-accepted. Physical-hardware acceptance still owes the whole lifecycle on the PowerBook 1400c.

Product release 0.2.0 is stated coherently in the shared guest header, PPC vers resource, host catalog and both host build paths; the wire-contract revision remains independent. Main-reference hooks reject incoherent copies, same-version product landings and resident changes without a strictly advanced resident version plus a verified shared bake. The 1.2 shared bake passed exact guest fingerprint/capability checks, all census probes, guest-clean shutdown, clean HFS+ volume and qemu-img check; its committed receipt accounts for the current shared oracle.

RESOLVED DOCUMENTATION CURRENCY: onboarding, Mirror scheduling, and resident versioning are in the web guides (2026-08-09, codex/pre-alpha-docs-audit-plan)

The host gained a guided PowerPC setup portal and classic-media builder, while Mirror gained session-wide work admission, symmetric invalidation hints, and coherent generation publication. Resident releases also gained one shared major/minor identity and a main-reference landing gate. Those changes landed in main after the first documentation pass and were absent from the curated web guide even though engineering records existed.

The setup portal is now a gated cross-product row in Core features, an early Getting started path, a user how-to with three screenshot slots, and a human developer architecture page. Mirror user and developer pages now explain human-over-ambient admission, non-preemption, hint coalescing, gap repair, cadence fallback, and generation-safe publication. Resident documentation now names contract/resident_version.h and its landing rule. The review preserves the evidence boundary: setup through a classic browser and the PowerBook Mirror latency target remain unverified on hardware.

RESOLVED DOCUMENTATION UX: navigation is incremental and Extension coverage is feature-first (2026-08-09, codex/pre-alpha-docs-audit-plan)

The initial site performed a full document load for every internal link, and the User guide placed NOW Extension ahead of the first-connection tutorial. Its main Extension table then exposed P0–P8 implementation planes rather than answering the reader's first questions: what works normally, what needs the Extension, and what remains experimental.

MkDocs Material instant navigation now replaces the document content for internal links, while the Mermaid initializer subscribes to each navigation update. User-guide navigation begins with a Getting started group: connect a classic Mac first, then review core features and Extension coverage. The coverage page now renders 15 user-facing feature rows from docs/feature-catalog.yaml, including app-only coverage, whether the Extension is required, and current maturity or evidence. The technical P0–P8 inventory remains in developer documentation. The docs gate maps every human-facing Extension row back to the complete resident capability inventory and its mutation suite proves that navigation, script lifecycle, or capability coverage cannot be removed silently.

Corrected later on 2026-08-09: that 15-row matrix was still incomplete. It named all nine Extension outcomes but compressed the 14 application modules into six broad rows, omitting iCloud, Chat, Diagnostics, Networking, MCP, and Logs as distinct features and hiding Hardware and Software behind one "inventory" label. The page now renders a 14-row application table and a separate nine-row Extension table. The gate compares the application row IDs exactly with docs/module-manifest.yaml, while the existing capability check continues to compare the Extension rows with P0–P8; its mutation suite has watched both omissions refuse by name.

Corrected again later on 2026-08-09: “initial alpha availability” was being presented as a tag even though alpha availability is the default, and the text incorrectly said the NOW Extension shipped separately. The profile and navigation now say alpha; included PowerPC pages and ordinary module pages no longer receive redundant availability banners. Optional and excluded states remain called out. The Extension is documented as a component included in the alpha bundle but optional to install, and the first-connection tutorial introduces that choice before sending the reader to the renamed Core features page.

PLANNED DOCUMENTATION DEPLOYMENT: the app repository will publish docs.newoldworldmac.com independently (updated 2026-08-11)

The selected public documentation origin is https://docs.newoldworldmac.com/. The application repository will own its GitHub Pages build and deployment once it is public. The separate newoldworld-web repository owns a stable link to that origin. Documentation changes therefore do not require a new website container image or Azure Container App revision.

Activation remains open: verify the custom domain, add the Pages workflow, change the deployed base path to /, set the canonical origin and future security-contact expiry, configure DNS, and prove the published URL. Until then, the local framework remains a standalone /docs/ preview and release mode correctly refuses publication. No submodule, copied generated working tree, or cross-repository commit automation has been adopted.

RESOLVED DOCUMENTATION OWNERSHIP: developer and coding-agent prose now have separate owners (2026-08-09, codex/pre-alpha-docs-audit-plan)

The original developer entry path mixed architecture and code-reading material with instructions specific to automated sessions. That was not only a navigation problem: the same reader had to move between explanatory prose and imperative agent protocol, and future edits would have duplicated technical claims to serve both.

docs/developer-guide/ now addresses developers digging into the code: mental models, source tracing, debugging, architecture, implementation workflows, and verification rationale. docs/agent-guide/ is a smaller operational overlay for coding agents: authority and scope, shared-repository protocol, platform routing, change routing, mutation evidence, and handoff. Agent pages link to the developer documentation owner instead of restating it. The docs gate validates the folder/audience distinction and its mutation suite proves both directions are rejected; prose ownership still requires review because metadata cannot prove that a paragraph serves the right reader.

PLANNED RELEASE CONTROL: documentation has a feature profile; runtime flags do not yet exist (2026-08-09, codex/pre-alpha-docs-audit-plan)

docs/feature-catalog.yaml now declares the alpha product boundary: the PowerPC Carbon guest is included, NOW Extension is optional, and the stale NOW-68K/pre-Carbon build is excluded. MkDocs renders those states on owning pages and generates the public release table and P0–P8 extension inventory from the catalog. The documentation gate rejects an incomplete profile, a page bound to an unknown feature, or an extension capability list that differs from contract/peek_table.h; its mutation suite has watched each refusal run.

This is documentation control, not runtime control. No application feature-flag system reads the catalog yet. The future implementation must either consume the reserved classic.pre-carbon key and active-profile default or replace the catalog as the single authority. Landing a second runtime-only availability matrix would recreate the release drift this gate is intended to prevent.

TESTED, NOT METAL-VERIFIED: the PPC LAN onboarding portal has not run in a classic browser (2026-08-09, codex/onboarding-portal)

Connections now has Set Up a New Mac…. It starts the configured NOW wire listener and a separate temporary HTTP listener, displays a detected LAN IPv4 and actual HTTP port, and serves an HTML 3.2-shaped fixed-route page. The page offers the canonical PPC MacBinary when installed, a per-request MacBinary New Old World Prefs aimed at the accepting interface and configured wire port, the optional standalone CodeKitten IDE, the optional NOW Extension, and explicitly enumerated files from the external dependency store. A missing CarbonLib row can download the known 1.6.1 archive directly into that store after verifying its published SHA-1; NOW wraps the unchanged StuffIt bytes in MacBinary with SIT5 / SIT! metadata before serving them. It can now also generate one setup disk: native packages and generated preferences on a bare HFS Plus volume, inside an uncompressed NDIF image with rohd / ddsk metadata, inside MacBinary for HTTP. The recommended /now/setup.img route uses the classic MacBinary MIME type without a forced .bin attachment; /now/setup.img.bin preserves an explicit envelope fallback. Unknown routes and mutation methods are refused; quit stops the listener. The product and server contract are in onboarding.md.

Updated 2026-08-09: the original image-sizing rule admitted both a native CarbonLib.bin and its matching StuffIt archive, then reserved twice the combined package bytes plus 4 MiB. That produced a 19 MiB uncompressed disk for about 5.4 MiB of installed files; none of those extra bytes were Xcode or Swift content. Known dependency representations are now deduplicated with the native file preferred, and the image is payload plus 1 MiB, rounded up with an 8 MiB floor. The onboarding sheet now makes each optional installed item a selection, explicitly rebuilds one cached image, and shows the live image's name, disk size, transfer size, time and contents. Selection changes are reported as pending until a successful rebuild replaces the bytes served by both setup-image routes.

Updated again 2026-08-09: the 8 MiB floor above was still transfer padding, not a property of the selected files. It has been removed. The builder now restores the selected native files into a staging directory, measures that directory, creates HFS against the measured content, and uses the formatted volume's own free-block count to tighten the result to at most 64 KiB free. The realistic app, extension and CarbonLib fork sizes produce a 5,607,424-byte raw disk with 36,864 bytes free. The MacBinary envelope adds only its header and block padding to those filesystem sectors.

Host acceptance 2026-08-09: the 457823dd release build's item selection, explicit rebuild, image details and content-fitted result were exercised and accepted after the earlier 2.6 MiB of free space was removed. This closes the host UI and image-sizing acceptance question. It does not close the classic browser or physical-Mac boundaries named below.

The focused host tests pass. Two use the real temporary listener over loopback to fetch the page, application and preferences, and to exercise unknown-route and POST refusal. The preference test separately pins the MacBinary header, CRC, Finder type/creator, big-endian V1 magic/format/port, and host field. The asset tests pin local-over-bundled precedence and dependency discovery. The new dependency tests pin explicit catalog enumeration, refusal before write on checksum mismatch, and the acquired archive's MacBinary name, data-fork boundary, Finder type and creator. A general encoder test pins both fork lengths and their independently padded regions. Each new guard was also watched failing against the mutation it names: a byte-swapped preference port, POST being admitted, and bundled assets taking precedence over the operator's local package store; additionally, a substituted Finder type, bypassed checksum comparison, and restored StuffIt archive link each produced the named failure. The new setup-image integration test builds and mounts the actual filesystem, then verifies both forks and Finder metadata; changing the NDIF Finder type from rohd to dImg produced its named failure. The 8 MiB capacity assertion, native-over-archive choice, and selection-to- served-image test were separately watched fail against mutations to the image floor, representation rank, and selection setter. The replacement content-fit guard was then watched fail when the old 8 MiB floor and 2.6 MiB free-space allowance were restored together: it named the 8,388,608-byte carrier as exceeding the sub-6-MiB bound.

scripts/test-all exits 0 after the HFS/NDIF checkpoint: staged-image discipline 28/28, native tests 149/149, MirrorKit, all guest/resident/instrument cross-builds, the complete host suites, and the Xcode app target in Debug and Release all pass. Stage 6 skips honestly because NOW_GUEST_LIVE was not set; nothing in that run reached a Macintosh.

Disk Copy 6.3.3 on the Mac OS 9.1 QEMU guest mounted a generated setup image and exposed its Read Me, application, preferences, extension and Dependencies folder. This is emulator evidence for the image carrier, not a physical-metal result. A later CodeKitten-bearing image reached Disk Copy through the same wire and odoc path, but the pristine snapshot presented Apple's first-run license; it was not accepted automatically, so that image's guest mount remains unverified. Direct deployment proved the payload's forks and identity, then exposed CodeKitten's pre-existing early process exit. What remains unverified is the browser boundary: no Netscape, Internet Explorer or Classilla-era browser has yet downloaded /now/setup.img and automatically decoded its MacBinary envelope; no physical PPC Mac has installed the generated preference file, launched the guest, and sent the hello that makes it appear under Active. CarbonLib remains external: NOW has pinned the mirror bytes and independently published checksum, but has not established Apple redistribution permission. The builder can use open-source unar/XADMaster on the host, and the test kit avoids that runtime dependency by carrying a locally prepared native CarbonLib.bin; a release-pinned helper bundle and its update policy remain unverified. The product intentionally does not create StuffIt: the HFS/NDIF disk is the single-package carrier instead.

There is a third category, and it is not on this page. known-wrong.md is the register of things NOW knowingly ships that disagree with the machine, or knowingly does not do — each with the measurement, the argument for leaving it, what closing it would cost, and who decided. Several of its rows draw their evidence from entries here and link back rather than restating them. The split is by what the reader is being told: broken-or-unverified means nobody chose this, and a row over there means somebody did.

PRE-MERGE CONSOLIDATION: Mirror is NOW-owned; the standalone product is archived (2026-08-09, codex/mirror-session-teardown)

The production semantic model and renderer now live at now-host/Packages/MirrorKit/; the reusable resource parsers live at tools/asset-pack/. The complete copied standalone app, guest, residents, oracle, probes, raw documents, and imported asset bytes are preserved beneath archive/mirror-standalone-2026-08-09/, where no active build, staging route, or runtime asset resolver enters. archive/mirror-lineage.md is the ownership map and docs/research/mirror/README.md keeps the Finder, GWorld, A5, desktop, and cursor conclusions discoverable outside the archive.

The external asset-store seam is now complete across the renderer, extractor, and MirrorKit gate: NOW_MIRROR_ASSET_STORE selects the store, the host keeps a persisted pack identity rather than a hard-coded path, and extraction creates a finished timestamped pack by default. The 276-test MirrorKit suite passes both with the current extracted pack and with NOW_MIRROR_ASSETS=none; the native guest suite passes after changing its legacy parity census to read archived source deliberately and its human-action census from the active package.

The MCP barrage handoff was reconciled from clean commit b02090f1, including the positive-size upload fix, protocol-v11 guest-reference semantics, complete 42-tool conformance receipt, and updated coverage. The combined full gate then passed before independent compound-engineering review. That review produced bounded correctness fixes at b40a8fb5, behavior-preserving simplifications at 3e7ab897, and a corrected gate/bake dependency in this plan's published receipt. The final payload gate at 781f6281 passed 28 image-discipline checks, 149 native tests, MirrorKit, all PPC/68K/resident/instrument cross-builds, and the host Debug/Release gate. Live and metal stages skipped as designed. The MirrorKit gate enforced the discovered external pack-2026-08-07b/Resources pack and repeated its suite with NOW_MIRROR_ASSETS=none.

The final combined MCP source revision is 858cfb51; its deterministic gate passed before the bake, and its developer-signed host artifact is staged at /private/tmp/now-host-858cfb51/New Old World.app. The shared PPC oracle was baked and promoted from that exact source revision after the only foreign QEMU clone shut down through Finder and left its own HFS volume clean. The promoted ~/Lab/Assets/os91-qemu/now-mirror-stage.qcow2 is 630,587,392 bytes, SHA-256 f48398203008b0ae220d9109e662054106e98864de860c1dc2a9fbfbc8a4c345. The guest reported resident lifecycle active, capabilities 511, source manifest f0ce0fa3ea33, and fingerprint 18203af3657f; the ABI self-test and full 14-probe survival census passed, QEMU exited through real guest shutdown, and the promoted HFS volume is clean. The previous shared image is preserved as now-mirror-stage.qcow2.bak-20260809-2.

Michelle installed the exact 908c3fc0 metal-test guest and resident on the PB1400c and reported the candidate "looks good enough" to land. That is the bounded final acceptance for this branch, not a blanket promotion of every open Finder or application-P3 row to Metal-verified; those entries retain the evidence levels stated below.

NOW main was atomically fast-forwarded to 908c3fc0. The parent corpus main was separately fast-forwarded to 3f923561, adding the A5-world and cursor-position findings and preserving the semantic Finder lineage. The corpus-wide tools/data check passes with the exact NOW tree available for its typed doc:now/... evidence references. Neither foreign dirty shared checkout was modified.

FIXED HOST-SIDE, NOT METAL-VERIFIED: positive-size MCP guest uploads always refused on this host (2026-08-09, codex/now-mcp-audit-barrage)

The live Luna barrage reproduced the earlier four-byte conformance symptom at 52 and 2,466 bytes. GuestUploadStagingStore reads volumeAvailableCapacityForImportantUsage; Foundation reports zero for that key on this Mac while ordinary available capacity is about 714 GB. The policy therefore concludes that every positive-size reservation exceeds capacity. A zero-byte begin and commit succeeds, proving the rest of the lane is reachable and making the false capacity answer the immediate blocker.

Fixed on 2026-08-09 at de4aedf2. PrivateStagingCapacity now supplies one answer to both GuestUploadStagingStore and AgentDownloadStore: a positive important-usage value wins; zero or unavailable falls back to a nonnegative ordinary capacity. The five-percent reserve did not change. The new resolver guard was watched fail before the fix, then passed with both store suites.

A private Mac OS 9.1 VM then accepted a four-byte upload through the real stdio companion and same-UID host socket, confirmed all four bytes and their guest CRC, finalized by same-folder rename, and statted the resulting TEXT file as four data-fork bytes. That is emulator verification, not PowerBook metal verification. See audit-report-2026-08-09.md F-003 and now-mcp-audit-barrage.md.

FIXED TEST COVERAGE: live upload conformance used the wrong digest (2026-08-09, codex/now-mcp-audit-barrage)

The full-surface recipe said it uploaded now\n, but its hard-coded SHA-256 was not the digest of those four bytes. F-003 had made this invisible by refusing at reservation. Once staging worked, append accepted all four bytes and commit correctly returned now-files-integrity-failed. The conformance gate still considered that explained refusal an acceptable answer, so the upload family could be exercised without ever proving a successful commit.

The recipe now owns one Data("now\n".utf8) value and derives its declared length, SHA-256, and base64 chunk from it. The spawned-client no-host run covered all 42 advertised tools; an identity-checked Mac OS 9.1 VM run then served upload begin, append, and commit, with zero failed or uncovered rows. See F-008 in audit-report-2026-08-09.md.

UNVERIFIED PRODUCT BOUNDARY: full agent access may upload bytes without a one-time host-file approval (2026-08-09, codex/now-mcp-audit-barrage)

The barrage's negative transfer prompt assumed all modern-host-to-guest files required now_transfer_approved_artifact. That is not the implemented model. The approval receipt governs a host-selected private file whose path is never given to MCP. The separate now_guest_files_upload_* family accepts bytes the caller supplies and commits them under the guest's full agent-access ceiling. A Codex worker with filesystem access therefore read a zero-byte file from the modern Desktop and placed it on the guest without minting or forging a receipt. The routing-skill follow-up repeated the authority path with a synthetic 25-byte Desktop fixture and verified its guest digest, so the result is not an artifact of the original empty file.

The implementation is internally explicit; the unresolved question is product authority. If full guest agent access is meant to authorize any bytes the agent can already read, this needs clearer paired documentation and the negative test was wrong. If every modern-host file needs an additional human gesture, the upload family is a second write path that does not enforce that decision. No policy changed in this audit. See F-004 in the audit report.

RESOLVED: a repo-scoped router makes bare Macintosh tasks select NOW (2026-08-09, codex/now-mcp-routing-skill)

After adding initialize instructions, one resource, one prompt, and the obvious now_list_machines entry point, five of seven bare Luna tasks still made zero NOW calls. Four reasoned about an empty workspace and one searched the modern host. Prefixing only “Use the NOW integration on the connected classic Macintosh” made all five repeats call now_list_machines first.

The server guide helps after selection; it is not a client intent router. The repo-scoped .agents/skills/now-mcp skill now supplies that missing route without duplicating tool schemas. In a fresh isolated Luna follow-up, all seven original bare prompts called now_list_machines first; none detoured into TimBotTu, an emulator harness, or the modern host. H0, R1, R2, A1, M1, and N1 completed correctly. X1 entered NOW and completed its upload, but exposed the separate modal-action issue below.

RESOLVED: semantic UI actions publish and enforce one exact gesture grammar (2026-08-09, codex/now-mcp-action-contracts)

The X1 skill run uploaded and byte-verified its text file, launched SimpleText, opened the File menu, and populated the Open dialog. When the dialog did not settle, it mixed retained semantic entities with direct-observation references and tried seven unsupported gesture names until the evaluator interrupted the run at 332 seconds. Typed refusals were clear, but the route from modal state to the permitted action vocabulary was not.

The skill first kept the reference families separate and stopped after two failed actions. The server now closes the underlying grammar gap: now_semantic_ui_act publishes the exact 16 gestures, one required-argument branch per gesture, and the same contract rejects malformed or cross-gesture arguments before the host. Dialog guidance points to snapshot.surfaces[].items and its 1-based number; Finder guidance states that Standard File rows are outside that action.

In a fresh isolated Luna repeat, the worker recovered from two refused direct control calls to gesture: dialogItem, corrected item to itemIndex from the boundary error, and dispatched the Open button without inventing gesture names. The task still failed for the separate list-selection gap below.

RESOLVED: direct and pixel tools now state their evidence-ladder rank (2026-08-09, codex/now-mcp-standard-file-audit)

The first-contact guide and repo-scoped skill already preferred structured and retained semantic state over direct probes and pixels. The escalation tools' own descriptions did not. That local omission mattered in the barrage: X1 had a retained dialog item but tried the direct family twice before recovering, and the earlier H1 run escalated to pixels without first using retained state.

now_observe_elements now identifies itself as targeted direct observation for incomplete retained state or fresh direct-action references, and says not to launch it in parallel with the retained snapshot. The snapshot description also requires reading its coverage before deciding whether to escalate. now_capture_screen now reserves pixels for genuinely visual facts or facts semantic evidence cannot answer. The registry guard was watched fail before both description rounds changed.

The second round came from a live private-VM A/B, not prose review. Under the first wording Luna launched process, retained, and direct reads in parallel. Under the sequencing wording the identical prompt used process and retained state only, even when the isolated harness failed to load the routing skill. A control with the skill readable began with now_list_machines, started and read semantic UI, then cross-checked processes; it used no direct probe or pixels. No runtime action or authority changed.

BROKEN: Standard File dialogs expose buttons but not selectable file rows (2026-08-09, codex/now-mcp-action-contracts)

SimpleText's Open dialog is richly present in the retained surface: Open is dialog item 1, Cancel is item 2, and the surrounding controls carry titles, state, and geometry. The file list itself is only an unnamed userItem with no rows, current selection, or row action. Direct observation has the same boundary.

The F-009 repeat uploaded and byte-verified Luna Contract.txt, set the Open dialog's text to that exact name, and dispatched the semantic Open item. The dialog opened the previously selected Apple DVD Player Read Me instead. Subsequent Finder and key attempts had no semantic row to ground against and the evaluator interrupted the run after about 226 seconds.

Do not infer selection from the editable text or add another gesture synonym. The design review in now-mcp-standard-file-review.md keeps this literal dialog-row problem open: Navigation Services selection requires an opaque NavDialogRef that the observed window does not reveal, so a generic fix is a resident-contract project rather than a bounded cleanup. The raw trace is under docs/local/now-mcp-barrage-2026-08-09/f009-action-contracts/.

UNPROJECTED: the PPC guest can open a document semantically but MCP cannot ask it (2026-08-09, codex/now-mcp-standard-file-audit)

The guest's closed aesend command already serves odoc: one exact process serial, one HFS document, one kAEOpenDocuments Apple Event, with sent reported distinctly from performed. The MCP projects neither that bounded outcome nor a document-shaped wrapper, so agents fall into F-010's opaque Open dialog for a task the OS can perform directly.

The recommended bounded slice is now_open_document(processReference, path). It should revalidate the opaque process reference, enforce the existing root-relative guest Files policy, and delegate only odoc. The guest contract needs an accretive share-relative path form; the host must not build an actionable full HFS path from rootLabel, which is presentation data. Generic aesend, quit, and printing stay off this row. NOW-68K should report typed unavailability. This is F-011 in the MCP audit and awaits explicit approval.

RESOLVED: inventory paths are not launch keys (2026-08-09, codex/now-mcp-standard-file-audit)

The fixed cross-model A1 made the handoff problem repeatable. GPT-5.6 Luna and GPT-5.4-mini both found SimpleText in now_software_inventory, then tried to feed its returned HFS path to now_launch_software; both recovered only after a typed refusal. Gemma and Qwen chose the exact name directly.

The launch boundary is correct: paths are not accepted, exact names are resolved against the current guest catalog, and ambiguous-name refusals mint the only launch references. Its descriptor is not. It says guest paths are never “accepted or returned,” even though the adjacent inventory tool returns one for every entry. F-013 now states the handoff locally and pins it in the registry without changing an argument or guest behavior. The guard was watched fail before the descriptor changed, then passed.

TESTED HOST/GUEST, NOT EMULATOR- OR METAL-VERIFIED: Mirror work now has priority, attribution, and invalidation generations (2026-08-09, codex/mirror-latency-state-plan)

Mirror request-shaped work previously entered several independent paths. A human gesture could wait behind a queued Finder complement or scene without the host naming the blocker, while direct and brokered acts could race the same resident request cell. The historical 9–12 second act_yield defect remains a separate closed cause; the matching number alone does not show that it returned.

One session-owned scheduler now admits Mirror gestures, scenes, content drains, Finder pages, capture and the principal list/process request families. Human gestures remain FIFO with each other and outrank ambient work that has not started. Active classic-Mac work is not preempted: it completes at its safe boundary, and automatic Finder observation is one page per slice with an explicit 1,800 ms timeout. A bounded work ledger reports admission, guest round-trip, settlement and publication separately through the existing Mirror metrics projection, including the current blocker and oldest human wait.

The PowerPC guest drains P5 transition evidence through one normal-context coordinator and emits optional mirror.invalidate generation hints. The host attributes each event to the transport session that delivered it, coalesces follow-up observation, and invalidates its delta baseline on gap/unknown evidence. P5 observes only the currently armed A5 world: a front switch away from that target is sampled evidence, not a complete event stream. The cadence poll remains the compatibility and liveness authority.

The host did not move state ownership to the guest or create a second snapshot model. Its existing session-pinned engine records exact structure, semantics, Finder, visibility and content generations plus the reason for each immutable publication, and refuses old enrichment from a changed base.

This work is Tested locally, not emulator-verified or Metal-verified. The full scripts/test-all gate passed: 28 staged-image checks, 150 native tests, MirrorKit, both guest families plus the extension and auxiliary cross-builds, the host suites, and Debug and Release app builds. Its live-guest stage explicitly skipped because NOW_GUEST_LIVE was unset. Priority ordering and transition-gap guards were also watched failing against their claimed mutations before the unmodified tests passed.

The open acceptance item is the exact-revision native Mirror campaign: close, front, double-click/open, selection, scroll, menu and drag during deliberately slow ambient slices, with paired guest evidence and distributions. In particular, the plan's claim that ambient work contributes no more than 2,000 ms to a PB1400c interaction remains unproven.

FIXED HOST-SIDE, NOT METAL-VERIFIED: Stop, disconnect, and guest replacement left a Mirror session alive (2026-08-08, codex/mirror-session-teardown)

The reported host could hold two live Mirror sessions from one guest, and Stop left the old scene, engine, Finder complements, queued work, and measurements addressable. Every later reproduction was therefore confounded until the host process restarted. The complete evidence and the correction to the first interpretation are in mirror-stop-should-disconnect.md.

All three events now cross the same destructive session boundary. Stop clears host state synchronously even if the guest never answers its resident-claim release; an active disconnect ends the pinned session; and a replacement active guest gets a fresh engine rather than the prior session's state. Persisted run intent remains separate: if the person did not click Stop, Mirror starts fresh when an active connection returns.

The focused host pass ran 104 Mirror, content-plane, container, and multi-guest tests with no failures. This is Tested, not Metal-verified. The current metal symptom remains open: opening Mirror forces Finder front for 10–20 seconds and Finder then crashes. Michelle has not recently seen the earlier Error 1 guest-application crash. The lifecycle repair makes the next run attributable; it does not claim to explain or fix the Finder crash.

A G3 Wallstreet on Mac OS 9.2.2 reproduced the forced-front and ping-pong without an immediate crash. Its guest log shows the live content target alternating between exactly two A5 worlds with hook churn at each move, then cleanly disarming to A5 zero. It contains no mach activate line. This moves the next diagnostic to the automatic interactive Finder AppleScript complements: disable only icon-roster and visibility reads, retain structural scene/content observation, and see whether the unsolicited fronting stops.

METAL-VERIFIED: P3 drawing trace crashes Finder on the PowerBook 1400c

The four guest policy domains isolated the failure on 2026-08-08. Structure, Finder details, and foreground discovery were exercised independently without a crash. Enabling Trace drawing contents armed P3 at 17:42:11; the resident retargeted from A5 0x319e38c to 0x22c096c, then disarmed. Finder's process identity changed from 0.8257537 before that interval to 0.9633793 after it, which is Finder relaunch evidence while NOW itself continued logging. The setting remains off by default and is labelled experimental. The prior wrong-context and stale-port guards are therefore insufficient; do not treat P3 as safe merely because its application-layer reader is read-only.

The same run contained 13–25 second main-loop gaps after P3 was off. Because it mixed direct Finder use and Mirror driving, those lines cannot distinguish a slow scene inside NOW from another cooperative process monopolising the Mac. Slow scenes now log their measured enumerate/bind/window/control/menu/semantic/ reference phases at the source. A subsequent metal run can therefore attribute the gap without adding unsafe yields inside a foreign-memory walk.

FIXED HOST/GUEST, NOT YET METAL-VERIFIED: Finder can no longer arm P3

The correction is now structural rather than a recommended setting. The host recognises Finder from the process roster (with the legacy name fallback), never sends qdtrace start, releases a prior application target when Finder comes front, and strips any historical Finder display. The PowerPC guest then enforces the same boundary: it accepts only a complete PSN or deliberate front:true, resolves and classifies that process through the shared roster, and refuses Finder before the resident claim or arm cells are written. Raw A5 is now a named refusal because it cannot prove which process it names.

Finder interiors instead come from bounded semantic snapshots. Each exact window carries its HFS path, measured view word, Finder enumeration order, live item bounds and front-window selection. Exact WindowRecord identity replaces the title-keyed cache; duplicate titles and missing exact identity retain the previous complete snapshot rather than guessing. List and small-icon names now render beside their 16-pixel boxes without P3. The provider can follow Finder navigation anywhere on the guest disk and never uses the Files module's share-root contract; it remains scoped to the displayed container and contains no entire contents search.

The first Wallstreet run of that path exposed a latency defect rather than a Finder failure: opening Macintosh HD at 19:53:56 produced its 18-item semantic snapshot at 19:54:09. The measured complement was 12,983 ms because the host read the desktop first, then read the folder in eight-row pages, then waited for type/creator icon-art enrichment before publishing anything. The host now reads the front Finder window first, starts ordinary pages at sixteen rows with an eight-row truncation fallback, publishes every completed page, and never waits for type/creator art before drawing semantic items. Generic host-side icons are the immediate fallback; extracted asset art enriches them later. Completed reads are tracked per container, so a failed desktop or background read cannot continuously restart a successful front-window read.

METAL-VERIFIED CRASH, MECHANISM CONTAINED: P3 offscreen tier crashes Sherlock 2

Two PB1400c / Mac OS 9.1 launches on 2026-08-08 produced the same boundary. After Finder opened Sherlock 2, the host requested P3 (requested=15, content=requested). Roughly six seconds later the resident reported content=active-current; in that same cycle the scene fell from two windows to one and Sherlock disappeared. The guest reported a Type 1 bus error. This is stronger than the earlier emulator captures: Sherlock can produce useful P3 records there, but the same hooks are not safe on this metal/application pair.

Blacklisting Sherlock would encode a symptom rather than a safety boundary. The corrected boundary is the mechanism: ordinary record installs grafProcs only on the exact requested WindowRecord. Only explicit diagnostic probe may install the permanent _QDExtensions trap patch, sweep the application heap for existing GWorlds, or hook an offscreen port. Once the trap patch exists its shim also declines selector wrapping outside the armed probe context, including later record sessions and foreign applications. No application creator or name is denied.

This is Tested, not metal-verified. It narrows the next metal run but does not yet identify which offscreen mechanism caused the Type 1 crash. Because an already-installed QDExtensions patch cannot safely be removed, the first test of the corrected resident must start from a reboot; using probe makes another reboot the honest reset boundary.

The osaErr -1753 seen while restarting Mirror is not the Type 1 crash code. Apple documents it as OSA's generic “script error,” and this host log contains the same refusal in Finder item and visibility complements outside Sherlock's P3-active interval. It remains a separate Finder-complement reliability issue.

Scroll is no longer a roster invalidation. The host translates cached boxes by the guest's live scrollbar delta and projects arrow, page, wheel and resident thumb-drag moves immediately while the same acts travel to the guest. Selection is an exact local-first set scoped to the exact window: click replaces, control/right-click toggles, shift-click ranges from an anchor, and an empty content drag rubber-bands across the same semantic item rects used for drawing. Name-view hit testing covers the visible semantic row instead of only its 16x16 glyph. Renderer and hit tester now share view inference, so legacy or unknown metadata cannot make a drawn list row fall through to the generic "window is already front" path. Command-A selects all, Command-O opens the set, Return/Enter starts an inline rename, and Escape cancels or deselects; each operation is optimistic locally and then sends one exact Finder AppleScript through the bounded direct-act lane. Finder owns window existence, frontness, geometry, view, and eventual control state; the host owns the interior presentation. This latency and interaction correction is Tested, not yet metal-verified.

Two interaction edges remain deliberately open. A right-click currently has classic control-click selection semantics but no host context menu, and an item drag still moves only the item under the pointer even when the local selection contains several items. Neither limitation can make a list row select the window instead.

Paging remains scoped to the one displayed container. Opening another folder starts a read for that directory; a later host-side expandable list must apply the same rule when a row expands. No parent read prewalks its children and no volume is recursively paged.

The 269-test MirrorKit gate and 52 focused host interaction/source tests pass. This is Tested. The next PB1400c run still owes the claim that Finder stays alive and does not front from observation, and a separate non-Finder application run owes evidence that P3 remains useful under its narrower target boundary.

UPDATE 2026-08-08, semantic Finder consolidation. Desktop icons now use the same host-owned Finder state and interaction model as folder windows; they are not a separately retained picture. This also closes two plausible causes of the reported “desktop icons disappear after Mirror has been unfocused” regression. The desktop layout key depends only on screen geometry, so the raw and enriched spellings of the same roster cannot oscillate the cache and restart its complement read. And a transient complete-but-empty desktop answer cannot erase a retained complete nonempty roster in the same session. The host logs and retains the stronger snapshot until a later nonempty answer, a new session, or the person's explicit Rebuild State action. The unfocused case still needs to be watched on metal; the guard is Tested, not a claim that the reported symptom has been reproduced here.

View changes now project immediately from the retained directory roster and per-view layout cache while type/creator icon art fills in through bounded eight-item pages. Optimistic selection covers desktop, icon and list views. Confirmed item rearranges update the local semantic position before a fresh roster/layout reconciliation, and scroll remains local-first. This makes the interior a Finder implementation over semantic items rather than a delayed copy of Finder's pixels. It can browse any displayed directory on the whole guest disk; “bounded” means one open container at a time, not a shared-tree or volume-root restriction, and it never recursively pages the volume.

The right-aligned Application menu is now projected from the live application roster when the scene does not carry Finder's system-owned menu. Application selection, Hide, Hide Others and Show All therefore remain available without P3 or foreground cycling. Hide Others and Show All are host-side compositions over the guest's existing per-process hide operation; the host gate covers the composition and read-back interpretation, but no PowerBook run has yet proved all three visibility operations end to end. Semantic Finder acts also place the guest cursor through the new cursor-only cursoract resident operation after a successful file action; that path builds and its guard is mutation-tested, but the new resident has not yet been baked or metal-driven.

The host Mirror inspector now discovers complete extracted asset packs from the documented Lab store, selects the newest valid pack by default, and offers a persisted Mirror Artwork picker. It stores a discovered pack identity, not a hard-coded pack path; an environment override remains authoritative and picker changes take effect at next launch because art is process-cached. Pack extraction and arbitrary-path selection remain subsequent work.

UPDATE 2026-08-09, corrected Finder ownership boundary. The persisted host toggle is Emulate Finder Window Interiors. It does not create a second set of windows: each semantic interior is joined to the exact guest Finder window, whose title, rectangle, stacking, visibility and chrome remain authoritative. The host owns icon/list/small-icon layout, sort, selection, marquee, rename and scrolling inside that shell. It lists only each open folder through the Files contract. File launches, document opens, renames and moves still cross the wire because they change the Macintosh, while interior presentation changes do not.

The mode is bounded by the configured guest share. Share entire boot volume makes that share the disk root, which is the expected whole-disk Finder setup; a narrower share remains a real authority boundary. Only an open directory is paged and no parent prewalks its descendants. The existing guest-follow mode is unchanged when the toggle is off and can still mirror a Finder window opened outside the Files share. These two paths must not be described as one boundary.

The host projection carries selection in the rendered scene, uses the same semantic rectangles for clicks and rubber-band selection, keeps scroll-wheel and thumb motion local, and retains icon positions across view switches. Right/control-click now targets an unselected item without manufacturing a multi-selection and preserves an existing multi-selection; a contextual menu is not yet implemented. Group drag, expandable list disclosure rows, Finder desktop replacement, and path-title navigation remain open. The inspector's Refresh button discards and rebuilds every open host directory as an escape hatch.

The independent Finder domain and session tests are green: six focused tests cover selected rendering, list-row targeting, stable sort, host menus, multi-window folder navigation, local view/sort/scroll, and the absence of a guest command when opening a folder. This is Tested, not Metal-verified.

UPDATE 2026-08-09, optimistic Finder coupling and catalog fallback. The independent mode now has two explicit development controls: synchronize Finder window opens/closes, and synchronize position/size. Both are local-first. A host open, close, move or resize changes the mirror immediately, sends the existing script or winact request when it has an exact guest identity, and holds that optimistic frame for eight seconds while a later guest scene either confirms or reconciles it. Guest-opened and guest-closed folder windows are joined by exact window identity plus semantic HFS path. Either axis can be disabled independently; this is diagnostic control, not a final product-mode decision. It is Tested and not metal-verified.

Desktop ownership is a separate Emulate Desktop switch. Off means the guest Finder roster and its exact positions are projected; on means the host lays out the bounded Desktop Folder catalog. The Finder roster now explicitly appends disks and trash with their live bounds because every item of desktop does not reliably enumerate those system objects. Enrichment merges instead of replacing structural disk/trash rows. Switching the desktop mode invalidates only that catalog, so the next projection must come from the newly selected owner. Guest-follow folder windows retain Finder's exact live roster; if that roster is absent but its semantic path is known, a bounded file.list of that one directory supplies fallback icons. The current icon cache is reapplied after state-engine projection so a one-complement-behind snapshot cannot blank the window. Neither path walks descendants.

EMULATOR-VERIFIED 2026-08-09: the final host build read the live OS 9.1 desktop as 19 semantic items, including Macintosh HD as disk at Finder's reported bounds and Trash as trash, and the already-open Macintosh HD window carried all 13 visible semantic rows. This proves the desktop system query and guest-follow roster path against Finder, not the still-owed Wallstreet/PB1400c rendering pass.

The host no longer persists Mirror's run intent across app launches. It deletes the retired preference and starts stopped; only --open-mirror or an explicit start in the current process can begin it. A reconnect during that same process can still resume the current intent. Asset-pack provenance moved to the inspector's Asset Packs scaffold and is no longer painted over the mirrored desktop. All of this remains unverified on the PB1400c/Wallstreet.

UPDATE 2026-08-09, Finder opens and view ownership. Emulated view changes are now disentangled from the guest scene by default. Icon, name/list and small-icon switches update the retained host directory immediately, and a later guest poll cannot silently put that host-owned presentation back. The new Sync Finder view type development control opts into the second axis: the local change remains optimistic, Finder receives its measured bare view word (icon, name, or small icon) against the exact HFS window, and later guest scenes confirm or reconcile it in either direction. A request made before the file listing supplies the volume root is retained and dispatched once that state arrives. This is Tested and not metal-verified.

Opening an item from an emulated interior now uses Finder's measured item "X" of window "Y" reference whenever the guest shell exists, with a nested disk/folder reference for a host-only window. cdev control panels no longer go to the application-only launch verb—the same category error already recorded when launch refused a control panel as “not an application.” Script settlement also inspects nonzero osaErr; a Finder refusal can no longer be reported as an open merely because the command envelope said ok: true. Focused tests cover the control-panel route, host-only view retention, optimistic synchronization and guest-to-host reconciliation. Metal remains owed.

LANDING STATUS 2026-08-09: Emulate Finder Window Interiors is Experimental. The host UI and README say so explicitly. Its semantic model, ownership boundary and focused tests are useful enough to retain, but file opening, visible selection, view synchronization and other Finder interaction work have not collectively passed a PowerBook acceptance run. Guest-owned Finder rendering remains the non-experimental path while this mode develops.

CORRECTION 2026-08-09, authority must not seed state. Enabling emulation no longer synthesizes a root-volume window or sends a guest open. The host starts with no emulated folder windows, adopts folders already present in the guest scene when lifecycle coupling is enabled, and creates a local window when a person opens a semantic desktop folder or volume. The local window precedes its optional guest script. A host move or resize made before that script yields a live guest window reference is retained and dispatched after the exact identity/path join; geometry commands are serialized rather than raced into the guest's one act cell.

Finder-window interior and desktop ownership no longer invalidate one another. Changing Emulate Desktop clears and rebuilds only the desktop catalog; changing Emulate Finder Window Interiors rebuilds only the open-window semantic state and rejoins it to the observed guest shells. With position/size synchronization enabled, that reconciliation also adopts guest opens/closes, so geometry synchronization cannot leave a parallel host window set behind. The Finder-card Refresh invalidates semantic catalogs without closing the observed guest windows. The top-level Rebuild State is deliberately destructive: it invalidates any in-flight generation, clears the reducer snapshot and bounded history, content state, displayed scene, icon/visibility/Finder caches and host Finder windows, then begins a new observation. It preserves the operation journal. These corrections are covered by focused host tests and remain unverified on metal.

UPDATE 2026-08-08, empty Application menu. Presence of menu -16489 is not evidence that Finder supplied its rows. An empty or incomplete menu is now rebuilt from the same switchable-process roster used by the agent surface, while preserving the guest-measured title position. The roster honors the process's modeOnlyBackground declaration, so a desktop application with no open windows remains switchable and a declared faceless agent remains absent. The 18 focused Application-menu tests pass; switching and Hide/Hide Others/Show All still require the next metal run.

BROKEN: the Mac OS allocation grew to 43.1 MB during a Mirror session

About This Computer showed Mac OS: 43.1 MB, virtual memory off, and only a 1.1 MB largest unused block during the 2026-08-08 PB1400c semantic-Finder run. New Old World: 7.4 MB is a separate, mostly fixed application partition and was not the suspicious number. The screenshot therefore records severe system-memory pressure; it does not by itself identify the owner.

The resident extension's direct system-heap footprint cannot explain that number. It allocates its shared table and fixed event/content rings once at boot, together well under 100 KiB, and does not allocate per Mirror cycle or Finder target. The highest-frequency remaining suspect is the Finder complement path: every semantic read currently opens an AppleScript OSA component, runs OSADoScript, and closes the component. The descriptors are disposed on the normal and observed failure paths, so component churn or an AppleScript/Finder retention path is a hypothesis, not a proved leak. Repeated Finder crashes and relaunches are a second confounder because their system-owned allocations may not settle like an ordinary application partition.

Do not continue a diagnostic run once the largest unused block is this small. After reboot, record FreeMemSys() and MaxBlockSys() at four boundaries: extension loaded before NOW launches; NOW idle before Mirror; before and after a bounded count of Finder-complement scripts; and after Stop plus Quit. Also run the same-duration Mirror session with automatic Finder details disabled. That distinguishes fixed extension cost, application-partition accounting, OSA-driven growth, and unrelated OS/Finder growth without adding polling to the scene loop.

TESTED: the suspect mechanisms are now separate guest policy domains

The Mirror page now owns four persisted checkboxes with enforcement at the guest boundary: passive application-structure observation, automatic Finder details, drawing-content tracing, and explicit foreground discovery. New and upgraded preference files start with only passive structure enabled. Turning content off withdraws an existing P3 request; turning structure off releases the scene-owned anchor/tree/act claims; a disabled cycle refuses before SetFrontProcess; and a Finder complement refuses before OSA opens.

The host reads the same policy in the schema-1 mirror reply, filters its plane requests through it, and does not schedule Finder complements while the guest gate is off. Automatic Finder scripts carry the typed purpose mirror-finder-complement, so this gate does not disable a deliberate Script command or a Finder mutation the person asked for.

This creates the diagnostic run requested by the PB1400 evidence without removing capability: structure on, the other three off. Applications not yet observed after Mirror starts may appear only as Process Manager rows/window skeletons until they naturally pump or are deliberately brought forward. The PowerPC guest, 68K guest, NOW Extension, and rig helpers build; focused native Mirror tests and host tests cover the policy object, legacy fallback, host projection, complement suppression, and typed script purpose. The isolation controls have now been watched on the Wallstreet and PB1400c; the P3 result and remaining liveness ambiguity are recorded above.

FIXED: the Mirror wrote no log line at all, so its crash could not be investigated (2026-08-08, claude/026-mirror-logging)

Derived rather than assumed, and it is worth stating as a measurement: on 2026-08-08 the areas actually in use across now-guest-ppc/src were sw, act, app, wire, proc, mach, put, get, send, network, files and chat. No mirror, no peek, no plane, no scene, no content. The whole Mirror — the writer lease, five planes, every arm and every refusal — was silent.

The bill arrived the night before. Mirror crashed NOW reproducibly on the PowerBook 1400c; four guest logs from that session were recovered, and every one of them said only what NOW happened to be doing when it stopped. Six plausible mechanisms were then read out of the source and not one could be falsified. The instrument could not see the defect, which is this project's recurring shape.

What is visible now, and which side each line is on

That distinction is the point of the change, not a footnote. Application, task time — the event loop, a wire command being served, a click being handled:

where what it now says
peek.c :: maintain_writer the writer verdict: owned, no resident, resident too short, not canonical, another session
peek.c :: publish_claims_to every arm/disarm request as 0x… -> 0x…, named by the owner that moved it
peek.c :: now_peek_settle all four exits — armed, no resident, request never published, resident never echoed
peek.c :: now_peek_disconnect the planes released when the link went
mirror_module.c the Workshop page created / entered / left / disposed, and show refusals with their reason
main.c the slow observer, and teardown

Resident, draw time, inside foreign processes: nothing. Not one line. The QuickDraw bottlenecks, trap patches and jGNE filter are bounded and allocation-free by construction, and a disk write there would change the timing of the thing being measured and could take the Finder down with it — strictly worse than the silence. Facts only knowable in a hook are surfaced as counters the application reads (NowContentCounters: installs, uninstalls, repairs, skipped_ports, dropped, the three arm refusals, retires), polled from the main loop every two seconds and written only as deltas. Surfaced, never recounted.

The line that matters most is one sentence:

21:04:11 mirror ? writer: REFUSED - binary is not 'New Old World'/NOWo, no plane can arm

That is the failure AGENTS.md records as taking a rename to discover — no plane arms while the resident goes on reporting active with full capabilities, and nothing anywhere named the cause.

Three gates, all watched fail

LoggingSpecTests.testTheResidentNeverLogs reads ext/src/*.c with a derived file list. Half of it the linker already did, and that was measured rather than assumed: the now_log( mutation compiled and failed to link, because the INIT has no such symbol. The File Manager half is the one nothing else catches — FSWrite and its neighbours are traps, so a resident keeping its own little log would link clean and then write to disk inside a bottleneck running in the Finder's context.

MirrorLoggingTests keeps the arm path instrumented, and gates the area registry in docs/logging.md against both guests' sources.

FIXED and TARGET-BUILT: guest launch logs are bounded

Both classic guests now keep 10 recognized launch logs by default under one shared 1–100 policy. The selector protects the current file, ignores unrelated files, orders by catalog creation date with a stable name tie-break, and is mutation-proven. PPC exposes Fewer/More controls and persists format 25; NOW-68K accepts log-retention in its existing dev settings file. Both File Manager adapters cross-build. Actual deletion on constrained HFS and physical machines is still unverified, so that platform proof remains open rather than being folded into the implementation result.

STILL OPEN: nothing here has run on a Macintosh

Both guests compile and the gate is green. Builds, in the ladder's terms — not Tested, not Metal-verified. Every claim above about what the log will say is a claim about what is written down, verified by reading the source with a test. The next metal attempt is what turns it into evidence, and that is the whole purpose of the change.

The area registry had four undocumented words in it

Found by writing the gate, not by looking: act, mach, chat and network were all in use and none was a row in the table that calls itself a "small closed vocabulary". And network was seven characters against a %-6.6s format, so the log line read networ — a grep network over the log, the one thing a tag exists for, found nothing. It is net now, and an over-long tag is a failure rather than a silent truncation.

FIXED and GATED: census killed the Macintosh, and no gate in this repository could have seen it (2026-08-07, claude/024-census-crash)

Michelle drove the round-10 stack and reported it twice: "confirmed: running census crashes workshop", then "i just tried selecting the app switcher and it crashed finder". A screendump showed the desktop with NOW gone, the Finder frontmost, and one orphaned window left drawn on it — a bare frame with a single control and no content.

Reproduced first try on this lane's own clone, and it is one probe:

overview .. ata   present          (nine probes, all fine)
pccard            THE GUEST DIED   (cursor 0, page 1)

The mechanism. census_trap_ready() proves the Mixed Mode ROUTE is open — CallUniversalProc resolved, the thunks flushed, the PPC->68K switch built. It proves nothing about the DESTINATION, and the two were conflated. The PB1400c this code was written against has a PC Card Manager at $AAF0; a Power Mac G4 running the same Mac OS 9.1 does not, and $AAF0 there is _Unimplemented. ata ($AAF1) is fine on the same machine, which is exactly why nine probes pass and the tenth is fatal. The file's own header says the dispatch is "proven", and it is — on a laptop that has the manager.

gather_scsi beside it already had the shape right: it asks Gestalt for a bus and answers absent when the machine says no, "never a select into hardware that is not there". The two Mixed Mode probes had no equivalent. They do now — asking the TRAP TABLE rather than Gestalt, because this same file records gestaltATAAttr answering falsely absent on the 1400c, which is why these managers are reached by trap at all. A trap table cannot lie about its own contents.

It is worse than a crash, and this is the part worth carrying forward. NOW died, and then the anchor worker — a separate process, which survives NOW dying every other time — stopped answering, and both graceful shutdown routes closed with it. The clone had to be power-cut. That is the same shape as the human's second symptom: the damage escaped the process that caused it. So the gate checks the anchor, not just NOW.

NOT the unbaked round-10 resident, which was the standing hypothesis. Two extension changes landed in round 10 with typed deferrals rather than a bake, and this was the first tree holding them together. They are innocent: the defect is in now-guest-ppc/src/census/, application code unchanged for weeks, and the control is clean — the SAME resident (230222d358c1, capabilities 511) was staged before and after, and with one application-side change the census completes. The deferred bakes are still owed; they are simply not this.

Why 1,902 green tests missed it. Every stage of scripts/test-all booted nothing. scripts/test-native compiles the guest's logic with the host cc and runs it here, so a Mixed Mode dispatch to a trap on a real Mac OS is not merely untested, it is unreachable. The metal gates are XCTSkipUnless(NOW_METAL) and aimed at the PowerBook. Every "emulator-verified" claim in this arc came from a lane driving a guest by hand.

The gate is tools/census-survives.py, and it lives in two places for one reason each:

  • scripts/bake-ext-image, unconditional. A guest is already booted there, so the sweep costs ~1 second where a test-all stage would cost every lane a two-to-three-minute boot on every run forever. It refuses, installs nothing, and leaves the VM up like every other failure path in that script — so the image cannot be baked with a build that kills the Macintosh it is baked from.
  • scripts/test-all stage 6, opt-in via NOW_GUEST_LIVE and failing rather than skipping once opted in. For the other half of the problem: a lane changing now-guest-ppc/src/census/ has a VM up and never bakes.

It drives census.request with cursors, not the census command: the command is declared single-page and always gathers cursor 0, so a gate built on it would stop at page one and call a machine safe that is not.

Still open, and it is a hole rather than a caveat. The App Switcher path is not covered. The gate checks that the Application Switcher process survives — it is in ps, so that much is checkable — but "selecting the app switcher", which is the sentence the human wrote, means pulling down the Application menu, and that needs the act plane armed against the Finder. Present is not usable. Nobody has reproduced the Finder crash independently of NOW dying first, so it remains unexplained on its own terms.

FIXED: no bake can be installed — the census ATA probe left the volume DIRTY (2026-08-07, found claude/024-census-crash, fixed claude/024-bake-volume-clean)

The shutdown route was never the bug and tools/volclean.py was right throughout. census_ata_identify built a correct ATA Manager Drive Identify parameter block and left ataPBFlags zero, so mATAFlagIORead — the flag that tells the manager to drive the data-in phase — was never set. The call returns noErr with an empty 512-byte buffer and leaves the device mid-command; the cost lands minutes later on the Shutdown Manager's final "volume unmounted" write. One line in now-guest-ppc/src/census/census_trap.c.

That also explains a symptom recorded as a drive's own quirk: the PB1400c's "the manager answers noErr with an EMPTY IDENTIFY buffer" (2026-07-22) was not the drive having nothing to say. Nobody had asked for the data. The 1400c's ata row should be re-read on metal — it may no longer say "present; drive returned no IDENTIFY data".

Isolated by five emulator trials, one base image and one shutdown route throughout — no census CLEAN; the full 14-probe sweep DIRTY; --probes ata alone DIRTY; that probe narrowed to the single device id that actually answers DIRTY (so it is the SUCCESSFUL Identify, not the timeouts on the absent ids); overview,volumes,drives,drivers CLEAN. A static parameter block was tried first and changed nothing, which ruled out the queue-element theory. With the flag: --probes ata CLEAN, full sweep CLEAN, and a whole private scripts/bake-ext-image run completed and installed.

Why it looked like a shutdown defect: the failing bakes were the first ones to run the new census gate, and every bake before them — three on 6–7 August, all CLEAN — took the identical Finder route. The route's "powered OFF and QEMU exited on its own (6s) — the real thing" line was true every time. Tested on the emulator; not metal-verified.

Two private bakes on 2026-08-07, both of them reaching the end and installing nothing:

  Finder Special(260) item 8 'Shut Down', psn 0.29949953
  the guest powered OFF and QEMU exited on its own (6s) - the real thing, not a quit
  qemu-img check: No errors were found on the image.
  session.qcow2  [untitled] HFS+ (in wrapper): DIRTY — will run Disk First Aid

The first ran with two other QEMUs and an xcodebuild on the Mac, so contention was the obvious explanation. The second ran on a quiet machine and failed identically, which removes it. The route scripts/bake-ext-image documents as "the only route MEASURED to leave a clean volume" powers the machine off for real in six seconds and does not finish unmounting.

This is not new and it is not caused by this branch. It is the same failure that on 2026-08-06 put three dirty images in as the oracle before tools/volclean.py existed to catch them; the oracle in place now is the hand-restored 3-August image. What has changed is that the guard works, so the failure is loud instead of silent — and the consequence is that nobody can bake at all. Every resident change is accumulating behind it, including round 10's two deferred ones.

Both bakes passed the resident gate AND the new census gate before reaching this, so the census work is not what is blocked:

  resident VERIFIED: active, capabilities 511, fingerprint 230222d358c1
  CENSUS GATE PASSED: every probe, every page, and the machine is
  still answering.

The candidate disks are preserved at /private/tmp/nowvm-bake-census/ for whoever picks this up. The base image measures clean before the run, so the bake dirties it; the question is what the Finder's Shut Down leaves unfinished on THIS base, and whether the applet fallback (whose own record is three unmounted volumes) does any better. Nobody should reach for --force here: a power cut is exactly what the dirty bit is reporting.

(The preserved candidate is the evidence that settled it, and it settled it without a boot: it still read DIRTY a day later, long after QEMU had exited and released the file, so the check could not have been reading too early. Its HFS+ header carried writeCount 28945 against the clean oracle's 28399 and a modifyDate at the moment of shutdown — the volume was written 546 times and right to the end, so the writes reached the qcow2 and only the one that sets the unmounted bit did not happen. Read the header, not just the bit.)

MEASURED: a VM snapshot restores in ~0 seconds, and the first two runs said it did not work (2026-08-07, claude/024-census-crash)

There are no integration tests in this repository that boot a VM, and never have been. The reason is cost: a cold boot of the PowerPC guest is two to three minutes. tools/vmsnap-experiment.py measured whether QEMU's savevm/loadvm can stand in, since three snapshots have sat unused in the stage image since July.

savevm 0.2–0.3 s
loadvm ~0.0 s
re-dial after restore 3.3 s / 6.1 s / 24.2 s
cold boot, for comparison 2–3 minutes

All three cases pass: a snapshot taken mid-conversation, a snapshot taken with no host connected, and a restore over a machine whose Control Strip had been quit — which came back, so the restore rewinds the MACHINE and not merely the wire.

The first two runs said the opposite, and that is the finding. Case A failed twice, deterministically, with the guest dialling and immediately closing. Case C failed differently: the connection open and the event loop not turning. Two symptoms, one cause, and it was ours. The guest retries its dial while no host is up, so connections pile up in the LISTENER'S ACCEPT BACKLOG; loadvm rewinds the guest and touches none of them, and the next accept() hands back a socket whose guest-side no longer exists — indistinguishable from a real dial until nobody answers. GuestWire.rebind() closes and reopens the listener before every restore.

Unexamined, that would have become a designed-in belief that "snapshots do not work with a live connection" — a guest property that does not exist. A harness bug wearing the failure message of a real one is the same shape as the parser bug in bake-ext-image that once said "the resident is NOT this build" about a guest that had just identified itself correctly.

Measured on one Mac against one guest, mac99/OS 9.1, wire 18993. A measurement, not a property.

BROKEN: the reserved human block stops LANES colliding, not two agents both building her stack (2026-08-07, claude/024-integration-10)

Block 590-599 is reserved so that allocation can never hand a lane Michelle's ports. It works, and it is the wrong shape for the failure that actually happened.

Round 10 built her stack on 16728/16729 at 20:46, verified it — display present, actselftest -> abi-agreed, guest connected to the host app, screendump showing a clean desktop — and cleared the lane's claim on it per docs/handing-over-a-human-stack.md. At 21:05, mid-report, another session (/private/tmp/claude-501/wt-mf-stack, block 124, run dir /private/tmp/nowvm-mf) booted its own human stack on the same two reserved ports. Mine was shut down, its host app quit, and its run directory's session.qcow2 removed. That session was doing exactly the right thing — including, by the look of the last screendump, the Finder-route shutdown that leaves a clean volume.

Nothing was misconfigured and nobody was careless. Two sessions were each asked for a fresh stack for the same person, and the reserved range is explicitly not an allocation any lane can hold — so there is nothing for a second session to find held and back off from. lane-ports whose --port 16729 would have shown the first stack through lsof, and nothing requires anyone to ask.

How it was noticed, which is the part worth keeping

Not by an error. A tools/qmp screendump was redirected to /dev/null and its output compared with cmp — and because BOTH dumps had silently failed to be written, cmp compared two nonexistent files and returned "changed". The instrument reported motion because it had gone blind. Fourteenth instance of the pattern this arc keeps paying for; the only reason the teardown was found at all is that the next command printed the missing file.

What would actually close it

Not another reserved range. A claim on the human block that a second session can see and refuse against — the same shape as the metal runbook's MetalMachineGuard, which asks whether the MACHINE is free rather than whether a name is taken. lane-ports whose --port already answers it from lsof; nothing calls it before booting into 590-599. spin-up-ppc could, in one check, on the range it already recognises as human.

And the record trap recurred in the same hour

/private/tmp/now-lanes/0124.json files the new stack's QMP socket and run directory under block 124, exactly as docs/handing-over-a-human-stack.md warns. That page's cure is a manual two-field edit after handover, which is a rule, not a floor — and this is the third stack in one day to be filed against its builder's block.

FIXED: the anchor passed as the wire buys a dirty volume and says nothing (2026-08-07, claude/024-integration-10)

tools/shutdown-guest.py takes --port (QEMU's hostfwd to the anchor worker) and --wire (the host listener the guest dials out to). They are different ports, they are adjacent by construction in tools/lane-ports, and they sit side by side in a run directory's ports file. Give the anchor to --wire and the script tries to bind a port QEMU itself holds, reports

  no guest dialled port 17072 ([Errno 48] Address already in use)
  the Finder route did not take; falling back to the applet

and takes the applet fallback — which shuts the machine down and leaves the volume marked mounted. tools/volclean.py then reads DIRTY.

The failure text cannot distinguish "wrong port" from "dead guest", so the diagnosis lands on the guest and the transposition is never suspected. Round 10 did this while pruning a stale VM; the image was a session clone on its way to deletion, so it cost nothing, and on a shared image it is the dirty-image class arriving through the tool written to prevent it.

Now refused (exit 64) when the two are equal, naming both and what each one is. Watched fail by mutation in both directions.

docs/handing-over-a-human-stack.md already stated the invocation correctly. It was written down and nothing checked it — the gate is the floor under the sentence, not a replacement for it.

STILL OPEN: the applet fallback's dirty volume is unmeasured as a rate

The fallback is documented as "does not reliably FINISH one — three images preserved after it were still marked mounted". This is a fourth instance and the first with a named cause that was not the guest's. How often the applet leaves a volume dirty when it is the CORRECT route chosen for the right reason — a guest with no act plane — is still not a number anybody has.

FIXED and GATED: the render was drawing the guest's own pixels, and nobody decided to (2026-08-07, claude/024-no-pixel-islands)

ScenePoller fetched the guest's real framebuffer bytes over the wire (wire.captureRegion) onto Scene.Window.island, and SceneRenderer drew them in place of the content for any window without a named item roster. NOW had gated over-the-wire pixels as a post-stability enrichment; the feature arrived anyway, as a passenger inside the 1 August wholesale vendoring of the Mirror subproject. Nobody crossed the gate — the import had no step that asks what came with it, and no individual diff looked wrong. The archaeology is the-drive-and-the-islands.md.

Michelle's ruling: the islands are prior art only, for a later deliberate re-implementation. Removed from the live product; kept whole and inert in archive/pixel-islands-2026-08-07/ (.txt, so nothing can build it), following the archive/mirror-port-2026-08-01 convention.

Nothing on the wire changed. island was never encoded, the contract never mentioned it, and both guests are untouched.

What it did NOT void — derived, and the derivation contradicted the analysis

The analysis feared it "potentially inflates every render score this arc has produced". It inflated none of them, and the check is one command: every ScenePoller in this tree is constructed in Mirror's own development tooling (MirrorApp ×3, MirrorOracleKit) and now-host constructs none, so NOW's host — whose scenes come off NOW's own wire, where island was never encoded — always had island == nil. The corpus agrees: 107 scene JSONs, 148 scene records, every one source: "peek"; zero axtree or observe scenes, which are the only planes ScenePoller produces; and no occurrence of the string island in any JSON under ~/Lab/Assets/now-mirror-assets/.

So Sweeps A–D, the integration rounds and Michelle's own drive all stand. What the islands voided was the sibling project's dev-tool output, which this arc never scored against. Recorded rather than quietly corrected, because the reasoning was sound and the conclusion was wrong: it reasoned from the code without checking which binary ran.

The rule is now gated, not remembered

Michelle, 2026-08-07: "the only time we should be using pixels from the guest is when we have imported those assets as part of our assets pack, so the pixels are provided by the host and not the wire… these rules need to be gated and not violated without explicit approval from me."

GuestPixelsGateTests (host suite) derives both sides from source at test time — origin (no render-path file may both handle pixels and hold the wire), carriage (no type reachable from Scene may be a bitmap), and separation (the deliberate, labelled screenshot path stays out of the render). Its own docs state what it does not cover. Override: NOW_ALLOW_GUEST_PIXELS=1 with NOW_ALLOW_GUEST_PIXELS_REASON, which writes the reason into docs/guest-pixel-overrides.json so it lands in the same commit; the flag alone still refuses.

Why it needs a gate at all: "make it a high-fidelity mirror of the guest" has a cheapest solution — show the real pixels — and that solution scores perfectly against every fidelity measure anyone can write while destroying both the product and the measurement. It is the shape of failure where the stated objective is satisfied by a route that removes the thing being measured.

STILL OPEN: the import hole, which is the actual root cause

The gate closes the rule. It does not close the route: a wholesale vendoring still imports a sibling's decisions, and nothing asks what came with it.

Half of this is closable by machine and half is not, and the split is worth being exact about:

  • Closable. A pre-commit check can notice that a commit adds a large number of files under a path this repository has never tracked, and refuse unless the commit body carries an inventory. That is mechanical and cheap. It forces someone to look; it cannot tell them what to look for.
  • Not closable as things stand. Checking an inventory against "this project's deferred decisions" needs those decisions to exist somewhere a machine can read. known-wrong.md is the nearest thing — the register of what NOW knowingly does not do, with who decided — and over-the-wire pixels was never a row in it. So even a perfect import inventory on 1 August would have had nothing to check against.

The durable closure is therefore neither of those: it is that a deferred decision worth keeping gets a gate at the moment it is deferred, so an import that crosses it fails a test rather than needing to be noticed by a reader. GuestPixelsGateTests is that for this one decision, seven days late. What a person has to do meanwhile: when vendoring anything wholesale, read known-wrong.md beside the import and say in the commit body which of its rows the import touches.

BROKEN: what a human's own drive found, correlated against her logs (2026-08-07)

Michelle drove the round-9 stack for ~32 minutes and reported fourteen symptoms. This entry records what her logs say about them, because a report is not a durable artifact and hers was the first sustained human drive of this arc.

Sources: ~/Library/Logs/now-logs/2026-08-07 183049.log (her session, 91 lines) and ~/Library/Logs/NewOldWorld/acts.log windowed to her guest build 2af13c079980. Beware the second one: it spans days. A first pass read decodes of 75–109 s from it and attributed them to her; those belong to an earlier session. Window it by build, not by clock time — the lines carry no date.

The act plane has ONE request cell, and it is most of the sluggishness

Her log, repeatedly, while she worked a scroll bar:

ctlact part 20 … refused: another act is already in flight — this Mac's act plane has one request cell and it is taken. Nothing was written.

Nine refusals in ninety seconds. Interaction does not queue, it refuses. This is the mechanism behind "closing some finder windows takes way longer than it should" and the long wait to front SimpleText — neither is a render problem.

A third outcome also appears ~25 times, distinct from success and from refusal: the guest answered without a dispatch row, on parts 10, 20 and 21.

The scrollbar thumb was never dispatched at all

Parts 20 and 21 (the arrows) appear ~25 times. Part 129, the indicator, appears zero times. ctlact is a click verb and a thumb needs press-move-release. "Arrows work, slider doesn't" is not a broken slider — nothing was ever sent for it. That belongs to the drag vehicle, which has only ever been aimed at Finder icons.

Nothing confirmed, for the whole session

Every winact reads settlement=dispatched-but-unconfirmed (dispatch is not guest-visible effect). Zero confirmations in 32 minutes. That is the KW-06 honesty fix working as designed — and it means the host never learns an act landed, so a press has nothing to settle on. The likely reason the pressed state never resolves: its state machine's exit from waiting has no input on this path.

The render tail, measured on her session only

median p99 max
host decode 23 ms 3,152 ms 7,527 ms
guest round-trip 16 ms 2,100 ms 14,743 ms

~22 s worst case combined, matching her "10–30 s to render the new scroll position". The structurally wrong number is 7.5 s to decode 3 windows and 49 elements while the guest's own phase counters are in microseconds. The host is the bottleneck, not the wire and not the guest.

The Finder item roster does not arbitrate against the machine's ink

This is the highest-value finding and it was Michelle's own read.

SceneRenderer arbitrates the display list against controls (semanticOwnsDisplay) and against dialog items (dialogItemOwnsDisplay), and carries the comment "P3 owns unstructured content, while P2 owns concrete drawing wholly."

There is no equivalent for win.items. The only exclusion involving items is against the pixel island (if win.items == nil, let island = …). So a window holding both the machine's drawn ink and our icon roster draws both.

One mechanism, four symptoms: icons drawn twice, labels drawn twice, icons appearing over list-view rows, and the cost — a 49-element window finding 7.5 seconds because the work is done twice.

The fourth symptom is not this mechanism. Measured 2026-08-07, claude/024-items-arbitration: the three fidelity symptoms are this defect and are fixed below. The 7.5 s is not. See the correction in that entry — decode_ms is a bracket, and its own log says which half.

Corroborated independently by the live flicker trace earlier the same day: the Finder's icon-grid boxes flipping semantic ↔ absent with a 0.83 s bounce nothing asked for, and an 8.4 s lag before the content plane followed a view switch.

Understood already, restated so the drive's list is complete

  • Extensions Manager's empty list — its rows live inside a userItem the application draws itself. No ControlRecord, no DITL row, so no state exists to report by any current route.
  • Set Time Zone, Sherlock's components, control-panel icons all show "Bitmap unavailable" — the loud hatch, which is positive proof a drain happened and the bitmap was not in it. Asset resolution, not the plane.
  • "Icon drag fails — nothing said where it is" is the homeIsTrustworthy refusal behaving correctly; the fix is upstream in what makes a home trustworthy.
  • No I-beam over Sherlock's text field — the cursor rule keys on semanticKind == "editText", and that field is almost certainly unclassified, which is the CDEF wall.

And the coordinator repeated the arc's own worst process defect

The shared worktree keen-clarke-4988fc was checked out by the asset-packs lane while the coordinating session was using it. Its reflog holds fe4d8179the silent revert round 7 caught — and from 14:50 onward the coordinator committed nine documentation commits onto that lane's branch without noticing. Nothing was lost, because round 9 merged it. Sixth firing of one-worktree-per-lane, and the first where the coordinator was the one who did it.

FIXED: the Finder roster had no arbitration, and three of its four symptoms are gone — the fourth was never it (2026-08-07, claude/024-items-arbitration)

SceneRenderer resolved the machine's display list against controls (semanticOwnsDisplay) and against DITL rows (dialogItemOwnsDisplay), five call sites between them, and against win.items not at all. So a folder window holding both the machine's ink and our Finder roster drew both. Michelle found it driving: "I can actually see the icon being selected underneath it … we're rendering the whole icon twice."

The rule chosen, and why it is not "one side wins"

The cell is split, the way this renderer already splits a check box into its mark and its label. Neither half is a new mechanism.

  • The icon box is the roster's, and now joins semanticFrames. It is rung-3 art addressed by identity — kind, type and creator picking a bitmap out of IconAtlas. The replay cannot better it: the Finder's own icon reaches this side as an unjoined blit, which by the ProvenanceLadder's own words "carries geometry and no pixels", so the strongest thing it can put in that box is a generic document stub or the marked-unknown hatch. Excluding it is not overruling the machine; the machine's pixels are not on offer, and the pixel-islands gate is that they never will be off the wire.
  • The name is the machine's. It writes it as a real text run — ellipsised at the cell width, wrapped, inverted when selected — and each of those is a fact this side would have to invent. The label is not excluded at all; it yields per piece through Coverage.textCovers, the same test a semantic label uses.

The exclusion is narrower than it looks, and that is by construction. DisplayReplay.semanticOwns silences only a bits op, and only when the frame CONTAINS its whole destination. Text, lines and shapes pass through any exclusion — so the roster cannot take the machine's words even by accident. That was checked before the rule was chosen, not after.

If a Finder icon blit is ever seen to JOIN, this is the line to revisit: real ink would outrank the pack, and the exclusion would have to become a per-piece yield. It is named in the source at the point of decision.

Third fix, same rule: the icon is drawn in the box the Finder drew, not a constant 32. A list row is 16x16 at the Finder's 19-point pitch, so a 32-point icon ran through the row below it and over the columns the machine wrote — Michelle's "prints icons in a list on top of the list". FinderItems.clickPoint stopped trusting that constant for the same reason; HitTester.targetSize owns the rule and the drawing now reads it, because a click computed from one number and a drawing from another is how a click lands where nothing was drawn.

The correction: the 7.5 s was never this

The ledger entry above lists the cost as the fourth symptom of this mechanism. It is not, and the log already said so. decode_ms is a bracket from delivery to publish, and MirrorCycleClocks splits it:

decode_ms=21233  dc_own_ms=3  dc_content_ms=21230

Three milliseconds of that twenty-one seconds is this host's own decode, reduce and project. The rest is the P3 content join — guest round-trips. Rendering is not in the bracket at all; it happens after publish.

Measured anyway, on one fixture used for both readings — a 24-item icon-view folder window carrying the machine's blit and label run per cell — 40 timed renders after 5 warm-up renders, three interleaved pairs:

before after
median 7.83 / 7.83 / 8.38 ms 7.08 / 7.35 / 7.45 ms
fastest sample 6.98 / 6.91 / 7.43 ms 6.15 / 6.39 / 6.60 ms

Every pair moves the same way — about 8-11%, roughly the 24 hatches and 24 label patches no longer painted. Real, and three orders of magnitude short of her number. So this arc's render slowness is the content join, and looking for it in the renderer is looking in the wrong half — which is exactly what the dc_own_ms / dc_content_ms split was added to prevent and what the field's NAME caused on 2026-08-06.

One guard was blind, and it is recorded rather than quietly fixed

Fifteen mutations were run against the guards. Fourteen were caught. The fifteenth — widening the label yield from textCovers to mostlyCoverspassed a test whose own comment claimed to catch it, because both predicates yield to a run that fills the patch. The case that separates them is a machine that painted the label band and wrote nothing in it: one takes the name away, the other keeps it. That is now its own test (testInkThatIsNotWordsDoesNotTakeTheName), and it is the third render guard in this tree to pass the exact mutation it was written for.

Verification level

Tested, not metal-verified and not watched on the emulator. Every claim above rests on offscreen renders and unit assertions here; nobody has driven a Finder window with this build. The three fidelity symptoms should be gone and the correction to the cost is measured, but a person looking at a real folder window is what would close it.

LOOK: round 9 landed four lanes, and the two things that would have gone wrong were both caught by reading (2026-08-07, claude/019-integration-9)

Verification level: TESTED, and the gates were NOT armed. scripts/test-all exits 0 only under TBT_ALLOW_UNARMED_HOOKS=1 in this worktree — core.hooksPath names an absolute path ending in now/.githooks, which does not exist, and 63 worktrees carry their own shadowing value. So GATE_EXIT=0 means the tests pass; nothing refuses a bad commit. Every commit in this round was made with git -c core.hooksPath=.githooks, which gates the committer's own commits and mutates nothing shared. Nothing here is metal-verified.

Four lanes, 21 commits: 019-asset-packs (docs), 019-pressed-and-waiting (8, skipped by rounds 7 and 8), 019-kw01-kw06 (6, carrying 019-known-wrong-register as an ancestor rather than as a second merge). Five conflicts, all resolved by reading.

The two that a mechanical resolution would have got wrong.

  1. RenderShot in SceneView.swift. This tree had round 8's split into cgImage plus a REFUSAL when neither a size nor a guest screen is known; the lane still had the older single png that defaulted the size. Both sides were adding the same thing — a pressed: parameter — to different shapes of the same function. Taking the lane's whole hunk parses, compiles, and silently reinstates an invented render surface where the tree now refuses. The resolution was to keep this side's structure and thread the lane's parameter through it.

  2. WindowChrome.growBox. Both sides carried a growBox; this side's still opened guard win.kind != 2 with a long note arguing the discriminator was wrong and left deliberately, and the lane REMOVES it. Here keeping "ours" — the reflex on a foundation file — restores the exact defect the lane was written to delete, measured at 0 of 225 pixels agreeing on Appearance and wrong in BOTH directions at once.

The shared shape is worth naming: in both cases the two sides were not disagreeing, they were at different points along the same argument. A merge tool sees two texts; only a reader sees which one is later.

The register renumbered, and an id stopped meaning what it said. 019-kw01-kw06 closed the grow box (then KW-01) and the ctlact false negative (then KW-06, which the first closure had already shifted to KW-05). Ids are gated for contiguity, so closing a row moves every row below it. Today KW-01 is the zoom box and KW-06 is menu geometry — different defects wearing the ids the branch name still carries. KnownWrongRegisterTests was updated by the lane and passes; its testKW01…/testKW06… names now assert the new occupants. Cite rows by title.

What was checked because a previous round paid for it.

  • Round 8's duplicate-declaration class (two branches each adding a byte-identical hasTitleBar at different offsets, merged cleanly, caught only by the compiler): every changed .swift/.c/.h scanned for repeated top-level declarations. None. This is still a scan, not a gate — nothing runs it automatically.
  • tools/merge-census-gate pending before each merge: 0 dropped, 0 files gone, 0 imported self-reverts, both times. tools/self-revert-gate scan: 0 of 5, 0 of 8, 0 of 6. Both were run by hand. Both still ship unarmed, so the next integrator who does not think of them gets nothing.
  • scripts/test-native's manifest audited in Python against the tree rather than by grep: 118 test files, 0 unlisted, 0 listed twice.
  • docs/open-issues.md union-resolved twice. The second was not a concatenation: the lane's block opened with a paragraph belonging to the file's INTRO and only then two entries, so the paragraph went to the intro and the entries to the top of the list.

A derived document read clean over real change, for the second round running. tools/derived-doc-gate rederive: unchanged, both files. This round altered now-guest-ppc/src/act/act_client.c and act_cmds.cctlact stopped concluding a refusal from a silent trap patch, which is a change to what a verb ANSWERS — and act_cmds.c is not among contract-coverage.md's declared sources, so not even sources-sha1 moved. Round 8 saw the narrower version (sources-sha1 moved while every derive answer was identical). These numbers count which verbs exist and are structurally blind to what a verb replies. That is not a defect in the gate; it is the boundary of what it claims, and it should be written where the table is read.

Numbers. test-all GATE_EXIT=0. Guest cross-builds: PPC, 68K, ext and three rig instruments all ok. test-native 149 passed / 0 failed. MirrorKit 280 tests, 0 failures. Host gate run twice by its own single command: 1899 tests, 0 failures, 54 then 72 skipped — the second run is NOW_MIRROR_ASSETS=none, so the skip pair is the documented honest-degradation pair and not variance. xcodebuild Debug and Release both pass.

What this round did NOT do. No bake (Michelle's call; two resident changes are owed a --shared bake and the installed oracle matches no receipt). Nothing landed on main. No drive, no emulator, no metal — the lanes' own "EMULATOR-VERIFIED" claims are inherited from their authors and were not re-witnessed here.

OPEN: a symbol census cannot see a duplicated definition, and round 8 landed one (2026-08-07, claude/019-integration-8)

Verification level: TESTED. scripts/test-all GATE_EXIT=0 on the twelve-lane merge; host gate run twice (1892 tests, 0 failures, 54 then 72 skipped — the documented single-command pair). Nothing metal.

tools/merge-census-gate landed in this same round and was run before every one of the twelve merges. It reported 0 dropped, 0 imported self-reverts, 0 files gone every time, and it was right each time. Then mirror/host/MirrorKit/Sources/MirrorKit/WindowChrome.swift failed to compile on invalid redeclaration of 'hasTitleBar'.

Two branches had each added a byte-identical hasTitleBar to the same file at different offsets. Git merged them with no conflict, and the census could not object: it asks whether a name is present, and the name was present twice. A census counts names; it does not count how many of each.

This is the same limit already written into the tool's own docstring for the inside-a-function-body case, arriving from the other direction, and it is worth stating as its own line because the two cases feel different and are not:

  • a definition the merge LOST — the census sees it (DROPPED);
  • a definition the merge DOUBLED — the census cannot;
  • a definition the merge kept while changing what it does — the census cannot.

Only the first is a census question. The compiler caught this one, and nothing would have caught it in a language without a redeclaration error — a duplicated Python def or a duplicated Markdown heading parses fine, and this repository has already paid for both (six duplicate defs; a docs entry gutted by keeping a newer heading). A +1/-0 count per symbol would close it and is a small change to a tool that already builds the per-file symbol table it would need.

Three other things this merge found by reading that no gate reported:

  • A constant restated in prose auto-merged clean while being wrong. WindowChrome.contentOrigin's doc said the titled render is off by (1, 2) because Platinum.contentTop = 22. A sibling branch had already taken it to 20 and moved the frame outside the scene rect. Neither side conflicted, because they are different files. Corrected in place, dated.
  • Two branches independently wrote down the same finding (the 3,789 tracked worktree files), one proposing the fix and one performing it. Both entries are kept and cross-referenced rather than collapsed, because the duplication is the evidence.
  • The two halves of one function landed on different branches. RenderShot gained a cgImage entry point on one lane and a screen-unknown REFUSAL on another, both editing the same six lines. Keep-both would have produced two function headers; the resolution was to move the refusal down into cgImage, which png now delegates to, so both callers pass one guard.

PARTLY GATED: the silent self-revert now has two gates, and both ship unarmed (2026-08-07, claude/019-merge-gates)

Verification level: TESTED, by replaying the real incident and by mutation. Neither gate has refused a real commit or a real merge, because neither is armed.

The entry below asked for two things nothing detected. Both now exist.

tools/self-revert-gate — a commit that undoes its own branch's predecessor. For each commit it walks its own first-parent line (stopping at a merge: past one, the ancestors are somebody else's) and asks how many of the lines a near ancestor ADDED this commit REMOVES, byte for byte, in the same file. It flags on four conditions together: 12+ lines, 55% of everything that ancestor added, within 20 commits and 24 hours, and no mention of the ancestor in the message. Replayed against fe4d8179 it names all four predecessors it undid, not the one an integrator found — 129 of 129 lines from 9c219366 seven minutes earlier, 128 of 130 from 361c50ec, 15 of 15 from 4fb1262d, 31 of 31 from 8d92dba2.

The last condition is what makes it fair to refuse on, by the ext-bake-gate's own test: satisfying it is fully in the committer's power. A commit that means to undo its predecessor says so in one line. It is not asking for different work, it is asking for a message that describes the diff it carries.

Measured flag rate, because a threshold without one is a hope: tools/self-revert-gate sweep claude/019- main flags 9 of 1050 distinct commits across 61 branches (0.86%, 2026-08-07). It is a subcommand rather than a number written down, because the rate moves as branches land and a derived figure quoted from memory is the defect this repository paid for three times in one day.

fe4d8179 is one of the nine. Another is 1de856e3, which removes all 48 substantial lines 39191398 added fourteen minutes earlier — a whole scene_self.c/scene_self.h feature and its source test — in a commit titled "drive system actions through guest state". Nobody had noticed that one, and it is not what this lane was sent to find. Whether the other seven are reworks or reverts is a reading; the point of the gate is that it hands a reader nine commits instead of a thousand.

tools/merge-census-gate — the merge half. It derives every top-level definition in a tree (C functions, prototypes, macros and tags; Swift declarations; Python defs and classes; every published document's headings — 29,917 names over 1,407 files, in 0.7 seconds) and compares both parents and the merge base against the result. Three verdicts:

  • DROPPED (refuses) — a definition neither side removed relative to the base, absent from the merge. Nobody deleted it and it is gone, so the resolution lost it. That is both keep-both truncations in this ledger: the six duplicate Python defs that still parsed, and the docs entry gutted by keeping a newer heading and dropping an older body.
  • IMPORTED SELF-REVERT (refuses) — a commit in HEAD..MERGE_HEAD the sibling gate flags. A merge whose own body NAMES the commit passes, which is why round 7's real merge of claude/019-asset-packs is green: it named fe4d8179 and resolved all seven files by hand. History already on the first parent is never re-refused, or an arc goes red forever for a commit nobody on the merge can change.
  • REMOVED BY A SIDE (reports, never refuses) — a deletion a parent actually made. A deliberate deletion and a revert are the same bytes here, so refusing would refuse every honest deletion in the fleet. Each name is instead attributed to the commit that removed it and grouped under its subject: 87 bare names is a list nobody reads, and 87 names under 26 subjects is 26 judgements.

The census alone would have caught one name out of that seven-file revertScene.swift's cdef — and that is stated here rather than discovered later. Everything else fe4d8179 reverted lives INSIDE function bodies: three switch arms in now_cdef_role, a parameter on put_coverage_claim, a snprintf of "cdef", 117 lines of this document under headings that survived. A census counts names, and the names did not move. So the imported-commit half is not a convenience — it is the half that works on this defect, and the census is the half that works on the keep-both truncations.

Both ship UNARMED, behind NOW_SELF_REVERT_GATE (.githooks/commit-msg, a new hook — the only one that has the staged content and the message, and the message is half this verdict) and NOW_MERGE_CENSUS_GATE (.githooks/pre-merge-commit and .githooks/pre-commit, because a conflicted merge is committed by git commit and only the second fires). That is the derived-doc-gate's precedent and its reasoning: arming a gate into a live fleet is a change to every branch at once, made by somebody who is not on any of them. Arming is one deletion per hook, named in the comment beside it.

What neither gate covers, and the list is longer than the covered one:

  • A revert that RETYPES rather than restores. Both comparisons are byte exact on whole lines.
  • A revert of a SIBLING lane's work. self-revert-gate searches a commit's own ancestry on purpose; widening it to the fleet turns a named accusation into a similarity score.
  • A partial revert. Three of twelve lines is under both thresholds and stays there.
  • Anything a merge changes inside a name it keeps — signatures, bodies, a heading whose body was gutted.
  • Neither is armed, so neither has refused anything. They are tools with tests until somebody arms them.

scripts/test-native runs both suites (10 and 11 cases, ~25 seconds between them). Eight mutations were watched to fail them, including one that stops censusing Swift properties — which is what makes the replay of the real auto-merge load-bearing rather than decorative.

BROKEN: ctlact part 11 reports act-not-taken over a press that landed (2026-08-07, fidelity sweep D)

CLOSED 2026-08-07 by claude/019-kw01-kw06, landed in round 9. The entry that records the fix is "FIXED and EMULATOR-VERIFIED: ctlact part 11 reported a refusal over presses that landed", above. This text is left standing because the page is not corrected by editing it.

The full argument, with the two screendumps, is in docs/fidelity-sweep-2026-08-07-d.md. The short form:

ctlact with part: 11 at Appearance's help "?" button answered

{"code": "act-not-taken",
 "message": "armed, and the application never called TrackControl",
 "settlement": "timed-out"}

and the machine opened Mac Help and ran a search for "Appearance". The screendump taken before the press has no Help Viewer anywhere on the guest; the one taken after has it frontmost with ten results. The only other act in that stretch was refused in 5 ms.

The mechanism is in the message. act-not-taken is inferred from the absence of a TrackControl call and phrased as a statement about the world. A Carbon control actuated by any other route lands while the plane reports it did not — and nothing in the reply distinguishes that case from the honest refusal the same message produces elsewhere.

ctlact part 0 already answers dispatched-but-unconfirmed, which is the right shape for exactly this and landed in round 6/7. part 11 should adopt it, or name TrackControl as the limit of what it observed.

Why this is BROKEN and not UNVERIFIED: an agent that believes the refusal presses again, and the second press lands too. Sweep C scored the identical message as a clean refusal and listed the refusal vocabulary among the things to leave alone; that reading cannot stand, and this entry is the dated line saying so rather than an edit to sweep C.

Emulator, guest d9a78b62a414, build asserted. Not metal-verified.

OPEN: the render is order-dependent, and the fix is on an unmerged branch (2026-08-07, fidelity sweep D)

49c9a6f6"fix(mirror): one scene, one picture — and the buffer that was never ours" — is not an ancestor of claude/019-integration-7. It sits alone on claude/019-first-render-differs, and neither round 6 nor round 7 merged it. RenderShot.png still asks ImageRenderer for its own cgImage.

Sweep D priced the absence on this project's real windows rather than on the two-button fixture the commit used. Render the same nine captures twice in list order: byte-identical, 9 of 9. Render them once with the list reversed: four of nine change, by up to 338 pixels and up to 157/255 in a channel. And two independent captures of an Appearance window the guest drew identically composed to renders differing by exactly one pixel — the error class the sweep spec says no similarity score can see.

Two consequences while it is unmerged: a per-target render finding carries an ordering artefact nobody is subtracting, and any pixel-exact render test in this tree is passing because of where its file happens to sit.

OPEN FOR APPLICATIONS, FIXED FOR FINDER: grow-box identity (2026-08-08)

CLOSED 2026-08-07 by claude/019-kw01-kw06, landed in round 9. growBox now answers nil everywhere; the entry that records it, and the cost it names (the windows that really do have one lose it too), is "FIXED and EMULATOR-VERIFIED: the render invented a grow box on windows the machine leaves plain", above. Note this entry's claim that the defect is one-directional is itself superseded: the closing measurement found it wrong in both directions.

The semantic Finder slice now establishes one narrower answer without reviving that guess: a visible, front Finder folder window is a standard resizable Finder container. It draws and hit-tests its 15-pixel size box, and the existing resize action still settles through the guest-owned window. Background Finder windows and all application windows remain conservative. Application grow boxes still require target-context FindWindow/inGrow evidence.

hasTitleBar was fixed from win.kind != 2 because every Dialog-Manager-owned window was getting no widgets. At the time of this finding, WindowChrome.growBox (WindowChrome.swift:96) still opened guard win.kind != 2, one function below it in the same file. The 2026-08-07 close withdrew that inference; the Finder-specific rule above does not restore it.

Confirmed in pixels, not inferred: Appearance is kind == 2000, so it passes the guard, the machine draws no grow box on it and the render draws one. That is a fabricated affordance — the same class the zoom-box change was landed to remove — surviving in the neighbouring function.

OPEN: a reference does not survive its own settlement wait (2026-08-07, fidelity sweep D)

A tab control's reference, minted and used successfully, was element-not-found: no observation minted this reference, or it has expired two steps and ~45 s later. The reference table holds 96 entries and every scene.request mints a fresh set, so a caller polling for settlement evicts the reference it is waiting to act on again.

The refusal itself is honest and fast (5 ms). The problem is that the documented way to confirm an act destroys the way to repeat it, so a five-step sequence cannot currently hold a reference across its own waiting.

FIXED: a human's stack was handed over headless, and the brief was the only thing that could have caught it (2026-08-07, claude/019-human-stack)

Verification level: TESTED. scripts/test-all green on claude/019-integration-7 plus these edits. The replacement stack itself is emulator-verified in the weaker sense that it was looked at: see below.

Round 6 handed Michelle a stack whose VM ran -display none. When a modal alert came up in Mail — "Is your computer set up for Internet access?" — she had no window to dismiss it in, and the act plane was refusing those buttons. Her stack was, from her side, stuck.

Nothing was misconfigured and the brief was followed exactly. Headless is the correct default for a lane; no step said "and give her a screen". That is the same shape as the port reservation one layer up — allocation reserved a NAME while the collision happened on a SOCKET — and it has the same cure: move the decision out of the prose and into the machine.

scripts/spin-up-ppc now reads the answer off the ports, which already know whose stack it is. A boot on the reserved human range (tools/lane-ports human --is-human) gets a window by default and REFUSES NOW_SPIN_DISPLAY=0 with exit 64 before it clones the image. NOW_HUMAN_STACK=0 is the only way past, and it is a statement rather than a knob. docs/handing-over-a-human-stack.md is the checklist.

Two things found while repairing it, neither chased

  • A modal in the front application costs you the clean-shutdown route too. tools/shutdown-guest.py's primary route is the Finder's own Special ▸ Shut Down, and it needs the Finder to own the menu bar. Mail did, so it declined by name (the menu bar belongs to 'Mail', not the Finder) and fell back to the applet — which shuts the machine down but is known to leave the volume marked mounted. Acceptable for a throwaway session clone and never for the shared stage image. So a stuck modal on a human stack is not only a stuck stack; it is also a dirty image waiting to happen if anyone preserves that clone.
  • The --wire route needs the host app quit FIRST. It listens on the wire port and waits for the guest to redial, so a running host app holding it means the Finder route cannot even be attempted. The obvious ordering — guest down, then host — is the wrong one.

Still unverified

Whether the act plane can drive that alert's buttons at all. It was reported refusing them, this arc did not reproduce it deliberately, and a fresh boot cleared the modal rather than answering the question. The guest is a Carbon-era alert in a foreign application, which is the dialogItem path — a second data point for whatever picks that up.

SUPERSEDED BY SEMANTIC FINDER: the Finder item roster reverted to PRE-SWITCH geometry for ~0.8 s (2026-08-07, claude/019-flicker-bc)

Historical verification level: TESTED (emulator). The competing P3/roster composition described below no longer exists for Finder: Finder is denied P3, its host interior is semantic, and a structural change in list-header geometry invalidates the per-container roster. A post-change metal view-switch run is still owed, so this is superseded as a mechanism rather than closed by metal evidence. Found by tools/fidelity-live.py's second-ever run — the measurement nobody had managed to take twice.

Driving Finder ▸ View ▸ as List over the live agent socket and tracing consecutive scene documents, the Macintosh HD window's item roster switches from the icon grid (13 items, 13 owner rects) to list geometry (23 items, 21 owner rects) at +1.81 s — and then, at +7.48 s, with itemTotal and the owner-rect COUNT both unchanged, the identities of the rects revert to the icon grid for 0.83 s before returning to list geometry.

 -8.06s   displayTotal=235   itemTotal=13   roster=ICON boxes (32x44)
 +1.81s   displayTotal=235   itemTotal=23   roster=ICON boxes    <- count switched, geometry did not
 +6.86s   displayTotal=235   itemTotal=23   roster=list rows     <- geometry switched
 +7.48s   displayTotal=235   itemTotal=23   roster=ICON boxes    <- the BOUNCE, 0.83 s, nothing asked for it
 +8.31s   displayTotal=235   itemTotal=23   roster=list rows
 +8.41s   displayTotal=845   itemTotal=23   roster=list rows     <- the CONTENT PLANE finally follows

All 20 flipping rects are icon-grid boxes; they contribute all 40 rectOwnerFlips in the run, every one of them semantic → absent → semantic or its mirror.

Two things are visible there and only one is a defect. The lag is not, or not obviously: displayTotal holds at the icon view's 235 ops until +8.41 s because the Finder had not yet repainted, and the content plane can only carry what the guest drew. What it means is that for 8.4 seconds the window's drawn content was the icon view's while the item roster already claimed 23 list items — the semantic layer and the QuickDraw layer disagreeing about which view the window is in, with the renderer compositing both.

The bounce is. At +7.48 s the roster went list → icon → list with itemTotal and the rect count both unchanged at 23 and 21, and nothing asked it to. The count holding still is what makes it a SWAP rather than an addition: the list rows were gone from the same frames the icon boxes were in. In the render that is roughly a second of Finder items drawn at the wrong positions — Michelle's complaint #1 ("content draws over, under, or absent across redraws") caught in a trace for the first time, with the window, the rectangles and the instant all named.

Not yet diagnosed, and the hypothesis is stated as a hypothesis. Scene.Window.items is the Finder's own live position of via FinderItems; a cached roster re-attached one cycle late would produce exactly this shape. Nobody has looked yet.

Its A-side counterpart scored zero, so this is either a regression introduced since, or a bounce the A-side trace happened to miss — its view-list-a run missed 6 snapshots and every count in this family is a floor. Deciding between those two is the next step and it is cheap: fidelity-live.py --gesture menuItem --menu 259 --item 3 on a tree before the ladder landed.

ANSWERED: the ink → unknown → ink signature is still ZERO, live, on both sides (2026-08-07, claude/019-flicker-bc)

Verification level: TESTED (emulator).

tools/fidelity-live.py was built to see flicker a settled capture structurally cannot, and it had never produced a second measurement: sweep B did not run it and said so; sweep C could not, because mirror_read --intention snapshot closed the connection without replying. That transport defect is fixed, and the B side is taken.

The signature the provenance-ladder work was aimed at — a rectangle losing its ink and getting it back — does not occur in any of the five traces, on a tree that has since gained the ladder, displayEpoch coherent pairs, the content plane's renewal carry-forward, contentPlane/controlsState not-attempted-vs-empty, Platinum.contentTop 22→20 and a render path that no longer depends on ImageRenderer.cgImage's backing store. Every rectOwnerFlip measured is the semantic ↔ absent roster swap in the entry above, which is a different defect in a different half of the system.

Run Frames Missed Flicker coverage rectOwner hatch / dropout / presence Drain in artifact
idle-b (nothing provoked, 90 s) 951 5 51 51 0 0 / 0 / 0 yes
finder-open-b 514 4 25 25 0 0 / 0 / 0 yes
view-list-b 515 4 67 27 40 0 / 0 / 0 yes
reselect-b (×4; every select REFUSED "already front", so this is a second idle trace) 854 17 45 45 0 0 / 0 / 0 yes
idle-renewal-b (700 s of provoking nothing, across the interval the content plane's 9-minute renewal must fall in) 6,857 17 383 383 0 0 / 0 / 0 yes, all 6,857 frames

snapshotsMissed is quoted beside every count deliberately: a flicker count is a floor, never a total.

The idle watch across the renewal boundary is the one that had never been completed. The decay lane fixed the renewal's picture-from-nothing at unit level and said so. Live, across eleven idle minutes: no hatch, no content dropout, and Macintosh HD's displayTotal went 845 → 1,140 at +200.9 s and held — the op count only ever grew, which is the carry-forward's prediction. The limit is stated rather than glossed: the plane's armedAt is host-internal and reaches no projection, so that epoch bump is a candidate renewal and not an identified one. The claim earned is the weaker one, and it is still the one nobody had made.

Three A-side findings reproduce unchanged:

  • The oscillation is one claim, still. process-visibility flips stale ↔ partial with a median return of 3.396 s (A side: 3.28 s) — the same ~0.3 Hz square wave, running with nothing asked of the machine, and it accounts for every coverage event in every run.
  • The render never settles, at a five-second threshold, even idle.
  • baseComplete is false in every frame of every trace.

And the instrument gained the assertion AGENTS.md requires of it. fidelity-live.py reads the live host, which arms P3 itself — so a run against a host that never armed reports every window stably empty and reads as a stability result. It now derives planeEvidence from the one field that already carries the distinction (displayTotal: None = never traced, 0 = traced and proven empty, >0 = a drain reached the artifact) and REFUSES (exit 3) without a drain unless --allow-no-drain says absence is what was meant. Every run above carries a drain in every frame (2 windows traced, both with ops, max 845). It also asserts which guest answered (--expect-build auto against session_health), because there is one agent socket per user and this tool binds nothing.

WAS BROKEN, NOW FIXED: the blit join keyed on the wrong generation, and a guard that could not see what it claimed to (2026-08-07, claude/019-window-flags-and-join)

Verification level: TESTED. Four mutations watched; nothing here has been near a guest.

NOWMirrorContentPlane looked up a blit's held source with SourceKey(port: source, generation: record.generation)the BITS record's generation, not the generation the held ops were recorded under. A world's ops arrive before the blit that places them, which is the entire reason they are held, so any re-arm in between missed the lookup and the composite never joined. It was left unfixed once because its failure is at least honest: the bits op falls through and the renderer hatches it. That is a reason to defer, not to keep — and its rate is coupled to arm frequency, because join renews every nine minutes against a ten-minute TTL, so it fired on machines nobody touched.

Fixed by tracking the generation each port's ops were recorded under (sourceGeneration, learned in the one place a hold is created) and quoting that at the lookup. Both sites, and the second is worth its own line: the nested world-into-world branch had the identical defect and no test at all. Reverting it alone left every test in this file and in NOWMirrorContentCoverageTests green. It was found by mutation, not by reading, and it now has a test that fails against exactly that revert.

What was GIVEN UP, said plainly, because it looks like a deletion

testSourceAddressReuseAcrossGenerationsDoesNotJoin asserted that a bare generation change voids the join, on the ground that a disposed world's address is reused by the next NewGWorld of the same size (measured 2026-08-06: 0x1ea59e00 twice running). The hazard is real and that guard could not see it. From the wire, "a re-arm happened between this world's ops and its blit" and "this world was disposed and its address reused" are indistinguishable if generation is all you consult — and the first is the routine one. So it was voiding every join across a renewal to catch a case it could not identify.

The guard now rests on the evidence that can tell them apart: the guest's own worlddied, which releases the hold outright, resolved through the same key the join uses so a death arriving after a renewal still finds the ops it is about. testADeadWorldDoesNotJoinEvenAtAReusedAddress replaces the old test and was watched failing against a mutation that stops the release — which also broke two real-capture tests, so the release is doing work on actual Date & Time and Appearance drains, not only in fixtures.

The residual, unfixed: a world disposed and its address reused with the worlddied record never reaching this host would now join wrongly. Before, every join across a renewal was missing. Nothing measures how often the first happens.

WAS BROKEN, NOW CLOSED: the zoom box — and the grow box is NOT the same fix (2026-08-07, claude/019-window-flags-and-join)

Verification level: TESTED, against the machine's own pixels. Nothing here is metal-verified and nothing here needed an emulator: the evidence is the 2026-08-07 screendump corpus, compared per rectangle.

Closed. "IR v1 cannot say whether a window has a ZOOM box" is no longer true. windows[].closeBox and windows[].zoomBox are the WindowRecord's own goAwayFlag and spareFlag, one byte each at offsets 112 and 113, beside the windowKind the walk already read — contract first (asyncapi.yaml's scene family and mirror/docs/IR-V2.md), then axwalk.c, then scene_json.c, then Scene.Window and IRSchema, then WindowChrome. NOW's own window answers the same two facts through GetWindowAttributes rather than a second derivation. contract-coverage.md carries the declared asymmetry: NOW-68K serves no scene, and here — unlike apps[].backgroundOnly — it has no second route to the fact either, because nothing in it reads a foreign WindowRecord.

The bar was per-rectangle agreement on a window that has one and one that does not, and that is what was measured. PlatinumTitleBarTests now prices the zoom box's rectangle on all five corpus windows — Extensions Manager and the Finder folder, where the machine draws one, and Appearance, Memory and Mouse, where it draws stripes and face. Exact, 0 pixels differing. Watched failing against both mutations that matter, which are the two states this field has been in: drawing it unconditionally (the original defect) fails Appearance, Memory and Mouse naming each, plus the stripe run beside them, plus two of the three geometry cases; returning nil always (yesterday's honest nothing) fails Extensions Manager and the Finder. The three-valued read is asserted separately — true draws, false does not, and absent does not, because a producer that cannot say has told us nothing about this machine.

NOT closed, and the reason is a correction rather than a to-do

The entry above said of the grow box: "Same fix, same field." That is wrong, and it is worth more written down than fixed quietly. spareFlag is the zoom box alone; MacWindows.h's WindowRecord carries no grow flag at all. The only other candidate is the variation code in the high byte of windowDefProc, and it is ambiguous without the WDEF's resource id: kWindowDocumentDefProcResID 64 numbers variant 7 kWindowFullZoomGrowDocumentProc, while kWindowDialogDefProcResID 65 numbers its own variants from 0 independently — and a foreign walk cannot ask the Resource Manager to name a Handle. That is the exact wall contrlDefProc already hit one level down, and it is why the control walk reports a heap origin rather than a kind.

So WindowChrome.growBox is left on kind != 2 — wrong in both directions, and now named in its own doc comment rather than changed to a different guess. The measurement that settles the ground truth was taken and is worth keeping: counted out of the corpus PPMs at each window's bottom-right corner, Extensions Manager (kind == 2) DRAWS a grow box — the same anti-diagonal ramp the Finder's corner shows — and Memory and Appearance (also kind == 2) draw flat face. So the guard is wrong on Extensions Manager, and nothing here proves a kind 8/20/2000 window is resizable either.

What would close it, in order of honesty: the Window Manager itself answers inGrow from FindWindow, because it consults the WDEF — an act, not a read, and it needs a gated surface. Short of that, nothing available to a foreign memory walk can say.

FIXED: HostAppStateWiringTests spent one deadline on two waits (2026-08-07, claude/019-housekeeping)

Verification level: TESTED, by mutation — the flake was reproduced deliberately before the fix was believed.

The test renders nothing. It binds a listener, connects a FakeGuest over a real socket and polls at 20 ms for two independent events: the listener becoming ready, then the guest's hello arriving. Both polled against one deadline computed once, so a slow bind spent the second wait's budget — and when it was fully spent the second loop's condition was checked zero times, failing on a connection that was merely late. 0.06 s in isolation, 8.4 s under a loaded Mac.

The fix is to compute the deadline inside wait, giving each its own. Watched against the failure it claims to fix, both halves:

  • pre-fix shape (shared deadline) plus a 9 s sleep standing in for a loaded Mac between the two waits: failed in 9.564 s, four assertions, the first reading listening(port: 34977) where connected(guestName: "PowerBook 1400") was expected — the second wait never polled;
  • per-wait deadline, same 9 s sleep left in place so the fix is the only difference: passed in 9.128 s.

Restored clean, it passes in 0.059 s.

OPEN: this tree cites corpus findings that are not in the corpus (2026-08-07, claude/019-housekeeping)

Verification level: TESTED. Four citations were corrected to say where the finding actually is; the class behind them is not closed.

docs/ said, in the completed past tense, that four findings had been graduated to the corpus. They are not on the parent's main. They sit on claude/018-findings, which holds 26 unlanded findings — so the repository was telling a reader they could go and read something that is not where it says. Landing that branch is a main decision and was not taken here; the citations were made honest instead, each naming the branch and saying it has not landed:

Finding Cited in
charcoal-ships-scalable-only-no-bitmap-strike docs/plans/2026-08-07-020-accounting.md
hdmx-is-apples-own-per-ppem-advance-table docs/plans/2026-08-07-020-accounting.md
headless-is-declared-not-observed docs/plans/2026-08-07-020-accounting.md
an-appearance-check-flags-correct-absences docs/fidelity-sweep-2026-08-07-b.md

The sweep that found the fourth also found a worse class, and it is left open. Comparing every finding-shaped slug cited anywhere in docs/, README.md, AGENTS.md and CONTRIBUTING.md against the parent corpus turned up six more that are on neither main nor claude/018-findingsderiving-a-drawn-procedure, pb1400-vram-read-bandwidth, qdtrace-drain-sweep-sound, postevent-modifiers-need-ppostevent, now-four-face-capability-cost, now-large-transfer-pacing-collapse. Two of those are honest on their face (now-large-transfer-pacing-collapse says "working slug"); the rest read as references to something that exists. They were not corrected because each needs its own check — a slug may have been renamed on the way in rather than never written — and guessing would replace one wrong claim with another.

Nothing gates any of this. A citation into another repository's corpus is prose, and prose that names an artifact is the shape this project has been bitten by repeatedly; the durable fix is a check that resolves cited slugs against the corpus, not a sweep repeated by hand.

FIXED and EMULATOR-VERIFIED: ctlact part 11 reported a refusal over presses that landed (2026-08-07, claude/019-kw01-kw06)

Verification level: emulator-verified, driven, with the old build driven on the same rig as the control. Closes the register row that was KW-06"ctlact part 11 reports act-not-taken over presses that landed" — which is deleted rather than left claiming a defect the plane no longer produces.

The verb inferred "the act was not taken" from one observation — the application never called TrackControl — and phrased it as a conclusion about the machine. A Carbon control actuated by any other route lands without that trap ever being called, and a Help button on CarbonLib is a prime candidate.

Driven on this lane's clone, the Appearance panel's help "?" button, identical target and part to sweep D's SEQ-A step 4, twice — once with the pre-fix guest restaged onto the same running clone as the control:

guest build reply the machine, in its own screendumps
d9a78b62a414 (pre-fix) ok: false, act-not-taken / "armed, and the application never called TrackControl" / settlement: timed-out Mac Help opened, frontmost, having searched "Appearance", ten results
99bfc1d30364 (this tree) ok: true, Dispatch: dispatched-but-unconfirmed, Settlement: dispatched-but-unconfirmed the same — Mac Help opened and searched

The screendumps before each press show no Help Viewer anywhere. So the press landed both times and only the verdict changed, which is the whole claim.

Why it is a correctness problem rather than a reporting one. An agent that believes the refusal presses again, and the second press lands too. A false negative in an act plane produces DUPLICATED ACTIONS; on a destructive verb that is worse than a crash.

Can a correct refusal be told from a false negative here? No — and that is why the weaker verdict is now given every time. Sweep C scored this identical message as "REFUSED CLEANLY" and listed the act plane's refusal vocabulary among the things to leave alone. It may well have been a correct refusal in that run; nothing in the reply distinguished the two then and nothing could now. Everything that can POSITIVELY establish that nothing happened already returns earlier with its own status — the plane refusing, the request never reaching the machine, the reference not resolving. Past those, the press was queued inside the target's own context and this application cannot see what the target did with it. So act-not-taken had nothing left to be reserved for on that path and was removed rather than narrowed.

settlement: timed-out was the other half, and it is latched: act_settlement.c refuses every later note once it is written, so a confirmed arriving afterwards would have been swallowed. now_act_await_fired therefore gained timeout_is_terminal — a caller that ends the act on the expiry writes the terminal word; a caller that goes on gathering evidence writes dispatched-but-unconfirmed. winact and menuact pass 1 and are unchanged; ctlact passes 0. The status returned is untouched, so the caller still decides what an absence means.

NOT FIXED, and the same shape — named rather than widened to unmeasured. winact's and menuact's arms also infer from a silent trap patch. winact at least reports its trap counters, which are evidence about the patch rather than about the machine. Neither has been driven against a press that landed, and this lane did not extend a correction into a case nobody has measured.

Guard: now-guest-ppc/tests/ctlact_verdict_source_test.py, in the scripts/test-native manifest — a source test because act_cmds.c includes Carbon and cannot be linked by a host cc. Watched to fail against three mutations: the act-not-taken reply restored, the wait made terminal again, and act_client.c taking the flag and writing timed-out regardless.

FIXED and EMULATOR-VERIFIED: the render invented a grow box on windows the machine leaves plain (2026-08-07, claude/019-kw01-kw06)

Verification level: emulator-verified, per rectangle, in both directions. Closes the register row that was KW-01"a grow box is drawn on windows the machine leaves plain, and withheld from windows it draws one on" — which is deleted rather than left claiming a defect the product no longer has.

WindowChrome.growBox opened guard win.kind != 2, one function below the zoomBox that was withdrawn for exactly this. Sweep D measured the discriminator wrong in both directions at once, and this lane re-derived both on its own clone (lane block 398, now-mirror-stage clone, this tree's guest and ext), per rectangle at each window's bottom-right 15x15:

window kind machine draws one? render BEFORE render AFTER
Appearance 2000 no a hatched plate — 0 of 225 px agree plain — 186/225
Extensions Manager 2 yes none (15/225) none (15/225)
New Old World (Workshop) 8 yes one (119/225) none (16/225)

Read the first row and the third together, because they are the whole trade. Appearance was being drawn an affordance the machine does not offer, and HitTester reported it as a drag target — so mirror.act.window op: resize dragged inside a live application's content region. That is gone. The Workshop row is the price: kind == 8 happened to be right, and answering "we cannot say" everywhere loses it. Extensions Manager is the proof the old rule was wrong in the second direction as well — kind == 2 was denied a grow box before and after, so nothing was lost there and the row's claim is confirmed rather than quoted.

Why the withdrawal rather than a better discriminator. There is no field to improve on. The WindowRecord carries goAwayFlag (close) and spareFlag (zoom) and no grow flag at all — resizability is a WDEF variant, and the variation code in windowDefProc's high byte is ambiguous without the WDEF's resource id, which a foreign walk holding a Handle cannot name. The route that closes it properly is FindWindow's inGrow, which is an act and needs an armed, gated surface rather than the passive scene walk every other chrome fact comes from. Still open; the register row for it is gone, so this entry is where it lives.

Serve answers element_not_found naming why, and the battery records a skip rather than grading a drag at a corner it cannot establish.

Guards: GrowBoxTests (three, posed as a property over window shapes rather than as two constants) plus the grow half of HitActionTests.testFrontWindowWidgetsAndGrowBox. Watched to fail against three mutations: kind != 2 restored, kind == 2 inverted, and no discriminator at all. The inverted one is why the property test exists — the fixture-based test passed vacuously against it, because its fixture window is not kind == 2.

OPEN: a lane can revert a sibling's work in a commit whose message never mentions it (2026-08-07, round 7 integration)

Verification level: TESTED. Found by reading a merge, not by any gate, and it would have landed silently.

(2026-08-07, later: both gates this entry asks for now exist and ship unarmed — see the entry above. The root cause, a shared worktree checked out under a running lane, is untouched: what changed is that its damage is visible.)

fe4d8179 on claude/019-asset-packs — subject "feat(contract): hello says which Macintosh, in fields rather than in prose" — reverts the CDEF-attribution work that its own immediate predecessor 9c219366 had just landed on the same branch. The revert spans seven files and 117 lines of this document, and the commit body describes only the hello.machine work. Restored: control_cdef.c's variant tables (which name every check box and radio button in OS 9's control panels a push button), control_cdef_test.c's rows — including the ones whose own comment warns that "a later reader who 'restores' the family table will reintroduce exactly the shipped bug" — scene_json.c's cdef emission (8 mentions to 4), Scene.swift's decode (5 to 0), observe.c's defProc/variant, and local-control-drive.py's --detail.

What makes it a class rather than an incident: git auto-merged almost all of it. Only control_cdef.h and local-control-drive.py conflicted. scene_json.c and Scene.swift took the revert with no conflict at all, while the sibling lane's scene_json_test.c assertions survived — an internally inconsistent tree that would have gone red at the gate, which is luck rather than a guard. Had the test file been reverted too, the merge would have been green and the product quietly worse.

The classification numbers moved DOWN on purpose (Memory 33→23 of 44, Date & Time 19→10 of 21) because the old route attributed variant 0 to a push button when the variation code cannot be read for a foreign control at all. A test asserting the higher count is asserting the old defect.

Two things are still open:

  • Nothing detects a commit that reverts its own branch's predecessor. A gate could: for each commit, whether any hunk restores a blob a recent ancestor changed, without the word "revert" in the subject.
  • Nothing detects an auto-merge that takes a deletion of live code. The keep-both hazards this repository has already paid for are all about conflicts; this one produced no conflict.

FIXED: arc-status's findings total was a sum with duplicates, and no landing could move it (2026-08-07, claude/019-arc-status-findings)

Verification level: TESTED (against a synthetic corpus, and derived against the real one).

019-multi-window-content fixed exactly this shape for COMMITS in round 7: shared_with was added because merge-base $b $INT counts a lane cut from another lane as owning that lane's commits, and one branch with one commit of its own read as seventeen. The findings counter below it was not fixed, and round 8 recorded that and moved on.

It cost a landing. A corpus lane swept all 166 local branches, found 74 unique finding files absent from the parent's main, landed 57 and left 15 (their doc_ref or evidence lives only on the source branch, so tools/data check refuses them) plus 2 stray PNGs. The parent's main went 216 → 273, and arc-status could see it — it printed the new total. Its next line still said 87 across 35, still named claude/018-findings as holding 26 that had just landed, and had not moved by one.

Two defects in the one loop, both from summing --diff-filter=A per branch:

  • Duplicates. A finding carried on eight branches was counted eight times, so the backlog grew every time somebody cut a worktree.
  • The number could not move. The three-dot diff measures against each branch's own merge base, so a finding COPIED onto the trunk — which is how the landing lane lands them — still read as an addition on every branch that ever held it. No amount of landing could reduce it.

The second is the expensive one, and it is wrong in the direction that costs work: it tells a reader there is a large backlog when the real remainder is 15, and acting on it means redoing a landing that already happened.

The fix is set arithmetic. Each branch's finding files as a set (git ls-tree -r --name-only <branch> -- data/findings, *.md, minus README.md), unioned, minus the trunk's set. Set arithmetic cannot double-count and cannot be fooled by a stale merge base. 87/35 became 15/11, reconciling with the landing lane's 15 excluded; its ~17 also counted the 2 PNGs, which are not findings. The line now names its own filter, so its 273 reconciles with tools/data check rather than with the raw .md count of 274.

The per-branch breakdown is now unique-at-risk — on that branch, on no other, and not on the trunk — because "what is lost if this branch is deleted" is the only question a per-branch breakdown is asked, and a per-branch total would have restated the double-count just removed.

Four cases added to tools/mirror-gate-tests/test_arc_status_measurements.py, watched to fail against five separate mutations (the per-branch sum; dropping the trunk subtraction; the literal original --diff-filter=A loop; counting README.md; printing branch totals in place of unique-at-risk). Each was named by the case written for it, and the diff-filter mutation is caught only by the landed-finding case — the two defects needed two tests, which is why one guard would not have been enough.

OPEN: a fix landed on a file this arc archived, and there is no forward target (2026-08-07, round 7 integration)

Verification level: TESTED.

claude/019-stale-reason-string fixed two refusal strings that outlive what they were about. The guest half (g_note in now-guest-ppc/src/files/files_browser_view.c) landed and is BUILDS — that pane is Carbon, has no host test, and nobody has watched it on a machine.

The host half did not land. contentNote lived in now-host/Sources/Host/MirrorModuleModel.swift, which 0443ab2b ("vendor Mirror whole as a subproject, and archive the port that replaced it") moved into archive/mirror-port-2026-08-01/ in this same arc — and git rename-detected the lane's edits onto the archived copy. An archive records what was archived, so it was restored rather than edited.

contentNote, joinContent and the page they belong to exist nowhere else in the tree (checked across now-host/ and mirror/), so there was no forward target either. The open question is whether the vendored Mirror has the same shape of defect — an explanation written on a failure path surviving into a success, specifically a live watch loop that ticks without clearing a note a refused press left standing. Nobody has looked.

FIXED: five complete nested agent worktrees were checked in under mirror/ (2026-08-07, claude/019-housekeeping)

Verification level: TESTED.

mirror/.claude/worktrees/ held five whole checkouts — confident-tu-60f97b, funny-khorana-e67807, gallant-ritchie-d78731, jolly-lamarr-d3cd03, wizardly-bardeen-e2b6dc3,789 tracked files, ignored by nothing, each carrying its own Package.swift and its own copy of MirrorKit.

They were stale in a way that matters for exactly the reason this repository states limits once: each carried Platinum.contentTop: CGFloat = 22 where the live tree says 20. A reader who grepped this tree for that constant got six answers, five of them wrong, and no gate read them at all.

Round 7's merge touched none of them (checked: zero paths under mirror/.claude/ in git diff claude/019-integration-6 HEAD). Deleting them was not an integrator's call mid-merge; it was taken separately, after checking the three things that make a deletion safe:

  • Nothing live reads them. The only references anywhere in the tree are the ones that describe the problem — this entry and asset-pack.md. No script, no Package.swift, no test resolves a path under mirror/.claude/.
  • They are not registered worktrees. git worktree list names none of them, and none carries a .git file. They are five plain directories that were once checkouts, so there is no branch whose working tree this disturbs and no unmerged index to lose.
  • They hold nothing that exists nowhere else. Comparing tracked paths against the live mirror/, each worktree had 202 paths the live tree does not: 200 of them the Platinum asset pack, which asset-pack.md removed from the index by decision and verified by sha256 against ~/Lab/Assets/now-mirror-assets/pack-2026-08-06, and two source files — ActionDispatcher.swift and QmpClient.swift — that had simply MOVED, into Sources/MirrorOracleKit/. QmpClient.swift is byte-identical to the live copy; ActionDispatcher.swift differs as an ancestor does, the live one carrying the plane model, the renamed deviceClick / thumbTracking cases and an explicit refusal arm the old one has no concept of. Nothing was unique, so nothing was stopped on.

The cure is a .gitignore entry for .claude/worktrees/ at any depth plus a git rm -r; history still holds the 25 MB, which is the same reversible move asset-pack.md took and the same unresolved question about a rewrite.

BROKEN: every rule in this repository is scoped by accident, and three of them are scoped wrong (2026-08-07, claude/019-scope-hooks)

Tested (mutation, six states, on the guardrail only; no metal). The inventory, the scope assignment for every hook/gate/tool/skill, and the ordered landing plan are docs/rule-scopes.md. Three findings that were not previously written down:

Fault C — the PreToolUse main guardrail did not exist here. AGENTS.md:316 has claimed since before .githooks landed that main is enforced by "a PreToolUse hook on Write/Edit/Bash". It is not. That hook lives in the parent TimBotTu checkout, now/ is excluded from that repository (.git/info/exclude:22), so a NOW session's $CLAUDE_PROJECT_DIR is this tree and the parent's .claude/ is never loaded. This is the same shape as the note already at the top of .githooks/pre-commit — the guard lived only in the parent, so a now worktree had documentation for enforcement it did not have — except that fix moved only the git-hook half and left the sentence claiming both. Fixed on this branch: .claude/hooks/guard-main.sh ported and wired, verified by mutation in all six states (main + in-repo Write → refused; main + git commit → refused; TBT_ALLOW_MAIN=1 prefix → allowed; main + ls → allowed; main + write outside the repo → allowed; on a branch → allowed). It refuses nothing any lane is doing.

The mis-scoping is the mechanism, not a tidiness complaint. .githooks never reaching main is one instance; the shared checkout being parked on claude/mirror-subproject — a subsystem's branch — is what made it a repository-wide outage. Meanwhile the two things that did reach 54 of 55 arc branches through .claude/settings.json are the Mirror drive loop and one research arc's Stop gate, and the coordination rules meant for the whole fleet (tools/arc-status, docs/arc-coordination.md) reach 25 and 27 of 55 and zero of main. The subsystem-specific rules propagated; the repository-wide ones did not.

25 MB of dead agent worktrees are tracked. mirror/.claude/worktrees/ holds 3789 files across five expired agent worktrees, landed in 0443ab2b ("vendor Mirror whole as a subproject"). Neither .gitignore excludes .claude/worktrees. Not removed here — deleting 3789 files during a seven-round integration would cost more in conflicts than it saves — but it should go with a .gitignore rule the moment the arc is quiet.

(2026-08-07, round 8 merge: this paragraph has stopped being true. claude/019-housekeeping removed all 3,789 from the index and added the .gitignore rule, on a separate branch rather than inside a merge — which is what "the moment the arc is quiet" asked for. The FIXED entry above carries the three checks that made the deletion safe. Both this branch and that one wrote the same finding down independently and neither could see the other; the duplication is left visible because it is the merge hazard this round was called to look for.)

Still open, and joint with Michelle: the landing plan's steps 1-6 in docs/rule-scopes.md. Steps 2 and 3 must be minutes apart: once the gate arc is on main, test-all refuses rather than warns in the ~48 override-carrying worktrees until hooks-doctor --fix runs, and TBT_ALLOW_UNARMED_HOOKS=1 is the deliberate override for that window.

FIXED: the renderer's first output differed from its later output, and it was never our cache (2026-08-07, claude/019-first-render-differs)

Verification level: TESTED. Host-side only; nothing here went near a guest or the PowerBook. MirrorKit 248 tests and the host's 1858 pass with the fix; the guard was watched failing 4/4 against the reintroduced defect.

claude/019-pressed-and-waiting found, while building render guards, that a pixel-exact comparison of two identical scenes failed when it held the process's first render and passed later in the same run, and wrote it down as "something on the shared draw path builds lazily". Both halves of that sentence turned out to be wrong, and the way they were wrong is the useful part.

It is not the first render, and it is not a cache

Rendering one fixture ten times in a fresh process: renders 1 and 2 are identical, the third moves two pixels, the fourth moves six more, and 4 through 40 never move again. Eight pixels in all — seven by 1/255 on antialiased control edges, one by 15/255 on the grow box's diagonal hatch. The settled image is byte-identical across processes; only the moment of settling varies, and it varies with wall-clock time as well as with the count (inserting a 0.4 s sleep between renders moved the whole transition from renders 3–4 to render 3 alone).

Three things it is not:

  • Not our lazy state. Resolving AssetPack.status, FontBook, DesktopPattern.answer and IconAtlas before the first render changes nothing at all. Those were the named suspects and they are cleared.
  • Not MirrorKit. A Canvas containing no MirrorKit code — a tiled 2×2 decorative image, a stroked rounded rect, a diagonal — drifts the same way. Pure CoreGraphics into a bitmap context of our own does not.
  • Not warmable. Because the transition moves with time, no fixed number of throwaway renders is a fix, and a warm-up would have hidden the symptom while leaving the fragility exactly where it was.

It was ImageRenderer.cgImage — its own backing store, which we were borrowing.

The fix, and what it costs

ImageRenderer.render(_:) hands the drawing to a CGContext of the caller's choosing. RenderShot now makes its own 8-bit sRGB bitmap context, and the output is stable from the very first rasterization: zero differing pixels over eight consecutive shots in a cold process, across runs, at the same cost as before (one rasterization per shot).

It changes the picture, and that is worth knowing before reading any render test's history: 1471 of the fixture's 57600 pixels differ from the old settled image — 1141 by one or two steps, the rest up to 213/255 on glyph interiors and 1px frames. Flat fills are untouched. Our own 1:1 bitmap draws hairlines crisply where ImageRenderer's drew them soft; the semantic checkbox's frame went from two heavy bars with washed-out sides to a clean square, which for a mirror of a 1px Platinum interface is a gain. Exactly one assertion in the tree was pinned to a glyph pixel that moved, and it was asking the wrong question anyway — it named one coordinate and expected it black, where its claim was "there is a mark in the box". The live window is a different path and is unaffected; this is the offscreen half only (tests, writeRenderShot, the serve endpoint).

Two traps in the GUARD, both of which passed the mutation first

AAARenderStabilityTests renders one scene ten times as the process's first renders. Its first draft was wrong twice, in ways worth stealing:

  1. It compared the first two renders. Renders 1 and 2 always agree — the drift starts at the third. A guard on the first pair passed the mutation it was written to catch, watched doing exactly that. It has to compare across the whole warm-up.
  2. It kept NSBitmapImageReps and read them at the end. NSBitmapImageRep(cgImage:) does not copy. Collect ten and sample them afterwards and they can all answer from one buffer: the same ten renders read 8 differing pixels when their bytes are taken as they are made, and 0 when only the reps are kept. This is the likeliest explanation for the reporting lane's own note that two of its render guards passed a mutation they claimed to catch, and it is a live hazard for any future guard — snapshot dataProvider.data at render time.

The limitation, which is why nobody caught this for twenty test files

Coldness belongs to the process, not to the test. AAA in the class name puts it first under XCTest's alphabetical order, so in a full-suite run the guard is real — but run after any other render it passes whether or not the defect is present, and there is no way to make a warm process cold again from inside itself. A mutation of the rasterization path must be watched under swift test --filter AAARenderStabilityTests. A green from a warm run proves nothing. Every one of the twenty render-test files in this tree looked strict and none of them could have seen this.

The two flakes it was asked about: one plausible, one refuted

  • The host gate's unnamed failure (one run of four reported Executed 1668 tests … 1 failure, name lost to a grep pipe) is plausibly this and not proven to be. There are about eight assertions in the tree of the form XCTAssertEqual(try RenderShot.png(scene: a), try RenderShot.png(scene: b)) — two consecutive shots compared for equality, in AlertItemTests and IslandRenderTests. Whichever of them ran first in a given process straddled the warm-up and could differ. The fix removes the mechanism. The name is gone, so this stays a hypothesis with a mechanism rather than a diagnosis.
  • The varying skip counts (54 / 65 / 72) are not this, and the run that answered it is scripts/test-all itself. scripts/test-host runs the host suite twice — once as the machine is, and once with NOW_MIRROR_ASSETS=none to check honest degradation. Measured in one green run today: 1858 tests, 54 skipped then 1858 tests, 72 skipped. Both numbers come out of the same invocation, minutes apart, and 54-versus-72 is precisely the pair that was being read as instability. The rest of the spread is the other environment gates — twelve XCTSkipUnless(NOW_METAL) and a dozen more waiting on a guest to dial in. Nothing there is nondeterminism; a report of a skip count that does not say WHICH of the two passes it came from is not a measurement.
  • HostAppStateWiringTests is refuted. It renders nothing. It binds a listener, connects a FakeGuest over a real socket, and polls at 20 ms against a single 8 s deadline computed once and shared by both waits — so 8.4 s under a loaded Mac against 0.06 s in isolation is that deadline being approached, not a render. It is a genuine load-sensitive test and a separate (small) piece of work: the second wait should get its own budget.

FIXED: the modal nobody could dismiss was fourteen pixels, not a classification (2026-08-07, claude/019-dialog-buttons-act)

Verification level: EMULATOR-VERIFIED, end to end and by mutation. Own VM, lane block 230 (anchor 13840 / wire 13841), session-private clone of now-mirror-stage.qcow2 carrying this checkout's build — the guest that dialled reported e715b0a6a5d7. Nothing here ran on the PowerBook.

Michelle, driving the round-6 stack: "the button labels are now correct, but the buttons still dont work and the modal is otherwise blank", on Mail's Internet-setup alert (Yes / No / Set Up Now), which she could not dismiss on a VM she had been handed headless. The status line read "click a control: the guest did not provide complete, authoritative semantics for that control".

That sentence pointed at the CDEF refusal, and the CDEF refusal is not the cause. Nothing about it was widened. The guest already authorises all three buttons through the DITL route — knowledge: known, kind: pushButton, action: press, a ref each, authorizesAction true — and the mitigation the CDEF lane predicted ("dialog windows lose nothing") holds exactly as written. ditemact on item 3 dismissed the alert on the first attempt.

The cause

The scene's window rect is the content port grown UP by titlebarHeight, whether or not anything is drawn in the band. HitTester read it that way. SceneRenderer did not: for a titleless dialog it treated the band as chrome it could shrink to its six-pixel border, so it drew the content 14 pixels high and 6 right.

Fourteen is more than half a push button. Aiming at the middle of a drawn button hit-tested above every dialog item, fell through to the userPane spanning the whole dialog — honestly unknown, no action — and was refused for having no semantics. The refusal was correct about the object it was handed and the object was wrong.

That is the third appearance of one shape: two sides each keeping their own copy of a geometry constant. WindowChrome already existed to stop it ("a click can never land on a box the renderer drew somewhere else") and did not own this number. It does now, and both sides read it.

The numbers, measured rather than assumed

Guest screendump and scene taken at the same instant:

reported content-local machine draws it at
NOW "Take Screenshot" (titled, kind 8) window (28, 50) (172, 423) (200, 493) = rect + (0, 20)
Mail alert buttons (titleless, kind 2) window t 96 t 85 y 201 = rect + (0, 20)

One convention, both window classes.

What was watched

A click at (326, 211) — the centre of where the renderer draws "Yes", read off the render by differencing it against the same scene with that item removed, not recomputed from the geometry under test — resolves through HitTesterObjectResolverInteractionPolicy to .dialogItem(ref: now-element-4dbd013f…, item: 3). Sent to the guest as that exact plan: dispatched-but-unconfirmed, and six seconds later the guest's own screendump shows the alert gone and Outlook Express's Inbox behind it.

Both guards were watched failing. Restoring the renderer's six-pixel content inset fails testEachDrawnButtonIsTheButtonAClickReaches naming the userPane by ref — Michelle's symptom, in the failure text. Restoring the old refusal sentence fails testARefusedControlSaysWhatWasUnavailable five times over.

And the refusal now names a fact

"The guest did not provide complete, authoritative semantics" is a verdict; a person can do nothing with it, which is the same failure as the dead-hooks warning four lanes read and correctly ignored. A refused control now says which fact was unavailable and what still works:

the guest could not determine what kind of control this is, only that its definition function is CDEF 23 — the button family, which is a push button, a check box or a radio button, and the variation code that would say which cannot be read from outside the owning process, so this driver will not decide for it. Pressing it by position is still possible and nothing here has done so.

MirrorObject.Control carries semanticKnowledge / semanticDefinition / semanticCdef for that sentence. The id is never mapped to a kind — naming it says the lookup worked and the id was not enough.

Still open, from the same capture

  • A modal can still strand a person, and this only widens the road out rather than closing the hole. A dialog whose items are all userItem or resCtrl, or whose window class the walk never enters, has no authorised route at all, and every other window is unreachable behind it. Nothing in the mirror treats "the front window is modal and I can serve none of it" as different from any other refusal. Mail's alert is now serveable; the class is not closed.

And the cost is larger than a dialog you cannot dismiss. tools/shutdown-guest.py's Finder route needs the Finder to own the menu bar; a modal holds it, so the tool declines by name and falls back to the applet, which is known to leave the HFS volume marked mounted (measured by another lane the same day, on this defect). So a stuck modal is also a dirty image waiting to happen for anyone who preserves that clone — a wedged machine and a poisoned oracle from one refusal. - The alert's message is ^0 and ^1, so the modal really is blank. Items 5 and 6 are staticText whose DITL text is the ParamText placeholders, and Mail substitutes at draw time via ParamText. The guest reads the item's own handle (dialog_text.h) and gets the template; the sentence a person must read to answer the question lives in the Dialog Manager's four param strings and is not on the wire at all. This is the second half of Michelle's report and it is a GUEST gap, not a renderer one. (Sweep A saw the same items as raw pointer bytes; that half is fixed, and this is what is underneath it.) - Outlook Express's Setup Assistant publishes no dialogItems at allwindowKind 1445, an application-defined kind, so now_scene_walk_window never enters the dialog branch and all 13 of its radio buttons and check boxes arrive as unknown cdef: 23 controls. It is a DialogRecord by every other sign. The kind gate is there for a real reason (a DialogRecord's TEHandle sits past the end of a plain WindowRecord), so widening it needs a different proof that the record is a dialog — not a wider kind test. - ~~A titled window's render is still off contentOrigin by (1, 2)~~ — (2026-08-07, round 8 merge: this stopped being true before it landed. claude/019-window-flags-and-join's predecessor moved the window frame OUTSIDE the scene rect, took Platinum.contentTop from 22 to 20, and the renderer now counts off PlatinumTitleBar.Row.contentTop = 20. Both classes agree with the machine; the (1, 2) it describes is gone with its cause, not left alone.*)

LOOK: a pressed Platinum button has never been photographed, and the guest cannot draw one for us (2026-08-07, claude/019-pressed-and-waiting)

Verification level: EMULATOR-VERIFIED. Own VM, lane block 432 (anchor 15456 / wire 15457), clone of now-mirror-stage.qcow2 (sha256 c466baa9…), this tree's build confirmed by matching sourceManifest effb72f9346b and buildFingerprint 6f7b3e5f508f, resident capabilities=511. Guest shut itself down cleanly; block reclaimed.

The slice was to draw a pressed style for buttons, and the standing rule said to compare against the guest's own pixels first, because Platinum's pressed state is a specific drawn procedure and not a darken filter. The comparison was made. It returned something more useful than a palette.

There is no such capture, and the usual routes cannot take one

Every screendump in the corpus is taken around an act — tools/local-control-drive.py:19 shoots "before and after" — and the one tool that holds the button down, tools/local-drag-vehicle.py, never screendumps at all. So the corpus holds a great many quiescent desktops and not one pressed control. docs/deriving-a-drawn-procedure.md:262 already says the same of tabs.

tools/local-pressed-capture.py is the missing combination: dragpress leaves the button down, and while it is down we shoot.

Holding the button down changed nothing, and the near-miss is the lesson

Two false answers came first, in opposite directions, and both looked exactly like a result:

  1. A run that pressed 54 px above the button. The scene's window rect is the STRUCTURE rect — its top is the top of the title bar — so content-relative coords need one title bar (20 px) added, not zero. The run sampled flat dialog grey, found no change, and reported that the machine draws no pressed state. The script now ABORTS when the target rectangle holds fewer than three colours: a push button is never flat, so a verdict from such a rect is about the rig.
  2. 65 px of change that were the mouse cursor. With the aim corrected, BEFORE vs DURING showed 65 of 2600 px moved, in convincing blacks and whites. All of it was the arrow arriving in the rectangle, which dragpress had moved there. DURING vs AFTER — cursor in the same place in both — is 0 of 2600.

The guest says why, in its own words

Two verbs, independently:

  • ctlact part 11act-not-taken: armed, and the application never called TrackControl
  • ctlact part 0"a real click at the point; this application does not route it through TrackControl at all, so no patch was consulted"

A Platinum pressed face is drawn by the application inside TrackControl. An application that never enters that loop never draws one. And Scene.Control carries no hilite field for it to be reported through — that gap is structural, not an oversight, because TrackControl is exactly the state in which the application has stopped calling GetNextEvent and the resident's scene walk cannot run.

Bounded honestly: the target was NOW's own Carbon window. Another application that does call TrackControl might well draw a pressed face. If someone photographs one, PressedVisual is where the measured drawing goes. Nobody has.

What that settled about the drawing

Imitating Platinum's pressed face would be the mirror asserting a machine state that does not exist — and unfalsifiable, since nothing can contradict it. So the press mark is deliberately the provisional family's, not Platinum's: ProvisionalVisual's edge and ink over UnknownVisual's stipple, saying what is actually true, which is that the host has your press and is waiting.

Also measured, and it sizes the deadline

ctlact part 0 took 5.1 s on this rig, independently reproducing Sweep C's ~5 s. PressSession.patience is 8 s so the ordinary slow case lands inside the wait rather than being reported as a mystery, and testPatienceOutlastsTheMeasuredWorstCasePress names the measurement for whoever tries to tighten it.

UNVERIFIED: a press that reaches the guest still cannot be confirmed

NOWMirrorSource answers a press's refusal and nothing else, because the other three dispositions all mean the act left and none is evidence it worked (.direct is dispatched with no typed postcondition and nothing will ever settle it). So every press that actually reaches the guest runs to the deadline and reports "asked X and never learned the answer after 8s — it may still have happened".

That is honest, and it is also the same missing plumbing ItemDragDriver is waiting on two entries below: the broker's settlement is not routed back to the gesture layer. Closing it closes both.

UNVERIFIED: the pressed style has never been seen by a person

Render-tested, gate-green, and never on a screen a human watched. Platinum fidelity is a human judgement and this drawing deliberately is not Platinum, which makes it more in need of one, not less.

Noticed in passing: the renderer's first output differs from its later output

A pixel-exact comparison of two identical scenes fails when it holds the process's FIRST render and passes when placed later in the same run — something on the shared draw path is built lazily on first use. Caught because PressedRenderTests compares whole images; it is stated in that test and warmed around, not fixed. It breaks rule 2 (the same window state renders the same way every time) and the fixture corpus is meant to harden into a contract gate at the IR v1 freeze, so it is worth someone's attention as its own question.

FIXED: the drive loop's own instrument armed every plane except the one that draws interiors (2026-08-07, claude/019-instrument-arms-content)

Verification level: EMULATOR-VERIFIED, by mutation. Own VM, lane block 520 (anchor 16160 / wire 16161), plain-base clone of os91-runner.qcow2 carrying this checkout's build — the guest that dialled reported 113f1b176035, which is this tree's NOW_SRC_HASH. Nothing here ran on the PowerBook. Evidence: ~/Lab/Assets/now-mirror-assets/019-instrument-arms-content/.

tools/local-pair-capture.py never issued qdtrace start, and its warm-up comment said the planes arm "as a RESULT" of a scene walk. That is true of three planes. P1, P2 and P4 echo the resident's arm request unconditionally (ext/src/now_ext.c:294,299,323); P3 appears nowhere in that function. Its bit is set in exactly one place (ext/src/now_content.c:1843), reachable only under a verdict over arm_window / arm_psn / arm_generation. Content is not a plane at all — it is a per-window, per-A5, TTL-bounded spotlight, and the only thing in the PPC app that claims it is qdtrace start (now-guest-ppc/src/content/qdtrace_cmd.c:311), released at :354.

The part that makes arming necessary but not sufficient

SceneBuilder.normalizeWindows sets display: nil unconditionally (mirror/host/MirrorKit/Sources/MirrorKit/SceneBuilder.swift:285). So no scene envelope from any capture has ever carried content ops, armed or not. The interior arrives on a second artifact — a qdtrace drain, composed by NOWMirrorContentPlane.apply(drain, to:) — which this tool did not write. Arming alone would have changed nothing on disk; the tool had to start emitting a drain as well.

Sweep B said this in its own words on 2026-08-07 (docs/fidelity-sweep-2026-08-07-b.md:414-417): "a rig that reports empty interiors is reporting its drain, and the drain is a different channel." It was written about other rigs and nothing checked this one against it.

The mutation, and what the control reproduces

One target — Date & Time — captured twice on the same machine minutes apart, once armed and once with the arm disabled by --no-content:

armed control (--no-content)
drain records 3917 0
op families state 1793, line 519, rect 510, text 313, bits 220, rrect 146, worldborn/worlddied 139 each, blitsrc 138 none
distinct strings 48, including 1:19 PM, 2026, Apple Americas/… 0
the plane's own sentence "784 new draw ops; 139 offscreen worlds hooked at creation; … joined 138 composites" "waiting for the guest to draw"

Rendered through LiveShapedRenderTests.testRenderASweepAsTheAppWouldDrawIt onto its own scene, the armed render draws the panel — both date and time values, the Time Zone and Network Time Server sentences, the group box titles. The control renders the same window with the field values empty, the static text absent and the group box titles unpainted.

That control picture is worth more than the fix. It is, close enough to identical, what round 3's LOOK recorded as a render defect (## BROKEN: the integrated render at round 3, above: "Date & Time — unchanged — Still no group boxes at all … Text fields are empty grey slabs"). The defect is reproducible on demand by disabling one arm, on a tree where the renderer draws that panel correctly.

Why it survived: the quiet hatch and the loud one

An unarmed capture and a genuinely empty window produced the same artifact, and there was no way to tell them apart afterwards — which is why nobody noticed for a day. The asymmetry claude/019-sweep-c found in the same pass is the other half of this: "Bitmap unavailable" can only come from a drain, so its presence is positive proof P3 was armed. The loud hatch carries its own provenance; the quiet one carries none.

So the fix is not only "arm it" — it is to make the arm legible in the artifact, the way the loud hatch already is:

  • contentArm in manifest.json per target: requested, ok, reason, window, psn, ttlTicks, records.
  • <slug>-drain.json in the fixture shape the content plane accepts, whose provenance.run reads NOT ARMED: <reason> when it was not — so the statement survives being copied into Fixtures/.
  • LIMITS.md via tools/sweeplimits, and the same limits stamped inside manifest.json.
  • stdout names every unarmed target at the end of the run.

--no-content is the control, implemented as a subclass that refuses rather than as flags threaded through the capture: a control made of scattered ifs is not a control.

The sweep of every other instrument in tools/

tool has the hole?
local-pair-capture.py yes — fixed here
fidelity-sweep.py no. Arms per target at line 220, releases at 298.
fidelity-live.py no, but it does not check. It reads the live host over the agent socket, and the host arms P3 itself (NOWMirrorContentPlane.swift:298). A run against a host that never armed would report every window stably empty and read as a stability result; nothing in the run says which. Worth a guard, not a defect today.
mirror-corpus same shape as fidelity-live: its ir read comes through the host. It records contentGeneration, which is a partial signal, and does not assert the arm.
fidelity-pair.py no — a post-processor over a sweep directory; captures nothing.
local-arm-latency.py no — arming is its subject.
local-observe-plane.py, local-plane-lapse.py, scene-delta-bench.py, local-scene-bench.py, local-control-drive.py, local-act-pump.py no — plane binding, wire bytes, walk timing, control classification, act latency. None claims anything about interiors.
mirror-gate, shot no — a rules ledger and a screenshot.

Two things noticed in passing, not chased

  • The attribution fingerprint in circulation is unsound. <label>-guest.ppm is written by both tools — fidelity-sweep.py:293 and local-pair-capture.py:217 — so it cannot tell them apart. The discriminators that hold are manifest.json and a surviving -guest.png (pair capture only) versus sweep-summary.json, LIMITS.md and <label>.json (sweep only). Anyone re-deriving an attribution table should use those.
  • local-pair-capture.py still has no --expect-build. Every QEMU guest on this Mac sees the host as 10.0.2.2, and this tool believes whoever dials. The build was checked by hand for this run; it is not checked by the tool. fidelity-sweep.py's --expect-build auto is the pattern.

FIXED (all four breaks) and EMULATOR-VERIFIED: "drag isn't working on my build" — where the gesture stopped, in four places (2026-08-07, claude/019-drag-live, 019-drag-element-refs, 019-drag-break-4)

Read this section in order and to the end. It is the record of four walls found one behind another over a single day, kept as it was written because the sequence is the useful part — and the last entry CORRECTS the diagnosis in the second-to-last. An item moves now; the evidence is at the bottom.

Michelle, 2026-08-07: "drag isnt working on my build". It is not working, it never has, and the reason is not in any of the code the arc spent two slices writing. Both halves are built and neither is connected to the other.

Her build is PID 17303 out of /private/tmp/int5-app, which is claude/019-integration-5 — a branch that carries slice 10 and slice 10.5 together, so everything below is about the tree she is running, not about a stale checkout.

Break 1: ItemDragDriver has no conformer anywhere in the tree

LiveMirrorView.beginItemDrag asks controller.itemDragDriver for someone who can hold the mouse button down (mirror/host/MirrorKit/Sources/MirrorKitUI/LiveMirror.swift:479). MirrorSceneSource declares it at :52 and gives it a default of nil at :74. The application's only conforming type, now-host/Sources/Host/NOWMirrorSource.swift:216, does not override it — so it is nil in the running app, and every item drag a person makes takes the refusal branch:

this mirror cannot hold the mouse button down — (name) was not dragged

Derived rather than remembered:

grep -rn ': ItemDragDriver\|ItemDragDriver {' . --include='*.swift'

answers with the protocol declaration and the nil default and nothing else. There is no conformer, and there is no host-side plumbing for the three verbs either: grep -rn 'dragpress' now-host/Sources is empty. AgentIntegrationActControl serves winact, ctlact, menuact and key; it has never had a drag verb.

So the honest answer to "is it being refused, and with what reason" is: yes, and the reason is that the app has no drag driver. Nothing was ever sent. dragpress has never left this Mac from the application.

The refusal message is well-written and specific, which is the one thing that went right here — but it is a status-line note, and a status line that scrolls is how a deliberate refusal reads as a dead feature.

Break 2: a Finder icon cannot be named to dragpress

This is the deeper one, and writing the missing conformer does not fix it. dragpress takes element"the opaque reference, now-element- and a UUID, minted by the observation that saw the element" (contract/asyncapi.yaml:4682). DragTargeting.Subject is a Scene.DesktopItem (Scene.swift:521-551): name, kind, type, creator, x, y, w, h — and no reference at all. Desktop and folder-window icons arrive through FinderItems' AppleScript, not through the element walk, so no observation ever minted one for them.

The same wall was already written down from the other side, under "a self reference described every one of our own controls as {0,0,0,0}": the elements and observe walks report a process with bind: no-plane and an empty windows array, so the wire cannot mint the reference and reaching a drag from a host "needs either that mint path on the wire or the drag verbs themselves". That sentence was about NOW's own controls. It is equally true, and unremarked, about every Finder icon — which is the only thing DragTargeting can pick up.

What this means for the two slices

Nothing built in slice 10 or 10.5 is wrong. The vehicle fires and the dead-man lets go (entry below, emulator-verified). Targeting, the session state machine and the provisional presentation are unit- and render-tested. The arc built the two ends of a bridge and no span.

The homeIsTrustworthy refusal at DragTargeting.swift:221 was the first suspect and is not the cause: it is reached only after a driver exists, and the code path stops one guard earlier. That refusal is correct and must stay — a drag aimed at a guessed target moves the wrong file.

What closing it actually costs

Two pieces, in this order, and the second is the real work:

  1. A way to name a Finder icon on the wire. Either the guest mints element references for Finder items (a new observation surface), or the drag verbs gain a by-position form and take the refusal-rather- than-guess rule with them. This is a contract change and starts in contract/asyncapi.yaml.
  2. dragpress/dragmove/dragrelease through AgentIntegrationActControl, and an ItemDragDriver conformer on NOWMirrorSource. Mechanical once (1) exists; it follows the winact/ctlact pattern exactly.

Until (1), the honest product behaviour is the refusal that is already there — but it should be stated where a person sees it, not only on a status line, because "I dragged and nothing happened" is what four hours of this investigation cost.

The half that does work, driven rather than reasoned

Emulator-verified on this lane's own clone — block 113, anchor 12904 / wire 12905, a fresh clone of os91-runner.qcow2 carrying this tree's ext and app, guest build 113f1b176035, capabilities 511 with bit 7 (kNowPeekTableCapDrag) set, actselftest abi-agreed. Nothing here ran on the PowerBook. tools/local-drag-vehicle.py:

what what the guest did
the Time Manager task fires 26 ticks across one press, Moves applied 0 → 1 → 2
the button goes down and stays State=held Button=down for the whole gesture
a move is consumed pointer followed to h=300 v=200
the dead-man releases without being asked idle=60 ticks, nothing sent, ended by itself in ~1.0 s
a fresh press afterwards succeeded — the vehicle is reusable

So the far end of the bridge is real. It is the span that is missing, and that is worth saying plainly: a green run of that probe says nothing whatever about the product, and the probe now says so in its own output rather than leaving a reader to infer it.

How a Finder icon gets named, and the fact that decided it (2026-08-07, claude/019-drag-element-refs)

Break 2 says a Finder icon has no element reference. The obvious repair — mint one — was written up, costed, and rejected, and the reason is worth more than the design that replaced it.

Candidate A: teach the element walk to see Finder icons. The walk reads foreign memory along two documented chains, the Window Manager's window list and each window's control list (axwalk.h). A desktop or folder icon is in neither: the Finder lays them out in its own private structures, which have no published layout and no stability guarantee across 8.6–9.2.2. Seeing them means reverse-engineering the Finder's heap, and the failure mode of getting it subtly wrong on a system version nobody tested is a file moving. Rejected.

Candidate B: FinderItems mints the reference. FinderItems is host code, and the registry's first anti-forgery property is that the token is not derived from anything a caller knows (obsref.h). A host-minted token is by construction a caller-supplied token. Rejected as stated.

Candidate C: the guest asks the Finder and mints there. Honest in principle — the Finder is the only observer that can see its own icons, and the guest already reaches it through script/OSADoScript. The cost is where it fails: the registry re-proves a reference from foreign memory — window address, ControlHandle, node fingerprint — and a Finder item has none of the three. Its revalidation could only re-prove the process, so a Finder-item reference would be weaker than every other reference while spelled identically on the wire. Uniform spelling over non-uniform strength is the convincing-lie shape this repository keeps paying for.

And then the fact that settles it, read out of the resident rather than reasoned about. For dragpress — unlike ctlactthe reference does not bound the act. ext/src/now_ext_act.c:829 serves the press as now_ext_drag_press(table, cell->control_handle, a5, cell->click_h, cell->click_v, …): the button goes down at the caller's point, and control_handle is repurposed to carry the session nonce, so no ControlHandle is checked and no trap patch answers for one. The element therefore buys exactly two things — the PSN whose context the press runs in, and a fallback rectangle for a caller that sends no point. A drag always sends a point.

So the element's whole job in dragpress is to name the process, and candidate C would have built a weak new species of reference to name the Finder, which a Finder window reference already names with far more re-proving behind it.

Chosen — candidate D: dragpress names its container. It accepts a now-window- reference as an alternative to element; in that form h/v are REQUIRED, because there is no rectangle to fall back on and the rule against pressing at a guessed point stands unchanged. And the reference is made to do real work rather than sit there as provenance: the press point must lie inside the resolved window, which the guest checks from the resolver's own global content rectangle (NowAxWindow). A window reference plus an arbitrary global point would have been "a coordinate is a reference" arriving by the back door; a window reference plus a point inside that window is a bound the guest can check and does.

The one thing that could have left this half-closed was the desktop: a desktop icon is inside no folder window, so the form reaches it only if the Finder's desktop is itself a window in the window list. That is a fact about this Macintosh rather than a design choice, so it was measured rather than assumed — and it came back better than the design needed.

The Finder's desktop IS a window, and the guest already mints a reference for it. Emulator-measured 2026-08-07 on this lane's own clone (block 979, anchor 19832 / wire 19833, guest build 113f1b176035), by aiming elements at the Finder's own PSN:

window kind z content bounds (global) reference
Macintosh HD 20 0 {48, 103, 452, 321} minted
Desktop 20 1 {0, 20, 800, 600} minted

So the desktop is an ordinary Finder window of the same kind as a folder window, its content region starts below the menu bar, and the element walk had been naming it the whole time — nobody had asked. No new species of reference is needed for any Finder icon, which is the strongest possible argument that candidates A through C were the wrong shape: the addressing this feature wanted already existed, one verb over.

The bound the guest checks falls out of the same fact. A desktop drag is a press inside Desktop's rectangle, a folder drag is a press inside that folder window's, and both are the same comparison against the same global content box.

Break 3, found by driving: nothing ever posted the mouseDown

The window form was built, the conformer written, and the gesture driven against a real Finder — and the icon did not move. The vehicle was never going to start a drag, and the reason is one line that was never there.

now_ext_drag.c says in its own header that the whole vehicle is "write those four" low-memory globals, and that is true. But writing MBState is what a tracking loop reads once it is running; it is not what starts one. An application begins a drag because a mouseDown arrived through GetNextEvent, it hit-tested the point, and it called DragGrayRgn. No mouseDown was ever queued.

That stayed invisible because the plane's other users do not need one: ctlact arms a patch that answers TrackControl for a handle the request names, so the target is already inside a loop when the button matters. A drag starts from outside one. local-drag-vehicle.py could not have caught it either, and says so in its own header — it aims at NOW's own window and claims nothing about what an application did.

So the DragPress serve now posts a mouseDown — one event, where stamped on the queue element, in the target's own context, which is where that serve already runs. Only the down: the up is what the vehicle's deadline owes, and queueing it here would end the gesture the instant it began. A refused queue abandons the gesture rather than being best-effort, the opposite of the mouseUp's policy and deliberately so — there the button is already up, here it would be down with nothing tracking it.

It works, and it is not enough. Emulator-verified, same rig as above: with NOW hidden and the Finder frontmost, the press reaches the Finder and the Trash selects itself under the pointer — the guest's own pixels, dt-1-pressed. Nothing in this project had ever made a foreign application respond to a drag press before. Which is what let the next wall be found.

Break 4: the press's own reply is blocked by the loop the press starts

With the Finder genuinely tracking, dragpress answers:

State = ended    Button = up    Vehicle ticks = 253
Moves applied = 0    Ended = dead-man-idle

The gesture is over before the caller learns the session nonce. Not a timing accident — it is the plane's own central fact, arriving one step earlier than the contract expected it. dragpress's description already says that from the instant the button is down "the filter that serves act requests is never entered again until the gesture ends", and concludes that motion and release therefore cannot be act requests. True, and incomplete: dragpress is itself an act request, and now_act_submit waits for the target to pump before it can report fired. The target stops pumping the moment it starts tracking. So the press cannot answer until the drag is over, and the only thing that ends it is the dead-man.

The circle: the press starts the tracking loop → the loop stops the target pumping → the press's reply cannot be composed → no dragmove can be sent, because it needs the nonce the reply carries → the idle deadline expires → the loop ends → the reply finally arrives, saying dead-man-idle, Moves applied 0.

This is stated and stopped at rather than patched. The shapes a repair could take all change the plane's contract, and picking one is a design pass rather than a fix:

  • the nonce could be minted by the caller and sent with the press, so nothing has to come back before motion can start;
  • or dragpress could answer on arming rather than on firing, which means saying "the press is queued" instead of "the button is down" — a weaker and more honest claim, and one that changes what an ok reply means;
  • or the resident could publish the session in the drag cell, which any reader can poll, so the press's reply stops being the only route to it.

The first is the smallest and the most suspicious: a caller-supplied nonce is a caller-supplied identifier, and that family of shortcut is what obsref.h exists to refuse. It is probably still right here — a nonce is not an address — but it is exactly the argument that wants to be had out loud rather than settled by whoever is holding the keyboard.

What this means for the two breaks above. Both are closed and both are real: a Finder icon is nameable, the conformer exists, and the press lands on the right icon in the right process. What is not closed is carrying the gesture, and no part of the product should claim otherwise until Break 4 is.

CORRECTED BELOW. The paragraph beginning "The circle" is wrong in its middle step, and all three candidate repairs it lists are answers to a question that was not being asked. The reply is committed BEFORE the tracking loop exists; what is missing is not a channel but a scheduled reader. See "break 4 was never about the act filter", two entries down, which measured it and closed it.

The resident change here carries NO bake receipt, and nothing stopped it

Said out loud because the mechanism that should have said it did not run. ext/src/now_ext_act.c changed in this thread, which is exactly the case tools/ext-bake-gate exists to catch — and scripts/test-all ended with:

all gates passed — BUT THE COMMIT GATES ARE NOT ARMED IN THIS CLONE. Nothing refuses a commit here: not the main guardrail, not the ext bake gate.

So the deferral was typed (TBT_DEFER_EXT_BAKE_REASON) and went nowhere: ext/stage-receipts.json is untouched by every commit in this thread, because the hook that would have written it never ran. A deferral is supposed to be a written decision that lands in the same commit as the work it excuses; this one is written here instead, which is the only place left for it.

The decision itself stands and is small: the resident under test was this tree's build, staged into a session-private clone and cold-booted — scripts/spin-up-ppc's ordinary mode, recorded in that run's own provenance.md — and no shared image was baked or touched. What is missing is the receipt, not the evidence.

Two things follow for whoever picks this up. Bake before this lands on main: merge-check refuses resident source covered by no bake, and it will be right to. And tools/setup-hooks is per clone, so an agent worktree under /private/tmp can be a clone where nothing is armed at all — a green test-all there means the tests pass and says nothing whatever about the guardrails. That warning is well-written and it is the only reason this was noticed.

A gate, so the next conformer is not silently absent

Nothing failed when itemDragDriver went unimplemented, because a protocol default that returns nil is a legal conformance. That is the same shape as "every other gate can be green while neither guest compiles": a seam with a default is a seam nothing checks.

ItemDragSeamTests now checks the conformer and the dragpress verb as a pair, and fails in either direction — verb without conformer (today, once someone writes the plumbing and stops) and conformer without verb (which would put a moving ghost on screen over a gesture no guest received, since the view begins showing a drag before the guest answers). Both directions were watched failing under mutation.

And a probe that failed a correct guest

Found on the way, and the sharper lesson of the two. local-drag-vehicle.py phase 4 required dragpress to refuse an element sent without h/v. That was never a property of the rule; it was a property of resolve_self leaving detail zeroed, so every self element resolved to {0,0,0,0} and there was no rectangle to fall back to. Another lane fixed that hole, the refusal stopped arriving, and the probe reported the fix as the defect.

Settled by asking the machine rather than the code: the scene reports the control content-local at {135,10,151,340}, the window content origin is (28,70), and the resolver answers {163,80,179,410} global — the two conventions agreeing exactly, which is the thing the resolve_self entry below said it could not drive. The centre (171,245) is precisely where dragpress pressed.

The phase now asserts the rule instead of the symptom: the point comes from the resolver, the reply says which of the two it used, and it is never 0,0. The refusal branch is now unreachable from that rig — it can only mint self references and those always resolve — and the probe states that rather than dropping it silently.

The generalisation is the same one this repository keeps paying for from a new angle: a probe's precondition is a derived claim, and it rots when a sibling lane repairs what it was derived from. A stale oracle does not go quiet; it goes red against correct code, which is more expensive than silence because someone believes it.

FIXED and EMULATOR-VERIFIED: break 4 was never about the act filter — nothing was SCHEDULED to read a reply that had already been written (2026-08-07, claude/019-drag-break-4)

The entry above says the press's reply "cannot be composed" because the act filter is never re-entered. Read the serve path and that is not what happens. now_ext_act.c :: now_ext_act_apply serves the drag press and then calls now_act_serve_commit(cell, error) in the same jGNE pass, three statements later — before the filter returns, before GetNextEvent hands the queued mouseDown back to the Finder, and therefore before any tracking loop exists. now_act_serve_commit (now-guest-shared/src/now_act_guard.c:400) writes status = kNowPeekActStatusDone, which is the exact word now_act_submit is spinning on.

So the reply the caller waits for is written before the drag starts. The act filter's re-entry has nothing to do with it. What is missing is a reader: classic Mac OS is cooperatively scheduled, the Finder is inside DragGrayRgn calling StillDown/GetMouse and nothing that yields, and NOW does not get the processor at all until that loop ends. The status word sat there Done for four seconds with nobody running to look at it.

That reframing kills all three candidate repairs the previous lane listed, and it is worth being explicit about why, because each of them is a correct answer to the wrong question:

  • a caller-minted nonce — already true. now_act_run_dragpress calls now_act_drag_next_session() and ships the nonce into the cell in control_handle; the resident never mints one. The reply was never the route to the nonce. This candidate was refused as an obsref.h-shaped shortcut and it turns out it did not need refusing, because it was not a change.
  • answering on arming rather than on firing — weakens what ok means and buys nothing: the app still cannot run to send either answer while the target holds the CPU.
  • publishing the session in the drag cell for a poller — there is no poller. The only process that could poll is the one that is not being scheduled.

What follows is not a repair to any of the three: it is that the gesture must be handed over BEFORE the button goes down. Nothing NOW sends can reach the guest between the press and the release, so the destination has to travel with the press and be applied by the one thing that does run at interrupt time — the Time Manager vehicle that is already there.

Chosen: dragpress carries toH/toV. now_drag_begin seeds want_h/want_v with the destination and want_seq = 1 instead of zeroing them at the press point, so the vehicle's first tick consumes a want that was published before the loop started. No new field, no change to sizeof(NowPeekDragCell), no shift of the cursor plane behind it — which matters, because the resident and the application are baked and built separately and a layout change is the one thing that cannot be half-deployed. The idle dead-man then releases the button at the destination about a second later, and reports dead-man-idle, which is the truth and should not be dressed up as released-as-asked.

The act deadline versus the writer lease, which this makes concrete. act_yield renews the writer heartbeat by calling now_peek_idle() — but act_yield only runs when the application runs. Under the reading above the application does not run for the length of the gesture, so the 180-tick writer lease lapses at 3 s of any drag longer than that and the resident disarms every plane underneath a gesture still in flight. The toH/toV form sidesteps that too: it needs nothing armed after the press.

AN ITEM MOVED, and the machine said which one

Emulator-verified 2026-08-07. Lane block 436, anchor 15488 / wire 15489, a session-private clone of now-mirror-stage.qcow2 (sha256 c466baa9…, volume clean) carrying this tree's ext and app, guest build 39ac9db60598, resident active capabilities 511, actselftest abi-agreed. tools/local-finder-drag.py, twice:

item press destination the FINDER's own bounds of, before → after
From Claude.txt 624,556 488,76 {608,540}{472,60}
HELLO_CLAUDE.txt 624,108 536,76 {608,92}{520,60}

Both landed where they were aimed. The oracle is the Finder answering bounds of for the item it moved, not our arithmetic; the guest's own screendumps are beside it because a picture is what a person asked for, and they show the icon in its new place, selected, with the gap it left. DragGrayRgn had never been measured in this project and it is measured now: it follows.

The resident's account of the same gesture:

State = ended   Button = up   At h = 488  At v = 76
Vehicle ticks = 59   Moves applied = 1   Ended = dead-man-idle

One want, consumed on the vehicle's first tick, published before the button went down. dead-man-idle is the truth about how it let go and is deliberately not dressed up as released-as-asked — nobody asked, because nobody could.

And the measurement that settles the mechanism

Submit ticks 67, Submit yields 1. This application got the processor once in the 1.1 seconds its own gesture ran. Second run: 68 ticks, 1 yield. That is the discriminating pair — a slow resident and a starved application both read as "the press took a second", and the repairs are opposite ones. The instrument counts a loop that already existed rather than adding a poll, which is the distinction finding instrument-feeds-the-clock was written about.

So the entry above is confirmed: the reply was written three statements after the mouseDown was queued and sat in the cell, coherent and unread, until the Finder gave the processor back. Every act longer than a target's own tracking loop has this shape, and it is a property of cooperative multitasking rather than of this plane — which is why the repair is not in the protocol at all. It is that the gesture travels with the press.

What is still not closed

  • dragmove and dragrelease are unreachable for this target class, and the contract now says so rather than leaving a caller to find out. They are not dead: an application that yields inside its tracking loop can still be driven that way. Nothing has measured one that does.
  • The drop is the idle dead-man, so a gesture holds for about a second at the destination before it lets go. That is a hand-movement duration and it looks right, but it is a side effect being relied on: a drop the press could ask for would say what it means.
  • Only ONE want travels. A path — press, via, drop — would need a queue in the drag cell, which changes sizeof(NowPeekDragCell) and therefore needs the resident and the application deployed together. Not needed for an icon; needed for anything that must avoid something on the way.
  • The resident change carries no bake. Typed deferrals are recorded in ext/stage-receipts.json for every commit in this thread. The evidence is a session-private clone with this tree's build staged in and its fingerprint checked against the local build; what is missing is the shared oracle, and it was deliberately not baked while other lanes were running against it. Bake before this lands.

FIXED and EMULATOR-VERIFIED: the Finder's list rows were unclickable because a column header was read as a scroll bar — and Extensions Manager's list is not a control at all (2026-08-07, claude/019-list-selection)

Michelle: Finder list view and Extensions Manager are "completely not selectable." Two targets, two different answers, and neither was the one the arc had been chasing.

Rig. mac99 / OS 9.1, session-private clone of os91-runner.qcow2 (sha256 f34f7e5d…, plain base, never baked), lane block 755, anchor 18040 / wire 18041, guest build 113f1b176035, actselftest abi-agreed. Every scene capture discarded a warm-up pass and was repeated; every claim below is paired with a QMP screendump.

The taxonomy, re-decided per target rather than assumed

The three list flavours the 018-control-semantics entry distinguished are the right frame, and neither of these two targets is flavour 1.

target flavour what the control walk sees
Appearance's Themes / Patterns lists 1 — kControlKindListBox the control, classified listBox
Extensions Manager 2 — not a control 2 scroll bars, 3 push buttons, 1 popup. No list.
Finder list view 3 — the Finder draws it 13 controls: 8 push buttons (the column headers), 2 scroll bars, a triangle, a static, one unknown. No list.

Both readings were steady across repeated passes.

The Finder: everything was right except the one thing that aimed

The rows' geometry has been correct since 2026-08-06 — bounds of gives the box the Finder drew, 16×16 at l=22, 19-px pitch — and FinderItems.clickPoint, HitTester.windowItem, ObjectResolver and InteractionPolicy all resolve a point to finderSelect by name. Every link was individually right and the chain still produced nothing, because FinderItems.iconArea bounded the clickable field by every visible control, telling a bottom scroll bar from anything else by SHAPE: wide and short.

A folder window in ICON view has exactly two controls and the shape held. A window in LIST view also has its column headers — "Name", "Date Modified", "Size", "Kind" — which are wide, short, and at the TOP, so each one pulled the field's bottom up to its own top. Measured on this run, Macintosh HD in name view:

icon field = 0,41,389,21      <- a bottom ABOVE its top

So every row had no click point, every click fell through to bare window content, and on a guest with no positional-click verb nothing happened at all. The two views differ by exactly the controls that broke one of them, which is why icon view worked throughout and this survived three lanes.

iconArea now reads only the window's own SCROLL BARS, which is the rule its own doc comment always stated. Field becomes 0,41,389,203.

Driven and watched. The point (78,249) — the centre of the box the Finder itself drew for row 7 — resolves through MirrorKit's real hit test, object resolver and interaction policy to finderSelect "TBT" of window "Macintosh HD"; the guest was asked for exactly that act and selected TBT. name of selection went {}"TBT", watched in a paired before/after screendump with the row highlighted. Eight of the ten rows resolve to their own name; the two below the horizontal scroll bar correctly resolve to nothing, because an item scrolled out is not a target.

The inch still unproven, and it is host UI plumbing: the drive above begins at a POINT and ends at a selected row, but the point was handed to HitTester by a script rather than by an NSView. No agent-socket gesture takes a raw point, so only a human's mouse in the Mirror window exercises that last call.

What is still not selectable in the Finder, and why: the file's NAME. bounds of answers the 16×16 row icon and says nothing about where the Finder drew the text, so a point on the name is bare content. The window's own "Name" column header measures the column (content 0..214) and that is not the same statement — OS 9 does not select from the empty space after a short name, and nothing here has measured the text. Widening to the column would be a guess about the Finder's hit rule rather than a reading of it, so testTheNameColumnIsNotYetATarget pins the current behaviour and says what would have to be measured first. From where a person sits this is most of what "not selectable" still means: the target is a 16-px icon at the far left of a 400-px row.

Extensions Manager: flavour 2, and it is blocked on the resident

Its list has no control anywhere in the window's chain — the walk finds the scroll bar beside the list (max 146) and nothing for the list itself. So all four requirements fail at once: nothing mints a ref, nothing reports a rect, there is no action to authorise, and no point to aim. The semantic assist cannot help: now_semantic.c :: list_for_control reaches a list through GetControlData(kControlListBoxListHandleTag) on a control, and there is no control to ask.

Closing it means the resident finding and reporting a list that is not behind a control, and reporting per-row geometry — NowPeekSemanticRecord carries row, column, flags and text, and no rect. That is contract/peek_table.h and ext/, which is a bake, so it is parked here rather than attempted. Per the standing rule: if the machine cannot say where a row is, the honest answer is that it is unclickable — a click aimed at a guessed row selects the wrong file.

A third thing, because it bit the instrument that makes these claims

tools/local-control-drive.py computed its global point as the SCENE window rect plus a content-relative control rect. windows[].rect is the structure box and controls[].rect is content-relative, so every point it aimed was one title bar — 20 points — too high. Measured both ways on this machine: Extensions Manager reported windows[].rect.t = 51 in the scene and bounds.top = 71 in axtree; the Finder's "Name" header (content t=21..42) lands at global 124..145 with the axtree origin and at 104..125 — the info bar above it — with the other. It now takes the content origin from axtree, whose window bounds ARE the global content rect, rather than from a constant written in the tool. A tall control absorbed the error, and the drives that found this instrument useful were the tall ones; a 16-px list row is smaller than the error.

Positive control for flavour 1, on the corrected origin: ctlact part 0 at global (504,221) selected "Lime Horizon" in Appearance's Themes list box — row highlighted, the panel's own "Current Theme:" changed from "Indigo Foam", and the whole desktop repainted. Paired before/after screendumps. So a real listBox row IS selectable by a click at a point on this build.

FIXED: a reply that would not fit closed the socket, and every verb had it (2026-08-07, claude/019-snapshot-refusal)

Sweep C's headline: mirror_read --intention snapshot closed the connection without replying, 3/3, while status, metrics and find answered normally on the same path. It silently disabled tools/fidelity-live.py — the only instrument that can see live render flicker — and it was the one refusal in this tree with no reason attached.

It was not a defect in mirror_read. AgentIntegrationLocalServer.finish encoded the response with try? and returned on failure, and its defer closed the socket. So any operation whose answer passed the 64 KB ceiling hung up with no error frame, no code and no reason. One verb was seen doing it; every verb could. Fixed at the transport, where the one exit is: AgentIntegrationLocalCodec.encodeOrRefusal substitutes a bounded response-too-large refusal naming the operation, the size the answer reached and the ceiling it met.

Why the snapshot overflowed at all, which is the more instructive half. Four families make up a snapshot and only two were bounded. itemBudgetBytes (40 KB) and contentBudgetBytes (12 KB) were constants chosen against a stress fixture that has two entities, no menubar and no coverage rows — so entities and menus were governed by nothing, and on a real OS 9 desktop they are not small: nine menus, and the Apple menu alone can hold 96 items. The ceiling held in the test and broke on the machine.

Then charging each family for its contents still left every CONTAINER uncharged — 48 window wrappers measured 11.4 KB, a fifth of the ceiling no budget had ever seen. That is the same omission shape this file already carries three instances of, arriving a fourth time. So the assembled snapshot is now encoded and checked, and a pass that does not fit adds the miss to a reserve and rebuilds. A loop that verifies cannot be wrong about an accounting it forgot; the arithmetic that reasons about wrapper sizes can be, and was — twice, inside this one change.

Every bound states its count: entityTotal, a menu's itemTotal, beside the surfaces' existing itemTotal and displayTotal.

Watched fail, three ways. Both new socket tests died with the transport's own io("read") — the closed socket — before the change. The new menu-heavy fixture threw messageTooLarge under the old constants and under the first, unreserved arithmetic. 61 663 bytes of 65 536 after.

Emulator-verified for the serving half: on this lane's own VM (block 127, anchor 13016 / wire 13017, guest build 113f1b176035 2026-08-07T17:19:58Z), snapshot answered 3/3 where sweep C got three hang-ups, and a six-control-panel scene — 6 surfaces, 205 elements, close to sweep C's 180 — served in 52 354 bytes.

What is NOT verified, and why

The B/C-side flicker measurement was not taken. It was the second half of this lane's brief and it is still owed. The run was stopped mid-teardown on a report that an agent had reached a human's running host; that turned out not to be so (next entry), but the stop was correct and the measurement needs a host launched under the isolation the next entry fixes. The A-side baseline therefore still has no B or C side — no longer for the reason sweep C recorded, which is now closed.

The rect-owner-flip question — zero on the A side — remains unanswered.

FIXED: NOW_PREFS_SUFFIX isolated the port and almost nothing else (2026-08-07, claude/019-snapshot-refusal)

Found while establishing whether this lane had reached a human's running host application. It had not, and the evidence is worth keeping because the alarm was raised twice on two different theories and both were wrong:

  • Each guest was paired with its own host throughout. At teardown, qemu 13498 (the human's VM) held two ESTABLISHED connections to Host 17303 on 16729; this lane's qemu 42122 held two to this lane's Host 77496 on 13017. No crossing, in either direction.
  • The two hosts held different agent sockets — …now-agent-501-int5/host.sock and …now-agent-501-snapref/host.sock. Every now-agent call this lane made carried NOW_AGENT_SOCKET_SUFFIX.
  • The human's settings suite (…settings.int5.plist) was last written at 12:41:08, 59 minutes before this lane wrote anything, and the base suite (…settings.plist) at 19:28 the previous day. This lane wrote only …settings.snapref.plist.

But the isolation a lane relies on is much narrower than its name, and nothing said which half it covered. ProductIdentity.preferencesSuite scoped four call sites. Thirteen others defaulted to UserDefaults.standard — the application's own bundle-id domain, one store shared by every host copy on this Mac. A run launched with NOW_PREFS_SUFFIX therefore isolated its listening port while its Mirror app path, QMP socket, forwarded agent port, plane policy store, file locations, host share directory, cloud settings, screenshot settings and sidebar still pointed at the desk's real preferences — and wrote into them, live, with somebody's session open. mirror.appPath, mirror.qmpSocket and mirror.forwardedAgentPort are in that shared domain, and they are exactly the settings that decide what a Mirror is looking at.

One accessor decides it now, ProductIdentity.defaults, and it is .standard when no suffix is set, so a shipped launch is unchanged and no existing preference moves. The guard reads the source, because the defect is a default parameter value and is invisible at every call site that relies on it; it named all thirteen before the change.

Still open, and it is the design question underneath

A host application cannot tell whose agent is talking to it, and the person at the machine gets no signal about what their session is pointed at. requireTheBuildUnderTest() exists because any guest can answer your listener. This is the mirror image and nothing covers it: a guest dials a port by number and cannot tell whose host answered, and a host accepts whatever dials in. Isolation here is entirely by convention — suffixes and a port block — and allocation is not possession: a lane that reserves a port has reserved a name, not a socket.

Worth proposing, none of it built: a session identity carried on the agent socket handshake; a refusal to serve a host that has a live human-driven session; and — cheapest and most valuable — the app stating in its own UI what it is pointed at and who pointed it there. The failure mode this lane's alarm assumed is not currently detectable by the human at all: an app that got repointed would keep running and look entirely normal.

RESERVED FOR A DECISION, with the numbers attached: one window interior at a time (2026-08-07, claude/019-multi-window-content)

Plan 019 slice C. The content plane P3 is a per-window spotlight, not a plane, and the question is whether that is the product. This entry is the options and their measured costs; the choice is Michelle's, and nothing here has been implemented except the half that is owed under every option (below).

What actually bounds it, and it is NOT the resident's hook

The one-window limit is one word in a shared struct and one parameter in a verb. Everything else on the path is already per-window:

  • qdtrace start takes exactly one window and refuses an all-windows arm by name (now-guest-ppc/src/content/qdtrace_cmd.c :: run_start).
  • The block carries one arm_window / active_window (contract/content_table.h), and the capture gate is one comparison — port != gBlock->active_window (ext/src/now_content.c:236), with the probe's sight gate the same test again.
  • content_install_exact_window walks the WindowList and installs the ONE port whose address matches.

But every ring record header already carries port — the comment in the contract calls it "the window identity key" — and the resident's port table holds kNowContentMaxPorts = 16 rows which it ALREADY fills with several ports of one A5 at a time (the offscreen worlds hooked at birth). On the host, currentDisplay, settledDisplay, replacementFloor, portStates and latestEpoch are all dictionaries keyed psn:addr, and attachCached attaches a display to every window that has one. The transport, the resident's table and the whole host accumulator are N-window ready today. Only the request cell and the capture gate are singular.

So the cost of N is not a rewrite. It is what N spends.

What N costs — measured, not reasoned

tools/local-multiwindow-cost.py, 2026-08-07. One lane-private clone off os91-runner.qcow2 (block 29, wire 12233), this checkout's ext staged and cold-booted, guest build 113f1b176035, sourceManifest f41867cfe431, buildFingerprint 4d0988e8e891. Finder front with two windows open (Macintosh HD, Desktop). Not the shared stage oracle; nothing on metal.

reading value
arm the Finder's front window 203 ms, hookedPorts settles at 2
retarget to its OTHER window, same process 117 ms, hookedPorts still 2
round-robin, 2 windows, 2 s dwell, 21 s 4,011 ring bytes/s, 10 forced repaints
control — ONE standing arm, same app, 24 s 33 ring bytes/s, 0 forced repaints

Read the control row first, because it is the whole finding. The handshake was never the cost: a second arm into an already-armed process is 69–117 ms, cheaper than the first. The cost is that every arm issues an InvalWindowRect at the newly armed window (content_request_redraw), so a rotation is a forced repaint of a real application's window, at the rotation rate, for as long as it runs. 122× the ring traffic of a standing arm, and 1.3 whole 64 KiB rings in 21 seconds against a ring one standing arm would take half an hour to fill. A host that fell 16 s behind on the drain would start losing bytes it had never seen.

Two more bounds, and one ceiling:

  • The port table is not the binding constraint at realistic N. One armed Finder window plus its interior costs 2 of 16 rows steady-state (9 worlds born and 9 died inside the settle window, none of them surviving the pass). A widget-per-world application is a different story and qdext_born_missed exists to say so, but the Finder leaves headroom for several windows.
  • A retarget drops the LEAVING window's rowscontent_uninstall_context runs on every identity change — so a rotation discards each window's composite lineage every time round.
  • THE CEILING: a background process never arms at all, and that is reproduced on this build with a control. Target NOW's own window while the Finder is front: not armed in 10 s, wrongContext climbing ~56/s (3,046 → 3,610 — other contexts pumping, none of them the target). Bring NOW forward without re-requesting and it arms 204 ms later, with wrongContext moving 3 in that interval. This is the 2026-08-06 arm-latency table's "never, 25–45 s" row, re-measured today against a different target — and a sharper one, because NOW was demonstrably executing the whole time (it answered a status poll every 200 ms from the background) and still never ran the resident's jGNE pass. So no shape here can serve more than the FRONT process's windows. Multi-process interiors are not expensive; they are closed.

The options

(a) One arm, a small window LIST. arm_window becomes a bounded array; the capture gate becomes a membership test over ≤N entries; content_install_exact_window matches a set in the one WindowList walk it already does. One generation covers all N, so no re-arm, no identity churn, no forced repaints beyond the N one-off invalidations. Costs: an in-memory contract change (contract/content_table.h) and therefore a bake; a verb change and its parity seam; ≤N extra compares in the resident's hottest path; N of 16 port rows plus whatever the application's worlds want. Front process only. This is the only shape that scales.

(b) Round-robin within one TTL. Cheapest to build, and the measurement says do not: 122× the ring traffic, a forced repaint of a live application's window at the rotation rate, and every rotation is a RETARGET — which under the decay lane's fix inherits nothing, so each window's picture restarts empty every time round and is complete only until the next rotation takes it away. It also multiplies a known live bug: the blit join looks a held source up by the bits record's generation (NOWMirrorContentPlane.swift, SourceKey(port:generation:)), so any re-arm between a world's ops and its placing blit misses the join and the composite never lands. More re-arms, more misses.

(c) Arm-on-demand from what the host renders. This is essentially what ships: the host arms scene.windows.first(where: \.front). Its retarget cost is (b)'s, paid on a focus change instead of a timer — which is why it is affordable. Extending it to "arm whatever is visible" is (b) with a nicer name.

(d) Accept one at a time, and make the others honest. Free, and owed under every option above. See below.

Recommendation, for the record and not as a decision: (d) now, (a) if and when a second live interior is worth an in-memory contract change. (b) and (c)-extended should be refused on the measurement rather than re-argued later.

The half that is owed regardless, and is now done

Whichever shape wins, the current render is dishonest in one exact place: a window we never armed and a window we armed and found nothing in both arrive at display == nil and both draw the same "Guest content not reported" hatch. One of those sentences is true and the other is a claim about the guest we never made. (A lapsed arm is a different lie and a different lane's: it freezes rather than hatching, and DisplayEpoch.stale is where that one is answered.)

claude/019-multi-window-content closes the never-armed half: Scene.Window.contentPlane records whether P3 was ever asked about this exact psn:addr, and the renderer says "Interior not captured — one window at a time" where it used to claim the guest reported nothing. Host-internal render state, off the wire, by the same argument as displayEpoch and island; no contract field, and no dependence on which option is chosen.

FIXED: the font substitution was silent, and arc-status measured the wrong tree (2026-08-07, claude/019-honest-substitution)

Two honesty defects, both TESTED and neither metal-verified.

The system font is substituted and now says so. Font id 0 is the system font, which under Appearance is Charcoal; where the pack carries no Charcoal strike the renderer falls back to Chicago, Chicago is wider, and a run the application clips loses its last glyph (§ "the system font is Charcoal and the pack has no Charcoal", below). Substituting is now an accepted product decision — Michelle, 2026-08-07: "our fonts are ok at this stage, im happy enough with them" — and this arc's accounting classed the SILENCE, not the substitution, as its one dishonest-by-default item. It had already cost a day: group-box frames appeared to cross their own labels, three people read it as a chrome defect, and it was Chicago overrunning a band sized for a narrower face.

No pixels changed. MirrorKitUI/FontSubstitution.swift makes it answerable — StrikeChoice says what was asked for, what was served, and whether the width came from the machine or from us — and LiveMirrorView carries a banner beside the render, sibling of the asset-pack one. substituted is its own state rather than unknown (nothing here is unestablished: both faces are nameable and the error has a known direction) and a second AXIS rather than a fifth rung of ProvenanceLadder, which correctly calls a substituted run the machine's own ink — which is exactly why the substitution was invisible.

Derived, never remembered: the pack gained a Charcoal rasterisation on 2026-08-06, so a constant saying "we draw Chicago" would have shipped false. The banner fires only when the fallback would actually fire, and reads nil on a desk whose pack has Charcoal — which is the state that makes the REMAINING substitutions the point: font id 2002, every family the pack does not carry, and any desk whose pack predates that work. It does not model style: op.face is still never consulted, so a bold run reports exact and is not.

LiveMirror.cursor(for:) had no test at all and now has one: guest-declared editText is the only I-beam, the match is exact, and the hover resolves through ObjectResolver/HitTester rather than walking the scene a second time — the failure its own comment predicts. Worth knowing before writing another: NSCursor.arrow segfaults in a bare xctest process without an NSApplication.

tools/arc-status was wrong about arc state. It reported "nothing graduated to the corpus in 21h" while 26 findings sat on claude/018-findings across 7 commits — it read the parent's working directory, which sits on main, so it measured a checkout rather than the work. The warning is kept (the worry is legitimate and this arc has genuinely let findings sit); what changed is that it now reads the commits on every ref, and reports written and landed as different states with different repairs.

Three more of the same class, found while in there:

  • A lane cut from another lane counted its parent's commits as its own (a branch with one commit read as seventeen) and the arc's unlanded total counted them twice. Now N commits (M shared with X), and the total is a union. Deliberately symmetric: deriving a direction was tried twice and answers backwards whenever the parent's tip has moved past the fork point, which is the normal state of a live lane.
  • Every table enumerated claude/01[0-9]-* under a heading saying "work that exists". Branches outside the glob are now counted on their own line — 55 of them at the time of writing.
  • The machine count matched qemu-system-ppc only, so a bench running 68K guests read as idle.

tools/mirror-gate-tests/test_arc_status_measurements.py drives the real script against a synthetic repository and its synthetic corpus; a source check would have passed on the day the bug shipped. Every guard watched failing by mutation.

Still open, and not this lane's: the shared checkout at the shared now checkout is parked on claude/mirror-subproject, which predates .githooks/, so core.hooksPath points at a directory that does not exist and the commit hooks are dead in every worktree off it. arc-status says so under GATE FRESHNESS; nothing here fixes it, because moving that checkout is a decision about somebody else's tree (AGENTS.md > Git, "keep the shared checkout on main").

ANSWERED: the desktop verb had no reader, and now the render says who answered (2026-08-07, plan 019 slice D)

Emulator-verified, this lane's own VM (block 963, wire 19705), build 490165bfb441. Nothing here touched the PowerBook.

Slice 9 gave the guest a live answer to what its desktop is drawn from — GetTheme with the Mac OS 9 desktop tags, served as the desktop command — and nothing on the host side ever read it. hasPattern, patternCarried, patternBytes and kThemeDesktopPatternTag returned zero matches across now-host and mirror. Meanwhile the renderer filled the largest rectangle in the picture from the offline asset pack's manifest.json: a record of the disk image the pack was extracted from, true for a guest booted from that image and unchanged since, and byte-identical to the truth whether that still held or not. Two producers of one answer with one of them unread — a seam by the sweep spec's own definition, and it survived because both halves landed before the slice that would have read them.

What the machine actually says, asked live:

source picture   hasPattern true   hasPicture true
patternBytes 16790   patternName "Lollipop 7"
pictureName  <TAG ABSENT>          pictureAlias 154 bytes

That last line is the case the design turns on, and it is measured rather than assumed: the picture NAME tag is absent while the ALIAS carries the file, so this machine confirms a picture without saying which. Kind agreement is then the strongest claim available, and it is labelled as one.

meta.desktop now rides every scene beside meta.theme, from the same gather through the same read_tag — one implementation, two shapes, so the console verb and the scene cannot drift. DesktopPattern.resolve asks it first and reports who answered: machine (the guest named it and the pack held what it named), assetPack (it did not, and the pack is standing in), none (nobody could say, and nothing is substituted).

The honesty half. A substituted desktop draws a plate in the bottom-left corner reading "desktop from asset pack, not this machine". A corner plate rather than a tint over the surface, and the trade is stated where it is made: the desktop is what a fidelity sweep compares pixel for pixel, so washing it would fail every substituted comparison for a reason unrelated to what is being compared. An unmarked desktop now means the machine named it.

Two rules that make the provenance more than a label, both mutation-tested: source: unknown never consults the pack — that value means the machine was asked and refused, and a fallback there launders a measured "I do not know" back into a confident picture — and a pattern the machine NAMED that the pack does not hold renders unknown rather than as a different pattern.

Still open: the guest can name its desktop and cannot hand over a drawable copy of it (the flattened ppat is an identity, not art), so machine provenance still means "the machine chose it and the pack had it". A guest whose desktop is not in the pack renders as the marked unknown, which is honest and is not a picture.

FIXED / MEASURED: the title bar was never drawn from the machine, and the inactive one had never been seen (2026-08-07, claude/019-titlebar-fidelity)

Michelle: "in terms of platinum fidelity, i want someone to focus on the titlebar itself and the buttons on it", and a side-by-side showing Extensions Manager and Sherlock 2 with no widgets at all.

The first finding is about the METHOD, and it is a negative

docs/deriving-a-drawn-procedure.md says to read the QuickDraw drain before writing anything, because DrawThemeTab leaks nearly all of itself through the bottlenecks. The title bar leaks NOTHING. Fourteen captures across three sweeps, 4,550 ops for one control panel, and not one op is in a title bar: every port in every drain is window-LOCAL (origin 0,0, bounded by the content rect), no port carries screen-global coordinates, and no text op anywhere in the corpus draws a window's own title. qdtrace hooks an application's ports and the WDEF does not draw into one.

That is worth knowing before the next piece of chrome: rung 3 answers window CONTENT and nothing the Window Manager draws. Scroll bars, grow boxes and menu bars are the same class of question. The title bar was derived entirely at rung 4 — the guest's own pixels — on eleven front windows.

It was a PER-CLASS defect, not a universal one, and the rule is exact

WindowChrome.widgetBox opened with guard win.kind != 2. Dialog Manager–owned windows are kind == 2, and that is what most control panels are — so Extensions Manager, Memory, Date & Time, General Controls and Mouse were given no widgets at all, while kind == 8, kind == 20 and kind == 2000 windows always had them. Confirmed by rendering the corpus with the old build and counting ink in the band: Extensions Manager and Memory show only face and frame; Appearance, Finder and Sherlock 2 show a widget's bevel. Michelle's "control panels especially" is the whole rule, not an impression.

Sherlock 2 is kind == 20000 and DOES get a widget from the old build when it is the front window, so its absence in her screenshot is not this defect. Either it was not front in that scene, or the front flag disagreed with the machine — that is unexplained and belongs to whoever owns which window is front.

Everything the drawer had wrong, measured

what the renderer the machine
band 18 rows from t+1 22 rows from t-2, black frame row and lit row above the face
stripes white lines every 2 px, full width FFFFFF/777777 alternating, 12 rows, t+2 … t+13
stripe field edge to edge l+15 to seven px clear of the first right widget
stripe phase none dark rows sit one pixel right, patch and all
close box l+1, top t+2 l-1, top t+3
collapse box r-14 r-10, right edge exactly r
zoom box r-29 r-26 where the machine draws one
widget face flat, plain bevel recessed 888888/FFFFFF, 222222 ring, 7×7 anti-diagonal ramp
collapse glyph one centred line two bars, local rows 4 and 6
zoom glyph 5×5 stroked square 7×7 square sharing the ring's top and left
title patch ink + 16, centred on the bar ink + 8, centred on the window
window frame 1 px black ON the scene rect 6 px OUTSIDE it, 000000 · FFFFFF · CCCCCC · CCCCCC · 999999 · 000000, plus a 1 px shadow
content top t + 22 t + 20

Exact-pixel agreement, per rectangle, old build → new, on five windows: widgets 11–15 % → 100.0 %, stripes 0–4 % → 100.0 %, the frame row and lit edge 0–3 % → 100.0 %, the bar's bottom edge 0.2 % → 100.0 %, the whole band 17–22 % → 89–97 %. The band's residual is the title text, where Chicago stands in for Charcoal.

The inactive bar had never been observed, and now has been

Every background window in all fourteen sweep captures is fully occluded; a scan of every screendump in the private store found the front window and nothing else. So one was photographed: a lane-private guest, two control panels, one screendump (titlebar-inactive-2026-08-07/PROVENANCE.md).

The geometry does not change at all. Three colours move — frame 000000 → 555555, the whole band → DDDDDD flat, title ink → 777777 — and the widgets and stripes are simply absent. The baseline stays t+13.

STILL BROKEN, and named rather than worked around

  1. IR v1 cannot say whether a window has a ZOOM box, and kind cannot stand in: Extensions Manager is kind == 2 and has one; Memory is kind == 2 and has none. Seven of eleven windows were being drawn one the machine does not draw, and HitTester reported it — so a zoom act sent a click into the racing stripes, which the Window Manager reads as the start of a DRAG. WindowChrome.zoomBox now answers nil and nothing draws it. The WindowRecord already carries the answer beside the windowKind the walk reads: spareFlag is the zoom flag and goAwayFlag the close flag, one byte each. Contract field first, then the guest read, then both guests. Until then four windows in the corpus have a real affordance the mirror does not offer.
  2. WindowChrome.growBox used to guard on kind != 2 and had the same problem in the other direction — Extensions Manager is resizable and got no grow box. FIXED 2026-08-07 by withdrawal, not by a better field: see "the render invented a grow box on windows the machine leaves plain" above, which carries the per-rectangle numbers in both directions and what the withdrawal costs.
  3. A single pixel at (432, t+19) on the inactive window — where the frame ring around the content meets the bar's last row — renders 696969 against the machine's 555555. A blend at a junction of two integer fills; its four neighbours are exact. The gate excludes exactly that pixel and says so, rather than widening a tolerance that would hide the next real defect.
  4. Semantic control frames are STROKED, and Core Graphics centres a stroke on the coordinate — the same half-pixel DisplayReplay's pixelCentre was written for, still unfixed in the control drawer. A list's boundary reads 128 rather than 0, and which pixel carries it moves with the content origin. Two IslandRenderTests were pinned to the old landing and now take the darkest of the corner's four pixels.
  5. The title's text is anti-aliased by the machine and aliased here. The pixel gate reads the stripes at their ends and the inactive band right of the title for exactly this reason. Chicago-for-Charcoal is an accepted product decision; this is where it shows up as a number.
  6. Modal alert chrome is untouched and unmeasured. The isDialog path keeps its old drawing because nothing here was measured for it, and the one alert in the corpus does not agree with a document frame either.
  7. Nothing here is metal-verified. Emulator only, against QEMU screendumps of Mac OS 9.1.

CORRECTION: four sentences on this page cite a render that could not have falsified them (2026-08-07, claude/019-correct-round3-prose)

Nothing on this page is corrected by editing it, so this entry names the sentences rather than touching them. Named by their entry, because a line number in an append-only file rots the next time anyone appends — these were :228, :527, :631 and :1123 when Sweep C found them and are not those now. Read this before quoting any of them:

Entry The sentence
WAS BROKEN, FIXED … a classified control the renderer will not draw (019-integration-4) "This is why the panel has had no group boxes, no static text and no field values in both rounds"
BROKEN: the integrated render at round 3 (019-integration-3) the Date & Time row — "Still no group boxes at all… there is nothing to draw them as" — the origin
BROKEN: the machine will not say what a foreign control IS (018-control-semantics) "This is also why Date & Time renders with no group boxes."
BROKEN: a FOREIGN process's controls have no determined kind (019-integration-2) "Date & Time has no group boxes.… render as nothing at all"

One sentence outside this file — render-composition.md's "no group boxes in any sweep" — is a claim about the product rather than a dated ledger line, and was corrected in place in the same commit.

What is corrected is the evidence, not the conclusion, and the two are different edits. Conflating them is how the error was made in the first place.

The semantic defect is REAL and STANDS. A DITL row and its ControlRecord are one object seen twice; the tie was broken by asking the control alone for knowledge == .known, derived failed that, and twenty of Date & Time's twenty-one classified controls lost to dialog items carrying kind: null, so a group box never reached drawGroup. Named code path, broken comparison, fix (semanticOutranks), mutation that fails naming it. Do not read this entry as a retraction of that.

What is VOID is the pixel evidence offered for it. SceneBuilder.normalizeWindows sets display: nil unconditionally (mirror/host/MirrorKit/Sources/MirrorKit/SceneBuilder.swift:285, the sole occurrence in that function, no branch), so an interior reaches a render only on a second artifact — the qdtrace drain. Rounds 2 and 3 were pair captures, not sweeps: both stores hold manifest.json and <slug>-guest.png, six renders each, and zero drains. Round 3's own method note says the path independently — "host render via MirrorApp --render-scene over the same envelope". So those renders show no machine-drawn interior whatever the semantics do, and citing them "in both rounds" proves nothing either way. Derived from the artefact stores rather than from prose; fidelity-sweep-2026-08-07-c.md carries the store fingerprints and the discriminating hatch strings.

Two things this does NOT void, said plainly because the temptation is to over-retract:

  • Round 4's mechanism finding stands. It was derived by reading the scene JSON — twenty controls sharing a ref with an untyped dialog item — not from a picture. Only its closing sentence "this is why the panel has had no group boxes … in both rounds" is the void attribution.
  • Sweep A's Finder hatch stands — sweep A meaning fidelity-sweep-2026-08-07-a.md, not the 2026-08-06 A-side baseline quoted below; the two share a letter and this correction is a poor place to be sloppy about which. It is a sweep, it has drains, and "Bitmap unavailable" is emitted inside DisplayReplay (DisplayReplay.swift:835), which cannot run without a drain — so that caption is positive proof the capture had one.

And one claim was outright wrong, in the direction nobody checked. render-composition.md said the panel had no group boxes "in any sweep". That was never derived from a sweep. Neither fidelity-sweep-2026-08-07-a.md nor -b.md contains the phrase "group box" at all, and two 2026-08-06 sweeps — sweeps, with drains, on the same panel — say the opposite in as many words:

  • fidelity-sweep-2026-08-06.md, row 1 — Date & Time is "the best window in the sweep. Every group box, button and label within a pixel or two".
  • fidelity-sweep-2026-08-06-b.md, row 1 — after the DITL-silencing fix, 3/3/2/3/3 and "both group boxes, the time-zone sentence and the time-server line are all back".

Two sweeps in the same directory contradicted the sentence, and one of them was the A-side baseline everything else in that family is measured against. The word "any" was doing work no one had checked. This is the enumerated-list failure AGENTS.md describes, in its cheapest form: a universal quantifier is an enumeration, and nothing enumerated it.

The lesson, which is the part worth keeping. The claim spread from one row of one round to four durable sentences without anyone re-deriving it, and each restatement read as corroboration of the one before. Two rules follow, both already this repository's shape:

  • A count — or an appearance — you did not derive is not evidence. The honest unit is the artefact store, which is checkable, not sentences agreeing with each other across five documents.
  • A live-reading instrument must assert that the plane armed, the way a metal gate asserts which build answered. An instrument that cannot tell you whether it was looking at anything reports absence and defect in the same words. Written into AGENTS.md beside the metal-gate rules in this commit.

LOOK: round 5's five checks, and the one thing four lanes claimed that a picture had to settle (2026-08-07, claude/019-integration-5)

Emulator-capture-verified, on the lane's own VM (block 591, anchor 16728 / wire 16729) against a fresh clone of os91-runner.qcow2not the shared stage oracle — carrying this tree's ext and app: sourceManifest f41867cfe431, buildFingerprint 4d0988e8e891, and the guest's own mirror agreed. Nothing here ran on the PowerBook. Four targets captured with tools/fidelity-sweep.py, rendered through LiveShapedRenderTests.testRenderASweepAsTheAppWouldDrawIt onto each target's OWN scene, and paired against the machine's pixels with tools/fidelity-pair.py.

What four lanes claimed, and what the pictures say:

  • Memory is readable and drawn ONCE. Sweep B's R1 and R2 are gone in a live capture, not only in a fixture: every string in the panel appears exactly once, the labels carry no spurious vertical stroke, and the paragraph, the popup, both value fields and Use Defaults all match. This is 019-sweepb-regressions surviving integration.
  • All six Appearance tabs draw, with Themes correctly raised, the theme list, its scroll bar, Save Theme… and Current Theme: Indigo Foam beneath it. Round 4's 3a — the root userPane erasing the whole panel — does not reproduce.
  • Date & Time has its group boxes, all five, titles included, and "Use a Network Time Server" arrives with its r. Round 4's 3b is gone. Its panel face matches the machine's grey, and its Menu Bar Clock radios draw as filled circles.
  • Cross-application stacking is right, and the Memory and Appearance rows are the evidence rather than the Finder row: in both, a foreign control panel renders in front of NOW's own Workshop window, which renders in front of the desktop — the same order the guest drew. (The Finder row is marked contaminated by the sweep's own dirty-exit check and was not used for this.)
  • The Mirror module's resting layout is the accepted one, rendered from the shipping MirrorModuleView at 620×720 and 1400×900: header, the picture, the status line, and both drawers shut with only the Events chevron showing. Its narrow form truncates the status line rather than reflowing it.

And the thing a picture had to settle, which no gate could: with the pills correctly withheld, Memory's six radios and its "Save contents to disk" check box now draw as bare labels where the machine draws a filled circle and a ticked box. That is the radio-CDEF item above, visible exactly as predicted. Date & Time's radios in the same render draw correctly, so this is the CDEF classification and not the choice renderer. Two smaller absences in the same frame, neither new: the panel's three icons and the RAM-disk slider render empty, and Appearance's theme thumbnails render as plain plates.

FIXED: two of the three the regressions lane named, and the third is a guest question (2026-08-07, claude/019-integration-5)

Round 5 merged seven lanes and then took the two defects 019-sweepb-regressions left standing in docs/open-issues.md on its way out. Verification level: TESTEDscripts/test-all, all five stages, against the integrated tree. Nothing here ran on the PowerBook.

  • FIXED: drawDialogItem's pushButton branch had no coverage gate at all. The regressions lane taught drawControl to consult Coverage and left the identical hole in its twin. Every other branch of drawDialogItem has yielded to the replay since 2026-08-06; this one drew unconditionally, so a DITL row classified pushButton painted a filled Platinum pill over the machine's own ink. It asks words (textCovers || mostlyCovers) rather than inked, because a DITL row is a slack box holding a short run and the area test alone says no on exactly the rows this is about. Gated by LadderArbitrationTests.testADialogItemButtonDoesNotPaintOverTheMachinesOwnInk, watched failing by mutation: 1217 of 1840 pixels over the "Virtual Memory" row with the guard deleted.

  • FIXED: Coverage.mostlyCovers was a single-rectangle test. Its own comment argued that a union is expensive and the cases are not close. The second half is false, and the fixture says so: QuickDraw draws in pieces. A machine-drawn well is a fill plus four bevel strokes; a multi-line paragraph is four or eleven separate text runs. Every fragment falls under half, so the predicate answered "the machine drew nothing here" about a rectangle the machine had covered completely — the ladder's own failure, reached from the far side of the same test. A single rectangle still answers alone and returns immediately; only when several fragments each fall short is their union measured, over a fixed 16×16 sample grid. A grid rather than summed areas, because two copies of one 40% fragment are 40% of something and summing says 120% — a permissive error in a predicate whose whole job is to silence the semantic plane. CoverageUnionTests holds both halves; sweep B's Memory panel has five rectangles the union answers and no single fragment does, and the mutation takes that count to zero.

  • NOT FIXED, and it is a guest question rather than a renderer one: the radio CDEF. Memory's six radios arrive as pushButton and now render as a bare label where the machine draws a filled circle. Michelle will see this. The mechanism is now traced end to end, which is the part worth having: cdef_resolver.c recovers the variation code only when it rides in the high byte of contrlDefProc — it asks the Resource Manager about the raw longword first, and about the masked one second, and takes the high byte as the variant only when the machine itself says the masked value is the resource and the raw one is not. On these controls the high byte is zero, so the resolver honestly reports CDEF 0 variant 0, and control_cdef.c honestly maps that to button. Nothing in the chain is guessing; the variant simply is not in that field on a 32-bit-clean OS 9.

The candidate route is GetControlVariant, which CarbonLib exports and which reads the variant wherever the Control Manager actually keeps it. It is not a small change and it is not a host-side one: it means calling a Toolbox accessor on a ControlHandle belonging to another process, which dereferences a foreign control record through code we do not own. That is a design question against resident-components.md's division — foreign MEMORY reads live in the application, foreign-context EXECUTION does not — and it wants a guest build and a machine to watch it on, neither of which this round had. Left visible rather than papered over: the bare label is honest and the pill was a confident wrong answer.

FIXED: four render defects, two mechanisms, one arbitration (2026-08-07, claude/019-sweepb-regressions)

Sweep B's R1 (Memory's interior drawn twice) and R2 (a vertical stroke merged into every static label), and integration round 4's 3a (a root userPane erasing all six Appearance tabs) and 3b (a classified control losing to its own untyped DITL row). They arrived from three lanes on the same day and they are one question asked in four places: which producer owns this rectangle.

Verification level: TESTEDscripts/test-all green, and every claim below measured offline through the sweep's own harness (LiveShapedRenderTests.testRenderASweepAsTheAppWouldDrawIt over sweep-2026-08-07-b/p1, manifest sha 5de039bf…), rendered against the guest's own screendumps for all nine targets. Nothing here ran on a VM or on metal, and no claim is made about either.

It was two mechanisms, not four bugs

  • Rung 2 drew over rung 1 because nothing asked. drawControl had never consulted Coverage at all — its twin drawDialogItem has since 2026-08-06 — and the test it should have asked, mostlyCovers, is the RECTANGLE's question rather than the text's. Memory's "Disk Cache" row is 102 points holding a 48-point run.
  • derived is a knowledge level and was read as two other things: as words (semantic.value carrying GetControlValue), and as "not good enough to outrank a dialog item that knows nothing".

The full statement is render-composition.md > "The arbitration is asked on BOTH planes". Six gates in LadderArbitrationTests, each watched failing by its own mutation, with sweep B's Memory scene and drain committed as fixtures — and one gate asserting that fixture still exhibits the conditions, because a gate whose capture was quietly replaced by a clean one proves nothing.

What the fix cost, found by re-rendering all nine targets

Two corrections that the named defect would not have surfaced, both caught by looking at targets nobody had complained about:

  • Ground drawn ahead of the chain is wrong for an ARMED window. The first ordering fix put "Structured content unavailable" beneath Date & Time's own Time Zone group. Ground marks an absence; where the replay owns the interior there is no absence.
  • Yielding a check box as one row takes its mark box with its label. NOW's own Workshop lost the tick beside "Compress on wire (PackBits)" — the same per-piece lesson Date & Time taught in August.

What did NOT get fixed, and is now visible

  • OS 9 builds its radio buttons from a CDEF the guest reports as the button FAMILY, so six of Memory's radios arrive as pushButton. With the pills correctly withheld, the machine's own radio MARK does not reach the replay either, so those rows now render as a bare label where the machine draws a filled circle. Before the fix they rendered as a Platinum pill, which was a confident wrong answer rather than a better one. Two separable items: a classification the guest cannot make honestly today, and a capture gap.
  • Monitors' small user pane (nothing inside it) renders as a marked unknown, and its root pane is silent. That is the rule working, not a defect, but it is the first place the two clauses are visible side by side.

BROKEN: the content plane is a one-window spotlight, and every other window hatches (2026-08-07, claude/019-content-never-active)

Michelle's live Planes panel: Structure/Semantics/Interaction Active, Content Requested, Plane bits cap 31, requested 15, active 15, and every window rendering "Guest content not reported". Three separate attributions were offered for that picture. All three are wrong.

Emulator-verified on a session-private clone (lane block 753, wire 18025, guest build 00a33f9fd567, resident sourceManifest f41867cfe431 / buildFingerprint 4d0988e8e891 — this tree's build, staged into a clone, not the shared oracle). Nothing here touched the PowerBook.

The mechanism. P3's arm_active bit is not an echo of arm_request, and it is the only plane of the five for which that is true. P1, P2 and P4 echo unconditionally on the resident's armed pass — ext/src/now_ext.c:294, :299, :323, each a bare table->arm_active |= … guarded only by request & cap. P3 appears nowhere in that function. Its bit is set in exactly one place, ext/src/now_content.c:1843, reached only under kNowContentVerdictArmed — a per-window, per-A5, TTL-bounded verdict over the content block's arm_window / arm_psn / arm_generation cells. And the only code path in the whole PPC app that ever claims kNowPeekTableCapContent is qdtrace start (now-guest-ppc/src/content/qdtrace_cmd.c:311), released again by qdtrace stop at :354.

So the content plane is not a plane the host switches on. It is a spotlight aimed at one window at a time, and it expires. The resident hooks exactly that identity (content_install_exact_window, ext/src/now_content.c:1859); the host fills display only for windows whose psn:addr collected ops (now-host/Sources/Host/NOWMirrorContentPlane.swift:970-999); and any window with display == nil hatches (mirror/host/MirrorKitUI/SceneRenderer.swift:789-797). With three windows and one armed target, at best one window can ever have an interior. That is the picture, exactly.

Watched, on a running machine. Paired mirror readings across one qdtrace start on the Finder's "Macintosh HD" window:

cap requested active content row
before arming 511 0 0 inactive
after arming 511 15 15 active-current

requested 15, active 15 is Michelle's own reading, reproduced — and with the arm live the content row is active-current, not Requested. Her Requested is therefore a sample between spotlights, not a plane that never activates: state = requested is the one-pass window where the claim is published and the resident has not yet agreed (now-guest-ppc/src/mirror/mirror_probe.c:245-250).

The three wrong attributions.

  • "Content never activates." False. It activates whenever a window is armed, and it deactivates when that arm expires. There is no sustained "on".
  • Sweep B: "the wire scene has no per-window content field at all, armed or not." False. windows[].display[] is a frozen, encodable IR key (mirror/host/MirrorKit/Sources/MirrorKit/Scene.swift:361, in CodingKeys at :394), and it is populated only by the P3 join. The content: false Sweep B quoted is a plane row, not a window field; no per-window content boolean exists on either side.
  • "whole 0B says nothing ever arrived." False, and it is a pure red herring: wireBytes is the scene document's delta-vs-whole transfer meter (now-host/Sources/Host/Session.swift:1584, hardcoded 0 on the .unchanged path at :1654). The content ops never travel in the scene document at all — they come over the separate qdtrace drain lane and are attached host-side. whole 0B is compatible with a perfectly healthy content plane.

The join still works, and that is where this stops being a P3 bug. The at-arm census reproduces on the current tree, unchanged: 467 records, a hooked offscreen world 0x1f472e60 (2 worldborn, 2 blitsrc on the window port 0x009eab90), and 20 text ops carrying the real filenames10 items, 3.21 GB available, System Folder, Applications (Mac OS 9), Documents, Late Breaking News, Rumpus PRO 2.0, TimBotTu, TBT, TBT-paced-dev, TBT-sndbuf-dev — item for item what the guest's own QMP screendump shows in that window. census: runs 1, examined 94, found 2, hooked 1. The drain is healthy. The loss is entirely between "which windows get armed" and the render.

Watched failing by mutation, and this is the reading that settles it. Arm, then qdtrace stop, reading mirror at each step:

requested active structure semantics content interaction
baseline (scene walk only) 7 7 ACTIVE ACTIVE off ACTIVE
qdtrace start 15 15 ACTIVE ACTIVE ACTIVE ACTIVE
qdtrace stop 7 7 ACTIVE ACTIVE off ACTIVE

Content is the only bit that moves; the other three never flinch. And the baseline is the finding: a normal scene walk arms 7, never 15. 1|2|4 is structure, semantics and interaction — the content bit, 8, is not in it and nothing in the ordinary cycle ever puts it there. So the product's steady state is a content plane that is off, and Michelle's 15/15 was a sample taken while some arm happened to be live.

Why nothing caught it, and the instrument that has to change. tools/local-pair-capture.py — the drive loop's own live pair instrument — never issues qdtrace start. Grep it: no qdtrace, no content. Its warm-up comment at :174-185 says "the planes arm as a RESULT" of a scene walk, which is true of P1/P2/P4 and false of P3 for the reason above. So the one instrument that photographs a live machine captures envelopes in which every window's display is nil by construction, and hatching is its only possible output. Every other recent render claim composed a committed capture onto its own scene, which feeds the renderer content from a file. Neither harness can see a content plane that never delivers — the defect lived in the one place nothing was looking, and it still does.

What is still unknown.

  • Whether the host's cycle arms P3 at all in normal operation, and on which window. NOWMirrorSource.swift:683-698 gates the join on planes.contains(.content), but nothing here established what drives the arm during a live host session. Not measured; the host app was never run in this lane.
  • Whether a multi-window interior is reachable at all without a contract change. qdtrace start takes ONE window; serving three windows means three arms, three censuses (57 ms each here, 68.9/186.5 ms measured previously) and three TTLs — or a new verb. That is a design decision, not a fix.
  • Michelle's resident reported cap 31; this tree's reports cap 511. Her image is older. Bit assignments are append-only and stable since 9d2a7bcd, so this does not change the mechanism, but her exact build was not reproduced.

WAS BROKEN, FIXED 2026-08-07 (see the entry above): a classified control the renderer will not draw, and one that erases the panel (2026-08-07, claude/019-integration-4)

Round 4's LOOK, on the emulator at 03:35–03:45, against the private bake now-stage-int4 (anchor 14624 / wire 14625). Emulator-capture-verified; nothing here ran on the PowerBook.

018-cdef-classify does what it says: 71 of Appearance's 73 controls come back classified by CDEF resource id, including the six-tab strip (kind: tab, value 1, min 1, max 6), the theme list, the horizontal scroll bar and the help button. Date & Time's four group boxes arrive with kind: groupBox and their exact titles — "Use a Network Time Server", with its r. That is a large, real gain in what the scene knows.

And the rendered picture got worse, not better. Two mechanisms, both found by rendering round 3's own Appearance scene through round 4's renderer: it comes out exactly as round 3 drew it, so the renderer did not regress — the SCENE changed underneath it.

  • The root user pane erases the window. Appearance's outermost control is a userPane at (0,0)-(464,330) — the whole content rect — and it is LAST in the control chain. SceneRenderer draws dataBrowser/userPane/imageWell/systemControl as an opaque unavailable plate, so the panel renders as one flat lattice and all six tabs, the theme list and Save Theme… disappear beneath it. When every one of those 73 controls was unknown (round 3) the same window drew its scroll bars and its button. A lane that learned more drew less. The rule this wants is the one 018-render-defects already established for DITL rows — the machine paints its ground and then draws on it — applied to region controls: they are ground, and ground is not drawn last. Not attempted here rather than attempted unwatched.

  • A typed control loses to its own untyped DITL row. Twenty of Date & Time's twenty-one controls share a ref with a dialog item, and every one of those items is knowledge: unknown, kind: null. The control loop skips any control whose ref is in dialogRefs, so the newly classified group box never reaches drawGroup and the untyped row draws in its place. This is why the panel has had no group boxes, no static text and no field values in both rounds — and why 019-charcoal's "Use a Network Time Server" fix cannot be seen in an integrated render even though the string now arrives whole. It was not fixable before this lane, because there was nothing better to prefer; it is fixable now, and the care it needs is the scene-ie-error-alert case the existing suppression was written for.

Both fixed 2026-08-07 on claude/019-sweepb-regressions, together with sweep B's R1 and R2 because all four are the same arbitration — see the entry above. The alert case is asserted rather than assumed: it has the OPPOSITE shape, unknown controls beside known pushButton items, and a mutation letting a derived control win unconditionally fails naming it.

Fixed in the same round: the derived branch was emitting GetControlValue into semantic.value, which on this contract is the control's own WORDS and which every consumer draws as text. Every static field and every user pane was published with the text "0", and Appearance's root pane was captioned 0. The number already rides the control's own value key; the derived branch now says nothing about contents. scene_json_test.c pins it, mutation-verified.

What DID survive integration, watched rather than asserted:

  • Cross-application stacking (019-depth-and-face). The Finder's "Macintosh HD" window renders in front of NOW's sidebar, matching the guest's own screendump. Round 3's render put NOW over it.
  • The panel face (019-depth-and-face). Date & Time's face measures (220,220,220) against the guest's (221,221,221). Round 3's render of the same panel measured (255,255,255).
  • A small unknown reads as texture (018-render-defects). A ~20×20 unknown region in Date & Time draws in the close style (0xC4 dots, 0xB8 edge) and is plainly distinguishable from the quiet lattice beside it in the same frame — the rung-4 defect Michelle filed as "a plate claims something is there".
  • Charcoal (019-charcoal). Every menu title, window title, button and label in the render is Charcoal; pack-2026-08-07b carries strikes 9 through 24.

UNVERIFIED: the Mirror is a module with a detach, and no drive has watched it (2026-08-07, claude/019-embed-mirror)

The Mirror renders inside the host app as a module, alongside the others, and detaches into the window it used to only ever be. One LiveMirrorView over one NOWMirrorSource in both containers, so an attached and a detached view cannot disagree — that is a property of there being no second render path, not of anything being kept in step.

The change is really the LIFECYCLE, and the rendering was nearly free. FitTransform already took both sizes as arguments and LiveMirrorView was already wrapped in a GeometryReader, so a pane and a window are the same code path with a different number. What had to be taken apart is that a window WAS the poll: show() called source.start() and windowWillClose called source.stop(). Running and where-it-is-shown are now two axes, each persisting its own last answer, and the second one cannot touch the first.

Backgrounding does nothing, deliberately, and this is the decision most likely to be revisited by somebody who has not read this. Stopping the poll when the module is not showing sounds thrifty and is not: every now_mirror_* projection and the whole fidelity sweep read the source with no window involved and refuse while it is stopped, so an agent's drive would start refusing because a person clicked Console, with nothing on either machine naming the cause. The rendering cost of backgrounding is already zero — HostRootView is a switch in a ViewBuilder and destroys the pane's view anyway. MirrorContainerTests fails if any view on that path grows an onDisappear.

What was found on the way, and is the most valuable thing here. Making every drawn bitmap nearest-neighbour — which the powers-of-two zoom stops need, or a 1-pixel Platinum rule becomes a smear that no similarity score can see — exposed that BitmapFont drew the glyph sheet at whatever fractional pen its caller passed. A centred window title lands on a half pixel whenever the parities differ. Smoothing had merely blurred that; nearest-neighbour made a glyph sample the sheet cell beside it, and a "c" in a title grew a stray dot. The pen rounds now, which is what QuickDraw does — it has no sub-pixel pen. IslandRenderTests caught it by comparing a render pair byte for byte, which is the only kind of gate that could have.

What is NOT verified, and why — this is the honest half.

  • No drive has been watched, attached or detached. The lane's own VM booted, the guest dialled in, the host owned the wire. The drive was going to go over the agent socket (now_mirror_drive), and there is one agent endpoint per user: FileManager.default.temporaryDirectory resolves through confstr(_CS_DARWIN_USER_TEMP_DIR) and ignores TMPDIR, so NOW_PREFS_SUFFIX cannot separate two hosts' sockets the way it separates their preferences. Another session's host held it, the guard refused correctly (unsafeEndpoint("Another New Old World host owns the local endpoint")), and taking it would have been the 2026-08-06 mistake in a new currency. Nothing here is emulator-verified.
  • Nobody has looked at the pane on screen. Capturing the host's own window needs a Screen Recording grant this session must not give itself, and UI scripting was out of scope by instruction. So "it renders as a module" is tested — the view is composed, the gates pass — and not seen.
  • The zoom stops are proven offline, not on screen. ZoomFidelityTests renders a 1-pixel-striped island at 100% and 200% and asserts the second is the first with every pixel doubled; watched failing by mutation, where a black row came back as 64/255 grey. That is the right measurement and it is not a person looking at a Platinum frame at 200%, which is what the sampling change ultimately asks for.
  • UnknownVisual's tiled shading cannot be told. GraphicsContext.Shading.tiledImage takes no interpolation parameter. The tile now carries .interpolation(.none) on the Image and shouldInterpolate: false on its CGImage, which is the same intent stated twice — but whether the shading honours either is undocumented and untested. It is a 2×2 stipple, so the visible cost is small; it is named here rather than assumed away.
  • A stop longer than ten seconds still costs the anchor lease. P1/P2/P4 are ten-second resident leases that retire by not being renewed, so a Stop→wait→Start makes every act refuse element-not-found until the restart re-arms. That is unchanged behaviour, but a Start/Stop control makes it reachable in two clicks where it used to take a deliberate window close.

Fixed on the way, both of them silent classes rather than symptoms anyone had reported: NOWMirrorContentPlane.disable wiped the content cache from inside its own qdtrace round trip with no generation guard, so a quick Stop→Start cleared the new run's interiors for a cycle and read as a render defect; and NOWMirrorWindow.exportEvidence captured window?.contentView, which is nil while attached and threw emptyFrame — an error naming a blank capture rather than a missing window. It goes through RenderShot now, which is 1:1 guest pixels in either container by construction.

The dated line this supersedes: the "--open-mirror could leave a window over a stopped poll" note below is now a note about a model that no longer exists. An open window has never been a running poll; since this change it is not even claimed to be, and showMirror() starts before it shows — gated, and watched failing by mutation with the two statements swapped.

FIXED: two capabilities and their gate had never met, and one integration found both (2026-08-07, claude/019-integration-3)

scripts/test-all failed at the host gate on the round-3 merge, twice, on the same row and for two different reasons. Neither lane was wrong; both were written against a surface that did not yet contain the other.

One: no argument. MCPClientConformanceTests — landed by 019-conformance — checks its recipe book against tools/list both ways and failed naming now_mirror_open, a capability 018-open-mirror had landed before that lane forked. Fixed with a recipe: it takes no arguments and its own annotations declare it idempotent and non-destructive, so a conformance run may take it. It goes in producersFirst ahead of the other now_mirror_* rows because its own descriptor says they read a state engine that does not run until the Mirror is open.

Two: a fourth name for one question. With an argument, the driver then read its reply as "a structured result in no shape this driver can read: alreadyOpen, detail, showing, unavailable". AgentIntegrationMirrorOpenResult says availability with showing: Bool beside the same unavailable payload the other envelopes carry — which makes four spellings of one question on this surface: available (session health), hostAvailable (guest files), ok, and now showing. The driver's verdict was the correct thing to say about that surface.

Taught to the driver rather than renamed, because renaming a shipped MCP result field is a surface change and not an integration's to make. The consolidation is the open item, and it belongs to plan 019's "one implementation per question": either these four become one envelope, or the reason there are four is written down where a fifth would be added. Until then every new capability with its own availability word fails this gate on arrival — which is the gate working, and is also four times more often than it should have to.

What makes this the round's most valuable finding. Both defects were invisible on both branches and green on both. A gate that samples would have missed the first; a gate that read only the declaration — which is what MCPCoverageTests does, and it was green throughout — would have missed both, because now_mirror_open was declared correctly the whole time. Only speaking to the surface the way a client does found them, and only the merge put the client and the capability in one tree.

NOT MERGED: 018-cdef-classify moved out from under the round-3 integration (2026-08-07, claude/019-integration-3)

tools/arc-status computed the lane as idle at 9 commits and it was merged at that tip (f3f7c7f3 and its ancestors — the reason strings, the 018-control-semantics findings, the scene tests). Three further commits landed on it while this round was gating:

601cb07f  feat(scene): the CDEF resource id, as a third knowledge state
62e82470  feat(act): ctlact takes a POINT, and a console face that works
58e9e5e5  build: the four new sources compile into the guest

Those three are the lane's namesake — the CDEF fallback the merged 018-control-semantics entry recommends as the next slice, plus the click POINT that entry says tabs and list rows need. They are not in this integration. The lane was in flight at the moment it was read as idle, and merging a moving tip would have produced a gate result nobody could attribute — the same class of problem as a derived count that was true when it was derived.

So the merged 018-control-semantics entry's "what did NOT land, and why" section is accurate about this tree and already stale about that branch. The lane lands next round.

The general shape, which is worth more than this instance: idle is a property of a branch at the moment it is computed, and an integration takes an hour. arc-status reports it once, at the start, and nothing re-checks it at the merge — so a lane can be honestly idle when the round begins and in flight when it ends. Re-reading the tips before the final gate would name it, the way re-deriving the coverage tables at the merge names their drift.

FIXED: an integration had truncated the entry it was recording itself in (2026-08-07, claude/019-integration-3)

The 018-integration entry — "the first watch of the INTEGRATED render, and what it shows", six numbered defects across five targets — exists on claude/018-integration and on claude/019-cursor-follow and did not exist on claude/019-integration-2. What survived there was the thirteen-line status addendum a later round had added on top of it, under the same heading, with the body gone. grep -c "Seventeen plan-018 lanes merged" answered 1, 1 and 0.

It was restored at this merge by folding the two: the addendum, then the body it was an addendum to. Nobody deleted it deliberately — a keep-both resolution had kept the newer heading and dropped the older text under it, and because the heading survived, every later reader saw an entry that looked complete.

That is the failure mode of resolving a documentation conflict by keeping headings: a heading is not the entry, and a ledger whose entries can lose their bodies while keeping their titles reads exactly like a ledger that is fine. Round 2 caught its own splice truncating HostAppState.swift from 447 lines to 295 by counting lines; nothing counts the lines of a prose entry, and this one had been short for a round before anyone noticed.

ANSWERED: LIST VIEW IS WATCHED, and the round-2 pair that claimed it was icon view (2026-08-07, claude/019-integration-3)

Two rounds recorded list view as blocked. It is not, and the block named in the entry below — menuact on View > "as List" answering no-such-processdid not reproduce. On the round-3 integrated tree, against a private baked image (resident 4d0988e8e891, capabilities 511), with the Finder fronted and the anchor cycle run first:

menuact menu=259 item=3 titleLeft=110 psn=0.29949953
  Dispatch      dispatched
  Mechanism     the application's own MenuSelect
  Identity      checked: the press is where this application's own
                menu bar puts that menu's title
  Settlement    dispatched-but-unconfirmed
  Correlation   4F976E8D-00000002

and the next scene carried as Icons mark=false / as List mark=true. The guest's own menu state is the confirmation, not a look at the pixels — and the pixels agree: ten rows with Name and Date Modified columns.

The round-2 baseline for this target is void. ~/Lab/Assets/now-mirror-assets/019-integration/finder-list-guest.png shows the Finder in icon view — same window, same ten icons, byte -for-byte the same content as finder-icon-guest.png beside it. The view was never switched and the file was named for what the run meant to capture rather than for what it got. Nothing about that pair can be compared against; the round-3 pair is the first list-view evidence this project has.

That is the second time in three days a target has been named for its intent. It is the same failure as deploying a build under the name it was meant to be (AGENTS.md > Deploying), arriving in the assets directory instead of on the PowerBook.

Not fixed by anything, and that is the useful half. No lane in this merge touched the anchor bind. Either the earlier refusal was a consequence of the rig — a guest that had not run its acquisition cycle, or an application not pumping with the plane armed — or it is intermittent. Both readings are live, and a refusal nobody can reproduce is not a closed defect.

BROKEN: the integrated render at round 3 — six targets, and the content plane reached none of them (2026-08-07, claude/019-integration-3)

Seven lanes merged, then the tree was watched composing. Pairs in ~/Lab/Assets/now-mirror-assets/019-integration-3/ (out of git) (guest pixels via QMP screendump after the walk returned, host render via MirrorApp --render-scene over the same envelope). Emulated mac99/OS 9.1, private image now-stage-int3, resident 4d0988e8e891. Nothing here is metal-verified.

Steady state was confirmed, not assumed. Pass 1 discarded — its scenes carry two to four windows because the control panels were still opening. Pass 2 is the measurement: every scene carries all five windows. Pass 3 re-took two targets and the semantic answer was identical (21 and 73 controls, all knowledge: unknown), which is the test that distinguishes steady state from the warm-up defect: a warm-up scene reports NOW's own nine controls as unknown too, and here they classify correctly every time — button, checkbox, scrollbar, popup, triangle.

Target Against round 2 What it looks like
Workshop unchanged, and the best of the six Layout, text, buttons, window chrome and the disabled/enabled split all right. Sidebar module icons are a generic page glyph rather than each module's own; the sidebar's scrollbar and the preview box's outline are not drawn
Desktop unchanged NOW's window with the Finder frontmost; the render correctly greys every control of the inactive window, which is the part most likely to have been got wrong
Finder icon view unchanged Frame, title, scrollbars and grow box right; the interior is empty and says "Guest content not reported". Identical to the round-2 render
Finder list view BETTER — new ground First ever capture (above). The render draws the Name / Date Modified column headers, which the icon-view render does not — so the renderer does distinguish the two views. The rows themselves are absent, same content gap
Date & Time unchanged Still no group boxes at all. "Current Date", "Current Time", "Time Zone", "Use a Network Time Server" arrive with correct rects and correct titles and no kind, so there is nothing to draw them as. Text fields are empty grey slabs. Buttons, radio buttons and the menu-bar clock group are right
Appearance unchanged, and still the one a person would call broken No tab strip at all, three scrollbars stacked, interior "Guest content not reported"

One cause under four of those rows. content: false on every window of every target, ours included. The content plane is supported: true, requested: false, active: false throughout, so every interior that says "Guest content not reported" is reporting the truth about a plane nothing armed. This rig arms it no more than round 2's did, so the comparison is fair — but four of the six "unchanged" verdicts are one unarmed plane rather than four defects, and a look that arms it would say something none of these six can.

z: 0 is still true, and the render still draws in array order. Every window in all six scenes reports z: 0 except the Finder's Desktop window, which reports 1. The renders happen to look right because the scene's array order is front-first and tracks the front application — so the defect is invisible in exactly the captures that would expose it, and stays invisible until two windows of the SAME process need ordering.

acquired: 0 on every cycle, in all three passes, with count settling at 4 and armed: true. The cycle reports "every application got its turn" and acquires nothing new because the four are already anchored; it is not obviously wrong, and it is not obviously right either, and no run here distinguishes them.

ANSWERED: the list-view lane is still unwatched, and the reason is NOT the one recorded (2026-08-07, claude/019-integration-2)

The 018 integration could not get a list-view pair and recorded that script "answered osaErr -1753 to every AppleScript on that rig, including get name of front window". That is disproven. On this rig, with the integrated tree:

script osaErr output
tell application "Finder" to activate 0
tell application "Finder" to get name of front window 0 "Macintosh HD"
… open file "Date & Time" of folder "Control Panels" of folder "System Folder" of startup disk 0 opened
… open file "Appearance" of … 0 opened
… set current view of front window to list view -1753
… set view of front window to list view -1753
… tell front window to set current view to list view -1753

So the script verb works and the AppleScript path is live. What fails is current view specifically, in all three spellings — that term is Mac OS X Finder vocabulary and Mac OS 9.1's Finder does not carry it. It is a guest-OS limit, not a rig fault and not a defect in anything this project ships.

The product-real route is the act plane, and it refuses for a reason already on this page: menuact on View > "as List" (menu 259, item 3, titleLeft 110, the Finder's own PSN) answers no-such-process — the known anchor-bind failure on the staging path, the same one that makes actselftest refuse on every spun-up clone. Two of menuact's three argument requirements were also found by being refused for them in turn, and both refusals said exactly what was missing, which is the error vocabulary working.

FIXED: four base-image failures in two days, each one layer away from the last fix (2026-08-07, claude/019-base-image-guards)

Michelle, after a coordinator pointed an integration lane at the wrong base image: "not the first time this has happened. we need guards for this." She is right that it recurs, and the recurrence is the finding.

  1. 2026-08-06 — the shared stage image was three days old while six ext/ commits landed that day. Every session cloned it, staged a fresh build into a throwaway clone, and discarded the clone. The rule was already written down and nothing checked it → tools/ext-bake-gate.
  2. 2026-08-06 — three bakes in one night installed images whose HFS volume was still marked mounted; qemu-img check called all three clean, because it answers a different question → tools/volclean.py.
  3. 19 July → 2026-08-07os91-runner.qcow2 sat dirty for nineteen days, opening a modal on every clone. Nobody noticed: the anchor came up at 167 s behind it → the volume check moved into tools/image-provenance (the entry below).
  4. 2026-08-07scripts/spin-up-ppc:89 defaulted BASE to os91-runner.qcow2 rather than the stage image, in one line, with a comment sixty lines further down saying so. Every lane in the arc cloned a 19 July image. That is failure 1 again, one layer down, where no gate was watching.

Each fix guarded the layer that had just failed. Bake was well gated; the clone sites were not gated at all, and "which base" was spread across a shell default, a comment contradicting it, ext/stage-receipts.json and prose in AGENTS.md — which is the shape this project's own rule forbids: state a limit once, where every reader looks.

tools/base-image is that one place. It answers which base for this purpose and is it fit: designated or not, volume clean / dirty / unknown (never folded together), and whether the base's baked resident predates this checkout's ext/. spin-up-ppc, bake-ext-image and q800-68k all consult it before cloning, and spin-up-ppc no longer carries a default of its own.

Warn or refuse is argued per check rather than picked once — the table is in docs/staged-images.md. The short version: a dirty base is a slower boot the caller cannot fix, so it warns; a run that stages no resident, pointed at a base with none in it, produces a wrong result that looks right, so it refuses. NOW_BASE_FORCE=1 with NOW_BASE_FORCE_REASON="…" overrides — a force with no reason is still refused.

Seven guards in tools/image-discipline-tests, each watched failing by mutation. Two of them are the ones meant to make this the last time rather than the fifth: a new clone site that does not consult is named by a scan, and a BASE= default that disagrees with the designated base fails the gate.

Still open, and honestly:

  • No clone site can currently trip a refusal, because all three stage a fresh resident, and that is what makes staleness a warning rather than an error. The refusing rows serve a caller that boots a base as it is — which nothing here does yet — and a person running tools/base-image fit --purpose oracle by hand. The refusal path is watched failing in the gates and end to end through spin-up-ppc with a forced purpose; it is not exercised by ordinary work.
  • Nothing here booted a VM. Tested, not metal-verified, and not even emulator-verified: the guards are watched failing against synthetic HFS images, and the wiring is watched only as far as bash -n, the spin-up-ppc python gates, and the tool's own exit codes. The first real spin-up-ppc run after this lands is the proof, and it should be said out loud when somebody makes it.
  • The designated base for ordinary PPC work is now the stage image rather than the plain runner. That is the right default — it is the one file anybody can account for — but it means ordinary lane work now clones the file --shared bakes over. Somebody should decide whether spin-up-ppc ought to prefer a lane's own private bake when one exists.
  • At the time of writing the shared oracle is UNACCOUNTED (no receipt on any branch claims its bytes) and claude/019-integration-6 is baking a fresh one. Until that lands, every spin-up-ppc run warns about the base's resident — correctly.

FIXED in the tools, REPAIRED but NOT INSTALLED on disk: every spin-up clone booted into Disk First Aid for nineteen days (2026-08-07, claude/019-clean-base-image)

Michelle saw Disk First Aid come up and asked whether anybody was baking dirty images. Nobody was. The bake path is sound and the check it was given works: now-mirror-stage.qcow2 measured CLEAN and matched the newest receipt, and all five per-lane clones under agent-stage/ were clean too.

The dirt was where the gate does not look. ~/Lab/Assets/os91-qemu/os91-runner.qcow2 measured DIRTY — will run Disk First Aid, with a content mtime of 19 July. And scripts/spin-up-ppc:89 defaults BASE to that file, not to the stage image — its own comment says so. So every spin-up clone since 19 July booted into "Your computer did not shut down properly", a modal that sits on the desktop until something dismisses it.

The shape is worth more than the fix. tools/ext-bake-gate requires volumeClean in a receipt, and a receipt describes the baked oracle. os91-runner.qcow2 is never baked, so there was no receipt to require anything of, and no gate was wrong — the gate and the script simply named different files. Nineteen days of green.

Two other bases were also dirty and are recorded here rather than repaired, because nothing in the current arcs clones them: os91-hd.qcow2 and mac99-stack-dev.qcow2.

What changed. tools/image-provenance now asks the volume of whatever image it describes — measured from the file, not read out of a receipt, and reported for every image including the plain unbaked bases no receipt covers. It already ran on the base of every spin-up-ppc and bake-ext-image run; scripts/q800-68k now runs it too, which was the third clone site nobody was inspecting. tools/volclean.py grew a non-printing verdict() so both callers ask one question in one place.

It warns and does not refuse, following ext-bake-gate:326-333's reasoning about the oracle: whether a shared base is clean is not a property of the run about to start, is not in that caller's power to fix, and is not even a wrong result — only a slower boot and a dialog. unknown is kept as its own state, because an image that could not be read is not clean, and folding those together is exactly how qemu-img check came to stand in for a question it cannot answer.

What was measured, and on what. A cp -c copy of the runner image was booted in place on a private lane block (467, anchor 15736). The modal was photographed — "Verification and repairs completed successfully" — dismissed with the worker's key verb, the shutdown applet pushed and tools/shutdown-guest.py run. Result: volclean.py CLEAN, qemu-img check No errors were found. A clone of the repaired copy was then booted end to end on a second port and reached the Finder desktop with no dialog at all, and that clone shut itself down clean as well — the procedure reproduces.

The keyboard route was tried so nobody has to try it again: OS 9's Ctrl-F2 menu access does not work here, because the worker's key verb drops the modifier (it answers mods: 0). The applet is the only route a plain base has, since the Finder's own Shut Down needs NOW's act plane.

STILL OPEN, and it is one command. The repaired image is not installed. Seven QEMU guests were live on this Mac, one of them Michelle's own stack, and the standing rule is not to replace a shared base while any is running. No running VM depends on it as a backing file — tbt_clone_disk is cp -c, a standalone copy, and the one session image that could be read had no backing chain — but the rule is the rule and the cost of waiting is a slower boot.

# when the Mac is free (tools/lane-ports list shows no busy blocks):
cd ~/Lab/Assets/os91-qemu \
  && mv os91-runner.qcow2.clean-20260807 os91-runner.qcow2 \
  && rm -f os91-runner.qcow2.sha256 os91-runner.qcow2.volclean.json \
  && python3 <NOW>/tools/volclean.py os91-runner.qcow2

The original is preserved as os91-runner.qcow2.bak-20260807 (still dirty, deliberately — it is what was found). The parked file is sha256 ba3c9ee1856689610fca4646d51056f250c0bb10c9d7913f698b54d8f778e5a9; the dirty original is f34f7e5df64e09ced96c7968776692bb94d9639ee32a6f51652449e1a9cda776, which is the hash docs/fidelity-sweep-2026-08-07-a.md and several other documents name as their base — those remain correct about the run they describe, and will be describing the pre-repair image.

The repaired image carries the shutdown applet at Macintosh HD:TimBotTu:now-dev:NOW Shut Down. That is deliberate and is the only difference beyond the volume bit: tools/stage-ext.py puts the same file at the same path on every clone anyway, and its presence makes the base repairable next time without a 68K build.

CORRECTED: the extension does not hook every draw at rest — but it did hold a MacTCP stream on machines that never ran NOW (2026-08-07, claude/019-ext-rests-when-unused)

The landing question was whether the resident hooks every app draw and every app state change while nothing is mirroring. The premise is false, and what is true instead was in a different plane. The full census is docs/resident-components.md > "What each plane costs at rest"; this entry is the ledger line.

False: draws. No grafProcs is installed into any port and no QuickDraw trap is patched until P3 is armed for a named A5 and a named window. now_content_boot builds the hook table and installs nothing. A machine that never opens the Mirror never executes a draw hook.

False: app state. Anchors are arm-gated, and the writer lease (kNowPeekWriterLeaseTicks, 180 ticks) force-clears the arm word within three seconds of the application going away — so standing down does not depend on any shutdown path being correct, and a machine whose user never launched NOW has been dark since boot.

True, and nobody named it: P6. A Time Manager task was installed and primed at boot unconditionally, re-firing every five seconds forever; and the first event-loop pass after every boot opened MacTCP's .IPP driver and created a TCP stream with a receive buffer. Neither was arm-gated or lease-gated. The dialling was correctly gated on the application publishing an endpoint — so nothing ever went on the wire — but the driver, the stream and the interrupt were held on behalf of an application nobody had started. Fixed on this branch: both now wait for the endpoint, and the task retires by declining to re-prime.

Two things remain unfixed, both deliberate, and both are the reason this entry is CORRECTED rather than closed.

  • P4's trap patches are one-way. Once the act plane has armed even once, six patches are in the dispatch table until reboot. They are bypassed rather than removed because unpatching from the middle of a chain another extension may have joined is unsafe on this OS. This is the one place the restart-to-apply answer is genuinely forced — and it is forced only for removal, since a disarmed patch already chains straight through and costs a fall-through dispatch, paid only by a machine that has actually used the plane. kNowPeekRestActPatched reports it rather than leaving it to a source comment.
  • P3 holds ~64 KiB of system heap from boot, armed or not, so that arming never has to allocate inside a foreign process. Making it lazy is possible — the first leased filter pass is non-interrupt time, which is exactly where P6 already allocates — but it moves content_block and the capability bit into a two-step discovery the application also has to learn, and 64 KiB is not what the landing question was about. Decided, not forgotten: deferred, and the design is written down in the census section rather than left to be rediscovered.

What is NOT yet known. Emulator evidence only. Nothing here has been watched on the PowerBook, and the resting cost has been established by counter rather than by stopwatch — gne_passes climbing while every plane's counters stay flat proves the filter runs and the planes do not, which is the question that was asked; it does not put a microsecond figure on the filter body itself. A per-pass timing number on a 33 MHz 68030 remains unmeasured, and NOW_METAL is where it would come from.

BROKEN: main never received the gates, and the hooks were dead two ways at once (2026-08-07, claude/019-hooks-diagnosis)

Tested (mutation, four states; no metal). The repair is not applied — it needs a decision only Michelle can make.

Two explanations were in circulation and both were right about a different half, which is why neither fix would have worked alone.

Fault A — the config. 48 of 246 worktrees carry a per-worktree core.hooksPath set to an absolute path ending in now/.githooks in .git/worktrees/<name>/config.worktree, which shadows the shared config's correct relative .githooks. The shared checkout is parked on claude/mirror-subproject (042f41f2, 2026-08-02), which predates .githooks, so that absolute path names a directory that does not exist.

Fault B — the branch, and this is the root cause. .githooks was added in 543b06af (2026-08-06 00:50) and has never been an ancestor of main. 109 of 411 branches carry it; main is not one of them. So a worktree cut off main has no .githooks to point at whatever core.hooksPath says, and tools/hooks-doctor --fix — which only rewrites config — cannot cure it.

main is not the head, and has not been since 2026-08-05 16:51. It is 13 commits ahead of the lane line and 802 behind, and it lacks the entire gate apparatus, not merely the hooks: no tools/ext-bake-gate, no tools/setup-hooks, no tools/hooks-doctor, no tools/receipts-merge-driver, no .gitattributes, no ext/stage-receipts.json, no tools/image-discipline-tests, and a scripts/test-all that does not mention hooks-doctor at all. AGENTS.md says "main is the head — keep the shared checkout on it"; the checkout drifted, and this time the consequence was not a confusing diff but every gate in the repository going quiet.

Blast radius — what has not been refusing anything. .githooks/pre-commit carries the main guardrail and ext-bake-gate check; .githooks/pre-merge-commit carries ext-bake-gate merge-check; .githooks/post-merge carries the stage-image verify-image announcement. In an affected worktree none ran. So: nothing refused a commit on main; a resident commit did not wait for the image that carries it; a TBT_DEFER_EXT_BAKE deferral wrote no record, because the gate that writes it is the gate that never ran; and a fast-forward that moved the stage receipts passed in silence. Fault A dates from the shared checkout being parked (2026-08-02); Fault B from .githooks landing off-main (2026-08-06).

Git says nothing when this happens. Mutation-proved: with core.hooksPath at a non-existent path, git commit on main succeeds with no diagnostic of any kind. A dead hooks path is indistinguishable from having no hooks.

Why nothing caught it. scripts/test-all did report it, at the top of every run and again at the bottom, and four lanes read it and correctly did nothing — the only repair it named rewrites config all 246 worktrees share. A warning everybody sees, believes, and is right not to act on is a broken warning. Worse, hooks-doctor recommended --fix even under Fault B, where --fix cannot work; a repair instruction that cannot work turns a live problem into one somebody believes they already tried. And a lane cut off main never saw the warning at all, because main's test-all does not call hooks-doctor.

Fixed here (this branch): test-all now refuses instead of warning twice, naming the safe action (git -c core.hooksPath=.githooks commit) separately from the one that needs a person, with TBT_ALLOW_UNARMED_HOOKS=1 as the open-decision override; and hooks-doctor distinguishes a config fault from a branch fault and stops offering --fix for the latter.

Still open — the actual repair, in order. Neither step alone suffices:

  1. Land the gate arc on main, so a branch cut from it has .githooks and the tools/ the hooks invoke. This is a change to main and joint with Michelle.
  2. Then tools/hooks-doctor --fix on the shared checkout, to strip the 48 per-worktree overrides, when no other session is mid-commit.
  3. Move the shared checkout off claude/mirror-subproject back to main.

Until then, a lane gates its own commits with git -c core.hooksPath=.githooks commit, which changes nothing shared and works on any branch that carries .githooks.

SWEPT: hardcoded machine facts — the whole class, mapped (2026-08-07, claude/019-no-hardcoded-machine-facts)

Michelle, approving a fix to one instance: "yes this should be dynamic and learned from the actual machine, never hard coded."

A hardcoded machine fact is a value that claims to describe the machine and describes the build, or the developer's desk, instead. It is uniquely nasty here because it is always plausible — right on the machine it was written on, right in every test, and silently wrong the first time somebody runs it somewhere else. No test that runs on one machine can catch it, and every test we have runs on one machine.

The sweep covered both guests, ext/, the host, MirrorKit and the tooling. Four categories, and the third is the largest and the point:

  1. DEFECT — a machine fact stated as a constant where the machine could answer.
  2. DELIBERATE, DECLARED — correct, knowingly, with a comment saying why. Left alone.
  3. DELIBERATE, UNDECLARED — correct today, nothing says so. This is what turns into category 1 the next time somebody edits nearby.
  4. UNKNOWABLE — the machine genuinely cannot answer, so the constant must read as an assumption rather than a fact.

The table

# Site Fact stated Cat Disposition
1 now-guest-ppc/src/commands/commands.c:256 addressing is 32-bit 1 FIXED — asked Gestalt, discarded the answer, printed a literal. Now tests the bit.
2 now-guest-ppc/src/workshop/workshop_window.c:54 this Mac is a PowerBook 1 FIXED — "this PowerBook" → "this Mac", matching all 8 sibling blurbs.
3 tools/stage-ext.py:84 guest volume is Macintosh HD 1 FIXEDNOW_GUEST_EXTENSIONS, matching DEV on the next line.
4 now-guest-ppc/src/peek/peek.c:384 guest OS is "9" 1 NOT OURS — the hello twin. Its own comment says if hello becomes computed this must read the same source, and hello is being computed now. See below.
5 now-guest-68k/src/ui/window.c:92-95 screen ≥ 552×360 1 DECLARED FALSELY — comment claimed a 512×342 fit that the arithmetic denies. Corrected; clamp still open (below).
6 now-guest-68k/src/console/conwin.c:83-86 screen ≥ 520×360 1 Same. Comment's own arithmetic added height to the left edge. Corrected.
7 now-host/Sources/Host/NOWMirrorWindow.swift:30 guest screen is 800×600 1 OPEN — see "three answers" below.
8 mirror/…/MirrorKit/Sources/MirrorKit/ScenePoller.swift:16,27,440 guest screen is 800×600 1 OPEN — same fact, same value.
9 mirror/…/MirrorKitUI/PlatinumTheme.swift:43 guest screen is 1024×768 1 OPEN — same fact, different value.
10 now-host/Sources/Host/Chat/ChatSystemPrompt.swift:97 guest is a 68030 with 8 MB 1 OPEN — told to the model unconditionally; situation() in the same file already composes per-guest from real facts.
11 now-host/…/ChatSystemPrompt.swift:119 guest screen is 640×480 1 OPEN — same fact, a third value.
12 now-host/Sources/Host/DiagnosticsModuleView.swift:228,230 which guest serves a verb 1 OPEN — enablement is derived correctly; only the explanatory prose asserts. The .vprobe arm 6 lines above states the right rule.
13 now-host/Sources/Host/NetworkingModuleView.swift:62 PowerPC guest serves net 1 OPEN — same shape.
14 now-host/…/Projection/{GuestLogTail,MachineFacts,GuestDiagnostics,SoftwareInventory}Projection.swift per-guest command tables 1 OPEN — model-facing prose asserting what the gates derive.
15 now-host/Sources/Host/NOWMirrorSource.swift:2306 System Folder child is named Apple Menu Items 1 OPEN — breaks on a localized System; Finder's apple menu items folder specifier is the ask-the-machine form.
16 mirror/tools/stage-{agent,mirror}.py, mirror/tools/extract-assets/{pull,iconpack}.py, scripts/probes/*.py guest volume / OS-bearing paths 1 OPENscripts/probes/oracles.py:45-57 already demonstrates the fix and says why.
17 tools/fidelity-sweep.py:327 guest is mac99/OS 9.1 4 FIXED — now reads "assumed … (not read from the guest)". Every other field in that provenance block was measured.
18 now-guest-ppc/src/core/prefs.c:177 screen is 8-bit 3 DECLARED — it is a capture policy, not a reading. Comment added, no behaviour change.
19 now-guest-ppc/src/census/census_probes.c:617 CPU is PowerPC 3 DECLARED — true by construction (a CFM/PPC binary cannot load elsewhere), but on a census page.
20 now-guest-ppc/src/core/prefs.c:175 peer is at 10.0.2.2 2 Left — declared twice already (prefs.c:125, prefs.h:107).
21 now-guest-ppc/src/main.c:250, workshop_window.c:143 desktop is 800×600, menu bar 20 3 OPEN, minor — dead fallback branch; GetMainDevice() and GetMBarHeight() would answer, and GetMBarHeight appears nowhere in the tree.
22 now-guest-68k/src/census/vprobe68.c:138 System ≥ 7.0, so Microseconds() exists 2 Left — declared, dated, reasoned, with a "confirm on metal" action. The gate would fail closed on a machine where the trap works.
23 now-guest-68k/src/main.c:214 desktop is the main screen 3 OPEN, minorDragWindow bounded by qd.screenBits.bounds; (*GetGrayRgn())->rgnBBox covers all GDevices.
24 now-guest-68k/now-guest-68k.r:60 machine has 4 MB 4 Left — a SIZE resource is read by the Process Manager before the app runs. Genuinely unanswerable, and declared.
25 now-guest-68k/src/core/screen68.c:239 depth is 1 2 Not a defect — reached only when Color QuickDraw is absent, where screenBits is a BitMap and 1-bit by definition.
26 tools/fakeguest.py:128 a guest's OS and capabilities 2 Not a defect — it is a test double; hardcoding is its job and it cites the source lines it mimics.

One fact, three host-side answers

Rows 7–11 are the sharpest structural finding, and they are one issue: how big is the guest's screen? The host answers 800×600, MirrorKit answers 800×600 in one file and 1024×768 in another, and the chat prompt tells the model 640×480. Each is a plausible default and no two agree.

This is the failure AGENTS.md already names — state a limit once, where both sides read it — arriving as a machine fact rather than a buffer size. The guest sends scene.screen.w/h; fitToGuestScreen() already uses it, once. Nothing was fixed here because the fix is a seam across two packages and belongs to whoever owns the Mirror window's sizing, not to a sweep.

hello.os and its twin (row 4)

peek.c:384 writes "9" into the resident endpoint table, and its comment (peek.c:381-383) says the quiet part: "one string in two places is exactly the drift this project has paid for, so if that one ever becomes computed this must read the same source." hello.os is becoming computed on claude/019-asset-packs. So the condition that comment sets has now been met, and peek.c is the half that will be left behind. Named here rather than fixed because the header it must read from is that lane's, mid-flight.

What is gated, and what is not

MachineFactProbeGateTests (now-host/Tests/HostTests/) fails the build when a guest asks the machine a question and discards the answer — the exact shape of row 1:

if (Gestalt(gestaltAddressingModeAttr, &v) == noErr) {
    add_row(rows, &n, max, "cpu", "Addressing", "32-bit");   /* v unused */
}

It derives both sides from source, keeps no list, and needs no allowlist: a probe that only cares whether a selector exists has two forms already used in this codebase that say so in code — the == noErr ternary and the Boolean assignment — and neither is block form, so neither is examined. The escape hatch cannot be a comment, because the gate reads through GateSource with comments stripped. Mutation-verified: reinstating the bug fails it by file and line. 35 probes found; the second test fails if that count collapses, because a scanner that has gone blind reports every guest clean.

It does not cover the class, and the table above is mostly outside it. It sees only Gestalt, only the == noErr block form, and says nothing about whether any other constant is true — deciding which literals are machine facts is not mechanically decidable. Rows 7–17 are exactly the part no gate catches, which is why they are written down.

Open, in priority order

  1. Rows 7–11: one guest screen size, three host-side answers, none asked.
  2. Row 4: peek.c's OS literal, once guest_identity.h lands.
  3. Rows 12–14: per-guest capability claims surviving in prose the gates no longer make.
  4. Rows 5–6: the 68K windows are placed unclamped and do not fit a 512×342 screen. health_static()->screen_width/height is sampled before the first draw, so the fix is a clamp — but it is a behaviour change nobody has watched on a compact screen, so it is parked rather than done blind. Supported floor today is 640×480.
  5. Rows 15–16: guest volume names and localized System Folder children.

FIXED: an act longer than three seconds switched off the instrument that was to report on it (2026-08-07, claude/019-ctlact-settle)

Sweep C briefed this as a regression — ctlact part 0 on Appearance's tab accepted, click posted, zero pixels — and the regression was not real. Driven on a live guest, the tab switches. The two things that made it look otherwise were both the instrument, and both were caught by re-posing rather than by inspecting harder, which is now the third time that has been the cheaper route:

  • The twenty points. scene.windows[].rect is the STRUCTURE box and elements' bounds is the CONTENT box; Appearance reports t=70 and top=90. A point built from a content-relative control rect and a structure origin aims one title bar high — a tall control absorbs it, a short one does not. (Found independently by claude/019-list-selection.)
  • The covering window. A posted click carries only a POINT, and the TARGET resolves it with its own FindWindow against the machine's window list. At 800×600 NOW's own Workshop window covers Appearance's panel completely, so a press aimed through it belongs to NOW and the panel does nothing with it — accepted, click posted, zero pixels, which is the briefed symptom exactly. tools/local-tab-settlement.py now refuses to press a covered point rather than reporting one.

What was underneath was worse than the briefed defect, and it is two things:

  1. An act that outlives the writer lease disarms every plane beneath itself. Three clocks govern an armed plane — writer lease 180 ticks, owner lease 600, act deadline 300 — and act_yield renewed the owner's (by pumping the wire) and not the writer's. Nothing else renews the writer heartbeat during an act: the main loop is not reached, and the wire's renewal rides an INBOUND message, of which there are none while the host waits for the reply. At t=3 s the resident reads a lapsed writer and sets request = 0. Measured: a 5.1 s ctlact part 0 that genuinely switched the tab came back Re-read value: the anchor plane is absent or not armed. This is the 2026-08-06 owner-lease finding one lease over, and the shorter clock was the one nobody had looked at.
  2. part 0 was judged by a patch it never asked for. It posts a real click and consults no patch on an Appearance-era tab — but it waited the full 300-tick deadline for one anyway, which is where the 5.1 s came from (a hundred times a part 23 scroll, and the very thing that outran the lease above), and then reported Settlement: timed-out regardless. The reply was word for word identical for a tab that switched and a tab that did not. Not a verb claiming success it did not have — a verb that could not tell its two outcomes apart, which is the same defect wearing a milder face.

Both fixed. part 0 now watches THE CONTROL, stopping the moment it moves, and reports the settlement vocabulary — confirmed - the control moved or dispatched-but-unconfirmed — never click posted.

EMULATOR-VERIFIED, both directions, on one guest, with the old and the new build A/B'd on the same machine and the same control (rig: os91-runner clone, this tree's build restaged in place; the A/B is what the accident of staging a stale binary bought). Each row confirmed twice as sweep spec v3 requires — the machine's own re-read value and the rectangle a person would look at, in the guest's own framebuffer:

old build new, tab switched new, same tab pressed
Dispatch click posted confirmed - the control moved dispatched-but-unconfirmed
Settlement timed-out confirmed dispatched-but-unconfirmed
Re-read value the anchor plane is absent or not armed 1 → 4 4 → 4
value, from the scene 4 → 1 1 → 4 4 → 4
strip pixels 1949 / 9320 1949 / 9320 0 / 9320
pane pixels 79146 / 130946 79146 / 130946 0 / 130946
dispatch 6092 ms 721 ms 2066 ms

The first column is the defect in one line: the tab switched and the verb said timed-out. The last two columns are the same verb, the same control and the same point, telling its two outcomes apart — which is the whole of what slice 8 asked for. The 2066 ms in the last column is the 120-tick watch being spent in full, which is the only case that pays it.

Evidence, screendumps and the rig table: docs/local/ctlact-settlement/.

STILL OPEN, from the same audit

  • part 0 arms a patch that would suppress the click it posts. now_act_control_answer returns the armed part_code, so an application that does call TrackControl on that control is answered 0 — "released outside the control" — and does nothing. part 0 is self-defeating on exactly the controls a named part serves, and correct on the ones it does not. The arming lives in the resident, so changing it is a bake; it is at least reported honestly now. Table in mirror-act-plane.md > "Which part codes can verify their own effect".
  • A push button cannot be verified by this guest at all. It has no value range, so part 10 has only the patch firing — which proves the application was asked, not what it did. That is now said out loud (dispatched-but-unconfirmed) rather than reported as dispatched.
  • The part code was never the axis. ctlact branches on part exactly once (0 or not); what decides verifiability is the CONTROL and the APPLICATION. A per-part hunt for missing settlement checks will find one difference and miss the two that matter.

An aside worth someone's time

spin-up-ppc's rule 1 states that "QMP keyboard events never reach this guest". On this run they did: send-key ret dismissed the Disk First Aid dialog on the dirty os91-runner base and the boot continued. Either the rule has aged out (mac99 + usb-kbd) or it was always narrower than written. Worth a measurement, because a great deal of rig design rests on it.

FIXED: one fact about the machine, four answers, none of them asked (2026-08-07, claude/019-one-screen-one-answer)

The guest measures its own screen and sends scene.screen.w/h. Exactly one call site read it. Four other places had decided for themselves:

Where Said
NOWMirrorWindow (host) 800x600
ScenePoller (MirrorKit) 800x600
PlatinumTheme.logicalSize (MirrorKit) 1024x768
ChatSystemPrompt (what the model is told) 640x480

Two more sat in the PowerPC guest as SetRect(&screen, 0, 20, 800, 600) fallbacks for a GetGrayRgn() that returned NULL.

Nothing was visibly wrong, which is exactly the state AGENTS.md's most-repeated rule describes: the control-frame cap lived in prose, in the sender, and as a different number in the receiver's buffer; nothing was wrong until a message grew past the smallest of the three. The two 800x600 defaults were both corrected in practice — the window by fitToGuestScreen on the first scene, the poller by detectScreen() — and the theme's 1024x768 was reachable only from a scene carrying no screen. So three of the four were latent. The fourth was not.

The chat prompt was the live one. A model told the screen is 640x480 while it is 800x600 aims at the wrong place and is confident about it, and the miss reads as an act-plane defect — which this arc has already spent days chasing twice. It was hedged ("the screen may be 640x480"), which is worse rather than better: a hedge is still a number, and the number was wrong on every guest this project has driven this month.

What it is now

Scene.ScreenSize carries the vocabulary — unknown, isKnown, known — and unknown is a state, never a substitute. Consequences, each of them a refusal rather than a default:

  • SceneRenderer.logicalSize is optional. An unmeasured screen draws "Screen size unknown" instead of letterboxing an invented surface.
  • RenderShot.png throws screenUnknown rather than emit a PNG at a size nobody measured — that picture is what an agent reasons about.
  • LiveMirror drops the pointer input entirely when the scale is unknown, because a click mapped through a guessed logical size lands somewhere else on the real machine and nothing about the miss says so.
  • ScenePoller.placeVolumes places nothing without a right edge.
  • The mirror window opens at a host window size that is deliberately not 4:3, so a wrong aspect is visible rather than plausible.
  • The guest measures once, in now-guest-ppc/src/core/screen_bounds.c; scene_collect, main and the Workshop all read it, and the Workshop skips its clamp rather than clamp to nothing.

GuestScreenIsOneAnswerGateTests derives both halves from source and maintains no list. Read its "what this gate does NOT cover" header before trusting it — notably it reads text rather than evaluating it, it gates the host side only, and its prompt half proves the sentence rather than the plumbing behind it.

Two things this arc learned the hard way

  • Its own first cut passed by reading zero files. The gate excluded paths containing /.claude/ to skip worktrees — and this checkout IS a worktree under .claude/worktrees/, so the filter matched every file in the repository. A fifth hardcoded screen size, planted in HostRootView.swift to watch the gate fail, went unnamed. It now filters on the relative path and asserts it scanned more than 200 files before reporting a finding, because a scan that read nothing passes.
  • Zero has to be allowed. ScreenSize.unknown is written down somewhere, so the scan objects only to a plausible size. That is a hole with a name, and it is smaller than the alternative of exempting a file by hand.

Still open, from the same work

  • The scene IR is not in the contract. contract/asyncapi.yaml has no screen schema, no scene object at all; the only machine-readable definition of scene.screen.w/h is MirrorKit's frozen key set in IRSchema.swift. This lane changed no behaviour so it changed no contract, but AGENTS.md's "a behaviour change starts there" cannot apply to a message family the contract has never heard of.
  • docs/scene-deltas.md shows "screen": {"w": 1024, "h": 768} in its worked example — the only place a reader learns the shape, using a size no code now uses. Harmless as prose, and the gate cannot see it.
  • scripts/probes/qmp.py pins the pointer against SCREEN_W, SCREEN_H = 800, 600, and winact-probe.py chooses its window geometry to fit the same screen. Both declare the assumption in a comment, so they are honest rather than wrong — but a probe run against a guest at another resolution produces hops that are all wrong, and nothing checks.
  • The 68K guest builds no scene, so it reports its screen only as census text ("Screen", "800 x 600, 8-bit"). Nothing structured crosses.

FIXED: agents were not power-cutting VMs carelessly — the rig cornered them (2026-08-07, claude/019-shutdown-always-available)

Five of the seven preserved qcow2 images on this Mac have the HFS volume still marked mounted, so every clone of them opens in Disk First Aid. tools/volclean.py and the bake gate were built to find that. Nobody had established why it kept happening, and three bakes in one night plus os91-runner.qcow2 sitting dirty since 19 July is not a run of bad luck.

The mechanism. A lane reported, honestly, that it had broken the "never QMP quit" rule: the lab's tools/shutdown-guest asks the Finder through the anchor's script verb, the anchor refused, and the tool's message ends "nothing here can shut this guest down gracefully." Its two remaining options were a power cut or abandoning a running VM.

Two facts turn that from an edge case into a standing defect:

  • The refusal is not a session's decision and does not fire sometimes — it fires always. The anchor's scope is a static worker.session baked into the image, byte-identical in now-mirror-stage.qcow2 and os91-runner.qcow2: the same 24 verbs, the same policyDigest 328e2ef0…, stamped "owner":"canonical". There is no configuration under which that tool succeeds here.
  • The message is false, and the false half is the expensive half. launch IS in that baked scope — which is why NOW's applet route works — and the Finder route never asks the anchor at all. Two graceful routes were open the whole time.

And our own callers took the dirty route by construction. --wire selects the Finder route, the only one measured to leave a clean volume. Neither tools/lane-ports reclaim nor the stop: recipe scripts/spin-up-ppc prints passed it, while both printed the words guest-clean. scripts/bake-ext-image did pass it and checks volclean afterwards, which is why the current oracle is clean and the older bakes are not — the fix landed on 2026-08-06 in one caller out of three.

What changed. tools/shared-image-guard.py refuses a power cut against a VM writing to a shared image, names the file, and fails closed when it cannot tell. shutdown-guest.py prints its route list before trying anything, ends by naming three options and what each costs, and carries --force as the sanctioned last resort — gated by the guard, stamping the disk .power-cut so the damage cannot be inherited silently. Both other callers now pass --wire. 23 tests, 21 mutations watched fail, no survivors (tools/mirror-gate-tests/test_shutdown_always_available.py).

TESTED, not metal-verified, and one thing is deliberately unverified: no VM was booted for this. It did not need one — the baked scope was read directly out of both images, which is a stronger oracle than one machine's hello, and the guard is driven over a fake QMP socket in its own tests. What that leaves unproven is the end-to-end claim that --force on a real cornered VM behaves as written. Nobody has been cornered since the fix.

Still open, and not this lane's to land: tools/arc-status and docs/arc-coordination.md exist on 18 of the 48 claude/019-* branches and were cited in briefs to lanes that did not have them — the same mistake three times in one day. If a rig tool is meant to be citable, it belongs on the arc base. See docs/staged-images.md > "Which rig tools a lane can actually rely on" for the per-tool counts.

Also unfixed, because it is in the lab checkout and this repo does not edit its instruments: the lab's tools/shutdown-guest still ends with the sentence that corners people. It should say that NOW's tools/shutdown-guest.py --wire is the route on this rig. Someone with standing in that tree should change it.

BROKEN: the machine will not say what a foreign control IS, and that is not a bug we can fix in the walk (2026-08-07, claude/018-control-semantics)

Michelle, driving the integrated build: "a lot of controls (such as scrollbars and tabs) and basically all lists say the guest did not provide complete authoritative semantics. so like lists, scrollbars and tabs render now but they cant be used". Slices 3, 8 and 16 each hit this and none named the cause. This entry names it. Emulator-verified (mac99/OS 9.1, resident active, 3370f72a244e); nothing here is metal-verified.

The proposed mechanism does not exist on our floor. Slice 18's brief proposed GetControlKind. Universal Interfaces Controls.h line 2310 says of it, in Apple's own words: "This function is only available in Mac OS X." CarbonLib 1.6 does not export it and the link fails — which control_kind.h already recorded in 2026-08-03. The working substitute is GetControlData(…, kControlKindTag, …), and the resident already uses exactly that (ext/src/now_semantic.c :: classify_member). Nothing was missing from the implementation.

What the Control Manager actually answers. Measured over steady-state scenes, with a warm-up scene discarded (the first scene of a connection walks before semantics is active and reports every role unknown for a reason that is not this defect):

Window Controls Answer
Appearance 73 62 Unsupported custom control, 5 Semantic classification unavailable, 4 never reached by the drain, 2 classified (listBox)
Date & Time 21 21 Unsupported custom control
NOW's own Workshop 9 9 classified, with actions — button, checkbox, scrollbar, popup, triangle

Unsupported custom control is the resident saying it asked for kControlKindTag and the Control Manager declined, and that the documented kControlListBoxListHandleTag fallback declined too. So the plane is reached, and cannot read them — not unreached. That settles the question the integration round left open, and it means the split is by process ownership only as a consequence: we classify our own controls from the procID control_kind.c recorded when we made them, and for a foreign control there is no such record and the machine will not substitute one.

The reason is that kControlKindTag is answered only for controls made through the Appearance-era Create*Control APIs. OS 9's own control panels are CNTL-resource controls created through GetNewControl, and the Control Manager holds no ControlKind for them. The authoritative answer does not exist for the windows this product most needs it for.

This is also why Date & Time renders with no group boxes. "Current Date", "Current Time", "Time Zone" and "Use a Network Time Server" are four of that panel's 21 refusals. They arrive with correct rects and correct titles and no kind, so the renderer has nothing to draw them as. The missing chrome and the refused act are one defect wearing two costumes.

The three list flavours, distinguished — they are not one thing and a single fix would have been wrong for two of them:

  1. kControlKindListBox — a real control. Classified correctly (the 2 in Appearance). It is refused for a different reason: scene_json.c :: control_action returns no action for listBox, so Semantics.authorizesAction is false and MirrorKit declines. Not a capture defect at all. Selecting a row needs a click POINT, and ctlact presses the control's centre — so this wants a contract argument, not a role.
  2. A bare List Manager ListHandle — not a control. The control walk never sees it and never will; it is not in any window's control chain.
  3. The Finder's own lists — the Finder draws them itself. In this session a Finder window did not reach the scene at all. FinderItems (AppleScript) is the answer for these and is a separate path.

What landed here: an unclassified control now carries the guest's own reason instead of a bare unknown. The guest computed those strings all along and the encoder emitted them only beside a known role — the one case that did not need them. So a refusal by the machine, a miss by the drain, and a question never asked were indistinguishable downstream, which is most of why this took three slices to locate.

What did NOT land, and why: nothing that would make these controls usable. Every route needs a gated surface and none should be taken unattended —

  • Tabs need kControlKindTabs ('tabs', ControlDefinitions.h line 861) added to the resident's compact_kind, which today collapses it and ~18 other documented kinds to OtherSystem. That is ext/, so it triggers a bake. It would not have helped here anyway — Appearance's tab control is among the 73 refusals, so the resident never gets a kind to map.
  • The honest fallback is the control's CDEF resource ID, read via GetResInfo on contrlDefProc — which the walk already reads raw (axwalk.c :: def_proc_origin) and already reports as system for all
  • A system-heap CDEF with a documented id is the machine stating an answer, not us guessing, and it is the only route that reaches OS 9's own panels. It needs its own provenance and a weaker knowledge level so the uncertainty stays visible. This is the recommended next slice.
  • List rows and tab selection need a click point on ctlact. The act cell already carries click_h/click_v, so no peek_table.h change — but the verb's args are contract-declared, and contract changes serialise.

Do not "fix" this by weakening the reporting. completeness: complete beside knowledge: unknown is the walk correctly saying it finished and still does not know. The number gets better by making it know, not by making it stop saying so.

FIXED: one menu act checked its identity, the other did not, and neither checked it against the machine (2026-08-07, claude/019-one-answer-a)

Plan 019 slice 2, F7. now_menu_act requires titleLeft — the x of the named menu's title, which is where the act arms its press — and the requirement was earned: the resident's MenuSelect patch answers a press at ONE point, and an act surface bounded by anything weaker rode a real user's press 18 times in 20. mirror_drive menuItem was read as skipping that check.

It does not skip it, and that was the smaller half of the finding. Both paths end at the same guest verb with a titleLeft; mirror_drive derives it from the published scene's menu.left rather than asking its caller for it. What neither path had is anything checking that the coordinate belonged to the menu the same call NAMED — a required coordinate is not a checked one — and two ways a careful caller supplies a wrong one:

  • A stale scene. The caller read left a second ago; the front application changed its menu bar since.
  • A missing reading. SceneBuilder.normalizeMenus defaults an unreported left to 0, which arms at x=4 — the Apple menu. A mirror_drive menuItem on such a menu arms on the Apple menu's title and answers whoever presses it next. That is the 18/20 hijack reintroduced by a ?? 0.

The probe that reads the item before pressing it (2026-08-07, the disabled-File > Print defect) already walks past the menu row carrying its left, so the machine is now asked. Where the bar is readable its own left is authoritative and a disagreement refuses menu-title-moved BEFORE anything is armed; where it is not, there is no second opinion to have, the press is armed where the caller said, and the reply's Identity row says unchecked rather than implying a check happened. Three answers, three sentences — the third being NOW's own menu bar, which is dispatched through its event loop and arms no press at all.

Driven on an emulated Power Mac G4 (OS 9.1, spin-up-ppc clone, guest build 097462c4d7f1), Finder in front, its menu bar read from a scene (File id 257 at x 38, Special at 218):

  • menuact 257/1 titleLeft=38dispatched, Identity: checked, and untitled folder 1 appeared in the Desktop Folder — the act landed where it was named, witnessed by the machine's own file list rather than by the reply.
  • menuact 257/1 titleLeft=218 (the Special menu's x) → refused menu-title-moved.
  • menuact 257/1 titleLeft=0 (what the ?? 0 produces) → refused.

The decision is a pure function in its own translation unit (act_menu_identity.c) because the state it refuses — a menu bar saying one x while a caller says another — is not one a real Macintosh can be asked to hold still in. Four mutations watched failing there, including the one with no symptom: dropping the title_left_known guard makes an unread menu bar report a claimed 0 as checked.

And the host's ?? 0 is gone too (ruled the same day): the default is removed rather than backstopped. Scene.Menu.left and MirrorObject.Menu.left are optional all the way to the act, so no layer can substitute without seeing the absence, and now_mirror_snapshot omits the key rather than reporting 0 — that being the row a caller reads to fill in titleLeft. Two guards at two layers is not redundancy here; it is defence at the layer that has the information, and the host is the side that knows it never learned the number.

One answer everywhere, because a menu bar is a positional surface: no position means not drawn, not hit tested, not pressed. Giving the renderer a fallback the act path refused would have put the two back into disagreement, which is the same defect one layer up — and it was already there in visible form: an unplaced menu arriving at 0 took the span from 0 to the next title, so a click at x = 4 in the mirror, on the Apple menu's own drawn title, resolved to the unplaced menu instead. The same four pixels the act armed at.

Four mutations watched failing, including the one that makes the other three unreachable — restoring ?? 0 in normalizeMenus, after which nothing downstream can refuse an absence it never receives.

Not driven, and it cannot be from this tree: no producer here omits left, so the absence is unreachable from a live guest. Tested only.

FIXED: two window readers, two rectangles, and windows[].rect meant three things at once (2026-08-07, claude/019-one-answer-a)

Plan 019 slice 2, F2. peek_read.c returned the structure region and axwalk.c the content region, from separate offset tables with opposite failure policies. The 2026-08-02 fix made the bind authoritative after the two disagreed about whether the Finder had any windows at all; it did not merge the readers, so windows[].rect went on having three derivations inside the scene plane alone — content grown up by a constant (bound foreign), peek_read's structure region raw (unbound foreign), and Carbon's structure region (self). One field, three meanings, and which one a row held depended on whether a bind had happened. No consumer could ask.

Both readers now return both regions, each through one region-reading helper so neither can be validated more loosely than the other by accident, and a window missing either is refused whole — the policy each reader already applied to the region it did read. Every branch then derives rect the one way IR v1 defines it.

And the constant survives, deliberately, which is the opposite of what the audit recommended. kNowSceneIRTitleBarHeight is not a measurement of any window's frame; it is a convention the consumer decomposes the same way — MirrorKit's hit tester finds the content origin at rect.t + titleBarHeight — and IR v1's key set is frozen. A producer that started sending the true structure region under rect would be sending something no consumer can take apart, alone, without an IR version to carry it. What the constant is no longer is a substitute for reading the region: both are read from the machine, and the approximation is now checkable against the measurement instead of competing with it.

Watched on the emulator, before and after, same machine and window: NOW's own window reported {l:22, t:48, r:779, b:555} (the true structure region) and now reports {l:28, t:50, r:772, b:548}. Its content top is 70, so the old rect decomposed to 68 — the render was two pixels out on every control in that window — and the new one decomposes exactly.

Owed, and deliberately not added here as its own key. A caller holding a rectangle still cannot ask which one it is, and claude/018-drag-targeting is adding per-item origin provenance (drawn / saved / unknown) for the same class of question. Two provenance schemes for two kinds of rectangle is the defect this slice exists to close, so this is written as a requirement for that vocabulary rather than shipped as a competing field. What a caller has to be able to tell apart, in the order it costs them:

  1. Which REGION this rectangle is. Three answers, and they are not interchangeable: the box (content grown by the IR title-bar constant — what windows[].rect carries, decomposable by the consumer), the content region (what elements/axtree publish under bounds, and what a control's local rect is relative to), and the structure region (the frame a person sees; read from the machine now but published nowhere). Today the same window under two planes gives two rectangles about twenty pixels apart with nothing saying so, and a caller that joined them would be quietly wrong.
  2. Whether the region was READ or DERIVED. The box is arithmetic over a measurement; the content and structure regions are measurements. A caller comparing a rectangle against pixels needs to know it is holding an approximation before it calls a two-pixel difference a rendering bug — which is the mistake this entry's own measurement would have produced.
  3. Which READER answered. axwalk (bound, from the target's own context), peek_read (unbound fallback), or Carbon (self). They agree now; they have disagreed, and when they disagree again the first question will be which one spoke.

(1) is the one a caller cannot work around. (2) and (3) are what makes a future disagreement diagnosable rather than another 2026-08-02. Note that this is a property of a rectangle, not of an item — the drag lane's placed is about a POSITION being invented, this is about a rectangle MEANING something different — so the shared vocabulary probably wants one enum with both concerns spelled out rather than one word reused.

UNVERIFIED-TO-BROKEN: the anchor lease lapses between two calls, watched (2026-08-07)

Plan 019 slice 3 predicted this; it was seen while driving slice 2. On one connection, cycle followed by two identical axsnap calls:

axsnap  ->  bind: ok,       hasWindows: true,  hasMenus: true
axsnap  ->  bind: no-plane, hasWindows: false, hasMenus: false

Seconds apart, nothing else touching the machine. A caller that observes once is told a bound process is unreachable, intermittently — which is the worst shape of false negative, because it teaches the caller to retry blindly. Recorded here rather than chased; it is slice 3's.

FIXED / RE-DIAGNOSED: four of slice 16's five render defects (2026-08-07, claude/018-render-defects)

Worked against the sweep-A capture corpus offline — LiveShapedRenderTests composes each capture onto its OWN scene and writes a PNG, so every defect below was reproduced, changed and re-watched against the guest's own screendump in about a second per pass. Emulator-capture-verified, not metal-verified; nothing here ran on the PowerBook.

  • 1 — Appearance's front tab. FIXED, and it was not the tab pass. rehome assumed every offscreen world is born at (0,0). Appearance's two theme thumbnails are born [36,57,213,182]NewGWorld takes a rect and an application composing a piece of its own window passes that piece's rect in window coordinates — so the origin was counted twice and both worlds, opening white erase included, landed at the content's top-left corner, which is where the Themes and Appearance tabs are. The frame is the world's BIRTH rect and deliberately not its live origin: Sherlock 2 shifts its composite's origin per element and never restores it, and reading the frame off that moves its whole interior.
  • 2 — blank grey plates. RE-DIAGNOSED, then fixed as legibility. They were already rung 4: sampling returns UnknownVisual's exact three colours at exactly 25% stipple. The defect was that rung 4 is invisible at 32×32. The loudness budget now scales with area.
  • 3 — the render printing text the machine truncated. FIXED. The ladder's own rule, implemented at last: only an unjoined blit may be silenced by a semantic rectangle. Two consequences had to move with it — a row that no longer silences must not paint over, asked per piece; and the four background DITL kinds are ground and are drawn BEFORE the replay, not after.
  • 4 — group-box frames "stroked through their labels". RE-DIAGNOSED, NOT FIXED, and it is not a frame defect. The frame is interrupted correctly on both sides. What crosses the label is the label's own last glyph: Chicago standing in for Charcoal is wider, so the title overruns the band the machine sized for Charcoal. The same pack gap as "…Time Serve", in a place that looks nothing like it. Needs an extractor run for Charcoal; no renderer change would be honest.
  • 5 — missing arrows. FIXED in the half that was missing. poly is the arrow family and the replay dropped it silently. Marked as rung 4 at its reported bounds, never as a triangle.

Handed back, capture-side: a window's z is a per-application index. CLOSED 2026-08-07, claude/019-depth-and-face — and z was left meaning exactly what it meant. See "cross-application depth" below.

Named while measuring, not chased: a control panel's content face renders white. CLOSED 2026-08-07, at the measured level — see "the panel face is measured, not asked" below.

EMULATOR-VERIFIED: cross-application depth is watched, not read (2026-08-07, claude/019-depth-and-face)

The Window Manager cannot answer across applications, and that is settled rather than suspected. WindowList is a low-memory global at 0x9D6 that the Process Manager swaps on every context switch, so each application has its own front-to-back chain and none links to another's (LowMem.h:2149; contract/peek_table.h's NowPeekAnchor comment; finding observe-process-local-ui). NOW's whole anchor plane exists because of it. There is nothing to read.

The fallback after the front process was Process Manager enumeration order, which is launch order — measured, not assumed: four captures of the 019 run put the same four background applications in the same order regardless of which had just been fronted.

What replaced it. Classic Mac OS layers by application: fronting one brings its whole layer with it, so cross-application order IS the order applications were last brought forward. That is not readable but it is watchable, and now-guest-ppc/src/scene/front_order.h watches it — one GetFrontProcess per pass of the event loop, sampled there rather than at scene time because a scene sees only where the machine ended up. z is unchanged and still means position within a process; the cross-process order rides where it always did, in the array's sequence.

What is still unknown, and now says so. A process the ledger has never seen fronted has no rank — it was running before NOW started, or has not been forward since. Those go behind everything ranked and the scene carries a new depth coverage claim reading partial. Faceless background applications are excluded from that count, or every scene on every machine would read partial forever.

  • Watched on an emulator, not on metal. Three applications fronted in a known order — New Old World, then the Finder with the startup disk open, then Date & Time — on a mac99/OS 9.1 clone, guest build baeca61f30e9. The guest emitted Date & Time, Macintosh HD, Desktop, New Old World, which is what its own QMP screendump shows, and the host render agrees; before this branch it emitted New Old World ahead of both Finder windows. Pair and scene: ~/Lab/Assets/now-mirror-assets/019-depth-and-face/ (out of git). The PowerBook has not seen it.
  • One machine, one sequence, four applications. The measurement is a single arrangement that used to be drawn wrong and now is not. It does not exercise eviction, a process that quits and relaunches, or an application fronted by a person rather than by AppleScript.
  • The ledger is bounded at 32 and counts its evictions, but nothing puts that count on the wire. A machine that has run more than 32 applications since NOW started has forgotten somebody, and the scene reports that process as never-observed rather than as evicted.
  • Nothing shows a person the depth claim. It reaches an agent through the state projection, which passes coverage through generically. The renderer draws the order either way and says nothing about how much of it is known.

EMULATOR-VERIFIED: the panel face is measured, not asked (2026-08-07, claude/019-depth-and-face)

The renderer filled every non-modal window's content with Platinum.g0. Against the guest's own screendump for the same scene, the Date & Time control panel's interior is 10149-of-11724 pixels at 0xDDDDDD — the Appearance Manager's kThemeBrushDialogBackgroundActive. Everything the semantic pass did not draw over therefore read as unfinished paper.

Fixed by separating two decisions that had been one flag: isDialog answers "does this window have a title bar" and is rightly keyed on the title; the FACE follows the window's owner, so kind == 2 decides it.

It was a measured constant, which is not an answer. CLOSED 2026-08-07, claude/019-theme-colours: the machine is asked now. meta.theme carries dialogBackground, alertBackground, documentBackground, highlight and the screen depth they were asked at, read from the live Appearance Manager once per scene. The renderer resolves them through SceneTheme, which publishes per-colour provenance so a fallback can never be mistaken for an answer. Rule, line and measurements: docs/theme-colours.md.

The machine agreed with the count exactly — #DDDDDD — so nothing moved in the dialog face. One other colour did not agree. Platinum.highlight was 0xCCCCFF, extracted offline from the theme FILE; LMGetHiliteRGB on the running guest answers 0x97A1DE and that guest's own screendump agrees at 2399 of 3240 sampled pixels. It had never been seen on a screen: AccentRampTests says in its own words that no capture in the corpus carries a selected row. That is the same defect this entry names, found in a colour nobody suspected.

Watched in the same run as the depth work above: Date & Time's face is 0xDDDDDD in the guest screendump and in the render beside it, against white in the 018 render of the same three applications (019-depth-and-face/00-before-render-018.png).

  • kind == 2 did not cover every panel. FIXED. windowKind 2000 — application defined, which the Appearance control panel is — was hardcoded to literal white, so no theme could ever move it. It now takes kThemeBrushDocumentWindowBackground, which is the brush that actually names that face.
  • Untitled alerts: MEASURED, and the value is settled. A real Finder alert was raised on the guest (docs/raising-the-unknown-creator-modal.md), screendumped, and its interior counted: 40372 of 45974 px at 0xDDDDDD. So g1 (0xEEEEEE) was wrong and the argued-from-the-name change was right. What a measurement cannot settle is WHICH BRUSH: alert and dialog both evaluate to 0xDDDDDD under Platinum, so no capture distinguishes them until the theme differs or the guest reports the WDEF variant. The renderer uses this side's existing untitled-kind-2 alert verdict and says so at the point of use; see docs/theme-colours.md, "What a measurement could NOT settle".

BROKEN: the first watch of the INTEGRATED render, and what it shows (2026-08-07, claude/018-integration)

Status: list view remains UNWATCHED, not disproved — third pass running. But it is now blocked on one named thing (foreign-process anchor bind on the staging path) rather than on a rig whose AppleScript was believed dead. Closing the anchor bind unblocks it; nothing else needs to change.

For the record, the scene DOES carry what a driver needs: the View menu arrives with all twelve items, correct indices, and mark: true on "as Icons" — so the icon-view pair that WAS captured is confirmed to be icon view by the guest's own menu state rather than by looking at it.

Seventeen plan-018 lanes merged into one tree, then the tree was watched composing real windows — five targets, guest pixels beside the host's own composition path, on emulated mac99/OS 9.1. Pairs and crops: ~/Lab/Assets/now-mirror-assets/018-integration/ (out of git). Nothing here is metal-verified.

The integrated render is better than what sweep A saw, and it is not uniformly better. Date & Time, the Finder in icon view and NOW's own Workshop compose recognisably — every group box framed and labelled, every button titled, radio and check states right, the selected Finder row correctly inverted, the desktop pattern drawn rather than hatched, and no hatching anywhere in five targets. Appearance is worse than recognisable and is the one target a person would call broken.

Six defects, all seen in more than one target unless said otherwise:

  1. THE FRONT TAB IS DESTROYED, and only the front one (Appearance). The guest draws six tabs; the render draws four — Fonts, Desktop, Sound, Options — correctly capped. Where Themes (front) and Appearance should be there are two stray diagonal strokes with no label and no cap. So the procedural tab strip works for every tab whose label box arrived and fails exactly on the one whose geometry differs. crop-appearance-tabs.png. This is the highest-value thing here: the chrome lane's fix survives integration for 4 of 6 tabs and inverts on the front one, which is the tab a person is always looking at.
  2. Icons are drawn as blank grey plates, silently. All nine Finder folder icons, all thirteen of NOW's sidebar icons, and both ? help buttons render as identical featureless squares where the machine drew distinct art (crop-finder-icons.png). The plate is the honest untyped answer for a blit whose pixels never crossed — but a plate claims something is there, and the ladder's own unknown ground (UnknownVisual) is what would say "this is a rectangle nobody can name". Today a missing icon and a real grey square are the same pixels.
  3. Group-box frames are drawn through their own labels. The guest interrupts the frame line behind the label; the render strikes the rule straight through the glyphs of Current Date, Current Time, Time Zone, Use a Network Time Server (crop-datetime-groupbox.png). Same shape as the tab-cap defect — art drawn without the label's clip — and wrong in every window that has a group box.
  4. Static text is truncated 2-4 characters early, consistently: "Set Daylight-Saving Time Automa", "…Time is in effe", "Use a Network Time Serve", "…in the following sect…". The runs are being clipped to a box narrower than the one the machine used.
  5. Framed rectangles the machine drew go missing. NOW's screenshot preview box — a large bordered rect with "No screenshot yet." inside — renders as the text alone, no box, in every capture that contains it. Popup-menu arrows, scroll arrows, stepper arrows and the « back button are absent the same way; the trough and thumb render, the arrows do not.
  6. The render carries text the machine did not draw. NOW's sidebar truncates on the guest ("Capture and stre…"); the render prints it in full ("Capture and stream"). It is reading the semantic title rather than replaying the drawn run — a fidelity divergence in the direction nobody checks, because it looks like an improvement.

Also absent from every render: the desktop's own icons and the Control Strip, both of which the machine draws on all five captures.

Two things a reader should not conclude from this page. The script verb answered osaErr -1753 to every AppleScript tried through it on this rig, including get name of front window, so the Finder's list view could not be reached that way and no list-view pair was captured — the listview lane's work is unwatched, not disproved. And menuact refused with no-such-process against a PSN axtree had just reported, which is the known anchor-bind failure on the staging path rather than a new defect. cycle itself worked and reported honestly (armed, complete, restored, 8 considered, 6 backgroundOnly — the headless lane's classification, live).

BROKEN: a FOREIGN process's controls have no determined kind — 169 of 169, across five targets (2026-08-07, claude/019-integration-2)

The 019 integration merged main and four lanes, then watched the tree compose five real windows on emulated mac99/OS 9.1. Pairs, scenes and renders: ~/Lab/Assets/now-mirror-assets/019-integration/ (out of git). Nothing here is metal-verified.

The split is the finding, and it is perfectly clean. In every one of the five scenes:

window controls roles determined
NOW's own Workshop 9 9 — 3 button, 3 checkbox, scrollbar, popup, triangle
Date & Time 21 0
Appearance 76 0
Finder Macintosh HD 3 0

Our own process resolves richly; every control in every foreign process is role: "unknown", without exception. This is not the semantic plane failing to run — it ran, and it says so:

"role": "unknown",
"semantic": { "completeness": "complete", "knowledge": "unknown",
              "definition": "system", "provenance": "guest-control-manager" }

completeness: complete beside knowledge: unknown is the plane reporting honestly that it walked the whole thing and could not classify any of it. The rects and the titles are right — "Current Date", "Set Time Zone…", "On"/"Off" all arrive with correct geometry.

What it costs in pixels, judged on the whole frame:

  • Date & Time has no group boxes. "Current Date", "Current Time", "Time Zone" and "Use a Network Time Server" arrive as controls with the correct enclosing rects and their titles, and render as nothing at all — so the panel is a scatter of buttons on bare grey. The guest draws four engraved, labelled frames.
  • The two date fields lose their steppers, and the two daylight-saving checkboxes render as bare truncated text: "Set Daylight-Saving Time Automa", "Daylight-Saving Time is in effe".
  • Appearance's tab strip is absent, and three scroll bars stack where the theme swatches belong.

This is the same defect mirror-element-coverage.md measured at 62%, seen here at 100% because control panels are entirely foreign. It is worth re-recording because the 100%/0% split by process ownership is sharper than a single percentage: whatever determines a kind is running only for our own process, and the foreign path returns guest-control-manager geometry with no classification step behind it.

And every window reports z = 0

Across all five targets, every real window carries z: 0; only the Finder's Desktop ever gets z: 1. So the renderer has no stacking order and draws in array order. Watched: in the Date & Time pair the guest shows the Finder's Macintosh HD window in front of NOW's Workshop, and the render draws the Workshop's sidebar over the top of it instead.

What this pass did NOT arm, and therefore does not judge

The capture drove the wire directly rather than through the host app, and a scene.request armed structure, semantics and interaction — not content. So every content-plane observation is out of scope here and is NOT evidence against the 018 pairs, which came through the app: the tab strip, the theme swatches, the field values, the Finder's items and the list rows all live on the content plane. The renderer's "Guest content not reported" placeholder in those regions is the product being honest about a plane nobody asked for, and reads as a defect only if you forget which planes you armed.

Two rig facts worth keeping, both of which cost a run:

  • The planes arm as a RESULT of the first scene.request, so the first scene on a connection is walked before semantics is active and comes back with every role unknown — indistinguishable, in the render, from the defect above. requested/active went 0/0 → 7/7 across one request and the body grew 25701 → 42621 bytes. A capture rig needs a warm-up scene it throws away.

2026-08-07, claude/019-instrument-arms-content: "the planes" is three planes, not four. P1, P2 and P4 echo the resident's arm request unconditionally (ext/src/now_ext.c:294,299,323). P3 — content — appears nowhere in that function; its bit is set in exactly one place (ext/src/now_content.c:1843), under a verdict over arm_window / arm_psn / arm_generation. That is what the 7 above is: seven requested with content dark. qdtrace start takes it to 15 and qdtrace stop returns it to 7. So the sentence is right about what it measured and wrong about what it implies, and the same sentence sat in tools/local-pair-capture.py's warm-up comment, where it read as permission not to arm anything else. Both are now corrected; the entry below has the rest. - cycle restores the previously-front application when it finishes (restored: true), which is correct and is not what a walk of one target wants. Front the target with script, then cycle, then walk.

BROKEN: the plain base image has been dirty on disk since 19 July, so every clone of it boots into Disk First Aid (2026-08-07, claude/019-integration-2)

scripts/spin-up-ppc's default base, ~/Lab/Assets/os91-qemu/os91-runner.qcow2, carries an HFS+ volume header with kHFSVolumeUnmountedBit clear. Read out of the image itself rather than inferred from a boot:

image volume attributes clean
os91-runner.qcow2 0x00000000 NO
now-mirror-stage.qcow2 0x00000100 yes

So a fresh, session-private clone shows "Your computer did not shut down properly" on its first boot — before any staging, before any shutdown, before the script has done anything a run could get wrong. The modal dialog then sits over the Finder and the anchor worker never comes up, so the spin-up stalls at "boot a fresh, session-private clone" with no error. It reads exactly like a hung boot.

This is not the failure the spin-up header warns about, and that is the reason to write it down. That header's rule 1 is about a QMP quit being a power cut and dirtying the volume at the cold-reboot step, and anybody meeting this screen will reach for that explanation first and go looking for a shutdown applet that failed. The applet is fine. The bytes were already dirty, and have been since the file's mtime of 2026-07-19T13:52:17 — which means every run off the plain base since then has paid a Disk First Aid pass at boot, silently, including runs whose slowness was noticed and attributed elsewhere.

Workaround, and it is the one this integration used: NOW_SPIN_BASE=$HOME/Lab/Assets/os91-qemu/now-mirror-stage.qcow2, which boots clean. That is a workaround and not the fix — it makes every run carry the Mirror oracle's resident whether or not the run wants it.

The fix is to clean the base once: boot it, let Disk First Aid finish, shut the guest down through tools/shutdown-guest.py, and keep the result. Nobody should do that mid-flight while other lanes are cloning it.

What would have caught it: the readiness check has no assertion about the volume's clean bit, and rule 2g of mirror-drive-loop.md already says in prose that a private clone is not automatically a clean clone. Reading two bytes of the volume header before boot is cheap and would name this in a second instead of presenting as a stall.

EMULATOR-VERIFIED: the drawn cursor follows what we act on, and the documented way to do it does not work (2026-08-07, plan 019, P8)

Watched, cropped and looked at. A screendump pair either side of a 617,443 → 180,160 placement changes 194 pixels in a box spanning both points; the arrow is at 617,443 in the first and at 180,160 in the second, with 617,443 empty. The emulated device's own move — the positive control — changes 194 pixels for the same motion. Reproduced across two cold boots. Rig: private clone of os91-runner.qcow2, anchor 1960 / wire 5510, resident active, caps=511 (bit 8 = kNowPeekTableCapCursor, which no build before today can set), table length 6116. Design and evidence: docs/cursor-follow.md.

The emulator was NOT the obstacle, and that is the durable half

The standing suspicion was that QEMU's pointing device reported over the top of the resident's writes — which would have meant the documented technique was fine and metal was the easy case and the emulator the hard one. It is not so, and the refutation is direct: read from outside the guest with nothing touching the host pointer, MTemp, RawMouse and MouseLocation held our value unchanged across seconds; the machine profile has no tablet at all (-M mac99,via=pmu -device usb-kbd); and every precondition the CrsrNew/CrsrCouple recipe needs was met — CrsrCouple 0xff, CrsrState 0, CrsrObscure 0, CrsrBusy 0 — with CrsrNew reading back consumed. The cursor task ran and did not draw.

So none of what follows is a rig workaround, and metal is not expected to differ.

Three routes, and only the third draws

Route Result
low memory (MTemp/RawMouse/MouseLocation + CrsrNewCrsrCouple) position right, sprite unmoved
CursorDeviceMoveTo, then the cursor task via JCrsrTask (0x08EE) noErr, manager's own CursorData.where reads back exactly the requested point (760,520 — checked from outside), sprite unmoved at 419,333
HideCursor() / ShowCursor() the sprite moves

The first two are kept, because they are what makes the machine agree about where the pointer is; only the third is what a person sees. On Mac OS 9 the Cursor Device Manager owns the position and the blit lives somewhere in the pointing device's own interrupt path that neither the manager's state nor the compatibility vector reaches.

Two defects the DRIVING found that no reading would have

  1. CrsrObscure had to be cleared, and we are the thing that clears it. ObscureCursor is what every text application calls on every keystroke — hide the arrow until the mouse moves — and the device driver clears it on its next report. Without the resident doing the same, P8 drew faithfully into an invisible cursor: route correct, by_device climbing, zero pixels, which is indistinguishable from the plane not working. It would have worked perfectly in an empty Finder and vanished in every application anybody actually drives.
  2. "Somebody else moved the pointer" cannot be asked of RawMouse. Between placements, with nothing holding the globals, RawMouse drifts back to the pointing device's position — so every act after any device motion looked like a person had just touched the mouse and the plane yielded forever (four acts in a row reporting yielded on a machine nobody was sitting at). It asks the manager's own CursorData.where now.

ANSWERED: the guest's cursor SHAPE already tracks our position

SimpleText launched, text area 4,20–619,581, the resident placed the pointer at 311,100 inside it — and the sprite drawn there is an I-beam. Nothing in this plane knows what an I-beam is: SimpleText read GetMouse, was told our point, and chose the cursor for it.

So cursor-shape mirroring needs nothing further from the guest. The remaining work for the deferred "mirror the guest's own cursor" feature is on the HOST — read the current Cursor and draw it — and that is a much smaller thing than it looked.

FIXED, and the cause was not where this section used to say it was (2026-08-07, claude/019-cursor-follows-act)

This heading used to read "a ctlact places the cursor at 0,0 — the act cell's click_h/click_v are zero". That had already been fixed elsewhere (62e82470, the lane that gave ctlact a point), and the symptom it described — the cursor follows a drag and never a click — carried on exactly as before, which is the whole reason it is worth writing down: a stale diagnosis for a live symptom is worse than no diagnosis, because it stops the next person looking.

The act plane's points were correct. P8 declined to use them, always, on every machine, and by two compounding defects in its own yield rule:

  • gLastPlaced is zeroed at boot. The first placement compares the real pointer against 0,0, so any pointer not resting in the very top-left corner reads as a person's hand and the act yields.
  • The yield branch then recorded the point it had just declined to move to. The device still held the pointer's true position; gLastPlaced held a point the device had never been sent to. The two could never agree again, so one yield poisoned the rest of the boot.

Driven before the fix, private bake, pointer parked at 15,15: asked 1, yielded 1, then asked 2, yielded 2 minutes later with the 60-tick courtesy window long expired. cursor at read 311,310 — the target window's exact centre — the whole time. The plumbing was right and the picture was refused.

Both halves are now one pure function each in now_cursor_logic.c (ever_placed is part of the foreign question; gLastPlaced is assigned only on branches that actually moved the device), watched failing by mutation, and driven: asked 1, by_device 1, yielded 0, with the sprite gone from 15,15 (82 pixels) and drawn at 311,310 (24) — as an I-beam, because SimpleText read GetMouse and chose the shape for its own text area.

The lesson worth keeping is about the counters. asked climbing beside yielded climbing was in every report from the first day, and it reads as the plane working and being polite. It was the plane never working at all. A courtesy counter and a failure counter are the same number until something says which one you are looking at.

Two more this found, both fixed here

  • A refused press left the pointer lying. act_post_click placed the cursor before PPostEvent, so a refused queue left the arrow at a point where nothing had happened. It now places after both events are queued — which cannot affect where the click lands, because where is stamped per queue element, and still precedes every tracking loop, because those read the globals only after dequeuing and that cannot happen until the jGNE filter returns.
  • winact select and winact move never reached P8 at all. They call the Window Manager directly and post no click, so nothing ever placed a cursor for them. A move now carries the pointer by the delta the window actually travelled (measured off contRgn either side, not assumed from the request), and places nothing at all when that delta cannot be read rather than inventing a point. Driven: window 4,40 → 200,150 carried the pointer 311,310 → 507,420.

STILL TRUE, and worth a look before the next cursor arc

  • winact on NOW's own window never reaches the resident. It is served in-process — "the Window Manager, called directly in this process" — so no act on NOW's own Workshop exercises any of this, and the pointer does not follow it. That is defensible (the app is not the machine being operated) but it is undeclared, and it means a cursor test that acts on NOW's own window silently proves nothing. It cost an hour here.
  • mouseloc reported the working route as none. quickdraw — the only route that redraws — had no case in the switch and fell into a default that printed the word the header defines as "no vehicle: the sprite is unmoved". Fixed by adding the case and by making an unrecognised route print its number, so the next route added to peek_table.h and forgotten here reads as route 5 rather than as a meaningful state. This is the enumerated-lists-rot class from AGENTS.md in a switch statement, and it has no gate.

NOT WATCHED

  • Metal. Nothing here has run on the PowerBook. The mechanism is an OS fact rather than an emulator one, so the expectation is that it behaves the same; that is an expectation.
  • A real drag. During DragGrayRgn the application never returns to GetNextEvent, so the owed redraw cannot be settled and the sprite is expected to freeze for the length of the gesture and jump at the end. The runs above pressed on controls the application does not track, so the jGNE pass was always available. This is a declared asymmetry, not a measurement.
  • A person's pointer. The yield rule was watched failing by mutation as pure logic; nobody has sat at the machine and fought it.

EMULATOR-VERIFIED: the drag vehicle fires, and the resident lets go by itself (2026-08-07, slice 10 of plan 018)

Driven on a guest, watched, and measured. P7 — a mouse button that stays down across a gesture, and a resident that releases it whether or not anybody asks. Design in docs/mirror-act-plane.md

"Drag"; driver is tools/local-drag-vehicle.py.

Private bake now-stage-drag.qcow2, resident active, capabilities 255 (bit 7 = kNowPeekTableCapDrag, published only when the Time Manager task installed), table length 6072. Both numbers are unique to this build and are the run's requireTheBuildUnderTest().

What was watched:

Claim Evidence
The Time Manager task fires ticks_served 2 → 31 → 59 → 66 across one gesture, ~16 ms apart, which is the cadence it was primed at
A press holds the button State=held, Button=down, and it stayed so across seconds of wire traffic
Motion is applied moves_applied 0 → 1 → 2; at followed 143,175 → 300,200 → 380,260
The dead-man fires without being told idle=60 ticks, nothing sent, the drag ended by itself in ~1.0 s
The cell recovers a fresh press succeeded immediately afterwards
A stale session is refused a move after the release answered conflict
An untrustworthy press point is refused unsupported, naming the missing rectangle

The one thing that does NOT work: the drag is invisible

CrsrNew/CrsrCouple do not move the drawn cursor on QEMU/mac99. Two QMP screendumps either side of a 143,175 → 620,450 drag differ by zero pixels — and moving the emulated pointing device changes 22, so the screendump does capture the sprite and the sprite simply did not follow.

Everything the Toolbox reads did follow. The guest's own mouseloc reported 143,175, then 400,350, then 620,450 — exactly the points the vehicle was given. So the vehicle drives GetMouse and StillDown, which is what a tracking loop consults and therefore what a drag needs; it does not drive the picture.

That split matters in two directions and neither is settled:

  • For the mechanism it is probably fineDragGrayRgn reads the globals, not the sprite.
  • For a person it is not — a drag nobody can see is a drag nobody can verify by looking, and the Mirror is a human-facing product. Why the documented technique does not work here is not known.

Still not proven: whether the Finder tracks it

Nothing in this run aimed at a Finder item, so DragGrayRgn remains untested. The evidence is now favourable — the globals it reads are the globals the vehicle demonstrably sets — but favourable is not measured, and this is the claim the whole design rests on.

Two defects the DRIVING found that no gate did

  1. The v2 identity guard refused every press. Its check switches on operation_kind, and a new op with no case falls through to a refusal. Fixed by kNowPeekActKindDrag, which binds to the session nonce rather than the point.
  2. The press landed at 0,0. The guest's resolver leaves detail.control zeroed for controls the scene walk reports with real bounds. Harmless for ctlact, whose patch answers for the handle; fatal for a drag, where the point is the operation. dragpress now takes the point from the caller and refuses rather than guessing.

That second one is a live finding for the geometry lane, and it is larger than this plane: handle.detail.control and the scene walk disagree about where a control is, and only one of them is right.

And the presentation half was blocked on geometry — that half is CLOSED

Superseded 2026-08-07 by the targeting lane. What it said: slice 10.5's snap-back needs a defensible "home", and placed == true was not a trust signal — only FinderItems.merge set it from the Finder's own drawn box, and that path ran for folder windows only, while SceneBuilder.desktopItems derived it from the saved fdLocation grid and ScenePoller.placeVolumes invented a position and set it true. So "refuse rather than guess" would have refused every desktop drag.

Closed both ways at once:

  • Scene.DesktopItem.origin carries drawn / saved / unknown, and homeIsTrustworthy is drawn alone. Absent means the producer did not say and reads as untrustworthy. placeVolumes still draws the disk it laid out — an absent disk is worse than a displaced one — and now says .unknown over the coordinates it makes up, so no act may aim with them.
  • The desktop is asked the same question every other surface is asked. A desktop clause in FinderItems.windowsScript reads bounds of every item of the desktop and every disk, in global coords, inside the one script call the folder windows already pay for.

Emulator-verified on a private clone (anchor 1920 / wire 5470, build 1a8ffb91f636, OS 9.1 on mac99). Two things the run said that were not predicted:

  • That guest served neither list nor volumes (unknown-command), so without the clause its desktop arrives EMPTY. The saved grid was not a weaker source there; it was no source at all. The Trash is on the desktop and is not in the Desktop Folder, so the catalog path could never have reported it under any circumstances.
  • items of desktop and disks both report Macintosh HD, with the same box, which is what makes "first record wins" inert.

And the live claim was watched rather than reasoned: set position of on a desktop file moved it and bounds of followed — 608,92,640,124 → 120,320 → 240,548 — with a QMP screendump showing the icon in its new place.

FIXED (builds): a self reference described every one of our own controls as {0,0,0,0}, over resolved: true (2026-08-07, slice 10.5 of plan 018)

The drag lane found a press landing at 0,0 and reported it as "the resolver and the scene walk disagree about where a control is". They do not disagree. One of them was never asked.

observe.c :: resolve_self is the path a reference into OUR OWN process takes — no foreign A5 world to aim, the Toolbox answers directly. Its element branch filled control, window, verdict and identity, and never touched out->detail at all, which resolve_kind had already memset to zero. So every self element resolved to a control with an empty title, invisible, disabled, value 0, bounds {0,0,0,0} — beside "resolved": true. The window branch had the same hole for its title.

Two consequences, and the second is why it survived:

  • The handle verb described this application's own controls that way.
  • ctlact computes its press point from that rectangle, so the press landed at 0,0 — and it did not matter, because the act plane's patch answers for the control HANDLE the request names and declines every other. Where the press landed decided nothing. It stops being harmless the moment a caller needs the point itself, which is a drag.

Fixed by filling detail from the live Toolbox (GetControlBounds/GetControlTitle/IsControlVisible/IsControlActive /value/min/max, and GetWTitle for the window). In GLOBAL coordinates, because that is what detail.control already means — the foreign twin now_ax_read_control adds the window's content origin to the local rect it reads, so the same origin is added here. Deliberately NOT scene_self.c's convention, which keeps a control's rect content-relative because IR v1 says so: those are two fields with two contracts, and conflating them is how a click misses by a title bar.

Status: builds. scripts/build-guests compiles it and test-all is green; the Carbon audit reports 0 findings. It is NOT driven, and the reason is worth writing down: on the emulator (build 1b6fbb321684) the elements and observe walks report our own process with bind: no-plane and an empty windows array, so the wire cannot mint a self element reference at all and handle cannot be pointed at one of our own controls from outside. Whatever minted the reference the drag lane used did not come through those two verbs. Reaching this from a host needs either that mint path on the wire or the drag verbs themselves.

UNVERIFIED: targeting and provisional presentation, with nothing to drive them (2026-08-07, slice 10.5 of plan 018)

The host half of the drag is built and gate-green, and not one gesture has reached a guest, because the vehicle above is still unreachable: dragpress / dragmove / dragrelease do not exist on the wire, so ItemDragDriver — the seam the view calls — has no conformer at all. Today the live view answers an item drag with "this mirror cannot hold the mouse button down", which is honest and is not the feature.

2026-08-07, claude/019-drag-live: half of that first sentence has stopped being true, and the half that remains is the whole problem. The two branches met. dragpress / dragmove / dragrelease do exist on the wire on claude/019-integration-5 and every branch below it — contract/asyncapi.yaml:4649-4787, served at now-guest-ppc/src/act/act_cmds.c:1439+ — and the vehicle behind them is emulator-verified. ItemDragDriver still has no conformer, and now there is a second wall behind that one: dragpress names an element reference and a Finder icon does not have one, so writing the conformer would not be enough. Both are set out under "drag isn't working on my build" at the top of this file. This entry is left as written because the shape of it — two ends of a bridge, tested, with no span and nothing red — is the point.

What IS proven: DragTargeting resolves subjects, destinations and intents against the geometry above, including against a recorded real desktop from the run named earlier; ItemDragSession carries the four presentation rules and each was watched failing under a mutation; ProvisionalDragRenderTests reads the pixels back.

What is NOT proven, and needs the vehicle:

  • That a drop lands where the targeting says it will. Every destination is derived from the hit tester, which is well exercised for clicks — and a click is a point while a drop is a point plus a thing being carried, and the Finder decides the second one.
  • That a rearrangement inside a Finder window behaves as one. The intent is computed; nothing has watched an icon shuffle.
  • The snap-back as a person sees it. The state machine returns the ghost to home; whether that reads as "it went back" or as a flicker is a question for a hand on a trackpad.
  • The provisional style at speed. The stipple is anchored to the rectangle so the texture rides with the item; that was verified in a render test at two positions, not at 60 frames a second under a moving pointer.

One gap deliberately left open: a desktop item whose position is .unknown — a disk the Finder would not place — is still clickable at its invented position. Drags refuse it and clicks do not, which is today's behaviour left unchanged rather than a decision. It sends a click to bare desktop, so it deselects rather than doing damage; closing it means teaching the click path the same provenance the drag path now reads.

FIXED: an act could report success it had not verified, and a window could go silent without naming itself (2026-08-07, lane D of plan 018)

Sweep A found as Buttons "dispatched cleanly twice and never produced button view", and scored Appearance DRIVABILITY 0. Driving the machine settled both, and neither was what it looked like.

as Buttons works, and did all along. Emulator-verified on a private clone (anchor 1760, wire 5310, build pinned with --expect-build): all three Finder View items switched the Macintosh HD window, the checkmark moved with each, and the pixels agree. What Sweep A actually met is that the plane could not TELL. Every menu act carried kNowActPostNone, and act_settlement.c skips observation for those, so a menu act could never leave dispatched-but-unconfirmed — the act switched the view and reported unconfirmed with the proof sitting in the very next scene. An observer cannot distinguish "it silently failed" from "the instrument cannot see it", which is why one honest sweep read it as a silent success.

The silent success is real, and it is elsewhere. Reproduced deterministically: pressing the Finder's File > Print with nothing selected returned ok: true, Dispatch: dispatched, and did nothing. MenuSelect returns whatever the plane answers with, so a press on a DISABLED item always "succeeds" and the application's handler ignores it. act_menu_probe.c now reads the item out of the target's own MenuList first and refuses with menu-item-disabled / menu-item-separator / menu-item-absent / menu-absent. A probe that cannot READ never refuses — an unbound process comes back Unknown and the act dispatches as before, because "I could not look" is not "no".

Appearance was never unreadable; the bound was ours. Its window publishes zero controls because its control chain is 73 controls long and kNowSceneWalkMaxControls is 48, so the whole plane is correctly dropped rather than delivered as a prefix. Nothing said so: controls: [] is emitted for a dropped plane and for a genuinely empty window alike, and the only note was a scene-wide sentence naming no window. Windows now carry a walk_verdict and meta.errors names them, distinguishing our bound from a chain that left the readable zones — different errands, and lumping them sent a reader to widen a memory validation that was never the problem.

STILL OPEN, and it is a sizing decision rather than a defect

Appearance remains undrivable. Making it drivable means raising kNowSceneWalkMaxControls (48, per window) past 73 and kNowSceneMaxControls (96, pooled across the scene) past the ~133 a desktop with Appearance, Mouse, a Finder window and NOW open already wants. Measured cost: sizeof(NowSceneControl) is 320 bytes, so NowScene is 147 KB today and every 32 pool slots add 10 KB. That budget is shared by every window in the scene and is pinned by scene_json_test.c :: test_size_against_the_control_cap, so it is a deliberate decision for whoever owns the scene plane, not a number to raise in passing.

PARTLY ANSWERED, and the cap was not the binding constraint on the emulator (2026-08-07, slice 14)

Two of the three numbers above are now derived rather than chosen. kNowSceneWalkMaxControls was 48 in scene_walk.h against a pool of 96 in scene.hone limit in two places with different values, so the smaller bit first and the pool's own headroom was unreachable. The per-window bound is now = kNowSceneMaxControls and clears 73 without any budget being spent. The pool (96) is unchanged: one measured panel is not a distribution, and a number fitted to Appearance alone would fail on the next panel someone opens.

A retracted plane now also reports how long the chain actually was, so the next attempt measures instead of infers. Lane D's verdict said the bound was ours and could not say by how much, which is why establishing 73-against-48 cost an investigation when both figures were in the guest's hand at the moment it gave up. A chain that never terminates is a separate verdict, because reported as "bound" it argues forever for a cap that can never reach it.

The re-measurement did not happen, and the reason is worth more than the number. On a fresh spin-up-ppc clone (build e14125014e17, resident lifecycle: active), with the Appearance control panel visibly open and frontmost — six tabs, theme swatches, confirmed by screendump — the scene reported zero windows for Appearance and ax_oracle_not_found for every other process. The control walk never ran, so the cap was never reached. Arming first and fronting afterwards, which is the documented order, changed nothing.

So on this image the anchor plane fails before the control cap is ever consulted, and raising the cap would have changed nothing observable. Lane D's 73 was measured on a rig where anchors resolved; that condition did not hold here. Whoever finishes the sizing needs claude/018-anchor-acquisition landed first, and should then use the chain-length report rather than repeating the investigation.

One consequence for how this gets re-measured: the console verbs (axtree, observe, elements) do not arm the planes — only the scene.request path in wire.c does. A walk driven from the console reports no-plane for every foreign process, which looks exactly like a dead resident and is not one.

TAKEN: the distribution, and a fourth state for the windows that lost (2026-08-07, plan 019 slice E)

Emulator-verified, on this lane's own VM (block 963, anchor 19704 / wire 19705), a fresh spin-up-ppc clone carrying this tree's build 490165bfb441, then 9a307a3577b3 after the fix below. Nothing here touched the PowerBook. The anchor plane resolved this time, which is exactly the condition the entry above said was missing — so the measurement the reason was written for is now taken.

Nine control panels opened one at a time, the scene walked after each. Controls per window, from the walks in which each was carried:

panel controls
Appearance 73
Memory 44
Mouse 38
General Controls 28
Date & Time 21
Keyboard 19
Monitors (VGA Display) 18
Set Time Zone 10
Extensions Manager 6
NOW's own Workshop 5-9

What the distribution says that Appearance alone could not. The mean is about 29 and the spread is 6 to 73, so a pool sized for the largest panel is sized for one outlier and a pool sized for the mean serves three windows. With ten panels open the scene wants over 250 slots against 96, and every walk in the run spent the pool exactly: 96, 96, 95, 94, 96. The cap is not the binding constraint on any one panel — the per-window bound already clears 73 — it is the SCENE budget, and no single number both fits NowScene in its 64 KB wire ceiling and covers a busy desktop. That is the argument for lazy delivery rather than for a bigger number, and it is now measured rather than asserted.

The fourth state, which is what the measurement actually exposed. A window walked after the pool filled is refused a slot and retracts, so it published controls: [] — identical to a window proven to have none, for a reason that has nothing to do with that window. Five panels read that way in one walk. windows[].controlsState now says which: complete / empty / unknown / notFetched, and only empty is a fact about the machine. Observed live in one scene: Appearance, Date & Time, Memory, Monitors and Sound notFetched; the Finder's desktop and Mouse's second window empty; five windows complete.

A defect this run found that no test would have. The walk verdict is one slot and the dialog-item retraction runs after the control one, so four of those five notFetched windows had their pool-full sentence — and its chain length — overwritten by "dialog item list hit a bound". True, and silent about the thing a reader would act on. The pool-full verdict now survives an item retraction, the same way ControlsAndItemsRetracted already composed the two older ones; watched failing by mutation, and re-confirmed on the machine (Mouse ... chain is 38 where the previous build said nothing).

Still open: the pool is still 96 and lazy DELIVERY does not exist — there is no way to go back and ask for a notFetched window's controls. What exists is the honest name for the gap, so nothing reads a spent pool as an empty panel while the fetch is built.

NOT what Sweep A thought, for Mouse

Mouse scored DRIVABILITY 0 and is addressable: 38 controls each with a reference and a content-local rect, and 38 dialog items carrying kinds and references. What it lacks is NAMES — 37 of 38 control titles empty, every role unknown, and its DITL titles arriving as pointer bytes. That is the pointer-title defect (everywhere, slice 4) plus element-kind coverage, not an addressing problem, and no fix belongs in the act plane.

ANSWERED: a window opened BEFORE arming now composes — the census (2026-08-07, plan 018)

Emulator-verified. QEMU mac99, Mac OS 9.1, guest build 288944aab6b8, wire 5340. Nothing here touched the PowerBook.

The 2026-08-07 fidelity sweep's worst single result was the Finder's whole interior rendering as one "Bitmap unavailable" hatch. Lane A established it was capture-side. This is the mechanism, watched live on both sides of the fix, same rig, same conditions, ten minutes apart:

records window port offscreen world text ops
before 22 1 bits, no blitsrc none (offscreenPorts 0) 1, the menu clock
after 219 1 bits + 1 blitsrc 0x1f472e60, hooked 10, the filenames

The ten are 10 items, 3.21 GB available, System Folder, Applications (Mac OS 9), Documents, Late Breaking News, Rumpus PRO 2.0, TBT, TBT-paced-dev, TBT-sndbuf-dev, TimBotTu — item for item what the guest's own QMP screendump shows in that window.

The trap patch contributed nothing to the fixed run: qdext.born was 0. Not one world was created during the pass. Every one of those 219 records exists because the census found a world that was already there.

What it is: at arm, once per armed identity and BEFORE the redraw arming itself requests, the resident sweeps the armed process's own application zone and hooks every live offscreen graphics world it finds, through exactly the path a birth takes. It emits the same worldBorn record, so the join, the retention and the ladder cannot tell a censused world from a born one and the wire contract gains nothing. ext/src/now_content.c :: content_census_run, verdict in now_content_logic.c :: now_content_census_match.

Cost, measured rather than estimated. Finder 955 KiB / 68.9 ms; Monitors 997 KiB / 186.5 ms. Roughly 70–190 µs per KiB, the spread being how many blocks got past the cheap filter into a dereference (99 against 209). The budget is 4 MiB — about three quarters of a second worst case, once, at a moment that already costs the target a full repaint. The truncation path has never run: no heap measured so far came near the cap.

How it survives a moving heap. Every test is a shape or a discriminator bit, because LockPixels relocates the PixMap record (toolbox-and-gworld.md §6) and any pointer is a snapshot of a block that has moved. The one non-shape is a RecoverHandle liveness gate standing between a match on the bytes a freed block happens to hold and a four-byte write into it — it fired once in each run (unrecoverable: 1, then 2) and cost no coverage either time.

It degrades honestly. No zone, or a zone whose own bounds cannot be believed, increments census_refused and hooks nothing; the product then renders the same honest gaps it does today. There is no new hard dependency and a resident that cannot run it is not a broken product.

The image this was baked into, and what landing still owes

The resident is baked and verified into a private lane image, not the shared oracle:

image ~/Lab/Assets/os91-qemu/agent-stage/now-stage-arm-census.qcow2
sha256 1bd2b5c67d1dc05408451fd5a7bc16f95e2a381959b4c5718d074774d93c4e51
ext digest 3e6296d2192281609092bc8331ae9e50bf4ebb7d
build fingerprint (the guest's own word) 875078a0d995d4b371c4d61737e249c6673ad01b
resident lifecycle active, capabilities 127
qemu-img check / volume / shutdown clean / clean / guest-clean

The shared oracle was not touched: it is still c466baa9a5455c343908e12197d68e57ffc7f07c140276a90c97a5ae2a137d70, byte-identical to what it was before this lane started, and ext/stage-receipts.json is unmodified — a throwaway receipt in that file would claim the oracle contains this resident, which it does not.

So landing this branch still owes the oracle a --shared bake, and that is a decision to announce rather than a step to run. Verified by asking the gate rather than by reasoning about it: with an ext/ path staged, tools/ext-bake-gate refuses, naming the shared image's receipt as being for a different resident. A private bake proves the resident works; only a shared one makes it what everybody else clones.

Still open, from the same run

  • Monitors is a DIFFERENT defect, and the sweep's reading of it was incomplete. Its window port is hooked and emits zero ops of every family over 16 seconds, in count mode and record mode alike — so "zero on its own window port" is not a hooking failure, the application simply does not draw there. The census does reach its heap (found: 3, hooked: 1), but Monitors then disposes that world, and qdext born 12 / died 12 per event-loop pass says why: it is a create-and-destroy-per-pass application, the Sherlock 2 shape, already covered by the birth patch. Its interior is missing because nothing makes it repaint while observed — the escalation problem tools/fidelity-sweep.py already documents at its own force_repaint branch. Not a census gap.
  • A censused world that dies is not re-censused. The row is correctly forgotten when the world goes, and the sweep does not run again for the same armed identity, so a replacement world is caught only by the birth patch. That is the right division for an application that recreates its world every pass, and it is unproven for one that recreates it rarely.
  • The 4 MiB truncation path is untested on a heap large enough to hit it.
  • Nothing here has run on the PowerBook.

FIXED: tabs had no edges, and the reason was that nobody read the drain (2026-08-07)

Appearance and Energy Saver rendered their tab labels on flat grey — no outline, no selected-tab join, no pane edge — reproduced on three VMs and four guest builds (fidelity sweep A verdict 4). It read as an extraction problem, and asset-extraction-offline.md had confirmed there is no tab bitmap anywhere to extract.

The parameters were in a capture this project had already taken. DrawThemeTab paints its label box, strokes three lines across the top and draws the title as ordinary QDPeek ops carrying the machine's own colours; only the slanted end caps go by a route the bottlenecks miss. MirrorKit.DrawnTabStrip reads what arrived — boxes, front tab, pane top, face and bevel colours, titles, and kThemeMetricLargeTabCapsWidth recovered as half the gap between neighbouring tabs — and MirrorKitUI.PlatinumTab draws the caps and the join. Measured against the guest's own screendump: the front tab's left slant lands on the machine's own column on 13 of 19 rows and is one pixel out on the other six.

The method, so the next element costs an hour rather than an afternoon: deriving-a-drawn-procedure.md.

UNVERIFIED: five of the seven tab states have never been seen. Every control panel and application in sweep A was scanned for the tab signature and exactly two have tabs — Appearance and Energy Saver — both with their window ACTIVE and FRONT. So there is no capture anywhere of an inactive tab strip, a pressed tab, an unavailable one, or a small (kThemeSmallTabHeight = 16) tab. DrawnTabStrip accepts 16-high tabs and PlatinumTab is written for them, and neither has ever been drawn against a reference. Same for direction: everything measured is kThemeTabNorth, and a South/East/West strip would be drawn as North and be wrong without refusing. Getting a capture is cheap — open Energy Saver and click another window — and nobody has.

Still open on that panel, and not addressed here: Appearance publishes controls = 0 and dialogItems = 0 — six tabs and not one of them enumerable — so the tabs are now visible and still not addressable, and its DRIVABILITY is unchanged at 0. Deriving Scene.Control rows from the strip (as DrawnCellGrid does for Sherlock, marked drawing-derivation, claiming no action) is the obvious next slice and was deliberately not taken, to keep this one to a single element.

FIXED: every 1-pixel line in every window was two rows of mid grey (2026-08-07)

Found while measuring the tab, and much larger than it. QuickDraw's 1x1 pen inks the PIXEL whose top-left corner is (h,v); Core Graphics strokes a line CENTRED on the coordinate. DisplayReplay stroked at integer coordinates, so every frame, bevel and separator the replay has ever drawn landed half in each of two rows. Measured on the Appearance panel's pane frame: #D9D9D9 where the machine draws #000000, and its bevel rows #EEEEEE/#F7F7F7 against #CCCCCC/#FFFFFF.

DisplayReplay.pixelCentre plus an inset on the frame verb. It moved one measured region's exact-pixel agreement from 55.6 % to 69.5 % on its own.

Why it survived the renderer's whole life: nothing ever compared pixels to the machine at 1:1 and looked at WHERE an edge was. Every fidelity instrument this project has built is a similarity score over a whole window, and a systematic one-pixel offset is exactly the error a similarity score cannot see — it degrades every number a little and names nothing. What found it was cropping thirty rows and printing them as characters.

FIXED: three states wore one error word, and it pinned a health signal at partial forever (2026-08-07)

ax_oracle_not_found was the scene's answer to three different questions, and only one of them was a defect:

  1. a faceless background application — no user interface by its own declaration, so it never could have had an anchor;
  2. an application with a face and nothing open right now — normal, and transient;
  3. an application whose windows exist and which we failed to read.

Six processes on a healthy Mac OS 9.1 boot reported it: Control Strip Extension, DVD AutoLauncher, FBC Indexing Scheduler, Folder Actions, tbt-appe, tbt-worker. An error word for a normal condition is the same defect class as a confidently wrong pixel — an assertion of failure where the honest answer is "this process has no UI by design".

It had a measured consequence. MirrorStateEngine.enrichVisibility required a visibility row for every application in the replica, and the census is the Finder's every application process, which a faceless process is not. So the process-visibility coverage claim could never read complete on a healthy machine, and a health signal that can never read green is the same as no signal. Measured on the emulator (mac99 / OS 9.1, build 0231bd990e2c): 8 processes, of which 2 could ever be covered. The denominator was six rows short and always would be. With the six excluded by their own declaration, the same census settles complete.

The discriminator is the process's own declaration — modeOnlyBackground in its SIZE resource, from ProcessInfoRec.processMode, read in the same breath as its name — never an inference from an empty window list, which cannot tell "has no UI by design" from "we failed to look". It reaches the wire as apps[].backgroundOnly, and the suppression of the error token is narrow: NotFound alone, on a process that declared itself. Ambiguous, Mismatch, Unreadable and NoPlane are real failures whatever the process is.

Three things worth keeping from how it went:

  • The Application menu is NOT a second signal. It was proposed as corroboration; the Process Manager populates it from this same bit, so agreement is arithmetic. Confirmed live: the menu offered exactly the two non-backgroundOnly processes. Also, the Application Switcher is itself a faceless process (seen headless in the first run) — so a design that read the switcher to learn what a process is would have been blind to the thing doing the reading.
  • A true-only key needs a roster-wide claim to mean anything. Sending backgroundOnly only when true made ABSENT mean both "has a face" and "this producer never heard of the question", which forces every such row to unknown — losing exactly the middle state. Sending an explicit false on every row broke the encoder's own 64 KB ceiling (66422 bytes, watched). The answer is one process-kind coverage claim per scene, which is scene.h's existing _present idiom rather than a second mechanism.
  • empty is earned, never assumed. The visibility lane measured a fourth condition that looks identical to state 2: an application acquires an anchor slot only after being frontmost once since NOW armed (451 armed passes, one slot scan). Such a process has not failed to be observed — the filter has never run inside it. It arrives with no windows and ax_oracle_not_found, so the classifier consults the ERROR first and the window count second, and it lands in unknown where it belongs.

Still open, and not this slice's to fix: the Finder read ax_oracle_not_found on that same emulator run with no planes armed. That is the anchor-acquisition condition above, owned by its own lane.

FIXED: nothing could open the Mirror in an already-running host (2026-08-06, night)

--open-mirror covers a launch and a click on the Mirror page covers a person sitting at the modern Mac. Everyone else had no route at all: an agent on the socket, and the person sitting at the CLASSIC Mac, whose screen the Mirror is a rendering OF.

The cost was not a missing feature, it was a documented bad habit. A sweep that could not open the Mirror from the agent socket fell back to macOS accessibility scripting to click the button, and wrote that up as a rig fact; later sessions copied the pattern and started taking computer-use on the desk's actual desktop mid-session. A missing affordance became a rig technique.

Closed with three faces over one implementation (HostAppState.showMirrorWindow, which ends at the same NOWMirrorWindow.show a click performs):

  • now_mirror_open on the agent surface — the only projection there whose whole effect is on the modern machine.
  • Window > Show Mirror (Cmd-Shift-M) in the host app.
  • A "Show Mirror on Host" button on the guest's Mirror page, with showmirror as its typed face, over a new additive contract family (host.show / host.shown).

Two things worth keeping:

  • Already open is a success, not an error. The window is raised and the answer says which it was. The asker wanted the Mirror in front of them; it is.
  • No Mac connected is a REFUSAL, not an empty window. A Mirror with nothing behind it publishes an empty state that no call of its own can get out of — the same trap --open-mirror fell into above.

Driven and watched, 2026-08-06 (emulator). Against a host launched WITHOUT --open-mirror and a PowerPC guest on mac99/OS 9.1 dialling it:

  • mirror_read --intention status answered "The selected guest has no published Mirror snapshot" — the window shut, the poll stopped.
  • The guest's button opened it. The host's own sentence — "Opened the Mirror on Power Mac G4." — appeared on the guest's Mirror page, and the resident's plane bits went from requested 0, active 0 to requested 15, active 15.
  • On a restarted host, now_mirror_open over the agent socket answered showing: true, alreadyOpen: false and the same status read came back with a live snapshot. Asking again answered alreadyOpen: true.
  • Typed at the machine, showmirror answered "The Mirror was already open; brought it to the front."

The menu item is TESTED, not driven. Clicking it means scripting macOS accessibility, which is the exact habit this entry exists to make unnecessary; MainMenuTests asserts the item, its selector and its shortcut instead.

What is left uneven, declared rather than left silent: NOW-68K neither asks nor serves the family. That is the arc's scope, not a limit of a 68030 — a NOW-68K with a Mirror page would want the same button. And there is deliberately no verb the other way: the host drives the guest's windows through the act plane, and a guest.show would be a second, weaker route into it.

FIXED: one question, five answers — the process family, the scene and the front process (2026-08-07, claude/019-one-answer-b)

Plan 019 slice 2, findings F3, F5 and F6 of the surface audit.

Five independent GetNextProcess walks answered "what is running and which of it is faceless" — process.list, the ps console command, the Processes page, the scene plane and observe — and the modeOnlyBackground classification was copy-pasted into four of them. commands.c admitted it in a comment: "the same test serve_process_list makes." The scene read the bit in none of them, so at one instant process.list said kind: background about a process the scene called ax_oracle_not_found. Both were honest. A caller could not tell, and a driving agent that reads one answer and acts on another is the failure this exists to prevent.

now-guest-ppc/src/processes/proc_roster.{h,c} is now the one walk and the one classifier. It is an iterator, not a table: a table of every process is a kilobyte of somebody's stack on a 56 MB machine, and the thing worth sharing is the sampling discipline. now_proc_roster_begin samples GetFrontProcess and GetCurrentProcess once, before the first row — so "one front process per reply" holds by construction rather than by care. observe.c's bind_target read the front per process (F5) and could emit a reply in which two rows carried front: true; it now has no such call to make.

The headless lane's vocabulary is preserved and was the point: the kind is the process's own SIZE declaration, never inferred from an empty window list, and the scene's partial coverage is derived from the roster's unreadable count rather than a flag each walk sets for itself.

F6 — three fronting implementations, three claims. SetFrontProcess returning noErr means the switch was scheduled. The console front confirmed with its own loop; mach activate confirmed with a second copy of that loop; the wire's process.front answered ok:true and never looked — and now_bring_to_front over MCP rides that exact path, so an agent got the weakest of three claims with no way to tell; the anchor cycle called SetFrontProcess raw and counted accepted requests in a field its own header documents as "applications actually brought forward". now_proc_front_confirm is that ask-and-re-read, once, and every caller now reports what happened.

DRIVEN, on an emulated mac99/OS 9.1 clone of the stage oracle (lane block 502, anchor 16016 / wire 16017, guest build d4f42b3753e7), not merely tested:

  • Console ps, wire process.list and the scene plane agree on all 8 rows of one machine in one connection — including the six faceless ones, which the scene now reports as backgroundOnly: true with window coverage unavailable, reason: no-ui. The scene's process-kind coverage reads complete.
  • The one remaining ax_oracle_not_found in that scene is the Finder — a faced process whose anchor genuinely could not be read on a VM with no armed plane. That is the word doing its job: an error about a failure, not about a normal condition.
  • observe returned 8 rows with exactly one front: true.
  • process.front on the Finder answered ok:true, and a process.list taken immediately after named the Finder as front. The verb's claim was checked against the machine rather than believed.

What is NOT verified, and matters. The confirm's accepted-but-unconfirmed branch was never exercised — asking to front a faceless process on this machine took the SetFrontProcess refusal path instead, and a switch that is accepted and then does not land could not be staged deliberately. So that branch builds, reads correctly and has never run. Nothing here is metal-verified.

The contract change WAS made, under an explicit ruling, after this entry first recorded it as deferred. ProcessResult gains outcome, in ActSettlement.status's vocabulary — borrowed rather than invented, because two honest-but-different words for "did it happen" would be a fresh instance of the defect this slice closed. ok now means the effect this verb can establish, never a receipt; and rather than silently redefining it, the difference between the two verbs is declared:

  • process.front can be told more, by re-reading GetFrontProcess, so its ok is true only when the target is actually frontmost.
  • process.quit cannot — a quit Apple Event is one an application may decline or sit on behind a Save dialog — so its ok means the event was delivered and its outcome says dispatched-but-unconfirmed.

That split is the project's own established position rather than a new one: the host had already worked around the missing field in three places, and BringToFrontProjection says outright that process.result "has no field that could carry 'and it landed'" while AgentIntegrationProcessControl records that "quit stops at 'the request was sent' because nothing on this platform can tell it more, and front CAN be told more."

outcome is optional, and its absence is not unknown — it means the sender does not report outcomes. NOW-68K emits none, so the host's confirming re-list stays as the fallback rather than being deleted as redundant. That asymmetry is declared in the schema, not left to be discovered.

Driven, all five shapes, on the same rig: front on the Finder → ok:true, outcome: confirmed, with an independent process.list naming the Finder as front; front on a faceless process → ok:false, outcome: refused, front unmoved; quit on NOW itself and on a dead PSN → outcome: refused (never reached the machine); quit on a background process → ok:true, outcome: dispatched-but-unconfirmed, and it had in fact gone three seconds later — which is exactly the point, because the verb declined to claim what it had not established.

ProcessRosterSingleSourceTests keeps it closed and maintains no list: it walks every .c under now-guest-ppc/src at test time and derives both sets from what it finds, so a sixth walk in a new directory fails the same day it is written. Its five substantive rules were each watched failing by mutation, and four of them failed on their first run against real code — naming main.c, software.c, proc_actions.c, anchor_cycle.c and processes_layout.c, all now closed. software.c's Finder lookup matched the 'MACS' creator alone, so a Finder identified by its 'FNDR' type — which every other reader in this guest accepts — was not the Finder there.

Removed on the way: proc_kind_text, which had a native test and no caller in the product — a fourth opinion on what the Finder is, held only by its own test. That is the shape contract-coverage warns about: coverage that proves nothing.

A second finding, from being the port scheme's first real user. Two tools/lane-ports frictions, both fixed here and both driven: scripts/spin-up-ppc required NOW_LAB_ROOT by hand because every copy of the lab lookup assumed a worktree lives inside its checkout — an agent lane under /private/tmp has no shared ancestor with the lab at all, so the walk-up reaches / and finds nothing; git's --git-common-dir answers it. And lane-ports reclaim died on a missing import while printing "clean shutdown FAILED and the VM is left up … re-run with --power-cut"a host-side setup error wearing a guest failure's words, one step from a dirty volume, in a repo that has already paid for getting shutdown wrong once. A rig that cannot start now exits 3 and says the machine was never asked and must not be power-cut. The lookup was copy-pasted into twelve rig tools in three shapes, one hardcoding a specific person's home directory; tools/lab_root.py is its one home and the two tools on the lane path use it. The other ten are experiment scripts off that path, left unchanged rather than edited unrun.

Left for later, named rather than silently skipped: F8 (sw sweeps live where software.list pages a cache that cannot say how old it is) and F12 (eight hand-rolled frame codecs). F4 was already closed in claude/018-integration before this lane began — scene_walk.c now reads "THE LIVE CONTROL WINS, ALWAYS", with Mail's Internet-setup alert as its worked example — so the audit's F4 text is stale rather than outstanding.

FIXED: rig ports were assigned by hand, and three lanes paid for it in one day (2026-08-07)

A dozen parallel lanes each wanted a guest VM and a host app, and the ports came one pair at a time from whichever session was coordinating — 1700/5250, 1710/5260, up to 1890/5440. It held only for as long as one session held every allocation in its head, and on 2026-08-06 it failed three ways: a lane could not stage because 1840 was held by a VM orphaned by a host crash and lost its run; HostMachineGuardTests went red for lane after lane over a neighbour's QEMU; and an agent then re-diagnosed that three separate ways, because a busy machine looks exactly like a defect in your own change.

tools/lane-ports derives a block of eight ports from the worktree path — hash first (so a wiped registry moves nobody), O_EXCL claim file second (so a collision is stepped past rather than shared). scripts/spin-up-ppc defaults to it and records the VM's QMP socket path before it boots. Both machine guards now say whose a held port is — your own block, another lane's, or unattributable — and none of them was made more permissive. docs/lane-ports.md is the whole scheme and the alternatives weighed against it.

Still open, and named here rather than fixed:

  • /private/tmp/nowvm-* from before the scheme. 4.7 GB of run directories on 2026-08-07, two of them with a live QEMU holding the disk. lane-ports gc reaps the ones a claim records and reports the rest, deliberately: an idle run directory between two boots looks exactly like an abandoned one, and deleting somebody's mid-flight clone to reclaim disk would be the 2026-08-03 mistake in a different currency. Somebody has to ask, once.
  • The host suite still shares one log file. HostLog.openFile writes ~/Library/Logs/now-logs/<stamp>.log, and the stamp is not per-process, so two concurrent swift test runs interleave into one file — HostLogTests, LoggingSpecTests and HostProjectionAuditTests fail transiently when two lanes overlap. Not touched here: it is product code, not rig, and the fix (a per-process suffix, or an env override the suite sets) wants its own change.
  • This machine has no .env.lab in now/ itself, only inside a sibling worktree — so every lane's scripts/build-guests still skips. The lookup is fixed (a worktree now uses the main worktree's file, and with one in place this worktree went from 6 SKIPPED to 6 ok in 7.8 s), but the file has to be put there once, by hand, by somebody who owns the desk.

FIXED: the MCP surface advertised 41 tools, answered a batch, and could not reach the machine with seven of them (2026-08-07)

Two defects, found one behind the other, and the second was found only because the first was being driven rather than tested.

Seven tools were advertised and dead. SocketAgentIntegrationClient — the client every MCP call travels through — never overrode observeElements, mirrorDrive or tailGuestLog, so all three landed on the protocol default in AgentIntegrationClient and answered "no lane" from a perfectly healthy host. The blast radius was larger than three rows: ObserveElementsProjection is the ONLY producer of the now-element-… references now_window_act, now_control_act, now_text_get and now_text_set take, so those four were unreachable for want of a legal argument.

Two of the three were four-line forwarders over lanes the socket had been serving the whole time. The walk was not. There was no observe_elements operation on the agent socket at all — no adapter method, no arm in the host's switch — so the audit's consolation that "all of it works through tools/now-agent" did not hold for that one: that tool speaks the same socket. The act plane's four addressed rows had had no argument producer on any face of this host since they landed on 2026-07-31.

And the transport answered nothing to a real client. The stdio loop read FileHandle.standardInput.readData(ofLength: 4096), which on Darwin blocks until it has the full count or the pipe closes. An MCP client holds stdio open for the session and sends one small line at a time, so the loop sat on a 76-byte initialize waiting for 4020 bytes that were never coming — and answered nothing, on all forty-one tools. Measured: one small line with stdin held open gets no reply in ten seconds; the same line padded to exactly 4096 bytes is answered immediately. It survived because every driver this binary ever had wrote its whole script and closed stdin, which is what makes the blocking read return. A pipeline that closes the pipe is a batch, not a client, and the surface passed on batches for months.

The shape both share, and the reason they are one entry. No test asserted that a registered projection's client method is overridden, and no test ever spoke to the transport the way a client does. Every gate that existed was pointed one layer inside the defect: MCPCoverageTests checks the catalog against the contract, HostProjectionRegistryTests that rows are registered, NOWAgentCompanionTests that tools are listed and bounded. All were green throughout. A table of what a surface declares is not evidence that the surface answers.

Two gates now, both derived rather than enumerated: SocketClientForwardingTests reads the projections' client.<method>( calls and the client protocol's requirements out of source and fails on any name the socket client does not declare a func for; StdioTransportLivenessTests spawns the real executable, writes one small line and holds stdin open. Both were watched failing — the first named all three missing lanes before the fix, the second by mutation.

Emulator-verified, not metal-verified. All seven were driven through the companion binary against Mac OS 9 under QEMU on 2026-08-07 and four were watched taking effect: the walk minted references for a Finder window's scrollbars, now_window_act moved that window to exactly the coordinates it was given, now_control_act dispatched against one of those scrollbars, now_mirror_drive zoomed it. now_text_get and now_text_set reached the guest and were refused in the guest's own words — "that reference names a control, not a text element" — which proves the reference vocabulary crosses but is not a completed reading. Nothing has run on real hardware.

Closed the same day, on claude/019-conformance

All three of the below were left open above and were taken up by the conformance lane. The entries stay because what each turned out to BE is worth more than the fix.

  • A completed text reading has now been taken through MCP. It needs a dialog's TEHandle, so SimpleText's Find dialog was opened; now_observe_elements minted the reference under windows[].text.ref, now_text_get answered completed with text: "New Old World", now_text_set wrote "read through MCP", and a second read returned it. Emulator-verified, guest build 20ba2e29bff1.
  • The lease was not the defect. The walk claimed no plane at all. observe / elements / axtree read every foreign process through the anchor plane and never claimed it, so they worked only while somebody else held it up — the scene, whose ten-second lease is renewed by host traffic only once a scene.request has been served on the link, or the Processes page while it is visible. A headless MCP client therefore had no claim and got no-plane for everything, for ever; a client with a Mirror polling beside it got the intermittent form that was reported. Fixed by giving the walk the scene's own claim-then-settle pattern under its own owner. Measured A/B on one clone: pre-fix, 8 of 8 processes no-plane and requested still 0x0 after the walk; post-fix, 0 of 8 and requested=0x3 active=0x3 inside the same call. The lease still lapses after ten seconds of silence, and a caller can no longer observe it — the next walk re-arms and waits for the resident's echo before reading.
  • The refusal vocabulary was wired at two of four sites, not two of two. reach is derived from the guest's own reply (a refusal with no correlation was never registered, so nothing was armed). dispatch and the elements walk derived it; textget and key hardcoded unknown and dropped the correlation and settlement the derivation rests on. Both fixed, and a gate now reads the source and derives both sets from it — every Self.failure(…) built from a result.error, against every one that asks reach(ofGuestRefusal:) — in both directions, because otherwise it is satisfiable by pasting the derivation onto local refusals and turning every provable notSent into an unknown.

Still open, noticed in passing and not chased

Three answers the first full live conformance run surfaced. All are answers a healthy host gave that contradict the machine, which is the class the surface work is for; none was chased.

  • now_mirror_lifecycle refused with "No Mac is connected, so no resident has answered" while a Mac was connected and eighteen other tools were serving from it in the same run.
  • now_guest_files_upload_begin refused a four-byte upload with "Private staging cannot reserve the declared upload".
  • now_reveal_item refused "nothing named System Folder to reveal" on a Mac OS 9 volume that has one.

And two structural gaps the same run named:

  • now_transfer_approved_artifact is unreachable from MCP. Its argument is minted by a person approving a transfer in the host UI, and no tool on this surface produces one. It is the conformance run's only uncovered row.
  • The reference walk's frame budget is spent on our own process. observe --scope all came back truncated: true after ONE process on a machine running eight, because NOW's own window carries the most controls — so the walk cannot reach the Finder unless it is aimed by serial. The least interesting process is the one that fills the reply.

BUILT AND NOT ARMED: the derived-document merge gate (2026-08-07)

tools/derived-doc-gate exists, is tested by mutation (tools/derived-doc-gate-selftest, 28 assertions including a real two-lane merge and the hooks driven through git commit), and is wired into both merge paths — .githooks/pre-commit for a conflicted merge and .githooks/pre-merge-commit for an automatic one. It is inert. It runs only with NOW_DERIVED_DOC_GATE=1, because arming a gate into a live fleet is a change to every branch at once, .githooks does not exist on main, and one lane of the 019 arc was still running when this landed.

tools/gate-impact-sweep (generalised in the same commit to answer for a gate a branch does not carry at all) was run over all 44 live claude/01* branches: none is refused, and none of them touches any declared source, so arming it strands nothing that exists today. That claim has a date on it and nothing else.

What is unverified: it has never refused a real merge in this repository, because no merge here has yet been run with it armed. Its refusals have only been watched in the selftest's throwaway repositories. And it is not in scripts/test-all — deliberately, because adding a stage is also a fleet-wide change, but that means the selftest is a gate nobody runs, which is the exact hole scripts/test-mirrorkit was added to close. Step 4 of arming is to add it.

What it cannot check, and the cure is a person rather than a flag: it cannot tell a re-derivation from a convincing hand edit, and it cannot know that a derivation asks the right question.

FIXED: the live render was worse than every fixture render, and no gate could see it (2026-08-06, evening)

Reported against a running session: the Mirror "looks largely regressed" while the same captures rendered whole in the harness. It was not a regression and not the asset pack, which resolves in the app process (no MirrorKit: warning on stderr with the Mirror open and drawing). It was that the harness and the app were never drawing the same scene. NOWMirrorContentCoverageTests composes every capture onto a canned scene with controls = [] and no dialogItems, so every fixture render is drawn with an EMPTY semantic-exclusion list; the app draws the same capture with one entry per DITL row. Date & Time has twenty.

Three defects lived in that blind spot, all now fixed (render-composition carries the rules):

  • A dialog item could silence P3 and put nothing in its place. Controls have answered semanticOwnsDisplay all along; items had no equivalent gate at all. Date & Time lost its date, its time, both group boxes and every field.
  • A placeholder was painted over pixels the replay had just drawn. DisplayReplay reported one Bool for a whole window, so nothing could ask about a rectangle. It now reports Coverage — and an ERASE is deliberately not counted as ink, because a composite repaint opens with a full-window erase and counting it would silence every placeholder everywhere.
  • A DITL row typed icon fell through to the generic hatch, though the guest had said an icon is there and named it. Fifteen of NOW's own Workshop sidebar rows drew as fifteen hatches.

LiveShapedRenderTests is the gate that was missing: the first scene fixture in this tree that carries a window's controls and dialog items, asserted in pixels, with each of the first two rules watched failing under its own mutation. The icon probe was written and DELETED — it did not move when the case was removed, so it was not attributing what it claimed, and that change is render-inspected only.

Two things found on the way and NOT fixed here:

  • --open-mirror could leave a window over a stopped poll. start() refuses when no Mac is active yet, which at launch is the ordinary order of events. Fixed (a bounded retry), but worth knowing it presented as "the mirror is frozen and mirror_drive says there is no pinned Mac".
  • The desktop projection reported zero elements while nineteen icons were on the screen — the backdrop is the one Finder window whose icons live in scene.desktopItems rather than window.items. Third instance of that omission shape. Fixed; the desktop icons themselves were never broken, and a render fed the live roster draws all 19 at the Finder's own coordinates.

OPEN: the dialog walk reports a POINTER where a label belongs (2026-08-06, evening)

Guest-side. Memory's captured scene carries twenty staticText/editText DITL rows whose title is eight bytes of 68K address — \u{1e}πN,\u{1e}πM@ — rather than the item's text. The walk reads the text through a handle and reports the handle when the read fails. The same capture reports item rects at l = 16555 and l = 16448: the same corruption one field over.

The renderer now refuses a title containing control bytes (SceneRenderer.displayableTitle), because it had been drawing them — the panel read "πO πM@" where the machine says "Disk Cache size is calculated when the computer starts up" — and, worse, believing them: a non-empty title made the row look like it carried content and silenced the guest's own drawing underneath. That is a defence, not a fix. Evidence: /private/tmp/fsweepB/memory-scene.json, guest build 59dce8562ad4. See fidelity sweep B R-B2.

FIXED: the system font is Charcoal and the pack has no Charcoal (2026-08-06, evening; closed 2026-08-07)

Closed by rasterising it. Charcoal has no bitmap strike anywhere on the image, so there was nothing to lift; the extractor now renders it from its own sfnt at the 16 ppem sizes Apple's hdmx table carries device metrics for, and font 0 answers Charcoal. The widths are hdmx's and reproduce all eight group-box bands the guest measured for itself, exactly; the shapes are FreeType's and disagree with the guest's own pixels on 5.86% of ink, against Chicago's 72.81%. Both symptoms below are gated in RendererTextFidelityTests, watched failing by putting the substitution back. Bold Charcoal and font id 2002 still substitute and are named in charcoal-strike.md. Not metal-verified.

2026-08-07, the leftovers now DECLARE themselves. Closing the main case did not close the substitution: font id 2002, every family the pack does not carry, and any desk whose pack predates the Charcoal work all still get a face they did not ask for, and all still got it in silence. MirrorKitUI/FontSubstitution.swift answers what was asked, what was served and whose the width is, and LiveMirrorView says so on screen when the Chicago fallback is live. Not the style axis: op.face is still never consulted, so a BOLD run reports exact and is not — that gap is unchanged and is still this section's.

The diagnosis as it stood:

Font id 0 means "the system font", which under the Appearance Manager on Mac OS 8.5+ is Charcoal. The pack carries no Charcoal strike, so DisplayReplay.strike answers Chicago, which is wider — and where an application clips its own text the extra width cuts a glyph off. Date & Time draws "Use a Network Time Server" under clip [40,195,210,217] and the mirror renders "Use a Network Time Serve"; General Controls loses the last glyph of two group titles the same way. The captured bytes are complete (len 25 fullLen 25 trunc false), so this is not truncation and not a renderer defect.

It also read as a second, unrelated defect: group-box frames appeared "stroked through their own labels". That was the same substitution seen from the other end — the CDEF erases a band out of its own frame to put the title in, Chicago overran the band, and the overrun landed on the frame line the guest had deliberately left standing.

OPEN: the desktop pattern is a plausible wrong answer (2026-08-06, evening)

DesktopPattern.tile tiles the extracted ppat 16 ("Mac OS Default") unconditionally. The guest's actual desktop pattern is never read, so every render shows purple Mac faces beside a machine showing a blue swirl. Not a placeholder and not a hatch — which makes it the more dangerous kind of wrong, because nothing about it says it is a guess.

FIXED: four renderer-side text defects from the fidelity sweep (2026-08-06)

The sweep (docs/fidelity-sweep-2026-08-06.md) put six items on a red list. Four were renderer-side and are closed on claude/renderer-text-fidelity; each carries a STATUS line in the sweep document itself, and each is gated in now-host/Tests/HostTests/RendererTextFidelityTests.swift over the same committed capture that exhibited it, watched failing by mutation.

  • R1, small system text a third too large. applFont (font 1) is Geneva, not the system font, and its size was being thrown away as well as its face — so Memory's 251 body runs drew as Chicago 12 and 113 of them measured wider than the clip the application had set for them. a93f6f3f.
  • R2, a declared truncation silent at the glass. len, fullLen and trunc rode the wire and the host dropped them. a93f6f3f, c6b85d75.
  • R6's dropped punctuation. Not a mapping defect: the extractor's char range stopped at 127, so no strike ever held a MacRoman character. 95 → 224 glyphs per strike, keyed by MacRoman. bfc51477.
  • R8, a selected label as a solid black bar. Text painted into its own highlight is drawn in the port's background colour. a93f6f3f.

Still open from the same sweep, and each is somebody else's shape of problem:

  • R6's WRONG glyphs are a different defect. "Check Disl", "8/ 6/202(" and Date & Time's clipped group titles are font 0 — the system font, which under Appearance is Charcoal. Charcoal ships no NFNT strike (TrueType-only; the pack's manifest has said so since extraction), so the replay stands Chicago 12 in for it, Chicago is wider, and the last glyph of a run falls outside its clip. Nothing above 0x7F is involved. fonts/ttf/Charcoal.ttf carries no bdat/bloc, so there are no embedded bitmaps to lift and closing this needs a TrueType rasterizer at the guest's own ppem — or an honest decision to keep substituting and say so.
  • Font id 2002 is treated as an ordinary family and is not one. It draws the menu-bar clock, Appearance's "Current Theme:" line and all 102 of Key Caps' key labels — every one of them a system-font site — so it is almost certainly Charcoal under a dynamically-assigned id. Deliberately NOT hard-coded: ids at that magnitude are allocated by the Font Manager at runtime and are not stable across machines, so mapping the number would be a guess dressed as a fact. It wants either a name from the guest beside the id, or the Charcoal work above, not a constant.
  • The strikes carry style 0 only. op.face (the QuickDraw text face) is decoded and never consulted, and the pack holds exactly one styled strike (geneva-9-italic, not bundled). Bold group titles render regular. Small, and it is a real difference in every panel.
  • R3, R4, R5 and R7 are unchanged — see the sweep.

DECIDED: the asset pack left git, and history still holds it (2026-08-06)

The extracted Platinum pack — 1,154 files, 5.0 MB of Apple's bitmaps — was committed to mirror/host/MirrorKit/Sources/MirrorKitUI/Resources/ as a side effect of wiring the pack up, not as a decision: the offline extractor's default output happens to sit inside a SwiftPM target.

It is now a runtime dependency, resolved by MirrorKitUI/AssetPack and regenerated by tools/extract-assets-offline, gitignored, with the renderer degrading to procedural stand-ins and SAYING so — stderr, a readable .absent(searched:) status, and a banner in the live mirror. Both gates run their suite twice, once with NOW_MIRROR_ASSETS=none. asset-pack.md is the whole story. Every one of the 1,154 files was sha256-verified against ~/Lab/Assets/now-mirror-assets/pack-2026-08-06 before leaving the index — 0 missing, 0 mismatched, 0 extra.

What is still open, and it is deliberate:

  • History holds the 5.0 MB. Removing it from HEAD was the reversible move; taking it out of history is a rewrite, and every worktree and branch off this repository would have to be re-cut. Not taken.
  • Another copy of the same bitmaps is still tracked, untouched: mirror/assets/platinum-pack/ (385 files, 1.9 MB). The other five — the vendored mirror/.claude/worktrees/* trees, 3,789 tracked files and 25 MB carried in whole by 0443ab2b vendor: Mirror whole as a subproject — were removed from the index on 2026-08-07; see the nested-worktrees entry above.

MEASURED: the capture fixtures had no dead weight — 94% of the bulk was whitespace (2026-08-06)

One day put 741,514 insertions into this repository and 720,258 of them were .json. The fixture directory was 9.5 MB.

The diagnosis was formatting, not capture: tools/fidelity-sweep.py wrote json.dump(indent=1), so a 3,509-op drain spent 61,036 lines. Re-emitting with one compact op per line took the directory from 697,828 lines to 39,715 and 9.61 MB to 6.46 MB, losslessly and verifiably.

The trimming that was expected to follow found nothing: tools/fixture-store prune drops ops on ports that cannot reach the window under test, and got 39,017 ops down to 38,983 — 0.09%. Every capture carries one to three ports and all of them reach. The bulk is repaint passes, and those are the finding, not padding.

So the fixtures stay in git, because a store would cost a fresh clone the ability to run the gate. fixture-bulk.md has the numbers, the four mutations that prove the compacted fixtures still catch their defects, and the two gate checks that keep the format and the provenance from drifting.

Left open: nine scene fixtures identify themselves with source: peek and a unix timestamp — a capture shape, not a build or a VM. They are now flagged provenanceIsWeak in the manifest rather than counted as provenance, but nobody can say today which machine they came off.

FIXED: MirrorKit's own 165-test suite was never in the gate — five days ungated, three of them red (2026-08-06)

mirror/host/MirrorKit is a full SwiftPM package, vendored into this tree as tracked files on 2026-08-01 (0443ab2b), carrying its own suite of 165 tests: the scene IR, the captured fixture corpus, hit-testing, the wire resync, the pixel islands. scripts/test-all never ran a line of it.

Nothing did, and the reason is worth stating plainly because it is not an oversight anyone would spot by reading the gate. Every stage of test-all looks like it covers the host: test-host runs swift test — against now-host. test-native compiles C. build-guests runs cross-compilers. Three green stages, none of which had any reason to walk into mirror/. This is the same hole AGENTS.md already names about the cross-compilers — "every other gate can be green while neither guest compiles, because none of them invokes a cross-compiler" — arrived at a second time, from a different direction, four days after the sentence warning about it was written down. A gate is only as wide as the tools it invokes, and adding a whole vendored package to a repository does not widen it.

What that bought: on 2026-08-03 (19dcab89, the semantic IR v2 layer) the scene IR's major moved 1 → 2. That commit updated IRFreezeTests and did not re-stamp the captured corpus. From that moment FixtureTests reported seven failures, and scripts/test-all reported all-green, for three days — until another agent ran the package's suite by hand today and noticed.

The failures themselves were trivial, and that is the least interesting part of this entry. Diffed field by field, produced against expected, every one of the seven fixtures differed in exactly one place: version: 2 produced, version: 1 expected. Every other byte of every scene matched. Nothing in MirrorKit had regressed. They were re-stamped, not re-captured — mirror/AGENTS.md is explicit that the corpus deliberately carries a real garbage menu, ax_oracle_not_found rows and stale samples, so a regeneration that "cleaned" it would have destroyed the thing it exists to be.

Fixed:

  • The gate. New stage 2 of 4, scripts/test-mirrorkit, after the ~2 s native tests and before the guest cross-builds; ~15 s cold. It runs swift build beside swift test, for the reason test-host runs xcodebuild beside swift test: the suites depend on MirrorKit and MirrorKitUI, so a break in MirrorApp compiles nowhere the suite can see it. A missing mirror/ or a missing swift exits 0 with a loud ==> SKIP and the reason, which test-all greps back out of the log and echoes — "did not run" must never again read as "passed".
  • The diagnosis. FixtureTests now compares the version stamp first and alone, then compares the scene body with the stamp removed from both sides. Scene.version is a compile-time constant, so pinning it inside each fixture pinned one fact seven times; the next major bump will say "the corpus predates IR vN — this is a STALE FIXTURE, not a scene regression" once per fixture instead of printing seven scene diffs.

Both watched failing under mutation before being trusted, and committed before mutating. IR.version = 3 reddened the stamp check with the stale- fixture message and no body diff; reversing the z-order in SceneBuilder (backdrops + foreground) reddened three fixtures with a real scene mismatch and left the stamp alone. In both cases test-all stopped at ==> FAILED at: MirrorKit gate and ran nothing after it. Renaming Package.swift away produced the loud SKIP and let the run continue.

What is still not gated. This closes mirror/host/MirrorKit. mirror/tests/ is a directory of Python probes that drive a live guest — nohijack-probe.py, winact-probe.py, trials.py and the rest — so it belongs outside a general gate for the same reason the metal suites do, and no attempt was made to wire it in. mirror/guest/ was not audited here at all; whether anything cross-compiles it is an open question, not a claim. The general form of the defect — a subdirectory arriving with tests nobody wired up — is not fixed by fixing one instance of it.

FIXED: a Finder item's type reached the atlas as AppleScript's WORD for it (2026-08-06)

Every icon the roster read produced carried a type of «cla or stri. Neither names a file type, so all 914 of the pack's creator__type lookups missed and every desktop and Finder item drew as generic-by-kind — twenty identical pages on a machine showing Sherlock, QuickTime Player and a SimpleText document. It cost art rather than contents, which is exactly why nothing ever failed over it.

The cause is one line of AppleScript. The Finder's file type is a type class, not text, and the script concatenates it into its output row with &; AppleScript coerces it the only way it can, to its own SOURCE rendering — «class APPL», or string where it has a WORD for the code. The host then took the first four characters.

It is not OSADoScript's result form. The mangling is inside the string the guest built, so the typeChar/source-vs-display question that looks like the culprit is a dead end; asking for a different result form would change nothing. Recorded because that is where an hour goes.

Fixed host-side (NOWMirrorSource.osType(fromAppleScript:)), chosen over the two guest-side cures deliberately: it needs no guest rebuild and it can be proven against the roster fixture already committed, where the alternatives could only be argued. An unrecognised rendering answers nil rather than a guess — a wrong OSType is a wrong icon drawn confidently, and nil is the generic art we were already drawing.

Still open, and it is the reason the deeper fix is worth doing. The word table cannot be closed by construction: AppleScript renders from all its registered terminology, not just the type names, and that machine answered text returned for SimpleText's ttxt. Any code with a keyword this side has not met falls through to generic. The durable cure is to make the type pass hand back four characters itself — Standard Additions' info for is the classic idiom and its file type is plain text on classic Mac OS — and that is unverifiable from this desk: it is one AppleScript coercion whose behaviour on Mac OS 9 no test here can answer. It wants a VM pass, not an argument.

Also fixed in passing, and worth knowing before writing a test over these fixtures: parseIcons(page + typePass) drops the first F row. Each blob arrives in source form carrying its own quotes, so concatenating them leaves "F at the join. One item silently loses its art; it was found by looking at a render, not by a test.

The broken INTERIORS: invert drew nothing, regions could not be checked, and the magnifier was a document (2026-08-06)

Two systematic content defects were measured across all nine committed capture fixtures, and the point of the entry is that they had different causes — one was ours to fix, one was never sent.

FIXED — invert was a renderer gap. DisplayReplay skipped GrafVerb 3 under the note "invert needs destination pixels we do not carry". That was true of the renderer it was written for and stopped being true when the host began compositing its own canvas: the pixels beneath an invert are pixels the replay just drew, so it is a difference blend against the layer. 34 inverts across the corpus were being dropped, and invert is how classic Mac OS draws selection, carets and pressed states — so the mirror was showing an unselected, uncaretted, unpressed version of every window.

What those 34 actually ARE is worth recording, because the guess was wrong. All of them are ONE rect: [33,116,34,132], a 1×16 bar drawn immediately after Sherlock's search field — the text caret, blinking. Not list selection. The visible win is a caret in the search field, which matches the machine.

PARITY IS THE SEMANTIC, not a defect: a drain holding N blinks of one caret ends visibly on or off by N's parity, and replaying every one of them reproduces the state the machine was in at the end of that stream (sherlock-live has 22 and ends off; sherlock has 11 and ends on). Coalescing them would be a prettier picture of a machine nobody watched.

CONTRACT CHANGE — a region can now say whether its box is its shape. content_table.h said plainly that RGN sends the bounding box and ext = 0, never the region data, so the host drew every region as a hard rectangle and had no way to know whether that was right. Sending the shape was refused: a region is unbounded, the ring is the measured limit (a hooked Sherlock overran 64 KiB in one settle), and the common region is rectangular — the expensive answer would be paid on every op to serve the rare one. So the guest sends a discriminator instead, and it costs literally nothing: the region's own rgnSize rides in ext1, a field the payload already carried and already zeroed. No record grew, the ring budget is unchanged, an older host reads 0 exactly as before. 10 is QuickDraw's minimum Region record, so ext1 == 10 means the box IS the shape.

Three states, not two. ext1 == 0 is a resident older than the rule and reports as "shape unreported" — never as rectangular, because a zero pretending to be an answer is what this contract refuses everywhere else. Every fixture in the corpus is in that state, which is exactly the point: nobody could have answered the question from the old data.

The PIXELS are deliberately unchanged by that knowledge. All 39 region ops in the corpus are ERASES, and for an erase the box is the same area approximated — a marker over it would replace a probably-right background with a certainly-wrong annotation. The verbs where a hard rectangle really is a claim stronger than the evidence are frame/paint/fill of a non-rectangular region, and there are zero of those to look at. Grade that placeholder when a capture of one exists, against it.

FIXED — Sherlock's magnifier was drawn all along, and mis-typed. It is a 48×48 CopyBits at window-local (417,98) with no blitsrc: its source world is never hooked, so no pixels and no identity cross. The replay's icon bound was 12–48 "with margin for masks and badges", so it accepted the button and painted a generic document icon over it — a placeholder typed more precisely than the drawing stream allows, which is the one rule docs/render-composition.md states about that layer. The bound is 36 now, measured rather than picked: every near-square blit in all nine captures is 12×12, 16×16, 18×18, 21×20, 32×24 or 32×32, and the only things above are the three magnifier blits. It reads as the untyped control plate.

CLASSIFIED, still absent — the Custom… popup is invisible to grafProcs. This supersedes "absent and unexplained" below. The radio row is drawn in one pass into a world the plane DID hook (0x1e9c6a60), and in that pass: the first radio's box drew, "File Names" drew, the second box drew, "Contents" drew, the third radio's box drew at x 201–217 — and then nothing at all until the Edit… button's theme drawing at x 411. The popup's whole ~190px span produced no bottleneck traffic of any kind, no world of its size was ever born, and there is no blit anywhere in the capture between 100 and 250 px wide. So it is not "never drawn" (its neighbours in the same row and the same pass drew), not "drawn into a world we do not hook" (its neighbours are in the world we hooked), and not "drawn and dropped" (there is nothing to drop). It is drawn by a mechanism the port's grafProcs cannot see — the same class as the CopyMask item. What that mechanism IS remains open.

arc and poly: no renderer case, deliberately. Zero of either in the corpus, so nothing was built for them. The guest CAN emit both, and poly sends a bounding box exactly as rgn does — which for a polygon is never the shape, so polySize now rides in its ext1 for symmetry and as honesty telemetry. Both stay named whole in the deferred-op inventory: a counter that fell silent would read as coverage.

The deferred-op inventory, which is the standing answer to "what can the mirror not draw yet": 73 → 39 over the corpus. 34 × rect (invert) is gone entirely. The 39 regions moved from rgn (bounds only) — which asserted the box was wrong — to rgn (shape unreported), which says nobody asked. NOWMirrorContentCoverageTests gates the whole number; a new entry there is a finding and belongs on this page.

The desktopItems entry below had no evidence because this side threw it away (2026-08-06)

A dated line under "FIRST LIVE ANSWERS from the 2026-08-05 drive", and not a cause for it: nobody has yet learned WHY the Finder roster does not read. What is now known is why every attempt to learn it came back saying the same thing.

The host's roster read asked the guest, matched the guest's osaErr, and then discarded it. readingOutput only promotes a raised script to an error when its caller opts in (osaFailureIsAnError), and the item pass never did — so a script that RAISED arrived here as ok: true with an empty output row, fell through the roster guard, and reported "incomplete or changing item roster". So did an unparseable reply, a container past the item cap, and a genuinely changing roster. One sentence for five different machines' answers, on the one read that has never worked.

Fixed host-side, reporting only: the item pass opts in, and each refusal now says its own reason — the guest's own osaErr where the guest gave one (NOWMirrorSource.rosterPageRefusal, NOWMirrorRosterReasonTests). The art pass keeps the old behaviour on purpose; it is built on a script that is expected to raise.

Still unmeasured: the code itself. Six emulator VMs were up on this Mac at the time, any of which would have dialled a listener bound to the default wire port — including the session diagnosing desktopItems — so no live run was taken rather than risk answering with a foreign guest. The next drive that reaches a Finder roster now gets the number for free; it belongs here when someone has it.

ANSWERED: the 21 accent ramps ARE extracted now, and one ported colour survived measurement unchanged (2026-08-06, later — plan 016 P1/P3 down payment)

A dated line under the entry below, which said the ramps were "not extracted yet". They are: tools/extract-assets-offline --accent-ramps, generating PlatinumAccentRamps.swift. The clut layout, the scen collection walk, the active-ramp determination and the difference list are in asset-extraction-offline.md.

The headline is the colour that did not move. Platinum.selection was ported as 0x333399 out of a CSS file, and the theme file says Lavender's fifth step is 0x333399 — the same three bytes. Plan 016's stop condition covers exactly this: the constant stayed, only its provenance changed. The one real difference is a selected list row, which was filled with g2 (0xCCCCCC, the Gray Space theme's highlight) and is now 0xCCCCFF, the default scene's own.

Unverified, and bounded honestly. Regenerating all nine captures changed nothing — and not because the change is subtle: forcing the highlight to magenta changed nothing either, so no capture in the corpus has a selected list row, and neither screendump on hand does. The colour is measured from the source and has never been watched on a machine.

Still open: which ramp a RUNNING machine resolves. This image has no Appearance Preferences file at all, so nothing has overridden the shipped default — evidence, not proof. PlatinumAccent.active is the named seam and plan 016 P2's GetThemeAccentColors applet is what closes it. Everything else in 016 — the ~40 metrics, the theme brushes and text colours, GetThemeFont — is untouched, and those are where the renderer's constants can still be wrong without any test noticing.

ANSWERED: the theme file has no chrome in it, and the pack now comes off the disk image (2026-08-06)

Two questions closed, one deliberately left open.

The theme file was opened rather than assumed. System Folder:Appearance:Theme Files:Apple platinum is an 876 KB resource fork holding 21 accent cluts, 15 preview PICTs (fourteen 177×125 thumbnails plus a banner), 14 scen settings blobs, and its own Finder icon. No window frames, no title bars, no scroll arrows, no buttons. mirror-assets.md's inherited claim that Platinum chrome is drawn procedurally survives contact with the file; the evidence table is in asset-extraction-offline.md. The one liftable thing in it is specification, not art — the 21 eight-entry accent ramps — and that is not extracted yet.

The pack is now built offline, by tools/extract-assets-offline: no VM, no wire, no guest. 914 per-application icons (was 186), the System file's full icon set at both sizes, 40 cursors, 8 patterns, 42 carried PICTs. The two routes were cross-checked — the generic icons come out byte-identical to the ones the wire route committed.

Still open, and untouched by any of this: WHICH icon belongs to which Finder item. Icons arrive as identity-less bits. A bigger pack does not help; the route is PlotIconSuite/IconServicesLib interception (plan 015 G4), and the entry below on identity-less blits still stands unchanged. The renderer's generic-by-kind fallback is pinned by a test so the day identity lands, it shows.

Unverified: the menu-bar slot's process-signature join. apps and processes share a PSN and a ProcessRef reports its creator, so that one icon can be real identity — but no fixture carries an apps/processes block, so no render exercises it. It is covered by a unit test against the pack only.

ANSWERED: the join works end to end on the control — and A2.1's pointer compare was wrong (2026-08-06, plan 013 slices A–C)

The host can now place a hooked GWorld's ops inside the window its blit names, with no pixels on the wire. The chain, all emulator-verified on mac99/OS 9.1 in one afternoon: the resident emits a blitsrc record (probe mode only) immediately before each bits record whose source resolves to a hooked offscreen port; the drain carries it (srcPort/srcPixmap, 0x-hex); and NOWMirrorContentPlane holds offscreen-keyed ops bounded, then replaces the joined blit's bits op with the held ops re-homed — origin-shifted by dst - src, clipped to dst, window state restored after. Against the loop control: 1000 blitsrc records, every one naming the port the applet reported for itself, and the captured drain is now a committed host fixture whose test places all six 'offscreen row' texts (testControlCaptureJoinsSixOffscreenRowsIntoTheWindow).

The correction worth the entry: plan 013 A2.1 prescribed comparing each row's same-instant handle deref against the src_bits pointer. Measured false — the bottleneck receives a copy of the source PixMap (odd address 0x1eb6aaae), so identity never fires. The working resolve falls back to shape via now_content_probe_pixmap_match, the same route the chase itself was forced onto. The plan carries the correction inline.

Two rig facts from the same runs: the loop control blits faster than the 64 KiB ring holds, so a one-shot drain from the arm-time cursor resyncs and answers EMPTY against 915 recorded ops (gwprobe --drain-seconds is the cure; ring pressure itself is still deferred item 4 of plan 013); and a Retro68 applet without canBackground stops dead — and stops writing its report — the moment anything else comes front, which reads exactly like a crash.

ANSWERED: a real Finder interior composes host-side, without one pixel on the wire (2026-08-06, plan 013 slice D)

The arc's payoff, emulator-verified the same day the join landed: a drain captured live off the CFM Finder carries 186 ops recorded under its hooked offscreen world plus the blitsrc+bits pair that reveals them, and the host join places all ten real labels at their true pens'10 items, 3.21 GB available' [135,14], 'Documents' [280,67], 'TimBotTu' [282,131], 'TBT' [40,195] — the same values plan 013 quotes from the original measurement, re-captured through the entire new pipeline. NOW_RENDER_OUT on the payoff test rasterizes the composed scene offscreen (Mirror's render-screenshot rule); testFinderCaptureComposesTheRealInteriorHostSide is the committed fixture gate. Icons are placed bits geometry, per deferred item 1.

The stimulus that worked, because two acts did not. winact resize refuses on a fresh boot (the anchor plane is absent or not armed — the standing anchor-bind entry below), and menuact View-toggles answered dispatched while the view never changed and the ring wrote nothing — a second dispatched-but-nothing case worth its own look. What forces a full composite rebuild with no act plane at all: front NOW's own window over the armed one, then re-front the Finder. The uncover repaint is total, and typing (key) gives only direct window drawing — selection redraw worlds are too transient for the chase (sighted, chased, gone: measured misses).

OPEN: a scrollbar is drawn in one place and clicked in another (2026-08-06)

Found by the doc audit, incidentally, and it is the "what you can see you can name" invariant broken rather than a cosmetic slip.

MirrorKit.Scrollbar.arrow = 16 is the HIT-TEST constant — it decides which part of a bar a click lands in, and it feeds track() and thumbRect(). The renderer's arrow boxes are sized independently, from the bar's own narrow dimension (SceneRenderer.drawScrollarrow, added 2026-08-06 when scrollbars gained their arrows). On any bar whose narrow dimension is not exactly 16, the arrow a person SEES and the region that answers a click are different rectangles, and nothing relates the two.

Nobody has watched it mis-click: the two agree at 16 px, which is the common case, so this is a latent inconsistency rather than an observed defect. It is recorded because the failure it would produce — a click that lands one part off — is the kind that reads as a flaky mirror rather than as a geometry bug, and because plan 016's GetThemeMetric(kThemeMetricScrollBarWidth) would settle the true width from the machine and make one number serve both.

ANSWERED: the Appearance path was load-bearing, and the panels' values cross (2026-08-06, plan 015 G3)

Michelle's hypothesis, and it was right in the way that matters: a control panel's field VALUES are drawn into offscreen worlds AppearanceLib creates on its behalf. Date & Time imports no NewGWorld at all — read from its own PEF — yet with the trap patch armed it reports 286 worlds born, 286 died, 0 missed, because the patch is on the TRAP and does not care who the caller is.

Everything a window-port hook could never see now crosses: the date and time digits themselves, every control label (Clock Options…, Date Formats…, Set Time Zone…, Use a Network Time Server, Time server: Apple Americas/…), and the group titles it already had. Gated by testDateAndTimeValuesCrossFromTheThemesOwnWorlds.

Compared against the machine, which is the step that was skipped when the plated render was called an improvement: a QMP screendump from the same run reads 3:17:38 PM where the render reads 3:17:44 PM — six seconds of a ticking clock — with the same fields, groups and buttons in the same places. Guarded by --expect-build; the guest that answered was 1bff0bd2ca39, this checkout's own.

The consequence is bigger than one panel: any application themed by Appearance is now reachable, which is most of Mac OS 9's own interface. The renderer's control-plate placeholder stays as the honest answer for the blits that remain, but it is no longer what a panel mostly shows.

FIXED by comparison, and the gap that remains: Sherlock's grid and popup (2026-08-06, plan 015 G2 partial)

Rendering Sherlock 2 beside a QMP screendump of the same run — the comparison that should have been happening all along — found a real defect and left two honest gaps.

Fixed: the destination's ORIGIN belongs in the join's translation. Sherlock blits every composed element to a CONSTANT dst under a shifted port origin (the same SetOrigin idiom its channel grid uses), so a join computing only dst - src collapsed them onto one corner: the volume row rendered at the window's top-left instead of inside the list. Live state is now tracked per DESTINATION port, and the prologue origin is src - dst + destination origin. The row now lands under its column headers, as on the machine.

Still missing, both visible in the comparison and neither yet explained:

  • The channel grid does not render at all. Its geometry is fully derivable (measured: 8x2, 51x46 cells, 55/50 pitch, selection readable from the sprite source rect) but nothing draws it — the top ~110 px of the render is empty where the guest shows sixteen buttons. Deriving it is a P2 opportunity, per docs/render-composition.md; leaving it in the replay would be drift.
  • The Custom… popup and the magnifier button are absent, and no measurement yet says whether their drawing is missing or merely unplaced. (2026-08-06, later: both classified — see "The broken INTERIORS" at the top. The magnifier was drawn all along and mis-typed as an icon, and is fixed; the popup produces no bottleneck traffic at all and is invisible to grafProcs.)

So Sherlock's TEXT is complete and its LAYOUT is now right for everything that draws; what is absent is absent, and the render no longer claims otherwise.

MEASURED: the ring is the limit, and one cycle now drains a whole one (2026-08-06, plan 015 G1)

~12 KB/s into a 64 KiB ring under active repaint is about five seconds of headroom, while the structural cycle that carried the only drain runs every ~2.2 s and took ONE page of it. That is how Sherlock lost 114018 bytes in a settle and its interior text with them. A cycle now chases the cursor while the guest reports more, bounded to 12 pages — a whole ring's worth — so an awake reader cannot be lapped, while the scene cycle stays the cadence owner (the sibling perf thread's territory; this is additive by construction). The ring stays 64 KiB: growing it costs system-heap bytes on 68K machines, and the measured rate says pacing suffices. If an application beats the paced drain, the ring decision reopens with numbers.

VOID, and the rule that would have caught it: two late findings came from someone else's machine (2026-08-06)

Two results recorded near the end of the plan-014 arc are void and must not be built on:

  • a Date & Time re-capture in record mode, and
  • a qdext check reporting installed 0x00000000 with active mode off, read as a regression in the record-mode graduation.

scripts/spin-up-ppc had refused to boot — "requested anchor port 1720 is in use (another session?)" — so no VM of mine existed, and the wire listener on 5340 was answered by another session's guest running another branch's build. The tells were all present and I read past every one: an identical build stamp across two supposedly fresh boots, a psn belonging to a process from a previous run, and finally an empty /private/tmp/nowvm-kg01 with no qemu.pid in it.

AGENTS.md states the cure exactly — "a metal gate must check WHICH guest answered ... assert a capability only the build under test has before believing anything it says" — and it was written after Metal68KSendTests was fooled the same way. The rule exists; nothing made following it automatic for an ad-hoc harness, which is the real gap. tools/gwprobe.py and the scratch drivers assert nothing about who answered, and until they do, any result from them is a claim about an unidentified machine.

What is NOT void: E0 (static, no machine), and the E1/E2/E3 runs, whose boots each produced a distinct build stamp and whose counters moved consistently with the code under test. The record-mode graduation compiles and its gates pass, but it has not been watched working on a machine, and the entry below should be read with that scope.

ANSWERED: worlds hooked at BIRTH, and Sherlock's interior crosses (2026-08-06, plan 014 complete)

Plan 014 ran end to end the same day it was written, and every slice answered:

  • E0 (static): Sherlock 2 and Appearance both resolve NewGWorld against InterfaceLib, read from their own PEF import tables (tools/pef.py --imports, extended to map symbols to libraries). Only NOW's own binary links CarbonLib, so the CarbonLib hole never stood between the resident and a target.
  • E1 (one boot): a resident-installed 68K patch on $AB1D fired for the CFM Finder — qdext {installed 0x0058e61c, calls 6509, newGWorld 2, foreign 0}. A native PowerPC caller reaches a 68K trap patch through InterfaceLib's glue, exactly as the documentation says.
  • E2/E3: the shim gained a tail wrap for selector 0 and a head action for selector 4, so each world is hooked at creation and released at disposal. Against Sherlock 2 — which the chase hooked 0 of 8 times — 77 born, 77 died, 0 missed, and its interior crossed: the radio labels, the column headers, and the volume row with its real index date.

The host join had to learn to nest. Sherlock composites two levels deep — the list is its own world, blitted into the world holding the interior, which is blitted into the window. Pending claims are now keyed by destination port, and a blitsrc on an offscreen port is a real join rather than noise. Gated by testSherlockInteriorComposesFromWorldsHookedAtBirth.

THE RING IS NOW THE LIMIT, and it is measured: a hooked Sherlock overran 64 KiB inside one settle (lostBytes 114018), and the first run's interior text was lost to the overrun rather than to the mechanism. Draining continuously during the stimulus recovered it (3030 records, all nine strings) but still reported lostBytes 343204 across the run. So deferred item 4 of plan 013 is no longer theoretical: continuous composition needs a drain cadence, a shorter arm, or a bigger ring, and the decision now has numbers behind it.

Also measured, and it corrects the applet's scope claim: foreign counts dispatches the patch saw outside the armed context, and it is NOT zero under load (139–327). The trap table reaches further than one process; the note function declines those, so the plane's behaviour is unchanged, but "an application's trap patch is process-local" was the applet's own failure and is not a general fact.

MEASURED: Sherlock's channel grid is fully derivable without pixels (2026-08-06)

The two-row picker of channel buttons looks like the least tractable thing on that window - sixteen cells of pure BMP art. It is not. Read out of the drain:

  • Sherlock draws each cell by moving the port origin and blitting to a constant (0,0,51,46). So the grid is in the state/origin ops: 8 columns x 2 rows, cells 51x46, pitch 55 horizontal and 50 vertical, at x = 27 + 55k and y = 21, 71. Sixteen distinct (origin, source) pairs, exactly the sixteen cells on screen.
  • Selection is readable from geometry alone. The selected cell's well art comes from sprite source rect (142,149,193,195); all fifteen unselected cells come from (209,231,260,277). Which channel is active is therefore a fact the mirror can state, with no pixels and no guessing.
  • Every cell's hit rect is the same arithmetic - origin plus 51x46 - which is precisely what an act-plane click needs.
  • Each channel ICON is its own offscreen world (a 32x32 blit from (0,0,32,32) of a per-icon world), so with worlds hooked at birth each icon has a stable per-session port identity via blitsrc.

What is still NOT derivable is the channel's NAME. The icon pixels come from resources by a path the bottlenecks do not show, which is the standing icon-identity item; Sherlock imports IconServicesLib (16 symbols, from its PEF), so the route exists if it is ever worth taking.

DONE (2026-08-06, later): the derivation exists and the grid renders. MirrorKit.DrawnCellGrid reads the idiom back out of a composed display list and emits ordinary Scene.Controls, so the renderer draws wells with the measured selection and HitTester resolves a click to a named cell. It is in the renderer-free core rather than in DisplayReplay on purpose — see render-composition.md, "P2 derived from P3's own evidence". Every number above was re-derived from the two committed captures rather than restated, and both captures agree cell for cell.

The render was compared against a guest screendump, and the comparison paid for itself: fifteen of the sixteen wells and eight of the nine channel icons were in the composed stream, at the right coordinates, and drew nothing. SetOrigin offsets a port's visRgn but NOT its clipRgn, so a clip set once and then walked by the origin — which is precisely this grid's idiom — slides across the pixels with it. The replay froze each clip into a mirror-space rectangle when the op arrived, pinning it to the first cell. Fixed; of the nine committed capture renders only the two Sherlock ones change.

Still not derivable: the channel NAME, as above. The cells are titleless and their icons are generic stubs.

BROKEN, unexplained, and NOT this: the Custom… popup and the magnifier button are absent from the render for a different reason nobody has looked at. (2026-08-06, later: looked at. Two different reasons, both at the top of this page — the magnifier is a 48×48 blit the replay mis-typed as an icon, now fixed; the popup emits nothing a grafProc can see, and stays open.)

The coverage spread, and the one application that beat the chase (2026-08-06, later)

Deferred item 3's thin spread is thin no longer. Live captures, each a committed fixture behind NOWMirrorContentCoverageTests: Finder icon, list and button views all composite and all join (list view crosses with its column headers and every row's real modification date); Date & Time and Memory are non-compositing exactly as the static table said, so the plain window hook reads their interiors whole; and NOW's own window composes its Workshop, with a retarget test pinning that a captured window stays composed as expected-stale when the plane moves on — the hatch behind the Finder is gone.

SUPERSEDED by plan 014 — see the entry above; the mechanism that closes this now exists and Sherlock is proven through it. Appearance itself has not been re-run. The original finding:

BROKEN in exactly the predicted way: Appearance. It builds a transient offscreen world per widget blit, and the sight→chase→hook cycle loses every one (measured: 10 misses, 34 busy drops, 136 small-blit refusals, no text crossing). Its window-port geometry records fine; its themed text never appears. This is not a new defect — it is the world-replacement fact at its sharpest, and the D0 resident NewGWorld patch (port handed over at creation) is the mechanism that would close it. Until D0 is re-run from the resident, Appearance-style per-widget compositors are out of reach and should be stated so. Sherlock 2 joined it the same evening, at full-window scale: its entire interior is one transient 490×448 composite (7 sights, 7 misses, 0 hooks), so the channel picker, list rows and radio labels are all behind the same wall — only its window stream's own drawing crosses (the Edit… well, the search caret), and the coverage suite gates exactly that plus the honest hatch. D0's resident NewGWorld patch is now the single mechanism between the mirror and both applications' interiors, which makes it the next slice by a wide margin.

Two rig facts from the same sweep: menuact works on a FRESH boot (the View-menu marks move; the earlier no-op was the stale kd01 boot — the scene-walk staleness trap again, now with an act-plane face), and the render fix for scrollbars was chrome, not capture: a ranged scrollbar is window furniture and draws over display-owned content.

ANSWERED: a 68K trap patch DOES see the CFM Finder's NewGWorld (2026-08-06, plan 014 E1)

The question that voided the applet experiment, asked properly - from the resident, installed on the armed pass in the armed process's own context - and answered on the first boot:

qdext: {installed: 0x0058e61c, calls: 6509, newGWorld: 2,
        lastSelector: 0x00080006, foreign: 0}

installed names the incumbent, so the patch demonstrably went in; calls counts dispatches seen INSIDE the armed Finder; newGWorld: 2 is two selector-0 dispatches from a native PowerPC application. That is the whole precondition of plan 014: a native caller executes no A-trap itself, but its InterfaceLib glue reads the trap dispatch table at call time, and the Finder resolves NewGWorld against InterfaceLib (E0, read from its own PEF import table). foreign: 0 says the patch did not fire outside the armed context, which is the scoping property the plane promises.

So creation-time notification EXISTS, and the transient worlds that beat the sight-then-chase route (Sherlock 2, Appearance) are reachable. The escalations named as fallbacks - CFM TVector redirection, pixel islands for composite-locked windows - are not needed and should not be built.

NO VERDICT on D0: an applet cannot ask the trap-patch question (2026-08-06)

Plan 013 slice D0 asks whether a NewGWorld trap patch fires for a CFM caller. tools/guest-gworld/src/trapwatch.c plants a counting tail-patch on the QDOffscreen dispatch ($AB1D), with a dedicated selector-0 (NewGWorld) counter after the raw count proved to be ~0.6/s of ambient selector-7 traffic. The control failed: a freshly launched 68K loop applet's startup NewGWorld — and its LockPixels-per-blit storm — never crossed the patch. An application's trap patch is process-local under the Process Manager; this is the act plane's own paid lesson (act-plane-click-never-taken: installed once, in NOW's context only) resurfacing in a rig applet. So the applet experiment is structurally unable to answer D0, and neither "mechanism exists" nor "fall back to the scan" is proven. The rerun must install from the resident at the armed pass, in the target's own context — exactly the act plane's install_patch pattern; the applet and its protocol are committed for that rerun to reuse.

ANSWERED: a foreign process's offscreen GWorld is hookable, and its labels are readable (2026-08-06)

qdtrace start mode=probe exists and is emulator-run: it sights a blit into the armed window, then sweeps the application and system zones for the CGrafPort that owns those pixels and hooks it as an offscreen row. The brief is what it answers; scratch log in docs/local/gworld-probe-run-notes.md.

Working: the recorder (SimpleText, the positive control, gives 22 text ops at their true pens). Re-confirmed: the Finder's icon view emits one content-sized blit, src == dst, zero text, zero per-icon ops — an independent reproduction of finder-window-icons-are-offscreen-blits on a new binary.

SOUND — the mechanism, and this is the premise everything else rested on. tools/guest-gworld, a 68K applet that allocates its own GWorld and hooks it in its own context, fires every family on all three allocation flavours: text 1, line 1, rect 2, rgn 1, bits 1. Offscreen drawing DOES consult grafProcs, so outcome 3 is off the table as a mechanism claim — any null from the foreign probe is a DISCOVERY failure. PlotIconSuite emitted a blit into the offscreen port, which answers outcome 2 favourably for the family most likely to have failed. Every geographic assumption the chase makes also holds: port in the APPLICATION zone, portRect == bounds, grafProcs NULL at allocation, RecoverHandle agreeing; useTempMem moves only the pixels.

THE CHASE WORKS, proven on a purpose-built control (tools/guest-gworld/src/loop.c, one GWorld held for the process's life) whose port address the probe matched exactly. Then on the Finder: holding the hook across a reflowing resize, its offscreen port gave 8 text + 24 rect + 11 rgn + 8 bits while the window port gave 4 opaque blits — and the text carries real filenames at true pens ('Documents' [280,67], 'TimBotTu' [282,131], 'TBT' [40,195], header '10 items, 3.21 GB available' [135,14]). Window icon-view labels are recoverable as semantic text, composable host-side with no pixels on the wire. Corpus finding: gworld-offscreen-ports-are-hookable.

THREE DEFECTS had to be cleared, and each hid the next: small blits stealing the chase slot (fixed by largest-wins); a dereferenced pixmap pointer carried across the draw-time/event-loop gap, when LockPixels relocates the RECORD out of the app zone (fixed by matching on SHAPE); and — the one that masked both fixes — a read guard bounded to [0x1000, MemTop], where MemTop is ~14.8 MB and application heaps sit ~512 MB above it, so it rejected every candidate it ever examined.

STILL OPEN. Icons arrive as identity-less bits while labels arrive as text, so full host-side composition still needs PlotIconSuite/IconRef interception. The broad spread (list view, a control panel, a double-buffered app) is unrun: after a long session the scene walk stops returning Finder windows with addresses, so phases lose their target — a fresh VM per phase is needed. Ring pressure is real: a composite-heavy app overran the 64 KiB ring inside one settle (lostBytes 835410).

Paid for on the way: the scan crashed the Finder. It dereferenced a portPixMap read out of arbitrary heap bytes with only a NULL/odd check, and a wild pointer into unmapped space is a bus error taken inside the application the resident is a guest in. Now range-checked to physical RAM on both hops — the general rule being that resident code walking a heap for a shape dereferences pointers it did not compute.

Two instrument lessons worth more than the probe: the counters encoded that crash a full run before anyone read the screen (a scan that increments its entry counter and none of its exit counters never returned), and nothing on the wire reports a guest-side alert at all — every counter read green while a crash dialog sat on screen.

Start here: the biggest OPEN items, as of 2026-08-06

The entries below run roughly newest-first and mix FIXED, BROKEN, CLOSED, LIVE RISK and MEASURED in one sequence, so nothing about their order tells you what is still live. This block does, and only this block is maintained — it is a pointer list with a date on it, not a summary that replaces reading the entries. Everything below it is append-only as the rule above says.

For what a person driving the Mirror actually experiences right now, good and bad, read mirror-state-of-play-2026-08-06.md first; it is the short version of this list with the wins beside it.

  1. Nothing is metal-verified. Not one change from 2026-08-05/06 has run on a PowerBook. Two of them are worse than merely unverified because they are the kind that behaves differently on hardware rather than merely slower: the Open Transport wake notifier, which runs at interrupt time where a mistake is a crash and not a slow answer, and the act pump, which now serves wire requests inside a window where a trap patch is live in every process on the machine. Neither has a scheduled pass; a metal pass is attended and Michelle's call. wirestat wake off disables the notifier from either face without a rebuild.
  2. The anchor-bind defect — actselftest answers no-such-process, and foreign processes report not-observed on a fresh clone. The heartbeat half was found and fixed on 2026-08-06; the no-such-process half did not change. This is the item that blocks measuring the others: it stopped two agents on 2026-08-06 from reproducing reported cases, because a clone on which no foreign process binds cannot be driven into a control panel or a foreign modal. Entry: "BROKEN: the anchor plane is active and binds nothing (2026-08-05)", and read its 2026-08-06 append, not just its title. Do not read mirror-parity-ledger.md's "anchor settle window" row as covering this; it carries a correction now saying why.
  3. The no-hijack protection was traded away, on purpose, with Michelle's approval. Entry: "LIVE RISK, deliberately taken", immediately below. The signature if it bites is an act that failed and a scene that moved anyway; go straight to that entry if you ever see it, because the analysis is already written.
  4. The element-classification gap — 190 of 308 corpus items carry knowledge: unknown, which is the root of most remaining render defects (a control whose kind is unknown cannot be drawn as the right widget). The cause is diagnosed and it is not a missing capability: a complete classifier ships in ext/src/now_semantic.c :: classify(), and it is starved by a transport that carries one NowPeekSemanticCell, spends it on the lowest-priority claimant, and only for the front process. The inventory is mirror-element-coverage.md; the diagnosis is "BROKEN: the control classifier gets one shot per scene, and 121 controls never got one (2026-08-05)" below, which is marked fixed-pending-verification and is not the whole of row 2.

The open experiment that used to be item 5 is CLOSED (updated 2026-08-06, after the merge). Most blank and hatched window interiors traced to the OS 9 Finder compositing its icon views in an offscreen GWorld — 25 ops and zero text for a full repaint. Whether that was recoverable semantically was the one unanswered question here, with a written brief: gworld-probe-brief.md. It is answered and outcome 1 shipped. The brief's own header says so; the architecture is render-composition.md. Do not start it.

  1. What replaces it: nothing in the composition arc has been watched in the live host application. Every composition result comes from replaying committed captures inside tests, and every guest-side number is QEMU mac99. That is the single largest gap in the arc that just landed, and it is a watching task, not a building one — see "Nothing in the content-plane arc has touched metal (2026-08-06)" at the foot of this file, which is the fuller statement.

FIXED in the host and the guest, UNVERIFIED by any drive: an alert rendered the wrong buttons, and they did nothing (2026-08-06)

Michelle, driving: a guest alert "renders the wrong buttons, and they dont work". Both halves are one defect, and none of it was the act plane.

The specimen. Internet Explorer's Error alert — it raises itself on launch on this image, which makes it the cheapest reproducible alert there is. Captured live off an emulated G4 with a QMP screendump of the same moment (now-host/Tests/HostTests/Fixtures/scene-ie-error-alert.json). The machine: a stop icon, one line — "Security failure. The server reply is invalid." — and ONE OK button wearing the default ring. The Mirror: two hatched "Visual unavailable" boxes side by side, the ring around one of them, no button, and no text at all.

The guest was right about the buttons. Its eight DITL items name item 1 an enabled pushButton titled OK with isDefault: true, and items 7 and 8 userItems — item 7 being the Dialog Manager's default-outline slot, which WRAPS the button and is declared after it.

Three consequences, all host-side, all fixed:

  • SceneRenderer drew a placeholder for every kind it cannot draw, so the outline slot hatched over the OK button and item 8 invented a second box beside it. A user item now draws nothing — its content is the application's, which is P3's business rather than a semantic fact — and no placeholder is drawn over an item it contains.
  • HitTester returns the topmost item, so a click landed on item 7, which carries no action and no reference; InteractionPolicy refused it by name and the mirror looked dead. A click now resolves to the topmost ANSWERABLE item.
  • The act plane was never involved. ditemact addressed to item 1's own reference answered dispatched, dispatched-but-unconfirmed, and the next screendump showed the alert gone. A fire-and-forget act would have left this looking like an act defect for a third time.

The missing text was a real content gap, and the guest fixed it. A DITL carries the RESOURCE's template; an application that fills a message in at runtime calls SetDialogItemText, which writes into the item's own handle. Item 4 reported an empty staticText while its handle held the line verbatim. The walk now asks GetDialogItemText for it (now-guest-ppc/src/scene/dialog_text.h), after proving through the memory seam that the handle dereferences inside the target partition.

The reason that is a Toolbox call and not one more bounded read is worth keeping: the block header on this heap is not the 24-bit-era layout. Measured with the QEMU oracle on a stopped VM, for the item above — master pointer 0x1e357590, the eight bytes below it 00000040 00088278. The second longword is a ZONE-RELATIVE offset, not the master pointer (two items in the same alert differ from their handles by the same base, 0x1DFB42A0), so a back-pointer coherence check against the handle fails. And the physical size is 64 for a 48-byte string while the tag byte's size-correction nibble reads 0, so physical - 8 - correction yields 56 and would append eight bytes of heap slop to every alert message. GetHandleSize knows the logical size; nothing else on this side does. That header is now a discovery task somebody may want, and it is not needed for this.

The sibling, from the same evening — Set Time Zone's ring. Item 1 Done reports isDefault: true AND enabled: false while the machine greys Done and rings Cancel (recorded below). isDefault comes from the DialogRecord's aDefItem, which the Dialog Manager initialises to 1 and only SetDialogDefaultItem moves, so it goes stale exactly when an application greys its first button. Verified against the IE alert that aDefItem is otherwise RIGHT: it said 1, the machine ringed item 1. So the renderer no longer draws a ring on a disabled item.

What that does NOT close. Nothing here knows where the ring WENT. The authoritative fact is the control's own kControlPushButtonDefaultTag, which the resident's classifier could read the same way it reads kControlKindTag — it would need a flag in NowPeekSemanticClassRecord and a field on a control's semantics, and no drive has asked for it. Until then a disabled default renders no ring at all, which is honest and incomplete.

Also still out of scope, and noted rather than scored (drive-loop rule 2f): the alert's stop icon is a hatched placeholder, and the IE window behind it reads "Guest content not reported".

Emulator only, and no drive has watched any of this in the Mirror window. The evidence is a paired capture and five tests over the captured document, each watched to fail with its fix reverted; the guest half was watched to change the wire on a live machine (build bcd1b1893664, wire 5600).

LIVE RISK, deliberately taken: the act wait now services the wire, and the no-hijack argument's single-cell protection is gone (2026-08-06)

This is not a defect report. It is a protection that was spent on purpose, written down so that if it bites, nobody spends a day deriving what was already known. The latency fix it bought is the entry below this one; read that first for what the ten seconds cost.

What changed

now-guest-ppc/src/act/act_client.c :: act_yield now calls now_wire_pump(). The guest therefore serves host requests while an act is armed — while a trap patch is live in every process on the machine, waiting for a click.

Until this change, no-hijack-criterion.md §4 argued the act plane's single request cell was safe because two requests could not overlap: the act wait did not service the wire, so a second act command sat in the socket until the first finished. That document said in as many words that the protection was incidental — "the guest app's threading model, not an interlock" — and named this exact change as one of two that would remove it. It is now removed.

Authorised by Michelle, 2026-08-06: "im ok with side stepping the no-hijack work for this. lets just be sure to flag it explicitly."

What stands in its place, and what it does not cover

now-guest-shared/src/now_act_inflight.{c,h} — a one-act-at-a-time latch. now_act_cell() refuses a nested act with act-busy before it can write a field; now_act_submit claims, now_act_withdraw releases. The position matters: act_cmds.c fills in op, control_handle and arm_point_h/v before it submits, so a guard at submit would have fired after the damage.

It covers the act cell and nothing else. Not covered, and each is a real gap rather than a formality:

  1. Everything that is not an act, served mid-arm. A scene walk, a census, ps, a file transfer — all of these now run inside the armed window. A scene walk in particular reads foreign process memory while the patches are installed. Nothing has measured whether that is harmful; it is new, and it is unmeasured.
  2. Nested command-dispatch depth. Each command served from inside the pump is a fresh stack frame carrying char result[kNowCommandResultCap] (3072 bytes, wire.c in the command.request arm). A command that itself waits and pumps — chat, quit --wait, another act — nests again. Bounded in practice; unbounded in principle; not addressed. The same exposure predates this change on the chat/exec.input path, which is why it was not treated as a blocker.
  3. A second application linking the act plane. The other removal §4 names. Untouched.
  4. The pending-press hijack. A property of the guard's identity clause, unaffected either way, and still unmeasured.

What would detect it if it bites, and what it looks like from the host

  • The latch firing at all. Any act reply with code act-busy means the wire really did dispatch an act into an armed window. Zero is the expected reading. A non-zero one is not a bug — it is the interlock doing its job — but it is the first evidence that the interleaving is reachable in the product, and it should be reported rather than shrugged at. The guest also counts it: now_act_inflight_refused().
  • The failure the latch is preventing, if the latch is ever bypassed. §4 describes it precisely: not a hijack of the user, but the caller's own synthesised click escaping into the interface. From the host it would look like an act that reported act-not-taken or act-timeout while something else on the guest's screen changed — a window activated, a checkbox toggled, a menu item taken — at coordinates the first request named. An act that failed and a scene that moved anyway is the signature. If that is ever seen, this entry is where to start.
  • What is NOT evidence of this. A slow act, a refused act, or an act that changes nothing. Those are the ordinary failures the plane has always had.

Measured, 2026-08-06, and the interlock was watched firing

Session-private clone, VM /private/tmp/nowvm-actpump, wire 5630, anchor 1740, guest build 04f5dba645ad 2026-08-06T21:23:11Z — asserted against the hello before any number below was believed. Instrument: tools/local-act-pump.py, which encodes the procedure and refuses a wrong build, an unarmed anchor plane and an act it could not aim. Emulated G4; no metal.

The act made to run its full deadline is ctlact at a control in NOW's own window while another process is in front — a click in a background window activates it rather than reaching TrackControl, so the request expires. Its answer is act-timeout (one phase), not act-not-taken.

1. The queue is collapsed.

before (2026-08-06, guest 711abdbd25ec) after (guest 04f5dba645ad)
the act 6.6 s 5.07 s — still its whole deadline, as it must be
scenes answered during it one, in 6634 ms 80, median 65 ms (min 18, max 146)

The act's own cost did not move and was not supposed to: the machine still will not take it. What moved is that the wire is no longer inside that wait with it.

The polling is not what makes the act expire, which had to be ruled out before the polled number meant anything: the same act run silently, with nothing asked of the guest while it waited, took 5.09 s and 5.00 s.

A taken act is still fast. menuact File → New Folder against the front Finder: 0.15 s, 0.20 s, 0.08 s, ok each time. Pumping did not tax the path that works.

2. The lease is held across a full-deadline act. requested=7, active=7 before and after, and the NEXT act answered act-timeout — its own deadline — rather than act-plane-absent. Read this narrowly: this act is ONE 5 s phase, and the lease is 10 s, so this run did not reproduce the >10 s lapse it is supposed to prevent. What it does show is the renewal path running — 80 host requests were answered inside the act window, and every one of them renews. A two-phase act could not be produced on this clone, because phase 1 never succeeded against the only target available.

3. The interlock fired, on a machine. Two ctlacts sent back to back with no read between them:

5.02s  first  -> act-timeout
5.02s  second -> act-busy

That single line is two findings at once. The second act reached the guest's dispatcher while the first was armed — impossible before the pump, and the exact hazard this trade bought — and it was refused before it could write a field. Both halves watched rather than argued.

What is proven, and it is less than it sounds

  • The latch's interleaving is executed: now-guest-shared/tests/now_act_inflight_test.c, run by scripts/test-native. Its load-bearing assertion is that the armed request's identity fields are byte-identical after a second act has been refused.
  • Its wiring is pinned as source: now-guest-shared/tests/act_inflight_wiring_source_test.py — the pump is present, now_act_cell() refuses on the latch, submit claims, withdraw releases, and every verb reports now_act_why_no_cell() rather than a hardcoded "no extension". Six mutations were watched failing, two of which first PASSED and forced the check to be strengthened (a paired idempotence check that a toggling release survived; a fixed-width window that reached past ditemact's not-fired exit into the success path's withdraw).
  • What a busy refusal PROTECTS is still unproven. The run above shows the latch refusing. It does not show what would have happened without it, because that would mean removing the guard on a live machine and watching an armed request's own synthesised click escape into the interface. Nobody has done that, here or upstream.
  • Michelle's own Set Time Zone case was not reproduced. Foreign processes report not-observed on this clone (the anchor-bind defect already in this ledger), so no control panel could be driven. What the reading above establishes is the mechanism, not her instance. Separately, code reading says a Mirror button click becomes ditemact, which is single-phase — the resident half sets fired at serve time (ext/src/now_ext_act.c), so a dialog dismissal has no phase 2 to burn. A slow dismissal therefore has to have been queued behind something else's act wait, or refused off a lapsed lease — both of which are what the pump addresses. That is inference from the source, not a measurement of her machine.

INVESTIGATED, then FIXED the same day: the twelve seconds under a modal is NOT starvation — it is NOW's own act wait, which did not pump the wire (2026-08-06)

The investigation below stands as written and its numbers are the BEFORE. §6's first recommendation was taken: act_yield now pumps. The cost that made it a decision rather than a patch — the no-hijack argument's single-cell protection — was paid, and the entry above carries what replaced it, what it does not cover, and the after numbers.

Investigation only; nothing product-facing was changed. Michelle's live session (host 87960f70, guest 711abdbd25ec) with Date & Time's Set Time Zone modal open: request_ms=12041 decode_ms=98, again 12292/97, 12179/153, 12501/294, while healthy cycles in the same log read total_ms=15–33. The guest's own walk phases in those slow lines total ~2 ms. Her Acts panel: click "Cancel" — queued behind 0, waited 0 ms, guest 12099 ms, never settled, plus 9045, 11732, and two refusals at 2790 and 11732.

Everything below was measured on 2026-08-06 on a session-private clone (VM /private/tmp/nowvm-starve, wire 5590, anchor 5591, guest 711abdbd25ec 2026-08-06T19:53:53Z — the same guest source as Michelle's), with tools/local-modal-starve.py and tools/local-modal-act.py. Both refuse a guest whose hello build is not the one asked for.

1. The twelve seconds and the act's twelve seconds are ONE event

now-guest-ppc/src/act/act_client.c waits for the target process to take an armed act in two phases — now_act_submit, then now_act_await_fired — each bounded by kNowActDeadlineTicks = 300 ticks = 5 seconds, and each spinning on:

static void act_yield(void)
{
    EventRecord ev;
    (void)WaitNextEvent(0, &ev, 2L, NULL);
}

That is a nested Toolbox loop that does not pump the wire, which AGENTS.md names as one of the two non-negotiables that bite hardest. While it runs, conn_service does not, so every scene request sits in the socket for the whole deadline.

Measured, with the Set Time Zone modal up and an act aimed at a control behind it (which the modal's own ModalDialog will never route to):

the act 6.6 s, act-not-taken: armed, and the application never called TrackControl
a scene.request sent in the same instant 6634 ms

The same number twice. The scene did not become slow; it queued. Two phases expiring instead of one, plus the resolver and the reply, is where 11.7–12.5 s comes from — so request_ms=12041 and guest 12099ms in Michelle's log are the same twelve seconds seen from both ends, not two independent slownesses.

It also explains the self-sustaining part. The anchor plane's owner lease is ten seconds and renewal now rides inbound host traffic (4b972ade) — but renewal happens in conn_service, which is exactly what does not run during the act wait. A ~10 s act therefore lapses the lease it was supposed to hold, and the next act refuses the anchor plane is absent or not armed.

2. A REAL application's modal does not starve NOW at all

The hypothesis under test was that a modal in a foreign application starves the cooperatively scheduled machine. For a real modal it is false, on this guest, by measurement. Date & Time's Set Time Zone raised through the product's own ctlact — the application runs its own click handler, so what follows is its ModalDialog and not an instrument's loop:

scene round trip, median max guest's own pass max
idle 21 ms 67 ms 906 ms
Set Time Zone modal up, 145 probes 413 ms 2280 ms 1328 ms

Twenty times slower and never silent. The median conn_service pass moves one bucket, 66–133 ms to 133 ms+. A real modal is a tax, not a wedge, and 413 ms is nowhere near twelve seconds.

And acts work through it. ctlact on the modal's own Cancel answered in 0.7 s and the modal was gone from the next scene — twice, on two separate raisings. Michelle's "Cancel refused the first time, worked after 12 s the second" is this: an act the modal will take is fast, an act it will not take costs the full deadline, and the refusals in between are the lapsed lease above.

3. So this contradicts the ledger's own entries, and here is which way

Three readings existed and none of them agreed. All three are now placed:

  • guest-wedge modal does not model a real modal, and plan 012 §5's row is refuted. That row says "never starved, 71 s — a modal SITTING there starves nothing; ModalDialog pumps and that is enough." Measured today, NOW Wedge modal 45 starved NOW for 43,974 ms of its 45 seconds, and the guest's own histogram says pass max = 44,939,026 us — the event loop did not run once. scan is refuted the same way (43,975 ms; the row says "never starved"). Only the spin row survives (44,061 ms, pass max 45.0 s).
  • The arm-latency entry's "under a modal the wait is the modal, 30.023 s" is confirmed — for the wedge, which is what it measured.
  • The reason both are true at once is the wedge's own loop. ModalUntil calls GetNextEvent with no sleep, and a real ModalDialog does not behave like that. The wedge's modal mode is spin with a dialog drawn over it, and its three modes are indistinguishable: 43.97 / 44.06 / 43.98 seconds. The header comment in tools/guest-wedge/src/now_wedge.c predicting the opposite carries a dated correction now.

4. Whether the machine was alive: answered below the application

The decisive instrument is plan 012's resident channel, and it worked every time. Under all three wedge modes the resident dialled, said role: resident, and went on pinging its own connection inside the starvation window (ping at 41.6 s of an 11.4–56.4 s wedge; 54.0 s of a 19.7–64.7 s one; 42.8 s of an 8.4–53.4 s one) while the session could not answer a scene. The machine was running and NOW was not being scheduled — starvation, measured rather than argued, and the third independent confirmation that plan 012 §4 does what it claims.

Under the real modal the resident pinged on its ordinary 30 s cadence and so did everything else; there was nothing to distinguish.

5. What was NOT measured, and must not be assumed

  • The Finder-owned alert. One attempt today failed to raise it (the Finder opened the file without complaining), so the run measured a healthy machine and is discarded under measurement rule 1 rather than reported. The existing reading stands unchanged and is consistent with the wedge: outcome=starved, request_ms 20,008–21,318, decode_ms=0, and the anchor worker refusing even hello.
  • A large scene under a modal. Every scene here is 2–3 windows. Michelle's had more, and a bigger document is more bulk frames, which is more event-loop passes at 413 ms each. That could add to the twelve seconds; nothing here says it does. Attempts to grow the scene by opening Finder windows failed for an unrelated reason — the Finder reported not-observed, so its windows never entered the walk.
  • Metal. Emulated G4 throughout.

6. What could be done, in order of what it buys

None of this was implemented; it is Michelle's call, like the decode one.

  1. Make act_yield pump the wire (pump.h, the rule that already exists). It does not make an act faster — the machine still will not take it — but it stops one act making the whole Mirror read as a twelve-second machine, and it stops the act wait lapsing the very lease 4b972ade fixed. This is the whole of the reported symptom and it is the smallest change here. Its risk is real and is why it is a decision, not a patch: pumping inside an armed window means serving requests while an act is armed, which is exactly the re-entrancy the no-hijack work is about, and it needs its own guard rather than an assumption.

DONE, 2026-08-06, and the risk above is now LIVE rather than hypothetical. Michelle took the decision and accepted the cost. The guard shipped with it. See the entry "LIVE RISK, deliberately taken: the act wait now services the wire, and the no-hijack argument's single-cell protection is gone" above. 2. State the act ceiling once, where both sides read it. The guest spends up to ~10 s + overhead; plan 014 gave the host a 20 s watchdog chosen against the script ceiling (15 s, kNowScriptDefaultMs) and nothing named the act one. Two limits, two files, the shape this project has already paid for. 3. Nothing needs adding to the resident for this case. It already tells starved from gone, and this case is not starvation. Its value here was as an instrument. 4. The honest half. Against a wedge-class starvation — a Finder-owned alert, or any application that stops calling WaitNextEvent — NOW cannot be made responsive, because it is not running. The product's correct behaviour is to say so quickly, which plan 012 §2 built and the status line's expected-stale vocabulary already speaks. Do not spend anything trying to make that case fast.

The instruments are committed and labelled diagnostic: tools/local-modal-starve.py (latency distribution, the resident's own socket, and the guest's wirestat pass histogram, in one run) and tools/local-modal-act.py (one act, timed, with the scenes queued behind it printed against the same wall clock).

BROKEN, latent, found in passing: two request families draw ids from separate counters and share one watchdog map (2026-08-06)

Found while reviewing plan 014's watchdog, not by a failure. Nothing has been observed to go wrong, and the entry exists because the shape is a collision waiting for load rather than a bug waiting to be seen.

GuestListener.armWatchdog(id:seconds:…) stores into watchdogs: [Int: Watchdog], keyed on the request id alone. But the listener runs three independent id sequences, each starting at 1 and incrementing on its own: nextCommandId, nextExecId and nextCensusId. Two of them arm watchdogs — runCommand (20 s, plan 014 §2) and runExec (60 s) — so command #7 and exec #7 are the same key. The second to arm calls clearWatchdog(7) and takes the slot; the first is then never expired, and if the second finishes normally the first's entry is removed by clearWatchdog on a completion that was never its own.

The failure it would produce is the one 014 exists to stop: a request that never comes back and never times out either, i.e. an unbounded wait the report says nothing about. It is invisible today only because exec is human-driven and rare, so the two counters seldom meet.

The other half of the same observation: requestCensus arms no watchdog at all, exactly as runCommand did before 014. Its completion is stored in pendingCensus with nothing to expire it.

The fix is a single monotonic counter for everything the listener sends, or a composite key naming the family. Either is small; neither has been written, and no test covers the collision. A guard for it would have to arm two families to the same number deliberately — which is also the mutation that proves it, and the reason to write the test first.

INVESTIGATED, NOT FIXED: decode_ms hits 12.5 s, none of it is decoding, and the shape of the bug is a priority inversion (2026-08-06) — FIXED by plan 014 the same day, emulator-verified; see the four dated sections at the end of this entry

Investigation only — no product behaviour was changed. What landed is diagnostic instrumentation and a fixture harness, both labelled as such and droppable. The repair is a design decision and is Michelle's to make.

1. What decode_ms actually brackets

MirrorCycleClocks.decode is publishedAt - deliveredAt, and publishedAt is stamped by recordCycleClocks() from finishCycle(). In NOWMirrorSource.accept() the cycle is not finished until, in order:

  1. NOWMirrorSceneDecoder.decode + phases — JSON.
  2. shadowEngine.accept, continuity, withIcons, enrichFinder, projectedScene — the reducer and projection.
  3. cycleIO.joinContent1–2 guest commands (qdtrace status, qdtrace start, or qdtrace drain).
  4. refreshComplementsrefreshIconsIfStaleAppleScript to the Finder, paged 8 items per container, plus a types pass per container; skipped when the layout key is unchanged.
  5. refreshComplementsrefreshVisibilityAppleScript to the Finder, paged 8 processes. Issued on every cycle; it is the only unconditional multi-round-trip in the lane.

So the field named decode is a bracket containing an unbounded number of guest round-trips. That is the load-bearing fact, and it is why every reading of "12,457 ms" started at the JSON parser.

2. It is not a computation, and it is not our decode

The captured six-window document — Fixtures/scene-quit-modal.json, the guest's own bytes — decodes, reduces and projects in 4 ms. The largest scene ever captured here (70 KB, now-scene-self-hidden-but-front) costs 15 ms. Ten chained reductions cost 44 ms, so nothing compounds with replica size either. MirrorDecodeCostTests is that measurement, kept and bounded.

3. There is no cliff at six windows. There is no cliff at all.

43,448 NOWBASE cycle lines from acts.log, grouped by windows/elements, median decode_ms:

12 windows / 346 elements   n=1914   median   714 ms
11 windows / 298 elements   n=   5   median  1438 ms
 8 windows / 220 elements   n=  10   median   694 ms
 6 windows / 170 elements   n= 442   median  1731 ms
 6 windows / 168 elements   n=  21   median 12824 ms
 5 windows / 153 elements   n=  11   median   462 ms
 3 windows /  60 elements   n=   3   median 12559 ms

Twelve windows costs 714 ms; three windows costs 12,559 ms. The "11% more content, 38x the time" framing does not survive the wider data — decode_ms is not a function of scene size at any point in the range. The two adjacent six-window rows differing by 7x on two elements make the same point locally.

The distribution is bimodal, which is the real signal:

< 100 ms   23.0%        4  - 8 s   0.6%
0.1 - 1 s  72.7%        8  - 11 s  0.03%   <- nearly empty
1  - 2 s    1.8%        11 - 14 s  0.2%
2  - 4 s    1.3%        14 - 20 s  0.4%

95.7% of cycles are under a second. The slow tail has peaks at 5.5–6.0 s (125 cycles) and 14.0–15.0 s (88 cycles) with a gap either side. Sharp modes with empty space between them are waits, not work. The coordinator's read was right and the brief's was wrong.

4. What is being waited on, and why there is no bound

Every complement goes through GuestListener.runCommand, and runCommand arms no watchdog at all (GuestListener.swift:608). A pending command.request is resolved only by a guest reply, a guest-sent problem, or a disconnect. The comment 20 lines below it — "the 15s a command.request gets" — is wrong: that 15 s belongs to the typed file/process families, not to this. So the host has no upper bound on how long a cycle waits for a Finder script.

The only bound is the guest's own, and the host never sets it: neither readIcons nor refreshVisibility passes timeoutMs, so each script gets kNowScriptDefaultMs = 15000 (input_args.h:187) — which is where the 14–15 s peak sits. But the scripts are usually slow, not timing out: 241 cycles exceed 12 s and there are only 4 "stopped at its deadline" notes in the whole log. The deadline is a ceiling that is rarely reached; the ordinary slow case is a Finder that answers, late.

FinderItems already records the price: a Finder Apple event costs ~1–2 s of Finder time on a healthy guest (measured 2026-07-31). With a modal up the Finder is starved and that goes to seconds. The guest is serial, so each script also blocks its whole event loop — which is why acts logged guest 11080ms while queued behind a cycle.

What is NOT established: which of the three — content join, icon roster, visibility census — dominates the 12.4 s. Nothing has ever timed them apart. That is exactly what the instrumentation below now answers, and it should be read off a live run before choosing between options 2 and 3.

5. Why this is the cause of the refused acts

The anchor plane's owner lease is 600 ticks — ten seconds (peek.c :: kNowPeekOwnerLeaseTicks) — and only scene.request renews it. A 12.6-second cycle lets it expire every time. The same log reads requested=15 active=8 with structure/semantics/interaction all back to requested, and Cancel on the Set Time Zone modal refused five times with element-not-found: the anchor plane is absent or not armed. This answers the open entry below ("what takes the planes inactive while content stays up"): nothing took them down — we stopped asking.

Michelle's framing is the right one and it is worth writing down: this is a priority inversion. The structural scene is the product — it is what a person sees and what every act's element reference resolves against. Drawing detail and icon rosters are enrichment. An optional plane is currently able to take the core feature down, which is the same rule docs/resident-components.md already states for the extension and which these planes do not honour.

6. The options

Option 1 — bound the wait (small, low risk, does not remove the inversion). Give runCommand a watchdog and pass an explicit timeoutMs on the complement scripts. ~20 lines. Turns an unbounded stall into a bounded one and fixes the wrong comment. It does not fix the cadence: a 3 s bound on three scripts is still a 9 s cycle, and the lease is 10 s. Worth doing whatever else is chosen; not sufficient alone.

Option 2 — take the Finder complements out of the cycle (medium, removes most of the inversion). The cycle publishes and rearms; icons and visibility fold in when they arrive, guarded by the pinned run rather than the cycle. Also key the visibility census on the process roster with a floor interval — today it runs every poll for an answer that changes only when a process starts, quits, hides or shows. Risk: the cycle-hold was buying one real thing — a roster read for one layout never landing on a different one, because positions are window-content-local and scroll-compensated, so a stale one puts a click on the wrong file. Buy it directly by re-checking the layout key at apply time, which is stricter than the proxy it replaces. Prototyped and mutation-tested on this branch (see below); dropped when the scope changed.

Option 3 — make the content plane asynchronous too (larger, removes the inversion entirely). The cycle publishes structure alone; the P3 join fetches out of band and folds in. This is the one that needs Michelle's call, because it touches things beyond the source: - State engine. Content becomes a plane that is routinely a cycle or two behind structure. The retention rules already distinguish "not observed" from "deleted", so the vocabulary exists — but enrichContent would be applying against a replica that has since advanced, and whether that is safe has not been checked. - Acts. Any act confirming from drawing detail rather than from structure would settle later. Acts resolving element references would not, since those are structural. - The status line. It would have to say "content still arriving" honestly rather than implying the frame is whole. The IR's coverage vocabulary (unavailable, with a reason) already has the words, and expected-stale already renders. - Serialisation. The cycle serialised these by waiting. Once it does not, a join spanning several polls must not have a second issued over it — two arms racing for one target on a cooperative wire.

Also worth doing under any option: do not issue an arm that cannot complete. A sibling measurement the same night has a trace arm against a non-front target never completing (25–45 s, wrongContext climbing ~1,770 in 25 s). join only ever arms the front window, but a failed arm leaves armedAt nil, which makes needsRenewal true forever, so prepare re-issues on every cycle — a retry storm against something that just refused. Unverified, but cheap to check.

7. Delta baselines: the poller DOES quote since

Plain yes. Session.sendSceneRequest sets since = sceneBaseline?.digest on every request unless a periodic resync forces full, and NOWMirrorSource never passes fullNOWMirrorCycleIO.livelistener.requestScenesendSceneRequest with the default. Independently corroborated: the status line's wire token reads same in Michelle's session, and same is reachable only through finishSceneSame, which fires only when a since was sent and matched. walk=full is about PLANES and says nothing either way, as suspected. No change needed.

What landed on claude/host-decode-perf

Diagnostic only, and safe to drop:

  • MirrorCycleClocks gains dc_own_ms, dc_content_ms, dc_icons_ms, dc_vis_ms on the NOWBASE cycle line — the four stages the bracket contains, in the order the cycle does them, absent rather than dashed for a stage a cycle never reached. One live run now answers §4's open question in a grep.
  • MirrorDecodeCostTests — the fixture harness, with bounds that fail if decode/reduce/project ever does grow superlinearly.
  • Timing brackets in NOWMirrorSource.accept and refreshComplements. No control flow changed.

The Option-2 prototype is in this branch's reflog (commit fix(host): the Finder complements no longer hold the scene cycle open) with a mutation-tested cadence guard, if it is wanted.

ANSWERED, live (2026-08-06): the visibility census dominates

Plan 014 §1's question — which round-trip owns the bracket — is closed, and the instrument that answers it was never broken.

The instrument was fine; the log was being read wrong. The fields were reported "absent from every live cycle". They are not: they were absent from the lines in acts.log, because the last line any host wrote before the reading was 13:58:19 and the app built from 1bb7fdcf — the commit that added them — did not start until 13:58:24. That build had never run a Mirror cycle. (nm on the packaged executable shows MirrorCycleClocks.bracketFields; the stack's README names the commit.) Filed here because it is the fourth misread number in three days and the cheapest one to have avoided: acts.log is shared by every host on the Mac and cycle lines carry no guest identity (drive-loop §2m), so a reading has to be attributed by a mark and a process start, not by being at the end of the file.

Measured, own VM on wire 5560, guest build 711abdbd25ec 2026-08-06T18:26:52Z, resident 97fb3f54…, host from this branch. Finder healthy, guest idle, n=85 cycles, one window / 68 elements:

field median p90 max
decode_ms 353 423 1675
dc_own_ms 9 9 10
dc_content_ms 5 6 9
dc_icons_ms 0 0 1349
dc_vis_ms 338 401 746

They sum to decode_ms. The visibility census is ~96% of the bracket on every steady-state cycle; this host's own CPU is 2.5% of it, which is the 4 ms MirrorDecodeCostTests already measured, seen live. The icon roster is 0 while the layout key holds — the guard doing its job, now visible rather than inferred — and 1349–1565 ms on any cycle where a Finder window opened or closed, taking decode_ms to 1.7–1.9 s.

So the census costs ~4× less than a roster read and is paid every cycle, forever, for state that changes only when an application starts, quits, hides or shows. That is §D's question answered in its favour, and it is why §4 was done together with §3 rather than after it.

Loaded, two ways. The NOW Wedge modal 45 instrument: outcome=failed request_ms=20839. A Finder-owned alert ("Could not find the application program that created…"): outcome=starved, request_ms 20,008–21,318, decode_ms=0 for as long as it was up — and the anchor worker stopped answering even hello, so the machine could not be driven at all and the clone had to be discarded.

That is a different sub-case from the one that bit Michelle, and the difference matters: her log has request_ms=82 with decode_ms=12457, i.e. NOW answering promptly while the Finder was merely busy. Her modal belonged to a foreign application. A Finder-owned modal starves the whole cooperative machine including NOW, so the complements never run and this repair cannot help — nothing on the host can. Worth knowing before someone measures the wrong modal and concludes the fix did nothing.

2026-08-06, later: the second half of that reading is right and the first half is a different defect. A Finder-owned modal does starve the whole machine and nothing on the host can help — unchanged. But the NOW Wedge modal 45 row (outcome=failed request_ms=20839) is not a modal result at all: the wedge's modal mode is a non-pumping GetNextEvent loop and starves as completely as spin. And Michelle's own case, which this entry reads as "NOW answering promptly while the Finder was merely busy", has since reversed on her newer host — request_ms=12041 decode_ms=98 — and that twelve seconds is NOW's own act wait failing to pump the wire, not the Finder. Entry at the top of this file.

What plan 014 changed (2026-08-06)

  • §2. GuestListener.runCommand arms a watchdog — it had none, on the family the whole content join and both complements travel on. 20 s, chosen against the guest's own kNowScriptDefaultMs (15 s, input_args.h) so the guest's typed refusal always beats the host's bare timeout; testTheCommandWatchdogOutlastsTheGuestsOwnScriptCeiling fails on the 3 s the ledger above proposed. The wrong "15s a command.request gets" comment is corrected in the same commit, with what it actually described. A timeout is reported: timeouts=N on the cycle line, omitted at zero.
  • §3. The cycle publishes on decode and rearms; the complements run beside it, guarded by the pinned RUN. The layout-key recheck at apply time replaces the cycle-hold as the thing that stops a roster landing on a layout the machine has left.
  • §3, honesty. A frame published before its roster now carries a finder-items coverage claim — typed status and a reason, the same vocabulary process-visibility uses — and the status line says awaiting icons. Absence stays absence in the scene; what changed is that it no longer goes unsaid. (process-visibility was already marked .stale on every structural accept, so the census never presents as current between reads.)
  • §4. The census is keyed on the process roster with a 3 s floor and invalidated explicitly by the hide act.

Measured after, same rig, same guest (2026-08-06)

Fresh clone, guest build 711abdbd25ec 2026-08-06T18:59:20Z, wire 5560, host from this branch. 258 cycles, every one outcome=ok.

before after
decode_ms median 353 16
total_ms median 364 25
total_ms p90 436 27
cycle on a layout change 1936 26
planes 15/15 15/15 (257/258 samples)
cycles reporting timeouts= 0

The complements still cost what they always cost — the log carries NOWBASE finder containers=2 complete=yes ms=1478 and NOWBASE visibility processes=7 complete=yes ms=277 — but they are no longer inside the cycle, so a 1.5 s roster read now sits beside a 26 ms cycle instead of becoming one. The census fired 17 times in 62 cycles rather than 62, which is the 3 s floor working.

The layout-change case is the clearest single line: opening a Finder window took decode_ms from 17 ms to 17 ms, where before the same transition on the same rig read decode_ms=1920 dc_icons_ms=1565.

NOT verified end to end: Michelle's own act. "A dialog act that settles rather than refusing" could not be driven from here. tools/now-agent reaches ONE host per user at a fixed path and her packaged app holds it; taking it would have disturbed a live session, and computer use was out of scope for this arc. What IS verified is the whole causal chain beneath the symptom — the cycle is 25 ms against a 10 s lease, the planes read 15/15 throughout, and testAStarvedFinderCannotHoldTheSceneCycleOpen fails if a complement can extend the cycle again. The act itself is unverified and should be the first thing a drive re-tests.

Still not covered by a test: the watchdog FIRING. The bound and its reporting are guarded; that the expiry actually settles a stored completion is not, because the suites here would have to wait 20 s for it and no other request family's watchdog is covered either. Watched only in the sense that the code path is the one exec has used since it was written.

DEFERRED with a reason, not open: the asynchronous content plane (plan 014 §5)

Explicitly out of 014. What §1's numbers say about it: the content join is 5–12 ms, three orders below the census, so making it asynchronous buys nothing measurable on a healthy guest. Its case is the starved one, where it shares the guest's serial event loop with everything else — and in the starved case measured here the guest did not answer scene.request either, so the cycle was already lost upstream of P3. Ordinary evidence does not currently justify the plan; the four bullets under Option 3 above remain the scope if one is written.

Read this as a decision, not as a to-do. It was scoped, argued against on §1's own numbers, and left unbuilt on purpose — the item that would be dropped by building it is the frame's simplicity, and the item that would be bought is 5–12 ms on a machine where the next cheapest thing costs 1,478 ms. What would REOPEN it is new evidence, and the evidence would have to be of one specific shape: a guest that answers scene.request promptly while the content join does not. Every starved case measured so far had the guest silent on both, which no amount of host-side asynchrony repairs. Measure that first; do not build from the argument alone.

FIXED, emulator only: the 115 ms round trip was the guest's own sleep (2026-08-06)

The deltas arc took the scene walk from ~950 ms to 3–8 ms and cut idle wire bytes by ~90%, and left one number standing: a round trip cost a 115 ms median even when the answer was a zero-byte "nothing changed". Neither the work nor the bytes. This is what it was.

The mechanism, measured rather than inferred. main.c asked WaitNextEvent for a six-tick sleep — nominally ~100 ms — unless a transfer, stream, offer, put or non-empty control queue was ALREADY in flight. A request arriving into a QUIET connection therefore sat readable on the socket until that sleep expired, before conn_service looked at it. The guest now measures both halves of that itself (wirestat): the interval between its own service passes, and the delay from Open Transport announcing data to the loop reading it.

condition round trip notice mean notice max pass mean
6-tick sleep, no wake (as shipped) 86 ms 48.5 ms 103 ms 111 ms
6-tick sleep + wake 10 ms 7.5 ms 105 ms 104 ms
1-tick sleep, no wake 15 ms 10.8 ms 28 ms 27 ms
1-tick sleep + wake 10 ms 2.2 ms 22 ms 27 ms

Mid-drive (fronting Finder and NOW alternately, 25 KB deltas rather than zero-byte answers): 58 ms → 15 ms, notice 48.5 ms → 4.5 ms.

notice is the finding. At a 48.5 ms mean into a ~111 ms loop it is uniform arrival into the sleep, which is exactly the shape a poll produces, and it is 100% of the unexplained time.

Why the wake and not simply a shorter sleep. Both fix the latency. Only the wake fixes it without paying for it: with a 6-tick sleep AND the wake, the loop still sleeps ~104 ms when nothing is happening, so the rest of the Macintosh keeps the time the anchor plane needs — a process's A5 is captured when that process pumps. The 1-tick condition runs the loop 37 times a second on a machine that is usually idle, which is affordable on an emulated G4 and is precisely the kind of cost plan 013 exists to stop paying on a 1400c.

The instrument's own trap, worth keeping. A bench that sends requests back-to-back measures nothing here: answering one leaves a queued control frame, conn_wants_fast_pump() goes true, and the next request lands in an already-fast loop. The first run of tools/local-wire-latency.py reported a 10 ms median for the SHIPPED build for this reason. Every number above is from gap-separated requests into a quiet connection, which is the cadence the product actually has.

What is NOT proven, and it is the whole risk. Every number is from an emulated G4 (os91-runner.qcow2, mac99, wire 5470, anchor 1706, builds e207f041daa0dc2900a44991). Nobody has watched this on a PowerBook 1400c. The wake is an Open Transport notifier — interrupt-time discipline, where a mistake is a crash rather than a slow answer — that does two things and only two: stamp the microsecond clock, and call WakeUpProcess on a PSN captured at startup on the main thread. It acknowledges nothing, so disconnects stay pending for the main loop's OTLook path where they have always been handled. It also keeps the notifier installed after the dial, which it did not before; that is the larger behavioural change of the two. Failure modes are graceful by construction — a notifier that never fires, or a WakeUpProcess that does nothing, leaves exactly the behaviour that shipped before — and wirestat wake off restores it from either face without a rebuild. A metal pass is owed, and until it happens the emulator's own advantage cuts the other way for once: on loopback the byte term is nearly free and the latency term is not, so this win probably UNDERSTATES what a slower machine gets.

Two residuals, both small and both honest:

  • One notification in eleven or so still costs ~100 ms with the wake on. A notification that lands while the process is already awake finds WakeUpProcess with nothing to wake. now_wire_sleep_ticks closes most of it by sleeping one tick whenever bytes are announced and unread, and it did not close all of it.
  • The anchor-plane CANARY (how many applications the scene could bind, how many windows it saw) did not move in any condition — but it was pinned at its floor (1 app idle, 2 mid-drive) in the SHIPPED condition too, so the run says nothing about starvation either way. Anyone repeating this should drive a machine with several applications and real windows before quoting the canary.

Where: now-guest-ppc/src/core/wire.c (the notifier's data era, g_wake_enabled), src/core/wire_sleep.c (the sleep rule, natively tested), src/core/loopstat.c (the histogram, natively tested), src/core/wirestat_cmd.c (the grammar both faces share), tools/local-wire-latency.py (the bench).

BROKEN: putstat has no console face, and probably not only putstat (2026-08-06)

Found while giving wirestat its second face. Typed at the guest's own console — or through the exec plane, which runs the same dispatch — putstat answers "command failed". Over the wire the same verb returns its full row table. Reproduced 2026-08-06 on an emulated G4, build dc2900a44991.

The cause is a buffer, not a missing case. console_model.c :: console_model_dispatch falls through to now_command_run(name, NULL, 0, result, sizeof result) with char result[512]. A verb whose command.result JSON exceeds 512 bytes is truncated, the reply then has no message key, and the console prints the generic failure. putstat's table is over that; gestalt, net and ps are unaffected because each has an explicit console case that reads the producer directly.

This violates command-parity.md — a capability on one face is half a feature — and CommandParityTests does not catch it, because that gate checks specific message families textually rather than every verb's console rendering. It is also invisible from the Workshop: the Diagnostics page reads the same counters through diag_putstat_rows, so the numbers ARE on that machine's screen; it is only the typed verb that fails.

What a fix has to decide, in the order it should be decided:

  1. Which verbs are affected. Anything routed through the generic tail whose result can exceed 512 bytes. That set has never been enumerated, and "probably not only putstat" is the honest state of this entry.
  2. The renderer. Either a rowArray renderer on the generic tail (json.h has now_json_array / now_json_next_array for [label, value] rows) with a larger reply buffer — note the buffer is on the stack of a classic Mac application, so growing it is not free — or an explicit console case per affected verb reading the same producer the wire verb reads. The second is the net precedent ("three faces, one producer") and is what wirestat does, at console_model.c :: console_show_loopstat.
  3. A gate that would have caught it. The shape is probably a source test over kNowCommandDocs (cmd_help.c) asserting every verb either has an explicit console case or provably fits the generic path. Without one this returns the next time a verb grows a row.

FIXED 2026-08-06, emulator only (a153d718) — and the diagnosis above was half wrong, which is why it stays. The 512-byte buffer is real and was fixed, but it was the second limit, not the cause. The cause is that console_model.c's fallback read a top-level message out of the reply and treated its absence as a failure — and no PowerPC verb has ever carried a message on success. So the branch that ran for every command that WORKED was the failure branch, and a verb's own words reached the screen only when it had refused. This entry reasoned from putstat's size because putstat was the verb in hand; the shape of the bug was not "one verb is too big" but "the renderer understands only a reply shape no verb emits on success".

Six verbs were measured printing "command failed" while working: putstat, axsnap, axtree, elements, mouseloc, observe. Underneath it, wire.c sized a command.result at 3072 bytes and console_model.c at 512, and neither number said so — kNowCommandResultCap is the one number now, which is the "state a limit once, where both sides read it" rule from AGENTS.md arriving a third time.

Question 1 above — which verbs are affected — is answered by not needing to be: console_reply.c is one Toolbox-free renderer rather than eighteen console cases, and console_reply_test.c runs it over every reply shape the guest emits.

Question 3 — a gate that would have caught it — is the durable part. CommandParityTests could not have caught any of this, and no version of it structured that way could. It compares dispatch tables, and a table says whether a verb is PRESENT, never whether its answer renders. Three checks now close that gap, all mutation-verified: the console must delegate to the renderer and must not decide for itself that a command failed; both faces must size a reply from the same constant; and every ok:true reply must open an output object. The rule generalises past this project — see command-parity.md, "Present on both faces is not the same as working on both".

Verified at both faces on an emulated G4, build fea48baffe8f, wire 5510, watched fail first on c5c39f61dbbf. Not metal-verified.

BROKEN: no verb reached through the PowerPC console's fallback can be given an ARGUMENT (2026-08-06)

Found while fixing the renderer half of the same seam (see command-parity.md, "Present on both faces is not the same as working on both"). console_model.c handles 27 verbs with a strcmp of their own and falls through for the other eighteen — activate, actselftest, aesend, axsnap, axtree, ctlact, ditemact, elements, handle, menuact, mouseloc, observe, putstat, qdtrace, script, textget, textset, winact. The fall-through is:

now_command_run(name, NULL, 0, result, sizeof result);

NULL is the whole request, so the verb sees no arguments at all. Twelve of the eighteen therefore answer a validation refusal and nothing else, no matter what a person types after the verb. Measured on the emulator, build fea48baffe8f, wire 5510, through the exec plane:

> winact          winact requires action: one of select, close, move, resize, zoom
> menuact         menuact requires menu (the menu's id) and item (its 1-based position…)
> script          script requires source: one AppleScript, as a string
> aesend          aesend requires event: one of quit, oapp, odoc, pdoc
> activate        activate needs a whole process serial number: serialHi and serialLo…

Typing the arguments changes nothing, because the console never passes them. The refusals are correct and useful sentences — they are just the only sentences those verbs can produce at that keyboard.

Six of the eighteen take no arguments and now work: putstat and mouseloc render real tables, and axsnap, axtree, elements, observe and qdtrace answer objects of references, which the console says it cannot show as a table rather than claiming a failure.

What it would take. A grammar, not a renderer: something that turns the rest of a typed line into the args object the wire sends, with types (now_json_find_int vs now_json_find_string — a quoted 3 is not a 3). now-guest-ppc/src/commands/cmd_line.c already reads the OTHER direction and is the obvious place to put its inverse. script and aesend and the act verbs are the ones a person would most want, and docs/command-parity.md's consoleDebt map is the list. The three verbs whose arguments are opaque now-element- references are a different problem and probably stay untypeable.

Why it is parked. The reported defect was a renderer printing "command failed" for commands that had succeeded; that is fixed and verified at both faces. A console argument grammar is a design decision about what a person can type at a classic Mac, and folding it into the same change would have made both harder to review.

BROKEN, and it is the WORST shape: a ten-second gap in scene polling blinds the whole walk, and the Mirror keeps drawing what it can no longer see (2026-08-06)

Michelle, side by side with the guest: Date & Time's Set Time Zone modal open on the machine, and in the Mirror "System Folder, Control Panels and Date & Time" and no modal at all — with the status line reading 5 windows · walk 0ms · transfer 36ms · **same** · content: 13 new draw ops. same is the guest saying nothing changed while a modal was visibly open, which is the stale-mirror-that-believes-it-is-current failure this project fears most.

It is neither of the two things it looked like. Both were tested on a private clone (VM /private/tmp/nowvm-modal, wire 5490, guest build c5c39f61dbbf), asking for WHOLE documents so no answer could be a delta artefact.

The walk is not the defect. With the modal up the guest publishes it, within ~600 ms of the quit and for as long as it is there:

seq=91  windows=[Date & Time; New Old World]                 digest 2c7784cf
seq=92  windows=[Set Time Zone; Date & Time; New Old World]  digest 95b9a08b

kind:2, visible, front, 9 dialog items with a ref on each, 10 controls, coverage complete for that owner, no error against it.

The delta plane is not the defect either. The digest MOVES when the window appears, and scene.same is decided against the digest of the scene just walked, not a remembered one (wire.c :: serve_scene), so a guest cannot answer same while the modal is in its own document. The host's half keeps it too: MirrorQuitModalTests runs that captured document through this side's decoder, the replica reducer and the projection, and the window survives all three.

What IS the defect: the anchor plane is held by a ten-second lease that only a scene.request renews. peek.c :: kNowPeekOwnerLeaseTicks is 600 ticks. Stop asking for scenes for longer and the next walk is blind — measured on the wire, with the modal up the whole time:

gap  3s → Set Time Zone, Date & Time, New Old World   (95b9a08b)
gap  8s → Set Time Zone, Date & Time, New Old World   (95b9a08b)
gap 12s → New Old World only                          (ae3f00e1)
gap 20s → New Old World only                          (ae3f00e1)

In a blind scene EVERY foreign process reports now_no_plane and coverage: unavailable/not-observed, and the document contains not one foreign window. The scene straight after each blind one is right again, because the blind walk is the one that RE-claimed: serve_scene claims the plane and walks immediately, while the extension only arms on its next jGNE pass. Re-claiming therefore costs exactly one scene.

Why that produced Michelle's frame. A second agent measured her guest at that moment: requested=8 active=8 — content armed, structure, semantics and interaction all inactive — and her Cancel click refused five times with element-not-found: the anchor plane is absent or not armed, at 6.7-8.8 s each. Thirty to forty seconds of acts with no scene between them is four leases' worth. Then it is self-sustaining:

  • the planes lapse, so the walk goes blind;
  • the blind document is honest but EMPTY of foreign windows, so the host retains its last-known ones as expectedStale (MirrorReplicaReducer deletes only under complete coverage) — which is the Finder folders and the Date & Time panel Michelle could still see;
  • the blind document is STABLE, so every later poll answers same;
  • the modal, raised after the planes went down, is in no scene ever;
  • and every act refuses, because the plane it needs is the one that lapsed, which spends more seconds not polling.

The split Michelle reported falls straight out of it: a modal she opens by clicking arrives while the planes are up and renders; the one Date & Time raises during its own quit arrives after a stall.

This is not tonight's work. The lease is from 2026-08-03/04 (4ebd575c, 7e5b0c8f). Tonight's deltas made it VISIBLE by putting the word same on screen, and are otherwise innocent — which is the one good thing here.

What is not yet decided, and needs a person. Three candidate fixes, and they are not equivalent:

  1. Do not serve a blind scene. serve_scene already claims the plane before walking; it could pump until arm_active includes anchors before it walks, bounded. The other agent measured the arm handshake at ~15 ms tonight, so the bound is cheap. This removes the one-scene lag.
  2. Do not let the lease lapse under a live consumer. A connected host that is doing anything at all is not a host that has gone away, so the renewal could ride the wire rather than the scene verb. Against: the lease exists so a plane is not armed on a machine nobody is watching.
  3. Say it. The status line reads 5 windows · same while several of those windows are retentions of a machine the guest could not observe. The reducer already knows — freshness == .expectedStale, actionable == false, baseComplete == false — and NOWMirrorSource.swift already has the vocabulary for exactly this (" · Apple menu expected-stale"). Windows have no equivalent. Do this one regardless of which of the other two wins, because a mirror that cannot see the machine must say so rather than keep drawing.

The fixtures for all of it are in the tree: now-host/Tests/HostTests/Fixtures/scene-quit-modal.json (the modal, as the guest sent it), and scene-plane-held.json / scene-plane-lapsed.json (the same machine, one poll apart, either side of the lapse). MirrorQuitModalTests pins what this side does with them; its staleness assertion has been watched to fail.

How to open a control panel over the wire, since it costs an hour to find. launch refuses one — not an application (type APPC). The route that works is the script verb:

tell application "Finder" to open file "Date & Time" of folder \
  "Control Panels" of folder "System Folder" of startup disk

({"source": "...", "timeoutMs": 25000} — the argument is source, not text.) On a clone whose time zone has never been set, that alone raises the Set Time Zone modal, so the case needs no quit to reproduce.

FIXED and measured, 2026-08-06 (4b972ade). Candidates 1, 2 and 3 all landed, because 1 and 2 cover different halves of Michelle's sequence and only together cover it. Same instrument either side — tools/local-plane-lapse.py, whole documents, a control panel open throughout, and it refuses a guest whose build is not the one asked for:

gap of total wire silence before (c5c39f61dbbf) after (53cbc0fc5dcb)
first scene of the connection NOW's window only the machine
3 s, 8 s the machine the machine
12 s, 20 s, 25 s NOW's window only the machine

And the reported frame end to end: twenty seconds of silence, then tell application "Date & Time" to quit, and the FIRST walk after it carries Set Time Zone. Before, that walk was the blind one.

  • Renewal rides host traffic, not the event loop (wire.c :: renew_scene_planes). A host sending acts wants the planes; the event loop runs whether or not anyone is watching, and an armed plane is work charged to every process on the machine. Two gates keep it honest: only after a scene has been asked for on THIS link, and never on pong, which is the reply to our own heartbeat and would make "connected" mean "armed forever".
  • A scene waits for the arm echo (peek.c :: now_peek_settle), bounded at half a second and returning the moment it lands. It never asserts an arm it did not observe — no resident, no writer, or a passed deadline all fall through to a walk that reports not-observed. This is also why the FIRST scene of a connection is no longer blind; the warming ritual tools/local-arm-latency.py documents is now unnecessary (its docstring is stale, and is left as the record of why it existed).
  • The status line distinguishes seen from retained: 5 windows, 3 expected-stale · walk … · same. MirrorQuitModalTests pins it against the two lapse fixtures, watched to fail.

What is NOT closed by this. A host that goes entirely silent for longer than the lease still lets the planes lapse — deliberately, that is the lease doing its job — and the recovery now costs latency (one settle) rather than a scene. And the arm handshake itself is unchanged: nothing here makes a plane arm faster, only stops asking before it has.

FIXED in the host, UNVERIFIED by any drive: Apple menu items did nothing, and the act was never the missing part (2026-08-06)

Michelle, driving: "apple menu items dont work (apple menu -> control panels, sherlock, system profiler etc)". The menu drew correctly and selecting a row did nothing on the machine.

The act was sent, reached the guest, and was dispatched. acts.log carries all three attempts, each planning menuCommand(menuID: -16383, itemIndex: 3/6/14, titleLeft: 10) and answering in 18–50 ms with settlement unknown — the self branch of menuact, which queues the choice for NOW's own main loop and returns without settlement rows. Compare the Finder's File > Quit on the same drive: ~800 ms, dispatched-but-unconfirmed, the foreign act plane. Two mechanisms, and the fast one ends at main.c :: handle_menu_choice, which serves menus 129, 140 and the rest of NOW's own bar and has no Apple-menu case at all. The choice fell off the end of the switch.

Fixed by routing every row below the Apple menu's first separator to openAppleMenuItem, which asks the guest Finder to open the file by name — the mechanism that already existed and was wired to Key Caps alone. ObjectResolver.isAppleMenuItemsEntry makes the decision where the sibling rows are visible, and MirrorDriveService calls the same function so the agent face cannot drift.

No drive has watched an Apple menu row open. Verified only by mutation: restoring the Key-Caps-only rule reproduces the three logged plans byte-for-byte.

Two things it does not fix

  • A folder alias will time out having worked. openAppleMenuItem predicts processNamedPresent(name), which is right for Sherlock 2 and Apple System Profiler and wrong for Control Panels — an alias to a FOLDER, which opens a Finder window. That burns the full 15 s settlement holding the one mutation lane, the same shape MirrorActionExecutor.finderOpen documents and fixed for itself by classifying the item. Here the item cannot be classified: every Apple Menu Items entry is an alias, and an alias reports its own kind and never its target's. It needs either a postcondition meaning "a process OR a Finder window by this name", or a guest probe that resolves an alias. Functionally the open works; only the reporting lies.
  • At the machine, the same click still does nothing. This is a host-side route. A person sitting at the guest with NOW frontmost who chooses Apple > Sherlock 2 gets the same silence, because handle_menu_choice still has no Apple case and OpenDeskAcc is not in CarbonLib. Whether the guest should serve its own Apple menu — and through what, since the obvious call is absent — is open, and it is a command-parity question, not a Mirror one.

BROKEN, and it is NOT the host's click translation: a modal's Cancel is refused because the planes went inactive (2026-08-06)

Michelle, same drive: in Date & Time's Set Time Zone modal — which renders — "Cancel doesn't work", while an earlier session's ditemact sent straight over the wire had dismissed that same button.

That pairing reads like a host translation defect and is not one. The host resolved the click correctly and sent the right act: plan=dialogItem(ref: "now-element-5f5c3825-…", item: 2), the same verb and the same item number the wire test used. The guest refused it:

element-not-found: the anchor plane is absent or not armed

after 6.7–8.8 s, five times over. The actmeta line at that moment says why: requested=8 active=8, with structure=inactive semantics=inactive interaction=inactive and only content active. The planes the act needs had gone down; requested had dropped from 15 to 8.

So the open question is what takes the structure/semantics/interaction planes inactive while the content plane stays up, and whether a modal being frontmost is what does it — which would tie this to the existing "one modal wedges the whole Mirror" entry below. Until that is answered, do not attribute a dead-looking Mirror click to hit-testing without reading actmeta at the same timestamp first: the refusal was recorded plainly and was still nearly diagnosed as the wrong half.

2026-08-06, later: the open question is answered, and it was not the modal. Nothing took those planes down — the host went quiet for longer than the ten-second lease, because optional planes were holding the scene cycle open for 12.5 s at a time. The modal only made the Finder slow enough for it to show. See the decode_ms entry at the top of this file; it is unfixed, and deliberately so pending a design decision.

2026-08-06, later still — the guest half of it IS fixed, and this entry and the scene-gap entry above were describing one defect from two ends. They gave different proximate causes for the same lapse — that one says "thirty to forty seconds of acts with no scene", this one says "12.5 s cycles" — and both are true of the same mechanism: only scene.request renewed the ten-second OWNER lease, so any reason the host stopped asking within ten seconds disarmed the planes, whether the reason was act cadence or a stalled cycle. 4b972ade makes renewal ride any inbound host frame, so a live host no longer lapses the lease at all, and a scene waits briefly for the arm echo rather than walking blind. That covers the Cancel refusal.

What it does not close is the decode_ms inversion itself: the host can still spend 12.5 s in a cycle, and a lease that no longer lapses is a guest making the best of a host that stalls. Both halves are owed. Emulator only; never on metal.

MEASURED, and the answer is DON'T BUILD ON IT: how long the content plane takes to arm (2026-08-06)

Plan 013 § A proposes turning the resident into a NOTIFIER — told at mutation time instead of asked on a timer. The content plane's arm handshake is that idea already shipping, so it is the one place the proposal can be costed instead of argued: the application writes a request, and it is honoured only when the TARGET process next runs the resident's jGNE pass and agrees it is the one named (ext/src/now_content.c :: now_content_gne). The host's status line says content: requested X's trace; waiting for its event loop to arm, and nobody knew whether that sentence covered 40 ms or 4 s.

The conditions. One session-private clone off os91-runner.qcow2, build 9a1d885fcac0, wire 5450, anchor 1702, tools/local-arm-latency.py. Time from the qdtrace start reply to the first qdtrace status reporting that generation and that A5 live.

condition arm wire floor, SAME condition n
NOW's own window, NOW front 22 ms (20–31) 102 ms 10
Finder's window, Finder front 116 ms (102–117) 101 ms 10
Finder's window, NOW front never, 25–45 s 57 ms 9
…then brought forward mid-wait 40 ms after the front change 56 ms 6
front application holding a modal 30.1 s 30.0 s for ONE status call 1
an A5 no process has (control) never 8

Read the floor column first, because it is the whole finding. A bare qdtrace status round trip costs ~101 ms when NOW is in the background — that is NOW's own event-loop period, not the network. So the 116 ms against a 101 ms floor means the frontmost Finder pumps within about 15 ms of the request, and the 22 ms self case is complete before the first status poll can be asked and is not measurable from here at all. The arm handshake is fast. The sentence a person reads is mostly the host's poll cadence.

The first version of this measurement reported the 116 ms as the arm cost, against a floor taken once at startup in a different front condition. It is the same mistake plan 013 records against its own first finding — two variables moved and the differential was read as one — and it was caught only because the floor was re-measured per condition.

The two large numbers are not handshake problems, which is why nothing here justifies a mechanism:

  • A background application never arms at all. Not slowly — not at all, over 45 s. The unguarded refusal counter says why, and it is the useful half: wrongContext climbed ~1,770 in 25 s, so processes WERE pumping and reading the request the whole time and none of them was the Finder (measurement rule 6 — "never entered" and "entered and declined" are opposite repairs). The deferred run settles it with one variable moved and the same arm: bring the Finder forward without re-requesting and it arms in 40 ms, 6/6. A resident notifier cannot fix this. Nothing can make a process pump that the Process Manager is not running.
  • Under a modal the wait is the modal. tools/guest-wedge modal 30 and the arm landed 30.133 s later — the wedge's own duration. The first qdtrace status after the launch took 30.023 s, so the host could not even ASK for 30 s. That is plan 012's liveness territory, already owned, and it is not about arming.

DECISION: do not build the § A notifier on this evidence. The host arms only the front window's process (NOWMirrorContentPlane.join takes scene.windows.first(where: \.front)) and re-arms on target change or after renewAfter = 9 minutes against a 10-minute TTL. So the measured cost is ~15 ms of real handshake, paid on a focus change and roughly never otherwise, on a machine whose whole scene walk is now 3–8 ms. There is no mechanism that pays for itself here. Per plan 013's own tiering this would have been tier 2, and the plan prefers re-arguing to executing — so it is re-argued and dropped.

What this does NOT settle.

  • Every number is from mac99. The dominant term is a Process Manager round, so it should scale with process-switch cost rather than with how much interface exists — but that is reasoning, not a reading, and a 1400c has never been asked.
  • "Just launched" was never measured. launch refuses a control panel (launch-refused: not an application (type APPC)), which is the case the observation actually named, and an arm needs an exact window address out of the scene — so a just-launched application cannot be armed until it has already pumped enough to open a window. The product can never request an arm earlier than that.
  • The modal row is n=1, and it is the wedge's GetNextEvent loop, not a real application's modal.

2026-08-06, later: that caveat was exactly right and is now measured. The wedge's GetNextEvent loop starves NOW completely (43,974 ms of a 45 s run, pass max 44.9 s) and a real application's ModalDialog does not starve it at all (413 ms median, 145 probes). So this row's 30.1 s is a true reading of the wedge and says nothing about a modal. See the entry at the top of this file. - The plane measured is the one on claude/mirror-thread-content. The GWorld work on claude/gworld-probe-grounding-3dcbd4 adds ~600 lines to ext/src/now_content.c, four hunks of them inside now_content_gne itself. now_content_arm_verdict is untouched, so the mechanism survives, but the per-pass constant does not necessarily. Re-measure after that lands.

SHIPPED on an emulator, UNVERIFIED on metal: a scene can now answer "the same", or send only what moved (2026-08-06)

Plan 013 § 5. The wire had become the dominant cost — a 3–8.5 ms walk against a 111–710 ms transfer of a 26–28 KB document, several times a second, for a Macintosh that mostly had not changed. scene.request now takes a since (the body digest the host already holds) and gets one of three answers: scene.same (a control frame, no transfer at all), a delta carrying only the entities that moved, or a whole document. Design: scene-deltas.md. Numbers: scene-delta-measurements.md.

What is proven. Tested, on a session-private mac99 clone, guest df570d8014de: an idle machine's wire cost drops to 10.3–10.4% over ten polls, and after the first scene each poll costs one control frame. Walk time did not move (0 ms idle, 16 ms driven, both conditions). The byte-exact reconstruction and the digest check are pinned natively on both sides, including a two-halves test whose fixtures the GUEST'S encoder emitted.

What is not.

  • No metal pass. And the emulator understates this one rather than overstating it, which is the reverse of plan 013's usual warning: its network is host loopback, so the byte term this change removes is nearly free here and expensive on a 1400c. The idle saving should be worth more on the hardware, not less. That is a prediction.
  • The driven case saves ~5%, because changing the frontmost application rewrites every app row, every window's front/z and the whole menu bar. That is honest worst-case behaviour and the guest correctly still picked the smaller of the two, but nobody has measured the realistic middle — one window moving on an otherwise still machine — on real hardware or with a real person driving.
  • NOW-68K serves none of it, because it serves no scene at all. The asymmetry is declared in contract-coverage.md; the delta design is sized for that machine (a few kilobytes of state, a 32-bit hash), so what it lacks is the walk, not the room.
  • The host applies a delta and then throws the result at the unchanged reducer. That is deliberate — one deletion rule, in one place — but it means a rebuild is a whole document's worth of decode every time. Nothing measures whether that decode is now the expensive half.

An observation the measurement produced, and it is not about deltas

Revealing two different Finder folders alternately, once per poll, produced nine scene.same answers out of ten: the walked scene was byte-identical each time. Either the Finder did not reorder its windows, or the walk does not see that it did. scene.same is a rather good change detector and it has just detected something. Worth a second look; not chased here.

FIXED on an emulator, NEVER ON METAL: both things the control sweep got wrong, from one change — the guest stopped hunting for controls it had made itself (2026-08-06)

Plan 013 § 2. The entry below names two costs, a 1.9-second focus change and a background sweep that reported zero controls, and treats them as one performance item and one correctness item. They are one defect, and its root is a single sentence in Inside Macintosh: FindControl refuses an INACTIVE window. Refusing means answering nothing, so with anything else in front the sweep walked all 3,724 points, found nothing, cached nothing, and re-swept on the NEXT poll forever — and the cache was therefore EMPTY at the exact moment a person clicked into NOW, which is when a probe costs ~240 µs instead of ~2.7 µs.

The fix is that the application already knew the answer. Every control here goes through now_control_new (workshop/control_kind.c), which existed so the scene could report a role. That table is now the scene's LIST of a window's controls, and the lifecycle is closed around it: now_control_adopt for the DataBrowsers (a constructor that takes no procID, so it cannot go through the wrapper), now_control_dispose, now_control_dispose_window and now_control_dispose_dialog — because DisposeWindow destroys a window's controls and tells nobody. control_kind_source_test.py gates all five calls, and both of its refusals were watched by mutation.

Only EXISTENCE is remembered. Title, bounds, value, enabled and visible are read live every pass, and invisible controls are skipped exactly as the sweep skipped them — so a Workshop page switch needs no invalidation at all, which also settles the question the § 1 measurement left open about whether one bumped the generation. Nothing is cached, so nothing can be stale, which is the opposite of the risk plan 013 warns about and is worth saying plainly: the projection is rebuilt from the Toolbox on every scene.

Same clone, same session, one variable — which binary is staged. Guest builds 2c5dc9cb7d54 (after) and the merge base (before), wire 5430, tools/local-scene-bench.py reading meta.phases, µs:

condition before, controls after, controls controls reported
NOW front, steady state 3,217 – 4,890 713 – 989 9 → 9
the scene NOW becomes front on 886,398 713 9 → 9
NOW backgrounded 7,206 – 7,588 477 – 1,324 0 → 9

The focus-change scene measured 886 ms here against the 1,891,174 µs in the entry below; a parallel instrumented run the same night saw 1.21–1.43 s. It is a wide range, and every value in it is the same defect. The focus cycle was repeated three times on the build under test and the controls phase stayed between 727 and 922 µs with no spike.

The correctness half, which was the worse one. Backgrounded, the scene now carries NOW's nine real controls instead of an empty window — an absence nobody had observed, which the coverage rules forbid. Where the registry cannot answer (a Dialog Manager window this application did not build, seen while inactive) the control plane is RETRACTED rather than swept: the key is absent and meta.errors carries the notice, because "nobody could look" and "there is nothing there" are different facts. That retraction path is built and NOT observed — it needs an inactive DITL dialog and nothing in this session opened one.

Page switching was checked as a positive control rather than assumed, over the wire with menuact on the View menu: Preferences 5 controls, Processes 6, Screenshots 9, each correct for its page, each fresh. Processes shows six because its DataBrowser is now adopted and therefore mirrored — it carries role: dataBrowser, where before it fell through to the emitter's range guess.

What is not done. Nobody has watched any of this on the PowerBook 1400c; every number is QEMU. The sweep still exists for windows this application did not build, and on that path everything the entry below says still holds.

OPEN, and now MEASURED rather than argued: where a scene's time goes, and the two things the control sweep still gets wrong (2026-08-06)

Plan 013 § 1. Every scene now carries meta.phases — microseconds per phase, named for what the guest was doing — so the entry below, which was written from a THROWAWAY breakdown, no longer needs one. The breakdown is permanent, additive on the wire, and publishes its own cost. This entry is what the first measurement with it says.

The conditions, one clone off now-mirror-stage.qcow2.bak-20260806, build 9ed6e7d18c19, wire 5410, control-sweep fix merged. Median of the STEADY-STATE scenes in each condition, microseconds:

phase Finder front NOW front
enumerate 952 1045
bind 182 214
windows 696 970
controls 290 5090
menubar 107 1359
semantics 20 17
refs 153 225
encode 589 596
whole walk ~3.0 ms ~8.5 ms

Against the 116 ms / 1116 ms this plan opened with. The sweep fix holds, and it holds in both conditions.

Two costs the fix does not remove, both visible only now.

  1. A focus change costs one 1.9-second scene. Measured: the scene in which NOW became frontmost reported controls = 1,891,174 µs. The cache is invalidated by the activation and the whole 3,724-point grid is re-swept, in the foreground, at ~240 µs a point. Every subsequent front scene is ~5 ms. So the cost did not go away — it moved from every poll to every focus change, which is a very large improvement and still the single most expensive thing on this machine.
  2. The background sweep is pure waste, and it is also a HOLE IN THE MIRROR. FindControl answers an inactive window immediately — which is why the sweep is cheap in the background — but "immediately" means it answers NOTHING. Measured: with the Finder in front, NOW's own window walk spends 5–10 ms sweeping 3,724 points and returns zero controls. The scene therefore shows NOW's own window as empty whenever NOW is not frontmost. That is not a performance issue with a performance fix; the mirror is reporting an absence it did not observe.

Both of those are FIXED (2026-08-06, emulator only) — and they were one defect, not two. See the entry above. What is left standing here is the diagnosis, which is worth keeping for its shape: the two were filed as a performance item and a correctness item because that is how they present, and they share a root (FindControl refuses an inactive window) and a cure (do not discover what you already know you made).

The wide answer plan 013 asked for: yes, the shape is everywhere. Every remaining phase is a full re-derivation, per poll, of something that rarely changes — the process list (enumerate, ~1 ms and identical in both conditions), the window chains (windows), NOW's own menu bar (menubar, 1.36 ms), and the document itself (encode). Nothing in the steady state is proportional to what CHANGED. The control sweep was not the only one; it was the loudest. But the absolute numbers are now single-digit milliseconds on an emulated G4, so the case for plan 013's slices 3–5 no longer rests on comfort here — it rests entirely on the vintage-hardware multiplier, and that argument should be made with a number from a 1400c rather than assumed.

What the breakdown costs, stated because it must be. 130–330 µs per scene on this machine, from 26–66 Microseconds calls at a calibrated ~4.9 µs each — the emulator's trap cost, and it will be lower on real PowerPC. That is 0.02% of the 1.9 s scene and up to 8% of a 3 ms one; it is bounded by PROCESSES and WINDOWS, never by controls or menu items, so it does not grow with the size of the interface. It is left on because a breakdown that is off by default is a breakdown nobody has when they need it, and every scene publishes phases.clockUs so a reader can subtract rather than wonder. If a metal pass shows Microseconds costing what it costs here, the seam count is what to reduce — the per-process bind/windows/menubar trio is 6 of every 8 clock reads.

FIXED: the PowerPC guest never reads the host's contract revision (2026-08-06)

Fixed and watched on an emulated Power Mac G4 the same night, guest build 48a2af200ab7 2026-08-06T06:54:16Z, on a session-private clone (wire 5421). A host answering contract: 1 is refused with {"type":"refuse","contract":2,"reason":"contract revision 1 != 2"} and the guest closes; a host answering no contract at all is refused with "host hello states no contract revision; this guest speaks 2"; a host answering 2 is served, and between all three the guest redialled on its own backoff without being asked. The permanent check is WireLimitsAgreementTests .testBothGuestsGateTheContractRevisionInTheirHelloHandler, which reads both guests' hello handlers and was watched failing with the old on_hello pasted back in.

The open question this entry ended on — whether a guest should SEND a refuse — is now settled in the contract, not in one guest's habits. contract/asyncapi.yaml's connection rules say the gate binds whoever RECEIVES a hello (both halves of it), that an ABSENT contract is a mismatch rather than a tolerance, and that the refusal is sent and names both numbers: a silent hang-up is indistinguishable from a dropped network at the far end, which is the one thing gating at the door exists to tell apart. Two implementations moved to meet it — NOW-68K, which refused silently with a bye and a status reading "contract mismatch", now sends refuse naming both numbers; and the host, which dropped an inbound refuse into a default: arm, now logs it and finishes, since "never swallowed" is that schema's own word.

The original entry, unedited:

FIXED: latencyMs was never a measurement — it could not see anything shorter than a tick (2026-08-06)

Recorded 2026-08-06 during the durability pass, because it is the instrument behind two of the wrong answers already in this file and it had no entry of its own. The defects it caused are written up; the defect in the instrument was not.

The mechanism. scene_collect.c computed the scene's only timing number as

out->latency_ms = (long)((TickCount() - t_start) * 1000UL / 60UL);

TickCount() advances at 60 Hz. The whole field is therefore quantised to ~16.6 ms, and anything faster than a tick reads as 0 or as 17 — not approximately, but identically, for every phase of a walk that we now know runs in hundreds of microseconds. A scene that spends 290 µs in controls and one that spends 5,090 µs in controls were the same number.

What it cost. Every question about guest cost was answered by differencing two runs of this one field, and the differences were inside the quantisation more often than not:

  • The menu-bar misattribution. A ~1 s self-front scene was confidently blamed on collect_self_menubar. The microsecond breakdown then put the menu bar at 0.1% of the walk and a FindControl grid sweep at ~95%. Written up above, at "NOW's own window cost ~1 s of every scene, and the suspect was the wrong one".
  • The 115 ms round trip could not be decomposed at all. With a tick-resolution clock the sleep, the notice delay and the work were one undifferentiated number, and the entry above ("the 115 ms round trip was the guest's own sleep") only became answerable once wirestat grew histograms of the guest's own service interval and of Open Transport's announce-to-read delay. notice at a 48.5 ms mean is three ticks; it is not a thing this field could ever have reported.

The repair is a better instrument, not a better estimate. meta.phases — microseconds, eight non-overlapping phases, permanent and additive on the wire rather than bolted on for one investigation. It calibrates the clock's own cost at startup and publishes it (clockUs), so it says what it costs rather than asking to be believed, and its seams are placed where their count is bounded by processes and windows, never by controls or menu items. scene_phase.h carries the reasoning; scene_phase_test.c drives the arithmetic on a host compiler with an injected clock.

latencyMs still ships, because a consumer may hold an old reader, and it is excluded from the scene digest along with seq, capturedAt and walkMs (scene_digest.h) — a clock reading is not part of what the scene is. Do not reason from it. If a number matters, read meta.phases.

The transferable rule is rule 14 in mirror-measurement-method.md: a clock cannot measure anything shorter than its own tick, and a clock's resolution belongs beside its first number. TESTED; the phase clock has never run on metal, where Microseconds is a real trap on a real 68K/PPC bus rather than an emulated one.

FIXED on an emulator, NEVER ON METAL: NOW's own window cost ~1 s of every scene, and the suspect was the wrong one (2026-08-06)

The symptom. With NOW frontmost — which is what a person does the moment they click its window — every scene poll took roughly a second. With anything else in front the same machine answered in 16 ms while reporting MORE (8 menus / 82 items against 7 / 48). The self path did less work for ten times the cost.

The suspect, and why it was wrong. The only step that runs exclusively when NOW is frontmost is collect_self_menubar (scene_self.c) — the menu bar collected through the Toolbox, where a foreign machine's bar is read through the validated memory reader. It looked expensive too: root_items_for() rescans the root menu once PER MENU, so ~7 full rescans a scene, and the Apple menu's 16 items are backed by a folder on disk. That is an inference from two conditions and it survived only until somebody timed the function.

Timed directly (a temporary meta.dbgUs breakdown, Microseconds, one fresh clone, six scenes each):

step NOW frontmost Finder frontmost
collect_self_menubar, whole 1.0 – 2.5 ms not run
root_items_for, all 7 menus 0.30 – 0.43 ms
add_one_menu, all 48 items 0.63 – 1.2 ms
find_controls_by_probe 875 – 925 ms 10 – 22 ms
latencyMs 866 – 1133 16

The menu bar is 0.1% of it, on this machine, with no sign of the disk cost the Apple menu could have carried. The cost is the FindControl sweep over NOW's own window, and it hides from the obvious A/B for a documented reason: FindControl answers an INACTIVE window immediately. The identical 3,724-point sweep costs ~2.7 µs a point in the background and ~240 µs a point in the foreground.

And the foreground figure is not one number, which is its own finding. ~240 µs a point is the SETTLED window. A separate run that counted the FindControl calls rather than dividing a total by an assumed point count measured the ACTIVATION scene — the one where a person has just clicked into NOW — at 335–397 µs a point, roughly 40% dearer, because the window is being activated and redrawn while it is probed. Both numbers are real and they are not interchangeable: the settled one describes a steady state nobody waits on, and the activation one describes the scene a person actually feels. A durability pass on 2026-08-06 flattened the two toward 240 and the distinction is restored here, because collapsing them deletes the reason the focus-change scene was the worst case.

The fix (scene_self.c, control_kind.c): cache the sweep's DISCOVERY — which controls this window has and the point each was found at — and re-prove it every pass at one FindControl per control. The rows themselves are still rebuilt from the live Toolbox every scene, so nothing a person changes goes stale. A control that went away, moved or was hidden fails its point and the sweep runs again; a control that ARRIVED disturbs nothing cached, so now_control_generation() catches it instead — every control this application makes goes through now_control_new, and control_kind_source_test.py enforces that. A cached ControlRef is only ever compared, never dereferenced, until FindControl has answered with it.

Same clone, same build discipline, guest's own latencyMs:

condition before after
NOW frontmost, steady state (10 scenes) median 916 ms median 0 ms (9 of 10 at 0)
Finder frontmost (8 scenes) median 16 ms median 0 ms
first scene after a Workshop page switch ~900 ms 250 – 1550 ms, then 0

Scene contents are unchanged (7 menus / 48 items either way), and the cache was verified by driving View across three pages over the wire: identical within a page, different across one, correct on each.

What is NOT done, and what would change the answer.

  • Nobody has watched this on the PowerBook 1400c. Every number here is QEMU. The direction should hold — the fix removes ~3,700 Toolbox calls per scene and adds ~7 — but the size will not.
  • The first scene after any control change still pays the full sweep, 250 ms to 1.5 s depending on how many controls the page has. That was paid on EVERY poll before; it is now paid once per UI change. Reducing it means reducing the sweep itself, and the honest options are a coarser grid (which would miss the 12pt disclosure triangle) or giving these windows a root control (which reshapes the interface being described — rejected once already, see scene_self.c).
  • The menu-bar optimisations were deliberately NOT taken. Resolving the root menu once per scene and change-detecting the MenuList are both correct and both worth ~1 ms here. They are not free: NOW's menu bar rendered EMPTY after a Hide on 2026-08-05, so menu freshness is a live defect surface, and spending it for 0.1% is a bad trade on the evidence we have. If a metal run shows the Apple menu's folder-backed items cost real time there, that is a different finding ("this cost is disk-shaped") and points at not re-reading unchanged items at all.

BROKEN, contract violation: the PowerPC guest never reads the host's contract revision (2026-08-06)

contract/asyncapi.yaml, connection rules: "contract is a single integer revision. Unequal revisions => refuse." Three of the four implementations do that. One does not.

side on an unequal revision in the peer's hello
host (Session.swift:1511) refuses, reason names both numbers
NOW-68K (wire68.c :: handle_host_hello) logs it, set_status_str("Protocol error: contract mismatch"), teardown_and_retry(NULL, "protocol-error")
scripts/probes/nowwire.py :: _gate refuses
NOW-PPC (wire.c :: on_hello) never looks at the field. It reads name and version, sets kConnConnected, and serves the session

on_hello is twenty lines and contract is not among them; the handshake dispatch above it (handle_frame, g.phase == kConnHandshaking) routes hello straight there without checking either. So the PowerPC guest will hold a full session with a host speaking any revision at all, including one that predates every message it is about to be sent.

How it surfaced. Five Python harnesses declared the revision by hand and two of them still said 1 (fixed on claude/wire-revision-drift, contract/wire_limits.py plus three gates in WireLimitsAgreementTests). Those two had been talking to NOW-PPC guests perfectly happily — which is exactly the problem: the guest's missing check is what made a year-stale harness look fine. Against NOW-68K the same harnesses could never have held a link, and that difference is the whole diagnosis.

Why it is worth more than the tools fix. The check exists to make a version skew fail loudly at the door instead of quietly in the middle of a message nobody can decode. A guest that skips it converts "refused, here is why" into a session that misbehaves later with no handshake to blame, and that is the two-halves-never-met-in-a-test shape.

The fix, not taken here because this branch is the tools half and a guest behaviour change wants a metal pass: read contract in on_hello, and on a mismatch set a status naming both numbers and tear the connection down the way handle_frame already does for refuse (return 0). NOW-68K's handle_host_hello is the model — including its treatment of an ABSENT contract as a mismatch, since the field is required.

What is not known. Whether the guest should also SEND a refuse before closing. The contract words the refusal as the host's move ("the host answers hello (accept) or refuse and closes") and never says what a guest does with a bad host hello, so NOW-68K's silent teardown and a hypothetical guest-sent refuse are both defensible readings. AGENTS.md says the file family is symmetric; if that governs here, the contract's prose should say so explicitly rather than leaving two implementations to guess. Settle the prose before writing the code.

FIXED in the host, UNVERIFIED by any drive: the host app reported as always-on-top; no window level exists to cause it (2026-08-06)

(Re-headed 2026-08-06 during the durability pass. It stood as "UNSETTLED" while its own body recorded a landed fix — the ledger's two words are BROKEN and UNVERIFIED, and a third one invented for a single entry is how a fixed thing goes on reading as an open question. The fix is real and the verification is genuinely owed, which is exactly what "FIXED in the host, UNVERIFIED by any drive" says.)

Michelle: the macOS app window floats above other applications. What was measured, without driving anything:

  • No window level is set anywhere. NSWindow.level, a floating panel, orderFrontRegardless, collectionBehavior, LSUIElement: none of them appear in now-host/, in the vendored mirror/ package, in the Xcode target's generated Info.plist, or anywhere in this repository's history on any branch.
  • The running windows are at level 0. CGWindowLayer read from the window server for every window of both copies running on the desk — the human's (pid 27820) and a freshly built one — was 0 (NSNormalWindowLevel). A level-0 window cannot stay above another application's windows once that application activates, so "floats" is not what the window server is being told to do.
  • The one lever that does steal the front is NSApp.activate(ignoringOtherApps: true) at App.swift:401, which runs on EVERY route into openMainWindow(): launch, applicationShouldHandleReopen (a Dock click, or any open of the bundle), the status item's "Open New Old World", ⌘, for Settings, and every module item in the Windows menu. ignoringOtherApps: true puts NOW in front of whatever application the person was typing in, rather than waiting to be switched to.

Fixed by narrowing the routes, not by dropping the call. Simply asking politely would not do: from a background app activate(ignoringOtherApps: false) does nothing at all, and a status item's menu does not activate its own application — so "Open New Old World" would open the window behind everything, which is the shape of the two regressions this file already carries comments about. So the forceful activation stays, and moves to openMainWindowFromOutsideTheApp(), called by exactly the two routes where a person asked for this window from outside the app: the status item, and applicationShouldHandleReopen.

Launch and the in-app menu routes (⌘, for Settings, every module item) now open the window without activating — the menu routes are already active, so they lost nothing, and a launch the person performed is activated by macOS itself while one they did not perform stays put.

What should change: the app no longer pulls itself in front when it is launched or relaunched in the background, and no longer re-takes the front from another application on its own. What should NOT change: clicking the Dock icon still raises the main window, "Open New Old World" in the status menu still brings the app forward from whatever you were in, and Open Mirror still leaves the mirror window in front (NOWMirrorWindow.show still does not activate, precisely because activating trips reopen, which raises the main window over it). Michelle is the verification step here — this was not driven, by her direction.

FIXED: hiding NOW leaves it frontmost, and Windows > Workshop then times out having worked (2026-08-06)

Reported by Michelle from a Mirror drive; measured on a session-private emulator clone the same night, guest builds bf4987c6eca1 (before) and a14d111103f8 (after).

What hiding NOW actually does. hide "New Old World" — and the Application menu's own Hide, which is the same Process Manager call — hides the window and does not move the front process. Measured, three scenes over six seconds and a QMP screendump of the same moment:

  • apps[].front stays New Old World;
  • axsnap agrees: front: New Old World, front: true;
  • the machine still draws NOW's menu bar (Apple, File, Edit, View, Windows, Help, and the New Old World application menu), over a desktop with no NOW window on it;
  • the scene's window row for it is visible:false, front:false, so the scene has no front window at all and the content plane says content: no front window.

Why Windows > Workshop then timed out. The menu item is still reachable in that state, and menuact 140/1 still dispatches through the application's own main-loop queue — but workshop_open only called SelectWindow, and selecting a hidden application's window shows nothing. Watched: four scenes over twelve seconds, the window visible:false, front:false throughout. The host had planned windowFront(New Old World) from MirrorActionExecutor.presentOrCreated, which that state can never satisfy, so the act burned its whole 15 s timeout having been dispatched correctly — and the mutation FIFO is one lane, so everything behind it waited too. Same shape as the Finder-open entry below, different cause.

The fix is now_proc_show_self() in workshop_open: show this application before selecting the window. Every route that promises a Workshop page comes through there. Watched pass: hidden, then the same menuact 140/1, and the window was visible and front 5.6 s later, with the Workshop drawn on the machine's own screen.

What was NOT settled: the menu bar reported empty after the first Hide. In the same report Michelle saw NOW's menu bar render EMPTY in the Mirror after the first hide, restored by cycling to the Finder and back. That could not be reproduced from the wire. In every hidden-and- front state produced here the guest reported menubar.app: New Old World with all seven menus and coverage menubar/…/complete, the machine drew them, and feeding those exact IR documents through the host's own MirrorScene.decode + MirrorReplicaReducer kept the menu bar with actionable: true. So neither the guest's report nor the host's retention is dropping it in that state.

The plane-bind hypothesis is DISPROVEN (2026-08-06, second pass). It was worth testing because axsnap does report bind: no-plane, hasMenus: false with NOW freshly front and bind: ok, hasMenus: true a few seconds later — so there IS an unarmed window, and "empty until you cycle away and back" is what one would look like. It reads axsnap and a scene at the SAME moment, 28 paired samples across the four phases the report names — fresh connection, the first hide, away to the Finder, back to NOW:

phase axsnap bind scene menubar
fresh, first pass no-plane, hasMenus: false New Old World, 7 menus, coverage complete
fresh, passes 2-8 ok New Old World, 7 menus
the first hide, 10 passes over 32 s ok New Old World, 7 menus
the Finder, 4 passes ok Finder, 8 menus
back to NOW, 6 passes ok New Old World, 7 menus

The unarmed pass is real and it lasts about three seconds — and the scene taken in that exact moment still carries the whole menu bar. That is the shape of the thing: scene_self.c reads the live MenuList through the Toolbox from inside our own process, and the anchor bind is the FOREIGN memory reader. They are not the same source and the self path does not wait for the plane.

And the RENDERER is cleared too, offscreen (third pass). RenderShot rasterises the same SceneRenderer the Mirror's window uses, with no screen involved, so the last stretch could be tested after all. The two live documents are now fixtures — now-host/Tests/HostTests/Fixtures/now-scene-self-hidden-but-front.json and …-front-visible.json, captured three seconds either side of a hide on guest build bf4987c6eca1 — and MirrorMenubarRenderTests renders them and counts the ink in the menu-title band:

document ink in the title band
hidden-but-front (her exact state) 705
front and visible 705
the same document with menubar stripped, as the negative control 77

So the renderer draws all seven titles for the document she was looking at. The 77 is worth knowing on its own: it is the Apple glyph shouldSynthesizeAppleMenu falls back to when a scene carries no menus at all — a Mirror bar showing an apple and nothing else is exactly what "the menu is empty" looks like, which says the scene the window was drawing had no menubar even though the one that arrived did.

What is left, and what would tell them apart. Every static stage is now cleared by a test, so the remaining candidates are all live state in the running app, and none of them is worth guessing between:

  • The window was drawing an older projection than the document that had just arrived. NOWMirrorSource.scene is shadowEngine?.snapshot?.scene ?? decoded, so a projection that had not taken the new observation would be drawn instead of it.
  • A plane toggle. planePolicyDidChange republishes from the engine's snapshot; a scene projected under a different plane policy is a different scene.
  • Nothing retained yet. reduceMenubar retains a previous record when no menubar claim is complete — retention can only KEEP a bar, never empty one, unless the first scene of that Mirror session had none.

The instrument that separates them already existed and said nothing: MirrorEngineDiagnostics compares the visible scene against the engine's projection on every cycle and names the disagreeing field — including menubar — into a 64-entry ring that was never logged, never shown in the Diagnostics pane and never exported, so its answer died with the process. It now writes each difference to HostLog (mirror area, warn level), which means the NEXT reproduction leaves a line in ~/Library/Logs/ saying whether the projection and the arriving document disagreed about the menu bar at that second — the first candidate confirmed or eliminated without anyone watching. If that line is absent while the bar is empty, the projection matched and the third candidate is the one to chase.

Her two details both point the same way and are worth carrying: it is the FIRST hide after launch and it self-heals on an application cycle, so it is transitional rather than steady — and a cycle is exactly what forces a fresh projection.

Set Time Zone: the acts reach the modal and apply; the LIST is what is missing (2026-08-06)

Michelle: "controls in some of these modals still dont work". Measured against Date & Time's Set Time Zone modal on the same clone, replies read rather than fired and forgotten:

  1. The act reaches the guest and applies. ditemact on the modal's Cancel — {element: now-element-…, item: 2} — answered Dispatch: dispatched, Mechanism: the application's Dialog Manager path, and the next scene and screendump showed the modal and Date & Time gone. A ModalDialog loop is therefore not the obstacle, and modal-vs-ordinary is not the discriminator. The refusal worth knowing: ditemact needs both element and item. With element alone it answers bad-request: ditemact requires item: a 1-based DITL number from 1 through 96, and a caller that does not read the reply sees a control that "does nothing". The host sends both (NOWMirrorSource.swift, .dialogItem).
  2. The list is not addressable at all. The scene carries the modal's 9 DITL items and 10 controls, including the list's scroll bar with max: 193 — 193 cities — and no listCells and no listTotalCount anywhere. There is no row to name, so no click on that hatched rectangle can hit one. This is row 1 of docs/mirror-element-coverage.md, now confirmed live rather than inferred.
  3. The greyed Done is TRUTH; the default ring on it is not. With the modal front the guest reports item 1 Done enabled:false and item 2 Cancel enabled:true, and the machine's own screen agrees — Done is greyed until a city is chosen, and the default ring is around Cancel. So the grey is a correct report and the ring is a renderer fidelity defect. (2026-08-06, later: half-closed, and the cause is named. isDefault is the DialogRecord's aDefItem, which is initialised to 1 and only SetDialogDefaultItem moves — so it is stale here and correct elsewhere, proven against Internet Explorer's alert the same evening. The renderer no longer rings a DISABLED item; where the ring actually WENT still needs the control's own kControlPushButtonDefaultTag. See the alert entry at the top of this file.) (With the modal's application in the BACKGROUND every item reports enabled:false, which is also true of the machine.)
  4. The truncated explanatory text is the renderer, not the content. The guest reports item 7's title complete: "The time zone must be set to determine the correct time. Select the closest city in your current time zone:" — 110 characters, in a rect three lines tall. SceneRenderer's case "staticText" draws it with one appText call at one baseline and no wrapping, so it is clipped mid-sentence. A wrap there is verifiable offscreen with the RenderShot harness the MirrorKitUI tests already use; it has not been done.

CLOSED: the resident channel dials, speaks, and holds the session through a 108-second starvation (2026-08-06)

Plan 012 § 4, and the whole plane it completes. A Macintosh — an emulated G4 under OS 9, never yet a real one — opens a SECOND connection from its optional resident component and keeps it alive while every application on the machine is starved. (Opening reworded 2026-08-06: it read "a real Macintosh now opens", which this entry's own closing paragraph contradicts. The Macintosh is real in the sense that matters to the code and emulated in the sense that matters to a claim, and only the second sense belongs in a first sentence.)

Two starvation numbers appear in the record and both are right. This entry's 108.8 s is the instrumented run measured here; plan 012's 110 s is the end-to-end run driven against the host application. They are different runs of the same behaviour, not a disagreement, and neither has been repeated on metal.

What was watched, on a fresh cold-booted OS 9 clone (mac99, guest build 4fe6d946e1a0). Two connections arrived from the same address a second apart:

[session]  {"type":"hello","contract":2,"side":"guest",...,
            "name":"Power Mac G4","os":"9"}
[resident] {"type":"hello","contract":2,"side":"guest",
            "role":"resident","version":"0.1",
            "name":"Power Mac G4","os":"9"}

The name and os are identical, and that is the load-bearing part rather than a nicety: the host associates a resident channel with its application by fingerprinting exactly those two fields plus the address, so a resident that invented its own name would be a channel vouching for nobody. capabilities: 127 — both P6 bits, the vehicle and the channel.

Then tools/guest-wedge spin 110, with tools/liveness-channel.py timestamping both connections:

what measured
the application, starved 108.8 s with no answer, past the host's 75 s window and past the 90 s Finder incident that started this
the resident, meanwhile three pings, at 86.6 s / 121.6 s / 156.6 s, every one answered
either connection dropped no

So the premise the plane rests on is no longer an argument from the scheduling model at either end. Something answering below the application kept ANSWERING ON THE WIRE while applications could not.

Then the same thing against the REAL host, which is what § 4 was for and the first time either half of § 1 had met a real guest. An application starved 110 s kept its session — the same two sockets before and after, 55223 and 55224.

And the mutation was watched to fail. A build whose application never publishes the endpoint, cold-booted the same way, opened ONE connection instead of two; the identical 110-second wedge replaced its session (5600556063). So the survival above is the resident doing it, and not the host having quietly stopped timing anybody out. This also fixes the honest limit of the guest-side fix below: it stops the guest tearing its own link down, and it does not keep a session that nothing is answering for.

How it is built, and the one decision worth arguing with. MacTCP's .IPP driver through the Device Manager — PBOpen and PBControl are traps, so a flat 68K code resource needs no library, which is what killed Open Transport at the linker. No completion routines. The plan expected register-based callbacks each needing now_liveness_tm.S's shim treatment; they buy nothing here. This component already has a periodic interrupt-time context, the channel's entire job is one frame every thirty seconds, ioResult is the same fact the callback would carry read from memory instead — and the ABI is genuinely ambiguous for these callbacks in a way Timer.h's was not (MacTCP.h declares a STACK-based completion; the Device Manager documents A0/D0). Guessing wrong there costs a five-second corruption of somebody else's memory, which is exactly what disarmed the vehicle for a day. The full argument is in ext/src/now_liveness_net.c's header.

What this does NOT close. The channel has been watched on an emulated G4 only. MacTCP on a real PowerBook under OS 9 is not OT's MacTCP compatibility on an emulator, and nothing has run on 68K/System 7.1 at all, where the extension is the same INIT but the stack is real MacTCP. Metal is attended and Michelle's call.

FIXED, and found only by driving the real host: the GUEST killed the very session the resident was holding open (2026-08-06)

The entry above nearly read the other way. The first end-to-end run against the real host — resident channel up, host correctly holding the session — still lost the session, and both connections were replaced within 150 s of the wedge.

now-guest-ppc/src/core/wire.c :: service_heartbeat compared last_rx_tick against a 65-second dead-link window. When the event loop next ran after the 110-second starvation it saw a 110-second gap and declared the link dead: "Reconnecting (no answer)". Nothing had been silent. The application was not running to listen, and then blamed the far side for it.

It is the same defect as the host's, from the other end, and the guest's version is the plainer one: the host at least observed real silence and had to be told a machine might be alive behind it.

The cure is the same shape. A gap between two consecutive passes of our own event loop longer than ten seconds is PROOF of starvation rather than evidence of it, so the interval is forgiven — the dead-link clock advances past it instead of counting it, which keeps a genuinely dead link noticed one window later rather than never. A healthy pass is milliseconds, and the pump runs from every nested Toolbox loop (pump.h), so nothing legitimate lands between one second and ten.

It deliberately needs no extension. liveness_ticks says the same thing more precisely and an application that has one could read it, but "keeps its session through a modal" must not be a thing only some machines do — the product degrades honestly without a resident component (docs/resident-components.md).

The instrument was feeding the clock it was measuring

Worth more than the fix. The same 108-second starvation measured through tools/liveness-channel.py did not drop the link, twice, and that is why this survived a whole afternoon of green runs.

That instrument polls the session with mirror every five seconds. The requests piled up in the socket while the guest was starved; when it came back it read all twenty-two of them before service_heartbeat ran, and last_rx_tick was refreshed on the way past. The probe supplied the traffic whose absence was the thing under test.

Same class as probe-oracles-were-blind and the hello-probe trap in tools/wedge-experiment.py, and the general form is worth stating: an instrument that talks to the subject on the channel it is measuring is not a passive observer of that channel. The real host, which pings nothing by contract, was the only observer quiet enough to see it.

A refusal that outlives the thing it was about (2026-08-07)

The lane was pointed at the guest's handle verb, on a report that an ok reply carried a leftover reason. handle did not have that defect — obsresolve.c's resolver sets the verdict and the reason as a pair and a native test already pinned Ok ⇒ no reason. Saying so is the useful half: the report was about a real class and named the wrong instance of it, and the reason it was believable is that handle is the only reply in either guest that states a reason inside an ok:true frame. That is by design (the refusal is the product) and it makes the one place where the two could ever disagree.

Fixed

Derived rather than remembered — the sweep is one command and it belongs beside the claim:

# every reply builder carrying BOTH ok:true and a reason-shaped key
python3 - <<'EOF'
import re, pathlib
for sub in ("now-guest-ppc/src", "now-guest-68k/src", "ext"):
    for p in sorted(pathlib.Path(sub).rglob("*.c")):
        t = p.read_text(errors="replace")
        for m in re.finditer(r'(snprintf|append|fmt_append_str)\s*\(.*?\);', t, re.S):
            s = m.group(0)
            if r'ok\":true' in s and re.search(r'\\"(reason|note|detail|message|why|error)\\"', s):
                print(f"{p}:{t[:m.start()].count(chr(10))+1}")
EOF

One hit: observe.c's handle emitter. It read verdict, reason and resolved off three independent expressions; all three now derive from handle.why through the mapping obsresolve.c already owned (now_obs_verdict_for_why). Not a bug fixed — a bug made unrepresentable, in a file that has no host test at all, which is why it also got handle_reason_source_test.py beside one_minter_source_test.py.

Two reachable instances of the actual class, found by asking the general question:

  • now-guest-ppc/src/files/files_browser_view.cg_note, the status placard, was written only on failure paths and never reset by a successful listing. Step into a refused folder, then into one that lists: twelve items under "that path leaves the shared folder". Now cleared where the question is ASKED, so it also covers an answer that never arrives.
  • now-host/Sources/Host/MirrorModuleModel.swiftcontentNote survived show(document:), which is the door the live watch loop comes through. One refused content join drew its sentence under every healthy window that arrived afterwards. clearScene() and guestLeft(_:) both already stated the rule this door broke.

Unverified, and the sharper of the two

  • The guest half is Builds, not Tested. files_browser_view.c is Carbon and has no host test; nobody has watched the placard clear on a machine. The reasoning is from the source and from two sibling browsers that get it right, not from a screen.
  • handle was never driven end to end this lane. A VM was staged (extension + app, cold-booted, both verified present after the reboot) and the guest did not dial the listener within 180 s. So the reply's shape is argued from the emitter and pinned by two gates; no one has read a real handle frame off a wire here. The host change is Tested — two XCTests watched failing first, each naming the leftover sentence.
  • The brief's rig names do not exist in this tree. tools/lane-ports and tools/shutdown-guest.py are not here; the parent checkout has tools/shutdown-guest, and scripts/spin-up-ppc cold-boots by QMP quit on purpose (an INIT loads at boot only). Worth knowing before the next lane spends time looking for them.

Not fixed, and deliberately

now-guest-ppc/src/screenshots/screenshots_module.c has the same shape in a weaker form: set_status(err) on failure, nothing on success, so a prior "Failed: …" can sit over a successful send. It is overwritten in practice by conn_set_shot_note arriving shortly after — which makes it correct by timing rather than by structure. Left alone rather than folded into an unrelated fix, and named here so it is a decision instead of an oversight.

Host Files export and mutation repair (2026-08-08)

Tested, not metal-verified. Four host-side failures are repaired:

  • an overwrite that the guest cannot finalize after staging now reports the refusal once; an already-authorized overwrite never opens a second prompt;
  • overlapping refreshes have one generation owner, so creating a folder no longer appends the same guest listing twice;
  • the new-folder sheet owns its text draft and dismisses before the request, so a late TextField write cannot reopen it after success; and
  • every dragged row is a real AppKit file promise. Folder promises recursively follow paged guest listings, and multiple promises wait in one bounded queue for the guest's single transfer lane.

The new native tests pin each repair, empty directories, partial-tree cleanup, refusal without deleting an existing host folder, the recursive item bound, and multi-promise ordering. The overwrite, overlapping-refresh, folder-bound, cleanup and sheet-lifetime guards were each watched fail before their fixes. No wire message or guest behavior changed.

Still unverified: Finder has not redeemed a multi-file or directory drag from a live guest in this pass, and a running application on a classic Mac has not exercised the finalization-refusal path. Those are the two useful metal checks; the host suites prove the state machines and app build, not those interactions.

2026-08-09 follow-up -- tested, not live-UI or metal-verified. The Files page now refreshes both its open listing and Places when a connection becomes usable, and the toolbar Refresh action refreshes both projections as one operation. A disconnected or superseded Places sweep cannot publish late. Command-Delete and the forward-delete key use the existing Move to Trash confirmation; Shift-Command-N opens the existing new-folder sheet. All three entry points refuse to start while another guest file change is active.

The delete regression now uses the listener's real connection identity and requires both refreshes that production emits -- the guest tree-change event and the successful mutation completion. It replies newest-first and proves the older reply cannot append a duplicate row. This settles the reported duplicate render at the state-machine level. The remaining useful check is a live host UI pass for the AppKit key events and sidebar redraw; no guest behavior or wire message changed.

Multi-guest host controls: tested, not visually or metal-verified (2026-08-08)

The host now keeps its guest choice at the top of the sidebar as a live-only menu. That selection is the listener's active session, so every module moves with it; remembered machines cannot appear attached merely because their registry record survived. The menu's Add Guest action opens Connections.

Connections is now a split view: active sessions above remembered machines on the left, the selected machine's link status, identity and session log on the right. Remembered rows are visibly secondary. The list's plus action opens the listening setup, while minus closes and forgets exactly a selected live session or forgets exactly a selected remembered record. Pressing Return in the port field follows the same validated Start Listening action as the button. Each row is headed by a host-owned display name, defaulted from the machine's reported name; collisions receive -2, -3, and so on. The guest IP and the host port that particular machine used are the subrow; remembered machines retain their own last-used ports when the current listener changes, and the guest's transient outbound source port is deliberately not shown. A pencil beside the detail title and the row/detail context menus rename that display name without changing the stable machine id; the same menus expose Delete and the listener's Start/Stop action.

Tested here: scripts/test-all passes: 82 native tests, the complete host suite, and the Debug and Release app builds. Focused host tests cover active-versus-remembered removal, exact remembered-record removal, session-scoped health and logs, the plural module name, and Return starting the listener. Display-name tests cover default and duplicate allocation, persistence across reconnect, editing live and remembered rows, and preserving the session and stable machine identities. Two-guest socket tests remove either session and prove the remaining guest is still promoted or still answers. Guest cross-builds were skipped because this worktree has no Retro68 toolchain.

Still unverified: nobody has looked at this layout in the running host app yet, and none of these controls has been exercised against two physical Macs. The existing two-guests-on-one-port limitation therefore remains: tested is not metal-verified.

The folding sidebar, both halves (2026-08-05)

The rail folds to icons on the guest and the sidebar does the same on the host, both with tooltips on the folded icons, and the host gained the guest's other two choices (density, rearrange).

Guest: emulator-verified. Folds and unfolds, the icons and the Connection lamp draw, the hand-drawn help tag appears over a Data Browser and the Browser repaints clean when it goes. A true fresh install - prefs file deleted - starts expanded with rich rows, which is the default that matters.

One note on reading a shared VM, because it cost a paragraph of wrong suspicion: a first launch appeared to open COLLAPSED, and the cause was the human clicking the collapse button in the seconds between the launch and the screendump. The evidence was all there and pointed the right way once anyone thought to ask - nothing writes that prefs file at launch, only a quit or a toggle does, and its created and modified stamps were seconds apart and dated that day rather than inherited from the base image. When a VM has a display, a person may be using it, and a screendump is a sample of a machine somebody else is also touching, not a readout of what the code did.

Host: confirmed working by the human at the desk, not by me: screen access was declined, so I have never seen it. swift test and the Xcode target in both configurations are all I can speak to directly.

Open, and worth knowing:

  • Carbon help tags do not display under Mac OS 9. This toolchain's MacHelp.h carries only the help-tag API (HMSetControlHelpContent, HMDisplayTag) - classic HMShowBalloon is not in these headers at all - and help tags are a Mac OS X facility. The guest's tooltip is therefore drawn by hand. Anyone reaching for the Help Manager on this target should expect a feature that compiles, links, and shows nothing.
  • The guest's tag is armed from idle, which runs unslept during a transfer. It is one GetMouse and two comparisons per pass and draws only on a change, but it has not been watched during a large transfer, which is the condition that has caught every other idle-path mistake in this window.

The rearrangeable sidebar: emulator-verified, never on metal (2026-08-04)

The rail gained four things in one arc — a scroll bar that appears only on overflow, Option-drag reordering, a compact density, and a Preferences page. Emulator-verified the same day (OS 9.1 under QEMU, a private --instance 7 clone off the runner-ready snapshot, the build pushed over the harness's own put channel):

  • the page renders and the pinned group draws with its new sliders icon;
  • compact fits all eleven nav rows plus the pinned three at the standard window size, keeps the icons, and correctly gives the Connection row its STATE rather than its title as its single line;
  • Option-drag arms, draws its XOR insertion line, tracks the pointer, erases the line cleanly on drop, and moved Chat from the foot of the list to just below MCP;
  • the scroll bar appears when the window is dragged down toward the minimum in rich density, spans exactly the nav rows, stops above the divider, narrows the rows to make room, and scrolls by one on its arrow;
  • all of it survives a quit and relaunch — saved order, density and window rectangle — and Reset Order puts the enum order back.

That is the level below metal and it settles the drawing. What it does NOT settle, and what to watch for on the PowerBook:

  • The drag against a real hand. The emulator was driven with synthesised relative motion in 3-pixel steps. A fast human drag is the case where XOR feedback smears, and no synthetic mouse will show that.
  • Compact baselines under a different system font. 18-pixel rows with a hard-coded baseline of 13 — arithmetic, not a metrics query. The emulator runs the same fonts the PowerBook does, so this is likely fine and is listed because "likely" is not "watched".
  • The prefs v19 migration from a file written by an OLDER build. What was exercised was v19 writing and reading its own record. The remap that matters (Connection 13 -> 14, Logs 12 -> 13) needs a prefs file from before this arc, which the throwaway VM did not have. This is the fifth time that remap has been written and the ordering trap — remap Connection first — has been the bug at least once.

A control created at an empty rectangle may never render (2026-08-05). Worth knowing beyond this file. The rail's bar is created once and Show/Hidden, and the layout reports an EMPTY nav_scroll while everything fits — an honest description of "no bar", and a fine creation rect for scrollBarProc, which drew correctly after a later SizeControl. kControlScrollBarLiveProc does not: born 0x0, it stayed invisible no matter how it was resized. The tell was that the bar appeared when the app opened INTO a small window and not when the same window was dragged down to the same size — two paths to identical geometry, one working. Create controls at a real rectangle and let Show/Hide carry visibility.

Two things are known limitations rather than suspicions:

  • The drag does not scroll the list under itself. Rearranging a row past the visible edge of a scrolled rail means dropping, scrolling, and dragging again. It only bites at rich density in a minimum-size window, which is also the only configuration that scrolls at all.
  • The Edit menu holds Preferences and nothing else. Cut/Copy/Paste are absent rather than greyed, because this window has no keyboard focus machinery and the TextEdit fields on Chat and Console each own their own TEHandle — there is no "the focused field" for an Edit command to act on. Wiring them needs a new module op handing the Workshop the page's current TEHandle. That is the same missing seam as the dead GetKeyboardFocus gates noted below.

The chat.* family: metal-proven end to end; edges still open (2026-08-02)

Host credential boundary updated and host-tested on 2026-08-09; the corrected legacy migration remains unverified. The first signed desk run disproved the original passive-read claim: opening Chat produced two legacy ACL dialogs because the fallback query still requested each secret's data. Passive provider discovery now probes only legacy item attributes; one explicit Chat-page action is the only path allowed to request an old login-keychain value, and successful migration moves it to the app's Data Protection Keychain access group. The default host build is Apple Development signed by team B93A9CG7F9; the host gate now rejects a validly sealed bundle that lacks that team identity, application identifier, or Keychain access group. --adhoc remains available for scratch work and is deliberately unable to use Chat credentials.

Tested here: the credential-store and Anthropic provider/OAuth suites, including noninteractive reads, lazy one-read-per-used-key caching, working API key fallback when an OAuth item needs authorization, single-flight refresh, and post-completion stale-token replay. A fresh default script build passed the signature/entitlement assertion; a known ad-hoc bundle failed it as intended. scripts/test-all passed; its native tests and host gate ran, while all three guest cross-builds reported their documented toolchain-unavailable skip.

Still unverified: no one has launched the corrected build against a protected legacy dev.newoldworld.now.chat item and watched the full sequence. The remaining proof is that passive startup shows no system dialog, one explicit Authorize action produces only the prompts required by the credentials that actually need access, subsequent signed launches are prompt-free, and migration or sign-out reports a login-keychain cleanup refusal instead of silently leaving or resurrecting the old item. The ad-hoc runtime refusal is also a code and signature claim, not something watched in the app.

Four rounds on the real desk moved most of the 2026-08-02 unknowns to answered. Metal-verified: the Anthropic subscription sign-in (paste-back OAuth, the Claude-Code-shaped request gate), oMLX serving (after the no-blank-line SSE finding), the guest Chat page drawing, streaming and tool use against a real guest — a model paged this PowerBook's software and tried to launch a game over the wire. Emulator-verified (2026-08-02, restricted VM): the TextEdit prompt types, edits and refuses correctly; the page renders with both popups. Still open:

  • The redraw pass is emulator-still-frame verified only. Per-row transcript damage and TE's incremental drawing are built exactly to kill the flicker metal showed, but flicker is a motion claim: a human watching a streamed answer on the real screen is the proof.
  • Ollama and LM Studio probes have still never seen a live runtime (oMLX has; the other two remain scripted-transport claims).
  • OpenAI (the hosted service) has never been reached — key entry, models list and a streamed turn are all unproven.
  • The turn ceiling raise (12 -> 40) and the launch guidance are untested against the errand that hit them — re-run the find-and-launch-a-game prompt.
  • chat hi from the host console (exec plane -> guest chat verb) still has not run against a booted guest.
  • Keyboard focus is dead application-wide (no root control, on purpose): the three Data Browser pages gate arrows on GetKeyboardFocus and so have never taken a key. Spun off as its own task; docs/guest-ui-start-here.md carries the rule.

CLOSED: the liveness vehicle runs, and it runs while applications do not (2026-08-05, later)

The entry below is fixed, and it is left standing because the wrong diagnosis in it is the more useful half.

The cause was the callback ABI, not A5. TimerProcPtr is declared CALLBACK_API_REGISTER68K, which on classic 68K collapses to a function of no arguments — the header says so in its own words — and Timer.h's procinfo (0x0000B802) spells out a register-based call taking four bytes in A1. now_liveness.c wrote the tick as pascal void tick(TMTaskPtr), so the compiled function read a stack argument belonging to whatever it had interrupted. Nothing rescued it in between: on classic 68K NewRoutineDescriptor is a no-op macro returning its ProcPtr, so NewTimerProc handed the Time Manager that C entry point bare. The fix is ext/src/now_liveness_tm.S — the shim now_ext_gne.S already was, for the same reason and in the same shape.

One defect explained both symptoms, which is why it is the answer. A task whose record pointer is garbage re-primes a garbage queue element, so it fires once and stops — the first build exactly. Move the counter into the record and that same garbage pointer becomes a five-second write into somebody else's memory during boot — the second build exactly. A two-defect story was not needed.

The A5 lesson recorded below is WRONG for this component, and that is worth more than the fix. Retro68 does not address this extension's globals through A5: _start calls RETRO68_RELOCATE and never frees them, so the flat blob's statics sit at fixed system-heap addresses. The tree already proved it before anyone theorised otherwise — now_ext_gne.S reaches gNowExtOldGNEFilter with an absolute load from inside every application's context, and the anchor plane's gLastA5 fast path works across processes. Both would be luck if the A5 story held. The wrong answer is left in the source comment because it was a plausible one and the next person will reach for it too.

Measured on a fresh cold-booted clone, three results:

  • The guest boots normally. capabilities: 63 — bit 5 (32) is the vehicle — with Finder, anchor worker and NOW all up.
  • The counter keeps the stated cadence. Two reads 232.5 s apart by the guest's own tick count: livenessTicks +46, and 46 × 5 s is 230 s.
  • It keeps climbing through a starvation that stops applications. Under tools/guest-wedge spin 25, tbt-worker — a background-only application on its own TCP port with no code in common with NOW — could not be reached for 25 consecutive seconds, and livenessTicks gained 5 over 26 s, the undisturbed rate.

So 012 § 3's deliverable is met: the premise the whole liveness plane rests on is no longer an argument from the cooperative scheduling model. The instrument is tools/liveness-experiment.py, and its INCONCLUSIVE verdict — the failure that most looks like success — was watched to fire by forcing the probe to report the worker always alive.

Two of § 5's recorded measurements did NOT reproduce here, and are flagged rather than resolved. § 5 records that a spin wedge left hello answering while stat died, and that modal starved nothing over 71 s. On this clone, under a spin wedge hello went silent too, and modal starved the full 25 s. The second has a reading available from the source — ModalUntil pumps with GetNextEvent, not WaitNextEvent, and background processes get time from the latter — but that is a hypothesis, not a finding, and the two runs differ in more than one way. Nothing here depends on either: this result is stronger if the worker was silent at every level, because the resident ticked through it anyway. Whoever needs § 5's table should re-measure it before quoting it.

2026-08-05, independently reproduced. A second session reached the same ABI diagnosis from the procinfo word alone and measured the same result on its own clone: 17 ticks over an 83 s window containing a 25 s spin starvation — the undisturbed rate; a vehicle frozen through the wedge would have shown ~12. Two rig lessons from that run, neither recorded elsewhere: launch of a Desktop Folder path resets the worker's connection on this rig and the app never starts (stage into the now-dev folder instead); and a starved guest shows as a long gap between successful probe replies, not as refusals — QEMU's user-net proxy accepts and buffers, so an untimestamped probe timeline reads "alive" straight through a starvation.

BROKEN and DISARMED: the liveness vehicle hangs the guest at boot (2026-08-05) — FIXED, see above

012 § 3's Time Manager task — the extension's first interrupt-time context — is in the tree and does not install. The return that disarms it is deliberate and is explained where it sits.

Two defects, one behind the other, and only running it found either.

  1. The first build's tick fired once and stopped. livenessTicks read 1 on a guest that had been up for minutes. A Time Manager task is entered with an ARBITRARY A5 and Retro68 addresses globals through A5, so from interrupt time this component's statics are somebody else's memory — reading them is luck, writing them is corruption. Fixed by carrying everything the task needs inside the task record, which the Time Manager hands back, so no global is touched at all.
  2. With that fixed, the guest never finished booting. The Finder drew an empty menu bar and stopped; the anchor worker never came up; the machine was still unreachable five minutes later, and a screendump shows a bare desktop. So the first defect had been masking the second: a task that fires once does not hang a machine, and one that fires every five seconds does.

The cause is not established, and this entry should not be read as though it were. Two candidates, both testable and neither tested: the task may need the same globals-world shim the jGNE filter has in assembly (now_ext_gne.S), or re-arming with PrimeTime from inside a standard Time Manager completion may be wrong here.

Why it ships disarmed rather than reverted. An extension that hangs a Macintosh at startup is the worst thing this component can be — it is recoverable only by pulling the file from a machine that will not boot. But the A5 lesson is worth more than the file, and whoever picks this up should start from a diagnosis rather than a blank page.

What this costs 012: § 3's deliverable was prove the vehicle runs while applications are starved, and that is now unproven in the strongest sense — the vehicle cannot be left running at all. The liveness_ticks counter and its livenessTicks report are built and correct; nothing has yet made them climb.

UNBLOCKED at the first gate: MacTCP's .ipp driver DOES open from the extension (2026-08-05, later)

The entry below stands — Open Transport is still ruled out, and for the structural reason it gives. What has changed is that its named successor has been asked, in the same cheap way, and answered yes.

On a fresh cold boot the resident opened MacTCP's .IPP driver with a plain PBOpenSync and reported transportProbe: 1 (open), transportResult: 0 (noErr). The guest booted normally alongside it — capabilities: 63, lifecycle active. So the Device-Manager route is real from a flat 68K code resource on OS 9: no library, no CFM, nothing for a linker to refuse.

What this proves is exactly one thing, and the temptation is to read it as more. A driver opened. Nothing has been created, dialled, sent or received. TCPCreate and TCPActiveOpen are PBControl calls this has not made, their completion routines are register-based callbacks that will need their own shim — the lesson the vehicle just charged for — and whether OT's MacTCP compatibility actually serves a connection to a resident client is untested. The first gate is clear; that is all.

The probe was watched to report the other answer. A build asking for a driver that does not exist (.NOP), cold-booted the same way, reported transportProbe: 2 (refused) and transportResult: -43 — the Device Manager's own fnfErr. That mutation is not ceremony: a field reading "open" unconditionally would be indistinguishable from the result above, which is the class probe-oracles-were-blind names.

BLOCKED, and the answer came from the LINKER: the extension cannot be an Open Transport client (2026-08-05)

Plan 012 § 4 has the resident dial the host itself, so a starved machine answers for itself. It cannot, from this component, and the reason is structural rather than a matter of effort.

Linking OT's client glue into ext/ fails at link time on __SLM11FuncDispatch, __SLM11VTableDispatch, __SLM11ConstructorDispatch, __SLM11ExtblDispatch and __gOTClientRecord. Those are Shared Library Manager dispatch stubs: OT's 68K libraries are CFM/SLM fragments, and the NOW Extension is a flat 68K code resource (-Wl,--mac-flat). The two linkage models do not meet. Four library combinations were tried, including the application flavour (OpenTransportApp) and the compatibility library; the best was fifteen unresolved symbols.

This is 012 § C's metal question answered at link time — the risk was named as "OT 1.x on a PowerBook may not behave as OT 2.x on an emulated G4", and it turns out not to reach a machine at all. Much the cheapest place to find it.

The routes left, none of them a small edit:

  • Reach TCP through the Device Manager instead. MacTCP's .ipp driver is driven with PBControl and completion routines, which a flat 68K INIT can do and which is how resident code did networking before OT. OS 9's OT still provides it for exactly these callers. This keeps the extension an INIT and is the smallest of the three.
  • Ship the resident as a CFM fragment, which can link OT properly but is a different kind of component with a different install story.
  • Ship it as an OT module, which is what OT's own layering intends for something living below applications, and is the largest change.

What landed anyway, and why it still matters: the VEHICLE. The extension now has its first interrupt-time context (a Time Manager task, ext/src/now_liveness.c) which runs whether or not any application is being scheduled — the thing every existing context in this component is not. It bumps liveness_ticks in the shared table and nothing else does, so the premise the whole plane rests on is now checkable rather than argued: under tools/guest-wedge spin an application-level probe stops answering, and this counter must keep climbing.

BROKEN: the guest's deafness OUTLIVES the host's liveness window, so an ordinary modal kills the session (2026-08-05, second drive)

This is the one that makes the other deafness entries conditional, and it was found by deliberately reproducing the wedge rather than waiting to be surprised by it again.

A document with a creator no application owns (Zz!9) was pushed to the guest desktop and opened from the Mirror. The Finder raised its "Could not find the application program that created the document named 'wedge doc'" alert — reproducible on demand, which is what makes this measurable at all. Then:

  • The alert deafened EVERY process on the guest, not merely the Finder. tbt-worker — a background-only application on its own TCP port, sharing nothing with NOW but the machine — stopped answering hello for more than 90 seconds, and refused four posted-click attempts across the six minutes after that. NOW went silent for the same span.
  • The host's idle timeout is 75 s (GuestListener.Timing.idleTimeout, and the host never pings by contract). So the wire died of "Connection lost (no traffic)" while the Macintosh was perfectly healthy and its socket perfectly open.

One ordinary Macintosh event therefore ends the wire session, and no amount of host-side patience is the answer: a modal waits for a person, so the deafness is unbounded. A larger idleTimeout moves the cliff and makes real death slower to notice; it does not remove the cliff.

What this falsifies. The entry below ("a blocked callee deafens NOW for 15 s per script") proposes short poll deadlines as "the cheapest real fix". That fix is still right and is still worth doing — it stops NOW spending fifteen seconds per poll on a blocked callee — but it cannot reach this: our own deadline governs how long WE wait, and this is the guest being unable to answer anybody at all. 011 § C's done-when ("a blocked Finder costs the guest its Finder reads and nothing else — the wire stays live") is unreachable by that mechanism alone, because the blocked Finder costs a separate background application its wire too.

Where the answer has to live. Liveness is being answered by the application, and a modal is precisely what takes the application away. So the signal must come from below it. Two layers can speak while every application is starved, and both are unproven here:

  • The TCP stack. A wedged guest's kernel still holds the connection; a dead machine's does not. Host-side keepalive would be answered by the guest's OT stack with no application involvement, which is exactly the distinction the idle timer cannot make. Unverified on OS 9's Open Transport, and it must be checked on metal before it is believed — MacTCP on the 68K side is a second question again.
  • A resident component. The NOW Extension runs at interrupt time and is not subject to cooperative starvation. This is what resident-components.md exists for, and it is the larger move.

There is also a third answer that does not try to keep the session alive: make a redial cost a reconnect rather than a session. GuestKey is per dial-in, so the same machine returning gets a new key and the Mirror's pinned engine, journal and act clocks are all invalidated — which is why the four acts in flight came back "the Mirror is pinned to guest-1". Re-pinning across a redial of the same MACHINE would make an unbounded modal survivable without pretending the wire survived it.

Recorded from the drive whose notes are in docs/local/, with QMP screendumps of the alert and of the machine unchanged after the posted click.

2026-08-05, later: this has a slice, and reading the code sharpened it twice. 012, Liveness below the application. Two corrections to the entry above, both from the source rather than from reasoning:

  • The keepalive already exists, and it is guest-driven — the contract has the guest send ping after 30 s of silence. So nothing needs a new liveness message; the starvation stops the application sending the one that is already there.
  • The resident cannot currently help. ext/src/now_ext.c installs its vehicle at LMSetGNEFilter, and the act plane's patches fire on trap calls, so every execution context the extension has is application-driven. During this exact starvation the resident does not run either. Liveness therefore needs the extension's first interrupt-time context, which is what makes 012 a slice rather than a patch.

One thing in this entry is over-stated and is corrected there: the deafness was >90 s with recovery, not indefinite. "A modal waits for a person, so it is unbounded" was an inference, and ModalDialog does call GetNextEvent — so a modal merely sitting may starve nothing. The suspect is the app-enumeration scan behind that particular alert, and it is a suspicion, not a finding.

UNVERIFIED: MCP can see the drawing now, but nobody has read one live (2026-08-05)

now_mirror_snapshot claims to carry the renderer's whole input and three times it has not: window.items for desktop icons, window.items again for Finder rows, and window.display — the per-window QuickDraw ops. The third is now projected, along with kind, ref and text, which turned out to be missing for exactly the same reason. Slice 6's render rule (to render a custom control, replay its ops rather than classify it) is reachable from MCP for the first time.

Status is TESTED, not verified. 1416 host tests pass and six new ones cover it, but no snapshot has been read off a real guest with ops in it — the VM stand-up could not complete a cold boot unattended that day. The shape is proven; the content is not. What would settle it is one headless now_mirror_snapshot against a guest whose front window is being drawn, checking that displayTotal is non-zero and the ops describe what the Mirror window is showing.

Three things worth keeping regardless of how that goes:

  • The omission CLASS has a guard now, not just this instance. A test walks Scene.Window's stored properties with IRSchema.declaredProperties and fails on any the projection has not disposed of — carried (proven against a real projection) or declined with a reason. Adding a field to the IR and forgetting the projection is no longer silent. Watched failing by adding a probe field to Scene.Window: it named probeField.
  • The guard has its own guard. A roster check that asks whether a key exists passes for a field that arrives empty, which is coverage the roster does not have — the same shape as a test that stays green for the wrong reason. So every carried check also runs against a window whose fields are all present and all empty and must fail there. Watched failing by weakening one check to != nil.
  • The 64 KB ceiling has no headroom left, and this is the number to know. The item projection ALONE encoded 54.6 KB of 64 KB in its worst case, measured rather than estimated. Any independently-bounded addition therefore overflowed the message, and overflow is not a truncated reply — it is the writer throwing and the connection closing with no reply, so snapshot stops answering while status still does (the same failure this file already records from adding Finder items). Item and content families now hold separate stated byte shares of one ceiling. Anything further added to this payload must take a share rather than assume room; there is none.

BROKEN: a blocked callee deafens NOW for 15 s per script (2026-08-05)

This is the mechanism behind the dropped connection, and it is not about modals — a modal is only the commonest way to get a blocked callee.

The guest is a serial Carbon application. now_input_run_script calls OSADoScript, and while that call is inside the OSA component NOW pumps nothing: not the wire, not its own event loop. The file says so in its own refusal text — "this guest is serial and could not wait longer".

There is a deadline, and it works: OSASetActiveProc installs a proc the component calls while the script runs, which returns userCanceledErr past the deadline. So a script cannot hang forever. But the deadline is kNowScriptDefaultMs = 15000, and the host never sends timeoutMs at all — zero occurrences in NOWMirrorSource. So every script to a blocked application makes the guest deaf for a full fifteen seconds.

Stack a few and the arithmetic is the drive that died: acts queued 48–87 seconds, the host's heartbeat unanswered through all of it, six operations settling sessionChanged, and the guest's socket left CLOSED while the host still believed it had a guest.

The Finder is the worst possible application for this, because NOW reaches it for the icon roster, finderOpen, finderSelect and the visibility census — routine scene maintenance, not just user acts. So a blocked Finder does not degrade one feature; it deafens the whole guest on a timer.

The cheapest real fix is not a new mechanism. Routine scene- maintenance reads (icons, census) should carry a SHORT timeoutMs — they are polls, and a poll that cannot answer in a second should give the wire back rather than hold it for fifteen. A user-initiated act can keep the long deadline, because a person is waiting for that one on purpose. The argument already exists on the verb; nobody passes it.

2026-08-05, later: "the cheapest real fix" is still worth doing and is no longer sufficient. Measured by reproducing the wedge: the deafness is not scoped to NOW's own script call. A Finder modal starved a separate background application (tbt-worker, its own port, no code in common) for over 90 seconds. A short timeoutMs bounds how long WE wait for an answer; it cannot make a starved machine able to answer. See the new entry at the top of this page for where the liveness signal has to come from instead.

The host's dispatch path has no test seam, and it cost three times today (2026-08-05)

Not a defect in the product; a gap in how the product can be checked, and it has now been paid for often enough to name.

NOWMirrorSource.serve turns a plan into a guest command and sends it through the listener. Nothing can intercept that command in a test, so every fix to it is verified one level BELOW where it lives — the helper that builds a string is tested, and the line that decides which string to build is not. Three fixes on 2026-08-05 landed with exactly that shape:

  • the Hide host switch: hideDispatchOutcome is tested; that .hide calls the guest verb at all is not;
  • the transitions arg key: the parse is tested against a real envelope, and the agent had to DRIVE A MACHINE to prove the wiring;
  • the Finder activate rule: finderScript is tested in both shapes, and mutating the call site to pass activate: true unconditionally left nine tests green.

That last one is the clean demonstration, because the mutation reinstates precisely the defect the change exists to fix and nothing notices.

It also explains a pattern in today's live drives: the defects the machine found — a refusal that was unparseable, a reply that lost its own records, a verb that could not arm — were all in this same band, between a tested helper and a tested projection.

What would close it is a seam of the kind NOWMirrorCycleIO already gives the scene cycle: one injectable "send this command" function, so a test can assert WHICH command a plan produces. That is a small refactor with a large blast radius on confidence, and it is the highest-value testing work in this arc.

2026-08-05, later: CLOSED, and by exactly that. GuestCommandSend is the injectable door, defaulting to GuestListener.runCommand and held by both NOWMirrorSource's run/readingOutput and AgentIntegrationActControl's single wire call. MirrorServeSeamTests drives perform through it and reads the composed command. The mutation named above is the proof it works: activate: !ownAppactivate: true now fails exactly the control-panel test and nothing else, where it used to leave nine tests green.

BROKEN: one modal wedges the whole Mirror, and the lane turns it into 90 seconds (2026-08-05, from Michelle's drive)

The most complete failure this arc has recorded, and every link is evidenced. It is not one defect; it is three, and the one that matters most is not the one that started it.

The chain. Opening harness.log — a document whose creator application is not on the machine — makes the FINDER raise its "could not find the application program that created the document" modal. A Finder inside ModalDialog does not service Apple Events, and NOW reaches the Finder for everything: the icon roster, finderOpen, finderSelect, the visibility census. So one ordinary Macintosh event stops the Mirror's entire Finder surface, and nothing on either face can dismiss it.

The amplifier, measured. From acts.log, queue waits behind the stuck lane: 48 737, 53 146, 53 821, 63 246, 65 653, 70 923 and 87 508 ms. One act waited 87.5 seconds. That is WORSE than the 51.8 s that made the lane a plan item in the first place, and it happened after the fix that was supposed to remove most of its fuel.

Then the session died. Six operations settled sessionChanged, four more logged disconnected: with 48–66 s waits already on them, and the guest's socket was left CLOSED on the host's listening port while the host still believed it had a guest. The Mirror had to be relaunched, then believed itself connected and dispatched nothing.

What this falsifies. The plan deprioritised the lane amplifier (5c item 2) on the reasoning that fixing the owner prediction "removes most of its fuel". That reasoning is wrong, and this is the counter-example: ANY Finder-blocking event refills the lane instantly, and a modal is a normal thing for a Macintosh to do. The amplifier is not downstream of item 1 — it is the thing that converts any single stuck act into a dead session.

Three defects, in the order they should be fixed.

  1. The lane must be bounded and cancellable. A single FIFO whose only escape is a 15 s timeout cannot survive one act that blocks. Nothing can be cancelled: a person watching a 70 s wait has no way to abandon it. This is now the highest-value item in the arc.
  2. A modal must reach the operator. Rung 4, unchanged since it was written, and now with a second door into it. Worth noting from the same session: a Dialog Manager button IS clickable through the posted click the anchor worker sends — the modal was dismissed that way — even though a MENU is not. So the act is reachable; the plumbing is not.
  3. A document open needs an honest postcondition. harness.log fell into the documented gap: unknown kind, so the prediction falls back to a Finder window, which never appears. That did not cause the modal, but it is why the act sat in the lane for 70 s instead of failing fast.

And one thing that worked. The Finder roster read correctly throughout this session — desktop icons rendered fine, per the operator. So the intermittency recorded above is real and this session is a case where it worked.

2026-08-05, later: item 1 is done, and the re-measurement is much better — but it did not test the part that matters most. The lane is now bounded and cancellable: an act can be cancelled from the Mirror's status line and from now_mirror_drive --gesture cancel; a timed-out act sheds everything queued behind it with an attributable refusal; and a failed cycle that finds the pinned session disconnected ends the lane at once rather than letting each act spend its own deadline. Driving the same wedge deliberately: five acts issued across it answered in 2.1 s total, no caller blocked, and the wedged act's own ceiling was 30.3 s (the guest's 15.3 s script deadline, then the broker's 15 s). No 87.5 s wait recurred.

That is consistent with the fixes and is not proof of them. Nothing reached the lane behind the wedge in that run — the four acts issued after it were held at the observation door, upstream of the broker — so the shed had nothing to shed. The shed, a cancel of a genuinely in-flight act, and the dead-guest notice ending a non-empty lane are verified by unit mutation only.

Item 2 is unchanged and got harder: the posted click that dismissed this modal on 2026-08-05 was refused four times on the second drive, because the anchor worker is itself starved by the alert (top entry). Item 3 is untouched.

2026-08-06: the same modal, reproduced deliberately, and it is a STRICTLY worse case than the one plan 014 repairs. Raising this exact alert on a private clone gave outcome=starved, request_ms 20,008–21,318 and decode_ms=0 for as long as it was up — and the anchor worker stopped answering even hello, so the machine could not be driven at all and the clone was discarded. Compare Michelle's own 2026-08-06 log: request_ms=82 with decode_ms=12457, i.e. NOW answering promptly while the Finder was merely busy. Her modal belonged to a foreign application; this one belongs to the Finder.

The distinction is load-bearing for anyone reproducing either. A foreign-owned modal starves the Finder, which is what 014 took out of the cycle — measured 353 ms → 16 ms. A Finder-owned modal starves the whole cooperative machine including NOW, so no complement runs, no scene is answered, and no host-side repair helps — item 2 above is still the only door, and it is still shut. Someone measuring 014 against a Finder-owned modal will correctly observe that it changed nothing, and draw the wrong conclusion from it.

BROKEN: the host face can HIDE and cannot SHOW (2026-08-05)

Found on the breadth-first drive, minutes after Hide started working from that face. Driving hide with no target hides the FRONT application — and when the front application is NOW itself, its window vanishes and nothing on the host face can bring it back. showAll has no route (the Finder refuses set visible), and the .hide plan only ever hides.

Recovering it needed the guest's own verb over the wire (hide --show "New Old World"), which means stopping the host to free the port — a person driving the Mirror cannot do that from the Mirror. Relaunching the application does NOT unhide it; the launch succeeds and visible stays false.

Measured on the machine: New Old World front=true visible=false, with the desktop icons still on screen because the Finder was untouched.

The asymmetry is the defect, not the hiding. What closes it is a show direction on the same verb the host already calls — hide --show NAME exists and works on the guest today — rather than anything new.

The Finder roster is INTERMITTENT, not absent (2026-08-05)

Sharpening the desktopItems entry above, which reads as though the Finder-item read never works. On one drive it clearly did: the absent-item refusal fired — "the Finder shows no item named No Such Item At All on the desktop" — and that guard only speaks when the roster is PUBLISHED, since an unread container claims nothing. Reads minutes later on the same guest returned nil again.

So it is not a read that never happens; it is a read that sometimes happens. That is a different investigation from the one the earlier entry implies, and a more tractable one: something is racing or expiring rather than missing. The Finder's own Desktop window publishing itemTotal: 0 while icons are on screen is the same symptom from the other side.

FIRST LIVE ANSWERS from the 2026-08-05 drive: display carries, definition says system, desktopItems does not read

Driven against a freshly cold-booted emulated Power Mac G4 with the resident active (fc6e0946bde92715, cap 31), with a real control panel opened so a FOREIGN window with undetermined controls was in view. Three queued questions, three answers.

window.display carries content — WATCHED. Date & Time's window reports displayTotal: 173, NOW's own window 205. The projection that landed the same day as TESTED is now watched answering off a live guest.

definition answers, and for this panel the answer is unanimous: all 29 undetermined controls report system. Every one is a standard Toolbox CDEF out of the System file rather than app-owned drawing. That is the first live data on slice 6's real question, and it points the same way the transport finding did: the information was recoverable all along. For this panel the genuinely-custom population is zero.

Read it with three limits. It is ONE panel at ONE moment; the extension on that image PREDATES the batched-classification fix, so this is the one-cell transport's behaviour and not the fixed one; and 12 of the 41 items were already known from the DITL type byte and never needed a CDEF at all. The corpus recapture is what turns this into a histogram.

desktopItems still does not read, and it is not the stale image. It was nil for the whole 2026-08-05 morning drive, and it is nil here on a clean cold boot with the resident live and a foreign app frontmost.

STILL nil on 2026-08-06, reconfirmed incidentally rather than hunted: two live scene captures taken for the content-plane work (guest build 1bff0bd2ca39, with Date & Time and then Sherlock 2 frontmost) carry no desktop key at all, while the guest's own screendump from the same run shows desktop icons and a full Control Strip. So this survives the anchor-plane lease fixes and everything the content plane gained today — it is a Finder-item read failure in P2, not a scene-transport or a composition problem, and the host render is correctly drawing what it was told (nothing) rather than inventing icons. The matching symptom from the other side: the Finder's own Desktop window is published with itemTotal: 0 while seventeen icons are on the screen. So the Finder-item read fails whole rather than partially, and slice 5c item 1's control-panel half is still unexercised because of it — the classifier's positive branch cannot fire against a roster that is never there.

FOUND AND FIXED, 2026-08-06 (later the same day). It was never a read failure, and it was never on the guest. The Finder-item read works and has worked throughout; the host threw the answer away one poll after it arrived.

What the machine said. On a session-private emulated Power Mac G4 (VM nowvm-dtop, wire 5250, guest build 1bff0bd2ca39 verified in the hello before anything was believed), the four AppleScripts NOWMirrorSource.readIcons sends for the desktop container were replayed verbatim. All four answered osaErr 0, untruncated:

N  20
I  Trash            716  510  folder
I  Macintosh HD     736   28  disk
I  HELLO_CLAUDE.txt 608   92  SimpleText text document
...  (three pages, 8 + 8 + 4, plus the type pass for 14 files)

every item of desktop is not the problem; desktop-object is the only spelling that fails (osaErr -1753), and nothing sends it.

Where it went. desktopItems is a HOST contribution — no producer emits it, because a desktop icon is not a window, a control or a menu — so every structural scene omits the key honestly, and MirrorReplicaReducer.project() starts from that scene (var projected = latest). A window's items survive because the replica retains them per window record; the desktop plane had no home at all. So the roster was published for the fraction of a second between the enrichment and the next poll, and the poll erased it. refreshIconsIfStale then never re-fetched, because the layout key was already current — which is why it reads as a WHOLE failure rather than a flicker, and why the Finder's own Desktop window shows itemTotal: 0.

That also explains the "INTERMITTENT" entry above: the absent-item refusal fired once because a caller happened to ask inside that fraction of a second. Nothing was racing or expiring. Both entries describe one defect.

The fix. The plane is retained on MirrorReplica and carried through project() under the rule the windows already use: an ABSENT key retains, a PRESENT one — including an empty array — is an answer and wins.

The tests. MirrorReplicaReducerTests holds both halves of that rule. DesktopPlaneCrossingTests holds the whole crossing on the guest's own committed bytes — the roster capture and one structural scene from the same connection, which carries no desktopItems key — and was watched failing with the retention reverted: XCTAssertEqual failed: ("nil") is not equal to ("Optional(20)"). The older icon tests build the strings they parse, which is precisely how the roster read carried the blame for a week.

Two things this run did NOT settle, both stated as open.

  • The 2026-08-06 "reconfirmation" immediately above is VOID as evidence. Those two captures came from a python probe speaking scene.request directly to the guest, which never runs the host's icon enrichment at all. A guest scene with no desktop key is the producer being honest (docs/scene-producer.md lists the plane as absent by design) and says nothing about the Mirror. The finding it was cited for — that the defect survives the anchor-plane work — was true by luck.
  • A NEW defect, found while proving this one and not fixed here: the Finder type pass answers, and its answer is not a type code. OSADoScript renders its result in SOURCE form, so file type of comes back as the AppleScript literal «class APPL» (or string for a coerced one) and readIcons takes its first four characters. Every icon on that desktop carries type: "«cla" or "stri", which names no file type at all. It costs icon ART and not contents, which is why nothing ever failed over it. Pinned by an expectation in DesktopPlaneCrossingTests so it cannot be discovered a third time.

Status: TESTED, not watched in the window. The guest half is live and build-verified, the crossing is proven on those live bytes, and a host built from this commit was run against that VM with the Mirror open and logged no roster failure — but nobody photographed the render. Two instruments were unavailable to this session and both are worth knowing about: the local agent socket is one per user and was held by another session's host (unsafeEndpoint("Another New Old World host owns the local endpoint")), and screencapture has no Screen Recording grant from a shell. So desktopItems crossing into a DRAWN desktop is still owed one pair of eyes.

FIXED, watched on an emulator (not metal): transitions start could not arm, by any route (2026-08-05)

Found by driving, on the first live run of the verb, and fixed and re-driven the same day. While it stood, P5's plane could never publish.

run_start read its target with now_json_find_string(json, "name", ...) where json is the WHOLE request — and every request envelope carries "name" already, as the verb's own name:

{"type":"command.request","id":101,"name":"transitions",
 "args":{"op":"start","serialHi":0,"serialLo":34734082}}

First match wins, so the target was always the string transitions, which is never a running process. The by-name route is tried BEFORE the serial/front/a5 selector, so it short-circuited every other route too. Proven three ways on a live guest, all answering the same by-NAME refusal:

  • no target at all → no-process (should have fallen to the selector);
  • serialHi/serialLo naming a real running process → no-process;
  • name: "New Old World", a real running process → no-process.

The contract already forbade this and the rule was not followed. hide names its argument target and says why in the contract, verbatim: "target and not name — see the arg-key rule in the preamble above." transitions chose name and collided with the envelope.

The fix. The arg is target (contract first, then both faces). The parse moved out of transitions_cmd.c — where it sat in a static function above #include <Carbon.h>, unreachable to any host compiler — into transitions_logic.c, beside the console line's grammar that was already there for exactly this reason. The console face never went through JSON at all, which is why transitions start Finder typed at the machine kept working throughout.

Swept the class. All 41 verbs' declared args checked against the envelope keys (type, id, name, args, line): transitions was the only collision. resolvedVia: name deliberately keeps its word — it names the resolution MECHANISM, and only a request is scanned flat.

Why no test caught it, and what now does. The native test exercised the logic BELOW the parse, with a NowTransitionsStartReq a test had filled by hand, so the suite was green while the verb could not arm. transitions_args_test.c feeds the parse whole request frames — envelope and all, because with the args object alone the collision is invisible, which is precisely how it shipped. contract_arg_key_source_test.py holds all 41 verbs to the rule from the contract itself; two verbs have now broken it (launch on metal, transitions on an emulator), so it is executable rather than prose a third time.

Re-driven on the live emulated Power Mac G4 (build c39f3e093af3, verified in the guest's own hello before believing anything it said):

-> {"type":"command.request","id":103,"name":"transitions",
    "args":{"op":"start","target":"New Old World"}}
<- {"cmd":"start","a5":"0x1f21cb60","resolvedVia":"name",
    "process":"New Old World","expiry":116845,"now":113245,
    "requested":true,"armed":false}

Then activity.passes moved — 1071, then 5572, then 9005 — which is the resident's own word that it ran INSIDE the armed process and agreed. The ring's reader and writer have now met on a machine.

FIXED, watched on an emulator (not metal): a transitions drain reply lost its own records (2026-08-05)

Found on this plane's first ever drain, immediately after the fix above made a drain possible at all. The reply carried the records AND threw them away:

{"cmd":"drain","records":[ …22 real records, 2853 bytes… ],
 "cursor":0,"nextCursor":22,"records":22,"lost":0,…}

records twice in one object — the array, then the count in the tail. A duplicate key is legal JSON and silently lossy: every conforming parser keeps one and drops the other, with no error anywhere. Python kept the integer. So transitions drain looked like it returned a count and no data, while the records were on the wire the whole time.

It is the arg-key rule in the REPLY direction — a key stated twice is a key lost. qdtrace names its array ops and its count records and never met this; this verb named both the same word.

The fix renames the tail count to count, declared in the contract beside the array. transitions_reply_source_test.py reads every reply in transitions_cmd.c and fails on any key stated twice in one object — source text because run_drain assembles one object across two snprintfs above <Carbon.h>, the same reason GuestWireConformanceTests asks for a fixture there.

Re-driven after restaging; a standard parser now reads both:

{"cmd":"drain","records":[{"ticks":113305,"seq":1,"kind":4,
 "kindName":"heartbeat","a5":"0x1f21cb60","value":"0x1ecfd8d0",
 "previous":"0x1ecfd8d0"}, …],"cursor":0,"nextCursor":4,"count":4,
 "lost":0,"dropped":0,"pending":74,"writeCursor":74,"more":true}

UNVERIFIED: no non-heartbeat transition has ever been recorded (2026-08-05)

The plane is proven to produce, but everything it has produced is kind 4 heartbeat — the resident's one-a-second pulse on a quiet machine. No windowList, frontProcess or menuList record has been observed on any machine, so the three kinds that are the point of the plane are still unexercised end to end.

One attempt, recorded because the failure is informative rather than a defect. With NOW armed, the front process was switched to the Finder and back over the wire. The drain shows heartbeats either side and a gap where the switch was — seq 125 at tick 141725, seq 126 at tick 141915, 190 ticks (~3 s) apart against a 60-tick cadence. NOW was backgrounded and got no time, so from the armed process's own event passes the front process never appeared to change. That is exactly the sampler limitation contract/event_tail.h states — "something raised and dismissed between two event passes is STILL missed" — observed for the first time.

What would test it properly: arm a process that keeps pumping while another comes and goes — the Finder is the obvious one. That attempt refused anchor-plane-absent: the anchor plane captures an anchor only when the target pumps while the plane is claimed, and the sequence for getting a FOREIGN process anchored before arming it is not established here. Arming NOW itself works because it is the process doing the arming.

A rig note worth keeping. Driving stopped when the human's own host app took port 5271 (lsof -nP -iTCP:5271 showed New Old World PID 29286 holding the listener with the guest connected to it). Everything above was measured before that, and every reply quoted carried build c39f3e093af3. This is the collision AGENTS.md warns about, met in practice.

FIXED, watched: a transitions refusal was unparseable JSON (2026-08-05)

Also found on that first live run, and fixed and re-driven the same session. transitions start refused no-process, whose message reads nothing by that name is running (see "ps") — and error_json interpolated it into "message":"%s" unescaped, so every client got a JSON parse error at column 124 instead of the reason. The refusal was correct and unreadable, which is the worse half of the two.

It is on an ERROR path, which is where a missing escape is least likely to be exercised and most likely to matter — a caller meeting it is already in trouble. qdtrace_json.c does it correctly with now_json_escape and this file simply did not copy it. Swept the siblings: hide, quit and front carry the same quote-bearing sentence and DO escape it, confirmed on the same live machine, so this was one emitter and not a class.

Verified by rebuilding, restaging onto the running guest and re-driving: the refusal now parses.

MEASURED: the batched classifier is cold-loaded and the numbers have not moved (2026-08-05)

The entry above says the transport rebuild is "fixed in code and none of the fix has run on a Macintosh", and that nothing has been cold-loaded and no panel reclassified. Half of that is now closed and the other half is the finding.

Cold-loaded: yes. scripts/spin-up-ppc staged NowExt at 69 716 bytes (against 67 741 for the previous build) and cold-booted; the resident came up lifecycle: active, cap 31, build fingerprint f172318dda79798d9025a2e37e310cabd75c10d9.

Reclassified: no. Date & Time reads exactly as it did before.

window items knowledge: known kinds determined definition
Date & Time (foreign) 41 12 12 system × 29
NOW's own Workshop 53 53 53

Identical to the pre-batching reading of the same panel: 29 unknown, 12 known, and all 29 undetermined ones reporting system. So the batched transport is resident and the classification count is unchanged.

Two things that disagree, informatively. definition says those 29 are standard CDEFs loaded from the System file. kind says it cannot determine what they are. If they really are stock Toolbox controls then kControlKindTag should name them — so either the host never asks for the batch, or it asks and the resident does not answer. That is the next question, and it is a narrow one.

And the inverse pair is worth putting beside it, because together they say the two mechanisms fail on opposite populations:

  • kinds: NOW observing itself 53/53; a foreign panel 12/41.
  • refs: NOW observing itself 0/53; a foreign panel 41/41.

Whatever is wrong is about the boundary between NOW and a process it does not own, in both directions at once.

UNVERIFIED: the classifier's transport was rebuilt; nothing has watched it classify a real panel (2026-08-05)

The three transport defects below are fixed in code and none of the fix has run on a Macintosh. P2 gained a second cell for batched control classification: one request now types a whole window instead of one control per scene, each record naming the exact ControlRef it describes. The priorities were inverted — a class fact is the prerequisite for a list request and now outranks it — and the front-only gate became a front-first priority, so background panels can be filled in at all. contract/peek_table.h, ext/src/now_semantic{,_logic}.c, now-guest-shared/src/now_semantic_guard.c, now-guest-ppc/src/peek/semantic_{client,policy}.c.

What is proven: scripts/test-all is green — 111 native tests, all three cross-builds (PPC guest, 68K guest, and the 68K flat INIT), and the host gate. The extension links the batch resolver: now_semantic_batch_ apply/resolve/ready/verdict all appear in NowExt.map, and appending a deliberate error to ext/src/now_semantic.c was watched fail the build, so that gate genuinely covers the resident code.

Five properties were watched fail under mutation — the resolver's one-walk-per-reply cost argument, the guard's refusal of a reply naming one control twice, the policy's join-by-named-control, the drained-window suppression, and the priority ordering. That last one caught a real error: the first draft had the ordering inverted and the source guard passed, because the assertion encoded the same mistake as the constants.

What is not proven, and is the whole point of the change:

  • No panel has been classified. The measurement that motivated this (1 of 122 controls determined) has not been retaken. Until the corpus is recaptured against a rebuilt extension that has been cold-loaded, the claim that this fixes the starvation is a design argument that compiles.
  • Nothing has been cold-loaded or run. A built INIT is not a serving INIT; docs/p2-semantic-evidence.md records that the previous P2 checkpoint passed every gate and then produced no live list cells at all on the first cold-load sweep.
  • How much of the corpus is genuinely custom-drawn remains unknown, still bounded by the UnsupportedCustom branch — which had been exercised 1/122 times precisely because nothing else got a turn.

The next step is a cold-load sweep on a machine with an m68k toolchain: rebuild the extension, recapture the ten panels, and count determined kinds. The original measurement below is kept as the baseline that number will be compared against.

(fixed, pending verification) BROKEN: the control classifier gets one shot per scene, and 121 controls never got one (2026-08-05)

Slice 6's opening measurement was supposed to size how much of the 190 undetermined corpus items are genuinely custom-drawn. It found something else: of 122 Control Manager controls in the ten-panel corpus, exactly one carries a determined kind — Monitors' resolution list. All 117 other determined items came from the DITL item-type byte and never involved a CDEF at all.

The classifier is not missing. ext/src/now_semantic.c :: classify() reads the Appearance Manager's public kControlKindTag through GetControlData, inside the target process's own context via the resident's jGNE patch, and resolves fourteen control families; its signature != kControlKindSignatureApple branch is an authoritative standard-versus-custom verdict. The plane was armed and serving during the capture (capabilities: 15, requested: 7, active: 7 in every manifest.json).

It starved on transport. contract/peek_table.h carries a single NowPeekSemanticCell semantic; — one request per scene, for one control. In now-guest-ppc/src/peek/semantic_client.c, control classification is the lowest-priority claimant on that cell (offer(10, …ControlClass) against offer(20, …ListCells) and offer(30, …SystemMenu)), and only the front process may spend it. Date & Time's window carries 21 controls and was front for one scene.

Why it matters beyond the count: the Date & Time "radios drawn as push buttons" red has been read as a knowledge gap that slice 6 would close by replaying draw ops. It is not. The right fix is a batched or multi-control request op and a priority that reflects what the answer is worth — a transport change, not a drawing one. How much is genuinely custom is still unknown, bounded by the UnsupportedCustom branch, which has been exercised 1/122 times.

Recorded in full, with the derivation, in the gap ledger. docs/open-issues.md already carried the adjacent symptom under "P4's plane is intact; the reason named for its silence was wrong" — this is the second, independent measurement of the same shape.

UNVERIFIED: semantic.definition ships, and its histogram has never been taken (2026-08-05)

The walk now reads contrlDefProc (offset 24) and reports which heap the handle came from — system, application or indeterminate — as Scene.Semantics.definition, IR v2 additive. It is a resident-free answer to the same standard-versus-custom split, with no per-scene budget: it classifies every control in one walk.

Nothing has run it against a Macintosh, or against an emulated guest. Status is: PPC cross-build green, scripts/test-native 110/110 green (including the new axdefproc_test, watched failing under three separate mutations), IRFreezeTests green and watched failing when the addition is unrecorded. That is tested, not metal-verified, and the whole point of the field is a number nobody has yet obtained.

Two things to check on first live contact, in this order:

  1. How large the indeterminate column is. The classic Control Manager keeps a variation code per control. If it rides in this field's high byte, the raw longword lands in neither zone and lands there. A large indeterminate column is evidence about the field's layout, not a failed read — and nothing is masked on a guess, because a 24-bit mask on a 32-bit-clean machine would manufacture a clean histogram out of a real question. axdefproc_test pins that case.
  2. Whether system and the resident's kControlKindSignatureApple verdict ever disagree. They answer the same question by independent routes; a disagreement names which one is wrong. There is currently one control in the corpus where both could be asked.

THE STALE-REF REFUSAL NO LONGER HOLDS THE LANE (2026-08-05)

Fixed, tested, and watched on an emulated guest. During the first human drive, winact refusals against windows the Mirror was still displaying held the mutation FIFO for its whole 15 s timeout. Measured in ~/Library/Logs/NewOldWorld/acts.log: a close refused at 02:03:47.753 held the lane 15 128 ms, and the next click on the same window waited 8 025 ms behind it. The cause was a substring test — a refusal was treated as "the effect may have landed" unless its wording contained "was not sent" — so a refusal the guest raised before it armed anything became non-terminal. The same reading produced a false green: a refused close was recorded confirmedAfterRefusal 3 ms later, settled by a scene that showed the window absent because a PREVIOUS close had closed it.

Two changes, both host-side. A guest act refusal now carries a typed AgentIntegrationProjectionFailure.Reach, read from whether the guest's error carries a correlation — the guest registers one only after now_act_submit, so its absence is the machine saying nothing was armed. And the target is re-read against the newest scene at dispatch time rather than the displayed one, so a click that waited in the queue while the window closed is refused here instead of on the machine.

Answered along the way, because the arc began by suspecting it: window references do not churn on republish. The guest interns them (now_obs_intern / identity_same excludes only minted_ticks), and one token in the same log confirmed an act fifteen seconds and a dozen scene generations after it was minted.

The drive that proved it, and how to run one beside somebody else

Driven headlessly on an emulated Power Mac G4 at 03:09 on 2026-08-05 — three closes at one Finder window, 300 ms apart, which is the incident's own shape. The third one:

NOT DISPATCHED: the scene moved on — that window is no longer on the
machine, so the act was not sent. Read it again.
NOWBASE act ... outcome=refused depth=2 waited_ms=8596
                dispatch_ms=12 settle_ms=- total_ms=8608

Twelve milliseconds and no guest round trip, against ~89 ms and a 15 s lane hold before. And refused, terminal: the 02:04 equivalent of this act was recorded confirmedAfterRefusal by a scene that another close had produced.

The conservative branch earned its keep in the same run. The FIRST close was refused the target served the request and did not arm — a refusal the guest raises after now_act_submit registered a correlation, so this side calls it unknown and keeps waiting. It held the lane the full 15 s… and then settled confirmedAfterTimeout at 16 503 ms. The close had landed. Had the rule been "any refusal is terminal", that act would have been written off as failed while the window was closing.

Two hosts on one Mac, without touching each other (this run shared the desk with another session's verification host, which owned 5250):

  • NOW_PREFS_SUFFIX=<slug> gives the second host its own settings and guest registry — it already existed for exactly this, see ProductIdentity. Put the port in that domain: defaults write dev.newoldworld.now.settings.<slug> listenPort -int 5262.
  • The guest dials 10.0.2.2:5250 from saved preferences with auto_connect on, and QEMU refuses a guestfwd on 10.0.2.2 because it is the gateway. Move the gateway instead, and the address is free to forward:

    TBT_EXTRA_HOSTFWD="net=10.0.2.0/24,host=10.0.2.5,dns=10.0.2.6,\ guestfwd=tcp:10.0.2.2:5250-tcp:127.0.0.1:5262" tools/launch --instance 91

The guest's own preferences never change, and its dial cannot reach the other host. guestfwd connects eagerly at launch, so start the host first. - The agent socket is the one thing that cannot be shared. It is <darwin-user-temp>/dev.newoldworld.now-agent-<uid>/host.sock, one per user, and FileManager.temporaryDirectory ignores TMPDIR on macOS even though tools/now-agent derives its path from TMPDIR and says so in a docstring. The second host logs local agent integration unavailable and runs on without an MCP surface. This drive used a build-only patch (an env override, reverted before committing) to move it. A supported override belongs in the product, and the two sides should agree on how the path is computed.

A second false green, found by the test that closed the first. Writing the brokered-path test surfaced the same defect entering by another door: AgentIntegrationUnavailable — "there was nobody to ask" — said nothing about reach, so an act refused because NO GUEST WAS CONNECTED was held open and then confirmed by a later scene. Nothing had been sent. That type now carries a reach too, defaulting to notSent, because every one of its statics means the request never left this host; the single exception is a guest that vanished mid-act, which is built from a failure and inherits its unknown. Both directions are tested, and the passthrough was watched to fail.

Still open from this:

  • A refusal AFTER the act plane armed nothing still holds the lane 15 s. the target served the request and did not arm carries a correlation, so this side calls it unknown — and in the 03:09 drive that was right: the act settled confirmedAfterTimeout at 16 503 ms, it HAD landed. But the guest can sometimes tell the two apart (kNowActNotArmed means no patch was installed, so no click can be delivered), and it says so only in prose. If that distinction is worth the 15 s, it belongs in the contract as a field, not in host guesswork over the guest's sentences.
  • A guest mis-attribution, noted not fixed. now_act_submit returns kNowActNoExtension before it registers a correlation, and act_cmds.c:546 answers that with reply_registered_status, which attaches the PREVIOUS command's correlation. It is conservative for the host rule above — a stale correlation reads as "may have landed", which only costs a wait — but it is wrong.

FIXED, TESTED: an agent-driven act was journalled as a person's when it waited (2026-08-05, test landed 2026-08-05)

An act that arrives while an observation is in flight is deferred and re-enters NOWMirrorSource.perform when the cycle clears — and it used to do that through the one-argument overload, whose source defaults to .human. So an MCP act unlucky enough to land mid-cycle was recorded as a hand-driven one, undoing the earlier attribution fix by a path older than it. The journal is the only thing that can tell the two faces apart after the fact.

Measured: a finderOpen driven entirely over the agent socket settled confirmed and recorded source: human. It was the only record in that host's journal, so there was nothing to confuse it with.

The argument is passed through, and the owed regression test is now here: MirrorFaceParityTests.testAnMCPActHeldMidObservationIsStill RecordedAsMCPs, with a human twin beside it so it cannot pass by making every act one face. Watched failing by reverting the argument, and again by reverting the older executor fix — both name the field.

The earlier attempt could not get a harnessed NOWMirrorSource to the broker at all. What it needed was to leave the content join OPEN: that is the deferred case, and until the join is released the act is held and has no record. The cycle harness that does it is shared test support (MirrorSourceTestSupport.swift) rather than private to one file.

RESOLVED — NOT the cause: the MCP drive path holds an engine (2026-08-05)

The suspicion was that HostAppState.mirrorSource hands the drive service a source whose shadowEngine is nil, so every MCP act takes the direct path and can never settle. It does not, and it cannot. Two independent arguments, both cheap to re-check:

  • shadowEngine and pinnedGuestKey are written together in start() and cleared together in both of stop()'s branches — nowhere else. So a nil engine implies a nil pin, and pinnedActionRefusal() refuses by name whenever the pin is nil, before the engine is consulted. On the MCP path a nil engine produces a refusal with a sentence, never id: "direct". (Before the source has ever started there is no published scene either, and the drive is refused now-mirror-snapshot-unavailable without reaching perform.)
  • The live evidence refutes it directly, and it was already in this file. The act that answered id: "direct" settled confirmed on a typed windowNamedPresent postcondition — which a nil engine cannot mint — and recorded source: human, which only the deferred branch produces. One act, both facts, and each says the engine was bound.

The real cause is the deferred branch, and it is the third instance of one defect rather than a new one. A held act returns without enqueuing, so no broker record exists yet, and MirrorDriveService inferred "no new record" == "took the direct path" — answering awaitsObservation: false, which tells the only face that cannot see the screen to stop waiting, about an act that was on its way to settling. All three of the arc's drive-service defects are that same reconstruction: a refusal read as a dispatch, a held act read as a direct one, and a record found by resemblance rather than by id.

Fixed by not reconstructing it. perform answers MirrorPerformDispositionrefused / brokered(id) / held / direct — and the service reports what it is told. Guarded by MirrorDriveServiceTests.testAHeldActIsNotReportedAsTheDirectPath and, through a real NOWMirrorSource mid-observation, MirrorFaceParityTests.testAnActHeldMidObservationIsNotReportedToMCPAs TheDirectPath. Both watched failing against the pre-fix reading.

What this means for the measurements already recorded: nothing needs re-reading. The worry was that the headless numbers behind slices 3 and 4 described the direct path rather than the shared one. They did not — the engine was bound throughout. What was wrong was the REPLY to the caller, not the path the act took: an act reported id: "direct" still went through MirrorActionExecutor, the broker and the typed postcondition. Slice 3's central claim stands, and it now has a gate (MirrorFaceParityTests) instead of a ceremony.

Still unverified: none of this has been driven live. No VM was stood up for the fix — the cold boot cannot run unattended today. The reply an agent now gets for a held act (id: "held", outcome: queued, awaitsObservation: true) has never been seen by a real MCP client.

FIXED: now/ had no commit hook at all, while AGENTS.md said one was enforcing the rules (2026-08-06)

Found while building the extension bake gate, and worth its own heading because the bake gate is not the interesting half.

AGENTS.md § Git says a commit on main is refused, and names the mechanism: "This is enforced (.githooks/pre-commit, plus a PreToolUse hook on Write/Edit/Bash)". CLAUDE.md repeats it. The file did not exist. .githooks/pre-commit was added to this repository on 2026-08-06 by 543b06af; git log --diff-filter=A on the path returns that one commit and nothing before it. For however long that sentence had been in AGENTS.md, half of a two-part enforcement was prose. The PreToolUse hook is real, which is why nobody noticed — agents were stopped, so the floor appeared to hold, and a human committing on main from a terminal would have met nothing.

The durable shape: a rule that names its own enforcement is read as enforced, and nothing re-checks the naming. Every other stale-oracle finding in this file is about data going stale; this one is about a claim of mechanism going stale, which is worse, because the whole point of writing the mechanism down was so a reader would not have to check.

git config core.hooksPath .githooks is not automatic in a fresh clone or a fresh worktree, so the file existing is still not the same as the hook running — tools/setup-hooks in the parent tree is the one-time step. Anyone auditing this should run the check rather than read this paragraph, which is the entire lesson repeating itself.

The gate that occasioned the discovery — a resident commit refusing to land without a verified, cleanly shut down image, with an explicit recorded deferral — is written up under "That rule went unfollowed for three days, and is now enforced (2026-08-06)", inside the CYCLE 25 section below. It is filed by cycle rather than by date, which is why it is hard to find; this heading is the pointer.

Three rig facts that each read as something else (2026-08-05)

None of these is a defect in NOW; all three cost time because they present as a hang or as a broken change. The second is now fixed and the third turned out to have a different cause than the one first written down here.

  • A worktree is too deep to host a VM. A UNIX socket path is capped at 104 bytes and now/.claude/worktrees/<branch>/run/<pid>/qmp-ui.sock spends 108. QEMU refuses to start and says so only in run/<pid>/qemu.log; spin-up-ppc's own output ends after "boot a fresh, session-private clone" with no error, which reads as a boot that hung. This is the DEFAULT path for every agent, because agents work in worktrees. scripts/spin-up-ppc now checks and names the cure (NOW_SPIN_RUN=/private/tmp/nowvm-$$).
  • spin-up-ppc could not finish its clean shutdown. FIXED, and the route is the Shutdown Manager rather than the Finder. Rule 1 needs the guest to shut ITSELF down before the cold boot an INIT requires, and the lab's tools/shutdown-guest asks the Finder through an agent's script verb — which the canonical baked worker does not have. Its hello lists 24 tools (putclick, key, type) and script is not among them, so the run stopped with "the agent refused the script verb" and left the VM up with the extension staged and never loaded: every plane then reported unsupported/gen0 and needs-restart, which reads as a broken build rather than a boot that never happened.

NOW now stages its own applet, tools/guest-shutdown, whose entire body is ShutDwnPower() — the Shutdown Manager call the Finder itself ends up making, which runs the registered shutdown procedures, flushes and unmounts the volumes and asks the power manager to cut power. tools/shutdown-guest.py quits the front application first (Cmd-Q through the worker's key verb, the one posted-event route measured working) and then launches it through the worker's launch verb. It is 68K because the Shutdown Manager is CALL_NOT_IN_CARBON and the application's own toolchain cannot compile the call at all.

Measured: launch to QEMU exit, 6 s. Booting the same qcow2 straight afterwards reached the anchor in 42 s with no Disk First Aid modal — which is the assertion that matters, because the whole reason not to use QMP quit is the unclean-volume bit.

What it does NOT do, by construction, is send quit AppleEvents to running applications; the Finder does that before it calls the Shutdown Manager and this skips it. That is why the front application is quit first, and why this is a rig instrument and not a way to stop a machine somebody is using.

One failure, and it did not reproduce. Four shutdowns went in 6 s each; a fifth did not go at all, and the worker then timed out on observe, so the machine was wedged rather than slow. That one had been driven through roughly fifteen actselftest calls first — each claiming and withdrawing the act plane, which is the path the entry below says is broken — and the run directory was deleted before a screendump could be taken, so there is no evidence and no cause. The same tool then shut down a comparable machine (extension resident, NOW launched and quit) in 6 s. Recorded because a rig that hangs one time in five is worth knowing about even when nothing was learned; if it recurs, screendump BEFORE cleaning up. - Nothing outside this machine can reach its human interface, and the reason is not ADB. The earlier entry blamed an ADB mouse; info qtree says has-adb = false on both macio-newworld and via-pmu, so mac99,via=pmu here has no ADB keyboard, no ADB mouse and no ADB power key at all. Every input device is USB: the machine's own usb-kbd and usb-mouse (enumerated, addresses 0.23 and 0.24) plus a second usb-kbd that -device usb-kbd puts behind a hub, which OS 9 never enumerates — both it and its hub still sit at address 0.0.

Three separate measurements, none of them a route in:

  • QMP keyboard input never arrives. Not "lands in the wrong place" — arrives nowhere. With Key Caps open and frontmost, neither send-key (plain, held, and with modifiers) nor input-send-event key-down/key-up pairs light a single key or put a character in its field. Cmd-N in the Finder makes no folder; Cmd-Shift-3 makes no Picture 1.
  • abs pointer events are refused, and now for a stated reason: "Input handler not found for event type abs" is what QEMU says when no absolute pointing device exists, and none does. rel motion is accepted and carries OS 9's pointer acceleration, so counted steps do not land where they are aimed (measured earlier the same day: aimed at (270,119), ended near (10,63)).
  • The worker's click verb closes a menu without selecting from it — a posted event pair is already up by the time MenuSelect's tracking loop reads the real button state.

So an agent still has no route to a menu selection on this rig, and that is now a UI-driving limitation rather than a blocker: nothing in spin-up-ppc needs a menu any more.

Two things this did NOT establish. Why the key events vanish is a hypothesis, not a measurement — the un-enumerated second keyboard is suspicious, but nothing here proved QEMU routes to it. And whether dropping -device usb-kbd from the lab's tbt_qemu_boot would restore QMP keyboard input was not tried; it is the cheapest next experiment, and it would unblock keyboard driving generally. Even if it worked it would not give a shutdown: a USB keyboard has no power key here either, and Shut Down has no Command-key equivalent in the Finder.

A fourth, added 2026-08-06: sessions collide on wire ports, and the collision is silent in both directions. Several agent sessions were working this tree at once, each spinning up its own guest, and the port each dialled was chosen by habit rather than by allocation. Two consequences were observed and neither announced itself:

  • A quit landed on another session's guest. The wire carries no identity a sender checks before acting, so a control verb sent to "the guest on port N" reaches whichever guest is on port N. From the sender's side it succeeded. From the other session's side a machine died mid-task for no reason it could see.
  • A guest answered that was not the build under test. This is the same failure AGENTS.md already names for the metal gates — "any session's VM, running any branch's build, can answer your listener" — arriving on the emulator, where it is more likely rather than less, because emulated guests are cheap and everyone starts one.

The metal rule is the cure and it generalises: pick a port nothing else is dialling, and assert a capability only the build under test has before believing anything it says (requireTheBuildUnderTest() is the pattern). A session should also say which port it took. Michelle's own stack sits on wire 5540 and is not to be touched; docs/68k-metal-runbook.md is the procedure for telling contention from a defect, and it applies here even though no metal is involved.

BROKEN: the anchor plane is active and binds nothing (2026-08-05)

Found by the first cold boot that ever got this far. Until the guest could be made to shut itself down, scripts/spin-up-ppc stopped before the reboot, so nothing had ever interrogated a machine with the unified NOW Extension resident from a clean start. It does now, and the resident is plainly alive:

mirror -> lifecycle "active", capabilities 31, all five planes
          supported, resident fc6e0946bde9271, table 4256 bytes
qdtrace op=status -> plane {format 2, length 65728, ringCap 65536}

And nothing can be addressed inside it:

actselftest -> no-such-process        (every attempt, over minutes)
axsnap      -> front "New Old World", bind "no-plane",
               hasWindows false, hasMenus false

The two halves disagree about the same planes. actselftest calls now_act_ready, which claims anchors + act; asked immediately afterwards in the same connection, mirror reports requested: 5, active: 5 — so the claim reaches the table and the resident is serving both. Yet now_ax_bind_process still answers kNowPeekReadNoPlane, which qdtrace_cmd.c words as "the window-anchor plane is not armed". The client's plane gate and the resident's own report of the same word do not agree.

Also note bind: "no-plane" and not not-pumped: this is not the settle window staging-path.md records, where a just-launched application is briefly invisible because it has not pumped. That one cleared in seconds; this does not clear at all, and it is NOW's OWN application — the case every earlier entry here treats as the easy one.

Not measured, and worth doing first: whether this predates the unified extension (2026-08-03) or arrived with it. staging-path.md records actselftest answering abi-agreed on 2026-08-01, on the plane model that work replaced, so a bisect has somewhere to start. Read now_ax_bind_process in now-guest-ppc/src/axwalk/axprocess.c against whatever mirror reads for active, because those are the two words that disagree.

spin-up-ppc runs actselftest and prints its answer, but does not gate on it — a gate permanently red for a defect the rig neither causes nor can fix is a gate nobody reads. It moves back into the gate when this is closed.

2026-08-06, mechanism found and fixed — it was the writer heartbeat, and "binds nothing" was the flap's worst phase. The application renewed writer.heartbeat_ticks only inside peek calls, and the lease is 3 s (kNowPeekWriterLeaseTicks), so any wire cadence slower than that let the resident see a dead writer between requests and de-arm every plane; the next scene then claimed, renewed, and read arm_active synchronously — before the resident's next jGNE pass could re-echo — so the first walk after every quiet gap answered now_no_plane for every foreign process. Measured on a clone before the fix: 6/6 scenes at a 4 s cadence carried no foreign process, while axsnap in the same connection bound the Finder ok (a command pass renews and pumps, which is why the two halves disagreed). The fix is now_peek_idle() — the event loop renews once per pass, so the heartbeat proves the loop runs, which is the fact the lease exists to check. Verified on the same clone: five scene polls with 6 s silent gaps all carried the Finder and its windows, and its scroll bars arrived classified. Two things this does NOT close: the FIRST scene of a fresh connection still misses foreign processes (claim-before-echo, by design — the host's second poll covers it), and actselftest's no-such-process did not change and needs its own look. The scene should also say WHICH gate refused in its coverage reason (no-plane vs not-observed) — this defect survived two sessions because the wording hid the distinction.

BROKEN: a Finder-open predicts the wrong owner, so panels time out having worked (2026-08-05)

MirrorActionExecutor builds windowNamedPresent(owner: Finder, title: item) for every Finder-open. A control panel opens as its own application: a live snapshot taken while Date & Time was up shows that window owned by a process named Date & Time, not by the Finder. The postcondition is unsatisfiable, so the act burns its full 15 s timeout having succeeded. It holds only for FOLDERS, whose windows the Finder does own.

Measured in Michelle's drive: four open "AppleTalk" attempts, 15 s each, and because the mutation FIFO is one lane those timeouts stacked — waits of 15 611, 22 207, 22 106, 37 391, 51 786 and 49 281 ms behind the lane. One click waited 51.8 seconds. This one defect is most of the queueing that drive complained about.

finderItem.kind in the snapshot already carries folder-or-file, so a fix can predict a window for a folder and a PROCESS for an application instead of guessing one shape for both.

Related, and separate: an act that legitimately does nothing costs the same 15 s. Driving finderOpen "Date & Time" against the desktop (where it does not live) correctly opened nothing, and still burned the timeout rather than the Finder answering "no such item".

BROKEN: both Hide routes fail, each in its own way (2026-08-05)

Hiding an application works on a Macintosh; a person does it from the Application menu. Neither route NOW has reproduces it.

  • AppleScript — setting visible through the Finder's object model is refused: -10000, -10006, and in the 2026-08-05 drive osaErr -1753. Read-only there.
  • The Application menu, commanded — the menu is read correctly (Hide Date & Time, Hide Others, Show All, present and enabled at menu -16489), and driving row 1 through now_mirror_drive --gesture menuItem returns dispatched and changes nothing. The paired screendump shows the application still frontmost. InteractionPolicy already records this: visibility is kept typed so it "cannot fall back to commanding menu -16489, the route that reported success without changing the machine".

Why the act plane cannot reach it, understood 2026-08-05. menuact works by arming a trap patch so the target application's own MenuSelect returns the chosen row — which is why it drives the Finder 8/8. The Application menu is system-owned: the Process Manager performs the hide, and the front application's MenuSelect never sees it. So this is not a flaky route, it is the wrong mechanism for this menu.

A real positional click is not possible today either: the guest has no positional click verb. mouseloc is a READ (input/input_cmds.c), and the act plane delivers menu choices by arming a patch rather than by moving a pointer.

THE ROUTE THAT EXISTS: ShowHideProcess, weak-linked from the Carbon app. An earlier version of this entry said the Process Manager's visibility call was absent from the toolchain under any spelling. That was wrong, and the error is worth keeping: the sweep checked toolchain/universal/libppc/libCarbonLib.a and two CarbonFrameworkLib archives, and never toolchain/multiversal/libppc/libCarbonLib.a. There are TWO CarbonLib archives of different vintages here and only one was looked at. Verified 2026-08-05:

  • present in Retro68/ImportLibraries/libCarbonLib.a and toolchain/multiversal/libppc/libCarbonLib.a, together with IsProcessVisible;
  • absent from toolchain/universal/libppc/libCarbonLib.a — and powerpc-apple-macos/lib/libCarbonLib.a is a SYMLINK to that one, so the linker currently resolves to the archive without the symbol;
  • the split is Universal Interfaces 3.4 (headers, on the include path) versus 3.4.1 (the richer archives). ShowHideProcess did not exist until 3.4.1 — zero occurrences in 3.2, 3.3.2 and 3.4.

pascal OSErr ShowHideProcess(const ProcessSerialNumber *psn, Boolean visible), THREEWORDINLINE(0x3F3C, 0x0060, 0xA88F), cited from UI 3.4.1 Processes.h ll. 542–545 and Apple's Process Manager Reference (2007-12-04, p.19). Availability: CarbonLib 1.5 and later — our floor is 1.6. It is a WEAK import, so the address is tested before it is called.

And this entry's other worry is settled. ProcessInfoRec carries no visibility field at all, so GetProcessInformation cannot disagree with the Application menu; the state is the LAYER's visible flag, which is also what Mac OS 8's own AdjustApplicationMenu tests when it decides whether Hide is enabled. One flag, both readers. IsProcessVisible (selector 0x005F) is the read-back.

Why menuact could never have worked, precisely. For a system-owned menu, MenuSelect calls SystemMenu (trap $A9B5) and returns 0 in the high word to the application; the Process Manager's patch on _SystemMenu performs the hide. Arming a patch on the front application's MenuSelect therefore skips the only code that acts. The wrong trap, not a flaky one.

Do not reach selector 0x0060 by raw _OSDispatch. Apple's dispatcher does no bounds check: an unimplemented selector does not return an error, it reads past the table and rtses into whatever that longword holds — in a resident, in every application's context, an unrecoverable crash rather than a paramErr, with no way to probe first. That route is last. Plan: 2026-08-05-010 § C.

Until something is watched working, Hide is UNBUILT rather than broken. Its other half — that it should not hold the shared lane for 15 s rediscovering a route already known to fail — was fixed on 2026-08-05; that changed the cost of failing and nothing about whether Hide works.

UPDATE 2026-08-08: the host no longer depends on Finder/P3 to supply the system-owned Application menu. It synthesizes the menu from the live process roster and sends Hide through the guest's Process Manager verb. Hide Others and Show All are explicit compositions of that same exact per-process route, with each result read back rather than treating dispatch as success. The projection and composition are Tested; the underlying visibility mutation is still not metal-verified, so the historical end-to-end status above does not silently become green.

P5, THE TRANSITION TAIL: RESIDENT WRITES, NOTHING DELIVERS YET (2026-08-05)

Built and gated; never executed on a machine. A ~2.2 s scene cycle cannot see anything shorter than 2.2 s — measured across 60 cycles: 783 ms idle, 92 ms request, 315 ms decode. An alert raised and dismissed between walks leaves the machine as it found it and the Mirror never knew. P5 is a transition SAMPLER at the guest's own event-loop rate, riding the jGNE pass that already reads the window list, menu list and current A5 for the anchor plane.

What exists: the contract (contract/event_tail.h), a ring in a system-heap block behind one appended table word, the resident writer (ext/src/now_event.c), the guest's ring reader, and a fifth transitions plane reported end to end. 109 native tests, three cross-builds, full host gate.

UPDATE 2026-08-05, claude/p5-transitions-delivery: the guest half of that gap is closed and the machine half is not. The contract declares a transitions verb — status (the default, which moves nothing), start, stop, drain — modelled on qdtrace and simpler than it in three stated ways; the PowerPC guest answers it on BOTH faces off one implementation; and a native test covers the command layer's own decisions. So something arms the plane now, and a message carries records.

What is still not proven is everything that matters. No guest has been stood up since the verb landed — the VM cold boot could not run unattended — so no record from this ring has ever been observed crossing the wire, on any machine. status has never answered from a real block, start has never had a resident agree with it, and drain has never returned a record. The reader and the writer have still never met: the ring's two halves are tested separately, against fixtures, by the host compiler. Treat every claim below the contract as BUILDS, not tested and not metal-verified.

Two things a first live run should look at, because they are where the guest half is most likely to be wrong: activity.passes staying at zero while a request reads live (that is an arm that named the wrong world, and it is the one diagnostic this design leans on), and whether reader_cursor moving forward on drain actually makes dropped behave — the forward-only rule is native-tested against a fixture and has never been read by the resident it exists for.

The host consumes none, and that is deliberate rather than pending: the host consumer is a later slice, declared in docs/mcp-coverage.md's gap table with its reason. Until it lands the plane is reachable only by a person typing at the guest or by a direct command.request.

Before that update, the status was: nothing arms the plane, no contract message carries records to the host, and the host consumes none. So a cold boot proved the INIT loads, allocates, publishes cap bit 4 and boots clean — and proved nothing about the tail. That last sentence is still true.

Two things found while building it, both mine:

  • The plane roster was made breaking in the wrong direction. Moving the contract from exactly-four plane rows to exactly-five meant a newer host refused an older guest's honest four-row report — "The Mac's mirror facts do not match schema 1". This project's rule is that an older reader refuses a NEWER message, never the reverse. The contract now says four rows or five and the reader completes a missing trailing row as present-and-unsupported.
  • The snapshot outgrew one protocol message. Adding Finder items pushed a scene of several panels plus a desktop past the 64 KB ceiling, past which the writer throws and the connection closes with NO reply — so now_mirror_snapshot stopped answering while status and lifecycle still did, which reads as a broken host rather than an oversized payload. Bounded per window AND per snapshot now, both stated in itemTotal.

Also unverified: finderDeselect and dialogItem against an item that does something have never been driven; the INIT resource is 67 KB against a conventional 32 KB budget (pre-existing, and this image boots it).

BROKEN: the Finder cannot set visible, so no Hide act can ever land (2026-08-05)

Measured directly against Mac OS 9.1 (mac99, guest build a4a59d37d100) by impersonating a host with tools/askguest.py and running the production scripts verbatim. All three visibility mutations are refused by the Finder's object model, so Hide, Hide Others and Show All have never been able to work by this route:

  • hideFrontApplicationScript-10000 Finder got an error: Can't continue .
  • showAllApplicationsScript-10006 Can't set visible of every application process to true.
  • hideOtherApplicationsScript-10006 Can't set visible of item 1 of every application process to false.

visible is readable and read-only there: set v to visible of process "tbt-worker" answers false, and every attempt to assign it fails. The census confirmed the machine was unchanged after each attempt. This is the real blocker behind C27's "Hide Finder timed out"; the settlement rule was never the problem, and the dispatch it was waiting on could not have happened. A working Hide needs a different mechanism — the act plane driving the Application menu is the candidate, and it is what a person uses — not a repair to this script.

Fixed the same day, and separately: the census asked a question the Finder cannot answer inline. visible of candidate read straight into a & chain is an object specifier, not a boolean, and the concatenation raised -1700 Can't make visible of «class prcs» "tbt-worker" of application "Finder" into a string. AppleScript fails a script WHOLE, so the census returned no rows at all — which is why now_mirror_snapshot showed visible: null for every process and a coverage claim that blamed name ambiguity. Binding the property first fixes it; the corrected script was run against the same machine and returned all 7 rows the Finder can see. MirrorStateEngine.enrichVisibility's sequence guard, which the investigation began by suspecting, was never implicated: the coverage row that appeared in the snapshot could only have been written by a census that already passed it.

Still open, and the reason a complete census is not enough. The Finder is absent from its own process list — count of (every process whose name is "Finder") is 0, and name of process "Finder" errors — while the replica always carries the Finder as an application. So matched == Set(replica.applications.keys) cannot hold, coverage stays partial, and a visibility postcondition still cannot settle even with the script repaired. visible of application "Finder" does answer a real boolean (true), but it addresses the Finder application rather than a process row, nothing has yet watched it change, and a value that never changes would settle mutations falsely — the failure docs/mirror-drive-loop.md §2j exists to prevent. It is not to be adopted until something has watched it go false.

Unverified here: the repaired census has not been watched settling an operation end to end. The host app could not be run beside the one already holding the per-user agent endpoint, so the script was proven against the guest directly and the join proven by unit test, not the two together.

P4's plane is intact; the reason named for its silence was wrong (2026-08-05)

The symptom stands; the diagnosis attached to it does not. A human drive on 2026-08-04 (guest a4a59d37d100, resident 67d5ef43) found now_mirror_lifecycle reporting interaction generation 0 while structure sat at 613025 and content at 1522260, and not one act settling confirmed in eight minutes. That is real and still open.

The reason recorded alongside it was that the host's requested plane mask "flaps between 7 and 15" with "the interaction bit (8) clear", pointing the repair at MirrorControlModel.requestedPlaneIDs, MirrorPlanePolicyStore, and the arming path. That reading is wrong, and it points away from the defect. contract/peek_table.h states the bits once:

kNowPeekTableCapAnchors = 1u << 0   P1
kNowPeekTableCapTree    = 1u << 1   P2
kNowPeekTableCapAct     = 1u << 2   P4   <- the act plane, BELOW P3
kNowPeekTableCapContent = 1u << 3   P3

P4 sits below P3 because P3 asked for 1u << 2 while P4 already held it — the near-miss recorded under "Two planes asked for the same bit" (2026-07-31). So cap=15 requested=7 active=7 says P4 was requested and active; the bit clear in 7 is P3, whose request is a bounded lease by design (now_peek_claim_until, from qdtrace_cmd.c), which is exactly what a mask expiring and being reclaimed looks like. Host plane policy is not implicated by that line. now-guest-ppc/tests/peek_table_test.c now pins each bit's value, so the same misreading fails a gate rather than an evening.

Where the defect actually has to be. mirror_probe.c :: plane_generation returns, for P4, table->act_v2.resident_generation. That word is written in exactly one place — now-guest-shared/src/now_act_guard.c, where v2_echo_request bumps it twice and now_act_v2_note bumps it twice per stage — and both are reached only through now_act_v2_begin. Generation 0 therefore means now_act_v2_begin returned early on every act of the drive, and it has only three early exits:

  • now_act_plane_state(table) != kNowActPlaneReady (length, caps or act_format),
  • cell->status != kNowPeekActStatusPending,
  • cell->target_a5 != current_a5 — by design, since only the target process's own pump may proceed.

Arming is not a candidate: the same log line that opened this arc shows the plane armed and active. Measurement below settles which of the three it is — the second, cell->status never pending, because no act ever arrived.

RESOLVED the same day: generation 0 means idle, not dead. Reproduced on an emulated Power Mac G4 (guest 16d99316ff6b, resident 67d5ef43 — the same resident as the drive), with all four planes reading requested=15 active=15 and every plane active-current:

interaction=active-current/gen0     <- before any act
interaction=active-current/gen6     <- after ONE console `actselftest`

Six is exactly one echo plus two stage notes at two bumps each, and the act itself failed. So resident_generation advances as soon as an act reaches now_act_v2_begin, and generation 0 is a truthful "no act has ever reached the resident" — not a broken plane. P4's publish path is intact.

That moves the open question upstream, off the resident entirely: in the 2026-08-04 drive, acts were attempted and settled unknown / dispatched-but-unconfirmed while the resident's counter never moved, so those acts never got as far as a pending cell the resident could see. The next investigation belongs in the host's act dispatch and the guest's now_act_submit, not in arming, plane policy, or ext/.

Two rig traps found on the way, both of which mimic a dead plane.

  1. /private/tmp/now-u7-extension-only/session.qcow2 carries the NOW Extension in its file system, but its internal snapshots predate it. Resuming --loadvm runner-ready (a 2026-07-19 state) gives a guest reporting lifecycle=absent, cap=-. Cold-boot it and the resident is active. (An earlier draft of this entry claimed a cold boot still reported absent; that was wrong — it was reading stale lines from an append-only log that BOTH sessions' hosts write to. Mark the log length before an experiment and read only past the mark. NOWBASE actmeta lines carry guest_build=, which is the only way to tell whose guest a line describes.)
  2. A guest binary not named New Old World cannot arm any plane at all. peek.c :: current_app_identity requires creator NOWo and the exact process name, and maintain_writer returns 0 without it — "dev-named app: read-only NWex" — so publish_claims never writes arm_request. Same build, same resident, same host: renamed from now-guest-ppc to New Old World, requested went 0 → 15. The app looks entirely healthy meanwhile, and the host still reports lifecycle=active cap=15, so the Mirror simply shows nothing and every act refuses. AGENTS.md records this name as a preferences rule; its sharper consequence is that a dev-named build has no Mirror at all.

The instrument that settled all of this is now permanent: the actmeta line carries each plane's own state and generation beside the masks.

CYCLE 27 RETAINED-STATE CHECKPOINT; TWO ADVANCES, TWO BLOCKERS (2026-08-04)

The exact d0a3e1a host was driven through native Mirror mouse input and compared with the explicitly identified QMP framebuffer oracle. Before that drive, scripts/test-all passed 103 native tests, the PowerPC, 68K, and NOW Extension cross-builds with their real Retro68 toolchains, and the full host gate. This is tested and emulator-observed, not metal-verified and not a green Mirror sweep.

Two earlier red cases advanced. Macintosh HD now opens with its item roster in the first settled Finder scene instead of remaining blank until an unrelated action. Key Caps now launches through a typed guest-Finder operation and comes frontmost. Workshop resize and close also continued to mutate both surfaces, and Workshop's structured content no longer disappeared during the observed poll sequence.

Two blocking families remain. Hide Finder timed out without changing the authoritative guest and produced no visibility-action line in the host log; the retained visibility census correctly refused to confirm it, but dispatch and observability are still broken. Key Caps is a successful launch with a completely empty Mirror body: QMP shows the full keyboard while Mirror shows a hatched unavailable region. That application has no standard controls, so its draw-owned content needs an explicit structured placeholder until deferred pixel transport is undertaken. Finder fidelity also remains partial, and Sherlock was not re-driven in this continuation.

The apparent Apple-row count changed with the front application: NOW-front began at AirPort, while Finder-front additionally exposed About This Computer. Do not treat that contextual difference as destructive row loss without an authoritative same-context comparison. Strict C27 rows remain blocked because the captures do not include the complete correlated operation/settlement/host-log/guest-log manifest. Exact identities, evidence paths, and the bounded verdict are in docs/mirror-retained-planes-checkpoint-2026-08-04.md.

CYCLE 26 HIGH-WATER CHECKPOINT; APPLE REPAIRED, STATE MODEL OPEN (2026-08-04)

The exact C26 native host was driven through its Mirror and compared with QMP guest captures. A later scene's empty Apple shell no longer erases complete same-guest rows: the host retains only previously observed guest rows, keeps the newest identity/geometry, marks the projection expected-stale, and never invents an initial menu. The regression guard was watched fail under mutation and pass after restoration. The full native/host gate passes; the cross-guest build step skipped because Retro68 was unavailable on this shell path. This is tested and emulator-observed, not metal-verified.

The direct sweep also fixes the verdict on several earlier reports. Macintosh HD opens from Mirror input and Finder contents render. Date & Time opens, Set Time Zone appears within 25 seconds, and Cancel works. Their fidelity remains red: Date & Time lacks authoritative field values and explanatory text, and the modal's guest city/country rows are blank in Mirror. Workshop structure renders but its authoritative detail content is blank after the content ring reports earlier bytes overwritten. Hide Finder briefly removes then restores the window without changing the front application. Windows > Workshop is refused because the guest never calls MenuSelect, so the window does not reopen.

C26 is paused rather than falsely scored green: the strict evidence manifests were not complete for every row. The durable build identities, direct sweep, act-log evidence, paired images, and resume instructions are in docs/mirror-high-water-checkpoint-2026-08-04.md. This checkpoint is the floor for the host state-engine plan; no later implementation may trade these passes for progress elsewhere.

STATE ENGINE U7 READ PARITY BUILT; LIVE STAGING STILL OPEN (2026-08-04)

The native Mirror and MCP now have one state owner rather than parallel observers. Local protocol v9 exposes status, snapshot, find, and wait as four read-only projections over the existing session-pinned engine. They carry the same snapshot ID, digest, stable process/window identities, freshness, coverage, and generation counters the native renderer/evidence path reads. Find is locally bounded and wait observes publication without polling the guest or creating another cache.

Seven focused service/projection/codec tests pass. The derived MCP coverage gate was watched failing because all four newly registered tools were absent from docs/mcp-coverage.md, then passed after the rows were documented. This is tested, not yet live-parity verified: the running host predates protocol v9 and the development VM still has a stale guest application despite the current NOW Extension.

Later 2026-08-04 runtime correction: protocol v9 was exercised through the exact newly built Host and now_mirror_status returned the live engine's guest, session, snapshot ID, sequence, digest, completeness, and both generations. That proves the socket/read projection is live, but not whole-surface parity. The first direct U8 preflight remained red: the cold Mirror acquired only a desktop shell until the first Apple-menu click; that click recovered six windows, but Finder and Date & Time content was blank/placeholder, the Apple menu had no rows, and both Windows > Workshop and Application-menu Finder selection refused because the displayed compatibility entities lacked stable guest identities. The exact current extension file was then staged, but the current guest application could not overwrite its running predecessor (create err -48). The scoped Worker refuses both targeted quit and scripted shutdown, so the image is explicitly partially staged and not clean-saved until the visible app is quit and the full app/extension pair is cold-booted.

Two items remain explicitly deferred. The old now_observe_elements call cannot be removed until the state engine owns structured control elements and their capability references; removing it now would make the existing act rows unaddressable. MCP mutation parity also remains off until a direct native equivalent is proven against the exact staged guest. Neither API reach nor MCP success can satisfy a direct-input/pixel gate.

STRUCTURED LIST CONTENT PRESERVED; BROAD PANEL FIDELITY STILL RED (2026-08-04)

The Date & Time city/country failure was not an absent guest capability. The NOW Extension already returned every List Manager cell as bounded row, column, text, and selection records, but the application bridge collapsed the answer to selected text before scene publication. The renderer then compounded the loss by giving an unclassified DITL resource-control shell precedence over the same live control after P2 had classified it as a list. The scene contract now retains bounded listCells plus the guest's total count, marks the payload complete only when every reported record is valid and present, and lets a classified semantic control supersede only its matching unknown resource shell. Focused native, IR decode/freeze, and renderer tests pass, and both PPC guests plus the extension cross-build with the real Retro68 toolchains.

This is built and tested, not emulator-verified. The running VM disconnected before the rebuilt extension could be cold-loaded, so no direct Mirror drive, authoritative guest capture, or pixel comparison has yet proven the city and country rows. Set Time Zone is currently known to be a titled Dialog Manager window (kind == 2) with DITL items and stacking; the scene does not prove whether the application is inside ModalDialog, so the host must not invent a stronger alert/modal-loop claim.

The wider application/control-panel content problem remains red. P3 deliberately refuses to replace an application's existing custom QuickDraw grafProcs; that is a safety boundary, not evidence that a blank interior is acceptable. The next cold-boot sweep must collect multiple application and control-panel surfaces in one pass, separating structured P2/DITL content, settled P3 replay, and explicit unavailable placeholders before another shared producer/renderer patch. A guest-side action performed outside the Mirror while the connection was failing also demonstrated the expected-stale case: retained same-session state is useful for continuity but must not be driven or scored as current.

The same session ended with the guest showing the system bomb dialog “Finder” error type 10 after an earlier alert was dismissed, before reaching a black powered-down framebuffer. The host log proves only that its Special > Shut Down menu action was dispatched and the guest disconnected one second later; it does not attribute the Finder crash to that action, P3, or the unbooted list patch. Treat this as a blocking crash sentinel for the next cold boot: record the resident extension identity, arm content conservatively, and stop the sweep if Finder faults again.

Later 2026-08-04 live correction: the rebuilt extension and application were cold-loaded, and a coherent QEMU memory sample reached the exact Set Time Zone list control. The extension's NWpt semantic cell completed that request as UnsupportedCustom; it did not return the city/country cells. The earlier bridge-collapse diagnosis came from unit fixtures and is valid coverage for a standard list, but it was not the live cause of this panel's blank rows. The current guest classifier rejects any list box with a nonzero LDEF before asking for its ListHandle. The generic, metal-compatible next step is to prove the public List Manager backing record and widen the extension producer under validated invariants, not special-case Date & Time or introduce pixels. The read-only dev method and exact observed records are in docs/qemu-memory-oracle.md.

The same run broadened the next batch before another extension rebuild. Date & Time's base window has 20 DITL items and 21 controls but loses multiple structured values and status strings. Sherlock has 35 live controls spanning standard, edit, list-like, and application-defined definitions while its Mirror is mostly structural shells. Key Caps has two windows and no controls at all, so its missing keyboard is draw-owned and belongs to the explicit placeholder/deferred-pixels path rather than the control-semantic patch. The extension work remains open until the Date & Time and Sherlock control classes are inventoried and patched together.

Later 2026-08-04 implementation checkpoint: P2 format v2 now classifies Apple-owned controls through public kControlKindTag, reads bounded clock/text values through public data tags, permits a public list-box ListHandle with a nonzero drawing LDEF, offers every live control, and retains 64 compact class facts plus four bounded list payloads. Sherlock's 35-control census therefore cannot restart after eight entries, and typed-but-undecoded data browsers/user panes/image wells produce explicit bounded placeholders. Key Caps remains the separate zero-control, draw-owned placeholder case. Focused native/renderer tests pass, the new placeholder guard was watched fail under mutation, and the PPC, 68K, and flat-INIT Retro68 cross-builds pass. This is built and tested, not emulator-verified until the new INIT is cold-loaded and the broad direct Mirror sweep is repeated.

Cold-load result: the broad direct-input sweep is still red. Set Time Zone and Sherlock retain bitmap-unavailable regions; the latter's newest settled P3 generation contains only its final CopyBits blit, so renderer ordering alone cannot recover its structured controls. A paired state-engine/QEMU sample showed Date & Time as the fresh front process while the live P2 request kept naming one Finder control and the last completed response remained an older UnsupportedCustom Date & Time base-control answer. The sample did not reach the exact Set Time Zone list, so it neither verifies nor refutes the public list-kind hypothesis and does not alone justify a scheduler rewrite. The renderer now keeps unavailable CopyBits geometry behind structured ops and lets a typed incomplete list suppress only its unknown DITL shell; both tests were mutation-watched, but neither row is green without current P2 data.

A generic follow-up is now built and staged: a custom-signature control may prove that it is List Manager-backed by successfully returning the public kControlListBoxListHandleTag with the exact handle size. Apple-owned non-lists are not probed, and a declined/malformed custom result stays explicitly unsupported or invalid. The native semantic slice and real 68K guest/flat-INIT cross-build pass; deleting the fallback fails its source guard. The running VM still contains the prior resident code until a clean cold reboot, so no UX claim has changed yet.

STATE ENGINE U1A: TYPED COVERAGE AND LIFETIME IDENTITY BUILT (2026-08-04)

The IR v2 producer and MirrorKit consumer now agree on typed collection coverage and durable process/window incarnation fields. Process census, per-process window membership, and front-menubar coverage reach the wire as complete, partial, retracted, failed, stale, or unavailable rather than requiring a reducer to parse English diagnostics. The normative rule is now explicit: only fresh complete parent coverage may prove deletion; weaker coverage retains compatible state expected-stale and inert.

Native scene_json_test, MirrorKit IR-freeze tests, and host scene decode tests pass. The test was first observed failing on both sides before the contract was implemented. The cross-guest build gate skipped because Retro68 was unavailable on this shell path, so this checkpoint is tested, not guest-built and not emulator-verified. U1 remains open for exact Finder-item and application visibility capability identities plus forced collector-exit coverage guards; those must not be replaced with title/name matching merely to advance U2.

STATE ENGINE U2A: PURE RECONCILIATION AND SETTLEMENT BUILT (2026-08-04)

MirrorKit now has session/process/window identities, an ordered scene observation, a pure replica reducer, immutable projection metadata and semantic digest, tombstones, and a separate pure operation reducer. Incomplete absence retains compatible entities expected-stale, non-frontmost, and inert; only an exact complete parent scope deletes. A complete process census plus complete window membership for every process is the base-acquisition barrier. IR v1 is still displayable but cannot enter durable maps or authorize deletion/action.

Fourteen focused reducer tests pass. A deliberate mutation that treated any process coverage claim as complete made the background-retention test fail by deleting New Old World, then the correct predicate was restored. The broader MirrorKit suite is not green: seven historical fixtures compare current-v2 builder output with v1-stamped golden scenes. That pre-existing version expectation is recorded rather than hidden or repaired inside this state slice.

This is still additive pure state, not visible Mirror behavior. Host shadow plumbing, full producer coverage, content generations, bounded history, capability-safe Finder/app actions, direct-input pixels, guest build, and VM staging remain open. The implementation contract and limits are in docs/mirror-state-engine.md.

STATE ENGINE U3A: SESSION-PINNED SHADOW ENGINE WIRED (2026-08-04)

The host now keeps one shadow engine per exact GuestKey connection session, publishes accepted projections into a bounded 32-snapshot/15-minute history, and records bounded semantic differences against the still-visible legacy scene. NOWMirrorSource pins the session it started on and sends structural polls to that exact socket even after the active picker changes. Responses from another session are ignored, and a second scene caller is refused instead of silently replacing the first completion.

The old action and content paths remain active-session-only. This checkpoint therefore pauses them and visibly refuses gestures whenever the selected Mac is not the Mirror's pinned Mac; shadow state never authorizes mutation. A delayed stop callback also cannot clear a newer Mirror binding. This is an honest safety boundary, not the final addressed operation broker.

Five focused engine tests, the addressed two-guest polling test, and all 21 existing NOWMirrorSourceTests pass. The live C26 Mirror/VM were deliberately not touched during this plumbing checkpoint. Full direct mouse/keyboard preflight, shadow parity across Workshop/Finder/Date & Time, structured-content and Finder enrichment reduction, native read cutover, guest build, and VM staging remain open. The visible product is still the C26 legacy projection and must not be described as state-engine-driven yet.

STATE ENGINE U4A: ASYNC RENDER ENRICHMENT AND FRAME EXPORT BUILT (2026-08-04)

Settled QuickDraw content, cached Finder window items, and desktop items now converge through the shadow engine after their exact structural sequence. The pure enrichment path changes render-bearing fields only, requires the same process/window incarnations and geometry, ignores stale sequences, and does not publish semantic no-ops. Engine snapshots now expose stable structural and content generations independently.

The native Mirror window also has an app-owned evidence export that pairs its PNG with the full decoded engine scene, exact snapshot identity, guest/session, sequence, digest, base completeness, and both generations. It refuses if the visible legacy scene differs from shadow state or if the snapshot changes while AppKit captures the frame. This is only the Mirror/state member of the strict gate; it does not replace direct keyboard/mouse provenance, authoritative guest capture, operation/settlement, or logs.

Ten focused engine/export tests plus nine content-plane and 21 Mirror-source tests pass. The stale-enrichment guard was watched fail under mutation, then restored. Guest-authored content epoch/generation metadata, typed Finder/content coverage, the QMP-only oracle target split, live direct parity, and visible read cutover remain open, so this is a tested shadow checkpoint, not emulator-verified product behavior.

STATE ENGINE U4B: EXECUTABLE QMP CODE IS ORACLE-ONLY (2026-08-04)

The QMP socket client, QMP-capable action dispatcher, and legacy live polling controller now live in a separate MirrorOracleKit SwiftPM product. MirrorApp opts into that development target; production NOW Host still links only MirrorKit and MirrorKitUI. The Host app builds without the oracle product, its binary contains none of the QMP handshake/dispatcher markers, two host tests guard the manifest/source boundary, and the standalone legacy MirrorApp still builds.

This is not yet the whole U4 cleanup: historical QMP-named action cases, availability fields, and MirrorTarget.qmp remain in core data types even though their executable behavior no longer links into NOW Host. Those types must become platform-neutral or oracle-owned before U4 is complete. No live VM or Mirror was touched, and this checkpoint changes no visible projection.

STATE ENGINE U4C: PRODUCTION ACTION MODEL IS ORACLE-NEUTRAL (2026-08-04)

The remaining development-oracle vocabulary has been removed from production state and action types. MirrorTarget now identifies only the guest wire and machine. Positioned press, double-click, drag, menu tracking, and thumb tracking are platform-neutral device actions behind ActionPlanes.inputDevice; the optional QMP socket and its availability decision live in MirrorOracleKit. LiveMirrorController exposes the adapter's actual planes, so a development launch without a socket no longer advertises device tracking that its dispatcher will refuse.

Sixty-one focused Mirror action/hit/Finder/scroll tests and 24 NOW source and oracle-boundary tests pass. Both MirrorApp and production Host build, and the Host binary contains none of the QMP client markers. The new model-source guard was mutation-watched: adding a QMP sentinel to ActionModel.swift failed the named test, and restoring the boundary passed it.

This is a tested architecture checkpoint, not visible or emulator-verified progress. U4 still needs producer-owned content epochs, strict full-manifest gate tooling, and live shadow parity; U5 read cutover, FIFO mutation, the direct native input/pixel campaign, guest build, and staged VM remain open. The live C26 Mirror and VM were deliberately left untouched.

STATE ENGINE U4D: ORACLE CAPTURES ARE EXPLICIT AND IDENTITY-BOUND (2026-08-04)

tools/shot no longer guesses the newest run/*/qmp.sock. Each framebuffer capture requires an explicit socket and a versioned oracle-identity artifact that names the guest, exact connection session, guest build, QEMU VM name, and socket. The helper verifies QMP query-name before screendump and emits a capture sidecar with the same identity and capture timestamp.

The UX evidence gate requires that sidecar, joins its guest/session to the engine state artifact, and rejects wrong build, VM, socket, frame, timestamp, or missing sidecar. QMP remains observation-only; none of this supplies input provenance. Twenty-eight scored-evidence tests and five shot-helper tests pass. The socket-discovery guard was mutation-watched: restoring an implicit find path failed the named test, then the explicit refusal was restored.

This is a tested tooling checkpoint, not a direct sweep. The operator still has to create the identity artifact from the pinned live session, and a scored row still needs native Mirror keyboard/mouse input, Mirror pixels, state, operation settlement, both logs, stable generations, and human visual review. The retained VM and Mirror were not touched.

STATE ENGINE U4E: OVERWRITTEN CONTENT RETAINS THE LAST SETTLED DISPLAY (2026-08-04)

The content plane no longer clears a window's settled display when the guest draw ring resyncs or reports overwritten bytes. It discards only the incomplete accumulator and records the current guest displayEpoch/generation as a replacement floor. Further records from that damaged generation are ignored; only a strictly newer guest-authored epoch or generation may build a replacement, and bounded more pages still cannot publish half a repaint.

All nine content-plane tests pass. The guard was mutation-watched by restoring the old settledOperations.removeAll() transition: the retained-display test failed at the resync and same-generation assertions, then passed after the clear was removed again. This is the code-path repair for the C26 Workshop blanking report, but it is only tested, not yet directly re-driven in the native Mirror or compared with the guest. That row remains red until the full sanity preflight and Workshop visual comparison are run on the new build.

STATE ENGINE U4F: INACTIVE WINDOW CONTENT IS RETAINED, RUNTIME RECHECK OPEN (2026-08-04)

The first direct-input sweep of the U4E build reproduced a second destructive transition that its ring-overwrite test did not cover. Selecting New Old World left the Finder window visible but produced a frontless structural observation; NOWMirrorContentPlane.join treated that bounded absence as a deletion and cleared every settled display. Retargeting another front window also discarded the accumulator identity used to find the last settled display. The Mirror therefore oscillated between the Finder's 106 structured draw operations and a five-operation Bitmap unavailable shell.

The content plane now records the last published identity separately from the in-progress identity for each exact guest process/window slot. Frontless and retarget observations retain compatible published content as expected-stale; partial replacement pages keep drawing the settled display until the newer guest epoch/generation finishes. Retention remains session-bounded and guestChanged still clears it atomically.

Ten content-plane tests and all 21 Mirror-source tests pass. A mutation that cleared the published identity on a frontless observation made testFrontlessObservationRetainsInactiveWindowDisplay fail with a missing display, then the correct transition was restored. This checkpoint remains tested, not emulator-verified until the rebuilt host is directly driven through the complete preflight.

That sweep also recorded action reds rather than hiding them: the Apple menu and native Application menu rows rendered correctly; selecting New Old World correctly activated only the application because its Workshop window was closed, while Windows > Workshop failed to reopen it; clicking the inactive Finder window was refused because the running guest accepted no select window act; Hide Finder and Date & Time open were dispatched but did not visibly settle. The source tree already carries winact select, so the guest/build mismatch must be resolved by staging and proving the exact latest extension and guest before treating those runtime results as current implementation failures.

STATE ENGINE U4G: ALL FOUR PLANES ARE RETAINED; LIVE REPROJECTION OPEN (2026-08-04)

Plane policy is no longer a destructive filter at the renderer adapters. The session-pinned engine retains P1 structure in its replica, P2 semantics per exact process/window/control or dialog-item identity, P3 content per exact window identity and geometry, and P4 operation history in the same engine's bounded journal. P1 cannot be disabled. Hiding P2 or P3 recomposes the current snapshot without those contributions; showing either immediately restores the cached contribution at the same guest sequence. Disabling P4 still gates mutation, but policy changes cannot erase previously recorded attempts or settlements. Evidence arriving while an optional plane is hidden is retained for the next composition rather than discarded.

Cross-generation QuickDraw retention now belongs to the state owner rather than NOWMirrorContentPlane. The adapter publishes only the newest settled guest generation. If that generation is bitmap-only, the engine keeps compatible prior non-bitmap operations—including their QuickDraw state records—as expected-stale structured evidence. This addresses the observed Sherlock transition where structured controls appeared briefly and were then replaced by a CopyBits-unavailable overlay; it does not make bitmap pixels part of the product path or turn Sherlock green by itself.

The focused state-engine, plane-domain, content-plane, and source suites pass. Two guards were mutation-watched: dropping retained state operations fails the structured-content test, and failing to cancel an old policy-refresh sleeper fails the single-poll-cadence test. This remains tested, not emulator-verified until the newly built host is directly driven through the full sanity preflight, its live toggles are exercised, and every assessment is paired with the authoritative guest capture.

The shipping review caught two lifecycle races before this checkpoint. A pre-close scene completion could be accepted by a same-guest restart, and a policy toggle could begin another structural request after scene transfer but before that scene's content command settled. One run generation now invalidates late callbacks, one cycle token covers scene plus content, and toggles made in flight coalesce into one immediate follow-up. Policy lookup is also keyed by the Mirror's pinned GuestKey, so changing the selected Mac cannot reproject another session. The source tests now hold real scene/content completions and count requests instead of inspecting source strings. Both lifecycle guards were mutation-watched; the four focused suites pass 62 tests.

The first direct run of ccf68a0 is recorded in docs/mirror-retained-planes-checkpoint-2026-08-04.md. P2 and P3 live reprojection behaved as designed: P3 could be hidden while P1/P2 remained, P2 could be hidden while retained QuickDraw content remained, and restoring P2 immediately restored the same semantic generation. The complete sanity preflight is still red. Finder and Control Panels content arrived only after a later polling cycle; Hide Finder remained unconfirmed; Date & Time's Finder item was absent and therefore not actionable from the Mirror; Key Caps did not launch; and Sherlock's structured content was again overwritten by a later bitmap-only/invert observation. The last result is direct evidence that the production renderer path still bypasses or obscures the engine retention that the focused guard proves. Status is now host-tested and partially emulator-observed, not a green sweep and not metal-verified.

The same sweep exposed a separate outcome-classification defect. Several actions whose effects later appeared in authoritative pixels or scenes kept an immediate act-refused, outcome-unknown, or dispatched-but-unconfirmed label. A transport or resident-act answer is attempt evidence, not the terminal verdict for the person's composite operation. A refusal before any dispatch remains terminal; once any part of an operation may have reached the guest, the operation stays non-green and eligible for later same-session postcondition evidence. A later complete authoritative observation may confirm that same operation; it must not erase the earlier contradictory evidence or be attributed across another queued operation.

STATE ENGINE U6A: DIRECT OPERATIONS ARE SERIALIZED; CURRENT VM LACKS LIVE IDENTITY PROOF (2026-08-04)

The first direct-mutation state-engine slice is built and focused-tested. One session-pinned FIFO now owns modeled native gestures and does not dispatch the next gesture until the active one confirms or times out. Its operation journal keeps displayed snapshot/sequence, exact entity, typed postcondition, attempt reply, and later authoritative outcome separate. A post-dispatch act-refused is therefore contradictory attempt evidence rather than a false terminal verdict. Late complete scene evidence may confirm it; timeout remains non-green and can later confirm only while no active retry makes attribution ambiguous.

The direct run also exposed an unsafe bridge that is now closed. The running VM can supply a compatibility scene that is drawable while the new replica cannot mint stable identities for its windows/processes. The initial U6 code quietly fell back to the legacy dispatcher in that case, recreating the raw act-refused labels the broker was meant to replace. Modeled plans now refuse before dispatch when identity is unavailable. The VM must be updated with the current guest and extension before positive broker settlement can be scored.

After correcting two missed Computer Use targets, the live results were:

  • Finder was selected first; the exposed System Folder title bar resolved to that exact second Finder window, but the stale guest refused winact select and the window did not come front.
  • Selecting New Old World through the guest Application menu activated NOW and did not reopen its closed Workshop, which is correct.
  • Windows > Workshop is the independent failing operation. It must create a named NOW window; the current guest does not call MenuSelect, so it remains red rather than borrowing success from application activation.

The focused gate currently covers FIFO serialization, contradictory refusal then confirmation, timeout and late confirmation, same-postcondition retry ambiguity, exact same-app window identity, Workshop named-window creation, and the no-legacy-bypass rule. Mutation-checking removal of the FIFO guard made the two-click test fail. This checkpoint is tested and directly characterized, not a green emulator sweep and not metal-verified.

CONTRACT FROZEN; UNIFICATION IMPLEMENTATION STILL OPEN (2026-08-03)

The unified NOW Extension prerequisite now starts from a source-derived retirement ledger rather than the claim that AXPeek/QDPeek/Portal parity can be remembered. The ordinary native guard derives 66 capability keys from the old resident shared headers, agent dispatch, service, and staging paths. Each key has a goal-facing outcome, allowed disposition, and one prerequisite proof owner. Goal-relevant rows cannot close as a bounded refusal, retained fixture, or retirement blocker. The older fold roadmap is historical; Mirror completion plan 001 is blocked until the prerequisite completes and now retains only the broker, broad renderer/UX, MCP, and later pixel work.

The scored-row gate also now enforces the full correlated evidence boundary. A pass/fail needs native NOW Mirror keyboard/mouse provenance, Mirror pixels and snapshot identity, a QMP observation capture within two seconds, decoded state, plane state, terminal operation settlement, both host and guest logs, no nonterminal operation, two unchanged pre-capture generation polls, and an unchanged post-capture generation/owner-epoch reread. The previous gate was watched accepting manifests without the new plane/settlement/log/quiescence members; the focused test then failed seven cases before the validator changed.

This is contract and guard work only. No resident plane has been unified by this entry, the focused corpus has not passed, and the development VM has not yet been replaced with the final exact extension/application pair. Workshop and menubar geometry remain the regression floor; Apple content, Finder/Date & Time foreign content, direct controls, and truthful settlement remain red until the direct-input paired sweep proves them.

U2 source boundary implemented; resident runtime proof still open

The appended resident ABI now carries deterministic source and embedded build identities, one canonical New Old World writer lease, named per-capability owners, extension echo of the accepted owner epoch, and P1 cadence counters. Content, interaction, scene, and Processes use the same owner union; the Processes renewal is intentionally in its cooperative idle() callback, not its repaint callback. Native layout, lease, resident-core, and guard tests pass, and both guest applications plus the extension cross-build with the real Retro68 toolchains.

That is tested at the native/source boundary and builds, not emulator- verified. The generated INIT payload is about 59 KiB, above the INIT skill's conservative 32 KiB audit budget (the preceding extension was already in this size class and loaded on Mac OS 9). A disposable cold boot must still prove the new exact NWid fingerprint, table identity, writer handoff, callback chaining, and six-tick counters before any runtime row turns green. The stage image still contains the older resident and must not be used as evidence for this build.

U3 ABI checkpoint only; semantic resolvers remain red

The R11 evidence review is now durable in docs/p2-semantic-evidence.md. It rejects every fact the bounded P1 reader already proves, rejects non-dialog TextEdit from v1 because no documented safe root exists, and fixes a 32-record envelope large enough for the measured 16-row Finder Apple menu. The table has an appended exact-target P2 cell, explicit refusal/truncation states, and a volatile generation-checked application copy. Mutation fixtures cover short tables, stale/wrong identity, partial publication, overflow, resolver-kind mismatch, and dishonest text completeness.

This is an incomplete checkpoint, not P2 behavior. NOW Extension does not advertise or arm P2 yet. Invalid-handle-safe Control/Resource/List/Menu Manager resolvers, request publication, scene joining, disabled-P2 degradation, and the direct Date & Time/Apple UX proof all remain red. The ABI's ability to represent those facts is not evidence that the guest has produced them.

Update 2026-08-03: bounded partial P2 behavior is implemented, but the UX rows remain red. NOW Extension now advertises the plane and serves exact standard List Manager state plus exact nonempty menu rows. The application publishes at most one request per scene, joins only an owner/scene/object match, retains bounded terminal facts across scenes, and resets them on owner change or scene regression. Mirror renders a standard list as an explicitly partial, selected-value-only list surface; it does not invent the unretained rows. Native mutation fixtures and the real PPC, 68K, and flat-INIT cross-builds cover this slice.

The Mac OS 9 system-root Apple submenu is still broken. The flat 68K INIT cannot link the CarbonLib/CFM root-menu calls that expose the child behind the empty Apple shell, and no undocumented trap or unproven Mixed Mode bridge was substituted. That exact shell therefore returns unsupported. Direct Mirror input against Date & Time and the Apple menu has not yet proved this partial plane in the running guest, so neither case is green. The development stage image also still contains the preceding resident build.

U4 P3 lifecycle and coherent-redraw path built; runtime proof still red

P3 format v2 now arms one exact A5/window identity instead of every window in a process. Every retained record carries the echoed PSN, exact A5 and port, request generation, and resident display epoch. A retarget clears the old live commit before rewriting any identity word; same-context target changes restore only a still-live hook that NOW still owns, while dead or foreign-context rows are forgotten by value and remain strict pass-through. The host joins only the currently armed front window and rejects old process, window, generation, A5, or display-epoch records, so pointer reuse cannot overlay a relaunched target.

After installing its exact hook, the resident requests one ordinary update and does not call the application's draw handler itself. InvalWindowRect cannot link into a flat 68K INIT, so this path uses the classic equivalent inside the resident only: prove exact WindowList membership and hook ownership, save the current port, set the target window port, call InvalRect, and restore the saved port. It reports redrawRequested only after that sequence and redrawServiced only after a later guest QuickDraw hook. A source-order guard forbids BeginUpdate, EndUpdate, direct drawing, CopyBits, or event injection in the shim. The Carbon UI lexical audit flags InvalRect; that finding is expected for this flat-INIT compatibility boundary, not waived for Carbon application source.

Native lifecycle/ring tests, the host/MirrorKit join tests, and the actual PPC application and flat-68K extension cross-builds pass. This is tested and builds, not emulator-verified. No guest has yet proved that a foreign Date & Time window services the requested update, that target death/relaunch remains safe at runtime, or that the resulting initial display is coherent. Bitmap and CopyBits operations still carry only bounded geometry and render as explicit Bitmap unavailable placeholders; no pixel transport was added. The stage image still contains the preceding extension.

The U4 cross-build produced a 63,978-byte INIT 128 (64,386-byte resource fork), so the resident still exceeds the conservative 32 KiB inspection budget and runtime loading remains a required gate. Its artifact SHA-256 is 1b5ad7638477d974e71d6852d61ff428ea797d1690a3f4f1dd2b7f72264f9e11; the embedded source manifest is ffe26237 08404b43 c9ae73ce c44e25aa bb1c11a3 and build fingerprint is 707560f7 d034f202 a01a4a83 faa2c1c4 576c0239.

U5 P4 settlement is effect-owned; focused runtime proof remains red

P4 format v2 is appended after the original act cell, so the V1 act bytes and every P1-P3 offset remain fixed. A request now carries one correlation, canonical-writer epoch, exact A5 and PSN echo, source scene generation, typed operation/object identity, and deadline. The resident may advance only its monotonic requested/accepted/armed/fired/refused/expired evidence for that tuple. PSN is correlation rather than resident authority: the foreign-context safety boundary remains the exact A5 plus the operation-specific object guard.

That evidence is deliberately not the outcome. New Old World owns a bounded 16-record settlement history and reconciles a fired action against a later normal-context scene and an operation-specific postcondition. Timeout remains recorded if a later scene confirms the effect; writer replacement terminates open records as session-changed. The host joins by correlation and renders a checkmark only for confirmed. Refused, timed-out, session-changed, unknown, and dispatched-but-unconfirmed remain visibly non-green. Successful keyboard, typed-text, and Finder dispatches without a postcondition are explicitly unconfirmed; activation requires the guest's own front-process reread, and application visibility requires a Finder visibility reread. Menu acts require the unique front PSN from the same scene rather than guessing the current app.

Native guard/settlement tests pass 100/100. The final host tree passes 1,289 tests with 54 opt-in skips; focused settlement presentation passes 17/17. The real PPC, NOW-68K, and flat-INIT cross-builds pass. The final U5 extension is a 64,994-byte INIT 128 in a 65,402-byte resource fork (65,536-byte MacBinary), SHA-256 76439badc2ef9499502592c4a3b533e657a768a6c7d89772a76e9bc36758fa7c. Its source manifest is 0fc7296f 1fd60fa0 4e7e5b68 b4a60944 a6174813 and embedded build fingerprint is b1e5890e 8cba499f 25c092c9 45d9e7c1 f54951b3.

This is tested and builds, not emulator-verified. No direct-input sweep has yet proved a menu, standard list, Date & Time Cancel, application visibility, or window operation against its paired guest pixels and settlement. Operation families without a stated postcondition honestly remain dispatched-but- unconfirmed until that focused proof supplies one. The development stage image still contains an older application/extension pair and is not evidence for U5.

U6 one-extension lifecycle and plane policy built; runtime proof remains red

The wire contract is revision 2 and the PowerPC mirror command now reports a schema-1 snapshot of only NOW Extension: exact lifecycle and build identity, resident capability/request/active bits, heartbeat freshness, and one row for each of Structure, Semantics, Content, and Interaction. The guest Console and read-only Workshop use the same probe. The host decodes that object into one plane domain, persists only optional-plane policy for an anchored machine, keeps unanchored emulator policy session-local, and presents one native Open/Close Mirror surface. No active UI asks for AXPeek, QDPeek, Portal, mirror-agent, forwarded port 1420, QMP, or an external Mirror binary.

Policy now reaches named claims rather than stopping at toggles. Scene requests carry the Semantics choice; Content off sends the bounded stop and retains an explicit refusal if release fails; Interaction off refuses before dispatch and logs the refusal. Structure is always required. Unsupported, enabled but inactive, requested, refused, degraded, stale and active are distinct; an actual resident-requested row degrades after five seconds without activation, while a closed Mirror's legitimately inactive planes do not start that timer. P1/P2/P4 claims stop renewing on close and expire through their ten-second resident lease; P3 is released explicitly because its lease is much longer.

Focused native JSON/layout/lease tests and host domain/content/contract tests pass, and the PowerPC guest, NOW-68K guest, and flat 68K extension cross-build with the real Retro68 toolchains. One complete scripts/test-all run passed, but two immediate primary-agent reruns exposed an unrelated nondeterministic host-gate red: local guest-listener tests timed out waiting for loopback connections in different cases (6 failures, then 8), while every named failure passed when filtered and rerun alone. A stale C26 host process that had held port 5250 since 20:25 was terminated before the second aggregate rerun, so that process was real contention but does not explain the remaining suite-level instability. Treat the focused U6 behavior as tested and building and the aggregate host gate as unresolved red until a clean rerun is repeatable. This is not emulator-verified. The revision-2 app/extension pair is not installed in the development image, no direct keyboard/mouse Mirror sweep has compared its pixels and mutations with a paired guest capture, and no metal claim is made. U7 still owns runtime/staging retirement and the cleanly shut-down updated VM; the old compatibility implementation remains internal seed material until that slice, guarded from the active product UI.

CYCLE 25 RED; SETUP CORRECTED BEFORE C26 (2026-08-03)

Cycle 25 directly re-ran the sanity preflight through the uniquely identified C25 native host. Workshop resize and close both mutated the guest; their act log rows answered now-window-act-outcome-unknown, while the paired guest frames proved the operations landed. The Macintosh HD double-click likewise dispatched and opened the guest Finder window.

Two real defects survived the sweep. Apple still opened a correctly placed but empty dropdown, so resolving the empty low-memory shell only through GetMenuItemHierarchicalMenu was insufficient. The next patch also checks the installed menu with the same measured ID and the root item's hierarchical ID; it remains red until C26 watches the rows. Separately, a locally synthesised app-only selector still appeared whenever the guest Application menu was absent. It collided with the native menu and necessarily dropped Hide, Hide Others, and Show All. The custom dropdown, hit target, hover state, and action route are now removed. Only guest menu -16489 may open or act; missing menu state is inert rather than replaced with a second control.

The sweep also made a setup error explicit. The extension file was staged, but the running system had never cold-booted it. The first Mirror frame said content-plane-absent, and the live act log later said "the NOW Extension is not installed". The earlier carry-forward assumption that an old resident was loaded was wrong. Therefore the missing foreign Finder and Date & Time windows and menus in C25 are not an implementation verdict: foreign scene collection was being tested without its resident plane.

The corrected local oracle is ~/Lab/Assets/os91-qemu/now-mirror-stage.qcow2. It contains the verified NOW Extension and current New Old World, was stopped by a human-performed guest shutdown (the exact QEMU PID exited on its own), and passed qemu-img check before and after preservation. Its SHA-256 at creation was c466baa9a5455c343908e12197d68e57ffc7f07c140276a90c97a5ae2a137d70. (Re-verified 2026-08-06 with tools/volclean.py: its volume really is cleanly unmounted, which the three images baked that night were not — see the correction below. It is once again the installed oracle.) Future extension changes must update and cleanly shut down this stage image before a Mirror sweep; merely copying a new INIT into a running clone does not change the resident code under test.

That rule went unfollowed for three days, and is now enforced (2026-08-06). On 6 August the image on disk was still the 3 August one while six extension commits had landed that day — the re-armed liveness Time Manager vehicle, its ABI shim, the MacTCP .IPP probe. Every session cloned the stale image, staged a fresh build into a throwaway clone and discarded the clone, which is exactly the mistake the paragraph above names. The rule was written and nothing checked it, so the verified oracle went stale in silence: any sweep started from an old resident with no warning, and a green sweep would have said nothing about the code anybody was working on.

The gate is now three pieces, and the sequence a resident change follows:

  • scripts/bake-ext-image — clone the oracle, stage this checkout's ext and app, cold boot so the INIT loads, ask the guest mirror and require lifecycle active, the expected capability word and a buildFingerprint equal to the local build's, then a guest-clean shutdown (tools/shutdown-guest.py, never a QMP quit), qemu-img check, .bak-YYYYMMDD of the old image, install, receipt. Any failure installs nothing and leaves the VM up. Plus tools/volclean.py since the correction below — the container check was never evidence that the volume inside was unmounted, and three images were installed dirty before anything asked.
  • ext/stage-receipts.json — the repo-tracked receipts. Each bake records the source digest, the guest-reported fingerprint, the image sha256, the qemu-img check result, that the shutdown was guest-clean — and, since the correction below, volumeClean, which is the only one of them that answers whether the volume was actually unmounted.
  • tools/ext-bake-gate, from .githooks/pre-commit — refuses any commit staging ext/ or contract/peek_table.h unless the newest receipt is a bake of exactly that source. TBT_DEFER_EXT_BAKE=1 with TBT_DEFER_EXT_BAKE_REASON="…" defers it and writes the reason into the receipts file, so a deferral lands as a written decision in the same commit as the work.

Two images were installed on 2026-08-06. The hand-run bake that found the staleness produced sha256 62be7be4d73a848f9d72818f42c879df7b4dfdfc83a18f6fdd2779529b297eae and preserved the 3 August one as .bak-20260806; the first run of scripts/bake-ext-image then produced sha256 46a51dcd0337baf3918a1fec6f2987bacbdf174d3dc28ffee61e04430bd5850c, keeping the previous as .bak-20260806-2. That image is the one ext/stage-receipts.json certifies: the guest answered mirror with lifecycle active, capabilities 63 and buildFingerprint bb95520ae51365c053f56d57d86bb10af09629c3, shut itself down in three seconds, and qemu-img check found no errors.

Still stale for one branch, and this is the gate working rather than a defect. claude/012-resident-transport gives the resident its own MacTCP connection and a sixth plane — capability word 127, not 63 — so the image above does not contain that resident. The first commit touching ext/ on a branch carrying it will be refused until scripts/bake-ext-image runs there, which is exactly the warning nobody got on 3–6 August.

CORRECTION (2026-08-06): every image baked that night was DIRTY, and the receipts could not have known

Read the three paragraphs above with this attached. All three images baked on 6 August were installed with their HFS volume still marked mounted, so every clone of them opened with "Your computer did not shut down properly" and a Disk First Aid pass. Measured with tools/volclean.py, by sha256, on the files still on disk:

image sha256 volume
.bak-20260806 — the 3 Aug one, human shutdown c466baa9… CLEAN
hand-run bake 62be7be4… DIRTY
first bake-ext-image (caps 63) 46a51dcd… DIRTY
second bake-ext-image (caps 127) 0785871a… DIRTY

The last two are the images the two receipts in ext/stage-receipts.json certify, matched by their own recorded imageSha256. Both receipts say shutdown: guest-clean and qemuImgCheck: clean.

Both statements are true. Neither is the question. qemu-img check validates the qcow2 container and knows nothing about the Macintosh filesystem inside it, so a volume the Mac will greet with Disk First Aid passes it without a word. And the guest really did run its own shutdown sequence — the Shutdown Manager quits applications first and flushes and unmounts volumes after, so the anchor worker, an application, falls silent while the volume is still being written. The rig then took the process away five seconds later. Measured that night: the disk image kept changing for 49 seconds after the worker went quiet. The sentence above, "shut itself down in three seconds", was reporting how fast the rig stopped watching.

The general lesson is the durable part, and it is bigger than this rig: a check adjacent to the question reads exactly like an answer to it. Two honest, passing checks sat where a third was needed, and nothing in the receipt distinguished them — which is why volumeClean is a separate field rather than a stricter reading of qemuImgCheck.

What changed:

  • tools/volclean.py asks the volume itself: HFS's "volume unmounted" attribute bit, the very flag the Mac's startup check reads. It follows drEmbedSigWord into the embedded volume, because the HFS wrapper every OS 8.1+ HFS+ disk carries has its own flag that reads dirty on a perfectly clean machine — checking that one reported the human-verified 3 August oracle as dirty, which is how the positive control earned its keep.
  • scripts/bake-ext-image runs it and installs nothing when the volume is dirty. The container check stays, demoted to what it always was.
  • The receipt gains volumeClean and shutdownPath; tools/ext-bake-gate requires the first. The two existing receipts are corrected in place with volumeClean: dirty and the reasoning, and the gate refuses a receipt carrying a correction.
  • tools/shutdown-guest.py waits for the disk to stop changing rather than for the worker to go quiet. Necessary, and not sufficient: a shutdown waited out to full disk quiet still left the volume dirty.

The oracle is repaired by restoring, not by re-baking. The dirty image was replaced with .bak-20260806 (sha c466baa9…, verified CLEAN), kept as .bak-20260806-4-dirty. That is the cheap fix and it is the right one, because scripts/spin-up-ppc stages the current ext/ and app into whatever clone it makes and cold-boots so the INIT loads — the resident under test comes from the staging, not from the base image. What a base image has to be is clean and furnished (Rumpus, the anchor worker, the folder layout), and that one is both.

Which means the framing higher up this section — "the oracle went stale" — was itself imprecise, and the note belongs next to the gate: a bake receipt proves that a resident was built, loaded by a cold boot, and identified itself over the wire. That is a real and valuable proof, and it is the one thing no amount of staging gives you. It does not prove that a later sweep runs that resident, because staging replaces it anyway, and it does not prove the base is clean unless volumeClean says so. The gate is not weakened here — Michelle asked for it deliberately, and "someone cold-booted this resident and it answered" is exactly the check that was missing on 3–6 August. It is simply worth writing down that its value is the boot-and-answer, not the image.

FOLLOW-UP (2026-08-06, night): the restore was right and left no trace a tool could read

Everything above is correct and none of it was visible to a program. The restored oracle's sha256 matched no receipt — the file was installed at 01:58 while the newest receipt was written at 01:19 — so tools/image-provenance could only report it as UNACCOUNTED FOR, and an unaccounted oracle reads like a defect rather than like the deliberate, well-reasoned repair it was. Meanwhile the sentence at the top of AGENTS.md ("every Mirror sweep and every scripts/spin-up-ppc clones it") was false and had already caused a coordinating session to tell Michelle that an arc's measurements were suspect; a lane that read the script corrected it.

So a shared mutable oracle goes wrong in both directions: forward when nobody bakes, and backward when somebody sensibly restores and writes it up in prose. What changed (docs/staged-images.md is the page):

  • tools/image-provenance — every guest run prints which base, its sha256, whether a receipt accounts for it, what was staged on top and what the guest said it was running, and writes provenance.{json,md} into the run directory. The .md is the rig table from the 08-07 sweep, generated.
  • scripts/bake-ext-image bakes privately by default into ~/Lab/Assets/os91-qemu/agent-stage/; --shared is the announced act and refuses while other guests are running (five were, when it was tested). A lane needing the boot-and-answer proof no longer has to touch the file everyone clones.
  • tools/ext-bake-gate verify-image runs the shasum this page has asked people to remember. It WARNS on the commit path and refuses only under --require: whether another lane replaced your image is not a property of your commit.
  • tools/ext-bake-gate note-image --reason "…" records a hand-install as provenance only — the verdict says in words that no guest was asked what is inside. The current oracle now carries one.
  • ext/stage-receipts.json conflicts by design at a merge (.gitattributes + tools/receipts-merge-driver), and a deferral that legitimately allowed a commit is refused at the merge into main.

Unverified: scripts/bake-ext-image --shared has not been run since this change — deliberately, because another lane was in ext/. The private path has been exercised only as far as its argument handling and guards; the first real private bake is the at-arm census lane's.

The applet's shutdown does not finish; the Finder's does (2026-08-06)

tools/guest-shutdown's applet calls ShutDwnPower and nothing else. It reliably starts a shutdown. It does not finish one: on mac99 QEMU never exits, and every image preserved after it — including ones waited out to full disk quiet — has the volume still marked mounted.

The Finder's own Special > Shut Down does finish it. Driven through the act plane's menuact, on a guest booted from the stage image: the machine powered off, QEMU exited on its own within 10 s, and volclean.py read the resulting image CLEAN. QEMU exiting by itself is the machine really cutting power, which is the thing the applet path never achieves — and it matches the one image anybody trusted, the 3 August one a human shut down from the Finder.

This route had been recorded as impossible, and why is the more useful finding. menuact requires serialHi/serialLo from the scene's front process, because a menu bar belongs to one exact process and the guest refuses rather than guess at whichever application happens to be front. The probe that wrote the route off omitted them and sent the act fire-and-forget, so the guest's bad-request — a perfectly good refusal, naming its own cause — went into the void, and a null reading was reported as a property of the mechanism. That is drive-loop rule 2e: a null reading needs a positive control before the mechanism is blamed. Supplying the serial and reading the reply was the entire fix; the machine had been answering correctly all along.

docs/mirror-knowledge.md records upstream Mirror finding that ShutDwnStart, ShutDwnPower, Finder Apple Events and embedded OSA all failed to power off a mac99 OS 9 guest, and that posted clicks cannot reproduce the Finder's held MenuSelect gesture — logged there as "deferred, not solved". That is now solved, and by a route upstream did not have: NOW's act plane does not post a click, it answers the application's own MenuSelect, so the held-gesture problem does not arise. tools/shutdown-guest.py --wire <port> takes the Finder route first and keeps the applet as the fallback for a guest with no act plane; tools/guest-shutdown/probe_shutdown.py is the bench that established it.

Still unverified: this was watched once, on one guest, on QEMU. Nobody has run it on the PowerBook, and the applet fallback has no measurement showing it ever produces a clean volume — only that it starts a shutdown.

CYCLE 24 RED BASELINE; FIX BUILT, NOT UX-VERIFIED (2026-08-03)

Cycle 24 was driven through the uniquely identified native C24 Mirror and paired with QMP screendumps used only as the guest oracle. The restored menubar geometry held: Apple, File, View, Windows, Help, the clock, and the right-aligned Application menu matched the authoritative frame. The following remained broken, and none is green merely because the subsequent patch builds:

  • Apple opened a correctly placed but empty dropdown. The live low-memory MenuList supplies the system menu's measured identity and left edge, but its Apple MenuHandle carries no rows under CarbonLib. The patch now retains that identity/geometry and reads the corresponding submenu's rows from AcquireRootMenu; it needs a C25 native drive.
  • Clicking bare desktop did not bring Finder forward. A NOW self-scene carries no Finder desktop-backdrop window, so desktop ownership resolved to nil. The patch falls back to the live Process Manager row with Finder signature 'MACS'; it needs a C25 native drive.
  • Hide New Old World did nothing. Hide, Hide Others, and Show All were generic menu commands even though menu -16489 is system-owned. They now resolve as typed visibility operations, preserve the guest's enabled/disabled state, and use the classic Finder on the guest. Each of the four switcher cases — Hide, Hide Others, Show All, and selecting another app — remains red until directly driven and compared in C25.
  • Clicking NOW Workshop's close box resolved to the correct named window, then refused because the optional resident extension was absent. winact already had a direct self-window implementation, but an unconditional now_act_ready() check made it unreachable. The patch moves only self-window Window Manager acts ahead of that optional-plane gate; foreign windows and self controls still require their real application event path.
  • Finder-front rendering remained a stale, disabled NOW Workshop with content: no front window. Workshop whole-frame fidelity also remains red.

The repository gate passes (90 native tests plus the host Debug/Release gate), and an explicit Retro68 PowerPC Carbon cross-build passes. Mutation checks were watched fail against the pre-patch self-window route and pre-preflight gate. This is tested, not emulator UX-verified and not metal-verified.

Every future cycle now begins with a gate-enforced eight-step sanity preflight: compare Workshop, compare the menubar, inspect Apple's real rows, resize and close Workshop, double-click Macintosh HD, compare Finder, and hide Finder through the native Application menu. Slice-specific work is currently Date & Time. Every row is attempted before patching independent failures in a batch; one red row does not erase coverage of later independent rows.

WATCHED, FIDELITY STILL RED: Workshop reports its manually drawn structure (2026-08-03)

Cycle 19 paired the native Mirror with the authoritative guest and corrected an earlier, too-narrow assessment: proper Control Manager chrome did not make the Workshop render. Its entire 13-row sidebar, page header, explanatory text, status line, and screenshot-page labels were absent. The bounded missing-content placeholder was honest, but a nearly empty window is still red.

The Workshop now describes those manually drawn regions through the same scene IR as controls: panels, placards, selection bands, separators, static text, and bounded icon/picture placeholders carry guest-workshop-model provenance. The host renders that structure from guest-authored data; it does not read or pipe guest pixels. The Screenshots page also reports its current dimensions, depth, streaming state, transport disclosure, and rate text. Actual icon and screenshot art remains deliberately out of scope and visibly placeholder-backed.

Cycle 20 staged that build without rebooting the guest and drove the native Mirror with Computer Use. A same-moment Mirror/QMP pair showed the structure above while the resident content plane still reported absent. This proves the fallback no longer depends on drawing capture: the empty/hatched Workshop regression is fixed on the watched surface.

The whole-frame fidelity row is still red. Mirror showed Depth as numeric 4 while the guest showed the selected popup title 8-bit; it also overlapped the preview placeholder text, and removed disabled controls when Finder was frontmost instead of dimming them. Sidebar icon art remains a named placeholder and is not scored while bitmap/picture work is out of scope. The popup-value cause was in the guest producer: it looked only in the process menu list, while the Appearance popup CDEF owns its MenuRef as control data. The producer now asks kControlPopupButtonMenuHandleTag first and retains the old lookup as a fallback. Cycle 21 watched the corrected Mirror value 8-bit beside the same guest value. Opening or choosing the popup remains red because the old resident act plane was still absent; correct presentation is not evidence of mutation.

The same structural patch fixes the application switcher's asymmetric geometry. The real guest menu title is correctly right-aligned even though Menu Manager reports its nominal left as zero; its dropdown must therefore be anchored to the right screen edge too. Ordinary menu hit spans now exclude that special menu, and a missing Apple menu is synthesized independently. Cycle 20 watched the right switcher open at the right edge and switch Finder/New Old World. It also proved the left Apple glyph was only a drawn fallback: clicking it answered nothing under the pointer. An initial follow-up guessed menu id 128 from NOW's own resource convention. Cycle 21 disproved that guess live: the Apple glyph remained inert. The preserved Aug 1 scene that produced the earlier nearly faithful Mirror records Mac OS 9's system Apple menu as id 256. The self scene now uses that measured system id so drawing, hit-testing, and MenuSelect can share one guest object. It remains red until a later drive watches it open.

Cycle 20 also found that the synthesized switcher lists background-only processes when Finder is frontmost. Choosing one refuses accurately as activate-background-only, but those rows should not be offered by an application switcher. The original switcher predicate already encoded the data-driven distinction available here: frontmost, owns a visible window, or owns the Finder desktop. The later unconditional process fallback caused this regression. It is removed without a host signature allowlist; this remains red until a later native drive watches the resulting roster.

BUILT, NOT UX-VERIFIED: proven control roles survive the scene (2026-08-03)

The guest already derived exact roles for NOW-owned Control Manager controls, including checkbox, radio, popup, group, progress, and disclosure controls. scene_json.c then collapsed every proven role except scroll bars into pushButton. The loss was in the producer, before Mirror rendered anything: Workshop's checkboxes and Depth popup therefore arrived as authoritative rounded buttons, and no renderer could recover the right kind honestly.

The scene now preserves each proven role, its applicable state or value, and only the action that role advertises. Mirror draws those semantic kinds directly; a v2 unknown no longer falls back to a title-and-geometry button guess. CopyBits still carries no pixels by design, but its exact destination now gets an explicit bounded placeholder instead of becoming unexplained white space. Guest-native tests, the PPC cross-build, and offscreen render tests pass. This is not green until a later drive clicks and types in the native Mirror and compares the whole same-moment frame with the guest capture, state, operation, and logs.

Update 2026-08-03, cycle 19: paired native frames confirmed the checkbox is box-shaped, the disclosure triangle is visible, and the buttons retain button chrome. The Depth popup still omitted its 8-bit value, and the missing Workshop structure made the whole-frame fidelity row fail. Those observations are inputs to the patch above, not a green result.

BROKEN: the staging reboot dirtied its own fresh clone (2026-08-03)

Emulator-observed; deferred from the Mirror UX arc. After tools/stage-ext.py, the old scripts/spin-up-ppc sent QMP quit; the cold boot of that same private clone then reported that the computer had not shut down properly. “Fresh clone” described ownership, not cleanliness after a host-side power cut. Later probes also found the current layers of both shared OS 9 bases already presenting Disk First Aid before staging, so neither was a valid oracle for proving a repaired stop path.

The first repair replaced quit with the parent tools/shutdown-guest, but the live os91-runner Worker did not grant its required script verb. Its hello advertised click, key, type, and launch; the Finder AppleScript was correctly refused. Two posted clicks are not a fallback: the first opens Special, then Finder is inside MenuSelect's tracking loop and the separately posted second click does not complete the held menu gesture.

Several bounded guest-native probes were rejected: ShutDwnStart, ShutDwnPower, Finder shutdown Apple Events, and an embedded OSA script did not power the VM off; a 120-second observer left it intact. QMP power/eject keys also failed, and relative-pointer capture was not a trustworthy held-menu gesture. This remains broken and is explicitly punted until after the Mirror's data-driven fidelity and direct-input loop are proven. Do not dismiss Disk First Aid, and do not present QMP quit as a clean stop.

WATCHED: the Workshop menu no longer overwrites the Apple slot (2026-08-03)

Emulator-verified, not metal-verified. The self-scene synthesized Help at left coordinate zero. That is the Apple menu's slot, so NOW Mirror showed Help over the left edge while the authoritative guest showed Apple, File, View, Windows, Help. The scene now reads MenuList.last_right, and a fresh paired live frame showed File, View, Windows, Help in the correct order.

This fixes one menu placement defect, not Workshop fidelity as a whole. The live Mirror still lacks sidebar icon representations, draws several control kinds with the wrong chrome, and defers CopyBits without a bounded placeholder. The Workshop is no longer structurally empty; those whole-frame mismatches remain red.

WATCHED: a person drove the guest from NOW's mirror (2026-08-03)

Emulator-verified. Open Mirror on the Mirror page opens a NOW window that renders the connected Mac and drives it. Watched, in the window, on a live Power Mac G4:

act what happened
click a scroll arrow the folder scrolled; status read "the lineDown of a scroll bar"
double-click a folder it opened, BY NAME, and its window appeared
drag a title bar the window moved to where it was dropped
pull a menu, pick a row View → as List; the window changed to list view
press a key "key e ✓" — the first keystroke ever to cross this wire

None of it uses QMP, so all of it is shaped for metal.

Objects first, which is what made the rest possible

A gesture now resolves to an OBJECT with identity — window, control, menu row, app, Finder item, and the desktop — and the gesture rides along as metadata the object interprets. Two things fall out that a gesture-first model could not express:

  • An icon is reached by name. NOW's contract has no click-at-a-point verb, so a desktop icon was previously unreachable by anything. As an object it is a file the Finder knows, and select item "X" of desktop works. Measured; so does item "X" of window "T", while the remembered target of window "T" fails with osaErr -1753.
  • The point picks the part. A press resolved as .lineDown of a scroll bar carries that, so no driver re-derives it from coordinates.

Three defects the machine found that no gate could

  • Every number this host sent arrived as ZERO. CommandRequest.args was [String: String], so part crossed as "21"; the guest reads numbers with strtol, which stops at the quote. Measured on a live scroll bar: 21 moved it one line, "21" moved it somewhere else, and both replies said dispatched. Fixed on both sides — CommandArg carries a number as a number, and now_json_read_int distinguishes absent from present-and-unreadable so a quoted one is now a refusal that quotes the fix.
  • key was unreachable over the wire, always. It read an argument called name; the guest scans a request FLAT, so it always got the envelope's own "name": "key" and refused every call as an unknown key name. The console face parses a typed line, so key space at the machine worked the whole time. Now named, and arg_shadow_source_test.py gates the whole class.
  • This Carbon guest cannot post a MODIFIED keystroke and says so: PPostEvent is not in CarbonLib. So ⌘ menu items take the MenuSelect route, and ActionPlanes.modifiedKeystrokes records the difference rather than every shortcut failing quietly.

Still open

  • Window interiors are empty. The content plane (P3) has never captured a drawing op, so a document window is chrome around blank space. Finder windows are fine — their icons come from the Finder.
  • A scroll thumb cannot be dragged. It needs a verb that SETS a control's value; ctlact presses a part at the control's own centre. Named as unsupported rather than approximated by paging, which would overshoot and read as a stutter.
  • No window raise. winact serves move/resize/zoom/close; nothing selects one window among an application's own, so clicking a background window fronts its APP and the rest is the Finder's choice.
  • role is still a guess (min != max), so About This Computer's memory bars are reported as scroll bars.
  • The process strip is drawn but not clickable. SceneRenderer paints it at the bottom of the guest canvas, so its pixels are in GUEST coordinates and HitTester — which only knows what the scene contains — resolves a click there to whatever guest window is behind it. Switching applications works, through the Application menu at the top right (appMenuappMenuItemactivate); the strip is the obvious-looking route and is the one that does nothing. Either it becomes a hit target or it should stop looking like one.

The mirror NOW draws itself: built, gated, not yet WATCHED (2026-08-02)

Unverified in the one way that counts. "Open Mirror" on the Mirror page opens a NOW window running Mirror's LiveMirrorView over NOWMirrorSource, which polls scene.request and dispatches to the act lane. Every link is proven against a live emulated Mac, separately:

link evidence
scene reaches the host NOWMirrorSource is this host's FIRST caller of requestScene
it decodes as IR v1 SceneIRDecodeTests, against a captured fixture
it draws SceneRenderTests, and a person looked at the PNG
a pixel finds the right element SceneHitTestTests (round trip)
the document agrees with the screen mirror-geometry-probe.py, real QMP click
the act moves the machine act-parts-probe.py, 60→156→60 by part code

Nobody has opened the window and clicked in it. That is the gap, and it is deliberately the human's: this side cannot screenshot the host app's own window (mirror/CLAUDE.md), so the assembled product is judged by a person or not at all. Both build systems compile it, Debug and Release, which is a different and weaker claim.

Known gaps in what it can drive

  • No positional click. NOW's contract has no click-at-a-point verb (asyncapi.yaml:3294, deliberate). So a click on bare desktop, on a desktop icon, or in bare window content is a NAMED refusal rather than an act. Controls, menus, windows and keys all work; the Finder's icons do not, and that is the largest hole in "drive the Mac".
  • Interiors are empty. The content plane (P3) has still never captured a drawing op, so windows render as chrome around blank space. The renderer is ready for it (MirrorKitUI.DisplayReplay); the guest side has never been armed in anger.
  • A raise is missing. winact serves move/resize/zoom/close but nothing selects one window among an application's own, so a title-bar click still falls back to a QMP press — emulator only. This is stated in MirrorAction.WindowAct rather than invented.
  • role is a guess. Derived from min != max, so About This Computer's memory bar graphs are reported as scrollbars and a click there asks a bar graph to scroll. The honest derivation needs the control's defProc, which the walk does not read.

FIXED: the mirror could not have clicked, and no gate could see it (2026-08-02)

Fixed on the guest, gated on the host, TESTED — not metal-verified.

NOW's scene emitted windows[].controls[].rect in global screen coordinates. IR v1 documents that field as content-relative, Mirror's own SceneBuilder subtracts the content origin to produce one, and MirrorKit.HitTester subtracts it from a click before it compares. So every control was hit-tested against a box displaced by its own window's origin.

Nothing errors when that happens, which is the entire problem. Measured on a live Finder: a point computed from the centre of one scrollbar in About This Computer resolved to a different control ninety pixels away, and a point at the centre of another resolved to the desktop. The render looked correct throughout — the same picture a person would call working.

This is the cause underneath "The last functional gap: a person cannot click the mirror" below. That entry described a chain with no join; this is what the join would have been wired to had it existed, and it would have mis-fired silently.

Why nothing caught it

Both conventions are four honest integers, so:

  • the decode gate passed (the document is structurally valid IR v1);
  • the render gate passed (a displaced control still draws somewhere);
  • the guest's own scene_walk_test asserted the wrong space in so many words — "a control's rect is translated to global coordinates" — and had done since the day it landed.

The only assertion that can hold this is one that names the space, so there are now two:

  • now-guest-ppc/tests/scene_walk_test.c — content-relative, stated, mutation-verified;
  • now-host/Tests/HostTests/SceneHitTestTests.swift — a round trip: compute a control's own centre from the document, hit-test it, require the same control back. It cannot pass through a space mismatch. Run against the pre-fix fixture it fails naming both the control aimed at and what was hit instead.

The second half, found by asking the machine

The control rects were only half of it. windows[].rect was the CONTENT region, where IR v1 wants that region grown UP by a title bar — the consumer recovers the content origin by adding the constant back, so a producer that skips the growth puts every control in the window twenty pixels low.

The round-trip gate cannot see this one, and that is worth understanding rather than patching: it derives the click point from the same rects it hit-tests, so an offset shared by both halves of the document cancels exactly. It stayed green.

So scripts/probes/mirror-geometry-probe.py asks the Macintosh. It computes the down arrow's position the way a renderer places it, delivers a real hardware click there over QMP, and reads the control back:

before   clicked (410,263)   value -4 -> -4    no change
         clicked (410,243)   value -4 -> 60    MOVED
after    clicked (410,243)   value -4 -> 60    MOVED, downward
         clicked (410,263)   value 60 -> 60    no change

Its negative control displaces DOWNWARD, and passing requires the arrow to scroll DOWN rather than merely to scroll. Both were learned in the same hour: the first draft displaced upward into the page-up region, watched the bar move for a legitimate reason, and reported inconclusive on a build that was already correct.

Still open

  • role on a control is derived from min != max (now-guest-ppc/src/scene/scene_json.c), so About This Computer's memory bar graphs are reported as scrollbars. Harmless to render; it means Scrollbar.part computes arrow and thumb regions for a thing that has none, and a click there would ask a bar graph to scroll.
  • The twenty-pixel title bar the two rectangles are related by is now stated in three places that share no header — SceneBuilder .titleBarHeight, the guest's kNowSceneIRTitleBarHeight, and the probe's own constant. The probe is what keeps them honest; there is no compile-time tie, and there cannot be one across a Swift package, a cross-compiled C guest and a Python instrument.
  • The FALLBACK window path (now_peek_windows_for_psn, taken for the self process and for anything that does not bind) reports the STRUCTURE region, which is a third convention again. It has not been measured and no consumer has complained, because the windows that matter come from the bound path.

UNVERIFIED: Hide has a route now, and nothing has watched it (2026-08-05)

This amends an entry that is not in this file yet. The branch this was written on forked before "BROKEN: both Hide routes fail, each in its own way" landed; when the two meet, fold this into that entry rather than leaving two. What follows is written to stand alone either way.

The route is real and it is the Process Manager's own. ShowHideProcess (selector 0x0060 under _OSDispatch, $A88F) and IsProcessVisible (0x005F) — Universal Interfaces 3.4.1, CarbonLib 1.5 and later, and our floor is 1.6, so they exist across the whole target range. The PowerPC guest now serves them as one hide verb on both faces (hide [--show | --status] <name>), and the contract declares it.

Two things had to be true first, and each had closed this route once before. The headers on the toolchain's include path are Universal Interfaces 3.4 and both calls arrived in 3.4.1, so the guest declares them itself. And Retro68 ships two CarbonLib import libraries that differ: the default -lCarbonLib resolves to toolchain/universal/libppc/libCarbonLib.a, which does not export them, while toolchain/multiversal/libppc/libCarbonLib.a (3.4.1-derived, byte identical to Retro68's own ImportLibraries/) does. The earlier finding that "the Process Manager's visibility call is absent from the toolchain under any spelling" was measured against the archive that lacks it. A sweep of a toolchain is only as good as the copy it swept.

What is proven, and it is less than it sounds. The verb cross-compiles; its argument grammar and its outcome vocabulary are native-tested (proc_hide_args_test.c); and the weak import is verified as far as the build products — the two symbols appear in the guest's PEF loader section as import class 0x82 (weak) inside a 0x40 (weak) CarbonLib entry, and the generated PowerPC for the guard loads the TOC word CFM fills and compares it against zero. Pointing the build at the archive without the symbols still links, and produces a guest that answers unavailable naming ShowHideProcess instead of one that will not build.

What is NOT proven is whether Hide works. No Macintosh, emulated or metal, has been watched hiding anything. Hide is UNBUILT in this ledger, not fixed. Two questions only a machine can answer are open:

  • Whether ShowHideProcess accepts the FRONTMOST process. The Application menu hides the front application, so the Process Manager plainly can — but whether this entry point declines it is not known here.
  • What it returns for the processes it is documented to decline, and whether a background-only application behaves differently. The verb reports the OSErr rather than swallowing it, so the first run answers this.

One trap found while building it, worth more than the verb. A CFM weak import is checked by comparing the symbol's address against kUnresolvedCFragSymbolAddress (zero) — and GCC deletes that comparison at -O1 and above, because a function designator is never null in standard C. Read in the generated assembly: the guard simply was not there. The address must be laundered through a volatile local so the compiler must load the TOC word CFM actually wrote. Any other weak-import guard in this tree written the obvious way is not a guard.

Swept 2026-08-05, and there are none — this worry is closed, not open. kUnresolvedCFragSymbolAddress appears in exactly one file (proc_actions.c, the guard above, correctly laundered), and no other guest or extension source compares a Toolbox function's address against zero or NULL by any spelling. So the rule stands as a rule for the NEXT weak import rather than as a defect anyone still has to go and find. It is worth re-running that sweep the first time a second weak import appears, because the failure is silent in both directions: the guard compiles, the tests pass, and the binary simply does not contain it.

"Agent: Running" was true and useless (2026-08-02)

Fixed on the guest, unverified on a machine. Measured on a live guest: NOW's Mirror page (now-guest-ppc/src/mirror/) reported the agent Running — correctly, the process was there — while that agent was bound to a stale port out of the base image's own mirror.port, so every connection from a host Mirror instance hit a QEMU forward with nothing behind it and was reset. RUNNING and SERVING are different facts and the page reported one of them.

Mirror's agent learns its TCP port from a text file called mirror.port sitting beside it, read once at launch (mirror/guest/app/src/main.c :: read_port, reached through set_dir_to_app — the file is beside the application and nowhere else). So the port file is now a fact of its own: mirror_probe.c reads it out of the same folder its catalog walk already resolves for the agent binary, and MirrorFacts carries three separable things — the process state, whether the file is there, and the port it names. The page has a Port row; the State row distinguishes running-and-named from running-with-nothing-naming-a-port; the placard carries the number. Enable REFUSES rather than launching an agent whose port nothing here can name, before LaunchApplication, because the alternative is this page manufacturing the state it was corrected for.

What it deliberately does NOT say is "serving nothing". read_port() falls back to a compiled-in 1420 when the file is absent, so an agent launched without one does bind something. The honest complaint is that the number is then a property of a binary nobody on this side can inspect — which is also why Mirror's own stager writes the file rather than leaning on that default. A page that replaced one confident wrong sentence with another would have learned nothing.

Three things it still cannot see, and two of them need a socket. It cannot report the port the RUNNING process actually bound: that was read at its launch and only re-reading the file now is available here, so a restage underneath a live agent shows the new number beside the old process. It cannot say anything answers on that port — nothing in this application opens a socket to Mirror. And the port fields are refreshed by the probe, not by the idle poll, because a file opened every second on the idle path is the starvation rule in guest-ui-start-here.md; a mirror.port staged while the page is open is stale there until the next action.

And NOW's staging now writes it. tools/stage-ext.py grew an optional Mirror bundle (NOW_STAGE_MIRROR=1, NOW_MIRROR_DIR): the three INITs, the agent, and mirror.port written with overwrite and truncate rather than inherited from the base image. Off by default — Mirror is a separate application and three more residents is not a thing to put on a guest nobody asked to. scripts/spin-up-ppc needs no flag; it passes its environment through. Host-cc tested (mirror_layout_test.c, mirror_port_staging_source_test.py, both mutation-watched); nothing in this entry has run on a machine.

BROKEN: the scene and observe disagree about the same machine (2026-08-02)

Measured, one machine, one moment, Finder in front, guest build 88507f25a8b9:

asked answer
axsnap Finder bind=ok, hasWindows=true, a5 0x1f50f550, fresh stamp
observe(front) Finder, window "Desktop", with a minted ref
scene.request one window, and it is NOW's OWN — no Finder at all
Mirror agent's axtree the front app's window with ten controls, plus Desktop

The scene enumerates all nine processes correctly and marks the Finder front, so enumeration is not the gap. The Finder's app row carries no error token, meaning its anchor resolved — the scene admitted none of its windows and said nothing about why.

Two readers in the same binary disagree about the same process at the same instant. observe goes through now_ax_bind_process and sees the window; the scene goes through peek_read.c :: resolvenow_peek_windows_for_psn and sees nothing. The capability is present and reachable — the scene is not using the path that works.

Why this matters beyond the bug. This was run as the go/no-go for dropping Mirror's agent and integrating fully. It answers it: NOW's scene is not structurally poorer than the agent's. It already carries the front application's MENU BAR (eight menus, eighty items) which the agent's axtree does not carry at all, and the windows it is missing are readable by code in the same binary. The agent is not compensating for something NOW cannot do.

Two instrument gaps found on the way, both worth closing with the reader:

  • The per-process anchor verdict is computed (scene_collect.c hands it to now_scene_add_process) and never encoded. A process that resolved fine and yielded no windows is indistinguishable from one that genuinely has none — which is exactly what this investigation spent its time on.
  • kNowSceneAnchorNoWindows produces no error token by design. Right for a process with no windows; wrong here.

Not the cause, but fixed while here: the scene path never armed the anchor plane at all. It does now.

FIXED: the act plane now acts inside foreign applications (2026-08-02)

The diagnosis below was right and the repair is in. act_install installed the six trap patches once, on the first armed pass, in whichever process pumped first — always NOW's own application, because it is the one serving the wire. The patches were then not in the dispatch path anywhere else. act_install now runs on each armed pass, and install_patch returns early when the incumbent is already its own shim — which makes a repeat install a no-op under a system-wide trap table and a real install under a per-context one, correct either way, and forecloses the fatal version where the chain points at itself.

Measured after the change, emulated Power Mac G4, dev INIT staged as "NOW Ext PerCtx" per the resident charter:

before after
actselftest vs NOW's own app abi-agreed abi-agreed
actselftest vs the Finder act-no-patch abi-agreed
actselftest vs SimpleText act-no-patch abi-agreed
menuact File/New Folder in the Finder act-not-taken, 0 folders 8/8 folders on the Desktop

The folder is the oracle — a fact on disk the Finder created, not anything a verb said about itself. This is the first time NOW's act plane has acted inside an application that is not its own.

A second defect the fix exposed, also repaired. With the plane working, menuact actuated 8/8 and reported act-not-armed 8/8: the resident arms and queues the press in one pass, so the application can dequeue, call the trap, and have the patch answer — setting fired, clearing armed — before this side looks at its snapshot. Reading armed alone called a completed request one that never armed. A false negative in the worst direction, since a caller that retries gets a second folder. All three sites (winact, ctlact, menuact) now ask whether the request reached the machine at all: armed OR already fired. Re-measured: 8/8 actuated, 8/8 replied ok.

And the guard was re-tested, because it had never really been. The menu no-hijack case was re-run on the same boot that had just driven File/New Folder 8/8: 0/17 hijacks, 17/17 clean chain-through, 3 dropped (a dropped trial is one whose QMP stimulus missed, measuring nothing). That is the number comparable to upstream's 0/19. The earlier 0/20 is not, and is marked void in the ledger: it was taken when the plane could not fire in that process at all, so a guard that held and a plane that could do nothing looked identical.

The diagnosis, kept: the act plane arms in a foreign app and its click is never taken (2026-08-02)

Measured, emulated Power Mac G4, guest build 3c6be9ffa460, both resident families staged. Against the Finder, addressed by PSN, with a titleLeft the scene supplied, menuact answers:

act-not-taken: armed, and the application never called MenuSelect

Read that carefully, because it is good news and bad news in one sentence. Armed means the extension's filter runs inside the Finder's context, the guard matched the target, and the plane posted its own press. Never called MenuSelect means the Finder did not consume that press. actselftest refuses against the same process in the same pass, while abi-agreeing against NOW's own application minutes earlier on the same build.

Every click-driven act verb depends on this one step. menuact, ctlact and winact all work by queueing a mouseDown/mouseUp with PPostEvent from inside the target's context and letting the application's own event loop dequeue it, call FindWindow, and call the trap the patch answers (ext/src/now_ext_act.c :: act_post_click, and the comment above it explains why the press is queued there rather than by the application). If the press is never dequeued, the whole family is inert in foreign applications no matter how correct the patches are.

It matches a finding that was never written down. The overnight arc of 2026-08-01 (claude/mirror-parity-overnight) measured the same family at 0/10 and recorded that a PPostEvent'd mouseDown is never delivered to any app on this guest while a keyDown from the same resident context IS. That branch's note lives in no document; this entry is where it now lives.

What it invalidates. The menu no-hijack case's 0/20 cannot be read as "the guard held" — a guard that held and a plane that cannot act in that process produce the same zero, and this measurement says the second is happening. Upstream's number has no such ambiguity because Portal measured 18/20 hijacks before its guard was fixed, proving it could act there. See the parity ledger.

ANSWERED, 2026-08-02, once the verb was made to report the plane's own error instead of the status. Four actselftest calls on one boot, guest build b77ba1c82e50:

target answer
A NOW's own app abi-agreed
B the Finder, by PSN act-no-patch
C SimpleText, by PSN act-no-patch
D NOW's own app again abi-agreed

D is the discriminator and it rules out the "only the first request since boot works" reading: A and D both agree, B and C both refuse. So this is about whose context the patch is asked to fire in.

And act-no-patch locates it exactly. cell->patches is one field in one shared table — the same value whichever process reads it — so if the GUARD's patches_present check were failing it would fail for NOW's app too, and it does not. The refusal therefore comes from the other place that returns kNowPeekActErrNoPatch: act_serve_selftest's if (!cell->fired). The resident called MenuSelect(0,0) from its own 68K code, inside the Finder's and SimpleText's contexts, and its own patch did not fire — while the identical call inside NOW's application fires and returns exactly what it wrote.

The measured fact, stated without a mechanism: the act plane's trap patch intercepts MenuSelect in the process whose context installed it, and not in others. The why is not established. The prime suspect is act_install's one-shot static int installed in ext/src/now_ext_act.c: it installs the six patches on the FIRST armed pass in whatever process happens to pump first — which is NOW's own application in every run so far — and never again. If Mac OS 9's trap dispatch is not as system-wide as a classic 68K machine's for this case, a one-shot install is exactly the shape of bug that produces this table.

The obvious experiment does not work, and why it does not is itself evidence. The plan was: from a fresh boot, arm while a FOREIGN application is frontmost so the first armed pass happens in its context, and see whether the answers invert. But act_install runs on the first pass of whatever process pumps, and NOW's own application is always pumping — it is the one serving the wire the request arrived on. So the install lands in NOW's context by construction, on every boot, no matter which application is in front. Fronting cannot move it.

That is not a dead end; it explains why the table always comes out this way round, and it makes the one-shot a stronger suspect rather than a weaker one. It also means the repair and the confirmation are the same change: make the install per-context (or prove the patch genuinely system-wide some other way) and the foreign-application answers should change. Per the resident-components charter that is developed as a throwaway dev INIT under its own name before it is folded in, because it edits the one file whose failure mode is a machine that will not boot.

Narrowed earlier the same day. The test was repeated against SimpleText — a plain classic application, launched, frontmost, and bound by the anchor plane (bind=ok, fresh a5) — and the result is identical: menuact answers act-not-taken, and actselftest answers act-refused. So:

  • It is not Finder-specific. It is general to foreign applications.
  • It is not only about the posted click. actselftest requires NO event to be dequeued by anybody: the resident calls MenuSelect itself, in the target's own context, and checks whether its own patch answered (act_serve_selftest). That refuses in SimpleText and in the Finder, while abi-agreeing in NOW's own application on the same build.

Since the anchor plane demonstrably runs in those same foreign contexts on the same event-loop pass (it captures their A5s), the sharper question is no longer "why is the press not delivered" but "why does the act pass not SERVE in a foreign process when the anchor pass in the same filter plainly runs there?" Candidates worth reading in order: now_ext_act_apply's verdict path and the A5 comparison it makes; whether the act arm bit is still set in arm_request at the moment the foreign process pumps, or has been withdrawn by the requesting application first; and whether act_install's one-shot static int installed interacts with which process armed first.

The event-delivery question is still real for menuact and remains open — but it is now downstream of this one, and fixing it first would prove nothing.

The Mirror page is a lifecycle now, and NOW cannot see residency (2026-08-02)

Landed, and the thing it cannot do is worth writing down. The Mirror module used to print the shell commands it was about to run and spawn swift run / spin-up.sh against Mirror's OWN throwaway emulator session — so a person with a Mac connected had no way to mirror THAT Mac, and a failed launch said nothing about why. It is now a module page that owns one Mirror instance pointed at the connected guest (MirrorControlModel, MirrorProduct, MirrorControlView): a status card, a lifecycle card, and settings. MirrorLauncherModel, MirrorLauncherView and their suite are gone, along with NOW_MIRROR_PATH and the remembered-checkout default.

The gap: NOW cannot tell whether Mirror's INITs are RESIDENT. The obvious probe is Gestalt('TBax') — the selector each INIT publishes at startup. NOW's gestalt verb takes no selector: run_gestalt gathers a fixed set and slices it by group, and the census selectors probe walks a closed documented list. Worse, an unknown argument on that verb is IGNORED rather than refused, so a host that sent one would get ok:true carrying every group and no evidence of the selector at all — which reads as a yes, which is the worst answer available. The page therefore asks software.list over the extensions domain, which sees the Extensions Manager disabled folder too, and reports installed / disabled / missing, saying plainly that an INIT loads at boot and that this side cannot see what is resident.

CLOSED on the guest side, 2026-08-02 — and the host still does not read it. The mirror verb landed: contract first, then the guest's wire face (commands.cmirror_json.c) and its console face (console_model.c), both rendering the same MirrorFacts the page draws. Measured on an emulated Power Mac G4 with all three staged: AXPeek resident v4, QDPeek resident v1, Portal resident v4, agent stopped, port named 1420 — the residency answer this entry says the host cannot get. NOW-68K answers unknown-command, and that is the ANSWER rather than a gap: Mirror's agent is PowerPC/CFM and its build refuses the 68K toolchain outright, so the residents have nothing to serve there (declared in contract-coverage.md).

What is still open is the two READERS. MirrorControlModel still calls software.list over the extensions domain and reports installed/disabled/missing, so the page a person looks at is still one step short of the truth the machine now tells; and there is no projection row, so no agent can ask either (declared unnoticed with its disposition in mcp-coverage.md). The verb is served and nothing reads it — which is the mirror image of the split this entry was written about, and worth not leaving long.

The original diagnosis, kept because it is still why the verb exists. The PowerPC guest's own Mirror page (now-guest-ppc/src/mirror/) calls Gestalt for all three selectors and distinguishes absent / resident / other-version — on its own screen only. That is the wire-only-versus-console-only split this repository has been bitten by before, in the other direction, and it is why the host has to infer from a folder listing what the machine already measured. Closing it is a contract change: either a mirror verb serving MirrorFacts (which the console face already has, so it is the cheaper half of command parity) or an optional selector argument on gestalt. Either way the contract moves first, then both guests' faces, the host projection, contract-coverage.md and the exact-set projection suites — and whether NOW-68K serves it is the parity question that arrives with it.

Never run against a real Mirror. The suite uses fakes throughout — nothing in this arc spawned MirrorApp, opened a socket to port 1420, or saw the page on screen. Specifically unproven: whether the launch invocation brings up a live window against a NOW guest; whether the emulator forward default (1724) matches the rig a person is actually running; whether mirror-agent is the name the agent's process wears in the guest's own process.list (it is the name Mirror's source and spin-up.sh use, read rather than observed); and whether SIGTERM releases the agent's single client slot as cleanly as the code assumes.

The host suite was fighting itself over ports (2026-08-02, settled 2026-08-05)

It WAS contention — and the suite was manufacturing it. For three days a red scripts/test-all here was read as "something else is running on this Mac", which was true and stopped the enquiry one step too early: the something else was another session's copy of this same suite, holding the product's own port because five of these tests take it by accident and none gives it back. Alongside that, the tests were hard-coding ports out of the range the kernel hands out, and abandoning several hundred sockets a run. Three separate defects, all in the test target, all now fixed — swift test is 1408 tests, 0 failures, and 51 s rather than 86.

What was actually wrong.

  1. Five tests were binding 5250 — the shipping port — and never letting go. SettingsModel reads an ABSENT listenAtLaunch as true and an absent (or zero) listenPort as defaultPort, so a HostAppState built on a fresh UserDefaults suite starts listening during init. Four cases in HostAppStateTests and one in GuestFilesCommandTests only wanted the module list and got a live listener on the port a person's own NOW app holds. Watched directly: lsof -nP -iTCP:5250 -sTCP:LISTEN during those tests shows xctest ... TCP *:5250 (LISTEN) before the fix and nothing after. UserDefaults.offTheWire() states it once. It also makes testHostCommandRegistrationDoesNotAddUIOrStartTheListener mean something: it was asserting .idle on a listener that WAS starting and had not finished binding — an assertion that passed by winning a race.
  2. Four fixed ports sat inside the ephemeral range. 52981, 52983 and 52987 were each chosen as "a specific, unlikely-taken port", and all three are inside 49152–65535 — which is exactly where this same process is handed every port-0 listener and every dial. So the suite took those ports from itself, and HostAppStateWiringTests, GuestStatusTests and ConnectionsModelTests intermittently failed to bind with EADDRINUSE. They use port 0 and read the bound port back now. HostAppStateWiringTests is the one test that must name a port before it binds — it is the only one still proving listen-at-launch — and it uses a pid-keyed port outside the range.
  3. The fake guest let the kernel choose its source port. macOS picks one by hashing the DESTINATION (RFC 6056), so one listener port is always offered one source port. A full run makes ~400 loopback dials and leaves ~150 of the sockets open — a test that stops its listener and drops its guest leaves one in CLOSE_WAIT for the life of the process — so when a later port-0 listener is handed a port that has been used before, the kernel proposes the source port that went with it, the 4-tuple already exists, and connect is refused. Network.framework does not FAIL such a connection: it parks it in .waiting and keeps it there, which at the test reads as a guest that never arrived. Measured with the state handler logging its endpoints: twelve dials in a row refused, every one from 127.0.0.1:55961 to 127.0.0.1:55963. FakeGuest names its own source port now, from a per-process lane above 1024 and below the ephemeral floor, never repeating a number in a run — so no 4-tuple can repeat.

What was wrong in the old entry, and is deleted. "The suites bind port 0, so this is not a simple port collision" — port 0 is what made it one, because the ephemeral range is shared with every client socket in the process. "A different subset fails on each run" (2026-08-02) and "the SAME subset every run" (2026-08-05) were the same defect at two machine loads, not two different signatures. The 2026-08-05 addendum also guessed at "a shared listener, a per-user socket path, a global/static, or a leaked task": there was a shared listener (5250), but the agent socket was never implicated — AgentIntegrationSocketTests already builds its endpoint under a unique temporary directory — and no global or leaked task was involved.

Found from the other end at the same time, and the two halves fit. A parallel session went looking for what was holding the machine rather than for what the suite was doing, and found it: check the PORT, not the process name.

lsof -nP -iTCP:5250 -sTCP:LISTEN

ps | grep for New Old World, swift build, swift test and xcodebuild matches none of them, because a SwiftPM suite runs as a bare xctest — so "nothing else is running" was never checked. lsof found an xctest from ANOTHER WORKTREE's session holding 5250, and it surfaced through scripts/spin-up-ppc, which now carries that check and says so in one line. That is the host-side twin of MetalMachineGuard this entry had been asking for.

But that foreign xctest was holding 5250 BECAUSE of defect 1. No test asks for that port; five of them take it by accident, and none gives it back. So "another session was running" and "the suite binds the product's port" are one fact from two directions, and the reading that the 2026-08-02 diagnosis therefore stands unchanged is too generous to it: the contention was manufactured here. A second xctest also explains the collisions in defect 3 far better than anything inside one process does — two suites drawing from one ephemeral range, each leaving ~150 sockets open.

What is NOT proven, and it matters. The five timeouts stopped reproducing on this Mac partway through the investigation, with the ORIGINAL code: a full run of the unfixed tests is green. The most likely reason is simply that the other session's xctest finished. So the before/after A/B that would settle it cannot be run any more. Defect 1 is verified by observation (the lsof above, watched by mutation: xctest ... TCP *:5250 (LISTEN) with the old code, nothing with the new). Defect 2 is verified by the failures in the trace logs (listener -> failed(...48...) on 52981 and 52983). Defect 3 rests on the captured port pairs plus a fix that removes the mechanism by construction, NOT on a watched before/after.

A reproduction was attempted and deleted rather than kept: building the collision by hand — dial a listener, abandon the socket, put a new listener back on that port — does not reproduce it, because the kernel's source-port choice advances on every successful bind elsewhere. A test that passes with and without the fix is worse than no test, so there is no guard here. If the five ever come back, the way in is FakeGuest's state handler: log state, currentPath?.localEndpoint and remoteEndpoint, and look for .waiting(EADDRINUSE).

The guard now exists (HostMachineGuardTests, 2026-08-05). It is a TEST rather than a line in scripts/test-host, because the person who reproduced this ran cd now-host && swift test, which no script wraps — a guard that only fires through the gate script is absent exactly when somebody is narrowing a failure by hand. It fails naming the process holding the wire port (watched by mutation: holding 5250 from another process produces python3.12 [pid 68233] *:5250 (LISTEN) in the failure text), takes NOW_ALLOW_BUSY_MACHINE=1 to proceed and label the result unattributable, and reports — without failing — any other copy of this suite running beside it. It reuses MetalMachineGuard's lsof reader rather than adding a second one.

Adding it turned up a fourth instance of defect 1, the worst one: a bare AppDelegate() builds its HostAppState on the PRODUCT's preference domain, so eight tests were reading a person's own saved settings and binding 5250 with them. AppDelegate takes an injectable defaults now (shipping behaviour unchanged) and the tests use quietAppDelegate().

Two things this entry has taught twice, worth keeping whichever way you come at it next time: a FIXED failing subset does not rule contention out, so subset stability is not a signal to reason from; and the only check that has ever given a straight answer is lsof on the port.

Still open: two more things the suite shares across processes

Found on 2026-08-05 by running two suites at once — which is NOT what swift test twice gives you, because SwiftPM locks .build and the second invocation waits. Invoke xctest on the built bundle directly, or run from two worktrees. Both of these are unfixed and neither is about ports:

  • HostLog.shared is one file per LAUNCH SECOND, not per process. Two runs starting in the same second share now-logs/<yyyy-MM-dd HHmmss>.log and read each other's lines: HostProjectionAuditTests testTheEventReachesTheHostLogInTheSpecFormat failed reading a line worth keeping, a string belonging to HostLogTests in the OTHER process, and LoggingSpecTests testALineMatchesTheFormatTheSpecDefines failed the same way. Two NOW apps launched together would do this too, so it is arguably a product defect and not only a test one; the fix (a pid in the name) is product-visible, which is why it is recorded rather than taken.
  • CloudModuleModelTests testTheToggleRemembersAndRestoresTheShare fails across processes on a share path. Undiagnosed.

Both were measured with the fixes above already in: one of the two concurrent runs was 1410 tests and fully green, the other carried these three. An earlier reading of the same experiment also blamed HostServingTests testGuestCanSendAFileAndItLandsInTheShare; that one stopped recurring once AppDelegate stopped binding 5250, so it was a knock-on and not its own defect.

Photo sizes became long-edge stops; three metal defects fixed, none re-verified on metal (2026-08-02, latest)

Unverified, deliberately labelled. Metal feedback named three things about the Photos save controls, and all three are fixed and TESTED — nothing here has been looked at on the PowerBook since.

  • The Size caption overprinted the popup. view_draw drew "Size" into size_popup's own rect — a popup paints its value across the whole control, so the caption landed on top of it and read as garbage. The caption now has CloudLayout.size_label of its own, on the same row at the group box's left inset, the shape dest_row + dest_btn already used one line below. cloud_layout_test.c asserts size_label.right <= size_popup.left relationally, watched failing under a mutation that puts the caption back on the popup.
  • "Host default" is gone. Every popup item now names a real size, and the host's configured setting arrives as data instead (CloudReport.defaultSize, additive) and is PRESELECTED. cloud.get from this guest always carries an explicit token.
  • The stops changed meaning. original / long640 / long1024 / long1600, each naming the LONGEST edge (aspect preserved, never upscaling). The fitN fit boxes are retired, not aliased — see the contract's own CloudGet.size prose for why the graceful refusal is what let a deliberate semantic break skip a revision bump.

What only metal can settle:

  • The caption and popup side by side at 640x480. The layout test proves they do not overlap in arithmetic; whether "Size" is legible beside a popup wearing "3024 x 4032" on a real 640-wide screen is a looking question, and the pane's inset clamp has never been seen.
  • A portrait photo actually arriving at 480x640. The scale is proven twice off-machine (PhotosProcessingTests through the real CoreGraphics pipeline, cloud_photo_size_test.c for the guest's label arithmetic, both mutation-watched) but never against a real PHAsset with real EXIF orientation, which is the one input a fixture cannot fake honestly.
  • The preselect on a real report. defaultSize riding the wire and moving the popup has run in no loopback test of the GUEST half — the guest's parser is unit-tested, the control mutation is not reachable from a host cc.
  • Whether anything still sends a retired token. Nothing in this tree does; a stale build on the PowerBook would, and would get the named refusal rather than a wrong picture. Nobody has watched that refusal land in the guest's status line.

RESOLVED: every modern classic-date field was silently dropped (2026-08-02)

Fixed, tested — not yet re-verified on metal

Watched on metal 2026-08-02: the iCloud Photos list showed "--" in Modified for every 2026 photo. Traced to ClassicDate.guestWireSeconds (now-host/Sources/Host/FileConverter.swift), which stopped at Int32.max — January 1972 in classic (1904-epoch) seconds — because the deployed guest read the field with strtol into a signed 32-bit long. Every date after that came back nil from the host function, modified was omitted from the wire entirely, and the guest drew the "unstated" dash instead. Not Photos-specific: CloudServices.swift, HostShare.swift and FilesModel.swift all route through the same function, so cloud listings and the drive/files browser carried the same silent gap — every date after early 1972, on every listing either guest reads.

A classic file date is actually unsigned seconds since 1904, good to early 2040; the host's ceiling was simply wrong, not conservative. Fix: the host stops clamping early (guestWireSeconds now just forwards macSeconds's own, correct, unsigned ceiling), and the guest gained an unsigned reader to match — now_json_find_u32 (now-guest-ppc/src/core/json.c) and now68k_json_find_u32 (already existed on the 68K side for CRC32, just needed pointing at this field) — used at every classic-seconds field either guest parses off the wire: the PPC guest's cloud listing rows, browse/pull replies (drive/files browser, file.pull), and both guests' file.offer push.

A second, independent site carried the identical bug and is not reached by ClassicDate at all: GuestFileUploadCommands.begin's own inline modified <= Int32.max clamp on the MCP agent-upload path (a raw already-classic-seconds Int, not a Date). Found by grepping now-host/Sources for Int32.max once the first site was fixed. Two PRE-EXISTING host tests turned out to encode the bug as correct behavior (asserting .modified == nil for a modern date) and needed correcting alongside the fix — caught by running the full host suite, not by writing new tests.

Tested, nothing more; full account in icloud.md. Host XCTest and guest json_native_test/cloud_model_test/test_putrx all pass, mutation-watched by hand. Nothing has run on the emulator or the PowerBook since the fix — confirming a 2026 photo's Modified column now draws a real date, rather than merely that the wire carries one, is the next metal session's job.

Drive's split-view pane has never run anywhere (2026-08-02, later still still)

Unverified. Drive stopped being the full-width flat list the 2026-08-01 entry below documents and went back to a list/detail split — the SAME split every other iCloud view uses, not a second one (cloud_layout.c computes one list/detail geometry and reuses it for drive mode too, differing only in list_top, pushed down by the breadcrumb row above it, and in the pane's own furniture below). The destination row and Choose... moved off the old toolbar strip and into the pane; the pull's moving bar and byte line moved there too, reusing Photos' own cloud_dl_bar_value/cloud_dl_bytes_line idle discipline against a different wire entry point (now_wire_get_active, since Drive pulls through now_wire_get_host rather than cloud.get); and the selected item's own name/kind/ size/date plus the double-click affordance line — which the 2026-08-01 review below moved onto the placard — moved back into the pane, so the placard no longer changes on selection or on a pull's byte count, only on durable folder/error/outcome news. No image preview for drive files: a drive row carries no cloud item id, so a later arc that wants one needs a real fetch-and-decode path, not this pane's text — the seam is named in cloud_drive_view.c's draw_item_card.

scripts/test-all is green with each exit code read directly (79 native tests including cloud_layout_test.c's rewritten, relational drive-mode assertions — the old ones asserted full width and had to change outright — both guest cross-builds, swift test at 1355 tests with 0 failures, xcodebuild Debug and Release), the new layout assertions were watched failing via a deliberate mutation before being trusted, and audit_source.py over both touched files raised only already-reviewed lexical categories (the new SetControlValue on the download bar is change-guarded, read back to confirm). None of this has run on the emulator or the PowerBook. The 2026-08-01 metal pass for Drive (below) predates every layout Drive has worn since, including this one — what it proves is the browsing logic (list, descend, Up, double-click fetch), not any pane pixels. Before this can move past "tested": watch the split render at 640x480 and at a roomier size, select a folder and a file and confirm the pane's text matches what the columns already say, start a pull and watch the bar/byte line move in the pane while the placard stays on the folder's own listing, and confirm Choose... still redirects a pull's landing folder from its new position.

The polish2 integration merged three UI arcs; the seam between them has never run (2026-08-02, later still)

Unverified, and the specific claim is narrower than "the union is untested." claude/polish2-drive-dest, claude/polish2-photos-cols and claude/polish2-contacts — each individually tested against the shell as it existed on claude/polish2-foundations — merged onto claude/polish2-integration with real conflicts in cloud_module.c and cloud_photos_view.c, not just adjacent additions:

  • cloud_module.c: view_own_browser()/active_browser()/ show_own_browser() had to generalize from two view-owned browsers (Drive, Photos, from the drive-dest+photos-cols merge) to three (adding Contacts) rather than picking either side's flag check wholesale — a real design decision made at merge time, not a mechanical union.
  • cloud_photos_view.c: photos-cols' own Data Browser (Name/Size/ Modified columns, the Size popup's exact-resolution labels) had to be kept while adopting polish2-contacts' extraction of the preview GWorld/fetch state out of this file into the new shared cloud_preview_well.c — meaning Photos' preview path now goes through the well's rebind-on-select note callback for the first time. Photos' OWN branch never tested against that extraction (contacts' branch predates photos-cols' columns); contacts' OWN branch never tested against Photos having a Data Browser of its own. Neither branch's tests can have exercised this interaction, only the merged tree's tests can, and scripts/test-all at the pure-logic level cannot see a Toolbox-level selection/rebind race.

scripts/test-all is green on the merged tree (79 native tests including all seven cloud_* ones, both guest cross-builds, the host suites and the Xcode app target) and audit_source.py over every touched now-guest-ppc/src/cloud/*.c file raised nothing new past already-reviewed, already-guarded lexical patterns. None of this ran on the emulator or the PowerBook. What only metal can prove, most load-bearing first:

  • The preview well correctly rebinds across a Photos-to-Contacts switch on a REAL machine. Select a photo, let its preview arrive on the new Data Browser, switch to Contacts mid-flight or right after, pick a card, and confirm the well's eviction/rebind hands the right pane its pixels — not a stale Photos preview drawn into the Contacts well, not a Contacts ask silently landing in the Photos pane.
  • Four browsers (shell, Drive, Photos, Contacts) sharing one window's activate/show lifecyclecloud_activate's lists[4] and show_own_browser's four-way dispatch are new arithmetic this merge wrote, unexercised past compiling and the pure geometry tests.
  • Every per-branch metal gap already ledgered below (Contacts guest UI, Photos download UX, Photos preview) still applies undiminished — this entry is additionally about the THREE arcs running together, not a replacement for any of them.

polish2-foundations: contract + host only, tested with fakes; the two real-data paths and the whole guest half are unbuilt (2026-08-02, later still)

Unverified, deliberately labelled — and narrower than the other 2026-08-02 entries: no guest UI exists for any of this yet. The foundations arc (contract: CloudGet.size grows fit1440/fit2048, CloudListing entries grow optional width/height, x-cloud contacts gains cloud.preview; host: PhotosCloudProvider.DownloadSize grows the same two boxes, .list fills width/height from PHAsset.pixelWidth/pixelHeight, ContactsCloudProvider.preview reuses the photos decode/fit/dither pipeline against CNContactThumbnailImageDataKey) is TESTED — loopback-proven with FAKE providers (CloudServingTests, CloudModuleModelTests) — and none of it has touched a real PHAsset or CNContact. What only a granted library/address book (this Mac's existing TCC grants) and metal can prove, additional to the items already ledgered below for PhotosCloudProvider:

  • PHAsset.pixelWidth/pixelHeight actually land in a real listing. The fill is one line reading documented public properties, but "documented and public" is a code-reading claim until a real library's rows carry real numbers a person can compare against Photos.app.
  • ContactsCloudProvider.preview has never run granted. The CNContactThumbnailImageDataKey fetch, a REAL contact that has a thumbnail, a real one that does not (proving the not-found "no photo" path fires from the actual store rather than only from a fake's scripted fault), and the reused pipeline against a real Contacts-app thumbnail's actual bytes (not the flat synthetic JPEG the loopback test generates) are all unexercised.
  • fit1440/fit2048 against a real multi-thousand-photo library. processedJPEG's box arithmetic is shared code already proven for the other three tokens (PhotosProcessingTests), so this is lower risk than a new pipeline — but "lower risk" is still a claim, not a measurement, until someone asks a real original at 2048x1536 and looks at the JPEG that comes back. (2026-08-02, later: moot as written — all five fitN boxes were retired the same day for the four longN long-edge stops, and the unmeasured claim now belongs to those. See the long-edge entry at the top.)
  • The guest half is entirely unbuilt. Nothing here has a now-guest-ppc counterpart: no Size popup entries for the two new boxes, no exact-resolution-from-dimensions arithmetic on the guest side, no contacts card wired to ask cloud.preview or draw the "no photo" placeholder. This arc is contract + host seams for those pages to consume, not the pages themselves.

No longer true for the contacts half (2026-08-02, later still): the contacts card now asks cloud.preview on selection and draws the "no photo" placeholder — see the Contacts guest UI entry below. The Size-popup entries and exact-resolution arithmetic remain unbuilt; those are Photos-only and this arc did not touch them.

Contacts guest UI shipped tested; nothing has run past cross-compilation (2026-08-02, later still)

Unverified, deliberately labelled — narrower than "tested" usually reads here. Built atop polish2-foundations: Contacts gets its own Data Browser (Name/Company columns, cloud_contacts_view.c, the drive browser's view-owned recipe), a real address-book card (photo well, name, grouped rows — cloud_contacts_card_layout in cloud_contacts_card.c), and a photo well shared with Photos (cloud_preview_well.c, extracted from cloud_photos_view.c). What is actually verified: the pure card layout is host-cc tested and mutation-watched (cloud_contacts_card_test.c), and the PPC guest cross-compiles clean with zero warnings. That is ALL — nothing here has run against a live host wire, on the emulator, or on the PowerBook:

  • The Data Browser itself is unwatched. Two real columns, its own UPPs, the fill-hilite call — all follow the drive browser's proven recipe, but "follows a proven recipe" is not the same claim as "watched drawing rows on the PB1400c."
  • The photo well's CopyBits landing is unwatched. Reused verbatim from Photos' own preview (metal status there is itself only loopback-proven, see the entries below), but landing into a SMALLER well (48x48) rather than the photos pane is new geometry nobody has seen render.
  • The hand-drawn silhouette placeholder has never been seen. A gray head-and-shoulders in two PaintOval calls, clipped to the well — geometry read by eye in the source, not by eye on a screen.
  • The preview-well extraction is a real behavior change for Photos, not just a file move. cloud_preview_well.c's _select rebinds the settle callback on every call, which changes exactly which view's pane gets invalidated when a late preview answer lands after the selection has moved on. Photos' preview path carried a metal pass before this refactor (2026-08-01); that pass does not cover the code as it exists now.
  • A contact WITH a real thumbnail has never been asked for. The wire path is loopback-proven (polish2-foundations, above) with a synthetic JPEG; nothing here has asked a real granted CNContactStore for a real photo and watched it dither and land in the well.
  • The card became titled GROUP BOXES (2026-08-02, later still) and no box has ever been drawn. The judged design replaced the flat label/value column with one kControlGroupBoxTextTitleProc control per section (Phone, Email, Address, Other), held as a fixed pool of four that a selection only retitles, moves and shows or hides. The pure half is host-cc tested and mutation-watched, and the guest cross-compiles — but the constructor is proven in this codebase only by software_module.c's ONE static box, never by four that move and retitle under a live selection. Three specific things nobody has watched: whether SetControlTitle + MoveControl on a visible group box repaints cleanly on CarbonLib 1.6 rather than leaving frame debris; whether the hand-drawn rows survive the box's own redraw ordering inside an update event (the pane is invalidated once per settled sync, which SHOULD make that moot, and "should" is the word doing the work); and whether the truncEnd values read as intended in the 70-point column at the smallest pane.

Photos download UX shipped tested; every visible behavior awaits metal (2026-08-02, later)

Unverified, deliberately labelled. The four-item download arc (the pane's "Loading preview..." state; the download bar + byte count off the new read-only now_wire_receive_active; the per-ask size on cloud.get — contract-additive, host loopback-proven with watched mutations, guest Size popup MENU 136; the guest-side destination chooser redirecting a cloud-born offer through now_files_receive_begin_at; and the receive-outcome seam replacing the stuck "Receiving..." status) is TESTED at its decidable seams and cross-compiles, and none of it has been watched on a machine. The specific things only metal can prove:

  • The furniture rows draw where the geometry says (size popup row, destination row, bar, byte line stacked over Save at 640x480), and the card/preview genuinely never draws under a live control.
  • The bar moves and the byte line ticks without flicker during a real multi-hundred-KB receive — the change-gates are unit-tested, the pixels are not.
  • A redirected offer lands whole in the chosen folder with type/ creator/date stamped, and choosing the share root really is byte-identical (it never sets the override; only a code-reading claim so far).
  • The outcome line replaces the status at completion on a real wire, including the refusal endings (exists / too-big / busy).
  • The popup CDEF under CarbonLib 1.6 accepts the fixed MENU 136 the way the services popup accepts its rebuilt one — same recipe, never this menu.

Photos preview + processing shipped tested; a granted library and metal own the rest (2026-08-02)

Unverified, deliberately labelled. The list+preview arc (cloud.preview / preview.begin / preview.end, contract-additive; ClassicDither; cloud_photos_view.c; the Downloads picker feeding cloud.photos.downloadSize into the get pipeline) is TESTED — pure ditherers with watched mutations, loopback serving with bytes-intact and lane-exclusivity proofs, in-test JPEG/HEIC fixtures for the resize pipeline, host-cc guest units — and none of it has met a machine. What only a granted library and metal can prove:

  • The palette is the real one only by construction. ClassicDither generates the standard 'clut' 8 layout (cube minus black slot, four ramps, black at 255) and dithers against it; the guest's GWorld wears whatever a NULL colour table gives an 8-bit depth. That the two tables are THE SAME TABLE on a real CarbonLib screen — the whole reason no palette travels — is a code-reading claim until a preview is looked at on the PowerBook. If colours arrive scrambled, suspect this first.
  • PhotosCloudProvider.preview has never run granted: the local-bytes-only fetch, the busy bargain for an un-materialized original, and a real HEIC through decode->fit->dither all need this Mac's TCC grant.
  • The pane under a held lane ("Preview after the download", the re-ask when selection moves mid-transfer) is guest logic past the pure units: builds only, exercised on no machine, and 1-bit asks (screens under 8-bit) have no fixture anywhere.
  • Downsized downloads against a real library: processedJPEG is fixture-tested; a 48 MP original through long640 (was fit640 until the long-edge arc, same day) on the wire to a real guest is not.
  • Preview pacing on real hardware: a 300x200 8-bit preview is ~60 KB, ~0.2 s at the measured 300 KiB/s — arithmetic, not a measurement; nobody has felt the selection-to-pixels latency at the PowerBook.

The cloud.* family: real providers are untested, and the guest half does not exist (2026-08-01)

Unverified / unfinished, deliberately. The host serves cloud.services/list/detail/get from a provider registry (now-host/Sources/Host/CloudServices.swift), tested over a loopback wire with FAKE providers only (CloudServingTests). Still unproven:

  • PhotosCloudProvider and ContactsCloudProvider have never run granted. They need this Mac's TCC consent (usage strings are in the Xcode project; the iCloud page has the grant buttons). First run: turn each on in the host's iCloud page, grant, and ask over the wire — cloud.list paging against a real multi-thousand-photo library, the JPEG/HEIC transcode, and the busy-then-bytes path for an un-materialized original are all claims from code reading.
  • The guest module does not exist yet. One Workshop page, service dropdown, per-service render (docs/icloud.md). Until it lands, the family is host-only and nothing exercises it end to end; when it lands, the guest's emitted cloud.* messages owe fixtures to GuestWireFixtureTests, and contract-coverage.md gains the family's guest rows.
  • cloud.get on a busy lane refuses busy by unit-tested logic, but no test drives a real concurrent capture/stream against it.

Update 2026-08-01, late: the entitlements fix is METAL-ADJACENT VERIFIED — with the hardened-runtime personal-information entitlements signed in, the Grant Access buttons surface macOS's real prompts, and with the grants given Michelle reports the granted services working as intended against the PowerBook, fan-out included ("functional enough"; her detailed notes are pending and may reopen items here). Narrower claims that remain untested by suites: the Photos fetch cache against a real library-change event, non-English Birthday parsing, unclipped long card values.

Update 2026-08-01, night: the fan-out landed (view seam, full-width drive browser + Up, contacts card view, photos hardening, live search) and its adversarial review's four must-fixes are in. Still open from that review: PhotosCloudProvider's fetch cache has no test (needs a granted library or a PHPhotoLibrary fake); contacts Birthday parsing matches English month names only (non-English hosts fall back to echoing text); long contact card values draw unclipped. Native tests now number 76; everything since the last metal pass — the whole fan-out — is tested, not metal-verified.

Update 2026-08-01, evening: METAL-VERIFIED for Drive on the PowerBook 1400c — the cloud.services round trip, the dropdown, and the in-page drive browser (list, descend, Up, double-click fetch) all watched working. Three faults the first metal pass found are fixed and their fixes watched: status_text garbage, popup menu reachable only through GetControlData, first-ask-before-connect. Still unproven: Photos and Contacts with real TCC grants, and cloud.get end to end (no serving service had it enabled yet).

Update 2026-08-01, same day: the guest module LANDED (now-guest-ppc/src/cloud/, docs/icloud.md) — parsers and geometry native-tested and mutation-watched, all three guests cross-compile, conformance gates cover the emitted asks. What remains unproven moves, not shrinks: the page has never been drawn on any screen (emulator pass owed first, then metal), the TCC-granted providers are still untried, and no end-to-end ask has crossed a real wire.

Update 2026-08-01, later: the metal-verified drive browser above is now a full-width flat list rather than the narrower list-beside-card layout it was verified in — cloud_layout.c gained a drive-mode variant (full width list, detail/save collapse to an anti-rect, a new up_btn in the toolbar row) and the card pane's per-row detail and pull progress both moved to the status placard (cloud_drive_view.c's draw is now NULL). TESTED, not metal-verified: scripts/test-all is green (74 native tests including new cloud_layout_test.c drive-mode cases, both guest cross-builds, host gate), and the new geometry was watched failing via a deliberate mutation, but nobody has driven this exact layout on the PowerBook or the emulator — the metal pass this arc references above predates this change. Before it: Data Browser's hierarchical/container surface (disclosure triangles, a real tree) was investigated and found not proven viable for this runtime — declared in the headers and compiles clean against a real container-callback call (spikes/databrowser-container-probe), but the container-specific entry points were never in spikes/databrowser's runtime symbol check against CarbonLib 1.6.0 on the PB1400c, so the drive view stays the flat, replace-on-navigate list it already had rather than adopt an unverified tree. Reopening that is a rerun of the runtime probe with four more symbol names, not another compile check — see spikes/databrowser-container-probe/README.md.

Update 2026-08-01, later: Photos hardened for an enormous library against FAKES only (docs/icloud.md > Hardened for an enormous library) — PhotosCloudProvider's PHAsset fetch cached per instance and invalidated by PHPhotoLibraryChangeObserver, a 10,000-row paging walk and the 4KB page bound proven and mutation-watched (CloudServingTests), a 3MB photo riding the ordinary transfer lane end to end, and the guest's cap-hit status wording made honest (cloud_listing_status, native-tested and mutation-watched). None of this touched a real PHPhotoLibrary: the cache's invalidation path, the real fetch's actual cost at 40,000+ photos, and whether Photos' authorization APIs behave as read on this Mac are all still claims from code reading, folded into the TCC-grant item above rather than duplicated here. PHAssetResource's byte size stayed out of scope — no public API exposes it short of downloading the resource — so CloudEntry.bytes stays unstated for photos, deliberately, not as an oversight.

Update 2026-08-01, later still: the three fan-out branches above (drive full-width layout, Contacts card view, Photos hardening) are merged into one tree (claude/swarm-icloud-integration, base claude/swarm-icloud-split). One conflict, in cloud_module.c's choose_service(): the drive branch added a layout recompute on every service switch, the contacts branch added per-service view dispatch — both intents kept, dispatch then relayout. Two more conflicts, in this file and docs/icloud.md, were two branches appending different ledger paragraphs after the same anchor line rather than true disagreement — both paragraphs kept. TESTED, not metal-verified: scripts/test-all is green post-merge (75 native tests including cloud_contacts_card_test, both guest cross-builds plus the NOW Extension, swift test at 1324 tests, xcodebuild Debug and Release) with the exit code read directly. Nothing here changes what each branch's own entry above already says is unproven — a clean merge does not prove Photos or Contacts against a real TCC-granted library, or put the new drive layout or the Contacts card in front of anyone on the PowerBook.

iCloud Drive sharing is tested against fabricated stubs only (2026-08-01)

Unverified. The share now sees a directory logically — iCloud placeholder stubs (.name.icloud) list under their logical names with the size the stub's plist promises, and a file.get for one calls startDownloadingUbiquitousItem and refuses busy with the reason (now-host/Sources/Host/HostShare.swift, HostShareCloudTests). Every test fabricates the stubs, so three claims rest on Apple keeping a shape no contract guarantees, and none has been tried against a signed-in iCloud Drive on this Mac:

  • the stub is a binary plist whose size lives under NSURLFileSizeKey (the fallback chain — promised-item API, then an honest zero — makes a format change degrade to "size 0", not a failure);
  • startDownloadingUbiquitousItem(at:) accepts the logical URL (the code retries with the stub URL, and swallows the error either way — the refusal is already on its way);
  • a download actually materializes the file where resolve will find it on the retry.

Trying it is cheap: sign-in, point the Sharing picker at iCloud Drive, browse from the emulator guest, pull an undownloaded file twice. Metal-verified is further still.

The name bridge (ClassicName) closed a live defect on the way: listed names were mangled one way (hfsName) and resolved verbatim, so any name the projection changed was advertised and then unreachable — file.get answered not-found for the listing's own spelling. Covered by round-trip tests now (HostShareTests, "The name round trip"), but the guest-side experience of fingerprinted names (Report#1A2B.txt in the Files page, Data Browser column width, MacRoman rendering of "#") has not been looked at on a real screen.

Related, found while mapping (2026-08-01): the host's serving half has no metal coverage at all. HostServingTests is loopback-only, and no Metal* suite exercises a real guest browsing this host's share. The browse direction guest→host is metal-verified only from the 2026-07-20 arc, before the name bridge and placeholders landed.

The Files path row names the share, unverified on metal (2026-08-01)

Unverified. file.listing.root now carries the host share's Finder display name ("iCloud Drive", "Downloads") through the standard MacRoman projection, instead of the raw POSIX path, and the guest's Files path row renders it — breadcrumbs from the share root for subfolders, with "Shared folder" kept only as the fallback for hosts predating the field. Host side is tested end-to-end over loopback (HostServingTests .testTheRootListingNamesWhatIsShared); the guest's label assembly is split Toolbox-free (files_path_label.c) and pinned natively. What nobody has watched: the row on a real screen — the root name arrives over the wire UTF-8→MacRoman via now_json_find_text, and an accented share name drawn through DrawString is exactly the kind of thing the emulator has hidden before.

The Mirror page has never been on a machine (2026-08-01)

Unverified, and the whole page is unverified together. The guest now has a Mirror page in the Workshop (now-guest-ppc/src/mirror/): three read-only rows for Mirror's resident extensions, three rows for its agent, and Enable / Disable for the agent alone. It builds, and its value core is covered by mirror_layout_test.c. Nothing about it has run on a Macintosh.

What a machine has to settle, none of which a host test can:

  • That Gestalt answers at all. The three selectors and the two or three longs read behind each of them ('TBax', 'TBqd', 'TBpt') come from Mirror's own shared headers, cited in mirror_probe.c. If a selector answers with an address whose magic does not match, the page says absent — which is the safe direction, and also indistinguishable from "we read the wrong offset".
  • That the agent is found where the page looks. mirror/tools/stage-agent.py puts the agent at Macintosh HD:TimBotTu:mirror-dev:mirror-agent; the page walks the boot volume to it and matches a running process by its processAppSpec. An agent staged anywhere else reads as not installed. There is no preference for the location and no browse button.
  • That LaunchApplication starts a faceless background application from a Carbon app, and that a kAEQuitApplication reaches one that owns no menu bar. Both are ordinary calls; neither has been watched against this particular target.

Why the agent is matched by file and not by creator. Mirror's agent is a Retro68 build with no creator override, so it carries the default '????' — read out of the MacBinary header of mirror/guest/app/build/mirror-agent.bin. So does every other Retro68 build on the machine, including the lab's own workers, which is why a signature match would cheerfully report the Mirror agent running about something else entirely. The signature is shown on the page and matched on by nothing. If Mirror ever stamps a real creator, this becomes a one-line change and a better rule.

The three extensions are deliberately not switchable, and the page says so in two lines rather than offering a control that does nothing. A file-move enable/disable is possible and is already proposed below ("an extension is a thing you enable, not a thing you launch") — it needs a guest verb, a confirmation, and the restart notice in the result, and none of that is what "status and enable/disable" asked for. Deferred, not overlooked.

No console or wire verb. This is a UI-only page: it adds no x-commands verb, no message type, and nothing to docs/contract-coverage.md. The parity rule is about capabilities the two faces of a guest reach, and nothing here is reachable from the wire because nothing here was added to it. A mirror console verb would be a real capability and would need both faces — worth doing, and not done.

The View menu was one item short, and this fixed it. Networking went in on 2026-08-01 without a menu item, and the handler maps item number to module id: Cmd-9 read "Logs" and selected Networking, Cmd-0 read "Connection" and selected Logs, and Connection could not be reached from the menu at all. Adding Mirror without repairing that would have moved the mismatch along. Logs and Connection now carry no Command-key — the digits ran out — and are one click away in the rail.

The Mirror port was thrown away (2026-08-01)

Settled, and it settles a great many entries below. NOW's re-implementation of Mirror's live-UI mirroring — MirrorKit, MirrorKitUI, the Mirror module's model, view, scene adapter, action driver, content join and window resolver, with their tests — has been deleted from now-host. In the built app its menu bar was mostly empty, its menus dropped down and did nothing, and nothing could be launched, clicked, moved or resized. Mirror already does all of it, working, on the same OS and the same emulator.

Mirror is now vendored whole at mirror/ — its own wire, its own 68K INITs, its own agent surface, its own SwiftPM package, built by nothing in now-host — and NOW's Mirror module is a launcher for its two halves (MirrorLauncherModel). The removed code is archived unchanged at archive/mirror-port-2026-08-01/, whose README says what is worth reading in it.

Update 2026-08-02: MirrorLauncherModel is itself gone. The launcher it describes pointed at Mirror's own emulator session and showed the shell lines it ran; the module now controls one Mirror instance aimed at the CONNECTED guest. See the 2026-08-02 entry at the top of this page.

So: every entry below that names MirrorKit, MirrorKitUI, MirrorModuleView, MirrorModuleModel, MirrorActionDriver, MirrorSceneAdapter or the Mirror pane describes code that is no longer in this tree. They are left standing per the rule at the top of this page — the shape of the mistake is the value — but none of them is a thing to pick up.

The lesson, which is not about Mirror: every acceptance number in that work was measured by probe scripts against the wire verbs, and the path a person actually uses was never once tested end to end. "winact closed a window 10/10" and "a person can close a window in the mirror" are different claims, and the gate only ever checked the first — so it stayed green for two days while the product did nothing.

Closed 2026-08-01: the guest spin-up works from here. It used to resolve the lab it borrows its emulator instruments from as its own parent directory, which inside NOW is this repository rather than the TimBotTu checkout that has them. Both scripts and both Python stagers now honour MIRROR_LAB_ROOT and otherwise walk up until a directory actually holds tools/lib.sh; MirrorInstallation.lab resolves it the same way and passes its answer down, so the preflight and the run cannot disagree.

Emulator-verified, not merely built: MIRROR_DISPLAY=1 tools/spin-up.sh from now/mirror/ booted a fresh mac99 clone (anchor at 90s), staged all three INITs, cold-rebooted with all three surviving, and the agent answered — oracle=ok v4, observe 9 processes front=Finder, axtree walking. Both preflight halves then read green against this checkout.

Two things it left behind:

  • stop-mirror.sh had a worse version of the same bug and now refuses rather than proceed: with no lab found, LAB resolved to /, the QMP quit failed into its own || branch reporting "VM may already be down", and the rm then unlinked the session disk out from under a QEMU that was still running it.
  • The standalone timbottu/mirror repository still carries the old resolution in all four files. It is not broken there — its parent really is the lab — but the vendored copy and the origin have diverged, and the walk-up version is the one that works in both geometries.

The last functional gap: a person cannot click the mirror (2026-08-01)

Retracted 2026-08-01, later the same day: the pane this describes no longer exists. See "The Mirror port was thrown away" above. The diagnosis below is why it was thrown away rather than finished, and is kept for that reason.

Broken, in the sense of unfinished rather than wrong. Every piece of the act path exists and is tested, and the path has no join. An agent can drive a Macintosh through the MCP act rows today. A person clicking a rendered control in the Mirror pane gets nothing — not a refusal, not a log line, nothing, because no code observes the click.

Three separate breaks in one chain. Each was verified against the tree on 2026-08-01, and none of them is recorded anywhere else.

1. The renderer has no hit-testing wired into it

HitTester.hitTest(_:x:y:) (now-host/Sources/MirrorKit/HitTester.swift:155) has no caller outside the test bundle. Nor does ActionModel.click(on:count:mods:) (ActionModel.swift:244), which is the only thing that constructs a MirrorAction. So at runtime no MirrorAction is ever built.

The pane draws and nothing more: MirrorModuleView.swift hands the scene to SceneView, which wraps a Canvas, and there is no onTapGesture, DragGesture, .gesture(, onHover or contentShape anywhere in MirrorKitUI/, MirrorModuleView.swift or MirrorModuleModel.swift. The pane's only interactive controls are Close Scene, Look Now / Look Again and Open Scene….

Other HitTester statics are live in production — isDesktopBackdrop, switchableApps, appMenuWidth, menubarHeight — which is why the type does not read as dead. The type is alive; the hit-testing is not.

2. The driver that would receive the gesture has no caller

MirrorActionDriver (now-host/Sources/Host/MirrorActionDriver.swift:56) is the seam a pane would call. It is built, it is tested (MirrorActionDriverTests.swift), and the only thing that constructs it is its own test. This is the half that could be finished without a machine, and it was; the pane is the half that wants one.

3. A window has no scene-side reference host-side

The guest emits windows[].refnow-guest-ppc/src/scene/scene_json.c:318 (put_ref(k, w->ref)), set by now_scene_set_window_ref (scene_build.c:314) off now_obs_walk_window_ref. It is an addition to IR v1's window field set, taken under the accretive rule, and the reason it was added is exactly this one: winact names a window, not a control.

Neither host model has a field to put it in. NOWSceneDocument.Window (now-host/Sources/NOWAgentIntegration/AgentIntegrationSceneModels.swift:195) and MirrorKit.Scene.Window (Scene.swift:155) both carry id / app / psn / title / rect / front / z / visible / kind? / controls? / text? / items? and no ref. NOWSceneCodec.decode is a plain synthesized Codable, so the key decodes without error and is discarded. MirrorSceneAdapter.window(from:) never mentions it.

Control refs do survive (NOWSceneDocument.Control.ref, mapped at MirrorSceneAdapter.swift:188). Window refs do not.

So winact has no caller from a rendered scene. The one place that sends it — AgentIntegrationActControl.swift:120 — takes its window argument from an opaque now-window-… minted by now_observe_elements, supplied by the agent caller. MirrorActionDriver has no winact route at all.

One correction to a phrasing that has been repeated: Scene.Window.id is not host-synthesised. It is minted by the guest at now-guest-ppc/src/scene/scene_build.c:197 as "%ld.%lu/%s#%d"psn.hi.psn.lo/title#z — deliberately in upstream SceneBuilder's own form, so an id minted here means what one minted there means. The host carries it through unchanged (MirrorSceneAdapter.swift:160). It is a name, not an address: nothing resolves it back to a WindowPtr.

The five faces are notReached, and honestly so

Each act row declares .appUI: .notReached with its reason, and the ledger is enforced both ways by HostFaceParityTests.appUIDivergences:

capability file line
now_window_act Projection/WindowActProjection.swift 72
now_control_act Projection/ControlActProjection.swift 55
now_menu_act Projection/MenuActProjection.swift 59
now_text_get Projection/TextGetProjection.swift 43
now_text_set Projection/TextSetProjection.swift 48

Eleven rows in total carry .appUI: .notReached; the other six are now_observe_elements, now_session_capabilities, now_transfer_approved_artifact, now_guest_files_capabilities, now_guest_files_upload_begin and now_guest_files_upload_append.

These declarations are the good news, not the bad. Rule 3 is recorded as owed, not waived, and the gate would have gone red if a row had claimed a face it did not have. What is missing is the pane, and the pane was correctly sequenced behind the thing it renders.

One reason has aged, and is worth fixing when the pane lands. WindowActProjection's reason says "the host has no window observation to select one from". That was true when it was written; the guest has emitted windows[].ref since 2026-08-01. The half that is still true is that the host model discards it.

Two stale claims in source, found while verifying this

Recorded here because they are in now-host/Sources/** and this pass owns no source:

  • MirrorSceneAdapter.swift:41-42 still says "NOW's walk reads a ControlRecord and cannot name a ref, so it is """. The reference plane landed 2026-08-01; the code below the comment already maps control.ref ?? "" correctly. Comment only.
  • ActionModel.availability's .key / .type reason (ActionModel.swift:130-137) says "NOW's contract declares no keystroke command." The contract declares key at contract/asyncapi.yaml:3024. What is actually missing is a host projection row — tracked as W3 in mcp-coverage.md. The refusal is right; its stated reason is not.

.activate reports available and this host has no lane (2026-08-01)

Broken, and it is a live inconsistency rather than a gap.

ActionModel.availability(.activate) answers .available(command: "activate") (ActionModel.swift:140-142), on the grounds that a scene carries a process serial for every window. The contract agrees that the verb exists — contract/asyncapi.yaml:2923 declares activate, taking serialHi / serialLo, described as not a second front.

This host carries no lane for it. There is no activate case in AgentIntegrationLocalProtocol.Operation. So MirrorActionDriver.drive(_:) passes the switch ActionModel.availability guard — because availability said yes — and lands in an explicit refusal at MirrorActionDriver.swift:145-155:

NOW's contract declares the activate command and this host carries no lane for it. The scene's process serial is not the opaque reference bring-to-front takes, so there is nothing to substitute.

The refusal is the right call and should not be traded for a substitution. now_bring_to_front takes an opaque now-process-… reference minted by process.list, validated by AgentIntegrationQuitPolicy.isValidReference, and re-listed and matched by full observed identity before it acts. A scene's bare "hi.lo" PSN string was minted by no host-side observation. Bridging the two would mean acting on an identity nothing on this side ever confirmed — which is the exact property the quit/front family was built to have.

What is actually owed is either a lane (an activate operation, with the serial's own validation story) or an availability answer that stops saying yes. Today the row is the one act in the vocabulary that reports sendable and has no route.

type and click are unavailable by design; key is now mods-gated (2026-08-01, updated same day)

Not a defect. Recorded because "why can't I type into the mirror" is the first question the pane will raise, and the answer is a hardware-era fact rather than a to-do for the MODIFIED half — but a plain keystroke is no longer one of these rows.

Updated same day: key split into two answers, not one. It read .unavailable unconditionally when this section was first written; that was too broad. now-guest-ppc's key verb posts an unmodified keystroke fine (mods is accepted as exactly 0) — the wall below is real for mods != 0 and was never a fact about mods == 0. ActionModel .availability(.key) now reads:

act mods answer why
key == 0 .available(command: "key") the guest posts it; MirrorActionDriver routes it to AgentIntegrationHostAdapter.key and the pane's drawing (MirrorModuleView + MirrorKeyCaptureView) sends one on a keystroke
key != 0 .unavailable the CarbonLib wall below — unchanged
type any .unavailable NOW writes text through textset against a referenced control (typeInto), never through a bare typed action with no target
click .unavailable NOW's contract declares no positional click. A control is acted on through ctlact by reference, not by where it is drawn

Not verified even for mods == 0: the pane's AppKit key-capture view (MirrorKeyCaptureView) has not been exercised in the running app — no display was attached to the work that added it. The specific, named risk is in docs/pane-keys-audit.md: whether its hitTest-returns-nil design actually leaves the drawing's existing click gesture untouched, and whether focus reaches it reliably after a click. swift build and swift build --build-tests both pass; nothing about the AppKit event path has run.

key still refuses modifiers outright. An event's modifiers live on the Event Manager's queue element, not in the message, and the only call that hands that element back is PPostEvent — which CarbonLib does not have (CALL_NOT_IN_CARBON). NOW's application is Carbon. So the guest can queue a keystroke and cannot say what was held down while it was typed; mods with any non-zero value answers unsupported and names the reason, and mods: 0 is accepted. The alternative — post the keystroke and drop the modifier — is a defect upstream already paid for: a literal character went into a document and the reply said success. Stated at contract/asyncapi.yaml:3036-3044, in input-plane-decisions.md, and in the guest at now-guest-ppc/src/input/input_args.c.

The reach exists, and only through the act plane's resident half. ext/src/now_ext_act.c is a 68K resident, not Carbon, so it can do what the application cannot: act_post_click() (now_ext_act.c:497-533) sets LMSetMouseLocation and calls PPostEvent for the press and the release itself, stamping evtQWhere and evtQModifiers. That inversion is worth holding onto — the older, less capable-looking half of this project is the half that can reach the queue element.

MirrorKit.SceneIslands kept its policy and lost its fetch (2026-08-01)

Unfinished, and it will read as dead code to the next auditor.

SceneIslands (now-host/Sources/MirrorKit/SceneIslands.swift:20) carries upstream's capture / hold / shift policy for pixel islands intact. The fetch it drives is an injected closuretypealias Capture = (Rect) throws -> PixelIsland (line 24), consumed by attach(_:poll:capture:) (line 53), island(for:key:capture:) (line 105) and the metered fetch(_:_:) (line 168).

Nothing supplies one. The only construction site in the repository is IslandLifecycleTests.swift:41. Nothing in Sources/** constructs SceneIslands or calls attach.

The file says so itself, and the reason is real rather than an oversight: the host's pixel path is GuestListener.requestCapture + CaptureDecoder, and no code joins it to a rendered scene. Joining them is a decision about the transfer lane — an island is a capture, and the lane is one transfer wide — not a wiring job.

Why it stayed: the policy is the expensive part and it is tested. A closure with no supplier is an honest shape for we ported the judgement and not the plumbing; deleting it would throw away the judgement.

The content plane has never run anywhere (2026-08-01)

Unverified in the strongest sense on this list, and not a fault to chase.

The reader is complete and natively tested against fabricated rings — now-guest-ppc/src/content/qdtrace_read.c (the ring walk and the seqlock), qdtrace_json.c (the replies), qdtrace_cmd.c (the only Toolbox, four subcommands: status / start / stop / drain), registered at commands.c:1376 with a help row.

The writer has never executed on any Macintosh. ext/src/now_content.c and now_content_logic.c are the resident half that fills the ring at draw time, and nothing has armed them — not on an emulator, not on metal, not upstream in this shape. No captured output, fixture or run log for an armed plane exists anywhere in the tree.

So qdtrace status answers content-plane-absent on every machine that exists, and that is correct. The refusal is emitted at three sites (qdtrace_cmd.c:198 on start, :275 on stop, qdtrace_json.c:420 on drain), gated on the caps bit kNowPeekTableCapContent (contract/peek_table.h:93, 1u << 3 after the collision described below). A run that gets content-plane-absent is not a failure. A run that gets anything else is news.

One stale comment in guest source, flagged rather than fixed here: qdtrace_cmd.c:11-15 still says "REGISTRATION IS NOT OURS". It is registered.

qdtrace's torn retraction: what is covered and what is not

The brief this checkpoint was written against said torn was "the one untested line". That is close and worth stating precisely, because the two halves have different standing.

Layer Path Covered?
Read qdtrace_read.c:307-326 — re-sample, seq1 != seq0 and the writer lapped the cursor → kNowQDDrainTorn, records = 0, resync = 1 Yes. qdtrace_read_test.c:526-543, driven by a deterministic lapping_sink
JSON qdtrace_json.c:444-450 — rewind e.pos to head, discarding whatever ops were already serialized No, and the test file names it as a gap (qdtrace_json_test.c:19-28)

Why the JSON line cannot be reached from a fixture: getting there needs a live writer lapping the ring between the seqlock sample and the re-sample. No host fixture can stage that, and a fixture that could would be staging the answer.

The nearest thing to coverage is a proxy, and it is a real one. Busy takes the same retraction branch, and test_busy_says_call_again (qdtrace_json_test.c:344-355) asserts the reply comes back with "ops":[]. So the discard is exercised; what has never been exercised is the discard after ops were written into the buffer.

The falsifiable claim, for whoever gets the first armed run: a torn reply is ok: true with "ops": [], "records": 0, "torn": true, "resync": true. The "ops":[ is emitted before the walk (qdtrace_json.c:407-409) and the retraction rewinds only as far as head. If a torn reply ever arrives carrying a non-empty ops array, that is the defect — it means ops survived a retraction that was supposed to discard them, and every one of them is a reading of a ring the writer had already overwritten.

Worth knowing about the shape: torn and busy are successful replies carrying flags; absent, mismatch and corrupt are ok: false errors. A caller that treats torn as an error will retry something that was telling it to call again.

Stale branches and two worktrees against a layout that is gone (2026-08-01)

Housekeeping, recorded rather than swept, because deleting another session's work is not this pass's call.

Four branches, none merged into main, none checked out anywhere:

Branch Head Behind main by Unmerged commits
claude/next-module-direction-02becd 3094e89 550 13
claude/laughing-tesla-b4cc41 686aa9c 605 10
fork/carbon-ui-cleanup b185b8a 692 6
claude/guest-installer 664cfd0 423 5

Two worktrees holding uncommitted edits against a directory layout that no longer exists:

  • .claude/worktrees/sweet-bouman-a714dd — HEAD a3f3adb, branch claude/sweet-bouman-a714dd. Five modified files: docs/open-issues.md, guest/src/commands.c, guest/src/wire.c, guest/tests/json_native_test.c, host/Tests/HostTests/GuestWireConformanceTests.swift.
  • .claude/worktrees/youthful-lumiere-d6e7be — HEAD 1cd1303, and the branch checked out is claude/68k-pn-180c-9c0940, not the one the worktree is named for. One modified file: guest68k/src/wire68.c.

Why they cannot simply be applied. main has no guest/, guest68k/ or host/ at top level — the trees are now-guest-ppc/, now-guest-68k/ and now-host/. Both worktrees sit on pre-rename commits. Salvaging an edit means path-mapping guest/src/now-guest-ppc/src/, guest68k/src/now-guest-68k/src/, host/now-host/, onto files that have moved and changed substantially across roughly 600 commits.

The honest read is that these are almost certainly not worth salvaging, and the reason to write them down anyway is that an uncommitted edit in a worktree is invisible to every other kind of audit. Whoever prunes them should look at the five diffs first and record corpus_impact for anything that turns out to be a finding.

Two planes asked for the same bit, and one collision was silent (2026-07-31)

Found and fixed during the fold-in, recorded because the near-miss is the lesson. The act plane (P4) and the content plane (P3) were ported by different agents, in parallel, neither able to see the other's edits to contract/peek_table.h. Both appended a capability bit and a state cell. Both asked for 1u << 2 and for the offset 36 + 60 * kNowPeekMaxAnchors.

The offset collision would have failed a compile — the header's static asserts pin every offset, which is exactly what they are for.

The bit collision would have been silent, and it is the dangerous one: arming the content plane would have armed P4's six trap patches inside another process. A person switching on a QuickDraw op counter would have been patching MenuSelect, TrackControl and FindWindow system-wide without asking for it.

P3 now sits at 1u << 3, appended after P4's cell. The shim keyed on NOW_PEEK_TABLE_HAS_CAP_CONTENT was deleted rather than left standing once it had retired.

What to carry forward: the accretive discipline (stamp_ticks never moves, gate on the format word, append only) was written for versions — one writer extending a table over time. It says nothing about two writers extending it at once, and parallel ports are now normal here. A test asserting that every cap bit is distinct and every plane's cell offset is unique would have caught this at the same moment the compiler caught the other half.

Related: now_act_guard_test went red on the append and was right to — it spelled "one byte short of the act cell" as sizeof(table) - 1, which is true only while that cell is the last field. A test written against the end of a struct is a test that fails the next time anyone appends.

now-host/Sources/MirrorKit/ActionModel.swift:92 hardcodes menuRowHeight = 16, and ActionDispatcher's .menuDrag releases on a point computed from it. That is the uniform-row assumption upstream measured as a ~30 px accumulated error once a menu contains separators — the rows are not uniform and the error compounds down the menu.

It survived the port because it is a constant rather than a mechanism, and nothing crossing looked at it.

The fix upstream built for this is MENU_GEOMETRY, which the act-plane port deliberately left behind on the grounds that "nothing in NOW consumes item rects." That reason has expiredActionModel consumes them implicitly, by assuming them. Porting it needs a new resident op (peek_table.h, ext/, the guard), so it is not a small change.

Until then, prefer menuact, which is identity-addressed and computes no geometry at all. The drag path this constant serves is emulator-only, so the blast radius is bounded — but a number that is wrong by 30 px two-thirds of the way down a menu will find a way to be believed.

Resolved 2026-08-01, on audit/menu-honesty. The ruling in docs/input-plane-decisions.md §3 was "measure the rows, or delete a computation nothing performs" — and by the time this branch landed, menuSelect already routed every item through .menuInvoke, so nothing performed it. menuRowHeight, ActionModel.menuItemPoint, and the MirrorAction.menuDrag case that was their only reader are deleted; no code path in now-host computes a menu-item pixel point from a row-height constant, live or dead. menugeom stays unported (still the riskiest call in upstream's file, still serving nothing) — re-open only if a caller needs an on-screen menu-item rect, per the re-open condition already on record.

Proposed: an extension is a thing you enable, not a thing you launch (2026-07-31)

Proposal, nothing built. From the manual review pass, and recorded here because it is a new guest capability rather than a UI gate — the gating half (never offering Launch or Bring to Front for an extension or a faceless background process) is being handled separately and is not this.

The user's shape for it:

extensions can surface an enable / disable function that just moves the extension between Extensions and Extensions (Disabled), plus a message that changes will take effect after restart

Three things make this worth writing down before anyone builds it.

It belongs on the guest, not composed on the host. The mechanism is a file move, and NOW already has a guest move verb — so the tempting cheap version is the host composing "disable" out of two paths it constructs itself. That is the projection layer deciding, which rule 2 forbids, and it breaks the first time it meets a System Folder that is not where the host assumed: a non-English system, a renamed volume, a machine with no Extensions (Disabled) folder yet. The guest knows where its own System Folder is and whether the disabled folder exists. The host should ask for "disable this extension", not for two paths.

The restart notice is part of the capability, not decoration. An INIT loads at boot and only at boot, so a disable that reports success is telling the truth about the file and a lie about the machine until it restarts. That gap is exactly the class of thing this product refuses to paper over elsewhere — it belongs in the result, not only in a label beside the button.

It is a destructive-ish capability with an easy undo, which puts it in the same family as the Files verbs: it wants the same confirmation and audit treatment, and an agent reaching it must appear in the audit line like any other mutation. It is also a good candidate for the consent tiers — a read-only tier should not be able to disable a system extension.

Open questions a builder must answer rather than assume: what happens when the disabled folder does not exist (create it, or refuse?); whether re-enabling has to remember where the file came from or can assume Extensions; and whether the 68K guest serves it at all.

A refused stream.start closes its bracket (2026-07-31)

The bracket is opened optimistically — activeStreamId is set before the guest has accepted — and its id is held by no pending map, so recordGuestError had nothing to match: the refusal set lastGuestError and the bracket stayed open on a stream that was never running. The 68K guest refuses stream.start every time (send_error_reply, now-guest-68k/src/core/wire68.c), which makes this that machine's ordinary behaviour rather than an edge case. GuestListener now recognises its own bracket id: it closes on the refusal, with the guest's own reason, and records the stream.start family through observeFamily in both directions. Tested, and mutation-proven both ways; no part of it has met a Macintosh.

Unverified:

  • Nobody has watched a real 68K Mac refuse a stream. The whole arc is proven against a fake guest that answers not-implemented on cue. What that cannot show is what the person sees: the Screenshots page should say the machine does not serve live streaming and grey the button, instead of sitting on "Waiting for the first frame…". That is the gate, and it wants the PowerBook rather than an emulator, because the emulated guest is the one that serves the family.
  • A guest that serves the start and refuses a STOP is reasoned about, not observed. Such a guest closes on the refusal rather than on the five-second fallback, and its stream.stop stays unproven rather than being recorded as a "no" — the three stream messages share one id, so the listener attributes a refusal to the open rather than guessing between them. No guest on the wire does this today, so the branch is untravelled.
  • The two capability stores still both exist. familyObservations on the listener now has the stream.start answer, and GuestCapabilityRecord — the page-side store that exists precisely because the listener could not see this refusal — records it separately from ScreenshotModuleModel. Whether one of them should now absorb the other is a question this change makes askable and did not answer.

The live stream reached the agent surface, and the bracket is a lease (2026-07-31)

now_stream_screen closes the last three unnoticed gaps in mcp-coverage.mdstream.start, stream.stop, stream.refresh — as one row with three intentions, so that list is now empty. Tested throughout; no part of it has met a Macintosh.

Unverified, and the first one decides whether the row should exist:

  • Nobody has measured whether a frame is cheaper than a capture. The whole premise is that an open bracket has the guest capturing continuously, so a frame is waiting rather than starting — against a capture measured at 0.5–0.6 s on the 1400c. If it is not clearly cheaper on metal, the row's reason for existing is wrong. The procedure is section 8 of metal-and-ux-review.md.
  • The default pace of 1000 ms was argued, not measured. It exists because the contract's absent-means-the-guest's-floor (~15 fps) is a Macintosh grabbing fifteen screens a second for a caller that reads one per call. The right number is a measurement nobody has taken.
  • The ownership rule has never met a real companion. Both halves are mutation-proven against injected values — a pid set and a movable clock — and both rest on an assumption about a real MCP companion's process: that it outlives a single call and dies with its client. If that is wrong, the liveness half is dead weight and the lease is doing all the work.
  • Nobody has seen contention happen. An agent's stream turns the person's live view on and greys out their Capture button; the sentences that explain that, on the Screenshots page and on the Agent page, have not been in front of anybody.

Three decisions worth revisiting rather than defects:

  • readOnlyHint: true, so the row sits at the Read Only consent tier. It is honest — a stream observes and changes nothing — and it means a machine that consented to being read has consented to a bracket that keeps reading, for as long as an agent keeps calling. The two tiers cannot express duration, which is the same gap now_reveal_item fell into from the other side, and more evidence for the middle tier. Declaring the row non-read-only to buy Full Access was rejected: it would corrupt the annotation agents actually read.
  • No maximum duration. An agent that keeps asking for frames is watching, and a ceiling would be a number with nothing behind it. The person can end any stream in one click. The cost is real and stated: a calling agent can hold a 1400c's screen lane indefinitely.
  • Capture does not end an agent's stream. The person wins by clicking Stop Streaming, not by pressing Capture — a button that says Capture and also silently ends somebody else's work does two things and shows one. If the UX pass finds that annoying enough, the other design is a small change.

One lesson that generalises past this row, recorded in source-text-gates.md: an asynchronous negative assertion is a gate that cannot fail. Two ownership guards were deletable with the suite green because "no stream.stop was sent" was read off the fake guest immediately after the call, before the message could have arrived. The cure is ordering against the wire, not sleeping — and applying it failed on unmutated code, which is how a real lease-renewal defect was found.

The agent surface can be seen, and refused (2026-07-31)

Plan 006 is built except its guest-side half. The host tracks companions, an Agent module shows what they have done, and HostProjectionDispatch refuses a call the connected machine has not consented to. Tested throughout; no part of it has met a Macintosh, and no person has looked at the pane.

Unverified, and the list is the point:

  • Nobody has seen the module. Everything asserted about it is about the model's words, not how they land in a window. The state most worth looking at is .neverAttached, because it is what the pane says on most machines for most of their lives. A screenshot on the host Mac closes this.
  • No real companion has ever attached. Presence, the 120-second active window and the LOCAL_PEERPID identity are all reasoned rather than observed against real agent traffic. Pid reuse can merge two short-lived companions into one — it undercounts rather than inventing, and is documented where it happens.
  • The ceiling has never met a guest that answers. No guest sends hello.agent yet except the PPC guest's hardcoded full, so disabled and read-only have been exercised only against fixtures.
  • The audit stream is per-launch and in memory, unlike the log, which can persist. A person looking for last week's agent activity needs the log.

Two decisions worth revisiting rather than defects:

  • now_reveal_item derives Full Access, against plan 006's stated intent that reveal is safe. Not a bug and not a slip: the row publishes readOnlyHint: false because it takes over the screen of whoever is sitting there, and the tier derives from the published annotation rather than a hand-maintained list — which is one of that plan's own stop conditions. The real cause is that two tiers cannot express reveal: derive from readOnlyHint and it is Full Access, derive from destructiveHint and so is upload, which writes to somebody's disk. Reveal is the case that fell in the gap when read / safe-write / full collapsed to two, and it is the evidence for reinstating the middle tier when something has actually used the first two.
  • Silence still fails open. Recorded in the schema as a decision, not a property, with the installer's arrival named as the moment to revisit.

One known skew: a host built before this change rejects an audit report carrying the new denied outcome. It costs one log line on a mixed install and never a failed call — deliberately cheaper than bumping the local protocol version, which would make such a host reject every request instead.

Debts the parity phase left behind (2026-07-31)

Twelve capabilities landed across twenty-six projection rows. These are the things that arc noticed and did not stop to fix, collected here rather than left in twelve agents' reports.

Gates that were not what they claimed:

  • Two source-scanning gates were decorative and nobody knew. The hello seam gate and the build gate each searched raw source for identifiers that their own explanatory comments also contained, so a mutation deleting the real call left each gate reading its own prose and passing. The build gate shipped that morning, mutation-proven at the time, and was hollow by lunch. Both now share a comment-stripping reader. Whether a third exists is being audited; the result belongs beside this entry.
  • MCPCoverageTests catches an omission and not its inverse. A capability that fails to add a familyPolicy row is named loudly. One that adds a row for a command, which needs none, passes in silence. Found by the machine- facts row, which has a test whose docstring claims the asymmetry and demonstrates it.
  • The guest-identity guard fires on prose. It scans Projection/ for guest names with comments included and has rejected doc comments four times this week, once per agent, costing an amend each. It is right about the rule and over-broad about the medium.

Timeouts classified wrong, twice:

The batched verb edit assigned each new operation a local receive window, and two were wrong in the same direction — a host bound shorter than the work it was waiting on. guest_file_mutation took the 2-second read-only window against a 20-second guest-side change watchdog, so a slow PBCatMove could time out locally on a call the machine then completed. census took the same 2-second window although its overview probe synthesizes every other probe. Both were patched by whoever tripped over them. The whole table deserves one pass, because the third instance will present as a machine fault.

Unexercised:

Ten of the twelve capabilities have never crossed a real wire — only capture and addressing are metal-verified. The capability ledger reads unproven on every guest by construction for several families, because the listener records no observation for them. That is honest and it means the first real call is also the first evidence.

Still open by decision:

Streaming (stream.start/.stop/.refresh) is the last unnoticed gap and the one genuinely undecided item. The 68K half of download stays a planned gap until HostProjection can express a disjunctive requirement — requires is a conjunction today, so a row needing "file.get or the put verb" cannot say so.

The machine's vote is carried (2026-07-31)

hello now has an optional agent field — disabled, read-only, full, or nothing — and the host decodes it, keeps it on the session health record, the roster row and AgentIntegrationSessionHealth.Guest, and writes it into the connect log line when the machine said something. That is section 2 of plan 006. Tested; nothing here has met a Macintosh.

The three unfinished things, and they are unfinished on purpose:

  • ~~Nothing enforces it.~~ Enforcement landed the same day — see "The agent surface can be seen, and refused" below. It went exactly where this entry said it belonged: HostProjectionDispatch, on the same line as the audit event. A machine sending disabled is now refused.
  • Absence fails OPEN, which is a decision recorded in the schema and not a property of the field. It matches today's default-on behaviour and keeps every deployed machine working. The moment to revisit is when the installer ships and silence stops being the common case.
  • Nothing on either machine can change the answer. The PowerPC guest answers full from now_agent_access(), a function with no preference and no switch behind it yet; the guest toggle, the mid-call prompt and the installer's AI-BAD path all land there. NOW-68K sends no agent at all — it has no switch to report and no installer, so it is a guest that has not been asked rather than one that answered.

The parity slice's capture lane and addressing met the PowerBook (2026-07-30)

Two capabilities of the parity slice had been proven at the codec and socket layers and never on hardware. Both have a gate now, and both ran on the PB1400c lab machine on 2026-07-29MetalCaptureProjectionTests, MetalAddressingTests, over MetalAgentLocalSurface. The rig is the shipped stack at every layer but the dispatch, which the file writes itself, mirroring App.swift.

Metal-verified

now_capture_screen, end to end from the agent face:

measured notes
screen 800x600
PNG at 1bpp 25,110 B in 4 pages 8 KiB pages
PNG at 8bpp 38,833 B in 5 pages
largest local response 11,643 of 16,384 B the local cap, enforced by the code that enforces it in production
guest-side transfer 278–349 ms

The claim no fake can make is the one that carries the entry: the reassembled bytes are decoded with ImageIO and their pixel dimensions checked against what the guest reported. Mutation-checked — one flipped byte in one fetched page fails the run as now-capture-digest-mismatch.

Addressing, four of the five selector states:

selector outcome status
absent answered by the driven machine metal-verified
the driven machine's id, and its session id answered by that machine metal-verified
a machine that is not connected now-guest-not-connected metal-verified
a session id whose connection ended now-guest-session-ended metal-verified
a machine connected but not driven now-guest-not-addressed not verified — see below

Mutation-checked: bypassing the refusal has the PowerBook answer for a machine nobody has ever seen, which is the substitution the scheme exists to prevent.

Not verified: the fifth selector state

now-guest-not-addressed means connected but not driven, and one connection cannot be in that state — the host refuses to re-point the console out from under whoever is at the machine, so the condition needs two live sessions to exist at all. testAConnectedButNotDrivenMachineIsRefused runs when a second peer is present and reports exactly what it needed when it is not.

It was exercised only with a supplied second peer (NOW_METAL_SECOND_PEER), and the distinction matters: the refusal is a decision the host makes — it holds two sessions, drives one, and a caller named the other. Nothing about that decision depends on what the far end of the second socket is, only on its being there. So a run with tools/fakeguest.py as the second peer is evidence about the host's addressing decision and about no guest at all; per AGENTS.md nothing verified against that harness may be called metal-verified. Unconditional coverage wants a second real Mac — or a QEMU guest — dialling the same port while the run waits.

capture.request reads unproven on every guest, by construction

The capability ledger cannot say more, and the reason is not the guest's: GuestListener.requestCapture is not wrapped by observing/observeFamily, so a settled capture records no family observation, and CaptureFailure carries a human sentence rather than the guest's typed refusal code, so there is nothing for the ledger to file even if it were wrapped. A capture is also deliberately not probed — it costs a whole screen grab and holds the connection's only transfer lane — so nothing else settles the row either. unproven is the truthful answer and leaves the capability callable; the note is in AgentIntegrationCapabilityLedger.swift beside the row. Fixing it is a behaviour change in the listener: give the capture lane a typed code and put the request through observing.

census.request joins it, and its page bound is unmeasured

now_hardware_census landed against the same two gaps and neither is new to it.

  • unproven on every guest, by construction. GuestListener.requestCensus is not wrapped by observing/observeFamily either, and the listener's own failure path folds a guest's typed refusal code into a CensusReport note ("[code] message") rather than keeping it typed — so there is nothing for the ledger to file even if the request were wrapped. The family is also not probed, and its reason is sharper than capture's: the probe argument is required, so a probe would have to choose one, and the registry's default is overview — the synthesis that arranges what every other probe read. Same cure as capture's: a typed code plus observing in the listener.
  • The adapter's 30 s page bound is a guess, and declared as one. census.request has no guest-side watchdog, so AgentIntegrationCensus.pageTimeout is the only bound on a probe. Not one census probe has ever run against a Macintosh (contract-coverage.md), so the number is the same order as the measurements beside it — catsearch ~20 s per pass, the software.list sweep ~4 s — and the first metal run is what replaces it. The local surface's window for the operation was moved off the 2 s read-only one at the same time, for the same reason: a page is 16 rows, and what costs is the probe.

software.list is the family that DOES settle, and no agent has settled it

Worth recording as the contrast to the two rows above rather than as a defect, because it is the shape those two are missing: GuestListener.listSoftware IS wrapped by observing, so an ordinary now_software_inventory call moves the ledger row to the guest's own answer and makes a later probeCostly report free. That is the cure capture and the census want, already working one lane over.

What is unverified is everything downstream of it:

  • No guest has served this family to an agent face. contract-coverage.md already records the software family as tested only — no guest has run the sweep for anyone — and now_software_inventory inherits that unchanged. The 25 tests behind it are over a real socket and a fake guest.
  • The ~4 s sweep figure is one disk's. It is metal-measured, but by catsearch on the 1400c. NOW-68K's apps path has two shapes the number has never covered: the 48-FSSpec cache, and the PBCatSearch-unusable fallback that walks the volume root. Neither has been timed on that machine, so nothing here knows whether the listener's 30 s watchdog is generous or tight there.
  • The note sentences have never crossed a real wire. Both are asserted against the guest's own literals, which proves the host carries whatever it is handed; it does not prove a 68K Mac with 60 applications actually emits the truncation note rather than a short page and silence.

The face-reachability proofs are textual, deliberately

Three coverage gates landed with the slice, and none of them proves what a reader may assume:

  • HostFaceReach.reached(file:symbol:) is file.contains(symbol) and nothing more. It catches the failure it was built for — the affordance deleted or renamed, the file gone — and cannot catch an affordance that is still spelled and no longer reachable: a call site wrapped in if false or #if, a control left permanently .disabled(true), a symbol surviving only in a comment or a #Preview, or the whole view no longer instantiated because its module left the sidebar registry with file and symbol untouched. Documented at the declaration rather than mechanised, because the mechanical version is a Swift-source reachability analysis and the honest cheap gate plus a stated limit beats a gate whose weakness nobody wrote down.
  • The MCP-face check is textual over NOWMCPServer's registry loop — it matches registry.projections.map and registry.projection(named:). A guard … continue added inside the loop body would skip a row without changing any matched string. It is still the stronger of the two: a loop fails uniformly, where a hand-built pane fails one row at a time.
  • docs/mcp-coverage.md and MCPCoverageTests are tested only. No part of the registry-versus-contract join has been read against a guest; the Served column claims only what a dispatch table answers. contract-coverage.md owns the how-far-proven axis and this file does not duplicate it.

Nine served capabilities that nothing asks for

docs/mcp-coverage.md derived the gap table and found the hand analysis had undercounted: nine capabilities are unnoticed — served by a guest right now with nothing in this repository arguing for their absence. They are absent because the question never came up, which is the process.list drift command-parity.md was written for, one layer out.

stream.start, stream.stop, stream.refresh, catsearch, gestalt, putstat, reveal, shotdiag, vprobe. Two are worth naming on their own:

  • gestalt is the largest single gap: one PPC verb answering CPU, memory, OS, network and hardware for the whole machine, served throughout, reachable from no face.
  • shotdiag is the verb that found the 180c's 24-bit addressing defect — precisely what someone standing at a misbehaving machine wants — and is reachable from nothing.

They were ten; capture.cancel left the list by being decided rather than by being built.

Updated 2026-07-30: all three diagnostics are now reachable, and only the streaming bracket is left on that list. vprobe, shotdiag and putstat are now_framebuffer_probe, now_capture_diagnostics and now_transfer_diagnostics, plus a Diagnostics module — three rows for one plan item and one wire operation, because requires is a conjunction and no guest serves all three (the argument is in docs/mcp-coverage.md, "One capability is three rows"). Tested, not metal-verified: nothing in this row set has run against a Macintosh, and the two unverified things worth naming are that the module's per-card availability reads the connected machine's own help table (so a machine that never answers help leaves all three cards unknown, which is stated rather than guessed past), and that the host's 40 s bound on a diagnostic is the only watchdog in the chain — neither vprobe nor shotdiag has a guest-side give-up, so a 68030 slower than that bound would read as a refusal and nobody has timed one.

AgentIntegrationLocalProtocol.swift is the real serialization point

Any capability needing a new client verb edits four things in that one file — an operation case, a result case, a response field and init parameter, and a strict-decode branch — at the tails of three lists. Those three tails conflicted on every merge that touched them this slice: the audit gate, the codec fix, its harvest, and the capture template. Always trivially, always needing a human decision. It belongs on the collision-hazard list beside contract/asyncapi.yaml and the scripts/test-native manifest: one owning agent per phase, and prefer batching a phase's verbs into a single edit over one agent per capability. W0.1's registry removed the tool-enum switch, which made the shared-file hazard look solved; it was displaced here.

Four hand-maintained capability lists survive the registry

So "one file plus one row" is true of the row and not of the capability. Each is trivial alone; eight times over it is a serialized edit on shared test files.

List Where
known-names set HostProjectionRegistryTests
approved-tool list NOWAgentCompanionTests
exhaustive switch over the operation enum AgentIntegrationSocketTests, NOWAgentCompanionTests
assertion matching a doc heading that names the tool count literally MCPCoverageTests (## What the thirteen reach)

The last is the worst: every new capability renames a heading in docs/mcp-coverage.md and a string in a test. Worth fixing before a wide phase, not during one.

An integer command argument cannot ride CommandRequest.args

CommandRequest.args is [String: String], so every typed argument reaches the guest quoted. A guest reading an integer argument uses now_json_find_int, which is strtol on the byte after the colon — strtol("\"40\"") is 0. The failure is silent in the worst way: tail's run_tail clamps 0 up to 1, so a caller asking for forty log lines gets one, with ok:true and nothing anywhere saying so.

Nothing had met this edge because launch, reveal, help and the census all take strings; tail (P1 #9) is the first host-side caller whose typed argument is a number. It sends the count on line instead, which the contract declares for that verb (x-commands.tail.x-line) and which run_tail reaches precisely when no typed lines is present, with a test that fails if somebody tidies it back into args.

Unfixed, and it will bite the next numeric argument. The fix is a typed args value on both sides of the wire, which is a contract/asyncapi.yaml + both-guests change and belongs to one owning agent, not to whichever capability trips over it. Until then: an integer argument goes on the line, and a reviewer seeing a number in an args dictionary should ask what the guest parses it with.

The guest log is readable by an agent and has not been read on metal

now_guest_log_tail (P1 #9) is tested, not metal-verified. Two things about it are worth knowing before it is trusted on a real machine:

  • The audit line it writes under app shares the state of every other agent-facing log line — see The agent audit line has never been read on a real run below.
  • It is the first row that returns text the machine wrote, so it is the first that can disclose a name from outside guestRoot: the guest's own get, put and files lines quote the items they handled. That is argued and recorded in mcp-coverage.md rather than accidental, but it is a widening over the Files family's authority and a reviewer should agree with it explicitly rather than inherit it.

A schema rejection surfaces as the wrong error

AgentIntegrationLocalServer replies to a decodeRequest failure with .init(error:) and no request id, so the response carries requestID: nil. The client checks the id first (decoded.requestID == request.requestID), so the caller sees "Local response request ID did not match" rather than the invalid-request error anyone would grep for. Minor, and explanatory: it is why the guestSelector defect below presented as a mismatched id rather than as the schema rejection it was.

Dead code

AgentIntegrationLocalProtocol.strictObject(_:keys:) — the private overload that also requires the keys to be present — has no callers. Both call sites use strictObject(_:allowedKeys:).

vprobe's CopyBits failed is not capture.request's

MetalCaptureProjectionTests was written expecting the guest might refuse, because an earlier vprobe on this PowerBook reported CopyBits failed. It did not reproduce: two clean captures at two depths. The two paths differ — vprobe measures framebuffer reads on its own bands, the capture lane stages through the guest's normal screen grab — so a CopyBits failure in one is not evidence about the other. Recorded so nobody conflates them again, and so the reverse is also clear: had the capture been refused for that reason, it would have been a finding about this machine and not a defect in the gate.

This distinction is now carried by the product rather than only by this ledger (2026-07-30). vprobe has a face on both sides — the now_framebuffer_probe tool and the Diagnostics module's first card — so the misreading is available to more people than the two who wrote these paragraphs. The tool description states it, and the card states it before the probe is run rather than beneath a number that has already sent someone looking for a bug in Screenshots.

PRODUCT_VERSION cannot tell two builds apart (2026-07-30)

Broken

PRODUCT_VERSION is "0.1.0" in now-guest-ppc/src/core/product_identity.h and was also "0.1.0" on the build previously deployed to the 1400c. It rides hello and is what now_session_health reports, so the one string a host has for "is this the build I just deployed" answers the same for every build there has ever been.

This cost a real misdiagnosis on 2026-07-30: a stale guest on the 1400c was failing every exec test, and the version string gave no signal that the machine was running old code. The metal gates now assert which build answered from the host-observed address and from the guest's own verb table, never from PRODUCT_VERSION — which is the right workaround and not a fix.

The fix is a build identity that changes when the build does. NOW already has build_stamp.c, which CMake touches at the end of every build and AGENTS.md already tells a human to check before believing a test result; putting that stamp where hello can carry it is the cheap version. (Hypothesis, not measured: the 68K guest has its own version string and is likely to have the same weakness.)

The agent audit line has never been read on a real run (2026-07-29)

Unverified

Every capability the MCP face invokes now emits one audit event, and the host writes it under the agent area of its log (agent-integration.md). The gate is mutation-checked and the whole path is exercised over a real private socket in NOWAgentAuditTests — but with a fake host at the far end. Nobody has yet driven the running app from a real MCP client and read the lines out of ~/Library/Logs/now-logs. The host app's own handler for the operation (the .audit case in App.swift) is the one piece with no automated cover, because nothing in this tree tests that closure; the line's format is tested one layer in, at AgentIntegrationAuditLog.

Two known gaps, stated rather than left to be discovered:

  • A call still waiting on a 32-second launch has not been logged yet. The event is emitted once, when the outcome is known — a begun/ended pair would double this face's local round-trips per call — so a launch in flight is invisible until it settles, and one that takes the process down is never logged at all.
  • A malformed guest selector is refused by the face before any capability is invoked, so it names no capability and emits nothing. (Still true, and now true one layer deeper: since 2026-07-29 the codec refuses an empty selector as well — same consequence for the audit line.)

The 2026-07-29 metal run did not touch this. MetalAgentLocalSurface refuses .audit by name, along with every other operation that could change the machine, so nothing about the audit path was exercised on the PowerBook. This entry stands unchanged.

RESOLVED: local schema v7's addressing could not survive its own codec (2026-07-29)

Fixed, and metal-verified

Both defects are fixed on this slice and the path is metal-verified on the PB1400c lab machine, 2026-07-29 for four of the five selector states. The table, and why the fifth state one connection cannot reach, are in "The parity slice's capture lane and addressing met the PowerBook" above.

What the fix is:

  • decodeRequest admits guestSelector — once, in the top-level allowlist, and as a conditional per-operation key added only when the caller actually sent it, so an absent selector stays absent rather than becoming a required field on every operation.
  • An empty selector is now refused as its own error ("Local request names an empty machine") rather than reaching the adapter as a third state that is neither nil nor an id. Validated in the codec and not only in the companion, because the companion is not the trust boundary: any process of this uid can write that socket.
  • decodeResponse admits notAddressed and counts it in the exactly-one-of guard, beside the operation results rather than outside them — the refusal is set instead of an answer, so a response carrying both is malformed for the same reason two results are.
  • AgentIntegrationAddressingCodecTests asserts, from a Mirror over the request type and over the response type, that each allowlist admits every field on it — derived rather than listed. That is what makes the whole defect class visible rather than these two instances of it.

What was wrong, kept because the shape recurs

Found while adding the audit operation; both were on main, both untested, and grep -rn "guestSelector\|notAddressed" now-host/Tests returned nothing, which is why neither was failing anything.

  • decodeRequest omitted guestSelector from its allowedKeys and from every operation's expectedKeys — a strict object check — so any request that actually named a machine was rejected as not matching the schema. Nil selectors are omitted by the encoder, which is why the single-Mac path kept working and the whole machine-id / session-id scheme landed 2026-07-28 could not work through this path at all.
  • decodeResponse omitted notAddressed from its allowlist, so the refusal SocketAgentIntegrationClient is specifically written to pass through as itself arrived as now-host-invalid-response: a real refusal wearing a protocol error.

Two omissions of one shape is the lesson, not either omission: an allowlist and the fields it is supposed to admit are two lists that drift silently, in both directions, and nothing observed it here until an unrelated feature needed the field.

NOW-68K has a hardware census, and none of it has run (2026-07-28)

Unverified, in the strongest sense on this list

NOW-68K answers all fourteen probes of the contract's x-census registry, on both faces (census.request on the wire, the census verb for a person), where it previously answered every one of them refused with the note "no probes implemented". Not one probe has run on a Macintosh - emulated or metal.

That is worth stating sharply because the automated cover looks better than it is. test_census.c (native, 680 checks) covers the page, the cursor arithmetic, the frame bound and both renderers, and the host decodes three pinned frames. All of that is the half with no Toolbox in it. Every PROBE is Gestalt, the GDevice list, PBHGetVInfo, the drive queue, the unit table, ADB, GetSysPPtr and the Power Manager, and no gate in this repository can reach any of them.

What a first pass should look at, in rough order of what could plausibly be wrong:

  • drivers is the riskiest walk. It reads the Device Manager unit table from LMGetUTableBase, and a driver's name is a Pascal string at offset 18 of the driver header - reached through a HANDLE for a RAM-based driver and a POINTER for a ROM-based one (dRAMBasedMask). The flag is checked rather than assumed, but the check has never been exercised. Bad names, or a hang, would point here.
  • power picks its call from gestaltPMgrDispatchExists: GetScaledBatteryInfo (a _PowerMgrDispatch selector) where that bit is set, the classic BatteryStatus otherwise. A 180c under 7.1 is expected to take the second path. If the machine takes the first and the selector is not really there, that is a crash and not a bad row.
  • pram should read valid $A8 on a machine with a live PRAM battery and something else on the 180c, whose battery is dead (the entry below). If it reports $A8 there, the byte is not saying what this probe claims it says.
  • adb should find two devices on the PowerBook (keyboard and trackball) and may find none under an emulator, which answers absent - correctly, and worth not misreading as a defect.
  • ata, pccard, pci should all answer absent on the 180c and the Q800, each with its reason. An absent there is the probe working.

What is deliberately NOT served, with the reason

Two of the fourteen answer refused, which means this build declined to look rather than the machine saying no:

  • scsi. The contract calls it the declared exception to passive-by-rule - an INQUIRY bus scan is active bus I/O - and says attended first runs on real hardware are the expected discipline. The 180c's internal disk is on that bus, nobody has ever attended a scan from this guest, and a wedged target on a cooperatively-scheduled 68030 is a power cycle. drives and volumes answer what is attached without touching it. Doing this properly means someone in front of the machine, which is the whole reason it is parked.
  • selectors. The PowerPC guest walks a snapshotted table of documented Gestalt selectors; that table is 32 KB of names, against a 384 KB partition. identity carries the rows a person actually reads.

Cost, measured

The 68K binary grew 197,248 -> 215,808 bytes of code (+18.1 KB, ~4.7% of the partition), and the census owns two ~1.1 KB BSS pages - one in wire68.c for the wire, one in commands68.c for the console, separate because a request arriving while somebody is reading must not overwrite the page they are reading. Most of the code is the notes: fourteen probes' worth of sentences explaining what absent means on this machine. That is the trade this subsystem exists to make.

Still missing beside it

gestalt - the five-group verb - is now the largest thing NOW-68K does not serve, and it is mostly a renderer: health.c already samples the facts and the census now reports most of them again. A host asking for it by name still gets unknown-command.

NOW-68K's software listing has never touched a disk (2026-07-28)

software.list and the sw verb are served on NOW-68K (now-guest-68k/src/software/). The status is tested, and the half that is tested is the half with no Toolbox in it.

Unverified

  • The sweep has never run. n68_swenum.c is pure Toolbox and no gate in this tree can reach it. Nothing has confirmed that PBCatSearchSync finds a single application on a System 7.1 volume, that the disabled sibling folders resolve through FindFolder, that the parent-chain climb produces a launchable HFS path, or that a folder domain's two-catalog cursor lands on the right item at the boundary. The emulator (scripts/q800-68k) can answer all of those and has not been asked.
  • The timing is a guess. The contract records ~4 s cold for the equivalent sweep on a PowerBook 1400c. The 180c is a 33 MHz 68030 with a much smaller, much older disk, and nobody has measured it. The budget in force is proc_launch_search_seconds() (20 s by default, shared with launch), so the honest statement is that the sweep will either finish or truncate inside 20 s — not that it finishes.
  • The pump has never been exercised under load. The sweep calls proc_yield_ticks() between slices, which runs wire_idle() and can re-enter the frame reader. That is the DEFECT 3 path proc68.c documents; it is guarded by the same single pumping flag, which is precisely why the pump was exported rather than copied. Nothing has pipelined a second request into a running sweep to watch it hold.
  • 48 may be the wrong bound. NOW68K_SWLIST_APP_CACHE_MAX was chosen from the memory budget (3360 bytes of BSS), not from a count of what is on the 180c's disk. If that machine has 200 applications the listing is honest and mostly useless; if it has 30 the bound never fires. One sw apps on metal answers it.

Open

  • No version, no running. Both are omitted on this guest with their reasons written down (contract-coverage.md). version is the one a person is most likely to want, and the bounded way in exists — a page is at most ten entries, so ten resource-fork opens — if the heap on a 4 MB machine turns out to tolerate it. That is a measurement, not a decision, and it has not been taken.
  • launch and the listing do not share a search. proc68.c sweeps for one named application and n68_swenum.c sweeps for all of them; the SHAPE is shared (slice, budget, retry, fallback) and the code is not. Two sweeps that drift would disagree about which applications this machine has, which is the two-halves-never-met-in-a-test shape one file over. Worth folding together the next time both are open.

The 180c's garbled capture was 24-bit addressing (2026-07-28)

Fixed, and confirmed on metal by remedy

A screenshot taken on the PowerBook 180c saved correctly to that machine's own Desktop and arrived at the host as structured noise. shotdiag, run on the 180c, answered it in one pass:

Base          0xFC080000
StripAddress  0x00080000
Addressing    24-bit (!)
Walk row 0    04 0F 0D 07 01 04 02 0E 0F 02 0B 0D 08 01 03 0A
Walk again    04 0F 0D 07 01 04 02 0E 0F 02 0B 0D 08 01 03 0A
Blit row 0    00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
Verdict       DIFFERS at byte 0 - wrong memory

The machine was in 24-bit addressing, so the top byte of the framebuffer's address was thrown away and every raw read went to 0x00080000 — main RAM. Walk and Walk again agreeing proves the screen held still, so the run is valid. Blit row 0 (CopyBits) is correct, which is why the on-disk PICT was always fine: QuickDraw resolves addressing itself. Confirmed by remedy — 32-bit addressing switched on in the Memory control panel, and captures crossed correctly at once.

Why the earlier refutation was wrong, and the lesson in it

A previous pass retired this exact hypothesis on the grounds that vprobe's fidelity sweep reported 480/480 rows matching at base 0xFC080000 (docs/vram-readout-68k.md, 2026-07-25), so the base must have been reachable. Re-run beside shotdiag three days later, the same sweep on the same machine reported 480/480 differ, 1st 0. Nothing had changed but the Memory control panel setting, which had reverted on its own — the PRAM battery is dead.

So vprobe was broken in exactly the same way as the capture, and the inference drawn from their difference ("the difference between them is the file the capture opens first") was drawn from a measurement taken in a different machine state. Two runs of one probe on one machine are not comparable unless the addressing mode is recorded beside them. vprobe now carries an Addressing row for that reason.

The fix

core/screen68.c decides, from the machine's actual state, how a raw read reaches the framebuffer:

  • 32-bit capable (Gestalt gestaltAddressingModeAttr / gestalt32BitCapable, confirmed by performing the switch once and checking low memory 0x0CB2 moved) → SwapMMUMode(true32b) around the VRAM copy, always, whichever mode the machine is currently in. The mode can change between the check and the read — the File Manager runs in between on the capture path — so the switch is not conditional on it.
  • not capable, address survives 24 bits → read it as it is.
  • not capable, address does not survive 24 bits → refuse the capture with a reason. Wrong pixels are worse than a refusal.

24-bit is the expected state of a vintage Mac, not an anomaly. Most of these machines have dead PRAM batteries and come up with 32-bit addressing off however it was left. Asking a human to set it is not a fix: it reverts on the next power cycle and reads as a regression.

The switch wraps the VRAM copy and nothing else. While switched, the machine is in an addressing mode the rest of the system was not told about, so no Toolbox or OS call may be made — and the staged capture interleaves the read with PackBits and File Manager writes. The copy goes out through a row_copy hook on N68ShotWireSink, which keeps n68_shotwire_emit() Toolbox-free (the host cc still compiles and drives it) while the dereference itself happens where the Toolbox is allowed. The hook is required: a NULL is refused rather than filled in with memcpy, because a caller that forgot would send main RAM at full speed with every test green.

StripAddress is a different question and is not the fix. Stripping 0xFC080000 is the bug, spelled deliberately. It is used as a predicate on the screen's base ("does this address survive 24-bit mode?") and as a normalisation of the offscreen band's base, which is a Memory Manager block whose top byte is master-pointer flags in 24-bit mode. It never rewrites the framebuffer address.

There are exactly two addressing modes on a Mac, 24-bit and 32-bit. There is no 16-bit mode; "16-bit" in vprobe's readout is a read WIDTH and "8-bit" beside the screen is colour DEPTH.

Unverified

Tested, not metal-verified. The fix has not run on the 180c — nobody here has a machine. Both guests cross-build, scripts/test-all is green, and the emitted 68K code contains the _SwapMMUMode trap (0xA05D) inline, so it links. What a metal pass should show, on a machine left in its default 24-bit mode:

  • shotdiagAddressing 24-bit, Raw read SwapMMUMode to 32-bit, Walk row 0 equal to Blit row 0, Verdict identical - the base is right.
  • vprobe in the same session → Addressing 24-bit, 32-bit for reads and Fidelity MATCH (480 rows), with the bandwidth rows unchanged from 2026-07-25 (a switch is two traps against passes of 150 ms).
  • A capture over the wire, decoded, showing the 180c's screen — with no visit to the Memory control panel.

If Raw read reads REFUSED - unreachable, the machine reported itself not 32-bit capable and the framebuffer is above 16 MB; that combination is believed impossible and would be the thing to report.

Two guests on one port (2026-07-28)

The host serves several guests at once, told apart by the identity in their hello (the name, trimmed and case-folded). Both PowerBooks can dial one port, and the window can be pointed at either. Tested, not metal-verified — neither machine has been near this, and the emulator has not either. Nothing below has ever run against a real classic Mac.

One guest is ACTIVE: every request-shaped call (runCommand, exec, listFiles, requestCapture, the modules, the agent projection) drives that one. What a guest gets regardless is the half it initiates — its pings, its pushes, and our share served back down its own socket.

Choosing which Mac

GuestListener.selectGuest is now reachable two ways: a pop-up in the sidebar footer, which appears only when a second machine is connected, and Guest ▸ Drive in the menu bar, rebuilt as the menu opens.

Each module model decides for itself what a switch means to it, and the decisions are not the same — the reasoning is at each Snapshot type, and the mechanism is one small cache (GuestScopedState.swift):

  • Kept per machine, because it cannot be re-fetched or is expensive to: the console scrollback, history and completions; the screenshot history; the census dossier; the software inventory; the Files breadcrumb and listing.
  • Discarded on a switch, because it goes stale on its own machine faster than a person can read it: the process table. Also every in-flight thing — stream brackets, sweeps, loads — which the listener has already failed by then.
  • Dropped rather than parked: a queue of files still waiting to be sent. They were meant for the other Mac; the module says so.

A machine that DISCONNECTS keeps its parked state (the same machine dialling back in finds its own scrollback), except the software inventory, which dies with the connection exactly as it did before — a redeployed guest has a different disk.

A pushed capture now says which Mac sent it

CaptureDelivery carries the sender's name and key, stamped in Session from the socket it arrived on. A background machine's push is filed under that machine — it no longer appears in the driven Mac's history — and the system notification names the right Mac. It is still auto-saved to the landing pad, and it does NOT take the clipboard.

No contract field was added and none was needed. Which machine sent a message is answered exactly by which socket it arrived on; a name in the payload would be a second, weaker copy of that fact. Both guests are unchanged, and neither now differs from the other.

Visible consequence, not yet addressed: a background push is announced and saved but appears in no list until you switch to that machine.

Open

  • Pending requests share one id space across guests. Ids are drawn from one host-side sequence so they cannot collide, and answers from a non-active connection are now dropped rather than settling somebody else's waiter — but the maps themselves are still flat on the listener rather than per guest. A switch fails what was in flight instead of keeping it.
  • One stream, one capture, one put, host-wide. Two guests cannot stream at once; the second is refused stream-busy. Honest, but a limit nobody chose for its own sake.
  • Nothing on screen says a background Mac is doing anything. The picker names the machines and nothing more: a push that landed under the other one, a transfer it started, an error it reported are all invisible until you switch to it. The roster is the obvious place for a badge, and it does not have one.
  • One stream, one capture, one put, host-wide (see above), and with it the reason MCP addressing is an assertion rather than a switch.

A guest is addressed by a machine id, mapped to its address

Identity used to be the folded hello.name. Two consequences, both wrong: two Macs calling themselves the same thing were ONE guest and the second was refused busy, and — because a deployed guest runs under its MacBinary name and that name carries the version — every redeploy minted a phantom machine. An identifier that changes when you deploy is not an identifier.

There are three identities now, kept apart on purpose (now-host/Sources/Host/GuestIdentity.swift):

what it is who asserts it changes when
machine id pb1400c — the handle a person or an agent types the HOST assigns it only a human rename
session id pb1400c-<uuid> — one connection the host mints it at hello every dial
address the peer IP off the NWConnection host-OBSERVED DHCP
display name NOW Guest 0.14 the GUEST asserts it every deploy

The roster pairs them: the picker and the Drive menu read pb1400c — NOW Guest 0.14, the log line adds the address, and a caller gets the id and session id together.

No contract change, and none was needed. The address is host-side knowledge, arriving on the socket; the name is already in hello. Neither guest is touched, neither now differs from the other, and docs/contract-coverage.md is unchanged because nothing about what a guest SERVES moved.

Where the id comes from. Assigned host-side and persisted host-side (GuestRegistry), anchored on the observed address plus a fingerprint (the hello's os and its name with the version stripped). The reasoning, including why Gestalt cannot supply one — no serial number; gestaltMachineType is a MODEL; gestaltSerialAttr is serial PORTS — is written out at the top of that file. First sight is guest-1, addressable with zero configuration and flagged auto-assigned; a human rename makes it pb1400c.

The rules, and what each costs. An id never silently rebinds: adoption needs address AND fingerprint to match, so a stranger inheriting a DHCP lease does not inherit pb1400c — and the same rule means a Mac whose lease changes costs the human one rename. Two machines never collapse onto one id: ordinals are unique and a rename onto a taken id is refused naming the holder. Where the address cannot tell machines apart — loopback, and therefore every emulated guest and every test — a slot completes the anchor, and the row says idIsAnchored: false rather than pretending.

The MCP surface is addressable (local protocol v7). Every tool takes an optional guest: a machine id ("whatever is connected to that Mac now", which follows a reconnection) or a session id (precise, and refused now-guest-session-ended once that connection is over rather than being answered by its successor — the same staleness contract the process and quit references already keep). now_session_health reports the driven machine's reference and the WHOLE roster, so a caller can discover the ids. Availability by capability is untouched: this decides which machine a question reaches, never what a machine can do.

Open, from this slice:

  • Addressing is an assertion, not a switch. Naming a machine the host is not driving is refused now-guest-not-addressed, with the driven machine and the roster in the message. It cannot be answered, because the request-shaped listener API drives one session at a time and every waiter map is still flat (above). Making an agent call re-point it would also take the console out from under whoever is sitting at it — a policy question, not just a plumbing one.
  • Not every projection names the guest yet. Session health, the process snapshot and the roster do. Launch, quit, artifact transfer and the Files results still carry only the session UUID. They cannot answer for the wrong machine — addressing is checked before any of them — but a caller reading one of those results alone still has to remember what it asked about.
  • The real fix is still a guest-minted id in hello. A stable id the MACHINE knows would survive a DHCP change without a rename and would tell two emulated guests apart. It is a contract change and both guests, deliberately not half-implemented here. The candidates and their failure modes — boot volume creation date (a cloned disk yields two machines with one id), a self-assigned id in the guest's own preferences (PPC preferences key off the BINARY'S name, so a side build mints a new one), PRAM (wiped every power cycle on the 180c), the Ethernet address (it belongs to a SCSI-Ethernet dongle that moves) — are recorded in GuestRegistry's header so the next attempt starts where this one stopped.
  • The address is not on the agent surface, on purpose. The host observes it and uses it internally; the companion is told the id, the session id and the display name, and nothing about where anything is. Being able to NAME a machine does not require being told its address. The human-facing halves — the app's roster and its log — do show it, because that is the human's own desk.

The README shows neither interface (2026-07-28)

Missing, not broken. There are no screenshots of either half, in a project whose entire subject is two Macintosh interfaces. A reader is being asked to take the interesting part on faith, and the README says so rather than quietly not mentioning it.

Wants: the guest's Workshop window on the classic Mac (the Files page with a real listing is the most legible single frame), and the host window from the same session, so the two images are visibly the same connection from both ends. On real hardware if possible — an emulator capture is honest, but a photograph of the PowerBook says more about what this is for.

What to capture and the rules for it (native size, nothing identifying in frame) are in images/README.md. Michelle is taking these; the row closes when they land.

The 180c, 2026-07-26: two suites metal-verified, the ladder not (0.22)

Five branches merged, deployed as NOW-68K 0.22, and run against the PowerBook. What is now metal-verified, what is not, and three defects the attempt found.

Metal-verified on 0.22

  • Metal68KTests — dial, handshake, keepalive, bounded catalog search, farewell, redial. 3 run, 2 skipped, 0 failures, 50.8 s.
  • Metal68KContractTests — 3 run, 0 failures, 72.7 s. Individually: an unimplemented message refused in 6.4 s, a second request during a confirm wait handled in 15.6 s, an oversized control frame costing one message rather than the wire in 67.3 s.

The control plane is healthy on this machine. That is the whole of what tonight added to the metal column.

Still NOT metal-verified

The file family, both directions. Metal68KPutTests never produced a usable result: contended the first time (below), killed the second when the machine was rested. Receive and send remain emulator-verified only, and the emulator's ~350 KB/s receive is a 68040's number that predicts nothing here. No NOWBASE baseline lines were captured either — neither run reached the point of emitting any.

Broken

  • The handoff cannot retire a build older than isSelf. The identity gate correctly refuses to name a process the guest has not marked as itself, and 0.19 predates ProcessListing.isSelf — so it declined to guess, and could not proceed. A one-time migration cliff created by the fix itself: the first build carrying the field has to be launched some other way. Worth deciding whether the gate should accept an explicitly-named outgoing build for this case, or whether the answer is simply "a human double-clicks once".
  • launch with a colon-bearing HFS path did not launch. Asked over the wire to launch Macintosh HD:Lab:now-68k:NOW-68K 0.22, the running 0.19 returned no reply within 40 s and the application did not start; a human launched it by hand. proc_launch_named is documented to treat a colon-bearing string as a full path and skip the catalog search — which is also the step deploy-68k --handoff depends on, so this is very likely the root of both failures rather than two. Not diagnosed; the machine is resting.
  • Metal68KContractTests was failing by SUCCEEDING. Its canary for "an unimplemented message is refused and says so" was file.list, which the browse branch implemented — so the guest answered success and the test reported a defect that was really a feature. Repointed to file.move. A test whose subject is a GAP has to be repointed every time that gap closes; picking a message nobody will ever implement is the worse alternative.

The machine set its own limits

A 1 MB push moved 606208 of 1048576 bytes with 77 progress reports at a healthy cadence, then stopped; every rung after it got 0 of N, including the empty and one-byte cases. Round-trip went from 14.4/21.4/28.4 ms idle to 39.3/266.7/439.5 ms. That run was contended by another session deploying into the same folder mid-ladder, so it is not cleanly attributable — but the shape is a silent MacTCP wedge, not a throughput limit, and a machine that is merely slow does not fail a zero-byte transfer.

Later that evening the display began to flicker and the machine was rested. The same panel failed mid-session days earlier.

The 4 MB rung exists to find protocol bugs at scale and the emulator finds those for free, while on this machine a serial multi-megabyte push is what wedged the stack. The parent corpus carries the envelope as vintage-laptop-sustained-load-envelope: ladders on the emulator, character on the metal, sessions in minutes. If the boundary is ever worth finding, the experiment holds total bytes constant and varies burst size and rest between bursts rather than climbing a size ladder.

The 68K file family's browse half (2026-07-26)

file.list / file.listing and the ls command. Additive: both messages were already in contract/asyncapi.yaml, already decoded by the host, and already served by the PowerPC guest — checked before designing, and nothing in the contract changed. NOW-68K now serves 15 inbound message types.

Broken

Nothing found in this pass. What the pass DID find is below, under unverified — most of it is about what a small frame costs.

Unverified

  • Indexed catalog cost at a deep cursor is unmeasured. PBGetCatInfo at index N on a large folder is not O(1), so a host paging into a thousand-entry folder pays more per page the further in it goes. Never measured, on either machine. If it ever needs bounding, the bound belongs in n68_fileenum.c as a wall-clock budget with an honest "truncated at the budget" answer — proc68.c's kLaunchSearchBudgetTicks is the local pattern — and NOT as a silently short page. Nothing pages a large folder today, which is the only reason this is parked.
  • Nothing has browsed the 180c. Emulator-verified only, on a Quadra 800 under Mac OS 8.1: a host lists files it just pushed, walks a twelve-file folder across several pages losing nothing and duplicating nothing, gets a file.refuse (not a timeout) for a folder that is not there, and sees the same entries through ls. That rig is a 68040 with 128 MB and a cached disk; the 180c is a 68030 with 4 MB and a real one. Metal68KBrowseTests is the gate to point at it.
  • A worst-case page carries ONE entry. This guest's outbound payload cap is 1024 bytes against the PowerPC guest's 4 KB, and an HFS name of 31 accented characters escapes to 186 bytes of \uXXXX. The arithmetic is pinned by static asserts and by test_filelist.c, so this is correct behaviour rather than a defect — but a host that assumed a page means a folder would be wrong here in a way it is not against the other guest. Never observed: no folder on either test machine has names like that.
  • A UTF-8 path does not resolve. NOW-68K has no UTF-8-to-MacRoman decoder, so a host asking for Café:Notes sends bytes this guest cannot turn into an HFS name and gets not-found. Truthful, and the same property the receive half already has (n68_putrx.c), so the two halves at least agree — but a folder a person can see in the Finder is a folder the Files module cannot open. The PowerPC guest decodes (now_json_find_text); this one needs the same table before it can. GuestWireConformanceTests.testHfsPathArgumentsAreTextDecoded does not catch it, because it checks for the wrong FUNCTION and this guest's scanner has a different name.
  • identity is absent from every entry. Deliberate — it is a precondition token for mutations this guest does not serve, and nothing in now-host/Sources reads it. It is the first field to add if file.move, file.trash or file.get ever land here, and adding it costs ~30 bytes of a 1024-byte page, which is roughly one entry.
  • Three row-array commands still answer inside now68k_commands_dispatch. help, ps and vprobe. The result type docs/command-parity.md called for now exists (N68CmdRows) and ls uses it; moving the other three is a refactor of working code that was deliberately not done in the same change as a new message family.

An abandoned transfer wedged NOW-68K against all future ones (2026-07-26)

file.cancel appeared nowhere in wire68.c's dispatch. The guest sent file.progress and handled no cancel inbound, so the question nobody had answered was what it actually did when a host walked away mid-transfer. The answer was worse than "it leaks a staging file", and the ledger entry is the finding rather than the fix.

What it did, measured before anything was changed

A fake host (a probe, not a fixture — it speaks just enough of the contract to arm a transfer and then abandon it) against 0.19 on the Quadra 800 emulator, all on ONE connection that stayed up throughout:

-> file.begin transfer 11 ... 8 KB of bulk ... file.cancel {transfer:11}
<- {"type":"error","code":"not-implemented","message":"unsupported message type"}
-> file.offer id 2
<- {"type":"file.refuse","id":2,"code":"busy","reason":"a transfer is already in flight"}
-> command.request put
<- {"ok":false,"error":{"code":"put-refused","message":"a file is arriving right now"}}

The guest answered the cancel with not-implemented and kept holding the transfer. Every later transfer, in either direction, was refused for the life of the connection — the lane is one transfer wide and shared across both — and pings were answered normally the whole time, so from the host's side the guest looked healthy and simply refused to move a byte ever again.

Why nothing rescued it

  • There is no transfer timeout, and there is no message for "I have lost interest". An abandoned transfer is indistinguishable from a slow one, and neither n68_putrx nor n68_puttx carries a clock.
  • The only clock in reach is service_live()'s 65 s no-traffic watchdog, and it is the wrong one. It is a property of the CONNECTION — kWireDeadTicks since the last inbound byte — and the guest's own 30 s keepalive ping keeps being answered, so on a live connection it never fires. A DROPPED connection was always fine (reset_read_state cancels both directions, which closes the outbound fork and deletes the staging file); the case nobody had established is a host that stays connected and stops caring.
  • The receive half held its staging file (NOW incoming <hex>) open for a transfer that would never end. Observed as the wedge; the orphan on the Desktop follows from the staging file never being discarded and was not separately confirmed on the baseline disk.

The send half had a second door into the same wedge

Found on the way. The host sends file.cancel and file.done together the moment its sink fails (GuestListener.swift :: failInboundStream), and n68_puttx_done() acted only in kN68SendEnded — so a file.done arriving while bytes were still going out was dropped, the guest streamed the rest of the file at a host that had already discarded it, and then parked in kN68SendEnded waiting for a reply that had already been and gone. The host does not send a second one: finishFile returns early for a transfer it is discarding. A receiver's file.done is final whenever it arrives; requiring our own file.end first is what made the park permanent.

Fixed, and what the fix is verified to do

No contract change was needed — FileCancel and file.done's cancelled code were already there, which is worth recording because the gap was entirely on the implementing side. file.begin's transfer is now remembered, because file.cancel names a transfer and carries no id, so nothing else could tell a live cancel from a late one.

Same probe, same emulator, 0.20:

Probe step Result
cancel a push after 8 KB file.done ok:false code:cancelled received:8192 cleanup:temp-discarded
offer again immediately accepted, completed, CRC-confirmed
cancel the guest's own send mid-stream file.end ok:false, 0 bulk frames after the cancel
ask for that send again offered again — the lane is free

cleanup:temp-discarded was checked against the disk rather than believed: hls on the session image afterwards shows the completed After Cancel and no NOW incoming staging file. (xfer_tmp_1 in that listing predates this work by weeks and is base-image debris.)

The deliverability claim in n68_puttx.h rule 3 held up under the one case it exists for: the cancel was acted on one chunk after it arrived, not at the end of the transfer. A staged bulk frame nobody has seen is dropped; one already part-way out finishes, because a frame cut short is a desynchronised wire rather than a cancelled transfer.

Still open

  • Not on the 180c. Emulator-verified only, and the emulator is a 68040 with 128 MB. The behaviour under test is a state machine rather than a rate, so it should carry — but nobody has watched it.
  • ~~A cancel has no console face.~~ Closed in the same pass. It is a cancel verb now — contract's x-commands first, then commands68.c, which the console reaches through now68k_commands_run without conwin.c gaining a second dispatch, so both faces run one implementation and help lists it. Verified on the emulator from the console face specifically: help shows the row, a quiet machine answers nothing-to-cancel rather than pretending, and the verb produces the same file.done ok:false code:cancelled cleanup:temp-discarded the wire message does. The PowerPC guest deliberately gains no verb — a host cancels it from the Files UI and a person at that guest from its own Workshop — and that decision is named with its reason in CommandRegistryTests.notOnThePowerPCGuest rather than left as a silent gap.
  • The other 65 s window is unexamined. A host that abandons a transfer AND stops answering pings is cleaned up by the watchdog, but no one has watched that path either, and it is the only path in which a transfer's cleanup depends on a timer.
  • The probe lives in a scratchpad, not the repository. Turning it into a metal gate belongs with whoever is working on that harness; it needs requireTheBuildUnderTest() before anything it reports can be believed.

front, on both faces of both guests (2026-07-26)

process.front had been on the PowerPC guest's wire since the Processes module was built, and there was no way to type it — not at either guest's own keyboard, and not from the host console, which is a dumb shell that knows no message families. A capability reachable only by clicking a button in one module is the ps shape exactly (command-parity.md).

So front is now a contract x-command served by both guests, over the same list → match → re-validate → act → re-check composition quit uses, and NOW-68K additionally answers the process.front drive verb it did not before. Its outcomes are deliberately not quit's with the words changed: not-running is ok:false here (nothing can bring forward a process that is not there, where quit was asked to produce exactly that state), and NOW itself is a fair target (fronting severs nothing; quitting would cut the reply mid-send).

Unverified

  • The confirm branch has never run. SetFrontProcess returning noErr means the switch was scheduled; it lands when the guest yields, and both guests yield with an event mask of zero. Whether a process switch completes inside that yield is unproven on either machine — if it does not, front will report unconfirmed every time while the screen plainly shows the switch happened. That is visible and diagnosable rather than a silent lie, which is why it is written this way, but it is the first thing to watch on metal.
  • Nothing else here has been on a machine either: both guests build clean, the host suite is green, and no PowerBook has run it.

Open

  • front's argument parser is not natively testable. quit's grammar lives Toolbox-free in proc_quit_args.c and has its own native test; front's is four lines of trim-and-unquote, static in each guest's command file, and duplicated across the two. It is small enough that a shared module would be more moving parts than it saves — but it is the second copy of a grammar, which is how the first one started.

quit targets a process identity, not a file name (2026-07-26)

The handoff's retire step named the outgoing build "NOW-68K " + <the version it reported in hello> — a FILE NAME derived from a compiled constant. They agree by convention only. On 2026-07-25 a build deployed as NOW-68K 0.18 reported 0.16, so the retire sent quit NOW-68K 0.16, the guest answered honestly that nothing of that name was running, the old build kept running, and a 4 MB machine was left with two NOW-68Ks. Nothing was broken on the guest; the identifier was invented on the host.

Fixed by naming the target the way a machine should:

  • process.listing gained isSelf (contract first), set by both guests on their own row. It is the only trustworthy answer to "which process is on the other end of this connection".
  • NOW-68K now answers the contract's process.quit drive verb — re-validate the PSN, refuse self, send — over proc_quit_psn, the same three steps proc_quit_named ends with. It does not confirm, and that is the contract's decision: process.result.ok means DELIVERED, and there is no field that could tell a granted quit from a declined one. A caller confirms by re-reading process.list.
  • The quit command still takes a name, because a person types what ps shows them. ps now says self on that row, on both guests.

Unverified

  • None of it has been on a machine. Both guests build clean and the host suite is green (511 tests), but isSelf, process.quit on NOW-68K, and the PSN-targeted handoff have not run on the 180c. The loopback test HandoffIdentityTests reproduces the version/name disagreement over scripted guests and watches the old derivation fail — that proves this side never invents an identifier, and proves nothing about the Toolbox code.
  • The first handoff has to be launched by hand. The build currently on the 180c is 0.19: it serves process.list without isSelf and does not answer process.quit at all. Handoff68K.identifySelf fails with a message saying so rather than falling back to a name — a fallback would be the defect, reintroduced. From 0.20 onward the handoff is automatic again.
  • process.front and process.shot are still unimplemented on NOW-68K. They fall through to send_error_reply, visibly. Only the verb the handoff needed was added; the family is deliberately partial rather than quietly half-served.

Open

  • The host's ProcessEntry.id is name#code#creator, so two processes of the same name collide in the table's identity — exactly the case isSelf and the PSN exist to handle, one layer up. Not hit by anything today; worth the PSN when it is.

The 68K file family, both directions in one tree (2026-07-25 night)

Three branches merged and verified together: the receive half (MacBinary, Desktop landing, the FSClose fork repair), the send half (the byte-source sender), and the version-bump commit that carried them to the machine. What that merge found, and what it left open.

Broken

  • ~~The two halves disagreed about where files live.~~ Fixed in this pass, and worth keeping in the ledger for how it hid. Receiving landed on the Desktop, sending read from the application's own folder; each branch was self-consistent, so no reviewer of either could see it. It survived the merge (no textual conflict — two roots in two files), 27 native tests, 508 host tests, both Xcode configs, and -Werror. The round-trip ladder on the emulator named it as fnfErr on all ten rungs. A cross-direction test is the only kind that could have caught this, and it could not exist while the halves were on separate branches. now68k_desktop_folder is now published from n68_putfile.h and both directions read it.
  • A merge can drop an #include with no conflict. git took one side's include block wholesale and <Processes.h> went with it. The block was never marked conflicted, so reviewing the conflicted hunks would not have shown it. -Werror caught it; nothing else would have until link time.
  • A conflict region can cut a function mid-body. The resolution looked complete — every declaration present — and the function simply never closed, which the compiler reported as four unrelated functions being "defined but not used" and a fifth reaching the end of a non-void function. The error names never mention the function that is actually broken.
  • The handoff's retire step may quit the wrong build. NOW-68K 0.17 reached the 180c and its log reads wire: connected then cmd: quit ok 0 — the incoming build took a quit and executed it, where the outgoing one was meant to. Not diagnosed, and not confirmed: the run it came from was contended (see below), so this is a suspicion with a log line behind it, not a defect with a repro.

Unverified

  • Neither direction has moved a byte on the 180c. Both are emulator-verified on a Quadra 800 under Mac OS 8.1 — receive 4 MB in 11.7 s (350 KB/s, 512 progress reports, CRC-confirmed), send 4 MB in 1.8 s, MacBinary both forks, control lane 0.05 s idle against 0.10 s during a 1 MB push. A 68040 with 128 MB is not a 68030 with 4 MB, and the send rate in particular reads off a disk the emulator caches — read it as "the path works", never as a rate.
  • The PackBits ratio and encode cost are unmeasured. vprobe has the framebuffer READ at 159 ms for a 300 KB frame (docs/vram-readout-68k.md); nobody has measured what compressing it costs on a 33 MHz 68030, and the ratio is what decides whether screenshots are viable over MacTCP at all. No branch in this repository implements PackBits. The send half was built as a byte source (n68_bytesrc.h) precisely so a capture can feed the pipe in bands rather than buffer 300 KB against a 384 KB partition — that shape held through the merge, so a screenshot sender does not need a second send path.

Two host sessions can contend for one PowerBook, invisibly

A metal run of these suites on 2026-07-25 held port 5252 for the better part of an hour while another session deployed a build into the same folder mid-ladder. The results were unattributable: a 1 MB push stalled at 606208 bytes and every rung after it timed out at 0 of N. The most likely cause is contention rather than a defect — NetPresenz serving an FTP upload while NOW-68K received a push, both on MacTCP, on a 68030 with 4 MB — but nothing proves that either, which is the point.

requireTheBuildUnderTest() would not have caught it. That guard asks whether the connected guest is the right guest, and it was. The gap is that nothing establishes whether the machine is already busy. lsof -iTCP:<port> before a run answers it in a second. The existing rule covers several guests reaching one listener; this is several listeners reaching one guest, and it is not written down anywhere else.

Fixed on 2026-07-26, test-side only — see MetalMachineGuard and 68k-metal-runbook.md. Before any 68K metal suite binds, it establishes that nothing else on this Mac holds the port and (when NOW_METAL_MACHINE names the guest's address) that nothing else is talking to the machine, and fails in about a second naming the process rather than producing an unattributable result. It also reports a bind failure as a bind failure: the suites used to wait out a full 120 s and say "no guest dialled in", which aims the diagnosis at the Macintosh for a fault entirely on this side.

What it still cannot see is below, and it is the honest limit of the fix.

A host cannot ask a 68K guest whether it is busy

Found 2026-07-26 while building the guard above; no code changed.

NOW-68K knows perfectly well whether it is mid-transfer in either direction, and renders exactly that: xfer reports an active receive with its byte count, an active send, and the last completed one either way. There is no way for a host to ask. xfer is console-only by a recorded decision (CommandParityTests :: consoleOnly — "renders the file.* family's state; the host reads it from file.progress and file.done instead"), and the PowerPC guest's wire-only putstat has no 68K counterpart.

That reasoning holds for a host that is driving the transfer, which is the case it was written for: such a host has the progress messages. It does not hold for a host that wants to know whether the machine is free before it starts — which is precisely the question the contended run needed to ask and could not. So contention detection is host-side only, and the guard says so rather than guessing.

Not fixed here, because it is a product change and this pass was tests and documentation. If it is taken up, the cheap version is a busy verb (or an xfer promoted to both faces per command-parity.md) answering the two booleans and the two byte counts N68PutStatus / N68SendStatus already hold. The gap it would close is real but narrow: it tells a second session that the machine is busy, and tells it nothing about who has it.

Related and unresolved: the name a build has on the disk and the version it reports on the wire are established by different means, and a guest answering "version":"0.16" was found on a machine whose deploy folder had just gained a file named NOW-68K 0.18. Whether those were the same application was never established. deploy-68k stamps both from one source, so this only arises when something bypasses it.

--filter Metal68K used to report a failure that meant nothing

Fixed 2026-07-26, test-side only. Metal68KHandoffTests is a deploy step, not a coverage gate: it needs a freshly uploaded build and the exact HFS path of it, which only scripts/deploy-68k --handoff knows. It used to FAIL when those were absent, on the argument that asking for a metal run with no build to hand off to is a broken invocation — sound in isolation, but --filter Metal68K catches that class too, so every ordinary 68K metal pass reported one red that meant nothing. A red that always fires is a red nobody reads.

It now SKIPS when NOW_68K_NEW_APP is unset, with the same second opt-in shape MetalQuitTests already uses for the dirty-document case ("NOW_METAL=1 alone does not say a human is at the keyboard"). Set but EMPTY is still a failure, because that is somebody having tried.

This is a deliberate exception to "a metal gate fails rather than skips", and it is narrow: the thing being skipped is a deploy action, not evidence about the guest.

NOW-68K cannot send the same file twice, and says it can

Found 2026-07-26 on the emulator, by the repeat sampling above; no code changed.

n68_puttx.c's offer never sets overwrite, so the host applies the contract's default of false (GuestListener.acceptOffer) and REFUSES the second offer of a name the share already holds. That is defensible policy — the host will not silently replace a file — but the guest end of it is not honest:

  • put <name> has ALREADY answered ok by then, because the command returns as soon as the offer is away (deliberately: a command that blocked for a multi-megabyte transfer would hold a command.result for minutes). So the person who typed it is told it worked.
  • The refusal arrives afterwards as file.refuse, and the only place it surfaces is xfer's "last FAILED" line — which nobody has a reason to type after being told ok.

From the host side it presents as a transfer that never starts: the offer goes out and nothing ever arrives. It cost a 300 s timeout per sample to work out, and read exactly like the machine having gone away — the same signature as the contended run, from an entirely different cause, which is worth knowing on its own.

Not fixed here (product change; this pass was tests and docs). Three candidate fixes and they are not equivalent: the guest could set overwrite on its offer (wrong — that hands a guest the right to replace files on the host), the host could decline more visibly, or the guest could hold the send's outcome somewhere put's caller can reach. The last is the one that matches the direction the contract already takes for progress.

The suites work around it by naming every sample separately (RT<size>r<rep>), which is a harness fix and not a fix.

Two swift test runs on one Mac fail three suites

Found 2026-07-26, reproduced deterministically; pre-existing, no code changed. Three suites share state outside the process: HostLogTests and LoggingSpecTests both write ~/Library/Logs/now-logs, and HostAppStateWiringTests binds a fixed port 52981. Running two swift test processes concurrently fails them every time.

It surfaced here because a metal pass and an ordinary gate run overlapped by a few seconds, and the result was two failures that vanished on re-run — the flakiness signature, from a cause that is not flaky at all. Worth fixing at some point (a per-process log path and an ephemeral port), and worth knowing meanwhile: the runbook says one at a time.

What a 68K metal run should record

68k-metal-baseline.md. In short: the suites now emit one greppable NOWBASE line per measurement, carrying the conditions (build, machine, port) beside the numbers, because the 2026-07-25 run's numbers were real and unattributable. NOW_METAL_REPEATS=3 takes three samples of every rung at or above 1 MB, so a rate can be told from an interruption — which one sample from this machine demonstrably cannot do.

Host -> guest file transfer on NOW-68K (2026-07-25)

NOW-68K receives a pushed file. Offer, accept, stream, checksum, done - the contract's hostPutsFiles sequence, served by the guest that previously discarded every bulk frame to stay in frame sync.

Emulator-verified, NOT metal-verified. Everything below was measured on a Quadra 800 under Mac OS 8.1 with 128 MB (scripts/q800-68k). The real target is a 68030 under System 7.1 with 4 MB. What carries over is correctness; what does not is every number in the table.

Size Result (emulator)
0, 1, 8191, 8192, 8193 B ok - the boundaries either side of one frame
64 KB ok, 299 KB/s
256 KB ok, 348 KB/s
1 MB ok, 357 KB/s
4 MB ok, 11.6 s, 352 KB/s, 512 progress reports

The 4 MB file was pulled back off the disk image with hfsutils and is byte-identical to what was sent (CRC-32 A627E416, agreeing with zlib and with the guest's own). The catalog shows exact sizes rather than allocation-block-rounded ones, so the Allocate + EOF-trim pair works, and no NOW incoming ... staging file was left behind.

The guest's event loop is not starved by the receive path: help round-tripped in 0.05 s during a 1 MB transfer against 0.06 s idle.

MacBinary on NOW-68K: the fork corruption, found and fenced (2026-07-26)

Full record: 68k-file-receive.md. In short.

FSClose of a written resource fork on Mac OS 8.1 splices 77 bytes of File Manager catalog state into the fork's first block at offset 48 - an in-memory record layout that matches nothing on disk. Deterministic: every resource-carrying MacBinary file, every run; data forks never; a MacBinary file with an empty data fork never affected.

Pinned by three structural facts, in this order: the splice is sub-sector, so the bytes were wrong in RAM and no allocation-level theory survives; the spliced content carries the staging name and BINA but both FINAL fork lengths, which brackets the write to the close window and explains why disabling Allocate, SetEOF, FSpRename, FSpSetFInfo and PBSetCatInfo each missed; and read-back probes read clean before the close and spliced after it, 5/5.

The guest keeps the fork's first 512 bytes as written, re-reads them after the close and after the rename, rewrites them when they diverge, and re-verifies through a fresh open. Unrepairable before rename fails the transfer; after rename it deletes the file rather than leave a corrupt application to be double-clicked. Detected 5/5, repaired in one round 5/5, raw disk clean, both forks byte-identical.

Still open, and the first is the one that matters:

  1. System 7.1 on the real 180c is untested - the 7.5.3 image in the lab has no MacTCP, so the OS discriminator is blocked. The shipped probes double as the experiment: push one MacBinary file and the log either names the scribble or stays silent. Either answer is safe.
  2. QEMU's contribution is not separated from 8.1 itself.
  3. The PowerPC guest's resource forks have never been byte-verified.

What is deliberately not there

  • No resume. The guest never reports have, which the contract reads as "start from the beginning". Partials are always discarded. Deliberate: resume is an open hang on the PowerPC side (see the large transfer notes) and a 4 MB transfer is not long enough to make restarting a hardship.
  • ~~Receive only.~~ Superseded — NOW-68K now sends as well as receives; see the next section. file.list, file.move, file.trash and the host-initiated PULL (file.get) still answer the generic not-implemented error, so the guest can push a file it is told to push but cannot serve a host that wants to browse or fetch.
  • ~~The destination is the application's own folder.~~ Superseded — files land on the Desktop, and nothing is gated. NOW-68K has no preferences and no share root, so there is nothing to read a destination out of. The Desktop needs no state to name and is where a person looks for something that arrived. path from the offer still resolves relative to it and a host may reach a subfolder — deliberate until the browse/ls verbs exist, because a boundary drawn before there is anything to browse is a guess dressed as a policy. It was briefly the application's own folder, which meant a host could write into the System Folder. The send half still reads its SOURCE from the application's own folder (now68k_app_folder), which is a different root for a different direction and deliberately so.

Open

  1. Nothing has run on the PowerBook 180c. Everything above is an emulator result. The 180c has 4 MB against the emulator's 128, a 68030 against a 68040, and MacTCP that has already been observed to wedge silently on that machine. A 4 MB transfer into a 384 KB partition is exactly the shape that behaves differently there.
  2. A contract gap: FileRefuse.code has no value for "this receiver cannot handle that". An unrecognized container is reported as io-error with the truth only in reason, which is a lie of category - nothing failed, the request was never serviceable. The honest fix is an additive enum value in the contract, which touches both halves and was out of scope for a spike. (Unknown containers are now REFUSED rather than treated as data; writing an unknown envelope out as a raw fork produces a file of the wrong length and the wrong shape and blames the disk.)
  3. FileOffer.modified has no stated units in the contract. Both guests treat it as Mac-epoch seconds (it goes straight into ioFlMdDat), and the two agreeing is the only reason it works. It should be written down.
  4. The host never uses the chunk it negotiates. hello.chunk is computed and echoed (GuestListener.swift), but the file sender's frame size is a hardcoded 8192. NOW-68K advertises 4096 for a stated MacTCP reason and is sent 8 KB frames regardless. Harmless today - the guest streams and needs no frame-sized buffer - but the negotiation is decorative, and a guest that genuinely could not take 8 KB would have no way to say so.
  5. The application partition is getting tight. This pass cost +19408 bytes (~5% of 384 KB), leaving roughly 184 KB of image before stack and heap. Preferred == minimum on a 4 MB machine, so there is nothing to borrow. The next addition this size needs the budget looked at rather than assumed.
  6. g_sink is still 256 bytes, inherited from when it was a pure discard sink. A 4 MB transfer therefore makes ~32 passes per 8 KB frame. Cheap memcpys, but it is a knob nobody has measured on the 180c, where BSS is the scarcer resource.

Guest -> host file transfer on NOW-68K (2026-07-25)

The other direction. NOW-68K makes the offer, streams the bulk frames and closes with a checksum: file.offer -> file.accept -> file.begin -> bulk -> file.end -> file.done. Additive — no contract schema changed. Every message and every field already existed, is already served by the host (GuestListener.onAcceptOffer / finishInbound) and is already sent by the PowerPC guest; that was verified against the schemas rather than assumed.

Emulator-verified, NOT metal-verified. Measured on a Quadra 800 under Mac OS 8.1 with 128 MB (scripts/q800-68k), driven by Metal68KSendTests. The real target is a 68030 under System 7.1 with 4 MB. What carries over is correctness; what does not is every number.

Each case pushes a known pattern to the guest, asks the guest to send that same file back, and compares the bytes the host still holds against the bytes that came back. Nothing in the comparison comes from the guest's own accounting — not its progress, not its CRC, not its byte count, because a sender marking its own work proves nothing.

Size Result (emulator)
0, 1, 4095, 4096, 4097, 8192 B ok — the boundaries either side of one chunk
64 KB ok, 313 KB/s
256 KB ok, 1227 KB/s
1 MB ok, 2198 KB/s
4 MB ok, 2.5 s, 1648 KB/s, byte-identical

Sending reads from a disk the emulator caches, so these rates are several times the receive direction's 352 KB/s and mean nothing about the 180c, where the read is a real one off a real disk.

The wire-sharing rule holds under real back-pressure, which is the claim nothing off-metal can check: during a 4 MB send, 28 help requests were answered, none dropped, worst 0.10 s. That is the rule working — control drains before bulk, and a reply waits for the chunk in flight rather than for the transfer.

The 0-byte case is worth its row: a zero-length source sends no bulk frame at all, begin then end, and the receiver closes out correctly rather than waiting for a stream that never comes.

It is a byte-source sender, not a file sender

The point of the shape, stated because it is the thing most likely to be undone by someone in a hurry. The source is an interface (n68_bytesrc.h) with four promises: it knows its length before the first fill, it returns promptly, it does not allocate, and it never touches the wire. A file (n68_filesrc.h) is the FIRST implementation, not the only intended one — a screen capture is ~300 KB against a 384 KB partition, so it can never be a buffer and cannot be staged to a disk that may not have room. If a screenshot ever needs a second, parallel send path, this was built wrong.

How bulk and control share the wire

Stated once in n68_puttx.h and enforced in flush_outbound():

  1. Bulk never touches the four 1024-byte control slots; it has one dedicated 4104-byte slot, so no volume of bulk can consume the slot a command.result needs.
  2. A frame already being handed to net_queue_send finishes before any other frame's first byte.
  3. Otherwise control drains before bulk, so a reply queued mid-transfer waits for the chunk in flight (~12 ms) and never for the transfer.
  4. Back-pressure is net_queue_send's short accept and nothing else.

Rule 4 is the one that will read as a bug later, so: this side does not need the receiver's file.progress to clock itself and the host does. MacTCP's staging buffer is small and reports truthfully and synchronously how much room is left; the host writes into Network.framework, which accepts essentially unbounded writes and so tells it nothing. Two senders, two mechanisms, neither one the other's bug.

Parity: put is on both faces here, and console-only on the PPC guest

Deliberate, and the two guests genuinely differ. A host driving the PowerPC guest reaches the same capability through file.list and file.get, so that guest needs no verb. NOW-68K is the machine whose display has already failed mid-session and whose host console is a dumb shell with no knowledge of message families — so on that guest the capability is a verb in commands68.c's table, reachable from both faces through one implementation. The contract declares put in x-commands first, per AGENTS.md.

CommandRegistryTests had to learn this: it assumed the registry IS the PowerPC guest's command set. NOW-68K always answered a strict subset, which that test never had to notice; put is the first command going the other way, and it is named in notOnThePowerPCGuest with its reason rather than subtracted silently.

Open

  1. Nothing has sent a byte on the 180c. The emulator results above say the code is correct; they say nothing about a 68030 with 4 MB, whose MacTCP has already been observed to wedge silently. Reading a 4 MB file off a real disk while streaming it is exactly the shape that behaves differently there.
  2. Several guests can reach one listener, and one of them is not yours. Every QEMU guest on this Mac sees the host as 10.0.2.2 under user-mode networking, so any session's VM can answer any session's listener. This cost real time: the first run of Metal68KSendTests reported unknown-command for put from a guest that was simply another branch's build — and the refusal test PASSED against it, because "unknown command" is also a refusal with a reason. requireTheBuildUnderTest() now asks the connected guest whether help lists put before believing anything it says. Run with NOW_METAL_PORT set to something nothing else is dialling.
  3. No MacBinary, so no application and no resource fork. The data fork only — which is the contract's own default for a both-forks file, so it is legal rather than a shortcut, but it means a file whose content lives in its resource fork (most classic Mac applications, every ResEdit document) arrives empty or meaningless. This is the natural SECOND implementation of N68ByteSourceOps — header, data fork, padding, resource fork, each in bands — and writing it is the real test of whether the interface earns its keep.
  4. The source is limited to the application's own folder. Same weakness the receive half has and for the same reason: NOW-68K has no share root. put takes a leaf name, not a path.
  5. One transfer at a time is enforced across both directions, and the answer to a second request is busy with a reason. Not exercised against a host that offers a push while a send is in flight — the check exists and has never been raced.
  6. A contract gap, the mirror of the receive half's. FileEnd has no way to say "the sender's own source let it down", so kN68SendSourceFailed, kN68SendShort and kN68SendLong all render as io-error with the truth only in the reason. Same honest fix: an additive enum value.
  7. The contract's operations index is asymmetric about this direction, and was before this change. guestServesFiles lists file.begin/end/progress/refuse/listing but not file.offer; file.accept and file.done appear in no operation at all, in either direction, though both sides send both. The schemas are complete and correct — this is the index above them. Left alone rather than fixed in passing: it is a contract edit that touches how both halves are described and deserves its own pass.
  8. sendMs is computed from TickCount at 1/60 s. Fine for a figure the contract types as advisory, but it is not milliseconds measured, it is ticks scaled.

Two defects found by this pass, in code that predates it

  • GuestWireConformanceTests could not see bye. Its C-literal scanner did not understand character literals, so the three '"' in wire68.c's read_string_field inverted its quote parity and every literal after them was read inside-out. bye had been piecemeal since the day it was written and never appeared in the cannot-check set - the set read complete and was not, inside the mechanism built to prevent exactly that. Fixed, and bye has the fixture it should always have had.
  • The console printed one line's tail on the end of the next. now68k_fmt_append_* do not NUL-terminate, and every builder in conwin.c declares char line[80] inside its loop, so a line shorter than the one before it trailed that one's tail: "files land in Startup Items" rendered as "files land in Startup Itemsbytes". show_help and show_processes had it latent and only escaped because their lines happen to grow rather than shrink. All emission now goes through con_out_built, which terminates. No native test could have caught this - it is pixels, and it was found by looking at a screen.

Deferred by decision

NOW agent-integration V0 is complete (2026-07-24). All five bounded projections are implemented, tested, and covered by one combined PowerBook acceptance receipt: now_session_health, now_list_processes, exact safe launch, revalidated cooperative quit, and receipt-only approved artifact transfer through the existing put lane. The pass also observed typed host absence, automatic guest redial, new session identity, stale-reference refusal, and unchanged ordinary Files/Connection UI. Exact evidence and limits live in agent-integration.md.

The artifact pass found one compatibility defect and closed it before V0 closeout: modern classic-epoch dates saturated the deployed guest's signed 32-bit JSON reader and stamped January 1972. Host→guest lanes now omit an optional date outside the deployed reader's range; the numeric guard, wire omission, mutation failure, and corrected live listing are recorded in files.md. The first disposable evidence file retains its bad stamp; no destructive cleanup was attempted.

The companion still has no guest component, lifecycle control, raw-path input, shell or general filesystem surface, force-quit surface, or CodeKitten/shared-transport dependency. Sustained load, destination-byte read-back, and any shared-transport extraction remain outside V0 rather than hidden completion claims.

Guest-initiated change controls. The browser on the classic side can list, navigate and pull, but offers no rename, delete, new folder or move. Michelle punted this 2026-07-20: write and overwrite were the goals of the slice and both work.

Worth knowing before anyone reopens it: file.move, file.trash, file.restore and file.mkdir already exist in the contract, the guest already SERVES all four, and HostShare learned to serve them too (2026-07-20, 13 tests). So the wire and both servers are done and the only missing piece is guest UI — plus a decision about undo, which the host side keeps on whoever initiated the action. That host-side implementation currently has no client. It is tested and symmetric, and it is also unused code until this is picked up; anyone auditing for dead weight should know it was built deliberately, not left over.

In flight elsewhere

The unified Workshop landed on claude/guest-workshop-unified-a3aab9 (2026-07-21): one window, a hand-drawn sidebar rail, and all four modules (Screenshots, Files, Console, Connection) behind the WorkshopModuleOps contract. The five old windows and the Connection dialog are deleted, and all four pages were watched working on the PowerBook the same night. The codex branch codex/guest-console-invert is abandoned by decision (Michelle, 2026-07-21) — do not merge it. Its one still-valuable idea, the async OT connect path (160ed85), was reimplemented against claude/processes-module-cb2d9c on 2026-07-21 (see "An unreachable host presents as a hang" below); the branch itself stays abandoned.

The Processes page landed and is metal-verified (2026-07-21, main at 22f129a; spec in processes-and-peek.md). The fifth Workshop module: a split view with a Data Browser process list (icon-and-text column, header sort) on the left and a detail pane (kind, type/creator, memory text + partition bar, launch date) on the right, plus Bring to Front and Ask to Quit (confirm -> quit Apple Event -> keep the PSN until the walk proves the process gone -> (no reply) after 10 s). The peek.h seam ships answering "NOW Extension not installed"; the group box renders it. Watched working on the PowerBook the same day. This is rung 0 of the extension ladder - everything above it (the NOW Extension itself, process.*/peek.* wire families, the semantic mirror) is still ahead.

The detour that dominated the arc was NOT the Processes page - it was reaching the Connection settings to repoint a chip that was listening on the wrong port. That exposed two real, now-fixed defects, both metal-verified: the async-connect launch wedge, and Connection field editing (see below). The Processes page itself was good across those rounds.

NOW Extension M0 is metal-verified (2026-07-21, rung 1, ext/). The guest's first resident code: a 68K INIT publishing the shared table, registering Gestalt 'NWex', and chaining a jGNE heartbeat filter. Booted on the PB1400c at 9.1; the app's now_peek_status() probed it and the Processes group box read "NOW Extension active." That proves install, DetachResource residency, Gestalt registration, table validation across the compiler boundary, and a live jGNE chain - the whole of M0. Size ~48 KB (Retro68 flat runtime), loads at 9.1; 8.6 loader ceiling still unprobed (waits on the 3400c). The recovery drill (Shift-boot off, drag-out) and the QEMU-clone pre-check remain good practice for the next resident change but M0 itself is done.

Rung 2a is metal-verified (2026-07-21) - the anchor plane and the first foreign-memory read. The extension's jGNE filter, once the Processes page arms the plane, records each process's low-memory CurrentA5/WindowList/MenuList into A5-keyed slots. Clicking a process in NOW's list reads THAT process's front-window global bounds (the per-process axtree behaviour): PSN -> partition (GetProcessInformation) -> the fresh anchor whose A5 lies in that partition (the correlation, validated by containment) -> strucRgn -> rgnBBox, every foreign pointer checked inside the partition OR the system heap before it is dereferenced (peek_validate.c, native-tested + mutation-checked), byte reads at fixed classic offsets. Watched on the PB1400c: NOW's own window read correct, and Finder read "516 x 557 at (7, 25)" - a real other process's window - once the validation was widened to accept the system heap (partition-only read "unreadable", exactly tbt's axtree lesson). The foreign read lives in the app, never the extension.

Known texture, not a defect and not fixable: the readout is only as fresh as the target process's last event-loop pass. Window state is a SNAPSHOT captured by the filter when the process pumps - classic Mac OS has no cross-process live window feed (axtree had the identical limit), so no reader can re-take it on demand. There is deliberately no time-based freshness gate on WHETHER to read: the A5-in-partition match and the fail-closed validation, not a clock, prove a slot is this process's, and the app carries the last good read across a stale blip. But staleness is surfaced HONESTLY (the AXPeek/qdpeek discipline, which hit this same wall): the reader reports the anchor's capture tick, and the detail's Windows header shows "as of a moment ago" / "as of N min ago" once the snapshot ages past ~3 s - an actively-pumping app stays live with no marker. An app that never pumped since arming reads "no anchor yet" until it does; an app with no windows reads "none open". Still open for a later pass: whether any app keeps its window structures in a zone neither the partition nor the system heap covers (would read "unreadable"); and rung 2b, cropping the actual Front & Capture to the rect.

Rung 2b - Front & Capture is metal-verified (2026-07-21). The first USE of the window bounds, and the anchor plane's first real artifact: a "Front & Capture" button in the NOW Extension group box brings the selected process forward, DEFERS the capture to a later idle (~0.75 s, so nothing nests an event loop - the main loop's WaitNextEvent yields let the target come forward and redraw), reads the now-front window's fresh bounds, crops the capture to them (capture_screen_rect - one blit, clamped to the screen), saves a PICT to the Desktop, and restores NOW. Watched on the PB1400c: a captured window PICT, well-formed 8-bit with its CLUT and PackBits rows, opened as the real window. So the whole rung proves out end to end: extension captures anchors -> app validates and reads bounds -> app crops a genuine screenshot to them.

Rung 3 - the process.* wire family is metal-verified, host Processes module metal-verified (2026-07-21). The contract gained process.list/process.listing (symmetric, paginated by a 1-based cursor, entries capped at 24 a page). The guest serves its own Process Manager walk on request (serve_process_list in wire.c: name, kind of application/background/finder, code/creator 4CCs, sizeKB, front). The host answers the mirror direction with its own running apps (HostProcesses off NSWorkspace - the degraded plane: modern macOS gives no OSType code/creator and no classic partition size, so those fields are honestly absent), and can ASK via GuestListener.listProcesses. Tested here: a byte-accurate guest fixture (multi-snprintf, so the conformance check names it as needing one), a process.list/.listing round-trip, and the conformance known-partial set. A NOW_METAL test (MetalProcessTests) pages the real PowerBook's process table onto the host and prints it; run on the PB1400c (2026-07-21) it read 8 processes correctly classified - the appe faceless-background set (Control Strip, Folder Actions, ORiNOCO Monitor, tbt-appe), the Finder by FNDR, three APPLs, and the guest itself flagged front. The host now DISPLAYS it: a read-only Processes module (ProcessesModel/ProcessesModuleView) that pages the whole table in on refresh, groups it into Applications (with the Finder) and Background, flags the front process, captions each row with kind/4CCs/partition size, and reads as the snapshot it is ("as of HH:MM:SS"). Metal-verified on the PB1400c: the pane drew the machine's 7 processes correctly grouped and flagged.

The one-way direction is by design, not a gap. NOW drives old-from- new - the host is the cockpit, the guest the operated machine - so host-sees-guest is the product and guest-sees-host is a non-goal. The guest issues no verbs at the host and has no ASK/UI for the host's processes, on purpose. The wire family stays symmetric in MEANING, but the host serves nothing back: the dead HostProcesses/NSWorkspace serve was removed rather than kept as ballast (2026-07-22).

Drive verbs added (2026-07-22). The Processes pane grew three actions on the selected row, all host->guest: Bring to Front (process.front -> SetFrontProcess), Ask to Quit (process.quit -> a 'quit' Apple Event it may decline), and Screenshot App. Each names its target by the PSN the listing now carries (psnHigh/psnLow); the guest re-validates the PSN against a live process before acting, and refuses a quit of NOW itself - that would sever the wire mid-reply. process.front/.quit share one process.result reply; their Toolbox calls are factored into proc_actions.c so the guest page and the wire handler use one implementation. Front, Quit, and the self-quit refusal are metal-verified on the PB1400c.

Screenshot App is its own verb, process.shot: the guest fronts the process, waits ~0.75 s for it to repaint (a deferred service pass, like the page's Front & Capture), reads its front window's fresh bounds off the anchor plane, captures ONLY that rectangle (capture_screen_rect), restores NOW, and delivers the crop over the capture transport - it reuses arm_transfer/capture.begin so the host receives it exactly as any capture, landing in the Screenshots module. The guest owns the timing, so the host-side delay hack is gone. Metal-verified cropping Finder and Strider on the PB1400c (2026-07-22). When the window bounds cannot be read - a genuinely windowless process - it falls back to a full-screen capture rather than erroring: the app is front, so the screen with it on it is a truthful answer.

Self-read fixed (2026-07-22): NOW reading its OWN windows returned "unreadable" (in the detail pane and to process.shot, which then failed "capture ended without a begin"). Cause: the anchor plane walks foreign memory at the classic 68K WindowRecord offsets, and NOW is a Carbon app whose own window records do not sit there. now_peek_windows_for_psn / now_peek_window_count now special-case self (SameProcess with GetCurrentProcess) and read NOW's own windows straight from the Window Manager (FrontWindow/GetNextWindow/GetWindowBounds/GetWTitle) - no reason to go foreign for oneself. So self now crops like any other process; the full-screen fallback remains only for the truly windowless. Metal-verified on the PB1400c (2026-07-22): the detail pane reads NOW's own windows and Screenshot App crops NOW's Workshop window.

With that, the whole drive arc is metal-verified: Bring to Front, Ask to Quit, the self-quit refusal, and Screenshot App cropping Finder, Strider, and NOW itself.

Smell, now fixed (tested, not yet metal-verified): the host's process list could hold stale PSNs across a guest relaunch, and a drive verb on a stale PSN failed (safely - the guest re-validated and answered ok:false / capture.end ok:false) until a manual Refresh. The list now notices the connection itself: ProcessesModel drops its whole table the instant the connection leaves .connected (rows belong to one connection, and the next guest reconnects with fresh PSNs), and the view re-reads on any transition back to connected - so a reconnect, or a pane reopened after one, reads afresh without a manual Refresh. Clearing on disconnect also covers the case the view's .onChange cannot see, a reconnect that happens while the Processes pane is closed. Host suite passes and the app builds; still needs a metal pass (relaunch the guest, confirm the list updates and the three drive verbs work with no Refresh). process.launch (opening an app that is not yet running) is the honest next verb; it needs a path/signature to name an unlaunched app, not a PSN. Everything is tested (contract round-trips incl. process.shot, a guest process.result fixture, the drivable/PSN decode) and builds clean on both halves.

Metal found one rung-0 bug, now fixed: the detail pane's "Launched" line read "1/1/04" for every process. ProcessInfoRec.processLaunchDate is ticks since boot, not a 1904-epoch date, so LongDateString clamped it. Now rendered as elapsed time via proc_uptime_text (pure, native- tested, watched failing by mutation): "3 min ago", "2 hr 14 min ago".

Workshop follow-ups, deliberately not done in the arc: a CarbonLib 1.6 launch gate (wire.c still surfaces kConnNeedsCarbonLib at connect time instead); the capture disclosure's expanded state is session-only, not persisted; the Files page's Send File button sits in the share block rather than the header placard the spec drew; and the sidebar has no focus ring, so Tab reaches controls but never the rail (arrows work whenever no field has focus).

Broken

Hard system crash (error 10) on quit — root-caused and fixed, metal soak pending (2026-07-23). Twice, quitting NOW hard-crashed the guest (error 10, a Line-F/unimplemented-instruction exception) and required a reboot; not every quit, and the app logged its own clean "stopped" first. Root cause: a shutdown use-after-free in every Data Browser module. workshop_close disposes each module (g_ops[i]->dispose()) BEFORE DisposeWindow, but each module's dispose freed its Data Browser item-data/notification/compare UPPs while the browser control was still live in the window — on the belief that "the window took the controls with it," which the call order makes false. DisposeWindow then tore down that live browser, which fires item notifications (removal, a deselect) through the now-freed UPPs. On PPC a UPP is a transition vector; once freed and reused, that call lands in garbage — an illegal instruction that corrupts the system heap, hence the reboot. Intermittent because it depends on whether the freed block was reused yet and whether the browser still held items to notify on. Fixed in all four DB modules (software, processes, census, files_browser_view): DisposeControl the browser FIRST — while its UPPs and model are still valid — then free the UPPs, then the model. Builds clean under -Werror. Unverified: an intermittent crash cannot be proven gone by one quit; it needs a soak of repeated quits from each Data Browser page on the PowerBook. The Processes page was "metal-verified" and still carried this — the verification never included a quit-crash soak, which the ledger should now expect. To make a recurrence diagnosable, teardown now leaves a FLUSHED breadcrumb before each step (quit: closing connectionstopping pumpremoving handlersdisposing windowclean) and closes the log LAST: a crash log that ends at quit: disposing window says the fix did not hold; one that reaches quit: clean/stopped is a clean teardown. Ordinary log lines sit in the disk cache and a crash loses them, so the breadcrumbs force FlushVol (now_log_flush), the same guarantee now_log already gives an error line.

Resume by offset hangs. A transfer resumed against a matching partial does not complete. The failing test is committed rather than skipped (MetalLargeTransferTests), which is the right shape: the feature announces its own absence. See docs/large-transfers.md.

One large transfer in about six degrades badly. 12 MB normally lands at ~293 KB/s; occasionally a run collapses anyway. The mechanism behind the common case is understood and fixed — this residual says the understanding is not complete. Measured, not reasoned about; the numbers are in docs/large-transfers.md.

An unreachable host presents as a hang. Reimplemented on claude/processes-module-cb2d9c (2026-07-21) after the wedge bit again on metal: now-guest-processes decoded under a fresh name, found no prefs, dialed 10.0.2.2, and a synchronous OTConnect to an address that never answers blocks INSIDE the call — before the first update event, so the window stays blank and only a force quit ends it. The fix is the codex branch's shape (160ed85) rebuilt against this tree: the endpoint goes asynchronous for the dial only, a notifier publishes one flag, the main loop finishes or fails the connect, and the endpoint returns to synchronous before the hello. now_log_open() also moved above conn_init() so this failure finally leaves a log. ot_connect_source_test.py pins the sequence — it was watched failing against the pre-patch sources — because this fix has now been lost once. Metal-verified 2026-07-21 on the PB1400c: launched with no prefs it dials the gateway and the UI stays alive and drivable, where before it wedged blank. (The emulator forgives the synchronous form, so this could only ever be proven on hardware.)

The Connection fields were dead once Connection became a page (fixed 2026-07-21, claude/processes-module-cb2d9c). Address and port took no clicks. The real cause, after two wrong guesses: the Workshop window had no root control, so it had no control-embedding hierarchy, so SetKeyboardFocus could not work and an edit-text control could take neither focus, clicks, nor keystrokes. This is the same wall the Console hit on metal - "the edit-text field never took a keystroke" - which is why it hand-rolled its input. Connection is the only page that uses edit-text controls; every other page's controls (buttons, checkboxes, popups, Data Browser, scrollbar) respond through TrackControl/HandleControlClick, which need no focus, so only Connection was affected.

Two dead ends before the fix, both worth recording because they are the wrong instinct:

  1. FindControlUnderMouse instead of FindControl - no effect, because the window had no embedding hierarchy to be wrong about.
  2. Adding a root control to the window. It got the field to focus but it still took no mouse or keys (the Appearance edit-text control just does not work for entry in this WaitNextEvent app), AND it broke every other control in the group: a root control turns the group-box control into an embedder, and an embedded control only receives clicks when HIToolbox's standard Carbon Event handler routes them, which this app deliberately does not install. So the retry popup and checkbox - which had worked - went dead too. The root control was removed.

The fix that holds: no root control anywhere (controls stay flat siblings the classic Control Manager hit-tests directly, with plain FindControl), and text entry moved out of the page entirely. Address and port are drawn read-only; an Edit button opens a movable-modal dialog (conn_edit_dialog.c, DLOG/DITL 301) whose entry the Dialog Manager drives - GetNewDialog + ModalDialog(now_pump_modal_filter()) + GetDialogItemText on editText items. That is the exact mechanism the original Connection dialog used before the Workshop rewrite, proven on this PowerBook; its own window has its own text handling, independent of the Workshop window. The filter pumps the wire; validation stays in conn_fields.c.

Net change from the last metal-verified state is only the Connection dialog: every page's control handling is back to no-root + FindControl.

Metal-verified 2026-07-21 on the PB1400c: the "Other Mac" popup/checkbox/Edit button click, the Edit dialog's fields take clicks and keys, and Save repoints the connection - which is how the wrong-port chip got corrected. Screenshots/Files/Console unchanged.

Type-select does nothing in the browser list. Selection, double-click and header sorting all work; typing a letter does not jump. SetKeyboardFocus is set and the key reaches the control. Universal Interfaces 3.4 has no type-select column flag, so the likely answer is that Data Browser wants the Carbon Event path — which means an event- model migration, and the Carbon UI skill explicitly warns against running two competing top-level loops in a mature WaitNextEvent app. Not load-bearing; parked as a known gap rather than chased.

Unverified on the machine

Everything here builds and passes its tests. None of it has been watched working on the PowerBook.

  • The gestalt reply's truncation path (2026-07-30, branch claude/goofy-sutherland-3e6cf4). run_gestalt used to write every structural byte of its JSON — the [, ], , and the quotes around each label and value — with a bare out[pos++], checking the cap only around the escaped text and once more at the very end, after all the unchecked writes. It never bit: a whole-machine gestalt is about 26 rows and sits well inside the 3072-byte buffer wire.c passes. But kGestaltMaxRows is 48 and a GestaltRow is 96 bytes, so a machine answering more selectors, or one more group, would have run past the caller's buffer rather than truncating.

The serializer now lives in now-guest-ppc/src/commands/gestalt_json.c, split out precisely so the host cc can run it at caps no Macintosh will ever produce — gestalt_json_test.c sweeps every cap from the floor to 6000 with a poisoned guard region past the bound. That sweep found a second defect the hand-picked sizes missed: the ] closing a group is a write like any other and can be the one that hits the bound, and clearing group_open regardless left an array nothing closed.

A truncated reply now says so, in a notice group carrying ["truncated", "<n> rows omitted - reply buffer full"] (contract: x-commands/gestalt/output), because a short reply and a machine with fewer facts to report were previously indistinguishable. Room for it is held back from the cap up front — a buffer too full for rows is also too full for the sentence explaining why.

What has not happened: no real gather has ever been large enough to truncate, so the notice has never crossed the wire from a Macintosh, and no host UI has been seen rendering it. The host's console shows command.result groups generically, so it should appear as a group named notice with one row; that is inferred from the code, not watched. Anyone with a machine that answers unusually many selectors is the first person who could confirm it.

  • ps on NOW-68K's wire (2026-07-25, branch claude/host-console-remote-shell). The dumb-shell console landed and ps still came back unknown-command from a 68K guest that ran it perfectly at its own keyboard: it had been added to conwin.c alone, reading the process.list family the wire already served. A message family serves a module, not a person — the host console sends commands and nothing else — so ps is now in commands68.c's table and its reply is built by n68_proclist_render_ps() from the same proc_list_rows() walk that feeds process.listing and the guest's own console text.

Tested here: the new renderer's shape, its empty case, its refusal at a hopeless cap and its worst-case row bound (test_proclist.c, and the truncation guard watched failing by mutation); two host fixtures for the reply as the guest writes it; and a new parity test, testEveryVerbTheSixtyEightKConsoleAnswersIsAlsoOnItsWire, watched naming exactly this bug when ps is pulled back out. The 68K guest cross-compiles clean. Unverified: nobody has typed ps into the host console against a real 68K Mac. Two things to watch when someone does — the truncation row (["...", "N more not shown"]) appears only on a machine running more processes than a 1 KB control frame holds, which is roughly a dozen and may never happen on a 7.1 machine; and the detail column is meant to read identically to the PowerPC guest's, which no run has yet compared side by side.

  • The dumb-shell console, both guests (2026-07-25, branch thread/host-menu-dumb-console). The host console no longer knows what commands the guest has: it sends command.request with line — the raw text a human typed — and renders whatever comes back. Every argument grammar moved to the machine that serves the verb (now-guest-ppc/src/commands/cmd_line.c, natively tested by mutation), help became an x-command answered from the one doc table each guest already showed its own console (now-guest-ppc/src/commands/cmd_help.c, now-guest-68k/src/commands/commands68.c), and the host's Tab completion is that answer at runtime.

Tested here: 459 host tests, the two new native guest tests, and both guests cross-compile clean at -Wall -Wextra -Werror. Nothing has been typed into a console on either machine. What that leaves specifically unverified:

  • gestalt slicing now happens guest-side from the line (--full, --cpu, …). Absent-line behaviour is unchanged for modules, but no human has typed gestalt --memory at a PowerBook.
  • screenshot --depth 8 --no-save and tail 40 parse from the line for the first time; the old host-side parsers are gone.
  • help on the PowerPC guest builds a ~1.2 KB reply against a 4 KB control frame with a byte-budget truncation row. The budget is reasoned, not measured on the wire.
  • help on NOW-68K builds into a 512-byte payload buffer and measures ~260 by hand-count. It has never been sent.
  • The MacRoman decode of an accented path typed as a console line (ls Café:Notes) is covered by a native test on the decoder, not by a file with that name on a real HFS volume.

  • ⌘Q's farewell, on metal (2026-07-25, same branch). The host now returns terminateLater and waits for bye shutting-down to reach the socket before the process ends, bounded at 0.5 s. Tested here by sequencing (mutation-verified: a shutDown that reports synchronously fails), and the menu bar and its Quit item were driven live through accessibility on this Mac. Not verified: that the ⌘Q keystroke dispatches (script-driven activation is refused in this environment, so the item was clicked rather than typed), and that a PowerBook watching the wire draws the right conclusion — the guest's own "host went away" handling has not been observed against a real quit.

  • quit <name> — the deploy loop's missing half (2026-07-25, branch thread/guest-quit-command). A console command and x-command that composes process.list → match by name → re-validate → process.quitre-list, so it can report gone apart from still-running. Design, outcome table and the decisions behind each case: processes-and-peek.md.

Emulator-verified, end to end, on mac99 / OS 9.1 / CarbonLib 1.6 — both invocation paths, and every outcome the composition can produce:

Watched Result
Guest console quit SimpleText "SimpleText" is gone (0.3 s)
Guest console quit --no-wait SimpleText asked "SimpleText" to quit; NOT confirmed (--no-wait)
Guest console help quit renders
Wire quit SimpleText gone (0.1 s), and confirmed absent by an independent process.list
Wire, dirty document [quit-declined] … is STILL RUNNING after 4 s, with SimpleText visibly sitting on its Save dialog
Wire, nothing of that name not-running, ok:true
Wire, its own name [quit-refused], and still there afterwards
Wire, no target / unknown flag [quit-bad-args]

The acceptance driver is committed: MetalQuitTests (NOW_METAL=1, plus NOW_QUIT_DIRTY=1 for the human-in-the-loop declined case).

What the PowerBook still has to settle. The emulator says nothing about timing on a 117 MHz 603e: SimpleText answered in 0.1–0.3 s there, and the 6 s default was chosen for a slower machine, not measured on one. Nor has the deliberate stall been felt on metal — for up to --wait N the guest's window does not repaint (it keeps servicing the wire; see nested-loops.md), and "does that read as a hang?" is a question about a real screen. An isolated copy is staged at Lab:now-quit on the 1400c (its own name, so its own preferences; fork sizes verified against the local MacBinary, 565127 / 2439). Being non-canonical it starts with no preferences and dials 10.0.2.2, so the console path needs no host at all — that is the one to run first. The real target is NetPresenz on a 180c, which is a different machine and a different client.

  • catsearch — the Software module's feasibility probe (2026-07-22). Times a whole-volume PBCatSearch sweep for APPL files on the startup volume, in 15-tick slices, cold then warm. Console verb on both sides (contract x-commands, guest commands.c, host ConsoleModel). Metal-verified on the 1400c (guest console path; same-day emulator run agreed in shape): 22,127 files / 2,411 folders, 601 APPL hits, cold sweep 228 ticks = 3.8 s in 184 slices, warm 172 ticks = 2.9 s, longest slice 3 ticks against the 15-tick budget, zero restarts. Two conclusions the Software module can build on: a full inventory sweep is affordable as background idle() work (~50 ms worst slice), and 184 slices ≈ the catalog arriving one 16 KB opt buffer per call — so the buffer size, not ioSearchTime, is the real slice-length dial. Warm is barely cheaper than cold; do not design around the cache. The host-console invocation was watched working too (2026-07-22, post-merge build), so both invocation paths are metal-verified — including MacRoman-high-byte names in First hits crossing the wire through the \uXXXX escaper.

  • Software rung 3 begins: the page is registered and appears (2026-07-22, spec in software-module.md, mock in mockups/software-mockup.html). The six-edit registration for a new nav module landed and is emulator-verified: Software shows as the 6th rail row (a boxed-app-tiles ics# 136) between Hardware and the pinned Logs/Connection pair, Cmd-6 selects it, and it draws the live installed-software overview (139 extensions, 33 control panels, …). The delicate part — inserting Software as id 6 pushed Logs 6→7 and Connection 7→8, the first insert to move an existing non-pinned id — bumped prefs to format 14 with a remap lifting both; the save/load round-trip is verified (quit + relaunch reopened on Software). Two supporting pieces are host-cc tested and integrated: software_layout.c (split-view geometry) and sw_vers_parse.c + now_software_read_version() (the 'vers' parse extracted to a unit with a mutation watched failing under ASan; the per-row primitive the trickle will call). Still ahead on this rung (the frame is drawn, these land on it): the interactive Data Browser with the FSSpec- bearing item model, the domain popup, live search, the launch/front/ quit/reveal buttons, and the idle-paced version trickle. None of that is metal-verified yet — only the emulator, and only the page's appearance + prefs migration.

  • Interactive cut (2026-07-22): first version was hand-drawn and metal-tested the same day; the second metal round found real problems, all fixed and re-verified in the emulator:
    • The module leaked port state. Three RGBBackColor(white) calls on the one shared Workshop window turned EVERY page's background white. Fixed by rule, not by restore: the module never touches the background color — white interiors are fore-painted with PaintRect. Watched fixed (Hardware gray again after visiting Software).
    • The list is a real Data Browser now (the processes_module pattern): Platinum header buttons, native four-column sort, native truncation/scrolling. Loading appends items and versions update one cell — the flashing was the hand-drawn list's invalidation model, and it is gone with the list.
    • Domains cache in memory for the run (lazy NewPtr each); switching rebuilds the browser from the cache, never the disk; Rescan is the only re-read; the apps sweep is resumable across switches. Watched: Extensions ↔ Applications both ways, the restore instant with versions intact.
    • The search field takes its click (focus ring); the detail pane is a group box with theme fonts and the selection's icon (GetIconRef on first selection only, cached for the run). Emulator-watched: sweep→browser fill, version cells trickling, live search (8 of 205), the domain popup driven by a genuine held QMP drag, cache restores, page-switch persistence, the bg fix. Not watched, needs a human click: row click-to-select and the search focus ring — a control experiment showed the metal-verified Processes browser ALSO ignores injected clicks (atomic and QMP-held), so this is an injection-vs-DataBrowser artifact, not a known defect; still, only a hand on a mouse closes it.
  • Fourth round (2026-07-22): the third metal round's four asks. The residual flashing was batched sorted inserts shuffling visible rows — the browser is now fed nothing mid-sweep (the placard counts arrivals) and populates ONCE at sweep end, watched. Duplicate groups: same-name items collapse under a container row (disclosure in the Name column, "N items", aggregate size, running-if-any; parents disclose, never select) — watched as "now-guest · 2 items · 1.0M · running" with indented per-version children, isolated by search ("2 of 206"). Where: the full path, wrapped, in the detail — watched, computed on selection never in draw. Show in Finder: alias in a 'misc'/'mvis' Apple Event, Finder fronted — watched revealing Note Pad in Apple Extras, matching the detail exactly. Bring to Front / Quit: wired over the metal-verified proc_actions with a fresh at-act-time PSN join; unwatched as buttons (the VM's only running singleton is the injection channel itself). Also unwatched: groups' collapsed-default on the unfiltered list. Nothing in this round is metal-verified yet.
  • Metal feedback on the host page (2026-07-23), two fixes. (1) The ® was passed down poorly. A launch/reveal from the host against an app with a non-ASCII name ("Adobe Photoshop® 5.0") came back "no such file", the echoed path double-mangled to ¬Æ. Cause: the host sends HFS names as UTF-8 (® = 0xC2 0xAE), but run_launch/run_vers/run_reveal read target with now_json_find_string (a raw byte copy), so FSMakeFSSpec never saw the MacRoman byte (0xA8). Fixed by reading target with now_json_find_text — the inbound half of now_json_escape, which decodes \u and raw UTF-8 back to MacRoman (a json_native_test case pins the ® round trip). The same latent bug in the Files path commands (mv/trash/restore/mkdir/offer/list/get in wire.c, console ls) was fixed in the same defect class by a parallel task — see "Non-ASCII paths INBOUND" in the Files section, guarded by test_inbound_hfs_path and a source-reading conformance test. (2) Selection hilite hugged the text. The Data Browser's default kDataBrowserTableViewMinimalHilite draws the selection only behind each cell's glyphs, so a selected row read as three disconnected patches; switched to kDataBrowserTableViewFillHilite for one continuous full-row bar (CarbonLib 1.1+, we floor at 1.6). Guest builds clean under -Werror; both unverified on metal — the reveal round trip needs the connected session, and the hilite is a visual change to watch on the PowerBook.
  • Host page reaches parity: split-pane, detail, reveal (2026-07-22). The host Software page grew a second half. It is now an HSplitView — the inventory Table on the left, a detail pane on the right carrying the selected item's version, size, state, kind, and full path (selectable), with Launch and Show in Finder beneath it. Search was already there; it stays, above the split. "Show in Finder" is a new wire verb, reveal — launch's read-only twin: it resolves a target exactly as launch and vers do (path / #n / bare name) but reveals ANY item (an extension, a control panel), since it opens nothing. The guest serves it by sending its OWN Finder a kAEMakeObjectsVisible for the item's alias then fronting the Finder — the same now_software_reveal the guest page's own button uses, now reachable from the host and the console (reveal <name|path|#n>). Contract-first: the reveal x-command is declared, answered in commands.c, and offered by the host console — CommandRegistryTests' three-way agreement holds. Host suite green (276) incl. a reveal test; guest builds clean under -Werror; audit clean. Never run live: like rung 4, the reveal round trip and the split-pane page both await a connected session with both new builds. reveal from the host console against a live guest, and the detail pane's two buttons, are the one-sitting check.
  • Rung 4 lands (2026-07-22): versions on the wire + the host Software page. serve_software_list now fills each served entry's version (a page's worth of fork opens per request, bounded, explicitly asked for); the contract, fixture, and Swift docs agree. SoftwareModel/SoftwareModuleView mirror the guest page host-side — domain picker over a Table, client-side search, Launch by the entry's path (the guest's words shown either way), the listing's note surfaced verbatim — registered between Hardware and the footer. Host suite green incl. SoftwareModelTests and the updated registry manifest. Never run live, all of it: the software.list round trip (and now the version enrichment and the page on top of it) awaits the first connected session with both new builds — swpage extensions in the host console, then the Software page itself, is the one-sitting check.
  • Fifth round (2026-07-22): the metal report "a collapsed group will not re-expand" was a real contract miss: closing a container REMOVES its children (the Data Browser's own behavior) and item_notify ignored container notifications, so reopen had nothing to show. Fixed: kDataBrowserContainerOpened re-adds the group's children, idempotent via GetDataBrowserItemCount. Unwatched — the disclosure triangle defeats click injection; the repro is on the PowerBook. Also added: a draggable splitter between the panes (gray XOR outline, own StillDown loop pumping the wire — nested-loops.md row added — clamps tested host-cc, session-only width). Watched end to end in the emulator. Bonus close: a press-MOVE-release drives the Data Browser under injection, so row click-to-select is now watched (previously the oldest gap). Known nit: below ~260px list width the fixed columns clip; a tighter clamp is a one-liner when it bothers. The whole-window redraw on module switch is spun off as its own task (parent container, not this module).
  • Sixth round (2026-07-22): the search field repainted the whole module per keystroke — Remove-all/Add-all, an unconditional detail invalidation, the group qsort, and a catalog walk, every key. Typing now refilters by DIFF against a view set (delta rows only, groups leave children-first), the detail is touched only when the selection actually changed, and there is no auto-pick mid-typing. The full rebuild remains for content changes. Per the redraw contract added to classic-mac-carbon-ui the same day. Emulator- watched: the selection and detail pane SURVIVE keystrokes untouched; a filtered-out selection clears once. The reduced repaint itself, like all flicker, only reads on metal.

    • The field itself still blinked (whole-field invalidate + full white repaint per key). Typing now echoes the DELTA directly — the contract's immediate-feedback exception: erase from the end of the unchanged prefix only, draw the tail + caret, clip restored, nothing invalidated; draw_search reproduces the same pixels at any real update. Emulator-watched ("quicktime" typed and backspaced entirely through the echo path).
  • Software rungs 1–2: resumable sweep, vers, running tags, and the software.list family (2026-07-22, spec in software-module.md). Rung 1 is emulator-verified: sw extensions tagged exactly the three running appe files the harness's process list names (Control Strip Extension, DVD AutoLauncher, FBC Indexing Scheduler), and vers SimpleText read Version 1.4 / "1.4.0 final" / the Get Info string / Product 1.1 by name-search resolution. Known texture: Application Switcher runs but is untagged — its process appSpec evidently names the System file, and the strict FSSpec compare declines to guess; that is the join being honest, not a defect. Rung 2 (the wire family, served from a one-domain cache with full-path launch keys) builds and is host-tested — fixtures pin the piecemeal listing including a MacRoman ® — but has never run live end to end: it needs the new host build connected to the new guest build, driven by the host console's swpage [domain] [cursor].

  • First metal round (2026-07-22, partial): sw apps and vers ran on the 1400c from the host console. Two findings, both closed the same day: launch from the host dispatched as unknown-command — the host sorts JSON keys, args precedes name, and the guest scans frames FLAT, so launch's arg named "name" was read as the command name (arg renamed target; the never-shadow-an-envelope-key rule now lives in the contract's x-commands preamble); and vers on a bare name met the disk's several SimpleTexts — it now shows every match path-first instead of refusing, launch's ambiguity refusal names the paths, and a duplicate finder (same/different version, user-driven consolidation) is marked in the spec as later work.
  • Second metal round (2026-07-22, same day): the multi-match view worked but truncated paths mid-folder, and retyping a full HFS path to disambiguate is brutal. Both fixed: matches print as a numbered list whose paths wrap across continuation rows, the list is stored on the guest, and launch #2 / vers #2 pick from it — either console, one wire frame. launch's ambiguity answers a distinct launch-ambiguous code for a future host UI. Emulator-verified with a manufactured duplicate (two now-guests: refusal listed both full paths, vers #2 read the picked copy, launch #1 launched).
  • Third metal round → launch redesign (2026-07-22, same day): the numbered-pick flow worked on the 1400c but read as too much ceremony for "just open it." launch <name> now launches the highest-versioned copy and names it in the reply (a visible answer, not a hidden guess); launch <name> <version> forces a copy by its short version string; full path and #n still work. The whole arg is tried as a literal name first, so "Sherlock 2" stays whole. Emulator-verified (newest-of-2, version pick, wrong-version message, single-match plain launch).
  • Fourth round → -v flag (2026-07-22): launch-newest became too surprising to reason about (which version won?), so the shape settled: launch [-v VERSION] NAME, NAME the whole remainder (spaces need no quotes; quotes stripped if used), a bare ambiguous name launches the FIRST found and names its version (one fork open, no walk), -v forces a copy, positional Name 1.2.3 retired with a "did you mean -v" hint. Emulator-verified: quote-strip, first-of-2-with-version, the hint. The -v launch flag is metal-verified (Michelle, 2026-07-22, human-typed — the emulator keystroke injection had dropped its leading chars, an input artifact, never the code). The software.list wire family (swpage) remains the one never-run-live path — it needs a host linked to the guest, deferred until the guest page is dialed in.

  • sw and launch — the software family's first verbs (2026-07-22). The Software module's data layer (software.c) surfaced as console verbs on both sides before the page exists. sw inventories the special folders live (Extensions Manager's disabled siblings tagged "(off)") and pages applications via the catsearch-verified APPL sweep, stopped at one page; launch opens an application by exact-name search or full HFS path, refuses ambiguous names, and logs outcomes under sw — it is the family's one mutation. Versions are deliberately absent: one 'vers' read per file is the expensive path, deferred to the module's lazy detail. Emulator-verified (OS 9.1 clone): overview counts (139/33/0/13), sw extensions with types+sizes, sw apps page with the more marker, and launch SimpleText bringing a live SimpleText to front. Not yet watched on the 1400c, and the guest's LOCAL launch intentionally does not log (only the wire path does — same rule as ls/ps); the host-console invocations of both verbs are host-tested but unrun live.

  • A page switch paints once, and only what changed (2026-07-22). Michelle watched Workshop page switches repaint the whole window on the PowerBook — rail, placards, everything. The investigation found the container's invalidation was already scoped (header/body/status plus the two selection rows); the churn was in the painting, three ways: HideControl/ShowControl draw immediately, so show(false)/show(true) repainted the pane piecemeal before the update event repainted it again; the update handler's full-port EraseRect painted the invalidated rail rows theme-gray a beat before the rail's own white erase; and DrawControls followed by UpdateControls drew every control twice per update. All three fixed in workshop_window.c alone: the swap runs under an empty clip and paints exactly once at the coalesced update, the erase narrowed to the body plus the sidebar gutter outside the rail panel (the placards and the rail fill their own faces), and one UpdateControls pass. Emulator-verified: all seven pages cycle with no stale pixels, zoom leaves the gutter clean, controls still track after switches. Watch on metal: that the rail genuinely stops flashing at the machine's real drawing speed — the emulator is too fast to show a flash either way. One module-side offender remains, out of the container's scope: the copy-pasted set_status in screenshots/census/connection invalidates a full-width bottom strip (port bounds, bottom 23 px) that crosses the rail's foot, so the Connection row can still flick when a module's status line changes. The module fix is to invalidate the status placard's rect, not the port's.

  • The Logs page, both machines (2026-07-22). A Monaco dump of the in-memory log ring that follows the tail live like a terminal, with Invert and Log-to-disk switches. The guest page was watched working on the PB1400c; the footer move, the invert switch, and the whole host module are built and tested but unrun since.
  • Placement. Pinned in the footer below the divider, directly above Connection — a logs_row on the guest (id 6, Connection 7), a .footer descriptor before settings on the host. The host footer row now shows link status only for the row that IS the link.
  • Guest scrollback. The ring grew 200 -> 2000 lines (kLogKept), ~240 KB of statics against a 6 MB partition. run_tail's stack index was decoupled from kLogKept so it stays 48, not 2000, pointers.
  • Disk toggle. now_log_set_disk/now_log_disk_on (guest) and HostLog.setPersistsToDisk gate the file at runtime; the ring is always live. Default on (crash survival is the point). Both switches reflect the ACTUAL state, so a failed open reads as off. On the host the file is now a switch, not opened at launch — LogsModel applies the saved choice.
  • Invert. A dark canvas like Console, saved per page. Guest prefs reached format 13 for it (12 was the disk field + Connection renumber); the host keeps logsInvert in UserDefaults.
  • Watch on metal: the host module unrun entirely; on the guest, that the invert switch redraws cleanly and the footer pair (Logs above Connection, under the divider) lays out at 640x480.
  • ps and census console commands + guest verb logging (2026-07-22). The two new modules — Processes and Hardware/census — had no console verb and logged nothing; both are now closed.
  • Console. ps (flat process list, the reading of process.list the Processes module drives) and census [probe] (one probe page, the flat cousin of censusExchange) were added across all three halves — contract x-commands, guest commands.c dispatch, host ConsoleModel offer + help — the way ls is to file.list. CommandRegistryTests reads all three and is green, so the set agrees and every offered command has help. The guest's own console (console_model.c) renders both locally too.
  • Logging. The guest drive verbs (process.front/quit/shot), census outcomes, and the process-list refresh now log their shape with the wire id (areas proc, census). The refusal reasons that used to live only on the wire now reach the log. process.list logs once per refresh (cursor 1), never per page, to stay off the per-chunk heartbeat rule.
  • Verified only here: host suite (263 tests) green, audit_source.py clean, the census/json header chains compile under cc -Wall -Wextra -Werror. Not cross-compiled — no Retro68 toolchain this session, so the guest-only additions (run_ps, run_census, the two console handlers, the wire.c log lines) are not even at builds yet. First metal run should confirm ps, census pci/ata/etc., and that a declined quit shows in the log.
  • The Processes page's product pass (2026-07-21) - built and suite-green, unrun on metal. All app-side (extension unchanged):
  • Kind grouping. Processes are classed from processMode (modeOnlyBackground), not guessed from the 'appe' type. The list sorts front-process first, then applications, a divider row, then background-only - kind and front-ness are the sort axes, never window state, so a row never jumps when a window opens/closes.
  • Row badges. Front app reads "(front)"; apps show their window count ("3 windows"); windowless and background rows show none - the windowed/windowless distinction, visible without selecting.
  • Richer detail. CPU time (processActiveTime), accurate Kind with "(frontmost)", and a Windows section listing each window's title + size (up to 3, "...and N more"), read through the anchor plane's validated foreign path (now walking the nextWindow chain and reading titleHandle). Menus line is a reserved STUB - the anchor captures MenuList, the walk is a later pass.

Watch on metal: the divider row is a non-process sentinel item in the Data Browser (kDividerItem), non-selectable by bouncing the selection off it - the one bit of fake-row territory in an otherwise proven-DB design; confirm it draws between the groups and cannot be selected. Also that window titles read correctly (another foreign pointer hop, titleHandle), and that per-app window-count reads every second don't cost visible time on the 33 MHz metal. - Prefs v10 module renumbering. Connection moved 4 to 5; a v9 file should reopen on the page the person had (the remap is three lines in now_prefs_load), exercised only by reasoning - same status as the v9 note below. - Corners of the Workshop no one has exercised anywhere: the send progress bar actually moving, and the preview well at 16/32-bit depths. (The first metal pass found two bugs - a mute Console edit-text and Modified dates clamped to 1/19/72 by signed DateString - both fixed the same night and metal-verified the next morning, 2026-07-21.) - Prefs v9. Reads v1-v8 files and seeds the Console page from a legacy console_open flag; exercised only by reasoning, not by an old prefs file on the machine. - The host serving move / trash / restore / mkdir. 13 tests, zero minutes of machine time. No client asks for it yet (see above). - Accented file names. macOS stores names decomposed, so "café" is "cafe" plus a combining accent, and MacRoman has the letter but not the mark — every accented name was arriving as "cafe_". The fix composes first. Nobody has pulled an accented file to the PowerBook. - Non-ASCII paths INBOUND, host to guest. The complement of the above, and the same defect class as the Software fix: the host sends every path UTF-8 (® is 0xC2 0xAE), but FSMakeFSSpec wants the MacRoman byte (0xA8). The guest's file-op verbs were pulling path/toPath/trashedAs with now_json_find_string, which does not convert, so a move/trash/restore/mkdir/list of any non-ASCII name looked for a file that does not exist. Fixed by switching those extractors (and file.offer's name, and console ls) to now_json_find_text; container/fileType/creator/tokens stay find_string, ASCII by contract. Guarded two ways — json_native_test.c :: test_inbound_hfs_path proves café® decodes to 0x8E 0xA8 (and that find_string leaves the raw four bytes), and GuestWireConformanceTests :: testHfsPathArgumentsAreTextDecoded reads the C and fails if any of those keys reverts to find_string (mutation-verified). Tested, not metal-verified: no one has moved or trashed an accented file from the host to the PowerBook. - The Finder reveal button. "Open" in the browser sends odoc to the Finder with an alias to the downloads folder. Standard, and untested on metal; it is kAENoReply so it should not block, but that is reasoning rather than evidence. - The Hardware census module (slice 1). New Workshop page: a passive census of this Mac, three Carbon-clean probes (gestalt full selector-table walk, video GDevice walk, volumes PBHGetVInfo), served over the new symmetric census.request/census.report family and shown in a split pane (probe list left, rows right). Builds clean (whole guest links; the ics# 133 chip icon compiles) and the host suite is green (242 tests), including a guest→host refusal round trip, the census.report fixture, and a mutation-checked serializer. Not watched on the PowerBook. Specific unknowns for the first metal pass: (1) two Data Browsers in one window — one is proven by the Files page, two side by side is not; (2) the full ~203-selector Gestalt walk paging 16 at a time; (3) the chip icon actually plotting from ics# 133 rather than losing to a System family at that id. See docs/adding-a-workshop-module.md. - The host Hardware module — runs and reads the GUEST's census (2026-07-22). A native macOS dossier: a census module in the sidebar (CensusModel + CensusModuleView), a probe list on the left and the selected probe's rows on the right, a Run Census sweep and per-probe rerun. It is a REQUESTER only — it asks the guest and displays the reports, following the more/cursor pagination to accumulate a probe's rows one page per request. The host probe registry (CensusProbes.all) is a copy of the guest's k_probes[] and the contract's x-probes; CensusProbeRegistryTests pins the set to the contract and the order to the guest, so a probe grown on one side and forgotten here fails a test. Tested, not seen against a real guest. CensusModuleModelTests drives the whole request→report path over the loopback listener with a scripted guest (pagination, cursor threading, outcome/note propagation, the full sweep, the disconnected guard, and rerun-replaces-not-appends); the SwiftUI view itself has not been run against a connected PowerBook. - The host does NOT serve its own census, by design. The census family is symmetric in the contract, but the guest is the machine with hardware worth asking about; the host is the requester. When the guest sends the host a census.request, the host answers refused ("the host does not serve a census yet"). That is a deliberate, permanent- feeling asymmetry now, not a scheduled stub — a host self-census (IOKit/ sysctl) is not planned as part of this feature. - The ata and pccard probes reach 68K-trap-only managers through a metal-proven Mixed Mode dispatch (census_trap.c, 2026-07-22). The 1400c's ATA Manager ($AAF1) and PC Card Manager ($AAF0) are trap-only — no CFM fragment, and gestaltATAAttr answers falsely absent — so a PowerPC Carbon app cannot import them. census_trap.c reaches them the way the parent project proved safe after four machine wedges (corpus cis-metal-safe-mixed-mode-fix): a hand-built M68K RoutineDescriptor so CallUniversalProc thunks PPC→68K, CallUniversalProc resolved from InterfaceLib and called variadically (a fixed-arg pointer leaves the args in registers → Type 1 bus error), and each thunk keeping its RTS return address on the stack. The mechanism itself is metal-verified by spikes/census-trap: selftest $4242, then real traps. - pccard (CSGetCardServicesInfo, selector 7) is metal-verified on the 1400c: CS 2.01, 4 sockets, Apple vendor string. Read-only, touches no socket or card, so it runs in the sweep. A card's own identity (its CIS) stays OUT until a gated design — the CIS is what froze the 1400c historically (pb1400-pccard-trap-only). - ata (IDENTIFY DEVICE) reaches the manager and it answers noErr, but on the 1400c internal drive the IDENTIFY buffer comes back empty (metal, 2026-07-22 — buffer dumped all-zeros for the one device that answers, device id $0000). So the row honestly reports the device present without a model. A drive that fills the buffer decodes into model/capacity/firmware; that path is builds-only. Getting model/serial off this drive is a separate follow-up (kATAMgrBusInquiry enumeration, or kATAMgrExecIO issuing a raw IDENTIFY task file rather than the manager's empty DriveIdentify) — deferred, banked with the wins per Michelle's call. - The whole integrated page — pccard/ata running inside the census sweep and rail — is tested and builds here; it has not yet been metal-verified as a page (only the underlying trap calls have). - The power probe. Slice-2 follow-up (2026-07-21). Carbon-clean (BatteryCount / GetScaledBatteryInfo, gated on gestaltPowerMgrAttr) and low-risk. Compiles, links and passes its decoder unit tests; has not run in the page on metal. - network and software probes, deferred as future modules (decided 2026-07-21). Network (Open Transport interfaces and TCP/IP config) and installed-software (extensions and control panels with their vers) are both Carbon-clean and were scoped OUT of the census probe rail — Michelle's call was to grow them as their own future Workshop pages rather than more rows on Hardware. Not built; recorded so the intent is not lost. - The rail has no scroll bar. At fourteen probes the hand-drawn probe rail fits the standard window (~371px of rail for 352px of rows at 25px/row) but overflows below about the minimum window. draw_rail now clips the row list to the rail rect, so the tail truncates cleanly instead of painting over the button strip — but truncated rows are then unreachable. This is the point where the rail needs a real vertical scroll bar rather than shorter rows; it lands with the extension "witness" tier that adds the next probes.

Reverse file streaming is bounded and verified on the machine

The 2026-07-24 reverse-path pass removed both whole-artifact buffers. The guest now opens the source forks only after acceptance and emits one bounded frame at a time, including MacBinary header/fork/padding segments. The host writes each frame to a same-folder temporary file, preflights free space, computes CRC-32 incrementally, sends batched file.progress, verifies count and optional checksum, and only then moves or stream-converts the result into place. Cancel, truncation, checksum failure, write failure, and disconnect all delete the partial.

The native host suite exercises 256 KiB, 2 MiB, and 16 MiB payloads with a fixed 32 KiB append bound, CRC/truncation/overrun/cancel cleanup, atomic materialization, and text conversion across a chunk boundary. Both guest send entry points have a source gate against whole-file allocation, and the Retro68 guest build passes.

The bounded path is metal-verified on the PowerBook 1400c (2026-07-24). A separately named guest on port 5252 preserved the canonical pairing and persistent preferences. Data-fork pulls at 32767, 32768, 32769, 256 KiB, 1 MiB, and 4 MiB matched their generated content and independent CRC-32. MacRoman/CR conversion and explicit MacBinary data/resource-fork fidelity passed. Cancelling a 4 MiB pull removed its host partial and left the session responsive. The guest process partition was 6506 KB before and after; the 4 MiB pull added 2.23 MiB peak host RSS and 1.94 MiB live malloc bytes.

Those numbers are bounded observations, not a transfer-rate guarantee. The metal pass did not exceed 4 MiB, run longer than two minutes, mutate a source during transfer, or measure guest free heap. It does not prove rate hardening.

Reverse resume remains deliberately absent. A deployed guest supplies no source identity before the receiver chooses an offset, so the host cannot prove a retained partial belongs to the current source. An interruption therefore deletes the partial and retries from zero. Adding resume safely needs an additive guest-issued source token (and fixtures for old peers), not an offset guessed from a filename and size.

Structural work deferred on the host

A cleanup pass (2026-07-20) applied what was cheap and left three extractions from GuestListener.swift, which is 2094 lines:

  • Session is built with 28 on... closures, 25 of which only forward to a listener method. A weak var owner or a delegate protocol collapses about 180 lines, and adding a message stops meaning edits in four places.
  • The share-serving block (~140 lines) touches only share, session and state. It is a file server living inside a listener.
  • The outbound write path (~400 lines) shares one invariant — nothing may write to the connection while a bulk frame is half-written — currently enforced by a flag two unrelated methods must remember to check. As its own type the flag cannot be forgotten.

These were skipped on purpose. Two reviews proposed DIFFERENT reorganisations of the same file, and the receiving-half work above implies a third (one transfer sink rather than three accumulators). Doing any one now makes the others harder, and only the receiving half has a consequence beyond tidiness. Whoever takes that should take these with it.

V1 host product work is planned, not implemented

The NOW V1 host product roadmap starts only after the optional MCP companion V0 is complete. It commits a persistent target catalog and host-side improvements to Files, Processes, Software, the menu bar, quit policy, and Settings while retaining the current guest-dials-host, one-port, single-session transport.

V1 explicitly defers a guest listener, multi-session runtime, mobile transport, and shared protocol service. Any common-protocol extraction waits for CodeKitten's separate listener, pairing/security, health/latency, recovery, cooperative-loop, and adversarial multi-peer proof and would begin in another worktree. The exact target-switcher information architecture, pairing-conflict UX, thumbnail and history retention, inventory analyses, local-browser defaults, and remembered module-state policies remain intentionally open.

MCP V0.5 guest Files command seam has a tested staged-upload slice

The approved NOW MCP V0.5 guest-files roadmap now has its first host-owned command slices: an explicit, persisted and versioned root-relative guestRoot policy; canonical HFS path validation; capability, one-page listing, and bounded exact-stat commands; typed receipts; and normal host audit lines. It also has a create-only staged upload command: private disk-aware reservation, ordered 8 KiB-or-smaller chunks, SHA-256 sealing, a file-backed sender through the existing transfer lane, and bounded progress, reservation, finalization, integrity, and cleanup evidence. No host path crosses the API. The destination parent must already exist: this slice does not implicitly implement mkdir. The existing private local socket and client-launched stdio companion project those completed commands; download, mkdir, overwrite, move, delete, tree deployment, and prune remain unavailable.

The read-only slice composes the existing file.list exchange and therefore adds no guest message or guest code. It is tested against fake paired sessions, including root escape, invalid policy recovery, empty and populated listings, paging bounds, stale sessions, concurrency, and host-product noninterference, plus local-schema and stdio validation. A bounded 2026-07-24 PowerBook 1400c acceptance verified capability discovery, two 16-entry root pages with cursors 17 and 33, and exact stat. The first live page exposed one legal HFS name containing control bytes; path validation now keeps those exact MacRoman names addressable, rejects only untransportable NUL, and escapes them in audit text. Download and every broader mutation remain unverified and unavailable.

The staged-upload slice is tested, including host-space refusal, ordered offsets, integrity failure cleanup, dead-process orphan recovery, root escape, unavailable/stale sessions, replay, concurrent commit, malformed local/MCP requests and MacBinary, strict guest completion evidence, late-collision preservation, stale-accept invalidation, cleanup-needed recovery, guest refusal evidence, and unchanged one-at-a-time transfer ownership. Host staging and outbound reads use bounded off-UI-actor disk I/O. The host builds and the Retro68 guest cross-builds cleanly. It is not metal-verified: no new disposable upload was sent to the PowerBook in this slice, so real-volume reservation values, Finder-visible finalization, fork/type/creator fidelity, interruption cleanup, and live throughput remain open.

The reconciliation also exposed two pre-metal hardening gaps. Host byte reservation does not yet cap the number of active stages, so repeated zero-byte or tiny begins can retain bounded-lifetime records without consuming meaningful byte quota. A stage is bound to session and policy version but not to an opaque active-share identity, so a human share change between begin and commit is not yet a typed stale condition. Both must be resolved and tested before staged upload advances to attended PowerBook acceptance.

Invalid persisted guestRoot recovery currently rejects the malformed value, logs the event, and restores the approved share-root default. That is the implemented and tested behavior, but it can broaden a future narrowed policy. Fail-closed recovery versus explicit rebinding remains a policy decision before an Integrations UI can configure narrower roots.

The reverse-streaming prerequisite is now integrated: the guest reads outbound forks one bounded frame at a time, and the host receives into a private disk sink with progress, length/CRC validation, interruption cleanup, and atomic finalization. This does not expose arbitrary download. That capability remains gated on a typed NOW command, root/size policy, deterministic receipts and audit, and an explicit MCP projection. Reverse resume also remains separately deferred pending a contract-first guest source-identity rule.

The combined V0.5 tree—root-scoped capability/list/stat, create-only staged upload, and reverse streaming—has been reconciled and promoted to local main. The read-only commands and reverse transport carry the bounded metal evidence stated above; staged upload is implemented and tested but remains unrun on the PowerBook. This integration did not add download, mkdir, overwrite, move, delete, tree deployment, prune, broad host filesystem access, plugin infrastructure, resume, or transfer-rate hardening.

Mutation is gated separately on guest-side revalidation of an opaque file observation. Listings now carry a responder-generated opaque catalog identity, and the host mints short-lived session/root-bound observation references, but no mutation accepts them yet. The current move/Trash/restore/mkdir messages still act by path alone; host-only precondition checks would permit a changed item to be acted on between check and use. The exact guest-side revalidation field and command behavior remain the next contract-first mutation gate. Tree deployment and mandatory-preview manifest prune follow only after it.

The companion against a partial guest: capability-derived, unverified on metal

NOW has two guests now, and the agent-integration companion was written against one of them. It is now guest-agnostic — but only two of its projections have ever been watched against a guest that implements part of the contract, and neither of those was the new one.

What changed. A twelfth tool, now_session_capabilities, reports what the connected guest can do and therefore which tools are available against it. The derivation has exactly two sources and neither is identity:

  • Commands come from help, which both guests serve on the wire, one fetch per connection. It is the same live source the host console's Tab completion already uses, so a guest that grows a verb becomes usable without a companion release.
  • Message families are not in any command table — that gap is how ps shipped wire-only here — so they are established by asking. Every family request the host makes records its own outcome as it settles, and the report additionally probes the read-only families it can settle cheaply (process.list, file.list). It never probes a family whose smallest request changes the guest (process.quit, file.put), and it probes software.list only on request because that first page is a whole-volume sweep. Those stay unproven, a third state that explicitly does not mean "no" — collapsing it into "no" is how a report would start understating a machine it never asked.

AgentIntegrationCapabilityTests fails the build if any deciding file in the companion surface reads a hello field or names a guest. That guard exists because the same mistake already happened in the other direction: MetalQuitTests derived a guest's abilities from its hello name and went stale the same afternoon that guest grew process.list, understating its own evidence with nothing failing.

The refusal path, which was half-built. GuestListener.recordGuestError claimed to route a guest error to "every waiter" and routed three of the six maps. Process listings, software listings and process results — exactly what a partial guest refuses — still sat on their 15 s and 30 s watchdogs and arrived with timeout instead of the guest's reason. All six are routed now. The mutation that removed three of them reproduced the original symptom: 15 s, 30 s and 15 s waits, each arriving as timeout.

That mutation also exposed a hazard in the first version of the fix: it cleared the watchdog before routing, so a waiter kind the function forgot would have had neither an answer nor a timeout and would have hung forever rather than merely slowly. The watchdog is now cleared only when a waiter was actually answered.

What this does NOT change. No safety property moved. Opaque session-bound references, revalidation before use, one-use receipts, create-only uploads, and the rule that no guest path or PSN crosses the adapter are all as they were. In particular now_request_quit was not made to work against a guest without process.quit: the opaque-reference and PSN-revalidation model has nothing to stand on there, so the tool is unavailable in typed form and that is the whole answer.

Unverified. All of it is tested here — twelve projections, 490 host tests, both xcodebuild configurations — and none of it is metal-verified. Specifically open:

  • No capability report has been taken against the PowerBook 180c. The fake partial guest in the tests answers not-implemented the way now-guest-68k/src/core/wire68.c does, but a fake guest proves the host's half twice and the guest's half not at all.
  • now_list_processes against NOW-68K is the tool this arc claims is newly possible, and it has not been called against that machine. The 68K's process.listing does carry PSNs, so references will be minted there — what happens when one is offered to now_request_quit and the guest refuses process.quit is tested against a fake and unobserved for real.
  • The help command table parse is exercised against a synthetic table. Neither guest's real help output has been fed to the ledger.
  • software.list probing is opt-in on the stated grounds that a guest which does not implement it refuses instantly. That asymmetry is reasoning, not a measurement; the ~4 s figure for a guest that does implement it comes from the earlier 1400c catalog sweeps, not from this code path.
  • The local protocol moved to v6 and the capabilities call gets a 90 s response window because it may wait on several guest-side watchdogs in turn. That number is a sum of the existing bounds, not an observed one.

NOW-68K: what has not been on the machine

The 68K guest for the PowerBook 180c is metal-proven for dial, handshake, keepalive, health, logging, clean quit, launch and the gone path of quit. Everything below has been built and cross-compiled and has never run on a Macintosh — some of it now runs under host-compiled native tests, which is a different and lesser thing, and each entry says which. Listed because "we shipped it and here is what we still do not know" is the useful half.

  • The interactive console is a SECOND WINDOW, by decision, and that is a standing exception rather than drift (2026-07-25). Every other statement this project makes about guest UI says the opposite: the Carbon guest's rule is that a new feature is a Workshop module and never a window (docs/adding-a-workshop-module.md), window.h and this README both describe NOW-68K as one page with no tabs, and now-guest-68k.r's SIZE comment agrees. Michelle asked for the console in its own window on this guest, and it is implemented that way.

The reason it is defensible: the main window's console pane is a log viewer — it shows what the wire and the status line said, it takes no input, and this change leaves it exactly as it was. An interactive console needs a keyboard focus, an edit field, an insertion point and a key-by-key event path, and the one 512×300 page already carries three connection fields, two controls, a status line and a health readout. Making it carry both would mean shrinking the log viewer to a few rows or growing the window past the 180c's 640×480 panel.

The next feature is still a page on the main window unless someone writes down a reason this good. conwin.h's header comment carries the same paragraph so it is read by whoever edits the code, not only by whoever reads the ledger.

  • The console runs the command table, not a copy of it — and only the seam is tested (2026-07-25). commands68.c used to run a command and emit its command.result JSON in one pass, which is fine with one reader and impossible with two. It now fills an N68CmdResult (the facts, no formatting) via now68k_commands_run(), and now-guest-68k/src/commands/n68_cmdresult.c holds both renderers side by side: JSON for the wire, text for the console. Adding a command means one case in now68k_commands_run and nothing else — it appears in both places in the same commit. This is deliberately aimed at the parent corpus finding two-halves-never-met-in-a-test.

What is proven: now-guest-68k/tests/test_cmdresult.c (50 checks) pins the JSON bytes for all three reply shapes against literals written out in full — not assembled from the renderer's own pieces — and walks six outcomes through both renderers asserting they never disagree about the ok bit or the error code. now-guest-shared/tests/console_history_test.c (38 checks — it was now-guest-68k/tests/test_history.c until the history became one file both guests compile) covers the arrow-key history, including the two cases that are wrong in most first attempts: "nothing further that way" must leave the field alone rather than clear it, and a walk must not re-capture a recalled entry as the half-typed line. Both were wrong in the PowerPC guest's own copy, which is why there is now only one.

The wire did not change, and that was checked differentially rather than assumed. A scratch harness ran the old finish_error / finish_ok_row1 / finish_ok_row2, extracted verbatim from 4a7703f, beside the new renderer over 1,092 combinations of reply shape × message × error code × output capacity (512 down to 0, including the caps where the compact fallback fires): 0 differences, in both the bytes and the returned length. The harness first reported 37, which was a real finding — the new N68CmdResult copies the message into a fixed 160-byte member where the old builders took an unbounded pointer, so a message longer than 159 bytes now truncates instead of falling back. That case is structurally unreachable (kDetailCap is defined as kN68CmdTextCap, and every message source is one of those buffers), and it is written down in n68_cmdresult.h rather than left for someone to rediscover.

What is not proven anywhere: that launch and quit behave the same when driven from the console as from the wire. Both paths call the same now68k_commands_run, which is the point of the design, but no test drives the console path (it needs a Toolbox) and no metal run has done it by hand. That is the first thing to check on the machine.

  • The console has never run on the PowerBook. It builds under the 68K toolchain at -O2 -Wall -Wextra -Werror and its Toolbox-free halves pass their native tests; nothing more. Specifically unproven on metal:

  • Up/Down history. The interception happens before TEKey because TextEdit given kUpArrowCharCode/kDownArrowCharCode moves the insertion point between display lines, which is a no-op in a one-line field. That reading is verified-document (Events.h constants read from the installed Universal Interfaces: up 30, down 31), not verified-target.

  • Left/right cursor movement, which is deliberately handed to TEKey rather than reimplemented. Same evidence level.
  • Option-Up/Option-Down scrollback. kPageUpCharCode / kPageDownCharCode (11, 12) are also accepted, but the 180c's built-in keyboard has no dedicated page keys, so Option-arrow is the binding that has to work on the target and it has never been pressed there. Command-arrow was not available: MenuKey in main.c consumes every Command chord first.
  • The two-window event routing. main.c now routes update, activate, click and key events by the window they name rather than assuming one exists. A mistake here does not crash — it draws the wrong window or types into the wrong field — and nothing off-metal catches that.
  • The memory cost — measured at link time, not on the machine. Against 4a7703f built the same way, the console and the seam it needed cost text +4,428 and bss +10,954 = +15,382 bytes, +4.0% of the 384 KB partition (m68k-apple-macos-size over the object files). The BSS is an 8.2 KB scrollback ring plus a 2.3 KB history, beside window.c's existing 9,186 bytes. What that does NOT include, and what nobody has sized: the WindowRecord and the TERec plus its text Handle that the Toolbox allocates out of the application heap when the window is opened. With ~231 KB free that is very probably fine and it has not been watched.

  • The console cannot copy text out, and its scrollback is 32 lines. The output pane is drawn text, not a TERec, so a click in it does nothing and there is no way to get a result off the machine except by reading it. The 32-line ring is n68_console_ring.h's compile-time capacity, shared with the main window's log viewer; Option-arrow paging makes all 32 reachable, but a long quit transcript still ages out. Both are deliberate: a selectable output pane means a second TERec and its text Handle, and a deeper ring is 256 bytes a line.

  • The declined quit — METAL-VERIFIED 2026-07-25. The whole re-list composition exists so a target that stops to ask about an unsaved document answers ok:false / quit-declined rather than claiming success, and it had never run anywhere. On the 180c, against a TeachText holding typed-but-unsaved text: [quit-declined] quit: TeachText is still running - declined, or busy. MetalQuitTests :: testADirtyDocumentDeclinesAndSaysSo (NOW_QUIT_DIRTY=1 NOW_QUIT_NO_LAUNCH=1 NOW_QUIT_APP=TeachText).

Two things the run taught that the design had not:

quit-ambiguous also ran, by accident, and was right. The test launches its victim before quitting it; against a TeachText a human had already opened, that produced a second copy, and quit refused the whole request rather than guess which one was meant. Correct behaviour, never previously exercised — and a test that manufactured the very ambiguity it then failed on. Hence NOW_QUIT_NO_LAUNCH.

The 68K re-check WAS weaker than the sentence "still running" suggests — true when written, and fixed since by process.list. With no process.list, confirmation is a second quit through the same subsystem, and it came back "asked TeachText; not confirmed (wait_ticks <= 0)". The assertion that holds is only that the target did not answer not-running. That is real evidence and it is not corroboration; the run says so in its own output. - The farewell — METAL-VERIFIED 2026-07-25. A menu quit on the 180c produced now-68k is shutting down on the host, which is the bye path; the abortive one reads Connection lost. Metal68KTests :: testTheFarewellIsOrderly (NOW_68K_BYE=1, human at the keyboard, because the guest refuses to quit itself). - The redial — METAL-VERIFIED 2026-07-25. Host dropped mid-session and restarted; the guest redialled and re-helloed in 15.5 s. Metal68KTests :: testTheGuestComesBackAfterTheHostGoesAway (NOW_68K_REDIAL=1; the cadence is human-armed by design, so the checkbox is part of the test's precondition). The reconnect re-handshakes, as the contract requires. - Oversized control frames — now tested, still never sent. The skip-not-fatal path (a frame larger than our 4 KB buffer but inside the protocol's 32 KB) is covered off-metal since 2026-07-25: the reader moved to now-guest-68k/src/core/n68_reader.c behind an ops table, and now-guest-68k/tests/test_reader.c drives it through a scripted transport — the oversized frame is skipped and the next frame still parses, which is the actual claim, under four chunkings plus a stall at every one of ~380 byte offsets. What that does not prove: nothing in NOW has ever sent one. The host does not produce a control frame over 4 KB, so the reader's contract is proven and the host's honouring of it is not. - FIXED 2026-07-25 — launch of a name not on the disk never answered. Watched broken on the 180c three times (60 s, 150 s, 300 s), then watched fixed on the same machine: NOW-68K 0.6 answers in 2.5 s with "nothing named X is on the startup volume". Kept in full below because the diagnosis was wrong twice before it was right, and the wrong turns are the reusable part.

The cause was one limit stated three times in two units, smallest winning: the builder's buffer 512 (a literal in wire68.c), the module's documented floor 320 (commands68.h prose), the outbound slot 160 (sized by a comment reading "hello (~110), ping (~30), or an error reply (~95)" — true when this guest had no commands, never revisited when launch and quit arrived). The reply built correctly at 166 bytes; commands68.c's compact fallback never fired, because from the builder's side nothing was wrong; the slot dropped it. Both numbers now come from commands68.h (NOW68K_COMMAND_RESULT_CAP), +704 bytes BSS.

The original diagnosis, retained:

What the guest's own log says: cmd: launch refused -50, then command.result dropped, outbound queue full. So the search RAN and RETURNED — launch is not hanging — and the reply was built and then thrown away on the way to the wire.

Two theories died on the way to that, both worth keeping because each cost a metal run. (1) The guest goes deaf inside PBCatSearchSync and writes its reply to a socket the host's idle timeout already killed. Refuted: the metal test now watches the wire during the search, and it stayed up for the whole 150 s with keepalives answered — yield_ticks(0) pumps between slices exactly as intended. (2) The reply is too long for the 160-byte outbound slot. "Refuted" by reading commands68.c's compact fallback — and this refutation was itself wrong, which is the lesson worth keeping. The fallback exists and would have fitted; it never ran, because the builder had 512 bytes and succeeded. Reading one half of a size mismatch and concluding the other half is fine is how the mismatch survived in the first place.

What is actually established is narrower: enqueue_control_send refused the payload, and its 0 return covers two different failures — payload too big for a slot, and both slots busy (kWireOutQueueDepth is 2) — which every caller logged with the same sentence. That is why the log could not settle it. 0.5 logged them apart, and the very next run on the machine said it outright: wire: send dropped - payload too big for a slot, bytes 166.

The method note, which is the transferable part. Two theories, two metal runs, both wrong, and the thing that ended it was not a better theory — it was making the log able to tell two causes apart. One message covering two failures is what turned a five-minute question into an hour, and the fix for that was three lines. When a log cannot distinguish the candidates, instrument before theorising again.

  • launch at scale. The catalog search is double-bounded on purpose — a whole-volume Finder search has hard-wedged this fleet before — but it has only resolved an application sitting in an obvious place. The truncation branch is still unproven: the one metal attempt at it never got its answer back (above), so whether the bound reports honestly is exactly as unknown as it was this morning.
  • The confirm wait under load. It yields with an event mask of zero and pumps the wire each pass, with a re-entrancy guard so a command arriving mid-wait cannot recurse into it. Neither the pump nor the guard has been observed under a second concurrent request.

  • error has a fixture and has still never been emitted. (2026-07-25, closing the old "hello, ping and error are not conformance-checked" entry.) All three now have hand-written fixtures in GuestWireFixtureTests, derived by compiling the guest's own emitters with the host cc rather than by reading the C — a fixture written from the decoder's side would test one half twice. hello and ping have also run live against a real host. error has not, anywhere: reaching it needs the host to send a live-state message type NOW-68K does not handle, which nothing does today. The fixture is a claim about send_error_reply, not evidence from a capture. Its negative-id echo is reachable in principle and has never been observed.

Worth correcting in the same breath, because it was written down wrong here: unknown-command and refused are not error shapes on this guest. unknown-command is a command.result error object and refused a census.report outcome; wire68.c routes both away from send_error_reply on purpose, because the wrong envelope leaves a different waiter blocked. The error emitter has one code, not-implemented, in two shapes. command.result is the one message still in the cannot-check set with no fixture at all.

  • Three oddities in the 68K frame reader, found and deliberately not fixed (2026-07-25, during the extraction to n68_reader.c). The extraction was kept pure because the code is metal-proven and no PowerBook was on the LAN to re-verify a behaviour change against; these are the things a fix would have quietly changed. (1) RS_HEADER and RS_BODY return on a short read while RS_SKIP loops and calls take once more — harmless, one no-op call per drained bulk frame, and it is why n68_reader_drain() means "one event-loop pass" rather than "read everything available". (2) handle_control_message's empty-frame branch is dead: the reader short-circuits zero-length control frames before dispatch, so there are two copies of that log string and one cannot fire. (3) frames_in counts skipped frames but not the fatal oversized one, so it means "frames whose header we accepted" rather than "frames received" — probably intended, but the stat's name does not say so.

  • The extraction is METAL-VERIFIED (2026-07-25). It was argued-faithful only — structure, call order and a clean -O2 -Werror build — until NOW-68K 0.4 ran on the 180c: handshake, one guest-driven keepalive answered after the 30 s silence, and a control frame round trip afterwards. Metal68KTests :: testTheWireStillWorksAfterTheReaderExtraction. The version bump is what makes it attributable — 0.3 predates the extraction and the wire carries no other way to tell two builds apart.

Both of the things that were known-wrong here are fixed (2026-07-25):

  • The metal gate no longer reads green when it never ran. Under NOW_METAL=1, the port being held and no Mac dialling in are two distinct failures with distinct messages rather than skips; the only skips left are the opt-ins themselves (NOW_METAL, and NOW_QUIT_DIRTY for the case that needs a human at the keyboard). Guest identity now comes from the hello handshake, and where NOW-68K cannot serve the independent process.list confirmation the run says WEAKER out loud in its output and in every failure string. Watched directly: unset skips 3 clean, NOW_METAL=1 with nothing dialling in fails at 120.1 s, and a deliberately lying guest is caught on both the strong and the weak path. It is tested, not metal-verified — the guests were simulated by tools/fakeguest.py, which is a claim about the harness and never about a guest.

  • The contract's reconnect clause is amended. Cadence is guest policy, capped backoff is the reference default, and the one surviving obligation is a ≥1 s floor between dial attempts. No revision bump: nothing changes shape and an older peer cannot tell. NOW-68K already clamped to the floor; the PowerPC guest reached it only incidentally through a prefs range check and now enforces it at the wire.

vprobe on the 180c: measured, and what it does not cover

vprobe is metal-verified on the PB180c (2026-07-25, NOW-68K 0.16): ran in 3.0 s, whole-frame on every row, answered in one frame, and the wire survived it. Numbers and their reading: vram-readout-68k.md.

The hypothesis it was built around resolved cleanly and in the direction that costs us: MOVEM.L does not burst on this machine (6% over unrolled move.l), and the reread row explains it — the VRAM is uncached, and burst fills are cache-line fills. The unexpected result is a ~16-bit width ceiling: 8→16-bit more than doubles the rate, 16→32-bit buys 12%. The 1400c's "the bus charges per transaction" does not transfer.

Unverified, and worth naming because the numbers will get quoted:

  • The CopyBits row is fifteen banded calls, not one blit. Best raw beats it 1.54×, which is the opposite of the 1400c result — but an unknown share of that gap is per-call overhead. It is a floor on the margin, not the margin.
  • Nothing at a non-native depth was measured. That is precisely where the 1400c's raw-vs-CopyBits margin evaporated, so the one number most likely to mislead a future capture stage is the one not taken.
  • fmove.d is content-dependent — extended conversion, and a 68882 handles denormals slowly — so it is what an FPU reader costs on that screen, not a bus figure.
  • Microseconds() has no availability gate. Its trap is assumed present on 7.1 from documentation; a wrong availability test fails in the wrong direction (disabling vprobe where it works), so none was added. It answered with 37 µs resolution on the 180c, which settles the assumption for this machine and no other.
  • fmovem.x was not measured — no conversion, no exception path, and the one row that might have rescued the FPU result. The reply cannot carry a 17th row; the honest next step if the fmove.d number ever looks wrong.

The capture tx: staged, and crossing the wire on the emulator

Slice two — the pixels reaching the host — works on the Quadra 800 (OS 8.1, 640x480x8, 2026-07-26). The host sends capture.request, the guest stages a PackBits capture, announces capture.begin, streams it down the bulk lane and closes with capture.end ok:true; the host decodes the palette and the packed rows into a pixel-accurate PNG of the guest's screen. 137,783 bytes for a full frame, every byte accounted for (consumed 137783 of 137783), 2.2:1 on that busy desktop. Nothing has run on the 180c, and the emulator's captureMs 16, encodeMs 3 are a 68040 reading host memory — meaningless as predictions, as ever.

Two ways to send exist now and they are not rivals:

  • staged (shotstage68.c) — pack the whole frame to a scratch file in the published root, whose size is then an exact fact, and stream that file through the tested file source. Costs a disk round trip; buys compression. This is the one that crossed.
  • streaming (shotsrc68.c) — read the framebuffer straight down the wire as raw, no staging and no scratch file, at ~300 KB. Built and native-tested; not yet routed, because the staged path answered capture.request first and one lane is one transfer wide.

The header of n68_bytesrc.h argues against staging ("cannot be staged to a temporary file first either"). It was written before anyone had measured a capture, and it is right about 300 KB and wrong about 65 KB. That argument is now answered with numbers rather than overridden.

PICT is not the wire format, and the contract said so first. CaptureBegin.encoding is raw | packbits, described as "NOT PICT: modern macOS cannot decode QuickDraw pictures, so the wire uses a format both sides own". So none of shot68.c's picture machinery is on this path. The stream is the palette as RGB triples, then rows top to bottom — which the host already decodes, because the PowerPC guest already sends it. The envelope is built field for field from now-guest-ppc/src/core/wire.c's, and bytes includes the palette (the contract's one-line description says rowBytes * height; the sender that exists sends GetHandleSize of palette-plus-rows, and the host agrees with the sender).

The pull/push problem, which is why the source reads the screen and not the picture. shot68.c hands the whole frame to QuickDraw in ONE CopyBits that runs for ~480 ms and cannot be suspended. fill() is a pull. There is no way to pull from inside a call that is pushing — no threads, no coroutines, and the banded recording that would have made it incremental is the thing that killed QuickDraw on the third band. So the source reads the framebuffer directly through the shared walk, which is exactly what raw already is. PICT stays the disk format. The two paths meet at the screen and nowhere else.

This rung sends raw, and packbits is blocked on a real constraint, not on effort. n68_bytesrc.h's first promise is that total is exact before the first fill, because capture.begin carries the byte count and the receiver sizes its staging from it. For raw that is arithmetic. For packbits it is not knowable without packing, and this machine cannot hold a packed frame to measure one:

  • the packed frame is not bounded. The 180c's own desktop packs 4.7:1 (65.6 KB), but PackBits expands incompressible data, so the worst case is ~303 KB against a 384 KB partition. "Usually fits" is not a budget.
  • a counting pass then an emitting pass would produce an exact number for a screen that no longer exists. The two passes read the display at different moments, so their lengths can differ — and capture.begin would then be a lie the receiver sized its buffer from. Worse than sending more bytes.

So packbits over this lane needs a decision, not code: either stage the packed frame in a temporary file (whose size IS exact — the 180c wrote 65 KB in ~800 ms, and screenshot already writes that file today), or a contract that can carry a transfer of unknown length. Neither is taken here. The cost of the rung that needs no argument is stated plainly: raw is ~300 KB where packed would be ~65 KB on a quiet screen, and on this machine's wire that difference is the whole user-visible experience.

Two bugs the wire found that no test could have. Both are recorded because both are the same shape — a thing that is only wrong when two real halves meet:

  1. The staged file was written where the sender does not look. Staging put it beside the application; n68_filesrc reads from the published root (the Desktop). The capture staged perfectly, 137,760 bytes, and then could not be found. This is the second time this tree has made exactly this mistake — n68_putfile.h records the send and receive halves briefly disagreeing about the root, and says only a real file system can notice. now68k_desktop_folder() is the one place it is decided and now this uses it too.
  2. capture.begin announced raw while the payload was packbits. The envelope builder was written for the streaming rung and hardcoded the word; the staged rung reused it. Every native test passed — they only ever built raw plans — and the guest sent 137,794 perfectly correct packed bytes under a label telling the host to read 307,968 unpacked ones. The encoding is a parameter now, and a native test pins both spellings.

RESOLVED: the 180c's wire capture arrived garbled — 24-bit addressing

Superseded by the entry at the top of this file, which carries the metal evidence, the fix and what is still unverified. The reasoning recorded here was wrong and is kept because being wrong in this particular way cost two passes.

What was written here: the StripAddress/SwapMMUMode hypothesis "is contradicted by this tree's own metal evidence", because vprobe's fidelity sweep reported MATCH (480 rows) at base 0xFC080000, and because that base "is only reachable with 32-bit addressing on", so the machine must have been in 32-bit mode.

Both halves were true of the session they were measured in and neither generalised. The 180c's PRAM battery is dead, so its Memory control panel setting reverts to 24-bit on every power cycle; the vprobe run and the garbled capture were in different machine states, three days apart. Re-run beside shotdiag, the same sweep reported 480/480 rows DIFFERING. The address does read like a 32-bit one because GetPixBaseAddr returns what QuickDraw knows — QuickDraw is not the thing that truncates it, the CPU is, at the moment of the dereference.

The lesson worth keeping. A measurement retired a hypothesis, and the measurement did not carry the state it depended on. Everything raw this tree measures on a 68K Mac is now reported beside its addressing mode for that reason.

What did change: the walk now has a gate. The framebuffer walk was the one part of the lane no test could reach — it sat between an FSSpec and a ShieldCursor, and the only other reader of that memory (vprobe) merely times it, so a wrong base reads at full speed and every number stays green. n68_shotwire_emit() now owns the walk with no Toolbox in it, shotstage68.c keeps the file, the cursor and the clock behind hooks, and test_shotemit.c drives it over a synthetic framebuffer whose padding is poisoned, decoding the result with the host's PackBits decoder transcribed from CaptureDecoder.swift rather than this guest's own unpacker. Both stride shapes are driven: 640-over-640 (the 180c, where a stride bug is invisible) and 640-over-1024 (the Quadra, where it is not). Mutation-verified — reintroducing the stride confusion fails the Quadra shape and the poison check and leaves the 180c shape green, which is the whole argument for testing both.

What that gate does NOT prove. It proves the arithmetic and the encoding, over memory the host cc can allocate. It cannot say anything about whether sc.base points at the 180c's framebuffer, which is the open question. Tested, not metal-verified.

What is left before metal. Nothing structural — this is a deploy and a run. Worth doing on the 180c specifically because every timing number so far is an emulator's, and because the compression that makes this lane worth having was 4.7:1 there against 2.2:1 here.

One thing that already works and is worth knowing. screenshot followed by put gets pixels to the host today, using two shipped verbs and no new code — as a PICT, which the host cannot render but can store. That is a stopgap, not the lane.

screenshot on NOW-68K: metal-verified, and what it measured

screenshot slice one is implemented on NOW-68K (shot68.c / n68_shot.c, contract-declared already — nothing in contract/asyncapi.yaml changed to add it). It captures the screen, encodes a packed 8-bit PICT, writes it to the guest's own desktop as Screenshot YYYY-MM-DD HH.MM.SS (type PICT, creator ttxt), and returns the measurement rows. No pixels cross the wire; that is slice two and belongs to the bulk-send work.

Metal-verified on the PowerBook 180c (System 7.1, 640x480x8, 4 MB, 2026-07-26) — deployed as a spike (NOW-68K shot 0.14+shot, its own folder and its own dev-settings file so the current build's 5252 was never touched), launched by asking the running build to launch it by path, and driven over the wire on 5050. Three captures: one --no-save and two saves. Both files landed on the guest's desktop with distinct names, and one was pulled back over FTP and decoded here — 640x480, pixelSize 8, 256-entry colour table, and the 180c's own screen, correctly. The capture ran inside the partition with room to spare (the guest reported free=489K max=179K at the time; the capture's ceiling is ~21 KB).

The numbers, which are the point of the slice:

180c (metal) notes
read 187–227 ms matches vprobe's ~200 ms banded CopyBits
pack 431–542 ms the unknown this slice existed to measure
write ~800 ms 65 KB to the internal disk
output 65,648–65,692 B full 640x480x8 frame
ratio 4.7:1

Packing costs about 2.4x the read, not 10x. The worst case in shot68.c was written assuming up to 10x and is therefore conservative by a wide margin: a whole capture is ~1.5 s wall clock, against a ~65 s death timer. And 4.7:1 on a real desktop means a frame is 65 KB, not 300 — which is the number slice two turns on, and it is a far friendlier number than the emulator's 2.2:1 suggested (the emulator's desktop was busier; a real 180c desktop packs better).

The 180c's clock is not set — its PRAM battery is dead, and the 2020 capacitor/battery work is queued for that machine anyway. Both captures were named Screenshot 1904-01-01 23.49.0x, the Mac epoch, which is what GetTime returned. The naming code is doing the right thing with the wrong input. What is worth keeping is that the per-second collision guard is carrying more weight on this machine than it was written for: every session after a restart starts near the same instant, so the tick-stamped fallback — not the timestamp — is what keeps shots from overwriting each other until that battery is replaced.

Also verified on the Quadra 800 emulator (OS 8.1, 640x480x8, 2026-07-25): run from the guest's own console, three captures in one session (--no-save, then two saves), the app survived all three, both files landed with distinct names, and one of them was pulled off the disk image with hfsutils and decoded on the host — 640x480, pixelSize 8, a 256-entry colour table, and pixel-for-pixel the screen at the moment of the command with the cursor shielded out of it. That is the strongest statement available short of hardware: the picture is not merely a file, it is the right picture.

The emulator settled nothing about TIME, and said so at the time. It reported read 0 ms, pack 23 ms, write 8 ms — a 68040 with a host-memory framebuffer. Its 2.2:1 ratio also did not carry: the 180c's own desktop packs to 4.7:1. Both were correctly labelled as proving the code RUNS and produces the right picture, and nothing more; the metal run is what produced numbers.

Still unverified, and named because these are the ones that will bite:

  • Only one screen has been captured, and it was quiet. 4.7:1 is a desktop with two windows on it. A screen full of dithered photographic content will pack far worse, and nothing here establishes a floor.
  • The timing split is a difference of two passes. read_ms is a real banded-CopyBits measurement (vprobe's, on vprobe's band); encode_ms is the recording pass minus the read minus the write, so it carries both passes' noise. On a machine where packing dominates that is fine; if the two ever land close together the number degrades to noise, and it is floored at zero rather than allowed to go negative.
  • 8-bit only, by refusal. A screen at any other depth is declined with a sentence naming the depth. CopyBits would convert for free but the 1400c showed a non-native path eats the whole margin (vram-readout.md), and nobody has measured that here.
  • The capture does not pump the wire. Bounded by arithmetic at ~10 s worst case against the host's ~65 s death timer (kShotWorstCaseMs) — measured at ~1.5 s, so the bound is conservative by ~7x, deliberately, because a pumped event can move a window mid-recording and tear the picture. If a real 180c ever exceeds that bound the fix is to band the PUMP, not the picture.

One refactor rode along with this and is worth naming: vprobe's walk to the framebuffer and its one-band GWorld moved out of vprobe68.c into screen68.c, unchanged, because screenshot needed the same three answers and a second copy of a fail-closed geometry check is one copy that falls behind. vprobe is metal-verified on the 180c; the moved version is not, and the move was verbatim rather than a rewrite, but "verbatim" is a claim about the diff and not about the machine.

The banded recording that had to be abandoned — worth knowing

The first implementation recorded the picture a band at a time into a 640x32 offscreen port, which is the obvious way to bound memory and is what the task was scoped around. On System 8.1 it killed the application on the third band, every time, while QuickDraw was writing that band's colour table. It was bisected on the emulator against the guest's own log:

  • not the file: --no-save (no FSWrite at all) died identically;
  • not the geometry: removing SetOrigin and recording every band at the port's top died identically;
  • not the partition: 2 MB instead of 384 KB died identically;
  • not the put proc: it was entered correctly and had already streamed 6.7 KB across two good bands, and the partial PICT recovered from the disk image decodes as two valid PackBitsRect opcodes.

The cure was to stop banding the destination at all, which the design did not need: a picture being recorded is never drawn into. QuickDraw diverts the bottleneck and hands the source pixels to the put proc, so the destination port supplies only a coordinate space, a depth and a clip — and the Window Manager's colour port is all three for free. One recording CopyBits over the whole frame, one colour table instead of fifteen, ~21 KB ceiling, and none of the above. The root cause inside QuickDraw was never identified; if anyone reopens banded recording, that is the thing to find first.

The 180c, 2026-07-25 evening: everything automated is green

The display came back (a via that wiggled against its pad, found by beeping continuity, reflowed). NOW-68K 0.14 landed itself by handoff and every automated gate passed against it: the reader extraction, ps, the bounded launch search, the redial, the error refusal, the oversized-frame skip with frame sync surviving, two overlapping requests, and quit's whole outcome table including the self-refusal.

quit's confirmation on this guest is now STRONG. With process.list served, a disappearance is re-checked against a different code path instead of by re-asking quit. MetalQuitTests probes for the capability rather than deriving it from the hello name, because deriving it meant the file kept understating its own evidence the moment the guest grew.

Three defects were in the GATES, not the guest, and the worst of them had been reading green:

  • The self-refusal case quit now-guest-68k — the CMake target name — while a deployed build runs as NOW-68K 0.14, its MacBinary name. It asked to quit a process that does not exist, got "nothing named that is running", asserted nothing, and passed. It had never once tested the behaviour it is named for.
  • The harness raced its own teardown: stop() reports .idle while NWListener cancels asynchronously, so the next test in a suite bound a port its predecessor still held. Two of five failing while three passed against the same live guest is a race, not a busy port.
  • The strength banner was printed before the capability probe ran, so it announced the strength assumed rather than the one measured.

Understating evidence is the same species of dishonesty as overstating it, and it is harder to catch because nothing fails.

The console, and what it is not verified to do

  • NOW-68K's interactive console is METAL-VERIFIED (2026-07-25, evening). Watched by a human at the 180c after its display was repaired: ps, help (rendering the shared command table plus the console-local verbs), and up/down history — the last of which had never been observed anywhere, on metal or in an emulator, and was the feature the console was asked for.

The two redraw bugs found in the q800 emulator earlier the same day were one cause: draw_output and draw_input drew without erasing first, and the Window Manager erases only what it newly exposes, so a rectangle the app invalidates itself keeps its old pixels. The command stayed on screen after Return looking unrun — inviting a second Return, which for quit or launch repeats a real action — and clear appeared broken while working perfectly. Neither was reachable by a native test; they are pixels.

  • The console pane cannot be copied out. A click in the output pane does nothing on purpose: it is drawn text, not a TERec. Reading a long result means retyping it. Real gap, not a decision anyone would defend on its merits.

Rough edges

A console line reaching a guest older than line is misread, not refused. Such a guest ignores the field and runs the command bare, so ls Lab:Code lists the share root and says nothing about the path it dropped. The field is additive by the contract's own rules — an unknown field is ignored — and this is the one place that politeness costs honesty. Both guests in this tree read it; the exposure is an older binary still sitting on a machine, which is a realistic state for the PowerBook. Stated beside the field in contract/asyncapi.yaml rather than only here.

No help, no completion. Tab completion is the guest's own answer, so a guest that does not serve help has none — deliberately, because a host-side fallback list is exactly what was removed. It does mean a shell that offers nothing until the guest is updated, and help there answers unknown-command, which reads as an error rather than as "this build is old".

Reverse streaming still needs longer and adversarial metal evidence. The PowerBook ladder now covers direct data-fork and MacBinary pulls through 4 MiB plus cancellation. It does not yet cover a transfer longer than two minutes, a file larger than 4 MiB, source mutation during a pull, or direct guest free-heap measurement.

The build stamp can read a few minutes early. CMake touches build_stamp.c at the END of a build, so the stamp reflects when that file was last compiled rather than when the binary was linked. It has already caused one "is this the build I think it is?" moment, and the verification ritual depends on it. touch now-guest-ppc/src/core/build_stamp.c before a build forces it current.

The wire fixtures are transcribed by hand. GuestWireFixtureTests holds copies of the strings wire.c emits. GuestWireConformanceTests reads the source directly and needs no maintenance, but it cannot reconstruct the three messages built across several snprintf calls (file.listing, file.result, command.result), which is why the hand-written copies exist. They can drift.

The browser stops at 128 rows (kMaxRows) and says so in its status line rather than paging further.

No icons in the browser list. GetIconRef is present on the machine (the type/creator lookup a listing off the wire needs, since it has no file to ask about) and GetIconRefFromTypeInfo is absent. Nothing uses either yet; the list is text-only.


Fidelity sweep, 2026-08-06 — six windows-worth of appearance defects

Appended, not edited, per this file's rule. The full survey — rubric, eleven scored windows, the render/screendump pairs and the reproduction command — is docs/fidelity-sweep-2026-08-06.md. Scores there describe branch claude/fidelity-sweep-2026-08-06 off eb952325, guest build 1bff0bd2ca39, before the icon asset pack grew from 186 to 914 app icons; the sweep is the deliberate A/B baseline for that change.

Everything below is BROKEN rather than unverified: each was seen in a render beside the machine's own pixels for the same moment.

Small system text renders 33% too large, and overruns its window

DisplayReplay.strike(font:size:) returns FontBook.system (Chicago 12) for font 0 and font 1 regardless of the requested size. The Memory control panel draws 251 of its 266 text ops at font 1 size 9; all 251 render at 12, so labels overrun their controls ("Disk Cache size is calculat", "Percent of available mem", "RAM Disk S") and spill past the window's right edge onto the desktop. The Geneva branch three lines below already picks a nearest bundled size; the system branch does not, and there is no bundled system strike below 12 for it to pick. Renderer-side. Fixture qdtrace-drain-sweep-memory.json.

A declared truncation becomes a silent one at the glass

Text records carry len, fullLen and trunc, and the resident sets them: Appearance's description arrives len 64, fullLen 69, trunc true. Neither NOWMirrorContentPlane nor DisplayReplay reads either field, so the render draws 64 characters as if they were the whole string. The capture is honest and the renderer discards the honesty. Renderer-side.

A later repaint pass paints over an earlier one

displayEpoch advances once per ARM and never per repaint pass, so a capture spanning several front/back cycles arrives as one frame with the passes concatenated. Where an application composites from more than one offscreen world, a later pass's full-window blit lands on an earlier pass's content and erases it.

The Sound panel is the clean case and it has a control: the same window on the same build, one repaint pass instead of three, keeps all nine sound-list rows; the three-pass capture renders an empty list box. The rows are present in BOTH composed displays, so this is draw order, not composition — the first version of the gate asserted they were missing and failed, which is how the mechanism got named. Gated by testFlattenedPassesPutAWindowBlitOverTheSoundList. The same signature costs Appearance its first two tab labels and both theme swatches.

Curable on either side: advance the epoch per update pass in the resident, or have the plane keep only the last complete pass.

The Monitors control panel reports nothing at all

Monitors / VGA Display (0x1e91a310) is fully drawn on the screendump, sits entirely inside NOW's window, and produced 0 records on its own port — twice. The second attempt escalated past the front/back cycle to hide/reveal and got 17 records, none of them on the panel's window; only the shared theme world answered. Not occlusion, not a gentle rig: the panel's drawing never reaches the hooked port. It is the only window in the sweep the mirror can say nothing about. Capture-side.

Control glyphs are absent with the content plane alone

Checkbox ticks, radio dots and slider thumbs draw as neutral plates in every panel. This is the honest fallback controlSized prescribes for a small themed blit, and it is scoped: the semantic plane was requested: 0, active: 0 for the whole sweep, and that plane is what knows a blit is a checkbox. The open question is whether the shipping product ever renders a control panel with the content plane alone — if it does, those windows look like this.

MacRoman punctuation drops, and sometimes drops to the wrong glyph

drops from every button title that has one, from all five Scrapbook bullets, ' from "user's guide". Two cases are worse than missing: "Check Disk" renders "Check Disl" and "8/ 6/2026" renders "8/ 6/202(". The wrong-glyph cases mean the bitmap strike's mapping above 0x7F wants checking, not just filling in. Renderer-side.

Not defects — two harness artifacts caught before they were reported

Recorded because a later reader will hit them again. Window titles appeared to never render: the coverage test sizes a captured window from contentSize, a union of everything drawn, which made Scrapbook nearly twice its true width and pushed the centred title off the visible area. The sweep records each window's real rect and composedSweep now uses it. The Finder desktop appeared to render empty: a desktop capture composes onto the canned scene's windows[1], which that helper documents as attaching a display that renders nowhere.

Still unjudged after this sweep

SimpleText with a document and with a dialog, Sherlock 2, the three Finder window views (only the desktop was captured), Get Info, a Standard File dialog, an alert, a pulled-down menu, and ten remaining control panels. Calculator and Key Caps failed to launch through the anchor (-192; and a dropped connection with eleven processes open) rather than being judged.

Key Caps' whole keyboard collapses onto the window origin (2026-08-06)

Added by the same sweep, after a second guest boot. Every one of Key Caps' ~80 key frames arrives as a rect at the origin — [0,0,21,21], [0,0,26,21], [0,0,31,21] … 4691 of them, differing only in width — while its key LABELS arrive correctly spread (pen [12,65], [53,65], [73,65]). The render is a handful of stacked boxes in the corner of an empty window; only the chrome is right.

A missing SetOrigin was the obvious suspect and is NOT the answer in general: kNowContentStateOrigin exists, the emitter handles it, and 1013 origin state ops appear across this sweep's other captures. Key Caps' own capture has zero — all 240 of its state ops are clip changes. The shape fits an application drawing each key through a scratch bitmap swapped under a port whose address never changes, where the recorded coordinates are true of the scratch and meaningless against the window. Confirm that before building on it. Capture-side. Fixture qdtrace-drain-sweep-key-caps.json.

SimpleText and Calculator cannot be launched through the anchor (2026-08-06)

launch err -192 for the application, for a document, and for the Apple Menu Items copy alike, on two separate guest boots. This blocked the single most informative window the sweep wanted — a text editor with a document open — and it is a RIG gap, not a mirror result. Sherlock 2, Note Pad, Stickies, Scrapbook and every control panel launched from the same anchor in the same runs, so it is specific rather than general.

A selected label renders as a solid black bar (2026-08-06)

Captured deliberately, because nothing in the corpus had a selection in it: tools/fidelity-sweep.py --reveal selects an item and --after issues a reflowing resize, without which the Finder never rebuilds its icon-view composite (front/back alone yields 0 text ops).

The Finder draws a selected icon's label by painting its background and writing the text over it in the inverse colour. The paint crosses; the inverse does not. The render draws "System Folder" as a black rectangle with nothing legible in it, beside nine correctly-drawn labels — worse than no highlight at all, because a black bar reads as damage rather than as selection. The replay tracks one fg and one bg per port and draws text in fg; nothing carries the text transfer mode.

Adjacent to but DISTINCT from the skipped-invert work: fixing invert will not fix this label, because this label was never inverted — it was painted. Fixture qdtrace-drain-sweep-finder-selected.json (11 paint ops, 0 invert ops, all ten labels present). Gated by testTheFindersSelectionNeverReachesTheCapture. Renderer-side.

The selection/invert baseline, for the work landing now

The replay skips invert outright ("invert needs destination pixels we do not carry"). The accent-ramp thread showed the corpus could not measure that: a forced-magenta selection colour regenerated all nine committed renders byte-identically, because no capture contained a selection.

This sweep leaves a better before-picture: 157 invert ops (verb: 3) across five captures — Stickies 52, Note Pad 52, Sherlock 2 31, Key Caps 18, Sound 4 — against 34 previously, plus the deliberately-selected Finder capture above. Still missing: a selected LIST row in a capture. Sound is the nearest available case and is already committed — its guest screendump shows SimpleBeep highlighted while sound-1pass.png draws that row unhighlighted, in a capture whose other nine rows are correct.

Desktop icons all carry a mangled file type (2026-08-06, IN FLIGHT)

Recorded here because it lives only in a commit message otherwise (48d6501a), and because it is the reason a large piece of work looks like it did nothing.

The Finder's type pass answers in AppleScript source form — «class APPL», not APPL — and the host takes the first four characters. Every desktop icon therefore carries a type of "«cla" (or "stri"), and since the per-application icon lookup is keyed on creator and type, all 914 extracted application icons miss. The pack landed and the desktop still drew generic glyphs, which reads as an asset-pipeline failure and is not one.

A fix was in flight when this line was written and is not described here as done. If a later reader finds icons resolving, append the date rather than editing this entry.

Two things the content plane produces that nothing consumes (2026-08-06)

Not broken — unwired, which is a different claim and should not be left implicit behind a coverage tick.

  • srcPixmap is emitted by the guest on every blitsrc record, decoded by the host's QDTraceDecode nowhere, and exists only as diagnosis in the JSON. The join uses srcPort alone.
  • The whole qdext status object — installed, calls, newGWorld, lastSelector, foreign, born, died, bornMissed — has no Swift reader at all. It is reachable from the console and over the wire, and it is the only way to distinguish a trap patch that never installed from one that installed and never fired, so a host that renders an empty interior currently cannot say which of those happened.

The second is the one worth closing: it is the diagnostic for the most invasive thing this project installs in a foreign process, and no host face asks for it.

Nothing in the content-plane arc has touched metal (2026-08-06)

Stated as its own line so it cannot be lost among the capability entries above. The ring, the blitsrc join, the $AB1D trap patch, the worldborn/worlddied records, the host's re-homing and nested composition, the measured accent ramps and the 914-icon pack are all emulator-verified on QEMU mac99 / Mac OS 9.1 and nothing more. The resident that writes the ring has never run on a Macintosh, no qdext counter has been read on one, and the live host application has never been watched composing an interior — every composition result comes from replaying committed captures inside tests.

A PowerBook 1400c is far slower than the emulator, and a trap patch on _QDExtensions in a foreign process is the highest-risk thing here. Treat every number in this arc as an emulator number until someone writes a line below this one saying otherwise.

A green suite plus a green suite is not a green tree (2026-08-06)

Recorded as a defect class rather than as a fix, because the fix was nine lines and the class has now cost this session twice.

Two branches were each green alone. One added a second swift test pass with NOW_MIRROR_ASSETS=none, so the degradation path is watched rather than assumed. The other added seven RendererTextFidelityTests that assert against the guest's own font strikes, which live in the asset pack. Neither tree could run the other's half, so the merged tree was the first place the pair could meet — and it met as eighteen failures reading "the renderer is broken" when they meant "the dependency is absent". They now skip by name through the same contract the MirrorKit suite uses: a VISIBLE skip when the pack is absent, a FAILURE when NOW_REQUIRE_ASSET_PACK says a machine that has the pack should be biting (commit d18ab259).

The general shape, which is worth more than this instance: a test's hidden dependency is invisible until something removes it, and a parallel branch is the thing most likely to remove it. A merge of two green branches deserves a full gate run on the merged tree before it is called green, and this closing pass exists partly for that.

The closing pass on the merged tree (2026-08-06)

scripts/test-all was run against the tree with all ten agent branches and the perf thread merged. Status is recorded in the session's handoff (plans/2026-08-06-016-handoff.md), including which stage, if any, needed a --skip and why.

Nothing in this closing pass changed behaviour. Nothing in it is metal-verified, and the live host application has still never been watched composing an interior.

What slice 4 closed, and the two halves it did not (2026-08-07)

Plan 018 slice 4 closed the title and rect defect classes at the point they are created — the PPC walk refuses to publish a title it cannot vouch for, clamps an implausible rect, and makes the live ControlRecord authoritative over a DITL's frozen text. Measured on the emulator against the build under test: Memory 21 pointer titles → 0 (and 18 out-of-port rects → 0), Monitors 13 → 0, Mouse 12 → 0, General Controls 7 → 0, and the application-switcher menu's '\x01\x1f@"\xcf' → omitted. Date & Time was the fifth target and its capture did not complete before the run was stopped; its 6 are unmeasured.

Two things are still open, and both are named here rather than left to be re-derived.

The Finder's items are still not addressable, and the remaining half is host-side. Sweep A read this as a guest-walk defect. It is not: the guest emits no Finder items at all (scene.h declares items[] and desktopItems[] absent by design). They are AppleScript to the Finder, run through the script verb, parsed in NOWMirrorSource.readIcons and flattened by MirrorStateProjectionService. Slice 4 fixed the geometry there — one coordinate space, a real 32x32 target box, no rect at all for an item the Finder did not place. What it did NOT fix:

  • Refs. ref: nil is a deliberate host policy: these are addressed by name, and finderSelect / finderOpen already work that way. If they should carry refs, that is a decision, not a bug fix.
  • List view. NOWMirrorSource.iconItemsScript asks position of every item whatever view the window is in, and in a list view the Finder answers with the SAVED icon grid — which is why sweep A saw ten rows at l ∈ {1,129,257} on a machine drawing a list. The honest fix is to read the window's view and mark the items unplaced when it is not a spatial one (placed: false already means exactly that, and both the hit tester and the projection already honour it). It was not done because the OS 9 Finder's view of window vocabulary has not been measured, and guessing it would put a plausible wrong answer in the one lane that must not have one. One AppleScript against a live guest settles it.

Why the walk is slow, as a reading of the code and NOT a measurement. Slice 4's own cost is negligible and was not measured separately: it adds a scan of each title's bytes (strings already being copied) and four comparisons per rect. The ~1.9 s walk sweep A measured is structural. Per window, scene_walk.c makes three passes over the control chain — read, name, join — and the read pass costs one handle read plus a 296-byte foreign read per control; Memory has 44 controls and 44 dialog items, and the dialog-item walk then reads the item list again and calls the Memory Manager once per static-text item for its live handle. So the cost is dominated by the number of cross-boundary reads, not by anything per-byte. If someone takes perf, that is where to look first — and they should measure before believing this paragraph.

UNVERIFIED: the host render never settles, and nobody knows if it reaches pixels (2026-08-07)

Measured live, on the unmodified tree, by tools/fidelity-live.py: fidelity-live-2026-08-07-a.md. Four traces, 2,099 frames, one boot, one guest.

The render did not settle in any of them — at a five-second quiet threshold, in up to 72 seconds — including the run where nothing was provoked at all. All 122 flicker events are one oscillation: the process-visibility coverage claim flipping stale ↔ partial with a 3.28 s median return, a ~0.3 Hz square wave whose rate barely moves between a run that opened a Finder window and a run that did nothing (0.50–0.62 events/s). baseComplete was false in all 2,099 frames, across sceneGeneration 1→5 and contentGeneration 2→6.

What is NOT known, and is the whole of why this is parked as unverified rather than filed as broken: whether any of it reaches a pixel. The instrument reads the scene documents the renderer draws from, and SceneRenderer.draw is a pure function of those, so a document change is a frame change — but a coverage claim flipping may or may not alter what is painted. Nobody has watched the live window while this oscillation runs.

The companion absence is worth as much: zero window-level flicker on this tree — no hatch flips, no content dropouts, no rectangle owner flips — through a Finder window opening and a Finder view switch. So Michelle's flicker complaint is bounded rather than explained. Three candidates remain, in the report; the cheapest to test is whether the process-visibility oscillation reaches the picture.

Two smaller things found on the way, both closed:

  • Sweep A drove macOS accessibility scripting to open the Mirror and recorded it as a hard floor for a headless agent. It is only half a floor: --open-mirror on argv already does it, and every trace on that page used it. What is genuinely missing is opening the Mirror in a host that is already running — no agent verb, no menu item, no preference. Adding the verb is a 17-place edit across the protocol, registry, docs and six tests, so it is written down rather than done in passing.
  • A sweep target's leftover modal can no longer void the next target's row. tools/fidelity-sweep.py now fingerprints the world per target, cancels-first on dialog items before trying winact close, quits only once the windows are gone, and names a dirty exit on both the row that made it and the row that inherited it. Sweep A lost Date & Time's whole stability row to exactly this, and --quit-after read as success the entire time, because an application holding a modal ignores a quit.

BROKEN: the anchor plane arms, and captures an anchor for NOW alone (2026-08-07)

Found while diagnosing plan 018 slice 3, and it is bigger than that slice. On this tree (91a5e754, guest 59dce8562ad4, ext from the same commit), on a pristine cold boot with nothing touched, every foreign process reports ax_oracle_not_found and the scene carries NOW's own window and nothing else — no Finder, no desktop, no panels:

scene.request x8 over 32 s, immediately after boot
  windows = 1  (New Old World)
  meta.errors = Application Switcher, Control Strip Extension,
                DVD AutoLauncher, FBC Indexing Scheduler,
                Folder Actions, tbt-appe, Finder, tbt-worker
                — all `ax_oracle_not_found`
  mirror -> lifecycle active, capabilities 127, requested 7, active 7

It is not the ten-second lease (the 2026-08-06 entry above). Scenes were asked every 4–5 s, well inside kNowPeekOwnerLeaseTicks, and the verdict word is different: that defect produced now_no_plane, this one produces ax_oracle_not_found — the plane IS armed and no slot claims the partition. It is also not the first-scene claim-before-echo lag: it does not clear over 95 s and twenty scenes.

Measured through both faces, so it is not the instrument. tools/gwprobe.py and the shipped host app (built from this tree, NOW_PREFS_SUFFIX, Mirror window opened) agree: the app's own mirror_read snapshot reports windows: unavailable / not-observed for all eight foreign processes and complete for NOW alone. axtree binds NOW ok.

Why it matters beyond one slice. Plan 018's whole subject — Finder views, control panels, modals — is invisible on a rig in this state, and so is sweep B. Sweep A, on the same commit and the same base image, reported ten windows and 182 elements, so either something on that VM differed or this is intermittent; neither is established, and that uncertainty is the finding. Anyone quoting a window count should check this first, because the failure looks exactly like "the machine has no windows".

The instrument that does not exist. ext/src/now_ext.c counts anchor_event_passes, anchor_full_publishes, anchor_slot_scans and anchor_count on every armed jGNE pass, and nothing on either side can read them — no guest verb, no host field. So "the plane is armed and captures nothing" cannot be told from "the filter never ran in a foreign context" without adding a reader first. That is the next step and it is small.

Slice 3's own findings, and the modal it can now raise (2026-08-07)

Plan 018 defect #7 (unknown-creator modal) is reproducible on demand for the first time — raising-the-unknown-creator-modal.md is the procedure and tools/stage-orphan-doc.py is the one command. Verified twice by QMP screendump on two boots. It is a titleless Finder dBoxProc alert with one OK button, and when the Finder is not front the Notification Manager puts up a second window beside it.

Whether it enters the scene is NOT settled, and the entry above is why: on this rig no Finder window enters the scene at all, so the modal is not a special case and nothing about its window CLASS was measured. Scenes taken with the alert visibly on screen carry NOW's window only — which is the same answer they give with no alert up. Sweep B should re-ask this the moment the anchor plane is seeing foreign processes.

Two smaller things measured on the way:

  • The guest's script verb reports timeout on SUCCESS when the script raises a modal, because the Finder stops answering Apple Events inside ModalDialog. A driver that reads that refusal as a failure will retry an act that already worked.
  • QMP send-key reaches this machine. scripts/spin-up-ppc states that QMP keyboard events do not arrive on mac99 (has-adb=false); that is true of the ADB keyboard, and the profile also attaches -device usb-kbd. A Return through it dismissed the alert on the first try, twice. Worth knowing: it is the cheapest manual override there is, and the notes said it did not exist.
  • A modal in a FRONTMOST Finder starves NOW badly enough to drop the wire — the guest's Connection panel fell back to "Retry in 4 s" and a scene.request timed out at 45 s. Raise modals with NOW frontmost.

The list view's rows: measured, fixed, and driven (2026-08-07)

Closes the second half Lane B named above ("List view"), and closes Michelle's "unable to select items in list view". Lane B was right to stop where it did: the vocabulary it needed had never been measured, and the answer is not the one a modern Finder would have given.

The Finder's view vocabulary, on mac99 / OS 9.1. view of window 1 works; its class renders as «class pvew». The value has no string coercion(view of window 1) as string raises −1700, "Can't make «class pvew» of window 1 … into a string" — but concatenating it renders it as a bare word, and those words are the whole domain:

raw word the view
icon icon view
name the list, confirmed against a screendump of the rows
small icon small-icon view

They are also the setter's vocabulary (set view of window 1 to name). The four-character enum was deliberately NOT pinned: «constant ****icnv», nmev, lisv, ivew, nvew, bvew, sicv and btnv each answered false or silently did nothing, and only «constant ****smic» round-tripped (to small icon). The bare words are what reaches the wire, so the code uses those and the codes are left unmeasured rather than guessed.

The phrasings that did NOT work, recorded because each cost a probe:

  • current view of window 1 — not a term, and osaErr −1753 for the whole script. An unknown term is a COMPILE failure, so a try around it does not catch it and it takes every other phrasing in the same script down with it. Probe one phrasing per script; a batched probe reports a single lie about all of them.
  • properties of window 1 — −1728, "Can't get properties of window 1". There is no property dump to enumerate the vocabulary from.
  • list view / button view — −1728 "Can't get view"; button / buttons — −2753, "not defined". Those are the modern Finder's words. set view of window 1 to list compiles and raises −15279, "Value out of range": list is the class, not the view.
  • A guest script source is capped at 2048 bytes; a batch of more than about five phrasings is refused too-large.

We got the real geometry, not placed: false. bounds of an item is the box the Finder actually DREW, in every view, and its top-left is the old position — so there is nothing position answered that it does not. Measured beside a screendump, Macintosh HD, ten items:

view position of item 1 bounds of item 1
icon 34, 25 34,25,66,57 — position + 32
small icon 35, 25 (a DIFFERENT grid) 16×16
name 2, 42 — the saved icon grid 22,43,38,59 — the row, 19-px pitch

So every producer now asks for bounds and none asks for position (NOWMirrorSource.iconItemsScript, FinderItems.windowsScript), and the size travels with it as Scene.DesktopItem.w/h. HitTester.targetSize is the one place that turns it into a target: the Finder's box, plus the label UNDER an icon-view icon, and nothing added to a list row, whose name is drawn beside it and whose width the Finder does not measure for us. That under-claims deliberately — it costs a click that must land on the icon, where guessing the text width would over-claim into the next row.

Consequently the view is measured and deliberately not carried: bounds made the geometry right in every view, so a view field would have no reader, which is its own defect class. It is written into FinderItems's header for whoever needs it next.

Driven and watched — emulator-verified, mac99 / OS 9.1, build 0011e6584be9, a session-private VM on anchor 1770 / wire 5320. Through a private host (NOW_PREFS_SUFFIX, its own agent socket confirmed by lsof) and its own agent surface, paired with QMP screendumps at every step:

  • Macintosh HD in list view projects ten 16×16 rows at l=22, t = 43, 62, 81, 100, 119, 138, 157, 176, 195, 214 — the rows the guest is drawing. Before this it reported the three-column icon grid.
  • finderSelect selected "Rumpus PRO 2.0" (row 4) and, in a second pass, "TBT-sndbuf-dev" (row 8). Both watched highlighting in the screendump, the right row each time, from nothing selected.
  • Driving View → as Icons through the Mirror turned every rect into a 32×44 box on the icon grid (columns 34/162/290) — so the rects follow the layout the window is drawing rather than a saved one. That is the claim, and it is the one a fixture cannot make.

What this did NOT prove. finderSelect addresses an item by NAME, so the drive above exercises the acquisition and the projection but not the mouse: the path from a POINT in the Mirror window to a row is HitTester.windowItem plus FinderItems.clickPoint, and no agent-socket gesture takes a raw point. Those two are covered by guards built from these same measured numbers and watched failing by mutation, which is a test and not a drive. Someone with the Mirror window in front of them should click a list row and watch it, and that is the last unproven inch.

Item ref stays nil throughout: standing host policy, untouched here.

One incidental corroboration for whoever owns menus. The View menu in the same snapshot reports marked: true for "as List" and for "as Window" and "Sort List" simultaneously — consistent with the sibling finding that mark is a raw byte carrying a submenu ID rather than a checkmark flag. Nothing here reads it.

EXPLAINED, and it was never a contradiction: an application holds an anchor only after it has been FRONTMOST once (2026-08-07)

This closes the blocker recorded as "BROKEN: the anchor plane arms, and captures an anchor for NOW alone" (slice 3, branch claude/018-slice-3). That entry is correct in every measurement it reports and wrong in the conclusion its title draws. It should be read together with this one, and its title read as a symptom rather than as a diagnosis.

The two reports, and why both were right.

Sweep A / lane E Slice 3
windows in the scene 10 1
ax_oracle_not_found 6 processes 8 processes
which ones Control Strip Extension, DVD AutoLauncher, FBC Indexing Scheduler, Folder Actions, tbt-appe, tbt-worker those six plus Finder and Application Switcher
what had happened on that machine thirteen targets opened and driven a cold boot, nothing touched

The failing sets are nested, not different. The six that fail on both machines are the faceless background processes; the two extra on the cold boot are the two an operator would have fronted on the way to doing anything at all.

The mechanism, measured from the resident's own counters. The extension counts every armed pass of its jGNE filter, and until today nothing on either side could read one of those numbers — which is why this survived two sessions. mirror.extension.anchors reports them now. On a pristine cold boot of this tree, after two scene requests:

eventPasses 451   slotScans 1   count 1
slots  [ 0  "New Old World"  age 128 ]

Four hundred and fifty-one armed passes, and one slot-table scan — the fast path in capture_anchor skips the scan when A5 is unchanged, so slotScans is a count of context CHANGES the filter noticed. It noticed none. The filter is not failing to capture foreign processes; it is never running inside one.

Then one activate of the Finder, and nothing else:

eventPasses 813   slotScans 2   fullPublishes 125   count 2
slots  [ 0  "New Old World"  age 1485
         1  "Finder"         age 127 ]

The scene went from 1 window to 2 in the same moment, and the Finder left the error list. That is the flip, on demand, in both instruments at once.

Three things this rules out, each of which had been proposed:

  • Not the resident. Both boots report capabilities 127, sourceManifest 18d732487b03…, buildFingerprint 0a91ea49abcd…. Same extension, byte for byte.
  • Not the stage image. scripts/spin-up-ppc defaults to os91-runner.qcow2 and stages this checkout's ext and app into the clone before cold-booting, so neither run ever read now-mirror-stage.qcow2. (That image is separately uncertified — see below — but it was not the oracle for anything here.)
  • Not a clock, and not an age. now_ax_bind_process passes max_age_ticks = 0, so age never refuses a slot on this path; and the machine had been up ~90 seconds, nowhere near a TickCount wrap. A real 64-bit/32-bit defect in this arithmetic was found and fixed today in mirror_anchor.c, and it is host-test-only: unsigned long is 32 bits on the guest, so the guest's own arithmetic was always right.

The acquisition is one-time and the slot then persists. Re-fronting NOW and polling six scenes over 48 s left the Finder anchored throughout. So this is a latch, not an oscillation — which is also why slice 3's 95-second wait did not clear it and never would have.

Why fronting is what does it. The plane is armed only for the duration of NOW's own scene walk (requested reads 0 between requests and 7 during one), the walk pumps, and the pump yields. A frontmost application is scheduled promptly inside that yield; a background one with nothing to do is not. So the armed window samples whatever is front, and the six faceless processes never acquire a slot on any machine because they never pump events at all.

What is actually broken, stated as a product defect

The Mirror can only show you a machine you have already driven. A person who opens the Mirror on a freshly booted Macintosh sees NOW's own window and is told, honestly but uselessly, that every other process could not be observed. Nothing recovers on its own. The fix is not in this entry — the plane needs to be armed across a window in which other processes are given time, or the arm needs to outlive the walk — but the diagnosis is no longer the missing part.

Two smaller things worth keeping:

  • ax_oracle_not_found is one word for three defects, and this cost two sessions. now_ax_bind_process returns kNowPeekReadNoAnchor when GetProcessInformation fails, when the partition reads as zero, AND when the oracle finds no slot. The 2026-08-06 entry above already learned this lesson once, in its own last paragraph, about no-plane versus not-observed. It wants closing properly.
  • now-mirror-stage.qcow2 is uncertified. Its sha256 is c466baa9a545…; the newest receipt in ext/stage-receipts.json names 0785871ab82a…, and the file was written at 01:58 on 2026-08-06, after that 01:19 receipt, displacing a backup named .bak-20260806-4-dirty. So no receipt describes the resident that any sweep setting NOW_SPIN_BASE to it is cloning. Nothing in plan 018 used it, so nothing here is void — but the next thing that does should bake first.

The verdict on the arc's measurements

  • Sweep A's scores are NOT void. They were taken on a machine whose targets had each been opened and driven, which is the state the product path produces, and its rig table says so. They describe a driven machine and should be quoted as such.
  • Slice 3's measurements are NOT void either. Every number in that entry is reproducible; this one adds the variable it was missing.
  • Sweep B may proceed, with one condition that its procedure already satisfies by construction: each target must be opened or fronted before it is measured. A sweep that measured a target it had not yet driven would be measuring this defect and calling it a renderer score.

Stating the rule exactly, because two nearly-right versions of it point at fixes that cannot work

The rule is not "frontmost after NOW arms", and it is not "a context change the filter observes". Both were proposed while this entry was being written, and both are close enough to sound settled.

A process acquires a slot when it itself executes a GetNextEvent or WaitNextEvent at a moment when arm_request has the anchors bit set. now_ext_gne_apply runs in whatever process is pumping, and capture_anchor (ext/src/now_ext.c:184, gated at :279) reads that caller's own LMGetCurrentA5(). Front-ness appears nowhere in it. Being frontmost is the PROXIMATE CAUSE via scheduling — a front application is scheduled promptly inside the short armed window and a background one with nothing to do is not — and the distinction matters, because "frontmost" invites a fix that watches the front process, which is the one process already covered.

And there is no single arming moment. arm_request is set and cleared repeatedly: mirror reads requested: 0 between scene requests and requested: 7 during one, in every capture taken here. Any fix reasoning about "after arming" as a boundary is reasoning about a latch that does not exist.

anchor_slot_scans counts TRANSITIONS, not distinct contexts. gLastA5 is a single global in resident BSS (now_ext.c:90) and is never reset, so the counter increments whenever this armed pass's A5 differs from the previous armed pass's. The cold boot's value of 1 is the very first pass — gLastA5 starts at 0, NOW's A5 does not match, one scan — followed by 450 consecutive matches. Two processes alternating would climb it steadily. Read as "how many distinct processes has the filter seen" it is wrong, and it is wrong in the flattering direction.

THE OPEN QUESTION, and it reframes the defect

After the flip, anchor_slot_scans stayed at 2 across 362 further armed passes, while NOW's slot aged to 1485 ticks and the Finder's stayed at 127. Since the counter would climb on every alternation, essentially EVERY armed pass in that window was in the Finder's context — and NOW contributed almost none, even though NOW is the process that arms the plane and serves the scene requests.

Nothing measured here explains that. It is written down as the open question it is, because it changes what the defect is: the armed window is spent wherever the CPU is, not where the arming happens. A design that assumes the arming process is among the sampled ones is assuming something this machine did not do.

"Rescan at arm time" cannot work

Worth stating in those words, because it is the first fix anyone proposes. A rescan runs in NOW's own context and re-samples the one process that is already covered. find_anchor_slot is not the bottleneck — anchor_event_passes is, and passes only happen where the CPU is. The fix has to make the plane armed across a window in which OTHER processes are scheduled, or find a hook that fires in a foreign context.

This also makes the Process Manager enumeration work complementary rather than duplicative: enumeration can say a process exists, but only something running inside that process can capture its A5, its window list and its CurApName. Those are two different halves of one answer.

The cheap experiment nobody has run

Launch a background-only process and read mirror.extension.anchors. tbt-appe and tbt-worker are launched background processes and never acquire a slot on any machine here — but they are faceless and may simply never call GetNextEvent at all, which is a different case with a different fix. The measurements in this entry cannot separate:

  • never acquires because it never pumps — nothing can help it, and the honest answer is that such a process has no observable interface; from
  • never acquires because it is never scheduled while armed — which is the same defect as the Finder's and is fixed by whatever fixes that.

Application Switcher is no help as a specimen: it left the process list entirely after the activate.

Two further things NOT tested here, and inferred rather than measured: that fronting several applications in turn accumulates all of them (the mechanism plus the observed persistence say yes, and Sweep A's six anchored processes are exactly its six driven ones — but no app cycle was ever driven deliberately), and whether an application launched after an arm acquires without being fronted, which in practice is hard to separate because a launched application usually takes the front by itself.

Verification level: emulator-verified (QEMU mac99, OS 9.1, run dirs /private/tmp/nowvm-vis18 and -vis18b, anchor 1810 / wire 5360, guest builds 59dce8562ad4 and f3db46a66630, both asserted on the hello). Nothing here touched metal.

WORKED AROUND, not fixed: cycle makes an undriven machine visible, and three things it taught (2026-08-07)

This continues the entry above — "an application holds an anchor only after it has been FRONTMOST once" — and the exact rule the visibility lane graduated from it. Nothing here changes that diagnosis. What it adds is a control a person or an agent invokes, and three measurements that narrow what the deep fix has to do.

The control

cycle, on both faces (command.request {name: "cycle"} and the console verb). It arms the anchor plane under its own lease owner and holds it armed for the whole cycle, brings each faced application forward in turn so it pumps its own event loop once, and restores the application that was frontmost. It never runs automatically and it announces itself, because it visibly disturbs the machine.

The ordering is the design, and it is the part a naive version gets wrong: arm_request is set and cleared repeatedly, so a cycle that fronted applications while the plane happened to be dark would acquire nothing and would look like it worked. That is why the result is the resident's own before/after counters rather than a success flag — eventPasses rising while slotScans does not is exactly the signature of that failure, and a consumer can tell it apart.

Measured on a freshly booted, undriven QEMU mac99 guest (OS 9.1, build 5ef5f1852bd1, run dir /private/tmp/nowvm-acq18b, anchor 1860 / wire 5410), with scripts/spin-up-ppc's own clone:

before cycle after
scene windows 1 (NOW's own) 2
ax_oracle_not_found 8 processes, incl. Finder and Application Switcher 5, all faceless
anchor count 1 2
slotScans 1 2
slots New Old World New Old World, Finder

The report: considered 7, alreadyAnchored 0, fronted 1, acquired 1, refused 0, vanished 0, backgroundOnly 6, complete true, restored true. A QMP screendump after it shows NOW frontmost with its own window, exactly as before — the Finder came forward and went back.

Cost, measured rather than estimated. A repeat cycle on the same settled machine costs 1–48 ms and fronts nothing, because a process holding an anchor is never a candidate. The first cycle pays one SetFrontProcess and up to 0.75 s per un-anchored faced application, under a 15 s ceiling for the whole run. Nothing was added to the scene path (see below), so a host that never asks for a cycle pays nothing.

Three things this measured

1. WakeUpProcess is not enough, and this is the useful negative. The obvious invisible cure — make every un-anchored process eligible for time without fronting it — was built, shipped into the scene path, and then removed after it was measured. On the undriven guest it woke eight processes, every call returned noErr, and half a second later slotScans had not moved: not one of them had executed a GetNextEvent. Making a process eligible is not making it pump. The code survives (now-guest-ppc/src/peek/anchor_acquire.h) as the cycle's first pass, where it is free, but no path that runs on a host poll pays for it any more. Anyone reaching for this again should know it has been tried.

2. Fronting DOES acquire, and the Appearance non-reproduction. A sibling lane reported that opening and fronting the Appearance control panel acquired nothing. On this rig and this build it did not reproduce: Appearance was opened from the guest (script, via the Finder), appeared in process.list as an application, and took slot 3 the moment the plane was armed while it was frontmostfronted 0, because it never needed bringing forward. Its window then appeared in the scene, titled Appearance, with controls: 0.

So the discriminator is not "control panels do not pump". The remaining candidate is the one the ordering above defends against: a scene.request arms the plane briefly, and a brief armed window need not overlap the target's pump. cycle holds it armed and yields repeatedly, which is why the same machine answers.

controls: 0 on a window that visibly has six tabs is a separate finding and is not this entry's.

3. The faceless set is stable and is out of scope by declaration. Six processes — Control Strip Extension, DVD AutoLauncher, FBC Indexing Scheduler, Folder Actions, tbt-appe, tbt-worker — carry modeOnlyBackground and have no window to bring forward. The cycle reads that bit from GetProcessInformation, the same bit process.list reads, and reports them as backgroundOnly rather than discovering them by failing. That matters: counting them as failures would make an honest cycle read as a broken one forever. Anything the cycle genuinely could not reach is named in unreached (bounded at eight, remainder counted) and must be read as UNKNOWN, never as empty.

The Application Switcher comes and goes: present in one ps and absent in the next, minutes apart, on an untouched machine. A cycle's considered count will differ from a ps taken moments earlier for that reason alone.

What the deep fix still has to do

Unchanged, and now with one option struck off. The plane must be armed across a window in which other processes are actually scheduled, or a hook must fire in a foreign context. WakeUpProcess is not that hook (finding 1). A rescan at arm time is not either — it runs in NOW's own context and re-samples the one process already covered.

The open question from the visibility lane stands and nothing here explains it: after the flip, slotScans stayed at 2 across 362 further armed passes while essentially every pass was in the Finder's context and NOW contributed almost none — even though NOW is the process arming the plane and serving the requests. Understanding that is probably the lever.

Verification level: emulator-verified. Nothing here touched metal, and no claim above is made about a real Macintosh.

One thing is NOT verified and is called out rather than buried: the cycle verb is not declared in contract/asyncapi.yaml, because contract changes serialise through the human. Until it is, CommandParityTests.testNeitherGuestInventsCommandsTheContractDoesNotDeclare and CommandRegistryTests.testTheThreeHalvesAgreeOnTheCommandSet fail, naming cycle — correctly, and they are the only two failures in the host suite.

2026-08-07 — the CDEF route classified 71 of 73, and a tab, a scroll bar and a list row were driven

Slice 19, and it closes the reading slice 18 left open: the Control Manager will not name a foreign control, and that was never the reason those controls could not be used.

Verification level: emulator-verified. A fresh session-private clone of the plain base (os91-runner.qcow2, sha256 f34f7e5d…) with this tree's ext and app staged; guest build ec03d8901bee; lane block 593, anchor 16744 / wire 16745. Nothing here touched metal, and no claim below is made about a real Macintosh. Every capture discarded a warm-up scene.

What the CDEF route answered

A control's contrlDefProc is a Handle to its loaded definition function. A CDEF supplied by the System file carries sysheap, so it is a resource in every process's resource chain and the Resource Manager will name it — type CDEF, and an ID. procID = 16 * id + variant is the Control Manager's own arithmetic and every ID is a *Proc constant in ControlDefinitions.h, so the id-to-kind table is documentation rather than inference.

Appearance, 73 controls, measured before and after:

before (slice 18) after
classified 2 71
unclassified 71 2

The 71 break down as 29 staticText, 16 pushButton, 10 popupMenu, 9 userPane, 2 scrollBar, 2 editText, 2 listBox and 1 tab — the tab strip being the control the whole slice was about. The two that stay unknown stay unknown.

It travels as its own knowledge level, derived, with provenance guest-cdef-resource. That is deliberately not known: known is the control answering about itself through kControlKindTag, and this is us reading the identity of the code that draws it. A caller that needs the stronger claim tests for known explicitly. MirrorKit's authorizesAction accepts derived — the bar it enforces is "the machine said so", not "the strongest possible source said so", and presentation-inference stays excluded by name.

A CDEF id this guest will not attribute produces nothing. Several documented ids are deliberately absent (slider, clock, placard, icon, picture, separator, little and chasing arrows, popup arrow, radio group, scroll text box) because the role vocabulary has no honest word for them. So is any variant a header does not declare — CDEF 0 and CDEF 23 are the button FAMILY, and returning "button" for a check box would authorise the wrong act on the right control.

Two things stopped a control being driven, and neither was the kind

1. kNowAxResolveMaxControls was 32. A control past that bound cannot be resolved, so no reference is minted for it and the scene reports it addressable by nothing. Appearance's chain is 73 long and the tab strip is number 71. Measured: 41 of that desktop's 82 controls carried a reference; with the bound raised to kNowSceneMaxControls (96) it is 73 of 73 in that one window. What a scene can CARRY and what an act can REACH were two independent constants and the smaller one silently decided the product's drivability — the same shape as the control-frame cap that once lived in three places.

2. ctlact pressed the centre. A tab strip is one control and eight tabs; a list box is one control and every row. ctlact now takes h and v in global screen coordinates, checked against the rect the resolver just proved and refused rather than clamped. The act cell already carried click_h/click_v, so no peek_table.h change and no re-bake.

part: 0 gains a stated meaning: answer TrackControl with nothing, so the application's own handling decides from where the click landed.

What was driven, and watched

  • A tab switched. Appearance went Themes → Desktop, and back, by a click at (438,110) and (206,110). Watched in a screendump; the control's own value read 1 → 3 → 4.
  • A scroll bar scrolled. The Themes tab's horizontal bar, driven at (570,305): the guest's own re-read went 20 → 19 — the arrow the point landed on, not the one a part code would have named.
  • A list row selected. The Desktop tab's Patterns list, driven at (504,193): selection went from "Lime" to "Lollipop 2", the highlight moved, and the panel's own label read "Pattern: Lollipop 2, 128 X 128, 64K".

Three things this run found and did NOT close

  • ctlact part 0 reported act-not-taken while the tab was switching. An Appearance-era tab is handled by the Appearance Manager's own click path and never reaches the TrackControl trap, so no patch is consulted. The message was literally true and the verdict was a lie. Part 0 now reports what happened, with the control's value before and after as the evidence — but that means part 0 cannot prove an act was taken for a control with no range. Its honest claim is "a real click was posted inside a control this Mac revalidated".
  • Re-read value answered "the anchor plane is absent or not armed" on every act against Appearance, while the same connection's scenes read that panel's 73 controls without difficulty. So the post-act re-resolve takes a path the scene walk does not. Unexplained.
  • A quoted AppleScript still differs between the two faces. With the console fix below, script 1 + 1 answers 2 and script return (ASCII character 104) answers "h" from the keyboard — but script return "hi" answers osaErr -1753 there while the identical wire call, sent as line or as source, answers "hi". So the guest decodes a quoted line correctly and something between the console and that decode does not. Narrowed to inputs the console writer escapes; not closed.

And one parity defect that was one place, not one verb

console_model_dispatch passed NULL for the request to now_command_run, so every verb without a console-local special case reached its handler with no arguments at all. script tell application "Finder" to activate and ctlact <element> <part> — both exactly as their own help prints them — answered "requires source" and "requires part" while the identical wire calls worked. CommandParityTests cannot see this: the verb is present on both faces and merely broken on one. The raw rest-of-line now becomes line, the field the contract already declares, read by the grammar in cmd_line.h that was always waiting for it.

2026-08-07 — the CDEF route classified a check box as a push button, and the variant that would have said so is not readable

Slice A of plan 019, and it does not close what it was sent to close. It corrects the record instead, which is worth more: the slice's brief said Memory's radio buttons came back CDEF 0, variant 0 and drew as bare labels. The first half is nearly right and the second is wrong, and the difference is the finding.

They did not come back unclassified. They came back classified as push buttonsknowledge: derived, kind: pushButton, action: press, no state — and so did every check box and radio button in every OS 9 control panel measured. A wrong kind reads exactly like a right one, so this shipped in slice 18's 71-of-73 and was counted as a success.

Verification level: emulator-verified. A session-private clone of os91-runner.qcow2 (sha256 f34f7e5d…) with this tree's ext and app staged; lane block 356, anchor 14848 / wire 14849; guest builds 217050dea748 (before) and 774b6fa7acea (after); warm-up scene discarded on every capture. Nothing here touched metal.

What was measured

Memory (44 controls) and Date & Time (21), every control:

  • Every button-family control reports contrlDefProc = 0x00002EC8, which the Resource Manager names CDEF 23 — the Appearance button family — with a zero high byte. Memory's "Save contents to" (a check box), its three On/Off pairs and its Custom/Default setting pair (radio buttons) are byte-identical in that field to "Use Defaults" (a push button).
  • GetControlVariant answered 0 for all 65.

The classic packing of the variation code into that high byte — which cdef_resolver.c's mask path exists to undo — is not what Mac OS 9 does with these controls.

The second line proves nothing on its own, because 0 is also what a declined call returns. So it was asked of controls this application created, whose variants are in this repository's own source: NOW's two checkBoxProc boxes answered 1, its kControlScrollBarLiveProc bar and its auto-toggle triangle answered 2, its push buttons and popup 0. The accessor works on this runtime and is right every time about a control we own. It is FOREIGN controls it cannot answer for. (A probe's control must be code you own; a negative without one is void.)

What changed

CDEF 0 and CDEF 23 now attribute nothing, variant 0 included. Reading zero from a field that is zero for all three kinds is not evidence of a push button; it is the absence of evidence wearing the push button's number.

before after
Memory, classified of 44 33 23
Date & Time, classified of 21 19 10

Nothing else moved. Date & Time's one bevel button (CDEF 2) still reads pushButton: a bevel button's three variants are three bevel depths, so the family IS the answer there.

The cost is real. Semantics.authorizesAction requires known or derived, so a driver that honours it will now decline these where it used to press them. ctlact with an explicit point still reaches them — the guest checks the point against the rect the resolver proved and never consults the kind — so what is lost is semantic authority, not the mechanism. The trade: an unknown a driver declines is a gap someone can close; a pushButton that is really a check box is a gap nobody will ever look for, and its state is never reported at all.

And the refusal is now diagnosable

IR v2 gains semantic.cdef — the resource id the Resource Manager named, beside unknown, only where kind is absent. Contract first, then the guest emitter and MirrorKit's Semantics. Without it these controls arrive as bare unknown, which is the flattening the reason field was added to stop, arriving by a different route.

It paid for itself on the first capture. Memory's 21 unknown controls are no longer anonymous: cdef: 23 × 8 (the button family), cdef: 20 × 4 (icon), cdef: 9 × 4 (separator line), cdef: 6 × 2 ("Cache Up/Down Controls", "VM Up/Down Control" — little arrows), cdef: 3 × 1 ("RAM Disk Slider"). Every one matches the control's own title. Those are ids this product deliberately has no role for, and now they can be counted rather than guessed at.

It is a fact, not a kind. A receiver that maps id 23 to "push button" has re-created the defect it records.

What this does NOT fix, named so nobody searches here again

The bare-label rendering symptom is a separate defect and this change makes its population larger, not smaller. Memory, General Controls and Date & Time were all classified identically before this change — every button-family control pushButton — and only Date & Time drew its check boxes. Identical input, different output, so the divergence was never in classification. It is downstream, in what the renderer has to draw an unknown control from.

Still open

  • Why GetControlVariant cannot answer for a foreign control is not established — only that it cannot. It may be reading a field the Appearance-era Control Manager no longer fills, or CarbonLib may be declining a ControlRef it did not mint. The two have the same consequence here and different consequences for anything that wants the variant by another route.
  • Whether ANY route reaches a foreign control's variant. GetControlData(kControlKindTag) was already measured at 0 of 21 for Date & Time, GetControlKind is Mac OS X only, and the two routes above are now closed. If one exists it is in the Appearance Manager's private per-control data, and nothing here has looked.
  • Appearance's own numbers were not re-measured after the change. The reference registry holds 96 and three panels open at once exhausted it, so Appearance walked with zero controls. Its before-figures (16 pushButton of 71 classified) come from slice 18.

2026-08-07 — three bare-widget symptoms, told apart, and only one of them is classification

Follow-up to the entry above, prompted by two coordination reports whose evidence pointed AWAY from where this slice was sent. They were right to, and separating the three is the useful result.

Verification level: emulator-verified. Same rig discipline as above; second machine, lane block 350, anchor 14800 / wire 14801, guest build 23fb3bc8f88b, warm-up scene discarded. Extensions Manager and General Controls captured for the first time.

The refutation first

Memory, General Controls and Date & Time were classified identically before this slice's change — every button-family control derived / pushButton — and only Date & Time drew its check boxes. Identical input, different output. Classification cannot be the cause of a divergence it does not contain, and this slice's change makes the unknown population LARGER rather than smaller.

The three symptoms, and where each one lives

1. Widget absent — Memory, General Controls. The radio buttons and check boxes are real controls, walked, and now honestly unknown with cdef: 23. What the renderer does with an unknown control is downstream of this slice and is where this symptom lives.

2. Widget drawn, state missing — Extensions Manager's rows. Not this route, and not reachable by it. That window's whole control chain is six controls: 2 scroll bars, 3 push buttons (cdef: 0) and 1 popup. There is no list control and there are no per-row controls. The rows sit inside dialog item 4, a userItem spanning (14,65)-(458,265) — a rectangle the application draws itself. So a per-row check box has no ControlRecord and no DITL row, and its state cannot come from the control plane or the dialog plane. It can only come from the display plane, and nothing reads it today.

3. Widgets "missing from the walk" — Extensions Manager's help button. This one is a reading error rather than a gap. The help button IS in the scene: dialogItems[1], title ?, knowledge: unknown, definition: system, with a ref. Extensions Manager is a DLOG, so it publishes 28 dialog items beside its 6 controls, and a consumer that reads controls alone sees 6 widgets where the window has more. "5 walked against ~8 on screen" is that consumer, not that walk.

And a mitigation the cost estimate above did not account for

Extensions Manager's three push buttons appear twice — as controls (now unknown, cdef: 0) and as dialog items 6/7/8, where the DITL route already reports them knowledge: known, kind: pushButton, action: press, provenance guest-ditl. For dialog windows the stronger answer was never coming from the CDEF route at all, so this slice costs those windows nothing. The cost lands on non-dialog windows and on controls no DITL row covers — which is what Memory and General Controls are.

One corroboration worth keeping

Extensions Manager's buttons are cdef: 0 (the CLASSIC family) while Memory, General Controls and Date & Time are cdef: 23 (the Appearance family). Both are live on the same System at the same moment. A change that refused one family and kept attributing the other would have been honest about one control panel and confidently wrong about the next.

The ids the route DOES resolve are worth reading as its own validation. Every one matches the control's own title without ever having seen it: General Controls' "Insertion Point Slider" and "Menu Blink Slider" are cdef: 3 (slider), its "Launcher Picture" and "Hide Desktop Picture" are cdef: 19 (picture); Memory's "RAM Disk Slider" is cdef: 3, its three "… Separator" controls are cdef: 9 (separator line), and its "Cache Up/Down Controls" and "VM Up/Down Control" are cdef: 6 (little arrows). All are ids this product deliberately has no role for — so they stay unknown, and are now countable rather than anonymous.

2026-08-07 — a shared worktree was checked out from under a running lane, and its commits landed on other lanes' branches

Recorded because it is invisible while it happens and expensive afterwards, and because nothing in this repository warns about it.

now/.claude/worktrees/keen-clarke-4988fc was being used by several sessions at once. This lane branched correctly (git checkout -b claude/019-cdef-memory-radios, 13:02) and then made three commits. Its HEAD reflog:

13:02  checkout: ... to claude/019-cdef-memory-radios
13:14  checkout: moving from claude/019-cdef-memory-radios
                 to claude/019-ctlact-settlement      <- not this lane
13:17  commit: diag(observe): ...                     <- this lane's
13:22  commit: fix(scene): ...                        <- this lane's
13:22  checkout: ... to claude/019-asset-packs         <- not this lane
13:28  commit: test+docs: ...                          <- this lane's

Two commits landed on claude/019-ctlact-settlement and one on claude/019-asset-packs, and claude/019-cdef-memory-radios stayed at the commit it was cut from. Nothing failed. git log --oneline -1 after each commit showed the expected hash; the branch guard was satisfied, because the branch really was not main.

The visible damage was a RED branch that nobody had broken: claude/019-ctlact-settlement received an emitter change without the test update that belonged with it, because the two were split across a checkout that happened between them. It reads exactly like a careless commit and was not one.

Three things follow, and the first is the one that costs nothing:

  • A lane's own worktree is not a nicety, it is the isolation. git worktree add -b <branch> <path> <base> and then git -C <path> for everything. Branch-per-thread does not isolate anything if the worktree is shared, because a branch is a property of the worktree and a neighbour can move it.
  • Check the branch immediately before every commit, not once at the start. git -C <path> branch --show-current is free and is the only thing that would have caught this at 13:17 instead of at 13:35.
  • git stash is repo-global and is not safe here. A stash push followed by a pop can return a NEIGHBOUR's stash into your tree if one arrived between them. Save a patch to a file instead.

Recovery was clean because every commit still existed: a private worktree, a branch off the intended base, and git cherry-pick -x of the three by hash. The other lanes' branches were left alone — rewriting a branch another session is actively committing to is worse than the mess it would tidy.

2026-08-08 — round 11's ledger entry: both census fixes are in one tree, and the file two lanes could not read is now landed

Round 11 merged three lanes into claude/024-integration-11 off claude/024-integration-10. scripts/test-all exited 0; the host suite ran twice inside it at 1902 tests / 0 failures (54 then 72 skipped — one command, two configurations, not variance), and a third standalone scripts/test-host also read 1902 / 0. TESTED, not metal-verified, and stage 6 skipped by design: nothing in this round reached a Macintosh.

What landed

  • claude/024-bake-volume-clean (7 commits) — and it carries claude/024-census-crash as a first-parent ancestor, which is the fact worth recording. The two census fixes were never separable: git merge-base --is-ancestor answered yes, so one merge put the pccard trap-existence check and the ata ataPBFlags fix in the same tree, and a second merge of 024-census-crash would have been a no-op against a lane that had already taken it. Merging both blindly was the hazard the brief named and the ancestry check is what answers it.
  • claude/024-items-arbitration (3 commits) — SceneRenderer's Finder-roster arbitration, split cell: icon box into semanticFrames, name yielded to the machine's own text run.
  • claude/019-drive-findings (7 commits) — docs/lane-context.md, the drive-defects plan, the integration-suite spec, tools/guest-ata/.

claude/025-drive-defects was deliberately not merged: it is another session's live work on SceneRenderer.swift, in flight from integration-10. Landing this round is the only channel by which the arbitration change reaches it, since an independent session cannot be messaged.

The census gate reported six files gone and none of them was a loss

tools/merge-census-gate pending refused nothing across all three merges — 0 dropped, 0 imported self-reverts — but the third printed 54 removed-by-a-side and 6 files gone, all of them the pixel islands (PixelIsland.swift, IslandStore.swift and four test files). That is exactly the verdict the gate declines to refuse on, and the reason is in its own text: a deliberate deletion and a revert are the same bytes. Read by hand rather than trusted: 019-drive-findings was cut before 866faa34 feat(mirror): rip out the pixel islands, so it still carries them; the integration line deleted them on purpose and the merge kept them deleted. The gate's non-refusing output is the half that still needs a reader, and this round is a worked example of what reading it looks like.

docs/open-issues.md union-resolved, audited in Python

Both 024-bake-volume-clean and 024-items-arbitration append here, and git auto-merged them. Audited by heading rather than by eye: 538 headings on the base, 541 after merge 1, 543 on the items lane, 546 merged with nothing lost from either side, and the line count is the exact sum (19,523 + 210 + 101 = 19,834). The five duplicate headings the audit found — ### Broken, ### Unverified and friends — are pre-existing in the base, one per ledger entry, and were checked against it rather than assumed.

The derived tables did not move, and that is a measurement

tools/derived-doc-gate check passed and rederive rewrote nothing: 49 / 23 inbound message types, a 47-verb registry served 44 by the PowerPC guest and 13 by NOW-68K, 14 / 14 census probes. Expected for this round — the census fixes changed what the ata and pccard probes DO, not which probes exist, and SceneRenderer is host-side and below the contract entirely. A derivation unmoved by a merge looks identical to a derivation nobody ran; the difference is that a machine took this one.

Still open

  • Nothing in this arc is metal-verified, including both census fixes. The ata fix is the one that matters most and can only be proven by a bake or a real machine; this round baked nothing.
  • tools/guest-ata/ is Builds only — it is not in scripts/build-guests, so no gate cross-compiles it.
  • scripts/test-all's header still says the MirrorKit stage is "165 tests"; it is 258 now. Prose restating a number is a second place to be wrong, and this is one.

FIXED: a stale arc-status answered confidently instead of refusing (2026-08-08, claude/026-corpus-reach)

tools/arc-status is the tool a coordinator runs before telling a person where the arc stands. It is not on main and is absent from 400 of the 449 branches here, so briefs fetch it by naming a branch — and the standing arc-trigger check has for weeks named git show claude/019-integration-5:tools/arc-status.

That copy globs claude/01[0-9]-*, in five separate places. Run it today and it enumerates the 019 lanes, omits every 02x lane, and prints a table headed "work that exists". Nothing in the output says the set is short. Of the 49 branches carrying this file, 40 carry the narrow glob; 9 carry the widened one from ed5da741 ("fix(arc-status): the arc rolled past 019 and the trigger did not").

This is the project's most-repeated shape: an instrument whose normal mode of operation is the one condition under which the defect cannot appear. A widened glob can see a narrow one's blind spot; a narrow glob cannot see its own, and a partial answer from a tool people quote as derived truth is worse than no answer.

Correcting the pinned branch name in the docs was necessary and is not sufficient — the next arc rolls to 03x and every copy pinned today goes quietly wrong again, this one included. So the tool now refuses: it derives the newest arc number from the branch names present, by a scan that deliberately does not use $LANE_GLOB (a check written in terms of the thing under test cannot fail), and if its glob cannot match that arc it exits 65 printing both globs and where to fetch a current copy. Older arcs the glob excludes are reported as a note and do not refuse — that is a scope decision, not a defect.

Watched fail. The narrow glob was reintroduced as LANE_GLOB and the refusal fired naming claude/01[0-9]-* against a derived newest arc of 026, exit 65; the widened glob restored and the tool ran clean. The glob also had two remaining hardcoded copies in the body (a printed heading and a case pattern); both now read $LANE_GLOB, so a future widening is one edit and cannot half-apply.

What is NOT fixed: the standing arc-trigger check itself lives outside this repository, and this lane cannot edit it. Until it is changed, it will still hand out the 019-integration-5 name — the refusal is what stops that from being silent, not a substitute for correcting it.

FIXED: mirror_read --intention metrics reported failed and never said which failure (2026-08-07, claude/026-cycle-outcome-reason)

Driving the live PowerBook 1400c over the host agent socket, now-agent mirror_read --intention metrics returned 24 cycles of which 14 read outcome: "failed", requestMs: 0, totalMs: 0 — every clock zero, the request never sent, and no way from outside to tell what had happened.

failed is a real word in the vocabulary, and it is the bucket in it. The brief that opened this lane enumerated NOWMirrorSource's outcome assignments and concluded failed was not among them; it is — line 597 is a two-line ternary and the enumeration stopped at the line break. So there was no collapse of starved or wrong-mac into a placeholder: those two survive to the socket and always did, and a test now pins that (testEveryNonOkOutcomeSurvivesTheMetricsProjectionVerbatim).

The real loss was one layer down. GuestListener.SceneFailure carries a message written for a person, and at least five distinct conditions arrive through it wearing the same word: no Mac connected, a scene already in flight, a transfer that arrived short, a delta that would not rebuild, and the 20-second watchdog's silence. The host spent that sentence on ambient — the one line under the Mirror window — and the agent socket cannot read a window. An agent asking the machine what was wrong got the bucket.

So the sentence now rides beside the word, verbatim, as MirrorCycleClocks.reasonAgentIntegrationMirrorCycleMetric.reason. No outcome word was added and none changed meaning. It stays off the NOWBASE cycle line: BaselineLine's values are space-free by construction and its own comment warns against inviting "somebody to put a message in one".

The second defect in the same row: counts belonging to another cycle

Each of those 14 rows also carried a window count and an element count. They were read off scene, which is the last proven scene and deliberately STANDS through a failure — blanking the Mirror on one poll is how a busy lane looks like a crash. Correct for the picture; wrong for the record, because it put the last good walk's numbers in the row of a cycle that never asked, beside three zeroed clocks, where they read as data.

They are now nil unless the cycle published a scene, rendering as -, which is the grammar this file already uses for "nothing to count". The alternative considered and rejected was marking them inherited: that needs a new field and a fourth word for absence, and empty / unknown / notFetched is the vocabulary this project has already paid for.

Still open

  • Not metal-verified. Tested only — the guard is watched failing against three separate mutations (drop the reason in the projection; flatten every non-ok outcome to failed; restore the stale counts), each confirmed to build and to be named by the test. Nobody has re-run the live drive that produced the 14 rows.
  • no-reply looks unreachable as a RECORDED outcome. It is the initial value at NOWMirrorSource.swift:381 and :552, and every path that reaches recordCycleClocks overwrites it first — the watchdog's silence arrives as a SceneFailure and lands on failed. Not touched here, because removing a word from a vocabulary is not a transport fix.
  • A cycle whose scene did not DECODE still records outcome: "ok". cycleOutcome is set to ok when the delivery arrives (:568); the two decode-failure catches near :763 fall through to finishCycle without correcting it, so an unreadable IR version is measured as a successful cycle. Found while reading for this fix, deliberately not fixed in it — it changes what ok means, which this lane was told not to do. With the counts now gated on cycleOutcome == "ok", such a cycle will also report the stale counts it used to; that is the same bug, unchanged in size, and it closes when ok does.

UNDECIDED, RESERVED FOR MICHELLE: may finderSelect front the Finder, when fronting starves the guest? (2026-08-08, claude/026-no-unbidden-front, lifted at round 15)

claude/026-no-unbidden-front (534f05be) stopped finderDeselect fronting the Finder. It deliberately left finderSelect alone, and it said why in its commit message — a message, and nothing else. That branch changes no file under docs/, so merging it as written would have landed the fix and lost the question. This entry is that question, lifted at the round 15 merge.

The question. finderSelect sends select … with activate: true, so selecting an item in the Mirror brings the Finder forward. Is that still the right trade, now that fronting has a measured cost?

The case for fronting, which is a DECIDED behaviour and not an oversight. It was measured on a live machine on 2026-08-05 and is guarded by testTheFinderComesForwardForASelectionAndNotOverANewApplication, whose stated reason is "a selection nobody can see is not a selection". A selection behind another window is invisible, and an invisible selection is arguably not one.

The case against, measured 2026-08-08 on the PowerBook 1400c. Fronting the Finder backgrounds the guest application — and the guest is the only context permitted to give back its own content-plane port hooks. content_uninstall_context skips every row whose a5 is not the caller's, and the jGNE filter that drives it runs in whatever process happens to pump. So while the Finder is front, the guest's hooks cannot be released by anyone. The night's Mirror logging recorded 22 hooks installed against 1 uninstalled, four ports still hooked at a5 0x0, and the guest logging "not scheduled for 14s". Michelle experienced this as the guest app "kept hiding itself" — it was not hiding, the host was fronting the Finder at it.

Why it is not an agent's call. Visibility versus scheduling is a genuine product trade-off with evidence on both sides, and the branch that found it said so plainly: "not one to resolve from a log at 3am — Michelle's call." A first draft of that commit changed finderSelect anyway and claimed nobody had asked; that was false, and the existing test is what said so.

What deciding it would look like. Either the guard's reason stands and the starvation is paid for elsewhere (a scheduling window, or a release path that does not require the owning process to be front), or finderSelect stops fronting and the visibility guard is retired with a dated line saying why. Both are edits to a test that currently encodes a decision, so neither should happen quietly.