# Percept embedding geometry is a measured property, and an always-on embedding tier summons the watcher — it may never suppress below the floor cadence

**Status:** Proposed (2026-07-15 — drafted in the owner percept hill-climb session; owner
acceptance + in-flight evidence pending). Sibling and extension of ADR-0044 (**percept-reading
comprehension is a first-class, measured property at every layer**): where 0044 established the
**reading** layer (can the model recover the facts the percept encodes?), this ADR establishes the
**representation** layer (does the render place decision-equivalent situations near each other in
embedding space, and separate decision-divergent ones?) and adds a **runtime** layer 0044 did not
touch (a cheap always-on embedder that decides *when to wake the expensive reader at all*). It
inherits ADR-0044's honest-measurement discipline verbatim in spirit (the JudgeCell / no-P&L /
publishable-but-never-blended pattern — named here **the A5 pattern**) and its citation-status
convention (A4). It grades on the **both-poles matched sets** of ADR-0042 (**season validity requires
matched, balanced, divergent-outcome sets**) — pole-separation is measured *on the very sets 0042
constructs* — and it obeys the **null-hurts / both-poles** admission law of ADR-0041 §2
(**template-as-hypothesis**). It slots into the coordinate system of ADR-0043 (**coordinate ascent
over five knob families**): geometry is an **F0 measurement concern** (audited, never climbed for
advantage), the cross-render screen prices **F2/F4** arms, and the runtime tier is a new **F3
delivery** mechanism (WHEN to wake). Follows the gate pattern of ADR-0040 (a session that cannot
measure a property can never ground a claim about it). Companion in-flight agents (external
provenance, per the A4 convention below): `embed-geom-1` (the geometry + cross-render-screen
instrument) and `embed-trigger-1` (the runtime cascade + custom-embedder path) — both **running,
not yet landed as LEDGER rows**. Cites `inner-loop-1` / `inner-loop-1c` (the economic and
frame-blindness ground truth), `percept-quiz-1` (the comprehension-regime sibling), and ADR-0042 /
ADR-0044.

## Context

**Reading was the prior property to judgment; geometry is the prior property to reading.** ADR-0044
decomposed the stack once: **reading** (recover the facts) is distinct from and prior to **judging**
(act well given the facts), and it built a quiz instrument to score reading directly. This ADR
decomposes one layer further down. Before a model can *read* a percept, the render must *represent*
it — must map the frozen Frame into a token stream whose geometry a reader (any reader, including a
tiny one) can traverse. A render can be locally legible (high quiz accuracy) and still be
**geometrically degenerate**: it can place a fakeout and a real breakout — situations that demand
*opposite* actions — at nearly the same point in embedding space. That collapse is not a cosmetic
defect. It **is** the passivity failure, expressed one layer earlier: if the two poles are
indistinguishable to the representation, no reader traversing that representation can reliably act
differently on them, and the economically-correct discrimination ADR-0042 designs the benchmark to
reward is at risk at the substrate. But the identity must not be overstated: a reader (an LLM) does
**not** traverse the *probe* embedder's geometry, so pole-collapse under one embedder does not
literally foreclose the reader's own discrimination. The honest claim is a **signature, not an
identity** — pole-collapse under a **capable embedder panel** is strong evidence the render
under-serves the decision-relevant distinction, the passivity failure's substrate-level *signature*
(Decision 5, Amendment 3) — and it is measurable without spending a single judgment dollar.

**Three facts force the runtime half of this ADR.**

1. **Most LLM wakes are economically wasted.** `inner-loop-1` (the floor-row hill-climb; doc
   `inner-loop-1-2026-07-15.md`) found that **60.2% of TRAIN wakes carry a perfectly flat advantage
   vector** (all candidate actions score identically — pre-entry / no-fire situations where *nothing
   the watcher could decide changes the outcome*) and **86% carry ties**. The expensive reader is
   summoned, deliberates, consumes tape-time (ADR-0040), and on three wakes in five there was nothing
   to decide. This is not a training artifact to be scheduled away — it is the shape of the job: the
   watcher is a **backstop** most of whose moments are flat.

2. **A cheap tier can tell flat from live.** An embedding model in the 100–600M class runs at
   roughly **1/500th the cost** of the watcher LLM, **20–100× faster**, and is **Mac/CPU-servable** —
   it does not compete for the scarce, OOM-prone GPU that the crash-safeguards memory forbids us from
   casually loading (the 2026-07-14 Jetsam incident). A per-bar embedding + nearest-neighbor lookup
   against a versioned exemplar bank is a pennies-per-hour standing computation. If it can separate
   "approaching a decision-relevant configuration" from "flat," it can **summon** the LLM only when
   summoning is likely to matter.

3. **A trigger that can *suppress* re-opens every trap this program has closed.** The seductive
   version of tier-2 is a gate that both wakes *and stands down* the watcher — "the embedder says
   nothing is happening, so skip the scheduled wake." That is the abstention trap (ADR-0042) and the
   null-policy-wins degeneracy (Phase-0) rebuilt in silicon: a cheap model that suppresses wakes is
   rewarded for *doing nothing*, and doing nothing already beats the frontier on a raw floor. The
   defense is not a reward-shaping trick; it is an **architectural constraint**, stated below as a
   hard rule.

**Geometry is the most Goodhartable instrument we have proposed.** A quiz answer is right or wrong
against kernel ground truth (ADR-0044's no-abstention-trap-by-construction). An embedding distance
is a **soft, continuous, aliasing-prone** number: two renders can be pushed close or far by
representation artifacts (shared boilerplate tokens, tokenizer quirks, positional priors) that have
nothing to do with decision-relevant state. A geometry number is therefore trivially gameable if
ever used as an *admission grade* — a render could be optimized to separate the poles in embedding
space while carrying no more decision information, exactly the render-fit-not-capability confound
PM-1b flagged for the leaderboard (ADR-0044 §Context.2). This hazard does not disqualify geometry;
it **fixes its role**. Geometry is admitted as a **diagnostic and a pre-screen**, never as a grade
that promotes a render or a model. That role restriction is Decision 5, and it is the load-bearing
guard of the whole ADR.

## Decision

**Percept embedding geometry is a measured property of the render; a cheap same-frame embedding
distance pre-screens every render arm before any judgment budget; an always-on embedding tier
*summons* the watcher and may never suppress it below a floor cadence; and a custom contrastive
embedder — deterministic on fp32 CPU — is the first artifact with a clean, outsider-recomputable
production story. Every claim is diagnostic-only under the A5 pattern and pins an embedder-regime
id.**

### 1. Percept geometry is a measured property — the representation-layer sibling of 0044's reading layer

A good render embeds percepts **CLOSE** when their decision-relevant state matches, **FAR** when it
diverges, and is **INVARIANT** to nuisance (ticker identity, price scale, absolute date/time,
cosmetic formatting). This is a *property of the render*, measured with a frozen embedder, exactly as
comprehension-per-token is a property of the render measured with a frozen quiz. The hard metric is
**pole-separation on the ADR-0042 matched sets**: for each matched, balanced, divergent-outcome set
(fakeout vs real-breakout; gap-fill vs gap-and-go; dovish-rally vs hawkish-fade vs whipsaw), embed
every episode's percept and measure whether the **action-correct pole separates from the
restraint-correct pole** in embedding space. A render on which the **fakeout collapses onto the
breakout** is a **geometric failure**, and — per §Context — a substrate-level *signature* of the
passivity failure ADR-0042 designs the judgment benchmark to punish (a signature under a capable
panel, not an identity — Decision 5, Amendment 3), caught one layer earlier and far cheaper.
Pole-separation is reported as a *pole-vector diagnostic per (cell, role)*, never a scalar (ADR-0041
§2's record-shape law: a representation that separates one pole pair while collapsing another must
not net to an innocent average).

**Nuisance-invariance doubles as an audit of the existing fences.** ADR-0041 makes date-blindness
*structural* (`PaneBuildContext` carries no clock) and keeps illustrative fixtures on generic tickers
(SPX/SPY/QQQ). Geometry gives those fences a **measured test**: if two percepts that differ *only* in
ticker symbol, or *only* in absolute price scale, or *only* in calendar date embed **far apart**, the
render is leaking a nuisance dimension the date-blindness / generic-ticker doctrine claims it does not
carry. Nuisance-invariance is therefore not a new fence — it is the **instrument that audits the
fences ADR-0041 already asserts**, upgrading "date-blind by construction" from morally true to
measured (the same upgrade 0044's quiz gave "comprehension").

### 2. The cross-render information-loss screen — a pennies-cost standing pre-check for every F2/F4 arm

`embed-geom-1`'s distinct contribution: a **same-frame, canonical-vs-arm embedding distance**. For a
single frozen Frame, embed the **canonical render** (ADR-0043 F4's permanent control) and the
candidate **F2/F4 arm** render of the *same frame*, and measure their distance under a frozen
embedder. A tokenizer-optimal or token-diet arm that has silently **dropped decision-relevant
information** relative to canonical shows up as an anomalous same-frame divergence *pattern* across the
matched set — cheaply, before a single quiz item is scored or a single judgment dollar is spent. This
is a **standing pre-check** that runs on every F2/F4 render arm in the ADR-0043 climb: it costs
pennies, it runs wide over the frozen replay corpus (F2/F4 are the cheap-and-wide families), and it
**precedes** any comprehension-quiz or judgment budget. It cannot *promote* an arm (that would be
Goodhart — Decision 5); it can only *flag* an arm as suspicious and *sequence it ahead of* the
expensive instruments — an ordering optimization, not a grade.

### 3. The runtime cascade — an always-on embedding tier that SUMMONS, with scheduled wakes as a floor-cadence backstop

At runtime, a small embedder runs **per-bar**, always on. Each bar it embeds the current percept and
computes cheap geometric signals against a **versioned exemplar bank** (KNN retrieval): **cluster
transition** (the situation has moved to a new neighborhood), **pole approach** (the embedding is
nearing a known action-correct or restraint-correct exemplar cluster), and **trajectory inflection**
(the embedding path is bending — the ADR-0041 percept-inflection notion, now measured in embedding
space rather than only in the Frame). Any of these **SUMMONS the LLM watcher** — a zoom-on-wake
trigger (ADR-0043 F3) sourced from geometry rather than only from the clock. Scheduled clock-honest
wakes (ADR-0040) **remain** as a **floor-cadence backstop**.

**The trigger enters through the EXISTING wake machinery, and its wake renders a receipt (Amendment
1).** Decision 3 defines the summon; it must also say *where it lands*. The trigger tier is a
**DETECTOR-class wake source** inside the existing wake / coalescing / attention-budget path — its
squelch is visible in the kernel like every other wake source, never a bespoke parallel channel
beside the attention system the language already governs. The summoned wake's rendered reason
**carries its evidence**: `wake [DETECTOR embed-knn-vN] cluster_transition … bank=<sha> conf=…`,
stamping the **embedder-regime id** and bank sha as an ADR-0038-style receipt. Without this the
cascade is a second attention system beside the governed one, and an agent reading the screen cannot
see WHY it was summoned — itself an absent-not-hidden violation.

> **HARD RULE (the abstention-trap defense, stated as architecture).** The trigger tier may only
> **SUMMON** the watcher. It may **never suppress a wake below the floor cadence.** The scheduled
> wakes are a floor, not a ceiling the embedder is licensed to lower. The embedder adds wakes when its
> geometry says a decision may be live; it *never removes* a wake the clock already owed. This makes
> the null-policy-wins / abstention degeneracy (ADR-0042, Phase-0) **structurally impossible for the
> trigger tier**: a tier that can only add attention can never be rewarded for withholding it. It is
> the exact architectural analogue of ADR-0042 designing the abstention trap *out of* the benchmark
> rather than detecting it post-hoc.

**Economic basis.** 60.2% of wakes are flat-advantage (`inner-loop-1`) — most LLM calls are waste —
and the embedder is ≈1/500th the cost, 20–100× faster, and Mac/CPU-servable, so it can run every bar
where the watcher cannot. The cascade's *upside* is spending the expensive reader on the ~40% of
wakes where something is decidable; its *downside is bounded by the floor* — in the worst case
(embedder is pure noise) the system degrades to exactly the clock-honest scheduled cadence we run
today, never worse. That asymmetry — unbounded upside, floor-bounded downside — is why the summon-only
rule is not a limitation but the design's safety proof.

### 4. The custom-embedder path — contrastive fine-tune, deterministic on fp32 CPU, the first clean production receipt

The always-on tier does not require a frontier embedder; it rewards a **custom, small, contrastive
one**. Fine-tune a 100–600M-class embedder on triplets/pairs drawn from artifacts this program
already produces: **positives** = cube oracle-class matches (percepts the oracle labels as the same
decision-relevant situation); **hard negatives** = the ADR-0042 matched-set poles (fakeout vs
breakout — the pair we most need pulled apart, and the pair a generic embedder most easily collapses);
**premise-transition pairs** = consecutive percepts across an ADR-0041 inflection (the trajectory the
runtime tier must detect). This is **cheap** — hours and dollars, not weeks and thousands — because
the training data is a byproduct of the cube and the matched-set library, and the model is small.

Crucially, the custom embedder is the **first artifact in this program with a clean production
story**: **fp32-CPU deterministic inference**. A small embedder pinned to fp32 on CPU produces
**bit-reproducible** embeddings — which means an outsider can **recompute the receipt**. A geometry
claim ("these two percepts embed at distance d under embedder E, layer L, pooling P") becomes an
*outsider-recomputable fact*, not a trust-us number from a nondeterministic GPU serving path. This is
the receipt property the program has wanted and the fp16/GPU serving path cannot give: determinism is
not a nice-to-have here, it is what makes the embedding tier *auditable* and therefore admissible as
even a diagnostic.

### 5. Authority and guards — the A5 pattern (diagnostic/pre-screen only), embedder-regime ids, versioned banks, practice-tier retrieval

Geometry is the most Goodhartable instrument in the program (representation aliasing — §Context), so
its authority is fixed by construction, generalizing ADR-0044's comprehension-JudgeCell discipline
(0044 §"the quiz is itself an instrument" + A3) into a named, reusable pattern:

- **The A5 pattern — diagnostic and pre-screen ONLY, never an admission grade.** Geometry is a
  JudgeCell with **no P&L, never blended into the P&L grade, and never a promotion criterion**. It may
  *flag*, *audit*, *sequence*, and *summon*; it may **never** admit a render, promote a model, or
  gate curriculum spend *on its own number*. Admission and promotion continue to rest on the harder
  instruments: comprehension-per-token (ADR-0044) and both-poles economics (ADR-0041 §2 / ADR-0042). A
  geometry number that ever appears on the right side of a promotion decision is a Goodhart breach, by
  definition of this pattern. (This is the exact sibling of 0044's rule that comprehension is
  publishable as its own tier but *never blended into or confusable with* the alpha/judgment rows —
  A3.) **The A5 pattern here is paired with the ADR-0044 A5/A6 amendments** (quiz authority stops at
  F1's border; gate quiz held-out from stage-0), landed 2026-07-17 per kestrel-wa0j.55, so the
  sibling pair is internally consistent — neither doc uses a soft instrument as an admission grade
  (Amendment 4).
- **The panel rule for diagnostic claims — a single embedder cannot audit a render (Amendment 2).** A
  geometry claim used for **render audit** is inadmissible on a single embedder: cross-embedder
  **convergence** across a declared panel (spanning frontier-API + local families) is evidence about
  the **RENDER**; **divergence** is evidence about the **EMBEDDER**, not the render (the ADR-0009
  measured-not-designed pattern, restated for embedders; already filed on kestrel-wa0j.60/.61). API
  embedders are **panel-only** — they may inform a diagnostic but never serve the runtime tier, which
  stays **deterministic-local** per Decision 4.
- **Every claim pins an EMBEDDER REGIME ID** = `{model, layer, pooling}` (which model, which layer's
  activations, which pooling over tokens). This is the direct sibling of ADR-0044's
  **comprehension-regime id** (quiz-generator version + threshold) and ADR-0043's F0 rule: a change to
  the embedder, the layer, or the pooling **mints a new geometry regime** and invalidates
  cross-regime comparability. No geometry number is legible without its regime id.
- **Exemplar banks are sha-versioned.** The runtime KNN bank and every matched-set embedding bank are
  content-addressed by sha256 (sibling to ADR-0041's `templateHash` / `elabHash` and ADR-0044's
  versioned quiz bank). A bank change mints a new regime; a claim cites the bank sha it retrieved
  against. This is what stops the runtime tier from silently drifting its own trigger surface.
- **Retrieval-based product panes are practice-tier until the receipt exists.** A "days shaped like
  today" pane (retrieve and show the nearest historical exemplars to the current percept) is
  compelling product, but it rests on the *softest* instrument. It is **practice-tier** (usable in the
  viewshop / for exploration) and **not admissible to a shipped default** until (a) the receipt story
  is real — deterministic fp32-CPU embedder (Decision 4) — and (b) it clears the A5 pattern's
  diagnostic-only bar (it may *surface* neighbors; it may not *decide* for the reader). This is the
  ADR-0041 template-as-hypothesis discipline applied to a retrieval pane: a seductive lens is a graded
  seed, never a blessed default, and geometry's softness raises the bar for fixation.

## Validation fixture — the guards ship with failing fixtures (guards-need-failing-fixtures)

Per the program's law that no guard ships without a fixture that makes it fire red on purpose (the
2026-07-14 silently-inert-instrument lesson), each instrument in this ADR ships with a deliberate
red-firing fixture:

- **Pole-separation (Decision 1) ships a COLLAPSED-POLE fixture.** A render (or an embedder regime)
  deliberately constructed to embed the fakeout pole onto the breakout pole MUST make the
  pole-separation diagnostic fire red; a matched sibling render that separates them MUST pass. A
  pole-separation instrument that cannot detect a hand-collapsed pole pair does not ship.
- **Nuisance-invariance (Decision 1) ships a LEAKED-NUISANCE fixture.** Two percepts differing *only*
  in ticker (or *only* in price scale, or *only* in date) that a broken render embeds **far apart**
  MUST fire the nuisance-leak audit red; the generic-ticker/date-blind sibling MUST pass. This is the
  fixture that makes the date-blindness fence *checked*, not merely asserted.
- **The cross-render screen (Decision 2) ships a DROPPED-INFORMATION fixture.** An F2 arm that
  deletes a decision-relevant field relative to canonical MUST show the anomalous same-frame
  divergence and be flagged; an information-preserving cosmetic arm MUST NOT be flagged (no false
  positive on pure formatting).
- **The runtime cascade (Decision 3) ships a SUPPRESSION-ATTEMPT fixture.** A test that attempts to
  drive the trigger tier to *skip* a floor-cadence wake MUST fail closed — the floor wake fires
  regardless. A trigger tier build on which a suppression path exists does not ship. This is the
  architectural hard-rule expressed as a red test.
- **The custom embedder (Decision 4) ships a DETERMINISM fixture.** The same percept embedded twice
  through the fp32-CPU path MUST be bit-identical; a nondeterministic serving path MUST fire red and
  is inadmissible for receipts.

## Consequences — what changes at each layer

- **Renderer.** The ADR-0009/0043 tournament gains a **representation-layer diagnostic** beside
  0044's reading-layer grade: pole-separation and nuisance-invariance *audit* F2/F4 arms (F0
  concerns), and the same-frame cross-render screen *sequences* expensive instruments behind a
  pennies-cost pre-check. No render is ever *promoted* on geometry (A5). Canonical render stays the
  F4 control and the screen's reference point.
- **Bench.** Geometry is an **F0 measurement concern** (ADR-0043) — audited, versioned by
  embedder-regime id + bank sha, never climbed for advantage. It sits beside the judgment instrument
  (ADR-0042) and the comprehension instrument (ADR-0044) as a third diagnostic axis, publishable
  under the A5 pattern (its own tier, never blended into alpha or reading rows).
- **Runtime.** A new **F3 delivery mechanism**: an always-on embedding tier that summons the watcher,
  with clock-honest scheduled wakes (ADR-0040) demoted to a *floor backstop* the tier may raise but
  never lower. The watcher's compute concentrates on the ~40% of wakes that are not flat; the flat
  60.2% (`inner-loop-1`) stop costing frontier tokens. Memory-first serving discipline (crash
  safeguards) is satisfied by construction — the tier is CPU/Mac-servable and does not contend for the
  OOM-prone GPU.
- **Training.** A new, cheap, high-value training target — the **contrastive embedder** (Decision 4)
  — trained on cube oracle-class positives + matched-set hard negatives + inflection pairs. It is the
  program's cleanest training story: small, hours-and-dollars, verifiable, and **deterministic on fp32
  CPU** so its outputs are outsider-recomputable receipts. It does not carry the judgment objective's
  degeneracies (no flat-advantage collapse, no null-policy-wins) because its loss is a contrastive
  right-here/far-there, not an economics sum.
- **Product.** Retrieval panes ("days shaped like today") become buildable but are **practice-tier**
  until the receipt + deterministic-embedder story exists and they clear the A5 diagnostic-only bar.
  The runtime cascade is a product feature (cheap always-on attention that wakes the expensive agent
  only when it matters) with a floor-bounded worst case.

## Falsifiability — what would overturn this ADR

This ADR is wrong, and one or more of its layers should be demoted, if any of the following holds on
real data:

1. **The trigger frontier is worse than the timer baseline.** If the geometry-summoned cascade, on
   the matched-set / live-cohort economics, does **not** beat a pure clock-honest timer baseline at
   equal watcher-call budget — i.e. the embedder's extra summons and skipped-flat savings net to zero
   or negative decision-quality-per-dollar — then the runtime tier buys nothing and Decision 3
   collapses to "just run the timer." (`embed-trigger-1` is designed to measure exactly this frontier;
   an aliased-control arm — trigger on a *scrambled* exemplar bank — is the pre-registered foil that
   must lose.)
2. **Geometry is uncorrelated with judgment, even directionally.** If pole-separation does **not**
   predict judgment quality across renders and models — renders whose poles separate better do not
   read/judge better even in sign — then geometry measures a property that does not matter, and even
   its diagnostic/pre-screen role is pure cost. (This is the geometric sibling of ADR-0044
   §Falsifiability.1; the frame-shuffle result — `inner-loop-1c`, TRAIN end-loss 1.51 NORMAL vs 2.95
   SHUFFLED with zero VAL transfer — is *evidence toward* representation mattering, but the geometry↔
   judgment correlation must be shown, not assumed.)
3. **The fine-tuned embedder overfits its exemplar bank.** If the custom contrastive embedder
   separates the poles only on its *training* exemplars and collapses them on held-out matched sets
   (or drifts catastrophically when the bank is refreshed), then the runtime tier is memorizing its
   bank rather than learning decision-relevant geometry — the embedder analogue of the RL-memorizes
   finding — and Decision 4 must be re-scoped to a frozen off-the-shelf embedder until generalization
   is demonstrated.
4. **The same-frame screen has no signal.** If the canonical-vs-arm distance flags cosmetic arms and
   information-dropping arms indistinguishably (the DROPPED-INFORMATION fixture cannot be separated
   from formatting noise on real renders), the pre-screen is a coin flip and Decision 2 is removed.

## Open questions

1. **Which layer/pooling is the geometry regime?** The embedder-regime id pins `{model, layer,
   pooling}`, but *which* interior layer and pooling best expose decision-relevant geometry is an
   empirical calibration (`embed-geom-1`'s first job), not a fiat — set it from pole-separation on the
   real matched sets, the way ADR-0044 sets its comprehension threshold from the roster-scale
   frame-shuffle re-run.
2. **What are the trigger signals' thresholds and dwell/hysteresis?** Cluster-transition, pole-
   approach, and trajectory-inflection each need a firing threshold and an anti-chatter dwell
   convention (sibling to ADR-0041's SHOCK dwell/hysteresis). Pin against real tape; until then the
   floor cadence carries the system and the tier only ever *adds*.
3. **Does the cross-render screen (Decision 2) subsume any comprehension-quiz spend, or only sequence
   it?** Leaning *sequence only* — the screen flags and orders, the quiz decides — until §Falsifiability.4
   shows the screen has enough signal to *retire* quiz items rather than merely prioritize them (the
   geometric analogue of ADR-0044 Open-2: does the cheap instrument stand in for, or only front, the
   expensive one?).
4. **Custom vs frozen embedder for the runtime tier.** Decision 4 argues for a custom contrastive
   embedder on receipt grounds, but §Falsifiability.3 is the risk; the fallback is a frozen
   off-the-shelf embedder pinned to fp32-CPU for the same receipt property. Which ships first is a
   cost/generalization tradeoff `embed-trigger-1` should settle.
5. **Should a positive trigger-frontier result become an L-gate?** If `embed-trigger-1` shows the
   cascade beats the timer, does "no launch claims cheap-always-on-attention without a measured,
   positive trigger frontier" join the platform's `docs/LAUNCH-GATES.md` beside the no-latency-blind
   (ADR-0040), season-validity (ADR-0042), and (proposed) Reading-Delta (ADR-0044 Open-3) gates?

## Citation status (sibling to ADR-0044 A4)

The founding evidence for the runtime and geometry claims is **in flight, not yet landed**:
`embed-geom-1` and `embed-trigger-1` are **agents currently running** in the watcher-pareto session's
worktree, and their result docs (and JSON) land on `main` with that session's next merge, exactly as
ADR-0044's founding docs do (A4). The ledger rows this ADR leans on — `inner-loop-1`
(`inner-loop-1-2026-07-15.md`, the 60.2%-flat economics), `inner-loop-1c`
(`inner-loop-1c-frameshuffle.md`, the frame-blindness / representation-matters evidence), and
`percept-quiz-1` (the comprehension-regime sibling) — likewise live in
`kestrel-wt/watcher-pareto/docs/research/watcher-pareto/` until that merge. **Until those docs land,
this ADR is their citation of record and that merge is owed.** The ADR is **Proposed**: it commits the
architecture, the role restriction (A5), and the falsifiers; it does not claim the in-flight verdicts.
