# LSP capability matrix — affirmative-clean-signal strategy

How pi-lens knows a just-edited file is **clean** (no diagnostics) vs the server
simply **hasn't answered yet** (cold/crashed/silent). pi-lens waits *synchronously*
for a verdict, so — unlike an editor, which renders asynchronously and never needs
to decide — it must have a positive signal. There is no signal for "silence", so we
classify each server and pick a per-server strategy. (Background: #240; mechanism
confirmed against Neovim's LSP client, which sidesteps this entirely by being async.)

Generate/refresh this matrix with `node scripts/characterize-lsp.mjs [--install]`
(the `mode` column) and `node scripts/probe-clean-signal.mjs [--install]` (the
`clean-behavior` column — 4-way: 2 / 2* / 3 / unknown — and the `first-publish`
column, #3310). Both **merge** in place: a server the
running host couldn't spawn keeps its prior row, so an ubuntu-poor run can't
regress a richer one (#390). The nightly **tool-smoke** workflow runs both (plus
`server-capabilities.mjs`) and opens/updates a single auto-PR
(`bot/lsp-docs-refresh`, "docs(nightly): refresh LSP capability docs") with the
regenerated docs — so this file self-populates from CI without manual copy-paste.

`probe-clean-signal.mjs` also runs a **drift check** (#529): it compares each
probed server's observed `clean-behavior` against the hand-set `silentOnClean`
marker in `clients/lsp/wait-policy/strategies.ts` and writes any mismatch to the
`## silentOnClean drift` section below. This is telemetry only — never a CI
gate — because the probe is a timing-based negative observation; a mismatch
just tells a human the marker may need updating. `unknown` observations are
never compared in either direction (a slow/absent server isn't evidence of
anything). The native TS7 launch variant (`typescript7`/`typescript7-clean`,
#524/#526) is deliberately excluded from comparison against classic's marker —
they share a server id but not a verified clean-signal behavior.

What IS gated is the pair of COMMITTED sources:
`tests/config/lsp-clean-behavior-census.test.ts` (#3347) compares the
`clean-behavior` column below against the `silentOnClean` markers in both
directions — every measured push row against its marker, and every marker
against a push row measured `silent` — so a nightly refresh cannot merge a
re-measured server until its marker moves with it, and a marker cannot outlive
the measurement that justified it. Push rows the probe has not classified are
admitted by name in that file, and an admission reds once its row becomes
measured.

## First publish: is the first push an answer? (#3310)

A second, ORTHOGONAL axis to the tiers below, measured from the same probe run
and written as the `first-publish` column:

| value | meaning | wait policy |
|---|---|---|
| `direct` | the first publish for a dirty file carries the findings | the first publish resolves the wait, as it always did |
| `empty-first` | the server answers `didOpen` with an EMPTY set while a one-time index builds, then publishes again once indexing ends | the client HOLDS that first empty publish (`emptyFirstPublish: "indexing"` in `clients/lsp/wait-policy/strategies.ts`); the server's own next publish releases it, and the hold is one-shot per client session |
| `empty-only` | every publish on the dirty fixture was empty | not classifiable on this axis — an undiagnosed dirty fixture looks exactly like an unfinished index, so nothing is inferred in either direction |
| `n/a (pull)` / `TBD` | pull-mode, or not measured on a host that could reach the server | no behavior change |

Why this axis had to exist at all: the wait needs a positive signal, and an
empty publish is a positive signal for a Tier 2/2\* server (a genuinely clean
file) and a MEANINGLESS one for a server that has not finished indexing. The two
are identical on the wire, so the difference is a measurement, not an inference.
intelephense is the measured member (2026-09-23, v1.18.5): `[]` at +295ms,
`indexingStarted` at +301ms — i.e. AFTER the empty publish, so no
work-in-progress signal is available at the moment the decision must be made —
`indexingEnded` at +689ms and the real 2-diagnostic set at +696ms. On a
genuinely clean file the same session publishes `[]` twice (at +304ms and again
at +663ms, right after indexing), which is what lets the held publish be
released by the server itself rather than by a timer; a WARM touch publishes the
real set first, so the hold costs nothing once the session is warm. pi-lens does
not advertise `window.workDoneProgress` (#974: advertising it crash-loops
opengrep's `--experimental` LSP mode), so `$/progress` is not an available
discriminator either.

`tests/config/lsp-first-publish-census.test.ts` compares this column against the
`emptyFirstPublish` markers and reds in BOTH directions, which is the expiry
check: when the nightly re-measures a server into a different class, the docs
refresh cannot merge until the marker is updated with it.

## The strategies

| Tier | Signal | Affirmative clean? | Example |
|---|---|---|---|
| **1 — pull** | `textDocument/diagnostic` returns an authoritative report (empty = clean) | YES, deterministic | rust-analyzer |
| **2 — push, publishes-versioned** | `publishDiagnostics([])` **with version** on every scan, incl. clean→clean | YES, currency-proven via version | ast-grep |
| **2\* — push, publishes-unversioned** | re-publishes on a clean scan but **version-less** — the wait still early-returns (the client accepts a version-less publish as fresh: it can't be proven stale), but currency is only *temporally correlated*, not proven. Since #3484 the client fences a server measured to answer a `documentSymbol` fence before it publishes (`diagnosticsFence: "reply-first"`: yaml, intelephense): a pre-edit publish is dropped until the fence's reply. Every other 2\* server (docker publishes before answering; prisma, taplo, zls, dart, gleam, clojure unmarked) keeps this behaviour | YES at runtime, with a staleness-risk caveat (not a latency cost) | opengrep |
| **3 — push, silent on clean** | server publishes nothing when nothing changed | **NO** — budget-wait floor (safe; a timeout is *not* a false clean). **This tier is #458's learned-deadline target set.** | typescript-language-server |
| **Navigation-only — custom, no evidence** | custom `lsp.servers.*` entry has no pull provider and has not published in this session | **NO** — diagnostics are unsupported; skip the wait and report navigation-only | Dexter |

Detection is **cached** at `initialize` (`detectWorkspaceDiagnosticsSupport` →
`state.workspaceDiagnosticsSupport.mode`, upgraded on `client/registerCapability`),
so the tier is free at collection time — no per-edit probe.

Custom servers without a pull provider begin with one bounded first-contact
probe. The probe uses the smaller of the server's push-wait budget and the live
hook deadline. If the hook cuts it off first, the result remains unconfirmed and
the next touch probes again. Only silence through the full push budget latches
that server id as navigation-only for the service session. A publish upgrades it
to the ordinary push-wait policy.

## Matrix (dev box + CI nightly; mode last refreshed 2026-06-17 from run 27713958681, clean-behavior probed in run 35914033696 — #460)

`mode` from cached capabilities; `clean-behavior` from the phase-aware publish-trace
probe. The probe attributes publishes to two phases — the **dirty touch** (proves the
server is live) and the **clean transitions** (the discriminator) — and classifies
4-way: `publishes-versioned` (tier 2: publish WITH version on a clean transition —
affirmative + currency-proven), `publishes-unversioned` (tier 2\*: version-less
publish on a clean transition — the wait still early-returns at runtime, currency
only temporally correlated), `silent` (tier 3: alive on dirty, silent on clean —
budget-wait, the #458 target), `unknown` (no publish at all — slow/absent,
conservatively not classified). `src` = where a row was measured: **ci** = the
nightly steps, **dev** = the dev box (a row measured on both reads `dev+ci`).
Merges never blank a prior good value, so a CI non-result leaves the dev
classification standing.

Two bounded guards keep that preservation from hiding a dead instrument
(#3401). A `direct` `first-publish` cell whose axis the nightly probe does not
re-observe is stamped once with the date of its first miss and degrades to
`unknown` once five calendar days have elapsed since it (skipped nights do not
stall it, and a re-observation clears the stamp). `empty-first` cells are never
expired: they back the live `emptyFirstPublish` markers (php, terraform) and
expiring one would erase the measurement behind a marker. A `clean-behavior`/
`tier` change is written only after two consecutive nightly runs observe the
same new value, so one flapping nightly (ast-grep went 2 → 2* → 3 → 2* across
four runs) cannot rewrite a cell; a night that measured nothing for the lang
resets the hold. A subset probe (`probe-clean-signal.mjs <langs>`) leaves the
bookkeeping of the langs it did not probe untouched.

The bookkeeping lives in the generated `## Capability matrix refresh state`
section at the end of this doc, the only state the refresh persists. The clock
is the nightly run, not the bot PR's merge: each nightly starts from the last
`bot/lsp-docs-refresh` doc when that branch is ahead of master and was built on
master's current doc (`scripts/seed-matrix-from-bot-branch.mjs`), and from
master's doc otherwise (branch absent, squash-merged, or master edited the doc
since). Closing the bot PR unmerged does not reset the bookkeeping: the branch
is kept and keeps seeding, so only deleting `bot/lsp-docs-refresh` resets it.

`vue`'s `clean-behavior` was hand-reset to `unknown` (#3390): its
`publishes-unversioned` cell came from 58/45 publishes that the shared
`extension.log` window had attributed to vue but that belonged to `tinymist`.
With the sink scoped per server, nightly 36046209160 measured vue 0/0, and an
`unknown` result is never written by the merge above — so the refuted value had
to be cleared by hand. It stays `unknown` until a run observes
`@vue/language-server` publish; `tests/config/lsp-clean-behavior-census.test.ts`
carries the named admission until then.

| lang | server | mode | clean-behavior | first-publish | tier | src |
|---|---|---|---|---|---|---|
| json | vscode-json-language-server | pull | — | n/a (pull) | 1 | dev+ci |
| css | vscode-css-language-server | pull | — | n/a (pull) | 1 | dev+ci |
| html | vscode-html-language-server | pull | — | n/a (pull) | 1 | dev+ci |
| rust | rust-analyzer | pull | — | n/a (pull) | 1 | dev+ci |
| svelte | svelte-language-server | pull | — | n/a (pull) | 1 | dev+ci |
| deno | deno (alt of typescript) | pull | — | n/a (pull) | 1 | dev+ci |
| ruby | ruby-lsp | pull | — | n/a (pull) | 1 | ci |
| csharp | csharp-ls | pull | — | n/a (pull) | 1 | ci |
| typescript | typescript-language-server | push-only | silent | direct | 3 | dev+ci |
| markdown | marksman | push-only | silent | direct | 3 | ci |
| lua | lua-language-server | push-only | silent | direct | 3 | dev+ci |
| python | pyright | push-only | publishes-versioned | direct | 2 | dev+ci |
| jedi | jedi-language-server (alt of python) | push-only | publishes-versioned | direct | 2 | ci |
| yaml | yaml-language-server | push-only | publishes-unversioned | direct | 2* | dev+ci |
| shell | bash-language-server | push-only | publishes-versioned | direct | 2 | dev+ci |
| dockerfile | docker-langserver | push-only | publishes-unversioned | direct | 2* | dev+ci |
| toml | taplo | push-only | publishes-unversioned | direct | 2* | dev+ci |
| terraform | terraform-ls | push-only | publishes-unversioned | empty-first | 2* | dev+ci |
| prisma | @prisma/language-server | push-only | publishes-unversioned | direct | 2* | dev+ci |
| php | intelephense | push-only | publishes-unversioned | empty-first | 2* | dev+ci |
| zig | zls | push-only | publishes-unversioned | direct | 2* | dev+ci |
| vue | @vue/language-server | push-only | unknown | unknown | 2/3? | dev+ci |
| dart | dart language-server | push-only | publishes-unversioned | direct | 2* | ci |
| gleam | gleam lsp | push-only | publishes-unversioned | direct | 2* | ci |
| clojure | clojure-lsp | push-only | publishes-unversioned | direct | 2* | ci |
| opengrep | opengrep (aux) | push-only | publishes-unversioned | direct | 2* | dev+ci |
| ast-grep | ast-grep (aux) | push-only | publishes-versioned | direct | 2 | dev+ci |
| cue | CUE Language Server (cue lsp serve) | push-only | publishes-versioned | direct | 2 | dev+ci |

**Unknown — fixture exists, mode not yet captured.** The toolchain-gated family
(no auto-install today; tracked in #241) — `go` (gopls), `java` (jdtls),
`kotlin`, `swift` (sourcekit-lsp), `cpp` (clangd), `haskell`, `elixir`,
`ocaml`, `nix` (nixd), `fsharp`. Their servers don't install in the nightly, so
characterize reports `unknown` (a non-failure ⚠). Once #241 lands they'll fill in
the same way clojure-lsp/gleam now do (both auto-install via the github strategy
and were characterized `push-only` in the run above).

## Key findings
- **Mode ≠ tier, and the split needs BOTH axes.** Push-only further splits along
  latency (does anything publish on a clean transition? — silence is the only
  budget-wait case, because pi-lens's publish handler emits and early-returns the
  wait on every publish, versioned or not, with the single #3310 exception of a
  held empty FIRST publish from an `empty-first` server) and currency-proof (is the publish
  versioned, i.e. provably about the live edit?). The 4-way
  `probe-clean-signal.mjs` measurement drives this: ast-grep → 2, yaml/opengrep →
  2\*, typescript (clean file) → 3.
- **opengrep is 2\*, not 3 and not plain 2.** It *does* re-publish on a clean scan
  (the wait early-returns at runtime — fast), but every push carries
  `pubVersion=undefined`, so currency is only temporally correlated, not proven —
  a staleness-risk note, not a latency cost. An earlier hand-note called it Tier 2
  on the "re-publishes" observation alone; the phase-aware probe refines it to 2\*.
- **typescript's clean behavior is diagnostic-set-dependent (major probe finding).**
  On a DIRTY file it re-publishes (version-lessly) after every change — the dirty
  fixture measures 2\*. On a genuinely CLEAN file (the `typescript-clean` fixture)
  it publishes nothing on a clean→clean edit — silent, Tier 3. The clean-file
  behavior is the production case (the observed budget-wait timeouts), so the
  matrix row records the clean fixture's verdict; the probe prefers `clean: true`
  fixtures for exactly this reason. Corollary: a 2\* measured only on a dirty
  fixture may overstate a server whose publishes stop when its set goes empty —
  langs without a clean fixture carry that caveat.
- **The probe's publish capture was dead, and the column looked alive anyway
  (#3310).** `probe-clean-signal.mjs` intercepted `console.error` for the
  `[lsp-pub]` trace, but #1333 moved that trace's sink to `extension.log` — so
  every phase counted zero publishes, every server classified `unknown`, and the
  #390 merge guard then preserved the July 2026 values verbatim. A dead
  instrument and a healthy one are indistinguishable when non-results are
  discarded by design. The probe now reads the sink the client actually writes;
  `first-publish` measurements on the dev box, 2026-09-23: php `empty-first`,
  typescript / opengrep / ast-grep / marksman `direct` (so the class does not
  extend to them on the measurement, whatever their comments suggest).
- **#458's learned-deadline target set = the tier-3 rows only.** 2\* rows resolve
  the wait at runtime and must NOT be given learned deadlines.
- **Tier 3 is budget-bound by necessity**, not laziness: a silent server's silence is
  ambiguous (clean-unchanged vs still-analyzing), so shortening the wait or reusing
  `lastKnownDiagnostics` would risk a false clean. The wait *is* the safety mechanism.
- **ast-grep (Phase 2 / #239) is Tier 2** — it self-signals clean on every scan, so it
  is not the bottleneck. The cost on a clean with-auxiliary touch is the *silent
  primary* (typescript), a pre-existing Tier-3 cost independent of ast-grep.

## Completing the matrix
The fixtures (`tests/fixtures/tool-smoke/<lang>/`) are durable and cover every
registered server. `mode` is read from the server's advertised capabilities at
`initialize`, so it is **content-independent** — for languages that already had a
tool-layer fixture we point `characterize-lsp.mjs` at the existing (deliberately
dirty) `bad.*` source rather than a colliding clean duplicate; new languages get a
minimal clean source. Either way the mode reported is the same.

The nightly **tool-smoke** workflow runs `characterize-lsp.mjs --install` **and**
`probe-clean-signal.mjs --install` (after the LSP handshake layer) on `ubuntu-latest`,
then opens/updates an auto-PR with the regenerated docs (#390) — so both the `mode`
and `clean-behavior` columns self-populate in CI without manual copy-paste. They fill
for servers that either auto-install (npm/pip/github — including clojure-lsp and gleam,
both github-strategy as of f263cf3) or whose toolchain the workflow provisions
(Ruby/Dart/Zig + .NET→csharp). The remaining `unknown` rows are the toolchain-gated
family (#241): until `runtimeInstall` + canonical-bin discovery land, their servers
don't install in CI and stay ⚠. The **merge guard** means a CI run that can't reach a
server never blanks its dev-measured row.

The `clean-behavior` split (4-way: 2 publishes-versioned / 2\* publishes-unversioned /
3 silent / unknown) is now measured per-server by `probe-clean-signal.mjs`, no longer a
manual one-off. Locally confirmed: ast-grep + ast-grep-baseline → 2, yaml + opengrep →
2\*, typescript → 3 on its clean fixture (2\* on the dirty fixture — see Key findings);
the rest fill in as the nightly reaches them.

## silentOnClean drift (nightly-generated)

Telemetry only — never a CI gate. Compares each probed server's observed
`clean-behavior` against `clients/lsp/wait-policy/strategies.ts`'s `silentOnClean`
marker; a mismatch means the marker may need a human update (#529). `unknown`
observations are never compared (a slow/absent server is not evidence either way).

_None observed as of the last probe run._

## Capability matrix refresh state (nightly-generated)

Bookkeeping for the date-based `direct` `first-publish` expiry (#3401), the
two-run `clean-behavior` hysteresis and the consecutive-night `idle-eviction`
counts (#3989). Regenerated every run; never a measurement.

```json
{"idle-eviction":{"docker":{"nights":[{"day":"2026-10-07","rssMb":64,"coldMs":566}]},"json":{"nights":[{"day":"2026-10-07","rssMb":67,"coldMs":1116}]},"powershell":{"nights":[{"day":"2026-10-07","rssMb":155,"coldMs":2601}]},"python-jedi":{"nights":[{"day":"2026-10-07","rssMb":51,"coldMs":1796}]},"zizmor":{"nights":[{"day":"2026-10-07","rssMb":58,"coldMs":608}]}}}
```
