# pi-decider

Decision backends as a tool call, for **pi** and **omp**. The model asks a backend for typed judgments over
state you hand it and gets back answers with probabilities instead of prose — so code (or the model) can
threshold them.

```
decide(state, questions[]) -> answers[] with probabilities + confidence
```

Backends are swappable per call and per question, so one request can combine several of them:

| backend | kind | endpoint | credential |
|---|---|---|---|
| `typesafe` | decisions | `POST https://api.typesafe.ai/v1/systemone` | `TYPESAFE_API_KEY` |
| `openrouter` | decisions | `POST https://openrouter.ai/api/alpha/decisions` | `OPENROUTER_API_KEY` |
| `llm` | chat | `POST <llm.baseUrl>/chat/completions` | `DECIDER_LLM_API_KEY` |

`typesafe` and `openrouter` both serve **Jev** (OpenRouter through its Decisions API, model `~typesafe/jev-latest`
— see [Jev through OpenRouter](#jev-through-openrouter)). `llm` is a stand-in: it asks an ordinary chat model and
enforces the same answer contract around it.

## Install

The plugin is a **package**: `package.json` declares its extension entry (`pi.extensions`) and its bundled skill
(`pi.skills`), so a harness that loads the package directory gets both. `typebox` and the harness API come from the
harness itself — there is nothing to `npm install`.

**pi — project mount** (`.pi/settings.json`; paths resolve relative to `.pi`):

```json
{
  "extensions": ["../plugins/pi-decider"],
  "skills": ["../plugins/pi-decider/skills"]
}
```

pi mounts extensions by path, so the bundled skill needs the explicit `skills` line. Project settings load only after
the project is trusted: pi asks once, or run `pi --approve` / save it with `/trust`.

**omp — package mount** (omp loads the package, including `skills/`):

```bash
omp config set extensions '["<abs path>/plugins/pi-decider"]'    # persistent
omp -e "<abs path>/plugins/pi-decider"                           # per run
```

omp's directory scan of `<cwd>/.omp/extensions/` loads *extension files* only — a package that way does **not** bring
its `skills/`, which is why the mount points at the package directory itself. Global alternative:
`cp -r plugins/pi-decider ~/.omp/agent/extensions/pi-decider` (then the skill needs a root omp scans, e.g.
`.omp/skills/decide/` or `skills.customDirectories`).

> Load the plugin **once per harness**. Two copies registering the same tool name make the second one fail to load
> (`Tool "decide" conflicts with …`), so drop the global copy when you use a project mount.

Then `/reload` in a running pi session (`omp` has no `/reload` — start a new session).

## Skill

`skills/decide/SKILL.md` ships inside the package and teaches the model when and how to use the tool: the three
primitives, the four shapes (single / fan-out / gate / composite), question-design rules, per-question backend
routing, threshold and weight handling, configuration commands, and limits.

Declared as `"skills": ["./skills"]` under `pi` in `package.json`, so any harness that loads the package directory
discovers it: pi via the `skills` settings line above, omp directly from the package mount. pi exposes it as
`/skill:decide` and lists it in the system prompt; omp additionally serves it as `skill://decide`.

## Configure

`/decide setup` walks through it inside the harness: pick the default backend, choose whether a key is an environment
reference (recommended) or a literal value, confirm model/base URL from their defaults, then write the file.
`/decide status` shows what resolved and probes the selected backend.

Everything lives in `<agentDir>/decider.json` and is re-read on every call, so edits need no `/reload`. The agent
directory comes from the harness itself (`getAgentDir()`), which means:

| harness | config file | override |
|---|---|---|
| pi | `~/.pi/agent/decider.json` | `PI_CODING_AGENT_DIR` |
| omp | `~/.omp/agent/decider.json` | `OMP_CODING_AGENT_DIR` |

Each backend block is self-contained — its own credential, endpoint, path, model, prices, limits:

```json
{
  "backend": "auto",
  "typesafe": {
    "apiKey": "$TYPESAFE_API_KEY",
    "baseUrl": "https://api.typesafe.ai",
    "path": "/v1/systemone",
    "model": "jev-latest"
  },
  "openrouter": {
    "apiKey": "$OPENROUTER_API_KEY",
    "baseUrl": "https://openrouter.ai/api",
    "path": "/alpha/decisions",
    "model": "~typesafe/jev-latest"
  },
  "llm": {
    "apiKey": "$DECIDER_LLM_API_KEY",
    "baseUrl": "https://openrouter.ai/api/v1",
    "path": "/chat/completions",
    "model": "",
    "temperature": 0,
    "maxTokens": 2048,
    "jsonMode": "json_object",
    "repairAttempts": 1
  },
  "modelsPath": "/v1/models"
}
```

`"$VAR"` (or `"${VAR}"`) resolves the environment at request time, so keys stay out of the file; a literal value is
used as-is. Per-backend `timeoutMs`, `maxRetries`, `costPerMTokInput`, and `costPerMTokOutput` follow their defaults
(60s, 2 retries for decisions backends, 0 for chat; Jev is $0.042/Mtok in and free out).

| env var | effect |
|---|---|
| `DECIDER_BACKEND` | `auto` \| `typesafe` \| `openrouter` \| `llm` |
| `TYPESAFE_API_KEY` / `OPENROUTER_API_KEY` / `DECIDER_LLM_API_KEY` | credentials, used when a block references nothing |
| `DECIDER_<ID>_API_KEY` / `_BASE_URL` / `_PATH` / `_MODEL` | per-backend overrides (`ID` is upper-cased) |
| `DECIDER_MODELS_PATH` | catalogue path (default `/v1/models`) |

**`auto` resolution:** explicit call argument → the `backend` field → the first usable backend, preferring real Jev
(`typesafe`, then `openrouter`) over the `llm` proxy. A backend is usable when it has a credential and, for chat, a
model id.

## Tool

```
decide({
  state: "Help! My payouts have been failing for 3 days.",
  questions: [
    { id: "is_urgent",  type: "noul",   instructions: "Does this convey urgency?",
      criteria: { true: "Explicitly time-sensitive", false: "No urgency expressed" } },
    { id: "department", type: "choice", instructions: "Which team should handle this?",
      criteria: { billing: "Payments, invoicing, refunds", technical: "Bugs, outages", none: null } },
    { id: "frustration", type: "score", instructions: "How frustrated is the customer?",
      criteria: ["Calm", "Frustrated", "Very angry"] }
  ]
})
```

| parameter | notes |
|---|---|
| `state` | string, structured JSON, or a JSON-encoded string |
| `questions[].id` | answer key, echoed back |
| `questions[].type` | `noul` (P(yes)), `choice` (one option + distribution), `score` (position across ordered levels) |
| `questions[].instructions` | one narrow judgment; string or structured JSON |
| `questions[].criteria` | noul `{true?, false?}`; choice `{option: rubric\|null}` or an array (≥2); score ordered levels (≥2) |
| `questions[].backend` / `questions[].model` | route this question to a specific backend/model |
| `backend` / `model` | call-level default for every question that does not override it |

Questions sharing a backend+model are sent as one request; different batches run in parallel. Failures are
per-question: one failing batch only marks its own questions `ERROR`.

Answer text carries one header per batch, then one block per answer:

```
openrouter · OpenRouter Decisions API · model ~typesafe/jev-latest · provider TypeSafe · 640 ms · tokens 321 in / 57 out · cost $0.000013
is_urgent [noul] 0.87
department [choice] billing · confidence 0.91
  probabilities: billing 0.8 · technical 0.1 · none 0.1
frustration [score] 1.5 · confidence 0.76
  levels: 0 Calm · 1 Frustrated · 2 Very angry
  probabilities: 0 0.33 · 1 0.33 · 2 0.33
```

Structured results (per-batch endpoint, provider, tokens, cost, plus every question/answer/issue/note) are
persisted on the tool result `details`, and tokens/cost flow into pi's usage totals.

## Commands

```
/decide                      status: resolved config, per-backend state, live probe (default)
/decide add [backend]        guided: key, model, base URL for one backend
/decide set [t] [f] [v]      one validated field; interactive when a part is missing
/decide unset [t] [f]        drop a field, or a whole backend block (confirms first)
/decide question [text]      ask your own question (interactive; Enter through = smoke test)
/decide models [b] [text]    models a backend catalogue reports, optionally filtered
/decide setup                guided setup incl. the default-backend choice
/decide help
/jev                         alias for /decide
```

Aliases: `setup` = `init`, `unset` = `remove` = `rm`, `question` = `ask`.

`set`/`unset` targets are `root` (plus the shorthand `set backend llm`) and the backend ids. Writable fields:
`apiKey`, `baseUrl`, `path`, `model`, `timeoutMs`, `maxRetries`, `costPerMTokInput`, `costPerMTokOutput`, and for `llm`
also `temperature`, `maxTokens`, `jsonMode`, `repairAttempts`. Arguments are optional everywhere — a bare `/decide set`
walks you through target, field (showing current values), and value, and literal credentials are masked in the report.
Everything is also scriptable: `/decide set openrouter apiKey $OPENROUTER_API_KEY`.

### Question shapes

`/decide question` covers the documented usage shapes rather than a single canned probe. Enter through the prefills
to get a smoke test; answer the dialogs to ask anything.

| shape | what it does |
|---|---|
| `single` | one typed question (noul / choice / score) |
| `fanout` | many independent questions in one request (speculative fan-out) |
| `gate` | confidence-gated routing: noul gates on its probability, choice/score on confidence, and the verdict is printed per question |
| `composite` | composite scoring: each score answer is normalized over its own levels, then combined by weight in code |

Intent routing needs no separate shape: it is a `choice` question whose answer selects a handler in the code that
calls the tool.

## Jev through OpenRouter

OpenRouter serves Jev on its **Decisions API**, not chat completions:

```
POST https://openrouter.ai/api/alpha/decisions
Authorization: Bearer $OPENROUTER_API_KEY
{ "model": "~typesafe/jev-latest", "state": "...", "questions": { "is_urgent": { "type": "noul", "instructions": "..." } } }
→ { "model": "~typesafe/jev-latest", "provider": "TypeSafe", "answers": { ... }, "usage": { "inputTokens": ..., "outputTokens": ... } }
```

Verified against the live service on 2026-09-18:

- `GET https://openrouter.ai/api/v1/models/~typesafe/jev-latest/endpoints` → `modality: "text->decisions"`,
  `output_modalities: ["decisions"]`; `POST /api/alpha/decisions` without a key returns 401 (the route exists).
- Pricing matches TypeSafe direct: $0.042/Mtok input, $0/Mtok output.
- Only the `~typesafe/jev-latest` family alias exists (pinned ids such as `~typesafe/jev-1.13.0` return 404), and it
  is absent from the public `/v1/models` listing — `/decide status` falls back to the model-route lookup above.
- Usage keys are camelCase on OpenRouter and snake_case on TypeSafe; both are accepted, and `provider` is reported
  when the response carries it.

## Contract enforcement

The plugin owns the input and output shapes for every backend:

- Questions are validated and canonicalized before any call (`normalizeQuestions`); aliases (`yes_no`, `rating`,
  `classify`), option arrays, and level maps are normalized, while duplicate ids, single-option choices, and
  single-level scores are rejected.
- Answers are validated against the question they belong to (`conformAnswers`): `noul` in `[0,1]`, `choice` must be
  one of the criteria keys, distributions must cover every option/level and are renormalized to sum 1, `score` must
  lie inside the level range.
- The `llm` backend additionally demands a single JSON object (`response_format: json_object`, retried without it if
  the endpoint rejects the parameter), recovers fenced or prose-wrapped output, and retries once with the validation
  errors fed back before reporting a per-question `ERROR`.
- A question that names an unusable backend (no credential, or a chat backend without a model id) is reported as
  that question's `ERROR`; the rest of the call still runs. Only an unusable call-level default is fatal, because
  then no question has a backend left.
- A chat model that spends its entire output budget before writing JSON (reasoning models do) is retried once with a
  budget raised 4× (capped at 16k). A batch that still cannot answer fails with the budget and the reasoning tokens in
  the message, and the tokens/cost of every attempt — including the rejected ones — are counted in the report.
- `confidence` is only ever relayed, never invented; `cost` appears per backend: computed from configured Jev prices,
  taken from the response when a chat endpoint reports it, otherwise omitted.

## Benchmark

`test/jev-suite/` is a scored benchmark for the backends, not a smoke test: **100 items** over **8 self-contained
state documents**, on two axes — domain (`reasoning`, `arithmetic`, `news`, `law`, `medicine`, `engineering`, `meta`)
× category (`judgment`, `calibration`, `abstention`, `arithmetic`, `architecture`, `meta`, `no-answer`).

Every item carries ground truth — a computed value, a normative best option, a forbidden assertion — or an invariant,
so a run reports **pass/fail per item, Brier over the `noul` items, and pair assertions**: language / option-order
items must return the same pick, and a changed constraint (`w1` vs `w2`, `l1` vs `l2`) must flip it. The law and
medicine documents are fictional and only ever test applying the clauses/protocol stated in the document.

```bash
bun run test/jev-suite/run.ts                                    # live; backend from decider.json
bun run test/jev-suite/run.ts --backend openrouter --repeat 2     # + per-pass stability
bun run test/jev-suite/run.ts --json out.json                     # record a run as a reference
bun run test/jev-suite/run.ts --replay out.json                   # re-score offline, no network
bun test test/jev-suite                                           # suite integrity + scorer only
```

`test/jev-suite/reference/jev-1.13.json` is a recorded run (`~typesafe/jev-latest`, 2026-09-18): 100/100 answered,
15.3k in / 4.1k out, **$0.00064**, Brier **0.079** over 37 `noul` items.

| domain | pass | what it says |
| --- | --- | --- |
| reasoning | 20/20 | determinate claims, routing, severity, language + option-order invariance |
| law | 9/9 | clause application, deadline arithmetic, "the document does not say" |
| medicine | 10/10 | protocol application, weight-based dose, contraindications, scope limits |
| news | 10/10 | source quality, headline-vs-body, missing-source detection, unknown follow-up |
| engineering | 24/26 | architecture calls, quorum/capacity/cache arithmetic, one latency + one retry error |
| arithmetic | 13/22 | **the weak spot**: multi-step and 6+-digit evaluation, plus the framing effect below |
| meta | 2/2 | self-contradicting requirement, absolute claim |

Two measured properties the suite pins down:

* **Framing matters.** The same fact scores very differently when the candidate answer sits in `state`:
  `num_span_claim` 0.21 ("is 1098.58 h correct?") versus `num_span_value` picking `1098.58` outright. Put candidates
  in `criteria`, keep `state` to raw facts.
* **No-answer options only help when the model is unsure.** With four wrong options and an escape hatch,
  `opt_span_escape` and `opt_ledger_escape` take the hatch, while `opt_comb_escape` walks past it at confidence 0.92
  (its own wrong value looks right). A confidence threshold is the only reliable guard.

## Limits

Jev takes text only (convert images/audio first), has a 64k request budget (32k for `state` plus the longest
question), streams nothing, and returns decisions rather than text or tool calls — which is why this is a tool and
not a pi model provider. A `llm` backend answer is a general model's best effort, not a calibrated System One answer.
