---
name: decide
description: Use the `decide` tool for calibrated typed judgments (Jev / TypeSafe System One, OpenRouter Decisions API, or a chat-model proxy) instead of asking a general model for prose. Covers the three question primitives, the four usage shapes (single, fan-out, confidence gate, composite score), backend routing per call and per question, question-design rules, thresholds, and configuration. Load when a task needs a probability, a pick among named options, a graded score, routing/ranking, moderation, or verification.
---

# decide — typed decisions as a tool call

`decide` returns **answers, not text**: a probability, a chosen option with its full distribution, or a position on
an ordered scale. Code (or you) thresholds the numbers. Reach for it whenever the right output is a decision rather
than a paragraph.

```
decide(state, questions[]) -> answers with probabilities + confidence
```

## Primitives (the three question types)

| type | answers | use for |
|---|---|---|
| `noul` | `noul`: P(yes) in [0,1] | one yes/no condition; → probability is the answer |
| `choice` | `choice` + `probabilities` + `confidence` | pick one of a named set (routing, intent, entity matching) |
| `score` | `score` + `probabilities` + `legend` + `confidence` | graded position across ordered levels |

- A `noul` near 0.5 means yes and no are equally likely — **not** "medium intensity". For intensity use `score`.
- `choice` probabilities compare the listed options with each other; `confidence` says how concentrated that
  distribution is. Two decent options can both be plausible — read the distribution, not just the winner.
- `score` may land between levels (probability-weighted). Levels are 0-indexed.

## Question design

- **One narrow, coherent judgment per question.** Split independent dimensions; don't destroy the relationship you
  are judging.
- **Criteria carry the definitions.** `choice` needs ≥2 options (`{option: rubric|null}` or an array of options);
  give every `choice` a no-match option when nothing may fit. `score` needs ≥2 ordered level descriptions, lowest
  first. `noul` criteria (`{true?, false?}`) are optional but sharpen the boundary.
- **Put everything the judgment needs in `state`** (source text, records, policies, current facts) — a string,
  JSON object/array, or a JSON-encoded string.
- **Batch independent questions into one call.** They run in parallel in a single request, which is far cheaper and
  faster than one call per question. Add a second call only when an earlier answer is needed to build the next
  state or to choose the next options.
- **Give each question an `id`** you will recognise in code; ids are not sent to the model.

## Shapes (compose them in code)

| shape | what it is | what your code does |
|---|---|---|
| single | one question | act on the number |
| fan-out | many independent questions, one request | consume only the answers that apply; keep speculative ones cheap |
| gate | confidence-gated routing | act when P/confidence ≥ threshold; otherwise hold, escalate, or ask a human |
| composite | several `score` questions | normalize each over its own levels, weight, combine — weights are code, not prompts |

Intent routing needs no special shape: it is a `choice` question whose answer selects a handler.

Thresholds belong to the caller. Validate them on real data, and remember that a low-confidence answer is often
still usable for a harmless preference while it must not silently drive a consequential action.

## Backends

| backend | kind | serves |
|---|---|---|
| `typesafe` | decisions | TypeSafe System One (`/v1/systemone`) |
| `openrouter` | decisions | OpenRouter Decisions API (`/api/alpha/decisions`, Jev) |
| `llm` | chat | any OpenAI-compatible chat model, wrapped in the same contract |

- `auto` (default) prefers real Jev when a credential exists, then the chat proxy.
- Override per call with `backend` / `model`, and **per question** with `questions[i].backend` / `.model`. Questions
  sharing a backend+model are batched into one request; different batches run in parallel — so one call can combine a
  cheap calibration question with a second opinion from another model.
- An `llm` answer is a general model's best effort, not a calibrated System One answer: treat it as lower trust, and
  note that `confidence` is only present when that model returned one.

## Configure

```
/decide                    status: resolved config, per-backend state, live probe
/decide add [backend]      guided: key, model, base URL
/decide set [t] [f] [v]    one field, e.g. set openrouter model ~typesafe/jev-latest
/decide unset [t] [f]      drop a field or a whole backend block
/decide question [text]    ask your own question (shapes: single | fanout | gate | composite)
/decide models [b] [text]  catalogue of a backend
/decide setup              guided setup incl. the default backend
```

Aliases: `/jev` = `/decide`; `setup` = `init`, `unset` = `remove` = `rm`, `question` = `ask`.

Credentials come from the config file (`<agentDir>/decider.json`: `~/.pi/agent/` under pi, `~/.omp/agent/` under
omp) or the environment — `TYPESAFE_API_KEY`, `OPENROUTER_API_KEY`, `DECIDER_LLM_API_KEY`. Prefer `$ENV_VAR`
references over literal keys in the file.

## Limits

Jev is text-only (convert images/audio first), has a 64k request budget (≈32k for `state` plus the longest question),
streams nothing, and cannot call tools — it answers, your code acts. Long states cost input tokens on every call
(Jev does not bill output).

## Examples

Route a support message (one request, two independent judgments):

```jsonc
{
  "state": "Help! My payouts have been failing for 3 days.",
  "questions": [
    { "id": "urgent", "type": "noul", "instructions": "Does this message convey urgency?",
      "criteria": { "true": "Explicitly time-sensitive", "false": "No urgency expressed" } },
    { "id": "team", "type": "choice", "instructions": "Which team should handle this?",
      "criteria": { "billing": "Payments, invoicing, refunds", "technical": "Bugs, outages", "none": null } }
  ]
}
```

Then: `if (urgent.noul > 0.8 && team.choice !== "none") route(team.choice)` — the threshold is yours.

Grade a candidate with a gate and a composite:

```jsonc
{
  "state": { "candidate": "…", "requirements": "…" },
  "questions": [
    { "id": "meets_bar", "type": "noul", "instructions": "Does this candidate meet the stated bar?" },
    { "id": "depth", "type": "score", "instructions": "How deep is the relevant experience?",
      "criteria": ["None", "Adjacent", "Direct", "Deep"] },
    { "id": "evidence", "type": "score", "instructions": "How well is it evidenced?",
      "criteria": ["Claim only", "Some evidence", "Strong evidence"] }
  ]
}
```

Gate on `meets_bar.noul`, then take `0.7 × depth/(levels-1) + 0.3 × evidence/(levels-1)`.
