# dsh-tool-result-guard

English | [中文](README.zh.md)

**Zero-loss tool-result pruning for the [DeepSeek Harness (DSH)](https://www.npmjs.com/package/@deepseek-ai/dsh).** A `dsh-plugin`.

Oversized plain-text tool results are **spilled to a file first**, then replaced with a bounded head/tail preview whose position-0 marker carries the exact elided span and the spill locator. The model can always recover the middle; nothing is ever dropped without a durable copy.

```
[pruned: kept chars [0, 4096) + [28976, 30000) of 30000; the elided middle
[4096, 28976) is NOT lost — full output saved at: /tmp/dsh-spill/.../bash.txt.
Recover any elided span with the read tool (offset/limit), grep, or sed on that file.]

<head: first 4096 chars>

[... middle elided ...]

<tail: last 1024 chars>
```

## Why

DSH ships two built-in mechanisms:

| mechanism | trigger | behavior |
|---|---|---|
| `dsh-spill-policy` | result > `maxInlineBytes` (default 50000 **bytes**) at post-execute | full text to `ctx.spillStore`, bounded preview + locator returned — **zero-loss** |
| `dsh-compaction-tool-result-pruner` | text > 8192 chars, only **when compaction pressure qualifies** | head 4096 + `[... tool result middle pruned ...]` + tail 1024. The original event stays in the append-only session log for replay, but the model gets **no locator, no offsets, and no tool to read the log** — the middle is unrecoverable for the model |

The gap: results between ~8K chars and 50K bytes that survive until compaction are pruned **with no model-facing recovery path**. This plugin closes the gap by pruning earlier (at `tools/post-execute`), always spilling the full text first, and embedding the exact elided span `[head, total-tail)` plus the locator in the marker.

If you install this plugin, the built-in pruner finds nothing over 8K chars on the surface and becomes a no-op; you may keep or remove its row. Same for `spill-policy` (its cap is never reached). Keeping both is harmless.

## Install

One command:

```sh
dsh plugin --profile web add dsh-tool-result-guard
```

Restart DSH. Done — it now applies to every agent preset in that profile. Remove with `dsh plugin --profile web remove dsh-tool-result-guard`.

Tune the budgets by id in your profile's `~/.dsh/profiles/web/cordis.patch.yml`:

```yaml
- id: tool-result-guard
  config:
    thresholdChars: 16384
```

<details>
<summary><b>Alternative: preset-level install without pnpm</b> (single agent preset instead of the whole profile)</summary>

```sh
npx dsh-tool-result-guard install --preset my-preset --from standard
```

copies the shipped `standard` preset to `~/.dsh/.agent-presets/my-preset/`, drops `dsh-tool-result-guard.js` beside its `agent.cordis.yml`, and appends the plugin row; start a **new** session on `my-preset`. Patch an existing user preset with `--preset <id>` (no `--from`), undo with `remove --preset <id>`, print the snippet with `--print`. Manual equivalent — copy `index.js` next to the preset's `agent.cordis.yml` and append:

```yaml
- id: tool-result-guard
  name: './dsh-tool-result-guard.js'
```

</details>

## Config

Unknown keys fail at load. All budgets are Unicode code points (surrogate pairs never split).

| Key | Default | Meaning |
|---|---|---|
| `thresholdChars` | `8192` | Prune when the flattened plain-text result exceeds this many code points. |
| `headChars` | `4096` | Leading code points kept inline. |
| `tailChars` | `1024` | Trailing code points kept inline. |
| `excludeTools` | `["read"]` | Tools whose results always pass through. `read` is excluded by default to prevent a `read spill file → prune → read again` loop. Add e.g. `["read", "subagent", "memory_search"]`. |
| `spillDir` | *(unset)* | Override the local fallback spill directory (default: a private dir under the OS temp dir). |

`headChars + tailChars` must be below `thresholdChars` so the marker always fits.

## Behavior contract

- **Fail-open.** No session owner, no reachable spill backend, a write failure, or a replacement that wouldn't be smaller/in-budget → the original inline result is kept, unchanged. A prune never hides output and never turns a successful call into an error.
- **Spill first, prune second.** The full text is durable (via the deployment's `ctx.spillStore` when loaded, else a private per-process local directory) before the middle leaves the model-facing result.
- **Pass-through.** Non-accept decisions, value replacements, nested composite sub-calls, excluded tools, results containing non-text blocks (images are never touched), and at/under-threshold results are returned byte-identical.
- **Composable.** Runs as a prepended `tools/post-execute` waterfall listener and delegates via `next()`, so tool-owned projection and other hooks run first; their replaced content is what gets pruned. `additionalContexts` survive.
- **Idempotent.** A replacement is always within `thresholdChars`, so a second pass never re-prunes.

## How the model recovers the middle

The marker names the spill file and the exact char span. Models recover with `read` (offset/limit), `grep -n`, or `sed -n 'X,Yp'` on that file — the `read` exclusion prevents recovery reads from being pruned themselves.

## Development

```sh
npm test          # node:test, zero dependencies
```

## License

MIT
