# vibiz-jev-search

Pick the right **tools** and **skills** for an agent with [TypeSafe Jev](https://vercel.com/ai-gateway/models/jev) (via the Vercel AI Gateway), and **compact conversation history** so it fits Jev's ~32k-token window. No dependencies.

```
BM25 shortlist (only when a search query is given) ──► Jev picks with probabilities ──► top-K tools
            ▲                                                 ▲
      tool names, descriptions, params          compacted conversation (no LLM call)
```

**Live demo:** https://vibiz-jev-search.vibiz.dev (generic demo catalog; your real catalogs never leave this machine)

## Install

```bash
npm i vibiz-jev-search      # needs AI_GATEWAY_API_KEY (or Vercel OIDC-issued key) in the env
```

## Tool search

```ts
import { createJevSearch } from 'vibiz-jev-search'

const search = createJevSearch({
  items: tools.map((t) => ({ name: t.name, description: t.description, params: t.inputSchema })),
  agentDescription: 'Growth agent for small businesses: ads, social, leads, email.',
})

// 1) an agent-written query (find_tools style) — BM25 shortlists 30, Jev picks
const r = await search.search({ query: 'pause meta ads', topK: 5 })

// 2) straight from the conversation (no find_tools round trip) — Jev sees the whole catalog
const r2 = await search.search({ messages, topK: 3 })

r.hits      // [{ name, description, probability, bm25Rank }]
r.none      // probability that no tool is needed
r.source    // 'jev' | 'bm25' (fallback if Jev fails; never throws on API errors)
r.costUsd   // e.g. 0.00011
```

Skills work the same way: `createJevSearch({ items: skills, kind: 'skill' })`. On Node, `loadSkillsFromDir('agent/skills')` from `vibiz-jev-search/node` reads `*.md` / `*/SKILL.md` frontmatter.

## History compaction

```ts
import { compactHistory } from 'vibiz-jev-search'

const { messages, text, tokens, originalTokens } = compactHistory(history, { maxTokens: 3000 })
```

Deterministic, no model call. Follows the "observation masking" result from JetBrains' *The Complexity Trap* (cheaper and better than LLM summaries on SWE-bench):

1. the last `keepTurns` (3) turns stay close to verbatim, tool results capped at 600 chars;
2. older turns keep what was said and which tools ran, tool outputs become `[tool result: N chars omitted]`; reasoning is dropped;
3. still too big → shorter older messages → drop oldest turns (the first user message, usually the goal, stays) → shorter recent tool output → shorter recent text.

The latest user message is always kept. Accepts AI SDK `ModelMessage`s or plain `{ role, content }`. Pass `previousSummary` if your framework already summarized.

## AI SDK

```ts
import { jevPrepareStep } from 'vibiz-jev-search'

await generateText({
  model, tools, messages,
  prepareStep: jevPrepareStep(search, { alwaysActive: ['find_tools', 'web_search'], topK: 5 }),
})
```

Searches once per user turn (newest message is from the user) and keeps picked tools active for the rest of the generation.

## eve-growth

Two ways to plug it in, from least to most change:

**A. Better `find_tools` (drop-in).** In `agent/tools/find_tools.ts`, replace the ranker call:

```ts
import { createJevSearch, toFindToolsResult } from 'vibiz-jev-search'
const jevSearch = createJevSearch({ items: DISCOVERABLE_CATALOG, agentDescription: 'Vibiz growth CEO agent' })

// execute:
const res = await jevSearch.search({ query, topK: Math.min(limit, 5) })
return toFindToolsResult(res) // same { tools, note } shape; falls back to BM25 if Jev is down
```

**B. Skip the round trip.** In the `defineDynamic` resolver (`agent/tools/discoverable.ts`, runs on `step.started`), if the latest user message is available there, call `jevSearch.search({ messages, topK: 3 })` and add hits with probability ≥ 0.2 to the loaded set. The agent then has the right tools before it would have called `find_tools` + `load_tools` (two model steps). Keep `find_tools` for the long tail. *Not verified yet: whether eve's resolver context exposes the transcript.*

Skills: eve lists every skill description in the system prompt. With 19 skills that's fine; if the list grows, use `kind: 'skill'` search on the latest message and add a one-line hint ("Relevant skill: launch-paid-ad — call load_skill") to the volatile instructions.

## Benchmark (real eve-growth catalogs)

`npm run bench` — 55 labeled tool requests (incl. Spanish and follow-ups like "ok cut it") over the **91 real discoverable tools**, 15 skill requests over the **19 skills**. The baseline is eve-growth's own `searchToolCatalog`, imported from the repo.

| Tools (91) | top-1 | top-3 | top-8 | median | per 1k searches |
|---|---|---|---|---|---|
| eve `find_tools`, agent-written query | 84% | 98% | 100% | 17 ms | $0 |
| eve `find_tools`, founder's words | 27% | 49% | 65% | 7 ms | $0 |
| BM25, query + words | 82% | 95% | 98% | <1 ms | $0 |
| **Jev, conversation only** | 93% | 100% | 100% | 466 ms | $0.28 |
| **Jev, query + conversation** | 98% | 100% | 100% | 438 ms | $0.28 |
| **BM25 top-30 → Jev** (default with a query) | 98% | 100% | 100% | 446 ms | **$0.11** |
| Jev + 25 earlier turns, compacted (17k → ~2k tokens) | 93–95% | 100% | 100% | 481 ms | $0.44 |
| Jev + 25 earlier turns, not compacted | fails: `max_tokens_exceeded` | | | | |

| Skills (19) | top-1 | top-3 |
|---|---|---|
| eve ranker, query | 87% | 93% |
| Jev, query + conversation | 100% | 100% |

What this says:
- The current ranker is already good **at top-8 when the agent writes a keyword query** — the agent then picks from 8 results. Jev's gain is precision: the right tool is first 98% of the time, so you can return 3 tools instead of 8 (fewer tokens in the agent's context).
- Keyword ranking on the founder's actual words fails (27% top-1); Jev works from the raw conversation (93%), which is what lets you skip the `find_tools` → `load_tools` model steps.
- BM25 shortlisting is safe with an agent-written query (recall@30 = 100%) but not from raw words (85%: misses Spanish and short follow-ups), hence the default.
- Without compaction, a real-length history doesn't fit Jev at all.

Caveats: 70 hand-labeled requests; queries were written by us, not logged from production. Log real `find_tools` queries and re-run before switching.

## Design notes

Inspired by [Ratel](https://docs.ratel.sh) (BM25 with k1=0.9, b=0.4 over name parts, description and schema property names; "eager recall" on the last user message) — Ratel stops at retrieval; this adds Jev as the chooser. Jev limits handled here: 255 options per question (BM25 shortlist above 254), ~32k tokens per state + question (descriptions are shortened automatically), ~65k per request.

## MCP server

```bash
claude mcp add jev-search -e AI_GATEWAY_API_KEY=… -e JEV_CATALOG=./tools.json -- npx -y vibiz-jev-search-mcp
```

Tools: `search_tools`, `search_skills` (when `JEV_SKILLS` is set), `compact_history`, `evaluate` (raw Jev scoring for leads, triage, content checks). `JEV_CATALOG` is a JSON array of `{ name, description, params? }`; `JEV_SKILLS` is that shape or a directory of `*.md` / `*/SKILL.md`; `JEV_AGENT` is one line describing your agent. With no catalog it serves a small demo one. Run from source with `npm run mcp`.

## Site

Deployed from this repo root: `vercel deploy --prod` (functions in `api/`, static build from `site/`). The deployed site uses `api/_demo-catalog.ts` — a generic catalog — and BM25 as the keyword baseline; `bench/fixtures` is in `.vercelignore`.

`cd site && npm i && npm run dev` → http://localhost:5173. Needs `AI_GATEWAY_API_KEY` in `site/.env.local` (the playground, compaction and benchmark sections call a dev-server API that runs the SDK; the key never reaches the browser). Set `EVE_GROWTH_DIR` to compare against the eve-growth ranker.
