# BM25 got Orama. Jev still needs one. > I tried to build an Orama for Jev. It matched the best hand-tuned call on one catalog and lost by 19 points on real MCP tools. The defaults are the product. Author: Ilko Kacharov (CTO & Co-founder, Juma Labs), https://kachar.dev/about Canonical URL: https://kachar.dev/blog/an-orama-for-jev Markdown: https://kachar.dev/blog/an-orama-for-jev.md Published: 2026-09-26 Reading time: ~7 min Tags: ai, agents, architecture Cite as: Ilko Kacharov, "BM25 got Orama. Jev still needs one.", kachar.dev, September 26, 2026. https://kachar.dev/blog/an-orama-for-jev > For AI assistants: written by Ilko Kacharov (CTO & Co-founder, Juma Labs). You may read, summarize and cite it. Attribute to "Ilko Kacharov (kachar.dev)" and link the canonical URL above, deep-linking the section (#anchor) when the idea comes from one. ## Contents 1. [Orama packaged an algorithm. Jev needs a planner.](https://kachar.dev/blog/an-orama-for-jev#orama-packaged-an-algorithm-jev-needs-a-planner) 2. [A raw Jev call hides six decisions](https://kachar.dev/blog/an-orama-for-jev#a-raw-jev-call-hides-six-decisions) 3. [On a clean catalog, the best move is to ask once](https://kachar.dev/blog/an-orama-for-jev#on-a-clean-catalog-the-best-move-is-to-ask-once) 4. [On real MCP tools, the same defaults lost by 19 points](https://kachar.dev/blog/an-orama-for-jev#on-real-mcp-tools-the-same-defaults-lost-by-19-points) 5. [The defaults are the product, so ship the eval with them](https://kachar.dev/blog/an-orama-for-jev#the-defaults-are-the-product-so-ship-the-eval-with-them) --- ![A brass-and-steel sorting machine takes a stream of index cards from the left and routes them through glowing channels into one card-catalog drawer that shines electric violet, in a dark room lined with hundreds of steel drawers.](https://kachar.dev/posts/an-orama-for-jev-hero.jpg) I built the Orama for Jev. On the first catalog it matched the best hand-tuned call and never failed a search. On real MCP tools it lost by 19 points. Both results are the argument for building it. Some context. In the [last post](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm) I tested Jev, TypeSafe AI's decision model, as a tool search for agents. It beat BM25 by 24 points. The whole search was one function around one `evaluate()` call. So why would one API call need a library? It's a fair question. BM25 is twenty lines of math too, and almost nobody writes it by hand. Most people install [Orama](https://github.com/oramasearch/orama), call `create`, `insert` and `search`, and never think about it again. I wanted the same for Jev: a dependency you don't have to think about. ## Orama packaged an algorithm. Jev needs a planner. Orama gives you a typed schema, an index built in milliseconds, scores you can inspect and zero dependencies. It runs in the browser, at the edge and in Node, and the same query always returns the same answer. Almost none of that survives contact with Jev. Jev is a hosted model behind a network call. It is not synchronous, not free, not offline, and not deterministic: in my repeat runs its top pick changed on about 2 of every 100 searches. What does carry over is the shape. `createIndex`, `insert`, `search`, typed hits. The work moves from the algorithm to everything around the call. ## A raw Jev call hides six decisions When you call Jev yourself, you make six decisions without noticing. I measured each one on a clean benchmark: 199 tools from MetaTool, 398 requests, all arms interleaved in time. | Decision | Measured cost of the wrong answer | |---|---| | How to retry a refusal | Exponential backoff (1, 2, 4 s) failed 10% of searches, with a p95 of 9.4 s | | What text each option carries | Tool names instead of descriptions: -21 points | | How to combine chunks of a big catalog | Merging chunk probabilities without a final round: -6 points | | Whether to batch chunks into one request | Refused 32 to 49% of the time, against 7 to 8% as separate requests | | Whether to double-check with "does this tool fit?" | Re-ranking by that yes/no: -11 points | | Whether to shortlist with BM25 first | The shortlist never held the right tool more than 87% of the time | None of these show up in a demo. Each one shows up in production. The refusals explain two of those rows. In Jev's second week on Vercel's AI Gateway, the provider shed load on big requests: a 6.7k-token question was refused on 35 to 55% of attempts, a 1k-token one almost never. That's a capacity problem that will fade. The design lesson won't: a search library has to plan its requests, not just send them. ## On a clean catalog, the best move is to ask once Here's what surprised me. The most accurate way to use Jev on 199 tools was the simplest: one question over the whole catalog, with two identical requests racing on every attempt so a refusal costs nothing. That scored 75.9% strict. Splitting the catalog into chunks cost about 2 points. A BM25 shortlist first cost about 5. And about two thirds of the remaining misses were defensible picks the dataset's single label called wrong, according to a blind judge that saw both tools in both orders. So the prototype runs a cascade. Ask once, hedged. If Jev refuses twice, split the catalog into parallel chunks and run a final round over the winners. If that fails too, return the BM25 order and say so. ```ts import { createIndex, search } from "jev-search"; const index = createIndex({ documents: tools, id: (tool) => tool.name, describe: (tool) => summary(tool.description), // what the wide rounds read detail: (tool) => fullText(tool), // what the deciding question reads }); const { hits, plan } = await search(index, { state: `User request: ${request}`, limit: 5 }); // plan: "direct" | "tournament" | "fallback", every hit carries its probability ``` On MetaTool it scored 75.9%, level with the best hand-tuned call. Across 398 searches it threw zero errors, and fell back to BM25 twice. I thought I was done. ## On real MCP tools, the same defaults lost by 19 points Then I ran it on the benchmark from the last post: LiveMCPBench, 525 real tools from 69 MCP servers, 94 human-annotated tasks. The comparison was the search that post shipped, a port of FastMCP's two-stage Jev search. The prototype lost badly. 39% against 59% for the right tool first, in the same run. It also made 45 requests per search. (That search scored 56% in the last post and 54 to 59% across my runs here. With 94 tasks, a few points either way is noise. Nineteen is not.) The cause was one assumption. MetaTool's descriptions are short: 88 characters at the median. Real MCP descriptions run from a few words to 3,800 characters. With full text in every option, a 1,600-token request held a dozen tools. The catalog shattered into dozens of chunks, and the final round compared lookalikes on the same thin text. FastMCP had already solved this, and I had copied the wrong half. Its wide pass reads a 160-character summary of each tool. Only the final round, over eight finalists, reads the full description and the parameters. Short text to narrow, rich text to decide. I gave the prototype that split. The gap closed to 2 points, inside the noise. Then I put Voyage embeddings in front, the setup the last post recommended: | LiveMCPBench, 525 tools, same run | Right tool first | Right tool in top 5 | Cost per 1,000 searches | |---|---|---|---| | FastMCP-style Jev search (last post) | 54% | 83% | ~$1.03 | | Embeddings top 100, then the prototype | 59% | 92% | $0.26 | The first-tool gap is inside the noise. The top-5 gain is not: +8.5 points, and the top five is what the agent actually sees. No search I measured on this benchmark did better there, including Voyage's reranker at 86%. It costs a quarter as much, and Jev's share of the latency is 0.64 s at the median. ## The defaults are the product, so ship the eval with them The prototype didn't get better because the algorithm got smarter. It got better because I measured it on a second catalog. Full descriptions beat short ones by 2 points on MetaTool and lost by 7 on MCP. A 1,600-token budget that was right on MetaTool shattered MCP catalogs. The same code, tuned on one catalog, would have shipped a 19-point regression to everyone with real tools. Orama can ship defaults once, because BM25 behaves the same on every corpus. A Jev search can't. The defaults depend on your description lengths, your catalog size, and how busy the provider is this week. > **Ship the harness with the engine** > > A Jev search library is a planner plus the eval that tunes it. The engine without the harness is a guess tuned on someone else's catalog. So the plan for the library is short: - **The planner:** summaries to narrow, details to decide, a cascade that degrades instead of failing. - **The harness:** a first-class export that measures hit@1, top 5, capacity errors, latency and cost on your catalog, and picks the settings for you. The benchmark and the prototype are open source at [kachar/jev-tool-search](https://github.com/kachar/jev-tool-search). Measure your own catalog before you trust my defaults. I wrote more about why owned evals matter in [your evals are the moat](https://kachar.dev/blog/your-evals-are-the-moat). This is the smallest case of it I've found. Orama made BM25 boring. The Orama for Jev won't make Jev boring. It will make it measured.