kachar.dev
The index
No. 19ai / agents / architecture

BM25 got Orama. Jev still needs one.

I tried to build an Orama for Jev. It matched the best hand-tuned call on one catalog and lost by 19 points on real MCP tools. The defaults are the product.

By , CTO & Co-founder, Juma Labs

Published
Words
1,349
Reading
7 min
Sections
05

A brass-and-steel sorting machine takes a stream of index cards from the left and routes them through glowing channels into one card-catalog drawer that shines electric violet, in a dark room lined with hundreds of steel drawers.

I built the for . On the first catalog it matched the best hand-tuned call and never failed a search. On real tools it lost by 19 points. Both results are the argument for building it.

Some context. In the last post I tested Jev, TypeSafe AI's decision model, as a tool search for agents. It beat by 24 points. The whole search was one function around one evaluate() call.

So why would one API call need a library? It's a fair question. BM25 is twenty lines of math too, and almost nobody writes it by hand. Most people install Orama, call create, insert and search, and never think about it again. I wanted the same for Jev: a dependency you don't have to think about.

Orama packaged an algorithm. Jev needs a planner.

Orama gives you a typed schema, an index built in milliseconds, scores you can inspect and zero dependencies. It runs in the browser, at the edge and in Node, and the same query always returns the same answer.

Almost none of that survives contact with Jev. Jev is a hosted model behind a network call. It is not synchronous, not free, not offline, and not deterministic: in my repeat runs its top pick changed on about 2 of every 100 searches.

What does carry over is the shape. createIndex, insert, search, typed hits. The work moves from the algorithm to everything around the call.

A raw Jev call hides six decisions

When you call Jev yourself, you make six decisions without noticing. I measured each one on a clean benchmark: 199 tools from MetaTool, 398 requests, all arms interleaved in time.

DecisionMeasured cost of the wrong answer
How to retry a refusal (1, 2, 4 s) failed 10% of searches, with a of 9.4 s
What text each option carriesTool names instead of descriptions: -21 points
How to combine chunks of a big catalogMerging chunk probabilities without a final round: -6 points
Whether to batch chunks into one requestRefused 32 to 49% of the time, against 7 to 8% as separate requests
Whether to double-check with "does this tool fit?"Re-ranking by that yes/no: -11 points
Whether to shortlist with BM25 firstThe shortlist never held the right tool more than 87% of the time

None of these show up in a demo. Each one shows up in production.

The refusals explain two of those rows. In Jev's second week on 's , the provider shed load on big requests: a 6.7k-token question was refused on 35 to 55% of attempts, a 1k-token one almost never. That's a capacity problem that will fade. The design lesson won't: a search library has to plan its requests, not just send them.

On a clean catalog, the best move is to ask once

Here's what surprised me. The most accurate way to use Jev on 199 tools was the simplest: one question over the whole catalog, with two identical requests racing on every attempt so a refusal costs nothing. That scored 75.9% strict.

Splitting the catalog into chunks cost about 2 points. A BM25 shortlist first cost about 5. And about two thirds of the remaining misses were defensible picks the dataset's single label called wrong, according to a blind judge that saw both tools in both orders.

So the prototype runs a cascade. Ask once, hedged. If Jev refuses twice, split the catalog into parallel chunks and run a final round over the winners. If that fails too, return the BM25 order and say so.

import { createIndex, search } from "jev-search";
 
const index = createIndex({
  documents: tools,
  id: (tool) => tool.name,
  describe: (tool) => summary(tool.description), // what the wide rounds read
  detail: (tool) => fullText(tool), // what the deciding question reads
});
 
const { hits, plan } = await search(index, { state: `User request: ${request}`, limit: 5 });
// plan: "direct" | "tournament" | "fallback", every hit carries its probability

On MetaTool it scored 75.9%, level with the best hand-tuned call. Across 398 searches it threw zero errors, and fell back to BM25 twice. I thought I was done.

On real MCP tools, the same defaults lost by 19 points

Then I ran it on the benchmark from the last post: , 525 real tools from 69 MCP servers, 94 human-annotated tasks. The comparison was the search that post shipped, a port of 's two-stage Jev search.

The prototype lost badly. 39% against 59% for the right tool first, in the same run. It also made 45 requests per search. (That search scored 56% in the last post and 54 to 59% across my runs here. With 94 tasks, a few points either way is noise. Nineteen is not.)

The cause was one assumption. MetaTool's descriptions are short: 88 characters at the median. Real MCP descriptions run from a few words to 3,800 characters. With full text in every option, a 1,600-token request held a dozen tools. The catalog shattered into dozens of chunks, and the final round compared lookalikes on the same thin text.

FastMCP had already solved this, and I had copied the wrong half. Its wide pass reads a 160-character summary of each tool. Only the final round, over eight finalists, reads the full description and the parameters. Short text to narrow, rich text to decide.

I gave the prototype that split. The gap closed to 2 points, inside the noise. Then I put in front, the setup the last post recommended:

LiveMCPBench, 525 tools, same runRight tool firstRight tool in top 5Cost per 1,000 searches
FastMCP-style Jev search (last post)54%83%~$1.03
Embeddings top 100, then the prototype59%92%$0.26

The first-tool gap is inside the noise. The top-5 gain is not: +8.5 points, and the top five is what the agent actually sees. No search I measured on this benchmark did better there, including Voyage's reranker at 86%. It costs a quarter as much, and Jev's share of the latency is 0.64 s at the median.

The defaults are the product, so ship the eval with them

The prototype didn't get better because the algorithm got smarter. It got better because I measured it on a second catalog. Full descriptions beat short ones by 2 points on MetaTool and lost by 7 on MCP. A 1,600- that was right on MetaTool shattered MCP catalogs. The same code, tuned on one catalog, would have shipped a 19-point regression to everyone with real tools.

Orama can ship defaults once, because BM25 behaves the same on every corpus. A Jev search can't. The defaults depend on your description lengths, your catalog size, and how busy the provider is this week.

Ship the harness with the engine

A Jev search library is a planner plus the eval that tunes it. The engine without the harness is a guess tuned on someone else's catalog.

So the plan for the library is short:

  • The planner: summaries to narrow, details to decide, a cascade that degrades instead of failing.
  • The harness: a first-class export that measures , top 5, capacity errors, latency and cost on your catalog, and picks the settings for you.

The benchmark and the prototype are open source at kachar/jev-tool-search. Measure your own catalog before you trust my defaults.

I wrote more about why owned matter in your evals are the moat. This is the smallest case of it I've found.

Orama made BM25 boring. The Orama for Jev won't make Jev boring. It will make it measured.

Glossary

Terms

BM25
A classic keyword search formula that ranks documents by how often the query words appear in them, adjusted for word rarity and document length. It needs no AI model, which makes it a common baseline.
embeddings
Lists of numbers that represent the meaning of text, so a computer can find items with similar meaning, even when they share no words.
Evals
Repeatable tests for an AI system: feed it inputs, score the outputs against a standard, and compare runs. They show whether a change actually helped or hurt.
Exponential backoff
A retry strategy where the wait before each new attempt grows, usually doubling, so a struggling service gets room to recover instead of being hammered with repeats.
hit@1
A search metric: the share of queries where the correct item is the very first result returned. It is the strictest way to score a ranking.
LiveMCPBench
A benchmark that tests whether AI agents can find and use the right tools among a large collection of MCP servers (a standard way to give models tools). Used to compare tool-selection methods.
p95
A speed measure: the response time that all but the slowest twentieth of requests beat. It shows how bad the slow cases get, which an average hides.
Token (LLM)
The chunk of text, often a word or part of a word, that a language model reads and writes. Context limits and pricing are counted in tokens, so they measure how much text a model handles.

Tools

FastMCP
A Python framework for building Model Context Protocol servers and clients, handling schema generation and protocol details.
Jev
TypeSafe AI's decision model. It writes no free text: given a situation and a question with named options, it returns a probability for each option.
MCP
Model Context Protocol, an open standard for connecting AI applications to external data sources, tools and workflows.
Orama
A JavaScript search engine and RAG library that runs in the browser, on a server or at the edge. It supports full-text, vector and hybrid search.
Vercel
A cloud platform for deploying web apps, APIs and AI agents. It builds from a Git repo or the CLI and runs server-side code as functions.
Vercel AI Gateway
A managed gateway from Vercel for calling AI models across providers. It centralizes credentials, request logs, spend budgets, routing and failover.
Voyage AI
A company that provides embedding models and rerankers for converting data into semantic vectors and scoring document relevance.