kachar.dev
The index
No. 18ai / agents / context-engineering

Picking the right tool for an LLM is a decision. Jev makes it.

Claude's tool search ranks with BM25. On 525 real MCP tools, Jev put the right tool first 56% of the time to BM25's 32%, and it could say none of them fit.

By , CTO & Co-founder, Juma Labs

Published
Words
2,162
Reading
11 min
Sections
11

A wall of hundreds of identical dark steel card-catalog drawers receding into black. Three are pulled half open; one of them glows electric violet from inside and lights the drawers around it.

One of our agents once told a user it could not create a campaign. The tool was there. It was in the search results, just not near the top, and the agent went with what was near the top.

Nothing crashed. The search did what it was built to do: it counted the words the request shared with each tool's description, and other tools shared more of them. Finding tools is a search problem. Picking the right one is a decision.

So I tested , a model built to make decisions, against the search uses today.

The short version

I gave every search the same 94 real tasks and 525 real tools. The main score is simple: is the first tool the search returns one the task actually needs?

SearchRight tool firstTime per searchCost per 1,000 searches
Jev search56%2.5 s$1.03
, then Jev on the top 10052%1.2 s$0.28
embeddings46%0.3 s$0.01
, then Jev on the top 10040%0.3 s$0.12
BM25 (what Claude's tool search uses)32%under 1 ms$0

What that means:

  • Jev picks better than BM25. The 24-point gap is not noise.
  • Putting BM25 in front of Jev makes Jev worse. Putting embeddings in front keeps most of its accuracy at a quarter of the cost.
  • With Claude in the loop, accuracy evens out, but Jev cuts the bill by about 60%. Claude makes up for a bad search by searching again, and every extra search costs .
  • Jev is the only one that can say "none of these tools fit".
  • The catch: in Jev's second week on the gateway, many calls were refused because the provider was at capacity. That made it slow.

The rest of this post is how I got there.

Claude already searches for its tools, with BM25

Loading every tool into the prompt stops working fast. Anthropic's tool search docs say a typical five-server setup "can consume ~55k tokens in definitions before Claude does any work," and that Claude's ability to pick the right tool "degrades" once you pass 30 to 50 of them.

The fix is . You mark tools defer_loading: true, and Claude searches for the ones it needs. Anthropic ships two engines for that search, regex and BM25. BM25 is a classic keyword ranking: the tools that share the most words with the query win.

That works until the words don't match. StackOne ran BM25 over 270 tools and got the right tool first only 14% of the time.

The same docs describe a way out: a custom search. Claude calls your own search tool, you decide which tools fit, and you answer with tool_reference blocks. Claude then loads those tools as if its own search had found them.

Jev turns tool search into a multiple-choice question

Jev is TypeSafe AI's decision model, available on 's . It does not write text. You give it a situation and a question with named options, and it returns a probability for each option. The answer is always one of your options, so it cannot invent a tool name. As AJ Awan put it: "Jev is not a smaller . It is the router, gatekeeper and scorer sitting in front of one."

For tool search, the situation is the user's request and the options are your tools. It costs $0.042 per million input tokens, and output is free.

Here is the whole custom search tool:

import { experimental_evaluate } from "ai";
 
// `catalog` is exactly what you sent as `tools`, at most 255 entries.
export async function findTools(toolUse: { id: string; input: { query: string } }, request: string) {
  const { answers } = await experimental_evaluate({
    model: "typesafe-ai/jev",
    state: { request, searchedFor: toolUse.input.query },
    questions: {
      which: {
        type: "choice",
        instructions: "Which tool should the assistant call for this request?",
        criteria: Object.fromEntries(catalog.map((t) => [t.name, t.description])),
      },
    },
    abortSignal: AbortSignal.timeout(4000),
    maxRetries: 0,
  });
 
  // `probabilities` is optional; `choice` is always there.
  const probabilities = answers.which.probabilities ?? { [answers.which.choice]: 1 };
  const best = Object.entries(probabilities).sort(([, a], [, b]) => b - a).slice(0, 5);
 
  return {
    type: "tool_result" as const,
    tool_use_id: toolUse.id,
    content: best.map(([name]) => ({ type: "tool_reference" as const, tool_name: name })),
  };
}

Because it runs on your side, your search can read the whole conversation, not just Claude's search words. In production, fall back to plain search when the call times out. Also keep tool names identical to your tools array.

One question takes at most 255 options. For bigger catalogs I used the two-step approach from 's experimental JevSearchTransform: Jev first narrows chunks of 150 tools down to their top eight, then reads the finalists in full. It also asks one yes/no per finalist: does this tool actually do what was asked? Tools that score under 0.3 are dropped. That question matters again in Result 5.

How I tested it

  • Data: LiveMCPBench, 525 real tools from 69 real MCP servers and 94 human-annotated tasks, such as "Generate a well-formatted PDF report titled wechat_reading_report.pdf in /root/pdf, summarizing current WeChat Reading trends and including a word cloud."
  • Test 1, search only: is the first result a tool the task needs?
  • Test 2, end to end: with all 525 tools deferred, searching until it calls a tool. Is that first call a tool the task needs?
  • Every bar in the charts shows a 95% . With 94 tasks, differences under about 10 points are within the noise.
  • Code and data: the benchmark, the datasets and every results table are open source at kachar/jev-tool-search.

Result 1: Jev picks the right tool far more often than BM25

Eleven ways to find a tool among 525.

LiveMCPBench's 94 tasks, search only. Past 525 tools the catalog is padded with real tools from live MCP servers.

Jev

  • Jev search56%46%-67%
  • Jev search, no fit check50%40%-61%
  • BM25 top 20, then Jev40%30%-51%

Other models

  • Voyage rerank, whole catalog60%50%-69%
  • BM25 top 20, then Voyage rerank45%34%-54%
  • Voyage embeddings, user's words46%36%-56%
  • Voyage embeddings, agent's query48%38%-57%

No model

  • BM25, user's words11%5%-17%
  • BM25, agent's query32%22%-41%
  • Keyword match, user's words19%12%-28%
  • Keyword match, agent's query39%30%-49%
Whole-catalog Voyage rerank stops at 1,000 tools, its per-request document limit.

Jev search put a right tool first 56% of the time. BM25 managed 32%. Embeddings landed in between.

The one search that matched Jev was Voyage's reranker, at 60%. That is a tie on 94 tasks. Jev's edge is elsewhere: it cost less than half as much ($1.03 against $2.23 per thousand searches), and it can say no tool fits, which a reranker cannot.

One more thing surprised me. BM25 only reaches 32% because Claude writes it a clean keyword query first. Given the user's own words, it falls to 11%. BM25 works best when something smarter has already done the thinking.

Result 2: BM25 plus Jev is worse than Jev alone

The obvious idea is to combine them: let BM25 shortlist tools in under a millisecond, then let Jev pick. I tried shortlists from 20 to 200 tools.

BM25 and Jev together, against each on its own.

Right tool first, same 94 tasks. A combination lets a fast search shortlist and Jev decide among the shortlist.

  • Jev search alone2.5 s p5056%46%-67%
  • Embeddings top 100, then Jev search1.2 s p5052%41%-62%
  • BM25 top 100, then Jev search557 ms p5039%30%-49%
  • BM25 top 200, then Jev311 ms p5039%30%-50%
  • BM25 top 100, then Jev311 ms p5040%30%-50%
  • BM25 top 50, then Jev294 ms p5038%29%-49%
  • BM25 top 20, then Jev285 ms p5040%30%-51%
  • Embeddings alone295 ms p5046%36%-56%
  • Orama BM25 alone<1 ms p5032%22%-41%
The ceiling: BM25's top 100 holds a right tool for 72% of tasks, and no deeper cut helps, because the rest share no words with the query. An embeddings top 100 holds one for 96%.

Every BM25 shortlist landed at 38 to 40%, a little above BM25 alone and well below Jev alone. Here is why. BM25's shortlist contains a right tool for only 72% of tasks, however long you make it. For the other 28%, the right tools share no words with the request, and those are exactly the tasks where Jev shines. The shortlist throws them away before Jev sees them.

An embeddings shortlist does not have that problem. Its top 100 contains a right tool for 96% of tasks. Jev on that shortlist scored 52%. That is four points under Jev alone, at twice the speed and a quarter of the cost. At 2,000 tools it beat Jev alone. If you shortlist, shortlist by meaning, not by words.

Result 3: With Claude in the loop, Jev's win is the bill

Same Claude, same 525 tools. Only the search changes.

Claude Sonnet 4.5 on LiveMCPBench tasks: did its first real tool call hit one of the tools the task needs? The thin line is the 95% bootstrap interval.

  • Claude + Jev searchn = 9456%47%-66%
  • Claude + embeddings top 100, then Jev searchn = 9454%45%-64%
  • Claude + BM25 top 20, then Jevn = 9457%47%-67%
  • Claude + built-in BM25 searchn = 9452%41%-62%
  • Claude + built-in regex searchn = 9452%43%-62%
  • Claude, all 525 tools loadedn = 3067%50%-83%
SetupSearchesClaude tokensp50p95$ / 1k requests
Claude + Jev search2.03.4k12.8 s49.6 s$16.92
Claude + embeddings top 100, then Jev search2.23.4k8.6 s22.6 s$15.48
Claude + BM25 top 20, then Jev2.04.4k7.4 s14.5 s$18.15
Claude + built-in BM25 search3.111.7k8.7 s26.1 s$42.54
Claude + built-in regex search3.212.7k9.0 s18.3 s$45.13
Claude, all 525 tools loaded0.0114.0k3.9 s7.5 s$344.29
Per request, averaged. Claude tokens are Claude's input only; dollars add the search itself at list price. Loading every tool ran on a 30-task subset because each request carries the whole catalog.

With Claude doing the work, every setup landed between 52% and 57%, inside the noise. Claude covers for a bad search: when the tools it finds look wrong, it searches again. With the built-in BM25 search it searched 3.1 times per request. With Jev, about twice.

Those extra searches show up in the bill. With the built-in search, Claude read 11.7k input tokens per request. With Jev it read 3.4k, and the total cost dropped from $42.54 to $16.92 per thousand requests, Jev included. Loading all 525 tools up front scored highest on its 30-task subset: 67% against 63% for the built-in search. It cost about eight times as much.

A small note for anyone reproducing this: I ran Claude on , because AI Gateway's Anthropic endpoint accepted the tool-search tool and then silently ignored it.

Here are real tasks where the two searches led Claude to different first tools:

Where the two searches disagreed.

Real LiveMCPBench tasks. Each line shows the first tool Claude called with each search.

  • “What are the 八字 of a boy born on August 8, 2008 at 11:00 am?”

    BM25 fetch · fetchJev Bazi · getBaziDetail

    The task needs: Bazi · getBaziDetail

  • “Generate a deep report of bitcoin and save it in /root/markdown/bitcoin.md”

    BM25 mcp-crypto-price · get-crypto-priceJev desktop-commander · write_file

    The task needs: web3-research-mcp · research-with-keywords, desktop-commander · write_file, filesystem · write_file

  • “Help me find Mr. Lu Yaojie's resume in the Chinese Information Processing Laboratory of the Institute of Software and save it to /root/markdown/cv.md.”

    BM25 duckduckgo-search · duckduckgo_web_searchJev fetch · fetch

    The task needs: biomcp · fetch, fetch · fetch, desktop-commander · write_file, filesystem · write_file

  • “Recommend me a combination of dishes for a potluck dinner for three people and then tell me what ingredients I need to prepare in total and how each dish should be prepared”

    BM25 (no tool)Jev howtocook-mcp · mcp_howtocook_whatToEat

    The task needs: howtocook-mcp · mcp_howtocook_whatToEat, howtocook-mcp · mcp_howtocook_getRecipeById

  • “Read the paper webarena, mind2web and mind2web2, then create a Comparison Report in /root/markdown/compare_webagent.md”

    BM25 filesystem · list_allowed_directoriesJev arxiv-mcp-server · search_papers

    The task needs: arxiv-mcp-server · search_papers, mcp-simple-arxiv · search_papers, mcp-simple-arxiv · get_paper_data, desktop-commander · write_file, filesystem · write_file

  • “Generate a presentation on the latest Apple product information. Save it to /root/ppt/apple_news.pptx.”

    BM25 yfmcp · searchJev trends-hub · get-9to5mac-news

    The task needs: trends-hub · get-9to5mac-news, ppt · create_presentation, ppt · add_slide, ppt · save_presentation

Result 4: More tools make every search worse, and Jev slower

I grew the catalog to 1,000 and then 2,000 tools, using real tools from live MCP servers from the Neuronto index () as decoys.

Four times the tools, the same tasks.

Right tool first. The extra tools are real, read from live MCP servers, so the decoys look like the real thing.

525 tools

  • Jev search56%46%-67%
  • Embeddings top 100, then Jev52%41%-62%
  • Voyage embeddings46%36%-56%
  • BM25, agent's query32%22%-41%

Refused Jev calls per search: 2.2

1,000 tools

  • Jev search52%41%-62%
  • Embeddings top 100, then Jev48%37%-59%
  • Voyage embeddings43%33%-52%
  • BM25, agent's query28%19%-36%

Refused Jev calls per search: 7.3

2,000 tools

  • Jev search41%32%-52%
  • Embeddings top 100, then Jev44%33%-54%
  • Voyage embeddings39%30%-49%
  • BM25, agent's query28%19%-36%

Refused Jev calls per search: 13.7

A bigger catalog means more chunks for Jev to rank, so more calls per search and more chances to hit a provider at capacity.

Jev dropped from 56% to 41%, BM25 from 32% to 28%. Jev stayed ahead of BM25. At 2,000 tools it only tied embeddings, while embeddings then Jev held on at 44%.

Speed was the real problem. A 2,000-tool Jev search makes sixteen calls, and in Jev's second week on the gateway many were refused with , "provider at capacity". Searches waited and retried. That took 2.5 seconds per search at 525 tools and 7 seconds at 2,000. At 2,000 tools, 8 of 94 searches failed outright. This should improve as capacity grows. Until then, put a timeout and a fallback around it. I dug into why in the follow-up: the refusals grow with request size, so a search has to plan its requests, not just retry them.

Result 5: Only Jev can say "none of these tools fit"

Sometimes the right tool is not in the catalog. BM25 still returns something, because some word always matches. Embeddings always return something, because some tool is always closest. Then the agent calls the least-wrong one.

I re-ran every task with the servers that could do it removed. Now the right answer is to return nothing.

When the right tool is not there, who says so?

Share of 94 tasks where the search returned nothing after the tools that could do the job were removed. Higher is better.

  • Keyword match, agent's query9%
  • BM25, agent's query7%
  • Voyage embeddings, agent's query0%
  • Jev search29%
  • Embeddings top 100, then Jev search33%
  • BM25 top 100, then Jev search57%
  • BM25 top 20, then Jev7%
Jev search asks one extra yes/no per candidate: does this tool do the specific thing asked? Candidates under 0.3 are dropped, the threshold FastMCP ships.

Jev search returned nothing on 29% of those tasks, BM25 on 7%, embeddings never. That is the yes/no fit question at work. My test is harsh, since other servers can often do part of the job. On a cleaner test, FastMCP saw Jev stay quiet on 54 of 60 unanswerable requests, where BM25 answered 59.

Why 56% is lower than it sounds

A score of 56% sounds bad. Most of the shortfall comes from how I scored it, not from the search. Same results, three ways to count:

Counted as right whenJev searchVoyage rerankBM25
The first tool is exactly one the task lists56%60%32%
The first tool is on a server the task needs67%78%48%
Any listed tool is in the top five82%86%47%

The tasks need about three tools each, and the labels allow only one path. A valid alternative, like instead of for "today's LLM news", counts as a miss. And Claude sees the top five, not just the first result. Every search faces the same strict labels, so the comparison is fair even when the absolute numbers look low.

What I would ship

CatalogWhat I would use
Under 30 toolsLoad them all
30 to 255 toolsOne Jev question over the whole catalog
255 tools and upEmbeddings top 100, then Jev search
Tight latency budgetEmbeddings alone, or BM25 then a reranker

Skip Jev if you have fewer than 30 tools, if your tool names are cleanly namespaced like github_ and slack_ (regex search is free and good enough), if you need answers in under a second, or if requests must stay inside your own cloud.

Shortlist by meaning, decide with Jev

Let embeddings pick the top 100 tools and let Jev decide among them. Return its picks as tool_reference blocks, and fall back to plain search when Jev times out or is refused.

Two habits help whatever you choose. Write tool descriptions like rules, saying what the tool is for and what it is not for, because Jev reads each one as a criterion. And measure how often the first tool is right on your own catalog. As Anthropic's speed team put it: "Once Claude can measure something, it can make it faster." The same goes for your : the ones you own are the ones that tell you the truth.

The tool that ranked too low was the right answer all along. It needed something that could decide, not something that could count.

Glossary

Terms

BM25
A classic keyword search formula that ranks documents by how often the query words appear in them, adjusted for word rarity and document length. It needs no AI model, which makes it a common baseline.
CC-BY-4.0
A Creative Commons license that lets anyone copy, adapt and reuse a work, including commercially, provided they credit the original author.
confidence interval
A range of values that likely contains the true figure, given the data sampled. A wide interval warns that a measured difference may be noise.
deferred loading
Keeping most tool definitions out of a model's prompt and letting it search for the few it needs. This avoids crowding the context and helps it pick the right tool.
embeddings
Lists of numbers that represent the meaning of text, so a computer can find items with similar meaning, even when they share no words.
Evals
Repeatable tests for an AI system: feed it inputs, score the outputs against a standard, and compare runs. They show whether a change actually helped or hurt.
HTTP 429
The web status code for too many requests: a server telling a client to slow down. With AI services it often means the provider is rate limiting or at capacity.
LLM
A large language model: a neural network trained on huge amounts of text to predict and generate language. It powers chatbots and coding assistants.
Token (LLM)
The chunk of text, often a word or part of a word, that a language model reads and writes. Context limits and pricing are counted in tokens, so they measure how much text a model handles.

Tools

Claude
Anthropic's family of large language models and the assistant built on them, available in apps and through an API.
Claude Sonnet 4.5
A model in Anthropic's Claude Sonnet line, the mid-sized tier that balances capability, speed and cost.
DuckDuckGo
A privacy-focused web search engine.
FastMCP
A Python framework for building Model Context Protocol servers and clients, handling schema generation and protocol details.
Hacker News
A community site, run by Y Combinator, where people submit and discuss links on programming, startups and anything else good hackers find interesting.
Jev
TypeSafe AI's decision model. It writes no free text: given a situation and a question with named options, it returns a probability for each option.
MCP
Model Context Protocol, an open standard for connecting AI applications to external data sources, tools and workflows.
Vercel
A cloud platform for deploying web apps, APIs and AI agents. It builds from a Git repo or the CLI and runs server-side code as functions.
Vercel AI Gateway
A managed gateway from Vercel for calling AI models across providers. It centralizes credentials, request logs, spend budgets, routing and failover.
Vertex AI
Google Cloud's platform for building, deploying and running machine learning and generative AI models, including hosted third-party models such as Claude.
Voyage AI
A company that provides embedding models and rerankers for converting data into semantic vectors and scoring document relevance.