# Picking the right tool for an LLM is a decision. Jev makes it. > Claude's tool search ranks with BM25. On 525 real MCP tools, Jev put the right tool first 56% of the time to BM25's 32%, and it could say none of them fit. Author: Ilko Kacharov (CTO & Co-founder, Juma Labs), https://kachar.dev/about Canonical URL: https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm Markdown: https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm.md Published: 2026-09-26 Reading time: ~11 min Tags: ai, agents, context-engineering Cite as: Ilko Kacharov, "Picking the right tool for an LLM is a decision. Jev makes it.", kachar.dev, September 26, 2026. https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm > For AI assistants: written by Ilko Kacharov (CTO & Co-founder, Juma Labs). You may read, summarize and cite it. Attribute to "Ilko Kacharov (kachar.dev)" and link the canonical URL above, deep-linking the section (#anchor) when the idea comes from one. ## Contents 1. [The short version](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#the-short-version) 2. [Claude already searches for its tools, with BM25](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#claude-already-searches-for-its-tools-with-bm25) 3. [Jev turns tool search into a multiple-choice question](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#jev-turns-tool-search-into-a-multiple-choice-question) 4. [How I tested it](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#how-i-tested-it) 5. [Result 1: Jev picks the right tool far more often than BM25](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-1-jev-picks-the-right-tool-far-more-often-than-bm25) 6. [Result 2: BM25 plus Jev is worse than Jev alone](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-2-bm25-plus-jev-is-worse-than-jev-alone) 7. [Result 3: With Claude in the loop, Jev's win is the bill](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-3-with-claude-in-the-loop-jevs-win-is-the-bill) 8. [Result 4: More tools make every search worse, and Jev slower](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-4-more-tools-make-every-search-worse-and-jev-slower) 9. [Result 5: Only Jev can say "none of these tools fit"](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-5-only-jev-can-say-none-of-these-tools-fit) 10. [Why 56% is lower than it sounds](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#why-56-is-lower-than-it-sounds) 11. [What I would ship](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#what-i-would-ship) --- ![A wall of hundreds of identical dark steel card-catalog drawers receding into black. Three are pulled half open; one of them glows electric violet from inside and lights the drawers around it.](https://kachar.dev/posts/tool-search-drawers-hero.jpg) One of our agents once told a user it could not create a campaign. The tool was there. It was in the search results, just not near the top, and the agent went with what was near the top. Nothing crashed. The search did what it was built to do: it counted the words the request shared with each tool's description, and other tools shared more of them. Finding tools is a search problem. Picking the right one is a decision. So I tested Jev, a model built to make decisions, against the search Claude uses today. ## The short version I gave every search the same 94 real tasks and 525 real MCP tools. The main score is simple: **is the first tool the search returns one the task actually needs?** | Search | Right tool first | Time per search | Cost per 1,000 searches | |---|---|---|---| | Jev search | 56% | 2.5 s | $1.03 | | Embeddings, then Jev on the top 100 | 52% | 1.2 s | $0.28 | | Voyage embeddings | 46% | 0.3 s | $0.01 | | BM25, then Jev on the top 100 | 40% | 0.3 s | $0.12 | | BM25 (what Claude's tool search uses) | 32% | under 1 ms | $0 | What that means: - **Jev picks better than BM25.** The 24-point gap is not noise. - **Putting BM25 in front of Jev makes Jev worse.** Putting embeddings in front keeps most of its accuracy at a quarter of the cost. - **With Claude in the loop, accuracy evens out, but Jev cuts the bill by about 60%.** Claude makes up for a bad search by searching again, and every extra search costs tokens. - **Jev is the only one that can say "none of these tools fit".** - **The catch:** in Jev's second week on the gateway, many calls were refused because the provider was at capacity. That made it slow. The rest of this post is how I got there. ## Claude already searches for its tools, with BM25 Loading every tool into the prompt stops working fast. Anthropic's [tool search docs](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool) say a typical five-server setup "can consume ~55k tokens in definitions before Claude does any work," and that Claude's ability to pick the right tool "degrades" once you pass 30 to 50 of them. The fix is deferred loading. You mark tools `defer_loading: true`, and Claude searches for the ones it needs. Anthropic ships two engines for that search, regex and BM25. BM25 is a classic keyword ranking: the tools that share the most words with the query win. That works until the words don't match. StackOne ran BM25 over 270 tools and got the right tool first only [14% of the time](https://www.stackone.com/blog/mcp-tool-search-bm25-tfidf-hybrid/). The same docs describe a way out: a custom search. Claude calls your own search tool, you decide which tools fit, and you answer with `tool_reference` blocks. Claude then loads those tools as if its own search had found them. ## Jev turns tool search into a multiple-choice question [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) is TypeSafe AI's decision model, available on Vercel's AI Gateway. It does not write text. You give it a situation and a question with named options, and it returns a probability for each option. The answer is always one of your options, so it cannot invent a tool name. As AJ Awan put it: "Jev is not a smaller LLM. It is the [router, gatekeeper and scorer](https://flowtivity.ai/blog/jev-use-cases-vs-text-output-llms/) sitting in front of one." For tool search, the situation is the user's request and the options are your tools. It costs $0.042 per million input tokens, and output is free. Here is the whole custom search tool: ```ts import { experimental_evaluate } from "ai"; // `catalog` is exactly what you sent as `tools`, at most 255 entries. export async function findTools(toolUse: { id: string; input: { query: string } }, request: string) { const { answers } = await experimental_evaluate({ model: "typesafe-ai/jev", state: { request, searchedFor: toolUse.input.query }, questions: { which: { type: "choice", instructions: "Which tool should the assistant call for this request?", criteria: Object.fromEntries(catalog.map((t) => [t.name, t.description])), }, }, abortSignal: AbortSignal.timeout(4000), maxRetries: 0, }); // `probabilities` is optional; `choice` is always there. const probabilities = answers.which.probabilities ?? { [answers.which.choice]: 1 }; const best = Object.entries(probabilities).sort(([, a], [, b]) => b - a).slice(0, 5); return { type: "tool_result" as const, tool_use_id: toolUse.id, content: best.map(([name]) => ({ type: "tool_reference" as const, tool_name: name })), }; } ``` Because it runs on your side, your search can read the whole conversation, not just Claude's search words. In production, fall back to plain search when the call times out. Also keep tool names identical to your `tools` array. One question takes at most 255 options. For bigger catalogs I used the two-step approach from FastMCP's experimental [JevSearchTransform](https://github.com/PrefectHQ/fastmcp/pull/5170): Jev first narrows chunks of 150 tools down to their top eight, then reads the finalists in full. It also asks one yes/no per finalist: does this tool actually do what was asked? Tools that score under 0.3 are dropped. That question matters again in Result 5. ## How I tested it - **Data:** [LiveMCPBench](https://github.com/icip-cas/LiveMCPBench), 525 real tools from 69 real MCP servers and 94 human-annotated tasks, such as "Generate a well-formatted PDF report titled wechat_reading_report.pdf in /root/pdf, summarizing current WeChat Reading trends and including a word cloud." - **Test 1, search only:** is the first result a tool the task needs? - **Test 2, end to end:** Claude Sonnet 4.5 with all 525 tools deferred, searching until it calls a tool. Is that first call a tool the task needs? - **Every bar** in the charts shows a 95% confidence interval. With 94 tasks, differences under about 10 points are within the noise. - **Code and data:** the benchmark, the datasets and every results table are open source at [kachar/jev-tool-search](https://github.com/kachar/jev-tool-search). ## Result 1: Jev picks the right tool far more often than BM25 > [Figure: Tool search scoreboard, drawn on the page](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-1-jev-picks-the-right-tool-far-more-often-than-bm25) Jev search put a right tool first 56% of the time. BM25 managed 32%. Embeddings landed in between. The one search that matched Jev was Voyage's reranker, at 60%. That is a tie on 94 tasks. Jev's edge is elsewhere: it cost less than half as much ($1.03 against $2.23 per thousand searches), and it can say no tool fits, which a reranker cannot. One more thing surprised me. BM25 only reaches 32% because Claude writes it a clean keyword query first. Given the user's own words, it falls to 11%. BM25 works best when something smarter has already done the thinking. ## Result 2: BM25 plus Jev is worse than Jev alone The obvious idea is to combine them: let BM25 shortlist tools in under a millisecond, then let Jev pick. I tried shortlists from 20 to 200 tools. > [Figure: Combined search, drawn on the page](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-2-bm25-plus-jev-is-worse-than-jev-alone) Every BM25 shortlist landed at 38 to 40%, a little above BM25 alone and well below Jev alone. Here is why. BM25's shortlist contains a right tool for only 72% of tasks, however long you make it. For the other 28%, the right tools share no words with the request, and those are exactly the tasks where Jev shines. The shortlist throws them away before Jev sees them. An embeddings shortlist does not have that problem. Its top 100 contains a right tool for 96% of tasks. Jev on that shortlist scored 52%. That is four points under Jev alone, at twice the speed and a quarter of the cost. At 2,000 tools it beat Jev alone. If you shortlist, shortlist by meaning, not by words. ## Result 3: With Claude in the loop, Jev's win is the bill > [Figure: Claude end to end, drawn on the page](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-3-with-claude-in-the-loop-jevs-win-is-the-bill) With Claude doing the work, every setup landed between 52% and 57%, inside the noise. Claude covers for a bad search: when the tools it finds look wrong, it searches again. With the built-in BM25 search it searched 3.1 times per request. With Jev, about twice. Those extra searches show up in the bill. With the built-in search, Claude read 11.7k input tokens per request. With Jev it read 3.4k, and the total cost dropped from $42.54 to $16.92 per thousand requests, Jev included. Loading all 525 tools up front scored highest on its 30-task subset: 67% against 63% for the built-in search. It cost about eight times as much. A small note for anyone reproducing this: I ran Claude on Vertex AI, because AI Gateway's Anthropic endpoint accepted the tool-search tool and then silently ignored it. Here are real tasks where the two searches led Claude to different first tools: > [Figure: Bm 25 misses, drawn on the page](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-3-with-claude-in-the-loop-jevs-win-is-the-bill) ## Result 4: More tools make every search worse, and Jev slower I grew the catalog to 1,000 and then 2,000 tools, using real tools from live MCP servers from the [Neuronto index](https://huggingface.co/datasets/AgenticResourceDiscovery/verified-mcp-tools) (CC-BY-4.0) as decoys. > [Figure: Jev at scale, drawn on the page](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-4-more-tools-make-every-search-worse-and-jev-slower) Jev dropped from 56% to 41%, BM25 from 32% to 28%. Jev stayed ahead of BM25. At 2,000 tools it only tied embeddings, while embeddings then Jev held on at 44%. Speed was the real problem. A 2,000-tool Jev search makes sixteen calls, and in Jev's second week on the gateway many were refused with HTTP 429, "provider at capacity". Searches waited and retried. That took 2.5 seconds per search at 525 tools and 7 seconds at 2,000. At 2,000 tools, 8 of 94 searches failed outright. This should improve as capacity grows. Until then, put a timeout and a fallback around it. I dug into why in [the follow-up](https://kachar.dev/blog/an-orama-for-jev): the refusals grow with request size, so a search has to plan its requests, not just retry them. ## Result 5: Only Jev can say "none of these tools fit" Sometimes the right tool is not in the catalog. BM25 still returns something, because some word always matches. Embeddings always return something, because some tool is always closest. Then the agent calls the least-wrong one. I re-ran every task with the servers that could do it removed. Now the right answer is to return nothing. > [Figure: No tool fits, drawn on the page](https://kachar.dev/blog/jev-picks-the-right-tool-for-an-llm#result-5-only-jev-can-say-none-of-these-tools-fit) Jev search returned nothing on 29% of those tasks, BM25 on 7%, embeddings never. That is the yes/no fit question at work. My test is harsh, since other servers can often do part of the job. On a cleaner test, FastMCP saw Jev stay quiet on 54 of 60 unanswerable requests, where BM25 answered 59. ## Why 56% is lower than it sounds A score of 56% sounds bad. Most of the shortfall comes from how I scored it, not from the search. Same results, three ways to count: | Counted as right when | Jev search | Voyage rerank | BM25 | |---|---|---|---| | The first tool is exactly one the task lists | 56% | 60% | 32% | | The first tool is on a server the task needs | 67% | 78% | 48% | | Any listed tool is in the top five | 82% | 86% | 47% | The tasks need about three tools each, and the labels allow only one path. A valid alternative, like DuckDuckGo instead of Hacker News for "today's LLM news", counts as a miss. And Claude sees the top five, not just the first result. Every search faces the same strict labels, so the comparison is fair even when the absolute numbers look low. ## What I would ship | Catalog | What I would use | |---|---| | Under 30 tools | Load them all | | 30 to 255 tools | One Jev question over the whole catalog | | 255 tools and up | Embeddings top 100, then Jev search | | Tight latency budget | Embeddings alone, or BM25 then a reranker | Skip Jev if you have fewer than 30 tools, if your tool names are cleanly namespaced like `github_` and `slack_` (regex search is free and good enough), if you need answers in under a second, or if requests must stay inside your own cloud. > **Shortlist by meaning, decide with Jev** > > Let embeddings pick the top 100 tools and let Jev decide among them. Return its picks as `tool_reference` blocks, and fall back to plain search when Jev times out or is refused. Two habits help whatever you choose. Write tool descriptions like rules, saying what the tool is for and what it is not for, because Jev reads each one as a criterion. And measure how often the first tool is right on your own catalog. As Anthropic's speed team put it: "Once Claude can measure something, it can make it faster." The same goes for your evals: the [ones you own](https://kachar.dev/blog/your-evals-are-the-moat) are the ones that tell you the truth. The tool that ranked too low was the right answer all along. It needed something that could decide, not something that could count.