How AI Coding Agents Find Code: codebase-memory-mcp vs jevgrep vs Grep, With the Benchmarks Read Closely
Coding agents spend much of a task finding the right files. Three approaches: grep-and-read, a code knowledge graph over MCP (codebase-memory-mcp) and model-judged retrieval (jevgrep). Architecture, the real benchmark numbers, the gotchas, and which to use on a customer's codebase.
BY ARJUN MEHTA · FDEINTERVIEWS EDITORIAL · UPDATED OCTOBER 7, 2026 · 10 MIN READ
PRACTICE THIS:Model Context Protocol (MCP), the concept ·RAG and agent design questions ·System design questions ·The must-know FDE questions
Coding agents find code in three ways: they grep and read files one by one, they query a prebuilt structural index of the repository, or they ask a model to judge which files and functions are relevant. Two open-source projects represent the newer two. codebase-memory-mcp (MIT, about 46,000 GitHub stars as of 7 October 2026) parses a repository into a knowledge graph and serves it over MCP. jevgrep (MIT, released 26 September 2026) answers a plain-English question with files and verbatim excerpts, using TypeSafe's Jev model to judge relevance. Both publish measurements, and both are more interesting read closely than read from the README banner. The graph agent used about ten times fewer tokens than a grep-and-read agent but scored 0.83 against 0.92 on answer quality. jevgrep matched its baseline's 8 of 10 SWE-bench solves at 25.8 percent lower total cost, with most of the saving coming from one task.
Why finding code is the expensive part
Watch a coding agent on an unfamiliar repository and most of its early turns are search: list a directory, grep for a name, open three files, grep for what those files call, open two more. Each read lands in the context window and on the bill. The Codebase-Memory paper frames the mismatch well: the questions developers ask are structural (who calls this, what breaks if I change it, where are the module boundaries), but the agent's tools are textual. Following a call chain with grep means one search and one read per hop, so tool calls and tokens grow with the size of the codebase.
The diagram shows the three approaches answering the same question.
codebase-memory-mcp: the repository as a graph
What it is. A single statically linked C binary with no runtime dependencies, for macOS, Linux and Windows. It indexes a repository into a graph stored in SQLite under ~/.cache/codebase-memory-mcp/ and exposes it to any MCP-capable agent. It contains no language model. When you ask your agent "what calls ProcessOrder?", the agent calls trace_path and phrases the result, so there is no extra API key or second model to pay for.
How the index is built. Parsing runs in two layers. A tree-sitter pass handles every supported language: the paper reported 66 languages, and the README now lists 158 to 162 grammars compiled into the binary. It extracts definitions, calls and imports. A second "Hybrid LSP" pass, a C reimplementation of type-resolution ideas from language servers such as pyright, gopls and rust-analyzer, refines call edges for Python, TypeScript and JavaScript, PHP, C#, Go, C, C++, Java, Kotlin, Rust and Perl. Without it, user.profile.display_name() is just a name. With it, the call resolves to the method three modules away. The paper describes a six-strategy call-resolution cascade in which the first three strategies resolve about 80 percent of calls in well-structured codebases. On top of the graph sit Louvain community detection (to find functional modules), git-diff impact mapping and HTTP route matching across services.
The tools an agent sees. Seventeen MCP tools in the current README. The ones that change how an agent works:
| Tool | What it answers |
|---|---|
search_graph | Find symbols by name pattern, label or degree; BM25 and semantic search over the graph |
trace_path | Who calls this, and what it calls, breadth-first up to depth 5 |
detect_changes | Which symbols your uncommitted diff touches, with a blast-radius risk class |
query_graph | A read-only openCypher subset, for example finding functions with no callers |
get_architecture | Languages, packages, entry points, routes, hotspots and clusters in one call |
get_code_snippet | The source of one function by qualified name |
A dead-code query, written in the supported subset, looks like this:
MATCH (f:Function)
WHERE NOT EXISTS { (f)<-[:CALLS]-() }
RETURN f.name LIMIT 20
Speed. The README reports, on one Apple M3 Pro machine, a full index of the Linux kernel (28 million lines, 75,000 files, 4.81 million nodes and 7.72 million edges) in about 3 minutes, Django in about 6 seconds, and graph queries in under a millisecond. The paper measured caller traversal at about 0.3 ms.
Reading the benchmark closely
The paper (arXiv 2603.27277, March 2026) is the evidence to quote, not the README badges. It ran 12 question categories against one real repository for each of 31 languages, with two agents on the same model: one with the graph tools, one that greps and reads files.
| Graph agent | Grep-and-read agent | |
|---|---|---|
| Answer quality (0 to 1) | 0.83 | 0.92 |
| Tool calls per question | 2.3 | 4.8 |
| Tokens per question | about 1,000 | about 10,000 |
| Query latency | under 1 ms | 10 to 30 s |
The graph agent reached 90 percent of the explorer's quality at about a tenth of the tokens. It did better on hub detection and caller ranking in 19 of 31 languages. The explorer did better on questions needing full source context (16 of 31) and exhaustive call-site search (10 of 31), and macro-heavy C was the weakest graph result at 0.58 against 1.00, because macros do not appear in a tree-sitter syntax tree. The authors list the limits plainly: one model, one repository per language, answers graded by the first author against reference answers, and no comparison yet against embedding-based retrieval or plain language servers. Their own recommendation is a hybrid: graph for structure, files for source.
Two README numbers need care. The "120x fewer tokens" and "99.2 percent reduction" figures come from five structural queries (about 3,400 tokens against about 412,000), not from the 31-repository study, which found about 10 times. And the README is internally inconsistent on counts (158 or 162 languages, 15 or 17 tools, 43 or 45 client surfaces), which is normal for a fast-moving project and a reason to quote a dated source.
Gotchas before you install it
- Static structure only. Reflection, dependency injection containers, dynamic dispatch and runtime-generated routes are not in the graph. On a Spring or Rails codebase, "nothing calls this" can mean "the framework calls this".
- The team artifact can bloat git history. You can commit a compressed graph snapshot so teammates skip reindexing. The README warns that committing every refresh turns a roughly 20 MB file into gigabytes of history, and cites one team that reached about 6 GB across about 350 commits of that single path. Commit on a cadence, or use Git LFS as documented.
- Every process must run the exact same build. One daemon coordinates all agent sessions, and a mismatched version is refused at startup. After an update, restart every open agent session.
- The watcher follows git projects by default. An index of a non-git directory stays frozen until you reindex.
- It is a binary with access to your source. The project publishes checksums, Sigstore signatures, SLSA level 3 provenance and VirusTotal scans per release, and
gh attestation verifychecks the provenance. That is a strong baseline, and the paper argues the MCP ecosystem should require it. Your customer's security team still has to approve it.
jevgrep: ask what the code does
What it is. A Node.js CLI (npm install -g @dzhng/jevgrep, Node 22 or newer) with an agent skill that teaches Claude Code, Codex, OpenCode and similar agents to call it. You ask a behavioral question, such as jg "Where is authentication checked before a request reaches a handler?" ., and it prints a summary, a compact file list, verbatim source excerpts with line references, then declaration and call locations. The README is explicit that this is evidence for the agent, not a generated answer and not a guarantee that every relevant file was found.
How it works. Its architecture document describes a hierarchical walk. Directory metadata and content previews decide where to explore, so it does not upload the whole tree first. At each level, Jev (the decision model we covered in What is Jev) judges relevance. Files are admitted above a 0.5 threshold, with no fixed top-N limit. Selected files are parsed into declarations: tree-sitter for Python, Go and Rust, the TypeScript compiler for TypeScript and JavaScript, bounded text chunks for everything else. Source is never executed, and when a search budget runs out the output says the discovery was partial instead of pretending to be complete.
Reading the benchmark closely
The jevgrep authors publish per-task costs, which lets you check the headline. In the total-cost rerun of version 0.4.3 on 28 September 2026, ten Python SWE-bench tasks were each run once with the coding agent at medium effort, against saved baseline runs without jevgrep:
Both runs solved the same 8 of 10 tasks. Total cost fell from $7.62 to $5.66, of which $1.48 was Jev itself, so 25.8 percent lower including the retrieval model. (The "~30 percent" banner is an earlier figure that excluded Jev's cost; the README says so.) Now break it down. The pytest task went from $2.34 to $0.53, a saving of $1.82 out of a total saving of $1.97. That one task is about 92 percent of the effect. Across the other nine tasks, total cost was $5.13 against $5.28, about 2.8 percent lower, and five of the ten tasks cost more with jevgrep than without it.
None of that is hidden. The results file says the tasks are tuned development tasks rather than a holdout, one sample per task does not establish a causal effect, and the aggregate saving is not a per-task guarantee. A later 0.5.0 evaluation cut Jev's own cost by about 59 percent against the 0.4.3 run, yet combined cost came out 2 to 3 percent higher than that run. The honest summary: the tool matched solve rate with lower cost on a small set, mostly by making one expensive task cheap. That is worth a pilot on your own repository, not a budget forecast.
Gotchas before you install it
- Your source goes to a model provider. Searches send eligible file content to Jev through Vercel AI Gateway, TypeSafe, OpenRouter, OpenCode Zen or a compatible endpoint. Default filters skip ignored, hidden, dependency, binary and obvious credential files, and the README says plainly that these filters do not guarantee sensitive data is removed. Run
jg filesfirst: it counts what a search would read, with no network call. Use--excludefor anything else. - The CLI alone does nothing for your agent. Run
jg skillto install the skill, and rerun it after upgrades, because npm updates do not overwrite skill files in your projects. - Declaration parsing covers four language families. Python, Go, Rust and TypeScript or JavaScript get declaration-level excerpts. Other languages fall back to text chunks.
- For exact symbols, grep is still faster. The README says so: when you know the name or the path, a direct read or
rgis all you need.
Which to use, and when
| Situation | Reach for | Why |
|---|---|---|
| Source may not leave the machine (regulated, air-gapped) | Graph tool plus grep | Everything runs locally; no model provider sees code |
| "Who calls this?", "what breaks if I change it?" | Graph tool | Pre-materialised call edges, about a tenth of the tokens |
| "How does retry work here?" in an unfamiliar repo | Model-judged retrieval | Behavioral questions span files whose names do not match the question |
| You know the symbol or string | grep or rg | Exact, free, already there |
| Macro-heavy C, heavy reflection or DI | Grep and read | Static graphs miss what the syntax tree does not show |
The FDE lens: code search on a customer's codebase
A forward deployed engineer usually works in code they did not write, under rules they did not set, which makes the first constraint legal rather than technical. Before choosing a tool, ask whether source can be sent to a third-party model endpoint at all. For many banks, insurers and public-sector customers the answer is no, which removes model-judged retrieval from the table and makes a local index plus disciplined grep the default.
Then measure on the customer's own repository, because both studies above use public repositories that may not resemble a 15-year-old monolith. Take ten questions the team actually asks in its first month (where is billing calculated, what calls the legacy auth module, which services read this table) and write down the right answer with the team's help. Run the agent with each tool configuration, and record correctness, tool calls, tokens and cost. That is the same protocol both projects used, scaled down, and it gives you a number your customer believes because it came from their code. codebase-memory-mcp's documentation includes a guide to measuring quality and savings on your own workload for exactly this.
Finally, tell the agent which tool is for what. A short line in the repository's AGENTS.md ("use trace_path for callers, search_code for literal strings, jg for how-does-this-work questions") prevents the agent from grepping for a call chain the graph already holds. That is harness engineering in miniature; our post on harness engineering for AI coding agents covers the rest.
In interviews this shows up as retrieval design: when to use structural retrieval versus embeddings, how to evaluate a code-search tool before rolling it out, and what data leaves the customer's network. The Model Context Protocol and GraphRAG and contextual retrieval concepts cover the foundations, and the RAG and agents bank has the design questions to practise.

Turn it into offers. Work the real questions and concepts this maps to:
FAQ
An open-source MCP server, written in C and MIT licensed, that parses a repository with tree-sitter into a persistent knowledge graph of files, functions, classes, calls, routes and cross-service links, stored in SQLite on your machine. A coding agent queries it with tools such as search_graph, trace_path, detect_changes and a read-only Cypher subset instead of grepping and reading files. It contains no language model; the agent is the query translator.
Discussion (5)
The paper's own conclusion is the one to repeat to a customer: hybrid. Graph for who-calls-what and blast radius, files for the actual lines. Every tool that pitches itself as replacing grep ends up as one more tool next to it.
Yes, and the agent needs to know which to reach for. That is an instructions problem as much as a tooling one: a line in AGENTS.md that says use trace_path for callers and search_code for literal strings saves more tokens than either tool on its own.
Reader question: is it safe to install an MCP server binary from GitHub on a work laptop?
Ask your security team, and bring evidence. This one publishes checksums, Sigstore signatures and SLSA level 3 provenance you can verify with gh attestation verify before running it. That is better than most. It still gets read access to your source and write access to your agent configuration files, which is exactly what a security review should look at.
The pytest number is the lesson. One task saved 1.82 dollars out of 1.97 total. If you only read the headline you would expect every task to get cheaper, and half of them did not. Ten tasks is a pilot, not a verdict, and the jevgrep authors say so in their own results file, which is to their credit.
