AI Agent Context Benchmark: Tokens Saved
How many tokens does an AI agent save by querying compiled knowledge instead of raw source? Real repository, standard tokenizer, every raw output published.
How much context does an AI agent save by asking EKOS instead of reading the source? We measured it on a real repository and published every command and raw output, including the case where EKOS loses.
The setup
- Repository: plausible/analytics, 2,022 files of Elixir, JavaScript and SQL, unmodified.
- Questions: two real ones. "What columns does the
imported_browserstable have?" is a single-object lookup. "What tables exist in the Postgres schema?" is an enumeration. - Raw source, at three tiers. Best case: the agent already knows the exact file and line. Realistic: it greps the whole repository with context, as coding agents do, which gives 32 matching files. Naive: it reads the right file whole.
- EKOS, counted including the "where is it?" step: an
ekos_searchcall plus anekos_statecall, or oneekos_eklquery for the enumeration. - Tokenizer:
tiktokencl100k_baseon both sides. Nocharacters / 4estimates.
Results
| Question | Raw source | Tokens | EKOS | Tokens | Result |
|---|---|---|---|---|---|
imported_browsers columns |
best-case grep (known line) | 122 | search + state | 1,186 | EKOS costs 9.7× more |
imported_browsers columns |
realistic grep (32 files) | 10,357 | search + state | 1,186 | 88.5% fewer |
imported_browsers columns |
naive (whole file) | 3,651 | search + state | 1,186 | 67.5% fewer |
| all Postgres tables | naive (2,738-line schema) | 18,995 | one EKL query | 1,258 | 93.4% fewer |
Cold ingestion: compiling the whole 2,022-file repository took 34 seconds.
What it means
If the agent already knows exactly where the answer is, nothing beats a direct grep, because there is nothing left to discover. That is rarely where an agent starts. In this repository 32 files mention imported_browsers and none of them is named after the schema. For realistic questions the compiled ledger is 67–93% cheaper, and its answer carries path:line evidence the agent can check.
Reproduce it
The full write-up, every raw command, every MCP response and the token-counting script are public: The First Benchmark Number. It needs only the same public repository and pip install tiktoken. Measured 2026-08-20; the dated record is devlog 62.
More measurements
- Business semantics on LedgerSMB (2026-10-05). From PostgreSQL, PL/pgSQL and Perl source, with no LLM, EKOS recovered 47 business concepts and 222 coded values (179 explained), resolving 458 of 649 predicate sites. The exported LinkML schema passes
linkml-lintwith 0 errors. The run is reproducible with one script: demo. - Answer quality.
ekos eval rungrades whole answers across 101 scenarios. See Performance.