EKOS
EKOS / AI Codebase Understanding for Large Repositories

AI Codebase Understanding for Large Repositories

Give AI coding agents a persistent, evidence-backed understanding of a large codebase: call graphs, schemas, lineage and impact analysis over MCP.

AI coding agents read code well and remember it badly. On a large repository they spend most of a session finding things: which of thirty files mentioning customer defines it, who calls this function, which views read this table. Then the session ends and the next one starts over.

EKOS gives an agent a persistent understanding of the codebase. It compiles the repository once into structured knowledge, updates it incrementally as the code changes, and serves it over MCP.

What the agent gets

  • Structure, not text. Modules, functions and their signatures, classes and inheritance, and a call graph (who calls what, from which line).
  • The database side too. Tables, columns, keys, views, stored procedures and triggers, plus which code reads and writes which table.
  • Lineage and impact. ekos_impact follows dependencies several hops out, so the agent can answer "what breaks if I rename this column?" before it renames it.
  • Evidence on every fact. path:line and the source fragment, so the agent can open the right file directly instead of searching for it.
  • History. Git authorship and co-change ("these files always change together"), and point-in-time queries over the knowledge itself.

Languages and sources

Rust, Python (including SQLAlchemy models and PySpark), JavaScript and TypeScript, Elixir and Perl. Six SQL dialects, plus PL/pgSQL, dbt and Pentaho. Git and GitHub history, and documentation. Compiled .NET and JVM binaries are supported through a separately licensed extension. See Indexing a Repository and Indexing a Database.

Built for large, old systems

EKOS was tested on LedgerSMB, a twenty-year-old accounting system in PostgreSQL, PL/pgSQL and Perl. The schema's 212 stored routines, its views and triggers were all recovered from source, linked to the tables they read and write, and documented from the schema's own comments. The LedgerSMB demo goes further and recovers business meaning: what entity_class = 2 means, which filters define an "open order". It exports the result as a LinkML schema.

Fewer tokens, measured

Answering real questions about a 2,022-file repository from the compiled ledger cost 67–93% fewer tokens than searching raw source. Compiling that repository from cold took 34 seconds. Methodology and the one case where grep wins: benchmarks.

Try it on your repository

curl -fsSL https://raw.githubusercontent.com/alexeyban/EKOS/main/install.sh | sh
cd /path/to/repo && ekos init --detect
ekos build && ekos recover && ekos resolve && ekos compile && ekos commit
ekos coverage          # which input kinds compiled, and which produced nothing

Then connect your agent.

Quickstart → Documentation → GitHub →