Performance
Measured numbers, benchmarks and what affects speed.
Measured, on a real 2,022-file open-source repository (plausible/analytics): cold ingestion took 34 seconds, and answering real questions from the compiled ledger cost 67–93% fewer tokens than grep-based search over the source, counted with a standard tokenizer. The one case where grep won is included in the published methodology.
What affects speed
| Factor | Effect |
|---|---|
| Observe scope | The biggest lever — exclude vendored and generated directories |
| Unchanged tree | build reuses cached artifacts by fingerprint |
recover --parallel |
runs independent passes concurrently |
| LLM stages | dominate wall-clock when enabled; cached in .ekos/llm-cache/ |
| Embeddings | one call per object; cached in .ekos/embed-cache/ |
| Storage layout | partition or distribute for very large ledgers |
Point lookups (RFC 0168)
Each fact-engine run carries a per-run Bloom filter of the entity ids it contains. A point lookup (ekos_state, ekos_neighborhood and the other by-id reads) skips any run whose filter rules the id out, and decodes only the block that holds the id instead of the whole run. Runs are compressed with zstd at level 9.
| Measurement | Before | After |
|---|---|---|
| Point lookup across an 8-run index (benchmark) | 1,083 µs | 41 µs |
| Point lookup on a real ledger | 1.43 ms | 0.171 ms |
| Rewriting the index for a real ledger | 150 s | 4.5 s |
Benchmarks
Criterion benchmarks live in benchmark/ (a separate Cargo workspace): cd benchmark && cargo bench (or cargo bench --bench ledger_write). CI runs them on every push to main.