Business Semantics and LinkML
Recover business concepts, code meanings and constraints from SQL, dbt, Perl and Python traces, review them, and export a LinkML schema.
Business meaning that nobody wrote down still leaves traces in code: WHERE NOT p.obsolete, entity_class = 2, the lookup row (2, 'Customer'), a COMMENT ON COLUMN legend, and the commit that last changed the line. RFC 0170 recovers those traces into reviewable hypotheses. A person confirms them, and EKOS exports the confirmed ones as a LinkML schema.
The feature is experimental and opt-in. Everything it produces is a hypothesis until a person reviews it, and only a person can confirm it.
Turn it on
[semantics]
enabled = true
# defaults: rationale = true, min-sites = 2, max-enum-values = 12,
# llm-definitions = false, llm-max-definitions = 50, ontology unset
Then run the normal pipeline. Synthesis runs at the end of ekos commit, is deterministic, and makes no LLM calls unless you set llm-definitions = true.
ekos build && ekos recover && ekos resolve && ekos compile && ekos commit
What it reads
Recovery records every column-vs-literal predicate. It normalizes them, so status IN (1,3) and status = 1 OR 3 = status count as the same predicate. It also records key columns, constraints and lookup seed rows.
| Source | Traces |
|---|---|
SQL views, routines, standalone SELECTs |
WHERE / HAVING / JOIN ON / CASE predicates, also through CTEs and derived tables |
| DDL | CHECK, NOT NULL, primary keys, UNIQUE, COMMENT ON legends (A=asset), INSERT seed rows |
| PL/pgSQL | IF NEW.col … / WHILE / EXIT WHEN, resolved to the trigger's table |
| dbt | rendered model SQL (ref / source / var), column lineage across models, schema.yml tests (accepted_values gives a declared domain, relationships gives a lookup link) |
| Pentaho | FilterRows conditions on the upstream table |
| Perl | SQL inside string literals, use constant EC_CUSTOMER => 2 groups |
| Python | Enum classes |
| Docs and Confluence | marked glossaries: a heading or file named glossary, terms, definitions, dictionary or terminology |
| git | git blame -w on each evidence line, giving rationale links |
An application constant names a code only when it agrees with something EKOS already knows: a known label, or the column's name.
What it produces
| Kind | Meaning |
|---|---|
BusinessConcept |
a filter that recurs (min-sites), or one that defines a view |
EnumMeaning |
a coded value and its label, with the label's source |
ConstraintCandidate |
CHECK, NOT NULL, PK/unique, dbt test |
SemanticGap |
a code no source explains, an undocumented concept, an unmatched glossary term |
ConceptConflict |
two thresholds on one column, or two concepts with one name |
RationaleLink |
the commit behind an evidence line |
Every item carries path:line evidence.
Review: only a human confirms
ekos semantics list [--kind concept|enum|constraint|gap|conflict|rationale] [--status hypothesis|confirmed|rejected|needs_review]
ekos semantics show PartsNotObsolete # definition, evidence lines, links, commits
ekos semantics gaps # open questions, conflicts, the needs_review queue
ekos semantics confirm PartsAssembly PartsNotObsolete --as ann --note "checked" # all or nothing
ekos semantics edit PartsWithInventoryAccnoId --name "Inventory part" --description "…"
ekos semantics reject UserPreferenceWithoutUserId --note "anti-join plumbing"
A review holds only while the item's assertion and evidence stay unchanged. Line numbers are not part of that check, so moving code does not reset a review. If the assertion or evidence changes, the item becomes needs_review. If a confirmed item's traces disappear, it is flagged as stale and kept.
The web console's Semantics tab has a review queue with bulk confirm/reject, a gaps view and a LinkML viewer and editor. Decisions there need the console's write role. They run through the same human-only CLI and are attributed to the signed-in user.
Export and round-trip LinkML
ekos export linkml --out schema.yaml # default --status confirmed
ekos export linkml --status all --out draft.yaml # every element annotated with its status
ekos import linkml schema.yaml --dry-run # show the review decisions an edit stands for
ekos import linkml schema.yaml --as ann # apply them, all or nothing
The export defaults to confirmed items only, so a hypothesis never leaves EKOS looking like a fact by accident. Each element carries an ekos_id annotation, plus ekos_status, ekos_confidence and ekos_evidence. ekos import linkml matches elements by ekos_id and turns renames, descriptions, labels and ekos_status: confirmed|rejected into review decisions. On LedgerSMB the exported schema passes linkml-lint with 0 errors, and gen-json-schema and gen-pydantic both run on it. EKOS supplies the semantics; LinkML's own tooling generates everything else.
Ontology suggestions
ontology = "vocab.yaml" points at your own vocabulary:
prefixes: {schema: https://schema.org/}
terms:
- {id: schema:Invoice, label: Invoice, synonyms: [bill]}
EKOS suggests a mapping only on an exact word match: an exact suggestion for the label, a close one for a synonym. Suggestions are exported as ekos_suggested_*_mappings annotations, never as LinkML exact_mappings.
For agents (MCP)
When [semantics] is on, ekos_semantics_lookup returns what a term, table, column or code value means, and ekos_semantics_gaps returns what is not known. Every answer states its status. Rejected definitions are left out, and an empty lookup returns no_semantics_found so the agent knows not to guess. No MCP tool can confirm or reject an item.
Optional AI summaries
llm-definitions = true adds a summary of at most two sentences to each undocumented concept. Every sentence must cite the concept's evidence or the known code meanings. Uncited and hedged sentences are dropped. The summary is a reading aid and never becomes the definition. With a cloud provider, each concept costs one metered call.
Measuring quality
ekos semantics gold-template --out gold.yaml # tables and columns only, no EKOS output
ekos semantics eval --gold gold.yaml # concept recall, label accuracy, gap recall, evidence validity
A domain expert should fill in the template before seeing any of EKOS's output. eval flags any gold set whose meta does not say it was written blind.
On LedgerSMB (SQL, Perl and UI observed, no LLM), EKOS found 649 predicate sites. About 455 of them resolved to a table column, giving about 46 concepts and 219 coded values, 177 of them explained. Against a starter gold set, which is not expert-written, concept recall was 0.68 and label accuracy 0.95.
Limit: no trace, no recovery. EKOS cannot recover a meaning defined only by comparing two columns or comparing against the current date, such as overdue or unpaid.