EKOSdocs
Docs / Guides Build

Business Semantics and LinkML

Recover business concepts, code meanings and constraints from SQL, dbt, Perl and Python traces, review them, and export a LinkML schema.

Business meaning that nobody wrote down still leaves traces in code: WHERE NOT p.obsolete, entity_class = 2, the lookup row (2, 'Customer'), a COMMENT ON COLUMN legend, and the commit that last changed the line. RFC 0170 recovers those traces into reviewable hypotheses. A person confirms them, and EKOS exports the confirmed ones as a LinkML schema.

The feature is experimental and opt-in. Everything it produces is a hypothesis until a person reviews it, and only a person can confirm it.

Turn it on

[semantics]
enabled = true
# defaults: rationale = true, min-sites = 2, max-enum-values = 12,
#           llm-definitions = false, llm-max-definitions = 50, ontology unset

Then run the normal pipeline. Synthesis runs at the end of ekos commit, is deterministic, and makes no LLM calls unless you set llm-definitions = true.

ekos build && ekos recover && ekos resolve && ekos compile && ekos commit

What it reads

Recovery records every column-vs-literal predicate. It normalizes them, so status IN (1,3) and status = 1 OR 3 = status count as the same predicate. It also records key columns, constraints and lookup seed rows.

Source Traces
SQL views, routines, standalone SELECTs WHERE / HAVING / JOIN ON / CASE predicates, also through CTEs and derived tables
DDL CHECK, NOT NULL, primary keys, UNIQUE, COMMENT ON legends (A=asset), INSERT seed rows
PL/pgSQL IF NEW.col … / WHILE / EXIT WHEN, resolved to the trigger's table
dbt rendered model SQL (ref / source / var), column lineage across models, schema.yml tests (accepted_values gives a declared domain, relationships gives a lookup link)
Pentaho FilterRows conditions on the upstream table
Perl SQL inside string literals, use constant EC_CUSTOMER => 2 groups
Python Enum classes
Docs and Confluence marked glossaries: a heading or file named glossary, terms, definitions, dictionary or terminology
git git blame -w on each evidence line, giving rationale links

An application constant names a code only when it agrees with something EKOS already knows: a known label, or the column's name.

What it produces

Kind Meaning
BusinessConcept a filter that recurs (min-sites), or one that defines a view
EnumMeaning a coded value and its label, with the label's source
ConstraintCandidate CHECK, NOT NULL, PK/unique, dbt test
SemanticGap a code no source explains, an undocumented concept, an unmatched glossary term
ConceptConflict two thresholds on one column, or two concepts with one name
RationaleLink the commit behind an evidence line

Every item carries path:line evidence.

Review: only a human confirms

ekos semantics list [--kind concept|enum|constraint|gap|conflict|rationale] [--status hypothesis|confirmed|rejected|needs_review]
ekos semantics show PartsNotObsolete            # definition, evidence lines, links, commits
ekos semantics gaps                             # open questions, conflicts, the needs_review queue
ekos semantics confirm PartsAssembly PartsNotObsolete --as ann --note "checked"   # all or nothing
ekos semantics edit PartsWithInventoryAccnoId --name "Inventory part" --description "…"
ekos semantics reject UserPreferenceWithoutUserId --note "anti-join plumbing"

A review holds only while the item's assertion and evidence stay unchanged. Line numbers are not part of that check, so moving code does not reset a review. If the assertion or evidence changes, the item becomes needs_review. If a confirmed item's traces disappear, it is flagged as stale and kept.

The web console's Semantics tab has a review queue with bulk confirm/reject, a gaps view and a LinkML viewer and editor. Decisions there need the console's write role. They run through the same human-only CLI and are attributed to the signed-in user.

Export and round-trip LinkML

ekos export linkml --out schema.yaml                  # default --status confirmed
ekos export linkml --status all --out draft.yaml      # every element annotated with its status
ekos import linkml schema.yaml --dry-run              # show the review decisions an edit stands for
ekos import linkml schema.yaml --as ann               # apply them, all or nothing

The export defaults to confirmed items only, so a hypothesis never leaves EKOS looking like a fact by accident. Each element carries an ekos_id annotation, plus ekos_status, ekos_confidence and ekos_evidence. ekos import linkml matches elements by ekos_id and turns renames, descriptions, labels and ekos_status: confirmed|rejected into review decisions. On LedgerSMB the exported schema passes linkml-lint with 0 errors, and gen-json-schema and gen-pydantic both run on it. EKOS supplies the semantics; LinkML's own tooling generates everything else.

Ontology suggestions

ontology = "vocab.yaml" points at your own vocabulary:

prefixes: {schema: https://schema.org/}
terms:
  - {id: schema:Invoice, label: Invoice, synonyms: [bill]}

EKOS suggests a mapping only on an exact word match: an exact suggestion for the label, a close one for a synonym. Suggestions are exported as ekos_suggested_*_mappings annotations, never as LinkML exact_mappings.

For agents (MCP)

When [semantics] is on, ekos_semantics_lookup returns what a term, table, column or code value means, and ekos_semantics_gaps returns what is not known. Every answer states its status. Rejected definitions are left out, and an empty lookup returns no_semantics_found so the agent knows not to guess. No MCP tool can confirm or reject an item.

Optional AI summaries

llm-definitions = true adds a summary of at most two sentences to each undocumented concept. Every sentence must cite the concept's evidence or the known code meanings. Uncited and hedged sentences are dropped. The summary is a reading aid and never becomes the definition. With a cloud provider, each concept costs one metered call.

Measuring quality

ekos semantics gold-template --out gold.yaml    # tables and columns only, no EKOS output
ekos semantics eval --gold gold.yaml            # concept recall, label accuracy, gap recall, evidence validity

A domain expert should fill in the template before seeing any of EKOS's output. eval flags any gold set whose meta does not say it was written blind.

On LedgerSMB (SQL, Perl and UI observed, no LLM), EKOS found 649 predicate sites. About 455 of them resolved to a table column, giving about 46 concepts and 219 coded values, 177 of them explained. Against a starter gold set, which is not expert-written, concept recall was 0.68 and label accuracy 0.95.

Limit: no trace, no recovery. EKOS cannot recover a meaning defined only by comparing two columns or comparing against the current date, such as overdue or unpaid.