kb ingest
— inline [[wikilink]] concept tagging — Design
Status (2026-09-13): Decisions confirmed. See wikilink-tagging-plan.md for the phased plan.
References: -
wikilink-tagging-feature-request.md — the filed idea this
design resolves. - narrative-documents-feature-request.md —
a second, independent consumer of the same
[[Name]] → concept mechanism, not built by this design, but
the reason the mechanism should live in knowledge rather
than being ingest-only from the start (see decision 3).
Motivation, and two corrections
Filed as wikilink-tagging-feature-request.md, inspired
by Build
a digitally sovereign second brain (Hinchliffe, Raspberry
Pi magazine — the article actually demonstrates Logseq + Syncthing,
not Obsidian as the feature request’s motivation section says; same
[[double-bracket]] convention either way, so the design
intent is unaffected). Two of the feature request’s open questions turn
out to have concrete answers once checked against the current code, not
just design choices:
“Does this ride on the observation that ingest already
creates per record?” No — checked cmd/kb/upsertAll
(cmd/kb/ingest.go:232-291) directly: kb ingest
never calls AddObservation. A record’s only existing path
to a concept today is linkInitiative
(cmd/kb/ingest.go:343-352), which links the record’s
project to a concept named after the initiative
field — one concept per record, scoped to the project, not to the
record. There is no existing record-to-concept relationship to reuse.
This settles the feature request’s schema question: a new join table is
required, not optional.
“Does [[Name]] conflict with the unused
Tags frontmatter field?” Checked
recordfile.go:76 and every reference to it:
Tags []string is parsed by ParseRecordFile and
carried through RecordFile/Record
(recordfile.go:219,368), but nothing in
cmd/kb/ingest.go ever reads it — it has been dead weight
since it was added. Rather than leave that question open, this design
resolves it: Tags and [[wikilink]] feed the
same mechanism (decision 3).
Decisions
New join table
record_concepts, added torecordsSchema(records.go), same shape asobservation_concepts/project_concepts:CREATE TABLE IF NOT EXISTS record_concepts ( record_id INTEGER REFERENCES records(id) ON DELETE CASCADE, concept_id INTEGER REFERENCES concepts(id) ON DELETE CASCADE, PRIMARY KEY (record_id, concept_id) );CREATE TABLE IF NOT EXISTSis idempotent on an existing database, same lazy-migration pattern the rest of the schema uses — noALTER, no backfill needed since this is a brand-new relationship with nothing to migrate.New method
LinkRecordConcept(recordID, conceptID int64) errorinknowledge.go, verbatim shape ofLinkObservationConcept/LinkProjectConcept—INSERT OR IGNORE, duplicate links a silent no-op.Two sources feed the same concept-linking call, not one:
[[Name]]tokens scanned from the record body only (not title, not other frontmatter), via a simple non-nested regex:\[\[([^\[\]]+)\]\].- Each entry in the frontmatter
Tagslist, resolved exactly the same way. This gives the deadTagsfield a purpose without new schema, rather than leaving “does this conflict with Tags” as an open question — they don’t conflict, they’re the same mechanism from two entry points (inline-while-writing vs. declared-up-front).
For each name found (from either source), trimmed of surrounding whitespace: resolve via the new
ResolveConceptName(decision 4), thening.kb.LinkRecordConcept(recordDBID, conceptID).Case-insensitive resolution, scoped to this path only — not a change to
AddConcept/kb concept add. Revised after review: a wikilink is embedded in ordinary prose, and English sentence position dictates capitalization independent of what the concept actually is —[[Computers]] are ... I use a [[computer]] regularlynames one concept twice, not two concepts. Folding this is correct for text scanned out of prose. It is not correct to widenkb concept add’s own semantics the same way: a human typingkb concept add Bugat the CLI is a single deliberate act, not an artifact of sentence grammar, and forcing case-insensitive uniqueness ontoconcepts.nameglobally means an actual schema migration (SQLite can’t widen aUNIQUEcolumn’s collation viaALTER TABLE; it requires a full table rebuild — new table, copy, drop, rename) for a problem this feature doesn’t have.So: a new method,
(kb *KnowledgeBase) ResolveConceptName(name string) (int64, error)—SELECT id FROM concepts WHERE name = ? COLLATE NOCASE LIMIT 1; on a hit, return the existing id (whatever casing it was originally stored with, however that concept was created — manually or via an earlier wikilink); on a miss,AddConcept(name, "")creates it verbatim, and that casing becomes canonical for every later mention regardless of case. Checked the liveagents/knowledge.db:SELECT LOWER(name)... HAVING COUNT(*) > 1returns nothing today, so there’s no pre-existing case-duplicate to worry about at the code level either — this is a clean, additive method with no migration.AddConcept/AddConceptWithIdentifierare untouched —kb concept addkeeps its exact-match behavior, no risk to existing callers or tests.Escaping is out of scope, per the feature request’s original decision — Markdown that legitimately needs literal
[[/]]is a known, unaddressed gap.New unexported
(ing *ingester) linkWikilinkTags(rf *knowledge.RecordFile, recordDBID int64), called unconditionally immediately aftering.linkInitiative(rf, projectID)inupsertAll— same call site, same unconditional-per-run shape (runs whether the record was added, updated, or skipped as unchanged), guarded only bying.dryRunandrecordDBID == 0(mirrors whylinkInitiativeguards onprojectID == 0). Because bothResolveConceptNameandLinkRecordConceptare idempotent, re-ingesting an unchanged file is a safe no-op rather than a duplicate-link error — no new idempotency mechanism needed.Portability parity is required, not optional.
record_conceptsmust be added everywhereobservation_concepts/project_conceptsalready are, or it becomes the one relationship in the schema that silently doesn’t survive a merge or export/import round-trip:knowledge_merge.go: a uuid-joinedINSERT OR IGNOREblock (mirroring the existingobservation_concepts/project_conceptsblocks atknowledge_merge.go:416-443), and an entry in theallTablessummary list (knowledge_merge.go:461-465) — the comment there is explicit that a table missing from that list is a table whose loss is never reported (DR-0013).jsonl.go: arecordConceptRecordJSON-L type, anexportRecordConceptsfunction, and an import branch — mirroringobservationConceptRecord/exportObservationConceptsexactly.
No
kb_ftschanges.record_conceptsis a pure join table; concepts are already indexed byAddConceptWithIdentifieritself. This feature touches none of the four-writer FTS mechanics.New read path:
kb record concepts ID, mirroringkb observation sources ID— lists the concepts linked to a record. A write-only feature with no way to see what it wrote is incomplete; this is the minimum visibility needed to verify the feature works and to debug a record’s tags later.Documentation.
kb-ingest.1.md/kb-record.1.mdare generated, not hand-maintained — checked the Makefile: each is produced by./bin/kb TOPIC -help >kb-TOPIC.1.md(thekb-topics-helptarget), and the actual source text isIngestHelpText/RecordHelpText, Go string constants incmd/kb/helptext.go(lines 563 and 633). So this decision is: edit those two constants (a paragraph on[[wikilink]]/Tagsresolution and the case-insensitivity behavior forIngestHelpText; the newconceptssubverb’s synopsis/description forRecordHelpText), then regenerate the.1.mdfiles by runningkb-topics-help— never hand-edit the.1.mdfiles themselves, they’d be overwritten.
Deferred, explicitly
- Scanning
kb observation addbodies for[[Name]]— the feature request scoped this tokb ingest(records) only for a first pass; nothing here changes that. - Widening
kb concept add/AddConceptto case-insensitive uniqueness globally (decision 4 scopes normalization to wikilink/Tags resolution only) — a real schema migration, not needed to solve the problem this design was asked to solve. - Escaping (decision 5) — known gap, not blocking.
- Everything in
narrative-documents-feature-request.md— that consumesResolveConceptNamethe same way but is a separate entity (documents, notrecords) and a separate design cycle.