━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Connected: Apertus (llamafile)
/help for commands · /exit to quit
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
25 session(s) unmined — /memory mine to extract learnings
RAG has 514 chunk(s) but is off — /rag on to enable context injection
harvey > /memory list
project_fact_019f37 project_fact - 0.5 safe_mode default allowed_commands now includes kb and man
workspace_profile_29074f workspace_profile - 0.5 Data Scientist — Laboratory
project_fact_84e77b project_fact - 0.5 Project: Laboratory
project_fact_4f8e21 project_fact pattern 1.0 harvey/INSTALL.md is hand-maintained, not cmt-generated; installer.sh/ps1 removed (pre-release, no binary distribution yet)
harvey > /memory mine
Extracting memories from /home/rsdoiel/Laboratory/agents/sessions/harvey-session-20260924-143905.spmd …
LLM proposed 1 memory candidate(s). Starting review…
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Proposed memory 1 of 1
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Type: tool_use
Kind: pitfall
Description: Never use the '+' character in filenames to avoid shell parameter expansion issues.
Action: Use '-' or '_' instead of '+' in filenames.
Tags: shell, filenames, parameters
Summary: The '+' character in filenames triggers parameter expansion, leading to incorrect file handling by shells. This is a permanent shell behavior.
[a]ccept [e]dit [s]how similar [r]eplace <id> [f]ull view [k]skip [q]uit
> a
Error saving memory: memory store: save: embed: ollama embed: HTTP 404: {"error":"model \"nomic-embed-text:latest\" not found, try pulling it first"}
Done. Accepted: 0 Skipped: 0
harvey > /memory list
project_fact_019f37 project_fact - 0.5 safe_mode default allowed_commands now includes kb and man
workspace_profile_29074f workspace_profile - 0.5 Data Scientist — Laboratory
project_fact_84e77b project_fact - 0.5 Project: Laboratory
project_fact_4f8e21 project_fact pattern 1.0 harvey/INSTALL.md is hand-maintained, not cmt-generated; installer.sh/ps1 removed (pre-release, no binary distribution yet)
harvey > /memory mine
Extracting memories from /home/rsdoiel/Laboratory/agents/sessions/harvey-session-20260713-171056.spmd …
LLM proposed 1 memory candidate(s). Starting review…
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Proposed memory 1 of 1
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Type: tool_use
Kind: pitfall
Description: Always use the chunk-size parameter when processing large files with Qwen3.5-4B-Q5_K_S.
Action: Increase the chunk-size parameter when processing large files.
Tags: model, processing, chunk-size
Summary: Files over 800 characters without chunk-size cause errors. This is a permanent model quirk.
[a]ccept [e]dit [s]how similar [r]eplace <id> [f]ull view [k]skip [q]uit
> a
Error saving memory: memory store: save: embed: ollama embed: HTTP 404: {"error":"model \"nomic-embed-text:latest\" not found, try pulling it first"}
Done. Accepted: 0 Skipped: 0
harvey > /memory mine
Extracting memories from /home/rsdoiel/Laboratory/agents/sessions/harvey-session-20260706-172458.spmd …
termlib bumped to the
v0.0.9+354195d pseudo-version (no
v0.0.10 tag exists; tagging it is optional, left to
RSDOIEL). Update 2026-09-27: RSDOIEL tagged and released termlib
v0.0.10; go.mod re-bumped from the pseudo-version to the
real tag,
go build/go vet/go test ./...
clean.
ThoroughProbeModel wired into /rag setup’s
embedder auto-pick (confirmEmbedder in
commands_rag.go): the keyword guess is confirmed or
corrected with one live /api/embed call before the store
commits to it. Not wired into every setOllamaModel
selection (rejected — extra request on every switch for a signal only
RAG setup uses).
--json (harvey.PrintJSONError,
harvey.ExtractJSONFlag in exitcode.go): both
binaries accept -json/--json anywhere on the
line and print a failing startup/flag-parsing error (or, for
harvey, a failed non-interactive session) as
{"error","class","code"} JSON on stderr, matching
kb’s shape.
Scripted harvey exit codes: a non-interactive session
(keepFirstFailure in terminal.go) now exits
with the class of its first failed slash command or chat turn instead of
always 0. plan_cmd.go’s two unclassified errors
reclassified (Unavailablef, Negativef) as the
first command family this covers.
Follow-up, not done here: most other command
handlers (/read, /read-dir, an unknown command
name, most of
commands_rag.go/commands_skill.go) still
print-and-swallow a failure or return an unclassified
fmt.Errorf, so a scripted session using them still exits 0
or 70. Reclassifying each is its own pass, named but not started.
Action Items
Status 2026-09-24 (end of day): both items above are
done and scheduled for v0.0.16, per
knowledge-learning-mode-design.md / -plan.md.
H0-H7 are done and committed:
/kb learn ingest|draft|review|concepts,
learn_model in harvey.yaml,
go.mod at knowledge fd588ef (unreleased
v0.0.13; RSDOIEL chose to ship on it). Harvey decisions/:
DR-0001 and DR-0002 accepted, DR-0003 (keep termlib) proposed. The
termlib gate is cleared and harvey/CLAUDE.md is fixed. What
remains before the tag is the release process itself (below), which is
RSDOIEL’s step.
Update next
to hold the release until kb ingest’s bugs and
[[wikilink]] tagging settled. Both shipped in
knowledge v0.0.4–v0.0.6 (wikilink-tagging,
concept-tag-retrieval, narrative-documents —
see ../knowledge/CHANGES.md), and go.mod had
no local replace for knowledge by the time
this was checked — it was already a plain tagged dependency. Bumped
go.modgithub.com/rsdoiel/knowledge v0.0.3 →
v0.0.6; go build/go test ./... clean. Also
rewired memory_unified.go’s recallKB to try
kb.MatchConceptNames/RecallByConceptNames
first (concept-tag match, not project-scoped, widened to
records and reviewed-status
document sections), falling back to the original
project-scoped substring scan over observations when no
concept matches — item 1–2 of
knowledge-learning-mode-feature-request.md. TDD: 6 new
tests in memory_unified_test.go, confirmed red before
implementation. Still open: items 3–6 of that
feature-request doc (the learning-mode UX itself, session-to-document
ingestion, decisions-directory realignment) — untouched by this
change.
observation update correction path, v0.0.9project rename/ concept rename + cross-machine
last-writer-wins + kb index --all — see
../knowledge/CHANGES.md). go build ./... and
go test ./... clean, no API changes touched harvey’s four
consumers (harvey.go, commands_kb.go,
memory_unified.go, terminal.go);
knowledge.go/ knowledge_merge.go no longer
exist in harvey (fully extracted into the module already).
go test -race ./... could not run on this machine (Pi:
“ThreadSanitizer: unsupported VMA range, Found 47 - Supported 48”) — a
platform limitation, not something this bump caused; unverified under
the race detector until run on hardware ThreadSanitizer supports. Items
3–6 of the feature-request doc remain open and untouched.
Update 2026-09-23: re-bumped v0.0.9 → v0.0.11
(v0.0.10concept suggest/document tag/density-linking/project rename
for record-owning projects; v0.0.11document fuzzy-tag/frontmatter, fuzzy
clustering, kb search punctuation fix).
go build, go vet, go test ./...
clean, no code changes needed; race detector still unrunnable on this
Pi. Items 3–6 are now designed: see
knowledge-learning-mode-design.md / -plan.md
and decisions/0001–0002 (all
proposed). Track B of that plan is gated on
knowledge v0.0.12
(../knowledge/library-lift-plan.md).
Bugs
Update 2026-08-08: Started the real per-model run
(Claude Code session, not manually at the terminal) — a standalone
throwaway Go program (chunkbench, built against this module
via a local go.modreplace, not checked in
anywhere) that calls ChunkDocument +
LlamafileBackend.Start/NewClient +
client.Chat directly per model, sequentially, timing 2
chunks each against natural_language_programming.md (12711
bytes, --chunk-size 800 → 23 total chunks, confirms the
same chunking as the 2026-07-05 run). Stopped by user request partway
through (3 of 7 models attempted) to free up the Pi; resume by rerunning
the remaining models below. Confirmed via ps that the
actual server invocation carries -ngl 0 -c 16384 as
configured — the GPULayers fix is genuinely in effect for this run,
unlike every prior timing attempt.
Results so far (context_length as configured in
harvey.yaml):
gemma-4-E4B-it-Q5_K_M (ctx 16384): chunk 1 = 4m40s,
chunk 2 = 4m36s (avg ~4m38s/chunk). Notably faster than the “~10
min/chunk” figure quoted above — that figure was a rough estimate from
the real 23-chunk run, not a tight back-to-back 2-chunk measurement;
treat ~4.5 min/chunk as the more reliable number for this model at this
chunk size.
gemma-4-E2B-it-Q5_K_M: no valid data —
both chunks errored with connection refused, and the
backend reported “server ready in 0s” (immediate, suspicious). Root
cause: a genuine race in the benchmark script, not a Harvey bug —
LlamafileBackend.Start adopts an already-listening server
instead of launching a fresh one (by design, for the case of an
externally-started server), and the script called Stop() on
the previous model then Start() on this one with no gap;
the prior process’s SIGINT hadn’t yet released port 8080, so Start
wrongly “adopted” the dying gemma-4-E4B server, which then actually
exited moments later. Needs a Detect()-poll-until-down guard
between Stop() and the next Start() before re-running E2B or
trusting any future back-to-back sequential benchmark script — worth
fixing in the script (not this package) since real interactive
/llamafile use switches are user-paced, not
back-to-back-instant.
Qwen3.5-4B-Q5_K_S (ctx 16384, the exact model/config
this TODO item’s “Update 2026-07-25” entry above wanted re-verified):
chunk 1 = 8m21s — nearly 2x gemma-4-E4B’s
per-chunk time even though both ran at the identical
context_length: 16384. This is a real, moderately
surprising data point: it means the earlier “large configured context
inflates CPU-only KV-cache setup cost” theory is not
the full explanation for Qwen’s slowness, since 16384 here is already
the smallest context of any model tested and it’s still the slowest —
something about this specific model/quant is just inherently heavier per
token on this CPU. Chunk 2 was in progress (interrupted by the stop
request) — re-run to get a second data point and confirm chunk 1 wasn’t
an outlier (e.g. one-time warmup cost).
Not yet run:Bonsai-8B-Q1_0 (ctx
65536), OpenELM-3B-Instruct-Q4_K_M (ctx 16384),
granite-4.1-8b-source-Q4_K_M (ctx 16384),
Apertus-8B-Instruct-2509 (ctx 49152) — all still queued, in
that order, in the chunkbench script’s model list.
Update 2026-08-08 (completed): Re-ran
chunkbench (v2, fixed: a waitForPortFree
poll-until-down guard between each model’s Stop() and the
next model’s Start(), closing the race that invalidated
gemma-4-E2B’s first attempt) for the 6 remaining/retry
models. All completed cleanly; full table below (2 chunks each,
--chunk-size 800, CPU-only -ngl 0, confirmed
via ps on every model this time):
model
context
avg/chunk
extrapolated, full 23-chunk doc
Apertus-8B-Instruct-2509
49152
1m51s
~42 min
granite-4.1-8b-source-Q4_K_M
16384
2m19s
~53 min
gemma-4-E2B-it-Q5_K_M
16384
2m24s
~55 min
Bonsai-8B-Q1_0
65536
3m36s
~83 min
gemma-4-E4B-it-Q5_K_M
16384
4m38s
~107 min
OpenELM-3B-Instruct-Q4_K_M
16384
6m13s
~143 min
Qwen3.5-4B-Q5_K_S
16384
7m51s
~180 min
Conclusion: parameter count does not predict per-chunk speed
on this CPU. The two fastest models (Apertus,
granite) are both 8B-class; the smallest model tested
(OpenELM, 3B) is the second-slowest, beaten only by
Qwen3.5-4B. Quantization scheme/architecture dominates raw
size for CPU-only inference here. Qwen3.5-4B-Q5_K_S is now
confirmed slowest across three independent chunk measurements (8m21s,
7m0s, 8m41s — consistently ~7-8.5 min/chunk), at the smallest
context length of any model tested, which rules out “large configured
context inflates KV-cache cost” as an explanation for its historical
slowness; something about this specific model/quant is inherently
heavier per token here. For an overnight/unattended full-document run on
this Pi 500, Apertus-8B-Instruct-2509 is the clear best fit
(~42 min vs. up to 3 hours for Qwen3.5-4B-Q5_K_S).
See agents/knowledge.db (Laboratory root), project
harvey, concept chunking, for the same summary
as a finding observation.
resolveDispatchTarget (dispatch_target.go)
now wires DebugLog onto whatever client it resolves, for
all three of its branches; (2) the now-redundant
dbg *DebugLog parameter was removed entirely from
RunChunkedAnalysis, along with all its internal logging
calls. See DECISIONS.md 2026-07-13 entry. Tests:
TestResolveDispatchTarget_RouteEndpoint_WiresDebugLog,
TestResolveDispatchTarget_LocalSwitch_WiresDebugLog.
Related finding, fixed separately 2026-07-13:builtin_tools.go’s read_file pre-read chunking
guard had the same cosmetic-only @mention bug already fixed
elsewhere for /read-chunks/injectOrChunk
during Direction D (Bug 1 in subagent-dispatch-design.md) —
it parsed @mention to relabel
ChunkAnalysisParams.Model but always dispatched via
a.Client, never the mentioned model. Missed during that
earlier work (only two of the three chunk-analysis call sites were found
at the time); now fixed the same way, via
resolveDispatchTarget. See DECISIONS.md 2026-07-13 entry.
Test: TestReadFile_MentionDispatchesToNamedModel.
Update 2026-07-25: confirmed via
gguf-dump against the GGUF embedded in
Qwen3.5-4B-Q5_K_S.llamafile (extracted with
unzip, since llamafile is a zip-appended APE binary, not a
raw GGUF) that the model’s own trained context length is
qwen35.context_length = 262144 — even larger than the
180224 this item suspected, and far above the 16384–65536 range of every
other registered model. This makes the KV-cache-allocation theory more
plausible, not less: llama.cpp/llamafile allocate KV cache proportional
to the configured -c value at server startup
(ActiveLlamafileContextLength() →
StartLlamafileService, see
backend_llamafile.go), independent of actual prompt/chunk
size, so a -c anywhere near this model’s native max would
explain slow, CPU-bound startup cost.
Also confirmed: agents/harvey.yaml’s registered entry
for Qwen3.5-4B-Q5_K_Salready reads
context_length: 16384 — the exact value this item
proposed testing. No record exists of when or why it was changed (no
matching git history in this repo — git log -p -S 180224
returns nothing), so this may already be applied but never
re-benchmarked. Remaining step: run
/read-chunks PATH --chunk-size 800 --max-chunks 2 against
the current (16384) config to confirm chunk time now falls in line with
the ~10 min/chunk baseline, then fold into the per-model timing table
above. Not run yet — this is a multi-minute, CPU-heavy operation and
deserves an explicit go-ahead rather than running unattended.
Resolved 2026-08-08, by the completed per-model
benchmark above: at context_length: 16384 — the exact
config this item wanted verified — Qwen3.5-4B-Q5_K_S
measured 8m21s, 7m0s, and 8m41s per chunk across three independent runs
(~7-8.5 min/chunk, consistent). This confirms the 16384 setting already
fixed the original 54+ minute outlier, but not the
KV-cache-allocation theory itself: 16384 is the smallest context of any
model benchmarked, yet Qwen3.5-4B-Q5_K_S is still the
single slowest model tested (nearly 2x the next-slowest,
OpenELM-3B). The remaining slowness is
model/quant-specific, not context-length-driven — no further action
planned here.