24 — Solomon: Retrieval & Agent Tier (Go)
Solomon 6 of 6 · Solomon/retrieval, Solomon/chains, Solomon/ensemble, Solomon/genengine
Stack: Go (4,757 lines), Kafka, vector index, HTTP fetchers, RSS, SQL migrations
Status: Deployed as phase2-retrieval in the Helm chart
Where this sits. The bridge between the ML work and the distributed-systems work. It is a Go service — the target language — doing ingestion, indexing, search, and a tool-calling agent loop with a grounding guard.
Structure
Solomon/retrieval/
index/ chunk.go embed.go search.go store.go service.go
agent/ loop.go tools.go parser.go calculator.go
grounding.go grounding_guard_test.go trace.go service.go
fetcher/ extract/ rss/ ingestion
kafkax/ kafkax.go event transport
llm/ searchclient/ model + search clients
authx/ config/ migrations/ cmd/ eval/
Plus chains/ (step/spec composition with a Go client), ensemble/
(orchestrator, members, serving, training, eval), and genengine/
(grammar-constrained generation, samplers, processors, context).
The grounding guard — the part worth leading with
agent/grounding.go has a dedicated grounding_guard_test.go. A retrieval
agent's characteristic failure is answering from the model's parameters while
appearing to answer from the retrieved documents. A grounding guard is the
check that the answer is actually supported by what was retrieved — and having
it under test means the failure mode was anticipated rather than discovered in
production.
"What stops your RAG system from confidently making things up?" is now a standard interview question. Most answers are hand-waving. This one is a file with a test next to it.
The agent loop, decomposed
| File | Responsibility |
|---|---|
loop.go | The iteration: think → call tool → observe → repeat, with a bound |
tools.go | Tool registry and dispatch |
parser.go | Parsing model output into tool calls — with its own test |
calculator.go | A deterministic tool, so the loop can be tested without a model |
trace.go | Execution tracing — observability for a non-deterministic system |
grounding.go | The answer-support check |
parser_test.go and loop_test.go exist because both are the parts that break:
a model emits malformed tool syntax, and a loop without a bound runs forever.
The index tier
chunk.go (with chunk_test.go), embed.go, store.go, search.go.
Chunking is tested because chunk-boundary bugs are silent — they do not error,
they just quietly degrade retrieval quality forever.
Why kafkax matters here
The retrieval tier consumes from Kafka. That makes this the point where the ML platform and the event-driven architecture meet: ingestion is not a cron job pulling feeds, it is a consumer off a log, with the same delivery-semantics questions as everything else in this portfolio.
Interview surface this opens
- RAG architecture: chunking strategy, embedding, retrieval, reranking
- Grounding/faithfulness checks and how you test a non-deterministic system
- Bounding an agent loop, and what you do when the bound is hit
- Deterministic tools as a testing seam for LLM systems
- Streaming ingestion off Kafka vs scheduled crawling