Building a Knowledge Layer for AI Agents: Embeddings, Signposts, and What Actually Worked
The problem: AI agents with amnesia
Every time an AI coding agent starts a new session, it knows nothing about what happened before. It doesn’t know that the settlement calculation was refactored last Tuesday. It doesn’t know that you tried a caching approach three weeks ago and abandoned it because of a race condition. It doesn’t know that the notification system requires both a database commit AND a WebSocket broadcast — a lesson learned the hard way across multiple sessions.
This isn’t a minor inconvenience. In a codebase with 1,700+ source files, 330+ planning documents, and 870+ raw agent session transcripts, the odds of an agent accidentally revisiting solved problems or missing critical context are high. We watched it happen repeatedly: agents proposing architectures that had already been rejected, missing gotchas that previous sessions had documented, duplicating work that was already half-done elsewhere.
The standard answer is RAG — Retrieval Augmented Generation. Chunk your documents, embed them, do a similarity search, stuff the results into the prompt. We tried that. It’s not enough. Here’s what we actually built, what failed along the way, and what we learned.
The inspiration: Karpathy’s “embedding into the tool layer”
Andrej Karpathy has talked about the idea of embedding capabilities directly into AI-accessible interfaces rather than building monolithic agents with hardcoded knowledge. The concept resonated with something we’d been thinking about: instead of trying to pre-load an agent with everything it might need, give it a fast way to search for and retrieve exactly what’s relevant.
But “semantic search over your docs” is table-stakes RAG. The real question was: what does the agent actually need when it’s trying to decide whether to read a document?
What we tried first: LangExtract + local LLM
Our first approach was structured extraction. We used LangExtract — a library for grounded information extraction — combined with Gemini Flash and a local LLM (Gemma via LM Studio) as a fallback.
The idea: process every document chunk through an LLM to extract typed knowledge — architectural decisions, rejected approaches, bug fixes, configuration changes, schema decisions. Store each extraction with source text, character offsets, confidence scores, and alignment quality. Ten extraction types total, each with its own schema.
It worked. Kind of. Gemini 3 Flash produced rich extractions with 0% parse errors. The local Gemma model worked offline but had a 40% parse failure rate due to thinking blocks leaking into the output. We processed 1,619 session doc chunks this way.
The problems showed up at scale:
- Cost of precision we didn’t need. LangExtract verifies that extracted text actually appears in the source document (grounding). That’s valuable for citation-quality work, but we just needed signposts — pointers to help an agent decide whether to read a file.
- Garbage data accumulation. Chunks with no extractable knowledge produced null rows. Over time we accumulated 5,319 null extractions and 3,145 duplicates. The extraction pipeline was creating data that degraded search quality.
- LM Studio blocked the pipeline. When we tried to process raw session transcripts (877 of them, some over a million characters), the local LLM summary step made per-session blocking calls. The pipeline stalled for hours before getting to the actual embedding work.
- Diminishing returns. The extraction types were too granular for how agents actually use context. An agent deciding whether to read a planning document doesn’t need to know there’s a
configuration_changeextraction at character offset 4,832. It needs to know: what was this session about, what decisions were made, and what files were touched.
We kept all the extracted data — 97,000+ rows in aa_knowledge_extractions — but moved on.
What we tried second: direct Gemini JSON
We stripped out LangExtract entirely and switched to direct Gemini Flash 2.5 calls with response_mime_type="application/json". No parsing library, no grounding verification, no extraction framework. Just a system prompt describing the JSON schema we wanted and a document as input.
This was dramatically simpler. A single function call replaces the entire LangExtract + provider abstraction + alignment verification stack:
model = genai.GenerativeModel(
model_name="gemini-2.5-flash",
system_instruction=SIGNPOST_PROMPT,
generation_config=genai.GenerationConfig(
temperature=0.1,
max_output_tokens=8192,
response_mime_type="application/json",
),
)
response = model.generate_content(content)
signpost = json.loads(response.text)
Gemini’s JSON mode guarantees valid JSON output. Parse errors dropped to near-zero. The 2% that did fail (truncated output on very long documents) were caught by a fallback to Flash Lite with higher token limits.
Cost: Processing the entire 332-document session docs corpus costs about $0.35. The 13 MCP reference docs cost $0.10. Source code indexing is even cheaper because the prompt is smaller. For a nightly incremental run where only a handful of documents changed, cost is negligible.
The architecture we landed on: two layers
Here’s what we actually use now. It’s two complementary systems that serve different purposes:
Layer 1: Voyage embeddings for semantic search
Every knowledge source — session docs, raw agent transcripts, MCP reference docs, developer guides, and the full source code tree — gets chunked and embedded using Voyage AI (voyage-3, 1024 dimensions).
Chunking splits documents on natural boundaries (headings, paragraph breaks). Each chunk gets a content hash for incremental processing — on subsequent runs, only chunks where the content actually changed get re-embedded. This is done via SQL-first filtering:
SELECT kc.id, kc.source_ref, kc.embed_text, kc.content_hash
FROM aa_knowledge_chunks kc
LEFT JOIN aa_knowledge_embeddings ke
ON ke.chunk_id = kc.id AND ke.source_table = 'aa_knowledge_chunks'
WHERE kc.source_key = :source_key
AND (ke.id IS NULL OR ke.content_hash != kc.content_hash)
Only chunks with no embedding or a changed hash get sent to Voyage. After the initial corpus processing (40,000+ chunks, 53,000+ embeddings), nightly runs process single-digit chunks in seconds.
Embeddings are written in streaming batches of 500 — if the process crashes mid-run, you lose at most one batch worth of work, and the SQL filter means the next run picks up exactly where it left off.
What this layer does: Given a query like “notification center integration,” it finds the 10-50 most semantically similar chunks across all sources. Fast, cheap, and it narrows 40,000 chunks to a handful of candidates.
Layer 2: Gemini signpost metadata for decision-making
This is where it gets interesting. Voyage tells you which documents are relevant. But an agent still needs to decide: should I actually read this document? Is it worth spending 5,000 tokens of context to open it?
For every document, Gemini Flash generates a structured signpost:
{
"rich_summary": "Refactored notification center to support both DB persistence and WebSocket broadcast...",
"work_types": ["backend", "architecture", "websocket"],
"key_decisions": [
{
"decision": "Require both session.commit() AND WebSocket broadcast for notifications",
"rationale": "DB alone doesn't update the UI; WS alone loses data on page refresh"
},
{
"decision": "Component-scoped alert system separate from user notifications",
"rationale": "Different lifecycle — alerts persist until withdrawn, notifications are fire-and-forget"
}
],
"artifacts": ["backend/component_alerts/sdk.py", "backend/component_alerts/router.py"],
"components": ["NotificationCenter", "ComponentAlerts"],
"tags": ["notifications", "websocket", "alerts"]
}
Key decisions include rationale — not just what was decided, but why. This is the highest-value signal for an agent: knowing that a previous session rejected approach X because of reason Y prevents the agent from re-proposing it.
Work types are open-ended rather than a fixed enum. The model decides what labels fit: planning, bug fix, frontend, backend, architecture, database, marketplace, task manager, and whatever else describes the work.
For source code files, we use a lighter-weight format — just purpose, key functions, tables touched, and API endpoints exposed. A 200-line router file doesn’t need the full signpost treatment.
How the two layers work together
When an agent invokes sybre://recall/{query}:
- Voyage narrows the field. Semantic search across 40K+ chunks returns the top candidate documents, ranked by similarity.
- Signpost metadata provides the summary. For each candidate, the agent sees: title, one-paragraph summary, work types, key decisions with rationale, and file paths.
- Source code files get their own section. Below the signpost list, the recall includes the top 10 relevant source files with their purpose descriptions — a separate signal from “what work was done” (session docs) vs. “what code exists” (source files).
- The agent decides what to read. No chunk text leaks into the response. The agent uses the signposts to decide which 2-3 documents are actually worth opening in full.
This is the critical insight: the agent itself is the second semantic engine. The first pass (embeddings) is cheap and mechanical. The second pass (deciding what’s relevant) requires the kind of nuanced judgment that LLMs are good at — and it’s already happening in the agent’s reasoning loop for free.
We don’t need to pre-compute “is this relevant to task X?” for every possible task. We give the agent enough information to make that call itself.
The output budget problem
Our first working recall returned 93KB of rich signpost data for 5 documents. That’s roughly 24,000 tokens — the signposts were so detailed they cost more context than just reading the source files would have.
We went through several iterations to find the right balance:
- Uncapped output (v1): Every key decision with full rationale, full rich summaries, all artifacts. 93KB for 5 results. Unusable.
- Hard truncation (v2): 600-character budget per record. Too lean — agents couldn’t tell whether a document was worth reading.
- Balanced budget (v3): 1,000-character budget per record with 300-char summaries, top 3 decisions at 140 chars each, and compact artifact lists. 20 results default. This landed at roughly 5,000-6,000 tokens total — enough for an agent to make informed decisions.
The key realization: recall is an investment, not a cost. 5,000 tokens of quality signposts prevents 30,000-50,000 tokens of blind grep-and-read. An agent that spends 5,000 tokens on recall and then reads 2-3 targeted files (10,000 tokens) uses ~15,000 tokens total. Without recall, the same agent would spend 25,000-50,000 tokens rummaging through files, often reading the wrong ones first.
It’s better to err on the side of slightly more per record — give the agent enough context to make a proper decision — than to be so compact that the agent has to open files just to figure out if they’re relevant.
Search strategy matters
Compound queries like “celery restart bid pipeline” force the embedding model to find chunks matching all concepts simultaneously. This biases results toward the handful of documents that happen to mention everything together, potentially missing highly relevant documents that focus on just one aspect.
We found that decomposing broad topics into focused queries produces much better coverage. “Celery restart” surfaces worker resilience and graceful shutdown patterns. “SP6 bid pipeline” surfaces pipeline architecture and stage orchestration. The union of both result sets gives the agent a more complete picture than either query alone.
This is agent guidance, not system logic — the agent knows its own intent. Sometimes a compound query IS what you want (when you need the intersection). But the default instinct should be: multiple focused queries for broad topics.
The diversity problem
Early testing revealed that the top 50 Voyage results for a specific query might come from only 5 unique documents. A single large session transcript with 90+ chunks could dominate the result set, pushing out other relevant documents entirely.
The fix was increasing the search window aggressively (fetching up to 500 chunks instead of 50) and deduplicating to the document level. Large sessions still rank high because their best-matching chunk scores well, but they don’t crowd out other documents by sheer volume of chunks.
The numbers
As of today, here’s what the system processes:
| Source | Documents | Chunks | Embeddings | Signposts |
|---|---|---|---|---|
| Session docs | 332 | ~3,200 | 3,200 | 332 |
| Raw agent sessions | 877 | ~24,400 | 24,400 | 877 |
| MCP reference docs | 13 | ~250 | 250 | 13 |
| Developer guides | 3 | ~50 | 50 | 3 |
| Source code | 1,746 | ~13,000 | 13,000 | 1,729 indexed + 17 skipped |
| Total | ~2,970 | ~40,000 | 53,000+ | ~2,950 |
Costs:
- Voyage embedding (initial full corpus): ~$2-3 total
- Voyage embedding (nightly incremental): <$0.01
- Gemini signpost generation (332 session docs): ~$0.35
- Gemini signpost generation (13 MCP docs): ~$0.10
- Gemini source code indexing (1,746 files): ~$0.50
- Total ongoing nightly cost: effectively zero (only changed content gets reprocessed)
Processing time (16 parallel workers, AdaptiveWorkerPool):
- Voyage embedding (40K chunks, initial run): ~35 minutes with streaming batches
- Gemini signposting (332 session docs): ~15 minutes
- Gemini signposting (877 raw sessions): ~38 minutes (some sessions split into 4-7 segments)
- Gemini code indexing (1,746 files): ~10 minutes
- Nightly incremental run: seconds to low single-digit minutes
What we learned along the way
Local LLMs aren’t ready for structured extraction at scale. Gemma via LM Studio works for single-shot tasks, but at 40% parse failure rate and 17 seconds per chunk, it’s not viable for processing thousands of documents. We kept it as infrastructure for future use, but removed it from the nightly pipeline.
Extraction granularity should match consumption granularity. We started with 10 extraction types and character-level offsets. Agents don’t consume context at that resolution. They consume it at the document level: “should I read this file?” Doc-level signposts with 3-5 key decisions turned out to be the right unit.
SQL-first filtering is non-negotiable at scale. Our first embedding implementation loaded all 24,000 chunks into Python memory, computed hashes, and filtered in a loop. It took 33 minutes and used 400MB of RAM before writing a single embedding. The SQL-first approach (LEFT JOIN where embedding missing or hash changed) reduces that to loading only the handful of chunks that actually need work.
Streaming batches matter for crash resilience. Writing 500 embeddings at a time means a crash loses 500 embeddings, not 24,000. Combined with the hash-based deduplication, recovery is automatic — just restart the job.
Push your API rate limits. Our initial implementation used 4 parallel Gemini workers. Gemini Flash allows 2,000 requests per minute on pay-as-you-go. We bumped to 16 workers with exponential backoff on 429 errors. The adaptive worker pool monitors throughput and CPU/memory, scaling workers up or down automatically. Processing time dropped by 4x.
Don’t delete what you’re replacing. We kept all 97,000+ LangExtract extractions, the full extraction service code, and the LM Studio integration. The new system runs alongside the old one. If signposts turn out to be insufficient for some future use case, the extraction infrastructure is still there. The cost of keeping it is a few hundred KB of code; the cost of rebuilding it would be days.
Exclude trivial files from the entire pipeline. Our initial source code scan included every Python file — including 17 __init__.py files that contained nothing but a one-line docstring. These made it through chunking, embedding, and indexing, then failed the Gemini call every run because there wasn’t enough content to summarize. They’d have been retried every night forever. We added a 100-character content threshold at the scan level and wrote “skipped” markers in the index table so they’re excluded permanently.
Separate code results from document results. Source code files and session documents serve fundamentally different purposes in recall. A session doc tells you what decisions were made and why. A code file tells you what exists and what it does. Mixing them in the same ranked list means one type always dominates. We split them into separate sections — session signposts ranked by relevance, then a dedicated “Relevant Source Files” list below.
What we don’t know yet
This system has been running for less than a day. The architecture feels right — the recall output looks genuinely useful in manual testing — but we haven’t validated the most important question: does agent performance actually improve in practice?
Specifically:
- Do agents make fewer redundant decisions when signpost recall is available?
- Is the signpost metadata rich enough to prevent agents from opening irrelevant documents?
- Does source code indexing (purpose + key functions for 1,746 files) help agents navigate the codebase faster, or is grep sufficient?
- Does the decomposed query strategy actually produce better results than single compound queries in real agent workflows?
- What’s the right balance between recall breadth (more results) and agent attention (more results = more to evaluate)?
We’ll report back after running this in production for a few weeks. The evaluation methodology will need to be qualitative more than quantitative — watching real agent sessions for patterns of “the agent knew about X because it found it via recall” versus “the agent reinvented Y because recall didn’t surface it.”
Whether the signpost architecture proves effective or not, the journey to get here — through LangExtract, local LLMs, extraction schemas, garbage data cleanup, output budget tuning, diversity fixes, and multiple architectural pivots — reflects the reality of building AI-for-AI systems in 2026. The tooling is powerful but the design space is wide open, and the right architecture often only becomes obvious after trying the wrong ones first.
The takeaway
If you’re building a knowledge layer for AI coding agents, here’s the shortest version of what we learned:
- Embed everything. Voyage (or equivalent) across all your knowledge sources. Hash-based incremental processing makes ongoing costs negligible.
- Don’t return chunk text. Return document-level summaries with enough metadata for the agent to decide what’s worth reading. The agent is the second semantic engine.
- Key decisions with rationale are the highest-value signal. An agent knowing “we tried X and rejected it because Y” prevents the most expensive category of wasted work.
- Direct LLM JSON beats extraction frameworks. For signpost-quality metadata, a well-crafted prompt with JSON mode is simpler, cheaper, and more reliable than a full extraction pipeline.
- Budget per record, not per response. The temptation is to cap total output size. The better approach is capping each record at a consistent budget (~1,000 characters) and letting the number of records be generous (20 default). This gives agents enough per-record context to make good decisions while keeping total cost predictable.
- Teach agents to decompose queries. Without explicit guidance, agents will dump their entire task description into one search. Multiple focused queries produce better coverage than one compound query. This is a documentation problem, not a code problem — but if you don’t document it, no agent will ever do it.
- Separate code from context. Source files and planning documents answer different questions. Don’t rank them in the same list.
- Start with more infrastructure than you need. We built Voyage embeddings, LangExtract, LM Studio integration, and Gemini extraction before we knew which combination would work. The ones we stopped using weren’t wasted — they were the experiments that pointed us toward the architecture that did work.
The code is in production. We’ll see if the theory holds.