The Problem With 'Just Fine-Tune It'
Every large org has the same bottleneck: the most valuable specialist knowledge lives in people's heads. Compliance reviews take days, the same questions recur across hundreds of product assessments, and inconsistency between experts creates real risk.
The instinctive fix — fine-tune the model on expert data — is expensive, opaque, and slow. A better approach treats the knowledge base as source code and the model as a compiler: keep complexity in text files readable by both humans and agents, not in model weights.
The architecture below comes from a production system built for a compliance domain, but the pattern generalizes to finance, security, and engineering standards review.
Four Layers, Each Solving One Problem
| Layer | Responsibility |
|---|---|
| Knowledge system | The organizational 'second brain' — curated positions, taxonomy, routing indexes |
| Reasoning layer | Composable 'recipes' that prescribe how to analyze, not what to know |
| Evaluation framework | Gates every proposed change with targeted replay + regression tests |
| Improvement loop | Compiles expert feedback into verified, version-controlled edits |
Remove any one layer and the others degrade. This is the key insight most RAG systems miss: retrieval alone doesn't capture reasoning.

The Knowledge Layer: Files Over Embeddings
Instead of stuffing documents into a vector store and hoping semantic search finds the right chunk, pre-extract knowledge into a strict taxonomy of 200+ files:
- Position files — authoritative organizational stances with constraints and machine-actionable routing implications
- Taxonomy / vocabulary files — single source of truth for entity types and classification tiers
- Routing indexes — deterministic mapping from input characteristics to applicable positions
- Gateway files — threshold tests the agent must pass before entering a domain
Every file declares its dependencies in YAML frontmatter:
# position_file.yaml
id: pos-2024-compliance-marketing-claims
version: 3.2
depends_on:
- taxonomy/entity-types.yaml
- vocabulary/claim-categories.yaml
referenced_by:
- recipes/assess-marketing-claim.yaml
- routing/index-marketing.yaml
constraints:
- "Applies only to consumer-facing claims"
- "Supersedes pos-2023-marketing-claims"
This forms a bidirectional dependency graph. When one file changes, you can trace exactly what breaks — which is what makes automated editing tractable.
Density-Based Partitioning
Not everything belongs in the wiki. Split on information density × usage frequency:
- High density + high frequency → wiki (consulted nearly every turn)
- Sparse + situational → RAG retrieval (deep but rarely needed)
The agent's core reasoning stays grounded in refined, current knowledge while still reaching for supporting evidence when a scenario demands it.

Recipes: Separating 'What' From 'How'
Knowledge files are declarative. Recipes are imperative. A recipe prescribes a multi-step analytical workflow — what to examine first, which knowledge to load at each step, what decision procedures to follow.
The critical design choice: recipes reference knowledge files but contain no domain facts; knowledge files state positions but prescribe no procedures.
This gives you clean failure attribution:
- Wrong conclusion despite correct source materials → recipe problem
- Correct procedure but missing facts → knowledge gap
- Experts disagree on the answer → ambiguity, escalate to human
Progressive Disclosure Cuts Tokens by ~80%
Early versions loaded a flat instruction file with everything via semantic search. After restructuring into recipe-driven stages, each query touches only a small targeted subset. Context windows are finite and attention degrades with volume — delivering the right instructions at the right time directly improves reasoning quality.
The Self-Improvement Flywheel
This is the part most teams skip. The loop has four phases:
- Diagnose — extract signals from expert conversation traces, then apply a single attribution test: Could the agent have reached the correct conclusion from its source materials?
- Compile — sub-agents analyze impact in parallel (cross-references, conflicts, token budget, test coverage). An independent adversarial reviewer runs in a fresh context with no knowledge of the rationale — it only sees the proposed diffs.
- Evaluate — targeted replay on the original scenario (with a blind judge) + regression tests across the domain benchmark.
- Land — a human reviews a proven fix, not a raw failure. The failing scenario is added to the regression suite permanently.
Every fix raises the bar for future changes. Expert effort compounds.
Limitations and Cautions
- Upfront cost is real. Building 200+ structured files with dependency graphs takes weeks. This only pays off if the domain has recurring, high-stakes questions.
- Not for open-ended creativity. This pattern fits domains governed by retrievable text and stable positions. It will fight you in fast-moving creative work.
- Human checkpoints are non-negotiable. The system accelerates experts; it does not replace their authority. Calibrate checkpoints to the domain's risk tolerance.
- Adversarial review can be gamed. If the reviewer agent shares any context with the proposer, blind spots propagate. Keep contexts strictly isolated.
Where to Go Next
If you're evaluating this for your org, start with the attribution test — it's the smallest useful unit. Pick one recurring expert correction and ask: was it a knowledge gap, a recipe flaw, or genuine ambiguity? That single question will tell you whether your system needs a knowledge layer, a reasoning layer, or a human escalation path.
For a complementary perspective on the limits of LLMs in evaluation contexts, see our analysis of the surrogacy assumption in LLM-based A/B testing. And if you're tracking where native platform features are replacing tooling layers, CSS @function is quietly doing to Sass what this architecture does to fine-tuning.
![]()
The Takeaway
The deeper principle is simple: keep the complexity in text files readable by both humans and agents, rather than in fine-tuned model weights. Every improvement is a diff a domain expert can review in 30 seconds. Every change is version-controlled, diffable, reversible.
The compilation pipeline is sophisticated, but its outputs are always transparent. That's the trade: engineering complexity in the pipeline, radical simplicity in the artifacts.
If your org has tribal knowledge that keeps walking out the door, this architecture is worth a serious look.