When Deterministic Pipelines Outperform Agentic Wandering

An engineered context distillation pipeline outperformed our best frontier-model agent loop.

Kanav Petkar

Member of Technical Staff

Matt Shu

Member of Technical Staff @ Brain Co

We benchmarked frontier LLM agent loops equipped with tool calls on the task of evaluating clinical rules over unstructured Electronic Health Records (EHR) data. A carefully engineered context distillation pipeline still won across accuracy, latency, and cost metrics. In this post, we explain why we replaced an agent-based approach with a deterministic context-distillation pipeline. We show why deterministic pipelines can outperform agents in closed-world domains.

The Problem

One project at Brain Co. is analyzing patient records to identify all the places in a patient’s care journey where they have diverged from their ideal care pathway (e.g., a diabetic patient missing a specific eye exam). Once these gaps are identified, patients can then be connected with the right clinicians to address them. This results in patients getting more targeted healthcare for their specific issues, alongside an operations benefit for providers.

This is a classic “needle in a haystack” problem. The clue for a diagnosis might be a single line in a PDF from three years ago, among decades of unstructured notes and labs (that might be incomplete!). Furthermore, as a healthcare task, both high precision and recall are required. We cannot hallucinate a diagnosis or miss a critical lab result.

The Agentic Allure

At first glance, this seemed a reasonable fit for an agentic approach, where the agent loop gets tools to search and filter (e.g., get_labs, search_notes) the EHR corpus, and lets it "think" its way to the answer. An approach like this has the benefit of not having to hand-craft logic and custom instructions for every new disease we roll out support for.

We built out a search layer over EHR data that enabled the fast retrieval of key fields (think lab results, medications, or demographics) and configured the agent to use it to retrieve the specific bits of the patient record that the agent thought might be useful to identify gaps.

With this layer built, we were ready to test out our agentic approach.

Expectation vs. Reality

Results were unexpectedly poor. Our baseline F1 score was 36%. After several engineering cycles spent on prompt tuning and tool definition improvements, we were able to push performance to 62.2% F1. but progress stalled well below what would be a usable system in production.

Across these iterations, a pattern emerged. Each agent improvement increased performance incrementally, but these gains were expensive. Every step added latency, cost, and variance.

When we went back to the drawing board, we asked why this strategy didn’t work. By handing the entire workflow to an agent, we forced the model to re-derive its search strategy every single time. The agent had to decide what to look for, how to look for it, and when to stop. Outside of accuracy issues, this flexibility also introduced variance and latency in inference results.

The Insight: Separating Search from Reasoning

Agents excel in tasks where they need to perform external retrieval or acquire new knowledge to make decisions. However, in our case, all the relevant information was already present in the patient record.

We realized that we were using this agentic approach as a roundabout way to prune the information we were passing to the model.

Through shadowing clinicians, we learned that clinicians follow a general order of operations when making these decisions manually. For a given rule, a doctor would:

  1. Identify the relevant lab results, demographics, and medications useful for making a determination.
  2. Search through the patient record for those specific fields.
  3. Synthesize the identified information into a clinical narrative to make a decision.

Agentic tool use allows models to query for this information, but turning this procedure over to an agent meant the agent had to "think through" which information was relevant each time it was called.

This gave the agent too much free rein. It meant the agent often failed to determine what was relevant, wasting compute on superfluous information. It introduced significant stochasticity and increased latency—two factors that made iterating on the approach nearly impossible.

The Solution: A Deterministic Context Distillation Approach

As we analyzed why the agentic approach wasn't working, a simpler approach emerged.

We realized that we had conflated the messiness of the data with the complexity of the reasoning. While a patient record is an “open world” of unstructured noise, the clinical rule itself is a closed-world problem.

For a diabetic eye exam rule the set of relevant evidence is finite and known in advance (e.g., ophthalmology notes, specific CPT codes). We didn't need an agent to discover which fields mattered; we only needed to extract the correct signals.

This gives the LLM a clean, compact context, makes evidence selection deterministic, and constrains the model’s decision context. Effectively, we split the problem into two distinct, independently reviewable tasks:

  1. Evidence Gathering (Pruning): Is the record pruning comprehensive (all necessary info) and concise (no noise)?
  2. Diagnosis (Reasoning): Is the model able to correctly make a determination given the proper evidence?

Figure 1 shows how our context distillation algorithm works:

Step 1: Extract Relevant Evidence For each rule, we identify the specific data types that matter: lab values, medication histories, diagnosis codes, and encounter summaries. We extract no more and no less.

Step 2: Normalize and Structure Snippets Once extracted, we normalize the structure through targeted, disease-specific evidence selection. We employ a combination of matching algorithms optimized for different resource types:

  • Token-set matching for medications.
  • Fuzzy matching for clinical terms (capturing variations like HbA1c versus hemoglobin A1C).
  • Partial ratio matching for free-text clinical notes.

Step 3: Temporal Fallbacks When these strategies do not yield results (perhaps because the patient never received the relevant treatment), we ensure the LLM still has data to work with through fallbacks filtered only by time.

We then sort by temporal relevance (most recent first), apply category-specific limits, and deduplicate objects to prevent redundancy. This exposes only rule-relevant evidence to the LLM while maintaining referential integrity.

Step 4: Pass Pruned Context to the Model Finally, we pass the pruned context and the rule to the model. The result is a single input and a single output, with no tool calls and no agent loops, which makes every decision straightforward to audit.

The Results

Once we integrated this context distillation approach, the impact was immediate. As seen in Figures 2 and 3, performance improved across experimental iterations for both approaches, but plateaued much earlier for agentic tool use. While our best agentic system peaked at 62.2% F1 after multiple rounds of tuning, the context-pruned pipeline surpassed 80% F1 in its first refined version and reached 94.5% F1 after incorporating higher-quality evaluation data.

These gains were not confined to a single disease or metric. The same architectural shift produced consistent improvements across accuracy, latency, and cost:

  • Accuracy: This new approach helped us more easily identify what stage of the pipeline caused a particular error; i.e., overly strict distillation leading to missed evidence, such as when fuzzy matching failed to capture clinical term variations like Hb1Ac versus hemoglobin A1C. When evidence was clearly present, we could then focus on tuning diagnosis prompts to further boost accuracy.
  • Latency: Eliminating multi-turn agent loops, tool-selection overhead, and "agent thinking time" between search iterations significantly reduced latency. As a result, the total runtime for our Type II Diabetes pipeline across the full patient cohort dropped from 28 minutes to 3 minutes, more than 9x improvement in speed.
  • Cost: Token usage was 20% lower with context distillation compared to the overhead of agent planning and formatting tool calls.

When You Should Use Agents

None of this is an indictment of agents. They remain enormously valuable. They just weren't right for our task. Agents shine when you truly need exploration or external knowledge like drug interaction lookups, payer formulary rules, or web search. They also fit domains that lack the structure or consistent terminology, workflows that are judgement-driven rather than rules-driven, and tasks that require long-horizon, multi-step reasoning.

As we expand into disease areas with more ambiguous clinical patterns (e.g., rheumatology, oncology), agents may again become the right tool. But for certain chronic conditions like diabetes, the protocols are well-defined even if the data is messy.

In this landscape, asking an agent to “explore” introduces unnecessary variance; a precise extraction pipeline is better.

Conclusion

Closed-world clinical rules benefit from a fixed evidence-gathering process. In our evaluation, context distillation improved F1, latency, and cost by preventing the model from rebuilding its search strategy on every run.

Our data was messy and inconsistently structured, so we thought agents were best equipped to explore it. However, clinical rules are closed-world problems with a finite set of relevant evidence. In this setting, building tools to deterministically retrieve this evidence outperformed asking agents to re-derive what clues to look for and how to find them. We learned that reliability doesn't come from making the model smarter; it comes from aligning the architecture with the strict requirements of the domain.

Key Takeaways:

  • Fixed evidence selection beat agentic search. In closed-world clinical rules, deterministic evidence selection consistently outperformed increasingly sophisticated agentic tool use. When you already know what evidence matters, don't ask the model to guess.
  • Justify added complexity. Agent tool use is powerful but it introduces unnecessary noise when the full solution already exists inside the record.
  • Separate the concerns. Separating evidence-gathering from clinical reasoning enables clearer debugging, better evaluations, and safer clinical behavior.
  • Problem-first design. Choose the approach based on the structure of the problem, not on the capabilities of the model.

Relevant and Related Resources

Join UsSee open roles

All posts