doubleo7/docs/case-study.md
Austin Schaefer f2c10783db
Some checks failed
CI / test (push) Failing after 6s
CI / test (pull_request) Failing after 7s
Extract swear_cleanup to its own repo, flatten deep_research to root
deep_research is the only project this repo is meant to showcase, so the
Cargo workspace wrapping it and an unrelated side project no longer earns
its keep:

- swear_cleanup moved to a new standalone local repo (~/dev/swear_cleanup,
  not pushed anywhere) via `git subtree split`, with its pre-workspace-
  split history (when it lived at src/swear_cleanup/ in a single shared
  crate) spliced onto its post-split history rather than starting from a
  single flattened snapshot. FINDINGS.md, which was sitting at this repo's
  root but was actually swear_cleanup's own build log, went with it.
- deep_research/{src,Cargo.toml,README.md,docs} moved to the repo root;
  the [workspace] table collapsed into a plain [package] manifest with
  dependency versions inlined from the old [workspace.dependencies].
- Cargo.toml keeps an explicit empty [workspace] table (not just omitted)
  so that checking this repo out as a nested git worktree — this
  project's own normal workflow — can't accidentally inherit a stale
  ancestor directory's workspace manifest, which is exactly what broke
  the build while testing this change from a worktree.
- .forgejo/workflows/deep_research-ci.yml -> ci.yml, dropping the now-
  meaningless -p deep_research scoping and path filters (redundant when
  it's the only thing in the repo).
- README.md and docs/case-study.md updated for the flattened commands
  (cargo run/test with no -p flag); their relative links to each other
  and to src/ were already correct since both moved together.

Verified: cargo build/test/clippy/fmt all clean from the new repo root.
2026-08-18 13:43:26 +02:00

9 KiB

Case study: hardening an agentic pipeline against real failure

This is a walkthrough of how deep_research's multi-agent pipeline went from "works in the happy path" to actually resilient — driven by two real failures it hit in normal use, not by imagining edge cases in the abstract. The throughline: root-cause failures instead of papering over them, and verify fixes against a live run, not just a green test suite.

The starting point

The pipeline is four agents with one job each — a researcher that tool-calls its way through web search and page fetches, a reviewer that checks the researcher's conclusions actually follow from its sources, a writer that turns approved findings into a report, and (added in the course of this work) a summarizer that only runs as a recovery path. See the README for the full shape.

Like any agent given a tool-calling budget, the researcher can run out of turns before it decides it's done. The original code treated that as fatal: whatever evidence had been gathered — search results, fetched pages, partial reasoning — was simply discarded when the run errored out.

Failure #1: turn-budget exhaustion, and why "just increase the limit" isn't the fix

Raising the turn cap doesn't solve this class of problem — it just moves the threshold. Any fixed budget can be exhausted by a topic that's slow to narrow down, and the failure mode (total loss of partial work) is the actual defect, not the specific number of turns.

The fix meant reading past the top-level API surface and into the actual error rig (the Rust agent framework this is built on) returns when the budget runs out: PromptError::MaxTurnsError, which — critically — carries the full chat history at the point of failure, not just an error string. That's the hook the recovery path needed.

First pass was a purely programmatic recovery: catch that specific error variant, walk the returned message history, and pull out whatever assistant commentary and tool-call results it contains into a flat findings dump, instead of raising.

Second pass made it genuinely agentic rather than just defensive: a dedicated summarizer agent — a fresh, tool-free model call — takes that same transcript (now annotated with which tool call produced which result, so a fetched page stays attached to its URL) and reconstructs the same footnote-style findings dump the researcher would have written itself, had it finished. This recovers real information that plain string concatenation would leave scattered and unattributed: deduped sources, correct citation numbering, a coherent narrative instead of a raw tool-output dump.

The summarizer call is itself just another model call that can fail or have nothing to work with — so it falls back to the programmatic extraction on an empty transcript or a failed call, rather than letting a second failure take down the one recovery path meant to be bulletproof. And regardless of what either path produces, a plain, hard-coded disclaimer is written directly to the output — not left to an LLM's discretion to remember to mention that the research is incomplete.

This was verified against a real run, not a mocked one: a genuinely under-specified research question was run end-to-end against local models until it actually exhausted its turn budget, with full tracing enabled, to confirm the recovery path fired, the summarizer produced a coherent partial report, and the disclaimer showed up exactly where expected.

Failure #2: the recovery path was masking a worse problem

That live verification run surfaced something the recovery path was built to survive, but shouldn't have had to: the researcher burned its entire turn budget re-running variations of the same search, and every single one came back "no results found." The web-search tool was scraping DuckDuckGo's HTML search page directly — no API key required, but no protection from rate limiting either. And a rate-limited response looks identical to a genuine empty result: the tool has no way to tell the model "you're being throttled" versus "this topic has no coverage," so the model just kept trying.

The recovery mechanism did exactly its job here — the run finished with an honest report saying "found nothing," rather than crashing — but that's a consolation prize. The actual bug was upstream: the search tool's failure mode was silent and indistinguishable from success.

Root cause fixed, not the symptom: the search tool now hits a self-hosted SearXNG instance's JSON API instead of scraping HTML. That's a straight upgrade on every axis that mattered here — a real API instead of parsing markup, results aggregated across multiple upstream engines instead of hammering one, and full control over request pacing since it's infrastructure already running on this machine. Before reaching for a third-party crate, the two SearXNG client crates on crates.io were checked and rejected: both single-maintainer, both v0.1.0, neither with any adoption signal — a ~20-line reqwest + serde call using dependencies already in the tree was the better bet for something this small.

Verification, again against a live run: the exact same research question that had previously burned its full turn budget on empty results was re-run, unchanged, against the new search backend. It completed normally on the second research round — approved by the reviewer, no max-turns event, no recovery path needed. The fix wasn't just plausible on paper; it was confirmed to actually change the outcome of the failure it was meant to fix.

Hardening: separation of concerns

With both failure modes fixed, the module that had accumulated all of this logic — CLI parsing, orchestration, the researcher phase, chat-history reconstruction, the summarizer, and the writer — had grown into a single 600+-line file mixing six unrelated concerns, while the rest of the codebase (review.rs, tools.rs, stream.rs, progress.rs) already kept one file per concern. That file was split to match: each agent phase, the orchestration loop, and the pure history-parsing helpers now live in their own module, with main.rs reduced to argument parsing, logging setup, and a single top-level call — see the project layout in the README for the resulting structure. Unit tests moved with the code they test (Rust convention keeps tests co-located, not split into separate files), and the full suite — now covering the history-reconstruction logic in isolation, independent of any model or network call — stayed green throughout.

A clean run, for reference

With both failures fixed, a normal research pass looks like this — the researcher gathered evidence, the reviewer approved it on the first pass, and the writer produced a cited report, no recovery path exercised:

Topic: trends in AI customer-support chatbots

# Trends in AI Customer-Support Chatbots (2026)

This report synthesizes current research notes detailing the major
technological, operational, and regulatory trends shaping the AI
customer-support chatbot market as of 2026.

### The Shift to Agentic AI
The industry is undergoing a fundamental transition from simple, scripted
chatbots toward "Agentic AI." This new paradigm focuses on autonomous AI
agents capable of performing complex tasks and orchestrating complete
workflows, moving beyond basic prompt responses [1, 4]...

### Governance, Security, and Regulation
...compliance with regulations such as GDPR and the EU AI Act is a primary
concern, demanding robust structures for security, transparency, and
governance [4, 6]...

### Operational and Economic Impact
The overall AI-powered customer service market is projected to reach
$15.12 billion in 2026 [1]. Adoption is widespread, with approximately 72%
of businesses across various industries having deployed AI-driven chatbots
for customer interactions [2]...

### Sources
[1] https://chatmaxima.com/blog/ai-customer-support-statistics-2026/
[2] https://www.itransition.com/ai/conversational
[3] https://salt.security/eu-ai-act-compliance
...

(Abbreviated here; the tool prints the full report, headings, open questions, and complete source list to the terminal as it streams.)

What this demonstrates

  • Designing a multi-agent pipeline as explicit, single-responsibility stages coordinated by plain control flow, not a single sprawling prompt or an opaque framework agent-graph.
  • Treating an agent's failure modes (turn-budget exhaustion, a tool's silent degradation) as defects to root-cause and fix, not edge cases to shrug off — including building a second, self-limiting agent as the recovery mechanism itself.
  • Verifying fixes against live runs of the actual failing case, not just unit tests in isolation.
  • A local-first architecture (Ollama + self-hosted SearXNG) with no cloud LLM API key and no query leaving the host — the same constraint a privacy-sensitive deployment would need.
  • Full observability via structured tracing on every agent phase, and a codebase organized so each concern has exactly one home.