Add CI/CD, README, and a case study documenting this session's work
- .forgejo/workflows/deep_research-ci.yml: build, test, clippy (-D
warnings), and fmt --check on push/PR, scoped to deep_research (not
workspace-wide — swear_cleanup has an unrelated pre-existing clippy
warning that would otherwise break CI on an unrelated project)
- README.md: what the project does, the four-agent architecture, why
it's local-first (Ollama + self-hosted SearXNG, no cloud API key, no
query leaves the host), project layout, and how to run/test it
- docs/case-study.md: narrative walkthrough of the max-turns recovery
path, the DuckDuckGo-rate-limiting root cause and SearXNG fix, and the
separation-of-concerns refactor — each step verified against a live
run of the actual failing case, not just unit tests. Uses a neutral
"AI customer-support chatbot trends" research run as the illustrative
clean-pipeline example rather than the personal topic used during
actual debugging.
2026-08-18 11:28:29 +00:00
|
|
|
# Case study: hardening an agentic pipeline against real failure
|
|
|
|
|
|
2026-08-19 10:11:47 +00:00
|
|
|
This is a walkthrough of how `deep_research`'s multi-agent pipeline went
|
|
|
|
|
from "works in the happy path" to an improved resiliency — driven by two real failures
|
|
|
|
|
it hit in normal use. The throughline: root-cause failure debugging via manual testing and
|
|
|
|
|
post-fix verification.
|
Add CI/CD, README, and a case study documenting this session's work
- .forgejo/workflows/deep_research-ci.yml: build, test, clippy (-D
warnings), and fmt --check on push/PR, scoped to deep_research (not
workspace-wide — swear_cleanup has an unrelated pre-existing clippy
warning that would otherwise break CI on an unrelated project)
- README.md: what the project does, the four-agent architecture, why
it's local-first (Ollama + self-hosted SearXNG, no cloud API key, no
query leaves the host), project layout, and how to run/test it
- docs/case-study.md: narrative walkthrough of the max-turns recovery
path, the DuckDuckGo-rate-limiting root cause and SearXNG fix, and the
separation-of-concerns refactor — each step verified against a live
run of the actual failing case, not just unit tests. Uses a neutral
"AI customer-support chatbot trends" research run as the illustrative
clean-pipeline example rather than the personal topic used during
actual debugging.
2026-08-18 11:28:29 +00:00
|
|
|
|
|
|
|
|
## The starting point
|
|
|
|
|
|
|
|
|
|
The pipeline is four agents with one job each — a researcher that
|
|
|
|
|
tool-calls its way through web search and page fetches, a reviewer that
|
|
|
|
|
checks the researcher's conclusions actually follow from its sources, a
|
|
|
|
|
writer that turns approved findings into a report, and (added in the course
|
|
|
|
|
of this work) a summarizer that only runs as a recovery path. See the
|
|
|
|
|
[README](../README.md#architecture) for the full shape.
|
|
|
|
|
|
|
|
|
|
Like any agent given a tool-calling budget, the researcher can run out of
|
|
|
|
|
turns before it decides it's done. The original code treated that as fatal:
|
|
|
|
|
whatever evidence had been gathered — search results, fetched pages, partial
|
|
|
|
|
reasoning — was simply discarded when the run errored out.
|
|
|
|
|
|
|
|
|
|
## Failure #1: turn-budget exhaustion, and why "just increase the limit" isn't the fix
|
|
|
|
|
|
|
|
|
|
Raising the turn cap doesn't solve this class of problem — it just moves the
|
|
|
|
|
threshold. Any fixed budget can be exhausted by a topic that's slow to
|
|
|
|
|
narrow down, and the failure mode (total loss of partial work) is the actual
|
|
|
|
|
defect, not the specific number of turns.
|
|
|
|
|
|
|
|
|
|
The fix meant reading past the top-level API surface and into the actual
|
|
|
|
|
error `rig` (the Rust agent framework this is built on) returns when the
|
|
|
|
|
budget runs out: `PromptError::MaxTurnsError`, which — critically — carries
|
|
|
|
|
the full chat history at the point of failure, not just an error string.
|
|
|
|
|
That's the hook the recovery path needed.
|
|
|
|
|
|
|
|
|
|
**First pass** was a purely programmatic recovery: catch that specific
|
|
|
|
|
error variant, walk the returned message history, and pull out whatever
|
|
|
|
|
assistant commentary and tool-call results it contains into a flat findings
|
|
|
|
|
dump, instead of raising.
|
|
|
|
|
|
|
|
|
|
**Second pass** made it genuinely agentic rather than just defensive: a
|
|
|
|
|
dedicated *summarizer* agent — a fresh, tool-free model call — takes that
|
|
|
|
|
same transcript (now annotated with which tool call produced which result,
|
|
|
|
|
so a fetched page stays attached to its URL) and reconstructs the same
|
|
|
|
|
footnote-style findings dump the researcher would have written itself, had
|
|
|
|
|
it finished. This recovers real information that plain string concatenation
|
|
|
|
|
would leave scattered and unattributed: deduped sources, correct citation
|
|
|
|
|
numbering, a coherent narrative instead of a raw tool-output dump.
|
|
|
|
|
|
|
|
|
|
The summarizer call is itself just another model call that can fail or
|
|
|
|
|
have nothing to work with — so it falls back to the programmatic extraction
|
|
|
|
|
on an empty transcript or a failed call, rather than letting a second
|
|
|
|
|
failure take down the one recovery path meant to be bulletproof. And
|
|
|
|
|
regardless of what either path produces, a plain, hard-coded disclaimer is
|
|
|
|
|
written directly to the output — not left to an LLM's discretion to
|
|
|
|
|
remember to mention that the research is incomplete.
|
|
|
|
|
|
|
|
|
|
This was verified against a real run, not a mocked one: a genuinely
|
|
|
|
|
under-specified research question was run end-to-end against local models
|
|
|
|
|
until it actually exhausted its turn budget, with full tracing enabled, to
|
|
|
|
|
confirm the recovery path fired, the summarizer produced a coherent partial
|
|
|
|
|
report, and the disclaimer showed up exactly where expected.
|
|
|
|
|
|
|
|
|
|
## Failure #2: the recovery path was masking a worse problem
|
|
|
|
|
|
|
|
|
|
That live verification run surfaced something the recovery path was built
|
|
|
|
|
to survive, but shouldn't have had to: the researcher burned its *entire*
|
|
|
|
|
turn budget re-running variations of the same search, and every single one
|
|
|
|
|
came back "no results found." The web-search tool was scraping
|
|
|
|
|
DuckDuckGo's HTML search page directly — no API key required, but no
|
|
|
|
|
protection from rate limiting either. And a rate-limited response looks
|
|
|
|
|
*identical* to a genuine empty result: the tool has no way to tell the
|
|
|
|
|
model "you're being throttled" versus "this topic has no coverage," so the
|
|
|
|
|
model just kept trying.
|
|
|
|
|
|
|
|
|
|
The recovery mechanism did exactly its job here — the run finished with an
|
|
|
|
|
honest report saying "found nothing," rather than crashing — but that's a
|
|
|
|
|
consolation prize. The actual bug was upstream: the search tool's failure
|
|
|
|
|
mode was silent and indistinguishable from success.
|
|
|
|
|
|
|
|
|
|
Root cause fixed, not the symptom: the search tool now hits a self-hosted
|
|
|
|
|
[SearXNG](https://docs.searxng.org/) instance's JSON API instead of
|
|
|
|
|
scraping HTML. That's a straight upgrade on every axis that mattered here —
|
|
|
|
|
a real API instead of parsing markup, results aggregated across multiple
|
|
|
|
|
upstream engines instead of hammering one, and full control over request
|
|
|
|
|
pacing since it's infrastructure already running on this machine. Before
|
|
|
|
|
reaching for a third-party crate, the two SearXNG client crates on
|
|
|
|
|
crates.io were checked and rejected: both single-maintainer, both v0.1.0,
|
|
|
|
|
neither with any adoption signal — a ~20-line `reqwest` + `serde` call
|
|
|
|
|
using dependencies already in the tree was the better bet for something
|
|
|
|
|
this small.
|
|
|
|
|
|
|
|
|
|
**Verification, again against a live run**: the exact same research
|
|
|
|
|
question that had previously burned its full turn budget on empty results
|
|
|
|
|
was re-run, unchanged, against the new search backend. It completed
|
|
|
|
|
normally on the second research round — approved by the reviewer, no
|
|
|
|
|
max-turns event, no recovery path needed. The fix wasn't just plausible on
|
|
|
|
|
paper; it was confirmed to actually change the outcome of the failure it
|
|
|
|
|
was meant to fix.
|
|
|
|
|
|
|
|
|
|
## Hardening: separation of concerns
|
|
|
|
|
|
|
|
|
|
With both failure modes fixed, the module that had accumulated all of this
|
|
|
|
|
logic — CLI parsing, orchestration, the researcher phase, chat-history
|
|
|
|
|
reconstruction, the summarizer, and the writer — had grown into a single
|
|
|
|
|
600+-line file mixing six unrelated concerns, while the rest of the
|
|
|
|
|
codebase (`review.rs`, `tools.rs`, `stream.rs`, `progress.rs`) already kept
|
|
|
|
|
one file per concern. That file was split to match: each agent phase, the
|
|
|
|
|
orchestration loop, and the pure history-parsing helpers now live in their
|
|
|
|
|
own module, with `main.rs` reduced to argument parsing, logging setup, and
|
|
|
|
|
a single top-level call — see the [project layout](../README.md#project-layout)
|
|
|
|
|
in the README for the resulting structure. Unit tests moved with the code
|
|
|
|
|
they test (Rust convention keeps tests co-located, not split into separate
|
|
|
|
|
files), and the full suite — now covering the history-reconstruction logic
|
|
|
|
|
in isolation, independent of any model or network call — stayed green
|
|
|
|
|
throughout.
|
|
|
|
|
|
|
|
|
|
## A clean run, for reference
|
|
|
|
|
|
|
|
|
|
With both failures fixed, a normal research pass looks like this — the
|
|
|
|
|
researcher gathered evidence, the reviewer approved it on the first pass,
|
|
|
|
|
and the writer produced a cited report, no recovery path exercised:
|
|
|
|
|
|
|
|
|
|
> **Topic:** trends in AI customer-support chatbots
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
# Trends in AI Customer-Support Chatbots (2026)
|
|
|
|
|
|
|
|
|
|
This report synthesizes current research notes detailing the major
|
|
|
|
|
technological, operational, and regulatory trends shaping the AI
|
|
|
|
|
customer-support chatbot market as of 2026.
|
|
|
|
|
|
|
|
|
|
### The Shift to Agentic AI
|
|
|
|
|
The industry is undergoing a fundamental transition from simple, scripted
|
|
|
|
|
chatbots toward "Agentic AI." This new paradigm focuses on autonomous AI
|
|
|
|
|
agents capable of performing complex tasks and orchestrating complete
|
|
|
|
|
workflows, moving beyond basic prompt responses [1, 4]...
|
|
|
|
|
|
|
|
|
|
### Governance, Security, and Regulation
|
|
|
|
|
...compliance with regulations such as GDPR and the EU AI Act is a primary
|
|
|
|
|
concern, demanding robust structures for security, transparency, and
|
|
|
|
|
governance [4, 6]...
|
|
|
|
|
|
|
|
|
|
### Operational and Economic Impact
|
|
|
|
|
The overall AI-powered customer service market is projected to reach
|
|
|
|
|
$15.12 billion in 2026 [1]. Adoption is widespread, with approximately 72%
|
|
|
|
|
of businesses across various industries having deployed AI-driven chatbots
|
|
|
|
|
for customer interactions [2]...
|
|
|
|
|
|
|
|
|
|
### Sources
|
|
|
|
|
[1] https://chatmaxima.com/blog/ai-customer-support-statistics-2026/
|
|
|
|
|
[2] https://www.itransition.com/ai/conversational
|
|
|
|
|
[3] https://salt.security/eu-ai-act-compliance
|
|
|
|
|
...
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
(Abbreviated here; the tool prints the full report, headings, open
|
|
|
|
|
questions, and complete source list to the terminal as it streams.)
|
|
|
|
|
|
|
|
|
|
## What this demonstrates
|
|
|
|
|
|
|
|
|
|
- Designing a multi-agent pipeline as explicit, single-responsibility
|
|
|
|
|
stages coordinated by plain control flow, not a single sprawling prompt
|
|
|
|
|
or an opaque framework agent-graph.
|
|
|
|
|
- Treating an agent's failure modes (turn-budget exhaustion, a tool's
|
|
|
|
|
silent degradation) as defects to root-cause and fix, not edge cases to
|
|
|
|
|
shrug off — including building a second, self-limiting agent as the
|
|
|
|
|
recovery mechanism itself.
|
|
|
|
|
- Verifying fixes against live runs of the actual failing case, not just
|
|
|
|
|
unit tests in isolation.
|
|
|
|
|
- A local-first architecture (Ollama + self-hosted SearXNG) with no
|
|
|
|
|
cloud LLM API key and no query leaving the host — the same constraint
|
|
|
|
|
a privacy-sensitive deployment would need.
|
|
|
|
|
- Full observability via structured tracing on every agent phase, and a
|
|
|
|
|
codebase organized so each concern has exactly one home.
|