From 9fa91b3da7123b914b07f4b07639ed7db33edee8 Mon Sep 17 00:00:00 2001 From: Austin Schaefer Date: Tue, 18 Aug 2026 13:28:29 +0200 Subject: [PATCH] Add CI/CD, README, and a case study documenting this session's work MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - .forgejo/workflows/deep_research-ci.yml: build, test, clippy (-D warnings), and fmt --check on push/PR, scoped to deep_research (not workspace-wide β€” swear_cleanup has an unrelated pre-existing clippy warning that would otherwise break CI on an unrelated project) - README.md: what the project does, the four-agent architecture, why it's local-first (Ollama + self-hosted SearXNG, no cloud API key, no query leaves the host), project layout, and how to run/test it - docs/case-study.md: narrative walkthrough of the max-turns recovery path, the DuckDuckGo-rate-limiting root cause and SearXNG fix, and the separation-of-concerns refactor β€” each step verified against a live run of the actual failing case, not just unit tests. Uses a neutral "AI customer-support chatbot trends" research run as the illustrative clean-pipeline example rather than the personal topic used during actual debugging. --- .forgejo/workflows/deep_research-ci.yml | 47 +++++++ deep_research/README.md | 111 +++++++++++++++ deep_research/docs/case-study.md | 175 ++++++++++++++++++++++++ 3 files changed, 333 insertions(+) create mode 100644 .forgejo/workflows/deep_research-ci.yml create mode 100644 deep_research/README.md create mode 100644 deep_research/docs/case-study.md diff --git a/.forgejo/workflows/deep_research-ci.yml b/.forgejo/workflows/deep_research-ci.yml new file mode 100644 index 0000000..7c1cdad --- /dev/null +++ b/.forgejo/workflows/deep_research-ci.yml @@ -0,0 +1,47 @@ +name: deep_research CI + +on: + push: + branches: ['**'] + paths: + - 'deep_research/**' + - '.forgejo/workflows/deep_research-ci.yml' + pull_request: + paths: + - 'deep_research/**' + - '.forgejo/workflows/deep_research-ci.yml' + +jobs: + test: + runs-on: docker + container: rust:1-bookworm + + steps: + - name: Checkout + uses: https://code.forgejo.org/actions/checkout@v4 + + - name: Cache cargo registry and build artifacts + uses: https://code.forgejo.org/actions/cache@v4 + with: + path: | + ~/.cargo/registry + ~/.cargo/git + target + key: cargo-${{ runner.os }}-${{ hashFiles('Cargo.lock') }} + restore-keys: | + cargo-${{ runner.os }}- + + - name: Install rustfmt and clippy + run: rustup component add rustfmt clippy + + - name: Check formatting + run: cargo fmt -p deep_research -- --check + + - name: Lint + run: cargo clippy -p deep_research --all-targets -- -D warnings + + - name: Build + run: cargo build -p deep_research + + - name: Test + run: cargo test -p deep_research diff --git a/deep_research/README.md b/deep_research/README.md new file mode 100644 index 0000000..aa2bd68 --- /dev/null +++ b/deep_research/README.md @@ -0,0 +1,111 @@ +# deep_research + +A local-first, multi-agent deep-research CLI: give it a topic, it searches +the web, cross-checks what it finds, and writes up a cited report β€” entirely +on infrastructure you control, with no cloud LLM API key and no query ever +leaving your machine. + +``` +$ deep_research "trends in AI customer-support chatbots" +πŸ”Ž Researching... +🧐 Reviewing findings... +✍️ Writing report... + +# AI Customer-Support Chatbots: 2025–2026 Trends +... +``` + +## Why this exists + +This started as a "does deep research actually work end-to-end" exercise and +turned into a small case study in building an *agentic* system that survives +contact with reality: models that hit their turn budget mid-task, search +providers that rate-limit, and reviewers that reject good-faith work. The +[case study](./docs/case-study.md) walks through what broke and how each +failure was fixed, not just papered over. + +## Architecture + +Four small agents, each with one job, coordinated by plain Rust control +flow β€” not a framework's agent graph, not an LLM deciding when to stop: + +``` +β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” approve/reject β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” +β”‚ researcher β”‚ ───────────────► β”‚ reviewer β”‚ +β”‚ (tool-using)β”‚ ◄─────────────── β”‚ β”‚ +β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜ gaps/feedback β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜ + β”‚ turn budget exhausted β”‚ approved, + β”‚ mid-investigation β”‚ or out of rounds + β–Ό β–Ό +β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” +β”‚ summarizer │──findings───►│ writer │──► report +β”‚ (recovery) β”‚ β”‚ β”‚ +β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ +``` + +- **researcher** β€” a tool-calling agent (`search_web`, `fetch_page`) that + gathers and cross-checks evidence, capped at a fixed model-call budget so + a confused model can't loop forever. +- **reviewer** β€” a separate, fresh-context agent that checks the researcher's + conclusions actually follow from its cited sources, and either approves + the findings or hands back concrete gaps for another pass. +- **writer** β€” turns approved (or partial) findings into a structured, + footnoted report, streamed to the terminal as it's generated. +- **summarizer** (recovery path) β€” only runs when the researcher exhausts + its turn budget before concluding on its own. It reconstructs a proper + findings dump from the raw tool-call transcript rather than the run + simply failing; see the case study for why this exists and how it + degrades gracefully if the summarizer call itself fails. + +Everything runs against local models via [Ollama](https://ollama.com) and a +self-hosted [SearXNG](https://searx.space) instance for search β€” no OpenAI/ +Anthropic/Google API key, no third-party search API, nothing about the +research topic leaves the host it runs on. That's a deliberate constraint, +not a limitation: it's the same shape a privacy-sensitive customer +deployment would need. + +## Running it + +Prerequisites: +- [Ollama](https://ollama.com) running locally with a tool-calling-capable + model pulled (the researcher and reviewer/writer models are configured in + [`src/models.rs`](./src/models.rs)) +- A local [SearXNG](https://docs.searxng.org/) instance with its JSON API + enabled (defaults to `http://localhost:8080`, overridable via + `SEARXNG_URL`) + +``` +cargo run -p deep_research -- "your research topic" + +# or, with tracing spans on stderr instead of the progress spinner: +cargo run -p deep_research -- -l info "your research topic" +``` + +## Project layout + +Split one concern per file rather than one large module: + +| File | Responsibility | +|---|---| +| `main.rs` | Argument parsing, logging setup, and the single top-level call β€” no orchestration logic | +| `research.rs` | The research/review round loop | +| `researcher.rs` | The tool-calling research phase | +| `review.rs` | The reviewer agent | +| `summarizer.rs` | Max-turns recovery: reconstructs findings via a model call | +| `writer.rs` | Turns findings into the final streamed report | +| `history.rs` | Pure, unit-tested helpers for parsing a rig chat history into usable text | +| `tools.rs` | `search_web` (SearXNG) and `fetch_page` tool implementations | +| `stream.rs` | Drains a streaming prompt response to the terminal | +| `progress.rs` | The terminal spinner and per-phase emoji | +| `models.rs`, `cli.rs`, `observability.rs` | Small shared config: model names, CLI args, tracing setup | + +## Testing + +``` +cargo test -p deep_research # unit tests β€” pure functions, no network +cargo test -p deep_research -- --ignored # + a live smoke test against SearXNG +cargo clippy -p deep_research --all-targets +``` + +CI (`.forgejo/workflows/deep_research-ci.yml`) runs formatting, lint, build, +and the unit test suite on every push and PR. diff --git a/deep_research/docs/case-study.md b/deep_research/docs/case-study.md new file mode 100644 index 0000000..85dc757 --- /dev/null +++ b/deep_research/docs/case-study.md @@ -0,0 +1,175 @@ +# Case study: hardening an agentic pipeline against real failure + +This is a walkthrough of how `deep_research`'s multi-agent pipeline went +from "works in the happy path" to actually resilient β€” driven by two real +failures it hit in normal use, not by imagining edge cases in the abstract. +The throughline: root-cause failures instead of papering over them, and +verify fixes against a live run, not just a green test suite. + +## The starting point + +The pipeline is four agents with one job each β€” a researcher that +tool-calls its way through web search and page fetches, a reviewer that +checks the researcher's conclusions actually follow from its sources, a +writer that turns approved findings into a report, and (added in the course +of this work) a summarizer that only runs as a recovery path. See the +[README](../README.md#architecture) for the full shape. + +Like any agent given a tool-calling budget, the researcher can run out of +turns before it decides it's done. The original code treated that as fatal: +whatever evidence had been gathered β€” search results, fetched pages, partial +reasoning β€” was simply discarded when the run errored out. + +## Failure #1: turn-budget exhaustion, and why "just increase the limit" isn't the fix + +Raising the turn cap doesn't solve this class of problem β€” it just moves the +threshold. Any fixed budget can be exhausted by a topic that's slow to +narrow down, and the failure mode (total loss of partial work) is the actual +defect, not the specific number of turns. + +The fix meant reading past the top-level API surface and into the actual +error `rig` (the Rust agent framework this is built on) returns when the +budget runs out: `PromptError::MaxTurnsError`, which β€” critically β€” carries +the full chat history at the point of failure, not just an error string. +That's the hook the recovery path needed. + +**First pass** was a purely programmatic recovery: catch that specific +error variant, walk the returned message history, and pull out whatever +assistant commentary and tool-call results it contains into a flat findings +dump, instead of raising. + +**Second pass** made it genuinely agentic rather than just defensive: a +dedicated *summarizer* agent β€” a fresh, tool-free model call β€” takes that +same transcript (now annotated with which tool call produced which result, +so a fetched page stays attached to its URL) and reconstructs the same +footnote-style findings dump the researcher would have written itself, had +it finished. This recovers real information that plain string concatenation +would leave scattered and unattributed: deduped sources, correct citation +numbering, a coherent narrative instead of a raw tool-output dump. + +The summarizer call is itself just another model call that can fail or +have nothing to work with β€” so it falls back to the programmatic extraction +on an empty transcript or a failed call, rather than letting a second +failure take down the one recovery path meant to be bulletproof. And +regardless of what either path produces, a plain, hard-coded disclaimer is +written directly to the output β€” not left to an LLM's discretion to +remember to mention that the research is incomplete. + +This was verified against a real run, not a mocked one: a genuinely +under-specified research question was run end-to-end against local models +until it actually exhausted its turn budget, with full tracing enabled, to +confirm the recovery path fired, the summarizer produced a coherent partial +report, and the disclaimer showed up exactly where expected. + +## Failure #2: the recovery path was masking a worse problem + +That live verification run surfaced something the recovery path was built +to survive, but shouldn't have had to: the researcher burned its *entire* +turn budget re-running variations of the same search, and every single one +came back "no results found." The web-search tool was scraping +DuckDuckGo's HTML search page directly β€” no API key required, but no +protection from rate limiting either. And a rate-limited response looks +*identical* to a genuine empty result: the tool has no way to tell the +model "you're being throttled" versus "this topic has no coverage," so the +model just kept trying. + +The recovery mechanism did exactly its job here β€” the run finished with an +honest report saying "found nothing," rather than crashing β€” but that's a +consolation prize. The actual bug was upstream: the search tool's failure +mode was silent and indistinguishable from success. + +Root cause fixed, not the symptom: the search tool now hits a self-hosted +[SearXNG](https://docs.searxng.org/) instance's JSON API instead of +scraping HTML. That's a straight upgrade on every axis that mattered here β€” +a real API instead of parsing markup, results aggregated across multiple +upstream engines instead of hammering one, and full control over request +pacing since it's infrastructure already running on this machine. Before +reaching for a third-party crate, the two SearXNG client crates on +crates.io were checked and rejected: both single-maintainer, both v0.1.0, +neither with any adoption signal β€” a ~20-line `reqwest` + `serde` call +using dependencies already in the tree was the better bet for something +this small. + +**Verification, again against a live run**: the exact same research +question that had previously burned its full turn budget on empty results +was re-run, unchanged, against the new search backend. It completed +normally on the second research round β€” approved by the reviewer, no +max-turns event, no recovery path needed. The fix wasn't just plausible on +paper; it was confirmed to actually change the outcome of the failure it +was meant to fix. + +## Hardening: separation of concerns + +With both failure modes fixed, the module that had accumulated all of this +logic β€” CLI parsing, orchestration, the researcher phase, chat-history +reconstruction, the summarizer, and the writer β€” had grown into a single +600+-line file mixing six unrelated concerns, while the rest of the +codebase (`review.rs`, `tools.rs`, `stream.rs`, `progress.rs`) already kept +one file per concern. That file was split to match: each agent phase, the +orchestration loop, and the pure history-parsing helpers now live in their +own module, with `main.rs` reduced to argument parsing, logging setup, and +a single top-level call β€” see the [project layout](../README.md#project-layout) +in the README for the resulting structure. Unit tests moved with the code +they test (Rust convention keeps tests co-located, not split into separate +files), and the full suite β€” now covering the history-reconstruction logic +in isolation, independent of any model or network call β€” stayed green +throughout. + +## A clean run, for reference + +With both failures fixed, a normal research pass looks like this β€” the +researcher gathered evidence, the reviewer approved it on the first pass, +and the writer produced a cited report, no recovery path exercised: + +> **Topic:** trends in AI customer-support chatbots + +``` +# Trends in AI Customer-Support Chatbots (2026) + +This report synthesizes current research notes detailing the major +technological, operational, and regulatory trends shaping the AI +customer-support chatbot market as of 2026. + +### The Shift to Agentic AI +The industry is undergoing a fundamental transition from simple, scripted +chatbots toward "Agentic AI." This new paradigm focuses on autonomous AI +agents capable of performing complex tasks and orchestrating complete +workflows, moving beyond basic prompt responses [1, 4]... + +### Governance, Security, and Regulation +...compliance with regulations such as GDPR and the EU AI Act is a primary +concern, demanding robust structures for security, transparency, and +governance [4, 6]... + +### Operational and Economic Impact +The overall AI-powered customer service market is projected to reach +$15.12 billion in 2026 [1]. Adoption is widespread, with approximately 72% +of businesses across various industries having deployed AI-driven chatbots +for customer interactions [2]... + +### Sources +[1] https://chatmaxima.com/blog/ai-customer-support-statistics-2026/ +[2] https://www.itransition.com/ai/conversational +[3] https://salt.security/eu-ai-act-compliance +... +``` + +(Abbreviated here; the tool prints the full report, headings, open +questions, and complete source list to the terminal as it streams.) + +## What this demonstrates + +- Designing a multi-agent pipeline as explicit, single-responsibility + stages coordinated by plain control flow, not a single sprawling prompt + or an opaque framework agent-graph. +- Treating an agent's failure modes (turn-budget exhaustion, a tool's + silent degradation) as defects to root-cause and fix, not edge cases to + shrug off β€” including building a second, self-limiting agent as the + recovery mechanism itself. +- Verifying fixes against live runs of the actual failing case, not just + unit tests in isolation. +- A local-first architecture (Ollama + self-hosted SearXNG) with no + cloud LLM API key and no query leaving the host β€” the same constraint + a privacy-sensitive deployment would need. +- Full observability via structured tracing on every agent phase, and a + codebase organized so each concern has exactly one home.