doubleo7/spike/NOTES.md
Austin Schaefer f47c30c9fb spike: attempt retry loop via states/transitions, document why it doesn't work
Tried reproducing MAX_GENERATION_RETRIES (critic rejects -> retry writer
-> re-check -> judge) using ai-agents' states:/transitions: instead of
pipeline:, since pipeline stages can't branch.

Two real blockers found and confirmed against the crate source:
- delegate: states have no per-turn input override (only pipeline:/
  concurrent: stages get input: templates), so the critic just echoed
  the document back instead of answering yes/no.
- extract: context extractors read the state's incoming user_message,
  not its generated/delegated response (runtime.rs:7320), so a guard
  meant to gate on "what the critic just said" has nothing real to
  read -- the loop-back transition never fires.

Documented as a genuine finding in NOTES.md rather than forcing a
fragile demo: this crate's state machine is built for turn-based intent
routing, not gating on a sub-agent's structured verdict.
2026-08-06 10:26:20 +02:00

7.4 KiB

Spike: ai-agents (declarative YAML) vs. hand-wired rig-core

What this proves

  • ai-agents (crates.io ai-agents v1.0.0, Rust-native, no Python) can load an agent purely from YAML and talk to both of this project's real backends:
    • spike/agents/writer.yaml -> Ollama (gemma4-e4b:latest), matches wire_gemma_client in src/main.rs.
    • spike/agents/critic.yaml -> Ollama (gemma4-e4b:latest), matches the self-review/quality-guard call in revise.rs's is_usable() — same provider as the writer, but its own system prompt and prompt template lifted verbatim from prompts.toml's [critic] section.
    • spike/agents/judge.yaml -> provider: openai-compatible against the local llama-server (Shieldstral), matches wire_shieldstral.
    • All three worked verbatim against this machine's real models/config, no mocking.
  • spike/pipeline.yaml expresses the three-stage generate -> critic -> judge flow (this project's revise::generate_below_threshold shape, including the format quality-guard before scoring) as one declarative pipeline: block with spawner.auto_spawn and {{ stages.<id> }} templating, no manual Rust orchestration code. Confirmed working end to end: writer runs, critic judges its format (yes/no), judge scores it independently — all three stages complete in a single agent.chat() call.

What it doesn't prove (real limitations found)

  1. VRAM ceiling Resolved — not a framework limitation at all. The initial failure (cudaMalloc failed: out of memory from Ollama's /api/chat) was from running Shieldstral's llama-server at its default -ngl 999 (full GPU offload) alongside Ollama's own gemma load — both fighting for the same 8GB card. Restarting llama-server with -ngl 0 (CPU-only, same flag this project already documents using for Shieldstral) puts it entirely on CPU, and the full three-stage pipeline: (writer -> critic -> judge) then runs cleanly end to end in one process — GPU usage stayed flat at ~6GB (all Ollama) throughout. This is a llama-server launch flag, not anything ai-agents-specific; ai-agents never touches GPU/CPU placement itself, it only talks HTTP to whatever backend is configured. src/server.rs's ensure_running() doesn't currently pass -ngl, so it would need an -ngl 0 addition (mirroring server.toml) to get the same behavior in the real app.

  2. No logprob-based scoring. revise.rs's real score() function reads token logprobs off Shieldstral's response (see models::ChatLogprobs) to get a continuous 0.0-1.0 score, not a yes/no string. ai-agents' Agent::chat() returns plain text content; there's no exposed hook for raw logprobs in the YAML/builder API surface I found. Reproducing the current scoring behavior would mean dropping to ai-agents' lower-level provider access (if any) or keeping rig-core for the judge call and only using ai-agents for orchestration/prompt config — a hybrid, not a clean swap.

  3. Critic stage runs but doesn't gate anything in pipeline:, and states:/transitions: doesn't cleanly fix that either. spike/agents/critic.yaml reproduces the self-review prompt from is_usable(), and pipeline: calls it after the writer — but its yes/no verdict is just an extra text output; pipeline: stages are linear/fire-and-forget, no branching back to retry the writer the way MAX_GENERATION_RETRIES does.

    Tried building the retry loop with states:/transitions: instead (spike/state_machine.yaml, src/bin/ai_agents_spike_states.rs): write -> critique -> (loop to write on "no", or advance to judge on "yes"), using a guard: expression on extracted context. Two real obstacles surfaced, both confirmed against the crate source (ai-agents-runtime-1.0.0, ai-agents-state-1.0.0) and by running it:

    • delegate: states have no per-turn input override. Only pipeline:/concurrent: stages get an input: Jinja template (ai-agents-state-1.0.0/src/config.rs PipelineStageEntry::Config). A bare delegate: critic state just forwards the parent conversation history, with no way to inject the "judge only the format" instruction — in practice the critic just echoed the writer's document back verbatim instead of answering yes/no. Switching the critique/judge states to single-stage pipeline: blocks (which do support input:) fixed this.
    • extract: context extractors read the state's incoming user_message, not its generated response. run_context_extractors_staged (ai-agents-runtime-1.0.0/src/runtime.rs:7320) builds its extraction prompt from user_message — the turn's input — never the assistant's (or a delegated/piped agent's) output. So an extract: block meant to capture "what the critic just answered" has nothing real to read; confirmed by RUST_LOG=debug showing no extraction activity at all around the state transition, and by the context.usable-gated transition to judge never firing even when the critic's actual answer was yes. The machine just stopped after write -> critique and returned the critic's raw response as chat()'s final output — i.e. the loop never proved out.
    • Separately, each agent.chat() call only appeared to advance one state transition (depth=1 in the logs) before returning, so even with working guards, driving the machine to a judge terminal state might require the caller to loop calling chat() per hop rather than getting one resolved answer per call the way pipeline: does.

    Bottom line: reproducing MAX_GENERATION_RETRIES declaratively isn't a matter of swapping pipeline: for states:/transitions: — the state-machine primitives here are built for turn-based conversational branching (routing user intent to sub-flows), not for gating on a sub-agent's structured verdict about text it just produced. That would need either a custom tool/hook that calls back into Rust to inspect stage output and decide the transition, or keeping the retry loop in hand-written Rust (as revise.rs already does) and using ai-agents only for the linear leg of the flow.

Verdict

The declarative-YAML story checks out for provider wiring and prompt config — that part is genuinely config, not code, and matches the CrewAI-style ergonomics from the earlier conversation, and the full two-stage pipeline now runs end to end against this project's real local models (see reproducing steps below). It does not cleanly cover this project's actual judge mechanism (logprob scoring), so adopting it wholesale would be a partial rewrite of revise.rs's scoring logic, not a drop-in replacement. Worth revisiting if a future judge model switches to yes/no-only verdicts, or if ai-agents grows raw-logprob access.

Reproducing

ollama serve                      # writer leg (gemma)

/home/austin/.local/share/llama.cpp/build/bin/llama-server \
  -m /home/austin/ai/Shieldstral-1.0-3B-BF16.gguf --jinja -c 32768 \
  --host 127.0.0.1 --port 8000 -ngl 0   # judge leg, CPU-only so it
                                         # doesn't fight gemma for VRAM

cargo run --bin ai_agents_spike           # linear pipeline: writer -> critic -> judge
cargo run --bin ai_agents_spike_states    # states:/transitions: retry-loop attempt (see limitation #3 above)