The earlier VRAM-exhaustion failure was llama-server defaulting to full GPU offload for Shieldstral while Ollama's gemma was also resident. Restarting llama-server with -ngl 0 frees the GPU for gemma and the full writer->judge ai-agents pipeline runs end to end with no VRAM contention. Not an ai-agents limitation -- it never touches GPU placement, only HTTP.
3.9 KiB
3.9 KiB
Spike: ai-agents (declarative YAML) vs. hand-wired rig-core
What this proves
ai-agents(crates.ioai-agentsv1.0.0, Rust-native, no Python) can load an agent purely from YAML and talk to both of this project's real backends:spike/agents/writer.yaml-> Ollama (gemma4-e4b:latest), matcheswire_gemma_clientinsrc/main.rs.spike/agents/judge.yaml->provider: openai-compatibleagainst the localllama-server(Shieldstral), matcheswire_shieldstral.- Both worked verbatim against this machine's real models/config, no mocking.
spike/pipeline.yamlexpresses the two-stage generate -> judge flow (this project'srevise::generate_below_thresholdshape) as one declarativepipeline:block withspawner.auto_spawnand{{ stages.<id> }}templating, no manual Rust orchestration code.
What it doesn't prove (real limitations found)
VRAM ceilingResolved — not a framework limitation at all. The initial failure (cudaMalloc failed: out of memoryfrom Ollama's/api/chat) was from running Shieldstral'sllama-serverat its default-ngl 999(full GPU offload) alongside Ollama's own gemma load — both fighting for the same 8GB card. Restartingllama-serverwith-ngl 0(CPU-only, same flag this project already documents using for Shieldstral) puts it entirely on CPU, and the full two-stagepipeline:(writer -> judge) then runs cleanly end to end in one process — GPU usage stayed flat at ~6GB (all Ollama) throughout. This is allama-serverlaunch flag, not anythingai-agents-specific;ai-agentsnever touches GPU/CPU placement itself, it only talks HTTP to whatever backend is configured.src/server.rs'sensure_running()doesn't currently pass-ngl, so it would need an-ngl 0addition (mirroringserver.toml) to get the same behavior in the real app.- No logprob-based scoring.
revise.rs's realscore()function reads token logprobs off Shieldstral's response (seemodels::ChatLogprobs) to get a continuous 0.0-1.0 score, not a yes/no string.ai-agents'Agent::chat()returns plain text content; there's no exposed hook for raw logprobs in the YAML/builder API surface I found. Reproducing the current scoring behavior would mean dropping toai-agents' lower-level provider access (if any) or keeping rig-core for the judge call and only usingai-agentsfor orchestration/prompt config — a hybrid, not a clean swap. - Multi-turn revision loop (
MAX_REVISION_ITERATIONS, feeding the previous score back into the next prompt) isn't attempted here — the pipeline stage in this spike is a single writer -> judge pass, not the full generate/score/revise loop with a threshold-driven exit condition. Thepipeline:construct is one-shot; the retry/threshold loop would likely needstates:/transitions:(state machine) rather thanpipeline:.
Verdict
The declarative-YAML story checks out for provider wiring and prompt
config — that part is genuinely config, not code, and matches the
CrewAI-style ergonomics from the earlier conversation, and the full
two-stage pipeline now runs end to end against this project's real local
models (see reproducing steps below). It does not cleanly cover this
project's actual judge mechanism (logprob scoring), so adopting it
wholesale would be a partial rewrite of revise.rs's scoring logic, not a
drop-in replacement. Worth revisiting if a future judge model switches to
yes/no-only verdicts, or if ai-agents grows raw-logprob access.
Reproducing
ollama serve # writer leg (gemma)
/home/austin/.local/share/llama.cpp/build/bin/llama-server \
-m /home/austin/ai/Shieldstral-1.0-3B-BF16.gguf --jinja -c 32768 \
--host 127.0.0.1 --port 8000 -ngl 0 # judge leg, CPU-only so it
# doesn't fight gemma for VRAM
cargo run --bin ai_agents_spike