spike: confirm full pipeline works with shieldstral on CPU (-ngl 0)

The earlier VRAM-exhaustion failure was llama-server defaulting to full
GPU offload for Shieldstral while Ollama's gemma was also resident.
Restarting llama-server with -ngl 0 frees the GPU for gemma and the
full writer->judge ai-agents pipeline runs end to end with no VRAM
contention. Not an ai-agents limitation -- it never touches GPU
placement, only HTTP.
This commit is contained in:
Austin Schaefer 2026-08-06 10:15:16 +02:00
parent c5c25c8b49
commit 4fbbe32d9d

View file

@ -16,14 +16,20 @@
## What it doesn't prove (real limitations found)
1. **VRAM ceiling, not a framework bug.** Running the full pipeline in one
process needs Ollama's gemma model and Shieldstral's llama-server (32k
ctx, ~5.6GB) resident at once. On this 8GB card that overflows CUDA
("out of memory" from Ollama's own `/api/chat`, not from `ai-agents`).
The writer and judge legs each work fine in isolation. This constraint is
identical for the existing rig-core code — nothing here is
`ai-agents`-specific — but it means an actual migration would need to
confirm the current app doesn't already skirt this same ceiling.
1. ~~VRAM ceiling~~ **Resolved — not a framework limitation at all.** The
initial failure (`cudaMalloc failed: out of memory` from Ollama's
`/api/chat`) was from running Shieldstral's `llama-server` at its
default `-ngl 999` (full GPU offload) alongside Ollama's own gemma
load — both fighting for the same 8GB card. Restarting `llama-server`
with `-ngl 0` (CPU-only, same flag this project already documents
using for Shieldstral) puts it entirely on CPU, and the full two-stage
`pipeline:` (writer -> judge) then runs cleanly end to end in one
process — GPU usage stayed flat at ~6GB (all Ollama) throughout. This
is a `llama-server` launch flag, not anything `ai-agents`-specific;
`ai-agents` never touches GPU/CPU placement itself, it only talks HTTP
to whatever backend is configured. `src/server.rs`'s `ensure_running()`
doesn't currently pass `-ngl`, so it would need an `-ngl 0` addition
(mirroring `server.toml`) to get the same behavior in the real app.
2. **No logprob-based scoring.** `revise.rs`'s real `score()` function reads
token logprobs off Shieldstral's response (see `models::ChatLogprobs`) to
get a continuous 0.0-1.0 score, not a yes/no string. `ai-agents`'
@ -44,21 +50,23 @@
The declarative-YAML story checks out for *provider wiring and prompt
config* — that part is genuinely config, not code, and matches the
CrewAI-style ergonomics from the earlier conversation. It does **not**
cleanly cover this project's actual judge mechanism (logprob scoring), so
adopting it wholesale would be a partial rewrite of `revise.rs`'s scoring
logic, not a drop-in replacement. Worth revisiting if a future judge model
switches to yes/no-only verdicts, or if `ai-agents` grows raw-logprob
access.
CrewAI-style ergonomics from the earlier conversation, and the full
two-stage pipeline now runs end to end against this project's real local
models (see reproducing steps below). It does **not** cleanly cover this
project's actual judge mechanism (logprob scoring), so adopting it
wholesale would be a partial rewrite of `revise.rs`'s scoring logic, not a
drop-in replacement. Worth revisiting if a future judge model switches to
yes/no-only verdicts, or if `ai-agents` grows raw-logprob access.
## Reproducing
```
ollama serve # writer leg
# and/or
ollama serve # writer leg (gemma)
/home/austin/.local/share/llama.cpp/build/bin/llama-server \
-m /home/austin/ai/Shieldstral-1.0-3B-BF16.gguf --jinja -c 32768 \
--host 127.0.0.1 --port 8000 # judge leg
--host 127.0.0.1 --port 8000 -ngl 0 # judge leg, CPU-only so it
# doesn't fight gemma for VRAM
cargo run --bin ai_agents_spike
```