diff --git a/spike/NOTES.md b/spike/NOTES.md index 482dfbf..7966eb7 100644 --- a/spike/NOTES.md +++ b/spike/NOTES.md @@ -16,14 +16,20 @@ ## What it doesn't prove (real limitations found) -1. **VRAM ceiling, not a framework bug.** Running the full pipeline in one - process needs Ollama's gemma model and Shieldstral's llama-server (32k - ctx, ~5.6GB) resident at once. On this 8GB card that overflows CUDA - ("out of memory" from Ollama's own `/api/chat`, not from `ai-agents`). - The writer and judge legs each work fine in isolation. This constraint is - identical for the existing rig-core code — nothing here is - `ai-agents`-specific — but it means an actual migration would need to - confirm the current app doesn't already skirt this same ceiling. +1. ~~VRAM ceiling~~ **Resolved — not a framework limitation at all.** The + initial failure (`cudaMalloc failed: out of memory` from Ollama's + `/api/chat`) was from running Shieldstral's `llama-server` at its + default `-ngl 999` (full GPU offload) alongside Ollama's own gemma + load — both fighting for the same 8GB card. Restarting `llama-server` + with `-ngl 0` (CPU-only, same flag this project already documents + using for Shieldstral) puts it entirely on CPU, and the full two-stage + `pipeline:` (writer -> judge) then runs cleanly end to end in one + process — GPU usage stayed flat at ~6GB (all Ollama) throughout. This + is a `llama-server` launch flag, not anything `ai-agents`-specific; + `ai-agents` never touches GPU/CPU placement itself, it only talks HTTP + to whatever backend is configured. `src/server.rs`'s `ensure_running()` + doesn't currently pass `-ngl`, so it would need an `-ngl 0` addition + (mirroring `server.toml`) to get the same behavior in the real app. 2. **No logprob-based scoring.** `revise.rs`'s real `score()` function reads token logprobs off Shieldstral's response (see `models::ChatLogprobs`) to get a continuous 0.0-1.0 score, not a yes/no string. `ai-agents`' @@ -44,21 +50,23 @@ The declarative-YAML story checks out for *provider wiring and prompt config* — that part is genuinely config, not code, and matches the -CrewAI-style ergonomics from the earlier conversation. It does **not** -cleanly cover this project's actual judge mechanism (logprob scoring), so -adopting it wholesale would be a partial rewrite of `revise.rs`'s scoring -logic, not a drop-in replacement. Worth revisiting if a future judge model -switches to yes/no-only verdicts, or if `ai-agents` grows raw-logprob -access. +CrewAI-style ergonomics from the earlier conversation, and the full +two-stage pipeline now runs end to end against this project's real local +models (see reproducing steps below). It does **not** cleanly cover this +project's actual judge mechanism (logprob scoring), so adopting it +wholesale would be a partial rewrite of `revise.rs`'s scoring logic, not a +drop-in replacement. Worth revisiting if a future judge model switches to +yes/no-only verdicts, or if `ai-agents` grows raw-logprob access. ## Reproducing ``` -ollama serve # writer leg -# and/or +ollama serve # writer leg (gemma) + /home/austin/.local/share/llama.cpp/build/bin/llama-server \ -m /home/austin/ai/Shieldstral-1.0-3B-BF16.gguf --jinja -c 32768 \ - --host 127.0.0.1 --port 8000 # judge leg + --host 127.0.0.1 --port 8000 -ngl 0 # judge leg, CPU-only so it + # doesn't fight gemma for VRAM cargo run --bin ai_agents_spike ```