| P0 |
Current Qwen3.6-35B-A3B APEX/MTP Local incumbent |
Already runs on the R9700 path. Keep it as the baseline, but still measure its real max context and MTP overhead. |
coding agentic chat |
Existing GGUF; MTP-on and matched MTP-off profiles. Run PERF-005 first. |
known load |
Local server |
| P1 |
Qwen3.6-27B Best 32GB Qwen target |
Latest dense Qwen3.6 target I verified. Stronger reported coding-agent numbers than the 35B-A3B row and a much more realistic 32GB fit. |
coding agentic reasoning |
Q4_K_M first, then Q5_K_M if headroom exists. MTP-off and MTP-on. |
likely 32GB |
HF |
| P1 |
Qwen3.6-35B-A3B Current Qwen MoE/generalist |
Still important because your existing data is Qwen-shaped and the model card calls out native MTP and 262K context. |
agentic coding long context |
IQ4_XS or Q4_K_M. Start 32K, then discover max context; MTP-on may lower ceiling. |
fit-gated |
HF |
| P1 |
OpenAI gpt-oss-20b Open-weight reasoning target |
OpenAI's practical local gpt-oss model: 21B total, 3.6B active MoE, Apache 2.0, 128K context, and a strong fit for coding, agentic, and chatbot comparisons. |
reasoning coding agentic |
Unsloth GGUF Q8_0 first, Q4_K_M as fallback/comparison. Treat MTP as N/A unless a compatible draft profile appears. |
32GB ready |
OpenAI |
| P1 |
Gemma 4 31B IT Largest practical Gemma |
Current dense Gemma flagship for chat, creative writing, and reasoning. It is near the upper edge for 32GB, so context must be discovered. |
reasoning chat creative |
Q4_K_M or official QAT quant. Start 16K/32K; add MTP drafter only after base fit is known. |
fit-gated |
Google |
| P1 |
Gemma 4 26B-A4B IT Gemma MoE efficiency candidate |
Worth testing only if the quantized total-weight footprint behaves in 32GB. Do not assume active-parameter savings solve memory. |
chat creative RAG |
Smallest official/QAT quant first. Base fit gate before MTP or quality tests. |
fit-gated |
Google |
| P1 |
Gemma 4 12B IT Primary everyday chatbot/RAG model |
Current mid-size Gemma, drafter-ready, and likely to give the best quality/speed balance for OpenWebUI and normal RAG. |
chat RAG creative |
Q6_K or Q8/QAT where available. MTP-off and MTP-on. |
32GB ready |
Google |
| P1 |
Gemma 4 E4B Fast small-model control |
Use as the "small but current" control for latency, RAG extraction, and local assistant responsiveness. |
chat RAG reasoning |
Official QAT or Q8_0. Test long context and MTP if drafters exist. |
32GB ready |
Google |
| P1 |
Gemma 4 E2B Ultra-fast edge baseline |
Cheap enough to run high-repeat reliability tests and useful as a drafter or summarizer candidate. |
chat RAG |
Official QAT or Q8_0. Use for throughput, summarization, and draft-model experiments. |
32GB ready |
Google |
| P1 |
Phi-4-reasoning-vision-15B Compact reasoning/VLM |
Current compact Microsoft reasoning target for math, UI/OCR-style tasks, document reasoning, and computer-use-style evaluation. |
reasoning coding computer-use |
Q5_K_M or Q6_K. Force think and force nothink profiles. |
32GB ready |
HF |
| P1 |
Phi-4-mini-reasoning Tiny reasoning control |
3.8B, 128K context, MIT license, and useful for latency-bound reasoning and judge/helper roles. |
reasoning chat RAG |
Q8_0 or Q6_K. Run as assistant, judge, and summarizer candidate. |
32GB ready |
HF |
| P2 |
Qwen3-VL-30B-A3B Document/vision option |
Use only if we want multimodal/document screenshots in the campaign. Text-only fit should be tested separately from image workloads. |
document RAG chat reasoning |
Q4_K_M/IQ4 first. Measure text-only context, then image-token overhead. |
fit-gated |
Paper |
| P2 |
Qwen3-Coder-30B-A3B-Instruct 32GB coder fallback |
Not the newest coder, but the newer Qwen3-Coder-Next 80B total model is not a sensible 32GB all-VRAM target. Keep this as the practical coder comparison. |
coding agentic |
Q4_K_M or IQ4_XS. Run only after Qwen3.6-27B and 35B-A3B. |
legacy fallback |
HF |
| P2 |
Command R7B RAG/tool-use specialist |
Small enough for 32GB and worth testing if GGUF packaging is clean. Command A is not a 32GB target. |
RAG agentic chat |
Q6_K or Q8_0 if available; compare to Granite and Gemma 4 12B for RAG. |
32GB ready |
Paper |
| P1 |
Granite 3.3 8B Instruct Enterprise/RAG control |
8B, 128K context, Apache 2.0, and explicitly aimed at summarization, extraction, RAG, code, and function calling. |
RAG chat tool use |
Q8_0 or Q6_K. Use as reliable/fast RAG and extraction baseline. |
32GB ready |
HF |
| Sidecar |
Granite Embedding Multilingual R2 RAG retrieval sidecar |
Not a chat model, but RAG quality depends on retrieval. Use it to benchmark dense retrieval and code/document search. |
RAG code retrieval enterprise |
Evaluate retrieval precision/recall separately from generation models. |
32GB ready |
Paper |