Models
Which model, by VRAM
CodeZaiku ships no model. It talks to any server that speaks the OpenAI chat API: llama.cpp, Ollama,
LM Studio, vLLM, or a hosted API with a key. codezaiku model serve install sets one up on this
machine for the card it finds: the row below for the card’s memory, downloaded once and checked against a
recorded sha256, behind a small proxy that starts the model when something asks and stops it after twenty
idle minutes. Linux uses llama.cpp’s CUDA container, macOS its Metal build, Windows its Vulkan build.
VRAM means the memory on your graphics card, not the computer’s RAM. nvidia-smi shows it on
Linux and Windows. On a Mac with Apple silicon the card shares the machine’s unified memory, so read the rows
against about two thirds of that.
The rows come from ResearchZosho’s measurement of eleven models on the same research questions and from CodeZaiku’s own coding probes on the two it drives with. A row marked research was measured reading and writing, not coding; it calls tools cleanly, which is what the harness needs first.
| VRAM | Model | File | What we saw |
|---|---|---|---|
| 24 GB or more | Qwen3.8-27B at 4-bit | Qwen3.8-27B-UD-Q4_K_M.gguf | the reference drive: 93% of tool-call probes acted on the right tool on time, the only local model green across the coding reverse-eval; 17 GB file, 3.6 s median turn |
| 16 GB | gpt-oss-20b | gpt-oss-20b-F16.gguf | research: right, fast (about 5 minutes a question), 13 GB in use with two 16k slots |
| 8 GB | Qwen3.5 9B at 4-bit | Qwen3.5-9B-Q4_K_M.gguf | the drive behind nearly every number in LIMITATIONS.md; calls tools cleanly; multi-file coding ceiling about 45% |
| 4 GB | Gemma 4 E4B at 4-bit | gemma-4-E4B-it-Q4_K_M.gguf | research: right, the best citation reader of the small models; 3.6 GB in use |
| 2 GB | Gemma 4 E2B at 4-bit | gemma-4-E2B-it-Q4_K_M.gguf | research: facts right, write-ups thin. A hosted API is the better answer this small |
The files are the ones Hugging Face lists under unsloth/<model>-GGUF. The full research
list, with second and third choices per tier and the models measured and not recommended, is on
ResearchZosho’s models page.
What the harness needs from a model
- Tool calling. The whole harness is a tool-calling loop. With llama.cpp that means
--jinja, so the server applies the model’s own chat template; without it tool calls come back as prose and nothing works.codezaiku smokeis the check. - Context. 8k for interactive use, 12k to be driven by another agent, 32k if you have it.
codezaiku doctorreports the window it detects. - Instruction-following on long, structured prompts. Size matters less than these: a 9B that calls tools cleanly beats a larger model that does not.
A bigger label is not more capability for a given job: a 30B-class coder that is a 3B-active mixture did no better than the 9B on knowledge-heavy work. The 27B dense model is the reference because it acts on tools reliably. Hosted APIs, two models at once, and embeddings are in MODELS.md.