Winnow-12B
Winnow-12B is available now for agents, vision, and json / tool use with open-source; evaluate it while provider pricing coverage matures.
- Family
- Winnow
- Released
- 2026-09-20
- Context
- 66k
- Parameters
- 12B
- Openness
- Open source
- License
- Apache 2.0OSI-approvedCommercial use: permitted
- Weights
- Available
- Code
- Unknown
- Training
- Fine-tuned
No tracked provider token pricing is available yet.
About
Winnow-12B is EldanRing's local decision model, a rank-32 LoRA fine-tune of Gemma 4 12B IT merged into the weights and released as GGUF files (BF16, Q8_0, and NVFP4) with an optional F16 vision projector. It answers typed noul, choice, and score questions against a shared state by prefilling the state once, forking question branches, and reading answer-token logits without generating text. The same loaded model also serves ordinary chat, streaming, and image input. Its llama.cpp-based inference server exposes /v1/systemone and /v1/chat/completions. The Q8_0 file was tested with a 64K context and vision on a 16 GB RTX 5070 Ti.
Provider price ladder
No tracked provider token pricing is available for this model yet.
Capabilities
Benchmark peer barsfor Agents
No task-mapped benchmark peers are available for this model yet.
Benchmark scores(1)
Migration checks
No linked migration route is available for this model yet.
API versions
EldanRing/Winnow-12BWinnow-12BWinnow-12B Q8winnow-12bNo tracked provider token pricing is available yet.