Last refreshed 2026-06-10. Next refresh: weekly.
Why use Gemma 4 12B IT on Hugging Face Inference Endpoints?
Hugging Face Inference Endpoints offers Gemma 4 12B IT with competitive pricing. Hugging Face is a leading AI community and platform dedicated to democratizing artificial intelligence.
Compare Gemma 4 12B IT across 2 providers to find the best fit for your use caseInput / 1M
-
Output / 1M
-
Cache
Not sourced
Batch
Not sourced
Setup recipe
Docs fallbackInstall
Use the provider REST API or SDKAuth
Create a provider API keyCall
model: google/gemma-4-12B-itModel ID
google/gemma-4-12B-itRequest example
Curated snippets for this provider are not sourced yet. Use Hugging Face Inference Endpoints documentation with model ID
google/gemma-4-12B-it.Gotchas
- Use provider model ID "google/gemma-4-12B-it", not the LLMReference slug "gemma-4-12b-it".
Compare Gemma 4 12B IT Across Providers
| Provider | Input (per 1M) | Output (per 1M) |
|---|---|---|
| Hugging Face Inference Endpoints | — | — |
| Kaggle Models | — | — |
Capabilities
VisionMultimodalReasoningJSON / Tool useStructured OutputsAudioFine-tuning
About Gemma 4 12B IT
Instruction-tuned version of Gemma 4 12B. Open weight (Apache 2.0), 12B parameters, encoder-free multimodal (text, image, audio). Optimized for chat and instruction-following. Runs on a 16GB laptop.
Get Started
Model Specs
Released2026-06-03
Parameters12B
Context256k
ArchitectureDecoder Only
Knowledge cutoff2025-01