Last refreshed 2026-06-11. Next refresh: weekly.
Why use Gemma 4 12B on Hugging Face Inference Endpoints?
Hugging Face Inference Endpoints offers Gemma 4 12B with competitive pricing. Hugging Face is a leading AI community and platform dedicated to democratizing artificial intelligence.
Compare Gemma 4 12B across 2 providers to find the best fit for your use caseInput / 1M
-
Output / 1M
-
Cache
Not sourced
Batch
Not sourced
Setup recipe
Docs fallbackInstall
Use the provider REST API or SDKAuth
Create a provider API keyCall
model: google/gemma-4-12BModel ID
google/gemma-4-12BRequest example
Curated snippets for this provider are not sourced yet. Use Hugging Face Inference Endpoints documentation with model ID
google/gemma-4-12B.Gotchas
- Use provider model ID "google/gemma-4-12B", not the LLMReference slug "gemma-4-12b".
Compare Gemma 4 12B Across Providers
| Provider | Input (per 1M) | Output (per 1M) |
|---|---|---|
| Hugging Face Inference Endpoints | — | — |
| Kaggle Models | — | — |
Capabilities
VisionMultimodalReasoningJSON / Tool useStructured OutputsAudio
About Gemma 4 12B
Google DeepMind's 12B open-weight multimodal model (Apache 2.0), designed to run on a 16GB laptop. First medium-sized model with native audio ingestion alongside text and image. Unified encoder-free decoder-only architecture. Supports 140+ languages. MMLU Pro: 77.2%.
Get Started
Model Specs
Released2026-06-03
Parameters12B
Context256k
ArchitectureDecoder Only
Knowledge cutoff2025-01