LLM Reference
Hugging Face Inference Endpoints

Using Gemma 4 12B IT on Hugging Face Inference Endpoints

Implementation guide · Gemma 4 · Google DeepMind

Open Source

Hugging Face Inference Endpoints exposes Gemma 4 12B IT through model ID google/gemma-4-12B-it. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-06-10. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at Hugging Face Inference Endpoints and generate an API key.
  2. 2
    Use the Hugging Face Inference Endpoints SDK or REST API to call google/gemma-4-12B-it — see the documentation for request format.

Code Examples

Pricing on Hugging Face Inference Endpoints

Capabilities

VisionMultimodalReasoningJSON / Tool useStructured OutputsAudioFine-tuning

About Gemma 4 12B IT

Instruction-tuned version of Gemma 4 12B. Open weight (Apache 2.0), 12B parameters, encoder-free multimodal (text, image, audio). Optimized for chat and instruction-following. Runs on a 16GB laptop.

Model Specs

Released2026-06-03
Parameters12B
Context256k
ArchitectureDecoder Only
Knowledge cutoff2025-01

More Models on Hugging Face Inference Endpoints

Provider

Hugging Face Inference Endpoints
Hugging Face Inference Endpoints

Hugging Face

New York City, New York, United States