LLM Reference
Hugging Face Inference Endpoints

Using Gemma 4 12B on Hugging Face Inference Endpoints

Implementation guide · Gemma 4 · Google DeepMind

Open Source

Hugging Face Inference Endpoints exposes Gemma 4 12B through model ID google/gemma-4-12B. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-06-11. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at Hugging Face Inference Endpoints and generate an API key.
  2. 2
    Use the Hugging Face Inference Endpoints SDK or REST API to call google/gemma-4-12B — see the documentation for request format.

Code Examples

Pricing on Hugging Face Inference Endpoints

Capabilities

VisionMultimodalReasoningJSON / Tool useStructured OutputsAudio

About Gemma 4 12B

Google DeepMind's 12B open-weight multimodal model (Apache 2.0), designed to run on a 16GB laptop. First medium-sized model with native audio ingestion alongside text and image. Unified encoder-free decoder-only architecture. Supports 140+ languages. MMLU Pro: 77.2%.

Model Specs

Released2026-06-03
Parameters12B
Context256k
ArchitectureDecoder Only
Knowledge cutoff2025-01

More Models on Hugging Face Inference Endpoints

Provider

Hugging Face Inference Endpoints
Hugging Face Inference Endpoints

Hugging Face

New York City, New York, United States