Using Gemma 2B Instruct on Cloudflare Workers AI

Implementation guide · Gemma · Google DeepMind

ServerlessOpen Weights

Cloudflare Workers AI exposes Gemma 2B Instruct through model ID @cf/google/gemma-2b-it-lora. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-04-24. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at Cloudflare Workers AI and generate an API key.
  2. 2
    Use the Cloudflare Workers AI SDK or REST API to call @cf/google/gemma-2b-it-lora — see the documentation for request format.

Code Examples

See Cloudflare Workers AI documentation for integration details.

Pricing on Cloudflare Workers AI

Capabilities

Structured Outputs

About Gemma 2B Instruct

Gemma 2B Instruct is a large language model developed by Google, designed to balance performance and accessibility with its 2 billion parameters. Derived from the Gemini family, it excels in tasks such as text generation, code interpretation, and mathematical problem-solving. Built on a transformer decoder architecture, it features multi-query attention, RoPE, GeGLU activations, and RMSNorm. Trained on approximately 6 trillion tokens, including web documents, code, and mathematical content, it uses SFT and RLHF for instruction-tuning. Notable for its lightweight design permitting deployment on consumer-grade hardware, it's open-source and optimized for dialogue applications.

Model Specs

Released2024-02-21
Parameters2B
Context2k
ArchitectureDecoder Only
Knowledge cutoff2023-04

Provider

Cloudflare Workers AI

Cloudflare

San Francisco, California, United States