Using Gemma 7B Instruct on Lepton AI API

Implementation guide · Gemma · Google DeepMind

ServerlessOpen Weights

Quick Start

1
Create an account at Lepton AI API and generate an API key.
2
Use the Lepton AI API SDK or REST API to call gemma-7b-it — see the documentation for request format.
3
You'll be billed $0.07/1M input, $0.07/1M output tokens. See full pricing.

API Portal Documentation Pricing

Code Examples

See Lepton AI API documentation for integration details.

About Lepton AI API

Lepton AI is a comprehensive cloud-native platform designed to simplify the development and deployment of AI applications. It offers a user-friendly interface that allows developers to build models natively in Python, eliminating the need for complex containerization or Kubernetes expertise. The platform supports local debugging, enabling users to test their models before deployment with a simple command. With a flexible API for easy integration into various applications and support for heterogeneous hardware, Lepton AI optimizes performance based on specific application needs. This flexibility allows for efficient scaling, accommodating workloads that can expand up to 1TB of memory. The platform provides a robust set of tools and infrastructure to enhance AI workflows. Its cloud-native architecture supports high-performance computing, featuring smart scheduling and dynamic batching to minimize downtime. Lepton AI enables continuous deployment through GitHub integration, facilitating rapid iteration and scaling of AI applications. The platform also includes built-in monitoring, logging, and autoscaling capabilities, ensuring that applications remain responsive and efficient in production environments. With these features, Lepton AI streamlines the entire AI development process, from model creation to deployment and maintenance, making it accessible for organizations of various sizes looking to innovate with AI technologies.

Lepton AI is building a scalable and efficient AI Application platform. Their platform aims to simplify the development and deployment of AI applications, making it easier for businesses to leverage artificial intelligence technologies. The company focuses on providing tools and infrastructure to streamline AI workflows, enabling faster development cycles and more efficient resource utilization. While specific details about their platform's features are not provided in the context, Lepton AI's mission is to make AI application development more accessible and efficient for developers and businesses alike.

View all models on Lepton AI API →

Pricing on Lepton AI API

Type	Price (per 1M)
Input tokens	$0.07
Output tokens	$0.07

Capabilities

Structured Outputs

About Gemma 7B Instruct

Gemma 7B Instruct is a cutting-edge large language model developed by Google DeepMind, boasting 7 billion parameters. As part of the Gemma family, it benefits from the advanced research underpinning Google's Gemini models. This model is optimized for text generation tasks, excelling in areas like question answering and summarization, and it is finely tuned to follow instructions effectively. Despite its compact size, Gemma 7B Instruct performs impressively on benchmarks, making it versatile for deployment across various hardware platforms, from laptops to cloud infrastructure. Moreover, it is open-source, with accessible weights and incorporates responsible AI practices, such as data filtering and human feedback, to ensure safe and ethical use.

Full model details →

Model Specs

Released2024-02-21

Parameters7B

Context8k

ArchitectureDecoder Only

Knowledge cutoff2023-04

Also available on(7)

Replicate API$0.05/1M GCP Vertex AI$0.10/1M Fireworks AI$0.20/1M

Compare all providers →

Provider

Lepton AI API

Lepton AI

Sacramento, California, United States