LLM Reference
Lepton AI API

Lepton AI API

Researched 93d agoInference PlatformTier 2

Lepton AI

CodingClassificationJSON / Tool useAI

Lepton AI API offers 14 tracked models (14 with output token pricing). This catalog covers coding, classification, and json / tool use; open any model detail page for benchmarks, batch tiers, and migration prompts.

Covers 3 workload areas across 14 tracked models; last verified 2026-06-01.

Use it for

  • Teams comparing token and batch pricing across this provider's models
  • Operators routing coding, classification, and json / tool use workloads through this API

Do not use it for

  • Final benchmark picks without opening the relevant model detail page

Tracked models

14

Models available through this provider

Priced output routes

14

Models with output token pricing tracked

Cheapest output

$0.070

WizardLM-2 7B on this route

Batch-ready models

0

No batch pricing tracked

Latest model release

2024-04-18

867d since newest release

Freshness

2026-06-01

Researched 93d ago

stale

Information

TypeInference Platform
TierTier 2
Models14
CompanyLepton AI
Founded2023
Sacramento, California, United States

Lepton AI is building a scalable and efficient AI Application platform. Their platform aims to simplify the development and deployment of AI applications, making it easier for businesses to leverage artificial intelligence technologies. The company focuses on providing tools and infrastructure to streamline AI workflows, enabling faster development cycles and more efficient resource utilization. While specific details about their platform's features are not provided in the context, Lepton AI's mission is to make AI application development more accessible and efficient for developers and businesses alike.

Where this host wins

  • Coding: 7 tracked models with SWE-bench / HumanEval-style scores.
  • Classification: 10 tracked models with MMLU-class moderation/safety coverage.
  • JSON/tool-use: 9 tracked models with BFCL / Nexus strict-JSON routing coverage.

Getting started

Verify: quotas and regions in the linked vendor documentation.

Platform Overview

Lepton AI is a comprehensive cloud-native platform designed to simplify the development and deployment of AI applications. It offers a user-friendly interface that allows developers to build models natively in Python, eliminating the need for complex containerization or Kubernetes expertise. The platform supports local debugging, enabling users to test their models before deployment with a simple command. With a flexible API for easy integration into various applications and support for heterogeneous hardware, Lepton AI optimizes performance based on specific application needs. This flexibility allows for efficient scaling, accommodating workloads that can expand up to 1TB of memory.

Available Models(14)

View all →

All models available as Serverless

ModelInput (per 1M)Output (per 1M)
Llama 3 70B Instruct$0.80$0.80
Llama 3 8B Instruct$0.07$0.07
Gemma 7B Instruct$0.07$0.07
WizardLM-2 7B$0.07$0.07
WizardLM-2 8x22B$0.50$0.50
OpenChat 3.5 (0106)$0.07$0.07
Dolphin 2.6 Mixtral 8x7B$0.30$0.30
Nous Hermes 13B$0.13$0.13
Mixtral 8x7B$0.30$0.30
MythoMax L2 13B$0.13$0.13
View full catalog →

Where else to run this