Baseten API
Researched 49d agoInference PlatformTier 2Baseten
Baseten API offers 15 tracked models (1 with output token pricing). This catalog covers coding, rag, and agents; open any model detail page for benchmarks, batch tiers, and migration prompts.
Covers 7 workload areas across 15 tracked models; last verified 2026-07-15.
Use it for
- Teams comparing token and batch pricing across this provider's models
- Operators routing coding, rag, and agents workloads through this API
Do not use it for
- Final benchmark picks without opening the relevant model detail page
Tracked models
15
Models available through this provider
Priced output routes
1
Models with output token pricing tracked
Cheapest output
$4.05
Inkling on this route
Batch-ready models
0
No batch pricing tracked
Latest model release
2026-07-15
49d since newest release
Freshness
2026-07-15
Researched 49d ago
Information
Baseten is an AI infrastructure platform that provides comprehensive tools for deploying and serving machine learning models efficiently and cost-effectively. The platform offers: 1. Rapid deployment: Users can deploy models in minutes, avoiding complex processes. 2. Support for open-source models: Baseten allows deployment of best-in-class open-source models. 3. Optimized serving: The platform provides optimized serving for custom models. 4. Scalability: Horizontally scalable services enable smooth transition from prototype to production. 5. High-speed inference: Baseten offers fast inference on infrastructure that automatically scales with traffic. 6. Cost-efficiency: The platform includes a scaled-to-zero feature to optimize costs. 7. Flexible deployment options: Models can be run on Baseten's cloud or the user's infrastructure. Baseten aims to simplify the ML deployment process while ensuring performance, scalability, and cost-efficiency for AI builders and developers.
Where this host wins
- Coding: 6 tracked models with SWE-bench / HumanEval-style scores.
- RAG: 1 tracked model with ruler / needle retrieval benchmarks.
- Agentic: 1 tracked model with BFCL, tau-bench, and SWE-bench tool-use coverage.
- Long-context: 2 tracked models with context-token or InfiniteBench-class signal.
Getting started
Platform Overview
The AI platform offers a comprehensive suite of features designed to streamline the development and deployment of machine learning models. At its core, the platform supports open-source models, allowing developers to leverage existing frameworks and tools for their AI applications. This flexibility is coupled with rapid deployment capabilities, enabling organizations to quickly bring their models into production environments. The platform's architecture is built for scalability, accommodating fluctuating workloads and user demands without compromising performance. A standout feature is its high-speed inference capabilities, crucial for applications that require real-time data processing and decision-making.
Available Models(15)
View all →| Model | Input (per 1M) | Output (per 1M) | Type |
|---|---|---|---|
| Inkling | $1.00 | $4.05 | ServerlessProvisioned |
| Phi-3 Mini 128K | Serverless | ||
| Phi-3 Mini 4k | Serverless | ||
| Llama 3 70B Instruct | Serverless | ||
| Llama 3 8B Instruct | Serverless | ||
| Mixtral 8x22B v0.1 | Serverless | ||
| NSQL 350M | Serverless | ||
| Mixtral 8x7B | Serverless | ||
| Zephyr 7B Alpha | Serverless | ||
| Mistral 7B v0.1 | Serverless |