Last refreshed 2026-07-15. Next refresh: weekly.
Why use Inkling on Baseten API?
Baseten API offers Inkling with pay-as-you-go pricing at $1.00/1M input tokens. Baseten is an AI infrastructure platform that provides comprehensive tools for deploying and serving machine learning models efficiently and cost-effectively.
Compare Inkling across 2 providers to find the best fit for your use caseInput / 1M
$1.00
Output / 1M
$4.05
Cache
read $0.17
Batch
Not sourced
Setup recipe
Docs fallbackInstall
Use the provider REST API or SDKAuth
Create a provider API keyCall
model: thinkingmachines/inklingModel ID
thinkingmachines/inklingRequest example
Curated snippets for this provider are not sourced yet. Use Baseten API documentation with model ID
thinkingmachines/inkling.Gotchas
- Use provider model ID "thinkingmachines/inkling", not the LLMReference slug "inkling".
Compare Inkling Across Providers
| Provider | Input (per 1M) | Output (per 1M) |
|---|---|---|
| Tinker | $1.87 | $4.68 |
| Baseten API | $1.00 | $4.05 |
Pricing
| Type | Price (per 1M) |
|---|---|
| Input tokens | $1.00 |
| Output tokens | $4.05 |
Capabilities
VisionMultimodalReasoningJSON / Tool useAudioFine-tuning
About Inkling
Inkling is Thinking Machines Lab's Apache-2.0 open-weight general-purpose multimodal model. It accepts text, image, and audio inputs and generates text with a 1M-token context window. Compare it for Coding, RAG, Agents, Long context, Vision, and JSON / Tool use.
Get Started
Model Specs
Released2026-07-15
Parameters975B total, 41B active
Context1m
ArchitectureMixture of Experts