LLM Reference
Baseten API

Inkling on Baseten API

Inkling · Thinking Machines Lab

ServerlessProvisionedOpen Source

Last refreshed 2026-07-15. Next refresh: weekly.

Why use Inkling on Baseten API?

Baseten API offers Inkling with pay-as-you-go pricing at $1.00/1M input tokens. Baseten is an AI infrastructure platform that provides comprehensive tools for deploying and serving machine learning models efficiently and cost-effectively.

Compare Inkling across 2 providers to find the best fit for your use case
Input / 1M
$1.00
Output / 1M
$4.05
Cache
read $0.17
Batch
Not sourced

Setup recipe

Docs fallback
Install
Use the provider REST API or SDK
Auth
Create a provider API key
Call
model: thinkingmachines/inkling
Model ID
thinkingmachines/inkling

Request example

Curated snippets for this provider are not sourced yet. Use Baseten API documentation with model ID thinkingmachines/inkling.

Gotchas

  • Use provider model ID "thinkingmachines/inkling", not the LLMReference slug "inkling".

Compare Inkling Across Providers

ProviderInput (per 1M)Output (per 1M)
Tinker$1.87$4.68
Baseten API$1.00$4.05

Pricing

TypePrice (per 1M)
Input tokens$1.00
Output tokens$4.05

Capabilities

VisionMultimodalReasoningJSON / Tool useAudioFine-tuning

About Inkling

Inkling is Thinking Machines Lab's Apache-2.0 open-weight general-purpose multimodal model. It accepts text, image, and audio inputs and generates text with a 1M-token context window. Compare it for Coding, RAG, Agents, Long context, Vision, and JSON / Tool use.

Get Started

Model Specs

Released2026-07-15
Parameters975B total, 41B active
Context1m
ArchitectureMixture of Experts