LLM Reference

Best LLMs for Customer Support (2026)

Last refreshed 2026-09-01. Next refresh: weekly.

Function-calling models for support bots, ranked by tau-bench service-task performance with BFCL fallback and a $25 per 1k conversation cost gate.

Cost gate. The list excludes models above $25.00 per 1k support conversations, using the cheapest public provider route and a 4k-input / 1k-output average turn over five turns.

Verdict

Use ByteDance Doubao Seed 2.0 Pro for support automation today.

LFM2-24B-A2B is the runner-up; compare τ-bench against Release.

Researched 54d agoWhy this pickMethodology

How we rank

Support bots prioritize τ-bench multiturn service scores, with BFCL fallback only when τ-bench is unavailable, after a cost gate removes high-throughput options.

  1. EligibilityModels with `function_calling` enabled, public list token pricing, and a cheapest public provider route under the support workload cost gate.
  2. Primary rankingτ-bench is the primary score; when retail and airline splits are both present, we average them. If τ-bench is unavailable, the row falls back to BFCL.
  3. Cost gateRows above $25.00 per 1k conversations are excluded using the cheapest public provider route. The category-local workload is 4k input + 1k output tokens per turn across five turns; it does not change the global cost calculator.
  4. Tie-breaksWhen primary scores match, lower blended support cost wins, then confirmed `function_calling = 1`, then newer `release`.
  5. Variant collapseWe keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell. This page also requires a public token-priced route that can be evaluated by the support cost gate.
  6. PricingSupport workloads are throughput-sensitive — compare batch/cache columns on provider pages.
#ModelInput $/1MOutput $/1M
1Ring-2.6-1T
ReasoningTools

Signal used: τ-bench 95.32%

$0.07$0.63
2ByteDance Doubao Seed 2.0 Pro
VisionTools

Signal used: τ-bench 90.4%

$0.47$2.37
3Qwen3.5-397B-A17B
ReasoningVisionTools

Signal used: τ-bench 86.7%

$0.39$2.34
4GLM-5
ReasoningTools

Signal used: τ-bench 82.1%

$0.60$2.08
5Qwen3.5-35B-A3B
ReasoningTools

Signal used: τ-bench 81.2%

$0.14$1.00
6Qwen3.5-122B-A10B
ReasoningVisionTools

Signal used: τ-bench 79.5%

$0.26$2.08
7Qwen3.5-9B
VisionTools

Signal used: τ-bench 79.1%

$0.10$0.15
8Qwen3.5-27B
ReasoningVisionTools

Signal used: τ-bench 79%

$0.20$1.56
9Qwen3.6-Plus
VisionTools

Signal used: τ-bench 76.8%

$0.33$1.95
10Kimi K2.5
VisionTools

Signal used: τ-bench 74.2%

$0.44$2.00
11Gemini 3 Flash
PreviewVisionTools

Signal used: τ-bench 71.5%

$0.50$3.00
12Mistral Small 4
VisionTools

Signal used: τ-bench 65.8%

$0.10$0.30
13Gemini 2.5 Flash
VisionTools

Signal used: BFCL 56.24%

$0.30$2.50
14GPT-5 Mini
ReasoningVisionTools

Signal used: BFCL 55.46%

$0.25$2.00
15GPT-4.1 Mini
VisionTools

Signal used: BFCL 50.45%

$0.40$1.60
16Mistral Large 2
VisionTools

Signal used: BFCL 38.37%

$0.48$2.40
17CoBuddy
ReasoningTools

Signal used: Release 2026-05-06

FreeFree
18Gemma 4 E2B
Tools

Signal used: Release 2026-03-31

FreeFree
19Gemma 4 E4B
Tools

Signal used: Release 2026-03-31

FreeFree
20Gemma 4 26B A4B IT
VisionTools

Signal used: Release 2026-03-31

FreeFree

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • Qwen3-Coder-30B-A3B-Instruct is Alibaba's efficient open-source code generation model in the Qwen3-Coder family, released December 3, 2025 under the Apache 2.0 license. The model has 30.5 billion total parameters with 3.3 billion active per forward pass, organized across 48 transformer layers with 128 experts and 8 activated per token. It uses Grouped Query Attention (GQA) with 32 query heads and 4 key-value heads. Native context window is 262,144 tokens, extendable to 1 million tokens via YaRN. The model supports multi-turn tool calling, function calling, repository-level code understanding, and structured outputs. It is compatible with vLLM, SGLang, Ollama, LM Studio, llama.cpp, and HuggingFace Transformers. Available via AWS Bedrock, Novita AI, and Vercel AI Gateway.

    2025-12-03

    Release

  • #5ERNIE 4.5 21B A3B

    ERNIE 4.5 21B A3B is a Mixture-of-Experts model from Baidu with 21B total parameters and 3B activated per token. It delivers strong performance across Chinese and English language tasks with efficient inference at 120K context.

    2026-01-01

    Release

  • #6GPT-5 Nano

    Fastest, cheapest GPT-5 variant for summarization and classification tasks. Also available via Realtime API.

    2025-08-07

    Release