LLM Reference

GPT-4o Audio Models by OpenAI

OpenAIProprietary
2 models2024Up to 128k ctx

Last refreshed 2026-05-19. Next refresh: weekly.

Details

ResearcherOpenAI
LicenseProprietary
Commercial useCommercial use: conditional
Models2
Released2024
Max context128k

Capabilities

VisionAll models
Code ExecutionAll models

Links

Website

About

GPT-4o Audio is a family of 2 AI models by OpenAI, released in 2024.

Decision facts

Best fit
audiovision and multimodal workcode execution
Capability starting point
GPT-4o Audio Preview (12-17) with 128k context and multimodal inputs
Lowest tracked input
Not tracked
Closest related family
GPT Realtime 2

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

2 in view

Use when the workload needs audio, 128k context, and code execution.

2024-12audio128k contextcode execution

Use when the workload needs audio, 128k context, and code execution.

2024-10audio128k contextcode execution

Release Timeline

2 release groups
2024-12
1 current
GPT-4o Audio Preview (12-17)
audio128k contextcode execution
Current
2024-10
1 current
GPT-4o Audio Preview (10-01)
audio128k contextcode execution
Current

Specifications(2 models)

GPT-4o Audio model specifications comparison
ModelReleasedContextVisionCode Exec
GPT-4o Audio Preview (12-17)2024-12128kYesYes
GPT-4o Audio Preview (10-01)2024-10128kYesYes