Last refreshed 2026-09-23. Next refresh: weekly.
Why use GPT-6 Luna on Azure OpenAI?
Azure OpenAI offers GPT-6 Luna with pay-as-you-go pricing at $0.10/1M input tokens. Azure OpenAI Service hosts OpenAI's GPT-4o, GPT-4, GPT-3.5, and embedding models on Microsoft Azure with enterprise SLAs.
Compare GPT-6 Luna across 4 providers to find the best fit for your use caseSetup recipe
Python + curlpip install openaiexport AZURE_OPENAI_API_KEY=...import os
from openai import AzureOpenAI
client = AzureOpenAI(
azure_endpoint=os.environ["AZURE_OPENAI_ENDPOINT"], # e.g. https://{resource}.openai.azure.comgpt-6-lunaRequest example
import os
from openai import AzureOpenAI
client = AzureOpenAI(
azure_endpoint=os.environ["AZURE_OPENAI_ENDPOINT"], # e.g. https://{resource}.openai.azure.com
api_key=os.environ["AZURE_OPENAI_API_KEY"],
api_version="2024-02-01"
)
response = client.chat.completions.create(
model="gpt-6-luna", # your deployment name
messages=[{"role": "user", "content": "Hello"}]
)
print(response.choices[0].message.content)Gotchas
- gpt-6-luna is your Azure deployment name, not the underlying model name. Deployment names are set when you deploy a model in Azure AI Foundry / Azure OpenAI Studio.
- The examples expect AZURE_OPENAI_API_KEY; rename it only if your application config maps the new variable.
Compare GPT-6 Luna Across Providers
| Provider | Input (per 1M) | Output (per 1M) |
|---|---|---|
| OpenAI API | $0.10 | $0.50 |
| AWS Bedrock | $0.10 | $0.50 |
| Azure OpenAI | $0.10 | $0.50 |
| OpenRouter | $0.10 | $0.50 |
Pricing
| Type | Price (per 1M) |
|---|---|
| Input tokens | $0.10 |
| Output tokens | $0.50 |
Capabilities
About GPT-6 Luna
OpenAI GPT-6 Luna, announced September 22, 2026 with GPT-6 Sol, is the most efficient GPT-6 tier for focused high-volume tasks below Sol and flagship Astra. First-party OpenAI API id gpt-6-luna. Docs: 1,050,000 context, 128,000 max output, knowledge cutoff May 18, 2026, text and image input to text output, reasoning effort none/low/medium/high/xhigh/max, function calling, structured outputs, prompt caching, Batch and Flex at 50% of Standard, Fast mode at 2x, and computer use and code interpreter via Responses tools.