GPT-5.6 Luna
OpenAI's fastest, most cost-efficient GPT-5.6 model — 1M context, vision, tool use, and web search for high-volume agentic work.
Get API key- Provider
- OpenAI
- Price
- $1.25 in · $7.50 out / 1M
- Context window
- 1M tokens
- Released
- July 9, 2026
- License
- Proprietary
What is GPT-5.6 Luna?
GPT-5.6 Luna is OpenAI's efficient, high-volume model in the GPT-5.6 family, released July 9, 2026. It offers a 1M-token context window, vision, function calling, reasoning, and web search, delivering strong coding performance at a fraction of the cost of flagship tiers.
Use it privately on Venice
On Venice, you run GPT-5.6 Luna with zero retention — your prompts are anonymized and not stored, profiled, or used for training. You get the full OpenAI capability stack including vision, tool use, reasoning, and web search, but you pay per token without a subscription and keep your data sovereign.
What can it do?
- •Optimized for speed and throughput, making it ideal for high-volume, latency-sensitive production workflows.
- •Strong coding performance with an 84.3% score on Terminal-Bench 2.1, reportedly beating the mid-tier Terra variant despite lower cost.
- •Full multi-modal and agentic stack on Venice — vision (including multiple image inputs), function calling, reasoning, web search, and structured JSON output.
- •Massive 1M-token context window with up to 128K tokens of output in a single generation.
- •Cost-efficient frontier intelligence — Artificial Analysis places Luna at the Pareto frontier of intelligence versus cost per task.
- •Closed and proprietary — no open weights, so self-hosting or fine-tuning is impossible.
- •Not uncensored — retains OpenAI's safety alignment and refusal patterns.
- •Privacy on Venice is anonymized, but the model does not run in a TEE or with end-to-end encryption.
- •Peak reasoning and intelligence still sit with the flagship GPT-5.6 Sol and Claude Fable 5; Luna trades absolute top-tier benchmark scores for efficiency.
- •The gpt-5.6 API alias routes to Sol by default, so integrations must explicitly target the luna slug.
GPT-5.6 Luna capabilities
- Tool use / function calling
- Vision (image input)
- Reasoning
- Web search
- Code-optimized
- Structured output (JSON schema)
- Audio input
- Video input
- Multiple image inputs
- Log probabilities
How to use it via API
Venice exposes an OpenAI-compatible API. Swap your base URL and call openai-gpt-56-luna.
curl https://api.venice.ai/api/v1/chat/completions \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai-gpt-56-luna",
"messages": [{ "role": "user", "content": "Explain quantum tunneling simply." }]
}'Specifications
Pricing
Billed per token on Venice: $1.25 per 1M input tokens and $7.50 per 1M output tokens.
New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.
GPT-5.6 Luna vs alternatives
| Model | Price (Venice) | Context window | Open weights | Best for |
|---|---|---|---|---|
| GPT-5.6 Luna | $1.25 in · $7.50 out / 1M | 1M tokens | No | High-volume agentic & coding workloads |
| Claude Sonnet 4.6 | $3.60 in · $18 out / 1M | 1M tokens | No | Balanced reasoning & safety |
| DeepSeek V3.2 | $0.33 in · $0.48 out / 1M | 160K tokens | Yes | Open-weight self-hosting |
| Kimi K2.6 | $0.75 in · $3.50 out / 1M | 256K tokens | Yes | Long-context open weights |
The efficiency-optimized GPT-5.6 model. It reportedly scores 84.3% on Terminal-Bench 2.1 and supports vision, tool use, reasoning, and web search on Venice — all at a much lower per-token cost than rival closed models.
What is it good for?
- •High-volume chatbots and customer-automation pipelines where token cost dominates the budget.
- •Agentic workflows that combine web search, tool calling, and reasoning across long documents.
- •Multi-modal applications that analyze images and text within the same 1M-token context.
- •Code generation and review at scale, leveraging Luna's reported strength on coding benchmarks.
- •Lightweight classification, summarization, and data extraction over large corpora.
Prompting tips
- •Explicitly select the gpt-5.6-luna model ID — the gpt-5.6 alias routes to Sol, which is priced higher.
- •Use the 1M context for long-document analysis, but remember output is capped at 128K tokens per response.
- •Enable reasoning or tool-calling modes for complex agentic tasks; Luna supports both natively on Venice.
- •Cache reusable prompt prefixes where possible: Venice offers discounted cached input at $0.13 per 1M tokens.
Version history
Flagship GPT-5.6 model for maximum reasoning.
Mid-tier balance of performance and cost.
CurrentCurrent — fast, efficient, high-volume tier.
Frequently asked questions
GPT-5.6 Luna is OpenAI's fastest, most cost-efficient model in the GPT-5.6 family, released July 9, 2026. It offers a 1M-token context window, vision, function calling, reasoning, and web search, making it suited for high-volume agentic and coding workloads.
Venice bills Luna at $1.25 per 1M input tokens and $7.50 per 1M output tokens, with cached input priced at $0.13 per 1M. You pay per token with no subscription required.
You can try Luna on Venice using the platform's free-tier credits. Beyond that, usage is billed per token in prepaid credits; there is no separate Luna subscription.
No. Luna is a proprietary OpenAI model with closed weights. It cannot be self-hosted or fine-tuned. If you need open weights, Venice hosts DeepSeek V3.2 and Kimi K2.6.
Yes. On Venice, Luna supports function calling, structured JSON output, vision (including multiple image inputs), reasoning, and web search — enabling complex agentic workflows.
Luna is significantly cheaper per token and offers native web search, vision, and tool use alongside a 1M context window. Claude Sonnet 4.6 is a closed alternative at a higher per-token price. Choose Luna for high-volume, multi-modal production work; Sonnet if you prefer Anthropic's safety stack and budget allows.
Sol is OpenAI's flagship GPT-5.6 model built for maximum reasoning and frontier intelligence. Luna is the efficiency tier, delivering strong performance at a much lower cost per task. Both share the same 1M context and modality support, but Sol leads on the hardest reasoning benchmarks.
Venice runs Luna under an anonymized privacy tier — your prompts are not stored, profiled, or used for training. However, it does not currently run inside a TEE or use end-to-end encryption. Your data is ephemeral and sovereign, but not hardware-isolated.
Related models
Run GPT-5.6 Luna privately.
No prompt logging. No data used for training. Free to start — no credit card.
