AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How OpenAI’s Price Cuts For GPT‑6 Sol And Luna Affect AI Benchmarking on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI has announced significant price reductions for its GPT‑6 Sol and Luna models, cutting costs by half compared to GPT‑5.6. This shift impacts AI benchmarking, making models more affordable while showing mixed results in quality and performance metrics.

OpenAI has announced a 50% price reduction for its GPT‑6 Sol and Luna models, effective immediately. The models, released on September 22, 2026, are now priced at half the cost of their GPT‑5.6 predecessors, marking a significant shift in AI deployment economics. This move aims to make advanced AI more accessible for a broader range of applications, from automation to research, by lowering the cost barrier.

The new GPT‑6 Sol and Luna models are designed with improved caching and inference efficiencies, enabling OpenAI to pass savings onto users. GPT‑6 Sol’s input cost drops from $4 to $2 per 1 million tokens, and output from $20 to $10, while Luna’s costs are halved from $0.20 to $0.10 for input, and from $1.20 to $0.50 for output. These reductions are part of OpenAI’s broader strategy to democratize AI access by offering more affordable models.

Independent evaluation by Artificial Analysis indicates that despite the price cuts, model performance remains roughly stable, with some metrics showing improvements and others slight regressions. For example, GPT‑6 Sol’s cost per task, at maximum effort, is now about $1.06—roughly half of GPT‑5.6 Sol’s $1.99—while Luna’s per-task cost drops to approximately $0.07, about 60% less. However, these models demonstrate mixed results in benchmarks measuring intelligence, coding, and knowledge work, with some scores improving and others declining.

One notable change is in hallucination rates, which have decreased significantly. GPT‑6 Sol’s hallucination rate on AA‑Omniscience fell from 92% to 60%, and Luna’s from 93% to 77%, partly because the models now refuse to answer more often, thus reducing false answers but also lowering overall accuracy. These adjustments reflect OpenAI’s focus on balancing cost, safety, and performance in the new models.

At a glance
updateWhen: announced September 22, 2026
The developmentOpenAI’s recent price cuts for GPT‑6 Sol and Luna models, announced on September 22, 2026, are reshaping AI benchmarking by lowering costs and influencing model evaluation standards.

GPT‑6 Sol and Luna: half the price, about the same intelligence

OpenAI’s September 22, 2026 release doesn’t raise the ceiling. It lowers the cost of everything below it, which changes what’s worth automating.

GPT‑6 Sol
$4 / $20 → $2 / $10
GPT‑6 Luna
$0.20 / $1.20 → $0.10 / $0.50

Per 1M input / output tokens. Cached input reads keep the 90% discount.

Cost per task, halved

Measured by Artificial Analysis as the weighted cost of one Intelligence Index task, at max effort.

GPT‑5.6 Sol
$1.99
GPT‑6 Sol
$1.06
GPT‑5.6 Luna
$0.18
GPT‑6 Luna
$0.07

The effort dial moves cost more than the model choice

Model and effortIntelligence IndexCost per task
GPT‑6 Sol (max)48$1.06
GPT‑6 Sol (low)34$0.13
GPT‑6 Luna (max)37$0.07
GPT‑6 Luna (low)21$0.0045
GPT‑6 Luna (non‑reasoning)18$0.01

Sol at low effort keeps about 70% of its max score for roughly an eighth of the cost, because it writes far fewer reasoning tokens. For reference, Claude Opus 5.5 leads the same index at 58.

What got better, and what got worse

Better

  • Hallucination rate on AA‑Omniscience: Sol 92% → 60%, Luna 93% → 77%
  • Coding Agent Index: Sol 57, up 2 points, at ~50% lower cost per task
  • OpenAI reports about half as many factual mistakes for Sol as its predecessor
  • Higher cache hit rates; GitHub reports over 50% fewer prompt tokens needing fresh processing

Sol gets there partly by declining more: it attempts 83% of questions vs 99%, and accuracy falls 59% → 54%.

Worse

  • GDPval‑AA v2.1: Sol down ~100 Elo, Luna down ~75
  • AA‑Briefcase v1.1: Luna down ~45 Elo
  • Coding Agent Index: Luna 41, down 2 points
  • Both models write more output tokens per task than their predecessors

Reviewers attribute the drops to weaker presentation and deliverables that omit required elements.

What to do about it

Already on GPT‑5.6 Sol or Luna? The move is mostly a price cut. Re‑test first if your output is a document someone reads, not data a system consumes.
Shelved an automation on cost? Token prices halved and the effort dial adds another order of magnitude. Re‑run the business case.
Choosing between labs? The question is no longer which model is smartest, but which clears your quality bar at the lowest cost per task.
ThorstenMeyerAI.comSources: OpenAI (pricing, vendor benchmarks) and Artificial Analysis (independent evaluation and model pages). Figures as of 23 September 2026.

Impact of Cost Reduction on AI Deployment and Benchmarking

The price cuts for GPT‑6 Sol and Luna are set to reshape the AI landscape by making large language models more affordable for a wider range of users and use cases. For businesses, this lowers the entry barrier for deploying advanced AI in customer service, content generation, and automation, potentially accelerating adoption across industries.

From a benchmarking perspective, the reduced costs mean that models can be tested at larger scales or with more extensive datasets without prohibitive expenses. This could lead to more comprehensive evaluations of AI capabilities and a shift in how performance and efficiency are measured in the field. However, the mixed results in quality metrics highlight that lower prices do not necessarily equate to better performance across all tasks, emphasizing the importance of context-specific testing.

Furthermore, the shift toward more refusals and lower hallucination rates suggests a focus on safety and factuality, which are critical for enterprise applications. Overall, these developments could influence future AI model design, evaluation standards, and the competitive landscape among AI providers.

Amazon

AI model cost management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Pricing and Benchmarking Trends

Prior to the September 2026 announcement, OpenAI’s GPT‑5.6 models were considered the benchmark for high-performance language models, with relatively high costs limiting widespread deployment. The release of GPT‑6 Astra earlier in 2026 introduced a new top-tier model with even greater capabilities, but at a premium price.

The recent introduction of GPT‑6 Sol and Luna marks a strategic shift, emphasizing cost efficiency over absolute performance. OpenAI’s focus on caching improvements and inference optimization reflects a broader industry trend toward reducing operational expenses while maintaining acceptable quality levels.

Independent evaluations by organizations like Artificial Analysis have historically tracked model performance across various benchmarks, including intelligence, coding, and knowledge work. These assessments inform enterprise decisions and influence industry standards, making the recent price reductions particularly impactful for the benchmarking community.

Amazon

AI benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Performance and Adoption

It is still unclear how these models will perform in large-scale, real-world deployments over time, especially regarding their long-term stability and safety. Although initial benchmarks show mixed results, the true test will be their effectiveness in diverse operational settings.

Additionally, the impact of increased refusals and lower hallucination rates on user experience and productivity remains to be fully understood. It is also uncertain how competitors will respond, whether through similar price cuts or new model innovations.

Further data from ongoing deployments and third-party evaluations will be necessary to confirm the models’ practical advantages and limitations.

Amazon

large language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Evaluation and Industry Response to Price Cuts

OpenAI is expected to continue refining GPT‑6 Sol and Luna, possibly releasing updates that enhance their performance and safety features. Industry analysts anticipate that other AI providers will respond with their own pricing strategies, potentially leading to a more competitive landscape.

Research organizations and enterprise users will likely increase testing of these models across various benchmarks and real-world tasks, providing more comprehensive data on their capabilities and limitations.

Additionally, the AI community may develop new standards for benchmarking cost efficiency versus performance, influencing future model development and deployment decisions.

Amazon

affordable AI development platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How much cheaper are GPT‑6 Sol and Luna compared to previous models?

GPT‑6 Sol’s input costs are halved from $4 to $2 per 1 million tokens, and output costs from $20 to $10. Luna’s costs are reduced from $0.20 to $0.10 for input, and from $1.20 to $0.50 for output, representing approximately 50-60% savings across the board.

Do the price reductions affect the models’ quality?

Initial evaluations show mixed results: some benchmarks indicate stable or improved performance, especially in hallucination reduction, while others reveal regressions in knowledge and complex reasoning tasks. The models also tend to refuse answering more often, which impacts certain workflows.

What are the implications for AI benchmarking?

The lower costs enable larger-scale testing and more diverse evaluations, potentially shifting benchmark standards. However, the mixed quality outcomes suggest that cost efficiency does not necessarily equate to overall better performance, emphasizing the need for careful, task-specific testing.

Will competitors follow OpenAI’s pricing strategy?

It is likely that other AI providers will consider similar price cuts or new model releases to remain competitive, especially as cost becomes a more critical factor in enterprise adoption and market share.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

732 Bytes to Root. One Hour of Scan Time.

A new Linux kernel flaw allows root access with a 732-byte script, discovered in one hour of scanning, collapsing security cost assumptions.

Microsoft’s Efforts In Making Windows 11 More Efficient Isn’t Just Due To The DRAM Shortage; Analyst Believes MacBook Neo Played A Pivotal Role In This Decision

Microsoft’s efforts to enhance Windows 11 performance are driven by multiple factors, including the influence of Apple’s MacBook Neo, not solely due to DRAM shortages.

Quantum Risk Monitoring For Government Contractors: A Guide

A new quantum risk monitor tool is being tested for government contractors and regulated firms to identify quantum-vulnerable assets and ensure compliance ahead of PQC deadlines.

How AI Is Redefining Workflow Automation In 2026

AI is redefining workflow automation in 2026 through new tools, integrations, and platform strategies, impacting businesses and developers alike.