share_log

What? GPT is now cheaper than DeepSeek? The price war among large language models reaches new heights.

Less than three weeks after unveiling its flagship model, Astra, OpenAI has once again rolled out GPT‑6 Sol and Luna, slashing API pricing by half. Luna's input rate is as low as $0.10 per million tokens—roughly RMB 0.67—and its output rate stands at $0.50. By contrast, DeepSeek V4.1 Flash charges $1 per million tokens for input and $4 for output during off-peak hours, jumping to $2 and $8 respectively during peak periods. Based on typical new request volumes, Luna is already more cost‑effective and avoids the price volatility associated with tiered, peak‑and‑off‑peak pricing.

Less than three weeks after unveiling its flagship model, OpenAI has struck again, slashing the API prices of two new models by half, aiming to erode DeepSeek's low‑cost competitive moat. The race among large‑model providers is shifting from a focus on "capabilities" to a full‑blown battle over "costs."

On September 22, Eastern Time, OpenAI officially unveiled two new models, GPT‑6 Sol and GPT‑6 Luna, and announced that their API pricing has been slashed by another 50% compared to the promotional rates for GPT‑5.6. Specifically, GPT‑6 Luna's input rate is as low as $0.10 per million tokens, with an output rate of $0.50, directly encroaching on the price range previously regarded as a competitive advantage by DeepSeek and mounting a direct challenge to the core market segments of Chinese rivals.

This price cut is not merely a commercial concession. OpenAI stated that the reduction stems from improved caching and inference efficiency, with the company passing the resulting cost savings directly to its users. Meanwhile, Anthropic unveiled Claude Opus 5.5 on the same day, emphasizing lower operating costs. With two leading AI companies entering the market on the same day at more competitive pricing, the industry has entered a new phase in which performance, speed, and cost are advancing in tandem.

Less than three weeks after Astra's launch, OpenAI swiftly expanded its product lineup.

On September 3, OpenAI unveiled its flagship GPT‑6 model, Astra, claiming it has reached new frontiers in capabilities across areas such as computer operations, software engineering, cybersecurity, and professional tasks. At the time, demand for Astra was so robust that OpenAI was forced to temporarily suspend sign-ups for new Pro subscriptions.

In just under three weeks, OpenAI has launched Sol and Luna under the Astra umbrella, rapidly expanding its product lineup to address diverse performance levels, cost requirements, and use cases. According to OpenAI, Sol and Luna "build on the technological advances underlying GPT‑6 Astra," bringing many of Astra's strengths to "faster, more affordable models" to support large‑scale deployments.

This pace aligns with the AI industry's increasingly clear product strategy: flagship models push the boundaries of capability, while faster, more affordable variants target broader, high‑frequency use cases.

API prices have been cut in half, and Luna has now entered DeepSeek's price range.

Price is the clearest commercial signal from this launch.

According to the pricing table released by OpenAI, the input price for GPT‑6 Sol has been reduced from $4 per million tokens—its promotional rate under GPT‑5.6 Sol—to $2, while the output price has dropped from $20 to $10. For GPT‑6 Luna, the input price has fallen from $0.20 to $0.10, and the output price has decreased from $1.20 to $0.50. Both models' API costs have been cut by 50% across the board.

OpenAI explained that the price reductions stem primarily from improvements in caching and inference efficiency, rather than mere commercial subsidies. The company has also further refined its prompt‑caching mechanism, enabling agents and long‑context conversations to reuse previously processed context more extensively, thereby reducing the per‑call cost.

Notably, Luna's pricing is particularly aggressive, entering the ultra‑low‑price segment previously dominated by DeepSeek's model lineup, leading observers to view this launch as OpenAI's direct challenge to DeepSeek's price‑based competitive moat.

By comparison, DeepSeek V4.1 Flash charges 1 yuan per million tokens for cache misses during idle periods and 4 yuan per million tokens for output, rising to 2 yuan and 8 yuan respectively during peak times. Based on typical new‑request pricing, Luna is already more cost‑effective and avoids the cost volatility associated with tiered pricing.

However, DeepSeek still enjoys a pricing advantage. Its in‑idle cache hit rate is priced at just 0.02 yuan per million tokens, significantly lower than Luna's. Moreover, V4.1 Flash boasts markedly superior overall capabilities, supporting a 1‑million‑token context and enabling multi‑modal understanding, tool invocation, and reasoning modes.

Sol focuses on complex tasks, while Luna targets high-frequency, low-cost use cases.

The two models are positioned differently and are not simply "the same model at different price points."

Sol is designed for tasks requiring higher levels of capability, including complex professional work, coding, and agent‑based scenarios. In AutomationBench's enterprise‑wide cross‑application workflow benchmark, GPT‑6 Sol scored 33.2% under "xhigh effort," surpassing GPT‑6 Astra's 30.3% at "low effort" and Claude Opus 5's 26.9% at "max effort." According to OpenAI, the cost of completing each task with Sol is just $0.27.

Luna, on the other hand, places greater emphasis on cost efficiency. According to OpenAI, in high‑intensity testing on AutomationBench, GPT‑6 Luna achieved a 5.4‑percentage‑point improvement over its predecessor, while reducing per‑task costs by 58%. In DeepSWE v1.1, Luna scored as high as 66.6%, matching the moderate‑intensity performance of Claude Opus 5 and Fable 5, yet at task‑level costs that are 93% and 96% lower, respectively.

In terms of factual accuracy, internal testing at OpenAI shows that GPT‑6 Sol generates roughly half the number of errors compared to its predecessor; Luna, when operating at higher reasoning intensities, can match the performance of GPT‑5.6 Sol while costing about one percent of the latter. OpenAI notes that this evaluation is based on anonymized ChatGPT conversations previously flagged by users as containing factual inaccuracies and does not necessarily reflect all real-world use cases.

Advanced Work and Codex—free users can try Luna.

In terms of product coverage, OpenAI did not immediately integrate Sol and Luna into the standard ChatGPT chat interface this time.

According to the plan, Plus, Pro, Business, Enterprise, and Edu users can access both models in ChatGPT Work and Codex, with API access also available; free‑tier and Go‑tier users can experience GPT‑6 Luna within the desktop app.

This arrangement clearly tilts the initial use cases toward work and development, rather than mass‑market consumers. Sol and Luna's low cost and high efficiency also make them better suited for coding tasks that require frequent model invocations, agent‑based workflows, and enterprise‑level automation. At the same time, directing some high‑frequency workloads to lower‑cost models could help OpenAI ease infrastructure strain as it scales usage.

Industry competition has shifted comprehensively from "competing on flagship products" to "competing on scale and cost."

Notably, as OpenAI launches Sol and Luna, the industry is simultaneously debating whether AI development should be slowed down.

Earlier this month, Anthropic CEO Dario Amodei publicly called for a slowdown in the pace of cutting-edge AI development, warning that rapid progress could pose significant risks. OpenAI CEO Sam Altman subsequently echoed this sentiment, emphasizing the need to decelerate the advancement of state-of-the-art models and strengthen safety safeguards. Meanwhile, OpenAI Chief Scientist Jakub Pachocki argued that coordinating a deliberate slowdown in future AI research is a crucial pathway to ensuring the safe deployment of self-improving AI systems.

Meanwhile, OpenAI has not slowed its own product‑iteration pace: less than three weeks after launching Astra, it unveiled two new models and slashed its API pricing by 50%. On the same day, Anthropic also rolled out Claude Opus 5.5, likewise touting lower operating costs as a key selling point.

While debating how to slow the pace of AI development, the industry is simultaneously accelerating model iteration and driving down deployment costs, ushering in a new phase in which both capabilities and costs are rapidly declining.

Editor/Deng

The translation is provided by third-party software.


The above content is for informational or educational purposes only and does not constitute any investment advice related to EleBank. Although we strive to ensure the truthfulness, accuracy, and originality of all such content, we cannot guarantee it.