Software & Apps

The trajectory of OpenAI reasoning models

OpenAI faces intense competition as DeepSeek-R1 offers costs 20 times lower than o1. While OpenAI shifts focus to o3-pro for science and math, DeepSeek prepares a 1.2 trillion parameter R2 model despite hardware challenges.

The trajectory of OpenAI reasoning models

Competition and pricing

OpenAI cut the price of its cheapest GPT-5.6 model, Luna, by 80% on July 9, 2025. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, down from $1 and $6. The mid-tier model, Terra, dropped in price. This price move follows competition from Anthropic and Google. Anthropic reduced Opus pricing by 67% in February, changing the rate from $15/$75 to $5/$25. Google restructured Gemini consumer subscriptions in May, dropping the top tier from $249.99 to $99.99. Anthropic also launched Claude Sonnet 5 on June 30 with introductory pricing of $2 input and $10 output. DeepSeek-R1 provides a significant cost advantage because it operates at 20 times lower cost than OpenAI o1. The AI startup Lindy moved 100% of its traffic from Anthropic to DeepSeek to save millions of dollars because the cost of using Chinese models is 60% to 90% cheaper than leading American rivals. American companies use Chinese models via OpenRouter at rates between 30% and 46% since February. OpenAI noted an internal 20% reduction in the end-to-end cost of serving its models and a 15% improvement in token-generation efficiency. The market just delivered a chunk of subsidized access for free because three companies are afraid of each other. Small businesses in places like Pawtucket or software shops on the East Side benefit from these cuts. The arithmetic changes when the cheapest tier drops fivefold.

Reasoning capability tiers

DeepSeek V4.1-Flash arrived on September 10, 2026, as the smallest model in its architecture family. It includes native multimodal visual understanding. DeepSeek also provides three thinking effort levels for its V4 models. Users choose between "low", "high" and "max" effort levels based on task complexity. OpenAI released o3-pro on June 10, 2025, for Pro and Team users. This model replaces o1-pro and focuses on science, education, and programming. The o3-pro model is designed to think longer and provide more reliable responses for challenging questions where reliability matters more than speed, making it ideal for science and math. The model can search the web, analyze files, reason about visual inputs, and use Python. Reviewers consistently prefer o3-pro over o3 in science, education, and programming. Users who prefer conversational warmth or creative ideation now use GPT-5.2 because OpenAI announced the retirement of GPT-4o. DeepSeek-R1 achieved a 97.3% score on MATH-500. It also achieved a Codeforces rating of 2029. The V4-Flash model scored 82.7 on the Terminal-Bench 2.1.

Model Release Date Type
o3-pro June 10, 2025 Reasoning
o4-mini April 16, 2025 Reasoning
DeepSeek V4.1-Flash September 10, 2026 Multimodal

Industry shifts and delays

DeepSeek R2 stays unreleased because CEO Liang Wenfeng is not satisfied with its performance. Reports from Reuters and The Information indicate that training runs on Huawei Ascend hardware failed. This forced a pivot back to Nvidia. DeepSeek closed its first external funding round in June 2026 for $7 billion. The company also accelerated data-center construction in Ulanqab, Inner Mongolia. Backers of the 50 billion yuan round include Tencent, CATL, NetEase, and JD.com. OpenAI scheduled the retirement of GPT-4o, GPT-4.1, and o4-mini for February 13, 2026. The gpt-5.4-cybermodel is also deprecated and will be removed from the API on October 1, 2026. DeepSeek R2 is expected to be a hybrid Mixture-of-Experts model with 1.2 trillion parameters. This model is expected to be trained on 5.2 petabytes of data. The anticipated model also includes multimodal vision capabilities. Frontier firms generate 8.3 times as many output tokens per active user as typical firms. You should check if your current deployment relies on deprecated models, as you likely already monitor your API versions. Will the high cost of training larger models force another round of massive price cuts?