Morgan Stanley believes that DeepSeek announced a price reduction of 11%–60% for its V4 Flash model amid competitive entries from Alibaba and Zhipu AI, as well as the imminent launch of its new V4.1 Flash product. However, the post-reduction prices remain higher than the initial launch prices and are the highest among mainstream lightweight models. While competition in the large language model (LLM) sector is intensifying under a dual-track system, gross margin constraints will curb disorderly price wars.
Amid intensifying competitive pressure and on the eve of new product launches, DeepSeek was compelled to reduce prices for its flagship lightweight model; however, this move did not alter its pricing disadvantage in niche markets.
On September 9, DeepSeek announced a price reduction for its V4 Flash model, effective the following day. The adjustment covers three pricing components: non-peak cached input, non-cached input, and output, with reductions of 60%, 33%, and 11%, respectively. After the adjustment, the price for non-cached input dropped to RMB 1.5 per million tokens, cached input to RMB 0.03, and output to RMB 6.
According to Zhuifeng Trading Desk, Morgan Stanley analysts Gary Yu, Lydia Lin, and Yang Liu believe that this price reduction is driven by two factors: first, from$Alibaba (BABA.US)$competitive pressure from its Qwen 3.8-Flash and$Z.AI (02513.HK)$AI GLM 5.3-Flash, both of which were launched in August this year; second, strategic repricing to make way for DeepSeek's upcoming V4.1 Flash product.
However, even after the price cut, DeepSeek V4 Flash remains priced above its April launch level and maintains the highest pricing tier among major competitors, meaning the market landscape has not been reversed as a result.

Behind the Price Cut: Squeezed by Old Rivals and New Pressures
In August this year, two lightweight competing models entered the market successively, directly impacting the market position of DeepSeek V4 Flash.
According to Morgan Stanley data,$Alibaba (BABA.US)$the non-cached input price for Qwen 3.8-Flash is RMB 1 per million tokens, and the output price is RMB 3;$Z.AI (02513.HK)$AI GLM5.3-flash adopts a more aggressive pricing strategy, with non-cached input priced at only CNY 0.4 and output at CNY 1.4—the latter being less than one-quarter of DeepSeek's post-price-cut quote.
In terms of model scale, DeepSeek V4 Flash has 284 billion parameters, significantly exceeding Qwen3.8-flash's 125 billion but falling short of GLM5.3-flash's 320 billion. However, the three models achieved similar scores on the AA Intelligence Index: DeepSeek scored 52, GLM5.3-flash took a slight lead with 57, and Qwen3.8-flash scored 56. This indicates that DeepSeek does not enjoy a significant performance premium, thereby accumulating pricing pressure.

Still the most expensive after price cuts: Pricing disadvantage remains unresolved
Despite substantial price reductions, DeepSeek V4 Flash remains the most expensive option among the three mainstream lightweight models.
By comparison, after the price cut, DeepSeek V4 Flash's output price is CNY 6 per million tokens, which is twice that of Qwen3.8-flash (CNY 3) and more than four times that of GLM5.3-flash (CNY 1.4). For non-cached input, DeepSeek's adjusted quote is CNY 1.5, higher than Qwen3.8-flash's CNY 1 and significantly above GLM5.3-flash's CNY 0.4.
Notably, the post-cut prices remain higher than the initial launch prices of DeepSeek V4 Flash in April this year, when non-cached input was CNY 1 and output was CNY 2. This implies that V4 Flash has experienced a pricing trajectory of initial increases followed by decreases; the current adjustment is only a partial correction and has not returned to the original baseline.
Industry Logic: Dual-track Pro+Flash strategy becomes mainstream
Morgan Stanley points out that the dual-track product strategy of "high-priced flagship + low-priced fast" is becoming the standard approach for large model vendors, and DeepSeek's recent adjustment aligns with this trend.
This model concurrently launches high-performance, high-cost Professional (Pro) models and low-latency, low-priced Lightweight (Flash) models. The former serves enterprise scenarios requiring high precision, while the latter targets high-frequency, cost-sensitive consumer and developer needs. As competition intensifies at the Flash layer, pricing games are expected to continue.
However, Morgan Stanley emphasizes that gross margin constraints will be a key factor in curbing disorderly price wars. Analysts believe that under profitability pressure, the market is unlikely to see cost-ignoring price competition; vendors are more inclined to compete for market share through differentiated performance and ecosystem advantages rather than solely through price.
Editor/melody