The Chinese AI developer DeepSeek has hit an annualised revenue run rate (ARR)of $1 billion according to The Information.
Citing two sources close to the matter, the publication reported that DeepSeek revealed the firm has more than doubled its revenue from $500 million earlier this year, with the figures shared with investors amid a push by DeepSeek for a $7.5 billion funding round.
The firm, which was first widely noticed for the disruptive release of its large language model (LLM) DeepSeek-R1 in January 2025, has become known as a provider of capable and cheap and performant LLMs that allow businesses to generate code or knowledge-based outputs for a fraction of the API costs Western labs charge.
Like many Chinese AI developers, DeepSeek made cheap and fast AI its focus during its rise to prominence. Its latest lightweight model, V4, is both its cheapest and most capable, launching at an API price of $0.15 per million input tokens and $0.60 per million output tokens.
Much of DeepSeek’s new revenue comes from price hikes, however, with the firm having shifted to a peak/off-peak pricing model in mid August. V4 Flash now costing up to $1.2 per million output tokens when used in thinking mode during peak hours.
The amount DeepSeek charges for cache hits, a term used for inference that draws on previously processed context such as a recently-referenced code repository, has increased significantly.
DeepSeek V4 Pro, the firm’s best model for coding, now costs between $0.022 and $0.044 per million cached input token hits, an increase of 507 to 1,114 per cent.
Since the release of V4 Flash 0731 and its flagship model V4 Pro 0813, however, DeepSeek has faced increased competition from Chinese competitors. Z.ai’s GLM-5.3, Moonshot AI’s Kimi K3 and Alibaba’s Qwen-3.8 perform better in benchmarks such as coding, though at a higher price point.
OpenAI’s GPT-6 Luna also now undercuts DeepSeek on price per task at just $0.10 per million input tokens and $0.50 per million output tokens.
With competition increasing, DeepSeek has allocated more time and spending to training more powerful models. Reuters reported that as much as 70 per cent of DeepSeek’s compute capacity is currently used for LLM training runs, with just 30 per cent left for customer inference.
Currently, the firm relies heavily on Huawei and Nvidia chips, with the former used primarily for inference. In July, Reuters reported that DeepSeek is looking to develop its own AI chip to lower its inference costs and reduce dependence on third-party suppliers.
This is a strategy used by several hyperscalers, such as Google with its TPU chip family and Amazon with its Trainium and Inferentia hardware.


.jpg)