Short answer: DeepSeek appears to be research-led, but not anti-commercial. Reports based on a circulated transcript of founder Liang Wenfeng’s investor meeting say the company’s main objective is advanced AI and AGI research, with commercialization—including API sales—used to support infrastructure and preserve strategic flexibility. That is a reported priority, not a fully authenticated corporate policy.
What “research over revenue” means
The phrase describes a choice to delay revenue maximization, not to reject revenue. In practical terms, DeepSeek may be prioritizing:
- new model capabilities over rapid user-growth targets;
- fundamental AGI research before building a broad consumer-product portfolio;
- open or relatively permissive releases instead of keeping every capability proprietary;
- API prices aimed at adoption or hardware-cost recovery rather than the highest possible margin; and
- a research organization less governed by quarterly sales goals.
Those choices can coexist with revenue generation, cost recovery, strategic distribution and long-term enterprise value. “Research-first” is therefore a description of sequencing: research comes before aggressive profit optimization.
What Liang Wenfeng reportedly told investors
The strongest current evidence comes from news reports describing a circulated transcript of a roughly four-hour investor meeting. TechNode reported that Liang presented advanced AI and AGI research as DeepSeek’s central goal and said the company did not want to prioritize short-term commercial growth: TechNode’s report. The Business Times account similarly attributed an AGI-first position to Liang while discussing a possible large funding round.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
- PLEASE NOTE: Exporting an NVIDIA RTX Pro 6000 GPU outside the US requires strict adherence to the U.S. Export Administration Regulations (EAR) and issuance of an export license from the Bureau of Industry and Security (BIS). Compliance and Know Your Customer (KYC) screening may be required as a condition of order acceptance. [NVIDIA Blackwell Streaming Multiprocessor] The new SM features increased processing throughput, and new neural shaders that integrate neural networks inside of programmable shaders | DLSS 4: Multi Frame Generation ensures ultra-smooth frame pacing for lifelike simulations.
- [Double-Flow-Through Design] The RTX PRO 6000 Blackwell features a double-flow-through cooling design, optimizing efficiency and airflow to sustain peak performance under 600W power loads. | [5th Gen Tensor Cores] Deliver up to 3X the performance of the previous generation and support for FP4 precision for faster AI model processing times with reduced memory usage, enabling local fine-tuning of LLMs and generative AI | [4th Gen Ray Tracing Cores] Double the ray-triangle intersection rate of the previous generation to create photoreal, physically accurate scenes and immersive 3D designs with RTX Mega Geometry, which enables up to 100X more ray-traced triangles.
- [PCIe Gen 5] Support for PCIe Gen 5 provides double the bandwidth of PCIe Gen 4, improving data-transfer speeds from CPU memory and unlocking faster performance for data-intensive tasks like AI, data science, and 3D modeling. | [GDDR7 Memory] With 96 GB of GPU memory and 1.8 TB ps bandwidth, it can tackle massive 3D and AI projects, fine-tune AI models locally, explore large-scale VR environments, and drive larger multi-app workflows.
- [DisplayPort 2.1] Achieve unparalleled visual clarity and performance, driving high resolution displays at up to 8K at 240 Hz and 16K at 60 Hz. Increased bandwidth enables seamless multi-monitor setups while HDR and higher color depth support ensures superior color accuracy for precision work, such as video editing, 3D design, and live broadcasting.
- [Universal MIG] Divide a single RTX PRO 6000 Blackwell into multiple isolated instances, each with dedicated resources, allowing for concurrent execution of multiple workloads, optimized GPU utilization, and secure isolation of different applications or users. [WARRANTY] 3 YR Manufacturer's Warranty. Bulk OEM Packaging. Retail Packaging is NOT included.
The underlying document is a circulated transcript, not a clearly authenticated DeepSeek filing or official strategy paper. The transcript archive is available at this URL, with another presentation at Open Market Notes. Accordingly, the remarks should be treated as reported and attributed.
The reported commercial implications
- API access was described as a viable baseline business rather than the company’s ultimate objective.
- Pricing was reportedly intended to recover hardware costs over approximately ten months, rather than maximize revenue or profit.
- DeepSeek’s eventual domestic business model was described as unsettled.
- Open development and commercial monetization were presented as compatible.
These claims support a research-led commercialization thesis, but they do not establish that DeepSeek is nonprofit, unprofitable or indifferent to cash flow.
Public releases provide stronger first-party evidence
DeepSeek’s own technical work is more verifiable than the investor-meeting account. Its R1 paper describes large-scale reinforcement learning intended to improve reasoning. The official R1 repository says DeepSeek released R1-Zero, R1 and six distilled models to the research community.
The repository lists the full R1 and R1-Zero models at 671 billion total parameters, 37 billion activated parameters and a 128K context length. Those are model specifications, not evidence of profitability or total development cost. The repository also states that the R1 series and weights use the MIT license for commercial use, modification, derivative works and distillation, while noting that particular distilled checkpoints derive from separately licensed Qwen and Llama models. An MIT label therefore does not eliminate obligations attached to upstream models.
Free tools Windows power users keep installed
One-click scans. No signup required.
DeepSeek’s V3 materials report a 671-billion-parameter mixture-of-experts model, 37 billion activated parameters per token, pretraining on 14.8 trillion tokens and approximately 2.788 million H800 GPU hours for full training. The repository describes Multi-head Latent Attention, DeepSeekMoE, auxiliary-loss-free load balancing and multi-token prediction. These are company-reported technical and compute figures, not independently audited accounts.
Why open models can still be commercial
Open releases are not necessarily charity. They can function as distribution:
Rank #2
- NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
- 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
- PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
- NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
- Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
- developers become familiar with DeepSeek’s tools and model behavior;
- third parties optimize serving, quantization and integrations;
- low-friction experimentation can create demand for hosted inference;
- organizations may standardize on DeepSeek and later purchase support or capacity; and
- research feedback can improve subsequent models.
DeepSeek’s January 2025 release announcement explicitly described R1 as open and said it could be distilled and commercialized: the official announcement. “Open source” should still be used carefully: weights, code, data, APIs and licenses are different things.
How High-Flyer changes the funding equation
DeepSeek grew out of High-Flyer, a quantitative hedge fund. TechCrunch reported in 2025 that DeepSeek had not announced conventional outside venture funding and that Liang was not rushing to accept it: TechCrunch’s account.
An affiliated financial firm could give a research lab more tolerance for low prices and long development cycles than a startup dependent on successive venture rounds. Venture investors commonly seek rapid revenue growth, product-market fit and a liquidity path. That does not mean High-Flyer can fund unlimited losses, but it may reduce immediate pressure to turn every model release into a high-margin product.
By 2026, reports described possible major external financing, including a 70 billion yuan round and a separate account of roughly $7.4 billion at an approximately $52 billion valuation. These figures may describe different stages or estimates; the available reporting does not establish that financing had definitively closed by August 18, 2026. See Business Times and Axios.
What the pricing signals—and what it does not
DeepSeek’s historical prices were unusually low, but they are dated and should not be treated as August 2026 prices. The January 2025 R1 announcement listed the following:
| Model and date | Cache-hit input | Cache-miss input | Output |
|---|---|---|---|
| R1, January 2025 | $0.14 per million tokens | $0.55 per million tokens | $2.19 per million tokens |
| V3, from February 8 after introductory pricing | $0.07 per million tokens | $0.27 per million tokens | $1.10 per million tokens |
Sources: R1 release and V3 release. Check the live pricing documentation before making a purchasing decision.
Recommended Free Tools
Rank #3
- Professional GPU with Blackwell Architecture
- Blackwell Architecture
- 24GB GDDR7 with PCIe 5.0 & Ray Tracing
- AI Workstation
Low prices can mean cost recovery, customer acquisition, high-volume utilization or an attempt to establish an ecosystem. They do not prove a lack of commercial ambition. The often-repeated $5.6 million V3 figure is a stated training-run or compute estimate, not the total cost of staffing, evaluation, infrastructure, serving and maintenance; see the congressional document and Associated Press analysis.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Can the model be financially sustainable?
A research-first model can work if several conditions hold:
- model and inference efficiency keep compute costs manageable;
- API demand produces enough volume to support infrastructure;
- open releases attract adoption without destroying all differentiation;
- High-Flyer or other backers tolerate long payback periods; and
- future research creates valuable options in APIs, hosting, specialized models or partnerships.
The main failure modes are equally concrete:
- Inference shocks: a demand surge can make subsidized pricing expensive.
- Funding pressure: new investors may demand higher prices, proprietary products or faster growth.
- Commoditization: competitors can fine-tune, distill or improve openly released models.
- Capacity constraints: outages, latency and rate limits can matter more than token price.
- Compute restrictions: chip access and infrastructure allocation can constrain both research and service capacity; congressional testimony discusses these variables at this link.
- Governance costs: privacy, security, moderation, data residency and enterprise compliance require continuing investment.
What this means for developers and financial decision-makers
DeepSeek’s hosted API
The API, documented at DeepSeek’s documentation and available through its platform, may suit cost-sensitive prototyping, research and high-volume text workloads. Buyers needing contractual uptime, indemnity, strict data residency or a predictable enterprise roadmap should evaluate those requirements separately from token price.
Self-hosting
Weights are available through the R1 model card, the V3 model card and DeepSeek’s GitHub repositories. Local deployment can improve data control and customization, but “free to download” does not mean free to operate. GPU capacity, electricity, serving software, monitoring, updates and engineering time become the cost.
Choosing between the two
| Priority | More suitable approach |
|---|---|
| Lowest variable token cost and quick experimentation | Hosted API, subject to current pricing and limits |
| Control over data and deployment | Self-hosted weights with license review |
| Guaranteed support, compliance and contractual service levels | Compare enterprise-focused providers rather than token price alone |
| Customization, distillation or research access | Openly released models, with upstream-license checks |
Bottom line: research before revenue maximization
The evidence supports a careful conclusion: DeepSeek appears to be choosing research before revenue maximization, not research instead of revenue. Reported investor comments describe AGI research as the priority and API pricing as a way to recover hardware costs. Its open releases, technical disclosures, low historical prices and High-Flyer connection are consistent with that model.
The decisive test will be whether those priorities survive large outside financing, rising inference demand and the obligations of serving major commercial customers. Until financing terms and strategy statements are formally documented, “research-led commercialization” is more accurate than the absolute claim that DeepSeek simply puts research over revenue.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




