RadixArk, the company commercializing the open-source SGLang inference project, was reportedly valued at approximately $400 million in an Accel-led financing. TechCrunch, citing two people familiar with the matter on January 21, 2026, said the financing amount was not confirmed. That distinction matters: a reported valuation is not the same as money raised, revenue, or proof of a durable business.
The spinout is part of a broader rush to fund the software and services that run AI models after training. SGLang remains open source, while RadixArk is reportedly charging for hosting and developing Miles, a reinforcement-learning framework.
What is confirmed about RadixArk?
| Item | What is reported |
|---|---|
| Company | RadixArk, formed around the SGLang project |
| Reported valuation | Approximately $400 million, according to two people familiar with the matter |
| Reported lead investor | Accel |
| Financing amount | Not confirmed by TechCrunch |
| SGLang origin | An open-source project started in 2023 in Ion Stoica’s UC Berkeley lab |
| Leadership reported by TechCrunch | CEO Ying Sheng, a former xAI engineer and former Databricks research scientist |
| Other reported product | Miles, a reinforcement-learning framework |
| Reported business model | Continued open-source development with paid hosting and related services |
These details come from TechCrunch’s January 21, 2026 report. The report does not establish the round size, whether the valuation was pre-money or post-money, the ownership sold, liquidation preferences, revenue, customer contracts, or the financing instrument. Consequently, “$400 million valuation” should not be read as “RadixArk raised $400 million” or “the company has a $400 million business.”
SGLang and RadixArk are different things
SGLang is the software layer
SGLang is an open-source inference and serving engine. Inference is the act of running a trained model to produce an answer to a request. The engine manages how requests are queued, combined, cached and executed on accelerators such as GPUs.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
The project is designed for large language and multimodal models, structured output and high-throughput serving. Its public repository is available at github.com/sgl-project/sglang.
RadixArk is the commercial company
RadixArk is the venture-backed entity built around that project. The available reporting says contributors and maintainers moved into the startup, but that SGLang is continuing as open source. The company therefore appears to be commercializing services around a public codebase rather than turning SGLang itself into a closed product.
Hosting is the only monetization method specifically reported. Enterprise support, dedicated deployments, control-plane software, security features, optimization work and cloud distribution are plausible ways to monetize an open-source engine, but they are not confirmed RadixArk offerings in the cited coverage.
Miles adds reinforcement-learning infrastructure
RadixArk is also developing Miles, described as a reinforcement-learning framework. That could broaden the company beyond serving already-trained models, although the available reporting does not provide product specifications, pricing or adoption data for Miles.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteWhy inference is attracting capital
Model training is an enormous but usually episodic expense. Inference begins when users arrive and continues with every request. A production service can incur recurring GPU, memory, networking, storage and electricity costs whether demand is steady or spiky.
Rank #2
Serving software can reduce those costs without retraining the underlying model. Better scheduling can keep accelerators busy; continuous batching can combine requests arriving at different times; caching can avoid recomputing repeated prompt material; and lower latency can allow a service to handle more users on the same hardware. Those improvements are close to the operating costs of an AI application, which makes the inference layer strategically valuable.
The investment case is not that every inference company will win. It is that applications are moving from demonstrations to production, where latency, predictable cost, data control and support become procurement requirements. Open-source projects can become widely adopted infrastructure before a commercial company forms around hosting, support or enterprise capabilities.
How SGLang improves serving performance
Scheduling and continuous batching
An inference server must decide which requests use limited accelerator memory and compute at each moment. Dynamic or continuous batching admits new requests while existing generations are still running, improving utilization compared with waiting for a fixed batch to finish.
Prefix-aware caching
Many requests share an instruction, system prompt or document prefix. SGLang’s RadixAttention-style approach stores reusable prefix states in a tree-like cache so later requests can avoid repeating that work. The benefit depends on how much prompt material is shared and whether the cache fits in available memory.
Structured generation
Applications often need valid JSON, grammar-constrained output or tool calls rather than unconstrained text. SGLang combines serving with structured generation controls, reducing the need for application-side retries and parsing work.
Rank #3
What the published benchmark does—and does not—show
The original SGLang paper reported up to 6.4× higher throughput than comparison systems in its experiments. That is a paper-specific, workload-specific result, not a universal production multiplier. Results can change with model architecture, prompt and output lengths, quantization, concurrency, GPU generation, drivers and configuration. The paper is available at arxiv.org/abs/2312.07104.
RadixArk compared with vLLM and Inferact
vLLM is the closest open-source counterpart in this discussion. It was also incubated in Stoica’s Berkeley lab, and its creators later formed Inferact. TechCrunch reported on January 22, 2026 that Inferact announced a $150 million seed financing at an $800 million valuation, led by Andreessen Horowitz and Lightspeed.
Recommended Free Tools
| Dimension | SGLang / RadixArk | vLLM / Inferact |
|---|---|---|
| Origin | UC Berkeley research project | UC Berkeley research project |
| Core role | Open-source serving and inference optimization | Open-source serving and inference optimization |
| Commercial vehicle | RadixArk | Inferact |
| Reported financing | Approximately $400 million valuation; round size unconfirmed | $150 million seed at an $800 million valuation |
| Reported lead investor | Accel | Andreessen Horowitz and Lightspeed |
| Reported emphasis | Structured generation, prefix-aware execution and reinforcement-learning tooling | Broad serving ecosystem and enterprise commercialization |
| Unresolved questions | Financing terms, traction and customer concentration | Final product and commercial strategy after the spinout |
Neither project is universally faster. A fair evaluation requires the same model, GPU, quantization, prompt distribution, output lengths, concurrency and reliability targets. The related financing report is at TechCrunch’s Inferact article.
Where RadixArk fits in the inference market
“Inference company” can describe several non-interchangeable layers:
- Open-source engines: SGLang and vLLM, which organizations can run themselves.
- Managed platforms: Baseten and Fireworks AI, which operate deployment and hosted serving services.
- Programmable cloud: Modal, which provides elastic infrastructure for developer-run workloads.
- Alternatives: cloud-provider endpoints, model-vendor APIs, Kubernetes deployments and in-house systems based on CUDA, TensorRT-LLM, Triton, vLLM, SGLang or custom code.
TechCrunch reported that Baseten had raised $300 million at a $5 billion valuation and Fireworks AI had raised $250 million at a $4 billion valuation, citing reported financing information. Those figures are attributed financing claims, not audited measures of revenue or profitability. A separate report said Modal was in talks to raise at a $2.5 billion valuation; talks are not a completed financing. See the Modal report.
Rank #4
How to evaluate SGLang, vLLM or a hosted service
Measure the workload, not the headline benchmark
- Time to first token and time between tokens at p50 and p99.
- Requests per second and tokens per second at realistic concurrency.
- Cost per million input and output tokens, including idle capacity, networking, storage and operations.
- Performance across the exact model family, quantization and hardware you intend to use.
- Cold-start time, autoscaling behavior, failure recovery and rolling upgrades.
Check technical fit
- Support for your model architecture, multimodal inputs, speculative decoding and quantization formats.
- Long-context memory use and prefix-cache efficiency.
- Structured output, JSON schemas and tool-calling behavior.
- NVIDIA, AMD, CPU and cloud-specific accelerator compatibility.
- OpenAI-compatible APIs, containers, Kubernetes and model-registry integration.
Check business and security terms
- Open-source license and commercial-use rights.
- Which capabilities remain in the public repository and which require a paid control plane.
- Support, service-level agreements, data residency and vendor lock-in.
- Model-file provenance, container and plugin isolation, authentication, secrets management and endpoint exposure.
- Whether the operational burden of self-hosting outweighs the infrastructure savings.
Open source removes a software subscription, not the cost of GPUs, power, monitoring, patching, capacity planning, security and on-call work. A managed endpoint may cost more per token while reducing those responsibilities.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Open-source commercialization creates governance questions
A spinout can fund maintainers and accelerate development, but it also changes incentives. Buyers and contributors should look for project documentation covering repository control, maintainer decisions, outside-contributor representation, license terms and the boundary between community and proprietary features.
The available reporting confirms continued open-source development but does not answer whether the most important features will remain upstream, whether commercial features will be exclusive, or how governance could evolve. Those are diligence questions rather than evidence that a split has already occurred.
What the valuation does—and does not—tell investors
A private-company valuation is the price implied by a particular financing transaction. Without the round size, security type, pre-money or post-money definition, ownership sold and investor rights, it cannot be converted into a public-market-style estimate of business value. It also says nothing by itself about revenue, gross margin, retention or customer concentration.
Reported users such as xAI and Cursor indicate interest in SGLang, but the coverage does not establish whether those organizations use it in production, pay RadixArk, use unmodified upstream code or deploy it only for particular workloads. Adoption should therefore not be treated as verified revenue evidence.
Best Value
For investors, the durable opportunity depends on whether RadixArk can turn technical adoption into recurring hosting or enterprise revenue while preserving a healthy project that users can trust. For technology buyers, the relevant question is narrower: whether the engine or service lowers total cost and meets reliability, security and compatibility requirements on the buyer’s own workload.
Frequently Asked Questions
Did RadixArk raise $400 million?
No confirmed source establishes that. TechCrunch reported an approximate $400 million valuation in an Accel-led financing, while the financing amount was not confirmed.
Is SGLang a paid commercial product now?
SGLang remains reported as an open-source project. RadixArk is the commercial company reportedly charging for hosting and related services.
Is SGLang faster than vLLM?
There is no universal winner. Performance depends on the model, hardware, quantization, context lengths, concurrency and serving configuration, so buyers should run matched benchmarks.
The Bottom Line
RadixArk’s reported $400 million valuation is a meaningful signal that investors view inference infrastructure as strategically important. It is not proof of $400 million in funding, revenue or product-market fit. The practical test is whether RadixArk can monetize hosting and enterprise services around an open-source engine while delivering measurable, workload-specific gains over vLLM, managed endpoints and in-house alternatives.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




