Fall ResetAmazon USFall reset deals: check better picks before checkoutAmazon US: today's deals, useful picks and quick comparisons.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanFall ResetAmazon USWork and home upgrades are worth comparing todayAmazon US: today's deals, useful picks and quick comparisons.See Picks×
Skip to content
Blog

Positron AI Enters Nvidia’s Turf With an Oracle Cloud Deal—but It’s Not Replacing Nvidia

By TheFinanceBase Team8 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Positron AI says Oracle is deploying tens of millions of dollars’ worth of its inference-focused systems and racks in Oracle Cloud Infrastructure (OCI). That is a meaningful commercial milestone for the startup and a new option for Oracle’s cloud customers—but it is not evidence that Positron is broadly displacing Nvidia.

The deal targets AI inference, particularly mixture-of-experts models, where power consumption, memory capacity and cost per token can matter more than maximum flexibility. For investors and infrastructure buyers, the most accurate interpretation is an early hyperscaler validation of Positron’s technology and a potential beachhead in a market still dominated by Nvidia.

What Oracle and Positron have actually agreed to

According to Positron CEO Mitesh Agrawal, Oracle is deploying multiple tens of millions of dollars’ worth of Positron systems and racks into OCI. The deployment is intended primarily for AI inference, including mixture-of-experts (MoE) workloads.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That description supports treating the arrangement as a commercial deployment or sale rather than only a research collaboration or laboratory evaluation. However, the public information does not disclose the exact purchase price, number of racks or chips, deployment locations, contract duration, exclusivity, committed capacity or revenue that Oracle may recognize from the relationship. The details come principally from Positron’s account in an EE Times interview.

#1 Best Overall
Sale
HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
  • NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
  • 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
  • PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
  • NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
  • Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads

Oracle’s own materials identify Positron as one of the accelerator options available through OCI alongside Nvidia, AMD and Cerebras. That makes Oracle more than a customer: it can act as a distribution channel through which enterprise users access emerging hardware without buying and operating the systems themselves.

Positron also identifies Jump Trading as a customer, while Cloudflare and Crusoe have been described as proof-of-concept customers. The public material does not establish that the latter relationships have become large-scale revenue deployments.

For financial readers, the distinction matters. A large hyperscaler deployment is stronger evidence than a benchmark slide or funding announcement, but it still does not prove recurring revenue, profitable production or broad customer adoption.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why inference is the target

Training creates or fine-tunes AI models. It generally rewards highly programmable, massively parallel hardware and a mature software ecosystem.

Inference runs an already-trained model to generate a response or prediction. Its economics are often measured using:

  • tokens per second;
  • latency and response consistency;
  • concurrent users or requests;
  • memory capacity and bandwidth;
  • energy consumed per token;
  • hardware utilization; and
  • total cost per request or token.

Inference can therefore favor specialized systems designed around a narrower set of transformer workloads. A system that is less flexible than a general-purpose GPU may still be attractive if it delivers lower operating costs for a predictable production workload.

This is not an “inference replaces training” story. The market is likely to remain a both-and environment: GPUs continue to matter for foundation-model training, post-training and mixed workloads, while inference demand grows as AI applications reach more users.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
MX3 M.2 AI Accelerator
  • High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
  • Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
  • Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
  • Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
  • Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.

What Positron sells today: Atlas

Positron’s first-generation product, Atlas, is an inference accelerator system based on FPGA technology. Positron says Atlas is shipping and supporting production inference workloads.

The company claims Atlas delivers:

  • 3.5 times better performance per dollar than Nvidia’s H100 in its target workloads;
  • up to 66% lower power consumption than an H100; and
  • up to three times more tokens per watt than existing GPUs.

These are company or investor-backed claims, not independently verified results established by the sources available for this article. A meaningful comparison would need to specify the models, model sizes, quantization, batch size, latency target, concurrency, software stack, system configuration and pricing assumptions. The claims should not be generalized to every AI workload or treated as evidence that Atlas is a faster or cheaper replacement for Nvidia across the market.

What Positron is planning: Asimov and Titan

Positron’s next-generation architecture is called Asimov, while Titan is the planned system built around that silicon.

In its February 2026 Series B announcement, Positron said Asimov is targeted for tape-out in late 2026 and production in early 2027. The company says the accelerator is designed to support roughly 2 terabytes—or, in the announcement’s configuration, more than 2.3 terabytes—of memory per device. Titan systems are described as offering approximately 8TB of memory per system.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Those figures are roadmap claims, not specifications of a generally available production product. Final silicon, pricing, availability, manufacturing yield and real-world performance could differ.

Positron has also described an Asimov target of approximately 450–500 watts per chip and a roughly 3-kilowatt system. The planned design is expected to use TSMC’s N3P process, an organic substrate and attached LPDDR memory rather than a conventional CoWoS-plus-HBM design.

The technical argument against Nvidia

Memory capacity can be as important as compute

Large models, long context windows and MoE architectures can require substantial memory. Positron argues that putting more memory close to the accelerator can reduce model sharding and improve inference economics.

Rank #3
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
  • ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
  • ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
  • ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
  • ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C

Its Series B materials compare a claimed more-than-2.3TB Asimov configuration with 384GB for Nvidia’s forthcoming Rubin GPU. That comparison requires caution: the memory technologies, bandwidth, interconnects and system configurations may differ. Capacity is not the same as usable bandwidth, and more memory does not automatically produce faster inference.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Realized bandwidth matters

Positron says its architecture can achieve more than 90% memory-bandwidth utilization, compared with roughly 20%–50% for typical best-in-class inference workloads on other systems. Utilization varies substantially with model architecture, sequence length, batch size, quantization, operator fusion, concurrency, compiler behavior and whether a workload is compute- or memory-bound.

As a result, a buyer should ask for independently reproducible results at realistic concurrency and an equivalent latency or quality target—not only peak bandwidth numbers.

Power and cooling can determine whether capacity is usable

Data-center operators may have available floor space or electrical capacity but be unable to support the rack densities and liquid-cooling requirements associated with newer high-end GPU systems. Positron is targeting racks in the 15–30kW range and says its planned systems are intended to remain air-cooled.

That is a practical competitive angle. The question is not simply which accelerator is fastest. It is whether a system can turn existing, underused data-center capacity into profitable inference capacity without expensive electrical and cooling upgrades.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why Oracle matters

Oracle is rapidly expanding its infrastructure cloud business. In its FY2026 results, Oracle reported $18.1 billion in cloud-infrastructure revenue, up 77%, while fourth-quarter cloud-infrastructure revenue reached $5.8 billion, up 93%. Oracle also said much of its large AI-contract growth involved customer prepayments or customer-supplied GPUs.

Those figures explain why OCI may want a broad accelerator portfolio. Offering multiple hardware options can help Oracle serve different price, availability, power and workload requirements while reducing dependence on a single supplier.

Rank #4

They do not show that Positron caused any specific portion of Oracle’s revenue growth. Nor do Oracle’s public materials establish that Positron hardware is replacing Nvidia hardware. The more measured interpretation is that OCI is using a multivendor strategy and Positron has earned a place in that portfolio.

Positron versus Nvidia: the comparison that matters

Category Positron Nvidia
Primary target Specialized AI inference, particularly predictable transformer and MoE workloads Training, inference and broad mixed workloads
Current product status Atlas is described by Positron as shipping; Asimov and Titan remain roadmap products Large, established portfolio of production accelerators and systems
Architecture thesis Specialization, high memory capacity and lower power for selected inference workloads General-purpose programmability, high performance and a broad platform ecosystem
Software position Must prove model compatibility, tooling and migration economics for customers used to CUDA CUDA, libraries, frameworks, networking and extensive developer familiarity
Cooling and density Targets lower-power, potentially air-cooled deployments High-end systems can require substantial rack power and advanced cooling
Customer access Enterprise inquiry or cloud availability, depending on deployment Direct systems, cloud providers and a large hardware partner ecosystem

This is not a like-for-like comparison. Positron is not attempting to reproduce Nvidia’s entire training, networking, software and support platform. Its stated strategy is coexistence: win workloads where specialization and efficiency matter more than maximum flexibility.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is this a serious Nvidia threat?

Commercially, yes—at an early stage. A hyperscaler deployment reportedly worth tens of millions of dollars is meaningful validation and a stronger signal than a prototype demonstration.

As a market challenge, not yet. Positron’s own CEO has described Nvidia, Google, AMD and AWS as controlling approximately 99.9% of data-center AI silicon, leaving startups with very small shares. That estimate is the company’s characterization, not an independently verified market measurement, but it illustrates the scale of the incumbent advantage.

Nvidia’s moat includes more than chips. It includes CUDA, libraries, compilers, networking, model support, developer tools, supply relationships, third-party support and operational familiarity. A specialized accelerator must deliver enough savings to compensate for software migration, model retuning, procurement risk and the possibility that a customer’s workload changes.

The Oracle relationship improves Positron’s credibility because a cloud provider can absorb some of that complexity and make the hardware available as a service. It does not eliminate the underlying platform challenge.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Funding shows confidence, not commercial proof

Positron announced a $51.6 million Series A in July 2025, bringing its disclosed capital raised that year to more than $75 million. In February 2026, it announced a $230 million Series B at a post-money valuation above $1 billion.

Best Value
ASRock Radeon AI PRO R9700 Creator 32GB Professional Graphics Card, 2920 MHz Boost Clock, GDDR6, AMD RDNA 4, AI-Accelerators, DisplayPort 2.1a, PCIe 5.0, Blower Cooler
  • Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
  • Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
  • Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
  • Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
  • Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.

The Series B investors included ARENA Private Wealth, Jump Trading, Unless, Qatar Investment Authority, Arm, Helena and existing investors. This financing gives Positron capital to develop Asimov, build systems and pursue deployments.

It does not establish sustainable revenue, gross margins, manufacturing yields, repeat orders or a durable competitive advantage. Those are the metrics investors will need to watch as the company moves from early deployments to scaled production.

Questions a serious buyer should ask

  • Which models and model sizes are supported today?
  • Which quantization formats and common inference runtimes are available?
  • How well does the software integrate with PyTorch, vLLM, Hugging Face or TensorRT-LLM?
  • How much application or model migration is required from CUDA?
  • Are performance claims measured at comparable latency, quality and concurrency targets?
  • What is the total cost of the system, including hosts, networking, memory, support and software?
  • What is Atlas’s production availability and support model?
  • Can customers purchase systems directly, or must they access them through a cloud provider?
  • What is the supply, yield and customer availability outlook for Asimov?
  • Does the Oracle deployment represent paid production capacity, a pilot or a broader supply agreement?

Oracle’s public materials show a multivendor accelerator strategy, but readers should not assume that a Positron-backed instance can be selected in the OCI console unless Oracle publishes a specific service or instance type. Public Positron hardware pricing was not disclosed in the cited materials.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What this means for investors and cloud buyers

For investors, the Oracle deployment is evidence that Positron has moved beyond fundraising and laboratory claims toward commercial infrastructure deployment. The next proof points are more demanding: repeat orders, production availability, software maturity, utilization, gross margins and revenue that scales beyond a small number of customers.

For cloud buyers, Positron could be attractive when the workload is inference-heavy, repeatable and constrained by power, cooling or memory capacity. Nvidia remains the safer choice for organizations that need broad programmability, training, rapidly changing models or a mature CUDA-centered operating environment.

The central question is therefore not whether Positron is “the next Nvidia.” It is whether specialized inference systems can lower the cost of serving AI enough to earn a growing share of cloud capacity.

Quick Recap

Bestseller No. 2
MX3 M.2 AI Accelerator
MX3 M.2 AI Accelerator
Software and Documentation can be accessed at the MemryX developer website
$169.00
Bestseller No. 3
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
✅Scalable, enabling simultaneous processing of multi-streams & multi-models; ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
$219.99
Bestseller No. 4
Tesla L40S 48GB AI HPC Graphics Accelerator
Tesla L40S 48GB AI HPC Graphics Accelerator
48GB AI graphics accelerator
$5,999.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Written by TheFinanceBase Team

The Team behind TheFinanceBase.

Add your note

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.