Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Speedata announced a $44 million Series B on June 3, 2025, alongside the commercial launch of its C200 Analytics Processing Unit (APU), a PCIe accelerator built for Apache Spark, batch ETL and AI data preparation. The Tel Aviv-based startup says the round brings its total funding to $114 million. Its product is aimed at a narrower job than Nvidia GPUs: speeding up analytics data processing, potentially alongside GPUs used for AI training and inference.
What Speedata announced
The $44 million Series B was announced on June 3, 2025, at the same time Speedata introduced its first commercial APU. Existing investors Walden Catalyst Ventures, 83North, Koch Disruptive Technologies, Pitango First and Viola Ventures participated, along with strategic investors Lip-Bu Tan and Eyal Waldman. Speedata was founded in 2019 and is based in Tel Aviv, according to TechCrunch.
The public announcement date is not necessarily the round’s closing date. Calcalist reported that the financing had been completed roughly six months earlier; that timing is a separate report, not the announcement date. Speedata did not disclose a specific allocation of the new capital in the cited launch materials, so it should not be assumed that the money is earmarked for any particular expansion.
What the APU is designed to accelerate
Speedata is targeting data-intensive work that often runs on CPUs in analytics clusters: Spark SQL and DataFrame operations, batch extract-transform-load (ETL), Parquet processing, database operations and preparation of data for AI systems. Data preparation is distinct from training or running a neural network. It can involve filtering, joining, decompressing and reshaping large datasets before they reach an analytics application or AI workload.
#1 Best Overall
- NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
- 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
- PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
- NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
- Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
The company’s argument is that a processor designed around these analytics operations can execute them more efficiently than a general-purpose CPU or a GPU adapted to the task. That is a product thesis, not proof that the APU will outperform either processor on every workload. The benefit depends on which operations a job contains, how much of the job can be offloaded, and whether data movement or storage becomes the limiting factor.
Inside the C200 and Dash software
Hardware and deployment
The C200 is a PCIe Gen 5 accelerator card built around Speedata’s Callisto ASIC. Speedata describes it as suitable for standard server integration and advertises a preconfigured 2U server containing two C200 cards. It also lists Dell and HPE as OEM procurement routes. These are advertised deployment options, not a guarantee that the card fits every server or cloud environment; buyers still need to confirm supported configurations, power, cooling and firmware requirements. The company’s product page offers a demo or sales contact rather than a public list price.
Rank #2
- Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
- Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
- Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
- Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
- Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
How Dash fits into Spark
Speedata’s Dash software integrates with Spark’s Catalyst optimizer. It identifies supported, compute-intensive operations for execution on the APU, while unsupported operations—including some user-defined functions—continue on the CPU. Speedata says existing Spark applications can use the accelerator without application-code changes. Its stated integration targets Apache Spark 3.x running with Kubernetes, YARN or standalone cluster managers; buyers should confirm the exact supported versions and configuration with the company. See the technology overview.
“No code changes” does not mean “no deployment work.” Installing and operating the system can still require compatible servers, drivers, runtime software, cluster configuration, monitoring, failure handling and a plan for mixed CPU/APU execution. The amount of work depends on the buyer’s existing environment.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #3
- HIGH COMPATIBILITY: The graphics card supports multiple displays and panels with a maximum resolution of 1920x1440, making it compatible with a wide range of systems for diverse applications.
- QUICK ROTATION: With the ability to quickly rotate screen images at 90°, 180°, and 270°, this graphics card enhances versatility in display orientation for improved user eerience and flexibility.
- POWERFUL 2D GRAPHICS ACCELERATION: Equipped with a robust 2D graphics accelerator, the card supports various graphic processing functions, ensuring efficient performance for demanding applications.
- VERSATILE APPLICATION: This accelerator card supports video display layers, making it ideal for a variety of applications, including industrial computers, POS systems, ensuring reliable performance across different fields.
- WIDE OPERATING TEMPERATURE RANGE: Designed for reliable operation in harsh environments, the card functions effectively within a wide temperature range of -40°C to +85°C, ensuring durability and stability in challenging conditions.
Why Nvidia is in the comparison—and where it stops
Speedata’s launch has been framed as a challenge to Nvidia, but the comparison is limited. Speedata is not presenting the C200 as a general replacement for Nvidia GPUs in model training, inference or the broader AI-computing market. Its pitch is that specialized hardware can accelerate analytics operations such as joins, filtering, decompression and column processing, while leaving GPUs available for workloads better suited to them.
That makes the APU a potential complement to CPUs and GPUs, rather than a like-for-like alternative across Nvidia’s product range. Whether it complements an existing GPU cluster economically is a separate question: a buyer needs comparable end-to-end results on its own workload, not just a fast result for an isolated operation.
Rank #4
- System Compatibility Note: This 2-slot card measures 271 x 112 x 39 mm and requires a single 12V-2x6-pin power connector. Please verify chassis and PSU compatibility before purchase.
- Dedicated Support: Please contact us directly through Amazon for any product questions or assistance you may require.
- Professional Intel Arc Pro B70 GPU: Built on the Intel Xe2-HPG architecture, it features 32 Xe cores and 256 XMX engines, designed to accelerate AI, rendering, and complex visualization workloads.
- Massive 32GB GDDR6 VRAM: Equipped with 32GB of high-speed GDDR6 memory on a 256-bit bus, running at 19 Gbps, which allows for handling large AI models and complex datasets locally.
- High-Performance Engine Clock: Delivers an engine clock of 2540 MHz, providing the compute power needed for demanding professional applications and AI inference.
What Speedata’s performance figures do—and do not—show
Speedata advertises large gains, but the figures below are company claims. Public coverage does not establish that they generalize across Spark workloads or that the comparisons use equivalent, current hardware and software configurations.
| Example | Speedata’s reported result | Evidence and qualification |
|---|---|---|
| Pharmaceutical workload | 19 minutes on the APU versus 90 hours on a nonspecialized processor; described as a 280× speedup | Company result reported by TechCrunch. The public account does not establish the exact comparison hardware or enough configuration detail for independent reproduction. |
| Spark example | 4 minutes 3 seconds versus 13 seconds, approximately 20× | Example on Speedata’s product page; not evidence of a general result across Spark jobs. |
| Selected workloads | Up to 100× faster | Speedata marketing claim on its product page; “up to” is workload- and baseline-dependent. |
| Total cost of ownership | Up to 90% lower | Speedata marketing claim on its product page; a public calculation method is not provided in the cited material. |
| Server consolidation | One deployment reduced 37 servers to 3 | Company-reported example on its product page; the workload, server specifications and utilization assumptions are not stated there. |
| Production Spark workloads | 62.7× acceleration; a separate example reduces a 90-hour job to 8 hours | Current company website claims, not independently validated in the available coverage: Speedata. |
| GPU comparison | 52× more performance per dollar than a GPU on an industry-standard benchmark | Current company website claim: Speedata. The claim should be assessed against the benchmark, GPU model, configuration and cost assumptions, which are not specified in the cited figure. |
To assess a result fairly, a prospective customer would need the dataset and format, Spark and hardware versions, CPU and GPU models, number of C200 cards, query mix, offload rate, and confirmation that the measurement includes storage, networking and data transfer. Power, cooling, software optimization, repeatability and the full cost calculation matter too. A comparison against an unoptimized baseline, or one that measures only accelerated operations, may not predict end-to-end job time or savings.
Best Value
- AI Performance: 767 AI TOPS
- OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode)
- Powered by the NVIDIA Blackwell architecture and DLSS 4
- Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
- A 2.5-slot design maximizes compatibility and cooling efficiency for superior performance in small chassis
Who might benefit—and what could limit the fit
Potentially suitable workloads
- Large, recurring Apache Spark jobs where CPU-based ETL or data preparation is a meaningful bottleneck.
- Clusters with material power, rack-space or server costs and enough sustained usage to justify specialized hardware.
- Applications with substantial supported operations and limited need for custom or unsupported execution paths.
- Organizations that can procure compatible PCIe systems and want to reserve GPUs for training or inference.
Reasons to be cautious
- Neural-network training and inference are not the C200’s stated primary purpose.
- Small or infrequent jobs may not run long enough to offset data-transfer and orchestration overhead or the cost of dedicated hardware.
- Unsupported operators or user-defined functions can remain on the CPU, reducing the share of work accelerated.
- Data skew, storage or network limits can dominate runtime even when compute-heavy operations improve.
- Organizations needing a fully managed public-cloud service should confirm an available deployment route; a PCIe card is not automatically usable in a managed Spark service.
- A specialized chip is less flexible than general-purpose infrastructure if workloads change or compatible software support is limited.
Any total-cost comparison should include the accelerator, host servers, storage and networking, power and cooling, support, software, integration, maintenance, utilization and staff time. A faster job is not necessarily a cheaper system if the hardware sits idle or the deployment adds substantial operating cost.
What remains to be established commercially
The funding announcement and product launch establish that Speedata is bringing a commercial accelerator to market; they do not establish broad production adoption. TechCrunch reported that large companies were testing the APU, but the companies were unnamed. The available public material also does not provide a price, delivery lead time, service-level commitment, detailed production-customer record or payback period.
For a buyer, a practical next step is a workload-specific evaluation. Speedata offers a Workload Analyzer for a browser-based assessment, a local CLI or comparison with TPC-DS benchmarks. It is a vendor evaluation tool, not independent certification. A useful pilot should compare the buyer’s own representative data and current CPU/GPU baseline, measure complete job time and power, and account for hardware, software and operational costs. Confirm supported Spark versions, operators, server configurations, cloud options and support arrangements before treating a speedup as a deployment forecast.
Speedata’s $44 million round funds the commercialization of a focused analytics accelerator, not a broad contest to replace Nvidia in AI computing. Its thesis is plausible for some Spark and ETL workloads; the scale of the business opportunity depends on independently reproducible, end-to-end results and economics in customers’ environments.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




