October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
The Finance Base
AI inference

DeepInfra Emerges From Stealth With $8 Million to Make AI Inference More Affordable

DeepInfra raised $8 million in 2023 to make open-model inference cheaper. Here is what it launched, how the cost thesis worked and what the company reports in 2026.

By TheFinanceBase Team 7 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DeepInfra emerged from stealth on November 9, 2023, with an $8 million seed round led by A.Capital and Felicis. Founded by former IMO Messenger engineers, the company set out to host open-source models such as Meta’s Llama 2 and Code Llama at a lower serving cost than leading proprietary APIs. Its underlying bet was that running models for real users—not only training them—would become a major infrastructure expense.

That bet has since expanded. DeepInfra’s 2026 materials describe an OpenAI-compatible inference cloud covering language, vision, embeddings, image, video and speech models, plus private deployments and GPU rental. The company reported a $107 million Series B in May 2026, nearly five trillion tokens processed weekly and more than 190 open-source models. Those scale figures are company claims; the original “cheaper inference” comparison was a dated 2023 snapshot, not a current guarantee.

What DeepInfra announced in November 2023

DeepInfra’s launch was reported on November 9, 2023, when it announced an $8 million seed financing. A.Capital and Felicis led the round, with Georges Harik and SVA also reported as participants. The founding team came from IMO Messenger, where it had worked on large distributed systems.

The initial product was hosted inference for open-source or open-weight machine-learning models. Launch coverage named Meta’s Llama 2, Code Llama, variants, tuned models and other newly released systems. Instead of asking every developer to acquire GPUs, load model weights and operate an inference stack, DeepInfra offered an API that handled serving.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
  • ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
  • ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
  • ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
  • ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C

VentureBeat’s launch report also quoted CEO Nikola Borisov saying prompts were not stored or used. That was a launch-era statement; it should not be read as a blanket description of every current product, log or account record.

Read the November 2023 launch report.

Inference is the cost of answering users

Training adjusts a model’s parameters. Inference runs those trained parameters to answer a prompt, classify a document, generate an image, transcribe audio or perform another production task.

A single demonstration may use one GPU briefly. A production service must handle simultaneous requests while meeting latency and availability targets. Large models consume substantial GPU memory; long contexts require more memory bandwidth; generated responses require repeated computation for each token. Applications can also make multiple chained calls for retrieval, tool use or agentic workflows.

The systems problem is therefore not simply “how many GPUs are available?” It is how effectively a provider places concurrent users and model executions on that hardware without creating queues or unacceptable tail latency. Scheduling, batching, model loading, networking, memory capacity and failure handling all affect the cost of a useful response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DeepInfra’s original cost thesis

The November 2023 coverage cited a headline rate of $1 per million input or output tokens for DeepInfra. The same report compared that figure with $10 per million tokens for GPT-4 Turbo and $11.02 per million tokens for Claude 2.

Provider or model Historical price cited in November 2023
DeepInfra $1 per million input or output tokens
OpenAI GPT-4 Turbo $10 per million tokens
Anthropic Claude 2 $11.02 per million tokens

These are historical figures from the launch coverage, not current quotes. They were not an independent, like-for-like benchmark: the models could differ in capability, context limits, input/output pricing rules, latency, support and reliability. A lower token rate also does not automatically mean a lower application bill. More retries, longer responses, weaker task performance or additional engineering can erase a nominal saving.

Rank #2
Hailo-8 M.2 AI Accelerator Module 26TOPS Hailo8 Support Linux/Windows
  • Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor.
  • 2.5W typical power consumption
  • Enabling real-time low latency and high-efficiency AI inferencing on the edge devices
  • Supports TensorFlow TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • Supports Linux and Windows.

DeepInfra’s proposed advantages were operational. It would run and optimize shared server fleets, use the founders’ distributed-systems experience, place many requests efficiently on the same hardware and make popular open models available through a common service. The launch material discussed concurrency, computation per token, memory bandwidth and avoiding redundant work, but did not publish an audited cost model, utilization rate or independent latency benchmark. Claims about particular techniques such as continuous batching or speculative decoding should not be inferred from that article.

Why open models were central to the strategy

Open or open-weight models can give developers more control over deployment, fine-tuning and model choice. They can reduce dependence on one proprietary provider and, in some workloads, lower serving costs. A hosted catalog also lets a team test a new model or a tuned variant without building a new GPU stack.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“Open source” is not one legal category. Weights, training data, commercial-use permissions, attribution duties and acceptable-use rules vary by model. A buyer must review the exact license and the hosted model version before using it in a commercial or regulated system.

How the platform changed after the seed round

Date Reported development
2022 DeepInfra says the company was founded.
November 9, 2023 Stealth launch and $8 million seed round.
May 4, 2026 $107 million Series B announcement.
August 2026 positioning OpenAI-compatible inference cloud, model hosting, private deployments and GPU infrastructure.

DeepInfra’s current documentation describes more than 100 language models and a broader catalog that includes vision and OCR, embeddings and rerankers, image and video generation, speech recognition and text-to-speech. It also documents private deployment for customer-owned or fine-tuned models and GPU instances or clusters.

In its May 4, 2026 Series B announcement, the company said it supported more than 190 open-source models, processed nearly five trillion tokens per week, operated GPU infrastructure across eight U.S. data centers and was expanding internationally. These are company-reported metrics, not independently audited measurements.

See the company’s Series B announcement.

Current pricing is model-specific

DeepInfra’s pricing page says language models may be billed by input and output tokens, while many other models are billed by inference execution time. It also says there are no long-term contracts or upfront costs. The following examples were visible on August 18, 2026 and can change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
  • NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
  • 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
  • PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
  • NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
  • Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
Model Input price per million tokens Output price per million tokens
DeepSeek-V4-Flash-0731 $0.08 $0.18
DeepSeek-V4-Pro $1.30 $2.60
Llama 4 Scout $0.10 $0.30
Llama 4 Maverick $0.20 $0.80
Qwen3.6-35B-A3B $0.10 $0.95
Gemma 4 26B A4B $0.07 $0.34

These rates do not establish that DeepInfra is the cheapest provider for every model or workload. Check current input, output, cached-token, execution-time, batch and dedicated-GPU charges before forecasting spend.

Using the OpenAI-compatible API

DeepInfra documents an OpenAI-compatible base URL: https://api.deepinfra.com/v1/openai. A simple migration can involve changing the API key and base URL, but compatibility is not guaranteed for every feature. Test structured outputs, tool calls, streaming events, audio, embeddings, batch operations, error formats and provider-specific parameters in the exact SDK path your application uses.

  1. Create a DeepInfra account and generate an API key in the dashboard.
  2. Store the key as an environment variable:
    export DEEPINFRA_TOKEN="your_token_here"
  3. Send a chat-completion request:
    curl "https://api.deepinfra.com/v1/openai/chat/completions" 
      -H "Content-Type: application/json" 
      -H "Authorization: Bearer $DEEPINFRA_TOKEN" 
      -d '{
        "model": "deepseek-ai/DeepSeek-V3",
        "messages": [{"role": "user", "content": "Hello!"}]
      }'

The official quick-start guide also shows OpenAI Python and Node.js clients configured with the same base URL.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Privacy, compliance and private deployments

DeepInfra advertises zero data retention, SOC 2 and ISO 27001 certification, secure U.S.-based data centers, dedicated deployments and autoscaling. These are company statements. A procurement review should establish the scope of each certification, the applicable contract terms, data residency and the controls available to your account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“Zero retention” should not be simplified to “DeepInfra stores no data.” Review how the policy treats prompts and outputs separately from account, billing, security, abuse-prevention and operational metadata, as well as subprocessors, support records, batch jobs and private endpoints. The data-privacy documentation is the appropriate starting point.

Shared public inference generally offers the simplest and lowest-friction path. A private deployment can run fine-tuned or customer-owned weights behind a private endpoint with autoscaling, but it can add GPU-hour charges, capacity planning, cold starts, packaging work and a larger minimum commitment. DeepInfra’s private-model documentation is at docs.deepinfra.com/private-models/overview.

Who should consider DeepInfra?

Potentially strong fits

  • Developers wanting many open models behind a familiar API.
  • High-volume applications where model-specific token rates materially affect costs.
  • Teams that need to switch among language, vision, embedding, speech or generation models.
  • Enterprises evaluating private endpoints for fine-tuned or proprietary weights.
  • Organizations that want hosted inference without operating GPUs.

Cases requiring caution

  • Applications dependent on one proprietary model’s exact behavior.
  • Highly regulated workloads that have not completed a contractual privacy and compliance review.
  • Latency-sensitive services that have not measured regional performance at production concurrency.
  • Teams unable to absorb model-version changes, deprecations or provider-specific API differences.
  • Buyers whose procurement policy requires a hyperscaler-only agreement.

How to evaluate the economics before switching

  1. Choose three representative prompts, including your longest realistic context and output.
  2. Pin the same model version where each provider allows it.
  3. Measure time to first token, total latency, tokens per second, error rates and cost.
  4. Repeat the test at expected concurrency, including bursts and long-running requests.
  5. Include retries, fallbacks, monitoring and engineering time in total cost.
  6. Confirm retention, training-use, encryption, access-control, residency and subprocessor terms.
  7. Check the model license, tool-calling behavior, structured-output support and deprecation policy.
  8. Test rate limits, 429 responses, 5xx recovery and regional failover before production launch.

Reasonable alternatives include Together AI, Fireworks AI, Replicate, Hugging Face Inference Endpoints, RunPod, AWS Bedrock, Google Vertex AI, OpenRouter and self-hosting with software such as vLLM. Their current prices and terms should be checked directly rather than assumed from DeepInfra’s rates.

What the $8 million seed round ultimately represented

The seed financing backed a specific infrastructure thesis: once capable open models became widely available, the difficult and recurring expense would be serving them reliably to many users. DeepInfra’s later funding, broader catalog and reported production volume suggest that inference grew into the larger market the founders anticipated. They do not, by themselves, prove that every workload is cheaper or faster on the platform.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a buyer, the durable lesson is to treat token price as one input to a systems decision. Model quality, concurrency, tail latency, reliability, licensing, privacy and operational control determine whether a low advertised rate becomes a real saving.

Frequently Asked Questions

When did DeepInfra launch?

DeepInfra emerged from stealth on November 9, 2023, announcing an $8 million seed round led by A.Capital and Felicis.

Is DeepInfra’s 2023 $1-per-million-token price still current?

No. That was a historical launch figure. Current prices vary by model and billing method; DeepInfra’s pricing page should be checked for a dated quote.

Does OpenAI compatibility make DeepInfra a drop-in replacement?

It can simplify basic chat-completion migration, but feature behavior varies. Test the SDK features, streaming, tools, structured outputs, limits and errors your application actually uses.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Money Desk

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.