October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

Modal Labs’ $16 Million Bet on Serverless AI Infrastructure

Modal Labs’ 2023 $16 million Series A backed a platform that abstracts cloud and GPU operations. Here is how Modal’s AI infrastructure, pricing and trade-offs look in 2026.
From TheFinanceBase Team7 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Modal Labs announced a $16 million Series A on October 10, 2023, led by Redpoint Ventures. Amplify Partners, Lux Capital and Definition Capital also participated, bringing the company’s reported funding total to $23 million. Modal’s original pitch was to let developers run data- and compute-heavy code without assembling containers, schedulers, cloud instances and GPU infrastructure themselves. By August 2026, that idea had expanded into a broader serverless AI platform for inference, batch jobs, training, notebooks, sandboxes and web services.

What Modal Labs raised in 2023

TechCrunch reported the $16 million Series A on October 10, 2023. Redpoint Ventures led the round, with Amplify Partners, Lux Capital and Definition Capital participating. The company said the financing brought its total raised to $23 million, following an earlier $7 million seed round.

Modal said it would use the money mainly to hire software engineers. It had 14 employees when the round was announced and planned to reach 17 by the end of 2023. The product was moving from beta toward general availability.

TechCrunch’s funding report also identified Substack and Ramp as customers at the time. Those are historical examples, not a current customer list.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Who founded Modal?

Modal was founded in 2021 by Erik Bernhardsson, who previously led data teams at Spotify and served as CTO of Better.com. Bernhardsson described fragmented, difficult-to-operate data tooling as the motivation for building Modal. That background explains the company’s original focus, but it should not be treated as independent proof that every team needs the same architecture.

What infrastructure Modal abstracts

In a conventional deployment, an engineering team may need separate systems for packaging code, acquiring GPUs, scheduling jobs, exposing HTTP services, collecting logs and storing credentials. Modal puts many of those controls behind a programmatic platform.

  • Container image construction and deployment
  • CPU, memory, disk and GPU selection
  • Scheduling, capacity acquisition and autoscaling
  • Job execution, retries and parallel fan-out
  • HTTP endpoints and long-running services
  • Logs, metrics and deployment controls
  • Secrets and persistent volumes
  • Queues, distributed coordination and notebooks
  • Usage reporting, budgets and spend controls

Modal’s current documentation describes these primitives as part of an AI infrastructure and serverless cloud platform. See the product guide, server documentation, notebook documentation and resource documentation.

“Abstracted away” does not mean infrastructure disappears. Customers still choose resources in code, package models and dependencies, design data flows, control permissions, instrument applications and handle failures. Modal manages much of the underlying operation; it does not remove the need for architecture and governance.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the developer workflow works

The current documented setup is Python-first:

  1. Install the client:
    pip install modal
  2. Create an account and authenticate the CLI:
    modal setup
  3. Write an application that defines an image, resources and functions.
  4. Run it remotely:
    modal run path/to/file.py
  5. Deploy it as an application:
    modal deploy path/to/file.py
  6. Use local hot reload during development:
    modal serve path/to/file.py

Modal’s CLI reference is the authoritative source for command behavior as the package changes.

import modal

app = modal.App("example")
image = modal.Image.debian_slim().pip_install("torch", "numpy")

@app.function(image=image, gpu="A100")
def run():
    import torch
    assert torch.cuda.is_available()
    return torch.cuda.get_device_name(0)

Resource declarations can also specify CPU, memory or multiple GPUs:

@app.function(gpu="H100:8")
def run_large_model():
    ...

@app.function(cpu=8.0, memory=32768)
def process():
    ...

The GPU guide lists hardware including T4, L4, A10, L40S, A100, H100, H200, B200 and B300. Availability, queueing and supported software can change.

How Modal’s product has expanded

The 2023 description centered on big-data workload infrastructure. As of August 18, 2026, Modal’s documentation presents a wider AI compute platform.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inference

Teams can expose GPU-backed model APIs for variable traffic and open-weight models. Modal positions some deployments around sub-second cold starts, but actual latency depends on image transfer, model loading, hardware, region, concurrency and application design.

Batch processing

Independent jobs such as transcription, image generation, embedding, evaluation and document processing can be distributed across many workers. Modal’s examples include batch Whisper transcription and parallel Parquet processing from Amazon S3.

Training and fine-tuning

Temporary access to multiple accelerators can suit experiments and irregular research demand. Distributed training still requires decisions about data locality, checkpointing, inter-node communication, reproducibility and failure recovery.

AI-generated-code sandboxes

Modal now documents isolated sandboxes for running code generated by AI systems. That creates additional requirements around isolation, network access, secrets, abuse prevention and cost limits.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Notebooks and services

Browser-based notebooks provide serverless CPU or GPU compute, collaborative editing and automatic idle shutdown. Servers, endpoints, volumes, queues and secrets extend the same model to longer-lived applications.

What “serverless GPU” means

Modal’s serverless model provisions infrastructure when workloads run instead of requiring customers to reserve a continuously operating GPU machine. It supports scale-to-zero behavior and usage-based billing, although plan fees, requested resources, storage, concurrency and premium execution modes still affect the bill.

  • Serverless execution: Modal provisions and scales the runtime.
  • Scale to zero: Idle workloads can stop consuming compute, subject to the product’s behavior.
  • Managed containers: Customers still define images and dependencies.
  • Serverless GPU: Ephemeral jobs can attach GPUs without the customer operating GPU instances.

This model is most compelling when demand is bursty or unpredictable. A continuously busy service may obtain a lower unit cost from reserved instances, dedicated servers or owned hardware.

Modal pricing as of August 18, 2026

Modal’s pricing page listed the following plan structure when checked on August 18, 2026. Prices, credits, hardware availability and limits are volatile.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Plan Plan fee Included compute credit Seats
Starter $0 per month, plus compute $30 per month Up to three workspace seats
Team $250 per month, plus compute $100 per month Unlimited seats
Enterprise Custom pricing Custom Enterprise limits and support

GPU, CPU, memory and volume charges are metered separately. The same page listed these example GPU rates:

GPU Listed rate per second
NVIDIA B300 $0.001972
NVIDIA B200 $0.001736
NVIDIA H200 $0.001261
NVIDIA H100 $0.001097
NVIDIA A100 80 GB $0.000694
NVIDIA A100 40 GB $0.000583
NVIDIA L40S $0.000542
NVIDIA A10 $0.000306
NVIDIA L4 $0.000222
NVIDIA T4 $0.000164

Modal states that selected regions can add a 1.5×–1.75× multiplier and non-preemptible execution can add a 3× base-price multiplier. Check the live pricing page before budgeting.

How to model the real cost

A GPU-hour comparison alone is incomplete. A practical estimate is:

total cost = GPU runtime
+ CPU runtime
+ memory allocation
+ storage
+ data transfer or external storage
+ plan subscription
+ region or premium-availability charges
+ warm-container time
+ engineering and migration cost

Measure average and peak concurrency, cold-start frequency, model initialization, accelerator utilization, data movement, volume use and regional requirements. Modal’s resource documentation says CPU and memory billing can reflect the higher of requested or actual usage, while storage and other categories are billed separately.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Budgets and reports are available through the platform. For example:

modal billing summary
modal billing report --start 2025-12-01 --end 2026-01-01

Report ranges use an inclusive start and exclusive end; dates default to UTC. See the billing CLI reference and billing guide.

Where Modal fits best

Workload Fit Primary question
Bursty inference Strong Are cold starts and GPU availability acceptable?
Large batch jobs Strong Can data move efficiently into workers?
Prototyping Strong Does the team prefer code over cluster administration?
Continuous, high-utilization GPU service Mixed Would reserved capacity cost less?
Specialized distributed training Mixed Is the required topology and control available?
Strict cloud-native deployment Mixed Is direct AWS, Google Cloud or Azure integration more valuable?
Highly regulated workload Case-specific Do the available controls and attestations meet policy?
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Operational limitations and failure modes

Large images and model files

Oversized images and weights increase build, transfer and cold-start time. Keep runtime images lean and use persistent or cached artifacts where appropriate.

Warm-container and notebook spending

Serverless does not mean free while a process remains active. A running notebook kernel or warm service can continue consuming resources. Modal documents automatic notebook idle shutdown, but teams should still set explicit policies.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

GPU and CUDA mismatch

Models may require a particular memory size, CUDA release or architecture. Modal notes that B300 requires CUDA 13.1 or newer and that Blackwell hardware can have different library support from Hopper GPUs.

Large multi-GPU requests

Modal says requests for more than two GPUs per container will usually wait longer. This matters for large-model serving and distributed training.

Unbounded fan-out

High concurrency, retries, oversized memory requests and accidental persistent services can multiply costs. Use budgets, alerts, bounded queues and workload-level attribution.

Public tunnels

Modal’s tunnel documentation warns that generated tunnel URLs are public on the internet. They should not be treated as authentication or private-network controls. See the tunnel guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Network-bound data jobs

If the bottleneck is object-storage or database transfer, adding GPUs may not help. Measure I/O, serialization, startup and queue latency as well as accelerator utilization.

Alternatives to evaluate

Hyperscaler services

  • AWS Lambda suits event-driven, AWS-native functions but is not a one-for-one replacement for Modal’s GPU, notebook and batch workflow.
  • Google Cloud Run offers managed containers and strong Google Cloud integration.
  • Azure Container Apps fits organizations standardized on Azure identity, networking and billing.

GPU-focused platforms

  • RunPod offers accessible GPU instances and serverless GPU options.
  • Baseten emphasizes production model serving.
  • Replicate provides a model-centric deployment and API experience.
  • CoreWeave targets larger-scale GPU infrastructure and dedicated capacity.

Self-managed infrastructure

Kubernetes with GPU operators, managed Kubernetes, Slurm, direct GPU instances and bare-metal deployments provide more node-level control and may lower unit costs at high utilization. They also transfer provisioning, upgrades, scheduling, security, observability and reliability work to the customer.

Bottom line for prospective users

Modal’s original insight was that teams should describe compute-heavy workloads in application code instead of operating a separate infrastructure stack. The $16 million 2023 Series A funded that early expansion; the current product is broader, covering AI inference, batch processing, training, notebooks, sandboxes and services.

Modal is strongest when demand is variable, GPU access is temporary or the team values operational simplicity over maximum infrastructure control. It is less obviously attractive for permanently saturated workloads, unusual networking or hardware requirements, strict regional constraints, or organizations that already operate an efficient hyperscaler or cluster platform. Evaluate it with a workload model that includes startup behavior, data movement, idle time, plan fees and engineering effort—not just the posted GPU rate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase07 MAR 2625 minWhat Is a 457 Plan?
  2. The Money DeskBlogTheFinanceBase07 MAR 2621 minTime Value of Money: What It Is and How It Works
  3. The Money DeskBlogTheFinanceBase07 MAR 2627 minAre You Living in One of These Top 10 Most Expensive Cities to Retire?
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.