Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

AWS S3 Vectors: What the 90% Cost-Savings Claim Means for Vector Databases

S3 Vectors can lower the cost of storing and querying large vector collections, but AWS’s “up to 90%” claim is workload-dependent. Here’s how to evaluate the real trade-offs.
From TheFinanceBase Team11 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AWS says Amazon S3 Vectors can cut vector upload, storage, and query costs by up to 90% compared with specialized vector-database solutions. That is an AWS claim, not a guaranteed saving or a like-for-like benchmark against every provider. S3 Vectors became generally available on December 2, 2025, and AWS’s own positioning points to a narrower conclusion: it can be a low-cost home for large vector collections, while a dedicated search service may still be needed for demanding applications.

What Amazon S3 Vectors is—and is not

S3 Vectors is a distinct vector-storage service within Amazon S3, not ordinary object storage that happens to hold vector files. It uses vector buckets and vector indexes, with APIs to insert, retrieve, delete, list, and search vectors. A search compares numerical embeddings for similarity and can apply metadata filters.

The service stores and searches embeddings; it does not generate them. An embedding model or another system must turn text, images, code, or other content into vectors before ingestion. AWS controls access through IAM and resource policies in the s3vectors service namespace. That makes the service AWS-specific: applications using it depend on AWS resource types, APIs, regions, and permissions.

What GA delivered, and what changed afterward

At general availability on December 2, 2025, AWS announced support for as many as 2 billion vectors in an index and 20 trillion in a vector bucket. AWS described latency as around 100 milliseconds or less for more frequent queries and subsecond for infrequent ones; these are product statements, not a universal latency guarantee or a stated p95/p99 service-level objective. GA also brought up to 1,000 single-vector PUT transactions per second, up to 100 results per query, availability in 14 Regions, and integrations with Bedrock Knowledge Bases and OpenSearch Service, plus CloudFormation, PrivateLink, and resource-tagging support. AWS’s GA announcement describes the launch capabilities.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AWS expanded the query-result limit in June 2026: an application can request up to 10,000 results, delivered in pages of up to 100. That is a maximum retrieval allowance, not a promise that returning thousands of matches is cheap or useful for a particular application. AWS also announced query data-processing charge reductions of up to 80% for indexes containing more than 10 million vectors. The change applies automatically in S3 Vectors Regions, and AWS still recommends distributing vectors across multiple indexes for query performance. See the result-limit announcement, documented limits, and query-charge announcement.

What “up to 90% lower cost” actually means

AWS’s claim covers uploading, storing, and querying vectors against specialized vector-database solutions. “Up to” matters: AWS does not say every workload will save 90%, and the published claim is not an independently measured comparison with Pinecone, Weaviate, Milvus, OpenSearch, or pgvector under identical conditions. AWS’s pricing page publishes service rates and illustrative calculations; the bill for a real system depends on its data shape and serving requirements.

The economic distinction is that S3 Vectors charges for logical vector data and API usage rather than requiring customers to provision a dedicated database cluster to hold all vectors in a hot serving tier. AWS manages the underlying service infrastructure, which can also reduce customer work provisioning, patching, resizing, and replicating database nodes. It does not eliminate the rest of the application’s costs or operating responsibilities.

The three direct billing components

  • PUTs: Charges are based on logical gigabytes uploaded. Logical size includes vector data, metadata, and the vector key.
  • Storage: Charges are based on total logical vector storage.
  • Queries: Charges include a per-query API fee, data processed (tied substantially to index and vector size), and data returned (based on returned key and metadata bytes).

As examples on the AWS pricing page for US East (N. Virginia), the listed rates are $2.50 per million query requests, $0.004/TB for the first 100,000 vectors, $0.002/TB for 100,000 to 10 million vectors, and $0.0004/TB above 10 million vectors for query processing. Returned data is listed at $0.01/GB, with the first 500 KB returned per query free; at least 256 bytes per result count toward returned-data accounting. AWS also lists $0.06 per GB-month for storage and $0.20 per GB of PUTs in its current pricing signal. These are AWS-published US East examples, not universal rates; check the live page for the Region and pricing applicable to your account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AWS’s pricing page illustrates one workload with 10 million vectors split among 40 indexes, one million monthly queries, and a six-month refresh cycle, totaling $11.38 per month for S3 Vectors under its stated assumptions. A separate AWS example with 500 million stored vectors and 10 million monthly queries totals $1,320.47 per month under that example’s assumptions. Neither figure is a competitor comparison or a forecast for another workload; the latter also shows why query processing can materially change the economics at scale.

The costs outside the headline

A complete bill may include embedding generation, application compute, data transfer, Bedrock or OpenSearch, replication, backups, exports, caching, index rebuilds, monitoring, and engineering labor. If S3 Vectors needs a second service to hit the application’s latency, throughput, or search-feature requirements, that service belongs in the comparison. A managed vector database may have a higher visible service charge but include serving capabilities that the S3 design would otherwise need to add.

Why AWS calls it “complementary”

“Complementary” describes a real architectural split, not simply a claim that S3 Vectors cannot compete. A vector system has separable jobs: retaining the corpus, maintaining indexes, executing similarity searches, filtering and ranking results, and serving application traffic. S3 Vectors is aimed at durable, economical storage and retrieval where queries are less frequent; other search systems can provide a hotter serving tier or richer search behavior. AWS explicitly contrasts S3 Vectors’ lower-throughput, sporadic-query profile with services aimed at high-throughput, low-latency operations in its integration guidance.

Keep the full corpus in S3 Vectors and serve hot data elsewhere

An organization can retain a large, durable corpus in S3 Vectors and keep the subset with strict response-time or throughput needs in a dedicated vector database or search service. This cold/hot pattern avoids paying hot-tier costs for every historical or rarely retrieved embedding. It adds synchronization, routing, and lifecycle work, so the storage savings should be compared with the cost of maintaining that second copy and serving path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use S3 Vectors with OpenSearch

AWS documents S3 Vectors as a lower-cost vector-storage engine for OpenSearch integrations. OpenSearch can supply search-oriented APIs and capabilities such as hybrid lexical/vector search, advanced filtering, aggregations, and faceting. The pattern can suit a team that needs those features but does not want the entire vector corpus to occupy the more expensive serving footprint. It is not equivalent to replacing OpenSearch with S3 Vectors: OpenSearch and its compute still have a cost.

Use it as a Bedrock Knowledge Bases vector store

S3 Vectors can serve as the vector store for Amazon Bedrock Knowledge Bases in an AWS RAG workflow. That can simplify a managed AWS architecture, but vector storage is only one cost line: include embedding models, ingestion, Bedrock retrieval or generation, and any application components in the estimate.

Use one service when the workload is simple enough

For modest-query-rate document retrieval or semantic search, S3 Vectors may be sufficient on its own. In that case, its value is not merely replacing an existing database; it may avoid adding a separately managed vector-search service. Validate actual response times and retrieval quality against the application’s needs before committing.

Which workloads are a plausible fit?

S3 Vectors is strongest where the corpus is large, long-lived, and not queried continuously at high rates. Examples include enterprise document collections, RAG corpora, semantic search across archival material, and image, video, code, or legal-document similarity use cases. Recommendation or personalization workloads may fit when their latency and QPS requirements are moderate rather than extreme. An AWS-native data platform may also benefit when S3, IAM, Bedrock, or OpenSearch already form part of the architecture.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Choose it when durable storage cost matters more than maximum search throughput.
  • Consider it when queries can tolerate AWS’s stated subsecond or roughly 100-millisecond-class performance profile rather than requiring consistently tighter latency.
  • It is a stronger candidate for straightforward similarity search and metadata-filtered retrieval than for a broad search application with complex ranking and analytics.
  • It can make sense as a low-cost source of truth even if a separate system serves the hottest records.

When a dedicated vector or search database may still be worth paying for

Look closely at alternatives if retrieval latency is a product differentiator, traffic is consistently high or bursty, or the application needs predictable tail latency. S3 Vectors’ stated performance figures should not be treated as a p99 guarantee. Also assess whether the application depends on hybrid lexical/vector ranking, aggregations, faceting, advanced ranking pipelines, rich index tuning, or database-specific query APIs.

Other warning signs include heavy concurrent writes near documented per-index limits, complex tenant-isolation needs, frequent large rebuilds, or a requirement for portability across cloud providers. These are selection risks, not proof that S3 Vectors cannot work; validate them with a representative proof of concept. AWS’s integration guidance is useful precisely because it places the service alongside other AWS options rather than presenting one search profile for every workload.

Limits and design choices that affect cost

Documented service limits shape what can be built and how it should be partitioned. AWS currently lists the following limits in its S3 Vectors limitations documentation.

Capability Documented limit
Vector buckets per Region, per AWS account 10,000
Vector indexes per vector bucket 10,000
Vectors per index 2 billion
Dimensions per vector 1–4,096; vectors in an index must share a dimension
Metadata per vector 40 KB total; up to 50 keys
Filterable metadata per vector 2 KB
Non-filterable metadata keys per index 10
Combined PutVectors and DeleteVectors requests per index 1,000 per second
Combined vectors inserted and deleted per index 2,500 per second
Request payload 20 MiB
Vectors per PutVectors call 500
Vectors per DeleteVectors call 500
Vectors per GetVectors call 100
Maximum topK requested 10,000
Results per query-response page 100

Vectors are 32-bit floating-point values, and an index uses one dimension and one distance metric (cosine or Euclidean). Keys can be up to 1,024 characters. These constraints matter when estimating storage and deciding how to map tenants, data types, or update patterns to indexes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Index partitioning changes both cost and performance

Because query processing depends substantially on the queried index’s size and vector size, two layouts containing the same total vectors can produce different query costs. One large index is simpler to organize and may avoid application-level fan-out, but each query can process a larger index. Multiple smaller indexes can reduce per-query processing and improve performance, at the cost of more index management and potentially querying multiple indexes. Tenant-per-index isolation has similar trade-offs. AWS’s June 2026 pricing update lowered processing charges for large indexes but retained its recommendation to distribute vectors across indexes for performance.

Metadata, pagination, and churn are not free details

Only filterable metadata can be used in filter predicates, and its per-vector limit is tighter than the overall metadata allowance. Keep large document chunks or contextual payloads out of filterable fields unless they are needed for filtering. AWS describes filtering as part of vector search, not merely a simple post-search operation; see its metadata filtering documentation. Exceeding supported metadata limits can cause a 400 Bad Request.

A request for many matches can require pagination: up to 10,000 results may be requested, but each response page contains at most 100. Applications need to handle continuation tokens, and returned key and metadata bytes contribute to returned-data charges. AWS’s query documentation describes the query flow.

Overwritten or deleted vectors disappear from query results immediately, but AWS says their storage may take up to a day to be reclaimed. Frequently overwriting the same keys can therefore temporarily affect storage and query-cost calculations; this behavior is noted on the S3 pricing page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to compare the bill fairly

Compare the least expensive architecture that meets the same workload requirements—not S3 Vectors storage alone against a production database that includes serving capacity, replication, and operational tooling. Likewise, do not add a full hot-serving system to an S3 estimate unless the application actually needs one.

  1. Describe the corpus: Count vectors, dimensions, precision, key length, filterable and non-filterable metadata, and indexes. State whether vectors are added, overwritten, or deleted regularly.
  2. Describe retrieval: Estimate monthly queries, average topK, returned metadata volume, QPS, concurrency, and latency targets. Include any filtering, hybrid ranking, or analytics requirements.
  3. Model the direct service charges: Use the relevant Region’s current PUT, storage, per-query, processing, and returned-data rates, including index partitioning and the effect of the 500 KB per-query free returned-data allowance.
  4. Add the surrounding architecture: Include embeddings, transfer paths and Regions, application servers, Bedrock, OpenSearch, caches, replication, disaster recovery, exports, monitoring, and index rebuilds.
  5. Price operations and migration: Account for engineering and on-call work, reindexing, API changes, metadata-schema conversion, and how embeddings and records can be exported or recreated if you later move providers.
  6. Test the workload shape: Benchmark representative data and queries for latency, throughput, retrieval quality, and failure behavior. Do not infer production performance from the maximum index size or AWS’s stated latency descriptions.

This framework prevents a misleading comparison in either direction. A high-availability vector database may look expensive if its replicas and serving compute are compared only with S3 storage. Conversely, an S3 design may lose its advantage when OpenSearch, caching, or another high-performance serving tier is required.

Which architecture fits which requirement?

Requirement Likely fit Why
Lowest durable cost for a large vector corpus S3 Vectors Purpose-built vector storage with usage-based query and upload charges.
Infrequent or moderate retrieval S3 Vectors Fits AWS’s stated lower-throughput, sporadic-query profile.
Bedrock-native RAG storage S3 Vectors Can act as the vector store for Bedrock Knowledge Bases.
Hybrid search, aggregations, or faceting OpenSearch or another full search engine These are broader search capabilities, not just vector storage.
High-QPS, latency-sensitive serving Dedicated vector or search database Evaluate serving performance and controls against the application’s targets.
Multi-cloud portability Vendor-neutral or self-managed database S3 Vectors uses AWS-specific APIs, resources, and access controls.
Large cold corpus plus a hot subset S3 Vectors plus a serving tier Retain long-tail vectors economically while serving latency-critical vectors elsewhere.

Does this threaten standalone vector-database vendors?

It can pressure vendors on the economics of storing very large collections that are queried infrequently: customers need a credible lower-cost tier if the alternative is paying to keep every vector continuously available in a specialized serving system. But storage price alone does not erase the value of fast serving, mature APIs, operational controls, hybrid retrieval, or portability.

A VentureBeat report described AWS’s complementary framing and differing analyst views of the competitive threat; that is market interpretation, not evidence that S3 Vectors has displaced a particular vendor or workload. See VentureBeat’s coverage. The more useful question for a buyer is whether the corpus needs to be hot all the time, and what the cheapest architecture is that still meets its latency, throughput, filtering, durability, and operational requirements.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase07 MAR 2625 minWhat Is a 457 Plan?
  2. The Money DeskBlogTheFinanceBase07 MAR 2621 minTime Value of Money: What It Is and How It Works
  3. The Money DeskBlogTheFinanceBase07 MAR 2627 minAre You Living in One of These Top 10 Most Expensive Cities to Retire?
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.