October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
The Finance Base
artificial intelligence

Google’s Gemini 1.5 API price cut exceeded 50%—but the models are now retired

Google’s Gemini 1.5 refresh delivered stable Pro-002 and Flash-002 models, higher paid-tier limits and targeted price cuts exceeding 50%. Those API models were retired on September 29, 2025.

By TheFinanceBase Team 4 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google’s September 24, 2024 Gemini 1.5 update introduced stable gemini-1.5-pro-002 and gemini-1.5-flash-002 models, raised paid-tier rate limits and cut selected Gemini 1.5 Pro API prices by more than 50%. The headline was not a universal 50% discount: for prompts under 128,000 tokens, input prices fell 64%, output prices 52% and incremental cached-token prices 64%. Google shut down the Gemini 1.5 API models on September 29, 2025, so these identifiers are historical rather than options for a new integration.

What Google announced on September 24, 2024

The announcement was a production refresh within the Gemini 1.5 family, not a wholly new model generation. Google released two stable versions:

  • gemini-1.5-pro-002
  • gemini-1.5-flash-002

The mutable aliases gemini-1.5-pro-latest and gemini-1.5-flash-latest were updated to point to those versions. Google also listed gemini-1.5-flash-8b-exp-0924 as the replacement for the earlier experimental 8B build. The model identifiers and alias changes are recorded in Google’s Gemini API release notes.

The same update added frequencyPenalty and presencePenalty support for Python and Node.js clients. Google described the releases as improved production models, but the announcement did not establish one universal benchmark percentage for quality gains.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Paid-tier throughput increases

Google announced higher paid-tier limits:

Model Announced paid-tier limit Previous limit
Gemini 1.5 Flash 2,000 requests per minute 1,000 RPM
Gemini 1.5 Pro 1,000 requests per minute 360 RPM

These were announced quota levels, not a guarantee of identical capacity for every account. Actual limits could depend on project, billing tier, region and Google’s quota policies. The announcement is documented at Google’s September 2024 blog post.

How much the Gemini 1.5 Pro price fell

The price change took effect October 1, 2024 and applied to Gemini 1.5 Pro prompts below 128,000 tokens. The reductions differed by billing component:

Billing component Announced reduction
Input tokens 64%
Output tokens 52%
Incremental cached tokens 64%

Calling this a “50% price cut” was shorthand for a reduction of more than 50% across several components. It did not mean Google halved every Gemini API price, every model or every request. The under-128K condition matters: applications sending longer prompts could not assume the same effective discount. Cached-token savings were most relevant to systems repeatedly sending the same instructions, documents or other context.

This was API pricing, separate from consumer Gemini subscriptions and not automatically a change to every Vertex AI or Google Cloud charge. The historical terms are in Google’s announcement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pro versus Flash: which workload each model targeted

Gemini 1.5 Pro

Pro was positioned for difficult reasoning, complex multimodal analysis, long documents and applications where answer quality justified higher latency or cost. A production team might have used it for contract analysis, research synthesis or workflows in which retries and human review were expensive.

Gemini 1.5 Flash

Flash was designed for speed, lower latency, cost efficiency and high volume. It was a better starting point for classification, extraction, routine summarization and other requests that could tolerate some capability trade-off.

Model choice was not determined by token price alone. Error rates, output length, latency, retries, review costs, function calling, structured output, tuning and multimodal requirements all affected total cost. A cheaper Flash call could be uneconomic if its mistakes required repeated calls or manual correction.

How this fit the wider Gemini 1.5 rollout

The September release followed several 2024 milestones rather than appearing in isolation:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Date Milestone
February 15, 2024 Gemini 1.5 Pro introduced for early testing.
May 10, 2024 Gemini 1.5 Flash entered API preview.
May 23, 2024 Pro-001 and Flash-001 reached general availability.
June 18, 2024 Context caching became available in the API.
June 27, 2024 Gemini 1.5 Pro’s two-million-token context window reached general availability.
September 24, 2024 Stable Pro-002 and Flash-002 released.
October 1, 2024 Pro price reductions became effective.
September 29, 2025 Gemini 1.5 Pro, Flash and Flash-8B API models shut down.

Other related 1.5 developments included code execution, PDF text-and-vision processing, video and plain-text File API support, JSON schema and function-calling improvements, Flash tuning, broader language support and AI Studio interface changes. Google documented those changes across its general-availability announcement, Flash and AI Studio update and release notes.

Who benefited most from the historical discount?

  • Production services making large numbers of Pro calls.
  • Text or multimodal workloads whose prompts usually stayed below 128,000 tokens.
  • Applications that reused context and could benefit from caching.
  • Startups whose margins were sensitive to inference costs.
  • Teams needing the higher paid-tier request rates.

The saving mattered less to free-tier hobby projects, workloads dominated by very long prompts, or systems where storage, retrieval, grounding, observability and other infrastructure outweighed token charges. It also provided no ongoing benefit to teams that needed a supported model after the 2025 shutdown.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Important caveats for interpreting the announcement

“50% cheaper” was not universal

The precise claim is that selected Gemini 1.5 Pro input, output and cached-token prices fell by more than 50% for prompts under 128K tokens. It is inaccurate to describe the change as a flat halving of all Gemini API prices.

Aliases could change behavior

-latest aliases were convenient for upgrades but mutable. Production systems generally had better reproducibility when they pinned an explicit version and ran regression tests before changing it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Free access and paid access were different

AI Studio and API free-tier access supported experimentation, while paid usage had different quotas and applicable data-use terms. Neither should be confused with a consumer Gemini subscription or with enterprise Vertex AI deployment.

What developers should use now

Gemini 1.5 Pro, Flash and Flash-8B are unavailable for new API use because Google shut them down on September 29, 2025. Check Google’s deprecation schedule, current model guidance and current pricing before selecting a replacement. Google’s newer Gemini families do not necessarily provide a one-to-one behavioral or pricing match, so migration should include representative prompts, latency and cost measurements, structured-output checks and safety evaluations.

For enterprise deployments requiring Google Cloud identity, regional controls and centralized billing, Vertex AI may be more appropriate than a direct API-key integration. Teams can also compare current offerings from OpenAI, Anthropic, Amazon Bedrock and Azure OpenAI; their prices and model availability change independently.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Money Desk

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.