DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

Cost and Model Complexity Remained Barriers to Enterprise AI in IBM’s 2024 Survey

IBM’s 2024 survey reported that 63% of executives cited model cost and 58% cited complexity as top concerns. The practical challenge is managing the full cost and risk of matching models to enterprise tasks.
From TheFinanceBase Team8 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

IBM’s 2024 survey found that executives were concerned about both the cost of generative-AI models and the work of managing them: 63% cited model cost as a top concern, and 58% cited model complexity. The findings describe surveyed executives’ views in 2024—not a measurement of the enterprise AI market in 2026. Their central implication still matters for planning: organizations need to choose models and deployment methods for particular tasks, then account for the full cost of making those systems reliable and governable.

What IBM’s survey found

The figures come from The CEO’s Guide to Generative AI: AI Model Optimization, published by the IBM Institute for Business Value (IBV) in collaboration with Oxford Economics. IBM describes it as proprietary research focused on U.S.-based executives and enterprise generative-AI decision-making. The public summary does not establish every methodological detail, including the full sample size, fieldwork dates, respondent composition, or margin of error.

The findings were reported in a VentureBeat article published July 31, 2024. They should be read as survey results and expectations from that period, not as independently audited usage data or a current 2026 market measure.

IBM-reported finding What it means—and what it does not establish
About 11 generative-AI models Average number used by organizations in the survey; not a universal enterprise count or telemetry-based audit.
About 50% portfolio growth Surveyed organizations expected their model portfolios to grow by approximately 50% over three years, described in IBM’s material as 2024–2027. This is an expectation, not a measured outcome.
63% cited model cost as a top concern A reported executive concern; it does not quantify enterprise spending or identify a cost threshold that blocks adoption.
58% cited model complexity as a top concern A reported concern, not a standardized complexity score or proof that complexity was the leading barrier in every industry.
42% consistently used fine-tuning and prompt engineering Reported in VentureBeat’s account of the IBM survey; this describes reported practice, not the share of workloads optimized.
63% expected open-model adoption to rise Surveyed organizations’ expectation for growth over three years, not evidence that adoption subsequently rose by that amount.

IBM also describes portfolios spanning commercial, open, embedded, and internally developed proprietary models. Any category shares shown in its report describe the surveyed model mix; they are not market-share estimates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why one model is rarely right for every task

Enterprise workloads differ in the errors they can tolerate, the response time they need, the information they can expose, and the volume they process. A model that works well for drafting marketing text may be too costly for millions of routine classifications, too unpredictable for a tightly controlled workflow, or inappropriate for a sensitive decision.

The useful selection question is not “Which model is best?” but “Which approach meets this task’s quality, latency, security, compliance, and cost requirements?” Depending on the work, the answer may not be a generative model at all.

Workload need Starting point to evaluate Important check
Deterministic rules or fixed calculations Conventional software or workflow automation Whether the process can be handled reliably without generative output.
Structured classification, forecasting, or prediction Traditional machine learning Whether structured features and a defined target outperform a language-model approach.
Narrow, high-volume language task Small or task-specific model Whether it meets the quality threshold at production volume, including exception handling.
Answers grounded in internal documents Retrieval-augmented generation with an appropriately sized model Document freshness, access permissions, retrieval quality, and whether answers can be traced to supporting sources.
Broad, difficult generation or reasoning Larger general-purpose model Whether the added capability justifies its cost, latency, and data-handling implications.
High-impact or legally sensitive decision Human-led process, with AI assistance only where appropriate Accountability, review, appeal, and rollback—not just model accuracy.

A smaller or cheaper model is not automatically less expensive overall. If it creates more incorrect outputs, escalations, or manual rework, its apparent savings can disappear. Compare cost per successful task rather than cost per API call.

What enterprise AI costs beyond the model call

For a hosted service, inference charges commonly depend on consumption such as input and output tokens. Long prompts, large context windows, high request volumes, retries, and multimodal inputs can change that bill. With internal hosting, the organization instead bears compute and storage costs, including the capacity needed for reliable service. These are different cost structures, not a simple free-versus-paid choice.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A useful total-cost view includes:

  • Inference: Input and output usage, request volume, context length, and real-time versus batch processing.
  • Training and adaptation: Fine-tuning, synthetic-data generation, evaluation runs, and retraining or preference optimization where used.
  • Infrastructure: Accelerators or CPUs, memory, storage, networking, orchestration, and high-availability capacity.
  • Data: Cleaning, labeling, storage, indexing, retrieval, security controls, and data transfer.
  • Integration: Connections to ERP, CRM, data warehouses, document systems, identity providers, and workflow tools.
  • Governance and operations: Monitoring, audit records, red-teaming, policy enforcement, privacy controls, regulatory documentation, and human review.
  • People: Data and application engineering, evaluation, security reviews, procurement, and change management.
  • Failure and dependency: Rework after incorrect answers, privacy incidents, downtime, and the cost of being locked into a provider or proprietary feature.

Costs can also accumulate in retrieval, embeddings and vector storage, agent loops and tool calls, evaluation and red-team traffic, logging and retention, and idle capacity in a self-hosted deployment. A budget based only on a model’s advertised per-token rate will miss those parts of the service.

Estimate the economics as total cost = inference + infrastructure + data preparation + integration + governance + monitoring + human review + failure and rework. Compare that with measurable benefit to determine net value. Track cost per successful task, task-completion quality, human escalation and correction rates, peak-volume latency, spending under longer contexts, model-change regressions, and the effort needed to approve and deploy a replacement.

What model complexity looks like in practice

Complexity is not just the number of models. Different providers and deployment types can bring distinct APIs, authentication, safety controls, context limits, output formats, licensing terms, and support arrangements. Each may require its own evaluation set, monitoring, security review, and update process.

Routing requests among models can improve fit and control costs, but it adds another decision layer: the router must select an appropriate model without sending sensitive information to an unsuitable endpoint or creating unpredictable output quality. Model updates can also change cost, latency, or behavior. Provider-specific features can make applications harder to move later.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These operational issues become more consequential when several models, prompts, agents, tools, and downstream actions participate in one workflow. Teams need an inventory, clear ownership, version records, consistent evaluations, data-flow visibility, and a way to trace which system produced or acted on an output.

Optimization methods—and the limits of the reported result

IBM’s coverage says fine-tuning and prompt engineering could improve accuracy by 25%, while 42% of executives reportedly said their organizations used those methods consistently. The 25% figure is a survey-reported result, not a universal performance guarantee. The public coverage does not fully specify the baseline, tasks, accuracy measure, evaluation method, or whether the change means relative improvement or percentage points.

Optimization should be treated as a measured choice among alternatives, not as a mandate to fine-tune every model:

  • Improve prompts when clearer instructions or structured outputs address the problem. Version prompts and test them against a fixed evaluation set; prompt behavior may shift when a model changes.
  • Use retrieval when answers need current or internal knowledge. It can reduce the need to encode changing facts in model weights, but depends on good documents, permissions, indexing, and retrieval.
  • Fine-tune when a stable, well-defined task and representative training examples justify adaptation. Poor or outdated data can encode bias or stale behavior, and fine-tuning adds update and evaluation work.
  • Route or substitute models when different tasks or request types have distinct quality and cost thresholds. Evaluate the router as well as the models it selects.
  • Reduce waste with appropriate context limits, caching, batching, and—where quality and hardware allow—distillation or quantization.

Retrieval is not a cure-all: stale or duplicated documents, incorrect access permissions, poor chunking, irrelevant results, and prompt injection in retrieved content can undermine an answer. Teams should be able to identify what evidence was retrieved and whether it supports the result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Open and proprietary models: a workload-level trade-off

IBM’s survey respondents expected open-model use to increase by 63% over the following three years. That is a 2024 expectation, not a confirmed later adoption figure. It is also not evidence that open models are universally cheaper, safer, or more private.

Open models may provide deployment control, customization options, and less dependence on a single proprietary provider. Some make model weights available, but “open” does not automatically mean that training data is open, commercial use is unrestricted, or compliance is straightforward. Licensing must be checked model by model.

Running a model yourself shifts responsibility toward the enterprise: infrastructure, security, patching, evaluation, support, and operational staffing can outweigh savings on hosted inference. Conversely, a hosted proprietary model may reduce infrastructure work but introduce provider, pricing, data-use, residency, and portability dependencies. Compare fully loaded cost and risk for the specific workload and deployment—not labels in isolation.

A practical way to choose and govern models

  1. Start with the process, not a model. Define the business outcome, users, systems affected, and whether AI is advisory or can trigger actions. IBM’s interview recommends beginning with the business process; it names customer service, IT operations, HR, and supply chain as areas to examine, while noting that conventional AI or automation may sometimes be more suitable.
  2. Set acceptance criteria. Record acceptable error and escalation rates, latency, cost per completed workflow, data sensitivity and residency, audit needs, integrations, and human review or rollback procedures.
  3. Choose the least capable approach that meets the criteria. Compare rules, conventional machine learning, smaller models, retrieval-based systems, and larger models. Do not default to the largest model merely because it is available.
  4. Evaluate on representative cases. Use realistic inputs, edge cases, and a fixed test set. Measure quality, latency, and fully loaded cost, including human correction and exception handling.
  5. Pilot with controls. Limit scope, log versions and outcomes, define escalation paths, and make a person responsible for review where the impact warrants it.
  6. Monitor and re-evaluate. Watch for changes in quality, latency, cost, data flows, and model versions. Re-test before updates or routing changes reach production.

For procurement and governance, document data-use and retention terms, processing region and residency, availability commitments, version-change notice, audit logging, fine-tuning rights, portability and exit options, peak-usage pricing, support, security assurances, and any performance commitments. Assess vendor-management, contract, and software-supply-chain exposure alongside technical performance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the survey can—and cannot—tell leaders now

IBM’s results capture executives’ concerns and plans at the time of the survey. They do not establish actual average enterprise AI spending, a universal cost threshold, which dimension of complexity respondents meant, or whether the stated barriers have eased since 2024. The 25% accuracy result also lacks enough published detail in the cited coverage to apply it as a forecast for a particular task.

The recommendations to match models to work, optimize, and govern a portfolio are useful decision principles, but the survey is not vendor-neutral proof of a particular platform choice: IBM also sells AI services and infrastructure. Leaders should test the principles against their own workloads and procurement constraints.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase07 MAR 2625 minWhat Is a 457 Plan?
  2. The Money DeskBlogTheFinanceBase07 MAR 2621 minTime Value of Money: What It Is and How It Works
  3. The Money DeskBlogTheFinanceBase07 MAR 2627 minAre You Living in One of These Top 10 Most Expensive Cities to Retire?
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.