AI coding startups can attract users and generate revenue at extraordinary speed without making attractive profits. The reason is a mismatch: customers often pay a fixed subscription while their use of models, codebase context, and autonomous agents can vary enormously. Reporting through August 16, 2026, offers examples of both the opportunity and the risk: Cursor reached an estimated $500 million annualized revenue run rate by June 2025, while sources familiar with Windsurf’s finances described its gross margins as “very negative.” Neither figure alone establishes whether a company is durable. The key question is whether it can earn more per useful, accepted software change than it spends producing one.
Why a coding assistant can be expensive to serve
A coding request is not always one prompt and one answer. A tool may search a repository, load files and documentation, plan a change, edit several files, run tests, inspect errors, and try again. An agent may also execute terminal commands or interact with other tools. Each model call can consume input and output tokens; long context and repeated retries multiply the bill.
Inference is only part of the cost. Serving a product can also require codebase indexing and embeddings, storage, sandboxed execution, build and deployment infrastructure, logging, security controls, abuse prevention, customer support, and enterprise implementation. A subscription’s economics depend on how these costs are accounted for, not just on the price of model tokens.
Usage is uneven. One subscriber might use occasional autocomplete; another might run agents against large repositories throughout the day. Flat-rate pricing makes those customers look alike on the invoice even when they impose very different serving costs. Free plans can add to the mismatch: The Information reported that some coding assistants exclude model costs for free users from cost-of-revenue calculations, which can make a reported gross margin look stronger than the product’s full blended economics (The Information).
#1 Best Overall
Why a monthly subscription does not reveal the margin
A $20 monthly price, for example, says little about profitability unless you also know the included usage, model mix, context limits, premium-model rules, overage charges, and whether background agents are included. The same price can be sustainable for a light user and loss-making for a heavy one.
Consider a purely illustrative example, not an estimate for any named company: a subscriber pays $20 in a month. A light user costs $5 to serve in model and infrastructure expenses, leaving $15 before other costs. A heavy agent user costs $35 to serve, so the company loses $15 on that subscription before research, sales, support, or corporate expenses. The average can look acceptable if heavy use is rare, but it can deteriorate as a product becomes more useful and its most enthusiastic customers use it more.
Vendors have several imperfect choices. They can raise prices or meter usage, which better aligns revenue with consumption but risks customer backlash and unpredictable bills. They can limit access to expensive models or agent runs, but restrictions weaken the product’s appeal. They can route simpler tasks to cheaper models, reserving stronger ones for difficult work, though quality may become less consistent. Cursor changed pricing in response to costs associated with newer Anthropic models, particularly for active users, and TechCrunch reported that the company later apologized for unclear communication about the change (TechCrunch).
What the reported company figures do—and do not—show
Private-company figures reported by media are not audited statements. They are still useful as illustrations if their dates, sources, and limits remain attached to them.
| Company and reported figure | What it indicates | Important qualification |
|---|---|---|
| Windsurf: gross margins described as “very negative” | People familiar with the company’s economics told TechCrunch that the direct cost of serving the product could exceed revenue. | An unnamed-source report, not company-confirmed or audited financial information. TechCrunch |
| Cursor/Anysphere: approximately $500 million annualized revenue run rate by June 2025 | Evidence of rapid commercial traction and willingness to pay. | A reported run-rate estimate, not $500 million of recognized annual revenue, profit, or evidence of healthy margins. TechCrunch; The Information |
| Lovable: approximately $5.7 million revenue and $3.7 million in LLM and other costs in May 2025 | The reported figures imply roughly 35% gross margin: ($5.7 million − $3.7 million) ÷ $5.7 million. | Figures attributed to a person familiar with company financials; accounting scope may vary. Gross margin is not operating margin. The Information |
These examples answer different questions. A large revenue run rate shows demand and sales momentum; it does not reveal cost per user, gross margin, retention, or cash burn. A reported negative gross margin for one company does not establish that every startup loses money on every customer. And a 35% gross margin, if calculated as reported, still leaves a company to fund research and development, sales, support, and general operations.
Windsurf: traction did not remove strategic pressure
Windsurf’s reported history illustrates why valuation and transaction headlines should not be mistaken for proof of a durable standalone business. TechCrunch reported that the company discussed financing at a $2.85 billion valuation in February 2025 and later had a proposed sale to OpenAI for roughly $3 billion that collapsed. These were reported discussions and a proposed transaction, not completed financing or acquisition. TechCrunch also reported that Windsurf’s leaders pursued strategic-sale options amid cost and competitive pressure, and that Cognition ultimately acquired the remaining business (TechCrunch).
That sequence is not enough to reduce the company’s outcome to “failure.” A strategic transaction can reflect product value, talent, technology, distribution, and the economics of operating independently. But the reported margin concern helps explain why rapid adoption and a high proposed valuation do not by themselves resolve whether a company can sustainably serve customers on its own.
Cursor: revenue velocity is not margin quality
Cursor’s reported annualized revenue run rate of about $500 million by June 2025 is a striking measure of demand. It is not proof of a $500 million annual profit, nor does it show how costs were distributed between light and heavy users. The pricing changes reported by TechCrunch point to the practical challenge: a tool can attract customers with a simple subscription, then face pressure when the most valuable workflows require expensive models and substantial usage.
TechCrunch also reported that Anysphere was attempting to build its own model. That may give a vendor more control over cost, latency, model routing, and product-specific behavior. It also requires model-serving expertise, evaluation systems, research, data, and potentially major infrastructure commitments. Replacing an API bill with fixed research and serving costs is a trade-off, not an automatic margin improvement.
The supplier can also be a competitor
Independent coding-tool companies may depend on model providers that sell competing products. TechCrunch identified Anthropic’s Claude Code and OpenAI’s Codex as direct competitors to startups using their models (TechCrunch). GitHub Copilot, Google’s developer offerings, and other large providers also compete for developer workflows.
Rank #3
This is more than ordinary vendor dependence. A model provider can set access and pricing terms, improve or change its model roadmap, and bundle coding features with products customers already use. It may also have broader distribution and resources to subsidize a coding product. A startup therefore needs reasons to remain valuable beyond access to a model: workflow design, integrations, reliability, enterprise controls, specialized data, or support.
Building a proprietary model is one possible response, but it brings high fixed costs and the risk that general-purpose providers improve faster. TechCrunch reported that Anysphere pursued this path while Windsurf’s leadership reportedly decided against it because of the expense and complexity. Neither decision proves a universal best strategy.
Why cheaper tokens may not mean cheaper coding
Some investors expect inference costs to fall. In TechCrunch’s reporting, GV general partner Erik Nordlander argued that current inference costs may be near a high point. But the same report noted that some newer models can cost more when they use additional computation for complex, multistep reasoning (TechCrunch).
The more useful economic measure is not just the price per million tokens. It is the cost to complete a useful coding task that a developer accepts and can ship. Token prices may decline while context grows, agents make more calls, users expect better results, and failed attempts trigger retries. The cost of a completed task can fall, stay flat, or rise depending on how those forces interact.
Nor is “AI coding” one uniform workload. Inline autocomplete, professional IDE assistance, terminal agents, and browser-based app builders have different serving costs, customer expectations, and routes to revenue. A flat-rate plan may work well for autocomplete but struggle with an open-ended agent that runs commands, reads a codebase, and iterates until a task is complete.
Rank #4
What makes customers stay—and what makes them switch
Individual developers can often try multiple tools, use several in parallel, or move to an open-source or bring-your-own-key (BYOK) option. Their repositories and code generally remain outside a vendor’s control, so switching may be easier than it is with deeply embedded business software. The Information has identified churn and weak differentiation as risks for coding startups (The Information).
That does not mean every product is interchangeable. Shared indexes, team policies, security and compliance controls, workflow integrations, usage analytics, and deployment pipelines can create real switching costs for an organization. Enterprise customers may pay more, but they can also require lengthy procurement, support, security reviews, and implementation work.
For a buyer, the monthly sticker price is therefore an incomplete comparison. A more useful question is how much the full workflow costs per accepted pull request, resolved bug, shipped feature, or hour of engineering time saved. Include usage, hosting or deployment, human review, rework, and the operational effort required to manage the tool.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to tell a durable business from a fast-growing wrapper
Revenue and user growth matter, but they are only part of the evidence. A founder, investor, executive, or developer assessing durability should look for economics by workload and customer segment rather than relying on one blended margin number.
- Cost per active user and per accepted code change: Track what it costs to deliver useful work, not just raw requests or token volume.
- Usage distribution: Measure heavy-user percentiles, agent runs, model mix, retries, and context consumption. A small group of power users can dominate variable costs.
- Plan-level contribution margin: Compare subscription and usage revenue with inference, execution, hosting, storage, and payment costs; show whether free users and trials are included.
- Task completion and rework: Track how often work is accepted without substantial human correction. Cheap output that requires extensive repair may be costly in practice.
- Retention after pricing or usage changes: A pricing change tests whether customers value the product enough to pay for the work it performs.
- Revenue quality: Separate recurring recognized revenue from annualized run-rate estimates, and assess customer concentration, enterprise mix, and renewal behavior.
- Operating needs: A positive contribution margin does not cover research, sales, support, recruiting, or other operating costs by itself.
High growth can still be strategically valuable when a company has a trusted brand, enterprise relationships, proprietary workflow data, or a position in the development process. Those assets can support partnerships or an acquisition even if standalone profitability remains uncertain. But they do not substitute for knowing whether serving customers becomes more economical as the business scales.
Best Value
Three plausible paths for the market
Inference gets more efficient
Cheaper models or better routing could reduce serving costs, particularly for routine tasks. This path works best if quality holds and the number of calls needed per accepted change does not rise faster than unit costs fall.
Pricing becomes more closely tied to usage
Vendors may add metered consumption, premium-model allowances, or enterprise commitments. That can make costs more predictable for the seller, but customers will need clear budgets and a visible connection between usage charges and delivered value.
Model providers absorb or consolidate the category
Large providers can bundle coding capabilities, use existing distribution, or subsidize products strategically. Independent startups may respond by specializing, partnering, selling, or focusing on workflows where their product and service are difficult to replace.
These paths can coexist. A company might route routine tasks to cheaper models, charge for intensive agent use, and focus its sales effort on enterprise teams at the same time.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




