U.S. companies are moving routine artificial intelligence work to open-weight and open-source models as the cost of operating premium systems rises with usage. The change is showing up in earnings calls, internal deployment targets and token-routing data, with companies increasingly reserving frontier models for difficult reasoning while assigning cheaper models to high-volume tasks such as classification, extraction, customer responses and code edits.
AlphaSense found that references to “open-weight” or “open-source” models in U.S. earnings calls and related meetings during August and September were six times higher than in the same period a year earlier. The language reflects a practical shift in enterprise AI strategy: businesses are no longer treating a single high-performing model as the answer to every workload.
Open-weight models publish their trained parameters, allowing companies to download, host and fine-tune them on their own infrastructure. That arrangement can give an organization more control over cost, data handling and customization than an externally hosted proprietary model, particularly when an application processes large volumes of internal documents or customer information.
Routine workloads move away from premium models
The cost issue has become more pressing as AI applications carry out longer tasks. A single user request may trigger multiple model calls, retrieve documents, process large context windows or use external tools. Per-token prices may be declining, but the number of tokens and steps consumed by enterprise applications is growing.
Uber experienced that pressure through internal use of AI coding tools. The company spent an AI budget expected to last through 2026 within four months, according to the supplied material, prompting it to introduce usage limits. The episode illustrates why finance teams are examining the full cost of AI workflows rather than relying on a model’s advertised price per token.
Tinder owner Match Group also began routing some non-technical queries to open-source models after its AI spending increased sharply. The company’s annualized AI spend rose from $1 million in January to $10 million in July, according to the supplied material. A lower-cost model may be sufficient for tasks where the output is short, structured or easy to verify, while premium systems can be held back for requests needing complex reasoning.
AT&T has put the strategy into measurable operational terms. About 40% of its AI workloads run on open-source models, and the telecom company aims to raise that share to 70% within a year. AT&T said it processes 45 billion tokens a day and fine-tunes open-source models with proprietary data for specific uses.
That scale helps explain why model selection has become an infrastructure decision rather than solely a product choice. Small differences in unit costs can become substantial when a company is handling billions of tokens each day.
Token data shows open models gaining share
Usage data from AI application platforms points in the same direction. Vercel said open-source models accounted for 56% of all tokens processed by its AI Gateway in August, compared with 7% in December. Separately, Citi data cited for model-routing platform OpenRouter showed open-source models’ token share rising from 34% in January to 65% in June.
Those figures do not mean proprietary frontier models are being displaced across the board. They suggest that enterprises and developers are becoming more selective, deploying each model where its performance-to-cost ratio makes sense.
The emerging architecture is often described as a multi-model setup. A frontier model handles planning, difficult coding, ambiguous requests or high-stakes analysis. A smaller or open-weight model executes repetitive steps after the work has been broken down. This structure can reduce spending without forcing companies to accept weaker results on their most demanding tasks.
Cursor’s testing offers one example of the potential gap. The company estimated that building a browser from scratch would cost more than $10,000 using OpenAI’s GPT-5.5 throughout, compared with $1,339 using Cursor’s Composer model alongside Anthropic’s Opus 4.8. The comparison put the all-premium approach at roughly 7.5 times the cost.
Telnyx, a communications technology company, had previously run 1,000 bot agents on Anthropic’s leading model and estimated daily costs could reach $100,000 under that approach. After moving work to open-source models, Telnyx expanded to 1,400 agents while lowering daily cost per agent to roughly $100, according to the supplied material. Legal AI company Harvey has adopted a similar division of labor, using open-source models for everyday work and top-tier systems for its hardest tasks.
Security and control influence deployment choices
Cost is not the only reason companies are adopting open-weight systems. Internal hosting can be attractive for organizations dealing with medical data, financial records, trade secrets or other sensitive material. Running a model inside company-controlled systems would allow an organization to apply its own access controls, maintain audit processes and fine-tune a model without sending every prompt to an external provider.
Digital Realty said it built an internal chat interface using open-source models and routes requests between open and proprietary systems according to sensitivity and complexity. That approach recognizes that a company’s data policies can be as decisive as model quality when choosing where a workload runs.
Citi pricing data cited in the supplied material placed some open-source models as low as $0.18 per million tokens, versus an average near $4 per million tokens for leading closed models. Performance is also narrowing in specialized areas. DeepSeek’s V4 Pro scored 80.6% on SWE-bench Verified and achieved the highest Codeforces rating among tested models, while its blended token rate was reported at roughly one-forty-sixth of Claude Opus 4.7’s rate.
The resulting market is less likely to be defined by one model winning every task. Companies are building systems that measure quality, latency, privacy requirements and cost request by request. Open-weight models are gaining a larger role where workloads are predictable, voluminous and suitable for customization, while premium proprietary systems retain an advantage in advanced reasoning and the most demanding general-purpose work.
That division places sustained pressure on AI providers to justify frontier-model pricing through clear performance gains. It also gives large enterprises more leverage: instead of accepting a single provider’s pricing and deployment model, they can route workloads across several systems and keep more of their AI operations under their own control.
Want to scale AI-driven trading without runaway costs? Explore AI copy trading for smarter, cost-efficient automation.
Disclaimer: The content on this page is provided for general informational purposes only and does not represent the views or financial advice of Toobit. We make no guarantees regarding the accuracy or completeness of this information and shall not be held liable for any errors, omissions, or outcomes resulting from its use. Investing in digital assets involves risk; users should independently evaluate their financial situation and the risks involved. For further details, please consult our Terms of Service and Risk Disclosure.
