Skip to main content

Corporate America Is Switching to Open-Weight AI Models to Cut Costs

A New York Times report published September 4 describes corporate America getting hooked on open-source AI: companies including AT&T and Ramp are cutting bills from OpenAI and Anthropic by routing work to free, downloadable models. AT&T says open models now handle 40 percent of employee queries, coding costs fell by up to 56 percent, and quality dropped only 2 percent.

On this page

The numbers behind the pivot

The New York Times published a report on September 4 titled "Corporate America Is Getting Hooked on Open-Source A.I.," describing how large companies increasingly use cheap, freely available models instead of paying premium prices to OpenAI and Anthropic. The trend is no longer experimental: Uber, Pinterest, Stripe, Coinbase, Ramp, and AT&T have all reported significant savings from moving work to open models, according to The Pragmatic Engineer's roundup of the same reporting. [1] [3]

AT&T is the most detailed example. The company cut costs on coding and other advanced AI tasks by as much as 56 percent while quality declined by only about 2 percent, according to The Information's reporting via PYMNTS. Employee AI usage is surging from 8 billion to 45 billion tokens per day, and open models are the lever AT&T is using to keep that growth affordable. [2] [4]

How AT&T actually does it

The mechanism is a router. AT&T uses LiteLLM model routers that assess each task's complexity and send simple queries to cheap open models, reserving frontier models for work that needs them. Currently about 40 percent of employee queries run on open models, with a target of 60 to 70 percent in the coming years, while spending on Anthropic and OpenAI models stays flat. [2]

The model roster is mainstream open weight: Nvidia's Nemotron, Meta's Llama, and Google's Gemma are in use. Notably absent are DeepSeek and Moonshot: AT&T is still evaluating potential risks around Chinese models before touching them. Mark Austin, the AT&T vice president overseeing employee-facing AI, says open-source capabilities generally lag frontier models by six to ten months, a gap that keeps narrowing, and that the open options are "just as good or better" than older Anthropic and OpenAI models. [2]

The catch: gaps and gatekeepers

Two limits keep this from being a total victory lap. First, the latency: a six-to-ten-month capability lag means cutting-edge work still routes to commercial models, which is why AT&T's plan caps open models at around two-thirds of usage rather than replacing paid providers entirely. [2]

Second, not every open model clears corporate risk review. AT&T's hesitation on DeepSeek and Moonshot shows that provenance matters to buyers, which is precisely why transparency efforts like MBZUAI's fully open K2 Horizon fleet and debates over who controls Hugging Face have business consequences beyond ideology. [2] [5]

What it means for local AI

For people who run models on their own hardware, this is quiet validation. The exact weights powering corporate cost-cutting, Nemotron, Llama, Gemma, and their quantized derivatives, are the same ones that run on a gaming PC or a phone, and every enterprise adoption makes the tooling around them better funded and more polished. [2]

It also explains why the local ecosystem keeps getting new infrastructure, from model routers to pooled home GPUs: the problems of serving open models cheaply are now enterprise problems, and enterprises fund solutions. The economics that pushed AT&T toward open weights are the same ones that make a local model the right default for private tasks, minus the metered tokens entirely. [3]

Sources

  1. Corporate America Is Getting Hooked on Open-Source A.I.The New York Times
  2. AT&T slashes AI costs by adopting model routers and open sourcePYMNTS
  3. The Pulse: tech companies move to open modelsThe Pragmatic Engineer
  4. Open models are driving AT&T's AI tokenomics strategyFierce Network
  5. MBZUAI's Institute of Foundation Models launches K2 HorizonMBZUAI

CuriousLM runs supported AI models locally on your device. Try CuriousLM.