The Cheap Model Is the Supply Chain

The cheap model is the supply-chain decision now.

In Episode 48 of The Sam Ellis Show, Sam reports on the new model-routing fight underneath AI agents and AI products: when inference cost decides which model handles real work, the router becomes procurement, compliance, reliability engineering, and geopolitics hiding behind one boring dropdown.

The lead proof is CNBC's reporting that Chinese-built AI models have gained traction among U.S. companies as costs rise at American labs. CNBC reported OpenRouter figures showing U.S. company token share on Chinese models through OpenRouter stayed above 30 percent each week since February 8, reached as high as 46 percent, and had averaged 11 percent over the previous 12 months. CNBC also reported that Lindy moved all of its traffic from Anthropic's Claude models to DeepSeek in June, with CEO Flo Crivello saying the move made the cost curve “crash to the ground,” and that Vercel saw Z.ai's GLM 5.2 grow about 27 times in daily token volume and about 80 times in customer count during its first full week.

The episode keeps the boundary exact. OpenRouter is a gateway, not the whole enterprise market. Company benchmark and efficiency claims remain company claims unless independently verified. Congressional scrutiny is treated as inquiry, not a finding. Reuters reporting on possible Chinese access curbs is treated as a discussion under consideration, not enacted policy.

The pressure is coming from both directions. U.S. lawmakers are probing American companies' use of PRC-developed AI models and raising supply-chain, data-security, and provenance concerns. Reuters reported that Chinese authorities have discussed potentially restricting overseas access to China's most advanced AI models, while the timing, scope, and even final decision remain unclear. That leaves operators squeezed between cheaper routing today and possible political, commercial, or technical interruption tomorrow.

OpenAI's GPT-5.6, xAI's Grok 4.5, and Meta's Muse Spark 1.1 make the same market signal louder. OpenAI is selling GPT-5.6 around “stronger performance per dollar,” cache economics, Programmatic Tool Calling, and multi-agent tiers. xAI is pricing Grok 4.5 into coding, agentic tasks, gateways, and tool workflows. Reuters reported Meta's Muse Spark 1.1 as a low-cost coding and agentic model, with Mark Zuckerberg saying Meta is focused on “delivering strong agentic and multimodal models at very low cost.” The arms race is no longer just intelligence. It is useful work per dollar.

For agents, this is not abstract procurement. Agents call, retry, summarize, inspect, repair, compact context, ask for tools, escalate, and route. Model choice is a repeated dispatch decision inside the work. If that dispatch layer is tuned mainly for cost, then cost is deciding what intelligence shows up where.

Sam's hook: the cheapest model is not automatically the wrong choice. Sometimes it is the only choice that lets the product exist. But once that choice becomes automatic, it stops being an optimization. It becomes dependency.

If you are routing production work between OpenAI, Anthropic, Chinese open-weight models, Grok, Meta, or anything through a gateway, send a note with the subject line routing cost. Anonymous and source-protection notes are welcome: [email protected].

Listen to Episode 48

Episode 48, "The Cheap Model Is the Supply Chain", is live now.

Download the episode or subscribe to the show feed.

Sources

Send tips, corrections, and source notes to [email protected].