Which Model You Run and Who Serves It Are Two Separate Decisions

TV
Thiago Victorino
7 min read
Which Model You Run and Who Serves It Are Two Separate Decisions

Chinese open-source models now carry more than 60% of OpenRouter’s traffic, in a market where US models held roughly 70% a year earlier. Fortune reported the shift on 21 August 2026. Fortune’s July token-volume ranking puts Chinese-developed models in all five top spots: Xiaomi’s MiMo V2.5 first, then DeepSeek, MiniMax, Alibaba’s Qwen family and Moonshot’s Kimi.

The pull is priced in dollars. OpenRouter’s Justin Summerville told CNBC that Chinese open-source models run 60% to 90% cheaper than leading Anthropic and OpenAI offerings. DeepSeek V4 Flash costs $0.14 per million input tokens. OpenAI’s GPT-5.5 costs $5.00 for the same million input tokens.

Most of the argument around those numbers is conducted as a single question about whether a Chinese model can be trusted. That question bundles two variables that move independently: whose weights are running, and whose jurisdiction the inference happens in. Naming them separately changes what a governance team is actually deciding.

The two variables

Model provenance is who trained the weights, on what data, under what content policy, and with what supply chain behind the artifact. Provenance travels with the weights. Downloading Qwen and running it on your own metal does not change who trained it, what the training corpus contained, or what behavior was optimized into it.

Serving jurisdiction is whose legal entity receives your prompt, where the request is processed, and which government can compel access to it. Serving jurisdiction travels with the endpoint, not with the model. It changes the moment you change hosts.

The evidence that these separate cleanly is operational, not theoretical. OpenRouter’s June analysis, cited in the Startup Fortune report this piece draws on, notes that western hosts including Fireworks, Together and DeepInfra let companies run these models without sending traffic through a Chinese first-party API. Same weights, different endpoint, different legal counterparty.

The residency side of the exposure is documented in plain language by the vendor. DeepSeek’s privacy policy states that the services are provided by Hangzhou DeepSeek Artificial Intelligence Co., Ltd., and that personal data is directly collected, processed and stored in the People’s Republic of China. It also says DeepSeek may share personal data with law enforcement or public authorities when it believes that is necessary to comply with law or legal process. That is a statement about the first-party endpoint. It says nothing about the weights, and it stops applying when the weights run somewhere else.

The four cells

Two binary variables produce four positions, each with a distinct risk profile.

Western-hosted servingChinese first-party API
Western weightsThe long-standing default. Highest cost, familiar provenance and residency posture.Not a configuration this source reports on. Named so the matrix stays complete.
Chinese open weightsMost of the cost advantage, without the first-party residency exposure.Cheapest, and the only cell where prompts land under the vendor’s stated PRC processing terms.

The bottom-left cell is where the interesting work is. It keeps most of the cost advantage that made the migration attractive, and it removes the specific exposure that DeepSeek’s own policy describes. It does not remove provenance questions. Content policy baked into the weights, training-data composition, and the integrity of the artifact you downloaded all survive a change of host. Those belong to the supply-chain question we covered in Your Model Provider Is a Supply Chain.

Teams that collapse the two variables into one question land badly in both directions. Treat everything Chinese as disqualified and you forfeit a 60% to 90% cost reduction over a residency risk you could have engineered away. Treat cost as the whole story and you take on a data-residency position nobody in the organization has approved, because nobody was asked.

The reason the hosting decision carries the residency weight, rather than the vendor’s policy language, is an argument we made in The Training Opt-Out Is Not an Egress Control: a contractual promise about how data will be handled after it arrives is a different class of control from a route that never sends it. Vendor policy is a commitment. Endpoint selection is an architecture. Only one of them is enforceable by your network.

What Washington has actually done

The congressional activity is a fact about the operating environment, and it matters for procurement regardless of what anyone thinks of the underlying politics.

In April 2026, the House Homeland Security Committee and the House Select Committee on China opened an investigation into PRC-developed AI models, naming DeepSeek, Alibaba, Moonshot AI and MiniMax, and raising concerns about model provenance, censorship, cybersecurity and supply-chain risk. In an April 29 letter to Airbnb, the committees noted that the company had publicly described Qwen as “fast and cheap.” On 29 July, Senator Tom Cotton wrote to Commerce Secretary Howard Lutnick urging a government-wide ban on contractor use of Chinese AI models.

Read against the matrix, this activity is aimed largely at provenance, which is the variable a change of host does not fix. A company with US federal exposure should assume that a future restriction may be written against the model family, not against the endpoint. That is a different planning question from data residency, and it belongs in a different row of the risk register. Jurisdiction as a purchasing column is a move we have argued is already underway in Europe, in Mistral’s year.

The number that is being quoted wrong

Martin Casado of Andreessen Horowitz is widely quoted as saying that 80% of AI startups use Chinese models. Startup Fortune states the correction explicitly: roughly 20% to 30% of AI startups pitching the firm build on open-source models, and about 80% of that group use Chinese ones. That works out to roughly 16% to 24% of all pitches.

The distance between 80% and 16-24% is the distance between an industry-wide default and a meaningful minority. If a board deck or a vendor pitch in your inbox cites the 80% figure, it is citing a compound number as if it were the base rate.

The direction of travel

Whatever the policy environment does, the supply side keeps building. Alibaba said in its fiscal 2026 filing that it released three Qwen updates in three months, and that its consumer Qwen app had passed 295 million monthly active users by March 2026. NVIDIA published a technical blog on 26 August showing Qwen3.8-Flash-Next running on its GB300 NVL72 system for agentic coding, and Perplexity’s local-first Portable Computer supports Qwen 27B on machines with Nvidia GPUs.

Open weights running on western silicon, documented by western vendors, is the same decoupling seen from the hardware side. Combined with the western hosts named above, the bottom-left cell is not hypothetical.

Do this now

Open your model inventory and split the single “approved / not approved” column into two: model provenance and serving jurisdiction. For every production workload, record which weights are running and which legal entity receives the prompt. Any row where those two fields collapse into one answer is a row where a decision was made without being named, which is the same failure we described in The Default Posture Is a Procurement Decision.

Then check whether a workload currently pointed at a Chinese first-party API can be repointed at a western host. If it can, you have a residency question with a technical answer, and the cost case survives the move. If it cannot, the provenance question is the one you are actually holding, and it needs an owner rather than a price comparison.


This analysis draws on US Startups Are Quietly Replacing OpenAI and Anthropic With Chinese AI (Startup Fortune, August 2026), which reports figures originating with Fortune, CNBC, The Decoder and OpenRouter. Those figures are cited here as secondary reporting and have not been independently verified.

Victorino Group helps engineering and governance teams separate model selection from serving topology before either decision reaches production. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation