Choosing an LLM: cost, capability, and sovereignty risk

Choosing an LLM provider is a three-variable decision, not one: cost, capability, and data-sovereignty and security risk. Pure cost/capability comparisons miss the third variable entirely, and it's the one with the most consequential recent evidence behind it.

Why cost and capability aren't the whole picture

US enterprise AI token usage on Chinese open-weight models (DeepSeek, Qwen, GLM, Kimi) has run roughly 30-46% of total weekly usage since February 2026, driven almost entirely by 60-90% cost savings over frontier US labs. That's a rational response to real cost pressure, but most teams are making the call on cost and capability alone, without pricing in the third variable.

The Chinese open-weight model tradeoff

A May 2026 Booz Allen Hamilton study found that a leading Chinese code-generation model produced measurably more vulnerable code under adversarial testing: roughly 130% more vulnerabilities than baseline when prompted under a "US government" persona. Separately, DeepSeek has confirmed data transfers to a Beijing-based cloud provider without explicit consent, has no GDPR adequacy decision, and is banned outright in Italy and several US states.

None of that means the cost savings are fake. They're real and often substantial. It means the savings need to be weighed against a specific, documented risk, not assumed away because the peg has held so far.

A scoring approach

Score each candidate model on three axes: cost relative to a frontier baseline, capability on benchmarks relevant to the actual task, and sovereignty/security (data residency and transfer transparency, regulatory status (GDPR adequacy, state or national bans), and demonstrated output security under adversarial testing). A model that wins on cost and capability but scores High risk on sovereignty isn't automatically disqualified, but it shouldn't be selected by default either.

What this means for procurement

Model selection should go through the same vendor risk review as any other data processor, with sovereignty and security scored explicitly as part of that review, not treated as a separate legal sign-off bolted on after the technical and cost decision is already made.

FAQ

Are Chinese open-weight models actually less secure, or is this overblown?

The evidence is specific, not hand-wavy: a May 2026 Booz Allen Hamilton study found measurably more vulnerabilities in code generated by a leading Chinese model under adversarial testing, and DeepSeek has confirmed data transfers to a Beijing-based cloud provider without explicit consent. Those are documented findings, not general suspicion of origin.

Does this mean companies should avoid Chinese open-weight models entirely?

Not necessarily. It means the decision should be made with the risk visible, not ignored because the cost savings are real. For low-sensitivity, non-regulated use cases, the tradeoff may be perfectly acceptable. For anything touching regulated data or security-relevant code generation, it usually isn't.

How much does getting this wrong actually cost?

It ranges from a compliance finding (GDPR exposure, an internal audit flag) to a real security incident if vulnerable generated code ships to production. The cost savings on the model itself are often small relative to either outcome, which is exactly why the sovereignty variable belongs in the decision, not as an afterthought.

Where does this fit into a procurement process?

As a formal scoring input alongside cost and capability benchmarks, not a separate legal sign-off bolted on at the end. Model selection should go through the same vendor risk review as any other data processor, with sovereignty and security scored explicitly rather than assumed.

This is the same lens I'd bring to an AI vendor risk review. Happy to talk through your current model stack. Get in touch.