Skip to content
ArticleAug 2026 · 6 min read

Should you be trying the new Chinese models?

The open-weight models are genuinely cheap per token. That saving reaches an API bill, not your monthly subscription. When a switch is worth making, and when it is just a distraction.

Generated Image August 02, 2026 – 8_02PM

Every few weeks a new open-weight model lands from DeepSeek, Moonshot, Alibaba or Z.ai, the benchmark charts do the rounds, and someone asks whether they should move their whole workflow across. The short answer is usually no, and the reason has almost nothing to do with the models.

We should say up front that we are firmly in favour of open weights. A world where the strongest models are all rented from four American companies is a worse world for everyone building on top of them. Weights you can download, run and audit are a real check on that, and the Chinese labs have been the ones actually shipping them. None of what follows is a knock on the models. It is a knock on the assumption that the cheap number on the pricing page is the number you will pay.

The pricing gap is real, and it is an API gap

The per-token numbers are not marketing. As of August 2026, DeepSeek V4 Flash runs at 0.14 dollars per million input tokens and 0.28 per million output. GLM-4.7-FlashX from Z.ai is 0.07 in and 0.40 out. GLM-5.2, their current flagship, is 1.40 and 4.40.

Now the other side of the table. GPT-5.6 Terra is 2.50 in and 15.00 out. Claude Opus 5 is 5.00 and 25.00. Claude Sonnet 5 sits at 2.00 and 10.00 on introductory pricing until the end of August, then moves to 3.00 and 15.00.

So on output tokens, DeepSeek V4 Flash is roughly ninety times cheaper than Claude Opus 5 and about thirty five times cheaper than Sonnet 5. Even comparing flagship to flagship, GLM-5.2 at 4.40 against GPT-5.6 Terra at 15.00 is a bit over three times cheaper. If you are running a pipeline that classifies ten thousand support emails a night, or generating product descriptions at volume, that difference is your entire margin. It is not a rounding error and it is not hype. Go and test it.

One footnote worth knowing before you build a spreadsheet: per-token prices are not directly comparable across vendors, because vendors tokenise differently. Anthropic’s own pricing documentation notes that the tokeniser introduced at Claude 4.7, and used by everything since including Opus 5, produces roughly 30 percent more tokens for the same text than the one before it. So a price cut can hide a token increase, and the same paragraph costs a different number of tokens on every platform. Compare cost per finished job, not cost per token.

The saving mostly does not reach you

Here is the part that gets skipped. Almost nobody reading this is billed per token.

Claude Pro is 20 dollars a month. ChatGPT Plus is 20 dollars a month. What you get for that is not twenty dollars of API usage. Twenty dollars at Opus 5 output rates buys you roughly 800,000 output tokens, and that assumes you spend the entire amount on output and nothing on the input, which is never how it works. Anyone using one of these tools seriously for a working day of coding or research goes through that in a sitting. The consumer plans are sold below what the usage costs, because the labs are buying habit and market share. You are already on the subsidised side of the deal.

Swapping to a Chinese model does not move you from expensive to cheap. It moves you from one flat monthly fee to another, because the Chinese labs sell subscriptions too, on the same logic. Z.ai’s GLM Coding Plan starts, at the time of writing, at 18 dollars a month, currently discounted to about 12.60 on an introductory promo. That is a few dollars a month against a plan you already know how to use, in exchange for relearning how a different model behaves. Put like that, the trade is obvious.

The distinction is simply which side of the meter you sit on. If your company is calling an API, tokens are your bill and the model choice is a finance decision worth taking seriously. If you are an individual on a plan, the model choice is a preference, and the pricing argument you read on X does not apply to you.

Wrappers are a second layer you cannot see through

There is a middle case: the generation platforms, the video and image tools, the agent builders. Higgsfield is a fair example, and the same shape applies across dozens of them.

You do not buy tokens there. You buy credits. At the time of writing, Higgsfield’s Starter plan is 15 dollars for 200 credits a month and Plus is 49 dollars for 1,000. Each generation costs a fixed number of credits depending on the model and resolution. That is a reasonable way to sell a product, and for a lot of people it is the right purchase.

But notice what the credit does. It sits between you and the token, and the exchange rate is set by the vendor. When the wrapper swaps in a cheaper model underneath, the credit price does not usually move. When usage patterns change, the definition of a tier can move instead: “unlimited” on these platforms has been redefined more than once this year, and fair-use caps tend to live in the terms rather than on the pricing page. None of that is fraud. It is just that you are pricing in a currency the seller mints.

So yes, if you are pushing large volumes of generated media through a wrapper, it is worth checking what it runs on and whether calling a cheaper model directly gets you the same result. That is the same question we ask about any subscription: what is actually underneath this, and am I paying for the wrapper or the work. Often the wrapper earns its keep. Sometimes it is a credit meter over an API you could call yourself.

The honest counter-argument

The case against everything above is that the gap at the top has genuinely narrowed. For a large share of ordinary work, summarising, drafting, classifying, straightforward code, a good open-weight model is close enough that you would struggle to pick the winner blind. That is true, and it is the strongest argument for switching. It is also why we think the choice matters less than people assume: when the models are close, the thing that differentiates your day is the tooling around them.

There are cases where we would move without hesitating. High-volume API work, where the token price is the business case. Anything with a data residency or privacy requirement that self-hosting solves. Workloads where you want no dependency on a single vendor’s roadmap. If one of those describes you, run the test properly and do not let brand loyalty make the decision.

Pick a harness and stop shopping

For everyone else, our position is boring and we hold it deliberately. Pick the harness you work best in and stay there long enough to get good at it. Claude is our daily driver, not because we think it wins every benchmark, but because the tooling around it fits how we work and we have stopped spending attention on the question.

Every switch has a cost nobody puts on the invoice. A new platform has its own quirks, its own failure modes, its own way of handling long context and tool calls, its own personality you have to learn to prompt around. A week of that is a week not spent shipping. Do it three times a year and you have spent a month becoming a beginner on purpose.

The way to make that safe is portability, not paralysis. Keep the things you have built in formats that travel: prompts and agent instructions as plain markdown files, tool access through open standards like MCP rather than a vendor’s proprietary plugin format, your context and reference material in a folder you own rather than inside someone’s cloud workspace. Build that way and switching harnesses becomes a weekend, not a rebuild. You get to hold your position and change your mind cheaply, which is the only kind of flexibility worth paying for.

A minimalist tool stack is not a lack of ambition. It is what lets you spend your attention on the work instead of on the tools, and the studios shipping the most are almost never the ones with the newest stack.

If you want a second opinion on what your AI stack is actually costing you, and whether any of it should move, that is a conversation we have most weeks. It usually ends with fewer tools rather than more.

Rabbit Hole Digital

AI, automation, and custom software for technical buyers.

Talk to us

Tell us what's eating your week.

Not sure what you need? Neither are most people when they call us. Tell us where the time goes and we'll tell you whether we can help. We take on a small number of projects each quarter, and we reply within two working days.