The richest model-maker in the world does not sell his models.
Liang Wenfeng, the founder of the Chinese lab DeepSeek, is now worth around $36bn. That puts him ahead of OpenAI's president Greg Brockman and well ahead of Anthropic's chief executive Dario Amodei on the list of people who have got rich building AI models. And his models are open-weight: anyone can download them, run them, and pull them apart, for nothing.
Read that back. The person who has made the most money building frontier AI gives the actual product away.
It is tempting to file that under "China does things differently" and move on. Don't. Because in the same fortnight, other signals appeared, including the one from Mira Murati - a well-known name in the Western world.

Mira Murati, former OpenAI CTO, founder of Thinking Machines. Her lab's first model, Inkling, is open-weight by design. Photo: Wired
Different cases, same signal
The first is Z.ai, the Beijing lab behind the GLM models. It is on track to become the first independent Chinese AI company to reach roughly $1bn in annual sales, and it is doing it while giving its best model, GLM-5.2, away as a free download. Its API revenue grew about sixtyfold in a year. Giving the model away did not cannibalise the business. It was the business.
The second is closer to home. Mira Murati ran product, and then briefly the whole company, at OpenAI. She left, went quiet for eighteen months, raised one of the largest seed rounds in history, and this month finally shipped her lab's first model, called Inkling. It too is open-weight. And in the launch notes her team did something no marketing department would choose: they said out loud that it is "not the most performant model available today." They built it to be adapted, not to top a leaderboard.
So: the richest builder gives it away. The fastest-growing challenger gives it away. And a founder who learned the closed-model playbook at OpenAI walked out and built an open one. Three different bets, same direction.
That is a good story. It is also, if you run a business, a bill about to change.

Liang Wenfeng, founder of DeepSeek and now the world's richest AI model maker, worth around $36bn. His models are free to download. Photo: AI Speakers
Give the brain away, sell the plumbing
Here is the model underneath, because it is not charity and it is not ideology.
When you publish the weights, the model itself stops being the thing you charge for. It becomes the thing that gets you in the door. Z.ai's revenue does not come from the free download. It comes from what a bank or a state-owned firm needs once it has decided to standardise on that model: running it on their own servers, support contracts, cloud capacity, help fitting it to their own data. The model is the free sample. The plumbing is the invoice.
That is the opposite of the deal most of us actually pay for today. ChatGPT, Gemini and Claude sell you access. A monthly seat, or a meter that ticks per token, or an enterprise licence. You are renting the intelligence, and the whole point is that you can never hold the weights yourself. Which is fine, and for a lot of work it is the right deal. But name the shape of it: one camp rents you the brain, the other gives you the brain and sells you the plumbing around it.
And here is the part that matters for you, which we will come back to: you do not have to own that free brain to win from it. You can rent it too, from someone who runs it well, for a fraction of what the premium one costs.
Free is not the same as cheap
Now the honest part, because there is a lie hiding in the word "free." A free download works on you the way a free sample works in a supermarket: by the time you have tasted it, you have half-decided to buy the trolley behind it.
When you download a model instead of calling an API, you have not removed the cost. You have moved it, off the licence line and onto your own people, your own servers, the specialists who keep the thing running. A couple of weeks ago I made the case that AI now runs on a live meter, a cost per outcome rather than a flat fee. Owning the model does not switch that meter off. It just changes whose meter it is.
For a two-person shop, that is a terrible trade. For an organisation of a thousand people, with a real IT function and data it cannot send to a third party, it can be a good one. This is not "go and train a model", you will not, and you should not. It is that for the first time you have a genuine choice on your most sensitive, highest-volume work: rent it, or own it. So what does each actually cost?
Put real money on it
Picture that thousand-person company. Not a tech firm, a manufacturer or an insurer or a professional-services outfit. Say three hundred of them lean on an AI assistant through the day, drafting, summarising, answering questions, and one back-office process, invoice matching or claims or document triage, runs a few thousand times a day on top. Round it off and you are looking at something like a billion tokens a month, the little chunks of text these models charge by. Real usage, not a toy.
Here is what that same billion costs three different ways, on 2026 list prices.
Rent it from a premium provider, ChatGPT or Claude or Gemini, and you are looking at roughly £5,000 to £8,000 a month. Rent the exact same capability as an open model from a cheap host, one of the providers already running DeepSeek or GLM on their own optimised hardware, and the bill drops to a few hundred pounds a month, for output most people cannot tell apart on everyday work. Run it yourself and it costs more, not less: a few thousand a month for rented GPUs, or north of £200,000 to buy the hardware outright.
Because the hardware is only the electricity. The thing that makes a model actually yours is the customisation, and that is a project, not a compute line: fine-tuning on your own data runs from £25,000 to well past £100,000, and it is never one-and-done. Then come the people nobody puts on the slide, two to four ML specialists at £90,000 to £100,000 each, which is what pushes running your own past £20,000 a month all in. It will not even feel faster: a single box slows under load while the big providers stay quick. Own the model and you own its salaries and its idle time too.
Start with the bill, not the headcount
Here is where most advice gets it wrong. The question is not how many people you employ. A hundred-and-fifty-person hedge fund can burn more tokens than a three-thousand-person builder, so headcount and turnover tell you almost nothing. The number that decides this is your monthly AI bill, what you already spend on inference, plus one override for data you are not allowed to move.
The bands, drawn from where the 2026 break-even analyses actually converge:
Under about £2,000 a month: rent premium and stop. ChatGPT, Copilot or Claude. Your volume will not keep a rented GPU busy, and one engineer to run it costs more than the bill you are trying to cut.
Two to twenty thousand a month: rent open. Take your heaviest, most repetitive workloads and move them to a cheap open-weight host, Together, Fireworks, DeepInfra, the firms that already run DeepSeek and GLM for you. You get most of the saving and none of the servers. This is the right answer for the large majority of mid-market companies, and the step almost all of them skip.
Twenty to sixty thousand a month: go hybrid. Now it is worth dedicating a tuned open model to your single biggest, steadiest workload, the invoices or the claims, and routing the clever tail back to a premium API. This is the point where tailoring earns its keep.
Above about sixty thousand a month, sustained: run the full private setup on dedicated cloud, with the fine-tuning and the people the earlier maths laid out, but only after a hard-headed total-cost check, because idle capacity and salaries, not the token price, are what usually sink it.
And the override that ignores every number above: if you hold data that legally cannot leave your walls, legal privilege, NHS or FCA records, defence, you go private for that workload whatever it costs, and keep everything else on the cheapest thing that fits.
Where Murati and the Chinese labs make their money
Which brings us back to the names from the top, because this is exactly what they sell.
The Chinese labs sell the model two ways. Straight off their API it is startlingly cheap: DeepSeek is about $0.27 for a million tokens in and $1.10 out, a fraction of the Western premium. Want it private, and you either self-host the free weights or buy a dedicated deployment, negotiated quietly, and for DeepSeek reportedly worth the call only once your spend heads past a million dollars a year. In between sits managed private cloud: one host will run GLM entirely inside your walls for a flat $7,500 a month, with no token meter at all.
Murati sells something narrower and cleverer. Thinking Machines does not meter Inkling the way OpenAI meters GPT. Its product is Tinker, a platform for fine-tuning the open model on your own data, billed by the token, with a fine-tune of a small model costing as little as twenty dollars of compute. The catch, in their own words: the customisation, and keeping it safe, is your job, and that needs real machine-learning talent. She is not selling you a brain. She is selling you the workshop to shape one.
To sum up: the model is the free sample, and the money, Murati's, Liang's, or yours, is in everything wrapped around it. What decides your move is not your size, it is your monthly AI bill and the sensitivity of your data. Under a couple of thousand a month, rent the finished article. Into the tens of thousands, tailor your own. In between, where most companies actually sit, rent the open model cheaply from someone who runs it well.
Call to action? Pull your real AI spend for last month, the invoice, not the licence count. And think, what should be your next action - based on it.