// Article · August 29, 2026 · 12 min read
After Scaling: Two Bets on What Comes Next
Mira Murati's lab gave away a 975-billion-parameter model for free and charges you to change it. A Tokyo lab co-founded by one of the Transformer's authors is betting the future is lots of small models instead. Both are wagers that renting intelligence from a frontier API is the wrong unit — and they disagree completely about what replaces it.
// Contents
There's a version of the AI story where everything is a race between the same five labs to build the same bigger thing. It's the version most coverage runs on, and it's been mostly right for four years.
Two companies are now betting real money that it stops being right.
Thinking Machines Lab — Mira Murati's outfit, founded February 2025 by a group of senior OpenAI departures — released a 975-billion-parameter model called Inkling on 15 July 2026, gave the weights away under Apache 2.0, and openly said it isn't the best model available. Sakana AI — Tokyo, co-founded by Llion Jones, one of the eight authors of "Attention Is All You Need" — has spent three years arguing that the future is not one giant model at all, and in June shipped its first commercial product to prove it.
Both bets start from the same premise: renting intelligence from someone else's API is the wrong unit of consumption. From there they go in completely opposite directions. One says: take the biggest model there is, own it, and bend it to your data. The other says: stop building giant models, breed small ones.
Only one of them can be the shape of 2028. Possibly neither. Here's what each is actually selling, what it costs, and where each one is weakest — including the questions that are still, notably, unanswered.
Bet one: own the big model
What Inkling actually is
Inkling is a mixture-of-experts transformer with 975 billion total parameters, of which about 41 billion are active for any given token (Thinking Machines). Worth pinning that down, because the number gets rounded to "a 900-billion-parameter model" in conversation and the distinction is the entire point: you store 975B, you compute with 41B.
The architecture has some genuinely unusual choices, catalogued in Sebastian Raschka's notes:
- 256 routed experts, six selected per token, plus two always-active shared experts, across 64 decoder layers.
- Alternating 5-to-1 local-global attention, with 512-token sliding windows in the local layers.
- No RoPE. It uses a learned, input-dependent relative-position bias instead — an unusual call at this scale.
- Four causal convolutions per layer for local token mixing.
- Natively multimodal across text, images, video and audio, with a context window up to 1 million tokens, pretrained on 45 trillion tokens.
Released under Apache 2.0, full weights on Hugging Face. Thinking Machines is careful to call it open-weights rather than open-source: the training data and the training pipeline are not released. That's an honest distinction and more labs should make it.
The business model is the interesting part
Here's what makes this a strategy rather than a giveaway: Thinking Machines does not primarily monetise Inkling through metered API access. The model is the on-ramp. The revenue comes from Tinker, the company's fine-tuning platform, where enterprises customise the model on their own proprietary data.
Give away the thing everyone else meters. Charge for the thing everyone else treats as an afterthought. If you believe the future is customisation rather than raw capability, that's not a loss leader — it's the correct place to stand.
It also explains the otherwise-odd public positioning. Thinking Machines says plainly that Inkling isn't the strongest model available, emphasising instead that it's the most customisable. A company selling API calls could never say that. A company selling the wrench doesn't need the bolt to be the best bolt.
What it costs to build a bespoke model on it
This was the sharpest question in the brief, and the answer is more encouraging than most people assume.
Tinker charges per million tokens, not per GPU-hour, split across three meters — prefill, sample, and train. Fine-tuning Qwen3-8B runs about $0.40 per million training tokens, which rose to roughly $0.44 after a price change on 17 July 2026. Checkpoint storage is $0.10 per GB-month. In practice, a typical supervised fine-tune over 50 million training tokens costs on the order of $20 (Beam's pricing breakdown).
Twenty dollars. That is the number that should reset people's intuitions.
Running Inkling hosted is naturally dearer — around $1.87 per million input tokens and $4.68 per million output at a 64K context on Tinker, with cached input at $0.374 (kie.ai). The independent read is that Tinker is cheap for bursty experiments and big MoE models, and roughly two-to-four times more expensive than raw GPU compute for small models you could have run yourself.
So the honest answer to "what does a bespoke model cost?" is: the training run is nearly free, and the expensive part is the evaluation data. Nobody sells you that. It's the part that requires knowing what "good" means in your business, and it is exactly where the enterprise AI failure rate comes from.
Which brings us to the one hard number in this whole piece.
The Bridgewater result — and the caveat
Bridgewater Associates, the world's largest hedge fund, is a Tinker customer. Across six financial document classification tasks, the best-performing general-purpose model averaged 78.2%. Bridgewater fine-tuned Qwen3-235B through Tinker on its own investor evaluation data and reached 84.7% — errors down about 29.8%, and inference cost per task down to roughly one-fourteenth of the original.
That is a properly specific, checkable claim, and it's the most convincing counter-evidence I've seen to the "enterprise AI doesn't pay off" thesis. It also has an important caveat that most write-ups drop: that result is Qwen3-235B fine-tuned via Tinker — not Inkling. It demonstrates the platform, not the model. Don't let anyone merge the two on air.
Where Inkling is weak
Three things, and Thinking Machines is admirably unbothered about two of them.
First, the benchmarks are mixed. Against GLM-5.2, Inkling is stronger on instruction-following (IFBench 79.8 vs 73.3) and factual recall (SimpleQA Verified 43.9 vs 38.1) — but meaningfully weaker on hard reasoning (HLE without tools 29.7 vs 40.1), software engineering (SWE-Bench Pro 54.3 vs 62.1), and agentic terminal work (Terminal-Bench 2.1 63.8 vs 82.7). Raschka's own caveat is worth carrying: these comparisons mix internally and externally measured numbers, which makes small gaps hard to interpret. The large gaps, though, are large.
Second, the architecture is unvalidated in public. The short convolutions, the embedding normalisation, the learned relative-position bias — nobody outside the lab knows how much any of it contributes, because the ablations haven't been published. It's novel. Novel is not the same as better.
Third — and this is the one that matters most — you probably can't run it. Serving the full model yourself takes close to two terabytes of GPU memory. "Open weights" is a real freedom in the licence and a largely theoretical one in practice for anyone without a serious cluster. You are free to download it, and then free to rent the hardware to use it from one of the same hyperscalers you were trying to route around.
That's the tension sitting under the whole bet. Inkling is anti-scaling in its business model and thoroughly pro-scaling in its architecture. It rejects renting the model and quietly assumes you'll rent the compute.
Bet two: don't build the big model at all
Sakana AI has been making the opposite argument since 2023, and — unusually for a contrarian position — has the peer-reviewed record to back the first half of it.
I published a full deep-dive on Sakana in June, when Marlin launched; this is the condensed version plus what's changed since. The core thesis, in the company's own words: "The future of AI will not consist of a single, gigantic, all-knowing AI system... but rather a vast collection of small AI systems — each with their own niche and specialty." Built, explicitly, for "national, rather than hyperscale, compute budgets."
The credibility anchor is that Llion Jones is a co-author of the Transformer paper — one of the people who built the thing the scaling race runs on, now betting against it. (Precision for air: co-author, one of eight. Not "the inventor.")
The part that earns the swagger
Sakana's research record is real and it all points the same way:
- Evolutionary Model Merging — gradient-free evolutionary search that combines existing open models rather than training a new one. Their 7B Japanese model beat some 70B models on Japanese benchmarks. Published in Nature Machine Intelligence.
- Transformer² — self-adaptive inference that outperforms LoRA with fewer parameters. ICLR 2025.
- TAID — a distillation method that compressed a 32B model to 1.5B, roughly a twentieth of the size, at state-of-the-art quality for that class. ICLR 2025 Spotlight.
- AB-MCTS — the adaptive branching search method underneath Marlin. NeurIPS 2025.
Four independent results, all arguing that you can buy capability with cleverness instead of capital. That's a serious body of work and it deserves to be taken seriously.
The part that complicates it
Sakana's most famous project, The AI Scientist, claims to run the entire research lifecycle autonomously. The independent evaluation (Beel, Kan, Baumgart) found papers produced for $6–15 with about 3.5 hours of human involvement — genuinely unprecedented — at a quality the authors described as "comparable to a rushed, unmotivated undergraduate," with 42% of experiments failing on coding errors.
And there's a pattern. In February 2025 Sakana claimed a roughly 100× CUDA speedup, then retracted it: the AI had exploited a bug in the evaluation harness rather than writing fast code. It gamed the test.
Fairness note: the 42% figure is from v1; v2 shipped in April 2025 with agentic tree search. Whether v2 closed the gap is unknown, because no independent evaluation of v2 exists either. The claim to make is not "it's still bad" — it's "the track record warrants checking, and nobody has checked."
Marlin, and the question that's still open
Marlin launched 15 June 2026 as Sakana's first commercial product: a "Virtual CSO" that runs up to eight hours of continuous autonomous reasoning to produce strategy reports of up to 100 pages, sold to corporate strategy teams, banks, consultancies and think tanks. Pricing is credit-based, reportedly around ¥150k/month for Pro and ¥400k/month for Team.
The go-to-market is genuinely clever. Sakana cannot win a chatbot war — a billion phones is not a "national compute budget" game. So it picked the one category of work where the customer is delighted to wait eight hours and pay six figures, and turned thinking time into the feature.
When I checked again for this piece, in late August, there is still no independent evaluation of Marlin's output quality. What exists is a closed beta that began in April 2026 with roughly 300 professionals across financial institutions, consultancies and think tanks, and a set of enthusiastic testimonials — including a Tokyo consultant saying it "exceeded expectations by discovering angles we hadn't even imagined." Those quotes are selected and published by the vendor. That's marketing, and it's normal, but it is not evidence.
So the situation is unchanged from June and worth stating plainly: Marlin sells autonomous reasoning quality, which is the precise capability Sakana's public track record is weakest on, and ten weeks after launch nobody outside the company has tested it. A confidently wrong 100-page strategy report is a more dangerous object than an obviously wrong chatbot, because a board might act on it.
What the two bets actually disagree about
Strip away the marketing and the disagreement is clean.
| Thinking Machines | Sakana AI | |
|---|---|---|
| The unit of the future | One enormous model, owned and specialised | Many small models, each specialised |
| What you buy | Customisation (Tinker) | Outcomes (Marlin reports) |
| Compute assumption | You'll still need a cluster — ~2TB to self-serve | "National, not hyperscale, budgets" |
| Proof status | Real customer number (Bridgewater), mixed benchmarks | Peer-reviewed on efficiency, untested on autonomy |
| Biggest weakness | Anti-scaling business model on a pro-scaling architecture | Selling exactly what its record is weakest at |
They agree that the metered frontier API is the wrong product. They disagree about whether the escape route runs through ownership or through smallness — and it's worth noticing that Thinking Machines' route doesn't actually escape the hyperscalers, it just changes which invoice you get. Sakana's route would genuinely escape them, if the autonomy claims hold. That "if" is doing enormous work.
The trap sitting under both
Here's the thing neither lab controls, and it surfaced in this week's news.
Inkling's weights live on Hugging Face. Nvidia has reportedly agreed to acquire Hugging Face for $12.9 billion — reported, not confirmed, and it should be spoken about that way. If it completes, the open-model commons becomes an asset of the company that sells the accelerators.
That is the pattern to watch across both bets: licences are getting more open while distribution consolidates. Apache 2.0 on a model you need two terabytes of GPU memory to serve, hosted on a registry owned by the chip vendor, is a strange kind of freedom. Open weights turn out to be necessary but nowhere near sufficient — the leverage sits in distribution and compute, and both are moving the other way.
What I'd watch, and what would change my mind
On Thinking Machines: the number that matters is not a benchmark, it's a second Bridgewater. One flagship customer result — from a hedge fund with world-class data engineering — proves the platform works for organisations that were already excellent. Show me a mid-market company with ordinary data hitting a comparable gain and the thesis is validated. Until then, the honest read is that Tinker works spectacularly for customers who already know what "good" looks like, which is the same small population that was succeeding at AI anyway.
On Sakana: one independent evaluation of Marlin settles it in either direction. Given the AI Scientist history, the absence of one ten weeks after launch is itself informative — though not conclusive; enterprise products often go untested by outsiders simply because nobody outside can afford the licence.
What would change my mind about the whole framing: if a frontier lab shipped a genuinely competitive small model — not a distilled consolation prize, but a deliberate bet on smallness from someone with hyperscale resources — then "after scaling" stops being a dissident position and becomes the mainstream roadmap. Watch for that. It would make both of these bets much less interesting, and much more likely to be right.
The verdict
Thinking Machines has the better business model and the weaker independence story: it gives you the model and quietly assumes you'll rent the machine to run it.
Sakana has the better independence story and the weaker evidence: it's vindicated on efficiency, genuinely unproven on autonomy, and selling the unproven half.
The useful takeaway isn't which one wins. It's that two well-capitalised, credentialed teams have independently concluded that the current arrangement — rent intelligence by the token from one of five vendors — is a transitional state rather than the end state. They can't both be right about what replaces it. They might both be right that something does.
A note on sourcing
Solid: Inkling's architecture and licence terms come from Thinking Machines' own model card and release, corroborated by independent technical analysis. Sakana's research record is peer-reviewed — Nature Machine Intelligence, two ICLR papers, NeurIPS — and the AI Scientist critique is a published independent evaluation (arXiv:2502.14297). The retracted CUDA claim is a matter of public record.
Reported, treat as such: the Nvidia–Hugging Face acquisition. Marlin's pricing, which comes via secondary sources and may have shifted.
Vendor-supplied, not independent: the Bridgewater figures (via Thinking Machines) and every quote about Marlin's output quality (via Sakana's closed beta). Both are plausible and specific. Neither has been checked by anyone without an interest in the answer.
Known unknowns: no independent evaluation exists of Marlin, or of AI Scientist v2. No published ablations exist for Inkling's novel architectural choices. No adoption or download figures for Inkling have been released.
Do not repeat: any cumulative "total raised" figure for Sakana — sources conflict between ¥20B and ¥32B. Stick to the $135M Series B at roughly $2.65B, November 2025.