Moonshot AI just moved the goalposts for open models. The Beijing lab's new Kimi K3 weighs in at 2.8 trillion parameters — its technical blog calls it the first open 3T-class system, and the largest open-weight model anyone has shipped. The download link is coming: full weights are promised by July 27.
Big model, light footstep
Don't picture all 2.8 trillion parameters firing at once. K3 is a mixture-of-experts design, and it wakes just 16 of its 896 experts per token — about 1.8% of the network. It takes a million tokens of context and reads images natively.
Moonshot isn't claiming the crown, either. By its own evaluation suite, K3 still trails Claude Fable 5 and GPT 5.6 Sol overall. But against everything else — Claude Opus 4.8 and GPT 5.5 included — the company says its model wins on coding and agentic work.
From #18 to first place
There's outside evidence for the coding claim. In Arena's Frontend Code evaluation, where developers rate output without knowing which model made it, K3 debuted at #1 with 1,679 points — ahead of Claude Fable 5. Its predecessor, K2.6, was ranked 18th. One release, 17 places. The July 16 results had it first in six of seven frontend domains.
The catch is the bill. Moonshot's API runs $0.30 per million input tokens when the cache hits, $3 per million when it misses, and $15 per million on output. Kimi K2 launched a year ago at $0.60 per million input — so uncached input now costs five times what it did a generation back. Scale isn't free, even from Beijing.
Built around the export wall
How K3 got built is arguably the bigger story. Moonshot credits a 2.5x scaling-efficiency gain over K2 to two architectural moves: Kimi Delta Attention, a hybrid linear attention scheme, and Attention Residuals, which rewire how information moves between layers. Quantization-aware training kicks in at the supervised fine-tuning stage — MXFP4 weights, MXFP8 activations — picked, the company says, for broad hardware compatibility.
"Broad hardware compatibility" is a polite way of saying: we can't count on Nvidia's best. The kernel benchmarks ran on Nvidia's H200 and on something the blog only calls a "GPGPU from an alternative vendor." MiniTriton, a compiler K3 built from scratch, gets charted on the L20 — the cut-down card Nvidia is allowed to sell into China. And the blog stays quiet about where those H200s physically are, which is worth noticing: in January, Congress closed the offshore cloud rental loophole that had let Chinese firms rent restricted chips remotely. To serve the thing, Moonshot recommends supernodes of 64-plus accelerators so expert traffic never leaves one high-bandwidth domain.
Bank of America analysts led by Alex Liu, in a note cited by CNBC, drew the obvious conclusion: heavyweight pre-training plus clever architecture can still produce step-change gains for China's flagship models, export controls or not.
The two-day chip designer
One case study stands out. Moonshot let K3 run autonomously for 48 hours and asked it to design a simulated inference chip for a nano model based on its own architecture, using open-source EDA tools and the Nangate 45nm library. It came back with a design that closed timing at 100 MHz inside four square millimeters — 1.46 million standard cells, an INT4 MAC array, and over 8,700 tokens per second of simulated decode.
The date to watch
Closed frontier models launch every few months. A near-frontier model you can download, inspect, fine-tune, and run yourself is rarer — and at 3T-class, unprecedented. If the weights land on July 27 as promised, the debate shifts. The question was whether China could engineer around U.S. compute limits. Now it's how much of the frontier ends up free.
Image: via Unsplash



