Kimi K3 Max is here: Moonshot's new frontier flagship
Moonshot AI quietly shipped Kimi K3 Max — a reported 2.8-trillion-parameter, 1M-context flagship — and it debuted at #1 on the Frontend Code Arena while pricing 40% below the competition. Here's what's confirmed, what's rumor, and why a model-agnostic scanner is watching.
By The Xalgorix Team
On July 16, 2026, Moonshot AI didn't hold a launch event or drop a 40-page technical report. It simply updated kimi.com — and anyone logging in found a new flagship waiting, in two flavors: K3·Max and K3 Cluster·Max. It is the most aggressive release the Kimi line has shipped, and the numbers are startling.
We build a model-agnostic scanner, so we watch frontier launches closely: the model is a component Xalgorix routes to, and a jump in long-horizon coding and tool-call reliability is exactly the kind of thing that changes what an autonomous pentester can do. Here's what's confirmed on day one, what's still rumor, and where it fits.
What exactly is Kimi K3?
Per Moonshot's announcement, the headline specs are:
- ~2.8 trillion parameters — described as the largest model shipped by a Chinese AI lab, roughly 2.8× the size of its predecessor, K2.
- 1 million token context window — enough to hold entire codebases, book-length documents, or days of agent trajectory in a single prompt.
- Native multimodality — text, image, and video understanding built in, not bolted on.
- Always-on reasoning — K3 thinks before it answers, with a configurable reasoning-effort level via the API.
For context, public estimates put Anthropic's Claude Opus 4.8 at roughly 1.5–2 trillion parameters — so on paper, K3 is the biggest frontier-class model anyone has publicly acknowledged shipping. (Parameter counts here come from the launch announcement, not a formal model card — see the caveat below.)
Max vs. Cluster·Max
- K3·Max — the flagship itself: the direct, single-model experience for coding, writing, analysis, and deep research.
- K3 Cluster·Max — K3 paired with Moonshot's agent-swarm infrastructure, which coordinates many sub-agents in parallel to decompose and execute large tasks. This is the “give it a project, not a prompt” mode.
Why it's fast despite being enormous
A 2.8T model with a 1M-token window should be unusably slow. Two pieces of engineering Moonshot has been developing in the open are why it isn't:
- Kimi Delta Attention (KDA) — a hybrid linear-attention design that, in published results on the same architecture, delivered large KV-cache reductions and roughly 6× decoding throughput at long context. In plain terms: K3 can actually use its million-token window without the speed falling off a cliff.
- Attention Residuals — a training-side change Moonshot credits with ~25% higher training efficiency at under 2% additional cost. You don't see it directly, but it's part of how a 2.8T model gets trained under real compute constraints.
The benchmarks — signal vs. noise
Confirmed: K3 debuted at #1 on the Frontend Code Arena with 1,679 points, taking the top spot in 6 of 7 categories — a seventeen-place jump from K2.6. Moonshot has also shown “vision-in-the-loop” demos: K3 iterating between writing code and inspecting live screenshots until a working interface emerges.
Reported but unverified: leaked benchmark data circulating in the developer community suggests K3's coding ability surpasses Claude Opus 4.8 and GPT-5.5. Treat those as rumor until Moonshot publishes a model card.
Pricing: flagship brains, mid-tier money
This is where K3 gets genuinely disruptive. API pricing, per 1M tokens:
| Model | Input | Output | Cache hit |
|---|---|---|---|
| Kimi K3 | $3.00 | $15.00 | $0.30 |
| Claude Opus 4.8 | $5.00 | $25.00 | $0.50 |
| Claude Fable 5 | $10.00 | $50.00 | $1.00 |
| Kimi K2.6 (previous gen) | $0.95 | $4.00 | $0.16 |
Read that twice: K3 lands around Claude Sonnet's tier while competing on capability with models that cost 40–70% more. If the early signals hold, that's Opus-class intelligence at Sonnet-class prices. And K2.6 isn't dead — at $0.95/$4.00 it remains one of the best price-performance deals in AI. K3 is the flagship; K2.6 is the workhorse.
What you can actually do with it
- Agentic coding — long, multi-hour sessions with thousands of tool calls are the Kimi line's signature strength, and K3 pushes tool-call reliability and vision-in-the-loop iteration further.
- Million-token knowledge work — feed it a monorepo, a legal archive, or a quarter of research notes and ask cross-document questions.
- Swarm-style projects — Cluster·Max for “build the whole thing” tasks where sub-agents divide the labor.
- Multimodal analysis — native video and image understanding for design reviews, UI debugging from screenshots, and video summarization.
Where Xalgorix fits
Xalgorix is model-agnostic by design: the engine routes across frontier models and runs against any OpenAI-compatible provider — you bring a key and pick a model, Moonshot included. So a stronger long-horizon coder with more reliable tool calls is directly useful to us: those are the exact bottlenecks in autonomous, multi-hour offensive work.
But the model is a component, not the product. As we've written before, a scanner's value isn't in what it detects — it's in what it proves. A better base model raises the ceiling on exploitation and reduces wasted iterations, but the discipline around it — exploit, then independently re-verify — is what turns a finding into evidence. New frontier models make that discipline cheaper to run, not optional.
Should you switch?
- Try K3 now if you're doing frontend/full-stack work, long-document analysis, or agentic workflows where tool-call reliability is the bottleneck. The free tier on kimi.com makes it a zero-risk experiment.
- Wait for the model card before an enterprise commitment — pricing is confirmed, independent benchmarks aren't.
- Stay on K2.6 if cost is the primary constraint; it's still absurdly cheap for its capability.
The bigger story: a Chinese lab just shipped the largest publicly acknowledged model on the planet, ranked it above Anthropic's best on a public leaderboard within 24 hours, and priced it well below the competition. The frontier isn't a two-country race anymore — and for anyone building on top of these models, that competition is the best news of all.
Sources
- Odaily — “Moonshot Kimi K3 Goes Live and Open for Use”: odaily.news
- @Kimi_Moonshot on X — official K3 announcement, Frontend Code Arena results, agent swarm: x.com/Kimi_Moonshot
- Kimi API Platform — K3 pricing and model capabilities: platform.kimi.ai
- Claude platform docs — model pricing: platform.claude.com
Figures marked reported/leaked are unverified and attributed to the sources above; this post will be updated when Moonshot publishes an official model card.
Ready to see it prove a bug?
Start a scan — from $1 →
xalgorix