Kimi K3 on ai&
Kimi K3 is being called “the clearest example yet of an open-weight model reaching frontier-level performance.” Moonshot AI’s newly released 2.8-trillion-parameter model scores 57 on the Artificial Analysis Intelligence Index — behind only Claude Fable 5 (60) and GPT-5.6 Sol (59), and ahead of every other model tested, including Claude Opus 4.8.
About ai&
ai& Inc. (HQ: Yokohama; CEO David Bennett; President Shimpei Hara) is a frontier AI company, vertically integrated across inference. Our platform, ai& Inference, is a domestic alternative to overseas services like Claude and GPT — cutting AI costs by up to 80%* while keeping all processing inside Japan. Today we're launching Kimi K3 — 初上陸, live now via our Inference API.
Frontier performance, fully open
K3 leads open-weight models across agentic search, reasoning, coding, and document understanding. Some highlights from Moonshot AI's published benchmarks:

Kimi K3 benchmark scores. Source: Moonshot AI, Kimi K3 Tech Blog (July 2026).
It still trails the very best closed models — Claude Fable 5 and GPT 5.6 Sol — on the hardest reasoning tests. But the gap between open and proprietary frontier AI has never been this narrow, and K3 is the first model this capable that anyone can run, inspect, and self-host.
How it stacks up
Head-to-head against the two leading closed models, K3 is at parity or ahead on most benchmarks — and honest about where it isn't yet:

Kimi K3 vs. Claude Fable 5 vs. GPT 5.6 Sol. Source: Moonshot AI / Kimi K3 benchmark suite (July 2026).
Independent third-party benchmarking backs this up. On the Artificial Analysis Intelligence Index, Kimi K3 scores 57 — just behind Claude Opus 5's 61 — while costing a fraction as much per task:

Kimi K3 vs. Claude Opus 5. Source: Artificial Analysis Intelligence Index (July 2026).
Frontier-tier intelligence, at a fraction of the cost.
Pricing
We serve K3 at $3.00 (¥400) / MTok input and $13.00 (¥2,080) / MTok output — well under Claude Opus 5's list price on both:

List price, $ per million tokens. Claude Opus 5 pricing per Anthropic's published rate card (July 2026).
Day-0 on heterogeneous silicon
We stood up K3 on both NVIDIA and AMD from day one, running real heterogeneous serving across two different GPU architectures simultaneously via vLLM and SGLang.
This is day one. If you're building with open-weight models — or want to start — let's talk.