Moonshot AI’s new Kimi K3 scored 1,527 points on the AA-Briefcase agentic reasoning benchmark, second overall and ahead of GPT-5.6 Sol Max. The Beijing company released the 2.8 trillion parameter model on July 27 as an open-weight system, which it calls the largest of its kind to date.
On the GDPval-AA v2 benchmark, K3 scored 1,687 points, placing third behind Claude Fable 5 Max and GPT-5.6 Sol Max. It outperformed Claude Opus 4.8 and GPT 5.5 specifically on coding and reasoning tasks.
K3 has a 1 million token context window and native vision capability. Moonshot used a mixture-of-experts design, activating 16 of 896 experts per token to cut computational cost relative to parameter count. No pricing or licensing terms were disclosed.
Sources
- TechScoop wire desk: benchmark scores, model specifications and release details.
Reporting compiled and contextualised by TechScoop Ireland. Figures reconciled across the sources above.
Originally reported by Silicon Republic.