Moonshot AI’s new Kimi K3 scored 1,527 points on the AA-Briefcase agentic reasoning benchmark, second overall and ahead of GPT-5.6 Sol Max. The Beijing company released the 2.8 trillion parameter model on July 27 as an open-weight system, which it calls the largest of its kind to date.

On the GDPval-AA v2 benchmark, K3 scored 1,687 points, placing third behind Claude Fable 5 Max and GPT-5.6 Sol Max. It outperformed Claude Opus 4.8 and GPT 5.5 specifically on coding and reasoning tasks.

K3 has a 1 million token context window and native vision capability. Moonshot used a mixture-of-experts design, activating 16 of 896 experts per token to cut computational cost relative to parameter count. No pricing or licensing terms were disclosed.

Sources

Reporting compiled and contextualised by TechScoop Ireland. Figures reconciled across the sources above.

Originally reported by Silicon Republic.