Kimi K3: Open Frontier Intelligence
Kimi Team, Jianfeng Cai, et al.
Kimi K3 is an open-weight, native multimodal Mixture-of-Experts model with 2.8 trillion total parameters, 104 billion activated parameters, and a context window of up to one million tokens. Its architecture combines Kimi Delta Attention for efficient long-sequence modeling, Attention Residuals for improved information flow across model depth, and Stable LatentMoE, which activates 16 of 896 routed experts per token. Together with refined data and training recipes, these advances provide an approximately 2.5× improvement in overall scaling efficiency over Kimi K2.
The post-training pipeline applies reinforcement learning across general, agentic, and coding domains at multiple reasoning-effort levels, then consolidates the specialized policies through multi-teacher on-policy distillation. Supporting infrastructure enables multi-trillion-parameter pre-training, million-token agentic reinforcement learning, persistent execution environments, and efficient online serving. Evaluations show strong performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. Although gaps to the strongest proprietary systems remain, the released model weights provide an open foundation for future research, deployment, and development.
Download paper from here or view the arXiv page. The Kimi K3 model weights are publicly available.
BibTeX formatted citation:
@misc{kimiteam2026k3,
title={Kimi K3: Open Frontier Intelligence},
author={{Kimi Team} and Jianfeng Cai and others},
year={2026},
eprint={2607.24653},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2607.24653}
}
