Umans Kimi K3 Experimental
prerelease, seat-gated via Labs
49.1tok/s
throughput · p50 · last 5 min
1.45s
TTFT · p50 · last 5 min
82.36%
uptime · 24h
Kimi K3 in prerelease: Moonshot's largest open-weight release, a 2.8T-parameter mixture-of-experts with a 1M-token context window and native vision, built for repository-scale code understanding and multi-step agentic work. It thinks by default at maximum reasoning effort; select none, low, high, or max to trade depth for speed. While we scale capacity, access is seat-gated through the Labs page, and availability is limited: expect occasional errors during the ramp. For production work today we recommend umans-coder or umans-glm-5.2.
Context
1049K
Max output
131K
Recommended
131K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/high/max
Weights
Trends
Speed over the last 90 days
now 41.6 tok/s
90 days agotoday
best 1.46s · Jul 28now 1.63s
90 days agotoday
Changelog
Events for Umans Kimi K3
Jul 62026
Resolved: back to full capacity Resolved
Hardware capacity was restored and both umans-glm-5.2 and umans-kimi-k2.7 are back to normal speed. During the outage the service ran in a reduced-capacity mode that favoured continuity over speed: requests kept flowing, at the cost of an uneven experience. The affected window is shaded on each model's speed trends.
Jul 22026
Incident: hardware outage, running at reduced capacity Incident
A hardware failure took part of our GPU fleet offline. We failed over to reduced capacity to keep the service available: umans-glm-5.2 and umans-kimi-k2.7 stayed up, but slower than usual and with an uneven experience under load. Live updates were posted on Discord throughout.