umans/status
Live · updated just now

All systems operational

Live status and speed of the Umans Code gateway and its models, refreshed every 30 seconds. For each model we show its output speed and median time to first token.

3 / 3production models operational
98.89%gateway uptime · 90 days
1model in testing
0active issues
Gateway

API endpoint

the front door every model is served through
API gateway Operational
api.code.umans.ai
98.75%
service uptime · 24h
1.66s
median TTFT · all models
90 days agotoday
Production

Models in production

3 listed

tok/s = output tokens per second · TTFT = time to first token · p50 = median over the last 5 minutes. On each gauge the midpoint is that model's target; the marker sits further right when it's beating target (faster TTFT, higher throughput).

umans-glm-5.2 · GLM
Operational
61.4tok/s
throughput · p50 · last 5 min
1.86s
TTFT · p50 · last 5 min
99.93%
uptime · 24h

GLM 5.2 is our best model for coding right now, with a 400K context window for large codebases. Vision is available on the Anthropic Messages API (`/v1/messages`) only, through a server-side handoff (GLM 5.2 generates the text, Kimi preprocesses the image); that handoff will be retired soon in favour of more efficient client-side image handling.

90-day speed trends & events →
90 days agoin production since Jun 21, 2026today
Context
406K
Max output
131K
Recommended
131K
Vision
Via handoff
Tools
Yes
Reasoning
Toggle · none/high/max
Weights
umans-kimi-k2.7 · Kimi K2.7-Code · Moonshot
successor to Kimi K2.6
also served as umans-coder
Operational
99.6tok/s
throughput · p50 · last 5 min
1.10s
TTFT · p50 · last 5 min
99.44%
uptime · 24h

Kimi K2.7-Code via Umans Code - Moonshot's strongest coding model and the successor to Kimi K2.6. Built for complex, tool-heavy agentic coding; it reasons more efficiently than K2.6, so agent sessions run faster at the same depth.

90-day speed trends & events →
90 days agoin production since Jun 12, 2026today
Context
262K
Max output
262K
Recommended
33K
Vision
Yes
Tools
Yes
Reasoning
Always on
Umans Flash Fastest
umans-flash · Qwen3.6-35B-A3B · Qwen
also served as umans-qwen3.6-35b-a3b
Operational
295.5tok/s
throughput · p50 · last 5 min
834ms
TTFT · p50 · last 5 min
99.08%
uptime · 24h

Our fastest model: a light workflow complement, not a standalone coder. Think Haiku next to Opus: not everything needs a frontier model, and Flash's speed (200+ tokens per second) compounds on the roles around umans-coder: gathering context, scout subagents, research, summaries, documentation, and quick edits.

90-day speed trends & events →
90 days agoin production since May 3, 2026today
Context
262K
Max output
262K
Recommended
33K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/medium/high
Weights
Testing

In the playground

1 · try before it's promoted
What "playground" meansPlayground models are short, experimental test runs: not permanent, offered at low capacity, and expected to be flaky under load. They're here so you can push them and tell us what you find. Uptime and speed are measured the same way as production, but don't build on them. For real work, use the production model each one is based on.
Umans Kimi K3 Experimental
umans-kimi-k3 · Kimi K3 · Moonshot
prerelease, seat-gated via Labs
In testing
47.7tok/s
throughput · p50 · last 5 min
1.82s
TTFT · p50 · last 5 min
82.36%
in testing

Kimi K3 in prerelease: Moonshot's largest open-weight release, a 2.8T-parameter mixture-of-experts with a 1M-token context window and native vision, built for repository-scale code understanding and multi-step agentic work. It thinks by default at maximum reasoning effort; select none, low, high, or max to trade depth for speed. While we scale capacity, access is seat-gated through the Labs page, and availability is limited: expect occasional errors during the ramp. For production work today we recommend umans-coder or umans-glm-5.2.

90-day speed trends & events →
Stage
Playground
Context
1049K
Max output
131K
Recommended
131K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/high/max
History

Gateway uptime

90-day window · daily worst status
90 days ago98.89% operationaltoday
Changelog

Recent events & model lifecycle

releases, retirements, and playground changes
Jul 262026
Resolved: umans-flash (Qwen3.6-35B-A3B-FP8) Resolved
The incident is resolved. Throughout the outage, some requests continued to be served, and a subset of users were affected. Only umans-flash (Qwen3.6-35B-A3B-FP8) model was affected. Incident duration 25min.
Jul 262026
Major outage: umans-flash (Qwen3.6-35B-A3B-FP8) Incident
We are currently experiencing a major outage affecting umans-flash (Qwen3.6-35B-A3B-FP8). Other models are unaffected. Our team is fully mobilized and working to restore normal operation as quickly as possible. We will post updates here as the situation evolves.
Jul 202026
Planned maintenance: Umans GLM 5.2 Maintenance
We're making a change to GLM 5.2 on 2026-07-20, from 11:15 to 12:00 CET. During the window, around 30 to 40 minutes, GLM 5.2 responses may noticeably be slower than usual. Requests still go through, and other models are unaffected. The work is aimed at making GLM 5.2 start responding faster once it lands.
Jul 162026
Playground closed: Umans DeepSeek V4 Pro DSpark Testing
The DSpark test window closed after two days. What we took from it: DSpark speculative decoding lets us serve more users at once while keeping each session fast enough, the DeepSeek V4 architecture is now mature enough to serve at scale, and the model itself is solid. V4 Pro is not joining the lineup though: it is still a preview build, and the issues testers hit (DSML leaks, language bleed, long-context artifacts) are model-side. DeepSeek confirmed a better version is coming this month, so we would rather roll the learnings into that. Thanks to everyone who tested.
Jul 142026
Playground opened: Umans DeepSeek V4 Pro DSpark Testing
umans-deepseek-v4-pro-dspark entered the playground for a short, seat-gated test window: DeepSeek V4 Pro served from the original weights with DSpark speculative decoding. Experimental and temporary; not for production.
Jul 102026
Resolved: Umans GLM 5.2 Resolved
The incident is resolved. Throughout the outage, most requests continued to be served, and a subset of users were affected. Umans GLM 5.2 is still running at reduced capacity, so responses may be slower than usual at peak times. We will post an update once full capacity is restored.
Jul 102026
Major outage: Umans GLM 5.2 Incident
We are currently experiencing a major outage affecting Umans GLM 5.2. Other models are unaffected. Our team is fully mobilized and working to restore normal operation as quickly as possible. We will post updates here as the situation evolves.
Jul 92026
GLM 5.2 performance stabilized, still on reduced capacity Resolved
GLM 5.2 performance has stabilized after deploying the prefill kernel update and selective prioritization for interactive requests. We’re still operating today in reduced-capacity mode while capacity is being restored, but the service is holding and requests are flowing normally. We’ll keep monitoring closely and update here if anything changes.
Jul 92026
Degraded performance on Umans GLM 5.2 Incident
We’re currently seeing degraded performance on Umans GLM 5.2, mainly higher time-to-first-token and uneven streaming speed. Requests are still going through, but the experience can feel slower than usual. We’re investigating and applying mitigations. We’ll update here once things are back to target.
Jul 62026
Resolved: back to full capacity Resolved
Hardware capacity was restored and both umans-glm-5.2 and umans-kimi-k2.7 are back to normal speed. During the outage the service ran in a reduced-capacity mode that favoured continuity over speed: requests kept flowing, at the cost of an uneven experience. The affected window is shaded on each model's speed trends.
Jul 22026
Incident: hardware outage, running at reduced capacity Incident
A hardware failure took part of our GPU fleet offline. We failed over to reduced capacity to keep the service available: umans-glm-5.2 and umans-kimi-k2.7 stayed up, but slower than usual and with an uneven experience under load. Live updates were posted on Discord throughout.
Jul 22026
Playground closed: Umans GLM 5.2 NVFP4 Testing
The short NVFP4 test window ended after four days. Thanks to everyone who pushed it and shared findings.
Jun 292026
Playground opened: Umans GLM 5.2 NVFP4 Testing
umans-glm-5.2-nvfp4 entered the playground for a short, low-capacity test window. Experimental and temporary; not for production.
Jun 242026
Retired: Umans GLM 5.1 Retired
umans-glm-5.1 was retired in favour of GLM 5.2. Requests to the old id now return a clear deprecation error pointing to umans-glm-5.2.
Jun 212026
Released to production: Umans GLM 5.2 Released
umans-glm-5.2 was released as the long-context model, with a 405K context window, after a pre-release period that started Jun 16.
Jun 182026
Retired: Umans Kimi K2.6 Code Retired
umans-kimi-k2.6 was retired and superseded by K2.7. Requests to the old id now return a clear deprecation error pointing to umans-kimi-k2.7.
Jun 122026
Released to production: Umans Kimi K2.7 Code Released
umans-kimi-k2.7 was released as the recommended coding model (also served as umans-coder).
May 132026
Retired: Umans Kimi K2.5 Retired
umans-kimi-k2.5 was retired and superseded by Kimi K2.6.
May 92026
Retired: Umans MiniMax M2.5 Retired
umans-minimax-m2.5 was retired without a direct replacement.
Past models 6 retired
umans-deepseek-v4-pro-dspark · DeepSeek
retired Jul 16, 2026
umans-glm-5.2-nvfp4 · GLM
retired Jul 2, 2026
umans-glm-5.2
umans-glm-5.1 · GLM
retired Jun 24, 2026
umans-glm-5.2
umans-kimi-k2.6 · Moonshot
retired Jun 18, 2026
umans-kimi-k2.7
umans-kimi-k2.5 · Moonshot
retired May 13, 2026
umans-kimi-k2.6
umans-minimax-m2.5 · MiniMax
retired May 9, 2026