umans/status/umans-flash
Live · refreshes every 30s
← all models
Umans Flash Fastest
umans-flash · Qwen3.6-35B-A3B · Qwen
also served as umans-qwen3.6-35b-a3b
Operational
404.4tok/s
throughput · p50 · last 5 min
315ms
TTFT · p50 · last 5 min
99.82%
uptime · 24h

Our fastest model: a light workflow complement, not a standalone coder. Think Haiku next to Opus: not everything needs a frontier model, and Flash's speed (200+ tokens per second) compounds on the roles around umans-coder: gathering context, scout subagents, research, summaries, documentation, and quick edits.

90 days agoin production since May 3, 2026today
Context
262K
Max output
262K
Recommended
33K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/medium/high
Weights
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
peak 521.6 tok/s · Aug 8now 355.3 tok/s
90 days agopre-release before May 3, 2026today
TTFT p50 · time to first token, lower is better
best 319ms · Aug 9now 409ms
90 days agopre-release before May 3, 2026today
Changelog

Events for Umans Flash

incl. gateway-wide announcements
No recent events.
Older events 2
Jul 262026
Resolved: umans-flash (Qwen3.6-35B-A3B-FP8) Resolved
The incident is resolved. Throughout the outage, some requests continued to be served, and a subset of users were affected. Only umans-flash (Qwen3.6-35B-A3B-FP8) model was affected. Incident duration 25min.
Jul 262026
Major outage: umans-flash (Qwen3.6-35B-A3B-FP8) Incident
We are currently experiencing a major outage affecting umans-flash (Qwen3.6-35B-A3B-FP8). Other models are unaffected. Our team is fully mobilized and working to restore normal operation as quickly as possible. We will post updates here as the situation evolves.