Live · refreshes every minute
partial outage
What our own probes see, unedited: per-model response time right now, 24-hour uptime from one-minute samples, and
gateway latency across every customer. Machine-readable at /status.json.
Generated 21:27:02 UTC.
Inference backends
Each backend is probed every minute. "Latency" is the time to first token of a one-token completion where the backend supports it, otherwise the model-list round trip. Overflow routing means a degraded backend does not usually mean failed requests for you.
| Backend | State | Latency now | Median 24h | Uptime 24h | Last check |
|---|---|---|---|---|---|
| Free-tier models (FastInfra · DEFAULT) | down | 10,001 ms | - | 0.00% | 21:26:31 UTC |
| Qwen3.6-27B (FastInfra GPU · H200) | operational | 266 ms | 526 ms | 100.00% | 21:26:31 UTC |
| Qwen3.6-27B (FastInfra GPU · H200B) | operational | 130 ms | 255 ms | 100.00% | 21:26:31 UTC |
| LTX-2.5 video generation | operational | 615 ms | 467 ms | 100.00% | 21:26:32 UTC |
| Whisper speech-to-text + Kokoro TTS · PRIMARY | operational | 63 ms | 61 ms | 100.00% | 21:26:32 UTC |