🚀 One API for every frontier model on FastInfra

Browse models

Transparent per-token pricing → compare providers on FastInfra

View pricing

OpenAI-compatible chat completions → start in minutes

Read the docs

⚡ Server Auction → Enterprise bare metal from $78.15/mo (173 in stock)

Browse deals

Live · refreshes every minute

partial outage

What our own probes see, unedited: per-model response time right now, 24-hour uptime from one-minute samples, and gateway latency across every customer. Machine-readable at /status.json. Generated 21:27:02 UTC.

Gateway requests (15 min)
0
Gateway latency p50
0 ms
end to end, including generation
Gateway latency p95
0 ms
Server error rate (15 min)
0.00%
gateway up 9h 4 m

Inference backends

Each backend is probed every minute. "Latency" is the time to first token of a one-token completion where the backend supports it, otherwise the model-list round trip. Overflow routing means a degraded backend does not usually mean failed requests for you.

Backend State Latency now Median 24h Uptime 24h Last check
Free-tier models (FastInfra · DEFAULT) down 10,001 ms - 0.00% 21:26:31 UTC
Qwen3.6-27B (FastInfra GPU · H200) operational 266 ms 526 ms 100.00% 21:26:31 UTC
Qwen3.6-27B (FastInfra GPU · H200B) operational 130 ms 255 ms 100.00% 21:26:31 UTC
LTX-2.5 video generation operational 615 ms 467 ms 100.00% 21:26:32 UTC
Whisper speech-to-text + Kokoro TTS · PRIMARY operational 63 ms 61 ms 100.00% 21:26:32 UTC

Related

Changelog · Rate limits · Pricing