API changelog
What changed
Additive changes only. Anything that alters limits or behaviour for existing integrations is announced here before it ships. Limits only go up.
Rate limits rebuilt around paying customers
27 September 2026- Paid accounts: 600 requests/min per key and 1,800/min per account (was 120 and an undocumented 180). Admin and internal accounts always get paid limits.
- The per-IP limit is gone. Several sites on one server never share a quota.
- Token budgets (1M tokens/min per key on chat) replace request counts as the real protection for GPU capacity; they never bind before the request limit at normal sizes.
- Every inference response now carries x-ratelimit-* headers (OpenAI/Groq names) with Together-style aliases, so SDKs pace themselves.
- 429 bodies carry a code naming the limit, the real Retry-After in seconds, your accepted and rejected counts for the last minute, and a plain statement when a client is retrying in a loop.
- Paying customers are admitted ahead of free-tier accounts when the server is near capacity. Published limits never shrink.
- A sustained storm of rejected requests now alarms us and emails the account owner within minutes.
Tool calling, embeddings, honest errors
27 September 2026- Tool calling round-trips correctly: tool_calls on assistant messages, tool_call_id on tool messages, finish_reason: tool_calls, and streamed delta.tool_calls with index.
- POST /v1/embeddings on self-hosted models at $0. text-embedding-3-small/-large and text-embedding-ada-002 are accepted as aliases; encoding_format: base64 and dimensions are supported.
- Every 4xx error carries a stable code and a docs_url. 401 and 402 responses include the URL to act on (api_keys_url, topup_url) and the model and minimum top-up involved.
- New Limits page in the dashboard: your tier, every limit, live headroom per key, 2xx vs 429 per hour, and a filterable request log with CSV export.
- Public status page at /status with per-model latency, 24-hour uptime, and gateway p50/p95.