Build on FastInfra
OpenAI-compatible REST API for chat, LTX-2.5 video, audio, and more. Swap the base URL to
https://api.fastinfra.ai/v1 — see the endpoint reference and
video API below.
Integration reference
Everything you need to integrate FastInfra — authentication, endpoints, rate limits, provider models, and copy-paste examples.
| Limit | Value |
|---|---|
| Requests per minute (per API key) | 600 |
| Requests per minute (per account, all keys combined) | 1800 |
| Concurrent requests (per API key) | 100 |
| Concurrent requests (server-wide) | 1024 |
Applies to every inference call (chat, video submit, audio, live transcription). There is no IP-based limit.
Accounts without a completed top-up use the smaller free tier
(60/min per key,
120/min per account,
4 concurrent).
See Rate Limits for error handling and retry guidance.
API Overview
FastInfra exposes an OpenAI-compatible REST API at https://api.fastinfra.ai/v1
(same paths as OpenAI: chat, models, audio). Video generation uses a dedicated async
job API — it is not available through POST /chat/completions.
| Capability | Method & path | Notes |
|---|---|---|
| Chat completions | POST https://api.fastinfra.ai/v1/chat/completions |
OpenAI-compatible; streaming supported |
| LTX-2.5 video | POST https://api.fastinfra.ai/v1/videos/generations |
HTTP 202 + job id — poll below. Model: lightricks/ltx-2.5 |
| Video job status | GET https://api.fastinfra.ai/v1/videos/jobs/{id} |
Poll until status is completed or failed |
| Embeddings | POST https://api.fastinfra.ai/v1/embeddings |
OpenAI-compatible; text-embedding-3-* ids accepted as aliases. See Embeddings |
| Speech-to-text | POST https://api.fastinfra.ai/v1/audio/transcriptions |
Multipart upload; see Audio |
| Kokoro TTS | POST https://api.fastinfra.ai/v1/audio/speech |
Standalone text→speech; model hexgrad/kokoro-82m (tts-1). See Kokoro TTS |
| Model catalog | GET https://api.fastinfra.ai/v1/models |
Includes lightricks/ltx-2.5 when video is enabled |
| Model metadata | GET https://api.fastinfra.ai/v1/models/{model_id} |
Validates a model id before use |
Error format
Every error is JSON in the OpenAI shape, plus two fields OpenAI does not send: a stable
code you can branch on, and a docs_url that points at the section of this
page that explains the fix. Billing and auth errors also carry the URL to act on
(topup_url, api_keys_url), and rate-limit errors carry
retry_after_seconds.
{
"error": {
"message": "'qwen3.6-27b' is a prepaid model. API access to prepaid models unlocks after one completed top-up of at least $5 ...",
"type": "payment_required",
"code": "payment_required",
"model": "qwen3.6-27b",
"minimum_topup_usd": 5,
"topup_url": "https://www.fastinfra.ai/Billing",
"docs_url": "https://www.fastinfra.ai/Docs#credits"
}
}
| HTTP | code | Meaning |
|---|---|---|
| 400 | invalid_request, invalid_input | Malformed body or unsupported parameter; the message names the field |
| 401 | invalid_api_key | Missing, invalid, expired, or deactivated key |
| 401 | key_not_linked | Key is not owned by a registered account |
| 402 | payment_required | Prepaid model before the first $5 top-up |
| 402 | insufficient_quota | Prepaid model with an empty wallet |
| 404 | model_not_found | Unknown model id |
| 413 | payload_too_large | Audio upload over 25 MB |
| 429 | rate_limit_*, server_capacity, queue_full | See Rate Limits |
| 502 / 503 / 504 | upstream_error, upstream_unavailable, upstream_timeout | Inference backend failed after failover; retry |
How to determine the API contract
LTX video is not part of OpenAI’s public spec — it uses FastInfra-specific
paths under https://api.fastinfra.ai/v1. Use any of these sources (they stay in sync):
- This page — Video Generations is the canonical human reference: request fields, HTTP 202 submit, poll loop, response JSON, pricing, and error codes.
-
OpenAPI —
/openapi.jsondocumentsPOST /videos/generationsandGET /videos/jobs/{id}for codegen and AI assistants. -
Live model catalog —
GET https://api.fastinfra.ai/v1/modelslistslightricks/ltx-2.5when video is enabled;GET https://api.fastinfra.ai/v1/models/lightricks/ltx-2.5validates the id before you integrate. -
llms.txt —
/llms.txtlinks back here and to OpenAPI for crawlers and LLM tooling.
POST https://api.fastinfra.ai/v1/chat/completions. That endpoint is chat-only; video always goes
through POST https://api.fastinfra.ai/v1/videos/generations → poll
GET https://api.fastinfra.ai/v1/videos/jobs/{id}.
Quick video smoke test
# 1) Submit (returns immediately with HTTP 202)
curl -sS https://api.fastinfra.ai/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"lightricks/ltx-2.5","prompt":"A red balloon floating over a city at sunset.","seconds":5,"size":"1280x704"}'
# 2) Poll (replace JOB_ID from the JSON "id" field; also in Location header)
curl -sS https://api.fastinfra.ai/v1/videos/jobs/JOB_ID \
-H "Authorization: Bearer YOUR_API_KEY"
Authentication
All API requests require an API key. Create one from the API Keys page.
Pass your key using either header:
Authorization: Bearer YOUR_API_KEY
# or
X-Api-Key: YOUR_API_KEY
Point your OpenAI SDK at base_url="https://api.fastinfra.ai/v1" (include the /v1 suffix).
Credits
Paid chat completions and video generations require a positive prepaid balance.
Usage is deducted after a successful response. Video jobs are billed on the first
poll that returns status: "completed" — submitting a job (HTTP 202) is not billed.
Free-tier models keep working at $0.
GET https://api.fastinfra.ai/v1/models is not gated.
When the wallet is empty, paid chat requests return HTTP 402:
HTTP/1.1 402 Payment Required
{
"error": {
"message": "Insufficient API credits. Add funds on the Billing page, or use a free-tier model.",
"type": "insufficient_quota",
"code": "insufficient_quota"
}
}
Rate Limits
Inference requests are rate limited to keep latency stable for every customer. The limits below are
the complete list: anything that can return 429 is on this page. Limits only go up;
they are never lowered under load. Production integrations must handle
429 Too Many Requests and honor the Retry-After response header.
Current limits
| Limit | Paid accounts | Free / unverified accounts | Scope |
|---|---|---|---|
| Requests per minute, per API key | 600 |
60 |
Rolling 60-second window for one key |
| Requests per minute, per account | 1800 |
120 |
Rolling 60-second window across every key the account owns |
| Tokens per minute, per API key | 1,000,000 |
100,000 |
Chat completions only. Prompt estimate + max_tokens (default 2048) is reserved when the request is admitted and settled to actual usage when it finishes |
| Tokens per minute, per account | 3,000,000 |
200,000 |
Across every key the account owns |
| Concurrent requests, per API key | 100 |
4 |
In flight at the same time; a streamed response holds its slot until the stream ends |
| Concurrent requests, server-wide | 1024 |
All API keys combined; a backstop, overflow routing absorbs GPU saturation first | |
| Per IP address | none | Several sites on one server never share a quota | |
Paid means the account has completed a top-up on the Billing page. Admin and internal accounts always get the paid limits. Need more? Per-key limits can be raised on request; a human answers within hours.
What is limited
- Rate limited:
POST https://api.fastinfra.ai/v1/chat/completions,POST https://api.fastinfra.ai/v1/videos/generations,POST https://api.fastinfra.ai/v1/audio/transcriptions,POST https://api.fastinfra.ai/v1/audio/translations,POST https://api.fastinfra.ai/v1/audio/speech, and each live transcription WebSocket session (one concurrency slot for its whole duration) - Not rate limited:
GET https://api.fastinfra.ai/v1/models,GET https://api.fastinfra.ai/v1/videos/jobs/{id}, account pages, and other non-inference routes - Under load: paying accounts are admitted ahead of free-tier accounts. Published limits never shrink; free-tier traffic simply queues behind paid traffic when the server is near capacity.
Rate-limit headers
Every rate-limited response, success or 429, carries your current headroom so your client can pace itself instead of discovering the limit by tripping it. The names are the ones the OpenAI and Groq SDKs, LiteLLM, LangChain and the Vercel AI SDK already read; Together AI's names are included as aliases of the same counters. Remaining values are the tighter of your key and your account.
| Header | Meaning | Alias |
|---|---|---|
x-ratelimit-limit-requests | Requests per minute allowed on this key | x-ratelimit-limit |
x-ratelimit-remaining-requests | Requests left in the current 60 s window | x-ratelimit-remaining |
x-ratelimit-reset-requests | Time until the oldest request leaves the window, e.g. 7.66s or 1m12.4s | x-ratelimit-reset (whole seconds) |
x-ratelimit-limit-tokens | Tokens per minute allowed on this key | x-tokenlimit-limit |
x-ratelimit-remaining-tokens | Tokens left in the current window (reservations included) | x-tokenlimit-remaining |
x-ratelimit-reset-tokens | Time until the oldest token reservation leaves the window | |
x-ratelimit-limit-concurrency | In-flight requests allowed on this key | |
x-ratelimit-remaining-concurrency | In-flight slots currently free | |
Retry-After | On 429 only: real seconds to wait before the request can succeed |
HTTP/1.1 200 OK
x-ratelimit-limit-requests: 600
x-ratelimit-remaining-requests: 599
x-ratelimit-reset-requests: 59.8s
x-ratelimit-limit-tokens: 1000000
x-ratelimit-remaining-tokens: 997640
x-ratelimit-reset-tokens: 59.8s
x-ratelimit-limit-concurrency: 100
x-ratelimit-remaining-concurrency: 99
x-ratelimit-limit: 600
x-ratelimit-remaining: 599
x-ratelimit-reset: 60
x-tokenlimit-limit: 1000000
x-tokenlimit-remaining: 997640
429 responses
Every 429 carries type: "rate_limit_error", a code naming the limit that fired,
the limit value, and a Retry-After header (also repeated as
retry_after_seconds) that is the real number of seconds until the window frees, usually
between 1 and 10. It also reports accepted_last_minute and rejected_last_minute
for your account, and sets client_retry_loop_detected: true when more requests were
rejected in a minute than your entire allowance: that means a retry policy on your side is ignoring
Retry-After, and the message says so in plain words.
Possible codes: rate_limit_key, rate_limit_account, rate_limit_key_tokens, rate_limit_account_tokens, rate_limit_concurrency, server_capacity.
rate_limit_key — too many requests on one key within 60 seconds:
HTTP/1.1 429 Too Many Requests
Retry-After: 3
{
"error": {
"message": "Rate limit exceeded. Maximum 600 requests per minute per API key.",
"type": "rate_limit_error",
"code": "rate_limit_key",
"limit": 600,
"retry_after_seconds": 3
}
}
rate_limit_account — too many requests across all keys on the account within 60 seconds. A request refused here does not count against the key window.
HTTP/1.1 429 Too Many Requests
Retry-After: 2
{
"error": {
"message": "Rate limit exceeded. Maximum 1800 requests per minute across all API keys on this account.",
"type": "rate_limit_error",
"code": "rate_limit_account",
"limit": 1800,
"retry_after_seconds": 2
}
}
rate_limit_concurrency — the key's in-flight cap stayed full for 30 seconds. Requests queue for a slot before this fires, so it is rare in normal operation:
HTTP/1.1 429 Too Many Requests
Retry-After: 5
{
"error": {
"message": "Too many concurrent requests on this API key (limit 100 in flight). Wait for in-flight requests to finish before retrying.",
"type": "rate_limit_error",
"code": "rate_limit_concurrency",
"limit": 100,
"retry_after_seconds": 5
}
}
server_capacity — the server-wide backstop was full for 30 seconds. This is our capacity, not your usage; overflow routing normally absorbs GPU saturation long before this fires. Retry after Retry-After:
HTTP/1.1 429 Too Many Requests
Retry-After: 5
{
"error": {
"message": "Server-wide capacity is temporarily full. This is not caused by your usage; retry after a few seconds.",
"type": "rate_limit_error",
"code": "server_capacity",
"limit": 1024,
"retry_after_seconds": 5
}
}
Video queue full — more than a handful of LTX jobs are already waiting on the GPU. Extra submits should stay queued; this 429 only fires when that queue itself is full. Poll existing job ids; retry the POST after Retry-After.
HTTP/1.1 429 Too Many Requests
Retry-After: 60
{
"error": {
"message": "Video generation queue is full. Retry shortly.",
"type": "rate_limit_error"
}
}
Integration best practices
- On 429, wait the number of seconds in
Retry-Afterbefore retrying, with exponential backoff and jitter if it repeats. Never retry a 429 immediately in a loop: every retry is another request against the same window. - Cache failures on your side for a short period. A page that fans out to many sections should not re-fire every failed section for every visitor.
- Disable your SDK's automatic 429 retry or configure it to honor
Retry-After; the OpenAI SDKs retry twice by default. - Use
"stream": truefor long generations — see Streaming. - Batch related prompts where your app allows, instead of many tiny back-to-back calls.
Retry example (Python)
import time
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
def chat_with_retry(messages, max_retries=3):
for attempt in range(max_retries):
try:
return client.chat.completions.create(
model="llama3.1:8b",
messages=messages,
)
except Exception as exc:
if getattr(exc, "status_code", None) != 429 or attempt == max_retries - 1:
raise
retry_after = int(getattr(exc, "response", {}).headers.get("Retry-After", 60))
time.sleep(retry_after)
Chat Completions
Create a chat completion using the OpenAI-compatible endpoint. Pass any model ID from GET https://api.fastinfra.ai/v1/models or the pricing catalog.
https://api.fastinfra.ai/v1/chat/completions
Request body
{
"model": "llama3.1:8b",
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "Hello!" }
],
"temperature": 0.7
}
Example (local / free-tier model)
curl https://api.fastinfra.ai/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.1:8b",
"messages": [{"role": "user", "content": "Summarize quantum computing in one sentence."}]
}'
Example (catalog model)
Any model ID from the catalog works the same way — routing is handled automatically. See Model Routing.
curl https://api.fastinfra.ai/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "agentica-org/DeepCoder-14B-Preview",
"messages": [{"role": "user", "content": "Explain model routing in one sentence."}]
}'
Response
{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1234567890,
"model": "llama3.1:8b",
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": "Quantum computing uses quantum bits..."
},
"finish_reason": "stop"
}],
"usage": {
"prompt_tokens": 12,
"completion_tokens": 24,
"total_tokens": 36
}
}
reasoning_effort: low). Pass "enable_thinking": true or a higher "reasoning_effort" to opt in. If you omit max_tokens, these backends cap at 2048.
Vision (image input)
Multimodal models accept images alongside text using the standard OpenAI content-parts
format: instead of a string, content is an array of parts, each with a
type of text or image_url. The gateway forwards
parts arrays to the upstream untouched, so anything the model supports works —
including data: URLs for local files.
qwen/qwen3.6-27b is natively multimodal: it reads photos, screenshots,
charts, and diagrams (OCR, chart understanding, visual reasoning) on both the dedicated
GPU route and its wholesale fallbacks. Plain-string content keeps working
exactly as before — the two formats can be mixed across messages in one request.
curl https://api.fastinfra.ai/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen/qwen3.6-27b",
"messages": [{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://example.com/architecture-diagram.png"}},
{"type": "text", "text": "Explain what this diagram shows."}
]
}]
}'
For a local file, base64-encode it into a data URL:
{"url": "data:image/png;base64,<BASE64>"}.
Send one image per request — the dedicated GPU route is tuned for
fast text throughput and caps multimodal input at a single image (video input is
not enabled).
qwen/qwen3.6-27b.
Streaming (recommended for production)
Streaming is a client choice, not something FastInfra forces.
Your app must send "stream": true in the JSON body (or use an SDK with streaming enabled).
FastInfra does not turn on streaming for you on non-streaming requests.
- Your side (API caller): set
"stream": trueonPOST /v1/chat/completions. Tokens arrive as Server-Sent Events (data: …lines). - FastInfra side (automatic): when self-hosted models omit
max_tokens, the gateway caps at 2048; setsreasoning_effort: lowon Qwen3.6 and gpt-oss-120b unless you opt into thinking; keeps the connection alive with SSE pings during slow streams; fails over to wholesale routes if the primary GPU hangs before the first token.
curl https://api.fastinfra.ai/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-N \
-d '{
"model": "gpt-oss-120b",
"stream": true,
"messages": [{"role": "user", "content": "Explain inference routing briefly."}]
}'
Use streaming for any response expected to take more than a few seconds — especially
gpt-oss-120b and qwen3.6-27b. It reduces silent wait time,
keeps proxies from treating the connection as idle, and lets your UI show partial output immediately.
600/min per key,
1800/min per account,
100 concurrent per key, and
1024 server-wide.
See Rate Limits for free-tier values and 429 handling.
Video Generations (LTX-2.5)
Public contract: this section, plus
/openapi.json and
GET https://api.fastinfra.ai/v1/models. See How to determine the API contract.
Text-to-video with synced audio, served on FastInfra GPUs.
Use model lightricks/ltx-2.5 (aliases: ltx-2.5, ltx-2.5-distilled, ltx-2-5-fast).
This is not a chat-completions model — do not send video prompts to
POST https://api.fastinfra.ai/v1/chat/completions. Use the dedicated video endpoints in the
API overview.
Generation takes minutes, so the API is asynchronous: POST returns HTTP 202 Accepted
with a job id in about a second, then poll GET until status is completed.
Waiting on POST until the MP4 is ready will 524 through Cloudflare (~100s).
| Field | Value |
|---|---|
| Model | lightricks/ltx-2.5 |
| Base URL | https://api.fastinfra.ai/v1 |
| Submit | POST https://api.fastinfra.ai/v1/videos/generations → HTTP 202 |
| Poll | GET https://api.fastinfra.ai/v1/videos/jobs/{id} |
| Auth | Authorization: Bearer YOUR_API_KEY |
| Default size | 1280x704 (width × height, divisible by 64) |
| Duration | 5 seconds default, 2–20 seconds |
Pricing (input / output tokens)
Billed like every other FastInfra model: USD per token. Video length is mapped to output tokens so the catalog stays consistent with chat SKUs. Competition (fal.ai LTX-2.5 fast) charges per second of video ($0.09/s at 720p, $0.13/s at 1080p). FastInfra’s default 1280x704 clip is priced at $0.10 per second of output.
| Direction | Price per 1M tokens | Video equivalent |
|---|---|---|
| Input | $0.10 | Prompt tokens (≈ 1 token per 4 characters) |
| Output | $100.00 | 1,000 tokens per second of video → $0.10/s |
Example: a 5-second clip is 5,000 output tokens → $0.50 plus a few cents of prompt tokens.
https://api.fastinfra.ai/v1/videos/generations
https://api.fastinfra.ai/v1/videos/jobs/{id}
Request body (JSON)
| Field | Required | Description |
|---|---|---|
prompt | Yes | Text description of the scene (spoken dialogue in the prompt is supported). |
model | No | Defaults to lightricks/ltx-2.5. Must be a known LTX alias. |
seconds | No | Clip length, 2–20 (default 5). |
size | No | WIDTHxHEIGHT, both divisible by 64 (default 1280x704). |
seed | No | Integer for reproducibility when supported by the worker. |
{
"model": "lightricks/ltx-2.5",
"prompt": "A woman looks at the camera and says, welcome to FastInfra, cinematic lighting.",
"seconds": 5,
"size": "1280x704",
"seed": 42
}
curl
curl https://api.fastinfra.ai/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "lightricks/ltx-2.5",
"prompt": "A woman looks at the camera and says, welcome to FastInfra, cinematic lighting.",
"seconds": 5,
"size": "1280x704"
}'
# HTTP 202 — copy "id", then poll:
curl https://api.fastinfra.ai/v1/videos/jobs/JOB_ID \
-H "Authorization: Bearer YOUR_API_KEY"
Python
import base64, json, time, urllib.request
headers = {
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
}
req = urllib.request.Request(
"https://api.fastinfra.ai/v1/videos/generations",
data=json.dumps({
"model": "lightricks/ltx-2.5",
"prompt": "A golden retriever running through a sunny meadow, cinematic, birdsong.",
"seconds": 5,
"size": "1280x704",
}).encode(),
headers=headers,
method="POST",
)
with urllib.request.urlopen(req, timeout=60) as resp:
job = json.load(resp)
job_id = job["id"]
while True:
time.sleep(2)
poll = urllib.request.Request(
f"https://api.fastinfra.ai/v1/videos/jobs/{job_id}",
headers={"Authorization": "Bearer YOUR_API_KEY"},
)
with urllib.request.urlopen(poll, timeout=60) as resp:
payload = json.load(resp)
if payload["status"] == "completed":
break
if payload["status"] == "failed":
raise SystemExit(payload.get("error") or "video job failed")
open("clip.mp4", "wb").write(base64.b64decode(payload["data"][0]["b64_json"]))
print(payload["id"], payload["usage"])
Submit response (HTTP 202)
Response header Location: /v1/videos/jobs/{id} points at the poll URL.
Save the JSON id field — that is your gateway job id (not the internal worker id).
HTTP/1.1 202 Accepted
Location: /v1/videos/jobs/video_abc123
{
"id": "video_abc123",
"object": "video.job",
"created": 1234567890,
"model": "lightricks/ltx-2.5",
"seconds": 5,
"size": "1280x704",
"status": "queued"
}
Failed poll (HTTP 200, job failed)
{
"id": "video_abc123",
"object": "video.job",
"status": "failed",
"error": "Video generation failed."
}
HTTP errors
- 400 — missing
prompt, unknown model, orsecondsout of range. - 402 — insufficient prepaid credits (same wallet as chat).
- 429 — gateway video queue full; honor
Retry-Afterand retry POST. - 503 — video backend not configured or unreachable.
- 404 on GET — unknown or expired job id (jobs expire after ~24h).
Completed poll (HTTP 200)
{
"id": "video_...",
"object": "video",
"created": 1234567890,
"model": "lightricks/ltx-2.5",
"seconds": 5,
"size": "1280x704",
"status": "completed",
"data": [{ "b64_json": "<mp4-bytes-as-base64>", "format": "mp4" }],
"usage": {
"prompt_tokens": 18,
"completion_tokens": 5000,
"total_tokens": 5018
}
}
queued, in_progress, completed,
or failed. Do not wait on the POST for the MP4; Cloudflare will 524 around
100 seconds. HTTP 429 with Retry-After means the gateway queue is full — wait
and retry the POST.
Tool Calling
Function/tool calling works exactly as with OpenAI: send tools and optionally
tool_choice, receive tool_calls on the assistant message with
finish_reason: "tool_calls", run the tool, and send the result back as a
role: "tool" message carrying the tool_call_id. Streaming delivers
delta.tool_calls fragments with an index, which every OpenAI SDK
accumulator already stitches together. Supported on the self-hosted flagship models
(qwen3.6-27b, gpt-oss-120b) and on wholesale-routed models that
support it upstream; see the compatibility matrix.
from openai import OpenAI
import json
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]},
},
}]
messages = [{"role": "user", "content": "What's the weather in London?"}]
first = client.chat.completions.create(model="qwen3.6-27b", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]
messages.append(first.choices[0].message)
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps({"temp_c": 18, "sky": "overcast"}),
})
final = client.chat.completions.create(model="qwen3.6-27b", messages=messages, tools=tools)
print(final.choices[0].message.content)
llama3.2:3b, phi3:mini, ...) do not
return structured tool calls through this gateway; use a flagship model for agents.
Embeddings
POST https://api.fastinfra.ai/v1/embeddings is OpenAI-compatible and served from self-hosted embedding
models at $0. The OpenAI model ids are accepted as aliases, so a client that only changes
base_url keeps working. encoding_format: "base64" (the OpenAI Python
SDK default) and dimensions (Matryoshka truncation with re-normalisation) are supported.
Input is a string or an array of up to 2,048 strings; token-id arrays are not accepted.
| Model | Dimensions | Aliases |
|---|---|---|
nomic-embed-text |
768 | text-embedding-3-small, text-embedding-ada-002, nomic-ai/nomic-embed-text-v1.5 |
mxbai-embed-large |
1024 | text-embedding-3-large, mixedbread-ai/mxbai-embed-large-v1 |
all-minilm |
384 | sentence-transformers/all-MiniLM-L6-v2 |
curl https://api.fastinfra.ai/v1/embeddings \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "nomic-embed-text", "input": ["FastInfra runs on its own H200s.", "Embeddings are free."]}'
{
"object": "list",
"data": [
{"object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, ...]},
{"object": "embedding", "index": 1, "embedding": [0.0789, 0.0012, ...]}
],
"model": "nomic-embed-text",
"usage": {"prompt_tokens": 17, "total_tokens": 17}
}
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
# Unchanged from an OpenAI integration: the id resolves to nomic-embed-text.
vectors = client.embeddings.create(model="text-embedding-3-small", input=["hello", "world"])
print(len(vectors.data[0].embedding)) # 768
Rate limited like other inference calls (requests and concurrency; no token budget). Errors:
400 invalid_input / model_not_found for a chat model id,
404 model_not_found when the backend lacks the model,
503 backend_unavailable while the embedding backend is down.
Compatibility Matrix
What each OpenAI request feature does on each class of model, as the gateway actually forwards it.
"Upstream" means the field is passed through unchanged and support is whatever that wholesale provider
offers. A feature the gateway drops is marked No even where the model could do it natively. Run
scripts/compat-check.py against your own key to verify any cell live.
| Model family | Stream | Tools | json_object | json_schema | Vision | logprobs | seed | stop | Reasoning |
|---|---|---|---|---|---|---|---|---|---|
Qwen3.6-27B (self-hosted vLLM, H200)qwen3.6-27b, qwen/qwen3.6-27bOne image per request, video disabled. Thinking is off by default; set reasoning_effort (medium/high) or enable_thinking: true to turn it on. |
Yes | Yes | Yes | Yes | Partial | Yes | Yes | Yes | Yes |
gpt-oss-120b (self-hosted vLLM, H100)gpt-oss-120bText only. reasoning_effort low by default; raise it for harder tasks. Reasoning text arrives in the reasoning field. |
Yes | Yes | Yes | Yes | No | Yes | Yes | Yes | Yes |
Free-tier small models (self-hosted Ollama)llama3.1:8b, llama3.2:3b, phi3:mini, gemma2:2b, qwen2.5:*$0 models for wiring and smoke tests. temperature, top_p, seed, stop and response_format are forwarded; tools, logprobs and images are not. Use a flagship model for agents. |
Yes | No | Yes | Yes | No | No | Yes | Yes | No |
Embedding models (self-hosted Ollama)nomic-embed-text, mxbai-embed-large, all-minilmPOST /v1/embeddings only. encoding_format float|base64 and dimensions are supported; OpenAI text-embedding-* ids are aliases. |
No | No | No | No | No | No | No | No | No |
Wholesale-routed models (400+ ids in GET /v1/models)anthropic/claude-*, meta-llama/*, mistralai/*, ...The full OpenAI request body is forwarded unchanged and the reply is passed back with tool_calls, logprobs and extra fields intact. Support is whatever the upstream provider offers for that model. |
Yes | Upstream | Upstream | Upstream | Upstream | Upstream | Upstream | Upstream | Upstream |
Not supported anywhere: n > 1 on streaming, token-id arrays as embeddings input, legacy
POST /v1/completions. Audio endpoints ignore chat fields such as max_tokens.
Framework recipes
One base_url and one key. Each snippet below is the complete change from an OpenAI integration.
OpenAI SDK (Python)
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1", max_retries=2)
# The SDK reads x-ratelimit-* headers and honours Retry-After on 429 automatically.
OpenAI SDK (JavaScript / TypeScript)
import OpenAI from "openai";
const client = new OpenAI({ apiKey: process.env.FASTINFRA_API_KEY, baseURL: "https://api.fastinfra.ai/v1" });
LangChain
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
llm = ChatOpenAI(model="qwen3.6-27b", api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
emb = OpenAIEmbeddings(model="nomic-embed-text", api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1", check_embedding_ctx_length=False)
# Tool calling: llm.bind_tools([...]) works as with OpenAI on qwen3.6-27b and gpt-oss-120b.
LiteLLM
import litellm
response = litellm.completion(
model="openai/qwen3.6-27b", # "openai/" prefix = OpenAI-compatible provider
api_key="YOUR_API_KEY",
api_base="https://api.fastinfra.ai/v1",
messages=[{"role": "user", "content": "hi"}],
)
Vercel AI SDK
import { createOpenAI } from "@ai-sdk/openai";
const fastinfra = createOpenAI({ apiKey: process.env.FASTINFRA_API_KEY, baseURL: "https://api.fastinfra.ai/v1" });
const { text } = await generateText({ model: fastinfra("qwen3.6-27b"), prompt: "hi" });
Continue / Cursor / any "OpenAI-compatible" provider slot
{
"provider": "openai",
"model": "qwen3.6-27b",
"apiBase": "https://api.fastinfra.ai/v1",
"apiKey": "YOUR_API_KEY"
}
n8n
Credentials → OpenAI → set Base URL to https://api.fastinfra.ai/v1. The OpenAI Chat Model and Embeddings nodes then list our models.
Open WebUI
Settings → Connections → OpenAI API: URL https://api.fastinfra.ai/v1, key YOUR_API_KEY. Models populate from GET /v1/models.
Audio: Speech-to-Text & TTS
OpenAI-compatible audio endpoints for transcription, translation, live streaming, and text-to-speech — all served on FastInfra's own GPU pods (Whisper STT and Kokoro-82M neural TTS on the same worker). Audio is processed in memory and never persisted (see data protection note below).
POST https://api.fastinfra.ai/v1/audio/speech for text-to-speech without ever uploading
audio or calling transcription. Each endpoint has its own model id and billing.
Full Kokoro reference: Kokoro TTS.
https://api.fastinfra.ai/v1/audio/transcriptions
https://api.fastinfra.ai/v1/audio/translations
https://api.fastinfra.ai/v1/audio/speech
wss://…/v1/realtime/transcription
Models & pricing
| Capability | Model id | Accepted aliases | Price |
|---|---|---|---|
| Transcription / translation / live streaming | openai/whisper-large-v3-turbo |
whisper-1, whisper-turbo, whisper-large-v3-turbo, whisper-large-v3, openai/whisper-large-v3 |
$0.006 / audio minute (OpenAI whisper-1 parity) |
| Text-to-speech | hexgrad/kokoro-82m |
kokoro, kokoro-82m, tts-1 |
$15 / 1M input characters (OpenAI tts-1 parity) |
The model field is optional on all audio endpoints — omit it and the default
model above is used. Live streaming sessions bill wall-clock duration at the same
per-minute rate.
Transcription (file upload)
Multipart upload, up to 25 MB. Common formats (wav, mp3, m4a, webm, ogg) are decoded automatically.
curl https://api.fastinfra.ai/v1/audio/transcriptions \
-H "Authorization: Bearer YOUR_API_KEY" \
-F file=@answer.wav \
-F model=whisper-1 \
-F response_format=json
{ "text": "Newton's second law states that force equals mass times acceleration." }
| Form field | Notes |
|---|---|
file |
Required. Audio file, max 25 MB. |
model |
Optional. Any alias from the table above. |
language |
Optional ISO code (e.g. en, hi); auto-detected when omitted. |
response_format |
json (default), text, or verbose_json (includes segments, language, duration). |
temperature |
Optional. Use 0 (the default) for maximum accuracy — higher values add sampling randomness and are not recommended for exams. |
prompt |
Optional context hint (names, technical terms) to guide spelling. |
max_tokens has no effect on audio
endpoints — the full audio is always transcribed and billed by duration, not tokens.
Translation (any language → English)
Same form fields as transcription (except language); the response text is always English.
curl https://api.fastinfra.ai/v1/audio/translations \
-H "Authorization: Bearer YOUR_API_KEY" \
-F file=@hindi-answer.wav \
-F response_format=json
# -> { "text": "My name is Rahul and I am a physics student." }
Live streaming (interactive viva)
Stream microphone audio over WebSocket and receive partial and final transcripts in real time. Send binary frames of PCM16 mono 16 kHz audio; receive JSON events. Browsers cannot set headers on WebSocket upgrades, so the API key may be passed as a query parameter on this endpoint only.
wss://api.fastinfra.ai/v1/realtime/transcription?api_key=YOUR_API_KEY
<- binary frames: PCM16 mono 16 kHz audio chunks
-> { "type": "partial", "text": "newton's second" }
-> { "type": "final", "text": "Newton's second law states that..." }
A session holds one rate-limit concurrency slot for its whole duration, and billing is per wall-clock minute of the session.
Kokoro text-to-speech
Self-hosted Kokoro-82M neural voices on FastInfra GPUs. OpenAI-compatible
POST https://api.fastinfra.ai/v1/audio/speech — JSON in, WAV (or raw PCM) out. No prior
transcription step required.
https://api.fastinfra.ai/v1/audio/speech
| Field | Required | Description |
|---|---|---|
input | Yes | Text to speak, max 4,096 characters. |
model | No | Defaults to hexgrad/kokoro-82m. Aliases: tts-1, kokoro, kokoro-82m. |
voice | No | Kokoro voice id (default af_heart). First letter selects language pipeline: a American, b British, h Hindi. |
speed | No | Playback speed, 0.5–2.0 (default 1.0). |
response_format | No | wav (default, 24 kHz mono PCM16) or pcm (raw PCM16 without WAV header). |
Quick start (TTS only)
curl https://api.fastinfra.ai/v1/audio/speech \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"input": "Can you explain Newton'\''s second law in your own words?",
"voice": "af_heart"
}' \
-o question.wav
Response: HTTP 200, Content-Type: audio/wav, body is the audio file.
Billed per input character ($15 / 1M characters, OpenAI tts-1 parity).
Available voices
| Voice id | Description |
|---|---|
af_heart | US female, warm (default) |
af_bella | US female |
am_michael | US male |
bf_emma | UK female |
bm_george | UK male |
hf_alpha | Hindi female |
hm_omega | Hindi male |
Python (TTS only)
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
speech = client.audio.speech.create(
model="tts-1",
voice="am_michael",
input="Welcome to FastInfra. This clip uses Kokoro on our GPU pods.",
speed=1.0,
)
speech.write_to_file("welcome.wav")
JavaScript (fetch)
const res = await fetch("https://api.fastinfra.ai/v1/audio/speech", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "tts-1",
voice: "bf_emma",
input: "Hello from Kokoro.",
}),
});
const wav = Buffer.from(await res.arrayBuffer());
// wav is 24 kHz mono PCM16 in a WAV container
Kokoro errors
- 400 — missing
input, unknown model, text over 4,096 chars, or invalidvoiceid. - 402 — insufficient prepaid credits.
- 503 — Kokoro worker still loading or temporarily unavailable; retry shortly.
Python (full voice loop: STT + TTS)
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
# Student's answer -> text
with open("answer.wav", "rb") as f:
transcript = client.audio.transcriptions.create(model="whisper-1", file=f)
print(transcript.text)
# Examiner's question -> natural voice
speech = client.audio.speech.create(
model="tts-1", voice="af_heart",
input="Good answer. What happens if the mass doubles?")
speech.write_to_file("question.wav")
Validating a model id
Gateway tools that verify a configured model can call the standard retrieve-model
endpoint. Unknown ids return a JSON invalid_request_error with HTTP 404.
curl https://api.fastinfra.ai/v1/models/whisper-large-v3 \
-H "Authorization: Bearer YOUR_API_KEY"
# -> { "id": "openai/whisper-large-v3-turbo", "object": "model", ... }
Common errors
- 401 — invalid or missing API key. Create a real key on the API Keys page; placeholder values are rejected.
- 400 — unknown model id, missing
file/input, undecodable audio, or unknown TTS voice. The error message states the exact problem. - 402 — insufficient prepaid credits (see Credits).
- 413 — audio file over 25 MB. Split long recordings or use live streaming.
- 503 — transcription backends temporarily unavailable; retry shortly.
GPU Pods
Rent dedicated GPU machines with any Docker image — training, fine-tuning, ComfyUI, vLLM, Jupyter, or your own stack. Pods are managed from the GPU Cloud dashboard (not the OpenAI API). Billing is per minute from your prepaid credit balance; stop the pod and compute billing stops immediately.
| Resource | URL |
|---|---|
| GPU catalog (public) | https://gpu.fastinfra.ai |
| Deploy a pod | https://gpu.fastinfra.ai/Deploy (sign in required) |
| My pods | https://gpu.fastinfra.ai/Pods |
| Add credits | Billing |
| Support | Support — choose GPU Pod issue for pod-specific help |
Getting started
- Sign in and add credits on the Billing page.
- Browse the GPU catalog for hourly rates (RTX 4090, A100, H100, H200, and more).
- Open Deploy, pick a GPU tier, Docker image, ports, and disk size.
- When status is
RUNNING, copy SSH or HTTP endpoints from the pod detail page and connect. - Stop when finished — billing for GPU compute ends the moment the pod stops.
Deploy options
| Option | Details |
|---|---|
| GPU tier | Secure — vetted datacenters. Value — lower price community hosts (when available for that GPU). |
| GPU count | 1–8 GPUs per pod (max depends on GPU type and tier). |
| Docker image | Any public image, e.g. runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404, ComfyUI, Ollama, or vLLM. |
| Ports | Comma-separated, e.g. 22/tcp, 8888/http. TCP ports get a public host:port; HTTP ports are exposed for web UIs. |
| Environment | One KEY=value per line (up to 50 vars). |
| Container disk | 5–1000 GB ephemeral disk — wiped when the pod stops. |
| Volume disk | 0 (none) or 10–4000 GB persistent volume mounted at /workspace — survives stops. |
| SSH / Jupyter | Enable SSH for shell access; optional JupyterLab (exposes port 8888). |
Pricing & billing
GPU pod prices include a 8% platform markup on underlying compute and storage. Catalog hourly rates are rounded up so displayed prices are always what you pay.
- Compute — billed per minute while the pod is
RUNNING(GPU hourly rate × GPU count, plus running disk). - Stopped volume — if you attached a persistent volume and stop (but do not terminate) the pod, a lower storage-only rate applies until you terminate or start again.
- Deploy gate — you need enough credits for at least 1 hour(s) of runtime at the pod's hourly rate before deploy or start.
- Auto-stop — if your balance hits zero while a pod is running, it is stopped automatically to prevent overdraft. Add credits and start it again.
- Terminate — permanently deletes the pod and its volume; all billing ends.
Per-minute charges appear on the pod detail page under Billing history. Inference API usage (chat completions) is billed separately — see Credits.
Pod lifecycle
| Action | Effect |
|---|---|
| Deploy | Provisions a new pod; status moves through PROVISIONING → STARTING → RUNNING (typically 1–2 minutes). |
| Stop | Halts compute billing. Container disk is discarded; volume data is kept if configured. |
| Start | Resumes a stopped pod (requires sufficient credit balance). |
| Restart | Reboots a running pod in place. |
| Terminate | Deletes the pod and volume permanently. Irreversible. |
Connecting
When a pod is RUNNING, the detail page lists connection endpoints:
- SSH —
ssh root@HOST -p PORT(when SSH is enabled and port 22 is exposed). - HTTP services — Jupyter, Gradio, or custom apps on exposed HTTP ports appear as
HOST:PORT. - Endpoints refresh — the detail page polls status every 15 seconds while provisioning or running; reload if endpoints are still assigning.
# Example after deploy (values shown on your pod detail page)
ssh root@203.0.113.42 -p 22001
# Jupyter in browser (if enabled)
http://203.0.113.42:8888
Popular Docker images
runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404— PyTorch + CUDA (default on deploy form)runpod/tensorflow:2.2.0-py3— TensorFlow- Community templates for ComfyUI, Ollama, vLLM — use any public registry image your workload needs
https://api.fastinfra.ai/v1 — a stopped GPU pod does not affect API keys or model routing.
For automation, use the inference API; for interactive GPU machines, use GPU pods.
CPU Servers
Rent Linux virtual machines from the CPU Cloud dashboard. Billing is per minute from your prepaid credit balance while the server exists — stopping a server does not stop charges; delete it when finished.
| Resource | URL |
|---|---|
| CPU catalog (public) | https://cpu.fastinfra.ai |
| Deploy a server | https://cpu.fastinfra.ai/Cpu/Deploy (sign in) |
| My servers | https://cpu.fastinfra.ai/Cpu/Servers |
| Add credits | Billing |
Getting started
- Sign in and add credits on Billing.
- Browse plans on the CPU catalog (vCPU, RAM, region, hourly rate).
- Deploy with a name, plan, region, and OS image (Ubuntu or Debian).
- When status is
RUNNING, copy the public IP and root password from the server detail page (password shown once). - Delete when finished — that is the only way to stop billing.
Prices include a 8% platform markup. Deploy requires ~1 hour(s) of credits at the server's hourly rate. CPU servers are dashboard-only today (no OpenAI-style provisioning API).
Model Routing
FastInfra serves hundreds of models through one OpenAI-compatible API. You send the same request shape for every model; the gateway automatically routes each call to the cheapest healthy capacity for that model, with automatic failover when a route degrades.
Discover models
GET https://api.fastinfra.ai/v1/models— full catalog of available model IDsGET https://api.fastinfra.ai/v1/models/count— total model count- Pricing page — searchable catalog with per-token pricing
GET https://api.fastinfra.ai/v1/models returns clean, stable model IDs
(e.g. deepseek/deepseek-v4-flash-0731). Vendor-specific path formats are normalized
automatically, so the ID you see in the catalog is the ID you send.
Routing behavior
For every request, the gateway resolves the route in this order:
- Free-tier models — served at $0 on self-hosted Ollama capacity (
llama3.1:8b,mistral:7b, etc.) - Dedicated GPU primary — e.g.
gpt-oss-120bon RunPod vLLM,qwen3.6-27b(multimodal — see Vision) on H200 - Fallback chain — wholesale providers tried automatically if the primary fails or is saturated
Routing is automatic — your integration stays identical as capacity changes. Enable wholesale API keys in the admin panel so fallbacks activate when self-hosted GPUs are busy or offline.
Wholesale fallbacks (cheapest first)
When a self-hosted route fails or hits its concurrency cap, these wholesale providers are tried in order (requires API keys in admin):
| Model | Fallback order | Typical wholesale input $/1M tokens |
|---|---|---|
gpt-oss-120b |
route-21 → route-07 |
~$0.15 → ~$0.04 |
qwen3.6-27b |
route-21 → route-18 → route-07 |
~$0.32 → ~$0.30 → ~$0.60 |
llama3.1:8b (free tier) |
route-07 → route-12 → route-03 |
varies |
You are billed per token only when traffic actually routes to a wholesale provider. Self-hosted capacity is used first whenever healthy.
Python example
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
response = client.chat.completions.create(
model="agentica-org/DeepCoder-14B-Preview",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
Errors
- 502 —
No inference provider available for the requested model.The model is not in the catalog or is not currently enabled. - 503 —
Inference provider is not reachable.Upstream capacity is temporarily offline; retry shortly.
If a model does not appear in GET https://api.fastinfra.ai/v1/models, it may not be enabled on this platform yet.
Browse the pricing catalog for models available on this deployment.
List Models
Returns the full model catalog available on this deployment. Chat models use
POST https://api.fastinfra.ai/v1/chat/completions. The LTX video model
(lightricks/ltx-2.5) appears here for discovery but must be called via
POST https://api.fastinfra.ai/v1/videos/generations.
https://api.fastinfra.ai/v1/models
Response shape
{
"object": "list",
"data": [
{ "id": "anthropic/claude-3.5-sonnet", "object": "model" },
{ "id": "deepseek/deepseek-v4-flash-0731", "object": "model" },
{ "id": "llama3.1:8b", "object": "model" }
],
"total_count": 3
}
Use GET https://api.fastinfra.ai/v1/models/count for a lightweight count without downloading the full list.
Currently available (784)
agentica-org/DeepCoder-14B-Previewaion-labs/aion-2.0aion-labs/aion-3.0aion-labs/aion-3.0-miniaion-labs/aion-3.5aion-labs/aion-3.5-miniaion-labs/aion-rp-llama-3.1-8balibaba/happyhorse-1.0-i2valibaba/happyhorse-1.0-r2valibaba/happyhorse-1.0-t2valibaba/happyhorse-1.1-i2valibaba/happyhorse-1.1-r2valibaba/happyhorse-1.1-t2vallenai/Molmo-7B-D-0924amazon/nova-2-lite-v1amazon/nova-lite-v1amazon/nova-micro-v1amazon/nova-premier-v1amazon/nova-pro-v1anthracite-org/magnum-v4-72banthropic/claude-fable-5anthropic/claude-fable-5.1anthropic/claude-fable-5.1:batchanthropic/claude-fable-5:batchanthropic/claude-haiku-4-5anthropic/claude-haiku-4.5:batchanthropic/claude-opus-4-7anthropic/claude-opus-4-8anthropic/claude-opus-4.1anthropic/claude-opus-4.1:batchanthropic/claude-opus-4.5anthropic/claude-opus-4.5:batchanthropic/claude-opus-4.6anthropic/claude-opus-4.6:batchanthropic/claude-opus-4.7:batchanthropic/claude-opus-4.8:batchanthropic/claude-opus-5anthropic/claude-opus-5-5anthropic/claude-opus-5.5anthropic/claude-opus-5.5:batchanthropic/claude-opus-5:batchanthropic/claude-sonnet-4anthropic/claude-sonnet-4-6anthropic/claude-sonnet-4.5anthropic/claude-sonnet-4.5:batchanthropic/claude-sonnet-4.6:batchanthropic/claude-sonnet-5anthropic/claude-sonnet-5-5anthropic/claude-sonnet-5.5anthropic/claude-sonnet-5.5:batch
Showing 50 of 784 models. Call GET https://api.fastinfra.ai/v1/models for the full list.
Code Examples
Chat examples use the OpenAI SDK with base_url="https://api.fastinfra.ai/v1".
Video generation uses raw HTTP (async job + poll) — see
Video Generations for the full Python/curl flow.
Kokoro TTS is standalone — see Kokoro TTS.
Kokoro TTS (Python, minimal)
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
client.audio.speech.create(
model="tts-1",
voice="af_heart",
input="Hello from Kokoro on FastInfra.",
).write_to_file("hello.wav")
LTX video (Python, minimal)
import json, time, urllib.request
BASE = "https://api.fastinfra.ai/v1"
KEY = "YOUR_API_KEY"
headers = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
req = urllib.request.Request(
f"{BASE}/videos/generations",
data=json.dumps({
"model": "lightricks/ltx-2.5",
"prompt": "Ocean waves at golden hour, cinematic.",
"seconds": 5,
}).encode(),
headers=headers,
method="POST",
)
with urllib.request.urlopen(req, timeout=60) as resp:
job_id = json.load(resp)["id"]
while True:
time.sleep(2)
poll = urllib.request.Request(f"{BASE}/videos/jobs/{job_id}", headers={"Authorization": f"Bearer {KEY}"})
with urllib.request.urlopen(poll, timeout=60) as resp:
body = json.load(resp)
if body["status"] == "completed":
print(body["usage"])
break
if body["status"] == "failed":
raise SystemExit(body.get("error"))
C# (.NET)
var client = new OpenAIClient(
new ApiKeyCredential("YOUR_API_KEY"),
new OpenAIClientOptions { Endpoint = new Uri("https://api.fastinfra.ai/v1") });
var chat = client.GetChatClient("llama3.1:8b");
var response = await chat.CompleteChatAsync("Hello!");
Console.WriteLine(response.Value.Content[0].Text);
Python (free-tier / local model)
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.fastinfra.ai/v1"
)
response = client.chat.completions.create(
model="llama3.1:8b",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
Python (provider model)
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
response = client.chat.completions.create(
model="agentica-org/DeepCoder-14B-Preview",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
JavaScript
const response = await fetch("https://api.fastinfra.ai/v1/chat/completions", {
method: "POST",
headers: {
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "agentica-org/DeepCoder-14B-Preview",
messages: [{ role: "user", content: "Hello!" }]
})
});
const data = await response.json();
console.log(data.choices[0].message.content);
Ready to ship?
Create your free account, generate an API key, and start calling frontier models in minutes.