🚀 One API for every frontier model on FastInfra

Browse models

Transparent per-token pricing → compare providers on FastInfra

View pricing

OpenAI-compatible chat completions → start in minutes

Read the docs

⚡ Server Auction → Enterprise bare metal from $78.15/mo (173 in stock)

Browse deals

Gemma 4 31b It:batch API

Run google/gemma-4-31b-it:batch through FastInfra's API. Pay per token. Video is billed as 1,000 output tokens per second of generated video.

Pricing

Gemma 4 31b It:batch pricing

Billed per token. 1,000 output tokens = 1 second of video ($0.00/s at this list price).

Direction Price per 1M tokens
Input$0.41
Output$1.02
Per second of video$0.00 (1,000 output tokens)
Availability

1 route available

Requests route to route-07 by default (lowest cost). FastInfra fails over automatically if a route is unavailable.

Route label Model ID
route-07 google/gemma-4-31b-it:batch
Quickstart

Call Gemma 4 31b It:batch in 30 seconds

Same FastInfra API key and base URL. POST /videos/generations (HTTP 202), then poll GET /videos/jobs/{id} — not chat completions.

Python

import base64, json, time, urllib.request

headers = {"Authorization": "Bearer YOUR_API_KEY", "Content-Type": "application/json"}
req = urllib.request.Request(
    "https://api.fastinfra.ai/v1/videos/generations",
    data=json.dumps({
        "model": "google/gemma-4-31b-it:batch",
        "prompt": "A woman looks at the camera and says, welcome to FastInfra.",
        "seconds": 5,
        "size": "1280x704"
    }).encode(),
    headers=headers,
    method="POST",
)
with urllib.request.urlopen(req, timeout=60) as resp:
    job = json.load(resp)

while True:
    time.sleep(2)
    poll = urllib.request.Request(
        f"https://api.fastinfra.ai/v1/videos/jobs/{job['id']}",
        headers={"Authorization": "Bearer YOUR_API_KEY"},
    )
    with urllib.request.urlopen(poll, timeout=60) as resp:
        payload = json.load(resp)
    if payload["status"] == "completed":
        break
    if payload["status"] == "failed":
        raise SystemExit(payload.get("error") or "video job failed")

open("clip.mp4", "wb").write(base64.b64decode(payload["data"][0]["b64_json"]))

curl

curl https://api.fastinfra.ai/v1/videos/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemma-4-31b-it:batch",
    "prompt": "A woman looks at the camera and says, welcome to FastInfra.",
    "seconds": 5,
    "size": "1280x704"
  }'
# HTTP 202 — copy "id", then poll:
curl https://api.fastinfra.ai/v1/videos/jobs/JOB_ID \
  -H "Authorization: Bearer YOUR_API_KEY"
FAQ

Gemma 4 31b It:batch — common questions

How much does the google/gemma-4-31b-it:batch API cost?

On FastInfra, google/gemma-4-31b-it:batch costs $0.4095/1M input, $1.0185/1M output tokens ($0 per second of video). Billing is per token used, with no subscription.

Is google/gemma-4-31b-it:batch compatible with the OpenAI SDK?

Video models use the same FastInfra API key and base URL (https://api.fastinfra.ai/v1). POST https://api.fastinfra.ai/v1/videos/generations returns HTTP 202 with a job id; poll GET https://api.fastinfra.ai/v1/videos/jobs/{id} until status is completed.

How does routing work for google/gemma-4-31b-it:batch?

google/gemma-4-31b-it:batch is available on 1 route(s): route-07. Requests use route-07 by default (lowest cost). FastInfra fails over automatically if a route is unavailable.