🚀 One API for every frontier model on FastInfra

Browse models

Transparent per-token pricing → compare providers on FastInfra

View pricing

OpenAI-compatible chat completions → start in minutes

Read the docs

⚡ Server Auction → Enterprise bare metal from $78.15/mo (173 in stock)

Browse deals

Llama Nemotron Embed Vl 1b v2 vs Nemotron 3.5 ASR Streaming Multilingual 0.6b

Live API pricing and availability, side by side. Both models run on FastInfra's OpenAI-compatible endpoint — switching between them is a one-line change.

Pricing

Price per 1M tokens

Prices are live and sync automatically from wholesale providers.

Llama Nemotron Embed Vl 1b v2 Nemotron 3.5 ASR Streaming Multilingual 0.6b
Input $0.01 —
Output — —
Default route route-21 route-21
Routes available 1 1
Quickstart

Try both in 30 seconds

Same endpoint, same SDK — only the model string changes.

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

for model in ["nvidia/llama-nemotron-embed-vl-1b-v2", "nvidia/Nemotron-3.5-ASR-Streaming-Multilingual-0.6b"]:
    response = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": "Hello!"}]
    )
    print(model, "->", response.choices[0].message.content)