Inference API
Pay per token across all platform models. OpenAI-compatible endpoints, multi-region routing.
Language
| Model | Input | Output | Notes |
|---|---|---|---|
DeepSeek-V4-Flash | $0.16/1M | $0.33/1M | — |
DeepSeek-V4-Pro | $1.46/1M | $2.93/1M | — |
GLM-4.7 | $0.33/1M | $1.33/1M | — |
GLM-5 | $0.67/1M | $3.00/1M | — |
GLM-5-Turbo | $0.83/1M | $3.67/1M | — |
GLM-5.1 | $1.00/1M | $4.00/1M | — |
Llama-3.1-8B-Instruct | $0.10/1M | $0.10/1M | Context: 128K |
Qwen2.5-7B-Instruct | $0.20/1M | $0.20/1M | — |
Qwen3-235B-A22B | $0.33/1M | $1.33/1M | — |
qwen3.8-max | — | — | — |
Vision
| Model | Input | Output | Notes |
|---|---|---|---|
Gemma-4-31B-IT | $0.50/1M | $0.50/1M | — |
Kimi-K2.6 | $1.08/1M | $4.50/1M | — |
Qwen2.5-VL-7B-Instruct | $0.30/1M | $0.60/1M | Context: 32K |
qwen3-omni-30b-a3b-instruct | $0.40/1M | $0.80/1M | — |
Qwen3.5-35B-A3B | $0.40/1M | $0.40/1M | — |
SenseNova-U1 8B | $0.30/1M | $0.60/1M | — |
Embedding
| Model | Input | Output | Notes |
|---|---|---|---|
Jina-Embeddings-V3 | $0.02/1M | — | — |
Jina-Embeddings-V4 | $0.12/1M | — | — |
qwen3-embedding-4b | $0.12/1M | — | — |
qwen3-embedding-8b | $0.12/1M | — | — |
Reranker
| Model | Input | Output | Notes |
|---|---|---|---|
BGE-Reranker-V2-M3 | $0.03/1M | — | — |
Image
| Model | Input | Output | Notes |
|---|---|---|---|
qwen-image | — | $30.00/1M out tok | Tokens / item: 1,000 |
Z-Image-Turbo | — | $1.00/1M out tok | Tokens / item: 2,000 |
Speech
| Model | Input | Output | Notes |
|---|---|---|---|
Fun-ASR-Nano | $0.05/1M in tok | — | Tokens / sec audio: 1,000 |
Kokoro-82M | — | $0.50/1M out tok | Tokens / sec audio: 100 |
Qwen3-ASR-1.7B qwen3-asr-1-7b | $0.05/1M in tok | — | Tokens / sec audio: 1,000 |
Qwen3-TTS | — | $10.00/1M out tok | — |
Whisper-Large-V3 | $4.00/1M in tok | — | Tokens / sec audio: 25 |
Video Generation
Billed on actual tokens used, at a rate set by output resolution and whether you supply a reference video. Higher resolutions bill at a lower per-token rate but consume substantially more tokens per second of video, so a 4K clip still costs more than the same clip at 720p.
Seedance 2.0
| Resolution | Text / image input | With reference video |
|---|---|---|
| 480p / 720p | $7.70/1M | $4.73/1M |
| 1080p | $8.47/1M | $5.17/1M |
| 4K | $4.00/1M | $2.40/1M |
Wan 3.0
| Resolution | Text / image input | With reference video |
|---|---|---|
| 1080p | $0.00/1M | $0.00/1M |
| 480p | $0.00/1M | $0.00/1M |
| 720p | $0.00/1M | $0.00/1M |
Wan2.1-T2V-1.3B
| Resolution | Text / image input | With reference video |
|---|---|---|
| 480p | $0.00/1M | $0.00/1M |
| 720p | $0.00/1M | $0.00/1M |
wan21-t2v-14b
| Resolution | Text / image input | With reference video |
|---|---|---|
| 480p | $0.00/1M | $0.00/1M |
| 720p | $0.00/1M | $0.00/1M |
wan3.0-video
| Resolution | Text / image input | With reference video |
|---|---|---|
| 1080p | $0.00/1M | $0.00/1M |
| 480p | $0.00/1M | $0.00/1M |
| 720p | $0.00/1M | $0.00/1M |
GPU Compute
Pay per GPU-hour. Same rate for dedicated inference, workspace instances, and clusters. Billed per second of running time.
GPU rates are momentarily unavailable — see the console for live rates.
Storage
Persistent storage attached to your GPU instances and clusters. Cloud Drives are single-instance (ReadWriteOnce); Shared Filesystems mount across multiple instances (ReadWriteMany). Billed per second of attached time.
| Type | Size | Monthly |
|---|---|---|
| Cloud Drive | 50 GB | $5.00/mo |
| Cloud Drive | 100 GB | $10.00/mo |
| Cloud Drive | 200 GB | $20.00/mo |
| Cloud Drive | 300 GB | $30.00/mo |
| Cloud Drive | 400 GB | $40.00/mo |
| Cloud Drive | 500 GB | $50.00/mo |
| Shared Filesystem | 50 GB | $5.00/mo |
| Shared Filesystem | 100 GB | $10.00/mo |
| Shared Filesystem | 200 GB | $20.00/mo |
| Shared Filesystem | 300 GB | $30.00/mo |
| Shared Filesystem | 400 GB | $40.00/mo |
| Shared Filesystem | 500 GB | $50.00/mo |