AQA
aqa
AQA for text generation, reasoning, tool calling, and live streaming responses.
Chat
Context window: 7,168 tokens
Max output: 1,024 tokens
Computer Use Preview
computer-use-preview
Computer Use Preview for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Streaming supported
Tool/function calling supported
Input per 1m tokens
$1.250000
Output per 1m tokens
$10.000000
Minimum hold
$0.010000
DeepSeek OCR
DeepSeek-OCR
DeepSeek OCR through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming Context window: 32,768 tokens
Max output: 8,192 tokens
Input per 1m tokens
$0.560000
Output per 1m tokens
$1.680000
Minimum hold
$0.010000
DeepSeek R1 0528
DeepSeek-R1-0528
DeepSeek R1 0528 through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 163,840 tokens
Max output: 32,768 tokens
Input per 1m tokens
$1.350000
Output per 1m tokens
$5.400000
Minimum hold
$0.010000
DeepSeek V3.1
DeepSeek-V3.1
DeepSeek V3.1 through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 131,072 tokens
Max output: 32,768 tokens
Input per 1m tokens
$1.230000
Output per 1m tokens
$4.940000
Minimum hold
$0.010000
DeepSeek V3.2
DeepSeek-V3.2
DeepSeek V3.2 through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 163,840 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.560000
Cached input per 1m tokens
$0.056000
Output per 1m tokens
$1.680000
Gemini 2.5 Flash
gemini-2.5-flash
Gemini 2.5 Flash for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 1,048,576 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.300000
Cached input per 1m tokens
$0.030000
Output per 1m tokens
$2.500000
Gemini 2.5 Flash-Lite
gemini-2.5-flash-lite
Gemini 2.5 Flash-Lite for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 1,048,576 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.100000
Cached input per 1m tokens
$0.010000
Output per 1m tokens
$0.400000
Gemini 2.5 Pro
gemini-2.5-pro
Gemini 2.5 Pro for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 1,048,576 tokens
Max output: 65,536 tokens
Input per 1m tokens
$1.250000
Cached input per 1m tokens
$0.125000
Output per 1m tokens
$10.000000
Gemini 3 Flash Preview
gemini-3-flash-preview
Gemini 3 Flash Preview for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 1,048,576 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.500000
Cached input per 1m tokens
$0.050000
Output per 1m tokens
$3.000000
Gemini 3.1 Flash-Lite
gemini-3.1-flash-lite
Gemini 3.1 Flash-Lite for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 1,048,576 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.250000
Cached input per 1m tokens
$0.025000
Output per 1m tokens
$1.500000
Gemini 3.1 Pro Preview
gemini-3.1-pro-preview
Gemini 3.1 Pro Preview for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 1,048,576 tokens
Max output: 65,536 tokens
Input per 1m tokens
$2.000000
Cached input per 1m tokens
$0.200000
Output per 1m tokens
$12.000000
Gemini 3.1 Pro Preview Custom Tools
gemini-3.1-pro-preview-customtools
Gemini 3.1 Pro Preview Custom Tools for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 1,048,576 tokens
Max output: 65,536 tokens
Input per 1m tokens
$2.000000
Cached input per 1m tokens
$0.200000
Output per 1m tokens
$12.000000
Gemini 3.5 Flash
gemini-3.5-flash
Gemini 3.5 Flash for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 1,048,576 tokens
Max output: 65,536 tokens
Input per 1m tokens
$1.500000
Cached input per 1m tokens
$0.150000
Output per 1m tokens
$9.000000
Gemini Flash Latest
gemini-flash-latest
Gemini Flash Latest for text generation, reasoning, tool calling, and streaming responses.
Chat
Streaming الأدوات Context window: 1,048,576 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.500000
Cached input per 1m tokens
$0.050000
Output per 1m tokens
$3.000000
Gemini Robotics-ER 1.6 Preview
gemini-robotics-er-1.6-preview
Gemini Robotics-ER 1.6 Preview for text generation, reasoning, tool calling, and streaming responses.
Chat
Streaming الأدوات Streaming supported
Tool/function calling supported
Input per 1m tokens
$1.000000
Output per 1m tokens
$5.000000
Minimum hold
$0.010000
Gemma 4 26B A4B IT
gemma-4-26b-a4b-it
Gemma 4 26B A4B IT through Cloudflare Workers AI with Omixa provider routing.
Chat
Streaming Context window: 128,000 tokens
Max output: 8,192 tokens
Input per 1m tokens
$2.000000
Cached input per 1m tokens
$0.500000
Output per 1m tokens
$8.000000
GLM 4.7
glm-4.7
GLM 4.7 through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 200,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$1.000000
Cached input per 1m tokens
$0.100000
Output per 1m tokens
$3.200000
GLM 4.7 Flash
glm-4.7-flash
GLM 4.7 Flash through Cloudflare Workers AI with Omixa provider routing.
Chat
Streaming الأدوات Context window: 131,072 tokens
Max output: 32,768 tokens
Input per 1m tokens
$2.000000
Cached input per 1m tokens
$0.500000
Output per 1m tokens
$8.000000
GLM 5
glm-5
GLM 5 through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 200,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$1.000000
Cached input per 1m tokens
$0.100000
Output per 1m tokens
$3.200000
GLM 5.2
glm-5.2
GLM 5.2 routed primarily through Azure AI Foundry with Cloudflare Workers AI fallback.
Chat
Streaming الأدوات Context window: 262,144 tokens
Max output: 128,000 tokens
Input per 1m tokens
$1.400000
Cached input per 1m tokens
$0.260000
Output per 1m tokens
$4.400000
GPT Chat Latest
gpt-chat-latest
GPT Chat Latest for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 128,000 tokens
Max output: 16,384 tokens
Input per 1m tokens
$5.000000
Cached input per 1m tokens
$0.500000
Output per 1m tokens
$30.000000
GPT OSS 120B
gpt-oss-120b
GPT OSS 120B through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 131,072 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.150000
Output per 1m tokens
$0.600000
Minimum hold
$0.010000
GPT OSS 20B
gpt-oss-20b
GPT OSS 20B through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 131,072 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.070000
Output per 1m tokens
$0.300000
Minimum hold
$0.010000
GPT-5
gpt-5
GPT-5 routed primarily through Azure OpenAI with Cloudflare fallback.
Chat
Streaming الأدوات Context window: 400,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$1.250000
Cached input per 1m tokens
$0.125000
Output per 1m tokens
$10.000000
GPT-5.1
gpt-5.1
GPT-5.1 for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 400,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$1.250000
Cached input per 1m tokens
$0.125000
Output per 1m tokens
$10.000000
GPT-5.3 Codex
gpt-5.3-codex
GPT-5.3 Codex for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 400,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$1.750000
Cached input per 1m tokens
$0.175000
Output per 1m tokens
$14.000000
GPT-5.4
gpt-5.4
GPT-5.4 for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 922,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$2.500000
Cached input per 1m tokens
$0.250000
Output per 1m tokens
$15.000000
GPT-5.4 Mini
gpt-5.4-mini
GPT-5.4 Mini for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 272,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$0.750000
Cached input per 1m tokens
$0.075000
Output per 1m tokens
$4.500000
GPT-5.4 Pro
gpt-5.4-pro
GPT-5.4 Pro for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 1,050,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$30.000000
Output per 1m tokens
$180.000000
Minimum hold
$0.010000
GPT-5.5
gpt-5.5
GPT-5.5 for language generation, reasoning, tool calling, and streaming chat responses.
Chat
Streaming الأدوات Context window: 922,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$5.000000
Cached input per 1m tokens
$0.500000
Output per 1m tokens
$30.000000
Grok 4.1 Fast (Non-Reasoning)
grok-4.1-fast-non-reasoning
Grok 4.1 Fast (Non-Reasoning) through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 128,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$0.200000
Cached input per 1m tokens
$0.050000
Output per 1m tokens
$0.500000
Grok 4.1 Fast (Reasoning)
grok-4.1-fast-reasoning
Grok 4.1 Fast (Reasoning) through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 128,000 tokens
Max output: 128,000 tokens
Input per 1m tokens
$0.200000
Cached input per 1m tokens
$0.050000
Output per 1m tokens
$0.500000
Grok 4.20 (Non-Reasoning)
grok-4-20-non-reasoning
Grok 4.20 (Non-Reasoning) through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 2,000,000 tokens
Max output: 8,192 tokens
Input per 1m tokens
$1.250000
Cached input per 1m tokens
$0.200000
Output per 1m tokens
$2.500000
Grok 4.20 (Reasoning)
grok-4-20-reasoning
Grok 4.20 (Reasoning) through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 2,000,000 tokens
Max output: 8,192 tokens
Input per 1m tokens
$1.250000
Cached input per 1m tokens
$0.200000
Output per 1m tokens
$2.500000
Kimi K2 Thinking
Kimi-K2-Thinking
Kimi K2 Thinking through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 131,072 tokens
Max output: 65,536 tokens
Input per 1m tokens
$1.045000
Cached input per 1m tokens
$0.176000
Output per 1m tokens
$4.400000
Kimi K2.5
Kimi-K2.5
Kimi K2.5 for text generation, reasoning, tool calling, and live streaming responses.
Chat
Streaming الأدوات Context window: 262,144 tokens
Max output: 262,144 tokens
Input per 1m tokens
$0.660000
Cached input per 1m tokens
$0.110000
Output per 1m tokens
$3.300000
Kimi K2.6
Kimi-K2.6
Kimi K2.6 routed primarily through Azure AI Foundry with Cloudflare Workers AI fallback.
Chat
Streaming الأدوات Context window: 262,144 tokens
Max output: 262,144 tokens
Input per 1m tokens
$1.045000
Cached input per 1m tokens
$0.176000
Output per 1m tokens
$4.400000
Kimi K2.7
Kimi-K2.7
Kimi K2.7 routed primarily through Azure AI Foundry with Cloudflare Workers AI fallback.
Chat
Streaming الأدوات Context window: 262,144 tokens
Max output: 262,144 tokens
Input per 1m tokens
$0.950000
Cached input per 1m tokens
$0.190000
Output per 1m tokens
$4.000000
Llama 3.1 70B Instruct
llama-3.1-70b-instruct
Llama 3.1 70B Instruct through Cloudflare Workers AI with Omixa provider routing.
Chat
Streaming Context window: 128,000 tokens
Max output: 8,192 tokens
Llama 3.3 70B Instruct FP8 Fast
llama-3.3-70b-instruct-fp8-fast
Llama 3.3 70B Instruct FP8 Fast through Cloudflare Workers AI with Omixa provider routing.
Chat
Streaming Context window: 128,000 tokens
Max output: 8,192 tokens
Meta Llama 3 405B Instruct
llama-3-405b-instruct
Meta Llama 3 405B Instruct through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 131,072 tokens
Max output: 8,192 tokens
Input per 1m tokens
$2.700000
Output per 1m tokens
$2.700000
Minimum hold
$0.010000
Meta Llama 3 70B Instruct
llama-3-70b-instruct
Meta Llama 3 70B Instruct through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 8,192 tokens
Max output: 8,192 tokens
Input per 1m tokens
$0.710000
Output per 1m tokens
$0.710000
Minimum hold
$0.010000
Meta Llama 3 8B Instruct
llama-3-8b-instruct
Meta Llama 3 8B Instruct through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 8,192 tokens
Max output: 8,192 tokens
Input per 1m tokens
$0.200000
Output per 1m tokens
$0.200000
Minimum hold
$0.010000
Meta Llama 3.2 90B Instruct
llama-3.2-90b-instruct
Meta Llama 3.2 90B Instruct through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 131,072 tokens
Max output: 8,192 tokens
Input per 1m tokens
$0.900000
Output per 1m tokens
$0.900000
Minimum hold
$0.010000
Meta Llama 3.3 70B Instruct
Llama-3.3-70B-Instruct
Meta Llama 3.3 70B Instruct through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 131,072 tokens
Max output: 8,192 tokens
Input per 1m tokens
$0.710000
Output per 1m tokens
$0.710000
Minimum hold
$0.010000
Meta Llama 4 Maverick Instruct
Llama-4-Maverick-17B-128E-Instruct-FP8
Meta Llama 4 Maverick Instruct through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 524,288 tokens
Max output: 8,192 tokens
Input per 1m tokens
$0.350000
Output per 1m tokens
$1.150000
Minimum hold
$0.010000
Meta Llama 4 Scout Instruct
llama-4-scout-17b-16e-instruct
Meta Llama 4 Scout Instruct through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 1,048,576 tokens
Max output: 8,192 tokens
Input per 1m tokens
$0.180000
Output per 1m tokens
$0.590000
Minimum hold
$0.010000
MiniMax M2
MiniMax-M2
MiniMax M2 through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 196,608 tokens
Max output: 196,608 tokens
Input per 1m tokens
$0.300000
Cached input per 1m tokens
$0.030000
Output per 1m tokens
$1.200000
Mistral Small 3.1 24B Instruct
mistral-small-3.1-24b-instruct
Mistral Small 3.1 24B Instruct through Cloudflare Workers AI with Omixa provider routing.
Chat
Streaming Context window: 128,000 tokens
Max output: 8,192 tokens
Nemotron 3 120B A12B
nemotron-3-120b-a12b
Nemotron 3 120B A12B through Cloudflare Workers AI with Omixa provider routing.
Chat
Streaming Context window: 128,000 tokens
Max output: 8,192 tokens
Input per 1m tokens
$2.000000
Cached input per 1m tokens
$0.500000
Output per 1m tokens
$8.000000
Phi 2
phi-2
Phi 2 through Cloudflare Workers AI with Omixa provider routing.
Chat
Streaming Context window: 2,048 tokens
Max output: 2,048 tokens
Input per 1m tokens
$2.000000
Cached input per 1m tokens
$0.500000
Output per 1m tokens
$8.000000
Qwen2.5 Coder 32B Instruct
qwen2.5-coder-32b-instruct
Qwen2.5 Coder 32B Instruct through Cloudflare Workers AI with Omixa provider routing.
Chat
Streaming Context window: 32,768 tokens
Max output: 8,192 tokens
Input per 1m tokens
$2.000000
Cached input per 1m tokens
$0.500000
Output per 1m tokens
$8.000000
Qwen3 235B A22B Instruct 2507
qwen3-235b-a22b-instruct-2507
Qwen3 235B A22B Instruct 2507 through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 262,144 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.220000
Cached input per 1m tokens
$0.022000
Output per 1m tokens
$1.800000
Qwen3 Coder 480B A35B Instruct
qwen3-coder-480b-a35b-instruct
Qwen3 Coder 480B A35B Instruct through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 262,144 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.220000
Cached input per 1m tokens
$0.022000
Output per 1m tokens
$1.800000
Qwen3 Next 80B A3B Instruct
qwen3-next-80b-a3b-instruct
Qwen3 Next 80B A3B Instruct through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 262,144 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.220000
Cached input per 1m tokens
$0.022000
Output per 1m tokens
$1.800000
Qwen3 Next 80B A3B Thinking
qwen3-next-80b-a3b-thinking
Qwen3 Next 80B A3B Thinking through Google Vertex AI Model Garden/MaaS with Omixa routing and streaming.
Chat
Streaming الأدوات Context window: 262,144 tokens
Max output: 65,536 tokens
Input per 1m tokens
$0.220000
Cached input per 1m tokens
$0.022000
Output per 1m tokens
$1.800000