Omixa API Docs

Audio and speech models

Developer documentation

Audio and speech models

Text-to-speech, GPT audio chat, realtime sessions, transcription-oriented catalog rows, and voice settings.

Model Reference

Audio and speech models

Text-to-speech, GPT audio chat, realtime sessions, transcription-oriented catalog rows, and voice settings. Endpoint: https://omixa.cloud/api/v1/audio

Chirp 2 Transcription

chirp-2

Google Cloud Speech-to-Text Chirp 2 for transcription and supported speech translation workflows.

صوتي
Audio per minute $0.016000
Minimum hold $0.010000
Integration docs

Chirp 3 Transcription

chirp-3

Google Cloud Speech-to-Text Chirp 3 for multilingual transcription, automatic language recognition, and speaker-aware results.

صوتي
Audio per minute $0.016000
Minimum hold $0.010000
Integration docs

Chirp Transcription

chirp

Google Cloud Speech-to-Text Chirp for speech recognition and transcription.

صوتي
Audio per minute $0.016000
Minimum hold $0.010000
Integration docs

Gemini 2.0 Flash Live

gemini-2.0-flash-live-001

Gemini 2.0 Flash Live for speech, transcription, translation, or voice generation workflows.

صوتي Streaming الأدوات Context window: 1,048,576 tokens Max output: 8,192 tokens
Input per 1m tokens $0.500000
Output per 1m tokens $2.000000
Audio per minute $0.018000
Integration docs

Gemini 2.5 Flash Live Preview

gemini-2.5-flash-live-preview

Gemini 2.5 Flash Live Preview for speech, transcription, translation, or voice generation workflows.

صوتي Streaming الأدوات Context window: 1,048,576 tokens Max output: 8,192 tokens
Input per 1m tokens $0.500000
Output per 1m tokens $2.000000
Audio per minute $0.018000
Integration docs

Gemini 2.5 Flash TTS

gemini-2.5-flash-tts

Gemini 2.5 Flash TTS for speech, transcription, translation, or voice generation workflows.

صوتي Streaming Streaming supported Reasoning controls: minimal, low, medium, high
Input per 1m tokens $0.500000
Output per 1m tokens $10.000000
Audio per minute $0.015000
Integration docs

Gemini 2.5 Flash-Lite TTS Preview

gemini-2.5-flash-lite-preview-tts

Gemini 2.5 Flash-Lite TTS Preview for speech, transcription, translation, or voice generation workflows.

صوتي Streaming Streaming supported Reasoning controls: minimal, low, medium, high
Input per 1m tokens $0.500000
Output per 1m tokens $10.000000
Audio per minute $0.015000
Integration docs

Gemini 2.5 Pro TTS

gemini-2.5-pro-tts

Gemini 2.5 Pro TTS for speech, transcription, translation, or voice generation workflows.

صوتي Streaming Streaming supported Reasoning controls: low, medium, high
Input per 1m tokens $1.000000
Output per 1m tokens $20.000000
Audio per minute $0.030000
Integration docs

Gemini 3.1 Flash Live Preview

gemini-3.1-flash-live-preview

Gemini 3.1 Flash Live Preview for speech, transcription, translation, or voice generation workflows.

صوتي Streaming الأدوات Context window: 1,048,576 tokens Max output: 8,192 tokens
Input per 1m tokens $0.750000
Output per 1m tokens $4.500000
Audio per minute $0.018000
Integration docs

Gemini 3.1 Flash TTS Preview

gemini-3.1-flash-tts-preview

Gemini 3.1 Flash TTS Preview for speech, transcription, translation, or voice generation workflows.

صوتي Streaming Streaming supported Reasoning controls: minimal, low, medium, high
Input per 1m tokens $1.000000
Output per 1m tokens $20.000000
Audio per minute $0.030000
Integration docs

GPT Audio

gpt-audio

GPT Audio for speech, transcription, translation, or voice generation workflows.

صوتي Context window: 128,000 tokens Max output: 16,384 tokens
Input per 1m tokens $2.500000
Output per 1m tokens $10.000000
Minimum hold $0.010000
Integration docs

GPT Audio 1.5

gpt-audio-1.5

GPT Audio 1.5 for speech, transcription, translation, or voice generation workflows.

صوتي Context window: 128,000 tokens Max output: 16,384 tokens
Input per 1m tokens $2.500000
Output per 1m tokens $10.000000
Minimum hold $0.010000
Integration docs

GPT Realtime 1.5

gpt-realtime-1.5

GPT Realtime 1.5 for speech, transcription, translation, or voice generation workflows.

صوتي Context window: 32,000 tokens Max output: 4,096 tokens
Input per 1m tokens $4.000000
Cached input per 1m tokens $0.400000
Output per 1m tokens $16.000000
Integration docs

GPT Realtime 2

gpt-realtime-2

GPT Realtime 2 for speech, transcription, translation, or voice generation workflows.

صوتي Context window: 32,000 tokens Max output: 4,096 tokens
Input per 1m tokens $4.000000
Cached input per 1m tokens $0.400000
Output per 1m tokens $24.000000
Integration docs

Chirp 3 HD TTS

chirp-3-hd

Google Cloud Text-to-Speech Chirp 3 HD for high-fidelity single-speaker voice generation.

صوتي
Minimum hold $0.010000
Integration docs
Copied Markdown