{
"audio": "UklGRqRQAgBXQVZFZm10IBAAAAABAAEAwF0AAIC7AAACABAAZGF0YYBQAgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAP//AAAAAAAA//8AAP//AAD/////AAACAAAAAAD//wEAAQABAP7/AAD9//3/AAD+//7////9//z//f/8/wAAAAAAAP7/AgABAP3/+//+/wAA//8BAP//+/8AAP///v/9//z///8DAP//AQAAAP3//P////v/AgD8//7/+//9//z//v/+/////f/+/wAA/v8DAP//AgAEAAMABwAGAAUABAAAAAIAAgADAAMAAQADAAMAAgABAAAAAAD///7//f////v//f8BAAIABAAAAAEAAgABAAIAAwAAAAAA//////v///8CAP7//P/+//////8AAPz///8BAAMABAAAAAAAAQADAAAAAQAAAP///v8AAP3/AAD+/wAA/f/9//3/AQABAAQAAgD+////AwD7//7///8DAAEAAQAAAAAAAgAAAAAAAAABAAQAAAAAAP//AwACAAIA//8CAAAAAgAEAP7/BAD7////CQD9/wYABgD+/wQA/v/7/wMAAgAAAAUAAAADAAoAAgADAAAA/v/+/wEA//8AAAYA/v/8//j//P/5//z/AwAFAAAABwD+////BgACAAcAAAACAP7///8KAAIAAgD9//z/AQADAAMABAABAPz//f8AAPr/9f8AAAUABwD///T/9v/6//r//f8EAAEA//8CAP//+v8CAP7//f8DAP7//v/8//r/+//9/wEAAgAAAP7/AQACAAEAAwABAAAAAAADAAQAAgD7//3//f/+/wEABAADAAIAAQADAAUA//8BAAEAAQD9//////8BAP///v/9...",
"mimeType": "audio/wav",
"chunksSynthesized": 1,
"_note": "Response truncated for documentation purposes"
}
curl --location --request POST 'https://zylalabs.com/api/13291/multilingual+speech+recognition+and+synthesis+api/27504/text+to+speech+tts+-+multilingual+synthesis' --header 'Authorization: Bearer YOUR_API_KEY'
--data-raw '{"text": "Hello world from Zyla Labs", "provider": "gemini", "voiceName": "en-US-Journey-F"}'
{
"text": "Xin lỗi, bạn chưa đính kèm file âm thanh nào vào yêu cầu. Vui lòng tải file âm thanh lên hoặc cung cấp đường dẫn (link) để tôi có thể hỗ trợ bạn chuyển đổi thành văn bản.",
"provider": "Google Gemini",
"model": "gemini-3.1-flash-lite",
"durationSecs": 1,
"chunksProcessed": 1
}
curl --location --request POST 'https://zylalabs.com/api/13291/multilingual+speech+recognition+and+synthesis+api/27505/speech+to+text+stt+-+multilingual+audio+transcription' --header 'Authorization: Bearer YOUR_API_KEY'
--data-raw '{
"audioBase64": "UklGRiQAAABXQVZFZm10IBAAAAABAAEAQB8AAEAfAAABAAgAZGF0YQAAAAA=",
"mimeType": "audio/wav",
"lang": "vi"
}'
After signing up, every developer is assigned a personal API access key, a unique combination of letters and digits provided to access to our API endpoint. To authenticate with the Multilingual Speech Recognition and Synthesis API simply include your bearer token in the Authorization header.
| Header | Description |
|---|---|
Authorization
|
Required
Should be Bearer access_key. See "Your API Access Key" above when you are subscribed.
|
No long-term commitment. Upgrade, downgrade, or cancel anytime. Free Trial includes up to 50 requests.
High-fidelity 24kHz Neural Text-to-Speech TTS and 99.8 percent accurate Speech-to-Text STT API supporting 70 global languages including Vietnamese, English, Japanese, Korean, Chinese, French, German, Spanish. Powered by Google Gemini and OpenAI ChatGPT.
The Text to Speech (TTS) endpoint returns high-fidelity WAV audio files synthesized from input text, while the Speech to Text (STT) endpoint returns verbatim transcriptions of uploaded audio files in text format.
For TTS, the key field is "audio," which contains the audio data. For STT, key fields include "text" (the transcribed text), "provider" (the service used), "model" (the specific model employed), "durationSecs" (length of audio), and "chunksProcessed" (number of audio segments processed).
The TTS endpoint accepts parameters such as "text" (the input text), "language" (the desired language for synthesis), and "voice" (the specific voice model). The STT endpoint accepts parameters like "audioFile" (the audio to be transcribed) and "language" (the language of the audio).
The TTS response contains a single field "audio" with the synthesized audio data. The STT response is structured with multiple fields: "text" for the transcription, "provider" for the service used, "model" for the transcription model, "durationSecs" for audio length, and "chunksProcessed" for the number of processed segments.
The TTS and STT functionalities are powered by advanced technologies from Google Gemini and OpenAI ChatGPT, ensuring high-quality audio synthesis and transcription accuracy.
Typical use cases include creating voiceovers for videos using TTS, generating audiobooks, and transcribing meetings or interviews with STT for documentation and accessibility.
Users can play the audio returned by the TTS endpoint directly in applications or save it as a file. For STT, the transcribed text can be used for documentation, analysis, or further processing in applications like chatbots or content creation.
Data accuracy is maintained through the use of advanced neural models from Google and OpenAI, which are continuously trained and updated to improve performance and adapt to various languages and accents.
To obtain your API key, first sign in to your account and navigate to the API you want to use. From the API's Pricing section, choose a plan and complete the subscription process. Once subscribed, return to the API page and you will see your API Access Key displayed at the top of the documentation page. You can use this key to authenticate your requests.
You can’t switch APIs during the free trial. If you subscribe to a different API, your trial will end and the new subscription will start as a paid plan.
The free trial lasts for 7 days and allows you to make up to 50 API requests.
No, the free trial is available only once, so we recommend using it on the API that interests you the most. Most of our APIs offer a free trial, but some may not include this option.
Yes. If the API offers a free trial, you will see a "Free 7-Day Trial" option in its Pricing section. The trial lasts for 7 days and allows up to 50 API requests, enabling you to evaluate the API before subscribing to a paid plan.
Zyla API Hub is like a big store for APIs, where you can find thousands of them all in one place. We also offer dedicated support and real-time monitoring of all APIs. Once you sign up, you can pick and choose which APIs you want to use. Just remember, each API needs its own subscription. But if you subscribe to multiple ones, you'll use the same key for all of them, making things easier for you.
You can monitor your API usage through the response headers included with every request:
x-zyla-api-calls-monthly-used: Shows the total number of API requests you have used during the current billing period.
x-zyla-api-calls-monthly-remaining: Shows the number of API requests you have remaining for the current billing period.
Yes, you can cancel your subscription at any time. Simply go to the Pricing section of the API you're subscribed to and click the "Unsubscribe" button.
Please note that upgrades, downgrades, and cancellations take effect immediately. Once your subscription is canceled, access to the service will end immediately, regardless of any remaining API calls in your quota.
Please have a look at our Refund Policy: https://zylalabs.com/terms#refund