Release Notes
2026-09-17
- Added support for the Qwen-Audio-3.0-ASR-Flash-Streaming speech recognition model.
ASR.VendorisAliyunQwenAudioASR, with hotword configuration supported, supporting Chinese (Mandarin and various dialects), English, Japanese, and other languages. For details, refer to Configure ASR. - Deprecated the qwen3-asr-flash-realtime speech recognition model (formerly with
ASR.Vendorset toAliyunQwenASR). If you are still using this model, it will be replaced with Qwen-Audio-3.0-ASR-Flash-Streaming later.
2026-04-07
-
Support for fallback ASR stream concurrency with billing. After enabling this capability, when the actual concurrency exceeds the purchased concurrency limit, the system automatically enables a backup ASR service for fallback, ensuring business continuity. Excess usage is billed on a daily basis, helping developers reduce decision-making costs. After enabling, the maximum concurrency can reach 2 times the subscribed monthly concurrency (for example, if the monthly concurrency is 50, it can scale up to 100).
NoteContact ZEGOCLOUD business support to enable. -
Support for dynamically adding and removing ASR recognition streams. During a running speech recognition task, you can dynamically add or remove streams to be recognized via API without recreating the task. Add stream API: specify a taskID to add new streams and ASR-related parameters; Remove stream API: specify a taskID to remove a stream being recognized. Suitable for scenarios such as participants joining/leaving mid-meeting, dynamic switching of multi-channel signals during live streaming, etc.
2026-01-21
- Support for translation capability based on completed speech recognition.
- Currently supported translation granularity: room level, stream level
- Currently supported translation models include doubao-seed-translation, Qwen-MT, etc. Note: Please purchase the authentication information for these translation models yourself before creating a task. For details, refer to Configure Translation
- Support for streaming subtitle delivery of recognition results and translation results through RTC room signaling. By configuring the SubtitleType field when creating a recognition task, you can get streaming recognition results or translation results from ZEGO Express SDK's RTC room messages. For details, refer to Display Subtitles
2025-12-04
- Support for specifying certain streams for recognition when creating a speech recognition task. This enables recognizing only certain specified users' audio streams within a room.
2025-11-12
-
Support for speech recognition with unlimited number of users in a single RTC room.
-
Added Alibaba Cloud Bailian speech recognition capability. Supports Chinese (Mandarin/dialects), Cantonese, English, Japanese, Korean, etc., with 2 types of models (contact ZEGOCLOUD business support to enable, configure vendor to select model):
- Paraformer: Suitable for noisy environments and Chinese dialect scenarios
- Gummy: Suitable for multi-language mixed scenarios, and German, French, Russian, Italian, Spanish scenarios
For details, refer to Configure ASR.
-
Added Microsoft real-time speech recognition capability. Supports English, French, German, Spanish, and a series of overseas languages. (Contact ZEGOCLOUD business support to enable)
For details, refer to Configure ASR.
2025-07-25
Brand new release. Real-time speech recognition for all audio streams in RTC rooms, converting speech to text, enabling scenarios such as online meeting real-time subtitles, multi-language voice chat room interaction, global live streaming subtitles, etc.
- Recognition latency around 600ms
- Recognition accuracy improved by 40%+
- Cost reduced by 50%+ compared to traditional recognition solutions
