On this page

Release Notes

2026-09-18

2026-09-17

v2.4.0
  1. Added support for the Qwen-Audio-3.0-ASR-Flash-Streaming speech recognition model. ASR.Vendor is AliyunQwenAudioASR, with hotword configuration supported, supporting Chinese (Mandarin and various dialects), English, Japanese, and other languages. For details, refer to Configure ASR.
  2. Deprecated the qwen3-asr-flash-realtime speech recognition model (formerly with ASR.Vendor set to AliyunQwenASR). If you are still using this model, it will be replaced with Qwen-Audio-3.0-ASR-Flash-Streaming later.

2026-04-07

v2.3.0
  1. Support for fallback ASR stream concurrency with billing. After enabling this capability, when the actual concurrency exceeds the purchased concurrency limit, the system automatically enables a backup ASR service for fallback, ensuring business continuity. Excess usage is billed on a daily basis, helping developers reduce decision-making costs. After enabling, the maximum concurrency can reach 2 times the subscribed monthly concurrency (for example, if the monthly concurrency is 50, it can scale up to 100).

    Note
    Contact ZEGOCLOUD business support to enable.
  2. Support for dynamically adding and removing ASR recognition streams. During a running speech recognition task, you can dynamically add or remove streams to be recognized via API without recreating the task. Add stream API: specify a taskID to add new streams and ASR-related parameters; Remove stream API: specify a taskID to remove a stream being recognized. Suitable for scenarios such as participants joining/leaving mid-meeting, dynamic switching of multi-channel signals during live streaming, etc.

2026-01-21

v2.2.0
  1. Support for translation capability based on completed speech recognition.
  • Currently supported translation granularity: room level, stream level
  • Currently supported translation models include doubao-seed-translation, Qwen-MT, etc. Note: Please purchase the authentication information for these translation models yourself before creating a task. For details, refer to Configure Translation
  1. Support for streaming subtitle delivery of recognition results and translation results through RTC room signaling. By configuring the SubtitleType field when creating a recognition task, you can get streaming recognition results or translation results from ZEGO Express SDK's RTC room messages. For details, refer to Display Subtitles

2025-12-04

v2.1.0
  1. Support for specifying certain streams for recognition when creating a speech recognition task. This enables recognizing only certain specified users' audio streams within a room.

2025-11-12

v2.0.0
  1. Support for speech recognition with unlimited number of users in a single RTC room.

  2. Added Alibaba Cloud Bailian speech recognition capability. Supports Chinese (Mandarin/dialects), Cantonese, English, Japanese, Korean, etc., with 2 types of models (contact ZEGOCLOUD business support to enable, configure vendor to select model):

    • Paraformer: Suitable for noisy environments and Chinese dialect scenarios
    • Gummy: Suitable for multi-language mixed scenarios, and German, French, Russian, Italian, Spanish scenarios

    For details, refer to Configure ASR.

  3. Added Microsoft real-time speech recognition capability. Supports English, French, German, Spanish, and a series of overseas languages. (Contact ZEGOCLOUD business support to enable)

    For details, refer to Configure ASR.

2025-07-25

v1.0.0

Brand new release. Real-time speech recognition for all audio streams in RTC rooms, converting speech to text, enabling scenarios such as online meeting real-time subtitles, multi-language voice chat room interaction, global live streaming subtitles, etc.

  1. Recognition latency around 600ms
  2. Recognition accuracy improved by 40%+
  3. Cost reduced by 50%+ compared to traditional recognition solutions

Previous

Pricing

Next

Quick Integration

On this page

Back to top