On this page

Configure ASR

2026-09-18

Feature Overview

To improve the recognition accuracy of speech recognition (or speech-to-text) in different scenarios, you can achieve this through the following methods:

  • Choose the appropriate vendor/recognition model: Supports Tencent ASR (including large model versions), Alibaba Bailian Paraformer, Alibaba Bailian Fun-ASR, Alibaba Bailian Qwen-Audio-ASR (recommended), Volcengine SeedASR (Doubao large model), Microsoft, etc.
  • Choose the appropriate language: By default, Tencent and Alibaba Bailian models are for Chinese recognition, and Microsoft is for English recognition.
  • Set recognition hotwords: In specific scenarios, there are usually some specialized vocabulary, such as character names, usernames, feature names, etc. You can set temporary hotwords when creating an agent instance to improve speech recognition accuracy.

Prerequisites

Currently, Tencent is the default enabled and supported speech recognition vendor. If you need Alibaba Bailian (including Fun-ASR, Qwen-Audio-ASR), Microsoft, Volcengine, or other recognition vendors, please contact ZEGOCLOUD business support to enable them.

Usage

When creating a real-time speech recognition task (StartRealtimeASRTask), you can set the vendor, language, hotwords, and other parameters through the ASR parameter.

ASR Parameter Description

ParameterTypeRequiredDescription
VendorStringNoASR vendor, defaults to Tencent:
  • Tencent: Tencent
  • AliyunParaformer: Alibaba Cloud Paraformer
  • AliyunFunASR: Alibaba Cloud Fun-ASR
  • AliyunQwenAudioASR: Alibaba Cloud Qwen-Audio-ASR
  • VolcSeedASR: Volcengine SeedASR (Doubao large model)
  • Microsoft: Microsoft ASR
HotWordStringNoThis parameter is deprecated.
Please set it through Params extension parameters. For specific usage, refer to the hotword setting instructions for each vendor below.
ParamsObjectNoVendor parameters. For specific usage, refer to the parameter setting instructions for each vendor below.
VADSilenceSegmentationnumberNoUsed to set how many milliseconds after the user stops speaking, two sentences will no longer be considered as one. Range [200, 2000], default is 500. For detailed explanation, refer to Speech Segmentation Control.

The Params parameter descriptions for each vendor are as follows:

Speech Segmentation Control

Determining whether the user has finished speaking can be influenced by the VADSilenceSegmentation parameter.

Scenario Example

asr_vad_example.png
ConfigurationQ&A Result
VADSilenceSegmentation = 500msUser is determined to have said 2 sentences:
1st sentence: The weather is really nice today. I want to go out and play
2nd sentence: How about you?
Note
Since 400ms < VADSilenceSegmentation, the first two segments are counted as the 1st sentence; 800ms > VADSilenceSegmentation, so the third segment is counted as an independent 2nd sentence.

Best Practice Configuration

Note
If you don't know which configuration works better, we recommend using Scenario 1 configuration.
ScenarioVADSilenceSegmentation
Scenario 1: Need to get recognition results as soon as possible. Used for displaying subtitles, etc.500ms
Scenario 2: Want the most accurate recognition results possible, can accept some delay. For example, real-time summarization.1000ms

Previous

Quick Integration

Next

1v1 Real-time Translation Subtitles