1v1 Real-time Translation Subtitles
Scenario Overview
In cross-language 1v1 voice call scenarios, two users communicate using different languages. To achieve barrier-free communication, each user's speech needs to be translated into their own language in real-time and displayed as subtitles.
This document uses a Mandarin Chinese and English 1v1 voice call as an example to demonstrate how to implement bidirectional language translation using ZEGO real-time audio/video products combined with cloud real-time speech recognition and translation features, and display only the other party's translated subtitles on the client side.

Core Flow
Sequence Diagram
The following diagram shows the core interaction flow for 1v1 real-time translation subtitles:
Flow Description
| Step | Description |
|---|---|
| 1 | Both users use ZEGO Express SDK to join the same real-time audio/video (RTC) room and publish streams with s1 and s2 as stream IDs respectively. |
| 2 | Business server calls StartRealtimeASRTask to create a recognition task, using StreamList to configure ASR and translation parameters for both streams separately. |
| 3 | ZEGO cloud ASR service joins the RTC room and pulls both audio streams. |
| 4 | Real-time recognition of user speech and translation, translation results are delivered to clients via room signaling. |
| 5 | After the call ends, the server calls StopRealtimeASRTask to stop the task. |
Prerequisites
- Cloud real-time speech recognition service has been enabled in ZEGOCLOUD Console.
- Translation vendor (such as Doubao, Qwen Machine Translation) API Key has been purchased and obtained.
- Latest ZEGO Express SDK has been downloaded and integrated.
- Client has implemented room joining and stream publishing functionality. This example assumes:
- User A (Chinese user):
userIdisu1,streamIdiss1 - User B (English user):
userIdisu2,streamIdiss2
- User A (Chinese user):
Implementation Steps
Step 1: Server Creates Recognition and Translation Task
The server calls the StartRealtimeASRTask interface, uses RecognitionRange: 1 to enable stream-level recognition, and configures ASR and translation parameters for both streams in StreamList:
- Stream
s1(Chinese user): Chinese recognition to English translation (for English user to view) - Stream
s2(English user): English recognition to Chinese translation (for Chinese user to view)
{
"RoomId": "your_room_id",
"RecognitionRange": 1,
"SubtitleType": 2,
"StreamList": [
{
"StreamId": "s1",
"ASR": {
"Vendor": "Tencent",
"Params": {
"EngineModelType": "16k_zh"
}
},
"EnableTranslation": true,
"Translation": {
"Vendor": "DoubaoSeedTranslation",
"SourceLanguage": "zh",
"TargetLanguage": "en",
"LLM": {
"Url": "https://ark.cn-beijing.volces.com/api/v3/responses",
"ApiKey": "your_doubao_api_key",
"Model": "doubao-seed-translation-250915"
}
}
},
{
"StreamId": "s2",
"ASR": {
"Vendor": "Tencent",
"Params": {
"EngineModelType": "16k_en"
}
},
"EnableTranslation": true,
"Translation": {
"Vendor": "DoubaoSeedTranslation",
"SourceLanguage": "en",
"TargetLanguage": "zh",
"LLM": {
"Url": "https://ark.cn-beijing.volces.com/api/v3/responses",
"ApiKey": "your_doubao_api_key",
"Model": "doubao-seed-translation-250915"
}
}
}
]
}Parameter Description
| Parameter | Description |
|---|---|
RoomId | Real-time audio/video (RTC) room ID, must match the room ID that clients join. |
RecognitionRange | Set to 1 to recognize streams specified in StreamList. |
SubtitleType | Set to 2 to deliver only translation results. |
StreamList | Stream configuration list, each stream can have independent ASR and translation parameters. |
StreamList[].ASR.Params.EngineModelType | ASR engine model, 16k_zh for Chinese, 16k_en for English. For more languages, refer to Configure ASR. |
StreamList[].Translation.SourceLanguage | Source language, the language spoken by the current stream user. |
StreamList[].Translation.TargetLanguage | Target language, the language after translation. |
- This example uses Doubao translation model
doubao-seed-translation-250915. You can also use Qwen Machine Translation (QwenMT), for details refer to Configure Translation. - In production environments, please make sure to fill in the correct
ApiKey.
Step 2: Client Displays Other Party's Translation Subtitles
The client receives and displays translation subtitles through the ZEGO subtitle component. To implement "display only the other party's translation subtitles", you need to filter out your own messages in the subtitle processing logic.
Integrate Subtitle Component
Please refer to the Display Subtitles document to download and integrate the subtitle component.
Filter to Display Only Other Party's Subtitles
In the subtitle message processing logic, compare the UserId field in the message with the local user ID to display only the other user's subtitles:
Step 3: (Optional) Stop Recognition Task
After the call ends, the server calls the StopRealtimeASRTask interface to stop the task:
{
"TaskId": "your_task_id"
}If the stop interface is not called proactively, the backend will automatically stop the recognition task when there are no real users in the RTC room for more than MaxIdleTime (default 120 seconds).
Notes
- SDK Version Requirement: You must use the Express SDK version optimized for Cloud ASR, otherwise you cannot properly receive subtitle signaling.
- Stream ID Consistency: The
StreamIdin the serverStreamListmust match exactly the stream ID used when the client publishes streams. - Translation Direction Configuration: Please configure the correct
SourceLanguageandTargetLanguagebased on the actual user language to ensure translation results meet expectations. - API Key Security: In production environments, the translation service's
ApiKeyshould be managed securely by the server to avoid leakage. - Message Ordering: Translation text received via room signaling may arrive out of order, the subtitle component has built-in processing logic to sort by
SeqId.
