Skip to main content
OT

Glossary

Key terms in cross-border live streaming and real-time translation, alphabetically ordered.

Automatic speech recognition (ASR)
Automatic speech recognition (ASR) is the technology that converts a speech signal into text. In live translation it is the first stage of the chain, and the text it produces sets the ceiling for everything that follows.
Also known as:ASR · speech-to-text · STT
Chroma key
Chroma key is the technique of making a region of an image transparent based on its colour. For live subtitles, a translation tool renders captions on a pure green background, and OBS keys the green out so only the text remains over the video.
Also known as:green screen · colour keying · chroma keying
End-to-end latency
End-to-end latency is the total time between the host speaking and an overseas viewer understanding the translation. Two definitions are common — speech to caption, and speech to the end of spoken translation — and they can differ by several seconds.
Also known as:E2E latency · total latency · translation delay
Text-to-speech (TTS)
Text-to-speech (TTS) is the technology that converts written text into audio. In live translation it reads the translated text aloud, and because its cost scales with text length it is one of the largest contributors to end-to-end latency.
Also known as:TTS · speech synthesis · voice synthesis
Virtual audio cable
A virtual audio cable is a paired set of software audio devices — one virtual output and one virtual input. Anything written to the output appears verbatim on the input, letting one application hand audio directly to another without physical speakers or microphones.
Also known as:VB-Cable · BlackHole · loopback device
Voice cloning
Voice cloning is the technique of reproducing a specific speaker from a short recording sample. In cross-border live streaming it makes the translated audio still sound like the host, preserving the personal continuity the format depends on.
Also known as:voice replication · speaker cloning · custom voice
Word error rate (WER)
Word error rate (WER) is the standard measure of speech recognition accuracy, equal to the number of insertions, deletions and substitutions divided by the total words in a human-labelled transcript. Chinese normally uses character error rate (CER) instead.
Also known as:WER · character error rate · CER