Skip to main content

Quick Start Video

Watch the quick demo, then follow the steps below.

AI Speech Recognition - Generate and Edit Subtitles from Audio or Video

Quantum Subtitle Speech Recognition converts spoken audio in a video or audio file into timestamped subtitle text. You can optionally identify different speakers and audio events, review the transcript in the subtitle editor, and download it as an SRT file.

What is Speech Recognition?​

Speech recognition (also called speech-to-text or ASR) transcribes spoken words from the media's audio track. It is different from Subtitle Extractor, which reads text that is already visible in video frames with OCR. Leave the audio language set to Automatic if you do not know it, or select the language yourself.

Key Features:

  • Speaker recognition: Optionally separates transcript segments by speaker. Speaker labels can be renamed in the editor.
  • Audio events: Optionally include recognized non-speech audio events in the transcript.
  • Punctuation: Choose whether punctuation appears in the transcript.
  • Timestamped output: Review and download subtitle cues in standard SRT format.
InputLimitOutput
MP3, WAV, M4A, MP4, MOV, AVI, MKVUp to 10 files per batch; 5 GB per fileEditable transcript and downloadable SRT

Recognition quality depends on the recording, background noise, overlapping voices, and names or specialist terms. Review the transcript before publishing it.

Step-by-step Guide​

Follow these simple steps to generate subtitles from your video or audio files:

  1. Upload media: Choose one or more MP3, WAV, M4A, MP4, MOV, AVI, or MKV files. Speech Recognition accepts up to 10 files in a batch, with a 5 GB maximum per file.

  2. Choose the audio language: Leave the field as Automatic when you are unsure, or select the language spoken in the recording.

  3. Set recognition options: Turn on Speaker recognition to separate voices, Audio events to include non-speech sounds, or turn punctuation on or off. These options are optional.

  4. Start recognition: Submit the task and wait for batch processing to finish.

  5. Review and export: Open the completed task in the subtitle editor. When the source media is loaded, click a subtitle row to play that segment, correct its text, and assign a speaker where needed. Rename speakers in the speaker list; changes are saved automatically. Download the reviewed subtitles as an .srt file. When speaker recognition is enabled, speaker names are included in the exported cue text.

Speech recognition or subtitle extraction?​

Choose Speech Recognition when you need captions for spoken dialogue, including audio from a video. Choose Subtitle Extractor when the video already contains visible, hardcoded subtitles and you need OCR to turn them into editable text. For a transcript that is ready to use as captions on a video, see Add subtitles to video.

Tips for Best Results​

  • Clear Audio Quality: Ensure your video/audio has clear speech without background noise
  • Single Language: For optimal accuracy, use content in one primary language
  • Appropriate Volume: Make sure the speech volume is adequate for recognition
  • File Size: Larger files may take longer to process but will still work effectively