01 / BEFORE YOU START
Use a clean timeline and intelligible audio
Automatic captions are created from speech recognition, so the audio matters more than the video resolution. Before transcribing, listen for music that overpowers the speaker, long silent sections, overlapping voices, or clips placed far away from their intended timeline position. A clean voice track produces a better first draft and makes the final review faster.
PeteCut can generate captions from the audio and video elements already arranged on the timeline. This is useful when you have combined narration, music, and several clips: the transcription starts from the project timeline instead of forcing you to choose and upload a separate audio file.
- Place every spoken clip at its final timeline position.
- Trim unwanted silence and accidental takes.
- Lower background music while people are speaking.
- Choose the language that matches the recording.
- Use headphones for the final text and timing review.
02 / GENERATE
Create captions from the PeteCut timeline
- Open or create a project.
Import your video and audio, then arrange the material on the timeline. If the narration begins later in the project, keep it at that position—PeteCut uses the timeline placement when it creates caption clips.
- Open Auto captions.
Select Auto captions in the editor header or open the Captions panel in the assets area.
- Use the timeline as the source.
Choose the option to generate captions from the timeline. Use the separate audio-file option only when you intentionally want captions beginning at 00:00.
- Select language and model.
Pick the spoken language. A smaller model usually starts faster, while a larger model can take longer and use more memory. Start with the default model, then retry a difficult recording with a larger one if your device can handle it.
- Generate the transcript.
The transcription runs in the browser and inserts caption text elements as a new track. The first model download can take longer than later runs.
03 / REVIEW
Correct words, timing, and line breaks
Every automatic transcript needs a human review. Play the project from the beginning and stop at names, numbers, brand terms, technical words, and fast dialogue. Correct the text directly, then drag a caption edge when it appears too early or remains visible after the phrase ends.
Readable captions are short enough to scan while the viewer watches the image. Split dense sentences at natural pauses. Avoid filling the frame with a paragraph, and do not leave a one-word caption on screen for an unnecessarily long time. If two people speak at once, simplify only when the meaning remains accurate.
| Problem | What to change |
|---|---|
| Captions begin too early | Move the caption clip right or regenerate after fixing source placement. |
| Captions drift after an edit | Regenerate from the updated timeline or realign the affected section. |
| Words are wrong | Edit names, numbers, jargon, and punctuation manually. |
| Text is hard to read | Use fewer words per caption and stronger contrast. |
04 / STYLE & EXPORT
Make subtitles readable on the final frame
Keep captions inside the safe area and away from interface elements that a social platform may place near the edges. Use a high-contrast text color, a restrained outline or background, and a consistent position. For a 9:16 project, preview at phone size before export because text that looks modest on a desktop monitor can feel oversized on a vertical screen.
Watch the complete video once with sound and once muted. The silent pass reveals whether the captions communicate enough by themselves. When the text, timing, and placement are correct, export the finished video from PeteCut.
05 / QUESTIONS
Automatic caption questions
Does PeteCut transcribe the audio already on my timeline?
Yes. Choose the timeline source in Auto captions. PeteCut mixes the audible timeline elements for transcription and preserves their project timing.
Why is the first transcription slower?
The browser may need to download and initialize the selected speech-recognition model. Later runs can be faster when the model remains cached.
Should I publish automatic captions without editing them?
No. Always review spelling, punctuation, speaker names, timing, and line breaks before publishing.