Back to blog

OpenChatCut caption workflow: transcription, style, and export

From speech transcription to word-level edits, caption styling, and export—how captions stay bound to an editable timeline.

Jul 21, 2026OpenChatCutOpenChatCut
OpenChatCut caption workflow: transcription, style, and export

Captions are central to talking-head and interview delivery. OpenChatCut keeps transcription, transcript editing, and caption tracks inside one project so agents and manual tools share the same timeline.

Recommended flow

  1. Import voiced media onto the right video/audio tracks.
  2. Transcribe (with a supported provider configured) to get word-level timings.
  3. Edit the transcript—remove fillers, fix recognition errors—so the timeline updates with the speech.
  4. Open the caption overlay, pick style, size, and safe margins.
  5. Add translation tracks when needed (depends on the current release).
  6. Export a master (burned-in or not) or caption files alone.

Transcript vs captions

The transcript is for editing by speech: delete words, tighten pauses, align speakers. Captions are for viewers: line breaks, style, in/out points.

Both should stay consistent with source media time and timeline frames. When an agent edits words or pauses, it should go through editor commands so captions do not drift from picture.

What agents can do

Typical prompts:

  • “Generate captions, keep lines under roughly 42 characters.”
  • “Remove um/uh fillers and compress pauses longer than 0.5s.”
  • “Restyle captions for vertical video.”

After the agent runs, scrub the preview. Fast speech, stacked lines, and safe-area clipping are the usual failure modes.

Export notes

  • Confirm the caption track is enabled and covers the export range
  • For NLE handoff, export project/caption files for further polish
  • On long form, spot-check more than the first ten seconds

Further reading