Auto Sync
Word Timestamp Transcript Format: What Your File Should Contain
A word timestamp transcript tells the editor exactly when each spoken word begins and ends. Once the structure is correct, Auto Sync can use it to build scene timing.
A normal transcript tells you what was said. A word-timestamp transcript tells you what was said and when each word was spoken. Auto Sync needs the second type because it uses those times to place scenes on the timeline.
The three pieces of information
- word — the spoken word, such as Imagine.
- start time — when that word begins.
- end time — when that word finishes.
Example structure
A JSON transcript can contain a words array. Each item can contain a word such as Imagine, a startOffset such as 0.200s, and an endOffset such as 0.500s. The next word then has its own start and end values.
Why start and end times both matter
The start time tells Auto Sync where a scene should begin. The end time is useful for knowing where the final spoken word finishes. Together they give the editor a real time reference instead of an estimated character count.
Keep the words in speaking order
The transcript should follow the audio from beginning to end. Do not sort words alphabetically or combine words from different parts of the recording. The order is part of the information Auto Sync uses.
Common mistakes
- Uploading a normal paragraph transcript with no timestamps.
- Using timestamps that are all zero.
- Putting milliseconds into a field that is expected to be seconds.
- Having an end time earlier than the start time.
- Using a transcript from a different voiceover recording.
- Leaving the word field empty for many entries.
How to test the file before a long edit
Use a short voiceover and a small script first. If the first few scenes line up correctly, the same structure can be used for a longer documentary or narrated video. This makes troubleshooting much easier than testing a large project first.