Auto-Generate Lyric Subtitles with Whisper (Groq, OpenAI or In-Browser)
Typing timestamps for a whole album is tedious. OpenAI's Whisper speech recognition model can listen to a song and produce lyrics with timestamps in seconds. This article compares the three ways to use Whisper in LyricsFlow MVMaker and explains how to get the most accurate subtitles from sung vocals.
How Whisper handles music
Whisper was trained mostly on speech, but it copes with singing reasonably well, especially when the vocal is clear and the mix is not too dense. It outputs text in segments with start and end times, which map naturally to SRT blocks. It has weaknesses you should know about:
- Long held notes and melisma may be transcribed as a shortened word or skipped.
- Repeated choruses are sometimes merged or, occasionally, "hallucinated" in instrumental parts.
- Invented words, names and mixed languages (for example English phrases inside a Japanese song) are often misspelled.
- Segment boundaries follow pauses in the audio, not your intended phrasing.
So treat the output as a draft: the timing skeleton is usually good, and the text needs a proofreading pass against your lyric sheet.
Three options compared
| Groq Whisper API | OpenAI Whisper-1 API | In-browser Whisper | |
|---|---|---|---|
| Speed | Fastest (seconds for a full song) | Fast | Depends on your PC; slower |
| Cost | Free tier available | Pay-as-you-go | Free |
| API key | Required (your own) | Required (your own) | Not needed |
| Models | large-v3-turbo / large-v3 | whisper-1 | Tiny (~75 MB) / Base (~140 MB) |
| Accuracy | High | High | Lower; best for clear vocals |
| Where audio goes | Sent from your browser to Groq | Sent from your browser to OpenAI | Stays on your device |
For most people, the Groq API with whisper-large-v3-turbo is the best balance. Choose whisper-large-v3 if the turbo result misses too many words. Use the in-browser option when the song is unreleased and you prefer not to send it anywhere, or when you have no API key.
Step by step in LyricsFlow MVMaker
- Open the editor and load your audio file.
- Open the automatic SRT generation panel.
- Choose the provider. For Groq, create a key at console.groq.com/keys and paste it. The key is saved in your browser only.
- Choose the transcription language when using the in-browser model. Setting the language explicitly avoids the model guessing wrong on the intro.
- Start the generation. When it finishes you see the elapsed time, phrase count and an SRT preview.
- Click Apply to Timeline, or Save SRT to edit it elsewhere first.
Tips for better accuracy
- Use a vocal-forward mix. If you produced the song, export a version with louder vocals (or the a cappella stem) just for transcription. The timing will match the final mix as long as the length is identical.
- Trim long silence carefully. Do not cut the intro — that would shift every timestamp. Instead, delete any hallucinated lines in instrumental sections afterwards.
- Proofread with the lyric sheet. Replace words in the Phrase Inspector while playing the song; the timing usually needs only small adjustments.
- Re-split long segments. Whisper sometimes returns a whole sentence as one segment. Split it into readable phrases so the animations have room to breathe.
- Check the first and last line. These are where timing errors are most common.
Privacy considerations
With the API options, your audio file is sent directly from your browser to the provider you chose; it does not pass through LyricsFlow MVMaker's servers. Read the provider's data policy if your song is unreleased. With the in-browser option, nothing leaves your computer. See our Privacy Policy for details.
Next steps
Once the lyrics are on the timeline, press 🎲 Random Apply to All for a quick first version, then refine the key lines. Our kinetic typography tips explain how.
Ready to make your own lyric video? It runs entirely in your browser — free, no sign-up.
⚡ Open the Editor