How to Make a Lyric Video Without an Editor
Timing words to music is the only hard part of a lyric video, and you do not need a timeline to do it.
A lyric video is two problems. One is the picture, which a visualizer solves by itself. The other is timing — getting each line on screen at the moment it is sung — and that is the part people open a video editor for, drag rectangles around for two hours, and abandon. There are three faster routes, all of them inside the browser.
Route one: a file you already have
If the song already has an .lrc (the karaoke format, timestamps in square brackets) or an .srt subtitle file, load it in Tune → Lyrics and you are finished. Timed lines start displaying against the track immediately. Plain .txt works too, but it arrives untimed and stays hidden until you time it — the status line under the box tells you how many lines are stamped and how many are still waiting, so nothing is ambiguous.
Route two: tap it in, in one pass
This is the one most people end up using, because it takes exactly as long as the song. Paste the words, one line per line, start Tap-sync mode, play the track, and press Enter as each line begins. Lines appear on screen as you stamp them, so you can see immediately whether you are landing early. If you fumble a verse, keep going and re-tap those lines afterwards; each line stores its own timestamp.
A trick worth knowing: tap slightly ahead of the vocal. A lyric that appears a beat early reads as anticipation, a lyric that appears a beat late reads as a mistake.
Route three: let Whisper do it locally
The Auto-transcribe button runs a Whisper model with word-level timestamps directly in your browser. The model downloads once — roughly 60 MB, cached after that — and then the transcription happens on your own machine. Your unreleased track is not uploaded to a captioning service, which is the usual price of automatic timing. Accuracy depends on the recording: a clear lead vocal comes back nearly perfect, a heavily processed or buried one needs editing. Every line stays editable afterwards, and any line can be re-tapped.
The whole job, in order
- Open BLOOB and load your audio file. The full track is analysed before the first frame, so scene changes will land on real section boundaries.
- Pick a scene with room in the middle of the frame. Aurora, Nebula and Deep Sonar leave the centre relatively calm; a dense particle scene will fight your text.
- Open Tune → Lyrics and choose your route: load a file, paste and tap-sync, or auto-transcribe.
- Play it back and fix the lines that drift. This is the step that separates a lyric video from a slideshow.
- Set the frame shape in Export: 16:9 for YouTube, 9:16 for Shorts, Reels and TikTok, 1:1 for a feed post.
- Add a title and subtitle if the platform needs one — track name and artist is usually enough — and set the palette.
- Export offline for a frame-perfect MP4. The lyrics are drawn into the picture, so they survive every upload and every platform that strips subtitle tracks.
Details that make it look deliberate
Keep lines short. Two clauses on screen at once is the ceiling for something moving; break a long line into two stamped lines rather than shrinking the type. Watch the safe area on vertical exports — the caption-safe guide shows where a phone's interface will sit, and text under the platform's own overlay is text nobody reads. And decide early whether the title stays up for the whole track or fades after the intro; both are defensible, but changing your mind after you have exported four versions is not.
If the song has a long instrumental, let the visual carry it rather than holding the last line on screen. Silence in the words is not silence in the picture.
Captions are the same job with a different purpose
Everything above works for spoken word too: an audiogram of a podcast clip with the transcript burned in, a teaching video, an accessible upload where the words have to be visible rather than optional. The mechanics do not change — load, tap or transcribe, then export — and the local transcription is the same model. For a listener who is deaf or hard of hearing, the words plus an honest rhythm on screen is a meaningfully different experience from a static waveform.
Keep reading
- Add captions to a video free — the three timing routes in detail.
- Audiogram maker — the podcast version of the same workflow.
- Seeing music — why the words and the rhythm belong on one screen.