Contented Manager

Video, Audio and Repurposing · Video writing

On-screen text and lower thirds

The words written into the picture itself: the opening title, the name band under a speaker, the point being made, and the card at the end that says what to do.

Illustrated character Nora Kerrigan“I know where every piece is.”

What on-screen text is

On-screen text is editorial writing placed inside the frame: the title card that opens a clip, the lower third that names a speaker and their role, the short line that states the point being made while somebody talks around it, and the end card that says what to do next. None of it was ever spoken aloud. It is written because most feeds start a video muted and a large share of viewers never turn the sound on.

It is not captions, which are a transcript of the speech burned into the picture in time with it. It is not subtitles, which are a separate timed file the viewer switches on. Both of those are the same words the speaker said. This is different text, written by a person deciding what a silent viewer needs to know that the speaker did not think to say.

When you need it

Any clip going to a social feed needs it, because the first three seconds are read rather than heard and a talking head with no text in the frame gives a scrolling viewer nothing to stop for. Any video with more than one speaker needs lower thirds, because a viewer who cannot hear the introduction has no idea who is talking or why their opinion counts.

Recorded webinars and interviews need it most of all. An hour of somebody explaining something is full of numbers, names and defined terms that are perfectly clear when heard and vanish when they are not. Those go in the frame, spelt correctly.

You do not need this if the video is a scripted piece to camera that already says everything plainly and is going on a page where people arrive intending to watch with sound. In that case burned-in captions alone will do the job, and they cost less. You also do not need it for a training video on an internal platform, where a subtitle file serves a wider set of viewers for less money.

What I write

  • The opening card. The four to eight words that appear before anything else happens, saying what the clip is about rather than announcing the brand. This is the part that decides whether the clip is watched.
  • Lower thirds. Name, role and organisation for every speaker, written once and applied consistently across every clip from the same recording, with job titles checked with the person rather than guessed.
  • Key-point text. Short lines pulled from what is being said and rewritten to be read: a figure, a term, a step number, the thing worth remembering. Written to fit two lines on a phone.
  • Transitions and labels. The words that mark a change of subject, number a sequence of steps, or label what is on screen when the picture is a product or a screen recording.
  • The end card. One instruction and the place to do it. Not three competing ones.

Everything is written to a timing, so the editor knows the second each line appears and the second it leaves, and nothing sits over a face or collides with a caption line.

What you get

A timed text sheet. One document per video: every piece of on-screen text with its in and out time, its position in the frame, and a note on what it must not cover. An editor can work straight from it.

A lower-third list. Each speaker's name, role and organisation as they should appear, checked for spelling and accents, reusable across every future recording with the same people.

A wording standard. After the first job, one page setting the house rules: capitalisation, how numbers are written, how long a line may be, where text sits in a tall frame and where it sits in a wide one.

A worked example

An illustration, not a client. A Victoria accounting firm records a forty-minute panel with two of its partners and an outside bookkeeper on payroll changes for small employers. Six clips are cut from it for LinkedIn. Watched silently, none of them makes sense: three different voices, no names, and a remittance deadline mentioned aloud four times and never shown.

The on-screen text is written as nineteen lines across the six clips. Each clip opens with a card naming the one question it answers. Every speaker gets a lower third the first time they appear, with the bookkeeper's firm named so her credibility is visible. The remittance date appears in the frame as a dated line each time it is said. Each clip ends with the same card pointing at the firm's payroll page. The recording is unchanged. The clips are now watchable with the sound off, which is how nearly all of them are watched.

How it runs

  1. A half-hour call. What has been recorded or is about to be, where the clips are going, and who appears in them. No charge.
  2. A fixed price in writing. Per video or per batch of clips, agreed before anything starts.
  3. The material. You send the footage or the rough cuts, a transcript if one exists, and the correct name and title for every person on camera.
  4. The writing. Three to five working days for a set of up to eight clips. The first clip is written and timed for you to approve before the rest are done.
  5. Delivery. The timed text sheet, the lower-third list and the wording standard, sent to you and to whoever is cutting the video.

What it costs

Quoted as a fixed price in CAD after the half-hour call, because a single clip and a series of twenty are different pieces of work. The price depends on how many clips there are, how many speakers appear, whether a checked transcript exists already, and whether the text has to work in both a tall and a wide version of the same clip. The pricing page sets out the four ways of working, and the figure is fixed in writing before anything begins.

What happens next

On-screen text is usually written at the same time as the cutting, so the next question is often who is doing that: short-form cuts if the source is one long recording. Most clips also want the speech itself burned in, which is captions, and the two are timed together so they do not fight for the same part of the frame. And the post that carries the clip is written separately, because a well-titled clip posted with no words around it still gets scrolled past. This job stands on its own, though, and nothing further is assumed.

Know the video, audio and repurposing vocabulary?

Four short games from the terms a proposal in this field uses. The full glossary is on the video, audio and repurposing page.

The word games need JavaScript. The glossary above has every term they use.