Skip to content
Audio Editing & Production

Voice-Over Recording: What Actually Works

Six common voice-over recording myths, debunked: why the room beats the microphone, how much headroom to leave, working distance, takes and room tone.

Tomas Lindqvist Post-Production Lead 26 min read 26 views
Voice-Over Recording: What Actually Works

Voice-over recording is the job of capturing a spoken performance cleanly, with a microphone, a room and a style of direction that suit the read the script actually needs. It sounds simple, and that is the trouble: because anyone can plug in a USB microphone and press record, voice-over has collected more confident folklore than almost any other part of audio production. Some of that folklore costs money, some of it costs time, and some of it quietly ruins recordings that nobody can rescue afterward.

This guide is for the people who commission, produce or record narration: marketing teams making explainer videos, course designers recording e-learning modules, podcast producers, agencies briefing talent, and in-house staff who have been handed a microphone and a deadline. It works through the misconceptions we see most often, explains why each one fails, and sets out what to do instead. The underlying principle runs through every section: a voice-over is one of the few production elements where the recording conditions set the ceiling entirely. Editing and processing can polish a good recording, but no processing recovers a bad room.

Each section below takes one myth, states the reality, and then gets into the practical detail: distances, levels, room choices, take management and the habits that make a session repeatable. At the end you will find a worked example with realistic settings and a checklist you can take into your next session.

Why Voice-Over Recording Quality Is Decided Before Anyone Speaks

Most production steps are forgiving. A color grade can be redone, a graphic can be re-exported, a music bed can be swapped. A voice recording is different because the sound that reaches the microphone is a mix of two things at once: the direct sound of the voice, and the room's reflections of that voice arriving a few milliseconds later. Once those two are summed on the recording, they are baked together. Modern tools can reduce steady background noise and soften some reverberation, but they work by guessing what the clean signal would have been, and the harder they work, the more they leave audible artifacts: a watery, hollow or lisping quality that listeners notice even if they cannot name it.

That is why experienced engineers spend most of their setup time on things that happen before the performer says a word: choosing and treating the space, placing the microphone, setting a sensible input level, marking up the script, and agreeing how takes will be named. The recording itself is the shortest part of a well-run session. The editing that follows, which covers breath control, timing, noise cleanup and loudness, is covered in detail in our guide to building voice processing chains properly, but every one of those steps is easier, faster and more transparent when the source is clean.

There is also a practical cost argument. Re-recording is expensive: talent fees, studio or booth time, scheduling, and the risk that the performer's voice on a different day will not match the original. A session that captures usable, consistent audio the first time is almost always cheaper than one that has to be patched later, even if the first one took an extra hour of preparation.

With that frame in place, here are the misconceptions that most often get in the way.

Myth One: A More Expensive Microphone Will Fix the Sound

Myth: If the recording sounds boxy, echoey or thin, the answer is to buy a better microphone.

Reality: Room treatment, and in particular absorption at the first reflection points, matters more than microphone choice. A better microphone in a bad room simply records the bad room in more detail.

This is the single most expensive misconception in voice-over. When a recording sounds wrong, the microphone is the most visible piece of equipment, so it gets the blame. But the "boxy" or "roomy" character that people hear is usually caused by early reflections: sound that leaves the performer's mouth, bounces off a nearby wall, desk, window or ceiling, and arrives at the microphone a fraction of a second after the direct sound. Those delayed copies interfere with the direct voice and create comb filtering, a series of peaks and dips across the frequency range that makes speech sound hollow and colored.

Why rooms with hard parallel surfaces are the worst case

Small rooms with bare, parallel walls are especially troublesome because sound bounces back and forth between the surfaces, producing flutter echo and strong resonances at particular frequencies. A kitchen, a bathroom, an empty office or a meeting room with a glass wall are classic examples. These are exactly the rooms people tend to choose because they are quiet, yet they are among the worst places to record a voice. A quiet room is not the same as a good-sounding room.

What actually fixes it

The first reflection points are the places on the walls, ceiling and any nearby surfaces where sound from the mouth reflects once on its way to the microphone. Placing broadband absorption there, typically dense panels several inches thick rather than thin acoustic foam tiles, removes the reflections that do the most damage. Soft furnishings help too: a room full of books, sofas, curtains and carpet often records far better than a bare "studio" space. For people working at home, a walk-in closet full of clothes is a well-known and genuinely effective option, because the clothing absorbs a broad range of frequencies and breaks up reflections.

Our detailed guide to room acoustics for recording walks through how to find reflection points and how much treatment a small voice space needs. As a rule of thumb for budgets: if you have a fixed amount to spend and the room is untreated, put the money into treatment before upgrading the microphone. A mid-range microphone in a well-treated space will beat a premium one in a bare room every time.

A quick diagnostic you can run today

Stand where the performer will stand and clap sharply once. If you hear a ringing, zingy tail after the clap, you have flutter echo from parallel surfaces. Then record a short test read, listen on good closed-back headphones, and pay attention to the gaps between phrases: if the end of each word seems to hang in the air, the room is contributing too much. Move the setup, add absorption behind and in front of the performer, and repeat until the tail is short and the voice sounds dry and present.

What to do and what to avoid with voice-over recording, side by side
Good practice against the usual mistakes, from the sources listed below.

Myth Two: A Condenser Microphone Is Always the Professional Choice

Professional voice studios are full of large-diaphragm condenser microphones, so it is natural to assume that is what you should buy. The reality is more nuanced, and the right answer depends on the room as much as on the voice.

Myth: Serious voice-over always uses a large-diaphragm condenser microphone; dynamic microphones are for live stages and podcasts on a budget.

Reality: A large-diaphragm condenser is chosen for detail in a treated room. A dynamic microphone is often the better tool in an untreated or partly treated room because it picks up less of the space and less background noise.

How the two types differ in practice

Condenser microphones are generally more sensitive and capture more high-frequency detail and air. In a well-treated booth, that sensitivity is exactly what you want: every nuance of the voice is preserved and the recording sounds open and polished. In an untreated room, that same sensitivity picks up computer fans, traffic, the refrigerator two rooms away and, above all, the room's reflections. Condensers also require power, usually 48-volt phantom power supplied by the audio interface; if you are new to that, our explainer on phantom power covers what it is and the handful of precautions worth taking.

Dynamic microphones are less sensitive and are typically worked closer to the mouth. Because the voice is so much louder than the room at that distance, the ratio of direct sound to reflected sound improves, and less of the space ends up on the recording. The trade-off is that dynamics need more gain from the preamp, can sound slightly less detailed at the top end, and demand more consistent technique from the performer because small changes in distance make a bigger difference to level and tone.

Choosing between them

SituationUsually the better fitWhy
Treated booth or studio, quiet environmentLarge-diaphragm condenserCaptures full detail; the room adds little that needs removing
Home office or untreated roomDynamicRejects more room sound and background noise at close range
Partly treated room, moderate noiseEither; test bothDepends on the voice, the noise floor and how close the performer can work
Performer who moves a lot or varies intensityCondenser at a steady distance, or dynamic with coachingCondensers tolerate small distance changes better; dynamics need discipline
Matching a previous sessionWhatever was used last timeChanging microphone type changes the tonal character audibly

The last row matters more than people think. If a series of training modules was recorded on one microphone, switching to a different model halfway through will be audible to anyone listening to the course in sequence, even if the new microphone is objectively better. Consistency often outranks quality once a project is underway.

Myth Three: Record as Hot as Possible for the Best Quality

This myth is a holdover from the days of analog tape and early 16-bit digital systems, where recording close to the top of the meter was a way of keeping the signal well above the noise. It no longer applies, and following it is one of the commonest ways to ruin a take permanently.

Myth: You should set the input level so the loudest peaks come close to 0 dBFS, because a louder recording is a higher-quality recording.

Reality: Leave substantial headroom and normalize afterward. Clipped peaks cannot be restored, while a recording made with generous headroom loses nothing meaningful at modern bit depths.

Why clipping is unrecoverable

Digital audio has a hard ceiling at 0 dBFS (decibels relative to full scale). Any part of the waveform that would exceed that ceiling is simply cut off flat. The result is distortion, often a harsh crackle on the loudest syllables: a laugh, an emphasized word, a sudden plosive. Declipping tools exist, but they reconstruct the missing peaks by estimation and work only on light, occasional clipping. A performer who gets excited on the key line of the script will often push past the level they gave during a polite sound check, and that key line is exactly the one you cannot afford to lose.

What sensible levels look like

Recording at 24-bit resolution gives an enormous usable dynamic range, so there is no penalty for leaving plenty of space above the voice. A common working practice for spoken word is to set the gain so that normal speech sits well below full scale, with the loudest expected peaks still leaving comfortable margin before 0 dBFS. Exact targets vary between engineers and between voices, and there is no single correct number; what matters is that the performer's most energetic moments still have room. Do the level check with the most intense line in the script, not the opening sentence, and ask the performer to deliver it at full performance intensity.

Recording software documentation is a good place to confirm how your specific tools meter and handle input levels. Both Steinberg's recording documentation and Adobe's Audition recording guides explain input metering, clip indicators and bit-depth settings for their applications, and it is worth reading the relevant pages for whichever software you use before a session rather than during it.

Normalizing and loudness afterward

After recording, the edited voice is brought up to its delivery level with gain, normalization, compression and a final loudness target appropriate to the platform. That is a mastering decision, made once, on the finished edit. Trying to achieve loudness at the recording stage gains you nothing and risks everything.

Myth Four: Getting Closer to the Microphone Always Sounds Warmer and Better

Radio presenters often work very close to the microphone, and the resulting rich, intimate sound has become shorthand for "professional voice." So performers lean in. Sometimes that helps. Often it causes problems that are hard to fix.

Myth: The closer you are to the microphone, the warmer, fuller and more professional the voice will sound.

Reality: Directional microphones exhibit proximity effect, a bass lift that grows as the source gets closer. A working distance of roughly a hand span, adjusted for how much proximity effect the read needs, gives a controlled, natural result.

Understanding proximity effect

Cardioid and other directional microphones boost low frequencies as the sound source gets closer. A little of this adds weight and intimacy to a voice, which is why it is used deliberately for late-night radio or a close, confiding read. Too much makes the voice boomy and muddy, reduces intelligibility, and, critically, makes the tone change every time the performer moves an inch or two. The Audio Engineering Society publishes extensive technical literature on microphone technique and directional microphone behavior for anyone who wants the physics in depth.

Setting the working distance

A typical starting point is roughly a hand span between mouth and microphone capsule, somewhere around six to eight inches for most adults. From there, adjust by ear:

  • If the voice sounds thin or distant, or the room is becoming audible, move a little closer.
  • If the voice sounds boomy, chesty or muddy, or the level jumps noticeably when the performer shifts, move back.
  • For dynamic microphones in untreated rooms, the working distance is usually shorter, which is part of why they reject the room so well; watch for proximity buildup and use a high-pass filter in post if needed.

Once the distance is set, it needs to stay set. Mark the position of the stand, measure the distance, and note it in the session log. A performer who drifts forward and back over a long read will produce a recording whose tone swings from line to line, and that is tedious, sometimes impossible, to even out with equalization.

Plosives and the off-axis trick

Close working also increases plosives, the bursts of air from "p," "b" and "t" sounds that thump into the capsule. Two simple measures handle most of them. First, use a pop filter, either mesh or fabric, a few inches in front of the microphone. Second, set the microphone slightly off-axis: angle it a little to one side of the mouth, or position it slightly above or below and aim it back toward the mouth, so the air stream passes the capsule rather than hitting it directly. The voice is still captured clearly, and the blasts of air go past. These two steps together prevent most plosive problems before they reach the recording, which is far better than trying to remove them in the edit.

Myth Five: One Good Take Is Enough Because You Can Fix It in the Edit

Time pressure makes this myth attractive. The performer delivered a clean read, the client on the call said it sounded fine, and everyone wants to move on. Then, three days later, the script changes by one word, or the editor discovers a stumble hidden under a breath, and there is nothing to cut to.

Myth: If a line sounded good once, you have it. Anything else can be fixed in editing.

Reality: Record several reads of difficult lines in the same session, while the performer is warm. A single take of anything that might need re-cutting leaves the editor with no alternatives and often forces a costly pickup session.

Why same-session alternates matter

A voice changes over the course of a day and between days. Pitch, energy, resonance, mouth noise and even the microphone position drift. Lines recorded in the same session, minutes apart, cut together seamlessly. Lines recorded a week later often do not: the new pickup can sound like a different person in a different room, even when the setup has been faithfully rebuilt. Collecting alternates while the performer is warm and the setup is untouched is the cheapest insurance in voice production.

Which lines deserve multiple takes

  • Brand names, product names, technical terms and anything with an unusual pronunciation.
  • Numbers, prices, dates and legal wording that the client might revise.
  • Emotionally key lines: the opening, the call to action, the final sign-off.
  • Long sentences where a mid-sentence breath or stumble is likely.
  • Any line the director or client hesitated over during the read.

It is also good practice to capture a few alternate reads with different emphasis on critical lines, for example stressing a different word, so that the edit can respond to feedback without another session.

Punch-ins versus full retakes

When a mistake happens mid-paragraph, there are two approaches: stop and re-read from the start of the sentence, or use punch-and-roll recording, where the performer hears a few seconds of the previous take and picks up seamlessly at the error point. Punch-and-roll produces a cleaner timeline and makes long-form narration much faster to edit; our guide to punch and roll recording explains how to set it up and when a full retake is still the better choice. Either way, clear slating and take naming are essential, because an editor facing forty unlabeled takes of the same paragraph will waste hours.

Mark up the script before the session

Multiple takes are most useful when everyone agrees in advance which lines are critical. A marked-up script shows emphasis, pronunciations, pauses and the lines flagged for alternates, and it gives the director and the performer a shared reference. Our practical guide to marking up narration scripts covers the notation that works well for commercial and e-learning reads.

Myth Six: Noise Reduction Will Clean Up Whatever the Room Adds

Modern noise reduction is impressive, and it is easy to believe it has made room choice and session discipline optional. It has not.

Myth: With today's noise-reduction and de-reverb plugins, background noise and room sound can simply be removed later, so room tone and a quiet environment no longer matter.

Reality: Processing can reduce steady noise and some reverberation, but every pass trades noise for artifacts. Clean capture comes first, and room tone recorded before the performer leaves is what makes invisible editing and gentle noise reduction possible.

What noise reduction can and cannot do

Broadband noise reduction works best on steady, consistent sounds such as a constant air-conditioning hum or preamp hiss. It does this by learning the noise profile and subtracting it from the recording. It struggles with noise that changes over time (traffic, voices in the next room, a chair creak) and with noise that overlaps the voice's frequency range. De-reverberation tools can take some edge off a lively room, but at the settings needed for a truly bad space they tend to produce a phasey, underwater sound. The cleaner the source, the lighter the processing needs to be, and the more natural the finished voice sounds.

Why room tone is non-negotiable

Room tone is a recording of the space with the microphone, gain and position exactly as they were for the performance, but with nobody speaking. Thirty seconds to a minute is typical. It matters for two reasons. First, editors use it to fill gaps: when a breath, click or cough is removed, the hole is filled with room tone rather than digital silence, so the edit is inaudible. Second, noise-reduction tools use it as the reference profile of what to remove. Without it, the editor has to hunt for tiny scraps of silence between phrases, which are often contaminated with breaths or mouth noise.

Record room tone at the end of the session, before the performer leaves and before anyone touches the gain or the microphone. Ask everyone in the room to stay still and silent. If the session changes location or setup partway through, record room tone for each setup.

Controlling noise at the source

  • Turn off air conditioning, fans and noisy computers during takes, and schedule breaks so the room does not overheat.
  • Silence phones, notifications and anything with a vibrate function; put them outside the room.
  • Use a paper or tablet script that does not rustle, and check for clothing or jewelry noise.
  • Keep water at room temperature nearby; mouth clicks increase as a performer dries out.
  • Record at a time of day when the building and street are quietest.

Keeping Voice-Over Recordings Consistent Across Sessions

Many voice projects are not one session. A course is recorded module by module, an app's prompts are updated for each release, a brand's ads are refreshed each quarter. A common assumption, closely related to the myths above, is that using the same performer and the same equipment is enough for the recordings to match. It frequently is not.

In practice, recordings only match when the whole setup is documented and reproduced, including microphone position, distance, angle, gain, room and processing. Changing the microphone position mid-session without noting it is one of the commonest causes of mismatched audio, and it is entirely avoidable.

What must be documented

A session log does not need to be elaborate, but it needs to exist. At minimum, write down the microphone model, the interface and its gain setting, the working distance and angle, the pop filter position, the room and where the performer stood or sat within it, the sample rate and bit depth, the software session template used, and any processing applied during recording. A couple of photos of the setup from fixed angles are invaluable when rebuilding it months later. Our guide to session templates explains how to save track layouts, routing and naming conventions so each new session starts from the same point.

Mid-session changes

Sometimes a change is necessary: the performer needs to sit rather than stand, a stand gets bumped, a noisy cable is swapped. That is fine, but it must be noted with the take number where it happened, and ideally a new room tone captured. The editor then knows that takes before and after the change may need slightly different equalization to match, rather than discovering the mismatch by ear and wondering what went wrong.

Remote performers on unknown equipment

Remote sessions multiply the risk. The performer may be in a different room each time, using a microphone you have never heard, with gain set by guesswork. Test recordings before the real session, a clear spec for file format and level, and a director on the call who knows what to listen for make the difference. Our guide on directing remote voice sessions covers the practical setup, from connection options to how to give notes when you cannot see the performer's face.

Worked Example: Planning a Narration Session for a Twelve-Module Course

The following example is illustrative, not a real client project. It shows how the principles above come together for a common job: a training course of twelve short modules, each with around 900 to 1,100 words of narration, to be recorded by one performer over two days, with further modules expected later in the year.

The starting situation

The organization has a spare office with a hard floor, a large window and two bare parallel walls. The initial plan is to buy a premium condenser microphone and record there. Applying the myths above, the plan changes.

  • Room first. Instead of the premium microphone, part of the budget goes to broadband absorption panels at the side-wall reflection points and behind the performer, a thick rug, and heavy curtains over the window. A clap test before treatment rings noticeably; after treatment the tail is short.
  • Microphone choice. With the room now partly treated but still exposed to some corridor noise, both a mid-range condenser and a dynamic microphone are tested on the same paragraph. The condenser sounds more open; the dynamic is quieter in the gaps. The team chooses the condenser at a hand-span distance with a pop filter and a slight off-axis angle, and records the choice.
  • Levels. The performer reads the most emphatic line in module seven at full intensity during the level check. Gain is set so that line still leaves generous headroom below 0 dBFS, recording at 48 kHz and 24-bit.
  • Script prep. Every product name and acronym is flagged for two or three alternate reads, and pronunciations are agreed with the subject-matter expert before day one.

Session time budget

At a natural narration pace, 1,000 words is in the region of seven minutes of finished audio. Raw recording time is always longer than finished time because of alternates, pickups, notes and breaks. A sensible plan assumes a multiple of the finished duration rather than a one-to-one ratio, and builds in breaks every hour to keep the voice fresh.

ItemPlanning figureNotes
Modules12Six per day across two days
Words per moduleAbout 1,000Range of 900 to 1,100
Finished audio per moduleAbout 7 minutesDepends on pace and pauses
Total finished audioAbout 84 minutes12 modules at roughly 7 minutes
Recording block per module35 to 45 minutesIncludes alternates, pickups and notes
Recording per dayAbout 4 hours of mic timePlus setup, breaks and room tone
Room tone60 seconds at the end of each dayPlus after any setup change

What the documentation looks like

At the end of day one, the log records the microphone and interface models, gain position, measured distance, angle, pop-filter position, the performer standing on a taped mark, the take-naming scheme (module, line, take), and a note that the stand was bumped and repositioned at module four, take 17. Photos are saved with the project. When the organization adds three more modules in the autumn, the setup is rebuilt from this log, a test paragraph from the original session is replayed against a new one, and adjustments are made until they match before the performer records anything for real.

The outcome to expect

Following this approach does not guarantee a perfect result, but it removes the most common failure modes: no clipped peaks on emphatic lines, no audible jump between modules, enough alternates to absorb late script changes, and clean room tone for the editor. The edit becomes a matter of selecting the best takes and polishing, rather than repairing.

When to Bring In Specialist Help

Plenty of voice-over work can be done well in-house with a treated space, sensible equipment and disciplined sessions. There are three situations, though, where outside expertise usually pays for itself.

Matching a previous session

If new lines must sit seamlessly alongside recordings made months or years earlier, and nobody documented the original setup, matching becomes detective work. An experienced engineer can analyze the original audio, rebuild a comparable chain, and use equalization and careful room choice to close the gap. That is a skilled job, and doing it badly tends to produce exactly the jarring tonal shift the project was trying to avoid.

Remote performers on unknown equipment

When the performer records from home and the equipment is a mystery, a producer who can quickly diagnose room, gain and microphone problems over a call saves days of back-and-forth. Casting also matters here; a performer with a reliable, well-treated home setup is a very different proposition from one who has never recorded remotely. Our guide to casting voice talent includes the questions to ask about a performer's recording environment before booking.

Consistent delivery across several languages

Multilingual projects add a further layer: several performers, often in different countries, need to sound as though they belong to the same production, and timing may need to fit the same picture. Consistent specs, shared reference audio, and centralized direction are what hold it together. Our overview of dubbing and localized voice tracks covers the workflow in more depth. For projects that need recording, direction and editing handled end to end, our voice and narration service covers session planning through to delivered, mastered files.

What to Do Instead: A Voice-Over Recording Checklist

The myths above share a pattern: each one looks for a shortcut, either in equipment or in post-production, that skips the unglamorous work of preparation. The alternative is a short list of habits. None of them is expensive, and together they account for most of the difference between a recording that edits easily and one that fights the editor at every step.

  • Treat the room before spending more on a microphone, starting with absorption at the first reflection points and avoiding spaces with hard parallel surfaces.
  • Choose the microphone for the room: a large-diaphragm condenser for detail in a treated space, a dynamic for untreated or noisy rooms.
  • Set a working distance of roughly a hand span, adjust for proximity effect by ear, then mark and measure it.
  • Use a pop filter and set the microphone slightly off-axis to keep plosives off the capsule.
  • Set levels with the most intense line in the script and leave generous headroom; normalize afterward.
  • Record at 24-bit so headroom costs nothing in quality.
  • Mark up the script before the session, flagging names, numbers and key lines for alternates.
  • Take several reads of difficult lines in the same session while the performer is warm.
  • Log every setup detail, and note any mid-session change with the take number where it happened.
  • Switch off fans, phones and notifications, and keep room-temperature water nearby.
  • Record room tone before the performer leaves, and again after any change to the setup.
  • Name and slate takes consistently so the editor can find alternates quickly.

If you work through that list before your next session, most of the problems described in this article will not occur. The ones that remain, such as matching old recordings, managing remote talent, or coordinating several languages, are the situations where a specialist earns their fee. Everything else is preparation, and preparation is the part of voice-over recording that no plugin can replace.

Where this comes from

The figures and practices above come from the sources listed.

Working on something like this?

We take on Audio Editing & Production work for teams who want it done once, properly. Tell us what you are building and we will tell you honestly whether we are the right studio for it. Start a project.

Where to go next

Spotted something wrong? Report an error on this page. We correct on the page and say what changed.

Frequently asked questions

The room. Reflections from nearby hard surfaces get baked into the recording and cannot be fully removed later. Absorbing the first reflection points and avoiding bare rooms with parallel walls usually improves the sound more than any microphone upgrade.
Use a large-diaphragm condenser when you have a treated, quiet room and want maximum detail. In an untreated or noisy room, a dynamic microphone worked at close range often sounds better because it picks up less of the space. If you can, test both on the same paragraph.
A common starting point is roughly a hand span, about six to eight inches for most adults. Move closer if the voice sounds thin or roomy and back if it sounds boomy from proximity effect. Once set, mark the position and keep it consistent.
Leave generous headroom below 0 dBFS rather than recording as loud as possible. Set the gain using the most intense line in the script at full performance energy, record at 24-bit, and bring the level up during editing and mastering.
Room tone is 30 seconds to a minute of the recording space with the same microphone and gain but nobody speaking. Editors use it to fill gaps where breaths or noises are removed, and noise-reduction tools use it as a reference profile. Record it before the performer leaves.
Clean, simple lines may need only one good take, but names, numbers, legal wording and key lines deserve several reads in the same session. Same-session alternates cut together seamlessly, while pickups recorded days later often sound noticeably different.
All services

The work behind this article, and what it costs.

Tomas Lindqvist

Picture and sound. Writes about editing, color, loudness and delivery specifications, including the ones that get deliveries rejected.

Keep reading

More in Audio Editing & Production