Skip to content
HomeRecordingSetup
Menu

How to Record Vocals at Home: Microphone, Distance and the Room

Updated 2026-08-15 By Glen Gomez Meade, Composer and mix engineer
Quick answer

Record home vocals with a cardioid dynamic at three to four inches if the room is untreated, or a large diaphragm condenser at six to ten inches if it is treated. Use a pop filter two to four inches in front of the capsule, set gain from the loudest phrase so peaks land near -8 dBFS, use closed back headphones with the cue mix as quiet as the singer will accept, and record four to six full passes to comp from.

Definition: Recording a vocal at home is the problem of capturing a voice with as little of the surrounding room as possible, because the room in the file cannot be separated from the voice in the file afterward.

A home vocal recording is one signal containing two things: the voice and the room. Everything else about the process is negotiable and reversible. That is not, because there is no plugin that removes a room from a take. So the entire craft of recording vocals at home is a series of decisions that shift the ratio between those two, and the biggest lever by far is working distance.

Which microphone, and why does the room decide it?

Not the voice. The room. This is the single most common mistake in home recording, and it costs people hundreds of dollars.

Direct sound obeys the inverse square law: it drops 6 dB every time you double the distance from the source. The reflected field in a small room does not, because it is arriving from everywhere and stays roughly constant across the room. So moving a singer from 3 inches to 12 inches costs 12.0 dB of direct level against a room level that has not changed at all. That is not a subtle difference, it is the difference between a dry take and a take with a bedroom in it.

A cardioid dynamic gets used at three to four inches because it can: its low sensitivity is fine when the source is that loud and that close, and its rolled-off top end happens to hide the splashy short-delay reflections that make an untreated room sound cheap. A large diaphragm condenser gets used at eight to twelve inches because at three inches most singers overload the proximity effect and the plosive handling. The polar pattern is not what differs; the working distance is.

Direct sound level change with working distance. The room level stays roughly the same, so every figure here is also a change in the direct-to-room ratio.
Move from To Direct level change What it sounds like
2 in12 in15.6 dBDry and chesty becomes airy and roomy
3 in6 in6.0 dBProximity warmth halves, room doubles
3 in12 in12.0 dBBroadcast intimacy becomes a normal vocal
4 in16 in12.0 dBThe point where an untreated room takes over
8 in16 in-6.0 dBNatural becomes distant, useful for a big room

The Audio-Technica AT2035 is the value benchmark for a home vocal condenser: an 80 Hz roll-off switch that you will use constantly and a 10 dB pad for loud singers. If the room is untreated, buy the Shure SM7B instead and budget for the gain to drive it. If the budget is tight and the room is bad, an Shure SM58 at three inches with a pop filter gets you far closer to a usable take than an expensive condenser in the same room does. The full reasoning is in condenser vs dynamic microphone and the models at each price are in the vocal microphone roundup.

How close should the singer be, and what does proximity effect do?

Every directional microphone gains bass as the source gets closer. This is not a defect, it is a direct consequence of how a pressure gradient capsule achieves directionality: the capsule responds to the difference in pressure between its two faces, and close to a source that difference grows disproportionately at low frequencies. Below about eight inches it becomes audible, and at two to three inches it can add 10 dB or more below 200 Hz.

Used deliberately it is one of the most useful tools in vocal recording. A thin voice recorded at three inches gets a chest and a weight that no EQ boost reproduces convincingly, because the EQ is amplifying what was captured and the microphone position changed what was captured. Used accidentally it produces a muddy, boomy vocal that sits badly in a mix and gets fixed with a high pass filter that also removes wanted body.

Three ways to control it. Back off an inch or two, which costs you room rejection. Engage the microphone's own high pass, which on the AT2035 rolls off from 80 Hz and is far gentler than a plugin filter set to the same frequency. Or have the singer work slightly off axis, angling the capsule past the corner of their mouth, which reduces proximity effect and plosive energy at the same time.

The practical downside of close work is consistency. At three inches, a two inch head movement is a large proportional change in distance, and therefore a large level and tone change. At ten inches the same two inches barely registers. Singers who move a lot need either a further distance or a visible reference point, and a strip of tape on the floor works better than telling someone to stand still.

What is a pop filter actually doing?

Breaking up moving air, not filtering sound. A plosive consonant, the p in "pull" or the b in "before", fires a slug of air at the diaphragm. The diaphragm is displaced bodily, which produces a large low frequency excursion that has nothing to do with the acoustic content of the word. That is why a plosive cannot be removed cleanly with EQ: it is not a frequency in the performance, it is the microphone being pushed.

A nylon mesh Pop filter on gooseneck mounted two to four inches in front of the capsule disrupts the airflow while passing the sound essentially untouched. Metal mesh filters work by deflecting the air sideways and are easier to clean. Both are effective. The common mistake is mounting the filter right against the microphone, which leaves the air no distance to disperse in.

A Shock mount solves a different problem: mechanical vibration travelling up the stand from the floor. If you work on a suspended timber floor, or if the singer taps their foot, the shock mount is not optional. And a proper boom stand or a desk-mounted arm lets you get the microphone at mouth height without the singer craning, which affects the performance more than any of the technical details on this page. The stands roundup covers what holds up over time.

The thing that took me longest to accept is that the room affects the singer more than it affects the recording. A person standing in a small, dead, foam-lined corner with a microphone six inches from their face sings differently, and worse. They pull in, they under-support, and they get quieter and more careful. I now record vocals standing up in the open part of the room facing a duvet on a stand rather than in a boxed-in corner, and the takes are better by a margin that has nothing to do with the acoustics of the capture. Give the person room to move their arms, get the lights right, and let them hear the track properly. The performance improvement outweighs a couple of decibels of room bleed every single time.

How do you make an untreated room work anyway?

Four moves, roughly in order of effect per dollar.

Get soft mass behind the singer. The most damaging reflection is the one directly behind the source, which comes back through the rear of the cardioid pattern where rejection is best but not perfect, and arrives very early. A thick duvet hung on a stand, a wardrobe with the doors open and clothes in it, or two panels on a stand all work. This is worth more than anything you can put in front of the microphone.

Aim into the room, not at a wall. Point the microphone so its rear rejection is toward the nearest hard surface and its front is looking down the longest available path. The reflection then has further to travel and arrives quieter and later.

Avoid the exact centre of the room and avoid the corners. The centre of a dimension is where modal nulls sit and the corners are where pressure piles up. Neither is a good place for a singer or a microphone.

Then consider a reflection filter. A sE Electronics RF Pro Reflexion Filter mounted behind the microphone helps with the reflections arriving from behind the capsule, which is a genuine effect, but it is a smaller effect than the duvet and it does nothing below a few hundred hertz. Buy it after the room work, not instead of it. The full sequence is in how to treat a room for recording, and the pattern behaviour these moves rely on is charted in microphone polar patterns.

How do you set the level and the headphone mix?

Set gain from the loudest phrase the singer will actually deliver, not from a spoken check. A belted chorus runs 15 to 20 dB above a speaking voice, so a level set from "check, check" clips on the first real line, and the first real line is very often the best one on the day. Ask for the biggest moment in the song, set peaks near -8 dBFS, and let the average land around -18 to -20 dBFS RMS. Twenty four bit means there is no cost to leaving that much headroom, and the full argument is in how to set recording levels.

The cue mix matters more than the gain. A singer who cannot hear themselves sings sharp and pushes; a singer hearing too much of themselves sings flat and under-supports. Start with the voice slightly quieter than feels right, add a short plate or hall reverb in the cue mix only, and ask before you change anything. Reverb in the headphones is the single most requested and most effective adjustment there is, and it never touches the recorded file.

Why does headphone bleed only show up later?

Because it is masked while the vocal is loud and exposed when it is not. Sound leaking out of the headphones is at a fixed level. During a full chorus it is 40 dB below the voice and completely inaudible. In the two beats of silence before the last line, with the vocal fader up and a compressor pulling the quiet parts forward, it is a thin ghost of the backing track sitting under the breath. Then you cut to that section in the mix and it is suddenly obvious.

Four fixes, in order. Use closed back headphones with real isolation: Beyerdynamic DT 770 PRO, 80 ohm in the 80 ohm version is the standard choice for this exact job and the tracking headphone roundup covers the alternatives. Turn the whole cue mix down, which the singer will resist and which works. Replace a bright metronome click with a softer sound such as a woodblock or a rimshot, because it is the high frequency transient that escapes the earcup most easily. And avoid the one-ear-off habit if you can, because it forces the level up on the other side and roughly doubles the bleed.

The one that people miss: mute the click entirely for sections that do not need it. If the song has a rubato intro or a final line that rides free, the click is not helping and it is the loudest thing leaking into the microphone.

How do you comp a vocal?

Record four to six complete passes. Not phrase-by-phrase punches, complete passes, because a full performance has an arc and a continuity between lines that a stitched one does not, and because a singer who knows they can run at it again performs more freely than one being punched every eight bars.

Then comp by phrase rather than by word wherever you can. Cross-fade at breaths, not in the middle of a sustained note, because a breath masks an edit and a vowel does not. Keep the breaths in: removing every breath is the most recognisable sign of an over-edited vocal, and breaths are how a listener hears the effort in a performance.

Two practical habits worth building. Name your takes as you go rather than sorting through numbered lanes afterward. And listen to the comp all the way through in one pass before you commit, because comps assembled phrase by phrase almost always have one line whose energy does not belong with its neighbours, and it is audible only in context.

The workflow habit that has saved me the most: I record the very first pass while pretending it is a rehearsal, with the gain already set and the recorder rolling. Nobody is told that it counts. Something like a third of the vocals I end up using are built mostly from that pass, because the singer had not yet started listening to themselves critically. Once someone knows they are being recorded, the second and third takes are usually technically better and emotionally flatter, and by take five they are managing their own performance instead of delivering it. Roll early, and keep everything.

Related guides and tools

Frequently asked questions

What microphone should I use for vocals in an untreated room?

A cardioid dynamic worked at three to four inches. The Shure SM7B is the standard answer and the SM57 is the budget version of the same idea. Working that close means the direct sound is enormous relative to the room, because direct sound drops 6 dB every time you double the distance while the reflected field stays roughly constant. A condenser at ten inches captures a much higher proportion of the room, permanently.

How far should a singer be from the microphone?

Three to four inches for a dynamic, six to ten inches for a large diaphragm condenser in a room with some absorption in it. Closer gives more proximity effect and more isolation from the room but punishes any movement, because at three inches a two inch head movement is a significant level change. Further gives a more natural tone and captures more of whatever the room is doing.

What is proximity effect and how do I use it?

Proximity effect is the low frequency boost every directional microphone produces as the source gets closer, caused by the pressure gradient design that makes the pattern directional in the first place. Below about eight inches it becomes noticeable, and at two to three inches it can add 10 dB or more below 200 Hz. Use it deliberately for warmth on a thin voice, and control it with a high pass filter or by backing off.

Do I need a pop filter?

Yes, and it is the cheapest problem you will ever solve. A plosive is a burst of moving air from a p or b sound hitting the diaphragm, which produces a low frequency thump that no plugin removes cleanly because it is not tonal content, it is the diaphragm being pushed. A nylon mesh filter placed two to four inches in front of the microphone breaks up the airflow without touching the sound.

How do I stop headphone bleed getting into the microphone?

Use closed back headphones, keep the cue mix as quiet as the singer will tolerate, and pull the click level down or replace a clicky metronome sound with a softer one. Bleed only becomes audible on quiet passages and in gaps between phrases, which is exactly where you notice it in a mix. Having the singer wear one ear off makes bleed dramatically worse, so trade that for a better cue balance instead.

How many vocal takes should I record?

Four to six full passes, then comp. More than about six and the singer usually starts getting worse rather than better, and you accumulate options faster than you can evaluate them. Record complete passes rather than punching phrase by phrase, because a full performance has continuity between lines that a stitched one loses, and comping from complete takes lets you keep that continuity where it exists.

Working out your own room and signal chain? The Home Studio Build Planner is the paid version of these pages: 8 printable worksheets you fill in with your own numbers, plus the full PDF, $29.