How to Record a Podcast at Home
Record a podcast at home with a cardioid dynamic microphone about 4 to 6 inches from the mouth, slightly off axis, one microphone per person on separate tracks, and master to about -16 LUFS integrated with true peaks no higher than -1 dBTP. A dynamic beats a condenser here because it hears far less of an untreated room, which is the single largest difference between a home podcast that sounds professional and one that does not.
Definition: Recording a podcast is spoken word capture, which prioritises intelligibility, consistency across episodes and rejection of the room over the wide frequency response and dynamic range that music recording is optimised for.
Podcasting inverts the usual home studio advice in one important way. For music, a condenser microphone is the default and a dynamic is the fallback for a bad room. For speech, a dynamic is the correct first choice for almost everybody, because a spoken voice in an untreated room is the situation dynamics were designed for and because nobody ever complained that a podcast lacked air above 15 kHz.
The second inversion is that loudness is not a taste question here. Podcast delivery converged on -16 LUFS integrated for stereo, so that an episode does not arrive twice as loud as the one a listener played before it. That is a measurement, not an opinion, and it is the easiest professional standard on this page to hit.
Why does a dynamic microphone beat a condenser for speech?
Three reasons, and they compound in a normal room.
- Lower sensitivity. A dynamic converts less of everything into signal, which means room reflections, computer fans, air conditioning and traffic all arrive proportionally quieter. Sensitivity that feels like a disadvantage on paper is the entire point here.
- Tighter effective pattern in practice. Combined with the close working distance a dynamic invites, the direct-to-room ratio it delivers is dramatically better than a condenser at twelve inches.
- Proximity effect as a feature. Working two to four inches away, the low frequency lift that would muddy a sung vocal produces exactly the chest and authority people associate with broadcast voice.
The category ladder is short and honest. The Rode PodMic at $85.00 is purpose built for this job and is the best value in the category. The Shure MV7+ at $299.00 has both USB and XLR outputs, which is genuinely useful because it lets you start on USB and move to an interface later without replacing the microphone. The Shure SM7B at $439.00 is the one you see on every podcast video, and the Electro-Voice RE20 at $449.00 is the broadcast alternative, with a variable-D design that reduces proximity effect so a presenter who moves stays tonally consistent.
Where should the microphone sit?
Four to six inches from the mouth, angled about 15 degrees off axis so you are speaking across the capsule rather than straight into it. That angle costs almost no level and removes most plosive energy, because a plosive is a blast of moving air rather than a loud sound and moving air is directional.
| Distance | Low end lift | Room heard | Sounds like |
|---|---|---|---|
| 1 to 2 in | +8 to +12 dB | Almost none | Late night radio, and easy to overdo |
| 3 to 4 in | +5 to +8 dB | Very little | The broadcast default |
| 6 in | +3 to +5 dB | A little | Natural and conversational |
| 12 in | +1 to +2 dB | Clearly audible | The room is now part of the sound |
| 24 in | None | Dominant | A recording of a bedroom with talking in it |
The practical consequence is that distance consistency matters more than any processing you can apply afterwards. A host who drifts between three and ten inches across an hour is producing a voice whose tone and level change constantly, and that is far harder to fix than it is to prevent. A boom arm at $139.99 is the practical answer: it holds a position, it takes no desk space, and it decouples the microphone from typing and footfall coming through the desk and floor.
How do you record two people in the same room?
One microphone each, on separate tracks, into one interface. Never share a microphone and never record both people into one merged file, because you lose the ability to fix one person's level, breathing, noise or timing without affecting the other.
Geometry does most of the isolation work and it is free. Point each microphone away from the other speaker so that person falls in its rear null, which on a cardioid is roughly 20 to 25 dB of rejection. The usual arrangement is the two hosts sitting at an angle to each other rather than directly opposite, with microphones crossing between them.
An interface with two real preamps is the right tool. The Focusrite Scarlett 2i2 at $224.99 covers two people. Three or four in the room needs the Scarlett 4i4 at $299.00, and at that point everyone also needs their own headphone feed, which a small headphone amplifier solves for the price of a cable.
Give every participant closed back headphones . It is not a luxury: without them, people talk over each other, nobody notices when a microphone has died, and open headphones spill the other person's voice back into your own microphone, which produces a hollow phasey quality that cannot be removed.
How do you record a remote guest?
Double-ender, every time. Each participant records their own microphone locally at full quality while the video call runs alongside purely for conversation. You then sync the local recordings in the edit.
The reason is not a preference for quality. Call audio is aggressively compressed, band limited and subject to packet loss, and none of that is recoverable, because the information was never transmitted. A local recording of the same person is full bandwidth and gapless.
- Everyone records locally, into whatever they have, at 48 kHz and 24 bit if possible.
- Everyone claps once, together, at the top of the session. That transient is your sync point.
- Record the call audio too, as a backup and as a safety net if a local file is lost.
- Align the local files to the claps in the edit, then mute the call track.
- Ask the guest to send the raw file, not an exported version, because a second lossy encode compounds.
Where each of those signals sits in the chain is laid out in home studio signal flow explained.
Two habits from doing this for real that cost nothing. First, record a thirty second test and actually listen back before the guest arrives, on the headphones you will be wearing. I have lost interviews to a cable that was fine until it was moved, and the test would have caught it. Second, leave the recording running through the small talk at the top and the goodbyes at the end. The genuinely good material is very often in the two minutes before someone thinks the session has started, and disk space costs nothing. The third habit is a stopped clock: I write the timecode down whenever anyone fluffs a line badly, because finding it later without a note takes ten times longer than writing it took.
What levels and settings should you record at?
Speech has a lower crest factor than music, which means the gap between average and peak is smaller, so you can aim for a slightly higher average without risking a clip.
- Sample rate 48 kHz, bit depth 24. 48 kHz is the video standard and podcasts frequently end up alongside video. There is no benefit above it for speech.
- Average around -18 dBFS with peaks at -10 to -6 dBFS. Set the gain from the loudest laugh, not from a level-check sentence, because people get louder once they relax.
- No compression while recording. It is not reversible and it is much easier to do well afterwards with the whole episode in front of you.
- A high pass filter at 80 Hz is reasonable for speech, since there is nothing below it in a human voice except rumble, handling noise and air conditioning.
The gain staging calculator turns those into meter positions, and how to set recording levels covers the method in full.
What does the edit and export chain look like?
Short and in this order. Doing it out of order is how a podcast ends up over-processed.
- Edit first. Cut before you process, always, since processing decisions made around a section you later delete are wasted work.
- Noise reduction, gently if at all. Broadband reduction pushed hard produces a watery artefact that listeners notice even when they cannot name it. A cleanly recorded track needs very little.
- Subtractive EQ. High pass at 80 Hz, then cut any specific resonance rather than boosting to compensate.
- Compression, 3:1 to 4:1 with 4 to 8 dB of reduction. Speech benefits from more consistency than music does, because listeners are often in cars and kitchens.
- De-ess if needed, targeted between 5 and 8 kHz. Close dynamic microphones and bright condensers both produce sibilance that a compressor makes worse.
- Loudness normalise to -16 LUFS integrated, stereo, true peak at -1 dBTP. Mono spoken word is commonly targeted at -19 LUFS instead, since a mono file is perceived louder at the same measured level.
| Delivery | Integrated loudness | True peak ceiling |
|---|---|---|
| Podcast, stereo | -16 LUFS | -1 dBTP |
| Podcast, mono | -19 LUFS | -1 dBTP |
| Streaming music platforms | -14 LUFS | -1 dBTP |
| Broadcast television, EBU R128 | -23 LUFS | -1 dBTP |
What about the room?
Treat it, do not try to soundproof it, and be clear about the difference. Treatment changes how the room sounds from the inside and is achievable in a rented bedroom. Soundproofing stops sound passing through walls, requires mass and decoupling, and is a construction project. Every listing that sells foam as soundproofing is selling the first thing under the name of the second.
For speech the return is concentrated in a small number of places: something absorbent directly behind you, something on the wall you face, and something soft underfoot if the floor is hard. Two 2 inch panels at the right two spots beat eight panels spread evenly around a room, because early reflections come from specific directions. A reflection filter behind the microphone is the portable version and it genuinely helps a spoken voice.
Read how to treat a room for recording and how to soundproof a home studio, which is mostly a page about why you probably cannot, and size the coverage with the acoustic treatment calculator.
Related
- Best dynamic microphones for untreated rooms: the right category for speech
- Audio interface versus USB microphone: when one box stops being enough
- Best microphone stands and boom arms: holding a consistent distance
- Microphone placement distance chart: every distance in one table
- Best headphones for tracking: closed back, one pair per person
- Gain staging calculator: the level to aim for
Frequently asked questions
What is the best microphone for a podcast at home?
A cardioid dynamic, not a condenser, and this is the opposite of the usual advice. A dynamic is less sensitive and more directional, so it captures far less of an untreated room, less keyboard noise and less of the person sitting opposite you. The Rode PodMic at about $85.00 and the Shure SM7B at $439.00 sit at either end of that category, and both beat a more expensive condenser in a normal room.
How loud should a podcast be?
Around -16 LUFS integrated for stereo and -19 LUFS for mono, with true peaks no higher than -1 dBTP. Those are the widely used podcast delivery targets and they exist so that episodes from different shows play back at a consistent level. Measure with a loudness meter rather than by eye, because peak meters tell you almost nothing about perceived loudness.
Do I need an audio interface for a podcast, or is USB enough?
A USB microphone is genuinely sufficient for one host recording alone. It becomes limiting the moment there are two people in the room, because most USB microphones cannot be sample-locked together and give you separate tracks reliably. Two guests in one room is the point where an interface with two preamps and one recording session becomes the simpler, cheaper answer.
How do I record two people in the same room?
Two cardioid dynamics, one per person, angled so each microphone points away from the other speaker, into one interface as two separate tracks. Never share a microphone, because you lose all independent control of level, noise and editing. Angle matters as much as distance: putting the other person in the rear null of your microphone is worth 20 dB or more of separation for free.
How do I make my podcast not sound like a bedroom?
Work closer, choose a dynamic microphone, and put something soft on the nearest hard surface. Getting from 12 inches to 4 inches raises the direct sound by about 10 dB relative to the room, which does more than any plugin. Soft furnishing behind and beside you handles the earliest reflections. Foam does not soundproof anything, so treat this as controlling how the room sounds, not what leaks in.
Should I record a remote guest locally or over the call?
Locally, always, and it is called a double-ender. Each participant records their own microphone to their own machine while the call runs alongside for conversation. Call audio is heavily compressed and drops out, and no processing recovers what a codec threw away. Have everyone clap once at the start so the tracks can be aligned, then sync in the edit.
Working out your own room and signal chain? The Home Studio Build Planner is the paid version of these pages: 8 printable worksheets you fill in with your own numbers, plus the full PDF, $29.