Mix Voice and Background Music So Your Reel Is Easy to Hear
Balance narration and music with a practical manual-ducking plan, an automatic-ducking workflow and a phone-speaker listening checklist.
Reviewed September 15, 2026. Includes original teaching examples and source-backed instructions. Third-party features and terms can change.
If viewers have to turn up the volume to understand a sentence and then turn it down when the music swells, the mix needs work. Raising the voice alone may not solve it. The music could contain a competing vocal, a bright instrument or a sudden hit that covers the important words.
Build the mix around the explanation. Get the voice clear first, introduce music quietly, then shape its level around the spoken phrases. A background track is optional; understandable speech is not.
This guide supplies an original cue-sheet example and a listening procedure. It does not include playable before-and-after audio or measured loudness results. Use your own authorised voice recording and music, and evaluate the actual exported file.
Start with the voice alone
Mute music and effects. Listen to the whole spoken track. Fix major jumps between takes and check that the voice is not distorted. If a word is unclear because of noise or a damaged recording, solve that before building a soundtrack around it.
Use the natural-voice cleanup guide when appropriate, but avoid processing merely because a control exists. A clear original voice does not need heavy enhancement to deserve a place in a Reel.
Leave room for the combined tracks. A voice that already pushes the output into clipping may distort when music is added. Watch the output meter if your editor provides one and listen for rough, overloaded peaks. Meter labels and loudness tools vary by application; do not confuse a clip's gain slider with a universal final-output target.
Choose music that leaves space for speech
Start with a section that does not fight the rhythm of the sentence. A track with lyrics can compete with narration because the listener is hearing two sets of words. A dense melody or sharp percussion can also mask speech even when the meter looks modest.
Check that you are allowed to use the chosen recording for the intended post and any planned commercial use. Save the source and permission details with the project. Do not assume a file is cleared merely because it is downloadable or described as background music.
Try the explanation without music as a genuine option. W3C's low-background-audio guidance explains why background sound can make speech harder to distinguish. That page describes a specific accessibility criterion for audio-only content; it is not an Instagram mastering specification or a one-number certification for your Reel.
Establish a quiet starting balance
Play a representative section of narration. Bring the music up from silence until it supports the mood, then check whether every word remains effortless to understand. Do not begin at a copied percentage such as “voice 100, music 20.” Different files can have very different recorded levels.
Listen to the quietest sentence and the busiest part of the music, not just the loud opening line. A mix that works for the first five seconds can fail when the speaker becomes softer or the track adds another instrument.
If the music becomes hard to hear when the voice is clear, that may be acceptable. Its job here is to support speech. If the music needs to be a featured moment, give it a deliberate gap instead of forcing both elements to dominate simultaneously.
Manual ducking: make a small level curve
Ducking means lowering one sound while another needs attention. In an editor with volume keyframes, place a point before the spoken phrase, another where the music has reached its lower level, another near the end of the phrase and a final point after the music returns.
Those four points create a down-ramp, a quieter section and an up-ramp. Listen to the transition rather than using the same timing everywhere. A sudden drop may feel like an error; an overly slow drop may leave the first important word masked.
For an illustrative tutorial, a cue sheet could look like this:
| Moment | Voice job | Music decision |
|---|---|---|
| Opening title | Introduce the question | Start restrained; do not cover the first word. |
| Three-step explanation | Carry the main information | Keep music low and stable. |
| Brief demonstration pause | Let the action breathe | Raise music only if it does not hide useful natural sound. |
| Final instruction | Make the next step clear | Lower music before the instruction starts. |
| End frame | Finish the thought | Use a controlled ending, not a surprise volume jump. |
This describes an editorial intention. Record the actual timecodes and levels after working with your files.
Automatic ducking still needs review
In Adobe Premiere, the automatic-ducking instructions describe assigning audio types in Essential Sound, enabling ducking on music or ambience, choosing what to duck against and generating adjustable keyframes.
That can provide a starting point when your timeline contains clear dialogue and music tracks. Review the resulting envelope around short pauses, breaths, overlapping clips and the first word of each sentence. The software does not know which instruction matters most to your viewer.
If you edit the dialogue afterward, revisit or regenerate the ducking as appropriate. Keyframes created for an earlier arrangement may no longer align with the spoken phrases. Keep manual changes understandable so another editor can see why a section differs.
Test the mix where it will be heard
Listen on headphones, then a phone speaker at a comfortable level. Test a quieter playback level too. If the message becomes hard to understand before the music disappears, lower or simplify the background.
Listen once without looking at the picture. You should still understand the spoken explanation. Then watch with sound off and review captions; readable captions help, but they are not an excuse for an unnecessarily difficult soundtrack.
Ask another person which words they missed without telling them what to listen for. If they repeat the same unclear number or phrase, fix that location instead of raising the entire voice track.
Check the export, not just the timeline
Export a short representative sample and listen to the actual file. Confirm the selected audio tracks are present, nothing was accidentally muted and the ending is complete. A correct timeline does not prove the export settings produced the intended result.
Save the voice-only track, music source, permission record and final mix where your workflow allows. Keep a versioned cue sheet so later revisions do not become guesses.
The finish line is a Reel whose words are easy to hear and whose music feels intentional. Sometimes that means a carefully shaped background. Sometimes it means no music at all.
Voice-and-music cue sheet
Plan the soundtrack by spoken moment and verify the actual mix on headphones and a phone speaker.
Preview the complete template
VOICE + MUSIC MIX PLAN No preset percentage or measured mix is supplied. Project/version: Voice source: Voice-only clarity checked: Music source: Permission/licence record: Editor/version: Final export filename: CUE ROW Start/end time: Spoken words: Music section: Intended music role: Duck before first word: Lower-level section: Return transition: Natural sound to preserve: Actual setting/keyframes: Listening result: ILLUSTRATIVE CUES Opening question | restrained music; first word clear. Main three-step explanation | low, stable background. Demonstration pause | optional small rise; preserve useful action sound. Final instruction | reduce background before speech. End frame | controlled fade or appropriate ending. AUTO-DUCKING CHECK Correct audio types assigned: Dialogue drives intended music: First words protected: Breaths/short pauses do not cause distracting jumps: Changes reviewed after dialogue edits: FINAL LISTEN Headphones: Phone speaker: Comfortably low playback: Quietest sentence: Busiest music section: Output distortion/clipping: Actual export checked: Captions match speech: Voice-only version retained: One change still needed:
Free plain-text file. Edit it in any notes app or document. You may adapt our original worksheet for personal and client work; linked third-party assets keep their own licences.