Add music to video

Put a track under the cut and mix it

A clip on the Morphic timeline with room below it to add background music to video
00:00:00:00
Music bed at 45. Press play to hear it under the picture.
V1
Picture, no sound of its own
A1
Music bed

The bed is its own clip on its own track, so its level is a decision you can revisit.

Add background music to video by putting the track on its own layer and deciding what it does against everything already playing. Music, narration and effects each keep their own level, so scoring a cut is a mix you build. All of it happens in your browser, next to where the footage was made.

Audio inMP3, WAV, AAC, M4A and OGG all place.
Non-destructiveLevels live on the clip, not in the file.
Your footageKept in your library, not on a countdown.
Clean exportsOne MP4 up to 4K on any paid plan.

How to add music to a video in three steps

  1. 01

    Bring in the track

    Drop in music you have, or describe the piece you want and generate it here.

  2. 02

    Set where it sits

    Slide it under the picture, then take the level down until speech reads over it.

  3. 03

    Export the scored cut

    One MP4 at 4K or below, rendered on your machine with the mix already inside it.

What scoring a video here involves

Stacked audio layers are a Compose V2 capability, and V2 is switched on per workspace.

A layer for every sound

Music, voiceover and effects each get their own audio track. Every clip on every track carries its own level, and that is what makes a score sit under a voice instead of fighting it.

Music written to order

Describe a style and a mood, give it a length to aim for, and generate the piece in the same workspace as the cut. Turn vocals on and it sings, either your lyrics or ones it writes.

Levels that move with the scene

Gain is automatable like any other property, so a bed can fall away under a line of dialogue and come back after it. That dip is the difference between a video with music on it and a video that was scored.

Mix by asking

Bring the music down under the voiceover is an instruction Copilot can carry out across the whole timeline, reading the current mix state before it touches anything. Whatever it changes is one undo.

Crossfades between cues

Where one piece of music gives way to another, an equal-power crossfade holds the level steady through the join. The same joins carry picture transitions when you want them.

Out as one file

The mix renders into the MP4 itself. At the project quality or one you pick, locally, with nothing queued.

A score is easier when the picture has somewhere to breathe. When a cue needs a few seconds of something and the footage has none, the shot can be generated in the same workspace, from the same roster the rest of the project draws on, and laid over the bar where the music opens up. Seedance 2.5 is one of the video models available for it.

Add audio to video

Three different things go by this name, and the timeline treats all three the same way:

  • A voiceover is an audio clip.
  • A sound effect is an audio clip.
  • A song is an audio clip.

Each lands on its own track with its own level. That matters when they overlap: a narrator over a bed over room tone is three faders, and the balance between them is the only thing you will come back to. Flatten it early and every later change means returning to the source. Keep it separate and one cut carries a different mix for a client, a platform or a language.

Add a song to a video

A song brings a structure the picture has to respect, so generate it before you cut rather than hunting for a track that nearly fits.

  • Describe it plainly. Pop, gospel, something slow and warm. Name the mood.
  • Give it a length close to the section it is scoring.
  • Switch vocals on to have it sung, with lyrics you paste or lyrics it writes.

What comes back is an original piece with no named artist behind it, and it drops onto the timeline as a normal clip to trim, fade and move.

Background music that sits under narration

The craft of a music bed is level, and level is rarely one number. Speech does not hold a constant volume, so a bed set once is too loud under the quiet lines and inaudible under the rest. Two routes, same cost:

  • By hand — a point before the line, a lower point under it, a return after.
  • By instruction — Copilot reads the mix, runs the pass everywhere it belongs, and lands it as one undo.

Either way the music keeps its own clip and its own track, and the narration underneath is untouched. For a door or a footstep instead, the sound effects library is the other half of this. When the track is meant to replace what was recorded, remove sound from video first and mix into the space that leaves.

Related video tools

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

1100 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3625 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6350 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24650 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

Does adding music to a video cost anything?
Placing a track, moving it and setting its level are hand edits, and hand edits are not billed however long the mix takes. Two things do cost credits. Generating a piece of music is a generation like any other, and handing the mixing to Copilot is billed per command it carries out. An instruction that fails applies nothing and costs nothing. Exports run on your own machine today at no credit cost.
Where does the music come from?
Either a file you bring or a track generated in the same workspace. Generation takes a description of the style and mood and a length to aim for, and it can sing: turn vocals on and either paste your lyrics or leave the field empty and let the model write them. What it will not do is imitate a named artist or a real singer, because there is no voice to clone and describing one in words is as far as the control goes.
Can I keep the original sound and add music underneath?
Yes, and that is the normal case. A clip and its audio are separate items on separate tracks, so the dialogue keeps its own level while the music takes another. Set the music low enough to sit under speech and drop it further under the lines that matter, either by hand or by saying so. Nothing about adding music overwrites the sound already there.
Why is my video silent after I drag it in?
A video clip placed on its own carries no sound until its audio is on an audio track beside it. Ask Copilot to bring the clip in and it does that placement for you, which saves a confusing five minutes. Once both are on the timeline they move together and mix independently.
Does the exported video have a watermark?
Every export on the free plan does, and it is burned into the picture, leaving the sound untouched. Paid plans export clean. The music makes no difference to this either way.
What audio and video formats can I bring in?
MP3, WAV, AAC, M4A and OGG all place on an audio track. Video arrives as MP4, MOV, WebM, MKV, AVI or MPEG, and stills sit alongside both. Everything renders out as MP4 at up to 4K, so a scored cut leaves as one file, with no stem to reunite somewhere else.
How long can the music or the video be?
Nothing stops you scoring a long timeline while you work. The ceiling is at the export, where a long render is split into parallel chunks and an extreme one can be refused outright, so a feature-length score is not a plan to make yet. Generated music has a per-model range for how long a single piece can run, and the picker shows it when you ask.
Can I generate music that matches a scene I already cut?
Ask for it by length and by feel, the way you would brief a composer. The generator takes a duration target, so a cue can arrive at roughly the size of the hole and save you trimming it into one. It still lands on the timeline as an ordinary audio clip, so the last few frames are yours to nudge.

You might also like