MiniMax H3 Open Weights | Try in Video Generator →

Music Video Generator | AI Digital Human API

wavespeed-ai/

AI Music Video Generator transforms audio + a single photo into a full music video with cinematic camera angles, smooth transitions, and perfect lip sync. Up to 10 minutes, 480p or 720p. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

audio-to-video
Input

Idle

$0.15per run·~66 / $10

Next:

ExamplesView all

Related Models

README

AI Music Video (MV) Generator

The world's best AI music video (MV) generator. Turn any song + a single photo into a professional-quality music video in minutes.

Why It's the Best

  • Blazing fast: Generate a full 1-minute music video in just a few minutes. No waiting hours.
  • Perfect lip sync: Vocal-aware segmentation ensures the singer's lips match the audio precisely throughout the entire video.
  • Cinematic quality: AI director plans each scene with different camera angles, compositions, and natural lighting — like a real music video shoot.
  • One photo is all you need: Upload a single portrait and the AI handles the rest — scene creation, angle variations, and smooth transitions.
  • Up to 10 minutes: Create full-length music videos, not just short clips.
  • Smart scene planning: Automatically detects vocal phrases and silence in the audio to create natural scene transitions at musically meaningful moments.

How It Works

  1. Upload your audio — any song, any genre, up to 10 minutes.
  2. Upload 1-3 reference images (optional) — the person who will appear in the video.
  3. Describe the scene (optional) — e.g. "A woman sings in a forest while playing a guitar".
  4. Choose aspect ratio — 16:9 (landscape) or 9:16 (portrait/vertical).
  5. Select resolution — 480p or 720p.
  6. Get your music video — fully rendered with transitions, multiple angles, and synced audio.

What Happens Behind the Scenes

  1. Vocal isolation — Separates vocals from instruments to analyze singing patterns.
  2. Smart segmentation — Splits the audio at natural phrase boundaries (not arbitrary fixed intervals).
  3. AI directing — A vision-language model plans each scene: camera angles, compositions, expressions, and camera movements.
  4. Scene generation — Creates unique starting frames for each segment from different angles.
  5. Video synthesis — Generates lip-synced digital human video for each segment.
  6. Cinematic assembly — Smooth crossfade transitions between scenes, with the original audio layered on top for perfect sync.

Pricing

Output ResolutionCost per 5 secondsMax Length
480p$0.1510 minutes
720p$0.3010 minutes

Billing Rules

  • Standard Rate: $0.03 per second
  • HD (720p) Rate: $0.06 per second
  • Minimum Charge: 5 seconds ($0.15 minimum)
  • Billing Cap: 600 seconds (10 minutes)

Parameters

ParameterRequiredDescription
audioYesURL of the audio/music file
imagesNoArray of 1-3 reference image URLs
promptNoScene/style description
aspect_ratioNo"16:9" or "9:16" (auto if omitted)
resolutionNo"480p" (default) or "720p"

Tips

  • Best results with vocals: The AI uses vocal patterns for scene timing. Songs with clear vocals produce the best-timed transitions.
  • Portrait photos work best: Clear, front-facing photos with visible face give the best identity preservation.
  • Be descriptive: A good prompt like "A rock singer performing on a neon-lit stage" gives much better results than just "singer".
  • No photo? No problem: If you don't provide images, the AI will generate a performer based on the detected voice (male/female).

Note

  • Max audio length: 10 minutes (600 seconds)
  • Processing speed: A 1-minute music video typically completes in 3-6 minutes
  • Supported audio formats: MP3, WAV, AAC, and most common formats
  • The AI automatically handles scene planning, you don't need to specify individual scenes
Note:This website uses AI models provided by third parties.

Music Video Generator API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/wavespeed-ai/music-video-generator with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Music Video Generator below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
    "aspect_ratio": "16:9",
    "resolution": "480p"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/music-video-generator" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL=$(printf '%s' "$TASK" | jq -r '.urls.get // empty')
if [ -z "$RESULT_URL" ]; then
  RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
fi

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    created|processing) sleep 2 ;;
    *) printf 'Unexpected status: %s
' "$STATUS" >&2; exit 1 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/wavespeed-ai/music-video-generator";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
        "aspect_ratio": "16:9",
        "resolution": "480p"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = task.urls?.get ||
  `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout"].includes(result.status)) throw new Error(JSON.stringify(result));
  if (!["created", "processing"].includes(result.status)) throw new Error("Unexpected status: " + result.status);
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
    "aspect_ratio": "16:9",
    "resolution": "480p"
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/wavespeed-ai/music-video-generator", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = task.get("urls", {}).get("get") or f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout"}:
        raise RuntimeError(result)
    if status not in {"created", "processing"}:
        raise RuntimeError(f"Unexpected status: {status}")
    time.sleep(2)

Music Video Generator API — Frequently asked questions

What is the Music Video Generator API?

Music Video Generator is a WaveSpeedAI model for AI inference, exposed as a REST API on WaveSpeedAI. AI Music Video Generator transforms audio + a single photo into a full music video with cinematic camera angles, smooth transitions, and perfect lip sync. Up to 10 minutes, 480p or 720p. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Music Video Generator API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/wavespeed-ai/music-video-generator.

How much does Music Video Generator cost per run?

Music Video Generator starts at $0.15 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Music Video Generator accept?

Key inputs: `prompt`, `images`, `audio`, `aspect_ratio`, `resolution`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/wavespeed-ai/music-video-generator.

How do I get started with the Music Video Generator API?

Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.

Can I use Music Video Generator outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (WaveSpeedAI). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.

Music Video Generator | AI Digital Human API on WaveSpeedAI