MiniMax H3 Open Weights | Try in Video Generator →

Sonilo Video-to-Music API

sonilo/

Sonilo Video-to-Music is a fast AI music generation model that creates background music synced to an input video’s mood, pacing, and scene transitions. Ready-to-use REST inference API for video soundtracks, social media content, advertising creatives, cinematic clips, product videos, creator workflows, and professional video-to-music generation with simple integration, no coldstarts, and affordable pricing.

audio-to-video
Input

Idle

$0.009per run·~111 / $1

Next:

ExamplesView all

Related Models

README

Sonilo Video-to-Music

Sonilo Video-to-Music generates music from a video input, with an optional style prompt to guide the soundtrack. It is designed for turning visual content into matching background music for short films, ads, social content, trailers, highlight clips, and other video-driven audio workflows.

Why Choose This?

  • Video-driven music generation
    Generate music that fits the pacing and feel of an uploaded video.

  • Optional style guidance
    Add a prompt to steer the mood, genre, instrumentation, or production style of the generated music.

  • Simple workflow
    Upload one video, optionally add a style prompt, and generate a matching music track.

  • Supports longer inputs
    Works with videos up to 360 seconds.

  • Production-ready API
    Suitable for trailers, branded content, social videos, cinematic edits, and background scoring workflows.

Parameters

ParameterRequiredDescription
videoYesInput video URL. Maximum supported video length is 360 seconds.
promptNoOptional style prompt for the generated music.

How to Use

  1. Upload your video — provide the source video you want to score with music.
  2. Add a style prompt (optional) — describe the mood, genre, instrumentation, or production feel you want.
  3. Submit — run the model and download the generated music.

Example Prompt

Cinematic emotional orchestral score with soft piano, warm strings, slow build, inspiring and modern trailer mood

Pricing

Pricing is based on the uploaded video duration.

Billing Rules

  • Pricing is $0.009 per billed second
  • Billing is based on the uploaded video duration
  • Billed duration is rounded up to the next whole second
  • Minimum billed duration is 1 second
  • Maximum billed duration is 360 seconds
  • prompt does not affect pricing

Example Costs

Video DurationCost
1s$0.009
5s$0.045
10s$0.090
30s$0.270
60s$0.540
120s$1.080
360s$3.240

Best Use Cases

  • Social media videos — Generate music beds for short-form clips.
  • Ads and promos — Create matching soundtrack material for branded content.
  • Trailers and highlights — Add cinematic or energetic music to visual edits.
  • Creator workflows — Quickly generate background music for uploaded video content.
  • Prototype scoring — Explore soundtrack directions before final post-production.

Pro Tips

  • Use a style prompt when you want stronger control over genre, mood, or instrumentation.
  • Keep the prompt focused and specific for more predictable results.
  • Shorter videos are useful for quickly testing soundtrack direction before scoring longer content.
  • Upload the cleanest final or near-final edit possible so the music better matches pacing and structure.

Notes

  • video is required.
  • Maximum supported video length is 360 seconds.
  • Pricing depends only on billed video duration.
  • Video duration is rounded up to the next whole second for billing.

Related Models

  • Sonilo audio generation workflows — Useful when you need prompt-first music generation instead of video-driven scoring.
  • Background music generation workflows — Useful when you need standalone music without a video input.
  • Video sound design workflows — Useful when you want synchronized effects instead of generated music.
Note:This website uses AI models provided by third parties.

Video To Music API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/sonilo/video-to-music with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Video To Music below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/sonilo/video-to-music" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL=$(printf '%s' "$TASK" | jq -r '.urls.get // empty')
if [ -z "$RESULT_URL" ]; then
  RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
fi

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    created|processing) sleep 2 ;;
    *) printf 'Unexpected status: %s
' "$STATUS" >&2; exit 1 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/sonilo/video-to-music";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = task.urls?.get ||
  `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout"].includes(result.status)) throw new Error(JSON.stringify(result));
  if (!["created", "processing"].includes(result.status)) throw new Error("Unexpected status: " + result.status);
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4"
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/sonilo/video-to-music", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = task.get("urls", {}).get("get") or f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout"}:
        raise RuntimeError(result)
    if status not in {"created", "processing"}:
        raise RuntimeError(f"Unexpected status: {status}")
    time.sleep(2)

Video To Music API — Frequently asked questions

What is the Video To Music API?

Video To Music is a Sonilo model for AI inference, exposed as a REST API on WaveSpeedAI. Sonilo Video-to-Music is a fast AI music generation model that creates background music synced to an input video’s mood, pacing, and scene transitions. Ready-to-use REST inference API for video soundtracks, social media content, advertising creatives, cinematic clips, product videos, creator workflows, and professional video-to-music generation with simple integration, no coldstarts, and affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Video To Music API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/sonilo/sonilo-video-to-music.

How much does Video To Music cost per run?

Video To Music starts at $0.009 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Video To Music accept?

Key inputs: `prompt`, `video`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/sonilo/sonilo-video-to-music.

How do I get started with the Video To Music API?

Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.

Can I use Video To Music outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (Sonilo). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.