Capability
Video to SFX
Generate sound effects that land on what happens in a video — footsteps, impacts, ambience, whooshes — as one full SFX bed or one clip per detected sound.
- Modes
- video_to_sfx — native payload
- standard — a full SFX bed
- plan → magic — one clip per sound
- Response
- Pricing
Native only — new capability
No compat route exists for Video to SFX. Call POST /v1/generations with the tool_id for model="v1", operation="video_to_sfx" (see Tool Catalog).
Modes
- standard — 1–4 takes of one continuous SFX bed for the whole window, each as long as the window.
- plan — detects the separate sounds in the window and quotes what generating them costs. Free; generates nothing.
- magic — one clip per detected sound, each timed to where it happens in the video. Call
planfirst, thenmagicwith itsplan_token.
Every time in the output is in milliseconds into the source video, so a clip can be placed on your timeline at its start_ms without any offset math.
video_to_sfx — native payload
POST/v1/generationstool_id for (model="v1", "video_to_sfx")
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| video | file input | Yes | — | The video to score, up to 1 GiB, as a URL, inline base64 or a Library file ref (see Providing audio/file input). A URL is the practical choice for video. |
| mode | string | No | "standard" | "standard", "plan" or "magic". |
| start_seconds | float | No | 0 | Where the scored window starts, in seconds into the video. |
| duration_seconds | float | No | to the end | How much to score. Up to 300 s per call; score a longer video in several windows. |
| takes | int | No | 1 | 1–4. Standard: takes of the bed. Plan/magic: takes per detected sound. |
| prompt | string | No | — | Standard only. Optional steer, up to 5000 characters. |
| negative_prompt | string | No | — | Standard only. Sounds to avoid (may be ignored). |
| seed | int | No | — | Standard only. 0–4294967295. |
| keep_dialogue | bool | No | false | Standard only. Keeps the clip's own dialogue and places effects around it. Turned off, with a note, when the video has no audio. |
| plan_token | string | No | — | Magic only. The plan_token a plan call on the same video returned. |
| budget_tokens | int | No | 0 | The most you agree to pay, in cents. Standard: ceil(seconds × takes × 3), plus seconds × 2.4 with keep_dialogue. Magic: the plan's quote_tokens. Plan: 0. |
budget_tokens is a ceiling, and it is enforced
The balance hold is exactly budget_tokens. A budget below the price is refused before anything is generated, and the error message names the exact price to send. You are charged only for what was delivered — a take or sound that doesn’t come back isn’t charged — and never more than budget_tokens.
standard — a full SFX bed
Two takes of a 20 s window: budget_tokens = 20 × 2 × 3 = 120 ($1.20).
curl -X POST https://apiv2.soundverse.ai/v1/generations \
-H "Authorization: Bearer sksoundverse_..." \
-H "Content-Type: application/json" \
-H "Idempotency-Key: vsfx-req-001" \
-d '{
"tool_id": "<video_to_sfx tool_id>",
"payload_json": "{\"video\": \"https://example.com/clip.mp4\", \"mode\": \"standard\", \"start_seconds\": 5, \"duration_seconds\": 20, \"takes\": 2, \"budget_tokens\": 120}"
}'plan → magic — one clip per sound
- Call with
mode: "plan"andbudget_tokens: 0. The completed output lists the detectedsoundsand returnsquote_tokens(the magic price, in cents), a signedplan_tokenandplan_expires_at. - Before
plan_expires_at, call withmode: "magic", the samevideo, window andtakes, thatplan_token, andbudget_tokens=quote_tokens. Magic generates exactly the sounds the plan listed.
import json, time, requests
API = "https://apiv2.soundverse.ai/v1"
HEADERS = {"Authorization": "Bearer sksoundverse_..."}
TOOL_ID = "<video_to_sfx tool_id>"
video = {"video": "https://example.com/clip.mp4", "start_seconds": 0, "duration_seconds": 30}
def run(payload, key):
task = requests.post(f"{API}/generations", headers={**HEADERS, "Idempotency-Key": key},
json={"tool_id": TOOL_ID, "payload_json": json.dumps(payload)}).json()
while True:
t = requests.get(f"{API}/generations/{task['task_id']}", headers=HEADERS).json()
if t["status"] in ("completed", "failed"):
return t
time.sleep(3)
plan = run({**video, "mode": "plan", "budget_tokens": 0}, "vsfx-plan-001")["output"]
print(len(plan["sounds"]), "sounds, magic costs", plan["quote_tokens"], "cents")
magic = run({**video, "mode": "magic", "plan_token": plan["plan_token"],
"budget_tokens": plan["quote_tokens"]}, "vsfx-magic-001")["output"]A plan_token is bound to the API key’s account and to the API: it can’t be spent by another account or in the Soundverse app. Past plan_expires_at, run plan again.
Response
Once the task completes, its output has this shape (magic shown):
{
"text": "Generated 2 sounds (2 files).",
"mode": "magic",
"model": "v1",
"window_start_ms": 0,
"window_duration_ms": 30000,
"source_duration_ms": 42000,
"tokens_charged": 27,
"notes": [],
"takes": [],
"sounds": [
{"sound_index": 0, "label": "door slam", "type": "sfx", "category": "foreground",
"start_ms": 2000, "end_ms": 5000, "sound_start_ms": 2500, "sound_end_ms": 3100,
"gain_db": -3.0, "credits": 30, "asset_indices": [0], "file_ids": ["018f..."]},
{"sound_index": 1, "label": "room tone", "type": "sfx", "category": "background",
"start_ms": 0, "end_ms": 6000, "sound_start_ms": 0, "sound_end_ms": 6000,
"gain_db": -12.0, "credits": 60, "asset_indices": [1], "file_ids": ["018f..."]}
]
}- Each file is FLAC and is also in
output.assets. Place a sound’s file at itsstart_ms;sound_start_ms/sound_end_msmark the sound itself inside it, andgain_dbis the suggested mix level. - standard fills
takesinstead: one entry per delivered take with itsfile_id,duration_msandgaps([start_ms, end_ms]ranges where that take is silent). - plan fills
soundswithout files, plusplan_token,quote_tokensandplan_expires_at. notesexplains ignored fields and automatic adjustments, one sentence each.
Pricing
| Mode | Price | Counted on |
|---|---|---|
standard | $0.03 / s / take | Seconds of video scored × takes. |
keep_dialogue | +$0.024 / s | Standard only. Once per call, however many takes. |
magic | $0.03 / s / take | Seconds of each detected sound × takes. plan returns the exact quote. |
plan | Free | Detects the sounds and quotes magic. Generates nothing. |
USD, the same on every license tier, rounded up to the whole cent per call. Rate limit: 20 calls an hour, 100 a day.