Skip to content

Latest commit

 

History

History
3231 lines (2751 loc) · 169 KB

File metadata and controls

3231 lines (2751 loc) · 169 KB

Task Commands

Tasks are utility operations that run outside of a pipeline step. Use them for preprocessing, assembly, data gathering, and analysis - most are plain code, and some run a small helper model (an upscaler, a face detector, a stem separator).

{
    "name": "step_name",
    "task": {
        "command": "command_name",
        "arguments": { ... }
    },
    "result": { "content_type": "image/jpeg" }
}

Any task that runs a model accepts a "device" argument to pin where it runs - useful for keeping a helper model (a captioner, an upscaler) off the accelerator a loaded pipeline is using, or on a second one. "device" is listed on every task's schema, since it is always safe to pass: a task that runs no model - slice_audio, compose_text, and the like - just ignores it.

Task argument schemas are discoverable: GET /api/tasks/{command} on the server returns each command's arguments read from its registered implementation's real signature, the web editor builds task forms from them, and workflow validation flags task-argument typos the same way it flags pipeline ones.

A signature carries no domain, though, so the numbers whose domain is not a judgement call are declared separately (dw/task_domains.py) and validation reports one outside it as an error at its JSON path: a count of frames or seconds to cut, and a sample rate or frame rate, have to be above zero, and an offset to start at zero or above. Those are refused rather than interpreted - num_frames: -10 used to answer with the track minus its last ten frames and target_sample_rate: 0 with the original samples under a 44100 Hz header, both reported as clean successes. The commands refuse the same values at run time, which is what catches one that arrived from a variable: or an earlier step rather than being written in the file.

A numeric argument is read the same way by every command and by validation (whole_number and real_number in dw/task_domains.py, the only numeric coercion in dw/tasks/). A number may arrive as a JSON number or as a numeric string - a variable: resolved from the command line is text - so 3, 3.0, "3" and "3.0" are all the whole number 3, and "23.976" is a frame rate. A whole-number argument (a frame count, an index, a sample rate, a pixel size) refuses a fractional value such as 3.5 or "22050.5" rather than truncating it, naming the argument and saying "whole number". Every numeric argument refuses true/false, text that is no number ("abc") and an infinite or NaN value. What validation refuses, the run refuses with the same sentence, and the reverse. Task.run reads every declared numeric argument before the command sees it (coerce_arguments), so a handler is handed a number, never the string it arrived as: exposure: "0.5" grades exactly as 0.5 does. An argument whose only rule is being a number, such as grade's exposure or a gain in dB, declares the finite domain, so validation reads it too.

Adding a task

A task is one function, registered with @register_command in dw/tasks/registry.py. Its signature is its argument schema, and everything validation knows about it beyond the signature is declared on the same decorator, so a new task cannot be half-registered (#692):

from .registry import register_command
from ..task_domains import NON_NEGATIVE, POSITIVE, my_task_errors


@register_command(
    "my_task",
    domains={"num_frames": POSITIVE, "start_frame": NON_NEGATIVE},
    choices={"mode": ("fast", "best")},
    static_check=my_task_errors,
    media_arguments=("source",),
)
def my_task(source, num_frames, start_frame=0, mode="fast"): ...
  • domains - each numeric argument's domain (dw/task_domains.py's POSITIVE, NON_NEGATIVE, ...): refused at validation at steps[i].task.arguments.<name>, and again at run time when the task calls check_arguments.
  • choices - each argument's literal choices; get_task lists them.
  • static_check - a function of the step's arguments that returns (argument, message) pairs for rules across arguments. It lives in dw/task_domains.py, named *_errors, and is registered to exactly one command - or, when it needs the task's own parser, in the task's module behind an import on use, as attribute_voices' does.
  • whole_numbers - the numeric arguments that take only a whole number; the rest of domains take any real number. Validation refuses 3.5 for one of these, as the task's own whole_number call does at run time.
  • media_arguments - an argument that reads a file under a name that does not say so (source, media, lut, clip, track). Validation confines it as it does image or *_video; an argument left out of it is not path-checked.

TASK_ARGUMENT_DOMAINS, TASK_ARGUMENT_CHOICES, TASK_STATIC_CHECKS, TASK_WHOLE_NUMBER_ARGUMENTS and TASK_MEDIA_ARGUMENTS are read-only views of these declarations; tests/test_task_registry_rules.py pins them to the registry and every name in them to the command's signature.

A new task is:

  1. a module in dw/tasks/ holding the function and its decorator;
  2. one import line for that module in dw/tasks/task.py - the registry is filled by importing dw.tasks.task, and modules are not discovered automatically, so a module it does not import registers nothing;
  3. a section in this file.

Image Processing

ControlNet Preprocessors

Generate control images for ControlNet pipelines:

Command Description
canny Canny edge detection
canny_cv OpenCV Canny (alternative)
depth Depth estimation (DPT)
midas Monocular depth (MiDaS)
zoe Zoe depth estimation
zoe_depth Zoe depth with colorization
leres Relative depth (LeReS)
normal_bae Surface normal estimation
openpose Pose estimation
dw_pose DW pose estimation
mlsd Line segment detection
lineart Line art extraction
lineart_standard Standard line art
hed HED edge detection
scribble Scribble-style edges
pidi Boundary detection
shuffle Content-preserving shuffle
teed TEED edge detection
anyline Anyline edge detection
sam Segment Anything
segmentation Semantic segmentation
depth_estimator Depth hint generation
depth_estimator_tensor Depth hint as tensor

All accept an image argument with processing parameters:

{
    "task": {
        "command": "canny",
        "arguments": {
            "image": {
                "location": "https://example.com/photo.jpg",
                "low_threshold": 50,
                "high_threshold": 200,
                "detect_resolution": 1024,
                "image_resolution": 1024
            }
        }
    }
}

Image Manipulation

Command Description Extra Arguments
remove_background Remove image background
resize_center_crop Crop to a centered square, then stretch to width x height - distorts a non-square target width, height
resize_resample Resample to nearest 64px multiple
resize_rescale Resize to exact dimensions width, height
resize_bucket Snap to closest model-native aspect ratio resolution, ratios, alignment
crop_square Center crop to square
recenter_crop Re-frame around a chosen point at a chosen scale, so a series of images registers on one feature; the window may run off the source center_x, center_y, crop, width, height, fill
add_border_and_mask Add border with alpha mask
add_border_and_mask_with_size Border with specific dimensions width, height
strip_exif Remove all EXIF/metadata from image
add_watermark Add visible text watermark text, position, opacity, font_size, color, margin
get_image_size Return {width, height} dict

EXIF Stripping

Remove all EXIF metadata, GPS coordinates, camera info, and timestamps from images for privacy-safe preprocessing:

{
    "task": {
        "command": "strip_exif",
        "arguments": {
            "image": "previous_result:input_image"
        }
    },
    "result": { "content_type": "image/png" }
}

Returns a clean copy with pixel data only — no embedded metadata. Useful as a first step when processing user-uploaded images.

Watermark Embedding

Add a visible text watermark to images for responsible AI compliance:

{
    "task": {
        "command": "add_watermark",
        "arguments": {
            "image": "previous_result:generate",
            "text": "AI Generated",
            "position": "bottom-right",
            "opacity": 128
        }
    },
    "result": { "content_type": "image/png" }
}
Argument Required Description
text No Watermark text (default: "AI Generated")
position No "bottom-right", "bottom-left", "top-right", "top-left", or "center" (default: "bottom-right")
opacity No Text opacity 0-255 (default: 128)
font_size No Font size in pixels, 0 = auto-scale ~3% of image height (default: 0)
color No RGB array for text color (default: white)
margin No Pixel margin from edges (default: 10)

Aspect Ratio Bucketing

The resize_bucket command snaps an image to the closest model-native aspect ratio, then resizes with 64-pixel alignment. This avoids distortion and ensures the model generates at a resolution it was trained on.

{
    "task": {
        "command": "resize_bucket",
        "arguments": {
            "image": "previous_result:input_image",
            "resolution": 1024
        }
    },
    "result": { "content_type": "image/png" }
}
Argument Required Description
resolution No Target short-side size in pixels (default: 1024)
ratios No Custom list of [w, h] ratio pairs (default: standard SDXL/Flux ratios)
alignment No Round dimensions to this multiple (default: 64)

Default ratios: 1:1, 4:3, 3:4, 3:2, 2:3, 16:9, 9:16, 21:9, 9:21

For example, a 1600x900 photo (16:9) at resolution 1024 becomes 1792x1024. A 800x600 photo (4:3) becomes 1344x1024.

grade

grade is the finishing pass: exposure, contrast, tonal range, local contrast, white balance, colour and a matte or vignette look, on CPU, for an image or a video. A video is graded one frame at a time and keeps its frame count, frame rate and audio. Unlike the other image commands, its input is named media, and a file path or asset:/output: reference to a video is accepted directly. An alpha channel passes through untouched.

{
    "task": {
        "command": "grade",
        "arguments": {
            "media": "previous_result:generate_video",
            "highlights": -0.3,
            "shadows": 0.2,
            "clarity": 0.25,
            "fade": 0.15,
            "vignette": 0.3
        }
    },
    "result": { "content_type": "video/mp4" }
}

Every argument is optional and defaults to identity, so a step with none returns the input pixels. The step's job events carry one "applied" log naming every argument that is not at its identity value.

Argument Range Identity Effect
exposure any 0.0 Stops to brighten (positive) or darken (negative): a multiply by 2^exposure
contrast 0 or above 1.0 Multiplier around mid grey; below 1 flattens
whites -1.0..1.0 0.0 Moves the white point: the top of the curve up or down, black held
blacks -1.0..1.0 0.0 Moves the black point: the bottom of the curve up (lift) or down (crush), white held
highlights -1.0..1.0 0.0 Lifts or pulls down tones above mid grey, by a smooth luma mask; tones at and below mid grey are untouched
shadows -1.0..1.0 0.0 The same for tones below mid grey; tones at and above mid grey are untouched
clarity -1.0..1.0 0.0 Adds (positive) or removes (negative) medium-detail local contrast in the midtones
temperature -1.0..1.0 0.0 Warmer (toward red) or cooler (toward blue)
tint -1.0..1.0 0.0 Toward magenta (positive) or green (negative)
saturation 0 or above 1.0 Multiplier around each pixel's luma; 0 is greyscale
fade 0.0..1.0 0.0 Lifts the black floor and flattens the shadows toward it, white kept: a matte look
vignette -1.0..1.0 0.0 Darkens (positive) or lightens (negative) toward the corners; the centre is untouched

Order of operations: exposure, contrast, whites/blacks, highlights/shadows, clarity, temperature/tint, saturation, fade, vignette. The order is fixed whatever order the arguments are written in: fade comes after the tonal controls so blacks cannot pull its floor back down, and vignette is last so it darkens the finished picture.

A value outside its range is refused by validate_workflow, naming the argument and the range, and again at run time when it arrives from a variable: or an earlier step. get_task("grade") reports each range.

sharpen

sharpen is an unsharp mask, on CPU, for an image or a video: the image is blurred with a Gaussian, and the difference between the image and its blur (the detail) is added back, scaled by amount. Like grade, its input is named media, and a file path or asset:/output: reference to a video is read with its audio. A video is sharpened one frame at a time and keeps its frame count, frame rate and audio. An alpha channel passes through untouched.

{
    "task": {
        "command": "sharpen",
        "arguments": {
            "media": "previous_result:generate_video",
            "amount": 0.8,
            "radius": 1.5,
            "threshold": 3
        }
    },
    "result": { "content_type": "video/mp4" }
}
Argument Range Default Effect
amount 0 or above 1.0 How much of the detail is added back; 0 is identity. Applied in whole percent, so it is rounded to 0.01
radius above zero 2.0 The Gaussian blur's radius in pixels: the scale of the detail that is sharpened
threshold 0..255 0 The smallest difference, in channel levels, between a pixel and its blur that gets sharpened; smaller differences are left alone. A whole number of levels: a fractional value is rounded

A value outside its range is refused by validate_workflow, naming the argument and the range, and again at run time when it arrives from a variable: or an earlier step. get_task("sharpen") reports each range.

film_grain

film_grain adds seeded film grain, on CPU, to an image or a video. Its input is media, as in grade and sharpen: a video is processed one frame at a time and keeps its frame count, frame rate and audio, and an alpha channel passes through untouched. The grain is Gaussian noise, strongest in the midtones and weaker toward black and white.

{
    "task": {
        "command": "film_grain",
        "arguments": {
            "media": "previous_result:generate_video",
            "amount": 0.15,
            "size": 1.5,
            "chroma": 0.2
        }
    },
    "result": { "content_type": "video/mp4" }
}
Argument Range Default Effect
amount 0.0..1.0 0.1 Grain strength; 0 is identity
size 1 or above 1.0 Grain size in pixels: the noise is generated at 1/size resolution and upsampled
chroma 0.0..1.0 0.0 0 puts the same grain on every channel (brightness only, hue unchanged); 1 draws independent grain per channel; values between mix the two
seed a whole number, 0 or above the workflow's or step's seed Seeds the grain

Seeding: one generator is made per step from the seed and consumed frame by frame, so every frame of a video gets different grain and the same seed reproduces the whole output byte for byte. A workflow run always has a seed (a random one is drawn when the workflow names none, and recorded in the run's manifest), so rerunning a job reproduces its grain. An explicit seed argument overrides the workflow's or step's seed. The step's job events carry a log naming the seed used.

A value outside its range is refused by validate_workflow, naming the argument and the range, and again at run time when it arrives from a variable: or an earlier step. get_task("film_grain") reports each range.

apply_lut

apply_lut colours an image or a video through a 3D lookup table (LUT), read from a .cube file or built from a palette, on CPU. Its input is media, as in grade: a video is processed one frame at a time and keeps its frame count, frame rate and audio, and an alpha channel passes through untouched. Each pixel's colour is looked up with trilinear interpolation, and the result is blended with the original by strength.

{
    "task": {
        "command": "apply_lut",
        "arguments": {
            "media": "previous_result:generate_video",
            "lut": "asset:looks/teal-orange.cube",
            "strength": 0.7
        }
    },
    "result": { "content_type": "video/mp4" }
}
Argument Range Default Effect
lut a .cube file none The LUT: an asset: reference (upload one with upload_asset, or POST /api/uploads), an output: reference, or a path inside the workflow's directory, the asset libraries or the output root
palette 2..16 #rrggbb colours none Instead of a .cube, a look built from colours ordered dark to light (below)
strength 0.0..1.0 1.0 0 returns the original, 1 the full LUT result; values between blend the two

Exactly one of lut or palette is given; both or neither refuses, naming the two.

A palette. palette is a list of 2 to 16 #rrggbb colours, ordered dark to light, spaced evenly along the luminance axis: the first colour sits at black, the last at white. Each pixel keeps its own luminance (Rec. 709 luma) and takes its hue and chroma from the palette at that luminance, interpolated between neighbouring colours, so shadows lean toward the first colour and highlights toward the last. Near black and white the chroma shrinks as far as it must to stay in range. The 33³ table is built in memory, never written to disk, and applied exactly as a .cube is (same interpolation, same strength blend). The same palette always gives the same result, so one palette passed as a variable: to every shot of a series gives them one look:

{
    "variables": { "look": ["#102030", "#e0c090"] },
    "steps": [
        {
            "name": "look",
            "task": {
                "command": "apply_lut",
                "arguments": {
                    "media": "previous_result:generate_video",
                    "palette": "variable:look",
                    "strength": 0.6
                }
            },
            "result": { "content_type": "video/mp4" }
        }
    ]
}

A palette that is not a list, a colour that is not #rrggbb, or fewer than 2 or more than 16 colours refuses, naming the bad entry - at validate_workflow, and again at run time for a palette that arrives from an earlier step.

Where the file may be. A literal path is held to the same roots as any media argument, and must end in .cube; anything else is refused, naming only what the workflow wrote. The file is read and parsed once per step, not once per frame.

The .cube the parser takes is the 3D subset of the format, read strictly. Anything else refuses the step, naming the file and the line:

  • at most 16 MiB of UTF-8 text;
  • only TITLE, LUT_3D_SIZE, DOMAIN_MIN, DOMAIN_MAX, # comment lines, blank lines and data rows. LUT_1D_SIZE (a 1D LUT) and every other keyword refuse;
  • every keyword before the data, and LUT_3D_SIZE exactly once, 2..65;
  • a domain of exactly 0 0 0 to 1 1 1 (the default when the file names none);
  • exactly size³ data rows of three finite numbers, each in 0..1, listed red fastest, then green, then blue, as the format specifies.

Some grading applications write values slightly outside 0..1 or a wider domain; those files are refused rather than clipped. strength outside 0..1 is refused by validate_workflow and again at run time when it arrives from a variable: or an earlier step.

A finishing chain

The three finishing commands chain through previous_result:, each taking the last one's output as its media. Grade first, sharpen the graded picture, and add grain last so it is not itself sharpened:

{
    "steps": [
        {
            "name": "graded",
            "task": {
                "command": "grade",
                "arguments": { "media": "previous_result:generate_video", "fade": 0.15, "vignette": 0.3 }
            },
            "result": { "content_type": "video/mp4" }
        },
        {
            "name": "sharpened",
            "task": {
                "command": "sharpen",
                "arguments": { "media": "previous_result:graded", "amount": 0.6 }
            },
            "result": { "content_type": "video/mp4" }
        },
        {
            "name": "grainy",
            "task": {
                "command": "film_grain",
                "arguments": { "media": "previous_result:sharpened", "amount": 0.12, "size": 1.5 }
            },
            "result": { "content_type": "video/mp4" }
        }
    ]
}

Video Processing

Command Description Extra Arguments
get_first_frame Extract first video frame
get_last_frame Extract last video frame
get_frame Extract frame at index frame_index
window_video One overlapping, fixed-length window of a long video, with its audio index, num_frames, overlap, fps - see window_video
join_windows Blend processed windows back into one video the source's length, with the source's audio videos, source, num_frames, overlap, curve, fps - see join_windows
fit_to_model Fit a video into a model's working size and exact frame count; returns {video, fit} width, height, num_frames, mode (letterbox, stretch, crop), downscale - see fit_to_model and restore_to_source
restore_to_source Put a model's output back at its source's size and length, from the fit record fit - see fit_to_model and restore_to_source

The frame commands accept videos in any shape a result carries them: PIL frame lists, numpy or torch frame arrays, and audio+video pairs (LTX-2, MiniMax H3). The extracted frame is always a PIL image.

concat_videos

Concatenate videos - and the audio generated with them - into one video. The standalone counterpart of a chained pipeline step's stitching (see "Chained video generation" in the workflow guide):

{
    "task": {
        "command": "concat_videos",
        "arguments": {
            "videos": ["previous_result:shot_1", "previous_result:shot_2"],
            "trim_frames": 1,
            "crossfade_ms": 75,
            "fps": 24
        }
    },
    "result": { "content_type": "video/mp4", "fps": 24 }
}
Argument Required Description
videos Yes The videos to join, in order - previous_result references, or the path or URL of a video file an earlier run wrote, which is read with the audio muxed into it; an entry may also be a {"location": ...} dict wrapping either. Every video must be the same frame size - unlike a sample-rate mismatch, this task does not resize one for you, so a statically-resolvable (asset:/output:/literal path, wrapped in a {"location": ...} dict or not) size disagreement is refused at validate; a previous_result: or other reference not yet resolved still fails only at run time (#504, #518). To fit the odd video: video_frames to get its frames, resize_rescale to the target size (resize_center_crop squares the frame first and then stretches it, distorting a non-square target), then pair_audio(fit="video") to put its soundtrack back before passing it here (#551)
trim_frames No Frames dropped from the head of every video after the first (default: 0)
crossfade_ms No Equal-power crossfade at each audio seam, drawn from the trimmed material - no effect when trim_frames is 0, and validation warns when one is written there (default: 75)
audio_bleed_ms No How long the outgoing video's tail rings on over the head of the next one, at seams with nothing trimmed to crossfade (default: 0, off)
audio_bleed_gain_db No Gain applied to the bled tail before it is added, in dB - negative ducks a tail that would otherwise push the seam over 0 dBFS (default: 0, unchanged)
seam_fade_ms No Fade on each side of a seam that gets neither a crossfade nor a bleed - for tonal material, not for a continuous bed (default: 3, just enough not to click)
fps No Frame rate of the videos - required to join audio when trimming, and the rate the joined file is written at unless result.fps overrides it
match_levels No Even the shots' loudness out before joining - "rms" for perceived level (the measurement get_gallery_metadata reports as mean_dbfs), "peak" for the loudest sample. Off by default
match_levels_dbfs No The level match_levels moves every shot to (default: -1 dBFS for peak, -20 dBFS for rms). A shot that would clip at the target is held at -0.5 dBFS peak instead, reported as a match_levels_held warning with a per-shot log event

A video may also be named by path or URL, which is how shots an earlier run already wrote are joined without regenerating them - the file is read with the audio muxed into it, and its track is fitted to the frames' own duration so the codec's block padding does not walk the sound off the picture over a dozen seams. A shot generated in memory through a previous_result: chain gets the same fit, applied where the file is written rather than where it is decoded, so per-shot drift does not accumulate across a cut the way it once did:

{
    "task": {
        "command": "concat_videos",
        "arguments": {
            "videos": [
                "/path/to/outputs/shot_01.mp4",
                "/path/to/outputs/shot_02.mp4",
                "previous_result:shot_03_rerendered"
            ],
            "trim_frames": 0,
            "fps": 24
        }
    },
    "result": { "content_type": "video/mp4", "fps": 24 }
}

Give each video its own entry. One previous_result reference naming a step that produced several videos does not hand them all over at once - it fans the step out over them, one concatenation per video, which is what makes the list form above the way to join a run's shots.

trim_frames and audio_bleed_ms address opposite situations. A chain carries its keyframe forward, so the trimmed head is material that covers the same stretch of time as the outgoing tail and the two can be crossfaded. A cut generates each shot independently, so there is nothing to fade with - and generated shots tend to open on near-silence and end mid-sound, leaving a butt-join that drops a running laugh track or a ringing room into a hole. audio_bleed_ms fills it the way an audience carries across a picture cut: a decaying copy of the outgoing tail is laid over the incoming head, added to whatever is already there, shortening neither side. Reach for more than the gap looks like it needs: a shot's head is silent for longer than the picture suggests, and the bleed has to outlast it. Measured on a five-shot H3 sitcom cut, 700 ms still left a 44 dB hole at the worst seam; 1800 ms brought it to 32 dB and 2500 ms gained almost nothing more, so the dialogue-short template defaults to 1800 and exposes it as audio_bleed_ms:

{
    "task": {
        "command": "concat_videos",
        "arguments": {
            "videos": ["previous_result:shot_1", "previous_result:shot_2"],
            "trim_frames": 0,
            "audio_bleed_ms": 1800,
            "fps": 24
        }
    },
    "result": { "content_type": "video/mp4", "fps": 24 }
}

A bleed works because it copies ambience, which has no pitch and no attacks to give the copy away. It is the wrong tool for anything tonal - a copied musical phrase or half-spoken word reads as a stutter whichever direction it runs. bleed_join checks the outgoing tail's spectral flatness and warns when it looks tonal or speech-like rather than noise-like, so this failure mode surfaces in the job's warnings list instead of only in the mix. When a shot ends on something tonal, either give the cut a continuous bed with slice_audio + loop_audio + mix_audio + pair_audio, which leaves no seam to treat at all, or fade the seam gracefully with seam_fade_ms (a hundred or so milliseconds) and accept the cut. That advice inverts on a continuous bed - a laugh track, room tone - where a longer fade only digs the hole deeper (the same sitcom cut measured 54-59 dB holes with a 250-500 ms fade and no bleed). audio_bleed_ms wins where both are set and there is material to bleed. A bleed covers the gap but cannot fill it: the silence is inside the incoming shot's own head, and the only complete fix is a continuous bed under the whole cut: slice_audio a few seconds of tone out of a shot, loop_audio it to the length of the cut, mix_audio it under the episode and pair_audio it back onto the picture.

Levels are the other thing a cut has to reconcile, and no fade control can touch it. Shots generated independently land wherever the model put them - two shots of one scene, same template, same cast, measured peak_dbfs -2.65 and -12.56 - and each reads as fine on its own, because a shot is only wrong relative to what it is cut against. Butt-joined, that is a 10 dB drop at the cut, and it is not an artifact at the seam that a fade could smooth: it is either side of it. match_levels scales each track before the join - "rms" matches perceived level, which is usually what "make these sound the same" means, and "peak" matches the loudest sample, which is the safer choice on material with big transients. A shot whose gain would clip at the target is held just below full scale and the log says so. Left off - the default, so nothing existing changes - a spread of 6 dB or more across the tracks being joined is reported as a warning rather than passing in silence: on the job's warnings and as a warning event in its stream, not only in the server's log, since the caller who can act on it is the one who asked for the run. The other end of the range warns too: a shot at or below -40 dBFS, or one that needs 20 dB or more of gain to reach the target, is noise floor rather than a quieter performance, and matching it up is reported as match_levels_near_silent. dissolve_videos takes the same pair.

The joined soundtrack is fitted to the joined frames. A track that comes out short of the frame grid - rounding in an input's own track, which otherwise compounds join after join - is padded with silence to it. A pad of a frame or more is a warning (joined_audio_padded_to_frames); less than a frame is rounding, and only logged. The file's AAC encode can then trim the track by a further handful of samples (typically 16-32, under a millisecond), which is logged the same way, or warned as joined_audio_short_after_mux if it reaches a frame. Either way the recorded shots are re-measured against the file as written, so media.shots stays accurate. Both apply to dissolve_videos the same way, and neither task warns about resampling inputs that already agree to a sample_rate the caller pinned.

dissolve_videos

Join videos with a cross-dissolve at every seam, and fade the whole piece in from and out to a colour. Where concat_videos cuts - right for shots that each carry their own sound - this melts one shot into the next, which is what a montage cut to a score wants:

{
    "task": {
        "command": "dissolve_videos",
        "arguments": {
            "videos": ["previous_result:shot_1", "previous_result:shot_2"],
            "dissolve_frames": 12,
            "fade_in_frames": 12,
            "fade_out_frames": 24,
            "fps": 24
        }
    },
    "result": { "content_type": "video/mp4", "fps": 24 }
}
Argument Required Description
videos Yes The videos to join, in order - previous_result references, or the path or URL of a video file an earlier run wrote, one entry per video as with concat_videos; an entry may also be a {"location": ...} dict wrapping either. Every video must be the same frame size - unlike a sample-rate mismatch, this task does not resize one for you, so a statically-resolvable (asset:/output:/literal path, wrapped in a {"location": ...} dict or not) size disagreement is refused at validate; a previous_result: or other reference not yet resolved still fails only at run time (#504, #518). To fit the odd video: video_frames to get its frames, resize_rescale to the target size (resize_center_crop squares the frame first and then stretches it, distorting a non-square target), then pair_audio(fit="video") to put its soundtrack back before passing it here (#551)
dissolve_frames No Frames of overlap at each seam, blended linearly (default: 12). 0 is a hard cut
fade_in_frames No Frames over which the first video rises out of fade_color (default: 0)
fade_out_frames No Frames over which the last video sinks into it (default: 0)
fade_color No The RGB colour the fades come from and go to (default: black)
fps No Frame rate of the videos - required to crossfade audio at a dissolve, and the rate the dissolved file is written at unless result.fps overrides it
match_levels No Even the shots' loudness out before joining - "rms" or "peak", as with concat_videos. Off by default
match_levels_dbfs No The level match_levels moves every shot to (default: -1 dBFS for peak, -20 dBFS for rms). A shot that would clip at the target is held at -0.5 dBFS peak instead, reported as a match_levels_held warning with a per-shot log event

Every seam shortens the result by one overlap, so eight 124-frame shots joined with 12-frame dissolves run 908 frames, not 992 - size a soundtrack slice to the joined length, not the sum. When every input carries audio, the tracks are crossfaded over exactly the seam's span so they stay in step with the picture; when any input is silent the result is, and pair_audio puts a score under it.

Example: dissolve-between-shots.json

join_into_song

Join a spoken scene into a musical number: dialogue shots keep their own audio, and the shots sung after them play over the unbroken song rather than the separate slices each was generated against. The song's entry point is dialogue length - cue_seconds, and a workflow cannot do that arithmetic itself - a hand-computed literal goes stale the moment one dialogue shot is regenerated at another length. This task measures the joined dialogue at run time and places the song from it, so the offset never goes stale:

{
    "task": {
        "command": "join_into_song",
        "arguments": {
            "dialogue": ["previous_result:line_a", "previous_result:line_b"],
            "song_shots": "gather:sung",
            "song": "asset:song.mp3",
            "cue_seconds": 1.5
        }
    },
    "result": { "content_type": "video/mp4" }
}
Argument Required Description
dialogue Yes The spoken shots, in order, each keeping its own audio - a non-empty list, the same entries concat_videos' videos takes: previous_result references, or the path or URL of a video file an earlier run wrote, each optionally wrapped in a {"location": ...} dict. A shot with no track is filled with silence for its length
song_shots Yes The sung shots, in order - a non-empty list in the same shapes as dialogue. Their own audio is discarded; they play over song
song Yes The unbroken track the song shots were sliced from - an audio result, a video or audio file's path, or anything carrying .audio and .sample_rate. Its rate is the output's
cue_seconds No The song time that lands on the first song shot's frame 0 - the start of the slice that shot was generated against. 0 (the default) starts the song exactly at the cut; above 0 the song enters that long before it, under the last spoken line. Longer than the dialogue is refused - the song would have to start before the film. Must be ≥ 0
dialogue_target_lufs No Integrated loudness (BS.1770) each dialogue shot is gained to, with one static gain per shot. Omitted (the default), the shots keep their own levels. A shot too short (under 400 ms) or too quiet to measure is left at its own level, with a dialogue_unmatched warning
duck_delay_ms No How long after the song enters the dialogue starts to duck (default: 0). At or past the dialogue's end, nothing ducks. Must be ≥ 0
duck_db No How far the dialogue ducks, in dB (default: -12). Must be ≤ 0
duck_ramp_ms No The length of the linear ramp into the duck (default: 250). Must be ≥ 0
fps No The rate the videos play at, needed only when none of them carries one of its own - a pipeline's frames carry none, a file brings its own. Must be > 0

The timeline, in samples at the song's rate: each dialogue shot's track is fitted to its own frames (trimmed or padded with silence, warning dialogue_fitted_to_frames past a frame's worth of difference), so the joined dialogue ends exactly where the first song shot's frame 0 is - call that sample D. The song is placed at D - cue_seconds * sample_rate, so song time cue_seconds lands on that frame, and it runs to the end of the picture; a song shorter than that warns song_short and is padded with silence rather than looped. The dialogue ducks by duck_db from duck_delay_ms after the song enters, over a linear duck_ramp_ms ramp. The song shots' own audio is discarded, and a dialogue shot with no track of its own is filled with silence for its length rather than skipped, so later shots do not land early.

Frames are joined one for one, so every video must share one frame size and one frame rate: a mismatch is refused rather than resampled, as is a join where no video carries a rate and fps is not given, or an fps that contradicts the rate the videos carry. A step that wrote its video with result.fps hands that rate on, so a shot written at 12 fps from 24 fps frames is a 12 fps shot to the join, as it is when read back with output:.

There is no final normalization here - what level a deliverable sits at is the workflow's to decide, with normalize_audio after this step; the templates that mux to video normalize to -3 dBFS peak, as with any pair_audio mux.

The result is one video/mp4: the dialogue's frames then the song shots', over the mix, with a shot record per input (shot@<key>, named from the step's dialogue then song_shots references, as for_each names a member), each shot's samples measured off the built waveform rather than derived from its frames.

stabilize_video

Remove a generated clip's accumulated framing drift - the slow wander a video model adds over a shot that was meant to hold still:

{
    "task": {
        "command": "stabilize_video",
        "arguments": {
            "clip": "variable:shot_1",
            "smooth": 0
        }
    }
}
Argument Required Description
clip Yes The video - a frame list, a frame array or tensor, an audio+video pair, or the path or URL of a video file, read with its audio, so a shot an earlier run wrote can be steadied without regenerating it
smooth No 0 (the default) locks the framing to the first frame, which is what a shot generated from a pinned keyframe wants. A window in frames instead removes only the wander faster than that window, so a slow deliberate camera move survives and the drift around it does not

The argument is clip, not video, on purpose: the engine loads an argument named video itself, as bare frames, which would strip the soundtrack off before the task ever saw it. Frames are shifted back and the result is cropped to the region every frame covers, then resized to the original size; a soundtrack passes through untouched.

It is a stabilization pass, not a format pass. smooth: 0 on a shot with a deliberate camera move fights the move - every frame is shifted back toward the first, cropped and rescaled - and nothing downstream will notice, since the duration, size and sample rate all survive. Run it on a shot that drifts, as its own step; do not run it on every shot before a cut, which is what the assembly templates once did and what made their output visibly wider than the source. The join tasks refuse shots of different sizes, so no normalization step is needed before them.

trim_video

Keep one span of a video's frames, and its audio over the same span:

{
    "task": {
        "command": "trim_video",
        "arguments": {
            "video": "previous_result:shot",
            "start_frame": 12,
            "num_frames": 72
        }
    }
}
Argument Required Description
video Yes The clip. A path, asset: or output: reference is read with its audio; an earlier step's video is used as it is. A URL is fetched as frames only
start_frame Yes The first frame kept, from 0
num_frames Yes How many frames are kept: 1 or more
fps No The clip's frame rate. Defaults to the rate the clip carries; needed only to cut its audio, which is refused without one

Frames [start_frame, start_frame + num_frames) are kept. The audio is cut to the same span at the track's own sample rate: each end is frames_to_samples of a frame index, rounded on its own, so consecutive trims tile the track with no sample lost or repeated. A clip with audio (or an audio+video pair) comes back as an audio+video pair with its frame rate and sample rate kept; a bare frame list comes back as frames.

The shots the clip carries are clipped to the span and re-based to start at 0, with their sample side cleared. A clip that carries none (a decoded file, a fresh render) comes back as one shot spanning the kept frames, named after the file it was read from (else video 1), its samples measured off the cut track. A span that reaches past the clip's end is refused, naming the clip's frame count - it is never shortened or padded - and so is a num_frames of 0 or less or a negative start_frame; validate_workflow catches these on literal values.

window_video

Cut one overlapping, fixed-length window out of a long video, so a video-to-video model that reads at most one bucket of frames (121 for LTX) can work through a longer source a window at a time:

{
    "name": "window",
    "for_each": [
        {"name": "w0", "index": 0},
        {"name": "w1", "index": 1},
        {"name": "w2", "index": 2}
    ],
    "task": {
        "command": "window_video",
        "arguments": {
            "video": "asset:long-take.mp4",
            "index": "item:index",
            "num_frames": 121,
            "overlap": 16
        }
    }
}
Argument Required Description
video Yes The long source. A path, asset: or output: reference is read with its audio; an earlier step's video is used as it is. A URL is fetched as frames only, so its window is silent
index Yes Which window, from 0
num_frames Yes Frames per window - the model's bucket, e.g. an 8n+1 LTX length
overlap Yes Frames each window shares with the one before it; 0 or more, and below num_frames
fps No The source's frame rate. Defaults to the rate its file was read at; needed only to cut its audio

With stride = num_frames - overlap, window i covers source frames i*stride - overlap up to (not including) i*stride + stride. The first window starts overlap frames before the source and repeats its first frame there; the last may run past the source's end and repeats its last frame. Every window is exactly num_frames frames, float32 in [0, 1] like loop_frames' output. A source of N frames takes ceil(N / stride) windows, indexes 0 to (N - 1) // stride.

The window's audio is the source's samples for its real frames, cut on the source's own frame boundaries (frames_to_samples, #401), so adjacent windows' strides tile the track exactly; the repeated frames carry silence of the length they would have had. A source with no track gives a silent window.

Memory (#695): a file source (a path, asset: or output:) is never decoded whole - the window reads only the source frames it covers, from the keyframe before them, and the soundtrack on its own, so a window of a long source holds about one window of float32 frames plus the decoded track. An earlier step's video is already in memory and is cut from there.

One step makes one window: a task cannot return a list of them, so drive it with for_each over {name, index} entries as above. It refuses an overlap that is not below num_frames, and a negative overlap or index

  • validate_workflow catches these on literal values - and, at run time, an index whose window would start at or past the source's last frame, naming the source's frame count and the last valid index.

join_windows

The other half of window_video: blend the processed windows back into one video the source's length. Give it the windows in order, the source they were cut from, and the same num_frames and overlap:

[
    {
        "name": "window",
        "for_each": "variable:windows",
        "task": {
            "command": "window_video",
            "arguments": {
                "video": "variable:source_video",
                "index": "item:index",
                "num_frames": 121,
                "overlap": 16
            }
        }
    },
    {
        "name": "joined",
        "task": {
            "command": "join_windows",
            "arguments": {
                "videos": "gather:window",
                "source": "variable:source_video",
                "num_frames": 121,
                "overlap": 16
            }
        }
    }
]

with "windows": [{"name": "w0", "index": 0}, {"name": "w1", "index": 1}, ...]. A step that processes each window goes between the two, and videos gathers that step instead.

Argument Required Description
videos Yes The processed windows, in window order: gather:<step> over the step that processed them
source Yes The video the windows were cut from, as for window_video's video. Its frame count sets the plan and its track is the output's
num_frames Yes Frames per window, as given to window_video
overlap Yes Frames each window shares with the one before, as given to window_video
curve No The blend's shape across a seam: cosine (the default, (1 - cos πt) / 2), smoothstep (3t² - 2t³) or linear (t)
fps No The source's frame rate. Defaults to the rate its file was read at; needed only to put its audio back

The result has exactly the source's frame count, each source frame once: window 0's real frames, then each later window's first overlap frames blended over the previous window's last overlap frames, and the last window's pad dropped. The incoming weight is w(t) at t = (k + 1) / (overlap + 1) for seam frame k - the open ramp dissolve_videos uses, so no seam frame is a bare copy of either side. The windows may be at a different size from the source (a 2x upscale); the output is at the windows' size.

Memory (#695): the output is built as uint8, with only a seam's overlap frames blended in float32, so a join holds the output at one byte a channel plus the windows it was given - never the whole video in float32. A file source is read for its frame count and soundtrack only, never its picture.

The window count must be exactly ceil(source_frames / (num_frames - overlap)) - the rule has one home, window_count in dw/task_domains.py, and window_video's last valid index is one below it. A different count is refused naming both numbers and the list entries (by index) to add or drop: by validate_workflow, before any window renders, when the count is knowable from the document - source an asset:/output: reference or a literal path (probed header-only), num_frames and overlap literal after substitution, videos a list (a gather: over a for_each step is one) - and otherwise at run time, once the source is decoded. The validate-time check is the task's (dw/window_count_errors.py), so any join_windows step gets it, and it applies the same window_count_problem the task does. It also refuses a window whose frame count is not num_frames (named by position) and windows of different sizes. An unknown curve and an overlap that is not below num_frames are refused by validate_workflow on literal values as well.

The output's audio is the source's own track over exactly frames_to_samples(source_frames), and none when the source has none; the windows' audio is discarded. Its shot records are one per window over the frames it owns (window i > 0 starts at its blended head), with overlap_frames on every seam and cumulative frames_to_samples sample spans (#401), so assess_output reads the seams as dissolves and its sync check measures against the source's timeline. Like concat_videos and dissolve_videos, its catalog shape is a cut task.

fit_to_model and restore_to_source

A video-to-video model works at its own size and frame count, not the source's. fit_to_model fits the source into the model's width x height and exactly num_frames, the pipeline runs on the fitted video, and restore_to_source puts the output back at the source's size and length:

[
    {
        "name": "fit",
        "task": {
            "command": "fit_to_model",
            "arguments": {
                "video": "asset:clip.mp4",
                "width": 768,
                "height": 512,
                "num_frames": 121,
                "mode": "letterbox"
            }
        }
    },
    {"name": "upscale", "pipeline": {"...": "reads previous_result:fit.video"}},
    {
        "name": "restore",
        "task": {
            "command": "restore_to_source",
            "arguments": {
                "video": "previous_result:upscale",
                "fit": "previous_result:fit.fit"
            }
        }
    }
]

mode is letterbox (default: scale to fit, centred on black), stretch (resize to fill exactly) or crop (scale to fill, centre-crop). The frame count is exact: a longer source is cut to its first num_frames frames and a shorter one holds its last frame. downscale (default 1) divides width and height, which must both be divisible by it: a 2x upscaler whose width x height is the output it renders takes its reference at half that, so it fits with downscale: 2 and passes the same width/height to the pipeline (templates/ltx2/upscale-clip). The step returns {video, fit}, read as previous_result:<step>.video and previous_result:<step>.fit. video is a float32 array of frames, height, width, RGB in 0-1 that carries its fps, so it goes straight to a pipeline's video or a reference condition's frames; saved, it is written as an mp4 at that rate. The fit record holds mode, source_width, source_height, source_frames, model_width, model_height, model_frames, content_box (where the source sits in the model frame) and source_box (the part of the source kept - all of it unless mode is crop), each box {x, y, w, h}.

restore_to_source takes the pipeline's output and that record; it has no scale argument. The scale is the output's size over the fitted size and must be the same across and down, so a 2x upscaler restores to twice the source. Letterbox bars are cropped off, and frames are trimmed to the source's count, dropping the held ones (an output shorter than that is returned as it is). crop cannot bring back the edges it cut, so it returns the kept region at the source's pixel density times the scale: 640x480 fitted into 512x288 restores to 640x360. A saved video keeps its exact size; only an odd side grows by one pixel, since the encoder needs even sides.

Neither task carries a soundtrack. Pair the source's with pair_audio after the restore.

crop_face_track

Follow one face through a clip and cut a steady square around it, for a face-detail pass that needs the face large and still:

{
    "name": "face_crops",
    "task": {
        "command": "crop_face_track",
        "arguments": {
            "clip": "asset:interview.mp4",
            "crop_size": 512
        }
    }
}
Argument Required Description
clip Yes The video - frames, an audio+video pair, or the path or URL of a video file. Named clip so the engine hands it over as read, with the shots it records
crop_size No Side of every crop in pixels, a multiple of multiple. Default 512
padding No Space added around the face on each side, as a fraction of its size, 0 to 3. Default 0.6
gate_full No Face width over frame width at or below which strength is 1. Default 0.06
gate_zero No Face width over frame width at or above which strength is 0; must exceed gate_full. Default 0.12
min_confidence No Detections scoring below this are ignored; below 1. Default 0.6
modulus No The crop count is padded to modulus * n + remainder frames - the frame grid of the model the crops feed. At least 1. Default 8
remainder No The padded count's remainder, 0 up to modulus - 1. Default 1
multiple No crop_size must be a multiple of this - the model's frame-size step. At least 1. Default 32
detector_repo No Hub repo holding the YuNet weights. Default opencv/face_detection_yunet, read at a pinned revision; any other must be a Hub repo id
detector_file No The detector file in that repo, a bare .onnx name. Default face_detection_yunet_2023mar.onnx
device No Where detection runs

modulus, remainder and multiple follow plan_cuts: the model's grid is the workflow's to declare. The defaults (8, 1, 32) are LTX's 8n+1 frames and 32-pixel steps, so a workflow that names none gets the output it always did; the face-repair template passes modulus: 8, remainder: 1, multiple: 32 explicitly. A modulus below 1, a remainder below 0 or not below modulus, and a crop_size that is not a multiple of multiple are refused at validation.

These defaults are the task's own. The face-repair template tunes them on lem for LTX-2.5 at padding 1.5, gate_full 0.03 and gate_zero 0.06: a medium face (7.5% of the frame's width when tuned) already renders cleanly, and the wider context keeps the refine from re-inventing hair and clothing.

Example: face-repair.json — find, refine and paste back a wide shot's far face.

Detection is OpenCV's YuNet, run on the full frame and on four overlapping enlarged tiles, merged by non-maximum suppression, so a face only a few dozen pixels wide is still seen. One track is kept and smoothed with an exponential moving average; across a short detector miss the last box is held at a decaying strength. The track restarts at each shot boundary the clip records, or, for a clip that records none, where the picture jumps: a sharp drop in its HSV colour histogram, or a spike in its difference from the previous frame (a cut between two framings of one picture keeps its colours). A padded square around the smoothed box is cut from every frame and resized to crop_size, then the crops are padded to modulus * n + remainder frames (8n+1 by default) with mirrored warm-up and cool-down frames.

Each frame carries a strength from 0 to 1: how much a face-detail pass should change it. It is 1 when face width over frame width is at or below gate_full, 0 at or above gate_zero, linear between, and multiplied by the hold decay on frames the detector missed.

It returns {crops, track}. crops is a video (no audio) at the clip's frame rate. track is a JSON record saved as its own .json file, readable by previous_result:<step>.track:

Field Description
source {width, height, frames, fps} of the input
crop_size, padding, gate_full, gate_zero, min_confidence The parameters used
detector {repo, file}
pad_before, pad_after Mirrored frames added to the start and end
crop_frames Frames in crops, including the padding
face_found Whether any frame had a face
message Present only when no face was found
frames One entry per source frame: box (the smoothed [x, y, w, h], or null), crop (the [x, y, side, side] square cut), strength, and state (tracked, held or none)
resets [{frame, reason, detail}] where the track restarted; reason is shot, cut or lost; a cut's detail is {histogram_correlation, frame_change}

A clip with no face still succeeds: every strength is 0, the crops are the centre square, and a no_face_found warning is raised. validate_workflow refuses a crop_size that is not a multiple of 32, a gate_zero at or below gate_full, a padding out of range, a detector_repo that is not a repo id and a detector_file that is not a .onnx name, before anything downloads.

paste_face_track

Put the crops of a face-detail pass back into the clip they were cut from, using the track record crop_face_track returned:

{
    "name": "pasted",
    "task": {
        "command": "paste_face_track",
        "arguments": {
            "clip": "asset:interview.mp4",
            "repaired": "previous_result:face_repair",
            "track": "previous_result:face_crops.track"
        }
    }
}
Argument Required Description
clip Yes The video crop_face_track tracked - frames, an audio+video pair, or the path or URL of a video file
repaired Yes The crops after a face-detail pass, still padded to the count crop_face_track produced; any frame size
track Yes The track record crop_face_track returned (previous_result:<step>.track), or the path of its saved .json, such as an output: reference
feather No Fraction of the paste's radius that fades out, 0 to 1. Default 0.3
color_match No Shift each crop's mean colour onto the source's inside the mask before blending. Default true

The pad frames are dropped: the crop at pad_before + i goes back on source frame i. Each crop is resized to the square it was cut from, then blended over the frame with a radial feathered mask times that frame's strength. The mask is the circle inscribed in the square; its outer feather fraction fades from 1 to 0 along a smoothstep, so the square's corners and everything past the circle keep the source. A square that ran past the frame's edge is clipped to the frame.

A frame with strength 0 (no face, or a face too large for the gate) is passed through exactly. With color_match, the crop's mean colour is shifted onto the source's inside the mask, so a repair pass that drifted in tone does not show as a patch.

It returns the source clip with its audio, frame rate and shots carried, the same frames one for one. It refuses a track whose frame count or frame size does not match the clip, a repaired with fewer frames than the track's crop_frames, and a track that is not a crop_face_track record. validate_workflow refuses a feather past 1 before the job runs.

Example: face-repair.json — the repaired crops pasted back over the source, its soundtrack kept.

video_frames

The frames of a generated video, as one (frames, height, width, channels) uint8 array. That is the shape an argument taking frames rather than a video wants - LTX-2's keyframe conditions, which are mapped from 0-255 - and it is one artifact where a list of frames would become one artifact per frame and multiply the step that consumed it:

{
    "name": "opening_frames",
    "task": {
        "command": "video_frames",
        "arguments": { "video": "previous_result:opening" }
    },
    "result": { "content_type": "video/mp4", "save": false, "fps": 24 }
}
Argument Required Description
video Yes The video - a frame list, a frame array or tensor, or an audio+video pair

An argument that goes through diffusers' video processor instead - LTX-2's IC-LoRA references - wants the [0, 1] frames the pipeline returned rather than this array; hand those over with previous_result:step.frames.

Example: extend-clip.json

pair_audio

Pair a video with an audio track, so the two are saved as one muxed file. A pipeline that generates its own soundtrack returns the pair together; anything working on the frames alone - a latent upsampler, an interpolator, an upscaler - returns frames without it, and this puts it back:

{
    "task": {
        "command": "pair_audio",
        "arguments": {
            "video": "previous_result:upscale",
            "audio": "previous_result:base"
        }
    },
    "result": { "content_type": "video/mp4", "fps": 24 }
}
Argument Required Description
video Yes The frames - a frame list, a frame array or tensor, or an audio+video pair whose own soundtrack is replaced; their own rate is carried through to the output, so result.fps is only needed to override it (frames that carry none are written at 8 fps)
audio Yes The soundtrack - a waveform, the earlier step whose video carried one, or the path or URL of an audio or video file; the last two bring their sample rate along. A mono track is fine: an mp4 audio stream takes stereo and nothing else, so saving duplicates the one channel into two and warns that it did
sample_rate No Sample rate of the waveform. Required unless audio carries one; given here it wins
fps No The rate the frames play at, only needed when they carry none of their own. Used solely to work out how long the video is - what fit and the length-mismatch check measure the track against - and is never written to the file; that's result.fps, which sets the rate the output plays at and defaults to 8 fps when the frames carry none
fit No "video" cuts or pads the track with silence to the length of the frames, warning either way (audio_padded_to_video / audio_trimmed_to_video). Left unset (the default) the track is used as it is, and a length that disagrees with the frames' is warned about rather than corrected (audio_video_length_mismatch). Any other value is refused at run time, not by validate_workflow

fit's guarantee is exact for the waveform handed to the encoder, not for the file the encoder writes: muxing is a lossy AAC encode, and it can still trim or pad the written track by a further handful of samples (#428 measured up to ~30, under a millisecond). That residual is logged, not warned; on a video with recorded shots it becomes a joined_audio_short_after_mux warning only if it reaches a frame. get_gallery_metadata's media.shots and assess_output's sync_length are measured against the written file, not the pre-encode prediction, so they are the number to trust for the track's actual length.

When the video carries recorded shots (from an earlier concat_videos, dissolve_videos or chain step), pair_audio remeasures each one's sample fields against the track it was handed. Every shot but the last is round(start_frame / fps * sample_rate); the last one runs to the track's actual end, and once the file is written it is measured again against what the file decodes to. So its num_samples can sit a few dozen samples off round(num_frames * sample_rate / fps): the encoder's trim, which the job's event log records. A real mismatch between the track and the video's length is a separate, thresholded warning (audio_video_length_mismatch, or audio_padded_to_video / audio_trimmed_to_video when fit: "video" corrected it), so a last shot short by less than a millisecond is expected, not a bug. get_gallery_metadata's media.shots reports the remeasured fields.

Example: assemble-and-score.json

slice_audio

Cut a slice out of an audio track, addressed in seconds or in video frames. Slices reaching past the end of the track are zero-padded — asking for more than the source holds returns a track of the length you asked for whose tail is digital silence, not a shorter track and not an error. Anything past a few milliseconds of that padding is reported as a slice_past_end warning on the job, because a score laid under a longer cut goes silent for the rest of the film without anything else saying so; to fill a cut longer than the recording, build a bed with loop_audio first and slice that. Either half of a pair may be left out - an omitted start begins at the head of the track, an omitted duration runs to the end of it - so a workflow that trims only when it is given a length still passes the whole track along:

{
    "task": {
        "command": "slice_audio",
        "arguments": {
            "audio": "./soundtrack.wav",
            "start_frame": 124,
            "num_frames": 124,
            "fps": 24
        }
    },
    "result": { "content_type": "audio/wav", "sample_rate": 44100 }
}
Argument Required Description
audio Yes Path or URL of an audio file (or of a video file, whose soundtrack is taken), a waveform from a previous step, or an earlier step's video generated with a soundtrack (which brings its sample rate along)
start_seconds / duration_seconds One pair The slice in seconds; either may be omitted
start_frame / num_frames / fps One pair The slice in video frames; fps is required, start and count may be omitted
lead_frames No Frame form only, default 0: extra audio before the cut. The slice starts at start_frame - lead_frames and still runs num_frames, so start_frame 48 with lead_frames 12 at 24 fps is audio from frame 36 for num_frames frames. A whole number, 0 or above; refused with the seconds form, and refused when start_frame - lead_frames is below zero (by validate_workflow for literals, at run time otherwise)
sample_rate With a waveform Sample rate of a directly passed waveform (files carry their own)

gain_audio

Apply a gain, in decibels, to a region of an audio track - the rest of the track passes through unchanged. The region is addressed the same way slice_audio's is, in seconds or in video frames, so ducking a scene under another (lowering a dialogue track between two timestamps) is one step instead of the slice_audio → gain (a whole-track normalize_audio on the slice) → mix_audio → rejoin → pair_audio chain that used to be the only way to gain part of a track rather than all of it. Unlike slice_audio, a region reaching past the end of the track is clipped to it rather than zero-padded - there is no silence there to gain, only the end of the real material:

{
    "task": {
        "command": "gain_audio",
        "arguments": {
            "audio": "./dialogue.wav",
            "gain_db": -12,
            "start_frame": 124,
            "num_frames": 48,
            "fps": 24
        }
    },
    "result": { "content_type": "audio/wav", "sample_rate": 44100 }
}
Argument Required Description
audio Yes Path or URL of an audio file (or of a video file, whose soundtrack is taken), a waveform from a previous step, or an earlier step's video generated with a soundtrack (which brings its sample rate along)
gain_db Yes Gain to apply within the region, in decibels - negative ducks it, positive boosts it
start_seconds / duration_seconds No The region in seconds; either may be omitted
start_frame / num_frames / fps No The region in video frames; fps is required if either is given, start and count may be omitted
sample_rate With a waveform Sample rate of a directly passed waveform (files carry their own)

No region argument is required: with every one of them omitted, the gain applies to the whole track (#395) - the same "no region means everything" reading mix_audio's gains use. To gain everything from some point on instead, give just start_seconds: 0 and leave duration_seconds unset (or start_frame: 0 + fps and leave num_frames unset), which runs to the end of the track without needing to already know how long that is.

crossfade_audio

Join audio tracks with an equal-power crossfade. Each seam overlaps the two tracks by the fade window:

{
    "task": {
        "command": "crossfade_audio",
        "arguments": {
            "audios": "previous_result:slices",
            "crossfade_ms": 75,
            "sample_rate": 44100
        }
    },
    "result": { "content_type": "audio/wav", "sample_rate": 44100 }
}

fade_audio

Fade a track in from silence and out to it. A slice cut out of the middle of a piece ends on whatever was sounding at the cut; a fade turns that into an ending. The curve is the equal-power cosine the seam joins use:

{
    "task": {
        "command": "fade_audio",
        "arguments": {
            "audio": "previous_result:soundtrack",
            "fade_in_ms": 500,
            "fade_out_ms": 2500,
            "sample_rate": 44100
        }
    }
}
Argument Required Description
audio Yes Path or URL of an audio file (or of a video file, whose soundtrack is taken), a waveform from a previous step, or an earlier step's video generated with a soundtrack (which brings its sample rate along)
fade_in_ms No Length of the fade in, from the head of the track (default: 0)
fade_out_ms No Length of the fade out, to the tail of the track (default: 0)
sample_rate With a waveform Sample rate of a directly passed waveform (files carry their own)

Example: audio-trim-fade.json — slice a generated track to length, then fade the cut into an ending.

normalize_audio

Scale a track so its loudest sample sits at a level. Generated music comes out wherever the model happened to land - a quiet take needs lifting before it sits under a picture, a hot one needs headroom before the encoder. Only the gain changes, so the dynamics survive:

{
    "task": {
        "command": "normalize_audio",
        "arguments": {
            "audio": "previous_result:faded",
            "peak_dbfs": -1.0,
            "sample_rate": 44100
        }
    }
}
Argument Required Description
audio Yes Path or URL of an audio file (or of a video file, whose soundtrack is taken), a waveform from a previous step, or an earlier step's video generated with a soundtrack (which brings its sample rate along)
peak_dbfs No The level the loudest sample is moved to, in dB below full scale (default: -1.0). 0 is full scale. With limit, a true-peak (4x oversampled, BS.1770) ceiling the output never crosses
target_lufs No An integrated loudness (BS.1770) to scale the track to instead, in LUFS (0 or below). Without limit the gain is capped so the peak stays under peak_dbfs, and target_lufs_capped warns when that cap wins. A track shorter than 400 ms (or silent throughout) cannot be measured, warns target_lufs_unmeasurable, and falls back to peak_dbfs
limit No Reach target_lufs past a transient instead of letting it cap the gain: a true-peak look-ahead limiter (5 ms look-ahead, 20 ms hold, 150 ms release, linked across channels) holds peak_dbfs, and the gain is searched for until the limited track lands within 0.1 LU of target_lufs - limiting takes some loudness back, more on dense material. The limiter itself never adds gain, and it stops at 12 dB of reduction - past that the gain stops too, and target_lufs_capped warns with limited: true and shortfall_lu (as it does for any track left more than 0.1 LU short). More than 6 dB warns limiter_heavy (pumping can be audible). The step's log carries constraint: "limiter", gain_db, max_gain_reduction_db, limited_fraction (0 when the limiter was not needed), output_true_peak_dbfs and output_lufs (default: false, which behaves exactly as without it)
sample_rate With a waveform Sample rate of a directly passed waveform (files carry their own)

A silent track is returned unchanged.

To level dialogue with a laugh or a shout on it, limit is the switch: a peak-only gain lets that one transient set the level of every line around it.

"arguments": {"audio": "previous_result:dialogue", "target_lufs": -16, "peak_dbfs": -1.0, "limit": true}

Example: dissolve-between-shots.json

Headroom and clipping warnings. Saving audio or a video with a muxed soundtrack checks the written level against two thresholds, reported in get_job's warnings and readable back afterward as media.peak_dbfs from get_gallery_metadata:

  • audio_no_headroom fires when a plain audio file's waveform, before encoding, peaks at or above -0.5 dBFS - encoding can push a level that already has no headroom over full scale.
  • audio_clipped fires when the file is decoded back after writing and measures at or above 0.0 dBFS - the ground truth of what a consumer's decoder will actually see, since an encoder's own overshoot varies by codec and is not reliably predictable from the pre-encode level.

A video mux only ever reports the second one: its pre-encode prediction is held rather than emitted, because a mux's overshoot is not reliably positive the way a plain audio encode's is. That leaves a real gap between the two thresholds - a video whose soundtrack decodes back between -0.5 and 0.0 dBFS produces no warning at all, because it predicted risk but measured clean. That is the file's own measured level, not a threshold bug: read media.peak_dbfs against -0.5 and 0.0 to judge a specific file rather than relying on the warning alone.

mix_audio

Layer tracks on top of one another. crossfade_audio puts tracks one after another; this puts them on top of each other - a score laid under a film's own sound, where the music runs unbroken while the world underneath it is replaced at every cut:

{
    "task": {
        "command": "mix_audio",
        "arguments": {
            "audios": ["previous_result:soundtrack", "previous_result:world"],
            "gains": [0.5, 1.0],
            "sample_rate": 44100
        }
    }
}
Argument Required Description
audios Yes The tracks to layer - waveforms, audio or video file paths, or videos generated with a soundtrack
gains No One plain multiplier per track, in the same order - not decibels. Defaults to unity on every track
sample_rate With a raw waveform Sample rate of the waveforms. Required unless every track brings its own; given here it wins

Tracks of different lengths are padded with silence to the longest, so a score shorter than the picture leaves the tail dry rather than cutting the picture down to fit. Summing can push peaks past full scale and the sum is not rescaled - follow it with normalize_audio to bring the peak back down.

Example: dissolve-between-shots.json — a generated score mixed under the shots' own audio.

loop_audio

Make a bed of a given length out of a short recording — the room tone laid under a whole cut, which is the only complete fix for the hole at a seam. Each shot in a cut carries its own room and nothing runs underneath the join; a continuous bed does, the way a location's room tone is laid under a dialogue scene so the edits stop being audible:

{
    "task": {
        "command": "loop_audio",
        "arguments": {
            "audio": "previous_result:room_tone",
            "target_frames": 620,
            "fps": 24,
            "crossfade_ms": 250
        }
    }
}
Argument Required Description
audio Yes Path or URL of an audio or video file, a video generated with a soundtrack (which brings its sample rate along), or a waveform
duration_seconds One of How long the bed should be, in seconds
target_frames / fps One of How long the bed should be, in video frames — how a bed is matched to a cut exactly
crossfade_ms No Crossfade at each loop point, clamped to the material available (default 250)
sample_rate With a waveform Sample rate of a waveform passed directly; given for a file or a video it overrides the rate they carry

Laps are joined with an equal-power crossfade rather than butted together, so the loop point itself is not a click. That only smooths the seam: a transient in the source (a hit, a swell) still recurs once per lap at full strength, so the loop still reads as a level pulse at the lap rate — measured at 9.3 dB on a source with one such transient. Pick a source with even internal level to avoid the pulse; the crossfade does not remove it. The source is used whole every lap and only the last one is trimmed, so the bed lands exactly on the requested length; a source longer than the request is trimmed to it.

The bed is laid under the cut with mix_audio and attached to the picture with pair_audio:

{ "name": "bed",   "task": { "command": "loop_audio",
                             "arguments": { "audio": "previous_result:room_tone",
                                            "target_frames": 620, "fps": 24 } } },
{ "name": "mixed", "task": { "command": "mix_audio",
                             "arguments": { "audios": ["previous_result:episode",
                                                       "previous_result:bed"],
                                            "gains": [1.0, 0.25] } } },
{ "name": "cut",   "task": { "command": "pair_audio",
                             "arguments": { "video": "previous_result:episode",
                                            "audio": "previous_result:mixed" } },
  "result": { "content_type": "video/mp4", "fps": 24 } }

The bed's source: find_loop_bed picks the stretch. Run it against the cut (or a stem of it) over the range that should hold room tone, copy the top candidate's start_seconds/duration_seconds into slice_audio, loop_audio the slice to the cut's length, and mix_audio it in at the candidate's gain — the room the shots were generated in, picked by measurement rather than by ear.

find_loop_bed

Pick the window of a recording worth looping into the room-tone bed above, measured as it will sound looped rather than as it sits in the source. A level check alone misses three things, each found on a real episode:

  • near-programme material — faint speech attenuated ~30 dB reads as quiet, and is audible once it repeats every lap.
  • lap-rate modulation — the loop's repeat rate beats against the source's own level movement, invisible in one pass through the source.
  • ticks — a 1 ms transient is invisible to 50 ms RMS and recurs once per lap.
{
    "task": {
        "command": "find_loop_bed",
        "arguments": {
            "audio": "output:episode/latest/final/cut.mp4",
            "start_seconds": 30.0,
            "end_seconds": 90.0
        }
    },
    "result": { "content_type": "application/json" }
}
Argument Required Description
audio Yes Path, asset:/output: reference of an audio or video file (a video is read audio-only - frames are never decoded), or a track or video from previous_result:
start_seconds / end_seconds No The range to search; the whole file if omitted
min_seconds / max_seconds No Shortest/longest window tried (default 0.5 / 2.0)
max_bin_dbfs No Every 50 ms bin of a window must be at or below this (default -55)
max_mean_dbfs No A window's mean (RMS) level must be at or below this (default -60)
max_spike_db No Most a window's largest 1 ms peak may sit above its median 1 ms peak (default 12)
crossfade_ms No The crossfade candidates are looped with - loop_audio's own default, so what is measured is what loop_audio will make (default 250)
loop_seconds No Length of the looped result that is measured (default 10.0)
target_bed_dbfs No The level each candidate's gain is computed to reach (default -60)
max_candidates No How many ranked candidates to return (default 5)
shots No Shot boundaries, the assessment probes' shape: [{name, start_frame, num_frames}], optionally with start_sample/num_samples; overrides the shots a video carries or its run records
fps No Frame rate the shots' frames count at, for a source with none of its own (an audio file); a video's own rate otherwise

Every window on a 50 ms grid, from min_seconds to max_seconds long, is judged against four rules in order and counted in rejected under the first one it fails, so rejected is a tally of the whole grid:

  • too_loud — a 50 ms bin above max_bin_dbfs, or a mean above max_mean_dbfs.
  • silent — digital silence: a mean of zero, or a median 1 ms peak of zero (more than half the window is exact zeros, whatever sits in the rest).
  • spike — its largest 1 ms peak more than max_spike_db above its median 1 ms peak. The largest peak is also looked for in the 5 ms just outside each end, unless that neighbouring bin is itself too loud, so a window ending on a click is thrown out rather than putting the click's onset at the loop's seam.
  • tonal — the flatness/harmonicity test bleed_join uses, taken over every 0.1 s and every 0.2 s block inside the window (one per 50 ms step; no longer than min_seconds). A window is tonal when any block in it is. Faint speech comes and goes, and over a whole window the pauses dilute a syllable below the threshold; in the block it sits in, it is not diluted. Two lengths, because 0.1 s sits inside one syllable and 0.2 s holds enough periods of a low hum. Flatness is measured over the band the source actually occupies: a source resampled up (a 16 kHz bed mixed at 24 kHz) has an empty band above its own Nyquist that reads as tonal whatever the material is, as bleed_join's tail does (#198). A candidate's flatness and harmonicity are the readings of its blocks closest to failing (lowest flatness, highest harmonicity).

On a cut, a bed must come from inside one shot: a window across a cut loops the seam's change of room as a once-per-lap step. Shots resolve in the probes' order - the shots argument (shots_source: "argument"), else the shots a video from an earlier step carries ("artifact"), else the ones the run manifest beside the file records ("manifest", which is why output: the joined file finds them unasked), else none (null, the search above unchanged). With shots, a window that crosses a boundary, or lies where no shot covers, is counted under rejected.shot_boundary before the four rules (so theirs count only in-shot windows), each candidate names its shot (shot@<name> from a manifest), and source.shots lists each shot's {name, start_seconds, end_seconds} as placed in the soundtrack. A shot with both frames and recorded samples is placed at their overlap - the picture's cut and the audio's can sit a few samples apart. A shot placed by frames on a source with no frame rate is refused; pass fps.

The survivors are thinned so no two overlap, steadiest source first, to a pool of up to 200 (LOOPED_POOL). The pool does not depend on max_candidates, which only cuts the final ranking, so asking for one candidate returns the default run's first. Each one in the pool is then looped with loop_audio's own crossfade to loop_seconds and measured: ripple_db (the 5-95% spread of the looped 50 ms bins), envelope_peak_db/ envelope_peak_hz (the strongest level wobble) and lap_component_db (the wobble at the lap rate). Candidates are ranked by looped ripple_db, lowest first; one whose envelope_peak_db is above -15 dB carries the lap_modulation warning.

{
    "source": { "duration_seconds": 620.4, "sample_rate": 44100, "searched": [30.0, 90.0], "shots_source": null },
    "criteria": { "min_seconds": 0.5, "max_seconds": 2.0, "target_bed_dbfs": -60.0 },
    "candidates": [
        {
            "rank": 1,
            "start_seconds": 41.28,
            "duration_seconds": 1.35,
            "end_seconds": 42.63,
            "shot": null,
            "mean_dbfs": -63.1,
            "max_bin_dbfs": -57.4,
            "spike_db": 4.2,
            "flatness": 0.61,
            "harmonicity": 0.08,
            "looped": {
                "ripple_db": 1.1,
                "envelope_peak_db": -24.0,
                "envelope_peak_hz": 0.4,
                "lap_hz": 0.096,
                "lap_component_db": -26.0
            },
            "gain_db": 3.1,
            "gain": 1.43,
            "warnings": []
        }
    ],
    "rejected": { "too_loud": 812, "silent": 0, "spike": 14, "tonal": 203 },
    "findings": []
}

start_seconds/duration_seconds are slice_audio's arguments and gain is mix_audio's multiplier for reaching target_bed_dbfs — the answer picks a window and builds nothing, so the remedy is a copy. Finding nothing is an answer, not an error: candidates is [], rejected counts the windows each rule threw out, and one no_loop_bed finding says which rule to relax.

Like the assessment probes it answers JSON and decides nothing, but it searches rather than checking a finished cut, so it is not one of them — it stays out of list_tasks' assessment list. The result must be saved as application/json.

{
    "variables": { "audio": "output:episode/latest/final/cut.mp4" },
    "steps": [
        { "name": "bed", "task": { "command": "find_loop_bed",
                                    "arguments": { "audio": "variable:audio" } },
          "result": { "content_type": "application/json" } }
    ]
}

run against the finished cut once it exists, then read back over MCP with get_output_text (JSON results are text) rather than get_output_audio.

Out-of-range literals (min_seconds/max_seconds/loop_seconds at or below zero, max_candidates at or below zero, a negative crossfade_ms, and the like) are refused at validation. end_seconds at or before start_seconds, min_seconds above max_seconds, a range past the end of the file, and a source shorter than min_seconds fail at run time — they depend on the file.

resample_audio

Convert a track to a different sample rate. A pipeline that conditions on audio wants it at its own rate (MiniMax H3 at its audio VAE's), and resampling a supplied recording once, up front, feeds it what it already wants:

{
    "task": {
        "command": "resample_audio",
        "arguments": {
            "audio": "previous_result:edit",
            "target_sample_rate": 44100
        }
    }
}
Argument Required Description
audio Yes Path or URL of an audio or video file, a video generated with a soundtrack (which brings its sample rate along), or a waveform
target_sample_rate Yes The rate to convert to
sample_rate With a waveform Sample rate of a waveform passed directly; given for a file or a video it overrides the rate they carry

A track already at the target rate is returned untouched. The conversion is PyAV's, which dw already needs for video - no torchaudio dependency.

Every audio task returns the waveform and the rate it is at, so one chains into the next without the rate being restated: a resample_audio fed previous_result: from a slice_audio takes the source rate from the slice. A sample_rate given on the step still wins, and one declared on the step's result still decides what is written to disk.

Example: assemble-and-score.json

compress_audio

Shape a track's dynamics with an envelope-follower - a compressor, a limiter and a gate are the same algorithm with different knob settings, so one task covers all three through mode:

{
    "task": {
        "command": "compress_audio",
        "arguments": {
            "audio": "previous_result:mixed",
            "threshold_dbfs": -18.0,
            "ratio": 4.0,
            "attack_ms": 10,
            "release_ms": 100,
            "sample_rate": 44100
        }
    }
}
Argument Required Description
audio Yes Path or URL of an audio file (or of a video file, whose soundtrack is taken), a waveform from a previous step, or an earlier step's video generated with a soundtrack (which brings its sample rate along)
threshold_dbfs Yes The level the envelope is measured against, in dB below full scale. Cannot be above 0
ratio No How hard the reduction is above the threshold, in compress/gate mode (default: 4.0). Ignored in limit mode, which always holds the signal at the threshold
attack_ms No How fast the envelope rises to a louder signal (default: 10.0). 0 means instantly
release_ms No How fast the envelope falls back after a louder signal ends (default: 100.0). 0 means instantly
mode No compress (turn down what's above the threshold), limit (hold the signal at the threshold), or gate (turn down what's below the threshold) (default: compress)
sample_rate With a waveform Sample rate of a directly passed waveform (files carry their own)

A silent track is returned unchanged.

mode: "limit" is a sample-peak limiter with no look-ahead: the envelope reacts to a transient as it arrives, so the transient's leading edge and any inter-sample peak get through. To hold a true-peak ceiling while reaching a loudness target, use normalize_audio with limit: true.

filter_audio

Run a track through a single biquad filter stage - trimming the frequencies a mix doesn't need, or carving out room for another element:

{
    "task": {
        "command": "filter_audio",
        "arguments": {
            "audio": "previous_result:world",
            "cutoff_hz": 120,
            "kind": "highpass",
            "sample_rate": 44100
        }
    }
}
Argument Required Description
audio Yes Path or URL of an audio file (or of a video file, whose soundtrack is taken), a waveform from a previous step, or an earlier step's video generated with a soundtrack (which brings its sample rate along)
cutoff_hz Yes The filter's corner frequency. Must be below the Nyquist frequency (half the sample rate)
kind No lowpass, highpass, bandpass, or notch (default: lowpass)
q No The filter's resonance/bandwidth (default: 0.707, a Butterworth response)
sample_rate With a waveform Sample rate of a directly passed waveform (files carry their own)

A silent track is returned unchanged.

analyze_audio

Measure a track without changing it - peak and RMS level, crest factor, and a rough low/mid/high spectral balance, the numbers a compress_audio or filter_audio step downstream is tuned against rather than guessed at:

{
    "task": {
        "command": "analyze_audio",
        "arguments": {
            "audio": "previous_result:mixed",
            "sample_rate": 44100
        }
    }
}
Argument Required Description
audio Yes Path or URL of an audio file (or of a video file, whose soundtrack is taken), a waveform from a previous step, or an earlier step's video generated with a soundtrack (which brings its sample rate along)
sample_rate With a waveform Sample rate of a directly passed waveform (files carry their own)

Returns a dict, not a track: peak_dbfs, rms_dbfs, crest_factor_db, low_dbfs (20-250 Hz), mid_dbfs (250-4000 Hz), high_dbfs (4000-20000 Hz). The three bands are each a share of the track's total power on the same scale as rms_dbfs (their powers sum to it), so the loudest band sits near rms_dbfs rather than tens of dB under it - comparable to compress_audio's threshold_dbfs. A silent track, or a band with no content at the track's sample rate, reads as null rather than -inf.

analyze_beats

Find where a song's beats fall, to cut picture to them - a Music 3 song carries no tempo map, and a cut that lands on the beat needs one. An onset envelope gives the tempo and the beats are tracked through it; the result is JSON and nothing is built:

{
    "name": "beats",
    "task": {
        "command": "analyze_beats",
        "arguments": {
            "audio": "previous_result:song",
            "anchors": [0.52, 31.9]
        }
    },
    "result": { "content_type": "application/json" }
}
Argument Required Description
audio Yes Path, asset:/output: reference of an audio or video file, a waveform from a previous step, or an earlier step's generated audio or video
sample_rate With a waveform Sample rate of a directly passed waveform (files carry their own)
tempo_bpm No The tempo, when known. With exactly one anchor it lays an exact grid and nothing is detected; otherwise it steers detection toward that tempo
anchors No Where beats are known to fall, ascending: a list of times in seconds, or of {beat_index, seconds} marks placing detected beat beat_index (from 0) at seconds. One kind per list
min_bpm / max_bpm No The tempo range searched (default 60 / 200); min_bpm must be below max_bpm

Returns {bpm, beats, downbeat_phase, method, calibration, duration_seconds, warnings}: beats ascend, in seconds. method is onset (tracked from the onsets), rms_peaks or grid. downbeat_phase is which of the first four beats starts a bar, null when none stands out. calibration is {offset_s, drift, anchors_used}, the anchors' shift at the first mark and how much it changes by the last. warnings repeats what was emitted.

Anchors correct the detection three ways:

  • Two or more warp the detected beats piecewise-linearly through them, so a detection that starts late or drifts lands on the marks.
  • One alone shifts every beat by the same amount.
  • One plus tempo_bpm lays an exact grid through it, with no detection: method is grid.

A bare-seconds anchor snaps to the nearest detected beat, so it corrects a drift of under half a beat; where the detection is further off than that, use {beat_index, seconds} marks, which name the beat instead of guessing it.

A track with no clear, regular onsets - a pad, a swell, a noise bed at any level - falls back to the peaks of its loudness (method rms_peaks, with a warning that no reliable beat was found and the peaks follow swells rather than a pulse). A level that drifts (a fade, a loud first second) is not a pulse. A pulse that tracks but is faint - its onsets repeat weakly at the period and its beats sit low against the rest of the envelope, as in a song with a long quiet build - is kept, with a warning to check a few beats against the song and pass anchors or tempo_bpm if they are off. The tempo is the track's strongest period, folded into the range: a half-time song whose pulse is 50 BPM reports 100, not a riff that happens to repeat at 95. A silent track returns empty beats and a warning. So does a near-silent one whose loudest 50 ms is under -40 dBFS - room tone, hiss, or a song mixed far too low - since the log-compressed onset envelope would find a pulse in it. Raise a real song's level first (gain_audio).

bpm is always within min_bpm-max_bpm, or null. A pulse found outside the range is folded by octaves into it (85.7 BPM searched at 140-200 reports 171.4, beating every half pulse). Where no octave fits, the beats keep to a tempo in the range and a warning names the pulse. Where the rms_peaks peaks have no octave in the range, bpm is null with a warning. The beats run to the song's ends: a hit at 0 s is a beat, and a noise bed starting at 0 s is not.

Refused: min_bpm at or above max_bpm, and malformed anchors (mixed kinds, out of order, a non-whole beat_index), at validate_workflow; an anchor past the song's end, or a beat_index past the detected count, at run start. It returns JSON, so result may only be application/json.

plan_cuts

Plan a music video's cuts from a song's lyrics and beats. transcribe_audio (with timestamps) says when each line is sung and analyze_beats says where the beats fall; this turns the two into shots - each a start frame and a length - that tile the song exactly, so one prompt per shot can be written and the list rendered. The result is JSON and nothing is built:

{
    "name": "plan",
    "task": {
        "command": "plan_cuts",
        "arguments": {
            "transcript": "previous_result:transcribe",
            "beats": "previous_result:beats",
            "duration_s": "previous_result:beats.duration_seconds",
            "lyrics": "variable:lyrics",
            "segment_by": "line",
            "fps": 24,
            "min_scene_s": 1.5,
            "max_scene_s": 10,
            "vocal_tail_s": 0.5,
            "snap_to_beats": true
        }
    },
    "result": { "content_type": "application/json" }
}
Argument Required Description
transcript Yes transcribe_audio's result with timestamps set ("segment" or "word"): a {text, chunks} dict of {start, end, text} chunks in seconds. A null end runs to the next chunk's start, or the song's end. Its result must be application/json
lyrics No The song's own lyrics, one line per line with a blank line or a lone section tag ([Chorus]) between stanzas, or a list of lines. The shots carry these lines verbatim and in order; the transcript only times them. null or "" means none: the transcript's text is used
beats For segment_by: "beat" analyze_beats' result, or a list of beat times in seconds
segment_by No line (default), stanza or beat
fps No Frames per second the shots are counted in (default 24)
duration_s No The song's length; analyze_beats' duration_seconds when beats is that whole dict, else the transcript's last end (with a warning). "beats": "previous_result:beats" passes only the result's beats list - the engine takes the key an argument is named for - so name the length too: "duration_s": "previous_result:beats.duration_seconds"
min_scene_s No The shortest a shot may be (default 1.0)
max_scene_s No The longest a shot may be (default none)
vocal_tail_s No Seconds a sung shot holds past its last word (default 0)
include_instrumental_gaps No Whether a long silence becomes its own instrumental shot (default true)
min_gap_seconds No The shortest silence that becomes its own shot (default 2.0)
snap_to_beats No Whether every cut moves to its nearest beat (default false)
modulus / remainder No The model's render grid: a shot renders modulus * n + remainder frames (H3: 17, 5). With none, a shot renders exactly its cut
min_frames / max_frames No The fewest and most frames a shot may render (H3: 124, 345)
lead_s No Seconds a shot starts before its cut, as a run-up the model renders and trim_video drops. Each shot's lead_frames is min(round(lead_s * fps), start_frame), so the first shot has none (default 0)

Returns {shots, bpm, fps, duration_s, total_frames, render_frames, warnings}. shots tile the song in order, each {name, start_frame, num_frames, cut_frames, lead_frames, lyric, kind}: start_frame and cut_frames are the cut's start and length on the timeline (the cut_frames sum to round(duration_s * fps)), lead_frames the run-up before the cut, and num_frames the length to render: max(min_frames, lead_frames + cut_frames) raised to the grid, so the render starts at start_frame - lead_frames and its cut is frames [lead_frames, lead_frames + cut_frames). With no grid arguments num_frames is cut_frames and lead_frames 0. render_frames is the sum of the num_frames - what the render costs, which is more than total_frames when shots lead or round up. lyric the lines sung in it joined by newlines or null, and kind "vocal" or "instrumental". A dict rather than a list, since a list result would become one artifact per shot. Its result may only be application/json.

How the plan is made:

  • With lyrics, those lines are the lines, and the transcript only lends them timings: they are aligned to it word by word, since Whisper mishears sung words and splits lines where it likes. A line never heard is placed between its neighbours with a warning, as a shot of its own: when its neighbours touch, it takes min_scene_s (at least half a second) from them, never more than half of either. Blank lines and section tags are stanza breaks, not sung lines. Without lyrics, each transcript chunk is a line.
  • segment_by makes a shot per line, per stanza (the lyrics' blank-line groups, or lines without a min_gap_seconds silence between them) or per few beats (a cut on the first beat at least min_scene_s after the last).
  • The silence between lines goes to the shot before it, or - when it lasts min_gap_seconds and include_instrumental_gaps is on - becomes its own instrumental shot, as can the silence before the first line and after the last.
  • A shot over max_scene_s splits evenly, each cut on its nearest beat when there are beats, and a sung shot split this way is warned about by its lyric, which every piece carries; one under min_scene_s merges into its shorter neighbour. A shot still outside the range is warned about by name.
  • A shot whose render would pass max_frames is split, on a beat when there is one, and warned about.
  • snap_to_beats moves every cut to its nearest beat; with no beats it warns and leaves the cuts where the lines put them.
  • Every boundary is rounded once, from its absolute time, so the frame counts sum to the song's and no rounding drifts.

Refused, at validate_workflow where the value is a literal and at run start otherwise: a bare-string transcript (the message names timestamps / return_timestamps), segment_by not one of the three, segment_by: "beat" without beats, min_scene_s above max_scene_s, and fps, duration_s, max_scene_s at or below zero or min_scene_s, vocal_tail_s, min_gap_seconds below zero, and a remainder outside 0 to below modulus (or one with no modulus) or a max_frames below the smallest grid length at or above min_frames. Also refused at run start: a song whose length can't be known (no duration_s, no beats result, and no transcript to end on) and lyrics with no sung lines. music-video-cuts.json chains transcribe_audio, analyze_beats and plan_cuts, passing H3's grid (17, 5, 124, 345) and a lead_s variable (default 0.5, 12 frames at 24 fps). Its shots are what music-video.json renders: the agent writes a prompt per shot, drops lyric and kind, and passes the list as shots.

Assessment Probes

Three read-only commands measure a finished cut and say where to look - analyze_shots, analyze_seams, analyze_sync_drift. Each is registered with assessment=True (register_command, dw/tasks/task.py), which is what makes a command a probe - not merely answering JSON, since attribute_voices and check_script answer JSON too and are not probes. Each takes a video (a stored file - asset:, output: or a path, read straight from disk rather than decoded first - or the video an earlier step returned; not a URL, whose download is a bare frame list with no soundtrack) and answers one JSON document: every measurement it took, plus findings (the measurements that crossed a rule in the table below), rules_applied (the rule names the probe checked) and shots_source (where the shot list came from). A probe reads the file streaming - a 64x36 grey thumbnail per frame and the soundtrack, never a full frame list - so it runs on a cut of any length.

Findings are places to look, not verdicts: nothing in the engine acts on one, no run fails for one, and a finding someone has looked at and accepted is simply left alone.

A probe's result must save as JSON:

{
    "task": {
        "command": "analyze_seams",
        "arguments": {
            "video": "output:<identity>/latest/final/cut.mp4"
        }
    },
    "result": { "content_type": "application/json" }
}

Any other content_type (or none) fails validation - a JSON document can only be saved whole under application/json; every other content type would explode it key by key or die trying to write a number.

Shot boundaries come, in order: the step's own shots argument, the shots carried by the video an earlier step returned, the run manifest beside the file, and otherwise the whole file is treated as one shot. shots_source reports which - argument, artifact, manifest, or none.

analyze_shots

Each shot's level and spectral balance, and how far apart the shots sit:

Field Meaning
shots[].name The shot's name
shots[].start_frame / num_frames The shot's frame range, as the shot record gave it
shots[].peak_dbfs Peak level within the shot
shots[].rms_dbfs RMS level within the shot
shots[].crest_db peak_dbfs minus rms_dbfs
shots[].low_dbfs / mid_dbfs / high_dbfs Spectral balance (20-250 Hz / 250-4000 Hz / 4000-20000 Hz), on the same scale as rms_dbfs
shots[].samples Whether the shot's sample span was recorded (carried by the shot record) or derived (scaled from its frames)
shots[].dead_air_seconds The longest run of 50ms windows inside the shot at or below the DEAD_AIR_FLOOR_DBFS threshold (-65 dBFS)
shots[].dead_air_at Where that run starts, in seconds into the file
shots[].dead_air_floor_dbfs The quietest 50ms window measured inside that run - not the threshold. Null when the run is pure digital silence, and null when there is no run at all
rms_range_db The spread between the loudest and quietest voiced shot
has_audio Whether the file carries a soundtrack at all

analyze_seams

Every seam between shots, audio and picture:

Field Meaning
seams[].seam The seam's index (1-based)
seams[].between [previous shot name, next shot name]
seams[].seconds Where the seam sits in the file
seams[].kind cut or dissolve (a dissolve has overlap_frames)
seams[].hard_cut Whether the incoming shot is marked hard_cut: true - concat_videos sets it on a shot opening a seam it butt-joined (a plain cut, or a cut with a seam_fade_ms fade). A seam that got an audio_bleed_ms bleed carries audio_bleed_ms instead of hard_cut
seams[].audio_bleed_ms The bleed the seam actually got (clamped to the material); absent where none ran. The picture is still a cut, so seam_frame_jump stays quiet
seams[].before_shot_rms_dbfs / after_shot_rms_dbfs RMS level of the whole shot either side of the seam
seams[].level_step_db The absolute difference between those two shot levels. Shot against shot, not the audio at the seam's edges: a take's own tail and head can sit 20 dB apart, which is not a step the cut made
seams[].before_rms_dbfs / after_rms_dbfs RMS level of the 0.25 s either side of the seam - what seam_hole's both-sides-voiced guard reads
seams[].floor_dbfs RMS level of the join itself (the fade, or a short window centred on a cut)
seams[].click_db How far a spike at the join peaks above its immediate neighbours
seams[].spectral_shift How much the low/mid/high balance shifts across the seam (0-1)
seams[].frame_delta The largest single-frame picture change across the seam
seams[].typical_delta The larger shot's own typical frame-to-frame change, floored
seams[].jump_ratio frame_delta divided by typical_delta

analyze_sync_drift

How far the soundtrack sits from the picture, shot by shot and over the whole file:

Field Meaning
shots[].name The shot's name
shots[].start_offset_ms How far the audio sits from the picture at the shot's start
shots[].end_offset_ms How far the audio sits from the picture at the shot's end
max_offset_ms The largest end_offset_ms across all shots, by magnitude
video_seconds / audio_seconds Each stream's own duration
length_delta_ms audio_seconds minus video_seconds

Rules

Each rule names the probe and field it reads, how the value is compared to its threshold, and the severity of a crossing:

Rule Probe Field Threshold Severity
shot_level_spread analyze_shots rms_range_db >= 6.0 dB warn
seam_level_step analyze_seams level_step_db > 3.0 dB warn
seam_click analyze_seams click_db > 12.0 dB warn
seam_hole analyze_seams floor_dbfs < -50.0 dBFS warn
seam_frame_jump analyze_seams jump_ratio > 25.0 info
shot_dead_air analyze_shots dead_air_seconds > 0.4 s warn
sync_drift analyze_sync_drift end_offset_ms > 40.0 ms (magnitude) warn
sync_length analyze_sync_drift length_delta_ms > 40.0 ms (magnitude) warn

Three rules carry a guard beyond the threshold: seam_hole only fires while both sides of the seam are voiced above -30 dBFS (a quiet join between two quiet shots is not a hole, it's a pause the shots themselves hold); seam_frame_jump is skipped at a seam whose incoming shot is marked hard_cut: true - a cut meant as a cut; and shot_dead_air is skipped inside a shot whose own rms is at or below -30 dBFS - a shot that is quiet throughout, on purpose, rather than one holding a gap. A gap the guard lets through is what slice_audio -> loop_audio -> mix_audio is for: cut a room-tone bed from the take, loop it to the gap's length, and mix it under the line rather than leaving the drop silent.

A shots record reaching past the file's own length is a separate finding, shot_span_overrun, on all three probes - not a threshold crossing, since the engine clips the record to the file before any of the rules above run. validate_workflow reports the same mistake ahead of the run when the video's length is already knowable (a shots argument against an asset:/literal video); a previous_result:/output: video not yet written is left to the finding.

list_tasks names the probes in their own assessment list, alongside commands, so a caller looking for a way to check a cut can find them without reading every command's schema. The list is exactly the commands declared assessment=True; attribute_voices stays in commands only, since it analyzes a song rather than checking a cut.

Stem separation

Split a mix into its htdemucs stems - vocals, drums, bass and other - each returned as its own audio result. It is the separator attribute_voices runs (one shared loader and cache), exposed for its own sake: transcribing the vocals stem aligns sung lyrics better than transcribing the full mix, and drums + bass + other is the instrumental to put under dialogue as a score bed.

{
    "name": "stems",
    "task": {
        "command": "separate_stems",
        "arguments": { "audio": "asset:song.mp3" }
    },
    "result": { "content_type": "audio/wav", "file_name": "stems" }
},
{
    "name": "lyrics",
    "task": {
        "command": "transcribe_audio",
        "arguments": { "audio": "previous_result:stems.vocals" }
    },
    "result": { "content_type": "application/json" }
}
Argument Required Description
audio Yes Path or URL of an audio file (or of a video file, whose soundtrack is taken), a waveform from a previous step, or an earlier step's video generated with a soundtrack
sample_rate With a waveform Sample rate of a directly passed waveform (files carry their own)

The result saves one file per stem, <file_name>-vocals, -drums, -bass and -other, and a later step reads one as previous_result:<step>.vocals. Stems are 44.1 kHz stereo, whatever the source was, and sum back to the mix. The model is fixed (htdemucs, run in fp32), its weights download on first use, and it needs the demucs package. On MPS a failed separation falls back to the CPU and warns (separation_cpu_fallback).

Voice attribution

Which reference voice sings each line of a song, by timbre - staging lip-sync shots for a generated song needs its section-to-singer map, and a generation model (MiniMax Music3 included) does not hand one back. Pitch cannot stand in for it: a tenor and a mezzo share a range, and a pitch heuristic has called a tenor female.

attribute_voices separates the vocal stem (htdemucs), reduces every line and every voice's reference to its voiced frames, embeds each with speechbrain's ECAPA speaker encoder (spkrec-ecapa-voxceleb), scores every line against every voice by cosine, and rolls lines up into named windows (shots, say) by the voiced seconds they overlap. It decides nothing: the argmax is always reported, alongside the margin and how much of the line was voiced, so a weak answer is visible as weak rather than silently accepted.

{
    "task": {
        "command": "attribute_voices",
        "arguments": {
            "audio": "asset:song.mp3",
            "voices": {
                "lena": [{"start_seconds": 4.0, "duration_seconds": 6.0}],
                "marcus": [{"start_seconds": 32.0, "duration_seconds": 6.0}]
            },
            "windows": [
                {"name": "shot-1", "start": 0.0, "end": 12.0},
                {"name": "shot-2", "start": 12.0, "end": 24.0},
                {"name": "shot-3", "start": 24.0, "end": 40.0}
            ]
        }
    },
    "result": { "content_type": "application/json" }
}
Argument Required Description
audio Yes The song - a path, asset:/output: reference, or an earlier step's audio or video
voices Yes Each voice's name mapped to its reference: a list of {start_seconds, duration_seconds} (or {start, end}) spans into audio, or a path/asset: of a separate clip. At least 2 voices; names follow the variable-name pattern; each voice's reference must total at least min_reference_seconds
lines No Spans to attribute, {start, end, text?} (a transcribe_audio transcript, or its chunks list, drops in directly) or {start_seconds, duration_seconds, text?}. A zero-length line (start = end, as Whisper's word timestamps sometimes give) comes back with voice: null rather than refusing the call. Omitted: the song is cut into fixed windows of window_seconds
windows No Named spans to roll lines up into, {name, start, end}. Omitted: mirrors the fixed windows when lines is also omitted, otherwise none
window_seconds No Length of the fixed windows used without lines (default 2.0)
min_reference_seconds No Least total reference length per voice; a shorter one is refused by name (default 3.0)
separate No Isolate the vocal stem with htdemucs before embedding (default true); false for audio that is already a dry vocal
device No Where the models run

A literal voices is checked at validation as well as on the step: fewer than two voices, a bad name, a malformed span, or a span-list reference under min_reference_seconds is refused at steps[i].task.arguments.voices, and a bare path as a voice meets the same location policy as audio. What needs the song itself - a span past its end, a clip's voiced length - is the step's to refuse.

The result: voices (the names), separated, duration_seconds, lines[] (start, end, text, scores, voice, margin, voiced_seconds, uncertain, reason), windows[] (name, start, end, voiced_seconds, share, voice, uncertain, reason), reference_similarity (pairwise cosine between the voices' references), voiced_floor_dbfs (the floor this song's stem was read against), warnings and the thresholds compared against.

The voiced floor is relative to the stem rather than a fixed level: a quiet sung verse under a loud chorus sits 20 dB or more beneath it, and a fixed -40 dBFS floor dropped such a verse as unvoiced and refused it as a reference. Separation leaves near-silence between phrases, so the floor sits well above that and well below the singing.

Threshold Value What crossing it does
voiced_level_percentile 95 The stem's level is this percentile of its 20 ms frames' rms
voiced_floor_below_level_db 35.0 dB A frame at or above the stem's level less this is voiced
voiced_floor_min_dbfs -60.0 dBFS The floor never drops below this, however quiet the stem
min_voiced_seconds 0.5 s A line or window under this much voiced time has no voice - voice: null, uncertain: true
uncertain_margin 0.05 A line whose best score beats the runner-up by less than this is uncertain - the argmax is still reported
uncertain_share_margin 0.2 A window whose leading voice's share beats the runner-up's by less than this is uncertain - a duet line, or a window straddling a hand-over
piece_seconds 2.0 s A line longer than this is also scored in pieces of this length
mixed_line_share 0.25 A line whose confidently attributed pieces give a second voice at least this share of their voiced time is uncertain, and its reason names each voice's seconds - one embedding of two singers can name either of them with a wide margin
voices_too_similar 0.8 Two references scoring above this against each other make every line between them a weak answer whatever its scores say

A voices_too_similar pair is reported in reference_similarity, added to warnings, and emitted onto the job's warnings (voices_too_similar) - picking better-separated reference spans is the fix, not reading past it.

Both models are fixed - there is no model-name argument - and both run in fp32 (their STFT front ends are fp32-only in practice, and both are small enough that it costs nothing). Weights download on first use - htdemucs from Hugging Face (adefossez/HTDemucs via demucs 4.1), ECAPA from speechbrain - and are cached between calls like any other model. On MPS a separate: true run that fails in htdemucs falls back to the CPU and warns (separation_cpu_fallback) rather than failing the step.

Example: attribute-lines.json — Which singer sings each line of a song, and when: transcribe_audio with timestamps, its transcript handed whole to attribute_voices as lines.

Checking the lip-sync target

A sung multi-shot piece can put the right song on the wrong mouth: the shot lip-syncs whoever the prompt or the reference made most prominent, not the singer the song gives that line to. Listening does not catch it, and one frame per shot does not either. This check lines each sung line up with the picture and looks at whose mouth is open on it.

  1. Run attribute-lines on the song with each singer's reference spans as voices. Every line comes back with numeric start and end seconds into the song, its voice, margin and uncertain flag.
  2. Map each line's song times onto the cut's timeline (the rules below).
  3. For every line with a non-null voice, take two moments: start + 0.3 s, where the mouth has opened on the line's first syllable, and the line's midpoint.
  4. Make one get_output_frames call per face: at the moments, crop the face's region, hear: 1.0. At most 32 moments per call, so a long song takes more than one call per face.
  5. On each line, the singer voice names should be the face with an open mouth. Another face singing it is the wrong lip-sync target; both mouths open on a solo line is a shot that lip-syncs everyone.
  6. Skip the lines that are uncertain or have a null voice, and name them in the report - a skipped line is unchecked, not passed.

Song time to cut time. attribute-lines answers in seconds into the song it was given; the frames are read from the cut. The offset between the two depends on how the piece was assembled:

  • minimax/music-video: song time is cut time unless trim_frames > 0. The template slices the song lead_frames before each shot's start_frame, renders num_frames, keeps the shot's cut_frames with trim_video, joins the trims with concat_videos at trim_frames: 0, and lays the whole song back with pair_audio at fit: video, so no offset applies. A trim drops frames at every seam, and the song no longer lines up past the first one.
  • join_into_song: add start_frame / fps − cue_seconds to every line time, where start_frame is the first sung shot's, from the step's shots output. That shot's frame 0 is where song time cue_seconds lands.
  • Any other assembly: transcribe the final cut's own track instead, so the times are already the cut's. Attribution is weaker there - dialogue ducked under the song is in the mix too.

Where it misleads:

  • A duet line - two singers on one line, or call and response inside one Whisper segment - comes back uncertain: its 2 s pieces name more than one voice, and the reason says how many seconds each holds. timestamps: "word" splits it into words, each attributed on its own, at the cost of noisier scores. A zero-length word comes back with a null voice, like any line too short to embed.
  • Whisper can return one segment over most of a song, its text a repeated hallucination, when the accompaniment drowns the words. Such a line is a duet line by the rule above whenever two singers are in it; re-run with timestamps: "word", or transcribe the vocals stem of separate_stems instead.
  • start + 0.3 s is a heuristic: a singer who opens late or holds a pickup note is still closed-mouthed there. That is why the midpoint frame is read too; trust the two together, not either alone.

ingredients_grid

Lay individual images - a character, a prop, a location - out on one canvas as a reference sheet. It is frame_grid's counterpart for separate images, and builds the still ltx2/reference-sheet conditions on. Pure PIL.

{
    "name": "sheet",
    "task": {
        "command": "ingredients_grid",
        "arguments": {
            "images": ["asset:hero.png", "asset:sword.png", "asset:castle.png"],
            "width": 768,
            "height": 448
        }
    },
    "result": { "content_type": "image/png" }
}
Argument Default Description
images required Images in reading order: references, paths or {"location": ...} dicts; gather: of a for_each step splices in
width, height 768, 448 Canvas size in pixels
layout auto rows, panels or auto
fit contain contain scales each image to fit whole and pads with background; cover fills its cell and crops the overflow
gap 8 Pixels between cells, and kept clear of the canvas edge
background white Canvas colour, a name or #hex
max_images 12 More images than this is refused, never silently dropped

rows keeps every image at its own aspect ratio and justifies each row to the canvas width. Which images share a row is searched: every way of cutting the list, in order, into rows is scored on wasted canvas, overflow past the height, and spread of row heights. panels gives every image an equal grid cell, the column count chosen to waste the least canvas, with a short last row centred. auto uses rows for up to 3 images and panels for more. Transparency is flattened onto white. Returns one RGB image of exactly width x height.

To build the sheet inside ltx2/reference-sheet, add this step first and point static_sheet's video at it ("video": "previous_result:sheet" in place of the reference_sheet asset). Keep width and height the clip's own, so the sheet reads at the size the model was trained on:

{
    "name": "sheet",
    "task": {
        "command": "ingredients_grid",
        "arguments": {
            "images": ["asset:hero.png", "asset:sword.png", "asset:castle.png"],
            "width": "variable:width",
            "height": "variable:height"
        }
    },
    "result": { "content_type": "image/png", "save": false }
}

Data Gathering

gather_images

Load images from URLs and/or file glob patterns:

{
    "task": {
        "command": "gather_images",
        "arguments": {
            "urls": ["https://example.com/a.jpg", "https://example.com/b.jpg"],
            "glob": "./images/*.jpg"
        }
    }
}

Returns a list of images that can be referenced by later steps with previous_result:.

gather_videos

Same as gather_images but for video files. Each video comes back as one artifact holding its frames and whatever audio was muxed alongside them, so a step referencing this one iterates over videos rather than over frames.

To join videos that are already on disk, give their paths to concat_videos directly rather than gathering them first: a previous_result reference to a gather step fans the consuming step out over the gathered videos instead of handing it all of them at once.

gather_inputs

Pass through arguments directly. Useful for organizing data flow.

Videos in image tasks

Every image command - the upscalers, face restoration, segmentation and the image processors - takes a video where it takes an image: an AudioVideo from a generation, concat_videos or dissolve_videos step, or a frame array from video_frames. The command runs over the frames one at a time and returns one video artifact, its soundtrack carried through untouched, so a generated clip can be upscaled without losing what was generated alongside it:

{
    "task": {
        "command": "upscale",
        "arguments": {
            "image": "previous_result:generate_video",
            "model_name": "Kim2091/UltraSharp"
        }
    },
    "result": { "content_type": "video/mp4", "fps": 24 }
}

Captioning (image_to_text) is the exception - describe a frame, taken with get_first_frame, rather than a video.

A file path or asset:/output: reference to a video is not accepted here (grade, whose input is media, is the exception), even though the route above takes a video from a step - route it through video_frames first: an image command's image argument must be an AudioVideo, a frame array, or a still image.

Image Upscaling

Upscale images using spandrel-compatible super-resolution models (ESRGAN, SwinIR, HAT, DAT, and 40+ other architectures). Models are auto-detected from weight files.

{
    "task": {
        "command": "upscale",
        "arguments": {
            "image": "previous_result:generate",
            "model_name": "Kim2091/UltraSharp",
            "filename": "4x-UltraSharp.pth"
        }
    }
}
Argument Required Description
image Yes PIL Image or previous_result: reference - a video runs frame by frame, see Videos in image tasks
model_name Yes HuggingFace repo ID or local file path
filename No Specific weight file in a HF repo (auto-detected if only one)
tile_size No Tile size for large images (default: 512)
tile_overlap No Overlap between tiles in pixels (default: 32)

Large images are automatically tiled to avoid GPU memory issues. Models can be loaded from HuggingFace Hub repos or local .pth/.safetensors files.

Examples:

Diffusion Upscaling

Upscale images using Stable Diffusion upscale pipelines. Text-guided upscaling with better detail recovery than traditional super-resolution, especially for faces and textures.

Two modes are available:

  • x4 (default): StableDiffusionUpscalePipeline — 4x upscale via stabilityai/stable-diffusion-x4-upscaler
  • x2: StableDiffusionLatentUpscalePipeline — 2x upscale via stabilityai/sd-x2-latent-upscaler
{
    "task": {
        "command": "diffusion_upscale",
        "arguments": {
            "image": "previous_result:generate",
            "prompt": "high quality, detailed",
            "negative_prompt": "blurry, low quality, artifacts",
            "mode": "x4"
        }
    }
}
Argument Required Description
image Yes PIL Image or previous_result: reference - a video runs frame by frame, see Videos in image tasks
prompt No Text guidance for upscaling (default: "")
negative_prompt No Negative text guidance (default: none)
mode No "x4" or "x2" (default: "x4")
model_name No Override the default model for the selected mode
num_inference_steps No Denoising steps (default: 25)
guidance_scale No Classifier-free guidance scale (default: 9.0)
noise_level No Noise level for x4 mode (default: 20, ignored for x2)

Examples:

  • upscale-diffusion.json — Upscale any existing image. mode selects which: x4 (the default) reaches 2048px, x2 reaches 1024px through the latent upscaler.
  • upscale-diffusion.json — Prompt-guided upscale of an image you already have, with no generation step.

Face Restoration

Restore and enhance faces in images using spandrel-compatible face restoration models (GFPGAN, CodeFormer, RestoreFormer). Uses facexlib for face detection and alignment, then runs each detected face through the restoration model.

{
    "task": {
        "command": "restore_faces",
        "arguments": {
            "image": "previous_result:generate",
            "model_name": "leonelhs/gfpgan",
            "filename": "GFPGANv1.4.pth"
        }
    }
}
Argument Required Description
image Yes PIL Image or previous_result: reference - a video runs frame by frame, see Videos in image tasks
model_name Yes HuggingFace repo ID or local file path
filename No Specific weight file in a HF repo (auto-detected if only one)
upscale_factor No Background upscale factor (default: 1, no upscaling)
face_size No Cropped face size in pixels (default: 512)
use_parse No Use face parsing for better blending (default: true)
only_center_face No Only restore the largest/center face (default: false)
detection_resize No Resize shorter side for detection speed (default: 640)
eye_dist_threshold No Skip faces with eye distance below this (default: 5)
upsample_img No Pre-upscaled background image (e.g., from a prior upscale step)

Models are loaded via spandrel, so any .pth/.safetensors face restoration weights work. CodeFormer requires pip install spandrel-extra-arches (non-commercial license).

Example: restore-faces.json — Generate a portrait, then restore faces with GFPGAN v1.4.

Combining with Upscaling

You can chain upscaling and face restoration. Generate first, upscale the background, then paste restored faces onto the upscaled image:

{
    "steps": [
        {
            "name": "generate",
            "pipeline": { "..." : "..." },
            "result": { "content_type": "image/jpeg" }
        },
        {
            "name": "upscale",
            "task": {
                "command": "upscale",
                "arguments": {
                    "image": "previous_result:generate",
                    "model_name": "Kim2091/UltraSharp",
                    "filename": "4x-UltraSharp.pth"
                }
            },
            "result": { "content_type": "image/jpeg" }
        },
        {
            "name": "restore",
            "task": {
                "command": "restore_faces",
                "arguments": {
                    "image": "previous_result:generate",
                    "model_name": "leonelhs/gfpgan",
                    "filename": "GFPGANv1.4.pth",
                    "upscale_factor": 4,
                    "upsample_img": "previous_result:upscale"
                }
            },
            "result": { "content_type": "image/jpeg" }
        }
    ]
}

This gives the best results: the super-resolution model handles background detail while the face model handles facial features, composited together at the upscaled resolution.

Object Segmentation

Detect and segment objects using text prompts via GroundingDINO + SAM2. Returns a binary mask image suitable for inpainting workflows.

{
    "task": {
        "command": "segment",
        "arguments": {
            "image": "previous_result:input_image",
            "prompt": "dog"
        }
    },
    "result": { "content_type": "image/png" }
}
Argument Required Description
image Yes PIL Image or previous_result: reference - a video runs frame by frame, see Videos in image tasks
prompt Yes Text description of object(s) to detect (e.g., "dog", "red car")
model_name No GroundingDINO model ID (default: IDEA-Research/grounding-dino-base)
sam_model_name No SAM2 model ID (default: facebook/sam2-hiera-large)
threshold No Detection confidence threshold (default: 0.3)
invert No Invert the output mask (default: false)

Returns a grayscale PIL Image (mode "L") — white (255) for detected objects, black (0) for background. Use with inpainting pipelines like FluxFillPipeline.

Examples:

Image Captioning

Generate text captions from images using a vision-language model.

Transformers 5 removed the dedicated image-to-text pipeline this task used to build, along with the BLIP/ViT-GPT2/GIT captioning models that ran on it. Captioning now goes through the same image-text-to-text pipeline as any other VLM, so model_name needs a vision-language model (SmolVLM, Qwen2.5-VL, LLaVA, etc.) and prompt is a question put to the model rather than a text fragment to continue.

{
    "task": {
        "command": "image_to_text",
        "arguments": {
            "image": "previous_result:input_image"
        }
    },
    "result": { "content_type": "text/plain" }
}
Argument Required Description
image Yes PIL Image, URL/path, or previous_result: reference
model_name No HuggingFace vision-language model ID (default: HuggingFaceTB/SmolVLM-256M-Instruct)
prompt No What to ask about the image (default: Describe this image.) — ask a narrower question for a narrower caption
system_prompt No System instruction for the model
max_new_tokens No Maximum tokens to generate (default: 50)

The default model is deliberately tiny, matching the footprint of the old captioning default; it produces short, plain captions. Point model_name at something larger for detail.

Returns a caption string. Save as text/plain for .txt output, or pass to a downstream step via previous_result: as a prompt for image generation.

For a detailed caption, hand the image to text_generation with a question and a larger vision-language model; that is what describe-and-regenerate.json does ahead of its prompt expansion.

Examples:

Composing Text

Assemble one block of text out of parts written once. A multi-shot workflow says the same things about its characters in every shot — who they are, what they are wearing, what their voice sounds like — and the engine deliberately has no string interpolation to splice them in with (see the no-interpolation rule in the workflow guide). Composition is the way round it: a part is a whole value, and compose_text joins parts in order.

{
    "name": "shot_1_prompt",
    "task": {
        "command": "compose_text",
        "arguments": {
            "parts": [
                "variable:character_a_bible",
                "variable:character_a_voice",
                "variable:shot_1_action"
            ],
            "separator": "\n\n"
        }
    }
},
{
    "name": "shot_1",
    "pipeline": { "arguments": { "prompt": "previous_result:shot_1_prompt" } }
}
Argument Required Description
parts Yes The parts to join, in order — each a whole value, usually a variable:, prompt: or previous_result: reference. Numbers are written out; null is dropped, so an optional part can be a variable left null
separator No What goes between the parts (default: a blank line, the paragraph break the prompt formats use)
skip_empty No Drop parts that are null or blank (default true). With it off, an empty part still contributes its separator

A part that is neither text nor a number is an error, not a coercion: it means the reference in that position resolved to something other than the text meant.

The parts are positional. A named form ("{bible} says {line}") would be the interpolation the engine does not have, one layer down — so a character bible is a variable named by every shot that needs it, and a voice string written once is checked by being the same value rather than by being compared.

Extracting Sections

Reduce generated text to a known set of labelled sections, dropping anything else:

{
    "task": {
        "command": "extract_sections",
        "arguments": {
            "text": "previous_result:expand",
            "sections": ["integrated_multimodal_description", "overall_soundscape", "non_diegetic_music"]
        }
    },
    "result": { "content_type": "text/plain" }
}
Argument Required Description
text Yes The generated text, usually a previous_result: reference
sections Yes Section labels to keep, in the order they should appear
keep_preamble No Keep any text before the first label (default: true)

A section runs from its label: to the end of that paragraph, so a blank line ends one and a single newline does not — a field holding one line per item stays intact. Repeats are dropped, missing sections are skipped, and text with no recognised label is returned unchanged.

This exists because a model asked for a rigid format usually produces it and then keeps going — restating the description, appending a summary, or looping until it runs out of tokens. Prompting against that is unreliable, and at small model sizes adding rules to an already long specification can make adherence worse. Trailing text is not free either: a prompt is conditioning, and a pipeline that does not truncate spends memory and attention on whatever arrives. Keeping the fields that were asked for is deterministic where prompting is not.

The built-in h3_context_ir workflow applies this to its own output, so a workflow delegating to it receives only the fields MiniMax H3 expects.

Text Generation / Prompt Expansion

Generate or expand text using a local language model. Useful for expanding short prompts into detailed image generation prompts, rewriting text, or other text-to-text tasks.

{
    "task": {
        "command": "text_generation",
        "arguments": {
            "prompt": "a cat on a windowsill",
            "system_prompt": "You are a helpful AI assistant that creates detailed prompts for text to image generative AI. When supplied input generate only the prompt, no other text."
        }
    },
    "result": { "content_type": "text/plain" }
}
Argument Required Description
prompt Yes The user message or short prompt to expand/transform
system_prompt No System instruction for the model (e.g., "expand this into a detailed image prompt")
model_name No HuggingFace model ID (default: Qwen/Qwen2.5-1.5B-Instruct, or HuggingFaceTB/SmolVLM-256M-Instruct when an image is supplied)
image No PIL Image, URL/path, or previous_result: reference — see below
repetition_penalty No Vision path only (default: 1.15) — see below
generate_kwargs No Anything else to pass to the model's generate() — no_repeat_ngram_size, top_p, min_new_tokens. Merged last, so it overrides the settings above
max_new_tokens No Maximum tokens to generate (default: 500)

Writing a prompt from a picture

Supplying image switches the task to a vision-language model, so the generated text describes what is actually in the picture instead of what the prompt guesses is there. model_name must then name a VLM — a text-only model cannot be loaded as one.

{
    "task": {
        "command": "text_generation",
        "arguments": {
            "prompt": "Write a video prompt that starts from this picture.",
            "image": "previous_result:input_image",
            "model_name": "Qwen/Qwen3-VL-4B-Instruct"
        }
    },
    "result": { "content_type": "text/plain" }
}

This matters most ahead of an image-conditioned generation step. Those pipelines pin the supplied picture as the first frame, so a prompt written without seeing it will describe a scene the keyframe contradicts and the two conditionings pull against each other. Pass the same image to both and the prompt agrees with the frame it opens on.

A vision model is large enough to be worth releasing before the generation model loads — see release_models in the workflow guide.

Generation stays greedy so a workflow reproduces, but greedy decoding against a long, rigid format specification makes these models loop — emitting a complete answer and then repeating its closing sections until the token budget runs out. The vision path applies a repetition_penalty of 1.15 to stop that. Measured on Qwen3-VL against the MiniMax H3 prompt spec, 1.05 still looped through the whole budget while 1.15 ended on its own at a length matching the format's own guidance. Raise it if a model still repeats itself, or set 1.0 to disable.

A penalty reins the looping in but does not guarantee the model stops where the format ends; for that, trim the output with extract_sections below.

There is a limit to what a small model will follow. Against the MiniMax H3 spec, neither Qwen3-VL-4B nor 8B produces the <d>[Language]...</d> dialogue tag or the (S1) speaker ids, whether the idea implies speech or supplies the line verbatim; the 8B is worse on layout, capitalising its section labels. Showing a complete worked example does produce them - by copying the example word for word, which is useless - and a placeholder skeleton does not produce them at all. The visual description these models write is grounded and usable; the dialogue markup is not. Write prompts by hand where a subject has to speak.

For anything the arguments above do not cover, generate_kwargs goes straight to generate():

"arguments": {
    "prompt": "a cat on a windowsill",
    "generate_kwargs": { "no_repeat_ngram_size": 25 }
}

It is merged after everything else, so it can override repetition_penalty and the sampling settings as well as add to them.

Examples:

Speech Generation

Speak a line of text with a local text-to-speech model. The result is a waveform carrying the rate its model generated at, so it composes with slice_audio, fade_audio and pair_audio directly (concat_videos and dissolve_videos join videos — pair the track onto a video first).

{
    "task": {
        "command": "generate_speech",
        "arguments": {
            "text": "The way ahead is longer still.",
            "voice_preset": "v2/en_speaker_6"
        }
    },
    "result": { "content_type": "audio/wav" }
}
Argument Required Description
text One of text/messages The line to speak
messages One of text/messages Chat-templated input for a model such as VibeVoice that takes a conversation rather than a bare string — a list of {"role": ..., "content": ...} dicts, passed straight through as the pipeline's text_inputs so the model's own chat template applies. A model with no chat template configured (Bark and friends) raises when handed this instead of text
model_name No HuggingFace model ID (default: suno/bark-small)
voice_preset No The speaker, for a model with presets — v2/en_speaker_0 through v2/en_speaker_9 for Bark. A model with no processor (a single-voice model such as facebook/mms-tts-eng) refuses a voice_preset with an error rather than ignoring it
speaker_embedding No A reference audio file (typically an asset: reference) whose voice a SpeechT5 model should speak in. Reduced to an x-vector with speechbrain's spkrec-xvect-voxceleb and injected into forward_params as speaker_embeddings. A model that isn't SpeechT5 refuses it the same way a single-voice model refuses voice_preset
forward_params No Passed to the model's forward/generate call
generate_kwargs No Ad-hoc generation settings for a generative model — temperature, do_sample

The command's own default is Bark because its voice presets give distinct speakers, which is what two characters in a scene need; facebook/mms-tts-eng is a quarter the size and a good override where one voice will do. generate-speech.json ships with facebook/mms-tts-eng as its own default instead, since a template is usually a single voice. voice_preset is a preprocessing argument — it selects the speaker before generation rather than parameterizing it — so naming it here is what makes it reach the processor. Passed through forward_params it would be dropped and every character would sound the same.

speaker_embedding is the same kind of preprocessing argument for a SpeechT5 model (microsoft/speecht5_tts), which conditions its voice on an x-vector rather than a preset name. A VITS model's speaker is different again — a plain speaker_id int, passed through forward_params unchanged, since it was never a preprocessing argument and needs no argument of its own here.

The result needs no sample_rate. A generated track carries the rate its model produced it at, and that beats the 44100 default; declaring one still wins over both, for a track whose rate was reported wrong. Every TTS model runs at a different rate, so a declared rate that does not match plays the speech at the wrong speed and pitch without ever failing.

Generating a voice to condition on

The role this earns its place in is voice timbre reference, not the track a mouth follows. A MiniMaxH3AudioReference is audio H3 reads, not audio it plays: it takes a few seconds of a voice to fix timbre, pitch and delivery while H3 still generates the line, and a mouth asked to follow a whole track through a reference follows it loosely. That reference is still the better lip sync: music-video and the match_audio chain templates pass each shot's slice as one. Holding the track with hold_audio (the workflows guide, "H3: generating to a held soundtrack") keeps it exact, but on a sung track the mouth followed it less well than the reference (#619's A/B, measured in that section), so hold is opt-in. Build the reference with from_previous_result and the clip's own sample rate comes across with it:

"references": [
    {
        "reference_type": "diffusers.modular_pipelines.minimax_h3.MiniMaxH3AudioReference",
        "from_previous_result": "voice"
    }
]

Referencing the same preset in every shot of a scene makes a character's voice a conditioning signal rather than a prose description that has to land identically a dozen times. The other honest uses are a voice that must be matched — a specific delivery H3 will not produce from description alone — and narration over shots where nothing has to lip-sync to it, muxed with pair_audio.

A speech model is worth releasing before a video model loads — set release_models on the step, as in the example below.

Examples:

Speech Transcription

Transcribe spoken audio to text with a local Whisper-class model. The word-correctness of a TTS deliverable — a dropped line, a mid-sentence truncation — can only be inferred from duration and timing arithmetic without this; transcribe_audio checks it directly against the text the deliverable was supposed to speak.

{
    "task": {
        "command": "transcribe_audio",
        "arguments": {
            "audio": "previous_result:speak"
        }
    },
    "result": { "content_type": "text/plain" }
}
Argument Required Description
audio Yes Path or URL of an audio file (or of a video file, whose soundtrack is taken), a video with a soundtrack, or a waveform — usually a previous_result: reference
sample_rate No Sample rate of a waveform passed directly
model_name No HuggingFace model ID of a Whisper-class ASR model (default: openai/whisper-base)
timestamps No "segment" or "word" to get chunk timings instead of plain text (see below)

Multi-channel audio is downmixed to mono and resampled to 16 kHz before transcription, since that is what a Whisper-class model is trained on; the source audio itself is untouched. By default the result is plain text, read with MCP's get_output_text.

Set timestamps to "segment" or "word" to get chunk timings instead — a music video cut to the lyric, or a dialogue shot checked against its line, needs the times Whisper already produces past 30 s rather than the collapsed string. The result becomes {"text": ..., "chunks": [{"start": ..., "end": ..., "text": ...}, ...]}, so the step's result.content_type must be "application/json" rather than "text/plain", and it's read with MCP's get_output_text (JSON results are text). A clip under 30 s asks Whisper for timestamps explicitly when timestamps is set — the 30 s long-form threshold is a separate, unrelated reason to ask. Every start and end is a number of seconds: a chunk the clip ends inside - a song cut mid-line, which Whisper leaves open-ended - ends at the clip's duration, so the transcript drops straight into attribute_voices' lines. "word" can give a zero-length chunk (start = end) on a short or clipped word; attribute_voices reads it as a line with no voice.

To compare a take against the lines it was meant to speak, use check_script rather than reading the transcript by eye.

Example: transcribe-audio.json — Transcribe an audio file to text.

Script check

Whether a dialogue take speaks its script. Confirming it used to mean transcribing the take and comparing the transcript to the script by eye, which missed a word Whisper dropped, an H3 tag read aloud, and a last word cut off by the end of the file. check_script does the comparison and reports where to look.

It transcribes the take with transcribe_audio word timestamps - the same code and model cache, no second ASR path. Every heard word is measured, and one at or below the dead-air floor is discarded as unheard, because Whisper invents words over silence. So is every word of a repetition loop - the same words over and over (Pre-pre-pre-...), which Whisper invents over music or room tone loud enough to pass the floor. H3 markup is stripped from each expected line, and the heard words are aligned to the expected words in order (difflib). Each line scores the similarity of its own expected and heard words. It decides nothing: findings are places to look, not verdicts. It is not an assessment probe (assessment=True); like attribute_voices it answers a JSON document.

{
    "task": {
        "command": "check_script",
        "arguments": {
            "audio": "output:take.wav",
            "lines": ["<d>[English] Where were you?</d>", "I was at the station."]
        }
    },
    "result": { "content_type": "application/json" }
}

The step's result.content_type must be "application/json"; read the answer with MCP's get_output_text.

Argument Required Description
audio Yes The take - a path, asset:/output: reference, or an earlier step's audio or video (its soundtrack is taken)
lines Yes The expected lines in order: a list of strings or {text, shot} objects, shot naming the take's shot the line belongs in. Markup is stripped (below). [] means no speech is expected
shots No The take's shot map, [{name, start_frame, num_frames, start_sample, num_samples}]. Default: the take's own .shots (an earlier step's video), else the run manifest or keep_output sidecar beside an asset:/output: file - resolved as the assessment probes resolve them (assess.resolve_shots). The argument wins over both
similarity No Least similarity, 0..1, of a line's heard words to its expected words before it is a line_mismatch (default 0.85)
model_name No HuggingFace ID of a Whisper-class ASR model (default openai/whisper-base)
sample_rate No Sample rate of a waveform passed directly
device No Where the ASR model runs

Markup stripped from an expected line: <d>[Language] ...</d>, <scenetrans>, <cutoff>, [unclear], speaker IDs (S1) and (S1,S2), and any other <tag> or [tag]. A plain parenthetical stays dialogue. The words of the stripped tokens (cutoff, unclear, english, s1) are what tag_spoken listens for; a tag word heard is reported there and left out of the line's similarity, so it is never a line_mismatch too.

Rule Fires when value / threshold at
line_mismatch A line's similarity is below similarity; a dropped line comes back with heard: "" The line's similarity / similarity {line, seconds, word} - the line's first heard word, or where it should have been
tag_spoken A word from the line's stripped markup was heard, and is not also a word of the line's dialogue The word heard / none {line, seconds, word}
line_clipped_at_end A line's last heard word ends inside the final CLIP_TAIL_SECONDS of its shot (when the line has one) or of the file, and that tail is still above GUARD_FLOOR_DBFS The tail's level in dBFS / GUARD_FLOOR_DBFS {line, seconds, word, shot} - shot when the line has one
speech_where_silent lines is [] and words were heard above the floor and outside a repetition loop Number of words heard / 0 {line: null, seconds, word} - the first; shot when shots are known
speech_in_silent_shot Shots are known, at least one line names a shot, and words were heard above the floor and outside a repetition loop in a shot no line names. A word is in the shot its midpoint falls in; one a line in another shot matched (a word straddling the cut) is that line's, not counted Number of words heard in the shot / 0 {line: null, seconds, word, shot} - the shot's first counted word, seconds no earlier than the shot's start

Every finding is severity warn; at.shot is set whenever the shot is known. A rule that does not apply to the call is listed in rules_skipped with its reason: speech_in_silent_shot is skipped when no shots are known (line_clipped_at_end then looks at the file's end only), when the shots carry neither a sample span nor a frame span the take's frame rate can place, or when no line names a shot (no shot is known to be meant silent).

The thresholds are module constants in dw/tasks/script_check.py, deliberately not in RULES (dw/assessment_rules.py): every rule there must fire on a synthetic file that holds no script. The answer echoes them under thresholds.

Constant Value What crossing it does
DEFAULT_SIMILARITY 0.85 A line below this similarity is a line_mismatch (the similarity argument overrides it)
GUARD_FLOOR_DBFS -65 dBFS (= DEAD_AIR_FLOOR_DBFS, shot_dead_air's floor) A heard word whose loudest window is at or below it is discarded as unheard; also the level a clipped tail must exceed
GUARD_WINDOW_SECONDS 0.05 s (= DEAD_AIR_WINDOW, shot_dead_air's window) The window a heard word is measured in - its loudest one, so a span overhanging a pause does not average a real word away
CLIP_TAIL_SECONDS 0.25 s The final stretch of the file a line's last word must end in to be line_clipped_at_end
REPEAT_RUN_MIN 6 Heard words repeating the same unit this many times over, back to back, are a repetition loop: every chunk holding one is discarded (reason: "repetition"), however loud
REPEAT_MAX_PERIOD 4 words The longest repeating unit a loop is looked for in - pre pre pre has period 1, thank you thank you period 2

The result: findings, lines[] (expected, heard, similarity, start, end, shot), discarded[] (text, start, end, level_dbfs, reason - below_floor or repetition), unmatched[] (heard words aligned to no line), transcript, model_name, rules_applied, rules_skipped[] (rule, reason) and thresholds. A line's shot is the shot it names, else the shot its heard words overlap most (null when none is known). shots is the map as placed - [{name, start, end}] in seconds on the take, or null when none is known - and shots_source says where it came from: argument, artifact, manifest or none.

How the alignment reads a take:

  • A dropped line comes back with heard: "" and a line_mismatch at where it should have been.
  • An added line's words go to unmatched and lower no line's similarity.
  • A line spoken out of order reads as a mismatch on that line.
  • Whisper mishears names and numbers (2 for two), so listen before re-rolling a take on a mismatch.
  • A quiet line can be discarded by the guard. Discarded words are reported under discarded, not hidden.

A malformed lines - not a list, an entry without text, an unknown key in a line object - is refused at validation, at steps[i].task.arguments.lines, as is a similarity outside 0..1. A line naming a shot the map lacks is refused, listing the map's shots, and so is a line naming a shot the map holds twice; a literal shots and literal lines are checked against each other at validation, and a shots or lines that is a reference is checked when the step runs.

Example: check-script.json — Check a take against its script and save the answer as JSON. Its defaults are openai/whisper-large-v3-turbo at similarity 0.6, not the task's openai/whisper-base at 0.85: on the two-model measurement (#609, job af5de241ca83), turbo scored the take's correct lines 0.71 to 1.00 and its wrong or missing lines 0.00, where base scored correct lines as low as 0.43 on mishearings (O'Connor, 415). Turbo writes numbers as words (four fifteen), so a script with digits costs similarity under it.

Frame Interpolation

Increase video frame rate using RIFE (Real-Time Intermediate Flow Estimation). Takes a video and inserts intermediate frames between each pair. The result is one video artifact without a soundtrack - the frame count changed, so pair_audio is how the original track comes back. interpolate-frames.json shows the interpolation itself.

{
    "task": {
        "command": "interpolate_frames",
        "arguments": {
            "video": "previous_result:generate_video",
            "multiplier": 2
        }
    },
    "result": { "content_type": "video/mp4", "fps": 60 }
}
Argument Required Description
video Yes The frames - a frame list, a frame array, or an audio+video pair from a concat or dissolve step (its audio is dropped) - usually a previous_result: reference
multiplier No Frame count multiplier: 2, 4, or 8 (default: 2)
model_name No HuggingFace repo with RIFE v4.13 weights (default: imaginairy/rife-interpolation)
filename No Weights filename within the repo (default: rife-flownet-4.13.2.safetensors)

Uses vendored IFNet v4.13 architecture. Weights are downloaded from HuggingFace Hub on first use.

Example: interpolate-frames.json — Generate video with Mochi, then 2x interpolate from 30fps to 60fps.

Metadata Embedding

Embed generation parameters in saved images. Enable by setting embed_metadata: true in a step's result configuration:

{
    "result": {
        "content_type": "image/png",
        "embed_metadata": true
    }
}
Format Storage Notes
PNG Text chunk (parameters key) Always available
JPEG/WebP EXIF UserComment Requires pip install piexif

Metadata includes step name, model name, and generation arguments (prompt, steps, guidance scale, etc.) as JSON.

Example: embed-metadata.json — Generate with Flux and embed parameters in PNG.

QR Code Generation

{
    "task": {
        "command": "qr_code",
        "arguments": {
            "qr_code_contents": "https://example.com"
        }
    }
}
Argument Required Description
qr_code_contents Yes Data to encode (URL, text, etc.)
height No Used with width to derive output resolution (default: 768)
width No Used with height to derive output resolution (default: 768)

The QR code is generated then resampled to max(height, width), aligned to the nearest 64px multiple.

Example: qr-code.json — QR code with artistic ControlNet

Chat/Dict Plumbing

These small tasks glue together multi-step pipelines that mix raw transformers components with task steps — for the cases text_generation does not cover.

format_chat_message

Build a text_inputs chat message list from a system and user message, in the shape a transformers.pipeline text-generation call expects:

{
    "task": {
        "command": "format_chat_message",
        "arguments": {
            "system_prompt": "You are a helpful assistant.",
            "user_message": "variable:prompt"
        }
    }
}
Argument Required Description
system_prompt Yes System instruction
user_message Yes User message content

Returns {"text_inputs": [{"role": "system", ...}, {"role": "user", ...}]}. Pass the result to a transformers.pipeline step's text_inputs argument via previous_result:.

get_dict_value

Extract a single value from a dictionary result (e.g., a transformers pipeline's output) for use in a later step:

{
    "task": {
        "command": "get_dict_value",
        "arguments": {
            "dict": "previous_result:augment_prompt",
            "key": "generated_text"
        }
    }
}
Argument Required Description
dict Yes Dictionary (or previous_result: reference) to read from
key Yes Key to extract

Returns the value at key, or None if the key is absent.

batch_decode_post_process

Decode generated token IDs and run model-specific post-processing (e.g., Florence-2's task-token parsing), using the processor from an earlier pipeline step:

{
    "task": {
        "command": "batch_decode_post_process",
        "pipeline_reference": "describe_image_processor",
        "arguments": {
            "generated_ids": "previous_result:describe_image_model.generated_ids",
            "task": "<DETAILED_CAPTION>"
        }
    }
}
Argument Required Description
pipeline_reference Yes Name of an earlier pipeline step whose processor to reuse (sibling of command/arguments, not inside arguments)
generated_ids Yes Token IDs to decode (e.g., a model step's generated_ids output)
task Yes Task token to post-process for (e.g., <DETAILED_CAPTION>)

Calls processor.batch_decode(...) then processor.post_process_generation(..., task=task) and returns parsed_answer[task].

Multi-Step Example

Canny edge detection followed by ControlNet generation:

{
    "steps": [
        {
            "name": "edges",
            "task": {
                "command": "canny",
                "arguments": {
                    "image": {
                        "location": "photo.jpg",
                        "low_threshold": 50,
                        "high_threshold": 200
                    }
                }
            },
            "result": { "content_type": "image/jpeg" }
        },
        {
            "name": "generate",
            "pipeline": {
                "configuration": {
                    "component_type": "FluxControlPipeline",
                    "offload": "sequential"
                },
                "from_pretrained_arguments": {
                    "model_name": "black-forest-labs/FLUX.1-Canny-dev",
                    "torch_dtype": "torch.bfloat16"
                },
                "arguments": {
                    "control_image": "previous_result:edges",
                    "prompt": "a watercolor painting",
                    "num_inference_steps": 50
                }
            },
            "result": { "content_type": "image/jpeg" }
        }
    ]
}

Examples