Four hard cuts in five seconds — scene-change handling

What this stresses: Editing inside a single generation. Tests whether hard cuts survive distillation, or smear into dissolves, and whether the audio bed cuts with the picture.
960×544 · 124f (5.17 s) · 24 fps · seed 424242 · identical prompt in every cell. Same seed does not mean same take — steps, sampler and LoRA all change the trajectory, so judge character rather than shot-for-shot identity.

prompt
integrated_multimodal_description:
[Shot 1] 0.0 to 1.2 seconds: a photorealistic extreme close-up of a match head striking a rough strip and bursting into flame, filling the frame, deep black background. HARD CUT.
[Shot 2] 1.2 to 2.4 seconds: a wide shot of a red hot-air balloon rising over green morning fields, cold blue sky, burner flaring. HARD CUT.
[Shot 3] 2.4 to 3.6 seconds: a low-angle close-up of muddy boots running fast across wet cobblestones, splashing a shallow puddle toward the lens. HARD CUT.
[Shot 4] 3.6 to 5.0 seconds: a static close-up of an old brass alarm clock on a bedside table, hammer striking the bells, vibrating hard. Each cut is instantaneous with no dissolve, no fade and no camera move across the boundary; the four shots share no colour palette, no location and no subject.

overall_soundscape:
The audio cuts hard with the picture, four times, with no crossfade. 0.0 to 1.2 seconds: a match rasp and a soft flame whoosh in a dead-quiet room. 1.2 to 2.4 seconds: open wind, birdsong, and a loud roaring burner blast. 2.4 to 3.6 seconds: fast heavy footfalls on wet stone with splashes and close breathing. 3.6 to 5.0 seconds: a piercing continuous mechanical alarm bell, close and bright.

non_diegetic_music:
N/A
baseline · 20 steps
No LoRA. The reference render.
166s · floor -32.7 · med -23.2 · peak -1.3 dB · jitter 282.7 · detail 0.144
ema-ckpt500 · 8 steps
larryvrh/drbaph Turbo LoRA, EMA weights. Round-3 winner.
77s · floor -37.9 · med -24.1 · peak 0.8 dB · jitter 298.4 · detail 0.123
ckpt500 (non-EMA) · 8 steps
Same checkpoint without EMA averaging. Dragged a noise bed in round 3 — re-tested here under the fixed audio path.
77s · floor -34.7 · med -28.5 · peak -2.1 dB · jitter 332.5 · detail 0.152
lightx2v · 4 steps · str 1.0
ModelTC reference settings.
45s · floor -39.2 · med -25.4 · peak -1.7 dB · jitter 355.2 · detail 0.137
lightx2v · 4 steps · str 1.0 · shift 12/6
Reference settings with our house audio shift.
45s · floor -47.6 · med -31.8 · peak -4.6 dB · jitter 347.2 · detail 0.127
lightx2v · 4 steps · str 1.0 · er_sde
Isolates the sampler at full strength.
45s · floor -43.1 · med -31.9 · peak -1.0 dB · jitter 301.9 · detail 0.130
lightx2v · 4 steps · str 0.75
Isolates strength on the reference sampler.
46s · floor -39.3 · med -28.0 · peak -2.0 dB · jitter 331.9 · detail 0.154
lightx2v · 4 steps · str 0.75 · shift 12/6
Lower strength with the house shift.
45s · floor -50.9 · med -32.3 · peak -3.1 dB · jitter 332.7 · detail 0.147
lightx2v · 4 steps · str 0.75 · er_sde
Kijai's recommended config. Pilot winner on jitter and luma.
45s · floor -48.7 · med -33.4 · peak -7.2 dB · jitter 292.6 · detail 0.131
lightx2v · 8 steps · str 1.0
Double the trained step count, reference settings.
77s · floor -30.9 · med -20.1 · peak 0.6 dB · jitter 353.6 · detail 0.144
lightx2v · 8 steps · str 1.0 · shift 12/6
Double steps with the house shift.
76s · floor -40.3 · med -25.9 · peak 0.1 dB · jitter 357.2 · detail 0.140
lightx2v · 8 steps · str 0.75 · er_sde
Kijai's config given twice the steps.
77s · floor -42.7 · med -31.1 · peak -6.6 dB · jitter 299.6 · detail 0.140
baseline · 20 steps + Spectrum
No LoRA, plus SpectrumApplyMiniMaxH3 (v0.1.9 defaults: degree 1, warmup 1, one-point bootstrap) skipping DiT evaluations.
108s · floor -44.7 · med -27.6 · peak -3.0 dB · jitter 313.6 · detail 0.155
ema-ckpt500 · 8 steps + Spectrum
Turbo LoRA and Spectrum stacked. Only possible since Spectrum v0.1.9 — the old warmup_steps=5 default left no forecastable window at 8 steps.
61s · floor -45.6 · med -33.1 · peak -7.6 dB · jitter 309.8 · detail 0.135
ema-ckpt500 · 8 steps + SageAttention
Turbo LoRA with SageAttention 2.2.0 (KJNodes patch, backend auto). Unlike Spectrum, Sage preserves the take — 0.955 frame correlation against plain emck8.
51s · floor -35.3 · med -26.4 · peak -4.1 dB · jitter 294.2 · detail 0.118
ema-ckpt500 · 8 steps + Sage + Spectrum
Everything stacked: the fastest usable config here, at the cost of a different take (the drift is Spectrum's, not Sage's).
41s · floor -46.2 · med -32.4 · peak -8.2 dB · jitter 300.8 · detail 0.135
baseline · 20 steps + SageAttention
Isolates the attention backend with no LoRA in play. Near-identical output to the reference render (0.974 correlation).
98s · floor -35.8 · med -23.8 · peak -1.4 dB · jitter 269.8 · detail 0.154
baseline · 20 steps + Sol-Attn
NVIDIA Sol-Attn sparse attention via the H3 zero-copy patch (tau 1.0, sink_conditioning exact_kv). Changes the take substantially — 0.723 correlation.
100s · floor -42.6 · med -28.6 · peak -5.5 dB · jitter 286.3 · detail 0.112
ema-ckpt850 · 8 steps
larryvrh's further-trained EMA checkpoint — the file he recommends — against the ckpt500 used everywhere else here.
77s · floor -36.7 · med -29.1 · peak -5.4 dB · jitter 319.4 · detail 0.138
v4-step600 EMA · 8 steps
larryvrh's new v4 training recipe (step 600, EMA) — a different training line, not a further ckpt of the 500/850 series. Single-variable swap against emck8, the round-4 winner: same steps, sampler, scheduler, shift and strength.
143s · floor -35.4 · med -23.7 · peak -0.9 dB · jitter 308.4 · detail 0.126
v4-step600 non-EMA · 8 steps
The same v4 checkpoint without EMA averaging. Re-runs the EMA axis on the new recipe — on the v1 line the gap was 6.3 dB of noise floor (emck8 −5.6 vs ck8 +0.7).
77s · floor -31.1 · med -20.5 · peak 0.1 dB · jitter 303.1 · detail 0.125
v4-step600 EMA · 8 steps · euler + beta
drbaph's recommended config for the pruned conversion we load: 8 steps, euler, beta scheduler. Varies sampler *and* scheduler against the rest of the grid, so it is only readable against v4e8, not against the res_multistep columns.
77s · floor -34.9 · med -24.3 · peak -0.6 dB · jitter 302.4 · detail 0.185
v4-step600 EMA · 4 steps
Probes the one regression larryvrh documents for v4: motion smear / trailing ghosting at 4 steps under fast motion, which 6–8 steps is said to remove. rapid-cuts, hands-dexterity and polyphony-foreground are where it should show.
45s · floor -37.2 · med -21.4 · peak 1.0 dB · jitter 410.2 · detail 0.178
v4-step600 EMA · 8 steps + SageAttention
Re-bases the production accelerator onto v4. emck8-sage (45 s, 3.66×, 0.955 correlation) is the current production config; if v4e8 dethrones emck8 this is what replaces it. Sage is an attention-backend swap and so should be weight-independent — this checks that.
51s · floor -32.3 · med -23.3 · peak -0.2 dB · jitter 304.7 · detail 0.127
ema-ckpt500 · 8 steps · euler + beta
Control for v4e8-eb, which varies LoRA, sampler and scheduler at once. Putting drbaph's euler+beta on the weights we already have 10 scenes of isolates the sampler/scheduler axis, so an eb win can be attributed to the config rather than to v4.
77s · floor -34.1 · med -29.7 · peak -6.6 dB · jitter 251.1 · detail 0.137

Measured

configtimefloormedianpeakjitterdetailraw detailluma
baseline · 20 steps166s-32.7-23.2-1.3282.70.14437080.9
ema-ckpt500 · 8 steps77s-37.9-24.10.8298.40.12336082.7
ckpt500 (non-EMA) · 8 steps77s-34.7-28.5-2.1332.50.15245882.9
lightx2v · 4 steps · str 1.045s-39.2-25.4-1.7355.20.13766288.1
lightx2v · 4 steps · str 1.0 · shift 12/645s-47.6-31.8-4.6347.20.12762887.8
lightx2v · 4 steps · str 1.0 · er_sde45s-43.1-31.9-1.0301.90.13040780.4
lightx2v · 4 steps · str 0.7546s-39.3-28.0-2.0331.90.15462289.9
lightx2v · 4 steps · str 0.75 · shift 12/645s-50.9-32.3-3.1332.70.14759089.1
lightx2v · 4 steps · str 0.75 · er_sde45s-48.7-33.4-7.2292.60.13135582.4
lightx2v · 8 steps · str 1.077s-30.9-20.10.6353.60.14454084.2
lightx2v · 8 steps · str 1.0 · shift 12/676s-40.3-25.90.1357.20.14053583.6
lightx2v · 8 steps · str 0.75 · er_sde77s-42.7-31.1-6.6299.60.14041887.7
baseline · 20 steps + Spectrum108s-44.7-27.6-3.0313.60.15531175.1
ema-ckpt500 · 8 steps + Spectrum61s-45.6-33.1-7.6309.80.13535175.8
ema-ckpt500 · 8 steps + SageAttention51s-35.3-26.4-4.1294.20.11835382.5
ema-ckpt500 · 8 steps + Sage + Spectrum41s-46.2-32.4-8.2300.80.13534378.0
baseline · 20 steps + SageAttention98s-35.8-23.8-1.4269.80.15439382.1
baseline · 20 steps + Sol-Attn100s-42.6-28.6-5.5286.30.11222765.7
ema-ckpt850 · 8 steps77s-36.7-29.1-5.4319.40.13846783.6
v4-step600 EMA · 8 steps143s-35.4-23.7-0.9308.40.12637285.6
v4-step600 non-EMA · 8 steps77s-31.1-20.50.1303.10.12539187.9
v4-step600 EMA · 8 steps · euler + beta77s-34.9-24.3-0.6302.40.18539780.7
v4-step600 EMA · 4 steps45s-37.2-21.41.0410.20.17864891.6
v4-step600 EMA · 8 steps + SageAttention51s-32.3-23.3-0.2304.70.12736385.8
ema-ckpt500 · 8 steps · euler + beta77s-34.1-29.7-6.6251.10.13732576.5
Rendered on ComfyUI 0.30.0 at commit 344b4398 (2026-08-07, "Support asym w4a8_int (#15308)") · torch 2.12.1+cu130 · NVIDIA GeForce RTX 5080, 15.4 GB
Custom nodes: ComfyUI-Spectrum-MiniMax-H3 1a06e20 (v0.1.9) — used by the two Spectrum columns only · comfyui-kjnodes 1.4.9 · ComfyUI-sol-attn 251c27f
All 14 columns rendered on this single ComfyUI commit. It postdates bdcb886a (MiniMax audio-sampler rework, merged 2026-08-06), which is why round-3 clips are archived separately rather than mixed in.