Overlapping dialogue — two voices talking across each other

What this stresses: Speaker separation under simultaneity. Yesterday's two-speakers scene alternated cleanly; this one deliberately collides the turns, which is where voice identity smears.
960×544 · 124f (5.17 s) · 24 fps · seed 424242 · identical prompt in every cell. Same seed does not mean same take — steps, sampler and LoRA all change the trajectory, so judge character rather than shot-for-shot identity.

prompt
integrated_multimodal_description:
[Shot 1] A photorealistic two-shot of a cramped newsroom desk: (S1) is a wiry white man in his fifties with wire-frame glasses, thinning grey hair and rolled shirtsleeves, seated left; (S2) is a South Asian woman in her late twenties with a dark ponytail and a denim jacket, seated right, both lit by cold overhead fluorescents and the blue spill of two monitors. From 0.0 seconds (S1) jabs a finger at his screen and, in a clipped irritated tenor, says: <d>[English] We do not have the second source, we cannot run it tonight.</d> Starting at 1.9 seconds, while he is still speaking, (S2) cuts across him in a faster, higher, urgent voice: <d>[English] We have the documents, that is the source.</d> Both voices overlap from 1.9 to 2.8 seconds. From 3.2 seconds they stop together, hold a beat of eye contact, and (S1) leans back in his chair. Handheld camera with a slow drift right.

overall_soundscape:
Open-plan newsroom: distant phones, keyboard clatter, an HVAC hum. (S1) is close and dry with a slight nasal edge; (S2) is brighter and more forward. Their overlap at 1.9 to 2.8 seconds keeps both voices individually intelligible. A chair creak at 3.6 seconds as he leans back.

non_diegetic_music:
N/A
baseline · 20 steps
No LoRA. The reference render.
166s · floor -39.8 · med -24.5 · peak -4.5 dB · jitter 169.5 · detail 0.158
ema-ckpt500 · 8 steps
larryvrh/drbaph Turbo LoRA, EMA weights. Round-3 winner.
77s · floor -46.7 · med -25.3 · peak -3.9 dB · jitter 163.3 · detail 0.161
ckpt500 (non-EMA) · 8 steps
Same checkpoint without EMA averaging. Dragged a noise bed in round 3 — re-tested here under the fixed audio path.
77s · floor -38.0 · med -23.5 · peak -4.1 dB · jitter 252.0 · detail 0.286
lightx2v · 4 steps · str 1.0
ModelTC reference settings.
45s · floor -29.9 · med -14.0 · peak 2.1 dB · jitter 281.7 · detail 0.161
lightx2v · 4 steps · str 1.0 · shift 12/6
Reference settings with our house audio shift.
45s · floor -36.7 · med -20.1 · peak 1.2 dB · jitter 275.7 · detail 0.154
lightx2v · 4 steps · str 1.0 · er_sde
Isolates the sampler at full strength.
45s · floor -36.4 · med -20.8 · peak -0.8 dB · jitter 250.7 · detail 0.099
lightx2v · 4 steps · str 0.75
Isolates strength on the reference sampler.
45s · floor -30.4 · med -15.2 · peak 2.4 dB · jitter 239.9 · detail 0.144
lightx2v · 4 steps · str 0.75 · shift 12/6
Lower strength with the house shift.
45s · floor -35.8 · med -22.6 · peak 1.6 dB · jitter 235.9 · detail 0.137
lightx2v · 4 steps · str 0.75 · er_sde
Kijai's recommended config. Pilot winner on jitter and luma.
45s · floor -38.7 · med -23.6 · peak -4.2 dB · jitter 222.6 · detail 0.104
lightx2v · 8 steps · str 1.0
Double the trained step count, reference settings.
77s · floor -38.8 · med -23.9 · peak -3.9 dB · jitter 320.1 · detail 0.122
lightx2v · 8 steps · str 1.0 · shift 12/6
Double steps with the house shift.
77s · floor -39.2 · med -23.9 · peak -5.8 dB · jitter 318.7 · detail 0.123
lightx2v · 8 steps · str 0.75 · er_sde
Kijai's config given twice the steps.
77s · floor -36.7 · med -23.5 · peak -6.6 dB · jitter 259.7 · detail 0.113
baseline · 20 steps + Spectrum
No LoRA, plus SpectrumApplyMiniMaxH3 (v0.1.9 defaults: degree 1, warmup 1, one-point bootstrap) skipping DiT evaluations.
108s · floor -42.1 · med -28.4 · peak -7.2 dB · jitter 162.5 · detail 0.168
ema-ckpt500 · 8 steps + Spectrum
Turbo LoRA and Spectrum stacked. Only possible since Spectrum v0.1.9 — the old warmup_steps=5 default left no forecastable window at 8 steps.
61s · floor -40.8 · med -28.3 · peak -10.2 dB · jitter 100.4 · detail 0.236
ema-ckpt500 · 8 steps + SageAttention
Turbo LoRA with SageAttention 2.2.0 (KJNodes patch, backend auto). Unlike Spectrum, Sage preserves the take — 0.955 frame correlation against plain emck8.
50s · floor -51.3 · med -24.4 · peak -4.7 dB · jitter 152.4 · detail 0.165
ema-ckpt500 · 8 steps + Sage + Spectrum
Everything stacked: the fastest usable config here, at the cost of a different take (the drift is Spectrum's, not Sage's).
41s · floor -42.1 · med -26.2 · peak -9.6 dB · jitter 113.0 · detail 0.216
baseline · 20 steps + SageAttention
Isolates the attention backend with no LoRA in play. Near-identical output to the reference render (0.974 correlation).
98s · floor -40.0 · med -25.0 · peak -3.8 dB · jitter 174.5 · detail 0.153
baseline · 20 steps + Sol-Attn
NVIDIA Sol-Attn sparse attention via the H3 zero-copy patch (tau 1.0, sink_conditioning exact_kv). Changes the take substantially — 0.723 correlation.
103s · floor -35.5 · med -21.4 · peak -3.1 dB · jitter 99.9 · detail 0.175
ema-ckpt850 · 8 steps
larryvrh's further-trained EMA checkpoint — the file he recommends — against the ckpt500 used everywhere else here.
77s · floor -39.3 · med -23.8 · peak -4.7 dB · jitter 252.5 · detail 0.232
v4-step600 EMA · 8 steps
larryvrh's new v4 training recipe (step 600, EMA) — a different training line, not a further ckpt of the 500/850 series. Single-variable swap against emck8, the round-4 winner: same steps, sampler, scheduler, shift and strength.
78s · floor -39.8 · med -23.4 · peak -6.2 dB · jitter 189.6 · detail 0.182
v4-step600 non-EMA · 8 steps
The same v4 checkpoint without EMA averaging. Re-runs the EMA axis on the new recipe — on the v1 line the gap was 6.3 dB of noise floor (emck8 −5.6 vs ck8 +0.7).
77s · floor -39.0 · med -23.0 · peak -3.8 dB · jitter 189.7 · detail 0.220
v4-step600 EMA · 8 steps · euler + beta
drbaph's recommended config for the pruned conversion we load: 8 steps, euler, beta scheduler. Varies sampler *and* scheduler against the rest of the grid, so it is only readable against v4e8, not against the res_multistep columns.
77s · floor -39.3 · med -22.8 · peak -1.8 dB · jitter 228.3 · detail 0.216
v4-step600 EMA · 4 steps
Probes the one regression larryvrh documents for v4: motion smear / trailing ghosting at 4 steps under fast motion, which 6–8 steps is said to remove. rapid-cuts, hands-dexterity and polyphony-foreground are where it should show.
45s · floor -41.9 · med -24.1 · peak 2.1 dB · jitter 187.8 · detail 0.195
v4-step600 EMA · 8 steps + SageAttention
Re-bases the production accelerator onto v4. emck8-sage (45 s, 3.66×, 0.955 correlation) is the current production config; if v4e8 dethrones emck8 this is what replaces it. Sage is an attention-backend swap and so should be weight-independent — this checks that.
50s · floor -39.7 · med -22.4 · peak -5.6 dB · jitter 182.1 · detail 0.180
ema-ckpt500 · 8 steps · euler + beta
Control for v4e8-eb, which varies LoRA, sampler and scheduler at once. Putting drbaph's euler+beta on the weights we already have 10 scenes of isolates the sampler/scheduler axis, so an eb win can be attributed to the config rather than to v4.
77s · floor -38.5 · med -23.3 · peak -3.6 dB · jitter 193.9 · detail 0.141

Measured

configtimefloormedianpeakjitterdetailraw detailluma
baseline · 20 steps166s-39.8-24.5-4.5169.50.15846776.2
ema-ckpt500 · 8 steps77s-46.7-25.3-3.9163.30.16144372.8
ckpt500 (non-EMA) · 8 steps77s-38.0-23.5-4.1252.00.28673372.5
lightx2v · 4 steps · str 1.045s-29.9-14.02.1281.70.16169686.7
lightx2v · 4 steps · str 1.0 · shift 12/645s-36.7-20.11.2275.70.15466386.2
lightx2v · 4 steps · str 1.0 · er_sde45s-36.4-20.8-0.8250.70.09936579.2
lightx2v · 4 steps · str 0.7545s-30.4-15.22.4239.90.14457982.0
lightx2v · 4 steps · str 0.75 · shift 12/645s-35.8-22.61.6235.90.13755282.7
lightx2v · 4 steps · str 0.75 · er_sde45s-38.7-23.6-4.2222.60.10434376.8
lightx2v · 8 steps · str 1.077s-38.8-23.9-3.9320.10.12242281.0
lightx2v · 8 steps · str 1.0 · shift 12/677s-39.2-23.9-5.8318.70.12343381.1
lightx2v · 8 steps · str 0.75 · er_sde77s-36.7-23.5-6.6259.70.11342083.0
baseline · 20 steps + Spectrum108s-42.1-28.4-7.2162.50.16835659.5
ema-ckpt500 · 8 steps + Spectrum61s-40.8-28.3-10.2100.40.23657461.6
ema-ckpt500 · 8 steps + SageAttention50s-51.3-24.4-4.7152.40.16545372.6
ema-ckpt500 · 8 steps + Sage + Spectrum41s-42.1-26.2-9.6113.00.21652462.3
baseline · 20 steps + SageAttention98s-40.0-25.0-3.8174.50.15345676.3
baseline · 20 steps + Sol-Attn103s-35.5-21.4-3.199.90.17554564.8
ema-ckpt850 · 8 steps77s-39.3-23.8-4.7252.50.23265472.2
v4-step600 EMA · 8 steps78s-39.8-23.4-6.2189.60.18257375.3
v4-step600 non-EMA · 8 steps77s-39.0-23.0-3.8189.70.22070676.8
v4-step600 EMA · 8 steps · euler + beta77s-39.3-22.8-1.8228.30.21646668.1
v4-step600 EMA · 4 steps45s-41.9-24.12.1187.80.19587090.6
v4-step600 EMA · 8 steps + SageAttention50s-39.7-22.4-5.6182.10.18056874.8
ema-ckpt500 · 8 steps · euler + beta77s-38.5-23.3-3.6193.90.14131569.3
Rendered on ComfyUI 0.30.0 at commit 344b4398 (2026-08-07, "Support asym w4a8_int (#15308)") · torch 2.12.1+cu130 · NVIDIA GeForce RTX 5080, 15.4 GB
Custom nodes: ComfyUI-Spectrum-MiniMax-H3 1a06e20 (v0.1.9) — used by the two Spectrum columns only · comfyui-kjnodes 1.4.9 · ComfyUI-sol-attn 251c27f
All 14 columns rendered on this single ComfyUI commit. It postdates bdcb886a (MiniMax audio-sampler rework, merged 2026-08-06), which is why round-3 clips are archived separately rather than mixed in.