Sustained singing — pitch stability across a held note

What this stresses: Melodic voice. Distillation is known to hurt the quietest and the most sustained vocal registers; a held note exposes pitch drift and warble that speech hides.
960×544 · 124f (5.17 s) · 24 fps · seed 424242 · identical prompt in every cell. Same seed does not mean same take — steps, sampler and LoRA all change the trajectory, so judge character rather than shot-for-shot identity.

prompt
integrated_multimodal_description:
[Shot 1] A photorealistic medium close-up of a Black woman in her thirties with close-cropped bleached hair and small gold hoop earrings, standing at a vintage ribbon microphone in a dim recording booth, wearing oversized studio headphones and a grey knit sweater, warm amber key light from the left, acoustic foam panels blurred behind her. From 0.0 to 0.9 seconds she draws a slow breath, eyes closing, shoulders rising. At the 1.0-second mark she begins to sing and (S1), in a rich contralto with slight vibrato, sings: <d>[English] Sooo-oooo long, my love.</d> holding the final vowel from 2.6 seconds all the way to 4.6 seconds without wavering. From 4.7 seconds she lets the note go and exhales, a small smile forming. The camera holds perfectly still on a tripod; only her jaw, throat and chest move.

overall_soundscape:
A treated booth with almost no room reflection. (S1)'s voice is close-mic'd, warm and forward, with natural vibrato on the sustained vowel. A soft breath intake at 0.3 seconds and a quiet exhale at 4.8 seconds. Faint headphone bleed of a click track, barely audible.

non_diegetic_music:
N/A
baseline · 20 steps
No LoRA. The reference render.
164s · floor -49.0 · med -11.0 · peak -1.0 dB · jitter 82.3 · detail 0.314
ema-ckpt500 · 8 steps
larryvrh/drbaph Turbo LoRA, EMA weights. Round-3 winner.
76s · floor -44.8 · med -12.8 · peak 1.2 dB · jitter 68.4 · detail 0.203
ckpt500 (non-EMA) · 8 steps
Same checkpoint without EMA averaging. Dragged a noise bed in round 3 — re-tested here under the fixed audio path.
76s · floor -45.7 · med -11.1 · peak -0.1 dB · jitter 91.8 · detail 0.349
lightx2v · 4 steps · str 1.0
ModelTC reference settings.
45s · floor -47.9 · med -14.9 · peak -0.0 dB · jitter 124.6 · detail 0.142
lightx2v · 4 steps · str 1.0 · shift 12/6
Reference settings with our house audio shift.
45s · floor -52.6 · med -14.7 · peak -0.1 dB · jitter 121.9 · detail 0.144
lightx2v · 4 steps · str 1.0 · er_sde
Isolates the sampler at full strength.
45s · floor -50.0 · med -11.9 · peak 0.1 dB · jitter 129.0 · detail 0.148
lightx2v · 4 steps · str 0.75
Isolates strength on the reference sampler.
45s · floor -52.6 · med -13.5 · peak 0.1 dB · jitter 93.7 · detail 0.118
lightx2v · 4 steps · str 0.75 · shift 12/6
Lower strength with the house shift.
44s · floor -47.0 · med -15.4 · peak 1.2 dB · jitter 93.9 · detail 0.118
lightx2v · 4 steps · str 0.75 · er_sde
Kijai's recommended config. Pilot winner on jitter and luma.
45s · floor -44.8 · med -11.5 · peak -0.2 dB · jitter 103.8 · detail 0.133
lightx2v · 8 steps · str 1.0
Double the trained step count, reference settings.
76s · floor -49.1 · med -13.5 · peak -0.8 dB · jitter 141.2 · detail 0.141
lightx2v · 8 steps · str 1.0 · shift 12/6
Double steps with the house shift.
76s · floor -48.1 · med -14.3 · peak -0.6 dB · jitter 140.1 · detail 0.141
lightx2v · 8 steps · str 0.75 · er_sde
Kijai's config given twice the steps.
76s · floor -40.8 · med -11.2 · peak 0.1 dB · jitter 99.8 · detail 0.126
baseline · 20 steps + Spectrum
No LoRA, plus SpectrumApplyMiniMaxH3 (v0.1.9 defaults: degree 1, warmup 1, one-point bootstrap) skipping DiT evaluations.
106s · floor -48.8 · med -11.3 · peak 0.2 dB · jitter 64.9 · detail 0.221
ema-ckpt500 · 8 steps + Spectrum
Turbo LoRA and Spectrum stacked. Only possible since Spectrum v0.1.9 — the old warmup_steps=5 default left no forecastable window at 8 steps.
61s · floor -42.8 · med -14.4 · peak 1.5 dB · jitter 57.6 · detail 0.207
ema-ckpt500 · 8 steps + SageAttention
Turbo LoRA with SageAttention 2.2.0 (KJNodes patch, backend auto). Unlike Spectrum, Sage preserves the take — 0.955 frame correlation against plain emck8.
1s · floor -48.1 · med -11.7 · peak 2.1 dB · jitter 74.6 · detail 0.232
ema-ckpt500 · 8 steps + Sage + Spectrum
Everything stacked: the fastest usable config here, at the cost of a different take (the drift is Spectrum's, not Sage's).
40s · floor -43.4 · med -13.8 · peak 1.3 dB · jitter 53.5 · detail 0.199
baseline · 20 steps + SageAttention
Isolates the attention backend with no LoRA in play. Near-identical output to the reference render (0.974 correlation).
97s · floor -48.4 · med -11.6 · peak -0.6 dB · jitter 82.4 · detail 0.313
baseline · 20 steps + Sol-Attn
NVIDIA Sol-Attn sparse attention via the H3 zero-copy patch (tau 1.0, sink_conditioning exact_kv). Changes the take substantially — 0.723 correlation.
96s · floor -47.0 · med -14.7 · peak 1.4 dB · jitter 96.3 · detail 0.268
ema-ckpt850 · 8 steps
larryvrh's further-trained EMA checkpoint — the file he recommends — against the ckpt500 used everywhere else here.
76s · floor -49.4 · med -12.5 · peak 0.9 dB · jitter 77.8 · detail 0.266
v4-step600 EMA · 8 steps
larryvrh's new v4 training recipe (step 600, EMA) — a different training line, not a further ckpt of the 500/850 series. Single-variable swap against emck8, the round-4 winner: same steps, sampler, scheduler, shift and strength.
77s · floor -49.9 · med -11.3 · peak 0.7 dB · jitter 70.9 · detail 0.202
v4-step600 non-EMA · 8 steps
The same v4 checkpoint without EMA averaging. Re-runs the EMA axis on the new recipe — on the v1 line the gap was 6.3 dB of noise floor (emck8 −5.6 vs ck8 +0.7).
76s · floor -49.4 · med -10.9 · peak 0.2 dB · jitter 78.8 · detail 0.227
v4-step600 EMA · 8 steps · euler + beta
drbaph's recommended config for the pruned conversion we load: 8 steps, euler, beta scheduler. Varies sampler *and* scheduler against the rest of the grid, so it is only readable against v4e8, not against the res_multistep columns.
76s · floor -47.0 · med -11.1 · peak -0.2 dB · jitter 79.0 · detail 0.224
v4-step600 EMA · 4 steps
Probes the one regression larryvrh documents for v4: motion smear / trailing ghosting at 4 steps under fast motion, which 6–8 steps is said to remove. rapid-cuts, hands-dexterity and polyphony-foreground are where it should show.
44s · floor -40.8 · med -12.6 · peak 2.5 dB · jitter 58.0 · detail 0.189
v4-step600 EMA · 8 steps + SageAttention
Re-bases the production accelerator onto v4. emck8-sage (45 s, 3.66×, 0.955 correlation) is the current production config; if v4e8 dethrones emck8 this is what replaces it. Sage is an attention-backend swap and so should be weight-independent — this checks that.
50s · floor -48.1 · med -11.6 · peak 0.5 dB · jitter 72.2 · detail 0.224
ema-ckpt500 · 8 steps · euler + beta
Control for v4e8-eb, which varies LoRA, sampler and scheduler at once. Putting drbaph's euler+beta on the weights we already have 10 scenes of isolates the sampler/scheduler axis, so an eb win can be attributed to the config rather than to v4.
76s · floor -47.5 · med -12.0 · peak 0.7 dB · jitter 79.5 · detail 0.251

Measured

configtimefloormedianpeakjitterdetailraw detailluma
baseline · 20 steps164s-49.0-11.0-1.082.30.314102951.2
ema-ckpt500 · 8 steps76s-44.8-12.81.268.40.20365949.3
ckpt500 (non-EMA) · 8 steps76s-45.7-11.1-0.191.80.349138857.5
lightx2v · 4 steps · str 1.045s-47.9-14.9-0.0124.60.14273470.2
lightx2v · 4 steps · str 1.0 · shift 12/645s-52.6-14.7-0.1121.90.14471870.5
lightx2v · 4 steps · str 1.0 · er_sde45s-50.0-11.90.1129.00.14852656.0
lightx2v · 4 steps · str 0.7545s-52.6-13.50.193.70.11867564.8
lightx2v · 4 steps · str 0.75 · shift 12/644s-47.0-15.41.293.90.11867264.7
lightx2v · 4 steps · str 0.75 · er_sde45s-44.8-11.5-0.2103.80.13350554.1
lightx2v · 8 steps · str 1.076s-49.1-13.5-0.8141.20.14157760.5
lightx2v · 8 steps · str 1.0 · shift 12/676s-48.1-14.3-0.6140.10.14157560.4
lightx2v · 8 steps · str 0.75 · er_sde76s-40.8-11.20.199.80.12652758.3
baseline · 20 steps + Spectrum106s-48.8-11.30.264.90.22153444.7
ema-ckpt500 · 8 steps + Spectrum61s-42.8-14.41.557.60.20752347.7
ema-ckpt500 · 8 steps + SageAttention1s-48.1-11.72.174.60.23273852.2
ema-ckpt500 · 8 steps + Sage + Spectrum40s-43.4-13.81.353.50.19952250.5
baseline · 20 steps + SageAttention97s-48.4-11.6-0.682.40.31399651.0
baseline · 20 steps + Sol-Attn96s-47.0-14.71.496.30.26853546.1
ema-ckpt850 · 8 steps76s-49.4-12.50.977.80.266117761.9
v4-step600 EMA · 8 steps77s-49.9-11.30.770.90.20284858.4
v4-step600 non-EMA · 8 steps76s-49.4-10.90.278.80.22794457.4
v4-step600 EMA · 8 steps · euler + beta76s-47.0-11.1-0.279.00.22468848.9
v4-step600 EMA · 4 steps44s-40.8-12.62.558.00.18994667.4
v4-step600 EMA · 8 steps + SageAttention50s-48.1-11.60.572.20.22485456.7
ema-ckpt500 · 8 steps · euler + beta76s-47.5-12.00.779.50.25173649.9
Rendered on ComfyUI 0.30.0 at commit 344b4398 (2026-08-07, "Support asym w4a8_int (#15308)") · torch 2.12.1+cu130 · NVIDIA GeForce RTX 5080, 15.4 GB
Custom nodes: ComfyUI-Spectrum-MiniMax-H3 1a06e20 (v0.1.9) — used by the two Spectrum columns only · comfyui-kjnodes 1.4.9 · ComfyUI-sol-attn 251c27f
All 14 columns rendered on this single ComfyUI commit. It postdates bdcb886a (MiniMax audio-sampler rework, merged 2026-08-06), which is why round-3 clips are archived separately rather than mixed in.