Hands close-up — ten fingers doing precise work

What this stresses: The classic generative failure. Full-frame hands at 0.5 MP with fine bimanual coordination, plus synchronised micro-foley.
960×544 · 124f (5.17 s) · 24 fps · seed 424242 · identical prompt in every cell. Same seed does not mean same take — steps, sampler and LoRA all change the trajectory, so judge character rather than shot-for-shot identity.

prompt
integrated_multimodal_description:
[Shot 1] A photorealistic overhead extreme close-up of a pair of weathered adult hands tying a fishing fly in a small steel vise, filling the frame against a dark green cutting mat, hard focused task lighting from the upper left, both hands fully visible with all ten fingers in frame. From 0.0 to 1.5 seconds the left thumb and forefinger pinch a wisp of red thread taut while the right hand makes three tight wraps around the hook shank. At the 1.8-second mark the right hand picks up fine-tipped scissors and makes a single clean snip at 2.3 seconds. From 2.6 to 4.0 seconds both hands work together whip-finishing the head of the fly, fingers crossing over one another. At 4.4 seconds the right index finger taps the finished fly once and it swings on the hook. No face, arms or body is ever visible; the hands never leave frame and never blur into one another.

overall_soundscape:
A quiet workroom, close-mic'd and intimate. Thread drawn taut with a fine squeak at 0.7 seconds. A single sharp metallic scissor snip at 2.3 seconds. Soft finger friction and thread whisper from 2.6 to 4.0 seconds. A tiny hook tap and a metallic ring at 4.4 seconds. A very low room tone underneath, no music, no voice.

non_diegetic_music:
N/A
baseline · 20 steps
No LoRA. The reference render.
166s · floor -55.1 · med -52.1 · peak -2.1 dB · jitter 224.7 · detail 0.036
ema-ckpt500 · 8 steps
larryvrh/drbaph Turbo LoRA, EMA weights. Round-3 winner.
77s · floor -58.3 · med -54.1 · peak 0.9 dB · jitter 289.5 · detail 0.027
ckpt500 (non-EMA) · 8 steps
Same checkpoint without EMA averaging. Dragged a noise bed in round 3 — re-tested here under the fixed audio path.
77s · floor -53.0 · med -45.4 · peak 3.0 dB · jitter 253.4 · detail 0.051
lightx2v · 4 steps · str 1.0
ModelTC reference settings.
45s · floor -63.1 · med -56.5 · peak 2.7 dB · jitter 302.4 · detail 0.049
lightx2v · 4 steps · str 1.0 · shift 12/6
Reference settings with our house audio shift.
45s · floor -73.9 · med -63.1 · peak 1.0 dB · jitter 298.2 · detail 0.048
lightx2v · 4 steps · str 1.0 · er_sde
Isolates the sampler at full strength.
45s · floor -63.6 · med -59.6 · peak 2.3 dB · jitter 287.1 · detail 0.042
lightx2v · 4 steps · str 0.75
Isolates strength on the reference sampler.
45s · floor -62.8 · med -56.2 · peak 3.1 dB · jitter 301.5 · detail 0.050
lightx2v · 4 steps · str 0.75 · shift 12/6
Lower strength with the house shift.
45s · floor -75.5 · med -60.9 · peak 3.0 dB · jitter 295.3 · detail 0.048
lightx2v · 4 steps · str 0.75 · er_sde
Kijai's recommended config. Pilot winner on jitter and luma.
45s · floor -64.8 · med -60.1 · peak 2.0 dB · jitter 250.5 · detail 0.034
lightx2v · 8 steps · str 1.0
Double the trained step count, reference settings.
77s · floor -59.3 · med -52.4 · peak 1.3 dB · jitter 319.1 · detail 0.045
lightx2v · 8 steps · str 1.0 · shift 12/6
Double steps with the house shift.
76s · floor -66.4 · med -55.8 · peak 2.2 dB · jitter 316.6 · detail 0.045
lightx2v · 8 steps · str 0.75 · er_sde
Kijai's config given twice the steps.
77s · floor -60.1 · med -57.0 · peak 2.2 dB · jitter 287.5 · detail 0.037
baseline · 20 steps + Spectrum
No LoRA, plus SpectrumApplyMiniMaxH3 (v0.1.9 defaults: degree 1, warmup 1, one-point bootstrap) skipping DiT evaluations.
107s · floor -54.1 · med -50.9 · peak -3.6 dB · jitter 288.1 · detail 0.027
ema-ckpt500 · 8 steps + Spectrum
Turbo LoRA and Spectrum stacked. Only possible since Spectrum v0.1.9 — the old warmup_steps=5 default left no forecastable window at 8 steps.
61s · floor -53.7 · med -51.9 · peak -4.1 dB · jitter 289.6 · detail 0.030
ema-ckpt500 · 8 steps + SageAttention
Turbo LoRA with SageAttention 2.2.0 (KJNodes patch, backend auto). Unlike Spectrum, Sage preserves the take — 0.955 frame correlation against plain emck8.
50s · floor -62.0 · med -56.5 · peak -2.1 dB · jitter 307.8 · detail 0.032
ema-ckpt500 · 8 steps + Sage + Spectrum
Everything stacked: the fastest usable config here, at the cost of a different take (the drift is Spectrum's, not Sage's).
41s · floor -56.6 · med -54.4 · peak -0.0 dB · jitter 283.2 · detail 0.026
baseline · 20 steps + SageAttention
Isolates the attention backend with no LoRA in play. Near-identical output to the reference render (0.974 correlation).
98s · floor -55.0 · med -52.1 · peak -4.7 dB · jitter 231.4 · detail 0.034
baseline · 20 steps + Sol-Attn
NVIDIA Sol-Attn sparse attention via the H3 zero-copy patch (tau 1.0, sink_conditioning exact_kv). Changes the take substantially — 0.723 correlation.
98s · floor -51.3 · med -48.7 · peak 0.3 dB · jitter 299.7 · detail 0.026
ema-ckpt850 · 8 steps
larryvrh's further-trained EMA checkpoint — the file he recommends — against the ckpt500 used everywhere else here.
77s · floor -56.0 · med -49.9 · peak -0.6 dB · jitter 297.8 · detail 0.039
v4-step600 EMA · 8 steps
larryvrh's new v4 training recipe (step 600, EMA) — a different training line, not a further ckpt of the 500/850 series. Single-variable swap against emck8, the round-4 winner: same steps, sampler, scheduler, shift and strength.
78s · floor -52.3 · med -49.0 · peak 1.9 dB · jitter 253.6 · detail 0.030
v4-step600 non-EMA · 8 steps
The same v4 checkpoint without EMA averaging. Re-runs the EMA axis on the new recipe — on the v1 line the gap was 6.3 dB of noise floor (emck8 −5.6 vs ck8 +0.7).
77s · floor -50.2 · med -46.9 · peak -1.2 dB · jitter 260.5 · detail 0.035
v4-step600 EMA · 8 steps · euler + beta
drbaph's recommended config for the pruned conversion we load: 8 steps, euler, beta scheduler. Varies sampler *and* scheduler against the rest of the grid, so it is only readable against v4e8, not against the res_multistep columns.
77s · floor -55.5 · med -51.4 · peak -0.1 dB · jitter 283.7 · detail 0.035
v4-step600 EMA · 4 steps
Probes the one regression larryvrh documents for v4: motion smear / trailing ghosting at 4 steps under fast motion, which 6–8 steps is said to remove. rapid-cuts, hands-dexterity and polyphony-foreground are where it should show.
45s · floor -54.6 · med -52.9 · peak -1.5 dB · jitter 338.5 · detail 0.126
v4-step600 EMA · 8 steps + SageAttention
Re-bases the production accelerator onto v4. emck8-sage (45 s, 3.66×, 0.955 correlation) is the current production config; if v4e8 dethrones emck8 this is what replaces it. Sage is an attention-backend swap and so should be weight-independent — this checks that.
50s · floor -51.7 · med -48.8 · peak 2.0 dB · jitter 266.8 · detail 0.035
ema-ckpt500 · 8 steps · euler + beta
Control for v4e8-eb, which varies LoRA, sampler and scheduler at once. Putting drbaph's euler+beta on the weights we already have 10 scenes of isolates the sampler/scheduler axis, so an eb win can be attributed to the config rather than to v4.
77s · floor -59.5 · med -57.3 · peak -5.2 dB · jitter 260.8 · detail 0.026

Measured

configtimefloormedianpeakjitterdetailraw detailluma
baseline · 20 steps166s-55.1-52.1-2.1224.70.03610696.0
ema-ckpt500 · 8 steps77s-58.3-54.10.9289.50.0279093.7
ckpt500 (non-EMA) · 8 steps77s-53.0-45.43.0253.40.05114986.9
lightx2v · 4 steps · str 1.045s-63.1-56.52.7302.40.04914198.1
lightx2v · 4 steps · str 1.0 · shift 12/645s-73.9-63.11.0298.20.04813998.4
lightx2v · 4 steps · str 1.0 · er_sde45s-63.6-59.62.3287.10.04210896.3
lightx2v · 4 steps · str 0.7545s-62.8-56.23.1301.50.05013396.0
lightx2v · 4 steps · str 0.75 · shift 12/645s-75.5-60.93.0295.30.04813095.8
lightx2v · 4 steps · str 0.75 · er_sde45s-64.8-60.12.0250.50.0349696.0
lightx2v · 8 steps · str 1.077s-59.3-52.41.3319.10.04511894.5
lightx2v · 8 steps · str 1.0 · shift 12/676s-66.4-55.82.2316.60.04511795.1
lightx2v · 8 steps · str 0.75 · er_sde77s-60.1-57.02.2287.50.037102100.8
baseline · 20 steps + Spectrum107s-54.1-50.9-3.6288.10.0277489.6
ema-ckpt500 · 8 steps + Spectrum61s-53.7-51.9-4.1289.60.0308991.7
ema-ckpt500 · 8 steps + SageAttention50s-62.0-56.5-2.1307.80.03210693.2
ema-ckpt500 · 8 steps + Sage + Spectrum41s-56.6-54.4-0.0283.20.0267490.8
baseline · 20 steps + SageAttention98s-55.0-52.1-4.7231.40.0349595.2
baseline · 20 steps + Sol-Attn98s-51.3-48.70.3299.70.0266890.6
ema-ckpt850 · 8 steps77s-56.0-49.9-0.6297.80.03912288.9
v4-step600 EMA · 8 steps78s-52.3-49.01.9253.60.0309391.7
v4-step600 non-EMA · 8 steps77s-50.2-46.9-1.2260.50.0359685.8
v4-step600 EMA · 8 steps · euler + beta77s-55.5-51.4-0.1283.70.0359590.0
v4-step600 EMA · 4 steps45s-54.6-52.9-1.5338.50.126386110.6
v4-step600 EMA · 8 steps + SageAttention50s-51.7-48.82.0266.80.03510290.0
ema-ckpt500 · 8 steps · euler + beta77s-59.5-57.3-5.2260.80.0267194.1
Rendered on ComfyUI 0.30.0 at commit 344b4398 (2026-08-07, "Support asym w4a8_int (#15308)") · torch 2.12.1+cu130 · NVIDIA GeForce RTX 5080, 15.4 GB
Custom nodes: ComfyUI-Spectrum-MiniMax-H3 1a06e20 (v0.1.9) — used by the two Spectrum columns only · comfyui-kjnodes 1.4.9 · ComfyUI-sol-attn 251c27f
All 14 columns rendered on this single ComfyUI commit. It postdates bdcb886a (MiniMax audio-sampler rework, merged 2026-08-06), which is why round-3 clips are archived separately rather than mixed in.