SparsePR results

Quality and efficiency across four models.

SparsePR is evaluated on text-to-video, image-to-video, and physical-world prediction models. Density includes routing and all exact probe pairs.

Table 1

Quality and efficiency

Reference fidelity, task quality, executed-pair density, attention PFLOPs, and end-to-end speedup.
ModelMethodPSNR ↑SSIM ↑LPIPS ↓ImgQual ↑SubCons ↑PBench ↑Density ↓PFLOPs ↓E2E ↑
HunyuanVideo-13BDense0.8500.976100%612.381.00×
HunyuanVideo-13BSpargeAttn†24.5890.7960.2320.90840.09%389.761.38×
HunyuanVideo-13BSVG2†30.4520.9100.1170.8520.92725.45%299.022.30×
HunyuanVideo-13BSVOO†24.8790.8430.2240.67930.97992.17×
HunyuanVideo-13BSVG-EAR†31.0430.9280.0920.8450.90322.17%281.861.93×
HunyuanVideo-13BSparsePR31.8440.9320.0870.8500.97621.92%255.952.61×
Wan2.2-I2V-A14BDense0.6890.974100%658.461.00×
Wan2.2-I2V-A14BSpargeAttn†27.1400.8830.1160.6800.95830.15%396.831.58×
Wan2.2-I2V-A14BSVG2†26.5620.8610.1380.6680.95931.28%393.951.59×
Wan2.2-I2V-A14BSVOO†29.6780.9130.0950.73370.97311.61×
Wan2.2-I2V-A14BSVG-EAR†29.7590.9180.0930.6800.95923.64%378.881.61×
Wan2.2-I2V-A14BSparsePR30.6580.9070.0440.6870.97321.97%328.701.80×
Cosmos-Predict2.5-14BDense0.7140.97677.76100%526.871.00×
Cosmos-Predict2.5-14BSVG220.0750.6240.3300.6780.89676.1428.81%286.511.24×
Cosmos-Predict2.5-14BSVOO22.0660.6850.2890.7010.90976.0337.63%315.381.03×
Cosmos-Predict2.5-14BSVG-EAR25.5490.9080.0620.7100.97677.7829.75%289.691.10×
Cosmos-Predict2.5-14BSparsePR26.3280.9420.0680.7140.97677.7522.14%253.611.51×
Cosmos3-Nano-16BDense0.7000.95077.31100.00%90.011.00×
Cosmos3-Nano-16BSVG222.4580.7350.2160.6770.91575.0337.29%57.691.16×
Cosmos3-Nano-16BSVOO16.6420.5730.3810.7070.96277.5967.32%69.511.02×
Cosmos3-Nano-16BSVG-EAR21.1670.7090.2610.6580.87272.8537.18%57.641.10×
Cosmos3-Nano-16BSparsePR24.4170.8010.1760.6990.94977.3025.96%43.221.48×

† Reported by prior work. Rows without † are reproduced under matched hardware, sequence shape, and timing protocols.

Error reduction from response-coupled partitioning and full-generation latency breakdown
At matched density, response-coupled partitioning reduces mean and p99 error. SparsePR reaches 1.80× end-to-end speedup on Wan2.2, with probe repair using 1.1% of total latency.

Table 2

Partitioning and reconstruction ablations

Mean and p99 normalized attention-output error at 22% total executed-pair density. Lower is better.
ConfigurationHunyuanVideoWan2.2Cosmos-Predict2.5Cosmos3-Nano
Semantic partition0.0887 / 0.71360.1634 / 1.73380.7903 / 7.53180.3590 / 3.3557
Key-response K/V partition0.0851 / 0.81210.1560 / 1.65910.7686 / 7.27000.3409 / 3.1738
Response-coupled partition0.0736 / 0.69670.1489 / 1.64790.7617 / 7.22710.3315 / 3.1502
Semantic + probe repair0.0527 / 0.35620.1041 / 0.81860.2622 / 0.82600.1720 / 0.8648
SparsePR0.0330 / 0.22850.0707 / 0.43050.0954 / 0.57690.0822 / 0.4951