SparsePR results

Quality and efficiency across four models.

SparsePR is evaluated on text-to-video, image-to-video, and physical-world prediction models. Density includes routing and all exact probe pairs.

Table 1

Quality and efficiency

Reference fidelity, task quality, executed-pair density, attention PFLOPs, and end-to-end speedup.
ModelMethodPSNR ↑SSIM ↑LPIPS ↓ImgQual ↑SubCons ↑PBench ↑Density ↓PFLOPs ↓E2E ↑
HunyuanVideo-13BDense–––0.8500.976–100.0%612.381.00×
HunyuanVideo-13BSpargeAttn†25.310.8320.2170.7630.928–40.19%389.761.38×
HunyuanVideo-13BSVG2†29.870.9070.1210.8500.927–25.45%299.022.30×
HunyuanVideo-13BSVOO†24.870.8430.2240.6790.976–33.26%345.812.17×
HunyuanVideo-13BSVG-EAR†30.540.9180.0980.8450.903–22.17%281.861.93×
HunyuanVideo-13BSparsePR31.840.9320.0870.8500.976–21.92%255.952.61×
Wan2.2-I2V-A14BDense–––0.6890.974–100.0%658.461.00×
Wan2.2-I2V-A14BSpargeAttn†26.740.8710.1160.6800.953–30.15%396.831.58×
Wan2.2-I2V-A14BSVG2†28.180.8780.1050.6680.970–31.28%393.951.59×
Wan2.2-I2V-A14BSVOO†29.670.9130.0950.6890.973–31.67%389.621.61×
Wan2.2-I2V-A14BSVG-EAR†29.750.9210.0860.6800.970–23.64%378.881.61×
Wan2.2-I2V-A14BSparsePR30.660.9070.0440.6870.973–21.97%336.551.80×
Cosmos-Predict2.5-14BDense–––0.7140.97677.76100.0%526.871.00×
Cosmos-Predict2.5-14BSVG220.070.6240.3300.6780.89676.1428.81%286.511.24×
Cosmos-Predict2.5-14BSVOO22.060.6850.2890.7010.90976.0337.63%315.381.03×
Cosmos-Predict2.5-14BSVG-EAR25.540.9080.0620.7100.97677.7629.75%289.691.10×
Cosmos-Predict2.5-14BSparsePR26.330.9420.0680.7140.97677.7522.14%253.611.51×
Cosmos3-Nano-16BDense–––0.7000.95077.31100.0%127.611.00×
Cosmos3-Nano-16BSVG222.450.7350.2160.6770.91575.0337.29%96.191.16×
Cosmos3-Nano-16BSVOO16.640.5730.3810.7070.96277.5967.32%108.011.02×
Cosmos3-Nano-16BSVG-EAR21.160.7090.2610.6580.87272.8537.18%96.141.10×
Cosmos3-Nano-16BSparsePR24.420.8010.1760.6990.94977.3025.96%81.791.48×
Error reduction from response-coupled partitioning and full-generation latency breakdown
At matched density, response-coupled partitioning reduces mean error by 38.7% to 71.4% and p99 error by 33.1% to 60.6%. SparsePR reaches 1.80× end-to-end speedup on Wan2.2, with probe repair using 1.1% of total latency.

Table 2

Partitioning and reconstruction ablations

Mean and p99 normalized attention output error at 22% total executed-pair density. Lower is better.
ConfigurationHunyuanVideoWan2.2Cosmos-Predict2.5Cosmos3-Nano
Semantic partition0.089 / 0.7140.164 / 1.7340.791 / 7.5320.359 / 3.356
Key-response K/V partition0.085 / 0.8120.156 / 1.6590.769 / 7.2700.341 / 3.174
Response-coupled partition0.073 / 0.6960.108 / 0.6470.761 / 7.2270.331 / 3.150
Semantic + probe repair0.053 / 0.3560.105 / 0.8190.262 / 0.8260.124 / 0.668
SparsePR0.030 / 0.2070.063 / 0.3230.075 / 0.4890.076 / 0.447