At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference

Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of inference for Transformer models. In the moderate-sparsity regime, Gustavson's dataflow provides a natural execution model for exploiting both activation and weight sparsity on vector processors through metadata-driven indexed accumulation. However, existing…

Paper

Similar papers

© 2026 NYSGPT2525 LLC