CUTIE: Beyond PetaOp/s/W Ternary DNN Inference Acceleration with\n Better-than-Binary Energy Efficiency

We present a 3.1 POp/s/W fully digital hardware accelerator for ternary\nneural networks. CUTIE, the Completely Unrolled Ternary Inference Engine,\nfocuses on minimizing non-computational energy and switching activity so that\ndynamic power spent on storing (locally or globally) intermediate results is\nminimized. This is achieved by 1) a data path architecture completely unrolled\nin the feature map and filter dimensions to reduce switching activity by\nfavoring silencing over iterative computation and maximizing data re-use, 2)\ntargeting ternary neural networks which, in contrast to binary NNs, allow for\nsparse weights which reduce switching activity, and 3) introducing an optimized\ntraining method for higher sparsity of the filter weights, resulting in a\nfurther reduction of the switching activity. Compared with state-of-the-art\naccelerators, CUTIE achieves greater or equal accuracy while decreasing the\noverall core inference energy cost by a factor of 4.8x-21x.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC