Leveraging Neural Graph Compilers in Machine Learning Research for Edge-Cloud Systems

This work presents a comprehensive evaluation of neural network graph compilers across heterogeneous hardware platforms, addressing the critical gap between theoretical optimization techniques and practical deployment scenarios. We demonstrate how vendor-specific optimizations can invalidate relative performance comparisons between architectural archetypes, with performance advantages sometimes completely reversing after compilation. Our systematic analysis reveals that graph compilers exhibit performance patterns highly dependent on both neural architecture and batch sizes. Through fine-grained block-level experimentation, we establish that vendor-specific compilers can leverage repeated patterns in simple architectures, yielding disproportionate throughput gains as model depth increases. We introduce metrics to quantify a compiler’s ability to mitigate performance friction as batch size increases. The resulting methodology turns graph-compiler benchmarking into design feedback for ML systems, allowing researchers to validate whether architectural, batching, and resource-utilization conclusions remain valid after compilation on target hardware.

Paper

Similar papers

© 2026 NYSGPT2525 LLC