Library

Subject
Tags

1,645 matches · cs.PF

#
001Bridging Compute- and Data-Optimal PretrainingarXivPaperTian Qin et al.3 days ago
002SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent ServingarXivPaperYihui Zhang et al.4 days ago
003A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, ForeverarXivPaperSietse Schelpe5 days ago
004Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery AlgorithmsarXivPaperYakov Kuzin et al.5 days ago
005FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple SiliconarXivPaperOm Mohite7 days ago
006Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAsarXivPaperIlia Sobakinskikh et al.7 days ago
007Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token ContextarXivPaperAlagappan ValliappanJul 23
008BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural AcceleratorsarXivPaperFabian Waschkowski, Prabod Rathnayaka et al.Jul 21
009Formulation-Level Auto-Tuning for QUBO-Based Machine Learning: A Case Study Across Multiple Quantum-Inspired AnnealersarXivPaperNaoya Mizuki, Takahiro Katagiri et al.Jul 21
010SALT: Salience-Aware Lexical Trie for Long-Context CompressionarXivPaperOteo Mamo, Hyunji Yi et al.Jul 20
011AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill WorkflowsarXivPaperTejas Singh Anand, Yuet Ying Christina Wang et al.Jul 16
012An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus MechanismarXivPaperMobina Kashaniyan, M. Ashtiani et al.Jul 16
013Toward Energy-Efficient and Low-Power Arrhythmia Detection for Wearable DevicesarXivPaperF. Bulten, Yawar Rasheed et al.Jul 16
014Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge FlywheelarXivPaperSietse SchelpeJul 15
015The Café in Amsterdam: When the Incumbent Becomes the OraclearXivPaperAugusto CamargoJul 15
016Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUsarXivPaperWeijia Han, Lisha QuJul 13
017Lightning Fast Matching Dependency Discovery with DesbordantearXivPaperAlexey Shlyonskikh, Michael Sinelnikov et al.Jul 12
018Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM ConfigurationsarXivPaperNada Zine, Tristan Coignion et al.Jul 10
019On-Device Adaptive Battery Power Prediction for Electric VehiclesarXivPaperAvik Bhatnagar, Anton Paule et al.Jul 10
020STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPUarXivPaperV. J. Jung, Gagandeep Singh et al.Jul 10
021Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvencyarXivPaperSatoshi MatsuokaJul 8
022Energy-Efficient GPU DVFS for Fine-Tuning of SLMs on Resource-constrained Embedded DevicesarXivPaperJurn-Gyu Park, S. Zholdybayev et al.Jul 7
023Think Before You Grid-Search: Floor-First Triage for LLM ServingarXivPaperYihua LiuJul 7
024Adaptive Inference Batching using Policy GradientsarXivPaperRuslan SharifullinJul 6
025Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS ProcessesarXivPaperDa-Hyun SonJul 6
026HiFA4: Training-Free 4-bit FlashAttention on Ascend HIF4 NPUs for LLM InferencearXivPaperHuilin Dong, Yanzhao Li et al.Jul 5
027Efficient Discovery of Conditional Dependencies with DesbordantearXivPaperI. Kozhukov, D. Fedoseev et al.Jul 4
028Energy-Aware System-Level Evaluation of Post-Quantum TLS on Embedded User Equipment over a Disaggregated 5G NetworkarXivPaperSanzida Hoque, Abdullah AydegerJul 4
029Scalable Maximal Frequent Episode Mining with DesbordantearXivPaperMaxim Ivanov, Matvei Smirnov et al.Jul 3
030BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native MetalarXivPaperPrabod Rathnayaka, Fabian Waschkowski et al.Jul 1
031LUMA: Benchmarking Segmentation via a Lightweight Universal Mask AdapterarXivPaperT. Nauen, Anosh Billimoria et al.Jul 1
032TraceLab: Characterizing Coding Agent Workloads for LLM ServingarXivPaperKan Zhu, Mathew Jacob et al.Jun 29
033KernelSight-LM: A Kernel-Level LLM Inference SimulatorarXivPaperXiteng Yao, Taeho Kim et al.Jun 26
034Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM ServingarXivPaperYasmin Moslem, Magdalena Kacmajor et al.Jun 25
035Compiler-Driven Approximation Tuning for Hyperdimensional ComputingarXivPaperXavier Routh, Abdul Rafae Noor et al.Jun 25
036Axon: A Synthesizing Superoptimizer for Tensor ProgramsarXivPaperAkash Kothari, Shaowei Zhu et al.Jun 24
037SOLAR: AI-Powered Speed-of-Light Performance AnalysisarXivPaperQijing Huang, S. Damani et al.Jun 24
038TileMaxSim: IO-Aware GPU MaxSim Scoring with Dimension Tiling and Fused Product QuantizationarXivPaperAshutosh SharmaJun 24
039AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set ApproacharXivPaperM. Gazzola, H. Souto et al.Jun 23
040Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted GenerationarXivPaperSijie Wang, Zhengyu Qing et al.Jun 23
041CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight DisaggregationarXivPaperZhuoren Ye, Tianyu Wo et al.Jun 23
042Power-Flexible AI Data Centers: A New Paradigm for Grid-Responsive ComputearXivPaperChris Williams, Philip Colangelo et al.Jun 23
043Learning Filters with CertaintyarXivPaperYuval Banoun, D. S. Menasché et al.Jun 22
044The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential ComputingarXivPaperHang Yin, Kevin WangJun 22
045Enabling Cloud-Level Accuracy in Edge AI through IoT Data PreprocessingarXivPaperAygun Varol, Katarzyna Kołodziej et al.Jun 21
046Load Testing for Machine Learning Model Serving Systems at ScalearXivPaperAmr S. Abdelfattah, Nakul Tirumalai et al.Jun 20
047Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical StudyarXivPaperAlfarizy Alfarizy, H. Nguyen et al.Jun 19
048UltraQuant: 4-bit KV Caching for Context-Heavy AgentsarXivPaperInesh Chakrabarti, David Limpus et al.Jun 18
049Balancing Bits and Drops: Stress-Adjusted Water Management for Data CentersarXivPaperZahidur Talukder et al.Jun 15
050Fractional Verkle Trees: A Hypertree Decomposition and Verified Proof Serialization Architecture for High-Performance Blockchain State AccumulatorsarXivPaperEkleen Kaur, E. FragaJun 15
051SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix ExtensionsarXivPaperFeiyang Chen, Haibo ChenJun 15
052MADAR: An Address-Free ProcessorarXivPaperMohamed Amine BergachJun 14
053Pseudonym Scheme Based on Hybrid Certificates for Security Credential Management System in Vehicular CommunicationsarXivPaperAbel C. H. Chen, F. Hwang et al.Jun 12
054A Modern Large-Scale Memory Characterization LaboratoryarXivPaperAtaberk Olgun, Haocong Luo et al.Jun 11
055GF-DiT: Scheduling Parallelism for Diffusion Transformer ServingarXivPaperXinwei Qiang, Yifan Hu et al.Jun 11
056The Price of Anarchy in Disaggregated InferencearXivPaperAthos GeorgiouJun 11
057AI Tokenomics: The Economics of Tokens, Computation, and Pricing in Foundation ModelsarXivPaperQuanyan ZhuJun 10
058XPR: An Extensible Cross-Platform Point-Based Differentiable RendererarXivPaperSteve Rhyner et al.Jun 10
059Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X ElitearXivPaperZhiyuan Cheng, Longying LaiJun 9
060Flash-GMM: A Memory-Efficient Kernel for Scalable Soft ClusteringarXivPaperGal Bloch, Ariel Gera et al.Jun 9

Showing 60 of 1,645 documents · scroll for more