001 Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories arXiv Paper Zhaoji Wang, Wanyu Si et al. Yesterday 002 SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions arXiv Paper Jennifer D’Souza, Sameer Sadruddin et al. Yesterday 003 Scientific Knowledge Discovery in the Age of Large Language Models arXiv Paper Eleni Adamidi, Serafeim Chatzopoulos et al. 2 days ago 004 F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill arXiv Paper F. Krebs 3 days ago 005 From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages arXiv Paper Théotime de la Selle 4 days ago 006 Making Mathematical Knowledge Explainable, Accessible and Interoperable Through Large Language Model Integration arXiv Paper Jan Range, B. Schembera et al. 4 days ago 007 Robust Interpretation of Historical Documents in Knowledge Graphs Through Query Inference and Execution arXiv Paper S. Nicolau, Adrià Molina et al. 4 days ago 008 Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature arXiv Paper Tanjin He, Aikaterini Vriza et al. 5 days ago 009 An Ontology for Machine Learning Interatomic Potentials arXiv Paper D. Hernández, Jong Hyun Jung et al. 6 days ago 010 From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics arXiv Paper Muhsen Hammoud Jul 23 011 Scientific exploration, collaboration and labor division in the large language model era arXiv Paper Xiang Zheng, Xi Hong et al. Jul 23 012 Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI arXiv Paper Deyu Jing Jul 23 013 TINY_SCHILLER: A Drop-In German Drama Corpus for Small Language Models arXiv Paper Mark Schutera Jul 22 014 Understanding Generative AI-mediated User Engagement with Academic Library Resources arXiv Paper Hae Min Kim, Stacy Stanislaw Jul 22 015 Using Hierarchical Controlled Vocabularies to Understand CLIP Retrieval Failures in Historical Photo Collections arXiv Paper Ratan J. Sebastian, Anett Hoppe et al. Jul 22 016 Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field arXiv Paper Tanay Aggarwal, Angelo Salatino et al. Jul 20 017 MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation arXiv Paper Hyewon Lee, Minkyung Song et al. Jul 20 018 Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries arXiv Paper Mohammad Arvan, Amber E. Osterholt et al. Jul 18 019 Translating AI into scientific impact: Field context, career position, and institutional capability in AI-enabled research arXiv Paper Zhiyong Tan, Hongkan Chen et al. Jul 18 020 ICAConfPubs: A Dataset and User Interface for ICA Conference Papers (2003-2018) arXiv Paper Hongtao Hao, Xin-Yen Chen et al. Jul 15 021 Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls arXiv Paper Michael Paris, Hande Celikkanat et al. Jul 15 022 Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026) arXiv Paper Olivier Martinez Jul 15 023 Towards Nexus-Score: Metadata Gaps Limit Scholarly AI Attribution arXiv Paper Aadi Narayana Varma Dantuluri, Sushrut Thorat et al. Jul 14 024 Characterising AI Models for Cataloguing arXiv Paper Miguel Arana-Catania, Neil Jefferies Jul 13 025 Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud arXiv Paper Balint Gyevnar, Atoosa Kasirzadeh et al. Jul 12 026 From peer review nuances to best practices arXiv Paper Sheng Lu Jul 12 027 Return of the solo author: The changing division of labor in science in the age of generative AI arXiv Paper Akira Matsui Jul 12 028 Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works arXiv Paper Miguel Arana-Catania, Gillian Pink et al. Jul 10 029 Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI arXiv Paper Miguel Arana-Catania, Catherine Conisbee et al. Jul 10 030 SetGo: Metadata Readiness for Scientific AI Datasets arXiv Paper Sean R. Wilkinson et al. Jul 10 031 An Autonomous Scientific Knowledge Generation Framework for AI-Driven Scientific Discovery arXiv Paper Dibakar Datta Jul 9 032 Conversational Retrieval and On-the-Fly Knowledge Modeling of Historical Penitentiary Repression Records arXiv Paper Paula Font Solà et al. Jul 9 033 A Word-Level Digital Reader of the Prasthanatrayi with Sankara's Bhasya: Corpus, Method, and an Open, Offline Reading Aid for the Advaita Vedanta Canon arXiv Paper Tamal Maharaj Jul 8 034 AAAI-26 Dual Submissions: Novel Challenges arXiv Paper K. Wagstaff, Joydeep Biswas et al. Jul 7 035 Whose fairness? Structural concentration in AI bias research arXiv Paper Abhash Shrestha, Subigya Gautam et al. Jul 6 036 Publishing Without Journals: An Open, Forkable Archive with Attributed Review arXiv Paper Matthew Lorig Jul 5 037 Large-scale dataset of automatically classified rhetorical sections in scientific papers arXiv Paper Daniel Verdi, Jacob Aarup Dalsgaard et al. Jul 3 038 Gender Differences in Research Topic and Method Selection in Library and Information Science: Perspectives from Three Top Journals arXiv Paper Chengzhi Zhang, Siqi Wei et al. Jul 2 039 Non-synchronism in Global Usage of Research Methods in Library and Information Science from 1990 to 2019 arXiv Paper Chengzhi Zhang, Liang Tian Jul 2 040 CHARLIE: An On-Premise Multi-Agent Retrieval-Augmented Generation System for Evidential Reasoning in Forensic Science arXiv Paper Leandro D. Carneiro, Andre L. S. Meirelles et al. Jul 1 041 Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences arXiv Paper M. Russinovich, Ram Shankar Siva Kumar et al. Jul 1 042 Reading Order Inference for Complex Document Layouts arXiv Paper Iddo Hakim, Sharva Gogawale et al. Jul 1 043 Automated High-Precision Extraction and Forensic Verification of Data-Bearing Vector Figures arXiv Paper Bowen Sun, Chaowei Xiao Jun 30 044 Building a Multimodal Dataset of Academic Paper for Keyword Extraction arXiv Paper Jingyu Zhang, Xinyi Yan et al. Jun 30 045 DEMUN: Fast and accurate discovery of music notation in very large collections arXiv Paper Vojtěch Dvořák et al. Jun 30 046 Exploring the relationship between team institutional composition and novelty in academic papers based on fine-grained knowledge entities arXiv Paper Ziling Chen, Chengzhi Zhang et al. Jun 30 047 Linking Hadith Narrator Identities Across Heterogeneous Arabic Biographical Databases: A Multi-Signal Entity Resolution Pipeline arXiv Paper Taufiq Wirahman Jun 30 048 Towards a foundational model for recognising diastematic Gregorian notation arXiv Paper Daniel Kurek et al. Jun 30 049 Usage frequency and application variety of research methods in library and information science: Continuous investigation from 1991 to 2021 arXiv Paper Chengzhi Zhang, Liang Tian et al. Jun 30 050 Exploring Motivations for Algorithm Mention in the Domain of Natural Language Processing: A Deep Learning Approach arXiv Paper Yuzhuo Wang, Yi Xiang et al. Jun 29 051 Research Entity Extraction and Topic Detection from UKRI Grant Proposals arXiv Paper Xingran Ruan, Angelo Salatino et al. Jun 29 052 Revealing the Technology Development of Natural Language Processing: A Scientific Entity-Centric Perspective arXiv Paper Heng Zhang, Chengzhi Zhang et al. Jun 29 053 Submission Responsibility Matters: Role-Aware Submission Quotas under Coauthorship arXiv Paper Furkan Mumcu, Yasin Yilmaz Jun 29 054 Unveiling Novelty Evolution in the field of Library and Information Science in China arXiv Paper Chen Yang, Yuzhuo Wang et al. Jun 29 055 Em-ergence of the em-dash: a population-level rise in em-dash frequency in medRxiv preprints at the dawn of the large-language-model era arXiv Paper Przemysław Czuma Jun 28 056 AICID: Unique Identifiers for AI Scientists arXiv Paper Clément Vidal, Martin Monperrus Jun 27 057 Attribution Bias in Philosophical Knowledge Graphs: Corpus Frequency versus Temporal Sourcing arXiv Paper Joy Bose Jun 27 058 Categorizing Mathematical Concepts with LLM Voting Ensembles in Mathswitch arXiv Paper Katja Bervcivc, Slobodan Stanojevikj Jun 27 059 Customized Generative AI Agent for Transportation Engineering Practice: A Development and Continued Pre-training Guideline arXiv Paper Dianwei Chen, Yuan-Zheng Lei et al. Jun 27 060 Mitigating LLM-based p-Hacking by Preregistering for the Next LLM arXiv Paper Mariam Thomas, Kristina Gligori'c et al. Jun 26