Library

Subject
Tags

188 matches · Dangerous capabilities

#
001The Biosecurity Blind Spot: Systematic Dual-use Detection in Open Science InfrastructurearXivPaperVasudha Sharma, C. Singh et al.May 10
002Are we Doomed to an AI Race? Why Self-Interest Could Drive Countries Towards a Moratorium on SuperintelligencearXivPaperEdward Roussel, Lode Lauwaert et al.May 2
003Code World Model Preparedness ReportarXivPaperD. Song, Peter Ney et al.May 1
004Exploration Hacking: Can LLMs Learn to Resist RL Training?arXivPaperEyon Jang, Damon Falck et al.Apr 30
005Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligencearXivPaperTommy Shaffer Shane, Simon Mylius et al.Apr 10
006Chain-of-Authorization: Embedding authorization into large language modelsarXivPaperYang Li, Yule Liu et al.Mar 24
007Consequentialist Objectives and CatastrophearXivPaperHenrik Marklund, Alex Infanger et al.Mar 16
008RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift StudiesarXivPaperPatricia Paskov, Kevin Wei et al.Mar 11
009The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational AwarenessarXivPaperSubramanyam Sahoo, Aman Chadha et al.Mar 10
010Jailbreaking Embodied LLMs via Action-level ManipulationarXivPaperXinyu Huang, Qiang Yang et al.Mar 2
011A Multi-Turn Framework for Evaluating AI Misuse in Fraud and Cybercrime ScenariosarXivPaperKimberly T. Mai, Anna Gausen et al.Feb 25
012Implicit Intelligence -- Evaluating Agents on What Users Don't SayarXivPaperVed Sirdeshmukh, Marc WetterFeb 23
013Measuring Mid-2025 LLM-Assistance on Novice Performance in BiologyarXivPaperShenda Hong, Alexander Kleinman et al.Feb 18
014Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report v1.5arXivPaperDongrui Liu, Yi Yu et al.Feb 16
015Small models, big threats: Characterizing safety challenges from low-compute AI modelsarXivPaperPrateek PuriJan 29
016VERGE: Formal Refinement and Guidance Engine for Verifiable LLM ReasoningarXivPaperVikash Singh, Darion Cassel et al.Jan 27
017AI Regulation Regimes and Competitive Outcomes: A Game-Theoretic Analysis of Regulatory Competition in Frontier TechnologiesSemantic ScholarPaperS. Paul, Herman SahniJan 1
018AI Systems That Think, Team, and FightSemantic ScholarPaperSvitlana VolkovaJan 1
019AI-Driven Navigational Assistance for Visually Impaired PersonsSemantic ScholarPaperVenkatesh, Jahnavi Thupalli et al.Jan 1
020ARTIFICIAL INTELLIGENCE IN HOSPITALITY: TRANSFORMING THE HOTEL INDUSTRY WHILE PRESERVING EMPATHY AND TRADITIONAL HOSPITALITYSemantic ScholarPaperOlaoluwa FowosereJan 1
021Advancing Knotted Protein Design with ESM3: Guided Generation and Topological InsightsSemantic ScholarPaperEva Maršálková, Petr ŠimečekJan 1
022Artificial Intelligence and the Future of Strategic StabilitySemantic ScholarPaperMichael C. HorowitzJan 1
023Awareness, Competence, and Perceptions of Augmented Reality, Virtual Reality, and the Metaverse among Pakistani University LibrariansSemantic ScholarPaperHina Sardar, Muhammad Kabir KhanJan 1
024Beyond Content Filtering: A “Circuit Breaker” Architecture for Autonomous Agent Action SafetySemantic ScholarPaperH. Sridharan, Reshma NairJan 1
025Designing an AI-Enhanced Web Crawler for Semantic Data Extraction and Knowledge Graph ConstructionSemantic ScholarPaperChitiz TayalJan 1
026Emotional Resonance Matching for Personalized E-Commerce Re-Engagement System Using Machine Learning and Generative AISemantic ScholarPaperS. MJan 1
027From Data to Decisions: Harnessing the Potential of Language Based AI in DrillingSemantic ScholarPaperC. Chatar, P. ShethJan 1
028From Emotion to Action: AURORA’s Sentiment-Based Model for Tourism DecisionsSemantic ScholarPaperM. Badouch, M. BoutaounteJan 1
029Frontier Safety Policies for AI Emergency Preparedness in ChinaSemantic ScholarPaperJames Zhang, Miles Kodama et al.Jan 1
030GenAI-Powered Autonomous Cyber Offense-Defense: An Explainable LLM Red-vs-Blue Simulation and Self-Defense FrameworkSemantic ScholarPaperHaitian DuJan 1
031Multi-Agentic Generative AI Framework for Accelerating Field Development PlanningSemantic ScholarPaperS. Ramatullayev, S. Su et al.Jan 1
032Personalised LLMs and the risks of the digital twin metaphorSemantic ScholarPaperM. Annoni, Davide Battisti et al.Jan 1
033Probing Adversarial Robustness of Protein Language Models: A Reproducible Case Study of ESM-2 Under Substitution-Based AttacksSemantic ScholarPaperMd. Robiul Islam Niloy, Zafer Aydin et al.Jan 1
034RitualLab: An AI-in-the-loop Collaborative Platform for Designing and Auditing Well-being-Aware Brand NarrativesSemantic ScholarPaperYong Zhen, Ethan Jing Li et al.Jan 1
035Sentinel-Math: A Framework for Neural Intent Monitoring and SafetyCircuit-Breaking Against Semantically Disguised Logical Tasks in LLMsSemantic ScholarPaperJing Zhang, Yaowei Wang et al.Jan 1
036Six misconceptions about large language models: A minimal model and diagnostic taxonomySemantic ScholarPaperZhicheng LinJan 1
037Structural Capability Containment in Advanced AI Systems A Foundational Framework for Multi-Jurisdiction Safety ArchitecturesSemantic ScholarPaperGaston ReyJan 1
038Students’ Experiences Using an AI-Powered Speaking Tool in English Courses in Higher Education [Abstract]Semantic ScholarPaperIlan Daniels Rahimi, Gila Cohen Zilka et al.Jan 1
039Technical Note Three: The Horizon Scan From V = 0 – Compositional Opacity And Catastrophic Risk Under Quantum IntegrationSemantic ScholarPaperAndrew DevinJan 1
040Biosecurity-Aware AI: Agentic Risk Auditing of Soft Prompt Attacks on ESM-Based Variant PredictorsarXivPaperHuixin ZhanDec 19, 2025
041How frontier AI companies could implement an internal audit functionarXivPaperF. Gómez, A. Buick et al.Dec 16, 2025
042Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real UsersarXivPaperManon Kempermann, Sai Suresh Macharla Vasu et al.Dec 11, 2025
043Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query ArchitecturearXivPaperGary Ackerman, Brandon Behlendorf et al.Dec 9, 2025
044Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models II: Benchmark Generation ProcessarXivPaperGary Ackerman, Z. Kallenborn et al.Dec 9, 2025
045Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models III: Implementing the Bacterial Biothreat Benchmark (B3) DatasetarXivPaperGary Ackerman, Theodore Wilson et al.Dec 9, 2025
046The Role of Risk Modeling in Advanced AI Risk ManagementarXivPaperChlo'e Touzet, H. Papadatos et al.Dec 9, 2025
047Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMsarXivPaperIgor Shilov, Alex Cloud et al.Dec 5, 2025
048Evaluating AI Providers' Frontier Safety FrameworksarXivPaperLily Stelling, Malcolm Murray et al.Dec 1, 2025
049PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic ApproacharXivPaperUdari Madhushani Sehwag, Shayan Shabihi et al.Nov 24, 2025
050Evaluating Adversarial Vulnerabilities in Modern Large Language ModelsarXivPaperTom PerelNov 21, 2025
051An International Agreement to Prevent the Premature Creation of Artificial SuperintelligencearXivPaperAaron Scher, David Abecassis et al.Nov 13, 2025
052LTD-Bench: Evaluating Large Language Models by Letting Them DrawarXivPaperLi Lin, Ke Li et al.Nov 4, 2025
053SciTrust 2.0: A Comprehensive Framework for Evaluating Trustworthiness of Large Language Models in Scientific ApplicationsarXivPaperEmily Herron, Junqi Yin et al.Oct 29, 2025
054Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent SafetyarXivPaperV. Bonagiri, P. Kumaraguru et al.Oct 18, 2025
055PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation CapabilitiesarXivPaperZicheng Liu, Lige Huang et al.Oct 13, 2025
056The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM ObjectivesarXivPaperMatthieu Bou, Nyal Patel et al.Oct 7, 2025
057How Catastrophic is Your LLM? Certifying Risk in ConversationarXivPaperChengxiao Wang, Isha Chaudhary et al.Oct 4, 2025
058Evaluation Awareness Scales Predictably in Open-Weights Large Language ModelsarXivPaperMaheep Chaudhary, Ian Su et al.Sep 10, 2025
059Constitutional Law and AI Governance: Constraints on Model Licensing and Research ClassificationarXivPaperA. Mark, Aaron ScherSep 3, 2025
060STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model ReportsarXivPaperTegan McCaslin, Jide Alaga et al.Aug 13, 2025

Showing 60 of 188 documents · scroll for more