Model evaluation for extreme risks
2023/05/24 by Toby Shevlane, Sebastian Farquhar, Shevlane, Toby +41 · 1 voice · 40 citations
Computer Science · #Information and Cyber Security #Software Engineering Research #Software Reliability and Analysis Research #cs.AI
paper · pdf · doi:10.48550/arxiv.2305.15324
arxiv published 2023/05/24 · arxiv updated 2023/09/22
Abstract
Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities. Further progress in AI development could lead to capabilities that pose extreme risks, such as offensive cyber capabilities or strong manipulation skills. We explain why model evaluation is critical for addressing extreme risks. Developers must be able to identify dangerous capabilities (through "dangerous capability evaluations") and the propensity of models to apply their capabilities for harm (through "alignment evaluations"). These evaluations will become critical for keeping policymakers and other stakeholders informed, and for making responsible decisions about model training, deployment, and security.
Cited by
- Solipsistic Superintelligence is Unlikely to be Cooperative
- ContainmentBench: Trace-Based Evaluation of Post-Injection Containment in Tool-Using LLM Agents
- Aligning Artificial Superintelligence via a Multi-Box Protocol
- A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing
- Adversarial machine learning :
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skills
- Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks
- Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
- PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
- Against racing to AGI: Cooperation, deterrence, and catastrophic risks
- The Controllability Trap: A Governance Framework for Military AI Agents
- Technical Requirements for Halting Dangerous AI Activities
- Domestic frontier AI regulation, an IAEA for AI, an NPT for AI, and a US-led Allied Public-Private Partnership for AI: Four institutions for governing and developing frontier AI
- From Turing to Tomorrow: The UK's Approach to AI Regulation
- On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
- The Singapore Consensus on Global AI Safety Research Priorities
- A Conceptual Framework for AI Capability Evaluations
- Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
- UCD: Unlearning in LLMs via Contrastive Decoding
- Benchmarking Misuse Mitigation Against Covert Adversaries
- Misalignment or misuse? The AGI alignment tradeoff
- Emergent Risk Awareness in Rational Agents under Resource Constraints
- Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks
- What Is AI Safety? What Do We Want It to Be?
- AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
- Evaluating Frontier Models for Stealth and Situational Awareness
- Capabilities Ain't All You Need: Measuring Propensities in AI
- Assessing LLM code generation quality through path planning tasks
- A Framework to Assess the Persuasion Risks Large Language Model Chatbots Pose to Democratic Societies
- Understanding and Mitigating Risks of Generative AI in Financial Services
- Accountability Asymmetry and Structural Trust in Autonomous AI Systems
- "Allow" to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents
- A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models
- Audit Cards: Contextualizing AI Evaluations
Discussions
Related