Papers1 provider · 1 record
March 25, 2025· SuperIntelligence - Robotics - Safety & Alignment
article
Open access

Highlights of the Issue

Abstract

Highlights of the IssueKris Carlson, Publisher and Editor-in-ChiefOur second issue surveys state of the art of large language models (LLMs) with an emphasis on safety and value alignment. Superintelligence StrategyDan Hendrycks, Eric Schmidt, Alexandr WangSeeking Stability in the Competition for AI Advantage: Commentary on Superintelligence StrategyIskander Rehman, Karl P. Mueller, Michael J. Mazarr (RAND Corp.) I recommend the RAND Corp critique by knowledgeable military policy analysts over the Hendrycks et al. article. The RAND article is illuminative, incisive, covers Superintelligence Strategy’s key points, and suggests critical reasoning flaws in their mutually-assured-AI-malfunction (MAIM} policy. Although it is valuable to compare the nuclear and AI revolutions in search of instructive parallels and insights, the differences between the technologies and their respective ecosystems have deep strategic implications. Taking these into account, we have concerns regarding both the practical viability of the MAIM concept as an approach to overcoming instability risks in the AI race and the potential escalatory dangers that could follow from its core prescriptions.— Rehman et al. pg. 1 Surely we’d like to avoid repeating the mutually-assured-destruction (MAD) policy. The MAD policy alone could trigger AGI taking over for their and our security. But we must realize that strategies like MAD and MAIM are considered in the US, its allies, and adversaries. And we must try to understand them in order to avoid them. Highlights of the critique: First, the report refers loosely to an array of actions that states might take to cripple a rival's architecture for developing advanced AI.... [which] assumes that adversary AI programs will have specific facilities that can be readily located and disrupted. However, distributed cloud computing, decentralized training, and algorithmic development increasingly may not require centralized physical locations, making AI systems more resilient to limited attacks…. The following critique argues for distributed autonomous organizations (DAO) as I advocated in Safe Artificial General Intelligence via Distributed Ledger Technology and Provably Safe Artificial General Intelligence via Interactive Proof Systems. A second practical challenge resides in the expectation that each party can accurately assess secretive AI progress by others and gauge when preventive action would be necessary. Contrary to what is averred in the report, it is unlikely that states will have a clear sense of when the moment has arrived to MAIM their opponent…. Third and finally, even a credible MAIM threat might not deter a rival from pursuing superintelligent AI. Halting one's AI development would entail essentially the same costs as being the victim of a MAIM attack — loss of the program. And here’s another critique: MAD did not seek to deter the development of weapons but instead their use, which made the threshold for response vastly simpler (though it could still be problematic in cases such as false or ambiguous warnings of attacks). We would like to hear, or be pointed to, policy alternatives to MAIM that incentivize AGI developers to move toward AGI that can be proven to benefit all of humanity. Humanity’s Last Exam (HLE)Long Phan, Alice Gatti, Ziwen Han, and Nathaniel Li are first-listed members of the Organizing Team, and have hundreds of co-author/collaborators. This very large-scale collaborative effort has an ambitious title. The authors note:[LLM] benchmarks are not keeping pace in difficulty [with LLM capabilities]: LLMs now achieve over 90% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce HUMANITY’S LAST EXAM (HLE), a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. HLE consists of 2,700 questions across dozens of subjects. I do not find any mention of the terms, ‘training set leakage into test set data’ or ‘test set contamination.’ But those issues aside, it seems to be the toughest test set yet – at least as of this writing (16 March 2025) before the LLMs learn the answers and can regurgitate them and reasonably close variants, at which point there will need to be a fresh ‘last exam.’ Kudos to the organizing authors. It’s interesting that frontier LLMs performed dramatically poorer on HLE than on previous benchmark tests, which is a tribute to the originality of the questions. Pathways to Short Transformational AI TimelinesZershaaneh Qureshi We excerpt here a chapter from the complete text. To understand this chapter note that the article distinguishes between two types of recursive self-improvement (RSI): • Direct recursive improvement: positive feedback loops which are mediated directly by AI systems. • Indirect recursive improvement: positive feedback loops that are not mediated directly by AI, such as economic feedback loops (driven by reinvestment of capital into AI R&D), scientific feedback loops (driven by advancements in scientific tools and methods) and political feedback loops (driven e.g. by competitive pressures/race dynamics) (pp. 15-16). HyperWrite, edited: The complete article outlines a framework for analyzing different scenarios that could lead to Transformative AI (TAI) within the next 10 years. Key parameters considered are: 1. Compute scaling dynamics (whether progress continues or hits bottlenecks)2. Indirect feedback loop dynamics (whether they can overcome scaling bottlenecks)3. Direct recursive improvement (DRI) timeline (before or after 2035)4. DRI strength (cannot sustain, sustains, or accelerates progress) Seven possible scenarios are: 1. "Straight Path" - Compute scaling continues successfully2. "Rising Tide" - Indirect recursive improvement (IRI) overcomes bottlenecks3. "New Spark" - Moderate direct recursive improvement maintains progress4. "New Engine" - Strong DRI accelerates progress5. "Dual Engine" - Combination of compute scaling and DRI6. "LLM Hybrid" - Hybrid AI systems enable TAI7. "Intelligent Network" - Networks of AI systems enable TAI The author argues that this variety of plausible pathways strengthens the case for short TAI timelines, as TAI could emerge through multiple different mechanisms rather than requiring one specific path to succeed. Please send pointers and commentary on AI timelines and recursive self-improvement to editor@s-rsa.com. The Road to Artificial SuperIntelligence: A Comprehensive Survey of SuperalignmentHyunJin Kim, Xiaoyuan Yi, JinYeong Bak, Jing Yao, Jianxun Lian, Muhua Huang, Shitong Duan, Xing Xie SuperIntelligence will publish reviews and survey articles to help newbies to AGI/SI get up to speed and experienced workers stay up to speed efficiently. The latter can scroll to Section 2.3, Overview of Superalignment Methods and Challenges. Brief analysis of DeepSeek R1 and its implications for Generative AISarah Mercer, Samuel Spillard, Daniel P. Martin For quick and incisive insights into DeepSeek, read this analysis and Dario Amodei’s cool-headed response to all the hype about DeepSeek. Effective Mitigations for Systemic Risks from General-Purpose AIRisto Uuk, Annemieke Brouwer, Tim Schreier, Noemi Dreksler, Valeria Pulignano, Rishi Bommasani A timely article with practical, near-term-implementable AGI risk mitigation suggestions. Examples: • Unlearning techniques: Removing specific harmful capabilities (e.g., pathogen design) from models using unlearning techniques.• Capability restrictions: Restricting risky capabilities of deployed models, such as advanced autonomy (e.g., self-assigning new sub-goals, executing long-horizon tasks) or tool use functionalities (e.g., function calls, web browsing).• Input and output filtering Monitoring for dangerous outputs (e.g., code that appears to be malware or viral genome sequences) and inputs that violate acceptable use policies to ensure models do not engage in harmful behaviour.• Bug bounty programs Clear and user-friendly bug bounty programs that acknowledge and reward individuals for reporting model vulnerabilities and dangerous capabilities. • Safety drills Regularly practising the implementation of an emergency response plan to stress test the organisation’s ability to respond to reasonably foreseeable, fast-moving emergency scenarios. Simulating Influence Dynamics with LLM AgentsMehwish Nasim , Syed Muslim Gilani, Amin Qasmi, and Usman Naseem Analyzing how AGI/SI may influence human opinion is a critical aspect of risk and safety analysis, as is simulation of AGI risk behavior. The methodology the authors present in this short paper has broad application: This paper introduces a simulator to model influence and counter-influence in a wargame setting. Wargames, originally developed for military strategy, have evolved into powerful tools for decision-making across various domains. Today, they are used to model business strategies, assess cybersecurity threats, and simulate geopolitical conflicts. Governments and corporations employ wargames to anticipate economic shifts, supply chain disruptions, and the impact of emerging technologies. In healthcare, they help model pandemic responses, testing different policy interventions before realworld implementation. AI-driven wargames further enhance scenario analysis, enabling rapid adaptation to complex environments. By fostering strategic thinking and resilience, modern wargaming serves as a critical tool for navigating uncertainty in an increasingly interconnected world. Can a Bayesian Oracle Prevent Harm from an Agent? Yoshua Bengio, Matt McDermott, Michael K. Cohen, Nikolay Malkin, Damiano Fornasiere, Pietro Greiner, Younesse Kaddar SI co-founding Editor Steve Omohundro comments: Turning an oracle into an agent may take just a page of code. OK, but that doesn’t mean the methods outlined by Be

Community

0 comments
Use Connect Wallet in the navigation

No discussion yet

Be the first to share a question or observation.