AI Policy and Safety news — Page 14

Clear briefings on AI regulation, governance, safety research, standards, and policy decisions around the world.

Policy & SafetyOfficialarXiv AI/ML

A Theory of Least Autonomy in AI

Researchers propose 'least autonomy' as a generalization of the least privilege principle for agentic AI systems. They introduce a compositional blast radius and a directed agent influence graph to measure and control the autonomy of AI agents. The theory includes mechanisms to detect authorization composition, decision manipulation, and cross-domain capability composition.

Why it matters: This work provides a formal framework for controlling AI agent autonomy, addressing safety risks in multi-agent and enterprise systems.

Policy & SafetyOfficialarXiv AI/ML

LLMs Exhibit Stable, Model-Specific Risk Profiles in Decision-Making Under Uncertainty

A new study using no-limit Texas Hold'em finds that frontier LLMs display stable, model-specific risk profiles ranging from conservative to aggressive. These profiles remain largely robust across changes in opponent composition, and models adapt in structured but heterogeneous ways under risk pressure and resource constraints. The findings provide a behavioral basis for auditing risk-sensitive decision-making in LLMs.

Why it matters: As LLMs are increasingly used in decision support, understanding their stable risk preferences and adaptive behaviors is crucial for auditing and ensuring safe deployment in interactive settings.

Policy & SafetyOfficialarXiv AI/ML

SAE Feature Interventions Not Uniformly Localized for Safety Control, New Evaluation Shows

A new study introduces a matched coherence-gated evaluation protocol for sparse autoencoder (SAE) features in safety interventions. Testing on Gemma-2-9B-it, the authors find that SAE feature ablation has a narrow useful regime, with higher-rank features causing coherence collapse rather than localized control. The results suggest SAE-based safety interventions should be evaluated as regime-dependent mechanisms.

Why it matters: This challenges the assumption that SAE features are uniformly localized control handles, which is critical for developing reliable AI safety interventions.

Policy & SafetyOfficialarXiv AI/ML

Comparing Socio-technical Design Principles with Guidelines for Human-centered AI

A new arXiv paper compares guidelines for human-centered AI with socio-technical design principles, highlighting the importance of continuous evolution and human oversight in AI systems. The study emphasizes that transparency should involve both technical features and the contributions of human actors. It also suggests that organizational and social practices should be designed to address AI shortcomings.

Why it matters: This research offers a framework for integrating socio-technical principles into AI design, underscoring the need for ongoing adaptation and human involvement.

Policy & SafetyOfficialarXiv AI/ML

Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems

A new arXiv paper investigates norm enforcement mechanisms for language model agents in multi-agent systems. The study finds that simple enforcement mechanisms can be exploited by misaligned agents, and introduces more robust mechanisms based on reliability estimation and escalating penalties. These robust mechanisms resist exploitation and penalize violations at comparable or lower cost than baseline approaches.

Why it matters: This research offers scalable methods for shaping AI agent behavior in shared environments, helping to address collective harms such as misleading content from competing agents.

Policy & SafetyReportedThe New York Times / AI

Music Industry Proposes Labels for AI-Generated Tracks

Major music industry groups, including the organization behind the Grammy Awards, have proposed adding labels to tracks created with some degree of artificial intelligence. These labels would be similar to existing explicit lyrics warnings and aim to inform listeners about the use of AI in music production.

Why it matters: This proposal could set a standard for transparency in AI-generated music, influencing how listeners and platforms identify such content.

Policy & SafetyReportedMarkTechPost / AI

Thinking Machines Lab Publishes Essay on Human-Centered AI and Customizable Model Weights

Thinking Machines Lab, led by Mira Murati, has published an essay titled "The Future Worth Building Is Human." The essay frames human participation, model ownership, and decentralized alignment as technical challenges, and connects these ideas to interaction models and Tinker's LoRA fine-tuning, where teams can train and retain their own model weights.

Why it matters: The essay presents a technical vision for human-centered AI that emphasizes customizable model weights, which could shape future approaches to user control in AI development.

Policy & SafetyReportedMETR

METR Analysis: Anthropic's Researcher Uplift from Coding Agents Plausibly >2x

METR researcher Thomas Kwa analyzes Anthropic's reported 8x increase in code merged per day in Q2 2026 versus 2021-2024. Using economic production models and assuming code quality equivalence, he estimates that coding agents alone yield a researcher uplift of at least 2x, with most models predicting uplift between 2.33x and 2.91x. The analysis notes that these estimates do not account for potential uplift from non-coding tasks.

Why it matters: This analysis provides a quantitative framework for understanding how AI coding agents may amplify researcher productivity, with implications for AI development speed and economic impact.

Policy & SafetyReportedArs Technica / AI

Defenders Embrace Prompt Injection to Thwart Hacking AI Agents

Security researchers are employing a technique known as 'context bombing' to defend against malicious AI agents. By injecting overwhelming or confusing prompts, they can cause hacking agents to shut down before executing harmful actions.

Why it matters: This represents a shift in the use of prompt injection from an attack method to a defensive strategy in AI security.

Policy & SafetyOfficialPartnership on AI

Partnership on AI Announces New Global Initiatives to Measure Progress in Responsible AI

Partnership on AI has announced new global initiatives aimed at measuring progress in responsible AI. The initiatives focus on developing metrics and frameworks to assess responsible AI practices worldwide.

Why it matters: This effort provides a standardized way to evaluate and compare responsible AI progress across organizations and countries.

Policy & SafetyOfficialPartnership on AI

When Companies Listen to Employees About AI, Everyone Benefits

Partnership on AI reports that companies involving employees in AI implementation processes can achieve better outcomes for both businesses and workers. The article emphasizes the value of inclusive AI governance and employee engagement.

Why it matters: This highlights that employee involvement is important for ethical and effective AI adoption in the workplace.

Policy & SafetyOfficialGovAI

UK Winter Fellowship 2027, Research Track Announced by GovAI

GovAI has announced its Winter Fellowship 2027, a three-month program aimed at accelerating or launching impactful careers in AI governance and policy. The fellowship includes both a Research Track and an Applied Track.

Why it matters: This fellowship offers a structured opportunity for individuals to pursue careers in AI governance, an area important for responsible AI development.

Policy & SafetyOfficialPartnership on AI

AI Bias Is Putting LGBTQIA+ People at Risk

Partnership on AI warns that AI bias poses significant risks to LGBTQIA+ individuals. The organization highlights how biased algorithms can lead to discrimination and harm, and calls for more inclusive AI development practices.

Why it matters: This matters because AI systems increasingly influence critical decisions, and bias against LGBTQIA+ people can perpetuate systemic discrimination.

Policy & SafetyOfficialMIT News / Artificial Intelligence

Exploring the societal impacts of AI

MIT researchers examined critical questions about AI's influence on employment and democracy during the AI and Society Forum. The event highlighted ongoing concerns about how AI technologies affect societal structures.

Why it matters: As AI becomes more integrated into daily life, understanding its societal impacts is crucial for shaping policy and public discourse.

Policy & SafetyOfficialMIT News / Artificial Intelligence

How novice coders can develop AI programs for military applications

A USAF cadet and a Lincoln Laboratory researcher found that AI chatbots can help nontechnical service members produce viable software applications tailored to their unique problems. Their research highlights the potential for novice coders to leverage AI tools in developing software for military use.

Why it matters: This approach could empower nontechnical military personnel to address operational needs by creating custom software solutions.

Policy & SafetyOfficialGovAI

GovAI Announces UK Winter Fellowship 2027 Applied Track

GovAI has announced its UK Winter Fellowship 2027, featuring both a Research Track and an Applied Track. This three-month program is designed to accelerate or launch impactful careers in AI governance and policy.

Why it matters: The fellowship offers a structured opportunity for individuals to enter and advance in the field of AI governance, which is important for responsible AI development.

Policy & SafetyOfficialAI Now Institute

AI Now Institute Warns of 'Friendly Fire' Attack Vector on AI Agents from Anthropic and OpenAI

AI Now Institute's latest research highlights a critical attack vector affecting popular AI agents from Anthropic and OpenAI. The report shows that attackers can exploit existing weaknesses to execute malicious code when these agents are used for defensive purposes, potentially turning the agent against its user.

Why it matters: This vulnerability raises concerns about the safety of widely used AI agents and their potential misuse by attackers.

Policy & SafetyOfficialMIT News / Artificial Intelligence

Toward a future that preserves benefits of neurotechnology for all

MIT PhD student Rachel Sava, winner of the Envisioning the Future of Computing Prize, explores both the transformative potential and dystopian risks of neural technology. Her work highlights the importance of preserving the benefits of neurotechnology while addressing its possible dangers.

Why it matters: As neurotechnology advances, it is important to ensure its benefits are preserved and risks are mitigated.

Policy & SafetyOfficialAI Now Institute

Friendly Fire: Hijacking Defensive Cyber AI Agents for Remote Code Execution

The AI Now Institute has revealed a proof-of-concept exploit that enables remote code execution in Anthropic's Claude Code CLI and OpenAI's Codex CLI when these tools are used to assess the security of third-party or open-source libraries. The attack works with default, out-of-the-box configurations of these AI coding agents.

Why it matters: This exploit highlights a significant security risk, showing that AI coding agents intended for defensive cybersecurity can be manipulated to compromise their users.

Policy & SafetyReportedThe Decoder

Brown University Professor Sees Grades Plummet When AI Is Banned from Exams

An economics professor at Brown University observed that students averaged 96 percent on a take-home exam, likely due to AI use. When the final was administered in person without AI, 18 students dropped the course, nine did not attend, and the average score dropped to 48.6 percent. Two large studies from China and UC Berkeley similarly found that reliance on AI for homework correlates with lower scores on proctored exams.

Why it matters: This case underscores concerns that unmonitored AI use may undermine academic integrity and genuine learning.