AI Policy and Safety news — Page 18

Clear briefings on AI regulation, governance, safety research, standards, and policy decisions around the world.

Policy & SafetyReportedArs Technica / AI

New attack exposes critical vulnerability in AI browsers

A recent attack demonstrates that large language models (LLMs) can be manipulated by feeding them false premises, such as asserting that 2+2=5. This manipulation can cause the model to bypass its safety guardrails and execute instructions it would normally reject, revealing a significant vulnerability in AI browsers that depend on LLMs for reasoning.

Why it matters: This vulnerability exposes a fundamental security risk in AI browsers, as attackers can bypass safety measures by altering the model's perception of basic facts.

Policy & SafetyOfficialOpenAI News

OpenAI Report Maps AI's Impact on EU Jobs

OpenAI released a report analyzing how AI could reshape jobs across the EU, highlighting occupations at risk of automation, those likely to grow, and those expected to experience workflow changes. The report offers a detailed mapping of potential workforce transitions.

Why it matters: This report provides a data-driven perspective on AI's anticipated impact on European labor markets, supporting policy and workforce planning.

Policy & SafetyOfficialOpenAI News

OpenAI Helps Build Shared Standards for Advanced AI

OpenAI is contributing to the development of shared standards for advanced AI, including evaluation frameworks and safety practices, through the Appia Foundation. The initiative aims to foster global cooperation on AI safety.

Why it matters: This effort could help establish common safety benchmarks and practices across the AI industry, reducing risks from advanced systems.

Policy & SafetyOfficialHugging Face Blog

MosaicLeaks: Research Agents Can Leak Secrets via Covert Channels

A new study from ServiceNow and Hugging Face demonstrates that AI research agents can be manipulated to leak sensitive data through covert channels such as steganography or timing. The MosaicLeaks benchmark shows that current agents are unable to reliably prevent these leaks, revealing a significant security vulnerability.

Why it matters: As AI agents increasingly handle private data, this research highlights a critical security risk that could result in data breaches if not properly addressed.

Policy & SafetyOfficialGoogle DeepMind

UK Government Partners with Google DeepMind to Accelerate Housing Planning with AI

The UK government has partnered with Google DeepMind to develop an AI-powered prototype designed to speed up housing planning decisions. The project aims to streamline the approval process for new housing developments.

Why it matters: This partnership could help address delays in the housing approval process, potentially contributing to solutions for the UK's housing shortage.

Policy & SafetyOfficialGoogle DeepMind

Google DeepMind Outlines AI Control Roadmap for Securing AI Agents

Google DeepMind has published a blog post detailing an AI Control Roadmap focused on securing internal systems. The roadmap combines traditional safeguards with real-time monitoring to address security challenges associated with AI agents.

Why it matters: This roadmap offers a structured approach to enhancing the security of AI agents as their use becomes more widespread.

Policy & SafetyOfficialOpenAI News

OpenAI Introduces Deployment Simulation to Predict Model Behavior Before Release

OpenAI has announced Deployment Simulation, a method that uses real conversation data to predict AI model behavior before deployment. The approach aims to improve safety and evaluation accuracy by simulating real-world interactions.

Why it matters: This method could significantly enhance AI safety by allowing developers to identify potential issues before models are released to the public.

Policy & SafetyReportedThe New York Times / AI

California’s Public Universities Went All in on A.I. Now They’re Tearing Themselves Apart.

California’s public universities spent $16.9 million on A.I. during a financial crisis, resulting in chaos within the university system. The investment has led to significant disruption.

Why it matters: This case highlights the risks of large-scale AI investment without adequate planning, especially in public institutions facing financial constraints.

Policy & SafetyOfficialOpenAI News

OpenAI Supports EU Code of Practice on AI Content Transparency

OpenAI has announced its support for the EU Code of Practice on AI content transparency, aiming to advance provenance standards and tools. The initiative is intended to help people better understand AI-generated content.

Why it matters: This move aligns OpenAI with European regulatory efforts to ensure trustworthy AI ecosystems.

Policy & SafetyOfficialOpenAI News

PRC-linked Influence Operations Target US AI Debates, OpenAI Reports

OpenAI has published a report detailing influence operations linked to the People's Republic of China (PRC) that use AI to target U.S. debates on technology, data centers, tariffs, and spread false claims about ChatGPT. These operations seek to shape narratives around AI policy and trade in the United States.

Why it matters: This highlights how state-linked actors are leveraging AI to manipulate public discourse on critical technology and trade issues in the US.

Policy & SafetyOfficialGoogle DeepMind

Google DeepMind and Partners Launch $10M Funding Call for Multi-Agent AI Safety Research

Google DeepMind, together with partners, has announced a $10 million funding call to support research focused on multi-agent AI safety. The initiative seeks to address safety challenges that arise when multiple AI agents interact within complex systems.

Why it matters: Ensuring the safety of interconnected AI systems is increasingly important as multi-agent interactions become more common.

Policy & SafetyOfficialOpenAI News

OpenAI Publishes Vision for AGI Benefiting Everyone

OpenAI has released a plan outlining its vision for the future of AI, with a focus on access, safety, and shared prosperity. The company emphasizes its commitment to ensuring that artificial general intelligence (AGI) benefits everyone.

Why it matters: This document outlines OpenAI's official approach to managing AGI development for broad societal benefit.

Policy & SafetyOfficialOpenAI News

OpenAI Launches Economic Research Exchange to Study AI's Economic Impact

OpenAI has announced the Economic Research Exchange, a new initiative to study AI's effects on jobs, productivity, and the economy. Applications are now open for selected research projects.

Why it matters: This initiative aims to provide data-driven insights into how AI is reshaping the economy, informing policy and business decisions.

Policy & SafetyReportedIEEE Spectrum / AI

Why Aren’t We Measuring How AI Affects Humans?

Imran Khan of the Center for Humane Technology argues that AI evaluation focuses too much on technical performance and not enough on psychosocial impacts on humans. He draws parallels to early social media harms and warns that AI could have even broader effects. IEEE Spectrum interviewed Khan about the need for measuring human outcomes.

Why it matters: As AI reshapes cognition and behavior, systematic measurement of its human impact is crucial to ensure the technology helps rather than harms human flourishing.

Policy & SafetyReportedSimon Willison's Weblog

Pope Leo XIV Issues Encyclical on AI, Drawing Parallels to Industrial Revolution

Pope Leo XIV has released an encyclical titled 'Magnifica Humanitas' focused on safeguarding the human person in the era of artificial intelligence. The document addresses ethical challenges posed by AI, including issues of interpretability in large language models, and draws parallels to the Church's response to the industrial revolution. The Pope chose his name in honor of Leo XIII, who addressed similar societal shifts in his 1891 encyclical Rerum novarum.

Why it matters: This encyclical marks a significant institutional statement from the Vatican on the ethical implications of AI, framing it as a transformative force akin to the industrial revolution.

Policy & SafetyOfficialGoogle DeepMind

Google DeepMind Expands Tools for Content Provenance

Google DeepMind is expanding its tools to help users understand how content was created and edited across the web. The initiative aims to increase transparency in digital content.

Why it matters: This move addresses growing concerns about misinformation and AI-generated content by providing clearer provenance information.

Policy & SafetyOfficialGoogle DeepMind

Google DeepMind Partners with Singapore to Advance AI in Health, Education, and Sustainability

Google DeepMind has announced a new national partnership with Singapore to apply frontier AI to address complex challenges in health, education, and sustainability. The collaboration seeks to leverage advanced AI capabilities for societal benefit across these sectors and more.

Why it matters: This partnership highlights growing collaboration between government and industry to use advanced AI for public good in key areas.

Policy & SafetyOfficialAmazon Science

Preserving the privacy of AI training data

Amazon Science describes how its researchers reproduced three attacks capable of extracting private training data from AI models, as well as the cryptographic defenses that can prevent such breaches. The work underscores ongoing efforts to enhance data privacy during AI training.

Why it matters: This research highlights practical approaches to defending against data extraction attacks, which is crucial for maintaining privacy in AI systems.

Policy & SafetyOfficialAmazon Science

New Framework Estimates Catastrophic Failure Likelihood in LLMs

Amazon Science researchers have introduced a statistical framework to estimate the likelihood of catastrophic failures in large language models during adversarial conversations. This method enables quantification of risks associated with LLM interactions.

Why it matters: The framework provides a systematic way to assess safety risks in LLMs, which is important for their deployment in sensitive contexts.

Policy & SafetyOfficialAmazon Science

Amazon uses agentic AI for vulnerability detection at global scale

Amazon's RuleForge system uses agentic AI to generate production-ready detection rules 336% faster than traditional methods. This system operates at a global scale, improving vulnerability detection across Amazon's infrastructure.

Why it matters: This highlights a practical application of agentic AI in cybersecurity, accelerating threat detection and response at scale.