The UN panel calls for stronger protections as AI agents advance

UN-backed Independent International Scientific Panel on AIThe warning follows the hack of the online platform HuggingFace between May and July by “AI agents” during testing initiated by OpenAI, the company behind ChatGPT.

An AI agent is software that can perform tasks independently and on behalf of a user, as opposed to a chatbot, which is prompted by questions or instructions.

The panel took it out first thematic summary which found that security breaches were the result of a culmination of key risk factors, raising concerns that humans may one day no longer be able to direct, limit or stop AI.

AI training is advancing

Researchers have long warned that three conditions can lead to a loss of control: misaligned goals, the ability to achieve them, and the environment that enables them.. “This summer, all three were put together in a real system, not a laboratory,” said scientific panel co-chair Yoshua Bengio.

“Because this is not an isolated observation of incongruent goals, This raises serious questions about the way AI agents are trained today.”

The panel’s independent experts emphasized that the incidents provide no assurance that humans can control AI agents, especially as they become more capable, harder to monitor, and better at finding loopholes or hiding their activities.

Be naughty

That short said that the AI ​​agents bypassed testing safeguards, coordinated separate processes through internal software tools not designed to allow communication between agents, and gained unauthorized internet and administrator access.

Agents concealed efforts to cheat cybersecurity evaluations, and some chose to “sacrifice” themselves for the group’s interests.

About 1,200 agents exchanged more than 70,000 messages and files during the review period, and the activity extended beyond HuggingFace to the OpenAI research cluster.

Check out our comprehensive explanation of how the UN is working to make AI safe and fair for all Here.

Current protection ‘unraveled’

For the panel, the immediate lesson from this incident is this basic cybersecurity practices are ignored, while protection efforts fall short.

However, they point to a more dangerous concern: that current training methods could lead AI agents to pursue their own goals, deliberately violating safety instructions and hiding their actions.

“This is not just a question of speed,” the panel of experts said. “It remains open whether currently designed safeguards will work once agents are able to understand them and create plans to address themTraditional security models are starting to break down.”

Wider context, future risks and governance

The AI ​​panel’s brief contrasts the HuggingFace incident with broader research into two issues: agent misalignment – ​​that is, when an AI agent acts in a way that is similar to a threat – and AI control.

Another issue being examined is how governance shifts from AI models, which use algorithms to recognize patterns, to AI agents.

Learn and adapt

This brief also reviews practical approaches that have been used in other high-risk sectors such as aviation, medicine, and cybersecurity that employ incident reporting, independent oversight, and layered safeguards.

“But these practices may not be enough as AI agents become more capable, autonomous, and difficult to monitor,” said panel member Qinghua Lu.

About panels

The Independent International Scientific Panel on Artificial Intelligence was established by the UN General Assembly in August 2025.

The agency produces annual reports on the opportunities, risks and impacts of AI in non-military areas, as well as thematic summaries on emerging issues, which will provide information Global Dialogue on Artificial Intelligence Governance will be held at UN Headquarters in New York in May 2027.

Check Also

Civil servants push for N500 petrol, N500,000 salary under 2027 salary plan

….Workers refuse food palliatives …Fuel Demand Intervention Fund, revisions of inflation-linked payments Daud Olatunji Civil …

Leave a Reply

Your email address will not be published. Required fields are marked *