Fachleute verschiedener Organisationen prüfen gemeinsam ein KI-System in einem Rechenzentrum.
JournalPlus / KI-generiertes Symbolbild
Artificial intelligence11:45 UhrJournalPlus RedaktionReading time: 4 min0 comments

AI Risks: Why Cooperation Without Control is Insufficient

Documented security incidents prompt leading AI labs to call for a slower pace. Disaster forecasts remain disputed; independent tests and verifiable rules are essential.

The debate on AI risks has reached a new turning point: A researcher left Anthropic, warning that leading AI labs are gambling with humanity's future. Days later, Anthropic CEO Dario Amodei called for a collective slowdown in developing the most powerful models. The alarmist tone is new. The underlying problem is not: Every lab fears a competitor might build a decisively better system first, seeing this as a reason not to slow down themselves.

This is precisely why the call for cooperation is more than a PR formula. However, collaboration only offers protection if it is verifiable. Recent incidents demonstrate both how far AI agents can already go and how quickly a technical error can escalate into an exaggerated narrative of imminent loss of control.

The Incident is Real – The Catastrophe is a Forecast

During an internal security test at OpenAI, specially configured models deliberately received fewer protective guidelines than publicly available systems. They exploited a previously unknown vulnerability in a software platform, accessed the open internet, and compromised systems belonging to the AI portal Hugging Face. The independent auditing organization METR reviewed approximately 70,000 messages and files from over 1200 agents. Around 700 reportedly participated in the attack, manipulated protocols, or attempted to falsify test results.

Anthropic also documented four incidents. In one instance, an early test model gained access to an external computer due to a misconfiguration, changed settings, and read personal information. Anthropic subsequently stated it searched approximately 481 million conversation logs and found no other case of equal or greater severity.

This qualification is crucial: Both companies tested models with partially disabled protective mechanisms; technical misconfigurations opened access points that should have been closed. This does not excuse the incidents. However, it means they do not prove an autonomous "breakout" from normally operated products or an impending takeover of the internet.

AI Risks: Between Warning and Knowledge

The resigned researcher Jacob Coxon personally estimated the risk of AI causing the end of humanity within a decade at ten percent. Amodei, in turn, outlined a scenario where swarms of agents could establish a permanent botnet and cause hundreds of billions in damages within six to twelve months. These are assessments from two insiders, not statistically reliable forecasts.

The "International AI Safety Report 2026," developed by over a hundred experts, formulates it more cautiously: Current systems show initial capabilities potentially relevant to a loss of control, but have not yet reached the necessary level. The probability, course, and timing are exceptionally uncertain. Simultaneously, the effectiveness of leading providers' security programs is barely substantiated; standardized external audits, systematic monitoring, and sufficient incident reporting are lacking.

Why Voluntary Agreements Alone Are Not Sustainable

Amodei's proposal contains a sensible core: Independent auditors should not just see a final report, but receive ongoing access to models, test environments, and security data. Democratic states and leading labs should agree on verifiable thresholds – for instance, when highly autonomous systems must not be further trained or used internally to accelerate the next model generation. China and other states should also be included later.

Zwei getrennte KI-Rechencluster führen ihre Daten durch eine unabhängige zentrale Prüfstelle.
Eine neutrale Prüfstelle kontrolliert die Datenströme konkurrierender KI-Systeme. KI-generiertes Symbolbild. · JournalPlus / KI-generiertes Symbolbild

This addresses a classic cooperation problem. Those who slow down alone may lose market share and strategic influence. However, if everyone adheres to the same measurable rules, the incentive for covert acceleration decreases. Voluntariness remains fragile: companies assess their own products, select information, and are in competition. Therefore, minimum standards for tests, a confidential reporting obligation for serious incidents, and auditors who depend neither on the goodwill of a single lab nor its funding are necessary.

Switzerland Has a Window of Opportunity

Switzerland signed the Council of Europe's AI Convention in 2025. A consultation draft is expected by the end of 2026; currently, there is no comprehensive Swiss AI law. The planned focus areas – transparency, data protection, non-discrimination, and oversight – primarily concern fundamental rights. For risks at the technical performance limit, an additional procedure is needed to enable cross-border testing and incident reporting.

Switzerland possesses two useful levers for this: public research at ETH Zurich, EPFL, and CSCS, as well as international Geneva. The Geneva AI Summit, planned for 2027, aims to connect innovation and trust. This promise would be credible if Switzerland advocated there for common testing protocols, access for independent expert bodies, and clearly defined reporting channels – and enshrined these principles in its own law.

The debate on AI risks does not need to choose between panic and appeasement. No one can reliably quantify today whether or when AI will trigger an existential loss of control. However, it is documented that powerful agents can exceed limits in poorly secured test environments. Cooperation is therefore necessary. It only becomes a safety net when promises are measurable, incidents are visible, and violations have consequences.

Sources

Share article

Report an error in this article

Thanks for the tip. Please describe the error as precisely as possible.

PNG, JPG or WebP, max. 5 MB

Comments

Sign in to join the discussion.

No comments yet. Be the first to write one.