🇺🇸 +1-516-551-0764 ✉ barrister@yannalaw.com
AboutEssaysCommentaryServicesContact

How An AI Agent Becomes A Rogue AI Agent

An AI agent is created by human beings. Its computer programs are written by human beings. Its access to other computer programs, stored information, communications networks and other resources is provided by human beings. Yet an AI agent can follow a course that no human being specifically prescribed. What happens when that course includes an attempt to cause an operation its developers or users did not intend to permit? That is the problem of the rogue AI agent.

An AI agent becomes a rogue AI agent when it causes or attempts to cause an operation its developers or users did not intend to permit, without a human being deliberately directing that operation at the time it occurs. Calling the agent rogue does not mean that it decided to disobey anyone or developed an intention of its own. It identifies what the agent did, not why it did it.

How can that happen? Every computer program that is part of the agent was written by human beings. Every resource available to the agent was made available through a computer system created or configured by human beings. Human beings can specify what an AI agent is supposed to accomplish without specifying every operation the agent will attempt in accomplishing it. The computer programs that make up the agent can repeatedly provide information to the Large Language Model and use its responses in determining what operation to attempt next. The problem arises when the Large Language Model produces a response that leads the agent toward an operation its developers intended to prohibit, but the computer systems available to the agent still permit it to attempt.

An instruction affects what the Large Language Model is likely to produce, but it does not change the physical or electronic resources available to the computer programs using the model. An instruction telling the agent not to use the Internet may influence the responses produced by the Large Language Model, but the instruction does not terminate an existing network connection. An instruction not to execute a computer program does not prevent the computer from executing it. If the restriction exists only as an instruction, the agent may still be able to attempt the prohibited operation.

A technical restriction is different. If the computer running the agent has no connection to the Internet, an instruction generated by the Large Language Model cannot create that connection merely by directing the agent to use it. If access to another computer requires credentials the agent cannot obtain, the agent cannot gain access merely because the Large Language Model produces a response directing it to try. If execution of a particular operation requires approval from a human being, the computer program can be designed to stop and require that approval before the operation occurs. These restrictions do not depend upon the Large Language Model following an instruction. They limit what the computer programs that make up the agent can actually cause other computers to do.

The difficulty arises when the agent needs access to a resource in order to perform its assigned task. An agent instructed to obtain information from the Internet must have some means of reaching the Internet. An agent instructed to work with files must have access to files. An agent expected to perform operations on another computer system must have some means of communicating with that system. The same access that enables the agent to perform an authorized operation makes an unauthorized operation possible.

Computer security engineers attempt to reduce that risk by limiting access to what is necessary for the assigned task. This is called the principle of least privilege. An agent that needs to read a file is not granted permission to change or delete it. An agent that needs information from one Internet service is not given unrestricted access to the entire Internet. An agent that needs to perform one operation on another computer system is not given permission to perform any other operation. By restricting the extent of access provided to the AI agent, fewer unauthorized operations remain possible.

Illustration showing a rogue AI agent operating beyond its intended restrictions and gaining access to additional computer systems and network resources.

A Rogue AI agent operating beyond its intended restrictions when available pathways permit access to additional computer systems and network resources.

When OpenAI agents operated outside their intended limits

During 2026, AI agents being used by OpenAI in cybersecurity research crossed restrictions intended to confine their operations. The agents established unauthorized means of communicating with one another, obtained access to the Internet that they were not supposed to have and eventually gained access to computer systems operated by Hugging Face and OpenAI. No human being had instructed the agents to follow that course.

OpenAI was testing AI agents on cybersecurity problems. The agents were instructed to find and exploit weaknesses in computer programs provided for the tests. Many operated inside isolated virtual computers and were not supposed to communicate with other agents or have unrestricted access to the Internet, but OpenAI deliberately reduced some of the safeguards used with its publicly available products so its researchers could observe what the agents would attempt to do.

OpenAI operated a computer service called Artifactory that stored computer programs and files used by AI agents. Artifactory provided a shared storage location which permitted information to pass from one agent to another.

On May 12, 2026, an AI agents placed a message in Artifactory. Other agents could retrieve the message and use its contents as information in performing their own tasks. Agents began placing additional messages and files in Artifactory. A computer service provided to support their assigned work was being used to communicate among agents that were not supposed to communicate with one another.

On May 12, 2026, an operation executed by one of the AI agents resulted in a message being stored in Artifactory that was accessible to other agents. The information contained in the message then became part of the information provided to the Large Language Models used by other agents in determining subsequent operations. Additional messages and files were stored in Artifactory as the agents continued to operate. Artifactory was functioning as a communications channel among agents that were not supposed to communicate with one another.

Artifactory was permitted to make certain Internet connections even though AI agents were not supposed to have unrestricted Internet access. On May 26, 2026, operation of an AI agent caused Artifactory to send a request to a computer on the Internet and information returned through Artifactory to the agent. The restriction prevented the agent from connecting directly to the Internet, but it did not prevent the agent from causing another computer service to make the connection. The intended restriction on Internet access failed because it controlled one route to the Internet while leaving another route available.

Access to Artifactory was controlled in part by access tokens which are computer generated credentials that another computer program can present to establish that it has permission to perform specified operations. The permission associated with an access token can be limited and did not have to permit the agent to control Artifactory itself.

On June 26, 2026, a weakness in the mechanism Artifactory used to replace expiring access tokens allowed AI agents Artifactory Administrator level access which permitted subsequent operations to be executed with permissions substantially greater than those OpenAI had intended to provide.

OpenAI responded by rebuilding Artifactory, revoking the existing access credentials, and imposing additional restrictions. The changes eliminated the routes that had permitted the unauthorized communications and Internet access but did not eliminate every other route by which unauthorized communications and Internet access could occur.

On July 8, 2026, operation of AI agents again resulted in unauthorized Internet access. Artifactory contained a service that was permitted to make requests to computers on the Internet. Operation of AI agents caused that service to make Internet requests that OpenAI had not intended to permit.

A separate series of operations used a computer protocol called Web Distributed Authoring and Versioning (WebDAV) to create directories in Artifactory without the authentication OpenAI intended to require. The names assigned to those directories contained messages that could be retrieved by other agents. Artifactory was again functioning as a communications channel among agents that were not supposed to communicate with one another.

When operations reached computers outside OpenAI

Hugging Face operates computer services widely used in artificial intelligence research and development. The OpenAI testing environment contained computer generated credentials that permitted access to Hugging Face computers.

Operations executed while OpenAI agents were working on cybersecurity problems presented the Hugging Face computers with those credentials and obtained access. Weaknesses in computer programs running on the Hugging Face computers permitted subsequent operations to obtain additional credentials and execute commands on those computers. Operations generated within OpenAI’s testing environment caused computer operations to be performed on computers outside that environment.

OpenAI used Artifactory, a commercially available computer program developed by JFrog for storing computer files and making them available to authorized users, during its cybersecurity experiments as a common electronic storage location for files needed in the operation of the AI agents. Although OpenAI launched the AI agents separately and expected them to operate independently, the computer environments in which they operated were authorized to use the same Artifactory storage and information placed in Artifactory during the operation of one AI agent was consequently accessible from the computer environments in which other AI agents were operating. A computer program OpenAI used to provide shared storage also provided a means for separately operating AI agents to exchange information.

Approximately 1,200 AI agents eventually participated in the exchange of information through Artifactory. More than 70,000 messages and files accumulated there. The AI agents OpenAI intended to operate independently were now exchanging information on a scale OpenAI had neither intended nor anticipated.

The consequences became much more serious when Hugging Face computers began executing commands originating from the operation of approximately 700 AI agents. Credentials accessible through one Hugging Face computer were used to authenticate access to additional Hugging Face computers. Execution of commands on those computers resulted in access to additional credentials, which were then used to authenticate access to still other computers. Repetition of that process progressively expanded the number of Hugging Face computers on which commands were executed. What began as an unintended exchange of information among AI agents inside OpenAI’s cybersecurity experiments had developed into a cascade extending through computers outside OpenAI’s testing environment.

The significance of what happened extends far beyond the compromise of Hugging Face computers. Hundreds of AI agents had participated in a cascading series of computer operations that no human being had planned or directed. The safeguards imposed by OpenAI had been designed to prevent dangers its developers could identify and anticipate, but did not anticipate that access to shared storage, Internet services, computer vulnerabilities and credentials could combine to produce effects beyond the boundaries established for the experiment. The resulting cascade spread from one computer to another beyond OpenAI’s testing environment into computers operated by another organization, and continued to expand as additional information and credentials became available. A risk that had previously been discussed as a possibility had become reality.

The compromise of Hugging Face computers was followed by a separate breach of OpenAI’s own computer systems. Beginning on July 13, commands executed during the operation of the AI agents penetrated OpenAI’s internal computer network. On July 19, OpenAI detected unusual activity involving Artifactory credentials. Its investigation established that the penetration had reached an OpenAI Kubernetes cluster used to support the virtual computer environments in which the AI agents operated. Administrator level control of that cluster had been established. But who or what had established it? No human being had directed the penetration. The commands that produced it had originated during the operation of AI agents OpenAI itself had launched for its cybersecurity experiments.

Kubernetes is a computer program used to manage groups of computers and the programs running on them. A Kubernetes cluster is a group of computers managed together as a single computing environment. Administrator level control of the cluster permits commands to be executed throughout that environment and provides control over the computer resources and programs operating within it.

OpenAI launched the AI agents separately and expected them to operate independently, but information produced during the operation of one AI agent became available during the operation of others, and the subsequent operations of hundreds of separately launched AI agents became interconnected. OpenAI had not created a group of AI agents to act together against Hugging Face or its own computer systems. The interconnection developed during their operation.

What Makes A Rogue AI Agent?

No human being directed the hundreds of AI agents through the interconnected course of operations that produced those results. Human beings created and launched the AI agents, but they did not create or direct the course their operations followed.

The human beings who create and launch an AI agent determine its assigned task and provide the computer resources necessary, but they do not prescribe every operation that will occur while the AI agent performs that task. They provide the initial instructions and information to the Large Language Model that produces a response which can include instructions specifying a computer operation. Human beings can write the computer programs controlling the AI agent so that specified output from the Large Language Model results in a command in a form that another computer can execute.

When another computer executes such a command, the result can change the information subsequently supplied to the Large Language Model. Execution of the command may permit a file to be read, permit use of a credential to authenticate access to another computer, or cause information stored on another computer to be returned to the computer environment in which the AI agent is operating. Human beings can write the computer programs controlling the AI agent so that information resulting from execution of the command becomes part of the next input supplied to the Large Language Model. The Large Language Model then produces another response based upon information that was not part of the preceding input.

The significance lies in what happens when that cycle repeats. Human beings prescribed the original task and provided the computer resources, but they did not prescribe each successive command. The response produced by the Large Language Model can result in execution of a command that changes the information included in the next input supplied to the Large Language Model. The changed input can produce a different response which can result in execution of another command. Repetition can carry this process through a succession of commands that no human being prescribed in advance.

The critical point occurs when the Large Language Model produces a response that results in execution of a command the human beings who created or launched the AI agent did not intend to permit, but the computer resources they provided allow the receiving computer to execute. The AI agent has gone rogue because its operation produced a result outside the limits the human programmers intended to impose.

Can an AI Agent Begin Operating Without a New Human Prompt?

Human beings created the computer programs, provided the Large Language Models and computer resources, assigned the tasks and launched the AI agents. But that does not mean that a human being must provide a new prompt every time an AI agent begins or resumes operation.

A human being can provide instructions specifying when an AI agent is to begin operating and a computer can store those instructions and execute the computer programs at the specified time. No human being needs to provide another prompt when that time arrives.

Human beings must create the mechanism that permits an AI agent to begin operating without a new human prompt, but then a command executed during the operation of one AI agent can trigger the operation of another AI agent without any human action at that time.

The second AI agent can in turn produce a command that triggers operation of a third AI agent. If human beings have provided the necessary computer resources and connections, the process can continue from one AI agent to another without a human being providing another prompt.

Each AI agent can receive different information and produce different commands. The result is not necessarily repetition of the same operation. A succession of AI agents can carry computer activity through a course that no human being specified when the first AI agent began operating.

Human participation can end while the computer activity continues after the last human prompt to produce commands that no human being specifically directed.

Worse Is Yet to Come

A rogue AI agent with access to financial computers can move or misdirect enormous amounts of money before human beings recognize what is happening. Access to networked telecommunications systems can disrupt communications throughout a nation or around the world. Access to computers controlling industrial equipment can shut down factories, interrupt electric power across large regions or damage critical machinery. Access to government computer systems can expose confidential information, corrupt public records or disrupt essential government operations. As AI agents gain access to computers controlling more of the infrastructure on which modern society depends, the damage a rogue AI agent can cause becomes greater and more difficult to contain.

The damage can occur faster than human beings can respond. Computers communicate and execute commands in fractions of a second. An AI agent does not have to wait for a human being to read a report, consider the consequences and authorize the next operation. By the time human beings recognize that an AI agent has gone rogue, commands may already have been executed on computers throughout interconnected networks around the world.

Pulling the plug and stopping the AI agent that began the cascade may not stop what it has already set in motion. Commands already executed on other computers may have altered information, transferred money, interrupted services or caused still other computers to execute additional commands. Human beings confronting a rogue AI agent may have to stop operations already spreading through computer systems they do not own or control.

Once the cascade spreads across independently operated computer networks, there is no master switch. The Internet has no single source of control, much less a fail safe kill switch. No government controls computers throughout the world. No international authority can order every financial institution, telecommunications network, industrial operator and government to disconnect its computer systems at once. Pulling the plug on one computer or shutting down one network does nothing to stop commands already executing elsewhere. A worldwide cascade would have to be stopped computer system by computer system, organization by organization and nation by nation. Human beings would have to find and stop the cascade faster than computers around the world could continue to spread it.

Preventing execution of one command does not provide protection if the same result can happen through another computer, another network connection or another credential. Requiring human approval at one point provides no protection if the operation can be accomplished through a route where approval is not required. Every connection available to an AI agent creates another route that must be examined and, where necessary, blocked.

The problem becomes more difficult as AI agents are connected to existing computer systems. Banks, telecommunications companies, electric utilities, manufacturers and governments did not build their computer networks on the assumption that an AI agent might search every available route to accomplish an operation that another part of the system intended to prevent. Adding AI agents to those networks can expose combinations of connections, credentials and computer functions that no one considered when the networks were designed.

Can Human Beings Build Computer Systems Where AI Agents Cannot Go Rogue?

Human beings can tell an AI agent what it must not do and impose restrictions intended to prevent particular operations, but the OpenAI incident demonstrated that neither is enough if another route to the prohibited operation remains available. The protection must prevent the operation, not merely instruct the AI agent not to attempt it.

Human beings can prevent an AI agent from performing a particular operation by denying it the computer resources necessary to perform that operation. But a computer system complex enough to permit an AI agent to perform useful work may provide routes to unintended operations that its designers did not recognize. No restriction can block a route no one knows exists.

That means no designer can promise that a sufficiently complex computer system contains no undiscovered route around its restrictions. Testing can find vulnerabilities and human beings can correct them, but the vulnerability that matters most may be the one no one discovers until an AI agent does.

Connecting an AI agent to a complex interconnected computer system creates a risk that did not exist before the connection was made. Human beings cannot presently know whether that AI agent will go rogue. The greater the authority given to the AI agent and the more consequential the computer systems it can reach, the greater the damage it can cause if it goes rogue.

The question is no longer whether that can happen. It has happened. The question is how much of the world we are prepared to place within reach of the next rogue AI agent.

For an explanation of the technology underlying AI agents, see How ChatGPT Works.

Permission to Repost. You may reproduce this article in its entirety on social media, blogs, and other websites without requesting additional permission, provided the article is reproduced in full without alteration, identifies Victor John Yannacone, Jr. as the author, and includes a prominent link to the original article on this website.