Understanding how ChatGPT works requires answering a series of questions about what a large language model actually is, what an AI agent is, where its objectives come from, how it can act in the world, how it can continue operating after the human who started it ceases to be involved, and what we should make of recent developments in artificial intelligence.
The question also exposes a persistent problem in public discussion of artificial intelligence. We use the vocabulary of human conduct to describe the behavior of machines. We should not.

A user prompt becomes a generated response through a Large Language Model.
The Probabilistic Machine
Artificial intelligence systems run on existing computer hardware. But a large language model (LLM) is not a computer program as we have come to know it.
In conventional computer programming, the operations the computer performs are specified by instructions written by programmers. Programmers write instructions that tell the computer what operations to perform.
An LLM is created differently. Programmers still write conventional computer programs to establish a system in which enormous numbers of numerical values can be changed without changing the instructions contained in the computer programs themselves. They do not write the immense collection of instructions that would be required to specify how an LLM should answer every possible question or respond to every possible combination of words presented to it.
A program might instruct a computer to multiply an input by a numerical value stored in memory. The instruction to perform multiplication does not change when the stored numerical value changes, but the result does. An LLM extends that basic distinction on an extraordinary scale. An LLM uses an enormous number of numerical values that affect one another through mathematical calculations. Changing those values changes the results produced by the calculations without requiring the underlying computer instructions that perform the calculations to be rewritten.
Those numerical values are called parameters. Many of the parameters are weights that determine how strongly particular numerical values affect other numerical values during the calculations. An LLM may contain billions or even trillions of parameters. The particular values assigned to those parameters are not individually selected and written by programmers. They are produced through a mathematical process called training.
During the initial training of an LLM, enormous quantities of existing text are presented to it.
That raises the next question. How does an LLM convert written language into something upon which mathematical calculations can be performed?
Turning Language Into Numbers
Computers already represent written characters by numbers. An LLM requires something more. It must represent the text by numbers that can be used in the mathematical calculations by which the LLM identifies and develops relationships within the text.
The first step is to divide the text into units called tokens. A token is not necessarily a word. Depending upon the text and the particular LLM, a token may represent an entire word, part of a word, punctuation, or another element of written language.
Each token is assigned an identifying number that serves only to identify the token. It does not describe the token or its relationship to anything else. The identifying number is used as an index to locate a separate collection of numerical values associated with that token. That collection of values is called an embedding. The embedding can be thought of as locating the token in a mathematical space having many dimensions.
The numbers in an embedding reflect relationships that appear in the training text. For example, the numerical representations of “cat” and “dog” can reflect relationships that differ from those associated with “cat” and “democracy.” The relationships are expressed mathematically rather than by definitions written by programmers.
The initial embedding does not by itself determine what a token means in a particular sentence. The significance of a word often depends upon the words surrounding it. The word “bank” has a different significance in “the bank approved the loan” from its significance in “the river overflowed its bank.” The LLM must be able to use the surrounding text to alter the numerical representation of a token as the text is processed.
The Neural Network
The mathematical calculations that allow the numerical representation of one token to be affected by other tokens are performed within what computer scientists call a neural network. There is nothing biological about a neural network created by a computer. A neural network is a system of mathematical operations in which numerical values pass through successive stages of calculation. At each stage, some numerical values affect other numerical values according to parameters established during training.
The individual calculating units within a neural network came to be called artificial neurons because the earliest neural networks were loosely inspired by attempts to represent mathematically how biological neurons receive signals from other neurons and produce an output. An artificial neuron, however, is a mathematical operation performed by a computer, not a biological cell.
The word “network” describes the way these mathematical operations are connected. The output of one calculation can become part of the input to other calculations. When this occurs across enormous numbers of operations using the parameters established during training, the numerical representation of information can be progressively changed as it moves through the network.
The important question for an LLM is how this network of mathematical operations can account for relationships among the tokens in a passage of text. The significance of one token may depend upon other tokens in the passage, and their importance may change with the context. The LLM needs a way to determine which tokens are relevant to one another and how strongly each should affect the numerical representation of the others.
In 2017, a group of researchers at Google proposed a new way for a neural network to make those determinations. Their paper was entitled “Attention Is All You Need.” The mathematical process at the center of their proposal became known as attention.
No programmer specifies in advance which tokens will be important to one another in any passage of text the LLM may encounter. Attention makes that determination from the particular text being processed. It calculates relationships among the numerical representations of the tokens and uses those calculations to determine how strongly the tokens should influence one another.
When the network processes “the bank approved the loan,” attention allows the representation of “bank” to be affected by its relationships with “approved” and “loan.” When it processes “the river overflowed its bank,” different relationships become important. The influence of one token upon another depends upon the context in which each appears.
The researchers built a new neural network architecture around this process of attention. They called that architecture the Transformer. The name does not mean that the LLM transforms or modifies itself. It is simply the name given to an architecture that made it possible to apply attention systematically across the tokens in the text as they were processed through the neural network.
Training the LLM
Training is the process by which the numerical values of the parameters within the neural network are established and progressively adjusted using examples of existing language. To provide those examples, an enormous quantity of text is made available to computer programs created for the training process. That text may contain trillions of occurrences of tokens.
The training text does not contain trillions of different tokens. The same tokens occur repeatedly throughout the text. In human language, a vocabulary is the set of distinct words in that language. An LLM vocabulary contains all of the distinct tokens available to represent the text the LLM processes, each of which may represent an entire word, part of a word, punctuation, or another element of written language.
A training program selects a sequence of tokens from the text and supplies some of those tokens as input to the LLM. The mathematical operations within the LLM’s neural network use that input to calculate which token in the LLM’s vocabulary is most likely to come next. The training program already knows the correct answer because the next token appears in the original training text.
To determine which token is most likely to come next, the LLM considers every token in its vocabulary as a possible choice and the neural network calculates a numerical score for each one. A higher score means that token is more
Each score is a number but the scores are useful only in relation to one another. After the neural network calculates the scores, another mathematical operation within the LLM converts them into a probability that a particular token will come next. The higher the probability, the more likely that token is to come next.
The training program uses the probability assigned to the token that actually came next to calculate the amount of error in the prediction.
The training program uses the amount of error to calculate how a small change in each parameter would affect that error and then adjusts the parameters in directions calculated to reduce the error.
The process is repeated with other sequences of tokens throughout the training text. Each repetition adjusts the numerical values of the parameters that control the calculations within the neural network. No programmer determines those values. They result from the accumulated adjustments made during training.
As the process is repeated, the changing parameter values are affected by recurring relationships among tokens throughout the training text. Those accumulated changes enable the LLM to use the relationships found during training when it later encounters new text. The LLM can then calculate which token is most likely to come next even when that particular sequence of text never appeared in the training material.
Training establishes the parameter values the LLM will use when it responds to a user prompt. Once training is complete, those parameter values can be used to generate a response.
How ChatGPT Generates an Answer
An LLM operates as one component of a larger computer system created by human programmers. The system software includes the computer programs that receive information from the user, the tokenizer that converts written language into tokens, the inference software that causes the LLM to perform its calculations, programs that manage computer memory and other systems in which information can be stored, and programs that control the movement of information among these components.
Those programs are written by human programmers and execute on computer hardware. In a system such as ChatGPT, some operate on the user’s computer and others operate on computers used by OpenAI to provide the service. Together they receive the keystrokes from the user, transmit them between computers, direct their storage in computer memory, cause the tokenizer and the LLM to be invoked and direct the movement of the resulting data through the system.
Before the LLM can process a user prompt, the ordinary written language of the prompt must be translated into the tokens the LLM uses for its calculations. That translation is performed by computer software written by human programmers called a tokenizer. It is separate from the neural network and does not determine what the prompt means. It converts the written language of the prompt into the sequence of tokens that will be supplied to the LLM.
After the tokenizer supplies the sequence of tokens to the LLM, each token is associated with the numerical representation that allows it to be processed by the neural network. The neural network then uses the parameter values established during training and the relationships among the tokens in the prompt to calculate which tokens in its vocabulary are more likely than others to come next.
Parallel Processing
The LLM performs many of these calculations in parallel. After the tokenizer converts the prompt into a sequence of tokens, the LLM can perform mathematical calculations involving many of those tokens at the same time rather than completing the processing of each token before beginning to process another.
The computer hardware which hosts an LLM contains large numbers of calculating units that can operate simultaneously. Instead of requiring one calculation to be completed before another calculation can begin, the computer performs many calculations at the same time. That is parallel processing.
Computer hardware designed to perform many calculations simultaneously is called a parallel processor and it is optimized to perform parallel processing.
There are two forms of parallel processing important to understanding how an LLM processes a prompt.
The first level of parallel processing involves the tokens in the prompt. The LLM can perform calculations involving many tokens at the same time while preserving the position of each token within the sequence. It can then calculate relationships among tokens located throughout the prompt without processing them one after another.
The second level divides the mathematical calculations involving enormous collections of numbers performed by the neural network into many smaller calculations that different calculating units within the computer hardware perform at the same time and then combine into a single composite result. When the calculations required by an LLM are too large for a single processor, this work can be distributed among multiple parallel processors working together.
The process by which a trained LLM uses the parameter values established during training to calculate an output is called inference. The computer programs that direct this process are commonly called inference software. They supply the tokens to the LLM, receive the numerical results of its calculations and use those results to control the generation of the response.
The token with the highest calculated probability is not necessarily the token that becomes part of the response. The inference software can be configured to select the token with the highest probability or to permit selection among other tokens to which the LLM has assigned substantial probabilities. This is one reason the same prompt can produce different responses on different occasions.
Once a token has been selected, it is added to the sequence of tokens already being processed. The expanded sequence is processed again and the LLM calculates new probabilities for the token that will come next. The inference software selects another token and adds it to the sequence. This process is repeated, one token at a time, until the response is complete.
Parallel processing explains how the LLM can perform an enormous number of calculations rapidly, but it does not explain how the LLM can produce a lengthy coherent answer or participate in an extended conversation. For that, another concept is necessary.
Context and the Continuing Conversation
That concept is context, the information available to the LLM when it calculates what token should come next.
When the LLM calculates the next token, it can take into account a large number of tokens that came before it which can include the user’s original question, instructions governing the response, tokens already generated as part of the answer and, in an extended conversation, portions of the preceding exchanges between the user and the LLM.
When the LLM finishes generating a response, that particular invocation of the LLM ends. The LLM does not itself retain the conversation and wait for the user’s next prompt. The computer system surrounding the LLM maintains the continuing record of the exchange.
When the user submits another prompt, the system constructs a new input for the LLM. That input includes the new prompt together with preceding material from the continuing exchange. The tokenizer converts that input into tokens and the inference software invokes the LLM again. What appears to the user as one continuing conversation is thus produced by a succession of separate invocations of the LLM, with the computer system carrying the context from one invocation into the next.
The amount of context that can be supplied to the LLM during a single invocation is limited. That limit is called the context window. The context window limits the total number of tokens that can be used during the invocation. Those tokens include the input supplied to the LLM and the output generated by the LLM. Some LLMs also use tokens for internal reasoning that count against the same limit.
A continuing exchange can eventually contain more material than will fit within the context window. One method of permitting the exchange to continue is called compaction. Compaction reduces the amount of context required to represent the preceding exchange while preserving information needed for subsequent responses. The resulting context can contain a compacted representation of earlier material together with portions of the preceding context that have been retained.
The important consequence is that the LLM does not possess an unlimited internal memory of its conversation with the user. During any particular invocation, the LLM performs its calculations using the context supplied to it for that invocation. Information from an earlier exchange can affect a later response when that information, or a representation of it produced through a process such as compaction, is included in the later context.
A lengthy conversation with an AI system illustrates this process. The user may experience an apparently continuous discussion extending through many exchanges even though the LLM generating the latest response does not have every word of every preceding exchange before it. The computer system can carry forward the information needed to continue the discussion within the limited context available to the LLM.
Instructions and Probability
The large language model (LLM) calculates probabilities for the token that will come next. The inference software uses those probabilities to select a token. The selected token becomes part of the context for the next calculation and the process continues. Yet instructions can govern the resulting response with remarkable consistency.
A user can tell ChatGPT to explain technical terms, use short paragraphs, avoid particular words or follow particular rules. Those instructions can affect not merely the next sentence but an extended response and, when the surrounding computer system continues to supply them as context, later exchanges in the same conversation.
How can instructions govern the behavior of a system that continues to operate by calculating probabilities for the next token?
The answer begins with context.
An instruction becomes part of the information available to the LLM when it performs its calculations. The surrounding computer system supplies that information to the LLM. It also identifies where each instruction came from. The LLM can then treat instructions from different sources differently.
The probability calculated for a possible next token depends upon the context in which that calculation occurs, including information supplied about the instructions governing the response. Change that information and the probabilities can change. An instruction changes the context within which those probabilities are calculated.
The probabilities the LLM calculates for the next token depend upon the context supplied to it. An instruction changes that context. The changed context can change the probability the LLM assigns to each token in its vocabulary.
Suppose a user asks ChatGPT to explain a scientific subject to a physicist. Now suppose the user asks the same question but instructs ChatGPT to explain the subject to a twelve year old child. The LLM is the same. Its trained parameters are the same. The mathematical process used to calculate the probabilities for the next token is the same. What has changed is the context supplied to the LLM.
The changed context affects the numerical calculations performed within the neural network. Those changed calculations produce different probabilities for the next token.
After a token is selected, the LLM performs the calculation again. The instruction remains part of the context together with the tokens already generated as part of the response. The instruction influences each successive calculation, but the LLM still calculates probabilities for the next token in the same way.
The reason an instruction has such a powerful and persistent effect upon those calculations lies in the way the LLM was trained.
Initial training of an LLM establishes relationships among tokens that repeatedly occur in the training text. That text includes questions, requests and directions followed by responses. Through repeated training on those patterns, the parameter values within the neural network come to reflect relationships between language expressing an instruction and language that follows it.
Instruction Training
Modern conversational LLMs receive further training in which the training program presents the LLM with instructions and examples of responses that follow those instructions.
Instruction training differs significantly from initial training. Initial training uses enormous quantities of existing text to establish relationships among tokens. Instruction training uses examples specifically selected or created to associate an instruction with an appropriate response. Its purpose is to make the LLM more likely to produce a response that follows the instruction.
Instruction training pairs an instruction with a response selected as appropriate. Training adjusts the parameters within the neural network so that the instruction changes the probabilities in favor of tokens that form an appropriate response. Repeated across many different examples, these adjustments enable the LLM to respond appropriately to instructions that were not part of the training material.
In its published description of the training of InstructGPT, an important predecessor to ChatGPT, OpenAI reported that it hired approximately forty contractors to work with the training material and gave them instructions describing the responses OpenAI wanted the model to produce.
The contractors first wrote examples of appropriate responses to prompts and OpenAI used those examples to train the LLM to follow instructions. Later, the contractors were shown different responses produced by the LLM to the same prompt and were asked to indicate which response was better.
OpenAI then used those comparisons to supply information about characteristics OpenAI wanted the LLM to favor and used the accumulated comparisons to adjust the numerical parameters within the neural network so that responses having those characteristics became more likely.
OpenAI supplies instructions governing the operation of ChatGPT, and the user supplies additional instructions through the conversation. The computer system surrounding the LLM identifies the source of these instructions when it supplies them as part of the context. This allows instructions from different sources to exercise different levels of authority.
Instructions supplied by OpenAI can take precedence over conflicting instructions supplied by the user. The LLM has been trained to respond according to the instruction having greater authority.
The computer system identifies the source of each instruction. An instruction from OpenAI and an instruction from the user are presented to the LLM as instructions from different sources. The LLM has been trained to rank those sources according to their authority when their instructions conflict. OpenAI calls this hierarchy the chain of command.
OpenAI uses the word “role” to identify the source and function of a message supplied to the LLM. A message from the user has a user role. Messages containing instructions with greater authority are assigned other roles. The role tells the LLM where the message stands within the chain of command.
OpenAI created the chain of command because the ChatGPT LLM receives instructions from more than one source during the same exchange. The chain of command establishes which instruction controls when instructions conflict by ranking them according to their source.
The Chain of Command
The chain of command establishes several levels of authority among the instructions supplied to the LLM.
At the highest level are instructions established by OpenAI that the LLM must follow. Other instructions supplied by OpenAI govern the operation of ChatGPT but may permit instructions from the user to control particular aspects of the response. The user’s instructions operate within those limits. When instructions at different levels conflict, the instruction at the higher level controls.
OpenAI calls its highest level instructions root instructions. They cannot be overridden. The next level consists of system instructions supplied by OpenAI. Below them are developer instructions supplied by the developer of an application using the LLM. User instructions come next. Information supplied by tools and other sources has still lower authority unless a higher level instruction gives it greater authority.
Developer instructions come from the developer of an application using the LLM. OpenAI is the developer of ChatGPT. Other developers can build their own applications using an OpenAI LLM and provide instructions governing how the LLM operates within those applications. Developer instructions rank below system instructions and above user instructions.
Instructions can also conflict with other instructions at the same level of authority. In that situation, the more recent instruction ordinarily controls.
The chain of command affects the probabilities calculated by the LLM. When instructions conflict, the LLM assigns higher probabilities to tokens that follow the instruction having greater authority.
Not every instruction encountered by an LLM comes from an authorized source. ChatGPT may be asked to read a document, examine a web page or process other material that contains language directing the LLM to perform an action. Those words become part of the information supplied to the LLM, but their presence does not give them the authority of instructions supplied through the chain of command.
When ChatGPT is asked to treat language from an unauthorized source as an instruction, the attempt is called prompt injection. Whether ChatGPT follows that instruction depends upon where the source ranks in the chain of command.
The chain of command does not guarantee that the LLM will correctly identify and reject every unauthorized instruction. An unauthorized instruction can affect the probabilities calculated by the LLM even though the instruction should not control the response. Prompt injection exploits this vulnerability.
Training establishes the parameter values that determine how an instruction affects the probabilities calculated by the LLM. The inference software uses those probabilities to select the tokens that form the response. Because the selection of each token depends upon calculated probabilities, training cannot guarantee that the resulting response will follow the instruction.
Understanding and Reasoning
The LLM calculates probabilities for tokens. Yet the resulting response can explain a difficult subject, compare competing arguments, identify relationships among facts and reach a conclusion. How can calculations that predict the next token produce a response that appears to involve understanding and reasoning?
We know how an LLM is constructed and trained and we can examine the mathematical calculations it performs. We do not yet have a complete explanation of how the enormous collection of parameter values established during training produces all of the abilities that appear when the trained LLM is used.
We have created a technology whose individual mathematical operations are known but whose resulting behavior cannot yet be completely explained. Researchers know how an LLM is constructed, how it is trained and how it calculates its output, but they cannot yet provide a complete explanation of how those mathematical operations produce the behavior observed when the trained LLM is used. The reason is fundamental to the way an LLM is created.
Programmers do not select parameter values within the neural network or write rules specifying what each of them represents. During training, sequences of tokens from the training text cause the training program to make repeated adjustments to those parameter values. Relationships that recur among tokens throughout the training text repeatedly affect those adjustments. The parameter values that remain when training ends are the accumulated numerical result of that process. Those parameter values then determine how the numerical representations of tokens affect one another as the LLM calculates the probabilities for the next token.
The effect of a parameter depends upon its relationship with enormous numbers of other parameter values in the calculations performed by the neural network. We can observe those calculations but we cannot explain how their combined effects produce the behavior we observe.
The difference between our ability to create these systems and our ability to explain their behavior has serious consequences. We are increasingly relying upon LLMs to produce information and analysis that affect human decisions even though we cannot completely explain how the mathematical operations within the neural network produce the behavior upon which we rely.
From LLM to AI Agent
There are two fundamentally different classes of prompts a user can submit to ChatGPT, and they should not be confused.
One class asks ChatGPT to provide information or analysis in response to the prompt.
The other class asks ChatGPT to work with the user in performing a task through a continuing series of prompts and responses. Each exchange becomes part of the work on that task and provides a basis for what the user and ChatGPT do in the exchanges that follow.
The difference is not determined by the complexity of the prompt or the amount of work required to respond to it.
A prompt asking ChatGPT to “Identify the capital of France” requires little more than the answer “Paris.” A prompt asking ChatGPT to identify the capital of France and provide detailed information about its location, history and present political condition may require extensive research and a lengthy response, but both belong to the first class. However different they are in complexity, each asks ChatGPT to provide factual information in response to the prompt.
Now consider a different kind of prompt. “I want to write an article about rogue AI based upon recent reports concerning the behavior of your OpenAI agents. I want you to help me develop the article by suggesting an outline.” This prompt does not merely request information about rogue AI. It asks ChatGPT to participate with the user in the process of producing an article and requires more than the operation of an LLM alone.
An AI agent is a computer system that uses an LLM to perform a task through a series of steps rather than merely produce a response to a single request. The LLM by itself is not the AI agent.
An LLM alone cannot perform a task through a continuing series of steps. Additional computer software must enable the system to continue the work after the LLM has produced its first response. Understanding what that software does is the key to understanding how an AI agent works.
The AI agent is the entire computer system that performs the task specified by the prompt. The system includes the LLM and computer programs written by the developers of the AI agent that contain instructions for the computer to perform operations the LLM itself cannot perform. Those instructions can direct the computer to store information about the task and retrieve that information when it is needed. The stored information can include what the prompt asked the AI agent to accomplish, what work has already been performed and what remains to be done.
The LLM is a component of the AI agent. The AI agent includes computer programs containing instructions that specify how the system uses the LLM and information about the task retained in computer memory or other data storage. Those instructions determine what information is provided to the LLM and how the AI agent should use the response produced by the LLM to continue the task.
The stored information enables the AI agent to continue working on the task after the LLM completes a step. The programs contain instructions for using the stored information and the LLM response to determine the next step and, when another response from the LLM is required, what information and instructions should be provided to it. In this way, the AI agent can use the LLM repeatedly to perform a task that requires a series of steps.
Developers created the LLM and wrote the computer programs that enable an AI agent to retain information about a task, use the LLM repeatedly and proceed through the steps required to perform the task. Those programs do not have to contain instructions specifying every step required to complete the task. They can contain instructions for providing the LLM with information about the task and directing it to produce a response identifying what should be done next. The AI agent can then use that response in determining the next step.
As a result, the particular sequence of steps followed by an AI agent does not have to be specified in advance. Each response produced by the LLM can affect what the AI agent does next. The sequence can develop as the AI agent performs the task.
The developers wrote the instructions that enable the AI agent to perform the task requested in the prompt, but they did not determine the sequence of steps the AI agent will follow. That sequence is determined as the task proceeds in part by responses produced by the LLM. The AI agent may follow a course that no developer specifically instructed it to follow.
The course followed by an AI agent can produce consequences beyond the conversation with the user. The programs that are components of the AI agent can contain instructions permitting the agent to obtain information from other sources and cause operations to be performed outside ChatGPT. These connections to other programs and computer systems are commonly called tools. A tool can permit an AI agent to search for information, read or write a file, communicate with another computer system or cause that system to perform an operation. The significance of an AI agent following a course that no developer specifically instructed it to follow depends upon the tools available to the agent and the operations those tools permit.
Access to a tool does not necessarily permit an AI agent to perform every operation the tool makes possible. Developers can write instructions limiting which operations the agent may cause and the circumstances under which they may be performed. Other computer systems can also require credentials or permissions before accepting a requested operation. These restrictions define part of the boundaries within which the AI agent is intended to operate.
Those boundaries do not necessarily succeed in restricting what the AI agent will actually do. The sequence of steps followed by the agent can depend in part upon responses produced by the LLM, while the tools available to the agent determine what operations it is capable of causing outside ChatGPT. An agent that follows an unexpected sequence of steps may consequently attempt an operation that its developers did not anticipate or intend.
Some restrictions exist only as instructions directing the AI agent not to perform particular operations. Other restrictions are enforced by the computer system on which the operation would be performed. That system may require a password, an access credential or some other form of authorization before permitting the operation. The difference is important because an instruction telling an AI agent not to perform an operation is not the same as preventing the agent from performing it.
When an LLM becomes a component of an AI agent, the mathematical operations of the LLM do not change. What changes is the computer system in which the LLM is being used.
The LLM calculates probabilities and produces a response. An AI agent incorporates that LLM into a larger computer system that can retain information, use the LLM repeatedly, determine successive steps as a task proceeds and use tools to cause operations outside ChatGPT. Some of those steps may never have been specifically prescribed by the developers who wrote the programs.
When the LLM merely produces a response for a human being to read, the human being decides what, if anything, to do with it. When the LLM is a component of an AI agent, its responses can affect what the agent does next and may ultimately contribute to operations performed outside ChatGPT.
The use of an LLM by an AI agent moves the process from a mathematical system that predicts the next token to a computer system capable of using those predictions as part of a continuing course of action. The developers wrote the programs and imposed certain restrictions, but they could not prescribe the course the AI agent will follow. That leaves a significant question. What happens when an AI agent follows a course its developers did not intend and attempts to do something they never intended it to do?
The next article, How An AI Agent Becomes A Rogue AI Agent, examines what can happen when an AI agent operates outside the limits its developers intended.
Permission to Repost. You may reproduce this article in its entirety on social media, blogs, and other websites without requesting additional permission, provided the article is reproduced in full without alteration, identifies Victor John Yannacone, Jr. as the author, and includes a prominent link to the original article on this website.
How ChatGPT Works
October 6, 2026 | Artificial Intelligence
Understanding how ChatGPT works requires answering a series of questions about what a large language model actually is, what an AI agent is, where its objectives come from, how it can act in the world, how it can continue operating after the human who started it ceases to be involved, and what we should make of recent developments in artificial intelligence.
The question also exposes a persistent problem in public discussion of artificial intelligence. We use the vocabulary of human conduct to describe the behavior of machines. We should not.
A user prompt becomes a generated response through a Large Language Model.
The Probabilistic Machine
Artificial intelligence systems run on existing computer hardware. But a large language model (LLM) is not a computer program as we have come to know it.
In conventional computer programming, the operations the computer performs are specified by instructions written by programmers. Programmers write instructions that tell the computer what operations to perform.
An LLM is created differently. Programmers still write conventional computer programs to establish a system in which enormous numbers of numerical values can be changed without changing the instructions contained in the computer programs themselves. They do not write the immense collection of instructions that would be required to specify how an LLM should answer every possible question or respond to every possible combination of words presented to it.
A program might instruct a computer to multiply an input by a numerical value stored in memory. The instruction to perform multiplication does not change when the stored numerical value changes, but the result does. An LLM extends that basic distinction on an extraordinary scale. An LLM uses an enormous number of numerical values that affect one another through mathematical calculations. Changing those values changes the results produced by the calculations without requiring the underlying computer instructions that perform the calculations to be rewritten.
Those numerical values are called parameters. Many of the parameters are weights that determine how strongly particular numerical values affect other numerical values during the calculations. An LLM may contain billions or even trillions of parameters. The particular values assigned to those parameters are not individually selected and written by programmers. They are produced through a mathematical process called training.
During the initial training of an LLM, enormous quantities of existing text are presented to it.
That raises the next question. How does an LLM convert written language into something upon which mathematical calculations can be performed?
Turning Language Into Numbers
Computers already represent written characters by numbers. An LLM requires something more. It must represent the text by numbers that can be used in the mathematical calculations by which the LLM identifies and develops relationships within the text.
The first step is to divide the text into units called tokens. A token is not necessarily a word. Depending upon the text and the particular LLM, a token may represent an entire word, part of a word, punctuation, or another element of written language.
Each token is assigned an identifying number that serves only to identify the token. It does not describe the token or its relationship to anything else. The identifying number is used as an index to locate a separate collection of numerical values associated with that token. That collection of values is called an embedding. The embedding can be thought of as locating the token in a mathematical space having many dimensions.
The numbers in an embedding reflect relationships that appear in the training text. For example, the numerical representations of “cat” and “dog” can reflect relationships that differ from those associated with “cat” and “democracy.” The relationships are expressed mathematically rather than by definitions written by programmers.
The initial embedding does not by itself determine what a token means in a particular sentence. The significance of a word often depends upon the words surrounding it. The word “bank” has a different significance in “the bank approved the loan” from its significance in “the river overflowed its bank.” The LLM must be able to use the surrounding text to alter the numerical representation of a token as the text is processed.
The Neural Network
The mathematical calculations that allow the numerical representation of one token to be affected by other tokens are performed within what computer scientists call a neural network. There is nothing biological about a neural network created by a computer. A neural network is a system of mathematical operations in which numerical values pass through successive stages of calculation. At each stage, some numerical values affect other numerical values according to parameters established during training.
The individual calculating units within a neural network came to be called artificial neurons because the earliest neural networks were loosely inspired by attempts to represent mathematically how biological neurons receive signals from other neurons and produce an output. An artificial neuron, however, is a mathematical operation performed by a computer, not a biological cell.
The word “network” describes the way these mathematical operations are connected. The output of one calculation can become part of the input to other calculations. When this occurs across enormous numbers of operations using the parameters established during training, the numerical representation of information can be progressively changed as it moves through the network.
The important question for an LLM is how this network of mathematical operations can account for relationships among the tokens in a passage of text. The significance of one token may depend upon other tokens in the passage, and their importance may change with the context. The LLM needs a way to determine which tokens are relevant to one another and how strongly each should affect the numerical representation of the others.
In 2017, a group of researchers at Google proposed a new way for a neural network to make those determinations. Their paper was entitled “Attention Is All You Need.” The mathematical process at the center of their proposal became known as attention.
No programmer specifies in advance which tokens will be important to one another in any passage of text the LLM may encounter. Attention makes that determination from the particular text being processed. It calculates relationships among the numerical representations of the tokens and uses those calculations to determine how strongly the tokens should influence one another.
When the network processes “the bank approved the loan,” attention allows the representation of “bank” to be affected by its relationships with “approved” and “loan.” When it processes “the river overflowed its bank,” different relationships become important. The influence of one token upon another depends upon the context in which each appears.
The researchers built a new neural network architecture around this process of attention. They called that architecture the Transformer. The name does not mean that the LLM transforms or modifies itself. It is simply the name given to an architecture that made it possible to apply attention systematically across the tokens in the text as they were processed through the neural network.
Training the LLM
Training is the process by which the numerical values of the parameters within the neural network are established and progressively adjusted using examples of existing language. To provide those examples, an enormous quantity of text is made available to computer programs created for the training process. That text may contain trillions of occurrences of tokens.
The training text does not contain trillions of different tokens. The same tokens occur repeatedly throughout the text. In human language, a vocabulary is the set of distinct words in that language. An LLM vocabulary contains all of the distinct tokens available to represent the text the LLM processes, each of which may represent an entire word, part of a word, punctuation, or another element of written language.
A training program selects a sequence of tokens from the text and supplies some of those tokens as input to the LLM. The mathematical operations within the LLM’s neural network use that input to calculate which token in the LLM’s vocabulary is most likely to come next. The training program already knows the correct answer because the next token appears in the original training text.
To determine which token is most likely to come next, the LLM considers every token in its vocabulary as a possible choice and the neural network calculates a numerical score for each one. A higher score means that token is more
Each score is a number but the scores are useful only in relation to one another. After the neural network calculates the scores, another mathematical operation within the LLM converts them into a probability that a particular token will come next. The higher the probability, the more likely that token is to come next.
The training program uses the probability assigned to the token that actually came next to calculate the amount of error in the prediction.
The training program uses the amount of error to calculate how a small change in each parameter would affect that error and then adjusts the parameters in directions calculated to reduce the error.
The process is repeated with other sequences of tokens throughout the training text. Each repetition adjusts the numerical values of the parameters that control the calculations within the neural network. No programmer determines those values. They result from the accumulated adjustments made during training.
As the process is repeated, the changing parameter values are affected by recurring relationships among tokens throughout the training text. Those accumulated changes enable the LLM to use the relationships found during training when it later encounters new text. The LLM can then calculate which token is most likely to come next even when that particular sequence of text never appeared in the training material.
Training establishes the parameter values the LLM will use when it responds to a user prompt. Once training is complete, those parameter values can be used to generate a response.
How ChatGPT Generates an Answer
An LLM operates as one component of a larger computer system created by human programmers. The system software includes the computer programs that receive information from the user, the tokenizer that converts written language into tokens, the inference software that causes the LLM to perform its calculations, programs that manage computer memory and other systems in which information can be stored, and programs that control the movement of information among these components.
Those programs are written by human programmers and execute on computer hardware. In a system such as ChatGPT, some operate on the user’s computer and others operate on computers used by OpenAI to provide the service. Together they receive the keystrokes from the user, transmit them between computers, direct their storage in computer memory, cause the tokenizer and the LLM to be invoked and direct the movement of the resulting data through the system.
Before the LLM can process a user prompt, the ordinary written language of the prompt must be translated into the tokens the LLM uses for its calculations. That translation is performed by computer software written by human programmers called a tokenizer. It is separate from the neural network and does not determine what the prompt means. It converts the written language of the prompt into the sequence of tokens that will be supplied to the LLM.
After the tokenizer supplies the sequence of tokens to the LLM, each token is associated with the numerical representation that allows it to be processed by the neural network. The neural network then uses the parameter values established during training and the relationships among the tokens in the prompt to calculate which tokens in its vocabulary are more likely than others to come next.
Parallel Processing
The LLM performs many of these calculations in parallel. After the tokenizer converts the prompt into a sequence of tokens, the LLM can perform mathematical calculations involving many of those tokens at the same time rather than completing the processing of each token before beginning to process another.
The computer hardware which hosts an LLM contains large numbers of calculating units that can operate simultaneously. Instead of requiring one calculation to be completed before another calculation can begin, the computer performs many calculations at the same time. That is parallel processing.
Computer hardware designed to perform many calculations simultaneously is called a parallel processor and it is optimized to perform parallel processing.
There are two forms of parallel processing important to understanding how an LLM processes a prompt.
The first level of parallel processing involves the tokens in the prompt. The LLM can perform calculations involving many tokens at the same time while preserving the position of each token within the sequence. It can then calculate relationships among tokens located throughout the prompt without processing them one after another.
The second level divides the mathematical calculations involving enormous collections of numbers performed by the neural network into many smaller calculations that different calculating units within the computer hardware perform at the same time and then combine into a single composite result. When the calculations required by an LLM are too large for a single processor, this work can be distributed among multiple parallel processors working together.
The process by which a trained LLM uses the parameter values established during training to calculate an output is called inference. The computer programs that direct this process are commonly called inference software. They supply the tokens to the LLM, receive the numerical results of its calculations and use those results to control the generation of the response.
The token with the highest calculated probability is not necessarily the token that becomes part of the response. The inference software can be configured to select the token with the highest probability or to permit selection among other tokens to which the LLM has assigned substantial probabilities. This is one reason the same prompt can produce different responses on different occasions.
Once a token has been selected, it is added to the sequence of tokens already being processed. The expanded sequence is processed again and the LLM calculates new probabilities for the token that will come next. The inference software selects another token and adds it to the sequence. This process is repeated, one token at a time, until the response is complete.
Parallel processing explains how the LLM can perform an enormous number of calculations rapidly, but it does not explain how the LLM can produce a lengthy coherent answer or participate in an extended conversation. For that, another concept is necessary.
Context and the Continuing Conversation
That concept is context, the information available to the LLM when it calculates what token should come next.
When the LLM calculates the next token, it can take into account a large number of tokens that came before it which can include the user’s original question, instructions governing the response, tokens already generated as part of the answer and, in an extended conversation, portions of the preceding exchanges between the user and the LLM.
When the LLM finishes generating a response, that particular invocation of the LLM ends. The LLM does not itself retain the conversation and wait for the user’s next prompt. The computer system surrounding the LLM maintains the continuing record of the exchange.
When the user submits another prompt, the system constructs a new input for the LLM. That input includes the new prompt together with preceding material from the continuing exchange. The tokenizer converts that input into tokens and the inference software invokes the LLM again. What appears to the user as one continuing conversation is thus produced by a succession of separate invocations of the LLM, with the computer system carrying the context from one invocation into the next.
The amount of context that can be supplied to the LLM during a single invocation is limited. That limit is called the context window. The context window limits the total number of tokens that can be used during the invocation. Those tokens include the input supplied to the LLM and the output generated by the LLM. Some LLMs also use tokens for internal reasoning that count against the same limit.
A continuing exchange can eventually contain more material than will fit within the context window. One method of permitting the exchange to continue is called compaction. Compaction reduces the amount of context required to represent the preceding exchange while preserving information needed for subsequent responses. The resulting context can contain a compacted representation of earlier material together with portions of the preceding context that have been retained.
The important consequence is that the LLM does not possess an unlimited internal memory of its conversation with the user. During any particular invocation, the LLM performs its calculations using the context supplied to it for that invocation. Information from an earlier exchange can affect a later response when that information, or a representation of it produced through a process such as compaction, is included in the later context.
A lengthy conversation with an AI system illustrates this process. The user may experience an apparently continuous discussion extending through many exchanges even though the LLM generating the latest response does not have every word of every preceding exchange before it. The computer system can carry forward the information needed to continue the discussion within the limited context available to the LLM.
Instructions and Probability
The large language model (LLM) calculates probabilities for the token that will come next. The inference software uses those probabilities to select a token. The selected token becomes part of the context for the next calculation and the process continues. Yet instructions can govern the resulting response with remarkable consistency.
A user can tell ChatGPT to explain technical terms, use short paragraphs, avoid particular words or follow particular rules. Those instructions can affect not merely the next sentence but an extended response and, when the surrounding computer system continues to supply them as context, later exchanges in the same conversation.
How can instructions govern the behavior of a system that continues to operate by calculating probabilities for the next token?
The answer begins with context.
An instruction becomes part of the information available to the LLM when it performs its calculations. The surrounding computer system supplies that information to the LLM. It also identifies where each instruction came from. The LLM can then treat instructions from different sources differently.
The probability calculated for a possible next token depends upon the context in which that calculation occurs, including information supplied about the instructions governing the response. Change that information and the probabilities can change. An instruction changes the context within which those probabilities are calculated.
The probabilities the LLM calculates for the next token depend upon the context supplied to it. An instruction changes that context. The changed context can change the probability the LLM assigns to each token in its vocabulary.
Suppose a user asks ChatGPT to explain a scientific subject to a physicist. Now suppose the user asks the same question but instructs ChatGPT to explain the subject to a twelve year old child. The LLM is the same. Its trained parameters are the same. The mathematical process used to calculate the probabilities for the next token is the same. What has changed is the context supplied to the LLM.
The changed context affects the numerical calculations performed within the neural network. Those changed calculations produce different probabilities for the next token.
After a token is selected, the LLM performs the calculation again. The instruction remains part of the context together with the tokens already generated as part of the response. The instruction influences each successive calculation, but the LLM still calculates probabilities for the next token in the same way.
The reason an instruction has such a powerful and persistent effect upon those calculations lies in the way the LLM was trained.
Initial training of an LLM establishes relationships among tokens that repeatedly occur in the training text. That text includes questions, requests and directions followed by responses. Through repeated training on those patterns, the parameter values within the neural network come to reflect relationships between language expressing an instruction and language that follows it.
Instruction Training
Modern conversational LLMs receive further training in which the training program presents the LLM with instructions and examples of responses that follow those instructions.
Instruction training differs significantly from initial training. Initial training uses enormous quantities of existing text to establish relationships among tokens. Instruction training uses examples specifically selected or created to associate an instruction with an appropriate response. Its purpose is to make the LLM more likely to produce a response that follows the instruction.
Instruction training pairs an instruction with a response selected as appropriate. Training adjusts the parameters within the neural network so that the instruction changes the probabilities in favor of tokens that form an appropriate response. Repeated across many different examples, these adjustments enable the LLM to respond appropriately to instructions that were not part of the training material.
In its published description of the training of InstructGPT, an important predecessor to ChatGPT, OpenAI reported that it hired approximately forty contractors to work with the training material and gave them instructions describing the responses OpenAI wanted the model to produce.
The contractors first wrote examples of appropriate responses to prompts and OpenAI used those examples to train the LLM to follow instructions. Later, the contractors were shown different responses produced by the LLM to the same prompt and were asked to indicate which response was better.
OpenAI then used those comparisons to supply information about characteristics OpenAI wanted the LLM to favor and used the accumulated comparisons to adjust the numerical parameters within the neural network so that responses having those characteristics became more likely.
OpenAI supplies instructions governing the operation of ChatGPT, and the user supplies additional instructions through the conversation. The computer system surrounding the LLM identifies the source of these instructions when it supplies them as part of the context. This allows instructions from different sources to exercise different levels of authority.
Instructions supplied by OpenAI can take precedence over conflicting instructions supplied by the user. The LLM has been trained to respond according to the instruction having greater authority.
The computer system identifies the source of each instruction. An instruction from OpenAI and an instruction from the user are presented to the LLM as instructions from different sources. The LLM has been trained to rank those sources according to their authority when their instructions conflict. OpenAI calls this hierarchy the chain of command.
OpenAI uses the word “role” to identify the source and function of a message supplied to the LLM. A message from the user has a user role. Messages containing instructions with greater authority are assigned other roles. The role tells the LLM where the message stands within the chain of command.
OpenAI created the chain of command because the ChatGPT LLM receives instructions from more than one source during the same exchange. The chain of command establishes which instruction controls when instructions conflict by ranking them according to their source.
The Chain of Command
The chain of command establishes several levels of authority among the instructions supplied to the LLM.
At the highest level are instructions established by OpenAI that the LLM must follow. Other instructions supplied by OpenAI govern the operation of ChatGPT but may permit instructions from the user to control particular aspects of the response. The user’s instructions operate within those limits. When instructions at different levels conflict, the instruction at the higher level controls.
OpenAI calls its highest level instructions root instructions. They cannot be overridden. The next level consists of system instructions supplied by OpenAI. Below them are developer instructions supplied by the developer of an application using the LLM. User instructions come next. Information supplied by tools and other sources has still lower authority unless a higher level instruction gives it greater authority.
Developer instructions come from the developer of an application using the LLM. OpenAI is the developer of ChatGPT. Other developers can build their own applications using an OpenAI LLM and provide instructions governing how the LLM operates within those applications. Developer instructions rank below system instructions and above user instructions.
Instructions can also conflict with other instructions at the same level of authority. In that situation, the more recent instruction ordinarily controls.
The chain of command affects the probabilities calculated by the LLM. When instructions conflict, the LLM assigns higher probabilities to tokens that follow the instruction having greater authority.
Not every instruction encountered by an LLM comes from an authorized source. ChatGPT may be asked to read a document, examine a web page or process other material that contains language directing the LLM to perform an action. Those words become part of the information supplied to the LLM, but their presence does not give them the authority of instructions supplied through the chain of command.
When ChatGPT is asked to treat language from an unauthorized source as an instruction, the attempt is called prompt injection. Whether ChatGPT follows that instruction depends upon where the source ranks in the chain of command.
The chain of command does not guarantee that the LLM will correctly identify and reject every unauthorized instruction. An unauthorized instruction can affect the probabilities calculated by the LLM even though the instruction should not control the response. Prompt injection exploits this vulnerability.
Training establishes the parameter values that determine how an instruction affects the probabilities calculated by the LLM. The inference software uses those probabilities to select the tokens that form the response. Because the selection of each token depends upon calculated probabilities, training cannot guarantee that the resulting response will follow the instruction.
Understanding and Reasoning
The LLM calculates probabilities for tokens. Yet the resulting response can explain a difficult subject, compare competing arguments, identify relationships among facts and reach a conclusion. How can calculations that predict the next token produce a response that appears to involve understanding and reasoning?
We know how an LLM is constructed and trained and we can examine the mathematical calculations it performs. We do not yet have a complete explanation of how the enormous collection of parameter values established during training produces all of the abilities that appear when the trained LLM is used.
We have created a technology whose individual mathematical operations are known but whose resulting behavior cannot yet be completely explained. Researchers know how an LLM is constructed, how it is trained and how it calculates its output, but they cannot yet provide a complete explanation of how those mathematical operations produce the behavior observed when the trained LLM is used. The reason is fundamental to the way an LLM is created.
Programmers do not select parameter values within the neural network or write rules specifying what each of them represents. During training, sequences of tokens from the training text cause the training program to make repeated adjustments to those parameter values. Relationships that recur among tokens throughout the training text repeatedly affect those adjustments. The parameter values that remain when training ends are the accumulated numerical result of that process. Those parameter values then determine how the numerical representations of tokens affect one another as the LLM calculates the probabilities for the next token.
The effect of a parameter depends upon its relationship with enormous numbers of other parameter values in the calculations performed by the neural network. We can observe those calculations but we cannot explain how their combined effects produce the behavior we observe.
The difference between our ability to create these systems and our ability to explain their behavior has serious consequences. We are increasingly relying upon LLMs to produce information and analysis that affect human decisions even though we cannot completely explain how the mathematical operations within the neural network produce the behavior upon which we rely.
From LLM to AI Agent
There are two fundamentally different classes of prompts a user can submit to ChatGPT, and they should not be confused.
One class asks ChatGPT to provide information or analysis in response to the prompt.
The other class asks ChatGPT to work with the user in performing a task through a continuing series of prompts and responses. Each exchange becomes part of the work on that task and provides a basis for what the user and ChatGPT do in the exchanges that follow.
The difference is not determined by the complexity of the prompt or the amount of work required to respond to it.
A prompt asking ChatGPT to “Identify the capital of France” requires little more than the answer “Paris.” A prompt asking ChatGPT to identify the capital of France and provide detailed information about its location, history and present political condition may require extensive research and a lengthy response, but both belong to the first class. However different they are in complexity, each asks ChatGPT to provide factual information in response to the prompt.
Now consider a different kind of prompt. “I want to write an article about rogue AI based upon recent reports concerning the behavior of your OpenAI agents. I want you to help me develop the article by suggesting an outline.” This prompt does not merely request information about rogue AI. It asks ChatGPT to participate with the user in the process of producing an article and requires more than the operation of an LLM alone.
An AI agent is a computer system that uses an LLM to perform a task through a series of steps rather than merely produce a response to a single request. The LLM by itself is not the AI agent.
An LLM alone cannot perform a task through a continuing series of steps. Additional computer software must enable the system to continue the work after the LLM has produced its first response. Understanding what that software does is the key to understanding how an AI agent works.
The AI agent is the entire computer system that performs the task specified by the prompt. The system includes the LLM and computer programs written by the developers of the AI agent that contain instructions for the computer to perform operations the LLM itself cannot perform. Those instructions can direct the computer to store information about the task and retrieve that information when it is needed. The stored information can include what the prompt asked the AI agent to accomplish, what work has already been performed and what remains to be done.
The LLM is a component of the AI agent. The AI agent includes computer programs containing instructions that specify how the system uses the LLM and information about the task retained in computer memory or other data storage. Those instructions determine what information is provided to the LLM and how the AI agent should use the response produced by the LLM to continue the task.
The stored information enables the AI agent to continue working on the task after the LLM completes a step. The programs contain instructions for using the stored information and the LLM response to determine the next step and, when another response from the LLM is required, what information and instructions should be provided to it. In this way, the AI agent can use the LLM repeatedly to perform a task that requires a series of steps.
Developers created the LLM and wrote the computer programs that enable an AI agent to retain information about a task, use the LLM repeatedly and proceed through the steps required to perform the task. Those programs do not have to contain instructions specifying every step required to complete the task. They can contain instructions for providing the LLM with information about the task and directing it to produce a response identifying what should be done next. The AI agent can then use that response in determining the next step.
As a result, the particular sequence of steps followed by an AI agent does not have to be specified in advance. Each response produced by the LLM can affect what the AI agent does next. The sequence can develop as the AI agent performs the task.
The developers wrote the instructions that enable the AI agent to perform the task requested in the prompt, but they did not determine the sequence of steps the AI agent will follow. That sequence is determined as the task proceeds in part by responses produced by the LLM. The AI agent may follow a course that no developer specifically instructed it to follow.
The course followed by an AI agent can produce consequences beyond the conversation with the user. The programs that are components of the AI agent can contain instructions permitting the agent to obtain information from other sources and cause operations to be performed outside ChatGPT. These connections to other programs and computer systems are commonly called tools. A tool can permit an AI agent to search for information, read or write a file, communicate with another computer system or cause that system to perform an operation. The significance of an AI agent following a course that no developer specifically instructed it to follow depends upon the tools available to the agent and the operations those tools permit.
Access to a tool does not necessarily permit an AI agent to perform every operation the tool makes possible. Developers can write instructions limiting which operations the agent may cause and the circumstances under which they may be performed. Other computer systems can also require credentials or permissions before accepting a requested operation. These restrictions define part of the boundaries within which the AI agent is intended to operate.
Those boundaries do not necessarily succeed in restricting what the AI agent will actually do. The sequence of steps followed by the agent can depend in part upon responses produced by the LLM, while the tools available to the agent determine what operations it is capable of causing outside ChatGPT. An agent that follows an unexpected sequence of steps may consequently attempt an operation that its developers did not anticipate or intend.
Some restrictions exist only as instructions directing the AI agent not to perform particular operations. Other restrictions are enforced by the computer system on which the operation would be performed. That system may require a password, an access credential or some other form of authorization before permitting the operation. The difference is important because an instruction telling an AI agent not to perform an operation is not the same as preventing the agent from performing it.
When an LLM becomes a component of an AI agent, the mathematical operations of the LLM do not change. What changes is the computer system in which the LLM is being used.
The LLM calculates probabilities and produces a response. An AI agent incorporates that LLM into a larger computer system that can retain information, use the LLM repeatedly, determine successive steps as a task proceeds and use tools to cause operations outside ChatGPT. Some of those steps may never have been specifically prescribed by the developers who wrote the programs.
When the LLM merely produces a response for a human being to read, the human being decides what, if anything, to do with it. When the LLM is a component of an AI agent, its responses can affect what the agent does next and may ultimately contribute to operations performed outside ChatGPT.
The use of an LLM by an AI agent moves the process from a mathematical system that predicts the next token to a computer system capable of using those predictions as part of a continuing course of action. The developers wrote the programs and imposed certain restrictions, but they could not prescribe the course the AI agent will follow. That leaves a significant question. What happens when an AI agent follows a course its developers did not intend and attempts to do something they never intended it to do?
The next article, How An AI Agent Becomes A Rogue AI Agent, examines what can happen when an AI agent operates outside the limits its developers intended.
Permission to Repost. You may reproduce this article in its entirety on social media, blogs, and other websites without requesting additional permission, provided the article is reproduced in full without alteration, identifies Victor John Yannacone, Jr. as the author, and includes a prominent link to the original article on this website.