Let's jump back one step further, to better understand how the LLMs which operate in harness software, actually work:
The brains behind AI products such as ChatGPT, Claude, and Gemini were created by training a statistical algorithm program to predict the next most likely word, which would most likely be expected to follow after any given series of input words. The results of this training routine are stored in billions/trillions of 'parameters', which are like little knobs which collectively adjust how any given input will be transformed into a selected output. This training routine basically finds patterns in gigantic collections of data, which are tuned to be more efficient/correct throughout a series of 'backpropagation' iterations. The algorithm uses repeated runs through the data to adjust parameter settings, based on a 'loss' evaluation which checks how well input values have produced predicted output values. This iteration process involves so much data and so many computed evaluations, that it would take millions of years complete on a normal desktop computer.
Companies such as OpenAI, Anthropic, Deepseek, Google, and others, spend many millions of dollars running this astronomically long statistical pre-training process, to build 'base' LLM models which are surprisingly capable of producing a next word which is actually meaningful within the full context of all the words that came before.
By 'meaningful' I mean, for example, that a base model will not just provide the response 'the butler did it' at the end of an input murder mystery text, because that's perhaps the statistically most common set of words, but instead will provide an answer such as 'the professor did it', in a case where that response makes more sense, given the meaningful arrangement of related words in a particular murder mystery text that come before the predicted output.
The output provided by the LLM represents a kind of understanding about how all the input words are related to one another, based upon patterns found in the trillions of curated words in its training text. The parameters in an LLM model are used to compute how every word relates to every other word in an input text, and how those deeply complex multi-dimensional relationships influence the likelihood of a given next word, based on the massively complex set of patterns stored in its parameters. The knowledge extracted from patterns in an LLMs training text is applied to a user's input text, to generate (or 'infer') a stream of statistically likely correct output text.
Pretrained base models which simply provide a stream of likely next text tokens, however, aren't really good at answering questions. Instead, they just continue the text which would be best expected after any given input text, by reflecting on all of the words that exist in an input submission, in the order/context they appear.
To turn a base LLM into a actually useful thinking machine that can help users solve problems, additional stages of 'post-training' are required. In the post-training phases, the algorithm program is basically fed millions of question and answer pairs, so that the model parameters learn the general shapes of question and answer inputs, and begin to naturally provide answers, whenever inputs in the shape of questions are submitted.
This tendency to produce answer responses to questions, after seeing many provided question-answer example texts, is simply an extension of how the training algorithm learns from patterns in input data.
Without post-training, if a base LLM saw the question 'The capital of France is ___", for example, repeatedly in quiz texts in training data, it might be more likely to simply output the rest of the surrounding quiz questions, instead of an answer, because it doesn't know to do anything else except continue text with the next most likely words it has seen in training data.
Post-training question and answer pairs traditionally are groomed by the scientists who create an LLM. Not only do these training steps make the LLM recognize the 'shape' of question and answer input patterns, they also train the model to answer with particular flavors, and with preferred sorts of responses, when a certain type of input is provided by the user.
Researchers provide question and answer pairs which train the model to provide safe responses, and responses which reflect a responsibly crafted 'constitution' - a general character/nature which the scientists intend the model to exhibit. They provide millions/billions of examples of that behavior, within the post-training text corpora.
Post-training also often involves automated 'reinforcement' learning stages, where the models learn to simply try millions/billions of trial and error self-play iterations, to solve problems. These iterations are rewarded positively when they find the correct solutions to math, programming, and other tasks which have verifiable, correct answers. This sort of self-play enables the models to learn in ways which go far beyond the data that humans prepare manually. It can in fact lead to capabilities which surpass human ability, because the models are able to learn from patterns which evolve through absolutely enormous volumes of trial and error tests (many more iterations that any human, or even large groups of humans, could perform in their lifetimes).
Post-training routines come as close to 'programming' a model to respond in a prescribed way, as is currently possible. This sort of training, based on providing a corpora of repetitive pattern shapes, however, still isn't the same as deterministic 'programming', in the way that traditional software is specifically formed from hard algorithmic rules defined by human engineers.
LLM parameter settings instead just 'emerge' stochastically from training, based on statistical pattern matching processes - instead of being built from set rules programmed by humans.
For that reason, we often say that LLMs are grown, rather than built. Emergent 'intelligent' capabilities simply come from the way AI models learn to respond, according to patterns that they find in training data. These patterns can be incredibly deep and meaningful, but they are still just predicted output based on enormously vast collections of carefully groomed input data.
So, a fully trained model simply accepts input data, and predicts the best most likely output data, based on its training. It does not remember anything about previous questions which have been entered by a user. This is a critically important concept to understand.
The 'emergent' capabilities which LLMs learn from their training data can be truly amazing to comprehend. They learn to translate languages, complete mathematical problems, and write perfectly functioning new software code, for example, without being programmed via any human-specified deterministic rules. As models grow in parameter count, and are trained on larger data sets, they tend to grow more complex and useful emergent capabilities.
No one knows exactly why or how all of a model's emergent capabilities are formed, because no human has ever written code to make them develop - they just appear within the output of a machine made to produce the next most likely token (word, character, etc.), following a sequence of input token. It's been keenly noted that useful emergent capabilities tend to grow with scale, so we just keep building bigger LLMs, with more parameters trained on more data.
And that's where the ability of models stop, and the agent harness applications pick up.
The chatbots which everyone got to know at the beginning of the LLM era (ChatGPT then Claude, Gemini,...) were simply pieces of software which enabled users to type in a prompt which was sent to an LLM, and after the model responded, it enabled the user to type in a new prompt, and that entire conversation history was sent back to the LLM, to provide another response. That loop just continued repeatedly. At the most basic level, chatbots can be very simple pieces of software, which just enable that loop.
The earliest agent harnesses were like chat loops which simply added the ability to work with files, operating-system commands, and other tools. They were programs which sent a layer of instructions to the LLM about how to format a tool call, and then executed that tool call whenever the model returned a properly formed tool call response (write (file) (content), for example).
Modern harnesses build much more orchestration around that basic idea, and the industry is working to expand those capabilities in many ways, but that's the basis of how agents work.
We have extremely intelligent models which, because of emergent capabilities born from enormous scale training routines, can reason though contexts of 1 million+ tokens (tokens are approximately 3/4 of a word), write code, and understand how to solve complex conceptual challenges - and we have many harness applications which give those brains some mix of abilities to work with files and the operating system where the harness runs. They provide well established ways of calling software tools, saving and loading instructions, and remembering information which needs to be recalled across prompt sessions.
Those are the basic pieces of every big chat and agentic system you've seen. Models cost millions of dollars to train, and require massive GPU computing power to produce, but you can build your own agent software with as little as a few hundred lines of code.
Learning how to interact with LLMs, so that they have all the tools required to respond to a prompt, with useful tool call output, generated code, and plans to work across many prompt iterations, is what makes it possible to solve very complex problems with LLMs - much more than can be accomplished in a chat loop.
A harness application is required to give the LLM agency to work with a surrounding operating environment - but you're still always reliant on the intelligence of the model to come up with intelligently reasoned responses, code, tool calls, documentation, etc., to get a job completed. You always need a smart enough brain to provide any sort of useful output. The harness just makes it possible to put that output to work.