HOW AGENTS WORK

261 views
Nick Antonaccio
Nick AntonaccioAdmin
Sep 01, 2026 at 22:27 (edited, 15 revisions)
#1

An LLM can be thought of like a brain. Smarter brains can understand and accomplish more challenging goals. But LLM brains can't do anything except 'think'. They just produce output tokens in response to input tokens.

A key concept to internalize about LLMs, is that every input to a model can be thought of almost like an entirely new, fully self-contained 'life experience' for that model. To the model, each input is something akin to being born again. The model has no memory of anything before or after that conversation prompt. It simply wakes up and receives a bunch of input, which it produces an output for.

The first goal of a harness (agent software program) is to provide some default instructions, which the harness application is programmed to send with every call to a model API, to provide some basic context about how to operate.

For example, something like this simplified prompt concept gets sent:

"You are a model being called from an agent harness running on a user's computer. If you want to write to a file on the user's computer, respond with the command 'write (filename) (content)'. If you want to read a file, respond with the command 'read (filename)'. If you want to run an operating system command, respond with the command 'os (command) (parameters)'. Read the contents of the agents.md file in your working directory now - it provides some basic rules you must follow, and some information about skills & tools you can use. You can edit your agents.md file at any time to save new instructions, whenever the user asks you to remember something about how to operate in future sessions, or when you need to remember important information in another session. Here is the user's prompt: (prompt text)"

The user's prompt text contains the entire conversation input/output history. That full conversation history, along with all the prompt instructions sent by the harness program, is called the 'context' of the session (in reality, the context can be more complex, but as a basic mental model, it can initially useful to think about the context as being a full copy of the entire conversation history).

The harness's default prompt text may also include some general instructions such as rules about creating written plans which can be followed across a number of separate sessions, to complete specific tasks for the user. There may be a default file in which plan information is stored, or a database table which the LLM can read to follow steps, for any plan/goal which has been initialized or partially/fully completed.

This sort of planning arrangement enables the user, for example, to tell the LLM to work on completing a long horizon goal, for which a plan gets created and then is updated/edited by the LLM periodically, as the plan progresses, to achieve steps toward that goal.

So, when a user types a request into an agent harness interface, the LLM model receives all the initial default instructions programmed into the harness software, it reads all the info in the current agents.md file, and it reads the entire conversation context (which contains every previous human request and LLM response in the conversation, along with the newest user prompt tacked on at the end) - and the LLM processes that input and responds - that single process forms an entire 'life' experience for the LLM.

The LLM has no other perspective or memory of anything but what's been submitted in that prompt context, and what's contained in its training. Every time a prompt call to the LLM API is sent, the entire life/death of the LLM's experience is contained fully in that single complex request/response.

So the harness is especially important, because it orchestrates not only the entire process which the LLM will go though for each request a user sends, but it also orchestrates everything the LLM knows about the environment it's working in, what tools are available, what steps in a workflow have been accomplished in previous sessions, what it should know about the user & their preferences, etc., and it helps the LLM communicate with it's next lifecycle, through all those well organized files, skills, tools, steps of a plan, subagent tasks which have been completed, etc., which are all stored and managed by the harness software,

If an LLM wants to store something to remember in a future session, which is not contained in the current conversation context, it needs to write down that information in a file, or save it in a database somewhere, along with instructions about where/how to find/use that information.

'Skills' are text files which explain steps in a complex process which have been previously worked out successfully, which may include information about 'tools' which have been developed to help complete a task.

The LLM, or a human, may write code to create a software tool which can be used to perform some task in the future. For example, it may write a program which interacts with the OS API that controls the windows of programs which run on the local operating system - call that program a 'computer control (tool', for example. Then the LLM or a human writes a bunch of instructions which explain how to use that program - call those instructions a 'computer control skill*'.

The LLM or a human could then write into the harness's agents.md file that instructions to use a 'computer control skill', to manipulate program windows on the local computer, are stored in a text file named ./computercontrol.skill, for example. A very short instruction is stored in the agents.md file to remember that that skill exists, along with a concise description of what it can be used to accomplish, and instructions to read that skill text file any time the LLM may want to control windows on the local computer.

The detailed skill file then contains all the complex instructions about what commands, in exactly what specified formats, the LLM should return, if a computer control tool call is required to complete a task (similarly to how default read, write, and OS tool call options are explained).

Importantly, the skill file only ever gets read if the LLM thinks that that skill may be useful for completing a specific task at hand. Otherwise, it never gets loaded, and thus never wastes any of the model's precious conversation context limit.

Skill files often include instructions to run another copy of the harness program (as a 'sub-agent' session, with its own separate context), to read the skill file and execute the tool call, so that the current session's conversation context doesn't get filled up with the technical details of that operation - that whole operation runs in another context that lives and dies, just to respond to that request...

So, a model does its thinking about how to solve a problem sent in a user request, and if, in the course of its solution, it thinks that writing a file to the hard drive is useful, for example, then it issues a write command, in the format the harness has defined and communicated. The harness application is simply programmed to write data to a file, in the specified format, whenever the model sends back a response in that format. The file that gets written, could be a Python code file, for example, which may be meant to be established as another custom tool that can be used later.

After that Python program has been created, then in the course of the model's thinking, it could at any time chose to run that Python file which was written to the user's hard drive. The model knows, from its training, how to call the Python interpreter to execute the code in that file, and how to interact with that program on the command line. LLMs have been specifically trained to be really good at performing this sort of work.

That's the basic architecture of how LLM harnesses work. Some harness apps come with a huge number of default instructions built in, along with a bunch of software tools and skill files that have been pre-created to perform operations such as controlling interactions in a web browser, or communicating with a user through a texting service, etc. But you can complete any task like that, using any basic LLM and agent harness combination, to create only the custom software tools, and to write only the custom skill instructions you need, to complete your own custom work - all from the ground up.

Pi comes with only the simplest tools set up by default, to read & write files, issue OS commands, connect with LLM provider APIs, load and unload skills/extensions, etc. Its default prompt includes instructions about where to find all the required documentation to edit not only its own configuration settings (such as how to hook up to other LLM model provider APIs which are not supplied in its default configuration settings), but also how to create extension code which modifies/extends its own source code. That enables the potential for the Pi harness to completely re-write/extend all its own internal capabilities, to enable completely new core software features (interaction with the operating system and programming SDKs, APIs, etc.).

Any user can alter/extend the Pi harness software itself, just by having the LLM write new harness code, create new tools, and write new skill instructions - all entirely from scratch, just by prompting the connected LLM brain It's the LLM brain that comes up with functional and creative solutions. The harness is just the full set of local tools the LLM can use to complete work on your local computer. A smarter LLM will engineer and implement better solution extensions, tools, skills, agents.md files, etc.

To ease your work load, https://pi.dev/packages is a free repository of shared skill packages with tools and extensions, which users have created to enable common agentic capabilities (computer user, browser control, subagent use, coding and design/styling capabilities, etc.). The packages are easy to install and remove, and you can disable/enable them at any point with:

pi config

Some harness makers try to pack every possible feature they can imagine, into their harness software. Those bundles may include many megabytes of tools and skills which come prepackaged with the harness app (browser control, computer control, subagent spawning skills, tools to communicate via text messaging systems, etc.). The Hermes agent and Openclaw, for example, contain skills, tools, and extensions which include many of the same sorts of features found in Pi's package repository. In fact, Openclaw is built directly on top of Pi - it just includes a pile of useful extensions and skills by default.

The new Deepseek harness is intended to make every feature of the harness software - including absolutely everything about how it operates internally, and how it interacts with LLMs, and everything in its working environment - all modular & swappable. That's a beautifully functional conceptual structure, which could potentially provide a more effective way to organize everything that Pi enables you to do, and what some harnesses messily and heavily provide by default, in an even more organized and efficient way.

You could also choose to go in a completely different direction, in terms of how you prefer to use LLMs, and build your own custom harness application from scratch, with just a few hundred lines of code that provide a way to connect with an LLM API, read & write files, issue OS commands, and send some default prompt text about the environment and how to use those tools, and work in a conversational loop which sends the entire previous conversation text along with each user prompt (and, for example, automatically summarize/compact the current conversation when the attached model's context size limit is reached). You can do that, because it's the LLM that has the capability - that smart brain just needs a little program to enable it to interact with your computing resources.

In fact, a tiny self-made harness application would be all that's required to make a fully functional LLM agent. The rest would be a matter of organizational and operational decisions: how do you choose to formalize a system of reading agent and skill files, calling tools, spawning sub-agent sessions (launching separate harness instances to run useful sub-conversations for tool calls that would otherwise fill up the human conversation context), etc. You could do all those things using any ad-hoc solution you and/or your LLM comes up with, to complete any task.

All the well known harness applications represent choices which their developers have made, to better engineer how each 'lifecycle' of a single LLM prompt can connect with all the other lifecycles that are involved in getting some task or long set of tasks completed.

A smarter LLM will generally engineer better tools, make better plans, author better skill instructions, etc., but it will never be able to do all that work on its own, if it doesn't have an effective set of agentic tools and an organized environment which enables its big brain to get work completed and to communicate efficiently with its future self (in future prompt 'lives'), about how to complete long-horizon tasks.

The creators of every harness app are just trying to engineer better standardized ways to organize all the potential tooling and interactions that an LLM may want to have with your local environment - but you don't have to rely on the preferences of a harness developer. You can just as easily choose to build your own custom solutions for your own custom workflows - and that's generally much easier to do now that the basics of a harness that give a very smart LLM access to complete work on your computer, already exist. In this way, really any harness can be bent to perform any workflow - you and your chosen LLM may just be required to build more pieces of tooling, and establish workflow instructions/practices to get the work completed. The most important pieces of the harness are the default abilities it provides to read & write files, run OS commands, get some basic context about how to interact with the environment, and where to find more information about work that has been completed in previous sessions & tools/skills that exist from previous work.

The smarter an LLM model you use, the more you can rely on it to enable any solution you imagine. The more effective and useful a harness, the tools, skills, and organized workflow patterns you give it to work on your local system, the more effectively it will have interfaces needed to achieve its goals.

Most harnesses will let you switch models in the middle of a conversation. When you do that, all that happens is, the entire context gets sent to that new brain, and that LLM goes through one complete 'lifecycle' with that entire input context. Dropping in a smarter brain can very often fix any problems, when a lesser LLM has troubles working out a task, and then afterward, you can go back to using the less expensive (and perhaps faster performing) LLM for more routine work.

Nick Antonaccio
Nick AntonaccioAdmin
Aug 28, 2026 at 11:34 (edited, 6 revisions)
#2

Let's jump back one step further, to better understand how the LLMs which operate in harness software, actually work:

The brains behind AI products such as ChatGPT, Claude, and Gemini were created by training a statistical algorithm program to predict the next most likely word, which would most likely be expected to follow after any given series of input words. The results of this training routine are stored in billions/trillions of 'parameters', which are like little knobs which collectively adjust how any given input will be transformed into a selected output. This training routine basically finds patterns in gigantic collections of data, which are tuned to be more efficient/correct throughout a series of 'backpropagation' iterations. The algorithm uses repeated runs through the data to adjust parameter settings, based on a 'loss' evaluation which checks how well input values have produced predicted output values. This iteration process involves so much data and so many computed evaluations, that it would take millions of years complete on a normal desktop computer.

Companies such as OpenAI, Anthropic, Deepseek, Google, and others, spend many millions of dollars running this astronomically long statistical pre-training process, to build 'base' LLM models which are surprisingly capable of producing a next word which is actually meaningful within the full context of all the words that came before.

By 'meaningful' I mean, for example, that a base model will not just provide the response 'the butler did it' at the end of an input murder mystery text, because that's perhaps the statistically most common set of words, but instead will provide an answer such as 'the professor did it', in a case where that response makes more sense, given the meaningful arrangement of related words in a particular murder mystery text that come before the predicted output.

The output provided by the LLM represents a kind of understanding about how all the input words are related to one another, based upon patterns found in the trillions of curated words in its training text. The parameters in an LLM model are used to compute how every word relates to every other word in an input text, and how those deeply complex multi-dimensional relationships influence the likelihood of a given next word, based on the massively complex set of patterns stored in its parameters. The knowledge extracted from patterns in an LLMs training text is applied to a user's input text, to generate (or 'infer') a stream of statistically likely correct output text.

Pretrained base models which simply provide a stream of likely next text tokens, however, aren't really good at answering questions. Instead, they just continue the text which would be best expected after any given input text, by reflecting on all of the words that exist in an input submission, in the order/context they appear.

To turn a base LLM into a actually useful thinking machine that can help users solve problems, additional stages of 'post-training' are required. In the post-training phases, the algorithm program is basically fed millions of question and answer pairs, so that the model parameters learn the general shapes of question and answer inputs, and begin to naturally provide answers, whenever inputs in the shape of questions are submitted.

This tendency to produce answer responses to questions, after seeing many provided question-answer example texts, is simply an extension of how the training algorithm learns from patterns in input data.

Without post-training, if a base LLM saw the question 'The capital of France is ___", for example, repeatedly in quiz texts in training data, it might be more likely to simply output the rest of the surrounding quiz questions, instead of an answer, because it doesn't know to do anything else except continue text with the next most likely words it has seen in training data.

Post-training question and answer pairs traditionally are groomed by the scientists who create an LLM. Not only do these training steps make the LLM recognize the 'shape' of question and answer input patterns, they also train the model to answer with particular flavors, and with preferred sorts of responses, when a certain type of input is provided by the user.

Researchers provide question and answer pairs which train the model to provide safe responses, and responses which reflect a responsibly crafted 'constitution' - a general character/nature which the scientists intend the model to exhibit. They provide millions/billions of examples of that behavior, within the post-training text corpora.

Post-training also often involves automated 'reinforcement' learning stages, where the models learn to simply try millions/billions of trial and error self-play iterations, to solve problems. These iterations are rewarded positively when they find the correct solutions to math, programming, and other tasks which have verifiable, correct answers. This sort of self-play enables the models to learn in ways which go far beyond the data that humans prepare manually. It can in fact lead to capabilities which surpass human ability, because the models are able to learn from patterns which evolve through absolutely enormous volumes of trial and error tests (many more iterations that any human, or even large groups of humans, could perform in their lifetimes).

Post-training routines come as close to 'programming' a model to respond in a prescribed way, as is currently possible. This sort of training, based on providing a corpora of repetitive pattern shapes, however, still isn't the same as deterministic 'programming', in the way that traditional software is specifically formed from hard algorithmic rules defined by human engineers.

LLM parameter settings instead just 'emerge' stochastically from training, based on statistical pattern matching processes - instead of being built from set rules programmed by humans.

For that reason, we often say that LLMs are grown, rather than built. Emergent 'intelligent' capabilities simply come from the way AI models learn to respond, according to patterns that they find in training data. These patterns can be incredibly deep and meaningful, but they are still just predicted output based on enormously vast collections of carefully groomed input data.

So, a fully trained model simply accepts input data, and predicts the best most likely output data, based on its training. It does not remember anything about previous questions which have been entered by a user. This is a critically important concept to understand.

The 'emergent' capabilities which LLMs learn from their training data can be truly amazing to comprehend. They learn to translate languages, complete mathematical problems, and write perfectly functioning new software code, for example, without being programmed via any human-specified deterministic rules. As models grow in parameter count, and are trained on larger data sets, they tend to grow more complex and useful emergent capabilities.

No one knows exactly why or how all of a model's emergent capabilities are formed, because no human has ever written code to make them develop - they just appear within the output of a machine made to produce the next most likely token (word, character, etc.), following a sequence of input token. It's been keenly noted that useful emergent capabilities tend to grow with scale, so we just keep building bigger LLMs, with more parameters trained on more data.

And that's where the ability of models stop, and the agent harness applications pick up.

The chatbots which everyone got to know at the beginning of the LLM era (ChatGPT then Claude, Gemini,...) were simply pieces of software which enabled users to type in a prompt which was sent to an LLM, and after the model responded, it enabled the user to type in a new prompt, and that entire conversation history was sent back to the LLM, to provide another response. That loop just continued repeatedly. At the most basic level, chatbots can be very simple pieces of software, which just enable that loop.

The earliest agent harnesses were like chat loops which simply added the ability to work with files, operating-system commands, and other tools. They were programs which sent a layer of instructions to the LLM about how to format a tool call, and then executed that tool call whenever the model returned a properly formed tool call response (write (file) (content), for example).

Modern harnesses build much more orchestration around that basic idea, and the industry is working to expand those capabilities in many ways, but that's the basis of how agents work.

We have extremely intelligent models which, because of emergent capabilities born from enormous scale training routines, can reason though contexts of 1 million+ tokens (tokens are approximately 3/4 of a word), write code, and understand how to solve complex conceptual challenges - and we have many harness applications which give those brains some mix of abilities to work with files and the operating system where the harness runs. They provide well established ways of calling software tools, saving and loading instructions, and remembering information which needs to be recalled across prompt sessions.

Those are the basic pieces of every big chat and agentic system you've seen. Models cost millions of dollars to train, and require massive GPU computing power to produce, but you can build your own agent software with as little as a few hundred lines of code.

Learning how to interact with LLMs, so that they have all the tools required to respond to a prompt, with useful tool call output, generated code, and plans to work across many prompt iterations, is what makes it possible to solve very complex problems with LLMs - much more than can be accomplished in a chat loop.

A harness application is required to give the LLM agency to work with a surrounding operating environment - but you're still always reliant on the intelligence of the model to come up with intelligently reasoned responses, code, tool calls, documentation, etc., to get a job completed. You always need a smart enough brain to provide any sort of useful output. The harness just makes it possible to put that output to work.

Nick Antonaccio
Nick AntonaccioAdmin
Aug 28, 2026 at 12:15 (edited, 9 revisions)
#3

I should point out that the explanations above are intentionally simplified, to be easily digestible. There are many details which blur the lines of how complex systems work in practice. For example, tool calls (read, write, etc.) are commonly represented using structured data, often JSON or a JSON-like schema, such as:

{
  "tool": "write_file",
  "arguments": {
    "path": "foo.py",
    "content": "..."
  }
}

JSON-like structures are better suited to represent deterministic tool call output, as opposed to relying on the harness to scrape loosely structured textual commands.

The concept to grok is that harnesses receive operation requests from the LLM output, and execute those structured requests. The important part to understand is that the harness is a deterministic software layer which gives an LLM controlled ways to interact with resources outside the model itself, such as files, operating-system commands, browsers, APIs, databases, and other tools.

Also, the concept of post-training was kept intentionally simple above (question-answer pairs and a little about reinforcement learning is easy enough to make sense). In reality, it can include combinations of supervised fine-tuning, preference optimization, reinforcement learning, rejection sampling, model-generated synthetic data, verifiable-reward training, tool-use training, reasoning-trace-related training methods, distillation, adversarial/safety training, and other approaches that would only pollute the main instructional concept here.

And of course, LLMs process tokenized representations of words (not words themselves), but the intent in the text above was for the verbiage to make sense conceptually. Use the tutorial above as a conceptual jumping off point, then perhaps copy/paste it into ChatGPT and ask for more details.

For more info and technically detailed instructional posts, see the tutorials category of this forum:

https://aibynick.com/category/8

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they're implemented in each of the well known harnesses.

Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness. No one harness or model is a magic pill - all the harnesses just represent different opinionated approaches to achieving the goal of enabling LLMs to interact with local hardware, saved data, and resources - all within the context limits of each stitched together independent 'lifetime' experience of an inferred response to a single input text.

Most users prefer to use an agent that has many tools, skills, and plan routines built in & automated, but if you understand how these capabilities are implemented, you can edit configuration files, working patterns, and even the code of any agent, to make it work in ways you prefer.

It tends to be easier to do all these things than you might initially expect, because you can simply ask the LLM to edit harness configuration files, write software tools & harness extension code, write skill instruction files based on successfully implemented workflows, write well reasoned plans, save memories to agents.md files & to database records, etc. Understanding how all those pieces work together - which really all just comes down to saving information, code and prompts to files, which can be read and used entirely within a single session - is what makes it possible to use any harness to do effective work. One of the most important things to get to know intuitively, is exactly what a chosen LLM can accomplish, what verbiage works to get it to achieve a desired goal, and how to leverage it to adjust everything about your harness configuration and workflow routines. The more you know about how your operating system, programming language, network, and other underlying computing tools work, the better you'll be able to instruct you AI system to accomplish a given goal.

Please login to post a reply.

© 2026 AI By Nick.