Post History

Current version by Nick Antonaccio

Current VersionAug 28, 2026 at 12:15

I should point out that the explanations above are intentionally simplified, to be easily digestible. There are many details which blur the lines of how complex systems work in practice. For example, tool calls (read, write, etc.) are commonly represented using structured data, often JSON or a JSON-like schema, such as:

{
  "tool": "write_file",
  "arguments": {
    "path": "foo.py",
    "content": "..."
  }
}

JSON-like structures are better suited to represent deterministic tool call output, as opposed to relying on the harness to scrape loosely structured textual commands.

The concept to grok is that harnesses receive operation requests from the LLM output, and execute those structured requests. The important part to understand is that the harness is a deterministic software layer which gives an LLM controlled ways to interact with resources outside the model itself, such as files, operating-system commands, browsers, APIs, databases, and other tools.

Also, the concept of post-training was kept intentionally simple above (question-answer pairs and a little about reinforcement learning is easy enough to make sense). In reality, it can include combinations of supervised fine-tuning, preference optimization, reinforcement learning, rejection sampling, model-generated synthetic data, verifiable-reward training, tool-use training, reasoning-trace-related training methods, distillation, adversarial/safety training, and other approaches that would only pollute the main instructional concept here.

And of course, LLMs process tokenized representations of words (not words themselves), but the intent in the text above was for the verbiage to make sense conceptually. Use the tutorial above as a conceptual jumping off point, then perhaps copy/paste it into ChatGPT and ask for more details.

For more info and technically detailed instructional posts, see the tutorials category of this forum:

https://aibynick.com/category/8

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they're implemented in each of the well known harnesses.

Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness. No one harness or model is a magic pill - all the harnesses just represent different opinionated approaches to achieving the goal of enabling LLMs to interact with local hardware, saved data, and resources - all within the context limits of each stitched together independent 'lifetime' experience of an inferred response to a single input text.

Most users prefer to use an agent that has many tools, skills, and plan routines built in & automated, but if you understand how these capabilities are implemented, you can edit configuration files, working patterns, and even the code of any agent, to make it work in ways you prefer.

It tends to be easier to do all these things than you might initially expect, because you can simply ask the LLM to edit harness configuration files, write software tools & harness extension code, write skill instruction files based on successfully implemented workflows, write well reasoned plans, save memories to agents.md files & to database records, etc. Understanding how all those pieces work together - which really all just comes down to saving information, code and prompts to files, which can be read and used entirely within a single session - is what makes it possible to use any harness to do effective work. One of the most important things to get to know intuitively, is exactly what a chosen LLM can accomplish, what verbiage works to get it to achieve a desired goal, and how to leverage it to adjust everything about your harness configuration and workflow routines. The more you know about how your operating system, programming language, network, and other underlying computing tools work, the better you'll be able to instruct you AI system to accomplish a given goal.

Previous Versions
Version 9Aug 28, 2026 at 12:15

I should point out that the explanations above are intentionally simplified, to be easily digestible. There are many details which blur the lines of how complex systems work in practice. For example, tool calls (read, write, etc.) are commonly represented using structured data, often JSON or a JSON-like schema, such as:

{
  "tool": "write_file",
  "arguments": {
    "path": "foo.py",
    "content": "..."
  }
}

JSON-like structures are better suited to represent deterministic tool call output, as opposed to relying on the harness to scrape loosely structured textual commands.

The concept to grok is that harnesses receive operation requests from the LLM output, and execute those structured requests. The important part to understand is that the harness is a deterministic software layer which gives an LLM controlled ways to interact with resources outside the model itself, such as files, operating-system commands, browsers, APIs, databases, and other tools.

Also, the concept of post-training was kept intentionally simple above (question-answer pairs and a little about reinforcement learning is easy enough to make sense). In reality, it can include combinations of supervised fine-tuning, preference optimization, reinforcement learning, rejection sampling, model-generated synthetic data, verifiable-reward training, tool-use training, reasoning-trace-related training methods, distillation, adversarial/safety training, and other approaches that would only pollute the main instructional concept here.

And of course, LLMs process tokenized representations of words (not words themselves), but the intent in the text above was for the verbiage to make sense conceptually. Use the tutorial above as a conceptual jumping off point, then perhaps copy/paste it into ChatGPT and ask for more details.

For more info and technically detailed instructional posts, see the tutorials category of this forum:

https://aibynick.com/category/8

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they're implemented in each of the well known harnesses.

Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness. No one harness or model is a magic pill - all the harnesses just represent different opinionated approaches to achieving the goal of enabling LLMs to interact with local hardware, saved data, and resources - all within the context limits of each stitched together independent 'lifetime' experience of an inferred response to a single input text.

Most users prefer to use an agent that has many tools, skills, and plan routines built in & automated, but if you understand how these capabilities are implemented, you can edit configuration files, working patterns, and even the code of any agent, to make it work in ways you prefer.

It tends to be easier to do all these things than you might initially expect, because you can simply ask the LLM to edit harness configuration files, write software tools & harness extension code, write skill instruction files based on successfully implemented workflows, write well reasoned plans, save memories to agents.md files & to database records, etc. Understanding how all those pieces work together - which really all just comes down to saving information, code and prompts to files, which can be read and used entirely within a single session - is what makes it possible to use any harness to do effective work. One of the most important things to get to know intuitively, is exactly what a chosen LLM can accomplish, what verbiage works to get it to achieve a desired goal, and how to leverage it to adjust everything about your harness configuration and workflow routines. The more you know about how your operating system, programming language, network, and other underlying tools work, the better you'll be able to instruct you AI system to accomplish a given goal.

Version 8Aug 28, 2026 at 12:15

I should point out that the explanations above are intentionally simplified, to be easily digestible. There are many details which blur the lines of how complex systems work in practice. For example, tool calls (read, write, etc.) are commonly represented using structured data, often JSON or a JSON-like schema, such as:

{
  "tool": "write_file",
  "arguments": {
    "path": "foo.py",
    "content": "..."
  }
}

JSON-like structures are better suited to represent deterministic tool call output, as opposed to relying on the harness to scrape loosely structured textual commands.

The concept to grok is that harnesses receive operation requests from the LLM output, and execute those structured requests. The important part to understand is that the harness is a deterministic software layer which gives an LLM controlled ways to interact with resources outside the model itself, such as files, operating-system commands, browsers, APIs, databases, and other tools.

Also, the concept of post-training was kept intentionally simple above (question-answer pairs and a little about reinforcement learning is easy enough to make sense). In reality, it can include combinations of supervised fine-tuning, preference optimization, reinforcement learning, rejection sampling, model-generated synthetic data, verifiable-reward training, tool-use training, reasoning-trace-related training methods, distillation, adversarial/safety training, and other approaches that would only pollute the main instructional concept here.

And of course, LLMs process tokenized representations of words (not words themselves), but the intent in the text above was for the verbiage to make sense conceptually. Use the tutorial above as a conceptual jumping off point, then perhaps copy/paste it into ChatGPT and ask for more details.

For more info and technically detailed instructional posts, see the tutorials category of this forum:

https://aibynick.com/category/8

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they're implemented in each of the well known harnesses.

Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness. No one harness or model is a magic pill - all the harnesses just represent different opinionated approaches to achieving the goal of enabling LLMs to interact with local hardware, saved data, and resources - all within the context limits of each stitched together independent 'lifetime' experiences of an inferred response to a single input text.

Most users prefer to use an agent that has all the tools, skills, and plan routines built in & automated, but if you understand how these capabilities are implemented, you can edit the configuration files, working patterns, and even the code of any agent, to make it work in ways you prefer. It tends to be easier than you may expect to do all these things, because you can simply ask the LLM to edit harness configuration files, write software tools & harness extension code, write skill instruction files based on successfully implemented workflows, write well reasoned plans, save memories to agents.md files & to database records, etc. Understanding how all those pieces work together - which all just comes down to saving information, code and prompts to files, which can be read and used entirely in a single session - is what makes it possible to use any harness to do effective work. One of the most important things to get to know intuitively, is exactly what a chosen LLM can accomplish, what verbiage works to get it to achieve a desired goal, and how to leverage it to adjust everything about you harness configuration and workflow routines.

Version 7Aug 28, 2026 at 12:11

I should point out that the explanations above are intentionally simplified, to be easily digestible. There are many details which blur the lines of how complex systems work in practice. For example, tool calls (read, write, etc.) are commonly represented using structured data, often JSON or a JSON-like schema, such as:

{
  "tool": "write_file",
  "arguments": {
    "path": "foo.py",
    "content": "..."
  }
}

JSON-like structures are better suited to represent deterministic tool call output, as opposed to relying on the harness to scrape loosely structured textual commands.

The concept to grok is that harnesses receive operation requests from the LLM output, and execute those structured requests. The important part to understand is that the harness is a deterministic software layer which gives an LLM controlled ways to interact with resources outside the model itself, such as files, operating-system commands, browsers, APIs, databases, and other tools.

Also, the concept of post-training was kept intentionally simple above (question-answer pairs and a little about reinforcement learning is easy enough to make sense). In reality, it can include combinations of supervised fine-tuning, preference optimization, reinforcement learning, rejection sampling, model-generated synthetic data, verifiable-reward training, tool-use training, reasoning-trace-related training methods, distillation, adversarial/safety training, and other approaches that would only pollute the main instructional concept here.

And of course, LLMs process tokenized representations of words (not words themselves), but the intent in the text above was for the verbiage to make sense conceptually. Use the tutorial above as a conceptual jumping off point, then perhaps copy/paste it into ChatGPT and ask for more details.

For more info and technically detailed instructional posts, see the tutorials category of this forum:

https://aibynick.com/category/8

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they're implemented in each of the well known harnesses.

Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness. No one harness or model is a magic pill - all the harnesses just represent different opinionated approaches to achieving the goal of enabling LLMs to interact with local hardware, saved data, and resources - all within the context limits of each stitched together 'lifetime' experience of an inferred response to a single input text.

Version 6Aug 28, 2026 at 11:46

I should point out that the explanations above are intentionally simplified, to be easily digestible. There are many details which blur the lines of how complex systems work in practice. For example, tool calls (read, write, etc.) are commonly represented using structured data, often JSON or a JSON-like schema, such as:

{
  "tool": "write_file",
  "arguments": {
    "path": "foo.py",
    "content": "..."
  }
}

JSON-like structures are better suited to represent deterministic tool call output, as opposed to relying on the harness to scrape loosely structured textual commands.

The concept to grok is that harnesses receive operation requests from the LLM output, and execute those structured requests. The important part to understand is that the harness is a deterministic software layer which gives an LLM controlled ways to interact with resources outside the model itself, such as files, operating-system commands, browsers, APIs, databases, and other tools.

Also, the concept of post-training was kept intentionally simple above (question-answer pairs and a little about reinforcement learning is easy enough to make sense). In reality, it can include combinations of supervised fine-tuning, preference optimization, reinforcement learning, rejection sampling, model-generated synthetic data, verifiable-reward training, tool-use training, reasoning-trace-related training methods, distillation, adversarial/safety training, and other approaches that would only pollute the main instructional concept here.

And of course, LLMs process tokenized representations of words (not words themselves), but the intent in the text above was for the verbiage to make sense conceptually. Use the tutorial above as a conceptual jumping off point, then perhaps copy/paste it into ChatGPT and ask for more details.

For more info and technically detailed instructional posts, see the tutorials category of this forum:

https://aibynick.com/category/8

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they're implemented in each of the well known harnesses.

Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness. No one harness or model is a magic pill - all the harnesses just represent different opinionated approaches to achieving the goal of enabling LLMs to interact with local hardware, saved data, and resources.

Version 5Aug 28, 2026 at 11:42

I should point out that the explanations above are intentionally simplified, to be easily digestible. There are many details which blur the lines of how complex systems work in practice. For example, tool calls (read, write, etc.) are commonly represented using structured data, often JSON or a JSON-like schema, such as:

{
  "tool": "write_file",
  "arguments": {
    "path": "foo.py",
    "content": "..."
  }
}

rather than relying on the harness to scrape a textual command. The concept to grok is that harnesses receive requested operations and execute them. The important part to understand is that the harness is a software layer which gives an LLM controlled ways to interact with resources outside the model itself, such as files, operating-system commands, browsers, APIs, databases, and other tools..

Also, the concept of post-training was kept intentionally simple above (question-answer pairs and a little about reinforcement learning is easy enough to make sense). In reality, it can include combinations of supervised fine-tuning, preference optimization, reinforcement learning, rejection sampling, model-generated, synthetic data, verifiable-reward training, tool-use training, reasoning-trace-related training methods, distillation, adversarial/safety training, and other approaches that would only pollute the main concept.

And of course, LLMs process tokenized representations of words (not words themselves), but the intent was for the verbiage above to make sense conceptually. Use the tutorial above as a conceptual jumping off point, then perhaps copy/paste it into ChatGPT and ask for more details.

For more info and instructional posts, see the tutorials section:

https://aibynick.com/category/8

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they are implemented in each of the well known harnesses. Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness.

Version 4Aug 23, 2026 at 15:37

I should point out that the explanations above are intentionally simplified, to be easily digestible. There are many details which blur the lines of how complex systems work in practice. For example, tool calls (read, write, etc.) are commonly represented using structured data, often JSON or a JSON-like schema, such as:

{
  "tool": "write_file",
  "arguments": {
    "path": "foo.py",
    "content": "..."
  }
}

rather than relying on the harness to scrape a textual command. The concept to grok is that harnesses receive requested operations and execute them. The important part to understand is that the harness is a software layer which gives an LLM controlled ways to interact with resources outside the model itself, such as files, operating-system commands, browsers, APIs, databases, and other tools..

Also, the concept of post-training was kept intentionally simple above (question-answer pairs and a little about reinforcement learning is easy enough to make sense). In reality, it can include combinations of supervised fine-tuning, preference optimization, reinforcement learning, rejection sampling, model-generated, synthetic data, verifiable-reward training, tool-use training, reasoning-trace-related training methods, distillation, adversarial/safety training, and other approaches that would only pollute the main concept.

And of course, LLMs process tokenized representations of words (not words themselves), but the intent was for the verbiage above to make sense conceptually. Use the tutorial above as a conceptual jumping off point, then perhaps copy/paste it into ChatGPT and ask for more details.

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they are implemented in each of the well known harnesses. Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness.

Version 3Aug 23, 2026 at 14:05

I should point out that the explanations above are intentionally simplified, to be easily digestible. There are many details which blur the lines of how complex systems work in practice. For example, tool calls (read, write, etc.) are commonly represented using structured data, often JSON or a JSON-like schema, such as:

{ "tool": "write_file", "arguments": { "path": "foo.py", "content": "..." } }

rather than relying on the harness to scrape a textual command. The concept to grok is that harnesses receive requested operations and execute them. The important part to understand is that the harness is a software layer which gives an LLM controlled ways to interact with resources outside the model itself, such as files, operating-system commands, browsers, APIs, databases, and other tools..

Also, the concept of post-training was kept intentionally simple above (question-answer pairs and a little about reinforcement learning is easy enough to make sense). In reality, it can include combinations of supervised fine-tuning, preference optimization, reinforcement learning, rejection sampling, model-generated, synthetic data, verifiable-reward training, tool-use training, reasoning-trace-related training methods, distillation, adversarial/safety training, and other approaches that would only pollute the main concept.

Use the tutorial above as a conceptual jumping off point, then perhaps copy/paste it into ChatGPT and ask for more details.

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they are implemented in each of the well known harnesses. Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness.

Version 2Aug 23, 2026 at 13:59

I should point out that the explanations above are intentionally simplified, to be easily digestible. There are many details which blur the lines of how complex systems work in practice. For example, tool calls (read, write, etc.) actually normally make use of structured json, such as:

{ "tool": "write_file", "arguments": { "path": "foo.py", "content": "..." } }

rather than relying on the harness to scrape a textual command. The concept to grok is that harnesses receive requested operations and execute them. The important part to understand is that the harness is a structured application which provides an LLM the capability to interact with your local operating system.

Also, the concept of post-training was kept intentionally simple above (question-answer pairs and a little about reinforcement learning is easy enough to make sense). In reality, it can include combinations of supervised fine-tuning, preference optimization, reinforcement learning, rejection sampling, model-generated, synthetic data, verifiable-reward training, tool-use training, reasoning-trace-related training methods, distillation, adversarial/safety training, and other approaches that would only pollute the main concept.

Use the tutorial above as a conceptual jumping off point, then perhaps copy/paste it into ChatGPT and ask for more details.

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they are implemented in each of the well known harnesses. Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness.

Version 1Aug 23, 2026 at 13:54

I'll continue this thread later by putting some concrete examples together, which demonstrate all those pieces and how they are implemented in each of the well known harnesses. Seeing all the ways people build harness features, and how you can accomplish the same goals with conceptually similar tools, implemented in slightly different ways, is one of the best ways to really grok how to work with LLMs and harnesses. You can, by the way, learn all that on your own, just by working with any capable LLM, from any API provider, in any capable harness.