
Agentic AI
Design autonomous AI agents. Tool use, planning, multi-agent cooperation, orchestration frameworks, memory systems.

After this course, you'll be able to:
Who this is for
- →Developers who have shipped an LLM feature and are now being asked for something that takes actions rather than producing text.
- →Engineers evaluating LangGraph, CrewAI, AutoGen or a hand-rolled loop who want to understand the underlying mechanics before committing.
- →Automation and platform engineers replacing brittle rule-based workflows with something that can handle variation.
- →Anyone who built an agent from a tutorial, watched it loop forever or call the wrong tool, and wants to know what the tutorial left out.
- →Skip it if your task is deterministic and well specified. A script or a state machine will be cheaper, faster and correct.
- →Skip it if you have never made a model call with a tool definition attached. Start with function calling first.
What you need first
- ·Solid Python or TypeScript, including async execution, since agents spend most of their time waiting on I/O.
- ·You have called a chat model with a tool or function definition and handled the resulting call, even once.
- ·Comfort with HTTP APIs, JSON Schema and reading API documentation, because tools are mostly wrappers over other systems.
- ·Some experience debugging systems whose output varies between runs. Determinism is the assumption you will have to give up first.
- ·No reinforcement learning background is needed. The planning material here is prompting and control flow, not policy optimization.
What an agent is once you remove the marketing
Strip away the terminology and an agent is a loop. A model receives a goal and the current state. It decides either to produce a final answer or to call a tool. If it calls a tool, the tool runs, its output is appended to the conversation, and the model is invoked again with the enlarged history. Repeat until the model stops or a limit is hit. That is the entire mechanism. Everything else in this field is a variation on what goes into the loop and how it is constrained.
The distinction that clarifies most architectural arguments is between a workflow and an agent. In a workflow, you write the control flow. Step one calls the model to extract fields, step two validates them, step three writes to the database, step four sends a notification. The model does cognitive work at fixed points, but the sequence is yours. In an agent, the model decides the sequence at runtime by choosing which tool to call next.
That difference is the whole trade. Workflows are predictable, testable, cheap and easy to debug. Agents handle variation you could not enumerate in advance. Almost every production system that works well is mostly workflow with agentic behaviour confined to the specific steps that genuinely need it. The common expensive mistake is granting the model control over the entire process when only one branch was actually unpredictable.
This also explains why the ambitious autonomous agent demos from a few years back, the ones given a vague goal and left to run, tended to disappoint. The failure was not intelligence but error compounding. If each step in a chain has a modest independent chance of going wrong, a twenty-step chain is unlikely to reach the end intact, and there is no recovery mechanism because nothing in the loop is checking. Reliability at each step matters more than capability, and the way you get reliability is by shortening chains, verifying intermediate results, and giving the loop somewhere to stop.
The useful mental model is not an autonomous colleague. It is a control loop with a nondeterministic decision function, which happens to be the description of a system that needs guardrails, timeouts, verification and observability. Building agents well is mostly systems engineering wearing an AI hat.
Tools are an interface designed for a reader who cannot ask questions
Tool use is where agents meet reality, and the quality of an agent is largely the quality of its tool definitions. The model sees only a name, a description and a JSON Schema. It cannot read your source, cannot ask a colleague what the third parameter means, and will not check the wiki. Everything it needs must be in the specification.
That reframes tool design as documentation writing with unusually high stakes. A few principles hold up consistently.
Design tools around tasks, not around your API surface. Wrapping forty REST endpoints one-for-one produces forty similar-looking options and a model that picks wrong. A single tool named search_customer_orders with a clear purpose beats five thin CRUD wrappers the model has to compose correctly.
Make descriptions say when not to use the tool. Overlapping tools are the most common source of wrong calls. If two tools could plausibly apply, each description should say what distinguishes it from the other.
Constrain the schema aggressively. Enums instead of free strings, required fields marked required, formats specified, sensible defaults supplied. Every degree of freedom in a schema is a way for the model to produce something your code has to handle.
Return errors the model can act on. "400 Bad Request" teaches it nothing and it will retry identically. "No customer found with that email; try search_customers by name first" tells it what to do next. Error messages are prompts, and treating them as such fixes a surprising number of loops.
Budget output size. A tool that dumps an entire result set as raw JSON will consume the context window, push earlier reasoning out of view, and cost real money on every subsequent turn. Paginate, summarize, truncate with an explicit marker, or return an identifier the model can use to fetch detail selectively.
Decide what a tool may do without permission. Reads are usually safe. Writes, deletions, payments, external messages and anything irreversible should require confirmation, run against a scoped credential, or be staged for review. The OWASP Top 10 for LLM Applications lists excessive agency as a distinct risk category for exactly this reason: the damage from an agent is bounded by what its tools can reach.
There is also a scaling limit worth reasoning about. Every tool definition occupies context, and the model must discriminate between options that grow more similar as the catalogue grows — two dozen tools means two dozen descriptions competing for attention, several of which plausibly fit any given step. Once selection starts going wrong, the fix is usually architectural rather than a better prompt: route to a relevant subset before the model sees the list, or split responsibilities across specialized agents that each hold a small toolset.
Planning and reasoning: what helps and what is theatre
A pile of named techniques sits between "call the model" and "the agent solves the task". Some of them earn their keep. Others mostly produce impressive-looking traces.
ReAct, from a paper by Yao and colleagues at Princeton and Google, interleaves reasoning with acting: the model articulates a thought, takes an action, observes the result, and reasons again. This is the default loop most frameworks implement. Its strength is that intermediate reasoning is visible, which makes debugging tractable. Its weakness is that the model narrating a plausible thought is not evidence the plan is sound.
Plan-and-execute separates the two phases. The model drafts a full plan first, then executes the steps. This reduces drift on long tasks because there is a stated goal to check against, and it produces an artefact a human can approve before anything runs. The cost is rigidity when reality diverges from the plan, which is why practical implementations include a replanning step.
Reflexion and self-critique add a review pass where the model evaluates its own output and retries. This helps when there is something concrete to check against, such as a failing test, a schema validation error, or a compiler message. It helps far less when the model is grading its own prose with no external signal, because the same weaknesses that produced the answer tend to produce the review.
Tree of thoughts and related search methods explore several reasoning branches and select among them. They can improve results on problems with a clear evaluation function and a genuine search space. On most business tasks they multiply cost for little benefit.
The pattern across all of these is that external verification is what makes reasoning reliable. An agent that writes code and runs the test suite has ground truth. An agent that writes code and asks itself whether the code looks right has an opinion. Wherever you can put a real check into the loop, a compiler, a validator, a database constraint, a unit test, a second system that must agree, do that before reaching for a more elaborate prompting scheme.
Newer reasoning-trained models have absorbed some of this into their own generation, which shifts the balance. Elaborate prompt scaffolding designed for older models sometimes hurts on models that already reason internally. That is a good argument for keeping the scaffolding thin and testable rather than architectural.
Memory, state, and the context budget
Everything an agent knows in a given call is what is in that call. The conversation is reconstructed and resent every turn, so memory is not a feature the model has, it is a system you build.
The immediate problem is growth. A loop that appends every thought, every tool call and every tool result will fill the context window on a moderately long task, and cost grows with it because you are paying for the whole history on every step. Long before you hit a hard limit, quality degrades: earlier instructions get buried, and models attend unevenly across very long prompts.
The techniques for managing this are straightforward but need deliberate design.
Trimming drops old turns. Simple, and it silently discards things the agent needed. Summarization compresses earlier history into a running digest, preserving more at the cost of an extra model call and some lossiness at the boundary. Scratchpads move state out of the conversation entirely: the agent writes findings to a file or a structured record, and the context holds a pointer instead of the content. Externalized retrieval stores past interactions and pulls back only what is relevant to the current step, which is retrieval-augmented generation applied to the agent's own history.
Then there is memory across sessions, which is a different problem. Users reasonably expect an assistant to remember stated preferences, prior decisions and established facts. Implementing that means deciding what is worth persisting, how it is retrieved later, how conflicting statements are reconciled, and how a user corrects or deletes something. Libraries exist for this, but the hard parts are policy questions rather than storage questions. Persisted memory is also a privacy surface with retention obligations, which is worth confronting before it is full of personal data.
One more distinction matters in production: the agent's memory and the application's state are not the same thing. A booking, an order, a case record and a permission grant belong in your database with transactions and audit trails. They should not live only in a conversation history. Agents are unreliable narrators of what they did; the system of record needs to be a system of record.
Multi-agent systems and the price of coordination
Splitting work across several agents is intuitively appealing. A researcher, a writer, a critic. A planner and a set of workers. Frameworks make it easy to express, which is part of why it is reached for early.
The honest position is that multi-agent architectures solve two real problems and create several new ones.
The genuine wins are context isolation and tool scoping. If one agent needs forty tools and a large volume of domain instruction while another needs three tools and a style guide, giving them separate contexts keeps each prompt focused and each toolset small enough for reliable selection. Parallelism is the second win: independent subtasks, such as researching five candidate suppliers, can run concurrently and cut wall-clock time substantially.
The costs are less discussed. Every handoff is a lossy serialization, because one agent's understanding must be compressed into a message. Errors propagate and are hard to attribute, since a bad final answer might originate three agents upstream. Cost multiplies, as each agent carries its own system prompt and history. Debugging gets considerably harder, because you are now reading interleaved traces from several nondeterministic processes. And coordination overhead is real: agents can deadlock waiting on each other, duplicate work, or converge on a shared mistake because they were all initialized from the same framing.
Several architectural patterns manage this. A supervisor or orchestrator holds control and delegates to specialists, which keeps the flow legible and gives you one place to enforce limits. A pipeline hands output forward through fixed stages, which is really a workflow with agentic steps and is often the right answer. Blackboard patterns have agents read and write a shared state store rather than message each other, which reduces coupling. Fully peer-to-peer conversation among agents is the most flexible and the least predictable.
Adversarial and competitive configurations are a distinct case with legitimate uses: a generator paired with a critic that is instructed to find fault, or a proposer and a verifier with different objectives. The value comes from the asymmetry, not from having two agents. A critic with the same prompt and the same information as the generator mostly agrees with it.
The default should be one capable agent with well-designed tools. Split when you can name the specific problem the split solves, and expect to pay for it in complexity.
What breaks in production
The gap between an agent that demos well and one that runs unattended is mostly operational.
Loops and runaway cost. An agent that cannot make progress will often keep trying. Without a hard cap on iterations, wall-clock time and token spend, a single stuck run can produce an alarming bill. Every loop needs limits, and hitting a limit should be a logged, alertable event rather than a silent truncation.
Nondeterminism defeats ordinary testing. The same input can produce different tool sequences on different runs, so assertion-based tests over exact outputs are useless. What works instead is a fixed suite of realistic tasks scored on whether the end state is correct, run repeatedly, with pass rates tracked over time. Public benchmarks illustrate the shape of this: SWE-bench from Princeton scores whether a patch makes real tests pass, and tau-bench from Sierra scores whether an agent reaches a correct final state under policy constraints. Your own version of that, built from your own tasks, is what tells you whether a prompt change was an improvement.
Observability is not optional. You need the full trace of every run: prompts, tool calls, arguments, results, timings and token counts. Without it, debugging is archaeology. Structured tracing should be built in from the first version, not added after the first incident.
Tool output is untrusted input. This is the security issue people underestimate. An agent that reads a web page, an email, a ticket comment or a document is ingesting text that someone else wrote, and that text sits in the same context as your instructions. Injected instructions in retrieved content can redirect an agent's behaviour. The exposure is highest when an agent has access to private data, ingests untrusted content, and can send information outward, a combination Simon Willison usefully named the lethal trifecta. What makes this hard is that it is a property of the loop, not of any one call: no individual step looks dangerous, and the capability set that creates the risk is assembled gradually as someone adds one more useful tool. The defences are correspondingly structural. Decide the capability set per run rather than per agent, so a run triggered by attacker-influenced content never holds an outbound tool at the same time as a privileged read. Keep untrusted material out of the turn that makes decisions, summarising or extracting it in a separate call whose output is data rather than instruction. Fix the plan before execution where the task allows it, so that content encountered mid-run cannot rewrite the goal. And put the irreversible steps behind a human, since an approval gate is the one control that does not depend on the model behaving.
Partial failure needs a story. Real agents fail halfway through. If three of five steps completed and two of them wrote to external systems, what happens? Idempotent tools, explicit compensation steps and resumable runs are the difference between a recoverable incident and manual cleanup.
Humans need a way in. The most reliable production designs are not fully autonomous. They pause at consequential steps, present what they intend to do, and act on approval. This is a design feature, not an admission of weakness, and it is usually what makes an agent deployable in a context where mistakes are expensive.
Frameworks, and whether you need one
The framework question comes up early and consumes more debate than it deserves.
Building the loop yourself is genuinely simple: send messages, parse tool calls, execute, append, repeat. A first working version is short enough to read in one sitting. The advantage is that you understand every part of it, you can put a breakpoint anywhere, and there is no abstraction between you and the model API. For a single agent with a handful of tools, this is frequently the right choice and the fastest path to something you can debug.
Frameworks earn their place when you need what they provide beyond the loop. Graph-based orchestrators such as LangGraph model an agent as an explicit state machine, which gives you persistence, checkpointing, resumability and human approval gates as first-class features rather than things you invent. Role-oriented frameworks such as CrewAI make multi-agent collaboration quick to express. Conversation-oriented frameworks such as AutoGen from Microsoft Research focus on multiple agents exchanging messages. Lighter typed libraries add structured output validation and dependency injection without taking over control flow.
What matters more than the choice is understanding what sits underneath it. Frameworks hide the prompt, and the prompt is where agent behaviour lives. When an agent misbehaves and you cannot see what was actually sent to the model, you cannot fix it. Whatever you adopt, learn how to dump the raw request.
Model choice is the other axis. Tool use is a trained capability and models differ substantially in how reliably they emit well-formed calls, respect schemas, and know when to stop. Open-weight models that run locally have improved considerably here, and a local model is attractive when tool arguments contain data you would rather not send anywhere. The practical approach is to test candidate models against your own task suite rather than a leaderboard, because tool-calling reliability on your specific schemas is what determines whether the system works.
Finally, tool interfaces are standardizing. The Model Context Protocol, published by Anthropic as an open specification, defines a common way for servers to expose tools and resources to any compatible client. Building integrations against a shared protocol rather than a framework's own abstraction means they survive a change of framework, which over a few years is a reasonable bet.
Common questions
What is the difference between an AI agent and a workflow?
In a workflow, you decide the sequence of steps and the model does bounded work inside them. In an agent, the model decides at runtime which tool to call next and when to stop. Workflows are cheaper, more predictable and far easier to test; agents handle variation you could not enumerate in advance. Most systems that work well in production are predominantly workflow, with agentic control confined to the specific steps where the path genuinely cannot be known ahead of time.
Can I run agents on a local open-weight model?
Yes, and it is increasingly practical. The capability that matters is reliable tool calling: emitting well-formed arguments against your schemas, choosing correctly among similar tools, and knowing when to stop. Models vary widely here and it does not track general chat quality, so test candidates against your own tool definitions. Local models are particularly attractive when tool arguments carry data you do not want leaving your network, since arguments and results both pass through the model context.
How do I stop an agent looping forever and burning tokens?
Hard limits first: maximum iterations, maximum wall-clock time, maximum token spend per run, enforced in your code rather than requested in the prompt. Then attack the cause. Most infinite loops come from a tool returning an unhelpful error the model cannot act on, so it retries the same call. Rewriting error messages to state what went wrong and what to try next resolves a large share of them. Detecting repeated identical calls and forcing a different branch handles the rest.
How do you test something that behaves differently every run?
Stop asserting on exact output and start scoring end states. Build a suite of realistic tasks with checkable success conditions, run each several times, and track the pass rate as your quality metric. That gives you a number that moves when a prompt, model or tool changes. Public agent benchmarks work this way, scoring whether tests pass or whether the final system state is correct. Alongside that, record complete traces of every run so a failure can be diagnosed rather than guessed at.
Is a multi-agent system better than one agent with more tools?
Usually not, and the default should be one agent. Multiple agents genuinely help in two situations: when contexts and toolsets are large enough that isolating them improves reliability, and when subtasks are independent enough to run in parallel for real time savings. Otherwise you pay in lossy handoffs, harder debugging, multiplied cost and coordination failures. Split only when you can articulate the specific problem the split solves.
What is the biggest security risk with agents?
Anything an agent reads is untrusted input sitting in the same context as your instructions, so retrieved web pages, emails, tickets and documents can carry injected instructions. Risk peaks when a single run combines access to private data, exposure to untrusted content, and some way of sending information outward. Note that this is a property of the combination, which is why it tends to appear by accident as tools accumulate. Auditing the capability set of each run, rather than each tool in isolation, is the check that catches it. Beyond that, the defences worth having are the ones that do not rely on the model cooperating: a fixed plan the run cannot renegotiate, untrusted content processed in a call that produces data rather than direction, and a person in front of anything irreversible.
Related reading
AI agent frameworks compared
Practical comparison of the orchestration options discussed in the frameworks section.
Build a local AI agent
A worked implementation of the agent loop from scratch, with no framework in the way.
Ollama function calling and tools
How tool definitions are actually wired up when the model is running on your own hardware.
Best local models for agents
Which open-weight models handle tool selection and multi-step control reliably enough to be worth testing.
Agent memory with Mem0
One concrete approach to the cross-session memory problem described in the memory section.
Prompt injection defence
The security material behind the untrusted tool output problem, with specific mitigations.
Full syllabus
Tool Use
Planning & Reasoning
Communication Protocols
Cooperation Patterns
Competitive & Adversarial
Orchestration Frameworks
Memory & Knowledge
Learning & Adaptation
Production Deployment
Unlock all 10 chapters
Plus 24 other courses — 551 more chapters included.
Every course, every future course, the Python Lab and eight downloadable kits, nothing to renew. Or subscribe: Pro $8.99/month