Looking into why graphs and multi-agent systems are getting attention, I was left with one question: just because AI can make a judgment, does that mean I delegated it?
Today I kept running into posts about graphs and multi-agent systems.
Several AIs split up the work, one picks up where another left off, and when something fails they fix it and test again. A person no longer has to pass each task along by hand.
I wondered why this is getting so much attention, so I looked into it a little.
What a graph actually changes
The core of a graph turned out to be less about using many AIs and more about turning the flow of work into a structure.
For example:
build → test → fix if it fails → test again → verify if it passes
When we work through chat, a person connects these steps. We ask for the build, look at the result, ask for tests, and ask for a fix if there is an error. AI does the actual work, but a person is still managing the flow of the work.
A graph turns that connection itself into a system. The flow a person used to hold in their head becomes visible inside the execution structure as elements like these:
- Nodes — units of execution such as analysis, building, testing, or human approval
- Edges and branches — where a task’s result goes, and which path to take on success or failure
- Loops — repeating fix and re-verify
- Fan-out and fan-in — running independent tasks at the same time and merging the results
- State — execution state such as results, variables, and whether something was approved
- Interrupt, checkpoint, and resume — stopping when a person or outside input is needed, saving state, and continuing from the same point
Deciding what to do next moves from a person’s instructions in a conversation into the execution structure. This also differs from an agent that reads project instructions and sets its own order. Instructions are rules an agent interprets; a graph puts part of the flow and its boundaries into the execution structure itself. The two are not mutually exclusive and can be used together.
Different ways of operating, not stages of progress
There are more dynamic approaches too. A lead agent looks at the situation, splits up the work, calls the specialist agents it needs, and gathers their results. Running the same task, say “improve this online store’s purchase conversion,” through different approaches shows the differences.
| Approach | Who decides the next task | The person’s role | Strength | Difficulty |
|---|---|---|---|---|
| Chat | The person | Directs continuously | Flexible | Management burden |
| Instruction-guided agent | AI interprets instructions and context | Goals and important judgments | Flexibility and automation | Depends on the model’s interpretation |
| Fixed graph | Nodes, edges, and conditions set by the system designer | Exceptions and approvals | Reproducibility and control | Unexpected situations |
| Dynamic multi-agent | A lead AI breaks down work during execution | Goals and oversight | Exploration and parallelism | Tracing, cost, and control |
| Hybrid | AI judges dynamically within fixed boundaries | Boundaries and important choices | Control and flexibility | How to design the boundaries |
It wasn’t just one company’s approach. Google released a graph-based workflow engine, built-in HITL, and dynamic orchestration together in ADK Go 2.0.1 Microsoft Agent Framework offers sequential and concurrent execution, handoffs between agents, group chat, and a manager agent that plans and coordinates, and it can have a person review that plan before execution.2 OpenAI distinguishes between a manager agent calling specialists as tools and handing the conversation off to a specialist.3 Anthropic’s Research feature uses a structure in which a lead agent runs several subagents in parallel in a real product.4
What mattered was that these approaches are not stages of a single line of progress. For some work it is better for a person to define the flow in advance; for other work it is better for AI to read the situation and decide the next step. Real systems mix the two.
More agents aren’t automatically better, either. As agents multiply, so do state, tracing, cost, and approval points. Anthropic reports that its multi-agent research system uses about 15 times the tokens of a chat,4 and OpenAI advises adding specialist agents when you actually need to split the work.3
So I quickly understood why people are interested in this. As AI becomes capable of doing more of the work, the question is shifting from “How do I get a better answer?” to “How do I structure the work so AI can keep going?” The human role isn’t disappearing; it’s moving. Task coordination shifts into the system, while people remain responsible for goals, boundaries, and decisions about new directions.
I realized I was already working this way
When I hand development to AI, I don’t want to check the code line by line.
I want it to build, test, fix what’s broken, verify again on its own, and bring me the result.
Rather than asking my permission at every step, I want it to stop only when my judgment is really needed.
But looking at graphs, one question came up.
Who decides when that “really needed” moment is?
Say I asked it to improve an online store’s purchase conversion, and I approved the overall plan.
While building and testing, the AI makes this judgment:
I think changing the product structure itself would produce better results.
It can do it technically. It also fits the original goal of improving conversion.
So should it just go ahead?
When a test fails, I don’t want to decide how the code should be fixed.
But whether to change the product structure itself is a little different.
To the AI, both may be “choosing the next action,” but to a person, they may not be the same kind of choice.
| Situation | What AI can do | What kind of choice it is for a person |
|---|---|---|
| A test fails | Fix the error and test again | An execution judgment. Fine to hand off |
| How to build a button | Choose a technique and build it | An execution judgment. Usually fine to hand off |
| Changing the product structure | Change the product’s direction | A new direction. Is it within what I handed over? |
| Changing the pricing policy | Change business policy | A new choice. Shouldn’t it come back to a person? |
| A past approval and what I say now differ | Judge which is valid now | Are approval points set in advance enough? |
Ways to hand decisions back already exist
As we saw, graphs already have a way to bring a person in: Human-in-the-Loop, or HITL, which stops execution at a specific point, waits for a person’s approval or input, and then continues from the same point.
For example:
Delete data → human approval
Payment above a set amount → human approval
Test fails → AI fixes it automatically
Structures like these can be built in advance.
So the question wasn’t whether a human could be put into the loop.
The harder question was: when should the system return the decision to the human?
Put three ways people take part in AI’s work side by side, and the structure I’m looking for becomes visible.
Heavy burden on people, little autonomy for AI.
Light burden, but important judgments that arise midway can slip through.
People step in only when a judgment is needed, and make only that judgment.
The third structure can also be built with a graph: add conditional branches, or have the system return the decision to the person based on the original context, the scope that was handed over, and the impact of the action. But the AI classifying something as a “new judgment” does not by itself complete the standard for returning it. The question “New judgment needed?” in the figure actually holds at least three questions.
- Grounds — Is there enough basis to continue with the decisions already made?
- Information — Is information missing or in conflict?
- Impact — Does the action affect others, or is it hard to undo?
The first two ask how uncertain the judgment is; the last asks whether there is authority to take the action. They are not the same thing and need to be handled separately. Even a clear judgment may need to return to the person if the action is hard to undo, and even a low-impact action may need a question if its grounds are shaky.
And what’s missing is not the mechanism but a general standard. These frameworks provide ways to stop and return control, but they do not by themselves tell us what should count as a “new judgment” in a particular human–AI relationship, or how that follows from the scope first handed over and from earlier judgments.
Risks you can foresee can be turned into rules.
But in real work, options appear that weren’t there at the start, situations change, and there are moments when it becomes unclear whether a judgment made a few days ago still holds.
Even in those moments, AI could probably produce a pretty good answer.
But another question remains.
Just because AI can make that judgment,
does that mean I delegated that judgment to it?
I find approaches like graphs encouraging because they provide increasingly concrete building blocks for implementing parts of what I have been describing as Mediation. Being able to specify, within the flow of execution, where to stop, what state to save, and where to resume after a person responds means that reusing judgments and returning them to people at the right moment can now be treated as something to design.
So the question has become clearer, not weaker. Now that we can build points where work stops, who decides where it should stop, and by what standard?
Reducing human intervention is not the same as reducing human choice.
What can be automated is not the same as what may be delegated.
What I want to preserve is not the right to intervene at every step. It is the ability to understand what matters as the work continues, and to make the choice myself when the work reaches a decision that differs from what I originally handed over.
The technology that extends how far AI can go is advancing quickly. Alongside it, I want to study further when a decision should return to the human, and what people should be able to see and weigh when it does.
I haven’t settled on an answer yet. But the more agents work for longer, the more often I think we’ll run into this question, and that is why I want to hold on to it now.
References
- Google Developers Blog (2026-06-30). Build reliable multi-agent applications with ADK Go 2.0. Source
Introduces a graph-based workflow engine, built-in HITL, and dynamic orchestration. - Microsoft Learn. Workflow orchestrations in Agent Framework. Source
Sequential, concurrent, handoff, group chat, and Magentic orchestrations, with human involvement. - OpenAI Agents SDK. Agent orchestration. Source
Manager (agents as tools) and handoff patterns. - Anthropic Engineering (2025-06-13). How we built our multi-agent research system. Source
An orchestrator-worker structure in which a lead agent coordinates parallel subagents.