ARTICLE 09 · AI PRODUCT · 2026-06-14

Multi-agent business platforms: when several AI roles work together with a clear workflow and clear accountability

Adding agents does not automatically make a system smarter. The payoff comes when each role has defined tools and data, and a clear point where it hands its work to the next.

Multi-agent business platforms: when several AI roles work together with a clear workflow and clear accountability
The short version
  • Giving one AI everything to do is like asking one employee to fill five positions. When something goes wrong, you can't find where the mistake happened.
  • If you split the work across several agents, you need a “supervisor” that hands out tasks and combines the results, plus a handoff sheet that states exactly what goes to whom. The agents should never be left to sort things out by chatting among themselves.
  • The yardstick is whether the work reaches the customer complete. The number of AIs tells you nothing. If splitting doesn't make the work safer or easier to check, don't split.

Where old websites and apps stop

Picture one employee who has to take customer calls, write quotations, check stock, review work and approve discounts, all at once. They might cope on a quiet day, but when something goes wrong you can't trace which step failed, and nobody dares change how they work because touching one thing affects everything. An AI that handles everything on its own runs into exactly the same problem.

A single agent that takes on every job ends up with a huge prompt, is hard to review and has access to more tools than it needs. Splitting the work across several agents without an orchestrator causes its own trouble: duplicate answers and nobody owning the result.

A single agent carrying every tool, compared with a team of agents that divide the work through one central coordinator
A single agent carrying every tool, compared with a team of agents that divide the work through one central coordinator

When a business has one agent read documents, plan, talk to other systems, check quality and approve work at the same time, the context gets long, the permissions get wide, and finding the cause of a wrong result gets hard. The team may keep adding to the prompt until nobody dares edit it, because a single instruction affects behavior across several functions.

A multi-agent platform splits roles when there is a real reason to: different responsibilities, tools, data or ways of evaluating the work. Each agent takes on a limited task and hands its output to the next one according to a contract. It works like a team with a coordinator, an analyst, an operator and a reviewer, and a person is still accountable for the final result.

The old wayThe new AI product approach
A single chatbot answering for every department, or several bots working separatelyA manager agent breaks the work down, calls specialists as tools or through handoffs according to their roles, combines the results, checks guardrails and records a trace of every step

Agents need to connect through task state and contracts that software can check, instead of passing loose messages back and forth. A tool gateway checks each role's permissions, and a combined trace lets the team see which agents and which data every output passed through.

New capabilities a business can put to work

The fix is to organize the AI like a real team: someone researches, someone plans, someone pulls data, someone executes and someone reviews, with one team lead who takes the brief, assigns the work and combines the results. What you, as the owner, should see on screen is how far the work has got, who is holding it, where the data came from and where it is waiting for your approval. A single chat box that hides everything behind it can't show you any of that.

What the project looks like

The platform has an orchestrator that takes a goal, breaks it into tasks and sends them to specialist agents such as Research, Planning, Data, Operation or Review. Users can see the plan, the status, the data sources and the points waiting for approval. The work should never be hidden behind a single chat screen.

Key features

For the development team · Feature list

Features may include a Task Board, Agent Registry, Handoff Contract, Shared Artifact, Approval Queue, Exception handling, Trace, Cost/Latency Monitor and Replay. Administrators can set which agent may use which tool or data, and can stop the whole job or a single step when something looks wrong.

The technology behind it

For the development team · System architecture

The system needs an Orchestration Layer, State Store, Event/Job Queue, Tool/API Gateway, Identity, Policy and Observability. Agents may use different models depending on the task, but every output has to pass a schema check and an evaluation before it moves on. Loose messages between agents with no contract make the system hard to inspect and hard to recover.

Benefits and the right timing

Splitting roles helps you limit permissions, test one part at a time and swap out one agent without affecting the rest. It suits multi-step work that needs different knowledge or different systems. It is a poor fit for short jobs that a single workflow can handle, because the complexity, cost and handoff time can outweigh the benefit.

  • Specialist agents whose knowledge and tools are limited to their role
  • Manager orchestration that combines the results and owns the final answer
  • Handoffs that let a specialist take over the conversation when it makes sense
  • Trace and evaluation that show which agent made a wrong decision, and where
For business owners: Add an agent only when splitting the role makes it easier to limit permissions, test or recover. If a job can be finished with one clear workflow, using several agents may be a cost you don't need.

What it looks like in practice

The easiest example to picture is work that today has to go around several departments before it's done, such as preparing a quotation for a major customer. Someone has to read the brief, someone has to dig up past data, someone has to work out the price and someone has to check it before it goes out.

Every agent work queue has a human owner who approves the key points and is accountable for the result
Every agent work queue has a human owner who approves the key points and is accountable for the result

A request to launch a new product is split among a Market Research Agent, a Content Agent and an Operations Agent, while the Manager Agent combines the plan, checks dependencies and sends the budget question to a person for approval. No agent has permission to use every system.

In a hypothetical tender response, an Intake Agent organizes the requirements, a Research Agent searches only the approved library, a Solution Agent matches them to the company's capabilities, a Pricing Tool calculates the price, and a Review Agent checks the claims and any gaps before everything goes to the person with authority to approve.

If the pricing data isn't ready, the orchestrator pauses only that branch and lets the other parts carry on. When a reviewer corrects a claim, the system records which agent created it, from which source, and through which contract it was passed. The team can then fix the actual fault without guessing whether the problem came from the model or from the handoff.

Least-capability Agent

Give each agent the fewest tools and the least data it needs to do its job.

This limits the damage when something goes wrong and makes evaluation clearer than handing out broad permissions.

Scope, risks and how to measure results

The most important caution is this: having many AIs is no achievement in itself. The more you split, the more places a handoff can fail, and the slower and more expensive the system gets. Ask one question: how often does the work make it all the way to the customer? If splitting doesn't improve that number, hold off.

The number of agents is no KPI. Measure End-to-end Completion, Handoff Failure, Human Correction, Tool Error, Cost and Recovery Time, and limit permissions on a Least Privilege basis. If agents simply pass messages to one another and nobody owns the state, the system adds risk and is harder to debug than a single agent.

A trace of each agent's work shows clearly where handoffs succeeded and where they failed
A trace of each agent's work shows clearly where handoffs succeeded and where they failed

Multi-agent setups add latency, cost and failure points. Use them when roles and permissions genuinely differ, and never just to make the architecture look modern.

For the development team · Technical metrics

Metrics to track: End-to-end Success, Handoff Error, Tool Rejection, Cost per Outcome and Trace Coverage

  1. Discover: follow the real work and collect examples of normal cases and exceptions
  2. Assist: let AI draft or recommend while people stay in control
  3. Act: switch on tools one at a time after the test set passes
  4. Scale: expand once monitoring, fallback, cost control and an owner are in place

Cut the number of agents when handoff failures are high, latency keeps piling up, or nobody owns the overall state. Merging roles, or turning some steps into fixed rules or tools, usually makes the system easier to maintain and more trustworthy.

Use several agents only when their responsibilities really differ, never to make the system look sophisticated

Before adding another AI, do what you would do before hiring a new employee: write down what this position handles, what data it may use, whom it hands work to and who checks it. If two positions turn out to do the same thing on every line, you don't need two positions.

Build a Responsibility Matrix: which agent receives what input, uses which tool, sends what kind of output, who checks it and who owns the outcome. If two agents use the same data, tools and criteria, you may not need to separate them. Multi-agent setups add cost and failure points, so a split needs a reason rooted in permissions, expertise or evaluation.

A team builds a Responsibility Matrix showing which agent receives what, which tools it uses and who checks it
Before adding an agent, fill in the responsibility table. Any empty cell marks a point that nobody owns.

Design a Handoff Contract that states what context is passed on, what is left out, the timeout and who owns errors. Run trace and evaluation for each agent and for the end-to-end flow, and cap concurrency and cost. The system should let people see the plan and stop flows that carry high stakes.

Build a Responsibility Matrix that sets out the input, output, tools, data, permissions, evaluation and owner for every role. If two agents use the same data, tools and criteria, ask whether they should be merged into one. A split has to actually reduce risk or make the work easier to check.

Start with a process that today involves specialists in several roles and has clear handoffs. Let agents help with one or two stages first, with a shared task state and human approval. Once trace and recovery are working, expand parallelism or autonomy. Don't open with a large team of agents in a single demo.

01
Which agent owns the final result?
02
Why does this agent need to be separate?
03
What is the minimum tool access it needs?
04
Who takes over when a handoff fails?
DNA MAKER · PRODUCT & ENGINEERING

Set up an agent team the way you set up a staff team: clear duties, clear permissions and a clear lead

DNA Maker starts with a process and responsibility workshop instead of drawing lots of agent boxes. We help you choose between manager-as-tools and handoffs, depending on who should control the conversation and the outcome, and we specify human approval points and guardrails for each tool.

A permissions map showing which agent can use which tools and data sets, color-coded for easy review
A permissions map showing which agent can use which tools and data sets, color-coded for easy review

We build a flow prototype that shows the plan, tool calls, handoffs and errors for the process owner to review. We then use an evaluation set to measure both the subtasks and the overall outcome, to prove that several agents really do better than one.

01 · Discovery02 · Product & UX03 · Engineering04 · Pilot & Improve

DNA Maker helps design agent responsibilities, handoff contracts and human control based on the customer's real processes. We pick out the points that should use an ordinary workflow, the points suited to AI, and the points that need a tool with fixed, predictable results. We also design the architecture for state, queues, permissions and audit so that everything can be traced back.

Development can cover the orchestrator, specialist agents, MCP/API tools, a task console, evaluation and observability, starting with one line of work and a set of failure scenarios we agree on together. If your organization has work that passes between several teams every week, bring the actual artifacts and the real waiting points to the conversation. We'll help you work out whether it should be multi-agent, a workflow or a mix of the two.

For the development team · What we can build

DNA Maker can build the agent platform, orchestrator, specialist agents, MCP/API integration, permissions, trace, evaluation and cost/latency monitoring, along with an admin console for switching tools and agents on and off as the work requires.

If your process involves specialists in several roles and context gets lost at each handoff, we can build an Agent Responsibility Map with you and pilot one flow, with a person still owning the final result.

Software engineering glossary

These terms cover the coordinator, handoffs, permissions and traces in a multi-agent system. Use them to check whether each role has a real scope and real pass criteria, or is just one more name on a diagram.

TermWhat it isA simple exampleWhat to ask the development team
OrchestratorThe controller that sequences agent work and pulls it together. The orchestrator owns the plan and the overall status, but it shouldn't hold every permission itself. Each tool still checks the permissions for its own task.A Manager Agent calls in the specialistsWho owns the final outcome?
HandoffPassing control to another agent. A good handoff carries the facts, the evidence, what has already been tried and why the case is being passed on, so whoever takes over can decide straight away.Sending a refund case to the Refund AgentWhat context is passed on, and what has to be left out?
Agent as ToolHaving an agent help with a subtask while the manager stays in control. The specialist agent is wrapped so it can be called like a tool with a clear input and output, which cuts down on uncontrolled conversations between agents.The Research Agent sends its findings to the managerHow are the partial results checked before they are combined?
Least PrivilegeGranting the fewest permissions necessary. Give only the access needed for that task and that time window, which limits the impact when an instruction is wrong or data gets used beyond its intended scope.The Content Agent can read data but can't send emailAre permissions reviewed when a role changes?
TracingRecording the system's sequence of reasoning steps and tool calls. A good trace links the input, model, prompt, tool, output, time and cost, so the team can track down causes and build regression tests.Seeing at which step a flow failedHow is trace data secured, and how long is it kept?

Further reading from the original documents: https://openai.github.io/openai-agents-js/guides/multi-agent/

Try this tomorrow: Pick one task where customers or employees have to switch between several screens. Write down the result you want and the points where a person has to approve. You'll end up with a clearer AI product idea than if you start from “we want a chatbot”.