It started with one agent. It watched the warehouse management system and flagged orders that looked likely to miss their dispatch cut-off. It worked well enough that operations asked for more. Soon there was a second agent forecasting stockouts, a third proposing re-allocations between distribution centers, and a fourth drafting messages to carriers and key customers when delays looked likely.
Six months later, the “multi-agent system” is really four agents joined by a collection of cron jobs, shared database tables and scripts that one engineer understands. The stockout agent polls the order-risk agent’s table every five minutes. The messaging agent sometimes notifies a customer twice because it can’t tell whether the previous run finished. When the re-allocation agent crashed during a peak week, nobody noticed for half a day.
If you’re a technology lead, architect or data team at a logistics provider, 3PL, retailer’s supply chain function or manufacturer, this article is about how to coordinate multiple AI agents reliably, without the tangle.
Why supply chains are natural multi-agent territory
Supply chain operations are a good fit for multi-agent AI because the work already divides into specialized roles that react to events:
- Monitoring orders, shipments and inventory positions
- Forecasting demand and detecting risk
- Planning re-allocations, expedites and substitutions
- Communicating with carriers, suppliers and customers
- Escalating to human planners when judgment is needed
Each role benefits from its own tools, data and prompts, so separate agents make sense. The difficulty is how they work together.
What goes wrong with ad-hoc coordination
Polling instead of reacting. Agents check shared tables on a timer. That adds latency (a risk spotted at 10:01 isn’t acted on until 10:05) and wastes compute when nothing has changed.
Tight coupling through shared state. When agents coordinate by reading and writing the same tables, every schema change ripples across all of them, and race conditions creep in.
Lost or duplicated work. Without durable execution, an agent that crashes mid-task either loses the work or repeats it on restart, which is how customers end up with two delay notifications or a re-allocation gets applied twice.
No end-to-end view. When something goes wrong, there’s no single place to see that an event triggered a risk assessment, which led to a plan, which was approved and executed. Debugging means correlating logs from four systems.
Long-running steps don’t fit. Some steps take hours or days, such as waiting for a supplier to confirm a revised date or for a planner to approve an expedite. Scripts and cron jobs aren’t designed to hold that kind of state reliably.
A better architecture: events plus durable workflows
Two patterns, used together, solve most of these problems.
Event-driven collaboration
Agents communicate by publishing and subscribing to events on a message broker, not by polling shared tables. The order-risk agent publishes OrderAtRisk. The re-allocation planner subscribes, evaluates options and publishes ReallocationProposed. The communications agent subscribes to ShipmentDelayed and notifies the right people.
This decouples agents from each other: each one knows only the events it consumes and produces. New agents can be added by subscribing to existing events, without changing the others. Reactions happen in seconds rather than at the next polling interval.
Durable workflows for each agent’s work
Events say what happened. Workflows make sure the response finishes. When an agent receives an event, it starts a durable workflow to handle it. Each model call, lookup and action is a recorded step, so a crash doesn’t lose progress, and a completed step, such as sending a carrier message, is never repeated.
Workflows also cover the long-running parts: waiting for a planner’s approval or a supplier’s confirmation becomes a durable wait that can last days without holding any compute, with a timeout and escalation path defined in code.
Orchestration where order matters
Not everything should be choreographed through events. Where a sequence must happen in a specific order, such as “assess, plan, get approval, execute, confirm”, a single orchestrating workflow that calls each agent in turn is easier to reason about. Many real systems mix both: events for loosely coupled reactions, orchestration for critical multi-step processes.
Building it with open-source components
Dapr, a graduated project in the Cloud Native Computing Foundation, provides both patterns as runtime building blocks. Its pub/sub API supports many brokers, including Kafka, RabbitMQ, Redis and several cloud messaging services, behind a single interface. Its workflow engine provides durable execution for each agent’s tasks. The open-source Dapr Agents framework builds on both, supporting deterministic workflow-based orchestration as well as event-driven collaboration between agents. Agents written with other frameworks, such as LangGraph or CrewAI, can also run on Dapr’s workflow engine without being rewritten.
A logistics company that published its experience took this route to automate warehouse supervision: a set of cooperating agents identifies at-risk orders and predicts stockouts, built on Dapr Agents and running on a managed Dapr-based platform.
Design principles that keep multi-agent systems manageable
Model events on business facts. Name events after things that happened in the operation (OrderAtRisk, StockoutPredicted, CarrierDelayConfirmed), not after agent internals. Business-level events stay stable as agents evolve.
Make every side effect idempotent. Use event IDs or business keys so that a message sent to a carrier, or a change applied to the WMS, has no extra effect if it’s delivered twice.
Give each agent a clear boundary. One responsibility, a defined set of tools, and explicit input and output events. If an agent’s description needs “and” three times, split it.
Put humans in the loop deliberately. Decide in advance which decisions need a planner’s approval, such as re-allocations above a certain value or customer-facing messages for key accounts, and model those as durable waits with timeouts.
Trace across agents. Propagate trace context through events and workflows so you can follow one incident from the first signal to the final action across every agent involved.
Secure agent-to-agent traffic. Give each agent its own identity and use mutual TLS between them, especially when agents can trigger operational changes or external communications.
Start small and add agents one at a time. Begin with one well-defined event, such as an order at risk of missing its cut-off, and one or two agents that react to it. Prove that the flow is reliable end to end, including crashes and duplicate events, before adding the next agent. Multi-agent systems are much easier to grow than to untangle.
Operating at scale
As the number of agents grows, so does the operational load: brokers, workflow state, upgrades, identity, monitoring. Teams can run open-source Dapr on their own Kubernetes clusters. Others prefer a managed platform so engineers can focus on the agents rather than the infrastructure. Diagrid is one provider in this space; its Catalyst platform offers durable workflows, event-driven messaging between agents and integrations with common agent frameworks, either hosted or deployed in the customer’s own cloud.
Conclusion
Multi-agent AI suits supply chains well, but only if the agents can coordinate as reliably as the operations they support. Replacing polling scripts and shared tables with events for collaboration and durable workflows for execution gives you agents that react in seconds, never lose or duplicate work, wait patiently for humans and suppliers, and leave a trail you can follow when something goes wrong. The result looks less like a tangle of scripts and more like the well-run operation it’s meant to support.