In 2014 you could get a Docker container running on your laptop in about five minutes. Getting fifty of them running reliably in production – restarting when they fell over, finding each other on the network, not leaking secrets everywhere – took the industry the best part of four years and a small war.
I think we’re at exactly that point with AI agents.
Building one agent is easy now. Claude Code, Codex, Copilot, or a Python script calling an API with a few tools bolted on. I run several most days. Getting a dozen of them to work together on something that matters to your business – with the right permissions, a record of what they did and a bill you can predict – is the hard bit. And that’s where the next big fight in AI is going to happen. Not the models. The orchestration layer that sits on top of them.
The land grab has already started
Look at what’s landed in the last eight weeks alone.
xAI launched Grok Bot in August: persistent, always-on agents that keep working on a cloud machine after you’ve shut your laptop. You can drop two to six of them into a group chat and they coordinate between themselves. On 1 October they followed it with an agent dashboard in Grok Build – “all your agents on one screen” – with each agent running in its own isolated Git worktree.
OpenAI used DevDay at the end of September to launch Dots, always-on agents with 4,000+ app integrations, alongside an Agents API with multi-agent support, computer use and context compaction built in. Tellingly, Dots are switched off by default for Enterprise workspaces until an admin turns them on. Somebody at OpenAI has clearly sat in a meeting with a CISO.
Microsoft has spent 2026 turning Foundry into a model-agnostic control plane for agents, with Agent 365 and Entra handling identity and policy. It’ll happily run GPT-6 or Claude Opus 5.5 behind the same governance layer, which tells you where Microsoft thinks the value sits.
And the plumbing has consolidated. MCP (how an agent talks to tools and data) and A2A (how agents discover each other and hand off work) now both sit under the Linux Foundation’s Agentic AI Foundation. A2A moved across in August.
Read that back and swap a few nouns. Docker. Swarm. Mesos. Kubernetes. OCI. CNCF. It’s the same shape.
How containers actually went mainstream
Containers weren’t new in 2013. LXC, chroot and Solaris Zones had been around for years. What Docker did was make them easy – a standard image format, a simple CLI, and “works on my machine” finally meaning something. Adoption went vertical.
Then came the hangover. Teams had hundreds of containers and no answers to basic questions. Which host is this running on? What happens when it dies at 3am? How does service A find service B? Who’s allowed to deploy what, where? I have seen organisations end up with a container estate that was harder to run than the VMs it replaced, because they’d adopted the unit without the control layer.
Orchestration was the answer, and for a couple of years everyone had one – a bit like everyone has an AI strategy now. Docker Swarm, Mesos with Marathon, Nomad, Kubernetes. Kubernetes won. Docker itself quietly conceded in late 2017 and shipped Kubernetes support in its own product.
Then something more important happened: Kubernetes got boring. The CNCF gave it neutral governance, OCI standardised the image and runtime specs, and AKS, EKS and GKE turned “run a cluster” into a tick box in the portal. That’s when containers properly went mainstream. Not when Docker launched – when orchestration made them safe and predictable enough for a normal business to run.
Containers were the unit. Orchestration made them useful.
Agents are at the 2015 point
An agent today is a container in 2014. A powerful unit, easy to start, genuinely useful on its own. Then you try to run several of them together and the questions start – and they’re almost identical to the ones we asked a decade ago.
That last row is the one that’ll bite. Container cost is broadly predictable: you pay for the node whether it’s busy or not. Agent cost scales with how much the agent thinks, and an agent stuck in a retry loop thinks a lot. If you’ve ever had an Azure bill surprise you, imagine one where the workload decides for itself how much compute it needs.
What an orchestrator actually has to do
This is where it gets technical, and where the Kubernetes comparison earns its keep. Strip away the marketing and every serious agent orchestration platform is building the same set of components. Most of them have a direct Kubernetes ancestor.
Routing and scheduling. kube-scheduler decides which node a pod lands on based on resources and constraints. An agent router decides which agent – and which model – gets a task. Classification and triage go to a small, cheap model; the hard reasoning goes to a frontier model. Get this wrong and you’re paying Opus prices to sort emails.
The supervisor loop. The clever bit of Kubernetes was never the scheduler, it was the reconciliation loop. You declare the state you want, and controllers keep nudging reality towards it – forever. Agent orchestrators need the same pattern: a supervisor that checks each step against the goal and decides whether to carry on, retry, re-plan or escalate to a human.
State and checkpoints. Kubernetes keeps cluster state in etcd and workload data on persistent volumes. A multi-agent workflow that runs for three hours needs durable checkpoints, so an interruption at hour two doesn’t mean starting again. Foundry already ships this; expect everyone else to follow.
Identity and policy. RBAC, admission controllers and resource quotas are what made Kubernetes acceptable to security teams. The agent equivalents are an identity per agent (Microsoft is doing this in Entra), scoped permissions, human approval gates for anything irreversible, and token budgets that act like resource quotas.
The protocols. Kubernetes stopped caring about the underlying runtime, network or storage once CRI, CNI and CSI gave everyone a standard plug. MCP is doing the same job for tools and data – write the MCP server once and any compliant agent can use it. A2A is closer to service discovery plus a contract: each agent publishes an “agent card” describing what it can do, and other agents delegate to it.
Observability and cost. Prometheus and OpenTelemetry made clusters visible. Agents need the same – every decision traced, every tool call logged, every token counted against a budget. This is FinOps all over again, just with a different meter.
Where the analogy breaks
It’s not a perfect fit, and the gap is worth understanding. Containers are deterministic. Same image, same inputs, same behaviour. Kubernetes only has to know whether a container is running – the liveness probe passes or it doesn’t.
Agents aren’t deterministic. An agent can be healthy, responsive and confidently wrong. So an agent orchestrator has to do something Kubernetes never did: judge whether the work is any good. Evaluation – automated checks, a second agent reviewing the first, a human sign-off at the right points – has no real container equivalent. I think whoever cracks that in a way normal businesses can configure wins the next round.
What happens next
My bet is the next 18 months look a lot like 2015–2017. Every vendor will call their product an orchestrator. Gartner reckons only about 130 of the thousands of companies selling “agentic AI” are the real thing, and predicts over 40% of agentic AI projects will be cancelled by the end of 2027 – because of cost, unclear value and weak risk controls. Every one of those three is an orchestration problem, not a model problem.
Then it consolidates. Kubernetes didn’t beat Swarm because it was easier – it absolutely wasn’t. It won because it was open, vendor-neutral and everyone could build on it. I’d expect the same here: the platforms that are model-agnostic and play properly with MCP and A2A will win over the walled gardens.
And then it gets boring. Managed agent orchestration becomes a tick box in Azure, AWS and Google Cloud, with sensible defaults for identity, budgets and audit. That’s the point where a 50-person business on the Costa del Sol can actually use this safely – and it’s closer than most people think.
What to do about it now
You don’t need to pick a winner yet. You do need to avoid building something you’ll have to rip out.
- Build on the standards. Use MCP for tool and data access, so whatever you build now moves with you to whichever orchestrator wins. Same logic as building to OCI images in 2016 – it didn’t matter whether you ended up on Swarm or Kubernetes.
- If you’re already on Microsoft 365 and Azure, Foundry with Agent 365 is the natural path. Same Entra identities, same Conditional Access, same audit trail you already have.
- Give every agent its own identity and a budget from day one. No shared service accounts, no unlimited API keys sitting in someone’s
.envfile. - Count tokens like you count cloud spend. Tag them, attribute them to a team or a workflow, set alerts. It’s the same discipline, applied to a new meter.
- Find out what’s already running. I’d put money on there being more agents connected to your M365 tenant, CRM and inbox than anyone in IT has signed off. Shadow IT has a new form factor.
If you want a second pair of eyes on where agents fit in your setup – and where they really don’t – drop me a message. It’s the kind of conversation I’m having with clients a lot at the moment.
Sources
- A Guide to Grok Bot – Composio
- Grok Agent Dashboard – Tesla North
- Agent Platforms, Manager-Worker Splits – Agent Brief
- OpenAI DevDay 2026 Highlights – Reworked
- Microsoft Foundry multi-agent governance – HyperFRAME Research
- A2A joins the Agentic AI Foundation – Forbes
- Gartner: 40% of agentic AI projects cancelled by 2027 – BigDATAwire