Blog

The Agent Org Chart: What DL4 Actually Demands

Dimitar Siljanovski
Published Jul 8, 20267min
 Diagram of a DL4 agent org chart: sixteen agents across seven layers reporting to one conductor

DL4 is the delivery level where you hand your shared context to a fleet of agents and step back to review. One request in, a finished feature out, you approving only the parts that matter. Done right, output passes 2.5x. The hard part is not the agents. It is the agent org chart.

Key Takeaways

  • DL3 is knowledge. DL4 is management - you stop driving every phase by hand.

  • DL4 is an org chart: at Intertec, sixteen agents, seven layers, one conductor.

  • Build the ten agents that close the loop first. Security, the learning loop, and discovery come last.

  • Teams stall by building coders and skipping the conductor and governance. The code was never the bottleneck.

DL3 was knowledge. DL4 is management.

DL3 is the knowledge layer - the CLAUDE.md, the context files, the team standard. If you missed it, that is the BRAIN you build before you go agentic. At DL3 you still drive every phase by hand, and output grows from 1.25x to 1.6x. Real, but capped.

DL4 is what happens when you hand that knowledge to a fleet of agents and step back to review. One request in, a finished feature out, you approving the parts that matter. Done right, that crosses 2.5x. The work changes shape: you are no longer the developer who prompts. You are the manager of a team you can't see.

Your job: drive every phase by hand
Unit of work: the prompt
Who executes: you, one session at a time
Realistic output: 1.25x to 1.6x

STATS ROW (3-UP)

  • 1.6x - Where DL3 tops out, with you still driving every phase by hand.
  • >2.5x - What DL4 unlocks when the org chart is built right, not just the coders.
  • 16 · 7 · 1 - Agents, layers, and one conductor in the fleet we run at Intertec.

The hard part was never the agents

An agent that writes code is the easy eighty percent. The moment you have more than one, the real problem shows up, and it is not technical. It is organizational: who hands work to whom, who is allowed to ship, who catches the security regression, who turns a single correction into a standing rule the whole fleet obeys.

That is an org chart. Not a metaphor for one - an actual chart of roles, authority, and review gates, except every box is an agent and you are the only human in the room.

The hard part was never the agents. It's the org chart.

This is not a whiteboard sketch. We built this fleet at Intertec, we run it now on our own products, and we are rolling it into client delivery. Sixteen agents, seven layers, one conductor.

Which agents to build first

You don't build all sixteen at once. You build the ten that close the loop first - request in, finished feature out, human review at the end - and you add the governance layers last. The order is the strategy. Build it backwards and you get a fleet that ships fast and wrong.

01

Close the loop

The ten agents you build first

One request goes in, a finished feature comes out, and a human approves the parts that matter. Nothing else ships until this loop is trustworthy on its own.

One request goes in, a finished feature comes out, and a human approves the parts that matter. Nothing else ships until this loop is trustworthy on its own.

02

The conductor

One interface, not sixteen

The orchestration layer that routes work, holds state, and gives you a single thing to manage. Skip it and you are back to babysitting sixteen sessions by hand.

The orchestration layer that routes work, holds state, and gives you a single thing to manage. Skip it and you are back to babysitting sixteen sessions by hand.

03

Security

Governance, added once the loop holds

The agent that catches the regression a coder would happily ship. A fleet without it does not move slower - it moves wrong, faster.

The agent that catches the regression a coder would happily ship. A fleet without it does not move slower - it moves wrong, faster.

04

The learning loop

Corrections feed the BRAIN

Every correction you make at review flows back into the shared context, so the fleet makes the mistake once instead of sixteen times. This is your DL3 brain, now feeding agents.

Every correction you make at review flows back into the shared context, so the fleet makes the mistake once instead of sixteen times. This is your DL3 brain, now feeding agents.

05

Discovery

Built last, on purpose

The agents that surface work rather than just do it. Powerful, and useless before the loop, the conductor, and governance are solid. Last for a reason.

The agents that surface work rather than just do it. Powerful, and useless before the loop, the conductor, and governance are solid. Last for a reason.

Where DL4 teams stall

The teams that stall at DL4 are the ones that built coders and skipped the conductor and the governance. They have sixteen agents writing code and nothing managing them, so velocity goes up and trust goes down, and they quietly roll the whole thing back.

Here is the trade-off the demos skip: a fleet amplifies whatever you give it. A good rule propagates to sixteen agents instantly. So does a bad one. Without the learning loop and the review gate from DL3, you are not scaling delivery - you are scaling your worst assumption sixteen times over.

The code was never the bottleneck.

Map your delivery to DL4

If you reached DL3, you already have the knowledge. DL4 is the harder, less glamorous work of turning that knowledge into roles, authority, and gates an agent fleet can run without you in every loop. Build the ten that close the loop. Add the conductor. Then govern. In that order.

So here is the question I keep asking our own teams: if you mapped your delivery to DL4 tomorrow, which layer would you keep a human on the longest?

Stop prompting. Start managing a team you can't see.

Frequently asked questions

DL4 is the delivery level where you hand your shared context to a fleet of agents and step back to review. One request goes in, a finished feature comes out, and you approve only the parts that matter. Done right it crosses 2.5x output. It sits one level above DL3, where AI understands your project but you still drive every phase by hand.

DL3 is knowledge: the CLAUDE.md, the context files, the team standard. You still drive every phase by hand, and output grows from 1.25x to 1.6x. DL4 is management: you hand that knowledge to a fleet of agents and review the output instead of producing it. The unit of work moves from the prompt to the request, and the ceiling moves past 2.5x.

An agent that writes code is the easy part. Once you have a fleet, the real problem is organizational: who hands work to whom, who is allowed to ship, who catches the security regression, and who turns a single correction into a standing rule. That is an org chart of roles, authority, and review gates, except every box is an agent and you are the only human in the loop.

Build the ten that close the loop first: request in, finished feature out, human review at the end, with a conductor to orchestrate them. Add the governance layers last, in this order: security, the learning loop that feeds corrections back into your shared context, and discovery. The sequence is the strategy. Build it backwards and you get a fleet that ships fast and wrong.

Because they build coders and skip the conductor and the governance. With sixteen agents writing code and nothing managing them, velocity rises while trust falls, and the team rolls the experiment back. A fleet amplifies whatever you give it, so a bad rule propagates to every agent instantly. The code was never the bottleneck.

DL3 takes a team from roughly 1.25x to 1.6x, with a human driving every phase. DL4, done right, crosses 2.5x by handing execution to an agent fleet and keeping the human on review. The gain is real but conditional: it only shows up once the conductor and governance layers exist. Coders alone do not get you there.

Dimitar Siljanovski

Founder and CEO of Intertec.io

Dimitar is the Founder and CEO of Intertec, a custom software development company serving the DACH region. He writes about AI in software engineering, context engineering, and the gap between AI hype and production reality.

Let's talk about your project

Legacy modernization, custom development, or AI, we help companies ship software that drives real business outcomes.

View all posts