Chapter 8

Workflow before agent

Agents inherit the workflow they are dropped into. Map it, own it and design its handoffs before you add the agent.

The easiest way to waste money on AI is to automate a workflow nobody understands.

A team sees a painful process and reaches for an agent. The demo looks impressive. It drafts, summarizes, routes, classifies, researches, or recommends. The first examples feel promising.

Then the workflow pushes back.

The data is inconsistent. Handoffs are unclear. Exceptions are common. Output is useful sometimes, but nobody knows who reviews it. The team argues about quality because the original standard was never written down.

The agent did not create the confusion. It revealed it.

The first rule of practical AI agent management is simple:

Workflow before agent.

Do not begin with “What agent should we build?” Begin with “What workflow are we trying to improve, and how does it actually operate today?”

Agents inherit the workflow

A useful agent depends on the system around it. If the workflow has a clear purpose, clean inputs, known decision points, defined owners, visible success metrics, and review cadence, an agent can become leverage.

If the workflow is vague, fragmented, political, undocumented, or dependent on tribal knowledge, the agent inherits that confusion. It may still produce output. That is the danger. The workflow can look more productive while becoming harder to manage.

This is not an argument against agents. It is an argument for management.

Leaders should not expect agents to compensate for unclear purpose, missing data ownership, undefined quality standards, or absent decision rights. AI can help once the workflow is understood. It should not be used to hide that the workflow is not understood.

Start with the outcome

Every workflow exists to produce an outcome. State it in plain language.

Not:

“We want to use AI for customer success.”

But:

“We want customer success managers to identify expansion risk two weeks earlier using support activity, product usage, account notes, and recent customer conversations.”

Not:

“We want an engineering agent.”

But:

“We want to reduce the time between bug report and verified root cause for high-priority production issues.”

Not:

“We want AI for marketing.”

But:

“We want to turn approved customer, product, and sales insights into useful field notes without adding another weekly writing bottleneck.”

The outcome clarifies the workflow. The workflow clarifies the agent.

Map the current workflow

A current-state map does not need to be fancy. It needs to be honest. For the workflow you want to improve, write down:

  1. Trigger: What starts it?
  2. Input: What information, data, request, or event enters?
  3. Actors and AI touchpoints: Which people, teams, systems, agents, automations, or models participate today?
  4. Decisions: What decisions are made?
  5. Handoffs: Where does work move?
  6. Outputs: What artifact, decision, action, or customer experience is produced?
  7. Review: Who checks quality, risk, completeness, or accuracy?
  8. Exceptions: What happens off the happy path?
  9. Metrics: How do you know whether it works?
  10. Owner: Who is accountable for performance?

The last question usually tells the truth. Many workflows have participants but no owner. An agent should not be introduced into an ownerless workflow. If no one owns the workflow, the first design task is ownership design.

Find the decision points

Agents are often described by tasks: summarize, classify, draft, research, route, analyze, monitor, write code, update records.

Tasks matter. Decisions matter more.

Every workflow has decisions, including hidden ones:

  • Is this customer issue urgent or routine?
  • Should this bug be escalated or batched?
  • Is this lead worth human follow-up?
  • Does this contract clause require legal review?
  • Should this generated response be sent, edited, or rejected?
  • Is this anomaly meaningful or noise?
  • Does this pull request meet review standards?

Before introducing an agent, decide which decisions it may assist, which it may recommend, and which require human approval. Also name who approves the use case, grants data access, launches production, handles exceptions, and can retire or roll back the workflow.

Separate assistance, recommendation, and action

Design boundaries in three levels:

  1. Assistance: The agent helps a human do the work.
  2. Recommendation: The agent proposes a decision or next action.
  3. Action: The agent executes inside a workflow or system.

These levels should not be governed the same way. An agent that assists research may need context but little permission. An agent that recommends account prioritization needs outcome evaluation and bias checks. An agent that updates records, sends messages, merges or deploys code, or changes operational state needs scoped tool permissions, a trace of every action, review rules, and rollback paths. With MCP and similar connectors, giving an agent a new tool can take minutes. Treat each grant as a decision about action, not a setup step.

Many companies jump from assistance to action because the demo makes it look easy. A better path is progressive delegation. Let the agent assist. Watch corrections. Improve context and evaluation. Then let it recommend. Compare recommendations to human decisions. Only after trust is earned should it act, and then within clear boundaries.

Delegation should be earned.

Design for exceptions, not demos

Demos show the happy path. Real workflows are made of exceptions: missing data, unusual customers, policy changes, incomplete bug reports, integration failures, conflicting instructions, confident wrong answers.

Before deployment, define exception paths. A simple anomaly register is often enough: missing data, contradictory data, unusual cases, integration failures, conflicting instructions, gameable metrics, and stop conditions.

For each anomaly, specify what the agent may do, when it must stop, who receives escalation, where the decision is logged, and how learning updates prompts, context, evals, policy, or the workflow.

A mature AI operating system assumes failure will happen. It makes failure visible, recoverable, and useful.

Design the handoffs between agents

More workflows now use several agents: one researches, one drafts, one checks, one acts. Each handoff is a place where work can be lost, duplicated, or run twice.

A message that says “done” is not a handoff. Write it down: what was handed over, who holds it now, what state it is in, and what the next agent may do with it. Give each piece of work one owner at a time, human or agent, so two agents do not act on the same item. Keep a record of the chain that a person can read, so that when something goes wrong, someone can see which step failed.

The rules that work for people still apply: one owner, a clear definition of done, and an escalation path.

Make review capacity part of design

Agents can create more work than teams can inspect. This is especially true for writing, coding, analysis, compliance-sensitive workflows, and customer communication.

Ask:

  • Who reviews output?
  • What standard do they use?
  • How long does review take?
  • What percentage needs review?
  • What can be sampled?
  • What must be reviewed every time?
  • What happens when reviewers disagree?
  • Is review replacing higher-value work?

A pilot with ten reviewed outputs may look fine. At one hundred or one thousand, the bottleneck moves from production to judgment. Coding agents make this concrete: when agents can open pull requests around the clock, review becomes the constraint on delivery. AI makes production cheaper. It does not make judgment cheap.

Workflow-before-agent map

Use this before approving a new agent or expanding an existing one:

  1. Outcome: What business or customer outcome should improve, and why now?
  2. Workflow: What produces that outcome today, and where are delays, rework, and handoff failures?
  3. Ownership: Who owns the workflow, agent component, data quality, review, and exceptions?
  4. Decisions: What decisions happen, and which can the agent assist, recommend, or act on?
  5. Data, context, and tools: What information is needed, where does it come from, what is the source of truth, and which tools or connectors may the agent use?
  6. Quality and risk: What does good output look like, what errors matter, which evals check it, and what anomaly log captures cases the evals miss?
  7. Review and cadence: Who reviews performance, what gets logged, and when should the agent be improved, paused, narrowed, rolled back, or retired?

If the team cannot answer, do not build the agent yet. Clarify the workflow.

Two short examples

For sales research, the weak version is “generate emails faster.” The workflow version is “create higher-quality conversations with accounts that match our business-model fit and show credible purchase intent.” That changes the design: account selection, fit scoring, research, signal detection, hypothesis, artifact selection, outreach review, CRM update, and follow-up learning.

For engineering bug triage, the weak version is “classify bugs.” The workflow version is “reduce the time between validated customer-impacting bug report and accountable engineering owner, without increasing false urgency or interrupt load.” That changes the design: evidence collection, duplicate checks, impact assessment, owner suggestion, escalation rules, and review against actual triage outcomes. If a coding agent then drafts the fix, the same workflow needs a reproduction test, a human reviewer, and a named person who approves the merge.

In both cases, the agent is useful because the workflow is designed.

Chapter 8 operating check

Before adding or expanding an agent, ask:

  1. What outcome are we trying to improve?
  2. What workflow creates that outcome today?
  3. Who owns the workflow end to end?
  4. What decisions happen inside it?
  5. Which decisions can an agent assist, recommend, or act on?
  6. What data and context does the agent need?
  7. Which tools can it call, and with what permissions?
  8. Does it hand work to other agents, and how is each handoff recorded?
  9. What does quality mean, and which evals check it?
  10. What failure modes matter?
  11. Who reviews output and performance?
  12. What cadence improves, pauses, or retires the agent?

If the team can answer clearly, an agent may be the right next step. If not, the most valuable work is workflow design.