Skip to main content
Book a call
All essays
Essay

AI Strategy

AI Strategy · · 14 min read

Is This Workflow Ready for an AI Agent? A 12-Point Assessment

Score any business workflow across 12 practical criteria to decide whether it is ready for an AI agent, needs redesign, or should remain human-led.

Hand-drawn twelve-point workflow readiness checklist

The worst way to choose an AI agent project is to ask, “What could we automate?”

The answer is almost everything. Given enough time and money, you can build a system that attempts nearly any workflow.

That does not mean you should.

Some workflows make excellent first agents. They happen frequently, use accessible information, produce a clear output, and tolerate a review step. Others look impressive in a demo but collapse in production because the process changes every week, the source data lives in people’s heads, or a single mistake can cause serious harm.

The difference is workflow readiness.

This assessment scores a workflow across 12 points. It helps you decide whether to build now, run a constrained pilot, redesign the process first, or leave it primarily human-led.

It evaluates the workflow, not your company’s entire AI maturity. For the organizational view, use our complete AI audit framework. Here, we are deciding whether one specific unit of work deserves an agent.

How to Use the Assessment

Choose one workflow with a specific beginning and end.

Good examples:

  • Prepare a pre-call account brief
  • Triage an inbound support ticket
  • Reconcile a supplier invoice against a purchase order
  • Monitor competitor ads and prepare a weekly change report
  • Turn discovery notes into a draft proposal

Bad examples:

  • Help with sales
  • Automate finance
  • Run customer success
  • Become an AI-first company

Score each of the following criteria from 0 to 2:

  • 0: The condition is absent or unknown.
  • 1: The condition is partly true, inconsistent, or needs work.
  • 2: The condition is clearly true and supported by evidence.

The maximum score is 24. Do not award points based on what the process owner hopes is true. Use samples, system data, and observation.

1. The Workflow Happens Often Enough

Agents have fixed costs: discovery, integration, testing, training, monitoring, and maintenance. A workflow needs enough volume for those costs to make sense.

Score 0: It happens a few times per year or volume is unknown.

Score 1: It happens monthly, or volume varies enough that the annual opportunity is uncertain.

Score 2: It happens weekly or daily at meaningful volume, and you can quantify that volume.

Frequency alone is not enough. A quarterly process performed by 500 people may be more valuable than a daily process performed by one. Measure annual units of work.

Annual workflow volume =
  frequency per person × number of people × working periods per year

If nobody can tell you how often the work occurs, score zero. Establish the baseline before evaluating the technology.

2. The Current Pain Is Measurable

“This is annoying” is a useful discovery signal but a poor investment case.

You need at least one measurable problem:

  • Human time per unit
  • End-to-end cycle time
  • Backlog size
  • Error or rework rate
  • Missed service-level agreements
  • Lost revenue or capacity
  • Customer or employee friction

Score 0: The pain is anecdotal and no baseline exists.

Score 1: The team agrees on the problem, but measurement is incomplete.

Score 2: The current cost, delay, quality, or business outcome is measured.

An agent can improve a workflow and still produce no visible ROI if the company never measured the starting point. Our AI agent ROI framework shows how to turn this baseline into a business case.

3. The Trigger and Output Are Clear

A reliable workflow has an observable start and a defined result.

For example:

Trigger:
  A qualified discovery call ends and the transcript is available.

Output:
  A draft proposal using approved scope language, pricing rules,
  and relevant case studies, ready for consultant review.

Score 0: People disagree about when the process starts or what it should produce.

Score 1: The general shape is known, but outputs vary by person or department.

Score 2: Trigger, completion state, required fields, and output format are documented.

If the output is “helpful insight” or “better decisions,” keep working. Define the artifact, decision, or system change the agent must create.

4. The Inputs Are Digital and Accessible

An agent cannot reliably use information it cannot reach.

The required input may live in a CRM, ERP, shared drive, email inbox, data warehouse, ticketing system, call recorder, or internal database. What matters is whether it is digital, permitted, and retrievable at the moment of work.

Score 0: Critical information lives in paper, private spreadsheets, undocumented conversations, or inaccessible systems.

Score 1: Most information is digital, but access is inconsistent or requires manual collection.

Score 2: Required inputs are available through stable APIs, exports, databases, or approved tools.

Do not confuse “the company has the data” with “the workflow can access the data.” A sales team may own years of account history while the notes required for a useful brief remain buried in personal inboxes.

5. The Source of Truth Is Known

Agents are good at combining information. They are not magically good at deciding which of three conflicting records is correct.

Score 0: Multiple systems conflict and no precedence rules exist.

Score 1: People know informally which source usually wins, but it is not documented.

Score 2: Each important field or decision has an authoritative source and a freshness rule.

Write rules such as:

Pricing:
  The approved rate card wins over previous proposals.

Customer status:
  The CRM account record wins over meeting notes.

Policy:
  The current policy library wins over email and archived documents.

Without source precedence, the agent will produce confident inconsistencies. The problem is not hallucination. The operating environment itself is ambiguous.

6. The Work Can Be Decomposed Into Decisions

Many workflows look like one task from the outside but contain several different kinds of judgment.

Consider support triage:

  1. Identify the customer.
  2. Retrieve account and product context.
  3. Classify the issue.
  4. Check whether a documented answer applies.
  5. Draft a response.
  6. Decide whether a human specialist is required.
  7. Route or send the response.

An agent may handle steps one through five reliably while a human owns six and seven.

Score 0: The work depends on tacit judgment that experienced people cannot explain.

Score 1: Some decisions are documented, but important branches remain intuitive or person-dependent.

Score 2: The process can be split into observable steps with rules, examples, and escalation conditions.

You do not need to eliminate judgment. You need to locate it. The most useful design may be a hybrid workflow where the agent handles preparation and the human owns the consequential decision.

7. Success Can Be Evaluated

If you cannot evaluate output, you cannot manage the agent.

Useful evaluation methods include:

  • Required-field checks
  • Policy or checklist compliance
  • Comparison with an approved answer
  • Factual accuracy sampling
  • Human correction rate
  • Escalation rate
  • Downstream outcome measurement

Score 0: Quality is subjective and no review standard exists.

Score 1: Experts can recognize good work but have not converted that judgment into a rubric.

Score 2: The output has a repeatable evaluation method and an acceptable threshold.

“A manager will look at it” is not an evaluation plan. What will the manager inspect? Which mistakes matter? What correction rate is acceptable? When should the agent stop and escalate?

8. Errors Are Detectable and Recoverable

Every agent will make mistakes. Readiness depends on whether the workflow can detect and recover from them.

Score 0: Errors can create irreversible financial, legal, safety, employment, or customer harm before anyone notices.

Score 1: Errors are usually recoverable, but detection depends on manual vigilance or occurs late.

Score 2: Outputs can be reviewed before action, changes are reversible, and high-risk conditions trigger escalation.

Drafting a customer response is more ready than autonomously issuing a refund. Preparing a candidate summary is more ready than rejecting a candidate. Flagging an unusual transaction is more ready than freezing an account.

This does not mean high-stakes workflows can never use agents. It means the first design should assign the agent the reversible preparation work and keep authority with a qualified person.

9. The Agent Can Act Through Stable Tools

An agent becomes valuable when it can leave the workflow in a better state: create the brief, update the record, prepare the draft, or route the exception.

Score 0: Required systems have no practical integration path, or policy prevents access.

Score 1: Some tools are accessible, but critical steps require fragile browser work or manual handoffs.

Score 2: Required read and write actions are available through supported APIs, reliable automation, or controlled interfaces.

Browser automation can be valid, especially when no API exists, but treat it as a maintenance cost. Interfaces change. Sessions expire. Anti-automation controls appear. A workflow that depends on five brittle browser steps needs a different cost and reliability assumption than one using a stable API.

10. A Human Owner Is Named

Every production agent needs a manager.

The owner is responsible for:

  • Output quality
  • Access and permissions
  • Escalation policy
  • Evaluation and monitoring
  • User feedback
  • Updating instructions and context
  • Deciding when to expand or stop

Score 0: The owner is “IT,” “the AI team,” or nobody.

Score 1: A department supports the idea, but no individual owns operating results.

Score 2: One named person owns the workflow metric and has authority to change the process.

The owner should usually come from the function where the work lives. Technology teams can support the system, but a sales leader should own a sales workflow and a finance leader should own a finance workflow.

An agent without a functional owner becomes a technical artifact. It may continue running, but nobody knows whether it still helps.

That owner-and-agent relationship is one of the defining differences between an AI-native company and a collection of individual AI users.

11. The People Doing the Work Have a Reason to Adopt It

Adoption fails when the system improves leadership’s spreadsheet while making frontline work worse.

Ask the people doing the workflow:

  • What part of this work do you want to stop doing?
  • What part must remain under your control?
  • Where would the output need to appear?
  • What would make you distrust it?
  • What happens to the time the system returns?

Score 0: Users were not involved, incentives conflict, or the agent threatens their role without a credible plan.

Score 1: Users see potential value but expect substantial workflow change or have unresolved concerns.

Score 2: Users helped design the workflow, the output appears where they already work, and the benefit is clear to them.

Do not launch by telling people that the agent will “free them for strategic work” unless managers have identified that strategic work and changed expectations accordingly.

12. The Feedback Loop Is Designed

The first production version will not be the final version.

A ready workflow captures what happened:

  • Did the agent complete the task?
  • Did the human approve, edit, reject, or escalate?
  • What changed during review?
  • Which input or rule caused the error?
  • Did the downstream outcome improve?

Score 0: The agent produces output, but usage and corrections disappear.

Score 1: Feedback can be collected manually, but no regular review cadence exists.

Score 2: Corrections are captured, failure patterns are reviewed, and someone updates the system on a defined cadence.

The goal is not for the model to mysteriously learn from every interaction. The goal is for the organization to learn. A repeated correction should become a better rule, a clearer source, a stronger evaluation, or a redesigned step.

Calculate the Score

Add the 12 scores for a total out of 24.

ScoreReadinessRecommended action
19 to 24ReadyBuild a measured pilot with production guardrails
13 to 18PromisingFix the weak criteria, then run a narrow pilot
7 to 12Process firstRedesign and document the workflow before building
0 to 6Not readyKeep it human-led or choose a different workflow

The total score is not the only rule. Four criteria are gates:

  • Inputs are accessible
  • Success can be evaluated
  • Errors are recoverable
  • A human owner is named

A zero on any gate blocks an autonomous workflow. You may still build a read-only assistant or a tightly supervised experiment, but do not pretend it is production-ready.

Worked Example: Weekly Competitor Monitoring

Suppose a marketing agency has analysts manually reviewing Meta, LinkedIn, and Google ad libraries for 40 competitors. Every Friday they compare the current ads with the previous week and prepare a client brief.

CriterionScoreReason
Frequency and volume2Weekly across 40 competitors
Measurable pain220 analyst hours per week
Clear trigger and output2Friday run, structured change brief
Accessible inputs1Public sources, but some require browser automation
Known source of truth2Platform libraries plus captured snapshots
Decomposable decisions2Collect, compare, classify, summarize, escalate
Evaluatable success2Coverage, extraction accuracy, change-detection precision
Recoverable errors2Analyst reviews before client distribution
Stable tools1Browser workflows require maintenance
Named owner2Director of Strategy
Adoption incentive2Analysts want research time back
Feedback loop2Corrections logged during weekly review
Total22/24Ready for a pilot

This is a strong candidate. The two weak points, input access and browser stability, belong in the cost and maintenance model. They do not invalidate the workflow because the output is reviewed and errors are recoverable.

That is exactly the kind of workflow behind our automated competitor ad intelligence case study.

Contrast Example: Autonomous Employee Performance Decisions

Now consider an agent that reads employee communications and performance data, then decides who should receive a poor rating.

It may have high volume, digital inputs, and a clear output. It still scores poorly where it matters:

  • The source data is incomplete and context-dependent.
  • Good performance requires nuanced managerial judgment.
  • Evaluation criteria may embed historical bias.
  • Errors can materially harm employment, compensation, and trust.
  • A human review after the decision may arrive too late.

The right design is not “improve the prompt.” Use agents for reversible preparation: collect evidence, check whether reviews include required examples, identify missing documentation, or summarize an employee’s stated goals. Keep the performance judgment with the manager.

Workflow readiness is about assigning the right work to the right actor, not maximizing the number of tasks labeled automated.

What to Do With a Low Score

A low score often reveals a valuable process-improvement project.

If the trigger is unclear, define ownership and service levels. If information conflicts, establish a source of truth. If success cannot be evaluated, create a rubric. If errors are hard to reverse, split preparation from action. If users resist, redesign the workflow with them.

This is an important pattern: preparing a workflow for an agent often improves the human process before the agent exists.

The company gains clearer ownership, better documentation, cleaner data, and measurable outputs. If the economics later say not to build, that work was not wasted.

Readiness Before Technology

Model choice comes later.

First decide whether the work is frequent, costly, observable, accessible, decomposable, evaluatable, recoverable, ownable, and capable of improving over time.

If it is, you have a candidate worth testing. Calculate the ROI of the proposed AI agent, define the target AI operating model, and run the smallest pilot that can prove or disprove the business case.

If it is not, do not force it. Redesign the workflow or pick a better one.

The fastest AI program is not the one that builds the most agents. It is the one that rejects weak projects early and concentrates effort where the work is ready to change.

If you want a structured map of the workflows across your business, ranked by readiness, impact, and implementation effort, our AI Audit produces that roadmap.

More to read

Keep reading.

Browse the rest of the essays, or if you'd rather talk, 30 minutes is enough to know whether a partnership makes sense.