The unit of adoption should not be “AI”

Saying “we are going to use AI in procurement,” “in quality,” or “in engineering” is too broad to design a responsible implementation.

Each function contains tasks with different data, consequences, and levels of risk.

A tool may help summarize documentation, classify information, or prepare an initial analysis and, within the same department, be unsuitable for a decision that requires physical interpretation, technical judgment, or regulatory responsibility.

That is why the useful unit of design is the specific workflow.

OpenAI Academy proposes evaluating AI opportunities by starting with an observable process problem: frequency, delays, rework, dependencies, systems involved, owners, and operational stability. The recommendation may be to test, validate further, postpone, or avoid automation when the process is not yet sufficiently defined.

AI capability is not uniform

Research by Harvard Business School and Boston Consulting Group on the jagged technological frontier documented an important behavior: AI can substantially improve performance on certain tasks and worsen it on others that appear similar in difficulty.

In the experiment with 758 consultants, participants who used GPT-4 achieved meaningful improvements on tasks inside the model’s capability frontier. Yet on a complex task outside that frontier, participants using AI were less likely to produce the correct answer.

The finding should not be extrapolated directly to an industrial plant. The tasks studied involved consulting and knowledge work.

The implication is transferable: an organization should not assume that performance observed on one task demonstrates reliability across all nearby tasks.

A reliable workflow needs visible boundaries

Before introducing AI, the process should answer at least five questions.

1. What is the expected result?

“Save time” is too imprecise. The workflow must define the output it produces and the condition that makes that output acceptable.

2. What information may it use?

Input data must be identified. Old versions, contradictory documents, or information without an owner turn an information problem into an automation problem.

3. Where can it fail?

Not every failure has the same impact. An error in an internal draft and an error in a specification sent to production require different controls.

4. Who validates the result?

Human review must have a real owner. “Someone will review it” is not governance.

5. What evidence will determine whether it should scale?

Adoption should not be measured only by user counts or speed. It may also require accuracy, reduced rework, quality, errors, overrides, escalations, and satisfaction among the people who use the process.

Apparent productivity also needs validation

The Anthropic Economic Index has shown that models can produce substantial time savings, especially on complex tasks, but its own analyses adjust productivity estimates when they account for the probability of task success.

That distinction matters.

A task can be completed much faster and then require more correction. It can also produce a fluent answer whose apparent quality exceeds its actual reliability.

Measuring only minutes saved can therefore overstate value.

OpenAI Academy recommends gathering evidence from the workflow in use: execution data, quality reviews, structured feedback, logs, observation, and metrics tied to the original problem.

A workflow does not prove value because it works once. It proves value when the evidence can explain what it does well, where it fails, and under what conditions it should be used.

Governance begins before deployment

Governing AI is not only about writing a general policy for permitted tools.

It also means designing each implementation with:

  • a workflow owner;
  • a defined purpose;
  • identified data and inputs;
  • known limits;
  • review proportional to risk;
  • an evidence record;
  • criteria to continue, correct, or stop.

In industry, this approach makes it possible to distinguish useful automation from premature automation.

The mature question stops being “what can AI do?” and becomes “what part of this process can we delegate, under what conditions, and with what evidence?”

Frequently asked questions

Which workflow should a company automate first?

One with an observable problem, sufficient stability, identifiable inputs, clear owners, and a realistic way to measure whether the intervention improves the work.

Does every AI output need human approval?

Not necessarily. The level of review should be proportional to the impact of an error, the system’s degree of autonomy, and the available evidence about its performance.

How is the value of an AI workflow measured?

With metrics tied to the original problem: total time, quality, accuracy, errors, rework, escalations, use, and qualitative evidence. Speed alone does not demonstrate value.

Was the “jagged technological frontier” studied in manufacturing?

No. The original study used consulting tasks. It is cited to show that AI performance can be uneven across tasks, not to attribute specific productivity figures to an industrial plant.

Sources

  1. Harvard Business School AI Institute. “Back to the Beginnings of AI at Work.” April 9, 2026. Review of the jagged technological frontier study. https://aiinstitute.hbs.edu/back-to-the-beginnings-of-ai-at-work/
  2. Dell’Acqua, Fabrizio et al. “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality.” Paper revision: March 17, 2026. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321
  3. OpenAI Academy. “Evaluate AI workflow readiness.” May 7, 2026; updated June 12, 2026. https://academy.openai.com/public/clubs/champions-ecqup/resources/ai-use-case-discovery-and-prioritizer-2026-05-07
  4. OpenAI Academy. “Gather appropriate evidence of value.” July 17, 2026. https://academy.openai.com/public/clubs/champions-ecqup/resources/gather-appropriate-evidence-of-value-2026-07-17
  5. Anthropic. “Anthropic Economic Index report: Cadences.” June 26, 2026. https://www.anthropic.com/research/economic-index-june-2026-report