In brief

A faster first pass does not by itself mean a work task has improved. Compare repeated, similar tasks with and without AI using the same input, acceptance standard, and review process. Record total person-time through checking, correction, and handoff; assess quality separately against criteria that matter for the task. Research finds cases where speed and quality improve together and cases where AI adds time or reduces correctness. The practical evidence is therefore local: keep AI in a particular workflow when accepted quality holds or improves and the full task takes less effort, or when another explicit benefit justifies the tradeoff. Do not turn that result into a claim about every task in a role or a person's job prospects.

What counts as a fair test of faster work?

Judge a workflow at the point an acceptable deliverable is ready, not when a draft first appears. The comparison should hold the task, input quality, acceptance bar, and review standard steady. Count the worker’s time to prepare context, operate the tool, inspect and correct its output, and hand off the result. Keep first-pass time separate: it shows where effort moved, not whether effort fell.

Quality also needs an independent reading. Depending on the work, assess correctness, completeness, usefulness to the next person, compliance, or acceptance without substantial revision. A fluent answer can omit a required condition; a longer answer can be less usable. Apply the same rubric to both workflows, ideally with a reviewer unaware of which produced the item. For consequential work, record error severity as well as error count.

Choose criteria before seeing results, and make them reflect what makes the deliverable useful rather than what the tool produces easily. A summary might be judged on whether it preserves required facts and caveats, while a customer response may need correct policy and account details. A single overall rating can hide a serious omission inside otherwise polished work. The acceptance bar should identify which defects require correction and which make the item unusable, so a reviewer is not forced to improvise standards after learning which workflow produced it.

Noy and Zhang’s online experiment illustrates that speed and assessed quality can improve together in a bounded setting. The accessible MIT working paper reports 444 college-educated professionals completing incentivized, occupation-specific writing tasks. On the post-treatment task, participants given ChatGPT took 10 fewer minutes than the control group’s 27-minute average and received evaluator grades 0.45 standard deviations higher. These were short tasks with limited context-specific knowledge; the working paper does not establish sustained workplace or organization-wide productivity. [The accessible MIT repository copy](https://economics.mit.edu/sites/default/files/inline-files/Noy_Zhang_1.pdf) reports the design and its limits.

Sources: Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence

Why can results reverse from one task to the next?

The contrast across experiments is a reason to read the setting before applying a headline. A 2026 Organization Science paper reports a preregistered experiment with 758 BCG consultants using GPT-4. Across 18 selected knowledge tasks, participants with AI completed 12.2% more tasks, worked 25.1% faster on average, and delivered solutions rated significantly higher in quality. On a separate complex managerial task categorized as outside the tested capability frontier, AI users were 19 percentage points less likely to produce a correct solution. This is evidence of contrasting outcomes in those selected tasks, not a reliable advance rule for identifying every task AI will handle well. [The open paper](https://www.hbs.edu/ris/Publication%20Files/dell-acqua-et-al-2026-navigating-the-jagged-technological-frontier_5c589c8c-fbb5-458f-b285-c944746cd717.pdf) describes the study and its scope.

METR supplies a different counterexample. Its randomized trial covered 246 tasks for 16 experienced open-source developers in mature repositories they knew well. With early-2025 AI tools available, task completion time increased by 19%, despite participants estimating afterward that AI had reduced their time. Screen recordings showed prompting, waiting, review, and cleanup as parts of the work that a draft-speed measure would miss. The authors caution that experimental artifacts cannot be entirely ruled out; the result does not establish a general slowdown for developers or later tools. [METR’s full report](https://metr.org/Early_2025_AI_Experienced_OS_Devs_Study-paper.pdf) details the trial.

These results are not directly comparable: the occupations, tasks, systems, and quality measures differ. Together they show why a single benchmark or impression cannot settle a local workflow decision. Verification burden can be material, and apparent speed may coexist with lower correctness; that is the practical distinction to carry into a particular task’s evaluation.

The METR result also warns against treating confidence as a time measure: participants’ retrospective estimate pointed in the opposite direction from recorded task completion. That discrepancy does not prove people systematically misjudge AI, but it shows why an impression should be reported separately from elapsed effort. Nor does the trial supply a general quality estimate to combine with the BCG quality ratings or the writing grades. Each study answers a narrower question under its own scoring method; their numbers should not be averaged into a pooled promise or penalty.

Sources: Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality; Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

What should you measure before changing your workflow?

Interpret the result using separate readings: person-time to accepted output, task-relevant quality, rework or error correction, and the share of cases meeting the existing standard. Count completed, accepted items if volume matters. A workflow is a productivity improvement only when its output remains fit for purpose and total effort falls; if it is slower, another benefit may still justify it, but name that benefit rather than calling the result faster work.

A useful decision rule distinguishes a hard constraint from a preference. If an output fails a mandatory accuracy or compliance requirement, a time saving does not compensate for that failure. If it meets the bar but requires more review, decide whether the added time buys a named benefit, such as a better starting point or a more consistent format, and whether that benefit matters enough for this task. A vague sense that the result is more polished is not a substitute for stating the tradeoff.

Small trials also have limited power to reveal rare but costly errors. Report how many cases were observed and preserve examples of misses, corrections, and accepted outputs, not just an average. A clean run on routine cases may justify continued testing, while an error with serious consequences may be enough to narrow use immediately. The evidence supports a proportionate decision under uncertainty, not a universal numeric cutoff: define what failure matters before the trial and revisit the choice when inputs, tools, or responsibilities change.

NIST’s voluntary AI Risk Management Framework offers measurement guidance, not evidence that AI raises productivity and not a prescribed personal trial. Its Measure function calls for appropriate metrics and benchmarks, documented limitations and uncertainty, repeatable assessment, and ongoing evaluation. The task comparison in this article is a practical synthesis of those principles, not a validated universal instrument. [NIST’s Measure guidance](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/) explains the framework’s role.

For a team decision, the useful next conversation is specific: “Can we agree what acceptable quality looks like and compare total time through review for a few similar cases?” If one task result sits within a wider career question, the free [task exposure checker](/ai-job-risk-checker) organizes task-level change-pressure signals; it does not estimate a probability of job loss.

Sources: AI RMF Core: Measure

Questions readers ask

Is a faster AI-generated first draft proof of higher productivity?

No. It is one intermediate measure. Include setup, checking, corrections, and handoff, then separately assess whether the accepted deliverable meets the same quality standard.

What quality measures should I use?

Choose criteria tied to the deliverable’s purpose, such as correctness, completeness, usability, required compliance, or acceptance without major revision. Record serious errors separately from minor edits.

Do task-level AI productivity results tell me whether my job is at risk?

No. They show how a particular workflow performed under tested conditions. They do not establish market-wide adoption, labor demand, job redesign, or an individual probability of displacement.

Sources and notes

  1. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence

    The accessible MIT working paper (identified as not peer reviewed) reports a preregistered online experiment with 444 college-educated professionals doing occupation-specific incentivized writing tasks. On the post-treatment task, the treatment group took 10 fewer minutes than the control group's 27-minute average and received evaluator grades 0.45 standard deviations higher. The paper describes 20–30-minute tasks; this supports a bounded finding about the experiment's writing tasks, not sustained workplace or organization-wide productivity.

  2. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality

    The accessible Organization Science paper reports a preregistered experiment with 758 BCG consultants. Its abstract reports improved task quantity, speed, and quality on selected tasks within the tested capability frontier, while on a separate task selected as outside that frontier AI users were 19 percentage points less likely to produce a correct solution. This is evidence from the paper's selected experimental tasks and setting.

  3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

    The accessible METR randomized controlled trial reports 16 experienced developers completing 246 tasks in familiar open-source repositories, randomly assigned to allow or disallow AI use. For the early-2025 tools and study setting, allowing AI increased completion time by an estimated 19%; the result is bounded to this sample, task set, repositories, and period.

  4. AI RMF Core: Measure

    NIST's voluntary AI Risk Management Framework Measure function calls for appropriate metrics, documented test and evaluation methods, repeated assessment and updates, deployment-relevant evaluation, documentation of limitations, and ongoing monitoring. This is risk-measurement guidance, not evidence that AI improves productivity and not a prescription for the article's particular task trial.

Apply this to your own work

See the whole job market at once.

Explore which occupations AI may reshape, then turn the signal into a practical response.

Explore the job map