In brief

Sometimes AI makes verification more prominent after it supplies a first draft or summary, but current research does not show that checking always takes longer or cancels time saved. A survey of 319 weekly users found that people described critical thinking shifting toward verifying information, integrating responses, and overseeing the task. Separate workplace studies found faster completion or changed work patterns, but did not measure verification time as a distinct task. For a workflow decision, compare human-only and AI-assisted work through the same accepted standard, including checking and rework. Exposure and changes in work patterns do not by themselves show how much verification a particular task requires.

What changes when a draft arrives before its evidence is checked?

AI can reduce the blank-page work of summarizing documents or arranging a first draft, while making checking more visible in what follows. That is a shift in the work sequence, not proof that every worker now spends more minutes verifying. The best direct evidence is a 2025 CHI study by Microsoft researchers: 319 knowledge workers who used generative AI weekly shared 936 examples of workplace use. Participants described critical thinking moving toward information verification, response integration, and task stewardship. The study also found that confidence in the tool was associated with less reported critical thinking, while task-specific confidence was associated with more. These are reported perceptions and associations, not timed observations or causal estimates of hours added.

The paper’s analysis also distinguishes source checking from broader cross-checking. It coded examples where 23 of 319 participants described checking cited sources and 114 described comparing outputs with external sources. Those counts show that verification took different forms in the survey material; they are not a population estimate of how many workers verify, how often they do it, or how long it takes. That distinction helps explain why a polished answer can still create work: a reviewer may need to locate evidence the draft did not supply and test whether the evidence supports the claim.

Consider an analyst preparing a monthly market note. Without AI, they might search reports, extract figures, and write a rough summary. With AI, they may receive a tidy first pass sooner, then check each number against the original report, confirm that a comparison uses the same period, remove unsupported claims, and make the conclusion fit the audience. Proofreading catches surface defects such as a typo. Verification asks whether a statement is true and appropriate; integration asks whether it belongs in this decision at all. The latter work can be mentally demanding even when it takes fewer clock minutes than the information gathering it replaced.

A repeatable retrieval or drafting step may be exposed to automation or assistance, while domain knowledge helps a worker spot a plausible but false figure, missing exception, or unsuitable recommendation. That remaining responsibility can be consequential, but the survey cannot establish how common the shift is across occupations, whether total workload grew, or how an employer will redesign a job. The useful conclusion here is narrower: task exposure is not a measure of verification minutes, and reported cognitive effort is not a clock-time total.

Sources: The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers

Does checking erase the time saved?

We do not have a general yes or no. The survey describes where participants felt effort moving, while other research measures outcomes such as time on email or document completion. Neither directly isolates the minutes spent checking an AI-generated first draft. Treating those measures as if they answered the same question would overstate the evidence.

A six-month randomized field experiment reported by Microsoft Research followed 6,000 knowledge workers across industries; half received access to a generative AI tool in their existing work applications. Among workers who used it, email time fell by three hours a week, or 25 percent, while the intent-to-treat estimate was 1.4 hours. Users also appeared to complete documents moderately faster; meeting time did not significantly change. This is evidence that some work patterns can shift in a particular setting. The study abstract does not say whether checking rose, whether outputs met an equal quality threshold, or whether the saved time went to other work.

A fair comparison ends at an accepted deliverable, not the instant a draft appears. For a recurring task, compare representative human-only and AI-assisted attempts through the same requirements, counting gathering, drafting, review, correction, integration, and rework. A quick but unusable report can mislead the first-draft stopwatch; a verifiable summary accepted after a short source check may save time even though verification remains essential. Record elapsed time for the whole attempt as well as the individual stages: a larger share spent checking does not necessarily mean more total time.

Make the comparison fit real conditions: use a representative batch for infrequent tasks, note unusual complications, and include time spent locating sources or repairing work for a colleague. Count work shifted to another reviewer too. The aim is a fair local comparison, not a perfect experiment. Keep the task, inputs, and acceptance criteria as similar as practical across attempts; otherwise a harder assignment or different source set may look like an AI effect.

Use that local result to make a bounded choice: retain AI when comparable accepted work takes less time without reducing quality; constrain it to a low-risk first pass when review costs vary; stop using it for that task when corrections routinely outweigh the benefit or reliable checking is unavailable. Review more carefully when errors carry greater costs. These results say nothing by themselves about whether employers are adding or reducing roles.

Sources: The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers; Shifting Work Patterns with Generative AI

What should a worker measure before changing course?

Start with the tasks you actually do, not a broad label for your work. The International Labour Organization’s 2026 brief reviews different ways researchers assess occupational exposure to generative AI and explains their limitations. Exposure indicators describe potential task effects; they do not establish whether a particular employer will adopt a tool, how a role will be redesigned, or how much verification a worker will perform. They are not measures of workflow time. Use them only as context for asking which tasks may change, then examine the work itself.

For a practical audit, choose one recurring task such as summarizing a policy update, preparing a client email, or assembling figures for a report. Define what an acceptable result must contain before comparing methods. Record the time spent gathering information, drafting, checking claims against available evidence, correcting errors, integrating the result, and handling rework. Note whether another person had to repair or approve the output. For infrequent work, use a representative batch and record complications that materially changed the task.

Then compare the human-only and AI-assisted attempts on the same quality standard. Keep AI in the workflow when accepted work improves on the measure that matters for that task; restrict it to a first pass or add checks when the result is uncertain or consequential; stop when correction and rework outweigh the useful gain or reliable checking is unavailable. The free task checker at /ai-job-risk-checker can help organize duties into task-level change-pressure signals and possible first actions, but those signals do not measure verification time or predict job loss.

Verdict: AI can shift attention from finding and composing toward checking and integrating, especially when a quick draft still needs domain judgment. Available evidence does not establish a universal increase in verification hours or prove that review cancels productivity gains. A fair answer for an individual workflow comes from observing accepted work end to end, including checks and rework. Reassess when the task, tool, stakes, evidence available, or checking method changes.

Sources: Workers’ exposure to AI: What indicators tell us and what they don’t

Questions readers ask

Is reviewing an AI draft the same as proofreading it?

No. Proofreading checks presentation, such as spelling or formatting. Verification checks claims against evidence, context, and task requirements; integration decides which material belongs in the final work. A draft can be polished and still fail those substantive checks.

Does AI exposure mean my job will require more checking?

No. Exposure measures describe potential task applicability, not actual adoption or your team’s workflow. Whether checking grows depends on the task, tool output, available evidence, quality standard, and consequences of an error. Measure recurring work through an accepted quality standard before drawing conclusions about that workflow.

Sources and notes

  1. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers

    The CHI 2025 survey reports that 319 weekly users described critical thinking shifting toward information verification, response integration, and task stewardship; it does not measure net verification time.

  2. Shifting Work Patterns with Generative AI

    This six-month randomized field experiment reports changes in email time and document completion among 6,000 knowledge workers, but its abstract does not report verification time or quality-adjusted total labor.

  3. Workers’ exposure to AI: What indicators tell us and what they don’t

    The ILO’s 2026 brief reviews differing methods for assessing occupational exposure to generative AI and their limitations. Exposure indicators concern potential task effects; they do not establish adoption, workplace redesign, verification workload, or job-loss probabilities.

Apply this to your own work

See the whole job market at once.

Explore which occupations AI may reshape, then turn the signal into a practical response.

Explore the job map