The ILO’s June 2026 review changes the useful question from “Can AI speed up a task?” to “Does that speed produce verified output, improve the wider workflow, and change how work is organized?” Its evidence includes a nationally representative survey of 5,512 Korean workers: 51.8% used GenAI at work and reported saving 3.8% of active working time, equivalent to 1.4% of total workforce hours. The weighted correlation between reported savings and reported output change was low. This is a self-reported survey result, not a direct measure of every worker’s productivity, and it does not settle longer-term effects or predict an individual role. The review also reports early demand signals for routine and entry-level work, with significant limits. Test a task carefully before deciding whether to redesign it or investigate a larger move.
What changed in the ILO’s 2026 verdict?
The review makes the evidence broader and more qualified at the same time. Published on 1 June 2026, it synthesizes experiments, firm-level data, platform studies, and worker and firm surveys across several countries. One result it reports comes from a nationally representative survey of 5,512 Korean workers: 51.8% used GenAI for work and reported saving an average 3.8% of active work time, equivalent to 1.4% of total working hours across the workforce. The weighted correlation between self-reported savings and self-reported output change was low. Because both measures are self-reported, this does not establish that each saved hour produced more usable work, nor does the survey show what happens over a longer period. The ILO review summarizes several evidence types, but it is not a causal estimate for every job or country. ([ILO review PDF](https://www.ilo.org/sites/default/files/2026-06/Research%20Brief_The%20impact%20of%20GenAi%20on%20jobs%20and%20work%20organization%20a%20review%20of%20empirical%20evidence.pdf))
“Faster” can mean a system produces a draft or a worker saves time; a team may not turn that time into more completed work. Wider adoption, labor demand, and displacement depend on service demand, staffing, and which tasks remain necessary. Evidence at one level cannot simply be promoted to the next. The ILO’s multi-source synthesis is richer than a task demonstration, but its public summary is not a causal estimate for every job or country.
One specific labor-demand signal comes from Teutloff and colleagues’ study of postings on a large online freelancing platform. They analyzed listings from January 2021 to September 2023, grouped them into 116 skill clusters with BERTopic, and used GPT-4o to classify clusters as substitutable, complementary, or unaffected by language models. Relative to the study’s counterfactual trend, demand for substitutable freelance skills such as writing and translation fell by 20–50%, with the sharpest decline in short projects; demand also shifted unevenly across complementary skills. This is evidence about posted freelance projects on one platform around ChatGPT’s release. It does not measure salaried hiring, total employment, or software jobs, and the platform-specific comparison cannot establish that GenAI alone caused a broad decline. It helps explain why the ILO treats labor-demand evidence as limited and context-specific, rather than supporting a forecast for a particular worker. ([Teutloff et al., open-access study](https://discovery.ucl.ac.uk/id/eprint/10204198/))
For a person deciding what to do next, the evidence favors curiosity with a short leash. Exposure means some tasks could be affected; it does not tell you whether your employer has adopted a system, whether customers will demand more output, or whether headcount will change. A task-level signal can justify learning a workflow or asking about redesign. It cannot by itself justify discarding accumulated experience or making an expensive retraining decision. For an early-career worker, the platform demand signals make it sensible to ask whether routine drafting or coding tasks are also the practice through which a role teaches judgment, and what supervised work could replace that practice.
Sources: The impact of GenAI on jobs, productivity and work organization: a review of the empirical evidence; Winners and losers of generative AI: Early Evidence of Shifts in Freelancer Demand
Why doesn’t a faster task settle the productivity question?
A task is only one part of a production process. Drafting a client update may take less time with a first-pass tool, yet the worker still needs to verify names, figures, commitments, and tone; revise unsupported claims; transfer the result into a customer system; and answer follow-up questions. If those steps grow, the draft got faster while the completed service may not have. Conversely, if the first pass is sound and reduces repetitive work without adding costly review, the saved time may become more completed work or more attention for exceptions. The relevant unit is usable output across the workflow, not keystrokes or generation speed.
The ILO’s May 2026 companion brief explains why observed gains at task or individual level do not automatically appear in firm, sector, or economy-wide productivity measures. Its summary points to uneven adoption and mixed firm evidence; diffusion, complementary investment, skills, reorganization, and measurement are among the conditions that affect whether local gains scale. A team can have a successful pilot without changing a whole process. A company can adopt a tool but lack the data, training, or authority to redesign handoffs. Official productivity figures can also fail to isolate the effect of one technology among many simultaneous changes. The brief outlines possible scaling conditions; it does not identify which one explains a given worker’s results. ([ILO aggregation brief](https://www.ilo.org/publications/aggregation-paradox-ai-why-do-micro-economic-productivity-gains-ai))
A specific workplace study shows why the broad and local findings can coexist. In a field study of 5,172 customer-support agents at one company, researchers examined a staged introduction of a generative AI assistant using recorded support conversations. Stanford reports an average 15% increase in issues resolved per hour, with larger gains for novice and lower-skilled workers. For the most experienced and highest-skilled agents, gains in speed were small and quality declined slightly. Thus even the positive average conceals different outcomes by experience and skill. This is a measured result for one deployment and outcome; it does not establish the effect in other companies or occupations, nor what happened to employment across the wider labor market. ([Stanford research summary](https://digitaleconomy.stanford.edu/publication/generative-ai-at-work/))
There is no contradiction between that result and the ILO’s caution. A local intervention can improve output for one group while adoption remains uneven and aggregate statistics show no clear AI-driven increase. Effects can vary within an occupation: new agents may benefit from suggested answers while veterans handle unusual cases where templates help less. The study supports the possibility of task-specific gains; the ILO sources caution against assuming they are universal, net of coordination costs, or shared with workers.
A practical comparison for your own work is therefore not simply tool versus no tool. Compare the current process with the proposed process on end-to-end time, correction or rework, accepted quality, and what happens to the time released. These are editorial decision measures, not a validated assessment instrument. For a routine report, record how long it takes to gather inputs, prepare a draft, verify it, and hand it off. If the draft is quicker but review and correction erase the saving, the workflow needs redesign or the tool is a poor fit. If quality holds and the freed time goes to analysis or client exceptions, there may be a real upgrade even if the role itself remains a mix of changed and unchanged tasks.
Sources: The impact of GenAI on jobs, productivity and work organization: a review of the empirical evidence; The Aggregation Paradox of AI: Why do micro-economic productivity gains from AI disappear at scale; Generative AI at Work
How can you test a task without treating it as a career forecast?
Start with one recurring task that has a clear input and a checkable result. Imagine a worker who writes a weekly client update from approved project notes. First record the usual time from reading the notes through final handoff, along with common edits or delays. If workplace rules allow, try a tool-assisted draft on a comparable update. Check every factual detail and commitment, count edits, and include the time spent preparing instructions and correcting the draft. Keep confidential information out of unapproved services. The question is whether the finished update is useful at an acceptable cost and quality, not whether the tool can produce fluent text.
Then distinguish a personal experiment from adoption. A single worker can learn that a particular step is faster for them; a manager or team must consider data protection, quality standards, integration, training, responsibility for errors, and what happens to workload. If the process changes, note whether the saved time is used for more volume, better service, difficult cases, or simply a faster pace. The ILO review’s discussion of autonomy and work organization makes that distribution relevant. A productivity gain does not automatically mean better job quality, a higher wage, or less work pressure. Nor does a failed trial prove that the occupation is untouched: the test may have been poorly matched to the task or tool.
Use the result to choose among proportionate moves. If a bounded task becomes easier and your broader work still depends on domain judgment, relationships, accountability, or exception handling, try redesigning the task and strengthen the checking or coordination skills that make its output usable. If several central tasks are changing, investigate adjacent work that uses your experience but has a different task mix. A larger retraining path deserves consideration when repeated changes affect the core of the role and a target path fits your salary floor, location, time, cost, health, and family needs. The review does not supply a universal threshold for that decision, so those constraints belong in the comparison rather than being waved away by a broad exposure label.
The immediate sequence can stay small: choose one permitted task; record baseline time and quality; run a limited comparison; include verification and handoffs; then decide whether to keep, revise, or stop the workflow. If your uncertainty is about which parts of your own role are exposed, the free task checker can organize those tasks into change-pressure signals and first actions. That signal is not a validated probability of displacement. If you need to compare a stay-and-redesign option with adjacent or larger-change scenarios against your actual constraints, the paid career roadmap is a next step after that task review, not a substitute for the evidence or a guarantee of work or income.
The verdict is measured experimentation before career abandonment, with attention to sustained change in the wider task bundle, especially if work redesign reduces learning opportunities or shifts accountability without support. A faster draft or broad review alone cannot resolve that question. Keep a short record of task trials, raise concrete workflow and training questions where appropriate, and revisit the decision when repeated evidence shows what has changed in your role.
Sources: The impact of GenAI on jobs, productivity and work organization: a review of the empirical evidence
Questions readers ask
Does the ILO’s 2026 review say AI is increasing productivity?
It reports uneven productivity gains and worker-reported time savings. One detailed example in the review is a nationally representative survey of 5,512 Korean workers: 51.8% said they used GenAI at work and users reported saving 3.8% of active work time. The correlation between reported savings and reported output change was low. Since these are survey reports, they do not directly measure every worker’s output or establish longer-term effects; the review is not a universal estimate for every job.
Should I change careers because my tasks are exposed to AI?
Exposure alone is not a job-loss probability and does not show employer adoption or labor demand. Test a bounded task, count verification and handoffs, and review whether several core duties are changing. Consider a pivot when that repeated evidence and your practical constraints support one, not because a task label sounds alarming.
Sources and notes
- The impact of GenAI on jobs, productivity and work organization: a review of the empirical evidence
The ILO's 2026 research brief synthesizes experimental, firm, platform, organizational, and worker and firm survey evidence. Its detailed account of a nationally representative survey of 5,512 Korean workers reports that 51.8% used GenAI for work and users reported saving 3.8% of active working time (1.4% of hours across the workforce), with a low weighted correlation between reported time savings and reported output change. The time and output measures are self-reported, so they do not directly establish how much usable output changed for each worker. The review characterizes labor-demand evidence as limited and context-specific and discusses work organization, autonomy, and job quality; findings apply to the populations, measures, and periods described, not universal or future outcomes.
- The Aggregation Paradox of AI: Why do micro-economic productivity gains from AI disappear at scale
The ILO brief says observed task- and individual-level AI productivity gains have not yet translated into measurable firm, sectoral, or macroeconomic productivity growth. Its summary describes mixed firm-level evidence and uneven adoption, and identifies broad diffusion, complementary workplace reorganization and skills, macroeconomic conditions, and other institutional conditions as relevant to scaling. It does not identify which condition explains results in a particular workplace.
- Generative AI at Work
Stanford's study summary reports that staggered access to a generative AI conversational assistant among 5,172 customer-support agents increased issues resolved per hour by 15% on average. Less-experienced and lower-skilled workers improved speed and output quality; the most experienced and highest-skilled workers had small speed gains and small quality declines. The result concerns this company, intervention, and measured work outcomes, not employment effects or other occupations.
- Winners and losers of generative AI: Early Evidence of Shifts in Freelancer Demand
The UCL Discovery record and abstract for Teutloff et al. (2025) describe analysis of job postings from a leading online freelancing platform, partitioned with BERTopic into 116 skill clusters and classified with GPT-4o as substitutable, complementary, or unaffected by LLMs. The abstract reports that demand for substitutable skills such as writing and translation declined 20–50% relative to the counterfactual trend, most sharply for short-term jobs; demand increased in complementary or unaffected clusters, with mixed results within complementary skills. This is evidence about posted freelance projects on one platform, not salaried hiring, total employment, or software jobs, and the comparison does not by itself establish that GenAI caused broad labor-market changes.
Apply this to your own work
See the whole job market at once.
Explore which occupations AI may reshape, then turn the signal into a practical response.
Explore the job map