In brief

A faster task or a larger pile of output shows throughput, not necessarily better work. Evidence of improved quality needs a separate outcome tied to the deliverable: correctness, completeness, usefulness, recipient success, or another standard that matters in the role. The strongest practical comparison holds the task and review standard steady, assesses the output independently where possible, and counts correction, verification, and downstream rework. Experiments show that AI can improve both speed and rated quality on bounded writing and consulting tasks, but other evidence is speed-only, narrowly tested, or points to worse results on tasks beyond a tool’s capability or inside context-heavy workflows. For your own work, define the quality floor first, compare a small set of similar tasks with and without assistance, and keep the use only if the deliverable meets that floor after review costs are counted.

Why is faster output not enough?

Speed and quality answer different questions. A timer asks how long a task took. A counter asks how many drafts, tickets, code changes, summaries, or analyses were produced. A quality check asks whether the completed work was correct, complete, useful, safe for its intended use, and acceptable to the person who relies on it. These measures can move together, but one does not prove the other. When a study reports that participants finished faster, the supported conclusion is about task time under that study’s conditions. Calling the result “better work” requires an additional measure of the work itself.

The distinction sounds simple until a dashboard collapses it. Imagine a support queue where an agent closes more conversations per hour. That can reflect faster handling, more cases reaching a resolution, easier cases being selected, or a change in what counts as closed. It does not, by itself, tell us whether customers received accurate answers, whether they had to reopen cases, or whether difficult problems were handed elsewhere. Likewise, more code lines may express a useful feature, duplication, or a verbose implementation that behaves exactly like the shorter version. The unit being counted matters, and the count needs a quality condition beside it.

A controlled Copilot experiment illustrates the limit of a speed result. Microsoft Research’s page describes recruited software developers implementing a JavaScript HTTP server as quickly as possible. Those given the tool completed that task 55.8% faster than the control group. The experiment is useful evidence that access changed completion time in this specific exercise. Its summary does not report a separate output-quality outcome for the task. It therefore cannot establish from that number alone that the server was more correct, easier to maintain, or more valuable to a user. Speed is a real finding, but it is not a synonym for quality.

The difference also matters to the worker whose apparent productivity rises. If a first draft takes ten minutes rather than thirty, but a colleague must spend twenty minutes checking facts and repairing omissions, the producer’s clock has improved while the full work system may not have. That example is a measurement problem, not a claim that AI routinely creates such rework. Conversely, a draft that arrives sooner and needs no extra correction may be a genuine improvement even if the final output count remains unchanged. A useful evaluation must follow the work far enough to see whether the saved effort survives contact with review and use.

For a career decision, this prevents a common leap: “The tool can produce this task” becomes “the task is now better” and then “the occupation is disappearing.” Each statement needs different evidence. Technical capability describes what a system can do under specified conditions. A controlled task result describes performance in an experiment. Workplace adoption asks whether organizations actually use it. Demand and displacement concern labor-market outcomes. None follows automatically from a timer. If a tool changes one task, that is a reason to inspect the task and its quality criteria, not a personal job-loss estimate.

The count is especially weak when tasks differ in difficulty or value. Ten routine answers and one careful answer to a high-impact question are not interchangeable units. If the tool changes which work gets attempted, the numerator itself may shift. A fair report states what counts as an output, whether the task mix changed, and what happened to work that remained unresolved.

Sources: The Impact of AI on Developer Productivity: Evidence from GitHub Copilot

What should a fair quality comparison hold constant?

A useful comparison begins by deciding what “good” means before seeing the results. For a client memo, that might include factual accuracy, coverage of the requested issues, clear reasoning, and a recommendation the client can act on. For a customer reply, it might mean the underlying issue is resolved, the answer follows policy, and the customer does not have to repeat the problem. For code, it can include passing relevant tests, understandable structure, security review, and fit with the existing system. The criteria should be visible in the actual job, not chosen afterward to make an appealing result look successful.

Next, compare like with like. Use the same task family, comparable inputs, information access, deadline, and deliverable definition. Randomly assign similar tasks to assisted and unassisted conditions if there are enough of them; for a small personal trial, alternate conditions or pair tasks that are genuinely comparable. Do not compare an AI-assisted routine request with a manual edge case and attribute the difference to the tool. If the work varies sharply, record the task type and evaluate each type separately. An average across unlike tasks can hide the very cases where assistance helps or harms.

Assessment should be independent of the production method where feasible. A reviewer who does not know which condition produced a document is less likely to reward polished phrasing simply because it sounds confident. Blinding is not always possible, and it does not remove every source of bias, but it can reduce one obvious cue. Use a qualified reviewer or the real recipient when that is appropriate. For operational work, an objective check may complement human judgment: whether a transaction is correct, a requested issue stays resolved, or a code test passes. No single method captures every quality dimension, so combine only the checks relevant to the task.

Record the complete workflow, not just the moment an answer appears. That includes task framing, prompt preparation, locating source material, checking claims, editing, handling failures, getting approval, and correcting work after delivery. Keep production time and review time distinct so you can see where effort moved. Record both mean or median quality and the share of outputs that clear a minimum acceptable floor. A high average can conceal a small number of serious failures; a floor can reveal whether the workflow is dependable enough for that use. The floor should reflect consequence: an awkward internal draft and a mistaken financial instruction do not carry the same cost.

Noy and Zhang’s preregistered writing experiment and GitHub’s controlled code-quality exercise show why comparison design matters: both assessed work beyond mere completion time, though with different populations, tasks, and rubrics. METR’s study, by contrast, focuses its primary outcome on task duration in developers’ familiar repositories. Reading those designs side by side suggests a practical protocol; it is not a protocol tested as a package by any one of them. Choose one recurring task, define two or three observable criteria and one unacceptable failure, collect comparable examples with existing approved materials, and have someone qualified review them without being told the condition if possible. That small test can guide a workflow choice. It cannot validate a job-level score or predict employment outcomes.

A baseline matters too. Compare assisted work with the quality people actually produce without assistance, not with an idealized manual process or an assumed best effort. If a team already uses templates, search, or peer review, those supports belong in the comparison. Otherwise the apparent treatment effect may partly reflect removing existing help from the control condition.

Sources: Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper); Does GitHub Copilot improve code quality? Here’s what the data says; Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

Can quality rise while time falls?

Yes. A preregistered online experiment by Shakked Noy and Whitney Zhang later appeared in Science. The accessible MIT working-paper version, which predates the journal publication and is marked not peer reviewed, reports occupation-specific, incentivized writing tasks with 444 college-educated professionals; half were randomly given ChatGPT access. It reports 37% less task time and blinded evaluator grades 0.45 standard deviations higher. Because the researchers measured both time and the work product, this result supports a narrower and more useful statement than “AI made people productive”: on those bounded writing exercises, participants completed work faster and evaluators rated the output more highly.

Random assignment helps interpret the difference between the study’s groups because access was not simply chosen by participants who were already more comfortable with the tool. The tasks were designed around participants’ occupations and included incentives, giving the exercise more work-related structure than a casual writing prompt. But “work-related” is still not the same as sustained performance in a real employer’s workflow. Participants completed short, defined tasks online. The study does not tell us whether the work held up through weeks of client use, whether a supervisor accepted it, or whether the added quality generated revenue, pay, more hiring, or fewer jobs.

The study also reports that lower-performing participants benefited more, with the quality gap between lower- and higher-performing participants narrowing. That finding describes variation within this task set. It should not be converted into a promise that AI will close skill gaps across occupations. Baseline differences, task design, the available model, and evaluation criteria shape the result. The finding does, however, caution against assuming that the most experienced worker necessarily receives the largest immediate benefit, or that everyone benefits by the same amount. Where a task has a clear brief and a reviewer can compare its result to a standard, assistance may help a worker get beyond a blank page or organize a first version. Whether that mechanism explains the result in a particular job needs its own evidence.

The right inference is an existence claim: there are bounded professional writing tasks where access to a generative tool improved both measured speed and assessed quality in a randomized experiment. It is stronger than a time-saved survey for this question because the study included an assessment of output. It remains weaker than evidence of consistently better organizational work across all stages. The quality measure belongs to the researchers’ task and rubric; a law firm, research team, or communications department may have different standards for accuracy, confidentiality, originality, and review.

For a worker, this is a reason to test a writing use case when the task is repeatable and the acceptance criteria are clear, not a reason to outsource judgment. A draft of a meeting summary may be easy to compare with notes and a transcript. A policy interpretation that affects an individual may require current authoritative sources and accountable review. Both are “writing,” but the cost of being wrong and the evidence needed to establish quality differ. The experiment shows that a quality gain is possible; it does not remove the need to specify whose quality, on what task, and over what period.

A second caveat is that a 0.45 standard-deviation increase is measured on the study’s evaluator-grade distribution, not a 0.45-point gain on every workplace quality scale. The abstract supports the reported result but does not make its measure interchangeable with a customer outcome or a code test. Keep the original unit and context attached whenever citing it.

Sources: Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper)

What counts as a direct quality measure?

A direct quality measure checks an output against a stated standard. It can be a functional test, an accuracy check, a blinded expert rating, a recipient outcome, or a rubric for completeness and usefulness. Each catches different failures. A test may confirm that a program returns the expected answer for selected inputs while missing confusing structure or behavior outside those cases. A reviewer may notice poor explanations but miss a subtle defect. A customer rating may reflect tone and experience, not the technical accuracy of every claim. Direct measures are stronger when the method matches the consequence and when their limits are visible.

GitHub’s report on a controlled Copilot experiment gives a layered example. It says the company recruited developers with at least five years of experience; 202 developers submitted valid work, with half randomly assigned access to Copilot. The task was to write API endpoints for a web server. The code was evaluated through ten unit tests, then blind developer reviews. GitHub reports that participants with access were 53.2% more likely to pass all ten tests. In blind review, Copilot-assisted code had fewer readability errors, and the company reports improvements across several rubric dimensions and approval likelihood. These outcomes ask more about the artifact than a timer does.

The layers still answer different questions. Passing ten tests is evidence of functionality on those tested cases, not proof that software is secure, accessible, or maintainable under every future change. A blind review can evaluate qualities that tests miss, but it depends on the reviewers, criteria, and code they see. GitHub reports a 13.6% increase in lines written without readability problems, as well as other small rating differences. A statistically detectable rating difference is not automatically important enough to change a release decision. Practical importance depends on the task, baseline failure rate, review burden, and whether users notice or benefit from the difference.

The source is also a vendor’s own research report about its product. That does not make its measurements unusable; it means readers should keep the sponsorship and setting attached to the claim. The task was a single fictional restaurant-review web server, not a mature application with years of dependencies and operational history. The blind-review phase included code that passed all tests, which makes it informative about readability among functional submissions but does not represent all generated attempts. The company’s proposed explanation for the improvements should be treated as its interpretation, not independent confirmation of a general mechanism.

Translate the same layered logic into your work. A market analyst can check whether a summary’s figures match the source data, whether the requested segments are covered, and whether a colleague can reproduce the conclusion. A recruiter can check whether a screening summary accurately reflects the stated criteria, while separately reviewing fairness and policy compliance. A project manager can assess whether a generated status note is current, includes dependencies, and helps a decision-maker act. The measure need not be elaborate. It does need to test what the recipient depends on, not merely whether the output looks finished.

A useful rubric also states how reviewers resolve disagreement. If one reviewer values brevity and another completeness, the average hides a standards problem. A short calibration discussion using sample work can clarify what “complete” means before the comparison begins. Keep examples anonymous and use only materials that can be handled under the organization’s data rules.

Sources: Does GitHub Copilot improve code quality? Here’s what the data says

Why does the task boundary matter?

A measured quality gain on one task does not transfer automatically to a task that looks similar. The 2026 Organization Science paper by Fabrizio Dell’Acqua and coauthors, developed with Boston Consulting Group, makes this boundary visible. Its preregistered experiment involved 758 knowledge workers completing realistic consulting tasks. Participants were assigned to no AI, GPT-4 access, or GPT-4 with a prompt-engineering overview. On 18 tasks selected as within the tool’s capability frontier, the authors report that participants with AI completed 12.2% more tasks, finished 25.1% faster on average, and delivered solutions of significantly improved quality.

The same paper tested a complex managerial task designated outside that frontier. Participants using AI were 19% less likely to produce correct solutions than those without AI. The contrast matters more than treating the positive average as a general return. The study describes a jagged boundary: assistance can improve performance on some tasks and worsen it on another task in the same broad knowledge-work setting. The teams used realistic and varied exercises, but these remained selected tasks in a controlled experiment. Neither result is an occupation-wide rate, and neither says what would happen in a company’s complete process with its own review, data, and incentives.

Why can a boundary appear inside one workflow? One task may have information that is easy to express in a prompt and an answer that can be checked. Another may depend on an implicit constraint, institutional history, or a subtle interaction among facts that the system does not handle reliably. If a mistake is easy to spot, a worker can reject the suggestion. If it is polished and plausible, the person may accept it without noticing what is missing. These are explanations consistent with the study’s contrast, not separate causal findings established by it. The general lesson is to test tasks at the level where their inputs, checks, and consequences differ.

This has a practical consequence for people whose job titles obscure a mixed task bundle. A financial analyst may draft a routine description, reconcile numbers, explain an anomaly, and advise a manager. An HR specialist may summarize a policy, handle an unusual accommodation request, and document a decision. A project coordinator may turn known updates into a status note, then negotiate a dependency that no document captures. A single “AI exposure” label or one successful trial on the easy part cannot answer whether the full role’s quality improved. Break the workflow into tasks with distinct standards and consequences.

When the boundary is uncertain, the response should be proportionate to error cost. Use AI to structure low-risk first drafts or generate questions to check, while keeping a qualified person responsible for consequential conclusions. Where correctness is easy to verify, test the tool against that check. Where errors are hard to detect or costly, demand stronger evidence, preserve independent review, or do not use the output for that purpose. The consulting study supports the claim that task fit matters; it does not supply a universal checklist that identifies the frontier in advance. A small, observable comparison can locate some boundary conditions in a worker’s setting, but managers and workers still need to monitor them as tasks and tools change.

The reported task-level reversal is not a reason to assume that any unfamiliar work should be avoided. It is a reason to treat confidence as an input to oversight rather than proof of correctness. The more difficult the output is to verify, the less a fluent answer should substitute for independent evidence.

Sources: Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality

What does service quality look like in live work?

In a live service operation, a completed case is closer to an outcome than a draft count, but it still needs interpretation. The widely cited customer-support study by Erik Brynjolfsson, Danielle Li, and Lindsey Raymond analyzed a staggered introduction of a generative AI conversational assistant among 5,172 customer-support agents at one software company. The Quarterly Journal of Economics abstract reports an average 15% increase in productivity measured as issues resolved per hour, with substantial differences among workers. The study’s published abstract also describes outcomes beyond volume, including resolution and customer experience. That allows a more useful question than “Did agents handle more?”: what happened to the resolution and experience of the people seeking help?

The study reports that less experienced and lower-skilled workers improved both speed and quality of their output, while the most experienced and highest-skilled workers saw small speed gains and small quality declines. It also reports customer conversations becoming more polite and customers becoming less likely to ask for a manager. These are distinct signals: productivity per hour, an assessment of the service interaction, and an observable customer response. The differences across worker groups complicate a simple average. An outcome can improve for some participants and fall for others, even inside a single organization using one deployed assistant.

This evidence is more workplace-like than an isolated writing prompt because the researchers observed a tool rollout in an operating support setting. Its staggered introduction supports comparison over time and across agents, but it is still one firm, one product context, and a particular customer-support workflow. “Issues resolved per hour” is the study’s productivity measure; it should not be paraphrased as a 15% improvement in every dimension of service. Customer politeness and requests for a manager are relevant experience indicators, but they do not certify that every answer was factually correct or that each problem stayed resolved after the interaction. A measure may be valuable without being complete.

For a support team, a fuller evaluation would pair handling volume with the proportion of cases resolved under a defined rule, later reopenings, escalation patterns, policy errors, and recipient feedback. Which measures exist depends on the service and its systems. If easy cases are resolved more quickly while difficult cases wait longer, a raw hourly rate can hide a distributional shift. If customers rate friendly language highly, the measure can still miss an incorrect instruction. Tracking the task mix and failure severity helps distinguish genuine service improvement from a change in what gets counted or who carries unresolved work.

The career implication is modest but important. An employee should look for evidence about recipient outcomes and the full work process, rather than infer job quality from a company’s productivity claim. A tool that helps newer agents may change how expertise is used, but this single study does not establish that employers will reduce headcount, expand support, or alter wages. Nor does it show what the same system would do in a regulated service with different consequences. The evidence shows why outcome measures belong beside throughput and why average results should be checked against the tasks and workers they combine.

This distinction between operational outcomes and quality guards against a misleading headline. Even an outcome close to the customer can be a proxy, and a proxy needs a clear definition. A reopened case may signal an unresolved issue, but some cases reopen for unrelated reasons; ratings can reflect wait time or tone as well as accuracy. Interpret several indicators together where available.

Sources: Generative AI at Work (Stanford Digital Economy Lab research record)

A woodworking bench displays illustrated steps for assembling and inspecting a wooden box, with a hand holding a magnifying glass beside blue plans and tools.
A woodworking bench displays illustrated steps for assembling and inspecting a wooden box, with a hand holding a magnifying glass beside blue plans and tools.

Whose quality standard is being measured?

A quality score is meaningful only in relation to someone’s need. The person producing the output may prize speed and polish; the colleague who checks it may care about traceable sources; a customer may need a clear resolution; an organization may require compliance and consistent documentation. These standards can overlap, but they can also conflict. If a manager evaluates only the volume an employee produces, the worker may have an incentive to create more material even when recipients must sort through it. If a customer rating is the only measure, employees may be pushed toward pleasing language at the expense of a complete or technically accurate answer.

A 2025 Stanford study provides context about worker preferences, not a causal test of output quality. Researchers surveyed 1,500 U.S. workers across 104 occupations about where AI might help or cause harm and interviewed 52 AI experts about current capabilities. The report says workers wanted automation particularly for repetitive tasks and preferred to retain agency and oversight. It describes a gap between what workers wanted and expert assessments of what systems could do. Those responses reveal what people value and expect; they do not show that AI-assisted work actually met those preferences or improved a recipient outcome.

That distinction is useful because “quality” can become a proxy for whoever has authority to define it. A productivity project may favor what is easiest to count. A worker may care that a tool removes repetitive formatting while leaving enough time to think. A client may care that the recommendation accounts for an exception. The Stanford survey cannot settle which standard should govern any particular job, and the university’s summary includes labor-market interpretations that are outside this article’s evidence question. It does support treating worker goals and oversight as relevant inputs to evaluation, while keeping preference evidence separate from performance evidence.

A minimal rubric can make the question concrete. Identify the deliverable’s recipient and the action they must take. Then write a few observable checks: is the information accurate against its source; does the output cover each requested point; can the recipient act without asking for missing context; did the work respect confidentiality and policy? Add a high-consequence failure condition that means the output cannot pass regardless of its style. The criteria should be specific enough that two reviewers can discuss a disagreement, not so broad that every polished response receives a high score.

Include the worker and reviewer in deciding what to measure when possible. A trial can improve one person’s pace by moving verification to a less visible colleague. That may be an organizational transfer of labor rather than an improvement in the whole process. Ask who has to check the output, who absorbs the risk, and whose feedback counts. Protect client, employee, and company information by using only tools and materials approved for that data. These questions do not prove whether a tool is good or bad; they make the quality claim answerable for the people affected by the work.

Preference information can also guide which tradeoffs deserve evaluation. If workers want repetitive formatting handled but want to retain control over advice, a quality rubric should distinguish those activities instead of scoring an entire role as one block. That does not establish what the tool can reliably do; it helps frame the task and the human role worth testing.

Sources: What workers really want from AI

What happens when the workflow is context-heavy?

A tool can make an isolated task look easier while adding time to a workflow that depends on accumulated context. In 2025, METR ran a randomized controlled trial with 16 experienced open-source developers completing 246 tasks in mature repositories they had worked in for an average of five years. The study assigned tasks to permit or prohibit AI tools, which primarily included Cursor Pro and Claude 3.5 or 3.7 Sonnet. Before starting, developers expected AI access to reduce completion time by 24%; after the study, they still estimated a 20% reduction. Measured completion time moved in the opposite direction: the authors report a 19% increase.

The authors discuss why the result differs from short, self-contained coding exercises. Participants worked in projects they knew well, with established conventions and quality expectations. The paper notes that lab tasks can omit prior context and familiarity, and that measures such as added code lines or task counts can be misleading when code is verbose or work is split into smaller pieces. A suggestion that looks plausible may take time to fit into a repository, test against dependencies, or reverse if it conflicts with how the system behaves. The trial is evidence about elapsed time in this setting, not a direct estimate of code quality.

Its limits are substantial. Sixteen developers are a small sample, and they worked in particular mature open-source projects using tools available during February to June 2025. The study’s own authors say experimental artifacts cannot be ruled out entirely, although they find the slowdown robust across their analyses. The finding is a boundary case, not an estimate for all programmers or all office workers. It is also not proof that AI degraded the final deliverables; the headline measure was time. What it does challenge is the assumption that a tool’s demonstrated ability on a benchmark or simple exercise guarantees a speed benefit inside expert work.

For experienced workers, context is part of the task even if it is absent from the prompt. A compliance analyst may know which version of a policy applies and which exception needs escalation. A researcher may know why a source is considered weak despite its polished summary. A software engineer may understand a dependency no task ticket mentions. The tool can contribute a draft or candidate answer, while the worker pays to reconstruct the background, check hidden assumptions, and integrate the result. Whether that extra effort is worthwhile depends on the complete deliverable and on benefits that matter to the worker, not on a generic claim that AI is faster.

A sensible test includes the part where context usually costs time. Do not benchmark only the first output. Measure how long it takes to understand the task, direct the tool, verify its result, fit it into existing work, and complete the normal review. Compare quality separately. If the tool takes longer but improves coverage or reduces a difficult error, that may still be a worthwhile trade for a particular task. If it adds inspection without raising quality or reducing a meaningful burden, the speed claim has not materialized for that workflow. The data should inform a local choice rather than discredit or endorse the technology in the abstract.

A time result can still be decision-relevant even when quality is not directly measured. If a task must meet the same external standard and the full process is faster, that may matter. But the quality condition then needs separate confirmation from ordinary acceptance checks. The METR trial helps demonstrate why participants’ expectations and their measured time can diverge.

Sources: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

How should review and rework enter the score?

The first output is an intermediate state, not always the finished work. A polished draft can contain a wrong date; a functional code sample can fail under an untested input; a resolved support ticket can reopen. To judge net quality, follow the artifact through the checks that the role normally requires and, where possible, into use. Record what passed at first review, what needed substantive correction, what failed after delivery, and who performed the additional work. Count total time across producer and reviewer when the question is whether the workflow improved, while preserving separate times so a transfer of effort remains visible.

It helps to separate consequential defects from taste. A reviewer may prefer a different heading or phrasing without the output becoming less useful. A missing data caveat, unsupported factual assertion, broken requirement, or policy violation is different. Decide in advance what counts as a substantive correction and how failure severity will be recorded. The goal is not to build a complex scoring system; it is to prevent every edit from being treated as evidence that the first draft was poor or every polished first draft as proof that the task succeeded. Keep the measure aligned with the real acceptance standard.

The studies in this article do not supply one common cross-industry rate of AI-related rework. GitHub’s code study examines unit tests and blind review in a bounded exercise. The customer-support research uses operational outcomes in one company. METR measures task duration in mature repositories. Noy and Zhang assess short writing outputs. Those endpoints cannot be combined into a universal “net quality” number. A practical accounting frame is therefore a recommendation drawn from their differences: look past initial output and define the checks that reveal whether the work remained good in its setting. It is not a claim that any specific quantity of hidden correction is occurring.

A team can often begin with existing quality-control signals rather than inventing new surveillance. For customer support, that could mean approved measures already used for case resolution and reopening. For document work, it may be normal editorial review or correction logs. For code, it could include test failures and ordinary review comments. Avoid collecting personal data or turning a short workflow trial into an employee ranking exercise. The evidence question concerns the tool and task under defined conditions; an individual’s output can vary with workload, experience, task difficulty, and access to information.

Use a short, bounded review period with a clear stop rule. If the tool produces an unacceptable error, increase supervision or remove that task from scope. If the same quality floor is met and total correction effort falls, the use case has stronger practical support. If a result is mixed, split the task into smaller parts and test those with distinct standards. This makes quality control part of using the tool responsibly, not a final formality. A worker can also document which review skill the workflow requires, since the ability to trace claims, test outputs, and recognize exceptions may become more relevant even when the tool handles the first draft.

A first-pass acceptance rate can be reported beside severity-weighted defects, but avoid combining them into one opaque score unless the weights have a defensible basis. A critical omission should not be canceled out by many minor style improvements. Simple counts and a short description of what each defect means are often more useful for a local decision.

Sources: Does GitHub Copilot improve code quality? Here’s what the data says; Generative AI at Work (Stanford Digital Economy Lab research record); Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

What can current studies support beyond the task?

A study can show that a defined task improved under tested conditions without proving that an entire organization became more productive. Moving from task performance to a lasting workplace result requires several additional steps: the organization must adopt the tool for relevant tasks, redesign a workflow or assign the saved capacity usefully, maintain quality, and see demand for the resulting work. These are separate questions. Capability is not adoption; access is not routine use; faster completion is not increased output that customers value; a better product is not automatically a higher wage or a safer occupation.

The International Labour Organization’s June 2026 review helps set that boundary. It synthesizes experiments, firm-level data, platform studies, and representative worker and firm surveys from Australia, Denmark, Germany, Korea, Kuwait, the United Kingdom, and the United States. Its summary says productivity gains are real but often unverified and uneven. It also says worker-reported time savings amounting to a few percent of working hours have not yet translated into higher measured output, earnings, or employment. This is a broad empirical review, but its landing page does not report a single common quality measure or pooled causal estimate. It cannot tell a particular worker whether their work improved.

The ILO finding is not a contradiction of the task experiments. A randomized study can detect a local treatment effect under a set design, while broad output, income, or employment depends on how many firms adopt the tool, which tasks they apply it to, what new work is created, what bottlenecks remain, and how gains are distributed. One level of evidence can be positive while another remains uncertain or unchanged. The article’s quality ladder is a way to avoid treating these different levels as interchangeable: output volume is one rung, artifact quality another, recipient outcomes another, and economy-level consequences a separate inquiry.

For an early- or mid-career worker, the implication is to treat a quality result as evidence for a bounded work choice. It may justify testing an approved tool on a recurring task, learning to verify outputs, or documenting where domain knowledge changes the result. It does not tell you to abandon experience, buy a particular credential, or move to a supposedly protected occupation. Those choices depend on your actual task mix, employer context, location, salary needs, available learning time, health, and family responsibilities. Labor-market changes deserve attention, but they require labor-market evidence, not a task experiment repurposed as an individual forecast.

The answer could change as evidence improves. Longer field studies that follow output through customers, errors, reviewer time, workflow redesign, and staffing would support stronger conclusions about durable work quality. Replications across employers, task types, countries, and tools would help show where results transfer. Until then, the well-supported statement is narrower: AI has improved assessed quality alongside speed in some bounded tasks and operational contexts, while capability boundaries and context costs can reverse the result. That is enough to motivate practical measurement; it is not enough to promise broad returns.

The review also reports uneven effects and broader organizational risks, but those claims should not be turned into an unsupported forecast for an individual. A multi-country synthesis can describe patterns across evidence types; it does not erase differences in occupation, employer, regulation, or economic conditions. A person’s own next step should stay within what the task evidence can support.

Sources: The impact of GenAI on jobs, productivity and work organization: a review of the empirical evidence

What is a proportionate next step?

Pick one recurring task whose quality you can inspect and whose failure cost is manageable. Name the recipient, describe the deliverable, and write two or three checks before using AI. For example, a research brief might need every key number to match its source, each requested question to be addressed, and an uncertainty to be stated rather than concealed. Decide what failure means the output cannot be used. If the work contains confidential or sensitive information, use only a tool and process your organization approves; a quality experiment does not justify exposing data.

Collect a small set of comparable tasks with and without assistance. Keep the input, deadline, and scope as similar as practical, and note which condition was used. A personal test does not need to be a formal trial to be useful, but it should avoid cherry-picking the easiest assisted example and the hardest manual one. Include the time spent directing, checking, editing, and handing off the work. If an independent reviewer is available, ask them to score outputs against the prewritten criteria without being told which process produced each one. If blinding is not practical, record that limitation rather than treating the review as neutral.

Look at the quality floor before the average. Did the work meet the required standard? Were there factual or functional mistakes? Did a recipient need clarification or correction? Did review work move to another person? Then look at total effort and speed. A use case can be worth keeping if it produces a better result at comparable effort, or the same reliable result with meaningfully less effort. It may also be useful if it gives a worker more time for a higher-value task, but that benefit should be named and checked rather than assumed. If the tool fails the quality floor, constrain it to a lower-risk step, add checks, or stop using it for that task.

After the test, write down a decision in plain terms: continue under these conditions, redesign and retest, or drop this use. Also note what could change the decision, such as a new task type, a revised policy, a different tool, or a change in the recipient’s needs. One worker’s trial can guide that worker’s workflow. It cannot establish an occupation-wide quality effect, an employer’s adoption plan, future demand, or a probability of redundancy. The most useful skill may be the ability to define the standard, check the result, and explain where the tool did or did not help, grounded in experience that the task actually requires.

If the next question is which parts of your own role are most exposed to task change, the free AI Proof Work checker can help organize that task-level review. Its result is a transparent change-pressure signal, not a validated probability of job loss. The evidence question in this article comes first: decide what better work means, then test whether the relevant task meets that standard. Start with one task, inspect its finished result and total review effort, and make a decision you can revisit when the work changes.

If no reviewer is available, use a verifiable check where possible and be candid about what remains subjective. Keep a record of examples rather than relying on a general impression after a busy week. A sample that includes normal variation is more informative than a demonstration built around a single unusually successful prompt.

Sources: Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper); The impact of GenAI on jobs, productivity and work organization: a review of the empirical evidence

Questions readers ask

Does completing more tasks prove AI improved work quality?

No. More completed tasks or shorter completion time establish throughput under the measured conditions. Quality needs a separate measure tied to the deliverable, such as correctness, completeness, recipient outcome, or a relevant expert rubric.

What is the strongest evidence that AI improved work quality?

Evidence is strongest when a controlled or carefully matched comparison assesses the actual output against a relevant standard, preferably with independent or blinded review, and also counts verification and downstream correction. No single metric proves every dimension of quality.

Can I use a task quality result to predict whether my job will be replaced?

No. Task performance does not establish employer adoption, labor demand, or displacement. A task-level result can guide what workflow to test or what verification skill to practice, but it is not a personal job-loss probability.

Sources and notes

  1. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot

    Supports the specific 55.8% faster result on a speed-focused JavaScript HTTP-server exercise; its summary reports no separate quality outcome.

  2. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper)

    Supports the accessible primary working-paper version reporting 444 professionals, 37% less task time, and blinded evaluator grades rising 0.45 standard deviations; it predates the final Science publication.

  3. Does GitHub Copilot improve code quality? Here’s what the data says

    Supports the vendor-reported controlled code-quality study using unit tests and blind developer reviews, with limits of one exercise and vendor authorship.

  4. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality

    Supports a preregistered 758-worker consulting experiment with quality gains on 18 within-frontier tasks and worse correctness on one outside-frontier task.

  5. Generative AI at Work (Stanford Digital Economy Lab research record)

    Supports the study description for 5,172 support agents, 15% more issues resolved per hour, heterogeneous worker quality, and customer-experience indicators.

  6. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

    Supports an RCT of 16 experienced contributors and 246 familiar-repository tasks reporting 19% longer completion time; it does not directly measure broad quality.

  7. What workers really want from AI

    Supports the description of a survey of 1,500 U.S. workers and interviews with 52 experts; preferences are not measured quality effects.

  8. The impact of GenAI on jobs, productivity and work organization: a review of the empirical evidence

    Supports the ILO's June 2026 synthesis that gains are uneven and often unverified and time savings have not yet translated into higher measured output, earnings, or employment.

Apply this to your own work

See the whole job market at once.

Explore which occupations AI may reshape, then turn the signal into a practical response.

Explore the job map