AI support clearly helps agents resolve more cases in the studied workflows, and it may help newer agents retain useful practices in familiar work. Evidence for learning is narrower than evidence for assisted throughput: outage-period speed gains were noisy and concentrated among agents who initially engaged with recommendations. Use AI as a case aid, then check explanation, verification, and comparable unaided work before inferring independent skill.
Does AI support help new agents learn, or mainly help them resolve cases faster?
Both, but the clearest evidence is that AI helps agents resolve more cases while it is available. Evidence that it teaches broader independent skill is narrower. In a staggered rollout involving 5,172 customer-support agents at one business-software firm, the Quarterly Journal of Economics field study reported a 15% average rise in resolutions per hour. Gains were larger for less experienced and lower-skilled agents. That measures assisted output in one operation; it does not show that a new agent can diagnose unfamiliar problems or work accurately without help. The study offers a separate, suggestive learning signal: during unexpected assistant outages, exposed agents handled chats faster than before adoption, with larger apparent gains after longer exposure. The authors describe these outage estimates as noisy; chat duration is not a test of correctness, policy judgment, or transfer to a different queue. The signal was also concentrated among agents who initially adhered more to suggestions, though adherence was not randomly assigned. So treat AI as a case aid with learning potential, not proof of mastery. To judge learning in your own task mix, look beyond the dashboard: ask what an agent can explain, verify, and handle unaided on comparable, low-risk cases.
Sources: Generative AI at Work
What did the strongest direct study actually measure?
The strongest direct workplace evidence comes from “Generative AI at Work,” a study of a staggered rollout of a conversational assistant among 5,172 customer-support agents at one Fortune 500 firm selling business-process software. The assistant monitored chats and suggested replies in real time; agents remained responsible for the conversation and could ignore or edit suggestions. The central outcome was not a test score or a measure of training completion. It was the number of customer issues successfully resolved per hour, a direct operational measure of work completed during the rollout. That makes the result useful for a team asking whether assistance changes case throughput, but it cannot on its own reveal what an agent can do without help. Across the studied operation, access to the assistant increased that measure by 15% on average. That average conceals meaningful differences. The study reports that less experienced and lower-skilled agents improved across productivity measures, including a 30% increase in issues resolved per hour. For newer agents, the researchers also report that treated workers with two months of tenure performed about as well as untreated agents with more than six months of tenure. The largest gains therefore appeared among workers who started with less experience or lower measured skill, rather than as an equal boost for everyone. At the other end, the most experienced and highest-skilled agents saw little productivity effect and small declines in conversation quality. This is a useful warning against reading an average as a universal result: an assistant can change the distribution of assisted performance, while its effect on an individual depends on the task and the worker. For the learning question, keep four rungs separate. First is access: the tool is available. Second is assisted performance: an agent resolves more cases while suggestions are present. Third is retained unaided performance: the agent does better when the tool is absent. Fourth is transfer: the agent can use the underlying skill on a new product, policy, or kind of case. The 15% headline establishes the second rung in this particular support operation. The paper also reports communication patterns moving toward those of higher-performing agents and improved English fluency, particularly among international agents; these are possible learning mechanisms, but they do not show broad independent mastery. Its separate outage analysis bears more directly on retention and needs its own careful interpretation. Neither the throughput figure nor the tenure comparison tells us how quickly every new agent learns, whether the same result would hold in another queue, or whether a team will hire fewer people. The paper describes medium-run effects in one firm and explicitly says its data cannot establish aggregate employment or wage effects. So take the productivity result seriously as evidence about assisted case resolution, while resisting the leap from more resolved cases to general ability, job security, or displacement.
Sources: Generative AI at Work; Generative AI at Work
Does better performance after an outage count as learning?
An unexpected outage offers a more demanding test than a busy shift with the assistant still on screen: the agent must continue the familiar queue without receiving its suggestions. In “Generative AI at Work,” the researchers compared chat duration during large outages with agents’ pre-adoption baseline. They report that exposed agents handled chats faster during outage periods, with declines in duration equivalent to about 15% to 25%. The advantage grew with exposure: little change after one month, but faster handling after three months. This is consistent with retained learning, since immediate assistance was absent. The comparison is informative, but it is not a controlled skills exam. The outage interrupted the tool, but not the familiar queue or agents’ accumulated experience. The researchers say outages were rare, their estimates are noisy, and the chats occurring during outage periods might differ in type from other chats. Moreover, at this individual-chat level they did not have information on whether each issue was resolved. They therefore used chat duration as the outcome. Faster handling could reflect fluency or familiarity; it does not establish correct answers, sound policy choices, or diagnosis of new issues. The study’s split by initial adherence adds an important qualification. “Generative AI at Work” defines adherence as copying or entering text highly similar to a suggestion when one was provided. In the outage analysis, agents in the high initial-adherence group showed significant and rapid reductions in chat processing time relative to their own pre-adoption baseline. The low-adherence group, whose members more often deviated from suggestions, showed no such reduction even after longer exposure. Trying a suggestion may let an agent observe a customer response and retain a useful communication practice. That finding does not mean supervisors should reward copying or treat acceptance as evidence of understanding. Initial adherence was not randomly assigned. The researchers note that adherent agents may differ in other ways, or workers who gain more may be more likely to follow suggestions. Recommendations can be wrong or irrelevant; the study also found small quality declines among its most experienced, highest-skilled workers. The useful distinction is between engaging with a suggestion and obeying it. Checking account facts and policy, explaining a correction, and noticing the response reveal more than acceptance rate. For a supervisor, outage performance is therefore one corroborating signal, not a verdict. Pair it with comparable low-risk cases: check diagnosis, policy use, and whether the agent caught mismatches. Rising errors or unexplained choices weaken the claim; accurate repeated examples with explained decisions strengthen it. This practical check is an inference, not a tested training protocol. The outage result supports a bounded conclusion: some engaged agents may retain useful practices in familiar work, but speed during rare interruptions cannot establish broad mastery or transfer to unfamiliar cases.
Sources: Generative AI at Work
Why does engagement with a recommendation matter more than access alone?
Access is only the starting condition. A suggestion can shorten a reply without teaching the agent why it fits. A plausible learning mechanism requires more: the agent compares the recommendation with the case, judges whether it matches the account facts and governing policy, makes or edits a response, and sees what happens next. That sequence can turn a tool output into feedback. It can also expose a mismatch that a worker must learn to catch. The useful distinction is between exposure to a recommendation, using it, adhering to it, and understanding the reason behind it. Those are related behaviors, not interchangeable measures of skill.
In “Generative AI at Work,” adherence means copying or entering text highly similar to a suggestion when one was offered. Average adherence was 38%, with an interquartile range of 23% to 50%; agents often disregarded suggestions. The paper reports that the outage-period speed improvement was concentrated among workers with high initial adherence: their chat processing times fell relative to their pre-AI baseline, while the low-adherence group showed no comparable reduction, . The authors suggest that engaging with a recommendation and observing customer responses may help workers learn. This is consistent with learning, but it does not establish that adherence caused it. Initial adherence was not randomly assigned; agents who followed suggestions may differ in other ways, or those who benefited most may have been more likely to follow them. The paper also raises a concern: some experienced, high-performing agents increased adherence despite small or negative quality gains, suggesting possible overreliance.
Consider an example, not a reported case: an assistant proposes a troubleshooting step for a customer's account. A new agent checks whether the account actually has the feature the step assumes and whether policy permits it. If either condition is missing, the agent rejects or revises the suggestion, explains why, and then watches whether the corrected response resolves the issue. If the agent simply pastes the text, a quick successful exchange would show that the workflow worked on that occasion; it would reveal little about whether the agent can diagnose a similar case next time. If the agent can explain the fit, spot a conflict, and learn from the customer’s response, the interaction offers more evidence of developing judgment, though still not proof of transfer to unfamiliar cases.
That is why a team should encourage explain-and-verify behavior, not a high acceptance rate. Following a useful suggestion can provide practice, but a recommendation that conflicts with account facts or authoritative policy should be rejected. A supervisor can ask what the agent noticed, what evidence controlled the decision, and what would change the answer. Usage counts show contact with the tool; adherence shows similarity to its output; neither alone shows understanding. A stronger signal is a defensible choice, including justified refusal, followed by accurate unaided handling.
Sources: Generative AI at Work
What can change in the work without proving transferable skill?
A cleaner message is evidence about the message; by itself, it does not show that an agent can solve the next case without help. The Quarterly Journal of Economics study “Generative AI at Work” reports that, after the assistant was introduced, lower-skilled agents’ communication patterns moved toward those of higher-skilled agents. It also reports improved English fluency, especially among international agents. These are observable changes that may help customers follow an explanation. But they describe conversations in an AI-supported workflow. They do not directly test whether an agent can diagnose an unfamiliar fault, identify a policy exception, or transfer a response pattern to a different product.
Several outcomes can move together without meaning the same thing. An agent may write a more organized reply, respond faster, and receive a more favorable customer reaction. None alone tells a supervisor what the agent understood, which source they relied on, or whether they would choose the same action without the suggestion. A fluent explanation can repeat a sound recommendation, but fluency does not establish that it fits the account facts.
The study describes communication patterns moving toward those of higher-performing agents as suggestive evidence. That is consistent with the assistant making effective communication practices more available to newer or lower-skilled workers. It may also reflect wording supplied during assisted chats. Because the evidence comes from conversation text, it cannot distinguish a habit that persists unaided from language shaped by the live system. Similar messages do not show that a worker can independently choose the right diagnostic question or recognize when a common answer does not apply.
The fluency result needs the same care. The paper reports improvement in English fluency, particularly among international agents, based on analysis of chat language. Better fluency may make an exchange easier to follow, Yet language is only one part of support competence. An agent may still need to locate the controlling policy, confirm what happened in the customer’s account, distinguish a routine symptom from an exception, and know when to escalate. A tool can help shape a clear sentence while leaving those decisions unresolved.
Consider a routine troubleshooting exchange in which the assistant proposes a concise explanation and next step. A newer agent might learn a useful way to explain it. The harder test comes when an account detail conflicts with the usual pattern: does the agent notice, check the relevant policy, and pause before sending a polished but unsuitable answer? This is an illustration, not a case reported by the study. It shows why surface quality and independent diagnosis need separate observations.
When someone says AI “trained” staff, ask what capability changed, whether it was observed without assistance, and whether correctness held on cases beyond the examples that shaped the suggestion. Improved wording during assisted chats supports a claim of better assisted communication. Accurate diagnosis and sound policy use in comparable unaided work would be stronger evidence of retained skill. Transfer to unfamiliar cases requires its own evidence.
Sources: Generative AI at Work
Why are speed, customer outcomes, and learning different comparisons?
Speed, customer experience, and agent learning answer different questions. In the randomized field experiment reported in “Engaging Customers with AI in Online Chats: Evidence from a Randomized Field Experiment,” agents at a meal-delivery company received AI-generated reply suggestions. The publisher's abstract reports faster responses, deeper engagement, and improvement in customer sentiment, with the largest benefits for less-experienced agents. This is evidence about assisted interactions. It does not show what agents could do after suggestions were removed, so it cannot establish retained skill.
The study shows why service metrics need context. AI improved efficiency and sentiment for subscription-cancellation conversations, while repeat complaints tied to systemic problems were least responsive. A well phrased reply may acknowledge a recurring delivery failure, but cannot resolve its operational cause. The study also reports a counterintuitive handoff: when a customer had experienced chatbot comprehension failure, rapid replies from an AI-assisted human agent were associated with worse sentiment. Customers could take the speed as a sign they were still talking to a bot. Faster response, then, did not necessarily feel like better service. This pattern comes from one meal-delivery operation; it does not show that speed generally harms trust or that every company should slow replies.
A second comparison comes from “Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations,” a 2026 arXiv preprint about e-commerce after-sales chat. Human agents were randomly assigned access to an assistant offering an opening diagnosis and proposed solution, which agents could adopt, modify, or ignore. Its abstract reports faster service and improved subjective quality measured by customer ratings, but no significant effect on objective quality measured by customer retrials. Lower-performing agents benefited most, while top performers had declines in subjective and objective quality, a pattern the authors describe as consistent with workflow disruption. This cautions against treating favorable ratings as interchangeable with changed customer actions, or assuming every agent benefits alike. As a preprint abstract in a different setting, it adds evidence to examine, not a settled rule for support teams.
The Harvard Business School AI Institute's “The Fast-Talking AI Chat Agent” explains the meal-delivery experiment; it is not a separate trial. It highlights 138 agents and more than 250,000 conversations, and summarizes faster messages and larger gains for newer agents. Those details describe study scale, but the tenure comparison still concerns assisted performance. It is not proof that a new agent acquired five months of independent competence.
A team should track response speed, correct resolution, customer sentiment or rating, repeat contact, and unaided performance on comparable work separately. The first four describe service outcomes; the last begins to address retention and needs checks of diagnosis and policy use. None of these studies tests transfer to unfamiliar products or exceptions. Pair throughput with resolution correctness, customer outcomes, and case context, then assess learning separately. A successful assisted workflow shows that it can help deliver service; it does not by itself show that an agent is independently skilled or that the role is redundant.
Sources: Engaging Customers with AI in Online Chats: Evidence from a Randomized Field Experiment; Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations; The Fast-Talking AI Chat Agent
What work does an assistant remove, and what checking work does it add?
An AI assistant can remove some of the effort involved in typing, recalling standard information, or drafting a routine reply. That does not make the whole support case automatic. An agent may still need to understand what the customer is asking, retrieve the relevant account facts, find the policy that governs the case, inspect the suggested answer, correct it when details do not match, communicate the decision, document the exchange, and escalate when the issue falls outside the agent’s authority. The balance depends on the task and the system; the available qualitative evidence describes one setting rather than a universal pattern. The study “Customer Service Representative’s Perception of the AI Assistant in an Organization’s Call Center” reports a field visit and semi-structured interviews with 13 customer service representatives at a power-grid service call center. Participants described AI as easing traditional burdens such as typing and memorizing, while also bringing new learning, compliance, and psychological burdens as they adapted to the system. This is evidence about how workers at that site perceived the change, not a measured estimate of how often those effects occur across support teams. The study does not establish that every assistant removes the same steps, or that the new burdens outweigh the old ones. That distinction matters when a team looks at faster replies or more cases handled. A shorter drafting step can coexist with additional checking. A suggestion may be fluent yet fail to reflect an account detail or the controlling policy. The work then includes noticing the mismatch, deciding what evidence or rule takes precedence, changing the response, and taking responsibility for what reaches the customer. For a new agent, checking is also part of the work they need time to learn. If speed targets leave no room to compare a generated suggestion with account facts and policy, an agent can produce an answer without demonstrating that they understand why it fits. Conversely, correcting a suggestion should not automatically count as poor performance: when the correction is grounded in the case and an authoritative rule, it can show that the agent is checking rather than copying. This is a practical interpretation of the reported adaptation burdens, not a validated training method. A team practice is to make that checking visible. On a routine, low-risk case, ask the agent to identify the relevant policy, explain whether the suggestion matches the facts, and say what would trigger escalation. Supervisors can review a small sample of those decisions while preserving existing safeguards for consequential or unclear cases. The aim is not to add a lengthy audit to every reply. It is to avoid assuming that faster drafting means training has become unnecessary. New agents need permission and time to verify, correct, document, and ask for help; otherwise a tool that reduces one burden may conceal whether the underlying judgment is developing.
Sources: Customer Service Representative’s Perception of the AI Assistant in an Organization’s Call Center
How can a team tell whether a new agent is learning?
A practical learning check should ask whether the agent can explain and verify a decision when the assistant is not supplying the next sentence. Choose a small set of comparable, low-risk cases from one routine category. For each assisted case, ask the agent to state the customer’s issue, identify the policy or account fact that governs the answer, and explain why the suggested response fits. Then review some similar cases with suggestions hidden. Use the same standards in both conditions: correct diagnosis, policy accuracy, appropriate escalation, and a clear explanation. This is a bounded practice inferred from the evidence, not a validated training program. The reason to compare assisted with unaided work is that the strongest workplace learning signal is also limited. In “Generative AI at Work,” a staggered rollout among 5,172 agents at one business-software support firm raised resolved issues per hour by 15% on average. During unexpected outages, exposed agents handled chats faster than before adoption, with stronger apparent gains after longer exposure. The authors describe these outage estimates as noisy, note that outages were rare and case mix may differ, and measure chat duration rather than a complete skills assessment at this level. It supports checking retention, not assuming transfer to unfamiliar policies or products. So the review should examine reasoning as well as outcomes. If a recommendation says to reset an account, the agent should be able to point to the relevant account condition and policy, and notice when either conflicts with the suggestion. Record a correction as evidence that the agent checked the output, then verify whether the final action was accurate. Repeating a polished sentence from the assistant is weaker evidence of learning than explaining why it applies and when it would not. Other support research helps keep the test proportionate. “Engaging Customers with AI in Online Chats” reports a randomized meal-delivery field experiment in which AI suggestions improved response efficiency and customer sentiment overall, while results varied by conversation type; it did not test later unaided skill. Interviews with 13 representatives at one power-grid call center, reported in “Customer Service Representative’s Perception of the AI Assistant in an Organization’s Call Center,” surfaced both relief from typing and memorization and new learning and compliance burdens. These interviews describe one site, so they support attention to checking time rather than a claim about prevalence. For stable, repetitive, low-risk work, a team may reasonably prioritize assisted throughput and reserve unaided checks for a sample, exceptions, or scheduled coaching. A new agent can ask a team lead to choose one routine case category and review a few assisted and unaided examples using the same rubric. If the agent cannot explain a recommendation or spot a mismatch, treat that as a prompt for coaching, not a verdict on overall readiness.
Sources: Generative AI at Work; Engaging Customers with AI in Online Chats: Evidence from a Randomized Field Experiment; Customer Service Representative’s Perception of the AI Assistant in an Organization’s Call Center
What is the next realistic move for a new support agent?
For a new support agent, use AI on a defined, low-risk case type while practicing how to check its work. Agree with a supervisor on the controlling policy source, what to do when a suggestion conflicts with account facts or policy, and how comparable cases will be reviewed. This supports service while building evidence of what the agent can handle independently. The study “Generative AI at Work” offers a reason to test learning, not assume it: its outage analysis suggests some agents retained faster handling in familiar work, especially those who initially engaged with suggestions. But the authors describe estimates as noisy, and chat duration is not a test of accuracy or judgment. Handling a routine case unaided, explaining the relevant policy, and identifying when to escalate would provide more meaningful evidence than speed or suggestion acceptance. This check is not a validated training protocol. If checks show accurate independent handling across routine and somewhat varied cases, confidence in retained learning can grow. If the agent cannot explain a recommendation, or errors carry material costs, treat that as a coaching signal and keep human review. These studies describe task-level changes; they do not estimate an individual’s chance of losing a job. The free task checker maps change-pressure signals in your role, but does not measure learning or predict personal displacement.
Sources: Generative AI at Work
Questions readers ask
What is the strongest evidence that AI support helps new agents learn?
In one field study, exposed agents handled chats faster than their pre-adoption baseline during unexpected assistant outages. The authors describe those estimates as noisy, and chat duration does not establish accuracy, policy judgment, or transfer.
Does higher AI-assisted case throughput prove an agent has learned?
No. Higher resolutions per hour measures assisted output while the tool is available. It does not show what an agent can do independently or on unfamiliar cases.
Should support teams reward agents for following AI suggestions?
Adherence was associated with the outage-period speed signal in one study, but it was not randomly assigned. Teams should encourage agents to explain and verify suggestions, including rejecting ones that conflict with account facts or policy.
How can a new support agent check whether they are learning?
Review a small sample of comparable, low-risk assisted and unaided cases using the same standards for diagnosis, policy accuracy, escalation, and explanation. This is a practical inference from the evidence, not a validated training program.
Sources and notes
- Generative AI at Work
Supports the staggered-rollout findings on assisted throughput, worker differences, communication and fluency, and noisy outage-period retention signals.
- Generative AI at Work
Provides the accessible research record for the field study of an AI conversational assistant used by customer-support agents.
- Engaging Customers with AI in Online Chats: Evidence from a Randomized Field Experiment
Publisher abstract supports findings on response efficiency, customer sentiment, conversation context, and differences by agent experience.
- Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations
Accessible preprint abstract reports randomized access and differing subjective ratings and objective retrial outcomes in after-sales support.
- Customer Service Representative’s Perception of the AI Assistant in an Organization’s Call Center
Accessible abstract reports interviews at one power-grid call center describing reduced typing and memorization burdens alongside new adaptation burdens.
- The Fast-Talking AI Chat Agent
University research explainer summarizes the meal-delivery experiment's scale and assisted interaction findings; it does not establish retained learning.
- How will AI change customer service?
Opened evidence review discusses the QJE outage signal and adherence condition, while leaving a practical unaided-skill check unresolved.
Apply this to your own work
See the whole job market at once.
Explore which occupations AI may reshape, then turn the signal into a practical response.
Explore the job map