The productivity debate surrounding generative AI is often reduced to a single question: Does an employee complete a particular task faster with AI? Recent research answers this question affirmatively for many tasks. But individual speed and organizational productivity are not the same thing. Writing an email in five minutes instead of ten can save time; however, if easier production leads to twice as many emails being written, recipients having to read more text, and inaccurate details requiring review, the system’s overall gain may shrink.
What does the research actually show?
In a six-month field experiment published by the NBER in 2025 and involving 7,137 knowledge workers at 66 organizations, some participants were given a generative AI tool integrated into email, meeting, and writing applications. Employees who actively used the tool spent about two fewer hours per week on email during the second half of the experiment and worked less outside normal business hours. However, the researchers found no clear overall change in the amount or composition of work performed. This result does not mean failure: Doing the same work with less spillover beyond regular hours is an important benefit. But it shows that time savings do not automatically translate into a greater volume of output.
The effect can be stronger in other tasks. A combined analysis of three controlled experiments involving a total of 4,867 software developers at Microsoft, Accenture, and a large company reported a 26.08 percent increase in tasks completed by developers using an AI tool. The results nevertheless varied across the experiments. Moreover, the number of completed tasks alone does not measure software reliability, maintenance costs, or the value delivered to customers.
Another experiment involving 776 professionals at Procter & Gamble found that individuals using AI could approach the performance of two-person teams working without AI in some product development activities. The notable point was not only speed but also the partial crossing of expertise boundaries: Technical and commercial employees were able to produce more balanced solutions concerning the other field. Even so, the outcome of a controlled innovation task cannot be directly generalized to the entire organization’s long-term performance.
Why do individual gains fail to benefit the organization?
The first reason is that the cost of production falls while the cost of consumption remains unchanged. AI can prepare a long report, presentation, or message in a few minutes, but the people on the receiving end still need time to read it, verify it, and turn it into a decision. The five minutes saved by the sender may cost ten recipients a total of fifty minutes. Local optimization can therefore become an organization-wide loss.
The second reason is verification debt. Text that appears fluent may contain incorrect figures, fabricated sources, or recommendations that do not fit the context. An employee produces the first draft quickly, but if review has not been clearly defined, the error is carried into later stages. Especially in legal, financial, health, safety, and customer commitments, speed indicators must be kept separate from quality indicators.
The third reason is that capacity is immediately filled with new demand. If a team can prepare 120 reports per week instead of 100, management may soon treat 120 as the new normal. Rather than being devoted to rest, learning, or more important problems, the time saved turns into an expectation of greater output. In this case, the tool does not ease the employee’s workload; it merely increases the pace of work.
The fourth reason is the coordination burden. Numerous AI tools, differing templates, and individual patterns of use can fragment shared working practices. If no one knows which model produced the output, what data was used, who verified it, and who holds final decision-making authority, faster draft production creates more meetings and corrections.
How should healthy productivity be measured?
Organizations should not track only “how many people use AI?” or “how many hours were saved?” Measurement should cover at least four layers: the time required to complete the task, the output’s error and rework rate, the review time the process imposes on other employees, and the value of the resulting business outcome. In a customer support team, first-contact resolution and repeat-contact rates matter as much as a fast response; in a software team, defects, rollbacks, and the maintenance burden matter as much as completed tasks.
The most reliable approach is to establish small, comparable pilots. First, a specific workflow is selected, and baseline values for time, quality, work outside regular hours, and waiting are recorded. The same measurements are then repeated during a four- to eight-week AI-assisted period. User satisfaction is assessed separately, because an apparently identical output may have been achieved with less strain on the employee. That, too, is a genuine gain in productivity and sustainability.
Both the stage delegated to AI and the decision point retained by humans should be documented clearly. For example, the tool may generate a draft list of actions from a meeting recording, but it should not transfer them to the task system until the participants have approved the responsible people and deadlines. It may draft an email, but high-impact commitments should not be sent automatically. Such a design manages the cost of errors as well as speed.
Conclusion: Who owns the time saved?
Current evidence does not say that generative AI is ineffective; it says that its impact depends on the task and the design of the work. In some jobs, completed output increases substantially; in others, the main benefit is less time spent on email or less work outside regular hours. Both outcomes are valuable, but they should not be assessed using the same unit of measurement.
The fundamental management question should therefore be not “How quickly does the tool write?” but “Where does the time saved go within the system?” If that time is redirected toward better decisions, less overtime, learning, and high-value tasks, AI becomes a genuine productivity tool. If it produces more messages, higher quotas, and a growing verification queue, it has merely moved the bottleneck.