In a preregistered experiment published in Science, Shakked Noy and Whitney Zhang assigned occupation-specific writing tasks to 453 college-educated professionals and randomly exposed half to ChatGPT. Average time fell by about 40%, and judged output quality rose by about 18%. Inequality between workers decreased. The result is a clean benchmark for assistive writing—not a claim about every knowledge-work domain.
What they measured
Participants completed incentivized, occupation-specific writing tasks designed to resemble real midlevel work—press releases, grant cover letters, sensitive emails, analysis plans, consultant-style reports. Half were given access to ChatGPT.
Time and quality were both measured. Quality came from independent evaluation of outputs, not self-report alone.
Who gained the most
Lower-performing workers saw larger quality gains, which compressed inequality on these tasks. That pattern matters for how teams roll out assistants: the average effect can hide who benefits.
Exposure also raised later self-reported use on the job, suggesting learning-by-doing—not only a one-off lab effect.
What the study does not claim
These were writing tasks with clear deliverables. Results do not automatically transfer to tool-using agents, regulated decisions, or long-horizon projects. Treat the paper as strong evidence for drafting and editing workflows with human judgment still in the loop.
How to use the finding
If your bottleneck is first-draft speed and baseline quality on similar writing, an assistant is a rational default. Still measure acceptance and review time in your own queue—the cost-of-retries note applies.
Cite the primary source when you brief stakeholders: Science, DOI 10.1126/science.adh2586.