What AI Drafting Does to Draft Quality
Written from NBER Working Paper 35720, September 2026. A working paper, not peer-reviewed, and not law. Six of its seven authors work at Google, whose drafting tool the trial studies. Current as of 29 September 2026.
In a pre-registered randomized trial, 133 patent lawyers at eleven firms were randomized, and 90 of them got a custom AI drafting assistant. Draft quality rose at ten days and at ninety. When the tool was taken away, the retained advantage belonged entirely to the senior lawyers.
Most claims about AI and patent drafting are anecdotes. This is a randomized controlled trial, pre-registered before the data existed, with every output scored by blinded expert patent attorneys. Its result is less comfortable than either side of the argument usually admits.
The design is why it is worth reading. It measured two different things: how good the work is while the AI is available, and how good the lawyer's own judgment is afterwards with the AI removed. Most workplace AI studies measure only the first.
The authors' own summary sentence is the one to sit with. "The largest gains from AI thus accrued to the lawyers who retained the least."
The Two Measurements
| What was measured | When | Result | Who gained |
|---|---|---|---|
| Work delivered with the assistant, on benchmark drafting tasks | 10 days | 0.34 SD (p = 0.03) | "larger gains among junior lawyers" |
| Same, later in the trial | 90 days | 0.38 SD (p = 0.01) | Same pattern |
| Redlining an application by hand, without the assistant | After 3 months | 0.32 SD (p = 0.04) overall | Senior lawyers only, 0.45 SD (p = 0.02) |
| Same test, junior lawyers alone | After 3 months | No average gain | Scores spread rather than shifted |
The final test is described by the authors as "a core task of patent practice requiring expert judgment." Everyone did it unaided, and the authors checked rather than assuming: they screened the submissions for signs of AI use and, "[e]rring toward over-classification," flagged "15 of 91 redlining submissions in our sample as possible non-adherence." Controlling for those flags "leaves the results intact"; dropping them entirely costs precision but, in their words, "does not overturn our main result."
Speed Points the Same Way
The abstract is about quality. The body of the paper also measures time, and the pattern is the same one.
| Task | Time saved with AI | Against a baseline of |
|---|---|---|
| 10-day drafting task | 10 minutes (p = 0.05) | 112 minutes |
| 90-day drafting task | 10 minutes (p = 0.27), not significant | 125 minutes |
| Junior lawyers, 10-day task | 18 minutes | Same task |
| Senior lawyers, either task | None. "Senior lawyers did not work faster with AI." | Same tasks |
So the juniors were the ones who got faster and the ones who scored higher while the tool was available, and they are the ones who kept nothing once it was gone. Two independent measurements, same group, same direction.
The Bifurcation Is the Interesting Part
That is a different finding from a null result, and it is the one with a practical edge. A tool that moves some people up and others down is a supervision question, not a purchasing question.
The authors offer an explanation and mark it as a hypothesis rather than a conclusion, using their own modal verb: "Foundational expertise may be a prerequisite for extracting durable skill from AI-assisted practice."
Who Ran It
Six of the seven authors give Google affiliations. The acknowledgements thank "as well as the Google InFlow team" and five named individuals at Google, and the lead academic author acknowledges support from "the Google Technology and Society Visiting Fellows Program." The assistant studied is Google's.
That is not a reason to discount the result. The study was pre-registered in the AEA RCT Registry in May 2025, which constrains what the authors could report after seeing the data, and the finding that lands hardest is the one least flattering to an AI drafting product. It is a reason to state who ran it whenever the numbers are quoted, which is why this page does.
We have our own interest to declare on the same page. We publish a local drafting tool and a benchmark of 51 models. That benchmark measures draft quality against a panel of scorers; it says nothing about what happens to a practitioner's judgment over three months, and this trial is the only evidence here that does.
What This Does Not Decide
- It is a working paper. NBER working papers are not peer-reviewed, and a published version may differ.
- One tool, one profession, ninety days, 133 subjects at firms that agreed to take part. It does not establish what a different tool, a local model, or a longer exposure would do.
- It measures drafting quality as scored by attorneys. It says nothing about whether anything is patentable or will be granted.
- It does not tell any practitioner what their professional obligations are. Those are set out by the Office, not by an economics paper.
- Senior and junior are the paper's own categories: it classifies "junior lawyers as those with fewer than seven years of experience, while those with seven or more years of experience are classified as senior lawyers." The authors vary that threshold and report that the choice does not change the conclusions.
Educational, not legal advice. Reported here as evidence about drafting, not as guidance about compliance.
Sources
- David Autor, Tanya Rodchenko, Josh Martin, Zanna Iscenko, Scott Strand, David Pearl and Melissa Ferere, Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting, NBER Working Paper 35720 (September 2026)
- Pre-registered in the AEA RCT Registry, AEARCTR-0015823, 21 May 2025
- Our own benchmark: drafting with a local model