Google's AI Patent Tool Sharpens Senior Lawyers but Leaves Juniors Behind
A Google field experiment with 133 patent lawyers shows AI boosts output for everyone, but only seniors convert that help into lasting skill.
- Google Research ran a 3-month RCT with 133 patent attorneys on an unreleased AI patent drafting tool.
- AI access raised drafting quality by 0.34-0.38 SD, roughly 10-11 percentile points, for everyone.
- On an unassisted redlining task, seniors improved by 0.45 SD, juniors showed no average gain.
- Junior scores split: more very low and more good scores, no new excellent ones.
- Seniors used AI as a "logic auditor"; juniors stayed stuck in surface-level copy-edits.
- Full paper on NBER; tool is now part of Gemini Notebook.
A Patent-AI Trial Finds Uneven Gains in Professional Judgment
A three-month randomized field experiment by Google Research found that an AI assistant improved patent drafting while the tool was available. A separate evaluation conducted without AI produced different results: senior lawyers showed stronger independent judgment, while junior lawyers recorded no statistically detectable average gain. For teams deploying professional copilots, assisted output alone cannot reveal whether workers are developing durable expertise.
MIT economist David Autor led the study while serving as a visiting fellow at Google. The findings are available in an NBER working paper. The research combines sustained use during ordinary professional work, blinded grading by domain experts, and a tool-free assessment designed to separate immediate productivity from independent skill.
AI and the apprenticeship bottleneck
Senior professionals develop judgment through years of drafting, revising, receiving feedback, and handling routine work under supervision. AI assistants can remove some of that practice by producing a plausible first draft or executing a revision before a junior worker has worked through the underlying problem.
Earlier research offers mixed evidence. Studies in radiology, legal education, business problem-solving, and job-seeker writing have found independent gains when AI is integrated into structured training. Experiments involving software engineers, consultants, students, and clinicians have also found that higher AI-assisted performance may fade when access ends.
Inside the three-month trial
Researchers randomized access to a then-unreleased Google Labs patent-writing assistant among 133 lawyers at 11 intellectual property firms with regular, non-exclusive business relationships with Google. About two-thirds of the lawyers at each firm received early access. The control group received AI training but could not use the assistant until the study ended.
Lawyers in the treatment group could use the assistant during normal work. The researchers also administered three controlled evaluations:
| Timing | Task | AI access | Purpose |
|---|---|---|---|
| Day 10 | Draft a patent from simulated inventor materials | Allowed for treatment group | Measure early assisted performance |
| Day 90 | Draft a patent from a different invention packet | Allowed for treatment group | Measure assisted performance after sustained access |
| Day 90 | Redline and correct a flawed hypothetical patent | Prohibited for everyone | Measure independent judgment |
Independent patent professionals graded the submissions without knowing each lawyer’s treatment assignment. They scored enforceability, accuracy, completeness, clarity, and strategic ambiguity, meaning the deliberate scoping of claims to preserve useful legal breadth.
The redlining exercise tested whether lawyers could identify and repair defects such as “patent profanity,” language that can unnecessarily narrow rights or weaken enforceability. Participants had to execute the corrections themselves rather than describe what an AI system should change.
AI raised drafting quality
| Measure | Estimated effect | Interpretation |
|---|---|---|
| Day-10 assisted drafting | +0.34 standard deviations | About 10 percentile places among control-group scores |
| Day-90 assisted drafting | +0.38 standard deviations | About 11 percentile places among control-group scores |
| Day-90 unassisted redlining, all lawyers | +0.32 standard deviations | Average effect across experience levels |
| Day-90 unassisted redlining, senior lawyers | +0.45 standard deviations | Clear gain among experienced practitioners |
| Day-90 unassisted redlining, junior lawyers | No detectable average effect | Scores became more dispersed |
A standard deviation expresses the performance gap relative to variation in the control group. In this study, the 0.34 and 0.38 estimates corresponded to gains of roughly 10 and 11 percentile places.
The assisted-drafting gains came from fewer poor submissions and more good ones. The share of excellent submissions remained unchanged, indicating that the assistant improved consistency without increasing the frequency of top scores.
Junior lawyers also completed the day-10 drafting task about 18 minutes faster than the control group’s 124-minute average, a reduction of roughly 15%.
Experience showed up without the tool
On the unassisted redlining task, lawyers who had received assistant access outscored controls by 0.32 standard deviations. Senior lawyers accounted for that result, recording a 0.45-standard-deviation gain. Junior lawyers showed no statistically detectable average improvement.
The junior score distribution also widened. The treatment group produced more very low scores, fewer middling scores, more good scores, and no increase in excellent scores. Those shifts suggest heterogeneous outcomes, although the design lacks a comparable pre-study redlining baseline that could identify which individuals improved or declined.
Senior lawyers rewrote the work
Qualitative review found that junior submissions in both groups often followed a rigid top-to-bottom sequence. Some lawyers spent substantial time copy-editing low-stakes introductory sections, reached the core claims late, and substituted synonyms without improving legal scope. Even when they identified serious defects, they sometimes left diagnostic comments instead of implementing the required correction.
Experienced lawyers allocated their effort differently during the tool-free exercise. Treated seniors spent longer on redlining than juniors, bypassed low-value prose, rebuilt claims, removed language that could narrow legal rights, and connected edits to specific legal doctrines.
In follow-up interviews, senior lawyers described using AI as a “logic auditor.” Reviewing generated text weakened their attachment to existing language and prompted them to explain the legal reasoning behind structural changes. Their prior expertise gave them a framework for interrogating the output and practicing judgment during assisted work.
A better copilot scorecard
Teams deploying AI in law, engineering, medicine, research, or analysis can use the study’s design to evaluate production gains and professional development separately:
- Add an unassisted benchmark. Run periodic, time-boxed exercises without AI and score them blind when feasible.
- Track distinct outcomes. Measure assisted quality, completion time, independent performance, error severity, and review effort.
- Segment results by experience. Report score distributions for junior and senior workers because a single average can conceal gains and losses at the tails.
- Preserve deliberate practice. Pilot workflows that require junior staff to draft, diagnose, or execute a revision before viewing generated suggestions.
- Instrument usage patterns. Where privacy and professional rules permit, record which suggestions workers accept, reject, rewrite, or escalate for review.
- Test workflow changes. The experiment did not evaluate specific training safeguards, so first-draft requirements and explanation prompts remain hypotheses to validate.
What the study cannot settle
Several features limit how broadly the results can be applied:
- Small, specialized sample: The trial involved 133 patent lawyers, whose work and training may differ from those of other professionals.
- Selected firms: All 11 firms had business relationships with Google, which may affect generalizability.
- Short duration: Three months captures early learning, while professional judgment develops over years.
- Simulated evaluations: The controlled tasks do not measure litigation outcomes, issued claim scope, client satisfaction, or long-term commercial value.
- Access-based estimate: Randomization covered access to the assistant; individual usage patterns could produce different learning effects.
- Changing technology: Model capabilities and lawyers’ baseline familiarity with AI continue to evolve.
Researchers did not release code, model weights, or the standalone experimental assistant, limiting direct replication. The paper says the tool’s capabilities have since been incorporated into Google’s NotebookLM.
AI teams can strengthen rollout evaluations by adding an independent-skill endpoint. Assisted quality and time savings describe current production, while blinded tool-free assessments measure whether judgment is accumulating across experience levels. This trial found clear production gains, stronger independent performance among senior lawyers, and unresolved learning outcomes for juniors.