Google Research trial: AI sharpens senior lawyers' judgment, not juniors'

Google Research trial: AI sharpens senior lawyers' judgment, not juniors'

On October 7, 2026, Google Research published a blog post by David Autor (Technology & Society Visiting Fellow) and Tanya Rodchenko (Principal, AI & Economy) on a three-month randomized controlled trial with practicing patent attorneys. The headline finding: routine AI use lifts work quality across the board, but its effect on on-the-job learning depends on seniority. Senior lawyers who used AI for 90 days showed stronger judgment than those who did not. Junior lawyers showed no average skill gain, and their scores split into more strong scores and more low scores.

The authors frame the question against mixed prior evidence. Studies in radiology, business problem-solving, job-seeker writing and legal education suggest that AI embedded in training workflows can help less-experienced workers perform better on their own. Experiments with software engineers, management consultants, high school students and clinicians suggest AI acts as a temporary "exoskeleton" whose gains often fail to persist once the tool is removed. The authors say a clean answer needs three things together: sustained AI use in ordinary work over months, a credible test of unassisted judgment, and blinded grading by domain experts.

The design is described in the NBER paper "Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting". The researchers randomized access to a then-unreleased Google Labs AI patent writing assistant, now part of Gemini Notebook, among 133 lawyers at eleven intellectual property law firms that do regular, non-exclusive business with Google. Two-thirds of the lawyers at each firm formed the treatment group and got early access. The rest, the control group, received some training on using AI but were withheld from the tool until the three months were over.

After 10 days, all lawyers received a packet of inventor materials for a hypothetical invention and were asked to draft a patent. They repeated this at 90 days with a different set of simulated materials. Third-party patent professionals scored submissions on five dimensions: enforceability, accuracy, strategic ambiguity (the tactical scoping of claims), completeness and clarity. The expected AI boost appeared both times. At 10 days, tool access raised drafting scores by 0.34 standard deviations, equivalent to a 10-point climb in percentile ranking among control group scores. By 90 days the gain was 0.38 SD, an 11-percentile point increase. It was consistent across all quality dimensions. Better performance came from fewer poor scores and more good scores, with no change in the frequency of excellent ones: AI-assisted lawyers moved from the bottom into the lower-middle and middle quintiles. The pattern was clearest for junior lawyers, whose quality gains came with time savings, for example 18 minutes faster than the control group's average of 124 minutes on the 10-day task.

To test judgment, the team added a task at 90 days. Subjects had to redline (manually mark up and correct) an existing hypothetical patent riddled with substantive and stylistic errors, including the kind of mistake sometimes called "patent profanity" that can make a patent unenforceable, such as overclaiming novelty or scope. No one was allowed to use AI on this task, so scores reflect how well skills had developed. With the assistant removed, the leveling effect vanished. Lawyers with AI access outperformed controls by 0.32 SD (a 9-point equivalent climb in percentile rankings), but this was driven entirely by senior lawyers, who outperformed controls by 0.45 SD (a 13-point climb). Juniors showed no discernible improvement. Their scores bifurcated instead: more very low scores, fewer mediocre ones, more good ones, and no more excellent ones.

The authors read this as follows. Seniors clearly gained professional judgment from three months of tool use. For juniors the picture is more nuanced: some scores improved, hinting that AI can leapfrog learning, while others spread to the lower rungs. In some scenarios AI helps juniors deliver adequate work, but better machine-assisted performance does not translate into faster acquisition of the expertise that makes senior professionals valuable.

Editing patterns help explain the gap. Across both groups, juniors worked top to bottom and often used up their time copy-editing low-stakes introductory sections before reaching the main claims. They focused on surface-level synonym swapping rather than improving commercial scope. Even when they spotted serious flaws, they often left comments diagnosing the defect instead of making the fix. That "diagnose-without-execute" posture was equally common among unassisted control juniors, so the authors call it a baseline junior deficit that three months of AI-executed rewrites did not remediate. Treated seniors spent longer on redlining than juniors, skipped low-stakes prose, rebuilt claims from scratch, removed language that could inadvertently narrow legal rights or protective scope, and tied many edits to explicit legal doctrines. In follow-up interviews they described AI output as a "logic auditor" rather than a finished product.

The authors conclude that improving performance and building durable judgment are different goals that, for junior professionals in particular, may be in tension. Foundational expertise may be the missing link that lets AI-assisted repetition become seasoned judgment over a career. They state the stakes as ensuring that tools built to produce better work today do not stop junior professionals from becoming the experts needed tomorrow. They also name limits: a limited sample of patent lawyers, only three months of observation, and rapid progress in models and AI familiarity over the past year that might change results if the trial were re-run now. Their clear takeaway is that measuring AI's effect on skill requires tests that separate the effect of using tools from the effect of unassisted performance.

Key facts

  • The trial randomized 133 lawyers at eleven IP law firms; two-thirds of the lawyers at each firm got early access to a then-unreleased Google Labs patent writing assistant, now part of Gemini Notebook.
  • Drafting quality rose with AI access: 0.34 SD at 10 days (a 10-point percentile climb) and 0.38 SD at 90 days (11 percentile points), via fewer poor scores and more good ones, with no change in excellent scores.
  • On the unassisted 90-day redlining task, AI-access lawyers beat controls by 0.32 SD (9 percentile points), but seniors alone drove it at 0.45 SD (a 13-point climb); juniors showed no discernible improvement.
  • Junior redlining scores bifurcated: more very low scores, fewer mediocre ones, more good ones and no more excellent ones.
  • The authors say performance and durable judgment are different goals that may be in tension for juniors, and note a limited sample and a three-month window.

Why it matters

The study targets a live question: does AI help people learn on the job, or does it hollow out the expertise that comes from routine work? Earlier studies point both ways. This one pairs months of real professional use with an unassisted test of judgment and blinded expert grading, which the authors say are rarely found together. Its result is split by seniority. AI raised the quality of what everyone produced, yet only seniors came away with better judgment when the tool was taken away. The authors' warning is that tools built for better work today could keep juniors from becoming the experts needed tomorrow.

Who it affects

Law firms and other expertise-heavy employers that rely on junior staff learning through routine work are the most direct audience. The study covers patent attorneys at eleven IP firms, with juniors and seniors responding differently. Juniors saw faster, better drafts but no average gain in unassisted judgment, and their redlining scores spread toward both ends. Seniors gained judgment, and described the AI as a "logic auditor" that weakened attachment to existing prose and pushed them to explain their structural edits. Researchers studying AI and skills are also addressed: the authors argue that tests must separate the effect of using tools from unassisted performance.

How to use it

This is a study, not a product release. The tool involved was a then-unreleased Google Labs patent writing assistant, now part of Gemini Notebook; the post does not discuss access or pricing. The practical takeaway the authors offer is about measurement: when judging AI's effect on professional skill, test people with the tool removed, as the redlining task did. Their findings on behavior also point to what to watch in juniors, such as spending time on low-stakes sections, swapping synonyms, and diagnosing flaws without fixing them. That last habit appeared equally in unassisted control juniors, so AI-executed rewrites did not fix it.

How solid is it

The design is strong on its own terms: random assignment within each firm, a control group withheld from the tool for the three months, different simulated materials at each checkpoint, and scoring by third-party patent professionals on five dimensions. The paper is published by the National Bureau of Economic Research. The post reports effects in standard deviations and percentile equivalents and gives no p-values, confidence intervals or standard errors. It does not give the number of seniors versus juniors or how "senior" and "junior" are defined, and it does not name the eleven firms. The tool was made by Google and the firms do regular, non-exclusive business with Google; the authors write as Google Research staff.

Risks and caveats

The authors list their own limits. The sample of patent lawyers is limited, a reflection of the high hourly billing rate of top-tier professionals. Three months is a small window compared with the years needed to build durable expertise. Model capability, baseline AI familiarity and AI product penetration in legal work have advanced over the past year, which might change results if the trial were re-run now. The post does not say the findings extend beyond patent law or to other AI tools. It also does not say AI harms junior lawyers' skills: juniors showed no discernible improvement and their scores bifurcated. The explanations about foundational expertise are framed by the authors as possibilities, with "may".

“It weakened attachment to existing prose and forced them to articulate the why and how of structural edits, activating and sharpening their foundational expertise.”

— David Autor and Tanya Rodchenko, Google Research, on how senior lawyers used AI output