Drafting with Local AI: What It Protects, and What It Costs
We had 29 models draft continuation claims on the same ten patent specifications, each scored 0 to 100 for drafting quality by a four-vendor panel. Whether a model's weights are open or closed barely matters. How much memory you can put under the desk decides almost everything.
So open weights cost about 4 points. Memory costs about 23. A 64 GB desktop sits between the two, running Gemma 4 31B at 80. Skip to the full results, or read on for what the confidentiality duties actually require.
Most guidance about AI and client confidentiality was written for lawyers generally. Patent practice has a sharper version of the problem, for reasons that are specific to the subject matter: the thing you would paste into a chat window is often an unpublished disclosure whose value depends on it staying unpublished, and which may be separately regulated for export.
This page covers what the duties actually say, why the patent context differs, and then what we measured when we ran real drafting work through open-weight models instead of closed cloud ones: how the output compares, and how much memory each option needs.
Written for licensed practitioners. Educational, not legal advice: every practitioner makes their own determination about what their duties require in their jurisdiction and on their matters. R&D-grade benchmark results, July 2026. Nothing here assesses patentability, novelty, or non-obviousness.
The Duty Landscape
Three sources bear on sending client technical material to a third-party model.
USPTO Practitioner Guidance
37 CFR 11.106
Practitioners must make reasonable efforts to prevent inadvertent or unauthorized disclosure of information relating to the representation of a client. The Office's AI guidance is explicit that using AI in practice before the USPTO can result in the inadvertent disclosure of highly sensitive technical information to third parties, and that practitioners should be especially vigilant. Two specific hazards it raises: the possibility of material being retained for training, and tools operated outside the United States.
ABA Formal Opinion 512
July 2024
It calls for a fact-specific assessment of the client, the matter, the task, the sensitivity of the information, the safeguards, and the particular tool. Depending on how a system uses or exposes matter information, informed client consent may be required, and generic engagement-letter language may not suffice. Enterprise contractual controls can reduce risk materially, but the label "enterprise" is not itself the analysis.
The Tool's Terms
Read before use, not after
Both sources land in the same place: the obligation is to understand a tool's terms of use, privacy policy, and security posture before client material goes into it. Vendor terms are commonly protective. The provisions that raise risk are broad internal reuse, training on customer inputs, cross-customer retrieval, publication rights, and unrestricted human review.
Why Patent Work Is a Harder Case
Unpublished applications are held in confidence
The USPTO generally must keep them confidential under 35 U.S.C. 122. Before filing, an invention disclosure is protected by the practitioner's confidentiality duties, the client's contractual and trade-secret controls, and other applicable law. The asset being protected is the non-publicness of the disclosure.
Export control reaches technical data, not just shipments
Filing in the United States is treated as including a petition for a foreign filing license under 35 U.S.C. 184. Whether export rules bear on a matter at all is a threshold question that comes first: the subject matter has to be classified as controlled under the EAR, ITAR, DOE Part 810, or a comparable regime. Where it is controlled, releasing it to a foreign person inside the United States counts as an export, so what matters is which people can reach the material, not only where the hardware sits. The model is not itself the foreign person in that analysis. The concern is human access to what you sent.
The material is the crown jewels by construction
A specification is written to enable the invention. It is, deliberately, the most complete technical description of the thing that exists anywhere.
Running inference on the practitioner's own machine removes the routine disclosure of matter information to a model provider. That is a large reduction in risk, and it is worth being precise about what it does not do.
So: no cloud-provider data-processing terms to audit for the inference itself, though model licenses, runtime software, and update channels still warrant review.
The reasonable next question is what that costs in output quality. We measured it.
What We Measured
Ten public parent and continuation patent pairs, specifications from 47 KB to 645 KB. Twenty-nine drafters: closed cloud models, open-weight models too large for any single machine, open-weight models that fit on a server, and open-weight models that fit on a desktop. Every model got the identical specification and prompt. The revise loop and best-practices pass were switched off, so this isolates raw first-pass drafting.
Scoring is the part that usually goes unexamined, so it was checked first. The scoring model is asked to quote verbatim specification passages supporting each claim, and those quotes are then validated in code as actual substrings of the specification. A model that cannot cite its evidence is disqualified from grading, no matter how confident its scores look. That test eliminated every candidate except one, which matters and is discussed below.
Results: Local versus Frontier
Scores are a 0-100 drafting-quality rubric (written-description support, definiteness, differentiation from the parent claims, claim craft). Scoring is done by a four-model, four-vendor panel. An earlier single-scorer pass agreed with the panel on cloud drafts to within about a point, then consistently scored local drafts too high, by up to 17 points. The panel figures below are the ones to rely on.
The colour of each bar is the smallest machine that can hold the model plus a context window big enough for a real specification. That, rather than who owns the weights, is what the results turn on.
Cloud only 512 GB server 64 GB desktop 32 GB and below
Selected rows. Ranks are positions in the full 29-model table below.
View full data table, all 29 models
| Drafter | Runs on | Weights | Specs | Score | Range |
|---|---|---|---|---|---|
| Claude Opus 5 | cloud only | closed | 9/10 | 92 | 88-94 |
| GPT-5.5 | cloud only | closed | 10/10 | 92 | 88-94 |
| Claude Opus 4.7 | cloud only | closed | 10/10 | 90 | 86-92 |
| Kimi K3 | cloud only | open | 10/10 | 89 | 85-93 |
| GLM-5.2 | 512 GB server | open | 10/10 | 88 | 84-94 |
| Nemotron 3 Ultra 550B | 512 GB server | open | 10/10 | 88 | 82-92 |
| GLM-5.1 | 512 GB server | open | 9/10 | 85 | 76-88 |
| MiniMax M2.5 | cloud only | open | 9/10 | 84 | 78-91 |
| Kimi K2.5 | cloud only | open | 10/10 | 81 | 74-86 |
| GLM-5 | 512 GB server | open | 8/10 | 80 | 70-85 |
| DeepSeek V4 | cloud only | open | 10/10 | 80 | 76-87 |
| Gemma 4 31B | 64 GB desktop, 30 GB | open | 9/10 | 80 | 73-84 |
| Qwen3 235B | 192 GB server | open | 10/10 | 77 | 65-84 |
| Mistral Large 2512 | 128 GB server | open | 5/10 | 76 | 52-87 |
| Qwen3-Next 80B | 64 GB desktop | open | 10/10 | 76 | 64-86 |
| Command A | 96 GB server | open | 10/10 | 68 | 52-82 |
| Qwen3-Coder 30B | 64-128 GB | open | 10/10 | 67 | 56-78 |
| Nemotron 3 Super 120B | 96 GB server | open | 10/10 | 66 | 43-84 |
| Gemma 3 27B | 32 GB desktop | open | 8/10 | 66 | 32-82 |
| Gemma 4 26B | 32 GB desktop, 20 GB | open | 10/10 | 65 | 36-81 |
| Gemma 4 12B | 16 GB, 12 GB | open | 9/10 | 63 | 48-78 |
| Nemotron 3 Nano 30B | 32 GB desktop | open | 10/10 | 62 | 48-74 |
| Gemma 3 27B (cloud copy) | cloud | open | 10/10 | 57 | 39-84 |
| GPT-OSS 120B | 96 GB | open | 9/10 | 55 | 37-72 |
| Llama 4 Scout | 96 GB | open | 10/10 | 50 | 22-74 |
| DeepSeek-R1 70B | 128-256 GB | open | 6/10 | 47 | 30-63 |
| Devstral Small 2 24B | 32-256 GB | open | 7/10 | 46 | 18-80 |
| Granite 4.1 8B | 16 GB | open | 5/10 | 40 | 31-54 |
| GLM-4.7 | 512 GB server | open | 1/10 | n/a | not rankable |
Read the Specs column before the score. A model that drafted only 5 of 10 specifications is averaged over the 5 it managed, and the one it failed is usually the hardest. Mistral Large's 76 and Granite's 40 are not comparable to a 10 of 10 result.
Several groups in that table are tied, not ranked. The leading cloud models (92, 92, 90, 89) sit within about a point of each other after accounting for scoring wobble, as do the two best self-hostable models (88, 88) and the cluster around 80. Repeated scoring moves models within those bands around by more than the gaps between them.
The hosting gap is roughly six times the open-versus-closed gap
Best closed weights score 92. The best open weights you can host yourself score 88, a gap of about 4 points. From that self-hosted model down to the best 32 GB desktop model is 23 points. What limits a practitioner is not the licence on the weights. It is how much memory sits under the desk.
Bigger is not reliably better
Gemma 4 31B needs 30 GB and scores 80. Qwen3 235B needs an estimated 132 GB and scores 77. Across most of this range, architecture and training decide the result more than parameter count does. Do not assume a model that needs a bigger machine will draft better on it.
A 30 GB desktop model matches a 1.6-trillion-parameter cloud model
Gemma 4 31B scores 80; DeepSeek V4 scores 80. For this task, parameter count is not the axis that decides the outcome.
Hybrid Mamba architectures buy coverage and memory, not quality
We tested four of them on the theory that near-flat memory growth would extend the workstation ceiling. Nemotron 3 Nano drafted all ten specifications including the 645 KB one that stopped five other models, at 30B parameters with 3B active, and still scored 62, below Gemma 4 26B. Only the 550B variant was competitive, and at that size the result is attributable to scale rather than to the architecture.
Whether a model returns clean structured output matters more than how well it drafts
The single most common reason a model was unusable was not poor claims. It was returning output the pipeline could not parse. Across one model family the discipline improved visibly generation over generation: 1 of 10 specifications completed, then 8, then 9, then 10, with the drafting score climbing alongside it. A model that drafts well but cannot hold a format is not a usable tool.
Read the range column, not just the score
Gemma 4 26B averages 65 but spans 36 to 81 across specifications: on some it drafts acceptably and on others poorly. Gemma 4 31B is both higher and steadier (73 to 84). For work where every specification has to come out usable, consistency matters more than the average.
What the extra points actually buy. Cloud models drafted 15 to 17 claims per specification against 8 to 12 locally, while both produced about the same number of independent claims (around 4). The additional volume is dependent claims: more fallback positions layered under each independent claim. That is a real drafting contribution and it is also the kind of thing a practitioner adds during review.
Claim count is not a quality signal on its own. Llama 4 Scout produced more claims than any other local model and scored second to last. Volume without support is padding, and only the support audit distinguishes them.
What Memory You Need
The constraint is not the model file size. It is the model plus the context window large enough to hold the entire specification, because the whole specification goes into the prompt.
512 GB Server
About 4 points behind the best closed cloud models, without routinely transmitting matter content to a model provider. This is a firm-level purchase rather than a desk-level one, but it is the configuration that makes the confidentiality argument without a meaningful quality concession.
64 GB Machine
About 12 points behind the best cloud models and 8 behind the 512 GB tier. 30 GB will not fit comfortably on a 32 GB machine once the OS and normal work are accounted for. Slower than the 26B by roughly a factor of four.
32 GB Machine
Handles every specification including the largest. The tradeoff is real: 15 points below the 64 GB tier and 23 below the 512 GB tier, with results varying more from specification to specification.
16 GB Machine
Gemma 4 12B fits in 12 GB but its quality is uneven across specifications, and the smaller models cannot hold a real specification at all.
What Local Cannot Currently Do
The pipeline can also score a draft: check every claim against the specification and report which limitations are supported and which are not. That scoring step could not be made to work locally.
The reason is worth stating because it is not obvious. Asked to quote the specification passages supporting each claim, local models produced quotes that were verbatim and real but taken from the claims being graded rather than from the specification. In one run, 81 of 81 quotes came from the wrong document. The claims were being offered as evidence for themselves.
One cloud model
33 of 33 cited quotes found in the specification. The only model that passed.
Best local model
Most quotes were real strings from the wrong document. Scale does not help: 400B+ models scored the same.
Honest Limits on These Numbers
- These are comparative scores, not absolute grades. Every draft for a specification is scored against its peers in a single pass, so a number means something only inside this benchmark. Adding one model to the comparison moved individual scores by as much as 12 points, and repeated runs move a typical model by about 2. Treat small gaps as ties rather than rankings. Scoring is done by a four-vendor panel because an earlier single scorer inflated local drafts by up to 17 points.
- Read coverage before the score. Six models could not draft the whole corpus, and the specification a model fails is usually the hardest one. Mistral Large's 76 is its average over the five it completed, not a result comparable to a model that finished all ten.
- The quantization penalty is unmeasured. Local models run at reduced numerical precision to fit in memory. Our attempt to isolate that cost returned a backwards result and told us nothing, so every local score carries an unknown precision penalty on top of everything else.
- What this did not cover. Raw first-pass drafting only, with the critic and revise loop switched off, so a real workflow needs a second model from a different family as critic and every memory figure here is for the drafting model alone. Each score comes from a single generation. All ten sources are public patents, and memory figures come from one Apple silicon workstation.
The Short Version
If confidentiality is the constraint, self-hosted drafting is a real option rather than a compromise gesture. Open weights cost about 4 points. Memory costs about 23. A firm with one 512 GB server runs a model roughly 4 points behind the best closed cloud models without routinely transmitting matter content to a model provider. A 64 GB desktop runs a model about 12 points behind. A 32 GB machine is about 27 points behind, and varies more from specification to specification.
In a verified local-only configuration, inference happens without routinely transmitting matter content to a model provider: no retention terms, no cross-border transfer, no vendor to audit. "Verified" is doing real work in that sentence, for the reasons in the data-path caveat above.
So the internal argument worth having is about the hardware budget, not about whether open-weight models are good enough. They are. The question is what you can afford to put them on.
If confidentiality is not a constraint on a given matter, because the material is already published, the leading cloud models still draft measurably better and the difference shows up mostly as more dependent claims.
Frequently Asked Questions
Does local-first mean worse quality?
It depends almost entirely on your hardware. A 512 GB server runs an open-weight model that scores 88, about 4 points behind the best closed cloud models. A 64 GB desktop runs one that scores 80, matching a trillion-parameter cloud model. A 32 GB machine drops to 65. The tradeoff is real, quantified, and set by memory rather than by the choice to go local.
Why not just use the cloud model?
For published material, you often should. For unpublished disclosures the question is what the vendor's terms allow. Terms are commonly protective, and where they are, sending a specification to a cloud API is not a public disclosure. Risk rises where terms permit broad internal reuse, training on customer inputs, cross-customer retrieval, publication, or unrestricted human review. Running locally removes the question rather than answering it, which is why it is worth knowing what that costs.
Are open-weight models worse than closed ones?
Barely. Across 29 models the gap from the best closed weights to the best open weights you can host yourself is about 4 points on a 100-point scale, which is close to the scoring noise. The gap from that self-hosted model down to a 32 GB desktop is 23 points. Open versus closed is nearly a non-issue; server versus desktop decides the outcome.
Why does a 30 GB model beat one that needs 132 GB?
Because parameter count is not what decides this task. Gemma 4 31B needs about 30 GB and scores 80; Qwen3 235B needs an estimated 132 GB and scores 77. Architecture and training matter more than size across most of this range. The practical consequence is that you cannot pick a model by how big a machine it demands.
Can a local model also score its own drafts?
Not reliably. Every local model tested cited evidence from the wrong document when asked to verify claim support. The best local judge scored 0.35 on anchor fidelity. Automated scoring currently requires a cloud model, so it cannot be used on privileged material.
What hardware should I buy?
For one practitioner, a machine with 64 GB of unified or GPU memory. On Apple Silicon that is an M-series Mac with 64 GB. The recommended model uses about 30 GB resident, leaving room for the OS and normal work. For a firm willing to make it a shared purchase, a 512 GB server closes most of the remaining gap to the cloud frontier and is the configuration that makes the confidentiality case without a real quality concession.
Do I need to disclose AI use to the USPTO?
There is no general duty to disclose AI involvement. The duty of candor creates a narrow exception where the use is material. The obligation is diligence about where the data goes, not disclosure of the tool itself. Consult your own assessment of the current guidance.
Sources and Related Resources
Duty Discussion Sources
- USPTO: Guidance on AI-Based Tools in Patent Practice (April 2024)
- ABA Formal Opinion 512 (July 2024)
- 37 CFR 11.106: Confidentiality of Information
- 35 U.S.C. 122: Confidential Status of Applications
- 35 U.S.C. 184: Filing of Application in Foreign Country