Skip to content

Drafting with Local AI: What It Protects, and What It Costs

What we found

We had 29 models draft continuation claims on the same ten patent specifications, each scored 0 to 100 for drafting quality by a four-vendor panel. Whether a model's weights are open or closed barely matters. How much memory you can put under the desk decides almost everything.

92 Claude Opus 5 Best closed-weight cloud model. The ceiling for this task.
88 GLM-5.2 Best open-weight model you can host yourself, on a 512 GB server. Four points off the ceiling.
65 Gemma 4 26B Best model that fits a 32 GB laptop. Twenty-three points below the server.

So open weights cost about 4 points. Memory costs about 23. A 64 GB desktop sits between the two, running Gemma 4 31B at 80. Skip to the full results, or read on for what the confidentiality duties actually require.

Most guidance about AI and client confidentiality was written for lawyers generally. Patent practice has a sharper version of the problem, for reasons that are specific to the subject matter: the thing you would paste into a chat window is often an unpublished disclosure whose value depends on it staying unpublished, and which may be separately regulated for export.

This page covers what the duties actually say, why the patent context differs, and then what we measured when we ran real drafting work through open-weight models instead of closed cloud ones: how the output compares, and how much memory each option needs.

Written for licensed practitioners. Educational, not legal advice: every practitioner makes their own determination about what their duties require in their jurisdiction and on their matters. R&D-grade benchmark results, July 2026. Nothing here assesses patentability, novelty, or non-obviousness.

The Duty Landscape

Three sources bear on sending client technical material to a third-party model.

USPTO Practitioner Guidance

37 CFR 11.106

Practitioners must make reasonable efforts to prevent inadvertent or unauthorized disclosure of information relating to the representation of a client. The Office's AI guidance is explicit that using AI in practice before the USPTO can result in the inadvertent disclosure of highly sensitive technical information to third parties, and that practitioners should be especially vigilant. Two specific hazards it raises: the possibility of material being retained for training, and tools operated outside the United States.

ABA Formal Opinion 512

July 2024

It calls for a fact-specific assessment of the client, the matter, the task, the sensitivity of the information, the safeguards, and the particular tool. Depending on how a system uses or exposes matter information, informed client consent may be required, and generic engagement-letter language may not suffice. Enterprise contractual controls can reduce risk materially, but the label "enterprise" is not itself the analysis.

The Tool's Terms

Read before use, not after

Both sources land in the same place: the obligation is to understand a tool's terms of use, privacy policy, and security posture before client material goes into it. Vendor terms are commonly protective. The provisions that raise risk are broad internal reuse, training on customer inputs, cross-customer retrieval, publication rights, and unrestricted human review.

None of this prohibits AI use. There is also no general duty to tell the USPTO that an AI tool helped draft a filing, though the duty of candor creates a narrow exception where the use is material. The obligation is diligence about where the data goes, not abstinence.

Why Patent Work Is a Harder Case

Unpublished applications are held in confidence

The USPTO generally must keep them confidential under 35 U.S.C. 122. Before filing, an invention disclosure is protected by the practitioner's confidentiality duties, the client's contractual and trade-secret controls, and other applicable law. The asset being protected is the non-publicness of the disclosure.

Export control reaches technical data, not just shipments

Filing in the United States is treated as including a petition for a foreign filing license under 35 U.S.C. 184. Whether export rules bear on a matter at all is a threshold question that comes first: the subject matter has to be classified as controlled under the EAR, ITAR, DOE Part 810, or a comparable regime. Where it is controlled, releasing it to a foreign person inside the United States counts as an export, so what matters is which people can reach the material, not only where the hardware sits. The model is not itself the foreign person in that analysis. The concern is human access to what you sent.

The material is the crown jewels by construction

A specification is written to enable the invention. It is, deliberately, the most complete technical description of the thing that exists anywhere.

Running inference on the practitioner's own machine removes the routine disclosure of matter information to a model provider. That is a large reduction in risk, and it is worth being precise about what it does not do.

"Local" has to describe the whole data path, not just where the model weights sit. A correctly configured local-only setup still leaves open: telemetry and update checks, crash reports, cloud backups, temporary files and swap, prompt and trace logs, who inside the firm can read those logs, and any automatic fallback to a remote model.

So: no cloud-provider data-processing terms to audit for the inference itself, though model licenses, runtime software, and update channels still warrant review.

The reasonable next question is what that costs in output quality. We measured it.

What We Measured

Ten public parent and continuation patent pairs, specifications from 47 KB to 645 KB. Twenty-nine drafters: closed cloud models, open-weight models too large for any single machine, open-weight models that fit on a server, and open-weight models that fit on a desktop. Every model got the identical specification and prompt. The revise loop and best-practices pass were switched off, so this isolates raw first-pass drafting.

Scoring is the part that usually goes unexamined, so it was checked first. The scoring model is asked to quote verbatim specification passages supporting each claim, and those quotes are then validated in code as actual substrings of the specification. A model that cannot cite its evidence is disqualified from grading, no matter how confident its scores look. That test eliminated every candidate except one, which matters and is discussed below.

Results: Local versus Frontier

Scores are a 0-100 drafting-quality rubric (written-description support, definiteness, differentiation from the parent claims, claim craft). Scoring is done by a four-model, four-vendor panel. An earlier single-scorer pass agreed with the panel on cloud drafts to within about a point, then consistently scored local drafts too high, by up to 17 points. The panel figures below are the ones to rely on.

The colour of each bar is the smallest machine that can hold the model plus a context window big enough for a real specification. That, rather than who owns the weights, is what the results turn on.

1
Claude Opus 5 92
Cloud only · closed weights · Range 88-94
2
GPT-5.5 92
Cloud only · closed weights · Range 88-94
4
Kimi K3 89
Cloud only · open weights · 2.8T, too large for any single machine · Range 85-93
5
GLM-5.2 88
512 GB server · open weights · Range 84-94 · Best model you can host yourself
6
Nemotron 3 Ultra 550B 88
512 GB server · open weights · Range 82-92 · Tied with GLM-5.2
7
GLM-5.1 85
512 GB server · open weights · Range 76-88
8
MiniMax M2.5 84
Cloud only · open weights · Range 78-91
9
Kimi K2.5 81
Cloud only · open weights · Range 74-86
10
GLM-5 80
512 GB server · open weights · 8 of 10 specifications · Range 70-85
11
DeepSeek V4 80
Cloud only · open weights · 1.6T total / 49B active · Range 76-87
12
Gemma 4 31B 80
64 GB desktop, 30 GB resident · open weights · Range 73-84 · Best model that fits on a workstation
13
Qwen3 235B 77
192 GB server · open weights · Range 65-84 · Scores below a model a fifth its size
15
Qwen3-Next 80B 76
64 GB desktop · open weights · 80B/3B active · Range 64-86
20
Gemma 4 26B 65
32 GB desktop, 20 GB resident · open weights · Range 36-81 · Best option at 32 GB
21
Gemma 4 12B 63
16 GB, 12 GB resident · open weights · Range 48-78

Cloud only   512 GB server   64 GB desktop   32 GB and below

Selected rows. Ranks are positions in the full 29-model table below.

View full data table, all 29 models
Drafter Runs on Weights Specs Score Range
Claude Opus 5 cloud only closed 9/10 92 88-94
GPT-5.5 cloud only closed 10/10 92 88-94
Claude Opus 4.7 cloud only closed 10/10 90 86-92
Kimi K3 cloud only open 10/10 89 85-93
GLM-5.2 512 GB server open 10/10 88 84-94
Nemotron 3 Ultra 550B 512 GB server open 10/10 88 82-92
GLM-5.1 512 GB server open 9/10 85 76-88
MiniMax M2.5 cloud only open 9/10 84 78-91
Kimi K2.5 cloud only open 10/10 81 74-86
GLM-5 512 GB server open 8/10 80 70-85
DeepSeek V4 cloud only open 10/10 80 76-87
Gemma 4 31B 64 GB desktop, 30 GB open 9/10 80 73-84
Qwen3 235B 192 GB server open 10/10 77 65-84
Mistral Large 2512 128 GB server open 5/10 76 52-87
Qwen3-Next 80B 64 GB desktop open 10/10 76 64-86
Command A 96 GB server open 10/10 68 52-82
Qwen3-Coder 30B 64-128 GB open 10/10 67 56-78
Nemotron 3 Super 120B 96 GB server open 10/10 66 43-84
Gemma 3 27B 32 GB desktop open 8/10 66 32-82
Gemma 4 26B 32 GB desktop, 20 GB open 10/10 65 36-81
Gemma 4 12B 16 GB, 12 GB open 9/10 63 48-78
Nemotron 3 Nano 30B 32 GB desktop open 10/10 62 48-74
Gemma 3 27B (cloud copy) cloud open 10/10 57 39-84
GPT-OSS 120B 96 GB open 9/10 55 37-72
Llama 4 Scout 96 GB open 10/10 50 22-74
DeepSeek-R1 70B 128-256 GB open 6/10 47 30-63
Devstral Small 2 24B 32-256 GB open 7/10 46 18-80
Granite 4.1 8B 16 GB open 5/10 40 31-54
GLM-4.7 512 GB server open 1/10 n/a not rankable

Read the Specs column before the score. A model that drafted only 5 of 10 specifications is averaged over the 5 it managed, and the one it failed is usually the hardest. Mistral Large's 76 and Granite's 40 are not comparable to a 10 of 10 result.

Several groups in that table are tied, not ranked. The leading cloud models (92, 92, 90, 89) sit within about a point of each other after accounting for scoring wobble, as do the two best self-hostable models (88, 88) and the cluster around 80. Repeated scoring moves models within those bands around by more than the gaps between them.

The hosting gap is roughly six times the open-versus-closed gap

Best closed weights score 92. The best open weights you can host yourself score 88, a gap of about 4 points. From that self-hosted model down to the best 32 GB desktop model is 23 points. What limits a practitioner is not the licence on the weights. It is how much memory sits under the desk.

Bigger is not reliably better

Gemma 4 31B needs 30 GB and scores 80. Qwen3 235B needs an estimated 132 GB and scores 77. Across most of this range, architecture and training decide the result more than parameter count does. Do not assume a model that needs a bigger machine will draft better on it.

A 30 GB desktop model matches a 1.6-trillion-parameter cloud model

Gemma 4 31B scores 80; DeepSeek V4 scores 80. For this task, parameter count is not the axis that decides the outcome.

Hybrid Mamba architectures buy coverage and memory, not quality

We tested four of them on the theory that near-flat memory growth would extend the workstation ceiling. Nemotron 3 Nano drafted all ten specifications including the 645 KB one that stopped five other models, at 30B parameters with 3B active, and still scored 62, below Gemma 4 26B. Only the 550B variant was competitive, and at that size the result is attributable to scale rather than to the architecture.

Whether a model returns clean structured output matters more than how well it drafts

The single most common reason a model was unusable was not poor claims. It was returning output the pipeline could not parse. Across one model family the discipline improved visibly generation over generation: 1 of 10 specifications completed, then 8, then 9, then 10, with the drafting score climbing alongside it. A model that drafts well but cannot hold a format is not a usable tool.

Read the range column, not just the score

Gemma 4 26B averages 65 but spans 36 to 81 across specifications: on some it drafts acceptably and on others poorly. Gemma 4 31B is both higher and steadier (73 to 84). For work where every specification has to come out usable, consistency matters more than the average.

What the extra points actually buy. Cloud models drafted 15 to 17 claims per specification against 8 to 12 locally, while both produced about the same number of independent claims (around 4). The additional volume is dependent claims: more fallback positions layered under each independent claim. That is a real drafting contribution and it is also the kind of thing a practitioner adds during review.

Claim count is not a quality signal on its own. Llama 4 Scout produced more claims than any other local model and scored second to last. Volume without support is padding, and only the support audit distinguishes them.

What Memory You Need

The constraint is not the model file size. It is the model plus the context window large enough to hold the entire specification, because the whole specification goes into the prompt.

Best self-hosted

512 GB Server

GLM-5.2 or Nemotron 3 Ultra 550B
~309 GB In use
88 Score
645 KB Max spec

About 4 points behind the best closed cloud models, without routinely transmitting matter content to a model provider. This is a firm-level purchase rather than a desk-level one, but it is the configuration that makes the confidentiality argument without a meaningful quality concession.

32 GB Machine

Gemma 4 26B
20 GB In use
65 Score
645 KB Max spec

Handles every specification including the largest. The tradeoff is real: 15 points below the 64 GB tier and 23 below the 512 GB tier, with results varying more from specification to specification.

16 GB Machine

Not recommended
12 GB In use
63 Score
9/10 Coverage

Gemma 4 12B fits in 12 GB but its quality is uneven across specifications, and the smaller models cannot hold a real specification at all.

The whole spread from a 512 GB server to a 32 GB laptop is 23 points. The spread from closed weights to open weights is 4. If confidentiality is driving the decision, the question worth arguing about internally is the hardware budget, not whether open-weight models are good enough.
Some architectures make long context nearly free. Both Gemma 4 variants hold a 645 KB specification in essentially the same memory as a 47 KB one (26B: 18 to 20 GB; 31B: 26 to 30 GB). Qwen3-Coder 30B climbs from 21 GB to 122 GB across the same range. This is a property of the attention architecture, not of model size, and it is the single most useful thing we learned about matching models to machines.
Some models cannot do this work at any memory size. Small models with context windows under about 45,000 tokens cannot hold a full specification, so no amount of RAM helps. They draft convincingly on short test documents and fail on real ones.

What Local Cannot Currently Do

The pipeline can also score a draft: check every claim against the specification and report which limitations are supported and which are not. That scoring step could not be made to work locally.

The reason is worth stating because it is not obvious. Asked to quote the specification passages supporting each claim, local models produced quotes that were verbatim and real but taken from the claims being graded rather than from the specification. In one run, 81 of 81 quotes came from the wrong document. The claims were being offered as evidence for themselves.

1.00

One cloud model

33 of 33 cited quotes found in the specification. The only model that passed.

0.35

Best local model

Most quotes were real strings from the wrong document. Scale does not help: 400B+ models scored the same.

Drafting can be private, and the automated support audit currently cannot. The audit is optional and the drafting path never invokes it, so this constrains a checking feature rather than the core work. For an unpublished disclosure, the honest options today are to draft locally and review the support yourself, or to use the automated audit only on material that is already public.

Honest Limits on These Numbers

  • These are comparative scores, not absolute grades. Every draft for a specification is scored against its peers in a single pass, so a number means something only inside this benchmark. Adding one model to the comparison moved individual scores by as much as 12 points, and repeated runs move a typical model by about 2. Treat small gaps as ties rather than rankings. Scoring is done by a four-vendor panel because an earlier single scorer inflated local drafts by up to 17 points.
  • Read coverage before the score. Six models could not draft the whole corpus, and the specification a model fails is usually the hardest one. Mistral Large's 76 is its average over the five it completed, not a result comparable to a model that finished all ten.
  • The quantization penalty is unmeasured. Local models run at reduced numerical precision to fit in memory. Our attempt to isolate that cost returned a backwards result and told us nothing, so every local score carries an unknown precision penalty on top of everything else.
  • What this did not cover. Raw first-pass drafting only, with the critic and revise loop switched off, so a real workflow needs a second model from a different family as critic and every memory figure here is for the drafting model alone. Each score comes from a single generation. All ten sources are public patents, and memory figures come from one Apple silicon workstation.

The Short Version

If confidentiality is the constraint, self-hosted drafting is a real option rather than a compromise gesture. Open weights cost about 4 points. Memory costs about 23. A firm with one 512 GB server runs a model roughly 4 points behind the best closed cloud models without routinely transmitting matter content to a model provider. A 64 GB desktop runs a model about 12 points behind. A 32 GB machine is about 27 points behind, and varies more from specification to specification.

In a verified local-only configuration, inference happens without routinely transmitting matter content to a model provider: no retention terms, no cross-border transfer, no vendor to audit. "Verified" is doing real work in that sentence, for the reasons in the data-path caveat above.

So the internal argument worth having is about the hardware budget, not about whether open-weight models are good enough. They are. The question is what you can afford to put them on.

If confidentiality is not a constraint on a given matter, because the material is already published, the leading cloud models still draft measurably better and the difference shows up mostly as more dependent claims.

The decision is per matter, not once for the firm. The same practitioner may reasonably send a published patent to a cloud model in the morning and keep an unfiled disclosure on the local machine in the afternoon. Publication status of the attached document is not the whole test: a prompt built around a public patent can still carry claim strategy, planned amendments, invalidity theories, or the attorney's own impressions. The question is what the entire input and purpose reveal, not whether the source document is public.

Frequently Asked Questions

Does local-first mean worse quality?

It depends almost entirely on your hardware. A 512 GB server runs an open-weight model that scores 88, about 4 points behind the best closed cloud models. A 64 GB desktop runs one that scores 80, matching a trillion-parameter cloud model. A 32 GB machine drops to 65. The tradeoff is real, quantified, and set by memory rather than by the choice to go local.

Why not just use the cloud model?

For published material, you often should. For unpublished disclosures the question is what the vendor's terms allow. Terms are commonly protective, and where they are, sending a specification to a cloud API is not a public disclosure. Risk rises where terms permit broad internal reuse, training on customer inputs, cross-customer retrieval, publication, or unrestricted human review. Running locally removes the question rather than answering it, which is why it is worth knowing what that costs.

Are open-weight models worse than closed ones?

Barely. Across 29 models the gap from the best closed weights to the best open weights you can host yourself is about 4 points on a 100-point scale, which is close to the scoring noise. The gap from that self-hosted model down to a 32 GB desktop is 23 points. Open versus closed is nearly a non-issue; server versus desktop decides the outcome.

Why does a 30 GB model beat one that needs 132 GB?

Because parameter count is not what decides this task. Gemma 4 31B needs about 30 GB and scores 80; Qwen3 235B needs an estimated 132 GB and scores 77. Architecture and training matter more than size across most of this range. The practical consequence is that you cannot pick a model by how big a machine it demands.

Can a local model also score its own drafts?

Not reliably. Every local model tested cited evidence from the wrong document when asked to verify claim support. The best local judge scored 0.35 on anchor fidelity. Automated scoring currently requires a cloud model, so it cannot be used on privileged material.

What hardware should I buy?

For one practitioner, a machine with 64 GB of unified or GPU memory. On Apple Silicon that is an M-series Mac with 64 GB. The recommended model uses about 30 GB resident, leaving room for the OS and normal work. For a firm willing to make it a shared purchase, a 512 GB server closes most of the remaining gap to the cloud frontier and is the configuration that makes the confidentiality case without a real quality concession.

Do I need to disclose AI use to the USPTO?

There is no general duty to disclose AI involvement. The duty of candor creates a narrow exception where the use is material. The obligation is diligence about where the data goes, not disclosure of the tool itself. Consult your own assessment of the current guidance.