The Legal AI RFP: Security and Verification Requirements Every Request for Proposal Needs
Most law firms evaluating an AI vendor collect a pile of yes-or-no answers to a security questionnaire, then pick the vendor that answered fastest. A request for proposal built that way rewards confidence, not proof. This guide sets out the security and verification requirements that belong in a legal AI RFP, and gives you a scoring matrix so you compare vendors on what they can demonstrate rather than what they assert.
A legal AI RFP should be a proof instrument, not a trust exercise. The single most useful change you can make to your legal AI RFP security requirements is to add, next to every question, a column that asks how the vendor will prove the answer. Adoption makes this urgent: in the American Bar Association's 2024 survey, 30.2 percent of respondents said their offices were using AI-based tools, up from 11 percent a year earlier, and accuracy was the top concern, named by 74.7 percent of respondents [1]. Firms are buying quickly, and the buying process is where privilege exposure is either caught or waved through.
The reason a generic IT RFP is not enough is that legal AI carries duties an IT questionnaire never contemplates. Under ABA Formal Opinion 512, a lawyer's duties of competence and confidentiality extend to the AI tools the firm deploys, which means the firm, not the vendor, answers to a client or a court if privileged material is mishandled [2]. The RFP is where you convert that duty into contract terms and verifiable evidence before you sign, not after an incident.
This guide is written from the perspective of a verification vendor, not a law firm, and it is informational rather than legal advice. It gives you five things: what a legal AI RFP must cover, the security requirements to demand, the data-handling terms to put in writing, how to turn vendor answers into verifiable proof, and a scoring matrix. What no competing checklist provides is the second column on every row: the proof that backs the claim.
What a legal AI RFP must cover that a generic IT RFP misses
A legal AI RFP must cover four things a generic IT RFP omits: how the vendor isolates one client's data from another's, whether it trains on your inputs, how it handles privileged material specifically, and how the firm can prove all of this to a court or insurer later. These map to the lawyer's confidentiality and supervision duties, which do not transfer to the vendor [2].
A standard IT security questionnaire is built for a generic SaaS purchase: uptime, encryption at rest and in transit, access controls, incident response. Those still matter, but they were never designed for the specific risk that a legal AI tool creates, which is that privileged client material passes through a third party's model and infrastructure.
The legal-specific layer asks different questions. Does the vendor isolate each client's or each matter's data, or does everything sit in one shared context where a prompt could surface another firm's material? Does the vendor use your inputs to train or improve its models? Can the firm produce a record, months later, showing what the tool did with a given document? An IT RFP rarely asks any of these, because they only matter when the data is privileged.
The organizing idea is that the duty does not move. Opinion 512 frames competence and confidentiality as the lawyer's obligations, not the vendor's, so the RFP has to extract terms and evidence strong enough that the firm can meet its own duty. Everything below is built to do that.
The security requirements to demand, and the proof that backs each
Demand five categories with proof attached: current independent audit (SOC 2 Type II or ISO/IEC 27001), an AI-specific management standard (ISO/IEC 42001), encryption and access controls, a tested incident-response process, and named sub-processors. For each, require the artifact that proves it, not a checkbox: the report, the certificate, the pen-test summary, the sub-processor list [3][4].
The requirements below are the core of the RFP. The value is in the second column. A vendor that answers yes to every question and cannot produce a single artifact has told you something important.
SOC 2 Type II is the baseline because it tests whether controls actually operated over a period, not just whether they exist on paper. ISO/IEC 42001, published in December 2023 as the first international standard for an AI management system, is the AI-specific counterpart, and a vendor pursuing it signals that AI governance is a managed process rather than an afterthought [3]. Require the actual report or certificate and read the scope, because a SOC 2 that excludes the product you are buying proves nothing about it.
| Requirement | What to require | How to verify (the proof) |
|---|---|---|
| Independent audit | Current SOC 2 Type II, ISO/IEC 27001, or both | The report itself; check the audit period and that scope covers the product you are buying [4] |
| AI management | ISO/IEC 42001 certification or a documented AI management program | Certificate or the program documentation and its owner [3] |
| Data isolation | Named isolation model (single-tenant, dedicated, or logical) for your data | Architecture description; ask which model and how it is enforced |
| Encryption and access | Encryption in transit and at rest; least-privilege access; SSO/MFA | Config summary; key-management description; access-review cadence |
| Incident response | Documented, tested IR plan with breach-notification timelines | The plan plus evidence of the last tabletop or test |
| Sub-processors | Full sub-processor list and data-residency map | The list and the countries your data touches |
The data handling and training terms to put in writing
Put four data terms in the contract, not the sales deck: a no-training commitment on your inputs and outputs, a data-processing agreement, a defined retention and deletion schedule, and named sub-processors with data residency. A verbal or marketing assurance is not enforceable; a contractual term is. Where matters must be provably walled off, require single-tenant or dedicated isolation rather than logical separation.
The most consequential term is training. A vendor that says it does not train on your data is making a claim that is only worth what the contract says. Require the commitment in the master agreement or data-processing agreement, covering both inputs and outputs, and treat any refusal to commit in writing as a material gap.
A data-processing agreement should define retention, deletion on request and at termination, and the vendor's obligations if it receives a subpoena for your data. Sub-processors and residency belong here too: you cannot assess privilege exposure if you do not know which downstream services touch the data or which countries it sits in.
This is where the isolation trade-off gets decided. Logical, multi-tenant isolation is cheaper and adequate for lower-sensitivity work. For matters that must be provably separated, single-tenant or dedicated infrastructure is the stronger posture, and the RFP should ask the vendor which model applies and let you require the stronger one for sensitive matters. The companion guide on reading a SOC 2 report covers how to check that these controls are real rather than asserted.
Turning vendor answers into verifiable proof
Do not accept a vendor's security answers; verify them. For each claim, ask for the artifact that proves it and confirm the artifact actually covers your use. A "we do not train on your data" claim becomes verifiable only when it is a contract term plus a technical control you can point to. This is the gap generic checklists leave open, and it is where most RFPs fail.
Verification is the difference between an RFP that protects the firm and one that documents a decision you cannot defend. A claim that a tool is accurate is not the same as evidence: independent testing by Stanford researchers found leading legal AI research tools returned inaccurate answers on a meaningful share of queries, with reported figures around 17 to 33 percent in 2024, which is exactly why "it is accurate" needs backing [5].
Build the verification into the RFP by requiring, for each security answer, the artifact that proves it and confirmation that the artifact covers the specific product and time period you are buying. A SOC 2 from two years ago, or one scoped to a different service, is not proof. The no-training claim in particular is one firms accept on faith and should treat as a contract-plus-control question.
The direction this points is attestation: a verifiable record that a control operated, rather than a promise that it exists. RankShield Legal is building an AI-tool attestation gateway toward exactly this, so that each AI-assisted action carries a signed record a firm can show a court or insurer. That capability is on the roadmap and under active design, not a shipped feature you can turn on today; the point for your RFP is to ask vendors how they would prove their claims, because the market is moving from assurance to evidence.
How to score competing legal AI proposals
Score each vendor on proof, not promises. Weight the categories, score each on a 0 to 2 scale where 0 is an unbacked claim, 1 is partial evidence, and 2 is a current artifact that covers your use, then set hard gates that disqualify regardless of total. A high score with no SOC 2 and no written no-training term is a fail, not a close call.
A scoring matrix keeps the decision defensible and comparable. Weight the categories by your firm's risk, score each vendor on the evidence they produced, and record the artifact you relied on so the file shows why you chose as you did.
Set decision triggers that override the total. If a vendor cannot produce a current SOC 2 Type II report covering the service you are buying, do not shortlist it. If the contract will not commit in writing to not training on your inputs and outputs, score that category zero and treat it as a gate, not a deduction. If the vendor cannot name its sub-processors and the countries your data touches, you cannot assess privilege exposure, so require the list before scoring rather than after.
The security questionnaire that anchors the RFP gives you the underlying questions; the matrix below converts them into a comparison a committee can sign.
One scoring discipline matters more than the weights you choose: record the artifact, not the impression. A matrix that says "Vendor A: 2" six months from now tells a reviewer nothing. A matrix that says "Vendor A: 2, SOC 2 Type II dated March 2026, scope covers the hosted drafting service" reconstructs the decision. If the firm is later asked why it selected a vendor whose control failed, the second version is a defensible record and the first is a number.
| Category | Suggested weight | Score 0 (assert) | Score 2 (prove) |
|---|---|---|---|
| Independent audit | High | Says "yes," no report | Current in-scope SOC 2 Type II or ISO 27001 [4] |
| No-training term | High | Marketing claim only | Written term covering inputs and outputs |
| Data isolation | High | "Isolated," undefined | Named model, enforcement described |
| Sub-processors and residency | Medium | Not disclosed | Full list with countries |
| Incident response | Medium | Plan on paper only | Plan plus evidence of a recent test |
Sequence the RFP so evidence arrives before the decision
Most RFPs fail on timing rather than on content. Artifacts like a current SOC 2 Type II, a named sub-processor list, and a redlined no-training term take weeks to produce and often require an NDA first. A firm that requests them after shortlisting has already made its decision and is collecting paperwork to justify it.
The practical failure looks like this. The committee runs a demo round, forms a preference, shortlists two vendors, and then asks for the audit reports. The reports arrive late, one is scoped to a different service, and by then the preference is established and the deadline is close. The gap gets waived.
Ordering the process differently costs nothing and changes the outcome. Send the artifact requests with the initial RFP, not after the shortlist, and make the NDA part of the opening package so the vendor can share a report under it without a second negotiation. Set a date by which artifacts must be in hand, and treat a vendor that misses it as having answered the question.
Give legal review of the contract terms its own track, running in parallel rather than after selection. The no-training commitment, the data-processing agreement, retention and deletion, and the sub-processor list are contract questions, not procurement questions, and discovering that a vendor will not commit in writing is far cheaper before you have chosen them than after.
Run the demo last rather than first. A demo is the most persuasive and least evidentiary part of an evaluation, and seeing it before the artifacts arrive is what produces a preference the evidence then has to argue against. Evidence first, demonstration second, is the same sequencing discipline the rest of this guide applies to individual claims.
An RFP is not a one-time event
Every artifact in the evaluation has an expiry. A SOC 2 Type II covers a defined audit window, sub-processor lists change without notice unless the contract requires it, and a product that was single-tenant at signing can be re-architected. A vendor that passed in 2026 is not a vendor that passes in 2028, and nothing in the file will tell you unless someone re-checks.
The evaluation produces a snapshot, and firms tend to treat it as a standing conclusion. The artifacts themselves say otherwise: an audit report covers a stated period and asserts nothing about the period after it.
Three things move between renewals. Audit reports lapse, so a report that was current at selection may cover a window that closed long ago. Sub-processors change, which matters because a new downstream service can alter where data sits and which jurisdictions can reach it. And products get re-architected, so an isolation model described at signing is a claim about the past unless the contract obliges the vendor to notify you of material changes.
The lightweight fix is a contractual notification duty plus an annual re-check. Require the vendor to notify you of new sub-processors and of material changes to data handling or isolation architecture. Then re-request the current audit report on a schedule, and score it the same way you scored the original: is it current, and does its scope cover the service you are actually using.
The same logic applies to the accuracy question. Independent testing found leading legal AI research tools returning inaccurate answers on a meaningful share of queries, with reported figures around 17 to 33 percent in 2024 [5]. Model behaviour changes with every version, so a performance claim validated at purchase is not a performance claim validated today. That is the argument for buying verification you can re-run rather than assurance you accepted once.
Test yourself on evaluating a legal AI vendor
Five questions on separating what a vendor can prove from what it asserts.
-
1A vendor says it does not train on your data. What makes that verifiable?
Answer: A written contract term covering inputs and outputs, plus a technical control
The claim is worth what the contract says. Require the commitment in the master agreement or data-processing agreement, covering both inputs and outputs, and treat a refusal to commit in writing as a material gap rather than a negotiating position.
-
2A vendor produces a SOC 2 Type II report. What still has to be checked?
Answer: That it is current and its scope covers the specific service you are buying
A report from two years ago, or one scoped to a different service, is not proof. An audit report covers a defined window and asserts nothing about the period after it, which is also why re-checking on a schedule matters.
-
3When should artifact requests go out?
Answer: With the initial RFP, before any preference forms
Artifacts take weeks and often need an NDA first. Requesting them after shortlisting means the decision is already made and the paperwork is justifying it. Evidence first, demonstration second.
-
4Why record the artifact alongside the score?
Answer: So the decision can be reconstructed later if a control fails
"Vendor A: 2" tells a later reviewer nothing. "Vendor A: 2, SOC 2 Type II dated March 2026, scope covers the hosted drafting service" is a defensible record of why the firm chose as it did.
-
5Why does a passing evaluation expire?
Answer: Audits lapse, sub-processors change, and products get re-architected
The evaluation is a snapshot. A contractual duty to notify you of new sub-processors and material changes to data handling or isolation, plus an annual artifact re-check, is what keeps the conclusion true.
Honest self-check. There is no sign-up, and nothing is stored.
Straight answers to the common questions
The questions readers ask about this topic, answered directly. No forms, no sales pitch.
Pick a question on the left, or search above. You will get the direct answer, the way an answer engine would give it.
References
- American Bar Association. 2024 Artificial Intelligence TechReport (Legal Technology Survey Report). 2025. https://www.americanbar.org/groups/law_practice/resources/tech-report/2024/2024-artificial-intelligence-techreport/
- American Bar Association. Formal Opinion 512: Generative Artificial Intelligence Tools. July 2024. https://www.americanbar.org/news/abanews/aba-news-archives/2024/07/aba-issues-first-ethics-guidance-ai-tools/
- International Organization for Standardization. ISO/IEC 42001:2023 Artificial Intelligence Management System. December 2023. https://www.iso.org/standard/81230.html
- AICPA. SOC 2 Trust Services Criteria. 2022. https://www.aicpa-cima.com/topic/audit-assurance/audit-and-assurance-greater-than-soc-2
- Stanford HAI. AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries. 2024. https://hai.stanford.edu/news/ai-trial-legal-models-hallucinate-1-out-6-or-more-benchmarking-queries
Check a citation against live case-law
Paste a citation from an AI-drafted brief and see whether the case actually exists, resolved against live case-law. Free, no sign-up. Then request early access to certify a full filing.
Try the citation checker