How to evaluate legal AI without grading it on features.
A feature checklist tells you what a tool does on a good day. This framework tells you what it costs you on a bad one.
6 axes, scored 0 to 3, with the specific test to run and the red flag that should stop the deal. Vendor-neutral. No gate, no email wall.
Features are commoditizing to zero.
Three things happened in 2026 that broke the feature checklist as a buying tool.
Records retrieval has been quoted under $0.40 per page. A feature that costs cents to deliver stops being a reason to pick a vendor.
The compute floor under most legal AI features fell hard in 2026. What one vendor ships this quarter, 3 more ship next quarter.
The companies supplying the models are building legal products directly. Feature parity arrives from above, on their schedule, not yours.
What survives commoditization is the part you carry: the malpractice exposure, the bar complaint, the discovery deadline you missed because a tool summarized around a gap. Score that.
6 axes. Score each 0 to 3.
Most eval guides test whether a tool invents facts. In legal work the expensive failure is the clause or precedent it fails to mention, because a missing thing trips no alarm.
Hand it a document with a known material clause removed. Does it flag the absence, or answer confidently around the hole?
The vendor publishes accuracy numbers and never publishes recall.
The practitioner version of this question is blunt: can you reproduce, under oath, the output this tool gave you 6 months ago?
Ask which model version served a given output, whether that version is pinned to your account, and whether you get notice before it changes.
"We're always on the latest model." That's a non-answer, and it's a hole in your malpractice defense.
Local and private execution is table stakes now. Firms building their own stacks in 2026 are blocked on privilege questions rather than on model quality.
Ask where the document sits at rest, in transit, and in the training pipeline. Get the subprocessor list. Get the retention window in writing.
Training on your data by default, with an opt-out buried in a settings pane.
You're squeezed from both sides. There's pressure to adopt AI and pressure not to file unreviewed output or leak client data. A tool either helps you supervise or it quietly makes supervision harder.
Does the tool make human review structurally easy through citations to source, diffs, and surfaced confidence? Or does it emit a finished-looking artifact that invites a rubber stamp?
Polished output with no provenance trail. Polish is a supervision hazard.
Rank every tool by whether a 4-attorney firm can run it without hiring an engineer. The gap for smaller firms is affordability and operability together, and operability is the half vendors ignore.
Ask who owns it when it breaks on a Friday, and walk the actual setup path end to end.
Enterprise pricing paired with enterprise assumptions about your IT staff.
The tool layer is commoditizing, so you will switch. Price the switch now, while you're still the one holding the signature.
Ask to export your matters, prompts, and history, then open the export and check what survived.
No export, or an export that loses the structure of the work product.
What each number means.
The vendor cannot answer the question, or answers it with marketing copy.
They claim it verbally. Nothing in the contract, docs, or product confirms it.
It's written down in docs or the agreement, and you could hold them to it.
You ran the test yourself on your own matter and watched it hold.
A 0 on any of axes 1 through 4 disqualifies the tool regardless of its total. Those 4 are liability axes, and a high total elsewhere doesn't buy back a hole in your defense.
Axes 5 and 6 are cost axes. A low score there opens a budget conversation rather than stopping the deal.
Common questions.
Features are converging fast. Records retrieval has been quoted under $0.40 per page, agent compute costs fell roughly 39x, and foundation model vendors are shipping legal products directly. The axes that predict whether a tool survives contact with a real matter are liability surface and operability, so those are what this framework scores.
Score each axis 0 to 3. Any 0 on axes 1 through 4 disqualifies the tool regardless of its total, because those 4 are liability axes rather than preference axes. Axes 5 and 6 are cost axes, so a low score there is a budget conversation instead of a stop.
Invented facts get caught because they look wrong. The missing clause, the precedent that never surfaced, the exhibit that was never requested: those read as clean output. That's why omission behavior is axis 1 and why you should test it with a document you've deliberately damaged.
Ask which model version served a specific output, whether that version is pinned to your account, and what notice you get before it changes. If the answer is that they're always on the latest model, you cannot reproduce a 6-month-old output, which matters the day someone asks you to explain how a filing was produced.
No. This page teaches the scoring method and stays vendor-neutral on purpose. Our tool directory is where individual tools get listed and compared.
That's a staffing constraint rather than a software one. A tool that scores 3 on supervision fit still assumes someone at the firm has the hours to do the review. If the review capacity isn't there, the tool transfers risk to you rather than reducing it.
Run the framework on real tools.
Our directory lists legal AI tools with vendor links, so you can score candidates against these 6 axes yourself.
If axis 4 is where your firm keeps landing, the constraint is review hours rather than software. We run dedicated paralegal pods for PI firms, and the team comes with the tools.