Buyer framework

How to evaluate legal AI without grading it on features.

A feature checklist tells you what a tool does on a good day. This framework tells you what it costs you on a bad one.

6 axes, scored 0 to 3, with the specific test to run and the red flag that should stop the deal. Vendor-neutral. No gate, no email wall.

Built for PI firms of 1 to 10 attorneys
Scores liability surface and operability
Every axis has a test you can run
No vendors named or ranked
The diagnosis

Features are commoditizing to zero.

Three things happened in 2026 that broke the feature checklist as a buying tool.

Per-page pricing is collapsing

Records retrieval has been quoted under $0.40 per page. A feature that costs cents to deliver stops being a reason to pick a vendor.

Agent costs dropped roughly 39x

The compute floor under most legal AI features fell hard in 2026. What one vendor ships this quarter, 3 more ship next quarter.

Foundation model vendors are going vertical

The companies supplying the models are building legal products directly. Feature parity arrives from above, on their schedule, not yours.

What survives commoditization is the part you carry: the malpractice exposure, the bar complaint, the discovery deadline you missed because a tool summarized around a gap. Score that.

The framework

6 axes. Score each 0 to 3.

01Omission behaviorLiability axis

Most eval guides test whether a tool invents facts. In legal work the expensive failure is the clause or precedent it fails to mention, because a missing thing trips no alarm.

Test to run

Hand it a document with a known material clause removed. Does it flag the absence, or answer confidently around the hole?

Red flag

The vendor publishes accuracy numbers and never publishes recall.

02Reproducibility and versioningLiability axis

The practitioner version of this question is blunt: can you reproduce, under oath, the output this tool gave you 6 months ago?

Test to run

Ask which model version served a given output, whether that version is pinned to your account, and whether you get notice before it changes.

Red flag

"We're always on the latest model." That's a non-answer, and it's a hole in your malpractice defense.

03Privilege and data handlingLiability axis

Local and private execution is table stakes now. Firms building their own stacks in 2026 are blocked on privilege questions rather than on model quality.

Test to run

Ask where the document sits at rest, in transit, and in the training pipeline. Get the subprocessor list. Get the retention window in writing.

Red flag

Training on your data by default, with an opt-out buried in a settings pane.

04Supervision fitLiability axis

You're squeezed from both sides. There's pressure to adopt AI and pressure not to file unreviewed output or leak client data. A tool either helps you supervise or it quietly makes supervision harder.

Test to run

Does the tool make human review structurally easy through citations to source, diffs, and surfaced confidence? Or does it emit a finished-looking artifact that invites a rubber stamp?

Red flag

Polished output with no provenance trail. Polish is a supervision hazard.

05Operability without a systems personCost axis

Rank every tool by whether a 4-attorney firm can run it without hiring an engineer. The gap for smaller firms is affordability and operability together, and operability is the half vendors ignore.

Test to run

Ask who owns it when it breaks on a Friday, and walk the actual setup path end to end.

Red flag

Enterprise pricing paired with enterprise assumptions about your IT staff.

06Exit costCost axis

The tool layer is commoditizing, so you will switch. Price the switch now, while you're still the one holding the signature.

Test to run

Ask to export your matters, prompts, and history, then open the export and check what survived.

Red flag

No export, or an export that loses the structure of the work product.

Scoring

What each number means.

0Absent

The vendor cannot answer the question, or answers it with marketing copy.

1Asserted

They claim it verbally. Nothing in the contract, docs, or product confirms it.

2Documented

It's written down in docs or the agreement, and you could hold them to it.

3Demonstrated

You ran the test yourself on your own matter and watched it hold.

The disqualifying rule

A 0 on any of axes 1 through 4 disqualifies the tool regardless of its total. Those 4 are liability axes, and a high total elsewhere doesn't buy back a hole in your defense.

Axes 5 and 6 are cost axes. A low score there opens a budget conversation rather than stopping the deal.

FAQ

Common questions.

Why score liability and operability instead of features?

Features are converging fast. Records retrieval has been quoted under $0.40 per page, agent compute costs fell roughly 39x, and foundation model vendors are shipping legal products directly. The axes that predict whether a tool survives contact with a real matter are liability surface and operability, so those are what this framework scores.

What does a disqualifying score look like?

Score each axis 0 to 3. Any 0 on axes 1 through 4 disqualifies the tool regardless of its total, because those 4 are liability axes rather than preference axes. Axes 5 and 6 are cost axes, so a low score there is a budget conversation instead of a stop.

Isn't hallucination the main risk?

Invented facts get caught because they look wrong. The missing clause, the precedent that never surfaced, the exhibit that was never requested: those read as clean output. That's why omission behavior is axis 1 and why you should test it with a document you've deliberately damaged.

How do I test reproducibility before I buy?

Ask which model version served a specific output, whether that version is pinned to your account, and what notice you get before it changes. If the answer is that they're always on the latest model, you cannot reproduce a 6-month-old output, which matters the day someone asks you to explain how a filing was produced.

Does this framework rank specific vendors?

No. This page teaches the scoring method and stays vendor-neutral on purpose. Our tool directory is where individual tools get listed and compared.

What if a tool scores well but we still can't supervise it?

That's a staffing constraint rather than a software one. A tool that scores 3 on supervision fit still assumes someone at the firm has the hours to do the review. If the review capacity isn't there, the tool transfers risk to you rather than reducing it.

Where to go next

Run the framework on real tools.

Our directory lists legal AI tools with vendor links, so you can score candidates against these 6 axes yourself.

If axis 4 is where your firm keeps landing, the constraint is review hours rather than software. We run dedicated paralegal pods for PI firms, and the team comes with the tools.