Inside the work
Why AI fails at legal reasoning
Law6 min read
Inside the work
Law6 min read
Ask a model a legal question and you will usually get something that reads like advice. Whether it is advice is a different matter.
Legal writing has a strong house style. Numbered reasoning, hedged conclusions, the confident register of someone who has considered the alternatives. Models learned that style thoroughly, and they reproduce it whether or not the underlying analysis is sound.
The result is output that passes the sniff test of anyone who is not a lawyer, and fails immediately for anyone who is. This is a much harder failure to catch than an obviously wrong answer, and it is why evaluation in this domain cannot be crowdsourced to generalists.
It passes the sniff test of anyone who is not a lawyer, and fails immediately for anyone who is.
The most frequent error we see is not misstating the law. It is answering confidently under the wrong law.
Asked about unfair dismissal without a jurisdiction specified, a model will typically produce a coherent answer drawn from wherever its training data was densest, and will not flag the assumption. A competent adviser's first move is to ask which jurisdiction. That instinct — knowing that the question is under-specified — is one of the most useful things a practitioner can teach an evaluation set.
In drafting, the interesting failures are quieter still. A clause can be grammatical, conventional, and allocate risk in a way no competent adviser would accept on those facts. Nothing about it looks wrong. It simply is.
Testing for that requires someone who has negotiated the clause. They know which words the other side will push on and what the market position is — knowledge that does not exist in any public corpus because it lives in the redlines nobody publishes.
It means legal evaluation cannot be reduced to a checklist. The task is to say why an answer is professionally unsafe, in terms a research team can act on, and that is a judgment only a practitioner can make.
It is also why we pay legal reviewers what we do. The scarce input is not time. It is the judgment.
Inside the work
What the work actually looks like, hour by hour, for a practising clinician doing six hours a week around a hospital job.
8 min readInside the work
How a submission is judged, with worked examples of a weak review and a strong one on the same answer.
5 min readInside the work
Where recall stops being enough, and engineering judgment starts — and why standards-bound work makes such good evaluation.
6 min readApplications reviewed within 48 hours