Anthropic's announcement of 20+ law firm integrations and Claude Opus 4.7's 90.9% score on Harvey's BigLaw Bench looks compelling on a slide deck. But read the real story: judges are already sanctioning law firms for AI hallucinations in filed documents, yet the industry is accelerating deployment into M&A due diligence and employment handbook drafting — precisely the workflows where a single missed clause or invented precedent causes commercial or employment law disasters. For UK regulated firms, this is not an innovation story. It is a governance failure waiting to happen, and your SRA Code obligations and FCA Consumer Duty PS22/9 expectations make you liable when it does.
This reflects a pattern we are watching across legal AI: vendors are optimising for benchmark scores and productivity metrics while the actual failure modes — hallucination, citation drift, context collapse in complex document sets — remain unsolved. Harvey, Luminance, and now Anthropic are all racing to embed deeper into core workflows before their limitations become undeniable. The message from Big Law is clear: the competitive pressure to deploy fast exceeds the willingness to wait for real trustworthiness. But mid-market UK firms do not have Big Law's liability insurance, reputational buffer, or in-house AI governance teams. When things break, they break harder.
Here is Trovix's direct view: benchmark scores are not governance. A system that scores 90.9% on legal reasoning still hallucinates, still invents facts, and still fails catastrophically on edge cases that matter most. If you are implementing AI into document review, due diligence, or compliance workflows, your starting point must be: what happens when this gets it wrong? How do you catch it before it reaches a client or a court? That is why we built Trovix Audit to sit *between* the AI system and your workflow — not to trust the AI, but to verify it continuously against your firm's governance standards and regulatory obligations. Your alternative is what Big Law is doing: shipping fast, hoping, and hiring outside counsel when the inevitable failure arrives. Trovix Sift handles document extraction and classification differently too — building confidence scores, transparency logs, and audit trails that actually survive FRC ISA UK scrutiny, not just internal testing.
If you are a mid-market law firm, insurer, or accountancy practice, do this now: audit what AI systems you have already deployed and where. For each one, ask: could a hallucination here cause client harm, regulatory breach, or litigation? If the answer is yes, you need governance, not just the tool. Do not wait for your regulator to ask. Do not assume your vendor's benchmark score means your client files are safe. The firms getting sanctioned thought they were fine too. Build a proper AI governance framework — start with Trovix Audit to map what you are running and where your real risks are, then implement verification at the point of use. Speed matters. Trustworthiness matters more.
Source: Fortune