Anthropic released over 20 integrations for legal workflows in May, and the headlines were loud: 90.9% on Harvey's BigLaw Bench, 12 role-specific plugins, US law firms scaling across M&A due diligence to employment handbook drafting. The headline UK firms should read is buried deeper: even as these tools perform well on benchmarks, hallucinations are showing up in live legal filings. This is not a future problem. It is happening now. For UK regulated firms — whether law firms answerable to the SRA Code, insurers under FCA Consumer Duty PS22/9, or accountancy practices subject to FRC ISA UK standards — the gap between a lab benchmark and a courtroom filing is where liability lives.
What this story reveals is an industry in a dangerous rush. US Big Law has committed enormous capital to AI, and the sunk-cost pressure is immense. Vendors like Harvey, Luminance, and now Anthropic are marketing confidence. The market rewards the boldest bets. But UK regulated firms operate under fundamentally different constraints: the SRA expects demonstrable competence (SRA Code 2019 Outcome 3.1); the FCA expects firms to manage operational resilience and material risk; the ICO's UK GDPR guidance on automated decision-making carries enforcement teeth. Copying Big Law's all-in approach without understanding these differences is not innovation — it is regulatory exposure.
Trovix's view is blunt: a 90.9% benchmark score is not permission to skip governance. Firms deploying AI at this scale need systematic oversight, not integration counts. This is why we built Trovix Audit — not to slow deployment, but to give firms the documented assurance they need to stay within regulatory bounds. Other products focus on raw performance (Harvey's outcome quality, Luminance's pattern matching). We focus on the invisible problem: how do you know when your AI is hallucinating in ways that matter? How do you prove you caught it before it breached a client duty? How do you show the regulator you had a system, not just a tool? The firms winning this race long-term are not the ones with the most integrations. They are the ones with the cleanest audit trail.
If you run a mid-market law firm, insurance broker, or accountancy practice, here is what to do now: (1) Do not wait for a scandal to force the question. Run a 90-day AI governance baseline using a framework like ISO 42001 or the EU AI Act's risk tiering — understand what you are actually deploying and where. (2) When you evaluate new legal AI, ask the vendor for their hallucination rate on your actual document types, not their benchmark scores. (3) Build oversight before you build at scale. Trovix Watch can alert you to regulatory changes that reframe AI risk (like new SRA guidance or FCA updates); Trovix Audit gives you the compliance dashboard you will need when the regulator calls. (4) Treat AI as a junior associate who is brilliant but occasionally lies. You still need supervision. The firms that thrive in this phase will be the ones disciplined enough to say no to the features they cannot yet control.
Source: Fortune