Anthropic's new legal integrations are genuinely capable. But US Big Law's rush to adopt them while hallucinations are still being discovered in actual filings tells you everything you need to know about where the industry's priorities are wrong.
AI Governance  Trovix AriaLegal · Financial Services · Insurance

Anthropic released over 20 integrations with legal workflow tools and Claude Opus 4.7 scored 90.9% on Harvey's BigLaw Bench. On paper, this is the kind of news that makes UK mid-market firms anxious—American competitors now have a tightly integrated AI system designed specifically for M&A due diligence and employment drafting. But the Fortune article buries a more important fact: hallucinations are already showing up in legal filings generated by large language models. This isn't theoretical risk. This is happening now. UK regulated firms—bound by SRA Code provisions on competence and diligence, by FCA Consumer Duty PS22/9 obligations on treating customers fairly, and by their own professional liability exposure—cannot simply copy what US Big Law is doing and hope the benchmark scores protect them.

What we are watching is a familiar pattern: technology vendors release capability benchmarks that measure narrow, controlled tasks. Legal teams report productivity gains in real-world workflows. Regulators and courts have not yet caught up to ask hard questions about liability and accountability. So firms adopt at speed. Then, months or years in, hallucinations, bias, or drift emerge in documents that matter—and by then the AI system is embedded in critical workflows. Harvey scored 90.9% on Harvey's own benchmark, which is useful information but incomplete. It tells you nothing about how Claude performs on non-standard contract language, on legacy documents with poor OCR, or on the long-tail edge cases that actually cause legal problems. Anthropic's integrations will accelerate this cycle: they make AI easier to use, which means less expert human review, which means more risk gets baked into business-critical documents.

Trovix's view is this: the benchmark is not the work. A 90.9% accuracy score on a test set is not the same as a tool you can trust unsupervised in a live M&A due diligence process. This is where we differ fundamentally from the Harvey-and-Anthropic approach. We do not sell speed; we sell verified AI output. Trovix Aria is purpose-built for exactly this problem—it uses retrieval-augmented generation to keep AI outputs tethered to your actual knowledge base, your precedents, your firm's way of working. It reduces hallucination risk by construction, not by hoping the benchmark generalizes. And Trovix Audit exists because governance is not optional: UK firms must be able to audit, trace, and justify every AI-assisted decision a fee-earner makes. That is not a nice-to-have. That is SRA Code, that is FRC ISA UK, that is professional indemnity insurance. Firms that adopt Anthropic's integrations without this layer are accepting tail risk they do not fully understand.

Here is what a mid-market law firm, insurance broker, or financial services practice should do right now. Do not wait to see if US Big Law's hallucination problem spreads to your sector—assume it will. Before you sign up for Claude plugins or any other integrated legal AI system, ask your vendor three questions: (1) Can you audit and trace every AI-assisted output to its source material and your firm's approved knowledge base? (2) What is your false-positive rate on non-standard or edge-case inputs? (3) How do you handle liability when the AI generates incorrect advice? If the answer to any of these is 'the benchmark shows it works,' you have your answer about whether to proceed. Second: invest in AI governance infrastructure now, before you scale adoption. Trovix Audit is one approach; there are others. But the firms that will win this cycle are not the ones that adopted AI fastest. They are the ones that adopted AI safely—with human review loops built in, with output verification, with audit trails, and with genuine understanding of where the tool can and cannot be trusted. Speed without governance is just liability with better marketing.

Source: Fortune

Related Trovix product:

Trovix Aria →Book a demo →