IAT Search now runs on Claude Opus 5.5: sharper legal reading, stricter citation checks
We are delighted to announce that, since 23 September 2026, every analysis on International AT Search runs on Claude Opus 5.5, Anthropic's newest Opus model, in place of Claude Opus 4.8. Research questions, cross-tribunal comparisons, follow-up questions and citation checks all move to the new engine, with no change to your plan or your quota.
We did not make the switch on the strength of a benchmark. We made it after putting both engines through the kind of questions our users bring to the platform, and checking their answers against the judgments themselves.
How we tested
We ran ten research questions through both engines, across five tribunals: the ILOAT, the UN Appeals Tribunal, the UN Dispute Tribunal, the World Bank Administrative Tribunal and the IMF Administrative Tribunal. We chose them to span different degrees of doctrinal consolidation: settled lines of case law, such as the right to be heard before a disciplinary measure or the treatment of hearsay evidence; developing ones, such as a former staff member's standing to challenge the handling of a harassment complaint; and narrow procedural questions on which the case law is thin or unsettled, from documents filed after the rejoinder to the receivability of cross-appeals.
For each question, both engines received the same retrieved judgments and the same instructions, so that the model was the only variable. Every memo was then checked automatically, each verbatim quotation against the source texts and each award in the outcome tables against the operative part of the judgment. Six of the ten went to a blind review: the memos were anonymised and shuffled, and a separate, more capable model (Claude Fable 5.1) verified their claims one by one against the full text of the judgments in our database, nearly 400 claims across the two engines. We spot-checked its findings ourselves.
What changed
| Opus 5.5 compared with Opus 4.8 | |
|---|---|
| Preferred in the blind review | 6 of 6 comparisons |
| Rated higher for accuracy, coverage and practical usefulness | 6 of 6 comparisons |
| Verbatim quotations of the judgments and rules | about twice as many |
| Quotations located in the source texts on first pass | 97% |
Both engines reported outcomes and awards accurately and cited only authorities found in the material in front of them. Where Opus 5.5 makes the difference is in the finer points of legal reading: telling the Tribunal's holding apart from the parties' submissions or an internal appeal body's views, placing each authority in its exact procedural setting, and surfacing the authority that cuts against the position being tested, so that a practitioner can prepare to meet it.
In practice: cross-appeals before the UNAT. Asked when the Appeals Tribunal holds a cross-appeal not receivable, Opus 5.5 placed each authority in its exact procedural setting (Wu, 2013-UNAT-306, for instance, as authority for the cross-appeal of the party against whom a default judgment had been entered), set out the adverse authorities on both sides of the question, and flagged the tension the case law has not resolved between Bagot (2017-UNAT-718) and the later judgments. The reviewer checked 36 of its claims against the judgments and confirmed every one.
Stricter citation checks
Verify, which checks a citation in an analysis against the judgment it cites, also runs on Opus 5.5, and it is more exacting. On a test of fourteen citations, its verdicts matched the previous verifier's on eleven; on the other three it drew finer distinctions, flagging overstatements in how a judgment had been characterised and keeping its assessment to the judgment actually cited.
In practice: Judgment 4502. An analysis cited this UNESCO classification case for the proposition that an internal appeal body lacking technical expertise confines itself to reviewing procedure. The new verifier found the citation only partially accurate and explained why, and the judgment bears it out: the Appeals Board had disclaimed competence in view of the complainant's retirement, it had in fact criticised the substance of the classification audit, and the Tribunal set that criticism aside on the evidence, not for want of competence.
Verification is also faster, at 14 seconds per citation instead of 18, and every check ends in an explicit verdict: when the verifier cannot reach one, the citation is clearly marked as not verified.
Under the hood: hybrid retrieval, now with a citation graph
A better model can only reason over what it is given, which is why the migration comes with a new layer in how we retrieve case law. Retrieval on IAT Search is hybrid. Semantic search runs over the full text of every judgment, lexical search catches exact terms of art, and an AI re-ranker keeps the judgments that actually discuss your issue rather than merely mention it; the model then works from the full text of the judgments selected, not from snippets. The new layer is a citation graph: the precedents most often cited by the selected judgments now join the analysis, because the authorities that anchor a line of case law are frequently the ones later judgments keep citing, whether or not they use the words of your question. On hearsay, for instance, the graph brings in Judgment 203 (1973) and Judgment 999 (1990), foundational due-process authorities that never use the word "hearsay" and that the search itself had not surfaced. The citation graph is live for the ILOAT, and we are extending it to the other tribunals.
What to expect
Analyses are more thorough, about 60% longer on average, and take somewhat longer to generate, typically around three minutes. Every quotation is still checked against the source texts, with anything that cannot be located highlighted for you, and every judgment cited is one click away from its full text. Try it on the question you are working on today.