We ran a structured test of seven AI legal research tools against 30 questions drawn from UAE statute, DIFC legislation, and ADGM rules. The questions were selected because each had a verifiable answer traceable to a specific article or section in a primary legal source. The goal was not to rank tools by accuracy alone, but to understand how each tool handles the citation layer: does it tell you where the answer came from, and is that source real?
The test design
Each of the 30 questions had a confirmed answer already in hand before the test began. We drew questions from four areas: corporate formation and ownership under UAE Federal Law No. 32 of 2021 on Commercial Companies; employment termination procedures under UAE Federal Decree-Law No. 33 of 2021 on the Regulation of Labour Relations; disclosure obligations under the DFSA Rulebook; and arbitration jurisdiction clauses governed by DIFC Arbitration Law. These are routine queries for any commercial practice in the region.
For each tool we recorded: (a) whether a citation was provided at all; (b) whether the citation was traceable to a real, publicly available source; (c) whether the cited provision actually supported the answer given; and (d) whether the citation included enough specificity for a lawyer to go directly to the relevant paragraph. A citation reading "UAE Commercial Companies Law" without an article number scores differently from "Federal Law No. 32 of 2021, Article 83." We counted only the latter as a passing citation.
What the results showed
Only two of the seven tools returned a locatable, article-level citation with every response across all 30 questions. Three tools provided citations on some responses and left others uncited, with no consistent pattern about which question types triggered the citation. Two tools provided answers without any citation, in some cases appending a note that the user should verify with a qualified lawyer.
The more revealing finding was citation quality among tools that did cite. Of the tools that cited on most responses, roughly a third of those citations pointed to the correct law but the wrong provision. One tool consistently cited the parent chapter heading rather than the operative article, meaning a lawyer following the citation would still need to search the chapter. Another tool cited a provision number that had been renumbered in a subsequent amendment, pointing to a now-outdated location.
We are not arguing that these tools are without value. A tool that provides a plausible starting point and flags its own uncertainty is preferable to one that gives a confident wrong answer with no trace. But the workflow implication differs from what the marketing copy typically suggests. "AI-assisted research" can mean very different things depending on how much verification the lawyer still needs to perform after the tool responds.
The partial-citation problem
One pattern that appeared repeatedly was what we started calling the partial-citation response. The tool provides a citation for the general rule but not for the exception or qualifier that governs the specific question. A question about grace periods for corporate filing might receive a correct citation for the base rule, with no citation for the ministerial decision that creates a different timeline in specific circumstances. The answer looked complete; it was not.
This matters more in the MENA context than a general legal AI test might suggest. UAE and DIFC legislation is frequently supplemented by Cabinet Decisions, Ministerial Decisions, and regulatory circulars that are not consolidated into the parent statute. A research tool that queries only the primary legislation layer will produce answers that are technically not wrong but operationally incomplete. In one of our test questions, the correct answer to an employment entitlement question under UAE law depended on a Cabinet Decision that amended the base regulation. Four of the seven tools missed this layer entirely.
What "citation" means for different purposes
Practitioners use citations for three purposes that do not always align: verifying the answer received, including a reference in a memo so the client can check the basis, and locating adjacent provisions that might affect the analysis. A citation format that works for one purpose may not serve the others.
A reference that reads "DFSA Rulebook, COB 3.4" works for internal verification by someone already familiar with the DFSA's module structure. In a client memo it needs expansion. A reference that reads "the Conduct of Business Module" is readable in prose but too imprecise to verify. The tools we tested applied inconsistent specificity standards across these purposes, often defaulting to a level that served internal use but would need significant expansion before appearing in any formal document.
Running your own version of this test
If you are evaluating an AI research tool for your practice, the 30-question structure is not difficult to replicate. Pick 10 questions from areas of law you know well, find the primary source for each answer before you run the tool, and then run it. Check not only whether the answer is correct but whether the citation provided goes directly to the relevant text and whether that text supports the answer.
The tools that performed best in our test shared a design orientation: they treated citation as part of the answer structure, not as an optional annotation. This reflects a product decision, not purely a technical difference. A model capable of citing but that defaults to not citing is making a choice about where to place the verification burden. For professional legal work, that choice carries responsibility.
At Qanooni, this test shaped how we designed the citation layer from the start. Every response returns with its source attached to article or section level, linked directly to the primary text. We built this because the test above represents a minimum bar for professional use, not a differentiator worth marketing.
A note on Arabic-language coverage
Our test covered English-language responses to questions framed in English. Several tools we reviewed have Arabic-language capabilities, and we did not include Arabic-language citation quality in this round. That evaluation is planned separately. Citation quality in Arabic responses is likely to differ meaningfully, because primary source materials across the Gulf are often available in Arabic only, and the Arabic legal corpora behind each tool vary substantially in their coverage of UAE, Saudi, and Egyptian legislation. We will publish those results when complete.