If you are evaluating an AI legal research tool, citation accuracy is the metric that matters most for professional use. Speed and coverage are table stakes. Citation accuracy is what determines whether the outputs are usable in a professional workflow or whether they require re-verification that consumes the time the tool was supposed to save. This post describes the internal testing protocol we use for Qanooni and publishes it in full so that our early-access partners can evaluate our methodology directly.
Why a citation accuracy protocol is necessary
A language model can produce plausible-sounding legal text that is factually incorrect. The text looks like a citation, references a real statute name, and uses the correct format for a provision number. It may be wrong in the article number, wrong in the paragraph, wrong in stating what the article says, or a reference to a provision that does not exist. Without testing, there is no way to distinguish a model with low hallucination rates on legal citations from one with high rates, because both will produce confident-looking output.
The standard accuracy benchmarks used for general language models do not translate well to legal citation use cases. General accuracy benchmarks typically measure whether the model's answer is broadly consistent with a reference answer. For legal citations, the relevant measure is much stricter: the citation must identify a real provision, the provision must say what the model claims it says, and the provision must be current rather than superseded. A model that scores 90% on a general accuracy benchmark might produce a hallucinated citation rate of 20% on legal provisions, which is not acceptable for professional use.
The structure of our accuracy protocol
Our internal testing protocol runs on a fixed question set that we maintain and update. Questions are structured to require specific statutory or regulatory answers. For each question we have a verified answer prepared in advance, sourced from the primary text, specifying the exact article and paragraph that answers the question. We know the right answer and the correct citation before we run the test.
We test against three failure modes. First, hallucinated provision: the model cites a provision that does not exist in the named statute. Second, misattributed answer: the model cites a real provision that does not actually say what the model claims. Third, outdated authority: the model cites a provision from a superseded version of the legislation where a subsequent amendment changed the applicable rule. Each failure mode has different implications for professional use, but all three produce responses that are not usable without independent verification.
For each question, we record whether the response would have been usable by a practitioner who followed the citation and relied on it without additional checking. This is a strict binary: pass or fail. An answer that is directionally correct but cites the wrong article fails, because a practitioner who follows that citation and quotes it in a document is citing an incorrect provision. A response that declines to answer and says the question requires professional advice is coded separately, not as a failure.
Our current question set
The current question set covers five areas of UAE and DIFC law that represent the highest-volume research categories for our user base. UAE company law, drawn primarily from Federal Law No. 32 of 2021 and its implementing Cabinet Decisions. UAE employment law, drawn from Federal Decree-Law No. 33 of 2021 and Cabinet Decision No. 1 of 2022. DIFC contract and company law, drawn from the relevant DIFC Laws published by the DIFC Authority. DFSA regulatory requirements, drawn from the DFSA Rulebook. UAE data protection, drawn from Federal Decree-Law No. 45 of 2021 on the Protection of Personal Data.
We run 50 questions per area, totalling 250 questions per testing cycle. We run a full cycle before any model update is deployed to users and on a quarterly cadence otherwise. Questions are reviewed and updated each cycle to account for legislative amendments; a question that was valid in a prior cycle may need revision if the applicable provision has been amended.
How we handle legislative amendment tracking
UAE and DIFC legislation is amended with some frequency. Cabinet Decisions can amend implementing regulations without touching the parent statute. The DFSA Rulebook is updated by instrument and the updates are not always reflected in secondary or database sources promptly. We maintain a monitoring process for the source bodies in our coverage: MOCCAE, Ministry of Human Resources and Emiratisation, DIFC Authority, DFSA, and the UAE Official Gazette.
When a source amendment is identified, we review the affected question set for accuracy and update the verified answers before the next testing cycle. We also mark affected questions as pending review in the interim period, so that any response touching an amended area includes a note that the user should verify currency. This is an imperfect solution for the period between an amendment and its incorporation into our test set, but it is preferable to treating potentially outdated provisions as verified.
Outdated authority is the failure mode that is hardest to catch at production time, because the underlying model may have been trained before the amendment and will produce responses that are correct for the prior law. We do not believe any production AI legal tool has fully solved this problem; we are honest with our users that our coverage of recent amendments depends on our monitoring process, which can have a lag. We state this explicitly in the product rather than assuming users will not encounter an amended provision.
Publication of test results
We share test results with our early-access partners on request. We do not publish aggregate pass rates publicly at this stage, because the test set is not standardised externally and we do not want pass rates on our own test to be read as externally validated accuracy claims. This is an important distinction: a tool can achieve a high pass rate on its own test set through test set curation rather than through genuine accuracy. We are designing our test set to be difficult enough to be informative rather than easy enough to look good.
If you are evaluating Qanooni for your practice, the most reliable test is to design your own question set in the areas of law you work in most frequently, run it against both Qanooni and any competing tool you are considering, and compare results using the same three failure-mode criteria we described above. A question set you design is harder to have been inadvertently optimised for than a question set published by the vendor.
What the results have changed in our design
Running this protocol continuously has changed several product decisions. We found early in development that our model had an elevated hallucination rate on DFSA Rulebook module references specifically, because the module naming convention differs from the article numbering used in statutes. We introduced a separate indexing layer for the DFSA Rulebook that treats module names and section numbers as a distinct citation format, which reduced the error rate on DFSA questions substantially.
We also found that questions involving the interaction between UAE federal law and DIFC law had a higher misattribution rate than questions confined to a single jurisdiction. The model was occasionally citing the correct rule for the wrong jurisdiction. We added a jurisdiction disambiguation step to the retrieval process for questions that could plausibly be answered under multiple legal systems. These are design changes driven directly by the accuracy protocol, not by theoretical reasoning about model behaviour.
The protocol is not a finished instrument. We refine the question set, the failure mode definitions, and the evaluation criteria each quarter. We publish the methodology because we believe transparency about how we test is part of what makes the tool trustworthy for professional use, not because we believe the methodology is optimal.