Contract review has a scaling problem that AI tools can help address, but only if the review output is structured well enough to be acted on. When a tool identifies a clause as non-standard, the next question from any senior lawyer is not "is this flag correct?" It is: "non-standard against what benchmark, and where does that benchmark come from?" If the tool cannot answer that question, the flag is not yet usable.
What makes contract review output actionable
We have been working with in-house legal teams in Dubai and Abu Dhabi on what they actually do with AI contract review flags. The workflow is relatively consistent. A junior associate uploads a contract, the tool runs and produces a list of flagged clauses with varying severity levels. The associate reviews the flags, escalates anything material to a senior lawyer, and the senior lawyer decides whether to push back on each point in negotiation.
The breakdown happens at the escalation step. When the associate presents a flag to the senior lawyer, the senior lawyer needs to understand the basis for the flag to make a useful judgment. "The indemnity scope is unusually broad" is the start of a conversation. "The indemnity scope extends to third-party claims arising from ordinary course activity, which departs from the standard in Article 282 of the UAE Civil Code and the precedent framing commonly used in commercial contracts governed by DIFC law" is a brief that the senior lawyer can work with.
The difference in usefulness is almost entirely about sourcing. The first flag requires the senior lawyer to apply their own benchmark. The second gives them the benchmark explicitly, with a traceable source, and the senior lawyer can assess whether that benchmark applies to this transaction.
The risk picture without sourced review
Unsourced contract review creates a specific risk pattern that is easy to miss until it materialises. A junior associate who trusts the AI flag but cannot explain its basis will face a choice when challenged by the counterpart's counsel. If the flag was correct but cannot be sourced, the negotiating position is weakened by the inability to cite the standard. If the flag was wrong, the inability to verify it means the error stays in play longer than it should.
This is not a hypothetical failure mode. In one M&A due diligence we supported for a regional family office acquiring a UAE mainland entity, the initial document review flagged a governing-law clause as problematic but the flag cited no specific source. The clause provided for arbitration at DIAC with substantive law of UAE. The flag was incorrect for the transaction type, but it was not until the sourced review layer was applied that the team had the clarity to set it aside confidently and move forward without requesting unnecessary amendments.
We mention this not to overstate the risk but to illustrate that the downstream cost of an unsourced flag is not just inconvenience. It can generate negotiation friction, introduce delay, or, worse, suppress a valid flag because the team loses confidence in the tool overall after encountering a few that could not be explained.
How sourced clause review works in practice
Sourced review means that each flag arrives with its benchmark explicitly stated and traceable. The benchmark might come from the applicable primary legislation, from standard drafting guidance such as that published by DIFC Authority for DIFC-law contracts, from common market practice in the region, or from the client's own internal contract standards if those have been configured into the tool.
For each flagged clause, the question the review output should answer is: "compared to what?" The comparison source determines the appropriate response. A departure from mandatory statutory language needs to be corrected. A departure from market practice needs to be assessed in context. A departure from the client's internal template requires escalation under the client's own governance. These are three different decisions requiring different inputs. Sourced review separates them.
In Qanooni's review mode, each flag includes the primary source it is measuring against. For a UAE-law governed supply contract, a liability cap clause will be measured against the applicable provisions of the UAE Civil Code on limitations of liability, rather than against a generic "standard indemnity" that is not located anywhere. The associate can follow the citation directly and read the operative text before escalating.
Scale effects on due diligence
The sourcing requirement becomes harder to satisfy manually at scale. A team reviewing 200 contracts in a mid-market M&A transaction cannot apply consistent benchmarks across every indemnity, warranty, and governing law clause manually. The result is drift: some flags reflect statutory standards, some reflect the most recent deal the reviewer worked on, and some reflect a benchmark no one can articulate.
This is where AI review with sourced output has a genuine advantage over manual processes, not just in speed. A tool that applies consistent benchmarks across all 200 contracts and cites the same source each time it flags a particular type of departure produces something more reliable than a team of tired associates applying slightly different implicit standards. The output is auditable in a way that manual review at scale is not.
That said, we are careful not to frame this as a replacement argument. The judgment that a departure from the statutory standard is acceptable in this transaction context, given the parties and the risk allocation, is not a judgment a review tool makes. The tool finds the departure and sources the benchmark. The lawyer decides what to do with it.
Setting up sourced review for a specific jurisdiction
One practical implication of the sourced-review approach is that the tool needs to know which jurisdiction it is benchmarking against. A contract governed by DIFC law is measured against DIFC Contract Law, DIFC Employment Law if employment provisions are present, and DIFC court precedent where relevant. A contract governed by UAE mainland law is measured against the UAE Civil Code, the UAE Commercial Transactions Law, and relevant Federal Decrees. These are not interchangeable, and a tool that treats them as equivalent will produce flags that are wrong in both jurisdictions about half the time.
Qanooni's review engine requires jurisdiction selection at the start of each review session, and the sourcing layer is populated from the applicable body of law. For dual-jurisdiction contracts where DIFC law and UAE mainland law interact, both source sets are active and the output indicates which jurisdiction each flag draws from. This is a design decision that adds a small amount of setup friction in exchange for substantially more useful output.
What sourced review does not solve
Sourced review addresses the "compared to what" question. It does not address the "what does this mean for my client's transaction specifically" question, which depends on facts the tool does not have. A sourced flag that says a warranty basket exceeds the standard under UAE commercial practice still requires a lawyer to assess whether the client's negotiating position, the deal economics, and the counterparty's flexibility make pushing back on that provision worthwhile. That assessment lives outside the review layer.
This boundary is important to state clearly. The value of sourced review is in structuring the conversation between the associate and the senior lawyer, not in replacing it. When the conversation is well-structured, with a precise benchmark and a traceable source, it is faster and more likely to produce a correct outcome. That is the improvement we are building toward.