It reads 5 documents out of 2,412. Which 5 is the whole question.
Everything an assistant like this gets wrong, it gets wrong in one of two places, and neither is visible in the answer. It retrieves the wrong documents, or it writes a sentence nothing supports. So both stages are shown here, including the one that fails.
One question, all four stages
Step through it. Stage two is a fluent draft; stage three is what survives being checked.
Asked
Can we terminate the MSA for convenience, and what does it cost us?
The question is matched against every document in the matter and the highest-scoring spans are pulled into context. Nothing else is read, which makes this the ceiling on the answer rather than a preliminary step.
Master Services Agreement
Clause 21.2 — Termination for convenience
0.94Master Services Agreement
Clause 21.5 — Consequences of termination
0.91Master Services Agreement
Schedule 4 — Exit fees
0.88Variation Agreement No. 2
Clause 3 — amendment to Clause 21
0.84Side letter
11 February 2024
0.79- cut-off — 0.75 · nothing below is read
Board minutes
Meeting of 3 March 2024
0.61 Email thread — renewal
R. Nasser to counsel, 9 Jan 2024
0.58Master Services Agreement
Clause 9 — Payment terms
0.55Variation Agreement No. 3
Clause 2 — notice provisions
0.52This one matters. It shortens the notice period again and never uses the word “terminate” — it says “bring the term to an end”. The ranking scored it on vocabulary and got it wrong.
Statement of Work 7
Annex B — service credits
0.44
The document it did not read
Variation Agreement No. 3, Clause 2 — notice provisions, scored 0.52 against a cut-off of 0.75. It never reached the model, so it left no mark on the answer — no hedge, no caveat, nothing to notice.
It was ranked on wording. The clause says "bring the term to an end" and never uses the word the question used, so it sorted below five documents that matter less.
Why this page shows you the list
Verification cannot save you here. Every sentence in that answer was matched to a real clause in a real document. It was supported, cited, and wrong — because the amending document was never in the room.
The only defence is a lawyer glancing at what was retrieved and noticing an absence. That takes about four seconds and it is the reason the retrieval list is attached to every answer instead of being hidden behind a developer setting.
5
documents reached the model
5
scored below 0.75 and were not read
1 of 5
sentences removed at verification
What it can reach
A short list, deliberately. Every item on the right is a question it will decline rather than approximate.
In scope
- This matter's documents
- All 2,412 of them, across 38,690 pages, in Arabic and English together.
- Your firm's precedent bank
- Agreements your firm has drafted and negotiated before, if you point it at them. Scoped per firm, never pooled.
- Published statute and regulation
- Nizam and implementing regulations, cited by subject. Kept current, with an honest account of what is not published.
Out of scope, by design
- Another client's matter
- Matters are isolated from each other. There is no configuration in which one client's file informs another's answer.
- Court judgments
- Saudi judgments are not comprehensively published. There is no corpus, so there is no search, and it says so rather than inventing one.
- The open web
- It does not browse. Anything it tells you came from a document you gave it or from published law, and both are cited.
Ask in one language about documents in the other
A Saudi matter is bilingual by default: the agreement in Arabic, the correspondence in English, the annexes in whichever the drafter preferred that week.
The question is answered in the language it was asked in, and each citation stays in the language of the document it points at. A clause is not translated on its way into a footnote — you get the words that were signed.
السؤال
ما هي مدة الإشعار المطلوبة لإنهاء العقد؟
مدة الإشعار مئة وعشرون يوماً كتابةً، وفقاً للتعديل الوارد في اتفاقية التعديل رقم ٢.
Cited
1 Variation Agreement No. 2 Clause 3 — English original
What the pipeline does not fix
Verification is a floor, not a guarantee. These four are the ones that survive it.
- Retrieval is the ceiling
- Everything downstream is bounded by what was pulled in. Verification catches sentences with no support; it cannot catch an answer that is confidently wrong because the amending document was never read.
- Verified is not correct
- A sentence matched to a clause is a sentence that reflects that clause. Whether the clause means what you think, and whether it is the operative one, remains a lawyer's judgement.
- It reads what it is given
- A matter with documents missing from it produces answers with those documents missing. The most common failure in the first month of any deployment is an incomplete file, not a bad model.
- Scanned paper is a real problem
- Handwriting, stamps and poor scans degrade badly, and Arabic scans worse than English. Pages that could not be read cleanly are flagged rather than silently skipped.
Taken together: Baz will not write you a sentence with nothing behind it, and that is worth having. It is a narrower promise than "it does the research", and the gap between those two is where the profession's accidents happen.
The rest of the platform
Research
Nizam, regulations and ministerial decisions.
Drafting
From your own precedents, not a generic bank.
Arabic and English
Two versions that say the same thing.
Contract review
Marked against your playbook, not somebody else's.
Document sets
Thousands of pages, ranked by what matters.
Matters
Every document, deadline and hour against the file.
Time and billing
Captured as you work, not reconstructed on Thursday.
Deadlines
Counted from the trigger, not from memory.