baz

Nobody reads 4,207 documents. The question is where you stop.

Every diligence exercise ends when the time runs out, and the report almost never says so. Ranking does not change that — it changes what is above the line when it happens, and it makes stopping a decision somebody made rather than something that occurred.

Findings against documents read

Drag the stopping point. The axis is logarithmic because everything interesting happens in the first few hundred.

app.baz.nml.sa/sets/diligence
69%of findings

after 250 of 4,207 documents

1101001,0004,207
0%25%50%75%100%1101001,0004,207documents read, ranked best-first (log scale)

3/4

Deal-relevant

4/5

Material

2/4

Noted

4 findings sit past your stopping point, including one deal-relevant item at position 1,640. That is the number worth writing into the report, because the alternative is a reader assuming you read everything.

Half the findings arrive inside 88 documents and ninety per cent inside 1,640. Reading 69% of them costs 250 documents out of 4,207 — which is the argument for ranking, and also the argument for being careful about what it lets you skip.

Position 1,640

Side letter varying the principal customer contract, filed with unrelated correspondence and unsigned in the data room copy.

It is the most consequential item in the set and it ranked below sixteen hundred other documents, because it was filed with unrelated correspondence and looks like nothing from the outside.

Why the curve is drawn honestly

A vendor's version of this chart reaches a hundred per cent by document three hundred and stays flat. Ours does not, because that flat line is the claim that ranking has solved diligence, and it has not.

What ranking does is make the first two hundred documents worth far more than the next two thousand. What it cannot do is tell you the tail is empty, and the moment a team believes that is the moment the tail becomes expensive.

Stop after Share of the set Findings surfaced Deal-relevant missed
50 1% 5/13 1
100 2% 7/13 1
250 6% 9/13 1
500 12% 10/13 1
1,000 24% 11/13 1
2,000 48% 12/13 none
4,207 100% 13/13 none

How the order is decided

Most of the value is in the first rule, which is not clever. The fourth is the one that matters most in practice, because it is the default everywhere else.

01 Document type first
A share purchase agreement outranks an invoice, always. Most of the lift comes from this and it needs no cleverness at all.
02 Then anomaly within type
Among sixty supply agreements, the useful ones are those departing from the other fifty-nine. Similarity is cheap to compute and it puts the odd one at the top.
03 Then the questions you asked
A diligence request list is a ranking signal. Documents that answer an open request move up, and requests nothing answers become a finding of their own.
04 Never by date or by folder
Data room structure reflects who uploaded what, not what matters. Reviewing in folder order is the default and it is close to reviewing at random.

What ranking is not

The first of these is the whole page in one line.

Ranking is not reading
Everything below your stopping point is unreviewed however good the order is. The value here is that the decision to stop becomes explicit and dated rather than being what happened when time ran out.
The tail is not empty
The most consequential finding in this set sat at position 1640. It was filed with unrelated correspondence, which is exactly why ranking struggled with it and why a clean curve would be a lie.
Duplicates are most of the pile
Data rooms are full of near-identical versions. Collapsing them is the single biggest reduction available and it is also where a materially different version gets hidden.
Scanned Arabic degrades worst
A document that cannot be read cannot be ranked. Pages that failed extraction are listed separately rather than sorted to the bottom, because the bottom is where they would stay.

The rest of the platform