How-to guideAug 14, 2026by Docurensic Team9 min read

How to Test a Document Forensics Tool Before You Buy It

A vendor-neutral evaluation protocol you can run in a week: four corpora, a blind submission, false positives measured separately from detection, and the five failure-mode files that reveal more than the other three hundred.

How to Test a Document Forensics Tool Before You Buy It
In this guide
  1. Key takeaways
  2. Before you start: the rules
  3. Step 1 — Build four corpora
  4. Step 2 — Run it blind
  5. Step 3 — Record four things per document
  6. Step 4 — Compute the numbers that matter
  7. Step 5 — Test the failure modes deliberately
  8. Step 6 — Interpret honestly
  9. Step 7 — Write the memo
  10. Frequently asked questions

Nearly every evaluation of a document-forensics product goes the same way. The vendor supplies a demo file that lights up beautifully. The buyer supplies three documents they already know are fake. Everything detects everything, everyone is pleased, and the first real disagreement happens in production four months later when the tool flags a genuine document from your largest customer.

The problem is not that anybody lied. It is that the test measured the wrong thing. Detection on known fakes is the easy half. What determines whether a tool is usable is what it does with the ordinary, boring, genuine documents that make up 99% of your volume — and almost nobody tests that, because it feels like testing nothing.

This is a protocol you can run yourself, on your own material, in about a week. It is deliberately vendor-neutral; run it against us and against whoever else is on your shortlist. We would rather compete on this than on a demo.

Diagram: the four-corpus benchmark structure
Four corpora, two error types, one decision

Key takeaways

Before you start: the rules

Two constraints, both non-negotiable.

Do not manufacture fraud using real identities. When you need altered documents for testing, alter documents belonging to your own organisation, or to consenting colleagues, or synthesise them entirely. Never use a customer's real statement as the base for a forgery, even internally. It is the kind of artefact that is impossible to explain afterwards.

Get the data handling straight before anything leaves the building. Test documents are usually real customer records. Whatever legal basis, retention terms and deletion commitments you would demand in production, demand them for the evaluation, in writing. A vendor who is casual about this during a sales process will not improve later.

Step 1 — Build four corpora

Corpus A: genuine, ordinary (aim for 200+). A representative sample of real documents you have already accepted, spanning your actual mix of document types, issuers, formats and quality. Include the ugly ones: phone photos of printouts, faxes, third-generation scans, documents produced by the one accounting package your smallest customers all use. This corpus is the point of the whole exercise.

Corpus B: genuine, unusual (aim for 30+). Real documents with legitimate oddities. A statement re-saved through three different tools. A contract signed on an e-signature platform, so it carries several incremental updates for entirely proper reasons. A scan from a fifteen-year-old office machine. These are your false-positive landmines, and they exist in production whether or not you test them.

Corpus C: known fraud (as many as you have). Anything you caught, anything a chargeback or a default later proved out, anything an investigator confirmed. Most teams have fewer of these than they expect — twenty is a respectable number, and if you have zero, that is itself informative.

Corpus D: constructed forgeries (aim for 40+). Since real confirmed fraud is scarce, build your own across a difficulty ladder:

Log exactly what you changed in each one. The log is what lets you tell a real detection from a lucky guess.

Step 2 — Run it blind

Shuffle all four corpora into one pile before submitting. Do not tell the vendor which is which; do not tell your own reviewers either. If the evaluation is run by the person who built Corpus D, they will unconsciously read the outputs generously, and this is not a criticism of anyone — it is why blinding exists.

Submit through the API you would actually use in production, at something near production concurrency. A tool that answers in four seconds for one document and ninety for a batch of fifty is a different product from the one in the demo.

Step 3 — Record four things per document

For every file, capture:

  1. The verdict or score the tool returned.
  2. The named findings — what specifically it says is wrong, and where on the page.
  3. The time taken, end to end.
  4. Whether a reviewer could act on the output without contacting the vendor. Score this crudely: yes, partially, no.

That fourth column is the one people leave out and the one that predicts whether the tool is still in use a year later.

Step 4 — Compute the numbers that matter

False-positive rate on Corpus A. Genuine documents wrongly flagged, divided by Corpus A size. This is the number your operations team will live with. At a 3% false-positive rate on 4,000 documents a month you have created 120 investigations a month out of nothing. Decide what rate you can afford before you see any vendor's result, and write it down, or you will rationalise whatever number appears.

False-positive rate on Corpus B, separately. If a tool is clean on A and noisy on B, it has learned "normal" too narrowly, and your legitimate edge cases will suffer.

Detection rate on C and D, broken out by difficulty tier. Aggregate detection across tiers is a meaningless average — 100% on crude and 20% on scanned is a completely different product from 60% across the board.

Evidence quality. The proportion of true detections where the tool named a specific, checkable finding rather than returning a bare score. We would argue this should be a pass/fail gate rather than a metric, but reasonable people weigh it differently.

Reviewer agreement. Have two people independently judge a sample of the flagged documents on the evidence provided. If they disagree often, the evidence is not doing its job regardless of what the score says.

Step 5 — Test the failure modes deliberately

Five files that reveal more than the other 300 combined:

The theme is that "clean" and "could not analyse" must never be the same answer. Silent degradation is the most dangerous behaviour a verification tool can have, because it is invisible exactly when it matters.

Step 6 — Interpret honestly

Some guidance on reading your own results, drawn from watching a lot of these evaluations:

Nobody detects everything, and any vendor claiming otherwise on your Corpus D has probably matched a signature rather than reasoned about the document. Test that suspicion by making a small variant of a detected forgery and resubmitting.

A tool that flags more is not better. Detection rate and false-positive rate move together; you can reach 100% detection by flagging everything. Compare vendors at a matched false-positive rate, or you are comparing sensitivity settings rather than products.

Weight the scanned tier by how your volume actually arrives. If 70% of your documents come in as phone photographs, a product that is superb on native PDFs and mediocre on captures is the wrong product for you, whatever the headline number says.

Look at what it says about the genuine documents it got right. A good report on an authentic file still tells you something — how it was produced, what it checked, what it could not check. Silence on a clean document is a missed opportunity to build the analyst's judgement.

Step 7 — Write the memo

One page, before anyone negotiates: the false-positive rate you observed and the operational cost that implies, detection by tier, evidence quality, the failure-mode results, and the two or three document classes where the tool was clearly weak. Circulate it to the people who will run the process, not just the people who will sign for it.

If a vendor will not let you run this — on a meaningful volume, through the real API, with your own documents — that refusal is your answer. We publish this protocol partly because we believe we do well on it, and partly because a market where buyers run real tests is a better market for the products that deserve to win.

Frequently asked questions

How many documents do I need to evaluate a fraud detection tool?

Roughly 200 genuine documents, 30 legitimately unusual ones, whatever confirmed fraud you have, and 40 or so constructed forgeries across a difficulty ladder. Below about 100 genuine documents, your false-positive estimate has too wide a margin to compare vendors on.

Should I create fake documents to test with?

Yes, and carefully. Alter your own organisation's documents or synthesise them entirely, never a real customer's records, and keep a log of every change you made. Build them across several difficulty levels, because crude and careful forgeries measure very different capabilities.

What false-positive rate is acceptable?

That depends on your volume and the cost of an investigation, which is why you should set the threshold before testing. Multiply your monthly document volume by the rate to get the number of unnecessary reviews per month, then decide whether your team can absorb it.

Can I trust a vendor's published accuracy figures?

Only as a starting point. Published rates are measured on corpora the vendor selected, using their own definition of a detection, and are rarely comparable between products. A blind test on your own documents is the only figure that describes your situation.

Check the PDF you are holding

Run a free PDF X-Ray in your browser — it recovers text from the file’s earlier revisions, so you can see what a value was before it was changed. No account needed.

Open the free PDF X-Ray

Keep reading

How-to guideAug 17, 20269 min

The Free Document Forensics Toolkit

ExifTool, pdfid, qpdf, mutool, FotoForensics and the rest — what you can genuinely establish about a suspicious document with software that costs nothing, in what order, and the four things free tooling cannot do.

How-to guideAug 26, 20267 min

Where to Report Document Fraud: Who Takes the Case

The bank first, within hours, because that is the only step that recovers money. Then the national centres in the US, UK, Canada and Australia — and the four reports almost nobody files that change outcomes most.