Documentation

Everything you need to run Docurensic

From your first scan to the forensic engines behind it, the standalone tools, the automation that runs after every verdict, the full API, real-world use cases, and how your account and data are handled. Pick a tab — every page here mirrors the running code, nothing is aspirational.

START HERE

What Docurensic is

Docurensic is a document forensics service. Upload a PDF, Office file, email, image, or archive and it inspects the file the way a document examiner would — metadata, fonts, revision history, embedded images, signatures, and the numbers on the page — then returns a plain-English report with a forensic evidence score and a clear operational action. Everything works from your browser — no install, no plugin; documents are analyzed server-side in isolated processing.

What it's good at

Catching altered invoices and bank statements, forged certificates and IDs, phishing / business-email-compromise messages, malware hidden in Office macros, and AI-generated text or images — the everyday fraud that arrives as a file.

What it does not claim

Docurensic surfaces forensic red flags; it does not certify a document as genuine. A low score means no evidence of manipulation was found, not a guarantee. Findings and actions support a human decision; they are not proof of authenticity.

Who it's for

Freight brokers vetting rate confirmations, lenders reviewing pay stubs and invoices, insurers checking certificates, HR teams onboarding with W-2/W-9 forms, and anyone who receives documents they can't take on trust.

01

Run your first scan

Three steps. Simple files often finish in seconds; large or complex files can take longer. Everything happens from the Documents page — the first item in the sidebar.

1

Upload a document

Drag a file onto the upload box, or click Browse files. Docurensic accepts PDF, Word / Excel / PowerPoint (including legacy and macro-enabled), .eml / .msg email, images, CSV/TSV, and .zip / .7z / .rar archives, up to 150 MB.

The file uses isolated temporary processing that is deleted after analysis. In storage mode, the analyzed original is automatically encrypted in File Vault when vault storage is available and quota allows. Storageless creates no new source copy; an existing vault or connected-storage file remains there. Its slim history record retains filename/basic scan metadata, forensic verdict, decision action, findings, safe document type, and verification hash.

2

Run one consistent analysis

There is no analysis-depth selector. The applicable forensic engines always run. When the document has bounded extractable or OCR text, the independent AI review runs in the same flow and stays advisory.

3

Read the report

You get an overall action (Accept, Review, or Reject), a separate forensic authenticity result and 0–100 evidence score, a separate threat level, and every finding ranked by severity with a plain-English explanation. Click Detailed report for the full breakdown, or find it again any time in Scan history.

02

AI analysis when document text is available

Every scan runs the applicable forensic engines. When bounded native or OCR text is available, it also runs an independent AI-powered review — there is no depth selector or AI tier to choose. A no-text document reports that the review was unavailable instead of calling a model with empty input. AI never changes the forensic score or verdict.

Call 1 · classification

Document type

An AI model classifies the extracted text into one of 33 canonical document types (invoice, bank statement, COI, deed, …) with a schema-enforced answer — or Undefined when nothing fits.

Call 2 · evidence graph

Type-aware checks + entities

The model runs only the integrity checks that apply to the classified type — financial arithmetic, dates and validity, party/issuer consistency, overall coherence — and returns one plain-language reason per check, its own red flags and fraud read, plus structured entities (emails, phones, IDs, accounts, addresses, domains, names) for case linking.

Privacy

Text only, never the file

Both calls see extracted text, metadata, and the engines' findings — the document file itself is never uploaded to the model.

The AI is advisory, always. Its result lands in a separate section of the report (document type, a 0–100 AI fraud score, a two-sentence summary, and red flags) and is never merged into the forensic scores or verdict. If the AI call fails, the forensic scan still stands — the report just carries an ai_analysis.error instead. Every AI call is logged to your usage page with tokens and cost.
03

Reading your report

The report leads with one operational action while keeping forensic authenticity, file threat, account policy, and advisory AI visibly separate.

Overall action & forensic evidence

The headline action is Accept, Review, or Reject. Directly below it, the forensic authenticity verdict and 0–100 evidence score explain the tamper assessment. Threat or account policy may make the action stricter without rewriting that evidence.

ACCEPT REVIEW REJECT

Threat level

A separate malware / phishing / active-content axis. A clean, authentic-looking document can still carry a high threat level if, say, it hides an auto-executing link or a macro payload.

Document identification

A single block reconciling every engine's type vote (keyword classifier, recognizer, field & type engine) into one headline document type, with per-engine confidence shown side by side.

Ranked concerns & evidence

Material findings are grouped into concerns with severity, location, plain-English evidence, benign context, and a next step. Lower-level signals remain available as drill-down, never as a bare score.

Field analysis

Typed business fields the engine extracted — invoice/PO numbers, amounts, dates, IBANs, MC/DOT numbers — each marked validated or unvalidated (e.g. a failing IBAN checksum).

AI review

When document text is available, Docurensic AI's independent read — document type, per-check reasons, red flags, and its own advisory score — appears in its own section with a token/cost footer, kept visually and computationally separate from the forensic verdict. A no-text report says the review was unavailable.

04

Finding your way around

The sidebar groups everything into Workspace, Analysis, and Organization.

Documents

Upload new files and search your full scan history. Your home base.

Cases

Group related documents (an application packet, a shipment's paperwork) so cross-document links between them are easy to follow.

Analytics

Volume, verdict mix, and the flags firing most often across your scans.

Review inbox

Documents whose final action is Review or Reject and still need a human disposition — approve, reject, or escalate.

File Vault

Your encrypted document storage — retained scan originals and files you upload directly, organized in folders, only you can open them.

Tools

The hub for every standalone tool: URL analyzer, Company analyzer, Metadata lab, Image forensics, Signature checks, Compare, PDF Studio, and the X-Ray Engine workbench.

Workflows

Automate what happens after a scan — "when a document scores high, alert the fraud channel and quarantine it". Built visually, no code.

Connections

Everything Docurensic talks to: watched Dropbox/Drive folders, Slack, Teams, Discord, Telegram, Google Chat, email lists, automation platforms, and your own webhooks.

Notifications

In-app alerts from your workflows, cases, and the system.

Settings

Storage mode, sensitivity, industry preset, detection controls, feature toggles, API access, data export, and account deletion. Covered in Account & privacy.

05

Automate what happens next

Docurensic doesn't stop at the verdict — you decide what happens next once, and it runs on every document after that, no code. In the visual Workflows builder you write rules as when a document event happens, if it matches your conditions, then run steps — alert a channel, tag, quarantine, open a case, run a tool, or hand off to another workflow. Save your Slack channels, email lists, cloud folders and webhooks once in Connections and reuse them by name, and let new files from a watched Dropbox or Drive folder analyze themselves.

The full trigger, condition and action catalog — with branching, schedules, formulas and connectors — is in the Automation tab.

The full technical reference: what each engine checks, how forensic authenticity, file threat, and account policy produce one operational action, and the live catalog of every signal the engines can raise. Jump to a section:

01

What Docurensic checks

Multiple forensic categories, dozens of individual signals. Most tampering leaves traces across more than one.

Structure & metadata

Counts revision markers, checks DocInfo for contradictions. The trailer /ID pair reveals re-saves a document claims never happened.

%%EOF count/ID pairModDate check

Font fingerprinting

Every embedded font carries a unique subset tag. Different prefixes for the same base font on one page signal spliced content from an editor.

Subset tagsEditor detectionFont catalog

Revision history

PDF stores edits by appending new content. Multiple %%EOF markers mean prior revisions exist that can be recovered — a clear sign of editing.

Incremental updatesPrior revisions

Content analysis

Extracts all text and checks for extractability — image-only scanned documents get flagged as needing separate OCR verification.

Text extractionScanned detection

Editor fingerprinting

Known PDF editors — Acrobat, Photoshop, GIMP, iText, PDF-XChange — leave traces in producer and creator fields that signal post-creation modification.

Producer fieldCreator fieldTool detection

Risk scoring

Forensic signals are weighted by severity, capped per category, then gated into an authenticity verdict. The final operational action separately accounts for file threat and explicit account policy.

0–100 evidence scoreSeparated axesOne action
02

Six engines, one report

Every upload is routed by file type to the engines that understand it. All write to the same normalized report shape. Findings remain isolated as authenticity, threat, AI-content, identity, or account-policy information before one decision contract produces the operational action. Each engine is import-guarded and failure-isolated — a missing dependency or runtime error degrades that one engine, never the whole scan.

pdf_forensics

The primary PDF engine — structural integrity, metadata, fonts, raster forensics, signatures, forms, embedded content, and semantic field checks.

axis · forgery + threat

multiformat

Office (OOXML + legacy + macro-enabled), email, archives, spreadsheet/tabular fraud, and barcode/MRZ cross-validation.

axis · forgery

ai_text / ai_image

Stylometric AI-written-text detection and synthetic-image detection — the content-trust axis, run on the relevant upload types and exposed directly via the API.

axis · ai_content

recognition

Multi-source ML document recognition and type fusion (magika, libmagic, layout, text) — the identity axis. Light sources run inline in every scan.

axis · identity

url_reputation

Passive URL / phishing reputation and typosquat analysis on any links a document or email contains.

axis · threat

intel

Cross-document intelligence layered on top of every scan — submission velocity, near-duplicate detection, entity-link fraud-ring detection, and the template registry.

Supabase-backed
Upload Routed to engines Signals separated by authority Decision + case file built Report saved (file deleted) Action + evidence returned
A heavier multi-source recognition engine is available on demand via POST /api/recognize — it fuses up to eleven optional ML / forensics sources (magika, DiT, docling, OCR, a vision-language model, pyHanko signatures) for a second opinion on document type and authenticity. See the API tab.
03

PDF forensics engine

Runs on every .pdf upload. Every check is isolated — a single failing check is recorded as an error and never crashes the request.

Metadata & timeline

Producer/Creator/Author/Subject/Keywords, custom Info-dictionary keys, XMP-vs-DocInfo consistency, CreationDate/ModDate logic (doctype-aware — a certificate's future expiration date isn't treated as back-dating).

Structural integrity

Incremental-update detection, object/stream/object-stream counts, xref-repair detection (the parser silently reconstructing a damaged table), trailer /ID consistency, linearization.

Font & text analysis

Font inventory, embedded-vs-subset, digit-glyph remapping (a font that visually draws one digit but declares another in its ToUnicode map), Unicode anomalies — directional-override controls, zero-width characters, non-Latin digit homoglyphs.

Image & raster forensics

Error-Level Analysis, copy-move/clone detection, JPEG-ghost/double-compression, quantization-table analysis, DPI consistency, embedded-image EXIF/GPS/editor-software extraction.

Revision history

Prior-revision recovery and diff, synthetic-scan detection (a born-digital PDF re-rendered to an image to dodge raster checks).

Digital signatures

Presence, cryptographic validity, byte-range coverage, shadow-attack detection, certification vs. approval signatures.

Forms

Full AcroForm field-tree walk — a field hidden from the viewer that still carries a value, or template placeholder text left unfilled.

Embedded content

Attachment bytes are extracted (not just listed), magic-byte-vs-extension disguise detection, and recursive re-scan through the multi-format engine for anything it understands.

URLs & links

Every URL classified as visible (a Link annotation or printed text) or hidden (reachable only via OpenAction/embedded action — never shown to a reader), plus IP-literal, punycode, and shortener detection.

Semantic & field logic

Cross-field arithmetic, IBAN/routing-number/VIN checksum validation, entity extraction (emails, phones, MC/DOT/policy numbers), template pHash matching against a registry.

OCR fallback

Image-only pages get OCR'd so the content/logic passes still run on scanned documents, not just born-digital ones.

Watermark & template

Expected-watermark presence by document type, and first-page perceptual-hash matching against registered issuer templates.

04

Multi-format engine

Handles everything that isn't a PDF: Office documents (including macro-enabled variants), email, archives, spreadsheets, and standalone images.

Office (OOXML + legacy)

.docx/.xlsx/.pptx/.docm/.xlsm/.pptm and legacy .doc/.xls/.ppt — core/app properties, track-changes recovery, hidden content, VBA macro detection, external data links, formula-vs-cached-value reconciliation.

Email (.eml/.msg)

SPF/DKIM/DMARC header parsing, From/Reply-To/Return-Path mismatch (the BEC redirect shape — high only when Reply-To goes to free webmail), display-name brand spoofing, punycode domain detection, attachment recursion.

Archives

zip/7z/rar entry listing with a compression-ratio guard against zip bombs.

Spreadsheet & CSV fraud

Benford's law (chi-square test on a real sample), balance reconciliation, duplicate/round-number clustering, sequence-gap detection, future-dated transactions.

Barcode / QR / MRZ

1D/2D decode, TD3 passport MRZ check-digit validation, and QR-payload-vs-printed-text reconciliation — a forger who edits the printed number and forgets the code is caught here.

Standalone images

Routes through the AI-image detector, barcode/MRZ decode, and OCR-based QR reconciliation.

Scoring discipline: like the PDF engine, verdicts here are gated on a decisive signal, corroboration across two or more categories, or a calibrated score — not any single noisy medium-severity finding.
05

AI / content-trust engine

A separate detector for synthetic content and unsafe links, run automatically on image uploads and available directly for text/URL analysis.

AI-image detection

ELA, noise-residual, and frequency-domain artifacts, plus known generator metadata signatures.

AI-text indicators

Stylometric scoring across seven weighted signals — sentence-length burstiness, AI-phrase marker density, transition-word openers, opener diversity, n-gram repetition, punctuation style, and contraction rate — plus an optional language-model perplexity + burstiness pass. Every signal is reported individually so a verdict is explainable, not a black box. Deliberately conservative — informational, never a standalone accusation, because of well-documented false-positive rates on formal and non-native-English writing.

URL reputation

Domain age heuristics, typosquat/brand-lookalike detection, and a hook for an external Safe Browsing lookup — applied to every URL a scan or the email analyzer surfaces.

Text analysis is exposed directly — POST /api/analyze/text — for use independent of a file upload (see the API tab); for standalone URL analysis use the URL analyzer page.
06

Recognition & fusion engine

A second opinion on document type and authenticity, fused from up to eleven independent sources — each optional, each skipped gracefully if its dependency isn't installed. A type verdict is corroboration-gated: two sources agreeing produces high confidence; a single source is capped and flagged uncorroborated_type.

Format identification

ML-based format ID (magika) cross-checked against libmagic — disagreement between the two, or against the declared file extension, is a masquerading signal.

PDF forensics front door

PyMuPDF renders every page natively, decides born-digital vs. scanned, and extracts provenance metadata and font-embedding data in a single pass shared by every other source.

Type classification

A 16-class document-type model (DiT) plus a vision-language read (Qwen2.5-VL) for free-form type, subtype, and key-field extraction.

Layout & tables

Reading order, layout regions, and table structure via docling, with deepdoctection as an independent full-pipeline corroborator.

Conditional OCR

Runs only on image-only pages by default; in always mode it also powers an embedded-text-vs-OCR tamper diff — a mismatch means the visible page doesn't match the content stream.

Signature integrity

pyHanko reports embedded PDF signature count, coverage (entire file vs. partial), and cryptographic intactness.

Exposed as POST /api/recognize. The vision-language and layout sources are the heaviest; pass disable=vlm (or any comma-separated source list) to skip them on latency-sensitive calls.
07

Field & type engine

Runs inside every scan. Extracts typed business fields, commits to a document type only when independent evidence corroborates (declining to guess otherwise), and tracks per-vendor term drift over time. Its output populates the report's document-identification and field-analysis blocks.

Field extraction

A broad set of field types — invoice/PO numbers, amounts, dates, IBANs, routing numbers, MC/DOT numbers, VINs, policy numbers, addresses, and more — each with a validated / unvalidated flag.

Type classification

Corroboration-gated: a document whose term/field/structure signals don't agree is reported as unclassified rather than forcing a low-confidence guess. Composite documents (multiple types in one file) are flagged as such.

3-way identification

Reconciles its own type vote against the keyword classifier and the recognizer into one headline type per report, with per-engine confidence shown side by side.

Vendor-drift baseline

Per-user, Supabase-backed term-frequency baseline per vendor; a document missing terms that vendor's history always includes, or carrying terms it's never used, is flagged as drift once enough history exists.

Cross-document field consistency

A shared invoice/PO/amount field that disagrees in value across two related documents is a core altered-document or double-brokering tell.

08

Cross-document intelligence

Runs after every scan, comparing the current document against prior submissions — intentionally not scoped to a single account, since fraud rings span customers.

Submission velocity

The same file (by SHA-256) submitted repeatedly in a short window.

Near-duplicate detection

SimHash comparison finds a lightly-edited version of a document that's byte-different but structurally the same.

Entity-link / fraud-ring detection

Shared MC/DOT/policy/IBAN/VIN/routing numbers, emails, or phones across otherwise-unrelated documents.

Template registry

Known-good issuer layouts (page-1 perceptual hash) that PDF template matching checks against.

Review queue

Flagged (review/reject) reports awaiting a human disposition — approve, reject, or escalate.

09

Format coverage

The formats the engines accept — routed automatically by extension and magic bytes, so a disguised file is scanned as what it really is, not what it claims to be.

PDF DOCX/XLSX/PPTX Macro .docm/.xlsm/.pptm Legacy .doc/.xls/.ppt Email .eml/.msg Images + EXIF HEIC/HEIF CSV / TSV Barcodes / QR / MRZ Archives zip/7z/rar
10

Forensic authenticity, threat & final action

Authenticity evidence, file danger, and account handling policy remain separate, so a phishing email, a forged invoice, and a routing rule are never conflated into one score.

ACCEPT REVIEW REJECT

Authenticity axis

Structural, metadata, content, and field evidence — is this document what it claims to be, and has it been tampered with?

Threat axis

Malware, phishing, and active-content risk — could opening or acting on this document harm you?

Account policy

Company, document-type, and custom handling rules may require review or rejection without changing the authenticity score or claiming tampering.

Corroboration gate

The forensic verdict escalates only on a curated decisive signal, corroboration across two or more distinct categories, or a calibrated score — never a single noisy medium-severity finding.

Per-category caps

Each signal category's contribution is capped before summing, so one noisy domain (say, ten low-severity font notes) can't out-score a genuinely cross-cutting problem.

11

Check catalog

Every signal category the engines can raise, with the number of distinct checks in the source that can raise it — computed live from the running code (GET /api/checks), not a hand-maintained figure that can drift.

Loading…

TOOLS

Standalone tools

Beyond the full scan, Docurensic ships focused single-file tools — small forensic labs and a PDF workbench you point at one document, link, or company. They all open from the Tools page (sidebar → Analysis), and most can pull a file straight from your File Vault so you don't have to re-upload. Jump to one:

Every tool here is advisory. They're for investigation and everyday document handling — none of them ever changes a forensic scan's authenticity score or verdict, and unless you explicitly save something to the vault, nothing you run through them is stored.

URL analyzer

One link, two lenses: is it dangerous, and how well-built is the site behind it?

Paste a URL or bare domain and Docurensic builds one shared picture of the site from real, live probes, then reads it two ways: a corroboration-gated safety verdict and a six-dimension A–F quality scorecard. It never convicts on one tell — a site is only flagged when the evidence agrees across at least two independent tiers and three categories. Reach for it before trusting a link in an email, invoice, or message — a supplier portal, a payment page, a "verify your account" link, or a vendor domain you've never dealt with.

  • Lexical link checks — insecure scheme, raw-IP hosts, "@" authority tricks, typosquat and brand-lookalike domains, punycode/homograph names, suspicious keywords, expanded URL shorteners, and open-redirect parameters.
  • Registration & infrastructure — domain age (a newly-registered domain is a strong fraud tell), registrar and locks, DNS and mail records (SPF/DMARC/CAA), and a live TLS certificate handshake.
  • Landing-page content — login forms on brand-new domains, credential forms posting elsewhere, hidden inputs, obfuscated scripts, invisible iframes, and brand-claiming titles on unrelated domains.
  • Quality scorecard — Setup/Security, SSL/TLS, Optimization, SEO, Tech stack, and Age/Trust, each graded, with a ranked "fix first" list and a tech-stack fingerprint (with end-of-life flags).

Good to know: an external reputation check runs only when configured; a clean read is not a guarantee a site is safe. Open the URL analyzer →

Company analyzer

Is this business real, who runs it, what's its reputation, and what to watch for?

Give it a company name, a domain, or a registration number and it returns one reconciled business report. Narrow data sources each verify a single fact, open-web discovery fills in the identity and story, and an AI analyst writes it up in plain English. The trust score and verdict are computed by fixed, auditable rules — the AI only writes the explanation around the numbers, never the numbers themselves. Use it during vendor or counterparty onboarding — before signing a new supplier, extending credit, or paying a first invoice.

  • Profile & people — legal/display name, industry, founded year, headquarters, phone, website, headcount estimate, and leadership names where discoverable.
  • Reputation — review snippets and a rating, sentiment themes, and dated, sourced news summaries.
  • Independent verifications — a real, reachable address, whether the phone is a live line (and its type), and whether the website domain is flagged as unsafe (a hard push toward high risk).
  • One clear read — a 0–100 trust score, letter grade, and verdict with confidence, a corroboration matrix (identity / address / phone / safety / reviews / leadership), and a "before you engage" checklist.

Good to know: results depend on which data sources are enabled; treat it as a fast starting point, not formal due diligence. Open the Company analyzer →

Metadata Lab

Everything a file says about itself — and every place its own story contradicts itself.

Deep-inspect one file's metadata without running a full scan. It lays out every property the file carries grouped by source, builds a chronological timeline of the file's own claimed dates, and surfaces a contradictions list — the places where the file disagrees with itself. Reach for it to interrogate one document's provenance fast: a scanned invoice whose dates feel off, a contract you suspect was edited after signing, a photo whose camera/GPS story you want, or an email whose headers you want laid out.

  • Grouped properties — PDF Info + XMP (raw and parsed), edit history, fonts and encryption; image EXIF/GPS/camera; Office core/app properties and macro presence; email headers and relay chain.
  • Date timeline — every claimed date in order, each tagged with where it came from.
  • Contradictions — dates that run backwards or sit in the future, Info-vs-XMP disagreements, content appended after a signature, and embedded macros, files, or JavaScript.
  • Revision history — a PDF's incremental-update map with which revision each signature covers.

Good to know: a contradiction is a lead to investigate, not proof of forgery — and a clean metadata story doesn't prove authenticity. Open any vault file with its "Metadata" action. Open the Metadata Lab →

Image forensics

Five independent visual lenses over one image to reveal edits, clones, and mismatched compression.

Run one image — or one rasterized PDF page — through five manipulation-analysis lenses, each shown as a heatmap or overlay. It deliberately gives you lenses and evidence to read, not a verdict or a score. Use it on a suspicious photo or image-based document — a receipt, an ID photo, a screenshot, a photographed invoice — when you want to see whether a number, stamp, signature, or region was edited, cloned, or pasted in.

  • Error-level analysis (ELA) at three qualities — heatmaps that highlight regions edited or pasted at a different compression history.
  • JPEG-ghost sweep — a per-block map that exposes a region carrying its own separate compression history.
  • Copy-move / clone detection — overlays lines and source/destination boxes where part of the image was duplicated to cover or repeat content.
  • Noise consistency — flags smoothed or inserted patches whose noise doesn't match the surrounding image; plus the JPEG quantization tables and an estimated original save quality.

Good to know: each lens has stated blind spots — heavily re-saved, resized, or screenshot images can wash the signals out. Read the lenses together as evidence, not a single answer. Open Image forensics →

Signature checks

Real cryptographic verification of a PDF's digital signatures — what they cover and what changed after.

Verifies the digital signatures inside a PDF without needing a trust store. For each signature it checks the underlying cryptography, reads the certificate the file carries about the signer, confirms the covered bytes are untouched since signing, and reports exactly what the signature covers and what was appended afterward. A second mode geometrically compares two signature images. Reach for it whenever a signed PDF matters — a contract, agreement, or certificate — to confirm the signature is intact and the document wasn't altered after signing.

  • Coverage & sealing — how much of the document each signature covers, which revision it seals, and whether bytes were added after signing.
  • Cryptographic integrity — the document digest still matches the signed byte range, and the signer's key actually produced the signature (RSA/PSS/ECDSA).
  • Signer certificate — subject/issuer, self-signed detection, validity window, and whether a trusted timestamp is present.
  • Signature compare — a geometric similarity score for two signature images (ink overlap, shape, stroke direction).

Good to know: it deliberately does not judge whether the signer's certificate chains to a trusted root or check revocation, and it says so. Compare is a similarity aid, not handwriting identification. Open Signature checks →

Compare

Overlay two versions of a document and see every change — down to the font.

Puts two documents side by side and surfaces every difference — text edits, moved content, and even font or size substitutions on otherwise identical text, so a changed amount typed in a slightly different font stands out. Use it when you have an original and a returned or countersigned copy of the same document — a contract, purchase order, invoice, or certificate — and need to confirm nothing was quietly altered.

  • Word-level diff with the changed regions located on the page.
  • Font forensics — flags where the same text appears in a different font or size, with a full font inventory of each side and what's new or missing.
  • Four visual modes — an X-ray channel overlay with blink, a pixel-difference heatmap, clustered difference boxes, and side-by-side highlighted changes.
  • Images welcome — either side can be an upload or a vault file, and an image is turned into a one-page document so you can compare a photo against a PDF.

Good to know: it shows you what differs — deciding whether a change is legitimate is up to you. Vault rows deep-link in with their "Compare" action. Open Compare →

PDF Studio

A full in-browser PDF workbench: edit, organize, sign, convert, redact, and protect.

Everyday PDF utilities built right into the product, so you don't need a separate app — all bytes-in, bytes-out, with nothing stored unless you save it. This is a productivity toolset, not a forensic analyzer: nothing here feeds scan scores or verdicts. Reach for it for the routine document handling around a case — combine exhibits, split out pages, redact personal data before sharing, OCR a scan, fill a form, or password-protect a file.

  • Pages & conversions — merge, insert, split, rotate, reorganize, crop, compress; convert to/from text, HTML, and images; and OCR a scan to make it searchable.
  • Markup & forms — shapes, arrows, highlight/underline/strikeout, investigation stamps, and list/fill/author form fields.
  • Permanent redaction by area, search term, or pattern (SSN/SIN, card numbers, email, phone) — the content under the box is genuinely removed, not just covered.
  • Sign & protect — sign, stamp, add text, watermark, edit metadata, flatten, and password-protect or unlock. Round-trips with the vault: open a vault PDF, save the edit back as a new file.

Good to know: search/pattern redaction finds text it can read — personal data baked into a scanned image needs an area redaction (run "Make searchable" first to check). Open PDF Studio →

X-Ray Engine

An evidence-first workbench that reconstructs a PDF's history and lays out competing explanations.

A standalone, PDF-only investigation workbench that goes deeper than a normal scan: it reconstructs a document's editing history, pinpoints exactly where and what changed, and lays out competing explanations with the evidence for and against each. It's explicitly a reasoning-and-evidence tool, not a pass/fail scorer. Reach for it on a high-stakes PDF where "what exactly changed, and where" has to be defensible. (A distilled Deep X-Ray version also runs automatically after every ordinary PDF scan and appears in the report.)

  • Recovered values — reconstructs recoverable prior revisions and shows before / after / difference crops of exactly what a field used to say.
  • Located findings — anchors each finding to an exact region on the page, with review crops of the key values to check against your source of truth.
  • Layers & construction — reconstructs document layers and hidden or covered content, and classifies how the document was built (native-digital, searchable scan, image-only capture, mixed).
  • Honest coverage — an amount-reconciliation and attention heatmap to focus review, with a complete / limited / skipped / failed note on every lens; missing coverage never means "clean".

Good to know: construction mode is not authenticity, tool markers don't identify the human editor, and it carries no risk score or verdict — it presents evidence, not intent. Open the X-Ray Engine →

WORKSPACE

Files, cases & your workspace

Where your documents live, how you group them, and where decisions and activity show up.

File Vault

Your private, encrypted home for documents. Every file is encrypted the moment you upload it and stays encrypted at rest under a key unique to your account, so only you can open your files. Organize them in folders; your storage allowance is set by your plan. Every scan automatically tucks the original into a "Scanned documents" folder (duplicates aren't stored twice) and badges it with its latest verdict, and any file offers "Send to analysis" to re-run the full scan. Open File Vault →

Cloud connections

Bring your own storage without moving files. Connect one Dropbox or Google Drive folder as a read-only mirror: files stay in your cloud, are indexed rather than copied, and use none of your vault quota. The mirror refreshes on a schedule and on "Sync now", and per-connection auto-scan runs any new file that arrives after you connect (never your initial import). Disconnecting just drops the index and access — your files stay put. Open Connections →

Cases

Group related documents — an application packet, a shipment's paperwork, a claim — and investigate them as one set. Docurensic suggests links when two documents share a strong identifier (email, IBAN, VIN, phone) or template, and you can link by hand; a connection graph draws the relationships, and the whole set can be reviewed and dispositioned together. Open Cases →

Review inbox

The human decision queue. Any document whose final action came back Review or Reject waits here for a person's call, so nothing that needs a second look slips through. Record a disposition — Approve, Reject, or Escalate — with an optional note; it's tracked separately from the automated verdict and feeds your review stats and workflows. Open the Review inbox →

Analytics

A bird's-eye view of your activity: how many documents you've scanned over time, the mix of trusted / review / reject with the trend against the prior period, your overall flag rate, and the forensic flags that fire most often across your recent reports. Adjust the time window to focus. Open Analytics →

Notifications

Your in-app alert feed (distinct from the Review inbox). It collects alerts raised by your Smart Workflows — a flagged document, a routed file — plus system messages from the Docurensic team, newest first. Mark one or all read, delete an alert, or clear everything you've already read. Open Notifications →

AUTOMATION

Smart Workflows — decide once, run on every document

Docurensic doesn't stop at the verdict. In the visual, no-code Workflows builder you write rules as WHEN something happens, IF it matches your conditions, THEN run these steps — and it runs automatically on every document after that. Jump to a section:

01

The WHEN / IF / THEN model

Every workflow reads the same way. WHEN picks the moment (a document finishing analysis, a human decision, a timer, or a manual run). IF is a set of optional filters, all AND-ed together — leave them empty to act on everything. THEN is a tree of steps that runs in order: each step is an action or a conditional split.

WHEN · a trigger fires IF · your conditions all hold THEN · the steps run in order
Workflows are automation glue, never forensic judgment. Actions can notify, tag, file, quarantine, note, and route — but they can never change a document's authenticity risk score or verdict. "Change status" writes only the operational review disposition, the same field a human reviewer sets. Everything is failure-isolated: if one step fails it's recorded and the rest still run, and workflow processing can never slow or sink the underlying analysis.
02

Triggers — the WHEN

The event that starts a workflow. Most derive from a finished scan; two run on their own.

TriggerFires when
Analysis completedA scan finished, whatever the outcome — the catch-all starting point.
Document uploadedA file entered the pipeline: uploaded here, sent via API, added to the Vault, or arriving from a connected Dropbox/Drive folder.
Document classifiedThe document type was identified (invoice, ID, certificate…). Pair with the "document type" condition to route by kind.
High fraud scoreThe forensic risk score met your threshold (defaults to 60; set your own with the minimum-risk condition).
Needs reviewThe final decision requires a human — including separate threat or advisory reasons.
Final action: rejectThe final decision is reject, whether from forensic evidence or your account policy.
Final action: acceptThe final decision is accept. Use it to auto-file documents that need no review.
Threat detectedAn active-content / payload threat level was raised (structural detection, not antivirus).
Signature detectedA digital signature, signature field, or visible signature mark was found.
Duplicate foundThe same or a near-identical document was analyzed before (possible double submission).
Document expiredAn expiry finding fired — lapsed insurance, an out-of-date certificate.
Disposable email foundA contact email on the document uses a throwaway / temporary-inbox domain.
Low analysis confidenceThe engines were unsure of their assessment or of the document's type.
OCR completedText was read off an image — a scan or photo that needed OCR.
QR / barcode readA QR code or barcode was decoded (shipping labels, tickets, certificates).
Manual review completedA reviewer set a disposition (approved / rejected / escalated). Fires from the Review inbox, not a scan.
Engine rule matchedOne of your Engine Detection custom rules fired. Narrow it with the "Engine rule names" condition.
On a scheduleRuns on a timer you set (hourly / daily / weekly / monthly in your time zone), not tied to any one document.
Run manuallyRuns only when you press Run — on a document you pick from history, or against the sample in test mode.
03

Conditions — the IF

Optional filters, all AND-ed together. Leave them empty and the workflow acts on everything the trigger catches; add a few to narrow it to exactly the documents you care about.

ConditionRuns only when
Risk score is at least / at mostThe forensic risk score is ≥ / ≤ your number (0–100). Combine both for a band.
Verdict is one ofThe forensic authenticity verdict is trusted, review, or reject (your pick). This can differ from the final action when policy forces a review.
Threat level is at leastThe active-content threat level is at least elevated, high, or critical.
Has an active-content threatThe document carries a payload / active-content threat.
Is a duplicateThis document matches one analyzed before.
Has an expiry / lapsed findingAn expiry or lapsed-coverage finding fired.
Has a signatureThe document is signed.
Amount is at least / at most ($)The document's parsed total is ≥ / ≤ your amount (e.g. invoices over an approval limit).
Document type is one ofThe document is one of the listed types (invoice, insurance…). Any match qualifies.
Filename containsThe file name includes this text (e.g. "contract").
Filename matches patternThe name matches a glob pattern where * is any run of characters and ? is one (e.g. INV-*.pdf).
Has any tag ofThe report carries at least one of the listed tags.
Came fromThe document's origin is an app upload, API upload, scheduled run, or manual run.
Engine rule namesOne of the named Engine Detection custom rules matched (pairs with the "Engine rule matched" trigger).
04

Actions — the THEN

What a workflow does, grouped into five families. Message actions can reuse a saved Connection instead of a pasted URL, and support formulas like {{filename}}.

Alerts & messages

Reach your team where they already work, or POST to any URL. In-app notification, Email, Slack, Microsoft Teams, Discord, Telegram, Google Chat, and Webhook (the full event JSON, signed when a secret is set — a bridge to Zapier, Make, n8n, Power Automate, or your own systems).

Tag, file & organize

Add tags to the report for later filtering, Move to a vault folder, Quarantine (move to a Quarantine folder and tag it), Create case (open an investigation and link the document in), and Delete document (permanently removes your own vault copy — the report keeps its findings).

Review & notes

Change status sets the human review disposition — approved / rejected / escalated. This is the operational disposition only; it never changes the forensic score or verdict. Add note appends a note (with formulas) to the report's reviewer notes.

Investigate with tools

Run metadata check and Run signature check run the Metadata Lab or Signature checks as part of the flow and tag/note the report on contradictions or a broken signature. Open in a tool posts an inbox alert with a one-click link that opens the document, pre-loaded from the vault, in Metadata Lab, Image forensics, Signature checks, Compare, or PDF Studio.

Flow control

Wait, then continue pauses and resumes the rest of the steps after 1 minute to 7 days. Stop the flow ends the workflow early (even inside a split). Start another workflow runs a second workflow's steps on the same document (chaining one level deep).

05

Branching — one flow, many paths

The THEN side is a step tree, not just a flat list. Any step can be a conditional split: several labelled paths, each with its own condition, plus an optional "Otherwise". Only the first path whose condition matches runs — exactly like a switch — then the flow continues past the split. Splits can nest up to three levels deep.

WHEN   Analysis completed
THEN   split on:
         ┌ High risk (score ≥ 75)   → create case + alert #fraud
         ├ Needs review             → tag "to-review" + email the queue
         └ Otherwise                → move to "Cleared" folder

A plain list of actions with no split validates and runs the same way, so branching is entirely optional. Limits keep flows readable: up to 10 top-level steps, up to 6 paths per split, and up to 40 actions across the whole tree.

06

Schedules & waiting

On a schedule

Run a workflow on a timer instead of on a document. Choose hourly, daily, weekly, or monthly and a time, all in an IANA time zone you pick (DST-correct). A background scheduler fires due workflows and re-arms the next run. Typical uses: a daily activity summary to your inbox, or a weekly review-queue reminder to Slack.

Activity digest

Scheduled runs carry a digest you can drop into messages: {{stats.scans}}, {{stats.high_risk}}, {{stats.rejected}}, {{stats.review_pending}}, and {{period}} (the window covered, e.g. "the last 24 hours").

Wait, then continue

The "Wait" step pauses a flow and resumes the remaining steps after anywhere from 1 minute to 7 days — so you can "alert now, then follow up tomorrow" or "quarantine now, wait 7 days, then delete". In a Test run it continues immediately instead of actually waiting.

07

Test it before you trust it

Three safe ways to try a workflow without waiting for a real document to trip it.

Test in the builder

The builder's Test button runs your unsaved draft against a built-in sample and shows each step's outcome — without saving the workflow or logging a run. Try splits and formulas freely before committing.

Test a saved workflow

Fire one saved workflow against the sample document. Its actions really execute, but outbound messages are marked as tests and document-changing actions skip (the sample has no real file). Scoped to just that workflow.

Run now

Run a saved workflow on demand against a real report you pick from your history — real execution, real side effects — or against the sample in test mode. Conditions still gate the run unless you choose to force it.

08

Start from a template

You rarely start from a blank canvas.

Suggested rules gallery

One-click, ready-made workflows grouped into Triage & routing, Scheduled digests, Investigate with tools, Essential alerts, and Integrations. Picking one opens it prefilled so you can adjust thresholds and recipients before saving — triage-by-risk (a three-way split), route-invoices-by-amount, daily/weekly digests, high-risk follow-up with a delay, quarantine-then-purge, and push-to-Slack/Teams/Discord/Telegram/Google Chat/webhook.

Default starters

A small safety net is seeded automatically the first time you open Smart Workflows — notify on completion, tag documents needing review, auto-open a case on very high risk, alert on rejected / duplicate / expired / threat, and label documents by risk tier. They're ordinary rules you can edit, disable, or delete, and a deleted default is never re-created.

09

Formulas — put document data in your messages

Every message field — notifications, emails, chat and webhook messages, notes, case titles — supports formula tokens of the form {{field|filter:arg}}. An insert-data picker in the builder means you don't have to type token names.

Worked example — a Slack message
🚩 {{filename}} — {{doc_type|title}} — scored {{risk_score}}/100
Verdict: {{verdict|upper}}   Amount: {{amount|currency:$|number}}
Review it: {{report_url}}

Fields include document facts ({{filename}}, {{doc_type}}, {{amount}}, {{tags}}), analysis facts ({{verdict}}, {{risk_score}}, {{threat_level}}, {{report_url}}), run facts ({{workflow}}, {{date}}), and the scheduled digest stats above. Filters chain after a pipe: upper, lower, title, round:1, currency:€, number (thousands separators), truncate:40, and default:none. A known field with no value shows an em dash, a typo'd token stays visible so you can spot it, and a broken formula never sinks the message.

10

Connections — set a destination once, reuse it everywhere

Save a URL or token once on the Connections page and every workflow can reference it by name — update the connection and every flow that uses it updates too. Each saved connection has a one-click test ping, and secret fields (like bot tokens) are write-only.

Messaging & alerts

Slack, Microsoft Teams, Discord, Telegram, Google Chat, and named email recipient lists. Mattermost and Rocket.Chat save as Slack-type (they speak Slack's webhook format).

File storage

Dropbox and Google Drive folders that mirror read-only into your vault and can auto-analyze new arrivals — the same cloud connections described under Files & workspace.

Automation bridges

Zapier, Make, n8n, and Power Automate as catch-hook webhooks — one connection reaches thousands of downstream apps (CRMs, ERPs, spreadsheets, ticketing).

Developer & scan-lifecycle webhooks

Your own custom webhook endpoints, plus platform webhooks that subscribe your systems to scan-lifecycle events (queued / processing / complete / failed), gated by risk level and HMAC-signed on every delivery.

Each gallery entry on the Connections page carries short "how to get your URL or token" steps, so wiring up a new destination is a copy-paste, not a hunt.
API

Integration API (API keys)

The supported, stable way to run Docurensic from your own systems. It is a small, analysis-only surface under /api/v1, authenticated with an API key — separate from the browser session endpoints documented further down this page.

Getting a key

Open Settings → API access, turn on Enable API access (the master switch — while it is off, every API request returns 403), then click Generate key. The key (drk_live_…) is shown once — store it securely; only a SHA-256 hash is kept server-side. Regenerate or revoke it in Settings at any time; the previous key stops working immediately. Send it on every request as Authorization: Bearer drk_live_… or X-API-Key: drk_live_….

curl -s https://docurensic.com/api/v1/analyze \
  -H "Authorization: Bearer $DOCURENSIC_API_KEY" \
  -F "[email protected]"

Same accepted types and 150 MB limit as the app. The call is synchronous — typical latency is tens of seconds for the full forensic + AI pass. The response is the fixed v1 document payload (verdict, score, the decision contract, document type, and the slim case-file answers). Full field-by-field reference: the API guide shipped with your deployment.

post/api/v1/analyzeapi key

Analyze one document. Multipart body: file. Returns the v1 document payload with a document_key you can re-fetch. Runs the identical pipeline as the app (report persisted to your history, auto-archived to File Vault unless storageless, webhooks fire).

get/api/v1/meapi key

Key check. Returns 200 with your user_id when the key is valid and API access is enabled.

get/api/v1/documents/{document_key}api key

Re-fetch a previous analysis (API- or app-created — it is one history). 404 if the key doesn't exist or belongs to another account.

get/api/v1/scans/{scan_id}/xrayapi key

The advisory X-Ray layer manifest attached to a completed PDF scan; overlay artifacts are served from /api/v1/scans/{scan_id}/xray/artifacts/{artifact_id}.

Rate limits. POST /api/v1/analyze shares your per-account scan budget (20 / minute); other v1 endpoints allow 60 requests / minute per account. Browser session tokens are not accepted on /api/v1/*, and API keys are not accepted on the application endpoints below — the two surfaces are deliberately separate, which is what makes usage attributable.
API

Application endpoints (browser session)

Every feature in the app is backed by a REST endpoint under /api. These are the app's own browser surface — authenticated with a Supabase session token, not an API key — documented here for reference. For programmatic integration from your own systems, use the Integration API above. Requests and responses are JSON unless a file upload is involved (then multipart/form-data).

Authentication

Docurensic uses Supabase Auth. Sign in to obtain an access token, then send it as a bearer token on every request. Tokens are verified server-side (with a 60-second in-process cache). A handful of reference endpoints are public — they're marked public below.

curl https://docurensic.com/api/scan \
  -H "Authorization: Bearer <your-access-token>" \
  -F "[email protected]"
Rate limits. Scans and standalone recognitions are each limited to 20 requests per 60 seconds per account, in separate buckets (so running the recognizer doesn't eat your scan quota). Uploads are capped at 150 MB. Exceeding a limit returns 429; an oversized file returns 413.

Scan & history

Run a full forensic scan and manage the reports it produces. A report is persisted to your workspace (unless your account is in storageless mode) and is only ever visible to you.

post/api/scanuser

Full forensic scan. Multipart body: file (the document). Returns the complete report — decision action, forensic verdict and score, threat level, identification, signals, fields, and, when document text is available, the AI review (classification + type-aware checks + entities). No-text documents return an explicit unavailable state.

get/api/scan/historyuser

List your reports. Query params: q (filename or SHA-256 prefix), verdict (trusted/review/reject), needs_review, limit (1–100, default 50), offset. Returns {rows, total, limit, offset}.

get/api/scan/{report_id}user

Fetch one full report by id.

delete/api/scan/{report_id}user

Delete a single report.

delete/api/scan/historyuser

Clear your entire scan history.

get/api/scan/statsuser

Aggregate counts for the workspace. See also /api/scan/analytics, /api/scan/analytics/overview, and /api/scan/analytics/flags for the Analytics page data.

get/api/tools/hash/{sha256}user

Find reports in your workspace by full SHA-256 or a 12+ character prefix.

Standalone analysis

Run a single engine directly, without a full scan — the endpoints behind the Tools page.

post/api/analyze/textuser

AI-generated-text detection on pasted text. JSON body: { "text": "…", "engine": "auto", "perplexity": false }. Returns the verdict, a score, and every stylometric signal individually. perplexity: true adds a language-model perplexity pass (slower).

post/api/analyze/text/fileuser

Same AI-text analysis on an uploaded file (multipart file) — text is extracted from PDF/DOCX/TXT/image (OCR when available) first.

post/api/analyze/metadatauser

Metadata & provenance only (no full scan). Multipart file; PDF and Office (docx/xlsx/pptx/doc/xls/ppt) only. Returns producer/creator, timestamps, origin classification, and any provenance flags — the exact functions a full scan uses, so the numbers can't drift.

post/api/analyze/emailuser

Email header & message analysis. JSON body: { "raw": "…" } — pasted raw headers or a full raw message. Returns key headers, parsed SPF/DKIM/DMARC results, the Received relay chain with hop delays, forensic signals (BEC reply-to redirect, display-name spoof, link-text/href mismatch), and every body URL with a passive reputation { risk, band } score.

post/api/analyze/email/fileuser

Same analysis for an uploaded .eml / .msg file (multipart file).

post/api/recognizeuser

Multi-source document recognition & forensics fusion (own rate-limit bucket). Multipart file; query params disable (comma-separated source list, e.g. disable=vlm), ocr_mode, ocr_engine. Returns document type with agreement count, born-digital/scanned, signature state, and the list of sources used.

Example — URL reputation response

{
  "domain": "paypal.com.account-verify.ru",
  "risk": 88,
  "band": "flag",
  "confidence": "high",
  "engine": "aidetect_single",
  "structural_flags": ["brand_token_outside_etld1", "suspicious_keyword"],
  "domain_age_days": 12,
  "evidence": [
    { "feature": "typosquat_distance", "score": 0.9, "note": "resembles paypal.com" }
  ]
}

Account

get/api/account/usageuser

Your scan counts, broken down by verdict and threat level, plus the timestamp of your last scan.

get/api/account/settingsuser

Your scan-configuration settings, merged over defaults.

put/api/account/settingsuser

Partial update. Keys: storage_mode, sensitivity, industry_preset, and a features map. Unknown keys are ignored; invalid values rejected. See the Account tab for the full list of accepted values.

get/api/account/exportuser

Everything held for your account — scan reports, settings, and vendor baselines — as one portable JSON download. Free on all plans.

delete/api/accountuser

Permanently delete your account and all associated data. Irreversible — the UI gates it behind a typed confirmation.

get/api/ai-usageuser

Your Docurensic AI usage log — tokens, model, latency, and estimated cost per call. /api/ai-usage/{row_id} returns a single entry.

Review & cases

get/api/review/queueuser

Documents whose final action is Review or Reject, with no disposition yet, awaiting human review. /api/review/queue/count returns just the count for the nav badge.

post/api/review/{report_id}/dispositionuser

Record an examiner decision. JSON body: { "disposition": "approved" | "rejected" | "escalated", "reviewer_note": "…" }.

post/api/cases · get/api/casesuser

Create a case ({ "title": "…" }) or list your cases with document counts. Fetch one with GET /api/cases/{case_id}.

post/api/cases/{case_id}/documentsuser

Link a document to a case: { "scan_id": "…", "link_reason": "manual" }. Remove with DELETE /api/cases/{case_id}/documents/{scan_id}.

Reference & health — public, no auth

get/api/checkspublic

The live check catalog — every signal category the engines emit, its human label, and how many distinct checks in the source can raise it. This is what powers the catalog table in the How-it-works tab.

get/api/reference/producerspublic

The producer-reliability tier table (which creation tools are trusted vs. suspect), ranked high → low. Also browsable at /producers.

get/api/auth/configpublic

The Supabase URL and anon key the browser client needs to sign in.

get/api/healthpublic

Liveness check.

Administrative endpoints require an administrator account; a regular user token is rejected. The template registry is global — any signed-in user can read it via GET /api/templates, and only administrators can write to it.
USE CASES

Docurensic in your business

Docurensic ships tuned profiles for the document types that carry the most fraud, and industry rule packs that layer sector-specific checks on top. You can pin your workspace to an industry in Settings → Industry preset (auto, freight, insurance, lending, legal, or hr), or leave it on auto-detect. The scenarios below are grounded in the actual doc-type profiles and cross-document intelligence the engines run.

Freight & logistics

Brokers, 3PLs, and carriers vetting shipment paperwork

Double-brokering and carrier-identity fraud run on altered rate confirmations and bills of lading. Docurensic's freight pack reads MC/USDOT numbers off the page and flags them for FMCSA verification, while cross-document intelligence catches the same carrier identity reused across unrelated loads.

rate_confirmationbill_of_lading customs_documentmc_authority
  • MC/USDOT extraction & FMCSA availability — the freight rule pack surfaces the carrier's authority number and flags it for verification.
  • Cross-document entity links — a shared MC/DOT number, email, or phone across otherwise-unrelated documents is a fraud-ring tell.
  • Rate / amount tampering — a rate-confirmation amount that was edited after issue leaves metadata and font-splice traces.
  • Near-duplicate detection — a lightly-edited copy of a real rate con, byte-different but structurally identical.

Lending & finance

Underwriters and finance teams verifying supporting documents

Loan and credit decisions rest on documents that are easy to doctor — invoices, purchase orders, pay stubs, bank statements. Docurensic checks the arithmetic and the account numbers, not just the look, and applies statistical fraud tests to any tabular data.

invoicepurchase_order payment_receiptbank_statement (CSV)
  • Checksum validation — IBAN, routing-number, and VIN check digits are verified, so a swapped payee account fails arithmetically.
  • Cross-field arithmetic — line items that don't sum to the stated total, or tax that doesn't reconcile.
  • Benford's law & balance reconciliation — spreadsheet/CSV statements get a chi-square Benford test, duplicate and round-number clustering, sequence-gap and future-date detection.
  • Cross-document field consistency — an invoice number or amount that disagrees between a PO and its invoice.

Insurance

Carriers and brokers checking coverage evidence

A certificate of insurance is trivial to edit in a PDF viewer — bump a coverage limit, extend an expiration date, swap an insured name. Docurensic's doctype profile is date-aware (a future expiration is expected, not treated as back-dating) and catches the edits themselves.

certificate_of_insurancepolicy_document
  • Doctype-aware timeline checks — expiration in the future is normal for a COI; a modified date after issue, or an impossible timestamp, is not.
  • Font-splice & revision detection — an edited coverage limit or policy number leaves a mismatched font subset or a recoverable prior revision.
  • Policy-number entity links — the same policy number appearing across unrelated certificates.
  • Template matching — page-1 perceptual hash against a registry of known issuer layouts.

HR & onboarding

Verifying identity and tax documents at hire

Onboarding packets are a common vector for synthetic identities — a doctored W-2, a fabricated W-9, a photoshopped ID. Docurensic checks the tax forms structurally and validates the machine-readable zone on ID scans.

w2_formw9 id_document (MRZ)
  • MRZ check-digit validation — TD3 passport / ID machine-readable zones are verified digit-by-digit; a forger who edits the printed number and forgets the MRZ is caught.
  • Tax-form field logic — W-2/W-9 required fields, EIN/SSN formatting, and cross-field consistency.
  • AI-image & synthetic-scan detection — an ID photo that's AI-generated, or a born-digital form re-rendered to an image to dodge raster checks.
  • Embedded-image EXIF — editor software and GPS traces left in a scanned document's images.

Legal & professional services

Contracts, signatures, and the email they arrive in

A signed PDF contract, and the message that delivered it, are both attack surfaces. Docurensic validates digital signatures cryptographically and reads email authentication headers to catch the business-email-compromise redirect shape.

contract (PDF)signed_agreement email .eml/.msg
  • Signature integrity — cryptographic validity, byte-range coverage, shadow-attack detection, and certification vs. approval signatures.
  • Hidden edits — an incremental update after signing, or a form field altered post-execution.
  • Email authentication — SPF/DKIM/DMARC parsing and From/Reply-To/Return-Path mismatch (the classic BEC redirect).
  • Link & brand spoofing — punycode domains, typosquats, and hidden auto-executing URLs in the document or message.
Group a whole packet with Cases. When several documents belong together — an application, a shipment's paperwork, a claim — add them to a Case so cross-document entity links and field-consistency signals between them are surfaced in one place, and the Review inbox can carry a single disposition for the set.
FINE-TUNE

Make detection fit your workflow

Every account scans with the same calibrated engine, but no two document flows look alike: a freight broker's rate confirmations, a lender's bank statements, and an HR team's diplomas carry different "normal". The Detection panel on the Settings page lets you tune what gets flagged, how strictly verdicts are drawn, and what your own business rules add on top — without ever weakening the core forensics. The controls below are read-only until you flip the Custom engine settings switch at the top of the Detection panel — the sensitivity dial applies either way, and while the switch is off the engine runs on its calibrated defaults.

Sensitivity dial

One dial that scales the review/reject thresholds. Turn it down when a noisy intake stream produces too many borderline reviews; turn it up when a missed forgery costs more than an extra manual look.

Settings → Detection → Sensitivity

Category switches

Advisory categories — AI-content analysis, web corroboration, and similar context checks — can be switched off per account. Signals in a disabled category are dropped before scoring, so a category that doesn't apply to your documents stops contributing noise entirely.

Advisory only

Verdict thresholds

Set your own review and reject score boundaries per account. Raising a threshold quiets borderline noise; it never overrides hard evidence (see the safety floor below).

Numbers you control

The safety floor

Core forensic checks — file structure, revision history, font forensics, signature coverage — always run and can never be disabled. And whenever a hard core-forensic signal stands, the engine's own verdict is a floor: your thresholds can quiet noise, but they can never silently auto-trust a caught tamper.

Always on
An untouched account changes nothing. Tuning is applied after the engine finishes, deterministically — if you never open the Detection panel, your scans behave exactly like the factory calibration, byte for byte.
CUSTOM RULES

Your business rules, as formulas

Custom rules let you express policy the engine can't know: which amounts matter to you, which vendors you trust, which document types deserve extra scrutiny. Rules are built in a visual formula builder — no code — from three ingredient types, and their outcome is always a separate policy action (tag the report, require review), never a change to the forensic evidence itself.

Document fields

Values extracted from the document — the total amount, dates, sender domain, page count. Example operand: document total.

Engine scores

The scan's own numbers — risk score, threat level, AI-content likelihood — usable as comparisons: risk score ≥ 40.

Detection flags

Did a family of checks fire? font anomaly fired, metadata contradiction fired — yes/no facts you can combine with AND / OR / NOT.

Worked example
WHEN   document type is Invoice
IF     document total ≥ 10,000
AND    (font anomaly fired  OR  revision history fired)
THEN   require review  +  tag "large-invoice-check"

This rule never claims the invoice is forged — it records that your policy wants human eyes on large invoices with any typography or edit-history finding. The tag it applies is visible on the report, filterable in your history, and usable as a Smart Workflows trigger ("my custom rule matched") to alert a channel or route the document to a tool automatically.

Start from a template. The rule builder ships a per-document-type starter gallery — high-value invoice checks, statement consistency guards, certificate scrutiny — so the usual rules are one click plus a number, not a blank canvas.
PROFILES

Company profile and per-type overrides

Two more layers tailor scans to your paperwork rather than paperwork in general.

Company profile checks

Tell Docurensic your own company details — name, tax id, bank account — and it checks documents that claim to be yours against them. A "your" invoice carrying an unknown bank account is exactly the payment-redirect fraud pattern; the mismatch requests review as a policy finding.

Per-document-type overrides

Different strictness for different types: run bank statements a notch more sensitive than general correspondence, or require review for every contract regardless of score. Overrides scope to the document type the scan identified.

Policy vs. evidence — kept apart

Everything on this page adjusts operational policy: what needs review in your workflow. The forensic risk score and verdict stay evidence-only, so a report always shows honestly which findings are engine evidence and which are your account's rules.

Tuning changes, done safely. Change one dial at a time, then re-scan a handful of documents you already know to be good and one you know to be bad — the score story on each report shows exactly how your settings shaped the outcome, so you can see the effect before trusting it in production.
ACCOUNT

Your account

Signing in, configuring how scans behave, and managing your data. Everything here lives on the Settings page.

Signing in

Docurensic uses email + password authentication (Supabase Auth). Create an account on the sign-up page; if you forget your password, use the reset link on the login page. Your session is what authorizes every scan and keeps your history private to you.

What's tied to your account

Every scan, report, case, and vendor-drift baseline is scoped to your user id and enforced by row-level security — only you can read or delete it. Your email, display name, and sign-in history come straight from your auth session.

01

Scan settings

Configure how every scan behaves. These are the exact keys the PUT /api/account/settings endpoint accepts.

SettingValuesWhat it does
storage_mode store · storageless store keeps the full report and automatically archives the analyzed original, encrypted, when File Vault storage is available and quota allows. storageless creates no new source copy; an existing vault or connected-storage file remains there. A slim history record retains filename/basic scan metadata, forensic verdict, decision action, findings, safe document type, and verification hash, without model, X-Ray-span, or provenance-value payloads.
sensitivity conservative · balanced · aggressive How readily the forensic authenticity verdict escalates. Balanced is the default; conservative reduces false positives, aggressive surfaces more borderline evidence for review. Account-policy actions stay separate.
industry_preset auto · freight · insurance · lending · legal · hr Pins the industry rule pack applied to your documents. auto detects it from the document type; only a tuned pack that matches the doc type ever applies.
features a map of on/off toggles Turn individual detectors on or off — ai_text, ai_image, recognition, url_reputation, cross_document. Off by default: the placeholder tools (tool_mrz and the extended lab features).
02

Privacy & security

Controlled retention

Every upload uses temporary processing that is deleted after analysis. Store mode automatically encrypts the original in File Vault when storage and quota are available. Storageless creates no new source copy; existing vault or connected-storage files remain there, and only a slim history record is added.

Authenticated access

Every scan, history entry, case, and report is tied to your account via Supabase Auth and scoped by row-level security so only you can see or delete it.

AI data handling

When bounded document text is available, the AI review sees only extracted text, metadata, and the engines' findings — the file itself is never sent to the model. Raw document text is not retained in the usage log by default, and never in storageless mode.

Rate limiting

20 scans (and 20 recognitions) per 60 seconds per account, in separate buckets, to keep the service usable for everyone.

Transport & headers

HSTS on HTTPS, X-Frame-Options: DENY, X-Content-Type-Options: nosniff, a strict Referrer-Policy, and a restrictive Permissions-Policy on every response.

Content-mismatch rejection

Uploads are sniffed by magic bytes; a file whose real content is an executable, or an archive disguised with a document extension, is rejected before any parser sees it.

03

Exporting & deleting your data

Export everything

From Settings (or GET /api/account/export) you can download everything held for your account — every scan report, your settings, and your vendor baselines — as one portable JSON file. Free on all plans, any time.

Delete your account

Account deletion is self-service and permanent. It removes your scan history, vendor-drift baselines, and settings, then deletes your login itself. There is no recovery window — the UI gates it behind a typed confirmation.

FAQ

Frequently asked questions

Can Docurensic detect if a PDF has been altered?

Yes. Because PDF stores edits by appending new content, we detect multiple revision markers (%%EOF) that indicate prior versions exist. Metadata timestamps that are physically impossible also flag tampering.

How long does a scan take?

Time varies by format, size, page count, OCR needs, and whether external AI is available. Structural checks are usually the fastest part; large, scanned, or complex documents take longer.

Is my document stored after the scan?

Every file uses temporary processing that is deleted after analysis. In storage mode, the analyzed original is automatically encrypted in File Vault when storage is available and quota allows. Storageless creates no new source copy; an existing vault or connected-storage file remains there. The slim history record keeps filename/basic scan metadata, forensic verdict, decision action, findings, safe document type, and verification hash.

Can Docurensic prove a document is genuine?

No — it detects forensic red flags, it doesn't certify authenticity. A low risk score means we found no evidence of manipulation, not a guarantee of genuineness.

Is my document sent to an AI model?

The file itself is never sent. When bounded extractable or OCR text exists, the AI review works from that text, metadata, and the forensic engines' findings only. With no extractable text, the external AI calls are skipped.

What file types are supported?

PDF (all versions), Word/Excel/PowerPoint (including legacy and macro-enabled), emails (.eml/.msg), images, CSV/TSV, and archives (.zip/.7z/.rar) — up to 150 MB.

Does it work on scanned paper documents?

Yes, with reduced signal coverage. Scanned-only PDFs (no extractable text) are flagged as image-only and OCR'd so content and logic checks still run — and "image-only" is itself useful for a document that claims to be born-digital.

Where do I report a problem or ask for help?

Open a case on the Help & support page — technical issues, questions, feedback, and feature requests all land in the same queue, and replies reach you in-app and by email. The Tools page lets you run any single engine on a one-off file or URL while you investigate.

Didn't find what you were looking for?

Open a support case and a human will get back to you — usually within one business day. The section you were reading rides along so you never have to explain where you got stuck.

Open a support case