Everything you need to run Docurensic
From your first scan to the forensic engines behind it, the standalone tools, the automation that runs after every verdict, the full API, real-world use cases, and how your account and data are handled. Pick a tab — every page here mirrors the running code, nothing is aspirational.
What Docurensic is
Docurensic is a document forensics service. Upload a PDF, Office file, email, image, or archive and it inspects the file the way a document examiner would — metadata, fonts, revision history, embedded images, signatures, and the numbers on the page — then returns a plain-English report with a forensic evidence score and a clear operational action. Everything works from your browser — no install, no plugin; documents are analyzed server-side in isolated processing.
What it's good at
Catching altered invoices and bank statements, forged certificates and IDs, phishing / business-email-compromise messages, malware hidden in Office macros, and AI-generated text or images — the everyday fraud that arrives as a file.
What it does not claim
Docurensic surfaces forensic red flags; it does not certify a document as genuine. A low score means no evidence of manipulation was found, not a guarantee. Findings and actions support a human decision; they are not proof of authenticity.
Who it's for
Freight brokers vetting rate confirmations, lenders reviewing pay stubs and invoices, insurers checking certificates, HR teams onboarding with W-2/W-9 forms, and anyone who receives documents they can't take on trust.
Run your first scan
Three steps. Simple files often finish in seconds; large or complex files can take longer. Everything happens from the Documents page — the first item in the sidebar.
Upload a document
Drag a file onto the upload box, or click Browse files. Docurensic accepts PDF, Word / Excel / PowerPoint (including legacy and macro-enabled), .eml / .msg email, images, CSV/TSV, and .zip / .7z / .rar archives, up to 150 MB.
The file uses isolated temporary processing that is deleted after analysis. In storage mode, the analyzed original is automatically encrypted in File Vault when vault storage is available and quota allows. Storageless creates no new source copy; an existing vault or connected-storage file remains there. Its slim history record retains filename/basic scan metadata, forensic verdict, decision action, findings, safe document type, and verification hash.
Run one consistent analysis
There is no analysis-depth selector. The applicable forensic engines always run. When the document has bounded extractable or OCR text, the independent AI review runs in the same flow and stays advisory.
Read the report
You get an overall action (Accept, Review, or Reject), a separate forensic authenticity result and 0–100 evidence score, a separate threat level, and every finding ranked by severity with a plain-English explanation. Click Detailed report for the full breakdown, or find it again any time in Scan history.
AI analysis when document text is available
Every scan runs the applicable forensic engines. When bounded native or OCR text is available, it also runs an independent AI-powered review — there is no depth selector or AI tier to choose. A no-text document reports that the review was unavailable instead of calling a model with empty input. AI never changes the forensic score or verdict.
Document type
An AI model classifies the extracted text into one of 33 canonical document types (invoice, bank statement, COI, deed, …) with a schema-enforced answer — or Undefined when nothing fits.
Type-aware checks + entities
The model runs only the integrity checks that apply to the classified type — financial arithmetic, dates and validity, party/issuer consistency, overall coherence — and returns one plain-language reason per check, its own red flags and fraud read, plus structured entities (emails, phones, IDs, accounts, addresses, domains, names) for case linking.
Text only, never the file
Both calls see extracted text, metadata, and the engines' findings — the document file itself is never uploaded to the model.
Reading your report
The report leads with one operational action while keeping forensic authenticity, file threat, account policy, and advisory AI visibly separate.
Overall action & forensic evidence
The headline action is Accept, Review, or Reject. Directly below it, the forensic authenticity verdict and 0–100 evidence score explain the tamper assessment. Threat or account policy may make the action stricter without rewriting that evidence.
Threat level
A separate malware / phishing / active-content axis. A clean, authentic-looking document can still carry a high threat level if, say, it hides an auto-executing link or a macro payload.
Document identification
A single block reconciling every engine's type vote (keyword classifier, recognizer, field & type engine) into one headline document type, with per-engine confidence shown side by side.
Ranked concerns & evidence
Material findings are grouped into concerns with severity, location, plain-English evidence, benign context, and a next step. Lower-level signals remain available as drill-down, never as a bare score.
Field analysis
Typed business fields the engine extracted — invoice/PO numbers, amounts, dates, IBANs, MC/DOT numbers — each marked validated or unvalidated (e.g. a failing IBAN checksum).
AI review
When document text is available, Docurensic AI's independent read — document type, per-check reasons, red flags, and its own advisory score — appears in its own section with a token/cost footer, kept visually and computationally separate from the forensic verdict. A no-text report says the review was unavailable.
Finding your way around
The sidebar groups everything into Workspace, Analysis, and Organization.
Documents
Upload new files and search your full scan history. Your home base.
Cases
Group related documents (an application packet, a shipment's paperwork) so cross-document links between them are easy to follow.
Analytics
Volume, verdict mix, and the flags firing most often across your scans.
Review inbox
Documents whose final action is Review or Reject and still need a human disposition — approve, reject, or escalate.
File Vault
Your encrypted document storage — retained scan originals and files you upload directly, organized in folders, only you can open them.
Tools
The hub for every standalone tool: URL analyzer, Company analyzer, Metadata lab, Image forensics, Signature checks, Compare, PDF Studio, and the X-Ray Engine workbench.
Workflows
Automate what happens after a scan — "when a document scores high, alert the fraud channel and quarantine it". Built visually, no code.
Connections
Everything Docurensic talks to: watched Dropbox/Drive folders, Slack, Teams, Discord, Telegram, Google Chat, email lists, automation platforms, and your own webhooks.
Notifications
In-app alerts from your workflows, cases, and the system.
Settings
Storage mode, sensitivity, industry preset, detection controls, feature toggles, API access, data export, and account deletion. Covered in Account & privacy.
Automate what happens next
Docurensic doesn't stop at the verdict — you decide what happens next once, and it runs on every document after that, no code. In the visual Workflows builder you write rules as when a document event happens, if it matches your conditions, then run steps — alert a channel, tag, quarantine, open a case, run a tool, or hand off to another workflow. Save your Slack channels, email lists, cloud folders and webhooks once in Connections and reuse them by name, and let new files from a watched Dropbox or Drive folder analyze themselves.
The full trigger, condition and action catalog — with branching, schedules, formulas and connectors — is in the Automation tab.
The full technical reference: what each engine checks, how forensic authenticity, file threat, and account policy produce one operational action, and the live catalog of every signal the engines can raise. Jump to a section:
What Docurensic checks
Multiple forensic categories, dozens of individual signals. Most tampering leaves traces across more than one.
Structure & metadata
Counts revision markers, checks DocInfo for contradictions. The trailer /ID pair reveals re-saves a document claims never happened.
Font fingerprinting
Every embedded font carries a unique subset tag. Different prefixes for the same base font on one page signal spliced content from an editor.
Revision history
PDF stores edits by appending new content. Multiple %%EOF markers mean prior revisions exist that can be recovered — a clear sign of editing.
Content analysis
Extracts all text and checks for extractability — image-only scanned documents get flagged as needing separate OCR verification.
Editor fingerprinting
Known PDF editors — Acrobat, Photoshop, GIMP, iText, PDF-XChange — leave traces in producer and creator fields that signal post-creation modification.
Risk scoring
Forensic signals are weighted by severity, capped per category, then gated into an authenticity verdict. The final operational action separately accounts for file threat and explicit account policy.
Six engines, one report
Every upload is routed by file type to the engines that understand it. All write to the same normalized report shape. Findings remain isolated as authenticity, threat, AI-content, identity, or account-policy information before one decision contract produces the operational action. Each engine is import-guarded and failure-isolated — a missing dependency or runtime error degrades that one engine, never the whole scan.
pdf_forensics
The primary PDF engine — structural integrity, metadata, fonts, raster forensics, signatures, forms, embedded content, and semantic field checks.
axis · forgery + threatmultiformat
Office (OOXML + legacy + macro-enabled), email, archives, spreadsheet/tabular fraud, and barcode/MRZ cross-validation.
axis · forgeryai_text / ai_image
Stylometric AI-written-text detection and synthetic-image detection — the content-trust axis, run on the relevant upload types and exposed directly via the API.
axis · ai_contentrecognition
Multi-source ML document recognition and type fusion (magika, libmagic, layout, text) — the identity axis. Light sources run inline in every scan.
axis · identityurl_reputation
Passive URL / phishing reputation and typosquat analysis on any links a document or email contains.
axis · threatintel
Cross-document intelligence layered on top of every scan — submission velocity, near-duplicate detection, entity-link fraud-ring detection, and the template registry.
Supabase-backedPDF forensics engine
Runs on every .pdf upload. Every check is isolated — a single
failing check is recorded as an error and never crashes the request.
Metadata & timeline
Producer/Creator/Author/Subject/Keywords, custom Info-dictionary keys, XMP-vs-DocInfo consistency, CreationDate/ModDate logic (doctype-aware — a certificate's future expiration date isn't treated as back-dating).
Structural integrity
Incremental-update detection, object/stream/object-stream counts, xref-repair detection (the parser silently reconstructing a damaged table), trailer /ID consistency, linearization.
Font & text analysis
Font inventory, embedded-vs-subset, digit-glyph remapping (a font that visually draws one digit but declares another in its ToUnicode map), Unicode anomalies — directional-override controls, zero-width characters, non-Latin digit homoglyphs.
Image & raster forensics
Error-Level Analysis, copy-move/clone detection, JPEG-ghost/double-compression, quantization-table analysis, DPI consistency, embedded-image EXIF/GPS/editor-software extraction.
Revision history
Prior-revision recovery and diff, synthetic-scan detection (a born-digital PDF re-rendered to an image to dodge raster checks).
Digital signatures
Presence, cryptographic validity, byte-range coverage, shadow-attack detection, certification vs. approval signatures.
Forms
Full AcroForm field-tree walk — a field hidden from the viewer that still carries a value, or template placeholder text left unfilled.
Embedded content
Attachment bytes are extracted (not just listed), magic-byte-vs-extension disguise detection, and recursive re-scan through the multi-format engine for anything it understands.
URLs & links
Every URL classified as visible (a Link annotation or printed text) or hidden (reachable only via OpenAction/embedded action — never shown to a reader), plus IP-literal, punycode, and shortener detection.
Semantic & field logic
Cross-field arithmetic, IBAN/routing-number/VIN checksum validation, entity extraction (emails, phones, MC/DOT/policy numbers), template pHash matching against a registry.
OCR fallback
Image-only pages get OCR'd so the content/logic passes still run on scanned documents, not just born-digital ones.
Watermark & template
Expected-watermark presence by document type, and first-page perceptual-hash matching against registered issuer templates.
Multi-format engine
Handles everything that isn't a PDF: Office documents (including macro-enabled variants), email, archives, spreadsheets, and standalone images.
Office (OOXML + legacy)
.docx/.xlsx/.pptx/.docm/.xlsm/.pptm and legacy .doc/.xls/.ppt — core/app properties, track-changes recovery, hidden content, VBA macro detection, external data links, formula-vs-cached-value reconciliation.
Email (.eml/.msg)
SPF/DKIM/DMARC header parsing, From/Reply-To/Return-Path mismatch (the BEC redirect shape — high only when Reply-To goes to free webmail), display-name brand spoofing, punycode domain detection, attachment recursion.
Archives
zip/7z/rar entry listing with a compression-ratio guard against zip bombs.
Spreadsheet & CSV fraud
Benford's law (chi-square test on a real sample), balance reconciliation, duplicate/round-number clustering, sequence-gap detection, future-dated transactions.
Barcode / QR / MRZ
1D/2D decode, TD3 passport MRZ check-digit validation, and QR-payload-vs-printed-text reconciliation — a forger who edits the printed number and forgets the code is caught here.
Standalone images
Routes through the AI-image detector, barcode/MRZ decode, and OCR-based QR reconciliation.
AI / content-trust engine
A separate detector for synthetic content and unsafe links, run automatically on image uploads and available directly for text/URL analysis.
AI-image detection
ELA, noise-residual, and frequency-domain artifacts, plus known generator metadata signatures.
AI-text indicators
Stylometric scoring across seven weighted signals — sentence-length burstiness, AI-phrase marker density, transition-word openers, opener diversity, n-gram repetition, punctuation style, and contraction rate — plus an optional language-model perplexity + burstiness pass. Every signal is reported individually so a verdict is explainable, not a black box. Deliberately conservative — informational, never a standalone accusation, because of well-documented false-positive rates on formal and non-native-English writing.
URL reputation
Domain age heuristics, typosquat/brand-lookalike detection, and a hook for an external Safe Browsing lookup — applied to every URL a scan or the email analyzer surfaces.
Recognition & fusion engine
A second opinion on document type and authenticity, fused from up to eleven independent sources — each optional, each skipped gracefully if its dependency isn't installed. A type verdict is corroboration-gated: two sources agreeing produces high confidence; a single source is capped and flagged uncorroborated_type.
Format identification
ML-based format ID (magika) cross-checked against libmagic — disagreement between the two, or against the declared file extension, is a masquerading signal.
PDF forensics front door
PyMuPDF renders every page natively, decides born-digital vs. scanned, and extracts provenance metadata and font-embedding data in a single pass shared by every other source.
Type classification
A 16-class document-type model (DiT) plus a vision-language read (Qwen2.5-VL) for free-form type, subtype, and key-field extraction.
Layout & tables
Reading order, layout regions, and table structure via docling, with deepdoctection as an independent full-pipeline corroborator.
Conditional OCR
Runs only on image-only pages by default; in always mode it also powers an embedded-text-vs-OCR tamper diff — a mismatch means the visible page doesn't match the content stream.
Signature integrity
pyHanko reports embedded PDF signature count, coverage (entire file vs. partial), and cryptographic intactness.
Field & type engine
Runs inside every scan. Extracts typed business fields, commits to a document type only when independent evidence corroborates (declining to guess otherwise), and tracks per-vendor term drift over time. Its output populates the report's document-identification and field-analysis blocks.
Field extraction
A broad set of field types — invoice/PO numbers, amounts, dates, IBANs, routing numbers, MC/DOT numbers, VINs, policy numbers, addresses, and more — each with a validated / unvalidated flag.
Type classification
Corroboration-gated: a document whose term/field/structure signals don't agree is reported as unclassified rather than forcing a low-confidence guess. Composite documents (multiple types in one file) are flagged as such.
3-way identification
Reconciles its own type vote against the keyword classifier and the recognizer into one headline type per report, with per-engine confidence shown side by side.
Vendor-drift baseline
Per-user, Supabase-backed term-frequency baseline per vendor; a document missing terms that vendor's history always includes, or carrying terms it's never used, is flagged as drift once enough history exists.
Cross-document field consistency
A shared invoice/PO/amount field that disagrees in value across two related documents is a core altered-document or double-brokering tell.
Cross-document intelligence
Runs after every scan, comparing the current document against prior submissions — intentionally not scoped to a single account, since fraud rings span customers.
Submission velocity
The same file (by SHA-256) submitted repeatedly in a short window.
Near-duplicate detection
SimHash comparison finds a lightly-edited version of a document that's byte-different but structurally the same.
Entity-link / fraud-ring detection
Shared MC/DOT/policy/IBAN/VIN/routing numbers, emails, or phones across otherwise-unrelated documents.
Template registry
Known-good issuer layouts (page-1 perceptual hash) that PDF template matching checks against.
Review queue
Flagged (review/reject) reports awaiting a human disposition — approve, reject, or escalate.
Format coverage
The formats the engines accept — routed automatically by extension and magic bytes, so a disguised file is scanned as what it really is, not what it claims to be.
Forensic authenticity, threat & final action
Authenticity evidence, file danger, and account handling policy remain separate, so a phishing email, a forged invoice, and a routing rule are never conflated into one score.
Authenticity axis
Structural, metadata, content, and field evidence — is this document what it claims to be, and has it been tampered with?
Threat axis
Malware, phishing, and active-content risk — could opening or acting on this document harm you?
Account policy
Company, document-type, and custom handling rules may require review or rejection without changing the authenticity score or claiming tampering.
Corroboration gate
The forensic verdict escalates only on a curated decisive signal, corroboration across two or more distinct categories, or a calibrated score — never a single noisy medium-severity finding.
Per-category caps
Each signal category's contribution is capped before summing, so one noisy domain (say, ten low-severity font notes) can't out-score a genuinely cross-cutting problem.
Check catalog
Every signal category the engines can raise, with the number of distinct checks in the source that can raise it — computed live from the running code (GET /api/checks), not a hand-maintained figure that can drift.
Loading…
Standalone tools
Beyond the full scan, Docurensic ships focused single-file tools — small forensic labs and a PDF workbench you point at one document, link, or company. They all open from the Tools page (sidebar → Analysis), and most can pull a file straight from your File Vault so you don't have to re-upload. Jump to one:
URL analyzer
One link, two lenses: is it dangerous, and how well-built is the site behind it?
Paste a URL or bare domain and Docurensic builds one shared picture of the site from real, live probes, then reads it two ways: a corroboration-gated safety verdict and a six-dimension A–F quality scorecard. It never convicts on one tell — a site is only flagged when the evidence agrees across at least two independent tiers and three categories. Reach for it before trusting a link in an email, invoice, or message — a supplier portal, a payment page, a "verify your account" link, or a vendor domain you've never dealt with.
- ✓Lexical link checks — insecure scheme, raw-IP hosts, "@" authority tricks, typosquat and brand-lookalike domains, punycode/homograph names, suspicious keywords, expanded URL shorteners, and open-redirect parameters.
- ✓Registration & infrastructure — domain age (a newly-registered domain is a strong fraud tell), registrar and locks, DNS and mail records (SPF/DMARC/CAA), and a live TLS certificate handshake.
- ✓Landing-page content — login forms on brand-new domains, credential forms posting elsewhere, hidden inputs, obfuscated scripts, invisible iframes, and brand-claiming titles on unrelated domains.
- ✓Quality scorecard — Setup/Security, SSL/TLS, Optimization, SEO, Tech stack, and Age/Trust, each graded, with a ranked "fix first" list and a tech-stack fingerprint (with end-of-life flags).
Good to know: an external reputation check runs only when configured; a clean read is not a guarantee a site is safe. Open the URL analyzer →
Company analyzer
Is this business real, who runs it, what's its reputation, and what to watch for?
Give it a company name, a domain, or a registration number and it returns one reconciled business report. Narrow data sources each verify a single fact, open-web discovery fills in the identity and story, and an AI analyst writes it up in plain English. The trust score and verdict are computed by fixed, auditable rules — the AI only writes the explanation around the numbers, never the numbers themselves. Use it during vendor or counterparty onboarding — before signing a new supplier, extending credit, or paying a first invoice.
- ✓Profile & people — legal/display name, industry, founded year, headquarters, phone, website, headcount estimate, and leadership names where discoverable.
- ✓Reputation — review snippets and a rating, sentiment themes, and dated, sourced news summaries.
- ✓Independent verifications — a real, reachable address, whether the phone is a live line (and its type), and whether the website domain is flagged as unsafe (a hard push toward high risk).
- ✓One clear read — a 0–100 trust score, letter grade, and verdict with confidence, a corroboration matrix (identity / address / phone / safety / reviews / leadership), and a "before you engage" checklist.
Good to know: results depend on which data sources are enabled; treat it as a fast starting point, not formal due diligence. Open the Company analyzer →
Metadata Lab
Everything a file says about itself — and every place its own story contradicts itself.
Deep-inspect one file's metadata without running a full scan. It lays out every property the file carries grouped by source, builds a chronological timeline of the file's own claimed dates, and surfaces a contradictions list — the places where the file disagrees with itself. Reach for it to interrogate one document's provenance fast: a scanned invoice whose dates feel off, a contract you suspect was edited after signing, a photo whose camera/GPS story you want, or an email whose headers you want laid out.
- ✓Grouped properties — PDF Info + XMP (raw and parsed), edit history, fonts and encryption; image EXIF/GPS/camera; Office core/app properties and macro presence; email headers and relay chain.
- ✓Date timeline — every claimed date in order, each tagged with where it came from.
- ✓Contradictions — dates that run backwards or sit in the future, Info-vs-XMP disagreements, content appended after a signature, and embedded macros, files, or JavaScript.
- ✓Revision history — a PDF's incremental-update map with which revision each signature covers.
Good to know: a contradiction is a lead to investigate, not proof of forgery — and a clean metadata story doesn't prove authenticity. Open any vault file with its "Metadata" action. Open the Metadata Lab →
Image forensics
Five independent visual lenses over one image to reveal edits, clones, and mismatched compression.
Run one image — or one rasterized PDF page — through five manipulation-analysis lenses, each shown as a heatmap or overlay. It deliberately gives you lenses and evidence to read, not a verdict or a score. Use it on a suspicious photo or image-based document — a receipt, an ID photo, a screenshot, a photographed invoice — when you want to see whether a number, stamp, signature, or region was edited, cloned, or pasted in.
- ✓Error-level analysis (ELA) at three qualities — heatmaps that highlight regions edited or pasted at a different compression history.
- ✓JPEG-ghost sweep — a per-block map that exposes a region carrying its own separate compression history.
- ✓Copy-move / clone detection — overlays lines and source/destination boxes where part of the image was duplicated to cover or repeat content.
- ✓Noise consistency — flags smoothed or inserted patches whose noise doesn't match the surrounding image; plus the JPEG quantization tables and an estimated original save quality.
Good to know: each lens has stated blind spots — heavily re-saved, resized, or screenshot images can wash the signals out. Read the lenses together as evidence, not a single answer. Open Image forensics →
Signature checks
Real cryptographic verification of a PDF's digital signatures — what they cover and what changed after.
Verifies the digital signatures inside a PDF without needing a trust store. For each signature it checks the underlying cryptography, reads the certificate the file carries about the signer, confirms the covered bytes are untouched since signing, and reports exactly what the signature covers and what was appended afterward. A second mode geometrically compares two signature images. Reach for it whenever a signed PDF matters — a contract, agreement, or certificate — to confirm the signature is intact and the document wasn't altered after signing.
- ✓Coverage & sealing — how much of the document each signature covers, which revision it seals, and whether bytes were added after signing.
- ✓Cryptographic integrity — the document digest still matches the signed byte range, and the signer's key actually produced the signature (RSA/PSS/ECDSA).
- ✓Signer certificate — subject/issuer, self-signed detection, validity window, and whether a trusted timestamp is present.
- ✓Signature compare — a geometric similarity score for two signature images (ink overlap, shape, stroke direction).
Good to know: it deliberately does not judge whether the signer's certificate chains to a trusted root or check revocation, and it says so. Compare is a similarity aid, not handwriting identification. Open Signature checks →
Compare
Overlay two versions of a document and see every change — down to the font.
Puts two documents side by side and surfaces every difference — text edits, moved content, and even font or size substitutions on otherwise identical text, so a changed amount typed in a slightly different font stands out. Use it when you have an original and a returned or countersigned copy of the same document — a contract, purchase order, invoice, or certificate — and need to confirm nothing was quietly altered.
- ✓Word-level diff with the changed regions located on the page.
- ✓Font forensics — flags where the same text appears in a different font or size, with a full font inventory of each side and what's new or missing.
- ✓Four visual modes — an X-ray channel overlay with blink, a pixel-difference heatmap, clustered difference boxes, and side-by-side highlighted changes.
- ✓Images welcome — either side can be an upload or a vault file, and an image is turned into a one-page document so you can compare a photo against a PDF.
Good to know: it shows you what differs — deciding whether a change is legitimate is up to you. Vault rows deep-link in with their "Compare" action. Open Compare →
PDF Studio
A full in-browser PDF workbench: edit, organize, sign, convert, redact, and protect.
Everyday PDF utilities built right into the product, so you don't need a separate app — all bytes-in, bytes-out, with nothing stored unless you save it. This is a productivity toolset, not a forensic analyzer: nothing here feeds scan scores or verdicts. Reach for it for the routine document handling around a case — combine exhibits, split out pages, redact personal data before sharing, OCR a scan, fill a form, or password-protect a file.
- ✓Pages & conversions — merge, insert, split, rotate, reorganize, crop, compress; convert to/from text, HTML, and images; and OCR a scan to make it searchable.
- ✓Markup & forms — shapes, arrows, highlight/underline/strikeout, investigation stamps, and list/fill/author form fields.
- ✓Permanent redaction by area, search term, or pattern (SSN/SIN, card numbers, email, phone) — the content under the box is genuinely removed, not just covered.
- ✓Sign & protect — sign, stamp, add text, watermark, edit metadata, flatten, and password-protect or unlock. Round-trips with the vault: open a vault PDF, save the edit back as a new file.
Good to know: search/pattern redaction finds text it can read — personal data baked into a scanned image needs an area redaction (run "Make searchable" first to check). Open PDF Studio →
X-Ray Engine
An evidence-first workbench that reconstructs a PDF's history and lays out competing explanations.
A standalone, PDF-only investigation workbench that goes deeper than a normal scan: it reconstructs a document's editing history, pinpoints exactly where and what changed, and lays out competing explanations with the evidence for and against each. It's explicitly a reasoning-and-evidence tool, not a pass/fail scorer. Reach for it on a high-stakes PDF where "what exactly changed, and where" has to be defensible. (A distilled Deep X-Ray version also runs automatically after every ordinary PDF scan and appears in the report.)
- ✓Recovered values — reconstructs recoverable prior revisions and shows before / after / difference crops of exactly what a field used to say.
- ✓Located findings — anchors each finding to an exact region on the page, with review crops of the key values to check against your source of truth.
- ✓Layers & construction — reconstructs document layers and hidden or covered content, and classifies how the document was built (native-digital, searchable scan, image-only capture, mixed).
- ✓Honest coverage — an amount-reconciliation and attention heatmap to focus review, with a complete / limited / skipped / failed note on every lens; missing coverage never means "clean".
Good to know: construction mode is not authenticity, tool markers don't identify the human editor, and it carries no risk score or verdict — it presents evidence, not intent. Open the X-Ray Engine →
Files, cases & your workspace
Where your documents live, how you group them, and where decisions and activity show up.
File Vault
Your private, encrypted home for documents. Every file is encrypted the moment you upload it and stays encrypted at rest under a key unique to your account, so only you can open your files. Organize them in folders; your storage allowance is set by your plan. Every scan automatically tucks the original into a "Scanned documents" folder (duplicates aren't stored twice) and badges it with its latest verdict, and any file offers "Send to analysis" to re-run the full scan. Open File Vault →
Cloud connections
Bring your own storage without moving files. Connect one Dropbox or Google Drive folder as a read-only mirror: files stay in your cloud, are indexed rather than copied, and use none of your vault quota. The mirror refreshes on a schedule and on "Sync now", and per-connection auto-scan runs any new file that arrives after you connect (never your initial import). Disconnecting just drops the index and access — your files stay put. Open Connections →
Cases
Group related documents — an application packet, a shipment's paperwork, a claim — and investigate them as one set. Docurensic suggests links when two documents share a strong identifier (email, IBAN, VIN, phone) or template, and you can link by hand; a connection graph draws the relationships, and the whole set can be reviewed and dispositioned together. Open Cases →
Review inbox
The human decision queue. Any document whose final action came back Review or Reject waits here for a person's call, so nothing that needs a second look slips through. Record a disposition — Approve, Reject, or Escalate — with an optional note; it's tracked separately from the automated verdict and feeds your review stats and workflows. Open the Review inbox →
Analytics
A bird's-eye view of your activity: how many documents you've scanned over time, the mix of trusted / review / reject with the trend against the prior period, your overall flag rate, and the forensic flags that fire most often across your recent reports. Adjust the time window to focus. Open Analytics →
Notifications
Your in-app alert feed (distinct from the Review inbox). It collects alerts raised by your Smart Workflows — a flagged document, a routed file — plus system messages from the Docurensic team, newest first. Mark one or all read, delete an alert, or clear everything you've already read. Open Notifications →
Smart Workflows — decide once, run on every document
Docurensic doesn't stop at the verdict. In the visual, no-code Workflows builder you write rules as WHEN something happens, IF it matches your conditions, THEN run these steps — and it runs automatically on every document after that. Jump to a section:
The WHEN / IF / THEN model
Every workflow reads the same way. WHEN picks the moment (a document finishing analysis, a human decision, a timer, or a manual run). IF is a set of optional filters, all AND-ed together — leave them empty to act on everything. THEN is a tree of steps that runs in order: each step is an action or a conditional split.
Triggers — the WHEN
The event that starts a workflow. Most derive from a finished scan; two run on their own.
| Trigger | Fires when |
|---|---|
| Analysis completed | A scan finished, whatever the outcome — the catch-all starting point. |
| Document uploaded | A file entered the pipeline: uploaded here, sent via API, added to the Vault, or arriving from a connected Dropbox/Drive folder. |
| Document classified | The document type was identified (invoice, ID, certificate…). Pair with the "document type" condition to route by kind. |
| High fraud score | The forensic risk score met your threshold (defaults to 60; set your own with the minimum-risk condition). |
| Needs review | The final decision requires a human — including separate threat or advisory reasons. |
| Final action: reject | The final decision is reject, whether from forensic evidence or your account policy. |
| Final action: accept | The final decision is accept. Use it to auto-file documents that need no review. |
| Threat detected | An active-content / payload threat level was raised (structural detection, not antivirus). |
| Signature detected | A digital signature, signature field, or visible signature mark was found. |
| Duplicate found | The same or a near-identical document was analyzed before (possible double submission). |
| Document expired | An expiry finding fired — lapsed insurance, an out-of-date certificate. |
| Disposable email found | A contact email on the document uses a throwaway / temporary-inbox domain. |
| Low analysis confidence | The engines were unsure of their assessment or of the document's type. |
| OCR completed | Text was read off an image — a scan or photo that needed OCR. |
| QR / barcode read | A QR code or barcode was decoded (shipping labels, tickets, certificates). |
| Manual review completed | A reviewer set a disposition (approved / rejected / escalated). Fires from the Review inbox, not a scan. |
| Engine rule matched | One of your Engine Detection custom rules fired. Narrow it with the "Engine rule names" condition. |
| On a schedule | Runs on a timer you set (hourly / daily / weekly / monthly in your time zone), not tied to any one document. |
| Run manually | Runs only when you press Run — on a document you pick from history, or against the sample in test mode. |
Conditions — the IF
Optional filters, all AND-ed together. Leave them empty and the workflow acts on everything the trigger catches; add a few to narrow it to exactly the documents you care about.
| Condition | Runs only when |
|---|---|
| Risk score is at least / at most | The forensic risk score is ≥ / ≤ your number (0–100). Combine both for a band. |
| Verdict is one of | The forensic authenticity verdict is trusted, review, or reject (your pick). This can differ from the final action when policy forces a review. |
| Threat level is at least | The active-content threat level is at least elevated, high, or critical. |
| Has an active-content threat | The document carries a payload / active-content threat. |
| Is a duplicate | This document matches one analyzed before. |
| Has an expiry / lapsed finding | An expiry or lapsed-coverage finding fired. |
| Has a signature | The document is signed. |
| Amount is at least / at most ($) | The document's parsed total is ≥ / ≤ your amount (e.g. invoices over an approval limit). |
| Document type is one of | The document is one of the listed types (invoice, insurance…). Any match qualifies. |
| Filename contains | The file name includes this text (e.g. "contract"). |
| Filename matches pattern | The name matches a glob pattern where * is any run of characters and ? is one (e.g. INV-*.pdf). |
| Has any tag of | The report carries at least one of the listed tags. |
| Came from | The document's origin is an app upload, API upload, scheduled run, or manual run. |
| Engine rule names | One of the named Engine Detection custom rules matched (pairs with the "Engine rule matched" trigger). |
Actions — the THEN
What a workflow does, grouped into five families. Message actions can reuse a saved
Connection instead of a pasted URL, and support
formulas like {{filename}}.
Alerts & messages
Reach your team where they already work, or POST to any URL. In-app notification, Email, Slack, Microsoft Teams, Discord, Telegram, Google Chat, and Webhook (the full event JSON, signed when a secret is set — a bridge to Zapier, Make, n8n, Power Automate, or your own systems).
Tag, file & organize
Add tags to the report for later filtering, Move to a vault folder, Quarantine (move to a Quarantine folder and tag it), Create case (open an investigation and link the document in), and Delete document (permanently removes your own vault copy — the report keeps its findings).
Review & notes
Change status sets the human review disposition — approved / rejected / escalated. This is the operational disposition only; it never changes the forensic score or verdict. Add note appends a note (with formulas) to the report's reviewer notes.
Investigate with tools
Run metadata check and Run signature check run the Metadata Lab or Signature checks as part of the flow and tag/note the report on contradictions or a broken signature. Open in a tool posts an inbox alert with a one-click link that opens the document, pre-loaded from the vault, in Metadata Lab, Image forensics, Signature checks, Compare, or PDF Studio.
Flow control
Wait, then continue pauses and resumes the rest of the steps after 1 minute to 7 days. Stop the flow ends the workflow early (even inside a split). Start another workflow runs a second workflow's steps on the same document (chaining one level deep).
Branching — one flow, many paths
The THEN side is a step tree, not just a flat list. Any step can be a conditional split: several labelled paths, each with its own condition, plus an optional "Otherwise". Only the first path whose condition matches runs — exactly like a switch — then the flow continues past the split. Splits can nest up to three levels deep.
WHEN Analysis completed
THEN split on:
┌ High risk (score ≥ 75) → create case + alert #fraud
├ Needs review → tag "to-review" + email the queue
└ Otherwise → move to "Cleared" folderA plain list of actions with no split validates and runs the same way, so branching is entirely optional. Limits keep flows readable: up to 10 top-level steps, up to 6 paths per split, and up to 40 actions across the whole tree.
Schedules & waiting
On a schedule
Run a workflow on a timer instead of on a document. Choose hourly, daily, weekly, or monthly and a time, all in an IANA time zone you pick (DST-correct). A background scheduler fires due workflows and re-arms the next run. Typical uses: a daily activity summary to your inbox, or a weekly review-queue reminder to Slack.
Activity digest
Scheduled runs carry a digest you can drop into messages: {{stats.scans}},
{{stats.high_risk}}, {{stats.rejected}},
{{stats.review_pending}}, and {{period}}
(the window covered, e.g. "the last 24 hours").
Wait, then continue
The "Wait" step pauses a flow and resumes the remaining steps after anywhere from 1 minute to 7 days — so you can "alert now, then follow up tomorrow" or "quarantine now, wait 7 days, then delete". In a Test run it continues immediately instead of actually waiting.
Test it before you trust it
Three safe ways to try a workflow without waiting for a real document to trip it.
Test in the builder
The builder's Test button runs your unsaved draft against a built-in sample and shows each step's outcome — without saving the workflow or logging a run. Try splits and formulas freely before committing.
Test a saved workflow
Fire one saved workflow against the sample document. Its actions really execute, but outbound messages are marked as tests and document-changing actions skip (the sample has no real file). Scoped to just that workflow.
Run now
Run a saved workflow on demand against a real report you pick from your history — real execution, real side effects — or against the sample in test mode. Conditions still gate the run unless you choose to force it.
Start from a template
You rarely start from a blank canvas.
Suggested rules gallery
One-click, ready-made workflows grouped into Triage & routing, Scheduled digests, Investigate with tools, Essential alerts, and Integrations. Picking one opens it prefilled so you can adjust thresholds and recipients before saving — triage-by-risk (a three-way split), route-invoices-by-amount, daily/weekly digests, high-risk follow-up with a delay, quarantine-then-purge, and push-to-Slack/Teams/Discord/Telegram/Google Chat/webhook.
Default starters
A small safety net is seeded automatically the first time you open Smart Workflows — notify on completion, tag documents needing review, auto-open a case on very high risk, alert on rejected / duplicate / expired / threat, and label documents by risk tier. They're ordinary rules you can edit, disable, or delete, and a deleted default is never re-created.
Formulas — put document data in your messages
Every message field — notifications, emails, chat and webhook messages, notes, case titles — supports
formula tokens of the form {{field|filter:arg}}. An insert-data picker in the
builder means you don't have to type token names.
🚩 {{filename}} — {{doc_type|title}} — scored {{risk_score}}/100
Verdict: {{verdict|upper}} Amount: {{amount|currency:$|number}}
Review it: {{report_url}}Fields include document facts
({{filename}}, {{doc_type}},
{{amount}}, {{tags}}), analysis facts
({{verdict}}, {{risk_score}},
{{threat_level}}, {{report_url}}), run facts
({{workflow}}, {{date}}), and the scheduled
digest stats above. Filters chain after a pipe: upper,
lower, title,
round:1, currency:€,
number (thousands separators), truncate:40, and
default:none. A known field with no value shows an em dash, a typo'd token stays
visible so you can spot it, and a broken formula never sinks the message.
Connections — set a destination once, reuse it everywhere
Save a URL or token once on the Connections page and every workflow can reference it by name — update the connection and every flow that uses it updates too. Each saved connection has a one-click test ping, and secret fields (like bot tokens) are write-only.
Messaging & alerts
Slack, Microsoft Teams, Discord, Telegram, Google Chat, and named email recipient lists. Mattermost and Rocket.Chat save as Slack-type (they speak Slack's webhook format).
File storage
Dropbox and Google Drive folders that mirror read-only into your vault and can auto-analyze new arrivals — the same cloud connections described under Files & workspace.
Automation bridges
Zapier, Make, n8n, and Power Automate as catch-hook webhooks — one connection reaches thousands of downstream apps (CRMs, ERPs, spreadsheets, ticketing).
Developer & scan-lifecycle webhooks
Your own custom webhook endpoints, plus platform webhooks that subscribe your systems to scan-lifecycle events (queued / processing / complete / failed), gated by risk level and HMAC-signed on every delivery.
Integration API (API keys)
The supported, stable way to run Docurensic from your own systems. It is a small,
analysis-only surface under /api/v1, authenticated with an API key — separate
from the browser session endpoints documented further down this page.
Getting a key
Open Settings → API access, turn on Enable API access (the master switch — while it
is off, every API request returns 403), then click Generate key. The key
(drk_live_…) is shown once — store it securely; only a SHA-256 hash is kept
server-side. Regenerate or revoke it in Settings at any time; the previous key stops working immediately. Send it on
every request as Authorization: Bearer drk_live_… or
X-API-Key: drk_live_….
curl -s https://docurensic.com/api/v1/analyze \ -H "Authorization: Bearer $DOCURENSIC_API_KEY" \ -F "[email protected]"
Same accepted types and 150 MB limit as the app. The call is synchronous — typical latency is
tens of seconds for the full forensic + AI pass. The response is the fixed v1 document payload (verdict, score,
the decision contract, document type, and the slim case-file answers). Full field-by-field
reference: the API guide shipped with your deployment.
Analyze one document. Multipart body: file. Returns the v1 document payload with a
document_key you can re-fetch. Runs the identical pipeline as the app (report persisted to your
history, auto-archived to File Vault unless storageless, webhooks fire).
Key check. Returns 200 with your user_id when
the key is valid and API access is enabled.
Re-fetch a previous analysis (API- or app-created — it is one history). 404 if the key doesn't exist or belongs to another account.
The advisory X-Ray layer manifest attached to a completed PDF scan; overlay artifacts are served
from /api/v1/scans/{scan_id}/xray/artifacts/{artifact_id}.
POST /api/v1/analyze shares your per-account scan budget (20 / minute);
other v1 endpoints allow 60 requests / minute per account. Browser session tokens are not accepted on
/api/v1/*, and API keys are not accepted on the application endpoints below — the
two surfaces are deliberately separate, which is what makes usage attributable.
Application endpoints (browser session)
Every feature in the app is backed by a REST endpoint under /api.
These are the app's own browser surface — authenticated with a Supabase session token, not an API key — documented
here for reference. For programmatic integration from your own systems, use the Integration API above. Requests
and responses are JSON unless a file upload is involved (then multipart/form-data).
Authentication
Docurensic uses Supabase Auth. Sign in to obtain an access token, then send it as a bearer token on every request. Tokens are verified server-side (with a 60-second in-process cache). A handful of reference endpoints are public — they're marked public below.
curl https://docurensic.com/api/scan \ -H "Authorization: Bearer <your-access-token>" \ -F "[email protected]"
Scan & history
Run a full forensic scan and manage the reports it produces. A report is persisted to your workspace (unless your account is in storageless mode) and is only ever visible to you.
Full forensic scan. Multipart body: file (the document). Returns the complete report —
decision action, forensic verdict and score, threat level, identification, signals, fields, and, when document text
is available, the AI review (classification + type-aware checks + entities). No-text documents return an explicit
unavailable state.
List your reports. Query params: q (filename or SHA-256 prefix), verdict
(trusted/review/reject), needs_review, limit (1–100,
default 50), offset. Returns {rows, total, limit, offset}.
Fetch one full report by id.
Delete a single report.
Clear your entire scan history.
Aggregate counts for the workspace. See also /api/scan/analytics,
/api/scan/analytics/overview, and /api/scan/analytics/flags for the Analytics page data.
Find reports in your workspace by full SHA-256 or a 12+ character prefix.
Standalone analysis
Run a single engine directly, without a full scan — the endpoints behind the Tools page.
AI-generated-text detection on pasted text. JSON body:
{ "text": "…", "engine": "auto", "perplexity": false }. Returns the verdict, a score, and every
stylometric signal individually. perplexity: true adds a language-model perplexity pass (slower).
Same AI-text analysis on an uploaded file (multipart file) — text is
extracted from PDF/DOCX/TXT/image (OCR when available) first.
Metadata & provenance only (no full scan). Multipart file; PDF and Office
(docx/xlsx/pptx/doc/xls/ppt) only. Returns producer/creator, timestamps, origin classification, and any
provenance flags — the exact functions a full scan uses, so the numbers can't drift.
Email header & message analysis. JSON body: { "raw": "…" } — pasted raw
headers or a full raw message. Returns key headers, parsed SPF/DKIM/DMARC results, the Received relay
chain with hop delays, forensic signals (BEC reply-to redirect, display-name spoof, link-text/href
mismatch), and every body URL with a passive reputation { risk, band } score.
Same analysis for an uploaded .eml / .msg file (multipart
file).
Multi-source document recognition & forensics fusion (own rate-limit bucket). Multipart
file; query params disable (comma-separated source list, e.g. disable=vlm),
ocr_mode, ocr_engine. Returns document type with agreement count, born-digital/scanned,
signature state, and the list of sources used.
Example — URL reputation response
{
"domain": "paypal.com.account-verify.ru",
"risk": 88,
"band": "flag",
"confidence": "high",
"engine": "aidetect_single",
"structural_flags": ["brand_token_outside_etld1", "suspicious_keyword"],
"domain_age_days": 12,
"evidence": [
{ "feature": "typosquat_distance", "score": 0.9, "note": "resembles paypal.com" }
]
}Account
Your scan counts, broken down by verdict and threat level, plus the timestamp of your last scan.
Your scan-configuration settings, merged over defaults.
Partial update. Keys: storage_mode, sensitivity,
industry_preset, and a features map. Unknown keys are
ignored; invalid values rejected. See the Account tab for the full list of accepted values.
Everything held for your account — scan reports, settings, and vendor baselines — as one portable JSON download. Free on all plans.
Permanently delete your account and all associated data. Irreversible — the UI gates it behind a typed confirmation.
Your Docurensic AI usage log — tokens, model, latency, and estimated cost per call.
/api/ai-usage/{row_id} returns a single entry.
Review & cases
Documents whose final action is Review or Reject, with no disposition yet, awaiting human review.
/api/review/queue/count returns just the count for the nav badge.
Record an examiner decision. JSON body:
{ "disposition": "approved" | "rejected" | "escalated", "reviewer_note": "…" }.
Create a case ({ "title": "…" }) or list your cases with document counts. Fetch one
with GET /api/cases/{case_id}.
Link a document to a case: { "scan_id": "…", "link_reason": "manual" }. Remove with
DELETE /api/cases/{case_id}/documents/{scan_id}.
Reference & health — public, no auth
The live check catalog — every signal category the engines emit, its human label, and how many distinct checks in the source can raise it. This is what powers the catalog table in the How-it-works tab.
The producer-reliability tier table (which creation tools are trusted vs. suspect), ranked high → low. Also browsable at /producers.
The Supabase URL and anon key the browser client needs to sign in.
Liveness check.
Docurensic in your business
Docurensic ships tuned profiles for the document types that carry the most fraud, and industry rule
packs that layer sector-specific checks on top. You can pin your workspace to an industry in
Settings → Industry preset (auto,
freight, insurance,
lending, legal, or
hr), or leave it on auto-detect. The scenarios below are grounded in the actual
doc-type profiles and cross-document intelligence the engines run.
Freight & logistics
Brokers, 3PLs, and carriers vetting shipment paperwork
Double-brokering and carrier-identity fraud run on altered rate confirmations and bills of lading. Docurensic's freight pack reads MC/USDOT numbers off the page and flags them for FMCSA verification, while cross-document intelligence catches the same carrier identity reused across unrelated loads.
- ✓MC/USDOT extraction & FMCSA availability — the freight rule pack surfaces the carrier's authority number and flags it for verification.
- ✓Cross-document entity links — a shared MC/DOT number, email, or phone across otherwise-unrelated documents is a fraud-ring tell.
- ✓Rate / amount tampering — a rate-confirmation amount that was edited after issue leaves metadata and font-splice traces.
- ✓Near-duplicate detection — a lightly-edited copy of a real rate con, byte-different but structurally identical.
Lending & finance
Underwriters and finance teams verifying supporting documents
Loan and credit decisions rest on documents that are easy to doctor — invoices, purchase orders, pay stubs, bank statements. Docurensic checks the arithmetic and the account numbers, not just the look, and applies statistical fraud tests to any tabular data.
- ✓Checksum validation — IBAN, routing-number, and VIN check digits are verified, so a swapped payee account fails arithmetically.
- ✓Cross-field arithmetic — line items that don't sum to the stated total, or tax that doesn't reconcile.
- ✓Benford's law & balance reconciliation — spreadsheet/CSV statements get a chi-square Benford test, duplicate and round-number clustering, sequence-gap and future-date detection.
- ✓Cross-document field consistency — an invoice number or amount that disagrees between a PO and its invoice.
Insurance
Carriers and brokers checking coverage evidence
A certificate of insurance is trivial to edit in a PDF viewer — bump a coverage limit, extend an expiration date, swap an insured name. Docurensic's doctype profile is date-aware (a future expiration is expected, not treated as back-dating) and catches the edits themselves.
- ✓Doctype-aware timeline checks — expiration in the future is normal for a COI; a modified date after issue, or an impossible timestamp, is not.
- ✓Font-splice & revision detection — an edited coverage limit or policy number leaves a mismatched font subset or a recoverable prior revision.
- ✓Policy-number entity links — the same policy number appearing across unrelated certificates.
- ✓Template matching — page-1 perceptual hash against a registry of known issuer layouts.
HR & onboarding
Verifying identity and tax documents at hire
Onboarding packets are a common vector for synthetic identities — a doctored W-2, a fabricated W-9, a photoshopped ID. Docurensic checks the tax forms structurally and validates the machine-readable zone on ID scans.
- ✓MRZ check-digit validation — TD3 passport / ID machine-readable zones are verified digit-by-digit; a forger who edits the printed number and forgets the MRZ is caught.
- ✓Tax-form field logic — W-2/W-9 required fields, EIN/SSN formatting, and cross-field consistency.
- ✓AI-image & synthetic-scan detection — an ID photo that's AI-generated, or a born-digital form re-rendered to an image to dodge raster checks.
- ✓Embedded-image EXIF — editor software and GPS traces left in a scanned document's images.
Legal & professional services
Contracts, signatures, and the email they arrive in
A signed PDF contract, and the message that delivered it, are both attack surfaces. Docurensic validates digital signatures cryptographically and reads email authentication headers to catch the business-email-compromise redirect shape.
- ✓Signature integrity — cryptographic validity, byte-range coverage, shadow-attack detection, and certification vs. approval signatures.
- ✓Hidden edits — an incremental update after signing, or a form field altered post-execution.
- ✓Email authentication — SPF/DKIM/DMARC parsing and From/Reply-To/Return-Path mismatch (the classic BEC redirect).
- ✓Link & brand spoofing — punycode domains, typosquats, and hidden auto-executing URLs in the document or message.
Make detection fit your workflow
Every account scans with the same calibrated engine, but no two document flows look alike: a freight broker's rate confirmations, a lender's bank statements, and an HR team's diplomas carry different "normal". The Detection panel on the Settings page lets you tune what gets flagged, how strictly verdicts are drawn, and what your own business rules add on top — without ever weakening the core forensics. The controls below are read-only until you flip the Custom engine settings switch at the top of the Detection panel — the sensitivity dial applies either way, and while the switch is off the engine runs on its calibrated defaults.
Sensitivity dial
One dial that scales the review/reject thresholds. Turn it down when a noisy intake stream produces too many borderline reviews; turn it up when a missed forgery costs more than an extra manual look.
Settings → Detection → SensitivityCategory switches
Advisory categories — AI-content analysis, web corroboration, and similar context checks — can be switched off per account. Signals in a disabled category are dropped before scoring, so a category that doesn't apply to your documents stops contributing noise entirely.
Advisory onlyVerdict thresholds
Set your own review and reject score boundaries per account. Raising a threshold quiets borderline noise; it never overrides hard evidence (see the safety floor below).
Numbers you controlThe safety floor
Core forensic checks — file structure, revision history, font forensics, signature coverage — always run and can never be disabled. And whenever a hard core-forensic signal stands, the engine's own verdict is a floor: your thresholds can quiet noise, but they can never silently auto-trust a caught tamper.
Always onYour business rules, as formulas
Custom rules let you express policy the engine can't know: which amounts matter to you, which vendors you trust, which document types deserve extra scrutiny. Rules are built in a visual formula builder — no code — from three ingredient types, and their outcome is always a separate policy action (tag the report, require review), never a change to the forensic evidence itself.
Document fields
Values extracted from the document — the total amount, dates, sender domain, page count. Example
operand: document total.
Engine scores
The scan's own numbers — risk score, threat level, AI-content likelihood — usable as comparisons:
risk score ≥ 40.
Detection flags
Did a family of checks fire? font anomaly fired,
metadata contradiction fired — yes/no facts you can combine with
AND / OR / NOT.
WHEN document type is Invoice IF document total ≥ 10,000 AND (font anomaly fired OR revision history fired) THEN require review + tag "large-invoice-check"
This rule never claims the invoice is forged — it records that your policy wants human eyes on large invoices with any typography or edit-history finding. The tag it applies is visible on the report, filterable in your history, and usable as a Smart Workflows trigger ("my custom rule matched") to alert a channel or route the document to a tool automatically.
Company profile and per-type overrides
Two more layers tailor scans to your paperwork rather than paperwork in general.
Company profile checks
Tell Docurensic your own company details — name, tax id, bank account — and it checks documents that claim to be yours against them. A "your" invoice carrying an unknown bank account is exactly the payment-redirect fraud pattern; the mismatch requests review as a policy finding.
Per-document-type overrides
Different strictness for different types: run bank statements a notch more sensitive than general correspondence, or require review for every contract regardless of score. Overrides scope to the document type the scan identified.
Policy vs. evidence — kept apart
Everything on this page adjusts operational policy: what needs review in your workflow. The forensic risk score and verdict stay evidence-only, so a report always shows honestly which findings are engine evidence and which are your account's rules.
Your account
Signing in, configuring how scans behave, and managing your data. Everything here lives on the Settings page.
Signing in
Docurensic uses email + password authentication (Supabase Auth). Create an account on the sign-up page; if you forget your password, use the reset link on the login page. Your session is what authorizes every scan and keeps your history private to you.
What's tied to your account
Every scan, report, case, and vendor-drift baseline is scoped to your user id and enforced by row-level security — only you can read or delete it. Your email, display name, and sign-in history come straight from your auth session.
Scan settings
Configure how every scan behaves. These are the exact keys the PUT /api/account/settings endpoint accepts.
| Setting | Values | What it does |
|---|---|---|
| storage_mode | store · storageless |
store keeps the full report and automatically archives the analyzed original, encrypted, when File Vault storage is available and quota allows. storageless creates no new source copy; an existing vault or connected-storage file remains there. A slim history record retains filename/basic scan metadata, forensic verdict, decision action, findings, safe document type, and verification hash, without model, X-Ray-span, or provenance-value payloads. |
| sensitivity | conservative · balanced · aggressive |
How readily the forensic authenticity verdict escalates. Balanced is the default; conservative reduces false positives, aggressive surfaces more borderline evidence for review. Account-policy actions stay separate. |
| industry_preset | auto · freight · insurance · lending · legal · hr |
Pins the industry rule pack applied to your documents. auto detects it from the document type; only a tuned pack that matches the doc type ever applies. |
| features | a map of on/off toggles | Turn individual detectors on or off — ai_text, ai_image, recognition,
url_reputation, cross_document. Off by default: the placeholder tools
(tool_mrz and the extended lab features). |
Privacy & security
Controlled retention
Every upload uses temporary processing that is deleted after analysis. Store mode automatically encrypts the original in File Vault when storage and quota are available. Storageless creates no new source copy; existing vault or connected-storage files remain there, and only a slim history record is added.
Authenticated access
Every scan, history entry, case, and report is tied to your account via Supabase Auth and scoped by row-level security so only you can see or delete it.
AI data handling
When bounded document text is available, the AI review sees only extracted text, metadata, and the engines' findings — the file itself is never sent to the model. Raw document text is not retained in the usage log by default, and never in storageless mode.
Rate limiting
20 scans (and 20 recognitions) per 60 seconds per account, in separate buckets, to keep the service usable for everyone.
Transport & headers
HSTS on HTTPS, X-Frame-Options: DENY, X-Content-Type-Options: nosniff, a strict Referrer-Policy, and a restrictive Permissions-Policy on every response.
Content-mismatch rejection
Uploads are sniffed by magic bytes; a file whose real content is an executable, or an archive disguised with a document extension, is rejected before any parser sees it.
Exporting & deleting your data
Export everything
From Settings (or GET /api/account/export) you can download everything held for your account — every scan report, your settings, and your vendor baselines — as one portable JSON file. Free on all plans, any time.
Delete your account
Account deletion is self-service and permanent. It removes your scan history, vendor-drift baselines, and settings, then deletes your login itself. There is no recovery window — the UI gates it behind a typed confirmation.
Frequently asked questions
Can Docurensic detect if a PDF has been altered?
Yes. Because PDF stores edits by appending new content, we detect multiple revision markers (%%EOF) that indicate prior versions exist. Metadata timestamps that are physically impossible also flag tampering.
How long does a scan take?
Time varies by format, size, page count, OCR needs, and whether external AI is available. Structural checks are usually the fastest part; large, scanned, or complex documents take longer.
Is my document stored after the scan?
Every file uses temporary processing that is deleted after analysis. In storage mode, the analyzed original is automatically encrypted in File Vault when storage is available and quota allows. Storageless creates no new source copy; an existing vault or connected-storage file remains there. The slim history record keeps filename/basic scan metadata, forensic verdict, decision action, findings, safe document type, and verification hash.
Can Docurensic prove a document is genuine?
No — it detects forensic red flags, it doesn't certify authenticity. A low risk score means we found no evidence of manipulation, not a guarantee of genuineness.
Is my document sent to an AI model?
The file itself is never sent. When bounded extractable or OCR text exists, the AI review works from that text, metadata, and the forensic engines' findings only. With no extractable text, the external AI calls are skipped.
What file types are supported?
PDF (all versions), Word/Excel/PowerPoint (including legacy and macro-enabled), emails (.eml/.msg), images, CSV/TSV, and archives (.zip/.7z/.rar) — up to 150 MB.
Does it work on scanned paper documents?
Yes, with reduced signal coverage. Scanned-only PDFs (no extractable text) are flagged as image-only and OCR'd so content and logic checks still run — and "image-only" is itself useful for a document that claims to be born-digital.
Where do I report a problem or ask for help?
Open a case on the Help & support page — technical issues, questions, feedback, and feature requests all land in the same queue, and replies reach you in-app and by email. The Tools page lets you run any single engine on a one-off file or URL while you investigate.
Didn't find what you were looking for?
Open a support case and a human will get back to you — usually within one business day. The section you were reading rides along so you never have to explain where you got stuck.