Quality-gated workflow

Scan Rescue and OCR Readiness

Use Scan Rescue when text extraction failed, a PDF appears scanned, or an image contains text that needs review.

Step 1

Classify the source first

Check for a PDF text layer, scanned-page signals, existing extractable text, page rotation, and image-density clues before launching recognition on a file that may already be digital.

Step 2

Inspect scan and image quality

Review scan quality, image contrast, blankness, dimensions, file format, metadata, and language-readiness signals before accepting OCR output as usable.

Step 3

Run bounded OCR jobs

Use temporary OCR jobs for scanned PDFs or images only after readiness checks make likely quality limits clear, then compare the output before downstream reuse.

Step 4

Route editable-document expectations

If the destination is editable Word, check PDF-to-Word readiness separately because OCR text and editable DOCX conversion are different workflows.

Step 5

Use guides before retrying

When the output is empty, noisy, or misread, use OCR readiness and scanned-document guides before launching another conversion or sending the result onward.

Workflow notes and help

Route scanned PDFs, screenshots, and image text through text-layer checks, scan-quality review, OCR readiness, PDF OCR, and image-to-text workflows.

Scan rescue readiness
  • Ready for OCR attempt: The file has no usable text layer, scan-quality signals are acceptable, rotation/language concerns are known, and you have a checklist for reviewing names, numbers, and dates.
  • Use text extraction first: If the PDF has selectable text or digital text-layer signals, PDF to Text is usually cleaner and cheaper than OCR.
  • Needs scan cleanup review: Low contrast, rotation, tiny text, blank pages, mixed page sizes, or photo glare should be reviewed before spending a server OCR job.
  • Needs output verification: OCR output should be checked against the source for names, totals, punctuation, line breaks, and table structure before reuse.