Skip to content
PDF ToolsPopular~2s

OCR PDF

Run OCR on image-based PDF pages to extract text. Quality depends on scan resolution.

Examples

See what this tool can do.

Make a scan searchable

Run OCR on a scanned contract so you can select and copy its text.

Get text out of a photographed page

Extract the words from an image-only PDF for pasting elsewhere.

Tips & tricks

Get the most out of it.

  • 1Higher-resolution scans give far more accurate text.
  • 2The first run downloads a ~2 MB English language model, then it is cached.
  • 3For clean, digitally-created PDFs you don't need OCR — use PDF Text Extractor instead.

About OCR PDF

What this tool does

OCR PDF reads the text inside image-based (scanned or photographed) PDFs and makes it selectable, so you can copy, search and reuse it.

How to use it

  • Upload a scanned or image-only PDF.
  • Run OCR and wait while each page is recognised.
  • Copy the extracted text or download it.

How it works

Each page is rendered to an image and passed to Tesseract.js, an open-source OCR engine that runs entirely in your browser. A small language model is downloaded once and cached for later runs.

Important considerations

  • Accuracy tracks scan quality — resolution, contrast and straightness matter.
  • Large documents take longer, as every page is processed on your device.

Privacy

Your PDF is processed entirely in your browser. The file is never uploaded to a server, which makes this safe for confidential contracts, statements and personal documents. The one-time language-model download is code, not your document.

Related PDF tools

Frequently asked questions

No. Recognition runs locally with Tesseract.js. The only network activity is a one-time download of the language data (~2 MB), not your file.

Related tools