OCR PDF
Run OCR on image-based PDF pages to extract text. Quality depends on scan resolution.
Examples
See what this tool can do.
Make a scan searchable
Run OCR on a scanned contract so you can select and copy its text.
Get text out of a photographed page
Extract the words from an image-only PDF for pasting elsewhere.
Tips & tricks
Get the most out of it.
- 1Higher-resolution scans give far more accurate text.
- 2The first run downloads a ~2 MB English language model, then it is cached.
- 3For clean, digitally-created PDFs you don't need OCR — use PDF Text Extractor instead.
About OCR PDF
What this tool does
OCR PDF reads the text inside image-based (scanned or photographed) PDFs and makes it selectable, so you can copy, search and reuse it.
How to use it
- Upload a scanned or image-only PDF.
- Run OCR and wait while each page is recognised.
- Copy the extracted text or download it.
How it works
Each page is rendered to an image and passed to Tesseract.js, an open-source OCR engine that runs entirely in your browser. A small language model is downloaded once and cached for later runs.
Important considerations
- Accuracy tracks scan quality — resolution, contrast and straightness matter.
- Large documents take longer, as every page is processed on your device.
Privacy
Your PDF is processed entirely in your browser. The file is never uploaded to a server, which makes this safe for confidential contracts, statements and personal documents. The one-time language-model download is code, not your document.
Related PDF tools
Frequently asked questions
No. Recognition runs locally with Tesseract.js. The only network activity is a one-time download of the language data (~2 MB), not your file.
