Extract text from scanned PDFs and images using OCR.
Your documents never leave your device. OCR processing happens entirely in your browser using Tesseract.js Γ’β¬β no uploads, no servers.
OCR (Optical Character Recognition) extracts text from scanned documents and images, making it searchable and editable. Upload a scanned PDF or photo of a document, and our OCR engine recognizes the text Γ’β¬β all in your browser with zero server uploads.
Perfect for digitizing printed documents, extracting text from screenshots, making scanned PDFs searchable, or converting image-based content into editable text. Powered by Tesseract.js, one of the most accurate open-source OCR engines.
A scanned PDF is little more than a picture on each page β readable to the eye but invisible to search. OCR, or optical character recognition, turns those images into real text you can copy, search, and save. This tool runs the Tesseract OCR engine directly in your browser, so your scans and photos are recognized on your own device and never uploaded anywhere. The result is plain text you can review on screen, copy, or download as a TXT file.
Once a document has been through OCR, the words on each page stop being pixels and become editable characters. That changes what you can do with the file. You can pull a quote out of a scanned contract, search an old report for a specific term, or paste a paragraph from a photographed page into an email without retyping a single line.
The tool works with PDFs and with common image formats, so a photo of a whiteboard or a JPG screenshot of a document is fair game. Results are organized page by page, and for PDFs the tool renders each page at a sharp scale before recognition to give the engine the best possible input. That extra rendering step matters: a cleaner image means fewer misread characters.
Recognition focuses on English text, which covers most business and academic scans. Multi-page PDFs are handled automatically, one page at a time, so even a long document flows through without you needing to split it first.
You can recognize another document any time without reloading the page. The tool clears the previous result when you pick a new file, keeping the workflow fast and uncluttered.
Note that the output is extracted text rather than a re-rendered PDF. If you need a searchable PDF with an invisible text layer, the plain text you get here still gives you the content to work with, and you can keep it as a record or feed it into the next step of your process.
Digitizing a printed form so its data can be typed into a database, pulling meeting notes out of a photographed whiteboard, and turning a scanned article into a quote you can paste are all common uses. Anything where the goal is the words rather than the page layout is a good fit for plain text OCR output.
Pro tip
For the clearest results, feed the tool scans at a reasonable resolution. Very low-resolution or heavily rotated pages are harder for any OCR engine to read, so a straight, in-focus scan will always beat a blurry phone photo. Straighten pages before recognition if you can.
Free users can recognize text in files up to 10MB and up to three pages per run, which suits a single short scan or a couple of images. Premium subscribers unlock unlimited pages, larger files, and no wait between runs. Either way, the engine runs locally β your documents never touch a server.