OCRmy PDF
Add OCR text layer to PDF files, making them searchable.
Open the official app on ocrmypdf.readthedocs.io
This tool is hosted by its maintainers. Click below to open ocrmypdf.readthedocs.io in a new tab — it's their official demo.
Browse pdf & tools tools →What's next with OCRmy PDF?
Choose how you want to get started.
Use it free
Open the official tool or demo — no account needed.
Self-host it
Run the open-source version on your own infrastructure.
What is OCRmy PDF?
OCRmyPDF is an open-source tool licensed under MPL-2.0 that adds an OCR (optical character recognition) text layer to scanned PDF files, enabling them to be searched and edited. Its primary purpose is to convert image-based PDFs into searchable, text-enabled documents while preserving formatting and quality. Users, including researchers, archivists, and businesses, face the challenge of unsearchable scanned documents, which hinder accessibility and data retrieval. OCRmyPDF solves this by leveraging OCR technology to extract text, making PDFs functional for digital workflows. The tool is particularly valuable for organizations needing to digitize paper records or comply with accessibility standards like Section 508.
How it works
OCRmyPDF is a command-line utility and web service wrapper designed to process PDFs by adding searchable text layers through OCR. It is maintained as an open-source project on GitHub, with over 34,400 stars, reflecting its widespread adoption. The tool is part of the broader ecosystem of document automation tools, offering a free alternative to commercial OCR solutions. Its core purpose is to transform scanned PDFs—created from physical documents into digital images—into searchable, editable formats. This capability is critical for users needing to archive, share, or analyze scanned documents without manual re-entry of text. OCRmyPDF supports advanced image processing, including noise reduction and layout analysis, to improve OCR accuracy. It integrates with Tesseract OCR for text recognition and allows customization of language packs for multilingual support. The tool also includes a JBIG2 encoder for compressing images without sacrificing quality, reducing file sizes while maintaining readability.
How to use it
- 1Install via pip: Run `pip install ocrmypdf` to download the latest version. 2. Prepare the PDF: Ensure the input file is a scanned image-based PDF, not a text-layered one. 3. Execute the command: Use `ocrmypdf input.pdf output.pdf` to process the file. 4. Specify options: Add parameters like `--language de` for German text or `--output-type pdfa` for PDF/A compliance. Practical tips include using Docker for cross-platform consistency, enabling language packs for non-English documents, and verifying system dependencies like Ghostscript for PDF rendering.
What it can do
- pdf ocr
Use cases
Assumptions and limitations
Assumptions
- source: https://github.com/ocrmypdf/OCRmyPDF
- license: MPL-2.0 — free to use
- privacy: Self-hosted — you control your data
Limitations
- Struggles with highly stylized or handwritten text
- Limited support for complex layouts (e.g., tables with merged cells)
- Requires manual intervention for PDF/A compliance
- OCR accuracy may degrade with low-resolution scans
- Lacks built-in PDF form field recognition
Understanding the result
Add OCR text layer to PDF files, making them searchable.
Tool details
- Clearly flagged when a network request is needed.
- No account, no sign-up, and no tracking of your content.
- Powered by (MPL-2.0).
- Built with
- (ocrmypdf/OCRmyPDF)
- License
- MPL-2.0
- Runs locally
- No — requires a network request
- Verification
- Not yet verified
- Input
- Query
- Output
- Text
Built with ocrmypdf/OCRmyPDF. OpenToolVault provides the discovery and browser interface while crediting the original project maintainers.
- Built with
- License
- MPL-2.0
Open-source project
OpenToolVault is an independent directory. We are not affiliated with or endorsed by this project.
References
- / — GitHub Repository
Upstream project · GitHub
- MPL-2.0 License
Upstream project
Frequently asked
How do I install OCRmyPDF on my system?
OCRmyPDF can be installed via pip by running `pip install ocrmypdf` in a terminal. For Linux/macOS users, ensure Ghostscript is installed first. Windows users may need to install additional dependencies like Poppler. Docker users can pull the official image with `docker pull ocrmypdf/ocrmypdf` and run it with `docker run -v /path/to/files:/data ocrmypdf/ocrmypdf /data/input.pdf /data/output.pdf`.
How does OCRmyPDF handle different languages?
OCRmyPDF uses the Tesseract OCR engine, which requires separate language packs for non-English text. These packs must be downloaded and configured via the `--language` flag. For example, to process a French document, run `ocrmypdf --language fra input.pdf output.pdf`. The tool supports over 100 languages, but accuracy depends on the quality of the language-specific training data.
How do I convert a scanned PDF to searchable text?
To convert a scanned PDF, run the command `ocrmypdf input_scanned.pdf output_searchable.pdf`. The tool automatically applies OCR to all pages, preserving original formatting. For best results, ensure the input PDF has clear, high-resolution scans. Use `--dpi 300` to increase scan resolution if needed. The output file will include both the original image and a searchable text layer.
How does OCRmyPDF compare to Adobe Acrobat OCR?
OCRmyPDF is open-source and free, while Adobe Acrobat offers proprietary features like PDF form field recognition and better handling of complex layouts. OCRmyPDF excels in multilingual support and customization via Tesseract, but lacks Acrobat’s advanced PDF editing tools. For users prioritizing cost and flexibility, OCRmyPDF is preferable; Adobe Acrobat is better for enterprise workflows requiring polished outputs.
What should I do if OCRmyPDF fails to recognize text?
If OCR errors occur, first verify the input PDF’s resolution and contrast. Use `--dpi 200` or higher for clearer scans. Check that the correct language pack is selected with `--language`. If issues persist, try re-scanning the document or using the `--force-ocr` flag to override existing text layers. For complex layouts, consider splitting the PDF into single-page files for processing.