textocr

Adds an invisible, searchable text layer to a PDF and re-exports it — the scanned pages themselves are untouched. A plain-text copy of what was read comes with every job.
Choose a file, or drop it here One PDF — or several page images, combined into one document
PDF, PNG, JPEG, TIFF, BMP or WebP.
Rasterising discards the original page image. Use “replace” to keep the scan intact. Surya ignores this and the language list — it detects script itself, and always writes a fresh layer.
Ctrl-click (⌘ on Mac) to pick several. Put the dominant language first — more languages is slower and slightly noisier.
Changes only what the engine sees, never the saved page.
Used only to decide between the two engines when they disagree. Pick the document’s language, or None — a wrong dictionary rejects every word and the arbiter stops deciding rather than deciding badly.
A spread OCRed whole gives each page half the resolution. Surya only.
Glyphless matches pre-v17 output, for the OCR-layer editor and the corrections applier.
DocumentWhen StatusPages TookChars Result
Nothing yet.