PDF Editor
OCR guide

How OCR works

Digital, scanned and mixed pages, Russian/English Tesseract, confidence, manual review and searchable output.

When to use it

Digital, scanned and mixed pages, Russian/English Tesseract, confidence, manual review and searchable output. The situations below show when this guide can help you choose the right workflow and avoid unnecessary processing.

  • A PDF has no text search or selection.
  • A mixed document contains scans and existing digital text.

How it works

This guide divides the topic into three consecutive stages: page classification, recognition, review and export. Follow them in order to understand the result and the points that require manual review.

  1. 01Page classificationDigital, scanned and mixed pages are identified before OCR.
  2. 02RecognitionServer-side Tesseract processes only required pages in Russian and English.
  3. 03Review and exportWords below the confidence threshold are reviewed before a separate searchable PDF is created.

What changes

Before starting, identify which parts of the document will change and which will remain intact. The list below describes changes in the output copy rather than hidden modifications to the source file.

  • A searchable text layer is added to scanned pages.
  • Existing digital text is preserved instead of being recognized again.
  • User corrections apply only to confirmed OCR words.

What can be lost

Some operations affect more than the visible page and can change text layers, forms, links or signatures. This list helps you decide whether the separate output copy is suitable for the next stage of your work.

  • The raster page remains visible, but a corrected OCR layer can differ from visual source text.
  • When a patch is confirmed, the previous recognized word is replaced in the result copy.

Limitations

Limitations describe the current implementation boundary and the cases that require additional review. Check them before uploading, especially when a document is complex, protected or contains important structure.

  • Russian and English are supported.
  • Low scan quality, complex tables or handwriting can reduce confidence.
  • Manual review is required; error-free recognition is not promised.

Related tools let you move directly from the explanation to the appropriate action. Each tool uses its own workflow and shows its specific limits before processing starts.