OCR guide

Free OCR: extract text from images and scanned PDFs without uploading them

Pulling the text out of a photo, a screenshot or a scanned PDF is one of those tasks that seem simple until you need them: most PDF websites reserve OCR for their paid plans and, on top of that, upload your document to their server to process it. If that document is a contract, an invoice or a scan of an ID, uploading it to a third party just to "read" it is exactly what you wanted to avoid.

This guide explains, with concrete numbers, where the most popular tools put the wall, and presents an alternative that is free, in 9 languages and uploads nothing: it recognizes the text inside your browser.

OCR (optical character recognition) is almost always a "premium" feature. As of this guide:

  • SmallPDF: OCR is not on the free plan (which also caps you at 2 tasks a day); it is a paid feature (Pro plan, from about 12 USD a month).
  • iLovePDF: OCR is Premium; the free tier limits the number of tasks and the file size.
  • PDFCandy and similar: OCR with hourly limits or waits between uses.

And the most important thing they share: they all upload the file to their server to process it. For a scanned document —which, by definition, is usually something you want to digitize, not publish— that is precisely the friction you can do without.

The alternative: OCR in your browser, free and without uploading

Colibrí's Extract text (OCR) does the recognition inside your browser, with the open-source Tesseract engine compiled to WebAssembly. It is free, with no sign-up, no watermark and no limits, and your images and PDFs are never uploaded anywhere.

  • Images and scanned PDFs: drop a JPG, a PNG, a screenshot or a PDF and get the text. PDFs are processed page by page.
  • 9 languages: Spanish, English, Portuguese, French, German, Italian, Japanese, Chinese and Arabic. The language downloads only once (1–3 MB) and is saved in your browser for next time.
  • In batches: add several files and process them all at once.
  • Editable text: the result appears in a box you can correct before copying it or downloading it as .txt.

Honesty: what to expect from OCR

No OCR tool is perfect, and it is worth saying so. Accuracy depends heavily on image quality: with printed, sharp, straight text the result is usually very good; with blurry, skewed or low-contrast photos there will be errors. Handwriting is unreliable. That is why the result is editable: fix whatever is needed and you are done.

One important case: if your PDF already has selectable text (it is not a scan), you do not need OCR. Use PDF to text, which extracts the text directly, faster and with no recognition errors. OCR is for when the text is "inside" an image or a scan.

Step by step

  1. Open Extract text (OCR) and drag your images or PDFs (or click to pick them). They open in your browser; nothing is uploaded.
  2. Choose the language of the document. The first time, that language is downloaded (only once).
  3. Press Extract text and wait for the recognition to finish; you will see the progress of each file.
  4. Review and correct the text, and copy it or download it as .txt (individually or all together).

If you work a lot with documents, the rest of the toolbox is right there —compress PDF, protect PDF or the PDF editor—, and they all work the same way: in your browser, nothing uploaded.

Frequently asked questions

Is it really free and unlimited?

Yes. No sign-up, no watermark and no limit on files or size. There is no paid plan that unlocks OCR: it is available to everyone, in all 9 languages.

Is my document uploaded to any server?

No. Unlike SmallPDF, iLovePDF and most OCR websites, here the recognition happens in your browser and neither the file nor the extracted text leaves your device. It is the best guarantee for confidential scanned documents.

How accurate is it compared to the paid ones?

It uses Tesseract, the most widely used open-source OCR engine. With printed, sharp, straight text the results are very good; with low-quality images or handwriting there will be errors, as with any OCR. Since the text is editable, you fix the little that fails before copying it.

Which languages does it work in?

Nine: Spanish, English, Portuguese, French, German, Italian, Japanese, Chinese (Simplified) and Arabic. Each language downloads only once (1–3 MB) and stays saved in your browser for next time.

Does it work for a PDF that already has text?

If the PDF is not a scan and already lets you select the text, "PDF to text" is better: it extracts it directly without OCR (faster and exact). OCR is for when the text is "inside" an image or a scanned document.

Tools used in this guide

Everything in this guide happens inside your browser: your files are never uploaded to any server. That is the Colibrí promise.

More guides