Extract text from images and PDFs (OCR)

Extract text from images (JPG, PNG, WebP…) and scanned PDFs directly in your browser, free and with no usage limits. Colibrí uses OCR (optical character recognition) with the Tesseract engine to turn a photo, a screenshot or a scanned document into text you can copy, edit and download. It recognizes text in 9 languages — Spanish, English, Portuguese, French, German, Italian, Japanese, Chinese and Arabic — processes several files at once, and adds no watermarks, sign-ups or limits. Unlike websites that upload your document to a server and charge for OCR, here everything happens on your device: your images and PDFs never leave your browser.

Drop images or PDFs here

or click to choose them

JPG · PNG · WEBP · scanned PDF

How to extract text from an image or PDF

  1. Drag in your images or PDFs, or click to choose them (you can add several).
  2. Choose the language of the text you want to recognize.
  3. Click “Extract text”. The first time, the language is downloaded once.
  4. Wait for the recognition to finish; you’ll see the progress of each file.
  5. Review and fix the text if needed, then copy it or download it as .txt.

Frequently asked questions

Are my images or PDFs uploaded to any server?

No. All the OCR happens in your browser with WebAssembly (Tesseract engine); your files never leave your device and the tool keeps working offline once the language has been downloaded. It’s safe even with private or confidential documents.

What is OCR?

OCR (optical character recognition) is the technology that “reads” the text inside an image or a scanned document and turns it into real text you can select, copy and edit. It’s exactly what you need when you have a photo or a scanned PDF whose text you can’t copy.

Which languages does it recognize?

Nine: Spanish, English, Portuguese, French, German, Italian, Japanese, Chinese (Simplified) and Arabic. Pick the language that matches your document for the best result. Each language is downloaded just once (between 1 and 3 MB) and saved in your browser for next time.

Does it work with scanned PDFs?

Yes. If your PDF is a scan (pages that are really images whose text you can’t copy), the tool converts each page into an image and runs OCR on it. If your PDF already has selectable text, it’s faster and more accurate to use “PDF to text”, which extracts the text directly without OCR.

How accurate is it?

It depends a lot on the quality of the image. With printed text that is sharp, well lit and straight, the result is usually very good. With photos that are blurry, skewed, low in contrast or handwritten, there will be more errors. That’s why the text is editable: you can fix any mistake before copying it.

Does it recognize handwriting?

The engine is optimized for printed or typed text. Handwriting may work in very clear cases, but in general it isn’t reliable; for handwritten notes the result will be limited.

Is it free? Any watermarks or limits?

It’s completely free, with no watermark, no sign-up and no limit on usage or file size. Many PDF websites lock OCR behind their paid plans and upload your document to their server; here it’s free and unlimited, and nothing is uploaded.

Can I process several images or pages at once?

Yes. You can add several files and they’re all processed in the chosen language. In PDFs the text is recognized page by page. When it’s done you can download each text separately or all together in a single .txt file.

How do I get better recognition?

Use the highest resolution you can, with the text straight and high in contrast (dark on a light background). Crop to the area you care about and avoid shadows or glare. Choosing the correct language also improves accuracy quite a bit.

Related guides