What the free OCR page does
WhatPDF Online offers a one-page OCR preview. Pick a scanned PDF, choose the page you want to test, and run recognition. The recognised text appears in the browser so you can see whether the page, scan quality, and chosen language are producing usable results.
The free preview intentionally stops short of creating a full searchable PDF. It’s there to answer the practical question first: will OCR work well enough on this document? WhatPDF Full Edition takes it further with full-document OCR and rebuilds scanned pages with a searchable text layer.
Why scanned PDFs are not automatically searchable
A PDF can contain real text, page images, or a mix of both. A digitally created report usually has text objects that a PDF reader can select and search. A scanned page, on the other hand, is often just a photograph wrapped inside a PDF. It looks normal on screen, but there may be no actual characters underneath to select.
Use the preview page strategically
Don’t just grab the easiest page by default. If the document has different kinds of pages, pick something more representative or challenging, like a dense paragraph page, one with small text, or a scan that matches the overall quality of the rest of the document. A solid result on that page is way more useful than a perfect score on a large-font cover sheet.
Language codes in WhatPDF OCR
The underlying OCR engine uses Tesseract language data. The Complete workflow accepts common Tesseract language codes such as eng for English, fra for French, deu for German and spa for Spanish. The Online preview is intended as a quick evaluation of recognition quality rather than a full multilingual document-conversion service.
OCR is recognition, not proofreading
Even strong OCR should be checked when accuracy matters. A recognised invoice total, legal clause, serial number or personal name can be wrong by one character and still look plausible. Searchable text is extremely useful for finding information, but it should not automatically be treated as a verified transcription of the original scan.
Making a scanned PDF searchable
In WhatPDF Full Edition, OCR can be applied across selected or complete scanned pages so the rebuilt PDF contains a hidden searchable text layer while retaining the visible page image. That lets a PDF reader search for words in a document that originally had no selectable text. The Complete edition also includes scan cleanup options for scanner borders, small skew corrections, contrast adjustment and optional orientation detection.
Privacy and OCR language data
OCR is more computationally demanding than rearranging pages. WhatPDF performs recognition in the browser using Tesseract.js, and language data may be downloaded when required. That means the browser needs network access to obtain the recognition resources if they are not already available. The design is still local-first with respect to the document-processing workflow: the PDF is not intentionally sent to UsefulKits simply to have OCR performed on a server.
When OCR is the wrong tool
If you can already select and search the text in the PDF, running OCR may add little value. Likewise, OCR will not magically repair a scan where the original text is unreadable to the human eye. In those cases, finding a better source document or rescanning at higher quality is usually more effective.
Frequently Asked Questions
Can WhatPDF Online make my whole PDF searchable?
The free Online edition provides a one-page OCR preview. Full-document searchable-PDF OCR is part of WhatPDF Full Edition.
Why can I see text but not search it?
The page may be a scanned image rather than real PDF text. OCR is used to recognise the characters in that image.
Will OCR be perfectly accurate?
No OCR system is perfect. Scan quality, language, layout, font and image clarity all affect the result, so important text should be checked against the visible page.
Does OCR send my scanned PDF to UsefulKits?
WhatPDF is designed to perform OCR in the browser. OCR libraries and language data can be fetched from external sources, but the document is not intentionally uploaded to UsefulKits for server-side recognition.