What Is OCR? How Scanned PDFs Become Searchable Text
Every day, millions of documents are scanned into PDFs — contracts, application forms, receipts, old archives. They look like documents. They open like documents. But try pressing Ctrl+F and searching for a word inside one, and nothing happens. Try selecting a paragraph to copy it into an email, and you can't. The reason is straightforward: your computer doesn't see text on the page. It sees a photograph of text. OCR — Optical Character Recognition — is the technology that bridges that gap, turning pictures of words back into real, selectable, searchable characters.
This article explains what OCR is, how it actually works under the hood, what determines its accuracy, and how you can use it to make your own scanned PDFs searchable — entirely in your browser, without uploading anything.
The Problem: Why Scanned PDFs Aren't Searchable
To understand OCR, you first need to understand that there are two fundamentally different kinds of PDF.
A born-digital PDF — one exported from Word, Google Docs, or a design program — contains an actual text layer. Every letter is stored as a Unicode character with a precise position on the page. When you select text, search with Ctrl+F, or use a screen reader, the software reads this text layer. The document "knows" what its words are.
A scanned PDF contains no text layer at all. A scanner (or a phone camera) captures the page as a grid of pixels — a raster image — and that image is placed on the PDF page. To your computer, the page is no different from a photograph of a landscape. There are no characters, no words, no sentences — only colored dots arranged in patterns that happen to look like letters to a human eye.
This creates real, practical problems:
- No text search. You can't find a clause in a 200-page scanned contract with Ctrl+F. You have to read it page by page.
- No copy and paste. Quoting a paragraph means retyping it by hand, introducing the risk of transcription errors.
- No accessibility. Screen readers used by visually impaired people cannot read image-only pages — the document is effectively blank to them.
- No indexing. Search engines, document management systems, and e-discovery tools can't index the content, so the document is invisible to search.
- Bloated file sizes. A page of text stored as an image is typically 10–50 times larger than the same text stored as characters.
OCR solves all of this by analyzing the image, recognizing the characters it contains, and writing them back into the PDF as a text layer — producing what is called a searchable PDF. The page looks identical, but underneath the image there is now real text you can select, search, and copy.
How OCR Works: From Pixels to Characters
Modern OCR is a multi-stage pipeline. Each stage solves a specific part of the problem, and the quality of the final result depends on every one of them.
Step 1: Image Preprocessing
Before any recognition happens, the raw scan is cleaned up. Scans are rarely perfect: pages are skewed, lighting is uneven, paper has speckles, and phone photos have perspective distortion. Preprocessing corrects these issues:
- Binarization converts the image to pure black and white, separating text from background. Adaptive thresholding handles pages where lighting varies across the scan.
- Deskewing detects the dominant angle of text lines and rotates the image so lines are horizontal. Even a 1–2 degree skew measurably reduces accuracy.
- Noise removal eliminates isolated dark pixels (dust, paper grain) that could be misread as punctuation.
- Contrast normalization evens out shadows and bright spots so characters have consistent edges.
This stage is invisible to the user, but it is responsible for a large share of OCR quality. A well-preprocessed mediocre scan often recognizes better than a raw high-resolution one.
Step 2: Layout Analysis
Next, the engine figures out the structure of the page: where are the blocks of text, where are images, tables, headers, footers, and page numbers? It segments the page into regions, then splits text regions into lines, and lines into individual words and characters. This matters because reading order must be reconstructed — a two-column academic paper has to be read down the left column first, then the right, not across the page.
Step 3: Character Recognition
This is the core of OCR. Each segmented character image is classified — "this shape is an 'a', this one is a '7'." Modern engines use two complementary approaches:
- Feature extraction and pattern matching. The engine measures geometric features of the character — strokes, loops, endpoints, aspect ratio — and compares them against trained models of each glyph.
- Recurrent neural networks (LSTM). State-of-the-art engines like Tesseract 4+ use Long Short-Term Memory networks that don't just look at one character in isolation but consider the sequence of characters around it. This is why modern OCR can correctly read a smudged "rn" as "m" in context, or distinguish "0" from "O" based on surrounding letters.
The neural-network approach was the single biggest leap in OCR accuracy in the last decade, cutting error rates roughly in half on difficult documents.
Step 4: Post-Processing and Language Correction
Raw recognition output still contains errors, so engines apply linguistic knowledge:
- Dictionaries flag character sequences that aren't valid words and suggest corrections.
- Language models use the probability of word sequences — "the quick brown fox" is far more likely than "the quiek brown fox" — to fix misread characters.
- Confidence scoring marks uncertain characters so downstream tools (or humans) know what to double-check.
Step 5: Building the Searchable PDF
Finally, the recognized text is written back into the PDF as an invisible text layer positioned exactly over the original image. The page looks pixel-identical to the scan, but now every word exists as real text underneath. You can select it, copy it, search it — and screen readers can read it. This is the file format most archives, courts, and government agencies require for scanned submissions.
Tesseract and Browser-Based OCR: How GroPDF Does It Locally
The OCR engine GroPDF uses is Tesseract, the most widely deployed open-source OCR engine in the world. Originally developed by Hewlett-Packard in the 1980s and later sponsored by Google, Tesseract has been refined over decades and supports over 100 languages. It is the engine behind countless document-processing pipelines, from Google Books digitization to enterprise archiving systems.
Traditionally, running Tesseract required installing software on a computer or sending documents to a server. GroPDF takes a different approach: the entire engine runs inside your web browser using a technology called WebAssembly.
WebAssembly (Wasm) is a binary instruction format that lets browsers execute code written in languages like C++ at near-native speed. Tesseract is written in C++; compiled to WebAssembly, it runs in your browser tab as fast as a desktop application, with no server involved. The language data files are downloaded once to your device, and from that point the recognition happens entirely on your machine.
The practical consequences are significant:
- Privacy. Your scanned documents never leave your device. There is no upload, no server copy, nothing to intercept or breach. For contracts, medical records, or ID documents, this is the safest possible architecture.
- No file size limits from servers. Processing happens on your hardware, so you're not constrained by a service's upload quotas.
- Works offline. Once the page and engine files are loaded, you can disconnect from the internet and OCR still works.
- Free. There is no per-page API cost because there is no API — it's your CPU doing the work.
The trade-off is a one-time download of the OCR engine (roughly 25MB including the English language data), which is why GroPDF loads it only when you open the OCR tool rather than on every page.
Accuracy Factors: What Determines OCR Quality
OCR is remarkably good on clean documents and remarkably bad on poor ones. Understanding what affects accuracy helps you get usable results and set realistic expectations.
Scan Resolution
Resolution is the single biggest factor. The widely cited minimum for reliable OCR is 300 DPI (dots per inch). At 300 DPI, a 10-point character is about 40 pixels tall — enough detail for the engine to distinguish fine features. Below 200 DPI, accuracy drops sharply because characters become ambiguous blobs. Above 400 DPI, gains are marginal and file sizes balloon, so 300 DPI is the sweet spot for scanning.
Contrast and Lighting
OCR needs clear separation between text and background. Faded print, yellowed paper, coffee stains, and shadows from phone photos all degrade results. Black text on white paper is ideal. If you're photographing documents with a phone, even lighting matters more than camera megapixels.
Fonts and Typography
- Standard fonts (Arial, Times New Roman, Helvetica) recognize at 98–99% accuracy on clean scans.
- Decorative or unusual fonts — script typefaces, heavy condensed fonts, dot-matrix print — can drop accuracy significantly because the engine's training data covers them less well.
- Very small text (below 8 points) and ALL CAPS headings are harder; capital letters have fewer distinguishing features than mixed-case text.
Language
OCR engines are trained per language. English recognition is the most mature; accuracy for other languages depends on the quality of their training data. Mixed-language documents (for example, an English contract with Turkish names, or Urdu text alongside English) are harder because the engine must switch language models mid-page. GroPDF currently supports English, with additional languages planned.
Handwriting
This deserves blunt honesty: standard OCR does not reliably read handwriting. Cursive script, inconsistent letterforms, and connected characters defeat the segmentation step that printed-text OCR depends on. Handwriting recognition (HTR) is a separate, harder problem requiring specialized models. If your document contains handwritten annotations alongside printed text, expect the printed portions to recognize well and the handwriting to come out as gibberish or be skipped. Do not rely on OCR for handwritten medical notes, filled-in forms with cursive answers, or historical manuscripts.
Layout Complexity
Multi-column layouts, tables, text wrapped around images, and forms with boxes and lines all challenge the layout-analysis stage. A simple single-column letter recognizes almost perfectly; a newspaper page or a complex tax form will have more errors, particularly in reading order — the words may all be correct but assembled in the wrong sequence.
As a rule of thumb: on a clean 300-DPI scan of standard printed English text, expect 98%+ character accuracy. On a phone photo of a crumpled receipt, expect to proofread carefully.
Try It: Convert Your Scan to a Searchable PDF
The fastest way to understand OCR is to run it on one of your own documents. GroPDF's OCR PDF tool processes your file entirely in your browser:
- Upload your scanned PDF.
- Select the recognition language.
- Click Start OCR and wait while each page is analyzed.
- Download your searchable PDF — it looks identical to the scan, but the text is now selectable and searchable.
If you only need the raw text rather than a searchable PDF — for pasting into a document, translating, or analyzing — use the PDF to Text tool, which extracts the text layer (from OCR-processed or born-digital PDFs) into a plain .txt file.
One practical tip: if your scan is a phone photo, try to capture it as flat and evenly lit as possible before uploading. Five seconds of care with the camera saves five minutes of correcting OCR errors afterward.