The Scanned Contract Problem: How to Make a Photographed Agreement Searchable
Someone signs a contract, photographs it with their phone, and emails you the PDF. It looks like a document. Open it, press Ctrl+F, search for “termination”, and nothing is found — because there is no text in the file at all. There is a picture of text, and pictures are not searchable.
This matters more than it sounds. A scanned contract cannot be searched, cannot be quoted without retyping, cannot be analysed by anything, and cannot be read by a screen reader. It is, in every practical sense, a photograph you are storing in a legal file.
Telling a scan from a document
Open the PDF and try to select a sentence with your cursor.
- A blue selection appears over the words — there is a text layer. Nothing to do.
- A blue box appears over the whole page, or nothing happens — it is a scan.
A file can be both: a contract exported from Word with a photographed signature page appended is text for nine pages and a picture for one. That last page is usually the one with the signatures and the date on it.
Fixing it
Make PDF searchable — upload the scan, get the same PDF back with an invisible text layer added underneath the image. The pages look identical; the difference is that Ctrl+F now works, text selects and copies, and anything downstream can read it.
This is the version to keep. It preserves the original appearance — including the signatures — while making the content reachable.
Extract the text if you want the words rather than the document, and image to text if what you have is a JPG or PNG rather than a PDF.
Up to 30 pages per document, no account.
What OCR gets wrong
Text recognition is very good and not perfect, and the errors cluster in the places that matter in a contract:
Numbers. 1 and l, 0 and O, 5 and S. A payment of $1,500 read as
$l,500 is not searchable and not obviously wrong at a glance.
Handwriting. Dates written by hand in a blank, initials in a margin, anything added after printing. Treat handwritten fields as unread.
Bad scans. A photograph taken at an angle, in poor light, or of a photocopy-of-a-photocopy. Every generation of copying costs accuracy.
Stamps and signatures over text. Where ink overlaps a printed word, one of them wins and it is usually the ink.
The practical rule: OCR is reliable enough to search and to analyse, and not reliable enough to quote from without checking the original page. Before you paste a figure into an email, look at the scan.
Getting a better scan in the first place
Most OCR failures are input failures, and the fix costs nothing:
- Flat surface, no shadow across the page, phone parallel to the paper rather than at an angle.
- Scan at 300 DPI if you have a scanner. Below 200 the error rate climbs sharply; above 400 you get a larger file and no more accuracy.
- Black and white text on white paper — colour adds size and nothing else.
- One page per image. A two-page spread photographed together halves the effective resolution of each page.
If the file is now too large to send, compress it afterwards rather than scanning at a lower quality.
Then read it
Once the text layer is there, the contract is a contract again. You can upload it for a free review — six scores out of ten with the reasoning for each, and the number of concrete weaknesses found — the same as for any other document.
That is the actual point of the exercise. A scan sitting in a folder is a liability nobody has read; a searchable one is a contract you can check.
Next step: make your scan searchable → — free, no account. Then see what it scores.
Ready to simplify your legal document review?
Start using LegalValidate.ai to instantly analyze, validate, and improve your contracts and agreements.
Get Started for Free No credit card required. Try it now!