What dpi to scan documents at
Scan at 300 dpi. That is the answer for anything with text on it that you may need to read again, search, send or keep — it is also the minimum written into the US federal rule for digitising paper records, so it is not a matter of taste. Drop to 200 for a copy you will throw away this afternoon, go up to 400 or 600 for fine print, engraved detail and photographs, and understand before you choose that every step up multiplies the file size while every step down is permanent.
OCR PDF →The short version, by what the scan is for
Resolution follows the job, not the machine's maximum. These are the settings worth using and roughly what one A4 page weighs afterwards.
| What you are scanning for | Resolution | Colour mode | One A4 page |
|---|---|---|---|
| A form to email today and forget | 200 dpi | grayscale | 100–250 KB |
| Ordinary paperwork you will keep | 300 dpi | grayscale or colour | 0.3–1.5 MB |
| Anything a machine will read | 300 dpi minimum | grayscale | 0.3–0.6 MB |
| Records held to a standard | 300 dpi, colour or grayscale, 8-bit | as specified | 0.5–1.5 MB |
| Fine print, seals, engraving, photographs | 400–600 dpi | colour | 2–8 MB |
| Reading on screen only | 150 dpi | grayscale | 60–150 KB |
One number to keep in your head: an A4 page at 300 dpi is 2480 × 3508 pixels, about 8.7 megapixels. That is the same order of magnitude as a phone camera, which is why a well-framed photograph of a page can stand in for a scan and a distant one cannot.
Anything below 200 dpi is a decision you cannot revisit. The detail is not compressed, it was never captured.
Where 300 comes from
The number is not folklore. In the United States, 36 CFR § 1236.50 — the rule governing digitisation of permanent federal records, in force since 5 June 2023 — requires image files at a minimum of 300 ppi «sized to the source document» for modern textual paper records, with a tolerance that allows ≥ 294 ppi, and a minimum of 400 ppi for photographic prints and paper records with fine details, tolerance ≥ 392 ppi. Colour or grayscale, 8 or 16 bits per channel.
«Sized to the source document» is the part people skip and then get wrong. Resolution only means something in relation to the physical page. A 3000-pixel-wide image of an A4 sheet is 363 dpi; the same 3000 pixels spread along the long side of an A0 drawing is 64 dpi and useless. If your scanner asks for a document size, tell it the truth, or the number it reports is fiction.
You are probably not filing with the National Archives. The reason to borrow their figure anyway is that it was chosen by people who had to make the scan usable decades later, for text of the kind offices actually produce — typeset, typed or laser-printed, with reasonable contrast between ink and paper. Your bank statements and lease agreements fall in exactly that category.
If software is going to read it
Text recognition is where resolution stops being an aesthetic question. The Tesseract engine — the same one behind a great many free recognition tools, including the one on this site — states it plainly in its own documentation: it «works best on images which have a DPI of at least 300 dpi». Below that, letters occupy too few pixels for the shapes to be distinguishable, and the engine starts inventing characters rather than reading them.
The practical threshold is the height of a lowercase letter. Around 20 pixels is the floor; a scan at 300 dpi gives an 8-point footnote roughly that, and 10-point body text comfortably more. Scan the same footnote at 150 dpi and there is nothing to work with.
Two honest notes about how our own OCR behaves, because they change what you should feed it. When you hand it a PDF, it draws each page at twice the PDF's internal grid — about 144 dpi of effective detail — which is fine for normal body text and thin for small print. When you hand it an image file, the picture goes to the recognizer at full resolution. So for tiny type, scan to JPEG or PNG and recognise the image directly rather than wrapping it in a PDF first. And what comes back is plain text you can copy and correct, not a searchable layer inside the original document.
Higher is not automatically better here either. Past roughly 400 dpi the accuracy curve flattens, and letters can become large enough to fall outside the range the recognition model was trained on. 300 to 400 is the useful window.
What each step up costs
Resolution is squared, which is why the jump feels sudden. Doubling from 300 to 600 dpi does not double the file — it quadruples the pixels.
The arithmetic for one A4 page at 300 dpi, 8.7 megapixels: 26 MB as uncompressed 24-bit colour, 8.7 MB as 8-bit grayscale, 1.1 MB as pure black and white. Compression brings those down to something usable — typically half a megabyte to a megabyte and a half for a colour page, three hundred to six hundred kilobytes in grayscale, and mere tens of kilobytes for bitonal text; one scanning vendor measures a standard office page at about 41 KB at 200 dpi and 62 KB at 300.
Now multiply. A twenty-page colour contract at 300 dpi lands around 10 to 30 MB, right at the twenty-five megabyte limit most mail gateways still enforce. The same contract at 600 dpi is 40 to 120 MB and will not travel by email at all. That, rather than image quality, is the reason to think about the setting before you press scan instead of afterwards.
Colour mode moves the number as much as resolution does, and it has its own trap. Bitonal — pure black and white — is wonderfully small and destroys a blue signature, a red stamp, a highlighted line and any faint pencil annotation, because every pixel has to choose a side. Grayscale at 300 dpi is the sensible default for paperwork. Reserve colour for pages where colour carries meaning, and expect to pay two or three times the size for it.
Down is easy, up is impossible
Scanning high and reducing later is the safe order, because reduction throws away detail that exists while enlargement invents detail that does not. A 150 dpi scan upsampled to 300 dpi is a 300 dpi file containing 150 dpi of information, and every recognition engine sees through it immediately.
So keep the original capture and make derivatives from it. If a portal demands a two-megabyte file, run the good scan through Compress PDF and send the result, keeping the original for the archive.
Know what that compression does, though. It re-draws every page as an image at a fixed resolution — 96 dpi on the strongest setting, 120 in the middle, 150 on the lightest, with JPEG quality falling in step. A 300 dpi scan sent through the strongest setting comes back as a 96 dpi document: perfectly readable on screen, no longer suitable for recognition or for archiving, and impossible to undo. Compress the copy, never the master.
When you need a page as an image rather than a document — for a report, a slide, a claim form that only accepts pictures — PDF to PNG exports at 72, 150 or 300 dpi without any lossy step, and PDF to JPG does the same with an adjustable quality setting. Choose 300 if the result may be printed, 150 if it will only be looked at.
Phone photographs and the same arithmetic
A camera has no dpi setting, which makes people assume the question does not apply. It applies exactly as much; it is just expressed in how much of the frame the page fills.
Work it through. A 12-megapixel photograph is around 4000 × 3000 pixels. If an A4 page fills the frame along its long side, those 4000 pixels cover 297 mm — 11.7 inches — which is about 340 dpi, comfortably past the threshold. Let the page occupy a third of the picture and you are at 113 dpi, below anything usable, with the rest of the resolution spent photographing your desk.
That single calculation explains why phone scans sometimes sail through a portal's checks and sometimes fail: not the camera, the framing. Fill the frame, hold the phone parallel to the page, and the file that comes out is genuinely comparable to a flatbed scan — the rest of the technique is in scanning documents with your phone.
Two last connections worth making. If a scan is unexpectedly heavy, resolution is only one of the possible causes and why is my PDF so large works through the others. And if you are scanning a stack with printing on both faces, the resolution decision is the second one to make — the first is getting the faces into one file in the right order, which is covered in scanning a two-sided document into one PDF.
Common questions
Is 600 dpi better than 300 for ordinary documents?
Not usefully. It quadruples the file size and adds detail that printed text does not contain. Save 600 dpi for small print, engraved or embossed detail, photographs and anything you intend to enlarge.
Can I scan at 150 dpi to keep the files small?
For a page you will read once on screen and delete, yes. For anything you might need to search, enlarge or print, no — 150 dpi is below what recognition engines can work with, and the detail cannot be added back later. Compressing a 300 dpi scan is the better way to get a small file.
Can I increase the dpi of a scan I already have?
You can resample it to a larger pixel count, but no detail is recovered — the file becomes bigger without becoming better. If the original paper is still available, rescanning is the only real fix.
Should I scan in colour or black and white?
Grayscale at 300 dpi is the sane default for paperwork. Use colour when colour carries information: a stamp, a seal, a signature in blue ink, a highlighted clause. Pure black and white is only for clean printed text where file size genuinely matters, because it erases anything faint.
Does higher resolution always improve text recognition?
Only up to a point. Tesseract asks for at least 300 dpi and gains little above about 400; very large letters can even fall outside the range its model expects. Even lighting and a straight, flat page improve accuracy far more than extra resolution does.
Is my scan uploaded when I use the tools here?
No. Recognition, compression and image export all run in the page on your own device, so a payslip or a medical letter is never transmitted. It also means the ceiling is your machine's memory rather than a server quota.