Compression
How to reduce PDF size without losing quality
File size in a PDF is almost always image data. "Quality" is not one dial but a different thing for each kind of content, and there is a point past which compressing further only costs you something.
A 40-page scanned contract arrives at 58 MB. A 300-page report exported from a design tool arrives at 120 MB. Both are PDFs, both feel unreasonably large, and the two have completely different causes. Compressing the wrong thing is how people end up with a blurry contract that is still too big to email.
This guide is about the diagnosis as much as the fix: what makes a PDF large, what each technique in a compressor actually does, and how to decide where to stop.
Find out what is actually big
Open the PDF and try to select a single line of text with your cursor. That one test splits every oversized PDF into two categories, and the two need opposite treatment.
- Text selects normally — the file is vector content. The pages are drawing instructions (place this glyph here, draw this line there), not photographs. Files like this are usually already small: a 200-page text document is typically 1–5 MB. If yours is far larger, the weight is in embedded images or whole font families.
- Text will not select, or the cursor drags a rectangle over the whole page — the file is a scan. Every page is one large photograph of paper, and this is where 50 MB files come from. A single A4 page scanned in colour at 300 dpi holds roughly 26 MB of raw pixels before any compression, so a modest stack adds up quickly.
| Document | Normal size | What a much larger one means |
|---|---|---|
| 20 pages, word-processed text | 0.2–1 MB | Full-resolution images pasted in |
| 300 pages, text only | 2–6 MB | Whole font families embedded, or images |
| 10 pages, colour scans at 300 dpi | 20–60 MB | Expected — this is raw image data |
| 50 photos exported to PDF | 30–150 MB | Camera JPEGs kept at full resolution |
Text is almost free; images cost everything. A page of text is a few kilobytes, a page of photograph is a few megabytes. Nearly every "why is this PDF so big" question is really "how many images are inside it, and how large is each one".
The four things that make a PDF large
Once you know which category you are in, the causes are short and specific.
- Raster images stored at print resolution. This is the big one. Resolution is measured in dots per inch, and the pixel count scales with the square: an A4 page at 300 dpi is about 2,480 × 3,508 pixels, and dropping to 150 dpi quarters that number. Scanners default to 300 dpi because that is the right choice for archiving and for optical character recognition — but it is four times more data than a screen can display.
- Images stored without efficient compression. Some capture tools embed page images with lossless Flate compression, and a few embed them nearly raw. The identical pixels re-encoded as JPEG at quality 80 are routinely 10 to 20 times smaller. Nothing about the page content changed; only the storage format did.
- Whole font families embedded rather than subsetted. Properly made PDFs embed only the glyphs actually used, which costs tens of kilobytes. PDFs that embed every weight and style of a family can carry several megabytes of glyphs you will never render — a common result of "print to PDF" from design software.
- Duplicated and orphaned objects. Every time a PDF is edited and re-saved, most tools append a new version rather than rewriting the file. Deleted pages, replaced images and old metadata often remain in the file as unreachable objects. Embedded thumbnail previews add a small copy of every page on top of the real one.
What "quality" means depends on what is on the page
Compression damage is not uniform. The same setting that is invisible on a photograph can destroy a signature, a barcode or a line drawing. Before choosing a level, decide which of these your document actually contains.
| Content on the page | What lossy compression does to it | What to use |
|---|---|---|
| Photographs, gradients | Blocky patches in smooth areas, colour banding | JPEG is the right tool; quality 75–85 is visually clean |
| Body text, signatures, charts | Grey fuzz and ringing around black strokes | Lossless (Flate/PNG), or a higher resolution |
| Barcodes, QR codes, fine line art | Strokes merge; the code stops scanning | Lossless only — never re-encode these |
| Already-JPEG photos | A second lossy pass on top of the first | Downsample instead of re-encoding |
A useful way to hold this: JPEG throws away high-frequency detail. Text and line art are made almost entirely of high-frequency detail. Photographs have very little of it. That is why the same setting is safe on one page and ruinous on the next.
The compression ladder: least damaging first
Apply these in order. The first rungs are free, and doing them first often means you never need the last one.
- 1. Strip what nobody needs. Remove embedded thumbnails, authoring metadata, editor history and unreachable objects. Typically 1–5% — small, but lossless, and it never softens a page.
- 2. Downsample to what the destination actually needs. The single largest lever. Screen reading needs about 110–150 dpi; a phone or laptop display is itself only around 100–150 pixels per inch, so more input resolution is literally not visible. Office printing wants 200–300 dpi; anything that will be archived, OCRed or submitted as evidence should stay at 300–400 dpi.
- 3. Re-encode images to JPEG at a quality floor. Quality 75–85 is where photographs stop looking different; below about 70 the blocking starts to show, and text-heavy scans should not be sent down this path at all.
- 4. Rasterise the page. Render each page to one image and rebuild the PDF from those images. This is the last resort, and it is what most browser-based compressors do.
That last point deserves to be said plainly, because it decides which tool you should use. A browser cannot surgically rewrite the image objects inside an arbitrary PDF — there is no full object-level image re-encoder available client-side. What a browser can do reliably is render each page and re-encode the result. That is genuinely excellent for scans: a 300 dpi colour scan of ten pages routinely falls from tens of megabytes to two or three, and it looks the same on screen. It is the wrong tool for a text PDF, where rasterising can make the file larger while also destroying selectable, searchable text.
Desktop tools built on Ghostscript or qpdf attack the problem at the object level instead: they identify each embedded image, downsample and re-encode it, and leave the text layer untouched. That produces better results on mixed documents. The trade-off is that you have to install it, and it does not help you on a phone.
Choosing settings in one pass
Nine times out of ten the document has one destination, and the destination decides the settings. Pick the row that matches yours.
| The file is going to… | Resolution | Encoding | Notes |
|---|---|---|---|
| be read on a screen or emailed | 150 dpi | JPEG 80 | Visually indistinguishable; usually a 60–90% cut on scans |
| be printed as a contract or form | 300 dpi | Lossless preferred | Trim the margins first instead of lowering resolution |
| be archived, OCRed or filed as evidence | 300–400 dpi | Lossless | Size is not the enemy here — do not compress at all |
| stay as a text document | unchanged | unchanged | Skip compression; clean up objects only |
Measure, then look at 100% before you keep
Any compressor worth using reports the size before and after. Keep both copies, and open the smaller one at 100% zoom on the smallest thing in the document — a footnote, a date, a signature, a barcode. If it is still legible and the strokes have not grown a grey halo, you are done. If not, step the quality back up one notch, or drop the compression entirely and send a link instead.
The compress tool on this site reports both numbers so the comparison is not a guess, and it is honest about the method: it renders and re-encodes pages, which is why scans shrink so well and why text stops being selectable afterwards. If you need the text layer intact, do not compress the original — compress a copy and keep the first one untouched.
A one-line test for any compressor: does it tell you the output size before you commit? If a tool only shows you a download button, you have no way to know whether you just degraded the file for nothing.
Frequently asked questions
Will compressing a PDF reduce its quality?
Why did my PDF get bigger after I compressed it?
What dpi should I use to shrink a scanned PDF?
Does compression remove a password or encryption?
Can I compress only the heavy pages?
Is there a size limit on the compression?
Do it here, in your browser
Related guides
Every tool on this site processes files on your device. Read the no upload policy to verify that yourself in the Network tab.