UseToolSuite UseToolSuite

Optimize PDF — Dark Mode, Compress, Grayscale, OCR & Repair

Five fixes for an existing PDF: convert it to dark mode for night reading, compress it, turn it grayscale for printing, make a scan searchable with OCR, or repair a file that will not open. All in your browser, nothing uploaded.

Dark mode in 3 themes Compress with DPI & JPEG quality control On-device OCR in 8 languages qpdf structure repair in WebAssembly Nothing uploaded

Recolor a bright PDF for night reading in True Black, Midnight Blue or Sepia Dark. Starts as soon as you drop the file.

Drop PDF file here or click to select

Smart inversion preserves images while darkening text and backgrounds.

Theme Variant

This page collects five things you do to a PDF you already have, rather than to its pages. Dark mode recolors it for reading at night. Compress makes it small enough to email. Grayscale makes it print cheaply and predictably. OCR makes a scan searchable. Repair rebuilds a file that won’t open. Every mode runs in your browser tab — the contracts, statements and records people usually run through these tools are never uploaded.

One distinction matters across all five. Dark mode, Compress and Grayscale render each page to an image and rebuild the PDF from those images, so the output looks right but its text is no longer selectable. OCR and Repair keep the original page content. Keep your original whenever you convert.

Dark mode: when dark documents actually help

Dark mode is about comfort and context, not universal superiority. In a dark room, a white PDF page is often the brightest object in view — pupils adjust to it and everything else vanishes; a dark page removes that mismatch, which is why night reading is dark mode’s strongest case. On OLED screens, dark pages also measurably extend battery life since black pixels are unlit. In bright daylight the advantage reverses: dark text on light backgrounds reads faster for most people, and astigmatism makes light-on-dark text harder for a meaningful minority.

Three themes are available: True Black, Midnight Blue and Sepia Dark. Gray and black content is recolored; colored pixels — photos, logos, colored charts — are kept and only dimmed slightly, so pictures don’t turn into negatives on digitally made PDFs.

Source PDFDark-mode result
Exported from Word/Docs/LaTeXExcellent — clean text recoloring
Code documentation, ebooksExcellent — long-form reading is the core use case
Slides with dark themes alreadySkip — double inversion helps nothing
Charts and data visualizationsMixed — colors can shift meaning; check legends
Scans and photographed pagesPoor — the whole page is one image, so paper texture inverts too

Compress: match the setting to where the file is going

Hitting a “file too large” wall when emailing or uploading a PDF is almost always an image problem. Phone scans are captured at far higher resolution than a screen or a printed page can show, so a five-page scanned contract can weigh 20 MB while a fifty-page text report weighs less than one. If your file is scan- or photo-heavy, there’s a lot to recover; if it’s mostly text, it’s already near its floor — and re-rendering it as images can even make it larger.

The compressor renders every page at the DPI you choose and re-encodes it as a JPEG at the quality you choose. The four presets run from 150 DPI at 85% quality down to 72 DPI at 35%.

  • Email and web upload — a middle preset usually clears the common 10–25 MB attachment limit while keeping the document legible on screen.
  • Print — stay at 150 DPI and high quality; lower resolutions show up as soft edges on paper.
  • Portals with a size cap — compress to comfortably under the limit so a slightly larger re-export still fits.

Grayscale: where it actually pays off

On metered office printers and print-shop pricing, a color page typically costs several times a black-and-white one, and a document with a single colored logo on each page can be billed entirely at the color rate. Converting the whole file to grayscale before printing guarantees every page meters as mono, and it prints identically on every device instead of depending on each printer driver’s color conversion.

Naive color removal averages the red, green and blue channels equally, which renders yellows too dark and blues too light and can collapse adjacent chart colors into the same gray. The Rec. 601 transform used here (0.299 R + 0.587 G + 0.114 B) weights channels by perceived brightness, so dark text stays dark, highlights stay light, and a pie chart’s slices remain distinguishable. A before/after preview of page 1 shows the result before you convert; choose 200 DPI for dense, small print.

OCR: why the text layer is invisible

A searchable scan has two layers: the original page image you see, and machine-readable text positioned at the exact coordinates of each printed word. The text is drawn with zero opacity, so a pixel-perfect scan of a signed contract stays pixel-perfect — but viewers search, select and copy against the hidden layer. Replacing the image with recognized text would be destructive, because OCR is never 100% accurate; with the invisible layer, recognition errors only affect search quality, never the document.

Recognition runs on your device with Tesseract compiled to WebAssembly. The only thing downloaded is the public language model (about 15 MB, cached after the first run); nothing about your document goes the other way. For best accuracy, scan at 300 DPI, keep the page straight, and pick the document’s language — each model is language-specific. For a single photographed page rather than a PDF, the Image OCR tool does the same job; to get plain text out of an already-digital PDF, the PDF Converter’s text output is instant.

Repair: what it can and can’t do

Usually the content of a PDF that “is damaged and could not be repaired” is fine; what’s broken is its internal scaffolding. A PDF keeps a cross-reference table that tells viewers where each object lives, and an interrupted download or save can leave that index incomplete or pointing to the wrong places. Repair runs qpdf — a long-established open-source PDF library, compiled to WebAssembly — to re-scan the file for valid objects and rebuild those references, falling back to a pdf-lib rewrite if qpdf can’t process the file. A live log shows each step.

  • Often recoverable — broken cross-reference tables, interrupted saves, minor structural damage where the page data survived.
  • Sometimes partial — a few damaged pages may drop while the rest are recovered.
  • Not recoverable — files truncated to a fraction of their size or overwritten; missing data can’t be reconstructed.

Encrypted files can’t be repaired until the password is removed, since encryption hides the structure the repair reads — use the Unlock mode of PDF Security first.

A sensible order when you need more than one

Repair first if the file won’t open. Run OCR on the original before anything that rasterizes pages, because Compress, Grayscale and Dark mode turn text into images. Compress last, once the content is final.

Optimize PDF — Dark Mode, Compress, Grayscale, OCR & Repair Powered by UseToolSuite — free browser tools

Last updated Built and maintained by Necmeddin Cunedioglu How tools are tested

How to Use This Tool

  1. 1

    Choose what to fix

    Pick Dark mode, Compress, Grayscale, OCR or Repair from the tabs above the drop zone.

  2. 2

    Load the PDF and set the options

    Each mode has its own settings: a theme, a DPI and quality level, a print resolution, a recognition language, or repair flags.

  3. 3

    Download the new copy

    The original file on your device is never changed. Dark mode and Repair start as soon as you drop the file; the others start from their button.

How helpful was this tool?

Click to rate

Key Concepts

DPI (Dots Per Inch)

A measure of image resolution indicating how many pixels are packed per inch. For screen display, 72–96 DPI is standard. For professional printing, 300 DPI is the industry standard. Reducing the DPI during PDF compression reduces the number of pixels per page, directly reducing file size. A 150 DPI page has 56% fewer pixels than a 200 DPI page of the same dimensions.

JPEG Quality Factor

A value (typically 1–100) controlling the aggressiveness of JPEG lossy compression. Higher values preserve more image data and produce larger files; lower values discard more data for smaller files. JPEG compression works by approximating image blocks using frequency components (DCT), discarding high-frequency details that the eye is less sensitive to. Quality 85 is generally imperceptible to human vision; below 50 produces visible compression artifacts.

Relative luminance (Rec. 601)

A weighted formula (0.299 R + 0.587 G + 0.114 B) that converts a color to the gray level the human eye perceives as equally bright. Green contributes most because the eye is most sensitive to it — a naive equal average makes yellows too dark and blues too light.

Colour inversion

Flipping each colour to its opposite — white becomes near-black, black becomes white — which is what turns a bright page into a dark one.

qpdf

A well-established open-source library for inspecting and repairing PDF structure. This tool runs a browser build of it, so nothing is uploaded.

Cross-reference index

The table at the end of a PDF that maps where each object lives. When it is damaged, readers cannot find the content even though it is still there.

Invisible text layer

Recognized words drawn at the exact position of each printed word but with zero opacity. The page looks unchanged, while viewers can search, select and copy against the hidden text — the same structure professional OCR software produces.

Frequently Asked Questions

How does browser-based PDF compression work?

The tool loads your PDF using PDF.js, renders each page to an HTML Canvas element at the target resolution and quality, then reconstructs the pages into a new PDF using jsPDF. The main compression gain comes from re-encoding page content as JPEG at a lower quality factor and reducing the render resolution (DPI). This approach works best for PDFs containing images or scanned documents.

What level of compression can I expect?

Compression results depend heavily on the PDF content. Image-heavy PDFs and scanned documents typically achieve 60–85% size reduction at Medium quality. PDFs containing mostly text and vector graphics may achieve 20–40% reduction. High compression with very aggressive settings can produce 90%+ reduction but at reduced visual quality. Use the quality slider to find the best size/quality balance for your use case.

Why is my compressed PDF larger than the original?

This can happen with PDFs that are already well-optimized, contain mostly vector text and graphics (no raster images), or were created with very efficient compression. In these cases, rasterizing and re-encoding adds overhead rather than reducing it. Very small PDFs (under 100 KB) often cannot be compressed further. For vector-only PDFs, dedicated PDF optimization tools like Ghostscript may perform better than image-based compression.

Why convert a PDF to grayscale instead of letting the printer do it?

Three reasons: color pages often cost several times more than black-and-white on metered office printers, printer drivers vary wildly in how they convert color (some make blue text nearly invisible), and a grayscale file guarantees the same output on every printer. Converting once, with a proper luminance transform, gives you a predictable, cheaper print everywhere.

Will colored text and charts stay readable?

Yes — the conversion uses the Rec. 601 perceptual luminance formula (30% red, 59% green, 11% blue) rather than a naive average. That preserves the relative brightness your eye perceives, so dark blue text stays dark, yellow highlights stay light, and adjacent chart colors keep distinguishable gray levels.

Is the dark background pure black?

It depends on the theme you pick. True Black turns white paper into pure black (#000000) with white text, which looks sharpest on OLED screens. Midnight Blue uses a deep navy (#0A192F) and Sepia Dark a warm brown-black (#2B2620), both with slightly off-white text that many people find gentler for long reading sessions.

Why did the photos come out looking like negatives?

If the PDF is a single flattened image of the whole page — text and pictures baked together — there is no way to invert the text but spare the photos. Inversion has to apply to the one image. PDFs that keep text and images as separate layers handle this much better.

How can a broken PDF be fixed at all?

A PDF keeps an index at the end that tells readers where everything is. If that index is damaged the file "won't open" even though the actual content is fine. This tool scans the file top to bottom, finds the intact pieces, and builds a fresh index — which is enough to recover most partly-damaged files.

How is this different from other online OCR tools?

The recognition itself runs in your browser using Tesseract compiled to WebAssembly — your document is never uploaded to a server. Practically every other online OCR service sends the file to their backend for processing. Since scanned PDFs are often contracts, medical records, or IDs, keeping recognition on-device removes the risk of the document sitting on someone else's infrastructure.

Does OCR change how my PDF looks?

No. The original pages are copied into the output untouched — they are not re-rendered or recompressed. The recognized words are added as an invisible (fully transparent) text layer positioned over the printed words, which is the same technique professional OCR software uses. Visually the file is identical; functionally, Ctrl+F and text selection now work.

Which of these modes keep the text selectable?

Repair and OCR do. Repair rewrites the file structure and leaves the page content as it was, and OCR copies the original pages untouched and adds an invisible text layer, so a scan becomes searchable. Compress, Grayscale and Dark mode work by rendering each page to an image and rebuilding the PDF from those images, so the output looks right but its text can no longer be selected or searched. If you need both, keep the original for search and use the converted copy for reading or printing.

Should I print a dark-mode PDF?

No — keep the original for printing. A dark page burns through ink or toner at enormous rates, and print contrast on paper works opposite to screens: dark text on white paper is strictly more readable. The right workflow is two artifacts: the dark version for screens, the original for paper.

What actually makes a PDF file large?

In most documents the bulk is images — high-resolution scans and embedded photos — not text, which is tiny. A text-only PDF is already small and won't shrink much; a scan-heavy PDF can often drop 50–90% when its pages are re-encoded at a lower resolution and JPEG quality.

Why does blue text sometimes disappear when a printer converts to black and white?

Because many printer drivers convert color with a crude average or a single channel instead of perceptual luminance. Pure blue carries very little perceived brightness (about 11% weight in the Rec. 601 formula), but a naive average treats it as one-third — so dark blue text that should print near-black comes out a washed-out mid-gray, and on some drivers nearly vanishes. Converting the PDF yourself with a proper luminance transform locks in correct gray levels before the driver ever gets involved.

How do I check whether a PDF already has a text layer before running OCR?

Open the PDF in any viewer and try to select a few words, or press Ctrl+F and search for a word you can see on the page. If selection works and the search finds it, the file already has a digital text layer and doesn't need OCR. If the cursor drags a rectangle instead of selecting words and search finds nothing, it's a pure scan and OCR is exactly what it needs.

What causes a PDF to become corrupted in the first place?

Most corruption comes from an interrupted write: a download that dropped, a file copied off a USB stick that was pulled early, a transfer cut by a lost connection, or an app that crashed mid-save. The result is a structurally broken file that viewers refuse to open. An interrupted transfer is often fixable; a file truncated to a fraction of its size may be beyond repair.

Troubleshooting & Technical Tips

PDF fails to load or shows blank pages

Ensure the file is a valid PDF (not a renamed image or document). Password-protected PDFs cannot be loaded — remove the password first with the Unlock mode of PDF Security. Very large PDFs (100+ pages or 100+ MB) may exceed browser memory limits; split the PDF into smaller parts first.

Text is blurry in the compressed PDF

Increase the quality slider to 80% or higher. For documents where text clarity is critical, use the "High quality" preset which renders at 150 DPI with 85% JPEG quality — this balances file size reduction with readable text. If text blurriness is unacceptable, the original PDF may already be using vector text that cannot be compressed further without quality loss.

The output file is larger than the original

A text-based PDF is extremely compact; re-rendering it as page images can increase size. Lower the resolution to 150 DPI or reduce the quality slider. For scanned (already image-based) PDFs, the grayscale output is usually similar in size or smaller.

Some text is still hard to read

Mid-grey text inverts to a similar mid-grey, so contrast stays low. It works best on documents with strong black-on-white text; heavily styled or coloured PDFs may need manual tweaking.

It cannot repair an encrypted file

Recovery works by reading the file's markers, which encryption hides. Remove the password first (the Unlock mode of PDF Security) and then run the repair.

Nothing is recovered

If the file was truncated to almost nothing or its bytes were wiped (a failing disk, an empty download), there is no surviving content to rebuild from and no tool can bring it back.

Recognized text is garbled or accented characters are wrong

Check the language selection — running a document through the wrong language model mangles diacritics and unusual characters. If the scan is rotated or skewed, fix the orientation first with the PDF page organizer; OCR accuracy drops sharply on rotated pages.

Very few words were found on a page that clearly has text

The page may be extremely low resolution, very noisy, or handwritten. Re-scan at 200-300 DPI if possible. Pages that already contain a digital text layer don't need OCR at all — try selecting text in your viewer first.

Related Guides

Related Tools

Embed this tool on your site

Paste this snippet into any HTML page or blog post to embed a live, fully working copy of Optimize PDF — Dark Mode, Compress, Grayscale, OCR & Repair. Free for any use.