Check whether your last converted PDF has real text
Open any PDF a browser-based converter produced for you and try to select a sentence. If the cursor draws a box around the whole page instead of highlighting words, the text is not text — it is a picture of text. It will not be searchable with Ctrl+F, it cannot be copied, a screen reader cannot read it, and the file is far larger than it needs to be.
This is not a rare defect. It is the normal outcome of how nearly every in-browser Markdown-to-PDF converter works, and almost none of them mention it.
Why it happens, and when it is the right trade
There are only two ways to put content on a PDF page from a browser, and they have opposite properties.
The layout route hands your content to the browser’s own rendering engine, screenshots the result with html2canvas, and places those screenshots into the PDF as JPEG images. Everything the browser can lay out — nested tables, floats, inline styles, web fonts, background colours — survives exactly as you saw it in the preview. Nothing that arrives is text any more.
The text route skips layout entirely and writes characters directly into the PDF’s content stream using one of the three fonts built into the format. The output stays selectable, searchable, copyable and small. The cost is that there is no CSS: no tables, no images, no colour, one font at one size.
This converter uses both, chosen by input type. Markdown and HTML take the layout route, because a Markdown table that lost its borders would be a worse outcome than one you cannot select. Plain text takes the text route, because a meeting note or a log excerpt has no layout to preserve and every reason to stay searchable.
The size difference is not subtle. Exporting the plain-text sample on this page produces a PDF just under 4 KB carrying a native Helvetica font reference and no image objects at all. The same content through the layout route arrives as a full-page JPEG at 2× scale — a different order of magnitude for identical words.
Picking the route deliberately
The input selector is the lever, so it is worth using on purpose rather than by whichever format your source happens to be in:
- Paste as plain text when the document will be searched, archived, indexed, fed to another tool, or read by assistive technology. Contracts, notes, logs, transcripts, anything going into a document management system.
- Paste as Markdown or HTML when the visual result is the point — a README going to a client, a formatted report, anything with tables or images where losing layout would be worse than losing selectability.
- Export to Word instead when the recipient needs to edit. That path produces a real
.docxwith actual paragraph and heading styles from any of the three inputs, and sidesteps the whole question.
If you need both fidelity and selectable text, no browser-only tool will give it to you — that requires a real PDF typesetting engine (LaTeX, Prince, WeasyPrint) running server-side, which means uploading the document. That is a genuine trade, and worth knowing before you assume a converter is simply broken.
Two things that surprise people about the layout route
External images need CORS. The canvas that captures your content cannot read pixels from an image served without permissive CORS headers — the browser taints it and the capture fails or comes back blank. Base64 data: URIs always work because nothing is fetched. Relative paths like ./diagram.png can never resolve, since there is no document directory in this context.
Web fonts do not embed. Only the three PDF base fonts — Helvetica, Times, Courier — are guaranteed present. A @font-face or Google Font renders in the preview, gets captured as pixels on the layout route, and is silently substituted on the text route. If typography matters to the output, check the exported file rather than the preview.