Drop your PDF
Any PDF with real, selectable text.
Drop a PDF and get its text as Markdown: text noticeably larger than the body is turned into headings, so the structure of the original document comes through instead of one flat block of text.
100% private — the PDF is processed in your browser and never uploaded.
Any PDF with real, selectable text.
Font sizes are compared to guess headings.
Ready to paste into a Markdown editor or a wiki.
Plain text extraction from a PDF gives you the words but loses the shape of the document: what was a heading looks the same as a paragraph. This tool takes one more step and compares the font size of each line of text against the most common size on the page, the body text, and turns noticeably larger lines into Markdown headings, so a document with a clear structure of titles and sections comes out more readable than a flat wall of text.
Text sized at roughly 1.15 to 1.35 times the body text becomes a level-three heading, 1.35 to 1.7 times becomes level two, and anything larger becomes a level-one heading. This is a size-based heuristic, not a true understanding of the document's structure, so it works best on documents with a clear, consistent visual hierarchy, like a report with obvious section titles, and less well on documents that use bold body text or unusual styling for emphasis rather than size.
Each page's Markdown is separated by a horizontal rule, so you can see where one page ended and the next began in the source document, useful context if headings span an awkward page break.
This does not detect bold, italics, links, tables or lists, only headings based on size; everything else comes out as plain paragraph text. For the exact text with no attempt at structure, use PDF to text instead. Scanned PDFs with no real text layer need PDF OCR first. Your file is processed entirely in your browser and is never uploaded.
More utilities that also run without leaving your browser.