What this converter does
A PDF is built for printing. It records where each piece of text goes on the page and in which font size, but it does not record paragraphs, headings or lists. Copying text out of a PDF often gives you broken lines, page numbers in the middle of sentences and no structure at all.
This converter reads the text with PDF.js and rebuilds the structure from the layout:
- Text pieces on the same line are joined, with spaces added where there is a visible gap.
- Lines that sit close together become one paragraph. A larger vertical gap starts a new one.
- Text that is clearly larger than the body text becomes a heading. The largest size becomes
#, the next##, and so on. - Lines that start with a bullet character such as a round dot become Markdown list items.
- Lines that hold only a page number at the top or bottom of a page are removed.
How to use it
- Press Choose a PDF or drop a
.pdffile on the box. - Wait while the pages are read. The status line shows the progress for long files.
- Adjust the options. The result updates straight away, without reading the file again:
- Detect headings turns large text into Markdown headings.
- Remove page numbers drops lone page numbers.
- Mark page breaks puts a
---rule between pages. - Output switches between Markdown and Plain text.
- Check the result. Preview shows it rendered, and you can edit the Markdown in the box.
- Press Copy or Download.
Example
We made a two page test PDF. Page one has a 22 point title, 15 point section headings, two 11 point paragraphs, two bulleted lines and a page number at the bottom. Page two has a heading, one sentence and a page number. The converter returned:
# Quarterly Report
## Overview
Revenue rose in all three regions this quarter. The largest gain came from online orders.
Costs stayed close to the plan.
## Next steps
- Hire two support staff
- Open the Leeds office
## Appendix
All figures are unaudited.
The first paragraph was two lines in the PDF. They were joined into one, because the gap between them was normal line spacing. “Costs stayed close to the plan.” starts a new paragraph because the gap above it was larger. Both page numbers were removed. With Output set to Plain text, the headings lose their # marks and the bullets keep their original dot characters.
What works well and what does not
Results depend on how the PDF was made.
Works well: PDFs exported from Word, Google Docs, LibreOffice, Pages or a web browser. These contain real text in reading order, with clear font sizes for headings.
Needs checking:
- Multi-column layouts. Text is taken in the order the PDF draws it. In some two column documents this is column by column, which is right. In others lines from both columns can interleave.
- Headings in the same size as body text. A heading that is only bold, not larger, is not detected, because bold is not reliably marked in PDF text data.
- Tables. Table text comes out as lines or short paragraphs. Rebuild the table with the Markdown table generator.
- Footnotes, headers and footers. Repeated running headers appear on every page. Delete them after converting.
- Hyphenated words at line ends are joined without a space, so “well-” followed by “known” becomes “well-known”. A word that was split only to fit the line keeps its hyphen too, as in “infor-mation”, so look for stray hyphens.
Does not work:
- Scanned PDFs. These are images. The tool warns you when it finds very little text. They need OCR first.
- Images, charts and form fields. They are skipped.
Text in Chinese, Japanese and Korean PDFs is supported. The character maps PDF.js needs for these are bundled with this site and load only when a file uses them.
Tips
- If you still have the source document, convert that instead. The Word to Markdown converter keeps tables, links and exact heading levels.
- Run the output through the Markdown formatter to wrap or unwrap paragraphs in one style.
- For a web page saved as PDF, the original HTML often gives a cleaner result in the HTML to Markdown converter.
Other ways to convert PDF to Markdown
- pdftotext from the Poppler tools extracts plain text on the command line. Its
-layoutoption keeps the page layout with spaces. - Python: libraries such as pdfplumber and PyMuPDF give you text with positions, so you can write your own rules for headings and tables.
- Pandoc does not read PDF files, so it cannot do this conversion directly.
Large PDFs with hundreds of pages work, but they take longer because every page is read on your device. Files over 100 MB are refused to keep the browser responsive.