Skip to content
Case Converter Online
en

PDF to Markdown Converter

Choose a PDF and get Markdown with headings, paragraphs and bullet lists rebuilt from the page layout. The file is read on your device and never uploaded.

Choose a PDF to start. Works best with PDFs made from Word, Google Docs or a web page.

What this converter does

A PDF is built for printing. It records where each piece of text goes on the page and in which font size, but it does not record paragraphs, headings or lists. Copying text out of a PDF often gives you broken lines, page numbers in the middle of sentences and no structure at all.

This converter reads the text with PDF.js and rebuilds the structure from the layout:

  • Text pieces on the same line are joined, with spaces added where there is a visible gap.
  • Lines that sit close together become one paragraph. A larger vertical gap starts a new one.
  • Text that is clearly larger than the body text becomes a heading. The largest size becomes #, the next ##, and so on.
  • Lines that start with a bullet character such as a round dot become Markdown list items.
  • Lines that hold only a page number at the top or bottom of a page are removed.

How to use it

  1. Press Choose a PDF or drop a .pdf file on the box.
  2. Wait while the pages are read. The status line shows the progress for long files.
  3. Adjust the options. The result updates straight away, without reading the file again:
    • Detect headings turns large text into Markdown headings.
    • Remove page numbers drops lone page numbers.
    • Mark page breaks puts a --- rule between pages.
    • Output switches between Markdown and Plain text.
  4. Check the result. Preview shows it rendered, and you can edit the Markdown in the box.
  5. Press Copy or Download.

Example

We made a two page test PDF. Page one has a 22 point title, 15 point section headings, two 11 point paragraphs, two bulleted lines and a page number at the bottom. Page two has a heading, one sentence and a page number. The converter returned:

# Quarterly Report

## Overview

Revenue rose in all three regions this quarter. The largest gain came from online orders.

Costs stayed close to the plan.

## Next steps

- Hire two support staff
- Open the Leeds office

## Appendix

All figures are unaudited.

The first paragraph was two lines in the PDF. They were joined into one, because the gap between them was normal line spacing. “Costs stayed close to the plan.” starts a new paragraph because the gap above it was larger. Both page numbers were removed. With Output set to Plain text, the headings lose their # marks and the bullets keep their original dot characters.

What works well and what does not

Results depend on how the PDF was made.

Works well: PDFs exported from Word, Google Docs, LibreOffice, Pages or a web browser. These contain real text in reading order, with clear font sizes for headings.

Needs checking:

  • Multi-column layouts. Text is taken in the order the PDF draws it. In some two column documents this is column by column, which is right. In others lines from both columns can interleave.
  • Headings in the same size as body text. A heading that is only bold, not larger, is not detected, because bold is not reliably marked in PDF text data.
  • Tables. Table text comes out as lines or short paragraphs. Rebuild the table with the Markdown table generator.
  • Footnotes, headers and footers. Repeated running headers appear on every page. Delete them after converting.
  • Hyphenated words at line ends are joined without a space, so “well-” followed by “known” becomes “well-known”. A word that was split only to fit the line keeps its hyphen too, as in “infor-mation”, so look for stray hyphens.

Does not work:

  • Scanned PDFs. These are images. The tool warns you when it finds very little text. They need OCR first.
  • Images, charts and form fields. They are skipped.

Text in Chinese, Japanese and Korean PDFs is supported. The character maps PDF.js needs for these are bundled with this site and load only when a file uses them.

Tips

Other ways to convert PDF to Markdown

  • pdftotext from the Poppler tools extracts plain text on the command line. Its -layout option keeps the page layout with spaces.
  • Python: libraries such as pdfplumber and PyMuPDF give you text with positions, so you can write your own rules for headings and tables.
  • Pandoc does not read PDF files, so it cannot do this conversion directly.

Large PDFs with hundreds of pages work, but they take longer because every page is read on your device. Files over 100 MB are refused to keep the browser responsive.

PDF to Markdown: questions and answers

Why is the result empty or almost empty?

The PDF is probably a scan or a photo of pages. Scanned PDFs hold images of text, not text, so there is nothing to extract. They need OCR (optical character recognition) first. Many scanner apps and PDF editors can add a text layer with OCR, after which this tool can read the file.

Are tables and images converted?

No. A PDF stores a table as separate pieces of text placed in a grid, with no record of rows and cells. The text of a table comes out as plain lines. Images are skipped. If you have the original Word file, the Word to Markdown converter keeps tables.

Is my PDF uploaded?

No. The PDF is read by PDF.js, the same engine Firefox uses to show PDFs, running inside your browser. The file is not sent to a server, so you can use it for contracts, reports and other private documents.

Can it open password-protected PDFs?

Yes, if you know the password. When a PDF is locked, a password box appears. Type the password and press Open. The password is only used in your browser to decrypt the file.

More code and data tools

See all tools