Claupt
📄 PDF → Markdown
📄

PDF/Office → Markdown Layout Parser

Stop feeding AI models expensive layout noise. Drop a PDF or .docx below — get editable Markdown with estimated # headings and bullets. PDF tables and complex layouts may need manual repair. Tool inputs are processed in your browser.

📂

Drop your PDF, DOCX, or TXT here

or click to browse — maximum 20 MB · processed locally

⚙️ Cleaning Options

1️⃣

Parse locally

pdf.js / mammoth.js run in your browser — your file never touches a server.

2️⃣

Strip the bloat

Footers, page numbers, repeated headers, and XML noise are removed.

3️⃣

Paste to any AI

Markdown preserves hierarchy with zero layout tokens — paste & go.

Convert documents into editable text

Choose a text-based PDF, DOCX, TXT or Markdown file up to 20 MB. For PDFs, select cleaning options before choosing the file. Options for page numbers and headings apply to PDFs; DOCX uses its document styles and TXT receives whitespace cleanup.

Example: a PDF line ending in “docu-” followed by “ment” can become “document”. Turn off page-number removal if standalone numbers are meaningful data.

Scanned PDFs require OCR elsewhere. Columns, equations and tables can lose structure. PDF file bytes are not a token count, so this tool does not claim a percentage saving from file size. Inspect the Markdown before using it.