PDF & Document to Markdown
In the Apify Store
PDF, Word, PowerPoint, Excel, CSV, HTML, EPUB and images to clean Markdown or text, with tables kept, OCR only where needed, and chunks for RAG.
One row per document, or one row per chunk when you want ready-to-embed pieces. Headings, lists and tables are kept. Pages that are only a picture are read with OCR, and only those pages, so text PDFs stay fast.
What you get
- Many formats, one tool. PDF, DOCX, PPTX, XLSX, CSV and TSV, HTML, EPUB, TXT and Markdown, and PNG, JPEG, TIFF, GIF, BMP and WebP images. The file type is read from the file's own bytes.
- Structure kept. Headings, lists and tables; two-column PDF pages; repeated headers and footers removed; one section per slide, sheet or chapter.
- OCR only where needed. Tesseract reads pages without a text layer, in 10 languages, and reports
pagesOcrand a confidence score. - Chunking for RAG. By headings or by size, with the heading path and page numbers of every chunk.
- Private by design. Documents are converted inside the run and deleted when it ends. No third-party AI or OCR service is called.
- Clear errors. A broken link, password-protected PDF or unsupported file becomes an error row with the reason. It does not stop the run and it is not charged.
Real output
{
"recordType": "document",
"status": "ok",
"url": "https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf",
"fileName": "NIST.CSWP.04162018.pdf",
"fileType": "pdf",
"pages": 55,
"pagesConverted": 5,
"pagesOcr": 0,
"title": "Framework for Improving Critical Infrastructure Cybersecurity, Version 1.1",
"metadata": {"author": "National Institute of Standards and Technology", "created": "2018-04-17T13:45:53Z"},
"language": "en",
"tables": 1,
"wordCount": 1189,
"billablePages": 5,
"markdown": "# Framework for Improving Critical Infrastructure Cybersecurity\n\nVersion 1.1\n\n...",
"text": null,
"warnings": ["The document has 55 pages; only the first 5 were converted (limit: Max pages per document)."],
"error": null
}
Example input
Paste this into the input form, or send it through the API.
{
"sources": [
"https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf",
"https://raw.githubusercontent.com/mwilliamson/python-mammoth/f3b7b9fc73fdbffe6ebac77e6b3ac6c418ee33d9/tests/test-data/tables.docx",
"https://raw.githubusercontent.com/tesseract-ocr/test/232ff181c66516116ec0e84c4963f70de15050fd/testing/phototest.tif"
],
"maxPagesPerDocument": 5
}
Run it from code
One request to the Apify API starts a run and returns the dataset rows when it finishes. Use your own Apify API token. The Apify client libraries and Apify's MCP server work too.
API=https://api.apify.com/v2/acts/tinlark~document-to-markdown
curl -X POST "$API/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"sources": ["https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf"], "maxPagesPerDocument": 5}'
Limits
- Files up to 50 MB by default (200 MB at most). Default 50 pages per document, up to 2,000.
- Public links or uploaded files only. Links to private or internal network addresses are refused. It fetches exactly the links you give it and does not crawl.
- Not supported: password-protected files, scanned handwriting, text in images inside Word or PowerPoint files, JavaScript-rendered pages, legacy
.doc,.xlsand.ppt, DRM-protected files. - Complicated layouts (magazines, posters) can come out in the wrong reading order, and tables without ruling lines are found only when their columns line up clearly.
Data source
Only the documents you give it: public http(s) links or files you upload to your own Apify account.
Questions before you start?
Write to [email protected]. Each product page lists what the tool does, what it does not do, and its exact price.