Extract text from a PDF
Choose a PDF to collect its embedded text. This tool does not perform OCR on scanned pages or text stored only as images.
No file bytes are sent to Lutrakit.
Loading
Loading
Extract the text already stored in a PDF
Turn a PDF's embedded text into a plain text file for searching, quoting or further editing. This is text extraction, not recognition of words in scanned pictures.
How to use it
- Select a PDF that contains a text layer.
- Enable page separators if you need to identify the source page of each passage; they are off initially. Choose the output filename.
- Extract and download the TXT file, then compare a sample paragraph, columns and special characters with the PDF.
Example
- Starting point
- A PDF with the extractable text Hello on page 1 and World on page 2, with page separators disabled.
- What to expect
- A TXT file containing Hello, a blank line, then World. Page-separator mode adds numbered page headings instead of relying only on blank lines.
Before you start
- The input limit is 100 MiB. An image-only scan needs OCR elsewhere before this tool can extract its words.
- TXT does not preserve fonts, pictures or table layout. Reading order depends on the PDF's text structure; review columns and unusual character encoding before reusing the result.