Convert from PDF

PDF to Text — copy the words out of a PDF

Pull the text layer out of a PDF into an editable .txt file, with optional page separators and layout preservation. Useful for notes, contracts, receipts and reports that need to go into another system.

No upload Files stay in your browser Free, no signup Works offline

How to extract text from a PDF

Add your PDFs

One file or several. Each is read separately, and you can combine them into a single output.

Choose the output style

Keep the original line layout for tables and forms, or switch it off to merge hyphenated line breaks into flowing paragraphs.

Download the .txt

Press “Extract text”. One .txt file downloads, or a ZIP if you chose separate outputs.

Scanned documents contain no text layer, so the tool tells you clearly instead of returning an empty file.

Features

Layout-aware extraction

Keeps the original line structure for tables and forms, or reflows it for prose.

Page separators

Optional markers such as “--- Page 2 ---” so you can find your way through a long extract.

Multiple PDFs at once

Process a folder of statements and get either one combined file or one file each.

Honest about scans

If a document has no text layer, you are told plainly rather than given a blank file.

Private by design

Text is extracted in your browser. Contracts and salary slips never leave your device.

Free, unlimited, no signup

No page limits, no account, no watermark, no daily quota.

What “text extraction” actually means

A digitally created PDF stores text as text: a font, a size and a position for every run of characters. That is why you can select and search such files. Extracting the text is a matter of walking the content stream and reading the character data in the right order.

A scanned PDF stores a photograph of a page. There is no text in it at all — only pixels. Extraction returns nothing, and no amount of cleverness changes that. Turning those pixels back into characters requires OCR (optical character recognition), which is a completely different technology.

This tool does the first job properly and reports honestly on the second. If your extract comes back empty, the PDF is a scan, and you will need an OCR tool. We would rather tell you that than hand you a blank file and let you assume the tool is broken.

Keeping the layout or reflowing the text

PDFs position every line independently, which is why extracting text naively produces odd results: lines run together, columns interleave, and words that were separate on screen end up adjacent in the output.

Keep the original line layout is the default and works by tracking vertical positions. When a run of text moves noticeably up or down the page, a line break is inserted. This preserves forms, two-column letters, itemised invoices and tabular data in a form you can still read. It is the right setting for anything you plan to read or copy by hand.

Merging into flowing text does the opposite: it removes line breaks that fall mid-sentence, joins words hyphenated across lines, and collapses runs of spaces. The result reads like a paragraph instead of a series of fragments, which is what you want when the text is going into a document, a search index or a translation tool.

Try the default first. If your output has lots of short broken lines that clearly belong together, switch the option off and extract again.

Practical uses for an extracted text file

Accounting and bookkeeping: pulling transaction lines out of a bank statement PDF so they can be pasted into a spreadsheet. Keep the layout on so that amounts stay on their own lines, then split by tabs or spaces in your spreadsheet tool.

Reviewing contracts: extracting clauses so they can be searched, compared with a template, or pasted into a redline. Switch the layout off for clean prose.

Study material: students copy definitions and formulas out of lecture PDFs into notes that they can actually edit and search.

Accessibility and translation: a text file can be read by a screen reader or pasted into a translation service, neither of which can work with a scanned image.

Data recovery: when a PDF is the only surviving copy of a document and you need the content in a format that will outlive it, plain text is the most durable format there is.

Common extraction problems and how to read them

Empty output: the PDF is a scan with no text layer. You need OCR. Common signs are a file created by a scanner or phone camera app, and a document where you cannot select text with your mouse.

Interleaved columns: the PDF stores a two-column newsletter as alternating lines. Turning the layout option off usually produces a readable single stream.

Missing bullet points and checkmarks: these are often drawn as vector graphics rather than characters, so they cannot be extracted as text. They will simply be absent.

Garbled characters: a PDF that has had its fonts subset with a broken encoding map. Copying text from such a file in any reader produces the same garbage, so the problem is in the original file rather than in the extraction.

Ligatures: some PDFs store “fi” as a single glyph. The extractor converts these back to two characters where it can; occasional oddities are normal in older documents typeset with unusual fonts.

Frequently asked questions

Why did I get an empty text file?

Because the PDF has no text layer — it is a scan or a photograph of a page. You need OCR software to convert those pixels into characters. A quick way to check is to try selecting the text with your mouse in a PDF reader: if nothing highlights, there is nothing to extract.

Can I extract text from several PDFs at once?

Yes. Add up to 50 files. With “combine all into one .txt file” ticked you get a single file with each document marked by its name; untick it and you get one .txt per PDF inside a ZIP.

How do I keep the original line breaks?

Leave “Keep the original line layout” ticked. That preserves the visual structure of tables and forms. Untick it to merge broken lines back into flowing paragraphs.

Are tables extracted properly?

Lines and cell text are extracted in reading order, which is often usable but rarely a perfect grid. For tabular data, keep the layout option on and then paste into a spreadsheet and split by delimiter.

Is the text extracted on a server?

No. Extraction happens in your browser using Mozilla PDF.js. Confidential contracts and financial statements stay on your own machine.

Can it extract text from an image inside a PDF?

No. Images contain no characters. Extracting images is a separate job, which the Extract Images from PDF tool handles, and recognising words inside them is OCR.

What encoding is the .txt file in?

UTF-8, which is the modern standard and is read correctly by Notepad, TextEdit, VS Code, Excel and every other tool you are likely to use.

Written and maintained by the PDFUtilise team. Last reviewed: 2026-10-07. Found a problem with this tool? Tell us — we fix reported bugs fast.