What “text extraction” actually means
A digitally created PDF stores text as text: a font, a size and a position for every run of characters. That is why you can select and search such files. Extracting the text is a matter of walking the content stream and reading the character data in the right order.
A scanned PDF stores a photograph of a page. There is no text in it at all — only pixels. Extraction returns nothing, and no amount of cleverness changes that. Turning those pixels back into characters requires OCR (optical character recognition), which is a completely different technology.
This tool does the first job properly and reports honestly on the second. If your extract comes back empty, the PDF is a scan, and you will need an OCR tool. We would rather tell you that than hand you a blank file and let you assume the tool is broken.
Keeping the layout or reflowing the text
PDFs position every line independently, which is why extracting text naively produces odd results: lines run together, columns interleave, and words that were separate on screen end up adjacent in the output.
Keep the original line layout is the default and works by tracking vertical positions. When a run of text moves noticeably up or down the page, a line break is inserted. This preserves forms, two-column letters, itemised invoices and tabular data in a form you can still read. It is the right setting for anything you plan to read or copy by hand.
Merging into flowing text does the opposite: it removes line breaks that fall mid-sentence, joins words hyphenated across lines, and collapses runs of spaces. The result reads like a paragraph instead of a series of fragments, which is what you want when the text is going into a document, a search index or a translation tool.
Try the default first. If your output has lots of short broken lines that clearly belong together, switch the option off and extract again.
Practical uses for an extracted text file
Accounting and bookkeeping: pulling transaction lines out of a bank statement PDF so they can be pasted into a spreadsheet. Keep the layout on so that amounts stay on their own lines, then split by tabs or spaces in your spreadsheet tool.
Reviewing contracts: extracting clauses so they can be searched, compared with a template, or pasted into a redline. Switch the layout off for clean prose.
Study material: students copy definitions and formulas out of lecture PDFs into notes that they can actually edit and search.
Accessibility and translation: a text file can be read by a screen reader or pasted into a translation service, neither of which can work with a scanned image.
Data recovery: when a PDF is the only surviving copy of a document and you need the content in a format that will outlive it, plain text is the most durable format there is.
Common extraction problems and how to read them
Empty output: the PDF is a scan with no text layer. You need OCR. Common signs are a file created by a scanner or phone camera app, and a document where you cannot select text with your mouse.
Interleaved columns: the PDF stores a two-column newsletter as alternating lines. Turning the layout option off usually produces a readable single stream.
Missing bullet points and checkmarks: these are often drawn as vector graphics rather than characters, so they cannot be extracted as text. They will simply be absent.
Garbled characters: a PDF that has had its fonts subset with a broken encoding map. Copying text from such a file in any reader produces the same garbage, so the problem is in the original file rather than in the extraction.
Ligatures: some PDFs store “fi” as a single glyph. The extractor converts these back to two characters where it can; occasional oddities are normal in older documents typeset with unusual fonts.