What is actually inside a PDF
A PDF page is a set of drawing instructions. Some of those instructions draw vector shapes — lines, curves, boxes. Others place an image object, called a form XObject, which holds the actual pixels.
Those image objects are what this tool finds. Each one has a width, a height and a stream of pixel data, compressed with JPEG, Flate, or occasionally something more exotic like JBIG2 or CCITT for scanned documents. When the same image is used on many pages — a company logo, a letterhead, a watermark — the PDF stores it once and references it repeatedly, which is exactly why de-duplication matters when extracting.
Because the extractor works from the stored pixel data rather than from a picture of the rendered page, the images you get back are the original ones at full resolution. A photograph that was placed at 20% scale on the page comes out at its full size, not tiny.
PNG or JPG, and why it matters here
The images inside a PDF have already been compressed once, by whatever produced the file. Extracting them and re-encoding involves a second compression, so the format choice affects quality.
PNG is lossless. The pixels you get are exactly the pixels that were decoded from the PDF, with no further loss. Choose it when the images will be edited, printed, or used as masters.
JPG re-compresses, which loses a little detail each time, but produces dramatically smaller files for photographs. At 90% quality the difference is invisible for most uses, and a 50-image ZIP becomes manageable rather than enormous.
If the source images were photographs, extracting as JPG is usually the pragmatic choice. If they were charts, logos or screenshots — anything with sharp edges and flat colour — PNG gives a cleaner result and often a smaller file too.
Where this gets complicated, and why the tool tells you
Not every PDF contains extractable pictures, and this is the most common source of confusion.
Vector-only PDFs have no images at all. A chart exported from a spreadsheet, a diagram from a drawing tool, an icon drawn as paths — all of these are instructions, not pixels. There is nothing to extract, and the tool will say so rather than returning an empty ZIP.
Images built from many small pieces happen in scanned documents where each page was assembled from strips or tiles. Extraction will return each tile, which looks like a strange fragment rather than a page. If you want the page as a picture, the PDF to JPG tool is the right choice.
Images stored in unusual formats such as JBIG2 (used in some compressed scans) may not decode cleanly in a browser. Those are reported as skipped rather than silently dropped.
Masked images — where a soft shape is combined with a transparency mask — come out as the base image without the mask in some files, which can leave a rectangular background where the original had a cut-out shape.
Practical uses for extracting images
Recovering a photograph from a PDF brochure or report where the original file is lost. Extraction gives you the full-resolution version, not the small version shown on the page.
Pulling a logo or letterhead out of a supplier’s document so you can check branding or reuse a graphic in a presentation. If the logo is vector art, extraction will not find it — take a high-resolution PDF to PNG export instead.
Getting charts and figures out of a research paper or a report to reuse in slides. Export as PNG at a good resolution to keep the text labels crisp.
Collecting product images from a catalogue PDF that contains fifty item photos. The size filter lets you skip icons and rules, and de-duplication stops the same placeholder appearing fifty times.
Checking what a document contains: occasionally the point is not the images but knowing what is in a file — especially useful when reviewing a document you did not create and want to understand what it embeds.