← All guides Workflow

Finding AI images inside PDFs and documents

Reports, decks and claim bundles carry images that nobody checks. How to extract them without losing quality, and which ones are worth the effort.

· 10 min read · Best AI Image Detector

Extract the embedded images rather than screenshotting pages. A PDF usually stores the original file inside it, and pulling that out preserves the encoding a check needs, which a page screenshot destroys.

Why documents are the blind spot

Verification effort concentrates on images that arrive as images. A photograph in a message gets checked; the same photograph inside a fifty-page report does not, because nobody thinks of a document as a container of pictures.

Yet documents are where images do their most consequential work. A claim bundle, a valuation report, a due diligence pack, a site survey, a compliance submission: each is a wrapper around photographic evidence that somebody will rely on.

The format also confers unearned authority. An image on a page, with a caption, a figure number and a letterhead around it, reads as more official than the same file sent on its own, and that framing discourages the question.

Volume compounds it. A bundle with two hundred photographs is not going to be checked one at a time by hand, so in practice it is not checked at all unless the extraction step is easy.

Extract, do not screenshot

This is the single technical point that matters. When an image is placed in a PDF, the file is usually embedded largely intact, which means the original encoding is sitting inside the document waiting to be recovered.

Screenshotting the page throws that away and substitutes a rendering by your PDF viewer at whatever zoom you happened to be at. It is the same loss described elsewhere on this site, applied to a case where the good copy was available all along.

Extract the embedded file

  • Original encoding preserved
  • Original dimensions
  • Metadata sometimes intact
  • Usable region map
  • A few clicks or one command

Screenshot the page

  • Re-encoded by your viewer
  • Whatever zoom you were at
  • Metadata gone entirely
  • Score pushed to the middle
  • Feels faster, answers less
Two ways to get an image out of a document, and what each leaves you with.

Most PDF readers offer an export or extract option, and command line tools do it in bulk. Office documents are simpler still: a modern Word or PowerPoint file is a zip archive, and renaming it exposes a media folder containing every image at full quality.

Which images are worth checking

A two-hundred-image bundle does not need two hundred checks. Prioritising by what an image is being asked to prove reduces the work by an order of magnitude.

  1. Images that evidence a condition. Damage, defects, wear, site conditions, medical presentations. These are the ones a decision turns on.
  2. Images of documents. A photographed certificate, licence or receipt inside a bundle is two layers of evidence and gets checked at neither.
  3. Before-and-after pairs. The most persuasive format in any report and the easiest to construct dishonestly.
  4. Anything dated or timestamped. Where an image is offered as proof of when something happened.
  5. Stock and decorative imagery. Skip entirely. Nobody is relying on the header photograph.

The third row deserves attention. A before-and-after pair is unusually easy to fake in one direction, because generating the before from the after is trivial with modern tools and nobody thinks to check the earlier image.

Reading a bundle as a set

The batch approach that works elsewhere works especially well here, because a document bundle usually claims a common origin. Photographs taken during one site visit, on one device, in one session, should cluster.

That gives a baseline for free. The question stops being whether 54 is high, which depends on the camera and the compression, and becomes why one image scored 54 when the other nineteen from the same visit scored between 18 and 26.

A bundle that scatters with no cluster is telling you something too: the images did not come from one session, whatever the document says. That is a finding about the report rather than about any picture in it.

Common document types and what to prioritise
DocumentCheck firstSkip
Insurance claim bundleDamage photographs, receiptsCover artwork
Property or valuation reportInterior and defect photosLocation maps
Due diligence packFacility and product imagesTeam headshots
Compliance submissionEvidence photographs, certificatesCharts and diagrams
Marketing deckProduct and case study imagesEverything else

Where this fits in a review process

The extraction step is worth automating even where the checking is not. A folder of images pulled from every incoming bundle, named after the document they came from, turns an occasional heroic effort into something a reviewer can glance at.

That also changes what gets noticed. A reviewer looking at twenty extracted photographs side by side sees the one from a different camera immediately, and that observation needs no analysis at all.

Where volume is genuinely high, the proportionate rule is to check bundles rather than images: run one batch per submission, look at the spread, and open individual results only when the spread has something in it.

None of this needs to be a project. A shared folder, an extraction command and a habit of running one batch per bundle covers most of the value, and it is the kind of process that survives the person who set it up leaving.

What to do with what you find

Extraction and checking produce a list of images and numbers, and the useful part is what happens next. The response that works is the same one that works everywhere else: ask, do not conclude.

A flagged photograph in a bundle is a request for the original file, addressed to whoever compiled the bundle. Most of the time the answer is that the image was enhanced, resized on export or photographed from a screen, and that ends it.

Where it does not end it, the finding to record is the question and its answer, not the score. A file note saying that an image was queried, an original was requested and one was or was not supplied is defensible in a way that a stored number never is.

One practical warning about bulk extraction: a single page can contain an image split into several tiles by the export process, and those tiles score badly on their own because each is a fragment. Where extracted pieces look unusually small, reassemble or work from the page image instead.

Questions people ask

How do I get images out of a PDF?
Most readers have an export or extract images option, and command line tools do it in bulk. What matters is that you extract rather than screenshot: the embedded file usually retains its original encoding, and a page screenshot replaces that with a rendering by your viewer.
What about Word and PowerPoint files?
Easier than PDFs. A modern Office file is a zip archive, so copying it and changing the extension to zip lets you open it and find a media folder containing every embedded image at full quality. No specialist tooling is required.
Do I have to check every image in a bundle?
No, and trying to is why bundles go unchecked. Prioritise images that evidence a condition, photographs of documents, before-and-after pairs and anything offered as proof of timing. Skip decorative and stock imagery entirely.
Why do extracted images sometimes score high?
Often the export settings. Some tools recompress images on the way into a document, which degrades the encoding a check reads and pushes scores upward. Where a result looks odd, ask for the original photographs rather than working from the report.
What is the most useful thing to look for?
The outlier. A bundle claiming one site visit should produce photographs that cluster, because they share a camera and a session. The image sitting well outside that cluster is the one to open first, and it is not always the highest scoring one.
Are before-and-after photos worth checking?
They are among the most worthwhile in any bundle. The pair is the most persuasive format in a report and the easiest to construct dishonestly, because generating a plausible before image from an after image is straightforward and nobody expects the earlier picture to be the fabricated one.