The problem with one number
Consider two images. The first is generated end to end and every part of it is synthetic. The second is a real photograph of a real room with a crack added to one wall.
A whole-frame model reduces each to a single figure. The first scores high, correctly. The second scores low, because eight ninths of it came from a camera, and the average buries the edit.
The second image is the one that costs somebody money. It is the insurance claim, the altered receipt, the doctored listing photo, the news image with a detail changed. A tool that reports one number will clear it.
The three patterns
Once the frame is scored in tiles, three shapes appear, and each means something different.
Every tile high
Consistency is the signature. Real photographs almost never score this evenly, because real scenes have varied texture.
One tile high, the rest cold
The cold tiles are genuine camera capture. The hot one is not. This is the pattern a single number destroys.
Evenly warm, nothing hot
Uniformly middling. Global processing raises everything a little and nothing a lot, which is the fingerprint of software rather than fabrication.
How to read a map properly
-
Look at the spread before the peak
The gap between the highest and lowest tile matters more than the highest value. A wide gap points to a local edit; a narrow one points to whatever happened to the whole frame.
-
Check whether hot tiles are adjacent
Real edits are contiguous, because objects are contiguous. Two hot tiles next to each other is a plausible insert. Two at opposite corners usually is not.
-
Compare against the picture
Look at what is actually in the hot region. If it contains the damage, the price, the face or the object under dispute, that is meaningful. If it contains empty sky, it probably is not.
-
Discount tiles with little detail
Flat areas of sky, wall or water carry less signal and score less reliably. A hot tile covering a blank wall deserves less weight than one covering a detailed object.
-
Treat conflict as uncertainty
A low frame score with several scattered hot tiles is usually compression rather than editing. That is a reason to find a better copy, not to split the difference.
When no map is produced
Below roughly 576 pixels on the shorter side, the frame cannot be divided into regions with enough detail to score independently. You get a whole-frame number only.
That is a real limitation on thumbnails, profile pictures and images pulled from chat apps, which is unfortunately where a lot of suspicious images arrive. Finding a larger copy of the same picture is usually the single most useful next step.
Why this matters more than accuracy
Improving a whole-frame model from 91 to 93 percent changes very little for somebody checking a claim photo. Being able to see that one corner of the frame behaves differently from the rest changes the finding entirely.
The region view also makes a result explainable, which matters wherever a decision has to be justified later. A moderator, a claims handler or an editor can point at a specific part of an image and say why they looked harder at it. Nobody can do that with a percentage.
Why overlapping tiles rather than a plain grid
A naive grid has a weakness that shows up immediately in practice. An edit sitting on a boundary is split between two tiles, and each one sees half the evidence against a full tile of genuine pixels. Both come back moderate, and the finding disappears.
Overlapping the crops fixes it. Every point in the frame falls fully inside at least one region, so a small insert is scored against a tile it dominates rather than one it shares. The cost is running the model more times, which on a local device is a second rather than a bill.
It also explains why adjacent tiles often light up together. A single inserted object that sits across two overlapping regions raises both, which is why adjacency is a supporting signal rather than a sign of two separate edits.
The practical consequence is that tile counts vary with the shape of an image. A wide panorama is divided differently from a square portrait, and the grid dimensions are reported with every result so you can see how finely the frame was cut.