What an upscaler actually does
A traditional resize interpolates. It calculates intermediate pixel values from the ones either side, produces a larger and softer image, and adds no information that was not already there.
A modern upscaler generates. It is a model trained on pairs of small and large images, and it produces detail that is plausible for the subject rather than derived from it: pores on a face, threads in fabric, leaves on a distant tree.
That distinction is the whole article. The pores were not photographed. They are a model best guess at what pores would look like there, which is the same operation a generator performs across a whole frame.
So an image that has been through an upscaler contains genuinely synthetic texture, and a detector reading texture reports exactly that. The finding is correct; the conclusion people draw from it is not.
Why this produces so many false positives
Upscaling is everywhere and mostly invisible to the person doing it. Phone cameras apply it to digital zoom, editing suites offer it as enhance, video calls use it to compensate for bandwidth, and platforms apply it when serving images at sizes above the original.
Archive and restoration work is the extreme case. An old photograph scanned at low resolution and enhanced for print is heavily rewritten by construction, and it will score like something a model produced, because much of what you are looking at is.
The top-right region is the problem. Those images depict real things and contain invented pixels, and no score can separate the two properties because the model is measuring only one of them.
How to tell them apart
The region map is the first tool, and it works because upscaling and generation distribute their invented texture differently.
- Upscaling is uniform. Every part of the frame was enlarged by the same process, so the grid is warm and flat.
- Generation is also uniform, which is why the map alone does not settle it.
- Selective enhancement is not. Face restoration applied to one person in a group photo produces localised heat, which is informative.
- Composition still holds. An upscaled photograph has the geometry, perspective and lighting of a real scene, because it was one.
- Detail can be checked. Enlarge the invented regions: upscaled text becomes plausible nonsense, and real text stays readable.
The last point is the practical one. Zoom into signage, labels or small print. An upscaler faced with unreadable text produces confident letter-shaped marks that spell nothing, and that failure is visible to a person even when a score is not.
| Signal | Upscaled photograph | Generated image |
|---|---|---|
| Whole-frame score | High | High |
| Region map | Warm and flat | Warm and flat |
| Small text on enlargement | Plausible nonsense | Plausible nonsense |
| Scene geometry | Consistent, it was real | Usually consistent now |
| An earlier, smaller copy exists | Often | No |
| Original file available | Yes, ask for it | There is none |
The bottom two rows are the ones that actually resolve it, and neither is a pixel measurement. A reverse image search finding a smaller earlier copy, or a source who can supply the pre-enhancement file, settles the question directly.
What this means in practice
For anybody assessing an image, the operating rule is to ask about enhancement before treating a high score as a finding. It is a single question, it costs nothing, and it explains a large share of confusing results.
For anybody publishing, it is an argument for keeping originals. A photographer, a claims handler or a journalist who can produce the file before enhancement has a complete answer to a question that otherwise has none.
And for archives and restoration work it is a labelling problem rather than a detection one. An enhanced historical photograph is a legitimate and useful artefact that will always score badly, so the honest response is to say what was done to it.
Where this causes real harm
Three settings turn this from a curiosity into a problem, and all three share a feature: somebody is making a decision about a person based on an image they did not capture.
Insurance claims are the clearest. A claimant photographs damage on a phone, the image is small, and somebody enlarges it to see the detail. The enlarged version scores high, and a genuine claim acquires a note about possible manipulation that nobody can now remove.
Journalism has the same shape. A photograph from a source is upscaled for print, the enhanced version circulates, and a later check on the circulating copy flags it. The original was fine and is no longer the version anybody has.
Legal and disciplinary contexts are the most serious, because the enhanced copy is often the one entered into the record. Where an image is going to be relied on, the file that should travel is the one before any enhancement, and the enhancement should be documented separately.
A note for anyone building a workflow
If images pass through your organisation and get assessed, the cheapest improvement available is a field recording whether a file has been enhanced. It costs a checkbox at intake and removes an entire category of unresolvable dispute later.
The second cheapest is a rule that enhancement happens to a copy. Keep the file as received, enhance a duplicate for viewing, and make sure the one that travels with the record is the original rather than the pretty one.
There is a broader lesson here about what a score measures. The model reports how much of the texture in front of it was produced by software, which is a real property and not the same as the property people want, which is whether the picture records something that happened.
Most of the confusing results in this field come from that gap. Upscaling is the clearest example because the two properties point in opposite directions: heavily synthetic texture, entirely genuine subject.