AI How to spot AI generated text, images and video
How to tell whether an image was generated by AI
The tells that still work on AI images, the ones that stopped working, and the provenance data that settles it when your eyes cannot.
The short answer
- Where an image came from tells you more than how it looks, so check the source before you zoom in on the pixels.
- The tells that still hold are text, counting, occlusion and lighting geometry, because none of them are things a generator reasons about.
- The famous tells, waxy skin and six fingers, are largely fixed, and a picture with clean hands is not evidence of anything.
- A provenance credential that traces back to a camera is real evidence, while the absence of one is meaningless because uploads strip it routinely.
- Reverse image search settles more cases than inspection does, because most misleading pictures are real photos with false captions.
- Detector scores are probabilistic and should never be the basis of an accusation against a person.
Start with where the picture came from, not with the picture. Who posted it, when, and does the same image exist anywhere earlier or from anybody else? That question resolves more cases than any amount of squinting, partly because generators keep getting better and partly because most misleading images circulating online are not generated at all. They are real photographs with a false caption attached. Visual inspection is still worth doing, and this article covers what actually survives, but treat it as the second step and keep your confidence proportionate. The only thing that reliably settles it is a record of where the file has been.
Check the source before the pixels
Four questions, in order, before you look at anything inside the frame.
- Who is showing it to you? Open the account. How old is it, what did it post before this, does it post about anything else? Accounts created recently that post only high emotion content are the single strongest signal available.
- Where is the earliest copy? Search a distinctive phrase from the caption and sort by oldest, then reverse image search the picture itself. A genuine newsworthy photo has a trail back to a photographer, an agency or a witness.
- Would anyone else have photographed this? A dramatic event in a public place yields several images from several angles. Exactly one picture, from one account, with no second source, is what a fabrication looks like.
- What is the caption claiming? Separate the image from the claim. Very often the image is authentic and the claim about what it shows is the lie.
These are the same habits that work on text and on any online claim. The broader version, including how to judge a publication you have not heard of, is in how to tell whether a source is worth believing.
What still shows in the image
The useful tells share a cause. A diffusion model enforces local plausibility: every patch of the picture has to look right next to its neighbors. Nothing enforces global consistency, because there is no scene, no geometry and no counting. So failures cluster wherever a picture has to be right at a distance, and that is where you look. The mechanism is set out in how a generator turns noise into a picture, and it predicts this list better than any collection of anecdotes.
| What to check | Why it fails | How strong is the signal |
|---|---|---|
| Small text, signs, labels, spines of books | Letters are treated as texture, not language | Still strong for anything beyond a few large words |
| Things that come in pairs, earrings, shoes, glasses arms | Left and right are distant regions decided independently | Strong, and easy to check quickly |
| Counting, fingers, teeth, chair legs, crowd members | No tally is kept of anything | Moderate on the subject, strong in the background |
| Objects that pass behind something | The two visible ends are denoised far apart | Strong, straps, railings, cables and necklaces misalign |
| Shadow directions and reflections | No light is simulated, only correlations learned | Strong, shadows from one source rarely agree |
| Depth of field and perspective | Blur is stylistic rather than optical | Moderate, look for a background that converges wrongly |
| Repeating texture, crowds, bricks, foliage, patterns | Detail is synthesized, not recorded | Good, patterns melt or repeat unnaturally at the edges |
| Where the subject meets the background | Composited or inpainted regions blend imperfectly | Moderate, and destroyed by heavy compression |
Practical method: view the image at full size, then zoom to about 200 percent and work around the edges and background rather than the face. Generators put their quality where the attention goes. The person in the center is usually excellent and the fourth person from the left is usually a mess.
The tells that stopped working
A lot of circulated advice is now a year or more out of date, and following it produces confident mistakes in both directions.
- Six fingers. Hands were the standing joke, so they were targeted specifically. Current systems get hands right most of the time. Correct hands are not evidence of a camera.
- Waxy, airbrushed skin. This came from tuning on human preference ratings, and it can be prompted away. Pores, blemishes and unflattering light are all available on request.
- Too perfect composition. Models now imitate phone snapshots, bad framing, motion blur, flash glare and low resolution, because those styles were requested.
- No metadata means AI. Wrong and backwards. Social platforms, messaging apps and screenshots strip metadata from real photographs every day.
- Obvious watermarks. Visible logos are cropped out in seconds. Their absence tells you nothing.
There is a bigger problem with pure inspection: the mixed image. A real photograph with one object removed, one person added or a sky replaced passes almost every test above, because most of it came out of a camera. Editing has been possible for decades. What changed is how little skill it now takes.
Metadata and provenance
Two different things get confused here. Ordinary metadata is the EXIF block a camera writes: make, model, lens, exposure, often GPS. It is easy to read, easy to edit and easy to forge, so a photo containing camera data is weak support and a photo containing none is no evidence at all.
Provenance is stronger. Content Credentials attach a cryptographically signed record to the file listing what created it and what edits followed, with each step signed by a tool that participates in the scheme. Some cameras sign at the moment of capture, major generators sign their output as AI generated, and some editors record the history of changes. Because it is signed, tampering breaks it visibly. Because platforms strip it inconsistently, a missing record is routine rather than suspicious. The full explanation of how to read one, and what it does not prove, is in the provenance label on a picture.
Several large generators also embed an invisible watermark in what they produce, a statistical pattern that survives cropping and mild editing and is readable by the maker's own tool. It helps platforms label content at scale. It does not help you personally, and an open model can be run with the marking removed, which is one of the trade offs discussed in what open weight models actually mean.
Reverse image search is still the best tool
Run the picture through more than one reverse search engine, because each indexes different parts of the web, and try a cropped version as well: cropping to the distinctive object sometimes finds a match that the full frame misses. You are looking for four outcomes, and three of them close the case.
- An earlier copy from a different event, usually years old. This is the most common answer by far, and it means the picture is real and the caption is not.
- An original from a news agency or a named photographer, which supports it.
- Appearance only on accounts that all posted within the same hour, which is consistent with something fabricated for this moment.
- Nothing at all. This is the ambiguous result, and it tells you only that the image is not widely indexed.
When you cannot resolve it, the honest position is that you do not know, and the useful move is to check the claim instead of the picture. Whether the event happened is usually answerable even when the photograph is not, using the approach in checking a viral claim before you pass it on.
About those detector tools
Image detectors look for statistical traces left by generation, such as characteristic frequency patterns in the noise. They work reasonably on clean files from the systems they were trained on and degrade sharply on anything else: a new generator, a screenshot, a heavily compressed re-upload, a photograph of a screen. They also produce false positives on real photographs that have been denoised, upscaled or run through a phone's computational photography pipeline, which is nearly all of them now.
Use a detector as one input among several, never as a verdict, and never as the basis of an accusation against a person. The same caution applies to the tools sold for writing, and the reasons are set out in why a detector score is not evidence. If what you are assessing is text rather than a picture, the equivalent judgment for writing covers what actually counts as proof there.
A five minute routine
- Do not share it yet. Almost all the harm comes from the reshare in the first minute.
- Open the account that posted it and look at its history.
- Reverse image search the full frame, then a crop of the most distinctive object.
- Zoom to the background and the edges, then check text, pairs and shadows in that order.
- Open the credentials panel if the platform offers one, and read a missing panel as no information.
- Check the claim independently of the image, since a real event will have other coverage.
- Say what you found rather than what you concluded. "I cannot find this anywhere before yesterday" is a fact. "This is AI" usually is not.
Over time, expect the visual half of this to matter less and the sourcing half to matter more. That is not a pessimistic prediction, it is a description of what has already happened, and the same shift is under way for AI generated video, where compression destroys the artifacts even faster.
Common questions
Is there an app that tells me if an image is AI?
Several exist and none are dependable enough to settle an argument. They score how likely generation is based on statistical traces, and those traces are weakened by screenshots, compression and re-uploading, which is how most images reach you. Treat the result as one weak input and put your effort into finding the earliest copy instead.
Do AI images have hidden watermarks?
Often, but not in a form you can check. Large providers embed invisible statistical markings readable by their own detection tools, which helps platforms label content automatically. Open models can be run without any marking, and marks can degrade under editing. Treat a detected mark as meaningful and an undetected one as inconclusive.
Can I still spot AI images by looking at hands?
Much less often than you could. Hands were the most publicized flaw, so they received the most attention in training, and they now come out correct in most pictures. A malformed hand is still a strong signal when you find one, but correct hands no longer tell you anything about how an image was made.
What does it mean if an image has no EXIF data?
Almost nothing. Social networks, messaging apps, screenshots and many editors strip EXIF from genuine photographs as a matter of course, often for privacy reasons because it can include location. Present camera data is weak support since it is easy to fake. Signed provenance credentials are the version worth trusting.
How do I check an image on my phone?
Take a screenshot, then use your phone's built in visual search or upload it to a reverse image search site in the browser. Crop to the most distinctive object and search that too. Pinch to zoom into the background and the edges of the frame, which is where generated detail is at its weakest.