AI AI images, video and voice: how they are made
AI video: how it is made and what still gives it away
How AI video generation works, why it is harder than images, the artifacts that still betray it, and what it means for anything you are asked to believe on video.
The short answer
- AI video is made by denoising a whole block of frames at once, so the model has to keep a scene consistent through time as well as make each frame look right.
- Nothing inside the model tracks objects, identity or forces, which is why failures cluster around occlusion, ground contact, hands in motion and background crowds.
- Compute rises roughly with the square of the number of patches, which is why clips are seconds long and why length is the expensive thing, not resolution alone.
- Every artifact that gets publicized becomes a training target, and platform compression destroys the fine detail you would need to see the remaining ones.
- The durable check is provenance and sourcing: who filmed it, where it first appeared, whether another angle exists, and what the file's signed history says.
- Much harmful video is not generated at all, it is a face swap on real footage or real footage from another event with a false caption.
AI video is made much the same way as an AI image, with one expensive addition: the model has to keep a scene consistent while it moves. A text to video system starts from random static and removes it in stages, but it denoises a whole block of frames together rather than a single picture, so that a face stays the same face and a car keeps traveling in one direction. That extra requirement explains almost everything about these clips: why they are short, why they cost far more than stills, and why what goes wrong now goes wrong in time rather than in space. A bag changes shape after passing behind a pole. A hand gains a finger mid gesture. Feet slide across the pavement instead of pushing against it.
How a video model builds a clip
If you already know how diffusion turns noise into a picture, you have most of this. A compression network squeezes the video down into a much smaller numerical grid, the model repeatedly predicts and subtracts noise from that grid while being steered by your prompt, and a decoder expands the finished result back into pixels.
What changes is what the grid represents. A still image model works on patches of one frame. A video model works on patches that have a duration as well as a position, small cubes of picture covering a handful of frames each, often called spacetime patches. The model looks at all of them together, which is how a decision about the first second reaches the last: when it judges what counts as noise in frame forty, it can see what it has already settled about frames one and eighty.
Three practical consequences follow from that design.
- The clip is resolved all at once, not played forward. The model is not simulating the next moment from the last one, it pulls the whole span out of static together. That is why you cannot usually add one more second without regenerating everything.
- Cost climbs steeply with length. The attention step scales roughly with the square of the number of patches, so doubling the duration more than doubles the work. Hence clips measured in seconds, with longer pieces assembled from several generations.
- The same engine runs from a starting image. Image to video adds noise to a still you supply but stops partway, then denoises forward. Animating a photograph you already have is cheaper and far more controllable than conjuring a scene from words.
Why time is the hard part
There is no scene inside the model. No three dimensional geometry, no list of objects with identities, no physics engine tracking momentum or contact. What the model learned is which pixel patterns tend to follow which, across an enormous quantity of real footage. That carries it a surprising distance, because most motion in real video is smooth and repetitive. It fails in specific, predictable places.
Occlusion is the classic one. Someone walks behind a lamp post. Nothing records that a striped shirt and a shoulder bag went in, so what emerges is whatever looks locally plausible: the stripes shift, the bag thins, a hand turns over. The same weakness appears whenever an object leaves the frame and returns.
Identity drift is the second. Over a long enough clip a face slowly becomes a slightly different face, a logo re-letters itself, the furniture in a room rearranges. Systems fight this with reference conditioning, feeding in a portrait alongside the prompt to anchor the subject, and it is much better than it was. It still degrades the longer the clip runs.
Contact physics is the third. Feet on the ground, fingers closing on a cup, cloth folding against a body, liquid leaving a bottle: all of these involve forces the model never represents, so it approximates them from appearance alone. You get the foot that slides a few centimeters, the fingertip that sinks into the handle, the fabric that moves as though gravity were somewhere else.
The artifacts that still give it away
Watch the clip several times, looking at one thing on each pass, and step through it slowly if your player allows. These failures live between frames, and normal speed playback hides most of them.
| Where to look | What goes wrong | How much it still helps |
|---|---|---|
| Hands in motion | Finger counts change between frames, grips pass through the object being held | Good, when the resolution survives |
| Writing and logos | Letters re-form from second to second, a sign reads differently at the start and end | Strong, text through time is still very hard |
| Ground contact | Feet slide, weight never shifts onto one leg, shadows do not track the feet | Strong and slow to improve |
| Objects after occlusion | Items change after passing behind something or returning to frame | Strong on anything over a few seconds |
| Edges of moving people | Hair and shoulders shimmer or smear where they meet the background | Weakening, face swap tools largely fixed this |
| Blinks and small expressions | Too regular, too rare, or absent on one person in a group | Weak now, and easy to misread in real footage |
| Background crowds | Distant faces melt, limbs merge, extras walk through each other | Good, because the effort goes into the subject |
| Lip sync on speech | Sounds that close the lips, b, m and p, land slightly off the mouth shape | Moderate, and destroyed by re-encoding |
The same reasoning underlies checking whether a still image was generated, except that a video hands you dozens of samples of one scene, so a disagreement between two moments is worth more than any single odd frame.
Artifact spotting is a shrinking advantage
Every giveaway that gets publicized becomes a target. Hands and teeth were the standing joke about generated pictures until enough people complained, at which point they were curated for and evaluated during training, and they improved. The same is happening to gait, hair edges and short bursts of on screen text. Any list of tells, including the one above, is a snapshot of where the weak points sit right now.
Compression works against you as well. By the time a clip reaches your phone it has been uploaded, re-encoded, cropped to a vertical frame, perhaps screen recorded and posted again. That pipeline eats exactly the detail you need: edge shimmer becomes blocky compression noise, and a six fingered hand is forty pixels across.
Then there is the awkward fact that a lot of harmful video is not fully generated. A face swap replaces one region of genuine footage: real camera, real room, real body, real physics, real lighting. Everything you were taught to inspect passes, because most of it was filmed. A real recording with a cloned voice laid over it has the same property, and that technique is covered in how a short audio sample becomes a fake voice. Cheapest of all, and still by far the most common, is authentic footage of a real event carrying a false caption. No generator involved.
Provenance is the durable answer
Because the pixels will stop telling you, the long term answer is a record of where a file came from. That is what Content Credentials are for: a signed manifest attached to the file recording what device or tool made it and what was done to it afterwards, which travels with the video if the platforms that handle it are willing to carry it. The details, including how to read one and how easily it can be stripped, are in how provenance labels on media work.
Some providers also embed an invisible watermark, a statistical pattern in the frames that survives moderate editing and is readable by their own checker. It is not something you can run yourself, and an open model can simply omit it. European transparency rules requiring synthetic content to be marked in machine readable form push the large providers this way, but none of it is retroactive and none of it covers everything.
Treat provenance as asymmetric evidence. A valid credential chain back to a camera is strong support. No credential at all means nothing.
How to check a clip that shocks you
- Wait before you send it on. Nearly all the damage from a fake clip comes from people who reposted within a minute of first seeing it.
- Find the earliest copy, not the one in front of you. Search a distinctive phrase from the caption, sort by oldest, and see whether the account posting it has any history.
- Ask who else would have filmed this. A newsworthy event in a public place produces several angles from several phones. One angle, one account, no second source is the single loudest warning sign.
- Screenshot a clear frame and reverse image search it. That often surfaces the original video, sometimes years old and from another country. Open the file's credential panel too, if the platform shows one.
- Check details that have nothing to do with the pixels: the language on the signs, the license plates, the weather, whether the trees match the season being claimed.
- Notice the emotional shape of it. Material engineered to spread is built to make you angry or frightened in the first two seconds, and the wider method for that is in how to check a viral claim before passing it on.
If someone offers a detector score as the answer, treat it as you would one for writing: the reasons those numbers mislead are in why a detector score is not evidence. Sourcing beats scoring, and it will keep beating it as generators improve.
Common questions
How long can AI generated video be?
A single generation is usually a few seconds up to a few tens of seconds, because the work grows faster than the length does. Longer pieces are built by chaining generations, feeding the last frame of one clip in as the starting image for the next, which is why you often see a subtle change in lighting or face at the joins.
Can AI video generate matching sound?
Some systems now produce audio alongside the picture, and others leave the clip silent so that music, effects or a synthetic voice are added afterwards. When sound is added separately, the mismatch between lip movement and consonants is often the most noticeable flaw, particularly in speech shot close up.
Is a deepfake the same thing as AI video generation?
Not quite. Generation builds a clip from scratch out of noise. A deepfake in the original sense swaps one face onto existing footage, so the camera work, the body and the background are all genuine. The second kind is easier to make, harder to spot, and responsible for most of the real harm.
Does a video without any metadata mean it is fake?
No. Metadata and provenance data are stripped routinely by messaging apps, social platforms, screenshots and screen recordings, so most authentic video you encounter has none. Present credentials can support a claim of authenticity. Absent credentials tell you nothing in either direction.
Can I tell by slowing the video down?
Often yes, and it is the cheapest test available. Stepping frame by frame exposes inconsistencies the eye smooths over at full speed: a finger that appears and goes, a shadow that lags, a sign whose letters rearrange. It works best on high quality uploads, since heavy compression hides the same detail you are hunting for.