I have been messing around with generated videos for a while, and did a bit of research trying to figure out why clips that look great on their own can still feel off when you put them in sequence.
It's mostly not the obvious stuff people used to complain about.
- The physics of small things,
Faces and bodies are getting pretty convincing, but smaller physical details still give away constantly.
Hair, clothes, water, smoke, loose objects, etc. are things that our brains are wired to notice when they're not behaving correctly. A coat that looks like it has no weight to it is the easiest tell.
Specifing material, weight and how it should move instead of just describing how it looks tends to somewhat do the trick. (Here's a CogVideoX tech paper on unnatural physics.)
- The camera has no operator,
Real hand held footage has breathing, hesitation, little corrections and changes in speed. Generated hand held gets the general movement right but smooths out the small human imperfections.
For example, a clip that's supposed to feel like a person running away from an SCP or something with a camera ends up feeling more like a gimbal or drone.
Can't figure out how to fix at generation, so I just add actual handheld movement in post.
- Everything sits at roughly the same distance from the lens,
This one kept showing up in my own generations.
You can have technically different shots, but they all end up with roughly the same framing and subject size, the sequence starts feeling weird before I can explain why.
Real coverage tends to jump between wides, mediums, close-ups, inserts, reaction shots etc. Six medium-ish shots in a row feels ... Wrong.
Specifying the shot size everytime tends to alleviate it.
(Link here for what I researched through)
- Light does not carry between shots,
This one's brutal.
Two shots can look completely believable on their own, but if one has the key light/sun coming from the left and the next cut shows it coming from somewhere completely different, it immediately seems artificial.
There are actual papers researching this so I am apparently not the only person annoyed.
Mostly seems like a workflow problem imo. Locking the time of day, general light direction and quality then keep repeating it seems like a good choice.
- Uniform shot length,
Everything comes out around the same length that the tool gives you, so it's really easy to just use the whole clip everytime.
Do that five or six times in a row and suddenly! You're watching a slideshow instead of a sequence.
Shot duration and editing rhythm obviously aren't an AI specific problem, but there's plenty of film research on how much they affect pacing and perception.
This one is completely fixable in post, but people really underestimate how much it matters.
(Link!)
These are just things I have started noticing the most. I'm curious as to see what others have found.
What's your AI-Video tell that you just can't unsee?
[link] [comments]