Back to blog
How AI video worksAug 2, 2026 · 7 min read

Why AI Video Still Falls Apart on Handshakes, Text, and Continuity

Every AI video tool shares the same three weak spots: hands, on-screen text, and objects that stay put. All three trace back to how these models learn to build a moving image, and once you see why, you can shoot around every one of them.

Two hands, one confused model

Hands are the classic failure case, and the training data explains why. Footage rarely holds a hand still and in full view for long. Hands move fast, get partially blocked by objects, cross each other, and change shape as fingers curl and extend. A model builds its sense of a hand from millions of blurry, partial, foreshortened glimpses, so the moment it has to render one clearly and hold it steady across ninety-plus frames, it’s working from a weak average. A handshake needs two hands to interlock at a precise angle and keep that grip through motion, exactly the kind of shot with the least reliable training signal behind it. Loose fists and gloved hands tend to hold up better than open fingers spread wide, since a mitten-shaped blob is a far easier average to hit than five independently moving digits.

Text is a language problem wearing a video costume

Ask for a street sign that reads OPEN and there’s a real chance it comes back reading OPFN, or something with no meaning at all. A model paints the shape of letters the same way it paints anything else, as texture, not as characters with one correct form. A brick wall can be approximately brick-shaped and still look right. A word is only ever exactly right or wrong, and nothing underneath the pixels is checking the spelling. Logos fare a little better than sentences, since a model has seen the exact same brand mark thousands of times and can reproduce it as one fixed shape instead of assembling it letter by letter.

Objects forget they existed

Blow out candles in frame one and by frame sixty they might be lit again, or the cake might have gained a slice nobody added. Models generate video in chunks, leaning on a short memory of recent frames rather than tracking every object as a persistent thing across the whole clip. The result behaves more like a flipbook that drifts a little on every page than footage from one continuous take.

The frame-count tax

Every extra second of video is another few dozen frames the model has to keep consistent with everything that came before, and the odds of a small drift compounding into something visible rise right along with it. A two-second shot of a coffee cup on a table is almost always clean. Stretch the same shot to eight seconds and there’s a real chance the handle quietly switches sides, or the steam stutters. Short, cut, repeat beats one long unbroken take most of the time, not because it looks more cinematic, but because it hands the model fewer frames to lose track of anything in.

2 seconds · stays clean

8 seconds · odds of drift climb

Each square is a frame. Short clips hold their line. Long ones accumulate small errors, frame by frame, until something visibly slips.

The trick that actually helps: give it a reference

Most video tools, including the one behind Zo, let you hand the model a still image instead of describing every detail in words. A clean product photo, a screenshot of a specific outfit, a portrait of a face you want to keep steady across shots, all of it locks details that a text prompt alone redraws slightly differently every time. It’s the single highest-leverage move for consistency, and most people never try it because the text box is right there and the upload button isn’t.

z
Zo's tip

Drop a photo in the chat before you describe the outfit or the face. I'll hold onto it as a reference instead of guessing from adjectives.

What to do about it

Play to the strengths instead of fighting the weak spots. Keep hands out of frame, or at the edges, when they aren’t the point of the shot. Skip on-screen text in the generation and add it afterward as a title card. For anything that has to stay consistent, an object, an outfit, a prop, keep the shot short and cut to a new angle rather than asking one continuous take to hold everything steady for ten seconds. Models are better at all three than they were a year ago, but for now, shoot around what the camera can’t hold onto yet.

Say it, and Zo makes it.

Chat with Zo