I generate a character sheet for an AI-generated animated show I make the same way every time: three-quarter body shot, front, side, back, one outfit, neutral pose. Clean and consistent, everything an animator would want on a model sheet. And somewhere around shot forty of a scene, the face starts drifting. The jaw gets a little squarer. The eyes sit a little wider apart. Nobody on the crew would notice if I hadn’t stared at every one of those faces for a full day of <a href="https://cinemamix360.com/2026/10/11/still-editing-photos-on-a-monitor-here-are-5-tablets-to-consider-instead/” title=”Still Editing Photos On A Monitor? Here Are 5 Tablets To Consider Instead”>editing.
The real fix is realizing that one reference image is being asked to do two unrelated jobs at once: hold the body’s proportions and hold the face’s identity, and most AI video models only really commit to one of those jobs when the face is small in the frame. A full-body reference photographs a face the way a passport photo photographs a landscape: technically present, functionally useless as the main
Why does an AI video character’s face drift across shots?
Because the model is pulling identity from a low-resolution patch of pixels. On a standard three-quarter body reference, the face occupies maybe 5 to 8 percent of the frame. Every other job the reference is doing, the outfit’s stitching, the belt buckle, the boots, competes for the same fixed token budget the model spends encoding that image. The body wins. The face, which is the one thing a viewer actually tracks shot to shot, gets the leftovers.
This isn’t a guess. Runway’s own Gen-4 References documentation recommends full-body prompting specifically to avoid getting a cropped result, which tells you the model’s default behavior already treats the face as secondary once a full figure is requested. Kling’s Image 3.0 guide makes the same trade explicit from the other direction, telling users a front-facing image gives the highest consistency, advice that only matters if a non-front-facing or partial reference is the common failure mode.
What is a headless character sheet, and why does it work?
The fix that’s been spreading through AI filmmaking circles this year is blunt: stop putting the face on the same panel as the body. Build a reference sheet with a full-body panel that has no visible head detail worth encoding, cropped tight at the neck or shown from behind, and a completely separate close-up panel that exists only to lock the face. One creator building a twelve-minute AI-animated short with identical twin characters, a genuinely brutal consistency test since the model has almost nothing but the face to tell the two apart, put it plainly after testing the technique: it sounds harsh, but it worked better than anything else tried.
The logic holds up once you say it out loud. A full-body panel’s job is proportions, outfit, and silhouette. A face panel’s job is identity. Merging them into one image doesn’t save a reference slot. It wastes one, because the model still has to choose which job to prioritize, and it usually picks the wrong one.
- One full-body panel, front-facing, neutral pose, face cropped out or turned away, locking proportions and outfit only
- One tight face close-up, neutral expression, even lighting, locking bone structure and features
- One or two three-quarter angle shots if the character appears in profile often
- A short written note of anything a still image can’t carry, a limp, a vocal tic, a scar that only shows in certain light
One image can’t hold both jobs at once.
One image can’t answer two different questions at once. Ask it what this person wears, and it answers well. Ask it who this person’s face belongs to from the same image, and it usually guesses.
How many reference images can you actually use?
This is the part most tutorials skip, and it’s the actual bottleneck. Every one of these platforms caps how many reference images a single generation can pull from, and that cap decides whether “just add more references” is even an option.
-
Runway Gen-4 References allows up to three active references per generation. Spend one on the body and one on the face, and there’s exactly one slot left for a prop, a second character, or an environment plate.
goes up to ten images or elements in a single generation, while its video-focused Element Library caps out lower depending on whether a video input is already in play: generally four with video already attached, seven without.
- Seedance 2.0, rebuilt from the ground up in February 2026, accepts up to twelve reference assets across images, video, and audio combined, tagged directly into the prompt.
Three references and twelve are not the same planning problem.
On Runway, splitting one character across two dedicated slots, body and face, is nearly half the entire budget before a second character or a prop even enters the shot. On Seedance, there’s room for face plus body on two characters at once. Decide the reference budget before building the sheet, not after the shot fails.
Not every shot needs both panels fed in. A tight close-up dialogue shot barely needs the body reference at all: feed the face panel alone and save the slot for something else. A wide establishing shot where the character is small in frame barely needs the face panel; the body silhouette and outfit are doing all the visible work. The mistake is treating the reference set as fixed for the whole project instead of picking the right panel or two per shot, the same way a cinematographer picks a different lens for a close-up than for a wide.
Common mistakes worth naming plainly:
- Reusing the exact same two-panel reference set for every shot regardless of framing, wasting reference budget on detail the current shot can’t even see
- Building the face panel at a three-quarter angle instead of straight-on, which several platforms’ own documentation flags as the weaker anchor
- Treating a locked seed as a substitute for a locked reference sheet; a seed controls initial noise, not identity, and won’t hold a face together the way a real reference will
- Skipping the written note entirely and expecting a still image to carry a limp, an accent, or a scar that only reads on camera in motion
Where this actually gets built
On Lost Garden, the group scenes are what forced this discipline. A shot with six hooded characters standing together under the same light source has six identities to hold at once, and there’s no reference budget left to be careless about any single one of them. Splitting each character’s sheet into a body panel and a dedicated face panel, rather than one photo asked to do both jobs, is the difference between a scene that reads as six distinct people and one that reads as six variations on the same generic face.
This is also the exact discipline AI production pipeline’s character stage is built to enforce: a character gets locked once, as a body reference and a face reference, before that character is attached to a single shot. It doesn’t generate the video for you, ScreenWeaver’s production workflow is still rolling out, but the writing and shot-planning stage already forces the sheet to exist and stay attached to every shot that character appears in, so nobody discovers the drift forty shots deep the way I did.
Does a headless character sheet mean literally cropping the head out of every image?
Just the full-body panel. The head is cropped at the neck, shown from behind, or angled away so there’s nothing for the model to mistakenly treat as the face’s, in close-up, on its own
Will this work with only one reference slot available?
Partially. With a single slot, prioritize the face close-up for any shot where the character’s expression or identity is the point, and switch to the body reference for wide shots where the face is small anyway. One slot forces a choice per shot rather than a single reference for the whole scene.
Does a locked seed fix character drift on its own?
No. A seed sets the initial noise pattern for a generation, not the subject’s identity, and the same seed can produce a different face after a model update or a change in sampler. Reference images, not seeds, are what actually anchor a face.
Does this apply to image generation too, or only video?
The same logic applies to any model that accepts multiple reference images, image or video. The face-versus-body budget problem is about how the model allocates attention across a fixed set of inputs, not about whether the output is a still frame or a moving clip.