Production

Keeping the same character across videos

It is the most visible flaw in generated video: the character in scene 3 does not look like the one in scene 1. The problem is not the model — it is how identity gets handed to it.

Why the face changes

A generative model remembers nothing between calls. Each generation starts from the text you provide. If the character’s identity is described only in words — "a brown-haired woman in her thirties" — the model reinvents a plausible face every time, drawn from the enormous space of faces matching that description.

No amount of textual detail fully solves this. Adding "green eyes, mole on the left cheek" narrows the variation space without closing it. Text is too loose a constraint to lock an identity.

Three approaches, ranked by reliability

In practice three methods coexist, with very uneven results:

  • The frozen prompt. Reuse exactly the same description. Simple, free, and insufficient the moment angle or lighting changes.
  • The visual reference. Provide one or more images of the character so later generations conform to them. Considerably more stable, provided the references cover several angles.
  • The persistent character. The reference is stored once, versioned, and reused automatically across every later production, with an adjustable identity strength.

Why storage changes the nature of the problem

The first two approaches put the burden back on the operator: find the right references, resend them on every generation, and hope nobody on the team uses a different version.

A library of persistent characters moves that responsibility into the system. In Viffly, a generated character is stored in your studio and reusable in later productions; the same logic applies to products, logos and brand palettes. Consistency becomes the default state instead of a discipline to maintain.

This matters most in automated production: a workflow generating twenty videos a week cannot depend on a human checking that the face stayed right.

What remains hard

It is worth being honest about the limits. Extreme angle changes, group scenes and very wide shots remain where likeness degrades most, whatever the method.

The practical workaround is to chain progressively similar shots — same framing type, same lighting — rather than cutting hard between close-up and wide, and to reserve your most precise references for shots where the face fills the frame.

Frequently asked questions

How many reference images does a stable character need?+

One is enough for shots close to the original framing. For varied angles, several references covering different orientations give noticeably more stable results.

Does consistency work for products too?+

Yes, and it is often more reliable than for faces: a product packshot acts as a reusable reference across every later scene.

Does a character stay consistent across two different Reels?+

Yes, as long as it is stored in the studio rather than re-described for each production. That is exactly what a character library is for.

Put it into practice

Read next