Back to blog

Hotel Lobby AI Face Consistency: Keep the Two People Distinct

To evaluate Hotel Lobby AI face consistency, keep the original portraits beside the complete result. Follow each person through turns, gestures and the ending. Record where a face changes or borrows the other person's features before deciding what to change. Two separate uploads identify the intended cast; they do not guarantee stable identities in every frame.

The example below is a real historical output made from fictional synthetic portraits. We inspected three frames extracted from its original file. This is an observation of one run, not an experiment proving that any input change fixes face drift. Current models can behave differently.

Keep the input identities beside the result

This friends case was generated on October 4, 2026 using Wan 3.0 with original AI audio. The recorded output measures 480 × 854 pixels and approximately 5.04 seconds. It has not been relabeled as a current-model result. The case record contains the topic and original settings.

Synthetic left reference for the recorded friends case, a man in a cream shirt over a navy top

Left input: fictional man, cream shirt and navy top.

Synthetic right reference for the recorded friends case, a woman in a terracotta jacket

Right input: fictional woman, curly hair and terracotta jacket. Both references were created for the demonstration.

The submitted topic was:

Two adult friends celebrating their Friday reunion, playful and upbeat

Use the portraits to assess identity, and the topic to assess content. A consistent jacket alone does not establish an exact facial match. Look at the hairline, eyes, face shape and each person's position through movement.

What do three sampled frames show?

The frames below come from the unchanged five-second output at approximately 0.5, 2.5 and 4.5 seconds. The local source file's SHA-256 prefix matches the public case video's identifier. No face retouching, generated replacement or before/after edit was applied to these images.

Original friends output around 0.5 seconds: cream-shirt performer left, terracotta-jacket performer right

About 0.5s: both people are separate. The left performer's moving hand is blurred; the faces remain unobstructed in this frame.

Original friends output around 2.5 seconds as the left performer turns toward the right performer

About 2.5s: the left person turns toward the partner. The dark hair, curly hair and two outfit colors remain distinguishable.

Original friends output around 4.5 seconds with both performers closer together beneath the microphone

About 4.5s: the pair move closer together. Both faces remain visible and assigned to separate people in this sampled frame.

These observations do not establish a similarity score or rule out brief changes between the samples. A still frame cannot establish lip synchronization, audible wording or a natural ending. Watch the complete result as well.

GENERATED DEMO October 4 friends demonstration; Wan 3.0, approximately 5.04 seconds, 480 × 854. Three frames above were extracted from this original output.

Distinguish face drift, blending and wrong sides

Face drift means one person's appearance changes during the clip. Record the timestamp and the feature that changed, such as the hairline or face shape. Repeatedly judging a cover image will not identify when the change occurred.

Blending means features from the two references appear combined or one identity takes over both people. Inspect moments where heads, hands or bodies overlap. Those moments are useful review targets; this article has not tested that removing overlap guarantees a correction.

Wrong sides can start with an unintended upload assignment. Check the thumbnail labels and Swap control before generating. If the starting sides are correct but change later, record the first crossing or identity switch in the output.

Use the photo-selection guide for sharpness, lighting and crop preparation. This page focuses on the video timeline rather than repeating the upload checklist.

Make the next comparison answer one question

If the reference hides a face, try one clearer portrait while keeping the partner, topic and available output settings the same. Keep both files and the quoted settings so you can inspect the same moment. If you change model, photos, topic and framing together, you cannot identify which change affected the result.

For a shared photo with a tiny or obstructed face, separate portraits can make the intended cast easier to inspect. That is a setup choice, not a measured reliability advantage from this example. We did not generate a shared-photo version of the same pair.

Before another paid attempt, write a short review record:

  • Which person changed, and at what time?
  • Which visible feature changed?
  • Was the issue present in the upload reference?
  • Which single input will change in the next request?
  • Which model, sound mode, length and framing will stay comparable?

Changing the input and generating another version starts a new quote. Refreshing a task or downloading its original result does not regenerate it. Use troubleshooting if the problem is task status or media access rather than identity.

Does keeping each person on the same side guarantee likeness?

No. Side assignment identifies the intended cast, but facial detail can still change through motion. Evaluate each face against its own input.

Do higher resolution and longer clips fix blending?

They change the request and cost; this case does not demonstrate that they fix blending. Choose those settings for your output needs and inspect the actual result.

Does stable identity prove lip sync?

No. A recognizable face can still have mistimed mouth movement. Review the sound and both mouths separately, using the intended audio mode. The workflow guide explains original voices versus reference recordings.

Sources and evidence limits

  • Friends case record, generated October 4, 2026; synthetic inputs and the complete provider output.
  • Original file checked October 6, 2026: SHA-256 prefix ed441fb0ad566846, 3,341,936 bytes, 480 × 854 pixels, 5.038005 seconds, with an audio stream. An audio stream does not itself prove audible quality.
  • Three extracted frames visually inspected October 6, 2026. No new render, input comparison, face-similarity metric or current-model quality benchmark was performed.