Back to blog

Hotel Lobby AI Pets: Make a Dog-and-Cat Rap Video

A Hotel Lobby AI pets video uses the same photo setup as a human duo: choose the intended left and right subjects, describe the idea and generate a complete performance. Our October 6, 2026 example uses a fictional golden retriever and a tabby cat. Seedance 2.0 Mini returned a roughly ten-second orange-studio clip with native generated music and vocals.

The complete result below is an actual API output. It shows what this pair and topic produced once, rather than proving that every pet photo will animate well. Inspect coat markings, ears, mouths and paws before posting your version.

Start with two recognizable pet photos

The golden retriever is the left input; the gray tabby with a white chest is the right input. Both references were generated with imagegen for the site, so no customer's pet photos are used here. The same reference files were retained when replacing the site's older Wan demonstration with this current-model result.

Fictional golden retriever with a red collar, assigned to the left
Left reference: golden retriever, red collar.
Fictional gray tabby cat with a white chest and blue collar, assigned to the right

Right reference: gray tabby, white chest and blue collar.

For your own references, select sharp pictures with the muzzle, ears and eyes visible. A face hidden behind a toy or cropped out of the image gives you less information to check in the result. Avoid a tiny animal inside a large group scene when you need a recognizable face. These are practical input choices, not advantages measured by this one example.

Two separate photos let you inspect each animal's role clearly. The creator also offers one shared photo with both performers, but we did not generate a shared-photo version of this pet pair. The photo guide explains the two modes.

Copy the actual topic and settings

This request submitted:

A golden retriever and a tabby cat celebrate their friendship with a playful original rap.

Keep the topic short enough to state the subjects, relationship and mood. It directs original generated content; it does not require the model to sing that sentence verbatim. A pet name or favorite activity can personalize another request, but this run did not test name pronunciation.

| Field | Recorded pet request | | ----------------------- | --------------------------------------------- | | Model and provider | Seedance 2.0 Mini through Kie | | Inputs | Two separate fictional pet photos | | Scene | Hotel Lobby orange background | | Sound | Native original AI vocals and music | | Requested output | 10 seconds, 480p, 9:16 | | Measured output | 496 × 864 pixels; approximately 10.08 seconds | | Extra processing | None; original supplier file | | Supplier receipt | 38 supplier credits; 124 seconds | | Equivalent public quote | 40 website credits, checked October 6 |

The supplier charge and site quote use separate credit systems. This direct paid API case deducted zero website account credits and does not demonstrate a free generation for a new customer. The supplier's time is specific to this request. Review the free and paid boundaries and your live quote before another attempt.

GENERATED DEMO Actual October 6 pet case: approximately 10.08s, 496 × 864. Original provider picture and generated audio retained.

What the returned pet video shows

At approximately 0.5 seconds, the dog raises a front paw while the cat remains seated. At 5.0 seconds, the cat has a widely open mouth while the dog's mouth is relaxed. At 9.5 seconds, both face nearer the camera with closed mouths. The dog stays left and the cat stays right in these samples; red and blue collars remain separate.

Pet result at 0.5 seconds with the dog raising a front paw
0.5s: lifted dog paw; cat seated.
Pet result at 5 seconds with the cat's mouth open and dog's mouth relaxed
5.0s: distinct mouth states.
Pet ending sample with both animals facing nearer the camera
9.5s: both subjects remain visible.

The native file contains an audio stream covering the complete video. A local automatic transcript recognized a couch-to-park friendship theme and a pet-family party phrase. This is a draft from a speech model, not a human-certified lyric transcript. It helps check topic relevance, but cannot certify which animal delivers each syllable or whether every mouth movement matches it.

Pet mouths and paws can look unusual during animation. Three stills cannot rule out short distortions between them. Play the complete file with sound and inspect the ending. We did not score animal anatomy, exact lip sync or the number of successful attempts across different pets.

Review anatomy separately from identity

Follow each animal's face and coat through motion. For identity, check the dog's muzzle and ears or the cat's stripes and white chest. For anatomy, inspect lifted paws, placement of visible limbs, teeth and mouth opening. A familiar collar can remain even when a paw looks wrong, so accessories alone do not establish a good result.

Save the timestamp of any problem before choosing another attempt. A clearer input is a reasonable next experiment when a feature is hidden. It is not a guaranteed fix. Changing the photo, model, framing and topic together makes the next output harder to interpret. The face-consistency comparison shows how to retain a baseline and change one reference file.

Choose original audio or a later music edit

This case generated new music and vocals with the picture; it did not use a HOTEL LOBBY recording or an input performance video. It is a new pet scene rather than a measured recreation of the source choreography.

If you replace the audio in an editor, keep a copy of the native result. Removing one soundtrack and adding another changes the saved audio, but leaves existing mouths in place. The same-picture lip-sync experiment provides two files that differ only in audio. Use it before treating a recognizable song as evidence of synchronized singing.

For a birthday version, add the recipient and one light detail, then review what was actually sung. The birthday case records a personal-topic request and explains why a written name is not a delivery guarantee. For a simple animal duet, the topic above is sufficient to start.

Can I use my own dog and cat?

Use suitable images that you are entitled to upload. Assign the animals to the intended slots, inspect the thumbnails and review your quote. This synthetic demonstration does not predict fidelity for every real pet.

Does the AI clone my pet's voice?

No personal voice sample was supplied. Original-song mode generates its own vocals; photos supply visual references.

Do I need a generator for each animal?

This case used one request with two references. No separate dog or cat generation was combined afterward. A new reference or regeneration is another attempt using the current quote.

Which format should I choose for a Short?

A portrait composition is a useful starting shape. Inspect the whole performance so neither animal is cut off. The requested ratio can be rounded to codec-aligned pixels; this 9:16 request returned 496 × 864. The publishing guide covers cropping and playback.

Sources and evidence limits

  • Pet case record, saved API settings, receipt and complete file, checked October 6, 2026.
  • Original imagegen references and three extracted frames inspected. No retouching was applied to output frames.
  • Native audio coverage checked from media packets; automatic transcription used as a draft. No human audio certification or pet success-rate benchmark was performed.
  • The older Wan result is preserved in the archive. The embedded video is the new Mini case.