NNSFWAITool
English

Production Case Study

How to Make an AI Adult Short Film with Wan 2.2: A 117-Shot Workflow

AI filmmaking edit suite with a dense shot timeline and local GPU workstation

The short answer: the interesting part of Huangguo is not that AI was involved. It is that the filmmakers did not ask one model to create an eight-minute film from one prompt. According to the production notes supplied for this case study, they divided the film into 117 short shots and controlled identity, keyframes, motion, dialogue, and finishing as separate problems. Wan 2.2 was the foundation because the project needed local weights, a modifiable pipeline, and sustained batch production—not because Seedance lacks visual quality.

About the evidence: the eight-minute runtime, 117-shot count, 3.2-second median shot length, and project-specific process come from Huangguo production material supplied for this article. External links verify public Wan, Seedance, Wan-S2V, and NVIDIA capabilities; those companies have not verified the film's private production log.

This article covers lawful, consent-based production involving clearly adult characters only. It does not endorse age-ambiguous content, minors, or non-consensual intimate deepfakes.

8 minutestarget finished runtime
117 shotsgenerated and processed separately
3.2 secondsmedian shot duration

阅读中文版:《黄果》的117镜头Wan 2.2完整工作流

Why Wan 2.2 instead of Seedance?

Seedance is not the weak model in this comparison. ByteDance's official page highlights image, audio, and video references plus control over performance, lighting, and camera movement. For a conventional project that wants polished hosted capability without maintaining local infrastructure, that can be a major advantage.

Huangguo had a different constraint set. More than one hundred shots needed repeated sampling. Character identity, movement, and dialogue had to be adjusted independently. A failed shot had to be recoverable from any stage. The studio also needed to own its training assets and generation queue. Wan 2.2 publishes model weights and inference code, with T2V, I2V, TI2V, and S2V paths that can be integrated into a private render system.

Production requirementLocal Wan 2.2 workflowPublic Seedance service
Model accessDownloadable weights and inference codeFinished capability through hosted products or APIs
Local inferenceOfficial local run instructionsPublic model page does not offer downloadable base weights
LoRA and targeted trainingCan connect to community LoRA/full-training toolsBase-model training stack is not exposed through the public interface
Batch productionSelf-managed queues, caches, and multi-node schedulingSubject to API quotas, pricing, and service rules
Adult contentLocal execution; the producer remains responsible for law, consent, and safetyVolcano Engine states that its safety system detects and blocks pornographic risk
Best fitTrainable, reversible private production lineConvenient multimodal generation for policy-compliant work

This is a workflow decision, not a universal image-quality ranking. Seedance offers a strong managed service. Huangguo needed ownership of weights, data, and scheduling logic. Volcano Engine's own safety-creation page also says its platform controls clearly non-compliant input and output, with specific attention to pornographic risk, which rules it out as the central generator for this case.

Why split eight minutes into 117 shots?

Longer generations multiply the number of things that can drift: faces, hands, clothes, background geometry, contact between characters, lighting, and camera movement. A small error in a complex interaction can compound across seconds. Short shots are not automatically faster-paced; they make each render testable.

Shot-length distribution: 84 shots at five seconds or less, 23 at five to eight seconds, 4 at eight to ten seconds, and 6 over ten seconds
84 plus 23 equals 107, so shots of eight seconds or less represent about 91%, not the 92% stated in an early draft of the production notes.

A three-to-eight-second shot can be assigned one job: an eyeline, a turn, one sentence, one movement phase, a change of angle, or a detail insert. Complex scenes can be rebuilt from wide shots, close-ups, inserts, and reactions. The sense of continuous performance is created in the edit instead of being demanded from one generation.

The production line starts with executable shots

Six-stage workflow from script and shot list through identity assets, keyframes, Wan generation, motion and dialogue, finishing and quality control
The value of the pipeline is not its complexity; it is that each stage can be rerun without rebuilding the whole scene.

1. Turn the script into a production database

An LLM can help organize plot, relationships, dialogue, and scene order. Its prose draft is not a usable render task. Each shot needs an ID, set, cast, framing, camera position, action, emotion, dialogue, duration, opening state, ending state, model route, and any required motion or audio reference.

“Two characters complete a complex interaction in a room” is too vague. The scene must be divided into entrance, approach, reaction, position change, detail, and end state. That makes failures diagnosable: composition, identity, and movement no longer collapse into the same problem.

2. Route drama, dialogue, and difficult motion separately

The production notes classify shots as regular narrative, dialogue, adult content, complex motion, environment/transition, or finishing/VFX. Each category gets its own models, adapters, sampling parameters, references, and acceptance tests. One universal preset looks efficient until retries and repairs erase the saving.

3. Make the character LoRA answer only “who is this?”

Lead characters need a clean identity pack: front, profile, three-quarter views, expressions, lighting changes, half-body and full-body views, and recurring wardrobe states. Age, face shape, makeup, and proportions must remain consistent across the data; otherwise the adapter learns a vague average.

Identity and specialized movement are trained separately. The character adapter establishes who is in frame. A content or motion adapter describes the required performance. That modularity lets the studio replace a character without retraining the entire motion stack, or improve movement without changing the performer's face.

One technical distinction matters: Wan's official release supplies weights and inference code. LoRA training generally comes through tools such as DiffSynth-Studio, which the official repository lists as a community integration supporting LoRA and full training. It is not a button built into a hosted Wan product.

4. Lock recurring sets before rendering

Bedrooms, living rooms, hotels, and hallways need reference packs of their own. Door and window positions, furniture orientation, key light, color temperature, wall material, and usable camera positions should be fixed. The more a performance interacts with beds, sofas, or walls, the more obvious spatial drift becomes. A set template is both an art-direction reference and a coordinate system for later pose and depth control.

Build the correct still image before asking it to move

Regular narrative shots begin as approved keyframes. The still establishes identity, expression, wardrobe, composition, light, set, and the starting pose. More complicated shots also solve the number of characters, relative position, body orientation, occlusion, environmental contact, camera angle, and frame boundary before video generation begins.

Every keyframe should be reviewed manually. Faces, eyes, hands, limb connections, proportions, clothing edges, furniture contact, reflections, background text, and extra objects are cheaper to fix once in the source image than across dozens of video frames.

The three control layers for specialized shots

Layer 1: the keyframe locks identity, pose, and camera

A purpose-built image model and adapters create the approved opening frame. This layer answers who is present, where they are, and how the scene is framed. It leaves less room for the video model to redesign the relationship between subjects.

Layer 2: Wan 2.2 I2V/TI2V handles time

The approved frame enters a Wan 2.2 I2V or TI2V route with the targeted adapters or fine-tune required by that shot. This layer handles motion continuity, identity retention, texture persistence, camera movement, and timing. Wan's official documentation says TI2V-5B supports both text-to-video and image-to-video at 720p and 24fps, and can run on a consumer GPU such as an RTX 4090.

Layer 3: licensed references describe how the action moves

Difficult movement can add authorized reference video, skeletal pose sequences, or depth information. Those inputs constrain trajectory, speed, weight shift, orientation, distance, and start/end state. Identity adapters control the performer; references control movement; the set template constrains space. Separating those responsibilities reduces character switching, body intersections, and unstable contact.

Regular shots, dialogue, and the 24-to-30fps finish

For regular shots, editability beats spectacle

Generate several candidates, then check identity, storyboard compliance, eyeline, wardrobe continuity, set continuity, and whether bad frames can be cut cleanly. A restrained shot that survives from first frame to last is more valuable than an ambitious camera move that mutates halfway through.

Dialogue: Wan-S2V or a separate lip-sync pass

The official Wan-S2V project describes a model that turns one image plus audio into synchronized video with natural expressions, body movement, and cinematic camera behavior. A useful dialogue shot still involves more than mouth shapes. Breath before speech, pauses, eye movement, posture, and the reaction after the line make the performance believable. When S2V is not the right route, the team can generate the visual shot first and apply lip sync in post.

24fps to 30fps: duplicated frames are not interpolation

The production notes say the Wan source clips were generated at 24fps and conformed to a 30fps master. A basic duplication conversion adds six output frames per second, or one repeated frame in every five output frames. Optical-flow or AI interpolation instead creates synthetic in-between frames and has a different artifact pattern.

Either method needs shot-level review. Hands, hair, overlapping figures, and fast motion can produce new structural errors between otherwise usable source frames. Finishing also included upscaling, deflickering, denoising, color matching, exposure correction, local repainting, bad-frame removal, and light stabilization.

Sound and editing turn 117 clips into one film

Completing video generation only means the raw material exists. TTS dialogue, ambience, action sound, music, transitions, mixing, noise reduction, and loudness normalization create auditory continuity. Subtitles, messaging graphics, system interfaces, titles, and restrained VFX carry the narrative packaging.

The final edit must preserve screen direction, eyelines, wardrobe, set lighting, action at the cut, lip sync, and sound timing while removing every failed frame. AI produced 117 separate pieces. Editing made them feel like an eight-minute short.

What hardware does this workflow require?

24GB VRAM: enough to validate, not necessarily to deliver at scale

Wan's official TI2V-5B command can run with at least 24GB of VRAM by using model offloading, dtype conversion, and CPU placement for T5. That is useful for character tests, prompt development, keyframes, and limited clips. A 117-shot delivery is a throughput problem: every approved shot may sit behind several rejected attempts.

Dual GPU or multiple nodes: parallel queues matter most

NVIDIA lists 96GB of GDDR7 ECC memory and 600W maximum board power for the RTX PRO 6000 Blackwell Workstation Edition. Two cards therefore represent 192GB aggregate VRAM and 1,200W of maximum board power; four represent 384GB and 2,400W before CPU, storage, memory, and cooling.

Aggregate VRAM is not automatically one shared memory pool. Two 96GB cards do not inherently behave like a single 192GB card. Software must use data parallelism, model parallelism, or separate render jobs to benefit. For 117 shots, distributing independent jobs across several nodes is often more flexible than building one extreme workstation.

Shot-level quality and rights checklist

  • Every depicted character is explicitly defined and presented as an adult.
  • Faces, voices, motion references, and training data have documented authorization and provenance.
  • No real person appears in non-consensual intimate deepfake content.
  • Identity, face, and body structure match the approved character design.
  • There are no fused limbs, intersections, unexplained anatomy, or physically incoherent movement.
  • Wardrobe, set geometry, lighting, and screen direction continue across adjacent shots.
  • Flicker, texture failures, false text, and destructive interpolated frames are removed.
  • Dialogue, facial performance, and action sound are synchronized.
  • Frame rate, resolution, color, and loudness meet the master specification.
  • The output follows the law and platform rules in both production and distribution locations.

Frequently asked questions

Why is Wan 2.2 a better fit for this adult AI short film?

It is not automatically better for every film. Huangguo needed downloadable weights, local inference, external training tools, and self-managed queues. For policy-compliant conventional content, Seedance's hosted multimodal workflow may be simpler.

Why not use Seedance?

Seedance is publicly offered as a hosted generation service, and Volcano Engine documents safety checks that include pornographic risk. Its public interface also does not expose the base-model training stack to end users. That makes it incompatible with this particular adult-content pipeline, without diminishing its strengths for ordinary filmmaking.

Can an eight-minute AI film be generated in one pass?

A shot-based workflow is currently easier to control and repair. Huangguo used 117 independently reviewed shots, then relied on editing and sound to build continuity.

Do Stable Diffusion or Flux still matter?

Yes. Image models remain useful for character turnarounds, set references, and keyframes. Video models and post-production handle motion, dialogue, and temporal continuity.

Should character and content LoRAs be merged?

Usually not. Separating identity from motion or content capability makes replacement, retraining, and debugging easier. The final design still depends on dataset size and training method.

Can a personal computer do this?

It can start with TI2V-5B on a 24GB GPU. “Can render a clip” and “can deliver 117 approved shots on schedule” are different thresholds, however. Storage, retries, queue management, and human review become equally important.

The real barrier is a private production line

The Huangguo method can be summarized as follows: structure the story as executable shots; lock identity with character adapters; maintain space with set templates; approve keyframes; animate short clips with Wan 2.2 and targeted variants; guide difficult motion with licensed references; route dialogue through S2V or lip sync; then finish sound, color, interpolation, and editing on a 30fps timeline.

Seedance represents powerful hosted video generation. Wan 2.2 provides weights, inference code, and more room to modify the system. For this lawful, authorized project—with its emphasis on privacy and repeated batch production—that difference was decisive. The comparison is intended to help readers judge which route matches their own content, budget, technical skill, and compliance obligations, not to hand down one universal “best model.”

Official sources

Keep this guide accurate

Found a changed model or specification?

Model versions, hardware specifications, and hosted-service rules can change. Send the source and we will review the update.

Report an update