Skip to main content
The WanImageToVideo node prepares conditioning and latent representations for video generation. It creates an empty latent space for the video and can optionally incorporate a starting image and CLIP vision output to guide the generation. Both the positive and negative conditioning inputs are updated with the provided image and vision data.

Inputs

Note: When start_image is provided, the frame sequence is encoded with the VAE and a mask is applied to the conditioning. The mask is set to 0 for the frames covered by the starting image and 1 for the remaining frames, so generation continues from the provided image. Only the first three color channels (RGB) of the image are used during encoding. Both positive and negative conditioning receive the same concatenated latent image, mask, and (if supplied) CLIP vision output.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 46779f9f2f3da16826b7b547761a96597a3b6b43ce51a9c13367987642f3d5b7