Inputs
Note: At least one view input must be provided for the node to function. The node only processes views that contain valid CLIP vision output data and skips views that are not connected. Each view receives a fixed positional encoding based on its slot (front, left, back, right), and the processed embeddings from all provided views are joined together along the sequence dimension.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
1492b51661d0bb8f2c142c1b1e8ef104beed1b9dae532a970e2928e27ad71d69