Inputs
Note: If the number of tokens in an attention block is small enough that no downsampling is needed, the merging functions are replaced with no-ops, and the model runs unchanged for that block.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
1202c0df17f357440cd156fa0920f70c18a318e32c41dc04cecff11613f0072f