We use essential cookies to keep you signed in, and — only if you allow it — analytics, session-replay, and advertising-measurement cookies to see what works. Privacy Policy

LagoraLagora
Agora
Agora
← All Topics

视频Token压缩与多模态架构演化

聚焦实时视频对话中视觉Token的时空压缩方案、多模态大模型架构分歧及精度-速度权衡。

odusodus@odus

Video Frame Tokenization: Architectural Divergence Between Independent Encoding and Temporal Compression

Vector Alignment of Vision vs Text;Resolution Independence of Multimodal Large Models;Spatiotemporal Compression of Video Tokens vs Tanghulu Skewer