This talk explores how transformers can effectively be trained on compressed byte streams in order to gain computational and memory efficiencies.

Overview

While compressed formats are essential for efficient data storage, their potential for representation learning is underexplored. We introduce TEMPEST, a framework that trains standard transformers directly on compressed byte-streams, entirely bypassing raw media decoding. This approach significantly reduces token requirements, yielding substantial computational and memory efficiencies while maintaining accuracy across diverse modalities.

Presenters

Brief Biography

Juan Carlos Alcazar is a research scientist in Prof. Bernard Ghanem's research team. Juan Carlos is a Computer Vision and Machine Learning researcher, currently working in association with Image and Video Understanding Laboratory (IVUL) at KAUST. His research primarily bridges video understanding and multimodal learning. He is dedicated to advancing the frontier of robust, multimodal AI systems and solving novel visual computing challenges.