CSE DSI Machine Learning Seminar with Eitan Richardson (Lightricks)
Building LTX-2: An Efficient Audio-Visual Foundation Model
LTX-2 is the first widely adopted open-source foundation model for high-quality audio-visual generation, used across industry, academia, and the broader video-generation community. Jointly generating coherent and synchronized audio and video while keeping inference fast and computationally efficient poses significant architectural and modeling challenges. This talk presents the key design choices behind LTX-2, focusing on its asymmetric dual-stream architecture, modality-specific latent representations, and mechanisms for cross-modal interaction.
The second part turns to rich multimodal control through in-context reference tokens, allowing a flexible set of audio, video, and image references to guide generation. This setting exposes a subtle training failure mode: the model can exploit a shortcut that lowers the training loss without learning the intended correspondence between language and references. We show how identifying and preventing this shortcut enables more reliable reference binding.
Eitan Richardson leads a research team in Lightricks’ LTX foundation-model group, developing LTX-2, an open-source foundation model for audio-visual generation. He received his Ph.D. from the Hebrew University of Jerusalem in 2021 under the supervision of Prof. Yair Weiss, focusing on generative models for natural images, and holds a B.Sc. from the Technion. His research has followed the evolution of generative image modeling, from Gaussian mixture models and GANs to diffusion models. Before joining Lightricks, he interned at Google Research, following many years of industry experience in computer vision.