The GAN is dead; long live the GAN! A Modern GAN Baseline Paper • 2501.05441 • Published 5 days ago • 72
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token Paper • 2501.03895 • Published 7 days ago • 47
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Paper • 2501.00958 • Published 13 days ago • 93
PERSE: Personalized 3D Generative Avatars from A Single Portrait Paper • 2412.21206 • Published 15 days ago • 15
Byte Latent Transformer: Patches Scale Better Than Tokens Paper • 2412.09871 • Published Dec 13, 2024 • 88
ColorFlow: Retrieval-Augmented Image Sequence Colorization Paper • 2412.11815 • Published 29 days ago • 26