Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Scaling up autoregressive models in vision has not proven as beneficial as in large language models. In this work, we investigate this scaling problem in the context of text-to-image generation, focusing on two critical factors: whether models use discrete or continuous tokens, and whether tokens ar...

Full description

Saved in:
Bibliographic Details
Published inarXiv.org
Main Authors Fan, Lijie, Li, Tianhong, Qin, Siyang, Li, Yuanzhen, Chen, Sun, Rubinstein, Michael, Sun, Deqing, He, Kaiming, Tian, Yonglong
Format Paper
LanguageEnglish
Published Ithaca Cornell University Library, arXiv.org 17.10.2024
Subjects
Online AccessGet full text

Cover

Loading…