Results 261 to 270 of about 152,555 (310)
Some of the next articles are maybe not open access.

Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

International Conference on Learning Representations
Scaling up autoregressive models in vision has not proven as beneficial as in large language models. In this work, we investigate this scaling problem in the context of text-to-image generation, focusing on two critical factors: whether models use ...
Lijie Fan   +8 more
semanticscholar   +1 more source

Unified Autoregressive Visual Generation and Understanding with Continuous Tokens

arXiv.org
We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture processes multimodal image and text inputs, generating discrete tokens for ...
Lijie Fan   +13 more
semanticscholar   +1 more source

Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens

arXiv.org
Are $n$-gram language models still relevant in this era of neural large language models (LLMs)? Our answer is yes, and we showcase their values in both text analysis and improving neural LLMs.
Jiacheng Liu   +4 more
semanticscholar   +1 more source

Rho-1: Not All Tokens Are What You Need

arXiv.org
Previous language model pre-training methods have uniformly applied a next-token prediction loss to all training tokens. Challenging this norm, we posit that"9l training".
Zheng-Wen Lin   +10 more
semanticscholar   +1 more source

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Computer Vision and Pattern Recognition
Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation by combining LLM and diffusion models, the state-of-the-art in each task, respectively. Existing approaches rely on spatial visual tokens, where
Kaihang Pan   +8 more
semanticscholar   +1 more source

Grounded 3D-LLM with Referent Tokens

arXiv.org
Prior studies on 3D scene understanding have primarily developed specialized models for specific tasks or required task-specific fine-tuning. In this study, we propose Grounded 3D-LLM, which explores the potential of 3D large multi-modal models (3D LMMs)
Yilun Chen   +7 more
semanticscholar   +1 more source

MaskBit: Embedding-free Image Generation via Bit Tokens

Trans. Mach. Learn. Res.
Masked transformer models for class-conditional image generation have become a compelling alternative to diffusion models. Typically comprising two stages - an initial VQGAN model for transitioning between latent space and image space, and a subsequent ...
Mark Weber   +6 more
semanticscholar   +1 more source

The token distribution problem

27th Annual Symposium on Foundations of Computer Science (sfcs 1986), 1986
Summary: A solution to the following fundamental communication problem is presented. Suppose that n tokens are arbitrarily distributed among n processors with no processor having more than K tokens. The problem is to specify a bounded-degree network topoloy and an algorithm that can distribute the tokens uniformly among the processors. The first result
David Peleg, Eli Upfal
openaire   +1 more source

DyCoke : Dynamic Compression of Tokens for Fast Video Large Language Models

Computer Vision and Pattern Recognition
Video large language models (VLLMs) have significantly advanced recently in processing complex video content. Yet, their inference efficiency remains constrained because of the high computational cost stemming from the thousands of visual tokens ...
Keda Tao   +4 more
semanticscholar   +1 more source

ImageFolder: Autoregressive Image Generation with Folded Tokens

International Conference on Learning Representations
Image tokenizers are crucial for visual generative models, e.g., diffusion models (DMs) and autoregressive (AR) models, as they construct the latent representation for modeling.
Xiang Li   +6 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy