Results 231 to 240 of about 152,555 (310)
Some of the next articles are maybe not open access.

LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

European Conference on Computer Vision, 2023
In this work, we present a novel method to tackle the token generation challenge in Vision Language Models (VLMs) for video and image understanding, called LLaMA-VID.
Yanwei Li, Chengyao Wang, Jiaya Jia
semanticscholar   +1 more source

Tokens and Tokenization

2020
The authors have carried out an examination of the status of tokens and tokenization in financial markets with regulatory problems, which lack proposals for solutions based on a generalized consensus. Overall, it seems to authors that cutting the Gordian knot of tokenization and tokens is the essential need to achieve a consensual and efficient ...
Carlos Fernandez-Herraiz   +2 more
openaire   +1 more source

Token Shifting on Graphs

International Journal of Computer Mathematics: Computer Systems Theory, 2021
We investigate a new variation of a token reconfiguration problem on graphs using the cyclic shift operation. A colored or labeled token is placed on each vertex of a given graph, and a “move” consists in choosing a cycle in the graph and shifting tokens by one position along its edges.
Win Hlaing Hlaing Myint   +2 more
openaire   +1 more source

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Neural Information Processing Systems
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs), while its mechanisms are not yet well understood.
Shenzhi Wang   +17 more
semanticscholar   +1 more source

An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

European Conference on Computer Vision
In this study, we identify the inefficient attention phenomena in Large Vision-Language Models (LVLMs), notably within prominent models like LLaVA-1.5, QwenVL-Chat and Video-LLaVA.
Liang Chen   +6 more
semanticscholar   +1 more source

Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

arXiv.org
Recent advancements in large language models (LLMs) have driven significant progress in zero-shot text-to-speech (TTS) synthesis. However, existing foundation models rely on multi-stage processing or complex architectures for predicting multiple ...
Xinsheng Wang   +24 more
semanticscholar   +1 more source

Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More

Conference on Empirical Methods in Natural Language Processing
Vision tokens in multimodal large language models often dominate huge computational overhead due to their excessive length compared to linguistic modality.
Zichen Wen   +7 more
semanticscholar   +1 more source

CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

arXiv.org
Recent years have witnessed a trend that large language model (LLM) based text-to-speech (TTS) emerges into the mainstream due to their high naturalness and zero-shot capacity.
Zhihao Du   +11 more
semanticscholar   +1 more source

Safety Alignment Should Be Made More Than Just a Few Tokens Deep

International Conference on Learning Representations
The safety alignment of current Large Language Models (LLMs) is vulnerable. Relatively simple attacks, or even benign fine-tuning, can jailbreak aligned models.
Xiangyu Qi   +7 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy