Results 251 to 260 of about 2,403,325 (304)
Some of the next articles are maybe not open access.

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

International Conference on Machine Learning
Harnessing the power of human-annotated data through Supervised Fine-Tuning (SFT) is pivotal for advancing Large Language Models (LLMs). In this paper, we delve into the prospect of growing a strong LLM out of a weak one without the need for acquiring ...
Zixiang Chen   +4 more
semanticscholar   +1 more source

ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

Robotics
Vision-Language-Action (VLA) models have shown substantial potential in real-world robotic manipulation. However, fine-tuning these models through supervised learning struggles to achieve robust performance due to limited, inconsistent demonstrations ...
Yuhui Chen   +5 more
semanticscholar   +1 more source

Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

International Conference on Machine Learning
Training models to effectively use test-time compute is crucial for improving the reasoning performance of LLMs. Current methods mostly do so via fine-tuning on search traces or running RL with 0/1 outcome reward, but do these approaches efficiently ...
Yuxiao Qu   +6 more
semanticscholar   +1 more source

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

arXiv.org
Reinforcement Learning (RL) benefits Large Language Models (LLMs) for complex reasoning. Inspired by this, we explore integrating spatio-temporal specific rewards into Multimodal Large Language Models (MLLMs) to address the unique challenges of video ...
Xinhao Li   +9 more
semanticscholar   +1 more source

SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

arXiv.org
Large language models (LLMs) have achieved remarkable progress in reasoning tasks, yet the optimal integration of Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) remains a fundamental challenge.
Yuqian Fu   +9 more
semanticscholar   +1 more source

ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning

Neural Information Processing Systems
We propose ReinFlow, a simple yet effective online reinforcement learning (RL) framework that fine-tunes a family of flow matching policies for continuous robotic control.
Tonghe Zhang   +3 more
semanticscholar   +1 more source

Parameter-Efficient Fine-Tuning for Foundation Models

arXiv.org
This survey delves into the realm of Parameter-Efficient Fine-Tuning (PEFT) within the context of Foundation Models (FMs). PEFT, a cost-effective fine-tuning technique, minimizes parameters and computational complexity while striving for optimal ...
Dan Zhang   +5 more
semanticscholar   +1 more source

Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

International Conference on Machine Learning
Current vision large language models (VLLMs) exhibit remarkable capabilities yet are prone to generate harmful content and are vulnerable to even the simplest jailbreaking attacks.
Yongshuo Zong   +4 more
semanticscholar   +1 more source

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

arXiv.org
Instruction-following Vision Large Language Models (VLLMs) have achieved significant progress recently on a variety of tasks. These approaches merge strong pre-trained vision models and large language models (LLMs).
Yiyang Zhou   +4 more
semanticscholar   +1 more source

Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

Conference on Empirical Methods in Natural Language Processing
When large language models are aligned via supervised fine-tuning, they may encounter new factual information that was not acquired through pre-training.
Zorik Gekhman   +6 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy