Results 251 to 260 of about 2,403,325 (304)
Some of the next articles are maybe not open access.
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
International Conference on Machine LearningHarnessing the power of human-annotated data through Supervised Fine-Tuning (SFT) is pivotal for advancing Large Language Models (LLMs). In this paper, we delve into the prospect of growing a strong LLM out of a weak one without the need for acquiring ...
Zixiang Chen +4 more
semanticscholar +1 more source
ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
RoboticsVision-Language-Action (VLA) models have shown substantial potential in real-world robotic manipulation. However, fine-tuning these models through supervised learning struggles to achieve robust performance due to limited, inconsistent demonstrations ...
Yuhui Chen +5 more
semanticscholar +1 more source
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
International Conference on Machine LearningTraining models to effectively use test-time compute is crucial for improving the reasoning performance of LLMs. Current methods mostly do so via fine-tuning on search traces or running RL with 0/1 outcome reward, but do these approaches efficiently ...
Yuxiao Qu +6 more
semanticscholar +1 more source
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
arXiv.orgReinforcement Learning (RL) benefits Large Language Models (LLMs) for complex reasoning. Inspired by this, we explore integrating spatio-temporal specific rewards into Multimodal Large Language Models (MLLMs) to address the unique challenges of video ...
Xinhao Li +9 more
semanticscholar +1 more source
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
arXiv.orgLarge language models (LLMs) have achieved remarkable progress in reasoning tasks, yet the optimal integration of Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) remains a fundamental challenge.
Yuqian Fu +9 more
semanticscholar +1 more source
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
Neural Information Processing SystemsWe propose ReinFlow, a simple yet effective online reinforcement learning (RL) framework that fine-tunes a family of flow matching policies for continuous robotic control.
Tonghe Zhang +3 more
semanticscholar +1 more source
Parameter-Efficient Fine-Tuning for Foundation Models
arXiv.orgThis survey delves into the realm of Parameter-Efficient Fine-Tuning (PEFT) within the context of Foundation Models (FMs). PEFT, a cost-effective fine-tuning technique, minimizes parameters and computational complexity while striving for optimal ...
Dan Zhang +5 more
semanticscholar +1 more source
Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models
International Conference on Machine LearningCurrent vision large language models (VLLMs) exhibit remarkable capabilities yet are prone to generate harmful content and are vulnerable to even the simplest jailbreaking attacks.
Yongshuo Zong +4 more
semanticscholar +1 more source
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
arXiv.orgInstruction-following Vision Large Language Models (VLLMs) have achieved significant progress recently on a variety of tasks. These approaches merge strong pre-trained vision models and large language models (LLMs).
Yiyang Zhou +4 more
semanticscholar +1 more source
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
Conference on Empirical Methods in Natural Language ProcessingWhen large language models are aligned via supervised fine-tuning, they may encounter new factual information that was not acquired through pre-training.
Zorik Gekhman +6 more
semanticscholar +1 more source

