Results 221 to 230 of about 776,676 (312)
Some of the next articles are maybe not open access.
Annual Meeting of the Association for Computational Linguistics
The"LLM-as-an-annotator"and"LLM-as-a-judge"paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans.
Nitay Calderon, Roi Reichart, Rotem Dror
semanticscholar +1 more source
The"LLM-as-an-annotator"and"LLM-as-a-judge"paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans.
Nitay Calderon, Roi Reichart, Rotem Dror
semanticscholar +1 more source
JudgeLRM: Large Reasoning Models as a Judge
arXiv.orgLarge Language Models (LLMs) are increasingly adopted as evaluators, offering a scalable alternative to human annotation. However, existing supervised fine-tuning (SFT) approaches often fall short in domains that demand complex reasoning.
Nuo Chen +6 more
semanticscholar +1 more source
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
International Conference on Machine LearningLLM-as-a-Judge models generate chain-of-thought (CoT) sequences intended to capture the step-bystep reasoning process that underlies the final evaluation of a response.
Swarnadeep Saha +4 more
semanticscholar +1 more source
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
Neural Information Processing SystemsAgentic search such as Deep Research systems-where agents autonomously browse the web, synthesize information, and return comprehensive citation-backed answers-represents a major shift in how users interact with web-scale information.
Boyu Gou +25 more
semanticscholar +1 more source
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
Conference on Empirical Methods in Natural Language ProcessingAssessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP). Traditional methods, usually matching-based or small model-based, often fall short in open-ended and dynamic scenarios ...
Dawei Li +12 more
semanticscholar +1 more source
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
Conference on Empirical Methods in Natural Language ProcessingLarge language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive environments. Instead of centering
Tongxin Yuan +11 more
semanticscholar +1 more source
On the Effectiveness of LLM-as-a-Judge for Code Generation and Summarization
IEEE Transactions on Software EngineeringLarge Language Models (LLMs) have been recently exploited as judges for complex natural language processing tasks, such as Q&A (Question & Answer). The basic idea is to delegate to an LLM the assessment of the “quality” of the output provided by an ...
Giuseppe Crupi +5 more
semanticscholar +1 more source
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
International Conference on Learning RepresentationsLLM-as-a-Judge has been widely utilized as an evaluation method in various benchmarks and served as supervised rewards in model training. However, despite their excellence in many domains, potential issues are under-explored, undermining their ...
Jiayi Ye +11 more
semanticscholar +1 more source
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Conference on Empirical Methods in Natural Language ProcessingLarge Language Models (LLMs) are rapidly surpassing human knowledge in many domains. While improving these models traditionally relies on costly human data, recent self-rewarding mechanisms (Yuan et al., 2024) have shown that LLMs can improve by judging ...
Tianhao Wu +7 more
semanticscholar +1 more source
Self-Preference Bias in LLM-as-a-Judge
arXiv.orgAutomated evaluation leveraging large language models (LLMs), commonly referred to as LLM evaluators or LLM-as-a-judge, has been widely used in measuring the performance of dialogue systems. However, the self-preference bias in LLMs has posed significant
Koki Wataoka +2 more
semanticscholar +1 more source

