Results 221 to 230 of about 776,676 (312)
Some of the next articles are maybe not open access.

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Annual Meeting of the Association for Computational Linguistics
The"LLM-as-an-annotator"and"LLM-as-a-judge"paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans.
Nitay Calderon, Roi Reichart, Rotem Dror
semanticscholar   +1 more source

JudgeLRM: Large Reasoning Models as a Judge

arXiv.org
Large Language Models (LLMs) are increasingly adopted as evaluators, offering a scalable alternative to human annotation. However, existing supervised fine-tuning (SFT) approaches often fall short in domains that demand complex reasoning.
Nuo Chen   +6 more
semanticscholar   +1 more source

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

International Conference on Machine Learning
LLM-as-a-Judge models generate chain-of-thought (CoT) sequences intended to capture the step-bystep reasoning process that underlies the final evaluation of a response.
Swarnadeep Saha   +4 more
semanticscholar   +1 more source

Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge

Neural Information Processing Systems
Agentic search such as Deep Research systems-where agents autonomously browse the web, synthesize information, and return comprehensive citation-backed answers-represents a major shift in how users interact with web-scale information.
Boyu Gou   +25 more
semanticscholar   +1 more source

From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge

Conference on Empirical Methods in Natural Language Processing
Assessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP). Traditional methods, usually matching-based or small model-based, often fall short in open-ended and dynamic scenarios ...
Dawei Li   +12 more
semanticscholar   +1 more source

R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Conference on Empirical Methods in Natural Language Processing
Large language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive environments. Instead of centering
Tongxin Yuan   +11 more
semanticscholar   +1 more source

On the Effectiveness of LLM-as-a-Judge for Code Generation and Summarization

IEEE Transactions on Software Engineering
Large Language Models (LLMs) have been recently exploited as judges for complex natural language processing tasks, such as Q&A (Question & Answer). The basic idea is to delegate to an LLM the assessment of the “quality” of the output provided by an ...
Giuseppe Crupi   +5 more
semanticscholar   +1 more source

Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

International Conference on Learning Representations
LLM-as-a-Judge has been widely utilized as an evaluation method in various benchmarks and served as supervised rewards in model training. However, despite their excellence in many domains, potential issues are under-explored, undermining their ...
Jiayi Ye   +11 more
semanticscholar   +1 more source

Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Conference on Empirical Methods in Natural Language Processing
Large Language Models (LLMs) are rapidly surpassing human knowledge in many domains. While improving these models traditionally relies on costly human data, recent self-rewarding mechanisms (Yuan et al., 2024) have shown that LLMs can improve by judging ...
Tianhao Wu   +7 more
semanticscholar   +1 more source

Self-Preference Bias in LLM-as-a-Judge

arXiv.org
Automated evaluation leveraging large language models (LLMs), commonly referred to as LLM evaluators or LLM-as-a-judge, has been widely used in measuring the performance of dialogue systems. However, the self-preference bias in LLMs has posed significant
Koki Wataoka   +2 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy