Results 241 to 250 of about 776,676 (312)
Some of the next articles are maybe not open access.

Improve LLM-as-a-Judge Ability as a General Ability

Conference on Empirical Methods in Natural Language Processing
LLM-as-a-Judge leverages the generative and reasoning capabilities of large language models (LLMs) to evaluate LLM responses across diverse scenarios, providing accurate preference signals.
Jiachen Yu   +5 more
semanticscholar   +1 more source

Judging the Judge

2012
“ The nation will judge both the offender and judges for themselves.” Jefferson to William B. Giles, April 20, 1807 “…His Honor did not for two days understand either the questions or himself…” Burr on Marshall, September 20, 1807 “Our Treason Laws may be defective, but I believe Marshall’s Conduct strictly and correctly legal as the Laws now ...
openaire   +1 more source

Judges, Judging and Humour

Comedy Studies, 2020
The courtroom is, perhaps surprisingly frequently, the site of interludes, interruptions – prone to frustration or humour when the not-so-consistently-well-oiled machinery of justice lets loose a p...
openaire   +1 more source

Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement

arXiv.org
This research introduces the Judge's Verdict Benchmark, a novel two-step methodology to evaluate Large Language Models (LLMs) as judges for response accuracy evaluation tasks. We assess how well 54 LLMs can replicate human judgment when scoring responses
Steve Han   +3 more
semanticscholar   +1 more source

The freedom of Judging

Iris : european journal of Philosophy and Public Debate : 3, 6, 2011, 2011
John McDowell and Christine Korsgaard have defended the claim that when human beings judge or believe that p, they are exercising a fundamental kind of freedom, the “freedom of judging.” David Owens has challenged the view: he argues that they offer us at best no more than a modest notion of freedom, which does not vindicate the claim that we are free ...
openaire   +3 more sources

MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation

arXiv.org
The LLM-as-a-Judge paradigm shows promise for evaluating generative content but lacks reliability in reasoning-intensive scenarios, such as programming. Inspired by recent advances in reasoning models and shifts in scaling laws, we pioneer bringing test ...
Yutong Wang   +6 more
semanticscholar   +1 more source

How to Judge a Book by its Cover


Let's read! We will often find out this sentence everywhere. When still being a kid, mom used to order us to always read, so did the teacher. Some books are fully read in a week and we need the obligation to support reading. What about now?
Kate Cuthbert
semanticscholar   +1 more source

Can LLM be a Personalized Judge?

Conference on Empirical Methods in Natural Language Processing
Ensuring that large language models (LLMs) reflect diverse user values and preferences is crucial as their user bases expand globally. It is therefore encouraging to see the growing interest in LLM personalization within the research community.
Yijiang River Dong   +2 more
semanticscholar   +1 more source

LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

arXiv.org
Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse tasks, driving the development and widespread adoption of LLM-as-a-Judge systems for automated evaluation, including red teaming and benchmarking.
Songze Li   +8 more
semanticscholar   +1 more source

Validating LLM-as-a-Judge Systems under Rating Indeterminacy

Neural Information Processing Systems
The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and standardizing GenAI evaluations.
Luke Guerdan   +5 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy