Results 241 to 250 of about 776,676 (312)
Some of the next articles are maybe not open access.
Improve LLM-as-a-Judge Ability as a General Ability
Conference on Empirical Methods in Natural Language ProcessingLLM-as-a-Judge leverages the generative and reasoning capabilities of large language models (LLMs) to evaluate LLM responses across diverse scenarios, providing accurate preference signals.
Jiachen Yu +5 more
semanticscholar +1 more source
2012
“ The nation will judge both the offender and judges for themselves.” Jefferson to William B. Giles, April 20, 1807 “…His Honor did not for two days understand either the questions or himself…” Burr on Marshall, September 20, 1807 “Our Treason Laws may be defective, but I believe Marshall’s Conduct strictly and correctly legal as the Laws now ...
openaire +1 more source
“ The nation will judge both the offender and judges for themselves.” Jefferson to William B. Giles, April 20, 1807 “…His Honor did not for two days understand either the questions or himself…” Burr on Marshall, September 20, 1807 “Our Treason Laws may be defective, but I believe Marshall’s Conduct strictly and correctly legal as the Laws now ...
openaire +1 more source
Comedy Studies, 2020
The courtroom is, perhaps surprisingly frequently, the site of interludes, interruptions – prone to frustration or humour when the not-so-consistently-well-oiled machinery of justice lets loose a p...
openaire +1 more source
The courtroom is, perhaps surprisingly frequently, the site of interludes, interruptions – prone to frustration or humour when the not-so-consistently-well-oiled machinery of justice lets loose a p...
openaire +1 more source
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
arXiv.orgThis research introduces the Judge's Verdict Benchmark, a novel two-step methodology to evaluate Large Language Models (LLMs) as judges for response accuracy evaluation tasks. We assess how well 54 LLMs can replicate human judgment when scoring responses
Steve Han +3 more
semanticscholar +1 more source
Iris : european journal of Philosophy and Public Debate : 3, 6, 2011, 2011
John McDowell and Christine Korsgaard have defended the claim that when human beings judge or believe that p, they are exercising a fundamental kind of freedom, the “freedom of judging.” David Owens has challenged the view: he argues that they offer us at best no more than a modest notion of freedom, which does not vindicate the claim that we are free ...
openaire +3 more sources
John McDowell and Christine Korsgaard have defended the claim that when human beings judge or believe that p, they are exercising a fundamental kind of freedom, the “freedom of judging.” David Owens has challenged the view: he argues that they offer us at best no more than a modest notion of freedom, which does not vindicate the claim that we are free ...
openaire +3 more sources
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
arXiv.orgThe LLM-as-a-Judge paradigm shows promise for evaluating generative content but lacks reliability in reasoning-intensive scenarios, such as programming. Inspired by recent advances in reasoning models and shifts in scaling laws, we pioneer bringing test ...
Yutong Wang +6 more
semanticscholar +1 more source
How to Judge a Book by its Cover
Let's read! We will often find out this sentence everywhere. When still being a kid, mom used to order us to always read, so did the teacher. Some books are fully read in a week and we need the obligation to support reading. What about now?
Kate Cuthbert
semanticscholar +1 more source
Can LLM be a Personalized Judge?
Conference on Empirical Methods in Natural Language ProcessingEnsuring that large language models (LLMs) reflect diverse user values and preferences is crucial as their user bases expand globally. It is therefore encouraging to see the growing interest in LLM personalization within the research community.
Yijiang River Dong +2 more
semanticscholar +1 more source
LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
arXiv.orgLarge Language Models (LLMs) have demonstrated exceptional capabilities across diverse tasks, driving the development and widespread adoption of LLM-as-a-Judge systems for automated evaluation, including red teaming and benchmarking.
Songze Li +8 more
semanticscholar +1 more source
Validating LLM-as-a-Judge Systems under Rating Indeterminacy
Neural Information Processing SystemsThe LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and standardizing GenAI evaluations.
Luke Guerdan +5 more
semanticscholar +1 more source

