Results 251 to 260 of about 776,676 (312)
Some of the next articles are maybe not open access.
The Judge as Reader, the Reader as Judge:
2017This chapter explores the links between reading and judgment in Machaut’s Jugement dou roy de Navarre and shows how the poem interweaves several literary genres into a vernacular “mirror for princes.” Like Dante’s Commedia and Gower’s Confessio amantis, Machaut’s poem echoes John of Salisbury’s association of reading, law, and good kingship in the ...
openaire +1 more source
arXiv.org
LLM-as-a-Judge has been widely applied to evaluate and compare different LLM alignmnet approaches (e.g., RLHF and DPO). However, concerns regarding its reliability have emerged, due to LLM judges' biases and inconsistent decision-making.
Hui Wei +5 more
semanticscholar +1 more source
LLM-as-a-Judge has been widely applied to evaluate and compare different LLM alignmnet approaches (e.g., RLHF and DPO). However, concerns regarding its reliability have emerged, due to LLM judges' biases and inconsistent decision-making.
Hui Wei +5 more
semanticscholar +1 more source
ACM Transactions on Software Engineering and Methodology
Evaluating the alignment of large language models (LLMs) with user-defined coding preferences is a challenging endeavor that requires a deep assessment of LLMs’ outputs.
M. Weyssow +3 more
semanticscholar +1 more source
Evaluating the alignment of large language models (LLMs) with user-defined coding preferences is a challenging endeavor that requires a deep assessment of LLMs’ outputs.
M. Weyssow +3 more
semanticscholar +1 more source
Human-Centered Design Recommendations for LLM-as-a-judge
HUCLLMTraditional reference-based metrics, such as BLEU and ROUGE, are less effective for assessing outputs from Large Language Models (LLMs) that produce highly creative or superior-quality text, or in situations where reference outputs are unavailable. While
Qian Pan +7 more
semanticscholar +1 more source
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
arXiv.orgEfficient and accurate evaluation is crucial for the continuous improvement of large language models (LLMs). Among various assessment methods, subjective evaluation has garnered significant attention due to its superior alignment with real-world usage ...
Maosong Cao +5 more
semanticscholar +1 more source
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
International Conference on Learning RepresentationsHigh quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefore become an important research ambition.
Florian E. Dorner +2 more
semanticscholar +1 more source
Abstract There are already some small-scale automated decision-making processes that have been introduced in the judicial arena. In addition, there are AI systems that can ‘nudge’, ‘prompt’, or ‘correct’ judges when making decisions, as well as generative forms of AI that could support judicial decision-making.
Tania Sourdin, Ella Sourdin Brown
openaire +1 more source
Tania Sourdin, Ella Sourdin Brown
openaire +1 more source
Soviet Law and Government, 1988
It happened on the eve of the election. A reader called the editorial office to ask: "Is it true that your correspondent Borin wrote an article defending his relative?" Generally, such "sensations" are nothing new to newspapermen; no sooner do we return from an assignment than the mud is already flying at our backs, faster than speeding bullets.
openaire +1 more source
It happened on the eve of the election. A reader called the editorial office to ask: "Is it true that your correspondent Borin wrote an article defending his relative?" Generally, such "sensations" are nothing new to newspapermen; no sooner do we return from an assignment than the mud is already flying at our backs, faster than speeding bullets.
openaire +1 more source
Self-rationalization improves LLM as a fine-grained judge
arXiv.orgLLM-as-a-judge models have been used for evaluating both human and AI generated content, specifically by providing scores and rationales. Rationales, in addition to increasing transparency, help models learn to calibrate its judgments.
Prapti Trivedi +9 more
semanticscholar +1 more source

