Results 231 to 240 of about 776,676 (312)
Some of the next articles are maybe not open access.

Judging the Judges’ Performance in Rhythmic Gymnastics

Medicine & Science in Sports & Exercise, 2015
Rhythmic gymnastics (RG) is an aesthetic event balancing between art and sport that also has a performance rating system (Code of Points) given by the International Gymnastics Federation. It is one of the sports in which competition results greatly depend on the judges' evaluation.
Flessas Konstantinos   +8 more
openaire   +3 more sources

Agent-as-a-Judge: Evaluate Agents with Agents

International Conference on Machine Learning
Contemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomes -- ignoring the step-by-step nature of agentic systems, or require excessive manual labour.
Mingchen Zhuge   +12 more
semanticscholar   +1 more source

One Token to Fool LLM-as-a-Judge

arXiv.org
Large language models (LLMs) are increasingly trusted as automated judges, assisting evaluation and providing reward signals for training other models, particularly in reference-based settings like Reinforcement Learning with Verifiable Rewards (RLVR ...
Yulai Zhao   +6 more
semanticscholar   +1 more source

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Annual Meeting of the Association for Computational Linguistics
Large Language Models (LLMs) have significantly advanced the state-of-the-art in various coding tasks. Beyond directly answering user queries, LLMs can also serve as judges, assessing and comparing the quality of responses generated by other models. Such
Hongchao Jiang   +4 more
semanticscholar   +1 more source

Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment

International Conference on Learning Representations
The performance of large language models (LLMs) is closely linked to their underlying size, leading to ever-growing networks and hence slower inference.
Gregor Bachmann   +8 more
semanticscholar   +1 more source

No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

arXiv.org
Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to evaluate response quality, is
Michael Krumdick   +4 more
semanticscholar   +1 more source

How Reliable is Multilingual LLM-as-a-Judge?

Conference on Empirical Methods in Natural Language Processing
LLM-as-a-Judge has emerged as a popular evaluation strategy, where advanced large language models assess generation results in alignment with human instructions.
Xiyan Fu, Wei Liu
semanticscholar   +1 more source

Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment

Conference on Empirical Methods in Natural Language Processing
Large Language Models (LLMs) are powerful zero-shot assessors used in real-world situations such as assessing written exams and benchmarking systems. Despite these critical applications, no existing work has analyzed the vulnerability of judge-LLMs to ...
Vyas Raina   +2 more
semanticscholar   +1 more source

Judge Anything: MLLM as a Judge Across Any Modality

Knowledge Discovery and Data Mining
Evaluating generative foundation models on open-ended multimodal understanding (MMU) and generation (MMG) tasks across diverse modalities (e.g., images, audio, video) poses significant challenges due to the complexity of cross-modal interactions. To this
Shu Pu   +12 more
semanticscholar   +1 more source

Improving LLM-as-a-Judge Inference with the Judgment Distribution

Conference on Empirical Methods in Natural Language Processing
Using language models to scalably approximate human preferences on text quality (LLM-as-a-judge) has become a standard practice applicable to many tasks. A judgment is often extracted from the judge's textual output alone, typically with greedy decoding.
Victor Wang   +2 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy