Results 231 to 240 of about 776,676 (312)
Some of the next articles are maybe not open access.
Judging the Judges’ Performance in Rhythmic Gymnastics
Medicine & Science in Sports & Exercise, 2015Rhythmic gymnastics (RG) is an aesthetic event balancing between art and sport that also has a performance rating system (Code of Points) given by the International Gymnastics Federation. It is one of the sports in which competition results greatly depend on the judges' evaluation.
Flessas Konstantinos +8 more
openaire +3 more sources
Agent-as-a-Judge: Evaluate Agents with Agents
International Conference on Machine LearningContemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomes -- ignoring the step-by-step nature of agentic systems, or require excessive manual labour.
Mingchen Zhuge +12 more
semanticscholar +1 more source
One Token to Fool LLM-as-a-Judge
arXiv.orgLarge language models (LLMs) are increasingly trusted as automated judges, assisting evaluation and providing reward signals for training other models, particularly in reference-based settings like Reinforcement Learning with Verifiable Rewards (RLVR ...
Yulai Zhao +6 more
semanticscholar +1 more source
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Annual Meeting of the Association for Computational LinguisticsLarge Language Models (LLMs) have significantly advanced the state-of-the-art in various coding tasks. Beyond directly answering user queries, LLMs can also serve as judges, assessing and comparing the quality of responses generated by other models. Such
Hongchao Jiang +4 more
semanticscholar +1 more source
Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment
International Conference on Learning RepresentationsThe performance of large language models (LLMs) is closely linked to their underlying size, leading to ever-growing networks and hence slower inference.
Gregor Bachmann +8 more
semanticscholar +1 more source
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
arXiv.orgReliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to evaluate response quality, is
Michael Krumdick +4 more
semanticscholar +1 more source
How Reliable is Multilingual LLM-as-a-Judge?
Conference on Empirical Methods in Natural Language ProcessingLLM-as-a-Judge has emerged as a popular evaluation strategy, where advanced large language models assess generation results in alignment with human instructions.
Xiyan Fu, Wei Liu
semanticscholar +1 more source
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
Conference on Empirical Methods in Natural Language ProcessingLarge Language Models (LLMs) are powerful zero-shot assessors used in real-world situations such as assessing written exams and benchmarking systems. Despite these critical applications, no existing work has analyzed the vulnerability of judge-LLMs to ...
Vyas Raina +2 more
semanticscholar +1 more source
Judge Anything: MLLM as a Judge Across Any Modality
Knowledge Discovery and Data MiningEvaluating generative foundation models on open-ended multimodal understanding (MMU) and generation (MMG) tasks across diverse modalities (e.g., images, audio, video) poses significant challenges due to the complexity of cross-modal interactions. To this
Shu Pu +12 more
semanticscholar +1 more source
Improving LLM-as-a-Judge Inference with the Judgment Distribution
Conference on Empirical Methods in Natural Language ProcessingUsing language models to scalably approximate human preferences on text quality (LLM-as-a-judge) has become a standard practice applicable to many tasks. A judgment is often extracted from the judge's textual output alone, typically with greedy decoding.
Victor Wang +2 more
semanticscholar +1 more source

