Results 291 to 300 of about 58,163,280 (350)
Some of the next articles are maybe not open access.
Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation
Neural Information Processing Systems, 2022We propose Active Surrogate Estimators (ASEs), a new method for label-efficient model evaluation. Evaluating model performance is a challenging and important problem when labels are expensive.
Jannik Kossen +3 more
semanticscholar +1 more source
Operational Evaluation Modeling
SIMULATION, 1990A graphical parallel process language is described that assists a system designer in visualization of system operation. Visual ization is essential when building an operational model to experiment with alternative system operation and architec tures and to evaluate design tradeoffs.
John R. Clymer +2 more
openaire +1 more source
COOPERATIVE MODELLING EVALUATED
International Journal of Cooperative Information Systems, 2005In any modelling activity, a framework to determine the maturity of a developed model before its use is highly advantageous. Such a framework would save modellers expensive time in many areas of information systems. It would also lower the risk of users relying on an incomplete or inaccurate model.
Ghassan Beydoun +5 more
openaire +1 more source
2020
We expect that the traffic will be almost optimal when the collective behaviour of autonomous vehicles will determine the traffic. The route selection plays an important role in optimizing the traffic. There are different models of the routing problem.
Vince Antal +4 more
openaire +1 more source
We expect that the traffic will be almost optimal when the collective behaviour of autonomous vehicles will determine the traffic. The route selection plays an important role in optimizing the traffic. There are different models of the routing problem.
Vince Antal +4 more
openaire +1 more source
Supplementary material to "A flux tower dataset tailored for land model evaluation"
Earth System Science Data, 2021. Eddy covariance flux towers measure the exchange of water, energy and carbon fluxes between the land and atmosphere. They have become invaluable for theory development and evaluating land models.
A. Ukkola, G. Abramowitz, M. D. De Kauwe
semanticscholar +1 more source
RewardBench 2: Advancing Reward Model Evaluation
arXiv.orgReward models are used throughout the post-training of language models to capture nuanced signals from preference data and provide a training target for optimization across instruction following, reasoning, safety, and more domains.
Saumya Malik +6 more
semanticscholar +1 more source
GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation
arXiv.orgTime series foundation models excel in zero-shot forecasting, handling diverse tasks without explicit training. However, the advancement of these models has been hindered by the lack of comprehensive benchmarks.
Taha İbrahim Aksu +7 more
semanticscholar +1 more source
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
Conference on Empirical Methods in Natural Language ProcessingExisting large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-linguistic reasoning abilities.
Weihao Xuan +17 more
semanticscholar +1 more source
Evaluation of Reference Models [PDF]
Evaluating a reference model is a demanding task. Not only that reference models inherit the problems well known from the evaluation of conceptual models in general. Furthermore, their claim for general (re-) usability implies to take into account the possible variety of requirements and specific constraints within the set of potential applications ...
openaire +1 more source
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
arXiv.orgMultiple choice benchmarks have long been the workhorse of language model evaluation because grading multiple choice is objective and easy to automate. However, we show multiple choice questions from popular benchmarks can often be answered without even ...
Nikhil Chandak +4 more
semanticscholar +1 more source

