Results 41 to 50 of about 1,541,664 (157)

Accelerating the Discovery of Proton Conducting Electrolytes via Machine Learning‐Enabled Literature Mining

open access: yesAdvanced Intelligent Discovery, EarlyView.
An end‐to‐end knowledge discovery framework is established to automate high‐precision property extraction from small, specialized literature corpora. Utilizing a domain‐specific bidirectional encoder representation from a transformer model and data augmentation, the system accurately extracts and structures electrolyte performance data, ultimately ...
Gaheun Shin   +4 more
wiley   +1 more source

Fixing Ill‐Formed UTF‐16 Strings With SIMD Instructions

open access: yesSoftware: Practice and Experience, EarlyView.
ABSTRACT Background We must fix ill‐formed UTF‐16 strings, where surrogate pairs are mismatched. Programming languages such as JavaScript provide functions that replace mismatched surrogates with the Unicode replacement character (U + FFFD). Objectives We seek to accelerate the correction of ill‐formed UTF‐16 strings.
Robert Clausecker, Daniel Lemire
wiley   +1 more source

MS4 - Multi-Scale Selector of Sequence Signatures: An alignment-free method for classification of biological sequences

open access: yesBMC Bioinformatics, 2010
Background While multiple alignment is the first step of usual classification schemes for biological sequences, alignment-free methods are being increasingly used as alternatives when multiple alignments fail.
Grasseau Gilles   +5 more
doaj   +1 more source

LLM‐based prior elicitation for Bayesian graphical modeling

open access: yesBritish Journal of Mathematical and Statistical Psychology, EarlyView.
ABSTRACT In the Bayesian graphical modeling framework, priors on network structure encode theoretical assumptions and uncertainty about the topology of psychological constructs under study. For instance, the Bernoulli prior specifies the probability of each pairwise interaction, the Beta–Bernoulli prior governs expected network density, and the ...
Nikola Sekulovski   +2 more
wiley   +1 more source

Word Predictability as a Measure of Second Language Proficiency

open access: yesLanguage Learning, EarlyView.
Abstract This study introduces predictabilityBERT, a novel metric for assessing second language (L2) proficiency based on the predictability of word choices in learner language production. Using BERT (Devlin et al., 2019), we calculated the conditional probability of each word in a text given its surrounding context.
Langdon Holmes   +3 more
wiley   +1 more source

AI foundation models in plant biology

open access: yesNew Phytologist, EarlyView.
Foundation models decode genomes, engineer proteins, phenotype crops, and drive AI agents across plant biology. This figure was created in BioRender (BioRender.com/xcmmkel). Summary Rapid technological progress has enabled plant biologists to accumulate unprecedented volumes of multi‐scale, multi‐modal data, yet this abundance of data has intensified ...
Haopeng Yu
wiley   +1 more source

Asymptotic Analysis of the Kth Subword Complexity

open access: yes, 2019
The Subword Complexity of a character string refers to the number of distinct substrings of any length that occur as contiguous patterns in the string. The kth Subword Complexity in particular, refers to the number of distinct substrings of length k in a
Ahmadi, Lida
core   +3 more sources

Machine Learning for Synthetic Organic Chemistry: Methods, Applications, and Best Practices

open access: yesAngewandte Chemie Novit, Volume 2, Issue 3, September 2026.
Artificial intelligence (AI) and machine learning (ML) are increasingly reshaping experimental chemistry. This review maps the challenges of synthetic organic chemistry to modern digital tools that can help address them. While focusing on the practical application of ML tools in real‐world laboratory settings, we outline prerequisites, emerging ...
Niklas Hölter   +3 more
wiley   +1 more source

An Evaluation Framework for Post‐Trained Bookkeeping Language Models

open access: yesIntelligent Systems in Accounting, Finance and Management, Volume 33, Issue 3, September 2026.
ABSTRACT Supervised fine‐tuning (SFT) enables large language models (LLMs) to acquire domain‐specific knowledge, yet the resulting trade‐off between specialization and general capability preservation remains poorly understood. This paper investigates this trade‐off systematically by fine‐tuning three open‐source LLMs on double‐entry bookkeeping posting
Mario Zupan
wiley   +1 more source

On scattered subword complexity

open access: yesCoRR, 2011
Special scattered subwords, in which the gaps are of length from a given set, are defined. The scattered subword complexity, which is the number of such scattered subwords, is computed for rainbow words.
openaire   +4 more sources

Home - About - Disclaimer - Privacy