Results 41 to 50 of about 1,541,664 (157)
An end‐to‐end knowledge discovery framework is established to automate high‐precision property extraction from small, specialized literature corpora. Utilizing a domain‐specific bidirectional encoder representation from a transformer model and data augmentation, the system accurately extracts and structures electrolyte performance data, ultimately ...
Gaheun Shin +4 more
wiley +1 more source
Fixing Ill‐Formed UTF‐16 Strings With SIMD Instructions
ABSTRACT Background We must fix ill‐formed UTF‐16 strings, where surrogate pairs are mismatched. Programming languages such as JavaScript provide functions that replace mismatched surrogates with the Unicode replacement character (U + FFFD). Objectives We seek to accelerate the correction of ill‐formed UTF‐16 strings.
Robert Clausecker, Daniel Lemire
wiley +1 more source
Background While multiple alignment is the first step of usual classification schemes for biological sequences, alignment-free methods are being increasingly used as alternatives when multiple alignments fail.
Grasseau Gilles +5 more
doaj +1 more source
LLM‐based prior elicitation for Bayesian graphical modeling
ABSTRACT In the Bayesian graphical modeling framework, priors on network structure encode theoretical assumptions and uncertainty about the topology of psychological constructs under study. For instance, the Bernoulli prior specifies the probability of each pairwise interaction, the Beta–Bernoulli prior governs expected network density, and the ...
Nikola Sekulovski +2 more
wiley +1 more source
Word Predictability as a Measure of Second Language Proficiency
Abstract This study introduces predictabilityBERT, a novel metric for assessing second language (L2) proficiency based on the predictability of word choices in learner language production. Using BERT (Devlin et al., 2019), we calculated the conditional probability of each word in a text given its surrounding context.
Langdon Holmes +3 more
wiley +1 more source
AI foundation models in plant biology
Foundation models decode genomes, engineer proteins, phenotype crops, and drive AI agents across plant biology. This figure was created in BioRender (BioRender.com/xcmmkel). Summary Rapid technological progress has enabled plant biologists to accumulate unprecedented volumes of multi‐scale, multi‐modal data, yet this abundance of data has intensified ...
Haopeng Yu
wiley +1 more source
Asymptotic Analysis of the Kth Subword Complexity
The Subword Complexity of a character string refers to the number of distinct substrings of any length that occur as contiguous patterns in the string. The kth Subword Complexity in particular, refers to the number of distinct substrings of length k in a
Ahmadi, Lida
core +3 more sources
Machine Learning for Synthetic Organic Chemistry: Methods, Applications, and Best Practices
Artificial intelligence (AI) and machine learning (ML) are increasingly reshaping experimental chemistry. This review maps the challenges of synthetic organic chemistry to modern digital tools that can help address them. While focusing on the practical application of ML tools in real‐world laboratory settings, we outline prerequisites, emerging ...
Niklas Hölter +3 more
wiley +1 more source
An Evaluation Framework for Post‐Trained Bookkeeping Language Models
ABSTRACT Supervised fine‐tuning (SFT) enables large language models (LLMs) to acquire domain‐specific knowledge, yet the resulting trade‐off between specialization and general capability preservation remains poorly understood. This paper investigates this trade‐off systematically by fine‐tuning three open‐source LLMs on double‐entry bookkeeping posting
Mario Zupan
wiley +1 more source
On scattered subword complexity
Special scattered subwords, in which the gaps are of length from a given set, are defined. The scattered subword complexity, which is the number of such scattered subwords, is computed for rainbow words.
openaire +4 more sources

