Contextual dependencies in unsupervised word segmentation [PDF]
Developing better methods for segmenting continuous text into words is important for improving the processing of Asian languages, and may shed light on how humans learn to segment speech. We propose two new Bayesian word segmentation methods that assume unigram and bigram models of word dependencies respectively.
Sharon Goldwater +2 more
openaire +3 more sources
Contrasting off‐line segmentation decisions with on‐line word segmentation during reading [PDF]
Chuanli Zang +2 more
exaly +2 more sources
The effect of morpheme positional frequency on tibetan novel word acquisition: An eye-tracking study. [PDF]
Word segmentation is crucial for reading in unspaced languages like Tibetan, where readers rely on high-level cues like morpheme positional frequency (the statistical likelihood of a morpheme appearing at the beginning or end of a word).
Dingyi Niu, Jing Tian
doaj +2 more sources
Semantic Segmentation Method of Tibetan Sentences [PDF]
Sentences are characters or words that are combined according to grammatical rules.Semantic segmentation is a decoding problem of sentence combination rules,that is,parsing the meaning of sentences.If the semantic analysis is performed directly after the
ROU Te, SE Chajia, CAI Rangjia
doaj +1 more source
On the Difficulty of Segmenting Words with Attention [PDF]
Word segmentation, the problem of finding word boundaries in speech, is of interest for a range of tasks. Previous papers have suggested that for sequence-to-sequence models trained on tasks such as speech translation or speech recognition, attention can be used to locate and segment the words.
Ramon Sanabria +2 more
openaire +2 more sources
Do Chinese readers follow the national standard rules for word segmentation during reading? [PDF]
We conducted a preliminary study to examine whether Chinese readers' spontaneous word segmentation processing is consistent with the national standard rules of word segmentation based on the Contemporary Chinese language word segmentation specification ...
Ping-Ping Liu +3 more
doaj +1 more source
Segmentation of Written Words in French [PDF]
Syllabification of spoken words has been largely used to define syllabic properties of written words, such as the number of syllables or syllabic boundaries. By contrast, some authors proposed that the functional structure of written words stems from visuo-orthographic features rather than from the transposition of phonological structure into the ...
Chetail, Fabienne, Content, Alain
openaire +2 more sources
An Algorithm Rapidly Segmenting Chinese Sentences into Individual Words [PDF]
This paper proposes an improved Trie tree structure. The tree node records the position information of the characters participating in the word formation, and the child node uses the hash search mechanism.
Xiong Zhibin
doaj +1 more source
Combining segmenter and chunker for Chinese word segmentation [PDF]
Our proposed method is to use a Hidden Markov Model-based word segmenter and a Support Vector Machine-based chunker for Chinese word segmentation. Firstly, input sentences are analyzed by the Hidden Markov Model-based word segmenter. The word segmenter produces n-best word candidates together with some class information and confidence measures ...
Masayuki Asahara +3 more
openaire +1 more source
Hybrid Feature Fusion Learning Towards Chinese Chemical Literature Word Segmentation
The rapid increase in the number of chemical science literature has brought challenges to researchers in search and data analysis. For many chemical scientific literature, extracting information from text and using knowledge is the focus of research ...
Xiang Li +4 more
doaj +1 more source

