Chinese word segmentation as LMR tagging [PDF]
In this paper we present Chinese word segmentation algorithms based on the so-called LMR tagging. Our LMR taggers are implemented with the Maximum Entropy Markov Model and we then use Transformation-Based Learning to combine the results of the two LMR taggers that scan the input in opposite directions. Our system achieves F-scores of 95.9% and 91.6% on
Nianwen Xue, Libin Shen
openaire +1 more source
Coupling Distant Annotation and Adversarial Training for Cross-Domain Chinese Word Segmentation [PDF]
Fully supervised neural approaches have achieved significant progress in the task of Chinese word segmentation (CWS). Nevertheless, the performance of supervised models always drops gravely if the domain shifts due to the distribution gap across domains ...
Ning Ding +6 more
semanticscholar +1 more source
Multi-Grained Chinese Word Segmentation [PDF]
Traditionally, word segmentation (WS) adopts the single-grained formalism, where a sentence corresponds to a single word sequence. However, Sproat et al. (1997) show that the inter-native-speaker consistency ratio over Chinese word boundaries is only 76%, indicating single-grained WS (SWS) imposes unnecessary challenges on both manual annotation and ...
Chen Gong 0004 +3 more
openaire +1 more source
A benchmark dataset and case study for Chinese medical question intent classification
Background To provide satisfying answers, medical QA system has to understand the intentions of the users’ questions precisely. For medical intent classification, it requires high-quality datasets to train a deep-learning approach in a supervised way ...
Nan Chen +4 more
doaj +1 more source
The role of format familiarity and word frequency in Chinese reading
For Chinese readers, reading from left to right is the norm, while reading from right to left is unfamiliar. This study comprises two experiments investigating how format familiarity and word frequency affect reading by Chinese people.
Mingjing Chen, Jiamei Lu
doaj +1 more source
Synthetic Word Parsing Improves Chinese Word Segmentation [PDF]
We present a novel solution to improve the performance of Chinese word segmentation (CWS) using a synthetic word parser. The parser analyses the internal structure of words, and attempts to convert out-of-vocabulary words (OOVs) into in-vocabulary fine-grained sub-words.
Fei Cheng 0002 +2 more
openaire +1 more source
The Trade-Off Between Format Familiarity and Word-Segmentation Facilitation in Chinese Reading
In alphabetic writing systems (such as English), the spaces between words mark the word boundaries, and the basic unit of reading is distinguished during visual-level processing.
Mingjing Chen +4 more
doaj +1 more source
Analysing the Methods of Dzongkha Word Segmentation
In both Chinese and Dzongkha languages, the greatest challenge is to identify the word boundaries because there are no word delimiters as it is in English and other Western languages.
Dhungyel Parshu Ram +1 more
doaj +1 more source
Word-Context Character Embeddings for Chinese Word Segmentation [PDF]
Neural parsers have benefited from automatically labeled data via dependency-context word embeddings. We investigate training character embeddings on a word-based context in a similar way, showing that the simple method improves state-of-the-art neural word segmentation models significantly, beating tri-training baselines for leveraging auto-segmented ...
Hao Zhou 0012 +5 more
openaire +1 more source
Introduction to CKIP Chinese word segmentation system for the first international Chinese Word Segmentation Bakeoff [PDF]
In this paper, we roughly described the procedures of our segmentation system, including the methods for resolving segmentation ambiguities and identifying unknown words. The CKIP group of Academia Sinica participated in testing on open and closed tracks of Beijing University (PK) and Hong Kong Cityu (HK).
Wei-Yun Ma, Keh-Jiann Chen
openaire +1 more source

