Results 31 to 40 of about 2,559,382 (283)
Neural Word Segmentation Learning for Chinese [PDF]
Most previous approaches to Chinese word segmentation formalize this problem as a character-based sequence labeling task where only contextual information within fixed sized local windows and simple interactions between adjacent tags can be captured. In this paper, we propose a novel neural framework which thoroughly eliminates context windows and can ...
Deng Cai 0002, Hai Zhao 0001
openaire +2 more sources
Multi-Grained Chinese Word Segmentation [PDF]
Traditionally, word segmentation (WS) adopts the single-grained formalism, where a sentence corresponds to a single word sequence. However, Sproat et al. (1997) show that the inter-native-speaker consistency ratio over Chinese word boundaries is only 76%, indicating single-grained WS (SWS) imposes unnecessary challenges on both manual annotation and ...
Chen Gong 0004 +3 more
openaire +2 more sources
Bootstrapping word alignment via word packing [PDF]
We introduce a simple method to pack words for statistical word alignment. Our goal is to simplify the task of automatic word alignment by packing several consecutive words together when we believe they correspond to a single word in the opposite ...
Stroppa, Nicolas, Way, Andy, Ma, Yanjun
core +2 more sources
A benchmark dataset and case study for Chinese medical question intent classification
Background To provide satisfying answers, medical QA system has to understand the intentions of the users’ questions precisely. For medical intent classification, it requires high-quality datasets to train a deep-learning approach in a supervised way ...
Nan Chen +4 more
doaj +1 more source
The role of format familiarity and word frequency in Chinese reading
For Chinese readers, reading from left to right is the norm, while reading from right to left is unfamiliar. This study comprises two experiments investigating how format familiarity and word frequency affect reading by Chinese people.
Mingjing Chen, Jiamei Lu
doaj +1 more source
Analysing the Methods of Dzongkha Word Segmentation
In both Chinese and Dzongkha languages, the greatest challenge is to identify the word boundaries because there are no word delimiters as it is in English and other Western languages.
Dhungyel Parshu Ram +1 more
doaj +1 more source
The Trade-Off Between Format Familiarity and Word-Segmentation Facilitation in Chinese Reading
In alphabetic writing systems (such as English), the spaces between words mark the word boundaries, and the basic unit of reading is distinguished during visual-level processing.
Mingjing Chen +4 more
doaj +1 more source
Synthetic Word Parsing Improves Chinese Word Segmentation [PDF]
We present a novel solution to improve the performance of Chinese word segmentation (CWS) using a synthetic word parser. The parser analyses the internal structure of words, and attempts to convert out-of-vocabulary words (OOVs) into in-vocabulary fine-grained sub-words.
Fei Cheng 0002 +2 more
openaire +2 more sources
Introduction to CKIP Chinese word segmentation system for the first international Chinese Word Segmentation Bakeoff [PDF]
In this paper, we roughly described the procedures of our segmentation system, including the methods for resolving segmentation ambiguities and identifying unknown words. The CKIP group of Academia Sinica participated in testing on open and closed tracks of Beijing University (PK) and Hong Kong Cityu (HK).
Wei-Yun Ma, Keh-Jiann Chen
openaire +2 more sources
Word-Context Character Embeddings for Chinese Word Segmentation [PDF]
Neural parsers have benefited from automatically labeled data via dependency-context word embeddings. We investigate training character embeddings on a word-based context in a similar way, showing that the simple method improves state-of-the-art neural word segmentation models significantly, beating tri-training baselines for leveraging auto-segmented ...
Hao Zhou 0012 +5 more
openaire +1 more source

