Results 31 to 40 of about 2,559,382 (283)

Neural Word Segmentation Learning for Chinese [PDF]

open access: yesProceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2016
Most previous approaches to Chinese word segmentation formalize this problem as a character-based sequence labeling task where only contextual information within fixed sized local windows and simple interactions between adjacent tags can be captured. In this paper, we propose a novel neural framework which thoroughly eliminates context windows and can ...
Deng Cai 0002, Hai Zhao 0001
openaire   +2 more sources

Multi-Grained Chinese Word Segmentation [PDF]

open access: yesProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, 2017
Traditionally, word segmentation (WS) adopts the single-grained formalism, where a sentence corresponds to a single word sequence. However, Sproat et al. (1997) show that the inter-native-speaker consistency ratio over Chinese word boundaries is only 76%, indicating single-grained WS (SWS) imposes unnecessary challenges on both manual annotation and ...
Chen Gong 0004   +3 more
openaire   +2 more sources

Bootstrapping word alignment via word packing [PDF]

open access: yes, 2007
We introduce a simple method to pack words for statistical word alignment. Our goal is to simplify the task of automatic word alignment by packing several consecutive words together when we believe they correspond to a single word in the opposite ...
Stroppa, Nicolas, Way, Andy, Ma, Yanjun
core   +2 more sources

A benchmark dataset and case study for Chinese medical question intent classification

open access: yesBMC Medical Informatics and Decision Making, 2020
Background To provide satisfying answers, medical QA system has to understand the intentions of the users’ questions precisely. For medical intent classification, it requires high-quality datasets to train a deep-learning approach in a supervised way ...
Nan Chen   +4 more
doaj   +1 more source

The role of format familiarity and word frequency in Chinese reading

open access: yesJournal of Eye Movement Research, 2023
For Chinese readers, reading from left to right is the norm, while reading from right to left is unfamiliar. This study comprises two experiments investigating how format familiarity and word frequency affect reading by Chinese people.
Mingjing Chen, Jiamei Lu
doaj   +1 more source

Analysing the Methods of Dzongkha Word Segmentation

open access: yesApplied Computer Systems, 2017
In both Chinese and Dzongkha languages, the greatest challenge is to identify the word boundaries because there are no word delimiters as it is in English and other Western languages.
Dhungyel Parshu Ram   +1 more
doaj   +1 more source

The Trade-Off Between Format Familiarity and Word-Segmentation Facilitation in Chinese Reading

open access: yesFrontiers in Psychology, 2021
In alphabetic writing systems (such as English), the spaces between words mark the word boundaries, and the basic unit of reading is distinguished during visual-level processing.
Mingjing Chen   +4 more
doaj   +1 more source

Synthetic Word Parsing Improves Chinese Word Segmentation [PDF]

open access: yesProceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), 2015
We present a novel solution to improve the performance of Chinese word segmentation (CWS) using a synthetic word parser. The parser analyses the internal structure of words, and attempts to convert out-of-vocabulary words (OOVs) into in-vocabulary fine-grained sub-words.
Fei Cheng 0002   +2 more
openaire   +2 more sources

Introduction to CKIP Chinese word segmentation system for the first international Chinese Word Segmentation Bakeoff [PDF]

open access: yesProceedings of the second SIGHAN workshop on Chinese language processing -, 2003
In this paper, we roughly described the procedures of our segmentation system, including the methods for resolving segmentation ambiguities and identifying unknown words. The CKIP group of Academia Sinica participated in testing on open and closed tracks of Beijing University (PK) and Hong Kong Cityu (HK).
Wei-Yun Ma, Keh-Jiann Chen
openaire   +2 more sources

Word-Context Character Embeddings for Chinese Word Segmentation [PDF]

open access: yesProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, 2017
Neural parsers have benefited from automatically labeled data via dependency-context word embeddings. We investigate training character embeddings on a word-based context in a similar way, showing that the simple method improves state-of-the-art neural word segmentation models significantly, beating tri-training baselines for leveraging auto-segmented ...
Hao Zhou 0012   +5 more
openaire   +1 more source

Home - About - Disclaimer - Privacy