Results 21 to 30 of about 1,323,874 (288)

Adaptive Chinese word segmentation [PDF]

open access: yesProceedings of the 42nd Annual Meeting on Association for Computational Linguistics - ACL '04, 2004
This paper presents a Chinese word segmentation system which can adapt to different domains and standards. We first present a statistical framework where domain-specific words are identified in a unified approach to word segmentation based on linear models. We explore several features and describe how to create training data by sampling.
Jianfeng Gao   +6 more
openaire   +2 more sources

Thai word segmentation for visualization of Thai Web sites [PDF]

open access: yes, 2011
Information overload is a problem in the Information Age and Information visualization is an approach to provide an overview of the content of a web site. Tag cloud is one of the ways to represent information as an image of a group of words.
Wigrai Thanadechteemapat   +3 more
core   +1 more source

Domain-Specific Chinese Word Segmentation Based on Bi-Directional Long-Short Term Memory Model

open access: yesIEEE Access, 2019
Most of the current word segmentation methods are rule-based and traditional machine learning methods. Universal word segmentation tools do not work well in the field such as metallurgy. Domain-specific Chinese word segmentation is rarely studied.
Dangguo Shao   +6 more
doaj   +1 more source

Feature Extraction and Analysis of Natural Language Processing for Deep Learning English Language

open access: yesIEEE Access, 2020
NLP (Natural Language Processing) is a technology that enables computers to understand human languages. Deep-level grammatical and semantic analysis usually uses words as the basic unit, and word segmentation is usually the primary task of NLP.
Dongyang Wang, Junli Su, Hongbin Yu
doaj   +1 more source

An efficient, font independent word and character segmentation algorithm for printed Arabic text

open access: yesJournal of King Saud University: Computer and Information Sciences, 2022
Characters segmentation is a necessity and the most critical stage in Arabic OCR system. It has attracted the interest of a wide range of researchers. However, the nature of the Arabic cursive script poses extra challenges that need further investigation.
Aziz Qaroush   +5 more
doaj   +1 more source

Segmenting Chinese Texts into Words for Semantic Network Analysis [PDF]

open access: yesJournal of Contemporary Eastern Asia, 2017
Unlike most languages, written Chinese has no spaces between words. Word segmentation must be performed before semantic network analysis can be conducted.
James A. Danowski
doaj   +1 more source

A Domain Feature Word Vector Description Method for Military Texts [PDF]

open access: yesJisuanji gongcheng, 2016
According to the large number of named entities and deep domain of feature words in military text information,this paper proposes a vector description method for domain feature words.It compresses the vector space through the optimization of word ...
QIN Jie,CAO Lei,PENG Hui,LAI Jun
doaj   +1 more source

A Dataset for Sanskrit Word Segmentation [PDF]

open access: yesProceedings of the Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature, 2017
La dernière décennie a vu une recrudescence des efforts de numérisation pour les manuscrits anciens en sanskrit. En raison de diverses particularités linguistiques inhérentes à la langue, même les tâches préliminaires telles que la segmentation des mots ne sont pas triviales en sanskrit.
Amrith Krishna   +2 more
openaire   +2 more sources

A Graph-based Model for Joint Chinese Word Segmentation and Dependency Parsing

open access: yesTransactions of the Association for Computational Linguistics, 2020
Chinese word segmentation and dependency parsing are two fundamental tasks for Chinese natural language processing. The dependency parsing is defined at the word-level.
Yan, Hang, Qiu, Xipeng, Huang, Xuanjing
doaj   +1 more source

Numerical Simulation of Ambiguity Resolution in Multiple Information Streams Based on Network Machine Translation

open access: yesComplexity, 2020
In natural language, the phenomenon of polysemy is widespread, which makes it very difficult for machines to process natural language. Word sense disambiguation is a key issue in the field of natural language processing.
Lei Wang, Qun Ai
doaj   +1 more source

Home - About - Disclaimer - Privacy