Results 31 to 40 of about 28,490,308 (324)
Chinese word segmentation (CWS) and part-of-speech (POS) tagging are two fundamental tasks for Chinese language processing. Previous studies have demonstrated that jointly performing them can be an effective one-step solution to both tasks and this joint
Yuanhe Tian, Yan Song, Fei Xia
semanticscholar +1 more source
Variable word rate N-grams [PDF]
4 pages, 4 figures, ICASSP ...
Yoshihiko Gotoh, Steve Renals
openaire +4 more sources
t-SS3: a text classifier with dynamic n-grams for early risk detection over text streams [PDF]
A recently introduced classifier, called SS3, has shown to be well suited to deal with early risk detection (ERD) problems on text streams. It obtained state-of-the-art performance on early depression and anorexia detection on Reddit in the CLEF’s eRisk ...
Sergio Gastón Burdisso +2 more
semanticscholar +1 more source
Learning Chinese Word Embeddings With Words and Subcharacter N-Grams
Co-occurrence information between words is the basis of training word embeddings; besides, Chinese characters are composed of subcharacters, words made up by the same characters or subcharacters usually have similar semantics, but this internal ...
Ruizhi Kang +4 more
doaj +1 more source
Automatically identifying code features for software defect prediction: Using AST N-grams
Context: Identifying defects in code early is important. A wide range of static code metrics have been evaluated as potential defect indicators. Most of these metrics offer only high level insights and focus on particular pre-selected features of the ...
Thomas Shippey, David Bowes, T. Hall
semanticscholar +1 more source
Stemming and n-grams in Spanish: An evaluation of their impact on information retrieval [PDF]
At some stage, most of the models and techniques implemented in IR use frequency counts of the terms appearing in documents and in queries. However, many words, since they are derived from the same stem, have very close semantic contents.
López de San Roman, Eva +2 more
core +2 more sources
DETEKSI PLAGIASI DOKUMEN SKRIPSI MAHASISWA MENGGUNAKAN METODE N-GRAMS DAN WINNOWING
Salah satu tantangan dalam bidang akademik adalah mencegah maraknya aktivitas plagiarisme. Salah satu cara yang bisa dilakukan adalah dengan melakukan deteksi dini plagiasi terhadap karya mahasiswa terutama skripsi. Penerapan deteksi indikasi plagiarisme
Fitri Ratna Ning Wulan +2 more
doaj +1 more source
It has been argued that most of corpus linguistics involves one of four fundamental methods: frequency lists, dispersion, collocation, and concordancing. All these presuppose (if only implicitly) the definition of a unit: the element whose frequency in a
Stefan Th. Gries
doaj +1 more source
Linguistic compositions highly volatile in Portuguese
In this paper we use a distance d between sequences of N-grams to identify N-grams that show a different performance when comparing two sequences of N-grams.
Jesús Enrique García +2 more
doaj +1 more source
Visualizing the development of prose styles in Horse Manuals from Early Modern English to Present-Day English [PDF]
This paper offers a data-driven analysis of the development of English prose styles in a single genre (instructive writing) dealing with a single topic (the correct way of feeding a horse) in 13 texts with publication dates ranging between 1565 to 2009 ...
Thijs Lubbers, Bettelou Los
doaj +1 more source

