22 papers · ranked by Valyu relevance
Yerai Doval, Carlos Gómez‐Rodríguez
Word segmentation is the task of inserting or deleting word boundary characters in order to separate character sequences that correspond to words in some language. In this article we propose an approach based on a beam search algorithm and a language model working at the byte/character level, the latter component…
Yerai Doval, Carlos Gómez‐Rodríguez
Word segmentation is the task of inserting or deleting word boundary characters in order to separate character sequences that correspond to words in some language. In this article we propose an approach based on a beam search algorithm and a language model working at the byte/character level, the latter component…
Tero Hakala, Tiina Lindh-Knuutila, Annika Hultén, Minna Lehtonen + 1 more
'Riitta Salmelin'] Title: Abstract This study extends the idea of decoding word-evoked brain activations using a corpus-semantic vector space to multimorphemic words in the agglutinative Finnish language. The corpus-semantic models are trained on word segments, and decoding is carried out with word vectors that are…
Yuanhao Liu, Sheng Yu
We propose a new approach to the Chinese word segmentation problem that considers the sentence as an undirected graph, whose nodes are the characters. One can use various techniques to compute the edge weights that measure the connection strength between characters. Spectral graph partition algorithms are used to group…
Song Nguyen Duc Cong, Hung Q. Ngo, Rachsuda Jiamthapthaksin
—Word segmentation is the first step of any tasks in Vietnamese language processing. This paper reviews stateof-the-art approaches and systems for word segmentation in Vietnamese. To have an overview of all stages from building corpora to developing toolkits, we discuss building the corpus stage, approaches applied to…
Liang Wang, Kaiyong Zhao
Unsupervised word segmentation methods were applied to analyze protein sequences. Protein sequences, such as "MTMDKSELVQKA…," were used as input to these methods. Segmented "protein word" sequences, such as "MTM DKSE LVQKA," were then obtained. We compared the "protein words" derived via unsupervised segmentation and…
Thasayu Soisoonthorn, Herwig Unger, Maleerat Maliyaem
Word segmentation is necessary for many natural language processing, especially Thai language, that is, unsegmented words. However, wrong segmentation causes terrible performance in the final result. In this study, we propose two new brain-inspired methods based on Hawkins' approach to address Thai word segmentation.…
Karl J. Friston, Noor Sajid, David Ricardo Quiroga-Martinez, Thomas Parr + 2 more
This paper introduces active listening, as a unified framework for synthesising and recognising speech. The notion of active listening inherits from active inference, which considers perception and action under one universal imperative: to maximise the evidence for our (generative) models of the world. First, we…
Yuxiao Ye, Yue Zhang, Weikang Li, Likun Qiu + 1 more
Cross-domain Chinese Word Segmentation (CWS) remains a challenge despite recent progress in neural-based CWS. The limited amount of annotated data in the target domain has been the key obstacle to a satisfactory performance. In this paper, we propose a semi-supervised word-based approach to improving cross-domain CWS…
Lucas Benjamin, Ana Fló, Marie Palu, Shruit Naik + 2 more
Since speech is a continuous stream with no systematic boundaries between words, how do pre-verbal infants manage to discover words? A proposed solution is that they might use the transitional probability between adjacent syllables, which drops at word boundaries. Here, we tested the limits of this mechanism by…
Hao Zhang, Jae Hun Ro, Richard Sproat
Breaking domain names such as openresearch into component words open and research is important for applications like Text-to-Speech synthesis and web search. We link this problem to the classic problem of Chinese word segmentation and show the effectiveness of a tagging model based on Recurrent Neural Networks (RNNs)…
Jake Ryland Williams
This work presents a fine-grained, textchunking algorithm designed for the task of multiword expressions (MWEs) segmentation. As a lexical class, MWEs include a wide variety of idioms, whose automatic identification are a necessity for the handling of colloquial language. This algorithm's core novelty is its use of…
Wang Liang
Since the completion of the human genome sequencing project in 2001, significant progress has been made in areas such as gene regulation editing and protein structure prediction. However, given the vast amount of genomic data, the segments that can be fully annotated and understood remain relatively limited. If we…
Zobia Rehman, Waqas Anwar, Usama Ijaz Bajwa, Wang Xuan + 2 more
'Zhou Chaoying' 'Randen Lee Patterson'] Text tokenization is a fundamental pre-processing step for almost all the information processing applications. This task is nontrivial for the scarce resourced languages such as Urdu, as there is inconsistent use of space between words. In this paper a morpheme matching based…
Neil Barrett, Jens Weber-Jahnke
Background Tokenization is an important component of language processing yet there is no widely accepted tokenization method for English texts, including biomedical texts. Other than rule based techniques, tokenization in the biomedical domain has been regarded as a classification task. Biomedical classifier-based…
Saquib Khushhal, Abdul Majid, Syed Ali Abass, Rabia Riaz + 3 more
Word embeddings are essential to natural language processing tasks because they contain a single word’s syntactic and semantic information. Word embeddings have been developed widely for numerous spoken languages across the globe like English. The research community needs to pay more attention to the Urdu language…
Peter Ford Dominey
During continuous perception of movies or stories, awake humans display cortical activity patterns that reveal hierarchical segmentation of event structure. Sensory areas like auditory cortex display high frequency segmentation related to the stimulus, while semantic areas like posterior middle cortex display a lower…
Ehsaneddin Asgari, Alice McHardy, Mohammad R.K. Mofrad
In this paper, we present peptide-pair encoding (PPE), a general-purpose probabilistic segmentation of protein sequences into commonly occurring variable-length sub-sequences. The idea of PPE segmentation is inspired by the byte-pair encoding (BPE) text compression algorithm, which has recently gained popularity in…
Friederike Tegge, Katharina Parry, Diego Raphael Amancio
The text-evaluation application Coh-Metrix and natural language processing rely on the sentence for text segmentation and analysis and frequently detect sentence limits by means of punctuation. Problems arise when target texts such as pop song lyrics do not follow formal standards of written text composition and lack…
Authors not listed
The exponential growth of chemical literature necessitates the development of automated tools for extracting and curating molecular information from unstructured scientific publications into open-access chemical databases. Current optical chemical structure recognition (OCSR) and named entity recognition solutions…
Alexander E. Siemenn, Eunice Aissi, Fang Sheng, Armi Tiihonen + 3 more
In materials research, the task of characterizing hundreds of different materials traditionally requires equally many human hours spent measuring samples one by one. We demonstrate that with the integration of computer vision into this material research workflow, many of these tasks can be automated, significantly…
Chonghuan Zhang, Adarsh Arun, Alexei Lapkin
Computer Aided Synthesis Planning (CASP) development of reaction routes requires understanding of complete reaction structures. However, most reactions in the current databases are missing reaction co-participants. Although reaction prediction and atom mapping tools can predict major reaction participants and trace…