23 papers · ranked by Valyu relevance
Lihao Wang, Xiaoqing Zheng
We present an unsupervised word segmentation model, in which the learning objective is to maximize the generation probability of a sentence given its all possible segmentation. Such generation probability can be factorized into the likelihood of each possible segment given the context in a recursive way. In order to…
Dingyi Niu, Jing Tian, Ramandeep Kaur
Word segmentation is crucial for reading in unspaced languages like Tibetan, where readers rely on high-level cues like morpheme positional frequency (the statistical likelihood of a morpheme appearing at the beginning or end of a word). Using a novel word learning paradigm, this study investigated whether initial and…
Tero Hakala, Tiina Lindh-Knuutila, Annika Hultén, Minna Lehtonen + 1 more
'Riitta Salmelin'] Title: Abstract This study extends the idea of decoding word-evoked brain activations using a corpus-semantic vector space to multimorphemic words in the agglutinative Finnish language. The corpus-semantic models are trained on word segments, and decoding is carried out with word vectors that are…
Dingyi Niu, Zijian Xie, Jiaqi Liu, Chen Wang + 2 more
'Daniela Zambarbieri'] This study utilized eye-tracking technology to explore the role of visual word segmentation cues in Tibetan reading, with a particular focus on the effects of dictionary-based and psychological word segmentation on reading and lexical recognition. The experiment employed a 2 × 3 design, comparing…
Matthew Rho, Yexin Tian, Qin Chen
We provide a detailed overview of various approaches to word segmentation of Asian Languages, specifically Chinese, Korean, and Japanese languages. For each language, approaches to deal with word segmentation differs. We also include our analysis about certain advantages and disadvantages to each method. In addition…
Zihong Zhang, Liqi He, Zuchao Li, Lefei Zhang + 2 more
Word segmentation stands as a cornerstone of Natural Language Processing (NLP). Based on the concept of "comprehend first, segment later", we propose a new framework to explore the limit of unsupervised word segmentation with Large Language Models (LLMs) and evaluate the semantic understanding capabilities of LLMs…
Yiu-Kei Tsang, Ming Yan, Jinger Pan, Megan Yin Kan Chan
The absence of explicit word boundaries is a distinctive characteristic of Chinese script, setting it apart from most alphabetic scripts, leading to word boundary disagreement among readers. Previous studies have examined how this feature may influence reading performance. However, further investigations are required…
Duc-Vu Nguyen, Ngan Luu-Thuy Nguyen
—To the best of our knowledge, this paper made the first attempt to answer whether word segmentation is necessary for Vietnamese sentiment classification. To do this, we presented five pre-trained monolingual S4-based language models for Vietnamese, including one model without word segmentation, and four models using…
Qing Xu, Zhiyou Wang
Chinese natural language processing tasks often require the solution of Chinese word segmentation and POS tagging problems. Traditional Chinese word segmentation and POS tagging methods mainly use simple matching algorithms based on lexicons and rules. The simple matching or statistical analysis requires manual word…
Ju-Xiao Zhang, Hai-Feng Chen, Bing Chen, Bei-Qin Chen + 2 more
'Jing-Hua Zhong' 'Xiao-Qin Zeng'] An important sign of the accessibility of Braille information is the realization of the mutual translation between Chinese and the Braille. Due to the irregularity and uncertainty of the Prevailing Mandarin Braille, coupled with the lack of a large-scale Braille corpus, the quality of…
Jindřich Libovický, Jindřich Helcl
We present three innovations in tokenization and subword segmentation. First, we propose to use unsupervised morphological analysis with Morfessor as pre-tokenization. Second, we present an algebraic method for obtaining subword embeddings grounded in a word embedding space. Based on that, we design a novel subword…
Hanna Ringer, Daniela Sammler, Tatsuya Daikoku
Listeners implicitly use statistical regularities to segment continuous sound input into meaningful units, e.g., transitional probabilities between syllables to segment a speech stream into separate words. Implicit learning of such statistical regularities in a novel stimulus stream is reflected in a synchronisation of…
Lucas Benjamin, Di Zang, Ana Fló, Zengxin Qi + 6 more
The debate over whether conscious attention is necessary for statistical learning has produced mixed and conflicting results. Testing individuals with impaired consciousness may provide some insight, but very few studies have been conducted due to the difficulties associated with testing such patients. In this study…
Wang Liang
Since the completion of the human genome sequencing project in 2001, significant progress has been made in areas such as gene regulation editing and protein structure prediction. However, given the vast amount of genomic data, the segments that can be fully annotated and understood remain relatively limited. If we…
Michel Godel, Ana Fló, Lucas Benjamin, Ghislaine Dehaene-Lambertz + 1 more
Delayed onset of canonical babbling and first words is often reported in infants later diagnosed with autism spectrum disorder. Identifying the neural mechanisms underlying language acquisition in autism is therefore critical to inform early diagnosis, prognosis, and intervention strategies. In this study, we…
Galit Agmon, Manuela Jaeger, Reut Tsarfaty, Martin G Bleichner + 1 more
Spontaneous real-life speech is imperfect in many ways. It contains disfluencies and ill-formed utterances that the brain needs to contend with in order to extract meaning out of speech. Here, we studied how the neural speech-tracking response is affected by three specific factors that are prevalent in spontaneous…
Benjamin Minixhofer, Jonas Pfeiffer, Ivan Vulić
Many NLP pipelines split text into sentences as one of the crucial preprocessing steps. Prior sentence segmentation tools either rely on punctuation or require a considerable amount of sentence-segmented training data: both central assumptions might fail when porting sentence segmenters to diverse languages on a…
Ole Bialas, Edmund C. Lalor
In recent decades, research on the neural processing of speech and language increasingly investigated ongoing responses to continuously presented naturalistic speech, allowing researchers to ask interesting questions about different representations of speech and their relationships. This requires statistical models…
Authors not listed
The exponential growth of chemical literature necessitates the development of automated tools for extracting and curating molecular information from unstructured scientific publications into open-access chemical databases. Current optical chemical structure recognition (OCSR) and named entity recognition solutions…
Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura
—Speech translation (ST) automatically converts utterances in a source language into text in another language. Splitting continuous speech into shorter segments, known as speech segmentation, plays an important role in ST. Recent segmentation methods trained to mimic the segmentation of ST corpora have surpassed…
Alexander E. Siemenn, Eunice Aissi, Fang Sheng, Armi Tiihonen + 3 more
In materials research, the task of characterizing hundreds of different materials traditionally requires equally many human hours spent measuring samples one by one. We demonstrate that with the integration of computer vision into this material research workflow, many of these tasks can be automated, significantly…
Yonghao Zhao, Zhouyuan Zhu, Sen Yang, Weihan Li
An essential step for quantitative image analysis is cell segmentation, which is the process of defining the outline of individual cells in microscopy images. Segmentation of budding yeast is challenging due to their asymmetric cell division and mother-bud morphology. As a result, a dividing cell is frequently…
Chonghuan Zhang, Adarsh Arun, Alexei Lapkin
Computer Aided Synthesis Planning (CASP) development of reaction routes requires understanding of complete reaction structures. However, most reactions in the current databases are missing reaction co-participants. Although reaction prediction and atom mapping tools can predict major reaction participants and trace…