23 papers · ranked by Valyu relevance
Evelina Bakhturina, Yang Zhang, Boris Ginsburg
Text normalization (TN) systems in production are largely rule-based using weighted finite-state transducers (WFST). However, WFST-based systems struggle with ambiguous input when the normalized form is context-dependent. On the other hand, neural text normalization systems can take context into account but they suffer…
Joanna Bitton, Maya Pavlova, Ivan Evtimov
Text-based adversarial attacks are becoming more commonplace and accessible to general internet users. As these attacks proliferate, the need to address the gap in model robustness becomes imminent. While retraining on adversarial data may increase performance, there remains an additional class of character-level…
Wenlin Dai, Changhe Song, Xiang Li, Zhiyong Wu + 3 more
'Xiulin Li' 'Helen Meng'] Text normalization, defined as a procedure transforming nonstandard words to spoken-form words, is crucial to the intelligibility of synthesized speech in text-to-speech system. Rule-based methods without considering context can not eliminate ambiguation, whereas sequence-to-sequence neural…
Stefano Lusito, Edoardo Ferrante, Jean Maillard
In this paper we examine the case of text normalization for Ligurian, an endangered Romance language. We collect 4,394 Ligurian sentences paired with their normalized versions, as well as the first open source monolingual corpus for Ligurian. We show that, in spite of the small amounts of data available, a compact…
Shahzad Nazir, Muhammad Asif, Mariam Rehman, Shahbaz Ahmad + 1 more
'Xiangjie Kong'] In text applications, pre-processing is deemed as a significant parameter to enhance the outcomes of natural language processing (NLP) chores. Text normalization and tokenization are two pivotal procedures of text pre-processing that cannot be overstated. Text normalization refers to transforming raw…
Sebastian Duesing, Jason Bennett, James A. Overton, Randi Vita + 1 more
ADP is a non-fully-automated normalization tool that enables a user to create standardization rules and apply them to datasets, which is available on GitHub . The ADP normalization scripts are written in Python version 3.10. The core normalization scripts import the libraries os, re, and sys from the Python Standard…
Sebastian Duesing, Jason Bennett, James A. Overton, Randi Vita + 1 more
'Bjoern Peters'] Background While unstructured data, such as free text, constitutes a large amount of publicly available biomedical data, it is underutilized in automated analyses due to the difficulty of extracting meaning from it. Normalizing free-text data, i.e., removing inessential variance, enables the use of…
Alvaro Barreiro-Garrido, Victoria Ruiz-Parrado, A. Belen Moreno, Jose F. Velez
'Jose F. Velez'] In the realm of offline handwritten text recognition, numerous normalization algorithms have been developed over the years to serve as preprocessing steps prior to applying automatic recognition models to handwritten text scanned images. These algorithms have demonstrated effectiveness in enhancing the…
Jiaxian Shen, Fangqiong Ling, Erica M. Hartmann
As the scientific literature grows exponentially and research becomes increasingly interdisciplinary, accurate and high-throughput reference deduplication is vital in evidence synthesis studies (e.g., systematic reviews, meta-analyses) to ensure the completeness of datasets while reducing the manual screening burden.…
Anton Ehrmanntraut
Language Modeling Authors: ['Anton Ehrmanntraut'] Historic variations of spelling poses a challenge for full-text search or natural language processing on historical digitized texts. To minimize the gap between the historic orthography and contemporary spelling, usually an automatic orthographic normalization of the…
Sina Ahmadi, Antonios Anastasopoulos
The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages. This, however, comes with certain challenges in script normalization, particularly where the speakers of a language in a bilingual community…
Kacper Dudzic, Filip Graliński, Krzysztof Jassem, Marek Kubis + 1 more
'Piotr Wierzchoń'] This paper discusses two approaches to the diachronic normalization of Polish texts: a rulebased solution that relies on a set of handcrafted patterns, and a neural normalization model based on the text-to-text transfer transformer architecture. The training and evaluation data prepared for the task…
Anh Thi-Hoang Nguyen, Dung Nguyen, Nguyet Nguyen, Khanh Ho + 1 more
'Kiet Van Nguyen'] Abstract. Social media data is a valuable resource for research, yet it contains a wide range of non-standard words (NSW). These irregularities hinder the effective operation of NLP tools. Current state-of-the-art methods for the Vietnamese language address this issue as a problem of lexical…
Erik Faessler, Udo Hahn, Sascha Schäuble
Knowledge about interactions between genes and proteins is vital for bio-molecular research. A large part of this knowledge is published in written text and not accessible in a structured way. To remedy this situation, several repositories of automatically extracted interaction facts were proposed over the years.…
Shadrack Barnabas, Timo Böhme, Stephen Boyer, Matthias Irmer + 4 more
The extraction of chemical information from documents is a demanding task in cheminformatics due to the variety of text and image-based representations of chemistry. The present work describes the extraction of chemical compounds with unique chemical structures from the open access CORE (COnnecting REpositories) and…
Shadrack Barnabas, Timo Böhme, Stephen Boyer, Matthias Irmer + 5 more
The extraction of chemical information from documents is a demanding task in cheminformatics due to the variety of text and image-based representations of chemistry. The present work describes the extraction of chemical compounds with unique chemical structures from the open access CORE (COnnecting REpositories) and…
Lis Arend, Klaudia Adamowicz, Johannes R. Schmidt, Yuliya Burankova + 7 more
Despite the significant progress in accuracy and reliability in mass spectrometry technology, as well as the development of strategies based on isotopic labeling or internal standards in recent decades, systematic biases originating from non-biological factors remain a significant challenge in data analysis. In…
Authors not listed
Iron, the most abundant element on Earth by mass (34.6%), primarily exists as iron minerals due to its inherent reactivity. The study of iron mineral phase transformations under changing environmental conditions remains an important research focus due to its geological, environmental, and industrial significance. Yet…
Vikas Singh, Nikhil Kirtipal, Byong-Sop Song, Sunjae Lee
The normalization of RNA sequencing data is a primary step for downstream analysis. The most popular method used for the normalization is the trimmed mean of M values (TMM) and DESeq. The TMM tries to trim away extreme log fold changes of the data to normalize the raw read counts based on the remaining…
Vikas Singh, Nikhil Kirtipal, Songwon Lim, Sunjae Lee
Normalization of single-cell RNA-seq (scRNA-seq) is a crucial step in downstream analysis, where raw data are adjusted to correct unwanted factors that prevent the direct comparison of expression measures. scRNA-seq data exhibits a multivariate relationship between transcript-specific expression and sequencing depth…
Annie W. Shieh, Sandeep K. Bansal, Zhen Zuo, Sidney H. Wang
Acute cellular stress is known to induce a global reduction in protein translation through suppression of cap dependent translation. However, selective translation in response to acute stress has been shown to play important roles in regulating the stress response. An accurate transcriptome-wide profile of acute…
Authors not listed
The scarcity and expense of fatigue data limits optimal design of components and constrains companies to a few well qualified materials when safety-critical applications are concerned. This research investigates different strategies to improve extraction of structured information from unstructured scientific…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…