23 papers · ranked by Valyu relevance
Salman Ahmad Ansari, Usman Zafar, Asim Karim
Text normalization is an essential task in the processing and analysis of social media that is dominated with informal writing. It aims to map informal words to their intended standard forms. In this paper, we present an automatic optimization-based nearest neighbor matching approach for text normalization. This…
Shahzad Nazir, Muhammad Asif, Mariam Rehman, Shahbaz Ahmad + 1 more
'Xiangjie Kong'] In text applications, pre-processing is deemed as a significant parameter to enhance the outcomes of natural language processing (NLP) chores. Text normalization and tokenization are two pivotal procedures of text pre-processing that cannot be overstated. Text normalization refers to transforming raw…
Sebastian Duesing, Jason Bennett, James A. Overton, Randi Vita + 1 more
'Bjoern Peters'] Background While unstructured data, such as free text, constitutes a large amount of publicly available biomedical data, it is underutilized in automated analyses due to the difficulty of extracting meaning from it. Normalizing free-text data, i.e., removing inessential variance, enables the use of…
Sebastian Duesing, Jason Bennett, James A. Overton, Randi Vita + 1 more
ADP is a non-fully-automated normalization tool that enables a user to create standardization rules and apply them to datasets, which is available on GitHub . The ADP normalization scripts are written in Python version 3.10. The core normalization scripts import the libraries os, re, and sys from the Python Standard…
Fenil Doshi, Jimit Gandhi, Deep Gosalia, Sudhir Bagul
Social media networks and chatting platforms often use an informal version of natural text. Adversarial spelling attacks also tend to alter the input text by modifying the characters in the text. Normalizing these texts is an essential step for various applications like language translation and text to speech synthesis…
Anurag Roy, Shalmoli Ghosh, Kripabandhu Ghosh, Saptarshi Ghosh
A large fraction of textual data available today contains various types of 'noise', such as OCR noise in digitized documents, noise due to informal writing style of users on microblogging sites, and so on. To enable tasks such as search/retrieval and classification over all the available data, we need robust algorithms…
Joanna Bitton, Maya Pavlova, Ivan Evtimov
Text-based adversarial attacks are becoming more commonplace and accessible to general internet users. As these attacks proliferate, the need to address the gap in model robustness becomes imminent. While retraining on adversarial data may increase performance, there remains an additional class of character-level…
Alvaro Barreiro-Garrido, Victoria Ruiz-Parrado, A. Belen Moreno, Jose F. Velez
'Jose F. Velez'] In the realm of offline handwritten text recognition, numerous normalization algorithms have been developed over the years to serve as preprocessing steps prior to applying automatic recognition models to handwritten text scanned images. These algorithms have demonstrated effectiveness in enhancing the…
Slobodan Beliga, Miran Pobar, Sanda Martinčić-Ipšić
This paper presents text normalization which is an integral part of any text-to-speech synthesis system. Text normalization is a set of methods with a task to write non-standard words, like numbers, dates, times, abbreviations, acronyms and the most common symbols, in their full expanded form are presented. The whole…
Stefano Lusito, Edoardo Ferrante, Jean Maillard
In this paper we examine the case of text normalization for Ligurian, an endangered Romance language. We collect 4,394 Ligurian sentences paired with their normalized versions, as well as the first open source monolingual corpus for Ligurian. We show that, in spite of the small amounts of data available, a compact…
Ke Wu, Kyle Gorman, Richard Sproat
In speech-applications such as text-to-speech (TTS) or automatic speech recognition (ASR), text normalization refers to the task of converting from a written representation into a representation of how the text is to be spoken. In all real-world speech applications, the text normalization engine is developed—in large…
Bharathi Raja Chakravarthi, Priya Rani, Mihael Arcan, John P. McCrae
Machine translation is one of the applications of natural language processing which has been explored in different languages. Recently researchers started paying attention towards machine translation for resource-poor languages and closely related languages. A widespread and underlying problem for these machine…
Fatima Zohra Smaili, Xin Gao, Robert Hoehndorf
Ontologies are widely used in biomedicine for the annotation and standardization of data. One of the main roles of ontologies is to provide structured background knowledge within a domain as well as a set of labels, synonyms, and definitions for the classes within a domain. The two types of information provided by…
Shadrack Barnabas, Timo Böhme, Stephen Boyer, Matthias Irmer + 4 more
The extraction of chemical information from documents is a demanding task in cheminformatics due to the variety of text and image-based representations of chemistry. The present work describes the extraction of chemical compounds with unique chemical structures from the open access CORE (COnnecting REpositories) and…
Lis Arend, Klaudia Adamowicz, Johannes R. Schmidt, Yuliya Burankova + 7 more
Despite the significant progress in accuracy and reliability in mass spectrometry technology, as well as the development of strategies based on isotopic labeling or internal standards in recent decades, systematic biases originating from non-biological factors remain a significant challenge in data analysis. In…
Zhenfeng Wu, Weixiang Liu, Haishuo Ji, Deshui Yu + 4 more
Data normalization is a crucial step in the gene expression analysis as it determines the validity of its downstream analyses. Although many metrics has been designed to evaluate the relative success of these methods, the results by different metrics did not show consistency. Based on the previous work, we designed a…
Authors not listed
Iron, the most abundant element on Earth by mass (34.6%), primarily exists as iron minerals due to its inherent reactivity. The study of iron mineral phase transformations under changing environmental conditions remains an important research focus due to its geological, environmental, and industrial significance. Yet…
Michael B. Cole, Davide Risso, Allon Wagner, David DeTomaso + 4 more
Systematic measurement biases make data normalization an essential preprocessing step in single-cell RNA sequencing (scRNA-seq) analysis. There may be multiple, competing considerations behind the assessment of normalization performance, some of them study-specific. Because normalization can have a large impact on…
Diem-Trang T. Tran, Aditya Bhaskara, Matthew Might, Balagurunathan Kuberan
The use of RNA-sequencing has garnered much attention in the recent years for characterizing and understanding various biological systems. However, it remains a major challenge to gain insights from a large number of RNA-seq experiments collectively, due to the normalization problem. Current normalization methods are…
Authors not listed
The scarcity and expense of fatigue data limits optimal design of components and constrains companies to a few well qualified materials when safety-critical applications are concerned. This research investigates different strategies to improve extraction of structured information from unstructured scientific…
Meng Wang, Lihua Jiang, Ruiqi Jian, Joanne Y. Chan + 3 more
Data normalization is an important step in processing proteomics data generated in mass spectrometry (MS) experiments, which aims to reduce sample-level variation and facilitate comparisons of samples. Previously published methods for normalization primarily depend on the assumption that the distribution of protein…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Authors not listed
Meteorological normalization is a key concept in studying anthropogenic effects on air pollutant concentrations and its temporal trends. While apparently successful in revealing anthropogenic effects and often used, there are downsides to the methods and limitations which should be taken into account when using it.…