12 papers · ranked by Valyu relevance
Jacob S. Berkowitz, Apoorva Srinivasan, Jose Miguel Acitores Cortina, Yasaman Fatapour + 1 more
1.## INTRODUCTION The global transition to digital clinical records has led to a drastic increase in electronically stored unstructured data, making up around 80%(1) of the information within healthcare. This data has the potential to improve patient care, deepen our understanding of diseases, and support research. Yet…
Shahzad Nazir, Muhammad Asif, Mariam Rehman, Shahbaz Ahmad + 1 more
'Xiangjie Kong'] In text applications, pre-processing is deemed as a significant parameter to enhance the outcomes of natural language processing (NLP) chores. Text normalization and tokenization are two pivotal procedures of text pre-processing that cannot be overstated. Text normalization refers to transforming raw…
Sebastian Duesing, Jason Bennett, James A. Overton, Randi Vita + 1 more
ADP is a non-fully-automated normalization tool that enables a user to create standardization rules and apply them to datasets, which is available on GitHub . The ADP normalization scripts are written in Python version 3.10. The core normalization scripts import the libraries os, re, and sys from the Python Standard…
Sebastian Duesing, Jason Bennett, James A. Overton, Randi Vita + 1 more
'Bjoern Peters'] Background While unstructured data, such as free text, constitutes a large amount of publicly available biomedical data, it is underutilized in automated analyses due to the difficulty of extracting meaning from it. Normalizing free-text data, i.e., removing inessential variance, enables the use of…
Sanaa Kaddoura, Ganesh Chandrasekaran, Daniela Elena Popescu, Jude Hemanth Duraisamy + 1 more
'Jude Hemanth Duraisamy' 'Vimal Shanmuganathan'] The presence of spam content in social media is tremendously increasing, and therefore the detection of spam has become vital. The spam contents increase as people extensively use social media, i.e., Facebook, Twitter, YouTube, and E-mail. The time spent by people using…
Alvaro Barreiro-Garrido, Victoria Ruiz-Parrado, A. Belen Moreno, Jose F. Velez
'Jose F. Velez'] In the realm of offline handwritten text recognition, numerous normalization algorithms have been developed over the years to serve as preprocessing steps prior to applying automatic recognition models to handwritten text scanned images. These algorithms have demonstrated effectiveness in enhancing the…
Alvin Subakti, Hendri Murfi, Nora Hariadi
Text clustering is the task of grouping a set of texts so that text in the same group will be more similar than those from a different group. The process of grouping text manually requires a significant amount of time and labor. Therefore, automation utilizing machine learning is necessary. One of the most frequently…
Lu Zhou, Shuangqiao Liu, Caiyan Li, Yuemeng Sun + 6 more
'Yuda Li' 'Huimin Yuan' 'Yan Sun' 'Fengqin Xu' 'Yuhang Li'] Background The modernization of traditional Chinese medicine (TCM) demands systematic data mining using medical records. However, this process is hindered by the fact that many TCM symptoms have the same meaning but different literal expressions (i.e., TCM…
Zainab Mansur, Nazlia Omar, Sabrina Tiun, Eissa M. Alshari + 1 more
As social media booms, abusive online practices such as hate speech have unfortunately increased as well. As letters are often repeated in words used to construct social media messages, these types of words should be eliminated or reduced in number to enhance the efficacy of hate speech detection. Although multiple…
Poluru Eswaraiah, Hussain Syed, Natalia Kryvinska
Multimedia data, which includes textual information, is employed in a variety of practical computer vision applications. More than a million new records are added to social media and news sites every day, and the text content they contain has gotten increasingly complex. Finding a meaningful text record in an archive…
Anna Persson, T. Florian Jaeger
Talkers vary in the phonetic realization of their vowels. One influential hypothesis holds that listeners overcome this inter-talker variability through pre-linguistic auditory mechanisms that normalize the acoustic or phonetic cues that form the input to speech recognition. Dozens of competing normalization accounts…
Deying Song, Douglas Ruff, Marlene Cohen, Chengcheng Huang
Neurons in higher-order visual areas integrate information through a canonical computation called normalization. The strength of normalization is highly heterogeneous across neurons, and this heterogeneity correlates with attention-mediated modulations in neural responses. However, the circuit mechanism underlying the…