14 papers · ranked by Valyu relevance
Jacob S. Berkowitz, Apoorva Srinivasan, Jose Miguel Acitores Cortina, Yasaman Fatapour + 1 more
1.## INTRODUCTION The global transition to digital clinical records has led to a drastic increase in electronically stored unstructured data, making up around 80%(1) of the information within healthcare. This data has the potential to improve patient care, deepen our understanding of diseases, and support research. Yet…
Shahzad Nazir, Muhammad Asif, Mariam Rehman, Shahbaz Ahmad + 1 more
'Xiangjie Kong'] In text applications, pre-processing is deemed as a significant parameter to enhance the outcomes of natural language processing (NLP) chores. Text normalization and tokenization are two pivotal procedures of text pre-processing that cannot be overstated. Text normalization refers to transforming raw…
Sebastian Duesing, Jason Bennett, James A. Overton, Randi Vita + 1 more
ADP is a non-fully-automated normalization tool that enables a user to create standardization rules and apply them to datasets, which is available on GitHub . The ADP normalization scripts are written in Python version 3.10. The core normalization scripts import the libraries os, re, and sys from the Python Standard…
Sebastian Duesing, Jason Bennett, James A. Overton, Randi Vita + 1 more
'Bjoern Peters'] Background While unstructured data, such as free text, constitutes a large amount of publicly available biomedical data, it is underutilized in automated analyses due to the difficulty of extracting meaning from it. Normalizing free-text data, i.e., removing inessential variance, enables the use of…
Alvaro Barreiro-Garrido, Victoria Ruiz-Parrado, A. Belen Moreno, Jose F. Velez
'Jose F. Velez'] In the realm of offline handwritten text recognition, numerous normalization algorithms have been developed over the years to serve as preprocessing steps prior to applying automatic recognition models to handwritten text scanned images. These algorithms have demonstrated effectiveness in enhancing the…
Bharathi Raja Chakravarthi, Priya Rani, Mihael Arcan, John P. McCrae
Machine translation is one of the applications of natural language processing which has been explored in different languages. Recently researchers started paying attention towards machine translation for resource-poor languages and closely related languages. A widespread and underlying problem for these machine…
Zainab Mansur, Nazlia Omar, Sabrina Tiun, Eissa M. Alshari + 1 more
As social media booms, abusive online practices such as hate speech have unfortunately increased as well. As letters are often repeated in words used to construct social media messages, these types of words should be eliminated or reduced in number to enhance the efficacy of hate speech detection. Although multiple…
Mohamed Osman Hegazi, Yasser Al-Dossari, Abdullah Al-Yahy, Abdulaziz Al-Sumari + 1 more
'Abdulaziz Al-Sumari' 'Anwer Hilal'] Currently, social media plays an important role in daily life and routine. Millions of people use social media for different purposes. Large amounts of data flow through online networks every second, and these data contain valuable information that can be extracted if the data are…
Alvin Subakti, Hendri Murfi, Nora Hariadi
Text clustering is the task of grouping a set of texts so that text in the same group will be more similar than those from a different group. The process of grouping text manually requires a significant amount of time and labor. Therefore, automation utilizing machine learning is necessary. One of the most frequently…
Danielle L. Mowery, Brett R. South, Lee Christensen, Jianwei Leng + 9 more
'Laura-Maria Peltonen' 'Sanna Salanterä' 'Hanna Suominen' 'David Martinez' 'Sumithra Velupillai' 'Noémie Elhadad' 'Guergana Savova' 'Sameer Pradhan' 'Wendy W. Chapman'] Background The ShARe/CLEF eHealth challenge lab aims to stimulate development of natural language processing and information retrieval technologies to…
Yoshimasa Tsuruoka, John McNaught, Sophia Ananiadou
Background One of the difficulties in mapping biomedical named entities, e.g. genes, proteins, chemicals and diseases, to their concept identifiers stems from the potential variability of the terms. Soft string matching is a possible solution to the problem, but its inherent heavy computational cost discourages its use…
Tianyong Hao, Buzhou Tang, Zhengxing Huang, Tianyong Hao + 7 more
'Zuofeng Li' 'Feichen Shen' 'Xiaoyi Pan' 'Boyu Chen' 'Heng Weng' 'Yongyi Gong' 'Yingying Qu'] Background Temporal information frequently exists in the representation of the disease progress, prescription, medication, surgery progress, or discharge summary in narrative clinical text. The accurate extraction and…
Yanshan Wang, Long Chen, Sérgio Matos, Sina Madani + 1 more
Background Clinical terms mentioned in clinical text are often not in their standardized forms as listed in clinical terminologies because of linguistic and stylistic variations. However, many automated downstream applications require clinical terms mapped to their corresponding concepts in clinical terminologies, thus…
Guocheng Wang, Yiwen Wang, Hui Li, Xuanqi Chen + 6 more
'Yanpeng Ma' 'Chun Peng' 'Yijun Wang' 'Linyao Tang' 'Rongrong Ji'] In this paper, some morphological transformations are used to detect the unevenly illuminated background of text images characterized by poor lighting, and to acquire illumination normalized result. Based on morphologic Top-Hat transform, the uneven…