26 papers · ranked by Valyu relevance
Rania Albalawi, Tet Hin Yeap, Morad Benyoucef
With the growth of online social network platforms and applications, large amounts of textual user-generated content are created daily in the form of comments, reviews, and short-text messages. As a result, users often find it challenging to discover useful information or more on the topic being discussed from such…
Yaakov HaCohen-Kerner, Daniel Miller, Yair Yigal, Weinan Zhang
Text classification (TC) is the task of automatically assigning documents to a fixed number of categories. TC is an important component in many text applications. Many of these applications perform preprocessing. There are different types of text preprocessing, e.g., conversion of uppercase letters into lowercase…
Marco Siino, Ilenia Tinnirello, Marco La Cascia
| 1 | | Introduction | | | | | | | | | | | | | | | | | | | | | | | | | | | 3 | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | | 1.1 | | Overview | and | | | contributions | | | . | .…
Antoine Ly, Benno Uthayasooriyar, Tingting Wang
Text is the most widely used means of communication today. This data is abundant but nevertheless complex to exploit within algorithms. For years, scientists have been trying to implement different techniques that enable computers to replicate some mechanisms of human reading. During the past five years, research…
Braga, Marco, Milanese, Gian Carlo + 2 more
—Text preprocessing is a fundamental component of Natural Language Processing, involving techniques such as stopword removal, stemming, and lemmatization to prepare text as input for further processing and analysis. Despite the context-dependent nature of the above techniques, traditional methods usually ignore…
José Camacho-Collados, Mohammad Taher Pilehvar
Text preprocessing is often the first step in the pipeline of a Natural Language Processing (NLP) system, with potential impact in its final performance. Despite its importance, text preprocessing has not received much attention in the deep learning literature. In this paper we investigate the impact of simple text…
Aijaz Ahmad Reshi, Furqan Rustam, Wajdi Aljedaani, Shabana Shafi + 9 more
COVID-19 pandemic has caused a global health crisis, resulting in endless efforts to reduce infections, fatalities, and therapies to mitigate its after-effects. Currently, large and fast-paced vaccination campaigns are in the process to reduce COVID-19 infection and fatality risks. Despite recommendations from…
Rita Kukafka, David Mimno, Daniel Low, Joseph Plasek + 9 more
'Stephanie S Merkouris' 'Gypsy A O’Dea' 'Lauren M Francis' 'Christopher J Greenwood' 'Matthew Fuller-Tyszkiewicz' 'Elizabeth M Westrupp' 'Jacqui A Macdonald' 'George J Youssef'] Background Topic modeling approaches allow researchers to analyze and represent written texts. One of the commonly used approaches in…
Zhangcheng Qiang, Kerry Taylor, Weiqing Wang
Matching? Authors: ['Zhangcheng Qiang' 'Kerry Taylor' 'Weiqing Wang'] Abstract—The generic text preprocessing pipeline, comprising Tokenisation, Normalisation, Stop Words Removal, and Stemming/Lemmatisation, has been implemented in many ontology matching (OM) systems. However, the lack of standardisation in text…
Helena Gómez-Adorno, Ilia Markov, Grigori Sidorov, Juan-Pablo Posadas-Durán + 2 more
'Juan-Pablo Posadas-Durán' 'Miguel A. Sanchez-Perez' 'Liliana Chanona-Hernandez'] We introduce a lexical resource for preprocessing social media data. We show that a neural network-based feature representation is enhanced by using this resource. We conducted experiments on the PAN 2015 and PAN 2016 author profiling…
Dezheng Zhang, Jing Li, Yonghong Xie, Aziguli Wulamu + 1 more
Text pre-processing is an important component of a Chinese text classification. At present, however, most of the studies on this topic focus on exploring the influence of preprocessing methods on a few text classification algorithms using English text. In this paper we experimentally compared fifteen commonly used…
Wilson Fearn, Orion Weller, Kevin Seppi
Text classification is a significant branch of natural language processing, and has many applications including document classification and sentiment analysis. Unsurprisingly, those who do text classification are concerned with the run-time of their algorithms, many of which depend on the size of the corpus' vocabulary…
Ayoub Bagheri, Daniel Oberski, Arjan Sammani, Peter G.M. van der Heijden + 1 more
With the increasing use of unstructured text in electronic health records, extracting useful related information has become a necessity. Text classification can be applied to extract patients’ medical history from clinical notes. However, the sparsity in clinical short notes, that is, excessively small word counts in…
J. Liao, S. Ananiadou, L. G. Currie, B. E. Howard + 5 more
The amount of published in vivo studies and the speed researchers are publishing them make it virtually impossible to follow the recent development in the field. Systematic review emerged as a method to summarise and analyse the studies quantitatively and critically but it is often out-of-date due to its lengthy…
Joseph Manning, Lev Sarkisov
With the continuously growing number of scientific articles on synthesis of nanomaterials, it becomes impossible for researchers to grasp and comprehend the landscape of synthetic protocols available for a particular material. The aim of this study is to explore the feasibility of extracting the collective knowledge on…
Jiayuan Ding, Zhongyu Xing, Yixin Wang, Renming Liu + 9 more
Preprocessing is a critical step in single-cell data analysis, yet current practices remain largely a black-box, trial-and-error process driven by user intuition, legacy defaults, and ad hoc heuristics. The optimal combination of steps such as normalization, gene selection, and dimensionality reduction varies across…
Long Qian, Xin Lu, Parvez Haris, Jianyong Zhu + 2 more
Clinical trials are crucial for drug development, but they require significant time and financial resources. Additionally, uncertainties may arise during these trials concerning their results due to concerns surrounding effectiveness, safety, or the enrollment of participants. If robust AI (artificial intelligence)…
Nadine S. J. Jacobsen, Daniel Kristanto, Suong Welp, Yusuf Cosku Inceler + 1 more
Preprocessing is necessary to extract meaningful results from electroencephalography (EEG) data. With many possible preprocessing choices, their impact on outcomes is fundamental. While previous studies have explored the effects of preprocessing on stationary EEG data, this research delves into mobile EEG, where…
Diogo de Jesus Soares Machado, Camilla Reginatto De Pierri, Letícia Graziela Costa Santos, Leonardo Scapin + 4 more
The large amount of existing textual data justifies the development of new text mining tools. Bioinformatics tools can be brought to Text Mining, increasing the arsenal of resources. Here, we present BIOTEXT, a package of strategies for converting natural language text into biological-like information data, providing a…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Authors not listed
The exponential growth of chemical literature necessitates the development of automated tools for extracting and curating molecular information from unstructured scientific publications into open-access chemical databases. Current optical chemical structure recognition (OCSR) and named entity recognition solutions…
Paola Bonizzoni, Tamara Ceccato, Gianluca Della Vedova, Luca Denti + 3 more
Recent advances in high throughput RNA-Seq technologies allow to produce massive datasets. When a study focuses only on a handful of genes, most reads are not relevant and degrade the performance of the tools used to analyze the data. Removing such useless reads from the input dataset leads to improved efficiency…
Julian Ivanov, Alan Lipkus, Haitao Chen, Chris Aultman + 3 more
A novel bibliometric methodology based on natural language data processing for identifying emerging topics in science is presented. Along with the usual practice of data collection and preprocessing, our method includes a natural language processing (NLP) technique and an innovative mathematical function data…
Sanjar Adilov
Generative neural networks have shown promising results in de novo drug design. Recent studies suggest that one of the efficient ways to produce novel molecules matching target properties is to model SMILES sequences using deep learning in a way similar to language modeling in natural language processing. In this…
Hasan M. Sayeed, Sterling G. Baird, Taylor D. Sparks
Capturing structure-property relationships of materials for property prediction using machine learning requires the representation or featurization of the structural aspects of materials at different levels, including atomic, crystal, and microscales. While crystal structure-based modeling techniques are effective for…
Authors not listed
Iron, the most abundant element on Earth by mass (34.6%), primarily exists as iron minerals due to its inherent reactivity. The study of iron mineral phase transformations under changing environmental conditions remains an important research focus due to its geological, environmental, and industrial significance. Yet…