26 papers · ranked by Valyu relevance
Varun Dogra, Sahil Verma, Kavita, Pushpita Chatterjee + 3 more
'Jaeyoung Choi' 'Muhammad Fazal Ijaz'] With the rapid advancement of information technology, online information has been exponentially growing day by day, especially in the form of text documents such as news events, company reports, reviews on products, stocks-related reports, medical reports, tweets, and so on. Due…
Sadia Zaman Mishu, S M Rafiuddin
The demand for text classification is growing significantly in web searching, data mining, web ranking, recommendation systems, and so many other fields of information and technology. This paper illustrates the text classification process on different dataset using some standard supervised machine learning techniques.…
Sunil Kumar Prabhakar, Harikumar Rajaguru, Kwangsub So, Dong-Ok Won
To classify the texts accurately, many machine learning techniques have been utilized in the field of Natural Language Processing (NLP). For many pattern classification applications, great success has been obtained when implemented with deep learning models rather than using ordinary machine learning techniques.…
Zhongwei Wan
In recent years, with the rapid development of information on the Internet, the number of complex texts and documents has increased exponentially, which requires a deeper understanding of deep learning methods in order to accurately classify texts using deep learning techniques, and thus deep learning methods have…
Hemn Barzan Abdalla, Awder M. Ahmed, Subhi R.M. Zeebaree, Ahmed Alkhayyat + 2 more
'Ahmed Alkhayyat' 'Baha Ihnaini' 'Xiangjie Kong'] Increasing demands for information and the rapid growth of big data have dramatically increased the amount of textual data. In order to obtain useful text information, the classification of texts is considered an imperative task. Accordingly, this article will describe…
Danyang Zheng, Shahid Mumtaz
In recent years, with the rapid development of the Internet and multimedia technology, English translation text classification has played an important role in various industries. However, English translation remains a complex and difficult problem. Seeking an efficient and accurate English translation method has become…
Mattias Wahde, Marco L. Della Vedova, Marco Virgolin, Minerva Suvanto
'Minerva Suvanto'] We investigate the differences between spoken language (in the form of radio show transcripts) and written language (Wikipedia articles) in the context of text classification. We present a novel, interpretable method for text classification, involving a linear classifier using a large set of $n-$gram…
Zhiqiang Wang, Yiran Pang, Yanbin Lin
—Retrained large language models (LLMs) have become extensively used across various sub-disciplines of natural language processing (NLP). In NLP, text classification problems have garnered considerable focus, but still faced with some limitations related to expensive computational cost, time consumption, and robust…
Karina Shyrokykh, Max Girnyk, Lisa Dellmuth, Nebojsa Bacanin
To analyse large numbers of texts, social science researchers are increasingly confronting the challenge of text classification. When manual labeling is not possible and researchers have to find automatized ways to classify texts, computer science provides a useful toolbox of machine-learning methods whose performance…
Mohsen Ahmadi, Matin Khajavi, Abbas Varmaghani, Ali Ala + 2 more
Detection with Robust and Context-Aware Text Classification Authors: ['Mohsen Ahmadi' 'Matin Khajavi' 'Abbas Varmaghani' 'Ali Ala' 'Kasra Danesh' 'Danial Javaheri'] Abstract—This study evaluates the effectiveness of different feature extraction techniques and classification algorithms in detecting spam messages within…
Xiaonan Xu, Xu Zheng, Zhipeng Ling, Zhengyu Jin + 1 more
between Natural Language Processing and System Recommendation Authors: ['Xiaonan Xu' 'Xu Zheng' 'Zhipeng Ling' 'Zhengyu Jin' 'Shuqian Du'] ABSTRACT Natural Language Processing (NLP) is an important branch of artificial intelligence that studies how to enable computers to understand, process, and generate human…
Lu Xiao, Qiaoxing Li, Qian Ma, Jiasheng Shen + 3 more
Text classification, as an important research area of text mining, can quickly and effectively extract valuable information to address the challenges of organizing and managing large-scale text data in the era of big data. Currently, the related research on text classification tends to focus on the application in…
Qinghua Wang, Jonathan Olshin, K. Vijay-Shanker, Cathy Wu
Chinese hamster ovary (CHO) cells are widely used for mass production of therapeutic proteins in the pharmaceutical industry. With the growing need in optimizing the performance of producer CHO cell lines, research on CHO cell line development and bioprocess continues to increase in recent decades. Bibliographic…
Simona Emilova Doneva, Sijing Qin, Beate Sick, Tilia Ellendorff + 3 more
The advent of large language models (LLMs) such as BERT and, more recently, GPT, is transforming our approach of analyzing and understanding biomedical texts. To stay informed about the latest advancements in this area, there is a need for up-to-date summaries on the role of LLM in Natural Language Processing (NLP) of…
İbrahim Karaman, Gülser Köksal, Levent Erişkin, Salih Salihoglu
Data Authors: ['İbrahim Karaman' 'Gülser Köksal' 'Levent Erişkin' 'Salih Salihoglu'] In real-world applications, as data availability increases, obtaining labeled data for machine learning (ML) projects remains challenging due to the high costs and intensive efforts required for data annotation. Many ML projects…
Junghoon Chae, David Heise, Keith Connatser, Jacqueline Honerlaw + 5 more
The demand for a comprehensive phenomics library, which requires identifying computable phenotype definitions and associated metadata from an ever-expanding biomedical literature, presents a significant, labor-intensive, and unscalable challenge. To address this, we introduce a transformer-based language model…
Hashim Ali, Raja Sarath Kumar Boddu, Umer Tanveer, Aamir Saeed + 4 more
Software requirements classification remains one of the important challenges in requirements engineering. Engineering that affects the smoothness of project success about software development life cycles. in this paper, a novel hybrid solution is being presented that beats the benchmarks set by previous approaches…
Pulan Yu
Associative classification mining (ACM) integrating association rule mining and classification has become a significant tool for knowledge discovery, especially in the chemical domain. Its major advantage is providing high accuracy as well as chemically interpretable models. Additionally, it is able to find…
Steven Tan, Sina Majidian, Ben Langmead, Mohsen Zakeri
The number of reference genomes is rapidly increasing, thanks to advances in long-read sequencing and assembly. While these collections can improve the sensitivity and specificity of classification methods, this requires highly efficient compressed indexes. K-mer-based approaches like Kraken 2 are efficient but limit…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
Nazila Ahmadi Daryakenari, Seyed Kamaleddin Setaredan
Schizophrenia (SZ) is a chronic and complex mental disorder associated with neurobiological deficits. The complexity and heterogeneity of schizophrenia symptoms pose challenges for objective diagnosis, which is currently based on behavioral and clinical manifestations. Furthermore, other psychiatric disorders such as…
Michael A. Zeller, Zebulun W. Arendsee, Gavin J.D. Smith, Tavis K. Anderson
Sequencing and phylogenetic classification have become a common task in human and animal diagnostic laboratories. It is routine to sequence pathogens to identify genetic variations of diagnostic significance and to use these data in real-time genomic contact tracing and surveillance. Under this paradigm, unprecedented…
Rachana Niranjan Murthy, Sai Teja Potu, Akhil Thomas, Lokesh Mishra + 2 more
Retrieving structured materials information from unstructured textual data is essential for data mining and automatically developing comprehensive ontologies. Information extraction is a complex task composed of multiple subtasks and thus often relies on systems of task-specialized language models. A foundation…
Long Qian, Xin Lu, Parvez Haris, Jianyong Zhu + 2 more
Clinical trials are crucial for drug development, but they require significant time and financial resources. Additionally, uncertainties may arise during these trials concerning their results due to concerns surrounding effectiveness, safety, or the enrollment of participants. If robust AI (artificial intelligence)…
Hasan M. Sayeed, Sterling G. Baird, Taylor D. Sparks
Capturing structure-property relationships of materials for property prediction using machine learning requires the representation or featurization of the structural aspects of materials at different levels, including atomic, crystal, and microscales. While crystal structure-based modeling techniques are effective for…