24 papers · ranked by Valyu relevance
Braga, Marco, Milanese, Gian Carlo + 2 more
—Text preprocessing is a fundamental component of Natural Language Processing, involving techniques such as stopword removal, stemming, and lemmatization to prepare text as input for further processing and analysis. Despite the context-dependent nature of the above techniques, traditional methods usually ignore…
K. M. Rani Krishna, K. Somasundaram, P. Arulmozhivarman, Sarah A. Immanuel + 1 more
Text Summarization, a vital aspect of natural language processing, aims to condense text while retaining its essential meaning. This process is achieved through extractive and abstractive methods. Deep Learning faces challenges in this domain, including semantic understanding, preservation of meaning, efficient…
Weiqing He, Bojian Hou, Amy Zheng, Yanbo Feng + 8 more
Despite careful initial data collection, social media content inherently contains noise, spam, and inconsistencies that could impact topic modeling quality. To address these challenges, we developed a comprehensive preprocessing pipeline that systematically improved data quality while preserving meaningful content. Our…
Darmono Darmono, Yanuar Agung Fadlullah, Khakam Ma’ruf, Muhamad Riyan Maulana + 5 more
Background The rapid growth of online learning platforms has transformed Technical and Vocational Education and Training, enabling broader access to skill based education through mobile applications. Understanding user acceptance is essential to ensure the sustainability and effectiveness of digital TVET platforms.…
Binxu Huang, Anasse Bari
We introduce a biologically inspired bird-flocking experimental framework for text summarization that identifies the most salient sentences using contextual information, sentence position, and thematic relevance. The bird-flocking-inspired algorithm, combined with large language models (LLMs), generates summaries with…
Saranzaya Magsarjav, Melissa Humphries, Jonathan Tuke, Lewis Mitchell
Sentiment analysis in Twitter datasets is important because it enables monitoring public opinion on products and analysis of political and social movements. One critical step is preprocessing: the automated processing of text for machine learning algorithms. Preprocessing plays a critical role in reducing noise and…
Virdio Samuel Saragih, Baruna Abirawa, Kartini Lovian Simbolon, Luluk Muthoharoh + 2 more
The rapid growth of electronic communication has necessitated more robust systems for email classification and sentiment detection. This study presents a comparative performance analysis between traditional machine learning algorithms and deep learning architectures, specifically focusing on Support Vector Machines…
Son Tran, Phuoc Tran, Manuel Herrador
This paper introduces a solution to the problem of detecting whether a sequence of text is Vietnamese based on its orthography and contextual features. For those unfamiliar with the language, it is known that understanding the meaning of certain texts can be challenging, since Vietnamese is a complex language that uses…
Shahzad Nazir, Muhammad Asif, Shahbaz Ahmad, Hanan Aljuaid + 2 more
This era has witnessed an enormous increase in textual corpus available in digital form. Therefore, an intelligent mechanism is required to extract the essential information. This task is performed using an automatic text summarization that converts the text into a shorter form while the semantics are preserved. The…
Md Sakhawat Hossen, Md. Zashid Iqbal Borshon, A. S. M. Badrudduza
The uprising of deep learning methodology and practice in recent years has brought about a severe consequence of increasing carbon footprint due to the insatiable demands on computational resources and power. The field of text analytics also experienced a massive transformation on this trend of monopolizing…
Yixiang Qu, Yifan Dai, Shilin Yu, Pradham Tanikella + 13 more
Large Language Models (LLMs) have demonstrated remarkable proficiency in automated text annotation within natural language processing. However, their deployment in clinical settings is severely constrained by strict privacy regulations and the prohibitive computational cost of processing voluminous unstructured…
Cem Rıfkı Aydın
| Prof. Tunga Güngör | | |--------------------------------|--| | (Thesis Supervisor) | | | | | | Assist. Prof. Tevfik Aytekin | | | | | | Prof. Fikret S. Gürgen | | | | | | Assoc. Prof. Arzucan Özgür | | | | | | Assist. Prof. Reyyan Yeniterzi | |
Authors not listed
Iron, the most abundant element on Earth by mass (34.6%), primarily exists as iron minerals due to its inherent reactivity. The study of iron mineral phase transformations under changing environmental conditions remains an important research focus due to its geological, environmental, and industrial significance. Yet…
Kenny Shao
Byte Pair Encoding (BPE) is widely used for subword tokenization, but standard BPE exposes every learned merge token to the downstream model, including tokens that mainly serve as intermediate construction units and rarely appear in the final encoded corpus. This paper proposes Pruned BPE, a post-training…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Jorge Radlowski Nova, Juan I. Lopez-Carbonero, Fernando J. Gómez-Márquez, José L. Risco + 2 more
Mixed-format lifestyle questionnaires contain both structured variables and free-text responses, but it remains unclear whether language-derived variables provide incremental predictive value beyond structured data, and under which representational condition. It was investigated whether variables derived from…
A. Sina Booeshaghi, Aaron Streets
Large language models excel at text extraction, but they sometimes hallucinate. A simple way to avoid hallucinations is to remove any extracted text that does not appear in the original source. This is easy when the extracted text is contiguous (findable with exact string matching), but much harder when it is…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…
Matthieu Vilain, Stéphane Aris-Brosou
The ever-growing amount of available biological data leads modern analysis to be performed on large datasets. Unfortunately, bioinformatics tools for preprocessing and analyzing data are not always designed to treat such large amounts of data efficiently. Notably, this is the case when encoding DNA and RNA sequences…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Md Rashedur Rahman
Zero-shot learning from functional magnetic resonance imaging (fMRI) data offers a principled approach to decoding conceptual knowledge without requiring training examples for every target concept. The Semantic Output Code (SOC) framework, introduced by Palatucci et al. [2009], operationalises this idea through a…
A.A. Poyda, V.A. Orlov, A.D. Zhemchuzhnikov, S.O. Kozlov + 4 more
The aim of the study was to analyze the influence of various pipelines for preprocessing raw functional magnetic resonance imaging (fMRI) data on the accuracy of classification of subjects into schizophrenia patients and healthy controls using machine learning methods, and to give recommendations for optimizing the…
Nadeen Kherbawy, Christine E Potter, Sagi Jaffe-Dax
Learning to read leads to widespread changes in brain organization, but it is not yet known when text first becomes a privileged stimulus. To test whether specialized neural responses to text appear prior to reading instruction, 31 monolingual toddlers in Israel (2.1-3.6 years) not yet enrolled in school were presented…
Authors not listed
Metal–organic frameworks (MOFs) represent a versatile class of porous materials, yet efficiently exploring their vast chemical space for target gas adsorption properties remains a major challenge. MOFid, a text-based encoding of MOF structures, has enabled large-scale data mining using natural language processing (NLP)…