24 papers · ranked by Valyu relevance
Diksha Khurana, Aditya Koli, Kiran Khatter, Sukhdev Singh
Natural language processing (NLP) has recently gained much attention for representing and analysing human language computationally. It has spread its applications in various fields such as machine translation, email spam detection, information extraction, summarization, medical, and question answering etc. The paper…
Sumaia Mohammed Al-Ghuribi, Shahrul Azman Mohd Noah
The rapid growth of the internet has increased the number of online texts. This led to the rapid growth of the number of online texts in the Arabic language. The enormous amount of text must be organized into classes to make the analysis process and text retrieval easier. Text classification is, therefore, a key…
Ingmar Böschen
JATSdecoder is a general toolbox which facilitates text extraction and analytical tasks on NISO-JATS coded XML documents. Its function JATSdecoder() outputs metadata, the abstract, the sectioned text and reference list as easy selectable elements. One of the biggest repositories for open access full texts covering…
Julia C. Lensing, John Y. Choe, Branden B. Johnson, Jingwen Wang + 1 more
'Toqir Rana'] Many practical disaster reports are published daily worldwide in various forms, including after-action reports, response plans, impact assessments, and resiliency plans. These reports serve as vital resources, allowing future generations to learn from past events and better mitigate and prepare for future…
Maxwell J. Farrell, Liam Brierley, Anna Willoughby, Andrew Yates + 1 more
Ecology and evolutionary biology, like other scientific fields, are experiencing an exponential growth of academic manuscripts. As domain knowledge accumulates, scientists will need new computational approaches for identifying relevant literature to read and include in formal literature reviews and meta-analyses.…
Roni Ramon-Gonen, Amir Dori, Shahar Shelly
Healthcare professionals produce abounding textual data in their daily clinical practice. Text mining can yield valuable insights from unstructured data. Extracting insights from multiple information sources is a major challenge in computational medicine. In this study, our objective was to illustrate how combining…
Maxwell J. Farrell, Nicolas Le Guillarme, Liam Brierley, Bronwen Hunter + 4 more
'Bronwen Hunter' 'Daan Scheepens' 'Anna Willoughby' 'Andrew Yates' 'Nicole Mideo'] In ecology and evolutionary biology, the synthesis and modelling of data from published literature are commonly used to generate insights and test theories across systems. However, the tasks of searching, screening, and extracting data…
Ziqi Zhang, Tomas Jasaitis, Richard Freeman, Rowida Alfrjani + 1 more
'Adam Funk'] Abstract. While text mining and NLP research has been established for decades, there remain gaps in the literature that reports the use of these techniques in building real-world applications. For example, they typically look at single and sometimes simplified tasks, and do not discuss in-depth data…
Wenke Xiao, Lijia Jing, Yaxin Xu, Shichao Zheng + 2 more
The amount of medical text data is increasing dramatically. Medical text data record the progress of medicine and imply a large amount of medical knowledge. As a natural language, they are characterized by semistructured, high-dimensional, high data volume semantics and cannot participate in arithmetic operations.…
Jay Kumar
A text stream is an ordered sequence of text documents generated over time. A massive amount of such text data is generated by online social platforms every day. Designing an algorithm for such text streams to extract useful information is a challenging task due to unique properties of the stream such as infinite…
Jérôme Dockès, Kendra Oudyk, Mohammad Torabi, Alejandro I de la Vega + 1 more
Automated analysis of the biomedical literature (literature-mining) offers a rich source of insights. However, such analysis requires collecting a large number of articles and extracting and processing their content. This task is often prohibitively difficult and time-consuming. Here, we provide tools to easily…
Muzamil Malik, Waqar Aslam, Zahid Aslam, Abdullah Alharbi + 2 more
People's lives are influenced by social media. It is an essential source for sharing news, awareness, detecting events, people's interests, etc. Social media covers a wide range of topics and events to be discussed. Extensive work has been published to capture the interesting events and insights from datasets. Many…
Wolfgang Emanuel Zurrer, Amelia Elaine Cannon, Ewoud Ewing, Marianna Rosso + 2 more
Systematic reviews, i.e., research summaries that address focused questions in a structured and reproducible manner, are a cornerstone of evidence-based medicine and research. However, certain systematic review steps such as data extraction are labour-intensive which hampers their applicability, not least with the…
Kishlay Jha
Recent progress in biological, medical and health-care technologies, and innovations in wearable sensors provide us with unprecedented opportunities to accumulate massive data to understand disease prognosis and develop personalized treatments and interventions. These massive data supplemented with rapid growth in…
Stephen Meisenbacher, Peter Norlander
We describe a method and new no-code software tools enabling domain experts to build custom structured, labeled datasets from the unstructured text of documents and build niche machine learning text classification models traceable to expert-written rules. The Context Rule Assisted Machine Learning (CRAML) method allows…
Joseph Manning, Lev Sarkisov
With the continuously growing number of scientific articles on synthesis of nanomaterials, it becomes impossible for researchers to grasp and comprehend the landscape of synthetic protocols available for a particular material. The aim of this study is to explore the feasibility of extracting the collective knowledge on…
Joseph Cornelius, Harald Detering, Oscar Lithgow-Serrano, Donat Agosti + 2 more
The fields of taxonomy and biodiversity research have witnessed an exponential growth in published literature. This vast corpus of articles holds information on the diverse biological traits of organisms and their ecologies. However, access to and extraction of relevant data from this extensive resource remain…
Authors not listed
Iron, the most abundant element on Earth by mass (34.6%), primarily exists as iron minerals due to its inherent reactivity. The study of iron mineral phase transformations under changing environmental conditions remains an important research focus due to its geological, environmental, and industrial significance. Yet…
Walid Bedhiafi, Véronique Thomas-Vaslin, Amel Benammar Elgaaied, Adrien Six
The automatic mining for bibliography exploitation in given contexts is a challenge according to the increasing number of scientific publications and new concepts. Several indexing systems were developed for biomedical literature. However, such systems have failed to produce contextualised research of genes and…
Akshansh Mishra, Vijaykumar S. Jatti, Vaishnavi More, Anish Dasgupta + 2 more
'Devarrishi Dixit' 'Eyob Messele Sefene'] Abstract: The ability to interpret spoken language is connected to natural language processing. It involves teaching the AI how words relate to one another, how they are meant to be used, and in what settings. The goal of natural language processing (NLP) is to get a machine…
Amjad Zia, Muzzamil Aziz, Ioana Popa, Sabih Ahmed Khan + 3 more
'Amirreza Fazely Hamedani' 'Abdul R. Asif' 'Yu-Feng Hu'] Understanding published unstructured textual data using traditional text mining approaches and tools is becoming a challenging issue due to the rapid increase in electronic open-source publications. The application of data mining techniques in the medical…
Authors not listed
The scarcity and expense of fatigue data limits optimal design of components and constrains companies to a few well qualified materials when safety-critical applications are concerned. This research investigates different strategies to improve extraction of structured information from unstructured scientific…
Pulan Yu
Associative classification mining (ACM) integrating association rule mining and classification has become a significant tool for knowledge discovery, especially in the chemical domain. Its major advantage is providing high accuracy as well as chemically interpretable models. Additionally, it is able to find…
Authors not listed
Artificial intelligence (AI) is reshaping scientific research by accelerating discovery and enabling the analysis of complex data that traditional methods struggle to handle. This review examines over 310,000 journal articles and patents from the CAS Content Collection (2015–2025), with a focus on, biomedical research…