24 papers · ranked by Valyu relevance
Veronika Plotnikova, Marlon Dumas, Fredrik Milani, Sebastian Ventura
The use of end-to-end data mining methodologies such as CRISP-DM, KDD process, and SEMMA has grown substantially over the past decade. However, little is known as to how these methodologies are used in practice. In particular, the question of whether data mining methodologies are used ‘as-is’ or adapted for specific…
M. Zanin, D. Papo, P. A. Sousa, E. Menasalvas + 3 more
The increasing power of computer technology does not dispense with the need to extract meaningful in-formation out of data sets of ever growing size, and indeed typically exacerbates the complexity of this task. To tackle this general problem, two methods have emerged, at chronologically different times, that are now…
Seyed Abbas Mahmoodi, Kamal Mirzaie, Seyed Mostafa Mahmoudi
Cancer is the leading cause of death in economically developed countries and the second leading cause of death in developing countries. Gastric cancers are among the most devastating and incurable forms of cancer and their treatment may be excessively complex and costly. Data mining, a technology that is used to…
Richa Gupta
Data is the collection of values and variables related in certain sense and differing in some other sense. The size of data has always been increasing. Storing this data without using it in any sense is simply waste of storage space and storing time. Data should be processed to extract some useful knowledge from it.
Sudhir B. Jagtap, B. G. Kodge
Data mining (also known as knowledge discovery from databases) is the process of extraction of hidden, previously unknown and potentially useful information from databases. The outcome of the extracted data can be analyzed for the future planning and development perspectives. In this paper, we have made an attempt to…
Neelamadhab Padhy
In this paper we have focused a variety of techniques, approaches and different areas of the research which are helpful and marked as the important field of data mining Technologies. As we are aware that many MNC's and large organizations are operated in different places of the different countries. Each place of…
Francisco Azuaje
In the early 1990s some sectors of the computer science community were developing the idea of data understanding as a discovery-driven, systematic and iterative process. This "data mining" research and development area was expected to take advantage of the expansion and consolidation of machine learning methodologies…
Peyman Mohammadi, Abdolreza Hatamlou, Mohammad Masdari
In recent years, applications of data mining methods are become more popular in many fields of medical diagnosis and evaluations. The data mining methods are appropriate tools for discovering and extracting of available knowledge in medical databases. In this study, we divided 11 data mining algorithms into five groups…
Alfonso de la Vega, Diego García‐Saiz, Marta Zorrilla, Pablo Sánchez
Context: Data mining techniques have demonstrated to be a powerful technique for discovering insights hidden in data from a domain. However, these techniques demand very specialised skills. People willing to analyse data often lack these skills, so they must rely on data scientists, which hinders data mining…
Amjad Zia, Muzzamil Aziz, Ioana Popa, Sabih Ahmed Khan + 3 more
'Amirreza Fazely Hamedani' 'Abdul R. Asif' 'Yu-Feng Hu'] Understanding published unstructured textual data using traditional text mining approaches and tools is becoming a challenging issue due to the rapid increase in electronic open-source publications. The application of data mining techniques in the medical…
Yu-Min Wang, Chei-Chang Chiou, Wen-Chang Wang, Chun-Jung Chen
With the continuous progress and penetration of automated data collection technology, enterprises and organizations are facing the problem of information overload. The demand for expertise in data mining and analysis is increasing. Self-efficacy is a pivotal construct that is significantly related to willingness and…
Joseph C. Mellor, Michael A. Stone, John Keane
Principles and Potential Authors: ['Joseph C. Mellor' 'Michael A. Stone' 'John Keane'] The ubiquity and cheapness of miniature low-power sensors, digital processing, and large amounts of storage contained in small packages has heralded the ability to acquire large amounts of data about systems during their course of…
Mikhail Moshkov, Beata Zielosko, Evans Teiko Tetteh, Przemysław Juszczuk + 1 more
'Przemysław Juszczuk' 'Jan Kozak'] In this paper, we deal with distributed data represented either as a finite set $T$ of decision tables with equal sets of attributes or a finite set $I$ of information systems with equal sets of attributes. In the former case, we discuss a way to the study decision trees common to all…
Tipawan Silwattananusarn, Kulthida Tuamsuk
Data mining is one of the most important steps of the knowledge discovery in databases process and is considered as significant subfield in knowledge management. Research in data mining continues growing in business and in learning organization over coming decades. This review paper explores the applications of data…
B. Radhakrishnan, G Shineraj, K M Anver Muhammed
One of the most important problems in modern finance is finding efficient ways to summarize and visualize the stock market data to give individuals or institutions useful information about the market behavior for investment decisions. The enormous amount of valuable data generated by the stock market has attracted…
Suruchi Jai Kumar Ahuja
A major objective of clustering is to identify groups in the data that maximizes the similarity between objects within the same cluster and minimizes the similarity between different clusters. A challenge for data clustering, and unsupervised learning in general, is that there is often no mechanism for feature…
Ahmed BaniMustafa, Nigel Hardy
This work demonstrates the execution of a novel process model for knowledge discovery and data mining for metabolomics (MeKDDaM). It aims to illustrate MeKDDaM process model applicability using four different real-world applications and to highlight its strengths and unique features. The demonstrated applications…
Pulan Yu
Associative classification mining (ACM) integrating association rule mining and classification has become a significant tool for knowledge discovery, especially in the chemical domain. Its major advantage is providing high accuracy as well as chemically interpretable models. Additionally, it is able to find…
Adarsh Arun, Zhen Guo, Simon Sung, Alexei Lapkin
Automated prediction of reaction impurities can be useful in facilitating rapid early-stage reaction development, synthesis planning and optimization. Existing reaction predictors are catered towards main product prediction, and are often black-box, making it difficult to troubleshoot erroneous outcomes. This work…
Ying Zhao, Charles C. Zhou
SARS-Cov-2, the deadly and novel virus, which has caused a worldwide pandemic and drastic loss of human lives and economic activities. An open data set called the COVID-19 Open Research Dataset or CORD-19 contains large set full text scientific literature on SARS-CoV-2. The Next Strain consists of a database of…
Rachel Lyne, Adrián Bazaga, Daniela Butano, Sergio Contrino + 10 more
HumanMine (www.humanmine.org) is an integrated database of human genomics and proteomics data that provides a powerful interface to support sophisticated exploration and analysis of data compiled from experimental, computational and curated data sources. Built using the InterMine data integration platform, HumanMine…
Authors not listed
Iron, the most abundant element on Earth by mass (34.6%), primarily exists as iron minerals due to its inherent reactivity. The study of iron mineral phase transformations under changing environmental conditions remains an important research focus due to its geological, environmental, and industrial significance. Yet…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Rebecca Brunk, Kriti Shukla, Bryant Hutson, Yue Wang + 7 more
Genomic sequencing and other big biological data is unquestionably of paramount value, however the success in recruiting highly skilled individuals with diverse backgrounds has been limited. A main reason for this deficiency could be due to the lack of educational resources and early exposure to the field. With the…