14 papers · ranked by Valyu relevance
Tlamelo Emmanuel, Thabiso Maupong, Dimane Mpoeleng, Thabo Semong + 2 more
'Banyatsang Mphago' 'Oteng Tabona'] Machine learning has been the corner stone in analysing and extracting information from data and often a problem of missing values is encountered. Missing values occur because of various factors like missing completely at random, missing at random or missing not at random. All these…
Anny K. G. Rodrigues, Raydonal Ospina, Marcelo R. P. Ferreira, Afnizanfaizal Abdullah
'Afnizanfaizal Abdullah'] Many machine learning procedures, including clustering analysis are often affected by missing values. This work aims to propose and evaluate a Kernel Fuzzy C-means clustering algorithm considering the kernelization of the metric with local adaptive distances (VKFCM-K-LP) under three types of…
Youran Zhou, Mohamed Reda Bouadjenek, Sunil Aryal
Missing data is a pervasive challenge spanning diverse data types, including tabular, sensor data, time-series, images and so on. Its origins are multifaceted, resulting in various missing mechanisms. Prior research in this field has predominantly revolved around the assumption of the Missing Completely At Random…
Monica Casella, Nicola Milano, Pasquale Dolce, Davide Marocco
Introduction Missing data in psychometric research presents a substantial challenge, impacting the reliability and validity of study outcomes. Various factors contribute to this issue, including participant non-response, dropout, or technical errors during data collection. Traditional methods like mean imputation or…
Jahan C. Penny‐Dimri, Christoph Bergmeir, Julian A. Smith
Most practical data science problems encounter missing data. A wide variety of solutions exist, each with strengths and weaknesses that depend upon the missingness-generating process. Here we develop a theoretical framework for training and inference using only observed variables enabling modeling of incomplete…
Youran Zhou, Sunil Aryal, Mohamed Reda Bouadjenek
Missing data poses a significant challenge in data science, affecting decision-making processes and outcomes. Understanding what missing data is, how it occurs, and why it is crucial to handle it appropriately is paramount when working with real-world data, especially in tabular data, one of the most commonly used data…
Muhammad Salar Khan, Larissa-Margareta Batrancea
Within the national innovation system literature, the low- and middle-income countries (LMICs) eligible for the World Bank’s International Development Association (IDA) support, are rarely part of empirical discourses on growth, development, and innovation. One major issue hindering empirical analyses in LMICs is the…
Sara Johansson Fernstad, Sarah Alsufyani, Silvia Del Din, Alison Yarnall + 1 more
'Alison Yarnall' 'Lynn Rochester'] This paper contributes a set of quality metrics for identification and visual analysis of structured missingness in high-dimensional data. Missing values in data are a frequent challenge in most data generating domains and may cause a range of analysis issues. Structural missingness…
Adrienne Kline, Yuan Luo
Most datasets suffer from partial or complete missing values, which has downstream limitations on the available models on which to test the data and on any statistical inferences that can be made from the data. Several imputation techniques have been designed to replace missing data with stand in values. The various…
Seongmin Kim, Jaewon Oh, Hye-Young Ko, Jeongmin Park + 1 more
Missing data is a common issue in various fields such as medicine, social sciences, and natural sciences, and it poses significant challenges for accurate statistical analysis. Although numerous imputation methods have been proposed to address this issue, many of them fail to adequately capture the complex dependency…
Muhammad Ishaq, Laila iftikhar, Majid Khan, Asfandyar Khan + 1 more
'Arshad Khan'] This study explored the use of machine learning algorithms for predicting and imputing missing values in categorical datasets. We focused on ensemble models that use the error correction output codes (ECOC) framework, including SVM-based and KNN-based ensemble models, as well as an ensemble classifier…
Karl Schweizer, Andreas Gold, Dorothea Krampen
In modeling missing data, the missing data latent variable of the confirmatory factor model accounts for systematic variation associated with missing data so that replacement of what is missing is not required. This study aimed at extending the modeling missing data approach to tetrachoric correlations as input and at…
Hugo Morvan, Jonas Agholme, Bjorn Eliasson, Katarina Olofsson + 3 more
Missing data is a prevalent issue in many applications, including large medical registries such as the Swedish Healthcare Quality Registries, potentially leading to biased or inefficient analyses if not handled properly. Multiple Imputation by Chained Equations (MICE) is a popular and versatile method for handling…
Jessica Ryan-Despraz, Amanda Wissler
Missing data is a prevalent problem in bioarchaeological research and imputation could provide a promising solution. This work simulated missingness on a control dataset (481 samples × 41 variables) in order to explore imputation methods for mixed data (qualitative and quantitative data). The tested methods included…