16 papers · ranked by Valyu relevance
Tlamelo Emmanuel, Thabiso Maupong, Dimane Mpoeleng, Thabo Semong + 2 more
'Banyatsang Mphago' 'Oteng Tabona'] Machine learning has been the corner stone in analysing and extracting information from data and often a problem of missing values is encountered. Missing values occur because of various factors like missing completely at random, missing at random or missing not at random. All these…
Anny K. G. Rodrigues, Raydonal Ospina, Marcelo R. P. Ferreira, Afnizanfaizal Abdullah
'Afnizanfaizal Abdullah'] Many machine learning procedures, including clustering analysis are often affected by missing values. This work aims to propose and evaluate a Kernel Fuzzy C-means clustering algorithm considering the kernelization of the metric with local adaptive distances (VKFCM-K-LP) under three types of…
Laila Mousafi Alasal, Emma U Hammarlund, Kenneth J Pienta, Lars Rönnstrand + 2 more
In the context of data science, the quality and completeness of data are critical factors that dictate the accuracy and reliability of analysis (). Missing values in datasets pose a significant challenge, as they can obscure underlying patterns and potentially lead to biased or incorrect conclusions. Missing values can…
Monica Casella, Nicola Milano, Pasquale Dolce, Davide Marocco
Introduction Missing data in psychometric research presents a substantial challenge, impacting the reliability and validity of study outcomes. Various factors contribute to this issue, including participant non-response, dropout, or technical errors during data collection. Traditional methods like mean imputation or…
Muhammad Salar Khan, Larissa-Margareta Batrancea
Within the national innovation system literature, the low- and middle-income countries (LMICs) eligible for the World Bank’s International Development Association (IDA) support, are rarely part of empirical discourses on growth, development, and innovation. One major issue hindering empirical analyses in LMICs is the…
Adrienne Kline, Yuan Luo
Most datasets suffer from partial or complete missing values, which has downstream limitations on the available models on which to test the data and on any statistical inferences that can be made from the data. Several imputation techniques have been designed to replace missing data with stand in values. The various…
Molly Ehrig, Garrett S Bullock, Xiaoyan Iris Leng, Nicholas M Pajewski + 2 more
'Nicholas M Pajewski' 'Jaime Lynn Speiser' 'Christian Lovis'] Title: Abstract Background Missing data in electronic health records are highly prevalent and result in analytical concerns such as heterogeneous sources of bias and loss of statistical power. One simple analytic method for addressing missing or unknown…
Huiting Ou, Anuradha Surendra, Graeme S V McDowell, Emily Hashimoto-Roth + 4 more
Missing data are a major problem for multivariate, machine learning (ML) and network analyses. For example, in large lipidomic or metabolic datasets, measurements for some analytes may not be available in every sample due to routine technical variability, low abundance, ion suppression from co-eluting analytes…
Asmaa Ahmad, Eric J. Rose, Michael S. Roy, Edward Valachovic + 1 more
Missing data in periodic time series can bias inference when temporal dependence is not preserved during imputation. We propose a framework that integrates the Variable Bandpass Periodic Block Bootstrap (VBPBB) with multiple imputation using Amelia II by incorporating statistically significant periodic components as…
Hu Pan, Zhiwei Ye, Qiyi He, Chunyan Yan + 7 more
'Jun Su' 'Ruihan Li' 'Grigore Stamatescu' 'Anatoliy Sachenko' 'Anna Romańska-Zapała'] Data are a strategic resource for industrial production, and an efficient data-mining process will increase productivity. However, there exist many missing values in data collected in real life due to various problems. Because the…
Karl Schweizer, Andreas Gold, Dorothea Krampen
In modeling missing data, the missing data latent variable of the confirmatory factor model accounts for systematic variation associated with missing data so that replacement of what is missing is not required. This study aimed at extending the modeling missing data approach to tetrachoric correlations as input and at…
Kritanat Chungnoy, Tanatorn Tanantong, Pokpong Songmuang, Bilal Alatas
'Bilal Alatas'] Existing missing value imputation methods focused on imputing the data regarding actual values towards a completion of datasets as an input for machine learning tasks. This work proposes an imputation of missing values towards improvement of accuracy performance for classification. The proposed method…
Aya El Mir, Eric Bezerra de Sousa, Ignacio Mesina-Estarrón, Leo Anthony Celi + 8 more
Missing, inaccurate, or poorly documented data in healthcare is often treated as a technical problem to be statistically resolved via imputation, deletion, or modeling assumptions about randomness. However, such inaccuracies relate to far more complex socioeconomic and geopolitical issues, rather than “errors of data…
Jay Darji, Nupur Biswas, Vijay Padul, Jaya Gill + 2 more
'Shashaanka Ashili'] Time series data are recorded in various sectors, resulting in a large amount of data. However, the continuity of these data is often interrupted, resulting in periods of missing data. Several algorithms are used to impute the missing data, and the performance of these methods is widely varied.…
Jessica Ryan-Despraz, Amanda Wissler
Missing data is a prevalent problem in bioarchaeological research and imputation could provide a promising solution. This work simulated missingness on a control dataset (481 samples × 41 variables) in order to explore imputation methods for mixed data (qualitative and quantitative data). The tested methods included…
Ala Alrawajfi, Mohd Tahir Ismail, Sadam Al Wadi, Saleh Atiewi + 2 more
'Ahmad Awajan' 'Stefano Cirillo'] Data imputation strategies are necessary to address the prevalent difficulty of missing values in data observation and recording operations. This work utilizes diverse imputation methods to forecast and complete absent values inside a financial time-series dataset, specifically the…