14 papers · ranked by Valyu relevance
Tlamelo Emmanuel, Thabiso Maupong, Dimane Mpoeleng, Thabo Semong + 2 more
'Banyatsang Mphago' 'Oteng Tabona'] Machine learning has been the corner stone in analysing and extracting information from data and often a problem of missing values is encountered. Missing values occur because of various factors like missing completely at random, missing at random or missing not at random. All these…
Anny K. G. Rodrigues, Raydonal Ospina, Marcelo R. P. Ferreira, Afnizanfaizal Abdullah
'Afnizanfaizal Abdullah'] Many machine learning procedures, including clustering analysis are often affected by missing values. This work aims to propose and evaluate a Kernel Fuzzy C-means clustering algorithm considering the kernelization of the metric with local adaptive distances (VKFCM-K-LP) under three types of…
Monica Casella, Nicola Milano, Pasquale Dolce, Davide Marocco
Introduction Missing data in psychometric research presents a substantial challenge, impacting the reliability and validity of study outcomes. Various factors contribute to this issue, including participant non-response, dropout, or technical errors during data collection. Traditional methods like mean imputation or…
Muhammad Salar Khan, Larissa-Margareta Batrancea
Within the national innovation system literature, the low- and middle-income countries (LMICs) eligible for the World Bank’s International Development Association (IDA) support, are rarely part of empirical discourses on growth, development, and innovation. One major issue hindering empirical analyses in LMICs is the…
Adrienne Kline, Yuan Luo
Most datasets suffer from partial or complete missing values, which has downstream limitations on the available models on which to test the data and on any statistical inferences that can be made from the data. Several imputation techniques have been designed to replace missing data with stand in values. The various…
Karl Schweizer, Andreas Gold, Dorothea Krampen
In modeling missing data, the missing data latent variable of the confirmatory factor model accounts for systematic variation associated with missing data so that replacement of what is missing is not required. This study aimed at extending the modeling missing data approach to tetrachoric correlations as input and at…
Martin Seifrid, Stanley Lo, Dylan Choi, Gary Tom + 12 more
Martin Seifrid 1 , Stanley Lo 2 , Dylan G. Choi 3 , Gary Tom 2 , My Linh Le 3 , Kunyu Li 3 , Rahul Sankar 3 , Hoai-Thanh Vuong 3 , Hiba Wakidi 3 , Ahra Yi 3 , Ziyue Zhu 3 , Nora Schopp 3 , Aaron Peng 3 , Benjamin Luginbuhl 3 , Thuc-Quyen Nguyen 3 , Alán Aspuru-Guzik 2
Jessica Ryan-Despraz, Amanda Wissler
Missing data is a prevalent problem in bioarchaeological research and imputation could provide a promising solution. This work simulated missingness on a control dataset (481 samples × 41 variables) in order to explore imputation methods for mixed data (qualitative and quantitative data). The tested methods included…
Diba Behnoudfar, Cory Simon, Joshua Schrier
Aqueous, two-phase systems (ATPSs) may form upon mixing two solutions of independently water-soluble compounds. Many separation, purification, and extraction processes rely on ATPSs. Predicting the miscibility of solutions can accelerate and reduce the cost of the discovery of new ATPSs for these applications. Whereas…
Bartłomiej Fliszkiewicz, Marcin Sajdak
The aim of the following research is to assess the applicability of calculated quantum properties of molecular fragments as molecular descriptors in machine learning classification task. The research is based on bio-concentration and QM9-extended databases. A number of compounds with results from quantum-chemical…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Authors not listed
Solute carrier (SLC) transporters constitute the largest family of membrane transport proteins in humans. They facilitate the movement of ions, neurotransmitters, nutrients, and drugs. Given their critical role in regulating cellular physiology, they are important therapeutic targets for neurological and psychological…
Sangjoon Lee, Clio Chen, Griheydi Garcia, Anton Oliynyk
Materials informatics uses data-driven approaches for the study and discovery of materials. Features or descriptors are the crucial components in generating reliable and accurate machine-learning models. While general data can be acquired through public and commercial sources, features must be tailored for a specific…
Chonghuan Zhang, Adarsh Arun, Alexei Lapkin
Computer Aided Synthesis Planning (CASP) development of reaction routes requires understanding of complete reaction structures. However, most reactions in the current databases are missing reaction co-participants. Although reaction prediction and atom mapping tools can predict major reaction participants and trace…