26 papers · ranked by Valyu relevance
Tlamelo Emmanuel, Thabiso Maupong, Dimane Mpoeleng, Thabo Semong + 2 more
'Banyatsang Mphago' 'Oteng Tabona'] Machine learning has been the corner stone in analysing and extracting information from data and often a problem of missing values is encountered. Missing values occur because of various factors like missing completely at random, missing at random or missing not at random. All these…
Peter C. Austin, Ian R. White, Douglas S. Lee, Stef van Buuren
Missing data is a common occurrence in clinical research. Missing data occurs when the value of the variables of interest are not measured or recorded for all subjects in the sample. Common approaches to addressing the presence of missing data include complete-case analyses, where subjects with missing data are…
Janus Christian Jakobsen, Christian Gluud, Jørn Wetterslev, Per Winkel
'Per Winkel'] Background Missing data may seriously compromise inferences from randomised clinical trials, especially if missing data are not handled appropriately. The potential bias due to missing data depends on the mechanism causing the data to be missing, and the analytical methods applied to amend the…
Dawei Liu, Hanne Oberman, Johanna Muñoz, Jeroen Hoogland + 1 more
'Thomas P. A. Debray'] This is a preprint of the following chapter: Liu D, Oberman HI, Muñoz J, Hoogland J, Debray TPA, "Quality control, data cleaning, imputation", published in "Clinical applications of artificial intelligence in real-world data", edited by [editor of the book], [year of publication], [publisher (as…
Youran Zhou, Mohamed Reda Bouadjenek, Sunil Aryal
Missing data is a pervasive challenge spanning diverse data types, including tabular, sensor data, time-series, images and so on. Its origins are multifaceted, resulting in various missing mechanisms. Prior research in this field has predominantly revolved around the assumption of the Missing Completely At Random…
Brett K. Beaulieu-Jones, Daniel R. Lavage, John W. Snyder, Jason H. Moore + 2 more
Missing data is a challenge for all studies; however, this is especially true for electronic health record (EHR) based analyses. Failure to appropriately consider missing data can lead to biased results. Here, we provide detailed procedures for when and how to conduct imputation of EHR data. We demonstrate how the…
Youran Zhou, Sunil Aryal, Mohamed Reda Bouadjenek
Missing data poses a significant challenge in data science, affecting decision-making processes and outcomes. Understanding what missing data is, how it occurs, and why it is crucial to handle it appropriately is paramount when working with real-world data, especially in tabular data, one of the most commonly used data…
Anny K. G. Rodrigues, Raydonal Ospina, Marcelo R. P. Ferreira, Afnizanfaizal Abdullah
'Afnizanfaizal Abdullah'] Many machine learning procedures, including clustering analysis are often affected by missing values. This work aims to propose and evaluate a Kernel Fuzzy C-means clustering algorithm considering the kernelization of the metric with local adaptive distances (VKFCM-K-LP) under three types of…
Daniel W.A. Noble, Shinichi Nakagawa
Ecological and evolutionary research questions are increasingly requiring the integration of research fields along with larger datasets to address fundamental local and global scale problems. Unfortunately, these agendas are often in conflict with limited funding and a need to balance animal welfare concerns. Planned…
Mikko Särkkä, Sami Myöhänen, Kaloyan Marinov, Inka Saarinen + 3 more
Modern clinical genetic tests utilize next-generation sequencing (NGS) approaches to comprehensively analyze genetic variants from patients. Out of these millions of variants, clinically relevant variants that match the patient’s phenotype need to be identified accurately within a rapid timeframe that facilitates…
Sara Johansson Fernstad, Sarah Alsufyani, Silvia Del Din, Alison Yarnall + 1 more
'Alison Yarnall' 'Lynn Rochester'] This paper contributes a set of quality metrics for identification and visual analysis of structured missingness in high-dimensional data. Missing values in data are a frequent challenge in most data generating domains and may cause a range of analysis issues. Structural missingness…
Ahmad R. Alsaber, Jiazhu Pan, Adeeba Al-Hurban
In environmental research, missing data are often a challenge for statistical modeling. This paper addressed some advanced techniques to deal with missing values in a data set measuring air quality using a multiple imputation (MI) approach. MCAR, MAR, and NMAR missing data techniques are applied to the data set. Five…
Barbora Rehák Bučková, Charlotte Fraza, Cecilie Koldbæk Lemvigh, Camilla Bärthel Flaaten + 11 more
Missing data remain a ubiquitous and critical challenge in large-scale clinical studies. Despite advances in imputation, most existing methods fail to address structured missingness, where data are missing according a deterministic pattern and which arise due to systematic patterns introduced by experimental design…
Neslihan Süzen, Evgeny M. Mirkes, Damian Roland, Jeremy Levesley + 2 more
'Alexander N. Gorban' 'Tim Coats'] Abstract— Electronic patient records (EPRs) produce a wealth of data but contain significant missing information. Understanding and handling this missing data is an important part of clinical data analysis and if left unaddressed could result in bias in analysis and distortion in…
Katya L Masconi, Tandi E Matsha, Justin B Echouffo-Tcheugui, Rajiv T Erasmus + 1 more
Missing values are common in health research and omitting participants with missing data often leads to loss of statistical power, biased estimates and, consequently, inaccurate inferences. We critically reviewed the challenges posed by missing data in medical research and approaches to address them. To achieve this…
Sara Johansson Fernstad, Jimmy Johansson
—This paper contributes a novel visualization method, Missingness Glyph, for analysis and exploration of missing values in data. Missing values are a common challenge in most data generating domains and may cause a range of analysis issues. Missingness in data may indicate potential problems in data collection and…
Gift Khangamwa, Terence L. van Zyl, C. J. van Alten
Missing data is a common concern in health datasets, and its impact on good decision-making processes is well documented. Our study's contribution is a methodology for tackling missing data problems using a combination of synthetic dataset generation, missing data imputation and deep learning methods to resolve missing…
Martin Seifrid, Stanley Lo, Dylan Choi, Gary Tom + 12 more
Martin Seifrid 1 , Stanley Lo 2 , Dylan G. Choi 3 , Gary Tom 2 , My Linh Le 3 , Kunyu Li 3 , Rahul Sankar 3 , Hoai-Thanh Vuong 3 , Hiba Wakidi 3 , Ahra Yi 3 , Ziyue Zhu 3 , Nora Schopp 3 , Aaron Peng 3 , Benjamin Luginbuhl 3 , Thuc-Quyen Nguyen 3 , Alán Aspuru-Guzik 2
María P. Fernández-García, Guillermo Vallejo-Seco, Pablo Livácic-Rojas, Ellian Tuero-Herrero
'Pablo Livácic-Rojas' 'Ellian Tuero-Herrero'] It is practically impossible to avoid losing data in the course of an investigation, and it has been proven that the consequences can reach such magnitude that they could even invalidate the results of the study. This paper describes some of the most likely causes of…
Diba Behnoudfar, Cory Simon, Joshua Schrier
Aqueous, two-phase systems (ATPSs) may form upon mixing two solutions of independently water-soluble compounds. Many separation, purification, and extraction processes rely on ATPSs. Predicting the miscibility of solutions can accelerate and reduce the cost of the discovery of new ATPSs for these applications. Whereas…
Xing Chen, Na Zhang, Xiaohui Yang, Chunyan Wang + 5 more
In daily life, two common algorithms are used for collecting medical disease data: data integration of medical institutions and questionnaires. However, these statistical methods require collecting data from the entire research area, which consumes a significant amount of manpower and material resources. Additionally…
Bartłomiej Fliszkiewicz, Marcin Sajdak
The aim of the following research is to assess the applicability of calculated quantum properties of molecular fragments as molecular descriptors in machine learning classification task. The research is based on bio-concentration and QM9-extended databases. A number of compounds with results from quantum-chemical…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Authors not listed
Solute carrier (SLC) transporters constitute the largest family of membrane transport proteins in humans. They facilitate the movement of ions, neurotransmitters, nutrients, and drugs. Given their critical role in regulating cellular physiology, they are important therapeutic targets for neurological and psychological…
Sangjoon Lee, Clio Chen, Griheydi Garcia, Anton Oliynyk
Materials informatics uses data-driven approaches for the study and discovery of materials. Features or descriptors are the crucial components in generating reliable and accurate machine-learning models. While general data can be acquired through public and commercial sources, features must be tailored for a specific…
Chonghuan Zhang, Adarsh Arun, Alexei Lapkin
Computer Aided Synthesis Planning (CASP) development of reaction routes requires understanding of complete reaction structures. However, most reactions in the current databases are missing reaction co-participants. Although reaction prediction and atom mapping tools can predict major reaction participants and trace…