25 papers · ranked by Valyu relevance
Krishnan Bhaskaran, Liam Smeeth
The terminology describing missingness mechanisms is confusing. In particular the meaning of ‘missing at random’ is often misunderstood, leading researchers faced with missing data problems away from multiple imputation, a method with considerable advantages. The purpose of this article is to clarify how ‘missing at…
Daniel W.A. Noble, Shinichi Nakagawa
Ecological and evolutionary research questions are increasingly requiring the integration of research fields along with larger datasets to address fundamental local and global scale problems. Unfortunately, these agendas are often in conflict with limited funding and a need to balance animal welfare concerns. Planned…
Zhang, Yunshu, Park, Chan + 10 more
Instrumental variable (IV) methods offer a valuable approach to account for outcome data missing not-at-random. A valid missing data instrument is a measured factor which (i) predicts the nonresponse process and (ii) is independent of the outcome in the underlying population. For point identification, all existing IV…
Daniel W. A. Noble, Shinichi Nakagawa
Ecological and evolutionary research questions are increasingly requiring the integration of research fields along with larger data sets to address fundamental local- and global-scale problems. Unfortunately, these agendas are often in conflict with limited funding and a need to balance animal welfare concerns. Planned…
Rui Duan, Chengcai Liang, Pamela Shaw, Cheng Yong Tang + 1 more
Practical problems with missing data are common, and statistical methods have been developed concerning the validity and/or efficiency of statistical procedures. On a central focus, there have been longstanding interests on the mechanism governing data missingness, and correctly deciding the appropriate mechanism is…
Mithilesh Prakash, Jussi Tohka
We introduce a new subtype of ‘Missing Not at Random’ (MNAR) data, where the missingness is correlated with the labels (y) to be predicted, termed (y)-dependent MNAR. We demonstrate that this subtype can significantly bias the estimation of performance metrics in typical machine learning tasks. Unbiased error…
Sara Johansson Fernstad, Jimmy Johansson
—This paper contributes a novel visualization method, Missingness Glyph, for analysis and exploration of missing values in data. Missing values are a common challenge in most data generating domains and may cause a range of analysis issues. Missingness in data may indicate potential problems in data collection and…
Mikko Särkkä, Sami Myöhänen, Kaloyan Marinov, Inka Saarinen + 3 more
Modern clinical genetic tests utilize next-generation sequencing (NGS) approaches to comprehensively analyze genetic variants from patients. Out of these millions of variants, clinically relevant variants that match the patient’s phenotype need to be identified accurately within a rapid timeframe that facilitates…
Runmin Wei, Jingye Wang, Erik Jia, Tianlu Chen + 2 more
Left-censored missing values commonly exist in targeted metabolomics datasets and can be considered as missing not at random (MNAR). Improper data processing procedures for missing values will cause adverse impacts on subsequent statistical analyses. However, few imputation methods have been developed and applied to…
Colby J. Vorland, Andrew W. Brown, John A. Dawson, Stephanie L. Dickinson + 11 more
'Stephanie L. Dickinson' 'Lilian Golzarri-Arroyo' 'Bridget A. Hannon' 'Moonseong Heo' 'Steven B. Heymsfield' 'Wasantha P. Jayawardene' 'Chanaka N. Kahathuduwa' 'Scott W. Keith' 'J. Michael Oakes' 'Carmen D. Tekwe' 'Lehana Thabane' 'David B. Allison'] Randomization is an important tool used to establish causal…
Tim P. Morris, A. Sarah Walker, Elizabeth J. Williamson, Ian R. White
'Ian R. White'] Background It has long been advised to account for baseline covariates in the analysis of confirmatory randomised trials, with the main statistical justifications being that this increases power and, when a randomisation scheme balanced covariates, permits a valid estimate of experimental error. There…
John C. Galati
Missing at Random (MAR) is a central concept in incomplete data methods, and often it is stated as P(R | Yobs, Ymis) = P(R | Yobs). This notation has been used in the literature for more than three decades and has become the de facto standard. In some cases, the notation has been misinterpreted to be a statement about…
Mark A. Eckert, Kenneth I. Vaden, Mulugeta Gebregziabher
Children with reading disability exhibit varied deficits in reading and cognitive abilities that contribute to their reading comprehension problems. Some children exhibit primary deficits in phonological processing, while others can exhibit deficits in oral language and executive functions that affect comprehension.…
D. M. Farewell, R. M. Daniel, S. R. Seaman
Title: Summary We offer a natural and extensible measure-theoretic treatment of missingness at random. Within the standard missing-data framework, we give a novel characterization of the observed data as a stopping-set sigma algebra. We demonstrate that the usual missingness-at-random conditions are equivalent to…
Dan Jackson, Ian R White, Morven Leese
When a randomized controlled trial has missing outcome data, any analysis is based on untestable assumptions, e.g. that the data are missing at random, or less commonly on other assumptions about the missing data mechanism. Given such assumptions, there is an extensive literature on suitable methods of analysis.…
Daniel Farewell, Rhian Daniel, Shaun R. Seaman
We offer a natural and extensible measure-theoretic treatment of missingness at random. Within the standard missing data framework, we give a novel characterisation of the observed data as a stopping-set sigma algebra. We demonstrate that the usual missingness at random conditions are equivalent to requiring particular…
Clémence Leyrat, James R. Carpenter, Sébastien Bailly, Elizabeth J Willamson
'Elizabeth J Willamson'] Background: Marginal structural models (MSMs) are commonly used to estimate causal intervention effects in longitudinal non-randomised studies. A common issue when analysing data from observational studies is the presence of incomplete confounder data, which might lead to bias in the…
Xijin Chen, Kim May Lee, Sofia S. Villar, David S. Robertson + 1 more
'Darrell A. Worthy'] When comparing the performance of multi-armed bandit algorithms, the potential impact of missing data is often overlooked. In practice, it also affects their implementation where the simplest approach to overcome this is to continue to sample according to the original bandit algorithm, ignoring…
Xijin Chen, Kim May Lee, Sofía S. Villar, D. S. Robertson
When comparing the performance of multi-armed bandit algorithms, the potential impact of missing data is often overlooked. In practice, it also affects their implementation where the simplest approach to overcome this is to continue to sample according to the original bandit algorithm, ignoring missing outcomes. We…
Marius Hofert, J R Jackson, Niels Hagenbuch
An approach to amputation, the process of introducing missing values to a complete dataset, is presented. It allows to construct missingness indicators in a flexible and principled way via copulas and Bernoulli margins and to incorporate dependence in missingness patterns. Besides more classical missingness models such…
Diba Behnoudfar, Cory Simon, Joshua Schrier
Aqueous, two-phase systems (ATPSs) may form upon mixing two solutions of independently water-soluble compounds. Many separation, purification, and extraction processes rely on ATPSs. Predicting the miscibility of solutions can accelerate and reduce the cost of the discovery of new ATPSs for these applications. Whereas…
Sangjoon Lee, Clio Chen, Griheydi Garcia, Anton Oliynyk
Materials informatics uses data-driven approaches for the study and discovery of materials. Features or descriptors are the crucial components in generating reliable and accurate machine-learning models. While general data can be acquired through public and commercial sources, features must be tailored for a specific…
David D. Hofmann, Gabriele Cozzi, John Fieberg
Integrated step-selection analyses (iSSAs) are versatile and powerful frameworks for studying habitat and movement preferences of tracked animals. iSSAs utilize integrated step-selection functions (iSSFs) to model movements in discrete time, and thus, require animal location data that are regularly spaced in time.…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Arni Sturluson, Ali Raza, Grant D. McConachie, Daniel Siderius + 2 more
Nanoporous materials (NPMs) selectively adsorb and concentrate gases into their pores, and thus could be used to store, capture, and sense many different gases. Modularly synthesized classes of NPMs, such as covalent organic frameworks (COFs), offer a large number of candidate structures for each adsorption task. A…