26 papers · ranked by Valyu relevance
Mahmudur Rahman Hera, David Koslicki, Conrado Martínez
With the surge in sequencing data generated from an ever-expanding range of biological studies, designing scalable computational techniques has become essential. One effective strategy to enable large-scale computation is to split long DNA or protein sequences into k-mers, and summarize large k-mer sets into compact…
Alberto Arletti, Maria Letizia Tanturri, Omar Paccagnella
Online data has the potential to transform how researchers and companies produce election forecasts. Social media surveys, online panels and even comments scraped from the internet can offer valuable insights into political preferences. However, such data is often affected by significant selection bias, as online…
Pavlos Msaouel, Juhee Lee, Peter F. Thall, Alan Hutson
Simple Summary We provide an extensive review of the fundamental principles of statistical science that are needed to accurately interpret randomized controlled trials (RCTs). We use these principles to explain how RCTs are motivated by the powerful but strange idea that flipping a coin to choose each patient’s…
Jae Kwang Kim
This textbook on survey sampling has its origins in a set of lecture notes prepared for a course on survey sampling at Iowa State University. Over the years, these notes have been refined and expanded into the comprehensive volume you now hold. It is designed to serve both as an introductory text for students and as a…
Matias López
The literature frequently recommends purposive sampling of elites based on the assumptions that random sampling negatively affects the response rate and that it induces bias. I test these assumptions drawing on metadata from 282 samples of political, economic, and social elites, and on microdata from 2,658 elites.…
Célia Landmann Szwarcwald
This article aimed to present an overview of national health surveys, sampling techniques, and components of statistical analysis of data collected using complex sampling designs. Briefly, surveys aimed at assessing the nutritional status of Brazilians and maternal and child health care were described. Surveys aimed at…
Antonio Maratea, Rita Perna
Adequate sampling space coverage is the keystone to effectively train trustworthy Machine Learning models. Unfortunately, real data do carry several inherent risks due to the many potential biases they exhibit when gathered without a proper random sampling over the reference population, and most of the times this is…
Ryan Covey, Lucca Buonamano
The use of big data in official statistics and the applied sciences is accelerating, but statistics computed using only big data often suffer from substantial selection bias. This leads to inaccurate estimation and invalid statistical inference. We rectify the issue for a broad class of linear and nonlinear statistics…
Hiroyuki Kuwahara, Xin Gao, Can Alkan
Reservoir sampling sequentially reads a streaming file and randomly samples k elements with the equal probability from the population of unknown size n. Waterman developed a direct approach of reservoir sampling called Algorithm R in the 1970s (). Algorithm R first places the first k elements in the selection set…
Sebahat Gok, Robert L. Goldstone
Interactive computer simulations are commonly used as pedagogical tools to support students’ statistical reasoning. This paper examines whether and how these simulations enable their intended effects. We begin by contrasting two theoretical frameworks-dual processes and grounded cognition-in the context of people’s…
Muhammad Azeem, Sundus Hussain, Musarrat Ijaz, Najma Salahuddin + 1 more
'Abdul Salam'] In survey sampling, systematic sampling design has attracted survey researchers in recent years due to its simplicity of use. We introduce a modified variant of systematic sampling scheme which improves the efficiency of a recently developed diagonal systematic sampling method. The suggested modification…
Prabhleen Kaur, Simone Ciuti, Michael Salter-Townshend, Damien Farine
Producing accurate and reliable inference from animal social network analysis depends on the sampling strategy during data collection. An increasing number of studies now use large-scale deployment of GPS tags to collect data on social behaviour. However, these can rarely capture whole populations or sample at very…
Kanwal Iqbal, Syed Muhammad Muslim Raza, Tahir Mahmood, Muhammad Riaz + 1 more
'Muhammad Riaz' 'Mohamed R. Abonazel'] Advancements in sensor technology have brought a revolution in data generation. Therefore, the study variable and several linearly related auxiliary variables are recorded due to cost-effectiveness and ease of recording. These auxiliary variables are commonly observed as…
Joseph Rich, Lior Pachter
Summary: fastQpick is a command-line tool and Python library for sampling FASTQ reads with replacement. Sampling with replacement turns a single FASTQ file into an arbitrary number of bootstrap replicates, which enables uncertainty quantification and statistical analysis at the level of raw reads. This process answers…
Daniel J. McGlinn, Shane A. Blowes, Maria Dornelas, Thore Engel + 6 more
There is considerable interest in understanding patterns of β-diversity that measure the amount of change in species composition through space or time. Most hypotheses for β-diversity evoke nonrandom processes that generate spatial and temporal within species aggregation; however, β-diversity can also be driven by…
Authors not listed
Developing generalizable machine learning models with minimal data remains a central challenge in materials informatics. Effective models can significantly reduce costly computational simulations and time-intensive experimentation by providing reliable predictions of material properties. In this work, we investigate…
Jaromı́r Antoch, Francesco Molà, Ondřej Vozár
In this paper, a new randomized response technique aimed at protecting respondents' privacy is proposed. It is designed for estimating the population total, or the population mean, of a quantitative characteristic. It provides a high degree of protection to the interviewed individuals, hence it may be favorably…
Piero Demetrio Falorsi, Stefano Falorsi, Vincenzo Nardelli, Paolo Righi
'Paolo Righi'] Abstract. The paper delineates a proper statistical setting for defining the sampling design for a small area estimation problem. This problem is often treated only via indirect estimation using the values of the variable of interest also from different areas or different times, thus increasing the…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…
Yizhuo Wang, Bing Z. Carter, Ziyi Li, Xuelin Huang
A key component for precision medicine is a good prediction algorithm for patients’ response to treatments. We aim to implement machine learning (ML) algorithms into the response-adaptive randomization (RAR) design and improve the treatment outcomes. We incorporated nine ML algorithms to model the relationship of…
Yuji Kaiya, Ryo Tamura, Koji Tsuda
Kinetic models are widely used in simulating the relationship between the input space and the outcome space of a chemical process. Ignoring the computational cost, complete profiling, i.e., performing simulation at all grid points in the input space, would be the best way to understand the model, because it provides us…
David A. Rasmussen, Madeline G. Bursell, Frank Burkhart
Inferences from population genomic data provide valuable insights into the demographic history of a population. Likewise, in genomic epidemiology, pathogen genomic data provide key insights into epidemic dynamics and potential sources of transmission. Yet predicting what information will be gained from genomic data…
PETER CARDEW, Keith Gregory, Adam Lechmere
In order to meet the EU requirement of compliance with the 10 µg/l lead standard utility companies in England and Wales implemented a large-scale plumbosolvency treatment programme based around the addition of orthophosphate. This was largely delivered by the end of 2003. This solution has resulted in a major…
Sterling Baird, Jason R. Hall, Taylor D. Sparks
Would you rather search for a line inside a cube or a point inside a square? Physics-based simulations and wet-lab experiments often have symmetries (degeneracies) that allow reducing problem dimensionality or search space, but constraining these degeneracies is often unsupported or difficult to implement in many…
Authors not listed
Metastable states and the conformational transitions in between them are key to understanding dynamical behaviour and function of large-scale molecular systems. By combining basic dimensionality reduction techniques with a state-of-the art approximation of the Koopman operator associated to molecular dynamics…
Megan N. Taylor, Nic M. Vega
Heterogeneity is ubiquitous across individuals in biological data, and sample batching, a form of biological averaging, inevitably loses information about this heterogeneity. The consequences for inference from biologically averaged data are frequently opaque, particularly when the underlying populations are…