22 papers · ranked by Valyu relevance
Mohamed Elfil, Ahmed Negida
Clinical research usually involves patients with a certain disease or a condition. The generalizability of clinical research findings is based on multiple factors related to the internal and external validity of the research methods. The main methodological issue that influences the generalizability of clinical…
Stephen Tyrer, Bob Heyman
Surveys of people's opinions are fraught with difficulties. It is easier to obtain information from those who respond to text messages or to emails than to attempt to obtain a representative sample. Samples of the population that are selected non-randomly in this way are termed convenience samples as they are easy to…
Mahmudur Rahman Hera, David Koslicki, Conrado Martínez
With the surge in sequencing data generated from an ever-expanding range of biological studies, designing scalable computational techniques has become essential. One effective strategy to enable large-scale computation is to split long DNA or protein sequences into k-mers, and summarize large k-mer sets into compact…
Alberto Arletti, Maria Letizia Tanturri, Omar Paccagnella
Online data has the potential to transform how researchers and companies produce election forecasts. Social media surveys, online panels and even comments scraped from the internet can offer valuable insights into political preferences. However, such data is often affected by significant selection bias, as online…
Jae Kwang Kim
This textbook on survey sampling has its origins in a set of lecture notes prepared for a course on survey sampling at Iowa State University. Over the years, these notes have been refined and expanded into the comprehensive volume you now hold. It is designed to serve both as an introductory text for students and as a…
Yves Tillé, Matthieu Wilhelm
The aim of this paper is twofold. First, three theoretical principles are formalized: randomization, overrepresentation and restriction. We develop these principles and give a rationale for their use in choosing the sampling design in a systematic way. In the model-assisted framework, knowledge of the population is…
Matias López
The literature frequently recommends purposive sampling of elites based on the assumptions that random sampling negatively affects the response rate and that it induces bias. I test these assumptions drawing on metadata from 282 samples of political, economic, and social elites, and on microdata from 2,658 elites.…
Célia Landmann Szwarcwald
This article aimed to present an overview of national health surveys, sampling techniques, and components of statistical analysis of data collected using complex sampling designs. Briefly, surveys aimed at assessing the nutritional status of Brazilians and maternal and child health care were described. Surveys aimed at…
Antonio Maratea, Rita Perna
Adequate sampling space coverage is the keystone to effectively train trustworthy Machine Learning models. Unfortunately, real data do carry several inherent risks due to the many potential biases they exhibit when gathered without a proper random sampling over the reference population, and most of the times this is…
Benyamin Ghojogh, Hadi Nekoei, Aydin Ghojogh, Fakhri Karray + 1 more
'Mark Crowley'] This paper is a tutorial and literature review on sampling algorithms. We have two main types of sampling in statistics. The first type is survey sampling which draws samples from a set or population. The second type is sampling from probability distribution where we have a probability density or mass…
Sanjar Adilov
Generative neural networks have shown promising results in de novo drug design. Recent studies suggest that one of the efficient ways to produce novel molecules matching target properties is to model SMILES sequences using deep learning in a way similar to language modeling in natural language processing. In this…
Muhammad Azeem, Sundus Hussain, Musarrat Ijaz, Najma Salahuddin + 1 more
'Abdul Salam'] In survey sampling, systematic sampling design has attracted survey researchers in recent years due to its simplicity of use. We introduce a modified variant of systematic sampling scheme which improves the efficiency of a recently developed diagonal systematic sampling method. The suggested modification…
Joseph Rich, Lior Pachter
Summary: fastQpick is a command-line tool and Python library for sampling FASTQ reads with replacement. Sampling with replacement turns a single FASTQ file into an arbitrary number of bootstrap replicates, which enables uncertainty quantification and statistical analysis at the level of raw reads. This process answers…
Andrew Hooyman, Matthew J. Huentelman, Sydney Y. Schaefer
Given the time- and resource-intense nature of human subjects research, we have developed a more intelligent approach to participant recruitment above and beyond random sampling that leverages pilot or preliminary results to reduce the overall number of participants needed for recruitment from an existing electronic…
Richard McGarvey, Paul Burch, Janet M. Matthews
Monitoring the density of natural populations is crucial for ecosystem management decision making and natural resource management. The most widely used method to measure the population density of animal and plant species in natural habitats is to count organisms in sample plots. Yet evaluation of survey performance by…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…
Simon R. White, Laura J. Bonnett
Title: Summary The statistical concept of sampling is often given little direct attention, typically reduced to the mantra “take a random sample”. This low resource and adaptable activity demonstrates sampling and explores issues that arise due to biased sampling.
PETER CARDEW, Keith Gregory, Adam Lechmere
In order to meet the EU requirement of compliance with the 10 µg/l lead standard utility companies in England and Wales implemented a large-scale plumbosolvency treatment programme based around the addition of orthophosphate. This was largely delivered by the end of 2003. This solution has resulted in a major…
Robert Arbon, Yanchen Zhu, Antonia S. J. S. Mey
Markov state models (MSM) are a popular statistical method for analyzing the conformational dynamics of proteins, including protein folding. With all statistical and machine learning (ML) models choices must be made about the modeling pipeline that cannot be directly learned from the data. These choices, or…
Authors not listed
Accurate and efficient calculation of alchemical free energies is a critical challenge in computational chemistry, frequently hindered by the inherent limitations of conventional Thermodynamic Integration (TI) methods. These limitations include poor phasespace overlap between discrete alchemical states, inefficient…
Philip Nega, Zhi Li, Victor Ghosh, Janak Thapa + 7 more
We use a data-driven approach to discover the influence of trace amounts of water on perovskite crystal formation. Statistical analysis of 8,470 inverse-temperature crystallization lead iodide perovskite synthesis reactions, performed over 20 months using a robotic system, revealed discrepancies between the empirical…
Pamela Reinagel
After an experiment has been completed and analyzed, a trend may be observed that is “not quite significant”. Sometimes in this situation, researchers incrementally grow their sample size N in an effort to achieve statistical significance. This is especially tempting in situations when samples are very costly or…