24 papers · ranked by Valyu relevance
Dingyi Wang, Haiying Wang, Qingpei Hu
Subsampling is a widely used and effective approach for addressing the computational challenges posed by massive datasets. Substantial progress has been made in developing non-uniform, probability-based subsampling schemes that prioritize more informative observations. We propose a novel stratification mechanism that…
Prateek Mittal, Jai Dalmotra, Joohi Chauhan
Non-Standard Data Environments Authors: ['Prateek Mittal' 'Jai Dalmotra' 'Joohi Chauhan'] This paper addresses the challenge of estimating high-dimensional parameters in non-standard data environments, where traditional methods often falter due to issues such as heavy-tailed distributions, data contamination, and…
Jasper B. Yang, Thomas Lumley, Bryan E. Shepherd, Pamela A. Shaw
Recent works have proposed optimal subsampling algorithms to improve computational efficiency in large datasets and to design validation studies in the presence of measurement error. Existing approaches generally fall into two categories: (i) designs that optimize individualized sampling rules, where unit-specific…
Gavin E. Arneill, Christopher M. Perrins, Matt J. Wood, David Murphy + 4 more
'Luca Pisani' 'Mark J. Jessopp' 'John L. Quinn' 'Camille Lebarbenchon'] Sampling approaches used to census and monitor populations of flora and fauna are diverse, ranging from simple random sampling to complex hierarchal stratified designs. Usually the approach taken is determined by the spatial and temporal…
Kayli Paterson, Michael Silverstan, Barbara Beckingham
Quantifying microplastics and other microparticles is a matter of interest in the field of environmental science. Stereomicroscopy is one of the most common methods to identify and enumerate micro-size particles. However, the process of enumerating an entire environmental sample can be tedious and time-consuming…
Dongyuan Song, Nan Miles Xi, Jingyi Jessica Li, Lin Wang
The number of cells measured in single-cell transcriptomic data has grown fast in recent years. For such large-scale data, subsampling is a powerful and often necessary tool for exploratory data analysis. However, the easiest random subsampling is not ideal from the perspective of preserving rare cell types. Therefore…
Fatimah A. Almulhim, Hassan M. Aljohani, Boris Ryabko
Most traditional estimators assume normality and remain sensitive to extreme observations, which limits their usefulness in practical applications. To improve accuracy, we introduce quintile-based median estimators using transformation methods in a stratified two-phase sampling technique. The design allows for…
Trong Duc Nguyen, Ming‐Hung Shih, Divesh Srivastava, Srikanta Tirthapura + 1 more
'Srikanta Tirthapura' 'Bojian Xu'] Stratified random sampling (SRS) is a fundamental sampling technique that provides accurate estimates for aggregate queries using a small size sample, and has been used widely for approximate query processing. A key question in SRS is how to partition a target sample size among…
Raphaël Jauslin, Esther Eustache, Yves Tillé
A balanced sampling design should always be the adopted strategies if auxiliary information is available. Besides, integrating a stratified structure of the population in the sampling process can considerably reduce the variance of the estimators. We propose here a new method to handle the selection of a balanced…
Anurag Kumar, Bhiksha Raj
In this paper we propose strategies for estimating performance of a classifier when labels cannot be obtained for the whole test set. The number of test instances which can be labeled is very small compared to the whole test data size. The goal then is to obtain a precise estimate of classifier performance using as…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…
Sixia Chen, David Haziza, Zeinab Mashreghi
Multi-stage sampling designs are often used in household surveys because a sampling frame of elements may not be available or for cost considerations when data collection involves face-to-face interviews. In this context, variance estimation is a complex task as it relies on the availability of second-order inclusion…
Mahmut Doğramaci, Sandra J. DeBano, David E. Wooster, Chiho Kimoto
Significant progress has been made in developing subsampling techniques to process large samples of aquatic invertebrates. However, limited information is available regarding subsampling techniques for terrestrial invertebrate samples. Therefore a novel subsampling procedure was evaluated for processing samples of…
Yan Tian, Jiaxin Song, Boris Ryabko
With the advancement of information technology, large-scale data have become increasingly common. Subsampling methods for the statistical analysis of such data require computing the sampling probability for each observation, a process that can be computationally intensive. In this paper, we extend the perturbed…
Nicolás Mongiardino Koch
Phylogenomic subsampling is a procedure by which small sets of loci are selected from large genome-scale datasets and used for phylogenetic inference. This step is often motivated by either computational limitations associated with the use of complex inference methods, or as a means of testing the robustness of…
David E. Hufnagel, Matthew B. Hufford, Arun S. Seetharam
PacBio sequencing is an incredibly valuable third-generation DNA sequencing method due to very long read lengths, ability to detect methylated bases, and its real-time sequencing methodology. Yet, hitherto no tool was available for analyzing the quality of, subsampling, and filtering PacBio data. Here we present…
Alexandre Wendling, Clovis Galiez
The analysis of binary outcomes and features, such as the effect of vaccination on health, often rely on 2 × 2 contingency tables. However, confounding factors such as age or gender call for stratified analysis, by creating sub-tables, which is common in bioscience, epidemiological, and social research, as well as in…
Authors not listed
Structural elucidation of unknown compounds using tandem mass spectrometry (MS/MS) is an ongoing challenge. Expert chemists will often use common product ion peaks and neutral losses or peak matching software to predict structural aspects from MS/MS spectra, a process which is time consuming and limited by the…
Authors not listed
The design-make-test cycle for drug discovery is highly dependent on the purification of synthesized compounds. Prior to evaluation of suitability, ultrahigh performance liquid chromatography is used for an initial standard analysis, where retention times of analytes are measured with a shorter standard gradient method…
Christopher R. John, David Watson, Dominic Russ, Katriona Goldmann + 4 more
Genome-wide data is used to stratify patients into classes for precision medicine using clustering algorithms. A common problem in this area is selection of the number of clusters (K). The Monti consensus clustering algorithm is a widely used method which uses stability selection to estimate K. However, the method has…
Authors not listed
Quantification is a challenge for non-targeted analysis (NTA) with liquid chromatography–high resolution mass spectrometry (LC–HRMS), due to the lack of analytical standards. Quantification via structure-based predicted ionization efficiency (IE) was found to provide the highest accuracy in estimating concentration.…
Joshua Hesse, Davide Boldini, Stephan Sieber
In the rapidly evolving field of drug discovery, High Throughput Screening (HTS) is a pivotal technique for identifying promising compounds. Despite its wide usage, the primary challenge remains in efficiently sifting through vast chemical libraries to discern true bioactive compounds from false positives. This study…
Wei Yang, Jacob Schreiber, Jeffrey Bilmes, William Stafford Noble
Analyzing and sharing massive single-cell RNA-seq data sets can be facilitated by creating a “sketch” of the data—a selected subset of cells that accurately represent the full data set. Using an existing benchmark, we demonstrate the utility of submodular optimization in efficiently creating high quality sketches of…
Zuguang Gu, Daniel Hübschmann
Consensus partitioning is an unsupervised method widely used in high throughput data analysis for revealing subgroups and assigns stability for the classification. However, standard consensus partitioning procedures are weak to identify large numbers of stable subgroups. There are two main issues. 1. Subgroups with…