27 papers · ranked by Valyu relevance
Yiquan Zou, Wenxuan Chen, Tianxiang Liang, Biao Xiong + 3 more
Prior studies on indoor LiDAR point-cloud semantic segmentation consistently report that sampling density strongly affects segmentation accuracy as well as runtime and memory, establishing an accuracy-efficiency trade-off. Nevertheless, in practice, the density is often chosen heuristically and reported under…
Bokun Sun, Ziyang Wang, Jiayun Huang, Yumeng Li + 4 more
While particulate matter (PM) instruments are widely used for air quality monitoring and policy development, there is limited research on how wind speed (U0) and contact angle (θ) affect the measurement accuracy of submicron PM, or particles with their diameters ≤ one µm (PM1). This study addresses this gap by…
Haoyang Lu, Hang Zhang, Li Yi
Efficient information sampling is crucial for human inference and decision-making even for young children. It is also closely associated with the core symptoms of autism spectrum disorder (ASD), since both the social interaction difficulties and repetitive behaviors suggest that autistic people may sample information…
Jacqueline E Rudolph, Yiyi Zhou, Karine Yenokyan, Xiaoqiang Xu + 4 more
A challenge to research in big data is the inherent computational intensity of analyses, particularly when using rigorous methods to address biases. We demonstrate the use of sampling methods in big data to estimate parameters using fewer resources. Our motivating question was whether lung cancer incidence differs by…
Guangya Wan, Zixin Stephen Xu, Sasa Zorc, Manel Baucells + 3 more
Sampling multiple responses is a common way to improve LLM output quality, but it comes at the cost of additional computation. The key challenge is deciding when to stop generating new samples to balance accuracy gains against efficiency. To address this, we introduce BEA-CON (Bayesian Efficient Adaptive Criterion for…
Nicholas Marco, Surya T. Tokdar
Elliptical slice sampling is a widely used gradient-free Markov chain Monte Carlo algorithm that is tuning-free and capable of adapting to local characteristics of the target distribution. However, its primary limitation is that sampling efficiency can quickly degrade when there is a mismatch between the prior…
Foo Hui-Mean, Yuan‐chin Ivan Chang
In large-scale statistical modeling, reducing data size through subsampling is essential for balancing computational efficiency and statistical accuracy. We propose a new method, Principal Component Analysis guided Quantile Sampling (PCA-QS), which projects data onto principal components and applies quantilebased…
Zhen Liu, Wenbo Dong, Lei Gu, Ruicheng Ge + 4 more
Single-cell proteomics (SCP) enables direct measurement of protein heterogeneity but remains constrained by throughput and limited applicability to primary tissues. Here, we present an integrated workflow developed to address both challenges. We engineered SPRINT, an AI-powered bioprinting platform that prepares more…
Jasper B. Yang, Thomas Lumley, Bryan E. Shepherd, Pamela A. Shaw
Recent works have proposed optimal subsampling algorithms to improve computational efficiency in large datasets and to design validation studies in the presence of measurement error. Existing approaches generally fall into two categories: (i) designs that optimize individualized sampling rules, where unit-specific…
Shaoyang Guo, Qian Sun, Xiaoyu Li
Introduction The No-U-Turn Sampler (NUTS), widely applied in psychometrics via the Stan platform, lacks algorithm-level systematic introduction for item response theory (IRT) models and tailored optimizations for specific models. This study systematically explicates the NUTS algorithm for the 4-parameter normal ogive…
Kosuke Morikawa, Jae Kwang Kim
Integrating probability and non-probability samples is increasingly important, yet unknown sampling mechanisms in non-probability sources complicate identification and efficient estimation. We develop semiparametric theory for dual-frame data integration and propose two complementary estimators. The first models the…
Eda Gizem Koçyiğit
Ranked Set Sampling (RSS) is known for its efficiency in parameter estimation, especially when ranking is more feasible than actual measurement. This study introduces a novel memory type estimator for RSS based on Hybrid Exponentially Weighted Moving Averages (HEWMA), using two auxiliary variables. The estimator aims…
Geunyong Kim, Molly L. Shen, Andy Ng, David Juncker
Ultrasensitive, quantitative and simple tests for proteins and nucleic acids could empower analysis and disease diagnosis, but slowness of analyte capture and detection prevent it. We introduce the microfluidic sieve-detector (MSD) based on microfluidic Brownian affinity traps (BATs). A BAT is a micro-conduit coated…
Joseph Rich, Lior Pachter
Summary: fastQpick is a command-line tool and Python library for sampling FASTQ reads with replacement. Sampling with replacement turns a single FASTQ file into an arbitrary number of bootstrap replicates, which enables uncertainty quantification and statistical analysis at the level of raw reads. This process answers…
Authors not listed
Thorough treatment of conformation in computational chemistry is required to capture the subtle energy differences that lead to experimental observations. Accurate quantum chemistry calculations are very expensive and evaluation of the entire ensemble found during a conformational search is often unachievable. This is…
Asaf Cohen
Traditional importance sampling (IS) is designed to estimate rare-event probabilities by minimizing estimator variance. However, many applications prioritize rapid discovery: the generation of a trajectory within a rare set $A_n$. This requires a shift from ensemble-based estimation to a design principle focused on the…
Ashwani Rajput, Neeraj Joshi
This paper investigates the problem of comparing the location parameters of two shifted exponential models through a novel double sequential sampling framework. The proposed hypothesis testing procedure is developed by controlling the type I error probability at a preassigned level while minimizing a loss function that…
Mikael Kubista, Amin Forootan, Michael W. Pfaffl, Stephen A. Bustin + 5 more
The quantitative polymerase chain reaction (PCR) standard curve is the central analytical tool for validating qPCR assays and can also be used to estimate target concentrations in test samples. This review explains how qPCR standard curves are constructed, validated, and analyzed for different purposes. We first…
Takashi Fukuzawa, Yanjie Zhao, Naofumi Nishizawa, Hisao Nagata + 1 more
Environmental DNA (eDNA) methodology is widely applied in the biomonitoring of organisms, but it requires the target DNA to be detected in a simple, stable, and highly sensitive manner. Detection sensitivity of eDNA measurement becomes particularly critical when monitoring species present at low abundance. In this…
Lia K. Domke, Kimberly J. Ledger, Jessica M. Whitney, Shannon M. Kachel + 3 more
Ecological studies aim to understand species distributions, yet the methods used to sample affect which species are detected (i.e., gear selectivity) and may be influenced by species traits. Interactions between traditional gear selectivity for marine fishes and their species traits have been studied, but such studies…
Authors not listed
Continuous manufacturing processes offer significant advantages over batch processes, including easier scalability, reduced costs, lower raw material and solvent consumption, and improved energy efficiency. A robust techno-economic assessment is therefore essential to evaluate and facilitate the adoption of such…
Ela Iwaszkiewicz-Eggebrecht, Emma Granqvist, Karol H. Nowak, Catalina Valdivia + 9 more
1. DNA metabarcoding—high-throughput sequencing of barcode regions from bulk samples—has become a key tool for insect biodiversity assessment. Yet, how methodological choices affect the accuracy of metabarcoding data remains insufficiently explored. In this paper, we ask: (1) How does the lysis method (non-destructive…
Authors not listed
Incorporating prior domain knowledge into Bayesian optimization (BO) remains difficult for statistical methods, which also typically suffer from limited interpretability. Large language models (LLMs) offer complementary strengths in reasoning and knowledge integration, but it remains unclear when and how they improve…
Authors not listed
The rapid growth of worldwide computing power has transformed in silico chemistry into a discipline that is integrated into the daily work of many chemists. Nowadays, researchers find it increasingly straightforward to predict a wide range of molecular properties and chemi- cal processes at reasonable computational…
James H. McVittie, Martin Lysy, Masoud Asgharian
Tolerance limits have received considerable attention in the statistical literature, with applications reaching far beyond their initial role in quality control. The well-known formula of Scheffé and Tukey (1944) establishes a simple, distribution-free relation between sample size and population coverage by two given…
Authors not listed
This paper addresses the challenge of decarbonizing global energy systems by proposing the Shibah Integrated Bio-Electro-AI CCUS-H2 Framework, a multidisciplinary approach that combines hydrogen production, storage, and utilization with carbon capture, utilization, and storage (CCUS). The framework tackles high costs…
Authors not listed
Biological processes underpin centralized wastewater treatment but are difficult to deploy at small scale. Thermomechanical and thermochemical approaches could enable household-level sanitation, yet their economic and environmental potential remains unclear. We assessed two prototype household reinvented toilets…