20 papers · ranked by Valyu relevance
Dominikus Noll
The EM algorithm assures monotone decrease of the incomplete data negative log-likelihood [30], but convergence of the iterates may fail in various ways [78]. Without coercivity iterates may escape to infinity while values converge. Even when iterates stay bounded, they may still fail to converge, cycle [76], or…
Vladimir Shakhov, Olga Sokolova, Paddy J. French
Air pollution monitoring systems use distributed sensors that record dynamic environmental conditions, often producing large volumes of heterogeneous and stochastic data. Efficient aggregation of this data is essential for reducing communication overhead while maintaining the quality of information for decision making.…
Xiaoru Huang, Tonghui Yu, Xiaoyu Liu
Interval-censored data commonly arise in medical studies when the event time of interest is only known to lie within an interval. In the presence of a cure subgroup, conventional mixture cure models typically assume a logistic model for the uncure probability and a proportional hazards model for the susceptible…
Guo-Liang Tian, Xuanyu Liu, Yuanfan Zhao
Although the $\textit{expectation-maximization}$ (EM) algorithm is a powerful optimization tool in statistics, it can only be applied to missing/incomplete data problems or to problems with a latent-variable structure. It is well known that the introduction of latent variables (or the data augmentation) is an art…
Annegret Seibt, Luc De Raedt, Giuseppe Marra
Neurosymbolic (NeSy) models integrate neural networks and symbolic reasoning for robust and interpretable AI. State-of-the-art NeSy models require that the symbolic component is expressed in a differentiable way, often complicating the use of approximate inference. We propose EM-NeSy which casts probabilistic NeSy…
Cabrera-Bean, Margarita, Vidal, Josep + 6 more
1 Dept. of Signal Theory and Communications Universitat Politecnica de Catalunya (UPC), Barcelona, Spain ` 2 Institut Universitari d'Investigacio en Atenci ´ o Prim ´ aria Jordi Gol, ` IDIAP Jordi Gol, Barcelona, Spain 3 Germans Trias i Pujol Research Institute (IGTP), Badalona, Spain 4 Red de Investigacion en…
Tianying Feng, Li Cai
The expectation-maximization (EM) algorithm is widely used for parameter estimation in item response theory (IRT) modeling. However, when applied to datasets with large numbers of individuals and items, the standard EM algorithm can be slow to converge, with computationally expensive E-steps. We propose a modified EM…
Joan Saurina-i-Ricos, Daniel Mas Montserrat, Alexander G. Ioannidis
Estimating genetic clusters from sequencing data is a fundamental task in population and medical genetics, enabling demographic inference and adjustment for population structure in association studies. ADMIXTURE, a widely used model-based clustering method, employs an accelerated Expectation–Maximization (EM) algorithm…
Zhiyuan Lu
The use of dual system estimation (DSE) is heavily used in Census Bureau operations. With DSE methods, it is important to implement methods to infer the population size among those with missing data from one or both data sources. The use of log-linear models, calculated through EM algorithms, promises a way for…
Xin Chen, Jiale Li, Heinrich Wörtche
A robust multi-sensor recursive Expectation-Maximization (RMSREM) algorithm is proposed in this paper for autoregressive eXogenous (ARX) models, addressing the challenges of heavy-tailed noise, as well as the difficulty in simultaneously processing multi-sensor information. First, for the potential outliers in…
David Silva-Sánchez, Erik H. Thiede, Roy R. Lederman, Pilar Cossio
Biomolecules are inherently dynamic, and understanding their conformational ensemble distributions is essential for understanding their dynamics and biological roles. Cryo-electron microscopy (cryo-EM), a technique that images individual biomolecules frozen in a thin layer of amorphous ice, has emerged as a leading…
Peter Carbonetto, Abhishek Sarkar, Zihao Wang, Matthew Stephens
In an effort to develop topic modeling methods that can be quickly applied to large data sets, we revisit the problem of maximum-likelihood estimation in topic models. It is known, at least informally, that maximum-likelihood estimation in topic models is closely related to non-negative matrix factorization (NMF). Yet…
Sumathi Subbarayan, G. Hannah Grace
Introduction Clustering high-dimensional and noisy data remains challenging for conventional expectation-maximization (EM) methods as overlapping clusters, sparse features, and outliers can lead to covariance degeneracy and unstable parameter estimates. This research aims to improve clustering performance in…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Andrea Polo-Rodríguez, David R. Penas, Julio R. Banga
Parameter estimation is a central challenge in systems biology, particularly for large dynamic models described by nonlinear ordinary differential equations (ODEs). These global optimization problems exhibit landscapes which are topologically heterogeneous, often exhibiting a pathological mixture of stiff, smooth…
Authors not listed
Two kinetic schemes of the general modifier mechanism have been analysed in a quasi-steady state approximation, assuming that the reaction product concentration is negligible (a natural assumption for the initial rate method) and without additional simplifying assumptions. The characteristic equations have been…
Authors not listed
We present a theoretical and computational framework for virtual mass spectrometry based on Molecular Maxwell Demons (MMDs) operating as information catalysts. Building on the biological Maxwell demon framework, we demonstrate that mass spectrometry data contain categorical state information that is fundamentally…
Authors not listed
Exploring the potential energy surface to sample transition state regions is crucial to understand the atomic processes that govern chemical reactivity. Ideally, the exploration does not require any collective variables that are based on prior chemical domain knowledge. With this in mind, we adapt the stochastic saddle…
Authors not listed
The political, social and economic consequences of climate change drastically influence the requirements of modern energy systems and its components. This includes not only energy production but also concepts and innovations for its storage, especially in magnitudes of gigawatt hours. Carnot batteries, which convert…
Alex N Popinga, Jack Forman, Dmitri Svetlov, Huy Vo + 1 more
Biological data is prone to both intrinsic and extrinsic noise and variability between experimental replicas. That same stochasticity and heterogeneity can carry information about underlying biochemical mechanisms but, if not incorporated in modeling and probabilistic inference, can also bias parameter estimates and…