25 papers · ranked by Valyu relevance
Dimitris Bertsimas, Vassilis Digalakis
We present the backbone method, a generic framework that enables sparse and interpretable supervised machine learning methods to scale to ultrahigh dimensional problems. We solve sparse regression problems with 107 features in minutes and 108 features in hours, as well as decision tree problems with 105 features in…
Xue Yu, Yifan Sun, Hai-Jun Zhou
High-dimensional linear regression model is the most popular statistical model for high-dimensional data, but it is quite a challenging task to achieve a sparse set of regression coefficients. In this paper, we propose a simple heuristic algorithm to construct sparse high-dimensional linear regression models, which is…
Eric Price, Sandeep Silwal, Samson Zhou
We explore algorithms and limitations for sparse optimization problems such as sparse linear regression and robust linear regression. The goal of the sparse linear regression problem is to identify a small number of key features, while the goal of the robust linear regression problem is to identify a small number of…
Xue Yu, Yifan Sun, Haijun Zhou
High-dimensional linear regression model is the most popular statistical model for high-dimensional data, but it is quite a challenging task to achieve a sparse set of regression coefcients. In this paper, we propose a simple heuristic algorithm to construct sparse high-dimensional linear regression models, which is…
Jianqing Fan, Zhipeng Lou, Mengxin Yu
We propose the Factor Augmented sparse linear Regression Model (FARM) that not only encompasses both the latent factor regression and sparse linear regression as special cases but also bridges dimension reduction and sparse regression together. We provide theoretical guarantees for the estimation of our model under the…
Keisuke Teramoto, Kei Hirose
Summary. In the field of materials science and engineering, statistical analysis and machine learning techniques have recently been used to predict multiple material properties from an experimental design. These material properties correspond to response variables in the multivariate regression model. This study…
Amber Srivastava, Alisina Bayati, Srinivasa M. Salapaka
— This work presents a new approach to solve the sparse linear regression problem, i.e., to determine a k-sparse vector w ∈ R d that minimizes the cost ∥y−Aw∥ 2 2. In contrast to the existing methods, our proposed approach splits this k-sparse vector into two parts — (a) a column stochastic binary matrix V , and (b) a…
Anthony Christidis, Stefan Van Aelst, Ruben H. Zamar
The two primary approaches for high-dimensional regression problems are sparse methods (e.g. best subset selection which uses the ℓ0-norm in the penalty) and ensemble methods (e.g. random forests). Although sparse methods typically yield interpretable models, they are often outperformed in terms of prediction accuracy…
Xinyu Zhou, Pengtao Dang, Xiao Wang, Laura Xianlu Peng + 7 more
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general statistical framework that…
Dongni Jia, Xiaofeng Zhou, Shuai Li, Haibo Shi + 1 more
Data-driven modeling of nonlinear industrial processes is often complicated by heterogeneous temporal dynamics, measurement noise, and fixed-rate data acquisition. Under such conditions, direct regression on raw time-series data may become sensitive to sampling imbalance and fast transient behavior, leading to degraded…
Xinyu Zhou, Pengtao Dang, Haixu Tang, Laura Xianlu Peng + 6 more
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general framework that estimates…
Hema Sri Sai Kollipara, Tapabrata Maiti, Sanjukta Chakraborty, Samiran Sinha
Genomics and other studies encounter many features and a selection of essential features with high accuracy is desired. In recent years, there has been a significant advancement in the use of Bayesian inference for variable (or feature) selection. However, there needs to be more practical information regarding their…
Xinyu Zhou, Pengtao Dang, Xiao Wang, Laura Xianlu Peng + 7 more
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general statistical framework that…
Nuray Sogunmez Erdogan, Deniz Eroglu
Mapping cell distributions across spatial locations with whole-genome coverage is essential for understanding cellular responses and signaling pathways. However, current deconvolution models often assume strong overlap between reference and spatial datasets, neglecting biological constraints like sparsity and cell-type…
Ruo-Hui Huang, Zi-Lu Ge, Gang Xu, Qing-Ming Zeng + 5 more
Background: Prostate cancer (PCa) is a malignant tumor of the male reproductive system, and its incidence has increased significantly in recent years. This study aimed to further identify candidate biomarkers with prognostic and diagnostic significance by integrating gene expression and DNA methylation data from PCa…
Ruofan Wang, Lei Fang, Yue Wang, Jin Jin
Leveraging observational data to understand the associations between risk factors and disease outcomes and conduct disease risk prediction is a common task in epidemiology. While traditional linear regression and other machine learning models have been extensively implemented for this task, the associations between…
Sanjar Adilov
Machine learning models for molecular-property prediction typically work with molecular representations in the form of fingerprints, descriptors, or graphs. In case of fingerprints and descriptors, molecular representations usually comprise thousands of features, which causes the curse of dimensionality for many…
Isaac Xoese Ocloo, Hanfeng Chen, Yuehua Wu
In this paper, the LASSO method with extended Bayesian information criteria (EBIC) for feature selection in high-dimensional models is studied. We propose the use of the energy distance correlation in place of the ordinary correlation coefficient to measure the dependence of two variables. The energy distance…
Kun Du
Likelihood Authors: ['Kun Du'] This paper compares convex and non-convex penalized likelihood methods in high-dimensional statistical modeling, focusing on their strengths and limitations. Convex penalties, such as LASSO, offer computational efficiency and strong theoretical guarantees, but often introduce bias in…
Alain J. Mbebi, Zoran Nikoloski
Despite extensive research efforts, reconstruction of gene regulatory networks (GRNs) from transcriptomics data remains a pressing challenge in systems biology. While non-linear approaches for reconstruction of GRNs show improved performance over simpler alternatives, we do not yet have understanding if joint modelling…
Alicia Zeng, Jack Gallant
Encoding models based on word embeddings or artificial neural network (ANN) features reliably predict brain responses to naturalistic stimuli but remain difficult to interpret. A central limitation is superposition—the entanglement of distinct semantic features along correlated directions in dense embeddings, which…
Arisa Toda, Misa Goudo, Masahiro Sugimoto, Satoru Hiwa + 1 more
Machine learnings such as multivariate analyses and clustering have been frequently used for metabolomics data analyses. In metabolomics data analyses, how much difference there is between the results calculated by supervised and unsupervised learning models is an interesting topic. Since metabolomics data include…
S. Park, E. Ceulemans, K. Van Deun
Principal component analysis (PCA) is an important tool for analyzing large collections of variables. It functions both as a pre-processing tool to summarize many variables into components and as a method to reveal structure in data. Different coefficients play a central role in these two uses. One focuses on the…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Kelsey Hatzell, Yanjie Zheng
X-ray Computed Tomography (CT) is a non-invasive, non-destructive approach to imaging materials, material systems and engineered components in two- and three- dimensions. Acquisition of 3D images requires the collection of hundreds or thousands of through-thickness X-ray radiographic images from different angles. Such…