25 papers · ranked by Valyu relevance
Hyonho Chun, Sündüz Keleş
Partial least squares regression has been an alternative to ordinary least squares for handling multicollinearity in several areas of scientific research since the 1960s. It has recently gained much attention in the analysis of high dimensional genomic data. We show that known asymptotic consistency of the partial…
Xue Yu, Yifan Sun, Hai-Jun Zhou
High-dimensional linear regression model is the most popular statistical model for high-dimensional data, but it is quite a challenging task to achieve a sparse set of regression coefficients. In this paper, we propose a simple heuristic algorithm to construct sparse high-dimensional linear regression models, which is…
Dimitris Bertsimas, Bart Van Parys
We present a novel method for exact hierarchical sparse polynomial regression. Our regressor is that degree r polynomial which depends on at most k inputs, counting at most ` monomial terms, which minimizes the sum of the squares of its prediction errors. The previous hierarchical sparse specification aligns well with…
Kuo-ching Liang, Ashwini Patil, Kenta Nakai, Niranjan Baisakh
The use of pathways and gene interaction networks for the analysis of differential expression experiments has allowed us to highlight the differences in gene expression profiles between samples in a systems biology perspective. The usefulness and accuracy of pathway analysis critically depend on our understanding of…
Junyang Qian, Yosuke Tanigawa, Ruilin Li, Robert Tibshirani + 2 more
In high-dimensional regression problems, often a relatively small subset of the features are relevant for predicting the outcome, and methods that impose sparsity on the solution are popular. When multiple correlated outcomes are available (multitask), reduced rank regression is an effective way to borrow strength and…
Hamed Firouzi, Alfred O. Hero, Bala Rajaratnam
This paper proposes a general adaptive procedure for budget-limited predictor design in high dimensions called two-stage Sampling, Prediction and Adaptive Regression via Correlation Screening (SPARCS). SPARCS can be applied to high dimensional prediction problems in experimental science, medicine, finance, and…
Xinyu Zhou, Pengtao Dang, Haixu Tang, Laura Xianlu Peng + 6 more
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general framework that estimates…
Basile Jumentier, Kevin Caye, Barbara Heude, Johanna Lepeule + 1 more
Association of phenotypes or exposures with genomic and epigenomic data faces important statistical challenges. One of these challenges is to remove variation due to unobserved confounding factors, such as individual ancestry or cell-type composition in tissues. This issue can be addressed with penalized latent factor…
Xinyu Zhou, Pengtao Dang, Xiao Wang, Laura Xianlu Peng + 7 more
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general statistical framework that…
Andreas Alfons, Christophe Croux, Sarah Gelper
Sparse model estimation is a topic of high importance in modern data analysis due to the increasing availability of data sets with a large number of variables. Another common problem in applied statistics is the presence of outliers in the data. This paper combines robust regression and sparse model estimation. A…
Shuichi Kawano, Hironori Fujisawa, Toyoyuki Takada, Toshihiko Shiroishi
'Toshihiko Shiroishi'] Principal component regression (PCR) is a widely used two-stage procedure: principal component analysis (PCA), followed by regression in which the selected principal components are regarded as new explanatory variables in the model. Note that PCA is based only on the explanatory variables, so the…
Zemin Zheng, Yang Li, Jie Wu, Yuchen Wang
Large-scale association analysis between multivariate responses and predictors is of great practical importance, as exemplified by modern business applications including social media marketing and crisis management. Despite the rapid methodological advances, how to obtain scalable estimators with free tuning of the…
Tzu‐Yu Liu, Laura Trinchera, Arthur Tenenhaus, Dennis Wei + 1 more
'Alfred O. Hero'] Abstract: Partial least squares (PLS) regression combines dimensionality reduction and prediction using a latent variable model. Since partial least squares regression (PLS-R) does not require matrix inversion or diagonalization, it can be applied to problems with large numbers of variables. As…
Xinyu Zhou, Pengtao Dang, Xiao Wang, Laura Xianlu Peng + 7 more
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general statistical framework that…
Shuichi Kawano, Hironori Fujisawa, Toyoyuki Takada, Toshihiko Shiroishi
'Toshihiko Shiroishi'] Abstract: Principal component regression (PCR) is a two-stage procedure that selects some principal components and then constructs a regression model regarding them as new explanatory variables. Note that the principal components are obtained from only explanatory variables and not considered…
Zhenqiu Liu, Gang Li
Variable selections for regression with high-dimensional big data have found many applications in bioinformatics and computational biology. One appealing approach is the L0 regularized regression which penalizes the number of nonzero features in the model directly. However, it is well known that L0 optimization is…
Donghwan Lee, Woojoo Lee, Youngjo Lee, Yudi Pawitan
Background Principal component analysis (PCA) has gained popularity as a method for the analysis of high-dimensional genomic data. However, it is often difficult to interpret the results because the principal components are linear combinations of all variables, and the coefficients (loadings) are typically nonzero.…
Sanjar Adilov
Machine learning models for molecular-property prediction typically work with molecular representations in the form of fingerprints, descriptors, or graphs. In case of fingerprints and descriptors, molecular representations usually comprise thousands of features, which causes the curse of dimensionality for many…
Kun Du
Likelihood Authors: ['Kun Du'] This paper compares convex and non-convex penalized likelihood methods in high-dimensional statistical modeling, focusing on their strengths and limitations. Convex penalties, such as LASSO, offer computational efficiency and strong theoretical guarantees, but often introduce bias in…
S. Park, E. Ceulemans, K. Van Deun
Principal component analysis (PCA) is an important tool for analyzing large collections of variables. It functions both as a pre-processing tool to summarize many variables into components and as a method to reveal structure in data. Different coefficients play a central role in these two uses. One focuses on the…
Liwei Cao, Danilo Russo, Vassilios S. Vassiliadis, Alexei Lapkin
A mixed-integer nonlinear programming (MINLP) formulation for symbolic regression was proposed to identify physical models from noisy experimental data. The formulation was tested using numerical models and was found to be more efficient than the previous literature example with respect to the number of predictor…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Kelsey Hatzell, Yanjie Zheng
X-ray Computed Tomography (CT) is a non-invasive, non-destructive approach to imaging materials, material systems and engineered components in two- and three- dimensions. Acquisition of 3D images requires the collection of hundreds or thousands of through-thickness X-ray radiographic images from different angles. Such…
Martin Seifrid, Stanley Lo, Dylan Choi, Gary Tom + 12 more
Martin Seifrid 1 , Stanley Lo 2 , Dylan G. Choi 3 , Gary Tom 2 , My Linh Le 3 , Kunyu Li 3 , Rahul Sankar 3 , Hoai-Thanh Vuong 3 , Hiba Wakidi 3 , Ahra Yi 3 , Ziyue Zhu 3 , Nora Schopp 3 , Aaron Peng 3 , Benjamin Luginbuhl 3 , Thuc-Quyen Nguyen 3 , Alán Aspuru-Guzik 2
Authors not listed
Quantitative Structure-Activity Relationship (QSAR) modeling is a pillar of computational drug discovery. However, standard machine learning (ML) models are often confounded by the high-dimensional and intensely correlated nature of molecular descriptors. A model may identify a "bulk" property (e.g., molecular weight)…