19 papers · ranked by Valyu relevance
Wen Bo Liu, Sheng Nan Liang, Xi Wen Qin, Seyedali Mirjalili
Gene expression data has the characteristics of high dimensionality and a small sample size and contains a large number of redundant genes unrelated to a disease. The direct application of machine learning to classify this type of data will not only incur a great time cost but will also sometimes fail to improved…
Hyunwook Koh
In high-dimensional omics studies, researchers often conduct kernel association testing to power-fully detect the relationship of the genetic or microbial composition with human health or disease. Especially, in human microbiome studies, its dimension reduction analysis follows to visually represent complex microbiome…
Nora K. Speicher, Nico Pfeifer
Personalized treatment of patients based on tissue-specific cancer subtypes has strongly increased the efficacy of the chosen therapies. Even though the amount of data measured for cancer patients has increased over the last years, most cancer subtypes are still diagnosed based on individual data sources (e.g. gene…
Zhongyuan Lyu, Ming Yuan
Principal component analysis (PCA) is arguably the most widely used approach for large-dimensional factor analysis. While it is effective when the factors are sufficiently strong, it can be inconsistent when the factors are weak and/or the noise has complex dependence structure. We argue that the inconsistency often…
Hui Dai, Yang Zhao, Cheng Qian, Min Cai + 7 more
'Juncheng Dai' 'Zhibin Hu' 'Hongbing Shen' 'Feng Chen' 'Joseph Devaney'] Genome-wide association studies (GWAS) are popular for identifying genetic variants which are associated with disease risk. Many approaches have been proposed to test multiple single nucleotide polymorphisms (SNPs) in a region simultaneously which…
Shimeng Huang, Elisabeth Ailer, Niki Kilbertus, Niklas Pfister + 1 more
Supervised learning, such as regression and classification, is an essential tool for analyzing modern high-throughput sequencing data, for example in microbiome research. However, due to the compositionality and sparsity, existing techniques are often inadequate. Either they rely on extensions of the linear…
Jérôme Mariette, Nathalie Villa-Vialaneix
Recent high-throughput sequencing advances have expanded the breadth of available omics datasets and the integrated analysis of multiple datasets obtained on the same samples has allowed to gain important insights in a wide range of applications. However, the integration of various sources of information remains a…
Benyamin Ghojogh, Milad Sikaroudi, Hamid R. Tizhoosh, Fakhri Karray + 1 more
'Mark Crowley'] Abstract. Fisher Discriminant Analysis (FDA) is a subspace learning method which minimizes and maximizes the intra- and inter-class scatters of data, respectively. Although, in FDA, all the pairs of classes are treated the same way, some classes are closer than the others. Weighted FDA assigns weights…
Sergio Rojas-Galeano, Emily Hsieh, Dan Agranoff, Sanjeev Krishna + 2 more
'Delmiro Fernandez-Reyes' 'Gustavo Stolovitzky'] Background The analysis of complex proteomic and genomic profiles involves the identification of significant markers within a set of hundreds or even thousands of variables that represent a high-dimensional problem space. The occurrence of noise, redundancy or…
Yunlong Jiao, Jean‐Philippe Vert
We propose new positive definite kernels for permutations. First we introduce a weighted version of the Kendall kernel, which allows to weight unequally the contributions of different item pairs in the permutations depending on their ranks. Like the Kendall kernel, we show that the weighted version is invariant to…
Md. Ashad Alam, Kenji Fukumizu, Yu‐Ping Wang
Many unsupervised kernel methods rely on the estimation of the kernel covariance operator (kernel CO) or kernel cross-covariance operator (kernel CCO). Both kernel CO and kernel CCO are sensitive to contaminated data, even when bounded positive definite kernels are used. To the best of our knowledge, there are few…
Mark F. Rogers, Colin Campbell, Yiming Ying
There is significant interest in inferring the structure of subcellular networks of interaction. Here we consider supervised interactive network inference in which a reference set of known network links and nonlinks is used to train a classifier for predicting new links. Many types of data are relevant to inferring…
Minta Thomas, Kris De Brabanter, Johan AK Suykens, Bart De Moor
Background Clinical data, such as patient history, laboratory analysis, ultrasound parameters-which are the basis of day-to-day clinical decision support-are often used to guide the clinical management of cancer in the presence of microarray data. Several data fusion techniques are available to integrate genomics or…
Matthieu Lesnoff
Partial least squares regression (PLSR) is a reference method in chemometrics. In agronomy, it is used for instance to predict components of chemical composition (response variables y) of vegetal materials from spectral near infrared (NIR) data X collected from spectrometers. The principle of PLSR is to reduce the…
Hyunwook Koh
There are numerous potential confounders, including genetic, environmental, technical, and demographic factors. These factors may be known or unknown, measured or unmeasured; hence, it is extremely challenging to capture them in downstream data analysis. However, randomized block design is an efficient design technique…
Nianxiang Zhang, Tod D. Casasent, Anna K. Casasent, Shwetha V. Kumar + 4 more
Principal component analysis (PCA), a standard approach to analysis and visualization of large datasets, is commonly used in biomedical research for detecting similarities and differences among groups of samples. We initially used conventional PCA as a tool for critical quality control of batch and trend effects in…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Murat Cihan Sorkun, Dajt Mullaj, J. M. Vianney A. Koelman, Süleyman Er
Visualizing chemical spaces streamlines the analysis of molecular datasets by reducing the information to human perception level, hence it forms an integral piece of molecular engineering, including chemical library design, high-throughput screening, diversity analysis, and outlier detection. We present here ChemPlot…
Authors not listed
Electrochemical impedance spectroscopy (EIS) is one of the most widely deployed methods to characterise electrochemical systems such as batteries, fuel cells or electrolyzers. The distribution of relaxation times (DRT) represents a technique to simplify EIS data by deconvolution with a suitable kernel, while with…