20 papers · ranked by Valyu relevance
Meelad Amouzgar, David R. Glass, Reema Baskar, Inna Averbukh + 4 more
Single-cell technologies generate large, high-dimensional datasets encompassing a diversity of omics. Dimensionality reduction enables visualization of data by representing cells in two-dimensional plots that capture the structure and heterogeneity of the original dataset. Visualizations contribute to human…
Werner van der Merwe, Herman Kamper, Johan A. du Preez
Latent Dirichlet allocation (LDA) is widely used for unsupervised topic modelling on sets of documents. No temporal information is used in the model. However, there is often a relationship between the corresponding topics of consecutive tokens. In this paper, we present an extension to LDA that uses a Markov chain to…
Authors not listed
Deciphering the correct mechanism governing certain phenomenon in polyelectrolyte (PE) brush grafted systems, revealed through atomistic simulations, is an extremely challenging problem. In a recent study, our all-atom molecular dynamics (MD) simulations revealed a non-linearly large electroosmotic flow (in the…
Maxat Tezekbayev, Arman Bolatov, Zhenisbek Assylbekov
We revisit Deep Linear Discriminant Analysis (Deep LDA) from a likelihood-based perspective. While classical LDA is a simple Gaussian model with linear decision boundaries, attaching an LDA head to a neural encoder raises the question of how to train the resulting deep classifier by maximum likelihood estimation (MLE).…
Cencheng Shen, Yuexiao Dong
Linear discriminant analysis (LDA) is a fundamental classification and dimension reduction method that achieves Bayes optimality under Gaussian mixture, but often struggles in high-dimensional settings where the covariance matrix cannot be reliably estimated. We propose LDA with gradient optimization (LDA-GO), which…
Nicolas Heintz, Tom Francart, Alexander Bertrand
—Linear Discriminant Analysis (LDA) is one of the oldest and most popular linear methods for supervised classification problems. In this paper, we demonstrate that it is possible to compute the exact projection vector from LDA models based on unlabelled data, if some minimal prior information is available. More…
Tuan L. Vo, Uyen Dang, Thu Hien Nguyen
for Classification with Missing Data Authors: ['Tuan L. Vo' 'Uyen Dang' 'Thu Hien Nguyen'] As Artificial Intelligence (AI) models are gradually being adopted in real-life applications, the explainability of the model used is critical, especially in high-stakes areas such as medicine, finance, etc. Among the commonly…
Anastasiia Kim, Sanna Sevanto, Eric R. Moore, Nicholas Lubbers + 1 more
'Georg Zeller'] Interactions between stressed organisms and their microbiome environments may provide new routes for understanding and controlling biological systems. However, microbiomes are a form of high-dimensional data, with thousands of taxa present in any given sample, which makes untangling the interaction…
Meelad Amouzgar, David R. Glass, Reema Baskar, Inna Averbukh + 4 more
'Samuel C. Kimmey' 'Albert G. Tsai' 'Felix J. Hartmann' 'Sean C. Bendall'] Title: Summary Single-cell technologies generate large, high-dimensional datasets encompassing a diversity of omics. Dimensionality reduction captures the structure and heterogeneity of the original dataset, creating low-dimensional…
Yiwei Zhang, Jiawei Han, Tengjun Liu, Zelan Yang + 2 more
Spike sorting is a fundamental step in extracting single-unit activity from neural ensemble recordings, which play an important role in basic neuroscience and neurotechnologies. A few algorithms have been applied in spike sorting. However, when noise level or waveform similarity becomes relatively high, their…
Liqian Zhou, Xinhuai Peng, Lijun Zeng, Lihong Peng
Introduction: Long non-coding RNAs (lncRNAs) have been in the clinical use as potential prognostic biomarkers of various types of cancer. Identifying associations between lncRNAs and diseases helps capture the potential biomarkers and design efficient therapeutic options for diseases. Wet experiments for identifying…
Alan Min, Timothy Durham, Louis Gevirtzman, William Stafford Noble
Single cell ATAC-seq (scATAC-seq) enables the mapping of regulatory elements in fine-grained cell types. Despite this advance, analysis of the resulting data is challenging, and large scale scATAC-seq data are difficult to obtain and expensive to generate. This motivates a method to leverage information from previously…
Etana Fikadu Dinsa, Mrinal Das, Teklu Urgessa Abebe
Afaan Oromo is a resource-scarce language with limited tools developed for its processing, posing significant challenges for natural language tasks. The tools designed for English do not work efficiently for Afaan Oromo due to the linguistic differences and lack of well-structured resources. To address this challenge…
Elizabeth Thomas, Ferid Ben Ali, Arvind Tolambiya, Florian Chambellent + 1 more
The aim of this study was to develop the use of Machine Learning techniques as a means of multivariate analysis in studies of motor control. These studies generate a huge amount of data, the analysis of which continues to be largely univariate. We propose the use of machine learning classification and feature selection…
Masahiro Yoshihara, Yoshihiro Itaguchi
In the semantic variant of verbal fluency tests (VFTs), clustering analysis has become popular for examining the semantic structure. While the computational psycholinguistics approach has recently drawn attention to increasing the reproducibility of clustering analysis, such an approach is not available in all…
Authors not listed
The equilibrium binding affinity has traditionally guided the drug discovery. Despite offering insight into the extent of binding, it is occasionally inaccurate in predicting biological response. Emerging evidence suggests that the lifetime of binary complex, known as residence time (RT), is more directly correlated…
Snigdha Sarkar, Md. Shahjaman, Sukanta Das
Supervised machine learning (SML) is an approach that learns from training data with known category membership to predict the unlabeled test data. There are many SML approaches in the literature and most of them use a linear score to learn its classifier. However, these approaches fail to elucidate biodiversity from…
Yiwei Zhang, Jiawei Han, Tengjun Liu, Zelan Yang + 2 more
'Shaomin Zhang'] Spike sorting is a fundamental step in extracting single-unit activity from neural ensemble recordings, which play an important role in basic neuroscience and neurotechnologies. A few algorithms have been applied in spike sorting. However, when noise level or waveform similarity becomes relatively…
Navid Ziaei, Behzad Nazari, Uri T. Eden, Alik S. Widge + 1 more
Decoder (LDGD) Model for High-Dimensional Data Authors: ['Navid Ziaei' 'Behzad Nazari' 'Uri T. Eden' 'Alik S. Widge' 'Ali Yousefi'] Extracting meaningful information from high-dimensional data poses a formidable modeling challenge, particularly when the data is obscured by noise or represented through different…
Authors not listed
Drug Discovery is a very lengthy and resource-consuming process. However, a variety of advanced Artificial Intelligence (AI) and Deep Learning (DL) techniques are being utilized to accelerate and advance DD, such as Large Language Models (LLMs). This survey is in aim of discovering and comparing the currently available…