27 papers · ranked by Valyu relevance
John P. Cunningham, Zoubin Ghahramani
Linear dimensionality reduction methods are a cornerstone of analyzing high dimensional data, due to their simple geometric interpretations and typically attractive computational properties. These methods capture many data features of interest, such as covariance, dynamical structure, correlation between data sets…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Carlos Óscar S. Sorzano, Javier Vargas, Alberto Pascual-Montano
—Experimental life sciences like biology or chemistry have seen in the recent decades an explosion of the data available from experiments. Laboratory instruments become more and more complex and report hundreds or thousands measurements for a single experiment and therefore the statistical methods face challenging…
Chun Kit Jeffery Hou, Kamran Behdinan
Surrogate modeling has been popularized as an alternative to full-scale models in complex engineering processes such as manufacturing and computer-assisted engineering. The modeling demand exponentially increases with complexity and number of system parameters, which consequently requires higher-dimensional engineering…
Subhrajyoty Roy
Different unsupervised models for dimensionality reduction like PCA, LLE, Shannon's mapping, tSNE, UMAP, etc. work on different principles, hence, they are difficult to compare on the same ground. Although they are usually good for visualisation purposes, they can produce spurious patterns that are not present in the…
Álvaro Huertas-García, Alejandro Martín, Javier Huertas-Tato, David Camacho
'David Camacho'] In scientific literature and industry, semantic and context-aware Natural Language Processing-based solutions have been gaining importance in recent years. The possibilities and performance shown by these models when dealing with complex Human Language Understanding tasks are unquestionable, from…
Alon Schclar, Lior Rokach, Amir Amit
We present a novel approach for the construction of ensemble classifiers based on dimensionality reduction. Dimensionality reduction methods represent datasets using a small number of attributes while preserving the information conveyed by the original dataset. The ensemble members are trained based on…
Nicholas J. Daras
We give two low-complexity algorithms, one for dimensionality reduction and one for dimensionality increase, which are applicable to any dataset, regardless of whether the set has an intrinsic dimension or not. The corresponding methods introduce chains of compositions of conformal homeomorphisms that transform any…
James W. Webber, Kevin M. Elias
High dimensionality, i.e. p > n, is an inherent feature of machine learning. Fitting a classification model directly to p-dimensional data risks overfitting and a reduction in accuracy. Thus, dimensionality reduction is necessary to address overfitting and high dimensionality. We present a novel dimensionality…
Nico Migenda, Ralf Möller, Wolfram Schenck, Chi-Hua Chen
“Principal Component Analysis” (PCA) is an established linear technique for dimensionality reduction. It performs an orthonormal transformation to replace possibly correlated variables with a smaller set of linearly independent variables, the so-called principal components, which capture a large portion of the data…
Zahra Moghaddasi, Hamid A. Jalab, Rafidah Md Noor, Saeed Aghabozorgi
Digital image forgery is becoming easier to perform because of the rapid development of various manipulation tools. Image splicing is one of the most prevalent techniques. Digital images had lost their trustability, and researches have exerted considerable effort to regain such trustability by focusing mostly on…
Brandon W Higgs, Jennifer Weller, Jeffrey L Solka
Background Accurate methods for extraction of meaningful patterns in high dimensional data have become increasingly important with the recent generation of data types containing measurements across thousands of variables. Principal components analysis (PCA) is a linear dimensionality reduction (DR) method that is…
Farid Saberi-Movahed, Kamal Berahman, Razieh Sheikhpour, Yuefeng Li + 1 more
'Shirui Pan'] Dimensionality Reduction plays a pivotal role in improving feature learning accuracy and reducing training time by eliminating redundant features, noise, and irrelevant data. Nonnegative Matrix Factorization (NMF) has emerged as a popular and powerful method for dimensionality reduction. Despite its…
Neda Pourali
Automatic image annotation is one of the most challenging problems in machine vision areas. The goal of this task is to predict number of keywords automatically for images captured in real data. Many methods are based on visual features in order to calculate similarities between image samples. But the computation cost…
Karaj Khosla, Indra Prakash Jha, Vibhor Kumar
Dimension reduction is often used for several procedures of analysis of high dimensional biomedical data-sets such as classification or outlier detection. To improve performance of such data-mining steps, preserving both distance information and local topology among data-points could be more useful than giving priority…
Benjamin J. Lengerich, Eric P. Xing
Dimensionality reduction is an important task in bioinformatics studies. Common unsupervised methods like principal components analysis (PCA) extract axes of variation that are high-variance but do not necessarily differentiate experimental conditions. Methods of supervised discriminant analysis such as partial least…
Elnaz Lashgari, Uri Maoz
Electromyography (EMG) is a simple, non-invasive, and cost-effective technology for sensing muscle activity. However, EMG is also noisy, complex, and high-dimensional. It has nevertheless been widely used in a host of human-machine-interface applications (electrical wheelchairs, virtual computer mice, prosthesis…
Huibert-Jan Joosse, Chontira Chumsaeng-Reijers, Albert Huisman, Imo E. Hoefer + 3 more
Background The routine diagnostic process increasingly entails the processing of high-volume and high-dimensional data that cannot be directly visualised. This processing may provide scaling issues that limit the implementation of these types of data into research as well as integrated diagnostics in routine care.…
Tümay Capraz, Wolfgang Huber
A fundamental step in many analyses of high-dimensional data is dimension reduction. Two basic approaches are introduction of new, synthetic coordinates, and selection of extant features. Advantages of the latter include interpretability, simplicity, transferability and modularity. A common criterion for unsupervised…
Christiane Ahlheim, Bradley C. Love
Recent advances in multivariate fMRI analysis stress the importance of information inherent to voxel patterns. Key to interpreting these patterns is estimating the underlying dimensionality of neural representations. Dimensions may correspond to psychological dimensions, such as length and orientation, or involve other…
Pitoyo Hartono
Visualizing high dimensional data by projecting them into two or three dimensional space is one of the most effective ways to intuitively understand the data's underlying characteristics, for example their class neighborhood structure. While data visualization in low dimensional space can be efficient for revealing the…
Authors not listed
The growing number and size of DNA-encoded libraries (DELs), together with the vast space of possible DEL designs, demand interpretable and scalable criteria for selecting which libraries to construct and screen against a given target. An ideal target-focused DEL shows both strong similarity with an active reference…
Murat Cihan Sorkun, Dajt Mullaj, J. M. Vianney A. Koelman, Süleyman Er
Visualizing chemical spaces streamlines the analysis of molecular datasets by reducing the information to human perception level, hence it forms an integral piece of molecular engineering, including chemical library design, high-throughput screening, diversity analysis, and outlier detection. We present here ChemPlot…
Anoop Kumar Tiwari, Rajat Saini, Abhigyan Nath, Phool Singh + 1 more
Fuzzy rough entropy established in the notion of fuzzy rough set theory, which has been effectively and efficiently applied for feature selection to handle the uncertainty in real-valued datasets. Further, Fuzzy rough mutual information has been presented by integrating information entropy with fuzzy rough set to…
Bartłomiej Fliszkiewicz, Marcin Sajdak
The aim of the following research is to assess the applicability of calculated quantum properties of molecular fragments as molecular descriptors in machine learning classification task. The research is based on bio-concentration and QM9-extended databases. A number of compounds with results from quantum-chemical…
José L. Medina-Franco, Ana L. Chávez-Hernández, Edgar López-López, Fernanda I. Saldívar-González
Technological advances and practical applications of the chemical space concept in drug discovery, natural product research, and other research areas have attracted the scientific community´s attention. The large- and ultra-large chemical spaces are associated not only with the significant increase in the number of…
Sterling Baird, Jason R. Hall, Taylor D. Sparks
Would you rather search for a line inside a cube or a point inside a square? This type of solution degeneracy often exists in physics-based simulations and wet-lab experiments, but constraining these degeneracies is often unsupported or difficult to implement in many optimization packages, requiring additional time and…