Search · four archives
Search · four archives
24 papers · ranked by Valyu relevance
Bei Zhu, Yi Xu, Pengcheng Zhao, Siu-Ming Yiu + 2 more
Many drugs can be metabolized by human microbes; the drug metabolites would significantly alter pharmacological effects and result in low therapeutic efficacy for patients. Hence, it is crucial to identify potential drug-microbe associations (DMAs) before the drug administrations. Nevertheless, traditional DMA…
Trang T. Le, Bryan A. Dawkins, Brett A. McKinney
Machine learning feature selection methods are needed to detect complex interaction-network effects in complicated modeling scenarios in high-dimensional data, such as GWAS, gene expression, eQTL, and structural/functional neuroimage studies for case-control or continuous outcomes. In addition, many machine learning…
Bryan A. Dawkins, Trang T. Le, Brett A. McKinney
The performance of nearest-neighbor feature selection and prediction methods depends on the metric for computing neighborhoods and the distribution properties of the underlying data. The effects of the distribution and metric, as well as the presence of correlation and interactions, are reflected in the expected…
Xinye Chen, Stefan Güttel
Fixed-radius near neighbor search is a fundamental data operation that retrieves all data points within a user-specified distance to a query point. There are efficient algorithms that can provide fast approximate query responses, but they often have a very compute-intensive indexing phase and require careful parameter…
Shahla Faisal, Gerhard Tutz
Missing values are a common phenomenon in all areas of applied research. While various imputation methods are available for metrically scaled variables, methods for categorical data are scarce. An imputation method that has been shown to work well for high dimensional metrically scaled variables is the imputation by…
Ning Liu, Jarryd Martin, Dharmesh D Bhuva, Jinjin Chen + 11 more
Understanding complex cellular niches and neighborhoods have provided new insights into tissue biology. Thus, accurate neighborhood identification is crucial, yet existing methodologies often struggle to detect informative neighborhoods and generate cell-specific neighborhood profiles. To address these limitations, we…
Suman Saha, Satya Prakash Ghrera
Nearest neighbor search is a basic computational tool used extensively in almost research domains of computer science specially when dealing with large amount of data. However, the use of nearest neighbor search is restricted for the purpose of algorithmic development by the existence of the notion of nearness among…
Shahin Pourbahrami, Leyli Mohammad Khanli
Finding neighbourhood structures is very useful in extracting valuable relationships among data samples. This paper presents a survey of recent neighborhood construction algorithms for pattern clustering and classifying data points. Extracting neighborhoods and connections among the points is extremely useful for…
Yakir Reshef, Laurie Rumker, Joyce B. Kang, Aparna Nathan + 3 more
As single-cell datasets grow in sample size, there is a critical need to characterize cell states that vary across samples and associate with sample attributes like clinical phenotypes. Current statistical approaches typically map cells to cell-type clusters and examine sample differences through that lens alone. Here…
Shaoyi Liang, Deqiang Han
Closeness measures are crucial to clustering methods. In most traditional clustering methods, the closeness between data points or clusters is measured by the geometric distance alone. These metrics quantify the closeness only based on the concerned data points’ positions in the feature space, and they might cause…
Martin Priessner, Anna Tomberg, Jon Paul Janet, Richard J. Lewis + 2 more
In the pursuit of improved compound identification and database search tasks, this study explores Heteronuclear Single Quantum Coherence (HSQC) spectra simulation and matching methodologies. HSQC spectra serve as unique molecular fingerprints, enabling a valuable balance of data collection time and information…
Sarah Chisholm, Andrew B. Stein, Neil R. Jordan, Tatjana M. Hubel + 5 more
'John Shawe‐Taylor' 'Tom Fearn' 'J. Weldon McNutt' 'Alan M. Wilson' 'Stephen Hailes'] Title: Abstract In recent years, there have been significant advances in the technology used to collect data on the movement and activity patterns of humans and animals. GPS units, which form the primary source of location data, have…
Albert Palleja, Lars J. Jensen, Yong Wang
Clustering algorithms are often used to find groups relevant in a specific context; however, they are not informed about this context. We present a simple algorithm, HOODS, which identifies context-specific neighborhoods of entities from a similarity matrix and a list of entities specifying the context. We illustrate…
Arthur Flexer, Dominik Schnitzer
The hubness phenomenon is a recently discovered aspect of the curse of dimensionality. Hub objects have a small distance to an exceptionally large number of data points while anti-hubs lie far from all other data points. A closely related problem is the concentration of distances in high-dimensional spaces. Previous…
Andrew Stokely, Lane Votapka, Marcus Hock, Abigail Teitgen + 3 more
We present the Netsci program - an open-source scientific software package that leverages GPU acceleration and a k-nearest-neighbor algorithm in order to estimate the mutual information (MI) between data in a set. The GPU acceleration presented here, as an improvement upon existing estimators, enables calculation…
Edgar López-López, Oscar Robles, Fabien Plisson, José L. Medina-Franco
Peptides are a re-emerged strategy to fight a plethora of diseases and their utility has been expanded to new areas. Now sequence-based peptide design opens up new possibilities to develop peptidic molecular entities. However, its methodological limitations (e.g., its inefficiency in designing large peptides and that…
Chris Zhang, Mary Pitman, Anjali Dixit, Sumudu Leelananda + 7 more
DNA-encoded libraries (DELs) provide the means to make and screen millions of diverse compounds against a target of interest in a single experiment. However, despite producing large volumes of binding data at a relatively low cost, the DEL selection process is susceptible to noise, necessitating computational follow-up…
Elzbieta Gralinska, Martin Vingron
In molecular biology, just as in many other fields of science, data often come in the form of matrices or contingency tables with many measurements (rows) for a set of variables (columns). While projection methods like Principal Component Analysis or Correspondence Analysis can be applied for obtaining an overview of…
Jovan Tanevski, Loan Vulliard, Felix Hartmann, Julio Saez-Rodriguez
Spatial omics data provide rich molecular and structural information about tissues, enabling novel insights into the structure-function relationship. In particular, it facilitates the analysis of the local heterogeneity of tissues and holds promise to improve patient stratification by association of finer-grained…
Steven Torrisi, Matthew Carbone, Brian Rohr, Joseph H. Montoya + 4 more
X-ray absorption spectroscopy (XAS) produces a wealth of information about the local structure of materials, but interpretation of spectra often relies on easily accessible trends and prior assumptions about the structure. Recently, researchers have demonstrated that machine learning models can automate this process to…
Ayman Taha
Intelligent geographic information system (IGIS) is one of the promising topics in GIS field. It aims at making GIS tools more sensitive for large volumes of data stored inside GIS systems by integrating GIS with other computer sciences such as Expert system (ES) Data Warehouse (DW), Decision Support System (DSS), or…
Lin Hu
This paper uses data mining technology to analyze students' English scores. In view of the influence of many factors on students' English performance, the analysis is realized by using the association rule algorithm. The thesis analyzes and applies students' English scores based on association rules and mainly does the…
Juryon Paik, Junghyun Nam, Ung Mo Kim, Dongho Won
With the advances of wireless sensor networks, they yield massive volumes of disparate, dynamic and geographically-distributed and heterogeneous data. The data mining community has attempted to extract knowledge from the huge amount of data that they generate. However, previous mining work in WSNs has focused on…
Pulan Yu
Associative classification mining (ACM) integrating association rule mining and classification has become a significant tool for knowledge discovery, especially in the chemical domain. Its major advantage is providing high accuracy as well as chemically interpretable models. Additionally, it is able to find…