27 papers · ranked by Valyu relevance
Martin Palazzo, Pierre Beauseroy, Patricio Yankilevich
Background Next generation sequencing instruments are providing new opportunities for comprehensive analyses of cancer genomes. The increasing availability of tumor data allows to research the complexity of cancer disease with machine learning methods. The large available repositories of high dimensional tumor samples…
Longlong Liao, Kenli Li, Keqin Li, Canqun Yang + 1 more
Background While there are a large number of bioinformatics datasets for clustering, many of them are incomplete, i.e., missing attribute values in some data samples needed by clustering algorithms. A variety of clustering algorithms have been proposed in the past years, but they usually are limited to cluster on the…
Chun‐Liang Li, Wei-Cheng Chang, Youssef Mroueh, Yiming Yang + 1 more
'Barnabás Póczos'] Kernels are powerful and versatile tools in machine learning and statistics. Although the notion of universal kernels and characteristic kernels has been studied, kernel selection still greatly influences the empirical performance. While learning the kernel in a data driven way has been investigated…
Nisar Wani, Khalid Raza
Computer aided diagnosis is gradually making its way into the domain of medical research and clinical diagnosis. With field of radiology and diagnostic imaging producing petabytes of image data. Machine learning tools, particularly kernel based algorithms seem to be an obvious choice to process and analyze this high…
Mitja Briscik, Gabriele Tazza, László Vidács, Marie-Agnès Dillies + 1 more
'Sébastien Déjean'] Background Advances in high-throughput technologies have originated an ever-increasing availability of omics datasets. The integration of multiple heterogeneous data sources is currently an issue for biology and bioinformatics. Multiple kernel learning (MKL) has shown to be a flexible and valid…
Xuehua Li, Lan Shu
Genomic microarrays are powerful research tools in bioinformatics and modern medicinal research because they enable massively-parallel assays and simultaneous monitoring of thousands of gene expression of biological samples. However, a simple microarray experiment often leads to very high-dimensional data and a huge…
Qi Mao, Ivor W. Tsang
Due to the growing ubiquity of unlabeled data, learning with unlabeled data is attracting increasing attention in machine learning. In this paper, we propose a novel semi-supervised kernel learning method which can seamlessly combine manifold structure of unlabeled data and Regularized Least-Squares (RLS) to learn a…
Abhishek Kumar, Alexandru Niculescu-Mizil, Koray Kavukcuoglu, Hal Daumé
'Hal Daumé'] With the advent of kernel methods, automating the task of specifying a suitable kernel has become increasingly important. In this context, the Multiple Kernel Learning (MKL) problem of finding a combination of prespecified base kernels that is suitable for the task at hand has received significant…
Daan Van Hauwermeiren, Michiel Stock, Thomas De Beer, Ingmar Nopens
In the pharmaceutical industry, the transition to continuous manufacturing of solid dosage forms is adopted by more and more companies. For these continuous processes, high-quality process models are needed. In pharmaceutical wet granulation, a unit operation in the ConsiGma $\text{TM}$-25 continuous powder-to-tablet…
Jérôme Mariette, Nathalie Villa-Vialaneix
Recent high-throughput sequencing advances have expanded the breadth of available omics datasets and the integrated analysis of multiple datasets obtained on the same samples has allowed to gain important insights in a wide range of applications. However, the integration of various sources of information remains a…
Christopher M. Wilson, Kaiqiao Li, Pei-Fen Kuan, Xuefeng Wang
Advances in medical technology have allowed for customized prognosis, diagnosis, and personalized treatment regimens that utilize multiple heterogeneous data sources. Multiple kernel learning (MKL) is well suited for integration of multiple high throughput data sources, however, there are currently no implementations…
Huan Song, Jayaraman J. Thiagarajan, Prasanna Sattigeri, Andreas Spanias
'Andreas Spanias'] Abstract—Building highly non-linear and non-parametric models is central to several state-of-the-art machine learning systems. Kernel methods form an important class of techniques that induce a reproducing kernel Hilbert space (RKHS) for inferring non-linear models through the construction of…
J. Emmanuel Johnson, Valero Laparra, Adrián Pérez-Suay, Miguel D. Mahecha + 2 more
Kernel methods are powerful machine learning techniques which use generic non-linear functions to solve complex tasks. They have a solid mathematical foundation and exhibit excellent performance in practice. However, kernel machines are still considered black-box models as the kernel feature mapping cannot be accessed…
Benyamin Ghojogh, Ali Ghodsi, Fakhri Karray, Mark Crowley
This is a tutorial and survey paper on kernels, kernel methods, and related fields. We start with reviewing the history of kernels in functional analysis and machine learning. Then, Mercer kernel, Hilbert and Banach spaces, Reproducing Kernel Hilbert Space (RKHS), Mercer's theorem and its proof, frequently used…
Masoud Badiei Khuzani, Liyue Shen, Shahin Shahrampour, Lei Xing
We propose a novel supervised learning method to optimize the kernel in the maximum mean discrepancy generative adversarial networks (MMD GANs), and the kernel support vector machines (SVMs). Specifically, we characterize a distributionally robust optimization problem to compute a good distribution for the random…
Dai Feng, Richard Baumgartner
BreimanâĂŹs random forest (RF) can be interpreted as an implicit kernel generator, where the ensuing proximity matrix represents the data-driven RF kernel. Kernel perspective on the RF has been used to develop a principled framework for theoretical investigation of its statistical properties. However, practical utility…
Christopher M. Wilson, Kaiqiao Li, Qiang Sun, Pei Fen Kuan + 1 more
The Cox proportional hazard model is the most widely used method in modeling time-to-event data in the health sciences. A common form of the loss function in machine learning for survival data is also mainly based on Cox partial likelihood function, due to its simplicity. However, the optimization problem becomes…
Ping Yang, E. Adrian Henle, Xiaoli Fern, Cory M. Simon
Pesticides benefit agriculture by increasing crop yield, quality, and security. However, pesticides may inadvertently harm bees, which are agriculturally and ecologically vital as pollinators. The development of new pesticides---driven by pest resistance to and demands to reduce negative environmental impacts of…
Martin Seifrid, Stanley Lo, Dylan Choi, Gary Tom + 12 more
Martin Seifrid 1 , Stanley Lo 2 , Dylan G. Choi 3 , Gary Tom 2 , My Linh Le 3 , Kunyu Li 3 , Rahul Sankar 3 , Hoai-Thanh Vuong 3 , Hiba Wakidi 3 , Ahra Yi 3 , Ziyue Zhu 3 , Nora Schopp 3 , Aaron Peng 3 , Benjamin Luginbuhl 3 , Thuc-Quyen Nguyen 3 , Alán Aspuru-Guzik 2
Lluís A. Belanche-Muñoz, Małgorzata Wiejacha, Jose C. Principe
Kernel methods have played a major role in the last two decades in the modeling and visualization of complex problems in data science. The choice of kernel function remains an open research area and the reasons why some kernels perform better than others are not yet understood. Moreover, the high computational costs of…
Authors not listed
Metastable states and the conformational transitions in between them are key to understanding dynamical behaviour and function of large-scale molecular systems. By combining basic dimensionality reduction techniques with a state-of-the art approximation of the Koopman operator associated to molecular dynamics…
Ping Yang, E. Adrian Henle, Cory M. Simon, Xiaoli Fern
Pesticides benefit agriculture by increasing crop yield, quality, and security. However, pesticides may inadvertently harm bees, which are valuable as pollinators. Thus, candidate pesticides in development pipelines must be assessed for toxicity to bees. Leveraging a data set of 382 molecules with toxicity labels from…
Joseph Redshaw, Darren Ting, Alex Brown, Jonathan Hirst + 1 more
Antimicrobial peptides (AMPs) represent a potential solution to the growing problem of antimicrobial resistance, yet their identification through wet-lab experiments is a costly and timeconsuming process. Accurate computational predictions would allow rapid in silico screening of candidate AMPs, thereby accelerating…
Sriram K Vidyarthi, Rakhee Tiwari, Samrendra K Singh
After harvesting almond crop, accurate measurement of almond kernel sizes is a significant specification to plan, develop and enhance almond processing operations. The size and mass of the individual almond kernels are vital parameters usually associated with almond quality, particularly head almond yield. In this…
Yinuo Yang, Shuhao Zhang, Kavindri Ranasinghe, Olexandr Isayev + 1 more
In the past two decades, machine learning potentials (MLPs) have driven significant developments in chemical, biological and material sciences. The construction and training of MLPs enables fast and accurate simulations and analysis on thermodynamic and kinetic properties. This review focuses on the applications of…
Authors not listed
We adapted an existing approach to identifying stabilisable crystal structures from prediction sets - the Generalised Convex Hull (GCH) - to improve its application to molecular crystal structures. This was achieved by modifying the Smooth Overlap of Atomic Positions (SOAP) kernel to define the similarity of molecular…
Souvik Manna, Diptendu Roy, Sandeep Das, Biswarup Pathak
Application of data science and machine learning (ML) techniques in the domain of materials science has been increasing by leaps and bounds recently. With the help of ML, through input features derived from available databases we can rapidly screen materials based on our desired output. Capacity is one of the important…