24 papers · ranked by Valyu relevance
Sergio Blanes, Fernando Casas, Ander Murua
This overview is devoted to splitting methods, a class of numerical integrators intended for differential equations that can be subdivided into different problems easier to solve than the original system. Closely connected with this class of integrators are composition methods, in which one or several low-order schemes…
Lisa Maria Kreußer, H. E. Lockyer, Eike H. Müller, Prabhdeep Singh
Splitting methods are widely used for solving initial value problems (IVPs) due to their ability to simplify complicated evolutions into more manageable subproblems. These subproblems can be solved efficiently and accurately, leveraging properties like linearity, sparsity and reduced stiffness. Traditionally, these…
Husam Abdulnabi, J. Timothy Westwood
Machine Learning (ML) models may perform inconsistently on individual classes on nominal outputs or ranges on continuous outputs, collectively referred to here as bins. Models should be assessed through metrics that consider each bin individually, called bin metrics. Inconsistent model performance is often due to model…
Husam Abdulnabi, J. Timothy Westwood
Machine Learning (ML) models may perform inconsistently on individual classes on nominal outputs or ranges on continuous outputs, collectively referred to here as bins. Models should be assessed through metrics that consider each bin individually, called bin metrics. Inconsistent model performance is often due to model…
V. Roshan Joseph, Akhil Vakayil
In this article we propose an optimal method referred to as SPlit for splitting a dataset into training and testing sets. SPlit is based on the method of Support Points (SP), which was initially developed for finding the optimal representative points of a continuous distribution. We adapt SP for subsampling from a…
Roman Joeres, David B. Blumenthal, Olga V. Kalinina
Information Leakage is an increasing problem in machine learning research. It is a common practice to report models with benchmarks, comparing them to the state-of-the-art performance on the test splits of datasets. If two or more dataset splits contain identical or highly similar samples, a model risks simply…
Winfried Auzinger, Wolfgang Herfort, Harald Hofstätter, Othmar Koch
This article is based on [1] and [2], where an approach based on Taylor expansion and the structure of its leading term as an element of a free Lie algebra was described for the setup of a system of order conditions for operator splitting methods. Along with a brief review of these materials and some theoretical…
Simona Reale, Pietro Di Stasio, Francesco Mauro, Alessandro Sebastianelli + 2 more
'Alessandro Sebastianelli' 'Paolo Gamba' 'Silvia Liberata Ullo'] Abstract—In this paper, a novel method for data splitting is presented: an iterative procedure divides the input dataset of volcanic eruption, chosen as the proposed use case, into two parts using a dissimilarity index calculated on the cumulative…
Eklavya Jain, J. Neeraja, Buddhananda Banerjee, Palash Ghosh
In machine learning, a routine practice is to split the data into a training and a test data set. A proposed model is built based on the training data, and then the performance of the model is assessed using test data. Usually, the data is split randomly into a training and a test set on an ad hoc basis. This approach…
Mohaddeseh Rahbaran, Ehsan Razeghian, Marwah Suliman Maashi, Abduladheem Turki Jalil + 7 more
'Abduladheem Turki Jalil' 'Gunawan Widjaja' 'Lakshmi Thangavelu' 'Mariya Yurievna Kuznetsova' 'Pourya Nasirmoghadas' 'Farid Heidari' 'Faroogh Marofi' 'Mostafa Jarahian'] Embryo splitting is one of the newest developed methods in reproductive biotechnology. In this method, after splitting embryos in 2-, 4-, and even…
Florian Privé
A few algorithms have been developed for splitting the genome in nearly independent blocks of linkage disequilibrium. Due to the complexity of this problem, these algorithms rely on heuristics, which makes them sub-optimal. Here we develop an optimal solution for this problem using dynamic programming. This is now…
Authors not listed
The effectiveness of machine learning (ML) in drug discovery hinges on evaluation and modeling approaches that align with how compounds are tested and compared in real experimental contexts. We observe that experimental data in public repositories like ChEMBL naturally clusters by assay origin, while retaining…
Authors not listed
Today, machine learning models are employed extensively to predict the physicochemical and biological properties of molecules. Their performance is typically evaluated on in-distribution (ID) data, i.e., data originating from the same distribution as the training data. However, the real-world applications of such…
Freimut Gebhard Herbert Hammer, Mateusz Buglowski, André Stollenwerk
A method for the anonymization of time-continuous data, which preserves the relation between the time- and value dimension is proposed in this work. The approach protects against linking- and distribution attacks by providing k-anonymity and t-closeness. Distributions can be generated from given sets using Distribution…
Nick Terry, Youngjun Choe, Enrico Scalas
Gaussian processes offer a flexible kernel method for regression. While Gaussian processes have many useful theoretical properties and have proven practically useful, they suffer from poor scaling in the number of observations. In particular, the cubic time complexity of updating standard Gaussian process models can be…
Elior Sulem, Omri Abend, Ari Rappoport
Sentence splitting is a major simplification operator. Here we present a simple and efficient splitting algorithm based on an automatic semantic parser. After splitting, the text is amenable for further fine-tuned simplification operations. In particular, we show that neural Machine Translation can be effectively used…
Munirat Yetunde Onireti, Raj Mani Shukla, Tapadhir Das
Split Federated Learning (SplitFed) has emerged as a decentralized method of training ML models that enables multiple healthcare parties to collaboratively share models without sharing their raw data. This method, however, is vulnerable to label inference attacks, which can compromise patient privacy. Previous research…
Matthew Witman, Peter Schindler
Machine learning (ML) models in the materials sciences that are validated by overly simplistic cross-validation (CV) protocols can yield biased performance estimates for downstream modeling or materials screening tasks. This can be particularly counterproductive for applications where the time and cost of failed…
Guido Pauli, G. Joseph Ray, Anton Bzhelyansky, Birgit Jaki + 18 more
Classical 1D 1H NMR spectra are prototypic for NMR spectroscopy in that they represent a wealth of chemical information encoded into convoluted graphs or patterns that contain complex features (aka multiplets), even for seemingly simple molecules. Accordingly, the utility of NMR depends on the theoretical and visual…
Lulu Wang, Shaohua Zhou, Wenrao Fang, Wenhua Huang + 4 more
'Chao Fu' 'Changkun Liu' 'Shengdong Hu'] This paper presents an automatic piecewise (Auto-PW) extreme learning machine (ELM) method for S-parameters modeling radio-frequency (RF) power amplifiers (PAs). A strategy based on splitting regions at the changing points of concave-convex characteristics is proposed, where…
Authors not listed
The process of label selection holds significant importance in the field of electrochemical biosensors, as it directly impacts the achievement of low detection limits and a wide dynamic range. To attain these objectives, it is necessary to take into account several aspects, including low electroactive potential, high…
Maria J.A. Creighton, Alice Q. Luo, Simon M. Reader, Arne Ø. Mooers
Species are the main unit used to measure biodiversity, but different preferred diagnostic criteria can lead to very different delineations. For instance, named primate species have more than doubled in number since 1982. Such increases have been attributed to a shift away from the ‘biological species concept’ (BSC) in…
Trevor Gokey, David L. Mobley
Molecular mechanics force fields require a chemical perception model to assign parameters to molecules. A recent advancement in force fields is the use of the SMARTS substructure query language as the perception model. Although it is straightforward to write SMARTS patterns to define new force field parameters, it is…
Mohaddeseh Rahbaran, Ehsan Razeghian, Marwah Suliman Maashi, Abduladheem Turki Jalil + 7 more
'Abduladheem Turki Jalil' 'Gunawan Widjaja' 'Lakshmi Thangavelu' 'Mariya Yurievna Kuznetsova' 'Pourya Nasirmoghadas' 'Farid Heidari' 'Faroogh Marofi' 'Mostafa Jarahian'] In the article titled “Cloning and Embryo Splitting in Mammalians: Brief History, Methods, and Achievements” [1], the incorrect affiliation was listed…