29 papers · ranked by Valyu relevance
Mohith Damarapati, Inavamsi Enaganti, Alfred Ajay Aureate Rajakumar
When people learn mathematical patterns or sequences, they are able to identify the concepts (or rules) underlying those patterns. Having learned the underlying concepts, humans are also able to generalize those concepts to other numbers, so far as to even identify previously unseen combinations of those rules. Current…
Cedric Foucault, Florent Meyniel
From decision making to perception to language, predicting what is coming next is crucial. It is also challenging in stochastic, changing, and structured environments; yet the brain makes accurate predictions in many situations. What computational architecture could enable this feat? Bayesian inference makes optimal…
Xueliang Leon Liu
As high-throughput biological sequencing becomes faster and cheaper, the need to extract useful information from sequencing becomes ever more paramount, often limited by low-throughput experimental characterizations. For proteins, accurate prediction of their functions directly from their primary amino-acid sequences…
Noémi Éltető, Dezső Nemeth, Karolina Janacsek, Peter Dayan
Humans can implicitly learn complex perceptuo-motor skills over the course of large numbers of trials. This likely depends on our becoming better able to take advantage of ever richer and temporally deeper predictive relationships in the environment. Here, we offer a novel characterization of this process, fitting a…
Boris Rubinstein
Recent researches demonstrate that prediction of time series by predictive recurrent neural networks based on the noisy input generates a smooth anticipated trajectory. We examine influence of the noise component in both the training data sets and the input sequences on network prediction quality. We propose and…
Wu Yan, Li Tan, Li Meng-Shan, Sheng Sheng + 3 more
'Daniel Fischer'] Biological sequence data mining is hot spot in bioinformatics. A biological sequence can be regarded as a set of characters. Time series is similar to biological sequences in terms of both representation and mechanism. Therefore, in the article, biological sequences are represented with time series to…
Moritz Wolter, Jüergen Gall, Angela Yao
Fourier methods have a long and proven track record as an excellent tool in data processing. As memory and computational constraints gain importance in embedded and mobile applications, we propose to combine Fourier methods and recurrent neural network architectures. The short-time Fourier transform allows us to…
Mustafa E. Aydın, Arda Fazla, Süleyman S. Kozat
—We investigate nonlinear prediction/regression in an online setting and introduce a hybrid model that effectively mitigates, via a joint mechanism through a state space formulation, the need for domain-specific feature engineering issues of conventional nonlinear prediction models and achieves an efficient mix of…
Munazah Andrabi, Andrew Paul Hutchins, Diego Miranda-Saavedra, Hidetoshi Kono + 3 more
'Hidetoshi Kono' 'Ruth Nussinov' 'Kenji Mizuguchi' 'Shandar Ahmad'] DNA shape is emerging as an important determinant of transcription factor binding beyond just the DNA sequence. The only tool for large scale DNA shape estimates, DNAshape was derived from Monte-Carlo simulations and predicts four broad and static DNA…
Vanessa Ferdinand, Amy Yu, Sarah Marzen
Organisms can solve complex tasks despite having limited cognitive resources when those resources are used optimally. Doing so optimally makes an organism “resource-rational”. In this paper, we show for the first time that humans are resource-rational at prediction. In a novel sequence learning experiment, participants…
Qingzhen Hou, Bas Stringer, Katharina Waury, Henriette Capel + 4 more
Antibodies play an important role in clinical research and biotechnology, with their specificity determined by the interaction with the antigen’s epitope region, as a special type of protein-protein interaction (PPI) interface. The ubiquitous availability of sequence data, allows us to predicting epitopes from sequence…
Jun Wang, Huiwen Zheng, Yang Yang, Wanyue Xiao + 1 more
DNA-binding proteins (DBPs) play vital roles in all aspects of genetic activities. However, the identification of DBPs by using wet-lab experimental approaches is often time-consuming and laborious. In this study, we develop a novel computational method, called PredDBP-Stack, to predict DBPs solely based on protein…
Yuxin Shen, Grzegorz Kudla, Diego A. Oyarzún
The growing demand for biological products drives many efforts to maximize expression of heterologous proteins. Advances in high-throughput sequencing can produce data suitable for building sequence-to-expression models with machine learning. The most accurate models have been trained on one-hot encodings, a…
Ruofeng Wen, Kari Torkkola, N. Balakrishnan, Dhruv Madeka
We propose a framework for general probabilistic multi-step time series regression. Specifically, we exploit the expressiveness and temporal nature of Sequence-to-Sequence Neural Networks (e.g. recurrent and convolutional structures), the nonparametric nature of Quantile Regression and the efficiency of Direct…
Steven Elsworth, Stefan Güttel
—Machine learning methods trained on raw numerical time series data exhibit fundamental limitations such as a high sensitivity to the hyper parameters and even to the initialization of random weights. A combination of a recurrent neural network with a dimension-reducing symbolic representation is proposed and applied…
Chunyan Ao, Shihu Jiao, Yansu Wang, Liang Yu + 1 more
With the rapid development of biotechnology, the number of biological sequences has grown exponentially. The continuous expansion of biological sequence data promotes the application of machine learning in biological sequences to construct predictive models for mining biological sequence information. There are many…
Wei Chen, Pengmian Feng, Hui Yang, Hui Ding + 2 more
'Kuo-Chen Chou'] Catalyzed by adenosine deaminase (ADAR), the adenosine to inosine (A-to-I) editing in RNA is not only involved in various important biological processes, but also closely associated with a series of major diseases. Therefore, knowledge about the A-to-I editing sites in RNA is crucially important for…
Wout Bittremieux, Varun Ananth, William E. Fondrie, Carlo Melendez + 5 more
Protein tandem mass spectrometry data is most often interpreted by matching observed mass spectra to a protein database derived from the reference genome of the sample being analyzed. In many application domains, however, a relevant protein database is unavailable or incomplete, and in such settings de novo sequencing…
Asma Ehsan, Khalid Mahmood, Yaser Daanial Khan, Sher Afzal Khan + 1 more
'Kuo-Chen Chou'] The molecular structure of macromolecules in living cells is ambiguous unless we classify them in a scientific manner. Signal peptides are of vital importance in determining the behavior of newly formed proteins towards their destined path in cellular and extracellular location in both eukaryotes and…
Varun Maher, Daniel Martin, David Spetzler, Zhan-Gong Zhao + 2 more
Systematic Evolution of Ligands through Exponential Enrichment (SELEX) was used as a model system to explore the evolution of DNA sequence information and function during enrichment of molecular recognition to a series of related target molecules. Using a Natural Language Processing (NLP) based approach, a model was…
Natsuki Iwano, Tatsuo Adachi, Kazuteru Aoki, Yoshikazu Nakamura + 1 more
'Michiaki Hamada'] Nucleic acid aptamers are generated by an in vitro molecular evolution method known as systematic evolution of ligands by exponential enrichment (SELEX). Various candidates are limited by actual sequencing data from an experiment. Here we developed RaptGen, which is a variational autoencoder for in…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Fuzhan Rahmanian, Robert M. Lee, Dominik Linzner, Kathrin Michel + 4 more
Predicting and monitoring battery life early and across chemistries is a significant challenge due to the plethora of degradation paths, form factors, and electrochemical testing protocols. Existing models typically translate poorly across different electrode, electrolyte, and additive materials, mostly require a fixed…
Authors not listed
Sequence is the critical determinant of macromolecular function, yet current polymer design approaches often optimize monomer composition and ratios while ignoring sequence. This creates poorly defined design spaces for active learning that miss the vast combinatorial landscape of sequence possibilities. We introduce…
Manie Tadayon, Greg Pottie
Time series and sequential data have gained significant attention recently since many real-world processes in various domains such as finance, education, biology, and engineering can be modeled as time series. Although many algorithms and methods such as the Kalman filter, hidden Markov model, and long short term…
Likun Wang, Hao Pei, Tong Zhu
DNA reactions are crucial in biology, synthetic biology, and DNA computing. Accurate prediction of thermodynamic and kinetic parameters is vital for understanding molecular interactions and designing functional DNA-based systems. Existing models have limitations due to simplifications and approximations that may…
Amit K. Chattopadhyay, Diar Nasiev, Darren R. Flower
Motivation: Within bioinformatics, the textual alignment of amino acid sequences has long dominated the determination of similarity between proteins, with all that implies for shared structure, function and evolutionary descent. Despite the relative success of modern-day sequence alignment algorithms, so-called…
Alex Lee, Joshua Rackers, William Bricker
One of the fundamental limitations of accurately modeling biomolecules like DNA is the inability to perform quantum chemistry calculations on large molecular structures. We present a machine learning model based on an equivariant Euclidean Neural Network framework to obtain quantum-accurate electron densities for…
Joseph Redshaw, Darren Ting, Alex Brown, Jonathan Hirst + 1 more
Antimicrobial peptides (AMPs) represent a potential solution to the growing problem of antimicrobial resistance, yet their identification through wet-lab experiments is a costly and timeconsuming process. Accurate computational predictions would allow rapid in silico screening of candidate AMPs, thereby accelerating…