25 papers · ranked by Valyu relevance
Haoyang Cao, Minshuo Chen, Yinbin Han, Renyuan Xu
Generating realistic synthetic sequential data is critical in real-world applications across operations research, finance, healthcare, energy systems, and scientific computing, where time-indexed observations are used for prediction, simulation, risk assessment, and data-driven decision-making. While diffusion models…
Katariina Perkonoja, Kari Auranen, Joni Virta
The rapid growth in data availability has facilitated research and development, yet not all industries have benefited equally due to legal and privacy constraints. The healthcare sector faces significant challenges in utilizing patient data because of concerns about data security and confidentiality. To address this…
Dan Liu, Samer El Kababji, Nicholas Mitsakakis, Lisa Pilgram + 5 more
Background Small datasets are common in health research. However, the generalization performance of machine learning models is suboptimal when the training datasets are small. To address this, data augmentation is one solution and is often used for imaging and time series data, but there are no evaluations on its…
Javier Geijo-Fernández, Alexander Pfundner, Carlos A Garcia-Perez
Microbial community profiling relies on comprehensive reference databases, yet full-length 16S rRNA amplicons remain sparse for many bacterial taxa. We present SGenerator, a neural network-based data augmentation method that generates biologically informative, full-length (1500 bp) 16S rRNA sequences for…
Marko Miletic, Murat Sariyar
Background Synthetic data generation (SDG) has emerged as a critical enabler for data-driven healthcare research, offering privacy-preserving alternatives to real patient data. Temporal health data - ranging from physiological signals to electronic health records (EHRs) - pose unique challenges for SDG due to their…
P Estève, Massimiliano Zanin
The generation of synthetic data is receiving increasing attention from the scientific community, thanks to its ability to solve problems like data scarcity and privacy, and is starting to find applications in air transport. We here tackle the problem of generating synthetic, yet realistic, time series of delays at…
Alessio Barboni, Massimiliano Lupo Pasini, Bishal Lakha, Edoardo Serra
Generating realistic and diverse graphs is a key problem in machine learning, with applications in molecular discovery, circuit design, cybersecurity, and beyond. However, current graph generative models remain limited by scalability and novelty. Diffusion-based methods often require costly full-adjacency operations…
Alberto Florez Prada, Darren J. Hart
Validating bioinformatics pipelines and benchmarking sequence processing algorithms requires reliable test datasets. Existing read simulation tools rely on reference genomes and empirical error profiles, lacking fine-grained control over specific targeted DNA constructs and controlled error injection.…
Chunkai Zhang, Jiarui Deng, Maohua Lyu, Wensheng Gan + 1 more
—Within the domain of data mining, one critical objective is the discovery of sequential rules with high utility. The goal is to discover sequential rules that exhibit both high utility and strong confidence, which are valuable in real-world applications. However, existing high-utility sequential rule mining algorithms…
Jaime Vale, Vanessa Freitas Silva, Maria Cristina de Almeida Silva, Fernando Silva
Time series data are essential for a wide range of applications, particularly in developing robust machine learning models. However, access to high-quality datasets is often limited due to privacy concerns, acquisition costs, and labeling challenges. Synthetic time series generation has emerged as a promising solution…
Gabrielle Josling, Ibrahima Diouf, Sankalp Khanna
Clinical and health research increasingly depends on rich, structured datasets such as electronic health records (EHRs), disease registries, and longitudinal cohort studies. These tabular datasets provide the foundation for epidemiological research, assessment of healthcare outcomes, and data-driven policy development.…
Authors not listed
The ability to generate crystal structures directly from textual descriptions marks a pivotal advancement in materials informatics and underscores the emerging role of large language models (LLMs) in inverse design. In this work, we introduce CrysText, a text-conditioned framework that generates crystal structures in…
Shujun He, Qing Sun
RNA molecules play critical roles in biology and therapeutics, with their function intimately tied to their secondary structure. Designing RNA sequences that reliably fold into desired secondary structures, especially those with complex pseudoknots, remains a fundamental challenge. Here, we present Struct2SeQ, a…
Alexandros Tzanakakis, Ioannis Mouratidis, Ilias Georgakopoulos-Soares
Large language models (LLMs) have shown remarkable success in natural language processing, prompting interest in their application to genomic sequence analysis. Genomic Language Models based on similar architectures offer a promising avenue for synthetic genome generation and characterization. However, their…
Rottenwalter, Georg, Tilly, Marcel + 4 more
Machine learning has significant potential for optimizing various industrial processes. However, data acquisition remains a major challenge as it is both time-consuming and costly. Synthetic data offers a promising solution to augment insufficient data sets and improve the robustness of machine learning models. In this…
Evan S. Hill, Ion R. Popescu, Jean Wang, William N. Frost
Behavioral sequences are essential for survival, yet the neural mechanisms that link one action to the next remain incompletely understood. In classical chain models, sequential behaviors arise through feedforward propagation of activity across distinct neuronal populations or network modules. Here, we identify a…
Alexander Grote, Anuja Hariharan, Christof Weinhardt
Introduction The analysis of discrete sequential data, such as event logs and customer clickstreams, is often challenged by the vast number of possible sequential patterns. This complexity makes it difficult to identify meaningful sequences and derive actionable insights. Methods We propose a novel feature selection…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Xutong Liu, Qixuan Zhao, Enyang Yu, Lijia Jia + 4 more
DNA-based information storage offers a promising alternative to conventional media due to its high density, long-term stability, and low energy requirements. However, its application remains hindered by synthesis costs, limited sequence length and poor scalability. DNA polymerase is a critical enzymatic tool in the DNA…
Sophie Seidel, Antoine Zwaans, Samuel Regalado, Junhong Choi + 2 more
CRISPR-based lineage tracing offers a promising avenue to decipher single cell lineage trees, especially in organisms that are challenging for microscopy. A recent advancement in this domain is lineage tracing based on sequential genome editing, which not only records genetic edits but also the order in which they…
Authors not listed
Three-dimensional molecular generative models have emerged that produce de novo molecules both unconditionally and conditionally, e.g., within protein pockets. However, steering those models in a specific region of the chemical space that satisfies a set of desired properties remains challenging. In this study, we…
Alev Kaya, İbrahim Türkoğlu, Boris Ryabko
Pseudorandom number generators (PRNGs) used in deoxyribonucleic acid (DNA)-oriented computational workflows often generate outputs in the bit domain and then map them to DNA symbols. This indirect strategy may treat DNA-specific constraints, including GC balance, homopolymer limits, and short-range sequence…
Authors not listed
Machine learning is increasingly used to predict reaction properties such as barrier heights, reaction energies, rates, or yields, as well as the underlying molecular geometries, including transition state structures. While such predictions have the potential to provide mechanistic insight for high-impact applications…
Authors not listed
Accurate prediction of chemical reaction yields remains essential for accelerating synthesis optimization, yet current machine learning models face critical limitations in capturing temporal dynamics, providing calibrated uncertainty estimates, and explicitly modeling reactant-to-product transformations. Here we…
Ibrahim Nawaz, Parv Agarwal, Thomas Heinis
DNA storage is a developing field that uses DNA to archive digital data owing to its superior information density and stability. Although DNA storage has been performed on a significant scale, challenges arise from the synthesis and sequencing of data-encoded oligonucleotides. Synthesis of DNA introduces significant…