27 papers · ranked by Valyu relevance
Kevin Zhang, Neha Patki, Kalyan Veeramachaneni
The goal of this paper is to describe a system for generating synthetic sequential data within the Synthetic data vault. To achieve this, we present the Sequential model currently in SDV, an end-to-end framework that builds a generative model for multi-sequence, real-world data. This includes a novel neural…
Haoyang Cao, Minshuo Chen, Yinbin Han, Renyuan Xu
Generating realistic synthetic sequential data is critical in real-world applications across operations research, finance, healthcare, energy systems, and scientific computing, where time-indexed observations are used for prediction, simulation, risk assessment, and data-driven decision-making. While diffusion models…
Zhimian Hao, Chonghuan Zhang, Alexei Lapkin
We propose a workflow for reduction in the time required for data generation during generation of statistical digital twins. This methodology is particularly relevant for real-world engineering problems when data generation is expensive. A prerequisite for building surrogates is sufficient input/output data, whereas…
Paul Tiwald, Ivona Krchova, Andrey Sidorenko, Mariana Vargas-Vieyra + 2 more
Generating High-Fidelity Synthetic Data Authors: ['Paul Tiwald' 'Ivona Krchova' 'Andrey Sidorenko' 'Mariana Vargas-Vieyra' 'Mario Scriminaci' 'Michael Platzer'] Synthetic data generation for tabular datasets must balance fidelity, efficiency, and versatility to meet the demands of real-world applications. We introduce…
Keyi Li, Sen Yang, Travis M. Sullivan, Randall S. Burd + 1 more
Process data with confidential information cannot be shared directly in public, which hinders the research in process data mining and analytics. Data encryption methods have been studied to protect the data, but they still may be decrypted, which leads to individual identification. We experimented with different models…
Ran He, Jie Cao, Tieniu Tan
Generative artificial intelligence (GAI) has recently achieved significant success, enabling anyone to create texts, images, videos and even computer codes while providing insights that might not be possible with traditional tools. To stimulate future research, this work provides a brief summary of the ongoing and…
Javier Geijo-Fernández, Alexander Pfundner, Carlos A Garcia-Perez
Microbial community profiling relies on comprehensive reference databases, yet full-length 16S rRNA amplicons remain sparse for many bacterial taxa. We present SGenerator, a neural network-based data augmentation method that generates biologically informative, full-length (1500 bp) 16S rRNA sequences for…
Sweta Kumari, C Vigneswaran, V. Srinivasa Chakravarthy
Sequential decision making tasks that require information integration over extended durations of time are challenging for several reasons including the problem of vanishing gradients, long training times and significant memory requirements. To this end we propose a neuron model fashioned after the JK flip-flops in…
Katariina Perkonoja, Kari Auranen, Joni Virta
The rapid growth in data availability has facilitated research and development, yet not all industries have benefited equally due to legal and privacy constraints. The healthcare sector faces significant challenges in utilizing patient data because of concerns about data security and confidentiality. To address this…
MohammadReza EskandariNasab, Shah Muhammad Hamdi, Soukaïna Filali Boubrahimi
Learning Authors: ['MohammadReza EskandariNasab' 'Shah Muhammad Hamdi' 'Soukaïna Filali Boubrahimi'] Abstract—Current Generative Adversarial Network (GAN) based approaches for time series generation face challenges such as suboptimal convergence, information loss in embedding spaces, and instability. To overcome these…
Yuxuan Li, Chenang Liu
Machine learning (ML) has been extensively adopted for the online sensing-based monitoring in advanced manufacturing systems. However, the sensor data collected under abnormal states are usually insufficient, leading to significant data imbalanced issue for supervised machine learning. A common solution is to…
Shuai Chen, Zhoujun Li
The research on intent-enhanced sequential recommendation algorithms focuses on how to better mine dynamic user intent based on user behavior data for sequential recommendation tasks. Various data augmentation methods are widely applied in current sequential recommendation algorithms, effectively enhancing the ability…
Mst. Fahmida Sultana Naznin, Swarup Sidhartho Mondol, Adnan Ibney Faruq, Ahmed Mahir Sultan Rumi + 2 more
The fast-growing amount of data needs reliable and long-lasting storage solutions. DNA has emerged as a promising medium due to its high information density and long-term stability. However, DNA storage is a complex process where each stage introduces noise and errors, including synthesis errors, storage decay, and…
Debapriya Hazra, Mi-Ryung Kim, Yung-Cheol Byun, Luca Agnelli
Nucleic acids are the basic units of deoxyribonucleic acid (DNA) sequencing. Every organism demonstrates different DNA sequences with specific nucleotides. It reveals the genetic information carried by a particular DNA segment. Nucleic acid sequencing expresses the evolutionary changes among organisms and…
Andrew B. Lehr, Arvind Kumar, Christian Tetzlaff
Neural activity in the brain traces sequential trajectories on low dimensional subspaces. For flexible behavior, these neural subspaces must be manipulated and reoriented within short timescales of tens of milliseconds. Using mathematical analysis and simulation of a recurrently connected neural circuit for sequence…
Nicola Mulberry, Tanja Stadler
A combination of recent advancements in molecular recording devices and sequencing technologies has made it possible to generate lineage tracing data on the order of thousands of cells. Dynamic lineage recorders are able to generate random, heritable mutations which accumulate continuously on the timescale of…
Sophie Seidel, Antoine Zwaans, Samuel Regalado, Junhong Choi + 2 more
CRISPR-based lineage tracing offers a promising avenue to decipher single cell lineage trees, especially in organisms that are challenging for microscopy. A recent advancement in this domain is lineage tracing based on sequential genome editing, which not only records genetic edits but also the order in which they…
Sohom Ghosh, Shefali Yadav, Xin Wang, Bibhash Chakrabarty + 1 more
'Serdar Kadıoğlu'] Sequential pattern mining remains a challenging task due to the large number of redundant candidate patterns and the exponential search space. In addition, further analysis is still required to map extracted patterns to different outcomes. In this paper, we introduce a pattern mining framework that…
D. Lin, Y. F. Ji, J. A. A. McArt, J. Li
While global medical research is poised to benefit from the rapid advance of artificial intelligence (AI) technologies, veterinary medicine research often faces significant limitations due to data scarcity and availability issues. To address this issue, we proposed a generative modeling framework, SynLS, for generating…
Isidoro J. Casanova, Manuel Campos, Jose M. Juarez, Antonio Gomariz + 3 more
'Bernardo Canovas-Segura' 'Marta Lorente-Ros' 'Jose A. Lorente'] Background Pattern mining techniques are helpful tools when extracting new knowledge in real practice, but the overwhelming number of patterns is still a limiting factor in the health-care domain. Current efforts concerning the definition of measures of…
Guillaume Bernas, Mariette Ouellet, Andréa Barrios, Hélène Jamann + 3 more
The discovery of the CRISPR-Cas9 system and its applicability in mammalian embryos has revolutionized the way we generate genetically engineered animal models. To date, models harbouring conditional alleles (i.e.: two loxP sites flanking an exon or a critical DNA sequence of interest) remain the most challenging to…
Qi Zhang, Chang Liu, Stephen Wu, Ryo Yoshida
In the last few years, de novo molecular design using machine learning has made great technical progress but its practical deployment has not been as successful. This is mostly owing to the cost and technical difficulty of synthesizing such computationally designed molecules. To overcome such barriers, various methods…
Authors not listed
Sequence is the critical determinant of macromolecular function, yet current polymer design approaches often optimize monomer composition and ratios while ignoring sequence. This creates poorly defined design spaces for active learning that miss the vast combinatorial landscape of sequence possibilities. We introduce…
Authors not listed
This research presents a novel approach to obstacle detection during navigation using a combination of Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks. The primary objective is to generate accurate image captions that describe the content of images, which is crucial for applications such…
Alexander Grote, Anuja Hariharan, Christof Weinhardt
Introduction The analysis of discrete sequential data, such as event logs and customer clickstreams, is often challenged by the vast number of possible sequential patterns. This complexity makes it difficult to identify meaningful sequences and derive actionable insights. Methods We propose a novel feature selection…
Kevin Kawchak
Large Multimodal Models (LMMs) possess the ability to analyze chemical spectra of an organic compound using state of the art conversational AI. These outputs can then be chained together and introduced as a text input for other LLMs or LMMs to predict the compound name. Here, a challenging 15 carbon molecule problem…
Wenyu Zhang, Mason Guy, Jerrica Yang, Lucy Hao + 5 more
Large Language Models (LLMs) have revolutionized numerous industries as well as accelerated scientific research. However, their application in planning and conducting experimental science, has been limited. In this study, we introduce an adaptable prompt-set with GPT-4, converting literature experimental procedures…