26 papers · ranked by Valyu relevance
Mauro Giuffrè, Dennis L. Shung
Data-driven decision-making in modern healthcare underpins innovation and predictive analytics in public health and clinical research. Synthetic data has shown promise in finance and economics to improve risk assessment, portfolio optimization, and algorithmic trading. However, higher stakes, potential liabilities, and…
Aldren Gonzales, Guruprabha Guruswamy, Scott R. Smith, Alistair Johnson
'Alistair Johnson'] Data are central to research, public health, and in developing health information technology (IT) systems. Nevertheless, access to most data in health care is tightly controlled, which may limit innovation, development, and efficient implementation of new research, products, services, or systems.…
Theodora Kokosi, Katie Harron
1. Synthetic data are artificial data that can be used to support efficient medical and healthcare research, while minimising the need to access personal data 2. More research is needed to determine the extent to which synthetic data can be relied on for formal analysis, the cost effectiveness of generating synthetic…
Iori Thomas, Bobby Stuijfzand
Objectives Synthetic data reproduces features of a dataset without disclosing sensitive information, allowing researchers to explore data structures and test code without requiring access to real, potentially sensitive, data. We produced a low-fidelity synthetic data generation tool, accompanied by extensive…
Randi E Foraker, Sean C Yu, Aditi Gupta, Andrew P Michelson + 10 more
Synthetic data enable data generation and sharing in support of research and precision healthcare. We demonstrated the creation and analysis of synthetic derivatives and validated the findings against original data. We used traditional statistics, machine learning approaches, and spatial representations of the data.…
Boris van Breugel, Mihaela van der Schaar
Generating synthetic data through generative models is gaining interest in the ML community and beyond. In the past, synthetic data was often regarded as a means to private data release, but a surge of recent papers explore how its potential reaches much further than this—from creating more fair data to data…
Krish Parikh
Vehicles through Synthetic Data Generation Authors: ['Krish Parikh'] Abstract—Smart vehicles produce large amounts of data, much of which is sensitive and at risk of privacy breaches. As attackers increasingly exploit anonymised metadata within these datasets to profile drivers, it's important to find solutions that…
Emiliano De Cristofaro
Sharing data can often enable compelling applications and analytics. However, more often than not, valuable datasets contain information of sensitive nature, and thus sharing them can endanger the privacy of users and organizations. A possible alternative gaining momentum in both the research community and industry is…
Jingchen Hu, Claire McKay Bowen
Synthetic data generation is a powerful tool for privacy protection when considering public release of record-level data files. Initially proposed about three decades ago, it has generated significant research and application interest. To meet the pressing demand of data privacy protection in a variety of contexts, the…
Vamsi K. Potluru, Daniel Borrajo, Andrea Coletta, Niccolò Dalmasso + 16 more
'Yousef El-Laham' 'Elizabeth Fons' 'Mohsen Ghassemi' 'Sriram Gopalakrishnan' 'Vikesh Gosai' 'Eleonora Kreačić' 'Ganapathy Mani' 'Saheed Obitayo' 'Deepak Paramanand' 'Natraj Raman' 'Mikhail Solonin' 'Srijan Sood' 'Svitlana Vyetrenko' 'Haibei Zhu' 'Manuela Veloso' 'Tucker Balch'] Synthetic data has made tremendous…
Gunther Eysenbach, Fida Dankar, Debbie Rankin, Michaela Black + 4 more
'Raymond Bond' 'Jonathan Wallace' 'Maurice Mulvenna' 'Gorka Epelde'] Background The exploitation of synthetic data in health care is at an early stage. Synthetic data could unlock the potential within health care datasets that are too sensitive for release. Several synthetic data generators have been developed to date…
Jörg Drechsler, Anna‐Carolina Haensch
The idea to generate synthetic data as a tool for broadening access to sensitive microdata has been proposed for the first time three decades ago. While first applications of the idea emerged around the turn of the century, the approach really gained momentum over the last ten years, stimulated at least in parts by…
Yingzhou Lu, Huazheng Wang, Wenqi Wei
—Machine learning heavily relies on data, but realworld applications often encounter various data-related issues. These include data of poor quality, insufficient data points leading to under-fitting of machine learning models, and difficulties in data access due to concerns surrounding privacy, safety, and…
Yefeng Yuan, Yuhong Liu, Liang Cheng
Generated by Large Language Models Authors: ['Yefeng Yuan' 'Yuhong Liu' 'Liang Cheng'] The rapid advancements in generative AI and large language models (LLMs) have opened up new avenues for producing synthetic data, particularly in the realm of structured tabular formats, such as product reviews. Despite the potential…
Raffaele Marchesi, Nicolò Lazzaro, Gianluca Leonardi, Federica Rignanese + 3 more
Synthetic data generation is emerging as an approach to overcome the limitations of real-world data scarcity in omics studies, especially in precision medicine and oncology. Omics datasets, with their high dimensionality and relatively small sample sizes, often lead to overfitting, especially in deep learning models.…
Micha Landoll, Yifei Huang, Filippo Follegot, Stephan Strassmann + 3 more
This study introduces a virtual patient generation model as online tool through the generation of high-quality synthetic data, addressing challenges like privacy cocerns and limited dataset sizes. Using a Conditional Tabular Generative Adversarial Network (CTGAN), we generated synthetic data from the Electronic Health…
D. Lin, Y. F. Ji, J. A. A. McArt, J. Li
While global medical research is poised to benefit from the rapid advance of artificial intelligence (AI) technologies, veterinary medicine research often faces significant limitations due to data scarcity and availability issues. To address this issue, we proposed a generative modeling framework, SynLS, for generating…
Ted Laderas, Nicole Vasilevsky, Bjorn Pederson, Melissa Haendel + 2 more
Our goal was to create a synthetic dataset and curricular materials to assist in teaching fundamentals of translational data science. A literature review was conducted to extract current cardiovascular risk score logic, data elements, and population characteristics. Then, clinical data elements in the models were…
Timothée Poisot, Dominique Gravel, Shawn Leroux, Spencer A. Wood + 5 more
The increased availability of both open ecological data, and software to interact with it, allows to rapidly collect and integrate data over large spatial and taxonomic scales. This offers the opportunity to address macroecological questions in a cost-effective way. In this contribution, we illustrate this approach by…
Authors not listed
Liquid chromatography-tandem mass spectrometry (LC-MS/MS) is an essential analytical technique in the pharmaceutical industry, used particularly for elucidating the structure of unknown impurities in the synthesis of active pharmaceutical ingredients. However, the interpretation of mass spectra is challenging and…
Yu-Chieh Huang, Pierre Tremouilhac, Stefan Kuhn, Pei-Chi Huang + 6 more
A method for data review in chemical sciences with a focus on data for the characterization of synthetic molecules is described. As current procedures for data curation in chemistry rely almost exclusively on manual checking or peer reviewing, a (semi-)automatic procedure for the evaluation of data assigned to…
Matthias Scheffler, Stefan Bauer, Peter Benner, Tristan Bereau + 57 more
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Matthias Scheffler
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Zhimian Hao, Chonghuan Zhang, Alexei Lapkin
We propose a workflow for reduction in the time required for data generation during generation of statistical digital twins. This methodology is particularly relevant for real-world engineering problems when data generation is expensive. A prerequisite for building surrogates is sufficient input/output data, whereas…
Nayeon Kim, Hyuk Jun Yoo, Daeho Kim, Heeseung Lee + 1 more
Autonomous laboratories hold great promise for accelerating material discovery but are often restricted by static, predefined experimental constraints. We present SPACESHIP, an AIdriven framework for dynamic, constraint-free exploration of synthesizable regions in chemical parameter spaces. SPACESHIP integrates…
Yi Luo, Saientan Bag, Orysia Zaremba, Jacopo Andreo + 3 more
Despite rapid progress in the field of metal-organic frameworks (MOFs), the potential of using machine learning (ML) methods to predict MOF synthesis parameters is still untapped. Here, we show how ML can be used for rationalization and acceleration of the MOF discovery process by directly predicting the synthesis…