21 papers · ranked by Valyu relevance
Weijieying Ren, Tianxiang Zhao, Yuqing Huang, Vasant Honavar
Based on the background discussed above, designing a stateof-the-art representation learning method for tabular data involves three fundamental elements: training data, network architectures, and learning objectives. To improve both the quantity and quality of training data, various data-related techniques, such as…
Jun-Peng Jiang, Si-Yang Liu, Hao-Run Cai, Qile Zhou + 1 more
—Tabular data, structured as rows and columns, is among the most prevalent data types in machine learning classification and regression applications. Models for learning from tabular data have continuously evolved, with Deep Neural Networks (DNNs) recently demonstrating promising results through their capability of…
Felix Wiegand, David Lähnemann, Felix Mölder, Hamdiye Uzuner + 3 more
Tabular data, often scattered across multiple tables, is the primary output of data analyses in virtually all scientific fields. Exchange and communication of tabular data is therefore a central challenge. We present Datavzrd, a tool for creating portable, visually rich, interactive reports from tabular data in any…
Felix Wiegand, David Lähnemann, Felix Mölder, Hamdiye Uzuner + 4 more
'Adrian Prinz' 'Alexander Schramm' 'Johannes Köster' 'Vivek Kumar'] Tabular data, often scattered across multiple tables, is the primary output of data analyses in virtually all scientific fields. Exchange and communication of tabular data is therefore a central challenge. We present Datavzrd, a tool for creating…
Lingxi Cui, Huan Li, Ke Chen, Lidan Shou + 1 more
of Embracing Generative AI Authors: ['Lingxi Cui' 'Huan Li' 'Ke Chen' 'Lidan Shou' 'Gang Chen'] Machine learning (ML) on tabular data is ubiquitous, yet obtaining abundant high-quality tabular data for model training remains a significant obstacle. Numerous works have focused on tabular data augmentation (TDA) to…
Chaithra Umesh, Manjunath Mahendra, Saptarshi Bej, Olaf Wolkenhauer + 1 more
'Markus Wolfien'] Recent advancements in generative approaches in AI have opened up the prospect of synthetic tabular clinical data generation. From filling in missing values in real-world data, these approaches have now advanced to creating complex multi-tables. This review explores the development of techniques…
Xiangjian Jiang, Nikola Simidjievski, Mateja Jamnik
Heterogeneous tabular data poses unique challenges in generative modelling due to its fundamentally different underlying data structure compared to homogeneous modalities, such as images and text. Although previous research has sought to adapt the successes of generative modelling in homogeneous modalities to the…
Gabriel Becker, Adrian Waddell
Tables form a central component in both exploratory data analysis and formal reporting procedures across many industries. These tables are often complex in their conceptual structure and in the computations that generate their individual cell values. We introduce both a conceptual framework and a reference…
Alok Sharma, Yosvany López, Shangru Jia, Artem Lysenko + 2 more
Tabular data analysis is a critical task in various domains, enabling us to uncover valuable insights from structured datasets. While traditional machine learning methods have been employed for feature engineering and dimensionality reduction, they often struggle to capture the intricate relationships and dependencies…
Miriam Remshard, Simon A. Queenborough
Tables and charts have long been seen as effective ways to convey data. Much attention has been focused on improving charts, following ideas of human perception and brain function. Tables can also be viewed as two-dimensional representations of data; yet, it is only fairly recently that we have begun to apply…
Nengneng Yu, Yuefan Wang, Lindsey Kathleen Olsen, Bing Zhang + 2 more
Machine learning applications in biomedicine such as omics data analysis are frequently hindered by datasets that are small, high-dimensional, and affected by batch effects across different patient cohorts. To address these challenges, we introduce TabSyM, a modular generative pipeline that synthesizes high-quality…
Gorka Fraga-González, Hester van de Wiel, Francesco Garassino, Willy Kuo + 4 more
Scientists are increasingly required by funding agencies, publishers and their institutions to produce and publish data that are Findable, Accessible, Interoperable and Reusable (FAIR). This requires curatorial activities, which are expensive in terms of both time and effort. Based on our experience of supporting a…
C. J. Lortie, Camila Vargas Poulsen, Julien Brun, Li Kui
Data support knowledge development and theory advances in ecology and evolution. We are increasingly reusing data within our teams and projects and through the global, openly archived datasets of others. Metadata can be challenging to write and interpret, but it is always crucial for reuse. The value metadata cannot be…
Jiayuan Ding, Jianhui Lin, Shiyu Jiang, Yixin Wang + 5 more
The ability to pre-train on vast amounts of data to build foundation models (FMs) has achieved remarkable success in numerous domains, including natural language processing, computer vision, and, more recently, single-cell genomics—epitomized by GeneFormer, scGPT, and scFoundation. However, as single-cell FMs begin to…
David F. Nippa, Alex T. Müller, Kenneth Atz, David B. Konrad + 3 more
Leveraging the increasing volume of chemical reaction data can enhance synthesis planning and improve suc- cess rates. However, machine learning applications for retrosynthesis planning and forward reaction prediction tools depend on having readily available, high-quality data in a structured format. While some public…
Authors not listed
Early-stage drug discovery often suffers from data scarcity and out-of-distribution (OOD) shifts, which constrain the reliability of predictive models. While deep learning has advanced representation learning from molecular and biological data, tabular modeling remains indispensable, particularly in small-sample and…
P. Travis Thompson, Hunter N.B. Moseley
In recent years, the FAIR guiding principles and the broader concept of open science has grown in importance in academic research, especially as funding entities have aggressively promoted public sharing of research products. Key to public research sharing is deposition of datasets into online data repositories; but it…
Tiqing Liu, Linda Hwang, Stephen K Burley, Carmen I Nitsche + 3 more
BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models, and computational chemistry…
Paul L. Soto
Data collection and analysis are central to scientific research, including in applied and basic behavior analysis. A substantial amount of attention has been given to how to rigorously collect and analyze data. Less attention has been paid to storing and maintaining research data, which becomes a critical step in the…
Charlotte Neidiger, Tarek Saier, Kai Kühn, Victor Larignon + 12 more
In this work, a concept for an open chemistry knowledge base was developed to integrate chemical research results into a collaboratively usable platform. To achieve this, we enhanced Semantic MediaWiki (SMW) to support the collection and structured summary of chemical data contained in publications. We implemented…
Authors not listed
The discoverability and reusability of data is critical for machine learning to drive new discovery in the chemical sciences, and the ‘FAIR Guiding Principles for scientific data management and stewardship’ provide a measurable set of guidelines that can be used to ensure the accessibility of reusable data. We…