26 papers · ranked by Valyu relevance
Ibna Kowsar, Kazi Akhter, Manar D. Samad
Transfer learning of tabular data is non-trivial due to heterogeneity in the feature space across disparate domains. The limited success of traditional deep learning in tabular knowledge transfer can be advanced by leveraging large language models (LLMs). However, the efficacy of LLMs often stagnates for mixed data…
Bahrul Ilmi Nasution, Floor Eijkelboom, Mark Elliot, Richard Allmendinger + 1 more
Synthetic data generation is an important tool for privacy-preserving data sharing. While diffusion models have set recent benchmarks, flow matching (FM) offers a promising alternative. This paper presents different ways to implement flow matching for tabular data synthesis. We provide a comprehensive empirical study…
Liane Vogel, Kavitha Srinivas, Niharika D'Souza, Sola Shirai + 2 more
Tabular foundation models aim to learn universal representations of tabular data that transfer across tasks and domains, enabling applications such as table retrieval, semantic search and table-based prediction. Despite the growing number of such models, it remains unclear which approach works best in practice, as…
Yu-Rong Lin, Han-Ming Wu, Ruriko Yoshida
Tabular data is the predominant format for statistical analysis and machine learning across domains such as finance, biomedicine, and environmental sciences. However, conventional methods often face challenges when dealing with high dimensionality and complex nonlinear relationships. In contrast, deep learning models…
Xiangjian Jiang, Mingxuan Liu, Nikola Simidjievski, Tassilo Klein + 1 more
Generative modelling is a demanding test of foundation models, because it requires robust, holistic representation learning for a given data modality, rather than optimisation for a supervised prediction target alone. While recent work on tabular foundation models has achieved remarkable progress in predictive…
Leng, Yunze, Ghosh, Rohan + 2 more
Supervised learning with tabular data presents unique challenges, including low data sizes, the absence of structural cues, and heterogeneous features spanning both categorical and continuous domains. Unlike vision and language tasks, where models can exploit inductive biases in the data, tabular data lacks inherent…
Shahab Ahmad Al Maaytah, Ayman Qahmash, Zeyar Aung
Predicting whether a patient will attend a scheduled medical appointment is essential for reducing inefficiencies in healthcare systems and optimizing resource allocation. This study introduces a local, LLM-assisted pipeline that uses LLaMA 7B solely to automate semantic preprocessing such as column renaming, datatype…
Nguyen Hong Tan, Tran Manh Tuan, Pham Minh Chuan, Nguyen Duc Hoang + 3 more
Artificial Intelligence (AI) has been dramatically applied to healthcare in various tasks to support clinicians in disease diagnosis and prognosis. It has been known that accurate diagnosis must be drawn from multiple evidence, namely clinical records, X-Ray images, IoT data, etc called the multi-modal data. Despite…
Myung Jun Kim, Maximilian Schambach, Frank Essenberger, Andre Sres + 1 more
Tabular data dominate the landscape of data science, increasingly attracting innovative machine learning models and tailored benchmarks. Yet, little is known for enterprise data, where tables constitute the backbone of business operations. To broaden the benchmarking landscape for business applications, this work aims…
Carl Du Plessis, Mike Wa Nkongolo
1## Introduction In modern data-driven environments, organisations frequently need to reconcile data across heterogeneous systems, where the same real-world entities are represented using different schemas, formats, and conventions. This task, commonly referred to as data reconciliation or entity matching, is critical…
Pei Guo, Enjie Liu, Yunzhi Tan, Mochi Gao + 5 more
Table-based reasoning with large language models (LLMs), which requires reasoning based on natural language questions and structured tabular data, has gained widespread attention. However, a series of issues still constrain the application of this task. The previous approaches suffered from significant performance…
Dan Liu, Samer El Kababji, Nicholas Mitsakakis, Lisa Pilgram + 5 more
Background Small datasets are common in health research. However, the generalization performance of machine learning models is suboptimal when the training datasets are small. To address this, data augmentation is one solution and is often used for imaging and time series data, but there are no evaluations on its…
Ivana Nanevski, Maryam Mohebi, Sebastian Jäger, Karen Otte + 5 more
Machine Learning (ML) research in healthcare remains challenging as large, privacy-preserving open datasets are lacking. Synthetic data could offer a solution, but the value of synthetic data depends on diverse and conflicting criteria such as utility, fidelity, and privacy, which are rarely evaluated comprehensively.…
Luna Xingyu Li, Carissa Bleker, Sylvain Soliman, Laurence Calzone + 13 more
Logical models are widely used to study regulatory and signaling systems, yet their reuse, annotation, and exchange across tools remain challenging. Although SBML Level 3 Qualitative Models (SBML-qual) provides a standard representation, its XML-based syntax is difficult to inspect and edit directly. Here we introduce…
Matthias Schonlau, Sandra Huang, Tiancheng Yang
For empirical studies, social and health scientists give background characteristics of their sample and summarize them in the famous "Table 1". When treatment/ control groups are present, this table gives summary statistics by group to see whether the background characteristics differ by group. We propose snapshot…
Paulo Lyra, Junhao Qiu, Khai Dang, Alyssa Pybus + 6 more
Machine learning is increasingly central to biomedical research, but using machine learning well often requires substantial computational expertise and methodological care to produce high-quality results. To make machine learning tools more accessible to biomedical researchers while supporting best-practice approaches…
Giovanni Palla, Alexander Hillsley, Yang-Joon Kim, Loic A. Royer
Predicting how cells respond to genetic and chemical perturbations is a central challenge in drug discovery and functional genomics. A growing ecosystem of specialized single-cell foundation models has been developed to address this problem, yet their practical advantage over domain-agnostic approaches remains unclear.…
Gabriel Bianchin de Oliveira, Fahad Saeed
Computational prediction of blood-brain barrier (BBB) permeability has emerged as a vital alternative to traditional experimental assays, which are often resource-intensive and low-throughput to meet the demands of early-stage drug discovery. While early machine learning approaches have shown promise, integration of…
Authors not listed
Solubility is a crucial property of each organic compound, impacting its potential applications in synthetic chemistry, materials science and drug design. Moreover, in technological processes mixtures of solvents are often utilized, making the solubility assessment more complicated. Predicting solubility values in…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Jasmin Walter, Carsten Kuenne, Noah Knoppik, Philipp Goymann + 1 more
Scientific research relies on transparent dissemination of data and its associated interpretations. This task encompasses accessibility of raw data, its metadata, details concerning experimental design, along with parameters and tools employed for data interpretation. Production and handling of these data represents an…
Anirban Das, Yan Cui
Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype–phenotype effect sizes across the ancestry continuum. While tabular…
Authors not listed
The analysis of metabolic profiles using high resolution mass spectrometry (MS) data gives deep insights into the biological processes. In metabolomics, MS generates a large number of features that represent metabolites. However, identifying specific metabolites from these features can be challenging. One of the major…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Authors not listed
Large Language Models have demonstrated impressive capabilities in natural language understanding and processing. However, as AI and LLMs continue to evolve, their ability to accurately and efficiently interpret data from scientific figures and plots remains obscure. In this study, we test and evaluate the ability of…