13 papers · ranked by Valyu relevance
Ander Cejudo, Yone Tellechea, Amaia Calvo, Aitor Almeida + 3 more
Background The increasing use of real-time health data from wearable devices and self-reported questionnaires offers significant opportunities for preventive care in aging populations. However, current health data platforms often lack built-in mechanisms for data and model traceability, version control, and coordinated…
Julien Barnier, Cassandra Bompard, Aurélie Siberchicot, Vincent Navratil + 2 more
The need to visualize data associated with NCBI Taxonomy Identifiers (taxids) is growing in various biological fields ranging from comparative genomics to metagenomics and metabarcoding, and even for outreach. No tool today allows to visualize such data while still keeping the full vision of the whole taxonomy…
G.P. Saggese, Paul Smith
| 1. | Introduction | 1 | |----|--------------------------------------------|----| | 2. | DataFlow at a Glance | 6 | | 3. | Challenges in time series machine learning | 7 | | 4. | Semantics | 11 | | 5. | DAGs | 29 | | 6. | Execution Engine | 38 | | 7. | Comparison to Related Work | 47 | | | References | 49 |
Kumar, Punit, Asif Imran, Tevfik Kosar
This paper presents a comparative performance analysis of three popular Python data manipulation libraries—Pandas, Polars, and Dask—within the context of deep learning training pipelines. The existing studies in this area do not embed the libraries inside a full deep-learning training pipeline where data loading…
Hongyi Dong, Yimeng Zhang, Yifan Chu, Hailing Zhou + 5 more
The rapid development of the Industrial Internet of Things (IIoT) generates massive heterogeneous sensor data, complicating data cleaning and normalization. Existing algorithmcentric methods often treat quality issues in isolation and lack unified governance. This paper proposes a governance-centered framework for…
Hugo Putuhena, Thomas J. Williams, Fraser Sturt, David White + 3 more
The rapid expansion of human activity in coastal and shelf seas provides impetus to investigate increased risks to ocean health and social-ecological resilience, but progress in understanding the role and relative importance of associated pressures is frustrated by a lack of a routinely available set of processed…
Sidney Shapiro, Daniel Pearson, Emiliano Sebastian Gonzalez Venegas
Spreadsheet-heavy analytical work remains common in business analytics, operations reporting, and applied research, yet workbooks that grow through formulas, manual edits, and copy-paste refresh are difficult to audit, reproduce, and govern at scale. When tabular work requires repeatability, validation, version…
Fuad Al Abir, Zongliang Yue, Ehsan Saghapour, Md Delower Hossain + 2 more
Gene Ontology encodes genes as a hierarchy, yet every enrichment visualization flattens it into a ranked list, discarding the ability to view the same process at different levels of abstraction. We present MondrianMap, a free interactive web application (https://mondrianmap.smartdrugdiscovery.org/) that organizes…
Authors not listed
Self-driving laboratories (SDLs) promise accelerated scientific discovery and product development by closing the loop between robotic execution and AI/ML-driven decision making. In practice, however, SDL orchestration remains fragmented; workflows are typically encoded as laboratory-specific scripts or bespoke…
Yang, Xunmo, Pospisil, Taylor + 4 more
This paper outlines a grammar of data analysis, as distinct from grammars of data manipulation, in which the primitives are metrics and dimensions. We describe a Python implementation of this grammar called Meterstick, which is agnostic to the underlying data source, which may be a DataFrame or a SQL database.
Yoonjin Cho, Min Seok Kim, Sangwoo Kim
Downstream use of genomic foundation models follows one of three conventions: aggregating representations across all layers (12), defaulting to the last hidden state as a fixed feature extractor (4), or picking a single intermediate layer via mechanistic-interpretability tooling (3). None of these examines which layer…
Authors not listed
Supervised deep learning has become a standard approach to deliver competitive predictive tools that allow relating the structure of molecules and their physicochemical features to properties such as binding to protein targets, performance as electronic materials, and reactivity. However, efforts to understand how…
Maxwell, David S, Darkoh, Michael + 8 more
- 1 Data Impact and Governance, The University of Texas MD Anderson Cancer Center, Houston, Texas, USA - 2 Department of Genomic Medicine, The University of Texas MD Anderson Cancer Center, Houston, Texas, USA - 3 The Institute for Data Science in Oncology, The University of Texas MD Anderson Cancer Center, Houston…