29 papers · ranked by Valyu relevance
Ander Cejudo, Yone Tellechea, Amaia Calvo, Aitor Almeida + 3 more
Background The increasing use of real-time health data from wearable devices and self-reported questionnaires offers significant opportunities for preventive care in aging populations. However, current health data platforms often lack built-in mechanisms for data and model traceability, version control, and coordinated…
Amin Beheshti, Boualem Benatallah, Hamid Reza Motahari‐Nezhad, Samira Ghodratnama + 1 more
'Samira Ghodratnama' 'Farhad Amouzgar'] Abstract. In modern enterprises, Business Processes (BPs) are realized over a mix of workflows, IT systems, Web services and direct collaborations of people. Accordingly, process data (i.e., BP execution data such as logs containing events, interaction messages and other process…
Jonas Cremerius, Mathias Weske
Understanding and improving business processes have become important success factors for organizations. Process mining has proven very successful with a variety of methods and techniques, including discovering process models based on event logs. Process mining has traditionally focussed on control flow and timing…
Gunther Eysenbach, Arcelio Benetoli, Nan Rothrock, Jennifer S Howard + 4 more
The collection and use of patient health data are central to any kind of activity in the health care system. These data may be produced during routine clinical processes or obtained directly from the patient using patient-reported outcome (PRO) measures. Although efficiency and other reasons justify data availability…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Aaron Liu, Myeongyeon Lee, Rahul Venkatesh, Jessica Bonsu + 4 more
Polymer-based semiconductors and organic electronics encapsulate a significant research thrust for informatics-driven materials development. However, device measurements are described by a complex array of design and parameter choices, many of which are sparsely reported. For example, the mobility of a polymer-based…
Jonas Cremerius, Mathias Weske
Process mining bridges the gap between process management and data science by discovering process models using event logs derived from real-world data. Besides mandatory event attributes, additional attributes can be part of an event representing domain data, such as human resources and costs. Data-enhanced process…
Fabian Haertel, Juergen Mangler, Nataliia Klievtsova, Celine Mader + 2 more
'Eugen Rigger' 'Stefanie Rinderle-Ma'] Abstract. The application and development of process mining techniques face significant challenges due to the lack of publicly available real-life event logs. One reason for companies to abstain from sharing their data are privacy and confidentiality concerns. Privacy concerns…
Yotam Evron, Arava Tsoury, Anna Zamansky, Iris Reinhartz-Berger + 1 more
Analysis Authors: ['Yotam Evron' 'Arava Tsoury' 'Anna Zamansky' 'Iris Reinhartz-Berger' 'Pnina Soffer'] Abstract. A business process model represents the expected behavior of a set of process instances (cases). The process instances may be executed in parallel and may affect each other through data or resources. In…
Noah Herrick, Susan Walsh
Processing raw genomic data for downstream applications such as imputation, association studies, and modeling requires numerous third-party bioinformatics software tools. It is highly time-consuming and resource-intensive with computational demands and storage limitations that pose significant challenges that increase…
Thiago Dantas, Vandeclécio Lira da Silva, André Faustino Fonseca, Diego A A Morais + 2 more
Data science is historically a complex field, not only because of the huge amount of data and its variety of formats, but also because the necessity of collaboration between several specialists to retrieve valuable information. In this context, we created Bio-DIA, an online software to build data science workflow…
Aichun Wu, Jitendra Yadav
Realty management relies on data from previous successful and failed purchase and utilization outcomes. The cumulative data at different stages are used to improve utilization efficacy. The vital problem is selecting data for analyzing the value incremental sequence and profitable utilization. This article proposes a…
Authors not listed
Photocatalytic overall water splitting is a promising pathway to green hydrogen but also presents unique research challenges due to the need to detect both gaseous products (H2 and O2). While gas chromatography (GC) is the most commonly employed method in this context, it faces multiple shortcomings: low time…
Maximilian Weisenseel, Julia Andersen, Samira Akili, Christian Imenkamp + 9 more
Major domains such as logistics, healthcare, and smart cities increasingly rely on sensor technologies and distributed infrastructures to monitor complex processes in real time. These developments are transforming the data landscape—from discrete, structured records stored in centralized systems to continuous…
Marta Pasquini, Marco Stenta
Background. The increasing amount of chemical reaction data makes traditional ways to navigate its corpus less effective, while the demand for novel approaches and instruments is rising. Recent data science and machine learning techniques support the development of new ways to extract value from the available reaction…
Tibor Horak, Peter Strelec, Michal Kebisek, Pavol Tanuska + 5 more
'Andrea Vaclavova' 'Arslan Musaddiq' 'Fredrik Ahlgren' 'Neda Maleki' 'Jorge L. Zapico'] Small- and medium-sized manufacturing companies must adapt their production processes more quickly. The speed with which enterprises can apply a change in the context of data integration and historicization affects their business.…
Authors not listed
Artificial intelligence (AI) is reshaping scientific research by accelerating discovery and enabling the analysis of complex data that traditional methods struggle to handle. This review examines over 310,000 journal articles and patents from the CAS Content Collection (2015–2025), with a focus on, biomedical research…
Gunther Eysenbach, Luca Toldo, Junfeng Gao, Weiqi Wang + 1 more
Background In the past few decades, medically related data collection saw a huge increase, referred to as big data. These huge datasets bring challenges in storage, processing, and analysis. In clinical medicine, big data is expected to play an important role in identifying causality of patient symptoms, in predicting…
B. Kamala
– Process mining is a new emerging research trend over the last decade which focuses on analyzing the processes using event log and data. The raising integration of information systems for the operation of business processes provides the basis for innovative data analysis approaches. Process mining has the strong…
Nishanthi Gangadharan, Richard Turner, Ray Field, Stephen G. Oliver + 2 more
'Nigel Slater' 'Duygu Dikicioglu'] There is a growing interest in mining and handling of big data, which has been rapidly accumulating in the repositories of bioprocess industries. Biopharmaceutical industries are no exception; the implementation of advanced process control strategies based on multivariate monitoring…
Joséphine Abi-Ghanem, Djomangan Adama Ouattara
The need of digital tools for integrative analysis is today important in most scientific areas. It leads to several community-driven initiatives to standardize the sharing of data and computational workflows. However, there exists no open agnostic framework to model and implement computation workflows, in particular in…
Janja Žagar, Jurij Mihelič
Advances in data science and digitalization are transforming the world, and the pharmaceutical industry is no exception. Multiple sensor-equipped manufacturing processes and laboratory analysis are the main sources of primary data, which have been utilized for the presented dataset of 1005 actual production batches of…
Ahmed BaniMustafa, Nigel Hardy
This work demonstrates the execution of a novel process model for knowledge discovery and data mining for metabolomics (MeKDDaM). It aims to illustrate MeKDDaM process model applicability using four different real-world applications and to highlight its strengths and unique features. The demonstrated applications…
Marieke Wesselkamp, Niklas Moser, Maria Kalweit, Joschka Boedecker + 1 more
Despite deep-learning being state-of-the-art for data-driven model predictions, it has not yet found frequent application in ecology. Given the low sample size typical in many environmental research fields, the default choice for the modelling of ecosystems and its functions remain process-based models. The process…
Sophia C. Tintori, Patrick Golden, Bob Goldstein
As the scientific community becomes increasingly interested in and committed to data sharing, there remains a need for tools that facilitate the querying of public data. Mining of RNA-seq datasets, for example, has value to many biomedical researchers, yet such datasets are often effectively inaccessible to…
Bohdan B. Khomtchouk, Kasra A. Vand, Thor Wahlestedt, Kelly Khomtchouk + 2 more
We propose a search engine and file retrieval system for all bioinformatics databases worldwide. PubData searches biomedical data in a user-friendly fashion similar to how PubMed searches biomedical literature. PubData is built on novel network programming, natural language processing, and artificial intelligence…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Ian L. Boyd
Comment When I think of data I think of binary or hexadecimal numbers. This betrays something of my background, but it was a surprise to me when in Defra, the UK Department of State with responsibility for food and the environment, we started to talk about data and I found that other people saw data very differently.…
Charlotte Neidiger, Tarek Saier, Kai Kühn, Victor Larignon + 12 more
In this work, a concept for an open chemistry knowledge base was developed to integrate chemical research results into a collaboratively usable platform. To achieve this, we enhanced Semantic MediaWiki (SMW) to support the collection and structured summary of chemical data contained in publications. We implemented…