23 papers · ranked by Valyu relevance
Nawsher Khan, Ibrar Yaqoob, Ibrahim Abaker Targio Hashem, Zakira Inayat + 4 more
'Zakira Inayat' 'Waleed Kamaleldin Mahmoud Ali' 'Muhammad Alam' 'Muhammad Shiraz' 'Abdullah Gani'] Big Data has gained much attention from the academia and the IT industry. In the digital and computing world, information is generated and collected at a rate that rapidly exceeds the boundary range. Currently, over 2…
Kang, Daniel
| 1 | | Introduction | 117 | | | |…
Jana Sedlakova, Paola Daniore, Andrea Horn Wintsch, Markus Wolf + 10 more
'Mina Stanikic' 'Christina Haag' 'Chloé Sieber' 'Gerold Schneider' 'Kaspar Staub' 'Dominik Alois Ettlin' 'Oliver Grübner' 'Fabio Rinaldi' 'Viktor von Wyl' '' 'Raymond Francis Sarmiento'] Digital data play an increasingly important role in advancing health research and care. However, most digital data in healthcare are…
Aaditya Prakash
--Self-Organizing Maps (SOM) are popular unsupervised artificial neural network used to reduce dimensions and visualize data. Visual interpretation from Self-Organizing Maps (SOM) has been limited due to grid approach of data representation, which makes inter-scenario analysis impossible. The paper proposes a new way…
Andry Castro, João Pinto, Luís Reino, Pavel Pipek + 1 more
The vast volume of currently available unstructured text data, such as research papers, news, and technical report data, shows great potential for ecological research. However, manual processing of such data is labour-intensive, posing a significant challenge. In this study, we aimed to assess the application of three…
Kornelia Batko, Andrzej Ślęzak
The introduction of Big Data Analytics (BDA) in healthcare will allow to use new technologies both in treatment of patients and health management. The paper aims at analyzing the possibilities of using Big Data Analytics in healthcare. The research is based on a critical analysis of the literature, as well as the…
Zihao Zhao, Z. Shen, Mingjie Tang
—Unstructured data (e.g., images, videos, PDF files, etc.) contain semantic information, for example, the facial feature of a person and the plate number of a vehicle. There could be semantic relationships among data items. For example, a person's face may appear in two irrelevant photos. Also, part of data is in…
Guy Tsafnat, François Remy, Bram Van Es, Yuxi Liu + 4 more
'Jan A Kors' 'Erik M van Mulligen' 'Peter R Rijnbeek'] Background Electronic health records (EHRs) consist of both structured data (eg, diagnostic codes) and unstructured data (eg, clinical notes). It is commonly believed that unstructured clinical narratives provide more comprehensive information. However, this…
Gerald Onwujekwe, Kweku-Muata Osei-Bryson, Nnatubemugo Ngwum
Mainstream knowledge management researchers generally agree that knowledge extracted from unstructured data and semi-structured data has become imperative for organizational strategic decision making. In this research, we develop a framework that captures and analyses unstructured data using machine learning techniques…
Oleg Stroganov, Amber Schedlbauer, Emily Lorenzen, Alex Kadhim + 3 more
The aim of this study was to make unstructured neuropathological data, located in the NeuroBioBank (NBB), follow FAIR principles, and investigate the potential of Large Language Models (LLMs) in wrangling unstructured neuropathological reports. By making the currently inconsistent and disparate data findable, our…
Chris Kimble, Giannis Milolidakis
Big data is one of the most discussed, and possibly least understood, terms in use in business today. Big data is said to offer not only unprecedented levels of business intelligence concerning the habits of consumers and rivals, but also to herald a revolution in the way in which business are organized and run.…
Valerie Restat
Data preparation, especially data cleaning, is very important to ensure data quality and to improve the output of automated decision systems. Since there is no single tool that covers all steps required, a combination of tools – namely a data preparation pipeline – is required. Such process comes with a number of…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Richard Jackson, Ismail Kartoglu, Asha Agrawal, Kenneth Lui + 11 more
Traditional health information systems are generally devised to support clinical data collection at the point of care. However, as the significance of the modern information economy expands in scope and permeates the healthcare domain, there is an increasing urgency for healthcare organisations to offer information…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Authors not listed
The exponential growth of chemical literature necessitates the development of automated tools for extracting and curating molecular information from unstructured scientific publications into open-access chemical databases. Current optical chemical structure recognition (OCSR) and named entity recognition solutions…
Avner Schlessinger, Jinfeng Liu, Burkhard Rost, Philip E Bourne
Natively unstructured or disordered protein regions may increase the functional complexity of an organism; they are particularly abundant in eukaryotes and often evade structure determination. Many computational methods predict unstructured regions by training on outliers in otherwise well-ordered structures. Here, we…
Qianxiang Ai, Fanwang Meng, Jiale Shi, Brenden Pelkie + 1 more
The popularity of data-driven approaches and machine learning (ML) techniques in the field of organic chemistry and its various subfields has increased the value of structured reaction data. Most data in chemistry is represented by unstructured text, and due to the vastness of the organic chemistry literature (papers…
Yanwei Huang, Yan Miao, Di Weng, Adam Perer + 1 more
8.1.1 Implications for managing semi-structured textual data. From our design process we learned several important lessons that may implicate semi-structured data analysis. Boundaries of data type classification. In the data science literature, data has long been classified as structured data, semistructured data, and…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
M. Zanin, D. Papo, P. A. Sousa, E. Menasalvas + 3 more
The increasing power of computer technology does not dispense with the need to extract meaningful in-formation out of data sets of ever growing size, and indeed typically exacerbates the complexity of this task. To tackle this general problem, two methods have emerged, at chronologically different times, that are now…
Authors not listed
The discoverability and reusability of data is critical for machine learning to drive new discovery in the chemical sciences, and the ‘FAIR Guiding Principles for scientific data management and stewardship’ provide a measurable set of guidelines that can be used to ensure the accessibility of reusable data. We…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…