13 papers · ranked by Valyu relevance
Carolyn Beth McNabb, Kou Murayama
Nested data structures create statistical dependence that influences the effective sample size and statistical power of a study. Several methods are available for dealing with nested data, including the summary-statistics approach and multilevel modelling (MLM). Recent publications have heralded MLM as the best method…
Jonas Hahnfeld, Jakob Blomer, Thorsten Sven Kollegger
> Abstract. High Energy Physics (HEP) experiments, for example at the Large Hadron Collider (LHC) at CERN, store data at exabyte scale in sets of files. They use a binary columnar data format by the ROOT framework, that also transparently compresses the data. In this format, cells are not necessarily atomic but they…
Igor Rozhkov, Natalia Loukachevitch
In this paper, we describe our participation in the RuTermEval competition devoted to extracting nested terms. We apply the Binder model, which was previously successfully applied to the recognition of nested named entities, to extract nested terms. We obtained the best results of term recognition in all three tracks…
Phillip P. A. Staniczenko, Debabrata Panja
Nestedness is a common property of communication, finance, trade, and ecological networks. In networks with high levels of nestedness, the link positions of low-degree nodes (those with few links) form nested subsets of the link positions of high-degree nodes (those with many links), leading to matrix representations…
Myles J Lewis, Athina Spiliopoulou, Katriona Goldmann, Costantino Pitzalis + 3 more
In summary, the nestedcv package implements fully k × l-fold nested CV while incorporating feature selection algorithms within the outer CV loops. It adds the capability of nested CV to the caret machine learning framework in widespread use. nestedcv is designed to help measure the performance and stability of…
Han Han, Tong Zhu, Xiang Zhang, Mengsong Wu + 2 more
Large Language Models Authors: ['Han Han' 'Tong Zhu' 'Xiang Zhang' 'Mengsong Wu' 'Hao Xiong' 'Wenliang Chen'] Large language models (LLMs) combined with tool learning have gained impressive results in real-world applications. During tool learning, LLMs may call multiple tools in nested orders, where the latter tool…
Xingming Liao, Nankai Lin, Haowen Li, Lianglun Cheng + 2 more
'Chong Chen'] Abstract—Nested Named Entity Recognition (NNER) focuses on addressing overlapped entity recognition. Compared to Flat Named Entity Recognition (FNER), annotated resources are scarce in the corpus for NNER. Data augmentation is an effective approach to address the insufficient annotated corpus. However…
Haiyan Gong, Jie He, Xiaotong Zhang, Lei Duan + 14 more
'Fuzhou Gong' 'Tong Liu' 'Zongguo Wang' 'Haifeng Zhao' 'Weipeng Jia' 'Lei Zhang' 'Xue Jiang' 'Wencong Chen' 'Shilong Liu' 'Hao Xiu' 'Wenjin Yang' 'Jiawang Wan'] National Materials Data Management and Service platform (NMDMS) is a materials data repository for the publication and sharing of heterogeneous materials…
Eugene Wu, Xiang Yu Tuang, Antonio Li, Vareesh Bainwala
Existing data visualization formalisms are restricted to single-table inputs, which makes existing visualization grammars like Vega-lite or ggplot2 tedious to use, have overly complex APIs, and unsound when visualization multi-table data. This paper presents the first visualization formalism to support databases as…
Panos Vassiliadis
In this paper, we provide a comprehensive rigorous modeling for multidimensional spaces with hierarchically structured dimensions in several layers of abstractions and data cubes that live in such spaces. We model cube queries and their semantics and define typical OLAP operators like Selections, Roll-Up, Drill-Down…
Connor Bernard, Gabriel Silva Santos, Jacques Deere, Roberto Rodriguez-Caro + 5 more
The ecological sciences have joined the big data revolution. However, despite exponential growth in data availability, broader interoperability amongst datasets is still needed to unlock the potential of open access. The interface of demography and functional traits is well-positioned to benefit from said…
Elizabeth Wenk, Payal Bal, David Coleman, Rachael Gallagher + 2 more
Trait databases have proliferated over the past decades, facilitating research on the ecology, evolution, and conservation of taxa across the Tree of Life. Typically, teams of independent researchers build these databases, and each must develop their own workflow and output structure. This divests research hours from…
Patrick Scheibe, Jana Schor
Scientific knowledge is increasingly captured in structured formats, such as knowledge graphs, yet it remains largely inaccessible to non-technical users. We present EcoToxFred, a prototype conversational AI agent that enables intuitive, natural language access to curated environmental toxicology data. Designed to…