13 papers · ranked by Valyu relevance
Jaclyn Smith, Michael Benedikt, Miloš Nikolić, Amir Shaikhha
While large-scale distributed data processing platforms have become an attractive target for query processing, these systems are problematic for applications that deal with nested collections. Programmers are forced either to perform nontrivial translations of collection programs or to employ automated flattening…
Ulrich Matter
The rise of the programmable web offers new opportunities for the empirically driven social sciences. The access, compilation and preparation of data from the programmable web for statistical analysis can, however, involve substantial up-front costs for the practical researcher. The R-package RWebData provides a…
Jeroen Ooms
A naive realization of JSON data in R maps JSON arrays to an unnamed list, and JSON objects to a named list. However, in practice a list is an awkward, inefficient type to store and manipulate data. Most statistical applications work with (homogeneous) vectors, matrices or data frames. Therefore JSON packages in R…
M. Trassinelli
We present here Nested_fit, a Bayesian data analysis code developed for investigations of atomic spectra and other physical data. It is based on the nested sampling algorithm with the implementation of an upgraded lawn mower robot method for finding new live points. For a given data set and a chosen model, the program…
Alan Geoffrey Hall, Michel Wermelinger, Tony Hirst, Santi Phithakkitnukoon
'Santi Phithakkitnukoon'] A spreadsheet is remarkably flexible in representing various forms of structured data, but the individual cells have no knowledge of the larger structures of which they may form a part. This can hamper comprehension and increase formula replication, increasing the risk of error on both scores.…
Foto Afrati, Matthew Damigos
In this paper, we consider a tree-structured data model used in many commercial databases like Dremel, F1, JSON stores. We define identity and referential constraints within each tree-structured record. The query language is a variant of SQL and flattening is used as an evaluation mechanism. We investigate querying in…
Xingming Liao, Nankai Lin, Haowen Li, Lianglun Cheng + 2 more
'Chong Chen'] Abstract—Nested Named Entity Recognition (NNER) focuses on addressing overlapped entity recognition. Compared to Flat Named Entity Recognition (FNER), annotated resources are scarce in the corpus for NNER. Data augmentation is an effective approach to address the insufficient annotated corpus. However…
Katarína Furmanová, Samuel Gratzl, Holger Stitz, Thomas Zichner + 3 more
'Miroslava Jarešová' 'Alexander Lex' 'Marc Streit'] Most tabular data visualization techniques focus on overviews, yet many practical analysis tasks are concerned with investigating individual items of interest. At the same time, relating an item to the rest of a potentially large table is important. In this work we…
Alexandr Savinov
The plethora of existing data models and specific data modeling techniques is not only confusing but leads to complex, eclectic and inefficient designs of systems for data management and analytics. The main goal of this paper is to describe a unified approach to data modeling, called the concept-oriented model (COM)…
Jeremy Kepner, Vijay Gadepally, Hayden Jananthan, Lauren Milechin + 1 more
'Siddharth Samsi'] Abstract—The AI revolution is data driven. AI "data wrangling" is the process by which unusable data is transformed to support AI algorithm development (training) and deployment (inference). Significant time is devoted to translating diverse data representations supporting the many query and analysis…
Eugene Wu, Xiang Yu Tuang, Antonio Li, Vareesh Bainwala
Existing data visualization formalisms are restricted to single-table inputs, which makes existing visualization grammars like Vega-lite or ggplot2 tedious to use, have overly complex APIs, and unsound when visualization multi-table data. This paper presents the first visualization formalism to support databases as…
Alexandr Savinov
We describe a new logical data model, called the conceptoriented model (COM). It uses mathematical functions as firstclass constructs for data representation and data processing as opposed to using exclusively sets in conventional set-oriented models. Functions and function composition are used as primary semantic…
Devin Petersohn, Stephen Macke, Doris Xin, William Ma + 6 more
'Xiangxi Mo' 'Joseph E. Gonzalez' 'Joseph M. Hellerstein' 'Anthony D. Joseph' 'Aditya Parameswaran'] Dataframes are a popular abstraction to represent, prepare, and analyze data. Despite the remarkable success of dataframe libraries in R and Python, dataframes face performance issues even on moderately large datasets.…