C-SPARQL Extension for Sampling RDF Graphs Streams

Abstract : Our daily use of Internet and related technologies generates continuously large amount of heterogeneous data flows. Several RDF Stream Processing (RSP) systems have been proposed. Existing RSP systems benefit from the advantages of semantic web technologies and traditional data flow management systems. C-SPARQL, CQELS, SPARQL stream , EP-SPARQL, and Sparkwave extend the semantic query language SPARQL and are examples of those systems. Considering that the storage and processing of all these streams become expensive, we propose a solution to reduce the load while keeping data semantics, and optimizing treatments. In this paper, we propose to extend C-SPARQL for continuously generating samples on RDF graphs. We add three sampling operators (UNIFORM, RESERVOIR and CHAIN) to the C-SPARQL query syntax. These operators have been implemented into Esper, the C-SPARQL's data flow management module. The experiments show the performance of our extension in terms of execution time and preserving data semantics.
Complete list of metadatas

Cited literature [15 references]  Display  Hide  Download

https://hal.archives-ouvertes.fr/hal-01663811
Contributor : Amadou Fall Dia <>
Submitted on : Tuesday, December 19, 2017 - 1:07:49 PM
Last modification on : Thursday, February 7, 2019 - 5:55:11 PM

Identifiers

  • HAL Id : hal-01663811, version 1

Collections

Citation

Amadou Fall Dia, Zakia Kazi-Aoul, Aliou Boly, Yousra Chabchoub. C-SPARQL Extension for Sampling RDF Graphs Streams. Advances in Knowledge Discovery and Management Volume 7, 2017. ⟨hal-01663811⟩

Share

Metrics

Record views

153

Files downloads

207