Optimizing XML Querying using Type-based Document Projection

Véronique Benzaken 1 Giuseppe Castagna 2 Dario Colazzo 1, 3 Kim Nguyễn 1
3 OAK - Database optimizations and architectures for complex large data
LRI - Laboratoire de Recherche en Informatique, UP11 - Université Paris-Sud - Paris 11, Inria Saclay - Ile de France, CNRS - Centre National de la Recherche Scientifique : UMR8623
Abstract : XML data projection (or pruning) is a natural optimization for main memory query engines: given a query Q over a document D, the subtrees of D that are not necessary to evaluate Q are pruned, thus producing a smaller document D'; the query Q is then executed on D', hence avoiding to allocate and process nodes that will never be reached by Q. In this article, we propose a new approach, based on types, that greatly improves current solutions. Besides providing comparable or greater precision and far lesser pruning overhead, our solution -unlike current approaches- takes into account backward axes, predicates, and can be applied to multiple queries rather than just to single ones. A side contribution is a new type system for XPath able to handle backward axes. The soundness of our approach is formally proved. Furthermore, we prove that the approach is also complete (i.e., yields the best possible type-driven pruning) for a relevant class of queries and Schemas. We further validate our approach using the XMark and XPathMark benchmarks and show that pruning not only improves the main memory query engine's performances (as expected) but also those of state of the art native XML databases.
Type de document :
Article dans une revue
ACM Transactions on Database Systems, Association for Computing Machinery, 2013, 38 (1), pp.1-45
Liste complète des métadonnées

https://hal.archives-ouvertes.fr/hal-00798049
Contributeur : Kim Nguyen <>
Soumis le : jeudi 7 mars 2013 - 19:48:42
Dernière modification le : lundi 13 novembre 2017 - 16:02:02
Document(s) archivé(s) le : samedi 8 juin 2013 - 10:05:08

Fichiers

main.pdf
Fichiers produits par l'(les) auteur(s)

Identifiants

  • HAL Id : hal-00798049, version 1

Citation

Véronique Benzaken, Giuseppe Castagna, Dario Colazzo, Kim Nguyễn. Optimizing XML Querying using Type-based Document Projection. ACM Transactions on Database Systems, Association for Computing Machinery, 2013, 38 (1), pp.1-45. 〈hal-00798049〉

Partager

Métriques

Consultations de la notice

414

Téléchargements de fichiers

181