Adjacency-constrained hierarchical clustering of a band similarity matrix with application to Genomics

Motivation: Genomic data analyses such as Genome-Wide Association Studies (GWAS) or Hi-C studies are often faced with the problem of partitioning chromosomes into successive regions based on a similarity matrix of high-resolution, locus-level measurements. An intuitive way of doing this is to perform a modified Hierarchical Agglomerative Clustering (HAC), where only adjacent clusters (according to the ordering of positions within a chromosome) are allowed to be merged. A major practical drawback of this method is its quadratic time and space complexity in the number of loci, which is typically of the order of 10^4 to 10^5 for each chromosome. Results: By assuming that the similarity between physically distant objects is negligible, we propose an implementation of this adjacency-constrained HAC with quasi-linear complexity. Our illustrations on GWAS and Hi-C datasets demonstrate the relevance of this assumption, and show that this method highlights biologically meaningful signals. Thanks to its small time and memory footprint, the method can be run on a standard laptop in minutes or even seconds. Availability and Implementation: Software and sample data are available as an R package, adjclust, that can be downloaded from the Comprehensive R Archive Network (CRAN).

Mots clés

hierarchical agglomerative clustering adjacency constraint segmentation Ward's linkage similarity min heap genome-wide association studies Hi-C

Domaines

Statistiques [math.ST] Bio-informatique [q-bio.QM]

Fichier principal

ambroise_etal_AMB2019.pdf (703.99 Ko)

RLGH_heap_label+linkage_0.pdf (6.28 Ko)

RLGH_heap_label+linkage_1.pdf (6.39 Ko)

algo-chac.pdf (4.51 Ko)

algo-chac_fusion-labels.pdf (4.5 Ko)

ambroise_etal_AMB2019-suppmat.pdf (783.52 Ko)

article_comptime_dots.png (71.03 Ko)

article_di_full.png (62.2 Ko)

pencil_large_h.pdf (33.56 Ko)

pencil_small_h.pdf (33.55 Ko)

snp_comptime_p.png (121.99 Ko)

snp_firstDiff.png (103.31 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Pierre Neuvial : Connectez-vous pour contacter le contributeur

https://hal.science/hal-02006331

Soumis le : dimanche 24 novembre 2019-21:57:59

Dernière modification le : lundi 22 avril 2024-13:20:02

Archivage à long terme le : mardi 25 février 2020-13:53:17

Dates et versions

hal-02006331 , version 1 (04-02-2019)

hal-02006331 , version 2 (24-11-2019)

Identifiants

HAL Id : hal-02006331 , version 2
ARXIV : 1902.01596
DOI : 10.1186/s13015-019-0157-4
PRODINRA : 488598
PUBMED : 31807137
WOS : 000497674400001

Citer

Christophe Ambroise, Alia Dehman, Pierre Neuvial, Guillem Rigaill, Nathalie Vialaneix. Adjacency-constrained hierarchical clustering of a band similarity matrix with application to Genomics. Algorithms for Molecular Biology, 2019, 14, pp.22. ⟨10.1186/s13015-019-0157-4⟩. ⟨hal-02006331v2⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-PARIS7 UNIV-TLSE2 CNRS UNIV-EVRY INSA-TOULOUSE INRA IMT UT1-CAPITOLE USPC LAMME IPS2 UNIV-PARIS-SACLAY INSA-GROUPE INRAE UP-SCIENCES GS-ENGINEERING INRAEOCCITANIETOULOUSE UNIV-UT3 UT3-TOULOUSEINP MATHNUM MIAT BIOLOGIE_ET_AMELIORATION_DES_PLANTES

180 Consultations

319 Téléchargements