DNNZip: Selective Layers Compression Technique in Deep Neural Network Accelerators

Habiba Lahdhiri; Maurizio Palesi; Salvatore Monteleone; Davide Patti; Giuseppe Ascia; Jordane Lorandel; Emmanuelle Bourdel; Vincenzo Catania

Communication Dans Un Congrès Année : 2020

DNNZip: Selective Layers Compression Technique in Deep Neural Network Accelerators

(1) , (2) , (3) , , , (4) , (1) ,

1
2
3
4

Habiba Lahdhiri

Fonction : Auteur
PersonId : 1062534

ASTRE [Cergy-Pontoise]

Maurizio Palesi

Fonction : Auteur

Università degli studi di Catania = University of Catania

Salvatore Monteleone

Fonction : Auteur

Equipes Traitement de l'Information et Systèmes

Davide Patti

Fonction : Auteur
PersonId : 1075156

Giuseppe Ascia

Fonction : Auteur
PersonId : 1075157

Jordane Lorandel

Fonction : Auteur
PersonId : 974427
IdRef : 196912512

Institut d'Électronique et des Technologies du numéRique

Emmanuelle Bourdel

Fonction : Auteur
PersonId : 884035

ASTRE [Cergy-Pontoise]

Vincenzo Catania

Fonction : Auteur
PersonId : 1075158

Résumé

In Deep Neural Network (DNN) accelerators, the on-chip traffic and memory traffic accounts for a relevant fraction of the inference latency and energy consumption. A major component of such traffic is due to the moving of the DNN model parameters from the main memory to the memory interface and from the latter to the processing elements (PEs) of the accelerator. In this paper, we present DNNZip, a technique aimed at compressing the model parameters of a DNN, thus resulting in significant energy and performance improvement. DNNZip implements a lossy compression whose compression ratio is tuned based on the maximum tolerated error on the model parameters provided by the user. DNNZip is assessed on several convolutional NNs and the trade-off inference energy saving vs. inference latency reduction vs. network accuracy degradation is discussed. We found that up to 64% energy saving, and up to 67% latency reduction can be obtained with a limited impact on the accuracy of the network.

Mots clés

Approximate Deep Neural Networks Deep Neural Network Accelerator Weights Compression Accuracy/Latency/Energy Trade-off

Domaines

Sciences de l'ingénieur [physics] Electronique

Fichier principal

DNNZip_DSD20.pdf (540.72 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Habiba Lahdhiri : Connectez-vous pour contacter le contributeur

https://hal.science/hal-02906973

Soumis le : dimanche 26 juillet 2020-22:39:10

Dernière modification le : jeudi 1 février 2024-10:34:04

Archivage à long terme le : mardi 1 décembre 2020-18:19:12

Dates et versions

hal-02906973 , version 1 (26-07-2020)

Identifiants

HAL Id : hal-02906973 , version 1

Citer

Habiba Lahdhiri, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Giuseppe Ascia, et al.. DNNZip: Selective Layers Compression Technique in Deep Neural Network Accelerators. Euromicro Conference on Digital System Design DSD, Aug 2020, Portorož, Slovenia. ⟨hal-02906973⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-NANTES UNIV-RENNES1 CNRS UNIV-CERGY INSA-RENNES IETR SUP_IETR ETIS ETIS-ASTRE CENTRALESUPELEC UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES INSA-GROUPE ETIS-CELL CY-TECH-SM UR1-MATH-NUM NANTES-UNIVERSITE

154 Consultations

252 Téléchargements

DNNZip: Selective Layers Compression Technique in Deep Neural Network Accelerators

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager