PGB: One-Shot Pruning for BERT via Weight Grouping and Permutation

Lim, Hyemin; Lee, Jaeyeon; Choi, Dong-Wan

Computer Science > Computation and Language

arXiv:2502.03984 (cs)

[Submitted on 6 Feb 2025]

Title:PGB: One-Shot Pruning for BERT via Weight Grouping and Permutation

Authors:Hyemin Lim, Jaeyeon Lee, Dong-Wan Choi

View PDF HTML (experimental)

Abstract:Large pretrained language models such as BERT suffer from slow inference and high memory usage, due to their huge size. Recent approaches to compressing BERT rely on iterative pruning and knowledge distillation, which, however, are often too complicated and computationally intensive. This paper proposes a novel semi-structured one-shot pruning method for BERT, called $\textit{Permutation and Grouping for BERT}$ (PGB), which achieves high compression efficiency and sparsity while preserving accuracy. To this end, PGB identifies important groups of individual weights by permutation and prunes all other weights as a structure in both multi-head attention and feed-forward layers. Furthermore, if no important group is formed in a particular layer, PGB drops the entire layer to produce an even more compact model. Our experimental results on BERT$_{\text{BASE}}$ demonstrate that PGB outperforms the state-of-the-art structured pruning methods in terms of computational cost and accuracy preservation.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2502.03984 [cs.CL]
	(or arXiv:2502.03984v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2502.03984

Submission history

From: Hyemin Lim [view email]
[v1] Thu, 6 Feb 2025 11:34:41 UTC (433 KB)

Computer Science > Computation and Language

Title:PGB: One-Shot Pruning for BERT via Weight Grouping and Permutation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:PGB: One-Shot Pruning for BERT via Weight Grouping and Permutation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators