Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA

Bae, Sangmin; Fisch, Adam; Harutyunyan, Hrayr; Ji, Ziwei; Kim, Seungyeon; Schuster, Tal

Computer Science > Computation and Language

arXiv:2410.20672 (cs)

[Submitted on 28 Oct 2024 (v1), last revised 28 Feb 2025 (this version, v3)]

Title:Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA

Authors:Sangmin Bae, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji, Seungyeon Kim, Tal Schuster

View PDF

Abstract:Large language models (LLMs) are expensive to deploy. Parameter sharing offers a possible path towards reducing their size and cost, but its effectiveness in modern LLMs remains fairly limited. In this work, we revisit "layer tying" as form of parameter sharing in Transformers, and introduce novel methods for converting existing LLMs into smaller "Recursive Transformers" that share parameters across layers, with minimal loss of performance. Here, our Recursive Transformers are efficiently initialized from standard pretrained Transformers, but only use a single block of unique layers that is then repeated multiple times in a loop. We further improve performance by introducing Relaxed Recursive Transformers that add flexibility to the layer tying constraint via depth-wise low-rank adaptation (LoRA) modules, yet still preserve the compactness of the overall model. We show that our recursive models (e.g., recursive Gemma 1B) outperform both similar-sized vanilla pretrained models (such as TinyLlama 1.1B and Pythia 1B) and knowledge distillation baselines -- and can even recover most of the performance of the original "full-size" model (e.g., Gemma 2B with no shared parameters). Finally, we propose Continuous Depth-wise Batching, a promising new inference paradigm enabled by the Recursive Transformer when paired with early exiting. In a theoretical analysis, we show that this has the potential to lead to significant (2-3x) gains in inference throughput.

Comments:	ICLR 2025; 49 pages, 17 figures, 19 tables
Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2410.20672 [cs.CL]
	(or arXiv:2410.20672v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2410.20672

Submission history

From: Sangmin Bae [view email]
[v1] Mon, 28 Oct 2024 02:15:45 UTC (1,497 KB)
[v2] Thu, 6 Feb 2025 03:23:11 UTC (1,517 KB)
[v3] Fri, 28 Feb 2025 16:44:24 UTC (1,481 KB)

Computer Science > Computation and Language

Title:Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators