0% found this document useful (0 votes)
10 views7 pages

SparseAdapter: Enhancing Adapter Efficiency

The document presents SparseAdapter, a method that enhances the parameter-efficiency of adapters in pretrained language models (PLMs) by utilizing network pruning techniques. SparseAdapter can achieve comparable or superior performance to standard adapters, particularly when the sparse ratio is high, and introduces a Large-Sparse setting to further improve model capacity without exceeding parameter budgets. Experimental results demonstrate that SparseAdapter consistently outperforms traditional fine-tuning methods across various benchmarks, highlighting its effectiveness in optimizing adapter performance.

Uploaded by

Nairouz Mrabah
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views7 pages

SparseAdapter: Enhancing Adapter Efficiency

The document presents SparseAdapter, a method that enhances the parameter-efficiency of adapters in pretrained language models (PLMs) by utilizing network pruning techniques. SparseAdapter can achieve comparable or superior performance to standard adapters, particularly when the sparse ratio is high, and introduces a Large-Sparse setting to further improve model capacity without exceeding parameter budgets. Experimental results demonstrate that SparseAdapter consistently outperforms traditional fine-tuning methods across various benchmarks, highlighting its effectiveness in optimizing adapter performance.

Uploaded by

Nairouz Mrabah
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

SparseAdapter: An Easy Approach for Improving the

Parameter-Efficiency of Adapters
Shwai He1, 4∗ Liang Ding1† Daize Dong4 Miao Zhang2 Dacheng Tao1, 3
1
JD Explore Academy
2
Aalborg University 3 The university of Sydney
4
University of Electronic Science and Technology of China
[Link]@[Link], dingliang1@[Link], dzdong2019@[Link],
miaoz@[Link], [Link]@[Link]

Abstract
 LS  +RXOVE\
Adapter Tuning, which freezes the pretrained )XOO Fine-tuning

$YHUDJH6FRUH
language models (PLMs) and only fine-tunes 
LS  /R5A
arXiv:2210.04284v5 [[Link]] 10 Nov 2022

a few extra modules, has become an appeal-


S  +RXOVE\
ing efficient alternative to the full model fine- 
tuning. Although computationally efficient, LS - PfeiffHU
S  /R5$ +RXOVE\
the recent adapters often increase parame-  /R5$
S  3IHLIIHr
ters (e.g. bottleneck dimension) for match- 3IHLIIHU
ing the performance of full model fine-tuning, 
which we argue goes against their original
    
)LQHWXQHG3DUDPHWHUV 
intention. In this work, we re-examine the
parameter-efficiency of adapters through the Figure 1: Performance of different parameter-efficient
lens of network pruning (we name such plug- tuning methods on tasks from GLUE benchmark with
in concept as SparseAdapter) and find that RoBERTa-base encoder. We report the performance of
SparseAdapter can achieve comparable or bet- Houlsby Adapters, Pfeiffer Adapters, LoRA as well as
ter performance than standard adapters when that used in our plug-in method SparseAdapter, where
the sparse ratio reaches up to 80%. Based we denoted the normal sparse (in Table 1 and 4) as “S-”
on our findings, we introduce an easy but ef- and Large-Sparse (in Table 3) as “LS-” in prefix.
fective setting “Large-Sparse” to improve the
model capacity of adapters under the same
parameter budget. Experiments on five com-
petitive adapters upon three advanced PLMs PLMs (Brown et al., 2020), full fine-tuning has
show that with proper sparse method (e.g. become prohibitively expensive, limiting the appli-
SNIP) and ratio (e.g. 40%) SparseAdapter can cability of PLMs to a broader range of tasks. Hence,
consistently outperform their corresponding various parameter-efficient fine-tuning approaches
counterpart. Encouragingly, with the Large- are explored (Houlsby et al., 2019; Hu et al., 2021;
Sparse setting, we can obtain further appeal- Zhong et al., 2022), among which Adapter Tuning,
ing gains, even outperforming the full fine-
that only tunes the extra light-weighted modules
tuning by a large margin. Our code will be re-
leased at: [Link] and keeps the original PLM frozen, has attached
SparseAdapter. great attention.
Despite the progress, existing adapters match the
1 Introduction performance of full fine-tuning by increasing the
The “pretrain-finetune” paradigm has become the bottleneck dimension (Houlsby et al., 2019; Wang
de facto standard for the community of natural et al., 2022). This increases the overall parame-
language processing (NLP) (Devlin et al., 2019; ters and FLOPs, violating the original intention
Liu et al., 2019). Given a pretrained language of adapters. In this work, we turn to investigate
model (PLM), the conventional fine-tuning man- the parameter-efficiency property (the nature of
ner is tuning the entire parameters, i.e., full fine- adapters) to answer the following questions: 1
tuning, for each downstream task (Devlin et al., Whether the current adapters can be further effi-
2019). Considering the ever-increasing size of cient? 2 How can we increase the representation

capacity of adapters within the original parameter
Work was done when Shwai was interning at JD Explore
Academy.
budget?

Corresponding author To this end, we examine the parameter-efficiency
of adapters through the lens of network prun- which requires more computation cost, violating
ing (Mozer and Smolensky, 1989; Janowsky, 1989), the original intention of adapters.
which reduces the model size of neural networks by To check whether augmenting adapters by in-
pruning redundant parameters and training the rest creasing the parameters is an optimal choice, we
ones, therefore, improving the network efficiency. decide to revisit the nature of adapters, i.e., parame-
We call such pruned adapters SparseAdapter. ter efficiency, by pruning the redundant parameters.
Specifically, we systematically investigate five rep- As shown in Figure 2, randomly pruned adapters
resentative pruning methods in §2.2 to check at can achieve comparable or even better performance
what sparse ratio can the adapters maintain the than standard adapters, which indicates the exis-
effectiveness. Note that to maintain the efficient tence of redundant parameters. The comparable
nature of adapters, we prune all adapters at ini- performance could even be held under 80% spar-
tialization such that there are no extra computa- sity. Such preliminary study urges us to investigate
tional costs. We find that 1 SparseAdapter the research questions 1 and 2 . We decide to
can achieve comparable (or even better) perfor- approach them by systematically investigating the
mance than standard adapters when the sparse ratio effects of different pruning methods.
reaches up to 80%. Such encouraging performance
Figure 2: The comparison between randomly pruned
could hold even using the random pruning method
adapters and standard adapters on datasets from GLUE.
(See Figure 2) on GLUE benchmark (Wang et al.,
2018). Based on these insights, we introduce a &R/$ 053& 676% 57(
frustratingly easy setting, namely Large-Sparse,  %(57/DUJH  5R%(57D/DUJH
for SparseAdapter. We find that 2 Scaling up  
 
3HUIRUPDQFH

3HUIRUPDQFH
the bottleneck dimension of SparseAdapter with a
 
correspondingly larger sparse ratio (to ensure the  
same parameter budget, for example, 2× dimen-  
 
sion scaling with 50% sparse ratio) can effectively  
yield significant improvement by augmenting the          
6SDUVLW\  6SDUVLW\ 
model capacity.
We validate the concept of our proposed Figure 3: Schematic comparison of (a) standard adapter
SparseAdapter upon five advanced adapters, i.e., and (b) our proposed SparseAdapter.
Houlsby (Houlsby et al., 2019), Pfeiffer (Pfeif-
fer et al., 2020b), LoRA (Hu et al., 2021), MAM
Fine-tune
Adapter (He et al., 2022) and AdapterFusion (Pfeif-
fer et al., 2021), spanning both natural language
Fine-tuned
Adapter
understanding (GLUE and SQuAD) and generation Adapter

(XSum) benchmarks. We show that with proper (a) Standard Adapter Tuning.
sparsity, e.g. 40%, SparseAdapter could consis-
tently outperform their correspondingly counter-
Prune Fine-tune
part baselines. And with our Large-Sparse setting,
Initialization
SparseAdapter could even beat the full fine-tuning Fine-tuned
Adapter SparseAdapter
method significantly, e.g. 79.6 vs. 79.0 in Fig- SparseAdapter

ure 1. (b) SparseAdapter Tuning.

2.1 Pruning Adapters at Initialization


2 Methodology
As is shown in Figure 3, we intend to prune
Motivation. Adapters are bottleneck modules out redundant parameters and then fine-tune the
plugged in PLMs, with bottleneck dimension r and SparseAdapter, instead of directly tuning all pa-
model dimension d. In standard Adapter Tuning, rameters (standard Adapter Tuning). By pruning
only adapter layers are trainable while the param- adapters at initialization, we can abandon the re-
eters of original parameters are frozen, where the dundant parameters at the early stage and avoid the
number of trainable parameters determines the ca- time-consuming iterative pruning process (Fran-
pacity of adapters. The common recipe to augment kle and Carbin, 2018). Specifically, considering
the capacity is to increase the bottleneck dimension, an adapter with weights wl inserted in the layer
l ∈ {1, · · · , L}, parameters can be pruned by a bi- 3 Experiments
nary mask ml as w̃il = wil mli , where w̃il denotes
the pruned parameters, wil and mli denote the i-th Setup. Experiments were conducted on three
element of wl and ml , respectively. Given the tar- widely-used benchmarks, spanning understanding
get sparsity s, we assign scores z to all parameters and generation tasks: (1) GLUE (Wang et al.,
w and then remove redundant parameters whose 2018), containing understanding tasks like natural
scores are below the threshold zs (the s-th lowest language inference, sentiment analysis, and sen-
percentile of z). The pruning process is shown in tence similarity evaluation; (2) XSum (Narayan
Algorithm 1. et al., 2018), a summarization dataset where the
models are required to generate a short summary
for a given article; (3) SQuAD v1.1 (Rajpurkar
Algorithm 1: Pruning on Adapters
et al., 2016), a pair-wise dataset for questions and
Require: adapter paramters w, sparse ratio s
Wikipedia paragraphs where models select the an-
1: w ← Initialization(w)
swer span to the question from the paragraph.
2: z = score(w)
We use Adam (Kingma and Ba, 2014) as the
3: Compute the s-th percentile of z as zs
optimizer with β1 , β2 = 0.9, 0.98. For regulariza-
4: m ← 1 [z − zs ≥ 0]
tion, we set the weight decay as 0.1 and grid-search
5: w̃ ← m w
the learning rate from {1e-5, 2e-5, 5e-5, 1e-4, 2e-
4}, where we warm up the learning rate in the
first 10% steps (of the total training steps). For
2.2 Pruning Methods
different data scales, we grid-search the training
Random. Random pruning assigns a random epoch and batch size from {5, 10, 15, 20}, and {8,
score z ∼ Uniform(0, 1) to each parameter and 16, 32, 64}, respectively. The maximum length is
removes parameters with the lowest scores. 512 for GLUE and 384 for SQuAD. For XSum,
we set the max length of source articles to be 512
Magnitude. Magnitude pruning assigns each pa-
and the max length of the target summary to be
rameter with its magnitude z = |w| as its score and
128. For the GLUE benchmark, we follow previ-
removes parameters with the lowest scores. Magni-
ous works (Phang et al., 2018; Lee et al., 2020;
tude pruning is a standard way to prune during (or
Dodge et al., 2020) to fine-tune the pretrained lan-
after) training (Janowsky, 1989; Han et al., 2015).
guage models, e.g. BERT (Devlin et al., 2019) and
Here we follow Frankle et al. (2020) to employ
RoBERTa (Liu et al., 2019), on the downstream
magnitude pruning at the initialization stage.
training set and report results on the dev set using
Erdős-Rényi (ER). Mocanu et al. (2018); Evci the last checkpoint. For the other tasks, we report
et al. (2020) specify each layer with a random topol- the test results.
ogy in which larger layers are allocated with higher
sparsity than smaller layers. The layer-wise spar- 3.1 Results
sity is scaled proportional to 1 − nninin+n out
·nout , where SparseAdapters with Different Pruning
nin and nout refers to the number of input and out- Methods. In Table 1, we carefully compare
put neurons, respectively. SparseAdapters (with aforementioned pruning
methods: “Rand.”, “Mag.”, “ER”, “SNIP”,
SNIP. Lee et al. (2018) compute the gradients gl
“GraSP”) to the standard adapter (Houlsby et al.,
for each layer with sampled mini-batch of training
2019) (“Adapter”) on GLUE benchmark for two
data, assign scores zl = −wl gl , and remove
backbone pretrained language models BERT
the weights with the highest scores in one iteration.
(Devlin et al., 2019) and RoBERTa (Liu et al.,
The method prunes the weights with the lowest
2019), where we set the bottleneck dimension
“effect on the loss” (either positive or negative).
to 64 for all adapter layers. As shown in Ta-
GraSP. Wang et al. (2020) compute the Hessian- ble 1, all SparseAdapters achieve comparable or
gradient product hl for each layer, issue scores even better performance compared to Houlsby
zl = −wl hl , and remove the weights with the Adapter (Houlsby et al., 2019) with lower com-
highest scores in one iteration. The method re- putational overhead. Notably, SNIP (Lee et al.,
moves weights that “reduce gradient flow” while 2018) based SparseAdapter could achieve up to
preserving weights that “increase gradient flow”. 0.6% average improvement compared to standard
Table 1: Experimental results of different SparseAdapters on GLUE benchmark, where we perform pruning
with the same sparsity ratio 40% for a fair comparison. CoLA is evaluated using Matthew’s correlation. STS-B is
evaluated using Pearson’s correlation coefficient. MRPC and RTE are evaluated using accuracy. Average scores
on all tasks are underlined. The best results are bold. We report the results of full fine-tuning “Fine-Tune” as
reference.

#Param. BERT RoBERTa


Method
(Trained) CoLA MRPC STS-B RTE Avg. CoLA MRPC STS-B RTE Avg.
Fine-Tune 100% 59.4 83.1 87.2 68.3 74.5 61.8 88.0 90.8 75.2 79.0
Adapter 2.0% 59.1 82.1 86.6 66.5 73.6 61.3 87.4 90.4 74.1 78.3
w/ Rand. 58.4 82.9 86.7 66.8 73.7 61.0 87.5 90.5 73.2 78.1
w/ Mag. 58.2 82.8 86.7 66.3 73.2 60.6 87.0 90.6 73.3 77.9
w/ ER 1.2% 58.6 82.2 86.8 67.0 73.7 60.9 87.2 90.2 73.6 78.0
w/ SNIP 59.4 82.3 87.0 68.2 74.2 61.4 87.6 90.3 75.0 78.6
w/ GraSP 59.0 82.7 86.9 67.2 74.0 61.2 87.1 90.7 74.4 78.4

adapter and nearly reach the performance of full ratios for SparseAdapter (Pfeiffer et al., 2021).
fine-tuning, which is therefore left as the default We use BERT-base (Devlin et al., 2019) and
setting in the following experiments. RoBERTa-base (Liu et al., 2019) as backbones.
SparseAdapters outperform the standard adapters
Table 2: Effect on different sparse ratios and dif- when s ≤ 40% and maintained stable performance
ferent tasks. Xsum and SQuAD are evaluated with while increasing the sparse ratio. Considering the
ROUGE-2 and F1 score, respectively. We denote
trade-off between performance and parameters, we
SparseAdapter with their sparse ratios.
set 40% as the default sparse ratio in our work.
GLUE XSum SQuAD
Method Effect on Different Adapter Variants. Since
#Para. Avg. #Para. R2 #Para. F1 SparseAdapter can be plugged into any adapter
Fine-Tune 100% 79.0 100% 21.9 100% 87.8 variants, we further validate its effectiveness on
other four variants besides Houlsby Apdaters in
Adapter 2.0% 78.3 4.5% 21.6 8.8% 87.4
the above experiments, including Pfeiffer (Pfeif-
s = 0.2 1.6% 78.7 3.6% 21.6 7.0% 87.5
fer et al., 2020a), LoRA (Hu et al., 2021), Mix-
s = 0.4 1.2% 78.6 2.7% 21.8 5.3% 87.7
And-Match Adapters (“MAM”) (He et al., 2022),
s = 0.6 0.8% 78.2 1.8% 21.5 3.5% 87.4
and AdapterFusion (“AF”) (Pfeiffer et al., 2021).
s = 0.8 0.4% 77.9 0.9% 21.3 1.8% 87.0
We choose RoBERTa-base (Liu et al., 2019) as
the backbone. Following previous experiments
Effect on Different Downstream Tasks. Utiliz- on GLUE benchmark for MAM Adapters (He
ing the proper sparse method, i.e., SNIP with 40% et al., 2022), we divide the trainable parameters
sparse ratio, we validate SparseAdapter on more equally into adapters in feed-forward layers and
downstream tasks, including GLUE, XSum, and Prefix-Tuning (Li and Liang, 2021) in attention lay-
SQuAD in Table 2. We use RoBERTa-base (Liu ers. Our SparseAdpater could consistently improve
et al., 2019) for GLUE (Wang et al., 2018), BART- the accuracy with 40% fewer training parameters,
large (Lewis et al., 2020) for Xsum (Narayan et al., showing the generalization of our plug-in method.
2018) and BERT-base (Devlin et al., 2019) for Experimental results are listed in Table 4.
SQuAD v1.1 (Rajpurkar et al., 2016). For XSum
and SQuAD, The bottleneck dimension is set to Augmenting SparseAdapter with Large-Sparse
512 and 256 respectively to match the performance Setting. One strength of SparseAdapter is the po-
of full fine-tuning. Clearly, SparseAdapter outper- tential to exploit large adapter (with a correspond-
forms the standard adapters in three tasks, showing ingly large sparse ratio) to augment the adapter
the universality of SparseAdapter. capacity under the same parameter budget, namely
Large-Sparse setting. To validate our claim, we
Effect on Different Sparse Ratios. In Table scale the bottleneck dimension by {2×, 3×, 4×}
2, we investigate the effect of different sparse with correspondingly {50%, 67%, 75%} sparse ra-
Table 3: Experimental results of scaling the bottleneck dimension. (2×, 3×, 4×) of SparseAdapters using
the same amount of parameters, coined as Large-Sparse setting (“LS-” in the prefix), on GLUE benchmark.
We correspondingly increase the sparsity to ensure the same number of parameters for SparseAdapters with larger
bottleneck dimensions.

Setting BERT RoBERTa


Method
r s CoLA MRPC STS-B RTE Avg. CoLA MRPC STS-B RTE Avg.
Adapter 64 0% 59.1 82.1 86.6 66.5 73.6 61.3 87.4 90.4 74.1 78.3
128 50% 59.9 82.3 87.6 67.5 74.3 61.7 88.2 90.3 75.5 78.9
LS-Adapter 192 67% 60.1 82.7 87.7 67.7 74.6 61.8 88.7 90.4 75.3 79.1
256 75% 60.6 83.3 88.2 68.2 75.1 62.1 89.5 90.5 76.2 79.6

Table 4: Effects on other different adapter variants. performance.


“S-” means equipped with our SparseAdatper.
4 Conclusion
Method CoLA MRPC STS-B RTE Avg.
In this work, we systematically reexamine the
Pfeiffer 61.2 85.8 89.2 74.7 77.7 parameter efficiency property of adapter Tuning
S-Pfeiffer 61.1 86.0 89.3 75.2 77.9 through the lens of network pruning. Based
LoRA 62.0 87.5 88.5 74.5 78.1 on our findings, we propose a plug-in strategy,
S-LoRA 62.1 87.7 88.8 74.6 78.2 i.e., SparseAdapter, for existing adapters. Our
MAM 61.3 86.5 89.7 74.6 78.0
study empirically indicates the potential to make
S-MAM 61.5 87.6 89.8 74.3 78.3 SparseAdapter (especially with the Large-Sparse
setting) a golden standard efficient transfer learning
AF 63.1 89.7 90.9 76.0 79.9 strategy for the NLP community.
S-AF 63.3 90.0 90.8 76.4 80.1 The future work includes applying our pro-
posed SparseAdapter to more tasks (e.g. multi-
Figure 4: The comparison between SparseAdapters lingual PLM based machine translation (Zan et al.,
with Large-Sparse setting and standard adapters. 2022a,b)) and benchmarks, and investigating the
parameter efficiency of of other neural network

&R/$5R%(57DEDVH 053&5R%(57DEDVH
 models, especially for scenarios where high effi-
3HDUVRQ&RUUHODWLRQ

 ciency is required, e.g. Prompt (Lester et al., 2021).




$FFXUDF\

  Acknowledgements


 U V 
 U V 
U V  U V  We are grateful to the anonymous EMNLP review-
 U V  U V 
U V   U V  ers and the area chair for their insightful comments

          and suggestions.
7UDLQLQJ3HUFHQWDJH  7UDLQLQJ3HUFHQWDJH 

tios. As shown in Table 3, while maintaining the Limitations


same amount of parameters, with bottleneck dimen- Despite the progress we made, there still exist limi-
sion increases, Large-Sparse could consistently tations in our work. On the one hand, we only in-
gain better performance, achieving up to +1.3% vestigated some classic pruning methods and found
and +0.6% average improvements against the stan- that SNIP (Lee et al., 2018) performs the best in
dard adapter and full fine-tuning, respectively. selected criteria. However, there may exist other ad-
Besides the encouraging performance, we com- vanced pruning methods that can further improve
pare SparseAdapters with Large-Sparse setting the performance, which deserves exploration in
to standard adapters on the training convergence future work. On the other hand, since we only con-
speed in Figure 4. SparseAdapters maintain a per- sider BERT, RoBERTa, and Bart in limited tasks,
formance advantage at the same training percent- it would be valuable to consider other architecture
age and converge at least 25% ahead in the train- families (e.g. XLNET (Yang et al., 2019), ELEC-
ing process. For both tasks, Large-Sparse setting TRA (Clark et al., 2020)) and tasks (e.g. machine
contributes to a faster convergence rate and higher translation).
References Gesmundo, Mona Attariyan, and Sylvain Gelly.
2019. Parameter-efficient transfer learning for nlp.
Tom Brown, Benjamin Mann, Nick Ryder, Melanie In International Conference on Machine Learning,
Subbiah, Jared D Kaplan, Prafulla Dhariwal, pages 2790–2799. PMLR.
Arvind Neelakantan, Pranav Shyam, Girish Sastry,
Amanda Askell, Sandhini Agarwal, Ariel Herbert- Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan
Voss, Gretchen Krueger, Tom Henighan, Rewon Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang,
Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, and Weizhu Chen. 2021. Lora: Low-rank adap-
Clemens Winter, Chris Hesse, Mark Chen, Eric tation of large language models. arXiv preprint
Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, arXiv:2106.09685.
Jack Clark, Christopher Berner, Sam McCandlish,
Alec Radford, Ilya Sutskever, and Dario Amodei. Steven A Janowsky. 1989. Pruning versus clipping in
2020. Language models are few-shot learners. In neural networks. Physical Review A, 39(12):6600.
Advances in Neural Information Processing Systems,
volume 33, pages 1877–1901. Curran Associates, Diederik P Kingma and Jimmy Ba. 2014. Adam: A
Inc. method for stochastic optimization. arXiv preprint
arXiv:1412.6980.
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and
Christopher D. Manning. 2020. ELECTRA: Pre- Cheolhyoung Lee, Kyunghyun Cho, and Wanmo Kang.
training text encoders as discriminators rather than 2020. Mixout: Effective regularization to finetune
generators. In ICLR. large-scale pretrained language models. ICLR.

Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Namhoon Lee, Thalaiyasingam Ajanthan, and
Kristina Toutanova. 2019. BERT: Pre-training of Philip HS Torr. 2018. Snip: Single-shot network
deep bidirectional transformers for language under- pruning based on connection sensitivity. arXiv
standing. In Proceedings of the 2019 Conference preprint arXiv:1810.02340.
of the North American Chapter of the Association
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021.
for Computational Linguistics: Human Language
The power of scale for parameter-efficient prompt
Technologies, Volume 1 (Long and Short Papers),
tuning. In Proceedings of the 2021 Conference on
pages 4171–4186, Minneapolis, Minnesota. Associ-
Empirical Methods in Natural Language Processing,
ation for Computational Linguistics.
pages 3045–3059, Online and Punta Cana, Domini-
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali can Republic. Association for Computational Lin-
Farhadi, Hannaneh Hajishirzi, and Noah Smith. guistics.
2020. Fine-tuning pretrained language models:
Mike Lewis, Yinhan Liu, Naman Goyal, Mar-
Weight initializations, data orders, and early stop-
jan Ghazvininejad, Abdelrahman Mohamed, Omer
ping. arXiv preprint arXiv:2002.06305.
Levy, Veselin Stoyanov, and Luke Zettlemoyer.
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel 2020. BART: Denoising sequence-to-sequence pre-
Castro, and Erich Elsen. 2020. Rigging the lot- training for natural language generation, translation,
tery: Making all tickets winners. In International and comprehension. In Proceedings of the 58th An-
Conference on Machine Learning, pages 2943–2952. nual Meeting of the Association for Computational
PMLR. Linguistics, pages 7871–7880, Online. Association
for Computational Linguistics.
Jonathan Frankle and Michael Carbin. 2018. The lot-
tery ticket hypothesis: Finding sparse, trainable neu- Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning:
ral networks. arXiv preprint arXiv:1803.03635. Optimizing continuous prompts for generation. In
Proceedings of the 59th Annual Meeting of the
Jonathan Frankle, Gintare Karolina Dziugaite, Association for Computational Linguistics and the
Daniel M Roy, and Michael Carbin. 2020. Pruning 11th International Joint Conference on Natural Lan-
neural networks at initialization: Why are we miss- guage Processing (Volume 1: Long Papers), pages
ing the mark? arXiv preprint arXiv:2009.08576. 4582–4597, Online. Association for Computational
Linguistics.
Song Han, Jeff Pool, John Tran, and William Dally.
2015. Learning both weights and connections for Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Man-
efficient neural network. Advances in neural infor- dar Joshi, Danqi Chen, Omer Levy, Mike Lewis,
mation processing systems, 28. Luke Zettlemoyer, and Veselin Stoyanov. 2019.
Roberta: A robustly optimized bert pretraining ap-
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg- proach. arXiv preprint arXiv:1907.11692.
Kirkpatrick, and Graham Neubig. 2022. Towards a
unified view of parameter-efficient transfer learning. Decebal Constantin Mocanu, Elena Mocanu, Peter
In International Conference on Learning Represen- Stone, Phuong H Nguyen, Madeleine Gibescu, and
tations. Antonio Liotta. 2018. Scalable training of artificial
neural networks with adaptive sparse connectivity in-
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, spired by network science. Nature communications,
Bruna Morrone, Quentin De Laroussilhe, Andrea 9(1):1–12.
Michael C Mozer and Paul Smolensky. 1989. Us- parameter-efficient tuning of large language models.
ing relevance to reduce network size automatically. arXiv preprint arXiv:2205.12410.
Connection Science, 1(1):3–16.
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Car-
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. bonell, Russ R Salakhutdinov, and Quoc V Le. 2019.
2018. Don’t give me the details, just the summary! Xlnet: Generalized autoregressive pretraining for
topic-aware convolutional neural networks for ex- language understanding. Advances in neural infor-
treme summarization. In Proceedings of the 2018 mation processing systems, 32.
Conference on Empirical Methods in Natural Lan-
guage Processing, pages 1797–1807, Brussels, Bel- Changtong Zan, Liang Ding, Li Shen, Yu Cao, Weifeng
gium. Association for Computational Linguistics. Liu, and Dacheng Tao. 2022a. On the complemen-
tarity between pre-training and random-initialization
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, for resource-rich machine translation. In COLING.
Kyunghyun Cho, and Iryna Gurevych. 2021.
AdapterFusion: Non-destructive task composition Changtong Zan, Keqin Peng, Liang Ding, Baopu Qiu,
for transfer learning. In Proceedings of the 16th Boan Liu, Shwai He, Qingyu Lu, Zheng Zhang,
Conference of the European Chapter of the Associ- Chuang Liu, Weifeng Liu, et al. 2022b. Vega-
ation for Computational Linguistics: Main Volume, mt: The jd explore academy translation system for
pages 487–503, Online. Association for Computa- wmt22. arXiv preprint.
tional Linguistics.
Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aish- Dacheng Tao. 2022. Panda: Prompt transfer meets
warya Kamath, Ivan Vulić, Sebastian Ruder, knowledge distillation for efficient model adaptation.
Kyunghyun Cho, and Iryna Gurevych. 2020a. ArXiv, abs/2208.10160.
Adapterhub: A framework for adapting transform-
ers. In Proceedings of the 2020 Conference on Em-
pirical Methods in Natural Language Processing:
System Demonstrations, pages 46–54.
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Se-
bastian Ruder. 2020b. MAD-X: An Adapter-Based
Framework for Multi-Task Cross-Lingual Transfer.
In Proceedings of the 2020 Conference on Empirical
Methods in Natural Language Processing (EMNLP),
pages 7654–7673, Online. Association for Computa-
tional Linguistics.
Jason Phang, Thibault Févry, and Samuel R Bowman.
2018. Sentence encoders on stilts: Supplementary
training on intermediate labeled-data tasks. arXiv
preprint arXiv:1811.01088.
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and
Percy Liang. 2016. SQuAD: 100,000+ questions for
machine comprehension of text. In Proceedings of
the 2016 Conference on Empirical Methods in Natu-
ral Language Processing, pages 2383–2392, Austin,
Texas. Association for Computational Linguistics.
Alex Wang, Amanpreet Singh, Julian Michael, Fe-
lix Hill, Omer Levy, and Samuel Bowman. 2018.
GLUE: A multi-task benchmark and analysis plat-
form for natural language understanding. In Pro-
ceedings of the 2018 EMNLP Workshop Black-
boxNLP: Analyzing and Interpreting Neural Net-
works for NLP, pages 353–355, Brussels, Belgium.
Association for Computational Linguistics.
Chaoqi Wang, Guodong Zhang, and Roger Grosse.
2020. Picking winning tickets before training
by preserving gradient flow. arXiv preprint
arXiv:2002.07376.
Yaqing Wang, Subhabrata Mukherjee, Xiaodong Liu,
Jing Gao, Ahmed Hassan Awadallah, and Jian-
feng Gao. 2022. Adamix: Mixture-of-adapter for

Common questions

Powered by AI

Experimental evidence demonstrating the effectiveness of SparseAdapters includes their performance on benchmarks such as GLUE, XSum, and SQuAD. SparseAdapters were shown to achieve comparable or better results than standard adapter tuning while maintaining lower computational overhead. For instance, SparseAdapters using the SNIP method outperformed standard adapters with a 0.6% improvement in average scores on GLUE . Additionally, the use of the Large-Sparse setting enabled SparseAdapters to exceed the performance of full fine-tuning methods, affirming their capability in augmenting model capacity under constrained parameters .

SparseAdapters maintain computational efficiency by employing parameter pruning to remove redundant parameters, thereby reducing the computational load. Despite this reduction, SparseAdapters achieve similar or enhanced performance due to strategic pruning methods that preserve critical connections within the model. Additionally, when scaled with appropriate sparse ratios, SparseAdapters extend the model's capacity without increasing parameters. This strategic pruning and scaling allow them to achieve performance levels that are comparable to, or even exceed, those seen in full fine-tuning settings across multiple NLP benchmarks .

The pruning strategy is critical in the effectiveness of SparseAdapters, as it directly impacts the model's capability to maintain performance while reducing computational overhead. By selectively pruning parameters, SparseAdapters are able to focus computational resources on the most impactful parameters, thereby maximizing efficiency without compromising model performance. Different pruning strategies, such as SNIP, have been shown to allow SparseAdapters to outperform standard adapters with fewer resources across various NLP tasks, demonstrating their adaptability and robustness. This underscores the importance of choosing appropriate pruning strategies for optimal balance between sparsity and model capacity, crucial for enhancing performance across different language processing benchmarks .

SparseAdapters address the challenge of balancing parameter efficiency with model capacity and performance. Traditional adapters tend to increase model parameters to improve representation capacity, leading to greater computational costs, which contradicts their intention of parameter efficiency. SparseAdapters overcome this by leveraging pruning techniques, allowing them to maintain a high level of performance even with a reduced set of parameters, thereby preserving computational efficiency . They achieve this by systematically investigating different pruning methods to determine optimal sparsity levels, minimizing computational overhead while sustaining adaptability and performance across diverse NLP tasks .

SparseAdapters introduce network pruning to improve the efficiency of adapter tuning in neural networks. By pruning redundant parameters, SparseAdapters maintain effectiveness at a high sparse ratio, such as up to 80%, which can achieve performance comparable to or even better than standard adapters . Additionally, by scaling up the bottleneck dimension while increasing the sparse ratio accordingly, SparseAdapters can significantly enhance model capacity, outperforming standard adapters without increasing the parameter budget . Furthermore, using methods like SNIP, SparseAdapters achieve similar or better performance with reduced computational overhead, as seen on benchmarks like GLUE .

Different adapter variants benefit from the SparseAdapter framework by experiencing performance improvements with fewer training parameters. Variants such as Pfeiffer, LoRA, MAM, and AdapterFusion show consistent enhancements in performance metrics such as accuracy when equipped with SparseAdapters, as demonstrated on tasks in the GLUE benchmark . This suggests that the SparseAdapter framework is versatile and can be effectively generalized across diverse adapter architectures, implying broad applicability in natural language processing applications where efficiency and performance are crucial .

The pruning methods evaluated for SparseAdapters include random pruning ('Rand.'), magnitude-based pruning ('Mag.'), edge removal ('ER'), single-shot pruning based on connection sensitivity ('SNIP'), and the gradient signal pruning approach ('GraSP'). Among these, the SNIP pruning method demonstrated the most improvement in performance, achieving up to 0.6% average improvement over standard adapters on the GLUE benchmark .

The bottleneck dimension in SparseAdapter models is a key factor that determines the capacity of the adapter layers, which are the only parts of the model that get trained. In experiments, this dimension is scaled up to increase model capacity without expanding the overall parameter count, achieved by corresponding adjustments in the sparse ratio (increase in sparsity). For instance, in Large-Sparse settings, bottleneck dimensions were scaled by factors like 2× with 50% sparsity to ensure the same parameter budget while enhancing performance . This manipulation allows SparseAdapters to optimize the balance between capacity and efficiency .

On the GLUE benchmark, SparseAdapters compare favorably to the fine-tuning baseline by using significantly fewer parameters while achieving comparable or better performance. For example, fine-tuning uses 100% of the parameters, whereas SparseAdapters achieve over 74.2% average scores using only 1.2% of the parameters with a SNIP-based pruning method and 40% sparsity. This showcases the parameter-efficient nature of SparseAdapters while maintaining a performance that nearly matches or exceeds that of the fully fine-tuned model .

The Large-Sparse setting enhances SparseAdapters' performance by scaling the bottleneck dimension and increasing the sparse ratio appropriately to maintain the same parameter budget. This setting significantly boosts model capacity, allowing SparseAdapters to outperform not only standard adapters but also full fine-tuning methods. For example, under the Large-Sparse configuration with bottleneck dimension scaled to 256 and 75% sparsity, SparseAdapters yielded better performance metrics than standard adapters, as evidenced in tasks like those on the GLUE benchmark .

You might also like