0% found this document useful (0 votes)
21 views13 pages

LLAMAFACTORY: Efficient LLM Fine-Tuning

L LAMA FACTORY is a unified framework designed for efficient fine-tuning of over 100 large language models (LLMs) with minimal coding effort, integrating various advanced training methods. It features a user-friendly web interface, L LAMA BOARD, and supports multiple efficient fine-tuning techniques, significantly reducing resource requirements and enhancing throughput. The framework has gained substantial popularity, receiving over 25,000 stars on GitHub and facilitating the customization of LLMs for diverse applications.

Uploaded by

zhangyiang.scu
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
21 views13 pages

LLAMAFACTORY: Efficient LLM Fine-Tuning

L LAMA FACTORY is a unified framework designed for efficient fine-tuning of over 100 large language models (LLMs) with minimal coding effort, integrating various advanced training methods. It features a user-friendly web interface, L LAMA BOARD, and supports multiple efficient fine-tuning techniques, significantly reducing resource requirements and enhancing throughput. The framework has gained substantial popularity, receiving over 25,000 stars on GitHub and facilitating the customization of LLMs for diverse applications.

Uploaded by

zhangyiang.scu
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

L LAMA FACTORY: Unified Efficient Fine-Tuning of 100+ Language Models

Yaowei Zheng1 , Richong Zhang1 * , Junhao Zhang1 , Yanhan Ye1 ,


Zheyan Luo1 , Zhangchi Feng1 , Yongqiang Ma2
1
School of Computer Science and Engineering, Beihang University, China
2
School of Software and Microelectronics, Peking University, China
{hiyouga,[Link],yeyanhan,akamya,zcmuller}@[Link], zhangrc@[Link], codingma@[Link]
Demonstration video: [Link]

Abstract To address the above problems, we develop L LA -


MA FACTORY , a framework that democratizes the
Efficient fine-tuning is vital for adapting large
language models (LLMs) to downstream tasks. fine-tuning of LLMs. It unifies a variety of effi-
cient fine-tuning methods through scalable mod-
arXiv:2403.13372v4 [[Link]] 27 Jun 2024

However, it requires non-trivial efforts to imple-


ment these methods on different models. We ules, enabling the fine-tuning of hundreds of LLMs
present L LAMA FACTORY, a unified framework with minimal resources and high throughput. In
that integrates a suite of cutting-edge efficient addition, it streamlines commonly used training
training methods. It provides a solution for approaches, including generative pre-training (Rad-
flexibly customizing the fine-tuning of 100+
ford et al., 2018), supervised fine-tuning (SFT)
LLMs without the need for coding through
the built-in web UI L LAMA B OARD. We em- (Wei et al., 2022), reinforcement learning from
pirically validate the efficiency and effective- human feedback (RLHF) (Ouyang et al., 2022),
ness of our framework on language modeling and direct preference optimization (DPO) (Rafailov
and text generation tasks. It has been released et al., 2023). Users can leverage command-line
at [Link] or web interfaces to customize and fine-tune their
and received over 25,000 stars and 3,000 forks. LLMs with minimal or no coding effort.
1 Introduction L LAMA FACTORY consists of three main mod-
ules: Model Loader, Data Worker and Trainer.
Large language models (LLMs) (Zhao et al., 2023) We minimize the dependencies of these modules
present remarkable reasoning capabilities and em- on specific models and datasets, allowing the frame-
power a wide range of applications, such as ques- work to flexibly scale to hundreds of models and
tion answering (Jiang et al., 2023b), machine trans- datasets. Concretely, we first establish a model reg-
lation (Wang et al., 2023c; Jiao et al., 2023a), and istry where the Model Loader can precisely attach
information extraction (Jiao et al., 2023b). Subse- adapters to the pre-trained models by identifying
quently, a substantial number of LLMs are devel- exact layers. Then we develop a data description
oped and accessible through open-source commu- specification that allows the Data Worker to gather
nities. For example, Hugging Face’s open LLM datasets by aligning corresponding columns. Fur-
leaderboard (Beeching et al., 2023) boasts over thermore, we provide plug-and-play implementa-
5,000 models, offering convenience for individuals tions of state-of-the-art efficient fine-tuning meth-
seeking to leverage the power of LLMs. ods that enable the Trainer to activate by replacing
Fine-tuning extremely large number of parame- default ones. Our design allows these modules
ters with limited resources becomes the main chal- to be reused across different training approaches,
lenge of adapting LLM to downstream tasks. A significantly reducing the integration costs.
popular solution is efficient fine-tuning (Houlsby L LAMA FACTORY is implemented with PyTorch
et al., 2019; Hu et al., 2022; Dettmers et al., 2023), (Paszke et al., 2019) and significantly benefits from
which reduces the training cost of LLMs when open-source libraries, such as Transformers (Wolf
adapting to various tasks. However, the commu- et al., 2020), PEFT (Mangrulkar et al., 2022), and
nity contributes various methods for efficient fine- TRL (von Werra et al., 2020). On the basis, we
tuning, lacking a systematic framework that adapts provide an out-of-the-box framework with a higher
and unifies these methods to different LLMs and level of abstraction. Additionally, we build L LAM -
provides a friendly interface for user customization. A B OARD with Gradio (Abid et al., 2019), enabling
* Corresponding author fine-tuning LLMs with no coding efforts required.

1
L LAMA FACTORY FastChat LitGPT LMFlow Open-Instruct Freeze-tuning GaLore LoRA DoRA LoRA+ PiSSA
LoRA ! ! ! ! ! Mixed precision ! ! ! ! ! !
QLoRA ! ! ! ! ! Checkpointing ! ! ! ! ! !
DoRA ! Flash attention ! ! ! ! ! !
LoRA+ ! S2 attention ! ! ! ! ! !
PiSSA ! Quantization % % ! ! ! !
GaLore ! ! ! !
Unsloth % % ! ! ! !
BAdam !
Flash attention ! ! ! ! !
S2 attention ! Table 2: Compatibility between the fine-tuning tech-
Unsloth ! ! niques featured in L LAMA FACTORY.
DeepSpeed ! ! ! ! !
SFT ! ! ! ! !
RLHF ! !
DPO ! ! 3 Efficient Fine-Tuning Techniques
KTO !
ORPO !
Efficient LLM fine-tuning techniques can be di-
Table 1: Comparison of features in L LAMA FACTORY vided into two main categories: those focused on
with popular frameworks of fine-tuning LLMs. optimization and those aimed at computation. The
primary objective of efficient optimization tech-
niques is to fine-tune the parameters of LLMs while
L LAMA FACTORY is open-sourced under the
keeping costs to a minimum. On the other hand,
Apache-2.0 license. It has already garnered over
efficient computation methods seek to decrease
25,000 stars and 3,000 forks on the GitHub, and
the time or space for the required computation in
hundreds of open-source models have been built
LLMs. The methods included in L LAMA FACTORY
upon L LAMA FACTORY on the Hugging Face Hub1 .
are listed in Table 2. We will present these efficient
For example, Truong et al. (2024) build GemSUra-
fine-tuning techniques and show the substantial ef-
7B based on L LAMA FACTORY, revealing the cross-
ficiency improvement achieved by incorporating
lingual abilities of Gemma (Mesnard et al., 2024).
them into our framework in the following sections.
Furthermore, dozens of studies have utilized our
framework to explore LLMs (Wang et al., 2023a; 3.1 Efficient Optimization
Yu et al., 2023; Bhardwaj et al., 2024).
Firstly, we provide an overview of the efficient op-
2 Related Work timization techniques utilized in L LAMA FACTORY.
The freeze-tuning method (Houlsby et al., 2019) in-
With the rapid increase in demand for fine-tuning volves freezing a majority of parameters while fine-
LLMs, numerous frameworks for adapting LLMs tuning the remaining parameters in a small subset
to specific purposes have been developed. LLaMA- of decoder layers. Another method called gradient
Adapter (Zhang et al., 2024) efficiently fine-tunes low-rank projection (GaLore) (Zhao et al., 2024)
the Llama model (Touvron et al., 2023a) using a projects gradients into a lower-dimensional space,
zero-initialized attention. FastChat (Zheng et al., facilitating full-parameter learning in a memory-
2023) is a framework focused on training and evalu- efficient manner. Similarly, BAdam (Luo et al.,
ating LLMs for chat completion purposes. LitGPT 2024) leverages block coordinate descent (BCD)
(AI, 2023) provides the implementation of genera- to efficiently optimize the extensive parameters.
tive models and supports various training methods. On the contrary, the low-rank adaptation (LoRA)
Open-Instruct (Wang et al., 2023d) provides recipes (Hu et al., 2022) method freezes all pre-trained
for training instruct models. Colossal AI (Li et al., weights and introduces a pair of trainable low-rank
2023b) takes advanced parallelism strategies for matrices to the designated layer. When combined
distributed training. LMFlow (Diao et al., 2024) with quantization, this approach is referred to as
supports training LLMs for specialized domains or QLoRA (Dettmers et al., 2023), which additionally
tasks. GPT4All (Anand et al., 2023) allows LLMs reduces the memory usage. DoRA (Liu et al., 2024)
to run on consumer devices, while also providing breaks down pre-trained weights into magnitude
fine-tuning capabilities. Compared with existing and direction components and updates directional
competitive frameworks, L LAMA FACTORY sup- components for enhanced performance. LoRA+
ports a broader range of efficient fine-tuning tech- (Hayou et al., 2024) is proposed to overcome the
niques and training approaches. We list the features sub-optimality of LoRA. PiSSA (Meng et al., 2024)
among representative frameworks in Table 1. initializes adapters with the principal components
1
[Link] of the pre-trained weights for faster convergence.

2
3.2 Efficient Computation LlamaBoard
In L LAMA FACTORY, we integrate a range of tech- Experiment Configurator Training Status Monitor
niques for efficient computation. Commonly uti-
lized techniques encompass mixed precision train-
ing (Micikevicius et al., 2018) and activation check- Trainer
pointing (Chen et al., 2016). Drawing insights from
Optimization Approaches
the examination of the input-output (IO) expenses
of the attention layer, flash attention (Dao et al., LoRA PiSSA Pre-train SFT
2022) introduces a hardware-friendly approach to
GaLore BAdam RLHF DPO
enhance attention computation. S2 attention (Chen
et al., 2024b) tackles the challenge of extended con-
text with shifted sparse attention, thereby dimin-
Model Loader Data Worker
ishing memory usage in fine-tuning long-context
LLMs. Various quantization strategies (Dettmers Initialization Patches Loading Aligning
et al., 2022a; Frantar et al., 2023; Lin et al., 2023;
Egiazarian et al., 2024) decrease memory require- Quantization Adapters Merging Preprocess

ments in large language models (LLMs) by uti-


lizing lower-precision representations for weights.
Pre-Trained Models Conversational Datasets
Nevertheless, the fine-tuning of quantized models
is restricted to the adapter-based techniques like
Figure 1: The architecture of L LAMA FACTORY.
LoRA (Hu et al., 2022). Unsloth (Han and Han,
2023) incorporates Triton (Tillet et al., 2019) for
implementing the backward propagation of LoRA, enabling users to configure and launch individual
which reduces floating-point operations (FLOPs) LLM fine-tuning instance codelessly and monitor
during gradient descent and leads to expedited the training status synchronously. We illustrate the
LoRA training. relationships between these modules and the over-
L LAMA FACTORY seamlessly combines these all architecture of L LAMA FACTORY in Figure 1.
techniques into a cohesive structure to enhance the
efficiency of LLM fine-tuning. This results in a 4.1 Model Loader
reduction of the memory footprint from 18 bytes
per parameter during mixed precision training (Mi- This section initially presents the four components
cikevicius et al., 2018) or 8 bytes per parameter in in Model Loader: model initialization, model patch-
half precision training (Le Scao et al., 2022) to only ing, model quantization, and adapter attaching, fol-
0.6 bytes per parameter. Further elaboration on the lowed by a description of our approach of adapting
components in L LAMA FACTORY will be provided to a wide range of devices by handling the parame-
in the subsequent section. ter floating-point precision during fine-tuning.

4 L LAMA FACTORY Framework Model Initialization We utilize the Auto Classes


of Transformers (Wolf et al., 2020) to load pre-
L LAMA FACTORY consists of three main modules: trained models and initialize parameters. Specifi-
Model Loader, Data Worker, and Trainer. The cally, we load the vision language models using the
Model Loader manipulates various model archi- AutoModelForVision2Seq class while the rest are
tectures for fine-tuning, supporting both large lan- loaded using the AutoModelForCausalLM class.
guage models (LLMs) and vision language models The tokenizer is loaded using the AutoTokenizer
(VLMs). The Data Worker processes data from dif- class along with the model. In cases where the vo-
ferent tasks through a well-designed pipeline, sup- cabulary size of the tokenizer exceeds the capacity
porting both single-turn and multi-turn dialogues. of the embedding layer, we resize the layer and
The Trainer applies efficient fine-tuning techniques initialize new parameters with noisy mean initial-
to different training approaches, supporting pre- ization. To determine the scaling factor for RoPE
training, instruction tuning and preference opti- scaling (Chen et al., 2023), we compute it as the
mization. Beyond that, L LAMA B OARD provides a ratio of the maximum input sequence length to the
friendly visual interface to access these modules, context length of the model.

3
Model Patching To enable the S2 attention, we Plain text [{"text": "..."}, {"text": "..."}]
Alpaca-like data [{"instruction": "...", "input": "...", "output":
employ a monkey patch to replace the forward com- "..."}]
putation of models. However, we use the native ShareGPT-like data [{"conversations": [{"from": "human", "value":
"..."}, {"from": "gpt", "value": "..."}]}]
class to enable flash attention as it has been widely Preference data [{"instruction": "...", "input": "...", "output":
supported since Transformers 4.34.0. To prevent ["...", "..."]}]

excessive partitioning of the dynamic layers, we Standardized data {"prompt": [{"role": "...", "content": "..."}],
"response": [{"role": "...", "content": "..."}],
set the mixture-of-experts (MoE) blocks as leaf "system": "...", "tools": "...", "images": ["..."]}
modules when we optimize the MoE models under
DeepSpeed ZeRO stage-3 (Rasley et al., 2020). Table 3: Dataset structures in L LAMA FACTORY.

Model Quantization Dynamically quantizing


4.2 Data Worker
models to 8 bits or 4 bits with LLM.int8 (Dettmers
et al., 2022a) can be performed through the bitsand- We develop a data processing pipeline, including
bytes library (Dettmers, 2021). For 4-bit quantiza- dataset loading, dataset aligning, dataset merging
tion, we utilize the double quantization and 4-bit and dataset pre-processing. It standardizes datasets
normal float as QLoRA (Dettmers et al., 2023). We of different tasks into a unified format, enabling us
also support fine-tuning the models quantized by to fine-tune models on datasets in various formats.
the post-training quantization (PTQ) methods, in- Dataset Loading We utilize the Datasets (Lhoest
cluding GPTQ (Frantar et al., 2023), AWQ (Lin et al., 2021) library to load the data, which allows
et al., 2023), and AQLM (Egiazarian et al., 2024). the users to load remote datasets from the Hug-
Note that we cannot directly fine-tune the quan- ging Face Hub or read local datasets via scripts or
tized weights; thus, the quantized models are only through files. The Datasets library significantly re-
compatible with adapter-based methods. duces memory overhead during data processing and
accelerates sample querying using Arrow (Apache,
Adapter Attaching We automatically identify 2016). By default, the whole dataset is downloaded
the appropriate layers to attach adapters through to local disk. However, if a dataset is too large to be
traversing the model layers. The low-rank adapters stored, our framework provides dataset streaming
are attached to all the linear layers for a better con- to iterate over it without downloading.
vergence as suggested by (Dettmers et al., 2023).
The PEFT (Mangrulkar et al., 2022) library pro- Dataset Aligning To unify the dataset format, we
vides an extremely convenient way to implement design a data description specification to charac-
the adapter-based methods such as LoRA (Hu et al., terize the structure of datasets. For example, the
2022), rsLoRA (Kalajdzievski, 2023), DoRA (Liu alpaca dataset has three columns: instruction, in-
et al., 2024) and PiSSA (Meng et al., 2024). We put and output (Taori et al., 2023). We convert the
replace the backward computation with the one dataset into a standard structure that is compatible
of Unsloth (Han and Han, 2023) to accelerate the with various tasks according to the data description
training. To perform reinforcement learning from specification. Some examples of dataset structures
human feedback (RLHF), a value head layer is ap- are shown in Table 3.
pended on the top of the transformer model, map-
Dataset Merging The unified dataset structure
ping the representation of each token to a scalar.
provides an efficient approach for merging multiple
datasets. For the datasets in non-streaming mode,
Precision Adaptation We handle the floating-
we simply concatenate them before the datasets are
point precision of pre-trained models based on the
shuffled during training. However, in streaming
capabilities of computing devices. For NVIDIA
mode, simply concatenating the datasets impedes
GPUs, we adopt bfloat16 precision if the computa-
data shuffling. Therefore, we offer methods to
tion capability is 8.0 or higher. Otherwise, float16
alternately read the data from different datasets.
is adopted. Besides, we adopt float16 for Ascend
NPUs and AMD GPUs and float32 for non-CUDA Dataset Pre-processing L LAMA FACTORY is de-
devices. In mixed precision training, we set all signed for fine-tuning the text generative models,
trainable parameters to float32 for training stability. which is primarily used in chat completion. Chat
Nevertheless, we retain the trainable parameters as template is a crucial component in these mod-
bfloat16 in half precision training. els, because it is highly related to the instruction-

4
following abilities of these models. Therefore, we Distributed Training We can combine the above
provide dozens of chat templates that can be auto- trainers with DeepSpeed (Rasley et al., 2020; Ren
matically chosen according to the model type. We et al., 2021) for distributed training. We adopt
encode the sentence after applying the chat tem- data parallelism to fully exploit the ability of com-
plate using the tokenizer. By default, we only com- puting devices. Leveraging the DeepSpeed ZeRO
pute loss on the completions, while the prompts optimizer, the memory consumption can be further
are disregarded (Taori et al., 2023). Optionally, we reduced via partitioning or offloading.
can utilize sequence packing (Krell et al., 2021)
to reduce the training time, which is automatically 4.4 Utilities
enabled when performing generative pre-training. Model Inference During inference time, we
reuse the chat template from the Data Worker to
4.3 Trainer build the model inputs. We offer support for sam-
Efficient Training We integrate state-of-the-art pling the model outputs using Transformers (Wolf
efficient fine-tuning methods, including LoRA+ et al., 2020) and vLLM (Kwon et al., 2023), both
(Hayou et al., 2024), GaLore (Zhao et al., 2024) of which support stream decoding. Additionally,
and BAdam (Luo et al., 2024) to the Trainer by re- we implement an OpenAI-style API that utilizes
placing the default components. These fine-tuning the asynchronous LLM engine and paged attention
methods are independent of the Trainer, making of vLLM, to provide high-throughput concurrent
them easily applicable to various tasks. We utilize inference services, facilitating the deployment of
the trainers of Transformers (Wolf et al., 2020) for fine-tuned LLMs into various applications.
pre-training and SFT, while adopting the trainers of
Model Evaluation We include several metrics
TRL (von Werra et al., 2020) for RLHF and DPO.
for evaluating LLMs, including multiple-choice
We also include trainers of the advanced preference
tasks such as MMLU (Hendrycks et al., 2021),
optimization methods such as KTO (Ethayarajh
CMMLU (Li et al., 2023a), and C-Eval (Huang
et al., 2024) and ORPO (Hong et al., 2024) from
et al., 2023), as well as calculating text similar-
the TRL library. The tailored data collators are
ity scores like BLEU-4 (Papineni et al., 2002) and
leveraged to differentiate trainers of various train-
ROUGE (Lin, 2004). This feature facilitates users
ing approaches. To match the input format of the
to measure the abilities of the fine-tuned models.
trainers for preference data, we build 2n samples in
a batch where the first n samples are chosen exam- 4.5 L LAMA B OARD: A Unified Interface for
ples and the last n samples are rejected examples. L LAMA FACTORY
Model-Sharing RLHF Allowing RLHF training L LAMA B OARD is a unified user interface based
on consumer devices is crucial for democratizing on Gradio (Abid et al., 2019) that allows users to
LLM fine-tuning. However, it is difficult because customize the fine-tuning of LLMs without writing
RLHF training requires four different models. To any code. It offers a streamlined model fine-tuning
address this problem, we propose model-sharing and inference service, enabling users to easily ex-
RLHF, enabling entire RLHF training with no more plore the potential of LLMs in their environments.
than one pre-trained model. Concretely, we first L LAMA B OARD has the following notable features.
train an adapter and a value head with the objec- Easy Configuration L LAMA B OARD allows us
tive function for reward modeling, allowing the to customize the fine-tuning arguments through
model to compute reward scores. Then we initial- interaction with the web interface. We provide de-
ize another adapter and value head and train them fault values for a majority of arguments that are
with the PPO algorithm (Ouyang et al., 2022). The recommended for most users, simplifying the con-
adapters and value heads are dynamically switched figuration process. Moreover, users can preview
through the set_adapter and disable_adapter the datasets on the web UI to validate them.
methods of PEFT (Mangrulkar et al., 2022) dur-
ing training, allowing a single pre-trained model Monitorable Training During the training pro-
to serve as policy model, value model, reference cess, the training logs and loss curves are visualized
model, and reward model simultaneously. To the and updated in real time, allowing users to monitor
best of our knowledge, this is the first method that the training progress. This feature provides valu-
supports RLHF training on consumer devices. able insights to analyze the fine-tuning process.

5
Gemma-2B Llama2-7B Llama2-13B
Method Trainable Memory Throughput PPL Trainable Memory Throughput PPL Trainable Memory Throughput PPL
Params (GB) (Tokens/s) Params (GB) (Tokens/s) Params (GB) (Tokens/s)
Baseline / / / 11.83 / / / 7.53 / / / 6.66
Full-tuning 2.51B 17.06 3090.42 10.34 6.74B 38.72 1334.72 5.56 / / / /
Freeze-tuning 0.33B 8.10 5608.49 11.33 0.61B 15.69 2904.98 6.59 0.95B 29.02 1841.46 6.56
GaLore 2.51B 10.16 2483.05 10.38 6.74B 15.43 1583.77 5.88 13.02B 28.91 956.39 5.72
LoRA 0.16B 7.91 3521.05 10.19 0.32B 16.32 1954.07 5.81 0.50B 30.09 1468.19 5.75
QLoRA 0.16B 5.21 3158.59 10.46 0.32B 7.52 1579.16 5.91 0.50B 12.61 973.53 5.81

Table 4: Comparison of the training efficiency using different fine-tuning methods in L LAMA FACTORY. The best
result among GaLore, LoRA and QLoRA of each model is in bold.

Flexible Evaluation L LAMA B OARD supports we set the rank and scale to 128 and 2.0, respec-
calculating the text similarity scores on the datasets tively. For LoRA and QLoRA, we attach adapters
to automatically evaluate models or performing to all linear layers and set the rank and alpha to
human evaluation by chatting with them. 128 and 256, respectively. All the experiments are
conducted on a single NVIDIA A100 40GB GPU.
Multilingual Support L LAMA B OARD provides We enable flash attention in all experiments and
localization files, facilitating the integration of new Unsloth for LoRA and QLoRA experiments.
languages for rendering the interface. Currently
we support three languages: English, Russian and Results The results about the training efficiency
Chinese, which allows a broader range of users to are presented in Table 4, where memory refers
utilize L LAMA B OARD for fine-tuning LLMs. to the peak memory consumed during training,
throughput is calculated as the number of tokens
5 Empirical Study trained per second, and PPL represents the perplex-
We systematically evaluate L LAMA FACTORY from ity of the model on the training corpus. Since full-
two perspectives: 1) the training efficiency in terms tuning Llama2-13B lead to a memory overflow, the
of memory usage, throughput and perplexity. 2) the results are not recorded. We observe that QLoRA
effectiveness of adaptation to downstream tasks. consistently has the lowest memory footprint be-
cause the pre-trained weights are represented in
5.1 Training Efficiency lower precision. LoRA exhibits higher throughput
leveraging the optimization in LoRA layers by Un-
Experimental Setup We utilize the PubMed
sloth. GaLore achieves lower PPL on large models
dataset (Canese and Weis, 2013), which comprises
while LoRA advantages on smaller ones.
over 36 million records of biomedical literature.
We extract around 400K tokens from the abstract
5.2 Fine-Tuning on Downstream Tasks
of the literature to construct the training corpus.
Then we fine-tune the Gemma-2B (Mesnard et al., Experimental Setup To evaluate the effective-
2024), Llama2-7B and Llama2-13B (Touvron et al., ness of different efficient fine-tuning methods, we
2023b) models using the generative pre-training compare the performance of various models after
objective with various efficient fine-tuning meth- fine-tuning on downstream tasks. We construct non-
ods. We compare the results of full-tuning, freeze- overlapping training set and test set using 2,000 ex-
tuning, GaLore, LoRA and 4-bit QLoRA. After amples and 1,000 examples from three representa-
fine-tuning, we calculate the perplexity on the train- tive text generation tasks, including CNN/DM (Nal-
ing corpus to evaluate the efficiency of different lapati et al., 2016), XSum (Narayan et al., 2018)
methods. We also incorporate the perplexities of and AdGen (Shao et al., 2019), respectively. We
the pre-trained models as baselines. select several instruction-tuned models and fine-
In this experiment, we adopt a learning rate of tune them following the sequence-to-sequence task
−5
10 , a token batch size of 512. We fine-tune using different fine-tuning methods. Then we com-
these models using the 8-bit AdamW optimizer pare the results of full-tuning (FT), GaLore, LoRA
(Dettmers et al., 2022b) in bfloat16 precision with and 4-bit QLoRA. After fine-tuning, we calculate
activation checkpointing to reduce the memory the ROUGE score (Lin, 2004) on the test set of
footprint. In freeze-tuning, we only fine-tune the each task. We also incorporate the scores of the
last 3 decoder layers of the model. For GaLore, original instruction-tuned models as baselines.

6
CNN / DM XSum AdGen
Model Baseline FT GaLore LoRA QLoRA Baseline FT GaLore LoRA QLoRA Baseline FT GaLore LoRA QLoRA
ChatGLM3-6B 18.51 22.00 22.16 21.68 21.70 16.14 26.25 26.34 26.50 26.78 14.53 19.91 20.57 20.47 20.49
Yi-6B 16.85 22.40 22.68 22.98 22.97 18.24 27.09 28.25 28.71 29.21 13.34 19.68 20.06 20.97 20.31
Llama2-7B 12.94 22.87 22.40 22.70 22.61 13.89 27.69 27.64 28.80 28.05 0.61 20.51 19.61 20.29 20.45
Mistral-7B 14.39 22.03 22.99 23.47 23.28 15.87 23.57 28.00 30.41 30.44 7.82 20.14 20.90 20.99 20.56
Gemma-7B 15.97 22.07 / 22.41 22.44 15.31 25.13 / 28.67 29.02 11.57 19.99 / 20.62 19.81
Qwen1.5-7B 15.40 22.46 21.76 22.71 22.52 19.27 26.68 26.64 27.77 27.60 14.49 20.42 21.08 21.31 21.34
Qwen2-7B 16.46 23.20 / 23.29 23.66 19.76 26.94 / 28.92 28.94 12.89 19.83 / 20.96 20.86
Llama3-8B 15.19 23.36 23.57 23.48 24.12 17.83 26.21 30.45 30.63 30.94 0.22 20.28 21.27 21.44 21.20

Table 5: Comparison of the performance (in terms of ROUGE) on specific tasks using different fine-tuning methods
in L LAMA FACTORY. The best result of each model is underlined, and the best result of each task is in bold.

In this experiment, we set learning rate to 10−5 , efficiency and effectiveness of our framework on
batch size to 4 and maximum input length to 2048. language modeling and text generation tasks.
We fine-tune these models using the 8-bit AdamW We will consistently keep L LAMA FACTORY syn-
optimizer (Dettmers et al., 2022b) in bfloat16 pre- chronous with the state-of-the-art models and effi-
cision with activation checkpointing. For GaLore, cient fine-tuning techniques. We also welcome con-
we set the rank and scale to 128 and 2.0, respec- tributions from the open-source community. The
tively. For LoRA and QLoRA, we attach adapters road map of L LAMA FACTORY including:
to all linear layers and set the rank and alpha to (1) Enabling fine-tuning for models that supports
128 and 256, respectively. All the experiments are a wider range of modalities, e.g., the audio and
conducted on NVIDIA A100 40GB GPUs. video modalities (Zhu et al., 2024a).
(2) Integrating more parallel training strategies,
Results The evaluation results on downstream e.g., sequence parallelism (Jacobs et al., 2023) and
tasks are shown in Table 5. We report the averaged tensor parallelism (Shoeybi et al., 2019).
scores over ROUGE-1, ROUGE-2 and ROUGE- (3) Exploring stronger fine-tuning methods for
L. Some results of the Gemma-7B and Qwen2- conversational models, e.g., self-play (Chen et al.,
7B (Bai et al., 2023) models are not included in 2024c; Yuan et al., 2024).
the table because the GaLore method may not be
applicable to them. An interesting finding from 7 Broader Impact and Responsible Use
the results is that LoRA and QLoRA achieve the
L LAMA FACTORY has attracted a large number of
best performance in most cases, except for the
individuals interested in LLMs to explore the pos-
ChatGLM3-6B (Zeng et al., 2024) and Llama2-7B
sibility of customizing models. This contributes
models on the CNN/DM and AdGen datasets. This
significantly to the growth of the open-source com-
phenomenon highlights the effectiveness of these
munities. It is gaining increasing attention and is
efficient fine-tuning methods in adapting LLMs
being featured in Awesome Transformers2 as a rep-
to specific tasks. Additionally, we observe that
resentative of efficient fine-tuning frameworks for
Llama3-8B achieves the best performance among
LLMs. We anticipate that practitioners build their
these models, while Yi-6B (Young et al., 2024) and
LLMs upon our framework that bring benefits to
Mistral-7B (Jiang et al., 2023a) exhibit competitive
society. Adherence to the model license is manda-
performance among models of the same size.
tory when using L LAMA FACTORY for fine-tuning
LLMs, thus preventing from any potential misuse.
6 Conclusion and Future Work
Acknowledgements
In this paper, we demonstrate L LAMA FACTORY, a
unified framework for the efficient fine-tuning of This work is supported partly by the National Sci-
LLMs. Through a modular design, we minimize de- ence and Technology Major Project under Grant
pendencies between the models, datasets and train- 2022ZD0120202, by the National Natural Science
ing methods and provide an integrated approach to Foundation of China (No. U23B2056), by the Fun-
fine-tune over 100 LLMs with a diverse range of damental Research Funds for the Central Universi-
efficient fine-tuning techniques. Additionally, we ties, and by the State Key Laboratory of Complex
offer a flexible web UI L LAMA B OARD, enabling & Critical Software Environment.
customized fine-tuning and evaluation of LLMs 2
[Link]
without coding efforts. We empirically validate the 0/[Link]#llama-factory

7
References Du Chen, Yi Huang, Xiaopu Li, Yongqiang Li,
Yongqiang Liu, Haihui Pan, Leichao Xu, Dacheng
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Zhang, Zhipeng Zhang, and Kun Han. 2024a. Orion-
Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, 14b: Open-source multilingual large language mod-
Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harki- els. arXiv preprint arXiv:2401.12246.
rat Behl, et al. 2024. Phi-3 technical report: A highly
capable language model locally on your phone. arXiv Shouyuan Chen, Sherman Wong, Liangjian Chen, and
preprint arXiv:2404.14219. Yuandong Tian. 2023. Extending context window of
large language models via positional interpolation.
Abubakar Abid, Ali Abdalla, Ali Abid, Dawood Khan, arXiv preprint arXiv:2306.15595.
Abdulrahman Alfozan, and James Zou. 2019. Gradio:
Hassle-free sharing and testing of ml models in the Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos
wild. In ICML Workshop on Human in the Loop Guestrin. 2016. Training deep nets with sublinear
Learning, Long Beach, USA. memory cost. arXiv preprint arXiv:1604.06174.

Lightning AI. 2023. Lit-gpt. Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai,
Zhijian Liu, Song Han, and Jiaya Jia. 2024b. Lon-
AI@Meta. 2024. Llama 3. gLoRA: Efficient fine-tuning of long-context large
language models. In International Conference on
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Al- Learning Representations.
shamsi, Alessandro Cappelli, Ruxandra Cojocaru,
Mérouane Debbah, Étienne Goffinet, Daniel Hess- Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji,
low, Julien Launay, Quentin Malartic, et al. 2023. and Quanquan Gu. 2024c. Self-play fine-tuning con-
The falcon series of open language models. arXiv verts weak language models to strong language mod-
preprint arXiv:2311.16867. els. arXiv preprint arXiv:2401.01335.

Yuvanesh Anand, Zach Nussbaum, Brandon Duder- Yiming Cui, Ziqing Yang, and Xin Yao. 2023. Efficient
stadt, Benjamin Schmidt, and Andriy Mulyar. 2023. and effective text encoding for chinese llama and
GPT4All: Training an assistant-style chatbot with alpaca. arXiv preprint arXiv:2304.08177.
large scale data distillation from GPT-3.5-turbo.
Damai Dai, Chengqi Deng, Chenggang Zhao, RX Xu,
Apache. 2016. Arrow. Huazuo Gao, Deli Chen, Jiashi Li, Wangding
Zeng, Xingkai Yu, Y Wu, et al. 2024. DeepSeek-
MoE: Towards ultimate expert specialization in
Viraat Aryabumi, John Dang, Dwarak Talupuru,
mixture-of-experts language models. arXiv preprint
Saurabh Dash, David Cairuz, Hangyu Lin, Bharat
arXiv:2401.06066.
Venkitesh, Madeline Smith, Kelly Marchisio, Sebas-
tian Ruder, et al. 2024. Aya 23: Open weight re-
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and
leases to further multilingual progress. arXiv preprint
Christopher Ré. 2022. FlashAttention: Fast and
arXiv:2405.15032.
memory-efficient exact attention with io-awareness.
Advances in Neural Information Processing Systems,
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, 35:16344–16359.
Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei
Huang, et al. 2023. Qwen technical report. arXiv DeepSeek-AI, Aixin Liu, Bei Feng, Bin Wang, Bingx-
preprint arXiv:2309.16609. uan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr,
Chong Ruan, Damai Dai, Daya Guo, Dejian Yang,
Edward Beeching, Clémentine Fourrier, Nathan Habib, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fuli
Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Luo, Guangbo Hao, Guanting Chen, Guowei Li,
Sanseviero, Lewis Tunstall, and Thomas Wolf. 2023. H. Zhang, Hanwei Xu, et al. 2024. DeepSeek-v2: A
Open LLM leaderboard. strong, economical, and efficient mixture-of-experts
language model. arXiv preprint arXiv:2405.04434.
Rishabh Bhardwaj, Do Duc Anh, and Soujanya Poria.
2024. Language models are homer simpson! safety Tim Dettmers. 2021. Bitsandbytes.
re-alignment of fine-tuned language models through
task arithmetic. arXiv preprint arXiv:2402.11746. Tim Dettmers, Mike Lewis, Younes Belkada, and Luke
Zettlemoyer. 2022a. GPT3.int8(): 8-bit matrix mul-
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, tiplication for transformers at scale. Advances in
Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Neural Information Processing Systems, 35:30318–
Qiushi Du, Zhe Fu, et al. 2024. DeepSeek LLM: Scal- 30332.
ing open-source language models with longtermism.
arXiv preprint arXiv:2401.02954. Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke
Zettlemoyer. 2022b. 8-bit optimizers via block-wise
Kathi Canese and Sarah Weis. 2013. PubMed: the quantization. In International Conference on Learn-
bibliographic database. The NCBI handbook, 2(1). ing Representations.

8
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Parameter-efficient transfer learning for NLP. In In-
Luke Zettlemoyer. 2023. QLoRA: Efficient finetun- ternational Conference on Machine Learning, pages
ing of quantized llms. Advances in Neural Informa- 2790–2799. PMLR.
tion Processing Systems, 36:10088–10115.
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu,
Shizhe Diao, Rui Pan, Hanze Dong, KaShun Shum, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen,
Jipeng Zhang, Wei Xiong, and Tong Zhang. 2024. et al. 2022. LoRA: Low-rank adaptation of large
LMFlow: An extensible toolkit for finetuning and in- language models. In International Conference on
ference of large foundation models. In Proceedings Learning Representations.
of the 2024 Conference of the North American Chap-
ter of the Association for Computational Linguistics: Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu
Human Language Technologies (Volume 3: System Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxi-
Demonstrations), pages 116–127, Mexico City, Mex- ang Huang, Weilin Zhao, et al. 2024. MiniCPM:
ico. Association for Computational Linguistics. Unveiling the potential of small language models
with scalable training strategies. arXiv preprint
Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, arXiv:2404.06395.
Elias Frantar, Artem Babenko, and Dan Alistarh.
2024. Extreme compression of large language Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei
models via additive quantization. arXiv preprint Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu,
arXiv:2401.06118. Chuancheng Lv, Yikai Zhang, Yao Fu, et al. 2023.
C-Eval: A multi-level multi-discipline chinese eval-
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, uation suite for foundation models. Advances in
Dan Jurafsky, and Douwe Kiela. 2024. KTO: Model Neural Information Processing Systems, 36.
alignment as prospect theoretic optimization. In In-
ternational Conference on Machine Learning, Vi- Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang,
enna, Austria. PMLR. Minjia Zhang, Leon Song, Samyam Rajbhandari,
and Yuxiong He. 2023. Deepspeed ulysses: Sys-
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and tem optimizations for enabling training of extreme
Dan Alistarh. 2023. GPTQ: Accurate post-training long sequence transformer models. arXiv preprint
quantization for generative pre-trained transformers. arXiv:2309.14509.
In International Conference on Learning Representa-
tions. Albert Q Jiang, Alexandre Sablayrolles, Arthur Men-
sch, Chris Bamford, Devendra Singh Chaplot, Diego
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bha- de las Casas, Florian Bressand, Gianna Lengyel, Guil-
gia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh laume Lample, Lucile Saulnier, et al. 2023a. Mistral
Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, 7b. arXiv preprint arXiv:2310.06825.
et al. 2024. OLMo: Accelerating the science of lan-
guage models. arXiv preprint arXiv:2402.00838. Jinhao Jiang, Kun Zhou, Wayne Xin Zhao, Yaliang
Li, and Ji-Rong Wen. 2023b. ReasoningLM: En-
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, abling structural subgraph reasoning in pre-trained
Kai Dong, Wentao Zhang, Guanting Chen, Xiao language models for question answering over knowl-
Bi, Y Wu, YK Li, et al. 2024. DeepSeek-coder: edge graph. In Proceedings of the 2023 Conference
When the large language model meets programming– on Empirical Methods in Natural Language Process-
the rise of code intelligence. arXiv preprint ing, pages 3721–3735, Singapore. Association for
arXiv:2401.14196. Computational Linguistics.

Daniel Han and Michael Han. 2023. unsloth. Wenxiang Jiao, Jen-tse Huang, Wenxuan Wang, Zhi-
wei He, Tian Liang, Xing Wang, Shuming Shi, and
Soufiane Hayou, Nikhil Ghosh, and Bin Yu. 2024. Zhaopeng Tu. 2023a. ParroT: Translating during chat
LoRA+: Efficient low rank adaptation of large mod- using large language models tuned with human trans-
els. In International Conference on Machine Learn- lation and feedback. In Findings of the Association
ing, Vienna, Austria. PMLR. for Computational Linguistics: EMNLP 2023, pages
15009–15020, Singapore. Association for Computa-
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, tional Linguistics.
Mantas Mazeika, Dawn Song, and Jacob Steinhardt.
2021. Measuring massive multitask language under- Yizhu Jiao, Ming Zhong, Sha Li, Ruining Zhao, Siru
standing. In International Conference on Learning Ouyang, Heng Ji, and Jiawei Han. 2023b. Instruct
Representations. and extract: Instruction tuning for on-demand in-
formation extraction. In Proceedings of the 2023
Jiwoo Hong, Noah Lee, and James Thorne. 2024. Conference on Empirical Methods in Natural Lan-
ORPO: Monolithic preference optimization without guage Processing, pages 10030–10051, Singapore.
reference model. arXiv preprint arXiv:2403.07691. Association for Computational Linguistics.
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Damjan Kalajdzievski. 2023. A rank stabilization scal-
Bruna Morrone, Quentin De Laroussilhe, Andrea ing factor for fine-tuning with LoRA. arXiv preprint
Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. arXiv:2312.03732.

9
Dahyun Kim, Chanjun Park, Sanghoon Kim, Wonsung Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae
Lee, Wonho Song, Yunsu Kim, Hyeonwoo Kim, Lee. 2023. Visual instruction tuning. Advances in
Yungi Kim, Hyeonju Lee, Jihoo Kim, et al. 2023. Neural Information Processing Systems, 36.
SOLAR 10.7b: Scaling large language models with
simple yet effective depth up-scaling. arXiv preprint Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo
arXiv:2312.15166. Molchanov, Yu-Chiang Frank Wang, Kwang-Ting
Cheng, and Min-Hung Chen. 2024. DoRA: Weight-
Mario Michael Krell, Matej Kosec, Sergio P Perez, and decomposed low-rank adaptation. In International
Andrew Fitzgibbon. 2021. Efficient sequence pack- Conference on Machine Learning, Vienna, Austria.
ing without cross-contamination: Accelerating large PMLR.
language models without impacting performance.
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Fed-
arXiv preprint arXiv:2107.02027.
erico Cassano, Joel Lamy-Poirier, Nouamane Tazi,
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei,
Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- et al. 2024. Starcoder 2 and the stack v2: The next
zalez, Hao Zhang, and Ion Stoica. 2023. Efficient generation. arXiv preprint arXiv:2402.19173.
memory management for large language model serv- Qijun Luo, Hengxu Yu, and Xiao Li. 2024. BAdam:
ing with PagedAttention. In Proceedings of the 29th A memory efficient full parameter training method
Symposium on Operating Systems Principles, pages for large language models. arXiv preprint
611–626. arXiv:2404.02827.
Teven Le Scao, Angela Fan, Christopher Akiki, El- Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut,
lie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Younes Belkada, Sayak Paul, and Benjamin Bossan.
Castagné, Alexandra Sasha Luccioni, François Yvon, 2022. PEFT: State-of-the-art parameter-efficient fine-
Matthias Gallé, et al. 2022. BLOOM: A 176b- tuning methods.
parameter open-access multilingual language model.
arXiv preprint arXiv:2211.05100. Fanxu Meng, Zhaohui Wang, and Muhan Zhang. 2024.
PiSSA: Principal singular values and singular vectors
Quentin Lhoest, Albert Villanova del Moral, Yacine adaptation of large language models. arXiv preprint
Jernite, Abhishek Thakur, Patrick von Platen, Suraj arXiv:2404.02948.
Patil, Julien Chaumond, Mariama Drame, Julien Plu,
Lewis Tunstall, et al. 2021. Datasets: A community Thomas Mesnard, Cassidy Hardin, Robert Dadashi,
library for natural language processing. In Proceed- Surya Bhupatiraju, Shreya Pathak, Laurent Sifre,
ings of the 2021 Conference on Empirical Methods Morgane Rivière, Mihir Sanjay Kale, Juliette Love,
in Natural Language Processing: System Demonstra- et al. 2024. Gemma: Open models based on
tions, pages 175–184. gemini research and technology. arXiv preprint
arXiv:2403.08295.
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gre-
Zhao, Yeyun Gong, Nan Duan, and Timothy Bald-
gory Diamos, Erich Elsen, David Garcia, Boris Gins-
win. 2023a. CMMLU: Measuring massive multitask
burg, Michael Houston, Oleksii Kuchaiev, Ganesh
language understanding in chinese. arXiv preprint
Venkatesh, et al. 2018. Mixed precision training. In
arXiv:2306.09212.
International Conference on Learning Representa-
Shenggui Li, Hongxin Liu, Zhengda Bian, Jiarui Fang, tions.
Haichen Huang, Yuliang Liu, Boxiang Wang, and Ramesh Nallapati, Bowen Zhou, Cicero dos Santos,
Yang You. 2023b. Colossal-AI: A unified deep learn- Caglar Gulcehre, and Bing Xiang. 2016. Abstractive
ing system for large-scale parallel training. In Pro- text summarization using sequence-to-sequence rnns
ceedings of the 52nd International Conference on and beyond. In Proceedings of the 20th SIGNLL Con-
Parallel Processing, pages 766–775. ference on Computational Natural Language Learn-
ing, pages 280–290.
Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie
Del Giorno, Suriya Gunasekar, and Yin Tat Lee. Shashi Narayan, Shay B. Cohen, and Mirella Lapata.
2023c. Textbooks are all you need II: phi-1.5 techni- 2018. Don’t give me the details, just the summary!
cal report. arXiv preprint arXiv:2309.05463. topic-aware convolutional neural networks for ex-
treme summarization. In Proceedings of the 2018
Chin-Yew Lin. 2004. ROUGE: A package for auto- Conference on Empirical Methods in Natural Lan-
matic evaluation of summaries. In Text Summariza- guage Processing, pages 1797–1807, Brussels, Bel-
tion Branches Out, pages 74–81, Barcelona, Spain. gium. Association for Computational Linguistics.
Association for Computational Linguistics.
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida,
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Carroll Wainwright, Pamela Mishkin, Chong Zhang,
Xingyu Dang, and Song Han. 2023. AWQ: Sandhini Agarwal, Katarina Slama, Alex Ray, et al.
Activation-aware weight quantization for llm 2022. Training language models to follow instruc-
compression and acceleration. arXiv preprint tions with human feedback. Advances in Neural
arXiv:2306.00978. Information Processing Systems, 35:27730–27744.

10
Kishore Papineni, Salim Roukos, Todd Ward, and Wei- CodeGemma Team. 2024a. CodeGemma: Open
Jing Zhu. 2002. BLEU: a method for automatic eval- code models based on gemma. arXiv preprint
uation of machine translation. In Proceedings of the arXiv:2406.11409.
40th annual meeting of the Association for Compu-
tational Linguistics, pages 311–318, Philadelphia, InternLM Team. 2023. InternLM: A multilingual lan-
Pennsylvania, USA. Association for Computational guage model with progressively enhanced capabili-
Linguistics. ties.

Adam Paszke, Sam Gross, Francisco Massa, Adam PaliGemma Team. 2024b. Paligemma.
Lerer, James Bradbury, Gregory Chanan, Trevor
Killeen, Zeming Lin, Natalia Gimelshein, Luca Philippe Tillet, Hsiang-Tsung Kung, and David Cox.
Antiga, et al. 2019. PyTorch: An imperative style, 2019. Triton: An intermediate language and com-
high-performance deep learning library. Advances in piler for tiled neural network computations. In Pro-
Neural Information Processing Systems, 32. ceedings of the 3rd ACM SIGPLAN International
Workshop on Machine Learning and Programming
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Languages, pages 10–19.
Sutskever, et al. 2018. Improving language under-
standing by generative pre-training. OpenAI blog. Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier
Martinet, Marie-Anne Lachaux, Timothée Lacroix,
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal
Ermon, Christopher D Manning, and Chelsea Finn. Azhar, et al. 2023a. LLaMA: Open and effi-
2023. Direct preference optimization: Your language cient foundation language models. arXiv preprint
model is secretly a reward model. Advances in Neu- arXiv:2302.13971.
ral Information Processing Systems, 37.
Hugo Touvron, Louis Martin, Kevin Stone, Peter Al-
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and bert, Amjad Almahairi, Yasmine Babaei, Nikolay
Yuxiong He. 2020. DeepSpeed: System optimiza- Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti
tions enable training deep learning models with over Bhosale, et al. 2023b. Llama 2: Open founda-
100 billion parameters. In Proceedings of the 26th tion and fine-tuned chat models. arXiv preprint
ACM SIGKDD International Conference on Knowl- arXiv:2307.09288.
edge Discovery & Data Mining, pages 3505–3506.
Sang Truong, Duc Nguyen, Toan Nguyen, Dong Le,
Jie Ren, Samyam Rajbhandari, Reza Yazdani Am- Nhi Truong, Tho Quan, and Sanmi Koyejo. 2024.
inabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Crossing linguistic horizons: Finetuning and com-
Zhang, Dong Li, and Yuxiong He. 2021. ZeRO- prehensive evaluation of Vietnamese large language
offload: Democratizing billion-scale model training. models. In Findings of the Association for Computa-
In USENIX Annual Technical Conference, pages 551– tional Linguistics: NAACL 2024, pages 2849–2900,
564. Mexico City, Mexico. Association for Computational
Linguistics.
Zhihong Shao, Minlie Huang, Jiangtao Wen, Wenfei Xu,
and Xiaoyan Zhu. 2019. Long and diverse text gen- Lewis Tunstall, Edward Beeching, Nathan Lambert,
eration with planning-based hierarchical variational Nazneen Rajani, Kashif Rasul, Younes Belkada,
model. In Proceedings of the 2019 Conference on Shengyi Huang, Leandro von Werra, Clémentine
Empirical Methods in Natural Language Processing Fourrier, Nathan Habib, et al. 2023. Zephyr: Di-
and the 9th International Joint Conference on Natu- rect distillation of LM alignment. arXiv preprint
ral Language Processing (EMNLP-IJCNLP), pages arXiv:2310.16944.
3257–3268, Hong Kong, China. Association for Com-
putational Linguistics. Leandro von Werra, Younes Belkada, Lewis Tunstall,
Edward Beeching, Tristan Thrush, Nathan Lambert,
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, and Shengyi Huang. 2020. TRL: Transformer rein-
Junxiao Song, Mingchuan Zhang, YK Li, Y Wu, and forcement learning.
Daya Guo. 2024. DeepSeekMath: Pushing the limits
of mathematical reasoning in open language models. Chenglong Wang, Hang Zhou, Yimin Hu, Yifu Huo,
arXiv preprint arXiv:2402.03300. Bei Li, Tongran Liu, Tong Xiao, and Jingbo Zhu.
2023a. ESRL: Efficient sampling-based reinforce-
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, ment learning for sequence generation. arXiv
Patrick LeGresley, Jared Casper, and Bryan Catan- preprint arXiv:2308.02223.
zaro. 2019. Megatron-LM: Training multi-billion
parameter language models using model parallelism. Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li,
arXiv preprint arXiv:1909.08053. Sen Song, and Yang Liu. 2023b. OpenChat: Advanc-
ing open-source language models with mixed-quality
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann data. arXiv preprint arXiv:2309.11235.
Dubois, Xuechen Li, Carlos Guestrin, Percy Liang,
and Tatsunori B. Hashimoto. 2023. Stanford alpaca: Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang,
An instruction-following llama model. Dian Yu, Shuming Shi, and Zhaopeng Tu. 2023c.

11
Document-level machine translation with large lan- Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho,
guage models. In Proceedings of the 2023 Confer- Sainbayar Sukhbaatar, Jing Xu, and Jason Weston.
ence on Empirical Methods in Natural Language Pro- 2024. Self-rewarding language models. arXiv
cessing, pages 16646–16661, Singapore. Association preprint arXiv:2401.10020.
for Computational Linguistics.
Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang,
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao,
Hessel, Tushar Khot, Khyathi Chandu, David Wad- Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun,
den, Kelsey MacMillan, Noah A Smith, Iz Beltagy, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, et al.
et al. 2023d. How far can camels go? exploring the 2024. ChatGLM: A family of large language models
state of instruction tuning on open resources. Ad- from GLM-130b to GLM-4 all tools. arXiv preprint
vances in Neural Information Processing Systems, arXiv:2406.12793.
36.
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei
Zihan Wang, Xinzhang Liu, Shixuan Liu, Yitong Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao,
Yao, Yuyao Huang, Zhongjiang He, Xuelong Li, and Yu Qiao. 2024. LLaMA-adapter: Efficient fine-
Yongxiang Li, Zhonghao Che, Zhaoxi Zhang, et al. tuning of language models with zero-init attention.
2024. Telechat technical report. arXiv preprint In International Conference on Learning Representa-
arXiv:2401.03804. tions.
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang
Adams Wei Yu, Brian Lester, Nan Du, Andrew M Wang, Anima Anandkumar, and Yuandong Tian.
Dai, and Quoc V Le. 2022. Finetuned language mod- 2024. GaLore: Memory-efficient llm training by gra-
els are zero-shot learners. In International Confer- dient low-rank projection. In International Confer-
ence on Learning Representations. ence on Machine Learning, Vienna, Austria. PMLR.
Tianwen Wei, Liang Zhao, Lichang Zhang, Bo Zhu, Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang,
Lijie Wang, Haihua Yang, Biye Li, Cheng Cheng, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen
Weiwei Lü, Rui Hu, et al. 2023. Skywork: A more Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen
open bilingual foundation model. arXiv preprint Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang,
arXiv:2310.19341. Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu,
Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023.
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien A survey of large language models. arXiv preprint
Chaumond, Clement Delangue, Anthony Moi, Pier- arXiv:2303.18223.
ric Cistac, Tim Rault, Rémi Louf, Morgan Funtow-
icz, Joe Davison, Sam Shleifer, Patrick von Platen, Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan
Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin,
Teven Le Scao, Sylvain Gugger, Mariama Drame, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023.
Quentin Lhoest, and Alexander Rush. 2020. Trans- Judging LLM-as-a-judge with MT-bench and chatbot
formers: State-of-the-art natural language processing. arena. Advances in Neural Information Processing
In Proceedings of the 2020 Conference on Empirical Systems, 36.
Methods in Natural Language Processing: System
Demonstrations, pages 38–45, Online. Association Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui,
for Computational Linguistics. WANG HongFa, Yatian Pang, Wenhao Jiang, Junwu
Zhang, Zongwei Li, et al. 2024a. LanguageBind: Ex-
Shaohua Wu, Xudong Zhao, Shenling Wang, Jiangang tending video-language pretraining to N-modality by
Luo, Lingjun Li, Xi Chen, Bing Zhao, Wei Wang, language-based semantic alignment. In International
Tong Yu, Rongguo Zhang, et al. 2023. YUAN 2.0: A Conference on Learning Representations.
large language model with localized filtering-based
attention. arXiv preprint arXiv:2311.15786. Qihao Zhu, Daya Guo, Zhihong Shao, Dejian Yang,
Peiyi Wang, Runxin Xu, Y Wu, Yukun Li, Huazuo
Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Gao, Shirong Ma, et al. 2024b. DeepSeek-coder-v2:
Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Breaking the barrier of closed-source models in code
Dong Yan, et al. 2023. Baichuan 2: Open large-scale intelligence. arXiv preprint arXiv:2406.11931.
language models. arXiv preprint arXiv:2309.10305.
Alex Young, Bei Chen, Chao Li, Chengen Huang,
Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng
Zhu, Jianqun Chen, Jing Chang, et al. 2024. Yi:
Open foundation models by [Link]. arXiv preprint
arXiv:2403.04652.
Hao Yu, Zachary Yang, Kellin Pelrine, Jean Fran-
cois Godbout, and Reihaneh Rabbany. 2023. Open,
closed, or small language models for text classifica-
tion? arXiv preprint arXiv:2308.10092.

12
Model Variant Organization Release Date
Llama (Touvron et al., 2023a) 7B/13B/33B/65B Meta AI Feb. 2023
Llama 2 (Touvron et al., 2023b) 7B/13B/70B Meta AI Jul. 2023
Llama 3 (AI@Meta, 2024) 8B/70B Meta AI Apr. 2024
Aya 23 (Aryabumi et al., 2024) 8B/35B Cohere For AI May 2024
Baichuan (Yang et al., 2023) 7B/13B Baichuan Inc Jun. 2023
Baichuan2 (Yang et al., 2023) 7B/13B Baichuan Inc Sep. 2023
BLOOM (Le Scao et al., 2022) 560M/3B/7.1B BigScience May 2022
BLOOMZ (Le Scao et al., 2022) 560M/3B/7.1B BigScience Sep. 2022
ChatGLM2 (Zeng et al., 2024) 6B Zhipu AI Jun. 2023
ChatGLM3 (Zeng et al., 2024) 6B Zhipu AI Oct. 2023
ChineseLLaMA2 (Cui et al., 2023) 3B/7B/13B HFL Jul. 2023
CodeGemma (Team, 2024a) 2B/7B Google Apr. 2024
DeepSeek-Coder (Guo et al., 2024) 6.7B/7B/33B DeepSeek AI Oct. 2023
DeepSeek-Coder-V2 (Zhu et al., 2024b) 16B/236B DeepSeek AI Jun. 2024
DeepSeek-LLM (Bi et al., 2024) 7B/67B DeepSeek AI Nov. 2023
DeepSeek-Math (Shao et al., 2024) 7B DeepSeek AI Feb. 2024
DeepSeek-MoE (Dai et al., 2024) 16B DeepSeek AI Jan. 2024
DeepSeek-V2 (DeepSeek-AI et al., 2024) 16B/236B DeepSeek AI May 2024
Falcon (Almazrouei et al., 2023) 7B/11B/40B/180B TII Apr. 2023
Gemma (Mesnard et al., 2024) 2B/7B Google Feb. 2024
Gemma 2 (Mesnard et al., 2024) 9B/27B Google Jun. 2024
GLM-4 (Zeng et al., 2024) 9B Zhipu AI Jun. 2024
InternLM (Team, 2023) 7B/20B Shanghai AI Lab Jul. 2023
InternLM2 (Team, 2023) 1.8B/7B/20B Shanghai AI Lab Jan. 2024
LLaVA (Liu et al., 2023) 7B/13B / Apr. 2023
MiniCPM (Hu et al., 2024) 2B ModelBest Inc Jan. 2024
Mistral (Jiang et al., 2023a) 7B Mistral AI Sep. 2023
Mixtral (Jiang et al., 2023a) 8x7B/8x22B Mistral AI Dec. 2023
OLMo (Groeneveld et al., 2024) 1B/7B Allen AI Jan. 2024
OpenChat (Wang et al., 2023b) 7B OpenChat Jul. 2023
Orion (Chen et al., 2024a) 14B OrionStar Jan. 2024
PaliGemma (Team, 2024b) 3B Google May. 2024
Phi-1.5 (Li et al., 2023c) 1.3B Microsoft Sep. 2023
Phi-2 (Li et al., 2023c) 2.7B Microsoft Dec. 2023
Phi-3 (Abdin et al., 2024) 3.8B/7B/14B Microsoft Apr. 2024
Qwen (Bai et al., 2023) 1.8B/7B/14B/72B Alibaba Cloud Sep. 2023
Qwen1.5 (Bai et al., 2023) 1.8B/7B/14B/72B Alibaba Cloud Jan. 2024
Qwen2 (Bai et al., 2023) 0.5B/1.5B/7B/72B Alibaba Cloud Jun. 2024
SOLAR (Kim et al., 2023) 10.7B Upstage AI Dec. 2023
Skywork (Wei et al., 2023) 13B Skywork Oct. 2023
StarCoder2 (Lozhkov et al., 2024) 3B/7B/15B BigCode Feb. 2024
TeleChat (Wang et al., 2024) 7B/12B Telecom Mar. 2024
Vicuna1.5 (Zheng et al., 2023) 7B/13B LMSYS Jul. 2023
Yi (Young et al., 2024) 6B/9B/34B [Link] Nov. 2023
Yi-1.5 (Young et al., 2024) 6B/9B/34B [Link] May 2024
Yuan2 (Wu et al., 2023) 2B/51B/102B IEIT Dec. 2023
Zephyr (Tunstall et al., 2023) 7B Hugging Face H4 Oct. 2023

Table 6: List of supported models by L LAMA FACTORY. Note that the models from unknown sources are not
included in this list.

13

Common questions

Powered by AI

LLAMAFACTORY consists of three main components: Model Loader, Data Worker, and Trainer. The Model Loader allows precise attachment of adapters to pre-trained models by identifying exact layers, facilitating fine-tuning across various models. The Data Worker manages dataset alignment using a data description specification that enables effective dataset gathering. The Trainer activates plug-and-play implementations of state-of-the-art efficient fine-tuning methods through module replacement, enhancing flexibility and scalability in model training. This modular design allows reuse across different training approaches, significantly reducing integration costs .

LLAMAFACTORY supports a broader spectrum of fine-tuning techniques and training approaches compared to other frameworks like LLaMA-Adapter, FastChat, and Open-Instruct. Its comprehensive suite of efficient fine-tuning methods, integration of reinforcement learning, and robust adapter models, such as LoRA and DoRA, set it apart in terms of feature diversity. Furthermore, the inclusion of a web interface (LLAMABOARD) enhances usability significantly, distinguishing it from others which might require more manual coding .

LLAMAFACTORY tackles the challenge of efficient fine-tuning of numerous LLMs through its unified framework that incorporates scalable modular components and state-of-the-art fine-tuning methods. The use of a model registry and the automatic attachment of adapters facilitate streamlined adaptation of pre-trained models. Efficient methods like LoRA and precision handling strategies reduce computational resource demands during fine-tuning processes, thus enabling high throughput with minimal resource consumption .

The modular design of LLAMAFACTORY, comprising adaptable components like the Model Loader, Data Worker, and Trainer, significantly enhances its adaptability to diverse machine learning tasks. This design permits seamless customization and configuration of different models and datasets, facilitating their integration across multiple environments. By supporting plug-and-play methods for fine-tuning and maintaining a minimal dependency framework, LLAMAFACTORY efficiently adapts to task-specific requirements while ensuring resource optimization and ease of updates or component upgrades .

LLAMAFACTORY enhances the fine-tuning experience for non-technical users by providing a no-coding-needed interface through LLAMABOARD, which allows easy customization and adaptation of LLMs. This is coupled with a high level of abstraction in its architecture, enabling users to perform complex fine-tuning tasks with intuitive controls and minimal technical intervention, thus broadening accessibility to advanced machine learning processes .

LLAMAFACTORY enhances its scalability by minimizing dependencies between its main modules and specific models or datasets, which allows the framework to adapt seamlessly to various LLMs and datasets. The architecture supports broad dataset types and model forms by establishing a model registry for precise adapter attachment and employing a data description specification for dataset alignment. This modular flexibility allows scaling across a diverse range of scenarios with minimal integration requirements .

LLAMAFACTORY utilizes PyTorch as its core implementation framework, integrating open-source libraries such as Transformers, PEFT, and TRL to build an out-of-the-box fine-tuning system for LLMs. These libraries provide advanced functionalities such as efficient fine-tuning mechanisms and user-friendly interfaces, which augment LLAMAFACTORY's ability to offer minimal code fine-tuning while allowing scalable and flexible adaptation of multiple language models .

In LLAMAFACTORY, adapter-based methods provide significant advantages in handling model quantization. Due to the inability to directly fine-tune quantized weights, these methods support fine-tuning by attaching low-rank adapters to linear layers, enhancing convergence. Adapter-based methods, such as LoRA and its variants, offer efficient handling of precision issues, allowing fine-tuning to be effectively conducted on dynamically quantized models. This ensures systematic adaptation while maintaining computational efficiency in resource-constrained environments .

LLAMABOARD serves as LLAMAFACTORY's web interface, enabling users to customize the fine-tuning of large language models without requiring extensive coding. It simplifies the user experience by allowing them to define and manipulate training parameters visually and intuitively. LLAMABOARD acts as a bridge, translating complex backend processes into manageable tasks, thereby democratizing access to advanced model adaptation techniques .

LLAMAFACTORY incorporates several innovations to effectively manage training costs for LLMs. Among these are scalable modules that employ adapter techniques to minimize the need for retraining entire models, saving computational resources. The integration of mixed precision training and dynamic model quantization further reduces hardware demands, while precision adaptation according to device capabilities enhances operational efficiency. These methods synergistically lower cost barriers associated with the training and deployment of large language models .

You might also like