Optimizing LLMs for Code Review Automation
Optimizing LLMs for Code Review Automation
Keywords: Context: The rapid evolution of Large Language Models (LLMs) has sparked significant interest in leveraging
Modern code review their capabilities for automating code review processes. Prior studies often focus on developing LLMs for code
Code review automation review automation, yet require expensive resources, which is infeasible for organizations with limited budgets
Large language models
and resources. Thus, fine-tuning and prompt engineering are the two common approaches to leveraging LLMs
GPT-3.5
for code review automation.
Few-shot learning
Persona
Objective: We aim to investigate the performance of LLMs-based code review automation based on two
contexts, i.e., when LLMs are leveraged by fine-tuning and prompting. Fine-tuning involves training the model
on a specific code review dataset, while prompting involves providing explicit instructions to guide the model’s
generation process without requiring a specific code review dataset.
Methods: We leverage model fine-tuning and inference techniques (i.e., zero-shot learning, few-shot learning
and persona) on LLMs-based code review automation. In total, we investigate 12 variations of two LLMs-based
code review automation (i.e., GPT-3.5 and Magicoder), and compare them with the Guo et al.’s approach and
three existing code review automation approaches (i.e., CodeReviewer, TufanoT5 and D-ACT).
Results: The fine-tuning of GPT 3.5 with zero-shot learning helps GPT-3.5 to achieve 73.17%–74.23% higher
EM than the Guo et al.’s approach. In addition, when GPT-3.5 is not fine-tuned, GPT-3.5 with few-shot learning
achieves 46.38%–659.09% higher EM than GPT-3.5 with zero-shot learning.
Conclusions: Based on our results, we recommend that (1) LLMs for code review automation should be fine-
tuned to achieve the highest performance.; and (2) when data is not sufficient for model fine-tuning (e.g., a
cold-start problem), few-shot learning without a persona should be used for LLMs for code review automation.
Our findings contribute valuable insights into the practical recommendations and trade-offs associated with
deploying LLMs for code review automation.
∗ Corresponding author.
E-mail addresses: [Link]@[Link] (C. Pornprasit), chakkrit@[Link] (C. Tantithamthavorn).
[Link]
Received 31 January 2024; Received in revised form 13 June 2024; Accepted 6 July 2024
Available online 11 July 2024
0950-5849/© 2024 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY license ([Link]
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
2
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
review automation approaches [5,6,9]. For example, Li et al. [5] pro- Table 1
The differences between our work and Guo et al.’s work [14].
posed CodeReviewer, a pre-trained LLM that is based on the CodeT5
model [8]. Prior studies found that LLMs-based code review automa- Guo et al. [14] Our work
tion approaches often outperform NMT-based ones [4–6]. For exam- LLMs/approaches GPT-3.5, GPT-3.5, Magicoder
CodeReviewer [21], CodeReviewer
ple, Li et al. [5] found that their proposed approach outperforms a
[5] [5], TufanoT5 [6],
transformer-based NMT model by 11.76%. Below, we briefly discuss D-ACT [4]
the general modeling pipeline of LLMs for code review automation
Include fine-tuning No Yes
presented in Fig. 1. LLMs?
Model Pre-Training refers to the initial phase of training a large Prompting Zero-shot Zero-shot learning,
language model, where the model is exposed to a large amount of techniques learning, Persona Few-shot learning,
unlabeled data to learn general language representations. This phase Persona
aims to initialize the model’s parameters and learn generic features that
can be further fine-tuned for specific downstream tasks. Recently, there
have been many large language models for code (i.e., LLMs that are
tasks [35–38]. In particular, zero-shot learning involves prompting
specifically trained on source code and related natural languages). For
LLMs to generate an output from a given instruction and an input.
example, the open-source community-developed large language models
On the other hand, few-shot learning [18,30,31] involves prompt-
such as Code-LLaMa [27], StarCoder [28], and Magicoder [21]; and the
ing LLMs to generate an output from 𝑁 demonstration examples
commercial large language models such as GPT-3.5.
{(𝑥1 , 𝑦1 ), (𝑥2 , 𝑦2 ), … , (𝑥𝑁 , 𝑦𝑁 )} and an actual input in a testing set, where
However, the development of large language models for code re- 𝑥𝑖 and 𝑦𝑖 are the inputs and outputs obtained from a training set,
quires expensive GPU resources and budget. For example, GPT-3.5 re- respectively. Persona [13] involves prompting LLMs to act as a specific
quires 10,000 NVIDIA V-100 GPUs for model pre-training.1 LLaMa2 [29] role or persona to ensure that LLMs will generate output that is similar
requires Meta’s Research Super Cluster (RSC) as well as internal pro- to the output generated by a specified persona.
duction clusters, which consists of approximately 2000 NVIDIA A-100
GPUs in total. Therefore, many software organizations with limited 2.3. GPT-3.5 for code review automation
resources and budgets may not be able to develop their large language
models. Thus, fine-tuning and prompt engineering are the two common Recently, Guo et al. [14] conducted an empirical study to investi-
approaches to leverage the existing LLMs for code review automation gate the potential of GPT-3.5 for code review automation. However,
when expensive GPU resources are not available for pre-training a their study still has the following limitations (see Table 1).
large language model from scratch, where these techniques are more First, the results of Guo et al. [14] are limited to zero-shot GPT-
desirable for many organizations to quickly adopt new technologies. 3.5. In particular, Guo et al. [14] conducted experiments to find the
Model Fine-Tuning is a common practice, particularly in transfer best prompt for leveraging zero-shot learning with GPT-3.5. However,
learning scenarios, where a model pre-trained on a large dataset (source there are other approaches to leverage GPT-3.5 (i.e., fine-tuning and
domain, e.g., source code understanding) is adapted to a related but few-shot learning) that are not included in their study. The lack of a
different task or dataset (target domain, e.g., code review automation). systematic evaluation of the use of fine-tuning and few-shot learning on
GPT-3.5 makes it difficult for practitioners to conclude which approach
Recently, researchers have leveraged model fine-tuning techniques for
is the best for leveraging LLMs for code review automation. To address
LLMs to improve the performance of code review automation ap-
this challenge, we formulate the following research question.
proaches. For example, Lu et al. [9] proposed LLaMa-Reviewer, which
is an LLM-based code review automation approach that is being fine- RQ1: What is the most effective approach to leverage LLMs for code
tuned on a base LLaMa model [10] using three code review automation review automation?
tasks, i.e., a review necessity prediction task to check if diff hunks need
a review, a code review comment generation task to generate pertinent Second, the performance of LLMs when being fine-tuned is
comments for a given code snippet, and a code refinement task to still unknown. In particular, Guo et al. [14] did not evaluate the
generate minor adjustments to the existing code. Lu et al. [9] found performance of LLMs when being fine-tuned. However, prior stud-
that the fine-tuning step on LLMs can greatly improve the performance ies [15–17] found that model fine-tuning can improve the performance
of the existing code review automation approaches. of pre-trained LLMs. The lack of experiments with model fine-tuning
Inference refers to the process of using a pre-trained language makes it difficult for practitioners to conclude whether LLMs for code
review automation should be fine-tuned to achieve the most effective
model to generate source code based on a given natural language
results. To address this challenge, we formulate the following research
prompt instruction. Therefore, prompt engineering plays an impor-
question.
tant role in leveraging LLMs for code review automation to guide
LLMs to generate the desired output. Different prompting strategies RQ2: What is the benefit of model fine-tuning on GPT-3.5 for code
have been proposed.2 For example, zero-shot learning, few-shot learn- review automation?
ing [18,30,31], chain-of-thought [32,33], tree-of-thought [32,33], self-
consistency [34], and persona [13]. Nevertheless, not all prompting Third, the performance of LLMs for code review automation
strategies are relevant to code review automation. For example, chain- when using few-shot learning is still unknown. In particular, Guo
of-thought, self-consistency and tree-of-thought promptings are not et al. [14] did not investigate the impact of few-shot learning on LLMs
applicable to the code review automation task since they are designed for code review automation. However, recent work [18–20] found
for arithmetic and logical reasoning problems. Thus, we exclude them that few-shot learning could improve the performance of LLMs over
zero-shot learning. The lack of experiments with few-shot learning on
from our study.
LLMs for code review automation makes it difficult for practitioners to
In contrast, zero-shot learning, few-shot learning, and persona prompt-
conclude which prompting strategy (i.e., zero-shot learning, few-shot
ing are the instruction-based prompting strategies, which are more
learning, and persona) is the most effective for code review automa-
suitable for software engineering (including code review automation)
tion. To address this challenge, we formulate the following research
question.
1
[Link] RQ3: What is the most effective prompting strategy on GPT-3.5 for
ChatGPT code review automation?
2
[Link]
3
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
Fig. 2. An overview of our experimental design (A persona is a part of zero-shot and few-shot learning).
Table 2 other hand, the testing set consists of only code submitted for review
Experimental settings in our study. We do not include experimental settings #3 and
and reviewers’ comments. Next, to fine-tune the studied LLMs, we first
#4 since LLMs already learn the relationship between input (i.e., code submitted for
review) and output (i.e., revised code).
randomly obtain a set of training examples from the training set since
Experimental setting Fine-Tuning Inference technique
using the whole training set is prohibitively expensive. Then, we use
the selected training examples to fine-tune the studied LLMs. On the
Prompting Use Persona
other hand, to use the inference techniques (i.e., zero-shot learning,
#1 ✗
Zero-shot few-shot learning and a persona), we first design prompt templates
#2 ✓
✓ for each inference technique based on the guideline from OpenAI.3 , 4
#3 ✗
Few-shot However, since few-shot learning requires demonstration examples, we
#4 ✓
select a set of demonstration examples for each testing sample from the
#5 ✗
Zero-shot training set. Then, we create prompts that look similar to the prompt
#6 ✓
✗ templates. Finally, we use the studied LLMs to generate revised code
#7 ✗
#8
Few-shot
✓
from given prompts. We explain the details of the studied datasets,
model fine-tuning, inference via prompting, evaluation measures, and
hyper-parameter settings below.
In this section, we provide an overview and details of our experi- Recently, Tufano et al. [1,2] collected datasets with the constraint
mental design. that revised code must not contain the code tokens (e.g., identifiers)
that do not appear in code submitted for review. Thus, such datasets do
3.1. Overview not align with the real code review practice since developers may add
new code tokens when they revise their submitted code. Therefore, in
The goal of this work is to investigate which LLMs perform best this study, we use the CodeReviewer [5], TufanoT5 [6], and D-ACT [4]
when using model fine-tuning and inference techniques (i.e., zero-shot datasets, which do not have the above constraint in data collection
learning, few-shot learning [18,30,31], and persona [13]). To achieve instead. The details of the studied datasets are as follows (the statistic
this goal, we conduct experiments with two LLMs (i.e., GPT-3.5 and of the studied datasets is presented in Table 3).
Magicoder [21]) on the following datasets that are widely studied in the
• CodeReviewerdata : Li et al. [5] collected this dataset from the
code review automation literature [4,9,14,39]: CodeReviewerdata [5],
GitHub projects across nine programming languages (i.e., C, C++,
Tufanodata [6] and D-ACTdata [4]. We use Magicoder [21] in our exper-
C#, Java, Python, Ruby, php, Go, and Javascript). The dataset
iment since it is further trained on high-quality synthetic instructions
contains triplets of the code submitted for review (diff hunk
and solutions.
granularity), a reviewer’s comment, and the revised version of the
In this study, we conduct experiments under six settings as pre-
code submitted for review (diff hunk granularity).
sented in Table 2. According to the table, when the LLMs are fine-tuned,
• Tufanodata : Tufano et al. [6] collected this dataset from Java
we use zero-shot learning with and without a persona. We do not use
projects in GitHub, and 6388 Java projects hosted in Gerrit. Each
few-shot learning with the fine-tuned LLMs since the LLMs already
record in the dataset contains a triplet of code submitted for re-
learn the relationship between an input (i.e., code submitted for review)
view (function granularity), a reviewer’s comment, and code after
and an output (i.e., improved code). On the other hand, when the LLMs
being revised (function granularity). Tufano et al. [6] created
are not fine-tuned, we use zero-shot learning and few-shot learning,
two types of this dataset (i.e., Tufanodata (with comment) and
where each inference technique is used with and without a persona.
Tufanodata (without comment)).
Finally, we conduct 36 experiments in total (2 LLMs × 6 settings × 3
datasets).
Fig. 2 provides an overview of our experimental design. To begin, 3
[Link]
the studied code review datasets are split into training and testing engineering-with-openai-api
sets. The training set consists of the code submitted for review and 4
[Link]
reviewers’ comments as input; and revised code as output. On the write-clear-instructions
4
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
Table 3
A statistic of the studied datasets (the dataset of Android, Google and Ovirt are from the D-ACTdata dataset [4]).
Dataset # Train # Validation # Test # Language Granularity Has Comment
CodeReviewerdata [5] 150,405 13,102 13,104 9 Diff Hunk ✓
Tufanodata [6] 134,238 16,779 16,779 1 Function ✓/✗
Android [4] 14,690 1,836 1,835 1 Function ✗
Google [4] 9,899 1,237 1,235 1 Function ✗
Ovirt [4] 21,509 2,686 2,688 1 Function ✗
5
[Link]
6
tuned-model [Link]
5
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
Table 4
The evaluation results of GPT-3.5, Magicoder and the existing code review automation approaches.
Approach Fine- Inference technique CodeReviewer Tufano (with comment) Tufano (without comment) Android Google Ovirt
Tuning
Prompting Use Persona EM CodeBLEU EM CodeBLEU EM CodeBLEU EM CodeBLEU EM CodeBLEU EM CodeBLEU
✗ 37.93% 49.00% 22.16% 82.99% 6.02% 79.81% 2.34% 74.15% 6.71% 81.08% 3.05% 74.67%
✓
✓ 37.70% 49.20% 21.98% 83.04% 6.04% 79.76% 2.29% 74.74% 6.14% 81.02% 2.64% 74.95%
Zero-shot
✗ 17.72% 44.17% 13.52% 78.36% 2.62% 74.92% 0.49% 61.85% 0.16% 61.04% 0.48% 56.55%
GPT-3.5
✓ 17.07% 43.11% 12.49% 77.32% 2.29% 73.21% 0.57% 55.88% 0.00% 50.65% 0.22% 45.73%
✗
✗ 26.55% 47.50% 19.79% 81.47% 8.96% 79.21% 2.34% 75.33% 2.89% 81.40% 1.64% 73.83%
Few-shot
✓ 26.28% 47.43% 20.03% 81.61% 9.18% 78.98% 1.62% 74.65% 2.45% 81.07% 1.67% 73.29%
✗ 27.43% 44.86% 11.14% 69.77% 1.97% 69.25% 0.27% 65.39% 0.57% 69.30% 0.30% 64.19%
✓
✓ 27.98% 45.36% 11.06% 69.60% 2.12% 68.84% 0.65% 65.41% 1.13% 69.30% 1.00% 64.16%
Zero-shot
✗ 9.75% 39.45% 8.65% 73.90% 0.81% 59.49% 0.16% 47.37% 0.08% 48.82% 0.04% 44.38%
Magicoder
✓ 9.93% 39.48% 8.71% 73.57% 1.51% 67.65% 0.27% 47.46% 0.08% 48.58% 0.11% 43.65%
✗
✗ 15.89% 36.24% 2.93% 4.36% 1.99% 7.49% 0.22% 37.61% 0.49% 42.50% 0.74% 40.33%
Few-shot
✓ 17.80% 38.93% 2.89% 3.70% 1.84% 6.96% 0.27% 16.59% 0.65% 18.83% 0.82% 19.55%
CodeReviewer [5] 33.23% 55.43% 15.17% 80.83% 4.14% 78.76% 0.54% 75.24% 0.81% 80.10% 1.23% 75.32%
– – –
TufanoT5 [6] 11.90% 43.39% 14.26% 79.48% 5.40% 77.26% 0.27% 75.88% 1.37% 82.25% 0.19% 73.53%
Table 5 4. Result
The statistical details of GPT-3.5, Magicoder
and the existing code review automation
approaches.
In this section, we present the results of the following three research
questions.
Model # parameters
GPT-3.5 175 B (RQ1) What is the most effective approach to leverage LLMs for
Magicoder [21] 6.7 B code review automation?
TufanoT5 [6] 60.5 M
Approach. To address this RQ, we leverage fine-tuning and inference
CodeReviewer [5] 222.8 M
techniques (i.e., zero-shot learning, few-shot learning, and persona) on
D-ACT [4] 222.8 M
GPT-3.5 and Magicoder (The details of GPT-3.5 and Magicoder are
presented in Table 5) . Then, we measure EM of the results obtained
from GPT-3.5, Magicoder and Guo et al.’s approach [14].
the generated revised code with the actual revised code, we first
Result. The fine-tuning of GPT 3.5 with zero-shot learning helps
tokenize both revised code to sequences of tokens. Then, we
GPT-3.5 to achieve 73.17%–74.23% higher EM than the Guo et al.
compared the sequence of tokens of the generated revised code
[14]’s approach. Table 4 shows the results of EM achieved by GPT-
with the sequence of tokens of the actual revised code. A high 3.5, Magicoder and Guo et al.’s approach [14]. The table shows that
value of EM indicates that a model can generate revised code when GPT-3.5 and Magicoder are fine-tuned, such models achieve
that is the same as the actual revised code in the testing dataset. 73.17%–74.23% and 26.00%–28.53% higher EM than the Guo et al.’s
2. CodeBLEU [22] is the extended version of BLEU (i.e., an n-gram approach [14], respectively.
overlap between the translation generated by a deep learning The results indicate that model fine-tuning could help GPT-3.5 and
model and the translation in ground truth) [43] for automatic Magicoder to achieve higher EM when compared to the Guo et al.’s
evaluation of the generated code. We do not measure BLEU like approach [14]. The higher EM has to do with model fine-tuning. When
in prior work [5,6] since Ren et al. [22] found that this measure GPT-3.5 or Magicoder is fine-tuned, such models learn the relationship
ignores syntactic and semantic correctness of the generated code. between inputs (i.e., code submitted for review and a reviewer’s com-
In addition to BLEU, CodeBLEU considers the weighted n-gram ment) and an output (i.e., revised code) from a number of examples in
match, matched syntactic information (i.e., abstract syntax tree: a training set. On the contrary, Guo et al.’s approach [14] only relies on
the instruction and given input to generate revised code, which GPT-3.5
AST) and matched semantic information (i.e., data flow: DF)
never learned during model pre-training.
when computing the similarity between the generated revised
(RQ2) What is the benefit of model fine-tuning on GPT-3.5 for code
code and the actual revised code. A high value of CodeBLEU indi-
review automation?
cates that a model can generate revised code that is syntactically Approach. To address this RQ, we fine-tune GPT-3.5 as explained in
and semantically similar to the actual revised code in the testing Section 3. Then, we measure EM and CodeBLEU of the results obtained
dataset. from the fine-tuned GPT-3.5 and the non fine-tuned GPT-3.5 with
zero-shot learning.
Result. The fine-tuning of GPT 3.5 with zero-shot learning helps
3.6. The hyper-parameter settings GPT-3.5 to achieve 63.91%–1100% higher EM than those that are
not fine-tuned. Table 4 shows that in terms of EM, the fine-tuning
In this study, we use the following hyper-parameter settings when of GPT 3.5 with zero-shot learning helps GPT-3.5 to achieve 63.91%–
using GPT-3.5 to generate revised code: temperature of 0.0 (as sug- 1100% higher than those that are not fine-tuned. In terms of CodeBLEU,
the fine-tuning of GPT 3.5 with zero-shot learning helps GPT-3.5 to
gested by Guo et al. [14]), top_p of 1.0 (default value), and max
achieve 5.91%–63.9% higher than those that are not fine-tuned.
length of 512. To fine-tune GPT-3.5, we use hyper-parameters (e.g., the
The results indicate that fine-tuned GPT-3.5 achieve higher EM and
number of epochs and learning rate) that are automatically provided by
CodeBLEU than those that are not fine-tuned. During the model fine-
OpenAI API.
tuning process, GPT-3.5 adapt to the code review automation task by
For Magicoder [21], we use the same hyper-parameter as GPT-3.5 directly learning the relationship between inputs (i.e., code submitted
to generate revised code. To fine-tune Magicoder, we use the following for review and a reviewer’s comment) and an output (i.e., revised code)
hyper-parameters for DoRA [40]: attention dimension (𝑟) of 16, alpha from a number of examples in a training set. In contrast, non fine-
(𝛼) of 8, and dropout of 0.1 . tuned GPT-3.5 is given only an instruction and inputs which are not
6
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
Fig. 4. (RQ3) Examples of the difference between code submitted for review and revised code generated by GPT-3.5 with zero-shot learning and few-shot learning.
presented during model pre-training. Therefore, fine-tuned GPT-3.5 can in Section 3. Then, similar to RQ2, we measure EM and CodeBLEU of
better adapt to the code review automation task than those that are not the results obtained from GPT-3.5.
fine-tuned. Result. GPT-3.5 with few-shot learning achieves 46.38%–659.09%
(RQ3) What is the most effective prompting strategy on GPT-3.5 higher EM than GPT-3.5 with zero-shot learning. Table 4 shows that
for code review automation? in terms of EM, the use of few-shot learning on GPT-3.5 helps GPT-3.5
Approach. To address this RQ, we use zero-shot learning and few-shot to achieve 46.38%–241.98% and 53.95%–659.09% higher than the use
learning with non fine-tuned GPT-3.5, where each inference technique is of zero-shot learning on GPT-3.5 without a persona and with a persona,
used with and without a persona, to generate revised code as explained respectively. In terms of CodeBLEU, the use of few-shot learning on
7
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
GPT-3.5 helps GPT-3.5 to achieve 3.97%–33.36% and 5.55%–60.27% automation approaches that require the whole training set to adapt to
higher than the use of zero-shot learning on GPT-3.5 without a persona the code review automation task.
and with a persona, respectively. Recommendations to practitioners. LLMs for code review au-
The results indicate that GPT-3.5 with few-shot learning can achieve tomation should be fine-tuned to achieve the highest performance. The
higher EM and CodeBLEU than GPT-3.5 with zero-shot learning. When reason for this recommendation is the results of RQ2 show that fine-
few-shot learning is used to generate revised code from GPT-3.5, GPT- tuned GPT-3.5 outperforms those that are not fine-tuned. In contrast,
3.5 has more information from given demonstration examples in a when data is not sufficient for model fine-tuning (e.g., a cold-start
prompt to guide the generation of revised code from given code sub- problem), few-shot learning without a persona should be used for LLMs
mitted for review and a reviewer’s comment. In other words, such for code review automation. The reason for this recommendation is
demonstration examples could help GPT-3.5 to correctly generate re- the results of RQ3 show that GPT-3.5 with few-shot learning outper-
vised code from a given code submitted for review and a reviewer’s forms GPT-3.5 with zero-shot learning, and GPT-3.5 without a persona
comment. outperforms GPT-3.5 with persona.
When a persona is included in input prompts, GPT-3.5 achieves
1.02%–54.17% lower EM than when the persona is not included in 5.2. The characteristics of the revised code that are correctly generated by
input prompts. Table 4 also shows that when a persona is included GPT-3.5
in prompts, GPT-3.5 with zero-shot and few-shot learning achieves
3.67%–54.17% and 1.02%–30.77% lower EM compared to when a The results of RQ2 and RQ3 demonstrate the benefits of model
persona is not included in prompts, respectively. Similarly, when a fine-tuning and few-shot learning on GPT-3.5 for the code review au-
persona is included in prompts, GPT-3.5 with zero-shot and few-shot tomation task, respectively. However, practitioners still do not clearly
learning achieves 1.33%–19.13% and 0.15%–0.90% lower CodeBLEU understand the characteristics of the code changes of the revised code
compared to when a persona is not included in prompts, respectively. that GPT-3.5 correctly generates. To address this challenge, we aim to
The results indicate that when a persona is included in prompts, qualitatively investigate the revised code that GPT-3.5 can correctly
GPT-3.5 with zero-shot and few-shot learning achieves lower EM and generate. To do so, we randomly select the revised code that is only
CodeBLEU. To illustrate the impact of the persona, we present an correctly generated by a particular model (e.g., we randomly obtain the
example of the revised code that GPT-3.5 generates when zero-shot revised code that only fine-tuned GPT-3.5 correctly generates while the
learning with and without a persona is used, and an example of the others do not.) by using a confidence level of 95% and a confidence
revised code that GPT-3.5 generates when few-shot learning with and interval of 5%. Then, we classify the code changes of the selected
without a persona is used in Fig. 4. revised code into the following categories based on the taxonomy of
Fig. 4(a) presents the revised code that GPT-3.5 with zero-shot code change created by Tufano et al. [1] (the examples of the code
learning (no persona) correctly generates. In this figure, GPT-3.5 sug- changes in each category are depicted in Fig. 5):
gests an alternative way to initialize variable logArg. In contrast,
Fig. 4(b) presents the revised code that GPT-3.5 with zero-shot learning • fixing bug : The code changes in this category involve fixing bugs
(use persona) incorrectly generates. In this figure, GPT-3.5 suggests in the past. The following sub-categories are related to this cate-
changing the variable type from Boolean to boolean and changing gory: exception handling, conditional statement, lock mechanism,
the variable name from sample_name to sampleName in addition method return value, and method invocation.
to suggesting an alternative way to initialize variable logArg. • refactoring : The code changes in this category involve making
changes to code structure without changing the behavior of the
Fig. 4(c) presents another example of the revised code that GPT-
changed code. The following sub-categories are related to this
3.5 with few-shot learning (no persona) correctly generates. In this
category: inheritance; encapsulation; methods interaction; read-
figure, GPT-3.5 suggests a new if condition. On the contrary, Fig. 4(d)
ability; and renaming parameter, method and variable.
presents the revised code that GPT-3.5 with few-shot learning (use
• other: The code changes that cannot be classified as neither fixing
persona) incorrectly generates. In this figure, GPT-3.5 suggests an
bug nor refactoring will fall into this category.
additional if statement and an additional else block.
The above examples imply that when a persona is included in Fig. 6 presents the characteristics of the code changes (i.e., Fixing
prompts, GPT-3.5 tends to suggest additional incorrect changes to Bug, Refactoring and Other) of the revised code that GPT-3.5 correctly
the submitted code compared to when a persona is not included in generates. According to the figure, for the Tufanodata (with comment),
prompts. CodeReviewerdata and D-ACTdata dataset, the fine-tuning of GPT 3.5
The above results indicate that the best prompting strategy when with zero-shot learning helps GPT-3.5 achieves the highest EM for the
using GPT-3.5 without fine-tuning is few-shot learning without a per- code changes of Refactoring and Other. In contrast, for the Tufanodata
sona. (without comment) dataset, non fine-tuned GPT-3.5 with few-shot
learning achieves the highest EM for the code changes of all categories.
5. Discussion According to Figs. 6(a) and 6(b), we also find that fine-tuned GPT-3.5
and non fine-tuned GPT-3.5 that zero-shot and few-shot learning are
In this section, we discuss the implications of our findings, the used achieve the highest EM for the code changes of type other across
additional results of GPT-3.5, and the cost and benefits of using GPT- all studied datasets. The reason for this result is that we do not specify
3.5. the characteristics of the revised code (i.e., fixing bug and refactoring )
in prompts. Thus, such models possibly generate revised code that is
5.1. Implications of our findings not specific to fixing bug or refactoring. Therefore, the majority of the
code changes of the generated revised code are categorized as other.
GPT-3.5 does not require a lot of training data for model fine-
tuning to adapt to the code review automation task since Table 4 5.3. The impact of the size of training dataset on fine-tuned GPT-3.5
shows that GPT-3.5 that is fine-tuned on a subset of a training set
outperforms the studied code review automation approaches [4–6]. The The results of RQ2 show that model fine-tuning can help increase
results imply that GPT-3.5 can adapt to the code review automation the performance of GPT-3.5. However, little is known whether fine-
task by learning from a small set of training examples (approximately tuned GPT-3.5 can achieve higher performance when being fine-tuned
20k training examples in this study), unlike the studied code review with larger training sets. Therefore, we conduct experiments by using
8
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
Fig. 6. The EM achieved by GPT-3.5Zero-shot , GPT-3.5Few-shot and GPT-3.5Fine-tuned categorized by the types of code change. Here, GPT-3.5Zero-shot and GPT-3.5Few-shot refer to non
fine-tuned GPT-3.5 with zero-shot learning and few-shot learning, respectively. On the other hand, GPT-3.5Fine-tuned refers to fine-tuned GPT-3.5 with zero-shot learning.
Table 6
The evaluation results of GPT-3.5 when being fine-tuned with different sizes of training sets.
Size of CodeReviewer Tufano (with comment) Tufano (without comment) Android Google Ovirt
training set
EM CodeBLEU EM CodeBLEU EM CodeBLEU EM CodeBLEU EM CodeBLEU EM CodeBLEU
6% 37.93% 49.00% 22.16% 82.99% 6.02% 79.81% 2.34% 74.15% 6.71% 81.08% 3.05% 74.67%
10% 37.72% 48.83% 22.31% 83.43% 5.37% 80.68% 2.51% 75.10% 5.98% 80.65% 2.71% 75.13%
20% 38.80% 49.33% 22.84% 83.44% 5.65% 80.42% 2.34% 76.04% 7.52% 81.40% 2.83% 75.46%
Table 7
The evaluation results of GPT-3.5 for different prompt templates. P1 refers to the prompt templates with simple instructions (Fig. 3). P2 refers to the prompt templates with
instructions being broken down into smaller steps.(Fig. 7). P3 refers to the prompt templates with detailed instructions (Fig. 8).
Prompt Prompting CodeReviewer Tufano (with comment) Tufano (without comment) Android Google Ovirt
design
EM CodeBLEU EM CodeBLEU EM CodeBLEU EM CodeBLEU EM CodeBLEU EM CodeBLEU
P1 17.72% 44.17% 13.52% 78.36% 2.62% 74.92% 0.49% 61.85% 0.16% 61.04% 0.48% 56.55%
P2 Zero-shot 14.47% 43.52% 11.24% 79.05% 2.25% 76.54% 0.54% 66.10% 0.16% 67.07% 0.33% 60.76%
P3 11.94% 41.18% 9.86% 76.18% 1.26% 72.31% 0.05% 53.03% 0.08% 46.22% 0.26% 41.20%
P1 26.55% 47.50% 19.79% 81.47% 8.96% 79.21% 2.34% 75.33% 2.89% 81.40% 1.64% 73.83%
P2 Few-shot 25.25% 48.45% 15.82% 80.16% 6.84% 76.50% 0.60% 75.94% 3.56% 81.40% 1.67% 74.18%
P3 25.14% 48.60% 15.07% 79.83% 5.81% 75.76% 0.38% 74.95% 2.91% 81.18% 1.49% 73.31%
10% and 20% of training sets to fine-tune GPT-3.5 (we do not use persona in these prompt designs since the results of RQ3 show that GPT-
persona in these experiments). 3.5 without a persona outperforms GPT-3.5 with a persona). First, we
Table 6 shows the results of EM and CodeBLEU that fine-tuned GPT- use the prompt design that a single instruction is broken into smaller
3.5 achieves across different sizes of training sets. The table shows that steps,7 as depicted in Fig. 7. Second, we use the prompt design that
GPT-3.5 that is fine-tuned with 20% of a training set achieves 2.29%– contains more detailed instructions,8 as depicted in Fig. 8.
12.07% higher EM and 0.54%–2.55% higher CodeBLEU than GPT-3.5 Table 7 shows the results of EM and CodeBLEU that GPT-3.5
that is fine-tuned with 6% of a training set. In addition, GPT-3.5 that achieves across different prompt designs. The table shows that for zero-
is fine-tuned with 20% of a training set achieves 2.38%–25.75% higher shot learning, GPT-3.5 that is prompted by the prompt with a simple
EM and 0.01%–1.25% higher CodeBLEU than GPT-3.5 that is fine- instruction achieves 16.44%–45.45% higher EM than GPT-3.5 that is
tuned with 10% of a training set, respectively. The results indicate that prompted by the prompt with an instruction being broken down into
fine-tuned GPT-3.5 achieves higher performance when being fine-tuned smaller steps. In addition, GPT-3.5 that is prompted by the prompt
with a larger training set. with a simple instruction achieves 37.12%–880.00% higher EM than
GPT-3.5 that is prompted by the prompt with a detailed instruction.
The table also shows that for few-shot learning, GPT-3.5 that is
5.4. The impact of prompt design on GPT-3.5
prompted by the prompt with a simple instruction achieves 5.15%–
290.00% higher EM than GPT-3.5 that is prompted by the prompt
In RQ3, we use the prompt templates in Fig. 3 that contain sim-
ple instructions to conduct experiments. However, prior work [44,45]
found that the design of prompts has an impact on the performance 7
[Link]
of LLMs. Thus, we further investigate the impact of prompt design on advanced-prompt-engineering?pivots=programming-language-chat-
GPT-3.5 for code review automation. To do so, we conduct experiments completions#break-the-task-down
8
by using the following two new prompt designs (we do not include a [Link]
9
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
a simple instruction is the most suitable for GPT-3.5 for code review
automation.
5.5. Cost and benefits of using GPT-3.5 for code review automation
6. Threats to validity
10
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
[1] Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, Denys
Threats to internal validity relate to the randomness of GPT-3.5 and
Poshyvanyk, On learning meaningful code changes via neural machine
Magicoder, and the hyper-parameter settings that we use to fine-tune translation, in: Proceedings of ICSE, 2019, pp. 25–36.
GPT-3.5 and Magicoder. The results that we obtain from GPT-3.5 and [2] Rosalia Tufano, Luca Pascarella, Michele Tufano, Denys Poshyvanyk, Gabriele
Magicoder may vary due to the randomness of GPT-3.5 and Magicoder. Bavota, Towards automating code review activities, in: Proceedings of ICSE,
However, doing the same experiments multiple rounds can be expen- 2021, pp. 163–174.
[3] Patanamon Thongtanunam, Chanathip Pornprasit, Chakkrit Tantithamthavorn,
sive due to large testing datasets. Finally, we do not explore all possible AutoTransform: Automated code transformation to support modern code review
combinations of hyper-parameter settings (e.g., the number of epoch or process, in: Proceedings of ICSE, 2022, pp. 237–248.
learning rate) when fine-tuning GPT-3.5 and Magicoder. We do not do [4] Chanathip Pornprasit, Chakkrit Tantithamthavorn, Patanamon Thongtanunam,
so since the search space of hyper-parameter settings is large, which Chunyang Chen, D-ACT: Towards diff-aware code transformation for code review
under a time-wise evaluation, in: Proceedings of SANER, 2023, pp. 296–307.
can be expensive. Nonetheless, the main goal of this study is not to find [5] Zhiyu Li, Shuai Lu, Daya Guo, Nan Duan, Shailesh Jannu, Grant Jenks, Deep
the best hyper-parameter settings for code review automation, but to Majumder, Jared Green, Alexey Svyatkovskiy, Shengyu Fu, et al., Automating
investigate the performance of GPT-3.5 and Magicoder on code review code review activities by large-scale pre-training, in: Proceedings of ESEC/FSE,
automation tasks when using model fine-tuning. 2022, pp. 1035–1047.
[6] Rosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella, Denys
Poshyvanyk, Gabriele Bavota, Using pre-trained models to boost code review
6.3. Threats to external validity automation, in: Proceedings of ICSE, 2022, pp. 2291–2302.
[7] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones,
Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin, Attention is all you need, in:
Threats to external validity relate to the generalizability of our find- Proceedings of NIPS, 2017, pp. 5999–6009.
ings in other software projects. In this study, we conduct the experiment [8] Yue Wang, Weishi Wang, Shafiq Joty, Steven C.H. Hoi, CodeT5: Identifier-
with the dataset obtained from recent work [4–6]. However, the results aware unified pre-trained encoder-decoder models for code understanding and
of our experiment may not be generalized to other software projects. generation, in: Proceedings of EMNLP, 2021, pp. 8696–8708.
[9] J. Lu, L. Yu, X. Li, L. Yang, C. Zuo, Llama-reviewer: Advancing code review
Thus, other software projects can be explored in future work. automation with large language models through parameter-efficient fine-tuning,
Another threat relates to the updates to GPT-3.5 made by OpenAI in: Proceedings of ISSRE, 2023, pp. 647–658.
in future. Due to the updates, reproduced experiment results may differ [10] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne
from those reported in this paper. Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal
Azhar, et al., Llama: Open and efficient foundation language models, 2023, arXiv
preprint arXiv:2302.13971.
7. Conclusion [11] Shuzheng Gao, Xin-Cheng Wen, Cuiyun Gao, Wenxuan Wang, Hongyu Zhang,
Michael R. Lyu, What makes good in-context demonstrations for code intelligence
tasks with llms? in: Proceedings of ASE, 2023, pp. 761–773.
In this work, we investigate the performance of LLMs (i.e., GPT- [12] Shuzheng Gao, Xin-Cheng Wen, Cuiyun Gao, Wenxuan Wang, Michael R. Lyu,
3.5 and Magicoder) for code review automation when using model Constructing effective in-context demonstration for code intelligence tasks: An
fine-tuning and inference techniques (i.e., zero-shot learning, few-shot empirical study, in: Proceedings of ASE, 2023.
learning, and persona). We also compare the performance of the LLMs [13] Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry
Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, Douglas C. Schmidt, A prompt
with the existing code review automation approaches [4–6]. Our re- pattern catalog to enhance prompt engineering with chatgpt, 2023, arXiv preprint
sults show that (1) fine-tuned GPT-3.5 performs best for code review arXiv:2302.11382.
automation and (2) the best prompting strategy when using GPT-3.5 [14] Qi Guo, Junming Cao, Xiaofei Xie, Shangqing Liu, Xiaohong Li, Bihuan Chen,
without fine-tuning is few-shot learning without a persona. Based on Xin Peng, Exploring the potential of ChatGPT in automated code refinement: An
empirical study, 2023, arXiv preprint arXiv:2309.08221.
the results, we recommend that (1) LLMs for code review automation [15] Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav
should be fine-tuned to achieve the highest performance; and (2) when Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton,
data is not sufficient for model fine-tuning, few-shot learning without Sebastian Gehrmann, et al., Palm: Scaling language modeling with pathways,
a persona should be used for LLMs for code review automation. J. Mach. Learn. Res. (2023) 1–113.
[16] Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian
Lester, Nan Du, Andrew M. Dai, Quoc V. Le, Finetuned language models are
CRediT authorship contribution statement zero-shot learners, 2021, arXiv preprint arXiv:2109.01652.
[17] Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de
Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg
Chanathip Pornprasit: Writing – review & editing, Writing – orig- Brockman, et al., Evaluating large language models trained on code, 2021, arXiv
inal draft, Methodology, Data curation, Conceptualization. Chakkrit preprint arXiv:2107.03374.
Tantithamthavorn: Writing – review & editing, Supervision. [18] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan,
Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda
Askell, et al., Language models are few-shot learners, Proc. NeurIPS (2020)
Declaration of competing interest 1877–1901.
[19] Sungmin Kang, Juyeon Yoon, Shin Yoo, Large language models are few-shot
testers: Exploring llm-based general bug reproduction, in: Proceedings of ICSE,
The authors declare that they have no known competing finan-
2023, pp. 2312–2323.
cial interests or personal relationships that could have appeared to [20] Mingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang, Ge Li, Zhi Jin,
influence the work reported in this paper. Xiaoguang Mao, Xiangke Liao, Large language models are few-shot summarizers:
Multi-intent comment generation via in-context learning.
[21] Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, Lingming Zhang, Magicoder:
Data availability Source code is all you need, 2023, arXiv preprint arXiv:2312.02120.
[22] Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel
Data will be made available on request. Sundaresan, Ming Zhou, Ambrosio Blanco, Shuai Ma, Codebleu: a method for
automatic evaluation of code synthesis, 2020, arXiv preprint arXiv:2009.10297.
[23] Supplementary material, [Link]
Acknowledgment review-automatiton.
[24] Laura MacLeod, Michaela Greiler, Margaret-Anne Storey, Christian Bird, Jacek
Czerwonka, Code reviewing in the trenches: Challenges and best practices, IEEE
Chakkrit Tantithamthavorn was supported by the Australian Re-
Softw. (2017) 34–42.
search Council’s Discovery Early Career Researcher Award (DECRA) [25] Peter C. Rigby, Christian Bird, Convergent Contemporary Software Peer Review
funding scheme (DE200100941). Practices, in: Proceedings of ESEC/FSE, 2013, pp. 202–212.
11
C. Pornprasit and C. Tantithamthavorn Information and Software Technology 175 (2024) 107523
[26] Rosalia Tufan, Luca Pascarella, Michele Tufanoy, Denys Poshyvanykz, Gabriele [36] Malinda Dilhara, Abhiram Bellur, Timofey Bryksin, Danny Dig, Unprecedented
Bavota, Towards automating code review activities, in: Proceedings of ICSE, code change automation: The fusion of LLMs and transformation by example,
2021, pp. 1479–1482. 2024, arXiv preprint arXiv:2402.07138.
[27] Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiao- [37] Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, Towards
qing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al., Code AI-assisted synthesis of verified dafny methods, 2024, arXiv preprint arXiv:
llama: Open foundation models for code, 2023, arXiv preprint arXiv:2308.12950. 2402.00247.
[28] Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, [38] Zhihan Jiang, Jinyang Liu, Zhuangbin Chen, Yichen Li, Junjie Huang, Yintong
Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al., Huo, Pinjia He, Jiazhen Gu, Michael R. Lyu, Llmparser: A llm-based log parsing
Starcoder: may the source be with you!, 2023, arXiv preprint arXiv:2305.06161. framework, 2023, arXiv preprint arXiv:2310.01796.
[29] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, [39] Xin Zhou, Kisub Kim, Bowen Xu, DongGyun Han, Junda He, David Lo,
Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Generation-based code review automation: How far are we? 2023, arXiv preprint
Bhosale, et al., Llama 2: Open foundation and fine-tuned chat models, 2023, arXiv:2303.07221.
arXiv preprint arXiv:2307.09288. [40] Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank
[30] Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, Weizhu Wang, Kwang-Ting Cheng, Min-Hung Chen, DoRA: Weight-decomposed low-rank
Chen, What makes good in-context examples for GPT-3? 2021, arXiv preprint adaptation, 2024, arXiv preprint arXiv:2402.09353.
arXiv:2101.06804. [41] Stephen Robertson, Hugo Zaragoza, et al., The probabilistic relevance framework:
[31] Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun BM25 and beyond, Found. Trends Inf. Retriev. (2009) 333–389.
Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al., Larger language models do [42] Zhiqiang Yuan, Junwei Liu, Qiancheng Zi, Mingwei Liu, Xin Peng, Yiling Lou,
in-context learning differently, 2023, arXiv preprint arXiv:2303.03846. Evaluating instruction-tuned large language models on code comprehension and
[32] Seungone Kim, Se June Joo, Doyoung Kim, Joel Jang, Seonghyeon Ye, Jamin generation, 2023, arXiv preprint arXiv:2308.01240.
Shin, Minjoon Seo, The cot collection: Improving zero-shot and few-shot learning [43] Kishore Papineni, Salim Roukos, Todd Ward, Wei-Jing Zhu, BLEU: A method for
of language models via chain-of-thought fine-tuning, 2023, arXiv preprint arXiv: automatic evaluation of machine translation, in: Proceedings of ACL, 2002, pp.
2305.14045. 311–318.
[33] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, [44] David OBrien, Sumon Biswas, Sayem Mohammad Imtiaz, Rabe Abdalkareem,
Quoc V. Le, Denny Zhou, et al., Chain-of-thought prompting elicits reasoning in Emad Shihab, Hridesh Rajan, Are prompt engineering and TODO comments
large language models, Proc. NeurIPS (2022) 24824–24837. friends or foes? An evaluation on GitHub copilot, in: Proceedings of ICSE, 2024,
[34] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, pp. 1–13.
Aakanksha Chowdhery, Denny Zhou, Self-consistency improves chain of thought [45] Soneya Binta Hossain, Nan Jiang, Qiang Zhou, Xiaopeng Li, Wen-Hao Chiang,
reasoning in language models, 2022, arXiv preprint arXiv:2203.11171. Yingjun Lyu, Hoan Nguyen, Omer Tripp, A deep dive into large language
[35] Junjielong Xu, Ruichun Yang, Yintong Huo, Chengyu Zhang, Pinjia He, Prompt- models for automated bug localization and repair, 2024, arXiv preprint arXiv:
ing for automatic log template extraction, 2023, arXiv preprint arXiv:2307. 2404.11595.
09950.
12
Fine-tuning is essential because it allows GPT-3.5 to adapt specifically to the code review automation task, learning the relationship between inputs and outputs from training examples. This results in a 63.91%–1100% increase in Exact Match compared to non-fine-tuned models, as it can better adapt to specific task characteristics and generate more accurate code revisions .
Code review automation can significantly benefit from LLM-based approaches by optimizing the accuracy and efficiency of code revisions. The study demonstrates that leveraging fine-tuned LLMs, particularly GPT-3.5 with few-shot learning, leads to higher Exact Match and CodeBLEU scores, ensuring more precise and contextually appropriate code outputs. These advancements allow for more robust automation, reducing the time and manual effort required by human reviewers .
The study expands on Guo et al.'s work by systematically evaluating fine-tuning and various prompting strategies, including few-shot learning, which were previously unexplored. Through experimental comparisons and assessments, this study provides a broader and more nuanced understanding of what constitutes the most effective approach for leveraging LLMs in code review automation, addressing gaps in previous research .
The guiding principles for prompting strategies in software engineering involve using instruction-based methods like zero-shot, few-shot learning, and avoiding persona prompting. Few-shot learning is particularly effective as it provides more context for the model, improving accuracy through examples. This strategy is preferred when data is insufficient for fine-tuning, as it balances learning efficiency and performance .
Exact Match (EM) and CodeBLEU are crucial evaluation measures for assessing LLM performance in code review automation. EM calculates the percentage of outputs that match the expected outputs exactly, focusing on precision. CodeBLEU, on the other hand, evaluates the quality of code synthesis based on BLEU metrics adapted for coding tasks. Together, they provide a comprehensive view of precision and the structural quality of generated code, enabling nuanced performance assessment .
The experiments show that while both GPT-3.5 and Magicoder are effective for code review automation, fine-tuning GPT-3.5 is crucial for achieving superior Exact Match scores in comparison. The fine-tuned GPT-3.5 shows a 73.17%–74.23% improvement over the non-fine-tuned approaches, highlighting the importance of tailoring models through fine-tuning for optimal performance .
GPT-3.5 with few-shot learning significantly outperforms zero-shot learning in generating revised code for code review automation. It achieves 46.38%–659.09% higher Exact Match scores and 3.97%–33.36% higher CodeBLEU scores compared to zero-shot learning without a persona. With a persona, these improvements are 53.95%–659.09% and 5.55%–60.27%, respectively. These results suggest that few-shot learning provides more context through demonstration examples, guiding the model to generate more accurate output .
When insufficient data is available for fine-tuning LLMs, the recommended approach is to employ few-shot learning without a persona. This method provides necessary guidance through demonstration examples, ensuring that LLMs maintain high performance in code review tasks without the need for extensive data typically required for full model fine-tuning .
Including a persona in input prompts has a negative impact on the efficacy of GPT-3.5, resulting in 1.02%–54.17% lower Exact Match and 1.33%–19.13% lower CodeBLEU scores compared to not using a persona. The results indicate that persona inclusion complicates the model's task of generating correct revised code, likely because it adds an unnecessary layer of abstraction or complexity not beneficial for the task .
Instruction-based prompting strategies, such as zero-shot and few-shot learning, play a pivotal role in software engineering by facilitating efficient language model guidance without extensive fine-tuning. These methods adapt models to specific tasks, enabling high-quality performance in environments where comprehensive training data may be unavailable. They enhance model adaptability and accuracy by providing contextually informative prompts tailored to the task requirements .