Efficiency
Efficiency
LLM Reasoning
006 often employ suboptimal reward shaping strate- degrading model performance (Turpin et al., 2024; 048
007 gies that erroneously penalize essential long- Cuadron et al., 2025; Shen et al., 2025). Therefore, 049
008 form reasoning. Consequently, these strate- improving the reasoning efficiency of LLMs has 050
009 gies hinder the deep cognitive processing re- emerged as an essential research focus. To mit- 051
010 quired to solve complex problems. Besides, igate this issue, existing approaches are broadly 052
011 these methods often struggle to adapt to the
categorized into two types: fixed-budget reasoning 053
012 varying cognitive demands of samples across
013 different difficulty levels. To address these lim-
methods and adaptive reasoning methods. Notwith- 054
014 itations, we present Cognition-Guided Policy standing their potential, these methods suffer from 055
015 Optimization (CGPO), which consists of Cog- two fundamental limitations: 056
016 nitive Utility Reward (CUR) and Cognition- First, suboptimal reward shaping strategies 057
017 Adaptive Regulation (CAR). CUR designs a erroneously penalize essential long-form rea- 058
018 multiplicative reward that scales correctness soning, misaligning efficiency with the cogni- 059
019 rewards using a non-linear length penalty to tive depth needed for complex problem-solving. 060
020 reduce redundancy. CAR adaptively adjusts
While current RL-based methods aim to promote 061
021 the strength of KL regularization based on real-
022 time sample difficulty. Extensive experiments conciseness, their reward shaping strategies often 062
023 on nine datasets across three reasoning tasks exhibit inherent flaws. A typical paradigm adopts 063
024 confirm that CGPO effectively balances effi- additive formulations wherein length-based penal- 064
025 ciency with reasoning accuracy. For instance, ties are subtracted from the correctness rewards. 065
026 on mathematical reasoning benchmarks using However, this design often yields negative learning 066
027 the DeepScaleR-Preview-1.5B model, CGPO signals for valid but verbose responses, particu- 067
028 surpasses other methods by 0.8 to 3.2 points
larly when integrated with group-relative normal- 068
029 in the average Pass@1 score, while reducing
030 token consumption by 0.7% to 38.4%.
ization (Xiang et al., 2025; Huang et al., 2025). 069
Moreover, some multiplicative structures applying 070
1
083 tain training stability, RL-based algorithms such as 2 Related Work 130
084 GRPO incorporate a KL-divergence penalty con-
Recent advancements in LLMs have spurred exten- 131
085 trolled by a fixed coefficient β. Nevertheless, this
sive research aimed at enhancing their reasoning 132
086 static strategy remains insensitive to variations in
efficiency, primarily by addressing issues of ex- 133
087 sample difficulty. For simple tasks with low cog-
cessive verbosity and unnecessary reasoning steps. 134
088 nitive demands, a fixed β is often insufficient to
Prevailing strategies can be broadly classified into 135
089 enforce the consistency needed to alleviate policy
two main categories: fixed budget reasoning meth- 136
090 degradation. Conversely, for challenging problems
ods and adaptive reasoning methods. 137
091 requiring extensive cognitive effort, this fixed coef-
Fixed Budget Reasoning Methods. These meth- 138
092 ficient becomes excessively restrictive, thereby hin-
ods aim to reduce token usage by imposing explicit 139
093 dering the exploration necessary to discover novel
constraints on output length. Common strategies 140
094 reasoning paths.
involve setting hard token budgets (Aggarwal and 141
095 Motivated by these limitations, this study centers
Welleck, 2025; Xu et al., 2025) or leveraging hi- 142
096 on the following core objective:
erarchical structures (Qi et al., 2025; Lyu et al., 143
2025; Hou et al., 2025) to constrain the reasoning 144
How to design a cognition-guided frame- process. Nonetheless, these methods often rely on 145
work that effectively balances efficiency manually specified budgets that fail to account for 146
with reasoning accuracy? problem complexity, complicating the selection of 147
097
appropriate limits for tasks of varying difficulty. 148
098 To achieve this goal, we present Cognition- Adaptive Reasoning Methods. In contrast, these 149
099 Guided Policy Optimization (CGPO), a novel approaches enable models to autonomously ad- 150
100 RL-based framework developed to balance reason- just their computational effort based on problem 151
101 ing accuracy with efficiency. CGPO integrates requirements. One prominent strategy is reward 152
102 two key components: (1) Cognitive Utility Re- shaping to discourage verbosity. A common imple- 153
103 ward (CUR): CUR formulates a non-linear mul- mentation is through an additive penalty, where a 154
104 tiplicative objective that treats token consumption length-based cost is subtracted from the task reward 155
105 as a cognitive cost, ensuring efficiency is achieved (Yi et al., 2025; Xiang et al., 2025; Huang et al., 156
106 without compromising accuracy. (2) Cognition- 2025). However, this additive structure is flawed as 157
107 Adaptive Regulation (CAR): CAR implements it can corrupt the learning signal by assigning neg- 158
108 a difficulty-aware mechanism that adaptively ad- ative advantages to valid but long reasoning paths. 159
109 justs the KL penalty to match varying sample Critically, even alternative designs such as the hard- 160
110 difficulty. This encourages exploration on com- threshold multiplicative bonus (Liu et al., 2025) can 161
111 plex problems while maintaining reasoning perfor- suffer from similar issues by creating a polarized 162
112 mance on simple ones. Collectively, these modules reward distribution that also leads to suppressive 163
113 empower CGPO to achieve a superior accuracy- signals. Another common approach is binary mode 164
114 efficiency trade-off, and extensive experiments on selection (Fang et al., 2025; Zhang et al., 2025; Tu 165
115 nine datasets across three reasoning tasks confirm et al., 2025), where models learn to switch between 166
116 its superiority over state-of-the-art baselines. “thinking” and “non-thinking” modes. While ef- 167
117 Our contributions are summarized as follows: fective for simple problems, this coarse-grained 168
118 (1) We identify two deficiencies in current meth- decision-making fails to accommodate the continu- 169
119 ods for efficient reasoning: suboptimal reward shap- ous spectrum of reasoning complexity. 170
122 (2) We introduce CGPO, a novel framework that The results in Table 1 and Figure 1 reveal criti- 172
123 integrates a multiplicative reward function with a cal limitations in existing reward shaping strate- 173
124 difficulty-adaptive mechanism to balance accuracy gies. While effective in mitigating verbosity, these 174
125 and efficiency. methods often rely on suboptimal reward structures 175
126 (3) Empirical results on nine datasets covering that inadvertently yield misleading learning signals. 176
127 three reasoning tasks showcase that CGPO estab- Specifically, additive penalties (e.g., ALP, HAPO) 177
128 lishes a superior accuracy-efficiency trade-off, out- tend to assign negative advantages to correct re- 178
129 performing other baselines. sponses solely due to length, thereby penalizing 179
2
Method Reward Formulation
ALP (Xiang et al., 2025) Ri − β · |oi | · max(mean{Ri }, K −1 )
|oi | |oi |
HAPO (Huang et al., 2025) Ri + w · max cos(min( π2 h(q) , π)), c Ri + w · min cos(min( π2 h(q) , π)), 0 (1 − Ri )
Table 1: Examples of reward designs for efficient reasoning, where Ri = I(oi is correct) ∈ {0, 1}.
3
Cognitive Utility Reward (CUR)
Length Sensor
Accuracy Conciseness
Cognitive
Utility Score
Reasoning
Path Correctness Multiplicative Efficiency Filter
Verifier
Policy Optimization
User Prompt
Update
LLM
Reasoning & Adaptive Cognitive Regulation (ACR)
Answer
HARD Looser
Constraint
(Explore More)
EASY HARD
Tighter
EASY
Constraint Adaptive
Task Difficulty
(Stay Stable)
Estimator Exploration
Batch Performance
Dynamic KL Regulator Constraint
Figure 2: Illustration of CGPO, which is composed of Cognitive Utility Reward (CUR) and Cognition-Adaptive
Regulation (CAR).
235 cay caused by noise accumulation (Shannon, 1948) To effectively guide policy updates, these util- 262
236 and error propagation (Lee and Cummins, 2004; ity scores Ui are used to compute the group- 263
237 Chen et al., 2015). The proposed linear marginal normalized advantage: 264
238 cost function offers a first-order approximation of
Ui − µU
239 the escalating cognitive load, modeling each suc- Âi = (6) 265
σU + ϵ
240 cessive reasoning step as increasingly computation-
241 ally demanding. where µU and σU are the mean and standard de- 266
242 The efficiency factor E(Li ) is derived by setting viation of {Ui }Gi=1 , and ϵ is a small constant for 267
243 its relative decay rate equal to γ(Li ): numerical stability. By using this utility-based ad- 268
vantage, the model learns to prioritize responses 269
dE(Li )
244 = −γ(Li )E(Li ) (2) that are both correct and efficient. Further analysis 270
dLi of CUR is shown in Appendix D.2. 271
245 Subject to the boundary condition E(0) = 1 (i.e.,
4.2 Cognition-Adaptive Regulation (CAR) 272
246 full efficiency at zero length), the solution to Eq.
247 (2) yields (see Appendix C for detailed derivation): A static KL-divergence coefficient β cannot adapt 273
248 to the varying cognitive demands of different sam- 274
1 2 ples. This leads to overly loose exploration for 275
249 E(Li ) = exp − γ0 Li + kLi (3)
2 simple problems and excessively restrictive explo- 276
250 The final utility Ui for oi is computed as the prod- ration for complex ones. To mitigate this, we pro- 277
251 uct of the correctness reward Ri and the efficiency pose Cognition-Adaptive Regulation (CAR) that 278
252 factor. For notational simplicity, the quadratic cost adaptively modulates KL regularization strength 279
253 coefficient is defined as η = k/2: based on real-time estimates of sample difficulty. 280
For each training batch, CAR operates as follows: 281
Ui = Ri · exp −(γ0 Li + ηL2i )
254 (4) Step 1: Adaptive Difficulty Estimation. To 282
approximate the cognitive demand of each sam- 283
255 In our implementation, γ0 is set to 0 to focus
ple, the difficulty d(qj ) is quantified as the model’s 284
256 solely on the accelerating cost component. More-
failure rate across G generated responses: 285
257 over, the length is normalized by the maximum
258 sequence length Lmax within each batch to account G
1 X
259 for variation in sequence lengths. This yields the d(qj ) = 1 − Ri (7) 286
G
260 final CUR formulation: i=1
4
289 a more informative signal than absolute difficulty 5.1 Settings 328
290 scores. For instance, a problem that is simple in ab- Models. In this study, DeepSeek-R1-Distill-Qwen- 329
291 solute terms may be the most challenging within a 1.5B (Guo et al., 2025), DeepSeek-R1-Distill- 330
292 batch of trivial examples. Accordingly, a standard- Qwen-7B (Guo et al., 2025), and DeepScaleR- 331
293 ized difficulty score zj is derived for each sample: Preview-1.5B (Luo et al., 2025) were chosen for 332
294
mathematical reasoning. For SQL generation and 333
d(qj ) − µd
295 zj = (8) multi-modal reasoning tasks, we applied Qwen2.5- 334
σd + ϵ Coder-3B-Instruct (Hui et al., 2024) and Qwen2.5- 335
296 where µd and σd are the batch mean and standard VL-3B-Instruct (Bai et al., 2025), respectively. 336
297 deviation of {d(qj )}, respectively. Datasets. In this work, the MATH dataset 337
298 Step 3: Difficulty-Aware KL Coefficient Map- (Hendrycks et al., 2021) was used for training in 338
299 ping. zj is mapped to an adaptive KL-divergence mathematical reasoning. Evaluation was carried 339
300 coefficient βada using a tanh function to establish out across widely-used benchmarks, encompass- 340
301 a smooth, bounded, and inverse relationship (i.e., ing MATH500 (Hendrycks et al., 2021), AIME24 341
302 reduced penalties for harder samples): (Li et al., 2024a), AIME25 (Codeforces), AMC23 342
(Ouyang et al., 2022), Minerva (Lewkowycz et al., 343
303 βada = βbase · (1 − tanh(zj )) (9) 2022), and OlympiadBench (Huang et al., 2024). 344
For SQL generation, we trained models using 345
BIRD-Train (Li et al., 2024b) and tested on Spider- 346
304 where βbase serves as a scaling factor. This mech-
Dev (Yu et al., 2018) and BIRD-Dev (Li et al., 347
305 anism allows CAR to effectively allocate the cog-
2024b). Multi-modal reasoning experiments relied 348
306 nitive budget. It relaxes constraints to encourage
on Geometry3K (Lu et al., 2021), which included 349
307 exploration on challenging samples while maintain-
dedicated training and test subsets. Additional de- 350
308 ing reasoning performance on simpler ones. Fur-
tails are provided in Appendix A. 351
309 ther analysis on Eq. (9) is shown in Appendix D.8.
Implementation Details. CGPO was compared 352
310 Step 4: Final Objective Function. Finally,
against the base model, GRPO (Shao et al., 2024), 353
311 βada is integrated into the standard GRPO objective
AdaptThink (Zhang et al., 2025), HAPO (Huang 354
312 (Shao et al., 2024) as follows:
et al., 2025), LASER (Liu et al., 2025), and ALP 355
G |oi | (Xiang et al., 2025) (See Appendix B for more de- 356
h1 X 1 X
L(θ) =E min ri,t (θ)Âi , tails about the baselines). η, βbase , batch size, the 357
G |oi | number of rollouts, and the sampling temperature 358
i=1 t=1
313 (10) were set to 1.0, 0.001, 512, 8, and 1.0, respectively. 359
clip(ri,t (θ), 1 − ϵ, 1 + ϵ)Âi AdamW (Zhou et al., 2024) was used with a learn- 360
ing rate of 1 × 10−6 and we trained each model for
i
361
− βada DKL (πθ ∥πref )
400 steps. For evaluation, the sampling tempera- 362
ture was 0.6. We used Pass@1 score and average 363
314 where ri,t (θ) is the importance sampling ratio. The token length as evaluation metrics. Notably, the 364
315 policy optimization is then guided by the CUR- Pass@1 score denoted execution accuracy in SQL 365
316 based advantage Âi and regularized by the adaptive generation. Computing resources comprised eight 366
317 KL-divergence penalty from CAR. NVIDIA GeForce A100 80GB GPUs. 367
319 Our experiments are designed to investigate four 5.2.1 Main Results (RQ1) 369
320 core research questions regarding the performance Table 3 showed that CGPO established a superior 370
321 and efficiency of CGPO: RQ1: How does CGPO accuracy-efficiency trade-off on mathematical rea- 371
322 perform against state-of-the-art RLVR baselines? soning. For instance, using DeepScaleR-Preview- 372
323 RQ2: Does CGPO learn to adapt its reasoning 1.5B, CGPO attained the highest average Pass@1 373
324 depth based on task difficulty? RQ3: How do the score of 49.8, surpassing GRPO by 1.3 points while 374
325 individual components (CUR and CAR) contribute reducing the average response length by 29.1%. 375
326 to the final results? RQ4: What is the impact of This improvement in accuracy and efficiency high- 376
327 hyperparameters η and βbase on model sensitivity? lighted the superior optimization strategy of CGPO. 377
5
Math500 AIME24 AIME25 AMC23 Minerva Olympiad Avg.
Methods
Pass@1 Length Pass@1 Length Pass@1 Length Pass@1 Length Pass@1 Length Pass@1 Length Pass@1 Length
DeepSeek-R1-Distill-Qwen-1.5B
Base 77.6 4086 16.7 8024 16.7 7921 62.5 7628 27.6 4821 42.4 7638 40.6 6786
GRPO 81.8 3106 20.0 7626 20.0 7231 67.5 5254 30.5 3845 47.6 5832 44.6 5482
AdaptThink 79.6 1728 23.3 6628 23.3 7014 70.0 3823 29.8 3248 45.5 4723 45.3 4527
HAPO 81.4 2601 23.3 7469 23.3 7217 72.5 4371 30.2 3472 46.6 4812 46.2 4990
LASER 81.8 2426 30.0 7225 26.7 7523 72.5 4285 31.3 2873 48.2 4626 48.4 4826
ALP 79.2 2241 26.7 6671 30.0 6892 70.0 4176 30.9 2447 46.7 5094 47.3 4587
CGPO 82.6 2180 30.0 6662 30.0 6826 72.5 4049 31.6 2835 48.7 4646 49.2 4533
DeepSeek-R1-Distill-Qwen-7B
Base 91.6 2836 50.0 9613 50.0 9425 77.5 5617 36.8 4673 63.1 6682 61.5 6474
GRPO 91.6 2654 50.0 8872 50.0 8545 80.0 5211 39.7 4153 64.2 6109 62.6 5924
AdaptThink 89.4 1952 53.3 7104 53.3 7315 80.0 3908 38.6 3105 64.5 4882 63.2 4711
HAPO 91.8 2413 53.3 7821 53.3 7534 82.5 4412 40.1 3519 65.9 5327 64.5 5171
LASER 92.0 2305 56.7 7618 56.7 7822 85.0 4306 40.8 3054 66.8 5113 66.3 5036
ALP 91.6 2108 53.3 7415 56.7 7429 82.5 4117 40.4 2956 66.5 4945 65.2 4828
CGPO 92.2 2083 56.7 7258 56.7 7243 85.0 4056 41.2 2912 67.2 4867 66.5 4737
DeepScaleR-Preview-1.5B
Base 86.4 3063 23.3 8924 23.3 9127 65.0 4716 34.6 4261 51.8 5268 47.4 5893
GRPO 85.6 2736 26.7 7612 26.7 8162 65.0 4124 34.2 3562 52.7 4528 48.5 5121
AdaptThink 84.8 1448 26.7 6129 26.7 6368 67.5 3127 32.7 2529 50.6 3612 48.2 3869
HAPO 85.0 2434 26.7 6836 26.7 7023 65.0 3756 30.9 3268 48.2 4176 46.6 4582
LASER 85.4 2263 30.0 5926 26.7 6147 70.0 3548 32.0 3111 53.3 3925 49.0 4153
ALP 85.2 2052 26.7 5457 30.0 5612 70.0 2852 31.6 2427 47.9 3551 48.6 3659
CGPO 85.6 1632 26.7 5398 30.0 5624 70.0 2811 33.5 2674 53.1 3660 49.8 3633
Table 3: Experimental results of different baselines across various mathematical reasoning benchmarks in terms of
Pass@1 score and reasoning length. The best results are highlighted in bold.
378 Moreover, the versatility of CGPO was showcased lution). Figure 5 revealed that CGPO achieved 401
379 on SQL generation and multi-modal reasoning lower CpA than GRPO. Crucially, this efficiency 402
380 tasks (Figure 3 and Table 5). Notably, it achieved gap diverged as task difficulty increased, indicating 403
381 the highest Pass@1 while maintaining superior to- that CGPO invested its computational budget more 404
382 ken efficiency compared to most baselines. These effectively than GRPO. 405
383 gains across diverse domains affirmed the broad Additionally, Figure 4(d) highlighted tangible 406
384 applicability and robustness of CGPO. benefits in inference efficiency, where the inference 407
385 5.2.2 Analysis of Adaptive Reasoning (RQ2) latency was lower for CGPO. This observation con- 408
firmed that shortened reasoning paths translated 409
386 Figure 4(a) showed that CGPO exhibited difficulty-
directly into lower end-to-end latency. Regard- 410
387 aware resource allocation, where the model adap-
ing training overhead, the total wall-clock times 411
388 tively adjusted its reasoning depth based on sam-
for GRPO and CGPO using DeepSeek-R1-Distill- 412
389 ple difficulty. This adaptability was attributed to
Qwen-1.5B were 649 and 654 minutes, respectively. 413
390 the synergy between CUR and CAR. CUR pro-
Consequently, CGPO achieved substantial gains in 414
391 vided a global pressure for conciseness, while CAR
both inference speed and reasoning performance 415
392 granted the local flexibility to tackle complex prob-
with negligible additional training costs. 416
393 lems. Meanwhile, the trend of decreasing length
394 and increasing reward (Figures 4(b) and (c)) con-
5.2.3 Ablation Study (RQ3) 417
395 firmed that CGPO effectively balanced accuracy
396 and efficiency. To quantify this behavior, we intro- Synergistic Effect of CUR and CAR. Fig- 418
397 duce the Cost per Accuracy (CpA) metric: ure 6(a) highlights the synergistic roles of CUR and 419
CAR. While both modules surpassed the GRPO 420
Average Length
398 CpA = (11) baseline, CUR exhibited a greater standalone im- 421
Pass@1 provement. This was because CUR structurally 422
399 where a lower value denotes higher reasoning effi- corrected the learning signal and produced a bet- 423
400 ciency (i.e., fewer tokens required per correct so- ter optimization landscape that benefited accuracy 424
6
6 6
3 6 B a s e B a s e
G R P O G R P O
A d a p tT h in k 6 4 A d a p tT h in k
3 4 H A P O H A P O
L A S E R L A S E R
3 2 A L P A L P
6 2
C G P O C G P O
1
1
P a s s @
P a s s @
3 0
6 0
2 8
5 8
2 6
2 4 5 6
3 0 0 3 5 0 4 0 0 4 5 0 5 0 0 5 5 0 3 2 0 3 4 0 3 6 0 3 8 0 4 0 0 4 2 0 4 4 0 4 6 0
L e n g th L e n g th
(a ) (b )
Figure 3: Experimental results on (a) SQL generation and (b) multi-modal reasoning tasks.
G R P O L e n g th C G P O L e n g th
G R P O P a s s @ 1 C G P O P a s s @ 1
4 5 0 0 4 4
R e w a rd 1 .0 R e w a rd
9 5 D S - G R P O
4 0 0 0 0 .8 4
0 .8
D S - C G P O
4 2
R e w a rd
R e w a rd
3 5 0 0 0 .6 3
D S R -G R P O
T e s tin g In fe r e n c e T im e
9 0
0 .6 D S R -C G P O
3 0 0 0
0 .4 2 4 0
8 5 0 .4
1
2 5 0 0
L e n g th
0 .2 1
P a s s @
2 0 0 0 2 7 6 0 L e n g th L e n g th 3 8
8 0 2 3 5 0
1 5 0 0 2 3 0 0
L e n g th
L e n g th
7 5 1 8 8 0 3 6
1 0 0 0 1 8 4 0
1 4 1 0
5 0 0 1 3 8 0 3 4
7 0
9 4 0
0
L e v e l 1 L e v e l 2 L e v e l 3 L e v e l 4 L e v e l 5 0 1 0 0 2 0 0 3 0 0 4 0 0 0 1 0 0 2 0 0 3 0 0 4 0 0 0 1 0 0 2 0 0 3 0 0 4 0 0
D iffic u lty L e v e l T r a in in g S te p T r a in in g S te p T r a in in g S te p
(a ) (b ) (c ) (d )
Figure 4: Evaluation of reasoning efficiency and training dynamics. (a) Average token length and Pass@1 on
MATH500 stratified by problem difficulty. (b-c) Evolution of average reward and response length during training for
DeepSeek-R1-Distill-Qwen-1.5B and DeepScaleR-Preview-1.5B. (d) Evolution of wall-clock inference latency on
MATH500. DS and DSR represent DeepSeek-R1-Distill-Qwen-1.5B and DeepScaleR-Preview-1.5B, respectively.
6 0 0 0
G R P O Impact of Reward Formulation. To validate 432
C G P O CUR’s design, we compared it with two alternative 433
5 0 0 0
reward formulations, including Additive Penalty 434
and Rational Quadratic Decay (Figure 6(b)): 435
4 0 0 0
L2i
C p A
7
D S -P a s s @ 1 D S R -P a s s @ 1 D S -P a s s @ 1 D S R -P a s s @ 1
D S -L e n g th D S R -L e n g th D S -L e n g th D S R -L e n g th
5 6 0 0 4 8 0 0
5 0 5 4 0 0 5 0
4 6 0 0
5 2 0 0
4 0 5 0 0 0 4 0
4 4 0 0
4 8 0 0
1
1
L e n g th
L e n g th
3 0 4 6 0 0 3 0 4 2 0 0
P a s s @
P a s s @
4 4 0 0
2 0 2 0 4 0 0 0
4 2 0 0
4 0 0 0 3 8 0 0
1 0 1 0
3 8 0 0
3 6 0 0
3 6 0 0
0 0
G R P O + A C R + C U R C G P O A d d itiv e R a tio n a l C U R
S e ttin g M e th o d
(a ) (b )
D S -P a s s @ 1 D S R -P a s s @ 1 D S -P a s s @ 1 D S R -P a s s @ 1
D S -L e n g th D S R -L e n g th D S -L e n g th D S R -L e n g th
4 8 0 0
5 0 4 6 0 0 5 0
4 6 0 0
4 4 0 0
4 0 4 0 4 4 0 0
4 2 0 0
4 2 0 0
1
1
L e n g th
L e n g th
3 0 3 0
P a s s @
P a s s @
4 0 0 0
4 0 0 0
2 0 2 0
3 8 0 0 3 8 0 0
1 0 3 6 0 0 1 0
3 6 0 0
0 3 4 0 0 0 3 4 0 0
A d d itiv e R a tio n a l C U R 0 .0 0 0 0 .0 0 1 0 .0 0 5 0 .0 1 0
(c ) (d )
Figure 6: Ablation studies and sensitivity analysis. (a) Ablation study on the components in CGPO. (b) Ablation
study on different reward formulations. (c) Sensitivity analysis on η. (d) Sensitivity analysis on βbase . Note that DS
and DSR indicate DeepSeek-R1-Distill-Qwen-1.5B and DeepScaleR-Preview-1.5B, respectively.
451 accounted for CUR’s superior performance. (For optimal performance. Concretely, setting βbase = 0 474
452 more results and analysis of reward formulations yielded suboptimal results, indicating the need for 475
453 and the normalization technique, see Table 9, Ap- baseline regularization to promote stable learning. 476
454 pendix D.5, and Appendix D.6). In contrast, a high value (e.g., βbase = 0.01) also 477
degraded performance by excessively constrain- 478
455 5.2.4 Sensitivity Analysis (RQ4) ing exploration. Notably, βbase = 0.001 achieved 479
456 Impact of η. Figure 6(c) examines the effects the highest average Pass@1, confirming it as an 480
457 of the efficiency coefficient η in CUR, revealing effective anchor for adaptive regulation scheme. 481
458 a clear accuracy-efficiency trade-off. Concretely, Additional results and analysis are shown in Table 482
459 a low value (η = 0.5), representing a weaker 7 and Appendix D.10, respectively. 483
460 penalty on length, yielded good accuracy but with
461 the longest response lengths. Conversely, a high
462 value (η = 1.5) aggressively shortened responses,
6 Conclusion 484
465 highest accuracy while reducing response length tion (CGPO) to resolve the tension between reason- 486
466 compared to η = 0.5, validating it as the optimal ing depth and computational cost in LLMs. Depart- 487
467 choice. Detailed results and analysis are available ing from the static exploration strategies of prior 488
468 in Table 8 and Appendix D.10, respectively. work, our approach harmonizes Cognitive Utility 489
Reward (CUR) and Cognition-Adaptive Regulation 490
469 Impact of βbase . Figure 6(d) shows the impact (CAR) within a unified RL-based framework. The 491
470 of the base KL coefficient βbase , which serves as experimental results conclusively demonstrate that 492
471 the anchor point for our CAR mechanism. While CGPO generates reasoning paths that are both more 493
472 CAR demonstrates robustness across a range of accurate and significantly more concise, offering a 494
473 values, the choice of βbase is crucial for achieving scalable path forward for efficient reasoning. 495
8
496 Limitations MAA Codeforces. American invitational mathematics 545
examination-aime 2024, 2024. 546
497 We acknowledge two primary constraints in our cur-
498 rent study. (1) CGPO utilizes token length as the Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao 547
Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, 548
499 primary proxy for cost, which may inadvertently Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, et al. 549
500 penalize necessary elaborations or self-corrections 2025. The danger of overthinking: Examining the 550
501 alongside redundancy. The method currently lacks reasoning-action dilemma in agentic tasks. arXiv 551
502 the semantic granularity to distinguish between su- preprint arXiv:2502.08235. 552
503 perfluous verbosity and valuable reasoning details, Gongfan Fang, Xinyin Ma, and Xinchao Wang. 2025. 553
504 potentially sacrificing interpretability for brevity. Thinkless: Llm learns when to think. arXiv preprint 554
505 (2) While effective in the difficulty estimation of arXiv:2505.13379. 555
506 CAR, this proxy may not fully capture a problem’s Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, 556
507 intrinsic complexity. Future research could inves- Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, 557
508 tigate more advanced difficulty estimators to im- Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: In- 558
509 prove the precision of cognitive regulation. centivizing reasoning capability in llms via reinforce- 559
ment learning. arXiv preprint arXiv:2501.12948. 560
510 Ethical Considerations Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul 561
Arora, Steven Basart, Eric Tang, Dawn Song, and Ja- 562
511 We follow strict ethical standards throughout this cob Steinhardt. 2021. Measuring mathematical prob- 563
512 work: (1) Regarding environmental impact, we lem solving with the math dataset. arXiv preprint 564
513 acknowledge that training LLMs involves signifi- arXiv:2103.03874. 565
514 cant energy expenditure. However, our method’s Bairu Hou, Yang Zhang, Jiabao Ji, Yujian Liu, 566
515 focus on token efficiency aims to mitigate the long- Kaizhi Qian, Jacob Andreas, and Shiyu Chang. 567
516 term inference cost. (2) To ensure transparency, all 2025. Thinkprune: Pruning long chain-of-thought 568
517 fine-tuned models are based on publicly released of llms via reinforcement learning. arXiv preprint 569
arXiv:2504.01296. 570
518 open-source architectures, avoiding the use of pro-
519 prietary or confidential models. (3) In terms of Chengyu Huang, Zhengxin Zhang, and Claire Cardie. 571
520 data usage, all benchmarks are publicly accessible, 2025. Hapo: Training language models to reason con- 572
521 ensuring that personal privacy is strictly protected. cisely via history-aware policy optimization. arXiv 573
preprint arXiv:2505.11225. 574
9
600 with 860k pairs of competition math problems and John Sweller. 1988. Cognitive load during problem 655
601 solutions. Hugging Face repository, 13(9):9. solving: Effects on learning. Cognitive science, 656
12(2):257–285. 657
602 Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua
603 Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Hugo Touvron, Louis Martin, Kevin Stone, Peter Al- 658
604 Geng, Nan Huo, et al. 2024b. Can llm already serve bert, Amjad Almahairi, Yasmine Babaei, Nikolay 659
605 as a database interface? a big bench for large-scale Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti 660
606 database grounded text-to-sqls. Advances in Neural Bhosale, et al. 2023. Llama 2: Open founda- 661
607 Information Processing Systems, 36. tion and fine-tuned chat models. arXiv preprint 662
arXiv:2307.09288. 663
608 Wei Liu, Ruochen Zhou, Yiyun Deng, Yuzhen Huang,
609 Junteng Liu, Yuntian Deng, Yizhe Zhang, and Junx- Songjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian, 664
610 ian He. 2025. Learn to reason efficiently with adap- Linjing Li, Xiangyuan Lan, and Dongbin Zhao. 2025. 665
611 tive length-based reward shaping. arXiv preprint Learning when to think: Shaping adaptive reasoning 666
612 arXiv:2505.15612. in r1-style models via multi-stage rl. arXiv preprint 667
arXiv:2505.10832. 668
613 Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan
614 Huang, Xiaodan Liang, and Song-chun Zhu. 2021. Miles Turpin, Julian Michael, Ethan Perez, and Samuel 669
615 Inter-gps: Interpretable geometry problem solving Bowman. 2024. Language models don’t always say 670
616 with formal language and symbolic reasoning. In what they think: unfaithful explanations in chain-of- 671
617 Proceedings of the 59th Annual Meeting of the Asso- thought prompting. Advances in Neural Information 672
618 ciation for Computational Linguistics and the 11th Processing Systems, 36. 673
619 International Joint Conference on Natural Language
620 Processing (Volume 1: Long Papers), pages 6774– Yikun Wang, Yibin Wang, Dianyi Wang, Zimian Peng, 674
621 6786. Qipeng Guo, Dacheng Tao, and Jiaqi Wang. 2025. 675
Geometryzero: Improving geometry solving for llm 676
622 Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, with group contrastive policy optimization. arXiv 677
623 William Y Tang, Manan Roongta, Colin Cai, Jeffrey preprint arXiv:2506.07160. 678
624 Luo, Tianjun Zhang, Li Erran Li, et al. 2025. Deep-
Violet Xiang, Chase Blagden, Rafael Rafailov, Nathan 679
625 scaler: Surpassing o1-preview with a 1.5 b model by
Lile, Sang Truong, Chelsea Finn, and Nick Haber. 680
626 scaling rl. Notion Blog.
2025. Just enough thinking: Efficient reasoning 681
627 Shangke Lyu, Linjuan Wu, Yuchen Yan, Xingyu Wu, with adaptive length penalties reinforcement learning. 682
628 Hao Li, Yongliang Shen, Peisheng Jiang, Weiming arXiv preprint arXiv:2506.05256. 683
629 Lu, Jun Xiao, and Yueting Zhuang. 2025. Hierarchi- Yuhui Xu, Hanze Dong, Lei Wang, Doyen Sahoo, Jun- 684
630 cal budget policy optimization for adaptive reasoning. nan Li, and Caiming Xiong. 2025. Scalable chain 685
631 arXiv preprint arXiv:2507.15844. of thoughts via elastic reasoning. arXiv preprint 686
arXiv:2505.05315. 687
632 Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida,
633 Carroll Wainwright, Pamela Mishkin, Chong Zhang, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, 688
634 Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, 689
635 2022. Training language models to follow instruc- Fei Huang, Haoran Wei, et al. 2024. Qwen2. 5 tech- 690
636 tions with human feedback. Advances in neural in- nical report. arXiv preprint arXiv:2412.15115. 691
637 formation processing systems, 35:27730–27744.
Yufan Ye, Ting Zhang, Wenbin Jiang, and Hua Huang. 692
638 Penghui Qi, Zichen Liu, Tianyu Pang, Chao Du, 2025. Process-supervised reinforcement learning for 693
639 Wee Sun Lee, and Min Lin. 2025. Optimizing any- code generation. arXiv preprint arXiv:2502.01715. 694
640 time reasoning via budget relative policy optimiza-
641 tion. arXiv preprint arXiv:2505.13438. Jingyang Yi, Jiazheng Wang, and Sida Li. 2025. Short- 695
erbetter: Guiding reasoning models to find optimal in- 696
642 Claude E Shannon. 1948. A mathematical theory of ference length for efficient reasoning. arXiv preprint 697
643 communication. The Bell system technical journal, arXiv:2504.21370. 698
644 27(3):379–423.
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, 699
645 Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingn- 700
646 Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan ing Yao, Shanelle Roman, et al. 2018. Spider: A 701
647 Zhang, YK Li, Y Wu, et al. 2024. Deepseekmath: large-scale human-labeled dataset for complex and 702
648 Pushing the limits of mathematical reasoning in open cross-domain semantic parsing and text-to-sql task. 703
649 language models. arXiv preprint arXiv:2402.03300. In Proceedings of the 2018 Conference on Empiri- 704
cal Methods in Natural Language Processing, pages 705
650 Yi Shen, Jian Zhang, Jieyun Huang, Shuming Shi, Wen- 3911–3921. 706
651 jing Zhang, Jiangze Yan, Ning Wang, Kai Wang, and
652 Shiguo Lian. 2025. Dast: Difficulty-adaptive slow- Zishun Yu, Tengyu Xu, Di Jin, Karthik Abinav 707
653 thinking for large reasoning models. arXiv preprint Sankararaman, Yun He, Wenxuan Zhou, Zhouhao 708
654 arXiv:2503.04472. Zeng, Eryk Helenowski, Chen Zhu, Sinong Wang, 709
10
710 et al. 2025. Think smarter not harder: Adaptive Appendices Content 721
711 reasoning with inference aware optimization. arXiv
712 preprint arXiv:2501.17974. A Appendix A: References to Models and 722
713 Jiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng, and Datasets 12 723
714 Juanzi Li. 2025. Adaptthink: Reasoning mod-
715 els can learn when to think. arXiv preprint B Appendix B: Baseline Implementation 724
716 arXiv:2505.13417. Details 12 725
11
755 A Appendix A: References to Models and • Spider: (1034 testing data) 799
756 Datasets [Link] 800
757 This section provides the links regarding the mod- • BIRD: (9428 and 1534 data for training and 801
758 els and datasets used in this work. These datasets testing 802
759 are publicly available under the CC BY-SA 4.0 [Link] 803
760 license.
761 (1) Models: • Geometry3K: (2100 and 601 data for training 804
771 • Qwen2.5-Coder-3B-Instruct: fair comparison. This includes the AdamW opti- 814
772 [Link] mizer with a learning rate of 1 × 10−6 , a batch size 815
790 • AMC23: (40 testing data) • HAPO (Huang et al., 2025): HAPO intro- 834
791 [Link] duces a cosine-based penalty that considers 835
792 math-ai/amc23 historical response lengths. We adopt the key 836
793 • Minerva: (272 testing data) hyperparameters from the original paper, set- 837
795 svc-huggingface/minerva-math
• LASER (Liu et al., 2025): This method uses a 839
796 • OlympiadBench: (674 testing data) simple yet effective length-based reward shap- 840
797 [Link] ing. Following its official implementation, we 841
798 knoveleng/OlympiadBench set α = 0.5 and LT = 800. 842
12
843 • ALP (Xiang et al., 2025): ALP proposes an D Appendix D: Further Analysis 877
844 additive penalty that adapts to the batch-level
D.1 Comparison between CUR and 878
845 correctness rate. We set β = 1 × 10−7 follow-
Suboptimal Reward Shaping 879
846 ing the original paper.
The primary advantage of CUR lies in its multi- 880
847 By adhering closely to the original configura- plicative structure, which fundamentally reshapes 881
848 tions, we ensure that our comparisons are fair and how correctness and efficiency interact. Although 882
849 accurately reflect the relative performance of each group normalization may still yield negative advan- 883
850 method against CGPO. tages for certain correct responses, the functional 884
form of CUR substantially reduces this issue, as 885
851 C Appendix C: Detailed Proof in CUR
evidenced in Table 4. 886
852 Theorem 1 (Closed-Form Solution for the Effi-
Deficiency 1: Decoupling Error in Additive For- 887
853 ciency Factor). Given the marginal cost function
mulations. Additive reward formulations (i.e., 888
854 in Eq. (1) and the decay principle in Eq. (2), the
Ui = Ri − penalty) are structurally flawed as they 889
855 unique closed-form solution for the efficiency fac-
decouple correctness rewards from length penalties. 890
856 tor E(Li ) is:
This separation enables length penalties to dispro- 891
portionately erode the utility of correct answers.
1 892
857 E(Li ) = exp − γ0 Li + kL2i (14) While the base reward denotes correctness, exces- 893
2
sive penalties often cause long but valid responses 894
858 Proof. We solve the differential equation from Eq. to fall well below the group mean. Consequently, 895
859 (2) using separation of variables. these valid solutions are assigned negative advan- 896
860 1. Separate Variables: We rearrange Eq. (2) to tages. This represents a critical alignment failure. 897
867 where C is the constant of integration. tion. In a group-normalization setting, the inflated 911
rewards of short answers raise the group average, 912
868 4. Solve for E(Li ): relegating longer correct responses to below-mean 913
1 2
status. As a result, valid but lengthy solutions re- 914
C
|E(Li )| = e · exp − γ0 Li + kLi ceive suppressive negative advantages. 915
2
869 (18) CUR’s Principled Multiplicative Solution. In 916
870 5. Apply Boundary Condition: Using E(0) = contrast, CUR’s multiplicative structure (i.e., Ui = 917
13
Method Hyperparameter Length Correctness Reward Advantage
−4
β = 4 × 10 [400, 600, 1500, 900, [1, 1, 1, 0, [0.88, 0.82, 0.55, -0.27, [1.10, 0.99, 0.48, -1.06,
ALP
K=8 2800, 250, 1200, 4500] 1, 1, 0, 1] 0.16, 0.925, -0.36, -0.35] -0.25, 1.19, -1.23, -1.21]
w=1
[400, 600, 1500, 900 [1, 1, 1, 0, [1.81, 1.59, 0.29, 0.00, [1.53, 1.22, -0.58, -0.99,
HAPO c = −0.8
2800, 250, 1200, 4500] 1, 1, 0, 1] 0.20, 1.92, -0.31, 0.20] -0.71, 1.68, -1.42, -0.71]
h(q) = 1000
α = 1.5 [400, 600, 1500, 900, [1, 1, 1, 0, [2.50, 2.50, 1.00, 0.00, [1.13, 1.13, -0.29, -1.25,
LASER
LT = 800 2800, 250, 1200, 4500] 1, 1, 0, 1] 1.00, 2.50, 0.00, 1.00] -0.29, 1.13, -1.25, -0.30]
[400, 600, 1500, 900, [1, 1, 1, 0, [0.992, 0.982, 0.895, 0, [0.97, 0.94, 0.72, -1.57,
CGPO η=1
2800, 250, 1200, 4500] 1, 1, 0, 1] 0.679, 0.997, 0, 0.368] 0.17, 0.98, -1.57, -0.63]
Table 4: An analysis of learning signals from different reward designs. Advantages are computed by normalizing
rewards within a group of 8 samples (see Eq. (6)), and negative values for correct responses are highlighted in blue.
926 answers cannot be suppressed merely due to 2. Monotonically Decreasing: For any length 961
927 length, CUR provides consistent optimization L > 0, U (L) is a strictly decreasing function 962
928 signals that reinforce correctness. of L. 963
929 • Mitigating the Polarization Error: Unlike Proof. The above two properties can be proven as 964
930 LASER’s discrete bonus, CUR employs a follows: 965
931 smooth and continuous decay as its efficiency
932 factor. This produces a fine-grained, mono- 1. Boundedness: Since L2 /L2max ≥ 0 and 966
2
933 tonic ranking of correct responses by length, η > 0, −η LL2 ≤ 0 and U (L) ≤ e0 = 1. 967
max
934 substantially reducing unjustified penalties As the exponential function is always positive, 968
935 for moderately long but valid answers under we have U (L) ∈ (0, 1]. This guarantees that 969
936 group normalization. all correct solutions receive a positive utility, 970
fundamentally distinguishing CUR from addi- 971
937 Summary. Adopting a multiplicative formula-
tive penalties. 972
938 tion is a prerequisite for maintaining the proper
939 hierarchical relationship between correctness and
2. Monotonicity: The first derivative of U (L) 973
940 efficiency. However, this measure alone is insuf-
with respect to L is: 974
941 ficient. The strength of CUR stems equally from
942 the functional design of its efficiency component. dU (L)
L2
2ηL
943 By introducing a smooth decay rather than a coarse = exp −η 2 · − 2
dL Lmax Lmax
944 threshold, CUR establishes a more stable and con- (20) 975
945 sistent foundation for policy optimization than both For L > 0, the exponential term is strictly 976
946 additive penalties and rigid multiplicative schemes. positive, while the term in the parentheses 977
947 D.2 Analysis of the Functional Form of CUR is strictly negative. Therefore, dUdL
(L)
< 0, 978
indicating that the function is monotonically 979
948 The design of CUR is based on the assumption that
decreasing. This ensures that shorter correct 980
949 the marginal cost of reasoning increases linearly,
solutions are always preferred. 981
950 resulting in a quadratic exponential decay. This
951 first-order approximation is principled and yields 982
952 distinct advantages over alternative cost models.
953 To formalize these properties, we present the These proven properties directly support our de- 983
954 following lemma regarding the final CUR utility sign choices. The justification for these choices, 984
955 function for correct responses (Ri = 1): from the underlying cost model to the final imple- 985
956 Lemma 1 (Properties of the CUR Function). For mentation, is detailed below. 986
957 any given batch with a fixed Lmax , the CUR utility
2
958 function U (L) = exp(−η LL2 ) is: Justification for the Linear Marginal Cost For- 987
max
mulation. The primary justification for a linear 988
959 1. Strictly Bounded: The utility U (L) is strictly marginal cost (i.e., γ(L) = γ0 + kL) lies in its 989
960 bounded within the interval (0, 1]. ability to represent escalating cognitive load in 990
14
991 an analytically tractable manner. As a reason- chain remaining valid as its length increases, par- 1040
992 ing process extends, the cognitive burden of man- ticularly under the assumption of cumulative error 1041
993 aging context and avoiding error propagation in- accumulation. 1042
994 creases. A linear increase offers the simplest for- Assumption 1 (Context-Dependent Error Drift). 1043
995 mal representation of this “harder-as-it-goes” phe- Consider a reasoning chain of length L. Within 1044
996 nomenon. In contrast, higher-order cost functions the effective reasoning where the model maintains 1045
997 (e.g., γ(L) ∝ L2 ) introduce additional parameters coherence, the probability of making an error at 1046
998 and complexity without compelling theoretical jus- token t (i.e., ϵt ) is not constant. Due to attention dis- 1047
999 tification. Therefore, inspired by cognitive science persion and hallucination drift, the instantaneous 1048
1000 principles (Sweller, 1988), the linear approxima- error risk increases linearly: 1049
1001 tion provides a robust and interpretable foundation.
ϵt ≈ λ · t (21) 1050
1002 Justification for γ0 = 0. In our implementation,
1003 γ0 = 0 is set to isolate the impact of the accelerat- where 0 < λ ≪ 1 is a small drift coefficient rep- 1051
1004 ing cost component (i.e., kL). This design focuses resenting the degradation of precision as context 1052
1005 the penalty on the growth of reasoning complexity grows. 1053
1006 rather than on a fixed per-token cost, which is a Remark 1 (Justification via Error Cascades). This 1054
1007 less critical factor in long-form reasoning. While a linear drift assumption is supported by the phe- 1055
1008 non-zero γ0 could be tuned to introduce a baseline nomenon of error cascades in LLMs. Specifically, 1056
1009 preference for brevity, experimental results indicate minor early hallucinations or logical deviations 1057
1010 that the accelerating quadratic term alone provides alter the context for subsequent tokens, increasing 1058
1011 sufficient regularization. Additional analysis on the the conditional probability of future errors. Conse- 1059
1012 design of CUR is shown in Appendix D.3. quently, the risk of reasoning collapse grows cumu- 1060
1013 Justification for Batch-wise Normalization by latively as the chain extends. 1061
1014 Lmax . The use of batch-wise normalization by Proposition 1 (Alignment with Linear Error Drift). 1062
1015 Lmax is a practical strategy for addressing the in- Under Assumption 1 and within the regime of small 1063
1016 trinsic non-stationarity of reasoning tasks. Abso- local error probabilities (ϵt ≪ 1), the probability 1064
1017 lute length penalties are context-sensitive, as a re- that a reasoning chain of length L remains entirely 1065
1018 sponse may be concise for one problem but ver- correct decays according to a quadratic exponen- 1066
1019 bose for another. In light of this phenomenon, tial function. Therefore, the CUR objective aligns 1067
1020 batch-wise normalization defines efficiency scores with maximizing the expected validity of the re- 1068
1021 relative to the current distribution of solution sponse. 1069
1022 lengths, introducing an adaptive baseline. This Proof. The probability that the entire chain is valid 1070
1023 approach ensures that penalties remain meaning- P (Valid|L) is the product of the success probabili- 1071
1024 ful and stable across batches with varying average ties (1 − ϵt ) at each token t: 1072
1025 response lengths. (See Appendix D.6 for further
1026 analysis). L
Y L
Y
P (Valid|L) = (1 − ϵt ) = (1 − λt) (22) 1073
1027 D.3 Theoretical Foundations of CUR t=1 t=1
1028 In this work, we introduced CUR with a quadratic To analyze the decay rate, we examine the log- 1074
2
1029 exponential decay form: rcur = R · e−ηL . While probability: 1075
1030 we provided an intuition based on cognitive load
L
1031 theory, this section derives CUR from two perspec- X
ln P (Valid|L) = ln(1 − λt) (23) 1076
1032 tives: (1) a probabilistic reliability model based on t=1
1033 cumulative error dynamics, and (2) an optimization
1034 stability analysis based on gradient sign consis- We focus on the high-fidelity regime where the 1077
15
1082 Using the summation formula for an arithmetic response yi relative to the correct response yc . The 1125
L(L+1) 2
series L ≈ L2 (for sufficiently
P
1083 t=1 t = 2
model actively avoids correct long-form reasoning 1126
1084 large L): to prevent length penalty. 1127
2
Case 2: CUR Stability. Let rcur = 1 · e−ηL . 1128
λ
1085 ln P (Valid|L) ≈ − L2 (25) For any finite L, rcur (yc ) > 0. Therefore, 1129
2
rcur (yc ) > r(yi ) is globally true. 1130
1086 Exponentiating back to the probability space
2
1087 yields: e−ηL − µ 0−µ
A(yc ) = > = A(yi ) (29) 1131
σ σ
λ 2
1088 P (Valid|L) ≈ exp − L (26)
2
Result: A(yc ) > A(yi ) always holds. 1132
λ
1089 By setting the hyperparameter η = 2,
we re-
1090 cover the exact form of the CUR scaling factor. Corollary 1 (Critical Breakdown Threshold). 1133
1091 This demonstrates that the L2 term provides a Specifically, for an additive penalty coefficient β, 1134
1092 theoretically motivated approximation for the cu- there exists a critical length difference beyond 1135
1093 mulative probability of success under linear error which the reward system fails. If a correct response 1136
1094 drift. has length Lc such that: 1137
1108 A(yc ) > A(yi ) ⇐⇒ r(yc ) > r(yi ) (27) than the advantage of an incorrect response A(yi ). 1149
Therefore, the optimizer correctly prioritizes: 1150
1109 Theorem 2 (Rank Reversal in Additive Penalties).
Efficient Correct ≻ Inefficient Correct ≻ 1151
1110 Additive penalty functions cause Rank Reversal,
Incorrect 1152
1111 where the optimizer is explicitly encouraged to pre-
1112 fer incorrect responses over sufficiently long cor- In contrast, additive penalties distort this hierar- 1153
1113 rect responses. CUR is immune to this pathology. chy to: 1154
1114 Proof. Assume the standard setting where incor- Efficient Correct ≻ Incorrect ≻ 1155
1115 rect responses receive zero base reward (r(yi ) = 0) Inefficient Correct 1156
1116 and correct responses receive a positive base re-
which fundamentally misaligns the optimization 1157
1117 ward of 1 scaled by efficiency.
objective. 1158
1118 Case 1: Additive Failure. Let radd = 1−βL. If
1119 a correct chain length L > 1/β, then radd (yc ) < 0. Summary. The dual perspectives above provide 1159
1120 Consequently, radd (yc ) < r(yi ). Applying the a holistic justification for CUR. Concretely, the 1160
1121 advantage formula: probabilistic derivation ensures the reward magni- 1161
tude reflects the intrinsic reliability of long chains, 1162
radd (yc ) − µ 0−µ
1122 A(yc ) = < = A(yi ) (28) while the rank analysis guarantees optimization 1163
σ σ stability. Collectively, they establish CUR as a the- 1164
1123 Result: A(yc ) < A(yi ). The policy gradient oretically sound objective for efficient reasoning. 1165
1124 update will increase the probability of the incorrect More analysis on CUR is shown in Appendix D.4. 1166
16
1167 D.4 Further Analysis of Optimization Remark 3. While A(ylong ) may become negative 1213
1168 Dynamics if R(ylong ) falls below the group mean µ (i.e., it is 1214
1169 In this section, we provide further analysis to justify less efficient than peers), it is structurally guaran- 1215
1170 the superiority of CUR from the perspectives of teed to remain ranked higher than any incorrect 1216
1171 Correctness Consistency and Pareto Efficiency. response. This ensures the negative signal targets 1217
inefficiency, whereas additive methods risk target- 1218
1172 D.4.1 Correctness Consistency ing correctness itself due to the rank flip. 1219
1173 A fundamental risk in efficiency-oriented RL is pol-
D.4.2 Pareto Optimality of the CUR Objective 1220
1174 icy degeneration, where the model unlearns valid
1175 reasoning paths since the length penalty outweighs We further clarify that the design of CUR is a solu- 1221
1176 the correctness reward. We define Correctness Con- tion to a constrained optimization problem on the 1222
1178 Definition 2 (Correctness Consistency). Let ypos Proposition 2 (Equivalence to Constrained Opti- 1224
1179 be any correct response and yneg be any incorrect mization). Maximizing the expected CUR objective 1225
1180 response. A reward shaping mechanism satisfies JCU R (π) is equivalent to solving the dual problem 1226
1181 Correctness Consistency if the reward for correct- of maximizing accuracy subject to a soft constraint 1227
1182 ness always strictly exceeds the reward for incor- on the reasoning length variance. 1228
1183 rectness, regardless of length:
Proof. Consider the primary objective of maximiz- 1229
R(ypos , Lpos ) > R(yneg , Lneg ), ∀Lpos , Lneg ing the expected correctness R and a constraint to 1230
1184 (31) keep the reasoning steps within a computational 1231
1185 Theorem 3 (Violation of Consistency in Addi- budget. We formulate this as: 1232
1190 response receives a lower reward than an invalid (as derived in Appendix C). The Lagrangian of this 1235
1192 Proof. Let Radd (y) = I(y) − βL(y). Consider L(π, λ) = Ey∼π [R(y)] − λ(Ey∼π [L(y)2 ] − C)
1193 ylong (correct) and yshort (incorrect). The condi- (34) 1237
1194 tion for rank reversal is R(ylong ) < R(yshort ): Rearranging the terms inside the expectation: 1238
1
1−βLlong < 0−βLshort =⇒ Llong −Lshort > L(π, λ) = Ey∼π [R(y) − λL(y)2 ] + λC (35) 1239
β
1195 (32)
While the standard Lagrange form is additive 1240
1196 When this occurs, the optimizer assigns a higher
(R − λL2 ), consider the optimization in the utility 1241
1197 value to the incorrect response. In the context of
space. For correct responses (R = 1), the CUR 1242
1198 advantage estimation, this can lead to A(yshort ) >
objective implies: 1243
1199 A(ylong ), causing the policy to reinforce the incor-
1200 rect behavior over the correct but verbose one. L2
ln(UCU R ) = ln(1) − η = −λL2 (36) 1244
1201 Theorem 4 (Consistency Guarantee of CUR). L2max
1202 CUR satisfies Correctness Consistency, ensuring
This mirrors the cost penalty term in the La- 1245
1203 that gradient signals always prioritize correctness
grangian. Therefore, maximizing the expected 1246
1204 over length.
CUR utility effectively imposes a soft constraint 1247
1205 Proof. The CUR formulation is Rcur (y) = I(y) · on the squared length variance to push the policy 1248
1206 ϕ(L). For any ylong (correct) and yshort (incor- towards the Pareto frontier, where correctness is 1249
1207 rect): maximized for a given computational budget de- 1250
1208 Rcur (ylong ) = ϕ(Llong ) > 0 fined by η. 1251
1209
1210 Rcur (yshort ) = 0 Remark 4 (Physical Meaning of η). From the per- 1252
1211 Since ϕ(L) is strictly positive, R(ylong ) > spective of Lagrangian multipliers, the hyperpa- 1253
1212 R(yshort ) holds universally. rameter η in CUR represents the exchange rate 1254
17
2
1255 between accuracy and computation on the Pareto As L → ∞, the term ecL grows infinitely large, 1298
1256 frontier. and the limit is 0. This proves that the exponential 1299
decay fcur (L) diminishes significantly faster than 1300
1257 L(π, η) = E[R] − η · E[L2 ] (37) the rational decay frq (L) for large L. 1301
1258 Tuning η allows us to traverse the Pareto fron- Although training operates within a bounded 1302
1259 tier, where a larger η prioritizes efficiency and a length L ∈ [0, Lmax ], this asymptotic dominance 1303
1260 smaller η favors accuracy. This insight offers a governs the curvature of the penalty function within 1304
1261 practical guideline for hyperparameter selection. the effective range. As depicted in Figure 7, the ex- 1305
1262 In particular, η should be aligned with the latency ponential function imposes a substantially stronger 1306
1263 constraints of the application. For instance, real- penalty as L → Lmax . This renders CUR partic- 1307
1264 time systems requiring rapid responses warrant a ularly effective at curbing excessive overthinking. 1308
1265 larger η, whereas correctness-oriented reasoning While its penalty is consistently stricter, the initial 1309
1266 tasks benefit from a smaller η. region remains sufficiently flat to safeguard essen- 1310
tial reasoning steps. However, its subsequent steep 1311
1267 D.5 Further Analysis on the Reward decline delivers a decisive signal against emerging 1312
1268 Formulation in CUR verbosity. 1313
1269 The superiority of CUR’s design over additive and Essentially, the exponential decay provides a 1314
1270 rational quadratic formulations derives from the closer approximation to a soft-threshold mecha- 1315
1271 distinct curvature profiles of their decay functions. nism. It permits complexity up to a certain point 1316
1272 The additive penalty (i.e., Ri − η(L/Lmax )2 ) is before becoming rapidly suppressive. In contrast, 1317
1273 fundamentally flawed since it decouples correct- the rational quadratic decay is overly gradual, mak- 1318
1274 ness from efficiency. Consequently, a correct but ing it less adept at discriminating between neces- 1319
1275 verbose solution may yield a lower raw reward than sary reasoning and unproductive verbosity. This 1320
1276 a short incorrect one, resulting in rank reversal and structural advantage enables CUR to strike a supe- 1321
1277 misleading advantage signals. rior balance, maintaining accuracy while achieving 1322
1278 The comparison with the rational quadratic for- greater token reduction. 1323
18
1330 non-normalized, absolute length penalty. Therefore, CAR analytically acts as an SNR- 1376
1331 Without normalization, CUR reduces to adaptive trust region scheduler, thereby optimizing 1377
the exploration-exploitation trade-off based on the 1378
1332 Ui = Ri · exp(−η · L2i ) (40) local difficulty estimate zj . 1379
1342 In contrast, batch-wise normalization by Lmax Lemma 3 (Properties of the CAR Mapping Func- 1389
1343 reframes penalties in terms of relative length rather tion). The CAR mapping function βada (z) = βbase · 1390
1344 than absolute length. This serves as an adaptive (1 − tanh(z)) exhibits the following properties for 1391
1345 and context-aware scaling mechanism. In batches any z ∈ R: 1392
1346 containing complex problems that elicit long re-
1347 sponses, a large Lmax naturally softens the penalty 1. Strict Boundedness: The output βada (z) is 1393
1348 for all solutions, accommodating necessary reason- strictly bounded within (0, 2βbase ). 1394
1357 D.7 Theoretical Justification for CAR 1. Boundedness: Since tanh(z) ∈ (−1, 1), it 1401
follows that (1 − tanh(z)) ∈ (0, 2). Multi- 1402
1358 We analyze the superiority of CAR via the Signal- plying by the positive constant βbase yields 1403
1359 to-Noise Ratio (SNR) of the policy gradient. βada (z) ∈ (0, 2βbase ), ensuring the KL coeffi- 1404
1360 Lemma 2 (SNR-Adaptive Regularization). Let the cient is strictly positive and bounded. 1405
1361 gradient SNR be defined as the ratio of the expected
2. Monotonicity: The first derivative of βada (z) 1406
1362 reward signal to the variance of the returns. For
with respect to z is: 1407
1363 difficult tasks, the success rate is low, implying high
1364 variance and low SNR. dβada (z)
= −βbase · sech2 (z) (41) 1408
1365 In Trust Region-based methods (e.g., GRPO), dz
1366 the KL penalty βDKL restricts the step size.
Since βbase > 0 and sech2 (z) > 0 for all 1409
z ∈ R, the derivative is always negative. This 1410
1367 • Easy Tasks (High SNR): The gradient points
proves that a higher difficulty score z strictly 1411
1368 reliably towards the optimum. A small trust
leads to a lower KL coefficient. 1412
1369 region (high β) is desirable to mitigate catas-
1370 trophic forgetting of the optimal solution. 3. Smoothness: The tanh(z) function is in- 1413
finitely differentiable, and since βada (z) is a 1414
1371 • Hard Tasks (Low SNR): A restrictive trust linear transformation of tanh(z), it inherits 1415
1372 region (high β) prevents the policy from ex- this property, guaranteeing a smooth response 1416
1373 ploring far enough to find the correct solution. to changes in difficulty. 1417
1374 To escape a local optimum, the policy requires
1375 a larger search radius (lower β). 1418
19
1419 Based on these proven properties, the design The results in Figure 6(a) highlight their com- 1466
1420 yields three practical advantages: plementary roles. Concretely, CUR provides a 1467
direct and fundamental incentive for conciseness. 1468
1421 1. Boundedness: As proven in Lemma 3, the By embedding a length-based penalty into the re- 1469
1422 output range is stable and predictable. This ward function, CUR imposes universal pressure 1470
1423 prevents the KL term from collapsing to zero to generate concise correct solutions regardless of 1471
1424 or exploding, mitigating risks of policy col- task difficulty. Consequently, the model is consis- 1472
1425 lapse or stagnation. tently rewarded for identifying efficient reasoning 1473
paths. This global pressure is the primary driver 1474
1426 2. Smoothness and Non-linearity: The smooth-
behind CUR’s substantial reduction in response 1475
1427 ness prevents abrupt shifts in the learning
length. In contrast, the efficiency gain observed in 1476
1428 objective. More importantly, the non-linear
CAR emerges as a byproduct of its adaptive regu- 1477
1429 shape naturally modulates sensitivity:
larization. Length reduction is achieved primarily 1478
1430 • For zj near 0 (ambiguous difficulty), by assigning higher β values to easier problems. 1479
1431 the gradient of tanh function is steep, This constraint discourages verbosity in tasks that 1480
1432 making βada highly responsive to small have already been mastered. While CAR permits 1481
1433 changes in difficulty. extended reasoning for difficult problems, the sup- 1482
1434 • For large |zj | values (very easy or very pression of verbosity in frequent simple tasks re- 1483
1435 hard problems), the function saturates, re- sults in a net reduction in average sequence length. 1484
1436 ducing the sensitivity of βada to extreme In conclusion, while both components enhance 1485
1437 scores, thereby enhancing robustness. efficiency, CUR’s effect is more pronounced due to 1486
its direct and global mechanism, whereas CAR’s 1487
1438 3. Inverse Relationship: The monotonic de- impact is indirect and conditional upon enhanced 1488
1439 crease proven in Lemma 3 ensures that higher learning stability. The complete CGPO framework 1489
1440 difficulty reduces regularization and encour- validates the efficacy of this synergy, achieving an 1490
1441 ages exploration, while lower difficulty in- optimal accuracy–efficiency trade-off. 1491
1442 creases regularization and reinforces stability.
D.10 Further Analysis on the Impact of η and 1492
1443 Accordingly, the tanh function is a natural fit βbase 1493
1444 for standardized scores due to its bounded, smooth,
The sensitivity analyses in Figures 6(b) and (c) 1494
1445 and non-linear mapping.
elucidate the distinct roles of η and βbase in mod- 1495
1446 D.9 Synergistic Effect of CUR and CAR ulating the accuracy-efficiency trade-off and the 1496
exploration-stability balance, respectively. 1497
1447 The superior performance of CGPO stems from the
1448 synergy between CUR and CAR, which operate at The Role of η in CUR The parameter η directly 1498
1449 distinct yet interrelated levels of the learning pro- governs the steepness of the efficiency penalty, 1499
1450 cess. Specifically, CUR shapes the global reward thereby regulating the equilibrium between cor- 1500
1451 landscape across a batch, while CAR modulates rectness and conciseness. 1501
1452 the sample-level exploration policy.
1453 Specifically, CUR defines a global optimization • A low η produces a gentle penalty slope, ren- 1502
1454 landscape, evaluating solutions under a unified cog- dering CUR more tolerant of verbosity. This 1503
1455 nitive utility standard to exert baseline pressure configuration enables the model to explore 1504
1456 toward conciseness. However, relying solely on extended reasoning paths that may enhance 1505
1457 static regularization proves suboptimal within this accuracy at the cost of reduced efficiency. 1506
1458 landscape. To address this, CAR acts as an adaptive
1459 moderator. For simpler tasks, it assigns a higher • A high η establishes a stringent penalty 1507
1460 β to tighten the KL constraint, promoting stability regime, rigorously penalizing any deviation 1508
1461 by precluding deviation from established efficient from brevity. This drives the model toward 1509
1462 paths. Conversely, for complex queries, it assigns the shortest possible correct solutions, maxi- 1510
1463 a lower β to relax the constraint, granting the pol- mizing efficiency but risking under-reasoning 1511
1464 icy greater latitude to explore intricate reasoning on complex tasks, which may ultimately im- 1512
1465 trajectories toward correctness. pair peak accuracy. 1513
20
1514 Experimental results indicate that η = 1 is an op- (mod 18). This allows for a direct and principled 1560
1515 timal balance between these competing pressures. simplification of the problem, leading to the cor- 1561
rect intermediate congruence (i.e., 11213141 ≡ 5 1562
1516 The Role of βbase in CAR The parameter βbase (mod 18)) and the correct final answer (i.e., n = 1563
1517 sets the anchor point for adaptive KL regulariza- 13). This case highlights CGPO’s ability to dis- 1564
1518 tion, thereby determining the exploration capacity: cover and reinforce reasoning strategies that are 1565
1519 • A low βbase anchors the adaptive βada at a correct and concise. The response is less than half 1566
1520 lower bound, promoting aggressive explo- the length of the GRPO response. This behavior is 1567
1521 ration across all difficulty levels. While con- fostered by the synergy of CUR and CAR, which 1568
1522 ducive to discovering novel solutions, this rewards efficient correct solutions and provides the 1569
1523 regime risks instability and policy degradation appropriate regularization to find such robust paths. 1570
21
Geometry3K Spider-Dev BIRD-Dev SQL Avg.
Methods
Pass@1 Length Pass@1 Length Pass@1 Length Pass@1 Length
Base 25.7 552 70.2 378 43.6 536 56.9 457
GRPO 35.6 418 73.8 283 55.7 482 64.8 383
AdaptThink 34.4 384 72.4 254 53.3 427 62.9 341
HAPO 33.9 386 72.8 297 54.7 446 63.8 372
Laser 35.6 351 73.5 242 55.4 458 64.5 350
ALP 34.6 338 73.2 228 54.2 431 63.7 330
CGPO 36.1 322 74.3 246 56.0 441 65.2 344
Table 5: Experimental results on multi-modal reasoning and SQL generation benchmarks. For each task, we report
Pass@1 score and average response length. The best results in each category are highlighted in bold.
Table 6: Ablation study of using different components of CGPO across various mathematical reasoning benchmarks
in terms of Pass@1 score and reasoning length. The best results are highlighted in bold.
Table 7: Analysis of the base KL coefficient (βbase ) in our CAR module. All variants use the full CGPO framework,
but with CAR anchored around a different βbase . The best results for each model are highlighted in bold.
22
Math500 AIME24 AIME25 AMC23 Minerva Olympiad Avg.
η
Pass@1 Length Pass@1 Length Pass@1 Length Pass@1 Length Pass@1 Length Pass@1 Length Pass@1 Length
DeepSeek-R1-Distill-Qwen-1.5B
0.5 82.8 2254 30.0 6706 30.0 6711 72.5 4102 31.3 2854 48.1 4676 49.1 4551
1 82.6 2180 30.0 6662 30.0 6826 72.5 4049 31.6 2835 48.7 4646 49.2 4533
1.5 82.2 2057 26.7 6511 30.0 6723 72.5 4009 31.0 2814 47.9 4655 48.4 4462
DeepScaleR-Preview-1.5B
0.5 85.6 1661 26.7 5417 30.0 5653 70.0 2801 33.1 2648 52.8 3686 49.7 3644
1 85.6 1632 26.7 5398 30.0 5624 70.0 2811 33.5 2674 53.1 3660 49.8 3633
1.5 85.4 1653 26.7 5348 26.7 5602 70.0 2647 32.7 2503 52.8 3631 49.1 3561
Table 8: Hyperparameter analysis of the efficiency coefficient η in our CUR formulation. The best results for each
model are highlighted in bold.
Table 9: Ablation study on different reward formulations. Our CUR (using quadratic exponential decay) is compared
against additive penalty and rational quadratic decay. The best results for each model are highlighted in bold.
Table 10: Ablation study on the normalization technique in CUR. We compare our default CUR against a variant
without normalization. The best results for each model are highlighted in bold.
23
Case study: GRPO vs. CGPO
Question:
What integer n satisfies n ≤ 0 < 18 and n ≡ −11213141 (mod 18)?
11 21 31 41
11 − 21 + 31 − 41 = −10 + 31 − 41 = 21 − 41 = −20
−20 + 18 = −2
−2 + 18 = 16
−16 + 18 = 2
Thus,
−11213141 ≡ 2 (mod 18)
24
The integer n that satisfies 0 ≤ n < 18 and n ≡ −11213141 (mod 18) is 2.
13
25