Option
Option
*Correspondence:
mczhao1986@[Link]; fqli@163.
Abstract
com The deployment of 5G networks has incorporated advanced multiple access technolo-
1
The 20th Research Institute gies like sparse code multiple access (SCMA) to address growing demands for high-
of CETC, Xi’an, Shaan xi, China speed connectivity and massive device access. As a Non-Orthogonal Multiple Access
2
China Mobile Communications
Group Shaanxi Co. Ltd, Xi’an, technique, SCMA enables multiple users to share identical time-frequency resources
Shaan xi, China through sparse codebook-based multiplexing. Nevertheless, achieving efficient
scheduling in SCMA networks remains challenging due to the inherent complexi-
ties in dynamic resource allocation. This paper proposed two artificial intelligence-
based approaches for resource scheduling in 5G SCMA networks: a multi-agent
deep reinforcement learning (MARL)-based approach and a large language model
(LLM)-empowered methodology. We systematically investigate these AI techniques
to develop adaptive resource scheduling policies capable of responding to diverse net-
work conditions. Simulation results validate that the proposed MARL-based and LLM-
based schedulers not only effectively learn optimal scheduling strategies but also out-
perform conventional algorithms, particularly in terms of system throughput and user
fairness metrics.
Keywords: Scheduling and resource allocation, Sparse code multiple access, Multi-
agent deep reinforcement learning, Large language model, Artificial intelligence
1 Introduction
The fifth generation (5G) of cellular networks is designed to support a wide range of
applications, including enhanced mobile broadband (eMBB), ultra-reliable low-latency
communications (URLLC), and massive machine-type communications (mMTC).
SCMA [1] is a promising multiple access technique for 5G networks, enabling high spec-
tral efficiency and massive connectivity by enabling multiple users to share the same
time-frequency resources through sparse codebooks. However, despite its advantages,
SCMA presents significant challenges in resource scheduling and interference manage-
ment due to its non-orthogonal nature.
Traditional scheduling algorithms, such as those based on orthogonal frequency-
division multiple access (OFDMA), are not well-suited for SCMA due to its unique
characteristics. The network operates in a time-slotted manner, with each time slot
corresponding to a scheduling decision. The objective of the SCMA scheduler [2] is to
© The Author(s) 2026. Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0
International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long
as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you
modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of
it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise
in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted
by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy
of this licence, visit [Link]
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 2 of 26
allocate codebooks and channel resources to user equipment (UEs) in a manner that
maximizes network performance metrics, including throughput and user fairness.
In wireless networks, the wireless channels are in a time-varying state, and tradi-
tional static algorithmic mechanisms are inadequate for addressing the complexities
of such environments. Consequently, researchers [3, 4] use deep learning methods to
tackle the challenges associated with wireless scheduling algorithms in cellular net-
works. Deep reinforcement learning (DRL) is proven to be a powerful tool for solv-
ing complex resource allocation optimization problems in dynamic environments.
For the complex multi-resource allocation problem, the authors [5] proposed uses
multi-agent deep reinforcement learning (MARL) method in D2D networks, where
each D2D pair acts as an independent agent making local decisions on subchannel
and power allocation.
Furthermore, with the rapid development of LLMs such as ChatGPT [6] and Deep-
Seek [7], researchers have begun exploring their innovative applications in wireless
communication resource allocation. By leveraging their powerful capabilities in natu-
ral language processing, pattern recognition, and decision optimization, LLMs offer a
new research paradigm for traditional resource allocation algorithms [8].
In this paper, we explore the adoption of multi-agent deep learning and LLM meth-
ods to address the multiple resource allocation challenges associated with wireless
scheduling algorithms in SCMA cellular networks. We propose a MARL-based and
LLM-empowered scheduling framework for 5G SCMA networks, aiming to optimize
key performance metrics such as throughput and user fairness. Our main contribu-
tions are highlighted as follows.
Some preliminary results of our proposed algorithms are reported in [9]. This paper
provides more comprehensive MARL algorithm, LLM prompts design and experimen-
tal results. Firstly, a more detailed description of MARL algorithm is provided, includ-
ing the L-W greedy algorithm. Secondly, we provide a more comprehensive description
of LLM-based scheduling algorithm, along with the design and a comparative analysis
of two different prompts. Thirdly, We conducted a more comprehensive experimental
validation, which included add the reference comparative algorithm MASR_SCMA, per-
forming a comparative analysis of the scheduling effectiveness of different LLMs with
different prompts, and analyzing system capacity, among other aspects.
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 3 of 26
The rest of this paper is organized as follows. Section 2 introduces the related work.
Section 3 introduces the proposed framework, the proposed MARL method and
LLM-empowered method. Simulation setup and results are presented in Sects. 4 and
5. We also discuss the challenges and open questions in Sect. 6. Followed by conclud-
ing remarks in Sect. 7.
2 Literature review
In resource allocation for wireless communication systems, parameters such as the
channel and transmit power are adjusted to optimize network performance, includ-
ing maximizing system throughput and fairness while satisfying various constraints.
Finding the optimal strategy is typically formulated as an optimization problem,
which can be addressed using conventional optimization-based approaches, DRL-
based approaches and LLM-based approaches as follows (Fig. 1):
Fig. 1 The comparison between Optimization-based, DRL-based and LLM-based methods for wireless
resource allocation problem
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 4 of 26
and power resource in order to maximize SCMA system throughput. The impact of the
factor graph matrix on the average sum rate was studied in [13] by assuming the known
link distances between users and the base station (BS), and a low-complexity iterative
algorithm is proposed to design the optimal graph matrix which maximizes the average
sum rate of the SCMA systems for the general parameters. The authors [14] studied to
maximize the NOMA system’s sum rate and energy efficiency (EE) using the proposed
iterative water-filling solution. An iterative algorithm in [15] was utilized in the resource
assignment problem to maximize energy efficiency.
present a consolidated review of the state of the art in network slicing resource manage-
ment modules and network slicing-enabled key industrial use cases in 5G core network
using AI-based methods, such as RL method. The author [23] proposes a DRL-based
rate adaptation algorithm for adaptive 360-degree video streaming, which maximizes
the quality of experience for viewers by adjusting the transmitted video quality to the
time-varying network conditions.
3.2 Problem formulation
We consider a centrally controlled cellular downlink communication system, which is
used in 5G cellular Network. A set of users m={1,…,M} requests services and the BS trans-
mits messages to these users via both the backbone network and the wireless cellular net-
work. The system operates over a total available bandwidth B, which is equally divided into
N={1,…,N} resource block groups, each with 12 consecutive subcarriers. All the chan-
nels are assumed to exhibit block fading characteristics. For each user m ∈ M and each
subchannel n ∈ N , the transmit power and the channel gain for user k in subchannel n are
denoted as pm,n and hm,n, respectively, and σ 2 represents the ambient noise variance. The
communication system is assumed to employ an M-QAM modulation scheme.
The SCMA resource allocation problem can be mathematically modeled as:
M
Problem 1 : maximize Rm + γ F , (1)
m=1
where Rm indicates the average rate of user m, F represents the fairness (indicated
by Jain’s fairness, which is popular in wireless scheduling index as [3, 4]) index. The
throughput in the reward function is the received instantaneous throughput, hence the
UE with the best channel condition tends to be chosen by considering throughput only.
The fairness component in the reward focuses on equity in average throughput among
different users. Consequently, resources may be allocated to users whose channel condi-
tions are not optimal, in order to ensure fairness, γ indicates the weight parameter which
controls the trade-off between rate and fairness. Jain’s fairness F can be calculated as:
M 2
[ m=1 Rm ]
F= M 2 (3)
M m=1 [Rm ]
Considering that γ is a number with a value range of [0, 1], we divided the rate Rm by 1
million when calculating resource allocation efficiency to bring its value range close to 1,
thereby effectively balancing the rate value with fairness and obtaining a more reason-
able reward value.
According to the resource allocation principle [32] of SCMA, a subcarrier can be allo-
cated to a maximum of dJ users, and a user can simultaneously receive data from a maxi-
mum of dV subcarriers. The total rate of user m across all subcarriers can be expressed as:
N
Rm = µm,n rm,n
n=1
N
(4)
B βpm,n |hm,n |2
= µm,n log2 1 +
N σ 2 + Im,n
n=1
where µm,n = 1 indicates that RB n is allocated to user m and µm,n = 0 indicates that
RB n is not allocated to user m. rm,n represents the transmission rate of user m on RB n.
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 9 of 26
Assuming the base station employs M-QAM modulation, Im,n represents the interfer-
ence received by user m on subcarrier RB n from other users. User m’s feasible rate on
subchannel n can be expressed as:
B βpm,n |hm,n |2
rm,n = log2 1 + . (6)
N σ 2 + Im,n
which is user m receives from other users Sn |(|hi,n |2 < |hm,n |2 ) on resource block n
[32]. Since RB n can be utilized by a subset of users, the signal of any user i will cause
the interference to other users in RB n. To demodulate the target signal, usually SCMA
receivers utilize the successive interference cancellation (SIC) decoding [34]. It firstly
decodes the signal of the user with better channel conditions, subtracts it, and then
decodes its own signal. Therefore, when user m decodes its own signal, the interference
it experiences comes from all users with poorer channel conditions than itself. In the
base station, the data transmitted to each user are mapped to the corresponding sparse
code of the codebook. The data from all m SCMA layers are multiplexed onto N shared
resource blocks. The signal on the resource block can be expressed as:
M
yn = hm,n xm,n + wn (8)
m=1
where hm,n represents the channel coefficient of user m on resource block n, xm,n denotes
the codeword information of user m on resource block n, and wn represents the Gauss-
ian white noise on resource block n.
In SCMA resource allocation, a subcarrier can be allocated to a maximum of dJ users,
which is expressed as:
M
µm,n ≤ dJ , ∀n ∈ N (9)
m=1
N
orthogonal RBs N. Moreover, the best number of users is calculated as M ∗ = .
dV
The total transmit power of the base station must not exceed the rated value P, so the
following constraint must also be satisfied:
M
N
pm,n ≤ P
m=1 n=1 (11)
pm,n ≥ 0, ∀m ∈ M, n ∈ N
Our goal is to determine the optimal binary variables µm,n for each episode. Moreover,
Eq. (1) is a complex non-convex problem to stably optimize a fixed objective for system
throughput or fairness, due to the binary constraint as well as the existence of the inter-
ference term, which is inherently challenging to solve as discussed in [32]. Furthermore,
the static policy can not adapt to the complex network conditions, sometimes it is even
opposed to the objective. Figure 3 illustrates the detailed framework.
The problem is NP-hard, which could be proved as follows.
Definition 1 The Multiple Knapsack Problem (MKP): Given a knapsack with capac-
ity V and N types of items, where each type n ∈ 1, . . . , N has a maximum availability of
sn items. Each item m of type n possesses a value wm,n and a volume vm,n. The problem
requires selecting quantities of items (up to sn for each type n) to pack into the knapsack
Fig. 3 RB allocated by different AI agents for SCMA. Wireless resource is allocated in each episode, and each
AI agent is responsible for the allocation of one RB
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 11 of 26
such that: (1) the total volume does not exceed V, and (2) the total value is maximized
[36].
Proof In problem 1, each RB n is analogous to a class. For each RB, it has a maximum
availability of sn = dJ , which is expressed as it can be allocated to a maximum of dJ user.
And the corresponding profit and the volume of the resource allocation item m, n are
wm,n = �rm,n + γ �Fm,n and vm,n = 1, respectively. For each user, it can simultaneously
receive data from a maximum of dV RBs, and thus the total volume should not exceed
V = 1 ∗ dV ∗ M . Thus, problem 2 is equivalent to MKP and is thus NP-hard.
• State: For wireless resource allocation problem, the wireless conditions, especially
instantaneous rate rn,t
ins
and average rate rn,tave
for each agent n, are contained in the
state st . Thus, each node’s state is represented in two dimensions by instantaneous
and average speed, st : (rn,t n,t . Reminding that the update of average rate involves
ins , r ave )
Fig. 4 The overall architecture of MARL algorithm for SCMA. Training each AI agent aims to obtain the
optimal neural network parameters to maximize the Q-value. At each scheduling time slot, based on the
current wireless network state and the current neural network state of each AI agent, resource allocation is
performed to obtain the optimal result
ave
ren,t = µm,n,t (�rm,n,t + γ �Fm,n,t ), ∀m ∈ M (12)
where µm,n = 1 indicates that RB n is allocated to user m. The execution part of the algo-
rithm could run continuously to get the wireless resource allocation result every TTI.
Moreover, the algorithm will insert the tuple (s, a, re, s′ ) to the experience learning mem-
ory to support the training tasks.
In the training phase, the algorithm samples a random mini-batch of transitions from
p
the experience learn memory Em. The Q-learning algorithm updates its policy follow-
ing the value iteration function that is proven to converge. It computes the expected
cumulative reward by combining the immediate reward RE and the Q-value function
of the next state. It improves the policy by greedily taking the action that maximizes the
Q-value in the future steps. Based on the centralized action-value function, the gradient
of the training network can be computed using Eq. (13), allowing gradient descent to be
performed on Eq. (14). Moreover, the smoothing factor η is set to 0.9.
ren,t , if episode terminates at step t+1
yn,t = ins , r ave ), a′ ), otherwise
ren,t + η max Q((rn,t (13)
n,t n
ins ave
descentn,t = (yn,t − Q((rn,t , rn,t ), an,t ))2 (14)
here (rn,t n,t is the state of the agent n at step t, an is the best action in the next step,
ins , r ave ) ′
method is used to achieve the optimal system scheduling result for all agents. The "L-W"
distinguishes weight-updating greedy from the basic greedy algorithm, and it focuses on
the weight variables and efficiency calculations compared with the basic greedy algorithm,
where µm,n = 1 indicates that RB n is allocated to user m. The total execution process is
showed as Algorithms 1, 2 and 3:
Algorithm 1 Multi-Agent Deep Reinforcement Learning Method
The complexity of this method focus on the L-W greedy method for each schedul-
ing episode t. In the algorithm 3, The complexity of sorting the actions matrix value
Am,n is O(MN ∗ lg(MN )), and the complexity of allocate µm,n is O(MN), where M
and N are the total number of the clients and resource blocks, respectively. Mean-
while, considering that the algorithm needs to be computed for each RB loop (with
a maximum of N calculations), the overall worst complexity of this method is
O(N ∗ (MN ∗ lg(MN ))) = O(MN 2 ∗ lg(MN )).
3.4 DeepSeek‑enpowered algorithm
In the proposed LLM-based resource allocation strategy, the resource allocation is
determined based on the instantaneous rate conditions of individual users. In this
approach, the instantaneous rate condition of each user is provided as input, and the
corresponding resource allocation outcome is generated as the output by the LLM.
Specifically, these rate conditions are fed into the model as prompts, and the resulting
resource allocation decisions are extracted from the text output of the LLM, as illus-
trated in Fig. 5. Finally, the scheduling module utilizes the proposed results generated by
DeepSeek to allocate wireless resources. Unlike traditional DRL-based methods, which
rely on custom-designed architectures, the LLM-based approach can leverage a general-
purpose language model for resource allocation. Additionally, while conventional DRL
approaches typically process numerical inputs directly, the LLM-based method requires
converting textual prompts into numerical values following a predefined format. If the
model’s output does not conform to this format, accurate resource allocation becomes
unfeasible. In this work, we employ DeepSeek [7] as the underlying LLM.
In our work, we employ two learning-based approaches to address the problem. The
first approach is a semi-automatic method, in which only the core resource allocation
logic is handled by the LLM. The logical relationships among constraints, as well as the
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 15 of 26
rate and fairness gains of RB allocation, are pre-computed. The constraints and optimi-
zation objectives related to RB allocation are provided as inputs to the LLM, and the
resulting RB resource allocation is generated as output.
Algorithm 4 Prompt 1 of SCMA resource allocation
prompt. The LLM must automatically compute the resource allocation result according
to the target function Eq. (1) and the given constraints Eq. (2) strictly.
Algorithm 5 Prompt 2 of SCMA resource allocation
As described in Algorithm 5, the LLM is required to determine the optimal resource allo-
cation strategy based on constraints 1, 2, and 3, and output the result in standard JSON
format. During each scheduling epoch, only the instantaneous rate value of one RB per user
and the current average rate of each user are provided as input to the LLM. The LLM then
computes the exponential moving average throughput, fairness value, RB allocation gain for
each user, and the final RB allocation result.
Prompt 2 requires the LLM to handle the complex relationships among the constraints
and the optimization objective and ultimately producing a result. In order to minimize pro-
cessing complexity and and enhance efficiency, we designed Prompt 1 to focus exclusively
on core resource allocation issues.
4 Experimental design
4.1 Experiment setup
To evaluate the performance of the proposed MARL and LLM-based algorithm, we imple-
ment it using a simulator system and Keras tools [37] using NVIDIA GPU A5000, and com-
pare it with the well-known PF scheduling rules [11] and MASR_SCMA method [12]. For
PF scheduling algorithm, UE should be chosen according to
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 17 of 26
ins
rm,t
m∗ = arg max ave . (15)
m∈M rm,t
with Tc as the window size for averaging, and is set to 10 in this experiment.
We consider a single cell containing multiple clients. A RB consists of several consecu-
tive subcarriers [38]. In our simulation, each subchannel comprises one RB. The simula-
tion parameters are summarized in Table 1. We use the pathloss model as [39].
For the MARL model, the neural networks (NNs) used in each DRL agent are fully
connected ones with 3 hidden layers, the first layer contains 1024 neurons, the second
layer contains 512 neurons, the third layer contains 256 neurons. ReLU function is used
as the activation function for all the hidden layers, similar to [3]. At last, the linear acti-
vation is used for value function. Moreover, we used DeepSeek reasoner model which
is upgraded to DeepSeek-V3.2-Exp, and API access method ([Link]
chat/completions) in this work.
In the experiment, we compare the PF, MARL, MASR_SCMA, MARL_SCMA and
DeepSeek scheduling algorithms 10 independent times and get similar results. The PF
and MARL algorithms primarily address resource scheduling, where each RB can be
allocated to only one user. In contrast, MASR_SCMA, MARL_SCMA_3_2 and MARL_
SCMA_2_1 focus on SCMA resource scheduling, which allows an RB to be shared
among multiple users for data transmission. MARL_SCMA_3_2 and MARL_SCMA_2_1
and (2, 1), respectively. Additionally, dV is set to 2,
correspond to the (L, W)pairs (3, 2)
N −1
dJ is calculated as dJ = = 3, where the Number of RBs N = 4 . Thus, the
dV − 1
) pairs
( d J , dV are set to (3, 2). The best number of users is calculated as
N
M∗ = = 6. We assume a cellular network with 6 users, three of whom are
dV
mobile with speeds of 3, 5, and 10 meters per second, respectively, while the remaining
users are stationary.
4.2 Evaluation metrics
We first define the evaluation metrics to assess transmission performance: (1) Average
Bitrate: This metric evaluates the efficiency of the scheduling by measuring the aver-
age bitrate. (2) Fairness: We use Jain’s fairness index to compare the variance among all
users.
Fig. 6 Jain’s fairness values using MARL, MARL_SCMA_3_2 and MARL_SCMA_2_1 method with different
numbers of parameter γ
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 19 of 26
Fig. 7 Average bitrates and CDF of PF, MARL, MASR_SCMA, MARL_SCMA_3_2 and MARL_SCMA_2_1
methods for the parameter γ = 1: a, b is the average bitrate value. c is the CDF of average bitrate. Results
for the parameter γ = 10: d, e is the average bitrate value. f is the CDF of average bitrate. Results for the
parameter γ = 100: g, h is the average bitrate values. i is the CDF of average bitrate
allocated more wireless resources, leading to a higher overall average throughput. Fur-
thermore, since SCMA enables multiple users to share wireless resources, the system
throughput of MASR_SCMA, MARL_SCMA_3_2 and MARL_SCMA_2_1 exceeds that
of both the PF and standard MARL scheduling algorithms. The MASR_SCMA has the
best performance of average bitrate and the worst performance of fairness, because it
only focuses on the average bitrate. When γ is increased to 100, as shown in Fig. 7i, the
fairness performance of the MARL algorithm improves significantly, with a noticeable
reduction in disparity among user average rates. At the same time, the average rate of
the MARL algorithm decreases by approximately 0.03 Mbps, as shown in Fig. 7d and g.
In contrast, the average rate of MARL_SCMA remains stable or even improves, while
fairness is enhanced.
Fig. 8 Jain’s fairness values using DeepSeek_prompt_1 and DeepSeek_prompt_2 methods with different
numbers of parameter γ
Fig. 9 Average bitrates and CDF of different deepseek prompts for the parameter γ = 1: a is the average
bitrate values. d is the CDF of average bitrate. Results for the parameter γ = 10: b is the average bitrate
values. e is the CDF of average bitrate. Results for the parameter γ = 100: c is the average bitrate values. f is
the CDF of average bitrate
Fig. 10 Jain’s fairness values using MARL_SCMA and DeepSeek methods with different numbers of
parameter γ
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 22 of 26
Fig. 11 Average bitrates and CDF of PF, MASR_SCMA, MARL_SCMA and DeepSeek methods for the
parameter γ = 1: a is the average bitrate values. d is the CDF of average bitrate. Results for the parameter
γ = 10: b is the average bitrate values. e is the CDF of average bitrate. Results for the parameter γ = 100: c is
the average bitrate values. f is the CDF of average bitrate
Fig. 12 Mean values and standard deviations of average rate and fairness for PF, MASR_SCMA, MARL_SCMA
and DeepSeek methods: a is the mean and standard deviation of average rate values. b is the mean and
standard deviation of fairness values
and DeepSeek begin to gradually decrease to about 1.3 Mbps when the number of users
is 10. They have better performance compared with PF method because of SCMA multi-
access gain and multiuser diversity gain. When the number of users increases to 14, the
average bitrate is about 0.3 Mbps, 1.0 Mbps, 0.9 Mbps, 0.9 Mbps for the PF, MASR_
SCMA, MARL_SCMA and DeepSeek methods correspondingly.
Fig. 13 Average rate of PF, MASR_SCMA, MARL_SCMA and DeepSeek methods under diverse user scenarios
6.2 Vulnerability
LLMs are highly sensitive to prompts. Inappropriate prompts may lead to responses in
incorrect formats or with non-compliant content. Therefore, during the training pro-
cess, it is necessary to integrate prompt engineering. By detecting and updating prompt
requirements, the correctness of the feedback results can be enhanced. Additionally,
methods such as redundant pre-sending and deduplication can be employed to improve
processing accuracy.
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 24 of 26
6.3 Privacy
Preserving the privacy of the users is the primary concern of mobile and service provid-
ers. The privacy concerns demand significant attention. LLMs are trained on massive
datasets and inadvertently memorize and regurgitate sensitive information from their
training data. Once they encounter adversarial attacks, they are quite likely to reveal the
sensitive information. Moreover, the complexity and lack of transparency in how LLMs
process and generate responses make it difficult to control and predict what information
they may disclose, thereby further obscuring privacy risks. To improve the model trans-
parency, we can develop explainable AI tools or white-box LLMs to provide insights into
the decision-making process, thereby improving its predictability and reliability. Alter-
natively, one might consider how to train models on datasets belonging to users without
sharing their input data or compromising the security of their personal information.
6.4 Security
The security of deep learning models itself in another challenge, as neural networks
are prune to adversarial attacks. Attackers can affect the training process by injecting
fake training datasets; such injection can lower the accuracy of the models and yield
wrong design, which may affect the network performance. Research in the security of
deep learning or machine learning, in general, remains shallow. To mitigate the poten-
tial harmful training data, robust data validation and sanitization techniques can be
employed to detect and remove anomalous or suspicious data before being incorporated
into the training dataset. Alternatively, adversarial training methods can be employed to
expose the model to adversarial examples during the training phase.
7 Conclusion
This paper formulates the optimization problem for sparse code multiple access in 5G
cellular networks and solves it using multi-agent reinforcement learning and large lan-
guage model-based approaches. Experimental results show that the proposed meth-
ods outperform existing resource allocation schemes, mainly due to the code-division
resource multiplexing mechanism of SCMA, and the advantages of intelligent algorithms
in perceiving and predicting environmental changes while enabling flexible adjustments.
In future work, we plan to optimize this problem in more realistic scenarios, such as pri-
oritizing how to reduce the inference latency of LLM-based resource scheduling meth-
ods by optimizing the LLM service architecture and network environment, or training
more specialized resource scheduling LLMs, thereby enhancing the practicality of the
approach. Additionally, more accurate models for 5G SCMA transmission characteris-
tics and multi-cell massive user scenarios can be established. The robustness of the pro-
posed methods in rapidly changing wireless mobile communication environments will
also be investigated.
Abbreviations
SCMA Sparse code multiple access
MARL Multi-agent deep reinforcement learning
LLM Large language model
eMBB Enhanced mobile broadband
URLLC Ultra-reliable low-latency communications
mMTC Massive machine-type communications
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 25 of 26
Acknowledgements
The authors thank the anonymous reviewers for their suggestions that have significantly enhanced the quality and
presentation of this paper.
Author contributions
Mincheng Zhao designed the scheme and algorithm and organized experiments. Mingqing Han provided SCMA
modeling methods and conducted MARL algorithm validation. Jiyuan Pan conducted the experiments, including the
implementation of the base algorithm. Ligang Du conducted the experiments, including data aggregation and process-
ing. Yongfeng Yan reviewed and revised the manuscript. Fuqiang Li reviewed and revised the manuscript, providing
access to experimental facilities.
Funding
Not applicable.
Data availability
Code was available at GitHub, [Link]
Declarations
Competing interest
The authors declare that they have no conflict of interest.
References
1. H. Nikopour, H. Baligh, in IEEE Proc. PIMRC. Sparse code multiple access (2013)
2. Z. Yuan, G. Yu, W. Li, Y. Yuan, J. Xu, in IEEE Proc. VTC. Multi-User Shared Access for Internet of Things (2016)
3. J. Wang, C. Xu, Y. Huangfu, R. Li, Y. Ge, J. Wang, in IEEE Proc. WCSP. Deep Reinforcement Learning for Scheduling in
Cellular Networks (2019)
4. C. Xu, J. Wang, T. Yu, C. Kong, Y. Huangfu, R. Li, Y. Ge, J. Wang, in IEEE Proc. WCNC. Buffer-aware Wireless Scheduling
based on Deep Reinforcement Learning (2020)
5. C. Kai, X. Meng, L. Mei, W. Huang, Multi-agent reinforcement learning based joint uplink-downlink subcarrier assign-
ment and power allocation for D2D underlay networks. Wireless Netw. 29, 891–907 (2022)
6. ChatGPT: A glimpse of the future. [Link]
7. deepseek Into the unknown. [Link]
8. J. Shao, J. Tong, Q. Wu, W. Guo, Z. Li, Z. Lin, J. Zhang, WirelessLLM: Empowering Large Language Models Towards
Wireless Intelligence (2024) arXiv:2405.17053 [[Link]]
9. M. Zhao, M. Han, L. Du, J. Pan, Y. Yan, F. Li, in IEEE Proc. WCSP. AI-based Resource Scheduling in 5G NOMA Cellular
Networks (2025)
10. P. McEnroe, S. Wang, M. Liyanage, in IEEE Proc. ITSC. Towards Latency Efficient DRL Inference: Improving UAV Obstacle
Avoidance at the Edge Through Model Compression, pp. 4242–4249 (2024)
11. G. Song, e.a., in IIEEE Proc. GLOBECOM. Adaptive resource allocation based on utility optimization in OFDM. (2003)
12. M. Cheraghy, W. Chen, in IEEE Proc. ICCC. Resource Allocation to Maximize the Average Sum Rate of the Uplink SCMA
Networks. (2021)
13. Y. Zheng, J. Cui, X. Lei, Z. Ding, P. Fan, D. Chen, Impact of factor graph on average sum rate for uplink sparse code
multiple access systems. IEEE Access 4, 6585–6590 (2016)
14. M. Zeng, N.-P. Nguyen, O.A. Dobre, Z. Ding, H.V. Poor, Spectral- and energy-efficient resource allocation for multi-
carrier uplink NOMA systems. IEEE Trans. Veh. Technol. 68(9), 9293–9296 (2019)
15. S. Jaber, W. Chen, in IEEE Proc. ICCT. Subcarrier assignment and power allocation for SCMA energy efficiency (2020)
16. X. Wang, Y. Wang, Q. Cui, K.-C. Chen, W. Ni, Machine learning enables radio resource allocation in the downlink of
ultra-low latency vehicular networks. IEEE Access 10, 44710–44723 (2022)
17. F.D. Rango, N. Cordeschi, F. Ritacco, in IEEE Proc. CCNC. Applying Q-learning approach to CSMA scheme to dynami-
cally tune the contention probability, pp. 1–4 (2021)
18. H.A. Akyildiz, O.F. Gemici, I. Hokelek, H.A. Cirpan, Hierarchical reinforcement learning based resource allocation for
RAN slicing. IEEE Access 12, 75818–75831 (2024)
Zhao et al. J. Wirel. Commun. Netw. (2026) 2026:39 Page 26 of 26
19. S. Bhardwaj, D.-S. Kim, Deep Q-learning based sparse code multiple access for ultra reliable low latency communica-
tion in industrial wireless networks. Telecommun. Syst. 83(4), 409–421 (2023)
20. X. Du, T. Wang, Q. Feng, C. Ye, T. Tao, L. Wang, Y. Shi, M. Chen, Multi-agent reinforcement learning for dynamic
resource management in 6G in-X Subnetworks. IEEE Trans. Wireless Commun. 22(3), 1900–1914 (2023)
21. D.C. Bikkasani, M.R. Yerabolu, AI-driven 5G network optimization: a comprehensive review of resource allocation,
traffic management, and dynamic network slicing. Am. J. Artif. Intell. 8(2), 55–62 (2024)
22. M. Dubey, A.K. Singh, R. Mishra, AI based resource management for 5G network slicing: history, use cases, and
research directions. Concurr. Comput.: Pract. Exp. 37(2), 66 (2024)
23. N. Kan, J. Zou, K. Tang, C. Li, N. Liu, H. Xiong, in IEEE Proc. ICASSP. Deep Reinforcement Learning-based Rate Adapta-
tion for Adaptive 360-Degree Video Streaming, pp. 4030–4034 (2019)
24. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser, I. Polosukhin, in ACM Proc. NIPS. Atten-
tion is all you need (2017)
25. D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X.B. al, DeepSeek-R1: Incentivizing Reason-
ing Capability in LLMs via Reinforcement Learning (2025) arXiv:2501.12948 [[Link]]
26. A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, e.a. Chenggang Zhao, Deepseek-v3 technical report (2025) arXiv:2412.
19437 [[Link]]
27. H. Zhou, C. Hu, D. Yuan, Y. Yuan, D. Wu, X. Liu, C. Zhang, Large Language Model (LLM)-enabled In-context Learning
for Wireless Network Optimization: A Case Study of Power Control (2025) arXiv:2408.00214 [[Link]]
28. W. Lee, J. Park, LLM-Empowered Resource Allocation in Wireless Communications Systems (2024) arXiv:2408.02944
[[Link]]
29. H. Noh, B. Shim, H.J. Yang, Adaptive Resource allocation optimization using large language models in dynamic wire-
less environments. IEEE Trans. Veh. Technol. 66, 1–6 (2025)
30. L. Bao, S. Yun, J. Lee, T.Q.S. Quek, LLM-hRIC: LLM-empowered Hierarchical RAN Intelligent Control for O-RAN (2025)
arXiv:2504.18062 [[Link]]
31. X. Peng, Y. Liu, Y. Cang, C. Cao, M. Chen, LLM-OptiRA: LLM-Driven Optimization of Resource Allocation for Non-
Convex Problems in Wireless Communications (2025) arXiv:2505.02091 [[Link]]
32. B. Di, L. Song, Y. Li, in IEEE Proc. ICC . Radio resource allocation for uplink sparse code multiple access (SCMA) net-
works using matching game (2016)
33. A.J. Goldsmith, S.-G. Chua, Variable-rate variable-power M-QAM for fading channels. IEEE Trans. Commun. 45(10),
1218–1230 (1997)
34. F.R. Kschischang, B.J. Frey, H.-A. Loeliger, Factor graphs and the sum-product algorithm. IEEE Trans. Inf. Theory 47(2),
498–519 (2001)
35. H. Zhang, S. Han, W.-X. Meng, in IEEE Proc. VTC-Fall. Multi-Stage Message Passing Algorithm for SCMA Downlink
Receiver, pp. 1–5 (2016)
36. C. Chekuri, S. Khanna, in ACM-SIAM Symposium on Discrete Algorithms. A PTAS for the multiple knapsack problem
(2000)
37. Keras: Deep Learning for humans. [Link]
38. Physical layer aspects for evolved universal terrestrial radio access (utra). 3GPP TS 25.814 bf7.1.0 (2006)
39. H. Nikopour, E. Yi, A. Bayesteh, K. Au, M. Hawryluck, H. Baligh, J. Ma, in IEEE Proc. GLOBECOM. SCMA for downlink
multiple access of 5G wireless networks (2014)
40. H. Zou, Q. Zhao, L. Bariah, M.D. Mehdi Bennis, Wireless Multi-Agent Generative AI: From Connected Intelligence to
Collective Intelligence (2023) arXiv:2307.02757 [[Link]]
41. B. Sun, Z. Huang, H. Zhao, W. Xiao, X. Zhang, Y. Li, W. Lin, in USENIX Symposium on OSDI. Llumnix: Dynamic Scheduling
for Large Language Model Serving (2024)
42. Y. Fu, L. Xue, Y. Huang, A.-O. Brabete, D. Ustiugov, Y. Patel, L. Mai, in USENIX Symposium on OSDI. ServerlessLLM: Low-
Latency Serverless Inference for Large Language Models (2024)
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.