0% found this document useful (0 votes)
31 views34 pages

Telecommunication Standardization Sector: Question(s) : Input Document Source: Title

Uploaded by

slo
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
31 views34 pages

Telecommunication Standardization Sector: Question(s) : Input Document Source: Title

Uploaded by

slo
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

INTERNATIONAL TELECOMMUNICATION UNION

FOCUS GROUP ON MACHINE LEARNING


TELECOMMUNICATION FOR FUTURE NETWORKS INCLUDING
STANDARDIZATION SECTOR 5G
STUDY PERIOD 2017-2020
ML5G-I-237-R3
Original: English
Question(s): N/A 9th meeting, (e-meeting) 2-3 June 2020
INPUT DOCUMENT
Source: FG ML5G
Title: A compilation of problem statements and resources for ITU Global Challenge on
AI/ML in 5G networks (formerly ML5G-I-223)
Contact: Xie Yuxuan Email: xieyuxuan@[Link]
China Mobile
[Link]
Contact: Jia Zihan Tel: +86 13810024426
China Mobile Email: jiazihan@[Link]
[Link]
Contact: Zhu Lin
China Mobile Email: zhulinyj@[Link]
[Link]
Contact: Mostafa Essa Email: [Link]@[Link]
Vodafone
Contact: AbdAllah Mahmoud-Eissa Email: [Link]-Eissa@[Link]
Vodafone, Egypt
Contact: Ai Ming
CICT Email: aiming@[Link]
[Link]
Contact: Francesc Wilhelmi Tel: +34 93 5422906
UPF, Spain Email: [Link]@[Link]
Aldebaro Klautau
Contact: Tel: +55 91 3201-7181
UFPA
Email: aldebaro@[Link]
Brazil
Contact: Tengfei Liu Tel: + 86 15652955883
China Unicom Fax: +010 68799999
[Link] Email: liutf24@[Link]
Contact: Wang Wei Tel: + 86 15510381035
China Unicom Fax: +010 68799999
[Link] Email: wangw200@[Link]
Contact: Jiaxin Wei Tel: + 86 13126813179
China Unicom Fax: +010 68799999
[Link] Email: weijx29@[Link]
José Suárez-Varela
Contact: Email: jsuarezv@[Link]
BNN-UPC
Spain
Albert Cabellos-Aparicio
Contact: Email: acabello@[Link]
BNN-UPC
Spain
Pere Barlet-Ros
Contact: Email: pbarlet@[Link]
BNN-UPC
Spain
Seongbok Baik
Contact: E-mail: [Link]@[Link]
KT
Contact: Dan Xu E-mail: xudan6@[Link]
China Telecom
P.R. China
Contact: Xin Guo E-mail: guoxin9@[Link]
Lenovo
P.R. China

Keywords: AI, Challenge, ML, Sandbox, Data, Resources


Abstract: This contribution compiles the list of problem statements and resources contributed
by the Focus Group members and partners towards the ITU AI/ML5G Global
Challenge. The resources are intended to be a reference list to be used for pointer
towards data, toolsets and partners to setup sandboxes for the ITU AI/ML5G
Challenge. The problem statements are intended to be analysed, short-listed and used
for the challenge to be solved by participants.

References
[ITU-T AI Challenge] ITU AI/ML in 5G Challenge website
[Link]
[ITU AI/ML Primer] ITU AI/ML 5G Challenge: Participation Guidelines (17th
April, 2020)
[ITU AI/ML Summary] ITU AI/ML 5G Challenge: Summary Slides (23rd April, 2020)

1. Introduction
[ITU AI/ML Participation Guidelines] described the proposal for ITU Global Challenge on
AI/ML in 5G networks.
Problem statements which are relevant to ITU and IMT-2020 networks are the backbone of
the challenge. They should be aligned with the theme/tracks of the challenge and should
provide enough intellectual challenge while being practical within the time period of the
challenge. They should address short term pain points for industry while pointing to long
term research directions for academia. In addition, many of them may need quality data to
solve them. This contribution collates the problem statements from our partners in a standard
format. Future steps for these problem statements are:
 analyse the submitted problem statements from our partners and colleagues,
present them for selection by the challenge management team
host the selected problem statements on the challenge website.
While discussing and disseminating the challenge with our partners, an important and
frequent question posed to us is about the relevant resources. This document contains a
collection of resources pointed to us by our members and partners in the context of ITU
ML5G global challenge. This is an attempt to compile and classify them so that it is useful to
all our partners. We invite our members and partners to add pointers to private as well as
public resources which may be of relevance to the Challenge.

2. Problem statements
NOTE 1- the structure of the list below is derived from the many discussions that we had
with partners across the globe.
NOTE 2- this list is in no specific order.
NOTE- some problem statements are “restricted problem statements”. These are available
in this document with red Title but the registration to the regional host’s website to such
problem statements and data are subject to conditions set forth by the Regional host. E.g.
currently the problem statements offered by AIIA-ITU challenge are restricted problem
statements and are available only to Chinese citizens with authorized Chinese identification.
NOTE- some problem statements use “restricted data” which is available only under a
certain conditions set forth by the Regional host.

Id ITU-ML5G-PS-TEMPLATE
Title Do not modify this particular table, this serves as a template, use
the one below.
Description NOTE 3- include a brief overview followed by a description about
the problem, its importance to IMT-2020 networks and ITU,
highlight any specific research or industry problem under
consideration.
Challenge Track NOTE 4- include a brief note on why it belongs in this track
Evaluation criteria NOTE 5- this should include the expected submission format e.g.
video, comma separated value (CSV) file, etc.
NOTE 6- this should include any currently available benchmarks.
e.g. accuracy.
Data source NOTE 7- e.g. description of private data which may be available
only under certain conditions to certain participants, pointers to
open data, pointers to simulated data.
Resources NOTE 7- e.g. simulators, APIs, lab setups, tools, algorithms, add a
link in clause 2.
Any controls or NOTE 8- e.g. this problem statement is open only to students or
restrictions academia, data is under export control, employees of XYZ
corporation cannot participate in this problem statement, any other
rules applicable for this problem, specific IPR conditions, etc.
Specification/Paper NOTE 9- e.g. arxiv link, ITU-T link to specifications, etc.
reference
Contact NOTE 10- email id or social media contact of the person who can
answer questions about this problem statement.

Id ITU-ML5G-PS-001
Title 5G+AI+AR (Zhejiang Division)
Description Background:
Augmented Reality, which enriches the real world experience through
digital means. Its realization depends on a variety of technical means
such as multimedia, three-dimensional modeling, real-time tracking and
registration, intelligent interaction, and sensing. It simulates computer-
generated virtual information such as text, images, three-dimensional
models, music, and videos, and then applies it to the real world. The
two kinds of information complement each other to achieve "augment"
of the real world.
The final breakthrough of AI technology comes from the rapid
development of big data and computing power. The combination of AI
and AR is based on data and hardware to improve perception
recognition, knowledge calculation, sameness and interaction fidelity,
so that virtual objects and real environment can have natural,
continuous and in-depth interaction with users. The deep integration of
AI and AR will enable the virtual world to be seamlessly connected to
the real world, and ultimately enable digital applications in various
industries.
Problems:
Focusing on the intelligence application demand of industry, the
artificial intelligence technology and augmented reality technology are
applied to the digital upgrade of the industrial Internet. It can be
expanded around the following two topics:
Direction 1: AI+AR entertainment application
"AI + AR Entertainment" combines 5G, AI, and AR technologies with
the consumption, entertainment, and business fields. It empowers the
entertainment market through technological means, changes existing
communication methods, strengthens the participation and interaction of
audiences, and brings people an immersive sensory experience.
"AI + AR Entertainment" includes rich industrial scenes such as city
landmarks, business district interaction, games, and digital venues.
Participants can choose any scene to play their creativity and
imagination and combine science and technology to achieve the purpose
of improving the audience experience, innovating the communication
and marketing methods, and enhancing the cultural and entertainment
content. This helps to ensure that the solution is innovative and
accessible and uses technology to help the development of the
entertainment industry.
[Link]+AR city landmark interaction:
The tourism supply side reform is shifting from relying heavily on large
resources, large capital and large commercial district to focusing on
differentiation, innovation, experience and operation. As the showcase
project of the city, the city landmark is not only the name card of the
city, but also the display window of city multiculturalism. In the city
landmark scene, the technologies of combination of virtual and reality
are introduced to provide rich and diverse interactive experience for
different groups of people, strengthen the digital operation value of
urban landmarks, rebuild the relationship between people and city, and
make the city identity more full and dynamic.
[Link]+AR commercial district interaction:
With the deepening of urbanization, the single shopping mall with a
large serving range has gradually disappeared. More and more shopping
zones and the impact of e-commerce makes it a new challenge for the
business complex to attract more young customers with strong
consumption ability and high consumption desire. In the era of 5G,
digital empowerment enables the effective connection between online
and offline. "Smart commercial district" will become a visible trend.
AI+AR technology is likely to break the space limitation of shopping
malls and create unprecedented experience upgrade and consumption
upgrade by using new interaction and communication methods.
[Link]+AR games:
Gaming is the most widely used area of AR technology at present.
Since Pokémon Go, the phenomenal-level AR interactive game, became
popular all over the world, AR games have become popular among
more and more players due to the high sense of immersion brought by
the combination of virtual and reality. When compared to the high
degree of homogeneity and repetitive patterns in traditional games,
AI+AR has great potential to bring fresh gameplay, visual expression
and new experience to games, realizing more creativity and
imagination.
[Link] digital venues:
In the era led by digital technology, more and more digital interactive
exhibition items are being used in the design of exhibition halls and
pavilions, which has also become the new vane of the industry. The
introduction of 5G and AR technologies further breaks the physical
space constraints of indoor pavilions, bringing possibilities for the
enhanced memory, experience and cognition of viewers, as well as new
market benefits.
Direction 2: AI+AR industrial Internet application
Driven by 5G technology, the Industrial Internet will develop rapidly,
and at the same time, it will bring opportunities for AI + AR
applications that are involved in multiple parts of the Industrial Internet,
and digital applications for vertical industries will emerge in succession.
This "AI + AR industry application" competition theme is closely
related to the theme of empowering the industry's digital upgrade and
improving production efficiency. It calls for solutions and products that
are innovative, useful and of practical value to industry needs.

1. AI+AR industry application - operation efficiency improvement


 AI+AR is applied to long-distance industrial maintenance,
intelligent maintenance, automation training, visual training and
other operation and maintenance fields
 AI+AR is applied to intelligent inspection, visual troubleshooting,
intelligent coordination and other inspection fields
 AI+AR is applied to intelligent research and development, remote
interaction design, 3D spatial information tracking, 3D content
interaction and other design and development display fields
 The application of AI+AR in intelligent storage, logistics transfer,
intelligent volume, intelligent sorting, automatic delivery and other
innovative applications
2. AI+AR industry application – new media
 Applying AI+AR to service media workers to improve work
efficiency
 AI+AR is applied to web-live/video/live events to achieve high-
quality mixed reality experience
3. AI+AR industry application – urban governance
 AI+AR is applied to urban security management, crisis
identification, population control, vehicle management, community
management and other fields
 AI+AR is applied to the daily operation and management of
transportation junctions (such as airports, stations and ports), such
as the innovative application of passenger guidance, public security
management, staff management, material and equipment
management, informed scheduling and other aspects.

Submitting:
Submission of works Our competition schedule is divided into two
stages: preliminary and final. The two stages need to submit different
competition works.

Challenge Track Vertical-track (invite participant to make solutions for 5G, AI and AR
application in vertical industries)
Evaluation criteria Evaluation Standard of preliminary:
Project ( full mark: 100) Evaluation Standard

Be concise, be able to effectively overview the


entire solution; have the distinct individuality,
Description of the have the creativity; Have clear ideas and goals;
project
Be able to highlight their own unique advantages;
(10 marks) The logic of the article is clear, the language is
fluent, the content is comprehensive, systematic
and scientific

Accurately describe the demand pain point,


Requirements analysis market opportunity and development orientation
and program design
of the project; The scheme involves the rationality
(40 marks) and feasibility, the completeness and the forward-
looking innovation
Operating Reasonable operation mode, clear goal planning, clear
mode/management focus; Accurately analyze the difficulty and resource
requirements in the process of project implementation
(20 marks)

Benefit evaluation The economic and social benefits of the project to the
industry
(10 marks)

Team (10 marks) Team members have relevant education and work
background; Reasonable division of work; Rigorous
organization; Proper division of property rights and equity
rights; The team has a strong ability to work under pressure,
and it is fully prepared for possible difficulties in starting a
business. The team has a strong interest in the industry

Can become China Unicom's business partner, or can well


Relevance with China
support China Unicom's existing business, or can combine
Unicom business(10
with China Unicom's key business, improve business
marks)
competitiveness

Total 100 marks

Evaluation Standard of
Project ( full mark: final:
Evaluation Standard
100)

Be concise, be able to effectively overview the entire


Data source
solution; have theNO
distinct individuality, have the creativity;
Description of the
project Resources Have clear ideas and
Notgoals; Be able to highlight their own
sure[TBD].
unique advantages; The logic of the article is clear, the
(5 marks) Any controls
languageoris fluent,This problem
the content statement is open
is comprehensive, to all
systematic, participants.
restrictions
scientific
Specification/Pape [1], [2], [3], [4] from Appendix I.
Requirements Accurately describe the demand pain points of the project,
r reference
analysis and program analyze the market opportunities, elaborate the business
design Contact model, and have certain quantitative data support;
liutf24@[Link]; Tel +86On the
15652955883; wechat: yudajiangshan
basis of preliminary scheme design, the key points and
(20 marks) wangw200@[Link]; weijx29@[Link];
details of the scheme implementation are detailed

Operating Id Reasonable operation mode, clear goal planning, clear focus;


ITU-ML5G-PS-002
mode/management Accurately analyze the difficulty and resource requirements
(10 marks) Title Fault Localization
in the process of project implementationof Loop Network Devices based on MEC Platform
(Guangdong Division)
Benefit evaluation (5 Estimate social benefits by combining with Demo project
marks) Description
examples Background:
As an information highway, the influence of network fault is expanding
Team members have relevant education
constantly. and work
The development of 5G technology brings the benefits of
background; Reasonable division of work; Rigorous
large bandwidth and wide access to this highway, but it also makes the
organization; Proper division of property
information highway rightsmore
and equity
complex. Moreover, multi-generation
Team (5 marks)
rights; The team has a strong ability to work under
technologies coexist for a long pressure,
time, which brings great challenges to
and it is fully prepared for possible difficulties in starting a
network operation. Similarly, the progress of science and technology
business. The team has a strong interest in the industry
also brings us MEC technology. MEC can be deployed in three
Can become China locations: eNodeB,
Unicom's business C-RAN
partner, and convergence ring. It can not only
or can well
Relevance with obtain the operation data of the equipment in the corresponding
support China Unicom's existing business, or can combine
China Unicom
location
with China Unicom's directly,
key business, but also
improve load the applications developed by the third-
business
business (5 marks)
competitiveness party developers. As a result, operators can provide IaaS / PaaS for the
development of special-purpose applications that need MEC features
the completion and experience
(such of the
as super DEMO
delay).
For AI+AR entertainment application, the adaptation of
On the
mobile phone terminal one hand,
experience is thethe
basicfault localization of loop network devices based
requirements,
the completion of smart glasses terminal adaptation can problem
on MEC platform solves the get of the decentralized resource
DEMO completion
5-20 points [Link] of network equipment. The decentralized devices do not
(50 marks) form an end-to-end support for business, and the basic foundation is
For AI+AR industry-Internet industry application, the
adaptation of smart glasses terminal experience is the basic
requirement, the adaptation of multiple terminals to achieve
cross-terminal platform applications can get 5-20 points plus

Total 100 marks


weak. The information technology level of the supporting process is
low, and the supporting work depends on an offline mode, with low
efficiency. On the other hand, this fault localization solves the problem
of large-scale network events will trigger a large number of single point
alarms at the same time, leading to great trouble to the fault repair
people, requiring engineers to check one by one, which is time-
consuming and labor-consuming. It is difficult to locate cross-domain
complex scenes, long fault handling time and low efficiency of cross
discipline linkage, which are the pain points of current operation and
maintenance attention. It is of great significance to enhance the
network usage awareness of MEC platform customers.

All network equipment will generate logs in the process of operation to


record the running status of the devices in real time. With the help of
MEC platform, the ability of data collection and analysis of edge
devices and the ability of AI to analyze network logs are very worthy
of study, especially for 5G network, collect the log from the terminal
and conduct real-time analysis, use AI technology to carry out
intelligent evaluation and decision-making on the operation state of the
network, and quickly and accurately define the hidden/display fault of
the current network. Thus enabling MEC platform can provide
customers with a better service.
Problems:
In order to find out the problem and find the root cause, the participants
are expected to focus on the analysis of the characteristics of the log
data provided. Combined with the network topology information
provided, it is necessary to analyze the association relationship
described in the network equipment log, extract the log template,
predict the Key log, search the keyword Association, find out the fault
points that affect the normal operation of the network, determine the
cause of the fault, and realize the network fault event playback through
the analysis of the fault transmission.
Submitting:
Preliminaries: participants need to submit two parts: one is the
algorithm model and analysis results (in csv format); the other is the
source code with annotations and descriptive documents (separately
attached with a file, in pdf format). All files are packed and
compressed into zip file, which is submitted through the email of
AIguangdong1@[Link].
1. Field description of submitted results
WEI
FIELD NAME MEANING
GHT
Flag Test data dentification A/B
RCF_device Root cause fault device 60
F_time Fault time 10
FC_log1 key log 1 15
FC_log2 key log 2 8
FC_log3 key log 3 4
FC_log4 key log 4 2
FC_log5 key log 5 1
2. Submit . csv format sample
Flag,RCF_device,F_time,FC_log1,FC_log2,FC_log3,FC_log4,FC
_log5
A,”XXX-
X”,20200XXX,"xxxxx","xxxx","xxxxx","xxxx","xxxxx"
B,”CSG-1,CSG-2”,20200219,"xxxxx","xxxx","xxxxx","xxxx","xxxxx"
All files (including csv\pdf\zip) are named in the format of participants’
title + team name, for example: " fault localization of loop network
devices based on MEC platform_China Unicom Network Research
[Link]".
Challenge Track Network-track(MEC)
Evaluation criteria The evaluation criteria are whether the prediction results of relevant
schemes are consistent with the real results. It is divided into three parts
for comprehensive scoring: The first part is the evaluation criteria F1 of
root cause fault device location; the second part is fault time point
evaluation criteria F2; the third part is fault critical log evaluation
criteria F3.
Where the root cause fault device is located accurately, F1 = 60, and
inaccurate F1 = 0. If the positioning time is within 5 minutes before and
after the standard time, then F2 = 10; if the positioning time is within 1
hour before and after, F2 = 4; if the positioning time is more than 1
hour before and after, F2 = 0. There are 5 key logs, 5 logs in the
standard answer are assigned scores according to the importance of 1,
2, 4, 8 and 15, and the corresponding scores are obtained when the
positioning results exist in the logs in the standard answer.
The analysis and processing data objects are divided into two parts: A
and B. the test data analysis results of the two parts are scored
respectively: FA = F1A + F2A + F3A, FB = F1B + F2B + F3B .
Final score: F = 0.5 * ( FA + F2B ).
Data source In this contest, A and B data are provided. These two data are
generated by network devices of different manufacturers, and the data
structure will be slightly different.
[Link] topology information
The occurrence of network fault usually has the characteristics of
propagation, and the topology related equipment will carry out fault
diffusion, which leads to the phenomenon that many devices have
faults, but usually the root cause of a fault is only one device, so it is
very necessary to analyze the fault for the network which is in constant
change.
[Link] training log + failure time log
The log is composed of unstructured text information. Although the
neighboring logs are not the same, there are always the same or similar
logs printed repeatedly. Moreover, there is a logical relationship
between different types of logs. Therefore, it is necessary to analyze the
similarity and relevance of historical logs. In addition, after the log is
transformed into structured data, statistical characteristics can be
analyzed, so as to grasp the change of equipment operation state, which
is very necessary for fault analysis. Most importantly, with the
occurrence of faults, some special logs are often printed, in which the
key information related to faults is stored.
Resources No
Any controls or restricted data
restrictions Data is under export control and employees of partners cannot
participate in this problem
Specification/Pape
No
r reference
Contact liutf24@[Link]; Tel +86 15652955883; wechat:
yudajiangshan
wangw200@[Link]; weijx29@[Link];

Id ITU-ML5G-PS-003
Title Configuration Knowledge Graph Construction of Loop Network
Devices based on MEC Architecture (Guangdong Division)
Description Background:
If knowledge is the ladder of human progress, knowledge graph is the
ladder of AI. In the past few years, Google, Microsoft, Facebook,
Alibaba, Baidu and other major companies have announced their own
knowledge graph products. Knowledge graph is the premise of
intelligence. The knowledge graph is trying to make the computer think
like human brain, which provides a new perspective and opportunity
for the interpretable AI.

By virtue of MEC's edge access capability and a large number of local


distributed computing capabilities, it is easier to build a "knowledge
graph of loop network devices configuration", "knowledge graph of
loop network devices configuration" integrates the unstructured data
information from multiple dimensions, and collects the status data of
network equipments based on the text analysis algorithm (Real time log
and network equipment alarms), configuration information, and
knowledge data (fault book, manufacturer's documents, alarm handling
book, etc.). By digitally cloning of real networks, abnormal events
driven by network changes, automatical event root cause
analysis , precise control of risks, both symptoms and treatment. The
network risks and hidden dangers can be mitigated significantly. So as
to provide high-quality network services for MEC platform customers.
Problems:
We hope that the participants will focus on the construction of network
operation knowledge graph, based on real network equipment operation
data. The framework of knowledge graph is designed according to the
logic of network structure. Analyze the relationship between network
devices, the internal protocol and business function of the devices.
According to the change of network state, the database of knowledge
graph is updated in real time, and the keyword search is supported for
knowledge interaction.
Submitting:
Preliminaries: participants need to submit two parts: one is the
algorithm model and analysis results (in csv format); the other is the
source code with annotations and descriptive documents (separately
attached with a file, in pdf format). All files are packed and compressed
into zip file, which is submitted through the email of
AIguangdong2@[Link].
1. Field description of submitted results
FIELD NAME MEANING WEIGHT
Flag Test data identification A/B
Device role classification -
Core_set
core device set
Device role classification -
Converge_set 0.4
converging device set
Device role classification -
Access_set
access device set
Relationship between and
Relations 0.6
within device
2. Submit . csv format sample
Flag,Core_set,Converge_set,Access_set,Relations
A,"A-23,A-14,…","A-09,A-16,…","A-25,A-32,…","A-23&A-
14,A-04&ospf,…"
B,"B-23,B-14,…","B-09,B-16,…","B-25,B-32,…","B-23&B-14,B-
04&ospf,…"
All files (including csv\pdf\zip) are named in the format of participants’
title + team name, for example: "configuration knowledge graph
construction of loop network devices based on MEC
architecture_China Unicom Network Research [Link]".
Challenge Track Network-track(MEC)
Evaluation criteria The evaluation criteria are whether the analysis results of relevant
schemes are consistent with real results, whether the role identification
of equipment and the relationship between them is correct. The
weighted mean value of the two aspects is used as the evaluation
criteria in this competition.
Based on the given equipment data, the participants need to classify
and identify the equipment roles. The specific calculation formula of
evaluation criteria F1 is as follows: P=TP/(TP+FP), R=TP/(TP+FN),
F2=2*P*R/(P+R). Where TP represents the set of devices identifying
the correct role, FP represents the set of devices discovering the wrong
role, FN represents the set of devices not discovering the role, P
represents the accuracy rate, and R represents the recall rate.
The specific calculation formula of evaluation criteria F2 is as follows:
P=TP/(TP+FP), R=TP/(TP+FN), F2=2*P*R/(P+R), where TP
represents the set of correct association relations, FP represents the set
of discovered incorrect association relations, FN represents the set of
undiscovered association relations, P represents Precise, and R
represents Recall.
The analysis and processing data objects are divided into A and B, and
the analysis results of the two data are scored respectively: FA =
Data source [Link] device configuration information

The configuration file contains the device instructions, which


guides a series of protection actions carried by the device to the
service, and saves all the parameter information that the device
follows during operation. It not only describes the relationship
between various business protocols within the device, but also
describes the logical and physical relationship between devices.
Through the extraction of key information and association
relationship in the configuration file, we can build a perfect
network knowledge graph and manage the network in the form
of graph database.

2. Data example

In this contest, A and B data are provided. These two data are
generated by network devices of different manufacturers, and
the data structure will be slightly different.

A data: network [Link] mask [Link]

This line of configuration command indicates: the IP address


range allocated dynamically. At the same time, when the
command in this line is under different interfaces, it indicates
the configuration restrictions on different interfaces.

2. B data: router-id [Link]

This line of configuration command indicates: configure the


router ID of OSPF process.

Resources No
Any controls or restricted data
restrictions Data is under export control and employees of partners cannot
participate in this problem
Specification/Pape
No
r reference
Contact liutf24@[Link]; Tel +86 15652955883; wechat:
yudajiangshan
wangw200@[Link]; weijx29@[Link];

Id ITU-ML5G-PS-004
Title Alarm and prevention for public health emergency based on telecom
data (Beijing Division)
Description Background:
In recent years, the worldwide outbreak of Covid-19, Ebola, MERS and
SARS posed grievous and global affects on human beings and seriously
challenged WHO as well as the health department of many countries.
Apart from the effort of health department, modern informational
technologies and data can help in health emergencies. In this problem
statement, competitors should use the tracking data of telecom users’
geographical movements and DPI information, technologies including
machine learning and big data, to propose comprehensive solutions,
product developing or advises on infrastructure for serious public
health emergencies. All these works can be considered on aspects of
epidemic surveillance, spread monitoring, precise prevention, resource
allocation, effect evaluation for health incidents.
Problems:
This topic focuses on epidemic surveillance, spread monitoring, precise
prevention, resource allocation, effect evaluation by telecom users’
tracking data and DPI information while the outbreak of Covid-19.
Participants should propose related products or solutions by using the
data, resources and developing environment provided by the
competition organizer. If participants use the data from anywhere else,
it should be taken in account that the accessibility and scalability of the
data.
Submitting:
Participants do mining and modeling based on the data provided by the
organizer and yield corresponding solutions or products. The final
submission should cover the following aspects:
Detailed introduction of the solutions or products.
The source code of mining and modeling, as well as the completed zip
file of applications; The model and explanations.
The product prototype, website or APP (optional, plus).
Challenge Track Vertical-track
Evaluation criteria Full marks 100

Problem analysis (10 marks): Whether it has a good understanding of


the core of the topic and key elements which affect the final results.
Application prospects: Whether there are demands, prospects and
potentials for the proposed solutions or products.

Solutions (25 marks): Whether the solutions are reasonable and


feasible, and meet the demand.
The use of data: Whether the data provided by organizer is fully used in
an effective way.
Innovation: Whether the works are innovative and different from
matured solutions in current industries, and whether it performs better.

Implementation (25 marks): Whether the solutions or products can be


implemented or used as a clear pattern in realistic situation and have
prospects in future.
Technical foundation: Whether it has a solid technical foundation to
carry out the solutions or products and improve them in future.
Social effect: Whether it has social effects and the ability to avoid the
risk of data breach.

Completion (40 marks): Whether the work is complete within the


allotted time and schedule and meet all the requirements.
Data source The tracking data including geographically locations and time
(directional offset) of sampled users (encrypted) in a city, the app use
data and the ownership information.
Detailed description: The format, parameter, field of the data, etc. More
details can be found in the zip file of the topic.
Resources None

Any controls or restricted data


restrictions Data is under export control and employees of partners cannot
participate in this problem
Specification/Pape None
r reference
Contact liutf24@[Link]; Tel +86 15652955883; wechat:
yudajiangshan
wangw200@[Link]; weijx29@[Link];

Id ITU-ML5G-PS-005
Title Energy-Saving Prediction of Base Station Cells in Mobile
Communication Network (Shanghai Division)
Description Background:
With the arrival of the era of mobile Internet + artificial intelligence,
Internet giants have occupied the forefront of AI in the era of AI and
IoT. Operators need to think deeply about how to give play to their
professional advantages, accelerate cross-industry integration and
enhance industry value.
Problems:
The service load of the base station is unevenly distributed in time and
space, and the power supply of the base station cannot follow the
service load of the base station, resulting in energy consumption waste.
Base station AI energy saving project is aimed at the accumulated
operation and maintenance data of operators. Taking AI as the starting
point, the base station is modeled and analyzed based on the historical
data of base station and base station cell, and the energy saving
optimization strategy is generated on the premise of ensuring the
service carrying capacity and coverage.
Submitting:
Contestants need to submit two parts of content in the preliminary
competition: one is to submit the algorithm model and the analysis
results (submitted in. CSV format); The second is the annotated core
code and documentation (a separate attached file submitted as [Link]
file). Finally, all the files are packaged and compressed into a zip file
for submission.
Challenge Track Network-track
Evaluation criteria TP (True Positive): 1 for True and 1 for prediction; FN (False
Negative): true 0, predicted 1; FP (False Positive): true is 1, prediction
is 0; TN (True Negative): 0 for True and 0 for prediction.
According to the following formula, the scores of the contestants are
calculated. According to the accuracy rate (formula 1) and recall rate
(formula 2), F1-score (formula 3) is calculated. Finally, all the
contestants are ranked according to F1-score.
P = TP/(TP+FP) (1)
R = TP/(TP+FN) (2)
F1-score = 2*P*R/(P+R) (3)

Data source This contest provides the resource data of the base station (eci, enodeb,
antenna, carrier frequency, etc.), the resource data of the base station
cell (flow, coverage, PRB, etc.), the cell phone bill information of the
base station cell, the perception data, etc.
In order to protect users' privacy and data security, the data has been
sampled and desensitized. There are null values or junk data in the data
table, and the participants need to handle it by themselves.
Resources None

Any controls or restricted data


restrictions Data is under export control and employees of partners cannot
participate in this problem
Specification/Pape None
r reference
Contact liutf24@[Link]; Tel +86 15652955883; wechat:
yudajiangshan
wangw200@[Link]; weijx29@[Link];

Id ITU-ML5G-PS-006
Title Core network KPI index anomaly detection (Shanghai Division)
Description Background:
The core network occupies a pivotal position in the entire mobile
operator network. Once the fault occurs, the service quality of the
whole network will be greatly affected. Therefore, it is necessary to
quickly discover the risk of the core network and timely eliminate the
fault before the influence scope is expanded.
Problems:
Key performance indicators (KPIs) reflect network performance and
quality. Analysis and mining of KPI can timely find the risk of network
quality deterioration. The organizer will provide the real data of a
certain operator's core network KPI during the competition, with
sampling interval of 1 hour. Contestants are required to train the model
and detect anomalies in the following 11 days (test data set) according
to the KPI data (training data set) with a history of two and a half
months, including normal labels and abnormal labels.
Submitting:
Contestants need to submit two parts of content in the preliminary
competition: one is to submit the algorithm model and the analysis
results (submitted in. CSV format); The second is the annotated core
code and documentation (a separate attached file submitted as [Link]
file). Finally, all the files are packaged and compressed into a zip file
for submission.
Challenge Track Network-track
Evaluation criteria TP (True Positive): 1 for True and 1 for prediction; FN (False
Negative): true 0, predicted 1; FP (False Positive): true is 1, prediction
is 0; TN (True Negative): 0 for True and 0 for prediction.
According to the following formula, the scores of the contestants are
calculated. According to the accuracy rate (formula 1) and recall rate
(formula 2), F1-score (formula 3) is calculated. Finally, all the
contestants are ranked according to F1-score.
P = TP/(TP+FP) (1)
R = TP/(TP+FN) (2)
F1-score = 2*P*R/(P+R) (3)
Data source [Link] of core network KPI and its meaning.
[Link] data set: data list file of 23 KPIs under different scenarios,
label 1 at abnormal moments.
[Link] data set: data list file of 23 KPIs in subsequent 11 days.
In order to protect users' privacy and data security, the data has been
sampled and desensitized. There are null values or junk data in the data
table, and the participants need to handle it by themselves.
Resources None

Any controls or restricted data


restrictions Data is under export control and employees of partners cannot
participate in this problem
Specification/Pape None
r reference
Contact liutf24@[Link]; Tel +86 15652955883; wechat:
yudajiangshan
wangw200@[Link]; weijx29@[Link];

Id ITU-ML5G-PS-008
Title Out of Service(OOS) Alarm Prediction of 4/5G Network Base
Station
Description At present, the operation and maintenance of 4/5G BS(base
station) follow a passive pattern, repairing orders will not be
generated until the out of service(OOS) fault occurs. Once the BS
is out of service, users will not be able to connect to the wireless
network, and their regular communication will be affected. In
general, there are some secondary alarms before the major alarm
(OOS alarm). Therefore, in this challenge, the participants are
expected to train an AI model using historical alarm data with
labels of major ones. By excavating the relationship between
alarms, one may use the secondary alarms to predict the
probability of the important alarm happening in a future period, so
that the operation and maintenance personnel can solve the fault in
advance and avoid network deterioration. Due to the similar
operation and maintenance mode of 4G/5G network, after the
large scale commercial use of 5G network, the AI model can be
smoothly transferred as a pre-trained model.
Challenge Track Network-track
Evaluation criteria Submit a comma separated value (CSV) file. The content includes
whether the given base station will have an out of service alarm in
the next 24 hours (or other period). The accuracy of the current
prediction model has reached 78%
Data source 4/5G network fault alarm data from China Mobile.
The data is fault alarm data of several months, including alarm
start time, alarm name, base station name, base station ID, vendor
name, city, etc.
Resources None

Any controls or restricted data


restrictions Data is under export control and employees of partners cannot
participate in this problem
Specification/Paper None
reference
Contact jiazihan@[Link]
Tel +86 13810024426

Id ITU-ML5G-PS-009
Title Radio signal coverage analysis and prediction based on UE
measurement report
Description Multiple frequency bands are usually deployed in the commercial
network to increase the network coverage and capacity. With the
increasing number of bands, inter-frequency measurements by UEs
may cause amount of signalling overhead and cost huge UE power
consumption and severely impact on running service by the data
interruption for inter-frequency measurement gap. It takes too long
time for UE to choose the proper cell to reside in. This will degrade
the network performance and UE experience. So quick inter-
frequency measurement is desired. One way to obtain the coverage
information of UEs' radio signal quickly is to divide the cell into the
grids by serving cell’s and neighbouring cell’s radio signal levels,
then locate the UE’s grid and perceive UE’s coverage information
based on statistical analysis or directly predict the inter-frequency
measurement based on the intra-frequency measurement, which can
largely reduce the numbers of UE inter-frequency measurement and
benefit for mobility based handover, load balancing, dual
connection and carrier aggregation.
Challenge Track Secure-track
Evaluation criteria Solution, criteria hasn’t been determined

Data source Training data from commercial LTE network with feedback on UE
MR data including RSRP,RSRQ,Earfcn,PCI of serving cell and
neighboring cells.
Resources No

Any controls or This problem statement is open to all participants.


restrictions

Specification/Paper No
reference
Contact xieyuxuan@[Link]

Id ITU-ML5G-PS-010
Title UE Moblity Analytics in 5G network
Description Background: In 3GPP, the NWDAF is the AI related network
function (NF), which collects data from NFs, OAM and to feedback
around 9 categories analytics to requested NFs (Please refer to
TS23.288). Within the category “UE related analytics”, the UE
mobility analytics or predications could be utilized by NFs, e.g.
AMF, SMF, EIR for some purposes, such as mobility management
parameter adjustment, detect UE been stolen, and etc.

The detailed content of “UE Mobility information” collected from


5G network, the output analytics including “UE mobility statics” and
“UE mobility predictions” could be found in TS23.288.

Problem: However, how 5GC NFs utilize aforementioned output


analytics in real 5G network would not be standardized in 3GPP
now, and has been leave to NF implementation (but how?), and the
benefits of such implementation for real network is still not clear.
It is very important to find out “how” and demonstrate the benefits.
This would help operator to deploy the NWDAF related and make
real 5G networks more intelligent.

Challenge Track Operator and vendor -track ?


1. Every team needs to provide output analytics including “UE
Evaluation criteria
mobility statics” and “UE mobility predictions”, according to the
input “UE Mobility information”.
Every team needs to provide the description of their
implementation on how to use the output analytics, and
corresponding benefits.

1. Every team itself needs to provide the “UE Mobility


Data source
information” from real 5G network or find equivalent from 4G
network.
Are there operators could possibly kindly provide the “UE
Mobility information” all the teams?
Resources TBD
Any controls or This problem statement is open to all participants.
restrictions
Specification/Paper TS23.288
reference
Contact
aiming@[Link]

Id ITU-ML5G-PS-011
Title Intelligent spectrum management for future networks
Description Background: Future networks are heterogeneous, e,g, Multi-RAT
(5G, 4G, licensed, unlicensed, fixed, mobile), Multiple platforms
(edge cloud vs. centralized cloud, VNF vs. PNF, Multiple
levels/domains (Access Network vs. Core, network slices with varied
KPI demands, various management and orchestration layers). Also
there several potential data sources e.g. (Peer-to-peer networks, NF,
applications, UEs.

Problem: In that context, spectrum management for future networks


is challenging. There is an expectation from end-customer for
coexistence and mobility across different networks (see above).
Interference management and seamless user experience across
different frequency bands used by the network is expected.
Power management in the basestation and UE is a challenge in future
networks with multi-bands.
Current methods for spectrum management has the following
disadvantages:
 The existing techniques for spectrum management are
technology specific, partly standardised + vendor-specific algorithms
implemented in scheduler.
Intra-RAT (radio access technology) standards available (e.g. X2)
Operator control is lesser, mainly driven by vendor differentiation
(scheduler and resource mangament algorithms).
Suited to less-dynamic network conditions of 4G than to future
networks of 5G and beyond.

In future networks, we would like schemes which:


1. exploit the upcoming open interfaces and data in RAN and
CN
flexible to optimize the on-demand spectrum access in tomorrow’s
networks.

In this context, the spectrum management for future networks is


proposed to be:
 Data-driven: Use data from different parts of the network
(based on VF contribution to ITU FG ML5G, Supplement 55 to
Y.3170 series)
Federated: Cross-domain exchange of data for ML (based on ITU
Y.3172, 3174)
Self-x: Adaptive, Distributed ML, decisions at the edge (to reduce
latency, communication overhead).
Level 5 intelligent: demand mapping, based on plug-in models from
operator ML marketplaces (based on ITU Y.3173).
Advantages of this approach:
 Data driven, at the same time, reduces latency,
communication overhead
Based on operator KPIs (e.g. interference reduction)
Standard (ITU-based) architecture and interfaces for interoperability
Take advantage of best ML mechanisms - Plugin models from
researchers
Challenge problem statement:
- Given a set of network bands for various types of future
networks, implement intelligent dynamic spectrum management for
future networks including IMT-2020 based on data from multiple
domains in the network.
Emphasises self-x strategy of VF.
Implements pluggable intelligence (AI models).
An optimal solution should have a model which reduces
interference between various networks, uses standard
interfaces (e.g. ITU), enables optimal operator KPIs and
imposes minimal communication overhead.

[More details, including the VF sandbox setup (lab), will be shared


later with interested participants]
Challenge Track Network track (private VF data)
Evaluation criteria In a testbed chosen by VF, shortlisted models and solutions will be
evaluated by:
1. Comparison with existing benchmarks for operator KPIs
Accuracy of models
Latency
Amount of communication overhead for the model
Data source Private data from VF (available only to VF approved candidates)
Resources Lab setsup / simulator (available only to VF approved candidates)
VF Sandbox will be setup using data and tools from VF. It will be
accessible only to selected participants nominated by VF. Data will
be hosted in a place of choice by VF. Only the data and tools
relevant to the VF problem statement will be hosted in the VF
Sandbox. Regular meeting and monitoring of participants having
access to the VF Sandbox will be done by ITU.

Any controls or Data privacy: No data should be moved from the region.
restrictions Private data from VF (available only to VF approved candidates)
Specification/Paper ITU-T Y.3172 and Y.3174
reference
Contact
[Link]-Eissa@[Link]

Id ITU-ML5G-PS-012
Title ML5G-PHY: Machine Learning Applied to the Physical Layer of
Millimeter-Wave MIMO Systems
Description The increasing complexity of configuring cellular networks
suggests that machine learning (ML) can effectively improve 5G
and future networks. One of the technologies for applications such
as vehicular systems is millimeter (mmWave) MIMO, which
enables fast exchange of data. A main challenge is that mmWave,
as initially envisioned for this application, requires the pointing of
narrow beams at both the transmitter and receiver. Taking into
account extra information such as out-of-band measurements and
vehicles positions can reduce the time needed to find the best beam
pair. Beam training is part of standards such as IEEE 802.11ad and
5G, and has also been extensively studied in the context of wireless
personal and local area networks. Hence, one of the tasks focuses
on beam-selection. Another task is channel estimation, which is
challenging due to mobility, strong attenuation in mmWave and
other issues. This challenge uses datasets obtained with the
Raymobtime methodology. The data consists of millimeter wave
(mmWave) multiple-input multiple-output (MIMO) channels,
paired with data from sensors such as LIDAR.
Challenge Track Network-track, as the challenge consists of use cases related to
signalling or management.
Evaluation criteria Top-K classification for beam selection and normalized mean
squared error for channel estimation
Data source Raymobtime datasets - [Link]
Resources None
Any controls or This Challenge is open to all participants.
restrictions
Specification/Paper [7] 5G MIMO Data for Machine Learning: Application to Beam-
reference Selection using Deep Learning, 2018 -
[Link]
[8] MmWave Vehicular Beam Training with Situational
Awareness by Machine Learning, 2018 -
[Link]
[9] LIDAR Data for Deep Learning-Based mmWave Beam-
Selection, 2019 - [Link]
[10] MIMO Channel Estimation with Non-Ideal ADCS: Deep
Learning Versus GAMP, 2019 -
[Link]
Contact Aldebaro Klautau – aldebaro@[Link]. Tel: +55 91 3201-7181

Id ITU-ML5G-PS-013
Title Improving the capacity of IEEE 802.11 WLANs through Machine
Learning
Description The usage of Machine Learning (ML) is foreseen to be a key
enabler to address the challenges podes by future wireless
networks. In IEEE 802.11 Wireless Local Area Networks
(WLANs), the major challenges will be the user’s density and lack
of coordination, which, given the current channel allocation
mechanisms, lead to sub-optimal performance. One potential
solution is the application of Dynamic Channel Bonding (DCB),
whereby an Overlapping Basic Service Set (OBSS) adapts the
spectrum to be used so that their performance is maximized.
Nevertheless, due to the complexity of massively crowded
deployments, choosing the appropriate channel width is not trivial.
Moreover, increasing the channel width entails a trade-off between
the link capacity and the quality of the link (using more bandwidth
entails a lower received signal strength and leads to a higher
contention). To address the abovementioned challenges, we
propose using Deep Learning (DL) to predict the performance that
will be obtained in an OBSS by using different channel bonding
strategies.
Challenge Track Network-track (students)
Evaluation criteria Participants should provide a .csv file containing the predicted
performance of each BSS (columns) in the different test
deployments (rows).
The evaluation of the proposed algorithms will be based on the
average squared-root error obtained from all the predictions
compared to the actual result in each type of deployment.
Data source To be provided
Resources The IEEE 802.11ax-oriented Komondor simulator [3] has been
used to generate both training and test datasets.
Any controls or This Challenge is open to all student participants.
restrictions
Specification/Paper [11] Barrachina-Muñoz, S., Wilhelmi, F., & Bellalta, B. (2019).
reference Dynamic channel bonding in spatially distributed high-density
WLANs. IEEE Transactions on Mobile Computing.
[12] Barrachina-Muñoz, S., Wilhelmi, F., & Bellalta, B. (2019). To
overlap or not to overlap: Enabling channel bonding in high-density
WLANs. Computer Networks, 152, 40-53.
[13] Barrachina-Muñoz, S., Wilhelmi, F., Selinis, I., & Bellalta, B.
(2019, April). Komondor: a wireless network simulator for next-
generation high-density WLANs. In 2019 Wireless Days (WD) (pp.
1-8). IEEE.
Contact Francesc Wilhelmi, [Link]@[Link] (+34 93 5422906)

Id ITU-ML5G-PS-014
Title Graph Neural Networking Challenge 2020
Description Network modelling is essential to construct optimization tools for
networking. For instance, an accurate network model enables to
predict the resulting performance (e.g., delay, jitter, loss) and
helps finding the configuration maximizes the network
performance according to a target policy.
Currently, network models are either based on packet-level
simulators or analytic models. The former are very costly
computationally while the latter are fast but not accurate. In this
context, Machine Learning (ML) arises as a promising solution to
build accurate network models able to operate in real time.
Recently, Graph Neural Networks (GNN) have shown a strong
potential to be integrated into commercial products for network
control and management. Early works using GNN have
demonstrated an unprecedented capability to learn from different
network characteristics that are fundamentally represented as
graphs, such as the topology, the routing configuration, or the
traffic that flows along a series of nodes in the network. In
contrast to previous ML-based solutions, GNN enables to produce
accurate predictions even in networks unseen during the training
phase. Nowadays, GNN is a hot topic in the ML field and, as such,
we are witnessing significant efforts to leverage its potential in
many different fields (e.g., chemistry, physics, social networks). In
the networking field, the application of GNN is gaining increasing
attention and, as it becomes more mature, is expected to have a
major impact in the networking industry.

Problem statement:
The goal of this challenge is to create a neural network model that
estimates performance metrics given a network snapshot.
Specifically, this model must predict the resulting per-source-
destination performance (delay, jitter, loss) given a network
topology, a routing configuration, and a source-destination traffic
matrix.

As a baseline, we provide RouteNet [5][6], a GNN architecture


recently proposed to model network performance. Participants are
encouraged to submit their own neural network architecture or
update RouteNet.
Challenge Track Network-track (design, train and test a neural network model for a
networking use case)
Evaluation criteria By means of an unlabelled dataset. Participants must label this
dataset with their neural network models and send the results in
CSV format. For the evaluation we will use a score that combines
the Mean Absolute Error (MAE) and the Mean Relative Error
(MRE) of the per-source-destination performance predictions
produced by the candidate solutions. The MAE indicates the
absolute error of the predictions with respect to the ground-truth
labels, while the MRE measures the relative distance between
them.
Data source Datasets are generated using a discrete packet-accurate network
simulator (OMNet++). The dataset contains samples simulated in
several topologies and includes hundreds of routing configurations
and traffic matrices.
The data is divided in three different sets for training, validation
and test. The validation and test datasets contain samples with
similar distributions.

You can find more details about the datasets at


[Link]
Resources - Paper, source code and tutorial of RouteNet, a reference GNN
model that can be used as a starting point for the challenge [5][6]
- User-oriented Python API to easily read and process the datasets
- Mailing list for questions and comments about the challenge [Challenge-KDN
mailing list]
- Website with a more detailed description of the challenge and the resources
provided ([Link]

Any controls or The following rules must be satisfied to participate in this


restrictions challenge:
 The solutions must be fundamentally based on neural networks
 The proposed solution cannot use network simulation tools.
 Solutions must be trained only with samples included in the
training dataset we provide. It is not allowed to use additional
data obtained from other datasets or synthetically generated.
 The challenge is open to all participants except members of the
organizing team and the research group “Barcelona Neural
Networking Center-UPC”.
 It is allowed to participate in teams. All the team members
should be announced at the beginning and will be considered
to have an equal contribution.
Final submissions must include the code of the neural network
solution proposed, the neural network model already trained, and a
brief document describing the proposed solution (1-2 pages).
Important notice: In the challenge, you may use any existing
neural network architecture (e.g., the RouteNet implementation we
provide). However, it has to be trained from scratch and it must be
clearly cited in the solution description. In the case of RouteNet it
should be cited as it is in [5].
Specification/Paper [5] Rusek, K., Suárez-Varela, J., Mestres, A., Barlet-Ros, P., &
reference Cabellos-Aparicio, A, “Unveiling the potential of Graph Neural
Networks for network modeling and optimization in SDN,” In
Proceedings of ACM SOSR, pp. 140-151, 2019.
[6] Source code an tutorial of RouteNet [6]

Contact José Suárez-Varela – jsuarezv@[Link]


Id ITU-ML5G-PS-015
Title DL-based RCA (Root Cause Analysis)
Description Background
 It is important for carriers to operate their complex network
stably.
 The stable operation includes locating and identifying the
root cause by looking at symptoms when some faults occur
on their networks.
 Vendors provide a variety of indicators (logical syslogs, or
physical LED indicators) to indicate the status of the
equipment when they release their equipment.
 When constructing a network with a small number of
equipment, it is easy to find the root cause and reasoning
the core problems.
 By making this reasoning process into a rule set, it is
possible to automate the whole inference logic, only under
the condition that the size of the network is moderately
large
 However, in a very large and complex environment of the
network, the rule-based inference method shows the very
limited performance.
 Especially in the 5G network, stability and speed are
emphasized to provide the new 5G services. Various brand-
new 5G equipment, which is physical and also virtual, is
deployed, resulting in the number of management points
increased exponentially.
 In this situation, the introduction of DL can be of great help
to the operators, because it is almost impossible to set up
the rules to pin-point the root causes in such a complex
environment.
Motivation
 For the introduction of DL technology, it is essential to
collect the training data
 However, it is almost impossible to acquire the fault
situation data much enough for training, because the fault
situations do not occur frequently in nature
 A promising alternative is to build a test-bed that simulates
5G network to simulate various fault situations and collect
data
 Using this collected data, a DL model for RCA can be
developed
 This DL model is developed in the form of a pre-trained
model through learning the characteristics of network
equipment on a test-bed
 In actual application, the characteristics of operator's
network can be fine-tuned to quickly increase accuracy and
be applied to the site
Objectives
 By implementing the following two items, the DL-based
RCA system can be implemented for complex 5G network
 1) Implement a Test-bed simulating 5G network (ML5G
test-bed)
 Composed of communication equipment common to
telecommunications operators providing 5G services
 Interworking with DB by adding data collection
function at the major management points in the
simulated network
 Configured to enable the fault scenario settings and
labeled data collection according to research needs
 2) Development of DL model optimized for RCA
 General DL model for RCA should be pre-trained
on this test-bed
 The pre-trained DL model will be fine-tuned to be
applied to the commercial environment
 Once constructed, the simulation test-bed can be
used for various purposes other than RCA

Challenge Track Network-track


Evaluation criteria
Data source TBD
Resources
Any controls or
restrictions
Specification/Paper
reference
Contact Seongbok Baik
[Link]@[Link]

Id ITU-ML5G-PS-016

Title Radio network traffic prediction

Description Background: In the 5G era, multiple new services are emerging, and
various Internet applications are constantly being enriched, which
has doubled Internet traffic. The rapid growth of traffic has brought a
lot of pressure to network bandwidth, computing, and storage. DPI
data records and presents key traffic information (data statistics start,
end time, and upstream and downstream traffic) in the application
dimension. The analysis of current network traffic models and traffic
service development trends through DPI data is the basis for solving
network congestion, improving user experience, and rationally
allocating and utilizing network resources to improve network
bandwidth utilization.
Problem: Based on the DPI traffic data collected by the big data
platform and the distance between base stations, artificial intelligence
technology can be used to analyse and predict base station traffic, in
order to provide guidance to subsequent network planning, operation
and maintenance. In this problem, we will provide a unified data set
for the participating teams. Each participating team can split the data
set into a training set, a test set, and a verification set, and use it for
training and testing of the AI algorithm model. The purpose of the
algorithm is to predict the traffic trend of base station in the future
through the historical DPI traffic data in the target area and the traffic
information in the surrounding area.
Submitting:
Competitors need to submit two parts in the preliminary competition:
one is to submit the algorithm model and analysis results (submitted
in .csv format); the other is the annotated complete code and
explanatory documents (separately attached files, submitted in .pdf
file format). Finally, all the files are packaged and compressed into a
zip file for submission.
Challenge Track Network-track

Evaluation
criteria
Evaluation criteria: (Mean Absolute Percentage Error,
MAPE),

Data source DPI traffic data collected from the current network and desensitized.

Resources No

Any controls or restricted data


restrictions Data is under export control

Specification/Pap No
er reference

Contact xudan6@[Link]

Id ITU-ML5G-PS-017
Title User-Specific Demand Prediction
Description Background:
In recent years, more and more research has pointed out that by
proactively caching content items, for which users may request, to
the edge of the network, the wireless network can reduce the
download time when users request the data. However, the benefits
of this approach relay heavily on the accuracy of user’s demand
prediction. The more accurate the user's demand prediction, the
greater the benefits of this approach.
Problem:
This topic focuses on user-specific mobile traffic demand
prediction. Competitors need to build mathematical models or
design algorithms to predict the time-varying requesting
probability of each user requesting each content item in the next 24
hours. The time-varying requesting probability can be modelled by
probability density function for continuous random variables and
probability mass function for discrete random variables. This
problem covers four sub-problems as follows.
1. Competitors need to collect datasets by themselves to solve the
problem. They can collect any dataset according to their needs,
e.g., the time spent by each user on TikTok.

2. Competitors need to predict the time-varying requesting


probability of each user requesting each APP (e.g., Youtube,
Bilibili, Baidu, Taobao, TikTok) in the next 10minutes, 1hour,
and 24 hours. As an example, the time-varying requesting
probability of each APP can be recorded as follows.
APP 00:00~01:00 01:00~02:00 … 23:00~24:00
APP 1 p1,1 p1,2 … p1,24
APP 2 p2,1 p2,2 … p2,24
… … … … …

3. Competitors need to predict the time-varying requesting


probability of each user requesting each content item in the
next 10minutes, 1hour, and 24 hours. Here the content item is
defined as a concrete file, such as a concrete video from the
Youtube platform or article from the Baidu platform. As an
example, the time-varying requesting probability can be
recorded as follows.
Content Item 00:00~01:00 01:00~02:00 … 23:00~24:00
Content Item 1 p1,1 p1,2 … p1,24
Content Item 2 p2,1 p2,2 … p2,24
… … … … …
4. Competitors need to decide the caching policy for each user.
Each user is assumed to be equipped a caching device, which
can cache 1GB data. Competitors need to design a caching
policy to determine the caching content items for next 10
minutes, 1hour, and 24hours. As an example, the caching
policy can be recorded as follows.
Content Item 00:00~01:00 01:00~02:00 … 23:00~24:00
Caching size Caching size Caching size
Content Item 1 …
x 1,1 x 1,2 x 1,24
Caching size Caching size Caching size
Content Item 2 …
x 2,1 x 2,2 x 2,24
… … … … …

Submitting:
Competitors need to solve the problem based on the data collected
by themselves. The final submission should cover the following
aspects:
1. The dataset. In order to facilitate the verification and repeat of
the experiment results, if the competitors solve the problem
based on a public dataset, they need to indicate the source and
download link for the public dataset; if the competitors solve
the problem based on the dataset collected by themselves, they
need to upload their dataset and a detailed report to explain
how they collect the data. (If the dataset is too large, a
download link for the dataset is acceptable.)
2. An annotated source code. In order to facilitate the verification
and repeat of the experiment results, competitors need to
submit all source code and corresponding explanatory
documents.
3. A detailed report. Competitors need to submit a detailed report
to explain how they process the data, build models, design
algorithms, and verify algorithm performance.
(All the files are packaged and compressed into a zip file for
submission.)
Challenge Track Network-track
Evaluation criteria 1. Competitors need upload a detailed report in PDF format to
explain how they process the data, build models, design
algorithms, and verify algorithm performance. The report will
be rated based on the innovation of solutions, the completeness
of implementation, the accuracy of results, and the writing
quality.
2. Competitors need upload a detailed file in CSV format to
record the prediction results and the caching policy.
3. Competitors can use the hit ratio, i.e., the amount of data the
user reads from the cache, to evaluate their caching policy.
Data source Competitors need to collect the data by themselves.
Resources None.
Any controls or None.
restrictions
Specification/Paper [1] M. Lee, A. F. Molisch, N. Sastry and A. Raman, "Individual
reference Preference Probability Modeling and Parameterization for Video
Content in Wireless Caching Networks," in IEEE/ACM
Transactions on Networking, vol. 27, no. 2, pp. 676-690, April
2019.
[2] B. Wu, W. Cheng, Y. Zhang, Q. Huang, J. Li, and T. Mei,
“Sequential prediction of social media popularity with deep
temporal context networks,” in Proceedings of the 26th
International Joint Conference on Artificial Intelligence
(IJCAI’17). AAAI Press, 3062–3068, 2017.
[3] S. D. Roy, T. Mei, W. Zeng and S. Li, "Towards Cross-Domain
Learning for Social Video Popularity Prediction," in IEEE
Transactions on Multimedia, vol. 15, no. 6, pp. 1255-1267, Oct.
2013.
Contact guoxin9@[Link]

2. Resources
NOTE 1- the structure of the list below is intentionally kept simple for our partners to easily
add or change it. The structure is as below:
<<type of resource: 1-line description, link, contact>>
NOTE 2- this list is in no specific order.
[RayMobTime] Data set: Raymobtime is a collection of ray-tracing datasets for wireless
communications. [Link] aldebaro@[Link]
[CUBE-AI] ML marketplace: It is an open source network AI platform developed by China
Unicom Network Technology Research Institute, which integrates AI model development,
model sharing. [Link] , liutf24@[Link]
[Adlik] Toolkit: an end-to-end optimizing framework for deep learning
models. [Link] , [Link]@[Link]
[KNOW] Challenge platform: a data challenge platform which lists several challenges and
competitions. [Link]
[SE-CAID] Data sets: An open AI research and innovation platform for networks and digital
infrastructures for industries, SMEs and academia to share a broad range of telecom data and
AI models. [Link]
[AIIA] Challenge: past competition, led by AIIA in China
[Link]
[Link]
[Link]
__biz=MzU0MTEwNjg1OA==&mid=2247487451&idx=1&sn=cb4370e9fa9d7f827dc632c7
9fe41d2d&chksm=fb2fb81ecc583108221592c69fdea3eb226da933859514dbd9fb8c15288c6f
cb392c65399ddc&mpshare=1&scene=1&srcid=&sharer_sharetime=1575542631509&sharer
_shareid=75fb4d5f665341fa1dafcbc554417e75&key=67a2c7aa29623c33d72ba777f7853d10
2e6f4db8ac8b23733613e267ce0dae54ca817de36bde651b3cf32c3a0daf055c432e46c3b8f43b
088f60edcdef801a54201eea05d0de9051201391ee19fd326f&ascene=1&uin=MjEzNjY3NDQ
5Mw%3D%3D&devicetype=Windows+7&version=62070141&lang=en&exportkey=AoB
%2BIuWyreUPRCOzxdLg0q0%3D&pass_ticket=fCmC%2FiTFfXlmGxvOLq
%2BdVPRElGBj59sZO2eVMyeABxg07Ve7tOfmRWTtKc1rmCRV
[DuReader] Challenge: past competition, includes data sets, including the largest Chinese
public domain reading comprehension dataset, DuReader
[Link]
[IUDX] Data and challenge: a research project for an open source data exchange software
platform, [Link]
[PUDX] Past challenge, Datathon  to develop innovative solutions based on India Urban Data
Exchange (IUDX), [Link]
[TI-bigdata] Data: a large dataset of 30+ kinds of data (mobile, weather, energy, etc. from
Telcom Italia big data challenge. [Link]
[TI-phone] Data: The Mobile phone activity dataset is a part of the Telecom Italia Big Data
Challenge 2014. [Link]
[MDC] Data: Mobile Data Challenge (MDC) Dataset,  restricted to non-profit organizations,
[Link] need to make a request to get a copy)
[MIRAGE] Data: MIRAGE-2019 is a human-generated dataset for mobile traffic analysis
with associated ground-truth, [Link]
[Urban-Air] Data: An air quality dataset that could be useful for
verticals [Link]
[UCR] Data: UCR STAR is built to serve the geospatial community and facilitate the finding
of public geospatial datasets to use in research and development. [Link]
[NYU] Data: NYU Metropolitan Mobile Bandwidth Trace, a.k.a. NYU-METS, is a LTE
mobile bandwidth dataset that were measured in New York City metropolitian area;
[Link]
[Omnet] Data: Challenge and dataset from comes from Omnet++ network simulator, contains
several topologies and thousands of labeled routings, traffic matrices with the corresponding
per-flow performance (delay, jitter and losses). [Link]
[GNN] Data: data sets for Unveiling the potential of GNN for network modeling and
optimization in SDN. This data set can be divided in two components: (i) the data sets used to
train the delay/jitter RoutNet models and (ii) the delay/jitter RouteNet models already trained
[Link]
network-modeling-and-optimization-in-SDN/tree/master/datasets
[Unity] [Link]
[ETSI ARF] ETSI GS ARF 003 V1.1.1 (2020-03) Augmented Reality Framework (ARF);
AR framework architecture
[Link]
df
[TH_COVID] COVID-19 Live Updates of Tencent Health is developed to track the live
updates of COVID-19, including the global pandemic trends, domestic live updates,
and overseas live updates. [Link]
[HW_NAIE] NAIE Learning Service Telecommunication scenario AI training solutions,
providing pre-consultation from now on. [Link]
[IBM_COVID] IBM has resources to share — like supercomputing power, virus tracking and
an AI assistant to answer citizens’ questions [Link]
[FB-COVID] public data sets from Facebook Data for Good [Link]
[GOOG_COVID] Google Cloud COVID-19 public dataset program: Making data freely
accessible for better public outcomes [Link]
analytics/free-public-datasets-for-covid19
Appendix I: Academic papers of interest
[1] ` "Very Long Term Field of View Prediction for 360-degree Video Streaming", Chenge
Li, Weixi Zhang, Yong Liu, and Yao Wang, 2019 IEEE Conference on Multimedia
Information Processing and Retrieval.
[2] "A Two-Tier System for On-Demand Streaming of 360 Degree Video Over Dynamic
Networks", Liyang Sun, Fanyi Duanmu, Yong Liu, Yao Wang, Hang Shi, Yinghua Ye, and
David Dai, IEEE Journal on Emerging and Selected Topics in Circuits and Systems (March
2019 )
[3] “Multi-path Multi-tier 360-degree Video Streaming in 5G Networks”, Liyang Sun,
Fanyi Duanmu, Yong Liu, Yao Wang, Hang Shi, Yinghua Ye, and David Dai, in the
Proceedings of ACM Multimedia Systems 2018 Conference (MMSys 2018),
[4] “Prioritized Buffer Control in Two-tier 360 Video Streaming”, Fanyi Duanmu,
Eymen Kurdoglu, S. Amir Hosseini, Yong Liu and Yao Wang, in the Proceedings of ACM
SIGCOMM Workshop on Virtual Reality and Augmented Reality Network, August 2017;
[5] Rusek, K., Suárez-Varela, J., Mestres, A., Barlet-Ros, P., & Cabellos-Aparicio, A, “Unveiling the
potential of Graph Neural Networks for network modeling and optimization in SDN,” In Proceedings
of ACM SOSR, pp. 140-151, 2019. [ACM SOSR] [arXiv]
[6] Source code and tutorial of RouteNet. (URL:
[Link]
[7] 5G MIMO Data for Machine Learning: Application to Beam-Selection using Deep
Learning, 2018 - [Link]
[8] MmWave Vehicular Beam Training with Situational Awareness by Machine Learning,
2018 - [Link]
[9] LIDAR Data for Deep Learning-Based mmWave Beam-Selection, 2019 -
[Link]
[10] MIMO Channel Estimation with Non-Ideal ADCS: Deep Learning Versus GAMP, 2019
- [Link]
[11] Barrachina-Muñoz, S., Wilhelmi, F., & Bellalta, B. (2019). Dynamic channel bonding in
spatially distributed high-density WLANs. IEEE Transactions on Mobile Computing.
[12] Barrachina-Muñoz, S., Wilhelmi, F., & Bellalta, B. (2019). To overlap or not to overlap:
Enabling channel bonding in high-density WLANs. Computer Networks, 152, 40-53.
[13] Barrachina-Muñoz, S., Wilhelmi, F., Selinis, I., & Bellalta, B. (2019, April). Komondor:
a wireless network simulator for next-generation high-density WLANs. In 2019 Wireless
Days (WD) (pp. 1-8). IEEE.
_____________

You might also like