1)Monitor Oracle Backups
Configure Zabbix alarm to ensure backups are run and successfully for the day.
Also configure to send status in our daily oracle report email.
We need to check for the column status, just RUNNING and COMPLETED are valid Status and we should
trigger an alarm for the other Status
2)Install automatic restart script for Zabbix servers.
When the servers are restarted or have a failure ensure that the agent automatically restarts
and monitors.
TangoA [Link]
TangoB [Link]
TangoC [Link]
TangoD [Link]
TangoE10.23.2.224
TangoF10.23.2.225
DB02 [Link]
DB01 [Link]
Bacula SVR [Link]
APP01 [Link]
APP02 [Link]
API0110.23.5.12
API0210.23.5.13
WSC01 [Link]
WSC02 [Link]
LBIVR01 [Link]
LBIVR02 [Link]
INGW02 [Link]
INGW01 [Link]
HAP02 [Link]
HAP01 [Link]
AS02 [Link]
ASO1 [Link]
Completed
3)Monitor the up status of specific key processes on
the following servers below
Specifications:
1. server
IP: [Link]
OS User: lng
Processes:
Process name Parameter to be used for distinguish between
similar processes.
[Link] name 11
[Link] name 11
[Link] name 12
[Link] name 14
2. server
IP: [Link]
OS User: lng
Processes:
Process name Parameter to be used for distinguish between
similar processes.
[Link] name 21
[Link] name 21
[Link] name 22
[Link] name 24
3. server
IP: [Link]
OS User: lng
Processes:
Process name Process parameter to be used for identify
Java Dappsubprocessname=Diameter
java Dappsubprocessname=Diameter2
4. server
IP: [Link]
OS User: lng
Processes:
Process name Process parameter to be used for identify
Java Dappsubprocessname=Diameter
java Dappsubprocessname=Diameter2
4)Monitor Thread count on the following servers
below for the lng user.
Specifications:
As a preventing step for further service outage and until the root cause for the latest production issues is
resolved, please add the following view for the following VMs:
1. View name: Number of lng threads
2. Commend line to receive this information: 'ps h -Led -o user | sort | uniq -c | sort -n | grep lng'
3. Information frequency: once a minute
4. VMs: API01, API02, APP01, APP02
Alarms:
1. Set following critical alarm:
2. API01, AP02: # lng Threads is above 1200 Threads
3. APP01, APP02: # lng Threads is above 2500 Threads
Actions to be taken:
1. Case API01, API02: Follow 'Frontend failure (GD/UI/IVR/VAS)' restart procedure
2. Case APP02, APP01: Follow 'Soft Kill & Restart Procedure'
5)Monitor the CPU% Utilization and Memory
Utilization of specific key processes on the
following servers below. Also create graph views to
give a visual representation of process utilization
Specifications:
1. server
IP: [Link]
OS User: lng
Processes:
Process name Parameter to be used for distinguish between
similar processes.
[Link] name 11
[Link] name 11
[Link] name 12
[Link] name 14
2. server
IP: [Link]
OS User: lng
Processes:
Process name Parameter to be used for distinguish between
similar processes.
[Link] name 21
[Link] name 21
[Link] name 22
[Link] name 24
3. server
IP: [Link]
OS User: lng
Processes:
Process name Process parameter to be used for identify
Java Dappsubprocessname=Diameter
java Dappsubprocessname=Diameter2
4. server
IP: [Link]
OS User: lng
Processes:
Process name Process parameter to be used for identify
Java Dappsubprocessname=Diameter
java Dappsubprocessname=Diameter2
6) Add process monitoring to the Pdp and smpp
processes to monitor the amount of threads.