0% found this document useful (0 votes)
5 views4 pages

NDM Delay Impacting Batch Processing

The 2SDS-WFAuto-Servicing Data Store batch processing was delayed from 08:00-11:30 ET on September 19, 2023 due to servers that were supposed to be decommissioned remaining connected to the application Network Data Mover. This caused the NDM process to halt when it couldn't connect to disconnected servers, stopping the batch processing. Support teams attempted to failover to BCP and place jobs on hold, but an upstream job was not correctly resumed until the next day. The root cause was identified as issues that began during a server decommissioning change request on September 18th.

Uploaded by

Srinivas Ragam
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views4 pages

NDM Delay Impacting Batch Processing

The 2SDS-WFAuto-Servicing Data Store batch processing was delayed from 08:00-11:30 ET on September 19, 2023 due to servers that were supposed to be decommissioned remaining connected to the application Network Data Mover. This caused the NDM process to halt when it couldn't connect to disconnected servers, stopping the batch processing. Support teams attempted to failover to BCP and place jobs on hold, but an upstream job was not correctly resumed until the next day. The root cause was identified as issues that began during a server decommissioning change request on September 18th.

Uploaded by

Srinivas Ragam
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Problem statement: WFAuto-SDS support reported the 2SDS-WFAuto- Servicing Data Store batch

processing was delayed from 08:00 until 11:30 ET Tuesday, September 19,2023

Impact: Data required for sending letters, and collections information was delayed being sent to our
downstream systems. Letters and Notices were delayed. Auto Loan collection dialling was put on
hold until the information could be synched. Dialling was stopped at 9:05 and resumed at 11:30.
Dialing was active beginning at 08:00 with stale data, which could have resulted in calls made
incorrectly to customers. Repossessions were not impacted.

Work Around: The batch process was restarted at 07:00 ET and completed at 14:00 ET. The WFAuto-
SRE-MMA support updated the NDM node list (removed the old NDM servers) the batch job
completed, and the dialler file was sent making the dialer available at 11:30.

Root Cause : The root cause occurred during the process of decomminsioning servers, that were still
connected to the application Network Data Mover (NDM) connecting to the data batch processing
servers. This happed during a Change Request (CHG1439543) to decommission a Dialler
server(ECTS7203W8). The NDM process stopped the data batch processing because it could not
connect to the disconnected Dialler servers (ECTS7203W8, ECTS7204W8, ECTX7202W8), when it
tried to connect with the first decommissioned server, leading support teams to attempt to failover
to BCP and place the batch job on hold.

Support teams began troubleshooting once errors occurred on Monday September 18, 2023, to
remediate batch processing issues experienced during the BCP event, one of those steps was to
place an upstream (dependent) batch job on hold. That job was not correctly set back to run until
Tuesday, September 19,2023. This was identified during planned monitoring for the long running job
from Monday/

roblem statement: WFAuto-SDS support reported the 2SDS-WFAuto- Servicing Data Store batch
processing was delayed from 08:00 until 11:30 ET Tuesday, September 19,2023 Impact: Data
required for sending letters, and collections information was delayed being sent to our downstream
systems. Letters and Notices were delayed. Auto Loan collection dialling was put on hold until the
information could be synched. Dialling was stopped at 9:05 and resumed at 11:30. Dialing was active
beginning at 08:00 with stale data, which could have resulted in calls made incorrectly to customers.
Repossessions were not impacted. Work Around: The batch process was restarted at 07:00 ET and
completed at 14:00 ET. The WFAuto-SRE-MMA support updated the NDM node list (removed the old
NDM servers) the batch job completed, and the dialler file was sent making the dialer available at
11:30. Root Cause : The root cause occurred during the process of decomminsioning servers, that
were still connected to the application Network Data Mover (NDM) connecting to the data batch
processing servers. This happed during a Change Request (CHG1439543) to decommission a Dialler
server(ECTS7203W8). The NDM process stopped the data batch processing because it could not
connect to the disconnected Dialler servers (ECTS7203W8, ECTS7204W8, ECTX7202W8), when it
tried to connect with the first decommissioned server, leading support teams to attempt to failover
to BCP and place the batch job on hold. Support teams began troubleshooting once errors occurred
on Monday September 18, 2023, to remediate batch processing issues experienced during the BCP
event, one of those steps was to place an upstream (dependent) batch job on hold. That job was not
correctly set back to run until Tuesday, September 19,2023. This was identified during planned
monitoring for the long running job from Monday/

Event Description

Problem Statement: On Tuesday, September 19, 2023, from 08:00 until 11:30 ET, WFAuto-
SDS support reported a delay in the 2SDS-WFAuto- Servicing Data Store batch processing.

Impact: The delay in batch processing resulted in the delayed transmission of data necessary
for sending letters and collections information to downstream systems. This, in turn, caused a
delay in sending letters and notices to customers. Auto Loan collection dialling had to be
temporarily put on hold until the data synchronization was completed. Dialling resumed at
11:30, but it had started at 08:00 with stale data, potentially leading to incorrect calls to
customers. Fortunately, repossessions were not impacted.

Work Around: To address the issue, the batch process was restarted at 07:00 ET and was
successfully completed at 14:00 ET. The WFAuto-SRE-MMA support team updated the
Network Data Mover (NDM) node list by removing the old NDM servers. The batch job was
subsequently completed, and the dialler file was sent, making the dialer available at 11:30.

Root Cause: The root cause of this incident was identified as the result of decommissioning
servers that were still connected to the application Network Data Mover (NDM), which
connects to the data batch processing servers. This occurred during a Change Request
(CHG1439543) to decommission a Dialler server (ECTS7203W8). The NDM process
stopped the data batch processing because it could not connect to the disconnected Dialler
servers (ECTS7203W8, ECTS7204W8, ECTX7202W8). When it attempted to connect with
the first decommissioned server, the support teams initiated a failover to BCP and placed the
batch job on hold. The issue began on Monday, September 18, 2023, during troubleshooting
for batch processing problems experienced during a BCP event. One of the troubleshooting
steps was to place an upstream (dependent) batch job on hold. Unfortunately, this job was not
correctly resumed until Tuesday, September 19, 2023. The issue was identified during
planned monitoring for the long-running job initiated on Monday.

Why 1: Why was the 2SDS-WFAuto-Servicing Data Store batch processing delayed from
08:00 until 11:30 ET on September 19, 2023?

Answer 1: The batch processing was delayed because servers that were supposed to be
decommissioned were still connected to the application Network Data Mover (NDM).

Why 2: Why were the servers that needed to be decommissioned still connected to the
NDM?

Answer 2: The servers remained connected to the NDM because the Change Request
(CHG1439543) to decommission a Dialler server (ECTS7203W8) caused the NDM process
to halt when it couldn't connect to the disconnected Dialler servers (ECTS7203W8,
ECTS7204W8, ECTX7202W8).

Why 3: Why did the NDM process stop the data batch processing when it couldn't connect to
the disconnected Dialler servers?

Answer 3: The NDM process stopped the batch processing because it relied on these Dialler
servers for data transfer, and when the connection was lost, it couldn't proceed with the data
transfer.

Why 4: Why did the support teams attempt to failover to BCP and place the batch job on
hold?

Answer 4: The support teams initiated the failover to BCP and placed the batch job on hold
in an attempt to address the data transfer issue and ensure that batch processing would
continue without interruption.

Why 5: Why was an upstream (dependent) batch job placed on hold not correctly resumed
until the following day?

Answer 5: The upstream batch job was not correctly resumed until the following day because
it was overlooked during the troubleshooting efforts on September 18, 2023, when it was
initially placed on hold to address the batch processing issues.

This 5-why analysis helps uncover the root causes of the incident, which can be used to
implement preventive actions and improve processes to prevent similar issues in the future.

To prevent similar incidents from occurring in the future, you can consider implementing the
following preventive actions based on the identified root causes:

1. Improved Change Management:


o Implement stricter validation procedures before executing changes to critical
systems. Ensure that data from business teams is thoroughly verified and
validated to prevent incorrect associations.
2. Documentation and Sign-off:
o Update change management procedures to include a documentation and sign-
off process for change requests. This step should involve comprehensive
validation and verification of data before changes are made.
3. Monitoring and Alerts:
o Implement robust monitoring systems that can quickly detect any irregularities
in batch processing or data transfer. Alerts should be configured to notify
teams when issues arise.
4. Training and Awareness:
o Provide training and awareness programs for staff involved in change
management and system administration to emphasize the importance of
thorough planning and validation.
5. Failover and Redundancy:
o Ensure that there are effective failover and redundancy mechanisms in place to
maintain critical operations when unexpected interruptions occur.
6. Corrective Action Plan:
o Develop and maintain a well-defined corrective action plan to address issues
promptly when they are identified. This should include clear procedures for
resuming batch jobs and verifying data integrity.
7. Timely Communication:
o Establish efficient communication channels between IT teams, support teams,
and management to ensure that issues are promptly addressed and resolved
without undue delay.
8. Continuous Improvement:
o Regularly review and update processes and procedures to adapt to changing
technology and business requirements. Continuously seek opportunities for
improvement in system management and data transfer procedures.
9. Root Cause Analysis:
o Conduct thorough root cause analyses for incidents and use the findings to
implement preventive actions. This should be an ongoing practice to prevent
reoccurrence of similar issues.
10. Testing and Validation:
o Prior to any significant change implementation, conduct comprehensive
testing to ensure the proposed changes will not negatively impact critical
systems and data transfers. This should include validation of connections and
data flows.

By implementing these preventive actions, you can significantly reduce the likelihood of
similar incidents occurring in the future and enhance the overall stability and reliability of
your systems and data transfer processes.

Based on the information provided, it does not appear that there was a control breakdown that
directly caused the incident. The incident seems to have been primarily driven by an
insufficient planning process during a change implementation. The breakdown occurred in
the validation of data prior to the change, and this is more related to a process issue than a
control breakdown.

However, it's essential to conduct a thorough review of existing controls and procedures to
identify any potential gaps or areas for improvement. It's possible that enhancing controls
related to validation, documentation, and change management could help prevent similar
incidents in the future.

You might also like