0% found this document useful (0 votes)
4 views41 pages

System Design

The document outlines a comprehensive guide on system design, covering algorithms, coding, business and functional requirements, and technical solution design. It details the steps in designing a solution, choosing the right database, and preparing for system design interviews, including HR and behavioral aspects. Additionally, it provides case studies of projects like Revverbank and PatientCarePortal, highlighting their business requirements and technical solutions.

Uploaded by

aswin
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views41 pages

System Design

The document outlines a comprehensive guide on system design, covering algorithms, coding, business and functional requirements, and technical solution design. It details the steps in designing a solution, choosing the right database, and preparing for system design interviews, including HR and behavioral aspects. Additionally, it provides case studies of projects like Revverbank and PatientCarePortal, highlighting their business requirements and technical solutions.

Uploaded by

aswin
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

System Design

Soundar Aswin
1 Contents
2 Algorithms and Coding.................................................................................... 3
2.1 Data structures........................................................................................ 3
2.2 Blind 75.................................................................................................... 3
2.3 Top 150.................................................................................................... 3
3 System Design................................................................................................ 3
3.1 Business Requirements............................................................................ 3
3.2 Functional Requirements..........................................................................3
3.3 Non Functional Requirements..................................................................3
3.4 Technical Solution Design........................................................................3
4 System Design Interview................................................................................ 3
4.1 Concepts.................................................................................................. 3
4.2 Solution Design........................................................................................ 3
4.3 HR............................................................................................................ 3
4.4 Behavioural.............................................................................................. 3
2 Algorithms and Coding
2.1 Data structures
2.2 Blind 75
2.3 Top 150

3 System Design
3.1 Business Requirements
Template

3.2 Functional Requirements


3.3 Non Functional Requirements
3.4 Technical Solution Design

4 System Design Interview


4.1 Concepts
4.2 Solution Design
4.2.1How to select a Technology ?
1. Project Requirements and Objectives
Choose technology that directly supports what the project needs to achieve. Make sure it fits the functional requirements and the overall
product goals.
2. Non-Functional Requirements
Check how well the technology supports key NFRs such as:
 Scalability
 Security
 Performance
 Reliability
 Latency
 Operational efficiency
 Consistency
Choose the option that meets these needs without creating bottlenecks.
3. User Experience
Select technologies that help deliver a fast, smooth, and responsive experience. Consider load times, responsiveness, and accessibility on all
devices.
4. Vendor Support and Community Strength
Pick technologies with good documentation, active community support, and strong vendor backing. This ensures you get help quickly, find solutions
easily, and avoid dead-end tools.
5. Team Skills and Expertise
Choose technologies your team already understands or can learn quickly. This reduces training time, speeds up delivery, and avoids steep learning
curve.
6. Interoperability and Integration
Ensure the technology integrates well with your existing ecosystems, APIs, data sources, and future tools. Good compatibility means fewer issues
during development and deployment.
7. Longevity and Future-Proofing
Select technologies that are widely used, actively updated, and likely to remain relevant. This helps avoid rework later and keeps your product
current with industry trends.

4.2.2Steps in a designing a solution


 Gather Business Requirement.(business goals the application needs to achieve.)
 Gather Functional Requirements(Specific features and functionalities the application must provide (e.g., user
management, payment processing, etc.).)
 Define and Agree Non Functional Requirements. ( performance, scalability, security, and compliance.)
 Choose an Application Architecture.
 Make Technology Choices.
 Ensure Security and Compliance
 Documentation and Governance
 Stakeholder Engagement and strategy
 Code using Cloud Native Design Patterns.
 Adhere to the Well Architected framework.

4.2.3Choose the right database


1. Project Requirements and Data Type
Whether the database will satisfy all your functional requirements.
 Start by understanding what type of data the system will handle:
 Structured data → SQL databases
 Semi-structured data → NoSQL databases
 Unstructured data → Object/document stores
Choose the database family that aligns with the core functional requirements.

2. Assess All Non-Functional Requirements

Evaluate each database option against your NFRs:

 Data Volume & Growth: Can it handle current and future scale?
 Performance: Is the workload read-heavy, write-heavy, or mixed?
 Concurrency: Expected number of concurrent users and operations.
 Consistency Model:
o Strict consistency (ACID)
o Eventual consistency (higher performance, distributed systems)
 Scalability:
o Vertical scaling (bigger server)
o Horizontal scaling (distributed nodes)
 Security & Compliance: Encryption, IAM, GDPR, PCI DSS.
 Integration Needs: Compatibility with your existing systems, APIs, and cloud.
 Operational Efficiency: Ease of backup, monitoring, availability, DR.
 Cost: Managed service cost, licensing, storage, operational overhead.
 Shortlist the databases (SQL, NoSQL, NewSQL, time-series) that meet these NFRs.

3. Team Skills and Operational Fit


 Choose a database your team can operate and troubleshoot effectively:
 Familiarity with SQL or NoSQL models
 Skills in performance tuning, indexing, partitions, sharding
 Ability to manage ops or use managed services (e.g., RDS, DynamoDB)
 This reduces learning effort and speeds up delivery.

4. Interoperability and Integration


Ensure the database integrates smoothly with:
 Existing apps and services
 API layers
 Analytics stack
 Cloud platform services
Good integration avoids unnecessary complexity during development.

5. Longevity and Future-Proofing


Choose databases that:
 Are actively updated
 Have strong community/vendor support
 Are widely used and stable
This ensures the solution remains reliable and scalable long-term.

6. Make the Final Decision


Weigh the trade-offs between:
 Scalability
 Performance
 Consistency
 Cost
 Operational overhead
Pick the database that satisfies the critical needs while accepting acceptable compromises.

4.3 HR
4.3.1Tell me about yourself
Thank you for asking. I work as a Principal Architect at Integrella, with around 15 years of experience in the industry with
around 10 years of being an architect, experience in cloud computing and as an architect. I have worked on multiple domains
like Core Banking, Healthcare, Digital Integration, Shipping, logistics and Retail. I am hands on across AWS and Azure GCP
and on prem platforms. My role is to lead the architecture and design for most of our client projects and also I have to lead
and collaborate with other solutions architect to define best practices, standards, naming common and architecture guidance
across out firm. My role extends to crafting technical proposals, for most of our sales pitches.

One of the best consulting firms out there,

JD matches with my current role,

And I can contribute from day one..

4.3.2Key Achievements
 Ø Created an low code no code tool named as AIS which minimised development time by 70% across for our
integration projects
 Ø Created an automated testing capability named as Comparator which minimised testing time by 50%
 Ø Principal Architect for a pro-active monitoring tool which created a revenue of around 1 million to the company
 Ø The CTO advisory started a Digital transformation strategy in which I played a key role - where we discussed
architecture styles and open source technology to save 30 to 40% Infrastructure costs for banking clients

4.4 Behavioural
1. How do you Handle Stakeholders?

1. Create a Stakeholder Mapping matrix

 Create a stakeholder mapping matrix assign role, relationship, power and influence among both internal and
external parties.

2. Create a Stakeholder Communication Plan

 Create a communication plan that defines how, when, and through which channels stakeholders will be updated.

 The plan should address Meeting schedules and formats, Project Reports, Project backlogs, decision logs, Change
Request and Approvals, Risk and issue log, and Financial reports

3. Create a Stakeholder Engagement Strategy

This will be unique for each engagement in general we should address,

 1. Identify Stakeholders Early

 Identify all stakeholders, both internal and external, Use stakeholder mapping to assess their influence and role.
High-influence and high-power stakeholders need the most attention.

 2. Understand Stakeholder Needs and Expectations

 Engage with stakeholders to understand their expectations, needs, and concerns.

 Setting mutual expectations helps avoid conflicts later.

 3. Maintain Clear and Open Communication


 Ensure transparency by providing regular updates on progress, challenges, and changes.

 4. Build Strong Relationships and Trust

 Establish trust by being reliable, responsive, and transparent.

 Engage stakeholders continuously, not just at the start.

 5. Manage Conflicts Proactively

 Conflicts among stakeholders are inevitable.

 Seek win-win solutions by understanding each party’s concerns and negotiating common ground.

 6. Involve Stakeholders in Decision-Making

 Involve key stakeholders in critical decisions. This fosters ownership and commitment to the project's success.

 7. Monitor and Adjust Stakeholder Engagement

 Regularly assess the effectiveness of your stakeholder management strategies. Are stakeholders satisfied? Are
their concerns being addressed?

 Be flexible and adjust your engagement plan as needed to respond to stakeholder.

2. Can you describe your career background and how it led you to Solution Architecture?

 I began in software development, which provided a strong technical foundation.

 During this time my mentor was a solution architect, gradually , my interest in system design grew, leading me to
Solution Architecture.

 Reasons I'm attracted to it because requires problem-solving skills and being innovative all the time.

 Also it’s a blend of Technical capability and leadership qualities.


 I am passionate about being an architect, and loved the growth from being a Technical, Solution, Senior solution and
now being a Principal architect. The journey is something I cherish

2. How do you approach problem-solving, especially in complex projects?

 My approach is methodical.

 I start by understanding the problem in depth,

 then brainstorm the possible solutions,

 communicating with key stakeholders to finalise the solution and

 Collaboration with stakeholders is crucial to ensure alignment with business objectives.

3. Can you explain the core principles of system design you follow?

Answer: I adhere to principles

 modularity,

 scalability,

 Secured and

 maintainability.

Modularity ensures system flexibility, scalability addresses growth capacity, and maintainability eases future changes.

4. How do you stay updated with emerging technologies and integrate them into your solutions?

Answer: I regularly attend industry conferences,

 ByteByteGo subscribed to their newsletter

 [Link] subscription

 Enrol to courses in Udemy


 Stay connected in Linkedin.

 Write articles in [Link]

5. Describe how you would handle a project with frequently changing requirements.

Answer: In such situations,

 I emphasize design flexibility, using Agile methodologies for iterative development with regular sprint plannning and
shorter sprint of around one to two weeks and

 frequent stakeholder feedback incorporation.

 Automating testing and CI/CD pipelines to accommodate frequent changes without compromising quality.

6. How do you ensure effective collaboration within a project team?

Answer: Effective collaboration starts with

 clear communication,

 regular team meetings,

 use of collaboration tools teams, slack, jira, confluence.

 Inclusive environment where everyone feels comfortable sharing ideas.

7. How do you manage stress in high-pressure situations?

Answer: I manage stress by

 prioritizing tasks,

 setting realistic deadlines,


 Following Kaizen approach

 taking breaks, and

 practicing mindfulness and stress-reduction techniques.

8. Can you describe your leadership style and its impact on your role as a Solution Architect?

Answer: My leadership style is

 collaborative and inclusive,

 focusing on empowering team members,

 encouraging autonomy, and being available for support,

 fostering innovation and problem-solving.

9. How do you keep up with the latest trends in cloud computing?

 ByteByteGo subscribed to their newsletter

 [Link] subscription

 Enrol to courses in Udemy

 Stay connected in Linkedin.

 Write articles in [Link]

5 Projects
5.1 Summary
Revverbank

 Business Requirements
The business requirement of the project is to consolidate transaction data obtained from the financial data cloud
platform and Starling Bank, and send it to a third-party reconciliation tool. The bank should be alerted in case of
discrepancies.

The goal is automate the reconciliation process and we need to move from their old monolithic approach to a microservices
approach.

 Solution Design
The solution design involves creating REST APIs to automate the reconciliation process. Within Integrella we categories
the API using Mulesoft Model.
We created a bunch of System APIs, Process APIs and Experience APIs. We used Apache came spring boots framework
to develop the [Link] used Kong as the API gateway. Azure AKS as the container orchestration platform and
cosmosDB as the document DB, where transactions were stored as JSON. Gitlab was used as the version control.
ARM templates for IAAS, Azure Container Registry is used for storing container images. D series VMs were hosted in a
muti-availability zone deployment.
Azure monitor and Application inisghts for monitoring,

PatientCarePortal

 Business Requirements

 The PatientCarePortal is an Integrated care record where we receive the patient details from Multiple trusts at real
time during ED admissions, we will use the information like NHS number, ODS code to pull all the details using the
GP-Connect and Summary Care record APIs.
 We use a portal to display the patient care record.
 The goal is to help the clinician in ED to view patient historical records, allergies, medical information to make real
time decisions.

Solution Design

The solution design involves creating a portal using Node JS, connectors to receive Hl7/FHIR messages in Intersytems IRIS,
Prometheus and Grafana for monitoring. We used Amazon EKS as the container orchestration platform. Gilab was used as the
version control and we used CI/CD. Terraform was used as the IAAS.
5.2 HSCNI
Business Requirements

To implement a robust Electronic Patient Record (EPR) integration for the HSCNI Encompass programme, ensuring
seamless interoperability between multiple healthcare systems. Achieve Integrated single Care Record among all the six
trusts in Northern Ireland

Functional Requirements

 Integrate EPR with Downstream Systems


 Integrate with medical devices

Non Functional Requirements

Scalability & High Availability:

 Support high transaction volumes across NHS Trusts.

 Implement active-active deployment to prevent downtime.

Performance & Latency:

 API response times should be <500ms for real-time interactions.

 Data replication must be near real-time (<2s delay).

Security & Compliance:

 NHS Data Security & Protection Toolkit (DSPT) compliance.

 GDPR-compliant patient data handling.

 Multi-factor authentication (MFA) for clinical users.

Resilience & Disaster Recovery:

 Geo-redundant architecture to ensure zero data loss.


 RPO (Recovery Point Objective): <5 minutes.

 RTO (Recovery Time Objective): <30 minutes.

5.2.1Technical Solution Design


Detailed Infrastructure Deployment

2.1 Azure Kubernetes Service (AKS) - Configuration & Deployment

Components

 Azure Kubernetes Service (AKS): Orchestrates microservices for API integration.

 Node Pools: Consists of 6 VMs for application workloads (auto-scalable to 12 nodes).

 Ingress Controller: Manages API traffic routing.

 Helm: Used for application deployment.

 Azure Container Registry (ACR): Stores Docker images.

 Azure Load Balancer: Balances requests across Kubernetes nodes.

 Azure Redis Cache: Caches API responses for performance improvement.

 Azure Managed Disks: Persistent storage for stateful workloads.

Configuration & Settings

 Ingress Controller Configuration:

apiVersion: [Link]/v1

kind: Ingress

metadata:
name: epic-api-ingress

spec:

rules:

- host: [Link]

http:

paths:

- path: /

pathType: Prefix

backend:

service:

name: epic-api-service

port:

number: 8080

 Auto-scaling Configuration:

apiVersion: autoscaling/v2beta2

kind: HorizontalPodAutoscaler

metadata:

name: epic-api-hpa

spec:

scaleTargetRef:
apiVersion: apps/v1

kind: Deployment

name: epic-api

minReplicas: 6

maxReplicas: 12

metrics:

- type: Resource

resource:

name: cpu

target:

type: Utilization

averageUtilization: 50

 Load Balancer Configuration:

resource "azurerm_lb" "hscni_lb" {

name = "hscni-load-balancer"

location = azurerm_resource_group.[Link]

resource_group_name = azurerm_resource_group.[Link]

sku = "Standard"

}
2.2 Azure API Management (APIM) - Configuration

Components

 Azure API Management (APIM): Manages API routing and security.

 OAuth 2.0 Integration: Uses Entra ID (Azure AD) authentication.

 Rate Limiting & Caching: Ensures API stability and efficiency.

Configuration & Settings

 APIM OAuth 2.0 Policy:

<authentication-managed-identity resource="[Link] />

<validate-jwt header-name="Authorization" />

 Rate Limiting Policy:

<rate-limit-by-key calls="1000" renewal-period="60" counter-key="@([Link])" />

 API Gateway Policies:

1. IP Filtering: Restricts API access to NHS Trusts.

2. Logging & Monitoring: Captures request metrics.

3. Response Caching: Uses Redis for improved response times.

4. Rewrite URL: Normalizes API endpoints.

5. CORS Policy: Allows controlled cross-origin API access.

6. Response Transformation: Standardizes API outputs.

2.3 Azure Entra ID (Azure AD) Configuration for Role-Based Access Control (RBAC)
Components

 Azure Entra ID (Azure AD): Manages authentication.

 RBAC Policies: Defines access control levels.

Configuration & Settings

 Custom Role Definition for API Access:

"Name": "HSCNI_API_Access",

"Description": "Grants read access to APIs",

"Actions": [

"[Link]/service/apis/read",

"[Link]/service/subscriptions/read"

],

"AssignableScopes": [

"/subscriptions/<subscription-id>"

2.4 Network Topology & NHS Trust Connectivity

Components

 ExpressRoute VPN Gateway: Secure NHS Trusts connection.


 Azure Private Link: Restricts external API exposure.

 VNET Peering: Connects multi-region deployments.

 Subnet Categorization:

o Application Subnet: Hosts AKS worker nodes.

o Database Subnet: Contains CosmosDB with Private Endpoint.

o API Gateway Subnet: Hosts APIM for API traffic management.

2.5 Monitoring & Alerting

Components

 Azure Monitor: Tracks system performance.

 Log Analytics: Captures API and system logs.

 Azure Alerts: Detects failures and security incidents.

Configuration & Settings

 CPU Utilization Alert:

az monitor metrics alert create --name high-cpu-alert --resource-group HSCNI-Integration --scopes


/subscriptions/<subscription-id>/resourceGroups/HSCNI-Integration --condition "avg Percentage CPU > 80" --window-size 5m
--evaluation-frequency 1m --action-group alert-action-group

2.6 Disaster Recovery (DR) & High Availability (HA)

Components
 Azure Traffic Manager: Handles regional failover.

 Cosmos DB Auto-Failover: Ensures database redundancy.

 Azure Backup Vault: Stores critical backups.

Configuration & Settings

 Disaster Recovery Testing Steps:

1. Simulate regional failure.

2. Validate auto-failover.

3. Verify data integrity checks.

4. Restore primary region and re-route traffic.

3. Conclusion & Next Steps

This Low-Level Design (LLD) ensures:

 Scalability: AKS-based microservices with auto-scaling.

 Security: OAuth 2.0 authentication and RBAC-based access control.

 High Availability: Traffic Manager ensures regional failover.

 Compliance: Meets NHS, DTAC, and GDPR standards.

Next Steps:

1. Validate security policies with NHS Trusts.

2. Conduct load testing for API performance.

3. Schedule quarterly disaster recovery drills.


End of Document

API Data Flow Diagram for HSCNI EPIC Patient API Solution

Here is the API Data Flow Diagram for the HSCNI EPIC Patient API Solution. This illustrates how data flows between the
NHS Trust System, APIM, AKS, Redis Cache, CosmosDB (FHIR), and Traffic Manager.

Diagram Explanation

1️⃣ NHS Trust System sends a request to Azure API Management (APIM).
2️⃣ APIM forwards the request to Azure Kubernetes Service (AKS).
3️⃣ **APIM authenticates the request via Azure Entra ID (OAuth 2.0).
4️⃣ AKS checks Azure Redis Cache for a cached response to reduce latency.
5️⃣ If no cached response is found, AKS queries Azure Cosmos DB (FHIR) for patient data.
6️⃣ Azure Cosmos DB (FHIR) processes the request and sends the response.
7️⃣ Azure Traffic Manager ensures high availability and directs requests across regions.
8️⃣ In case of a failure, Traffic Manager redirects traffic to a backup region, activating Azure Backup Vault (DR).

Disaster Recovery (DR) Environment for HSCNI EPIC Patient API Solution

The Disaster Recovery (DR) environment ensures high availability, resilience, and minimal downtime for the HSCNI
EPIC Patient API Solution in case of failures. It is designed using multi-region redundancy, automated failover, and
backup restoration.

1️⃣ DR Environment Components

The DR environment is spread across two Azure regions:

 Primary Region: UK South (Active)

 Secondary Region: UK North (Passive, Failover-ready)

Each region has a replica of critical components:

🔹 Compute & API Infrastructure

Component Primary (UK Secondary (UK Failover Mechanism


South) North)
Azure Kubernetes Active Standby Traffic Manager redirects API
Service (AKS) requests

Azure API Management Active Standby Switches API requests to UK


(APIM) North

Azure Entra ID (OAuth Active Active No failover needed (global


2.0) service)

Azure Load Balancer Active Standby Redirects traffic to AKS in UK


North
🔹 Data & Caching Infrastructure

Component Primary (UK Secondary (UK Failover Mechanism


South) North)

Azure Cosmos DB Read/Write Read-only Auto-failover via multi-region


(FHIR API) replication

Azure Redis Cache Active Standby Reloads data from Cosmos DB

Azure Backup Vault Active Standby Restores missing data in case of


corruption
🔹 Networking & Traffic Management

Component Primary (UK Secondary (UK Failover Mechanism


South) North)

Azure Traffic Manager Active Active DNS-based routing to the available


region
ExpressRoute VPN Active Standby NHS Trust traffic redirected via
Gateway alternative VPN

Azure Private Link Active Standby Secure data exchange remains


available

2️⃣ How DR Environment Works (Normal vs. Failure Scenarios)

✅ Normal Operations (No DR Event)

1️⃣ API requests from NHS Trust Systems reach Azure API Management (APIM) in UK South.
2️⃣ APIM authenticates via Azure Entra ID (OAuth 2.0).
3️⃣ AKS microservices in UK South process API requests.
4️⃣ If data is cached, Azure Redis Cache serves responses.
5️⃣ If not, Azure Cosmos DB (UK South) retrieves and stores patient data.
6️⃣ Azure Traffic Manager directs requests to UK South (Primary Region).

📌 Primary Region handles all API traffic unless a failure occurs.

🚨 DR Event - Primary Region Failure (UK South Unavailable)

If UK South fails due to network outage, power failure, or service degradation, the DR environment reacts as follows:

Step 1: Azure Traffic Manager Detects the Failure

 Azure Traffic Manager performs health checks on API endpoints.

 If UK South is unresponsive, Traffic Manager updates DNS records to point traffic to UK North.

 NHS Trusts automatically start sending API requests to APIM in UK North.

Step 2: Cosmos DB Automatic Failover


 Azure Cosmos DB is multi-region replicated.

 Read/Write operations shift from UK South → UK North automatically.

 Data consistency is maintained using session-level consistency policies.

Step 3: Azure API Management (APIM) Switches Backend

 APIM in UK North starts handling API requests.

 All API policies (OAuth 2.0, rate limiting, logging) remain intact.

 NHS Trust systems continue to interact with FHIR APIs without manual intervention.

Step 4: Redis Cache Reload

 Azure Redis Cache in UK North reloads frequently accessed data from Cosmos DB.

 This minimizes database load and speeds up API responses.

Step 5: ExpressRoute VPN Redirects NHS Trust Traffic

 If VPN Gateway in UK South fails, NHS Trust traffic is rerouted via the UK North VPN Gateway.

 Azure Private Link ensures secure communication remains intact.

3️⃣ DR Testing & Validation Process

To ensure the DR environment works efficiently, quarterly DR testing drills are conducted.

✅ Step 1: Simulate Region Failure

 Disable AKS, APIM, and Cosmos DB in UK South.

 Force Cosmos DB failover to UK North:

sh
CopyEdit

az cosmosdb failover-priority-change --resource-group HSCNI-Integration --account-name hscni-cosmosdb --failover-policies


"UK North=0" "UK South=1"

 Monitor how quickly Traffic Manager redirects traffic to UK North.

✅ Step 2: Monitor API & Database Failover

 Send API requests to APIM and validate latency & availability.

 Check if Cosmos DB in UK North is handling all queries.

✅ Step 3: Validate Data Consistency

 Compare patient data responses between UK South backups & UK North active instance.

 Run data integrity checks to ensure data was not lost.

✅ Step 4: Restore Primary Region

 Restart AKS, APIM, and Redis Cache in UK South.

 Use Azure Backup Vault to restore missing data.

 Revert Cosmos DB failover priority to make UK South the primary:

sh

CopyEdit

az cosmosdb failover-priority-change --resource-group HSCNI-Integration --account-name hscni-cosmosdb --failover-policies


"UK South=0" "UK North=1"

4️⃣ Key Benefits of DR Environment


✅ 99.99% uptime – thanks to multi-region failover.
✅ Automated failover – No manual intervention needed.
✅ Traffic rerouting in seconds – DNS-based failover via Traffic Manager.
✅ Continuous data availability – thanks to Cosmos DB replication.
✅ Proactive DR Testing – to ensure failover readiness.

5️⃣ DR Environment Summary

Component Primary (UK Secondary (UK Failover Mechanism


South) North)

Azure API Management Active Standby DNS routing by Traffic


(APIM) Manager

Azure Kubernetes Active Standby Traffic redirected to UK


Service (AKS) North AKS

Azure Cosmos DB (FHIR Read/Write Read-only Automatic failover


API)

Azure Redis Cache Active Standby Reloads from Cosmos DB

Azure Load Balancer Active Standby Redirects API requests

Azure Traffic Manager Active Active Handles API endpoint


failover

Azure Backup Vault Active Standby Restores missing data


5.3 Revverbank
 5.1 Revverbank Reconciliation Solution

 Business Requirements

The business requirement of the project is to consolidate transaction data obtained from the financial data cloud
platform and Starling Bank, and send it to a third-party reconciliation tool. The bank should be alerted in case of
discrepancies.

 Functional Requirements

 Reconciliation Reports: Generate and send daily reconciliation reports (future goal: near real-time).

 Discrepancy Alerts: Notify the bank if there are issues in the reconciliation report.

 Third-Party Integrations:

o General Ledger: Integrate with the third-party banking platform.

o Starling Bank API: Retrieve payment details.

o Reconciliation Tool: Generate reconciliation reports.

 5.1.1 Technical Solution Design

 Compute Considerations

 D-Series VMs for scalable compute.

 Azure Kubernetes Service (AKS): Used for container orchestration.

 Database Considerations

 Azure Cosmos DB: Used for API-related storage and semi-structured data.
 Azure SQL Server: Used for core banking data storage.

 Automation

 ARM Templates: Used to create infrastructure dynamically.

 Infrastructure Scripting: Scripts to manage spinning up/down across Dev, Test, and Staging environments.

 Disaster Recovery (DR): Automated DR script for triggering during failures.

 Development

 Apache Camel & Spring Boot: Core development framework for integration and processing.

 Integration

 Azure Logic Apps: For integration between services.

 Azure Event Grid: Used for loading files from Blob Storage.

 Azure Service Bus: Manages messaging between components.

 Networking

 Dedicated VNet with separate subnets per service.

 Security

 Azure Key Vault: Used for securing secrets and sensitive data.

 Monitoring

 Prometheus & Grafana: For real-time monitoring and alerting.

 5.1.2 Disaster Recovery & Infrastructure Management

 Azure Infrastructure Model


Azure provides four levels of management and organization:

 Management Groups: Policies and access control across multiple subscriptions.

 Subscriptions: Logical units for managing costs and resource limits.

 Resource Groups: Logical containers for Azure resources.

 Resources: Individual services (e.g., VMs, storage, databases).

 Step 1: Management Groups & Subscriptions

 Separate Management Groups for Development, Testing, Staging, and Production.

 Dedicated Subscriptions per Management Group for better governance.

 Step 2: Resource Groups & Core Infrastructure

 Create dedicated Resource Groups for each environment.

o RG-Network-Dev, RG-Network-Test, RG-Network-Prod for network resources.

o RG-Compute, RG-Storage, RG-Data for respective infrastructure components.

 Step 3: Networking & Security

 Virtual Networks (VNets): Segregated VNets per environment (e.g., VNet-Dev, VNet-Test, VNet-Prod).

 Network Security Groups (NSGs): Enforced at the subnet level.

 Azure Firewall: Used for external traffic control.

 ExpressRoute/VPN Gateway: For private banking connectivity.

 Step 4: Identity & Access Management

 Azure Entra ID (Azure AD): Centralized identity provider.

 RBAC Policies: Implemented at Resource Group level.


 Privileged Access Management (PIM): Just-in-time access control for admin roles.

 Step 5: Disaster Recovery Strategy

 Azure Site Recovery (ASR): Automates DR failover.

 Multi-Region Replication: Geo-redundancy across Azure regions.

 Quarterly DR Testing: Ensure preparedness and compliance.

 Traffic Manager Failover: DNS-based routing for high availability.

 Summary of RPO and RTO Strategies

Service DR Strategy RPO RTO

Logic Near- Minute


Active-active or redeployment
Apps zero s

Service Geo-DR with secondary Near- Minute


Bus namespace zero s

Event Near- Second


Multi-region topics & retries
Grid zero s

Cosmos Near- Second


Multi-region replication
DB zero s

 5.2 Conclusion & Next Steps

This High-Level Design (HLD) ensures: ✅ Scalability: AKS-based microservices with auto-scaling. ✅ Security: OAuth 2.0
authentication and RBAC-based access control. ✅ High Availability: Traffic Manager ensures regional failover. ✅
Compliance: Meets FCA, PCI DSS, and GDPR standards.

 Next Steps:
11️⃣Validate security policies2️⃣
2with banking regulators. 2 Conduct performance testing3️⃣
3for API performance. 3 Schedule
quarterly disaster recovery drills.

5.4 iSentry
 5.1 iSentry AWS Implementation

 Business Requirements

The iSentry solution is designed to proactively monitor and alert support engineers on integration performance issues.
The solution must be migrated from GCP to AWS, ensuring:

 Seamless monitoring of Rhapsody Integration Engine logs and metrics.

 Proactive alerting based on queue/message metrics and wrapper logs.

 Secure communication with NHS Trusts through VPN tunnels.

 Scalable and resilient architecture to handle increasing integration workloads.

 Functional Requirements

 Log Collection: Collect logs from multiple NHS Trusts for monitoring and alerting.

 Alerting System: Generate alerts based on predefined message queue parameters.

 Storage & Analysis: Store logs in Elasticsearch (Amazon OpenSearch Service) for indexing and querying.

 Secure Data Transfer: Ensure secure data flow through an AWS Site-to-Site VPN.

 User Interface: Provide a dashboard (Amazon QuickSight) for visualizing integration status and alerts.

 5.1.1 Technical Solution Design


 Infrastructure Deployment in AWS

 1. Compute & Orchestration

 Components

 Amazon Elastic Kubernetes Service (EKS): Manages microservices for log collection and alerting.

 Fargate for EKS: Provides serverless container execution.

 Amazon EC2 (D-Series): Used for running high-performance monitoring agents.

 2. Data Processing & Storage

 Components

 Amazon Kinesis Data Streams: Ingests real-time logs from NHS Trusts.

 AWS Lambda: Processes logs and filters relevant event triggers.

 Amazon OpenSearch Service (Elasticsearch): Stores indexed logs for querying.

 Amazon S3: Stores raw log files for long-term retention.

 3. Alerting & Visualization

 Components

 Amazon CloudWatch: Collects performance metrics and triggers alerts.

 Amazon SNS & EventBridge: Sends alerts to support engineers via email and SMS.

 Amazon QuickSight: Provides a dashboard for visualization and reporting.

 4. Security & Networking

 Components

 AWS Site-to-Site VPN: Ensures secure data transfer between NHS and AWS.
 AWS PrivateLink: Secures communication between microservices.

 AWS WAF & Shield: Protects against DDoS attacks.

 AWS Secrets Manager: Secures authentication credentials.

 5.1.2 Monitoring & Disaster Recovery

 Monitoring & Logging

 Components

 Amazon CloudWatch Logs: Tracks system performance and errors.

 AWS X-Ray: Provides distributed tracing for debugging.

 AWS Config & GuardDuty: Ensures compliance and detects security anomalies.

 Disaster Recovery (DR) & High Availability (HA)

 Components

 Multi-Region Deployment: Active in EU-West-1 (Primary) & EU-West-2 (Failover).

 Amazon RDS Multi-AZ: Ensures data availability for structured data.

 Amazon S3 Cross-Region Replication: Maintains backup logs in a secondary region.

 AWS Backup & Restore: Provides scheduled snapshots and point-in-time recovery.

 Disaster Recovery Testing Steps

1. Simulate AWS region failure.

2. Validate auto-failover for OpenSearch and RDS.

3. Monitor failover recovery times using Route 53 health checks.


4. Restore traffic to the primary region upon resolution.

 5.2 Conclusion & Next Steps

This AWS-based iSentry HLD ensures: ✅ Scalability: EKS-based architecture with auto-scaling. ✅ Security: Enforced
through AWS Site-to-Site VPN & PrivateLink. ✅ Resilience: Multi-region failover ensures 99.99% uptime. ✅
Compliance: Meets NHS DSPT, GDPR, and security best practices.

 Next Steps:

1️⃣Validate AWS security policies with NHS Trusts. 2️⃣Perform performance testing on Kinesis log ingestion. 3️⃣
Schedule quarterly disaster recovery drills.

5.5 TTK Comparator

6 AWS
6.1 Identity and Access
Feature AWS IAM AWS Cognito AWS SSO AWS STS AWS AWS Directory
Organizations Service

Explanati Centralized user User Single Sign-On Temporary Multi-account Managed


on access authentication service for AWS credentials management and Microsoft Active
management for and authorization accounts and for federated governance Directory
AWS resources service for applications access service
applications
Security Policy-based Token-based Federated access Temporary Centralized LDAP/Active
Modelss permissions authentication management token-based account Directory
security governance integration

Scalability Supports large- Handles millions Scales across Scales with Manages multiple Scales for
scale of users multiple AWS IAM roles AWS accounts corporate
enterprises accounts directory
services

Integratio Works with AWS Integrates with Integrates with Works with Integrates with Works with on-
n services via IAM social identity AWS IAM and all AWS accounts premises Active
roles providers Organizations external Directory
identity
providers

Cost Free with AWS Pay-as-you-go for Included with Free with Free for basic Pay for
usage user AWS AWS usage features, paid for managed
authentication Organizations advanced directory
governance services

Use Case Fine-grained Secure Unified access Short-lived Centralized LDAP


permissions for authentication for management credentials account and authentication
AWS resources apps and mobile across AWS for external billing for corporate
users accounts users management workloads

6.2 Networking
Feature VPC Peering Transit VPC Private Site-to- Client Direct Network
Gateway Endpoints Link Site VPN VPN Connect Firewall
Explanation Enables direct Centralized Allows Enables Secure Provides Dedicated Stateful,
communicatio hub for private private connection remote private managed
n between two connecting access to connecti between on- users connection firewall
VPCs multiple VPCs AWS vity prem and secure from on- service for
and on- services between AWS via access to prem to controlling
premises without VPCs and IPSec AWS AWS with traffic in/out of
networks traversing AWS tunnels resources low latency VPC
the services via VPN
internet via ENIs

Connectivity One-to-one One-to-many Used for Used for Connects Remote High- Controls and
Scope connection and many-to- AWS third- on-prem to access speed, filters
between two many services party AWS VPCs VPN for dedicated inbound/outbo
VPCs connectivity like S3, SaaS and via VPN clients connection und traffic for
for multiple DynamoDB internal tunnels from on- security
VPCs services prem to
AWS

Performance Low latency High- No Low Medium Medium Low Minimal


but does not performance additional latency latency latency latency performance
scale well for and scalable latency and high (depends on (depends and high impact but
large networks throughp internet) on client bandwidth adds security
ut network) overhead

Security No encryption Provides Uses AWS Uses IAM Encrypted Encrypted More Deep packet
by default, centralized IAM and (IPSec) (TLS) secure inspection,
requires security policies Security than VPN rule-based
additional management Groups (private traffic control
setup link)

Cost Charges for Pay for data Pay per Pay per VPN Pay per Pay per Pay per rule
cross-region transfer and endpoint PrivateLi connection client VPN dedicated and GB
data transfer attachments usage nk charges connectio connection inspected
connecti n
on
Management Simple for Easier Low Low Requires Requires Requires Needs
Complexity small management complexity complexi VPN setup & VPN client provisionin continuous
networks, for multiple ty maintenanc software g and rule tuning
complex as VPCs and e & maintenan
scale hybrid setups managem ce
increases ent
Scalability Limited, Scalable, Scales with Scales Limited by Scales High Highly
manual supports service with VPN tunnels with user scalability scalable
configurations 5,000+ usage service (max 10 per connectio for security layer
needed attachments usage VGW) ns enterprise
needs

6.3 EC2
6.4 EKS
Amazon Elastic Kubernetes Service (Amazon EKS) is a managed Kubernetes service that makes it easy for you to run
Kubernetes on AWS and on-premises. Kubernetes is an open-source system for automating deployment, scaling, and
management of containerized applications. Amazon EKS is certified Kubernetes-conformant, so existing applications that run
on upstream Kubernetes are compatible with Amazon EKS.

Amazon EKS automatically manages the availability and scalability of the Kubernetes control plane nodes responsible for
scheduling containers, managing application availability, storing cluster data, and other key tasks.

Amazon EKS lets you run your Kubernetes applications on both Amazon Elastic Compute Cloud (Amazon EC2) and AWS
Fargate. With Amazon EKS, you can take advantage of all the performance, scale, reliability, and availability of AWS
infrastructure, as well as integrations with AWS networking and security services, such as application load balancers (ALBs)
for load distribution, AWS Identity and Access Management (IAM) integration with role-based access control (RBAC), and AWS
Virtual Private Cloud (VPC) support for pod networking.
6.5 ECS
6.6 AWS Network
6.7 AWS Databases
6.8 Load Balancer
6.9 IAM

7 Azure

You might also like