Abhay Choudhary
Gurugram, Haryana, India
+91-8628831793 # abhaychoudharyqrt@[Link] ï LinkedIn § GitHub
PROFESSIONAL SUMMARY
Results-driven DevOps and Cloud Engineer with 2+ years of experience managing cloud infrastructure across AWS and Azure.
Specialized in container orchestration, observability, and proactive incident management. Reduced production disruptions by
35% through strategic monitoring and alerting implementation. Proven track record of enhancing system reliability,
accelerating incident resolution (MTTD/MTTR), and delivering automation solutions that save 200+ hours annually. Strong
expertise in Kubernetes, Docker, Grafana, and Infrastructure as Code practices.
TECHNICAL SKILLS
Cloud Platforms: AWS (EC2, S3, VPC, IAM, CloudWatch, EKS), Microsoft Azure (VMs, Blob Storage, Monitor, AKS)
Container & Orchestration: Docker, Kubernetes (K8s), Amazon EKS, Azure AKS, Helm, Container Registry
Observability & Monitoring: Grafana, Prometheus, Azure Monitor, CloudWatch, ELK Stack
Infrastructure as Code: Terraform, CloudFormation, ARM Templates
CI/CD Tools: Jenkins, GitLab CI/CD, GitHub Actions, Azure DevOps
Operating Systems: Linux (RHEL, Ubuntu, CentOS), Windows Server
Scripting & Automation: Python, Bash, YAML
Version Control: Git, GitHub, GitLab, Bitbucket
Configuration Management: Ansible
DevOps Practices: Incident Management, Root Cause Analysis, SRE Principles, Agile/Scrum
PROFESSIONAL EXPERIENCE
Solution Integrator July 2023 – Present
Ericsson Gurugram, India
– Cloud Infrastructure Management: Deployed multi-tier cloud infrastructure on AWS and Azure supporting 50+
microservices, achieving 99.8% uptime SLA for business-critical applications serving 30K+ daily active users.
– Container Orchestration & Managed Kubernetes: Managed production-grade Kubernetes clusters on Amazon
EKS and Azure AKS, orchestrating 15+ containerized applications with auto-scaling, self-healing capabilities, and
multi-AZ deployment for high availability.
– Observability & Monitoring: Engineered comprehensive monitoring solution using Grafana and Prometheus with
75+ custom dashboards and 120+ metrics, enabling real-time visibility into application performance and infrastructure
health across EKS/AKS clusters.
– Proactive Incident Prevention: Conducted advanced log analysis using Grafana Loki to identify critical integration
failures between microservices, preventing 15+ major production outages and saving an potential revenue loss.
– Production Disruption Reduction: Achieved 35% reduction in unplanned production disruptions through proactive
monitoring, alert tuning to identify weaknesses before customer impact.
– Linux System Administration: Managed 40+ Linux servers (RHEL/Ubuntu) performing regular tasks including user
access management, disk space monitoring, log rotation configuration, service restarts, and package updates using
yum/apt, ensuring system stability and 99.5% uptime.
– Automation & Efficiency: Developed Python automation scripts for automate operational tasks and process
Excel-based reports, reducing manual effort and improving accuracy. approximately 200+ hours annually (5 hours/week).
– Cross-functional Collaboration: Partnered with development, QA, and security teams to establish CI/CD best
practices, conduct post-incident reviews, and implement preventive controls, reducing recurring incidents by 65%.
– Root Cause Analysis: Led RCA investigations for 40+ critical production incidents, identifying systemic issues and
implementing permanent fixes including code refactoring, infrastructure improvements, and process enhancements.
– Infrastructure as Code: Transitioned 70% of manual infrastructure provisioning to Terraform-based IaC, ensuring
consistency, version control, and reducing provisioning time from hours to minutes.
– Cost Optimization: Identified and implemented cloud cost optimization strategies including rightsizing, reserved
instances, and automated resource cleanup, achieving 25% reduction in monthly cloud spending.
EDUCATION
National Institute of Technology Hamirpur 2019 – 2023
Bachelor of Technology in Electrical Engineering (CGPA: 8.48/10.0) Hamirpur, Himachal Pradesh
CERTIFICATIONS & TRAINING
AZ-900 Azure Fundamentals — Docker and Kubernetes — Google Cloud Digital Leader — Python - Udemy
— Linux — React js - Udemy — Nodejs - Udemy —
KEY ACHIEVEMENTS
• Prevented Production Outage Through Log Analysis: Identified integration failure between payment and
notification services by analyzing error patterns in Grafana logs, implemented fix before peak traffic hours, preventing
potential disruption affecting 20K+ users.
• Selenium Test Automation Framework: Developed Python-based Selenium automation framework for end-to-end
regression testing of critical user workflows, reducing manual testing effort by 4 hours per release cycle and improving
deployment confidence with 95% test coverage.
• WebDriverIO Automation Implementation: Built comprehensive WebDriverIO test automation framework for UI
validation across multiple browsers, eliminating 3-4 hours of manual testing per release and enabling parallel test
execution for faster feedback loops.
• Operational Automation Scripts: Created Python automation scripts saving approximately 5 hours weekly (260
hours annually) of manual operational work.