
Deepanshu Pandey
5+ years of experience in Production Support and Site Reliability Engineering, specializing in end-to- end support for distributed systems and derivi…
5+ years of experience in Production Support and Site Reliability Engineering, specializing in end-to- end support for distributed systems and deriving immediate resolutions for incidents within Insurance and Annuities applications. Proven expertise in supporting batch operations, proactive incident management, and performance enhancement recommendations, ensuring minimal business impact and quick restoration of service degradation. Proficient in leveraging tools like Puppet, GitLab, Axway, Splunk, Dynatrace, ServiceNow, and Atlassian Tools (Jira, Confluence, Bitbucket) to maintain high availability and optimize application performance. Deep understanding of Agile methodologies and ITIL framework, with experience in Incident, Problem, and Change Management.
Production SupportIncident ManagementProblem ManagementBatch Operations SupportTWS/Tidal SchedulerServiceNowDynatenceSumologicSQL ServerOperating Systems: Windows/Linux/UnixJAVA/.NET/XMLMiddlewareAPIsAWS servicesETL frameworksVersion Control: GitCloud Platforms: AzureCI/CD & Automation: JenkinsPuppetAnsibleShell ScriptingMonitoring & Observability: DynatraceSplunkXymonPrometheusGrafanaSolarWindsProject Management: JiraEnterprise Integration: AxwayITIL Framework
Site Reliability Engg. (Production Support)
NetSpend
Jan 2024 – Present
- Orchestrated many rapid release deployments annually across PROD and CERT environments, achieving 100% success rate and maintaining zero production downtime.
- Managed a Puppet-based configuration management system for a fleet of 500 Linux nodes, ensuring 100% configuration consistency across Dev, QA, and Production.
- Implemented GitOps workflow for puppet using GitLab, enabling automated code deployments and peer-reviewed infrastructure changes.
- Troubleshoot issues related to puppet.
- Deployed Multiple microservices and F5 using puppet.
- Application Support: Ensuring that deployed applications function correctly and efficiently, providing end- to-end production support for distributed systems.
- Incident & Problem Management: Conducting analysis, investigation, diagnosis, and problem-solving to identify, troubleshoot, and resolve production issues, driving immediate resolutions and root-cause analysis.
- Service Level Management: Assisting Service Managers in delivering services while meeting service level agreements (SLAs).
- Performed hotfix deployments to resolve critical production issues.
- Ensures that issues are fully documented within the relevant reporting systems.
- Managed and updated property and configuration files for production and certification environments.
- Handled secure property configurations using Vault.
- Conducted application restarts and coordinated server reboots for scheduled patching activities.
- Proactively monitored applications and infrastructure using tools such as Dynatrace, Splunk, Xymon, and other relevant production support channels.
- Configured F5 load balancers, vanity URL redirects, proxy passes, and Apache setups.
- Managed job lifecycle operations (promotion, pausing, restarts) within Tidal Scheduler to support batch operations.
- Upgraded tools Tidal and Axway on Linux machines.
- Created new Axway Setups for different clients and resolved issues in case of transmission failures or partner requests.
- Performed disk space cleanup and provisioned additional storage as required.
- Set up new applications, including configurations for F5, port management, property and conf files, and monitoring dashboards.
- Acted as a liaison with external partners (e.g., PFG, associations, banks) to resolve issues and coordinate maintenance activities.
- Provided on-call support and worked during shifts accommodating US workdays, including weekends/holidays.
Atlassian DevOps
Publicis Resources
May 2021 – Jan 2024
- Experience with CI/CD tools like Jenkins.
- Experience with configuration management tools with Ansible.
- Resolved critical application outages through in-depth root cause analysis and strategic bug fixes, reducing system downtime by 25% and ensuring continuous service for 10,000+ users.
- Provided patching support for applications of Linux.
- Installation, setup, maintenance, Upgrade, backup & recovery on Linux server applications.
- Upgraded applications to latest versions.
- Migration support for applications from one server to another.
- Monitored application / server using Prometheus/Grafana and SolarWinds.
- Hosted applications environments includes Nginx.
- Basic understanding of Database concepts & SQL queries.
- Incident Management support using ServiceNow and JSM, aligning with ITIL framework practices.
- Changed DNS of Applications.
- Resolved LDAP connectivity issues in applications.
- Sonar administrator.
- Debugged Sonar issues and encryption in Sonar.
- Set up complete application from scratch on server.
Aligarh Muslim University
MBA
2019
Aligarh Muslim University
Bachelor’s in science
2016
Microsoft
Microsoft Certified Azure Fundamentals-Az 900, Azure Fundamentals