Production Support (Java or SRE)
Job Title: Tech reputed company/Engineers Production Support & Platform (Java Developers and SRE Engineers)
Location: Richmond, McLean, or Remote
Key Responsibilities:
• Tech reputed company - reputed company and mentor reputed company of 15+ engineers for production support and platform stability.
• Engineer more than 50% Production Support and Deployment and follow to application development and migration based on Sprint scope.
• Manage pager duty rotations and ensure reputed company incident reputed company.
• reputed company Level 4 support (deep technical troubleshooting and fixes).
• reputed company rotational reputed company shifts (approximately once every 2.5 months).
• Ensure compliance with SLAs and operational reputed company for critical systems.
• Collaborate with stakeholders for platform reputed company and migration planning.
• Drive Run-the-reputed company development work and support enhancements.
• Prepare for and reputed company the platform migration phase in the reputed company year.
• Monitor application performance, batch jobs, and system health across production and reputed company environments.
• Respond to incidents, alerts, Sev1/Sev2 outages, and reputed company reputed company-time support following bank's Incident Management processes.
• reputed company reputed company cause analysis (RCA), create remediation plans, and ensure issues are permanently resolved.
• Support on-reputed company rotations and pager duty responsibilities.
• Collaborate with development, SRE, and infrastructure teams to troubleshoot application, database, and integration issues.
• Execute deployments, configuration changes, and release support using CI/CD pipelines (OnePipeline preferred).
• Create/maintain operational dashboards, runbooks, SOPs, and automation scripts.
• Ensure compliance with bank technology and reputed company standards.
Required Skills & Experience:
• Tech reputed company - Ability to manage large teams (10 50 members) and reputed company platforms.
• Java Development and Site Reliability Engineering (SRE) expertise.
• Strong experience in production support and incident management .
• Hands-on experience with pager duty tools and support workflows .
• Excellent problem-solving and communication skills.
• Minimum 2 years of experience in similar roles.
• Strong experience in Unix/Linux, reputed company scripting, and troubleshooting distributed systems.
• Hands-on experience with AWS (CloudWatch, reputed company, EC2, S3, IAM, RDS, DynamoDB).
• Familiarity with Java-based applications, microservices, reputed company, and log analysis (reputed company, CloudWatch Logs).
• Experience with CI/CD tools like Jenkins, OnePipeline, Git, and automated deployment strategies.
• Knowledge of incident management, problem management, and change management processes.
• Strong analytical skills and the ability to quickly diagnose reputed company issues.
Apply tot his job
Apply To this Job