JPMorganChase· Chapter 02

L2/L3 Production Support Engineer

Production incident management, application troubleshooting and database change engineering across a 24×7 financial application environment.

January 2014 - December 2014P1 / S1 incidentsITSM-controlled changeIndia + US operations

From implementation to production ownership

The question changed from “How do we implement this client?” to “Why is the system behaving this way?”

After working on implementations, my responsibility moved closer to the live production environment. I supported production issues for multiple clients across application, database and integration layers in a 24×7 operating model involving teams in India and the US.

The visible error was the symptom. The work was finding the underlying cause.

L2/L3 application support

Tracing reported symptoms through the application stack.

  • Application errors and incorrect user behaviour or configuration.
  • Data inconsistencies and database issues.
  • Integration failures and application defects.
  • Performance-related problems.
  • Production deployment issues.
  • Fix validation and confirmation of production behaviour after implementation.

Root-cause investigation

A layered troubleshooting process kept the fix matched to the actual failure.

UserApplicationDatabaseIntegrationDeployment

I analysed application behaviour, Java code, database records, PL/SQL objects, logs and error messages, recent deployments, data changes, integration behaviour and configuration. The goal was to distinguish between a data problem, application defect, database defect, configuration problem, deployment problem or integration problem before choosing the fix.

Application fixes and ITSM-controlled database change

Production access became more controlled as the support responsibility matured.

Application fixes

When the root cause was in the application, I identified the affected component, understood the defect, developed or supported the required Java change, validated the fix, coordinated deployment and confirmed the production result.

Database fixes

With read-only access to production databases, database writes were executed through formal ITSM-controlled implementation processes. I prepared plans for data corrections, database object changes, PL/SQL updates, production fixes and emergency changes.

What every change plan covered

Change rationaleAffected objects and dataImplementation stepsValidation stepsRollback considerationsExpected production impact

P1 / S1 incident management

A severe incident needed a durable answer, not only a quick workaround.

DetectTriageStabilizeCommunicateResolvePrevent recurrence

For P1/S1 incidents, I worked with stakeholders and technical teams to establish priority, identify business impact, coordinate investigation, provide updates, establish a workaround where necessary, implement the permanent fix, document root cause and define additional alerts or preventive controls.

24×7 operations and team coordination

Operational ownership extended beyond the individual incident.

  • Participated in 24×7 support coverage and weekend rotations.
  • Handled production incidents, priority management and emergency changes.
  • Managed handover between geographic teams and stakeholder communication.
  • Prioritised daily production work and assigned tasks based on urgency and capability.
  • Tracked open incidents and fixes.
  • Prepared daily progress updates and weekly status reports.
  • Managed support coverage and escalation coordination.

What production support added

Engineering changed when real users depended on the system.

Production support made me more conscious of production risk, data integrity, backward compatibility, change impact, rollback planning, observability, incident prevention and operational ownership.

That experience later influenced how I designed automation: not simply to test functionality, but to identify risk before it became a production incident.

Next chapter
Automation CoE Lead