Bachelor’s degree in engineering or equivalent practical experience supporting production equipment and systems.
3 to 5 years of experience in manufacturing, equipment troubleshooting, internal production support,or closely related technical support environments.
Hands-on experience using production monitoring and incident troubleshooting tools such asITRS, Splunk, and Datadog.
Demonstrated ability to perform incident management, advanced troubleshooting, and service restoration for internal production systems and equipment.
Proficiency in root cause analysis (RCA) and implementing permanent corrective actions to prevent recurring production incidents.
Experience scripting and automation usingShell and Pythonto reduce manual work and repetitive tasks.
Strong ownership during incidents, adaptability, and ability to remain calm under pressure.
Responsibilities:
Lead incident management for internal production systems and equipment, including monitoring, troubleshooting, and restoring service to minimize downtime.
Perform tool bring-up and execute qualification activities, including qualification levels 600, 700, and 800.
Conduct root cause analysis for recurring production issues and implement permanent fixes and preventive actions.
Automate repetitive operational tasks through scripting and maintain procedures and documentation to support consistent execution.
Support system maintenance and change activities including batch processes, upgrades, release deployments, and user acceptance testing (UAT), coordinating across development, operations, and business teams.