Job Description
Role: Senior Production Support Engineer
Location: Austin, TX (Onsite)
Exp. Level:8+ yrs
Role PurposeOwns AI-augmented incident triage, runbook execution, and root-cause analysis, and separately owns AI-powered anomaly detection and spend/usage alerting.
Key Responsibilities:
•Own AI-augmented incident triage, runbook execution, and root-cause analysis across the 3 core FinOps applications and 50+ services.
•Build and maintain AI-powered anomaly detection and spend/usage alerting.
•Support the 99.99% uptime target during core PST business hours (8AM-5PM)
•Operate within the support portion of delivery (L1/L2/L3 tiers) within the 35% run-and-support split.
•Contribute to baselining incident response/resolution times post-transition.
Must-Have:
•8–13 years in production support or SRE, including incident management.
•Experience with monitoring/alerting tooling, ideally GCP-native.
•Demonstrated root-cause analysis discipline at production scale.
Nice-to-Have:
•Experience with AI-based anomaly detection tools.
•Node.js or Python for support automation/scripting.