JOB SUMMARY
Supports the reliability, resilience, performance, and continuous improvement of Vanguard Charitable's technology services. The role combines operational engineering, monitoring, automation, technical incident response, and service support responsibilities while helping the organization mature toward a more proactive and data-driven operating model. Working closely with service owners, engineering teams, vendors, and the IT Service Management & Support Lead, the analyst helps improve observability, reduce operational risk, strengthen disaster recovery readiness, support technical troubleshooting, and automate repetitive operational activities. The position also serves as a catalyst for operational innovation by applying Microsoft Copilot, Claude, automation technologies, and modern operational practices to improve productivity, reduce manual effort, strengthen knowledge sharing, and enhance business outcomes.
ESSENTIAL JOB FUNCTIONS
- Evaluate the health, availability, performance, and reliability of technology services; develop and improve monitoring, alerting, dashboards, observability practices, and operational reporting that improve visibility and early issue detection. (20%)
- Support incident triage, troubleshooting, technical investigation, and service restoration activities by analyzing alerts, logs, telemetry, configuration data, and system dependencies in partnership with engineering teams, vendors, and service owners. (15%)
- Develop, maintain, and improve operational runbooks, recovery procedures, synthetic testing, resilience validation activities, and technical readiness standards that improve supportability and operational consistency. (15%)
- Leverage Microsoft Copilot, Claude, automation tools, scripting, and workflow technologies to improve productivity, documentation, reporting, troubleshooting, knowledge management, and operational effectiveness; develop reusable prompts, templates, and automation solutions that reduce manual effort across Platform Operations. (20%)
- Support disaster recovery, resilience testing, recovery validation, capacity planning, and operational readiness efforts; identify reliability risks and recommend actions that strengthen recovery capabilities and service continuity. (10%)
- Provide secondary Help Desk and operational support coverage, support service-request fulfillment, assist with escalation triage, maintain support knowledge, and help ensure continuity during colleague absences, periods of elevated demand, and business-critical events. (15%)
- Maintain operational documentation, dependency records, support information, and configuration data while participating in special projects and continuous-improvement initiatives. (5%)
QUALIFICATIONS
Education & Training
- Bachelor's degree in Information Technology, Computer Science, Engineering, Business, or a related discipline required; equivalent experience may be considered.
- Training or coursework related to systems administration, cloud technologies, platform operations, automation, reliability engineering, AI, or IT operations preferred.
Licensure & Certification
- Microsoft, AWS, ITIL, observability/monitoring, automation, cybersecurity, AI, or related certifications preferred.
Knowledge, Skill, and Ability
- Strong understanding of technology operations, system reliability, monitoring, troubleshooting, and operational support.
- Working knowledge of observability, incident response, automation, disaster recovery, and operational resilience concepts.
- Experience using monitoring, logging, reporting, and service management platforms.
- Ability to analyze technical issues, identify patterns, and recommend practical improvements.
- Experience leveraging Microsoft Copilot, Claude, automation tools, scripting, or similar technologies to improve productivity and operational effectiveness.
- Strong written and verbal communication skills with the ability to explain technical concepts to technical and non-technical audiences.
- Curious, self-directed learner with a passion for continuous improvement, automation, and operational innovation.
- Strong organizational skills with the ability to manage multiple priorities and operate effectively during incidents and business-critical events.
Experience
- Four or more years of experience in technology operations, technical support, systems administration, platform operations, reliability engineering, or related disciplines.
- Experience supporting production technology services and participating in incident resolution activities.
- Experience developing operational documentation, runbooks, monitoring solutions, dashboards, or automation capabilities.
- Experience working with vendors, managed service providers, and cross-functional technology teams preferred.
- Experience applying AI capabilities, automation, or workflow improvements to real business or operational challenges preferred.
- Experience participating in disaster recovery, resilience testing, platform monitoring, or operational readiness activities preferred.