Operations Incident Management

200000

On-siteOperationsDigital InfrastructureUSPermanentDallas, TX, United States

About the Role

Our client is seeking a Director of Operations Resilience in Dallas to lead the strategy, development, and continuous improvement of our incident management program. This role is responsible for creating a proactive incident management framework that improves operational resiliency, strengthens response processes, and drives continuous improvement across our 24x7x365 operations.

The ideal candidate is a strategic leader who can build and mature an incident management program — not just manage individual incidents. This person will establish best practices, develop response strategies, improve workflows, and partner across Operations, IT, Engineering, customers, and strategic partners to ensure effective incident prevention, response, and recovery.

This role requires strong leadership, technical understanding, and the ability to think beyond immediate issues to identify trends, reduce risk, and create long-term operational improvements.

Responsibilities

  • Develop and lead Compass’ incident management strategy, including governance, processes, escalation procedures, and continuous improvement initiatives.
  • Own the full incident lifecycle from detection and response through resolution, root cause analysis, and corrective actions.
  • Drive operational maturity by creating playbooks, improving workflows, establishing KPIs, and implementing best practices.
  • Partner with Operations, IT, and technology teams to leverage automation and AI to improve efficiency, reduce errors, and enhance incident response capabilities.
  • Provide leadership during critical incidents and ensure effective communication across internal teams, customers, and strategic partners.
  • Analyze incident trends and develop proactive strategies to improve reliability and reduce recurring issues.
  • Support audits, reporting requirements, and ongoing improvements to the Incident Management Program.

Skills and Experience

  • Senior-level experience leading incident management, operations, service reliability, or command center programs in a mission-critical environment.
  • Proven ability to create and implement incident management strategies, processes, and operating models.
  • Strong leadership skills with experience influencing cross-functional teams and driving organizational change.
  • Experience with ServiceNow or similar ITSM platforms, along with knowledge of AI and automation solutions.
  • Technical understanding of data center infrastructure, including HVAC, electrical, life safety, security, and building management systems.
  • Ability to create technical documentation, analyze operational data, and communicate effectively with both technical teams and leadership.

Location

Dallas, TX - in office

Need to know more about this role?

NEED MORE HELP? REACH OUT TO A MEMBER OF OUR TEAM.

We're on hand to take your call if you have any questions regarding available positions, or steps involved in the process. We will connect you with a dedicated member of our recruitment team, to help guide you towards your ideal career.