Head of SSv Reliability Engineering

Location: 

Guangzhou, GD, CN, 510620


Brand:  HSBC
Area of Interest: 
Closing Date:  Hybrid Worker
Date:  26 Aug 2026

Job description

Some careers have more impact than others.
If you’re looking for a career where you can make a real impression, join HSBC and discover how valued you’ll be.

We are currently seeking an experienced professional to join our team in the role of POSITION TITLE.

Business:CIB Technology

Job ID:57523

About HSBC Securities Services Technology

Securities Services (SSv) is a profitable, scalable, capital-light business at the heart of HSBC’s transaction banking, institutional client and wealth ambitions, enabling clients to safely hold, move and administer investments across global markets. It supports critical post-trade and servicing capabilities across custody and safekeeping, settlement support, corporate actions, and high-quality data and reporting, helping clients operate efficiently and meet regulatory obligations. SSv also provides services across Fund Administration, Middle Office and Transfer Agency, as well as supporting evolving client needs in Private Assets. It plays a critical role for HSBC by reinforcing institutional connectivity, supporting sustainable fee income, and capitalising on the Group’s global reach and operational strength.

SSv Technology is a 1,200-person global engineering organisation partnering with SSv business and engineering colleagues across HSBC to deliver resilient, secure and industry-leading technology services that underpin these client outcomes.

About the Job

Job Description

Assume a founding leadership role in shaping the reliability strategy for HSBC SSv Technology, driving lasting improvements to the resilience of business‑critical platforms across a global organisation spanning London, India, Hong Kong and Guangzhou.

As Head of Site Reliability Engineering (SRE), you will stand up a new SRE pillar within SSv Technology from the ground up. This is a build mandate, not a maintain mandate: you will design the operating model, set the roadmap, establish the culture, and build the team — while remaining close enough to the work to lead major incidents and set the technical bar yourself.

This is the first SRE leadership role of its kind in SSv Technology. Success depends on a strong Reliability Engineering mindset: designing for failure, prioritising automation, and building self-healing systems. You will think proactively about reliability risks and reliability-by-design, not just react to incidents.

You will report into the CIO of SSv Technology and act as the primary architect and evangelist for reliability engineering across custody, fund services, and post-trade platforms.

Job Responsibilities

Organisational Design, Implementation Strategy & Governance

  • Lead a global Reliability Engineering initiative for a 1,200-person engineering organisation, setting the strategy, outcomes and adoption plan to raise reliability maturity across teams and regions.
  • Design and stand up the SRE operating model for SSv Technology — evaluate central, embedded and hybrid structures, and implement the model best suited to custody, fund accounting, transfer agency and post-trade platforms across regions.
  • Establish the SRE function’s charter, mandate and governance, and secure alignment from SSv Technology and CIB Technology leadership.
  • Set the multi-year roadmap and priorities for SRE maturity — observability, incident management, automation, resiliency engineering, and toil reduction — with clear measures of progress (e.g., SLO adoption, incident trends, toil reduction, and reliability improvements).
  • Partner with regional technology leads, principal engineers and application engineering leads to align global practice while adapting to local regulatory and operational realities in China and across the globe.

Reliability Engineering Team Building Strategy

  • Hire and lead a small central SRE team of 6–7 people to provide enablement, patterns, tooling direction, coaching and incident leadership across SSv Technology.
  • Build an extended SRE community by identifying reliability-focused engineers across the wider engineering population and mobilising them as an embedded/virtual team within application squads.
  • Define how SRE engages with teams (e.g., reliability clinics, embedded rotations, design reviews, readiness assessments) to get the practice off the ground quickly and sustainably.
  • Implement an operating model where reliability practices are progressively internalised by application teams, reducing dependency on the central team over time.

Culture, Education & Enablement at Scale

  • Establish SRE culture and practice from zero: error budgets, blameless postmortems, toil reduction, reliability-as-code, and a shared operational language across application teams.
  • Create an education and enablement programme to upskill ~1,200 engineers on Reliability Engineering concepts and processes (e.g., SLI/SLO thinking, error budgets, incident learning, automation/toil reduction), supported by playbooks, training materials, communities of practice and targeted coaching.
  • Embed Reliability Engineering into day-to-day delivery by introducing pragmatic design reviews and reliability guardrails for critical changes, ensuring teams build resilient services without creating unnecessary bureaucracy.
  • Define career paths, skills frameworks and training for SRE talent, and act as the internal sponsor for the discipline.
  • Hire, grow and mentor a reliability engineering team based in China; build a durable talent pipeline in a competitive local market.

Hands-On Reliability Engineering

  • Personally lead major incident responses, root-cause analysis and remediation for SSv Technology platforms, coordinating across teams and regions to restore service rapidly and drive lasting fixes.
  • Drive self-healing-by-design: automated detection, safe auto-remediation patterns, resilience patterns and engineering guardrails that prevent repeat incidents and reduce manual intervention (toil).
  • Own stability, availability and end-to-end operational performance of business-critical SSv Technology platforms, with clear accountability for outcomes.
  • Drive adoption of observability and monitoring tooling (e.g., Dynatrace, Splunk, Grafana), automation and scripting (Python, Shell, Ansible, Terraform), and modern distributed platforms (containers, microservices, Kubernetes/OpenShift).
  • Champion enterprise-authorised AI capabilities to accelerate SRE workflows — incident triage, investigation support and knowledge capture — with rigorous validation habits and data-sensitivity awareness.
  • Establish capacity and performance engineering practices (capacity planning, load testing, latency/throughput SLIs, saturation monitoring) for critical services.
  • Own and mature resilience validation (DR/BCP planning and regular failover testing; resilience game days where appropriate).

Stakeholder Partnership

  • Serve as a trusted advisor to SSv business and technology stakeholders, providing concise, business-focused communication during incidents and strategic initiatives.
  • Partner closely with Production Services, Application Development, Infrastructure, Non-Functional Risk, and Compliance/Regulatory functions to improve platform stability and operational maturity.
  • Represent the SRE function in governance forums, service reviews, and China-specific regulatory or technology-risk discussions.

Required Qualifications, Capabilities, and Skills

  • Formal training or certification in SRE concepts, extensive applied experience, including prior experience building or scaling a reliability engineering/SRE function from the ground up, ideally in a regulated financial-services environment.
  • Demonstrated Reliability Engineering mindset with evidence of designing self-healing systems and driving automation-first reliability improvements proactively (not only via process).
  • Proven experience leading new initiatives from scratch at scale (strategy, operating model, mobilisation, adoption and delivery across large engineering populations).
  • Demonstrated experience designing organisational models (central vs embedded vs hybrid) for reliability, platform, or infrastructure engineering functions.
  • Proven track record hiring, growing and retaining technical teams, ideally within the China/APAC technology talent market.
  • Strong systems thinking and problem-solving capability to assess complex, cross-domain, cross-regional production issues and drive end-to-end resolution.
  • Demonstrated major incident leadership, coordinating effectively across global teams to restore service rapidly and drive root-cause remediation.
  • Hands-on technology experience across cloud (AWS/Azure/GCP), automation and scripting (Python, Shell, PowerShell, Ansible, Terraform), and modern distributed platforms (containers, microservices, Kubernetes/OpenShift).
  • Strong observability and service management expertise (Dynatrace, Splunk, Grafana; ITIL Incident/Problem/Change/Availability), with excellent verbal and written communication in English
  • Working knowledge of custody, fund services, and/or post-trade workflows, or strong transferable capital markets technology experience.
  • ITIL v4 certification (or equivalent service management experience) preferred.

Preferred Qualifications, Capabilities, and Skills

  • Experience integrating support models and standardising operating practice across multiple regions or legal entities.
  • Background in organisational or cultural change leadership, particularly introducing a new engineering discipline into an established technology organisation.
  • Knowledge of Securities Services domains: custody, fund accounting, transfer agency, middle office, or post-trade processing would be an advantage.
  • Experience partnering with Front Office or client-facing teams to support business-critical, high-availability platforms.