Associate Director, Software Engineering

Location: 

Guangzhou, GD, CN, 510620


Brand:  HSBC
Area of Interest:  Technology
Closing Date:  Hybrid Worker
Date:  1 Sept 2026

Job description

Some careers have more impact than others.

If you’re looking for a career where you can make a real impression, join HSBC and discover how valued you’ll be.

We are currently seeking an experienced professional to join our team in the role of Associate Director, Software Engineering.

 

Business: Fin, Supervision and TR Tech

Job ID: 58280

Principal responsibilities:

1. Service Reliability and Operations

•       Own the service-reliability strategy and operating model for the CCR SRE Pod, aligned with Traded Risk IT and business priorities.

•       Define, agree and monitor service-level indicators, service-level objectives and operational measures, including availability, alert quality, incident volume, mean time to detect, mean time to recover, change failure rate and recurring operational toil.

•       Ensure the team monitors production services and responds appropriately to alerts, xMatters calls, service requests and incident queries.

•       Establish effective T1 duty-engineer and T2 engineering/support responsibilities, with clear coverage, handover, backup and escalation arrangements.

•       Maintain operational readiness through current runbooks, knowledge articles, support contacts, dependency maps, recovery procedures and component-specific escalation routes.

•       Ensuring production risks, capacity constraints, technical debt, control weaknesses and single points of failure are identified, prioritized and remediated.

•       Lead readiness reviews for material releases and changes, ensuring monitoring, rollback, support and recovery arrangements are adequate before production implementation.

 

2. Incident and Problem Management

•       Provide leadership during major or complex production incidents, establishing clear technical coordination, decision-making, ownership and stakeholder communication.

•       Ensure incidents are correctly assessed and recorded in the relevant service-management systems, including the creation and maintenance of incident, SOTN, INS and related recovery records where required.

•       Coordinate recovery across CCR application pods, business teams, infrastructure and cloud services, vendors, and upstream/downstream technology teams.

•       Ensure unresolved component issues are routed promptly to the correct engineering pod or xMatters support group and escalated according to the agreed support model.

•       Own the quality and timeliness of post-incident reviews, root-cause analysis, corrective actions and lessons learned.

•       Identify incident patterns and ensure preventive actions are converted into prioritized, owned and tracked engineering work.

•       Maintain clear and timely communications to business and technology stakeholders throughout service disruption and recovery.

 

3. SRE Engineering, Automation and Continuous Improvement

•       Build an engineering-led SRE culture in which recurring operational activity is automated wherever practical and safe.

•       Drive improvements in monitoring, alerting, logging, tracing, dashboards, health checks, capacity management and automated recovery.

•       Reduce noisy or unactionable alerts and ensure alerts provide clear ownership, business context and response guidance.

•       Sponsor automation for routine support, deployment verification, service recovery, data-quality checks and operational reporting.

•       Promote resilient architecture, graceful degradation, failure isolation, disaster recovery, controlled change and secure-by-design engineering practices.

•       Partners with application engineering leads to improve non-functional requirements, production operability and reliability throughout the software-development lifecycle.

•       Use incident, change, availability and operational-toil data to prioritize continuous improvement and demonstrate measurable outcomes.

•       Encourage appropriate adoption of cloud, DevOps, AI-assisted engineering and other emerging technologies while maintaining security, data and model-risk controls.

 

4. Delivery and Process

•       Own and prioritize the SRE backlog, balancing production support, risk remediation, automation, technical debt and strategic delivery.

•       Ensure committed deliverables are achieved within schedule and budget, with resources, dependencies, risks and issues actively managed.

•       Establish quality controls and ensure engineering, testing, release, change and service-management processes are defined and followed.

•       Work with Technical Product Managers, service owners and application pod leads to align reliability investment with value-stream outcomes.

•       Review policies, practices and procedures continuously to improve delivery speed, service stability and control effectiveness.

 

5. Customer and Stakeholder Management

•       Act as the primary leadership contact for the CCR SRE service and build trusted relationships with business stakeholders, application pods and internal technology partners.

•       Represent service health, operational risk, reliability priorities and investment needs clearly to senior stakeholders.

•       Ensure stakeholders receive transparent reporting on incidents, service performance, recurring risks, remediation progress and reliability outcomes.

•       Maintain effective engagement with business, transformation, infrastructure, cloud, cybersecurity, service-management, vendor and upstream/downstream teams.

•       Resolve conflicting priorities through evidence-based decisions that balance customer impact, operational risk, regulatory obligations and delivery value.

 

6. People Leadership

•       Lead, inspire and develop a geographically distributed, high-performing and inclusive SRE team.

•       Set clear objectives and behavioral expectations and manage performance against measurable engineering and service outcomes.

•       Develop team capability across software engineering, production support, automation, observability, incident leadership and CCR domain knowledge.

•       Establish succession plans, role backups and cross-training to minimize key-person and location risks.

•       Plan shift and on-call coverage sustainably, protecting team wellbeing while meeting service obligations.

•       Lead recruitment, onboarding, career development and retention for permanent and contingent colleagues.

•       Foster a blameless learning culture that encourages ownership, disciplined execution, constructive challenge and continuous improvement.

 

7. Financial Management

•       Manage the Pod within the agreed budget, including forecasting, monitoring, reconciliation and transparent reporting.

•       Produce an accurate resource and financial forecast aligned with service coverage and the SRE delivery roadmap.

•       Identify opportunities to reduce support and delivery costs through automation, simplification, platform reuse and improved estimation.

•       Work with the Technical Product Manager and service leadership to plan sustainable growth and capability investment for the Pod.

 

8. Risk, Compliance and Controls

•       Identify, assess, record and manage technology, operational-resilience, cybersecurity, data and third-party risks affecting CCR services.

•       Ensure decisions demonstrate appropriate commercial judgement, customer focus and risk awareness.

•       Ensure compliance with applicable HSBC policies, controls, change governance, incident-management processes and regulatory expectations.

•       Maintain sufficient evidence for controls, audit, incident review, operational acceptance and risk remediation.

•       Escalate material service risks, control weaknesses and overdue remediation clearly and promptly.

 

Knowledge & Experience/Qualifications:

Essential Experience

•       Degree or equivalent professional experience in Information Technology, Computer Science, Engineering or a related discipline.

•       Significant experience in software engineering, production engineering, SRE, DevOps or application support for business-critical distributed systems.

•       Proven experience leading a medium-to-large technical team across multiple locations or time zones.

•       Demonstrable leadership of major incidents, service recovery, stakeholder communications, root-cause analysis and preventive remediation.

•       Experience establishing or improving observability, alerting, service-level objectives, operational metrics, runbooks and on-call practices.

•       Strong knowledge of the complete software-development lifecycle and the production operation of complex applications.

•       Experience applying Agile, DevOps, CI/CD, infrastructure automation and continuous-improvement practices.

•       Experience managing operational risk, production change, service continuity and control obligations in a regulated environment.

•       Proven ability to manage multiple senior stakeholders with competing priorities and to make evidence-based decisions under pressure.

•       Experience in workforce planning, performance management, talent development and succession planning.

•       Experience in financial planning, budgeting, forecasting and vendor-outcome management.

•       Strong architecture and engineering judgement, with a focus on resilience, scalability, security, maintainability and operability.

•       Excellent written and verbal communication skills, including the ability to explain technical incidents and risks to non-technical stakeholders.

•       Fluency in English.

 

Essential Technical Skills

•       Strong understanding of Java and/or Python-based distributed applications, APIs and microservices.

•       Practical knowledge of observability and application-performance monitoring, including metrics, logs, traces, dashboards and actionable alert design.

•       Working knowledge of relational and distributed databases such as Oracle, ClickHouse, MongoDB, Hadoop or HBase.

•       Experience with test automation, CI/CD pipelines, source control and automated deployment practices.

•       Working knowledge of at least one major cloud platform, preferably GCP, AWS or Azure.

•       Understanding of containers and orchestration technologies such as Docker and Kubernetes.

•       Ability to analyze production behavior, dependencies, capacity and failure modes across complex systems.

 

Desired Experience

•       Knowledge of counterparty credit risk, traded risk, market risk or related banking and financial services domains.

•       Experience supporting batch and data-processing platforms with complex upstream and downstream dependencies.

•       Familiarity with ITIL-aligned incident, problem, change and service-request management, including ServiceNow and xMatters or equivalent tooling.

•       Experience with public-cloud platforms, Kubernetes-based services, distributed data platforms and hybrid infrastructure.

•       Experience designing disaster-recovery, capacity-management, chaos-testing or resilience-testing programmes.

•       Knowledge of HSBC policies, processes and technology-control frameworks.

 

What additional skills will be good to have?

•       SRE, DevOps, cloud architecture, Kubernetes, ITIL, Scrum, PMP or equivalent certification.

•       FRM or CFA qualification.

•       Experience with infrastructure as code, policy as code and automated operational controls.

•       Experience applying AI and machine learning to engineering productivity, observability or incident management with appropriate governance.

Job Board Tags:

/WX/51/LP

 

You’ll achieve more when you join HSBC.

 

HSBC is an equal opportunity employer committed to building a culture where all employees are valued, respected and opinions count. We take pride in providing a workplace that fosters continuous professional development, flexible working and opportunities to grow within an inclusive and diverse environment. We encourage applications from all suitably qualified persons irrespective of, but not limited to, their gender or genetic information, sexual orientation, ethnicity, religion, social status, medical care leave requirements, political affiliation, people with disabilities, color, national origin, veteran status, etc., We consider all applications based on merit and suitability to the role.

 

Personal data held by the Bank relating to employment applications will be used in accordance with our Privacy Statement, which is available on our website.

 

***Issued By HSBC Software Development (GuangDong) Limited***