Job Description
About FWD Group
FWD Group (1828.HK) is a pan-Asian life and health insurance business that serves approximately 40 million customers across 10 markets, including BRI Life in Indonesia. FWD’s customer-led and tech-enabled approach aims to deliver innovative propositions, easy-to-understand products and a simpler insurance experience. Established in 2013, the company operates in some of the fastest-growing insurance markets in the world with a vision of changing the way people feel about insurance. FWD Group is listed on the main board of the Hong Kong Stock Exchange under the stock code 1828.
For more information, please visit www.fwd.com
Purpose
Serve as the hands-on technical and delivery contributor within the IT Resilience and Modernization function — building, automating, and hardening the platforms, tooling, and runbooks that keep New Business (NB) and Customer-Facing systems reliable.
Provide front-line engineering support during P1/P2 incidents — troubleshooting across infrastructure, cloud, application, and network layers, and driving fixes to resolution alongside SME virtual teams.
Execute the technical building blocks of modernization and resilience-by-design — observability instrumentation, automation of toil, DR/chaos test execution, and platform-standard adoption.
Key accountabilities
Hands-on engineering & delivery
Implement and maintain observability tooling — dashboards, alerts, telemetry, and correlation/AIOps rules — to improve MTTD and alert quality.
Build and automate runbooks and remediation scripts to reduce toil and enable self-healing.
Configure and tune SLI/SLO instrumentation and error-budget measurement in monitoring platforms.
Incident response & troubleshooting
Act as hands-on responder in P1/P2 war rooms — diagnosing across code, infra, cloud, and network; identifying where fixes are needed and applying them.
Execute the P1/P2 escalation protocol at working level and support parallel troubleshooting with GO and local sub-teams.
Contribute to blameless PIRs with technical root-cause detail and implement the resulting systemic fixes.
Modernization & resilience execution
Execute modernization tasks — refactoring, replatforming, and decommissioning steps — under the Director’s target-state architecture and guardrails.
Hands-on application modernization — refactor and replatform legacy application components, containerize/redeploy services, build and update APIs and integration/middleware, and execute application decommissioning and data-migration cutovers.
Run and support DR / chaos-engineering exercises and capture recoverability evidence (RTO/RPO).
Maintain the consolidated resilience knowledge base — architecture info, runbooks, and RCA records — for fast retrieval during incidents.
Modernization enablement
Wire CI/CD pipelines, automated tests, and observability/telemetry into modernized applications, and validate resilience patterns (health checks, retires, failover) post-migration
Qualifications / experience
Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
5+years hands-on in SRE, infrastructure, cloud, DevOps, or platform engineering, with direct incident-response experience.
Demonstrated experience building observability, automation, or CI/CD tooling in a production environment.
Practical exposure to modernization work — refactoring, replatforming, or legacy remediation — is an advantage.
Knowledge & technical skills
Hands-on with observability / AIOps stacks (e.g., Elastic, Datadog, Dynatrace, OpenTelemetry) — dashboards, alerting, correlation, anomaly detection.
Practical SRE skills — SLI/SLO instrumentation, error-budget measurement, and capacity/reliability tuning.
Strong scripting / automation ability (e.g., Python, Bash, PowerShell) and infrastructure-as-code (e.g., Terraform, Ansible).
Working knowledge of cloud platforms (AWS/Azure/GCP), Kubernetes / container orchestration, CI/CD pipelines, and API integration.
Solid troubleshooting across application, infrastructure, network, and cloud layers, including log/trace analysis.
Hands-on application-modernization skills — refactoring/replatforming legacy code, containerization (Docker/Kubernetes), API and integration/middleware development, and strangler-fig-style incremental migration.
Familiarity with ITIL 4 incident management, NIST 800-61 response, and BCP/DR practices.
Good written and verbal communication for technical documentation and war-room updates; fluent in English; Chinese written proficiency preferred.