Job Description
Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production environment. This role focuses on troubleshooting complex trading infrastructure, managing observability, and collaborating across trading and technology teams to ensure system reliability and performance.
Role Overview
We are seeking a highly motivated Site Reliability Engineer (SRE) to join the Equity Trading Platform Engineering team, supporting critical trading applications and Fidessa flows used by internal business users and external clients. This role is responsible for ensuring the stability, resiliency, performance, and availability of front-office equity trading systems across normal trading hours, late evening processing, and weekend support rotations.
The ideal candidate will have strong technical troubleshooting skills, hands-on experience with trading platforms, FIX connectivity, Linux/Unix and Windows environments, monitoring and observability tools, and automation. The role requires close collaboration with business users, application development, infrastructure, QA, middle office, and global support teams to resolve issues quickly, improve platform reliability, and support ongoing business growth.
This position requires a strong team player who can work independently when needed, communicate clearly under pressure, and contribute to a collaborative, high-energy production support environment.
Key Responsibilities
- Site Reliability Engineer (SRE) who works on production support for Equity trading systems, including Fidessa flows used by internal users and external clients.
- Develop and maintain automation, scripts (Python, Perl etc), and operational tools to reduce manual effort, improve reliability, and accelerate issue resolution.
- Support day-to-day trading activity, including order routing, FIX connectivity, client onboarding, drop-copy flows, DMA, algorithmic trading, and high-touch / low-touch workflows.
- Investigate, troubleshoot, and resolve complex application, infrastructure, connectivity, and trading workflow issues.
- Build, enhance, and maintain monitoring and observability solutions for application, infrastructure, database, messaging, and trading flow components.
- Gather and analyze metrics from operating systems, applications, FIX flows, databases, and infrastructure to support performance tuning, capacity planning, and fault isolation.
- Lead or participate in incident response, root-cause analysis, post-mortems, and permanent remediation follow-up.
- Partner with business, development, QA, infrastructure, and global support teams to improve platform stability through disciplined testing, release procedures, and production readiness checks.
- Support CI/CD pipelines and deployment automation to streamline release management and production checkouts.
- Collaborate with application teams to define, document, and enforce production support standards, operational procedures, and development best practices.
- Maintain clear documentation, including runbooks, troubleshooting guides, FAQs, production checkout procedures, and support handover notes.
- Handle complex operational tasks and recommend improvements to processes, tooling, monitoring, and platform architecture.
- Support compliance, legal, audit, and regulatory queries related to trading activity, client flows, logs, and production events.
- Participate in global support coverage, including late evening support, weekend availability, production checkouts, and on-call rotations to ensure system availability.
· Support evening / weekend production checkouts (On-call rotations) and releases
Required Experience and Skills
- Strong experience supporting front-office trading systems in a production environment, preferably within Equities or Capital Markets.
- Hands-on experience with Fidessa, OMS/EMS platforms, FIX flows, order routing, client connectivity, and post-trade processing.
- Strong understanding of the trade lifecycle, including order capture, routing, execution, allocation, confirmation, clearing, settlement, and downstream reporting.
- Strong knowledge of FIX Protocol, including FIX 4.2, 4.4, and 5.0, message analysis, session behavior, sequence resets, rejects, cancels, replaces, drop copy, and client onboarding flows.
- Experience supporting asset classes such as Cash Equities, Futures, Options, ETFs, Swaps, Bonds, and Derivatives.
- Strong troubleshooting experience across Linux/Unix, Windows Server, networking, databases, messaging, and application layers.
- Strong scripting and automation skills using Python, Bash/Shell, SQL, PowerShell, or Perl.
- Experience with monitoring and observability tools such as Splunk, Grafana, Prometheus, Nagios, Datadog, OpenTelemetry, CloudWatch, or ELK Stack.
- Experience with databases such as MS SQL Server, MySQL, or PostgreSQL.
- Working knowledge of messaging and integration platforms such as IBM MQ, Tibco EMS, Kafka, ActiveMQ, AWS SQS/SNS, or equivalent technologies.
- Good understanding of network concepts including TCP, UDP, multicast/unicast, latency, connectivity troubleshooting, firewall flows, and packet-level diagnostics.
- Experience with CI/CD, DevOps, and release tools such as Jenkins, Bamboo, Bitbucket, Git, GitHub Actions, JFrog, Ansible, Terraform, ArgoCD, ServiceNow, JIRA, and Confluence.
- Familiarity with cloud and container platforms such as AWS, Azure, Docker, Kubernetes, OpenShift, Terraform, Helm Charts, and Lens.
- Ability to analyze complex production issues quickly while carefully assessing business and operational impact.
- Excellent communication skills with the ability to interact directly with traders, business users, technology teams, infrastructure teams, and senior stakeholders.
- Strong time management, ownership, prioritization, and follow-through skills.
- Ability to work effectively both independently and as part of a global team.