Atlas Systems

Senior MySQL DBA (Replication Specialist)-BK

Atlas Systems  •  Bengaluru, IN (Remote)  •  4 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

https://atlas.bamboohr.com/careers/607Senior MySQL DBA (Replication Specialist)

About Us:

Atlas Systems Inc. is a Software Solutions company headquartered in East Brunswick, NJ. Incorporated in 2003, Atlas provides comprehensive range of solutions in the area of GRC, Technology, Procurement, Healthcare Provider and Oracle to customers across the globe. Combining our unparalleled experience of over a decade in the software industry and global reach, we have grown with extensive capabilities across industry verticals.

For more information, please visit our website https://www.atlassystems.com/

Please click on the link below to apply for this position:

https://atlas.bamboohr.com/careers/608

Job Title: Senior MySQL DBA – Replication Specialist

Location: [Client Location] (Remote / Hybrid as applicable)

Work Timing: 6AM EST – 7PM EST

Experience: 8+ Years

Engagement Type: Long-term / Multi-year (Contract / Full-time)

We are seeking an experienced Senior MySQL Database Administrator with deep expertise in MySQL replication to support and improve a client's production database environment. The client currently runs MySQL 5.7, with an active migration to MySQL 8 underway. The client's master/primary database is managed by a separate vendor; this role is responsible for the client's three MySQL read replicas, one of which serves live read traffic for the client's application.

This role requires strong hands-on experience designing, monitoring, and troubleshooting MySQL InnoDB replication at scale (~2.5TB), including diagnosing replication delays, rebuilding failed replicas efficiently, and advising the client on ways to modernize and harden the overall replication architecture.

Key Responsibilities

Replication Monitoring & Incident Response

Perform daily monitoring of replication health and lag across all read replicas, proactively identifying and responding to delays before they impact the application.

Diagnose and resolve replication failures on the ~2.5TB production database, minimizing time-to-recovery.

Coordinate with the client's infrastructure team when a replica failure requires a DNS change to route application read traffic to a healthy replica.

Rebuild failed or broken replicas, and identify ways to speed up the rebuild process (e.g., parallelized data copy, Percona XtraBackup-based provisioning, snapshot/clone-based rebuilds, network and disk I/O tuning).

Replication Architecture & Improvement

Design & Implement Topologies: Architect, deploy, and manage advanced MySQL replication environments, including traditional Asynchronous, Semi-Synchronous, and Group Replication.

Evaluate the client's current replication setup (externally managed master, three read replicas) and recommend and implement improvements to resiliency, failover speed, and rebuild time.

Apply deep replication knowledge — binary logging (Row-Based vs. Statement-Based Replication), GTID (Global Transaction Identifiers), and multi-source replication — to troubleshoot and optimize the environment.

Support the client's MySQL 5.7 to MySQL 8 migration, ensuring replication compatibility and minimal disruption across all replicas.

Cluster Management & Disaster Recovery

Administer and monitor production MySQL InnoDB Clusters, ClusterSets, and Galera/Percona XtraDB Clusters where applicable.

Build, maintain, and test backup and point-in-time recovery (PITR) strategies using Percona XtraBackup, MySQL Enterprise Backup, or cloud-native snapshots.

Performance Optimization

Monitor, diagnose, and resolve multi-threaded replication delays caused by long-running transactions or disk I/O bottlenecks.

Identify and optimize slow queries impacting the master node to prevent replica performance degradation.

Monitoring & Tooling

Configure and monitor database metrics using Percona Monitoring and Management (PMM) and related observability tooling.

Use MySQL Router, ProxySQL, or HAProxy for intelligent connection pooling and read/write splitting across replicas.

Required Qualifications

8+ years of hands-on MySQL DBA experience, with a strong specialization in replication architecture and operations.

Proven experience administering production MySQL InnoDB replication environments at multi-terabyte scale (2TB+).

Experience operating in environments where the primary/master database is managed by a third-party vendor, with DBA ownership limited to read replicas.

Demonstrated experience rebuilding failed replicas on large databases and reducing rebuild time.

Experience with MySQL 5.7 to MySQL 8 upgrade/migration projects.

Strong troubleshooting skills for replication lag, replication failures, and DNS-based failover/traffic rerouting.

Technical Skills

MySQL & Replication

MySQL 5.7/8.0, InnoDB Replication, Asynchronous/Semi-Synchronous/Group Replication, GTID, Row-Based vs. Statement-Based Replication (binary logging), Multi-Source Replication, MySQL InnoDB Cluster/ClusterSet, Galera/Percona XtraDB Cluster.

Backup & Recovery

Percona XtraBackup, MySQL Enterprise Backup, Point-in-Time Recovery (PITR), cloud-native snapshots.

Monitoring & Proxy Tools

Percona Monitoring and Management (PMM), MySQL Router, ProxySQL, HAProxy.

Performance & Operations

Query optimization and tuning, disk I/O and OS-level performance tuning, DNS-based failover/traffic routing, large-scale (2.5TB+) database rebuild and provisioning.

Core Competencies

MySQL Replication Architecture

Disaster Recovery & Backup Strategy

Performance Troubleshooting & Optimization

Incident Response & Root-Cause Analysis

Monitoring & Observability

Process Improvement & Automation

Stakeholder Communication

Attention to Detail Under Production Pressure

Success Measures

Reduced replica rebuild time following failures, measured against current baseline.

Reduced frequency and duration of replication lag/delay incidents.

Improved mean-time-to-recovery (MTTR) for replica failures, including DNS-based failover events.

Disruption-free completion of the MySQL 5.7 to MySQL 8 migration across all replicas.

Strengthened monitoring coverage and proactive alerting for replication health.

Documented, repeatable rebuild and failover runbooks adopted by the client team.

Atlas Systems

About Atlas Systems

With offices in the US and India, Atlas Systems is a trusted partner helping companies on their digital transformation journeys – expanding their capabilities and delivering added value. Leveraging innovative technologies, such as AI and Cloud, Atlas works closely with clients to provide technology solutions that seamlessly enhance in-house teams and systems.

Atlas’s offerings include:

- IT services & software development: Collaborating extensively with in-house teams and bringing the benefits of AI and other innovations

- PRIME: The premiere service for verifying health insurance provider directory data

- Provide-Payer Connect - A breakthrough in the instant transfer of health provider data

- ComplyScore®: An AI-driven approach to third-party risk management (TPRM)

Industry
Unknown
Company Size
201-500 employees
Headquarters
East Brunswick, New Jersey
Year Founded
2003
Social Media