Job Description
We are seeking a hands-on Senior Platform Engineer to support the deployment, validation, and operation of large-scale solutions. This role focuses on system bring-up, hardware and firmware validation, and operational reliability across storage, compute, and networking infrastructure.
You will work closely with senior engineers to ensure systems are correctly configured, stable, and performant: from BIOS settings and firmware versions to OS behavior and device health.
What You’ll Do
- Perform system bring-up and validation, including BIOS configuration, firmware updates, and OS-level checks.
- Configure and verify platform settings (BIOS, BMC, OS) to ensure systems meet performance, reliability and consistency standards.
- Execute and validate firmware updates across components, including NVMe drives, NICs, and system firmware.
- Identify and troubleshoot issues related to:
- Device initialization and visibility.
- Firmware mismatches or upgrade failures.
- Misconfigurations impacting performance or stability.
- Understand, define and test procedures for:
- Drive swaps and rebuilds.
- Data integrity and system health checks.
- Minimizing impact to running systems.
- Monitor and validate storage device health, including basic performance characteristics and failure indicators.
- Work with Linux systems to verify configuration, collect diagnostics, and assist in debugging issues.
- Collaborate with senior engineers to escalate and help triage complex cross-layer issues.
- Contribute to documentation and standardization of bring-up, validation, and operational procedures.
What will you bring to DDN
- 5+ years of experience in systems, infrastructure, or platform-focused engineering roles.
- Comfortable working directly with hardware and firmware, including BIOS, BMC, and device-level tools.
- Solid understanding of Linux systems administration and troubleshooting.
- Familiarity with storage and networking fundamentals (NVMe, NICs, basic I/O concepts).
- Experience performing and supporting firmware updates, with awareness of risk and failure scenarios.
- Good scripting skills; familiarity with Go is a plus.
- Strong attention to detail, able to verify configurations and catch subtle issues.
- Willingness to learn and grow into deeper system, performance, and debugging responsibilities.
- Curiosity about why systems perform the way they do: investigate, rather than just run scripts blindly.