Sprout

Server Engineer - AI Systems

Sprout  •  Onsite  •  2 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

Locations: Garland, Texas

Categories: Operations

Req ID: 2632

Overview

Sprout is a global IT hardware retirement provider for hyperscaler and enterprise clients. We leverage a nationwide footprint (and international partner network) combined with proprietary software to enable efficient end-to-end IT asset disposition with a focus on data-bearing devices from the client to the cloud. The company is headquartered in Charlotte, NC with additional operations near Sacramento, and Dallas. Sprout provides software and services to clients in the form of our IT Asset Disposition, Certified Destruction, and Responsible Recycling solutions.

Since our founding as an electronic waste startup from a Duke University dorm room in 2014, we have been expanding at an average rate of >66% each year. By adhering to our 3 values (One Sprout, Deliver Excellence, and Integrity Matters), we are proud of our culture to move at #SproutSpeed to become the emerging leader in our industry. For more information, please visit: www.sproutup.com.

Velerity Compute builds and recertifies AI-capable GPU servers: DGX and HGX-class A100 and H100 systems, InfiniBand fabric, and the hyperscaler and enterprise decom gear that feeds both. Today, every technical question a customer or broker asks routes through people whose primary job is running the floor or closing the deal. Answers get slow, configs get quoted before anyone confirms they are buildable, and deals stall on questions that take an experienced tech ten minutes.

This role is the technical answer. You are the person Sales calls when a customer asks whether four HGX baseboards can be made into two sellable systems, what the fabric needs to look like, whether a rack of R750s can be reconfigured to spec, or whether the units on the floor will pass certification. You spend most of your week in front of Sales and customers, and enough of it in front of the hardware to keep your answers honest.

Responsibilities

Technical support to Sales (roughly 60 percent)

  • Run GPU diagnostics and interpret results, including NVIDIA field diagnostics, DCGM, and XID fault classification, and make disposition calls on that basis
  • Troubleshoot faults across GPU, baseboard, NVLink and NVSwitch, PSU, cooling, and fabric
  • Configure and validate InfiniBand and Ethernet fabric, including adapter firmware, cabling, and transceiver compatibility
  • Own configuration and assembly quality for systems being built to a customer order
  • Own burn-in and certification testing, and sign off on units as sellable
  • Feed accurate technical attributes back into the master data and inventory record so what Sales sees is what is on the floor

Hands-on technical work (roughly 40 percent)

  • Perform BIOS, firmware, and BMC audits on inbound AI systems and hyperscaler gear
  • Run GPU diagnostics and interpret results, including NVIDIA field diagnostics, DCGM, and XID fault classification, and make disposition calls on that basis
  • Troubleshoot faults across GPU, baseboard, NVLink and NVSwitch, PSU, cooling, and fabric
  • Configure and validate InfiniBand and Ethernet fabric, including adapter firmware, cabling, and transceiver compatibility
  • Own configuration and assembly quality for systems being built to a customer order
  • Own burn-in and certification testing, and sign off on units as sellable
  • Feed accurate technical attributes back into the master data and inventory record so what Sales sees is what is on the floor

Making the answer repeatable

  • Document recurring technical questions into reusable reference material so Sales stops asking the same thing twice
  • Maintain the internal config catalog: what Velerity builds, what it costs to build, and what it takes to build it
  • Train sales staff on enough technical baseline to qualify opportunities without escalating everything

Qualifications

  • 5 or more years in data center hardware engineering, field engineering, systems integration, or a hyperscaler or ODM environment
  • Direct hands-on experience with NVIDIA AI systems: DGX or HGX A100 and H100 platforms, SXM and PCIe form factors, NVLink and NVSwitch topology
  • Working knowledge of InfiniBand fabric: HDR and NDR, ConnectX adapters, Quantum switches, subnet management, cabling and transceiver rules
  • Broad enterprise server fluency across the major OEM lines, not just accelerated platforms. You should be able to spec, configure, and troubleshoot across:
  • Dell EMC PowerEdge, including the R and XE series, iDRAC, PERC controllers, and OpenManage
  • HPE ProLiant and Apollo, including iLO, Smart Array, and Intelligent Provisioning
  • Supermicro, including SuperServer and GPU chassis lines, IPMI, and the SuperMicro Update Manager
  • Lenovo ThinkSystem, including XClarity
  • Cisco UCS, including B and C series, fabric interconnects, and UCS Manager or Intersight service profiles
  • ODM and hyperscaler gear: OCP form factors, Open Rack, and high-density sled and blade platforms from the major ODMs
  • Able to read a mixed manifest of enterprise gear and say quickly what it is, what it is worth configuring into, and what it will take
  • Firmware and low-level configuration: BIOS, BMC, IPMI, Redfish, and OEM-specific management tooling and firmware bundles
  • GPU diagnostics and fault interpretation, including XID codes, ECC and row remap behavior, bus and link faults, and firmware level faults
  • Data center power and cooling fundamentals: three-phase power, rack level power budgeting, PDU configuration, high-density air cooling, exposure to liquid cooling
  • Demonstrated ability to explain technical constraints to a non-technical audience without either dumbing it down or burying them

EEO – Equal Employment Opportunity

The Company is an equal opportunity employer and does not discriminate on the basis of and all qualified applicants will receive consideration for employment without regard to race, creed, color, sex, affectional or sexual orientation, gender identity or expression, gender, ethnicity, religion, national origin, ancestry, nationality, age, disability, marital status, veteran status, genetic information, or on any other basis prohibited by law (except where an attribute is a bona fide occupational qualification)

Sprout

About Sprout

Sprout is the intelligent IT asset management platform that unifies data, workflows, and reporting across the full technology lifecycle, giving enterprises real-time asset visibility, one-click compliance, and measurable value recovery across global operations. Trusted by 60% of the top 10 S&P 500, Sprout partners with IT, finance, and sustainability leaders at every stage, from first deployment to final disposal and everything in between. With world-class NPS scores, 90% logo retention, and value recovery that outperforms the market by 20–50%, Sprout is the platform - and the partner - chosen by the world's most forward-thinking enterprises.

Industry
IT & Software
Company Size
201-500 employees
Headquarters
Charlotte, North Carolina
Year Founded
2014
Social Media