GroqCore is building a high-performance software stack that converts bare metal compute into a token-producing engine This role is for a seasoned engineer with practical experience making LLMs fast, efficient, and reliable in production. This individual will work on an inference serving software stack from a trained model to a high-throughput delivery of output tokens. This will require experience on understanding/working with model internals, runtime software, and the supporting GPU & accelerator hardware underpinning it all.
Location: This role will be based in one of our three hiring hubs: the Dallas, San Francisco, or New York City area. The person hired for this role must be based in one of these three areas. You’ll have the flexibility to work remotely while we establish our local Groq office, with the expectation that this role will transition to onsite once the office opens.
If this sounds like you, we’d love to hear from you!
Groq is committed to providing competitive compensation through our Total Cash philosophy, which incorporates potential bonus value directly into base pay. The total cash salary range for this position, which is inclusive of the potential bonus value, is $341,400 - $401,600, with individual placement determined by your geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates will vary based on local market dynamics. Beyond cash compensation, Groq also offers a Long-Term Incentive (LTI) Program and a robust suite of employee benefits.
#LI-MS1

Inference is the engine that powers AI, and Groq was built from the silicon up to deliver the world's fastest inference at scale. We pioneered the LPU—the first processor designed specifically for AI inference—and are transforming that innovation into a global cloud platform powering production AI workloads.
With the capital, infrastructure, and team to execute, we're uniquely positioned to define the next era of AI infrastructure. The opportunity is massive, and it's still wide open. Now let's go build it!