Cerebras

Senior Cloud Quality Engineer

Cerebras  •  Bengaluru, IN (Onsite)  •  15 days ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of a single device. This approach allows Cerebras to deliver industry-leading training and inference speeds and empowers machine learning users to effortlessly run large-scale ML applications, without the hassle of managing hundreds of GPUs or TPUs.

Cerebras' current customers include top model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

Thanks to the groundbreaking wafer-scale architecture, Cerebras Inference offers the fastest Generative AI inference solution in the world, over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

About The Team

The Cloud Quality team is responsible for the confidence behind every production release shipped to Cerebras Inference Cloud.

We work closely with platform, infrastructure, ML systems, and product engineering teams to ensure that rapid iteration never comes at the expense of customer trust. Our environment spans distributed cloud systems, multi-region deployments, APIs, orchestration layers, and hardware-backed inference services.

We are scaling quickly. The systems are growing in complexity, traffic is increasing rapidly, and release velocity remains high. We need engineers who can build quality systems that scale with the business.

About The Role

We are hiring a Senior Quality Engineer to own the quality of our weekly cloud releases end to end and to build the test infrastructure that lets the team scale. This is a hands-on senior IC role for someone who treats quality as a first-class engineering problem - not a downstream gate.

You will drive every release from branch cut to sign-off, build scalable test infrastructure that grows with customer load, and push back when quality is at risk. You will operate effectively across timezones, async-first, with clear written communication.

This is a role for someone who self-drives. You will frequently work without complete product specifications, decide under ambiguity, and ask the right questions before code lands - not after.

Responsibilities

  • Release Quality Ownership: Drive weekly cloud release qualification end to end. Read every PR in the release branch first-hand; understand what changed; decide where the risk is; and design the qualification that exercises the actual risk. Be the final voice before a release ships.
  • Test Infrastructure at Scale: Build and evolve the test infrastructure - functional, integration, performance, and fault for the Inference Cloud platform. Plan for 20x growth in coverage, environments, and traffic. Today's setup will not survive tomorrow's load; design for the next horizon.
  • End-to-End System Understanding: Reason through the full stack — client SDK, API, gateway, inference software, driver, hardware. Know enough to debug from any layer and to test the right thing.
  • Code Review with Intent: Read and review developer PRs with genuine understanding of what each change does and what its blast radius is. Test the change's actual impact, not its surface area.
  • Automation Expansion: Increase automation coverage continuously. Fix flaky tests rather than tolerate them. Use AI tooling effectively to accelerate test creation, debugging, and analysis.
  • Quality Discipline: Choose high-value tests over volume metrics. Drive the team's standards for what "tested" means and what "ready to ship" means.
  • Cross-Team Operation: Work with platform, ML, infrastructure, and product teams across timezones. Influence quality outcomes without owning every team's roadmap.

Skills & Qualifications

  • 5+ years of experience in quality engineering, test engineering, or a closely related role, with substantial individual contributor experience on large-scale distributed systems or cloud infrastructure.
  • Deep cloud platform experience, preferably AWS - networking, compute orchestration, container platforms, and multi-region production services. You can reason about what is happening at the cloud layer when something fails.
  • Track record of building scalable test infrastructure - frameworks, harnesses, environments, and automation that scale with the system under test rather than fighting it.
  • Strong systems debugging and reasoning. You can take an unfamiliar failure and follow it through layers of the stack to a root cause.
  • Strong proficiency in at least one backend language (Python, Go, or C++), sufficient to read production code, write production-grade tests, and contribute infrastructure code directly.
  • Excellent written and async communication. You operate effectively across time zones and in environments where most decisions get made in writing.
  • Self-direction under ambiguity. You frame problems, make trade-off decisions, and push back when quality is at risk - without waiting to be asked.
  • Experience with Cloud infrastructure, model serving systems, or GPU accelerated workloads is a strong plus.
  • Experience using AI tooling (LLMs, coding assistants, agents) to accelerate test development, triage, or analysis is a plus.

Why Join Cerebras

People who are serious about software make their own hardware. At Cerebras we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

  1. Build a breakthrough AI platform beyond the constraints of the GPU.
  2. Publish and open source their cutting-edge AI research.
  3. Work on one of the fastest AI supercomputers in the world.
  4. Enjoy job stability with startup vitality.
  5. Our simple, non-corporate work culture that respects individual beliefs.

Read our blog: Five Reasons to Join Cerebras in 2026.

Apply today and become part of the forefront of groundbreaking advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Cerebras

About Cerebras

Cerebras Systems is the world's fastest AI inference. We are powering the future of generative AI. Follow us for model breakthroughs and real-time AI results.

We’re a team of pioneering computer architects, deep learning researchers, and engineers building a new class of AI supercomputers from the ground up.

Our flagship system, Cerebras CS-3, is powered by the Wafer Scale Engine 3—the world’s largest and fastest AI processor. CS-3s are effortlessly clustered to create the largest AI supercomputers on Earth, while abstracting away the complexity of traditional distributed computing.

From sub-second inference speeds to breakthrough training performance, Cerebras makes it easier to build and deploy state-of-the-art AI—from proprietary enterprise models to open-source projects downloaded millions of times.

Here’s what makes our platform different:

🔦 Sub-second reasoning – Instant intelligence and real-time responsiveness, even at massive scale

⚡ Blazing-fast inference – Up to 100x performance gains over traditional AI infrastructure

🧠 Agentic AI in action – Models that can plan, act, and adapt autonomously

🌍 Scalable infrastructure – Built to move from prototype to global deployment without friction

Cerebras solutions are available in the Cerebras Cloud or on-prem, serving leading enterprises, research labs, and government agencies worldwide.

👉 Learn more: www.cerebras.ai

Join us: https://cerebras.net/careers/

Industry
Hardware & Semiconductors
Company Size
501-1,000 employees
Headquarters
Sunnyvale, California
Year Founded
Unknown
Social Media