qode.world

Senior Software Engineer (Performance)

qode.world  •  Socialist Republic of Vietnam (Onsite)  •  3 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

We are looking for a Senior Inference Engineer with a strong foundation in software engineering, distributed systems, and performance optimization to build and optimize inference engines for large-scale LLM serving systems. You will work across both research and production environments, ensuring our LLM serving systems are fast, scalable, and efficient. The role spans the entire inference stack — from kernel and runtime to scheduling, memory management, and distributed execution

Key Responsibilities:

  • Profile, benchmark, and analyze bottlenecks for LLM inference workloads across multiple layers: kernel, memory, networking, and scheduler
  • Optimize inference engines (vLLM, SGLang, TensorRT-LLM) for throughput, latency, memory efficiency, GPU utilization, and cost
  • Implement and fine-tune inference optimization techniques including batching, KV-cache management, quantization, speculative decoding, parallelism strategies, and disaggregated serving
  • Build instrumentation and profiling tools to identify bottlenecks
  • Ensure the reliability of the inference pipeline through A/B launches, rollback, model versioning, and fault tolerance
  • Collaborate with the Platform Engineering team to improve serving architecture based on performance findings
  • Document and share knowledge, contributing to internal best practices and AI open-source projects whenever possible

Requirements

1 - Mandatory:

  • At least 5 years of experience as a Software Engineer, Performance Engineer, or equivalent.
  • Strong foundation in Software Engineering, Software Architecture, and Distributed Systems.
  • Proficiency in at least one of the following languages: Python, Go, or C++.
  • Experience developing or optimizing distributed systems, high-throughput backends, or large-scale serving systems.
  • Experience with benchmarking, profiling, and performance tuning in production environments.
  • Ability to analyze CPU, Memory, Network, or Storage bottlenecks.
  • Strong systems thinking, Root Cause Analysis capabilities, and the ability to solve complex performance problems.
  • Strong ownership mindset and the ability to work independently.

2 - Nice to Have:

  • Experience with Linux internals, kernel tuning, or custom Linux kernel.
  • Understanding of GPU Architecture or CUDA Programming.
  • Experience with AI/ML Serving Systems or LLM Inference.- Have worked with one of the inference engines such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server.
  • Understanding of batching, KV Cache, quantization, speculative decoding, tensor/pipeline parallelism, or disaggregated serving.
  • Experience with the NVIDIA inference stack (TensorRT, Triton, CUTLASS, NCCL, cuBLAS, cuDNN).
  • Experience with observability stacks such as Prometheus, Grafana, or OpenTelemetry.
  • Open-source contributions or research related to AI Infrastructure, ML Systems, or Performance Optimization.
qode.world

About qode.world

We revolutionize how talent finds meaningful careers by harnessing the power of data and automation. Our platform utilizes LLMs to parse resumes and reconstruct queries, transforming unstructured data into actionable insights. This enables us to build robust data moats, such as creating 'Private Talent Pools' for recruiters where autonomous agents enrich candidate profiles.

By automating high-volume recruiting workflows, we reduce the marginal cost of work to zero. Agents match profiles to job descriptions, find contact information, and send personalized messages and schedule interviews automatically, significantly decreasing the time to close. Additionally, we transcribe the interviews and make the data searchable, making hiring decisions more objective.

We drive confidence by raising the quality bar for job seekers. We automate technical exercises such as coding tests, evaluate candidates on merit, providing recruiters with pass/fail scores and qualitative feedback.

We also provide Exclusive or Retained Recruitment services, offering specialized recruitment with no upfront cost or a retained model with a partial fee, ensuring exclusivity and dedicated support throughout the hiring process.

Our Fractional Head of People and HR Advisory services offer flexible, strategic support through part-time or interim roles, as well as comprehensive advisory services to guide crucial HR decision-making.

Lastly, our HR Due Diligence process provides thorough insights into the HR frameworks of target companies, helping mitigate risks across the board.

How do you envision the future of recruiting with the integration of such advanced technologies?

Industry
IT & Software
Company Size
51-200 employees
Headquarters
Singapore , SG
Year Founded
2023
Social Media