Job Description
Our team builds the ML-inference stack that powers generative AI for Apple Intelligence's Private Cloud Compute — running on Apple Silicon in the datacenter, distributing work across on-SoC acceleration hardware and multi-node clusters. Built on Private Cloud Compute's privacy guarantees, we're growing the team to scale across more platforms and support a widening set of features.
As part of the team you will help engineer continuous improvements in stability and performance for Private Cloud Compute, help implement entirely new functionality as it emerges from the research community, and help bring our inference stack up on new generations of SoCs and hardware acceleration IP as we extend to more platforms — in collaboration with hardware, product and research teams throughout Apple.
We write performant and scalable frameworks (primarily in Swift, with C++ where we bridge to the hardware) to distribute and coordinate ML inference across the acceleration IP blocks of different SoCs, and to move data and coordinate work reliably across multi-node inference clusters. You will integrate inference code into a full service stack so that user traffic is served reliably and performantly, with a strong focus on code that is easy and safe to develop, update, and monitor in production.
We're a collection of highly skilled and friendly engineers who value each other's opinions and experience. We strive for excellence and believe strongly in the quality of our output. We are a team of domain experts, each specializing in specific core subject areas, with broad collective experience across cloud software services and platforms.
Preferred Qualifications
Low-level or close-to-the-metal work — systems programming, performance, or hardware/SoC bring-up.
Server-side Swift, or RPC/networking stacks (gRPC, Protocol Buffers).
Strong debugging and observability instincts — fluent in logs, metrics, and traces, with observability-stack (OpenTelemetry, Splunk), SLO/error-budget, and on-call experience, and able to drive a production incident to root cause.
Depth in ML inference serving optimizations — quantization, sparsity, batching, KV cache, tokenization, GPU acceleration — enough to optimize the system and reason about the trade-offs and requirements it places on models (you won't be training them).
Solid grasp of concurrency, async/streaming, resource lifecycle, and error/cancellation handling; bonus for Apple platform experience (XPC, Instruments, Swift Concurrency).
Minimum Qualifications
2 Years practical experience plus Bachelor's degree in Computer Science, Computer Engineering, or a related field — or equivalent practical experience.
Experience building large-scale distributed systems that serve ML inference, reasoning across processes, hosts, and service tiers as well as model behavior under load.
A ML performance-centric mindset — able to reason about latency/throughput trade-offs and to distinguish what can be solved at the system level from what requires model–system co-design.
Strong in a systems or server language — Python, Go, Rust, Java, C++, or similar. Swift is what we write, but we don't expect it going in.