
Generating one image with a diffusion model means evaluating a large neural network tens to hundreds of times. That repetition is where the energy goes, and it is part of why data-centre demand is now visible at grid scale: the Lawrence Berkeley National Laboratory puts US data-centre electricity use at 176 TWh in 2023, 4.4 % of national consumption, a figure that more than doubled since 2017 largely because of AI servers, and projects 325–580 TWh by 2028. Emerging hardware attacks exactly this primitive: analog in-memory computing, silicon photonics, ferroelectric devices and stochastic computing perform the underlying matrix–vector products at lower energy than digital CMOS, under conditions on scale and precision, but return a perturbed result. The usual objection is that nobody can say in advance how much perturbation a model tolerates. In EMMA, this opportunity is pursued through a PCM-based photonic matrix–vector engine as the primary demonstrator, complemented where appropriate by FeFET-based analog computing, so that device-level non-idealities are treated as explicit design parameters rather than as after-the-fact implementation losses. Diffusion models are an unusual case in which that question has a mathematical answer: their convergence theory bounds the discrepancy between the generated and target distributions by three terms, of which only one, the L 2 error of the learned score, depends on the hardware. That side is now in good shape: for stochastic samplers, bounds linear in the dimension under minimal assumptions on the data; for deterministic samplers, a matching theory under regularity. What is missing is the bridge. The theory is parameterised by a score error, the hardware literature by conductance noise, effective number of bits and photons per multiply–accumulate; nobody has written the map between them, in either direction. The quantisation literature for diffusion models reports image-quality scores; the induced L 2 score error, in the form the convergence theory needs, is directly measurable and goes unreported. One open sub-question sets the direction: the distinction between an error redrawn at each evaluation, as in stochastic computing, and an error frozen at write time, as in deterministic quantisation, has been settled in one setting only, and not the one hardware lives in; carrying it over is what this position exists to do. The work is judged on one concrete outcome. For at least one architecture, we want to show that theenergy gain from moving the score evaluation off digital CMOS is real, measured where it survives integration and not extrapolated from a single tile, and that it is paid for by a loss in sample quality that is small and, above all, tunable: a knob the designer sets, whose position the theory predicts rather than discovers afterwards.
The work is organised in four packages, summarised below; the full research programme is available from the supervisors on request.
A clear negative answer is more useful to us than a vague positive one: a negative result on point 3 leaves points 1 and 2 intact and turns point 4 from a prediction into a measurement.
Prior work on diffusion or score-based generative models; numerical analysis of finite-precision arithmetic. Familiarity with hardware-aware machine learning, approximate or stochastic computing, analogue in-memory computing, device modelling, circuit simulation, photonics, ferroelectric devices, or energy modelling. Experience linking model accuracy or robustness to implementation non-idealities, such as quantisation, noise, drift, variability, finite precision, or correlated errors. Evidence of interdisciplinary research, for example work at the interface of mathematics, machine learning, devices, circuits, architectures, or electronic-design automation. No hardware experience is expected or needed to apply.
We are seeking a researcher who wants to explore how to turn physical constraints into mathematical and architectural design parameters. The project does not require prior laboratory hardware experience: INL provides device expertise, measured characteristics and modelling infrastructure. However, the successful candidate must be prepared to engage critically with non-ideal analog, photonic and ferroelectric accelerators, and to work with both supervisors to connect device-level error mechanisms, score-error guarantees and the energy–quality trade-off of diffusion-model inference. Candidates with a primarily mathematical background and strong curiosity for hardware are welcome, as are candidates from AIhardware, device-modeling, circuit, architecture or EDA backgrounds who have the stochastic-analysis maturity needed to develop the theory. The decisive criterion is the ability to make the mathematics and hardware inform one another, producing predictions that are quantitative, testable and relevant to an energy-efficient accelerator demonstrator.
Twelve-month fixed-term contract under French public law, full time (1 607 h annual reference), 25 working days of annual leave plus the RTT days of the establishment’s scheme. Affiliation to the régime général for health and to IRCANTEC for supplementary pension; employer contribution to complementary health cover under the scheme in force at the establishment; 75 % reimbursement of publictransport season tickets, up to the statutory ceiling. The exact number of RTT days is set by the establishment.
