ActiveFence

Senior AI Researcher

ActiveFence  •  Ramat Gan, IL (Onsite)  •  5 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

We are seeking a Senior AI Researcher to lead post-training evaluation, red-teaming, and reinforcement learning (RL) gym audits on open-weight models. The ideal candidate will establish rigorous benchmarking methodologies, evaluate large language models (LLMs) against complex threats like Indirect Prompt Injections (IPI), and construct post-training evaluation pipelines that accurately measure realistic frontier-level security capabilities.

Key Responsibilities

RL Post-Training & Benchmarking

  • Execute post-training runs (e.g. GRPO) using mainstream open-weight generalist models against security-focused RL environments, targeting threat vectors like Indirect Prompt Injection (IPI).
  • Reward Diagnostics & Trace Analysis - Analyze live loss curves and rollout traces to identify reward hacking, lazy policy convergence, and flawed or over/under-specified verifiers.

Task & Environment Auditing

  • Review tasks and multi-turn environments (including tool use, web navigation, and computer use) for realism, threat model accuracy, data distribution, and dataset balance.

Performance Reporting (Gym Cards)

  • Generate comprehensive evaluation cards detailing hill-climbing performance uplift across checkpoints, failure modes, tokens/turns per rollout, and task-level success rates.

Integration & Orchestration

  • Integrate dockerized environments (e.g., Harbor format) into internal training frameworks, optimizing reset/statefulness semantics, concurrency, and throughput ceilings.

Requirements

Required Qualifications

  • Technical Background: M.S. or Ph.D. in Data Science, Machine Learning, Computer Science, or equivalent practical experience in deep learning.
  • RL & Post-Training Expertise: Strong hands-on experience training large-scale models using RL algorithms (e.g. GRPO, PPO) on open-weight architectures.
  • AI Security Expertise: Solid understanding of LLM vulnerabilities, red-teaming methodologies, and defensive alignment against IPI attacks.
  • Infrastructure Skills: Proficiency in PyTorch, Docker containerization, and distributed training architectures.
  • Diagnostic Skills: Ability to analyze agent rollout traces, craft deterministic rubrics/verifiers, and debug complex reward shaping flaws.

Preferred Qualifications

  • Prior experience working with standard RL gym formats, such as Harbor.
  • Experience evaluating complex agentic workflows in tool-use or web-browser environments.
  • Familiarity with evaluating open-weight models similar to Llama or Mistral against adversarial workloads.

About Alice

Alice is a trust, safety, and security company built for the AI era. We safeguard the communicative technologies people use to create, collaborate, and interact- whether with each other or with machines.

In a world where AI has fundamentally changed the nature of risk, Alice provides end-to-end coverage across the entire AI lifecycle. We support frontier model labs, enterprises, and UGC platforms with a comprehensive suite of solutions: from model hardening evaluations and pre-deployment red-teaming to runtime guardrails and ongoing drift detection.

Alice is widely considered a global leader in online safety and AI security. We have some of the most forward-thinking and passionate minds in the world working to safeguard over 3 billion users across the largest AI and tech platforms.

If you're creative and driven to secure the future of AI, we want to hear from you!

ActiveFence

About ActiveFence

ActiveFence is the leading provider of AI security and safety solutions, protecting online experiences and AI applications for over 3 billion users, top foundation models, and the world’s largest enterprises and tech platforms.

As a trusted partner to major technology companies and Fortune 500 brands, we secure user-generated and GenAI products against prompt injection, adversarial attacks, and harmful content through Real-Time Guardrails, continuous Red Teaming, and the industry’s most advanced threat intelligence.

With unmatched detection capabilities in 117+ languages, ActiveFence empowers organizations to deliver engaging, safe, and trustworthy experiences globally, helping them innovate responsibly while staying ahead of emerging threats.

Industry
IT & Software
Company Size
201-500 employees
Headquarters
New York
Year Founded
2018
Social Media