We are seeking a Senior AI Researcher to lead post-training evaluation, red-teaming, and reinforcement learning (RL) gym audits on open-weight models. The ideal candidate will establish rigorous benchmarking methodologies, evaluate large language models (LLMs) against complex threats like Indirect Prompt Injections (IPI), and construct post-training evaluation pipelines that accurately measure realistic frontier-level security capabilities.
RL Post-Training & Benchmarking
Task & Environment Auditing
Performance Reporting (Gym Cards)
Integration & Orchestration
Alice is a trust, safety, and security company built for the AI era. We safeguard the communicative technologies people use to create, collaborate, and interact- whether with each other or with machines.
In a world where AI has fundamentally changed the nature of risk, Alice provides end-to-end coverage across the entire AI lifecycle. We support frontier model labs, enterprises, and UGC platforms with a comprehensive suite of solutions: from model hardening evaluations and pre-deployment red-teaming to runtime guardrails and ongoing drift detection.
Alice is widely considered a global leader in online safety and AI security. We have some of the most forward-thinking and passionate minds in the world working to safeguard over 3 billion users across the largest AI and tech platforms.
If you're creative and driven to secure the future of AI, we want to hear from you!

ActiveFence is the leading provider of AI security and safety solutions, protecting online experiences and AI applications for over 3 billion users, top foundation models, and the world’s largest enterprises and tech platforms.
As a trusted partner to major technology companies and Fortune 500 brands, we secure user-generated and GenAI products against prompt injection, adversarial attacks, and harmful content through Real-Time Guardrails, continuous Red Teaming, and the industry’s most advanced threat intelligence.
With unmatched detection capabilities in 117+ languages, ActiveFence empowers organizations to deliver engaging, safe, and trustworthy experiences globally, helping them innovate responsibly while staying ahead of emerging threats.