Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI SDET (Software Development Engineer in Test) based in Brazil.
This role sits at the intersection of software quality engineering, artificial intelligence, and regulatory assurance.
You will own and evolve an evaluation harness designed to rigorously test AI agents and their non-deterministic outputs.
Your work will help ensure AI systems remain accurate, safe, explainable, and compliant across a wide range of scenarios.
You will design automated evaluations, adversarial test cases, and benchmark comparisons that expose edge cases and logic failures.
The position also involves validating data quality, human-in-the-loop checkpoints, population segmentation, and model performance.
You will collaborate closely with engineering teams to translate evaluation findings into reliable improvements.
This is an ideal opportunity for a highly analytical SDET who enjoys challenging AI systems and applying a rigorous, evidence-driven approach to quality.
Accountabilities:
- Own and manage the AI evaluation harness, including the creation, maintenance, and continuous improvement of automated evaluation workflows.
- Develop and maintain “Golden Path” scenarios representing expected AI behavior and critical user journeys.
- Design adversarial and edge-case test scenarios to identify weaknesses in AI agents and expose failures that conventional testing may overlook.
- Validate regulatory quality gates and compliance requirements, including frameworks such as SR 26-2 and NYDFS Regulation 187.
- Execute shadow-mode validation to compare AI system performance against established human benchmarks and analyze differences statistically.
- Test for data quality degradation, population segmentation errors, inconsistencies, and other risks that could affect model reliability.
- Evaluate the reliability of human-in-the-loop checkpoints and verify that required intervention points operate correctly.
- Confirm that AI-generated outputs consistently meet defined standards, including explanation blocks, signal types, and other required output structures.
- Produce clear, traceable evaluation evidence and communicate findings to engineering and other technical stakeholders.
- Partner closely with engineers to investigate failures, identify root causes, and improve AI system logic and quality.
Requirements:
- Advanced proficiency in Python 3.11+, Pytest, and modern automated software testing frameworks.
- Hands-on experience with AI evaluation platforms and methodologies, including Langfuse, RAGAS, or comparable tools.
- Practical knowledge of LLM-as-a-judge evaluation patterns and approaches for assessing non-deterministic AI outputs.
- Understanding of model risk management principles and regulatory compliance testing for AI-driven systems.
- Experience with shadow-mode deployment strategies, statistical validation, and comparative model evaluation.
- Strong understanding of automated testing principles and the ability to design robust, repeatable test scenarios.
- Exceptional attention to detail, particularly when producing evidence related to regulatory requirements and quality gates.
- A naturally skeptical and adversarial mindset, with the ability to question expected behavior and proactively identify edge cases.
- Strong analytical and problem-solving abilities, with a structured approach to investigating complex AI behavior.
- Excellent collaboration and communication skills, particularly when working with engineers to diagnose and resolve logic failures.
- Ability to work effectively in an environment where AI systems, testing methodologies, and requirements evolve rapidly.
Benefits:
- Health and dental insurance.
- Meal and food allowances.
- Childcare assistance.
- Extended paternity leave.
- Access to gyms and health and wellness professionals through partner programs such as Wellhub and TotalPass.
- Profit Sharing and Results Participation (PLR).
- Life insurance.
- Continuous learning opportunities through a dedicated corporate learning platform.
- Partnerships with online education and professional development platforms.
- Language learning resources and programs.
- Access to digital resources focused on physical, mental, and overall well-being.
- Pregnancy and responsible parenting courses.
- Inclusive workplace practices and dedicated support for employees with disabilities.
- Access to health, well-being, inclusion specialists, and employee affinity initiatives.
- Opportunities to work on innovative AI and software quality challenges with multidisciplinary teams.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1