AI Evaluation Engineer (Python, QA or Security)
Mindrift · AI/ML
Source: Himalayas
This role’s source listing does not include Hungary as an eligible location.
Mindrift is seeking an AI Evaluation Engineer for project-based work to create and test coding scenarios for AI agents, specifically for candidates based in Japan.
This role involves building realistic developer environments, designing complex coding tasks, and writing automated evaluation tests to challenge frontier AI models. You will be responsible for iterating on these tasks based on feedback and analyzing model failures to ensure robust evaluation criteria.
Why You’ll Love This Role
What You’ll Bring
What You’ll Get
About the Company
Mindrift connects specialized professionals with project-based opportunities to help leading tech companies test, evaluate, and improve their AI systems. They focus on building sophisticated datasets to benchmark AI coding agents against real-world development tasks.