Freelance Agent Evaluation Engineer

Mindrift · AI/ML

Apply
FreelanceUSD 40 / hourlyPosted Sep 7, 2026

Source: Himalayas

RemoteJobSearch Check
Specific countries

Poland — no hybrid days in this listing.

Not open from Hungary

This role’s source listing does not include Hungary as an eligible location.

Category
AI/ML
Employment type
Freelance
Experience level
Senior
Job salary
USD 40 / hourly

How this role compares to the AI/ML market

Across all remote AI/ML roles on RemoteJobSearch. Refreshed every 30 minutes.

Pay, compared

63 of 265 AI/ML roles publish a salary

Typical AI/ML role$212,444 / year

Median of published salaries. Full reported spread in this market: $75,000–$351,000, outliers included.

265
Similar jobs
120
Companies hiring
68
Added this week

26% of these roles

Browse remote AI/ML roles worldwide →

Original employer: MindriftSource ATS: HimalayasJob status: Active

Mindrift is hiring a freelance Agent Evaluation Engineer in Poland to create challenging AI coding agent tasks and tests.

As a Freelance Agent Evaluation Engineer, you will build complex, realistic developer environments and design tasks to evaluate how well AI models handle real-world programming challenges. You will iterate on evaluation criteria based on QA feedback to ensure robust performance analysis of frontier models, working on a project-based, flexible basis.

Why You’ll Love This Role

Location: Poland
Pay: USD 40 / hourly
Freelance project-based role
Focus on AI coding agent evaluation
Remote flexibility

What You’ll Bring

5+ Years in software development
Proficiency in Python, FastAPI, JavaScript, React, Docker, Postgres, Kafka, and Redis
Experience writing functional and integration tests
B2+ English proficiency
Master’s Degree in CS/Software Engineering/Data Science or Bachelor’s with 5 years experience
3 Years of professional experience in QA-automation or cybersecurity

What You’ll Get

Benefits not specified in the listing

About the Company

Mindrift connects specialists with project-based AI opportunities for leading tech companies. They focus on testing, evaluating, and improving AI systems through high-quality dataset creation.

Listing checked for remote eligibility · Posted Sep 7, 2026 · HimalayasApply
EXPLORE MORE
Apply