Job Description

Role Overview

We're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments. You'll design tasks, write tests, and iterate on tasks based on QA feedback.

What You Will Do

Create tasks and tests for AI agents, review agent solutions, analyze failures, and refine until the evaluation is fair and robust.

Why It Might Be a Fit

You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.

Requirements

  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Benefits

  • Up to $50/hr equivalent
  • Flexible schedule
  • 20 hours per task, set your own pace
Apply now
Report job

More job openings