Role Overview
We're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments. You'll design tasks, write tests, and iterate on tasks and tests based on QA feedback.
What You Will Do
Create tasks and evaluation criteria, write tests, and iterate on tasks and tests to ensure the evaluation is fair and robust. You'll work on a project-based basis, not permanent employment.
Why It Might Be a Fit
You'll need 5+ years in software development, experience writing tests, and English proficiency (B2+). You'll work with a core stack of Python, JavaScript/TypeScript, Docker, Postgres, Kafka, and Redis.
Requirements
Benefits