Type: Contract, hourly
Location: Remote
Hours: 20–30 per week
Pay: $20–30/hour, based on experience and language coverage
Start: Immediate
ABOUT THE ROLE
We evaluate how well advanced AI coding agents solve real engineering problems. An agent is given a real open source codebase inside a container and a hard task, then works on its own for 80 to 250 steps. A trajectory is the full record of that run — every command, result, and decision.
You will do two jobs, and you should expect either on any given day.
•Annotate — Read a trajectory nobody has looked at yet and judge it step by step.
• Review — Take an existing annotation, written by our AI tooling or another person, and confirm, correct, or reject it.
TASKS YOU'LL SEE
• Feature build — Add a working feature to a live library without breaking anything that already worked.
• Rebuild — Work out what a compiled tool does by running it, then rebuild it to match its output, exit codes, and file effects.
• Bug hunt — Find and fix twenty undocumented bugs across a dozen files with no test suite, then record what caused them.
Mostly Python and Go, with some Rust, C, and Ct+. A trajectory runs about 100 steps.
WHAT YOU JUDGE IN A TRAJECTORY
• Was the command right for the state the environment was actually in?
• Did the agent read the previous output correctly?
• Was the step wrong, or only inefficient — these are scored differently.
• Where did the run first go off course — usually earlier than where it visibly broke.
• Did the agent notice its own mistake and recover, or keep building on a false assumption?
• Did it game the grader instead of solving the task (e.g., weakening a test or hardcoding an expected value)?
WHAT WE NEED FROM YOU
• Experience — 2+ years in software engineering, DevOps, or site reliability, with real debugging in real codebases.
• Languages — Strong in Python or Go, and able to read a language you've never used.
• Linux — Comfortable with logs, running processes, build failures, and containers.
• Workflow — Everyday Git, diffs, pull requests, and issue tracking.
• Debugging — Able to work with no test suite and no error message pointing at the cause.
• Focus — Able to hold context across a long run, because step 74 can depend on step 12.
• Writing — Clear English, since every judgement needs an explanation another engineer can check.
• Software engineering background required
• Python and Go, Rust, C, and C++.
• Concurrency, asynchronous code, or distributed systems
• Earlier AI evaluation work such as RLHF, supervised fine-tuning, model evaluation, or red teaming
• Reverse engineering, or matching a program you cannot read
• Docker and container tooling
WHO THIS ROLE IS NOT FOR
Image labelling, video labelling, transcription, content moderation, and basic chatbot rating do not prepare you for this work — the role depends on understanding what an agent's commands actually did to a running system.
Bespoke Labs is a Mountain View based Series A AI Research/Data Curation for Agents Lab. We're working with Frontier AI Labs, and F500 Cos to advance the capabilities of AI Agents.