AI Agent Trajectory Annotator and Reviewer

Learn more about Bespoke Labs
Bespoke Labs

Bespoke Labs

AI Agent Trajectory Annotator and Reviewer

Remote
Full Time
Paid
  • Responsibilities

    Type: Contract, hourly

    Location: Remote

    Hours: 20–30 per week

    Pay: $20–30/hour, based on experience and language coverage

    Start: Immediate

    ABOUT THE ROLE

    We evaluate how well advanced AI coding agents solve real engineering problems. An agent is given a real open source codebase inside a container and a hard task, then works on its own for 80 to 250 steps. A trajectory is the full record of that run — every command, result, and decision.

    You will do two jobs, and you should expect either on any given day.

    •Annotate — Read a trajectory nobody has looked at yet and judge it step by step.

    • Review — Take an existing annotation, written by our AI tooling or another person, and confirm, correct, or reject it.

    TASKS YOU'LL SEE

    • Feature build — Add a working feature to a live library without breaking anything that already worked.

    • Rebuild — Work out what a compiled tool does by running it, then rebuild it to match its output, exit codes, and file effects.

    • Bug hunt — Find and fix twenty undocumented bugs across a dozen files with no test suite, then record what caused them.

    Mostly Python and Go, with some Rust, C, and Ct+. A trajectory runs about 100 steps.

    WHAT YOU JUDGE IN A TRAJECTORY

    • Was the command right for the state the environment was actually in?

    • Did the agent read the previous output correctly?

    • Was the step wrong, or only inefficient — these are scored differently.

    • Where did the run first go off course — usually earlier than where it visibly broke.

    • Did the agent notice its own mistake and recover, or keep building on a false assumption?

    • Did it game the grader instead of solving the task (e.g., weakening a test or hardcoding an expected value)?

  • Qualifications

    WHAT WE NEED FROM YOU

    • Experience — 2+ years in software engineering, DevOps, or site reliability, with real debugging in real codebases.

    • Languages — Strong in Python or Go, and able to read a language you've never used.

    • Linux — Comfortable with logs, running processes, build failures, and containers.

    • Workflow — Everyday Git, diffs, pull requests, and issue tracking.

    • Debugging — Able to work with no test suite and no error message pointing at the cause.

    • Focus — Able to hold context across a long run, because step 74 can depend on step 12.

    • Writing — Clear English, since every judgement needs an explanation another engineer can check.

  • Desired skills

    • Software engineering background required

    • Python and Go, Rust, C, and C++.

    • Concurrency, asynchronous code, or distributed systems

    • Earlier AI evaluation work such as RLHF, supervised fine-tuning, model evaluation, or red teaming

    • Reverse engineering, or matching a program you cannot read

    • Docker and container tooling

    WHO THIS ROLE IS NOT FOR

    Image labelling, video labelling, transcription, content moderation, and basic chatbot rating do not prepare you for this work — the role depends on understanding what an agent's commands actually did to a running system.

  • Compensation
    20$ - 30$ USD/hr based on experience
  • Industry
    Information Technology and Services
  • About Us

    Bespoke Labs is a Mountain View based Series A AI Research/Data Curation for Agents Lab. We're working with Frontier AI Labs, and F500 Cos to advance the capabilities of AI Agents.