Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts
Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with
Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases
Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt
Identify failure modes, reward hacking, and degenerate solutions before they reach training
Must-Have Skills
Advanced-level Python
Hands-on RL and model training experience: RLVR, or preference-based training, familiarity with an RL framework such as Gymnasium or dm_env and how environments, action spaces, and reward signals are structured, plus practical experience fine-tuning or post-training LLMs and reading eval signal to know whether a change actually helped
Able to design coding challenges and tasks from scratch, and calibrate difficulty so a task is hard enough to make the model fail
Strong software engineering fundamentals: debugging, writing test cases, and reasoning about edge cases and failure modes
Strong problem-solving on unfamiliar problems, and a fast learner who adapts across domains and projects
Comfortable working in code with standard tooling: git, terminal, Docker
Good to Have
Built RL or preference-data pipelines, not just individual tasks: data curation, rollout collection, reward model training, and the infrastructure around them
Synthetic task generation: programmatically producing task families, verifiers, and difficulty ladders rather than hand-writing each one
Synthetic worlds and simulated environments: procedural or seeded generation, multi-step agentic sandboxes, stateful digital worlds an agent can act in over long horizons
Experience with ops: DevOps, MLOps, or LLMOps
Familiarity with CI/CD pipelines
Understanding of how LLMs behave: prompt engineering and spotting failure patterns to target in tasks
Experience with agentic LLM systems: tool use, function calling, agent loops
Experience authoring evaluation rubrics, benchmarks, or structured task specs, especially at scale
Prior work on AI training, evaluation, or data platforms (Mercor, Turing, Scale, Outlier, Surge, AfterQuery, or similar)
Bespoke Labs is a Mountain View based Series A AI Research/Data Curation for Agents Lab. We're working with Frontier AI Labs, and F500 Cos to advance the capabilities of AI Agents.