Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts
Shape environments end to end: stateful, resumable systems with snapshotting, checkpointing, and branching rollouts, at multi-node scale where needed
Generate and refine ideas for tasks and agentic trajectories across multi-turn, tool-using, and computer-use agent loops with persistent state across turns
Design reward structure for sparse-reward settings: milestone and process rewards, subgoal and task decomposition, and credit assignment across long trajectories
Validate that the work is correct, hard, and covers the right edge cases, including rubric design for partial credit, contamination avoidance, and defenses against reward hacking
Design curricula that ramp task difficulty rather than shipping fixed-difficulty tasks
Build automated pipelines that curate high-quality long-horizon RL environments and tasks at scale
Write reliable, well-tested Python infrastructure rather than one-off research scripts
Work directly with our research team with high ownership, building novel systems rather than maintaining legacy ones
Must-Have Skills
Good to Have
verifiers/prime-rl, Harbor, OpenEnv)Bespoke Labs is a Mountain View based Series A AI Research/Data Curation for Agents Lab. We're working with Frontier AI Labs, and F500 Cos to advance the capabilities of AI Agents.