Long-Horizon Coding Task Expert

Learn more about Bespoke Labs
Bespoke Labs

Bespoke Labs

Long-Horizon Coding Task Expert

Remote
Full Time
Paid
  • Responsibilities
    • Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts

    • Shape environments end to end: stateful, resumable systems with snapshotting, checkpointing, and branching rollouts, at multi-node scale where needed

    • Generate and refine ideas for tasks and agentic trajectories across multi-turn, tool-using, and computer-use agent loops with persistent state across turns

    • Design reward structure for sparse-reward settings: milestone and process rewards, subgoal and task decomposition, and credit assignment across long trajectories

    • Validate that the work is correct, hard, and covers the right edge cases, including rubric design for partial credit, contamination avoidance, and defenses against reward hacking

    • Design curricula that ramp task difficulty rather than shipping fixed-difficulty tasks

    • Build automated pipelines that curate high-quality long-horizon RL environments and tasks at scale

    • Write reliable, well-tested Python infrastructure rather than one-off research scripts

    • Work directly with our research team with high ownership, building novel systems rather than maintaining legacy ones

  • Desired skills

    Must-Have Skills

    • Hands-on experience of Long horizon coding RL work
    • Experience building automated pipelines to curate high-quality long-horizon RL environments and tasks at scale
    • Practical experience with sparse-reward problems: milestone/process reward design, subgoal or task decomposition, credit assignment across long trajectories, with familiarity in process reward models (PRMs) or hierarchical RL literature
    • Experience building or operating stateful, resumable environments (snapshotting, checkpointing, branching rollouts, ideally at multi-node scale)
    • Strong Python; comfortable building reliable, well-tested infrastructure, not just research scripts
    • Familiarity with RL environment frameworks (Gymnasium, PettingZoo, or custom equivalents), multi-turn rollout design, and experience with agentic, tool-using, or computer-use LLM systems (persistent state across turns)
    • Understanding of evaluation methodology: held-out sets, contamination avoidance, rubric design for partial credit
    • Experience designing curricula that ramp task difficulty rather than presenting fixed-difficulty tasks

    Good to Have

    • Experience building coding agents, agentic products, or RL infrastructure at a company shipping real coding/agent tools, not only research settings
    • Familiarity with reward hacking and "fuzzy verifier" design, scoring quality beyond pass/fail correctness
    • Contributions to open RL-environment infrastructure (Prime Intellect's verifiers/prime-rl, Harbor, OpenEnv)
    • Track record shipping a benchmark or environment adopted outside your own team
    • Publications or public writeups on long-horizon RL, credit assignment, or reward shaping
  • Industry
    Information Technology and Services
  • About Us

    Bespoke Labs is a Mountain View based Series A AI Research/Data Curation for Agents Lab. We're working with Frontier AI Labs, and F500 Cos to advance the capabilities of AI Agents.