Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch
Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues
Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy
Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams
Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics
Implement methods from recent ML papers quickly and turn them into production-grade systems
Must-Have Skills
3+ years of ML engineering experience: model training, fine-tuning, or post-training pipelines in research or production
Strong Python and deep learning proficiency (PyTorch preferred; familiar with training loops, optimizers, mixed precision)
Hands-on experience with LLM post-training (SFT, RLHF, PPO, DPO, or reward model training) and understanding of how training data quality affects model behavior
Hands-on experience with long-horizon RL work: designing tasks and environments that span many steps, with persistent state across turns and multi-turn rollout design
Familiarity with RL frameworks (Gymnasium, PettingZoo, dm_env, or custom equivalents) and the ability to design or modify reward functions for agent training objectives
Practical experience with sparse-reward problems: milestone or process reward design, subgoal and task decomposition, and credit assignment across long trajectories
Experience running experiments at scale on cloud or HPC (AWS, GCP, SLURM, or Ray)
Solid understanding of evaluation methodology: held-out sets, benchmark design, rubric design for partial credit, and avoiding train/eval contamination
Good to Have
Experience building automated pipelines to curate high-quality long-horizon RL environments and tasks at scale
Familiarity with preference data collection and annotation pipelines
Experience building or operating stateful, resumable environments (snapshotting, checkpointing, branching rollouts, ideally at multi-node scale)
Experience with multi-turn or agentic LLM systems (tool use, function calling, agent loops, computer use)
Experience designing curricula that ramp task difficulty rather than presenting fixed-difficulty tasks
Familiarity with reward hacking and "fuzzy verifier" design: scoring quality beyond pass/fail correctness
Familiarity with process reward models (PRMs) or hierarchical RL literature
Prior work on RL-from-human-feedback or model-based RL at scale
Contributions to open-source training or RL-environment frameworks (trlX, OpenRLHF, verl, Prime Intellect's verifiers/prime-rl, Harbor, OpenEnv)
Track record shipping a benchmark or environment adopted outside your own team
Experience reading and implementing methods from recent ML papers quickly
Bespoke Labs is a Mountain View based Series A AI Research/Data Curation for Agents Lab. We're working with Frontier AI Labs, and F500 Cos to advance the capabilities of AI Agents.