Research Engineer, RL Environments and Infrastructure

  • San Francisco, CA, USA
  • Full-Time
  • On-Site

Job Description:

We are looking for a hybrid Systems Engineer and AI Researcher to lead the development of our agent evaluation framework and post-training data pipelines. You will design sandboxed execution environments, high-throughput reinforcement learning feedback loops, and robust evaluation setups that handle non-deterministic agent outputs.

Core Responsibilities

  • Build Stateful Agent Environments: Design deterministically verifiable, stateful sandboxes (web, OS, API, database) where agents can execute 50+ step action trajectories safely.
  • Scale Post-Training & RL Pipelines: Implement high-throughput post-training infrastructure (RLHF, Direct Preference Optimization, Process-Supervised Reward Models) for dynamic policy optimization.
  • Design Enterprise Benchmarks: Formulate evaluation metrics and automated grading harnesses that catch agent drift, hallucination, and loops in realistic enterprise environments.
  • Systems Optimization: Keep latency low and compute efficiency high across distributed GPUs and sandboxed runtime environments.

Requirements

  • Production AI Experience: Track record of deploying evaluations, RL loops, or sandboxed agent environments into production.
  • Systems Polyglot: Deep systems background (Python, Rust, C++, Go, TypeScript, CUDA)—you pick up new tools and frameworks within days.
  • San Francisco On-Site: In-person collaboration at our SF office to iterate quickly with founders and domain experts.
  • Pragmatic Problem Solver: Comfortable navigating raw paper implementations, undocumented SDKs, and custom distributed training setups.