Research Engineer, RL Environments and Infrastructure
Job Description:
We are looking for a hybrid Systems Engineer and AI Researcher to lead the development of our agent evaluation framework and post-training data pipelines. You will design sandboxed execution environments, high-throughput reinforcement learning feedback loops, and robust evaluation setups that handle non-deterministic agent outputs.
Core Responsibilities
- Build Stateful Agent Environments: Design deterministically verifiable, stateful sandboxes (web, OS, API, database) where agents can execute 50+ step action trajectories safely.
- Scale Post-Training & RL Pipelines: Implement high-throughput post-training infrastructure (RLHF, Direct Preference Optimization, Process-Supervised Reward Models) for dynamic policy optimization.
- Design Enterprise Benchmarks: Formulate evaluation metrics and automated grading harnesses that catch agent drift, hallucination, and loops in realistic enterprise environments.
- Systems Optimization: Keep latency low and compute efficiency high across distributed GPUs and sandboxed runtime environments.
Requirements
- Production AI Experience: Track record of deploying evaluations, RL loops, or sandboxed agent environments into production.
- Systems Polyglot: Deep systems background (Python, Rust, C++, Go, TypeScript, CUDA)—you pick up new tools and frameworks within days.
- San Francisco On-Site: In-person collaboration at our SF office to iterate quickly with founders and domain experts.
- Pragmatic Problem Solver: Comfortable navigating raw paper implementations, undocumented SDKs, and custom distributed training setups.