Jobs › US jobs › Member of Technical Staff - Research & Post-training
San Francisco
ABOUT US Preference Model is building automated ML research engineering. Existing frontier models are brittle when applied to real-world ML tasks. The present bottleneck is the lack of high-quality RL training environments. Our first step is to build RL environments that reflect real-world complexity, with diverse tasks and robust reward functions. Our founding team has previous experience on Anthropic’s data team building data infrastructure, and datasets behind Claude.
Read the full advert and apply on the employer's site →
Collected from the employer's own careers site (Ashby). Posted 11 Sep 2026. TUNAI shows an excerpt and the facts it read from the advert; the employer's page has the full description and the application.
Does your CV match this job? Paste both and see exactly which of these skills it shows.
Check my CV against this job →