{"version":"0.1","company":{"name":"YubHub","url":"https://yubhub.co","jobsUrl":"https://yubhub.co/jobs/skill/reward-robustness-evaluation"},"x-facet":{"type":"skill","slug":"reward-robustness-evaluation","display":"Reward Robustness Evaluation","count":1},"x-feed-size-limit":100,"x-feed-sort":"enriched_at desc","x-feed-notice":"This feed contains at most 100 jobs (the most recently enriched). For the full corpus, use the paginated /stats/by-facet endpoint or /search.","x-generator":"yubhub-xml-generator","x-rights":"Free to redistribute with attribution: \"Data by YubHub (https://yubhub.co)\"","x-schema":"Each entry in `jobs` follows https://schema.org/JobPosting. YubHub-native raw fields carry `x-` prefix.","jobs":[{"@context":"https://schema.org","@type":"JobPosting","identifier":{"@type":"PropertyValue","name":"YubHub","value":"job_64176983-af0"},"title":"Research Engineer, Reward Models Platform","description":"<p>You will work as a Research Engineer on Anthropic&#39;s Reward Models Platform. Your primary responsibility will be to design and build infrastructure that enables researchers to rapidly iterate on reward signals. This includes tools for rubric development, human feedback data analysis, and reward robustness evaluation. You will also develop systems for automated quality assessment of rewards, including detection of reward hacks and other pathologies. Additionally, you will create tooling that allows researchers to easily compare different reward methodologies and understand their effects. You will collaborate with researchers to translate science requirements into platform capabilities and optimize existing systems for performance, reliability, and ease of use.</p>\n<p>You will have the opportunity to contribute directly to research projects yourself and have a direct impact on our ability to scale reward development across domains. You will work closely with researchers and translate ambiguous requirements into well-scoped engineering projects.</p>\n<p>To be successful in this role, you should have prior research experience and be excited to work closely with researchers. You should have strong Python skills and experience with ML workflows and data pipelines, and building related infrastructure/tooling/platforms. You should be comfortable working across the stack, ranging from data pipelines to experiment tracking to user-facing tooling.</p>\n<p>Strong candidates may also have experience with ML research, building internal tooling and platforms for ML researchers, data quality assessment and pipeline optimization, experiment tracking, evaluation frameworks, or MLOps tooling. They may also have experience with large-scale data processing, Kubernetes, distributed systems, or cloud infrastructure, and familiarity with reinforcement learning or fine-tuning workflows.</p>\n<p style=\"margin-top:24px;font-size:13px;color:#666;\">XML job scraping automation by <a href=\"https://yubhub.co\">YubHub</a></p>","url":"https://yubhub.co/jobs/job_64176983-af0","directApply":true,"hiringOrganization":{"@type":"Organization","name":"Anthropic","sameAs":"https://www.anthropic.com/","logo":"https://logos.yubhub.co/anthropic.com.png"},"x-apply-url":"https://job-boards.greenhouse.io/anthropic/jobs/5024831008","x-work-arrangement":"hybrid","x-experience-level":"mid","x-job-type":"full-time","x-salary-range":"$350,000-$500,000 USD","x-skills-required":["Python","ML workflows","data pipelines","infrastructure/tooling/platforms","rubric development","human feedback data analysis","reward robustness evaluation","automated quality assessment","reward hacks","pathologies","experiment tracking","evaluation frameworks","MLOps tooling"],"x-skills-preferred":["ML research","building internal tooling and platforms for ML researchers","data quality assessment and pipeline optimization","Kubernetes","distributed systems","cloud infrastructure","reinforcement learning","fine-tuning workflows"],"datePosted":"2026-04-18T15:42:43.065Z","jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY"}},"jobLocationType":"TELECOMMUTE","employmentType":"FULL_TIME","occupationalCategory":"Engineering","industry":"Technology","skills":"Python, ML workflows, data pipelines, infrastructure/tooling/platforms, rubric development, human feedback data analysis, reward robustness evaluation, automated quality assessment, reward hacks, pathologies, experiment tracking, evaluation frameworks, MLOps tooling, ML research, building internal tooling and platforms for ML researchers, data quality assessment and pipeline optimization, Kubernetes, distributed systems, cloud infrastructure, reinforcement learning, fine-tuning workflows","baseSalary":{"@type":"MonetaryAmount","currency":"USD","value":{"@type":"QuantitativeValue","minValue":350000,"maxValue":500000,"unitText":"YEAR"}}}]}