Be the first to hear about new sota jobs + exclusive salary research + career cheatsheets.
Anthropic · San Francisco/New York City/Seattle · Hybrid
Staff+ Research Scientist, Multi-Agent
1/6/2025
Description
Multi-Agent systems are becoming an increasingly important part of how AI is deployed, whether via fast small-model subagents inside a product, or large groups of agents solving very large problems. Training Claude to be maximally effective and safe within large groups is a challenging new area of reinforcement learning, and represents a new axis for scaling test time compute.
We are looking for researchers who have experience training multi-agent systems at the largest scale and an appreciation for the incentives and mechanism design that come into play.
Responsibilities:
Help create and optimize environments and data for model training that maximize Claude’s performance or ease of use on agentic tasks
Ideate, develop, and compare the performance of different agent harness configurations (eg memory, context management, communication architectures for agents)
Design and implement rigorous quantitative benchmarks for large scale agentic tasks
Work with our product org to find solutions to our most vexing challenges in applying agents to our products
Qualifications
Have experience with large-scale RL on language models
Have experience training multi-agent systems
Enjoy going deeply into the roots of a problem and understanding its foundations, rather than its surface.
Have good communication skills and an interest in working with other researchers on difficult tasks
Have a passion for making powerful technology safe and societally beneficial
Are excited for a mission-driven org with fast-paced, impactful work
Design and build reinforcement learning environments to train groups of Claudes how to solve problems together efficiently
Design and build agent affordances that unlock new capabilities and scales of agents, while keeping the Bitter Lesson in mind
Design and build a novel eval that measures how large teams of agents interact in groups to solve problems