Senior Software Engineer: Agentic Coding RL Environments
Frontier labs train and evaluate their coding agents on the RL environments and benchmarks you'll build. Make $200–300k+/year.
Coding agents are strong today in many ways, but frustratingly bad in others. They miss edge cases, make brittle changes, and often need close supervision from an experienced engineer.
That's where you come in. To train better agents, we need the taste, judgment, and wisdom you've built over a career shipping scaled production systems.
You’ll build coding RL environments and benchmarks that test how agents perform on realistic software engineering tasks.
That includes:
- Designing tasks based on real engineering work
- Reviewing and red-teaming model output
- Writing tests, rubrics, and success criteria
- Giving structured feedback that improves future models
We provide the framework, tools, examples, and feedback you need to get started. You’re in charge of the when, where, and how you complete projects.
Frontier labs train on the agentic coding RL environments you create.
Labs rely on expert SWEs to pressure-test models and surface failures.
The industry uses the benchmarks you build to know which models are best.
The environments you build are used to train, evaluate, and compare frontier coding models. See our research and benchmarks for examples.
You’ll also develop a deep understanding of where coding agents succeed, where they fail, and how experienced engineers can use them effectively.
Professional software engineering experience shipping real, scaled production systems.
Strong judgment about what "good code" and a "model failures" look like
Comfort working independently, and giving technical feedback
Expertise in at least one common tech stack, such as TypeScript, Python, Ruby, PHP, Java, Rust, Go, or C++
You do not need prior AI or machine-learning experience. Strong engineering judgment and attention to detail matter most.
This is a fully remote contractor role paying $100–$150+/hour. Rates scale with your experience and depth in the area. Some projects go well past the top of that range.
You choose when and where you work, use your own equipment, and communicate primarily through Slack, with occasional video calls.
You're judged on the quality of what you produce. Quality matters more than speed.
Most contributors choose to spend at least 20-40 hours per week.
This opportunity celebrates your SWE expertise, but it isn't for everyone. Those who love it really love it.
There are no sprint ceremonies, routine one-on-ones, office requirements, or feature roadmaps. Instead, you’ll focus on one goal: creating rigorous environments that teach coding agents to perform better on real software engineering work.
The role is best suited to engineers who enjoy independent work, technical review, experimentation, and high standards.
You won't have: a manager, 1:1s, workplace politics, OKR ceremonies, or Jira backlog grooming.
You will have: regular feedback on your work; full autonomy over how you expose model failures; deep fluency with what makes coding agents succeed and fail.
Application process
A couple of hours. This replaces a multi-round interview loop.
Do well, and you're on the team.