RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers
Abstract
Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agents' broader engineering capabilities. Real-world robotics extends beyond control: agents must build, integrate, diagnose, and improve heterogeneous artifacts under resource constraints and reason from multimodal feedback. To evaluate these broader capabilities, we introduce RLE-Bench, a benchmark of robot-learning tasks spanning four representative robotics development workflows: interactive control, policy learning, perception and estimation, and mechanical design. We use diverse task-specific metrics to evaluate the artifacts submitted by the coding agents, from the success rate the agents achieved to the policy agents trained, the harness agent built, and the mechanical structures the agent designed. We aggregate these metrics into an overall RLE Index and report workflow-specific capability profiles, enabling systematic comparison of coding agents' capabilities across multiple capability dimensions. Beyond performance ranks, we also conduct in-depth case studies examining agent behavior on representative tasks, highlighting both current capabilities and limitations, and pointing to the opportunities robotics tasks have to offer for future agent training.
Community
RLE-Bench is a benchmark evaluating coding agents on general robot learning tasks, inlcuding interactive control, policy learning, perception & estimation, and mechanical design.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation (2026)
- LIBERO-Agent: Evaluating General-Purpose Agents for Direct Embodied Manipulation (2026)
- RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning (2026)
- MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation (2026)
- HarnessPAI: An Evolving Harness for Physical AI (2026)
- SimEX: Simulation-Integrated Robotics AutoResearch (2026)
- Self-Evolving Embodied Agents via Skill-Harness Evolution (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.34210 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper