Conference on Robot Learning (CoRL) 2026 · Half-Day Workshop · Nov 12, 2026 · Austin, Texas, USA
A community-wide effort to unify simulation benchmarking and real-world robot deployment.
Simulation has become the engine of modern embodied AI. Physics-grounded environments like BEHAVIOR provide the scale, safety, and reproducibility that physical data collection cannot, enabling agents to develop high-level reasoning, long-horizon planning, and dexterous bimanual manipulation across thousands of everyday scenarios. Rigorous diagnostic frameworks like RoboEval further deepen what we can learn from simulation by instrumenting every task with stage-level metrics that reveal not just whether a policy succeeds, but precisely how and why it fails. Yet evaluation alone is not enough: RoboChess puts simulation-trained policies to the ultimate test by requiring them to transfer directly to physical robots, perceiving a real chessboard, planning long-horizon move sequences, and executing dexterous manipulation across multiple hardware embodiments. The frontier question is therefore not whether to train in simulation, but how to ensure that the capabilities learned at scale in simulation carry over robustly to the physical world. This workshop unites benchmark designers, robot-learning researchers, foundation-model developers, and real-robot practitioners to address that challenge together.
Rich, physics-grounded environments enable agents to practice thousands of household tasks at scale — from high-level reasoning and long-range navigation to dexterous bimanual manipulation — far beyond what physical data collection alone can provide.
Task completion alone is not enough. Stage-level metrics that measure outcome, efficiency, bimanual coordination, and safety reveal precisely where and why policies fail, enabling targeted improvement rather than opaque leaderboard scores.
Many policies succeed on curated demonstrations but fail to generalize to diverse initial states and long-horizon settings. Closing this gap requires evaluation protocols that stress-test policies across a wide range of conditions in simulation before real-world deployment.
Translating simulation-trained policies to physical robots remains an open challenge. Robustness to lighting, object position, sensor noise, and embodiment differences must be systematically evaluated so that strong simulation performance predicts real-world capability.
The workshop is organized around the CoRL 2026 Embodied AI Benchmark and Sim-to-Real Challenge Suite, consisting of three complementary tracks. Each track retains its canonical technical infrastructure, evaluation protocol, leaderboard, and prizes. Across tracks, participants will report a shared set of outcome, efficiency, robustness, safety, computational-cost, and reproducibility measures, synthesized into a community white paper.
Leading researchers from academia and industry sharing perspectives across robot learning, simulation, and real-world deployment.

Half-day workshop — Nov 12, 2026, Austin, Texas. All times local (CST).
| Time | Type | Session |
|---|---|---|
| 8:50 – 9:00 | Opening | Opening Remarks Workshop organizers |
| 9:00 – 9:30 | Keynote | Yashraj Narang NVIDIA |
| 9:30 – 10:00 | Keynote | Guannan Qu Carnegie Mellon University |
| 10:00 – 10:45 | Awards |
Track 1 & 2 Challenge Awards & Spotlight — BEHAVIOR & RoboEval Challenge results announced · Challenge Winner Spotlights 1 & 2 — top-performing teams present their methods |
| 10:45 – 11:15 | Keynote | Yilun Du Harvard University |
| 11:15 – 11:45 | Keynote | Chen Tang UC Los Angeles |
| 11:45 – 12:15 | Awards + Demo |
Track 3 Challenge Awards & Spotlight & Live Demo — RoboChess RoboChess results, winner presentations & live demonstration of winning policies on physical hardware |
| 12:15 – 13:00 | Discussion | Group Discussion & Poster Marketplace Open discussion on cross-benchmark findings · interactive poster session |
We invite short papers on simulation, benchmarking, and real-world robot learning, with particular interest in work that connects scalable evaluation in simulation to meaningful robot capabilities and real-world performance.
We welcome new methods, benchmarks, challenge reports, empirical studies, negative results, and position papers that help clarify when and how simulation provides reliable evidence for real-world robot learning.
Submission portal available on OpenReview.
Submit on OpenReviewIndependent leaderboards, codebases, and prize structures — unified by a shared evaluation protocol.
Generalist embodied AI in house-scale everyday scenes
1,000 household activities across 50 scenes and 10,000 objects, powered by the Isaac Sim–based BEHAVIOR simulator. Agents must reason, navigate, and execute dexterous bimanual manipulation over long horizons.
Evaluation: Task completion rate, efficiency (total movements), 20,000 near-optimal human-demonstrated trajectories with stage-level language annotations.
Stage-level failure diagnosis for bimanual manipulation
A structured evaluation framework for bimanual robotic manipulation that augments binary success with behavioral and outcome metrics. Every task is instrumented with stage definitions to localize failure modes and test robustness to spatial variation.
Evaluation: Outcome (stage completion), Efficiency, Bimanual Coordination, Safety & Stability. 3,000+ VR-teleoperated expert demonstrations.
From simulation training to real-world robot deployment
Policies are trained in GPU-accelerated simulation and must generalize to physical robot hardware. Agents perceive a chessboard, plan long-horizon move sequences, and execute dexterous piece manipulation across multiple real-world robot embodiments.
Evaluation: Standardized manipulation metrics measuring planning accuracy, perception robustness, and execution reliability under real-world perturbations.
For details on challenge prizes, compute credits, and travel grants, please visit the individual challenge pages.
A diverse team spanning academia and industry across multiple institutions.
NVIDIA
NVIDIA
Lambda
NVIDIA

Northwestern University

Stanford University

NUS

Stanford University

Stanford University

Stanford University

Northwestern University

Stanford University

Stanford University

Stanford University

Google DeepMind

Stanford University

Stanford University

Stanford University

USC

Carnegie Mellon University

Stanford University

OpenAI

OpenAI

OpenAI

University of Washington

UW / Allen AI

Northwestern University

UW / Allen AI

University of Washington

University of Washington

University of Washington
University of Washington
University of Washington
University of Washington
University of Washington
University of Washington

University of Washington
University of Washington
Allen Institute for AI
NVIDIA
NVIDIA
Lambda
NVIDIA

UW–Madison

Adobe Research