CoRL 2026 Workshop
From Contact-Rich Execution to Generalist Robot Intelligence
Half-day Workshop • Austin, Texas, USA • November 9 2026
Dexterous manipulation has made rapid progress in grasping and pick-and-place, but robots still struggle to use tools with the flexibility and physical competence of humans. Humans not only use general tools across kitchens and household work, but also wield specialized tools for surgery, sports, music, laboratories, repair, and manufacturing. These behaviors are not simply longer versions of pick-and-place: the tool becomes a physical mediator between the body and the world. Successful tool use requires sustained contact, force regulation, evolving hand–tool–object relationships, and often bimanual coordination — where one hand stabilizes, repositions, or constrains the object while the other applies the tool.
Recent progress in dexterous manipulation, imitation and reinforcement learning, tactile sensing, teleoperation, humanoid control, VLA, and world models suggests the field is ready to treat tool use as a central robot learning problem. Yet tool use exposes a fundamental conflict: practical systems require fast, physically grounded, contact-rich execution, while general-purpose systems require high-level reasoning over affordances, future states, and generalization across tools, objects, tasks, and embodiments. Existing systems rarely satisfy both requirements at once.
"How can robot learning bridge physically reliable, contact-rich tool execution with high-level generalizable reasoning?"
Tool use requires representing the evolving relationship between hands, fingers, tools, target objects, contacts, and forces. Many useful tasks further require one hand to stabilize, guide, regrasp, or constrain while the other manipulates the tool. What representations make this structure learnable and transferable, and how should robots learn role assignment, hand switching, synchronization, and recovery in long-horizon bimanual tool use?
How can robots execute tool-use skills at real-time speed while maintaining stable contact under uncertainty in force, friction, compliance, deformation, and object motion? Which parts of tool use can be learned directly from data, and which parts still require explicit planning, feedback control, tactile/force sensing, or physics-based constraints to achieve stable real-world execution?
Vision alone is often insufficient for tool use because contacts may be occluded, force thresholds may be subtle, and failure can depend on slip, pressure or deformation. What sensing and hardware capabilities are needed for robust tool-mediated interaction (e.g., vision, touch, force, proprioception, audio, tactile-enabled hands, and teleoperation interfaces)? How should these signals be fused and exposed to policies, controllers, and VLA/world-model systems at the latency and reliability required for real-time tool use?
How can robots learn tool use from human videos, motion capture, teleoperation, tactile demonstrations, simulation, and robot self-practice despite differences in morphology, sensing, compliance, and control frequency? What benchmarks and metrics should evaluate not only task success, but also contact stability, force regulation, robustness, generalization, and real-world transfer?
How can world models, video prediction, VLA policies, and language-conditioned planners infer tool affordances, predict future physical effects, and generalize across unseen tools, objects, and tasks? Where do these high-level models fail when success depends on precise contact rather than semantic understanding alone? How can their predictions or planning be fast, physically-grounded, and actionable enough under real-time contact uncertainty?
A challenge-driven, half-day format organized around active discussion: invited challenge talks, contributed spotlights, posters & demos, structured problem-solving, and a fishbowl-style synthesis panel.
| Time | Event |
|---|---|
| 2:00 PM | Welcome & Framing: bridging contact-rich execution and generalizable reasoning |
| 2:15 PM | Invited Challenge Talk 1 |
| 2:45 PM | Invited Challenge Talk 2 |
| 3:15 PM | Contributed Spotlights |
| 3:30 PM | Poster & Demo Session & Coffee Break |
| 4:00 PM | Invited Challenge Talk 3 |
| 4:30 PM | Invited Challenge Talk 4 |
| 5:00 PM | Structured Problem-Solving / Breakout Discussion |
| 5:30 PM | Fishbowl Synthesis Panel |
| 6:00 PM | Workshop Ends |
We invite non-archival submissions on tool-mediated physical interaction, including bimanual and single-hand tool use across general, surgical, sports, musical, and household settings; tactile and contact-rich policies; human-to-robot transfer; robot world models; VLA policies; simulation; hardware; and evaluation — as long as the work addresses tool-mediated physical interaction.
Submissions are accepted in three formats: full workshop papers (up to 8 pages), short papers or extended abstracts (up to 4 pages), and demo or position abstracts (1–2 pages), excluding references and appendices. We welcome mature works as well as early-stage ideas, negative results, benchmark proposals, hardware platforms, teleoperation pipelines, and real-world failure cases. Accepted contributions are presented as posters, with a subset selected for oral spotlights and demos.