Papers on interpretability, alignment, control, and evaluation of robot foundation models. Up to 4 pages, CoRL template, non-archival.
Every accepted paper is presented as a poster, and the top submissions are invited to give lightning talks.
Deadline · October 1, 2026 · AoERobot foundation models (RFMs) are leaving the lab.
Vision-language-action models (VLAs) and world action models (WAMs) are increasingly capable at accepting open-ended language instructions and outputting physical actions, and billions of dollars are racing them into homes, warehouses, and streets.
RFMs inherit the transformer and diffusion substrates that AI safety research has spent years dissecting, and they may inherit the failure modes too: goal misgeneralization, jailbreaks, and specification gaming can become physical harms when the model controls a robot.
What the field lacks is a shared science of interpreting, aligning, and controlling RFMs. This workshop invites short papers (up to 4 pages, using the CoRL template) on interpretability, alignment, control, and evaluation of RFMs to help build one.
We welcome submissions related to RFM safety, with topics including but not limited to:
Given the emerging nature of this area, we explicitly welcome position pieces, rigorous negative results, replications, and open-source tooling alongside algorithmic advances.
The workshop is non-archival. All accepted papers will be presented as posters, and the top submissions will be invited to give lightning talks.
Submit on OpenReview ↗