SPAIS 2026 Submit ↗
CoRL 2026 · November 12 · Austin, Texas

Call for
Papers

Papers on interpretability, alignment, control, and evaluation of robot foundation models. Up to 4 pages, CoRL template, non-archival.

Every accepted paper is presented as a poster, and the top submissions are invited to give lightning talks.

Deadline · October 1, 2026 · AoE
Submissions Open
Submit on
OpenReview
DeadlineOctober 1, 2026 · AoE
ReviewDouble-blind · October 2–16
NotificationOctober 17, 2026
WorkshopNovember 12, 2026
LengthUp to 4 pages
ArchivalNon-archival
Submit Here ↗
SPAIS 2026 submission portal
The Call

A shared science of RFM safety

Robot foundation models (RFMs) are leaving the lab.

Vision-language-action models (VLAs) and world action models (WAMs) are increasingly capable at accepting open-ended language instructions and outputting physical actions, and billions of dollars are racing them into homes, warehouses, and streets.

RFMs inherit the transformer and diffusion substrates that AI safety research has spent years dissecting, and they may inherit the failure modes too: goal misgeneralization, jailbreaks, and specification gaming can become physical harms when the model controls a robot.

What the field lacks is a shared science of interpreting, aligning, and controlling RFMs. This workshop invites short papers (up to 4 pages, using the CoRL template) on interpretability, alignment, control, and evaluation of RFMs to help build one.

Scope

Topics of interest

We welcome submissions related to RFM safety, with topics including but not limited to:

Area 01
Interpretability
  • Probing, sparse autoencoders, activation patching, and steering applied to RFMs
  • Localizing safety-relevant concepts in model internals (e.g. human proximity, fragile objects, force limits, task boundaries)
  • What changes when interpretability moves from text to continuous, closed-loop action, and how to ground internal findings in physical outcomes
Area 02
Alignment & Robustness
  • Specifying and instilling safe behavior: robot constitutions, RL and supervised fine-tuning, guardrails, and whether LLM safety training survives the transfer to embodiment
  • Jailbreaks, prompt injection, and adversarial attacks across the embodied attack surface (observations, instructions, and the physical environment), plus defenses against them
Area 03
Control
  • Runtime monitoring, anomaly detection, and competence estimation for deployed policies
  • Extending classical guarantees (reachability analysis, safety filters, constrained control) to learned policies and latent spaces
  • Oversight of embodied agents: how monitoring, sandboxing, and shutdown change when the policy acts in the physical world
  • Understanding and improving sim-to-real transfer of safety properties
  • Security of deployed robot fleets
Area 04
Evaluation
  • Benchmarks, datasets, red-teaming, and hardware-in-the-loop protocols for RFM safety beyond collision avoidance
  • Evaluation methods for open-ended, general-purpose physical systems
Explicitly Welcome

Given the emerging nature of this area, we explicitly welcome position pieces, rigorous negative results, replications, and open-source tooling alongside algorithmic advances.

Submission

Science of
Physical AI Safety

The workshop is non-archival. All accepted papers will be presented as posters, and the top submissions will be invited to give lightning talks.

Submit on OpenReview ↗
Up to 4 pages, CoRL template. Submissions close October 1, 2026 — Anywhere on Earth.