RI PhD Thesis Defense - Bardienus Duisterhof
Price
Category
OtherDuration
3 hours
Sign up free to save events, follow performers, and get reminders.
Join FreeDate: 9/23/2026 Time: 2:00 PM - 4:00 PM Zoom: https://cmu.zoom.us/s/98474543738 Location: NSH 4305 Type: PhD Thesis Defense Who: Bardienus Pieter Duisterhof Title: Spatiotemporal World Models for Robot Manipulation Abstract: Robot manipulation aims to automate tasks that are too dull, dirty, or dangerous for humans. This future requires remarkable resource efficiency: robots must adapt to new tasks with limited data and compute while meeting stringent performance requirements. Current systems can succeed at dexterous tasks but require substantial resources to meet industrial standards. One possible explanation for this inefficiency is that frontier models learn directly from RGB images, which are typically dominated by content irrelevant to robots. Previous work has addressed this problem through task-specific feature engineering, improving efficiency by focusing on task-relevant information. Can we construct similarly focused representations that remain scalable and broadly applicable? This thesis explores spatiotemporal representations--representations of 3D geometry and motion over time--focusing on their reconstruction, generation, and robot applications. The first part of this thesis addresses unconstrained spatiotemporal reconstruction. Robots may benefit from representations that place past observations in a precise spatial and temporal context. We contribute methods that improve 3D reconstruction and calibration from arbitrary image sets and lens configurations. We show that calibrated multi-camera setups and neural rendering yield precise reconstruction in dynamic scenes, including highly deformable objects such as cloth. In the second part of this thesis, we investigate learning spatiotemporal generative priors for robot manipulation. Humans can infer plausible geometry and dynamics from a single observation. Can we instill similar priors into robots? We contribute methods that can infer depth maps, complete object geometry, and predict object dynamics. With Modality Forcing, we investigate text-to-image pre-training as a way to learn geometric priors. With PointZero, we use 3D point track completion to learn spatiotemporal priors without any robot data. The final part of this thesis considers spatiotemporal world models applied to robot manipulation. First, we show that PointZero improves performance in robot manipulation tasks, including imitation learning and action-conditioned dynamics prediction. Next, we investigate how predicting future scene states can guide action generation in world-action models (WAMs). In 3PoinTr, we show that 3D point tracks can serve as compact task plans that support transfer from human demonstrations to robot execution. In ModAR, we autoregressively denoise multiple future modalities and robot actions, achieving the best performance among the WAM formulations tested. We systematically study which modalities contribute most to manipulation performance and find that predicting RGB images provides no consistent additional benefit at the scale studied. Together, these works connect reconstruction, learned geometric and dynamics priors, and action generation to study how spatiotemporal representations can support efficient, scalable robot learning. Committee: Jeffrey Ichnowski (Chair) Deva Ramanan Shubham Tulsiani Abhishek Gupta (University of Washington)
