We present DWM, a scene-action-conditioned video diffusion model that simulates dexterous human interactions in static 3D scenes.
Hi, I'm a 3rd year MS/PhD student at Seoul National University, where I am advised by Prof. Hanbyul Joo. My current research interests lie in simulating the physical world with agents using video generative models. I received a B.S in Industrial Engineering and a B.S. in Computer Science and Engineering from Seoul National University.
I'm currently seeking research internship opportunities starting anytime.
Feel free to reach out if you're interested in collaboration!
We present DWM, a scene-action-conditioned video diffusion model that simulates dexterous human interactions in static 3D scenes.
project page code & data arXiv openreview
Our target-aware video diffusion model generates a video in which an actor accurately interacts with the target, specified with its segmentation mask.
We factor action realization and robot appearance out of the world model, presenting actions as explicitly rendered robot geometry.
We leverage VLMs to generate executable 3D manipulation plans from calibrated multi-view imagery without task-specific training.
We present an automated real-world system that collects dexterous grasp data and build a retrievable database for execution in novel scenes.
We present a framework that takes as input a single-layer clothed 3D human mesh and decomposes it into complete multi-layered 3D assets.
project page code & data arXiv
We present a framework for learning a compositional generative model of humans and objects (backpacks, coats, and more) from real-world 3D scans.