Rare find

DiT4DiT. This is the official code repo for DiT4DiT, a Vision-Action-Model (VAM) framework that combines video generation model with flow-matching-based action prediction for generalizable robotic manipulation.

github.com/Mondo-Robotics/DiT4DiT

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.