Multimodal Agents
Autonomous agents that perceive across language, vision, video, and audio, grounding reasoning, planning, and tool use in real digital environments.
SNU PI Lab investigates multimodal agentic AI and world foundation models.
We study how agents can perceive, reason, and act reliably across digital and physical systems.
Autonomous agents that perceive across language, vision, video, and audio, grounding reasoning, planning, and tool use in real digital environments.
Generative models that learn how the world evolves under actions, providing the basis for prediction, planning, and embodied decision making.
World action models and vision-language-action learning that couple prediction with control, enabling agents to act reliably on system.
Quantization, distillation, and efficient multimodal architectures that bring agents to edge devices and robots, enabling real-time intelligence beyond the cloud.
Safety evaluation and alignment of foundation models, toward AI that stays reliable and beneficial across diverse users and languages.
Advancing research through government-funded projects and industry collaborations.











