Shiqi Yang
Director, Multimodal AI Department · SB Intuitions
I lead multimodal AI research across vision and speech while remaining hands-on in model architecture, training, experimentation, and technical implementation.
At SB Intuitions, a SoftBank R&D company, I lead the Multimodal AI Department, translating multimodal research into practical applications across vision and speech. The department brings together the Creative Vision Team and Conversational Speech Team.
Before joining SB Intuitions, I was an audio-visual research scientist at Sony Group Corporation. I received my Ph.D. (2023) from the Learning and Machine Perception (LAMP) team at the Computer Vision Center, Autonomous University of Barcelona, advised by Joost van de Weijer. I also serve the community as an area chair for ICML/NeurIPS/ICLR, guest editor for an IJCV special issue, and organizer of the workshop series EVG and AVGenL. My earlier research also includes transfer learning and continual learning.
Project focus:
-
Visual and Multimodal Generation
- Real-time and interactive autoregressive audio-video generation
- Image generation pre-training and post-training
-
Video World Models and World Action Models
- Spatial consistency and long horizon
- Model Efficiency
Selected Projects & Tech Blog
Selected hands-on and team work (including working prototype), spanning model development, training, real-time systems, and behind-the-scenes engineering insights.
Latest News
- Jul. 2026 We are organizing two workshops at ECCV 2026: the 1st workshop on Efficient Visual Generation (EVG) and the 3rd workshop on Audio-Visual Generation and Learning (AVGenL). Across the two workshops, we will feature industrial demos from Cantina, Google DeepMind, Tavus, Runway, Kyutai, Black Forest Labs, and Reactor.
- Jun. 2026 We are organizing the official Runway Local Meetup Tokyo on July 16. Please see the official Runway meetup page and register here.
- Sep. 2025 2 papers are accepted by NeurIPS 2025.
- May 2025 We will host the 2nd workshop on Audio-Visual Generation and Learning (AVGenL) in ICCV 2025. We will have 1 industrial session this year: Veo 3 from Google DeepMind. Stay tuned for more details.
- Feb. 2025 "One-way ticket" is accepted by CVPR 2025.
- Jan. 2025 "Mine Your Own Secrets", "InterLCM" and "One-Prompt-One-Story" (spotlight) are accepted by ICLR 2025.
- Oct. 2024 Have visiting talks in MICC Lab (Prof. Andrew Bagdanov) in University of Florence and MHUG Lab (Prof. Nicu Sebe) in University of Trento.
- Sep. 2024 Our paper "Faster Diffusion: Rethinking the Role of UNet Encoder in Diffusion Models" is accepted by NeurIPS 2024.
- Apr. 2024 We are organizing an ECCV 2024 workshop "AVGenL: Audio-Visual Generation and Learning". Please check the site for CfP and speakers.
- Dec. 2023 My doctoral thesis received "Pioneer Awards 2023 - CERCA".
- Sep. 2023 Our paper "Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image Editing" is accepted by NeurIPS 2023.
- Aug. 2023 Extended version of "NRC" is accepted by IEEE TPAMI.
- Jun. 2023 "Casting a BAIT for Offline and Online Source-free Domain Adaptation" is finally accepted by CVIU.
- Jan. 2023 Have a visiting talk in Prof. Maria Brbic's group in EPFL.
- Nov. 2022 I present our work on model adaptation under domain and category shift on TrustML Young Scientist Seminars (hosted by RIKEN AIP) on Dec. 7.
- Sep. 2022 "Attracting and Dispersing: A Simple Approach for Source-free Domain Adaptation" is accepted by NeurIPS 2022 as Spotlight, and our paper "Positive Pair Distillation Considered Harmful: Continual Meta Metric Learning for Lifelong Object Re-Identification" is accepted by BMVC 2022.
- Sep. 2021 "Exploiting the Intrinsic Neighborhood Structure for Source-free Domain Adaptation" is accepted by NeurIPS 2021.
- Jul. 2021 "Generalized Source-free Domain Adaptation" is accepted by ICCV 2021.
Industry Experience
-
SB Intuitions, SoftBank, Tokyo, JapanApr. 2026 – Present Director, Multimodal AI Department Management roleJun. 2026 – Present Principal Research Scientist & Research Manager, Creative Vision Team Concurrent hands-on research and team-management roleApr. 2025 – May 2026 Chief Research Scientist & Research Manager, Creative Vision TeamDec. 2024 – Mar. 2025 Lead Research Scientist
-
Sony Group Corporation, Tokyo, JapanOct. 2023 – Nov. 2024 Research Scientist
-
OMRON SINIC X, Tokyo, JapanJan. 2023 – Jun. 2023 Research Intern
-
Kyoto University, JapanOct. 2018 – Mar. 2019 Guest Research Associate
Publications
Latest Publications
International Conference
- EchoDistill: Bidirectional Concept Distillation for One-Step Diffusion Personalization CVPR 2026 Findings. [arXiv]
- Free-Lunch Color-Texture Disentanglement for Stylized Image Generation Advances in Neural Information Processing Systems (NeurIPS), 2025. [arXiv]
- From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face Aging NeurIPS, 2025. [arXiv]
- One-way ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025. [arXiv]
- Mining Your Own Secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion Models International Conference on Learning Representations (ICLR), 2025. [arXiv] [openreview] [project]
- One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt ICLR, 2025. (Spotlight) [arXiv] [openreview] [project]
- InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration ICLR, 2025. [arXiv] [openreview] [project]
- Faster Diffusion: Rethinking the Role of UNet Encoder in Diffusion Models Advances in Neural Information Processing Systems (NeurIPS), 2024. [project] [arXiv] [code]
- SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond International Society for Music Information Retrieval (ISMIR), 2024. [arXiv]
- Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image Editing Advances in Neural Information Processing Systems (NeurIPS), 2023. [paper] [arXiv] [code]
- Positive Pair Distillation Considered Harmful: Continual Meta Metric Learning for Lifelong Object Re-Identification British Machine Vision Conference (BMVC), 2022. [arXiv] [code]
- Attracting and Dispersing: A Simple Approach for Source-free Domain Adaptation NeurIPS, 2022. (Spotlight) [project] [paper] [arXiv] [code]
- Exploiting the Intrinsic Neighborhood Structure for Source-free Domain Adaptation Advances in Neural Information Processing Systems (NeurIPS), 2021. [project] [paper] [arXiv] [code]
- Generalized Source-free Domain Adaptation International Conference on Computer Vision (ICCV), 2021. [project] [paper] [arXiv] [code] [video]
- Parallel Convolutional Networks for Image Recognition via a Discriminator Asian Conference on Computer Vision (ACCV), 2018. [paper] [arXiv]
- Attention to Refine Through Multi Scales for Semantic Segmentation Pacific-Rim Conference on Multimedia (PCM), 2018. [paper] [arXiv]
Journal
- Training-free image inversion for one-step diffusion models Pattern Recognition, 2026. [paper]
- GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models Transactions on Machine Learning Research (TMLR), 2025. [arXiv]
- Trust your Good Friends: Source-free Domain Adaptation by Reciprocal Neighborhood Clustering IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2023. [paper] [arXiv]
- Casting a BAIT for Offline and Online Source-free Domain Adaptation Computer Vision and Image Understanding (CVIU), 2023. [paper] [arXiv] [code]
- On Implicit Attribute Localization for Generalized Zero-Shot Learning IEEE Signal Processing Letters, 2021. [paper] [arXiv]
Preprint and workshop paper
- Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling World Modeling Workshop 2026. [arxiv]
- OpenMU: Your Swiss Army Knife for Music Understanding preprint, 2024. [arXiv] [code]
- Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation preprint, 2024. [arXiv] [demo]
- MaTe3D: Mask-guided Text-based 3D-aware Portrait Editing preprint, 2023. [arXiv]
- A Critical Look at the Current Usage of Foundation Model for Dense Recognition Task preprint, 2023. [arXiv]
- OneRing: A Simple Method for Source-free Open-partial Domain Adaptation preprint, 2022. [project] [arXiv] [code]
Talks, Awards & Activities
- Organizer of the official Runway Local Meetup Tokyo, Tokyo, Japan, Jul. 2026.
- Visiting talks at MICC Lab (Prof. Andrew Bagdanov), University of Florence, and MHUG Lab (Prof. Nicu Sebe), University of Trento, Italy, Oct. 2024.
- Pioneer Awards 2023, CERCA Research Center of Catalonia, Spain, Dec. 2023.
- Visiting talk with Prof. Maria Brbic's group at EPFL, Switzerland, Jan. 2023.
- Invited talk at the TrustML Young Scientist Seminars, RIKEN AIP, Japan, Dec. 2022.
- Participated in the ICVSS Summer School, Sicily, Italy, Jul. 2022.
- Invited talk at the AI Time Seminar on NeurIPS 2021 (virtual), China, Feb. 2022.
Academic Service
- Conference Area Chair: ICML (2026); NeurIPS (2026); ICLR.
- Guest Editor: IJCV Special Issue "Audio-Visual Generation".
- Workshop Organizer: ECCV 2026 "Workshop on Efficient Visual Generation (EVG)"; ECCV 2024 / ICCV 2025 / ECCV 2026 "Audio-Visual Generation and Learning (AVGenL) Workshop".
- Conference Reviewer: ICLR; ICCV; NeurIPS; ECCV; ICML; CVPR; WACV.
- Journal Reviewer: IEEE TKDE; TPAMI; TAI; IJCV.
Education
-
Sep. 2012 – Jun. 2016
Bachelor in Automation, Wuhan University of Science and Technology, China. -
Sep. 2016 – Jun. 2019
Master in Control Science and Technology, Huazhong University of Science and Technology, China. -
Oct. 2019 – Jul. 2023
Ph.D. in Computer Science, Computer Vision Center, Autonomous University of Barcelona, Spain.