Shiqi Yang

Shiqi Yang

杨 诗琪   ·   Ph.D.
Principal Research Scientist
Director, Multimodal AI Department · SB Intuitions

I lead multimodal AI research across vision and speech while remaining hands-on in model architecture, training, experimentation, and technical implementation.

Tokyo, Japan Email LinkedIn X

At SB Intuitions, a SoftBank R&D company, I lead the Multimodal AI Department, translating multimodal research into practical applications across vision and speech. The department brings together the Creative Vision Team and Conversational Speech Team.

Before joining SB Intuitions, I was an audio-visual research scientist at Sony Group Corporation. I received my Ph.D. (2023) from the Learning and Machine Perception (LAMP) team at the Computer Vision Center, Autonomous University of Barcelona, advised by Joost van de Weijer. I also serve the community as an area chair for ICML/NeurIPS/ICLR, guest editor for an IJCV special issue, and organizer of the workshop series EVG and AVGenL. My earlier research also includes transfer learning and continual learning.

Project focus:

  • Visual and Multimodal Generation
    • Real-time and interactive autoregressive audio-video generation
    • Image generation pre-training and post-training
  • Video World Models and World Action Models
    • Spatial consistency and long horizon
    • Model Efficiency
Selected Work

Selected Projects & Tech Blog

Selected hands-on and team work (including working prototype), spanning model development, training, real-time systems, and behind-the-scenes engineering insights.

Updates

Latest News

Career

Industry Experience

  • SB Intuitions, SoftBank, Tokyo, Japan
    Apr. 2026 – Present Director, Multimodal AI Department Management role
    Jun. 2026 – Present Principal Research Scientist & Research Manager, Creative Vision Team Concurrent hands-on research and team-management role
    Apr. 2025 – May 2026 Chief Research Scientist & Research Manager, Creative Vision Team
    Dec. 2024 – Mar. 2025 Lead Research Scientist
  • Sony Group Corporation, Tokyo, Japan
    Oct. 2023 – Nov. 2024 Research Scientist
  • OMRON SINIC X, Tokyo, Japan
    Jan. 2023 – Jun. 2023 Research Intern
  • Kyoto University, Japan
    Oct. 2018 – Mar. 2019 Guest Research Associate
Research Output

Publications

The latest work is shown first. Expand the full archive to browse by year and category.
* Project lead

International Conference

  • EchoDistill: Bidirectional Concept Distillation for One-Step Diffusion Personalization Yixiong Yang, Tao Wu, Senmao Li, Shiqi Yang, Yaxing Wang, Joost van de Weijer, Kai Wang CVPR 2026 Findings. [arXiv]
  • Free-Lunch Color-Texture Disentanglement for Stylized Image Generation Jiang Qin, Senmao Li, Alexandra Gomez-Villa, Shiqi Yang, Yaxing Wang, Kai Wang, Joost van de Weijer Advances in Neural Information Processing Systems (NeurIPS), 2025. [arXiv]
  • From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face Aging Tao Liu, Dafeng Zhang, Gengchen Li, Shizhuo Liu, Yongqi Song, Senmao Li, Shiqi Yang, Boqian Li, Kai Wang, Yaxing Wang NeurIPS, 2025. [arXiv]
  • One-way ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models Senmao Li, Lei Wang, Kai Wang, Tao Liu, Jiehang Xie, Joost van de Weijer, Fahad Shahbaz Khan, Shiqi Yang, Yaxing Wang, Jian Yang IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025. [arXiv]
  • Mining Your Own Secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion Models Saurav Jha, Shiqi Yang*, Masato Ishii, Mengjie Zhao, Christian Simon, Muhammad Jehanzeb Mirza, Dong Gong, Lina Yao, Shusuke Takahashi, Yuki Mitsufuji International Conference on Learning Representations (ICLR), 2025. [arXiv] [openreview] [project]
  • One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt Tao Liu, Kai Wang, Senmao Li, Joost van de Weijer, Fahad Shahbaz Khan, Shiqi Yang, Yaxing Wang, Jian Yang, Ming-Ming Cheng ICLR, 2025. (Spotlight) [arXiv] [openreview] [project]
  • InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration Senmao Li, Kai Wang, Joost van de Weijer, Fahad Shahbaz Khan, Chun-Le Guo, Shiqi Yang, Yaxing Wang, Jian Yang, Ming-Ming Cheng ICLR, 2025. [arXiv] [openreview] [project]
  • Faster Diffusion: Rethinking the Role of UNet Encoder in Diffusion Models Senmao Li, Taihang Hu, Fahad Shahbaz Khan, Linxuan Li, Shiqi Yang, Yaxing Wang, Ming-Ming Cheng, Jian Yang Advances in Neural Information Processing Systems (NeurIPS), 2024. [project] [arXiv] [code]
  • SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond Marco Comunità, Zhi Zhong, Akira Takahashi, Shiqi Yang, Mengjie Zhao, Koichi Saito, Yukara Ikemiya, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji International Society for Music Information Retrieval (ISMIR), 2024. [arXiv]
  • Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image Editing Kai Wang, Fei Yang, Shiqi Yang, Muhammad Atif Butt, Joost van de Weijer Advances in Neural Information Processing Systems (NeurIPS), 2023. [paper] [arXiv] [code]
  • Positive Pair Distillation Considered Harmful: Continual Meta Metric Learning for Lifelong Object Re-Identification Kai Wang, Chenshen Wu, Andrew D. Bagdanov, Xialei Liu, Shiqi Yang, Shangling Jui, Joost van de Weijer British Machine Vision Conference (BMVC), 2022. [arXiv] [code]
  • Attracting and Dispersing: A Simple Approach for Source-free Domain Adaptation Shiqi Yang, Yaxing Wang, Kai Wang, Shangling Jui, Joost van de Weijer NeurIPS, 2022. (Spotlight) [project] [paper] [arXiv] [code]
  • Exploiting the Intrinsic Neighborhood Structure for Source-free Domain Adaptation Shiqi Yang, Yaxing Wang, Joost van de Weijer, Luis Herranz, Shangling Jui Advances in Neural Information Processing Systems (NeurIPS), 2021. [project] [paper] [arXiv] [code]
  • Generalized Source-free Domain Adaptation Shiqi Yang, Yaxing Wang, Joost van de Weijer, Luis Herranz, Shangling Jui International Conference on Computer Vision (ICCV), 2021. [project] [paper] [arXiv] [code] [video]
  • Parallel Convolutional Networks for Image Recognition via a Discriminator Shiqi Yang, Gang Peng Asian Conference on Computer Vision (ACCV), 2018. [paper] [arXiv]
  • Attention to Refine Through Multi Scales for Semantic Segmentation Shiqi Yang, Gang Peng Pacific-Rim Conference on Multimedia (PCM), 2018. [paper] [arXiv]

Journal

  • Training-free image inversion for one-step diffusion models Tao Wu, Senmao Li, Yaxing Wang, Shiqi Yang, Kai Wang, Joost van de Weijer Pattern Recognition, 2026. [paper]
  • GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models M. Jehanzeb Mirza, Mengjie Zhao, Zhuoyuan Mao, Sivan Doveh, Wei Lin, Paul Gavrikov, Michael Dorkenwald, Shiqi Yang, Saurav Jha, Hiromi Wakaki, Yuki Mitsufuji, Horst Possegger, Rogerio Feris, Leonid Karlinsky, James Glass Transactions on Machine Learning Research (TMLR), 2025. [arXiv]
  • Trust your Good Friends: Source-free Domain Adaptation by Reciprocal Neighborhood Clustering Shiqi Yang, Yaxing Wang, Joost van de Weijer, Luis Herranz, Shangling Jui, Jian Yang IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2023. [paper] [arXiv]
  • Casting a BAIT for Offline and Online Source-free Domain Adaptation Shiqi Yang, Yaxing Wang, Luis Herranz, Shangling Jui, Joost van de Weijer Computer Vision and Image Understanding (CVIU), 2023. [paper] [arXiv] [code]
  • On Implicit Attribute Localization for Generalized Zero-Shot Learning Shiqi Yang, Kai Wang, Luis Herranz, Joost van de Weijer IEEE Signal Processing Letters, 2021. [paper] [arXiv]

Preprint and workshop paper

  • Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling Saurav Jha, M Jehanzeb Mirza, Wei Lin, Shiqi Yang, Sarath Chandar World Modeling Workshop 2026. [arxiv]
  • OpenMU: Your Swiss Army Knife for Music Understanding Mengjie Zhao, Zhi Zhong, Zhuoyuan Mao, Shiqi Yang, Wei-Hsiang Liao, Shusuke Takahashi, Hiromi Wakaki, Yuki Mitsufuji preprint, 2024. [arXiv] [code]
  • Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation Shiqi Yang, Zhi Zhong, Mengjie Zhao, Shusuke Takahashi, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji preprint, 2024. [arXiv] [demo]
  • MaTe3D: Mask-guided Text-based 3D-aware Portrait Editing Kangneng Zhou, Daiheng Gao, Xuan Wang, Jie Zhang, Peng Zhang, Xusen Sun, Longhao Zhang, Shiqi Yang, Bang Zhang, Liefeng Bo, Yaxing Wang preprint, 2023. [arXiv]
  • A Critical Look at the Current Usage of Foundation Model for Dense Recognition Task Shiqi Yang, Atsushi Hashimoto, Yoshitaka Ushiku preprint, 2023. [arXiv]
  • OneRing: A Simple Method for Source-free Open-partial Domain Adaptation Shiqi Yang, Yaxing Wang, Kai Wang, Shangling Jui, Joost van de Weijer preprint, 2022. [project] [arXiv] [code]
Recognition

Talks, Awards & Activities

  • Organizer of the official Runway Local Meetup Tokyo, Tokyo, Japan, Jul. 2026.
  • Visiting talks at MICC Lab (Prof. Andrew Bagdanov), University of Florence, and MHUG Lab (Prof. Nicu Sebe), University of Trento, Italy, Oct. 2024.
  • Pioneer Awards 2023, CERCA Research Center of Catalonia, Spain, Dec. 2023.
  • Visiting talk with Prof. Maria Brbic's group at EPFL, Switzerland, Jan. 2023.
  • Invited talk at the TrustML Young Scientist Seminars, RIKEN AIP, Japan, Dec. 2022.
  • Participated in the ICVSS Summer School, Sicily, Italy, Jul. 2022.
  • Invited talk at the AI Time Seminar on NeurIPS 2021 (virtual), China, Feb. 2022.
Community

Academic Service

Background

Education

  • Sep. 2012 – Jun. 2016
    Bachelor in Automation, Wuhan University of Science and Technology, China.
  • Sep. 2016 – Jun. 2019
    Master in Control Science and Technology, Huazhong University of Science and Technology, China.
  • Oct. 2019 – Jul. 2023
    Ph.D. in Computer Science, Computer Vision Center, Autonomous University of Barcelona, Spain.