Skip to content

Site search

Search by research topic, paper title or author.

HOMEMultimodal AI

Multimodal AI

We research AI that links different types of information such as images, videos, language, and sound to understand the world and create new expressions. Utilizing models trained on image-language correspondences and generative models, we tackle tasks including image recognition from limited data, language-guided segmentation, and integrated understanding and generation of audio and video. Through complementary use of diverse information and effective utilization of pre-trained models, we aim to develop AI that flexibly adapts to unknown objects and situations.

Scalable Semantic Gesture Generation Utilizing Multimodal Prior Knowledge (MIRU2026)

Gestures are an important means of conveying intent and emotion while complementing speech. Previous gesture generation from speech has mainly learned movements synchronized with the rhythm of speech and did not adequately handle "semantic gestures" linked to word meanings, such as "drinking" or "pointing." Although there has been research addressing semantic gestures, all rely on manual motion recording and annotation, making large-scale data collection difficult. In this study, we propose a framework called "SeGA" that fully automates the generation of semantic gesture data by combining prior knowledge from foundational models. First, a large language model generates words involving actions and their textual descriptions, and these movements are obtained through video generation models and 3D human pose estimation. Next, phoneme alignment identifies the relevant segments in the speech, and a foundational human motion model replaces them to ensure a natural connection with preceding and following movements. Using this method, we have constructed a dataset called "SeGA-4.5k," consisting of 158 types and 4,484 clips (approximately 6.6 hours).

Zero-Shot Music-to-Image Generation Based on Common Textual Embedding Representations Using Textual Inversion (MIRU2026)

Generative technologies for images and music have advanced significantly through the development of diffusion models and flow matching. However, generating images from music—a cross-modal generation task—typically requires extensive paired data to align features. This research focuses on music and image generation models that share the same text encoder (T5-XXL) and proposes a zero-shot Music-to-Image generation method that requires no additional training. First, Textual Inversion is performed on the input music using the music generation model to obtain a “music token” representing the characteristics of the music. Next, this token is directly embedded into the prompt of the image generation model to produce images. The generated results show natural light and plants for ambient and piano music, intense red-based expressions for rock, and electronic representations for technopop. This confirms the ability to visualize abstract characteristics such as musical mood rather than specific information like instruments or song titles.

DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers (Accepted at NeurIPS 2026)

Transformers have become widely used in 3D vision, and when handling multi-view geometry, it is common practice to provide the camera's position and orientation as positional encodings in Attention mechanisms. However, when training novel view synthesis models at a large scale, a problem was identified where performance plateaus during the later stages of training. This research analyzes the cause of this issue. Because camera rotation and translation are embedded into the same dimension, they cannot be distinguished, which hinders scaling. To address this, we proposed the Decoupled Pose Positional Encoding (DPPE), which separates rotation and translation and provides positional encodings to Attention according to their respective geometric roles. DPPE enables stable training over long periods even in large-scale settings and demonstrates high generalization performance under conditions different from training, such as an increased number of viewpoints and zooming.

Overlap-aware ViT Harmonizer for Automatic Color Tone Adjustment of Different Optical Satellite Images (SSII2026)

Optical satellite images exhibit varying color tones depending on the capturing satellite and sensor, which reduces the accuracy of time-series analysis for environmental monitoring and disaster response. However, reliable pixel-to-pixel correspondence between images from different satellites exists only in overlapping areas capturing the same location. Therefore, simply applying uniform correction to the entire image cannot fully compensate for location-specific color differences. This study proposes a method that uses overlapping regions as clues to perform pixel-level color correction even in non-overlapping areas. First, a global correction estimates color transformation for the entire image from overlapping regions. Next, a Vision Transformer estimates pixel-wise color transformation parameters to correct residual local color differences. Evaluation on the Sen2VenμS dataset showed consistent improvement over global correction at seven untrained sites (PSNR +0.31 dB, MSE −7.4%).

Reference-Free Image Quality Assessment for Virtual Try-On via Human Feedback(ECCV2026)

In practical applications of virtual try-on (VTON), ground truth images are unavailable, making reference-based metrics such as SSIM and LPIPS unusable and individual generated image quality assessment challenging. This research proposes predicting quality scores consistent with human subjective evaluations without ground truth images. VTON-IQA We constructed the largest subjective evaluation dataset in this field, containing 431,800 evaluations of 62,688 try-on images generated by 14 VTON models. VTON-QBench We also introduced a model capturing the relationship between try-on images and garment/person images. Interleaved Cross-Attention As a result, we achieved a correlation with human evaluations far exceeding existing metrics and a preference judgment accuracy approaching human inter-rater agreement (joint research with ZOZO Research).

Paper: https://arxiv.org/abs/2603.13057
Code / Dataset: https://github.com/litelightlite/VTON-IQA

Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
(CVPR 2026)

Humanoid robots are expected to play active roles in assistance and service fields; however, existing motion tracking methods are limited to non-contact social interactions or single-person motions and are unsuitable for assistance scenarios requiring immediate response to a partner's posture and dynamics. This research formulates imitation of close interpersonal interactions involving force exchanges as a multi-agent reinforcement learning problem and proposes AssistMimic, a framework that simultaneously learns cooperative policies for both supporter and recipient in a physics simulator. By introducing Partner Policies Initialization, which transfers prior knowledge from single-person motion tracking policies; Dynamic Reference Retargeting, which dynamically retargets assistance motion references according to the partner's real-time posture; and Contact-Promoting Reward, which encourages physically meaningful assistance, we realized a physics-based controller capable of successfully tracking highly contact-intensive interpersonal motions for the first time, achieving state-of-the-art performance on the Inter-X and HHI-Assist benchmarks.

Project page: https://yutoshibata07.github.io/AssistMimic-projectpage/

Listening without Looking: Modality Bias in Audio-Visual Captioning(ICIP 2026)

Audio-Visual Captioning integrates audio and video to generate descriptions of entire scenes. Recent advances in multimodal fusion have improved performance, yet the actual complementarity and robustness to degradation in either modality have not been thoroughly examined. This research conducted systematic modality robustness tests by selectively suppressing or degrading audio and video streams for the state-of-the-art model LAVCap, quantitatively evaluating sensitivity and complementarity, revealing a strong bias toward audio stream inference. Additionally, we constructed a novel dataset AudioVisualCaps with text annotations describing both audio and video extending AudioCaps, demonstrating that training LAVCap on this dataset reduces modality bias compared to training on AudioCaps alone.

4D Reconstruction from Sparse Dynamic Cameras(CVPR 2026 Workshop 4DV)

Dynamic 3D (4D) reconstruction from monocular moving cameras has recently advanced significantly but remains fundamentally limited by depth ambiguity. This research focuses on a sparse dynamic camera setting where multiple independently moving cameras capture the same subject. By introducing multi-view geometric constraints while controlling shooting costs, we aim for practical 4D reconstruction suitable for real-world productions such as sports, concerts, and TV programs. Since naive extensions of existing methods cannot resolve complex spatiotemporal inconsistencies, we integrated inter-camera feature matching and intra-camera point tracking to ensure spatiotemporal consistency with a 3D track initialization method, introduced a noise-robust depth ordering regularization loss, and employed spatially and temporally diverse batch sampling strategies. Moreover, we newly developed a real video dataset, LetCamsGo, for this task, demonstrating that the proposed framework substantially improves 4D reconstruction quality in dynamic regions.

Sign-to-Speech Prosody Transfer (ICPR2026)

Sign language is an essential communication medium for individuals with hearing impairments. Recent advances in deep learning have improved translation accuracy from sign language to text, making sign language messages more accessible to non-signers. However, sign language includes prosodic features like emphasis and intonation that text alone cannot express, and existing systems fail to capture these adequately. The current mainstream two-stage pipeline (sign-to-text followed by text-to-speech) mediates only via text, resulting in substantial loss of prosodic nuances inherent in sign movements. Therefore, this research proposes a new task, "Sign-to-Speech Prosody Transfer," to directly integrate the prosodic nuances of sign language into synthesized speech. This task poses three major challenges: (1) sign language translation itself is complex, and no high-quality paired datasets of sign language and speech exist; (2) speech prosody is complex, adding difficulty to the translation process; (3) the correspondence between sign language prosody and speech prosody is subtle, and no existing methods map this directly. To address these, we propose S2PFormer (Sign-to-Prosody Transformer), which leverages sign language prosody reconstruction to enable training on unpaired datasets without requiring direct correspondence between sign and speech. Additionally, applying cross-attention to human joint data and text allows capturing fine prosodic details. Extensive experiments confirm that our approach synthesizes speech reflecting sign language prosody, opening new possibilities for more natural sign language communication.

Pre-training with Synthetic Patterns for Audio

This paper proposes a method to pretrain speech encoders using synthetic patterns as substitutes for real speech data. The proposed framework consists of two main components. The first is a Masked Autoencoder (MAE), a self-supervised learning framework that reconstructs original data from randomly masked input, focusing on low-level visual patterns and regularities within data. Thus, the content itself is not critical whether the input is images, speech mel spectrograms, or synthetic patterns. The second component is synthetic data, which unlike real speech, does not involve privacy or licensing issues. By combining MAE with synthetic patterns, we can learn generalized feature representations without relying on real data and avoid issues inherent in real speech. To validate this framework's effectiveness, extensive experiments across 13 speech tasks and 17 synthetic datasets were conducted to analyze which synthetic pattern types are effective for speech. Results show that our method achieves performance comparable to models pretrained on AudioSet-2M and partially surpasses image-based pretraining methods.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to enhance audiovisual representation learning. LG-CAV-MAE integrates a pretrained text encoder with contrastive learning of audio and video and a Masked AutoEncoder, enabling learning across audio, video, and text modalities. To train LG-CAV-MAE, we introduce a method that automatically generates triplets of audio, video, and text from unlabeled videos. First, frame-level captions are generated using an image captioning model, followed by filtering based on CLAP to ensure high consistency between audio and captions. This approach obtains high-quality audio-video-text triplets without requiring manual annotations. Evaluations of LG-CAV-MAE on audio-video retrieval and classification tasks show significant improvements over existing methods, achieving up to a 5.6% increase in recall@10 for retrieval and a 3.2% improvement in classification accuracy.

Text-guided Synthetic Geometric Augmentation for Zero-shot 3D Understanding

Sufficient amounts of training data are indispensable for zero-shot recognition models to achieve adequate generalization performance. However, collecting 3D data and captions necessary for zero-shot 3D classification is costly, posing a significant barrier. Recent generative models have greatly improved the realism of synthetic data, suggesting potential for using generated data as training data. This study addresses the question: "Can synthetic 3D data generated by generative models be used to augment limited 3D datasets?" To clarify this, we propose a synthetic 3D dataset augmentation method called Text-guided Geometric Augmentation (TeGA). TeGA is designed to align with state-of-the-art language-image-3D pretrained models excelling in zero-shot 3D classification, supplementing and expanding scarce 3D data using text-to-3D generative models. Specifically, TeGA applies a consistency filtering strategy to remove noisy samples whose text and geometric shapes do not semantically match from the synthetically generated 3D data produced according to text. Experiments doubling the size of the original datasets show performance improvements surpassing baselines, with zero-shot accuracy increases of +3.0% on Objaverse-LVIS, +4.6% on ScanObjectNN, and +8.7% on ModelNet40. These results demonstrate that TeGA effectively addresses 3D data scarcity and achieves robust zero-shot 3D classification even with limited real data.

Formula-Supervised Sound Event Detection: Pre-Training Without Real Data

In Sound Event Detection (SED) tasks, the lack of precise time-stamped labels and noise from subjective annotations have hindered learning. This study proposes Formula-SED, a dataset that enables large-scale, noise-free pretraining by synthesizing acoustic signals solely from mathematical formulas, using synthesis parameters as ground truth labels. Experiments on the DESED dataset show that pretraining with the proposed dataset is effective in improving both accuracy and convergence speed.

Rethinking Image Super-Resolution from Training Data Perspectives

Accepted to ECCV2024

In the field of image super-resolution, it has traditionally been believed that high-resolution images with minimal compression noise are essential for successful super-resolution learning. This study demonstrates, through image quality evaluation by blockiness distribution measurement and measurement of object diversity via image segmentation counts, that these factors are fundamental to super-resolution learning success. The proposed DiverSeg dataset, despite consisting of low-resolution web-collected images, achieves higher performance than existing super-resolution datasets.

Proto-Adapter: Efficient Training-Free CLIP-Adapter for Few-Shot Image Classification

Paper(https://www.mdpi.com/1424-8220/24/11/3624)

In applications where obtaining large amounts of data is difficult, image recognition through few-shot learning is required. While the large-scale vision-language model CLIP can recognize images of arbitrary classes in a zero-shot manner, there remains room for improvement in performance on downstream tasks. We propose a new method, Proto-Adapter, to adapt CLIP to downstream tasks using a small amount of training data. Our method constructs a lightweight adapter using class-specific prototype representations, enabling significant performance improvements on downstream tasks with minimal additional cost. Experiments using 11 types of image recognition benchmarks confirmed the effectiveness of the proposed method.

Construction of a large-scale dataset for image recognition targeting retail product shelves

We created a large-scale dataset for image recognition targeting retail product shelves. Due to the characteristic that many products appear simultaneously in a single image of a retail shelf, annotation costs are extremely high. Therefore, we used 3DCG to generate photorealistic images of retail product shelves and automatically annotated them based on the 3D coordinate data of objects, resulting in a large-scale dataset with 200 classes and 100,000 images.
The dataset is available for download on the linked page (use restricted to research purposes).

CG Retail Shelves Dataset – A Massive-Scale, Photorealistic, Rich Annotated CG Dataset for Retail Image Processing –

https://yukiitoh0519.github.io/CG-Retail-Shelves-Dataset

MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation

MaskDiffusion is an open-vocabulary semantic segmentation method leveraging pretrained diffusion models without requiring additional training or annotations. We demonstrated that MaskDiffusion excels at handling open-vocabulary categories including fine-grained proper noun-based categories, expanding the applications of segmentation. MaskDiffusion shows significant qualitative and quantitative improvements compared to other comparable unsupervised segmentation methods on datasets such as Potsdam (+10.5 mIoU) and COCO-Stuff (+14.8 mIoU).

Arxiv: https://arxiv.org/abs/2403.11194
Code : https://github.com/Valkyrja3607/MaskDiffusion

TAG: Guidance-free Open-Vocabulary Semantic Segmentation

We propose a novel approach called TAG to realize open-vocabulary semantic segmentation that requires no training, annotation, or guidance. TAG utilizes pretrained models like CLIP and DINO to segment images into semantically meaningful categories without additional training or dense annotations. It acquires class labels from external databases, providing flexibility to adapt to new scenarios. TAG achieves state-of-the-art results in open-vocabulary segmentation without specifying class names on PascalVOC, PascalContext, and ADE20K datasets.

Arxiv: https://arxiv.org/abs/2403.11197
Code : https://github.com/Valkyrja3607/TAG

Boosting Semantic Segmentation by Conditioning the Backbone with Semantic Boundaries

Paper

The Semantic Boundary Conditioned Backbone (SBCB) framework is an effective learning method that improves semantic segmentation performance especially at mask boundaries, while maintaining compatibility with various segmentation architectures. This framework performs multi-task learning of semantic boundary detection (SBD) using multi-scale features obtained from the backbone of segmentation architectures, resulting in features with enhanced boundary awareness. This approach achieved an average improvement of 1.2% in IoU and 2.6% in boundary F-score on the Cityscapes dataset, enhancing segmentation accuracy. The SBCB framework adapts well to a range of backbones, including vision transformer models, demonstrating its potential to advance semantic segmentation without adding complexity to the models.

Efficient Video Recognition Based on Important Patch Selection

Paper(Sensors, 2022)

Videos are computationally much more expensive to process than images, while each frame image is similar and highly redundant. This study proposes a video recognition method that reduces processing costs by identifying patches with small temporal motion or change as redundant and excluding them from input. Since the temporal motion and changes are derived as by-products during decoding of compressed videos, the additional processing cost required is minimal compared to the processing cost saved by excluding patches. Applying the proposed method to a Transformer-based video recognition model achieved over 70% reduction in processing costs with less than one point drop in accuracy for action recognition.

Peripheral Completion of 360-Degree Images for Efficient 3DCG Background Production

※Accepted to CVPR2022 : Paper, Code

360-degree images are used as background images representing the entire surroundings to efficiently produce scenes in 3DCG production. In this research, we address the problem of generating 360-degree images by inputting a single standard field-of-view image and completing its surroundings. The proposed method using Transformers obtains output images of higher resolution and more natural appearance compared to previous methods. Additionally, it can output diverse result images from a single input, providing users with many options. Thus, this research aims to support efficient and original 3DCG production for users.

Shadow Removal from Document Images by Learning with Fully Synthetic Images

※Accepted to ICIP2022 : Paper

Removing shadows from document images is an important application for improving the quality of digitized documents. Recent studies have proposed many deep learning-based shadow removal methods that learn from sets of images with and without shadows. These conventional supervised learning methods require large paired datasets of document images, which are costly to create. Therefore, in this study, we use 3DCG rendering to create large and diverse datasets without capturing real documents. Experiments show that a deep neural network trained solely on the proposed dataset performs well on real data and that pretraining with this dataset improves performance.

Introduction of Temporal Cross-Modal Attention for Event Estimation Using Audio and Visual Information in Videos

Precision Engineering Journal (2022): Paper

Previously, recognition and detection of audio events from sound and recognition and detection of events or actions from visual images were performed independently. However, videos contain both audio and visual components, and in audio-visual events where both modalities represent the same event, using both modalities is considered to improve recognition accuracy. Moreover, in audio-visual events, one modality can serve as supervisory data for the other. This enables learning using self-supervised data without requiring labeled training data. This study aims at recognizing audio-visual events. To accurately grasp the entire video by learning relationships between segments, we propose Temporal Cross-Modal Attention based on self-attention and achieved accuracy surpassing previous studies.

Fast Soft Color Segmentation

※Accepted to CVPR2020 : Arxiv , OSS

This research addresses the problem of decomposing a single image into multiple RGBA layers, each containing similar colors. Our proposed neural network-based method can perform this decomposition 300,000 times faster than existing optimization-based methods. The speed advantage enables new applications, such as changing colors in videos.

Attribute Transformation of Object Images Based on Natural Language Instructions

With the advent of image editing software, image editing has become more active. This has made simple edits, such as slight adjustments of shape or color, easier; however, natural editing of complex objects like human faces still requires advanced techniques. This research focuses on human face images and aims to transform their attributes conditioned solely on English instruction sentences. We newly establish evaluation metrics for this task of image transformation based on natural language instruction-driven attribute changes in face images and evaluate the proposed method accordingly.

Graph Convolutional Neural Networks on Superpixels for Segmentation

※IEICE : Paper

A drawback of CNN-based image segmentation is the loss of spatial information caused by downsampling in pooling layers, which reduces segmentation accuracy around object contours. As an alternative approach to prevent information loss from pooling, we propose graph convolution on superpixels. Furthermore, as an extension of graph convolution, we propose Dilated Graph Convolution to more effectively expand the receptive field. In segmentation tasks on the HKU-IS dataset, the proposed method outperformed conventional CNNs with comparable configurations.

Super pixel pooling

Saliency Map Generation for Image Classification Using CNN Classifiers

It is generally difficult to explain why a certain output is obtained when an image is input to a CNN. This study proposes a saliency map generation method by applying the framework of Generative Adversarial Networks (GANs). In this system, two neural networks learn competitively. The first network is trained to perform image classification. The second network is trained to generate images similar to a given image but that cause misclassification by the first network. To efficiently generate such images, the second network modifies the important image regions for classification significantly. Through this learning, important regions for image classification can be explicitly identified and output as saliency maps.

Saliency Map Generation for Image Classification Using GANs

Simultaneous Color Adjustment and Image Completion Using GANs

This research proposes a method that simultaneously performs natural paste compositing by color adjustment and image completion with context consideration. To explicitly require the inserted object to appear in the completion area, we utilize CNN and Generative Adversarial Network (GAN) for context-aware completion, extracting context features from the entire background image. Moreover, these context features are used not only for image completion but also for color adjustment, enabling context-aware color adjustment. In this way, we realize a network that simultaneously addresses the challenges of color adjustment and image completion.