Recent Research Achievements
Phase Detection of Golf Swings Using Combined Action Spotting and Action Segmentation
In golf form analysis, accurately extracting specific moments (phases) during the swing, such as address and impact, from video is crucial. Conventional methods have treated this task as "action spotting," detecting each phase instant directly. However, these methods have limited ability to capture temporal context and may mistakenly confuse similar postures, such as interpreting post-impact poses as backswing frames. This study proposes a method that combines "action segmentation," which estimates action intervals and their order, with focused spotting near phase transitions. Our approach outperformed existing methods on the public benchmark GolfDB and, on an original high-frame-rate dataset shot at 120/240 fps, improved accuracy from 62.5% to 73.7% compared to the most accurate existing method, an increase of 11.2 percentage points.

Surgical Instrument Detection in Open Surgery Using Feature Fusion Based on Hand and Instrument Interaction (MIRU2026)
The technology for detecting surgical instruments within surgical videos serves as a foundation for quantitatively understanding the progress of surgery and the surgeon's techniques. Videos of open surgery, recorded from the surgeon's first-person perspective, often have instruments frequently occluded by the hand, which made conventional image-based instrument detection methods insufficiently accurate. This research focuses on the fact that surgeons hold instruments differently depending on the instrument type and proposes an instrument detection method that uses hand information as auxiliary data. First, surgical instruments and hands are simultaneously detected with a single detector, and the hands holding the instruments are associated based on overlapping detection boxes. Then, hand features are incorporated to refine the instrument class predictions. Evaluation on the first-person surgical video dataset EgoSurgery-tool showed an improvement in mean average precision (mAP) for instrument detection from 0.646 to 0.676 compared to the baseline (Deformable DETR), reducing misrecognition between instruments of similar shapes and false detections of the background.

Scalable Semantic Gesture Generation Utilizing Multimodal Prior Knowledge (MIRU2026)
Gestures are an important means of conveying intent and emotion while complementing speech. Previous gesture generation from speech has mainly learned movements synchronized with the rhythm of speech and did not adequately handle "semantic gestures" linked to word meanings, such as "drinking" or "pointing." Although there has been research addressing semantic gestures, all rely on manual motion recording and annotation, making large-scale data collection difficult. In this study, we propose a framework called "SeGA" that fully automates the generation of semantic gesture data by combining prior knowledge from foundational models. First, a large language model generates words involving actions and their textual descriptions, and these movements are obtained through video generation models and 3D human pose estimation. Next, phoneme alignment identifies the relevant segments in the speech, and a foundational human motion model replaces them to ensure a natural connection with preceding and following movements. Using this method, we have constructed a dataset called "SeGA-4.5k," consisting of 158 types and 4,484 clips (approximately 6.6 hours).

Zero-Shot Music-to-Image Generation Based on Common Textual Embedding Representations Using Textual Inversion (MIRU2026)
Generative technologies for images and music have advanced significantly through the development of diffusion models and flow matching. However, generating images from music—a cross-modal generation task—typically requires extensive paired data to align features. This research focuses on music and image generation models that share the same text encoder (T5-XXL) and proposes a zero-shot Music-to-Image generation method that requires no additional training. First, Textual Inversion is performed on the input music using the music generation model to obtain a “music token” representing the characteristics of the music. Next, this token is directly embedded into the prompt of the image generation model to produce images. The generated results show natural light and plants for ambient and piano music, intense red-based expressions for rock, and electronic representations for technopop. This confirms the ability to visualize abstract characteristics such as musical mood rather than specific information like instruments or song titles.

Improving Inference Efficiency in Dense Event Camera Tasks (MIRU2026)
Event cameras are sensors that asynchronously output per-pixel brightness changes. Due to their features of low latency, high dynamic range, and low power consumption, they are expected to be applied in autonomous driving and robotics. Dense tasks that predict at the pixel level, such as semantic segmentation, incur high computational costs. Nevertheless, conventional methods perform inference at every time step regardless of scene changes. This research proposes a method that dynamically switches inference based on scene changes to improve processing efficiency. First, a decision network determines the necessity of inference based on changes in features obtained by a lightweight encoder. When inference is skipped, past prediction results are transformed using optical flow and reused. The computation of this skipping process requires approximately one-seventh of the cost of normal inference. Evaluation on a driving scene dataset showed that, with the same computation, our method achieved higher segmentation accuracy compared to a method that randomly switches between inference and warping.

DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers (Accepted at NeurIPS 2026)
Transformers have become widely used in 3D vision, and when handling multi-view geometry, it is common practice to provide the camera's position and orientation as positional encodings in Attention mechanisms. However, when training novel view synthesis models at a large scale, a problem was identified where performance plateaus during the later stages of training. This research analyzes the cause of this issue. Because camera rotation and translation are embedded into the same dimension, they cannot be distinguished, which hinders scaling. To address this, we proposed the Decoupled Pose Positional Encoding (DPPE), which separates rotation and translation and provides positional encodings to Attention according to their respective geometric roles. DPPE enables stable training over long periods even in large-scale settings and demonstrates high generalization performance under conditions different from training, such as an increased number of viewpoints and zooming.

Overlap-aware ViT Harmonizer for Automatic Color Tone Adjustment of Different Optical Satellite Images (SSII2026)
Optical satellite images exhibit varying color tones depending on the capturing satellite and sensor, which reduces the accuracy of time-series analysis for environmental monitoring and disaster response. However, reliable pixel-to-pixel correspondence between images from different satellites exists only in overlapping areas capturing the same location. Therefore, simply applying uniform correction to the entire image cannot fully compensate for location-specific color differences. This study proposes a method that uses overlapping regions as clues to perform pixel-level color correction even in non-overlapping areas. First, a global correction estimates color transformation for the entire image from overlapping regions. Next, a Vision Transformer estimates pixel-wise color transformation parameters to correct residual local color differences. Evaluation on the Sen2VenμS dataset showed consistent improvement over global correction at seven untrained sites (PSNR +0.31 dB, MSE −7.4%).

Event Representation Learning Considering Communication Bandwidth (MIRU 2026)
Event cameras are sensors that asynchronously output pixel-wise luminance changes, offering high temporal resolution and a wide dynamic range. However, their data rate can reach tens of Gbps, making direct handling on edge devices infeasible. This research proposes a method that assumes placing an FPGA between the camera and the edge device to reduce and re-encode events according to communication bandwidth constraints. By voxelizing event streams and using four types of trainable kernels and thresholds that capture different features, the method re-represents events while preserving information necessary for downstream tasks. Evaluation on the automotive dataset DSEC showed that communication bandwidth could be reduced by about 40% while maintaining image reconstruction performance equivalent to raw events. Furthermore, performance exceeding that of raw events was achieved at the same transmission rate.

Reference-Free Image Quality Assessment for Virtual Try-On via Human Feedback(ECCV2026)
In practical applications of virtual try-on (VTON), ground truth images are unavailable, making reference-based metrics such as SSIM and LPIPS unusable and individual generated image quality assessment challenging. This research proposes predicting quality scores consistent with human subjective evaluations without ground truth images. VTON-IQA We constructed the largest subjective evaluation dataset in this field, containing 431,800 evaluations of 62,688 try-on images generated by 14 VTON models. VTON-QBench We also introduced a model capturing the relationship between try-on images and garment/person images. Interleaved Cross-Attention As a result, we achieved a correlation with human evaluations far exceeding existing metrics and a preference judgment accuracy approaching human inter-rater agreement (joint research with ZOZO Research).
Paper: https://arxiv.org/abs/2603.13057
Code / Dataset: https://github.com/litelightlite/VTON-IQA

Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
(CVPR 2026)
Humanoid robots are expected to play active roles in assistance and service fields; however, existing motion tracking methods are limited to non-contact social interactions or single-person motions and are unsuitable for assistance scenarios requiring immediate response to a partner's posture and dynamics. This research formulates imitation of close interpersonal interactions involving force exchanges as a multi-agent reinforcement learning problem and proposes AssistMimic, a framework that simultaneously learns cooperative policies for both supporter and recipient in a physics simulator. By introducing Partner Policies Initialization, which transfers prior knowledge from single-person motion tracking policies; Dynamic Reference Retargeting, which dynamically retargets assistance motion references according to the partner's real-time posture; and Contact-Promoting Reward, which encourages physically meaningful assistance, we realized a physics-based controller capable of successfully tracking highly contact-intensive interpersonal motions for the first time, achieving state-of-the-art performance on the Inter-X and HHI-Assist benchmarks.
Project page: https://yutoshibata07.github.io/AssistMimic-projectpage/

EMARS: Event-based Motion-Aware Correction, Deblurring and Interpolation of Rolling Shutter Images(ICIP 2026)
Rolling shutter (RS) CMOS sensors are low cost but suffer significant geometric distortion and motion blur under fast motion, limiting applications. Event cameras with high temporal resolution provide useful motion cues for correcting these artifacts; however, existing methods using Implicit Neural Representation (INR) face challenges with residual blur and loss of fine details due to implicit compression. This study proposes a novel framework that explicitly utilizes time-conditioned optical flow as a key kinematic constraint to jointly perform RS correction, deblurring, and frame interpolation. By extracting optical flow at query times and imposing geometric consistency through Flow-Constrained INR, the method learns physically plausible motion trajectories, achieving state-of-the-art PSNR and SSIM across all temporal upsampling factors.

Listening without Looking: Modality Bias in Audio-Visual Captioning(ICIP 2026)
Audio-Visual Captioning integrates audio and video to generate descriptions of entire scenes. Recent advances in multimodal fusion have improved performance, yet the actual complementarity and robustness to degradation in either modality have not been thoroughly examined. This research conducted systematic modality robustness tests by selectively suppressing or degrading audio and video streams for the state-of-the-art model LAVCap, quantitatively evaluating sensitivity and complementarity, revealing a strong bias toward audio stream inference. Additionally, we constructed a novel dataset AudioVisualCaps with text annotations describing both audio and video extending AudioCaps, demonstrating that training LAVCap on this dataset reduces modality bias compared to training on AudioCaps alone.

4D Reconstruction from Sparse Dynamic Cameras(CVPR 2026 Workshop 4DV)
Dynamic 3D (4D) reconstruction from monocular moving cameras has recently advanced significantly but remains fundamentally limited by depth ambiguity. This research focuses on a sparse dynamic camera setting where multiple independently moving cameras capture the same subject. By introducing multi-view geometric constraints while controlling shooting costs, we aim for practical 4D reconstruction suitable for real-world productions such as sports, concerts, and TV programs. Since naive extensions of existing methods cannot resolve complex spatiotemporal inconsistencies, we integrated inter-camera feature matching and intra-camera point tracking to ensure spatiotemporal consistency with a 3D track initialization method, introduced a noise-robust depth ordering regularization loss, and employed spatially and temporally diverse batch sampling strategies. Moreover, we newly developed a real video dataset, LetCamsGo, for this task, demonstrating that the proposed framework substantially improves 4D reconstruction quality in dynamic regions.

Geometric-Photometric Event-based 3D Gaussian Ray Tracing (CVPR 2026 Highlight)
We propose GPERT, the first framework to achieve high-quality 3D reconstruction using 3D Gaussian Splatting (3DGS) by maximally leveraging the event camera's data characteristics of being spatially sparse and temporally dense. Our method separates event-level geometric (depth) rendering using ray tracing and snapshot-level luminance rendering, then integrates them via Warped Events images, resolving the traditional trade-off between accuracy and temporal resolution in event-based 3DGS methods. Without relying on pretrained models or COLMAP initialization, we achieve state-of-the-art performance, one of the fastest training times, and sharp scene edge restoration on real-world datasets.
Authors: Kai Kohyama, Yoshimitsu Aoki (Keio University), Guillermo Gallego (TU Berlin et al.), Shintaro Shiba (Keio University / University of Tokyo)
Project: https://e3ai.github.io/gpert/
Code: https://github.com/e3ai/gpert
Paper: https://arxiv.org/abs/2512.18640

Simultaneous Motion And Noise Estimation with Event Cameras(ICCV2025)
We propose the first method to simultaneously estimate motion (pose changes of autonomous moving bodies and optical flow) and noise solely from raw event camera data. Previously, noise removal from events and motion estimation were conducted separately, but our research integrates both, actively leveraging that event data is inherently linked to motion.
Methodologically, we extend the standard framework of event-based motion estimation, Contrast Maximization (CMax), by quantifying each event's contribution to contrast and iteratively optimizing event signal/noise discrimination and motion parameters. This flexible framework can be combined not only with conventional one-step CMax but also with any motion estimator, including deep learning models.
Experiments achieve state-of-the-art performance on the representative event denoising benchmark E-MLB and competitive results on DND21. Furthermore, the method robustly estimates rotational ego-motion and optical flow, and reduces artifacts in intensity image reconstructions such as E2VID and EVILIP.
Paper: https://arxiv.org/abs/2504.04029
Code: https://github.com/tub-rip/ESMD
Video: https://www.youtube.com/watch?v=iJZsIEWinXk

Sign-to-Speech Prosody Transfer (ICPR2026)
Sign language is an essential communication medium for individuals with hearing impairments. Recent advances in deep learning have improved translation accuracy from sign language to text, making sign language messages more accessible to non-signers. However, sign language includes prosodic features like emphasis and intonation that text alone cannot express, and existing systems fail to capture these adequately. The current mainstream two-stage pipeline (sign-to-text followed by text-to-speech) mediates only via text, resulting in substantial loss of prosodic nuances inherent in sign movements. Therefore, this research proposes a new task, "Sign-to-Speech Prosody Transfer," to directly integrate the prosodic nuances of sign language into synthesized speech. This task poses three major challenges: (1) sign language translation itself is complex, and no high-quality paired datasets of sign language and speech exist; (2) speech prosody is complex, adding difficulty to the translation process; (3) the correspondence between sign language prosody and speech prosody is subtle, and no existing methods map this directly. To address these, we propose S2PFormer (Sign-to-Prosody Transformer), which leverages sign language prosody reconstruction to enable training on unpaired datasets without requiring direct correspondence between sign and speech. Additionally, applying cross-attention to human joint data and text allows capturing fine prosodic details. Extensive experiments confirm that our approach synthesizes speech reflecting sign language prosody, opening new possibilities for more natural sign language communication.

Pre-training with Synthetic Patterns for Audio
This paper proposes a method to pretrain speech encoders using synthetic patterns as substitutes for real speech data. The proposed framework consists of two main components. The first is a Masked Autoencoder (MAE), a self-supervised learning framework that reconstructs original data from randomly masked input, focusing on low-level visual patterns and regularities within data. Thus, the content itself is not critical whether the input is images, speech mel spectrograms, or synthetic patterns. The second component is synthetic data, which unlike real speech, does not involve privacy or licensing issues. By combining MAE with synthetic patterns, we can learn generalized feature representations without relying on real data and avoid issues inherent in real speech. To validate this framework's effectiveness, extensive experiments across 13 speech tasks and 17 synthetic datasets were conducted to analyze which synthetic pattern types are effective for speech. Results show that our method achieves performance comparable to models pretrained on AudioSet-2M and partially surpasses image-based pretraining methods.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to enhance audiovisual representation learning. LG-CAV-MAE integrates a pretrained text encoder with contrastive learning of audio and video and a Masked AutoEncoder, enabling learning across audio, video, and text modalities. To train LG-CAV-MAE, we introduce a method that automatically generates triplets of audio, video, and text from unlabeled videos. First, frame-level captions are generated using an image captioning model, followed by filtering based on CLAP to ensure high consistency between audio and captions. This approach obtains high-quality audio-video-text triplets without requiring manual annotations. Evaluations of LG-CAV-MAE on audio-video retrieval and classification tasks show significant improvements over existing methods, achieving up to a 5.6% increase in recall@10 for retrieval and a 3.2% improvement in classification accuracy.

Text-guided Synthetic Geometric Augmentation for Zero-shot 3D Understanding
Sufficient amounts of training data are indispensable for zero-shot recognition models to achieve adequate generalization performance. However, collecting 3D data and captions necessary for zero-shot 3D classification is costly, posing a significant barrier. Recent generative models have greatly improved the realism of synthetic data, suggesting potential for using generated data as training data. This study addresses the question: "Can synthetic 3D data generated by generative models be used to augment limited 3D datasets?" To clarify this, we propose a synthetic 3D dataset augmentation method called Text-guided Geometric Augmentation (TeGA). TeGA is designed to align with state-of-the-art language-image-3D pretrained models excelling in zero-shot 3D classification, supplementing and expanding scarce 3D data using text-to-3D generative models. Specifically, TeGA applies a consistency filtering strategy to remove noisy samples whose text and geometric shapes do not semantically match from the synthetically generated 3D data produced according to text. Experiments doubling the size of the original datasets show performance improvements surpassing baselines, with zero-shot accuracy increases of +3.0% on Objaverse-LVIS, +4.6% on ScanObjectNN, and +8.7% on ModelNet40. These results demonstrate that TeGA effectively addresses 3D data scarcity and achieves robust zero-shot 3D classification even with limited real data.

Formula-Supervised Sound Event Detection: Pre-Training Without Real Data
In Sound Event Detection (SED) tasks, the lack of precise time-stamped labels and noise from subjective annotations have hindered learning. This study proposes Formula-SED, a dataset that enables large-scale, noise-free pretraining by synthesizing acoustic signals solely from mathematical formulas, using synthesis parameters as ground truth labels. Experiments on the DESED dataset show that pretraining with the proposed dataset is effective in improving both accuracy and convergence speed.

Rethinking Image Super-Resolution from Training Data Perspectives
In the field of image super-resolution, it has traditionally been believed that high-resolution images with minimal compression noise are essential for successful super-resolution learning. This study demonstrates, through image quality evaluation by blockiness distribution measurement and measurement of object diversity via image segmentation counts, that these factors are fundamental to super-resolution learning success. The proposed DiverSeg dataset, despite consisting of low-resolution web-collected images, achieves higher performance than existing super-resolution datasets.

Data Collection-free Masked Video Modeling
Pretraining video transformers generally requires large amounts of data, but collecting and using vast quantities of videos is costly and raises issues such as privacy, licensing, and bias. Using synthetic data is a promising approach to address these challenges, but pretraining solely on artificial data remains difficult. This paper proposes a self-supervised learning framework for video recognition models that leverages easily obtainable, low-cost still images. The method employs a Pseudo-Motion Generator (PMG) module that recursively applies image transformations to still images to generate videos with pseudo motion. These generated videos are then used for training via Masked Video Modeling. This approach is applicable to both natural and synthetic images, addressing cost and other concerns associated with real video data for video model pretraining. Experiments on action recognition tasks demonstrate that this framework effectively learns spatiotemporal features through videos with pseudo motion. The proposed method significantly outperforms existing approaches using only still images and surpasses some methods using both real and synthetic videos.

Proto-Adapter: Efficient Training-Free CLIP-Adapter for Few-Shot Image Classification
In applications where obtaining large amounts of data is difficult, image recognition through few-shot learning is required. While the large-scale vision-language model CLIP can recognize images of arbitrary classes in a zero-shot manner, there remains room for improvement in performance on downstream tasks. We propose a new method, Proto-Adapter, to adapt CLIP to downstream tasks using a small amount of training data. Our method constructs a lightweight adapter using class-specific prototype representations, enabling significant performance improvements on downstream tasks with minimal additional cost. Experiments using 11 types of image recognition benchmarks confirmed the effectiveness of the proposed method.

Guided by the Way: The Role of On-the-route Objects and Scene Text in Enhancing Outdoor Navigation
https://2024.ieee-icra.org
In outdoor environments, Vision-and-Language Navigation (VLN) requires an agent to rely on multi-modal cues from real-world urban environments and natural language instructions. While existing outdoor VLN models predict actions using a combination of panorama and instruction features, this approach ignores objects in the environment and learns data bias to fail navigation. According to our preliminary findings, most instances of navigation failure in previous models were due to turning or stopping at the wrong place. In contrast, humans intuitively frequently use identifiable objects or store names as reference landmarks, ensuring accurate turns and stops, especially in unfamiliar places. To address this insight gap, we propose an Object-Attention VLN (OAVLN) model that helps the agent focus on relevant objects during training and understand the environment better. Our model outperforms previous methods in all evaluation metrics under both seen and unseen scenarios on two existing benchmark datasets, Touchdown and map2seq.

Non-invasive estimation of air convection using schlieren imaging technology with event cameras
https://ieeexplore.ieee.org/document/10301562
Schlieren imaging is a method to visualize density variations in transparent media such as air using a camera. Event cameras, which record only changes, offer features of high speed and high dynamic range compared to conventional frame cameras. Leveraging these characteristics, we developed for the first time in the world a schlieren imaging technique using event cameras. We theoretically and experimentally demonstrated the capability to estimate temporal changes in density associated with air thermal convection and other phenomena. Furthermore, the properties of event cameras eliminate the need for lighting equipment required by conventional frame-based imaging and enable slow-motion analysis.

Video:https://www.youtube.com/watch?v=Ev52n8KgxIU
Code:https://github.com/tub-rip/event_based_bos
Paper:https://ieeexplore.ieee.org/document/10301562
Three-dimensional human body scanning using event cameras
Project page: https://florpeng.github.io/event-based-human-scan/
Arxiv: https://arxiv.org/abs/2404.08504
Conventional 3D pose estimation and human mesh reconstruction methods are limited by the temporal resolution and dynamic range of cameras, constraining the scene. We proposed a method that uses event cameras to perform 3D scanning of the human body solely from events without relying on frame images.
The proposed method achieved higher reconstruction accuracy than conventional frame-based methods and demonstrated effectiveness against severe camera movements that cause blur.

Construction of a large-scale dataset for image recognition targeting retail product shelves
We created a large-scale dataset for image recognition targeting retail product shelves. Due to the characteristic that many products appear simultaneously in a single image of a retail shelf, annotation costs are extremely high. Therefore, we used 3DCG to generate photorealistic images of retail product shelves and automatically annotated them based on the 3D coordinate data of objects, resulting in a large-scale dataset with 200 classes and 100,000 images.
The dataset is available for download on the linked page (use restricted to research purposes).

CG Retail Shelves Dataset – A Massive-Scale, Photorealistic, Rich Annotated CG Dataset for Retail Image Processing –
https://yukiitoh0519.github.io/CG-Retail-Shelves-Dataset
MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation
MaskDiffusion is an open-vocabulary semantic segmentation method leveraging pretrained diffusion models without requiring additional training or annotations. We demonstrated that MaskDiffusion excels at handling open-vocabulary categories including fine-grained proper noun-based categories, expanding the applications of segmentation. MaskDiffusion shows significant qualitative and quantitative improvements compared to other comparable unsupervised segmentation methods on datasets such as Potsdam (+10.5 mIoU) and COCO-Stuff (+14.8 mIoU).

Arxiv: https://arxiv.org/abs/2403.11194
Code : https://github.com/Valkyrja3607/MaskDiffusion
TAG: Guidance-free Open-Vocabulary Semantic Segmentation
We propose a novel approach called TAG to realize open-vocabulary semantic segmentation that requires no training, annotation, or guidance. TAG utilizes pretrained models like CLIP and DINO to segment images into semantically meaningful categories without additional training or dense annotations. It acquires class labels from external databases, providing flexibility to adapt to new scenarios. TAG achieves state-of-the-art results in open-vocabulary segmentation without specifying class names on PascalVOC, PascalContext, and ADE20K datasets.

Arxiv: https://arxiv.org/abs/2403.11197
Code : https://github.com/Valkyrja3607/TAG
Quality evaluation of figure skating jumps using expert gaze information
We proposed a method to predict jump scores using kinematic features obtained from performance videos and tracking systems. In the proposed method, in addition to weighting temporal information, spatial directional weightings based on the gaze patterns of human judges and skaters are applied to extract features influencing jump quality from the performance videos, achieving improved accuracy over baseline models.

Quality evaluation of figure skating jumps using expert gaze information
We proposed a method to predict jump scores using kinematic features obtained from performance videos and tracking systems. In the proposed method, in addition to weighting temporal information, spatial directional weightings based on the gaze patterns of human judges and skaters are applied to extract features influencing jump quality from the performance videos, achieving improved accuracy over baseline models.

Boosting Semantic Segmentation by Conditioning the Backbone with Semantic Boundaries
The Semantic Boundary Conditioned Backbone (SBCB) framework is an effective learning method that improves semantic segmentation performance especially at mask boundaries, while maintaining compatibility with various segmentation architectures. This framework performs multi-task learning of semantic boundary detection (SBD) using multi-scale features obtained from the backbone of segmentation architectures, resulting in features with enhanced boundary awareness. This approach achieved an average improvement of 1.2% in IoU and 2.6% in boundary F-score on the Cityscapes dataset, enhancing segmentation accuracy. The SBCB framework adapts well to a range of backbones, including vision transformer models, demonstrating its potential to advance semantic segmentation without adding complexity to the models.

Boosting Outdoor Vision-and-Language Navigation with On-the-route Objects
The Vision-and-Language Navigation (VLN) task aims for robots to navigate real-world environments using natural language instructions. Although many VLN models have recently been proposed, they typically combine instructions with panorama features to predict action sequences, often overlooking specific semantics and object recognition. Preliminary experiments revealed that existing models do not sufficiently attend to object tokens. This trend contrasts with how humans rely on landmarks when navigating unfamiliar areas. Therefore, this study proposes the Object-Attention VLN (OAVLN) model designed to enhance the agent's environmental understanding by focusing on objects within navigation instructions. Experiments across multiple datasets demonstrated OAVLN's superiority, confirming its ability to leverage objects as primary navigation landmarks and accurately control the agent.

Efficient Video Recognition Based on Important Patch Selection
Videos are computationally much more expensive to process than images, while each frame image is similar and highly redundant. This study proposes a video recognition method that reduces processing costs by identifying patches with small temporal motion or change as redundant and excluding them from input. Since the temporal motion and changes are derived as by-products during decoding of compressed videos, the additional processing cost required is minimal compared to the processing cost saved by excluding patches. Applying the proposed method to a Transformer-based video recognition model achieved over 70% reduction in processing costs with less than one point drop in accuracy for action recognition.

Automatic Procedure Classification in Plastic Surgery Considering Surgical Instrument Information
This study proposes a model for automatically classifying surgical procedures in plastic surgery videos. Compared with endoscopic surgery videos, which have been extensively studied, plastic surgery videos exhibit more diversity in surgical types and body parts, necessitating a more generalizable model. To perform procedure classification in such plastic surgeries, incorporating information about surgical instruments, which are important for understanding surgical scenes, enabled achieving high accuracy. Additionally, the study demonstrated that scene understanding, previously performed only for endoscopic surgeries, can be extended to other types of surgical procedures.

Human Posture Estimation Using Acoustic Information
A novel framework for 3D posture estimation using acoustic information is proposed.
The proposed method mainly consists of
① Sensing using TSP signals
② Creation of acoustic features (Log Mel Spectrum, Intensity Vector)
③ Joint coordinate regression network employing 1D CNN
By capturing amplitude changes and arrival directions of acoustic signals reflected by the human body, it became possible to accurately obtain 3D coordinates of joint points.

Self-Supervised Noise Removal for Event-Based Optical Flow Estimation
Paper
Event cameras output pixel-wise brightness changes asynchronously at high temporal resolution.
By assuming local linearity in spatiotemporal events and fitting them to a plane, normal flow can be estimated.
However, events contain a lot of noise, and outliers degrade the fitting quality.
In response, a method was proposed that introduces a neural network capturing 3D structure to judge whether an event is noise, and performs self-supervised learning while sampling.
Compared to rule-based event selection, the accuracy of estimated flow improved.

Optical Flow and Ego-Motion Estimation Using Event Cameras
Event data significantly differs in nature from conventional image data, possessing especially asynchronous and spatiotemporal characteristics. Therefore, simply applying recent image-based deep learning methods is not necessarily effective. Our research analyzes these spatiotemporal properties in detail and develops egomotion estimation and optical flow estimation methods that achieve high accuracy across various datasets and scenes. In particular, for optical flow estimation, by extending contrast maximization methods, we achieved performance surpassing other machine learning methods using an optimization-based approach.

Peripheral Completion of 360-Degree Images for Efficient 3DCG Background Production
360-degree images are used as background images representing the entire surroundings to efficiently produce scenes in 3DCG production. In this research, we address the problem of generating 360-degree images by inputting a single standard field-of-view image and completing its surroundings. The proposed method using Transformers obtains output images of higher resolution and more natural appearance compared to previous methods. Additionally, it can output diverse result images from a single input, providing users with many options. Thus, this research aims to support efficient and original 3DCG production for users.

Shadow Removal from Document Images by Learning with Fully Synthetic Images
Removing shadows from document images is an important application for improving the quality of digitized documents. Recent studies have proposed many deep learning-based shadow removal methods that learn from sets of images with and without shadows. These conventional supervised learning methods require large paired datasets of document images, which are costly to create. Therefore, in this study, we use 3DCG rendering to create large and diverse datasets without capturing real documents. Experiments show that a deep neural network trained solely on the proposed dataset performs well on real data and that pretraining with this dataset improves performance.

Simultaneous Detection of Objects and Relationships in Dynamic Scene Graph Generation
Dynamic scene graphs are a framework that provides comprehensive recognition within videos by representing objects and their interrelationships in each scene as graph structures. Conventionally, two-step methods detect objects first and then detect relationships, but such methods depend on the object detector and face speed bottlenecks in inference. Our research improves processing speed by concurrently detecting objects and relationships, enabling mutual learning between the object detector and relationship detector.

Non-Deep Active Learning for Deep Neural Networks
Active Learning is an approach that designs label-efficient algorithms by sampling the most representative samples for labeling when creating training data. In this research, we propose a model that derives the most informative unlabeled samples from the output of a task model. The tasks handled include classification, multilabel classification, and semantic segmentation. The model consists of an uncertainty indicator generator and a task model. After training the task model on labeled samples, it predicts unlabeled samples. The uncertainty indicator generator then outputs uncertainty indicators for each unlabeled sample based on the predictions. Samples with high uncertainty are regarded as informative and selected. Experiments using multiple datasets showed that our model achieves higher accuracy than conventional Active Learning methods and reduces execution time by up to approximately one-tenth.

Detection of Osteoarthritis Using Multimodal Hand Data
As global aging progresses, the number of people with osteoarthritis (OA) of the hand, a joint disease with increased risk due to aging, is steadily rising. Current OA diagnosis relies on ultrasound and X-ray examinations conducted by trained physicians, necessitating technologies to reduce patient burden and improve diagnostic efficiency. In this study, we propose a pipeline to automatically detect OA at the finger joint level using multimodal data consisting of videos, RGB images, and thermal images collected from over 200 patients at Keio University Hospital. This is the first study diagnosing hand OA at the individual joint level.

Diagnosis of Scoliosis Using Depth Images
In recent years, the discontinuation of the manufacture and sale of moiré cameras has made the diagnosis of scoliosis from moiré images difficult. In this study, moiré images, pseudo-moiré images, and bilateral correlation images were created through preprocessing of distance images to enhance the features of the distance images. Using deep learning on the preprocessed images, the spinal alignment was estimated. The estimation results from the preprocessed images demonstrated higher accuracy compared to those from the original distance images. In particular, the estimation results from moiré images reconstructed from distance images showed the highest accuracy.

