Skip to content

Site search

Search by research topic, paper title or author.

HOMEPhysical AI

Physical AI

We conduct research to understand and predict human and robot actions, connecting this to generating appropriate behaviors and control. Based on person detection and tracking, pose estimation, action recognition, and environmental understanding from video, we work on motion generation considering physical consistency and control via reinforcement learning. We also explore methods to efficiently learn motion representations, such as utilizing pseudo-videos generated from still images, aiming for intelligence that enables humans and AI to cooperate in the real world.

Data Collection-free Masked Video Modeling

Accepted to ECCV2024

Pretraining video transformers generally requires large amounts of data, but collecting and using vast quantities of videos is costly and raises issues such as privacy, licensing, and bias. Using synthetic data is a promising approach to address these challenges, but pretraining solely on artificial data remains difficult. This paper proposes a self-supervised learning framework for video recognition models that leverages easily obtainable, low-cost still images. The method employs a Pseudo-Motion Generator (PMG) module that recursively applies image transformations to still images to generate videos with pseudo motion. These generated videos are then used for training via Masked Video Modeling. This approach is applicable to both natural and synthetic images, addressing cost and other concerns associated with real video data for video model pretraining. Experiments on action recognition tasks demonstrate that this framework effectively learns spatiotemporal features through videos with pseudo motion. The proposed method significantly outperforms existing approaches using only still images and surpasses some methods using both real and synthetic videos.

Guided by the Way: The Role of On-the-route Objects and Scene Text in Enhancing Outdoor Navigation

Accepted to ICRA2024
https://2024.ieee-icra.org

In outdoor environments, Vision-and-Language Navigation (VLN) requires an agent to rely on multi-modal cues from real-world urban environments and natural language instructions. While existing outdoor VLN models predict actions using a combination of panorama and instruction features, this approach ignores objects in the environment and learns data bias to fail navigation. According to our preliminary findings, most instances of navigation failure in previous models were due to turning or stopping at the wrong place. In contrast, humans intuitively frequently use identifiable objects or store names as reference landmarks, ensuring accurate turns and stops, especially in unfamiliar places. To address this insight gap, we propose an Object-Attention VLN (OAVLN) model that helps the agent focus on relevant objects during training and understand the environment better. Our model outperforms previous methods in all evaluation metrics under both seen and unseen scenarios on two existing benchmark datasets, Touchdown and map2seq.

Boosting Outdoor Vision-and-Language Navigation with On-the-route Objects

Paper(Embodied AI Workshop in CVPR2023)

The Vision-and-Language Navigation (VLN) task aims for robots to navigate real-world environments using natural language instructions. Although many VLN models have recently been proposed, they typically combine instructions with panorama features to predict action sequences, often overlooking specific semantics and object recognition. Preliminary experiments revealed that existing models do not sufficiently attend to object tokens. This trend contrasts with how humans rely on landmarks when navigating unfamiliar areas. Therefore, this study proposes the Object-Attention VLN (OAVLN) model designed to enhance the agent's environmental understanding by focusing on objects within navigation instructions. Experiments across multiple datasets demonstrated OAVLN's superiority, confirming its ability to leverage objects as primary navigation landmarks and accurately control the agent.

Automatic Procedure Classification in Plastic Surgery Considering Surgical Instrument Information

※Accepted to CARS2023

This study proposes a model for automatically classifying surgical procedures in plastic surgery videos. Compared with endoscopic surgery videos, which have been extensively studied, plastic surgery videos exhibit more diversity in surgical types and body parts, necessitating a more generalizable model. To perform procedure classification in such plastic surgeries, incorporating information about surgical instruments, which are important for understanding surgical scenes, enabled achieving high accuracy. Additionally, the study demonstrated that scene understanding, previously performed only for endoscopic surgeries, can be extended to other types of surgical procedures.

Human Posture Estimation Using Acoustic Information

CVPR2023 Paper Project page Video

A novel framework for 3D posture estimation using acoustic information is proposed.
The proposed method mainly consists of
① Sensing using TSP signals
② Creation of acoustic features (Log Mel Spectrum, Intensity Vector)
③ Joint coordinate regression network employing 1D CNN
By capturing amplitude changes and arrival directions of acoustic signals reflected by the human body, it became possible to accurately obtain 3D coordinates of joint points.

Simultaneous Detection of Objects and Relationships in Dynamic Scene Graph Generation

Presented in SSII2022

Dynamic scene graphs are a framework that provides comprehensive recognition within videos by representing objects and their interrelationships in each scene as graph structures. Conventionally, two-step methods detect objects first and then detect relationships, but such methods depend on the object detector and face speed bottlenecks in inference. Our research improves processing speed by concurrently detecting objects and relationships, enabling mutual learning between the object detector and relationship detector.

Retrieving and Highlighting Action with Spatiotemporal Reference

※Accepted to ICIP2020 : Arxiv

This study proposes Action Highlighting, which visualizes when and where human actions occur in videos, using a cross-modal retrieval framework based on deep learning. For this task, we focus on the co-occurrence of verbs and nouns from video-description pairs, perform representation learning on local regions within the video, and learn appropriate embeddings for each spatiotemporal block using a 3D CNN. By leveraging these learned feature representations, it becomes possible to perform Action Highlighting referencing novel verbs.

Moment-Sentence Matching from Temporal Action Proposals

Vision and language is a multimodal research field that fuses visual video information with linguistic information obtained from natural language. In this study, we focus on the task of localizing the temporal segment within an untrimmed video corresponding to a natural language input describing a part of the video. We propose a two-stage model that first obtains candidate temporal segments using existing Temporal Action Proposals, then determines the optimal time segment by approach ranking using Video Grounding based on Natural Language, solving the problem as a proposal and ranking framework.

Daily Action Recognition Using Human and Object Presence Probabilities

Daily actions that occur over medium to long durations of several minutes to hours often comprise multiple fine primitive actions, making action recognition from video alone challenging. In this research, we built a system that generates human and object presence probability maps along with a medium- to long-term daily human action dataset, using information about "when," "where," "with what," and "how" a person acted as features, and evaluated it on the dataset. Utilizing Human-Object Maps—feature maps obtained by estimating presence probabilities of humans and objects—enabled highly accurate recognition of medium- to long-term daily actions.

Detailed Action Recognition in Factory Production Lines

This research addresses detailed action recognition on production line workplaces, aiming to recognize tasks segmented into even finer primitive actions. Since the work videos mainly consist of detailed actions involving only the arms, capturing differences between tasks from overall image information is difficult. Therefore, we propose a method focusing on upper-body pose information to capture arm movements, combined with hand-area information capturing tool and hand motions. Our approach achieved high recognition rates and robustness to variations in workers and work environments on a small homemade dataset. We also developed a task analysis tool for visualizing recognition results.

Person Re-identification Using Distance Learning with CNN

We propose a novel method for person re-identification by learning similarity of individuals in videos using convolutional neural networks. Each person video is feature-extracted by a CNN and mapped to a Euclidean space where the distance between embedding vectors directly corresponds to a measure of difference between individuals. An improved parameter learning method called Entire Triplet Loss updates parameters by considering all possible triplets in a mini-batch in one step. This simple change in parameter updates greatly improves network generalization performance, allowing embeddings to be more easily separated by person. Evaluation experiments achieved state-of-the-art re-identification rates on an international dataset.

Convolutional Neural Network Architecture

Online Multiple Object Tracking Using Re-identification of Tracking Trajectories

Many existing methods for online multiple object tracking adopt a tracking-by-detection approach, which assigns object bounding boxes obtained through object detection on each video frame in a time series manner. However, these existing methods cannot track targets that become undetected by the object detector due to occlusion or other reasons. Therefore, we propose a method that transitions targets that have once disappeared back into the tracking state by re-identifying tracking trajectories. High-dimensional appearance features of objects are represented as embedding vectors obtained using a convolutional neural network, and re-identification decisions between tracking trajectories are made based on the distances between their embedding vectors. By using mask images of objects obtained through segmentation as input to the network, re-identification decisions robust to background changes can be achieved. Moreover, since re-identification decisions between tracking trajectory pairs are based on distances between low-dimensional vectors, the computational cost increase due to re-identification is minimal.

Overview of Online Tracking Processing Method

Time-Series Action Recognition in Action Transition Videos

This study addresses action recognition in videos where multiple actions transition continuously, proposing the use of hierarchical LSTMs for diverse time-series analysis. We also propose learning contextual information additively while focusing on posture information robust to environmental changes, and filtering contextual features based on posture features to utilize the contextual features more effectively. Applying this to the action recognition task in action transition videos using a dataset, improvements over conventional methods were achieved.

Action Recognition System Using Filtering of Contextual Features

Calibration-Free Gaze Estimation

Existing gaze estimation methods often require special devices such as infrared LEDs or distance sensors, or a prior calibration process. In this study, aiming to realize gaze estimation suitable for practical societal use, we propose a calibration-free gaze point estimation method capable of handling a wide range of head positions relative to the camera. Based on a resolution-independent robust iris tracking method, the gaze estimation consists of face landmark detection, iris tracking, and gaze point estimation, demonstrating that calibration-free gaze estimation is feasible within a wide spatial range. We aim to apply this to various fields.

Calibration-Free Gaze Estimation System