Field AI
We advance research to deliver sensing and AI research outcomes to actual fields such as space, sports, fashion, and medical care. Facing unique challenges in each domain—such as satellite image analysis, sports motion evaluation, virtual fitting, and recognition of surgical procedures and diagnostic support—we refine our technologies through collaboration with experts and companies. Considering constraints of real environments and user perspectives, we aim to create systems useful on-site and generate new value.
Phase Detection of Golf Swings Using Combined Action Spotting and Action Segmentation
In golf form analysis, accurately extracting specific moments (phases) during the swing, such as address and impact, from video is crucial. Conventional methods have treated this task as "action spotting," detecting each phase instant directly. However, these methods have limited ability to capture temporal context and may mistakenly confuse similar postures, such as interpreting post-impact poses as backswing frames. This study proposes a method that combines "action segmentation," which estimates action intervals and their order, with focused spotting near phase transitions. Our approach outperformed existing methods on the public benchmark GolfDB and, on an original high-frame-rate dataset shot at 120/240 fps, improved accuracy from 62.5% to 73.7% compared to the most accurate existing method, an increase of 11.2 percentage points.

Surgical Instrument Detection in Open Surgery Using Feature Fusion Based on Hand and Instrument Interaction (MIRU2026)
The technology for detecting surgical instruments within surgical videos serves as a foundation for quantitatively understanding the progress of surgery and the surgeon's techniques. Videos of open surgery, recorded from the surgeon's first-person perspective, often have instruments frequently occluded by the hand, which made conventional image-based instrument detection methods insufficiently accurate. This research focuses on the fact that surgeons hold instruments differently depending on the instrument type and proposes an instrument detection method that uses hand information as auxiliary data. First, surgical instruments and hands are simultaneously detected with a single detector, and the hands holding the instruments are associated based on overlapping detection boxes. Then, hand features are incorporated to refine the instrument class predictions. Evaluation on the first-person surgical video dataset EgoSurgery-tool showed an improvement in mean average precision (mAP) for instrument detection from 0.646 to 0.676 compared to the baseline (Deformable DETR), reducing misrecognition between instruments of similar shapes and false detections of the background.

Quality evaluation of figure skating jumps using expert gaze information
We proposed a method to predict jump scores using kinematic features obtained from performance videos and tracking systems. In the proposed method, in addition to weighting temporal information, spatial directional weightings based on the gaze patterns of human judges and skaters are applied to extract features influencing jump quality from the performance videos, achieving improved accuracy over baseline models.

Quality evaluation of figure skating jumps using expert gaze information
We proposed a method to predict jump scores using kinematic features obtained from performance videos and tracking systems. In the proposed method, in addition to weighting temporal information, spatial directional weightings based on the gaze patterns of human judges and skaters are applied to extract features influencing jump quality from the performance videos, achieving improved accuracy over baseline models.

Detection of Osteoarthritis Using Multimodal Hand Data
As global aging progresses, the number of people with osteoarthritis (OA) of the hand, a joint disease with increased risk due to aging, is steadily rising. Current OA diagnosis relies on ultrasound and X-ray examinations conducted by trained physicians, necessitating technologies to reduce patient burden and improve diagnostic efficiency. In this study, we propose a pipeline to automatically detect OA at the finger joint level using multimodal data consisting of videos, RGB images, and thermal images collected from over 200 patients at Keio University Hospital. This is the first study diagnosing hand OA at the individual joint level.

Diagnosis of Scoliosis Using Depth Images
In recent years, the discontinuation of the manufacture and sale of moiré cameras has made the diagnosis of scoliosis from moiré images difficult. In this study, moiré images, pseudo-moiré images, and bilateral correlation images were created through preprocessing of distance images to enhance the features of the distance images. Using deep learning on the preprocessed images, the spinal alignment was estimated. The estimation results from the preprocessed images demonstrated higher accuracy compared to those from the original distance images. In particular, the estimation results from moiré images reconstructed from distance images showed the highest accuracy.

Estimation of Object Manipulation Methods Based on Learning of Stable Arrangements
We propose a method to estimate the operations needed to restore unnaturally placed objects to natural arrangements. The proposed method utilizes an encoder-decoder type network equipped with a special layer for estimating object manipulations. By providing only the layout of a stable scene, it can self-supervise the learning of procedures for changing object arrangements. Experiments on real images confirmed that operations to change an input scene into a stable state can be generated in real time.

Estimation of Functional Parts of Objects Based on Operation Task Input
We propose task-oriented functional parts, a method for describing the functional parts of an object according to the task, allowing a robot to handle a single object in multiple ways. Unlike affordance, task-oriented functional parts allow multiple handling methods to be described for a single object because functional parts exist for each task. Additionally, we created a dataset for learning task-oriented functional parts and proposed a method for their estimation. On our original dataset containing 1,200 images with 6,000 labels, a mean IOU of 0.80 was achieved.

Tactile Logging: A Method for Describing Operation History on Object Surfaces Based on Human Motion Analysis
We propose a method to analyze demonstrations of human tool operation recorded as RGB-D videos. The method tracks the 3D pose of the human posture and the operated object, estimating interactions occurring on the object. The results are recorded as a time-series usage history (Tactile Log) on the surface of the object's 3D model. The Tactile Log is a novel data representation that visualizes ideal usage methods of objects and can be used to generate 'natural' tool grasping and handling motions by robot arms.

Six Degrees of Freedom Pose Estimation of Similar-Shaped Objects Focusing on Spatial Arrangement of Functional Attributes
We propose a six degrees of freedom pose estimation method that works even when an identical 3D model of the target object does not exist. Tools in the same category typically share the spatial arrangement of functional attributes despite differences in design. Our method uses this as a clue for pose estimation. By simultaneously optimizing the consistency of the arrangement of functional attributes and shape consistency, we confirmed improved pose estimation reliability. In practical use, associating functional attributes and grasp methods with even a single 3D model per object category allows direct handling of real objects without preparing model data for each target object.

Scoliosis Screening by Estimating Spinal Alignment from Back Moiré Images
This study proposes a method that inputs back moiré images of subjects without X-ray exposure and automatically calculates the Cobb angle and VR angle necessary for scoliosis screening. Using moiré images and X-ray images, the method trains a CNN with supervised data from specialist-extracted spinal feature points on X-rays to accurately estimate spinal alignment coordinates from moiré images alone, and automatically calculates the Cobb and VR angles from spinal alignment information. The effectiveness of the proposed method was demonstrated on an original dataset. Currently, we are investigating a method for estimating 3D spinal alignment from 3D back scan data.

Shot Detection in Tennis Match Videos
This study proposes a method to identify player 'shots' at the frame level during tennis matches. Using deep learning to consider the movements of the player, racket, and ball, the method achieves more accurate shot detection than previous ball detection-dependent approaches. This technology can be applied to other sports involving ball shots such as table tennis and volleyball, and can also be adapted to recognize actions like touching or striking objects for user interface applications.

Rugby Video Analysis System
We developed a hybrid video analysis technology combining feature-design-based ball detection/tracking and deep learning-based player detection/tracking, accurately mapping the movement trajectories of balls and players onto a 2D field from a single camera video. Additionally, automatic play classification by deep learning was implemented to explore automating the tagging tasks of major plays traditionally done manually. This technology is applicable not only to rugby but various sports, with potential applications in industrial and other non-sports domains.

American Football Video Analysis System
In American football videos, where player occlusion is significant among team sports and there are many player action patterns during plays, we perform play time estimation using Global Motion Features such as player positions and overall field movement information. Furthermore, by calculating the positions of two characteristic points such as play start and end locations, we classify plays into American football play patterns: Pass, Run, and Kick. After classification, we estimate ball trajectories—information vital for game analysis—without directly detecting the ball itself. This approach enables automatic creation of a game analysis database by acquiring play duration, play classification, and ball trajectory information.

Swimmer Tracking System
Targeting competitive swimming videos, we propose a player tracking and stroke estimation method robust against noise like splashes and independent of shooting conditions. Players are detected and tracked within the video, and detected player images are input into a CNN to extract feature vectors. From these features, a temporal sequence is created and input into a Multi-LSTM network to estimate strokes. Based on the final player position and stroke information, we visualize player speed and strokes, overlaying this on broadcast footage to enhance the live viewing experience.

