Skip to content

Site search

Search by research topic, paper title or author.

HOMEEdge AI

Edge AI

We research sensing technologies that quickly capture real-world changes by combining new vision sensors such as event cameras with lightweight and efficient AI. In addition to image restoration and estimation of motion and three-dimensional structures that leverage the characteristics of sensors, we also work on reducing inference processing and communication volume, as well as improving labeling efficiency through active learning. Our goal is to realize AI that operates with high accuracy and in real time under limited computing resources and data.

Improving Inference Efficiency in Dense Event Camera Tasks (MIRU2026)

Event cameras are sensors that asynchronously output per-pixel brightness changes. Due to their features of low latency, high dynamic range, and low power consumption, they are expected to be applied in autonomous driving and robotics. Dense tasks that predict at the pixel level, such as semantic segmentation, incur high computational costs. Nevertheless, conventional methods perform inference at every time step regardless of scene changes. This research proposes a method that dynamically switches inference based on scene changes to improve processing efficiency. First, a decision network determines the necessity of inference based on changes in features obtained by a lightweight encoder. When inference is skipped, past prediction results are transformed using optical flow and reused. The computation of this skipping process requires approximately one-seventh of the cost of normal inference. Evaluation on a driving scene dataset showed that, with the same computation, our method achieved higher segmentation accuracy compared to a method that randomly switches between inference and warping.

Event Representation Learning Considering Communication Bandwidth (MIRU 2026)

Event cameras are sensors that asynchronously output pixel-wise luminance changes, offering high temporal resolution and a wide dynamic range. However, their data rate can reach tens of Gbps, making direct handling on edge devices infeasible. This research proposes a method that assumes placing an FPGA between the camera and the edge device to reduce and re-encode events according to communication bandwidth constraints. By voxelizing event streams and using four types of trainable kernels and thresholds that capture different features, the method re-represents events while preserving information necessary for downstream tasks. Evaluation on the automotive dataset DSEC showed that communication bandwidth could be reduced by about 40% while maintaining image reconstruction performance equivalent to raw events. Furthermore, performance exceeding that of raw events was achieved at the same transmission rate.

EMARS: Event-based Motion-Aware Correction, Deblurring and Interpolation of Rolling Shutter Images(ICIP 2026)

Rolling shutter (RS) CMOS sensors are low cost but suffer significant geometric distortion and motion blur under fast motion, limiting applications. Event cameras with high temporal resolution provide useful motion cues for correcting these artifacts; however, existing methods using Implicit Neural Representation (INR) face challenges with residual blur and loss of fine details due to implicit compression. This study proposes a novel framework that explicitly utilizes time-conditioned optical flow as a key kinematic constraint to jointly perform RS correction, deblurring, and frame interpolation. By extracting optical flow at query times and imposing geometric consistency through Flow-Constrained INR, the method learns physically plausible motion trajectories, achieving state-of-the-art PSNR and SSIM across all temporal upsampling factors.

Geometric-Photometric Event-based 3D Gaussian Ray Tracing (CVPR 2026 Highlight)

We propose GPERT, the first framework to achieve high-quality 3D reconstruction using 3D Gaussian Splatting (3DGS) by maximally leveraging the event camera's data characteristics of being spatially sparse and temporally dense. Our method separates event-level geometric (depth) rendering using ray tracing and snapshot-level luminance rendering, then integrates them via Warped Events images, resolving the traditional trade-off between accuracy and temporal resolution in event-based 3DGS methods. Without relying on pretrained models or COLMAP initialization, we achieve state-of-the-art performance, one of the fastest training times, and sharp scene edge restoration on real-world datasets.
Authors: Kai Kohyama, Yoshimitsu Aoki (Keio University), Guillermo Gallego (TU Berlin et al.), Shintaro Shiba (Keio University / University of Tokyo)

Project: https://e3ai.github.io/gpert/
Code: https://github.com/e3ai/gpert
Paper: https://arxiv.org/abs/2512.18640

Simultaneous Motion And Noise Estimation with Event Cameras(ICCV2025)

We propose the first method to simultaneously estimate motion (pose changes of autonomous moving bodies and optical flow) and noise solely from raw event camera data. Previously, noise removal from events and motion estimation were conducted separately, but our research integrates both, actively leveraging that event data is inherently linked to motion.
Methodologically, we extend the standard framework of event-based motion estimation, Contrast Maximization (CMax), by quantifying each event's contribution to contrast and iteratively optimizing event signal/noise discrimination and motion parameters. This flexible framework can be combined not only with conventional one-step CMax but also with any motion estimator, including deep learning models.
Experiments achieve state-of-the-art performance on the representative event denoising benchmark E-MLB and competitive results on DND21. Furthermore, the method robustly estimates rotational ego-motion and optical flow, and reduces artifacts in intensity image reconstructions such as E2VID and EVILIP.

Paper: https://arxiv.org/abs/2504.04029
Code: https://github.com/tub-rip/ESMD
Video: https://www.youtube.com/watch?v=iJZsIEWinXk

Non-invasive estimation of air convection using schlieren imaging technology with event cameras

Accepted to T-PAMI
https://ieeexplore.ieee.org/document/10301562

Schlieren imaging is a method to visualize density variations in transparent media such as air using a camera. Event cameras, which record only changes, offer features of high speed and high dynamic range compared to conventional frame cameras. Leveraging these characteristics, we developed for the first time in the world a schlieren imaging technique using event cameras. We theoretically and experimentally demonstrated the capability to estimate temporal changes in density associated with air thermal convection and other phenomena. Furthermore, the properties of event cameras eliminate the need for lighting equipment required by conventional frame-based imaging and enable slow-motion analysis.

Video:https://www.youtube.com/watch?v=Ev52n8KgxIU
Code:https://github.com/tub-rip/event_based_bos
Paper:https://ieeexplore.ieee.org/document/10301562

Three-dimensional human body scanning using event cameras

Accepted to CV4MR Workshop in CVPR2024
Project page: https://florpeng.github.io/event-based-human-scan/
Arxiv: https://arxiv.org/abs/2404.08504

Conventional 3D pose estimation and human mesh reconstruction methods are limited by the temporal resolution and dynamic range of cameras, constraining the scene. We proposed a method that uses event cameras to perform 3D scanning of the human body solely from events without relying on frame images.
The proposed method achieved higher reconstruction accuracy than conventional frame-based methods and demonstrated effectiveness against severe camera movements that cause blur.

Self-Supervised Noise Removal for Event-Based Optical Flow Estimation

※Accepted to BMVC2022
Paper

Event cameras output pixel-wise brightness changes asynchronously at high temporal resolution.
By assuming local linearity in spatiotemporal events and fitting them to a plane, normal flow can be estimated.
However, events contain a lot of noise, and outliers degrade the fitting quality.
In response, a method was proposed that introduces a neural network capturing 3D structure to judge whether an event is noise, and performs self-supervised learning while sampling.
Compared to rule-based event selection, the accuracy of estimated flow improved.

Optical Flow and Ego-Motion Estimation Using Event Cameras

※Accepted to ECCV2022 : arxiv, video, code
Sensors : Paper

Event data significantly differs in nature from conventional image data, possessing especially asynchronous and spatiotemporal characteristics. Therefore, simply applying recent image-based deep learning methods is not necessarily effective. Our research analyzes these spatiotemporal properties in detail and develops egomotion estimation and optical flow estimation methods that achieve high accuracy across various datasets and scenes. In particular, for optical flow estimation, by extending contrast maximization methods, we achieved performance surpassing other machine learning methods using an optimization-based approach.

Example of Optical Flow Estimation

Non-Deep Active Learning for Deep Neural Networks

※Accepted to Sensors(2022): Paper

Active Learning is an approach that designs label-efficient algorithms by sampling the most representative samples for labeling when creating training data. In this research, we propose a model that derives the most informative unlabeled samples from the output of a task model. The tasks handled include classification, multilabel classification, and semantic segmentation. The model consists of an uncertainty indicator generator and a task model. After training the task model on labeled samples, it predicts unlabeled samples. The uncertainty indicator generator then outputs uncertainty indicators for each unlabeled sample based on the predictions. Samples with high uncertainty are regarded as informative and selected. Experiments using multiple datasets showed that our model achieves higher accuracy than conventional Active Learning methods and reduces execution time by up to approximately one-tenth.

Optical Flow Estimation Using Vehicle-Mounted Event Cameras

We propose a regularization tailored for vehicle-mounted camera scenes for optical flow estimation using event cameras, which utilizes the vehicle's motion characteristics and properties of the Focus of Expansion (FOE). The FOE is defined as the intersection of the camera's translation axis and the image plane. When removing the rotational component from the optical flow of surrounding environmental objects caused by the vehicle's motion, the optical flow radiates outwards from the FOE. The proposed regularization constrains the direction of the optical flow using this property. By evaluating the rotation parameters estimated during the process, the usefulness of this regularization was demonstrated.

Left: Event camera output (light green indicates negative changes, red indicates positive changes). Right: Optical flow estimation results (hue represents flow direction, brightness represents magnitude).

Robust 2D Code Recognition Using Event Cameras Under Varying Lighting and Motion Blur Conditions

In factory automation, reading QR codes on production lines is essential, but lighting conditions and conveyor belt speed cause motion blur. To address this issue, event cameras, which asynchronously capture changes in brightness at each pixel and have high temporal resolution and high dynamic range, are utilized. This research proposes a method to robustly estimate QR codes from event data by representing images with QR codes and affine transformations and performing optimization in the QR code space, which is more constrained than image space.