Publications
publications by categories in reversed chronological order.
2025
- under review
PipeNN: Energy-Efficient and Predictor-Free CPU-GPU Pipeline for Mobile DAG-DNN InferenceYukun Tian, Tianwei Jiang, Kun Tian , and 6 more authors2025The deployment of DAG-based deep neural networks (DNNs) on mobile devices has become increasingly common, enabling advanced AI applications such as real-time object detection and face recognition. However, running these com- plex models on resource-constrained mobile hardware presents significant challenges, particularly in optimizing inference latency and energy consumption. Existing frameworks either limit execution to a single processor or fail to efficiently utilize heterogeneous computing resources, often overlooking energy efficiency in DAG-based DNNs. To address these issues, we propose PipeNN, the first energy-efficient and predictor-free heterogeneous inference framework for mobile devices. Unlike conventional intra-operator parallelism, PipeNN adopts subgraph-level pipeline parallelism, enabling fine-grained task partitioning without extensive profiling overhead. It introduces a custom execution control mechanism, lightweight memory management, and optimized inter- processor data transfer to achieve efficient CPU-GPU execution while minimizing storage and synchronization overhead. Additionally, a coarse-to-fine energy-latency optimization algorithm ensures optimal DAG partitioning, balancing inference speed and energy consumption. We implement PipeNN on various platforms and evaluate it on multiple DNN models and mobile devices. Compared to state-of-the-art inference framework, PipeNN achieves up to 3x speedup and reduces energy consumption by 59.7% on DAG DNN inference.
- under review
SAPE: Efficient End-to-End Dense Representations for Event-based Learning via Spatio-Temporal Adaptive Perception in Multi-scale Temporal ContextYukun Tian, Yongjian Deng, Wei You , and 1 more author2025Event cameras offer unique advantages, such as high dynamic range and low latency. However, they produce sparse, non-uniform 4-D event streams of arbitrary length, which are inherently incompatible with modern deep neural networks (DNNs). Existing methods convert event streams into dense representations using handcrafted rules or coarse learnable transformation, leading to significant spatio-temporal information loss. In this work, we bridge this long-standing gap between event data and off-the-shelf DNNs by introducing a powerful end-to-end representation method, termed Spatio-Temporal Adaptive Perception Event (SAPE). Our method features attention-driven spatio-temporal learnable kernels and group polarity interaction modules, leveraging intrinsic event properties (e.g., sparsity, polarity) to extracting task-relevant spatio-temporal information adaptively. Additionally, we introduce a parallel multi-branch representation learning architecture to capture multi-scale features in diverse temporal contexts efficiently by processing multi-scale features in parallel. The proposed method demonstrates exceptional generalization and robustness across various tasks, backbones and datasets. Experiments on object recognition and semantic segmentation achieve state-of-the-art performance. Furthermore, our method exhibits remarkable adaptation ability by achieving competitive results with only 0.11M (0.53%) of parameters fine-tuned.
2024
- under review
EvAug: Integrating Hierarchical and Adaptive Spatio-Temporal Augmentations into Event-Based Data by Mimicking Real-World Object PatternsYukun Tian, Yongjian Deng, Wei You , and 1 more author2024Event cameras have shown great potential in various applications due to their low latency and high dynamic range. However, challenges such as data scarcity and limited diversity hinder model generalization, and research on event-specific data augmentation remains limited. This work aims to address this gap by introducing a systematic augmentation scheme named EvAug, which is inspired by real-world object patterns. In particular, we first propose Multi-scale Temporal Integration (MSTI) to diversify the motion speed, then introduce Spatial-Salient Event Mask (SSEM) and Temporal-Salient Event Mask (TSEM) to enrich object variants by emulating occlusion and interruption. Our EvAug can facilitate models learning with richer motion patterns, object variants and local spatio-temporal relations, thus improving model robustness and generalization capabilities. Experiment results demonstrate that EvAug consistently yields significant improvements across different tasks, representations, backbones and in multi-modal scenarios (e.g., 4.87% accuracy gain on gesture recognition, 12.03% gain on object classification and 1.8% mIOU gain on semantic segmentation).