Silver Reviewer Award, ICML, Seoul, South Korea
I am a fourth-year Ph.D. student in Electrical and Computer Engineering at the University of Pittsburgh, advised by Prof. Jingtong Hu. I received an M.Eng. from Zhejiang University in 2020, advised by Prof. Cheng Zhuo, and a B.S. from Southwest Jiaotong University in 2017.
My research interests lie in efficient machine learning algorithms and systems. I am currently working on improving the inference efficiency of large-scale vision and language foundation models, with a focus on exploiting structural heterogeneity through sparse attention, model quantization, and KV-cache optimization.
I expect to graduate in Summer 2027 and am actively seeking industry opportunities. Please feel free to reach out if you are interested in my work.
Selected Publications
SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers
A training-free dynamic sparse attention framework that exploits head-level heterogeneity to accelerate feed-forward 3D reconstruction transformers while preserving camera pose and geometry quality.
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
A mixed-precision quantization framework for MoE LLMs that jointly optimizes expert bit-width allocation and router adaptation, enabling ultra-low-bit quantization with minimal accuracy loss.
Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction
A quantization-robust LLM unlearning framework that flattens the forget-loss landscape and updates only forget-critical layers, preserving unlearning effectiveness while maintaining model utility.
Dynamic Mixed-Precision Routing for Efficient Multi-step LLM Interaction
A dynamic mixed-precision routing framework for long-horizon LLM agents that switches between full- and low-precision models at each decision step to improve the accuracy–cost trade-off.
FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference
A KV cache retrieval method that uses 1-bit quantized keys for fine-grained token selection, accelerating long-context LLM inference while matching full-cache accuracy.
Enhancing Facial Expression Recognition in Head-Mounted Displays with Synthetic Data
A training data synthesis framework that transforms abundant frontal-view facial images into head-mounted-camera-view data, improving facial expression recognition in AR/VR scenarios.
STHVC: Spatial-Temporal Hybrid Video Compression for UAV-Assisted IoV Systems
A hybrid video codec that combines a space-time super-resolution network with conventional video codecs to achieve faster encoding and better rate–distortion performance.
STDF: Spatio-Temporal Deformable Fusion for Video Quality Enhancement on Embedded Platforms
An efficient video enhancement and super-resolution system for embedded platforms that leverages fast spatio-temporal deformable fusion for alignment and aggregation of neighboring frames.
Spatio-Temporal Deformable Convolution for Compressed Video Quality Enhancement
A compressed-video quality enhancement method that introduces spatio-temporal deformable convolution for robust single-pass alignment and context aggregation across multiple frames.
Energy-Efficient Real-Time UAV Object Detection on Embedded Platforms
An energy-efficient real-time UAV object detection system on NVIDIA Jetson TX2, achieving 28 FPS at 3 FPS/W and ranking 2nd among 53 GPU-track teams in the 2018 DAC System Design Contest.
Open Source
Awards
3rd Place, System Design Contest (GPU track), DAC, Las Vegas, NV
2nd Place, System Design Contest (GPU track), DAC, San Francisco, CA
Silver Medal, ACM-ICPC Asia Regional Contest, Shanghai, China
Experience
Research Associate, Zhejiang University
Hangzhou, China
Research Intern, Hikvision Research Institute
Hangzhou, China
Service & Teaching
Conference Reviewer: NeurIPS, ICML, ICLR, ICCV, AAAI, MLSys
Journal Reviewer: IEEE TIP, IEEE TNNLS, ACM TECS
Teaching Assistant, University of Pittsburgh