CV

Experience

  • Research Fellow

    Hyundai–NTU–A*STAR Corporate Lab, NTU

    Foundation models for industrial inspection and robotic part picking.

    • Developed DriftAD for few-shot anomaly detection with CLIP, achieving up to 1.5-point gains in 1-shot AUROC and PRO. First-author paper at ACM Multimedia 2026.
    • Extended to zero-shot anomaly detection with TRACE, using CLIP and DINOv3; submitted to AAAI 2027.
    • Built an open-vocabulary perception and grasping pipeline for unfamiliar object categories. Validated on industrial parts and demonstrated in on-site robotic grasping; filed a related patent.
  • Research Associate

    School of EEE, Nanyang Technological University

    Image and video restoration, efficient super-resolution, and multimedia data understanding.

    • Developed a two-stage method to restore damaged JPEG images, reaching 38.92 dB PSNR on 2K images (+5.52 dB over EPDN). First-author CVPR 2023 paper; contributed to a related video-recovery study at NeurIPS 2023.
    • Designed PromptSR, a 0.78M-parameter super-resolution model with gains of up to 0.48 dB over baselines (IEEE TMM 2026). Extended this work to adaptive superpixel token aggregation in AsTaSR.
    • Built Byte2Image and ByteNet to classify raw file fragments without file extensions or metadata. Fused byte and image representations, improving accuracy by up to 12.2% across two benchmarks (IEEE TMM 2024).
  • Graduate Researcher

    College of Computer Science, Chongqing University

    Energy-efficient video decoding and deep learning hardware.

    • Parallelized H.264 decoding with dynamic voltage and frequency scaling: 27–32% faster decoding and 23–25% lower energy use across 720p–2160p video (ICPADS 2018).
    • Contributed to HolyLight, a nanophotonic CNN accelerator achieving 13× higher throughput per watt than a ReRAM baseline (DATE 2019). Second author, with my master’s advisor listed first; 140+ citations.

Education

  • Ph.D. in Electrical and Electronic Engineering

    Nanyang Technological University

  • M.Eng. in Computer Science

    Chongqing University

  • B.Eng. in Internet of Things Engineering

    Chongqing University

Selected Publications

  • DriftAD: Visually-Guided Text Drift for Few-Shot Industrial Anomaly Detection

    Proceedings of the ACM International Conference on Multimedia (ACM MM, CCF-A)

    First author. CLIP adaptation for few-shot anomaly detection with up to 1.5-point gains in 1-shot AUROC and PRO.

  • Bitstream-Corrupted JPEG Images Are Restorable: Two-Stage Compensation and Alignment Framework for Image Restoration

    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR, CCF-A)

    First author. Restored 2K bitstream-corrupted JPEG images at 38.92 dB PSNR.

  • PromptSR: Cascade Prompting for Lightweight Image Super-Resolution

    IEEE Transactions on Multimedia (CCF-A, CAS Q1 Top, JCR Q1)

    First author. A 0.78M-parameter model improving lightweight super-resolution baselines by up to 0.48 dB.

  • ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives

    IEEE Transactions on Multimedia (CCF-A, CAS Q1 Top, JCR Q1)

    First author. Multimodal byte-visual fusion improving file-fragment classification accuracy by up to 12.2%.

  • SSH-Net: A Self-Supervised and Hybrid Network for Noisy Image Watermark Removal

    Journal of Visual Communication and Image Representation (JVCIR, CCF-C, CAS Q3, JCR Q2)

    First author. Self-supervised noisy watermark removal without paired clean targets.

  • A Byte Sequence Is Worth an Image: CNN for File Fragment Classification Using Bit Shift and n-Gram Embeddings

    IEEE International Conference on Artificial Intelligence Circuits and Systems (AICAS)

    First author. Encoded bit-level byte patterns as images for file-fragment classification.

  • HolyLight: A Nanophotonic Accelerator for Deep Learning in Data Centers

    Design, Automation & Test in Europe Conference (DATE, CCF-B)

    Second author; my master’s advisor is the first author. A nanophotonic accelerator achieving 13× higher throughput per watt than a ReRAM baseline; 140+ citations.

  • Fine-Grained Task-Level Parallel and Low-Power H.264 Decoding in Multi-Core Systems

    IEEE International Conference on Parallel and Distributed Systems (ICPADS, CCF-C)

    First author. Parallel H.264 decoding with DVFS, improving speed by 27–32% while reducing energy consumption by 23–25%.

  • MindReading: An Ultra-Low-Power Photonic Accelerator for EEG-Based Human Intention Recognition

    Asia and South Pacific Design Automation Conference (ASP-DAC, CCF-C)

    Co-first author. Photonic acceleration for EEG-based intention recognition with 62.7% lower power consumption.

Skills

Foundation & Multimodal Models: CLIP, DINOv3, SigLIP, ImageBind, SAM, Grounding DINO, LLaVA, Qwen
Vision & Content Understanding: Object detection, Instance segmentation, Object tracking, Anomaly detection, Open-vocabulary perception, Image/video understanding
Learning Methods: Few-/zero-shot learning, Multimodal learning, Self-supervised learning, Prompt learning, Distillation, PEFT
Programming & Tools: Python, C/C++, PyTorch, TensorFlow, Hugging Face, NumPy, OpenCV, Git

Languages

Chinese : Native
English : Professional working proficiency