MiCoPro Optimizes Mixed-Precision AI for Edge Devices.

Zijun Jiang, Yangdi Lyu· August 10, 2026 View original

Key takeaways

  • MiCoPro is an end-to-end framework for mixed-precision quantization.
  • It optimizes neural networks for efficient deployment on edge devices.
  • A Hardware-Aware Proxy model enhances prediction accuracy and versatility.
  • MiCoPro achieves significant latency reduction with minimal accuracy loss.

Who benefits

Edge AIIoTAutomotiveConsumer ElectronicsIndustrial Automation

Summary

MiCoPro is an end-to-end framework for mixed-precision quantization (MPQ) that optimizes neural networks for edge AI applications, achieving significant latency reduction with minimal accuracy loss. It features a novel optimization algorithm and a Hardware-Aware Proxy (HAP) model for accurate prediction and hardware versatility.

Quantized Neural Networks (QNNs) using low-bitwidth data are crucial for efficient storage and computation on edge devices. Mixed-precision quantization (MPQ), which applies different bitwidths to different layers, is a popular strategy to balance accuracy and speed. However, existing MPQ exploration algorithms often lack flexibility and efficiency, struggling to predict the complex impact of various MPQ schemes on model performance. Furthermore, a comprehensive end-to-end framework for MPQ optimization and deployment has been missing. To address these challenges, researchers developed MiCo, a holistic framework for MPQ exploration and deployment in edge AI. MiCo incorporates a novel optimization algorithm designed to find accuracy-optimal quantization configurations while adhering to strict latency constraints. This framework was further extended to MiCoPro, which introduces a robust Hardware-Aware Proxy (HAP) model. The HAP model significantly enhances prediction accuracy and hardware versatility by leveraging target-specific latency modeling. MiCoPro enables rapid exploration of MPQ schemes and direct deployment from PyTorch models to bare-metal C code. Demonstrations on both the BitFusion accelerator and SIMD-extended RISC-V processors show that MiCoPro can achieve up to 40% latency reduction with less than a 3% accuracy drop, making it a powerful tool for efficient edge AI.

Why it matters

Professionals can use MiCoPro to deploy high-performance AI models on resource-constrained edge devices, significantly reducing latency and power consumption while maintaining accuracy.

How to implement this in your domain

  1. 1Evaluate MiCoPro for optimizing existing neural networks for deployment on edge AI hardware.
  2. 2Utilize the Hardware-Aware Proxy (HAP) model to predict performance and latency for various mixed-precision configurations.
  3. 3Integrate MiCoPro into your AI development pipeline for end-to-end optimization from PyTorch to bare-metal code.
  4. 4Benchmark the latency and accuracy improvements on your target edge devices, such as RISC-V processors or custom accelerators.

Original post by Zijun Jiang, Yangdi Lyu

"arXiv:2608.06916v1 Announce Type: new Abstract: Quantized Neural Networks~(QNN) with low-bitwidth data have proven promising in efficient storage and computation on edge devices. To mitigate accuracy degradation while maximizing speedup, layer-wise mixed-precision quantization~(M…"

View on X

Originally posted by Zijun Jiang, Yangdi Lyu on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses