NANQ Quantization Boosts Analog AI Compute-in-Memory Accuracy.

Yizhe Chen, Wenshuai Yao, Saiya Wang, Yuannuo Feng, Wenbo Qi, Kechao Tang, Ngai Wong, Wenyong Zhou, Wang Kang· August 5, 2026 View original

Key takeaways

  • Analog CIM performance is degraded by device variation and read noise in low-bit models.
  • NANQ is a noise-aware quantization framework for analog CIM.
  • It adaptively allocates precision based on hardware noise profiles.
  • NANQ significantly improves accuracy and reduces perplexity on eFlash CIM with ultra-low bits.

Who benefits

Edge AIConsumer ElectronicsAutomotiveIoTSemiconductor

Summary

This paper introduces NANQ, a noise-aware mixed-precision non-uniform quantization framework for analog compute-in-memory (CIM). NANQ models hardware noise to allocate precision efficiently, improving vision model accuracy by 8.05% and reducing language model PPL by 54.7% on eFlash CIM.

Analog compute-in-memory (CIM) offers significant energy efficiency for neural network inference, but its performance can be severely hampered by device variations and read noise, especially with low-bit quantized models. Existing quantization methods for CIM primarily focus on minimizing ideal quantization error, often overlooking the practical limitations imposed by hardware noise. This oversight leads to inefficient precision allocation. To address this, researchers propose NANQ, a novel noise-aware mixed-precision non-uniform quantization framework specifically designed for analog CIM. NANQ models the magnitude-dependent weight noise observed in eFlash CIM arrays and translates this noise profile into an adaptive quantization density. This allows for finer resolution in low-noise regions while avoiding wasted precision in noise-dominated areas. Furthermore, NANQ assigns layer-wise bit-widths by identifying each layer's precision saturation point under hardware noise. On-chip experiments demonstrate that NANQ significantly improves vision model accuracy by 8.05 percentage points and reduces language model perplexity by 54.7% on average compared to other methods, achieving these gains with only 3.2-3.8 equivalent bits.

Why it matters

Professionals developing energy-efficient AI hardware and deploying models on edge devices can leverage NANQ to achieve higher accuracy with ultra-low-bit quantization, making AI inference more practical and performant in resource-constrained environments.

How to implement this in your domain

  1. 1Evaluate current quantization strategies for AI models deployed on analog compute-in-memory hardware.
  2. 2Integrate noise-floor modeling into your quantization pipeline, specifically for hardware noise profiles.
  3. 3Implement adaptive quantization density, assigning precision based on noise characteristics of different regions.
  4. 4Apply layer-wise bit-width assignment by determining precision saturation points under hardware noise.
  5. 5Benchmark NANQ against existing quantization methods on your target analog CIM hardware for accuracy and efficiency gains.

Original post by Yizhe Chen, Wenshuai Yao, Saiya Wang, Yuannuo Feng, Wenbo Qi, Kechao Tang, Ngai Wong, Wenyong Zhou, Wang Kang

"arXiv:2608.02700v1 Announce Type: new Abstract: Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade low-bit quantized models. Existing CIM-oriented quantization methods mainly minimize ideal qu…"

View on X

Originally posted by Yizhe Chen, Wenshuai Yao, Saiya Wang, Yuannuo Feng, Wenbo Qi, Kechao Tang, Ngai Wong, Wenyong Zhou, Wang Kang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses