On-Premise AI Agents Optimized for Factory Hardware

Vasileios Rizeakos, Georgios Paisios, Alexandros Machairas, Michael Birbas, Athanasios Bachoumis· September 3, 2026 View original

Key takeaways

  • Model size is not a reliable predictor of retrieval-augmented answer quality on edge devices.
  • Measurement-driven sub-network selection optimizes for quality and throughput within hardware constraints.
  • Weight-shared supernetworks and distillation make sub-network selection efficient.
  • On-premise AI assistants can provide high-quality information with low power consumption.

Who benefits

ManufacturingIndustrial IoTLogisticsEnergyField Services

Summary

This research presents a method for deploying retrieval-augmented AI assistants on factory shop-floor hardware by selecting optimized sub-networks. It shows that after compression and adaptation, model size is not a reliable predictor of answer quality, and a measurement-driven selection process can achieve high quality and throughput within device constraints.

Deploying conversational AI assistants on factory shop floors to provide workers with access to machine documentation presents a significant challenge: powerful models often exceed the capabilities of on-premise hardware. This paper introduces a novel approach to address this by focusing on sub-network selection after model compression and retrieval-grounded adaptation. The study reveals a counter-intuitive finding: once a model undergoes structural compression and is adapted for retrieval-augmented generation, its size (parameter count) no longer reliably predicts the quality of its answers. While general capabilities tend to decrease linearly with parameter reduction, the quality of answers when augmented with retrieval does not follow the same pattern. This suggests that optimizing for size alone is insufficient. Consequently, the authors propose treating deployment as a post-adaptation selection problem. For each device, a specific sub-network is chosen based on its judged answer quality and measured on-device throughput, while adhering to a configurable general-capability floor and memory budget. Strategies that solely optimize for size, speed, or quality individually lead to compromises in other areas. To make this selection process cost-effective, a weight-shared supernetwork is trained using sandwich-style in-place distillation. A case study involving manufacturing manuals demonstrated that initial quality loss from extraction was largely recovered after retrieval-grounded distillation, and the same assistant could run efficiently across various edge tiers with low power consumption.

Why it matters

For professionals in manufacturing, industrial IoT, or edge computing, this research offers a practical solution for deploying high-performing, retrieval-augmented AI assistants on resource-constrained, on-premise hardware, improving worker efficiency and access to critical information.

How to implement this in your domain

  1. 1Assess your on-premise hardware constraints for deploying AI models in industrial settings.
  2. 2Explore structural compression and retrieval-grounded adaptation techniques for large language models.
  3. 3Implement a measurement-driven sub-network selection process that considers both answer quality and on-device throughput.
  4. 4Utilize weight-shared supernetworks and in-place distillation to efficiently generate deployable sub-networks.
  5. 5Pilot retrieval-augmented AI assistants on the shop floor to provide workers with instant access to documentation.

Original post by Vasileios Rizeakos, Georgios Paisios, Alexandros Machairas, Michael Birbas, Athanasios Bachoumis

"arXiv:2609.02760v1 Announce Type: new Abstract: On-premise assistants can give factory workers conversational access to machine documentation, but models capable of the task rarely fit shop-floor hardware. We show that after structural compression and retrieval-grounded adaptatio…"

View on X

Originally posted by Vasileios Rizeakos, Georgios Paisios, Alexandros Machairas, Michael Birbas, Athanasios Bachoumis on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses