DRL Optimizes Edge-Cloud Networks with Multi-Timescale Latent Actions

Vo Phi Son, Van-Dinh Nguyen, Ngoc Hung Nguyen, Trinh Van Chien, Symeon Chatzinotas· July 22, 2026 View original

Summary

This paper introduces a two-timescale multi-layer deep reinforcement learning framework with a latent action space (2T-MDRL-LA) for joint optimization in hierarchical edge-cloud computing (HECC) systems. It minimizes end-to-end latency by optimizing service placement, computational delegation, and power control, adapting to dynamic network conditions.

Hierarchical edge-cloud computing (HECC) systems often suffer from load imbalance, leading to high latency and inefficient resource use, especially with dynamic task arrivals and heterogeneous resources. The challenge lies in jointly optimizing service placement, computational delegation, and power control (JSCP), which is a complex, mixed-integer, and NP-hard problem due to tightly coupled discrete and continuous variables. To address this, the research proposes decomposing the JSCP problem into long-term system configuration and short-term resource allocation subproblems, leveraging the inherent differences in decision dynamics. Based on this decomposition, a novel two-timescale multi-layer deep reinforcement learning framework with a latent action space (2T-MDRL-LA) is introduced. This framework jointly optimizes various aspects, including service placement, user association, computational delegation, task offloading, and user transmit power. A variational autoencoder is used to create a latent action representation, effectively compressing the high-dimensional combinatorial action space. Simulations demonstrate that 2T-MDRL-LA adapts well to dynamic network conditions, achieving near-optimal performance, up to a 20.8% reduction in average end-to-end latency, and a 13% improvement in resource utilization compared to schemes without computational delegation, while converging 50% faster than conventional PPO.

Why it matters

Network architects and cloud/edge infrastructure managers can significantly improve the performance and efficiency of their HECC systems, reducing latency and optimizing resource utilization in dynamic environments.

How to implement this in your domain

  1. 1Analyze current edge-cloud network architectures for bottlenecks related to load balancing, latency, and resource utilization.
  2. 2Investigate the application of multi-timescale deep reinforcement learning for joint optimization of service placement and resource allocation.
  3. 3Experiment with latent action space representations (e.g., using VAEs) to manage high-dimensional action spaces in complex network control problems.
  4. 4Pilot a DRL-based controller in a simulated edge-cloud environment to evaluate its ability to adapt to dynamic task loads.
  5. 5Collaborate with network engineers to integrate DRL solutions for real-time optimization of computational delegation and power control.

Who benefits

TelecommunicationsCloud ComputingIoTSmart CitiesAutomotive

Key takeaways

  • A DRL framework optimizes edge-cloud networks by decomposing long-term and short-term problems.
  • It uses a latent action space to manage high-dimensional combinatorial actions efficiently.
  • The framework significantly reduces end-to-end latency and improves resource utilization.
  • It adapts effectively to dynamic network conditions, outperforming conventional methods.

Original post by Vo Phi Son, Van-Dinh Nguyen, Ngoc Hung Nguyen, Trinh Van Chien, Symeon Chatzinotas

"arXiv:2607.18288v1 Announce Type: new Abstract: Load imbalance across edge and cloud layers degrades latency performance in hierarchical edge-cloud computing (HECC) systems under dynamic task arrivals and heterogeneous resources, leading to severe queuing delays and inefficient r…"

View on X

Originally posted by Vo Phi Son, Van-Dinh Nguyen, Ngoc Hung Nguyen, Trinh Van Chien, Symeon Chatzinotas on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses