Resource Allocation in Energy Efficient Heterogeneous IoT Networks using Multi-Agent Deep Reinforcement Learning

Authors

  • Nimra Hassan Ghazi University, Dera Ghazi Khan, Pakistan
  • Muhammad Afzal Ghazi University, Dera Ghazi Khan, Pakistan.
  • Ameer Hamza Ghazi University, Dera Ghazi Khan, Pakistan.
  • Tayyaba Altaf Ghazi University, Dera Ghazi Khan, Pakistan.
  • Shehzadi Hubba Ghazi University, Dera Ghazi Khan, Pakistan.
  • Hafiz Gulfam Ahmad Umar Ghazi University, Dera Ghazi Khan, Pakistan

Keywords:

Reconfigurable Intelligent Surface (RIS), Energy Efficiency, Heterogeneous IoT Network, Multi-Agent Deep Reinforcement Learning (MADRL), MADDPG, Beamforming, Resource Allocation, 6G, Non-Convex Optimization, Intelligent Reflecting Surface (IRS)

Abstract

Reconfigurable Intelligent Surfaces (RIS) provide a revolutionary passive beamforming technique that can dynamically reshape wireless communication channels. By incorporating RIS into heterogeneous Internet of Things (H-IoT) networks, spectral efficiency and energy efficiency can be improved through intelligent reflection of incoming electromagnetic waves toward desired users. However, the optimization of RIS phase-shift matrices, subcarrier allocation, transmit power allocation, and device scheduling in H-IoT networks constitutes an NP-hard problem due to its high dimensionality and non-convexity, making conventional resource management techniques insufficient for real-time large-scale deployment.

Introduction: The Internet of Things (IoT) is experiencing rapid growth and is expected to exceed 29 billion connected devices by 2030 worldwide by 2030, resulting in heterogeneous IoT (H-IoT) networks with diverse QoS requirements. Energy efficiency (EE) has become the primary design challenge as most IoT devices operate on limited batteries, and conventional resource management techniques fail to scale in complex, high-dimensional H-IoT environments.

Novelty statement: This work proposes RIS-MADDPG, a novel Multi-Agent Deep Reinforcement Learning framework that jointly optimizes RIS passive beamforming phase shifts, subcarrier allocation, and transmit power allocation in RIS-enabled H-IoT networks — a problem not previously addressed through cooperative MADRL with an attention-based centralized critic in distributed heterogeneous IoT networks.

Materials and Methods: The proposed RIS-MADDPG employs the Centralized Training with Decentralized Execution (CTDE) paradigm, where each base station (BS) and RIS controller acts as an independent agent using an extended MADDPG algorithm with an attention-based centralized critic and prioritized experience replay. Simulations are conducted in Python 3.11 with PyTorch 2.1 and OpenAI Gymnasium, using QuaDRiGa 2.6 for spatially consistent channel modeling, Results are averaged over 20 independent Monte Carlo runs across three H-IoT deployment scenarios (Massive IoT, V2X, IIoT).

Results and Discussion: Simulation results demonstrate that RIS-MADDPG achieves an average EE gain of 47.31% over non-RIS baselines, 31.6% over single-agent DRL, and 22.8% over convex alternating optimization approaches. The framework converges within 1,200 training episodes, maintains inference latency below 2 ms for up to 20 agents, achieves an overall QoS satisfaction rate of 98.6%, and degrades by only 15.3% under 20% CSI estimation error, compared to 29.1% degradation for convex optimization methods.

Concluding Remarks: RIS-MADDPG provides an effective, scalable, and robust solution for energy-efficient resource allocation in heterogeneous IoT networks, offering significant performance gains over both classical optimization and single-agent DRL methods.

References

“IoT Connections Forecast to 2030 | GSMA Intelligence.” Accessed: Jul. 06, 2026. [Online]. Available: https://www.gsmaintelligence.com/research/iot-connections-forecast-to-2030

Q. Shi, M. Razaviyayn, Z. Q. Luo, and C. He, “An iteratively weighted MMSE approach to distributed sum-utility maximization for a MIMO interfering broadcast channel,” IEEE Trans. Signal Process., vol. 59, no. 9, pp. 4331–4340, Sep. 2011, doi: 10.1109/TSP.2011.2147784.

Q. Wu and R. Zhang, “Towards Smart and Reconfigurable Environment: Intelligent Reflecting Surface Aided Wireless Network,” IEEE Commun. Mag., vol. 58, no. 1, pp. 106–112, Jan. 2020, doi: 10.1109/MCOM.001.1900107.

Q. Wu and R. Zhang, “Intelligent Reflecting Surface Enhanced Wireless Network via Joint Active and Passive Beamforming,” IEEE Trans. Wirel. Commun., vol. 18, no. 11, pp. 5394–5409, Nov. 2019, doi: 10.1109/TWC.2019.2936025.

Y. Li, “Reinforcement Learning Applications,” Aug. 2019, Accessed: Sep. 16, 2024. [Online]. Available: http://arxiv.org/abs/1908.06973

Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, Igor Mordatch, “Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments,” arXiv:1706.02275, 2020, [Online]. Available: https://arxiv.org/abs/1706.02275

Y. Wang, X. Li, X. Yi, and S. Jin, “Joint User Scheduling and Precoding for RIS-Aided MU-MISO Systems: A MADRL Approach,” IEEE Trans. Commun., vol. 73, no. 6, pp. 3880–3893, 2025, doi: 10.1109/TCOMM.2024.3496745.

P. S. Aung, L. X. Nguyen, Y. K. Tun, Z. Han, and C. S. Hong, “Deep Reinforcement Learning-Based Joint Spectrum Allocation and Configuration Design for STAR-RIS-Assisted V2X Communications,” IEEE Internet Things J., vol. 11, no. 7, pp. 11298–11311, Apr. 2024, doi: 10.1109/JIOT.2023.3329893.

H. Muhammad Fahad Noman, K. Dimyati, K. Ariffin Noordin, E. Hanafi and A. Abdrabou, “FeDRL-D2D: Federated Deep Reinforcement Learning- Empowered Resource Allocation Scheme for Energy Efficiency Maximization in D2D-Assisted 6G Networks,” IEEE Access, vol. 12, pp. 109775–109792, 2024, doi: 10.1109/ACCESS.2024.3434619.

K. Guo, M. Wu, X. Li, H. Song, and N. Kumar, “Deep Reinforcement Learning and NOMA-Based Multi-Objective RIS-Assisted IS-UAV-TNs: Trajectory Optimization and Beamforming Design,” IEEE Trans. Intell. Transp. Syst., vol. 24, no. 9, pp. 10197–10210, Sep. 2023, doi: 10.1109/TITS.2023.3267607.

Manzoor Ahmed, Fang Xu, Yuanlin Lyu, Aized Amin Soofi, “RIS-Driven Resource Allocation Strategies for Diverse Network Environments: A Comprehensive Review,” arXiv:2501.03075, 2025, [Online]. Available: https://arxiv.org/abs/2501.03075

B. Di, H. Zhang, L. Song, Y. Li, Z. Han, and H. V. Poor, “Hybrid Beamforming for Reconfigurable Intelligent Surface based Multi-User Communications: Achievable Rates with Limited Discrete Phase Shifts,” IEEE J. Sel. Areas Commun., vol. 38, no. 8, pp. 1809–1822, Aug. 2020, doi: 10.1109/JSAC.2020.3000813.

C. Huang, A. Zappone, G. C. Alexandropoulos, M. Debbah, and C. Yuen, “Reconfigurable intelligent surfaces for energy efficiency in wireless communication,” IEEE Trans. Wirel. Commun., vol. 18, no. 8, pp. 4157–4170, Aug. 2019, doi: 10.1109/TWC.2019.2922609.

X. Guan, Q. Wu, and R. Zhang, “Intelligent Reflecting Surface Assisted Secrecy Communication: Is Artificial Noise Helpful or Not?,” IEEE Wirel. Commun. Lett., vol. 9, no. 6, pp. 778–782, Jun. 2020, doi: 10.1109/LWC.2020.2969629.

C. Huang, G. C. Alexandropoulos, A. Zappone, M. Debbah, and C. Yuen, “Energy Efficient Multi-User MISO Communication Using Low Resolution Large Intelligent Surfaces,” 2018 IEEE Globecom Work. GC Wkshps 2018 - Proc., Jul. 2018, doi: 10.1109/GLOCOMW.2018.8644519.

W. Jin, J. Zhang, C. K. Wen, S. Jin, X. Li, and S. Han, “Low-Complexity Joint Beamforming for RIS-Assisted MU-MISO Systems Based on Model-Driven Deep Learning,” IEEE Trans. Wirel. Commun., vol. 23, no. 7, pp. 6968–6982, 2024, doi: 10.1109/TWC.2023.3336742.

H. Jiao, H. Liu, and Z. Wang, “Reconfigurable Intelligent Surfaces aided Wireless Communication: Key Technologies and Challenges,” 2022 Int. Wirel. Commun. Mob. Comput. IWCMC 2022, pp. 1364–1368, 2022, doi: 10.1109/IWCMC55113.2022.9824117.

R. Zhong, X. Mu, and Y. Liu, “Machine Learning Empowered Large RIS-assisted Near-field Communications,” IEEE Veh. Technol. Conf., 2023, doi: 10.1109/VTC2023-FALL60731.2023.10333609.

S. S. Hassan, Y. M. Park, L. X. Nguyen, Y. K. Tun, Z. Han, and C. S. Hong, “Optimizing Connectivity for Remote Users with RIS-Enabled UAVs: A Game-Theoretic and Whale Optimization Synergy,” Sep. 2024, doi: 10.36227/TECHRXIV.172668484.41027786/V1.

C. Huang, R. Mo, and C. Yuen, “Reconfigurable Intelligent Surface Assisted Multiuser MISO Systems Exploiting Deep Reinforcement Learning,” IEEE J. Sel. Areas Commun., vol. 38, no. 8, pp. 1839–1850, Aug. 2020, doi: 10.1109/JSAC.2020.3000835.

Z. Han, D. Niyato, W. Saad, T. Başar, and A. Hjørungnes, “Game theory in wireless and communication networks: Theory, models, and applications,” Game Theory Wirel. Commun. Networks Theory, Model. Appl., vol. 9780521196963, pp. 1–535, Jan. 2011, doi: 10.1017/CBO9780511895043.

S. Wang, H. Liu, P. H. Gomes, and B. Krishnamachari, “Deep Reinforcement Learning for Dynamic Multichannel Access in Wireless Networks,” IEEE Trans. Cogn. Commun. Netw., vol. 4, no. 2, pp. 257–265, Feb. 2018, doi: 10.1109/TCCN.2018.2809722.

Mahmoud Abbasi, Amin Shahraki, “Deep Reinforcement Learning for QoS provisioning at the MAC layer: A Survey,” Eng. Appl. Artif. Intell., vol. 102, 2021, doi: https://doi.org/10.1016/j.engappai.2021.104234.

“Nonorthogonal Multiple Access for 5G and Beyond | IEEE Journals & Magazine | IEEE Xplore.” Accessed: Jul. 06, 2026. [Online]. Available: https://ieeexplore.ieee.org/document/8114722

Junhui Zhao, Fajin Hu, “Multi-agent deep reinforcement learning based resource management in heterogeneous V2X networks,” Digit. Commun. Networks, vol. 11, no. 1, pp. 182–190, 2025, doi: https://doi.org/10.1016/j.dcan.2023.06.003.

S. Nimmala, P. V. Sena, S. Inturi, S. Janbhasha, P. Narsimhulu, and J. Manoranjini, “Multi-Agent Deep Reinforcement Learning for Intelligent Industrial Iot Networks,” Proc. 9th Int. Conf. Inven. Syst. Control. ICISC 2025, pp. 455–459, 2025, doi: 10.1109/ICISC65841.2025.11187915.

“AI-Driven Resource Management for Heterogeneous Industrial IoT: Challenges and Opportunities | IEEE Journals & Magazine | IEEE Xplore.” Accessed: Jul. 06, 2026. [Online]. Available: https://ieeexplore.ieee.org/document/11278729

V. Mnih et al., “Human-level control through deep reinforcement learning,” Nat. 2015 5187540, vol. 518, no. 7540, pp. 529–533, Feb. 2015, doi: 10.1038/nature14236.

C. He, Y. Hu, Y. Chen, and B. Zeng, “Joint Power Allocation and Channel Assignment for NOMA with Deep Reinforcement Learning,” IEEE J. Sel. Areas Commun., vol. 37, no. 10, pp. 2200–2210, Oct. 2019, doi: 10.1109/JSAC.2019.2933762.

Y. S. Nasir and D. Guo, “Multi-Agent Deep Reinforcement Learning for Dynamic Power Allocation in Wireless Networks,” IEEE J. Sel. Areas Commun., vol. 37, no. 10, pp. 2239–2250, Oct. 2019, doi: 10.1109/JSAC.2019.2933973.

H. Zhang, X. Li, N. Gao, X. Yi, and S. Jin, “A Deep Reinforcement Learning Approach to Two-Timescale Transmission for RIS-Aided Multiuser MISO systems,” IEEE Wirel. Commun. Lett., vol. 12, no. 8, pp. 1444–1448, Aug. 2023, doi: 10.1109/LWC.2023.3278171.

F. Liang, C. Shen, W. Yu, and F. Wu, “Towards optimal power control via ensembling deep neural networks,” IEEE Trans. Commun., vol. 68, no. 3, pp. 1760–1776, Mar. 2020, doi: 10.1109/TCOMM.2019.2957482.

Y. Xiao, Y. Song, and J. Liu, “Multi-Agent Deep Reinforcement Learning Based Resource Allocation for Ultra-Reliable Low-Latency Internet of Controllable Things,” IEEE Trans. Wirel. Commun., vol. 22, no. 8, pp. 5414–5430, Aug. 2023, doi: 10.1109/TWC.2022.3233853.

Sadhvi Parashar, Rajeev Arya, “MA2CL: : Multi‐Agent Actor‐Critic Learning Scheme for Efficient Resource Management in 5G‐Enabled NB‐IoT Networks,” Internet Technol. Lett., vol. 8, no. 3, 2025, doi: https://doi.org/10.1002/itl2.70011.

Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine, “Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,” arXiv:1801.01290, 2018, [Online]. Available: https://arxiv.org/abs/1801.01290

G. Lee, M. Jung, A. T. Z. Kasgari, W. Saad, and M. Bennis, “Deep Reinforcement Learning for Energy-Efficient Networking with Reconfigurable Intelligent Surfaces,” IEEE Int. Conf. Commun., vol. 2020-June, Jun. 2020, doi: 10.1109/ICC40277.2020.9149380.

Tom Schaul, John Quan, Ioannis Antonoglou, David Silver, “Prioritized Experience Replay,” arXiv:1511.05952, 2016, [Online]. Available: https://arxiv.org/abs/1511.05952

Downloads

Published

2026-06-27

How to Cite

Hassan, N., Afzal, M., Hamza, A., Altaf, T., Hubba, S., & Hafiz Gulfam Ahmad Umar. (2026). Resource Allocation in Energy Efficient Heterogeneous IoT Networks using Multi-Agent Deep Reinforcement Learning. International Journal of Innovations in Science & Technology, 8(3), 1462–1484. Retrieved from https://journal.50sea.com/index.php/IJIST/article/view/1938