Leveraging Clustering and Reinforcement Learning for a Product Recommender System

Authors

  • Mubbashir Ayub University of Engineering and Technology Taxila
  • Amna Obaid School of Information Technology, Whitecliffe Aucklan Campus, New Zealand
  • Tasawer Khan School of Information Technology, Whitecliffe Aucklan Campus, New Zealand

Keywords:

Clustering, Policy, Q-Learning, Markov Decision Process, Standard Deviation, K-Means

Abstract

The rapid development of digital platforms and e-commerce necessitates the adoption of strong recommender systems that can accommodate a wide range of user preferences. Traditional recommender systems often rely on collaborative filtering and content-based methods, which cannot grasp the complex relationships between users and products. Systems struggle with issues including sparse data, little flexibility, and inadequate user customization. The static nature of collaborative filtering and content-based filtering restricts their efficacy in dynamic situations, where item availability and user preferences can change. Here we are proposing a reinforcement-learning-based method that operates on the clustered data, where the agent learns optimal strategies to recommend products by navigating a defined action space. The proposed method employs a grid-based environment where actions determine the agent's progression through clusters. A variety of clustering algorithms were applied, and the one with the minimum standard deviation was selected for the formulation of states of a grid-based environment. Rewards are assigned based on user satisfaction with the recommended items, encouraging the agent to make relevant and impactful recommendations. Evaluation is done on two public datasets, Amazon and Movie Lens. The evaluation parameters used were start state count analysis, number of states visited by each user, return earned to reach a goal state from a predetermined start state, and number of steps taken in each episode during the learning process.

References

M. Ayub, M. A. Ghazanfar, M. Maqsood, and A. Saleem, “A Jaccard base similarity measure to improve performance of CF based recommender systems,” Int. Conf. Inf. Netw., vol. 2018-January, pp. 1–6, Apr. 2018, doi: 10.1109/ICOIN.2018.8343073.

T. S. Madhulatha, “An Overview on Clustering Methods,” IOSR J. Eng., vol. 02, no. 04, pp. 719–725, May 2012, doi: 10.9790/3021-0204719725.

K. Sivamayil, E. Rajasekar, B. Aljafari, S. Nikolovski, S. Vairavasundaram, and I. Vairavasundaram, “A Systematic Study on Reinforcement Learning Based Applications,” Energies 2023, Vol. 16, Page 1512, vol. 16, no. 3, p. 1512, Feb. 2023, doi: 10.3390/EN16031512.

A. Iftikhar, M. A. Ghazanfar, M. Ayub, S. Ali Alahmari, N. Qazi, and J. Wall, “A reinforcement learning recommender system using bi-clustering and Markov Decision Process,” Expert Syst. Appl., vol. 237, p. 121541, Mar. 2024, doi: 10.1016/J.ESWA.2023.121541.

H. Yin, A. Aryani, S. Petrie, A. Nambissan, A. Astudillo, and S. Cao, “A rapid review of clustering algorithms,” Array, vol. 30, Jul. 2026, doi: 10.1016/j.array.2026.100904.

M. Bin Tariq and H. A. Habib, “A Reinforcement Learning Based RecommendationSystem to Improve Performance of Students in Outcome Based Education Model,” IEEE Access, vol. 12, pp. 36586–36605, 2024, doi: 10.1109/ACCESS.2024.3370852.

J. L. Herlocker, J. A. Konstan, A. Borchers, and J. Riedl, “An algorithmic framework for performing collaborative filtering,” Proc. 22nd Annu. Int. ACM SIGIR Conf. Res. Dev. Inf. Retrieval, SIGIR 1999, pp. 230–237, 1999, doi: 10.1145/312624.312682.

M. J. Pazzani and D. Billsus, “Content-based recommendation systems,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 4321 LNCS, pp. 325–341, Jan. 2007, doi: 10.1007/978-3-540-72079-9_10/Save-Research.

V. Vellaichamy and V. Kalimuthu, “Hybrid collaborative movie recommender system using clustering and bat optimization,” Int. J. Intell. Eng. Syst., vol. 10, no. 5, pp. 38–47, 2017, doi: 10.22266/IJIES2017.1031.05.

W. Hassan, S. E. Hosseini, and S. Pervez, “Real-Time Anomaly Detection in Network Traffic Using Graph Neural Networks and Random Forest,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 14542 LNCS, pp. 194–207, 2024, doi: 10.1007/978-3-031-60994-7_16/SAVE-RESEARCH.

R. Ghous, S. E. Hosseini, and S. P. Chattha, “Lung Cancer Survival Period Prediction: Exploring Machine Learning Approaches,” Int. J. Online Biomed. Eng., vol. 21, no. 04, pp. 45–60, Mar. 2025, doi: 10.3991/IJOE.V21I04.52889.

GanganiParth, AlyaseriSana, and HosseiniSeyed, “Machine Learning Approaches for Predicting Diabetes Onset: A Comparative Study of XGBoost, Random Forest, and Traditional Models,” Mar. 2025, doi: 10.36227/TECHRXIV.174317792.24675569/V1.

S. Mansur et al., “Sales forecasting for retail stores using hybrid neural networks and sales-affecting variables,” PeerJ Comput. Sci., vol. 11, p. e3058, 2025, doi: 10.7717/PEERJ-CS.3058/SUPP-4.

J. Wilson, S. E. Hosseini, and S. Pervez, “Identification of Fake News in Social Media Using Sentimental Analysis,” IEACon 2023 - 2023 IEEE Ind. Electron. Appl. Conf., pp. 220–224, 2023, doi: 10.1109/IEACON57683.2023.10370300.

L. Acosta, S. E. Hosseini, and S. Pervez, “Ethical Challenges Associated with Security Vulnerabilities and Data Privacy in Social Networking,” Proc. - Int. Conf. Dev. eSystems Eng. DeSE, pp. 647–652, 2023, doi: 10.1109/DESE60595.2023.10469464.

M. Ali, S. E. Hosseini, and S. Pervez, “Assessment and Enhancement of Real-Time Recognition of Sign Language Alphabets Through Diverse Machine Learning Techniques,” Proc. - 2025 8th Int. Conf. Data Sci. Mach. Learn. Appl. CDMA 2025, pp. 85–90, 2025, doi: 10.1109/CDMA61895.2025.00020.

M. Ali, S. E. Hosseini, S. Pervez, and M. Ahmad, “Real-Time Recognition of NZ Sign Language Alphabets by Optimal Use of Machine Learning,” Bioeng. 2025, Vol. 12, Page 1068, vol. 12, no. 10, p. 1068, Sep. 2025, doi: 10.3390/bioengineering12101068.

R. Ahuja, A. Solanki, and A. Nayyar, “Movie recommender system using k-means clustering and k-nearest neighbor,” Proc. 9th Int. Conf. Cloud Comput. Data Sci. Eng. Conflu. 2019, pp. 263–268, Jan. 2019, doi: 10.1109/Confluence.2019.8776969.

K. Arai and A. R. Barakbah, “Hierarchical K-means: an algorithm for centroids initialization for K-means,” 2007.

X. Xin, A. Karatzoglou, I. Arapakis, and J. M. Jose, “Self-Supervised Reinforcement Learning for Recommender Systems,” SIGIR 2020 - Proc. 43rd Int. ACM SIGIR Conf. Res. Dev. Inf. Retr., pp. 931–940, Jun. 2020, doi: 10.1145/3397271.3401147.

M. Addanki, S. S, D. B. SLAVAKKAM, R. B. Challagundla, and R. Pamula, “Integrating Sentiment Analysis in Book Recommender System by using Rating Prediction and DBSCAN Algorithm with Hybrid Filtering Technique,” Jul. 2023, doi: 10.21203/RS.3.RS-3173405/V1.

M. M. Afsar, “Personalized Recommendation Using Reinforcement Learning,” 2022, doi: 10.11575/PRISM/39785.

D. M. Shakoor, V. Maihami, and R. Maihami, “A machine learning recommender system based on collaborative filtering using Gaussian mixture model clustering,” Math. Methods Appl. Sci., vol. 49, no. 5, pp. 3589–3602, Mar. 2026, doi: 10.1002/MMA.7801.

X. Li, Z. Wang, R. Hu, Q. Zhu, and L. Wang, “Recommendation algorithm based on improved spectral clustering and transfer learning,” Pattern Anal. Appl. 2017 222, vol. 22, no. 2, pp. 633–647, Nov. 2017, doi: 10.1007/S10044-017-0671-2.

L. Cui, X. Wang, and T. Gu, “A Generic Data Synthesis Framework for Privacy-Preserving Point-of-Interest Recommender Systems,” 2023 Res. Adapt. Converg. Syst. RACS 2023, vol. 1, Aug. 2023, doi: 10.1145/3599957.3606241;Ctype:String:Book.

S. K. Mann and S. Chawla, “A proposed hybrid clustering algorithm using K-means and BIRCH for cluster based cab recommender system (CBCRS),” Int. J. Inf. Technol. 2022 151, vol. 15, no. 1, pp. 219–227, Oct. 2022, doi: 10.1007/S41870-022-01113-6.

S. Hassan, M. Ayub, M. Waqar, and T. khan, “Contrasting Impact of Start State on Performance of a Reinforcement Learning Recommender System,” Int. J. Innov. Sci. Technol., pp. 565–581, May 2024, doi: 10.33411/IJIST/202462565581.

M. Waqar and M. Ayub, “A personalized reinforcement learning recommendation algorithm using bi-clustering techniques,” PLoS One, vol. 20, no. 2, p. e0315533, Feb. 2025, doi: 10.1371/JOURNAL.PONE.0315533.

C. Fan and S. Fujita, “Reinforcement Learning-Based Recommender Systems Enhanced With Graph Neural Networks,” IEEE Access, vol. 13, pp. 150228–150243, 2025, doi: 10.1109/ACCESS.2025.3598092.

S. Choi, H. Ha, U. Hwang, C. Kim, J.-W. Ha, and S. Yoon, “Reinforcement Learning based Recommender System using Biclustering Technique,” Jan. 2018, Accessed: May 19, 2024. [Online]. Available: https://arxiv.org/abs/1801.05532v1

J. Wang, A. Karatzoglou, I. Arapakis, and J. M. Jose, “Reinforcement Learning-based Recommender Systems with Large Language Models for State Reward and Action Modeling,” SIGIR 2024 - Proc. 47th Int. ACM SIGIR Conf. Res. Dev. Inf. Retr., vol. 1, pp. 375–385, Mar. 2024, doi: 10.1145/3626772.3657767.

S. Amin, “An Introduction to Reinforcement Learning and Its Application in Various Domains,” https://services.igi-global.com/resolvedoi/resolve.aspx?doi=10.4018/979-8-3693-1738-9.ch001, pp. 1–25, Jan. 1AD, doi: 10.4018/979-8-3693-1738-9.CH001.

F. M. Harper and J. A. Konstan, “The movielens datasets: History and context,” ACM Trans. Interact. Intell. Syst., vol. 5, no. 4, Dec. 2015, doi: 10.1145/2827872;Wgroup:String:Acm.

C. Tran, J. Y. Kim, W. Y. Shin, and S. W. Kim, “Clustering-Based Collaborative Filtering Using an Incentivized/Penalized User Model,” IEEE Access, vol. 7, pp. 62115–62125, 2019, doi: 10.1109/ACCESS.2019.2914556.

Downloads

Published

2026-07-17

How to Cite

Ayub, M., Amna Obaid, & Tasawer Khan. (2026). Leveraging Clustering and Reinforcement Learning for a Product Recommender System. International Journal of Innovations in Science & Technology, 8(4), 1660–1681. Retrieved from https://journal.50sea.com/index.php/IJIST/article/view/1943