• International Journal of Technology (IJTech)
  • Vol 17, No 5 (2026)

A Deep Reinforcement Learning Scheduling Approach for Production Lines in Series

A Deep Reinforcement Learning Scheduling Approach for Production Lines in Series

Title: A Deep Reinforcement Learning Scheduling Approach for Production Lines in Series
Zhijing Zhou, Jiajia Wang, Gang Yin, Jiachen Zhou, Yongxiang Zhang, Jian Guo, Junhao Zhang

Corresponding email:


Cite this article as:
Zhou, Z., Wang, J., Yin, G., Zhou, J., Zhang, Y., Guo, J., & Zhang, J. (2026). A deep reinforcement learning scheduling approach for production lines in series. International Journal of Technology, 17 (5), 1879–1897


27
Downloads
Zhijing Zhou 1. College of Mechanical and Electrical Engineering, Wenzhou University, Wenzhou 325035, China 2. Zhejiang Haozhonghao Health Products Co., Ltd, Wenzhou, 325400, China
Jiajia Wang College of Mechanical and Electrical Engineering, Wenzhou University, Wenzhou 325035, China
Gang Yin Zhejiang Haozhonghao Health Products Co., Ltd, Wenzhou, 325400, China
Jiachen Zhou Academy of Interdisciplinary Studies, Hong Kong University of Science and Technology, Hong Kong, 999077, China
Yongxiang Zhang Zhejiang Haozhonghao Health Products Co., Ltd, Wenzhou, 325400, China
Jian Guo College of Mechanical and Electrical Engineering, Wenzhou University, Wenzhou 325035, China
Junhao Zhang College of Mechanical and Electrical Engineering, Wenzhou University, Wenzhou 325035, China
Email to Corresponding Author

Abstract
A Deep Reinforcement Learning Scheduling Approach for Production Lines in Series

Mixed-model assembly lines (MMAL) are increasingly replacing traditional single-type lines to meet mass customization needs as product demand diversifies. This paper addresses the series-parallel production line sequencing problem (SPPLSP) within MMAL, addressing the complexity introduced by fluctuating orders, asynchronous machinery, and resource imbalances. This study introduces a deep reinforcement learning (DRL) approach to address these challenges, focusing on the impact of machine selection and job sequencing on order tardiness and energy consumption during transitions. Initially, a mixed-integer linear programming model (MILPM) is expressed with constraints, such as conditions related to serial-parallel production lines and machines, to make a clear problem description. Subsequently, a double deep Q-network (DDQN) algorithm is employed to solve the SPPLSP, utilizing an innovative state and action space representation to maintain dynamic balance across serial and parallel lines and adapt to changes in machine and workpiece data. Additionally, an improved agent reward function integrates a multi-agent hierarchical distributed scheduling structure to approximate the actual objectives. Finally, a simulation-based numerical study, grounded in real-world case scenarios, validates the performance of the proposed method across various scales of problems. The numerical results reveal that the DDQN method consistently achieves faster convergence and superior stability values across all test instances and yields higher-quality Pareto frontier solutions in 49/50 test instances.

Double deep Q-network; Mixed-model assembly line; Multi-agent; Series-parallel production line sequencing

Supplementary Material
FilenameDescription
R3-IE-8191-20260828151112.docx ---
References

Alla, K. R., & Thangarasu, G. (2023). Supply chain management of drug products in blockchain using reinforcement learning. International Journal of Technology, 14 (6), 1256–1265. https://doi.org/10.14716/ijtech.v14i6.6626

Azadeh, A., Shoja, B. M., Ghanei, S., & Sheikhalishahi, M. (2015). A multi-objective optimization problem for multi-state series-parallel systems: A two-stage flow-shop manufacturing system. Reliability Engineering & System Safety, 136, 62–74. https://doi.org/10.1016/j.ress.2014.11.009

Chen, C., Xia, B., Zhou, B., & Xi, L. (2015). A reinforcement learning based approach for a multiple-load carrier scheduling problem. Journal of Intelligent Manufacturing, 26, 1233–1245. https://doi.org/10.1007/s10845-013-0852-9

Du, Y., Li, J., Li, C., & Duan, P. (2024). A reinforcement learning approach for flexible job shop scheduling problem with crane transportation and setup times. IEEE Transactions on Neural Networks and Learning Systems, 35 (4), 5695–5709. https://doi.org/10.1109/TNNLS.2022.3208942

Farouk, M. M., Chin, C. G., & Roslee, M. (2026). Hybrid reinforcement learning-enabled scheduler for 5G burst traffic. International Journal of Technology, 17 (3), 866–882. https://doi.org/10.14716/ijtech.v17i3.8208

Fuqing, Z., & Yu, J. (2026). A multi-agent reinforcement learning with multi-task learning and graph edge attention networks for dynamic flexible job shop scheduling. Applied Soft Computing, 187, 114277. https://doi.org/10.1016/j.asoc.2025.114277

Gao, K., Huang, Y., Sadollah, A., & Wang, A. (2020). A review of energy-efficient scheduling in intelligent production systems. Complex & Intelligent Systems, 6, 237–249. https://doi.org/10.1007/s40747-019-00122-6

Gholizadeh, H., Chaleshigar, M., & Fazlollahtabar, H. (2022). Robust optimization of uncertainty-based preventive maintenance model for scheduling series–parallel production systems (real case: Disposable appliances production). ISA Transactions, 128, 54–67. https://doi.org/10.1016/j.isatra.2021.11.041

Han, X., Liu, H., Sun, F., & Zhang, X. (2019). Active object detection with multistep action prediction using deep q-network. IEEE Transactions on Industrial Informatics, 15 (6), 3723–3731. https://doi.org/10.1109/TII.2019.2890849

Hu, H., Jia, X., He, Q., Fu, S., & Liu, K. (2020). Deep reinforcement learning based AGVs real-time scheduling with mixed rule for flexible shop floor in industry 4.0. Computers & Industrial Engineering, 149, 106749. https://doi.org/10.1016/j.cie.2020.106749

Jing, X., Yao, X., Liu, M., & Zhou, J. (2024). Multi-agent reinforcement learning based on graph convolutional network for flexible job shop scheduling. Journal of Intelligent Manufacturing, 35 (1), 75–93. https://doi.org/10.1007/s10845-022-02037-5

Kayhan, B. M., & Yildiz, G. (2023). Reinforcement learning applications to machine scheduling problems: A comprehensive literature review. Journal of Intelligent Manufacturing, 34 (3), 905–929. https://doi.org/10.1007/s10845-021-01847-3

Khadivi, M., Charter, T., Yaghoubi, M., Jalayer, M., Ahang, M., Shojaeinasab, A., & Najjaran, H. (2025). Deep reinforcement learning for machine scheduling: Methodology, the state-of-the-art, and future directions. Computers & Industrial Engineering, 200, 110856. https://doi.org/10.1016/j.cie.2025.110856

Li, L., Yang, X., Yang, S., & Xu, X. (2023). Optimization of oxygen system scheduling in hybrid action space based on deep reinforcement learning. Computers & Chemical Engineering, 171, 108168. https://doi.org/10.1016/j.compchemeng.2023.108168

Li, Y. F., & Peng, R. (2014). Availability modeling and optimization of dynamic multi-state series–parallel systems with random reconfiguration. Reliability Engineering & System Safety, 127, 47–57. https://doi.org/10.1016/j.ress.2014.03.005

Liu, X., Yang, X., & Lei, M. (2021). Optimisation of mixed-model assembly line balancing problem under uncertain demand. Journal of Manufacturing Systems, 59, 214–227. https://doi.org/10.1016/j.jmsy.2021.02.019

Mao, H., Liu, Z., & Qiu, C. (2023). Adaptive disassembly sequence planning for VR maintenance training via deep reinforcement learning. The International Journal of Advanced Manufacturing Technology, 124, 3039–3048. https://doi.org/10.1007/s00170-021-08290-x

Mosadegh, H., Fatemi Ghomi, S. M. T., & Süer, G. A. (2017). Heuristic approaches for mixed-model sequencing problem with stochastic processing times. International Journal of Production Research, 55 (10), 2857–2880. https://doi.org/10.1080/00207543.2016.1223897

Mosadegh, H., Ghomi, S. M. T. F., & Süer, G. A. (2020). Stochastic mixed-model assembly line sequencing problem: Mathematical modeling and q-learning based simulated annealing hyper-heuristics. European Journal of Operational Research, 282 (2), 530–544. https://doi.org/10.1016/j.ejor.2019.09.021

Nourelfath, M., & Yalaoui, F. (2012). Integrated load distribution and production planning in series-parallel multi-state systems with failure rate depending on load. Reliability Engineering & System Safety, 106, 138–145. https://doi.org/10.1016/j.ress.2012.06.006

Paeng, B., Park, I. B., & Park, J. (2021). Deep reinforcement learning for minimizing tardiness in parallel machine scheduling with sequence dependent family setups. IEEE Access, 9, 101390–101401. https://doi.org/10.1109/ACCESS.2021.3097254

Rai, R., Tiwari, M. K., Ivanov, D., & Dolgui, A. (2021). Machine learning in manufacturing and industry 4.0 applications. International Journal of Production Research, 59 (16), 4773–4778. https://doi.org/10.1109/OJIES.2024.3431240

Sikora, C. G. S. (2024). Balancing mixed-model assembly lines for random sequences. European Journal of Operational Research, 314 (2), 597–611. https://doi.org/10.1016/j.ejor.2023.10.008

Taube, F., & Minner, S. (2018). Resequencing mixed-model assembly lines with restoration to customer orders. Omega, 78, 99–111. https://doi.org/10.1016/j.omega.2017.11.006

Van Hasselt, H., Guez, A., & Silver, D. (2016). Deep reinforcement learning with double q-learning. Proceedings of the AAAI Conference on Artificial Intelligence, 30 (1). https://doi.org/10.48550/arXiv.1509.06461

Wang, J., Wang, W., Hu, X., Qiu, L., & Zang, H. (2024). Black-winged kite algorithm: A nature-inspired meta-heuristic for solving benchmark functions and engineering problems. Artificial Intelligence Review, 57 (4), 98. https://doi.org/10.1007/s10462-024-10723-4

Wang, K., Li, X., Gao, L., Li, P., & Gupta, S. M. (2021). A genetic simulated annealing algorithm for parallel partial disassembly line balancing problem. Applied Soft Computing, 107, 107404. https://doi.org/10.1016/j.asoc.2021.107404

Wei, Q., Li, Y., Zhang, J., & Wang, F. Y. (2022). VGN: Value decomposition with graph attention networks for multiagent reinforcement learning. IEEE Transactions on Neural Networks and Learning Systems, 35 (1), 182–195. https://doi.org/10.1109/TNNLS.2022.3172572

Wen, G., Fu, J., Dai, P., & Zhou, J. (2021). DTDE: A new cooperative multi-agent reinforcement learning framework. The Innovation, 2 (4). https://doi.org/10.1016/j.xinn.2021.100162

Xi, S., Smith, J. M. G., Chen, Q., Mao, N., & Zhang, H. (2022). Simultaneous machine selection and buffer allocation in large unbalanced series-parallel production lines. International Journal of Production Research, 60 (7), 2103–2125. https://doi.org/10.1080/00207543.2021.1884306

Yamsani, N., & Reddy, P. C. (2026). EdgeSched-DQN: An intelligent deep reinforcement learning-based framework for optimized task scheduling in edge-cloud environments. Array, 29, 100645. https://doi.org/10.1016/j.array.2025.100645

Yang, W., & Cheng, W. (2020). Modelling and solving mixed-model two-sided assembly line balancing problem with sequence-dependent setup time. International Journal of Production Research, 58 (21), 6638–6659. https://doi.org/10.1080/00207543.2019.1683255

Yelles-Chaouche, A. R., Gurevsky, E., Brahimi, N., & Dolgui, A. (2022). Minimizing task reassignments under balancing multi-product reconfigurable manufacturing lines. Computers & Industrial Engineering, 173, 108660. https://doi.org/10.1016/j.cie.2022.108660

Yuan, M., Li, Y., Zhang, L., & Pei, F. (2021). Research on intelligent workshop resource scheduling method based on improved NSGA-II algorithm. Robotics and Computer-Integrated Manufacturing, 71, 102141. https://doi.org/10.1016/j.rcim.2021.102141

Zhang, B., Xu, L., & Zhang, J. (2020a). A multi-objective cellular genetic algorithm for energy-oriented balancing and sequencing problem of mixed-model assembly line. Journal of Cleaner Production, 244, 118845. https://doi.org/10.1016/j.jclepro.2019.118845

Zhang, K., He, F., Zhang, Z., Lin, X., & Li, M. (2020b). Multi-vehicle routing problems with soft time windows: A multi-agent reinforcement learning approach. Transportation Research Part C: Emerging Technologies, 121, 102861. https://doi.org/10.1016/j.trc.2020.102861

Zhang, M., Liu, Y., Chen, H., & Cai, W. (2025). Double DQN-based efficient quality of service routing protocol in Internet of Underwater Things with mobile nodes. Ad Hoc Networks, 175, 103856. https://doi.org/10.1016/j.adhoc.2025.103856

Zhao, C., & Melkote, S. N. (2024). Learning the manufacturing capabilities of machining and finishing processes using a deep neural network model. Journal of Intelligent Manufacturing, 35 (4), 1845–1865. https://doi.org/10.1007/s10845-023-02134-z

Zhao, M., Guo, X., Zhang, X., Fang, Y., & Ou, Y. (2020). Aspw-drl: Assembly sequence planning for workpieces via a deep reinforcement learning approach. Assembly Automation, 40 (1), 65–75. https://doi.org/10.1108/AA-11-2018-0211

Zhou, B. H., & Shen, C. Y. (2018). Multi-objective optimization of material delivery for mixed model assembly lines with energy consideration. Journal of Cleaner Production, 192, 293–305. https://doi.org/10.1016/j.jclepro.2018.04.251