Published at : 30 Sep 2026
Volume : IJtech
Vol 17, No 5 (2026)
DOI : https://doi.org/10.14716/ijtech.v17i5.8191
| Zhijing Zhou | 1. College of Mechanical and Electrical Engineering, Wenzhou University, Wenzhou 325035, China 2. Zhejiang Haozhonghao Health Products Co., Ltd, Wenzhou, 325400, China |
| Jiajia Wang | College of Mechanical and Electrical Engineering, Wenzhou University, Wenzhou 325035, China |
| Gang Yin | Zhejiang Haozhonghao Health Products Co., Ltd, Wenzhou, 325400, China |
| Jiachen Zhou | Academy of Interdisciplinary Studies, Hong Kong University of Science and Technology, Hong Kong, 999077, China |
| Yongxiang Zhang | Zhejiang Haozhonghao Health Products Co., Ltd, Wenzhou, 325400, China |
| Jian Guo | College of Mechanical and Electrical Engineering, Wenzhou University, Wenzhou 325035, China |
| Junhao Zhang | College of Mechanical and Electrical Engineering, Wenzhou University, Wenzhou 325035, China |
Mixed-model assembly lines (MMAL) are increasingly replacing traditional single-type lines to meet mass customization needs as product demand diversifies. This paper addresses the series-parallel production line sequencing problem (SPPLSP) within MMAL, addressing the complexity introduced by fluctuating orders, asynchronous machinery, and resource imbalances. This study introduces a deep reinforcement learning (DRL) approach to address these challenges, focusing on the impact of machine selection and job sequencing on order tardiness and energy consumption during transitions. Initially, a mixed-integer linear programming model (MILPM) is expressed with constraints, such as conditions related to serial-parallel production lines and machines, to make a clear problem description. Subsequently, a double deep Q-network (DDQN) algorithm is employed to solve the SPPLSP, utilizing an innovative state and action space representation to maintain dynamic balance across serial and parallel lines and adapt to changes in machine and workpiece data. Additionally, an improved agent reward function integrates a multi-agent hierarchical distributed scheduling structure to approximate the actual objectives. Finally, a simulation-based numerical study, grounded in real-world case scenarios, validates the performance of the proposed method across various scales of problems. The numerical results reveal that the DDQN method consistently achieves faster convergence and superior stability values across all test instances and yields higher-quality Pareto frontier solutions in 49/50 test instances.
Double deep Q-network; Mixed-model assembly line; Multi-agent; Series-parallel production line sequencing
| Filename | Description |
|---|---|
| R3-IE-8191-20260828151112.docx | --- |
Alla, K. R., & Thangarasu, G. (2023). Supply chain management of
drug products in blockchain using reinforcement learning. International Journal
of Technology, 14 (6), 1256–1265. https://doi.org/10.14716/ijtech.v14i6.6626
Azadeh, A., Shoja, B. M., Ghanei, S.,
& Sheikhalishahi, M. (2015). A multi-objective optimization problem for
multi-state series-parallel systems: A two-stage flow-shop manufacturing
system. Reliability Engineering & System Safety, 136, 62–74. https://doi.org/10.1016/j.ress.2014.11.009
Chen, C., Xia, B., Zhou, B., & Xi, L.
(2015). A reinforcement learning based approach for a multiple-load carrier
scheduling problem. Journal of Intelligent Manufacturing, 26, 1233–1245. https://doi.org/10.1007/s10845-013-0852-9
Du, Y., Li, J., Li, C., & Duan, P.
(2024). A reinforcement learning approach for flexible job shop scheduling
problem with crane transportation and setup times. IEEE Transactions on Neural
Networks and Learning Systems, 35 (4), 5695–5709. https://doi.org/10.1109/TNNLS.2022.3208942
Farouk, M. M., Chin, C. G., & Roslee,
M. (2026). Hybrid reinforcement learning-enabled scheduler for 5G burst
traffic. International Journal of Technology, 17 (3), 866–882. https://doi.org/10.14716/ijtech.v17i3.8208
Fuqing, Z., & Yu, J. (2026). A
multi-agent reinforcement learning with multi-task learning and graph edge
attention networks for dynamic flexible job shop scheduling. Applied Soft
Computing, 187, 114277. https://doi.org/10.1016/j.asoc.2025.114277
Gao, K., Huang, Y., Sadollah, A., & Wang,
A. (2020). A review of energy-efficient scheduling in intelligent production
systems. Complex & Intelligent Systems, 6, 237–249. https://doi.org/10.1007/s40747-019-00122-6
Gholizadeh, H., Chaleshigar, M., &
Fazlollahtabar, H. (2022). Robust optimization of uncertainty-based preventive
maintenance model for scheduling series–parallel production systems (real case:
Disposable appliances production). ISA Transactions, 128, 54–67. https://doi.org/10.1016/j.isatra.2021.11.041
Han, X., Liu, H., Sun, F., & Zhang,
X. (2019). Active object detection with multistep action prediction using deep
q-network. IEEE Transactions on Industrial Informatics, 15 (6), 3723–3731. https://doi.org/10.1109/TII.2019.2890849
Hu, H., Jia, X., He, Q., Fu, S., & Liu,
K. (2020). Deep reinforcement learning based AGVs real-time scheduling with
mixed rule for flexible shop floor in industry 4.0. Computers & Industrial
Engineering, 149, 106749. https://doi.org/10.1016/j.cie.2020.106749
Jing, X., Yao, X., Liu, M., & Zhou,
J. (2024). Multi-agent reinforcement learning based on graph convolutional
network for flexible job shop scheduling. Journal of Intelligent Manufacturing,
35 (1), 75–93. https://doi.org/10.1007/s10845-022-02037-5
Kayhan, B. M., & Yildiz, G. (2023).
Reinforcement learning applications to machine scheduling problems: A
comprehensive literature review. Journal of Intelligent Manufacturing, 34 (3),
905–929. https://doi.org/10.1007/s10845-021-01847-3
Khadivi, M., Charter, T., Yaghoubi, M.,
Jalayer, M., Ahang, M., Shojaeinasab, A., & Najjaran, H. (2025). Deep
reinforcement learning for machine scheduling: Methodology, the
state-of-the-art, and future directions. Computers & Industrial
Engineering, 200, 110856. https://doi.org/10.1016/j.cie.2025.110856
Li, L., Yang, X., Yang, S., & Xu, X.
(2023). Optimization of oxygen system scheduling in hybrid action space based
on deep reinforcement learning. Computers & Chemical Engineering, 171,
108168. https://doi.org/10.1016/j.compchemeng.2023.108168
Li, Y. F., & Peng, R. (2014).
Availability modeling and optimization of dynamic multi-state series–parallel
systems with random reconfiguration. Reliability Engineering & System
Safety, 127, 47–57. https://doi.org/10.1016/j.ress.2014.03.005
Liu, X., Yang, X., & Lei, M. (2021).
Optimisation of mixed-model assembly line balancing problem under uncertain
demand. Journal of Manufacturing Systems, 59, 214–227. https://doi.org/10.1016/j.jmsy.2021.02.019
Mao, H., Liu, Z., & Qiu, C. (2023).
Adaptive disassembly sequence planning for VR maintenance training via deep
reinforcement learning. The International Journal of Advanced Manufacturing
Technology, 124, 3039–3048. https://doi.org/10.1007/s00170-021-08290-x
Mosadegh, H., Fatemi Ghomi, S. M. T.,
& Süer, G. A. (2017). Heuristic approaches for mixed-model sequencing
problem with stochastic processing times. International Journal of Production
Research, 55 (10), 2857–2880. https://doi.org/10.1080/00207543.2016.1223897
Mosadegh, H., Ghomi, S. M. T. F., &
Süer, G. A. (2020). Stochastic mixed-model assembly line sequencing problem:
Mathematical modeling and q-learning based simulated annealing
hyper-heuristics. European Journal of Operational Research, 282 (2), 530–544. https://doi.org/10.1016/j.ejor.2019.09.021
Nourelfath, M., & Yalaoui, F. (2012).
Integrated load distribution and production planning in series-parallel
multi-state systems with failure rate depending on load. Reliability
Engineering & System Safety, 106, 138–145. https://doi.org/10.1016/j.ress.2012.06.006
Paeng, B., Park, I. B., & Park, J.
(2021). Deep reinforcement learning for minimizing tardiness in parallel
machine scheduling with sequence dependent family setups. IEEE Access, 9,
101390–101401. https://doi.org/10.1109/ACCESS.2021.3097254
Rai, R., Tiwari, M. K., Ivanov, D., &
Dolgui, A. (2021). Machine learning in manufacturing and industry 4.0
applications. International Journal of Production Research, 59 (16), 4773–4778.
https://doi.org/10.1109/OJIES.2024.3431240
Sikora, C. G. S. (2024). Balancing
mixed-model assembly lines for random sequences. European Journal of
Operational Research, 314 (2), 597–611. https://doi.org/10.1016/j.ejor.2023.10.008
Taube, F., & Minner, S. (2018).
Resequencing mixed-model assembly lines with restoration to customer orders. Omega, 78, 99–111. https://doi.org/10.1016/j.omega.2017.11.006
Van Hasselt, H., Guez, A., & Silver, D. (2016). Deep
reinforcement learning with double q-learning. Proceedings of the AAAI
Conference on Artificial Intelligence, 30 (1). https://doi.org/10.48550/arXiv.1509.06461
Wang, J., Wang, W., Hu, X., Qiu, L., & Zang, H. (2024). Black-winged
kite algorithm: A nature-inspired meta-heuristic for solving benchmark
functions and engineering problems. Artificial Intelligence Review, 57 (4), 98.
https://doi.org/10.1007/s10462-024-10723-4
Wang, K., Li, X., Gao, L., Li, P., &
Gupta, S. M. (2021). A genetic simulated annealing algorithm for parallel
partial disassembly line balancing problem. Applied Soft Computing, 107,
107404. https://doi.org/10.1016/j.asoc.2021.107404
Wei, Q., Li, Y., Zhang, J., & Wang,
F. Y. (2022). VGN: Value decomposition with graph attention networks for
multiagent reinforcement learning. IEEE Transactions on Neural Networks and
Learning Systems, 35 (1), 182–195. https://doi.org/10.1109/TNNLS.2022.3172572
Wen, G., Fu, J., Dai, P., & Zhou, J. (2021). DTDE: A new
cooperative multi-agent reinforcement learning framework. The Innovation, 2
(4). https://doi.org/10.1016/j.xinn.2021.100162
Xi, S., Smith, J. M. G., Chen, Q., Mao,
N., & Zhang, H. (2022). Simultaneous machine selection and buffer
allocation in large unbalanced series-parallel production lines. International
Journal of Production Research, 60 (7), 2103–2125. https://doi.org/10.1080/00207543.2021.1884306
Yamsani, N., & Reddy, P. C. (2026).
EdgeSched-DQN: An intelligent deep reinforcement learning-based framework for
optimized task scheduling in edge-cloud environments. Array, 29, 100645. https://doi.org/10.1016/j.array.2025.100645
Yang, W., & Cheng, W. (2020).
Modelling and solving mixed-model two-sided assembly line balancing problem
with sequence-dependent setup time. International Journal of Production
Research, 58 (21), 6638–6659. https://doi.org/10.1080/00207543.2019.1683255
Yelles-Chaouche, A. R., Gurevsky, E.,
Brahimi, N., & Dolgui, A. (2022). Minimizing task reassignments under
balancing multi-product reconfigurable manufacturing lines. Computers &
Industrial Engineering, 173, 108660. https://doi.org/10.1016/j.cie.2022.108660
Yuan, M., Li, Y., Zhang, L., & Pei,
F. (2021). Research on intelligent workshop resource scheduling method based on
improved NSGA-II algorithm. Robotics and Computer-Integrated Manufacturing, 71,
102141. https://doi.org/10.1016/j.rcim.2021.102141
Zhang, B., Xu, L., & Zhang, J.
(2020a). A multi-objective cellular genetic algorithm for energy-oriented
balancing and sequencing problem of mixed-model assembly line. Journal of
Cleaner Production, 244, 118845. https://doi.org/10.1016/j.jclepro.2019.118845
Zhang, K., He, F., Zhang, Z., Lin, X.,
& Li, M. (2020b). Multi-vehicle routing problems with soft time windows: A
multi-agent reinforcement learning approach. Transportation Research Part C:
Emerging Technologies, 121, 102861. https://doi.org/10.1016/j.trc.2020.102861
Zhang, M., Liu, Y., Chen, H., & Cai,
W. (2025). Double DQN-based efficient quality of service routing protocol in
Internet of Underwater Things with mobile nodes. Ad Hoc Networks, 175, 103856. https://doi.org/10.1016/j.adhoc.2025.103856
Zhao, C., & Melkote, S. N. (2024).
Learning the manufacturing capabilities of machining and finishing processes
using a deep neural network model. Journal of Intelligent Manufacturing, 35
(4), 1845–1865. https://doi.org/10.1007/s10845-023-02134-z
Zhao, M., Guo, X., Zhang, X., Fang, Y.,
& Ou, Y. (2020). Aspw-drl: Assembly sequence planning for workpieces via a
deep reinforcement learning approach. Assembly Automation, 40 (1), 65–75. https://doi.org/10.1108/AA-11-2018-0211
Zhou, B. H., & Shen, C. Y. (2018). Multi-objective optimization of material delivery for mixed model assembly lines with energy consideration. Journal of Cleaner Production, 192, 293–305. https://doi.org/10.1016/j.jclepro.2018.04.251