Quantum reinforcement learning in continuous action space
Institute of Fundamental and Frontier Sciences, University of Electronic Science and Technology of China, Chengdu, 610051, China
| Published: | 2025-03-12, volume 9, page 1660 |
| Editor: | Roohollah Ghobadi |
| Eprint: | arXiv:2012.10711v5 |
| Doi: | https://doi.org/10.22331/q-2025-03-12-1660 |
| Citation: | Quantum 9, 1660 (2025). |
Find this paper interesting or want to discuss? Scite or leave a comment on SciRate.
Abstract
Quantum reinforcement learning (QRL) is a promising paradigm for near-term quantum devices. While existing QRL methods have shown success in discrete action spaces, extending these techniques to continuous domains is challenging due to the curse of dimensionality introduced by discretization. To overcome this limitation, we introduce a quantum Deep Deterministic Policy Gradient (DDPG) algorithm that efficiently addresses both classical and quantum sequential decision problems in continuous action spaces. Moreover, our approach facilitates single-shot quantum state generation: a one-time optimization produces a model that outputs the control sequence required to drive a fixed initial state to any desired target state. In contrast, conventional quantum control methods demand separate optimization for each target state. We demonstrate the effectiveness of our method through simulations and discuss its potential applications in quantum control.

Featured image: The quantum reinforcement learning model. Each iterative step can be described by the following loop: (1) at step $t$, the agent receives $|s_t\rangle$ and generates the action parameter $\boldsymbol{\theta_t}$ according to the current policy; (2) the agent generates $|s_{t+1}\rangle \equiv U_a(\boldsymbol{\theta_t}) |s_t\rangle$; (3) based on $|s_t\rangle$ and $|s_{t+1}\rangle$, a reward $r_{t+1}$ is calculated and fed back to the agent, together with $|s_{t+1}\rangle$; (4) based on $|s_{t+1}\rangle$ and $r_{t+1}$, the policy is updated and then used to generate $\boldsymbol{\theta_{t+1}}$.
Popular summary
To overcome this, we developed a quantum Deep Deterministic Policy Gradient (DDPG) algorithm that perform well in problems with continuous action spaces. In simple terms, our method uses quantum neural networks (implemented via variational quantum circuits) to learn the best sequence of quantum operations (or control pulses) needed to drive a quantum system from any starting state to any target state. One of the key breakthroughs is that after a single round of training, the algorithm can generate the proper control sequence for any desired target state—no need to start from scratch every time a new state is required. The paper demonstrates the effectiveness of this approach through simulations on one-qubit and two-qubit systems and applies it to tackle quantum eigenvalue problems, which are central to understanding the properties of quantum systems.
Overall, this research introduces a quantum reinforcement learning framework that combines ideas from artificial intelligence and quantum physics to control quantum systems more effectively. For future work, it is interesting to explore how RL could play a pivotal role in optimizing quantum experiments, automating quantum circuit design and improving quantum error suppression technique.
► BibTeX data
► References
[1] Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. The MIT Press, second edition, 2018. URL http://incompleteideas.net/book/the-book-2nd.html.
http://incompleteideas.net/book/the-book-2nd.html
[2] David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. Mastering the game of go without human knowledge. Nature(London), 550 (7676): 354–359, 2017. 10.1038/nature24270. URL https://doi.org/10.1038/nature24270.
https://doi.org/10.1038/nature24270
[3] Mnih Volodymyr, Kavukcuoglu Koray, Silver David, Graves Alex, Antonoglou Ioannis, Wierstra Daan, and Riedmiller Martin. Playing atari with deep reinforcement learning. 2013. 10.48550/ARXIV.1312.5602. URL http://arxiv.org/abs/1312.5602.
https://doi.org/10.48550/ARXIV.1312.5602
arXiv:1312.5602
[4] David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. Mastering the game of go with deep neural networks and tree search. Nature(London), 529 (7587): 484–489, 2016. 10.1038/nature16961. URL https://doi.org/10.1038/nature16961.
https://doi.org/10.1038/nature16961
[5] Jan Peters, Sethu Vijayakumar, and Stefan Schaal. Reinforcement learning for humanoid robotics. In Proceedings of the third IEEE-RAS international conference on humanoid robots, pages 1–20, 2003. 10.1109/LARS/SBR/WRE51543.2020.9307084. URL https://ieeexplore.ieee.org/document/9307084.
https://doi.org/10.1109/LARS/SBR/WRE51543.2020.9307084
https://ieeexplore.ieee.org/document/9307084
[6] Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel. Benchmarking deep reinforcement learning for continuous control. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML'16, page 1329–1338, 2016. 10.5555/3045390.3045531. URL https://dl.acm.org/doi/10.5555/3045390.3045531.
https://doi.org/10.5555/3045390.3045531
[7] Hendrik Poulsen Nautrup, Nicolas Delfosse, Vedran Dunjko, Hans J. Briegel, and Nicolai Friis. Optimizing Quantum Error Correction Codes with Reinforcement Learning. Quantum, 3: 215, December 2019. ISSN 2521-327X. 10.22331/q-2019-12-16-215. URL https://doi.org/10.22331/q-2019-12-16-215.
https://doi.org/10.22331/q-2019-12-16-215
[8] Philip Andreasson, Joel Johansson, Simon Liljestrand, and Mats Granath. Quantum error correction for the toric code using deep reinforcement learning. Quantum, 3: 183, September 2019. ISSN 2521-327X. 10.22331/q-2019-09-02-183. URL https://doi.org/10.22331/q-2019-09-02-183.
https://doi.org/10.22331/q-2019-09-02-183
[9] Pantita Palittapongarnpim, Peter Wittek, Ehsan Zahedinejad, Shakib Vedaie, and Barry C. Sanders. Learning in quantum control: High-dimensional global optimization for noisy quantum dynamics. Neurocomputing, 268: 116 – 126, 2017. ISSN 0925-2312. https://doi.org/10.1016/j.neucom.2016.12.087. URL http://www.sciencedirect.com/science/article/pii/S0925231217307531.
https://doi.org/10.1016/j.neucom.2016.12.087
http://www.sciencedirect.com/science/article/pii/S0925231217307531
[10] Zheng An and D. L. Zhou. Deep reinforcement learning for quantum gate control. EPL (Europhysics Letters), 126 (6): 60002, jul 2019. 10.1209/0295-5075/126/60002. URL https://doi.org/10.1209/0295-5075/126/60002.
https://doi.org/10.1209/0295-5075/126/60002
[11] Marin Bukov, Alexandre G. R. Day, Dries Sels, Phillip Weinberg, Anatoli Polkovnikov, and Pankaj Mehta. Reinforcement learning in different phases of quantum control. Phys. Rev. X, 8: 031086, Sep 2018. 10.1103/PhysRevX.8.031086. URL https://doi.org/10.1103/PhysRevX.8.031086.
https://doi.org/10.1103/PhysRevX.8.031086
[12] Murphy Yuezhen Niu, Sergio Boixo, Vadim N Smelyanskiy, and Hartmut Neven. Universal quantum control through deep reinforcement learning. npj Quantum Information, 5 (1): 1–8, 2019. 10.1038/s41534-019-0141-3. URL https://doi.org/10.1038/s41534-019-0141-3.
https://doi.org/10.1038/s41534-019-0141-3
[13] Han Xu, Junning Li, Liqiang Liu, Yu Wang, Haidong Yuan, and Xin Wang. Generalizable control for quantum parameter estimation through reinforcement learning. npj Quantum Information, 5 (82): 1–8, 2019. 10.1038/s41534-019-0198-z. URL https://doi.org/10.1038/s41534-019-0198-z.
https://doi.org/10.1038/s41534-019-0198-z
[14] Xiao-Ming Zhang, Zezhu Wei, Raza Asad, Xu-Chen Yang, and Xin Wang. When does reinforcement learning stand out in quantum control? a comparative study on state preparation. npj Quantum Information, 5 (85): 1–7, 2019. 10.1038/s41534-019-0201-8. URL https://doi.org/10.1038/s41534-019-0201-8.
https://doi.org/10.1038/s41534-019-0201-8
[15] Matteo M. Wauters, Emanuele Panizon, Glen B. Mbeng, and Giuseppe E. Santoro. Reinforcement-learning-assisted quantum optimization. Phys. Rev. Research, 2: 033446, Sep 2020. 10.1103/PhysRevResearch.2.033446. URL https://doi.org/10.1103/PhysRevResearch.2.033446.
https://doi.org/10.1103/PhysRevResearch.2.033446
[16] Christopher J. C. H. Watkins. Learning from delayed rewards. PhD thesis, University of Cambridge, 1989. URL https://doi.org/10.1016/0921-8890(95)00026-C.
https://doi.org/10.1016/0921-8890(95)00026-C
[17] Christopher J. C. H. Watkins and Peter Dayan. Q-learning. Machine Learning, 8 (3-4): 279–292, 1992. 10.1007/BF00992698. URL https://doi.org/10.1007/BF00992698.
https://doi.org/10.1007/BF00992698
[18] Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level control through deep reinforcement learning. Nature(London), 518 (7540): 529–533, 2015. 10.1038/nature14236. URL https://doi.org/10.1038/nature14236.
https://doi.org/10.1038/nature14236
[19] Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. 2015. 10.48550/ARXIV.1509.02971. URL https://arxiv.org/abs/1509.02971.
https://doi.org/10.48550/ARXIV.1509.02971
arXiv:1509.02971
[20] Michael A. Nielsen and Isaac L. Chuang. Quantum Computation and Quantum Information. Cambridge University Press, USA, 10th edition, 2011. ISBN 1107002176. https://doi.org/10.1017/CBO9780511976667.
https://doi.org/10.1017/CBO9780511976667
[21] Jacob Biamonte, Peter Wittek, Nicola Pancotti, Patrick Rebentrost, Nathan Wiebe, and Seth Lloyd. Quantum machine learning. Nature(London), 549 (7671): 195–202, 2017. 10.1038/nature23474. URL https://doi.org/10.1038/nature23474.
https://doi.org/10.1038/nature23474
[22] Peter W Shor. Algorithms for quantum computation: discrete logarithms and factoring. In Proceedings 35th Annual Symposium on Foundations of Computer Science, pages 124–134, 1994. 10.1109/SFCS.1994.365700. URL https://ieeexplore.ieee.org/document/365700.
https://doi.org/10.1109/SFCS.1994.365700
https://ieeexplore.ieee.org/document/365700
[23] Lov Kumar Grover. Quantum mechanics helps in searching for a needle in a haystack. Phys. Rev. Lett., 79: 325–328, Jul 1997. 10.1103/PhysRevLett.79.325. URL https://doi.org/10.1103/PhysRevLett.79.325.
https://doi.org/10.1103/PhysRevLett.79.325
[24] Nathan Wiebe, Daniel Braun, and Seth Lloyd. Quantum algorithm for data fitting. Phys. Rev. Lett., 109: 050505, Aug 2012. 10.1103/PhysRevLett.109.050505. URL https://doi.org/10.1103/PhysRevLett.109.050505.
https://doi.org/10.1103/PhysRevLett.109.050505
[25] Patrick Rebentrost, Masoud Mohseni, and Seth Lloyd. Quantum support vector machine for big data classification. Phys. Rev. Lett., 113: 130503, Sep 2014. 10.1103/PhysRevLett.113.130503. URL https://doi.org/10.1103/PhysRevLett.113.130503.
https://doi.org/10.1103/PhysRevLett.113.130503
[26] Seth Lloyd and Christian Weedbrook. Quantum generative adversarial learning. Phys. Rev. Lett., 121: 040502, Jul 2018. 10.1103/PhysRevLett.121.040502. URL https://doi.org/10.1103/PhysRevLett.121.040502.
https://doi.org/10.1103/PhysRevLett.121.040502
[27] Sankar Das Sarma, Dong-Ling Deng, and Lu-Ming Duan. Machine learning meets quantum physics. Physics Today, 72 (3): 48–54, Mar 2019. ISSN 1945-0699. 10.1063/pt.3.4164. URL http://dx.doi.org/10.1063/PT.3.4164.
https://doi.org/10.1063/pt.3.4164
[28] Seth Lloyd, Masoud Mohseni, and Patrick Rebentrost. Quantum principal component analysis. Nature Physics, 10 (9): 631–633, 2014. 10.1038/nphys3029. URL https://doi.org/10.1038/nphys3029.
https://doi.org/10.1038/nphys3029
[29] Nico Meyer, Christian Ufrecht, Maniraman Periyasamy, Daniel D. Scherer, Axel Plinge, and Christopher Mutschler. A survey on quantum reinforcement learning, 2024. URL https://arxiv.org/abs/2211.03464.
arXiv:2211.03464
[30] Daoyi Dong, Chunlin Chen, Hanxiong Li, and Tzyh-Jong Tarn. Quantum reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38 (5): 1207–1220, 2008. 10.1109/TSMCB.2008.925743. URL https://ieeexplore.ieee.org/document/4579244.
https://doi.org/10.1109/TSMCB.2008.925743
https://ieeexplore.ieee.org/document/4579244
[31] Vedran Dunjko, Jacob M. Taylor, and Hans J. Briegel. Quantum-enhanced machine learning. Phys. Rev. Lett., 117: 130501, Sep 2016. 10.1103/PhysRevLett.117.130501. URL https://doi.org/10.1103/PhysRevLett.117.130501.
https://doi.org/10.1103/PhysRevLett.117.130501
[32] Giuseppe Davide Paparo, Vedran Dunjko, Adi Makmal, Miguel Angel Martin-Delgado, and Hans J. Briegel. Quantum speedup for active learning agents. Phys. Rev. X, 4: 031002, Jul 2014a. 10.1103/PhysRevX.4.031002. URL https://doi.org/10.1103/PhysRevX.4.031002.
https://doi.org/10.1103/PhysRevX.4.031002
[33] Vedran Dunjko, Jacob M Taylor, and Hans J Briegel. Advances in quantum reinforcement learning. In 2017 IEEE International Conference on Systems, Man, and Cybernetics (SMC), pages 282–287, 2017a. 10.1109/SMC.2017.8122616. URL https://ieeexplore.ieee.org/document/8122616.
https://doi.org/10.1109/SMC.2017.8122616
https://ieeexplore.ieee.org/document/8122616
[34] Vedran Dunjko and Hans J Briegel. Machine learning & artificial intelligence in the quantum domain: a review of recent progress. Reports on Progress in Physics, 81 (7): 074001, jun 2018. 10.1088/1361-6633/aab406. URL https://doi.org/10.1088/1361-6633/aab406.
https://doi.org/10.1088/1361-6633/aab406
[35] Sofiene Jerbi, Lea M. Trenkwalder, Hendrik Poulsen Nautrup, Hans J. Briegel, and Vedran Dunjko. Quantum enhancements for deep reinforcement learning in large spaces. PRX Quantum, 2: 010328, Feb 2021a. 10.1103/PRXQuantum.2.010328. URL https://doi.org/10.1103/PRXQuantum.2.010328.
https://doi.org/10.1103/PRXQuantum.2.010328
[36] Samuel Yen-Chi Chen, Chao-Han Huck Yang, Jun Qi, Pin-Yu Chen, Xiaoli Ma, and Hsi-Sheng Goan. Variational quantum circuits for deep reinforcement learning. IEEE Access, 8: 141007–141024, 2020. 10.1109/ACCESS.2020.3010470. URL https://ieeexplore.ieee.org/abstract/document/9144562.
https://doi.org/10.1109/ACCESS.2020.3010470
https://ieeexplore.ieee.org/abstract/document/9144562
[37] Owen Lockwood and Mei Si. Reinforcement learning with quantum variational circuits. In Proceedings of the Sixteenth AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, AIIDE'20. AAAI Press, 2020. ISBN 978-1-57735-849-7. URL https://dl.acm.org/doi/abs/10.5555/3505464.3505499.
https://dl.acm.org/doi/abs/10.5555/3505464.3505499
[38] Andrea Skolik, Sofiene Jerbi, and Vedran Dunjko. Quantum agents in the Gym: a variational quantum algorithm for deep Q-learning. Quantum, 6: 720, May 2022. ISSN 2521-327X. 10.22331/q-2022-05-24-720. URL https://doi.org/10.22331/q-2022-05-24-720.
https://doi.org/10.22331/q-2022-05-24-720
[39] Owen Lockwood and Mei Si. Playing atari with hybrid quantum-classical reinforcement learning. In Luca Bertinetto, João F. Henriques, Samuel Albanie, Michela Paganini, and Gül Varol, editors, NeurIPS 2020 Workshop on Pre-registration in Machine Learning, volume 148 of Proceedings of Machine Learning Research, pages 285–301. PMLR, 11 Dec 2021. URL https://proceedings.mlr.press/v148/lockwood21a.html.
https://proceedings.mlr.press/v148/lockwood21a.html
[40] Samuel Yen-Chi Chen. Quantum Deep Q-Learning with Distributed Prioritized Experience Replay . In 2023 IEEE International Conference on Quantum Computing and Engineering (QCE), pages 31–35, Los Alamitos, CA, USA, September 2023. IEEE Computer Society. 10.1109/QCE57702.2023.10180. URL https://doi.ieeecomputersociety.org/10.1109/QCE57702.2023.10180.
https://doi.org/10.1109/QCE57702.2023.10180
[41] Sofiene Jerbi, Casper Gyurik, Simon Marshall, Hans Briegel, and Vedran Dunjko. Parametrized quantum policies for reinforcement learning. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 28362–28375. Curran Associates, Inc., 2021b. URL https://proceedings.neurips.cc/paper/2021/file/eec96a7f788e88184c0e713456026f3f-Paper.pdf.
https://proceedings.neurips.cc/paper/2021/file/eec96a7f788e88184c0e713456026f3f-Paper.pdf
[42] Nico Meyer, Daniel Scherer, Axel Plinge, Christopher Mutschler, and Michael Hartmann. Quantum policy gradient algorithm with optimized action decoding. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 24592–24613. PMLR, 23–29 Jul 2023a. URL https://proceedings.mlr.press/v202/meyer23a.html.
https://proceedings.mlr.press/v202/meyer23a.html
[43] Nico Meyer, Daniel D. Scherer, Axel Plinge, Christopher Mutschler, and Michael J. Hartmann. Quantum Natural Policy Gradients: Towards Sample-Efficient Reinforcement Learning . In 2023 IEEE International Conference on Quantum Computing and Engineering (QCE), pages 36–41, Los Alamitos, CA, USA, September 2023b. IEEE Computer Society. 10.1109/QCE57702.2023.10181. URL https://doi.ieeecomputersociety.org/10.1109/QCE57702.2023.10181.
https://doi.org/10.1109/QCE57702.2023.10181
[44] André Sequeira, Luis Paulo Santos, and Luis Soares Barbosa. Policy gradients using variational quantum circuits. Quantum Machine Intelligence, 5 (1): 18, 2023. ISSN 2524-4914. 10.1007/s42484-023-00101-8. URL https://doi.org/10.1007/s42484-023-00101-8.
https://doi.org/10.1007/s42484-023-00101-8
[45] Valeria Saggio, Beate E Asenbeck, Arne Hamann, Teodor Strömberg, Peter Schiansky, Vedran Dunjko, Nicolai Friis, Nicholas C Harris, Michael Hochberg, Dirk Englund, et al. Experimental quantum speed-up in reinforcement learning agents. Nature, 591 (7849): 229–233, 2021. 10.1038/s41586-021-03242-7. URL https://doi.org/10.1038/s41586-021-03242-7.
https://doi.org/10.1038/s41586-021-03242-7
[46] Vedran Dunjko, Jacob M. Taylor, and Hans J. Briegel. Advances in quantum reinforcement learning. In 2017 IEEE International Conference on Systems, Man, and Cybernetics (SMC), page 282–287. IEEE Press, 2017b. 10.1109/SMC.2017.8122616. URL https://doi.org/10.1109/SMC.2017.8122616.
https://doi.org/10.1109/SMC.2017.8122616
[47] El Amine Cherrat, Iordanis Kerenidis, and Anupam Prakash. Quantum reinforcement learning via policy iteration. Quantum Machine Intelligence, 5 (2): 30, 2023. ISSN 2524-4914. 10.1007/s42484-023-00116-1. URL https://doi.org/10.1007/s42484-023-00116-1.
https://doi.org/10.1007/s42484-023-00116-1
[48] Daochen Wang, Aarthi Sundaram, Robin Kothari, Ashish Kapoor, and Martin Roetteler. Quantum algorithms for reinforcement learning with a generative model. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 10916–10926. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/wang21w.html.
https://proceedings.mlr.press/v139/wang21w.html
[49] Hans J. Briegel and Gemma De las Cuevas. Projective simulation for artificial intelligence. Scientific Reports, 2 (1): 400, 2012. ISSN 2045-2322. 10.1038/srep00400. URL https://doi.org/10.1038/srep00400.
https://doi.org/10.1038/srep00400
[50] Alexey A. Melnikov, Adi Makmal, Vedran Dunjko, and Hans J. Briegel. Projective simulation with generalization. Scientific Reports, 7 (1): 14430, 2017. ISSN 2045-2322. 10.1038/s41598-017-14740-y. URL https://doi.org/10.1038/s41598-017-14740-y.
https://doi.org/10.1038/s41598-017-14740-y
[51] Giuseppe Davide Paparo, Vedran Dunjko, Adi Makmal, Miguel Angel Martin-Delgado, and Hans J. Briegel. Quantum speedup for active learning agents. Phys. Rev. X, 4: 031002, Jul 2014b. 10.1103/PhysRevX.4.031002. URL https://doi.org/10.1103/PhysRevX.4.031002.
https://doi.org/10.1103/PhysRevX.4.031002
[52] V Dunjko, N Friis, and H J Briegel. Quantum-enhanced deliberation of learning agents using trapped ions. New Journal of Physics, 17 (2): 023006, jan 2015. 10.1088/1367-2630/17/2/023006. URL https://dx.doi.org/10.1088/1367-2630/17/2/023006.
https://doi.org/10.1088/1367-2630/17/2/023006
[53] Th Sriarunothai, S Wölk, G S Giri, N Friis, V Dunjko, H J Briegel, and Ch Wunderlich. Speeding-up the decision making of a learning agent using an ion trap quantum processor. Quantum Science and Technology, 4 (1): 015014, dec 2018. 10.1088/2058-9565/aaef5e. URL https://dx.doi.org/10.1088/2058-9565/aaef5e.
https://doi.org/10.1088/2058-9565/aaef5e
[54] Martijn Van Otterlo and Marco Wiering. Reinforcement learning and markov decision processes. In Reinforcement Learning, pages 3–42. Springer Berlin Heidelberg, 2012. 10.1007/978-3-642-27645-3_1. URL https://doi.org/10.1007/978-3-642-27645-3_1.
https://doi.org/10.1007/978-3-642-27645-3_1
[55] Gavin A Rummery and Mahesan Niranjan. On-line q-learning using connectionist systems. Technical report, 1994. URL http://mi.eng.cam.ac.uk/reports/svr-ftp/auto-pdf/rummery_tr166.pdf.
http://mi.eng.cam.ac.uk/reports/svr-ftp/auto-pdf/rummery_tr166.pdf
[56] Sham M Kakade. A natural policy gradient. In Advances in Neural Information Processing Systems, volume 14, pages 1531–1538. MIT Press, 2002. URL https://proceedings.neurips.cc/paper/2001/file/4b86abe48d358ecf194c56c69108433e-Paper.pdf.
https://proceedings.neurips.cc/paper/2001/file/4b86abe48d358ecf194c56c69108433e-Paper.pdf
[57] Alberto Peruzzo, Jarrod McClean, Peter Shadbolt, Man-Hong Yung, Xiao-Qi Zhou, Peter J. Love, Alán Aspuru-Guzik, and Jeremy L. O'Brien. A variational eigenvalue solver on a photonic quantum processor. Nature communications, 5: 4213, 2014. 10.1038/ncomms5213. URL https://doi.org/10.1038/ncomms5213.
https://doi.org/10.1038/ncomms5213
[58] Edward Farhi, Jeffrey Goldstone, and Sam Gutmann. A quantum approximate optimization algorithm. 10.48550/ARXIV.1411.4028. URL https://arxiv.org/abs/1411.4028.
https://doi.org/10.48550/ARXIV.1411.4028
arXiv:1411.4028
[59] Marcello Benedetti, Erika Lloyd, Stefan Sack, and Mattia Fiorentini. Parameterized quantum circuits as machine learning models. Quantum Science and Technology, 4 (4): 043001, nov 2019. 10.1088/2058-9565/ab4eb5. URL https://doi.org/10.1088/2058-9565/ab4eb5.
https://doi.org/10.1088/2058-9565/ab4eb5
[60] Long-Ji Lin. Reinforcement Learning for Robots Using Neural Networks. PhD thesis, USA, 1992. URL https://dl.acm.org/doi/10.5555/168871.
https://dl.acm.org/doi/10.5555/168871
[61] Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Proceedings of International Conference on Learning Representations, 2015. URL http://arxiv.org/abs/1412.6980.
arXiv:1412.6980
[62] Xiaokai Hou, Guanyu Zhou, Qingyu Li, Shan Jin, and Xiaoting Wang. A universal duplication-free quantum neural network. URL https://arxiv.org/abs/2106.13211.
arXiv:2106.13211
[63] Dorit Aharonov and Tomer Naveh. Quantum np-a survey. arXiv preprint quant-ph/0210077, 2002. 10.48550/ARXIV.QUANT-PH/0210077. URL https://arxiv.org/abs/quant-ph/0210077.
https://doi.org/10.48550/ARXIV.QUANT-PH/0210077
arXiv:quant-ph/0210077
[64] John Watrous. Quantum computational complexity. 10.48550/ARXIV.0804.3401. URL https://arxiv.org/abs/0804.3401.
https://doi.org/10.48550/ARXIV.0804.3401
arXiv:0804.3401
[65] Sevag Gharibian, Yichen Huang, Zeph Landau, and Seung Woo Shin. Quantum hamiltonian complexity. pages 7174–7201, 2009. 10.1007/978-0-387-30440-3_428. URL https://doi.org/10.1007/978-0-387-30440-3_428.
https://doi.org/10.1007/978-0-387-30440-3_428
[66] Julia Kempe, Alexei Kitaev, and Oded Regev. The complexity of the local hamiltonian problem. In FSTTCS 2004: Foundations of Software Technology and Theoretical Computer Science, volume 35, pages 372–383, 2006. URL https://doi.org/10.1007/978-3-540-30538-5_31.
https://doi.org/10.1007/978-3-540-30538-5_31
[67] Daniel S. Abrams and Seth Lloyd. Quantum algorithm providing exponential speed increase for finding eigenvalues and eigenvectors. Phys. Rev. Lett., 83: 5162–5165, Dec 1999. 10.1103/PhysRevLett.83.5162. URL https://doi.org/10.1103/PhysRevLett.83.5162.
https://doi.org/10.1103/PhysRevLett.83.5162
[68] Navin Khaneja, Timo Reiss, Cindie Kehlet, Thomas Schulte-Herbrüggen, and Steffen J. Glaser. Optimal control of coupled spin dynamics: design of nmr pulse sequences by gradient ascent algorithms. Journal of Magnetic Resonance, 172 (2): 296–305, 2005. ISSN 1090-7807. https://doi.org/10.1016/j.jmr.2004.11.004. URL https://www.sciencedirect.com/science/article/pii/S1090780704003696.
https://doi.org/10.1016/j.jmr.2004.11.004
https://www.sciencedirect.com/science/article/pii/S1090780704003696
[69] Warwick Masson, Pravesh Ranchod, and George Konidaris. Reinforcement learning with parameterized actions. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, AAAI'16, page 1934–1940. AAAI Press, 2016. URL https://ojs.aaai.org/index.php/AAAI/article/view/10226.
https://ojs.aaai.org/index.php/AAAI/article/view/10226
[70] Kenji Doya. Reinforcement learning in continuous time and space. Neural Computation, 12 (1): 219–245, 2000. 10.1162/089976600300015961. URL https://ieeexplore.ieee.org/document/6789455.
https://doi.org/10.1162/089976600300015961
https://ieeexplore.ieee.org/document/6789455
[71] Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. Openai gym. 10.48550/ARXIV.1606.01540. URL https://arxiv.org/abs/1606.01540.
https://doi.org/10.48550/ARXIV.1606.01540
arXiv:1606.01540
Cited by
[1] Amani Lamine, Ameni Mejri, and Bilel Charfi, Lecture Notes in Networks and Systems 1907, 134 (2026) ISBN:978-3-032-21146-0.
[2] Gerardo Iovane, "An Extended Epistemic Framework Beyond Probability for Quantum Information Processing with Applications in Security, Artificial Intelligence, and Financial Computing", Entropy 27 9, 977 (2025).
[3] Abhilash Nelson, R. S. Shaji, and S. V. Ashikaa, "Quantum autoencoder with quantum reinforcement learning for breast cancer detection", The European Physical Journal Plus 140 12, 1200 (2025).
[4] Belkacem Chikhaoui, "GroverAttention: Quantum-Enhanced Attention Mechanism for Effcient Transformers", (2025).
[5] Chi-Sheng Chen and En-Jui Kuo, ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 22422 (2026) ISBN:979-8-3315-6701-9.
[6] G. Sathish Kumar, G. Uma Maheshwari, B. Arun Kumar, A. Andrine Dinola, and P. Jenita, Sustainable Artificial Intelligence-Powered Applications 1 (2026) ISBN:978-3-032-15187-2.
[7] Marco Wiedmann, Maniraman Periyasamy, and Daniel D. Scherer, 2025 IEEE International Conference on Quantum Computing and Engineering (QCE) 1785 (2025) ISBN:979-8-3315-5736-2.
[8] Wenxia Wang, Jinchen Xu, Fudong Liu, Bei Zhou, Weilong Wang, Hanyun Wang, Qiming Du, Benzheng Yuan, Yizhen Huang, Yifan Hou, Xiaodong Ding, and Zheng Shan, "Lightweight Quantum Neural Networks Intelligent Generation", Advanced Quantum Technologies 8 11, e00302 (2025).
[9] Shaojun Wu, Shan Jin, Abolfazl Bayat, and Xiaoting Wang, "Enhancing the reachability of variational quantum algorithms via input-state design", Communications Physics 9 1, 194 (2026).
[10] Eustache Muteba A. and Nikos E. Mastorakis, "Quantum Intelligent Agent of Medical Decision: The Theoretical Machine", DESIGN, CONSTRUCTION, MAINTENANCE 5, 44 (2025).
[11] Georg Kruse, Rodrigo Coelho, Andreas Rosskopf, Robert Wille, and Jeanette Miriam Lorenz, 2025 IEEE International Conference on Quantum Computing and Engineering (QCE) 1640 (2025) ISBN:979-8-3315-5736-2.
[12] Shiiv R S, Yuven Senthilkumar, and V. Karthick, "BB84 a new hope enhanced QKD for secure email communication with additional quantum gates", EPJ Quantum Technology 13 1, 34 (2026).
[13] Oliver Sefrin, Manuel Radons, Lars Simon, and Sabine Wölk, "Quantum reinforcement learning in dynamic environments", Quantum Machine Intelligence 8 1, 58 (2026).
[14] P Ashok, Ravi Gorli, S Lakshmi Sridevi, and Harishchander Anandaram, Reinforcement Learning 295 (2026) ISBN:9783111633596.
[15] Jawaher Kaldari and Saif Al-Kuwari, "Challenge-response quantum reinforcement learning with application to quantum-assisted authentication", Quantum Machine Intelligence 8 2, 85 (2026).
[16] Ahmad Alomari and Sathish A. P. Kumar, "GPA: Grover Policy Agent for Generating Optimal Quantum Sensor Circuits", IEEE Transactions on Artificial Intelligence 6 10, 2722 (2025).
[17] Xuegang Wu and Pinle Qin, "Adaptive Lighting and Thermal Comfort Control Strategies in Digital Twin Classroom via Deep Reinforcement Learning", Electronics 15 4, 873 (2026).
[18] Amin Masoumi and Mert Korkali, 2026 IEEE Texas Power and Energy Conference (TPEC) 1 (2026) ISBN:979-8-3315-5720-1.
[19] Thet Htar Su, Shaswot Shresthamali, and Masaaki Kondo, "Quantum framework for reinforcement learning: Integrating the Markov decision process, quantum arithmetic, and trajectory search", Physical Review A 111 6, 062421 (2025).
[20] Mayank Shekhar Jha, Sameul Yen-Chi Chen, Chetan Kulkarni, and Joongheon Kim, "A comprehensive review on quantum deep neural networks for prognostics and health management: Fundamentals, challenges and opportunities", Engineering Applications of Artificial Intelligence 177, 114991 (2026).
[21] Michael Broughton, Guillaume Verdon, Trevor McCourt, Antonio J. Martinez, Jae Hyeon Yoo, Sergei V. Isakov, Philip Massey, Ramin Halavati, Murphy Yuezhen Niu, Alexander Zlokapa, Evan Peters, Owen Lockwood, Andrea Skolik, Sofiene Jerbi, Vedran Dunjko, Martin Leib, Michael Streif, David Von Dollen, Hongxiang Chen, Shuxiang Cao, Roeland Wiersema, Hsin-Yuan Huang, Jarrod R. McClean, Ryan Babbush, Sergio Boixo, Dave Bacon, Alan K. Ho, Hartmut Neven, and Masoud Mohseni, "TensorFlow Quantum: A Software Framework for Quantum Machine Learning", arXiv:2003.02989, (2020).
[22] En-Jui Kuo, Yao-Lung L. Fang, and Samuel Yen-Chi Chen, "Quantum Architecture Search via Deep Reinforcement Learning", arXiv:2104.07715, (2021).
[23] Nico Meyer, Christian Ufrecht, Maniraman Periyasamy, Daniel D. Scherer, Axel Plinge, and Christopher Mutschler, "A Survey on Quantum Reinforcement Learning", arXiv:2211.03464, (2022).
[24] Samuel Yen-Chi Chen and Shinjae Yoo, "Federated Quantum Machine Learning", Entropy 23 4, 460 (2021).
[25] Marco Pistoia, Syed Farhan Ahmad, Akshay Ajagekar, Alexander Buts, Shouvanik Chakrabarti, Dylan Herman, Shaohan Hu, Andrew Jena, Pierre Minssen, Pradeep Niroula, Arthur Rattew, Yue Sun, and Romina Yalovetzky, "Quantum Machine Learning for Finance", arXiv:2109.04298, (2021).
[26] Andrea Skolik, Sofiene Jerbi, and Vedran Dunjko, "Quantum agents in the Gym: a variational quantum algorithm for deep Q-learning", Quantum 6, 720 (2022).
[27] Sofiene Jerbi, Casper Gyurik, Simon C. Marshall, Hans J. Briegel, and Vedran Dunjko, "Parametrized quantum policies for reinforcement learning", arXiv:2103.05577, (2021).
[28] Samuel Yen-Chi Chen, Tzu-Chieh Wei, Chao Zhang, Haiwang Yu, and Shinjae Yoo, "Hybrid Quantum-Classical Graph Convolutional Network", arXiv:2101.06189, (2021).
[29] William M. Watkins, Samuel Yen-Chi Chen, and Shinjae Yoo, "Quantum machine learning with differential privacy", Scientific Reports 13, 2453 (2023).
[30] Esther Ye and Samuel Yen-Chi Chen, "Quantum Architecture Search via Continual Reinforcement Learning", arXiv:2112.05779, (2021).
[31] Samuel Yen-Chi Chen, Daniel Fry, Amol Deshmukh, Vladimir Rastunkov, and Charlee Stefanski, "Reservoir Computing via Quantum Recurrent Neural Networks", arXiv:2211.02612, (2022).
[32] Samuel Yen-Chi Chen, Chih-Min Huang, Chia-Wei Hsing, Hsi-Sheng Goan, and Ying-Jer Kao, "Variational quantum reinforcement learning via evolutionary optimization", Machine Learning: Science and Technology 3 1, 015025 (2022).
[33] Qingfeng Lan, "Variational Quantum Soft Actor-Critic", arXiv:2112.11921, (2021).
[34] Sofiene Jerbi, Arjan Cornelissen, Māris Ozols, and Vedran Dunjko, "Quantum policy gradient algorithms", arXiv:2212.09328, (2022).
[35] Samuel Yen-Chi Chen, "Asynchronous training of quantum reinforcement learning", arXiv:2301.05096, (2023).
[36] Samuel Yen-Chi Chen and Shinjae Yoo, "Federated Quantum Machine Learning", arXiv:2103.12010, (2021).
[37] Han Xu, Lingna Wang, Haidong Yuan, and Xin Wang, "Generalizable control for multiparameter quantum metrology", Physical Review A 103 4, 042615 (2021).
[38] Nico Meyer, Jakob Murauer, Alexander Popov, Christian Ufrecht, Axel Plinge, Christopher Mutschler, and Daniel D. Scherer, "Warm-Start Variational Quantum Policy Iteration", arXiv:2404.10546, (2024).
[39] Seyed Shakib Vedaie, Archismita Dalal, Eduardo J. Páez, and Barry C. Sanders, "Framework for learning and control in the classical and quantum domains", Annals of Physics 458, 169471 (2023).
[40] Andrea Skolik, Stefano Mangini, Thomas Bäck, Chiara Macchiavello, and Vedran Dunjko, "Robustness of quantum reinforcement learning under hardware errors", EPJ Quantum Technology 10 1, 8 (2023).
[41] Marco Wiedmann, Maniraman Periyasamy, and Daniel D. Scherer, "Fourier Analysis of Variational Quantum Circuits for Supervised Learning", arXiv:2411.03450, (2024).
[42] Dániel T. R. Nagy, Csaba Czabán, Bence Bakó, Péter Hága, Zsófia Kallus, and Zoltán Zimborás, "Hybrid Quantum-Classical Reinforcement Learning in Latent Observation Spaces", arXiv:2410.18284, (2024).
[43] Hailan Ma, Bo Qi, Ian R. Petersen, Re-Bing Wu, Herschel Rabitz, and Daoyi Dong, "Machine Learning for Estimation and Control of Quantum Systems", arXiv:2503.03164, (2025).
[44] Georg Kruse, Theodora-Augustina Dragan, Robert Wille, and Jeanette Miriam Lorenz, "Variational Quantum Circuit Design for Quantum Reinforcement Learning on Continuous Environments", arXiv:2312.13798, (2023).
[45] Nico Meyer, Daniel D. Scherer, Axel Plinge, Christopher Mutschler, and Michael J. Hartmann, "Quantum Policy Gradient Algorithm with Optimized Action Decoding", arXiv:2212.06663, (2022).
[46] Xianchao Zhu and Xiaokai Hou, "Quantum architecture search via truly proximal policy optimization", Scientific Reports 13, 5157 (2023).
[47] Nico Meyer, Julian Berberich, Christopher Mutschler, and Daniel D. Scherer, "Robustness and Generalization in Quantum Reinforcement Learning via Lipschitz Regularization", arXiv:2410.21117, (2024).
[48] Samuel Yen-Chi Chen, "Quantum deep recurrent reinforcement learning", arXiv:2210.14876, (2022).
[49] Samuel Yen-Chi Chen, Chih-Min Huang, Chia-Wei Hsing, and Ying-Jer Kao, "An end-to-end trainable hybrid classical-quantum classifier", arXiv:2102.02416, (2021).
[50] Yu-Xin Jin, Hong-Ze Xu, Zheng-An Wang, Wei-Feng Zhuang, Kai-Xuan Huang, Yun-Hao Shi, Wei-Guo Ma, Tian-Ming Li, Chi-Tong Chen, Kai Xu, Yu-Long Feng, Pei Liu, Mo Chen, Shang-Shu Li, Zhi-Peng Yang, Chen Qian, Yun-Heng Ma, Xiao Xiao, Peng Qian, Yanwu Gu, Xu-Dan Chai, Ya-Nan Pu, Yi-Peng Zhang, Shi-Jie Wei, Jin-Feng Zeng, Hang Li, Gui-Lu Long, Yirong Jin, Haifeng Yu, Heng Fan, Dong E. Liu, and Meng-Jun Hu, "Quafu-RL: The cloud quantum computers based quantum reinforcement learning", Chinese Physics B 33 5, 050301 (2024).
[51] Dániel Nagy, Zsolt Tabi, Péter Hága, Zsófia Kallus, and Zoltán Zimborás, "Photonic Quantum Policy Learning in OpenAI Gym", arXiv:2108.12926, (2021).
[52] Qibing Xiong, Xiaodong Ding, Yangyang Fei, Xin Zhou, Qiming Du, Congcong Feng, and Zheng Shan, "A hybrid quantum ensemble learning model for malicious code detection", Quantum Science and Technology 9 3, 035021 (2024).
[53] William M Watkins, Samuel Yen-Chi Chen, and Shinjae Yoo, "Quantum machine learning with differential privacy", arXiv:2103.06232, (2021).
[54] Yanxuan Lü, Qing Gao, Jinhu Lü, Maciej Ogorzałek, and Jin Zheng, "A Quantum Convolutional Neural Network for Image Classification", arXiv:2107.03630, (2021).
[55] Maximilian Zorn, Jonas Stein, Maximilian Balthasar Mansky, Philipp Altmann, Michael Kölle, and Claudia Linnhoff-Popien, "Quality Diversity for Variational Quantum Circuit Optimization", arXiv:2504.08459, (2025).
[56] Ahmad Alomari and Sathish A. P. Kumar, "GPA: Grover Policy Agent for Generating Optimal Quantum Sensor Circuits", arXiv:2502.13755, (2025).
[57] BAQIS Quafu Group, "Quafu-RL: The Cloud Quantum Computers based Quantum Reinforcement Learning", arXiv:2305.17966, (2023).
[58] Manuel Guatto, Gian Antonio Susto, and Francesco Ticozzi, "Improving robustness of quantum feedback control with reinforcement learning", Physical Review A 110 1, 012605 (2024).
[59] Tailong Xiao, Jingzheng Huang, Hongjing Li, Jianping Fan, and Guihua Zeng, "Quantum generative adversarial imitation learning", New Journal of Physics 25 3, 033034 (2023).
[60] David M. Bossens, Kishor Bharti, and Jayne Thompson, "Quantum Policy Gradient in Reproducing Kernel Hilbert Space", arXiv:2411.06650, (2024).
[61] Shumin Zhou, Hailan Ma, Sen Kuang, and Daoyi Dong, "Auxiliary Task-based Deep Reinforcement Learning for Quantum Control", arXiv:2302.14312, (2023).
[62] Samuel Yen-Chi Chen, "Quantum deep Q learning with distributed prioritized experience replay", arXiv:2304.09648, (2023).
[63] Xinliang Wei, Xitong Gao, Kejiang Ye, Cheng-Zhong Xu, and Yu Wang, "A Quantum Reinforcement Learning Approach for Joint Resource Allocation and Task Offloading in Mobile Edge Computing", IEEE Transactions on Mobile Computing 24 4, 2580 (2025).
[64] Ahmad Alomari and Sathish A. P. Kumar, "HCQA: Hybrid Classical-Quantum Agent for Generating Optimal Quantum Sensor Circuits", arXiv:2508.21246, (2025).
[65] Aueaphum Aueawatthanaphisut and Nyi Wunna Tun, "Hybrid Quantum-Classical Policy Gradient for Adaptive Control of Cyber-Physical Systems: A Comparative Study of VQC vs. MLP", arXiv:2510.06010, (2025).
[66] Nutkritta Kraipatthanapong, Natthaphat Thathong, Pannita Suksawas, Thanunnut Klunklin, Kritin Vongthonglua, Krit Attahakul, and Aueaphum Aueawatthanaphisut, "Lyapunov-Aware Quantum-Inspired Reinforcement Learning for Continuous-Time Vehicle Control: A Feasibility Study", arXiv:2510.18852, (2025).
[67] Roopa Ravish, Nischal R. Bhat, N. Nandakumar, S. Sagar, Sunil, and Prasad B. Honnavalli, "Optimization of Reinforcement Learning Using Quantum Computation", IEEE Access 12, 179396 (2024).
[68] Yuheng Xie, Yuanchen Hao, Yuefeng Lin, Yuchen Sun, Ding Wang, Cong Guo, Na Chen, Yang Liu, and Jianjun Tang, "Efficient routing algorithm for trusted relay quantum key distribution networks via quantum reinforcement learning", Optics Express 33 22, 46545 (2025).
[69] Shumin Zhou, Hailan Ma, Sen Kuang, and Daoyi Dong, "Auxiliary Task-Based Deep Reinforcement Learning for Quantum Control", IEEE Transactions on Cybernetics 55 2, 712 (2025).
The above citations are from Crossref's cited-by service (last updated successfully 2026-08-08 04:25:36) and SAO/NASA ADS (last updated successfully 2026-08-08 04:25:38). The list may be incomplete as not all publishers provide suitable and complete citation data.
This Paper is published in Quantum under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Copyright remains with the original copyright holders such as the authors or their institutions.