Hybrid Genetic Programming and Q-Learning for Behavior Tree Evolution inMicroRTS Tactical Scenarios
Abstract
Genetic Programming (GP) can be used to evolve human-interpretable Behavior Tree controllers for realtimestrategy games. Current Behavior Trees-based GP approaches compute fitness only at the end ofeach episode, which does not allow learning from finer-grained tactical decisions during gameplay. In thiswork, we integrate tabular Q-learning within the BT controller to gate actions at terminal nodes based onlearned action-values, and collect additional rewards at each tick of the game that are used to augmentthe final fitness signal. Since the Q-table is reset at the start of each episode, our approach allows thelearned information to affect fitness, while keeping the learned values from one generation separate fromsubsequent generations (a form of Baldwinian learning). Applied to plain-terrain MicroRTS micromanagementchallenge (population of 100 BTs, 2000 generations, against a deterministic rush opponent), ourapproach achieves a maximum fitness of 26.44 compared to GP’s max fitness of 22.0, a relative improvementof 20.2%. Mean best-fitness was also increased by 17.8% compared to the GP-only baseline. Theseresults are reported from a single representative run per configuration; multi-seed replication is identifiedas a priority for future work. The additional per-episode RL reward signal is strongly correlated with elitefitness (r = 0.89), confirming that it provides informative guidance for the evaluation. Unit coordinationexhibited by the hybrid agents is also found to be more structured, with defined front-line and support roles.The agent maintains its interpretable Behavior Tree structure.References
[1] J. E. Baker. Reducing bias and inefficiency in the selection
algorithm. In Proceedings of the Second International
Conference on Genetic Algorithms (ICGA),
pages 14–21. Lawrence Erlbaum Associates, 1987.
[2] J. M. Baldwin. A new factor in evolution. American
Naturalist, 30:441–451, 536–553, 1896.
[3] C. Berner, G. Brockman, B. Chan, V. Cheung,
P. Dębiak, C. Dennison, D. Farhi, Q. Fischer,
S. Hashme, C. Hesse, et al. Dota 2 with large
scale deep reinforcement learning. arXiv preprint
arXiv:1912.06680, 2019.
[4] Michele Colledanchise, Diogo Almeida, and Petter
Ögren. Towards blended reactive planning and acting using behavior trees. In 2019 IEEE International Conference
on Robotics and Automation (ICRA), pages
8839–8845. IEEE, may 2019.
[5] Michele Colledanchise and Petter Ögren. Behavior
Trees in Robotics and AI: An Introduction. CRC Press,
Boca Raton, FL, 2018.
[6] Sam Michael Devlin and Daniel Kudenko. Dynamic
potential-based reward shaping. In Proceedings of the
11th International Conference on Autonomous Agents
and Multiagent Systems (AAMAS 2012), volume 1,
pages 433–440. IFAAMAS, 2012.
[7] Rahul Dey and Chris Child. Ql-bt: Enhancing
behaviour tree design and implementation with qlearning.
In 2013 IEEE Conference on Computational
Intelligence in Games (CIG), pages 275–282, New
York, NY, USA, sep 2013. IEEE.
[8] Razan Ghzouli, Thorsten Berger, Einar Broch
Johnsen, Swaib Dragule, and Andrzej Wąsowski. Behavior
trees in action: A study of robotics applications.
In Proceedings of the 13th ACM SIGPLAN International
Conference on Software Language Engineering
(SLE 2020), pages 196–209, Virtual Event,
USA, nov 2020. Association for Computing Machinery
(ACM).
[9] G. E. Hinton and S. J. Nowlan. How learning can
guide evolution. Complex Systems, 1(3):495–502,
1987.
[10] S. Huang, S. Ontañon, C. Bamford, and L. Grela.
Gym-μRTS: Toward affordable full game real-time
strategy games research with deep reinforcement
learning. In 2021 IEEE Conference on Games (CoG),
pages 1–8. IEEE, 2021.
[11] Matteo Iovino, Edvards Scukins, Jonathan Styrud,
Petter Ögren, and Christian Smith. A survey of behavior
trees in robotics and ai. Robotics and Autonomous
Systems, 154:104096, aug 2022.
[12] Matteo Iovino, Jonathan Styrud, Pietro Falco, and
Christian Smith. Learning behavior trees with genetic
programming in unpredictable environments. In 2021
IEEE International Conference on Robotics and Automation
(ICRA), pages 4591–4597. IEEE, jul 2021.
[13] Matteo Iovino, Jonathan Styrud, Pietro Falco, and
Christian Smith. A framework for learning behavior
trees in collaborative robotic applications. In
2023 IEEE 19th Conference on Automation Science
and Engineering (CASE), pages 738–745, Kampala,
Uganda, aug 2023. IEEE.
[14] Shauharda Khadka and Kagan Tumer. Evolutionguided
policy gradient in reinforcement learning. Advances
in Neural Information Processing Systems 31
(NeurIPS 2018), pages 1–12, 2018.
[15] J. R. Koza. Genetic Programming: On the Programming
of Computers by Means of Natural Selection.
MIT Press, 1992.
[16] Martin Masek, Chiou Peng Lam, Luke Kelly, and
Martin Wong. Discovering optimal strategy in tactical
combat scenarios through the evolution of behaviour
trees. Annals of Operations Research, 320(2):901–
936, 2023.
[17] V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P.
Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu.
Asynchronous methods for deep reinforcement learning.
In Proceedings of the 33rd International Conference
on Machine Learning (ICML), pages 1928–
1937, 2016.
[18] Volodymyr Mnih, Koray Kavukcuoglu, David Silver,
Andrei A. Rusu, Joel Veness, Marc G. Bellemare,
Alex Graves, Martin Riedmiller, Andreas K. Fidjeland,
Georg Ostrovski, Stig Petersen, Charles Beattie,
Amir Sadik, Ioannis Antonoglou, Helen King,
Dharshan Kumaran, Daan Wierstra, Shane Legg, and
Demis Hassabis. Human-level control through deep
reinforcement learning. Nature, 518(7540):529–533,
feb 2015.
[19] Andrew Y. Ng, Daishi Harada, and Stuart J. Russell.
Policy invariance under reward transformations: Theory
and application to reward shaping. In Proceedings
of the Sixteenth International Conference on Machine
Learning (ICML 1999), pages 278–287, San Francisco,
CA, USA, 1999. Morgan Kaufmann.
[20] S. Ontañon, N. A. Barriga, C. R. Silva, R. O. Moraes,
and L. H. S. Lelis. The first MicroRTS artificial intelligence
competition. AI Magazine, 39(1):75–78,
2018.
[21] S. Ontañon, G. Synnaeve, A. Uriarte, F. Richoux,
D. Churchill, and M. Preuss. A survey of real-time
strategy game AI research and competition in Star-
Craft. IEEE Transactions on Computational Intelligence
and AI in Games, 5(4):293–311, 2013.
[22] R. Poli, W. B. Langdon, and N. F. McPhee. A Field
Guide to Genetic Programming. Lulu.com, 2008.
[23] M. Samvelyan, T. Rashid, C. S. de Witt, G. Farquhar,
N. Nardelli, T. G. J. Rudner, C. Hung, P. H. S. Torr,
J. Foerster, and S. Whiteson. The StarCraft multiagent
challenge. Proceedings of the 18th International
Conference on Autonomous Agents and Multi-Agent
Systems (AAMAS), pages 2186–2188, 2019.
[24] Olivier Sigaud. Combining evolution and deep reinforcement
learning for policy search: a survey. arXiv
preprint, 2022.
[25] Jonathan Styrud, Matteo Iovino, Mikael Norrlöf,
Mårten Björkman, and Christian Smith. Combining planning and learning of behavior trees for robotic
assembly. In 2022 IEEE International Conference
on Robotics and Automation (ICRA), pages 11511–
11517. IEEE, may 2022.
[26] Richard S. Sutton and Andrew G. Barto. Reinforcement
Learning: An Introduction. MIT Press, Cambridge,
MA, 2nd edition, 2018.
[27] O. Vinyals, I. Babuschkin, W. M. Czarnecki,
M. Mathieu, A. Dudzik, J. Chung, D. H. Choi,
R. Powell, T. Ewalds, P. Georgiev, et al. Grandmaster
level in StarCraft II using multi-agent reinforcement
learning. Nature, 575:350–354, 2019.
[28] Christopher J. C. H. Watkins and Peter Dayan. Qlearning.
Machine Learning, 8(3–4):279–292, may
1992.
[29] G. N. Yannakakis and J. Togelius. Artificial Intelligence
and Games. Springer, 2018.
DOI:
https://doi.org/10.31449/inf.v50i14.14678Keywords:
Reinforcement Learning (RL), Behavior Trees, Genetic Programming, Q-learning, RTS Games, MicroRTS, Hybrid Soft ComputingDownloads
Published
Issue
Section
License
Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.
All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.
Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.







