Hybrid Genetic Programming and Q-Learning for Behavior Tree Evolution inMicroRTS Tactical Scenarios

Abstract

Genetic Programming (GP) can be used to evolve human-interpretable Behavior Tree controllers for realtimestrategy games. Current Behavior Trees-based GP approaches compute fitness only at the end ofeach episode, which does not allow learning from finer-grained tactical decisions during gameplay. In thiswork, we integrate tabular Q-learning within the BT controller to gate actions at terminal nodes based onlearned action-values, and collect additional rewards at each tick of the game that are used to augmentthe final fitness signal. Since the Q-table is reset at the start of each episode, our approach allows thelearned information to affect fitness, while keeping the learned values from one generation separate fromsubsequent generations (a form of Baldwinian learning). Applied to plain-terrain MicroRTS micromanagementchallenge (population of 100 BTs, 2000 generations, against a deterministic rush opponent), ourapproach achieves a maximum fitness of 26.44 compared to GP’s max fitness of 22.0, a relative improvementof 20.2%. Mean best-fitness was also increased by 17.8% compared to the GP-only baseline. Theseresults are reported from a single representative run per configuration; multi-seed replication is identifiedas a priority for future work. The additional per-episode RL reward signal is strongly correlated with elitefitness (r = 0.89), confirming that it provides informative guidance for the evaluation. Unit coordinationexhibited by the hybrid agents is also found to be more structured, with defined front-line and support roles.The agent maintains its interpretable Behavior Tree structure.

References

[1] J. E. Baker. Reducing bias and inefficiency in the selection

algorithm. In Proceedings of the Second International

Conference on Genetic Algorithms (ICGA),

pages 14–21. Lawrence Erlbaum Associates, 1987.

[2] J. M. Baldwin. A new factor in evolution. American

Naturalist, 30:441–451, 536–553, 1896.

[3] C. Berner, G. Brockman, B. Chan, V. Cheung,

P. Dębiak, C. Dennison, D. Farhi, Q. Fischer,

S. Hashme, C. Hesse, et al. Dota 2 with large

scale deep reinforcement learning. arXiv preprint

arXiv:1912.06680, 2019.

[4] Michele Colledanchise, Diogo Almeida, and Petter

Ögren. Towards blended reactive planning and acting using behavior trees. In 2019 IEEE International Conference

on Robotics and Automation (ICRA), pages

8839–8845. IEEE, may 2019.

[5] Michele Colledanchise and Petter Ögren. Behavior

Trees in Robotics and AI: An Introduction. CRC Press,

Boca Raton, FL, 2018.

[6] Sam Michael Devlin and Daniel Kudenko. Dynamic

potential-based reward shaping. In Proceedings of the

11th International Conference on Autonomous Agents

and Multiagent Systems (AAMAS 2012), volume 1,

pages 433–440. IFAAMAS, 2012.

[7] Rahul Dey and Chris Child. Ql-bt: Enhancing

behaviour tree design and implementation with qlearning.

In 2013 IEEE Conference on Computational

Intelligence in Games (CIG), pages 275–282, New

York, NY, USA, sep 2013. IEEE.

[8] Razan Ghzouli, Thorsten Berger, Einar Broch

Johnsen, Swaib Dragule, and Andrzej Wąsowski. Behavior

trees in action: A study of robotics applications.

In Proceedings of the 13th ACM SIGPLAN International

Conference on Software Language Engineering

(SLE 2020), pages 196–209, Virtual Event,

USA, nov 2020. Association for Computing Machinery

(ACM).

[9] G. E. Hinton and S. J. Nowlan. How learning can

guide evolution. Complex Systems, 1(3):495–502,

1987.

[10] S. Huang, S. Ontañon, C. Bamford, and L. Grela.

Gym-μRTS: Toward affordable full game real-time

strategy games research with deep reinforcement

learning. In 2021 IEEE Conference on Games (CoG),

pages 1–8. IEEE, 2021.

[11] Matteo Iovino, Edvards Scukins, Jonathan Styrud,

Petter Ögren, and Christian Smith. A survey of behavior

trees in robotics and ai. Robotics and Autonomous

Systems, 154:104096, aug 2022.

[12] Matteo Iovino, Jonathan Styrud, Pietro Falco, and

Christian Smith. Learning behavior trees with genetic

programming in unpredictable environments. In 2021

IEEE International Conference on Robotics and Automation

(ICRA), pages 4591–4597. IEEE, jul 2021.

[13] Matteo Iovino, Jonathan Styrud, Pietro Falco, and

Christian Smith. A framework for learning behavior

trees in collaborative robotic applications. In

2023 IEEE 19th Conference on Automation Science

and Engineering (CASE), pages 738–745, Kampala,

Uganda, aug 2023. IEEE.

[14] Shauharda Khadka and Kagan Tumer. Evolutionguided

policy gradient in reinforcement learning. Advances

in Neural Information Processing Systems 31

(NeurIPS 2018), pages 1–12, 2018.

[15] J. R. Koza. Genetic Programming: On the Programming

of Computers by Means of Natural Selection.

MIT Press, 1992.

[16] Martin Masek, Chiou Peng Lam, Luke Kelly, and

Martin Wong. Discovering optimal strategy in tactical

combat scenarios through the evolution of behaviour

trees. Annals of Operations Research, 320(2):901–

936, 2023.

[17] V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P.

Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu.

Asynchronous methods for deep reinforcement learning.

In Proceedings of the 33rd International Conference

on Machine Learning (ICML), pages 1928–

1937, 2016.

[18] Volodymyr Mnih, Koray Kavukcuoglu, David Silver,

Andrei A. Rusu, Joel Veness, Marc G. Bellemare,

Alex Graves, Martin Riedmiller, Andreas K. Fidjeland,

Georg Ostrovski, Stig Petersen, Charles Beattie,

Amir Sadik, Ioannis Antonoglou, Helen King,

Dharshan Kumaran, Daan Wierstra, Shane Legg, and

Demis Hassabis. Human-level control through deep

reinforcement learning. Nature, 518(7540):529–533,

feb 2015.

[19] Andrew Y. Ng, Daishi Harada, and Stuart J. Russell.

Policy invariance under reward transformations: Theory

and application to reward shaping. In Proceedings

of the Sixteenth International Conference on Machine

Learning (ICML 1999), pages 278–287, San Francisco,

CA, USA, 1999. Morgan Kaufmann.

[20] S. Ontañon, N. A. Barriga, C. R. Silva, R. O. Moraes,

and L. H. S. Lelis. The first MicroRTS artificial intelligence

competition. AI Magazine, 39(1):75–78,

2018.

[21] S. Ontañon, G. Synnaeve, A. Uriarte, F. Richoux,

D. Churchill, and M. Preuss. A survey of real-time

strategy game AI research and competition in Star-

Craft. IEEE Transactions on Computational Intelligence

and AI in Games, 5(4):293–311, 2013.

[22] R. Poli, W. B. Langdon, and N. F. McPhee. A Field

Guide to Genetic Programming. Lulu.com, 2008.

[23] M. Samvelyan, T. Rashid, C. S. de Witt, G. Farquhar,

N. Nardelli, T. G. J. Rudner, C. Hung, P. H. S. Torr,

J. Foerster, and S. Whiteson. The StarCraft multiagent

challenge. Proceedings of the 18th International

Conference on Autonomous Agents and Multi-Agent

Systems (AAMAS), pages 2186–2188, 2019.

[24] Olivier Sigaud. Combining evolution and deep reinforcement

learning for policy search: a survey. arXiv

preprint, 2022.

[25] Jonathan Styrud, Matteo Iovino, Mikael Norrlöf,

Mårten Björkman, and Christian Smith. Combining planning and learning of behavior trees for robotic

assembly. In 2022 IEEE International Conference

on Robotics and Automation (ICRA), pages 11511–

11517. IEEE, may 2022.

[26] Richard S. Sutton and Andrew G. Barto. Reinforcement

Learning: An Introduction. MIT Press, Cambridge,

MA, 2nd edition, 2018.

[27] O. Vinyals, I. Babuschkin, W. M. Czarnecki,

M. Mathieu, A. Dudzik, J. Chung, D. H. Choi,

R. Powell, T. Ewalds, P. Georgiev, et al. Grandmaster

level in StarCraft II using multi-agent reinforcement

learning. Nature, 575:350–354, 2019.

[28] Christopher J. C. H. Watkins and Peter Dayan. Qlearning.

Machine Learning, 8(3–4):279–292, may

1992.

[29] G. N. Yannakakis and J. Togelius. Artificial Intelligence

and Games. Springer, 2018.

Authors

  • Mohamed Amine Chikh Touami LISYS Lab, University of Mascara, Algeria
  • Mohammed Salem LISYS Lab, University of Mascara, Algeria
  • Mohamed Fayçal Khelfi ESGEE Oran, Algeria

DOI:

https://doi.org/10.31449/inf.v50i14.14678

Keywords:

Reinforcement Learning (RL), Behavior Trees, Genetic Programming, Q-learning, RTS Games, MicroRTS, Hybrid Soft Computing

Downloads

Published

08/06/2026

How to Cite

Chikh Touami, M. A., Salem, M., & Khelfi, M. F. (2026). Hybrid Genetic Programming and Q-Learning for Behavior Tree Evolution inMicroRTS Tactical Scenarios. Informatica, 50(14). https://doi.org/10.31449/inf.v50i14.14678