Publications
Scientific publications
Е.Е. Васильева, А.В. Леонидов, А.С. Титов.
Влияние матрицы выплат на больцмановское Q-обучение в повторяющихся матричных играх
// Математическая Теория Игр и ее Приложения, т. 18, в. 2. 2026. C. 23-49
Ekaterina E. Vasilyeva, Andrey V. Leonidov, Alexey S. Titov. The impact of payoff matrix on Boltzmann Q-learning // Mathematical game theory and applications. Vol 18. No 2. 2026. Pp. 23-49
Keywords: reinforcement learning, Q-learning, prisoners’ dilemma, payoff matrix
This study investigates the influence of the payoff matrix structure on the dynamics of Boltzmann Q-learning in repeated 2 × 2 matrix games, using the prisoner’s dilemma, the stag hunt, and an interpolation between them as examples. The payoff matrices are parameterized by characteristic gaps between elements. It is shown that a formal definition of the game class is insufficient to describe the structure of stationary states: the phase diagram of the number of asymptotic solutions (which are quantal response equilibria) depends on the specific values of the gaps. The paper also demonstrates the effect of “strategic cooling”, manifested as a shift of the phase boundaries of the solution-count diagram into the region of higher temperatures as the discounting factor increases, due to the emergence of additional metastable states. A direct dependence of the agents’ residence time in these metastable states on the sizes of the characteristic gaps is demonstrated. The obtained results show that the specific values of parameters of the payoff matrix significantly affect the dynamics and the resulting states of Boltzmann Q-learning.
Indexed at RSCI, RSCI (WS)
Last modified: July 15, 2026


