- AutorIn
- Sabine Haberland
- Titel
- Bridging Machine Learning and Cognitive Neuroscience: Modeling Stimulus-Response Transformations with Deep Neural Networks
- Zitierfähige Url:
- https://nbn-resolving.org/urn:nbn:de:bsz:14-qucosa2-1055656
- Erstveröffentlichung
- 2026
- Datum der Einreichung
- 21.01.2026
- Datum der Verteidigung
- 10.06.2026
- Abstract (EN)
- Every moment, the human brain makes countless decisions. Each action is the result of the continuous processing of sensory inputs. Through these stimulus-response (S-R) transformations, the brain interacts quickly and effortlessly with dynamic environments. However, how the neural representations underlying the steps of S-R processing are represented in the human brain remains only partly understood. While the standard paradigm in cognitive studies relies on factorial designs with temporally discretized events and controlled stimuli, it has advanced our understanding of brain function by isolating neural correlates of cognitive mechanisms. However, its ecological validity remains limited. To investigate brain function in more real-world scenarios, this dissertation introduces a deep neural network (DNN)-based encoding model to characterize S-R transformations while capturing the complexity and temporal dynamics of sensory processing. Studies have shown that stimulus features derived from DNNs can be mapped onto functional magnetic resonance imaging (fMRI) activity patterns recorded during naturalistic tasks using encoding models. This DNN-based modeling approach has proven to be a powerful tool for gaining deeper insights into how natural visual information is represented along the ventral visual stream. While participants in these studies passively viewed stimuli without performing motor responses, the development of deep Q-networks (DQNs), which are DNNs combined with reinforcement learning, has enabled the characterization of S-R processing and their neural representations along the dorsal stream in time-continuous, complex, and interactive environments, increasing the ecological validity of such experiments. In recent years, the research fields of artificial intelligence, cognitive science, and computational neuroscience have increasingly engaged in the exchange of concepts and empirical findings. In particular, machine learning has benefited from advances in the other two fields. However, it remains unclear whether progress in machine learning can be transferred back. DNNs have become increasingly powerful, outperforming humans across a wide range of tasks. DNNs, originally inspired by the human brain, have become increasingly powerful, but their optimization techniques are no longer biologically plausible and deviate substantially from brain function. Can advanced DNNs provide better cognitive models for investigating S-R transformations? This thesis lies at the intersection of these research domains. Studies 1 and 2 examined human behavior and the fine-grained neural representations underlying S-R transformations in high-dimensional, time-continuous environments such as video games. In these studies, we used a DQN-based encoding model approach to predict motor responses (Study 1) and voxel-wise brain activity measured with fMRI (Study 2) during arcade gameplay. The DQN was used to generate feature representations of the visual stimuli that the participants observed during video gameplay. The activations generated by the DQN were then used as predictors in an encoding model, implemented here as a general linear model. To examine whether recent advances in machine learning can be transferred back to neuroscience, we analyzed the DQN models Ape-X and SEED and compared their encoding performance with that of a baseline model, the original DQN algorithm that first achieved human-level performance on Atari games. These models differ not only in their training algorithms but also in their network architectures, which have significantly improved gaming performance. While both Ape-X and SEED incorporate a dueling network architecture, SEED additionally includes a long short-term memory (LSTM) module as a key innovation. We hypothesized that these enhancements would lead to improved encoding performance. By comparing the models, we aimed to gain deeper insights into the functional organization of the human brain and to identify which of the distinct inductive biases best align with the functional mechanisms of different brain regions. In Study 1, we first analyzed whether human motor responses during arcade gameplay can be predicted from the activations of the output layer of a DQN, representing the Q-values, in order to obtain an initial validation of this approach and to assess whether advances in DQNs improve prediction accuracy. Our results demonstrate that the DQNs can predict human behavior significantly above chance. SEED, the most complex DQN model, significantly outperformed the others, providing initial evidence for the hypothesis and highlighting an important role of the LSTM in improving prediction accuracy. Since we investigated a time-continuous task, we also aimed to determine the temporal resolution at which accurate predictions remain possible, as a resolution matching the original task would not be meaningful. We examined how smoothing the time series affects prediction accuracy and found that a Gaussian kernel of 0.79 seconds optimally balances precise prediction of human motor responses while also maximizing the information content of the human time series. These findings highlight the potential of DQNs as cognitive models of human behavior, demonstrating that recent advances in machine learning can be leveraged to enhance our ability to model human motor responses in visuomotor tasks at a fine-grained temporal scale. These results motivated us to extend this approach to the neural level by testing whether the networks generate features capable of encoding task-relevant neural representations underlying the S-R processing steps. Therefore, in Study 2 we used the activations of the hidden layers of the DQNs within regularized encoding models to predict fMRI data. We showed that the features generated by the DQNs can be used to capture variance in human brain activity along the dorsal stream toward the posterior parietal cortex, further into motor-related areas, including the primary motor cortex and the premotor cortex, up to the lateral prefrontal cortex. Building on the validation of this approach at the neural level and the precise localization of the brain regions involved in processing these stimuli, we further investigated whether the more advanced DQNs can provide more accurate models of brain activity. We demonstrated that this hypothesis also holds at the neural level, with SEED outperforming the other models in explaining voxel activations. Furthermore, we aimed to identify which brain regions are associated with distinct S-R transformation steps and how they are represented within these regions. Comparing the predictive capacity across individual DQN layers showed that the early hidden layers contribute to the prediction of brain activity to a comparable extent. Differences in prediction accuracy only emerged in the deeper layers, where SEED showed particularly strong predictive performance in regions associated with higher-level cognitive functions, including the posterior parietal cortex, premotor cortex, and lateral prefrontal cortex. The LSTM component appeared to play a key role in this, due to its ability to capture temporal dependencies. In the layers of SEED, we additionally observed brain-like hierarchical processing stages of S-R transformations. In combination with insights into their underlying mechanisms, this underscores the potential of DQNs as cognitive models of human brain activity in visuomotor tasks. In summary, these studies validated the DQN-based encoding model approach at the behavioral and neural levels, showing that the representations learned by DQNs are informative for encoding task representations in real-world scenarios. From a neuroscientific perspective, this thesis makes an important contribution to the investigation of human S-R transformations by extending the analytical toolbox. It complements traditional trial-based study designs with a computational modeling framework that enhances ecological validity and may reveal new patterns in brain activity. At the same time, this approach contributes to an ongoing debate on whether DNNs constitute suitable scientific models for cognitive processes, providing supporting evidence for their validity in this context. More broadly, this work aligns with a growing interdisciplinary trend toward transferring methods and concepts between machine learning and computational neuroscience. The proposed framework highlights the mutual benefits between these fields and suggests avenues for methodological refinement on both sides. The DNN-based encoding model approach may pave the way for broader applications of state-of-the-art machine learning algorithms in the analysis of neuroimaging data.
- Verweis
- Advances in deep reinforcement learning enable better predictions of human behavior in time-continuous tasks
DOI: 10.1371/journal.pone.0338034 - Encoding neural representations of time-continuous stimulus-response transformations in the human brain with advanced deep neural networks
DOI: 10.1162/IMAG.a.1142 - Freie Schlagwörter (DE)
- Tiefe neuronale Netze, Enkodierungsmodelle, Arcade Spiele, Neurobildgebung, fMRT
- Freie Schlagwörter (EN)
- Deep neural networks, Encoding models, Arcade games, Neuroimaging, fMRI, LSTM
- Klassifikation (DDC)
- 150
- Klassifikation (RVK)
- CP 5000
- WC 7782
- GutachterIn
- Prof. Dr. Hannes Ruge
- Prof. Dr. Lune Bellec
- Den akademischen Grad verleihende / prüfende Institution
- Technische Universität Dresden, Dresden
- Förder- / Projektangaben
- Deutsche Forschungsgemeinschaft ID: 445383113
- Version / Begutachtungsstatus
- publizierte Version / Verlagsversion
- URN Qucosa
- urn:nbn:de:bsz:14-qucosa2-1055656
- Veröffentlichungsdatum Qucosa
- 01.07.2026
- Dokumenttyp
- Dissertation
- Sprache des Dokumentes
- Englisch
- Lizenz / Rechtehinweis
CC BY 4.0