Conferences and seminars

Attention and transformers explained for mathematicians


Bilde
Attention
Attention Photo: Created by Francesco Ballerin

Francesco Ballerin, PhD @ Mathematical Department at UiB

Abstract: 

Transformers are the neural network architecture behind large language models (LLM) and a growing range of applications in vision, biology, and beyond. At their core lies the attention mechanism, a construction that is usually described in the language of computer science rather than mathematics. This talk presents attention and transformers from a mathematician's point of view. No deep and specific background in machine learning is required, though familiarity with the basics of neural networks, and convolutional neural networks in particular, can provide useful context to the talk.