Conferences and seminars
Attention and transformers explained for mathematicians
Bilde
Francesco Ballerin, PhD @ Mathematical Department at UiB
Abstract:
Transformers are the neural network architecture behind large language models (LLM) and a growing range of applications in vision, biology, and beyond. At their core lies the attention mechanism, a construction that is usually described in the language of computer science rather than mathematics. This talk presents attention and transformers from a mathematician's point of view. No deep and specific background in machine learning is required, though familiarity with the basics of neural networks, and convolutional neural networks in particular, can provide useful context to the talk.