Skip to main content

Foundation models, self-supervised learning and transfer learning

Foundation models are large-scale neural networks trained on broad and diverse datasets, designed to serve as reusable backbones for a wide range of downstream tasks. The central principle is to learn general-purpose representations that can subsequently be adapted efficiently to new problems. While related ideas such as transfer learning and representation reuse have long been established in machine learning, recent advances in data availability, computational infrastructure, and scalable architectures have given rise to the modern concept of foundation models. In particular, transformer-based architectures enable models with minimal task-specific inductive bias to learn complex relationships directly from data. When trained at sufficient scale, such models can support many applications through lightweight adaptation rather than task-specific training from first principles.

A key enabler of this paradigm shift is large-scale self-supervised or weakly supervised learning. In self-supervised settings, training objectives are derived directly from the data, allowing representation learning without manual annotation. Weak supervision instead exploits indirect or noisy signals, such as image-text pairs or metadata, which provide partial labelling at scale. These approaches enable models to learn from vast quantities of raw data using surrogate objectives such as masked prediction, contrastive learning, or autoregressive modelling. Such strategies allow the extraction of statistical regularities and the construction of representations that capture general semantic and structural information. For example, masked language model pretraining in BERT illustrates how self-supervised learning on unlabelled texts can produce representations that generalise effectively across diverse tasks with minimal modification. Similar principles extend across modalities, including vision-language learning through contrastive alignment.

Once pretrained, these representations can be adapted through various transfer mechanisms. Classical approaches involve full fine-tuning, while more recent methods reduce computational cost by updating only small subsets of parameters via lightweight adapters. Retrieval-augmented approaches incorporate external knowledge at inference time, and prompt-based conditioning enables task adaptation without modifying model weights. At sufficient scale, such systems may also exhibit emergent behaviours, including zero- and few-shot learning. These properties make foundation models particularly attractive in settings characterised by limited labelled data or the need for robust cross-domain generalisation.