GPT-1 established the pre-training and fine-tuning paradigm, GPT-2 demonstrated emergent zero-shot learning capabilities at scale, and GPT-3 introduced in-context learning with 175 billion parameters.
Open the full topic