Foundations of Transformer Models - Absolute Positional Embeddings
Have you ever felt confident about the big picture of transformer architectures but found yourself scratching your head when it came to the nitty-gritty details? If so, you’re not alone. Whether you’re a student diving into NLP for the first time or a seasoned professional brushing up for your next interview, understanding positional embeddings is essential. This article aims to break down the intricate mechanics of transformers, starting with absolute positional embeddings—the foundation of how models like BERT understand token order. Note that I assume you are aware of what an Encoder/Decoder is and what are embeddings. If not, there is plenty of resources that will do more justice to these topics than I can. So feel free to come back to this article once you get through with that.
Read more