Allowing tokens to attend to both past and future context, used in encoder models like BERT.
This concept is essential for understanding large language models and forms a key part of modern AI systems.
Related Concepts
- Attention
- BERT
- Encoder
Allowing tokens to attend to both past and future context, used in encoder models like BERT.
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
A collection of data examples used for training, validating, or testing machine learning models.
A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.
A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.
Learning useful features or representations of data automatically, rather than hand-crafting them.
A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.
A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.
A mechanism where each token attends to all other tokens in the sequence to understand contextual relationships.
Allowing tokens to attend to both past and future context, used in encoder models like BERT.
This concept is essential for understanding large language models and forms a key part of modern AI systems.