What fundamental limitation of Recurrent Neural Networks (RNNs) did the Transformer architecture primarily resolve by introducing the self-attention mechanism?
Answer and explanation
Correct answer: B
RNNs process sequences step-by-step, making parallel computation impossible and causing the vanishing gradient problem over long distances. The Transformer's self-attention mechanism evaluates all tokens simultaneously, allowing massive parallelization and direct connections between distant words.
Question 2
During the computation of scaled dot-product attention in a Transformer, the dot product of the Query (Q) and Key (K) matrices is divided by the square root of the dimension of the key (sqrt(d_k)). What is the primary mathematical reason for this scaling factor?
flowchart LR
Q[Query] --> Dot[Dot Product Q*K^T]
K[Key] --> Dot
Dot --> Scale[Scale by 1/sqrt d_k]
Scale --> Softmax[Softmax]
Softmax --> Mult[Multiply with Value]
V[Value] --> Mult
Mult --> Out[Attention Output]
Answer and explanation
Correct answer: A
As the dimension of the key vectors (d_k) increases, the variance of the dot product increases, leading to very large values. Feeding these large values into a softmax function pushes the outputs to 1 or 0, resulting in extremely small (vanishing) gradients during backpropagation. Scaling by the square root of d_k normalizes the variance.
Question 3
Multiple answers
Which TWO of the following statements accurately describe the role and implementation of Positional Encoding in the standard Transformer architecture? (Select TWO)
Answer and explanation
Correct answers: B, C
Self-attention operations treat the input sequence as a "bag of words." Without positional encoding, the model would not be able to distinguish the order of tokens, which is critical for understanding language syntax and semantics.
The original "Attention Is All You Need" paper proposed using continuous sinusoidal functions (sine and cosine) of varying frequencies to inject positional information, allowing the model to easily learn to attend by relative positions.
Question 4
When comparing different foundational LLM architectures, which structural approach is primarily utilized by models like BERT to achieve deep bidirectional context understanding?
Answer and explanation
Correct answer: B
BERT (Bidirectional Encoder Representations from Transformers) relies on an encoder-only architecture. It uses masked language modeling (MLM) during pre-training, allowing it to look at context from both the left and right simultaneously, unlike autoregressive decoder models (like GPT) that only look at past tokens.
Question 5
What is the primary function of a Vector Database in a Generative AI application architecture?
Answer and explanation
Correct answer: B
Vector databases are specialized systems designed to store, manage, and index high-dimensional vector embeddings generated by AI models. Their primary function is to perform extremely fast similarity searches (like approximate nearest neighbor search) to find data semantically related to a user's query.
Question 6
An AI engineer is setting up an unstructured data ingestion pipeline for semantic search. Which step must immediately precede the insertion of text chunks into the Vector Database?
Answer and explanation
Correct answer: B
Before text can be stored and searched semantically in a vector database, it must be converted into numerical representations (vectors). This is done by passing the text chunks through an embedding model (like text-embedding-ada-002 or MiniLM).