NVIDIA-Certified Associate Generative AI LLMs Free Sample Questions

20 free sample questions203 in the full practice test

Try simulator

NCA-GENL Sample Questions

  1. Question 1

    A data science team is preparing a large text dataset for fine-tuning a Llama 3 model. The dataset consists of 500GB of raw text files. The team needs to perform tokenization and data cleaning as quickly as possible. Which NVIDIA library is specifically designed for GPU-accelerated data manipulation and would be most suitable for this task?

    Answer and explanation

    Correct answer: C

    NVIDIA RAPIDS cuDF is the correct choice. It provides a pandas-like API for data manipulation that runs on GPUs, making it ideal for accelerating data preprocessing tasks like cleaning and tokenization on large datasets. Triton is for inference serving, TensorRT-LLM is for optimizing inference, and NeMo is a framework for building and training models.

  2. Question 2

    A developer is implementing a Retrieval-Augmented Generation (RAG) system to answer questions about internal company documents. They have already generated embeddings and stored them in a vector database. Which step in the RAG pipeline immediately follows the retrieval of relevant document chunks from the vector database?

    Answer and explanation

    Correct answer: B

    In a standard RAG pipeline, after the system retrieves the most relevant document chunks (context) from the vector database based on the user's query, the next step is to combine this retrieved context with the original user prompt. This new, augmented prompt is then sent to the LLM to generate a contextually-aware answer. Generating the final answer happens after the prompt is augmented.

  3. Question 3

    An MLOps engineer is deploying a large language model using NVIDIA Triton Inference Server. They observe that under high load, requests with long sequences are causing head-of-line blocking, increasing latency for all subsequent requests. Which Triton feature is specifically designed to mitigate this issue by processing requests out of order?

    Answer and explanation

    Correct answer: C

    In-flight batching (also known as continuous batching) is the correct feature. Unlike dynamic batching, which waits to form a complete batch before processing, in-flight batching can add new requests to a batch that is already being processed. This allows shorter requests to be processed and return while longer ones are still running, effectively eliminating head-of-line blocking and improving overall throughput and latency.

  4. Question 4

    Multiple answers

    A hospital is developing an internal chatbot to help doctors quickly summarize patient histories. To ensure patient privacy and prevent the model from discussing off-topic subjects like celebrity gossip or financial advice, which TWO NVIDIA technologies or techniques should be implemented? (Select TWO)

    Answer and explanation

    Correct answers: A, B

    NeMo Guardrails is specifically designed to control the topics an LLM can discuss, prevent it from executing unsafe actions, and ensure it stays within predefined conversational boundaries. This directly addresses the need to prevent off-topic conversations.

    Using a RAG architecture that only pulls from the hospital's internal, secure patient database ensures that the LLM's responses are grounded in factual, relevant data and do not hallucinate or access external information. This enhances both accuracy and privacy.

  5. Question 5

    True or False: Using LoRA (Low-Rank Adaptation) for fine-tuning a large language model involves updating all of the original model's weights.

    Answer and explanation

    Correct answer: B

    The statement is false. LoRA is a Parameter-Efficient Fine-Tuning (PEFT) method. It works by freezing the original pre-trained model weights and injecting small, trainable rank-decomposition matrices (adapters) into the layers of the Transformer architecture. Only these new, much smaller matrices are updated during training, which drastically reduces the number of trainable parameters and computational cost.

  6. Question 6

    A research team is fine-tuning a 70-billion parameter model on a single DGX node with 8 GPUs. The full model requires more VRAM than is available on a single GPU. To overcome this, they decide to split the model's layers across the 8 GPUs. What is this distributed training technique called?

    Answer and explanation

    Correct answer: C

    Pipeline Parallelism is the technique of splitting the layers of a neural network model across multiple GPUs. Each GPU processes a subset of the model's layers (a stage) in a pipeline fashion. This is used when a model is too large to fit into a single GPU's memory. Data Parallelism replicates the model on each GPU, and Tensor Parallelism splits individual layers/tensors across GPUs.

  7. Question 7

    When evaluating a text summarization model, a team calculates a score based on the overlap of n-grams between the machine-generated summary and a human-written reference summary. This metric is known as:

    Answer and explanation

    Correct answer: C

    ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is a set of metrics used for evaluating automatic summarization and machine translation. It works by comparing an automatically produced summary against a set of reference summaries (typically human-written) and measures the overlap of n-grams, word sequences, and word pairs.

  8. Question 8

    A developer is using the NVIDIA NeMo Framework to create a custom conversational AI application. They need to define rules for how the AI should respond to inappropriate user queries and ensure the conversation stays on a specific topic. Which NeMo component is specifically designed for this purpose?

    Answer and explanation

    Correct answer: B

    NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable, rule-based controls to LLM-based applications. It allows developers to define specific conversational paths, control topics, prevent the model from accessing certain tools, and filter out undesirable content, making it the correct choice for ensuring safe and on-topic conversations.

  9. Question 9

    What is the primary function of the self-attention mechanism in the Transformer architecture?

    Answer and explanation

    Correct answer: C

    The self-attention mechanism allows the model to associate each word in the input sequence with other words in the same sequence. It calculates attention scores that determine how much focus to place on other words when encoding a specific word. This allows the model to capture long-range dependencies and understand context within the sequence, which is a key advantage over recurrent architectures.

  10. Question 10

    A financial firm is using a generative AI model to create market analysis reports. They are concerned that the model, trained on public data, might inadvertently generate text that is too similar to copyrighted articles, creating a legal risk. Which AI safety problem does this scenario describe?

    Answer and explanation

    Correct answer: C

    This scenario describes output regeneration, also known as plagiarism or regurgitation. It occurs when a generative model reproduces verbatim or near-verbatim excerpts from its training data. If the training data includes copyrighted material, this can lead to copyright infringement. This is a key concern in responsible AI development.

Register free to unlock 10 more sample questions

Lifetime One

Own this practice test forever.

$79.99
$75.99
one-time
  • Full access to 203 questions
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • Brainy AI Assistant
  • Lifetime updates

Two

Any 2 exams per month.

$20.00/exam
$39.99
/month
  • 2 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 1,000 Brainy AI Credits
  • Cancel anytime

Premium Twelve

Any 12 exams over 3 months.

$15.00/exam
$179.99
/3 months
  • 4 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 15,000 Brainy AI Credits
  • Dedicated support
  • Friend seat included — full access

Trusted by professionals at

NvidiaSupabaseGitHubOpenAITursoClerkClaude AIAmazon