Ncp - AI Operations Free Sample Questions

Description

Covers deploying Mission Control and Base Command Manager, job scheduling, cluster administration and GPU virtualization, training and inference workload deployment, and container troubleshooting.

19 free sample questions247 in the full practice test

Try simulator

NCP-AIO Sample Questions

  1. Question 1

    Intermediate

    Workload Management · Inference Workload Deployment

    An MLOps team is deploying a new Triton Inference Server on a Kubernetes cluster managed by NVIDIA Base Command Manager (BCM). They observe that during traffic spikes, pod startup latency is high, impacting the auto-scaler's effectiveness. The investigation reveals that the delay is caused by downloading a large, multi-gigabyte model from a remote S3 bucket every time a new pod is created. Which strategy provides the MOST efficient solution to reduce this model-loading latency for new inference pods?

    Answer and explanation

    Correct answer: C

    The most efficient and scalable solution is to implement a caching layer. A DaemonSet can be configured to run on each node, pre-warming a local cache with the required models. New Triton pods can then mount this local cache and load models almost instantaneously, drastically reducing startup latency. Pre-baking models into the container image is inflexible, making model updates cumbersome. Increasing network bandwidth helps but does not eliminate the latency of downloading large files for every new pod.

  2. Question 2

    AdvancedMultiple answers

    Troubleshooting and Optimization · GPU and Fabric Manager

    A research institution is running a multi-node, multi-GPU deep learning training job on a Slurm cluster composed of NVIDIA DGX A100 nodes. A junior administrator reports that the job is running, but dcgmi diag -r 1 shows NVLink bandwidth is significantly lower than the expected 600 GB/s bidirectional bandwidth. All GPUs are healthy. Which TWO of the following are the most likely causes for the degraded NVLink performance? (Select TWO)

    Answer and explanation

    Correct answers: A, B

    Fabric Manager is essential for initializing and maintaining the high-speed NVSwitch fabric in DGX systems. If it's not running, the NVLinks may operate in a degraded, non-optimal state. Additionally, incorrect process-to-GPU affinity, often due to poor Slurm task binding, can force inter-GPU communication to traverse the slower CPU interconnect (UPI) instead of the direct NVLink paths, severely impacting bandwidth.

  3. Question 3

    Intermediate

    Installation and Deployment · Network Configuration

    You are tasked with deploying a new bare-metal Kubernetes cluster on a set of servers equipped with NVIDIA ConnectX-6 Dx SmartNICs. To maximize network performance and offload the host CPU, you plan to use the ASAP² (Accelerated Switching and Packet Processing) feature. Which component is essential for enabling and managing ASAP² in this environment?

    Answer and explanation

    Correct answer: D

    NVIDIA DOCA (Data Center on a Chip Architecture) is the software framework required to unlock and program the capabilities of BlueField DPUs and ConnectX SmartNICs. ASAP² is a feature managed through DOCA, which allows for offloading the virtual switch (OVS) data plane to the SmartNIC's hardware, freeing up CPU cores and accelerating network packet processing.

  4. Question 4

    Beginner

    Administration · BCM Administration

    True or False: When using NVIDIA Base Command Manager (BCM) to provision a cluster, the nvsm-health command is the primary tool used from the head node to verify the health and status of all compute nodes and their GPUs after an OS image has been deployed.

    Answer and explanation

    Correct answer: B

    False. The primary command used within the BCM environment for cluster-wide health checks is pdsh combined with dcgmi. A typical command would be pdsh -g all dcgmi diag -r 1. While nvsm-health is a valid NVIDIA System Management command, it's generally used on a single node. BCM administration relies on parallel shell tools like pdsh to execute commands across the entire cluster efficiently.

  5. Question 5

    Intermediate

    Administration · MIG Configuration

    A hospital's data science team uses a shared NVIDIA DGX H100 server for developing medical imaging AI models. To ensure resource isolation and predictable performance for multiple concurrent users, the lead AI Operations engineer decides to partition the GPUs using Multi-Instance GPU (MIG). The requirement is to create the maximum possible number of isolated, compute-capable instances on a single H100 GPU. Which nvidia-smi command correctly achieves this?

    Answer and explanation

    Correct answer: A

    The NVIDIA H100 GPU supports up to seven MIG instances. The smallest compute-capable profile is 1g.10gb. This command correctly creates seven instances of this profile, maximizing the number of isolated environments on a single GPU. The other options either use invalid profile names for H100, an incorrect number of instances, or a mix of profiles that does not yield the maximum number.

  6. Question 6

    Intermediate

    Administration · Data Center Architecture

    An AI infrastructure team is managing a large cluster with hundreds of nodes. They need a way to perform out-of-band management, including power cycling, BIOS configuration, and monitoring hardware sensors, without relying on the node's primary operating system. The cluster nodes are equipped with NVIDIA BlueField-3 DPUs. How can the team achieve this level of remote management?

    Answer and explanation

    Correct answer: B

    NVIDIA BlueField-3 DPUs include an integrated Baseboard Management Controller (BMC). This allows for true out-of-band management of the host server. Administrators can connect to the DPU's management port to perform tasks like power control, firmware updates, and sensor monitoring, completely independent of the host server's OS state.

  7. Question 7

    Intermediate

    Troubleshooting and Optimization · Network Performance

    A system administrator is troubleshooting a performance issue with a distributed training job running across several nodes connected via an InfiniBand fabric. They suspect a faulty cable or switch port is causing excessive errors on a specific link. Which command-line utility is the MOST appropriate first step to check the error counters and status of all links in the InfiniBand fabric from a single node?

    Answer and explanation

    Correct answer: D

    ibdiagnet is a comprehensive InfiniBand fabric diagnostic tool. It scans the entire fabric topology, checks for connectivity, and reports on the health of links, including detailed error counters like symbol errors and link recovery events. Running ibdiagnet is the standard procedure for identifying physical layer issues such as bad cables or failing ports.

  8. Question 8

    Advanced

    Installation and Deployment · Slurm Installation and Configuration

    An administrator is configuring a new Slurm cluster with NVIDIA H100 GPUs. They have enabled MIG on all GPUs and created various MIG instance profiles. The goal is to allow users to request a specific MIG instance profile in their sbatch scripts. For example, a user should be able to request two instances of the 2g.20gb profile. How must the gres.conf file be configured on the Slurm controller to enable this functionality?

    Answer and explanation

    Correct answer: D

    To make Slurm aware of MIG instances and allow users to request them by profile, you must use the MIG_CONFIG flag for the GPU resources in gres.conf. This tells Slurm to query the MIG configuration for each GPU and make the available profiles (e.g., 2g.20gb, 3g.40gb) available as schedulable resources. Users can then request them using --gpus=mig:2g.20gb:2.

  9. Question 9

    Advanced

    Administration · vGPU Administration

    A financial services company is using NVIDIA AI Enterprise to run critical, low-latency inference workloads on a vSphere cluster. During a routine audit, a security team member raises a concern that a compromised VM could potentially access the memory of other VMs running on the same GPU via a side-channel attack. Which NVIDIA vGPU feature, when enabled on the host, mitigates this specific security risk?

    Answer and explanation

    Correct answer: C

    Frame buffer (VRAM) scrubbing is a security feature of NVIDIA vGPU. When enabled, it ensures that whenever a VM's vGPU instance is destroyed or its memory is deallocated, the corresponding physical VRAM region on the GPU is overwritten with a pattern (typically zeros). This prevents a subsequently scheduled VM from potentially reading residual data left by the previous VM, mitigating the risk of data leakage between tenants.

Register free for 10 more questions

Or unlock all 247 NCP-AIO questions with explanations, timed mode and flashcards.