Question 1
IntermediateWorkload Management · Inference Workload Deployment
An MLOps team is deploying a new Triton Inference Server on a Kubernetes cluster managed by NVIDIA Base Command Manager (BCM). They observe that during traffic spikes, pod startup latency is high, impacting the auto-scaler's effectiveness. The investigation reveals that the delay is caused by downloading a large, multi-gigabyte model from a remote S3 bucket every time a new pod is created. Which strategy provides the MOST efficient solution to reduce this model-loading latency for new inference pods?
Answer and explanation
Correct answer: C
The most efficient and scalable solution is to implement a caching layer. A DaemonSet can be configured to run on each node, pre-warming a local cache with the required models. New Triton pods can then mount this local cache and load models almost instantaneously, drastically reducing startup latency. Pre-baking models into the container image is inflexible, making model updates cumbersome. Increasing network bandwidth helps but does not eliminate the latency of downloading large files for every new pod.
