Question 1
A data science team is preparing a large text dataset for fine-tuning a Llama 3 model. The dataset consists of 500GB of raw text files. The team needs to perform tokenization and data cleaning as quickly as possible. Which NVIDIA library is specifically designed for GPU-accelerated data manipulation and would be most suitable for this task?
Answer and explanation
Correct answer: C
NVIDIA RAPIDS cuDF is the correct choice. It provides a pandas-like API for data manipulation that runs on GPUs, making it ideal for accelerating data preprocessing tasks like cleaning and tokenization on large datasets. Triton is for inference serving, TensorRT-LLM is for optimizing inference, and NeMo is a framework for building and training models.