Azure Data Scientist Free Sample Questions

20 free sample questions303 in the full practice test

Try simulator

DP-100 Sample Questions

  1. Question 1

    A financial services company is developing a Retrieval-Augmented Generation (RAG) solution to answer questions about internal compliance documents. The documents are a mix of short policy statements (1-2 paragraphs) and long procedural guides (10-20 pages). The goal is to ensure that answers are precise and source attribution is accurate. Which data preparation strategy is most suitable for this scenario?

    Answer and explanation

    Correct answer: B

    A recursive character text splitter is ideal for documents with varied structure. It attempts to split along semantic boundaries (paragraphs, sections) first before falling back to smaller units. This preserves the context within chunks, which is crucial for accurate retrieval from both short policies and long guides. A moderate chunk size with overlap ensures that sentences or ideas are not awkwardly split between chunks.

  2. Question 2

    A data science team is using Azure Machine Learning pipelines to orchestrate a complex training workflow. A custom component in the middle of the pipeline frequently fails due to transient network issues when accessing an external data source. The team wants to make the pipeline more resilient without modifying the component's internal code. How should they configure the pipeline job to handle these intermittent failures?

    Answer and explanation

    Correct answer: C

    Azure Machine Learning components can be configured with retry settings directly in their YAML definition. By adding a retry_settings block with properties like count, delay, and backoff, you instruct the pipeline orchestrator to automatically re-run the component if it fails. This is the correct approach for handling transient errors without altering the component's source code.

  3. Question 3

    Multiple answers

    You are designing a secure environment for a multi-team data science project. You need to ensure that each team can manage its own compute resources and data assets but cannot access the resources of other teams. All teams must use a centrally-managed set of curated Docker environments and foundation models. Which combination of Azure Machine Learning features should you use? (Select TWO)

    Answer and explanation

    Correct answers: B, C

    Creating separate workspaces provides the strongest isolation boundary. Each team can manage its own compute, data, and experiments independently, satisfying the requirement that they cannot access each other's resources.

    An Azure Machine Learning registry is designed for sharing assets like models, environments, and components across multiple workspaces. This allows a central MLOps team to manage and distribute curated assets to all the individual team workspaces.

  4. Question 4

    A data scientist is using the Azure Machine Learning SDK v2 to submit a hyperparameter tuning job for a classification model. The goal is to maximize the 'AUC_weighted' metric. The search space is large, and the compute budget is limited. They need to configure the sweep job to efficiently find good parameters by terminating underperforming runs early. Which early termination policy is most appropriate for this goal?

    Answer and explanation

    Correct answer: A

    The Bandit policy is an aggressive early termination policy that terminates runs whose primary metric falls outside a specified slack factor/amount compared to the best-performing run. This is highly effective for efficiently exploring a large search space with a limited budget by quickly discarding unpromising trials.

  5. Question 5

    You are developing a prompt flow that orchestrates multiple calls to a language model to generate a marketing campaign proposal. You need to ensure that the output of an early step, which generates a target audience description, is correctly passed as input to a later step that writes ad copy. Which Prompt flow feature allows you to define this data dependency?

    Answer and explanation

    Correct answer: C

    In Prompt flow, you use Jinja2 templating syntax to reference the outputs of previous nodes. For example, in the ad copy node's prompt, you would write {{generate_audience.output}} to insert the output from the generate_audience node. This creates the explicit data dependency and chains the steps together.

  6. Question 6

    A hospital is using an AutoML for tabular data job to predict patient readmission risk. The Responsible AI dashboard for the best model reveals that the model has a significantly lower prediction accuracy for a minority demographic group compared to the majority group. This indicates a potential fairness issue. What is the most appropriate first step to mitigate this bias?

    Answer and explanation

    Correct answer: C

    Performance disparities often arise from imbalanced data where the model has insufficient examples for minority groups. The most effective initial step is to address this root cause by collecting more representative data or using sampling techniques (like SMOTE or random oversampling) to create a more balanced training dataset. This allows the model to learn the patterns for the minority group more effectively.

  7. Question 7

    An MLOps engineer is packaging a trained MLflow model for deployment. To ensure seamless inference, they must include information about the required input data schema directly within the model artifacts. Which file within the MLflow model directory should be modified to include this schema definition?

    Answer and explanation

    Correct answer: D

    The MLmodel file is the primary metadata file for an MLflow model. It contains essential information, including the model's flavor, creation time, and most importantly, the signature. The signature explicitly defines the input and output schemas (data types, names, and tensor shapes), which is used by deployment tools to validate requests and format data correctly for inference.

  8. Question 8

    True or False: When using an Azure Machine Learning compute cluster for training, you are only billed for the compute time when a job is actively running on the nodes.

    Answer and explanation

    Correct answer: A

    This is true. Azure ML compute clusters can be configured with a minimum number of nodes (typically zero). The cluster automatically scales up when a job is submitted and scales down to the minimum count when idle. You are only charged for the nodes when they are allocated and running, not for the cluster resource itself when it's scaled to zero nodes.

  9. Question 9

    A team has deployed a machine learning model to a managed online endpoint with a blue-green deployment strategy. 80% of the traffic is currently directed to the stable 'blue' deployment, and 20% is directed to the new 'green' deployment for testing. After monitoring, the team confirms the green deployment is performing well and decides to route all traffic to it. Which command should be used to achieve this without causing downtime?

    Answer and explanation

    Correct answer: B

    The az ml online-endpoint update command is used to modify the properties of an existing endpoint, including the traffic allocation. To route all traffic to the green deployment, you would use this command with the --traffic "green=100" parameter. This updates the endpoint's routing rules in place, ensuring a smooth transition with no downtime.

  10. Question 10

    A data scientist is working on a local machine with the Azure Machine Learning SDK and needs to access data stored in an Azure Blob Storage container for interactive analysis in a Jupyter notebook. The workspace is configured with a datastore named blob_datastore. Which code snippet correctly accesses a file named data/customers.csv from this datastore as a Pandas DataFrame?

    Answer and explanation

    Correct answer: A

    The azureml:// URI scheme is the standard way to reference data assets and paths within datastores in Azure ML SDK v2. By creating a data asset with a path pointing to this URI and then using pd.read_csv(my_path), the SDK handles the authentication and data access seamlessly, loading the CSV into a Pandas DataFrame.

Register free to unlock 10 more sample questions

Lifetime One

Own this practice test forever.

$79.99
$75.99
one-time
  • Full access to 303 questions
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • Brainy AI Assistant
  • Lifetime updates

Two

Any 2 exams per month.

$20.00/exam
$39.99
/month
  • 2 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 1,000 Brainy AI Credits
  • Cancel anytime

Premium Twelve

Any 12 exams over 3 months.

$15.00/exam
$179.99
/3 months
  • 4 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 15,000 Brainy AI Credits
  • Dedicated support
  • Friend seat included — full access

Trusted by professionals at

NvidiaSupabaseGitHubOpenAITursoClerkClaude AIAmazon