Databricks Machine Learning Professional Free Sample Questions

20 free sample questions205 in the full practice test

Try simulator

ML-PRO Sample Questions

  1. Question 1

    A machine learning team is developing a model to predict customer churn. They are using Databricks Asset Bundles (DABs) to manage their project environments. They need to define separate configurations for development, staging, and production, including different cluster policies and secrets scopes. Which section of the databricks.yml file is specifically designed to manage these environment-specific overrides?

    Answer and explanation

    Correct answer: B

    The targets block in a databricks.yml file is used to define different deployment targets, which typically correspond to environments like development, staging, and production. This section allows for overriding default configurations specified in the main bundle or resources sections, enabling environment-specific settings for compute, secrets, and other parameters.

  2. Question 2

    An MLOps engineer is implementing a canary deployment for a new version of a demand forecasting model using Databricks Model Serving. The goal is to route 10% of the inference traffic to the new model version (version 2) while the remaining 90% goes to the stable version (version 1). Which configuration snippet correctly implements this traffic split within a model serving endpoint definition?

    Answer and explanation

    Correct answer: A

    Databricks Model Serving endpoints support traffic splitting for canary and blue-green deployments. The served_models array in the endpoint configuration is where you define which model versions are active. The traffic_percentage key for each model version entry specifies the percentage of requests that should be routed to it. This option correctly assigns 90% to version 1 and 10% to version 2.

  3. Question 3

    A data scientist is building a SparkML pipeline to process text data for sentiment analysis. The pipeline needs to tokenize text, remove stop words, and then convert the tokens into numerical feature vectors using TF-IDF. Which sequence of SparkML transformers is correct for this task?

    Answer and explanation

    Correct answer: C

    The correct logical sequence for this NLP preprocessing task is to first break the text into tokens (Tokenizer), then remove common stop words from the token list (StopWordsRemover), then convert the cleaned tokens into term frequencies (HashingTF), and finally re-weight the term frequencies based on their importance across the corpus (IDF).

  4. Question 4

    Multiple answers

    A team is building an automated retraining pipeline for a credit risk model. The pipeline should trigger a new training job whenever significant drift is detected in the model's key input features. They are using Lakehouse Monitoring to track drift. Which of the following components are essential for implementing this automated retraining workflow? (Select THREE)

    Answer and explanation

    Correct answers: A, B, C

  5. Question 5

    An ML engineer is tasked with creating a custom PyFunc model in MLflow. This model needs to load a pre-trained tokenizer from Hugging Face and a custom-trained scikit-learn classifier. The entire model, including the tokenizer, must be packaged together for deployment to a sandboxed environment without internet access. Which MLflow feature should be used to package the tokenizer along with the model?

    Answer and explanation

    Correct answer: C

    The artifacts parameter in mlflow.pyfunc.log_model() is designed for this exact purpose. It allows you to specify a dictionary of local file paths that will be packaged with the model. Inside the custom model's load_context method, you can then access these artifacts using the provided context object, ensuring the model is self-contained and portable.

  6. Question 6

    True or False: When using Databricks Feature Store, point-in-time correctness is automatically guaranteed for batch inference jobs without any specific configuration required in the create_training_set method.

    Answer and explanation

    Correct answer: B

    False. To ensure point-in-time correctness and prevent data leakage, you must provide a timestamp lookup key in your primary key DataFrame when calling fs.create_training_set(). The Feature Store uses this timestamp to join features that were valid at that specific point in time, preventing the model from being trained on data that would not have been available at the time of prediction.

  7. Question 7

    Multiple answers

    A large-scale image classification model is being trained on Databricks. The team observes that the training process is bottlenecked by the single-node driver's ability to coordinate the workers. They decide to explore distributed hyperparameter tuning to find optimal learning rates. Which two of the following technologies are natively integrated with Databricks for distributed hyperparameter tuning and can effectively manage this workload? (Select TWO)

    Answer and explanation

    Correct answers: A, B

  8. Question 8

    An ML team has configured Lakehouse Monitoring on an inference table. They receive an alert that the Population Stability Index (PSI) for a critical categorical feature, 'customer_segment', has exceeded the defined threshold. However, the drift analysis for the model's prediction and label columns shows no significant change. What is the most likely interpretation of this situation?

    Answer and explanation

    Correct answer: B

    A high PSI for an input feature indicates that the distribution of that feature has changed significantly between the baseline and current data (feature drift). The fact that prediction and label distributions are stable suggests that this change has not yet affected the model's output or the underlying data relationships. This could be because the feature has low importance, or the model is robust to this specific change. It serves as an early warning that requires investigation.

  9. Question 9

    A financial institution is building a real-time transaction fraud detection system. Latency is critical, as predictions must be returned in under 50 milliseconds. The features for this model require complex, on-the-fly calculations based on the user's recent activity, which is not available in the batch feature store. Which Databricks solution is best suited for this requirement?

    Answer and explanation

    Correct answer: B

    Databricks Feature Serving is designed for use cases that require ultra-low latency and on-demand feature computation. It allows you to define functions that compute features at inference time using data provided in the request. This avoids the need to pre-compute and store all possible feature values, making it ideal for real-time applications with dynamic feature requirements.

  10. Question 10

    Case Study:

    A global logistics company, ShipFast, wants to build an MLOps platform on Databricks to manage hundreds of models that predict package delivery times. Their key requirements are strict separation of development, staging, and production environments, auditable model transitions, and automated testing before any model is promoted to production.

    The current process is manual, with data scientists promoting models via the UI, leading to inconsistent testing and accidental deployments. The MLOps team has been tasked with designing a fully automated, code-driven CI/CD pipeline. The pipeline must enforce that any model version proposed for the 'Staging' stage must first pass a suite of integration tests, including performance evaluation on a holdout dataset and a bias check. If the tests pass, the model version should be automatically transitioned to 'Staging' with a comment linking to the CI job results.

    The MLOps team decides to use Databricks Jobs and Model Registry webhooks. They create a multi-task job that checks out the code, runs the tests, and on success, transitions the model. They need to ensure this job is triggered securely and reliably whenever a data scientist registers a new model version.

    What is the most secure and robust way to architect the trigger mechanism for this validation pipeline?

    Answer and explanation

    Correct answer: B

    This approach is the most secure, integrated, and aligned with the requirements. Using a Job webhook directly links the Model Registry event to a Databricks Job without external middleware. Triggering on TRANSITION_REQUEST_CREATED instead of MODEL_VERSION_CREATED correctly implements the business logic: the tests run when a promotion is requested, acting as a gatekeeper. This allows data scientists to register many experimental versions without triggering a pipeline for each one, only for those they wish to promote.

Register free to unlock 10 more sample questions

Lifetime One

Own this practice test forever.

$79.99
$75.99
one-time
  • Full access to 205 questions
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • Brainy AI Assistant
  • Lifetime updates

Two

Any 2 exams per month.

$20.00/exam
$39.99
/month
  • 2 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 1,000 Brainy AI Credits
  • Cancel anytime

Premium Twelve

Any 12 exams over 3 months.

$15.00/exam
$179.99
/3 months
  • 4 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 15,000 Brainy AI Credits
  • Dedicated support
  • Friend seat included — full access

Trusted by professionals at

NvidiaSupabaseGitHubOpenAITursoClerkClaude AIAmazon