Databricks Machine Learning Associate Free Sample Questions

20 free sample questions210 in the full practice test

Try simulator

ML-ASSOC Sample Questions

  1. Question 1

    A large financial services company operates multiple Databricks workspaces for different business units. They need to develop a centralized repository of customer features (e.g., credit score, transaction frequency) that can be securely shared and reused across all workspaces, with strict access controls managed by a central governance team. Which approach best meets these requirements?

    Answer and explanation

    Correct answer: C

    Creating Feature Store tables within Unity Catalog at the account level is the correct approach. Unity Catalog provides a centralized governance and access control layer that spans across all workspaces within an account. This allows the central governance team to manage permissions on a single set of feature tables, which can then be securely accessed by authorized users and services from any workspace, fulfilling the core requirements of centralization and secure sharing.

  2. Question 2

    True or False: In the context of the bias-variance tradeoff, increasing a model's complexity (e.g., adding more layers to a neural network or increasing the depth of a decision tree) will generally decrease its bias but increase its variance.

    Answer and explanation

    Correct answer: A

    This statement is true. Increasing model complexity allows the model to learn more intricate patterns from the training data, which reduces its bias (the error from erroneous assumptions in the learning algorithm). However, a more complex model is also more likely to learn the noise in the training data, making it more sensitive to variations in the data. This increased sensitivity is known as higher variance, which can lead to overfitting and poor generalization to new, unseen data.

  3. Question 3

    A data scientist is cleaning a dataset containing employee salary information for a large corporation. They observe that the 'salary' column has a number of extreme outliers, including several C-level executive salaries that are orders of magnitude higher than the rest of the employees. For a feature engineering task, they need to impute a few missing salary values. Which imputation method should they prefer and why?

    Answer and explanation

    Correct answer: B

    The median is the correct choice because it is a robust measure of central tendency, meaning it is not significantly affected by extreme outliers. The mean, on the other hand, is sensitive to outliers; the very high executive salaries would pull the mean upwards, making it an unrepresentative value for the typical employee. Using the median ensures the imputed value is closer to the center of the majority of the data points.

  4. Question 4

    A leading e-commerce company wants to deploy a new product recommendation system. The system has several complex requirements:

    1. Real-time Personalization: The model must provide recommendations within 200ms of a user's action (e.g., viewing a product).
    2. Dynamic User Profiles: User feature vectors, which are used for inference, must be updated in near real-time based on their clickstream data.
    3. A/B Testing: The MLOps team must be able to deploy a new 'challenger' recommendation algorithm and route 10% of live traffic to it for evaluation against the current 'champion' model.
    4. Scalability: The system must handle traffic spikes during holiday seasons, scaling automatically without manual intervention.

    Which Databricks deployment architecture best fulfills all these requirements?

    graph TD subgraph User Interaction WebApp[Web Application] --> API_GW[API Gateway] end subgraph Real-time Processing Clickstream[Kafka: User Events] --> DLT[Delta Live Tables: User Profile Update] DLT --> OnlineStore[Online Feature Store] end subgraph Model Serving API_GW --> Endpoint[Model Serving Endpoint] Endpoint -- 90% --> Champion[Champion Model] Endpoint -- 10% --> Challenger[Challenger Model] Champion --> OnlineStore Challenger --> OnlineStore end subgraph Batch Processing ProductCatalog[Batch: Product Catalog] --> OfflineStore[Offline Feature Store] end

    Answer and explanation

    Correct answer: C

    This architecture correctly addresses all requirements. A serverless Model Serving endpoint provides low-latency (<200ms) inference and automatic scaling. Deploying multiple models to a single endpoint with traffic splitting directly enables A/B testing. An Online Feature Store, updated by a streaming process like DLT, provides the dynamic, low-latency feature vectors needed for real-time personalization.

  5. Question 5

    A machine learning engineer is using the FeatureEngineeringClient to create a new feature table in Unity Catalog. The table will store user features, and it's critical that each user is uniquely identified and that features can be looked up efficiently for online serving. Which parameter in the fe.create_table method is used to specify the unique identifier column(s) for the entities in the table?

    Answer and explanation

    Correct answer: B

    The primary_keys parameter is used to specify the column or columns that uniquely identify each row or entity in the feature table. This is crucial for the Feature Store as it uses these keys for joining features during training set creation and for efficient lookups in online stores.

  6. Question 6

    A data scientist is using Hyperopt to perform Bayesian hyperparameter optimization for a machine learning model. They need to supply the correct search algorithm to the algo parameter of the fmin function. Which of the following options implements a Bayesian approach, specifically Tree-structured Parzen Estimator?

    Answer and explanation

    Correct answer: C

    tpe.suggest stands for Tree-structured Parzen Estimator suggestion algorithm, which is a form of Bayesian optimization available in Hyperopt. It intelligently chooses the next set of hyperparameters to evaluate based on the results of previous trials, making it more efficient than random or grid search.

  7. Question 7

    Multiple answers

    An MLOps team is managing a critical fraud detection model registered in Unity Catalog. The current production model is aliased as 'Champion'. A new version has been validated and is ready to be promoted. To minimize risk, the team needs a strategy that allows for an immediate rollback to the previous version if the new model underperforms. Which steps should they perform? (Select TWO)

    Answer and explanation

    Correct answers: B, D

    Setting the alias on the new version makes it the active production model.

    Applying a new, descriptive alias to the old version (like 'Previous_Champion' or 'Archived_Champion') maintains a reference to it, making it easy to find and re-alias as 'Champion' for a quick rollback.

  8. Question 8

    During exploratory data analysis for a demand forecasting model, a data scientist observes that a key feature, user_daily_logins, has a strong positive skew, with most users logging in 1-2 times a day but a small number of power users logging in over 50 times. This skew could negatively impact the performance of a linear regression model. Which feature transformation is most appropriate to apply to this feature?

    Answer and explanation

    Correct answer: D

    A logarithmic transformation is the most appropriate method for handling features with a strong positive skew and a wide range of values. It compresses the range of the large values more than the small values, pulling the long tail in and making the distribution more symmetrical and closer to a normal distribution. This often improves the performance and stability of linear models.

  9. Question 9

    A machine learning team is using GridSearchCV from scikit-learn to tune a Gradient Boosting model. The parameter grid is defined as follows:
    param_grid = {'n_estimators': [100, 200], 'learning_rate': [0.01, 0.1, 0.2], 'max_depth': [3, 5, 7]}
    They are using 5-fold cross-validation (cv=5). How many individual models will be trained during this entire hyperparameter tuning process?

    Answer and explanation

    Correct answer: C

    The total number of models trained is the product of the number of parameter combinations and the number of cross-validation folds. First, calculate the number of combinations in the grid: 2 (for n_estimators) * 3 (for learning_rate) * 3 (for max_depth) = 18 combinations. Then, multiply this by the number of folds: 18 combinations * 5 folds = 90 models.

  10. Question 10

    A logistics company needs to deploy a model for real-time anomaly detection in its package delivery event stream. The pipeline must process a high volume of events with fluctuating loads and requires a solution that simplifies infrastructure management and automatically handles cluster scaling. What is the primary advantage of using Delta Live Tables (DLT) for this streaming inference task compared to a manually configured Structured Streaming job?

    Answer and explanation

    Correct answer: C

    The primary advantage of Delta Live Tables is its declarative nature. Developers define the desired outcome (the dataflows), and DLT manages the underlying infrastructure, including automatic cluster scaling, error handling, and data quality checks. This simplifies the operational burden compared to a standard Structured Streaming job, which requires manual configuration and management of the cluster and job settings.

Register free to unlock 10 more sample questions

Lifetime One

Own this practice test forever.

$79.99
$75.99
one-time
  • Full access to 210 questions
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • Brainy AI Assistant
  • Lifetime updates

Two

Any 2 exams per month.

$20.00/exam
$39.99
/month
  • 2 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 1,000 Brainy AI Credits
  • Cancel anytime

Premium Twelve

Any 12 exams over 3 months.

$15.00/exam
$179.99
/3 months
  • 4 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 15,000 Brainy AI Credits
  • Dedicated support
  • Friend seat included — full access

Trusted by professionals at

NvidiaSupabaseGitHubOpenAITursoClerkClaude AIAmazon