Why Study with PlanetCert?
The Latest Questions
Practice questions and exam topics aligned with the current exam objectives.
Detailed Explanations
Go beyond the answer. Master the material with comprehensive learning and professional explanations for every concept.

AI-Powered Insights
Personalized preparation guidance that adapts to your performance and identifies weak spots automatically.
Exam Information
Official specifications published by Databricks
Exam Format
Registration
Validity
ML-ASSOC Exam Topics and Domains
ML-ASSOC is organized into 4 weighted domains. Expect to work with Spark ML, scikit-learn, Delta Live Tables, Feature Store, and more.
Databricks Machine Learning
MLOps Strategy and Best Practices
- Identify the best practices of an MLOps strategy
- Identify the advantages of using ML runtimes
AutoML
- Identify how AutoML facilitates model/feature selection
- Identify the advantages AutoML brings to the model development process
Feature Store
- Identify the benefits of creating feature store tables at the account level in Unity Catalog vs at the workspace level
- Create a feature store table in Unity Catalog
- Write data to a feature store table
- Train a model with features from a feature store table
- Score a model using features from a feature store table
- Describe the differences between online and offline feature tables
MLflow
- Identify the best run using the MLflow Client API
- Manually log metrics, artifacts, and models in an MLflow Run
- Identify information available in the MLflow UI
- Register a model using the MLflow Client API in the Unity Catalog registry
- Identify benefits of registering models in the Unity Catalog registry over the workspace registry
- Identify scenarios where promoting code is preferred over promoting models and vice versa
- Set or remove a tag for a model
- Promote a challenger model to a champion model using aliases
Data Processing
Exploratory Data Analysis
- Compute summary statistics on a Spark DataFrame using .summary() or dbutils data summaries
- Create visualizations for categorical or continuous features
- Compare two categorical or two continuous features using the appropriate method
Data Cleaning
- Remove outliers from a Spark DataFrame based on standard deviation or IQR
- Compare and contrast imputing missing values with the mean or median or mode value
- Impute missing values with the mode, mean, or median value
Feature Engineering
- Use one-hot encoding for categorical features
- Identify and explain the model types or data sets for which one-hot encoding is or is not appropriate
- Identify scenarios where log scale transformation is appropriate
Model Development
ML Foundations and Algorithm Selection
- Use ML foundations to select the appropriate algorithm for a given model scenario
- Identify methods to mitigate data imbalance in training data
Spark ML Pipelines
- Compare estimators and transformers
- Develop a training pipeline
Hyperparameter Tuning
- Use Hyperopt's fmin operation to tune a model's hyperparameters
- Perform random or grid search or Bayesian search as a method for tuning hyperparameters
- Parallelize single node models for hyperparameter tuning
Model Validation
- Describe the benefits and downsides of using cross-validation over a train-validation split
- Perform cross-validation as a part of model fitting
- Identify the number of models being trained in conjunction with a grid-search and cross-validation process
Model Evaluation
- Use common classification metrics: F1, Log Loss, ROC/AUC, etc
- Use common regression metrics: RMSE, MAE, R-squared, etc
- Choose the most appropriate metric for a given scenario objective
- Identify the need to exponentiate log-transformed variables before calculating evaluation metrics or interpreting predictions
Model Complexity
Assess the impact of model complexity and the bias variance tradeoff on model performance
Model Deployment
Model Serving Approaches
Identify the differences and advantages of model serving approaches: batch, realtime, and streaming
Model Deployment Implementation
- Deploy a custom model to a model endpoint
- Use pandas to perform batch inference
- Identify how streaming inference is performed with Delta Live Tables
- Deploy and query a model for realtime inference
- Split data between endpoints for realtime inference
How do I earn this certification?
Passing ML-ASSOC earns the Databricks Certified Machine Learning Associate certification. It sits in the Machine Learning track.
- Databricks-Certified-Data-Engineer-Associate - Data Engineer AssociateFoundation for ML data pipelines
- Databricks-Certified-Data-Engineer-Professional - Data Engineer ProfessionalAdvanced data engineering for ML workflows
- Databricks-Certified-Data-Analyst-Associate - Data Analyst AssociateSQL and analytics skills for ML insights
- Databricks-Certified-Generative-AI-Engineer-Associate - Generative AI Engineer AssociateSpecialized in GenAI and LLM applications
- Databricks-Certified-Apache-Spark-Developer-Associate - Apache Spark Developer Associate Deep Spark expertise for distributed ML
Practice with Precision
The PlanetCert Simulator mirrors the real exam environment with authentic questions and timed pressure.
How to study for this exam?
The most effective way to prepare for ML-ASSOC is by using the PlanetCert Simulator to practice questions and review detailed explanations.
What's changed on this exam?
- ACTIVE
- Last content update: 2025-03-01
- Announcement date: 2025-01-15
- Unity Catalog Latest Critical - 38% of exam focuses on UC integration • Release date: 2025-06-12
- MLflow 2.x High - Core component for experiment tracking and model registry • Release date: Ongoing updates
- AutoML Latest Medium - Understanding when to use AutoML vs manual development • Release date: Continuous improvements
- Feature Store With FeatureEngineeringClient High - New API emphasis in exam v2.0 • Release date: 2024-2025
- Delta Live Tables Latest Medium - Streaming inference implementation • Release date: Ongoing
Who should take this exam?
This exam is typically taken by Data Scientists and Machine Learning Engineers.
- 6+ months of hands-on experience performing machine learning tasks
- Working knowledge of Python and major ML libraries like scikit-learn and SparkML
- Working knowledge of Unity Catalog and Databricks data management features
- Familiarity with Databricks ML documentation