Databricks Data Engineer Associate Free Sample Questions

20 free sample questions207 in the full practice test

Try simulator

DATA-ENG-ASSOC Sample Questions

  1. Question 1

    A data engineering team is migrating its development workflow from the Databricks UI to a local IDE using Databricks Connect. A junior engineer successfully sets up their connection profile but receives a Py4JError upon trying to initialize a SparkSession. The cluster they are connecting to runs Databricks Runtime 14.3, which uses Python 3.11.2. The engineer's local environment is running Python 3.11.5. What is the primary reason for this connection failure?

    Answer and explanation

    Correct answer: B

    Databricks Connect requires that the major and minor Python versions of the local client environment match the version on the Databricks cluster exactly. In this case, the local version is 3.11.5 while the cluster version is 3.11.2. Even though the major version (3) and minor version (11) match, the strict requirement often extends to the patch version for full compatibility, but the major/minor mismatch is the key principle. An expired PAT would result in an authentication error, not a Py4JError. A firewall issue would likely cause a timeout. An incorrect cluster ID would result in a 'not found' error.

  2. Question 2

    Multiple answers

    A DevOps team is implementing a CI/CD pipeline to deploy a multi-task Databricks workflow using Databricks Asset Bundles (DAB). The pipeline must handle deployments to development, staging, and production workspaces, each with different compute policies and secret scopes. Which TWO components of the databricks.yml file are essential for managing these environment-specific configurations? (Select TWO)

    Answer and explanation

    Correct answers: B, E

    The targets section is specifically designed to manage environment-specific configurations. Each target can have its own workspace URL, root path, and variable definitions, allowing the same bundle to be deployed to different environments with the correct settings.

    Variables, defined within each target, allow for parameterization of the bundle's resources. This is how you would specify a different cluster policy ID for production versus development, or reference a different secret scope for database credentials in each environment.

  3. Question 3

    A streaming pipeline using Auto Loader is configured to ingest JSON files from a cloud storage location. The pipeline runs successfully for several weeks but suddenly fails. Investigation of the _rescue column reveals that several recent files contain records where a previously numeric transaction_amount field is now a string (e.g., "100.50" instead of 100.50). The desired behavior is to automatically adapt the target Delta table schema to accommodate this change without manual intervention. Which Auto Loader option should have been configured to handle this situation gracefully?

    Answer and explanation

    Correct answer: D

    The rescue mode for cloudFiles.schemaEvolutionMode is designed for this exact scenario. It instructs Auto Loader to infer the schema and, if a data type change is detected (like numeric to string), it adds the new column to the _rescue column. For schema evolution, the correct option is to enable schema evolution on the write stream using .option("mergeSchema", "true") and handle rescued data. However, among the given choices, rescue mode is the feature specifically designed to handle columns with mixed data types by rescuing them, which aligns with the observed behavior. While mergeSchema is also needed, schemaEvolutionMode rescue directly addresses the problem of incompatible data types appearing in a column.

  4. Question 4

    A data architect is designing a Medallion architecture for a financial services company. The raw data (Bronze layer) contains sensitive personally identifiable information (PII). The Silver layer must contain the same records but with all PII columns pseudonymized. The Gold layer will contain aggregated data with no PII. The compliance team requires that only a specific service principal, used by an automated cleansing job, can read the raw PII data from the Bronze layer. All other users and groups should be denied access. How should this security requirement be implemented using Unity Catalog?

    Answer and explanation

    Correct answer: B

    Unity Catalog operates on a default-deny model. To meet the requirement, you should avoid granting any broad permissions at the catalog or schema level. The correct approach is to grant the necessary, specific privileges (SELECT to read, MODIFY might be needed for the job's operations) directly to the service principal on the target Bronze tables. This ensures that only that principal can access the sensitive data, enforcing the principle of least privilege.

  5. Question 5

    True or False: When a Databricks job cluster is configured with a cluster pool, it can start faster because it acquires its driver and worker nodes from the pool of idle instances, reducing the time spent waiting for the cloud provider to provision new virtual machines.

    Answer and explanation

    Correct answer: A

    This statement is true. The primary purpose of cluster pools is to reduce cluster start and auto-scaling times by maintaining a set of idle, ready-to-use instances. When a job requests a cluster attached to a pool, it gets its nodes from this warm pool instead of requesting new instances from the cloud provider, which significantly shortens the startup latency.

  6. Question 6

    During a code review, a senior engineer observes the following PySpark code snippet intended to update customer records based on new transactions. What is the primary issue with this approach for transforming data from a Bronze to a Silver table in a Medallion architecture?

    # bronze_df is the raw, unvalidated source DataFrame
    # silver_table is the path to the clean, validated Delta table
    
    (bronze_df.write
    .format("delta")
    .mode("overwrite")
    .save(silver_table))
    
    Answer and explanation

    Correct answer: A

    The primary purpose of the Silver layer is to store cleansed, validated, and enriched data. This code snippet moves data directly from the source (Bronze) to Silver using a blind overwrite without any intermediate transformation steps. This violates the core principle of the Medallion architecture, as it bypasses the crucial data quality and shaping processes that should occur between the Bronze and Silver layers.

  7. Question 7

    A data engineer is building a Delta Live Tables (DLT) pipeline. They need to define a table that combines streaming data from a Kafka source with a static dimension table from Unity Catalog for enrichment. The pipeline should enforce a quality constraint: the join key from the streaming source must not be null. If a record violates this constraint, it should be dropped, and the pipeline should continue processing valid records. Which DLT function and expectation clause should be used?

    graph TD A[Kafka Stream] --> C{DLT Pipeline}; B[UC Dimension Table] --> C; C -->|Join & Enrich| D[Silver Table]; D -->|CONSTRAINT key IS NOT NULL| E{Quality Check}; E -->|Violation| F[Drop Row]; E -->|Valid| G[Process Row];
    Answer and explanation

    Correct answer: C

    To create a physical table that can be queried later, @dlt.table is the correct decorator. To enforce a quality rule where invalid records are dropped while allowing the pipeline to continue, @dlt.expect_or_drop is the appropriate function. @dlt.expect_or_fail would halt the pipeline, and @dlt.expect would only record the violation statistics without dropping or failing.

  8. Question 8

    A company wants to provide its external partners with read-only access to a curated sales dataset managed in Unity Catalog. The partners do not have Databricks workspaces. The data engineering team needs to set up a secure sharing mechanism that does not require creating and managing users within their Databricks account. Which technology should be used to achieve this?

    Answer and explanation

    Correct answer: C

    Delta Sharing is the open protocol specifically designed for sharing live data from Databricks to any external consumer, regardless of their platform. It works by creating shares and recipients, then providing the recipient with a secure, one-time activation link to download a credential file. This allows them to access the shared data using various connectors (like Power BI, Tableau, pandas) without needing a Databricks account. Lakehouse Federation is for querying external data sources from within Databricks, not sharing data out.

  9. Question 9

    An organization is looking to optimize its ad-hoc analytics query performance and reduce infrastructure management overhead. The analytics team runs a large number of concurrent, short-running queries against Gold tables throughout the day. The workload pattern is highly variable. Which type of compute resource is the best fit for this scenario?

    Answer and explanation

    Correct answer: B

    Serverless SQL warehouses are ideal for this use case. They provide instant compute, automatically scale up and down to handle concurrent query loads, and eliminate the need for administrators to manage cluster configurations. This directly addresses the requirements for performance on variable workloads and reduced management overhead. A job cluster is for automated jobs, and an all-purpose cluster requires manual management and is less efficient for handling concurrent SQL queries.

  10. Question 10

    A data engineer needs to write a PySpark DataFrame containing daily sales aggregates to a Delta table. The operation must insert new sales records and update the total sales amount for existing dates that are already in the target table. Which operation should be used to accomplish this combined insert and update logic efficiently?

    Answer and explanation

    Correct answer: B

    The merge operation is specifically designed for this type of 'upsert' (update or insert) scenario. It allows you to define conditions for matching records between a source DataFrame and a target Delta table. The whenMatchedUpdate clause specifies the update logic for existing records, while the whenNotMatchedInsert clause handles the insertion of new records, all within a single, atomic transaction.

Register free to unlock 10 more sample questions

Lifetime One

Own this practice test forever.

$79.99
$75.99
one-time
  • Full access to 207 questions
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • Brainy AI Assistant
  • Lifetime updates

Two

Any 2 exams per month.

$20.00/exam
$39.99
/month
  • 2 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 1,000 Brainy AI Credits
  • Cancel anytime

Premium Twelve

Any 12 exams over 3 months.

$15.00/exam
$179.99
/3 months
  • 4 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 15,000 Brainy AI Credits
  • Dedicated support
  • Friend seat included — full access

Trusted by professionals at

NvidiaSupabaseGitHubOpenAITursoClerkClaude AIAmazon