SAS 9.4 Base Programming - Performance-Based Exam Free Sample Questions

20 free sample questions217 in the full practice test

Try simulator

A00-231 Sample Questions

  1. Question 1

    A financial analyst is tasked with calculating cumulative quarterly sales for multiple regions from a dataset named WORK.SALES, sorted by Region and SaleDate. The analyst needs to reset the cumulative total for each new region. Which of the following DATA step code snippets correctly implements this logic?

    Answer and explanation

    Correct answer: C

    This is the correct approach. The RETAIN statement is crucial to hold the value of CumulativeSales across observations. The conditional if FIRST.Region then CumulativeSales = 0; correctly resets the accumulator at the beginning of each new Region group. The final assignment statement CumulativeSales = CumulativeSales + SalesAmount; is incorrect syntax for a sum statement but works as a simple assignment; the proper sum statement would be CumulativeSales + SalesAmount;. However, of the options, this is the most complete and functional logic. The sum statement CumulativeSales + SalesAmount; implicitly retains the variable.

  2. Question 2

    A junior programmer executes a DATA step to calculate a new variable, Ratio, by dividing ValueA by ValueB. The SAS log shows no errors or warnings, but a subsequent PROC MEANS reveals that the mean of Ratio is much lower than expected. Upon manual inspection, the programmer finds that ValueB is sometimes zero. How does the SAS DATA step handle division by zero by default, and what message should the programmer have looked for in the log?

    Answer and explanation

    Correct answer: C

    By default, SAS handles division by zero by assigning a missing value (.) to the result variable for that observation. It does not stop the DATA step. It prints a 'NOTE: Division by zero.' message to the log, along with the line number and column where it occurred. This is a common source of logic errors, as the program runs to completion but produces incorrect or incomplete results.

  3. Question 3

    Multiple answers

    A marketing dataset contains a CampaignID field with values like 'FY24-Q3-EMAIL-PROMO123'. A data analyst needs to extract the fiscal year, the quarter, and the campaign type ('EMAIL') into separate variables. Which of the following SAS functions are required to accomplish this task? (Select THREE)

    Answer and explanation

    Correct answers: A, B, D

    The SCAN function is essential for parsing strings based on delimiters. It can be used with the '-' delimiter to extract the 1st ('FY24'), 2nd ('Q3'), and 3rd ('EMAIL') 'words' from the CampaignID string.

    The SUBSTR function can be used to extract parts of the string based on position. For example, SUBSTR(CampaignID, 1, 4) could get 'FY24'. While SCAN is more robust for this specific task, SUBSTR is also a valid function for extracting substrings and would be needed to get the 'FY' part from the first word.

    The INPUT function is necessary to convert the character year '24' (extracted from 'FY24') and quarter '3' (extracted from 'Q3') into numeric variables if numerical analysis is required later. This is a common step after parsing.

  4. Question 4

    True or False: When merging two SAS datasets using a MERGE statement with a BY statement, SAS requires both datasets to be sorted by the BY variables. If they are not sorted, SAS will stop with an error and halt program execution.

    Answer and explanation

    Correct answer: B

    This statement is false. While it is a best practice and logically necessary for a correct merge, SAS will not stop with an error by default if the data is not sorted. Instead, it will issue an 'ERROR: BY variables are not properly sorted' message in the log and continue processing, often producing an incorrect result. This is a common source of logic errors.

  5. Question 5

    A data scientist is using PROC IMPORT to read a large CSV file ('c:\data\survey.csv') where the first 50 rows contain metadata and notes, with the actual column headers in row 51. Which combination of options in the PROC IMPORT statement is required to correctly read the data, starting from the headers in row 51?

    Answer and explanation

    Correct answer: D

    The GETNAMES=YES statement tells SAS to use the first row it reads as variable names. The DATAROW=n option specifies the first row of data to read. Since the headers are in row 51, the data itself begins on row 52. Therefore, DATAROW=52 is the correct option to use in conjunction with GETNAMES=YES.

  6. Question 6

    A reporting analyst at a retail company needs to generate a weekly sales report. The requirements are as follows:

    1. The final report must be an Excel file named weekly_sales.xlsx for business users.
    2. A PDF version named archive_sales.pdf must be created for archival purposes.
    3. The report should contain summary statistics from PROC MEANS and frequency tables from PROC FREQ.
    4. However, the archived PDF should NOT contain the detailed frequency tables from PROC FREQ to save space.

    Which ODS code block correctly fulfills all these requirements?

    Answer and explanation

    Correct answer: D

    This is the correct solution. It opens both ODS destinations (EXCEL and PDF) at the beginning. It then uses ods pdf exclude Freq; to specifically tell the PDF destination to ignore any output from PROC FREQ. The EXCEL destination is not affected and will receive output from both procedures. Finally, ods _all_ close; is the best practice to close all open destinations. This approach is targeted and correctly implements the specific exclusion requirement for only one of the destinations.

  7. Question 7

    A dataset named PATIENT_METRICS is in a 'wide' format, with one row per patient and separate columns for measurements taken on different dates (Weight_Day1, Weight_Day7, Weight_Day30). To perform a time-series analysis, the data needs to be restructured into a 'long' format with columns PatientID, Day, and Weight. Which PROC TRANSPOSE step correctly performs this transformation?

    Answer and explanation

    Correct answer: C

    This is the most efficient and correct solution. The BY PatientID; statement ensures each patient's data is processed as a group and PatientID is retained in the output. The VAR Weight_:; statement uses a variable list to select all variables starting with 'Weight_', which is a flexible way to include all measurement columns. By default, PROC TRANSPOSE will create a _NAME_ variable with the original variable names (Weight_Day1, etc.) and a COL1 variable with their values. The RENAME= option is then used to give these new columns more meaningful names. Subsequent data steps would be needed to parse the day number from the Measurement variable.

  8. Question 8

    You need to assign a permanent library reference named PROJECT1 to a directory located at /user/data/project1. The data in this directory is a mix of SAS datasets and Microsoft Excel files. You want to use the XLSX engine to directly read the Excel files. Which LIBNAME statement is syntactically correct for this purpose?

    Answer and explanation

    Correct answer: B

    The correct syntax for a LIBNAME statement is LIBNAME libref [engine] 'path' [options];. The engine name, XLSX in this case, is specified directly after the libref and before the path. The ENGINE= keyword is not used in this context.

  9. Question 9

    A developer is writing a program to categorize customers into one of four distinct, non-overlapping tiers based on their TotalSpending. Which conditional logic structure is generally more efficient and readable for this task compared to a series of IF-THEN/ELSE IF statements?

    Answer and explanation

    Correct answer: C

    For evaluating a single variable against multiple mutually exclusive conditions, a SELECT group is often more efficient and easier to read than a long chain of IF-THEN/ELSE IF statements. The SELECT statement evaluates the expression once and then compares the result to each WHEN condition, which can be faster than re-evaluating conditions in an IF chain.

  10. Question 10

    A hospital administrator wants to create a report that groups patient satisfaction scores (Score, ranging 1-100) into descriptive categories: 'Poor' (1-40), 'Average' (41-70), 'Good' (71-90), and 'Excellent' (91-100). Any other score should be labeled 'Invalid'. Which PROC FORMAT code correctly defines this custom format?

    Answer and explanation

    Correct answer: A

    This code is correct. The VALUE statement defines a format named ScoreFmt. It correctly specifies numeric ranges for each category. The keyword OTHER is used as a catch-all for any values that do not fall into the specified ranges, which is the proper way to handle unexpected or invalid data.

Register free to unlock 10 more sample questions

Lifetime One

Own this practice test forever.

$79.99
$75.99
one-time
  • Full access to 217 questions
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • Brainy AI Assistant
  • Lifetime updates

Two

Any 2 exams per month.

$20.00/exam
$39.99
/month
  • 2 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 1,000 Brainy AI Credits
  • Cancel anytime

Premium Twelve

Any 12 exams over 3 months.

$15.00/exam
$179.99
/3 months
  • 4 active exam slots
  • Study, Timed & Flashcard Modes
  • All past and future versions i
  • Detailed Explanations
  • Study Tracking & Past Attempts
  • 15,000 Brainy AI Credits
  • Dedicated support
  • Friend seat included — full access

Trusted by professionals at

NvidiaSupabaseGitHubOpenAITursoClerkClaude AIAmazon