AI Testing Free Sample Questions

20 free sample questions167 in the full practice test

Try simulator

CT-AI Sample Questions

  1. Question 1

    Advanced

    Methods and Techniques for the Testing of AI-Based Systems · Metamorphic Testing Application

    A financial institution has developed an AI model to assess loan application risk. The model is a deep neural network, making it a "black box" where the logic for a specific decision is not easily explainable. The test team faces a significant test oracle problem, as calculating the "correct" risk score for a new applicant profile is infeasible. They decide to use Metamorphic Testing (MT). Which of the following represents the MOST effective Metamorphic Relation (MR) for testing the logical consistency of this loan risk model?

    Answer and explanation

    Correct answer: C

    This is the most effective Metamorphic Relation because it tests a fundamental logical property of a loan risk model without needing a precise expected output. The property is that, all other factors being equal, a higher income should not increase financial risk. This allows the team to verify the model's logical consistency even when the exact risk score calculation is unknown. The other options are less effective: a tiny income change may not trigger a score change, duplicating the test case only checks for determinism, and assuming proportionality is not a guaranteed property.

  2. Question 2

    IntermediateMultiple answers

    Testing AI-Specific Quality Characteristics · Testing for Inappropriate Bias

    A QA team is testing a new AI-powered hiring tool that screens resumes to identify top candidates. The company is concerned about introducing unintentional bias against protected groups. The test data includes demographic information, but this data is NOT used as a feature for the model's prediction. Despite demographic data not being an input feature, the model may still exhibit bias. Which TWO of the following testing activities are most crucial for uncovering this hidden bias? (Select TWO)

    Answer and explanation

    Correct answers: B, C

    This method, known as disparate impact analysis or statistical parity testing, is a primary technique for detecting outcome-based bias, regardless of the input features. It directly measures whether the model's decisions are fair in practice.

    This is a critical step in identifying the root cause of hidden bias. Proxy variables are features that are not explicitly sensitive but act as stand-ins for protected characteristics, introducing bias through the training data itself.

  3. Question 3

    Intermediate

    Using AI for Testing · AI for Test Case Generation

    A DevOps team wants to improve its testing of a complex REST API with hundreds of endpoints and intricate dependencies. Manual test case creation is slow and often misses complex interaction bugs. They decide to use an AI-based tool for test case generation that uses a search-based algorithm (like a genetic algorithm) to explore the API's behavior. What is the primary challenge the team will face when integrating this AI-driven test generation tool into their CI/CD pipeline?

    Answer and explanation

    Correct answer: C

    This is the fundamental limitation of most AI-based test generation tools. While excellent at exploring the state space and generating novel inputs (the 'how to test'), they lack the domain-specific knowledge to determine what the correct output should be. This is known as the test oracle problem. The most common use for such tools is to find crashes or server errors (5xx), for which the oracle is simple, rather than to find subtle logical bugs.

  4. Question 4

    Intermediate

    Quality Characteristics for AI-Based Systems · Explainability (XAI)

    A hospital is deploying an AI system to predict the likelihood of sepsis in ICU patients. Due to the critical nature of the decisions, regulations require that the model's predictions be explainable to clinicians. The development team chose a deep neural network (DNN) for its high accuracy. Which technique would be most appropriate for the testing team to use to validate the explainability requirement for this high-stakes, black-box model?

    Answer and explanation

    Correct answer: B

    For a complex black-box model like a DNN, direct inspection of weights is not human-interpretable. LIME is designed specifically for this use case: providing local, understandable explanations for individual predictions of any black-box model. This allows a clinician to ask 'Why did the model flag this specific patient?' and get an answer based on the most influential features (e.g., 'because of high heart rate and low blood pressure'), making it ideal for validating explainability in a clinical setting.

  5. Question 5

    Beginner

    ML - Neural Networks and Testing · Coverage Measures for Neural Networks

    True or False: Achieving 100% neuron coverage in a deep neural network guarantees that all logical paths within the model have been tested and that the model is free from defects.

    Answer and explanation

    Correct answer: B

    The statement is false. Neuron coverage is a very weak structural coverage criterion, analogous to statement coverage in traditional code. It only ensures that each neuron has produced an output above a certain threshold at least once. It does not test the complex interactions between neurons, the different activation ranges, or the combinatorial logic of the network. Therefore, it provides no guarantee that the model is free of logical flaws or defects.

  6. Question 6

    Intermediate

    ML - Data · Data Labelling

    A startup is developing a supervised learning model to identify defective products on an assembly line from camera images. They have 1 million images but lack the in-house staff to label them. They need to get the data labeled quickly and cost-effectively, while managing the risk of incorrect labels. Which data labeling strategy offers the best balance of speed, cost-effectiveness, and quality control for this scenario?

    Answer and explanation

    Correct answer: C

    This is a standard industry best practice for large-scale labeling tasks. Crowdsourcing is fast and cost-effective for a large dataset. The key to quality control in this approach is redundancy: having multiple independent workers label the same data item. A consensus mechanism, like a majority vote, is then used to establish a higher-confidence ground truth label, effectively filtering out random errors and individual worker biases. This balances speed, cost, and quality.

  7. Question 7

    AdvancedMultiple answers

    Methods and Techniques for the Testing of AI-Based Systems · Adversarial Attacks and Data Poisoning

    A security testing team is evaluating the robustness of a traffic sign recognition model for an autonomous vehicle. They are concerned about both data poisoning during training and adversarial attacks in production. Which TWO of the following test activities should the team perform to assess the model's vulnerability to these specific threats? (Select TWO)

    Answer and explanation

    Correct answers: A, D

    This activity directly simulates a data poisoning attack. By intentionally corrupting a portion of the training data, testers can evaluate the model's resilience and determine if it can be manipulated into making specific, incorrect classifications.

    This describes the process of an adversarial attack. The goal is to create inputs that are visually indistinguishable from legitimate ones to a human but are specifically crafted to fool the model. Testing with such examples is crucial for assessing robustness against malicious attacks in production.

  8. Question 8

    Beginner

    Machine Learning (ML) – Overview · Forms of ML

    An e-commerce company wants to analyze its customer purchase history to discover which products are frequently bought together (e.g., 'customers who buy hot dogs also tend to buy hot dog buns'). The goal is to use these findings for product placement and marketing campaigns. The dataset contains transaction records but no pre-defined labels. Which form of Machine Learning is most suitable for this task?

    Answer and explanation

    Correct answer: C

    The scenario describes market basket analysis, which is a classic example of association rule mining. Since the goal is to discover hidden patterns and relationships in unlabeled transactional data, it falls under the category of Unsupervised Learning, and more specifically, Association. Classification and Regression are supervised techniques that require labeled data, and Reinforcement Learning involves an agent learning through trial and error with rewards, which is not applicable here.

  9. Question 9

    Intermediate

    Test Environments for AI-Based Systems · Virtual Test Environments

    An aerospace company is developing an AI-based collision avoidance system for drones operating in dense urban environments. Testing the system with real drones in a city is expensive, dangerous, and not reproducible. What is the primary benefit of using a high-fidelity virtual test environment (a simulator) for this type of system?

    Answer and explanation

    Correct answer: C

    The core value of simulation for autonomous systems is the ability to test high-risk scenarios without real-world consequences. A virtual environment allows testers to create and repeat dangerous edge cases (like near-misses or sensor failures) thousands of times to ensure robustness. This is impractical, unsafe, and prohibitively expensive to do in a physical environment. This makes simulation an indispensable tool for testing safety-critical AI systems.

  10. Question 10

    Intermediate

    ML Functional Performance Metrics · Selecting ML Functional Performance Metrics

    A bank uses an AI model to detect fraudulent credit card transactions. The cost of a missed fraud (a False Negative) is very high, as the bank has to cover the financial loss. The cost of incorrectly flagging a legitimate transaction as fraud (a False Positive) is an inconvenience to the customer but is relatively low. The testing team needs to choose the primary metric to optimize during model evaluation. Based on the business requirements, which metric from the confusion matrix should be prioritized?

    graph TD subgraph Model_Prediction Fraud Legitimate end subgraph Actual_Transaction Is_Fraud Is_Legitimate end Is_Fraud -- True Positive --> Fraud Is_Fraud -- False Negative (High Cost!) --> Legitimate Is_Legitimate -- False Positive (Low Cost) --> Fraud Is_Legitimate -- True Negative --> Legitimate

    Answer and explanation

    Correct answer: C

    The scenario explicitly states that the cost of a False Negative (missing a fraudulent transaction) is very high. Recall is calculated as TP / (TP + FN). To maximize Recall, the number of False Negatives must be minimized. Therefore, Recall is the most important metric to prioritize when the business cost of missing a positive case (in this case, fraud) is high. Precision would be prioritized if the cost of False Positives was the main concern.

Register free for 10 more questions

Or unlock all 167 CT-AI questions with explanations, timed mode and flashcards.