QA / Software Testing

Test Case Design Techniques: How QA Teams Build Better Test Coverage with AI

JIN

Sep 30, 2026

Table of contents

Table of contents

    AI can generate thousands of test cases in minutes, but volume has never been the real challenge in testing. The difficulty lies in designing cases that actually find defects and provide actionable feedback.

    A 2026 study of six frontier coding agents on SWE-bench Verified makes the point well. Claude Opus 4.5 wrote tests in 84.4% of the tasks it solved, GPT-5.2 wrote tests in just 0.8%, and the two resolved almost the same share of issues: 74.4% and 71.8%. When the agents did write tests, lines that printed values outnumbered real assertions by about five to one. The tests helped the agent look around the code. They did very little checking.

    When Meta evaluated TestGen-LLM, its LLM-based unit test generator, on Instagram’s Reels and Stories products, 75% of the generated test cases built correctly, 57% passed reliably, and only 25% increased coverage. Meta still put the tool to work, and engineers accepted 73% of the improvements it recommended. It worked because every generated test had to clear automated filters before anyone looked at it. Generating tests was cheap. Proving each one added something was the real gate.

    The classic test case design techniques (equivalence partitioning, boundary value analysis, decision tables, state transition testing, pairwise testing) still decide whether a suite finds defects. What AI changes is who does the listing. Current models are fast at producing values, steps, and combinations, and they now write assertions about as strong as a developer’s. What they cannot do on their own is know what the requirement says the result should be, and they rarely tell you what they skipped.

    This guide explains how each test design technique fits into an AI-assisted workflow, why headline coverage metrics often overstate what AI-generated suites actually detect, and how to report coverage in a way that stands up to scrutiny in a release review.

    What Are Test Case Design Techniques?

    Test case design techniques are structured approaches for selecting which inputs, conditions, and sequences to test. The main techniques, like equivalence partitioning, boundary value analysis, decision table testing, state transition testing, and pairwise testing, remain central. In an AI-assisted workflow, QA engineers use these techniques to define the model’s structure (such as partitions, rules, states, or parameters) and set the standard for evaluating output. People define the expected results; AI generates the cases. Teams then measure coverage using requirement tracing and mutation testing, not just line coverage. Since exhaustive testing is impossible, these techniques provide a defensible rationale for which values are tested and which are not.

    The ISTQB Certified Tester Foundation Level v4.0 syllabus groups techniques into four families: black-box, white-box, experience-based, and collaboration-based. The international standard ISO/IEC/IEEE 29119-4:2021 defines the techniques more formally and adds combinatorial methods such as pairwise and each-choice testing.

    The more practical question is how responsibilities should be divided between people and AI within each technique family.

    Technique Examples Test case Where AI can help Human own
    Black-box (specification-based) Equivalence partitioning, boundary value analysis, decision tables, state transition Requirements, business rules, UI specs Listing values, pulling conditions out of prose, drafting state models Where the partitions are and what the expected result is
    White-box (structure-based) Statement and branch testing Source code Finding unreached branches and writing unit tests to reach them Whether each assertion checks intended behavior
    Experience-based Error guessing, exploratory testing, checklist-based testing Tester knowledge, defect history Drafting charters and checklists from past defects Judgment during the session
    Collaboration-based User stories, acceptance criteria, ATDD Conversations between business, development, and test Proposing acceptance criteria and flagging vague wording Agreement on what “done” means
    Combinatorial Pairwise, each choice, t-way testing Parameters and their values Proposing parameters, values, and constraints Confirming constraints; the math belongs to a covering-array tool

    Why Does AI-Generated Test Coverage Often Look Better Than It Is?

    A suite built mostly by AI can post strong numbers and still miss the defects that matter. Research from Meta, Google, and several university teams points to three reasons. Each one maps back to a gap that a test design technique is meant to close.

    The Expected Result Comes from the Code, Not the Requirement

    Every test has two halves: the steps and the check. Early studies with GPT-3.5 found that AI-written assertions often copied whatever the code did. Newer models have closed much of that gap. In an ASE 2025 study of 13,866 test oracles from 135 Java projects created after the models’ training cutoff, LLM-generated oracles achieved a mutation score of 43%, close to the 45% achieved by developer-written oracles. A 2026 follow-up using GPT-4o-mini, Llama 3.3, and DeepSeek V3 found that simple prompting now outperforms four specialized test-generation tools in line coverage, branch coverage, and mutation score.

    So the assertions are strong. The open question is what they are checked against. The ASE 2025 authors found that the test setup and the methods the code calls gave the model enough to write good oracles, and that extra context added little. The model works out “correct” by reading the code. Mutation score cannot flag this, because mutants are small changes to the current code, and an assertion that locks in today’s behavior (bug included) still kills them. The 2026 study adds a practical warning: only about 56% to 59% of generated tests passed on average, and the authors named incorrect oracles as the main reason the rest failed. We looked at the code-generation side of this problem in Does AI-Generated Code Reduce the Need for Testing, or Demand More?. The test design fix is to give the model a second source of truth. When the requirement is written down as partitions, rules, or transitions with expected results, a gap between what the code does and what the spec says shows up as a failing test instead of hiding inside a passing one.

    Line Coverage Counts What Ran, Not What Was Checked

    Line coverage tells you a statement was executed. It says nothing about whether a test would fail if that statement were wrong. Google has run mutation testing inside code review for years, across more than 24,000 developers and over 1,000 projects. In a study of almost 15 million mutants, Google found that developers who saw surviving mutants wrote more tests over time, and that mutants were coupled to real high-priority faults. Meta’s numbers make the gap concrete. Of the 571 tests its ACH system generated, 277 would have been thrown away if line coverage had been the only acceptance criterion. Almost half of the tests that caught simulated faults added no new lines.

    Output Quality Follows Specification Quality

    Researchers from Tokyo City University and VeriServe, a Japanese testing company, tested a prompt-only method in which the model first selects test design techniques that suit a requirements document and then writes high-level test cases for each technique. For detailed Bluetooth profile specifications, the generated cases achieved a macro recall of 0.84 on the human-written set. On shorter Mozilla feature descriptions, recall fell to 0.37. The speed gain was real in both cases: about 20 minutes to produce roughly 100 test cases, whereas the paper puts manual work at around 120 minutes for 30.

    Taken together, these findings highlight a clear pattern: AI does not eliminate the need for thoughtful test design. Instead, it shifts design work earlier, clarifying specifications and expected results, and later, through filters and adequacy checks. Manual effort to write each case decreases, but the need for human judgment at both ends remains. Teams that focus only on automating the middle generate more tests, while those that reinforce both the design and review stages achieve stronger coverage.

    How Should Each Test Design Technique Change When AI Is in the Loop?

    1. Equivalence Partitioning and Boundary Value Analysis: People Draw the Lines, AI Fills Them In

    Take a checkout rule: orders from $50.00 to $500.00 ship free, orders below $50.00 pay a flat fee, and orders above $500.00 need a freight quote. A QA engineer should define the partitions (invalid negative totals, below threshold, free shipping band, above ceiling, non-numeric input, wrong decimal precision) and the boundaries, including 3-value checks such as $49.99, $50.00, and $50.01.

    Once that structure exists, AI is very good at the rest: realistic data for each partition, currency and locale variants, and formatted test cases in your team’s template. The expected result for each partition should come from the business rule and be written down before generation. If a model decides where the partitions sit, it will usually find the obvious ones and quietly miss the rule that appears only in a footnote in the pricing policy.

    2. Decision Table Testing: Let AI Extract the Conditions, Then Check the Collapsed Rules

    Business rules written in prose are where AI saves the most time. A model can read an insurance eligibility policy or a promotion rule and turn it into a first draft of conditions, actions, and rule columns in minutes. The risk sits in the next step, when rules are collapsed using “don’t care” entries to shrink the table. A seemingly harmless merge can combine two rules whose outcomes actually differ. For pricing, eligibility, and compliance logic, have a reviewer check every collapsed column against the source text, or keep the full table and accept a few extra tests.

    3. State Transition Testing: Use AI to Draft the Model, Not to Approve It

    Give a model the user stories for an order workflow, and it will produce a reasonable state diagram: created, paid, packed, shipped, delivered, returned. What it tends to underweight are invalid transitions, such as canceling after shipment, refunding twice, or paying for an order that has already expired. Those are often where the costly defects live. Ask for an explicit table of invalid transitions, aim for valid transitions coverage first, and then add invalid ones in order of business risk. Treat the model as a fast drafter, and still have someone who understands the domain sign off on the state model.

    4. Pairwise and Combinatorial Testing: Let the Algorithm Do the Math

    NIST’s research on real-world failures found that most are triggered by a single parameter or by the interaction of two parameters. In one NASA application, 67% of failures were due to a single parameter value, 93% to two-way combinations, and 98% to three-way combinations; none of the systems studied had failures involving more than six parameters.

    Consider a checkout page with 4 browsers, 3 operating systems, 3 languages (English, Japanese, Vietnamese), 2 payment methods, and 2 user types. That is 144 full combinations, and a standard pairwise generator covers every pair of values in about 12 tests. Do not ask a language model to build that covering array itself, because it can drop a pair without any sign that it did. Use a proven generator such as NIST ACTS or Microsoft PICT, and use AI to propose parameters, values, and constraints (for example, Safari does not run on Windows) that a tester then confirms. SHIFT ASIA’s earlier post on all-pairs testing walks through a worked example.

    5. Mutation-Guided Generation: Point AI at the Faults Your Suite Misses

    Meta’s ACH system reverses the usual order. Instead of asking a model to write tests and hoping they catch something, it first creates realistic faults tied to a specific concern, then asks the model to write tests that catch them. Applied to 10,795 Android Kotlin classes across seven platforms, ACH generated 9,095 mutants and 571 privacy-hardening tests, and engineers accepted 73% of those tests. Each accepted test provides proof that it detects something the suite used to miss. Teams without Meta’s internal tooling can apply the same idea with open-source mutation tools such as PIT for Java or Stryker for JavaScript and .NET, pointing AI at surviving mutants in high-risk modules.

    6. Exploratory Testing and Error Guessing: AI Writes the Charter, People Run the Session

    Defect logs and incident reports are full of patterns that a model can quickly surface: time zone bugs at midnight, duplicate submissions on slow networks, rounding errors in currency conversion. Use AI to turn that history into exploratory charters and error-guessing checklists for each release. The session itself still needs a person, because exploratory testing depends on noticing that something feels wrong before anyone can say why. Whatever the session finds should flow back into the structured techniques as a new partition, rule, or transition.

    Across all six techniques, the division of labor is consistent: people define the structure and expected results, while AI handles the enumeration. When proven deterministic tools exist, such as covering-array generators or mutation engines, use them for checking. Relying on a language model for tasks that established algorithms can guarantee is a common pitfall in AI-assisted test design.

    What Does an AI-Assisted Test Design Workflow Look Like in Practice?

    The workflow below is our recommended approach for teams transitioning from manual to AI-assisted test design. It preserves the speed benefits of automation while adding the checks research and industry experience show are essential.

    1. Tighten the test basis: Before generating anything, have AI flag vague words, missing limits, and conflicting rules in the requirements, then resolve them with the product owner. The Masuda results show why this step pays for itself.
    2. Choose techniques and write the structure: For each requirement, pick the technique that fits (ranges suit boundary value analysis, business rules suit decision tables, workflows suit state transition) and record partitions, rules, states, or parameters along with the expected results.
    3. Generate within that structure: Give the model the structure, the technique name, your test case format, and an instruction to list anything it could not cover. A template is shown below.
    4. Filter automatically: Keep only tests that build, pass consistently across repeated runs, and either cover a requirement item not yet covered or kill a surviving mutant. This mirrors the filter approach Meta used for TestGen-LLM.
    5. Review by exception and trace every test: Human review should focus on expected results and anything linked to a high-risk requirement. Every test should trace back to a requirement ID and a technique so coverage can be reported by requirement, not by file.

    A generation prompt built this way looks less like a question and more like a work order:

    Technique: Boundary value analysis (3-value) + equivalence partitioning
    Requirement: CHK-114 Free shipping for order totals from 50.00 to 500.00 USD
    Partitions (with expected results):
    P1 total < 0 -> validation error "Invalid amount"
    P2 0.00 to 49.99 -> flat fee 7.50
    P3 50.00 to 500.00 -> free shipping
    P4 > 500.00 -> freight quote required
    P5 non-numeric -> validation error "Invalid amount"
    Boundaries: 49.99 / 50.00 / 50.01 and 499.99 / 500.00 / 500.01
    Output: one test case per row in our table format (ID, precondition, steps, data, expected result)
    Rules: Use only the expected results listed above. List any partition or boundary you could not cover and why.

    This workflow forms the foundation of the AI-driven testing framework we are developing at SHIFT ASIA. The framework is built to take a feature from test case design through to execution in a single, integrated process: requirements are defined, engineers specify the techniques, partitions, and expected results, and AI generates and automates the test cases within that structure. Results are always traced back to the original requirements. The division of responsibilities remains as described above: our QA engineers make the key test design decisions, while AI accelerates the work between design and execution. This approach increases speed without compromising on the definition of ‘correct.’

    How Do You Measure Test Coverage That Holds Up in a Release Review?

    No single percentage can determine whether a test suite is effective. Each metric provides a different perspective and comes with its own limitations.

    Metric What it tells you What it misses Best use for
    Requirement coverage Which requirements have at least one linked test Whether those tests are any good Release readiness and audit trails
    Partition and boundary coverage Share of defined partitions and boundaries exercised Partitions nobody identified Input-heavy features and forms
    Decision rule and transition coverage Share of rules or transitions exercised Rules missing from the table Business logic and workflows
    Pairwise (2-way) coverage Share of value pairs covered Faults needing three or more parameters Configuration and compatibility testing
    Branch coverage Share of code branches executed Whether assertions would catch a wrong result Finding untested code paths
    Mutation score Share of seeded faults the suite detects Whether assertions match the requirement, and faults unlike the mutation operators used Judging the strength of the suite, especially AI-generated tests

    For most enterprise teams, a practical approach to reporting includes requirement coverage for the entire release, technique-level coverage for the highest-risk features, and mutation score on recently changed code. Limiting mutation analysis to code changes, as practiced by Google, helps manage compute costs while still identifying weak tests where they have the greatest impact.

    What Are the Challenges and Governance Needs of AI-Assisted Test Design?

    Challenges

    The first challenge is confidence: AI-generated test cases can appear complete and professional even when the expected results are incorrect, leading reviewers to trust the output more than they should.

    The second is volume: generating hundreds of test cases quickly can double maintenance and increase flaky tests, so suites should only grow with tests that pass established filters.

    The third is data exposure: requirement documents, source code, and defect logs often contain sensitive information, and sending them to external models requires the same scrutiny as any other data transfer.

    The fourth challenge is skill development: if junior testers never learn to define partitions by hand, they may struggle to review AI output, weakening the critical control that catches incorrect expected results.

    Governance

    Effective governance in AI-assisted test design centers on traceability and ownership. Each test should be linked to a requirement and a specific technique, using terminology from ISTQB and ISO/IEC/IEEE 29119-4, so that auditors and new team members can understand its purpose. Acceptance criteria for AI-generated tests should be clearly documented: the test must build cleanly, remain stable across repeated runs, and either add requirement coverage or detect a surviving mutant. For high-risk areas like payments, eligibility, and personal data, expected results should have a named human owner. Teams should also maintain a concise model and data policy specifying which models are approved, what data they can access, and how prompts and outputs are stored for future review.

    More Test Cases Are Easy to Get. Coverage You Can Defend Takes Test Design.

    Engineered by humans. Accelerated by AI. Japanese quality methodology, delivered by QA engineers in Vietnam.

    For most teams, generating test cases is no longer the main challenge. The real issue is coverage; specifically, ensuring that increased generation does not obscure gaps in what is actually tested. At SHIFT ASIA, we apply the proven test design methodology of SHIFT Inc., one of Japan’s largest software quality companies, through our ISTQB-certified QA engineers in Vietnam. This guide outlines the techniques we use to design tests for functional, regression, and automation projects, always tying expected results to your requirements rather than to the code’s current behavior. We are also embedding this methodology into our AI-driven testing framework, which connects test case design to execution in a single flow. If your test suite keeps growing but you are unsure whether it is catching more defects, our QA team can help assess your coverage and show how AI can enhance your testing without lowering standards.

    Talk to Our QA Team


    Frequently Asked Questions

     

    The main test case design techniques are equivalence partitioning, boundary value analysis, decision table testing, state transition testing, pairwise (combinatorial) testing, statement and branch testing, error guessing, exploratory testing, and acceptance test-driven development. ISTQB groups them into black-box, white-box, experience-based, and collaboration-based families, and ISO/IEC/IEEE 29119-4 formally defines them.

    No. AI can take over most of the typing: listing values, formatting test cases, and drafting decision tables or state models. Current models write assertions nearly as strong as a developer's, but research shows they work out the expected result mainly from the code itself. If the code is wrong, the test can confirm the wrong behavior. QA engineers still need to define the partitions, rules, and expected results from the requirement and review what the model produces.

    AI improves test coverage when it works inside a defined structure. Given partitions, rules, or states, it can fill in cases quickly and propose edge cases people forget. It improves coverage most reliably when paired with mutation testing, where it writes tests aimed at faults the current suite misses, as Meta did with its ACH system.

    Code coverage measures which lines or branches ran during testing. Mutation score measures how many small, deliberately introduced faults the tests actually detect. A suite can reach high code coverage with weak assertions, but it cannot reach a high mutation score that way, which makes mutation score a better check on the strength of AI-generated tests. Mutation score does not tell you whether an assertion matches the requirement, so it works best alongside expected results written from the spec.

    It depends on the number of parameters and values, but the savings grow quickly. A checkout page with 4 browsers, 3 operating systems, 3 languages, 2 payment methods, and 2 user types has 144 full combinations, while a pairwise set that covers every pair of values needs about 12 tests. NIST research found that most real-world failures involve one or two parameters, which is why pairwise testing catches a large share of defects with a small suite.

    It is better not to. A covering-array tool such as NIST ACTS or Microsoft PICT guarantees that every pair is covered, while a language model can silently leave pairs out. Use AI to suggest the parameters, values, and constraints, confirm them, and let the tool build the set.

    SHIFT ASIA is developing its own AI-driven testing framework that covers the path from test case design to test execution. Our QA engineers define the test design techniques, partitions, rules, and expected results based on SHIFT Inc.'s quality methodology, and AI generates, automates, and runs the test cases within that structure, with results traced back to the original requirements.

    Share this article

    ContactContact

    Stay in touch with Us

    What our Clients are saying

    • We asked Shift Asia for a skillful Ruby resource to work with our team in a big and long-term project in Fintech. And we're happy with provided resource on technical skill, performance, communication, and attitude. Beside that, the customer service is also a good point that should be mentioned.

      FPT Software

    • Quick turnaround, SHIFT ASIA supplied us with the resources and solutions needed to develop a feature for a file management functionality. Also, great partnership as they accommodated our requirements on the testing as well to make sure we have zero defect before launching it.

      Jienie Lab ASIA

    • Their comprehensive test cases and efficient system updates impressed us the most. Security concerns were solved, system update and quality assurance service improved the platform and its performance.

      XENON HOLDINGS