Software testing theory for interviews: reasoning for QA engineers and developers
Knowing the definition of regression is useful. Explaining what to test after a change to lead conversion, and at which level, is more useful. Connect theory to explicit conditions, observations and decisions you can defend in an interview.
This is a learning guide, not a leaked question bank or an employment guarantee. Interview prompts and business cases are editorial exercises; their rules do not promise the current behavior of simulator modules. Developer sections cover testing, not a complete algorithms, language or system-design curriculum. Core terminology was checked against ISTQB CTFL 4.0.1, whose authors and copyright holder are credited in the linked ISTQB document. This is independent material, not an accredited course.
- Clarify the rule and risk
- Choose data and an action
- Specify the expected result
- Explain the limits of the check
1. Testing, QA and evidence of quality
Testing provides information about a product and its risks, not a certificate that it contains no defects. QA extends beyond checking a finished feature: a team can improve requirement reviews, test data and feedback processes in advance. Debugging investigates the cause of an observed problem and fixes it. One person can perform these activities, but their purposes remain distinct.
Imagine a CRM: a manager clicks Convert lead and a customer appears. That is an observation, not yet a correctness claim. The requirement says that one lead produces at most one customer and a repeat action returns the existing customer. The test oracle, meaning the basis for judging the outcome, is not the green notification. It is the customer count and its relationship to the original lead. Your test must observe those properties.
Distinguish a human error, a defect in an artifact and a failure during execution. An incorrect assumption about retries can leave the code without duplicate protection; two customers after a retry are an observable failure. A defect report does not need to guess what the developer was thinking: a reproducible violation of the rule is enough.
Interview question
Lead conversion created a customer once. Can you say the feature has been tested?
Explore the answer and follow-up
One successful path has been checked with particular data. I record the lead and customer IDs and verify their relationship. I separately consider a repeated request, missing permissions, incomplete data and interruption. Completeness depends on agreed scope and risk: working once does not cover those conditions. My report distinguishes what was checked from the remaining limitations.
The interviewer asks: The requirement says nothing about retries. Should you immediately report a duplicate as a bug?
First record the observation and risk, then clarify the rule with the product owner. The duplicate may be serious, but I should not present my preferred outcome as an agreed requirement without a basis. Existing data constraints or documented team decisions may provide that basis.
The developer perspective
State an invariant before implementation: a condition that must remain true. Here, a lead_id is associated with at most one customer. Review, database constraints and tests can then protect one business idea instead of three unrelated technical details.
Common trap: No defects found does not mean no defects exist. A passing result is bounded by the data, environment and observations of the test.
2. Testable requirements and acceptance criteria
The statement “a task becomes overdue on time” hides at least three decisions: how the deadline is stored, which clock is authoritative and whether equality counts as overdue. Our exercise contract stores due_at in UTC, uses server time and marks a task overdue only when now > due_at and status != done. The expected result can now be calculated before execution.
For a deadline of 12:00:00 UTC, an open task is not overdue at 11:59:59, is still not overdue at 12:00:00 and is overdue at 12:00:01. A completed task is not overdue at any of those times. Changing the display timezone changes the representation, not the instant. These cover distinct risks, not three equivalent clicks.
Acceptance criteria describe required properties of a particular feature. A team-wide definition of readiness to ship may also require reviews, documentation and checks. Do not substitute one for the other: “all tests pass” does not explain the business rule for overdue tasks.
Interview question
How would you test the requirement “customer search is fast”?
Explore the answer and follow-up
I first clarify users, data volume, queries, load and where time is measured. I might propose agreeing on p95 server response time at most 500 ms at 50 requests per second over a million records, with an explicit error-rate limit. p95 means 95% of measurements do not exceed the selected value. These are example targets, not universal standards. I measure the distribution and errors, not one unusually fast request.
The interviewer asks: The average is 200 ms. Has the requirement been met?
Not necessarily: an average can hide a slow tail, and failures can be fast. We need the agreed percentile, error rate and matching load conditions. I preserve the workload, service version, duration and warm-up rules so the result can be reproduced.
The developer perspective
An injectable clock and explicit timezone make boundary tests possible without waiting for real time. Do not calculate the expected result using the same implementation being tested: an identical mistake on both sides can pass the assertion.
Common trap: A convenient 500 ms target does not become an agreed requirement just because it is easy to automate.
3. Levels, types, confirmation testing and regression
A test level concerns the boundary of the object under test: a component, cooperating components, a complete system or interactions between systems. A test type concerns a property, such as functionality or performance. Acceptance examines suitability for an agreed need. A test can belong to both a level and a type; these are not competing lists.
After fixing duplicate lead conversion, confirmation testing reproduces the original failure and checks the fix. Regression looks for unwanted effects elsewhere: ordinary conversion, manager permissions and handling an existing customer. Smoke testing is a short build-health check that helps decide whether deeper testing is worthwhile.
| Check | What it observes | What it does not prove |
|---|---|---|
| Component | A function rejects an empty lead_id | The real database prevents two customers |
| Database integration | Two requests preserve the unique relationship | The browser explains the rejection clearly |
| System through the UI | A manager completes conversion and sees the customer | Every external service failure mode |
| Acceptance | The agreed workflow serves the manager’s need | The absence of all technical defects |
Interview question
Why do integration checks matter when all unit tests pass?
Explore the answer and follow-up
A unit test may replace the database with a stub that accepts everything. The real database has constraints, types and transactions. I choose an integration check for a risk at that boundary, such as concurrent writes. Running every combination through a browser is more expensive and less diagnostic, so end-to-end tests complement rather than replace lower-level checks.
The interviewer asks: Should the same test set be repeated at every level?
Not mechanically. Many rule combinations can be checked close to the logic, boundary contracts separately, and critical user paths through the UI. Repetition should cover a new risk, such as a field lost during serialization, rather than merely increase the test count.
The developer perspective
The test pyramid is a way to discuss feedback cost, not mandatory percentages. Choose the narrowest level capable of exposing the target defect while retaining checks of the components working together.
Common trap: Integration does not simply mean slow, and automated is not a test level.
4. Equivalence classes, boundaries and combinations
Our exercise rule: quantity is required and must be an integer from 1 through 99. Strings, null and fractions are rejected without changing the cart. There is a valid class of 1–99, values below and above that range, and separate invalid-type and missing-field classes. One value cannot represent every reason for rejection.
Useful boundary values are 0, 1, 2 and 98, 99, 100. The value 50 exercises an ordinary case. null, an omitted field, the string "2" and the number 1.5 exercise other contract rules. Every rejected request also needs a check that the cart stayed unchanged. If the contract allows string-to-number conversion, the expectation for "2" is different.
Pairwise coverage selects combinations so every pair of parameter values occurs at least once. Two browsers, two roles and two languages have eight full combinations. The four rows below cover every pair but not every triple. If a known risk only affects a viewer using Firefox in Russian, add that exact triple explicitly.
| Browser | Role | Language |
|---|---|---|
| Chrome | admin | en |
| Chrome | viewer | ru |
| Firefox | admin | ru |
| Firefox | viewer | en |
Interview question
Is testing quantity = 1 and quantity = 99 enough?
Explore the answer and follow-up
Those values test inclusion of the boundaries, not rejection of neighboring invalid values. I add 0 and 100, nearby values, types and the missing field as specified. With limited time I prioritize incorrect quantities and hidden cart mutations after rejection. I explain deferred checks rather than claim complete coverage.
The interviewer asks: Why not try every number from 1 to 99?
If the same rule applies across the range, that repetition adds less information than testing distinct rejection rules. But the assumption must change if there are discount or package thresholds: these create new classes and boundaries.
The developer perspective
A parameterized test should expose the input, expectation and purpose of each case. A hundred similar data rows do not replace an assertion that a rejected request performed no write.
Common trap: Pairwise coverage does not guarantee detection of failures involving three or more parameters.
5. State transitions: an unambiguous login-lockout exercise
Define the entire contract. For an existing account, five consecutive incorrect-password attempts trigger a 15-minute lock starting at the fifth attempt, measured by server time. A successful login before locking resets the counter. While locked, any password is rejected and the timer is not extended. At now >= locked_until the lock is removed, the counter resets and the incoming request is processed normally. All devices share the account counter.
State includes not just locked or unlocked, but the count and time. Reset the account before each scenario, otherwise a previous failure changes what the next test means. A controlled test clock lets you reproduce 14:59 and 15:00 without waiting. In a manual investigation, record server timestamps and the acceptable precision.
| Starting state | Action | Result |
|---|---|---|
| 4 incorrect attempts | Correct password | Login succeeds; count is 0 |
| 4 incorrect attempts | Incorrect password | Login denied; 15-minute timer starts |
| Locked for 14:59 | Correct password | Login denied; expiry unchanged |
| Locked for exactly 15:00 | Correct password | Login succeeds; count is 0 |
| Lock has expired | Incorrect password | Login denied; new count is 1 |
Interview question
Which checks distinguish lockout testing from checking a single error message?
Explore the answer and follow-up
I check count accumulation, the transition on the fifth incorrect attempt, reset after a successful login before locking, rejection of a correct password while locked and the unlocking boundary. I then change devices to verify the shared counter. Each case records initial state, action, result and the count or deadline change. The text “incorrect password” alone cannot support those conclusions.
The interviewer asks: What changes if two incorrect attempts run simultaneously after three previous failures?
The specified rule must count both and lock the account. A sequential test cannot establish this. We need two synchronized requests and an observation of the final state: a lost counter update could leave it at four.
The developer perspective
The counter update and lock transition must remain consistent under concurrency. Separately discuss the product risk: account-wide lockout can let an outsider disrupt the owner. That is a security-design question, not a reason to silently change the exercise contract.
Common trap: “Five errors” specifies neither the error type, accumulation window nor timer start. Without those rules, the test is ambiguous.
6. HTTP and APIs: why a 200 response is not enough
An API is an agreement between a consumer and a service. Testing needs the method, path, permissions, headers, request body, response and side effects. The browser is one consumer: a working screen does not prove direct requests have the same protection.
Exercise contract: POST /enrollments accepts the course_id of an existing published course and creates one enrollment for the current user. Check response structure and ID, the user-course relationship, a repeated request, an unknown course, an unpublished course, a missing field and somebody else’s user_id in the body. Rejections also require evidence that no enrollment was created.
Idempotency concerns the same intended effect when an identical request is repeated, not necessarily identical responses. POST is not automatically idempotent. An application-level idempotency key needs a contract for scope, retention and reuse with a different body. A timeout means the client did not receive a response, not that the server performed no write.
Interview question
POST returned 200. What else should you check before declaring success?
Explore the answer and follow-up
I compare the status with the contract and verify required fields, types, values and the absence of unexpected sensitive data. I then independently read stored state: exactly one record belongs to the intended user and course. Separately I check rejection and absence of side effects. The number 200 describes the HTTP response, not fulfillment of every business condition.
The interviewer asks: Can you simply repeat the POST after a timeout?
First establish the retry contract. With idempotency-key support, repeat the same key and body; otherwise use the documented operation identifier to check the original result. A new random key may create a second operation. Concurrent duplicates also matter, not only sequential retries.
The developer perspective
Separate request syntax validation, business validation and persistence. A contract test can expose incompatible consumer-provider formats, but does not replace checks of real permissions and stored data.
Common trap: Matching status codes do not imply matching state. Conversely, a retry can return a different status while preserving the same effect.
7. Authentication, authorization and access boundaries
Authentication establishes who is requesting; authorization determines permitted actions. A hidden button is not server-side protection. Prepare two owned test manager accounts A and B in an authorized environment, customer 101 owned by A and customer 202 owned by B. Contract: managers may read and change only their own customers.
With A’s token, read 101 as a positive control. Then request 202, try an update and export if they are within the authorized scope. A rejected GET is not enough: writes may follow a different permission path. Verify a forbidden update left the record unchanged using the authorized B account. Never use real third-party data or test without permission.
401 usually concerns missing suitable credentials, while 403 is refusal to fulfill a request; a service may conceal a resource using 404. The expected status therefore comes from the contract. Essential properties are no disclosure and no forbidden effect.
Interview question
Why is a valid token insufficient for GET /customers/202?
Explore the answer and follow-up
The token identifies the subject, not their permission for this object. I check role, customer ownership and tenant boundaries where applicable. I compare allowed and forbidden requests and inspect the response body. Replacing numeric IDs with UUIDs is not authorization: knowing an identifier must not grant access.
The interviewer asks: The server returns 403 but includes the other customer’s name. Did the test pass?
No. A rejection status does not cancel a disclosure in the body. I record exposed fields, role, resource and redacted evidence. I also investigate exports and lists within the authorized scope because several routes may share the same flaw.
The developer perspective
Server-side object checks belong on every applicable action. A useful regression matrix is subject × object × action: owner, other manager, administrator; read, change, export.
Common trap: A successful login test says nothing about horizontal access between two managers with the same role.
8. Interview SQL: schema, query and result verification
This is explicitly a SQL task, not a verbal summary. PostgreSQL has customers(id PRIMARY KEY, email TEXT NULL) and orders(id PRIMARY KEY, customer_id NOT NULL REFERENCES customers(id), status TEXT NOT NULL). Find customers with no orders of any kind, then separately find duplicate nonempty email groups. In this exercise email matching is case-sensitive, spaces are not trimmed, and NULL and empty strings are excluded from duplicates. Cancelled orders still count as orders.
Test data: customers contains (1, a@example.test), (2, b@example.test), (3, a@example.test), (4, NULL), (5, empty string). orders contains (10, 1, paid), (11, 1, cancelled), (12, 2, cancelled). Expected customer IDs without orders: 3, 4, 5. Expected duplicate: a@example.test with count 2. Customer 1 deliberately has two orders to expose accidental multiplication when joining and counting customers.
SELECT c.id
FROM customers AS c
WHERE NOT EXISTS (
SELECT 1 FROM orders AS o WHERE o.customer_id = c.id
)
ORDER BY c.id;
SELECT email, COUNT(*) AS customer_count
FROM customers
WHERE email IS NOT NULL AND email <> ''
GROUP BY email
HAVING COUNT(*) > 1
ORDER BY email;
Interview question
How do you explain why these queries solve the actual problem?
Explore the answer and follow-up
NOT EXISTS retains a customer with no matching order; status is intentionally not filtered. The second query excludes missing and empty addresses, groups customers by email and retains groups larger than one. I verify exact IDs 3, 4, 5, not just three result rows. For duplicates I verify both address and count 2. Ordering is explicit, making comparisons reproducible.
The interviewer asks: Now find customers without paid orders. What changes?
Add AND o.status = 'paid' inside NOT EXISTS. The result becomes 2, 3, 4, 5: customer 2 has an order, but it is cancelled. This is a different business question. With LEFT JOIN, blindly moving a condition into WHERE can eliminate the NULL-extended rows you needed.
The developer perspective
PRIMARY KEY identifies a row, FOREIGN KEY maintains a relationship and NOT NULL disallows a missing value. These do not establish that a business query is correct. Clarify email normalization before introducing uniqueness; lower(trim(email)) changes this task’s semantics rather than merely speeding it up.
Common trap: Three returned rows could be three wrong customers. You need expected values and control cases, not just a row count.
9. Transactions, concurrency and integration failures
A transaction groups database changes into a unit of completion: the agreed set commits or unfinished changes are rolled back. It does not automatically include an external service. If the CRM saved a customer and sent an email, rolling back a local record cannot unsend that message. First define success and how partial outcomes are recovered.
ACID groups four properties: atomicity guards against partial persistence, consistency concerns defined constraints, isolation determines permitted interaction between concurrent transactions, and durability concerns committed changes surviving. State the boundaries: a database does not know every business rule, and isolation is not always equivalent to serial execution. For a transfer, specify permitted balances, related ledger entries and concurrent behavior separately.
Exercise rule: one customer per lead_id; a mail-service failure does not cancel customer creation, but leaves a pending notification that can be retried. Make the mail service time out after accepting the message. Verify one customer, a persisted notification job, bounded retries and no second customer after repeated conversion. Exactly one email needs its own provider agreement and duplicate-effect protection.
A race requires overlapping operations. If both requests read “no customer” and then both insert, a sequential test can always pass. A unique constraint helps protect the invariant, but the application must handle the conflict and return a consistent outcome. Isolation level and retry strategy depend on the database and operation.
Interview question
A transfer debited one account but crediting failed. What do you check?
Explore the answer and follow-up
I clarify architecture and contract. For two writes in one local transaction, I inject failure between steps and verify both roll back. Across services, I check the pending operation state and prescribed recovery or compensation. I preserve the operation ID and reconcile ledger entries, not just the final balance. A retry must not cause another debit, and the permitted reconciliation delay must be defined.
The interviewer asks: Does a timeout mean the external operation never happened?
No. Its response may have been lost after execution. Use the contract’s status lookup, operation identifier and safe retry mechanism. Simulate failure before acceptance separately from response loss after success because their consequences differ.
The developer perspective
Discuss transaction boundaries, conflict handling and event delivery. An outbox can persist the intent to publish alongside the business change; repeated delivery still requires consumer protection against a repeated effect. A queue alone does not justify an exactly-once promise.
Common trap: ACID in one database does not guarantee atomicity across a bank, email provider and queue.
10. Bug reports, severity, priority and release decisions
A useful report makes the discrepancy reproducible: version and environment, role, initial data, exact actions, expected result with its basis, actual result and evidence. “Checkout sometimes crashes” identifies neither the step nor the symptom. Checkout means placing an order; say, for example, that confirmation returns API status 500, the order is saved and the page offers to retry.
Severity describes impact; priority describes urgency in context. A visual defect on a campaign launch page may be urgent, while rare data corruption remains serious despite affecting few users. Scales and decision ownership are team agreements, not a universal internet ranking.
For time-limited regression, identify the changed path, related data, permissions and irreversible effects. After changing CRM conversion, duplicates, lost lead relationships and unauthorized access may come first. Report completed checks, known defects, environment limitations and untested risks. Release decisions need that evidence, not just a green-test percentage.
Interview question
A failure happens once in ten attempts. How do you make the report useful?
Explore the answer and follow-up
I describe the ten attempts and differing conditions, the failing request time, build, request ID, response and order state afterwards. I attach redacted logs or a trace if available. I separate fact from hypothesis: another concurrent request might be involved, but that is not proven. I identify the duplicate-order risk and the exact reproduction procedure.
The interviewer asks: The developer says “works for me.” What next?
Compare versions, configuration, roles, data and action order. Use the request ID to locate the specific attempt and compare component observations. Do not replace evidence with an argument; if the conditions are not yet understood, keep the investigation status and observations explicit.
The developer perspective
A correlation identifier connects records of one operation across components. Logging should support investigation without recording passwords, tokens or unnecessary personal data.
Common trap: Reproduction frequency is not impact. A rare duplicate debit does not become a cosmetic issue.
11. Developer interview theory: doubles, coverage, TDD and CI
An automated test needs controlled setup, an action and an assertion about observable behavior. For the CRM, create a fresh lead and its owner, convert it and assert one customer with the correct relationship. Isolate the data: the test must not rely on a neighboring test running first. Cleanup or an isolated namespace is more reliable than a random name that only reduces collision probability.
A test double replaces a real participant. A stub provides prepared responses, a mock typically also verifies expected interactions, and a fake provides a simplified working implementation. Libraries vary in terminology. Explain what is replaced and which failures become invisible. A mail stub helps trigger a timeout but cannot establish compatibility with the real provider.
TDD uses a short cycle: a test expresses required behavior and fails for the expected reason, minimal implementation makes it pass, then the code is improved while preserving checks. It cannot guarantee good requirements. CI, continuous integration, runs checks automatically for changes; reproducible environments and useful failure artifacts make the feedback trustworthy.
Interview question
Code coverage is 100%. Why can defects remain?
Explore the answer and follow-up
Coverage measures executed code elements, not complete requirements or strong assertions. A test can execute a branch without checking anything useful. An omitted requirement may not exist in the code at all. I inspect value assertions, negative cases, interactions and concurrency. Mentally changing a rule, or using mutation testing, helps ask whether a meaningful behavior change would be detected.
The interviewer asks: A test occasionally fails in CI and passes on retry. Add a sleep?
First investigate shared data, execution order, asynchronous conditions, clocks and network dependencies. A fixed pause can hide a race and slow the entire suite. UI tests should wait for a specific observable condition with a bounded timeout. Passing on retry is diagnostic evidence, not proof the instability is fixed.
The developer perspective
Review whether the test would fail for the real defect, explain its failure and work both alone and in parallel. Do not assert only internal calls when the user-visible result matters. Keep CI traces and failure reports, not secrets.
Common trap: A mock interview is a practice conversation. A mock in test automation is a test double: the same English word has different contexts.
12. Performance, recovery and mid-level reasoning
Load testing evaluates a stated workload; stress testing explores limits and failure; a sustained run can expose resource accumulation. Users and requests per second are not interchangeable: one user may wait while another issues many requests. Specify operations, intensity, data, duration and environment.
For an exercise LMS, assume 100 active learners reading lessons, each making a request every ten seconds on average, plus 20 learners submitting answers simultaneously. Replacing this with a hundred identical GET requests does not prove readiness. Alongside latency, check errors, saved attempts, processing queues and recovery when load decreases.
RPO is the acceptable data-loss window; RTO is the target restoration time. If an incident happens at 10:00 and the usable backup is from 09:50, the missing window is ten minutes. Restoration at 10:40 takes forty minutes. Compare these with agreed objectives. A backup file existing does not prove consistent data can actually be restored.
Interview question
How would you answer “when can testing stop” without saying “when there are no bugs”?
Explore the answer and follow-up
I describe agreed exit criteria: critical scenarios checked, acceptable defect status, addressed risks and known limitations. Then I present actual results and residual risks. Scope or environment changes may require revisiting the conclusion. No discovered defects does not mean no defects exist; shipping is a decision made with known uncertainty.
The interviewer asks: Release is in 30 minutes and full regression takes a day. What do you do?
Clarify changes and failure impact, prioritize a short set of critical-risk checks and state what remains untested. An irreversible migration or financial operation may justify postponement. Rollback planning, post-release monitoring and limited rollout help manage risk, but do not turn untested behavior into verified behavior.
The developer perspective
For schema changes, check compatibility of old and new application versions, migration order and recoverability. Reverting code does not necessarily reverse a data transformation. Demonstrating recovery on an isolated copy is more useful than saying recovery is planned.
Common trap: A strong mid-level answer is not necessarily faster. It makes assumptions, priorities and evidence limitations explicit.
Study without memorizing scripts
- First pass: explain one topic in your own words and list unfamiliar terms. Do not race through twenty definitions in a row.
- Practice: turn an example into setup → action → expectation → evidence. For SQL, execute the queries in your own learning database using the specified data.
- Self-check: answer the main and follow-up questions before expanding the explanation. Compare the rule, example, observations and limitations, not exact wording.
- Transfer: replace the shop with a CRM or LMS. Explain what stays the same in your method and which business rules change. This tests understanding better than reciting text.
- Return: at your next session, investigate the same risk with different data and no hint. Explaining why an alternative answer is wrong is a useful sign of stronger understanding.
Review four things: a clear rule, concrete data, an observable result and honest limits. This is a self-check framework, not an automatic qualification score. Article explanations are available without registration; this page does not separately save answers.
Interview preparation FAQ
What should a junior QA engineer study first?
Start with requirements, expected results, test design, testing levels and defect reports. Connect them to HTTP, permissions and simple SQL. Give every term a worked example of your own. Adjust priorities to the job description, not the length of a generic checklist.
How does a mid-level answer differ from a junior answer?
Titles vary between companies. A useful distinction is that a junior explains a check under clear conditions, while a mid-level engineer also discusses dependencies, risk, test cost, concurrency and result limitations. This is not an official scale or a demand to know every technology.
Do developers need testing theory?
Yes. It helps select test boundaries, define expected behavior independently of implementation and recognize what a stub cannot prove. Developer interviews can benefit from explanations of transaction boundaries, test-data isolation, flaky tests and CI feedback. This guide does not replace preparation for a specific programming language.
Should I memorize answers and obtain a certificate?
Memorized wording breaks down under follow-up questions. Reconstruct the reasoning with different data instead. A certificate may be required by a particular vacancy but does not replace practice; you do not need to buy one to use this guide.
Theory becomes a working tool when it helps choose the next check and explain the result. Build one investigation of your own with data and evidence, then move to the next risk.