A request form displays “Saved.” Reload the page and find the record. Then try reading it as a user who should not have access. Those checks examine different promises made by the same screen.
For a request application, completing the task may mean creating a record, assigning an owner, recording a decision, and showing that decision to the requester. A confirmation message alone is insufficient evidence of the entire flow.
Cover the rules around the successful case
Use a small set of representative cases:
| Case | Result to inspect |
|---|---|
| Valid request | Correct stored fields and visible confirmation |
| Missing required information | Clear validation and no unintended record |
| Duplicate submission | Behavior consistent with the documented duplicate policy |
| Unauthorized user | Refused operation and no disclosure or change |
| Unavailable integration | Recognizable failure and a defined recovery path |
| Reload or later visit | Persistent state that agrees with the saved result |
Choose cases from the actual business rules. Adding tests that repeat the implementation's assumptions will not reveal a misunderstood requirement.
Check boundaries directly
Exercise permissions through the backend operation as well as the screen. Check invalid inputs and limits at the point where data is accepted. Test integrations with known responses and failures so the application does not mistake an upstream error for an empty result.
If an action can be retried, establish what happens after a timeout. Use a controlled test to determine whether the original action completed and whether a repeated request changes the result. The user interface should give the operator enough information to avoid guessing.
For AI output, keep examples with expected properties. A summary should preserve the source facts. A suggested resume bullet should not invent an employer or credential. Check the result and the workflow around human review, rather than treating fluent text as a passing test.
Make a failing case reproducible
For an approval workflow, write the initial state, acting identity, input and expected result. For example: a pending request belongs to Team A; a manager from Team B calls the approval operation; the operation refuses the change, the request remains pending, and the response reveals no private request details.
If the test fails, keep the observed response and stored state with the case. “Permissions seem wrong” is difficult to reproduce. A specific user role, record and operation gives the developer a direct path to investigate. Use synthetic fixtures and remove sensitive values from shared reports.
Separate implementation tests from acceptance review. A function test can check a validation rule. An integration test can show that a saved record is returned through the API. A person completing the workflow can identify a confusing decision note or a missing step. Each contributes evidence the others may not provide.
After a fix, rerun the failed case and the nearby behavior it could affect. A change to team authorization deserves checks of permitted and refused access. A text-label correction usually needs a much smaller review. Match the verification effort to the behavior that changed.
Test the conditions people will use
Check the supported devices, input methods, and relevant accessibility behavior. Evaluate performance with representative data and concurrency where those affect the task. Record the conditions so the result can be interpreted correctly.
OWASP ASVS can help organize security verification alongside functional testing. Use it with an appropriate review scope; passing a few selected checks is not a blanket security claim.
When an agent builds on OBTO, request read-back evidence from the actual destination and open the application yourself. Keep local demonstrations and live-system checks clearly labeled.
The release record should name the version checked, cases exercised, results, and remaining limitations. Assign an owner to each unresolved issue before people depend on the application.