Skip to main content

Findings

Findings are confirmed bugs, failures, or vulnerabilities that Qodex records with evidence. A failed run should not always become a bug report. Qodex first decides whether the failure is a real product issue, a stale test, or an environment problem.

What happens when a test fails

When a scenario fails or a security probe lands, Qodex does more than show a red X. It analyzes the failure, classifies it, and writes a finding only when the failure looks like a real issue. A finding includes severity, reproduction steps, evidence, and the affected endpoint or page. Qodex deduplicates findings against the existing set so the same bug does not pile up across nightly runs.

Severity model

Five levels. Severity is the worst realistic outcome, capped by who can reach it: A person can change a finding’s severity from the Findings page. A severity change needs a reason, and the row is marked as set by a person.

Failure classification

Every failed scenario in a run is classified from its evidence: the error, the step results, screenshots, and the HTTP response. There are three outcomes: On the test run page these read as Functional regression or Vulnerability present (a real bug), Still present (a real bug that was already open), Test needs repair, Environment, or Not yet judged while classification has not run. When the evidence cannot tell which it is, Qodex says so instead of guessing. This classifier keeps regression suites usable. Without it, flaky selectors and temporary outages would look like product bugs.

Deduplication

Before recording a new finding, Qodex compares it with the findings the project already has. If a matching open finding exists, the new occurrence is recorded as a re-observation on the existing finding rather than a duplicate. The Findings page shows First seen, Last seen, and a count such as ×3 when an issue was seen more than once. A code scan finding you marked False positive or Won’t fix is not filed again on the next scan.

Evidence and basis

Qodex will not file a high or critical finding from exploration without captured proof of the failure. For access-control claims (“user X should not have been able to do that”), the finding must rest on a rule you wrote, an endpoint note, or something you said. Otherwise Qodex asks you first, and the candidate waits as Needs review until someone chooses File as a bug or Dismiss.

Status lifecycle

Findings carry one of five statuses:
Status moves are recorded with the person who made them, and the Findings page shows who reviewed each finding. Triage happens on the Findings page, in chat, or from an MCP client.

What evidence includes

Every finding ships with:
  • The exact HTTP request that triggered the failure, redacted
  • The response that proves the vulnerability or bug
  • A screenshot (UI) or response snippet (API)
  • Reproduction steps a human can follow without the agent
  • For security findings, the OWASP category (for example, A01:2021 Broken Access Control)
  • For code scan findings, the file and line in the repository

When to use it

  • Promote any real bug to a tracked finding for the team
  • File a finding when a security scenario fails, since pass means blocked and fail means vulnerable
  • Triage findings through the Findings page, in chat, or from an MCP client

When not to use it

  • Stale tests. Those are scenario maintenance, not bugs. Use Fix in chat on the test run instead.
  • Environment issues. Surface those to the team that owns the environment.

On the roadmap

Planned: flaky-test detection and auto-repair proposals for stale tests.
Planned: Jira and Linear ticket creation from findings, and SARIF export.

Findings reference

The deeper reference and data model.

Failure classification

Real bug vs stale test vs environment issue in depth.

Triage workflow

Filters, review, and the status lifecycle.

Security testing

Where the inverted-semantics finding rule lives.