COOKED
The modal QA analyst writes test cases from requirements, builds and maintains Selenium/Cypress/Playwright scripts, runs regression suites, triages failures, and files structured bug reports — all text-and-screen work that coding models now do at usable quality and at volume. What resists is the thin senior tier: exploratory testing that finds the bug nobody specified, owning release go/no-go, designing test strategy across flaky distributed systems, and performance/security testing on real devices and hardware. There is no license, no signature requirement, and buyers do not pay for a relationship with a tester, so nothing outside the work itself slows displacement.
Mixed — a routine tier and a judgment tier. Writing test cases from a Jira ticket, generating Playwright selectors, and triaging a red CI build are exactly the text-in/text-out loops LLMs handle now, which caps this at 9 rather than mid-teens; the 9 instead of 4 comes from work that still needs a human in the loop — reproducing a heisenbug that only appears on a specific Android build, deciding which of 300 failing assertions are real versus a selector drift after a UI refactor, and exploratory sessions where the oracle is your own judgment about what the product should do, not a written requirement.
Fully desk- and screen-based. A 4 reflects that the job is a laptop, an IDE, and a browser for most sprints, with the physical component limited to the device lab — plugging a phone into a USB hub for real-device testing, checking a kiosk or POS terminal build, or verifying a hardware peripheral integration — none of which happens in an uncontrolled environment or on most teams' sprints.
No licence, no signature requirement. There is no state license, no ISTQB requirement in any statute, and no signature on a release; when a defect reaches production the postmortem lands on engineering management and the release owner, and QA's sign-off is an internal Jira transition with no legal weight, which is why this sits at 2 and not 5.
Anonymous artifact production. Bugs are consumed as tickets by developers who mostly never talk to the filer, and no customer selects a vendor because of who tests it — the 5 rather than 0 is the internal capital a long-tenured tester builds from knowing which subsystems historically break and which PM will fight a severity rating, which is real but transfers to nobody outside the team.
Meaningful discretion. Most days are bounded by an acceptance criteria list and a defined severity/priority matrix, which holds this under the discretion band; the 10 comes from calls that genuinely are yours — assigning severity on an ambiguous edge case, deciding a flaky test is masking a race condition rather than muting it, and recommending go/no-go into a release meeting where you can be overruled and usually are.
Has AI actually changed your work?