Job Overview:
You’ll own quality across Invictus end to end — the platform itself (web apps, APIs, and data and document pipelines) and the agentic AI layer that runs on top of it (multi-agent workflows, retrieval, and copilot features). This is a hands-on ownership role for a QA engineer who is equally at home automating regression suites for a SaaS product and bringing structure and rigor to testing non-deterministic AI systems — treating evaluation as an engineering discipline rather than an afterthought.
Requirements:
-
3–5 years of QA / test engineering experience, with real ownership of quality for systems running in production.
-
Strong test automation across the stack — API-level testing of backend services (Python/pytest, Postman/REST) and end-to-end UI automation (Playwright, Cypress, or similar), with the discipline of maintainable, reviewable test code.
-
Solid QA fundamentals — test strategy and planning, regression and integration testing, defect triage, and release sign-off for a live SaaS product.
-
Demonstrable experience testing LLM-powered or agentic systems — designing eval datasets, golden sets, and regression suites for non-deterministic outputs. This is what sets this role apart from a standard QA position.
-
A systematic approach to LLM evaluation — LLM-as-judge, rubric design, human-in-the-loop review, and structured iteration backed by metrics rather than ad-hoc spot checks.
-
Working knowledge of agentic patterns — tool use, multi-agent orchestration, and RAG — enough to design meaningful test scenarios: agent trajectory checks, tool-call correctness, groundedness, and hallucination detection.
-
Hands-on experience wiring quality gates into CI/CD (GitHub Actions) and running tests in Docker-based environments.
-
Strong debugging instincts in complex, non-deterministic systems — comfortable following a failure end-to-end through logs and traces before guessing.
-
Fluency with modern AI-assisted development tools — Claude Code, Cursor, or similar.
Nice to Have:
-
Hands-on experience with eval frameworks and tooling — RAGAS, DeepEval, promptfoo, or the eval features of LangSmith / Langfuse.
-
RAG retrieval evaluation metrics (recall@k, MRR, context precision/recall) and how they hold up against real-world data.
-
Adversarial / red-team testing — prompt injection, jailbreaks, PII leakage, and guardrail verification.
-
Performance and load testing — both classic API/UI load (Locust, k6, JMeter) and LLM-specific dimensions like latency and token cost under concurrency.
-
Testing data pipelines or document-processing workflows — extraction, chunking, and indexing quality.
-
Experience with Azure or similar cloud platforms, and observability tooling (Application Insights, OpenTelemetry, Langfuse).
-
Exposure to financial services, wealth management, or other regulated domains.