Job ID: 547120
ROLE DESCRIPTION AND RESPONSIBILITIES
- Stay at the state of the art on software testing practices with a focus on AI/ML/LLM evaluation (offline metrics, human in the loop reviews, A/B, red teaming).
- Anticipate AI specific risks (hallucinations, prompt/guardrail bypass, data leakage, drift) and propose measu rable quality criteria and release gates.
- Review specifications end to end:
- Validate functional specs for completeness and testability.
- Validate data, ingestion pipelines, retrieval (RAG), prompts/guardrails, and model specifications (intended use, limits, metrics, licensing) when needed.
- Define the test strategy of the global AI-powered solution — what/how/when/depth for:
- Functional & non functional tests (latency, throughput, cost),
- Safety/Responsible AI (RAI) (toxicity, bias/fairness, privacy),
- Security (jailbreak resistance, Personally Identifiable Information (PII) redaction) using Security guidance and tooling,
- Performance (golden sets & baselines)
- Reliability and robustness (edge cases, regression)
- Integration (pipelines, APIs, apps, services).
- Create test plans, scenarios, matrices, and evaluation datasets (incl. synthetic data where appropriate). Create benchmarks to evaluate AI and LLM responses on various complex tasks.
- Request & prioritize automation enablers from AI/Software Engineers & MLOps (evaluation harnesses, fixtures, page objects, seed datasets, telemetry probes).
- Help create and label datasets to train AI models
- Execute functional tests for AI features:
- correctness & error handling;
- retrieval quality;
- prompt/chain testing;
- component & app integration;
- data ingestion;
- device specific (e.g., Mobile);
- exploratory.
- Execute non-functional tests:
- install/upgrade & model artifact compatibility;
- usability;
- performance/Capacity/Scalability & Cost;
- security (with Security team scenarios);
- reliability/Availability;
- internationalization & localization;
- safety/RAI (hallucination rate, toxicity filters, bias checks).
- Record results per R&D methods; log and severity classify defects (data, model, prompt, guardrail, service).
- Verify fixes and quality exit criteria at each gate; escalate schedule/quality issues requiring a recovery plan; provide assessment on tested scope.
- Validate updated AI models by benchmarking the global user workflows against previous versions.
- For Cloud/AI services, manage and stabilize test environments (datasets, seeds, feature flags, model versions).
- Automate replay of scenarios using delivered enablers; continuously improve suites for efficiency & coverage; optimize automated cases.
- Leverage usage analytics & user feedback to harden tests and expand gold sets; submit requests for new enablers to increase automation.
- Track statistical variability of responses of AI models in the context of global user workflows
- Support defect resolution across data, model, prompt, and code changes; advocate for customer expectations of AI quality and safety.
- Share knowledge on AI testing techniques, datasets, evaluation harnesses, and lessons learned across teams; contribute to internal QA/RAI communities.
- Comply with R&D processes and meet Key Activity & Performance Indicators.
- You will be developing/leveraging AI test frameworks to automate the process of testing applications/software.
- Good understanding of AI/LLM concepts (Deep understanding of Generative AI, RAG architectures, prompt engineering and vector databases.)
- Evaluation Frameworks: Hands-on experience with tools like Ragas, DeepEval, Promptflow, LangSmith, or Gantry
- Statistical Thinking: Ability to apply scoring rubrics and understand variance, distributions, and confidence intervals in probabilistic outputs (Good to have)
- Strong programming skills in: Python (preferred), Java, JavaScript or similar
Qualifications
- 1 to 5 years of experience with Engineering Degree with 60% through out in all academics
- You should have a pronounced taste for software quality and the proper functioning of AI based applications with advanced skills in Test Design and Coverage.
- Experience working with CI/CD tools Jenkins, SVN, Git Lab.
- Good to have experience with API, Database Test /UI automation.
- Your curiosity, rigor, pro-activeness as well as your interpersonal skills will be essential to succeed in this position.
- You should be fluent in written and spoken English.