The practice

A full offensive practice, specialised in models.

Model and LLM assessments, application and cloud testing, and red-team readiness — delivered by operators who finish the chain. We do not sell a platform.

01 — Assessment

Model & LLM security assessment

Direct and indirect prompt injection, multi-turn and encoded jailbreaks, system-prompt extraction, guardrail bypass, PII and tenant regurgitation, model extraction, and denial-of-wallet. Typical duration: 2–3 weeks for a single production assistant.

02 — Assessment

Agentic systems & tool-use

Function-calling, MCP servers, plugins, multi-agent orchestration, goal hijacking, and excessive agency. Aligned to the OWASP Top 10 for Agentic Applications.

03 — Assessment

RAG & data-pipeline security

Indirect injection via documents, email and tickets; corpus and embedding poisoning; connector over-permission; cross-tenant retrieval.

04 — Assessment

Hybrid application penetration testing

The surrounding web and API estate: authentication, object-level authorisation, business logic, and output handling where model text is executed or stored.

05 — Assessment

Cloud & inference infrastructure

GPU clusters, model registries, vector databases, inference gateways, secrets, isolation patterns, and trust boundaries.

06 — Assessment

Model lifecycle & supply chain

Fine-tune and training-data poisoning, CI/CD gate tampering, model provenance, dependency risk, secrets in notebooks and weights.

07 — Operations

AI-focused red team & readiness

Multistep adversary emulation against the model pipeline. Purple-team variants work live with your detection staff. Tabletop drills for AI incidents.

08 — Review

GenAI architecture & governance review

Model and tool deployment, data flows, human-in-the-loop, vendor risk, and control coverage against NIST AI RMF, ISO/IEC 42001, and EU AI Act evidence.

09 — Retainer

Continuous intrusion programme

A named operator retains scope over your models and agents, retesting after releases, keeping an attack library current.

10 — Review

Secure code & prompt review

System prompts, tool schemas, retrieval filters, output parsers, and the glue that turns a completion into an action.

How engagements are sized

Three depths

Snapshot

5–8 days

One assistant or agent. OWASP LLM baseline and a written verdict on whether you can ship.

Standard

2–4 weeks

Full-stack assessment: model, RAG or tools, application, and cloud perimeter. Retest window included.

Campaign

6–8 weeks

AI red team plus purple team. Lifecycle and human layers in scope. Board brief and detection-gap register.