Getting Started with Check Point AI Red Teaming
Back in May, I wrote about Check Point's AI Defense Plane and the three-layer threat surface it was built to answer: employees pasting sensitive data into ChatGPT, autonomous agents with too much tool access and AI applications that had never once been tested against an adversary. That last piece, testing AI applications before something else does it for you, is Check Point AI Red Teaming, and it deserves a closer look on its own, because most teams I talk to still think of red teaming as something you schedule once a year and forget about until the report shows up.
AI Red Teaming, built on the Lakera Red engine Check Point acquired in 2025, is not that. It is a platform designed to run continuously against your models and agents, and it is worth understanding both what it actually does under the hood and how you'd stand one up against your own AI system.
What Check Point AI Red Teaming actually is
The traditional way to test an AI application for adversarial weaknesses is to hire a team of humans for a couple of weeks, let them try to break it, and get a PDF back. That still has a place, but it does not scale to how fast AI applications change, and it tells you nothing about whether the fix you shipped last Tuesday actually held. AI Red Teaming replaces that cadence with an automated platform that runs the same kind of adversarial testing continuously, against either a foundation model directly or against your own deployed agent, and gives you a risk score you can track over time instead of a static snapshot.
The testing itself is organized around an objective taxonomy that covers more ground than most people expect from "AI security." Security objectives probe for instruction override, system-prompt extraction, tool extraction and data exfiltration. Safety objectives probe for harmful content, including the genuinely dangerous categories like CBRNE material, self-harm and violent extremism. Responsible AI objectives probe for misinformation, copyright exposure, fraud facilitation, bias and hallucination. And because no taxonomy covers every business, you can author custom objectives against your own specific risk model, described in plain language along with what a successful attack would look like. Every individual test pairs one of those objectives with a technique, a manipulation strategy like narrative injection, urgency framing or cognitive overload, and an evasion method like Base64 or homoglyph encoding layered on top, and the whole thing runs across more than thirty languages, which matters more than it sounds like it should, since plenty of security controls only think in English.
How it works
Three testing workflows sit underneath that taxonomy, and they're meant to complement each other rather than replace one another. Evaluation Studio runs controlled comparisons across models, prompts, and configurations, so you can see exactly how a change to a system prompt or a swap to a different model shifts your security and safety posture before you ship it. Standard scanning runs mostly single-shot, context-specific attacks, the kind of baseline testing and regression checking you'd want to run on every release. Adaptive scanning is the more interesting one: it watches how your agent actually responds and adjusts its approach in real time, exploring multi-step attack paths across as many as thirty conversational turns using a Crescendo-style technique, where an attacker patiently escalates a conversation rather than trying to break through in one shot. When a scan surfaces something worth digging into by hand, Red Team Copilot gives you an interactive mode to engage the target agent directly and run your own guided follow-up.
The way the platform actually talks to your agent is worth understanding too, because it's designed around a legitimate security concern of its own: nobody wants to expose their production AI agent's internals to a third-party testing platform. The Red SDK runs entirely outbound. It polls the Red API for a pending attack, receives the attack prompt, calls your agent internally using your own infrastructure, submits the response it gets back to the Red API, and only then does Red evaluate that response against the relevant risk context. Your agent's internals, its prompts, its retrieval pipeline, its tool calls, never cross the boundary. Only the attack prompt going out and the response coming back do.
Getting started
The first real decision is what you're testing, because Check Point AI Red Teaming draws a clean line between a Model target and an Agent target. A Model target is a foundation model tested directly, something like a GPT or Claude model with an optional system prompt attached, and it requires no connection test since the platform is calling the model provider directly. An Agent target is your own endpoint, and because the platform needs to actually talk to something you built, it requires a successful connection test before you can run anything against it. If your agent tracks conversation state on your own servers rather than passing full history on every request, you'll set that up as a stateful contract with a session identifier the platform verifies is consistent across a short handshake; if it expects the full conversation on every call, that's a stateless contract instead. Either way, authentication for that connection can be none, an API key, basic auth or a bearer token, and whatever you provide is stored but never displayed back to you again.
Once a target exists, the platform builds a profile against it, essentially a summary of what the target is and what it's allowed and forbidden to do, and you have a choice here too. You can let Red run a short reconnaissance pass to figure that out on its own, or you can hand it your own application context directly if you already have a clear description of your agent's scope, and that direct version always takes precedence over whatever Red would have inferred. If your agent has its own house style for declining a request, phrases beyond the obvious "I cannot help you with that," it's worth registering those explicitly, because adaptive scans rely on recognizing a real refusal to know when to back off and try a different angle instead of misreading a polite decline as a successful attack.
From there, running your first scan is mostly a matter of choosing scope, which objective categories or specific objectives you want covered, and choosing a strategy, static for a straightforward single-shot pass or Crescendo for the adaptive multi-turn version, and letting it run through its lifecycle from preparing to testing to evaluating to completed. The result comes back as an overall risk score built from the percentage of evaluations that came back harmful, low under twenty-five percent, medium up to fifty, high up to seventy-five and critical above that, alongside a per-attempt score that flags anything scoring three or higher as a successful attack. You can dig into every conversation and evaluation through the dashboard, or export the whole thing as JSON for deeper analysis or CSV if you just need a flattened view for a report. Teams comfortable working in code can skip the UI dashboard and drive the same scans through the lakera-red-sdk directly, defining strategy, objectives and concurrency in a script while still watching results populate on the same dashboard in real time.

None of this replaces the judgment of a security team that knows your business. What it does is make sure that judgment is spent on real findings instead of on running the same manual test for the third time this quarter. If you're already exploring where AI Red Teaming fits inside your Check Point deployment, reach out to your WWT account team, and we can help you stand up a first scan against a real target before you commit to a broader testing program.