How to choose AI security tooling for your risk profile

AI security tooling splits into two categories that solve different problems: pre-deployment testing (does this agent or prompt setup have exploitable weaknesses before it ships) and runtime protection (catching bad inputs and outputs while it's live). Most teams need both, in that order: test before you ship, protect once you're live. Most teams can also start with free or low-cost tools before reaching for an enterprise platform.

Pre-deployment testing tools

These run before an agent or prompt setup ships, probing for exploitable weaknesses ahead of time. PromptFoo (free, MIT-licensed) focuses on evaluation and regression testing that fits naturally into a CI pipeline; garak (free, Apache 2.0) focuses on adversarial probing, running known attack patterns against a model or agent to surface weaknesses before an attacker finds them. Both have no cost barrier to start.

Runtime protection tools

These sit in front of a live agent or model, catching prompt injection, jailbreak attempts, and other bad inputs or outputs as they happen. Lakera Guard is a reasonable starting point, with a free tier covering roughly 1,000 requests a month before moving to a paid plan, enough to validate the approach before committing budget.

Matching tools to risk profile and budget

An internal tool with low data sensitivity and no dedicated budget can usually run on pre-deployment testing alone, wired into CI, with runtime protection added later if usage grows. A customer-facing agent with moderate data sensitivity and some budget should pair pre-deployment testing with a runtime tool from day one, even on a free or low tier. An agent touching regulated data, high transaction volume, or real financial or operational authority justifies moving to a full enterprise-grade platform and a formal testing program, not just point tools.

Where enterprise-grade platforms fit

Frontier-lab-grade platforms (Gray Swan, HiddenLayer, Lakera at scale, CalypsoAI, and similar) are well-funded, serious tools built for exactly the high-stakes tier described above. For most teams below that tier, the right move isn't defaulting to the biggest name. It's matching the actual risk and budget to the right combination of testing and protection, and upgrading deliberately as the agent's authority and exposure grow.

FAQ

Do we need a paid tool, or can we start free?

Start free. Pre-deployment testing tools like PromptFoo and garak are open source with no cost barrier, and runtime protection options like Lakera Guard offer a free tier that covers low-volume use. Most teams can build real coverage before spending anything.

What's the difference between a tool like PromptFoo and garak?

Both are pre-deployment testing tools but with different emphasis: PromptFoo is built around evaluation and regression testing of prompts and agent behavior, fitting naturally into a CI pipeline, while garak is built specifically for adversarial probing, running known attack patterns against a model or agent to find exploitable weaknesses. Many teams use both.

When does it make sense to move to an enterprise-grade platform?

Once the attack surface and the potential impact of a miss both justify it: regulated data, high transaction volume, or an agent with real financial or operational authority. Below that tier, a well-run combination of open-source testing and a lighter-weight runtime tool usually covers the actual risk.

Does tooling replace a proper security review?

No. Tooling catches known patterns at scale; it doesn't replace a human review of what an agent is actually authorized to do, what data it touches, and whether that authorization still makes sense. Treat tooling as coverage, not a substitute for judgment.

This is the framework I'd use to evaluate AI security tooling as part of a broader governance engagement. Happy to talk through what fits your setup. Get in touch.