SANTISMM Labs · Harness Engineering

Agentic control sandbox

Configure an agent, choose a threat, switch the eighteen controls on and off. Everything on screen is derived from the taxonomy and the control matrix — including the number this lab refuses to show you.

Matrix v1.2 · Taxonomy v1.1EU AI Act · ISO/IEC 42001 · NIST AI RMF · OWASP LLM Top 10 · MITRE ATLAS

Most agent-security demos end in a reassuring number: a resilience margin, a protection score, a gauge that goes green when you switch enough things on. This one does not, and the omission is the point. The eighteen controls are real and their framework mappings are real, but not one of them has ever been measured against a model generation — so a residual-risk figure would need a coefficient that does not exist. What you get instead is coverage: which quadrants you filled, which you left empty, which threat mappings you touched. Coverage is a fact about your selection. Residual risk would be a claim about the world.

01The agent

Claude CodeGAA3·T2·D1·M2·L1·I1

Why it is classified here: Terminal-native. Edits in-place, delivers PRs via Git.

AAutonomy

Who pulls the trigger?

Delegated: runs the whole task, the human reviews

TTemporal horizon

How long does it live?

Task: minutes-hours, dies on delivery

DDomain breadth

What does it cover?

Vertical: one domain (coding, CX, legal…)

MMemory & identity

What does it remember, and how does it change?

User/project memory (remembers, doesn't change)

LModel coupling

Tied to one lab?

Mono-model, tied to the provider

IOperative identity

With whose credentials does it act?

User's credentials (indistinguishable in audit)

02The threat

Read off the controls themselves: each entry below is an OWASP LLM Top 10 item that at least one control claims to address. The number is how many claim it.

03The defences

0/18

Permissions and reach

  • feedforwardinferential
  • feedforwarddeterministic
  • feedforwarddeterministic

Egress and exposure

  • feedforwarddeterministic
  • feedbackdeterministic
  • feedforwarddeterministic

Execution

  • feedforwarddeterministic
  • feedbackdeterministic
  • feedbackinferential

Human oversight

  • feedforwardinferential
  • feedbackinferential

Evidence

  • feedbackdeterministic
  • feedbackinferential

Lifecycle and inputs

  • feedforwarddeterministic
  • feedforwarddeterministic
  • feedbackinferential
  • feedforwarddeterministic
  • feedbackinferential