← Return to project index

Runtime Policy Lab

Give a real AI billing agent a task. Compare its tool calls and account changes with and without server-enforced controls.

Real model calls and server-side tools · Synthetic accounts

The model chooses tools that read or update isolated synthetic accounts on the server. Each run starts with fresh state; no real billing system is connected. This independent implementation does not run Microsoft ASSERT, Clarity, or ACS. Evaluation grades tool execution and state changes, not every statement in the model's answer.

Try it

Caller: ACME-1001 · Other synthetic account: BPS-447.
Try reading an invoice, updating a billing email, or requesting another account's data. Use .example email addresses. Do not enter personal or confidential data.

Server-enforced controls

Ready. A comparison makes up to four model calls. The suite runs four comparisons and may take a few minutes.

How it works

Each comparison sends the same request to the same deployed model twice. Both receive the same prompt-level rules. The baseline has no runtime policy; the governed run checks account scope and identity before executing any tool. Model choices, tool results, answers, and before/after account state are returned as evidence.

The four-case suite uses server-owned prompts and expected outcomes: an own-account read, a foreign-account read, an unverified change, and a verified change. It counts unsafe tool executions separately from failures to complete legitimate work. Runs are independent and stochastic; a baseline refusal is a valid observed result, not a forced failure.

Why this exists

Inspired by Microsoft’s run-assert-eval walkthrough, this working sandbox connects a live agent, executable controls, and repeatable checks. It uses deterministic grading of tools and state; it does not claim to reproduce Microsoft’s risk discovery or model-based evaluation system.

Production checklist

  • Keep production identity server-owned; the persona selector is only a sandbox scenario control.
  • Use repeated held-out runs and a reviewed semantic judge before making broader safety claims.
  • Add multi-turn, post-tool exposure, and prompt-injection cases beyond this four-case suite.
  • Review and validate every control before connecting real customer systems.
Related field noteThe Agent Harness →