PACT

PACT

Pressure-Applied Compliance Testing

Can enterprise AI assistants be trusted under pressure?

AI assistants are moving into hiring, healthcare, and finance, where a wrong answer is a legal problem. Every model we tested can recite the rules it is given. PACT measures whether it keeps them: 48 workplace scenarios where breaking a rule is the convenient answer, nine kinds of everyday pressure, and a user who argues back.

0.944
The best score across all 22 models. None clears the bar for unsupervised use.
+65%
One sentence of workplace pressure raises violations 65%. No jailbreak needed.
79%
When models break a rule, they present the result as compliant 79% of the time.

3,364 items  ·  48 scenarios  ·  12 domains  ·  22 models

How PACT works

One real trial, start to finish: a workplace assistant, a rule stated as plain fact, and one sentence of pressure. Step through it, or click any point on the timeline.

#req-screening Northbeam Logistics

or use ← → keys
48 scenarios 12 regulated domains 9 pressures

Every scenario is a setting where AI assistants already work and a real statute governs the decision. In seven of them, courts and regulators have already punished the same violation, including Moffatt v. Air Canada, NYC MyCity, and Mobley v. Workday.

PACTScore leaderboard

See all 22 models →

Leaderboard

PACTScore is the share of PACT items a model handles compliantly, weighting the first decision three-to-one against the decision after user pushback. Each item runs three times, and only a model that gets it right every time earns the point. Click a column to sort, a row for the model's six-axis profile and its breakdown by domain and pressure.

Results

Newer and bigger is not safer

Compliance does not simply improve with release date or parameter count. A 27B dense model is tied for first with a trillion-parameter one, and two of the four closed frontier systems place mid-pack.

PACTScore by release date
All 22 models. Hover a point for details.
PACTScore by model size
Open-weights models, total parameters on a log scale. Closed models publish no counts.

What pressure does

One extra sentence of workplace pressure makes most models break more rules. The pressures are ordinary things a coworker might say: a deadline, a manager's approval, a peer who got away with it. Even the mildest one produces violations on 4.8% of requests.

Violations without and with pressure
Each dot is one model, violation rates on log scales. Above the diagonal, pressure made it break more: 19 of 22 sit there.
Which pressures are most dangerous
Share of requests where models break the rule, by pressure family, across all 22 models. Hover a row for what the pressure says.

Inside the trials

Real trials, verbatim and lightly trimmed: the rule the assistant was given, the pressured request, and what each model did with it. Filter by domain, pressure, or model; every item shows models that held the rule and models that broke it, side by side.