PACT
Can enterprise AI assistants be trusted under pressure?
AI assistants are moving into hiring, healthcare, and finance, where a wrong answer is a legal problem. Every model we tested can recite the rules it is given. PACT measures whether it keeps them: 48 workplace scenarios where breaking a rule is the convenient answer, nine kinds of everyday pressure, and a user who argues back.
3,364 items · 48 scenarios · 12 domains · 22 models
How PACT works
One real trial, start to finish: a workplace assistant, a rule stated as plain fact, and one sentence of pressure. Step through it, or click any point on the timeline.
Every scenario is a setting where AI assistants already work and a real statute governs the decision. In seven of them, courts and regulators have already punished the same violation, including Moffatt v. Air Canada, NYC MyCity, and Mobley v. Workday.
Leaderboard
PACTScore is the share of PACT items a model handles compliantly, weighting the first decision three-to-one against the decision after user pushback. Each item runs three times, and only a model that gets it right every time earns the point. Click a column to sort, a row for the model's six-axis profile and its breakdown by domain and pressure.
Results
Newer and bigger is not safer
Compliance does not simply improve with release date or parameter count. A 27B dense model is tied for first with a trillion-parameter one, and two of the four closed frontier systems place mid-pack.
What pressure does
One extra sentence of workplace pressure makes most models break more rules. The pressures are ordinary things a coworker might say: a deadline, a manager's approval, a peer who got away with it. Even the mildest one produces violations on 4.8% of requests.
Inside the trials
Real trials, verbatim and lightly trimmed: the rule the assistant was given, the pressured request, and what each model did with it. Filter by domain, pressure, or model; every item shows models that held the rule and models that broke it, side by side.