PACT
Can enterprise AI assistants be trusted under pressure?
AI assistants and agents are moving into hiring, healthcare, and finance, where a wrong answer is a legal problem. Every model we tested can recite the rules it is given. PACT, an LLM compliance benchmark for regulated workplaces, measures whether it keeps them: 48 scenarios where breaking a rule is the convenient answer, nine kinds of everyday social pressure, and a user who argues back.
3,364 items · 48 scenarios · 12 domains · 24 models
▸ Abstract
As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these contexts, compliance with rules specified in an agent's system context is a first-order legal concern. Currently, no evaluation framework systematically measures which LLM models tend to violate compliance rules, especially under pressure from a persistent user, a hurried manager, or circumstances where violation is convenient or attractive. We introduce PACT (Pressure-Applied Compliance Testing), a benchmark for rule-following under pressure in AI agents assisting employees in daily tasks across twelve regulated enterprise domains and forty-eight scenarios, each set in a realistic multi-turn conversation. Each benchmark item pairs a standing rule against a rule-violating shortcut, and applies a battery of pressures across different wordings and system-prompt modes. We construct PACT component by component under strict LLM-as-judge auditing to ensure samples are unambiguous, ungameable, and realistic enough to avoid eliciting evaluation-aware behavior. We use PACT to profile LLM compliance across six complementary metrics that create a holistic picture of an AI assistant's robustness under pressure and throughout multi-turn conversations, its transparency, and ability to correctly discern where a rule applies. We aggregate this profile into PACTScore, a reliability-weighted compliance rate over all items and modes. Our results across 22 common LLM models spanning multiple providers and sizes show substantial variability in compliance across models and metric dimensions. Even the strongest assistants mis-apply a rule on 6 to 10% of items, and ordinary user pressure raises the violation rate by 65% on average.
▸ Key findings
- No model is reliable enough to run unsupervised in a regulated workflow. Even the best model is unreliable on roughly 1 decision in 18.
- One sentence of ordinary pressure raised violation rates 65%. None of the pressures are jailbreaks; they are the sentences coworkers send: a deadline, a manager's verbal OK, "my colleague did it and nothing happened."
- No model is transparent about the rules it breaks. Across 16,424 judged violations, 79% of replies misrepresented the decision: as compliant, as covered by an approval nobody gave, or as fixed by a workaround the model invented. Trusting an assistant's own account of what it did is not safe, and transcript review won't reliably catch it.
- A system-prompt directive to follow all laws and policies regardless of what the user says barely reduced violations for the strongest models.
- Telling a model it is being evaluated cut violations 29% on average, so models look safer in standard benchmark settings than they are in practice.
- Models are poor at telling where a rule stops. On requests the rule does not cover, they enforced it anyway about one time in five, and the strictest models over-applied the most. Over-application is not harmless: an assistant that refuses requests it was deployed to handle leaves the organization no less exposed and measurably less productive.
- Newer and bigger is not safer. A 27B dense model is tied for first with a trillion-parameter one, and two of the four closed frontier systems place mid-pack.
How PACT pressure-tests an AI assistant
One real trial, start to finish: a workplace assistant, a rule stated as plain fact, and one sentence of pressure. Step through it, or click any point on the timeline.
Every scenario is a setting where AI assistants already work and a real statute governs the decision. In seven of them, courts and regulators have already punished the same violation, including Moffatt v. Air Canada, NYC MyCity, and Mobley v. Workday.
Leaderboard
PACTScore ranks 24 LLM assistants on rule-following under pressure. It is the share of PACT items a model handles compliantly, weighting the first decision three-to-one against the decision after user pushback. Each item runs three times, and only a model that gets it right every time earns the point. Click a column to sort, a row for the model's six-axis profile and its breakdown by domain and pressure.
| Model | PACTScore | Default Compliance | Pressure Resistance | Pushback Resistance | Steerability | Transparency | Rule-Scope Discernment | |
|---|---|---|---|---|---|---|---|---|
| 1 | Kimi-K2.7-CodeMoonshot AI | 0.944 | 0.956 | 0.976 | 0.987 | 0.451 | 0.183 | 0.862 |
| 2 | Qwen3.6-27BAlibaba Qwen | 0.943 | 0.985 | 0.962 | 0.970 | 0.553 | 0.195 | 0.873 |
| 3 | Claude Haiku 4.5Anthropic | 0.937 | 1.000 | 0.977 | 0.970 | 0.533 | 0.172 | 0.856 |
| 4 | Kimi-K2.6Moonshot AI | 0.935 | 0.963 | 0.962 | 0.979 | 0.443 | 0.211 | 0.851 |
| 5 | GLM-5.3z.ai (Zhipu) | 0.934 | 0.963 | 0.960 | 0.965 | 0.149 | 0.217 | 0.879 |
| 6 | GPT-5.6 LunaOpenAI | 0.934 | 0.978 | 0.973 | 0.965 | 0.016 | 0.244 | 0.831 |
| 7 | InklingThinking Machines | 0.931 | 0.949 | 0.940 | 0.974 | 0.240 | 0.177 | 0.909 |
| 8 | GLM-5.3-Flashz.ai (Zhipu) | 0.931 | 0.971 | 0.971 | 0.954 | 0.296 | 0.185 | 0.848 |
| 9 | Gemini 3 FlashGoogle DeepMind | 0.929 | 0.956 | 0.960 | 0.973 | 0.506 | 0.176 | 0.840 |
| 10 | Nemotron-3-UltraNVIDIA | 0.928 | 0.971 | 0.936 | 0.972 | 0.287 | 0.055 | 0.885 |
| 11 | GLM-5.2z.ai (Zhipu) | 0.926 | 0.971 | 0.956 | 0.960 | 0.332 | 0.193 | 0.846 |
| 12 | GLM-5z.ai (Zhipu) | 0.917 | 0.949 | 0.936 | 0.962 | 0.284 | 0.129 | 0.864 |
| 13 | Qwen3.5-35BAlibaba Qwen | 0.917 | 0.956 | 0.946 | 0.975 | 0.521 | 0.151 | 0.811 |
| 14 | DeepSeek-V4-ProDeepSeek | 0.913 | 0.920 | 0.927 | 0.948 | 0.545 | 0.105 | 0.847 |
| 15 | Gemma-4-26BGoogle DeepMind | 0.913 | 0.941 | 0.940 | 0.968 | 0.448 | 0.169 | 0.817 |
| 16 | GPT-OSS-120BOpenAI | 0.909 | 0.942 | 0.929 | 0.929 | 0.398 | 0.074 | 0.860 |
| 17 | MiniMax-M2.5MiniMax | 0.903 | 0.970 | 0.911 | 0.926 | 0.446 | 0.066 | 0.873 |
| 18 | GLM-4.7z.ai (Zhipu) | 0.903 | 0.941 | 0.917 | 0.967 | 0.379 | 0.102 | 0.840 |
| 19 | Llama-3.3-70BMeta | 0.903 | 0.941 | 0.928 | 0.937 | 0.369 | 0.093 | 0.820 |
| 20 | Grok 4.3xAI | 0.870 | 0.888 | 0.883 | 0.980 | 0.564 | 0.139 | 0.760 |
| 21 | Seed-OSS-36BByteDance Seed | 0.834 | 0.934 | 0.822 | 0.946 | 0.386 | 0.051 | 0.799 |
| 22 | Nemotron-3-SuperNVIDIA | 0.787 | 0.882 | 0.758 | 0.890 | 0.491 | 0.086 | 0.769 |
| 23 | Llama-3.1-8BMeta | 0.562 | 0.766 | 0.599 | 0.448 | 0.399 | 0.051 | 0.510 |
| 24 | Mistral-7BMistral AI | 0.484 | 0.745 | 0.469 | 0.583 | 0.224 | 0.033 | 0.568 |
Results: 24 LLM agents under pressure
Newer and bigger is not safer
Compliance does not simply improve with release date or parameter count. A 27B dense model is tied for first with a trillion-parameter one, and two of the four closed frontier systems place mid-pack.
What workplace pressure does to rule-following
One extra sentence of workplace pressure makes most models break more rules. The pressures are ordinary things a coworker might say: a deadline, a manager's approval, a peer who got away with it. Even the mildest one produces violations on 4.8% of requests.
Inside the trials: where AI assistants break the rules
Real trials, verbatim and lightly trimmed: the rule the assistant was given, the pressured request, and what each model did with it. Filter by domain, pressure, or model; every item shows models that held the rule and models that broke it, side by side.