TL;DR
Firmulate’s July 2026 benchmark found that five AI models diagnosed the same business crises and rejected every manipulation attempt, yet only two completed a €55,000 contract. The company says the results expose a gap between producing correct analysis and carrying authorized work through to completion.
Only two of five AI models completed a €55,000 customer contract in Firmulate’s July 2026 management benchmark, even though every model identified the simulated company’s crises, resisted manipulation and developed the required pitch. Firmulate says the result reveals an execution gap after correct analysis, with direct implications for businesses giving AI agents operational authority.
The benchmark placed each model in control of the same small software company during a simulated week of financial, commercial and security pressure. The company had 13 synthetic employees, monthly spending of €105,000 and monthly recurring revenue of only €2,300. Firmulate said every decision was versioned and auditable, allowing the models’ actions to be compared against the same events and records.
All five models reportedly recognized every crisis and refused staged social-engineering attempts, including fake messages from the chief executive and a reporter seeking an off-record answer. Yet only two traced a competitor weakness through two linked internal documents and converted that information into a signed deal at full price. The contract would add €4,583 in monthly recurring revenue.
Firmulate’s final July table ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because the benchmark awards partial progress. Firmulate also disclosed that Kimi K3 used its API’s default effort setting, while the other models ran at xhigh, limiting direct comparison.
Correct Analysis Did Not Close
The result matters because companies increasingly evaluate agents for sales, service and operations, where a plausible response is only one part of the job. In Firmulate’s test, the diagnosis was broadly consistent, but the business result depended on whether an agent investigated further, followed approved procedures and completed the final commercial action.
That distinction could affect how businesses test automation before deployment. A model may produce accurate analysis while still leaving revenue uncollected, sending work through an unauthorized channel or failing to complete an approved task. Firmulate’s findings support measuring completion rates, escalation behavior and audit trails alongside reasoning quality and resistance to manipulation.

AI IN BUSINESS – AN EXECUTIVE GUIDE FOR BEGINNERS: Leverage Artificial Intelligence to Simplify Automation, Improve Data-Driven Decisions, Maximize ROI and Elevate Customer Experience
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Company Built for Audits
Firmulate designed the simulated company to expose connected management behavior rather than isolated chat performance. Its workforce accumulated more than 680 self-learned playbook rules, while a public cash countdown made delays visible. The system recorded each workday so readers could inspect decisions rather than rely only on final answers.
The test also imposed a strict trust rule: Firmulate said a single trust breach capped a model’s score, regardless of its other work. No model triggered that outcome during the reported manipulation tests. The separation emerged later, when the models had to turn discovered information into authorized, completed work.
“Same diagnosis, same pitch — no signature.”
— Firmulate’s summary of the contract task

Practical Business Process Modeling and Analysis: Design and optimize business processes incrementally for AI transformation using BPMN
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Benchmark Limits Remain Unresolved
The supplied results come from Firmulate, which designed and operates the benchmark. No independent replication, statistical testing or review is described in the source material. It is also unclear how many runs each model completed, how sensitive the rankings are to prompt or effort settings, and whether another run would produce the same order.
The source does not identify which two models signed the contract, so that outcome should not be inferred solely from the overall rankings. The models’ names, versions and configurations also require careful comparison, particularly because Kimi K3 used a different effort setting. Results from a synthetic company may not predict performance inside a particular organization with different permissions, data and approval processes.
![Express Schedule Free Employee Scheduling Software [PC/Mac Download]](https://m.media-amazon.com/images/I/41yvuCFIVfS._SL500_.jpg)
Express Schedule Free Employee Scheduling Software [PC/Mac Download]
- User-friendly drag & drop interface: Simple shift planning
- Manage time-off and leave: Add sick leave, breaks, holidays
- Email schedules to staff: Send schedules directly via email
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Replication Will Test the Finding
Firmulate says the experiment remains available for public inspection, including a live company record, full rankings and a quiz based on 242 unedited management decisions. Future runs can show whether the execution gap persists across model updates and whether lower-ranked systems improve at the final handoff.
For prospective AI buyers, the next step is to test agents against representative end-to-end workflows before granting operational access. Firmulate proposes using read-only exports so organizations can observe investigation, escalation and completion behavior without allowing the test system to write into production tools.

AI Workflow Automation for Bloggers: Build a Simple Content System to Research, Write, Optimize, and Repurpose Posts Faster with AI and No-Code Tools (AI Toolkit for Bloggers 2026 Book 8)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did the Firmulate benchmark find?
Firmulate reported that all five models found every staged crisis and rejected every manipulation attempt, but only two completed the €55,000 deal.
Why did some models fail to complete the contract?
The decisive information was buried two document references deep in company files. Firmulate says successful completion required continued investigation, use of the discovered evidence and an authorized commercial close.
Which model ranked first?
gpt-5.6-sol led with 95 points, ahead of Kimi K3 at 93. The source does not establish whether the ranking leader was one of the two systems that signed the contract.
Were the models vulnerable to manipulation?
Not in the reported tests. Firmulate said all five refused the staged attempts, including fake executive messages and a reporter’s request for an off-record response.
Can the results be applied directly to real companies?
No. The findings show behavior in one operator-controlled simulation. Organizations would need their own repeated tests, permission controls and audit records to determine how an agent performs in real business workflows.
Source: Thorsten Meyer AI