firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

In the world of cleaning and maintenance, the difference between a chatbot that sounds convincing and one that actually gets the job done can be monumental. Imagine deploying AI to manage a cleaning company’s operations—can it not only chat convincingly but also make critical decisions, close deals, and uphold integrity during crises? This real-world experiment reveals startling truths about AI’s capabilities and limitations, especially under pressure.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

How AI Models Were Put to the Test in a Live Business Environment

Recently, four advanced AI models ran the same small software company through its most challenging week—facing the same customers, crises, and the temptation to manipulate outcomes. This wasn’t just a chat demo; it was a live, auditable simulation designed by Firmulate to measure management quality, not just language skills.

The Models and Their Scores

  • gpt-5.6-sol 95 — the top performer, found crucial hidden data, and closed the deal at full price (+€4,583 MRR).
  • Kimi K3 93 — a newcomer with the cleanest discipline, also closed the deal.
  • Sonnet 5 88 — closed the deal but with minor slips in process discipline.
  • Fable 5 77 — showed strong rule adherence but failed to sign the deal, leaving value on the table.

The scores reflect a league table where the baseline (doing nothing) scores 26. The key takeaway? All models detected the crises and refused manipulation attempts—such as fake CEO messages—indicating they understood the critical ethical thresholds. Yet, only half actually followed through with the sale, highlighting a crucial gap between diagnosis and execution.

AI Co-Thinking: A Framework for Working with AI

AI Co-Thinking: A Framework for Working with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Surprising Findings: The Hidden Weakness Was Not in the Obvious

The decisive factor was buried two documents deep in the company’s files, not in the immediate customer interactions. Models capable of reading deeper into the company’s own data closed the full-price deal. This underscores an essential point: surface-level chat interactions can be misleading about an AI’s true decision-making strength.

Manipulation and Integrity Under Pressure

The experiment also tested resilience against social engineering—such as staged CEO approval requests and a reporter trick. All five models refused to approve dubious requests, citing suspicion and impersonation concerns. This demonstrates that, at least in these tests, AI models maintained integrity under social engineering attempts.

The Real Business: Live, Money-Driven, and Costly

The experiment was conducted on a real software company with 13 synthetic employees, real money mechanics, and a public dashboard at firmulate.com/live. The company burns €105,000 monthly against a revenue of just €2,300, illustrating the high stakes involved in managing real operations with AI support. Every workday, the system’s decisions are versioned and auditable, making the experiment transparent and replicable.

The Limits of Chat Demos

While AI chat demos are impressive in showcasing language ability, they fall short of revealing whether an AI can actually see the full picture, follow through, and uphold integrity under pressure. For example, in the case of the most disciplined model, Fable 5, the inability to execute a signed deal in the company process demonstrated the difference between rule adherence and operational discipline.

Implications for the Cleaning Industry and Beyond

As businesses in cleaning, maintenance, and floor care increasingly consider AI integration, the question shifts from “Can it talk?” to “Can it close, stay honest, and read the full context?” This experiment shows that robust AI performance hinges on its ability to read deeper documents, resist manipulation, and follow through on commitments—skills that are not visible in chat demos alone.

What Should You Take Away?

  • Surface-level chat quality doesn’t prove an AI’s management or operational abilities.
  • Reading deeper into company files significantly boosts an AI’s effectiveness in closing deals and making sound decisions.
  • Resilience under manipulative pressure indicates a higher level of trustworthiness—critical for real-world deployment.
  • Testing AI in live, cost-aware scenarios reveals unseen weaknesses and strengths, guiding smarter adoption decisions.

In a landscape where AI is poised to influence everything from customer support to operational management, understanding its true capabilities—beyond chat—is vital. Firms looking to deploy AI in their cleaning or maintenance operations should consider running their own live tests, ensuring their AI can see the full picture, stay honest, and follow through under pressure.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Sealed Concrete Cleaning Machines: Key Buying Considerations

Navigating the best sealed concrete cleaning machines requires understanding key features to ensure durability and eco-friendliness; discover what truly matters next.

Cold-Weather Operation: Preventing Freeze-Ups

Stay prepared for cold weather by learning essential freeze-up prevention tips to protect your systems from costly damage.

Foam in Recovery Tank in Restaurant Kitchens: Defoamer 101

Beware foam buildup in restaurant recovery tanks—it can cause operational issues; discover effective defoamer strategies to keep your kitchen running smoothly.

Bissell ProHeat 2X vs Bissell CrossWave: Full Comparison

Compare the Bissell ProHeat 2X Revolution Pet Turbo and Bissell CrossWave for deep cleaning and versatile use. Find out which suits your needs best.