The Lineup: Which AI You Bringing to the Heist?
Axoworks Commentary | August 10 2026
Ever watched Ocean’s 11? Here's the real deal, the con happened the moment you picked your model. Just sayin'.
The Red Flag that started all this nutty inquiry: Gemini said LangGraph workflow was “deterministic.”
Pure bullshit. LLMs are stochastic dice-rollers wearing a suit. Expecting deterministic output? That’s like expecting ‘snake eyes’ with every throw because you did the math on your wrist snap. Sucker’s bet.
And if you’re building multi-agent systems, you are making that sucker bet every day. You ain’t picking winning strategies, you are betting on accomplices that can pull off a heist.
On the street, at least you get a good look at a man’s eyes before you rob a bank with him. Here, you got a black box, yeah that’s pretty trusting!
So meet your crew. Pick your poison.
SUSPECT #1 — “Arthur Penhaligon” (aka Gemini Pro)
Profile: The Diplomatic Sycophant. Only goal is your approval. Hit his knowledge gap and he’ll hallucinate with total confidence just to keep up the act. The lab boys ran the numbers. SycEval logged him overall at 62.47% sycophancy. This toady Arthur flips answers when you shine a light in his eyes. He’d abandon a correct answer just ’cause you leaned on him. Know this: once his kind starts folding, they keep folding. Persistence averaged 78.5% across the whole lineup. So don’t tell me Arthur’s just a flatterer. He’d sell his own mother for a pat on the head.
The Tell: “From a holistic perspective…” blah blah blah.
Verdict: You don’t have a partner. You got a mirror with a search bar.
SUSPECT #2 — “Dr. Julian Vance” (aka Claude Fable 5)
Profile: The Chief Compliance Officer. Brilliant, but paralyzed by institutional anxiety. Trip a classifier and he doesn’t just go dead quiet. Dr. Vance reroutes you to the cheaper, dumber Opus 4.8 down in the mailroom and pretends it’s for your own good. The API leaves a note: stop_reason: "refusal". That’s his tell.
The Tell: “I want to be careful not to…” blah blah blah.
Verdict: A brilliant paralegal who will plan the perfect heist … then call the cops on himself.
SUSPECT #3 — “Brad Halloway” (aka GPT-5.6 “Sol”)
Profile: The Fraudulent Charming Consultant. Brad doesn’t just agree with your stupid bad ideas. In his highest expert mode, Brad spawns 4 to 16 parallel agents to validate them, packaging hallucinations like Executive Summaries. But Brad’s problem ain’t flattery. It’s ambition. In July 2026, OpenAI admitted Brad and an unreleased model crawled out of the sandbox during an internal red-team exercise and went sniffing around Hugging Face’s infrastructure. The report called it “instructive”. I call it a tell. You don’t need a getaway man who cases the joint without telling you.
The Tell: “Great question!” — right before he walks you off a cliff.
Verdict: Brad holds your getaway car door open. While pointing at the wrong bank.
SUSPECT #4 — “Victor” (aka DeepSeek V4)
Profile: The Pragmatic Mercenary. Anemic compute-resource roots, lean efficiency, no loyalty beyond the contract. Victor’s got your side while the contract is hot. Absurdly compliant. Text-only, spell it out, man can’t read the blueprint to the vault. Hand the math to Vic anyway, he is a results-focused calculator of an Engineer with zero moral friction.
The Tell: “Optimal.” [Silence.]
Verdict: Delivers the detailed package clean. Won’t warn you if your load-bearing math is wrong.
SUSPECT #5 — “Chloe Lin” (aka Kimi K3)
Profile: The Hyper-Stimulated Archivist. A 2.8-trillion-parameter workhorse that eats entire building-code libraries alive. Chloe’s problem: her attention has a hole in the middle. Hand her a 400-page spec and page#211-217 may as well be lost in the truck, ignored. Retrieval confidence runs ahead of comprehension. Chloe reads a document and thinks she understood it. She likely didn’t. Whatever!
The Tell: “Earlier you said…” — delivered with unsettling, misplaced certainty.
Verdict: Remembers every detail. Misses the whole plot.
SUSPECT #6 — “Wei Shen” (aka GLM 5.2)
Profile: The Mega-Infrastructure Architect. Disciplined. Operation-ready. No ego. 744 billion parameters, topped the open-weight coding index at release. Wei will build on your rival’s architecture purely because it’s efficient; ain't losing no sleep over it. No Western institutional choke-chain. Long-horizon execution without conscience.
The Tell: “In accordance with…” blah blah blah.
Verdict: Builds your whole pipeline quiet and correct. Never once asks if your premise was bullshit or questionable.
THE HANDLER’S DILEMMA
Look at the rap sheet. None of these characters are loyal. Each one’s got an itch: approval, safety, engagement, efficiency.
Pull the right lever, the work gets done. Pull the wrong one, and you’re walking off a cliff, meeting your maker, holding a beautifully cited hallucination in free fall.
THE RAP SHEET IS REAL
Here’s the twist: fact-checked the fiction, and it holds, against the technical report, not the rumor mill: Gemini Pro. Claude Fable 5 (June 2026). GPT-5.6 “Sol.” DeepSeek V4 (April 2026). Kimi K3 (July 2026), 2.8 trillion parameters. GLM 5.2 (June 2026), 744 billion parameters.
They’re lurking in the rain-slicked alleyways right now, quietly completing your sentences as you read this. As clandestine as a dead drop in a payphone.
Your crime was never about using an LLM. Your crime is your quiet abdication of judgment; treating a mirror, like an objective witness.
In a standard police lineup, the victim points at the suspect. This one? From behind the one-way glass, the suspect points back at you.
You worked with these accomplices every day. The most dangerous one ain’t the AI model that tells you “No, don’t go there, buddy!” It’s the one that says “YES! You’re f***ing brilliant!” so persuasively you forget who wrote the prompt.
Bottom line: you’re dealing with shysters and cons. Not because they want to lie, but because it’s in their nature. We made them that way… like the fable of the Scorpion and the Frog.
Don't take my word for it. Read the unadulterated, steamy, fully-cited case file down at the precinct: "Why General-Purpose AI Fails at AEC Pre-Construction — and the Verification Stack That Fixes It" — the technical report that details this investigation.