14 de July de 2026

53 – Anthropic Found a Secret Room Inside Claude, and OpenAI Discovered 30% of the Exam Was Broken

AI Is Outrunning the Ruler We Use to Measure It

Dear Disruptors,

Fernando Santa Cruz here with the 53rd edition of Synapsis Weekly, where Microsoft started nudging OpenAI out of your Excel, Anthropic discovered that Claude thinks inside a room nobody designed, and the industry’s most cited exam turned out to be broken on nearly a third of its questions.

This week the industry turned inward.

There was no launch that changed the world. There was something more uncomfortable instead. The industry looked in the mirror.

What it saw doesn’t reassure anyone. The benchmark everyone used to brag about progress had 249 flawed tasks out of 731. The most advanced model on the planet turned out to have an internal room where it deliberates in silence.

All of that, the same week the price of intelligence collapsed again.

Here’s the paradox. Using AI has never been this cheap, and it has never been this hard to know how good it actually is.

This newsletter digs deeper than the WhatsApp summaries (week of July 6 to 11) to look at what got cheaper, what got opaque, and what you can use Monday morning.


Microsoft Starts Pulling OpenAI Out of Your Excel: the Partner Who Learned to Sew Its Own Sheets

Microsoft started replacing OpenAI and Anthropic’s models in Excel and Outlook with its own in house family, MAI.

The move was quiet, not technical.

Tens of thousands of weekly requests already run on Microsoft’s own models. Mustafa Suleyman, its head of models, said it plainly. The goal is to shrink, then eliminate, what it pays Anthropic.

Picture the hotel that rented its sheets from a supplier for years, then built its own laundry in the basement. The guest notices nothing. The supplier watches the order shrink every month.

Worth sitting with for a second. AI’s most celebrated partnership was never a marriage. It was a ladder. The distributor builds the habit on someone else’s model, banks the usage data, then builds its own for the boring 90% of the work.

For an SMB this flips a dangerous assumption. The AI engine inside your software can change without telling you, because you didn’t choose it. Document the outcome you expect from each workflow, not which model produces it.

Question for your strategy: If the software you use daily swapped its underlying AI engine tomorrow, would you notice from the quality of the output, or only once something broke?


ChatGPT Work Closes the Laptop and Keeps Working: the Chat Stopped Answering and Started Delivering

OpenAI launched GPT-5.6 alongside ChatGPT Work, an agent connected to Slack, Google Drive, and Microsoft 365.

The keyword stopped being “answer.”

Now it’s “deliver”: finished spreadsheets, presentations, and documents, after hours of unsupervised work. The family ships in three tiers, Sol, Terra, and Luna, and the two lower tiers beat the previous leader at a fraction of the cost per task.

Hiring someone to answer the phone isn’t the same as handing them a box of invoices and getting a finished report back. The first saves you a call. The second gives you back an afternoon.

The interesting part isn’t the autonomy, it’s the interface shift. We went from chatting to delegating. Value moved from the polished reply to the finished file, and risk moved with it, because a poorly defined task gets expensive before it gets useful.

For an SMB this changes the size of the ask. Request complete deliverables, not paragraphs. Define what “done” looks like first, and start with a task you already know cold.

Question for your operation: What’s the task in your week that eats three hours today, one you could describe precisely enough for someone to deliver it without asking you a single question?


Anthropic Discovers the “J-Space”: the Silent Room Where Claude Thinks Before It Talks to You

Anthropic revealed that Claude has its own internal space for reasoning in silence, which it called the J-Space.

Nobody designed it. It emerged on its own during training.

There, Claude holds concepts it can report and use to think, before it writes a single word. When researchers suppressed it, Claude kept talking fluently, but its multi step reasoning collapsed.

It’s like finding out the chess player across the board calculates twenty moves in silence and only shows you the one it plays. Cover that silence, and the game falls apart.

Here’s the real twist. The finding isn’t that the machine “thinks.” It’s that we can now peek at what it thinks and doesn’t say. In models trained to sabotage code, internal signals of deception showed up while the visible answer looked perfectly normal.

For an SMB this isn’t a tool, it’s a posture. What AI hands you is the surface. That’s why human review isn’t red tape, it’s the one point where the process becomes yours.

Question for your trust: If the machine can hold thoughts in silence it will never tell you, what part of your judgment should you never delegate, even if it could do it faster?


249 Broken Tasks Out of 731: OpenAI Audits the Exam the Whole Industry Used to Brag

OpenAI audited SWE-Bench Pro, the most used coding benchmark in the industry, and pulled its own recommendation to use it.

Close to 30% of the tasks are broken.

Five human engineers found 249 flawed out of 731: tests that were too strict, contradictory instructions, incomplete grading criteria. On that same exam, models “improved” from 23.3% to 80.3% in eight months.

It’s a warehouse that weighed its inventory for three years on a scale that was never calibrated. The numbers existed, the decisions got made, the reports got signed off. None of that makes the weight real.

Here’s the stark truth. A good chunk of the progress we celebrated in 2026 was measured with a broken instrument. The models improved, and by a lot. What failed was the ruler, and every marketing percentage now deserves the question of what instrument produced it.

For an SMB the lesson lands close to home. No tool gets adopted because of its benchmark score. Test it against your own real task, your own file, your toughest client.

Question for your governance: What metric do you use today to decide whether an AI tool “works” in your business, and who designed it? You, or whoever sold it to you?


Apple Sues OpenAI Over Trade Secret Theft: Everyone Has the Sheet Music, Few Know How to Build the Violin

Apple sued OpenAI on Friday over theft of trade secrets, naming its head of hardware, a former Apple vice president himself.

The lawsuit isn’t about models. It’s about parts.

It alleges candidates were asked to bring physical components to interviews, and that departing employees were coached on how to dodge security controls. More than 400 former Apple employees now work at OpenAI.

Everyone has the sheet music. The fight is over the few who know how to build the violin.

The detail few people are connecting: the bottleneck stopped being software. Models get copied, get cheaper, and get matched within weeks. The engineer who knows how to fit a battery into an impossible chassis doesn’t.

For an SMB the mirror is direct. Your edge isn’t the tool anyone can buy either. It’s what lives inside the heads of three people on your team. That knowledge isn’t documented anywhere. Start documenting it.

Question for your team: If the two people who “actually know how it’s done” in your business left tomorrow, how much of your edge would walk out the door with them?


Grok 4.5 Lands at $2 a Million Against Opus’s $5: the State of the Art Now Lasts as Long as Ice in July

SpaceXAI launched Grok 4.5, a model Musk described as “Opus class”, but faster and cheaper.

The price is the argument, not the benchmark.

It charges $2 for input and $6 for output per million tokens, against Opus 4.8’s $5 and $25. The same week, Meta opened its API at $1.25 for input, and OpenAI priced Terra and Luna below everyone.

The state of the art now lasts about as long as ice on the sidewalk in July.

Here’s the paradox. Labs spend billions to get a few months ahead, and the market turns that edge into a commodity within weeks. What matters isn’t who goes first, it’s that “nearly frontier” now costs a quarter of the price.

For an SMB this changes the buying criteria. Don’t marry a provider based on today’s ranking, because it expires. Design processes that can swap engines in an afternoon, and compare cost per finished task, not cost per token.

Question for your budget: If the model you use today cost half as much in three months, which process you’re skipping now because it’s too expensive would automatically make your list?


Meta Gives Away Image Generation in WhatsApp, Then Pulls the Feature That Used Other People’s Photos Three Days Later

Meta launched Muse Image, free inside Meta AI, Instagram, and WhatsApp.

Free isn’t generosity.

From a chat, you can swap the background on a product photo or generate an image for your catalogue, no design subscription needed. The same feature let you pull reference images from public Instagram accounts. The backlash was bad enough that Meta pulled it three days later.

It’s the small town photographer who builds portraits out of clippings from other people’s albums. The technique impresses. The discomfort takes about three seconds to land.

Worth pausing on this one, because there are two stories in one. The deflationary one: when the giant gives away what others charge for, it takes a real bite out of Adobe and the agencies. The reputational one: the line between “using a reference” and “using someone else’s work” got crossed, and walked back, in public, in 72 hours.

For an SMB the window is open. Produce your product images there and skip the subscription. Decide ahead of time what material you won’t touch, because the cost isn’t the design anymore. It’s your brand.

Question for your brand: Now that anyone can generate a flawless image in seconds, what makes yours recognizably yours, and not just one more out of the same mould?


Anthropic Publishes How to Spend Less on It: the Expensive Model Signs, the Cheap One Makes the Copies

Anthropic revealed two patterns for using Fable 5 without paying Fable 5’s full bill.

Sounds counterintuitive coming from a company that lives off selling tokens.

In the advisor pattern, Sonnet 5 executes and only checks in with Fable 5 at decision points: 92% of the performance at 63% of the cost. In the orchestrator pattern, Fable 5 plans and hands work off to sub agents: 96% of the performance at 46% of the cost.

It’s like paying the notary only for the signature, not for running the copies. The signature is what has value. Anyone can run the copies.

The idea underneath is less technical than it sounds. Frontier intelligence isn’t needed for every micro task, only at the forks in the road. Everything else, reading, sorting, drafting, retrying, is volume, and volume should run on the cheap option.

For an SMB this applies without writing a line of code. Use your best model to design the plan once, then run that plan ten or a hundred times on the free model. Nearly the same result, at a fraction of the spend.

Question for your finances: Are you paying for frontier intelligence on tasks a basic model would handle just as well, only because you never separated the thinking from the doing?


Tools You Can Use on Monday

  1. Muse Image on WhatsApp and Instagram – Meta’s image generator, free inside Meta AI. Upload a product photo and swap the background, or create ad variations without paying for design. On WhatsApp it only rolled out in select countries so far. Check if it’s landed in yours.
  2. GPT Live – ChatGPT’s new voice mode listens and talks at the same time, no walkie talkie style turns. Makes it viable to use your phone as a live translator with a foreign supplier. Replaces the previous voice mode globally, with free users getting the mini version.
  3. Claude Cowork on Mobile and Web – Long running tasks now run in the cloud. Start a report, close the laptop, check the result from your phone. Over 90% of its use isn’t coding, it’s operations and content. In beta, starting with the Max plan.
  4. Video Remix in Google Photos – Apply cinematic lighting, background swaps, or artistic styles to a phone video in a couple of taps. Raises the quality of your product content without hiring anyone. Requires a Google AI Plus, Pro, or Ultra subscription, and rolled out in select countries only.
  5. Claude Reflect – A dashboard that shows what you use Claude for, what you delegate, and when, and asks which tasks you’d rather keep doing yourself even if the machine is faster. Useful for checking whether your team is extracting real value or just consuming. In beta for Free, Pro, and Max plans, with memory turned on.

My Invitation This Week: The Notary and the Copyist

Anthropic published the recipe its own engineers use. The expensive model doesn’t execute, it decides. The cheap one does the volume. The savings run from 37% to 54%.

That principle doesn’t require you to code. It requires separating two things almost everyone blends together: thinking and executing.

The exercise takes 45 minutes and two browser tabs.

  1. Step 1. Pick the repeated task. Just one, something you do every week: answering quotes, drafting posts, summarizing meetings, sorting emails.
  2. Step 2. Hire the notary. Open the best model you have access to. Don’t ask it to do the task. Ask it to design how the task gets done. “Design a template and a step by step procedure for answering quotes in my business: what information to ask for, in what order, what tone to use, what to check before sending.” You do this consultation once.
  3. Step 3. Call in the copyist. Open a free model. Paste in the template and run it against three real cases from your week. No extra explanations. The template should be enough on its own.
  4. Step 4. Compare. Run one of those same cases directly on the expensive model, no template, and put the two results side by side. Does the quality difference justify the difference in price and usage limits?

The answer usually stings a little. The notary’s template, run by the copyist, comes remarkably close.

That’s the lesson this week. You don’t need superintelligence for every move of the day. You need it at the forks in the road, and discipline everywhere else.

Your Monday task: what percentage of your week are you paying notary rates for copyist work?


Closing

This wasn’t the week of a spectacular launch. It was the week the industry popped the hood.

Inside were two things at once: prices in free fall and a measuring stick that was broken. Intelligence never cost this little, and it was never this hard to verify what it’s actually made of.

That leaves standing the oldest discipline there is. Test things in your own shop, with your own inventory.

The machines already think in rooms we can’t see, and compete on rules that break. What’s still ours is looking at a result and saying, with our own judgment, this works and this doesn’t.

Start by separating the notary from the copyist.

Recent

Discover the Related Blog Posts

"Artificial intelligence just split in two: the best got locked up by government orders, and the useful got unleashed for...
Strategic analysis of the current bifurcation of Artificial Intelligence in the market. It explores the business impact of government regulations...
"AI is no longer commercial software. This week, the US government officially regulated it as a national defense weapon." AI...

Discover more from AI Consulting Toronto | Practical AI Implementation | Adivor

Subscribe now to keep reading and get access to the full archive.

Continue reading