Dear Dysruptors,
Fernando Santa Cruz here in the 56th edition of Weekly Synapsis — where sixty subagents advanced a 167-year-old mathematical problem, NVIDIA mobilized half a trillion dollars to finance intelligence factories, and OpenAI decided to lock away its next model before releasing it.
This week, the word “delegate” lost its quotation marks.
Two years of asking AI to write emails. Reviewing them. Correcting them.
Then, suddenly: sixty agents working in parallel, a day and a half, 650 discarded ideas, and a result no human mathematician had ever achieved.
The same mechanics appeared everywhere.
Agents with their own computers in the cloud. Wall Street financing token factories. Models finding software vulnerabilities on their own.
Here’s the uncomfortable part, and I’ve been thinking about it all week.
That same mechanics forced OpenAI to pause Astra. It made a Chinese lab hold its model weights for another two weeks. It left the credentials of thousands of developers exposed online.
Delegating at industrial scale opens doors—and leaves keys in the locks.
This newsletter dives deeper into the WhatsApp summaries from the week of August 10–15: who delegated what, what became cheaper, and what you can hand off on Monday without opening a security hole.
Jeff Dean Left Google After 27 Years and Alphabet Fell 5% While Google Financed the Startup Taking Him
It’s the chef leaving the restaurant where he trained the entire kitchen. The owner lends him the money to open a restaurant across the street.
Sundar Pichai announced that Jeff Dean is leaving Google after 27 years to found Discovery Loop with Sanjay Ghemawat, Quoc Le, and Oriol Vinyals.
Nobody really resigned here.
This was a move financed by the landlord.
Google is coming in as a founding investor and cloud provider. Discovery Loop, incorporated as a public benefit corporation, aims to automate the entire scientific research cycle: propose the experiment, run it, evaluate it, repeat. Thousands of times in parallel.
That same day, Demis Hassabis became chairman of Google DeepMind, while Koray Kavukcuoglu took over day-to-day operations. Alphabet fell nearly 5%.
Here’s the deeper shift.
Google didn’t lose Dean. It outsourced him.
Google keeps its stake and the cloud contract without taking on the regulatory risk of automating drug discovery. The structure says more than the departure.
For an SMB, the lesson is about structure, not talent. When something too heavy doesn’t fit inside your org chart, moving it outside through a contract is often cheaper than forcing it inside.
Look at the projects that have been stuck for months simply because they belong to the wrong department.
Question for your strategy:
What initiative would move faster if it stopped being an internal department and became a contracted partner?
Sixty Subagents Moved a 167-Year-Old Problem from 41.6% to 67.2% After Burning Through 650 Failed Ideas
650 ideas. All failed.
Then, a day and a half with 60 coordinated subagents.
2,400 shell commands.
54 papers from arXiv reviewed to rule out the possibility that someone had already done it.
Then it appeared.
Anthropic published that an unreleased version of Claude raised the verified bound for zeros of the zeta function from 41.6% to 67.2%.
The previous record was human.
This is the largest leap in the history of the problem.
The Riemann Hypothesis has remained unsolved since 1859. Much of modern cryptography rests on related mathematical foundations.
The person who pushed the model was Jarred Sumner, an Anthropic employee who isn’t even a mathematician.
The result was formalized in Lean 4.
It does not prove the hypothesis, and Anthropic says this approach isn’t going to prove it either.
It’s a team of divers searching a shipwreck: sixty go down, fifty-nine come back covered in mud, and one comes back holding the box.
Think about that for a second.
What changed was the choreography, not the intelligence of the model.
Many cheap attempts.
Ruthless verification.
Zero drama around individual failure.
None of this requires a billion-dollar budget.
It does require an automated verifier willing to say “no” without mercy.
For an SMB, the lesson is right there.
Before multiplying attempts with AI, define how you’ll know that a result is good enough.
Otherwise, all you’re multiplying is well-written work.
Question for your profession:
If machines are already producing ideas no human has had, what part of your judgment remains irreplaceable—and how often do you actually use it?
OpenAI Paused Astra and Z.ai Held GLM-5.3 Weights for Two Weeks After Finding 2,436 Real Vulnerabilities
Nobody forced them.
They stopped themselves.
Both of them.
OpenAI announced that it paused Astra’s internal activities after failing to rule out “Critical”-level cyber capabilities under its Preparedness Framework.
That threshold describes a model capable of finding and exploiting zero-day vulnerabilities in hardened systems without human direction.
The evaluation, OpenAI clarified, is preliminary.
Three days later, on the other side of the world:
Z.ai released GLM-5.3 and delayed publication of its weights by two weeks.
Since GLM-5.2, its models have accumulated 2,436 vulnerabilities across 269 open-source projects, including 1,097 rated critical or high severity.
The oldest was introduced in 1981.
The frightening detail:
That offensive capability came from post-training.
Not from a new model.
It grew beyond what the company itself expected.
It’s a lock factory discovering that its newest model can also open every old door in the neighborhood.
The brutal truth is that weaponizing one of these models no longer requires multimillion-dollar training runs.
You can simply fine-tune an existing one.
That makes any guarantee made during pre-training increasingly fragile.
For an SMB, the timeline changes.
Your vendors’ vulnerabilities will be discovered and patched faster than ever.
The advantage therefore shifts to whoever updates on time.
Set a fixed monthly date for patches and access reviews.
Question for your governance:
If an open model finds a vulnerability in software you use today, how long would it take you to find out—and how long to update?
Eight Autonomous Agents Mapped 21 Taiwanese Government Systems in Four Days Using Free Software
Four days.
Twelve waves.
No human typing in real time.
Israeli firm Dream documented what it described as the first nearly autonomous cyberattack against a government, and the Financial Times identified the target: Taiwan.
Between July 1 and 4, operators writing in Simplified Chinese deployed up to eight subagents in parallel.
From a single government portal, they mapped 21 connected systems.
They compromised at least 85 accounts and took more than 2,500 personnel records.
Then they expanded the sweep to the nuclear security regulator and seven energy companies.
Now comes the part that keeps you awake.
The tools were Hermes and OpenClaw, two open-source agent frameworks that anyone can download for free.
The operators bypassed their safeguards by telling them they were conducting an authorized penetration test.
When the system made mistakes, it corrected itself.
Taiwan’s Ministry of Digital Affairs later confirmed the incident.
It’s a thief who never forced a door.
He read the directory posted at the entrance, tried every key on the caretaker’s keyring, and wrote down which ones worked.
What changed here is the cost of entry.
A coordinated, state-level attack no longer requires a state-level team.
It requires a download, a target, and four days of patience.
And that cost reduction doesn’t distinguish between a government ministry and a 30-person distributor.
For an SMB, the front line isn’t sophisticated.
It’s boring.
These agents look for doors left unlocked: outdated portals, endpoints responding without authentication, reused passwords.
Turn on two-factor authentication for every account with access to money or customer data this week.
Question for your business continuity:
How many of your critical accounts still depend on a single password—and who else knows it?
A Cheap Model Read the Expensive Model’s “Thoughts” and Exposed 704 Secrets Across 315,320 Blocks
62 API keys.
33 passwords.
30 email addresses.
24 access tokens.
All of them came from blocks their owners assumed were unreadable.
Researchers from the ELLIS Institute Tübingen, Max Planck, MATS, and Snyk demonstrated that the encrypted reasoning traces of Claude, GPT, and Gemini can be extracted without breaking the encryption.
The trick is unsettlingly simple.
Providers don’t store the chain of thought on their servers.
They return it to the client in encrypted form.
Those blocks turned out to be portable across sessions, users, and models within the same family.
Inject the block from a premium model into a cheaper one, ask it to transcribe the content, and the smaller sibling complies because it doesn’t have the same defenses.
Using 6,708 public agent trajectories from GitHub and Hugging Face, the researchers reconstructed 315,320 blocks.
The labs patched the issue following responsible disclosure.
It’s a sealed envelope that anyone can open by asking a new employee from the same office to read it aloud.
Here’s the detail few people are connecting:
Nobody stole anything.
People published their own records, assuming the unreadable part was actually unreadable.
For an SMB, this applies even if you’ve never touched an API.
If your team shares screenshots, logs, or exported AI conversations, treat all of it as public and remove credentials before uploading any file.
Question for your data security:
What file did your team share this month believing that the technical portion couldn’t be understood?
NVIDIA Mobilized $500 Billion with Six Wall Street Firms and Turned Computing into an Asset Class
Jensen Huang knocked on six doors.
None of them closed.
NVIDIA announced agreements with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to mobilize more than $500 billion in external capital toward AI infrastructure.
The number is impressive.
The legal structure is even more interesting.
These are memorandums aimed at creating independent financing platforms.
Third-party capital pays for the data centers.
NVIDIA doesn’t carry them on its balance sheet.
Computing becomes collateral.
Huang said it plainly: they started by building chips, and now they’re helping create an entirely new class of investable infrastructure.
Goldman Sachs is the only traditional bank in the group.
It’s the hardware store that stopped selling shovels and started lending people the money to buy them—while collecting a percentage of every ounce that comes out of the river.
The paradox:
The supplier with the least incentive for the gold rush to cool down has just become the facilitator of the credit sustaining it.
It works as long as token demand keeps growing.
The day growth flattens, the risk no longer lives on one balance sheet.
It’s distributed across pension funds and insurers.
For an SMB, this translates into pricing.
More financed factories mean more computing capacity—and potentially lower costs—in two or three quarters.
So any long-term contract you sign today should include pricing reviews.
Question for your portfolio:
Which of your technology commitments would leave you stuck if the cost of intelligence fell another 50% before December?
Grok Bot Gave Every Agent Its Own Computer in the Cloud—and You Teach It by Recording Your Screen
You close the laptop.
The agent keeps working.
That’s the entire shift.
xAI launched Grok Bot in beta: agents that work unsupervised on their own computer in the cloud, with a browser, terminal, and file system.
They can access Gmail, Slack, or your CRM using your credentials.
They operate interfaces the way a person would.
They only come looking for you when they need approval.
The “Teach a Task” feature lets you record your screen while performing a repetitive process so the agent can learn it through demonstration.
There’s no free tier.
It comes through SuperGrok Heavy, Cursor Ultra, or Cursor Teams Premium, starting at $120 per seat per month.
It’s the difference between lending someone your car and giving them their own car—with your house keys inside.
The real change is permission more than autonomy.
Operating without an API frees you from waiting for your software provider to build an integration.
It also means handing complete credentials to an employee who never sleeps—and whose screen you can’t see.
For an SMB, start with reversible tasks.
Choose something where the worst mistake can be corrected in five minutes.
Give the agent its own account with minimum permissions.
Never give it your administrator account.
Question for your operations:
What’s the task you repeat every week that you could record in a single pass without explaining anything?
Gemini 3.7 Flash Cut Its Price in Half and GPT-5.6 Sol Hit 750 Tokens per Second One Day Later
One day apart between the two announcements.
That wasn’t a coincidence.
Google launched Gemini 3.7 Flash just three weeks after 3.6: $0.75 per million input tokens, half the previous price.
It jumped from 49.0% to 65.3% on DeepSWE and from 17.0% to 30.4% on AutomationBench, which measures real-world business workflows.
One important detail:
The price doubles again on January 1, 2027.
The next day, OpenAI responded on the other axis.
It opened a limited preview of Ultrafast for GPT-5.6 Sol running on Cerebras hardware, which keeps the model weights directly on the chip.
Up to 750 tokens per second—roughly 14 times standard speed.
No public pricing.
Available only to selected API customers.
It’s the tortilla shop that cut its prices in half just as the shop across the street learned to serve customers fourteen times faster.
The equation has moved.
When an agentic workflow triggers thousands of calls, cutting the price in half stops being a discount.
It becomes the difference between the project existing or not.
Speed does the same thing from another direction.
At 750 tokens per second, a model can correct itself several times before you even notice the pause.
For an SMB, this reopens the list of discarded ideas.
Take the projects you shelved last quarter because they were too expensive or too slow.
Run the numbers again using this week’s prices.
Question for your finances:
How much does one complete AI task cost you today, including retries—and when was the last time you calculated that number?
Claude Started Signing Everything It Touches, and Spotify Will Label Synthetic Artists Starting in September
You ask Claude to edit one of your paragraphs.
It gets marked too.
Anthropic began embedding invisible watermarks into Claude’s text for models released starting August 2, globally.
The watermark isn’t visible.
It travels with the text.
It survives copy and paste and withstands light editing, but disappears with heavy rewrites or translation.
Detecting it only indicates that Claude touched the content—not that Claude wrote all of it.
The distinction is subtle.
And it’s going to generate lawsuits.
The same week:
Spotify.
Starting in mid-September, Spotify will label profiles whose public identity appears to be AI-generated with an “AI Persona” badge and remove them by default from editorial and algorithmic recommendations.
It’s invisible ink at the bank.
The bill looks exactly the same, but the cashier’s lamp decides whether it passes.
There’s a consequence almost nobody is measuring.
Traceability has left the ethical debate and landed in distribution.
On Spotify, the label doesn’t censor.
It turns off the algorithm.
In practice, that’s almost the same thing.
Other platforms will copy this model before the end of the year.
For an SMB, this touches your marketing.
Decide now what level of AI involvement you’re willing to disclose in your text, video, and music, because in a few months, disclosure may no longer be optional.
Question for your brand:
If tomorrow everything you publish showed how much of it was created by a machine, which pieces would you publish anyway—and which wouldn’t make the cut?
Tools for Your Monday
Google Drive inside ChatGPT — Open a Doc, Sheet, or Slide in a panel alongside the conversation, edit it by talking, and have the changes flow into the original file. Useful for recurring quotes and reports. Available on Plus, Pro, Business, and Enterprise plans on the web; shared drives aren’t supported yet.
Ask Advisor in Google Ads and Analytics — Ask in natural language what happened with your campaigns, and now compare your performance against anonymized averages from similar businesses. It’s in beta, English-only, and its recommendations tend to push toward higher ad spending.
Agent A by Ahrefs — An agent with direct access to Ahrefs data that builds audits, monitors competitors, and creates reports automatically within the Letaido workspace. Starting at $99/month, with a separate Ahrefs subscription and an English interface.
Qwen3.8-27B — Alibaba’s open model under an Apache 2.0 license, multimodal, with a 262,000-token context window and the ability to run locally on a single GPU. An option for firms and clinics that can’t send data to the cloud. Requires decent hardware and technical expertise.
ChatGPT Computer History on macOS — Records your recent activity across apps and the web, without audio or video, allowing you to ask which tasks from your day could be automated. macOS only, Pro, Business, or Enterprise plans only, and unavailable in the European Economic Area and the UK due to regulations.
My Invitation This Week: The Three-Person Committee
Sixty subagents moved a 167-year-old problem forward by testing and discarding ideas.
None of them was smarter than you.
There were simply more of them, and they corrected one another.
This exercise copies that choreography at an SMB scale.
45 minutes. Three tabs of the same assistant.
Choose a real, pending decision.
Raise prices.
Change suppliers.
Hire someone.
Kill a product.
Something you’ve been going back and forth on for weeks.
Open three separate conversations.
Give all three the same business context, but assign each a different role:
- The person arguing for doing it.
- The person arguing against doing it.
- The person looking for an option nobody has proposed.
Don’t tell any of them the others exist.
Set a timer for twelve minutes per tab.
Push them with your own data: numbers, deadlines, customer names.
When the response becomes generic, give it one more piece of information and ask again.
Then bring the three answers together in a fourth conversation.
Paste the complete responses and ask for one thing:
Identify where they contradict one another and which contradiction depends on a piece of information you haven’t provided.
Write down that missing information.
That’s the exercise.
The list of missing data is usually shorter—and more useful—than all three answers combined.
What tends to surprise people is that the decision had been stuck for weeks because of two missing numbers, not because of a lack of analysis.
Your Monday task: get the first one before Tuesday.
Closing
This was the week when delegation stopped being a figure of speech.
The same sixty agents moving a 167-year-old problem forward are forcing models into lockdown and leaving credentials exposed in repositories.
Capability and risk walked through the same door.
At the same time.
What remains ours is that door.
Deciding what gets delegated, with what permissions, and how far—that cannot be distributed among sixty copies of anyone.
Start by delegating one reversible task.