The Most Important Takeaways from This Edition
- AI-verified mathematics: OpenAI solved the Navier-Stokes problem (90 years without a solution) with 10,000 agents in 88 hours; Anthropic formalized Fermat’s Last Theorem with Claude in 11 days.
- Security alarm: researcher Jacob Coxon resigned from Anthropic, accusing the company of an irresponsible race toward superintelligence; alignment chief Evan Hubinger publicly backed an extinction risk estimate of more than 10% over the next decade.
- Regulation: California signed the country’s first laws (SB 813 and AB 1405) requiring independent audits of AI systems on September 9.
- Capital: Mistral AI raised €3 billion ($3.5 billion), reaching a $21 billion valuation—the largest technology funding round in Europe, led by Samsung.
- AI costs: DeepSeek launched V4.1-Flash, reducing the required cache memory by 75% and retiring its own flagship model because it was more expensive.
- Personal agents: Meta launched Muse, an agent that books, pays, and manages tasks through WhatsApp, and acquired startup Stilla.
Dear Dysruptors,
Fernando Santa Cruz here in the 59th edition of Synapsis Weekly — where a swarm of agents solved a problem that had been open for 90 years, a researcher quit saying we are moving too fast, and California signed the first law requiring someone from outside the system to review the work.
All three things happened in five days.
Almost nobody connected them.
This newsletter is going out on Monday this week because I spent the weekend in Campeche at the Digital Trends for SMBs 2026 event, and I didn’t want to let the week go by without sharing this.
The machine produced new knowledge. Good knowledge, verified and public.
13 million lines of code in one case. 130 billion tokens in the other.
No human can read all of that. Ever.
Think about that for a second.
For 300 years, the seal of quality in mathematics was that another person could understand the argument. This week, that seal stopped being enough.
Not because the argument is wrong. Because it no longer fits inside one person’s head.
This newsletter goes deeper into the WhatsApp summaries from September 7–13: who verifies, with what, and what happens when nobody can.
10,000 Agents Took Down a 90-Year-Old Problem in 88 Hours—and Another Wrote 13 Million Lines
88 hours. Around 10,000 agents running in parallel. 2.7 million messages between them.
130 billion output tokens for a single problem.
On September 8, OpenAI published a solution to the Navier-Stokes problem, one of the seven Millennium Prize Problems, with an analytical proof and formalization in Lean.
It’s worth reading the fine print. They did not solve fluid flow in general.
They constructed a counterexample: a fluid that starts at rest, receives a smooth force, and develops a singularity in finite time.
The model that did it is not for sale. It is an internal system that OpenAI describes as more capable than GPT-6 Astra.
The Lean verification took another 17 hours.
Four days earlier, Anthropic had announced that Claude formalized Fermat’s Last Theorem in 11 days with almost no supervision: 13 million lines of Lean and 29,500 intermediate theorems in the final proof.
It’s a building erected over a weekend and delivered with 13 million blueprints. The inspector is still outside, tape measure in hand, deciding where to start.
Here’s the deeper shift. The bottleneck in science is no longer having the idea. It is verifying it at the speed at which it is produced. Lean stopped being a technical detail and became the only thing holding trust together.
For an SMB, the translation is straightforward. Your AI already produces more than you can review, and the gap grows every month. Pick one deliverable you currently approve at a glance and write the test that declares it acceptable.
Question for your craft: What work are you signing off on this week without verifying it, simply because it turned out well the last few times?
An Anthropic Researcher Quit on Tuesday—and the Head of Alignment Agreed with a Number
He quit on a Tuesday. He posted a thread. His colleague, who is still inside the company, came out and said he was right.
Jacob Coxon, who spent three years conducting pretraining research between OpenAI and Anthropic, left the company accusing both of moving irresponsibly toward self-improving superintelligence.
The unusual part came afterward.
Evan Hubinger, who leads alignment work at Anthropic, publicly responded that Coxon is right and put forward his own number: more than a 10% probability that AI could end humanity within the next decade.
Two clarifications got lost in the noise. It is a personal estimate, not Anthropic’s official position.
Hubinger was explicit that today’s deployed models represent comparatively low risk. His concern is recursive self-improvement, which does not yet exist.
It’s the pilot getting off the plane at the gate and saying it out loud in front of the line. The plane takes off anyway. The difference is that now everyone heard him.
Brutal truth: nobody is lying in this story. The people building these systems believe what they are saying, and they keep building them because they are convinced that slowing down only hands the lead to someone else. It’s not denial. It’s a prisoner’s dilemma with a salary.
For an SMB, the lesson isn’t the apocalypse. It’s the asymmetry. The people who know a tool best are the ones who see its limitations most clearly. Ask the person operating it on your team, not the person who sold it to you.
Question for your leadership: When someone on your team tells you out loud that something is moving too fast, what does the way you react tell you about yourself?
California Signed the Country’s First Mandatory AI Audits as Congress Prepared to Leave Until November
Wednesday: Sacramento signs.
Friday: four federal lawmakers ask them not to go on vacation.
On September 9, Gavin Newsom signed SB 813 and AB 1405, the first laws in the country to establish a framework for independent verification organizations and a state registry of AI auditors.
The sentence that sums it all up came from Assemblymember Bauer-Kahan: we cannot wait for the industry to grade its own homework.
Anthropic and OpenAI supported both laws. Newsom used the signing to call on the federal government to regulate AI.
Meanwhile, the House returned Monday for a single week of session before leaving until November 9. A letter led by Sam Liccardo called for canceling the recess and legislating on AI. The agenda included a vote on the electricity costs of data centers. On AI, nothing.
It’s a building code conducting its first inspection when the building is already 40 stories tall and still going up.
There’s a date in the fine print that changes the entire reading. The state agency has until January 2028 to define the criteria, and the auditor registry must exist by January 2029. The law passed this week; it bites in two or three years.
For a Mexican SMB, this is not foreign news. California sets the standard that your software providers will eventually adopt globally, just as happened with privacy. Ask your AI provider today what evidence of external evaluation they can provide you in writing.
Question for your governance: If a client asked you tomorrow for documentation showing how you review your automated processes, would you have anything to show them?
Mistral Raised €3 Billion for European Sovereignty—with a Korean Company Leading the Round
It’s a house with its own deed, a Korean mortgage, and all its wiring bought in California. It’s still yours. You just need to know who you’re paying every month.
On September 8, Mistral announced that it raised €3 billion in a Series D at a valuation above $21 billion, the largest capital raise ever for a European technology company. The company is just three years old.
Samsung Electronics led the round. The Scaleup Europe Fund, managed by EQT, and PSG Equity joined as co-leads.
New investors include Advent, BlackRock funds, and the Grand Duchy of Luxembourg. Existing investors include NVIDIA, a16z, Lightspeed, and Salesforce Ventures.
Mistral operates in 20 countries and serves more than 125 companies, including Airbus, ASML, and HSBC. It has said it wants up to one gigawatt of its own computing capacity in Europe by 2030.
The paradox: the company selling European technological independence is financed with Asian and American capital, as well as by the chipmaker Europe wants to depend on less. Real sovereignty is not autarky. It is having an alternative. Nobody builds alone.
For an SMB, Mistral’s sales pitch is what matters, not the funding amount. Corporate customers aren’t buying a better model. They’re buying the ability to move that model between clouds, countries, or providers. Before signing your next annual contract, ask what happens to your data and workflows the day you want to leave.
Question for your portfolio: Which of your software providers has you without a realistic way out today, and how much is that dependency worth on your annual bill?
DeepSeek Shut Down Its Flagship Model This Monday After Releasing One That Needs One-Quarter of the Memory
Today, at 04:00 UTC, DeepSeek began redirecting all requests from its flagship model to the cheaper one.
Its own flagship, retired by a model that costs less.
On September 10, the company introduced DeepSeek-V4.1-Flash, a 552-billion-parameter model with a new architecture that activates only 8 billion parameters when reading and 16 billion when writing.
The number that matters isn’t the size. It’s the memory.
Its cache requires one-quarter of the HBM memory and one-eighth of the storage required by the previous generation. The weights were released openly under an MIT license.
Outside peak hours, one million input tokens cost $0.15 without caching and $0.003 with caching. Output costs $0.60. During peak hours, prices double.
They didn’t buy a bigger truck. They learned how to fold the clothes, and suddenly everything fit in the same closet.
The detail few people are connecting: an agent rereads its context dozens of times per task, so the real bill isn’t generated by thinking. It’s generated by remembering. Compress the memory, compress the bill. That’s why DeepSeek can retire its expensive model without losing anything.
For an SMB, this changes a concrete decision this quarter. If you postponed an automation because the cost per task didn’t make sense, the numbers you used are already outdated. Run the calculation again before you finalize this year’s budget.
Question for your finances: What automation project did you shelve because it was too expensive, when the cost of compute has fallen twice since then?
Meta Released an Agent That Pays with Your Card Through WhatsApp—and Bought the Startup It Was Missing
You give it your email. You give it your calendar. You give it your card.
On September 8, Meta introduced Muse, its personal agent, which doesn’t answer questions—it executes tasks: booking, filling out forms, drafting, and paying.
It runs inside Muse Secure VM, a dedicated virtual machine in the cloud where the agent and the credentials for everything you connect live. A second agent, Sentinel, monitors it from the same machine, isolated at the system level.
It is available through the app, the web, and WhatsApp. It is currently available only in the United States and only to adults.
One detail circulated incorrectly all week. Muse was not the #2 most-downloaded app in the United States. It was #2 within the Productivity category.
It’s the assistant you gave a copy of your keys and its own room inside your house, with a security guard in the hallway whom you never hired or interviewed.
It’s worth looking at where the lock is, not who installed it. Meta’s bet isn’t a smarter model. It’s packaging complex infrastructure inside a WhatsApp conversation. Whoever wins the habit wins the category, and Meta already has the habit.
For an SMB, what matters isn’t Muse. It’s what Meta bought last week: Swedish startup Stilla, to strengthen its Business Agent, which according to the company is already used by more than 1 million businesses on WhatsApp and Messenger. This week, review what your WhatsApp Business currently does without anyone attending to it.
Question for your operations: If an agent answered your sales WhatsApp tomorrow, what three things would you prohibit it from saying or promising?
Anthropic Modeled 17.9% Office Unemployment by 2030 While Construction Wages Rise 33%
It’s not a forecast. Anthropic wrote it twice so nobody would confuse it for one.
Its economics team published three scenarios for the 2030 economy, without assigning any probability to them.
In the modest scenario, AI looks like the internet: GDP ends up 1.6% higher and employment barely moves. In the substantial scenario, growth doubles its normal pace.
In the extreme scenario, GDP reaches $44.4 trillion, 32.4% higher, growing at 15% annually. The economy would double every 4.5 years.
That same scenario brings 17.9% unemployment among knowledge workers and 11.9% across the entire workforce.
Anthropic surveyed 10,980 Americans in August. The typical response points toward the middle scenario, not the extreme one.
It’s a construction site that speeds up because the architect draws faster. The person laying the bricks ends up earning more. The person who used to draw them doesn’t.
What’s interesting isn’t the unemployment. It’s the distribution. In that scenario, office wages fall 11.5% while wages for everyone else rise 33.6%, and the share of national income going to labor drops from 60% to 45.2%. The model’s proposed driver is concrete: faster permitting means more projects.
For an SMB in construction, services, or manufacturing, this flips the usual conversation. Your risk isn’t that AI replaces you. It’s that you won’t have enough people when demand rises. Start documenting what your field personnel know in writing before the market competes for them.
Question for your team: What knowledge about your operation currently lives only in the head of someone who could leave six months from now?
OpenAI Connected ChatGPT to Company Databases—and Let It Show You What’s Missing First
You ask why sales fell this month. Someone has to run the report. They get back to you Thursday.
On September 10, OpenAI launched the Data agent inside ChatGPT Work, which investigates what changed in your metrics and builds interactive dashboards you can share.
It connects to Amazon Redshift, Google BigQuery, Snowflake, Databricks, MongoDB, ClickHouse, and Datadog, and pulls files from Google Drive and SharePoint.
It respects existing permissions, down to the row and column level. The administrator decides who can use it.
Nobody writes a query. You ask in Spanish and refine the question in the same conversation.
It’s the simultaneous translator hired for a meeting where nobody has yet agreed on exactly what “closed sale” means.
The uncomfortable part comes before the software. The tool interprets data using your company’s metric definitions, drawn from semantic layers that someone had to write. If your numbers live across four Excel files with different criteria, this agent will give you fast and contradictory answers.
For an SMB, the step before the software costs more than the license and is worth more than the license. Choose your five headline metrics and write the exact definition of each one on a single sheet, including who owns the data. Without that, no AI sitting on top of your data will work.
Question for your data: If you asked three people in your company how many active customers you have, would they give you the same number?
Tools for Your Monday
- ChatGPT Images 2.5 — New image engine with more precise editing and the Sketch feature, which turns a hand-drawn sketch inside the chat into the final image. Useful for product mockups, flyers, and catalogs without a designer. Available across all ChatGPT plans, ChatGPT Work, and Codex on desktop, mobile, and web.
- Gemini for Windows — Desktop app that opens with Alt + Space over whatever you’re using and pulls context from Gmail and Drive. Useful for drafting and verifying without switching windows. The download is free and global for Windows 10 and 11, but the Gemini Spark agent and Gemini Omni video require a Google AI subscription, are available only to users over 18, and vary by country.
- Suno v6 — Suno’s first generation of music models developed using licensed catalogs from Warner Music, BMG, and Believe. Useful for radio spots, bumpers, and social-media sonic branding. Watch the fine print: the free v6-mini model does not allow downloads or commercial use; those require paid plans, and Universal and Sony’s lawsuits against the company remain ongoing.
- DaVinci Resolve 21.1 — The video editor can now connect to assistants such as Claude or Codex to organize footage, assemble highlight cuts, and send renders using natural-language instructions. The base editor remains free, but assistant integration lives in the Studio version, priced at $295 for a perpetual license.
- Data agent in ChatGPT Work — Query your company’s databases in Spanish and receive shareable dashboards. It only works inside ChatGPT Work, is installed by an administrator from the plugins directory, and requires your data to already live in a supported connected warehouse. If your operation runs on scattered spreadsheets, this isn’t for you yet.
My Invitation This Week: The Cost of Verification
This week, two labs delivered new mathematics, and in both cases the serious question wasn’t whether the machine could do it. It was who could review it, with what, and how quickly.
You have exactly the same problem, just in miniature. Let’s measure it. Forty minutes and a stopwatch.
- Pull three real deliverables. Three things AI produced for you last week and that you approved: a quote, a customer email, a meeting summary, a calculation.
- Time the verification. Take each one and check it against the source of truth, not your memory. Write down the actual minutes it took, not the number you think it took.
- Calculate the ratio. Divide verification minutes by the minutes AI saved you by generating it. If the number is greater than 1, that task isn’t saving you anything.
- Sort them into three boxes. Cheap to verify. Expensive to verify. Impossible to verify with what you have today.
- Write one rule. Anything that lands in the third box does not go out with your name on it until there is a way to verify it.
What you’ll almost always discover is surprising: the tasks you’re most proud of having automated are often the most expensive to review, while the boring ones turned out to be the truly profitable ones.
Your Monday task: Take the task from the “expensive to verify” box and design a one-line test that would make it cheap.
Closing
It was the week when the machine produced knowledge no human can read in full, and several of the people building it publicly asked someone to review them.
A 90-year-old problem. 13 million lines of code. A resignation. A signature in Sacramento.
Beneath it all is a single question, and it isn’t whether AI can.
It’s who verifies, with what evidence, and how quickly.
The good news is that we still write that question ourselves. Lean didn’t decide what counted as a valid proof: mathematicians did, years ago, by writing rules before they needed them.
That is still the job. Not producing faster than the machine, but defining before it does what we are willing to accept as true. It is an act of judgment, and judgment is still built among people who correct one another.
Start by choosing one task and writing down how you will know it was done correctly.
Fernando Santa Cruz
Head of AI & Automation @ Adivor Consulting
Verification is the new way of thinking slowly.