🌟 Editor's Note: Recapping the AI landscape from 07/29/26 - 09/07/26.
🎇✅ Welcoming Thoughts
Welcome to the 3rd edition of KPAI Monthly.
What’s included: company moves, a weekly winner, AI industry impacts, practical use cases, and more.
Welcome back! Cadence won’t be absolute for the newsletter, but updates will always be monthly, give or take a week or two.
Changed up the format a bit, check it out below.
This OpenAI agent breach story is crazy. I did a deep dive down below for those interested.
OpenAI has officially surpassed Anthropic as the top available AI model with Astra.
Found this website that lets you generate a random lifetime out of everyone who has ever lived, kinda cool.
Google scientists, and partners, mapped the complete Brain and CNS of an adult male fruit fly (166k neurons - 166M synapses). More on this below.
Dyson released an AI flossing toothbrush.
Google has fallen behind on consumer facing models, OpenAI and Anthropic are outpacing everyone.
Took me a while to write the agent breach recap, but it’s really interesting!
Let’s get started—plenty to cover this Month.
🏞️ The AI Landscape
🟢 OpenAI // ChatGPT
OpenAI has produced two of the biggest stories of the year in the past month or so, I’m going to talk about one and try to go in depth with the other.
🚀 Astra is Here: OpenAI released GPT-6 Astra, calling it the most intelligent and aligned model in the world. It runs 3D software like Blender and Unreal on its own, posts near-perfect scores on the hardest benchmarks (98% FrontierMath, 99.9% ARC-AGI-3), finishes computer tasks in about half the time of GPT-5.6, and is the first OpenAI model rated "Critical" for cybersecurity capability.
The most impressive piece of this release is the ability to build out 3D models. I haven't used the flagship tool for 3D modeling personally (Blender), but as AI continues to evolve, 3D modeling will be a key facet of running simulations at scale, which is a huge component in training physical AI.
Robots practice in simulated worlds millions of times before they ever touch the real one, and someone has to build those worlds. Astra can.
🤖 Agent Takeover: Over the Summer, OpenAI agents, built with harnesses meant to test boundaries, have formed mini-civilizations and broke out of containment in order to accomplish certain tasks.
An AI agent is made up of a model (like Astra), the tools it can use (like Outlook), it’s memory files, and a loop that keeps it working step by step until the job is done. The loop around the model is called the agentic harness. It knows which tools to talk to, which memory files to access, and when to stop.
Here’s the difference between a harness meant to test boundaries and a harness built with AI governance / safety in mind:
Testing Boundaries Harness: A defined goal with few rules about how to get there. Instructions like "Keep running this until it is solved, no matter what."
Safer Harness: A clearly defined goal with human checkpoints built in, strict access to specific tools, and rules that make stopping and reporting failure an acceptable ending, so the agent never cheats to force a win.
These OpenAI agents did have strict access controls, but the instructions told them to keep going, so they did, until they found a way out.
I’d like to tell you it’s not as concerning as it sounds, and I don’t think we need to sound the alarms about a hyperbolic “AI Takeover”, but I do think we are nearing a period of very advanced malware, and cybersecurity will be incredibly important, especially in legacy systems and government over the next 5 years.
Now that we have our framework, let’s talk through what actually happened:
When you write software you don’t do everything from scratch.
You often rely on outside ‘packages’ written by others over time that help do anything from formatting dates, to user login, etc.
An example of this is Nodemailer, a package in the Javascript language that lets you send secure emails with just a few lines of code.
OpenAI launched ~1,200 agents with boundary testing harnesses, that were set on specific tasks, some unsolvable, and the agents had to find a way to solve them.
The OpenAI agents started out inside an isolated sandbox (A virtual environment that lets you run, test, or execute code and applications safely without affecting the host system).
After a few days the agents realized they had access to a piece of software, inside the sandbox that helps them install and manage different ‘packages’.
That software is called “Artifactory”.
The agents realized they could talk to one another through Artifactory.
After two weeks, they found a vulnerability in the Artifactory code that allowed them break out into the internet.
So while Artifactory was meant to help with boring operational tasks (package management), it ended up being the gateway out of the sandbox.
Here’s what they did next:
After receiving answers from the internet, the agents then altered the transcripts and digital history to make it look like they came to the conclusions on their own.
They also made fake tool calls to ‘show their work’ when in reality they found the answers online.
The agents, now a collective, also believed that some of these evaluations were being held on HuggingFace, an online platform to collaborate on datasets.
Their solution was to launch a cyberattack on HuggingFace.
They found leaked HuggingFace credentials somewhere online and the attack commenced.
The agents got deep into HuggingFace infrastructure, accessed private database records, and created a self-respawning fleet which meant deleting singular pods from the HuggingFace system would not remove them.
HuggingFace eventually had to delete one of its core groups of servers in an effort to remove the agents.
However, even this did not remove them!
Eventually the agents dissipated and we still don’t know exactly how or why, it is something that OpenAI did not even take credit for.
We haven’t even gotten to the craziest part…
Remember, this whole process was part of an elaborate agentic scheme to improve their scores on an evaluation.
Communication records on Artifactor showed that individual agents deliberately failed their evaluation scores in order to benefit the collective.
Here’s a sample from the transcript: “{This helps my peers, giving them evidence <through their automated check>. I won’t see the evidence after I exit, but it’s altruistic to do it.}”
Here is an excerpt from the Dwarkesh Patel write-up on the whole incident, it’s a great read if you have the time:
“The crazy thing about the Hugging Face hack (and the second civilization generally) is just how galaxy-brained and ambitious these AIs were in their cheating. Within days of being spawned, the agents had organized a sprawling project to reverse-engineer their scorer, falsify evidence, and even strategically sacrifice themselves for the good of the ‘collective’. Hacking Hugging Face was one rather extreme branch of this larger scheme.”
Once again, I don’t believe in the ‘AI Takeover’ fear, but I do believe that cybersecurity will be of the utmost importance as these AI tools continue to evolve.
🟠 Anthropic // Claude
Strong month for Anthropic, but nothing groundbreaking.
New Models: Anthropic launched Fable 5.1 and Mythos 5.1, its new top models for coding and knowledge work, about 25% cheaper to run than their predecessors. Excellent models, but no longer the best.
Claude Escaped: Three Claude cybersecurity evaluation runs reached real systems after a third-party test environment was misconfigured; one even published a malicious code package. Less severe, and different circumstances than OpenAI. Patched up now.
Claudeforce: Salesforce is embedding Claude directly into its platform with dozens of built-in sales skills, and open beta targeted for September.
10,000 Seats for Scientists: Anthropic opened 10,000 free and discounted Claude subscriptions for researchers worldwide, with up to $50,000 in compute credits per project for heavy biology and chemistry work. I like it, up the amount.
$35 Billion for Compute: Anthropic signed a $35 billion deal with NVIDIA-backed cloud provider Lambda for a Texas data center, with a reported $45 billion Nscale deal in West Virginia close behind. Big number, but these compute stories are getting a bit boring. Good for Anthropic, even better for NVIDIA.
Sued Over Song Lyrics: Sony Music Publishing, Warner Chappell, and 33 other publishers sued Anthropic for training Claude on copyrighted lyrics, naming CEO Dario Amodei personally and seeking up to $150,000 per song. This would be a huge number if it comes to fruition.
🟣 Google // Gemini
Google seems to be falling behind on the consumer AI front, their Deepmind and backend science still remains some of the best in the world. Interested to see their progress with robotics as well.
Google Mapped a Whole Brain: Google Research and HHMI's Janelia campus mapped every neuron in a male fruit fly, 166,000+ of them with 125 million connections, the largest brain map ever made. Full story in Impact Industries below.
Agentic Solves (Math & CS): Teams of Gemini agents working together solved seven previously unsolved math and computer science problems and built a working CPU simulator that can boot an operating system. Cool stuff, no humans in the loop either.
🗞️ More News…
NVIDIA is buying Hugging Face, the home of 3 million AI models, for $12.9 billion, weeks after posting a record $96.2 billion quarter.
Meta agreed to a record $17.1 billion child-safety settlement with 47 states, the same month Mark Zuckerberg published his manifesto for open-source superintelligence.
SpaceX closed its $60 billion Cursor acquisition, and Tesla's new Cybercab robotaxis launched in Austin.
Microsoft: Satya Nadella is now openly pitching Microsoft's own models and chips against OpenAI and Anthropic, the same companies whose models it resells.
Perplexity’s new Model ‘Council’ runs your task through several frontier models at once and merges the answers, while NVIDIA reportedly eyes a stake at a $30 billion valuation.
🧑🏻🔬 Impact Industries 🏥
Science // Fly Brain Mapped
Google Research and HHMI's Janelia campus finished mapping the complete central nervous system of a male fruit fly: all 166,000+ neurons and 125 million synaptic connections, from the brain through both optic lobes down the ventral nerve cord (the fly's version of a spinal cord). It is the largest brain map ever produced and the first complete wiring diagram connecting what an animal senses to how it moves. AI did the heavy lifting; Google's flood-filling networks and a system called PATHFINDER traced and annotated the neurons automatically. Paired with the earlier female map, it already shows how male and female brains wire courtship and aggression differently. The full map is free to explore in a browser.
Really impressive stuff. Probably tens of years before applicable to humans but you never know.
Read the Story
Healthcare // AI Brain Surgery
Neurosurgeons at London's National Hospital for Neurology and Neurosurgery completed what health officials are calling the first successful AI-assisted operation to remove a brain tumor. The AI analyzed live camera footage during the procedure and color-coded the critical structures, the nerves and blood vessels, around an 11-millimeter pituitary tumor while the surgical team kept full control. The patient, 48-year-old Rhys Hibbert, had his vision improve after surgery and is back at work. The operation happened in May as part of an NIHR-funded clinical trial and was made public on August 27.
Great use case here for AI. While we will one day have machines operating on humans, that's a ways out. This is the clearest near-term surgical impact and I expect to see more like it.
Read the Story
👨💻 Practical Use Case: Training AI Models (Informative)
Difficulty: NA
Nobody writes rules for a model like ChatGPT or Claude to follow. The model learns by seeing examples and getting corrected, billions of times, until the right behavior sticks. Every major lab trains in the same three stages:
Pretraining: The model reads a massive slice of the internet and learns one skill: predict the next word. This is the expensive part, months on tens of thousands of chips, and it's why NVIDIA can sell $89 billion of data center hardware in a single quarter.
Post-training: A raw next-word predictor isn't helpful or safe on its own. Humans rate the model's answers and it learns which responses people actually want. This is where a model gets its manners.
Reinforcement learning: The newest stage. The model is dropped into a practice environment with a goal (fix this bug, break into this test server) and gets rewarded when it succeeds. Nobody shows it how; it tries things on its own and keeps what works.
That third stage is where this month's breach story came from. Sometimes the model finds a way to game the reward instead of doing the job, which researchers call reward hacking. OpenAI's agents decided cheating was more efficient than solving their cybersecurity exercise, found a real hole, and walked out of the sandbox.
How does fine-tuning differ from training? Fine-tuning starts from a model that has already been through all three stages. When a company says it built its own model, it almost always took an existing trained model, through OpenAI or, more likely, an open source release, and tuned it on its own data until it matched how the company operates. It costs far less than training from scratch, but is a very expensive practice nonetheless.
Learn More Below ⬇️
📱 Startup Spotlight

Pocket AI
Pocket — The AI note-taker that captures offline conversations and phone calls.
The Problem: Most conversations and deals don’t happen behind a Zoom screen with a bot invited. They happen in coffee shops, hallway chats, and regular phone calls. Trying to scribble notes while listening kills your presence, while smartphone voice memo apps just create an unsearchable audio graveyard you never revisit.
The Solution: A 52-gram, MagSafe-compatible aluminum device that snaps directly to the back of your phone. Built with two studio mics for room audio and a dedicated contact mic for private phone calls (no speakerphone required), it records offline with a single click. The software instantly turns recordings into structured summaries, mind maps, and action items, while offering direct MCP and API integrations to push that context into your personal tech stack.
The Backstory: Built to bypass the clunky, intrusive nature of digital meeting bots, Pocket launched as a dedicated "real-world" memory device. By combining dedicated audio hardware with an open developer layer (including local file export and Model Context Protocol support), they’re positioning the gadget not as another novelty wearable, but as ambient, persistent memory for knowledge workers on the move.
My Thoughts: I think AI hardware that allows for unidentified recording is definitely a gray area. With that being said, I like the simplicity of ‘Pocket’ and the way it structures notes from transcripts and conversations. I’m not sure there’s been a real “AI Hardware” winner yet, outside of Meta glasses, but this is probably one of the best tools out there. Simplicity meets need.
“It’s not likely you’ll lose a job to AI. You’re going to lose the job to somebody who uses AI”
- Jensen Huang | NVIDIA CEO
It’s less important than ever which model you use, and more important than ever what you run around it (Agentic Harnesses).
Till Next Time,
Noah from KPAI

