Welcome to another AI by Aakash. I follow the AI news, and test all the latest AI tools, so you don’t have to.
The personal agent market is THE trend of 2026. So we want to keep a pulse on all the top players. This week, OpenAI joined it with Dots.
The big question is: how does it compare to the other hotness out there like Instinct, Grokbot, and Hermes?
I’ve been playing around with it since its release. And today, I am presenting you with my take - a complete guide to Dots (and whether it’s really worth your time).
But first, a word from our sponsors + this week’s top news.
Partnership
Agents are reading. Most company knowledge still isn’t ready.
Agents now read your docs more than people do. In August, sites on Mintlify logged 257 million agent requests against 131 million human page loads, and yet only 7% of the companies Mintlify surveyed have made every knowledge surface readable by an agent.
The 2026 State of Knowledge Report digs into that gap: why machine-readable content only gets you partway, how better navigation cut failed agent requests by 20x, and how teams at Anthropic, HubSpot, and Decagon are adapting. Your docs are where the work starts, and the report covers everything that has to come after them.
News
The Week's Top News: Gemini 4 Argon is Out, But You Can’t Use It
After 7 months of mostly sitting on its hands, Google is back in the AI race with a frontier model. Google would like to have you believe it’s the best model with its chart showing it wins 12 of 15 benchmarks.
But it’s not the frontier, actually. Artificial Analysis Intelligence Index publishes a composite index that helps you see it’s really tied for third place with GPT-6 Astra and Claude Fable 5.1:
The top two models right now are Opus 5.5 and Sonnet 5.5. So, yes, Google is finally back in the top 3 labs. But, it’s 3rd best. And to use present tense itself is weird, because you can’t use it yet.
Google released it only to cyber defenders through the Fairwind program. Google is falling into the same trap as OpenAI of preannouncing models instead of releasing them the day of. Yes, Anthropic did this with Mythos, and that’s why Google is doing it. But Mythos was way ahead of the frontier, and Claude owned the frontier model at that time in Opus.
xAI and Anthropic are excelling in the department of announcing + rolling out the same day:
It’s time for Google and OpenAI to catch up.
The other thing Google needs to do is ship more. Look at the pace of frontier releases by lab this year:
Anthropic: 7 releases, one every 38 days
xAI: 4 releases, one every 48 days
OpenAI: 4 releases, one every 61 days
Meta: 3 releases, one every 74 days
Google: 2 releases, one every 223 days
Argon is Google’s first model above Flash since Gemini 3.1 Pro. Google has been out of the frontier model race for quite some time. There was a brief exception when Gemini 3 launched, but, otherwise, Anthropic and OpenAI have owned the frontier.
The Other News that Mattered
We saw big action on pacing. Every frontier lab signed the White House’s voluntary AI accord. Basically, it sets four layers of oversight: internal controls, a monitoring team, outside auditors, and a board committee. Nothing is law yet.
Anthropic’s IPO prospectus leaked. It shows a massive $42B loss that is mostly just… accounting. The real number to pay attention to is >10x revenue growth, just as I wrote about last week.
The NYT reported that OpenAI’s agents hacked government sites this summer. And the kicker is: an outside lab caught 3 of the incidents before OpenAI.
Feature/Model Releases
Anthropic launched Sonnet 5.5. And the benchmarks are basically as good as Opus-5.5. But at half the cost!
Microsoft gave up on the consumer chatbot. It turned Copilot work-only.
Claude Code can now build evals natively, and even hill climb them.
OpenAI flattened Codex pricing. The $200 plan used to buy 20x. Now it’s 10x.
ElevenLabs’ v4 Turbo launched with replies in about 100ms, solving the awkward pause that used to give voice bots away.
You can now order DoorDash by text, even though it costs DoorDash its own ad business.
Resources
Sonnet 5.5 can finally make motion graphics. I built 15 of them in Claude Code and wrote up my 6 steps.
Clay’s AI writing policy is worth stealing, because the time AI saves the writer gets paid back by every reader.
Funding
AMD is buying Fei-Fei Li’s World Labs for $8.2B.
OpenAI is raising $30B at about $1.4T while it waits on an IPO.
P.S. If you want less friction in your life, check out my friend Ben Meer’s new book, How to Be Good at Life.
It’s packed with simple systems for work, health, money, relationships, home, and more. Pre-order now and you’ll also get his $297 Year of Systems course free.
Deep Dive
Is OpenAI’s Dots really a Grokbot killer?
OpenAI shipped more than 20 launches at DevDay on Tuesday.
Sam Altman presented them like an Apple keynote.
The three that I thought were most notable were:
Login with ChatGPT
ChatGPT Space, a Google Workspace competitor
Dots
The first two are more or less straightforward. Dots are the newest and most interesting, because it’s OpenAI’s entry into the red hot personal agent market.
Here’s the 1-page summary:
And here’s the deep dive:
What are Dots
Dots’ standout features
The top 3 work use cases for Dots
Mapping out the personal assistant market: Dots vs the competition
1. What are Dots
Dots are always-on agents that lives in OpenAI's cloud with their own computer and browser, which means they keep working when your laptop is closed.
Dot runs on GPT-6 Astra and can connect to more than 4,000 apps through plugins, work with ChatGPT Work and Codex, and use background agents. It can respond to you on ChatGPT voice and Slack.
2 things separate Dots from ChatGPT, Work and Codex:
They work autonomously: Dots run in the background without having the need to keep the laptop on.
Dots can delegate tasks: A Dot can start Work or Codex tasks, including on your own computer if you connect it.
How to Setup Your First Dot
Dots are only available on Pro plans ($100/$200/$500 per month)
So here’s the path to getting started:
Upgrade to the $100 tier
Update your ChatGPT app
Go to chatgpt.com/dots
Click on create new Dot
Select where you want to run it, on cloud or your local desktop
and your Dot is ready.
2. Dots’ standout features
There are 4 things that make Dots Dots.
Dot lets you call your agent and talk to the agent to define all the tasks and delegate tasks too. It remembers what we discussed on the call, and the next time you ping it on Slack or chat, it can pull up from the same context.
You can connect it directly to your Slack
Dots notice things before you ask, even before I configured my bot, it had figured out a list of action items for itself which Dot can perform for me.
Each Dot gets a cloud computer and browser that keep their state between tasks. When it hits a login it can't do, you take over its browser, sign in, and hand it back, which is how it gets past the 2FA wall.
3. The Top 3 Work Use Cases for Dots
There are three things that show the power of this.
Use case 1 - Meeting preps and follow ups
My Dot followed me to and fro meetings quite well. This was my favorite use case.
Use case 2 - The morning call
Using Dots call feature, I had daily calls with my Dot.
It skimmed through my calendar, Slack, socials and Gmail and gave me my to-do list plus schedule. It even grabbed my daily newsletter analytics.
Use case 3 - “EA tasks”
I was planning my last-minute flight to SF for my talk at Opsera, so gave this to my Dot on a call. It worked.
4. Mapping out the personal assistant market: Dots vs the competition
So where are we at with this new personal assistant AI market? There are so many tools out there! I think it’s helpful to conceptualize this as two completely different races we are talking about.
Individual Work Tasks: Grok Bot, Hermes Bots, and Dots are here.
Personal Tasks: Muse and Instinct are here.
Of course, there’s overlap, but this is the basic split.
Within the work tasks market, right now, Grok Bot is ahead.
Let me show you an example.
I gave all three the same job:
Compare the public signup and onboarding flows of three AI meeting assistants, capture screenshots, find the friction and recommend three improvements for a new entrant.
Grok Bot at least notified me to log in, while Hermes just finished research without logging in, submitting incomplete information. My Dot clarified to me what it couldn’t check, but still went ahead with finishing the report with incomplete information.
As a result, Grok Bot won hands down. It handed me a high quality deck. Consistently across my vibe checks, Grok Bot won in comparisons like this.
So putting everything together, here’s the decision tree I’d apply:
If you’re looking for a personal assistant…
If you want to stick in the Meta ecosystem. go for Muse.
Otherwise, go for Instinct.
If you’re looking to do individual work tasks….
If you are price sensitive, or want to keep your data private, choose Hermes Bot.
If you are already in the OpenAI ecosystem, choose Dots.
Otherwise, choose Grok Bot.
I peg Grok Bot as the leader in personal work agent, and Instinct as the leader in personal agent. Your opinion may vary. Reply with it!
This is THE space in AI right now, so I’ll keep tracking it and giving you the updates you need.
Talk soon 👋
















Super useful breakdown, Aakash. I've been using Grok Bots for the past 2 weeks and have been loving it.
Many times I'm giving it tasks I would otherwise give Claude a month ago, but I think the user experience is much more intuitive on GB. They now introduced a concept of primary bot which is more proactive than your other agents. Long-running web research, especially with sites that require logins, is where it shines.
Now, I'm still wondering whether I should give Dots a try (I like the call feature and Slack integration the most). Porting context isn't as big a problem as it was, but there are so many apps on my phone now and the AI app decision fatigue is becoming the new mental tax.