Now that we've had a few days with it, how are people differentiating between Dot and regular ChatGPT?
I'm having trouble deciding when I should prompt Dot vs using ChatGPT - my Dot seems to afford a single conversation, but I like controlling my context across multiple threads
Your dot learns how you work, keeps context across apps, coordinates Codex tasks, and flags what needs your attention.
引用@dkundelSaw some asks about whether you should use dots or Codex. In reality they work great together! My dot is now my favorite way of using Codex. Less mental overload and the proactiveness is a lifesaver https://t.co/5KQQi89KEA↗
Pro Account user here. I’m failing to understand what the point of dots is. What is it supposed to do that we haven’t already been able to do in ChatGPT? Literally every time I’ve tried to do anything with it, getting it to connect to anything has been an …
All right, we’re locking in. Only things being worked on are simplifications, more efficiency for more usage, groundbreaking features or new models.
Sometimes you have to invest ahead of the curve, but feedback is clear that you all want things to get simpler. On it.
There are so many approaches to agent memory. I believe the simple solution is usually the best, and giving an agent access to a ticketing system like jira, linear, or even just any kind of popular database for it to manage locally like SQLite, will yield …
▲16 wreck_of_u: Any study regarding this? My instinct tells me it's more efficient that I give it raw text to bring home to latent wonderland. I could be wrong
SQLite与工单方案来自作者实践主张;评论区也明确追问是否有对照研究。
Matt Pocock提供了一个更轻的反馈环:让Agent读取最近10次编码会话,找出它在哪里花太久找信息、在哪里依赖过期文档。记忆不一定先做成一个新系统,也可以从“旧会话暴露了哪些仓库导航问题”开始。
引用@mattpocockukPrompt of the day:
/retro read my last 10 coding agent sessions and find ways to make my repo easier to navigate. Find where agents take too long to find relevant information, or rely on out-of-date docs.
Improving navigability is such an underrated way to … ↗
Wait, Google Cloud finally added hard caps on spending per service? I've been wanting that for so many years! They sure took their sweet time. https://cloud.google.com/blog/topics/cost-management/new-ear... Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or …
Hi guys, I have been trying out Claude code on pro plan for about 12 days now, , but I have a few questions. 1. Is this a bug or do I just use Claude a lot 2. Is this a lot of tokens 3. How did I not get limited
Dot, Muse, Grok Bot, and Hermes are all EXCELLENT agents
They all have their own strengths and weaknesses, so you need to know which to use for each task
Here is each agent, and when you should be using them:
GROK BOT- best agent for heavy knowledge work
Strengths: Because you can create many different agents with many different contexts, skills, and memories, you are able to get a wide plethora of work done really really well
Weaknesses: Grok 4.7 is a great model, but not frontier. Also not much usage on cheaper plans.
Who it's for: People who have a diverse set of responsibilities and tasks they need to get done
META MUSE: best agent for getting personal tasks done online
Strengths: Lightning fast, free, and by far the best agent at navigating the internet. Can just 'figure it out' for getting tasks done in browsers
Weaknesses: Not great at coding. Model is lightning quick, but not the smartest or most creative. Weaker at knowledge work because of this.
Who it's for: The best 'personal assistant' agent. Booking reservations, finding flights, filling out forms, unsubscribing from websites. The more 'normie' type use cases. Also it's free. So anyone on a budget.
HERMES: Open source, fully customizable, do anything assitant
Strengths: Fully open source and customizable. Anything you want to change in the agent, you can just change it. Use any model or subscription. Does almost anything you ask without many guardrails. Also a smaller decentralized team supporting it, so it's always great to support the rebels
Weaknesses: No mobile app. Because it's open source and fully customizable, it's not always the most reliable if you tinker too much
Who it's for: tinkerers who want full control over their agent. Also great for the local model community and people with their own home AI labs. My Hermes maintains all my networks and equipment for me better than any other agent
DOT: Your ultimate AI agent project manager
Strengths: The smartest model in the best harness. You get Astra capabilities with best in class ChatGPT tools like Voice, Spaces, and Codex. Can manage all your Codex threads for you and see your projects to completion. Also has ALL of your ChatGPT memories out of the box, so no onboarding needed
Weaknesses: Feels rushed. Lots of rough edges at the moment. Slowest agent. Has trouble with computer and browser use more than you'd like. Tokens in the Dot chat are free, but can heavily burn tokens when managing Codex threads.
Who it's for: Builders who do a lot of coding. ChatGPT power users. Anyone who has long running tasks they'd like an agent to project manage for you
If you put the right agents on the right tasks, you'll have super powers. I use all 4 of these every single day, and switching between them is a breeze
They're all also EXCELLENT agents, so if you can only use one, using any of them is good enough. You can choose literally any of these 4 as daily drivers and you'll be way more productive than the average person
Just make sure you're USING them, and not just reading about them.
If you're stuck at "AI is just another tool", you're making a category error. This is not "same computer, but better". It's far more akin to hiring a bunch of extremely intelligent and capable coworkers with whom you might occasionally disagree.
It's just too good. I got a decade of commercial software dev experience in C#. I am a CTO at my company and have coached a lot of coders from total junior. I hate hate hate hate how good this thing is and what it has done to my passion. I used to love coding …
热评 2 条
▲501 nora_sellisa: LLMs have laid bare how much of software we were writing for decades was bullshit. In an efficient world we would have libraries for everything years ago and only a handful of people would be needed to handle design and critical business logic algorithms. LLMs code well because we've been implementing the same three apps and websites for the last three decades. We should have been mad at the state of the industry a l …
▲161 Lost-Air1265: I also hate it. Twenty years of dopamine shots when solving issues. I have noticed that just because you push out features like crazy it doesn’t mean it gives you the same satisfaction. Development as I knew it is dead.
This is it: "No challenge"
I already do a lot of sports, I play video games, I see my friends often, I do gardening etc
But I have no intellectual challenges anymore, since AI took over my work
What's the solution 🤔? https://t.co/KHDfyrhHQK
引用@T_ZahilI can feel my brain frying since early 2026 because of AI
I need more offline hobbies ↗
▲1 WithoutReason1729: Your post is getting popular and we just featured it on our Discord! Come check it out! You've also been given a special flair for your contribution. We appreciate your post! *I am a bot and this action was performed automatically.*
▲264 FullstackSensei: Trained on 20T tokens using a 768 B300 cluster over 4 weeks. Assuming an $8/hr/GPU, that's ~$4M training cost for a 78B MoE model. Crazy to think how far training prices have fallen in the past two years.
I'm a GP (family doctor) in training in Australia, and I've built a game where you play the GP: you talk to the patient in your own words, examine them, order tests, prescribe and refer. Code scores every consultation against a hand-written answer key, the …
病例、评分代码和成本均来自作者自建系统;样本只有5个病例,不能外推临床能力,更不能用于医疗决策。
Peter Gostev的另一项实验不考一次答题,而让模型在200盘国际象棋里自行选难度、记笔记并尝试提高。作者更新称Opus 5.5大约从1500升到1750 Elo,Astra表现恶化;他猜测可能与上下文长度有关,但明确仍在继续测试。
This is actually super impressive from Opus - it went from ~1500 elo to ~1750 elo.
One possible explanation for Astra's deterioration is the context window - it natively has 272k in Codex, so perhaps with more notes & files it just cracks. I will see if I can run a 1m context window, but probably GPT-6.1-Sol only.
Watch here: https://t.co/7NyccxNgys
引用@petergostevI have a Continuous Learning benchmark where models attempt to learn to play chess. They are given a /goal of learning and improving playing against a Stockfish opponent in 200 games. They can choose the difficulty, take notes, whatever they like - except … ↗
this little experiment i posted last friday was an overwhelming success. have spent the past week up to my ears onboarding companies.
i'm doing one more quick round at $1500/mo.
just 9 slots left. then will be launching the service publicly at $2000/mo.
DM me!
引用@ShpigfordDoing a little experiment. Looking for 5 software companies to run SEO for, 1-on-1.
I built an AI SEO system that runs on my own products every day. Now I want to run it on a few sites that aren't mine.
What you get each month:
• A keyword map of your … ↗
gahhhh i can't wait to show you all this next week.
based on current demand, i'm about 97% certain this thing will be at $1m ARR by end-of-month, especially after i flip the switch.
absolute madness how fast it's already growing just via twitter DMs. https://t.co/RqvZfI5pUo
I've built 12 products. Every one of them flopped. Most of them weren't bad products but they died the same way every time. No idea how to get it in front of people. Ads burned through my budget. Agencies wanted more per month than what I spent even building …
I keep going more niche with reverse engineered web-based versions of games
👉 https://t.co/MohU4A8Zel 👈
6 months ago I tried to reverse engineer Urban Terror
Urban Terror was kind of a Counter-Strike style mod on top of the Quake 3 Arena engine, I think @Erwin_AI got me to play this during … 全文↗
Matt Pocock主张AI时代反而需要更多高杠杆抽象和严格lint,用更小的决策空间约束Agent;这是工程观点,效果仍要在具体仓库验证。
My hot take from chatting to @poteto is that we should use MORE abstractions in the AI age
You can use them (combined with harsh lint rules) to reduce the design space available to the agent and constrain them only to good decisions.
Combined with the fact that high-leverage abstractions let you … 全文↗
Building with the Agents API?
Catch up on what’s new:
引用@stevendcoffeyThe team behind Agents API continues to cook. Here's what's new this week:
Computer use: Spin up a browser + agent with 1 API call 🖥️
Bedrock Managed Agents: Agents API except AWS
GPT‑6.1 Sol 🌞
Environment sizes: Spin up a light or beefy oai-hosted … ↗
Screen Studio can now automatically mask sensitive data of any kind.
It scans every frame of your video looking for emails, credit card numbers, API keys, etc.
It takes ~6s to analyze an entire 20s-long recording on an M3 Mac.
Face masking (avatars, etc.) is coming soon too. … 全文↗
When calculating shadows and optics, Screen Studio actually has an imaginary light source shared between every stage element.
I believe it creates better consistency and makes the loupe just a little more believable as a real thing you shouldn't really analyze while watching. … 全文↗
trq212展示Claude Code内置插件“You should know”:扫描模型输出并提示可能漏看的重要信息,说明插件层开始接管注意力管理。
"You should know" is a great way of staying on top of what's possible with Claude and also a great example of the kind of things mods enable.
You can really make Claude Code yours.
引用@ClaudeDevsWe're adding a new plugin to Claude Code: You should Know.
It scans Claude's output for important information you might miss to help keep you in the loop.
Enable it with:
/plugin enable cc-plugin-you-should-know@builtin https://t.co/A4byi4Q9Df↗
Finances in ChatGPT is rolling out to Free and Go users in the U.S.
Securely connect your accounts with @Plaid and @Experian to make sense of your money and credit, with answers based on your own financial information. https://t.co/I0drGHpkoQ
You start to become aware of how much everyone else is looking at their phones. Everywhere. The worst place is playgrounds.
Parents locked into their phone and their kid runs up pulls on their leg or arm. Gets no response and runs off. Parent unaware they just rejected their kid.
And the instant … 全文↗
There seems to be a whole category of extremely narrow inference runtimes appearing: Strata, ninfer, DwarfStar, Splash, llamAmpere, gufo, etc. They deliberately give up the thing llama.cpp/vLLM are great at - generality - and optimize around a small number of …
I've used Opus 5.5 to help me recreate 7 of the top Adobe apps in 100% pure native Rust: * Photoshop * Illustrator * Premiere * Lightroom * Acrobat Pro * After Effects * InDesign No more subscriptions with cancellation fees. These are yours to own forever. …
▲83 bensyverson: This is absolutely wild. Having built Photoshop readers/writers in the past, I know the challenge just in dealing with the file format. Then you have the sheer amount of UI/UX to implement well. And this looks exceptionally well-executed.
I implemented the feedback I got. Added a rep counter, and a rep by rep breakdown of how you performed in each set. I added a voice coach. Last week it was a coach after your session. Now it is a coach during your session, telling you out loud when your form …
热评 2 条
▲8 Background_Daikon300: Very cool, looks like a unique and helpful product.
▲7 AdPristine1358: Form declines after too many reps, and bad form is what leads to injury. This app could be to incredibly not just for learning, but for people doing high performance training and top athletes. You might want to talk to some personal trainers whose job is to teach form and guide workouts. Going from these form judgments to actual correct posture and form is challenging. You will want this to be actionable, like at the …
First off, yes, this is kinda an ad. But bear with me, you may glean something from my story. I've spent the past month working on snapgains. Early mornings before work. …
热评 2 条
▲2 kyrdotjs: i felt like i did this project and succeed it after read that. w on you bro good job and keep it up!
▲2 divvychugsbeer: Ah, so proud, man. I just took mine to market and got crickets with warm outreach and paid ads. So happy for you; I bet u feel a million bucks! happy days
Opus 5.5的独立“降智”跟踪完成10天基线,作者准备观察到第30天;当前波动仍可能只是统计噪声。
Last week I made LiveNerf, a project to independently track Opus 5.5’s performance day to day and see whether there’s actually evidence of models getting “nerfed” after release. If interested in tracking the performance of Opus 5.5, here’s the repo: …
热评 2 条
▲1 ClaudeAI-mod-bot: **TL;DR of the discussion generated automatically after 100 comments.** Okay, let's get this straight. The top comments are sarcastically screaming "OMG THEY BUFFED IT!" but the thread is actually a three-way brawl. The consensus from the more skeptical (and upvoted) users is that **the performance fluctuations are likely just statistical noise.** They point out that all the data points fall within the confidence int …
▲294 zarrathustraa: OMG THEY BUFFED IT!!! WOWOWOWOW!!!111!!11
Being wealthy is better than being poor, and being able to pay for college is better than working in construction. However, when choosing a specialty it is important to consider one's preferences. If your child don't want to be a lawyer it doesn't matter whether you can afford paying for the degree, and if he insists he would rather be fixing cars than dealing with documents, helping with a student loan is not really the best thing you can offer as a parent. The part about blue collar jobs being …
Because this is a technical article with nonstandard jargon, I'm going to state the most important piece of tacit knowledge that I know psychology professionals have but do not communicate to laymen (I personally came across it somewhere when reading the complete works of Eric Berne, but I can't recall exactly where). Any apparent disorder in a person or persons has to be decomposed into the innate parts that existed prior to birth and the environmental experiences that happened afterwards as a …
I really want to understand the perspective of the other side of this case. Why did this museum care so much about this issue? They appear to have put an enormous legal effort into preventing the release of these point cloud scans. Why?
FTL v0.1.0 was just released, adding async Rust support (multi-thread Tokio runtime) and lots of missing pieces in the Linux compatibility layer. (not my project)
▲6 BocciaChoc: I've always wanted something like a vestaboard, but they're far too expensive, I found some digital options like SplitFlapTV who charge monthly fees Open www.maclaine.se/en/split-flap on any screen and it flips through whatever you set: messages, the time, public transport departures (Swedish focused for now), the weather, electricity prices, countdowns, Stock trackers and so on (this is growing). Every flap turns on …
▲4 khevmoore: This is going to be a blast to play around with. ;)
Didn't expect this at all. GymMane is free, open source, works offline, no ads and no account. Thanks to everyone who tried it or reported bugs, it really helped. https://github.com/InlitX/GymMane
NullPlayer skin support is clean room reverse engineered. Development is coded with claude and controlled by a real working Software Engineer who personally stands behind the app. The app has accumulated 31 releases and 124 github stars over the 11 months. …
I built this after a horrible experience with my own HOA. There was no way to check an association before buying without paying for a title search or digging through court sites. hoaspy.com Github: …
Choosing what to buy often takes hours of digging through reviews and tech specs. Using ChatGPT or similar tools doesn't really help: they give generic answers, aren't real-time, and weren't built to guide a purchase decision. So I created a dedicated tool …
热评 2 条
▲5 Creative_Ambition_: I wasn't happy with the fact that I can't see my search results untill I have signed up. perhaps you could put that at the next stage where I've seen the results and want to proceed further.
▲4 Omnipisix: what do you need today? i typed chromebook, it asked me to login, i closed and left. thats the user experience you will get.
I have a small site with free SaaS calculators. Traffic is fine but not growing. Read somewhere about citation hubs — basically pages that exist so other bloggers can link to them when they need a quick stat. Decided to try it. Built a 2026 SaaS benchmarks …
Trying to build B2B SaaS software, played by the book, and get stuck at the step of talking to your customer. Damn that's harder than I thought. The fact is, your prospect are busy, very busy (that's why you want to sell to them too), and they rarely talk to …
The product is published and I already have around 2-3 visitors per day. I even got like 10 people signed up on the app. But that's not enough for my app type (its an alternative apps directory) I know people will use this app. Everyday there's another post …
There’s a commonly repeated statement in this subreddit: “Distribution is harder than building.” And for most people, I’m sure that’s true to some extent. However, marketing is such a broad field, with countless channels, strategies, and approaches, such that …
I've been looking at alternatives to Google Analytics recently and I'm a bit overwhelmed by the options. I've seen people mention Plausible, Matomo, Umami, Fathom, Simple Analytics, and Keen. I'm mostly looking for something that's relatively simple, …
It's like the entire sub has become that scene from Konosuba where the cult keeps making up fake scenarios saying the only solution is to join their religion
热评 2 条
▲1 WithoutReason1729: Your post is getting popular and we just featured it on our Discord! Come check it out! You've also been given a special flair for your contribution. We appreciate your post! *I am a bot and this action was performed automatically.*
▲344 722e672e722e: This is the third time I’ve found out about something new in this sub from the bitch/complaining post someone made about it.
Hey everyone, I’ve been working on a game called Anyworld. It’s a browser-based multiplayer (single player also supported) text adventure inspired by the early days of AI Dungeon, especially its browser-based free version AI Dungeon 2. The setup is pretty …
Given the commentary on the Q3.8FN release page here https://qwen.ai/blog?id=qwen3.8-flash-next I assume/hope that all the work that's going on to optimise the hell out of running it will be useful when Qwen4 drops?
▲1 ClaudeAI-mod-bot: **TL;DR of the discussion generated automatically after 50 comments.** **The consensus is that this is awesome.** People are blown away by the wallpapers and are especially impressed that you used Claude to **plot real scientific data** instead of just using standard image generation. The community sees this as a much more creative and high-effort use of the AI. Of course, a top-voted comment thread is sarcastically …
▲162 SelectivePro: This is phenomenal! Art-wise, I use Claude to make fictional sci-fi UI's (uispace.org) and so I really appreciate all the hard work you put into this
I took a Java course in college, so I understand the basics of coding and can generally follow how it works. I also have some experience with Python. For example, on my YouTube channel, I transcribe videos and feed the transcripts into AI to generate …
Trying to repost this as the past one tripped reddit's filters and I'm not sure why, and it had a good discussion going on: The entire workflow happened inside claude with a Seedance 2.5 skill that I've used to turn natural language prompts into detailed …
I was wondering if any of you bother with Fable now that Opus 5.5 is out? From my own personal testing I'm finding Opus 5.5 is great but wanted to see what the rest of you thought? Opus 5.5 all the way? Do you know of any advantages of Fable over Opus at this …
A novel way to intervene existing video through diffusion, particularly abstract visuals *\[in this case, audio-reactive geometries\]*: taking its movement and form as the starting point, and reinterpreting its textures, materials, and visual language. I’ve …
▲28 ElatedPyroHippo: The prompt: "Create a single level of a "Mario 64" style platformer game. The level should have a fantasy theme." It took roughly 30 minutes and used 6% of my weekly budget on the $20/mo plan. Sorry about the blank space around the video, first time using OBS Studio and I set something up wrong.
▲12 Lazy-Background-7598: The music makes me wanna kill myself
A novel way to intervene existing video through diffusion, particularly abstract visuals *\[in this case, audio-reactive geometries\]*: taking its movement and form as the starting point, and reinterpreting its textures, materials, and visual language. I’ve …
热评 1 条
▲2 ElatedPyroHippo: Reminds me of winamp visualization plugins from the 90's... milkdrop I think was the name of one of them
Right now, the AI space feels entirely focused on massive datacenter clusters and renting H100s by the hour. But after spending way too much time looking at the actual footprint of these models, I realized that 90% of use cases are completely over engineered. …
I am very uncomfortable about people trying to ascribe religious force or a surrender of human judgment to AI models, and think it is a real safety issue.
The contrast between hanging up the last phone call of the day of high-stakes decisions and then a minute later walking in the door to the pure love of small children who want to laugh and play and show you every discovery they made that day is the strangest and best thing I have ever experienced in life.
There is some speculation about our partnership with Cerebras.
Cerebras is a close partner, and we have a deep engagement pushing on the frontiers of speed.
dot is my favorite openai product so far!
it is amazing to me that each day it feels noticably better as it learns more of my workflow and style.
having it do the stuff i don't like doing--and usually just builds up as a gravity well of dread--has me very happy.
Lots of inventive things made with Claude these past few weeks. A few of our favorites:
A working watermill, built with Opus 5.5.
https://t.co/s9WP51q5os
All right, we’re locking in. Only things being worked on are simplifications, more efficiency for more usage, groundbreaking features or new models.
Sometimes you have to invest ahead of the curve, but feedback is clear that you all want things to get simpler. On it.
Pretty sure there are more dots than bots already in this little world.
Cool thing is that we’re improving them everyday and they learn directly from your feedback.
引用@thsottiauxGlobal reset landing tomorrow 10am PST for all paid ChatGPT accounts. Apologies for the slow start with GPT-6.1 Sol, it's now back to running at expected speeds after the massive load spike in the first two days. ↗
Matt Pocock@mattpocockukAI362.1K粉 · 4条主张用抽象和严格lint收窄Agent决策空间,并用旧会话审计仓库导航。
引用@mattpocockukMy hot take from chatting to @poteto is that we should use MORE abstractions in the AI age
You can use them (combined with harsh lint rules) to reduce the design space available to the agent and constrain them only to good decisions.
Combined with the fact … ↗
My hot take from chatting to @poteto is that we should use MORE abstractions in the AI age
You can use them (combined with harsh lint rules) to reduce the design space available to the agent and constrain them only to good decisions.
Combined with the fact that high-leverage abstractions let you do more with less code - so, more token efficient.
Plus, unwinding the damage from a bad abstraction is much cheaper with agents.
This runs counter to a lot of folks thinking that agents just want to read the raw code. They can, but they're not maximally efficient that way.
Be braver! Design abstractions.
引用@mattpocockukSuper excited to be interviewing @poteto, live, in ~48 hours on YouTube.
We'll be nerding out about skills, high-velocity software factories, and learning SOTA techniques for shipping with agents.
This Friday - 9AM PT. Don't miss … ↗
引用@mattpocockukPrompt of the day:
/retro read my last 10 coding agent sessions and find ways to make my repo easier to navigate. Find where agents take too long to find relevant information, or rely on out-of-date docs.
Improving navigability is such an underrated way to … ↗
用最近10次Agent会话反查仓库导航和过期文档。
▸ 折叠1条(转推/噪音)
转推2026-10-03 RT @poteto: i had a lot of fun chatting with @mattpocockuk today about how i was able to land 2,500 PRs last month! Matt is a wonderful int…
"You should know" is a great way of staying on top of what's possible with Claude and also a great example of the kind of things mods enable.
You can really make Claude Code yours.
引用@ClaudeDevsWe're adding a new plugin to Claude Code: You should Know.
It scans Claude's output for important information you might miss to help keep you in the loop.
Enable it with:
/plugin enable cc-plugin-you-should-know@builtin https://t.co/A4byi4Q9Df↗
Decision models now run on device in llama.cpp. Free, fast, private!
llama serve -hf ggml-org/Kev-4B-GGUF https://t.co/ApnQQaKB4Z
4B决策模型进入llama.cpp端侧运行,性能仍需具体任务验证。
▸ 折叠14条(转推/噪音)
转推2026-10-04 RT @pham_blnh: I just made the most expressive real-time motion harness for the Reachy Mini.
It was created by:
> synthesizing a massive m…
转推2026-10-03 RT @MilksandMatcha: all this just to get us to pay for your newsletter
转推2026-10-03 RT @Aleph__Alpha: Small bird, fast wings, Kolibri is here.
78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe.
Now…
转推2026-10-02 RT @guohao_li: Agentic RL is hard at every layer: environments, training infra, and algorithms. So we’re open-sourcing all of it. For envir…
转推2026-10-02 RT @guohao_li: SETA got accepted to NeurIPS 2026! 🎉
We released the full training scripts for DeepSeek-V4-Flash, GLM-5.2 (LoRA), Inkling-S…
转推2026-10-02 RT @Hesamation: babe, wake up.
Hugging Face just dropped a new banger article about RL training coding agents on different harnesses. http…
转推2026-10-02 RT @natolambert: Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're buil…
转推2026-10-02 RT @_brightmirror: Made with Claude Opus 5.5.
The fearmongering about AI always makes us forget that NOTHING WENT FOOM, as was always pre…
转推2026-10-02 RT @huggingface: The same model, with the same weights, scores 62% in one agent harness and 33% in another.
@adithya_s_k and the @huggingf…
转推2026-10-02 RT @ggerganov: Decision models in llama.cpp are now available
The `/v1/systemone` endpoint is available in the latest llama builds. Use it…
转推2026-10-02 RT @Thom_Wolf: Karpathy: disappears from X
Ben Affleck: alright, gather round, so you'll want to freeze the base weights first, learning r…
Building with the Agents API?
Catch up on what’s new:
引用@stevendcoffeyThe team behind Agents API continues to cook. Here's what's new this week:
Computer use: Spin up a browser + agent with 1 API call 🖥️
Bedrock Managed Agents: Agents API except AWS
GPT‑6.1 Sol 🌞
Environment sizes: Spin up a light or beefy oai-hosted … ↗
Your dot learns how you work, keeps context across apps, coordinates Codex tasks, and flags what needs your attention.
引用@dkundelSaw some asks about whether you should use dots or Codex. In reality they work great together! My dot is now my favorite way of using Codex. Less mental overload and the proactiveness is a lifesaver https://t.co/5KQQi89KEA↗
转推2026-10-03 RT @coreyching: Got to take the main stage at OpenAI DevDay to talk about plugin extensions!
A few takeaways for developers building plugi…
转推2026-10-03 RT @stevendcoffey: The team behind Agents API continues to cook. Here's what's new this week:
Computer use: Spin up a browser + agent with…
转推2026-10-02 RT @simpsoka: Awesome Sites has voting. Go vote for your favorites! Personally I think Jelly Hockey needs some love. #awesomesites https://…
转推2026-10-02 RT @metrox_eth: Today was mostly about modularity.
I’m building MOSS to be modular from day one, so getting started doesn’t mean buying ex…
转推2026-10-02 RT @abstrakt314: GPT-6.1 Astra is extremely good at creating simple Gaussian Splats.
I asked it to use Blender MCP + Lichtfeld Studio to r…
转推2026-10-02 RT @NicholasRuncie: In a self driving car using codex to install a game my dot made for me while I was asleep. This is wild!
@OpenAI@Ope…
Simon Willison@simonwAI229.6K粉 · 2条呼吁默认硬预算上限,并追问Dots与普通ChatGPT的使用边界。
Now that we've had a few days with it, how are people differentiating between Dot and regular ChatGPT?
I'm having trouble deciding when I should prompt Dot vs using ChatGPT - my Dot seems to afford a single conversation, but I like controlling my context across multiple threads
重度用户仍分不清Dots与普通ChatGPT的入口边界。
▸ 折叠1条(转推/噪音)
转推2026-10-04 RT @biilmann: @simonw That was one of the components in our switch to credit based pricing at Netlify a year ago.
转推2026-10-03 RT @MadisonJRickert: Anthropic released native mods for customizing Claude Code yesterday, so I did the obvious thing and built a Jev-based…
Finances in ChatGPT is rolling out to Free and Go users in the U.S.
Securely connect your accounts with @Plaid and @Experian to make sense of your money and credit, with answers based on your own financial information. https://t.co/I0drGHpkoQ
It's time to get serious about Security x AI.
One of the joys of doing AIE is helping others start their own high quality conferences in their countries and domains. Every single 2025 partner has come back — and the 2nd AI Security Summit is now way more important than before. Proud to be back as a founding partner with @snyksec!
Join the event: https://t.co/2IeM4a3MGO
The landscape has completely transformed since we started this a year ago. An exploding number of rogue agents, more breaches, more AI-driven attacks. now more than ever, security must be at the center of the AI conversation. With so much more hacks and vulnerabilities, this stuff is no longer theoretical...
转推2026-10-02 RT @altryne: sneak peak : new something something from @swyx 👇
转推2026-10-02 RT @LaudeInstitute: @a1zhang sees academia as a place to take big research bets. PhD researchers might not have endless compute, but advant…
this is what I like about OpenAI too
flat org w/ a huge bias for action, if you’re willing - just do what piques your interest
引用@nikitabierOne of the most compelling parts about working at an ElonCo was the agility to pivot to whatever the most important problem was.
One day I said: I wasn’t hired to go after spam, but a tsunami is coming and I’m going to drop everything to focus on that.
The … ↗
“Eat at a local restaurant tonight. Get the cream sauce. Have a cold pint at 4 o’clock in a mostly empty bar. Go somewhere you’ve never been.
Listen to someone you think may have nothing in common with you. Order the steak rare. Eat an oyster. Have a Negroni. Have two.
Be open to a world where you may not understand or agree with the person next to you, but have a drink with them anyways. Eat slowly. Tip your server.
Check in on your friends.
Enjoy the ride.”
引用@ryolu_convergence to the mean
when everyone looks at everyone else to decide what to make, we eventually stop making anything of our own.
we copy what works. borrow the language, the aesthetic, the features. we call it learning from the market, following best … ↗
+1 the last time I felt PMF (internal & external) for a product this good was the Codex App launch
Specially it proactively fixing things for me and recommending me stuff I should be on top of
And the experience continues getting better the more I use it
引用@samadot is my favorite openai product so far!
it is amazing to me that each day it feels noticably better as it learns more of my workflow and style.
having it do the stuff i don't like doing--and usually just builds up as a gravity well of dread--has me very … ↗
Dots强PMF评价来自OpenAI生态内部视角。
▸ 折叠1条(转推/噪音)
转推2026-10-02 RT @ajambrosino: we are going to simplify the app
Peter Gostev@petergostevAI27.6K粉 · 4条更新连续学习棋局基准,并比较Sol在Agent Arena的任务成本。
This is actually super impressive from Opus - it went from ~1500 elo to ~1750 elo.
One possible explanation for Astra's deterioration is the context window - it natively has 272k in Codex, so perhaps with more notes & files it just cracks. I will see if I can run a 1m context window, but probably GPT-6.1-Sol only.
Watch here: https://t.co/7NyccxNgys
引用@petergostevI have a Continuous Learning benchmark where models attempt to learn to play chess. They are given a /goal of learning and improving playing against a Stockfish opponent in 200 games. They can choose the difficulty, take notes, whatever they like - except … ↗
Continuous Learning benchmark - now with GPT-6.1-Sol and Fable 5.1 added. Looks like Opus 5.5 actually found some way to increase Elo, so let's see if it can keep going up.
Live games here: https://t.co/6QXRMNQi1Ihttps://t.co/9LpbkRqm5D
引用@petergostevI have a Continuous Learning benchmark where models attempt to learn to play chess. They are given a /goal of learning and improving playing against a Stockfish opponent in 200 games. They can choose the difficulty, take notes, whatever they like - except … ↗
On the Agent Arena leaderboard, we have GPT-6.1-Sol very close to Astra and it is 80% cheaper per task, based on real world prompts
引用@arenaExciting news: GPT-6.1 Sol (Max) by @OpenAi just landed in the Agent Arena at #5 (+11.23%) and reshaped the Pareto frontier!
At a $0.56 median cost per task, it delivers performance within 2 percentage points of GPT-6 Sol and GPT-6 Astra for substantially … ↗
'Hi ChatGPT' - I generated the song over a year ago, still to this day my favourite bit of AI generated media.
The lyrics were by the magnificent GPT-4.5, song by Suno v4 - now the video re-made with Opus 5.5. https://t.co/zK5gf6n8mm
▸ 折叠1条(转推/噪音)
转推2026-10-04 RT @OriolVinyalsML: @petergostev Nice one. Been telling anyone I can this would be a good benchmark for ICL 👏
Philipp Schmid@_philschmidAI122.9K粉 · 1条征集决策模型的真实用例。
can someone catch me up on the whole anthropic/religion thing that's been going around?
@grok or someone terminally online? I had a bit of a week didn't follow everything
引用@trycua1/ We're reimagining what it means for your agents to work with all your computers.
Today we're excited to share Cua Spaces, built on Cua Driver and our virtualization stack, rolling out on macOS today.
Download it now, it's free and source-available: … ↗
The best fun about coming to SF for a week is the people.
Too many to name but I cannot believe how many excellent people I get to call friends in this industry!
From excellent time at OpenAI DevDay (4th time covering) to an absolute banger @thursdai_pod stream from Moscone at CoreWeave FullyConnected, this week had so much!
New models, new friends, a new assistant (hey Dot), and some amazing personal achievements!
Most important, I got to meet and chat with many fans of the podcast, who gave feedback, told me what is valuable to them and that just keeps me going!
One important thing I noticed that even in the middle of the AI SF bubble, there are tons of folks who are still not imagining big enough! Not tons of folks still use personal AI assistants to their full potential, many imagination gaps!
I'm more energized than ever to keep doing AI Evangelism, as humanity steps into its next phase of AI adoption (AI Assistants) and continuing on my mission to bring positivity about AI!
🫡
What's going on tonight in SF?
As if ... 3 conferences, 2 podcasts, 1 hackathon wasn't enough for me... if there's stuff happening, do lmk know, last night here
This shit.. every 2-3 days.
I have to ssh and relogin, every 2-3 days with Claude Code!
Remote-control is all but useless like this, I know there's a push to the cloud but come on @AnthropicAI ! https://t.co/YPNkAEnFGE
If you use Muse and want to plug it into your house, you can claim (for free) one these devices while supplies last! https://t.co/JHiS8f2ESi
引用@natfriedmanWe built a gadget of our own, too: Muse Home Link is a little usb-c powered device that allows Muse to connect to your home network to talk to smart home devices like TVs and speakers and anything else that exposes https.
We manufactured a batch of 5000 Home … ↗
Whatever OpenAI is paying @pvncher - it's not enough
▸ 折叠3条(转推/噪音)
转推2026-10-03 RT @thursdai_pod: FDA-certified... glasses??
Meta just made a genius play that no one is talking about. @altryne on why Ray-Ban Gen 3 gett…
转推2026-10-03 RT @NVIDIAAIInfra: Congratulations to @Cognition for being the first customer live on NVIDIA Vera Rubin NVL72, powered by @CoreWeave.
Wit…
转推2026-10-02 RT @altryne: Asked Sam Altman at DevDay if they'd consider slowing down. They said no.
So nobody's pacing, and this week's ThursdAI came L…
Ben Tossell@bentossellAI201.3K粉 · 7条以会场碎片和Agent使用感受为主,原创信息密度较低。
转推2026-10-03 RT @maria_rcks: queue should be removed from every harness just like plan mode is.
转推2026-10-02 RT @kiwicopple: @supabase has acquired @tursodatabase
i'm been a huge fan of the team and what they've built. we have big plans together.…
One of the most compelling parts about working at an ElonCo was the agility to pivot to whatever the most important problem was.
One day I said: I wasn’t hired to go after spam, but a tsunami is coming and I’m going to drop everything to focus on that.
The org is so flat and agile that your title is basically meaningless. Anyone can sound the alarm about a problem and solve it themselves—or rally the team to join them.
This is how startups work but it’s very hard to make it work in a larger company.
Two hours ago I was officially disconnected from my X laptop.
Good night my sweet little app.
Thank you for giving me the most interesting chapter of my life.
Onto the next one.
We are now entering into an era where any product can be created exactly to a consumer's preferences & needs.
I vibe-fabricated a dog door with a wifi-controlled lock, perfectly to the specifications & design of my house.
I know nothing about metal fabrication or electrical engineering.
It's now getting manufactured and delivered in 2 weeks -- for almost the same cost if I bought a mass-produced item off-the-shelf.
引用@MarioJoos🚨 MAJOR ANNOUNCEMENT: YouTube will start prioritizing original content and reducing re-uploads (RIP clipping)
YouTube has just announced one of the most impactful changes to its short-form algorithm. Its recommendation systems (the algorithm) will start to … ↗
let's start a 30-day series
AI will teach us every day
today: JPEG
I thought I know, I actually didn't
drop next topic ideas below
stuff we use every day but couldn't explain https://t.co/bUgJ3EFHv0
引用@tibo_maker"explain how Shazam works"
I had no idea
Karpathy is so right about this
custom explainer videos will change education https://t.co/L3xjjdGEPn↗
I asked Squad to use Revid CLI and make a 2-min video on how humans have worked over the last 300 years
it did the research, sourced visuals from the internet, made a soundtrack, and fed it all to Revid
here's the result https://t.co/TxCCarOLNE
"explain how Shazam works"
I had no idea
Karpathy is so right about this
custom explainer videos will change education https://t.co/L3xjjdGEPn
引用@karpathyWe'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification … ↗
If you're stuck at "AI is just another tool", you're making a category error. This is not "same computer, but better". It's far more akin to hiring a bunch of extremely intelligent and capable coworkers with whom you might occasionally disagree.
Grotesque. Denmark should step out of the ECHR until such a point that the articles are compatible with deporting criminals, denying citizenship to malignant individuals, and allowing the Danes full sovereignty to decide who is welcome and who is not in their country.
引用@jonatanpallesenArticle 8 of the European Convention on Human Rights says that you have a right to private and family life.
This has over time drifted to mean that we can't deport foreign violent criminals. This is not what people had in mind when the Convention was written … ↗
I didn't realize that many estimates put the Mac at less than 200M active users worldwide! Much smaller than you'd intuitively think if you spend most of your time around tech people.
Apple is going to make it even harder to productively use macOS in the age of agents? Bold move. Let's see how it plays out!
引用@natlungfyShots fired?
Apple plans to introduce new privacy controls for Mac users, warning of the growing risks around granting broad data access to third-party software, including AI agents https://t.co/UYnNQSFdum↗
Peterson on the virtue of becoming dangerous is not just a plea for the individual man, but for the Leviathan too. Well-functioning societies need the capacity for formidable violence, such that they may keep the peace, but paired with the control to use that power sparingly. https://t.co/DLlEmWslVm
It took the Reconquista nearly eight hundred years to reclaim Spain. Don't give me that loser talk that it's "too late for Europe."
The next generation coming up is unbelievably based. They've seen what suicidal empathy has wrought upon their nations.
Never lose faith. https://t.co/8FLfyA9K7r
Omarchy now has an official bug bounty program on HackerOne! The Omacom Foundation has funded it with $100,000 for bounties, and we've promoted @mdisec to Omarchy Core as our Head of Security. https://t.co/CEODf0leTA
▸ 折叠7条(转推/噪音)
转推2026-10-03 RT @Panagiotou90St: Notice how those who are telling you that burning a bystander’s face constitutes protesting against insufficient fundin…
转推2026-10-02 RT @bryan_johnson: I’m watching three friends lose 11 lbs of body weight in a race car tomorrow
+ Tobias Lutke (founder shopify)
+ DHH (cr…
转推2026-10-02 RT @JoshuaSWarren: Omarchy on Apple Silicon makes local AI fast with my MLX and Apple Neural Engine drivers.
@dhh inspired me to believe…
转推2026-10-02 RT @MrMesserschmidt: Lad mig på 60 sekunder forklare, hvorfor ikke-vestlig indvandring er det største problem i Danmark. https://t.co/Tz2rK…
转推2026-10-02 RT @JoshuaSWarren: Got an M1 Pro, M1 Ultra, base M2, M2 Pro or M2 Ultra? Or any M3 or newer Mac?
I'm bringing the Apple Neural Engine up o…
This is great and another example of city states attracting all the skilled Europeans with better living, working and regulatory conditions!
引用@suraj_sharma14Up to $5.5M to move your startup to Qatar.
Just an MVP gets you in the door.
The Startup Qatar Investment Program (backed by QDB) funds tech startups to launch or expand in Qatar.
Two tracks:
START: up to $1.1M if you have a proof of concept or MVP
GROW: … ↗
I keep going more niche with reverse engineered web-based versions of games
👉 https://t.co/MohU4A8Zel 👈
6 months ago I tried to reverse engineer Urban Terror
Urban Terror was kind of a Counter-Strike style mod on top of the Quake 3 Arena engine, I think @Erwin_AI got me to play this during COVID
Because it runs on Quake 3's engine like RTCW and Q3 itself, I thought I could just use ioquake3 and WASM to run it on web, then add a multiplayer server and it'd just work...
But no it didn't, I got stuck 6 months ago, it was buggy and wouldn't work, I kept trying and kinda got it to load but not really play
Today I tried it again with Opus 5.5 and again, I couldn't get it to work, giving up I did a last try with Fable 5.1 on high effort and it after awhile it did figure out the issue!
"The engine's struct was 4 bytes longer than what UT's bytecode expects, everything after that point (the entity count and the entity list) was read 4 bytes off"
It rewrote the bytecode and made it work!
So now we have another game preserved, running 100% the web via WASM, with a multiplayer server tied to it, so you can play with your friends 😊
Every model that comes out makes a bit more things possible and you can just keep trying old things that didn't work before
I see lots of people mod modern games now too which looks very cool too
Here I'm playing it with @marckohlbrugge a few minutes ago! Anyway you can play it with everyone else at https://t.co/MohU4A8Zel now :D
I think this is super accurate
AGI might just create a two-tier split in society, and I'm kinda starting to see it around me
One part of society will succumb to AI slopscrolling and free money (basic income) from government as work is replaced by AGI
The other part will fight that instead and become an elite counter movement that focuses on IRL, family, local community, and health
I think many will be in kind of smaller communities (kind of like eco villages in the pic below), and my friend @csonotes is a good example of that and we try do the same (so I'm probably biased)
@okaythenfuture below thinks it'll be in the central parts of big cities though
Maybe both?
For both of them you need real estate and land though
Real estate and land also one of the few scarce resources left in a post-AGI world where almost every good becomes almost free
If you told me a decade ago I'd want to buy real estate and land I'd call you crazy, but post-AGI it makes a lot of sense I think
引用@okaythenfutureTechnology like this makes me even more bullish on tier 1 global city real estate.
The internet is about to be ABSOLUTELY NUKED by all this AI garbage on the internet, including realistic rendered AI companions that sadly a lot of lonely and isolated people … ↗
After trying the Reverse Osmosis (RO) filter and making V60 coffee with it, it tasted a bit flat
I was kind of expecting that because even with the remineralizer installed (yes we have a remineralizer!) it only gets to 25 TDS, which is quite a low amount of mineral content for water
TDS means total dissolved solids, so like what's left after you remove all the water, essentially the mineral content
The plastic 6L water bottles we used before had a TDS of about 60
So I asked Claude what specialty coffee places do? They usually have an RO filter without any remineralizer. Then they use an industrial machine to remineralize it to the perfect level. That's why coffee of specialty coffee chains like @the_coffee (my fav) taste the same from Brazil to Japan. And their TDS level is high around 150
How can I do get my RO 25 up to 150 though? Claude suggested getting a mineral water coffee kit by Lotus Coffee, you get 4 drippers for Calcium, Magnesium, Potassium and Sodium, so I got it today and tried it. There's a specific amount of drops to add based on your current TDS
The TDS went up to 156 so at specialty coffee water level, and the V60 looked different and way more bubbly
And the flavor is great, it's not bitter, it's pure kind acidic like the Brazil beans we have should be
Adding these drops everytime might be annoying though, maybe I can mix them and make my own solution, or I can get a remineralizer for coffee that gets it to 150 TDS by itself after RO
Anyway I've become a coffee water snob now 🧔♂️ 🫣
引用@levelsio💧 For years I was brought up in the Netherlands being told we had the "best tap water in the world"
"The Americans are silly for buying bottles of mineral water, you can just drink from the tap!"
I remember the same sentiment with my German friends
When I … ↗
✨ I have made coworking meetups a special category on https://t.co/xQI1chWXtQ now
So you can organize one in any city in the world to meet people there
I've also added a [☑️] Recurring Meetup option now so if you do a weekly cowork or event, you don't need to keep adding the same event, it does it for you!
引用@levelsioI've been actively fighting almost guaranteed age-related loneliness for years, studies show after age 30 you start seeing friends less and when you get older especially men don't have any social life left!
It's preventable, but it takes effort and … ↗
引用@marclouI did my yearly body fat test this morning.
5.5%
The doctor used a skinfold tool and measured 7 areas on my body.
5.5% is very low, and I'm still a bit skeptical because bodybuilders have 3-4%, and I look nowhere near that.
I'm 181cm and 84 kg, so my lean … ↗
I did my yearly body fat test this morning.
5.5%
The doctor used a skinfold tool and measured 7 areas on my body.
5.5% is very low, and I'm still a bit skeptical because bodybuilders have 3-4%, and I look nowhere near that.
I'm 181cm and 84 kg, so my lean mass is good. I'm happy not to have excess fat because it can disrupt metabolism and increase cardiovascular risk.
I finally got a Reverse Osmosis filter!
For the past 3 years, I've been traveling the world looking for a place to live. Now I have a yearly rental in Cyprus; I've started upgrading my life 🤗
The RO filter purifies drinking water by removing up to 99% of heavy metals and chemicals in tap water.
Mine is the Cross 90 by Ecosoft. It costs $800, including installation.
My longevity stack is now:
💨 2 air filters
💧 Water filter
🛏️ @eightsleep pod
引用@marclouI just bought 2 air purifiers 💨 for my new place in Cyprus 🤗
I added a big one in the living room where I work, and a smaller one in the bedroom for nighttime.
Around 7M people die every year because of air pollution, so I try to keep my home’s air quality … ↗
转推2026-10-04 RT @trust_mrr: ACQUIRED ✅ 🎉
A mobile app just got acquired on TrustMRR.
💰 Sold for $20K
🤑 $3.0K revenue last 30 days
📈 0.6x multiple
It…
转推2026-10-03 RT @alexwestco: Building a $1M business for the second time.
I made $1,409 in September 2026.
🚀 CyberLeads - $1,083
🤣 Modeling - $250
🌐 B…
转推2026-10-02 RT @marclou: I did my yearly body fat test this morning.
5.5%
The doctor used a skinfold tool and measured 7 areas on my body.
5.5% is v…
转推2026-10-02 RT @marclou: I finally got a Reverse Osmosis filter!
For the past 3 years, I've been traveling the world looking for a place to live. Now…
We are living in the greatest 30 day time period in technology history
In the last month all of these have been released/are expected to be released:
• Opus 5.5: greatest AI model ever
• Meta Muse: The AI agent that will get a billion new people into AI agents
• ChatGPT 6 Astra: First model that could be considered AGI
• Grok 4.7: affordable version of Opus + Elon’s back
• GPT 6.1 sol: Basically free intelligence
• Sonnet 5.5 Dirt cheap intelligence you basically can't tell apart from Opus
• ChatGPT Dot: Incredibly powerful AI agent that's capable of building out large scale applications
The way you work 30 days from now will be dramatically different than the way you work today
You will be significantly more productive with significantly less effort required
Thank your God he/she decided you should be alive right now
Dot, Muse, Grok Bot, and Hermes are all EXCELLENT agents
They all have their own strengths and weaknesses, so you need to know which to use for each task
Here is each agent, and when you should be using them:
GROK BOT- best agent for heavy knowledge work
Strengths: Because you can create many different agents with many different contexts, skills, and memories, you are able to get a wide plethora of work done really really well
Weaknesses: Grok 4.7 is a great model, but not frontier. Also not much usage on cheaper plans.
Who it's for: People who have a diverse set of responsibilities and tasks they need to get done
META MUSE: best agent for getting personal tasks done online
Strengths: Lightning fast, free, and by far the best agent at navigating the internet. Can just 'figure it out' for getting tasks done in browsers
Weaknesses: Not great at coding. Model is lightning quick, but not the smartest or most creative. Weaker at knowledge work because of this.
Who it's for: The best 'personal assistant' agent. Booking reservations, finding flights, filling out forms, unsubscribing from websites. The more 'normie' type use cases. Also it's free. So anyone on a budget.
HERMES: Open source, fully customizable, do anything assitant
Strengths: Fully open source and customizable. Anything you want to change in the agent, you can just change it. Use any model or subscription. Does almost anything you ask without many guardrails. Also a smaller decentralized team supporting it, so it's always great to support the rebels
Weaknesses: No mobile app. Because it's open source and fully customizable, it's not always the most reliable if you tinker too much
Who it's for: tinkerers who want full control over their agent. Also great for the local model community and people with their own home AI labs. My Hermes maintains all my networks and equipment for me better than any other agent
DOT: Your ultimate AI agent project manager
Strengths: The smartest model in the best harness. You get Astra capabilities with best in class ChatGPT tools like Voice, Spaces, and Codex. Can manage all your Codex threads for you and see your projects to completion. Also has ALL of your ChatGPT memories out of the box, so no onboarding needed
Weaknesses: Feels rushed. Lots of rough edges at the moment. Slowest agent. Has trouble with computer and browser use more than you'd like. Tokens in the Dot chat are free, but can heavily burn tokens when managing Codex threads.
Who it's for: Builders who do a lot of coding. ChatGPT power users. Anyone who has long running tasks they'd like an agent to project manage for you
If you put the right agents on the right tasks, you'll have super powers. I use all 4 of these every single day, and switching between them is a breeze
They're all also EXCELLENT agents, so if you can only use one, using any of them is good enough. You can choose literally any of these 4 as daily drivers and you'll be way more productive than the average person
Just make sure you're USING them, and not just reading about them.
四种个人Agent横评;Dots消耗Codex额度为个人体验。
▸ 折叠1条(转推/噪音)
转推2026-10-03 RT @AlexFinn: Dot, Muse, Grok Bot, and Hermes are all EXCELLENT agents
They all have their own strengths and weaknesses, so you need to kn…
Danny Postma@dannypostmaindie185.6K粉 · 2条公开Hetzner、Tailscale与多端Agent的远程环境,准备补Kanban控制层。
my dogs favorite toys:
- plastic bag on a rope
- plastic bag with a ball inside
- plastic bag with my hand in it
toys my dogs doesnt touch
- all the toys i bought him
Finally found my ideal remote agent setup.
Hetzner + Ocra
Orca runs on my laptop, iOS and Hetzner server.
I can start a project at home and continue in the go, without being afraid it breaks anything.
All are connected via Tailscale so the server is secure + can serve me preview builds no matter where I am.
I’m building a wrapper around it now that gives me a kanban task board and many more learnings from my previous AgentOS.
Fun!
Adam Lyttle@adamlyttleappsindie59.0K粉 · 6条复盘一周不用手机、只靠Apple Watch的体验,另有家庭与公司纠纷内容。
Receipts
And actually, I find myself reaching less for my phone now
Recommended https://t.co/JxdNjC8u2h
引用@adamlyttleappsI went a whole week without my phone
Relied only on my Apple Watch Ultra 4
First thing I noticed is how much I reached for the phone as a distraction or because I was even slightly bored.
I never realised how much my brain looks for distractions so it … ↗
You start to become aware of how much everyone else is looking at their phones. Everywhere. The worst place is playgrounds.
Parents locked into their phone and their kid runs up pulls on their leg or arm. Gets no response and runs off. Parent unaware they just rejected their kid.
And the instant realisation: “oh no! That’s me!”
The phone has become so embedded into life
Drive a car? No podcast, no navigation
Go to shop? Cant pay
Realised something important though:
Apple is going the wrong direction
The iPhone Duo is what a focus group would want when asked “what do you want from your iphone?”
Bigger screen
More screens!
But Steve jobs I think would have looked at today’s world and perhaps realised people want less screens
The Apple watch is the perfect answer… but it falls short
watchOS 27 is buggy
Siri works sometimes not others
And worse: you need your phone present to use it. Even when you set reminders and calendar events. But watchOS 26 worked more freely without the phone present
Apple seem to be looking at the watch as an extension of the phone. Not an escape from it.
I went a whole week without my phone
Relied only on my Apple Watch Ultra 4
First thing I noticed is how much I reached for the phone as a distraction or because I was even slightly bored.
I never realised how much my brain looks for distractions so it doesn’t need to deal with its own thoughts
For the last 2 weeks I’ve been desperately trying to save my family. But sometimes you just gotta accept what you don’t want to accept and push forward.
引用@adamlyttleappsThere’s this weird dichotomy of being hyper focused/high agency with something this rewarding (that it literally lifts your families quality of life)
But also being so absorbed in it that people around you think you’re emotionally unavailable. Or … ↗
This company now seems to be attempting to intimidate me into not working with authorities
Can’t make this up
引用@adamlyttleappsAll it took was pressure from the media and 4 years of waiting to get this fixed by the company
I went out of my way to find the CEO contact details. Followed up multiple times.
Tried everything I possibly could to get this attended to.
Glad it's fixed … ↗
When calculating shadows and optics, Screen Studio actually has an imaginary light source shared between every stage element.
I believe it creates better consistency and makes the loupe just a little more believable as a real thing you shouldn't really analyze while watching. https://t.co/XfBEyEPXMP
Almost done adding 'automatic human face masking' to Screen Studio.
I believe this has a lot of real use cases. For example, I wanted to post a demo on how we train our background removal model, but I didn't want to record all the faces of people from the training dataset. https://t.co/k58MUeVZUn
Screen Studio can now automatically mask sensitive data of any kind.
It scans every frame of your video looking for emails, credit card numbers, API keys, etc.
It takes ~6s to analyze an entire 20s-long recording on an M3 Mac.
Face masking (avatars, etc.) is coming soon too. https://t.co/xLAOaaCj7y
逐帧敏感信息扫描的6秒/20秒为M3 Mac作者自测。
Jon Yongfook@yongfookindie172.9K粉 · 2条两条短推,调侃AI致富焦虑与现实银行流程。
It’s 2026 and we have infinite AI intelligence but I’m still making screenshots of a QR code then uploading it to my banking app to transfer money around.
Marc Köhlbrugge@marckohlbruggeindie89.1K粉 · 3条主要吐槽频繁登出,并回应.si域名注册热潮。
I’ll save you the trouble
Wrote a script the few days ago checking all dictionary words and all the good ones are taken
引用@PolymarketBREAKING: Slovenia sees an unprecedented +2,100% surge in .si domain registrations after Trump officially renames AI to “Super Intelligence,” with 44,000 registered in September alone. ↗
If you trim down your prompts, you’re making it worse.
At best, you repeat yourself, you re-describe ideas in several ways, you go on tangents, and you explore way beyond the initial scope. More context means more clarity, at least for the LLM.
Right now, we're in the "omg AI can write code and solve math problems, what a miracle" stage.
But the true Cambrian explosion will be near-instantaneous smart inference at the edge. Real autopilot for any vehicle. No spam ever again. And that's just the obvious stuff.
It looks like taking a stroll through Azeroth while watching your agents do the work has similar effects to actually taking a walk, at least on the level of “distract your conscious mind so that your subconscious can work on the other stuff”.
Interesting.
引用@airkatakanais anyone in tpot actually going to commit to world of warcraft
im sitting on three level 20s in the beta and did this with no drop to my productivity. in fact my productivity may have increased because now im at the computer watching codex all the … ↗
引用@Shpigfordgahhhh i can't wait to show you all this next week.
based on current demand, i'm about 97% certain this thing will be at $1m ARR by end-of-month, especially after i flip the switch.
absolute madness how fast it's already growing just via twitter DMs. … ↗
gahhhh i can't wait to show you all this next week.
based on current demand, i'm about 97% certain this thing will be at $1m ARR by end-of-month, especially after i flip the switch.
absolute madness how fast it's already growing just via twitter DMs. https://t.co/RqvZfI5pUo
this little experiment i posted last friday was an overwhelming success. have spent the past week up to my ears onboarding companies.
i'm doing one more quick round at $1500/mo.
just 9 slots left. then will be launching the service publicly at $2000/mo.
DM me!
引用@ShpigfordDoing a little experiment. Looking for 5 software companies to run SEO for, 1-on-1.
I built an AI SEO system that runs on my own products every day. Now I want to run it on a few sites that aren't mine.
What you get each month:
• A keyword map of your … ↗
My dad wrote a letter to a professional studio with an attached file about some technical stuff and asked them to double check it because he'd made it quickly with Claude.
The owner replied that the file was really well done and that he's looking for staff so he asked if by any chance my dad could introduce him to this Claude guy for a collab.
Simon Høiberg@SimonHoibergindie163.5K粉 · 3条继续抱怨前沿模型拒绝任务,并把Qwen、DeepSeek称为合规模型。
This is getting completely out of hand.
GPT-6-sol will outright refuse to do work over the slightest little details, completely unreasonably.
First acting like a petty lawyer then giving me a small moral speech when I ask them to stop.
I now added a switch on all task items where I can set Qwen 3.8 Max or DeepSeek 4.1 Pro with a single click. These are my "compliance models" - they'll do what I tell them.
引用@SimonHoibergThere's nothing more infuriating than GPT-5.6 or Fable acting "morally superior" and refusing to do the work I assign them.
And it happens often...
GPT-5.6 refusing:
- To create content the way I say.
- To create ad campaigns the way I want.
- To advise or … ↗
Switzerland, UAE, and Singapore all have this in common. High-trust and super safe societies.
引用@levelsioA good example @johnonolan sent me today
An Apple Store in Dubai where you have Apple Watches lying around freely with the door open, not secured, not behind glass
Right by the front door!
A high trust society like we used to have too … ↗
Very sad to see this comment section. How the people I know must be "boring" and "not cool" while bragging about their own friends competing in who can drink more and labeling themselves "alcoholics" like it's a flex.
I'm glad most gen z's and younger seem better wired than this. They'll likely become the influence of my children instead of this sad bunch.
Yesterday we had friends over, and we had a crazy idea
We asked Claude to create a murder party for us 😂
It built the whole game with characters, backstories and clues, and I just had to share the link
everyone opened it on their phone and discover their secret role
The quality was amazing, and it was super fun 🥳
引用@bigaiguyA French engineer who lives quietly in Paris has spent 30 years writing software that the entire internet now runs on without knowing his name.
He wrote the code that streams every YouTube video, every Netflix show, every TikTok clip. He wrote the code that … ↗
引用@noahkaganEvery founder should do this today. ✋
Paste this into Claude:
"Go to [your site]. Sign up as a brand new user. Go through onboarding and use the product like a real customer would. Then give me the top 5 things to fix, with what you saw at each step."
It … ↗
This is it: "No challenge"
I already do a lot of sports, I play video games, I see my friends often, I do gardening etc
But I have no intellectual challenges anymore, since AI took over my work
What's the solution 🤔? https://t.co/KHDfyrhHQK
引用@T_ZahilI can feel my brain frying since early 2026 because of AI
I need more offline hobbies ↗
AI接管工作后失去智力挑战的个人情绪样本。
Daniel Vassallo@dvassalloindie204.8K粉 · 0条本期仅转发Anthropic仍用Typeform做调查的观察。
▸ 折叠1条(转推/噪音)
转推2026-10-03 RT @nickgraynews: I feel like the fact that Anthropic still uses Typeform for their email marketing and customer surveys is testament to th…