dex
Re great ep @DavidOndrej1 @swyx https://www.youtube.com/watch?v=EWk9PBbKqzc
Suhail
Perplexity is very good at building these benchmarks. Whenever I study them in-depth, it’s always the strongest in the field. DRACO was excellent compared to all the other Deep Research ones, for instance.
Perplexity: We’re open sourcing WANDR.
WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer.
https://research.perplexity.ai/articles/wandr-benchmark-evaluating-research-agents-that-must-search-wide-and-deep
Peter Yang
Tomorrow, I’m sharing a new video on how I use ChatGPT Work (also known as Codex 😅) to do almost everything on my computer.
I’ll cover my complete setup in 7 steps below, from choosing the right GPT-5.6 model to managing email, calendar, and recurring tasks.
📌 Subscribe to get the tutorial tomorrow:
https://www.youtube.com/@PeterYangYT?sub_confirmation=1
Tomorrow, I’m sharing a new video on how I use ChatGPT Work (also known as Codex 😅) to do almost everything on my computer.
I’ll cover my complete setup in 7 steps below, from choosing the right GPT-5.6 model to managing email, calendar, and recurring tasks.
📌 Subscribe to get the tutorial tomorrow:
https://www.youtube.com/@PeterYangYT?sub_confirmation=1
with chatgpt work & sol, i'm finding it incredibly joyful to just ask any question about the business and have it be thoroughly researched and answered.
realizing i have so many questions i wouldn't have bothered asking because they would be too burdensome to answer.
this is just cool
Derrick Choi: Try out the Visualize plugin (in preview) in Codex.
Some ideas click faster when you can see and interact with it.
Here’s a fun example where I asked Codex to build a planet simulator with different controls. 🌎
shyam
Over the last month, I built a long-term memory system for my AI agent.
Not a folder of notes.
Not a vector database with random chunks.
A structured knowledge system that my agent can actually query, reason over, and improve.
The stack combines three ideas:
1. Karpathy’s LLM Wiki pattern
Raw information should not stay raw.
Meetings, notes, links, transcripts, decisions, and documents get compiled into a clean, persistent, human-readable Markdown wiki.
Customers, products, decisions, experiments, and concepts become dedicated pages with citations and backlinks.
The goal is not summarization.
The goal is compounding knowledge.
2. GBrain-style graph retrieval
Once the knowledge lives in structured Markdown, it can become a graph.
Vector search is useful for finding semantically similar text.
But company knowledge is often relational.
“What did customers say about onboarding?” is a semantic question.
“Which customer pain led to which feature, decision, roadmap item, and launch note?” is a graph question.
That requires entities, relationships, timelines, backlinks, and traversal.
My current system has:
1,257 pages
8,778 chunks
951 indexed Markdown documents
semantic search
graph retrieval
typed relationships
timelines
citations
backlinks
3. Hermes as the agent layer
The final step is giving the agent access to this memory.
Instead of starting from zero every session, Hermes can query the wiki and retrieval layers to find old decisions, surface patterns, detect contradictions, update pages, and improve its own workflows over time.
The architecture looks like this:
raw inputs
→ meetings, notes, docs, links, transcripts
markdown wiki
→ structured pages, citations, backlinks
gbrain
→ chunks, entities, graph edges, timelines
qmd/vector index
→ semantic retrieval over markdown
hermes agent
→ long-term memory, synthesis, self-improvement
The biggest lesson:
Don’t start with a vector database.
Start with a source of truth humans can read and edit.
Then add search.
Then add graph structure.
Then give agents access to it.
Most company knowledge bases are digital graveyards.
I think the next generation will look more like operating memory:
human-readable, machine-queryable, graph-connected, and agent-operated.
Hey
learned something useful from @petergyang:
not every annoying workflow needs an api or a custom app.
i wanted to:
• look through the last 10 videos in my youtube history
• review the last 10 bookmarks i saved on x
• extract the useful ideas and information from both
i previously tried turning this into an app or automation.
then i ran into x and youtube access restrictions, api permissions, and all the usual integration headaches.
today i just asked codex to use my computer and handle the administrative work through the interfaces i already use.
it worked like a charm.
thanks Peter!
Parker Conrad
Huge congrats to my cofounder Prasanna on the launch of his new co!
Prasanna S: Launching @vorfluxai : The autopilot for software engineering. I was prev co-founder / CTO of @Rippling ($10B) and #1 coder in India. Vorflux is my high octane Ferrari.
Every AI coding tool still makes you fly the plane. That's the copilot model: you stay in the seat, approving
T Wolf 🌁
Drug-free supportive housing passed the San Francisco Board of Supervisors 8-3. Congratulations @mattdorsey! Another win despite the best efforts of the Coalition on Homelessness who opposed this. We're done listening to them. Drug-free is the way!
https://www.sfchronicle.com/sf/article/drug-free-policy-city-funded-supportive-housing-22344024.php
What do you love about Sol, or why did you switch to it?
Tibo: Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched?
Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens!
https://switch-to-codex.openai.chatgpt.site/
Prasanna is the truth
Prasanna S: Launching @vorfluxai : The autopilot for software engineering. I was prev co-founder / CTO of @Rippling ($10B) and #1 coder in India. Vorflux is my high octane Ferrari.
Every AI coding tool still makes you fly the plane. That's the copilot model: you stay in the seat, approving
Prakash
GPT5.6 one shot solved (by counter example) the central open question of one of the two most important developments in statistics since 1950.. a paper with 130,000 citations.
1) awesome
2) also great that other scientific fields outside ML are moving fast enough for people to public on X immediately (the models can do peer review too after all)
3) yet another result using combinatorial innovation, using a tool from another field in a way a trained human expert might not have
4) yet another result by finding a counter example to a conjecture…
5) wow the pins have started falling eh..
Edgar Dobriban: AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the
Equity over rigor is destroying our public math education in California and increasingly everywhere in the United States
CheesemonkeySF: @garrytan ESSENTIAL BACKGROUND ON THIS STORY:
NOW @thevosf 👇🏼
How California’s math establishment built a generation of students who don’t know what they don’t know.
https://thevoicesf.org/the-layers-of-learning-they-cant-backfill/
Marco Foster
Raphael Warnock: “Donald Trump lost Georgia in 2020. That’s not my opinion, it’s a fact. The votes were counted, recounted, audited, and litigated. He lost, he lost, he lost. He’s trying to sow doubt on the integrity of our elections in Georgia so he can create the pretext to interfere in 2026. This president is a liar, a cheater, and a fraud and he has shown us over and over again that staying in power matters more to him than anything else”
Todd Davis 帅猛男
We all need to remember who is against having drug free housing in San Francisco.
"Supervisors Shamann Walton, Connie Chan, Jackie Fielder and Chyanne Chen voted against the legislation."
Thank you, @MattDorsey. I know this was a time suck, but well worth it.
T Wolf 🌁: Drug-free supportive housing passed the San Francisco Board of Supervisors 8-3. Congratulations @mattdorsey! Another win despite the best efforts of the Coalition on Homelessness who opposed this. We're done listening to them. Drug-free is the way!
https://www.sfchronicle.com/sf/article/drug-free-policy-city-funded-supportive-housing-22344024.php
Mukund Jha
Most businesses don’t need just any software. They need a way to turn how they actually work into software
That’s what @emergentlabs is becoming: the operating system for businesses
Today we’re announcing our $130M Series C at a $1.5B valuation to put that future in their hands
Kevin Blumson
Thanks to @threejs and @Replit, I made a tamagotchi virtual pet 😀
No 3D models, no textures. Every joint and all 523,800 hairs are generated from pure code in ~250ms. Each hair is individually physics-simulated on the GPU.
https://blumson-surprise.replit.app/
Evil Martians
The first edition of the @sfrubyconf was a complete success.
This year, we're doing it again, applying all the learnings, and making it even better.
Brad called it "one hell of a conference." We say it's an unmissable one:
- A keynote by @garrytan
- Talks by Chris Oliver (@excid3) and Rosa Gutiérrez (@rosapolis)
- Organized networking sessions as part of the agenda
- A full day dedicated to community events
Early bird tickets are officially out. Available for 24 hours (or when they sell out).
Get yours: https://luma.com/sfrubyconf2026
Brad Gessler: It was one hell of a conference.
Not mentioned: thanks everybody @evilmartians who took on the financial risk of this conference and brought it to life.
All I need to do now is sell a few hundred @terminalwire licenses so I can help sponsor the next one. 😅
My smoke Canadian
My salad parasitic
My building's in a nosedive
NYC in five!
Peter Yang
ChatGPT Work and Codex are growing by 1M users EVERY FEW DAYS and will likely hit 10M this week.
Here's my new tutorial that walks through how I use Work to do (almost) everything, including how to:
→ Choose the right GPT-5.6 model
→ Organize long-running threads
→ Manage your email and calendar
→ Automate recurring work
→ Build and publish Sites
As usual, I've included real examples and prompts for each.
📌 Watch my full 28-min breakdown here: https://www.youtube.com/watch?v=WLg9qWOf6zw
Aaron Rupar
OMGGGGG -- Clayton is reduced to sitting in silence instead of acknowledging Joe Biden won the 2020 election
OSSOFF: Who won the 2020 election?
CLAYTON: Uhm, you know, I'm not going to do this with you
OSSOFF: This is a job interview. You have an obligation to be honest with the committee. Who won the 2020 election?
CLAYTON: I'm not gonna get into that with you
OSSOFF: You're not being honest or forthright
CLAYTON: ...
Vercel Agent is excellent for optimization questions. Ask it to optimize your build, performance, your bill.
Adam Killam: Vercel's agent just cut our build times by 10x. @rauchg 🙏
as someone who does a lot of dev community and dev youtube this pace of growth has been one of the greatest mysteries to me because clearly there's something to learn here
@sytses i'd love to talk to whoever was the growth guy/gal at Kilo!!
Kilo: 🚨 BIG NEWS: Kilo Code has been acquired by Anaconda (@anacondainc)!
We've grown our agentic engineering platform from zero to a thriving open-source community of 3M developers in only 16 months.
Now, we're joining Anaconda's trusted foundation to cover the full AI-native dev
our models are built to provide the best price for any given task.
if you're able to get better price/perf on any workload, would love to hear the details and look at it together — gdb@openai.com.
Some really cool usecases of Web Analytics API:
▪️ Ask your agent to correlate visitors, custom events (“purchase”, “checkout”), with the evolution of your deployments and performance
▪️ Build custom frontends, and plot this data alongside e.g.: Stripe’s and Resend’s
Vercel Developers: Web Analytics API is now public.
You can build custom reports and live user-facing metrics with the same data that powers the Web Analytics dashboard. https://vercel.com/changelog/web-analytics-api
calling it now
FDE -> ODE -> PDE
the existence of Forward Deployed Engineering and now Ordinary Deployed Engineering implies that the next pokemon evolution is Partial Deployed Engineering
Andrew Curran: Anthropic, Blackstone, Hellman & Friedman, and Goldman Sachs launched their standalone AI enterprise services firm today. It is named Ode. And it just went live at http://Ode.com. No announcement from Anthropic yet, probably forthcoming.
ok this might be AI Woodstock 2.0
i'd love to do this actually outdoors, in a park with a small sound stage, but dont have contacts. has anyone organized an outdoors event in the Presidio, City Hall, or GGP before? even a small one
cc @NaderLikeLadder @TheAhmadOsman youre gonna have to extend ur stay lmao
clem 🤗: Going to be in San Francisco next week. Should we organize some sort of a meetup or march in support of open-source and local AI?
OpenAI
Introducing GPT-Red
An internal automated red teamer on a mission to find our models’ prompt injection vulnerabilities at scale, helping us build stronger defenses before wider deployment.
https://openai.com/index/unlocking-self-improvement-gpt-red/
AI Engineer
🆕This Year In Claude
https://www.youtube.com/watch?v=uU5Gv2h8-9g
@simonw chats with @_catwu and @trq212 about the state of:
- @claudeai Code
- Claude Fable
- @anthropicai culture & product strategy
- Claude Tag & multiplayer collaboration
- The surprising succcess of Remote control
- HTML artifacts for code review
- Why Anthropic uses Auto Mode as the standard for long-running tasks at Anthropic, not --dangerously-skip-permissions
- using Claude for Video editing
- how to be more ambitious as a developer: refusing to "negotiate against oneself"
Timestamps
0:00 Introductions and Claude Code overview
1:22 How coding agents have changed daily workflows
3:51 Shifting focus: Product sense over manual implementation
5:09 Why modern rewrites are now beneficial
6:37 Introducing Claude Tag and team collaboration
11:38 Prioritization and internal "dog-fooding" culture
13:06 The surprise success of remote control features
14:17 Evolving code review processes and automation
17:16 Building trust in new model generations
19:18 Optimizing for capability and user experience
21:23 Reducing system prompts for frontier models
28:05 The philosophy of tool design
30:57 Safety, security, and using Auto Mode
37:53 The human element and developer ambition
41:50 Surprising use cases for Claude (e.g., video editing)
43:35 Limitations and future design aspirations
45:09 Cultural hacks for productivity
46:42 Absurd, fun projects built with Claude
49:03 Audience Q&A
Mira Murati
Our first model, Inkling. Trained from scratch, weights are open, fine-tunable on Tinker today.
Thinking Machines: Today, we are introducing Inkling.
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
https://thinkingmachines.ai/news/introducing-inkling/
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Thariq
this was such a fun talk with Cat and Simon, hope you get to check it out if you missed it
AI Engineer: 🆕This Year In Claude
https://www.youtube.com/watch?v=uU5Gv2h8-9g
@simonw chats with @_catwu and @trq212 about the state of:
- @claudeai Code
- Claude Fable
- @anthropicai culture & product strategy
- Claude Tag & multiplayer collaboration
- The surprising succcess of Remote control
- HTML
Thariq
this was such a fun talk with Cat and Simon, hope you get to check it out if you missed it
AI Engineer: 🆕This Year In Claude
https://www.youtube.com/watch?v=uU5Gv2h8-9g
@simonw chats with @_catwu and @trq212 about the state of:
- @claudeai Code
- Claude Fable
- @anthropicai culture & product strategy
- Claude Tag & multiplayer collaboration
- The surprising succcess of Remote control
- HTML
GPT-Red — improving model security through automated red teaming of prompt injection vulnerabilities:
OpenAI: Introducing GPT-Red
An internal automated red teamer on a mission to find our models’ prompt injection vulnerabilities at scale, helping us build stronger defenses before wider deployment.
https://openai.com/index/unlocking-self-improvement-gpt-red/
something special is happening with Sol:
invincibleHunter: GPT-5.6 Sol is the most impressive model ever. Outside of agentic coding, I find its capabilities in mathematics, especially in vision to be insanely good. I tested both Terra and Sol (Max) on visual mathematics and they are far superior to other models. Huge improvements.
6x price efficiency (!!) with Sol for react/frontend dev:
Aiden Bai: @gdb our benchmark shows that Sol ranks #1 is 6x more cost efficient than Fable across React/frontend work
Ryan Petersen
My favorite launch video I've seen in a while. I don't know quite why but I find it so charming and badass at the same time.
Prasanna S: Launching @vorfluxai : The autopilot for software engineering. I was prev co-founder / CTO of @Rippling ($10B) and #1 coder in India. Vorflux is my high octane Ferrari.
Every AI coding tool still makes you fly the plane. That's the copilot model: you stay in the seat, approving
someone just told me about this take* on CUA
this is one of those gell mann moments for me lol. i've been watching computer use since World of Bits (Shi, fan, karpathy, hernandez & liang 2017). we were the first technical pod to interview @jluan about Adept's work three years ago, we were there in the @AnthropicAI building when they first launched Computer Use 2 years ago, I fanboyed over Claude Cowork in our @felixrieseberg pod 3 months ago, and we ran our first full computer use track at @aidotengineer ft. @DhruvBatra_ @proceduralia @francedot 3 weeks ago.
GPT 5.6 + Superapp is even better at CUA than everything i just mentioned. excited for our @AriX podcast to discuss the @skybysoftware story and Codex progress.
if you actually use these things as intensely as we do, CUA is progressing so, so incredibly fast. i have asked my nontechnical team to CUA as much as possible, all their knowledge work with signing up for random payment and invoicing portals and speaker and sponsor and attendee and vendor and union data requests. if you found yourself nodding along to this take below, you are so not up to date that you don't know what you don't know, and underestimating capabilities is quite a dangerous category error if you are doing any ai decisionmaking.
*i admire dwarkesh alot; only criticizing one single take, not the message nor the overall enterprise, screenshot only to share
This game is TENSE it's clear the teams don't like each other lol
It's not over until it's over
amazing to me that some people want the silent version
https://openai.com/supply/co-lab/work-louder/
Tyler Bosmeny
This is just crazy.
Poetiq has built harnesses that make Opus 4.8 / GPT-5.5 outperform Fable / GPT 5.6 Sol across every benchmark.
Poetiq: Benchmarks are dead (for us).
Our RSI system makes benchmarks too easy.
Give ours a benchmark and it builds its own solution–then beats SOTA. We just did it on 6 at once: math, coding, planning, long-context, tool use, web apps. No human tuning.
I’m very excited by this test time training work for robotic learning! It’s an awesome collaboration between @StanfordSVL and @NVIDIARobotics !
Jim Fan: We scaled a robot model natively to 8,000 timesteps of context, 5 minutes worth of muscle memory, with constant inference cost. Robot policies used to live their lives a few frames at a time (< 0.1 sec), instantly forgetting what just happened. We pushed to 3 orders of magnitude
Argentina almost lost to every single team in the knockout stage yet they kept their composure and didn't give up.
I don't care what people say they're goated.
Easy money (British pounds)
Zach Kram
Timing of Argentina's winners in the knockout rounds:
vs. Cape Verde: 111th minute
vs. Egypt: 92nd minute (stoppage time)
vs. Switzerland: 112th minute
vs. England: 92nd minute (stoppage time)
They haven't led after 90 minutes of any knockout game. But they're in the final.
the new pelican test
Alex: Tell Codex to open Microsoft Paint and try to draw you
This guy led England to lose
FOX Sports: England's manager speaks on his decision to make defensive subs late in the match
Gokhan Egri
We just made Thinking Machines Inkling runnable on Claude Code, Codex and OpenCode. We're making it free for the next 24 hours for the first 1000 people that interact with this post, try it on @BrainbaseHQ!
🐐
Kane 謝凱堯
Legendary run by New York:
> shut down Indian Point reactor
> increase taxes to cover structural fraud, shrinking tax base
> ban data centers
Addicted to mediocrity.
Governor Hochul Press Office: HOCHUL ENACTS NATION’S FIRST STATEWIDE DATA CENTER MORATORIUM
It has been written 😂
dilemma: Argentina just beat Spain at the 2026 World Cup final, 3-2.
Diana
very cool from @lmsysorg so you can get Inkling OSS model to run fast up to 71.7 tok/s input throughput and 171 tok/s decoding speed
LMSYS Org: Inkling, @thinkymachines' first open model, dropped today: 975B total / 41B active MoE, up to 1M context, reasoning natively over text, images, and audio.
Serving and RL support are already live: you can run and shape it on an open stack, starting now.
Day 0 support on SGLang