← 2026-07-08

Daily Edition

2026-07-09

2026-07-10 →

AI Builders 日报 — 7月9日

追踪 AI 领域真正在做事的人,而不是空谈者。

今日思考

今天最值得关注的信号是 OpenAI 的两条发布同时落地:GPT-5.6 Sol 将于周四公开,而 GPT-Live 作为新一代语音模型已在 ChatGPT 中上线。Sam Altman 明确表示他一直以来更倾向于打字而非与 AI 语音交流,但现在这个习惯可能会改变。如果语音真的成为主流交互方式,这会是对现有 AI 产品设计逻辑的一次根本性调整。

与此同时,Sol 与 Fable 之间的竞争是另一条主线。Matt Shumer 的早期测试显示 Fable 在大多数任务上更胜一筹,而且更具 agentic——一次 Fable 交互往往完成 Sol 需要多次才能做到的事。这与 Ethan Mollick"两者都大幅领先"的说法形成了有趣的对话:差距确实存在,只是领先者可能不是 Sol。


产品与发布

GPT-5.6 Sol (OpenAI)

Sam Altman 今天确认,GPT-5.6 Sol 将于周四公开,同时还有 Terra 和 Luna 两个变体发布,OpenAI 正在全球范围内扩大预览访问的覆盖范围。Greg Brockman 分享了他测试 Sol 开发 Next.js 的体验,称其在日常的 Next.js 开发中表现极其出色,能理解架构权衡,能调查复杂的 issue 报告,在修复 bug 时会考虑代码库其他区域。faviconx.com faviconx.com

GPT-Live (OpenAI)

Altman 宣布 GPT-Live 当天在 ChatGPT 中上线,称其"感觉像魔法、像真人"。Greg Brockman 补充说:GPT-Live 是智能语音 AI,感觉像在一次自然的对话中,测试中仍有大量可能性有待挖掘。目前正在推进 API 和 Codex 的集成。faviconx.com faviconx.com

Grok 4.5 (xAI / Vercel)

Grok 4.5 现已向所有 Vercel 客户开放,通过 AI Gateway 一条命令即可切换到 xAI 旗舰模型。faviconx.com

Eve Chat SDK (Vercel)

Vercel 推出 Eve Chat SDK,支持通过 iMessage、WhatsApp、Telegram 等渠道交付 AI 代理。Guillermo Rauch 称之为"所有 agent 技术栈完美协同的声音",并表示这将驱动他的个人生产力 agent。faviconx.com

Restate BYOC

Restate 推出自带云(BYOC)方案:在用户自有账户和专用 VPC 中提供全托管服务,支持 10 万次/秒的持久操作,已在多个客户的生产环境中运行。faviconx.com


观点与判断

Greg Brockman (OpenAI)

GPT-5.6 Sol 在 Next.js 开发工作中极其出色。它理解架构权衡,能调查复杂的 issue 报告,在修复 bug 时会考虑代码库其他区域,需要的人类干预极少。faviconx.com

Matt Shumer (Hebbia)

GPT-5.6 Sol 是令人惊艳的模型,但在几乎所有测试任务上,Fable 表现更优,且更具 agentic 能力——一次 Fable 交互完成 Sol 需要多次交互才能做到的事。faviconx.com

Peter Yang (Notion)

已获得 GPT-5.6 Sol 早期访问权限,对其评价为"相当不错,甚至可能超越 GTA 5"。同时透露希望采访一位 AI 原生设计师,展示用 design.md 和组件构建的流程 vs 传统流程的区别。faviconx.com

Guillermo Rauch (Vercel)

当前 agent 技术栈的所有组件正在完美协同,这让他感到愉悦,并将驱动他的个人生产力 agent 开发。faviconx.com

swyx (swyx)

大多数 agent 实验室不愿意承认使用中国模型,因为它们需要向政府/国防客户销售。Cognition 团队真正完成了艰难的工作:构建了多语言宣传与审查评估体系,在后训练中成功修正,并在 1000 token/s 的速度下低成本部署。faviconx.com

Taylor Blau 即日加入 OpenAI,开始从事 SCM 工具方面的工作。faviconx.com

Garry Tan (Y Combinator)

不在旧金山建住房,旧金山将把一个巨大的经济 W 变成 civic 历史上最大的 L。faviconx.com

SF Board of Supervisors 已投票将公共银行列入11月公投。Garry 认为这是由 Jackie Fielder 的亲信运营的巨型贪腐机器,会让这座城市破产。faviconx.com


技术动态

Cognition SWE-1.7

Cognition 推出 SWE-1.7,自述为"迄今训练的最强模型",性能接近最强前沿模型但成本只是零头,现已支持 1000 token/s。RL 尚未遇到瓶颈——完善配方后持续看到扩展带来的收益。faviconx.com

Peter Yang (Notion)

已体验 GTA 6 近一个月,整体评价"相当不错",甚至可能比 GTA 5 更好。faviconx.com

X / Twitter

76
ylecun
ylecun @ylecun
Retweeted
Alfredo Canziani @ ICML Alfredo Canziani @ ICML
Come and learn about «Anatomy of Massive Activations and Attention Sinks», 📆 today (Thu) at 🕥 10:30, 📍 HALL A #3002. #ICML2026
garrytan
garrytan @garrytan
Retweeted
Chef Andrew Gruel Chef Andrew Gruel
There you have it.
DSA Watch: DSA Members: "The most important thing we can do is take America down from within"
rauchg
rauchg @rauchg
AI will make all software Native.

Uncompromising performance and platform affinity.

Chris Tate: Introducing Native SDK

The toolkit for building native apps

→ Hot reload
→ Markup + Zig
→ Instant launch
→ macOS, Windows and Linux
→ GPU engine built from scratch
→ Built-in design system and themes
→ Custom components + design tokens

petergyang
petergyang @petergyang
Retweeted
🍓🍓🍓 🍓🍓🍓
i don’t love that anthropic and openai are starting to release models much later publicly than they do privately to employees, friends, and influential people.
as these models become increasingly capable, it creates a dangerous precedent. how is it going to play out when a handful of individuals and companies had gpt 10 five months before you?
garrytan
garrytan @garrytan
Retweeted
Matt Dorsey Matt Dorsey
If “All Housing is Recovery Housing,” why do 26% of all drug overdose fatalities in S.F. happen in Permanent Supportive Housing?
It’s time move beyond ghoulish drug-tolerant mandates that are failing! Come to City Hall Thursday at 10:30 a.m. to speak for Drug-Free Housing options.
Here are the details...
Drug-Free Supportive Housing Ordinance Hearing
Public Safety and Neighborhood Services Committee
10:30 a.m. – Thursday, July 9, 2026
Board Chambers, Room 250
San Francisco City Hall
swyx
swyx @swyx
Retweeted
Akshat Bubna Akshat Bubna
Fun conversation with @swyx on our journey building the cloud for true elastic inference, sandboxes, and more. And of course, how we're evolving Modal's dev experience to be better for agents.
Latent.Space: Modal's Agent-Native Cloud: DX→AX, sandboxes, elastic inference, and 100,000 rollouts https://www.latent.space/p/modal2026
@modal CTO @akshat_b explains why developer experience is becoming agent experience, why agents need infra they can operate instead of YAML they have to reason through,
sama
sama @sama
i do love rottweilers

Peter Gostev: My view of: Fable 5 vs GPT-5.6-Sol. They are not easy models to compare, these are my vibes - take them as you will.

My overall feel is that Fable is a 'wise owl' who is very thoughtful and very well spoken, GPT-5.6-Sol is like a rottweiler who will grab the problem by the
sama
sama @sama
🫶

dax: i've never hyped a model release, we're generally conservative with how we use these things

but gpt-5.6 has had a massive impact on our team, we're using 5x the tokens as we used to

it's not even smarter than fable or anything, but it's just so reliable and fun to use
sama
sama @sama
it surely doesnt

eric: 5.6 solves depression
amasad
amasad @amasad
Retweeted
Emma Emma
I really enjoyed this piece. It’s a rare article that explains, at an intuitive level, how “self-evolving agent products” actually work today.
My takeaway is this: a system needs to build an effective data feedback loop from real user interactions by controlling context; move beyond static benchmarks by creating accurate end-to-end behavior evals that can become automated optimization loops for agent iteration; and use the model’s own reasoning and coding capabilities to drive product improvements at scale (i.e. automatically cluster large volumes of production logs to identify systemic failures, then have AI write the patch, run tests, and submit the PR on its own.)
Through mechanisms like these, developers can build agent systems that can self-repair and continuously improve.
Hopefully one day I’ll be able to share my learning and observation like this one too.
Amjad Masad: Many are asking how Replit is improving so rapidly—we closed the loop and the agent is self-improving.
Technical details here:
ylecun
ylecun @ylecun
Retweeted
Xavier Bresson Xavier Bresson
I will present tomorrow our recent work on Crys-JEPA, where we use Joint Embedding Predictive Architectures to learn crystal distributions for materials discovery.
https://www.linkedin.com/posts/zhongpc_ai-energymaterials-batteries-share-7464854007417581570-drdn
sama
sama @sama
tbh i dont think sol gets that many dates either

Mitchell Hashimoto: I had early access to 5.6/Sol for ~month. Sol is my default. It is faster, plans/judges just as good as Fable, and I think produces better overall work. I’ll reach for Fable still for highly targeted debug or performance work with clear reward functions.

A cheeky way I describe
ylecun
ylecun @ylecun
Retweeted
Kenneth Roth Kenneth Roth
Trump’s cuts in funding for higher education is reducing the number of PhD candidates and “raising fears that the nation’s capacity to produce new science could be diminished.” https://trib.al/TzuHUm6
sama
sama @sama
what a good video

OpenAI: Introducing GPT-Live, a new generation of voice models for natural human-AI interaction.

Rolling out in ChatGPT starting today.

You’ll want to turn the sound on for this one.

amasad
amasad @amasad
When do we stop comparing autonomous agents to hand-written code? You don’t see compilers comparing to engineers hand-writing assembly.

Mehul Mohan: It took $165,000+ of Fable 5 API to port Bun from Zig to Rust.

Wow!

garrytan
garrytan @garrytan
Retweeted
Kane 謝凱堯 Kane 謝凱堯
Personally I don’t think @sfgov, a government that spends >$2B a year, half of Denver’s entire budget, on homelessness but still manages to have some of the worst homelessness in the country, should run a bank.
ConnectedSF: At the First Reading for the "Charter Amendment - Municipal Finance Corporation or Public Bank" at today's Board of Supervisors meeting. The Board voted 10-1 in favor to file a motion for this item be read a second and final time at the July 7 meeting. Bravo @alankennywong for
ylecun
ylecun @ylecun
Retweeted
Yann LeCun Yann LeCun
Re @KenRoth The biggest risk of AI is the concentration of power in a few dominant providers of proprietary AI assistants.
The only solution to AI sovereignty is open source foundation models.
sama
sama @sama
Retweeted
Psyho Psyho
All problems have been solved by OpenAI!
garrytan
garrytan @garrytan
Retweeted
Anjney Midha Anjney Midha
an interesting property of traditional capitalism is that businesses often have to pick between scale and culture
but with AI, small teams are scaling revenue exponentially without losing their culture
a new pareto frontier is emerging in this regard
garrytan
garrytan @garrytan
Retweeted
Daniel Priestley Daniel Priestley
Socialists imagine a class struggle. In their made-up fantasy the CEO is in competition with low level workers, the wealthy entrepreneur is stealing from the underpaid nurse.
In reality, workers do not compete vertically they compete horizontally.
Entrepreneurs compete with entrepreneurs. Investors outbid each other. CEOs are benchmarked against other CEOs. Nurses are hired from a pool of nurses. Etc.
The CEOs pay has no correlation to the entry level workers. The Football star on £300K a week isn’t linked to the person selling drinks in the stadium. A biotech entrepreneur raising VC capital isn’t paid relative to a cleaner.
What is linked is the demand and supply dynamic of each role.
If a company places an ad for a qualified truck driver and 150 people apply for the role, then the company knows it does not need to increase wages for that role. If the company has an open role for months, it is forced to look at the compensation package.
Same for a CEO. A board representing shareholders would like to hire a CEO for a lot less if they could. Their dream scenario would be to hire a CEO who brings in institutional investors, attracts top executives, drives innovation and growth, keeps margins steady and is a good public face for the business even under pressure. It turns out there aren’t a lot of these people looking for work and if you want one you have to pay more than other companies are offering.
The class struggle isn’t vertical it’s horizontal. CEOs are in competition with CEOs. Retail workers are in competition with retail workers. Demand and supply dynamics set the price.
Sure you can say that a CEO want’s profitability and would like wages to be lower BUT it’s not up to the CEO - demand and supply tension sets the price of workers. An Airline like RyanAir would like free pilots if they could get them but they can’t… so they pay the market rate.
The reason incomes are rising at the top and falling at the bottom is not class warfare. It’s technology and globalisation.
Technology makes basic jobs simple, remote or fully automated. At the same time tech makes executive roles more leveraged, more important and more valuable.
A CEO used to run a smaller organisation. Today a CEO who’s 2% better on a $5B company is generating $100M more. Seems sensible to try and pay a few million to get $100M.
Globalisation has put workers from all over the world in completion with each other - downward pressure on wages. Globalisation has given CEOs more market opportunities to explore - upside opportunity to unlock.
The rich are not very interested in buying houses that poor people own. The poor are not buying up the homes the rich want. They are separate groups living separate lives. Try finding the genuinely rich people whose strategy is to hoard normal residential homes - it barely exists as a thing. About 85% of landlords are people who own 1-4 properties. Super-landlords (100+ properties) are 0.2% of landlords and own a tiny fraction of the 30M homes in the UK… and they’re heavily taxed.
Class warfare isn’t real. It’s an imagined war in the minds of socialists.
Demand and supply dynamics are real. To the degree it is measured in class, it’s a horizontal competition not a vertical one.
Gary Stevenson: There's a difference between normal people spending money and really rich people spending money. And it explains why our economy is failing.
ylecun
ylecun @ylecun
Retweeted
David Williams David Williams
I am coining a new word - Yanntificate. To espouse and communicate sensible AI policy.
Yann LeCun: @KenRoth The biggest risk of AI is the concentration of power in a few dominant providers of proprietary AI assistants.
The only solution to AI sovereignty is open source foundation models.
gdb
gdb @gdb
typing into a phone or computer is painfully unnatural

John Collison: It’s really good — feels like the moment where the voice modality for AI apps will finally take off. Upgrade to the latest version of the ChatGPT App to get it. (Also great launch video.)
garrytan
garrytan @garrytan
Retweeted
Mark Zuckerberg Mark Zuckerberg
(1) Today we're releasing Muse Spark 1.1 -- a strong agentic and coding model at a very low price. It's available through our new Meta Model API and in Meta AI.
garrytan
garrytan @garrytan
Retweeted
Y Combinator Y Combinator
Most founders obsess over dashboards and aggregate metrics, but some of the best product insights come from understanding how individual users actually use their product.
In this episode of Startup School, YC's @dflieb walks through one of his favorite tools for better understanding your users, the dot plot. It's a simple two-dimensional grid that reveals usage patterns no aggregate chart can show you.
He’ll cover why it gives founders a better sense of product health, what patterns to look for, and real-world examples of how dot plots helped teams at Google Photos and PayPal.
00:52 — Why DAUs Lie to You
01:39 — What is a Dot Plot and How Does it Work?
02:50 — Picking the Right Event to Track
03:34 — Reading Patterns in the Dots
05:17 — Tracking User State & Attributes
06:16 — The PayPal Fraud Insight
07:59 — Dot Plot vs. DAU Graph
08:56 — Finding Features That Drive Retention
10:30 — Scaling Dot Plots to Billions of Users
11:13 — The $80K Contract That Churned
11:57 — Common Dot Plot Mistakes
12:41 — Dot Plots + Cohort Curves
gdb
gdb @gdb
If you’d like to be a design partner as we test our GPT-Live API (either by making a creative new app or integrating into your existing product), please reach out — gdb@openai.com. Very excited to see what new ideas this technology can unlock for developers.

OpenAI: Introducing GPT-Live, a new generation of voice models for natural human-AI interaction.

Rolling out in ChatGPT starting today.

You’ll want to turn the sound on for this one.

garrytan
garrytan @garrytan
Retweeted
T Wolf 🌁 T Wolf 🌁
Best District Attorney in maybe the entire country. 👇
Brooke Jenkins 謝安宜: Four years ago, I was sworn in as San Francisco’s District Attorney. While there is still more work ahead, I’m proud of the progress the @SFDAOffice has made in standing up for victims, reducing crime, and building a more balanced justice system that holds offenders accountable.
garrytan
garrytan @garrytan
Retweeted
Matt Dorsey Matt Dorsey
Who wants Drug-Free Supportive Housing? San Franciscans do — and by an overwhelming margin!
It’s time to center the needs of supportive housing residents THEMSELVES, by offering them the choice to have the same lease provision banning illicit drug use found in every standard residential lease in S.F.
Speak out today at City Hall at 10:30 a.m.
Information below...
Matt Dorsey: If “All Housing is Recovery Housing,” why do 26% of all drug overdose fatalities in S.F. happen in Permanent Supportive Housing?
It’s time move beyond ghoulish drug-tolerant mandates that are failing! Come to City Hall Thursday at 10:30 a.m. to speak for Drug-Free Housing
ylecun
ylecun @ylecun
Retweeted
Adam Thierer Adam Thierer
We stand at a critical crossroads in the debate over AI governance in the United States, and it feels like we are inching closer to a very serious battle over whether or not open source models will even be allowed in an environment where a new de facto licensing regime has been taking shape.
Lacking formal congressional statutory frameworks or clear administration rules (like the diffusion rule revision), we appear to be left with a sporadic, arbitrary, non-transparent process for model review. The fiction of “voluntary” agreements hangs over this debate, and some large model developers are already showing an incredible willingness to bend over backwards to accommodate national security-related officials / orders that the rest of us are not privy to. It's a very opaque process. And those model developers are expected to play ball with those officials, or else their models get pulled from the market or held up for long periods. Or they will lose any government procurement contracts they have. There is nothing “voluntary” about it when that Sword of Damocles hangs in the room.
As this mess worsens, at some point the question of how to handle open source models will come into sharper focus because it will have to. I've even heard some rumors lately that something may be coming from the admin on this front to address this.
Needless to say, if this informal new AI model review regime expands and takes on more pre-vetting characteristics / requirements, it is hard to see how open source players could comply with such quasi-licensing of AI models. Specifically, if this ambiguous new regime is accompanied by a general presumption of ‘restrict-until-permitted,’ then that would spell doom for open source. That is a very dark path for our country.
Worse yet, of course, would be a move by national security officials to more directly restrict open source models and capabilities. If that happens, then we would be right back in the thick of a Clipper Chip-like battle along the lines of what we saw in the late 1990s. That is a much darker path for America.
Meanwhile, open source developers have no “golden shares” or other goodies to offer the government to make their problems go away.
Let’s be clear: If our government takes the dark path, it will become the single most important battle over computational freedom of modern times. It is time for people to make a stand in defense of open source before it is too late.
garrytan
garrytan @garrytan
Retweeted
h100envy h100envy
Ex-NVIDIA engineer who built Unsloth explained RL, kernels, reasoning, quantization, and agents in 2 hours 42 minutes - better than $5000 fine-tuning bootcamps.
pick the base model -> write triton kernels for 2x faster fine-tune -> quantize to 4-bit -> run GRPO/DPO -> ship a reasoning model on your single GPU.
That loop is why Unsloth is the default way to fine-tune Llama, Qwen, Gemma, and Phi on hardware you already own.
Unsloth + Triton kernels + 4-bit quantization + GRPO/DPO + single-GPU fine-tuning - that's the stack.
Watch and save it, then fine-tune your first model tonight.
h100envy: http://x.com/i/article/2068798868930920448
garrytan
garrytan @garrytan
Retweeted
Ahmad Sadeddin Ahmad Sadeddin
🚀 We open-sourced Sighthound today.
Sighthound is a Rust-based static vulnerability scanner for source code. It runs locally or in CI, uses Tree-sitter parsers, and supports pattern-based detection plus taint flow analysis. It also comes with all it's rulesets with no paid account or signup needed.
https://github.com/Corgea/Sighthound/
Why build another static scanner in 2026?
Semgrep and OpenGrep are useful, but we kept running into a few constraints: the OCaml core raises the contribution bar for many developers, adding language support is not as simple as we wanted, and some higher-signal rule content in the ecosystem sits behind accounts or paid offerings.
We wanted something fast, inspectable, easy to extend, and fully open.
A few implementation details:
- Rust binary for local and CI use
- Tree-sitter parsing
- Pattern matching and taint analysis by default
- Cross-file taint propagation and dependency tracking
- Parallel file discovery and scanning
- Text, JSON, and CSV output
- Rules written in RON and deserialized into typed Rust structs/enums
- MIT licensed
Current support includes Python, JS/TS, Java, Go, C#, HTML, PHP and Ruby.
It focuses on source-code vulnerability classes like command injection, SQL injection, XSS, path traversal, code injection, unsafe deserialization, and crypto issues. It is not a secrets scanner.
We're still have a lot of work to getting to where we want it to be so, we would love the community to test it, break it, report issues, and tell us where it falls short. PRs for rules, language support, fixtures, and false-positive tuning are very welcome.
petergyang
petergyang @petergyang
Competition is great for consumers and developers.

It's great to see Grok 4.5 and Muse Spark 1.1 both be Opus level viable models. Hopefully Google launches Gemini Pro 3.5 soon too.

Having 5 labs compete vs. 2 will keep all of our AI subscription plans alive and prices reasonable :)
ylecun
ylecun @ylecun
Retweeted
Daniel Jeffries Daniel Jeffries
Open source powers everything in America. We must fight any and all attempts to restrict it in the age of AI.
We must never surrender to the short-sighted safetyists and hawks.
Open source is the entire reason American tech dominates across the globe. If you can't see this it is largely because it is invisible but your blindness to it is no excuse. You are protected by a smoothly running invisible shield the powers everything and that shield is open source.
It powers every major cloud. The router in your house. Your phone. The servers at your work. The NASDAQ and NY Stock Exchange. Radar systems. Just about every website you know and love. The SaaS you're reading this on right now, called X.
Restricting it in the AI era is idiotic, short-sighted and foolish at an almost mind boggling level. It is civilizational suicide.
Take Microsoft. In the early days of Linux they pushed hard against it and called open source "communism" and "cancer." Imagine if their short-sighted stupidity had won the day? Their Azure cloud now predominately runs Linux.
They would have destroyed *their own* future revenue with their foolishness and greed.
In this fight we must never surrender.
We shall fight on the beaches, we shall fight on the landing grounds, we shall fight in the fields and in the streets, we shall fight in the hills.
We shall never surrender!
Adam Thierer: We stand at a critical crossroads in the debate over AI governance in the United States, and it feels like we are inching closer to a very serious battle over whether or not open source models will even be allowed in an environment where a new de facto licensing regime has been
sama
sama @sama
Retweeted
OpenAI OpenAI
Today. 10am PT.
rauchg
rauchg @rauchg
X is the arena. Always has been
petergyang
petergyang @petergyang
Get your Hermes agent (or any other agent) to do this and thank me later :)
garrytan
garrytan @garrytan
Retweeted
Kane 謝凱堯 Kane 謝凱堯
In San Francisco, the public school district (55% reading, 46% math proficiency) will block downtown commercial revitalization projects unrelated to schools.
San Francisco Chronicle: NEW: The prospective buyers of the shuttered San Francisco Centre mall have walked away from the pending deal to buy the downtown shopping center after months of due diligence, according to sources familiar with the talks. https://www.sfchronicle.com/realestate/article/sf-centre-mall-downtown-sale-dead-22337271.php?taid=6a4e7befeafdd900014de862&utm_campaign=trueanthem%2B3988&utm_medium=social&utm_source=twitter
sama
sama @sama
5.6 livestream going now.

in addition to the model, 3 major product things.

1. ChatGPT Work--really big deal!
2. new ChatGPT desktop app
3. hosted sites
gdb
gdb @gdb
happening now!

OpenAI: Today. 10am PT.

sama
sama @sama
obviously the best model we have ever produced, but also one of the best blog posts we have ever produced:

https://openai.com/index/gpt-5-6/
sama
sama @sama
we have heard enterprises on their concerns about AI costs, and 5.6 sol is a huge step forward for dollars-per-task, as are terra and luna
mattshumer_
mattshumer_ @mattshumer_
http://x.com/i/article/2074916560041349120
mattshumer_
mattshumer_ @mattshumer_
I've had access to GPT-5.6 since May 27.

Everyone is going to tell you it's incredible. They're right.

But this isn't another post glazing 5.6.

For 2 weeks, it was the best model I'd ever used.

Then Fable came out, and I stopped using GPT-5.6 overnight.

Here's my review:

Matt Shumer: http://x.com/i/article/2074916560041349120
petergyang
petergyang @petergyang
Retweeted
Peter Yang Peter Yang
OpenAI just made GPT-5.6 available to everyone, so I wanted to answer the most burning question:
Is GPT-5.6 better than Claude Fable 5?
I tested both models across 6 real use cases:
→ Frontend design (5.6 has closed the gap)
→ Build a 3D Star Fox-style game
→ Edit and publish clips with browser use
→ Add a feature to a mobile app
→ Get life and business advice
→ Improve my personal AI OS
📌 Watch my full head to head comparison here: https://www.youtube.com/watch?v=8mY9wx_iMSU
petergyang
petergyang @petergyang
Retweeted
Peter Yang Peter Yang
OpenAI just made GPT-5.6 available to everyone, so I wanted to answer the most burning question:
Is GPT-5.6 better than Claude Fable 5?
I tested both models across 6 real use cases:
→ Frontend design (5.6 has closed the gap)
→ Build a 3D Star Fox-style game
→ Edit and publish clips with browser use
→ Add a feature to a mobile app
→ Get life and business advice
→ Improve my personal AI OS
📌 Watch my full head to head comparison here: https://www.youtube.com/watch?v=8mY9wx_iMSU
mattshumer_
mattshumer_ @mattshumer_
GPT-5.6-Sol one-shotted this voxel-based Manhattan.

Just look at the precision... it's insane.

It ran for almost a week, completely autonomously, to get the job done.
mattshumer_
mattshumer_ @mattshumer_
Most people know I usually prefer OpenAI's models over Anthropic's.

But for this generation, it's the opposite.

That said, a larger pretrained base + OpenAI's RL stack will create an absolute beast of a model, so I expect OpenAI's next release will bring me back.

Matt Shumer: I've had access to GPT-5.6 since May 27.

Everyone is going to tell you it's incredible. They're right.

But this isn't another post glazing 5.6.

For 2 weeks, it was the best model I'd ever used.

Then Fable came out, and I stopped using GPT-5.6 overnight.

Here's my review:
sama
sama @sama
Retweeted
ARC Prize ARC Prize
GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8%
Sol is the first verified frontier model to ever beat an ARC-AGI-3 game
It is the best model at orienting in a situation it's never encountered
garrytan
garrytan @garrytan
Retweeted
Y Combinator Y Combinator
Congrats to @jmorgan, @mchiang0610 and @ollama on their $65M Series B!
They built the easiest way for developers to get up and running with open models, and it's become the leading platform for exactly that, with 8.9 million developers and 85% of the Fortune 500 using it today. All with just 14 employees.
https://ollama.com/blog/all-aboard-open-models
garrytan
garrytan @garrytan
Retweeted
Garry's List Garry's List
Reminder that Rep. Ro Khanna defended Hasan Piker—a man who celebrates the murder of the wealthy—even while quietly running one of the most profitable stock portfolios in Congress.
Kane 謝凱堯: Rep @RoKhanna has been playing populist w performative policies like banning Waymo.
Instead of Congress's electronic financial disclosures, he hand-files papers that are not searchable.
I burned some tokens and OCR's all >1000 pages and >$300M of his finances! Site shortly.
gdb
gdb @gdb
We’ve brought together ChatGPT and Codex, in the form of ChatGPT Work: an agent for your most ambitious work.

Use it from mobile or web, in addition to desktop — no need to leave your laptop cracked open!

OpenAI: Introducing ChatGPT Work, a new agent in ChatGPT powered by Codex and GPT-5.6.

It can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work.

It’s a whole new way to get work done.

mattshumer_
mattshumer_ @mattshumer_
You will 100% be able to one-shot a GTA-scale game by this time next year.

It’ll cost a ridiculous amount, but it will be possible.

There will be a creativity explosion unlike anything we’ve ever seen.

Obviously, this assumes that we still have access to frontier models.

Matt Shumer: GPT-5.6-Sol one-shotted this voxel-based Manhattan.

Just look at the precision... it's insane.

It ran for almost a week, completely autonomously, to get the job done.

alexalbert__
alexalbert__ @alexalbert__
More Fable!

ClaudeDevs: We've reset 5-hour and weekly rate limits for all users.
petergyang
petergyang @petergyang
GPT 5.6 has closed the gap with Fable on frontend design.

I asked both models to make a travel site for my upcoming Japan trip:

→ GPT 5.6 created a beautiful site with a 3D Torii gate
→ Fable 5 made a nice site too with animated snowflakes

Tip: Ask AI to “add a 3D WebGL element to the hero” when you want animated 3D graphics like this.

📌 Watch me walk through both sites live here: https://youtu.be/8mY9wx_iMSU?si=4Mu5bZHRyWHbnQ3y&t=116


Peter Yang: OpenAI just made GPT-5.6 available to everyone, so I wanted to answer the most burning question:

Is GPT-5.6 better than Claude Fable 5?

I tested both models across 6 real use cases:

→ Frontend design (5.6 has closed the gap)
→ Build a 3D Star Fox-style game
→ Edit and
petergyang
petergyang @petergyang
Retweeted
Peter Yang Peter Yang
GPT 5.6 has closed the gap with Fable on frontend design.
I asked both models to make a travel site for my upcoming Japan trip:
→ GPT 5.6 created a beautiful site with a 3D Torii gate
→ Fable 5 made a nice site too with animated snowflakes
Tip: Ask AI to “add a 3D WebGL element to the hero” when you want animated 3D graphics like this.
📌 Watch me walk through both sites live here: https://youtu.be/8mY9wx_iMSU?si=4Mu5bZHRyWHbnQ3y&t=116
Peter Yang: OpenAI just made GPT-5.6 available to everyone, so I wanted to answer the most burning question:
Is GPT-5.6 better than Claude Fable 5?
I tested both models across 6 real use cases:
→ Frontend design (5.6 has closed the gap)
→ Build a 3D Star Fox-style game
→ Edit and
sama
sama @sama
Retweeted
Alexander Embiricos Alexander Embiricos
Massive day for us @OpenAI:
- GPT-5.6 SOTA at ~everything & by far most token efficient
- Agents for everyone in the new ChatGPT app Work and Codex modes
- Work mode available on desktop (most powerful), web & mobile
- Ultra mode
- Sites out to everyone
- Artifact templates
swyx
swyx @swyx
Retweeted
AI Engineer AI Engineer
🆕 Congrats to @OpenAI on the highly anticipated launch of GPT 5.6 Sol, Terra, and Luna!
https://youtu.be/pMggiOb18tc
Here's @romainhuet, @embirico, and @steipete on why this is the dawn of the Golden Age of AI Engineering:
mattshumer_
mattshumer_ @mattshumer_
Workbench template to allow many GPT-5.6-Sols and Fables to coordinate as a team:

https://workbench.md/templates/agent-team-hq#from=twitter
sama
sama @sama
check this out! you can get some amazing things done.

codex is the core of our new work product and what makes it so good. codex is not going anywhere.

OpenAI: Introducing ChatGPT Work, a new agent in ChatGPT powered by Codex and GPT-5.6.

It can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work.

It’s a whole new way to get work done.

sama
sama @sama
Retweeted
OpenAI Developers OpenAI Developers
OpenAI Build Week is here.
The challenge starts July 13. Across the week, join live sessions and community events with builders around the world.
Registration is open.
Bring the idea sitting in your backlog.
Build it with Codex.
mattshumer_
mattshumer_ @mattshumer_
My guide to prompting Fable also applies to GPT-5.6-Sol.

The techniques I describe will allow you to get outputs like this.

https://workbench.md/pub/IbaCrTjLJT?key=uQOQ2NPO3TTUSXyYDjyLf

Matt Shumer: GPT-5.6-Sol one-shotted this voxel-based Manhattan.

Just look at the precision... it's insane.

It ran for almost a week, completely autonomously, to get the job done.

garrytan
garrytan @garrytan
Retweeted
SemiAnalysis SemiAnalysis
The Future of Meta Superintelligence: A 1 Year Progress Update
A top tier RL environment startup spawns out of thin air,
the most aggressive compute ramp we've ever seen,
2000km+ scale-across, and some advice for Google DeepMind
https://semianalysis.substack.com/p/the-future-of-meta-superintelligence
garrytan
garrytan @garrytan
🇺🇸

Sriram Krishnan: between Grok from @elonmusk / @mntruell and Muse Spark from @finkd this week very excited to see an ecosystem of frontier models in America.
garrytan
garrytan @garrytan
Retweeted
Sam Lyman Sam Lyman
NEW: The New York Times reports that "China, Russia and, to a lesser extent, Iran have sought to use state media outlets to turn the controversy over data centers in the United States into 'a domestic fracture point.'"
Between Jan. and June, state media from these 3 countries mentioned data centers 700 times.
The NYT report cites BPI's research and highlights the work of Chinese propagandists and "covert Russian information operations" to foment anger against data centers.
petergyang
petergyang @petergyang
I asked both GPT 5.6 and Fable to build a 3D Star Fox-style level complete with enemies, power-ups, and even a boss battle.

Here are the results. It’s incredible that both models can one-shot a working game from a simple prompt in about 10 minutes.

However, I have to give the slight edge to Fable on this one for remembering to support barrel rolls 😅

📌 Watch me play each game here: https://www.youtube.com/watch?v=8mY9wx_iMSU&t=344s


Peter Yang: OpenAI just made GPT-5.6 available to everyone, so I wanted to answer the most burning question:

Is GPT-5.6 better than Claude Fable 5?

I tested both models across 6 real use cases:

→ Frontend design (5.6 has closed the gap)
→ Build a 3D Star Fox-style game
→ Edit and
petergyang
petergyang @petergyang
Retweeted
Peter Yang Peter Yang
I asked both GPT 5.6 and Fable to build a 3D Star Fox-style level complete with enemies, power-ups, and even a boss battle.
Here are the results. It’s incredible that both models can one-shot a working game from a simple prompt in about 10 minutes.
However, I have to give the slight edge to Fable on this one for remembering to support barrel rolls 😅
📌 Watch me play each game here: https://www.youtube.com/watch?v=8mY9wx_iMSU&t=344s
Peter Yang: OpenAI just made GPT-5.6 available to everyone, so I wanted to answer the most burning question:
Is GPT-5.6 better than Claude Fable 5?
I tested both models across 6 real use cases:
→ Frontend design (5.6 has closed the gap)
→ Build a 3D Star Fox-style game
→ Edit and
petergyang
petergyang @petergyang
Re Btw the video above was made only using GPT 5.6 using ffmpeg and browser + computer use.

GPT played the game and recorded it.
petergyang
petergyang @petergyang
Ugh France is way too stacked
mattshumer_
mattshumer_ @mattshumer_
Retweeted
Josh Josh
Ran for almost a week.
Looks right in line with METR’s time-horizon graph.
Matt Shumer: GPT-5.6-Sol one-shotted this voxel-based Manhattan.
Just look at the precision... it's insane.
It ran for almost a week, completely autonomously, to get the job done.
swyx
swyx @swyx
Re u guys were clowning on @greptile but turns out they were just the inspo for @openai to go so, so much harder https://x.com/JangLawrenceK/status/2075204015890325703

Lawrence Jang: guys I just cancelled my Claude plan I don’t know what happened

petergyang
petergyang @petergyang
Some praise and feedback about OpenAI's launches:

1. More than any other lab, OpenAI has the opportunity to make working with agents mainstream. ChatGPT with images, live voice, browser/computer use, and plugins for all your favorite apps is the closest thing to working with a super-capable coworker who can learn and do any white collar work.

2. I didn't get to test GPT 5.6 weeks early, but after a day of testing, I think my best compliment is: "It's got that dog in it."

It basically never gives up, no matter how complicated a task you give it, and is very reliable thanks to browser use and the other features above.

It does seem to burn more tokens than GPT 5.5, although maybe I don't have the right settings. I'm using Sol on High.

3. GPT Live is arguably a more important launch than today's updates for the masses.

We've been talking to each other using voice for 100K+ years, while the keyboard and mouse weren't introduced until the 1950s/60s. Maybe voice will soon make these two tools look like old computer punch cards.

Also, the best thing about GPT Live is that it lets me pick the intelligence level. I care about latency, but I care more about talking to someone who's smart and gets me.

Speaking of which, it'd be cool if I could have a live conversation with ChatGPT voice and see the tasks it's doing for me.

Ok, now for some constructive feedback:

4. I think the ChatGPT Work vs. Codex thing is confusing. It raises questions like, "So developers aren't working?" and "What if I'm using it to plan my vacation, is that work?"

I think the whole thing should just be called ChatGPT, Codex, or ChatGPT Codex and there should be no tabs or toggles. IMO, Codex is not that complicated for a normie to understand and use. You're still chatting, except now it can do stuff.

The sooner this is all unified, the better.

5. I'm very confused about when I should use Sol, Terra, or Luna, and whether I should set effort to Light, Medium, High, etc.

There's nothing in the UI that walks through what to use when. I understand power users might want these options, but regular people will just get confused.

6. Another thing that's confusing is tasks vs. chat. Ideally, your past chats should still be there in the left nav, except now those chats can do something with plugins. Normies will not understand why their chat history is suddenly harder to find.

7. When I was a PM, I loved building with a community of trusted users who gave me feedback along the way before general availability. But I also tried to make it transparent to users what qualifications were needed to join this community. I don't mind that OpenAI is doing the former, but the latter is not clear at all.

Thanks for listening. Codex has changed how I work and for that I'm very grateful.

OpenAI: Introducing ChatGPT Work, a new agent in ChatGPT powered by Codex and GPT-5.6.

It can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work.

It’s a whole new way to get work done.

garrytan
garrytan @garrytan
Retweeted
Y Combinator Y Combinator
YC is coming to NYC next week! 🗽
We're hosting a happy hour for designers with YC's @aaron_epstein, @raphaelschaad, and @eve_bouff.
If you're a designer who wants to build, come hang.
RSVP: https://events.ycombinator.com/nydesignhh-summer26
swyx
swyx @swyx
Retweeted
Sarah Chieng Sarah Chieng
Technical writing is the #1 top-of-funnel motion at a $12B company
@philipkiely broke down exactly how he does it, his full writing process, the hindsight 20/20 lessons, and how one article drove 500K+ views
This [technical] Write and Learn workshop #5 is part of our Independent Studies series, hosted with @swyx @KernelLabs_ai.
All sessions are recorded. Our next one will be in two weeks :)
sama
sama @sama
i am really sad about this and very grateful for all fidji has done for openai, and even grateful for her friendship and who she is as a person.

we all wish her the best for a speedy recovery. this sucks.

Fidji Simo: Today, I shared with the OpenAI team that I have decided to leave my full-time role at OpenAI and transition to being a part-time advisor.

Three months ago, I had to go on medical leave after a severe exacerbation of a chronic illness I’ve lived with for seven years. During that
mattshumer_
mattshumer_ @mattshumer_
Sup

Polymarket: NEW: Vibe coder claims GPT-5.6 “one-shotted” a complete 3D Manhattan build after running autonomously for nearly a week.
amasad
amasad @amasad
If you haven’t tried Fable in Replit…

waqas: having a bit of an out-of-body experience seeing claude (fable) talk to the replit agent while i can see their convo happening live in the browser...
gdb
gdb @gdb
easy to overlook the implications for how Sol can accelerate your engineering workflows

Andrew Curran: Quote from OpenAI on the livestream.
'Already, Sol has been transforming our research program. As one example, GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna.'

YouTube

0

No recent videos fetched on this date.