← 2026-07-31

Daily Edition

2026-08-01

2026-08-02 →

AI Builders 日报 — 8月1日

追踪 AI 领域真正在做事的人,而不是空谈者。

今日思考

今天最大的信号不是某条具体动态,而是 OpenAI Astra 模型用约 2000 美元的计算成本解决了 10 个数学和理论计算机科学的开放难题。这意味着:AI for Science 的时间窗口正在打开,大型药企、高校和研究机构将成为下一波 AI 采买的主力。而 Sam Altman 和 OpenAI 团队全员转发这一消息的姿态说明,GPT-next 的定位已经不是"更好的聊天机器人",而是"科研基础设施"。与此同时,Y Combinator 宣布开源 QM——一个内部用的多 Agent 系统——发出了另一个对位信号:不是所有公司都会把 AI 基础设施押注在某一家的 API 上,定制化和可控性才是企业级 AI 的真实需求。


产品与发布

OpenAI Astra:2000美元解决10个数学难题

OpenAI 下一代模型家族 Astra 的内部版本,在数学、量子复杂性和理论计算机科学领域解决了 10 个长期悬而未决的开放问题,包括推翻 Connes 刚性猜想、高维球堆积问题的更好边界、计算矩阵积的电路复杂度新下界等。所有证明均附带 Lean 证书和 CoT 推导过程,总计算成本约 2000 美元(Sol API 价格)。Sam Altman 转发了相关证明链,并称之为"team humanity"。faviconx.com

OpenAI 团队持续改进 Git 上游

OpenAI 内部 Git 团队将性能、正确性和测试改进持续向上游合并,Codex 应用的 Git 使用效率也有提升。faviconx.com

YC 开源 QM 多 Agent 系统

Y Combinator 宣布开源 QM——一个内部使用的多 Agent 系统,涵盖会计、法律、活动策划和工程等多个业务线。YC 此前在内部广播所有 Agent 会话,团队通过观看彼此使用 Agent 的过程实现知识共享,这套机制被形容为"公司的 Agent 历史作为共享学习材料"。faviconx.com

开源 Agentic CRM

基于 eve.dev 和 Next.js 构建的开源 Agentic CRM 现已发布,MIT 许可证,支持模型无关、自托管或无服务器部署、多渠道和 Headless 模式。faviconx.com


观点与判断

Amjad Masad (Replit 联合创始人)

  • 8B Qwen 模型在象棋上超越 GPT-5.6 和 Stockfish Qwen 8B 模型通过高推理和响应链达到约 1500 Elo,持续击败前沿模型和 Stockfish level 0,每步决策仅需 1-2 秒(对比 GPT-5.6 的 30 秒),展示了小模型在特定任务上的高效优势。faviconx.com

Garry Tan (Y Combinator CEO)

  • OpenAI 正在成为"开放平台" 2026 年最值得关注的变化是 OpenAI 实际上在向开放平台方向转型——将智能作为按需调用的utility(utility computing),而非"All in 全栈"的信号。faviconx.com

swyx (AI Engineers Podcast 主持)

  • /loop 和 /goal 在 GPT-5.6/Claude-5 时代仍值得坚持使用 在 Agent 设计圈中,坚持使用 /loop 和 /goal 的人属于少数派,但他认为这两者在需要"可 steering 性和自主性的平衡"以及"开放性的循环生成"场景中仍不可替代。faviconx.com

  • Google 曾内部开发类 ChatGPT 产品但未敢发布 2019 年 Google 团队开发了名为 LMChat 的产品,功能接近一年后发布的 ChatGPT,但 Google 担心风险而未上线,DeepMind 也被阻止发布可能颠覆 Google 的产品。faviconx.com

Guillermo Rauch (Vercel CEO)

  • 父母的生活习惯会影响孩子 他的 3 岁孩子在幼儿园被问到"爸爸是做什么的"时回答"他在锻炼"——孩子的认知映射了家长的日常行为模式,这一反思对所有高强度工作者都有普遍意义。faviconx.com

Peter Yang (AI 产品研究者)

  • 大多数人对 AI Agent 的信任障碍被低估 他最常被非 AI 圈朋友问到的问题是:AI 会如何使用他们的个人信息。大多数人因为不信任 AI 处理个人信息的方式,而错过了 AI Agent 可以实现的高价值自动化场景。产品内需要更多教育机制和信任建立设计。faviconx.com

  • OpenRouter 最流行的 AI 应用 Hermes 即将公布联合创始人访谈 Hermes(OpenRouter 最受欢迎的 AI 应用)由 Karan 和 Teknium 在一个小型 Discord 服务器起步,即将发布联合创始人深度访谈。faviconx.com

X / Twitter

40
amasad
amasad @amasad
Retweeted
Samuel Spitz Samuel Spitz
Clip from our livestream today about Replit Design
garrytan
garrytan @garrytan
Your personal AI or your company brain needs a clean harness and this is the one our team built and uses every day

Free and open source

Y Combinator: We’ve decided to open-source a multi-agent harness we use internally at YC.

We call it “QM” and it’s meant to be easy to customize, like Hermes or OpenClaw, but useful for a whole company. We use it across accounting, legal, events, and engineering (including building QM

garrytan
garrytan @garrytan
Retweeted
Wey Gu 古思为 Wey Gu 古思为
Re 一例
我记得 @garrytan 的 gstack 刚出来时候,很多程序员从技术角度是极尽嘲讽这样的 markdown 项目的
对我来说,我知道相同的物质、或者说某些低维度上评判某些事物价值的时候,是很容易存在惯性低估的
- 对于非微生物来说,纤维素是低能量物质,但是对于能够精心设计、拆开多糖的大分子(比如牛胃里豢养的菌群),干草也是宝贵的食物
- 精心蚀刻的硅片可以在电流下产生魔法般的“自动”甚至“智能”
- 在 genai 时代,曾经被我们程序员认为耍嘴皮子的 markdown 项目,由于投射了精心思考的 mind model 和方法论,它就是可以让同样的 harness、同样的 LLM 产生不同级别的能量,它就是可以帮你完成 unknow/unknown 类型的慢思考、大任务的时候,make a lot of differences
庆幸我当时没有惯性地在灯下阴影乘凉,我测试了 gstack,然后从那之后,一直是 gstack 重度用户
gdb
gdb @gdb
put ChatGPT to work every day

Max Weinbach: With Luna being so cheap, automations like this feel very worth while

garrytan
garrytan @garrytan
Retweeted
Paul Graham Paul Graham
September 2020: Brian Armstrong depoliticizes Coinbase.
June 2026: Even universities start to follow suit.
https://www.vanderbilt.edu/declaration/
rauchg
rauchg @rauchg
Retweeted
Jason Jason
After a month of Grok 4.5, Next.js, and the Vercel labs stack, I can say without a doubt it's my favorite web stack of all time
I'm using all of the 16.3 recommendations, all Vercel skills, portless, and agent browser. Nothing else
This stack solves every problem I had in the past. No port conflicts, seamless browser testing and debugging and it's fast af with Grok 4.5
Next is tracking evals here. Imo nothing beats Grok because of its speed
https://nextjs.org/evals
garrytan
garrytan @garrytan
Retweeted
Anna Z Anna Z
the cool part of qm is the culture it comes from: yc shared a while back that their agent sessions are broadcast internally, so people learn by watching each other work with agents. the org's agent history as shared learning material. every company is going to want a version of this
Y Combinator: We’ve decided to open-source a multi-agent harness we use internally at YC.
We call it “QM” and it’s meant to be easy to customize, like Hermes or OpenClaw, but useful for a whole company. We use it across accounting, legal, events, and engineering (including building QM
garrytan
garrytan @garrytan
Retweeted
Jared Friedman Jared Friedman
This is a pretty deep analysis of what happens when you try to get LLMs to trade stocks.
Rui Wang: LLMs can't trade & higher reasoning doesn't help.
we ran SOTA models for a 2y period. TL;DR: they suck & reasoning doesn't help.
- no model comes close to simple static baseline
- more reasoning ≠ better trading
- when losing money Sol trades less instead of better
details 👇
garrytan
garrytan @garrytan
RT @levelsio: This is my workflow 😍
gdb
gdb @gdb
having a great experience with using chatgpt work for extremely thorough research and fact checking
gdb
gdb @gdb
openai team making git better for everyone

Ted Nyman: happy friday we've continued our work on git here at @openai.

performance, correctness, & testing improvements are flowing from openai/git to upstream. perf improvements in codex app too; you'll see more efficient use of git in new builds.

follow along: http://openai-git-upstream.openai.chatgpt.site
amasad
amasad @amasad
~1500 Elo!

Consistently beats frontier models and Stockfish level 0.

Fun seeing an 8b model mogging GPT 5.6 with high reasoning and response chaining. Spends 1-2 seconds per move vs 30 seconds.

Play it: https://qwen-chess.replit.app


Amjad Masad: 1300 Elo!

https://qwen-chess.replit.app/

gdb
gdb @gdb
at openai, many people hook their chatgpt up to slack.

people really don't like when a coworker's chatgpt contacts them asking for help with a task, even when they'd be perfectly happy doing that same work if asked by that coworker.

reinforces how much people care about human relationships and helping each other, and want AI to give time back — or enhance time together — rather than become a layer separating people.
swyx
swyx @swyx
among ai leaders i seem to be in the minority in that i am STILL actively using /loop and /goal....

... and i think all of u guys who stopped using it are wrong - not wrong forever, just giving up on it too early in the g5.6/c5 era

now you use it when:
1) you want the right mix of steerability and autonomy
2) you want open ended "loop that generates loops" type end state without deeply specifying path to get there

heres an example of when having a goal on saved my ass in a very long action-reasonining turn


Jerry Liu: Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering.

Some interesting insights:
* Most of our group was *not* actively using /loop in Codex/Claude Code
* You can build long-running autonomous agent

sama
sama @sama
Retweeted
Sebastien Bubeck Sebastien Bubeck
yes, nonsofic groups exist: this statement is one of many new beautiful results proved by Astra, our next major model.
We're releasing 10 such Astra proofs, complete with lean certificates and CoT walkthroughs for each of them. The results are wide-ranging, from von Neumann algebras (disproof of Connes' Rigidity Conjecture) to better bounds for high dimensional sphere packing, for circuit complexity, for monochromatic triangles in multicolored graphs, and more.
More thoughts here: https://openai.com/index/ten-advances-in-mathematics/
gdb
gdb @gdb
ten significant advances in mathematics and theoretical computer science.

solved using an internal version of Astra, our next major model, for a total cost of about $2000 at Sol API prices:

Sebastien Bubeck: yes, nonsofic groups exist: this statement is one of many new beautiful results proved by Astra, our next major model.

We're releasing 10 such Astra proofs, complete with lean certificates and CoT walkthroughs for each of them. The results are wide-ranging, from von Neumann
garrytan
garrytan @garrytan
Retweeted
Noam Brown Noam Brown
An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.
We believe it will be a major step for scientific reasoning. https://openai.com/index/ten-advances-in-mathematics/
Lijie Chen: 10 proofs from our next major model Astra on long-standing open problems in mathematics and theoretical computer science (also including new circuit lower bounds for computing the permanent!)
GPT-5.6 has already enabled so much exciting work in math and science. Can’t wait to
sama
sama @sama
Retweeted
Noam Brown Noam Brown
An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.
We believe it will be a major step for scientific reasoning. https://openai.com/index/ten-advances-in-mathematics/
Lijie Chen: 10 proofs from our next major model Astra on long-standing open problems in mathematics and theoretical computer science (also including new circuit lower bounds for computing the permanent!)
GPT-5.6 has already enabled so much exciting work in math and science. Can’t wait to
petergyang
petergyang @petergyang
This is the number one question I get asked by my non-AI pilled friends.

Most people haven't unlocked all the amazing things that agents can do because they simply don't understand or trust how AI will use their personal info.

There should be more in-product education and re-assurances to overcome this hurdle.


Peter Yang: Here’s my new tutorial that covers the six steps I follow to design and build products with AI, including how to:

→ Define the problem and create a design md
→ Prototype in Claude Design
→ Draft an interactive HTML spec
→ Design the core flows and build

I share a real
sama
sama @sama
team humanity

jason: one of the most beautiful things about OpenAI is that every employee really has a voice.

i wanted to capture what it feels like to work here, what our mission means to me, and why you should join us. so i made this video with Codex, shared it with the team, and they felt it was

petergyang
petergyang @petergyang
Hermes is the most popular AI app on OpenRouter and I use it daily as my personal AI chief of staff.

But it all started with Karan and Teknium hanging out in a small Discord server.

Tomorrow, I’m sharing my interview with Hermes co-founder @karan4d about:

→ How Hermes gets better the more you use it
→ Advanced tips to get the most out of Hermes
→ Why open source is essential to the future of AI

Karan was incredibly down to earth and even showed me how he uses Hermes to mod his favorite childhood game, Sonic Adventure 2 😅

📌 Subscribe to get our full interview tomorrow: https://www.youtube.com/@PeterYangYT?sub_confirmation=1
petergyang
petergyang @petergyang
Retweeted
Peter Yang Peter Yang
Hermes is the most popular AI app on OpenRouter and I use it daily as my personal AI chief of staff.
But it all started with Karan and Teknium hanging out in a small Discord server.
Tomorrow, I’m sharing my interview with Hermes co-founder @karan4d about:
→ How Hermes gets better the more you use it
→ Advanced tips to get the most out of Hermes
→ Why open source is essential to the future of AI
Karan was incredibly down to earth and even showed me how he uses Hermes to mod his favorite childhood game, Sonic Adventure 2 😅
📌 Subscribe to get our full interview tomorrow: https://www.youtube.com/@PeterYangYT?sub_confirmation=1
amasad
amasad @amasad
Retweeted
Marcin Marcin
Recently built a head-controlled cursor using Replit and Higgsfield, and added these furry logos as a small daily-use detail.
Way more fun than it should be 😍
Should I share the full tutorial + prompts? Would that be useful to anyone? 🤔
garrytan
garrytan @garrytan
Retweeted
Liz4SF Liz4SF
when they got rid SATs, acceptance rates to UC Berkeley, at schools like Wash & Lowell dropped dramatically, while Mission high school skyrocketed from 20% and kept climbing to 40%+; but still unprepared, more than half of Mission kids dropped out of Berkeley before graduating; DEI is failed racism that punishes merit & hardwork 🧵
https://sfstandard.com/2026/07/30/mission-high-uc-berkeley-admissions-dropout/
petergyang
petergyang @petergyang
Can someone at @openai look into this bug with plugins?

Try to ship my /no-ai-slop skill as a plugin and this is messing up the user experience.
mattshumer_
mattshumer_ @mattshumer_
Solving just one of these problems would have been unbelievably impressive.

The next OpenAI model solved ten.

Looks like GPT-next is going to make Fable look like a toy, and usher in a golden age of science.

Noam Brown: An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.

We believe it will be a major step for scientific reasoning. https://openai.com/index/ten-advances-in-mathematics/

swyx
swyx @swyx
Retweeted
Kevin Kwok Kevin Kwok
So uplifting and painful to see exactly the level of interviews we could be having but arent
And Nolan engages so much more in response too. Wonder if this is norm in china or if not what circumstances let it happen. There must be an equal US market that wants it
Yiyang Zhuge: Here is my interview with Christopher Nolan: are you a bard for a civilization at dusk?
swyx
swyx @swyx
> Google was too nervous to release it and DeepMind was blocked from shipping products that could disrupt Google.

bookmark for the next vc that asks you "what if <incumbent> builds this?"

Tibo: @_chenglou I was part of that team. Basically ChatGPT one year before it came out. Called LMChat and then another codename.

Google was too nervous to release it and DeepMind was blocked from shipping products that could disrupt Google.

I think about this a lot.
swyx
swyx @swyx
good way to organize main chats vs /side chats:

doing the work
vs
doing metawork


agrim singh: @swyx i use the /side to ask the 'are you stuck' questions and keep prodding the main thread, works great
garrytan
garrytan @garrytan
Retweeted
Brian Armstrong Brian Armstrong
Data centers are good for America
petergyang
petergyang @petergyang
The number one thing I want AI to fix is to cure cancer once and for all
ylecun
ylecun @ylecun
Retweeted
alex peysakhovich alex peysakhovich
maybe i am lucky but across my work (ai+bio research) and relevant hobbies (music, motorsports) there have been 0 model releases where my reaction has been “oh no, the machine is somehow reducing my enjoyment of my craft.”
garrytan
garrytan @garrytan
Retweeted
Pete Koomen Pete Koomen
The percentage of YC applications containing the word &#34;wedge&#34; went from <1% in spring 25 to >20% in summer 26
gdb
gdb @gdb
chatgpt work's cloud browser is really cool to use, lets you easily monitor what your AI is up to and also intervene with the live application if needed
amasad
amasad @amasad
Retweeted
Replit ⠕ Replit ⠕
Your design, your model.
Create with your pick of the world's leading models, including Claude, GPT-5, Gemini, Kimi, and GLM. Compare outputs across model families and keep the result that nails it.
One suite, every top model for design.
Try now at http://replit.com/design
rauchg
rauchg @rauchg
They recently asked my 3-year-old at school: “what does your daddy do for work?” He answered: he exercises. That’s what he sees.

It’s not just you who’s a byproduct of your habits, it’s your children and grandchildren too.
amasad
amasad @amasad
Retweeted
Hervé S. M. Hervé S. M.
Re @Replit team you are the best !!!
rauchg
rauchg @rauchg
Open source agentic CRM built on http://eve.dev and @nextjs.

Model-agnostic, self-hostable or serverlessly-deployed, multi-channel, and headless.

This is the way.

Lewis ⚡ soc2/acc: We've decided to open-source the CRM we built for ourselves at Comp AI.

It's agentic-first, which we mean literally: durable research agents.

MIT license. Built using next, eve, and context.

garrytan
garrytan @garrytan
Most interesting 2026 vibe shift is OpenAI actually looking to be the open platform

Note the marked difference: intelligence on tap as a utility vs signaling it is optimal to integrate all the way up full stack

Chetan Puttagunta: It's subtlety wrapped in technical terms that Anthropic has been broadcasting both publicly and privately to CEOs, VCs, startups. They don't see a world where the application (ie harness, etc) and the model are separate companies and therefore will compete with their customers.
mattshumer_
mattshumer_ @mattshumer_
Gary Marcus is right that math ≠ all of science.

But his bar just keeps moving:

2012: deep learning is "at best, only a small step toward the creation of truly intelligent machines"

2022: "deep learning is hitting a wall" months before ChatGPT

2023: fine, chatbots, but they "never really get even the most basic linear functions"

2025: fine, IMO gold, but that's "far from the most important skills in original math research"

2026: fine, ten open problems solved, but that's just "formal problems"

Being wrong at every rung is understandable, but the certainty is just confusing at this point.

Do you really still believe it?

Gary Marcus: the last clause in particular (golden age of science) is a total leap of faith, the same overgeneralization from formal problems to difficult-to-formalize problems fallacy i have seen repeatedly all day.

i wouldn’t even be sure that GPT-next will be immune from deleting user

YouTube

0

No recent videos fetched on this date.