← 2026-07-14

Daily Edition

2026-07-15

2026-07-16 →

AI Builders 日报 — 7月15日

追踪 AI 领域真正在做事的人,而不是空谈者。

今日思考

今天的信号很清晰:AI 编程工具正在从"copilot"向"autopilot"跃迁。swyx 警告那些低估 CUA(Computer Use Agent)进展的人"不知道自己不知道什么",而 gdb 引用的数据显示 Sol 在 React/前端开发上的成本效率是竞品的 6 倍。这不只是数字——这意味着 AI 编程的临界点正在到来,autopilot 即将成为默认工作方式,而不是少数 early adopter 的玩具。与此同时,Mira Murati 的新公司 Thinking Machines 发布 Inkling 并开放权重,证明开源多模态模型的竞争格局仍在快速演变。


产品与发布

Inkling

Thinking Machines(Mira Murati 创办)发布首个模型 Inkling,从零训练,完全开放权重,支持文本、图像、音频多模态推理,即日起可在 Tinker 上微调,并在 Inkling Playground 直接体验。这家公司的起点直接对标最前沿的开源多模态竞争。faviconx.com

GPT-Red

OpenAI 推出内部自动化红队工具 GPT-Red,专注于规模化发现模型的 prompt injection 安全漏洞,在更广泛部署前构建更强防御。这是模型自我改进循环的又一个工程里程碑。faviconx.com

Vorflux AI

前 Rippling CTO Prasanna 创办 Vorflux,定位为"软件工程的 autopilot"。核心论点是:现有所有 AI 编程工具仍让人类"坐在驾驶座上审批",Vorflux 要把这个角色彻底拿走。Garry Tan 评价"Prasanna is the truth",YC 背景加持。faviconx.com

Vercel Web Analytics API

Vercel 公开 Web Analytics API,允许开发者构建自定义报表和实时用户指标面板,可与 Stripe、Resend 等数据联动分析。rauchg 指出一个酷用例:让 agent 关联访问量、自定义事件与部署记录和性能表现之间的相关性。faviconx.com

Kilo Code 被 Anaconda 收购

Kilo Code 在 16 个月内从零做到 300 万开发者的开源社区,现加入 Anaconda 旗下,覆盖 AI 原生开发的完整技术栈。这是开源 AI 开发工具整合的又一个信号。faviconx.com


观点与判断

swyx(AI Engineer / swyx.io)

  • CUA 正在以难以置信的速度进步 这位从 World of Bits 时代就跟踪 computer use 的开发者警告:如果你还在点头同意"CUA 还早"的观点,你已经严重落后了。他让非技术团队全面使用 CUA 处理各种后台行政工作(注册支付平台、开发票、处理供应商数据),结论是 GPT 5.6 + Superapp 的 CUA 能力已经超过之前所有他能观察到的实现。faviconx.com

  • FDE → ODE → PDE 的玩笑正在成真 Anthropic、Blackstone、Hellman & Friedman 和 Goldman Sachs 联合推出了独立 AI 企业服务公司 Ode(ode.com),直接对应他此前的命名玩笑——Forward Deployed Engineering 之后是 Ordinary Deployed Engineering,下一个应该是 Partial Deployed Engineering。faviconx.com

  • Kilo Code 16 个月 300 万开发者的增长是个谜 swyx 表示自己长期做开发者社区和 YouTube,这个速度对他来说是个谜,很想找 Kilo 的增长负责人聊聊。faviconx.com

garrytan(Garry Tan / Y Combinator)

  • AI 智能体的长期记忆系统三件套 Garry 转发了 shyam 构建的 AI agent 记忆系统,架构组合:Karpathy 的 LLM Wiki 模式(原始信息编译成人类可读、可编辑的 Markdown wiki)+ GBrain 风格图检索(语义相似之外还有关系查询)+ Hermes agent 层(在记忆上查询、推理、自我改进)。最大教训:不要从向量数据库开始,先从一个人类和机器都能读和编辑的事实来源开始。faviconx.com

rauchg(Guillermo Rauch / Vercel CEO)

  • Vercel Agent 是优化问题的答案 问它优化构建时间、性能、账单,它直接动手改。rauchg 推荐把它当作优化类问题的默认工具。faviconx.com

petergyang(Peter Yang / Replit)

  • 不是每个烦恼的工作流都需要 API 或定制 App 他分享的经验:与其把"查看最近 YouTube 视频 + 整理书签 + 提取有价值信息"做成一个自动化工具,不如直接让 Codex 用自然语言操作现有界面。遇到 API 权限和集成麻烦时,换个思路。faviconx.com

技术动态

Fei-Fei Li(Stanford / drfeifei)

  • 机器人模型上下文窗口扩展到 8000 步 Stanford SVL 与 NVIDIA Robotics 合作,将机器人策略模型的上下文扩展到 8000 步(相当于 5 分钟的肌肉记忆),同时保持推理成本不变。过去机器人策略只能在几帧的尺度上生活,这是质的飞跃。faviconx.com

Garry Tan(Garry Tan / Y Combinator)

  • GPT-5.6 一击解决了统计学重大问题 GPT-5.6 通过找到一个反例,解决了 Benjamini-Hochberg (1995) 论文中关于假发现率(FDR)控制的核心开放问题——该论文有 13 万次引用。Garry 指出这又是组合式创新的例子:用另一个领域的工具以训练有素的人类专家可能不会采用的方式解决了问题。faviconx.com

X / Twitter

53
swyx
swyx @swyx
Retweeted
dex dex
Re great ep @DavidOndrej1 @swyx https://www.youtube.com/watch?v=EWk9PBbKqzc
garrytan
garrytan @garrytan
Retweeted
Suhail Suhail
Perplexity is very good at building these benchmarks. Whenever I study them in-depth, it’s always the strongest in the field. DRACO was excellent compared to all the other Deep Research ones, for instance.
Perplexity: We’re open sourcing WANDR.
WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer.
https://research.perplexity.ai/articles/wandr-benchmark-evaluating-research-agents-that-must-search-wide-and-deep
petergyang
petergyang @petergyang
Retweeted
Peter Yang Peter Yang
Tomorrow, I’m sharing a new video on how I use ChatGPT Work (also known as Codex 😅) to do almost everything on my computer.
I’ll cover my complete setup in 7 steps below, from choosing the right GPT-5.6 model to managing email, calendar, and recurring tasks.
📌 Subscribe to get the tutorial tomorrow:
https://www.youtube.com/@PeterYangYT?sub_confirmation=1
petergyang
petergyang @petergyang
Tomorrow, I’m sharing a new video on how I use ChatGPT Work (also known as Codex 😅) to do almost everything on my computer.

I’ll cover my complete setup in 7 steps below, from choosing the right GPT-5.6 model to managing email, calendar, and recurring tasks.

📌 Subscribe to get the tutorial tomorrow:
https://www.youtube.com/@PeterYangYT?sub_confirmation=1
gdb
gdb @gdb
with chatgpt work & sol, i'm finding it incredibly joyful to just ask any question about the business and have it be thoroughly researched and answered.

realizing i have so many questions i wouldn't have bothered asking because they would be too burdensome to answer.
gdb
gdb @gdb
this is just cool

Derrick Choi: Try out the Visualize plugin (in preview) in Codex.

Some ideas click faster when you can see and interact with it.

Here’s a fun example where I asked Codex to build a planet simulator with different controls. 🌎

garrytan
garrytan @garrytan
Retweeted
shyam shyam
Over the last month, I built a long-term memory system for my AI agent.
Not a folder of notes.
Not a vector database with random chunks.
A structured knowledge system that my agent can actually query, reason over, and improve.
The stack combines three ideas:
1. Karpathy’s LLM Wiki pattern
Raw information should not stay raw.
Meetings, notes, links, transcripts, decisions, and documents get compiled into a clean, persistent, human-readable Markdown wiki.
Customers, products, decisions, experiments, and concepts become dedicated pages with citations and backlinks.
The goal is not summarization.
The goal is compounding knowledge.
2. GBrain-style graph retrieval
Once the knowledge lives in structured Markdown, it can become a graph.
Vector search is useful for finding semantically similar text.
But company knowledge is often relational.
“What did customers say about onboarding?” is a semantic question.
“Which customer pain led to which feature, decision, roadmap item, and launch note?” is a graph question.
That requires entities, relationships, timelines, backlinks, and traversal.
My current system has:
1,257 pages
8,778 chunks
951 indexed Markdown documents
semantic search
graph retrieval
typed relationships
timelines
citations
backlinks
3. Hermes as the agent layer
The final step is giving the agent access to this memory.
Instead of starting from zero every session, Hermes can query the wiki and retrieval layers to find old decisions, surface patterns, detect contradictions, update pages, and improve its own workflows over time.
The architecture looks like this:
raw inputs
→ meetings, notes, docs, links, transcripts
markdown wiki
→ structured pages, citations, backlinks
gbrain
→ chunks, entities, graph edges, timelines
qmd/vector index
→ semantic retrieval over markdown
hermes agent
→ long-term memory, synthesis, self-improvement
The biggest lesson:
Don’t start with a vector database.
Start with a source of truth humans can read and edit.
Then add search.
Then add graph structure.
Then give agents access to it.
Most company knowledge bases are digital graveyards.
I think the next generation will look more like operating memory:
human-readable, machine-queryable, graph-connected, and agent-operated.
petergyang
petergyang @petergyang
Retweeted
Hey Hey
learned something useful from @petergyang:
not every annoying workflow needs an api or a custom app.
i wanted to:
• look through the last 10 videos in my youtube history
• review the last 10 bookmarks i saved on x
• extract the useful ideas and information from both
i previously tried turning this into an app or automation.
then i ran into x and youtube access restrictions, api permissions, and all the usual integration headaches.
today i just asked codex to use my computer and handle the administrative work through the interfaces i already use.
it worked like a charm.
thanks Peter!
garrytan
garrytan @garrytan
Retweeted
Parker Conrad Parker Conrad
Huge congrats to my cofounder Prasanna on the launch of his new co!
Prasanna S: Launching @vorfluxai : The autopilot for software engineering. I was prev co-founder / CTO of @Rippling ($10B) and #1 coder in India. Vorflux is my high octane Ferrari.
Every AI coding tool still makes you fly the plane. That's the copilot model: you stay in the seat, approving
garrytan
garrytan @garrytan
Retweeted
T Wolf 🌁 T Wolf 🌁
Drug-free supportive housing passed the San Francisco Board of Supervisors 8-3. Congratulations @mattdorsey! Another win despite the best efforts of the Coalition on Homelessness who opposed this. We're done listening to them. Drug-free is the way!
https://www.sfchronicle.com/sf/article/drug-free-policy-city-funded-supportive-housing-22344024.php
gdb
gdb @gdb
What do you love about Sol, or why did you switch to it?

Tibo: Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched?

Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens!

https://switch-to-codex.openai.chatgpt.site/
garrytan
garrytan @garrytan
Prasanna is the truth

Prasanna S: Launching @vorfluxai : The autopilot for software engineering. I was prev co-founder / CTO of @Rippling ($10B) and #1 coder in India. Vorflux is my high octane Ferrari.

Every AI coding tool still makes you fly the plane. That's the copilot model: you stay in the seat, approving

garrytan
garrytan @garrytan
Retweeted
Prakash Prakash
GPT5.6 one shot solved (by counter example) the central open question of one of the two most important developments in statistics since 1950.. a paper with 130,000 citations.
1) awesome
2) also great that other scientific fields outside ML are moving fast enough for people to public on X immediately (the models can do peer review too after all)
3) yet another result using combinatorial innovation, using a tool from another field in a way a trained human expert might not have
4) yet another result by finding a counter example to a conjecture…
5) wow the pins have started falling eh..
Edgar Dobriban: AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the
garrytan
garrytan @garrytan
Equity over rigor is destroying our public math education in California and increasingly everywhere in the United States


CheesemonkeySF: @garrytan ESSENTIAL BACKGROUND ON THIS STORY:

NOW @thevosf ⁦👇🏼

How California’s math establishment built a generation of students who don’t know what they don’t know.

https://thevoicesf.org/the-layers-of-learning-they-cant-backfill/
ylecun
ylecun @ylecun
Retweeted
Marco Foster Marco Foster
Raphael Warnock: “Donald Trump lost Georgia in 2020. That’s not my opinion, it’s a fact. The votes were counted, recounted, audited, and litigated. He lost, he lost, he lost. He’s trying to sow doubt on the integrity of our elections in Georgia so he can create the pretext to interfere in 2026. This president is a liar, a cheater, and a fraud and he has shown us over and over again that staying in power matters more to him than anything else”
garrytan
garrytan @garrytan
Retweeted
Todd Davis 帅猛男 Todd Davis 帅猛男
We all need to remember who is against having drug free housing in San Francisco.
"Supervisors Shamann Walton, Connie Chan, Jackie Fielder and Chyanne Chen voted against the legislation."
Thank you, @MattDorsey. I know this was a time suck, but well worth it.
T Wolf 🌁: Drug-free supportive housing passed the San Francisco Board of Supervisors 8-3. Congratulations @mattdorsey! Another win despite the best efforts of the Coalition on Homelessness who opposed this. We're done listening to them. Drug-free is the way!
https://www.sfchronicle.com/sf/article/drug-free-policy-city-funded-supportive-housing-22344024.php
garrytan
garrytan @garrytan
Retweeted
Mukund Jha Mukund Jha
Most businesses don’t need just any software. They need a way to turn how they actually work into software
That’s what @emergentlabs is becoming: the operating system for businesses
Today we’re announcing our $130M Series C at a $1.5B valuation to put that future in their hands
amasad
amasad @amasad
Retweeted
Kevin Blumson Kevin Blumson
Thanks to @threejs and @Replit, I made a tamagotchi virtual pet 😀
No 3D models, no textures. Every joint and all 523,800 hairs are generated from pure code in ~250ms. Each hair is individually physics-simulated on the GPU.
https://blumson-surprise.replit.app/
garrytan
garrytan @garrytan
Retweeted
Evil Martians Evil Martians
The first edition of the @sfrubyconf was a complete success.
This year, we're doing it again, applying all the learnings, and making it even better.
Brad called it "one hell of a conference." We say it's an unmissable one:
- A keynote by @garrytan
- Talks by Chris Oliver (@excid3) and Rosa Gutiérrez (@rosapolis)
- Organized networking sessions as part of the agenda
- A full day dedicated to community events
Early bird tickets are officially out. Available for 24 hours (or when they sell out).
Get yours: https://luma.com/sfrubyconf2026
Brad Gessler: It was one hell of a conference.
Not mentioned: thanks everybody @evilmartians who took on the financial risk of this conference and brought it to life.
All I need to do now is sell a few hundred @terminalwire licenses so I can help sponsor the next one. 😅
mattshumer_
mattshumer_ @mattshumer_
My smoke Canadian
My salad parasitic
My building's in a nosedive
NYC in five!
petergyang
petergyang @petergyang
Retweeted
Peter Yang Peter Yang
ChatGPT Work and Codex are growing by 1M users EVERY FEW DAYS and will likely hit 10M this week.
Here's my new tutorial that walks through how I use Work to do (almost) everything, including how to:
→ Choose the right GPT-5.6 model
→ Organize long-running threads
→ Manage your email and calendar
→ Automate recurring work
→ Build and publish Sites
As usual, I've included real examples and prompts for each.
📌 Watch my full 28-min breakdown here: https://www.youtube.com/watch?v=WLg9qWOf6zw
ylecun
ylecun @ylecun
Retweeted
Aaron Rupar Aaron Rupar
OMGGGGG -- Clayton is reduced to sitting in silence instead of acknowledging Joe Biden won the 2020 election
OSSOFF: Who won the 2020 election?
CLAYTON: Uhm, you know, I'm not going to do this with you
OSSOFF: This is a job interview. You have an obligation to be honest with the committee. Who won the 2020 election?
CLAYTON: I'm not gonna get into that with you
OSSOFF: You're not being honest or forthright
CLAYTON: ...
rauchg
rauchg @rauchg
Vercel Agent is excellent for optimization questions. Ask it to optimize your build, performance, your bill.

Adam Killam: Vercel's agent just cut our build times by 10x. @rauchg 🙏
swyx
swyx @swyx
as someone who does a lot of dev community and dev youtube this pace of growth has been one of the greatest mysteries to me because clearly there's something to learn here

@sytses i'd love to talk to whoever was the growth guy/gal at Kilo!!


Kilo: 🚨 BIG NEWS: Kilo Code has been acquired by Anaconda (@anacondainc)!

We've grown our agentic engineering platform from zero to a thriving open-source community of 3M developers in only 16 months.

Now, we're joining Anaconda's trusted foundation to cover the full AI-native dev

gdb
gdb @gdb
our models are built to provide the best price for any given task.

if you're able to get better price/perf on any workload, would love to hear the details and look at it together — gdb@openai.com.
rauchg
rauchg @rauchg
Some really cool usecases of Web Analytics API:

▪️ Ask your agent to correlate visitors, custom events (“purchase”, “checkout”), with the evolution of your deployments and performance

▪️ Build custom frontends, and plot this data alongside e.g.: Stripe’s and Resend’s

Vercel Developers: Web Analytics API is now public.

You can build custom reports and live user-facing metrics with the same data that powers the Web Analytics dashboard. https://vercel.com/changelog/web-analytics-api
swyx
swyx @swyx
calling it now

FDE -> ODE -> PDE

the existence of Forward Deployed Engineering and now Ordinary Deployed Engineering implies that the next pokemon evolution is Partial Deployed Engineering

Andrew Curran: Anthropic, Blackstone, Hellman & Friedman, and Goldman Sachs launched their standalone AI enterprise services firm today. It is named Ode. And it just went live at http://Ode.com. No announcement from Anthropic yet, probably forthcoming.

swyx
swyx @swyx
ok this might be AI Woodstock 2.0

i'd love to do this actually outdoors, in a park with a small sound stage, but dont have contacts. has anyone organized an outdoors event in the Presidio, City Hall, or GGP before? even a small one

cc @NaderLikeLadder @TheAhmadOsman youre gonna have to extend ur stay lmao

clem 🤗: Going to be in San Francisco next week. Should we organize some sort of a meetup or march in support of open-source and local AI?
sama
sama @sama
Retweeted
OpenAI OpenAI
Introducing GPT-Red
An internal automated red teamer on a mission to find our models’ prompt injection vulnerabilities at scale, helping us build stronger defenses before wider deployment.
https://openai.com/index/unlocking-self-improvement-gpt-red/
swyx
swyx @swyx
Retweeted
AI Engineer AI Engineer
🆕This Year In Claude
https://www.youtube.com/watch?v=uU5Gv2h8-9g
@simonw chats with @_catwu and @trq212 about the state of:
- @claudeai Code
- Claude Fable
- @anthropicai culture & product strategy
- Claude Tag & multiplayer collaboration
- The surprising succcess of Remote control
- HTML artifacts for code review
- Why Anthropic uses Auto Mode as the standard for long-running tasks at Anthropic, not --dangerously-skip-permissions
- using Claude for Video editing
- how to be more ambitious as a developer: refusing to "negotiate against oneself"
Timestamps
0:00 Introductions and Claude Code overview
1:22 How coding agents have changed daily workflows
3:51 Shifting focus: Product sense over manual implementation
5:09 Why modern rewrites are now beneficial
6:37 Introducing Claude Tag and team collaboration
11:38 Prioritization and internal "dog-fooding" culture
13:06 The surprise success of remote control features
14:17 Evolving code review processes and automation
17:16 Building trust in new model generations
19:18 Optimizing for capability and user experience
21:23 Reducing system prompts for frontier models
28:05 The philosophy of tool design
30:57 Safety, security, and using Auto Mode
37:53 The human element and developer ambition
41:50 Surprising use cases for Claude (e.g., video editing)
43:35 Limitations and future design aspirations
45:09 Cultural hacks for productivity
46:42 Absurd, fun projects built with Claude
49:03 Audience Q&A
chipro
chipro @chipro
Retweeted
Mira Murati Mira Murati
Our first model, Inkling. Trained from scratch, weights are open, fine-tunable on Tinker today.
Thinking Machines: Today, we are introducing Inkling.
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
https://thinkingmachines.ai/news/introducing-inkling/
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
petergyang
petergyang @petergyang
Retweeted
Thariq Thariq
this was such a fun talk with Cat and Simon, hope you get to check it out if you missed it
AI Engineer: 🆕This Year In Claude
https://www.youtube.com/watch?v=uU5Gv2h8-9g
@simonw chats with @_catwu and @trq212 about the state of:
- @claudeai Code
- Claude Fable
- @anthropicai culture & product strategy
- Claude Tag & multiplayer collaboration
- The surprising succcess of Remote control
- HTML
swyx
swyx @swyx
Retweeted
Thariq Thariq
this was such a fun talk with Cat and Simon, hope you get to check it out if you missed it
AI Engineer: 🆕This Year In Claude
https://www.youtube.com/watch?v=uU5Gv2h8-9g
@simonw chats with @_catwu and @trq212 about the state of:
- @claudeai Code
- Claude Fable
- @anthropicai culture & product strategy
- Claude Tag & multiplayer collaboration
- The surprising succcess of Remote control
- HTML
gdb
gdb @gdb
GPT-Red — improving model security through automated red teaming of prompt injection vulnerabilities:

OpenAI: Introducing GPT-Red

An internal automated red teamer on a mission to find our models’ prompt injection vulnerabilities at scale, helping us build stronger defenses before wider deployment.

https://openai.com/index/unlocking-self-improvement-gpt-red/
gdb
gdb @gdb
something special is happening with Sol:

invincibleHunter: GPT-5.6 Sol is the most impressive model ever. Outside of agentic coding, I find its capabilities in mathematics, especially in vision to be insanely good. I tested both Terra and Sol (Max) on visual mathematics and they are far superior to other models. Huge improvements.
gdb
gdb @gdb
6x price efficiency (!!) with Sol for react/frontend dev:

Aiden Bai: @gdb our benchmark shows that Sol ranks #1 is 6x more cost efficient than Fable across React/frontend work
garrytan
garrytan @garrytan
Retweeted
Ryan Petersen Ryan Petersen
My favorite launch video I've seen in a while. I don't know quite why but I find it so charming and badass at the same time.
Prasanna S: Launching @vorfluxai : The autopilot for software engineering. I was prev co-founder / CTO of @Rippling ($10B) and #1 coder in India. Vorflux is my high octane Ferrari.
Every AI coding tool still makes you fly the plane. That's the copilot model: you stay in the seat, approving
swyx
swyx @swyx
someone just told me about this take* on CUA

this is one of those gell mann moments for me lol. i've been watching computer use since World of Bits (Shi, fan, karpathy, hernandez & liang 2017). we were the first technical pod to interview @jluan about Adept's work three years ago, we were there in the @AnthropicAI building when they first launched Computer Use 2 years ago, I fanboyed over Claude Cowork in our @felixrieseberg pod 3 months ago, and we ran our first full computer use track at @aidotengineer ft. @DhruvBatra_ @proceduralia @francedot 3 weeks ago.

GPT 5.6 + Superapp is even better at CUA than everything i just mentioned. excited for our @AriX podcast to discuss the @skybysoftware story and Codex progress.

if you actually use these things as intensely as we do, CUA is progressing so, so incredibly fast. i have asked my nontechnical team to CUA as much as possible, all their knowledge work with signing up for random payment and invoicing portals and speaker and sponsor and attendee and vendor and union data requests. if you found yourself nodding along to this take below, you are so not up to date that you don't know what you don't know, and underestimating capabilities is quite a dangerous category error if you are doing any ai decisionmaking.

*i admire dwarkesh alot; only criticizing one single take, not the message nor the overall enterprise, screenshot only to share
petergyang
petergyang @petergyang
This game is TENSE it's clear the teams don't like each other lol
petergyang
petergyang @petergyang
It's not over until it's over
sama
sama @sama
amazing to me that some people want the silent version

https://openai.com/supply/co-lab/work-louder/
garrytan
garrytan @garrytan
Retweeted
Tyler Bosmeny Tyler Bosmeny
This is just crazy.
Poetiq has built harnesses that make Opus 4.8 / GPT-5.5 outperform Fable / GPT 5.6 Sol across every benchmark.
Poetiq: Benchmarks are dead (for us).
Our RSI system makes benchmarks too easy.
Give ours a benchmark and it builds its own solution–then beats SOTA. We just did it on 6 at once: math, coding, planning, long-context, tool use, web apps. No human tuning.
drfeifei
drfeifei @drfeifei
I’m very excited by this test time training work for robotic learning! It’s an awesome collaboration between @StanfordSVL and @NVIDIARobotics !

Jim Fan: We scaled a robot model natively to 8,000 timesteps of context, 5 minutes worth of muscle memory, with constant inference cost. Robot policies used to live their lives a few frames at a time (< 0.1 sec), instantly forgetting what just happened. We pushed to 3 orders of magnitude

petergyang
petergyang @petergyang
Argentina almost lost to every single team in the knockout stage yet they kept their composure and didn't give up.

I don't care what people say they're goated.
rauchg
rauchg @rauchg
Easy money (British pounds)
petergyang
petergyang @petergyang
Retweeted
Zach Kram Zach Kram
Timing of Argentina's winners in the knockout rounds:
vs. Cape Verde: 111th minute
vs. Egypt: 92nd minute (stoppage time)
vs. Switzerland: 112th minute
vs. England: 92nd minute (stoppage time)
They haven't led after 90 minutes of any knockout game. But they're in the final.
gdb
gdb @gdb
the new pelican test

Alex: Tell Codex to open Microsoft Paint and try to draw you

petergyang
petergyang @petergyang
This guy led England to lose

FOX Sports: England's manager speaks on his decision to make defensive subs late in the match

garrytan
garrytan @garrytan
Retweeted
Gokhan Egri Gokhan Egri
We just made Thinking Machines Inkling runnable on Claude Code, Codex and OpenCode. We're making it free for the next 24 hours for the first 1000 people that interact with this post, try it on @BrainbaseHQ!
petergyang
petergyang @petergyang
🐐
garrytan
garrytan @garrytan
Retweeted
Kane 謝凱堯 Kane 謝凱堯
Legendary run by New York:
> shut down Indian Point reactor
> increase taxes to cover structural fraud, shrinking tax base
> ban data centers
Addicted to mediocrity.
Governor Hochul Press Office: HOCHUL ENACTS NATION’S FIRST STATEWIDE DATA CENTER MORATORIUM
petergyang
petergyang @petergyang
It has been written 😂


dilemma: Argentina just beat Spain at the 2026 World Cup final, 3-2.
garrytan
garrytan @garrytan
Retweeted
Diana Diana
very cool from @lmsysorg so you can get Inkling OSS model to run fast up to 71.7 tok/s input throughput and 171 tok/s decoding speed
LMSYS Org: Inkling, @thinkymachines' first open model, dropped today: 975B total / 41B active MoE, up to 1M context, reasoning natively over text, images, and audio.
Serving and RL support are already live: you can run and shape it on an open stack, starting now.
Day 0 support on SGLang

YouTube

0

No recent videos fetched on this date.