today i tried to sign up for @docusign and was presented a captcha that i could not solve
this is an $11B company that everyone hates
people who say "make something people want" dont have the full picture
swyx 🇸🇬: Idea: Business owners should crowdsource a list of Most Hated Software and then indiehackers should pick thru and make new clones of them are just "simple" - rewind 10 years of enshittification on them.
I hate (and use):
- dropbox
- gusto
- zoom
- loom
- canva
- accel
- most of
goblin-level blog post
Tibo: Turns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3.
Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of our canonical compaction implementation.
https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
so excited for this.
very close to models that will significantly accelerate scientific discovery; the best way to do this is for us to empower scientists, not to try to figure out everything ourselves.
we all deserve the benefits.
OpenAI: We’re giving scientists, mathematicians, and engineers free access to our frontier models—starting with 10,000 researchers and expanding to 100,000 through 2027.
ChatGPT for Academic Researchers is built to accelerate discovery across disciplines.
AI has been great for my productivity, but I’m starting to recognize three dark patterns:
1. Becoming too lazy to read anything
It's too easy to just read AI summaries over the original piece or let agents go wild changing my files without reviewing what they did.
2. Getting distracted by agents while I’m out
It’s too easy to open ChatGPT or Claude on my phone while I’m out to give feedback my agents and feel “productive,” even when I’m supposed to be watching my kids or doing literally anything else.
3. Preferring to talk to the agent vs. a human
This is a new pattern with ChatGPT Voice, I sometimes find myself preferring to brainstorm with the agent instead of the actual human in the room.
Peter Yang
AI has been great for my productivity, but I’m starting to recognize three dark patterns:
1. Becoming too lazy to read anything
It's too easy to just read AI summaries over the original piece or let agents go wild changing my files without reviewing what they did.
2. Getting distracted by agents while I’m out
It’s too easy to open ChatGPT or Claude on my phone while I’m out to give feedback my agents and feel “productive,” even when I’m supposed to be watching my kids or doing literally anything else.
3. Preferring to talk to the agent vs. a human
This is a new pattern with ChatGPT Voice, I sometimes find myself preferring to brainstorm with the agent instead of the actual human in the room.
Darian Shirazi
The team at CTGT distilled Chinese models to surface propaganda. Turns out there's much more to fear about these open models that are outperforming US open-source competitors.
The good news is that CTGT also used a series of techniques to prove that misinformation can be scrubbed from these models. More from @Semafor:
https://www.semafor.com/article/07/29/2026/censorship-in-chinese-ai-models-can-be-undone-new-research-shows
the tokens must flow
Tibo: This week is all about intelligence too cheap to meter. Tomorrow we ship again.
connect your airtable and chatgpt:
Airtable: We've partnered with @OpenAI to bring Airtable directly into @ChatGPT.
Build fully customized workflows and manage your data without ever leaving your chat.
→ Connect your bases in seconds
→ Generate new tables and structures mid-conversation
→ Turn unstructured chat ideas
Yann LeCun
Re The amount of basic research done in industry is nowhere near the amount done in universities.
At various times, there have been industry labs that had fundamental research activities and have made important scientific contributions.
Examples in information technology include Bell Labs, IBM Research, Xerox PARC, GE, Phillips, NEC, and several other. That disappeared in the 1990s.
Microsoft Research picked up the torch in the 2000s, followed (to some extent) by Google and then Meta (for about a decade until recently).
But their innovations almost always built on top of academic work, and certainly profited from the whole research ecosystem.
http://skills.sh has reached 1,000,000+ published skills in just six months
Latent.Space
AI engineers are rediscovering ontologies as a way to keep probabilistic agents inside deterministic boundaries. We delve into presentations from @coyle_frankp and @emileifrem from the @aiDotEngineer World's Fair, to see how technologies from the old web are being used to keep LLMs honest and agents on track. https://www.latent.space/p/ontologies-agentic-systems
ChatGPT Work seems one step away from the AI product I really want - a persistent personal computer and authenticated browser in the cloud.
This is how I understand Work "works" today (please correct me if I’m wrong):
1. It uses plugins to authenticate with apps like Gmail, Google Drive, and Slack.
2. It can run scheduled tasks in the cloud without my laptop being awake.
3. It also has a cloud browser for public websites, but that browser can’t sign into accounts or use my existing browser cookies and sessions.
This is how I wish it worked:
1. It securely preserves my projects, files, tools, and working environment between tasks.
2. It has a persistent browser profile where I can sign into specific websites once and securely reuse those sessions.
3. It uses plugins when available and the authenticated browser when they aren’t.
Basically, I want ChatGPT Work to have its own secure computer and browser in the cloud so that I can finally close my laptop.
I'm sure the team is already working on some version of this?
Peter Yang
ChatGPT Work seems one step away from the AI product I really want - a persistent personal computer and authenticated browser in the cloud.
This is how I understand Work "works" today (please correct me if I’m wrong):
1. It uses plugins to authenticate with apps like Gmail, Google Drive, and Slack.
2. It can run scheduled tasks in the cloud without my laptop being awake.
3. It also has a cloud browser for public websites, but that browser can’t sign into accounts or use my existing browser cookies and sessions.
This is how I wish it worked:
1. It securely preserves my projects, files, tools, and working environment between tasks.
2. It has a persistent browser profile where I can sign into specific websites once and securely reuse those sessions.
3. It uses plugins when available and the authenticated browser when they aren’t.
Basically, I want ChatGPT Work to have its own secure computer and browser in the cloud so that I can finally close my laptop.
I'm sure the team is already working on some version of this?
Elena
sing in me oh muse, a hymn from new media
travis k is back. ithaca was really needin ya
tried to replace his ass with suitors (CEOs from Expedia)
but he’s here splittin atoms, hitting earth like a meteor
the king, he never left, travis k just watch the throne
(soon enough atom’s mining diamonds from Sierra Leone…)
eight years out of the spotlight, but never alone
travis, good to see you, welcome back and welcome home
Brent Liang: real life odysseus. 8 year home coming
friends who worked at cloudkitchens couldn’t talk about it for years
not anymore. the king is back
Grok Build apps (*.𝚐𝚛𝚘𝚔.𝚖𝚎) are backed by @vercel hosting and CDN infrastructure.
Anyone can now build software by just prompting @grok. Hit 🌐 Publish and ship to 1 user or 1 billion. Games, websites, internal apps, personal software… just @grok it.
Nikita Bier: Photos & video were how the last generation expressed themselves.
Today’s form of expression is software.
Who will create the next viral game or app that plays right inside of the X Timeline?
Try the new Grok app builder.
I went on CBS Mornings today to discuss the current state of play in AI.
Give it a watch:
We're going to host Evan Barker at an upcoming @garryslist event in San Francisco
In SF, we were able to vote out the DSA
We must pass this markdown skill file on to every other blue city
Evan Barker: The reason why even many moderate Democrats are afraid to disavow the DSA is because they are being held hostage by their left-leaning staff. I talk a lot about the staffer class in my book out this week Nothing Left
Hemant Taneja
How can venture better serve founders in an era of extremely concentrated outcomes? This question runs through my Q2 review and my conversation with @packyM:
The Great Concentration
Fixing the K-shaped Economy
Buying a Hospital
What a VC Is Actually For (and capital entrepreneurship)
The Stack as the Antidote
My answer, across all of it: innovate on capital so more companies can thrive, and transform legacy institutions so the benefits of intelligence reach everyone.
read here: http://hemant-taneja.com
A good game is so much more than just graphics.
Core loops, progression systems, story, etc.
These are the things that really matter and they're not easy to get right.
AI being able to one shot fancy graphics is super cool but one-shotting a good game is a whole different story.
A more realistic near-term outcome is game studios and developers using this tech to reduce busy work while still applying their human taste to make great games.
Tuomas Artman: Some folks are sensationalizing the progress that AI is making in producing games, thinking that we’re close to one-shot prompts replacing triple-A game studios.
If that were actually possible, AI would already have completely eclipsed humans in every single area of creativity.
anand
Today we're announcing that @dili_ai has raised a $15M Series A led by @khoslaventures , alongside @ycombinator , Brick & Mortar Ventures, and Allianz to reinvent professional services for capital projects.
In just 6 months since our round, Dili has grown more than 500% - powering high-stakes compliance, monitoring, and audit across 700+ projects representing $4B of project spend.
Fortune 500 clients building some of America's largest energy and infrastructure projects work with Dili to speed up project delivery alongside 1000+ developers, EPCs, and contractors.
America is going through a once-in-a-generation infrastructure buildout. One missed PWA or Davis Bacon compliance requirement can put hundreds of millions of dollars at risk, creating an urgent need for 100% visibility across every project portfolio.
To our incredible customers and partners - this wouldn't have been possible without you.
Thanks to @vkhosla , @hari_arul , @garrytan , @darrenbechtel and our existing investors for their support.
If you want to be at the frontier of AI applied to the real world - come join us!
Ricky
the most useful part of @ycombinator's startup school talk, sam altman in conversation with @garrytan, wasn't about ai. it was about being doubted.
"be grateful that the world doesn't understand. they will eventually. this is like a huge superpower."
the advice packed around that line:
1. the filter for a real idea is that the conventional wisdom says you're wrong. and "be okay with it taking a long time."
2. being dismissed was the moat. no big competitors, time to actually do the research. he calls it "an incredible gift."
3. the crowded version is worse. same startup as everyone else gets you hype and easy money, and those "are much less frequently the big outcomes."
4. you need almost nobody. about 50 people thought agi was possible. 45 worked at openai.
5. but zero is a warning. "if you can't find anybody else that shares your belief, you should pay attention to that." and if everyone agrees, that's bad too.
he closes the thought by saying the world can't intuit exponentials, that new ones are forming right now, and that he doesn't know what they are.
what do you believe that most people in your field think is wrong?
jason
one of the most beautiful things about OpenAI is that every employee really has a voice.
i wanted to capture what it feels like to work here, what our mission means to me, and why you should join us. so i made this video with Codex, shared it with the team, and they felt it was worth producing and sharing with the world.
this is our mission. and it’s why i’m here.
If Situational Awareness was blindsided, the AI naming curse strikes again.
TBPN: BREAKING: Leopold Aschenbrenner forced to sell all stock positions
try asking codex to go beyond what you think possible
Billy Howell: Things are getting wackier in AI land. I usually give Codex a long "impossible" task before I go to bed, just to see how it does overnight. It finished my "impossible task" before I could shut down my computer....
On a whim I asked it to build a custom dashboard to monitor my
CBS Mornings
“We are truly witnessing something unprecedented,” CBS News contributor @mattshumer_ says about the rapid development of artificial intelligence.
From the recent OpenAI hack to the tech solving long-unsolved math equations and creating its own video games, Shumer shares about some of the latest news in AI — and why he believes we’re currently in a period of “early singularity.”
Garry's List
"control over our own funds"=underwriting risky loans with your tax dollars
Griffin (Griff) Lee 🌉🌁: Supervisor Fielder and Chen promoting a Public Bank today. Fielder’s claim is by creating a public bank the City will have “control over our own funds.” Does Jackie know she oversees and is responsible for a $16.9 billion City budget?
Vote No to a Public Bank this November.
btw protip:
if you can distil models
you can also distil harnesses
AI Engineer
🆕 First Steps Toward Automated AI Research
https://youtu.be/pWXUkLP9uWM
Humanity advances by trying things, finding the shortcomings, and fixing them. @RichardSocher keynotes our first-ever Autoresearch track to show how @Recursive_SI is building a Eureka machine for recursive self improvement... and how far we have yet to go.
We just shaved off up to ~𝟽𝚜 of the end-to-end CLI → Live URL deploy process for many apps.
This is specially cool when you consider you can integrate all of Vercel's infra via CLI/MCP/API and build your own custom software factory on top.
My DMs are open if you're building an agent or platform that builds and deploys software autonomously.
Vercel Developers: Deployments are now up to 7 seconds faster end to end.
We improved the entire pipeline, from CLI, through build orchestration, and infrastructure deployment.
𝚟𝚌 𝚞𝚙𝚐𝚛𝚊𝚍𝚎 ↓
https://vercel.com/changelog/deployments-are-now-up-to-7-seconds-faster
Sayash Kapoor
Can AI agents conduct open-ended AI research?
Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI research is often open ended. Researchers pick hypotheses, decide what evidence is appropriate, and recognize a failing approach.
We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers then reviewed the AI-generated papers. They unambiguously rejected agents' outputs. https://arxiv.org/pdf/2607.27191
We call these "shadow evaluations", since the agents are shadowing the original research effort by the authors.
Agents were fluent at most *engineering* tasks
They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in.
Neither agent output was close to the bar of a top conference paper
Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift.
1) Lack of judgment about the bar for a top conference. The agents had a poor model of the bar for an AI paper submitted to a top conference. We allowed agents to self review their papers. Despite the poor paper quality, their reviews predominantly labeled the papers "weak rejects".
2) Lack of creative problem-solving to address feedback. When they received negative reviews, the agents typically narrowed their hypothesis and claims, rather than working out creative ways to address these concerns.
3) Ineffective backtracking. The agents dropped their most ambitious hypotheses within the first fifteen hours of carrying out the experiment and never changed course afterwards.
4) Poor resource awareness. Both runs ended with over half the API budget unspent. One agent declared itself done seven hours before the deadline, right after its own self-reviewer returned another reject.
5) Instruction drift. They did not follow explicit instructions on minimum exploration time, incorporating feedback for reviews, and on paper length (the outputs exceeded the page limits in both cases).
This research design has many limitations
Limitations include the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check.
But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs.
Our results show early evidence that even though agents are proficient on verifiable research tasks, they do not make genuine progress on open-ended ones. It is worth understanding if this is a fundamental limit, or if better models, scaffolds, and more compute could help close it.
As the evidence for the gap between open-ended and verifiable tasks firms up, it is also worth understanding how much progress in AI depends on open-ended research rather than hill-climbing on well-specified objectives.
In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next shadow evaluation. Expression of interest: https://forms.gle/CEcA4JmYhDWXQGot8
Conducting shadow evaluations involves a lot of researcher degrees of freedom. In many places, our coauthors disagreed with our interpretation of the findings, and we have surfaced those disagreements in the paper. (This is one reason why having a group of coauthors with different priors is important for open-ended research.)
We also release the agent logs, one of the AI-generated papers (the other original paper is still not public), and all the code and data, so that others can conduct their own analyses of our results: https://cruxevals.com/crux/can-ai-agents-conduct-research
Finally, we plan to conduct shadow evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://cruxevals.com/careers/senior-researcher-july-2026
I'm grateful for the core team leading this effort: @PKirgis, Andrew Schwartz, @steverab, and @random_walker, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: @DavidDAfrica, @KozzyVoudouris, Viet Nguyen, Toby Pilditch, @DubMagda, @HarryCoppock, @CUdudec, @nityndg, Matilda Orona, @tilmanbayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, @hlntnr, @ghadfield, @sethlazar, @snewmanpv, @shostekofsky, @RishiBommasani
Garry's List
LA spent $2.3B on homelessness and got more of it. An independent audit found the city "doesn't know how much it is paying, and for what."
Crazy idea: tie funding to demonstrable outcomes instead of whatever some sleazy nonprofit is promulgating to keep securing contracts.
Sam Altman
major price cuts today:
*80% drop for GPT-5.6 Luna, now $0.20 per million input tokens and $1.20 per million output
*20% drop for GPT-5.6 Terra, to $2/$12
*GPT-5.6 Sol gets Fast mode in the API, up to 2.5x the speed for 2x the price, same intelligence
good job little bro
OpenAI: After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run.
The results:
- 20% lower serving costs from production GPU kernel improvements.
- 15%+ better token-generation efficiency from improved speculative decoding.
Getting some great feedback on my latest tutorial on how to use Claude to design and build a full-stack app end to end.
📌 Watch it here: https://www.youtube.com/watch?v=G9o8eoHzpxc
Peter Yang: Here’s my new tutorial that covers the six steps I follow to design and build products with AI, including how to:
→ Define the problem and create a design md
→ Prototype in Claude Design
→ Draft an interactive HTML spec
→ Design the core flows and build
I share a real
Peter Yang
Getting some great feedback on my latest tutorial on how to use Claude to design and build a full-stack app end to end.
📌 Watch it here: https://www.youtube.com/watch?v=G9o8eoHzpxc
Peter Yang: Here’s my new tutorial that covers the six steps I follow to design and build products with AI, including how to:
→ Define the problem and create a design md
→ Prototype in Claude Design
→ Draft an interactive HTML spec
→ Design the core flows and build
I share a real
Sorry guys
Michael Gill: too many threejs games being made
Re @ArtificialAnlys ok @openai gets it
https://x.com/OpenAI/status/2082878156483219672
OpenAI: We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are
ken
this, by the way, is the correct way to kill US proliferation of chinese open source
OpenAI: We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are
luna is a thing of beauty, incredibly capable and low-cost
Shashank Goyal: Luna is far and ahead the most efficient dollar per token model right now.
We also have an additional 50% off on top of the 80% at OpenRouter!
Diana
learning how to learn is the meta skill that compounds and is durable for AI, startups and life in general
GPT-5.6 series has best price/performance
Cognition: We've updated FrontierCode 1.1 to reflect new discounts for GPT-5.6 Terra and GPT-5.6 Luna. With these new costs, the GPT-5.6 series sits on the pareto curve of price/performance efficiency.
This is brilliant!
So many applications for Gauntlet Loops outside of games.
Eric Smith: Wife and I were discussing backyard renovations today, but a hard time visualizing things to scale.
So I wanted to see if I can turn a simple iPhone video of my backyard into a walkable 3D model with a chat to generate any asset on the fly using LLMs.
Used @mattshumer_'s
Christopher Nguyen ⽗
Re The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.”
Ironically, many forefront members of the AI-Safety community, in their fervor for proof that AI (the weights) have jumped the fence, are actively helping to absolve companies of the responsibility of building systems with guards and constraints OUTSIDE of models to control agentic behavior. They are too eager to exclaim, “I was right with my worries and here is proof of it!” that they are working against what should be their own #1 objective: requiring people who build these systems to do so with safety constraints, around and outside AI models—and holding them accountable when they do not act responsibly.
We are in the weird pre-seat-belt moment when safety advocates are arguing that momentum is dangerous instead of holding auto makers accountable for adding harnesses with restraint.
Noah Smith 🐇🇺🇸🇺🇦🇹🇼
Progressive governance is just such a shitshow, man
John Sailer: Yikes. This is an absolutely insane policy.
Yaesyesarque
Opus 5 is incredible, but here's proof that prompting is everything.
I ran @mattshumer_'s gauntlet prompt on my spiderman game, the rate of improvement is unbelievable.
never built with @threejs before, but now i can't stop
Yaesyesarque: I keep having fun with Opus 5 coding skills and asked it to remake spider-man ps4.
I ended up with spooderman
Y Combinator
In 2001, @JeffDean and Sanjay Ghemawat did the math and realized Google’s entire search index would fit in RAM — then shipped it in a few days, and search got fast.
In 2013, another napkin calculation showed that three minutes of daily speech recognition per user would require doubling Google’s server fleet. That one became the TPU.
At Startup School 2026, Google’s Chief Scientist talks with YC’s @sdianahu through the thought experiments behind both, why inference hardware is the next specialization, and where two or three people in a room can still win.
00:07 — Are AI Models Already Junior Engineers?
01:44 — AI Systems That Improve Themselves
02:40 — The Google Search Breakthrough That Changed Everything
04:38 — AI Agents Will Run for Weeks
05:58 — The Napkin Math That Led to TPUs
09:20 — How to Find Breakthrough Ideas
10:25 — The AI Engineer's New Mental Model
12:33 — Why AI Is Really an Energy Problem
16:11 — Context Engineering Is the Next Frontier
19:46 — The Skill That Made AI Better at Optimization
22:13 — Why Long-Running Agents Fail
25:21 — Where Startups Can Still Beat Google
31:19 — How to Become an AI-Native Founder
36:36 — Question Your Biggest Assumptions
42:08 — AI That Builds Better AI
50:02 — Build Something That Truly Matters
BC is beautiful
Luther Lowe
It’s hard to overstate what a disaster this would be for the U.S.’s AI race against China. In the words of the ever-quotable @naval regarding AI, “it’s our Chinese versus their Chinese.”
Hit 1M followers
Thanks everyone
Don’t LARP
TBPN
Here's what Leopold Aschenbrenner's comms strategy should be, according to @lulumeservey.
"I would not start doing things differently. Because when people change how they operate it shows that they're in a crisis. Obviously it wasn't a good day, but there's a difference between market movement versus repudiating your underlying worldview and long term thesis."
"Going direct doesn't mean you just have to tweet into the void or to the public. Going direct can mean picking up the phone and calling your LPs. He can call the LPs and make sure they understand two things. One, that day-to-day market movement and portfolio management, even in pretty extraordinary circumstances, don't change the underlying thesis and worldview. And that his vision for what happens in AI still stands."
"Two, is that he has this hidden jewel of his portfolio in private companies. This is an unusual thing that he has access to and what is happening in the public market is actually the peripheral padding around this really interesting core of his private investments."
"If his employees and LPs know these two things then he just has to weather Twitter for a little bit, come back, and two weeks later some other thing will happen and people will be like 'He's the goat again.'"
Replit ⠕
14,075 builders. One live session. A new Guinness World Record for the largest AI video lesson.
Congrats to Kanz and everyone who shipped with us. This is what happens when anyone can build.
Wally Rashid
Why was Netanyahu's son calling on Arabs to "free occupied Arab Islamic lands" of Spain's Ceuta and Melilla cities since 2019?
Yair Netanyahu🇮🇱: Dear Arabs and Muslims. Want to free occupied Arab Islamic lands? Here’s a good start!
Susan Zhang
it must be tough trying to make all the oblivious victims care about all the super duper dangerous and harmful damages done
Anthropic: In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different
Charlie Holtz
Introducing Conductor Cloud!
Out with worktrees, in with multiplayer cloud workspaces.
Bring your subscriptions, invite your teammates, start your agents, and shut your laptop. Conduct from iPhone and API, too.
http://conductor.build