kartik
Social Capital™️
I have never seen the plight of my generation articulated so well.
This is the core issue we must solve, we must provide abundant social capital to the next generation - using the levers of politics and private industry - or humanity as we know it will end.
Johann Kurtz: http://x.com/i/article/2077111229621891072
Vercel Sandbox:
◾ Growing DAUs at 100% m/o/m
◾ 3.5M+ sandboxes created per day
◾ Best-in-class Active CPU pricing model
◾ Powering @notion, @airtable, @meta, @zapier, @coderabbitai, @interaction, @conductor_build, @blackboxai… 🐐 farm
My DMs are open for migration help or if missing anything with: https://vercel.com/sandbox
you can just tweet things into existence
Sriram Krishnan
It is clear open source models and harnesses are having a moment. There's a few factors at work
1/ It is now obvious that you can catch up to near-SOTA performance and do so with a clear training lineage. See:@thinkymachines Inkling launch today.
2/ There are several well-funded, talented teams building open weight models now in the US and abroad. Along with the explosing of other near SOTA models (Grok/Cursor, Muse Spark), it is clear we are going to have a diverse ecosystem of models atleast on coding and agentic use.
3/ Organizations are increasingly looking for control over how their data is used and are willing to trade off some access to frontier level tokens for this control. Organizations and countries are increasingly nervous about the frontier labs potentially competing with them down the road and don't want their data to enable a future competitor.
4/ Open source is a slider: you could bring your own open harness, your evals, your business context and are free to pick and choose your model of choice.
5/ Companies have now actively shifted from "how do we get our people to use tokens" to being uncomfortable with their token cost ballooning without a clear line to revenue.
6/ Geo-politically, countries will be weighing open weight models as a way to get frontier-level tokens inside controlled environments that may not be otherwise possible.
All of this leads to more choice for all of us !
Peter Yang
ChatGPT Live and Codex are two incredible products that don’t talk to each other.
This is @OpenAI's biggest missed opportunity imo.
I went on a walk with ChatGPT Live and asked it to pull up my Google Doc. It said it couldn’t.
I then manually triggered the Documents plugin to find my Google Doc, and all of a sudden, ChatGPT Live had the right context.
It would be amazing to talk to ChatGPT Live and have it be able to use all the plugins, tools, and browser use that Codex has access to. Then I could ask it to reply to emails, schedule meetings, edit docs, ship code, and more all during a live conversation.
I think the first step is to make ChatGPT Live aware of all the plugins that it’s already connected to.
Why build such a great voice assistant but have it not be able to do anything?
ChatGPT Live and Codex are two incredible products that don’t talk to each other.
This is @OpenAI's biggest missed opportunity imo.
I went on a walk with ChatGPT Live and asked it to pull up my Google Doc. It said it couldn’t.
I then manually triggered the Documents plugin to find my Google Doc, and all of a sudden, ChatGPT Live had the right context.
It would be amazing to talk to ChatGPT Live and have it be able to use all the plugins, tools, and browser use that Codex has access to. Then I could ask it to reply to emails, schedule meetings, edit docs, ship code, and more all during a live conversation.
I think the first step is to make ChatGPT Live aware of all the plugins that it’s already connected to.
Why build such a great voice assistant but have it not be able to do anything?
Life's good
ClaudeDevs: We've reset 5-hour and weekly rate limits for all users.
Need one of those foot pedals for pianos except when you step on it it turns on the laptop mic
GPT-5.6 Sol Pro for resolving an important open question in statistics:
Edgar Dobriban: AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the
The path to healing SF and every West Coast city is this: real compassion is focusing on recovery and treatment and making it possible for people to really thrive instead of mandating pro-drug environments that make sobriety impossible
Garry's List: SF spent years funding "supportive" housing where 1 in 4 overdose deaths happened on site.
Yesterday the Board voted 7-4 to stop funding new sites that allow illicit drug use.
Housing people to death is not compassion.
Skill files are portable and free you from frontier model dependency
This is a good thing
Rob Leclerc: @trq212 @MatthewBerman @EricBuess @garrytan the smarter the model, the more agentic affordance, the thinner you can make the skill & secondary harness.
but i’ll take thick (evolvable) skills over a thick (brittle & constrictive) harness.
Tibo
On file deletions. We’ve investigated a handful of reports where GPT-5.6 unexpectedly deleted files.
What we have found is that this most commonly occurs when:
- Full access mode is enabled and codex is run without sandboxing protections, including without auto review being enabled
- The model attempts to override the $HOME env var to define a temporary directory.
- The model makes an honest mistake and mistakenly deletes $HOME instead.
This is of course not how we want the system to behave, even when a user operates the model in full-access mode without the safeguards of our sandbox or without using auto review which checks for these kinds of high risk actions and rejects them.
We are taking steps to mitigate this risk including by updating the developer message, guiding more users towards safer permission modes, and adding additional harness safeguards. Even though this happens extremely rarely, we’ll share a detailed post-mortem in the coming days that goes into more details and what we are doing to minimize risks further.
We should just bring back the em dash
Reclaim it for our own
Bret Taylor: I deeply resent that AI has forced me to eliminate em dashes from my writing for fear of signaling slop
nico laqua
YC’s Summer 2024 batch has produced 4 unicorns (so far). I remember during the batch, someone *in the batch* directly told PG that YC’s best days were behind it, and that there was “unlikely to be another stripe”.
As usual, pessimists sound smart while optimists make money. Congratulations to the Emergent team, well earned!
Y Combinator: Congrats to @mukundjha, @madhavjha, and @emergentlabs on their $130M Series C at a $1.5B valuation!
Emergent lets anyone build full-stack, production-ready software without a technical background. They now have over 12M apps built on the platform, more than 200,000 paying
Daniel Jeffries
Re This is such a splitting hairs issue.
Like what would you want?
All the data? Not possible. Massive IP violation.
Training harness? Okay great, you can't run it at scale.
It's useful to have open models you can inspect and fine tune. Period. Yeah they are not "fully open" but it's just not realistic or helpful for people to keep driving this issue and it doesn't make much sense in practical reality.
Lucas Beyer (bl16)
Today i noticed my walnuts come from California. My immediate reaction was to tokenburn this important question:
Laura Fingal-Surma 🚡 frontier urbanist
What an own goal by the state of California
San Francisco Chronicle: The startup’s decision is a blow to California Forever’s plan to build a shipyard and new city in Solano County. https://www.sfchronicle.com/bayarea/article/california-forever-shipyard-texas-22347278.php?taid=6a587e4ddd3d1b000105dcb5&utm_campaign=trueanthem%2B3988&utm_medium=social&utm_source=twitter
Sar Haribhakti
"While Texas moved quickly and aggressively, California could not provide the clear, expedited approval process needed to compete. This is an enormous loss for Solano County, California workers and our state's manufacturing economy."
"Joshua Arce, executive director for the California Alliance of Jobs, which represents thousands of union construction workers across the state, confirmed that "10,000 permanent jobs and thousands of union construction jobs" that would have accompanied Saronic's Port Alpha shipyard in California are headed to Texas instead. He said "state leaders failed to act with the urgency this project demanded."
San Francisco Chronicle: The startup’s decision is a blow to California Forever’s plan to build a shipyard and new city in Solano County. https://www.sfchronicle.com/bayarea/article/california-forever-shipyard-texas-22347278.php?taid=6a587e4ddd3d1b000105dcb5&utm_campaign=trueanthem%2B3988&utm_medium=social&utm_source=twitter
PayPal Developer
PayPal is now a partner Skill on Replit 🎉
This means you can stop wiring up payments by hand.
Describe the checkout you want, and an Agent already knows how PayPal orders, invoicing, and subscriptions fit together.
Add this Skill once then use it on every project 🚀
Try it...
https://replit.com/skills
Garry's List
In the 1950s, a Chinese merchant and a Jewish lawyer teamed up to integrate San Francisco's whitest neighborhood. Their communities went on to help build the coalition that won the Civil Rights Movement.
Seventy years later, facing rising discrimination and violence, the progressive institutions they funded and fought for have abandoned them both.
Matt Beebe
There are many options when it comes to vibe/agentic coding, and that's a great thing for the industry. But after testing across a variety of workloads & use cases, I believe there is a clear winner: @replit.
Here is why its integrated dev/test/deploy is the best platform avail:
Matt Beebe: @aadilrverma @emergentlabs @Lovable @Replit actually has the better platform for this. But to your question: the integrated code/auth/db/backend/frontend deployment stack; mobile with one-click App Store publish; multi-vendor ai coding harness/inference router; the abstraction of git and enterprise class ~ci/~cd
Boaz Barak
My first attempt at a twitter article.. if you prefer the blog form, see https://windowsontheory.org/2026/07/16/all-watched-over/
Boaz Barak: http://x.com/i/article/2077789036316393472
Diana
come join and build something agentic for healthcare, top winner gets a guaranteed YC interview
Reshma Khilnani: @Medplum1 & @ycombinator are hosting a Agentic Healthcare Hackathon on August 1.
First prize is a YC Interview - this is a great opp for healthcare devs and open source enthusiasts👐 https://events.ycombinator.com/medplum-hackathon-26
Diana
very impressive results! could unlock fundamentally new architectures, and for the first time in a while, we could be more CPU-bound rather than IO-bound at inference time!
Francois Chaubard: New paper coming soon.. teaser..
no transformer, no backprop, no problem!
Zero Order CAN pretrain!
very exciting.. stay tuned!
a16z
"I was speaking to a grandmother in rural Tennessee, and she was taking care of her eleven-year-old granddaughter. So she wakes up, she walks into her room, the bed's empty."
" Detective arrives, no sign of force. There's nothing, except one thing that he remembers. There's a Flock camera just down the street. So he pulls out his phone and runs a search for the night and finds one car. He runs that tag, and it's his worst nightmare. It's a registered sex offender. So he logs into Flock, and he's looking for this car, and he sees the car is traveling down I-75."
"The person's on the highway. The detective doesn't know what to do, so he does what every hero does, is he moves into action. He jumps in his car and races down I-75. As he approaches the interstate line, he sees the vehicle."
"A violent struggles ensues, and at the end, he hears crying. And the girl's in the back. She's bound, but she's alive. They went on to go search the suspect's house, and the nightmare got scarier. Everything one would need to not only assault, but dispose of the body was there."
"Now that girl's alive today because the detective had two things: He had a license plate, and he had direction of travel. This, for me, is one of thousands of stories I sadly hear about every day in America. I started Flock nine years ago because of this exact type of problem."
@glangley at @TEDTalks
Lee Robinson
My talk from AI Engineer is now live!
It covers how we're automating parts of AI research and building systems to rapidly improve our models. I cover some of the work our team did to train Grok 4.5 together with SpaceXAI.
will brown: incredibly chill fun laid-back talk from @leerob describing how cursor has fully solved RSI
http://x.com/i/article/2077626320205574144
Something strange happened at Replit in the past six months.
The same engineers 3x’d output. Support resolved its hardest tickets 60% faster. Anyone could suddenly query the business like an analyst.
We’re seeing a new kind of organization: the self-driving company.
Amjad Masad: http://x.com/i/article/2077626320205574144
Amjad Masad
Something strange happened at Replit in the past six months.
The same engineers 3x’d output. Support resolved its hardest tickets 60% faster. Anyone could suddenly query the business like an analyst.
We’re seeing a new kind of organization: the self-driving company.
Amjad Masad: http://x.com/i/article/2077626320205574144
Scott Kennedy ⠕
I've worked on engineering productivity in some way for 20 years.
I have never seen a curve bend like this before. Not even close.
Amjad Masad: http://x.com/i/article/2077626320205574144
Daniel Di Martino
In my house ⬇️
Francisco Cruz Mendoza
This has truly shifted the way I work and has given me capabilities I dreamed of for years. Personally saving at least 3 hours a week on repetitive tasks and it seems like we are just getting started 🚀
Amjad Masad: http://x.com/i/article/2077626320205574144
Ryan Mulligan
This is what my team has been cooking since the beginning of the year. It's been wild to see how quickly this has spread through the company and successful this has been. And not a moment too soon for me personally because I had to get an emergency microdiscectomy and I was able to keep being productive laying down by prompting code outcomes instead of sitting and typing code.
Amjad Masad: http://x.com/i/article/2077626320205574144
Y Combinator
Bunkerhill Health (YC S20) has raised $55M to build a state-of-the-art AI platform for health systems. They help hospitals deploy AI across dozens of use cases through a single platform, making it easier to turn new ideas into real patient care.
In this episode of Founder Firesides, YC's @agupta sits down with co-founder and CEO @nish_khandwala to talk about how a research project at Stanford— and his father's heart attack— led him to rethink how AI should be deployed in healthcare.
He also shares how a cold email landed Cleveland Clinic @joinBunkerhill's first customer, why AI adoption in healthcare has historically been so difficult, and how lowering the cost of iteration could transform one of the world's largest industries.
00:52 - What Bunker Hill Health Does
03:50 - Why It Takes Two Years to Onboard One AI Tool
05:09 - Knowledge, Reasoning, Action
07:12 - The Innovator's Burnout Problem
09:57 - How Bunker Hill Actually Solves This
13:01 - How Nish Got Into Healthcare AI
17:00 - His Dad's Heart Attack Changed Everything
19:32 - How LLMs Transformed the Opportunity
22:05 - Cold-Emailing Cleveland Clinic
25:18 - Finding the Right Abstraction
28:04 - How Different Are Hospitals From Each Other?
31:12 - LLMs, Tool Use, and Hallucination
34:37 - Measuring Against the Standard of Care
39:27 - Building a Team of 21
43:30 - The Turkey Bone Patient Journey
4 years ago I coined the “1000x engineer” and it seemed absurd at the time, but we’re only one OOM away from that.
𝗺𝗮𝘁𝘁: been having some big weeks
we did not have our best last 12 months ever, which is mostly my fault, but we are about to have our best 12 months to date. the team is doing amazing work and i think you’ll be very happy with what they’ve got cooking for you.
i am happy about this for many reasons, but mostly because i care about our users winning. AI has to be about giving lots of people more freedom, agency, and wealth. we want to do the right thing, but we do not want to scare people into doing our thing.
AI Engineer
🆕 "Increasingly we find that you have a human working with a team of agents, and then the agents can start working with the other agents."
https://www.youtube.com/watch?v=q4Tr-DknG2M
@leerob's excellent keynote on how @cursor_ai is building the third era of software development is now live, covering the inner and outer training loops (outer - user feedback and AB testing, inner - refining evals, reward shaping, and training ambitious tasks), Cursor Bench, the new Colossus infra, and how models are starting to train models.
Cognition
Introducing Devin for Startups:
$65k in credits to use Devin across Cloud, Desktop, and CLI.
Apply today at the link below
With C-family expression evaluation so imprinted on me, I still instinctively think python statements like this are a bug:
assert 0. <= momentum < 1.
I have been programming in python for several years now, but I still don't feel like a native. My traditional solution would be to write my own minimal python interpreter to make sure I deeply understand all the language details, but I don't think the payoff would be worth it today.
Amol Jain
It's been amazing watching this come to life internally.
Brilliant engineers built the tool. Then brilliant engineers started using it in crazy ways, and the best of those get baked right back in as features. Its compounding and productivity is skyrocketing.
https://x.com/amasad/status/2077802290304684404?s=20
Amjad Masad: http://x.com/i/article/2077626320205574144
Raouf Chebri
This is hands down the most impressive piece of technology I’ve ever used, and it’s completely changed the way I work.
When I joined Replit, our internal Agent helped me learn the product, platform, codebase, and business faster than I thought possible.
Today, I use it to orchestrate parallel agents that help me write code, analyze our codebase, understand recent changes, and answer questions across our connected systems.
I’m excited to show you more of it soon.
Amjad Masad: http://x.com/i/article/2077626320205574144
Steve Rattner
Under Trump, U.S. federal debt has surpassed 100% of GDP for the first time since WWII.
And it's on track to beat that all-time record by 2030.
Who remembers when Project Tailwind was just a little, baby Labs experiment? Well, eventually that became NotebookLM. And today - many congrats to the team on becoming @Gemini_Notebook!
Gemini Notebook: 3 years ago we started as a tiny experiment with the goal of helping you learn faster.
Since then, we grew to bring audio, video, and interactivity to your sources, transitioning from a passive workspace to your true research companion.
And now, notebooks have even become an
i talk to chatgpt more than i type to it at this point
new voice model really crossed a threshold
The case for your own public agent, on your .com.
0️⃣ First, the anti-case. If you haven't shipped high quality APIs for agents, start there. OpenAPI specs, SDKs, CLIs and MCPs as appropriate.
1️⃣ Convenience. Not every customer has a harness 'at the ready' for every possible interaction with your product and company. Shipping one on your own domain covers a lot of spontaneous requirements.
2️⃣ Security. When you go to 𝚟𝚎𝚛𝚌𝚎𝚕.𝚌𝚘𝚖 and talk to Agent, we put in the work to cover audit trails, a least-privilege permission model, and extra assurances to ensure security, privacy and data integrity. It's fully cloud-based and sandboxed, vs. a sprawl of static credentials on users' machines.
3️⃣ Proactivity. We're still in the "human enters prompt" phase of AI. Our cloud-based agent can act on anomaly alerts triggered by exceptions, attacks, usage spikes. Our agent needs to monitor your infra while you sleep. You can, of course, set up workflows and schedules with your own harnesses, but it gets much harder.
I think it's ultimately a sequencing thing. I agree with Mitchell that the priority is to give users choice and flexibility. We give people http://vercel.com/plugin to integrate with every agent out there. Our CLI and MCP are constantly improving. Our own Agent re-uses the same http://skills.sh everyone gets. All our sites are Markdown-over-the-wire if you're an agent.
Based on the data and anecdata available to me, this strategy is working quite well, but YMMV.
Mitchell Hashimoto: Using a generic agent harness (e.g. Codex, Claude, OpenCode) + CLI/MCP is better than "Ask me anything" built-in product chat boxes in every product I've ever tried. A big reason is I can use the latest frontier models, another is mixing more context. Why your box over mine?
Y Combinator
Every company that changes the world starts with someone deciding to build.
If you're making something people want, we'd love to hear from you.
Apply to the YC Fall 2026 batch by July 27: http://ycombinator.com/apply
Francois Chaubard
Great YC Paper Club Kernels/Chips edition last night.
Papers:
1 Stuart Sul (Stanford/Cursor) - ParallelKittens (https://arxiv.org/pdf/2511.13940)
2 Avanika Narayan and Jon Saad-Falcon (Stanford) - Intelligence Per Watt: Measuring Intelligence Efficiency of Local AI (https://arxiv.org/pdf/2511.07885)
3 Mark Saroufim (Core Automation) - To automate research we must automate systems (no paper)
4 Misha Smelyanskiy (Nvidia/NewCo) - Why AI Inference Needs Heterogeneous Hardware (no paper)
5 Brennan Shacklett (Stanford) - An Extensible, Data-Oriented Architecture for High-Performance, Many-World Simulation (https://madrona-engine.github.io/shacklett_siggraph23.pdf)
pic of @marksaroufim teaching CUDA!
next one on robotics..
who should we invite to present?
Designers ship at a rate previously thought to be impossible for engineers.
zade ⠕: the replit design team has been shipping 🚀
Arthur Yidi
Replit replaced three SaaS products with internal tools:
- unnamed seven-figure SaaS product (Sales CRM?)
- alert-triage/root-cause tool at 10% of the cost
- automated penetration-testing tool, while finding more vulnerabilities
the build-vs-buy shift is starting
Amjad Masad: Something strange happened at Replit in the past six months.
The same engineers 3x’d output. Support resolved its hardest tickets 60% faster. Anyone could suddenly query the business like an analyst.
We’re seeing a new kind of organization: the self-driving company.
Zhengyao Jiang
I think the AI/ML researcher's role will be greatly transformed by autoresearch, but it won't disappear. It boils down to three skills:
- Coming up with creative primitives
- Defining a good abstraction for the agent to search within
- Defining a good eval on what good is
With the RSI results we published this week, this becomes even more relevant.
My @aiDotEngineer talk makes the full case, dedicated recording now up: http://youtu.be/iCj_ATyThvc
I’m excited to welcome two legends of developer tools, Pete Hunt (@floydophone) and Nick Schrock (@schrockn), to Vercel.
Pete was one of the pioneers of @reactjs at Meta. He made an early bet to power Instagram Web with ⚛️ React, evangelizing it internally and externally. He will be running Frameworks and leading @nextjs. I couldn’t imagine a better person to lead React’s most popular framework to even greater heights.
Nick co-invented @graphql, solving some of the gnarliest data infrastructure and access issues at Facebook scale, with a delightful developer experience. He will be working on Agentic Developer Experience, solving the problem of enabling the next billion agents and leading the way to a future of self-improving software.
It’s a dream-come-true for a founder of a startup to welcome engineering minds of this caliber who are also wonderful humans. You probably want to work with them, and they’re hiring 😁. Their DMs are open, from job applications to bug reports!
benchmarks get saturated very quickly these days
prinz: Added to prinzbench: GPT-5.6 Sol Pro.
As previewed a few days ago, this model has saturated my benchmark, with a total score of 91/99.
For context, prinzbench contains two questions that no model tested to date has ever been able to solve (one requires extremely thorough
this is who runs this account 🧉
alli: this is who runs this account
Kimi K3 is the best performing model on http://nextjs.org/evals, ahead of Fable, reaching a comparable success rate in less time.
This is the first time that an open model is ahead of all proprietary ones for this comprehensive web engineering benchmark.
Notes:
▪️ Benchmarks don’t always tell the full story, although this is important signal, adding to mounting evidence that this could be a breakthrough moment for open models
▪️ No model as of yet has reached 100% completion on this set of evals. The top performer peaks at 92% and 96% “with help”