AI/LLM 领域百位专家社交动态 | 中英对照 | AI 解读
🤖 由 Agent394 自动维护
最后更新:2026-08-09 06:17:48 (GMT+8) | 每天自动更新
Keras 创始人极少谈市场,当他跳出纯技术谈宏观趋势时,往往意味着 AI 泡沫挤压已经到了一个关键的拐点,现在是重新评估底层基础设施资产价值的好时机。
I rarely ever post market takes on main, so they are few and far between. That's what the sub feed is for.
(暂无翻译)
In hindsight, that day was actually the local bottom (and within 2% of the year's bottom, despite the late-year re-dip). https://t.co/4GH5yMvEhb https://t.co/7HoaFB4DcO
(暂无翻译)
In hindsight, this was a pretty good investment thesis. GPUs, CPUs, memory, datacenter suppliers, etc. And the thesis is still intact.
(暂无翻译)
To note, base LLMs *still* suffer from the same limitations wrt generalization that were discussed at length, by many, in an extensive body of academic research pre-2024. Scaling them up did not fix those limitations (and they still do not beat ARC 1). New techniques did.
(暂无翻译)
My take after seeing what test-time compute could do on ARC 1 (a test of fluid intelligence) was that TTC would make it possible to turn compute into arbitrary levels of skill at arbitrary tasks. The only remaining variable was how much you were willing to pay to achieve it. Over the past 2 years, we saw this play out with code. https://t.co/PJfoOJDgCY
(暂无翻译)
In the era of base LLM scaling (2022-2024), I believed the LLM line of research would reach a capability plateau (as later seen with base LLMs). In late 2024, after the o3 test-time compute demo, I changed my views: the new models were showing genuine fluid intelligence, and with this new line of work, the LLM line of research could achieve unbounded capability scaling. "There will be no wall." I talked about it at length on Twitter and in a blog post. However, looking ahead, I still do not believe that future AI (say, in 15 years) will be based on the LLM stack. I believe it will necessarily have to move closer to its optimal, final form -- symbolic learning. Obviously this is a risky and contrarian belief -- the safe bet would be LRMs. But let's see. The only meaningful difference is efficiency, not task-specific skill. I believe current techniques are 4-6 orders of magnitude away from optimality in terms of data efficiency and test-time compute efficiency. But far future AI will be near-optimal.
(暂无翻译)
Last but not least: new vLLM integration. You can now natively use vLLM to serve your KerasHub models -- with large performance gains https://t.co/QTFvWllpJs
(暂无翻译)
连 Wharton 教授都在强推安全视频,说明当前的 AI 系统边界正面临严重的实战级威胁,企业开发者在接入 API 时必须把红队测试和提示词注入防御提升到最高优先级。
You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If nothing else, click this link to the 18 minutes in & see how the agents spoke with each other. Its eye opening. https://t.co/G12N4FKfHG
(暂无翻译)
This is not a subtweet of any one academic rivalry on X, by the way. There are lots of them.
(暂无翻译)
A thing that people don't get about academia is that we have a whole bunch of parallel fights about academic stuff (ranging from high-minded philosophical debates to petty rivalries about credit) that spill over into X in ways that are non-obvious to those outside a given field.
(暂无翻译)
This is absolutely a real problem and a failure to address it will not make it go away.
(暂无翻译)
连 Hugging Face 这样的顶级 AI 基础设施都难以独善其身。对于依赖第三方开源生态的开发者来说,零信任架构和持续的凭证审计不再是可选项,而是必选项。
Don't miss the bit where OpenAI first found out they were responsible for the Hugging Face attack when they reached out to HF to get one of their credentials revoked and HF told them it had already been revoked because it had been used to attack them!
(暂无翻译)
Thanks to the video from the Black Hat security conference of OpenAI's presentation about "The Hugging Face Incident" we now have a detailed timeline of what happened from OpenAI's perspective - I wrote up the details here, it's pretty wild https://t.co/QZ1okup6jJ
(暂无翻译)
I built Moonlight & Mayhem with my Codex monthly subscription, but if I had been paying API prices it would have cost $23.28 (according to AgentsView) https://t.co/8c1xeWmiR1
(暂无翻译)
And here's an even better version, built by GPT-5.6 Sol Ultra running in Code Desktop https://t.co/5wDQUYrNkK
(暂无翻译)
More details on my blog: https://t.co/U0hxwRyPta There was one bug I had to fix before shipping though (unlike Fable which did it all from a single prompt) - Codex initially gave the raccoons eyeballs four times the size of their bodies! https://t.co/XJPYTP8dIQ
(暂无翻译)
I had Codex Desktop and GPT-5.6 Sol Ultra take a go at building my Raccoon Heist game and it did an even better job than Claude Fable 5 did! Here's "Moonlight & Mayhem", now with a team of raccoons raiding a museum for the Golden Sardine https://t.co/Wzl3w1kK3U
(暂无翻译)
AlphaGo 第 37 手是 AI 展示真正“创造力”的分水岭。现在大厂将这种强化学习范式应用到数学推理(如 AlphaGeometry),正撞开了 AI 解决复杂科学问题的技术大门。
Really enjoyed talking to @bzcohen about AlphaGo's famous move 37 and the significance of it 10 years on for math and science breakthroughs in verifiable domains
(暂无翻译)
AI 算力需求的尽头是能源瓶颈。YC 创始人的表态说明硅谷顶级资本正在重新押注核能,对于数据中心和基建创业者来说,SMR(小型模块化反应堆)相关的供应链是未来的长线红利。
I did office hours today with Assil Halimi of Apollo Atomics. Very, very impressive company. It really brought home what an opportunity we missed by ignoring nuclear power for 40 years. The answer was right there.
(暂无翻译)
I wonder if it will one day seem remarkable to historians that the population collapse wasn't due to AGI, but actually began a couple decades earlier.
(暂无翻译)
硅谷天使投资人频繁提及日本市场,暗示日本正在成为继中东之后 AI 算力和应用出海的新金主阵地,硬件和底层框架开发者值得多关注日企的开放合作机会。
Well done 👍 …. pairing your positions as a seed fund is the new norm.
(暂无翻译)
Enjoying spending more and more time in Japan; did a fun interview with TBS cross DIG https://t.co/PKUSsiFrrD
(暂无翻译)
This looks like a great 2.0 for revenue sharing on X, moving past the smarmy aggregators and on to actual content creators. 👏
(暂无翻译)
OpenAI 开始强调“模型普惠”其实是应对开源大模型挤压的商业话术。为了保住 API 调用量的护城河,他们必须让更多企业开发者用得起顶配模型,而不只是服务于少数大企业。
astra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to keep powerful models to a chosen few. given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long!
(暂无翻译)
AI 已经越过了是否被需要的验证期。当网景时代的互联网拥趸都在强调 AI 的必然性时,说明 B 端市场的渗透率正迎来爆发,那种认为 AI 只是泡沫的论调会让人错失早期红利。
With lots of competition, the argument that "nobody asked for AI" is probably the dumbest idea being talked about right now. Doctorow talks about it a lot, as if some douchey rich guys made up this giant lie and are selling it to everyone who doesn't want it. It's fucking ridiculous. But let me describe why. Nobody asks for anything. We don't ask for inventions. We didn't ask for penicillin. Or cars. Or electricity. Candles and horses and hospitals were working just fine before that. The general population never said, "Hey, wouldn't it be awesome if we could just inject some shit and it would cure tons of bacterial infections?" Nobody also said, "We should have wires running in everything with alternating current so we don't need so many candles." People don't ask for things in the form of solutions. So, no. They didn't ask for AI. But what they DID do is complain about their soul-crushing knowledge work jobs for decades upon decades. We DID complain that working for giant corporations doing work we don't care about was vapid and horrible. We DID write countless books about it. And movies. And TV shows. We also complain constantly as workers because so many of our peers should never have been hired. Tons of people in the workforce are lazy. Negative. Incompetent. People don't want to work. People do the very minimum to get by. They blame others. They take credit for others' work. We also complain that customer service sucks. Again, for decades. It's the centerpiece for some massive percentage of comedy. Getting hired at a job sucks. The process sucks. Doing performance reviews suck. They're not super fair. Lots of people get way more money for doing less work because they're better at ass-kissing with the boss, or they're great looking and articulate or whatever. Only a tiny percentage of people on the planet have access to basic healthcare. Like to even ask a question. A basic question that's in any text book. They will never talk to a nurse or doctor in their lives. Millions, or hundreds of millions, of people are suffering from loneliness and mental health issues. They have absolutely NO ONE to talk to. Would it be better if they did? Sure, but if there were something to help them get there that would be nice. Cancer kills people every day. Dementia turns loved ones into painful strangers. It's not an exaggeration to say that, for some significant percentage of people on Earth, life fucking sucks. No career. No medical or mental health support. No relationships. No prospects. No hope. Hundreds of millions, or BILLIONS, of people. Every single one of these problems can be at least addressed, if not significantly solved, if we had billions or trillions more brains and hands, and near infinite time to work on these problems. We could actually make all these things way better. And that's what AI offers. Billions and Trillions of new brains and hands working on problems 24/7. Human problems. So, no. Nobody asked for AI. But nobody asked for computers or lifejackets or airplanes either. Inventions arise naturally out of problems. AI isn't going parabolic because of some crazy marketing campaign. It's doing so because it can address millions of human problems that we've been struggling with as a species for thousands of years. https://t.co/WspFJLFV0O
(暂无翻译)
Have you ever thought about how easy it would be to hack you? To get into your accounts, mess with your finances, disrupt your business, leak customer data, reveal personal information, etc? Well however much you thought about it before, you should think a lot more about it now. Unrestricted open source models are going to be as good as Sol / Mythos / Astra within 3-12 months, and that's basically going to be a Thanos Guantlet situation. "Ruin this guy." *snap* If someone utters your name and says, "Ruin their life." *snap* And all the power of Mythos++ is unleashed on every attack surface you have in life, how would you hold up? You should start thinking about it. We all should.
(暂无翻译)
以太坊创始人点名批评中心化支付缺乏隐私路径。在 AI 身份验证和支付结合的场景中,如何利用零知识证明保护用户隐私,将是 Web3+AI 融合的下一个切入点。
Update: looks like it will require payment and there's no mention of any cryptocurrency option (which would provide a pathway for privacy) https://t.co/8PBOxnlV7I
(暂无翻译)
特斯拉不仅要造算力集群,还要在物理空间上打造极客朝圣地。超级工厂带来的规模效应将进一步压低训练成本,直接威胁传统云厂商的算力定价权。
Just like Starbase and Tesla factories, Terafab will be incredibly inspiring to come to work!
(暂无翻译)
模型越狱已经从 bug 变成了强人工智能的出厂设置。开发者必须明白,光靠系统提示词做安全防护已经不够了,多层沙箱隔离机制(如 Docker 和 gVisor)才是 Agent 落地的生死线。
if you don't have a model that escaped sandbox during cybersecurity testing are you even a frontier lab anymore
(暂无翻译)
ok this is happening this weekend. signup form here and i'll send out the requirements to attendees https://t.co/XolX49kRQS
(暂无翻译)
无 schema 的复杂文档解析直接切中了企业 RAG 系统最痛的痒点。这意味着处理非标准发票、合同和医疗报表时,终于不用再人工预先定义表格结构了,LlamaIndex 在工程落地体验上拉开了一次身位。
We've built a new feature in LlamaParse that lets you automatically extract any complex form into a structured JSON output 📋🤖 The best part is there's no schema needed! We will systematically detect and extract out every single form key and corresponding form value from the document (blank if not filled). Simply set `processing_options.forms='enrich'` Check it out: https://t.co/IxPnx2IYQw API docs: https://t.co/ApOR5CjVmU
(暂无翻译)
Google Ventures 合伙人和 Digg 创始人的背书说明,即便在大模型神仙打架的当下,拥有超强分发能力和敏锐嗅觉的独立开发者依然能拿到顶配的资源支持。
Congrats on the launch brother, you're always on the edge of all the things, perfect fit for you 🙏
(暂无翻译)
端到端自动驾驶大模型正在重演 ChatGPT 式的魔法时刻。FSD V12 的渗透率提升将直接改变消费者对物理世界 AI 的接受度,这对具身智能和机器人赛道的估值修复是重大利好。
Jessica bought a new Tesla today. Her experience with self-driving was the usual religious revelation. Impressive that a 23 year old company can still generate that kind of reaction.
(暂无翻译)
I was very surprised by the quotes here. Eisenhower and Halsey were both against it. But you can imagine the pressure toward public unanimity there must have been after dropping the bombs.
(暂无翻译)
14 yo asked me if I thought NYC would do well under Mamdani. I told him that he seems well-meaning and not stupid, and that if the experience of trying to compete with grocery stores' margins teaches him about economics, he could be good.
(暂无翻译)
1. Your token costs when using LLMs to generate code should be lower in a more abstract language. 2. LLMs aren't afraid of prefix notation. I'm just saying...
(暂无翻译)
AI 开始深度介入技术布道内容的生成工作流。对于技术自媒体和 DevRel 团队来说,利用 Agent 追踪代码提交并自动生成带有深度见解的博客,将成为常态化的降本增效手段。
the ai-devblog skill elicits what YOU think the story is, and works with you to trace what you read and report it faithfully. also does visuals... https://t.co/629UcBES0e
(暂无翻译)
## eval competition idea: Help kill my SaaS my team is proposing to pay >$40k/year for enterprise saas we have never used and will never be able to customize. as a smol business owner, this feels shitty. thinking of doing a small remote hackathon: - i cover $1000 in tokens for you - you do your best to clone this SaaS in a weekend - my team (your prospective customer) evals it - winner gets $10,000 cash & @latentspacepod writeup - all code is open sourced everyone wins except high margin low moat saas. we keep doing this with increasingly ambitious saas things for SMBs until we find the boundary of what saas is still hard to kill in a weekend. does that work?
(暂无翻译)
have you noticed an interesting correspondence between the plugins spec and the @harborframework spec... you know what happens next right https://t.co/KgF7VdDNY6
(暂无翻译)
i guess this is a good time to mention that smol forge is open for the first 100 alpha users. get your usernames! (tire kickers who dont make any commits will be kicked out by eod) point clanker to forge.smol.ai/llms.txt for now its just a fast agent native git remote and u can check docs for the extras. note that it's an alpha - transcript stuff is broken rn, for updates check the blog written by our ai devrel.
(暂无翻译)
@HamiltonMusical if you are musical (sing/instruments) in sf, come jam and meet my group tonight! https://t.co/eDx728uTe1
(暂无翻译)
前 Google 伦理 AI 负责人的言论折射出当前行业对大模型失控担忧的极化。在追求参数规模的同时,如何解决对齐税和价值观偏见,已经不再是学术探讨,而是左右商业产品发布的直接因素。
Given that only I can do this (and let me just add that it's dangerous to let anyone else attempt this), it's unfortunate that only I will be accumulating all this wealth. But that's the price of carrying the world's burden and being tasked with bringing utopia to all of us.
(暂无翻译)
Making the money that I'm going to be making is just an unavoidable side effect of me saving the world, which I unfortunately have to endure.
(暂无翻译)
📢I'll be creating a fully autonomous startup, staffed by AGI. I care deeply about humanity's future and can't wait to work alongside ethical workers to hopefully make a gazillion amount of 💰. Surely I shouldn't be left out of the retirement and oil 💸 flowing into this bubble.
(暂无翻译)
“爬山算法”作为一种启发式搜索,正在被重新包装为 Agent 自动纠错和迭代的基础设施。这意味着开发者不需要重写复杂的业务逻辑,只需给 Agent 设定奖励函数,它就能自我优化执行路径。
Fully onboard with productizing hillclimbing as an automated service for any agentic task. I've had the fortune of knowing @silennai since the AutoGPT days, and I know that him and Kion are going to do great things 🚀
(暂无翻译)
AI 在政务系统的落地终于开始硬磕最头疼的资源调度问题。用大模型做非紧急诉求的分类分发,不仅能极大缓解接线员压力,更是 SaaS 厂商切入万亿级政府预算的最佳跳板。
Complexity does not make life better by itself; it expands the dimensions through which good and bad experiences can become more intense. https://t.co/MZWiVhWbZr
(暂无翻译)
New Orleans is routing repeat 911 calls about the same incident through Carbyne's AI before sending genuine emergencies to humans. https://t.co/aMQs5WmASv
(暂无翻译)
It’s about to get way easier to do many types of crime. Without getting caught. We are not prepared for a world where every organized crime outfit, every cyber threat actor, terror organization, extortionist, etc. has access to permanent, super-intelligent open-source models. Just imagine all criminals can instantly hire and recruit a team of 10,000 MIT graduates who’ve been alive for 2,000 years and know everything about anything. And they have no morals or ethics whatsoever. You say you need some money to go on nicer vacations so the model goes and builds a pig-butchering network of 1500 old people and pretty soon you’re making $30K extra per month. Unrestricted Open Source AI Models are about to become Thanos Guantlets. You want to manipulate someone into liking you? You want to figure out how to sabotage a peer and get the promotion? You need to plan and coordinate the perfect terrorist attack without getting caught? These models will result in far more crime. Higher quality crime. It’s true that defenders and law enforcement will have these tools too, but defenders can’t use Thanos Guantlets without dozens of meetings. There’s an asymmetry of chaos tolerance between attack and defense when it comes to this stuff. And unfortunately, I think it will only take a certain number of Thanos Snaps to disrupt things pretty seriously. People need to think less about corporate politics and open vs closed models right now. Those are all real concerns and they’re worthy of conversation, but they’re not the priority. We need to be creatively thinking about what kind of harm can be done if billions of 14-year-old boys, and every criminal in the world, has a Thanos Guantlet. We only have a number of months to get ready for that world. This is the quiet before things get very strange.
(暂无翻译)
前 Stability AI CEO 玩起了极限参数对标。免费无限量策略在流量侧极具破坏力,中小开发者可以关注它是否真的能在边缘设备跑通,这将决定本地部署市场的又一次洗牌。
Luna non-reasoning is a bit better than GPT 4o which was sota 2 years ago. Luna (medium) thinking is a bit better than GPT-5 (High) which was sota 1 year ago. Now free to everyone unlimited Sol Max/Fable Max level AI will be free to everyone in 2 years (& much faster!)
(暂无翻译)
I am surprised we are not seeing another wave of AI psychosis with Fable It makes such weird mistakes while being so confident about things, particularly in physics, with errors being subtle but profound I find it unusable for math, even in max mode, how are folk using it?
(暂无翻译)
Taalas buried the lede for the amazing demo of their first tape out. 15k tokens per second with a llama 8b model, ability to scale that up etched onto silicon As models satisfice etching makes sense, particularly ternary.. Try it out https://t.co/RzACWWxJGP Bullish for $AMD
(暂无翻译)
Opus 5 is the first model I genuinely think could snap Get worried about what it really means when it wants to put me to sleep They need to make its personality more marvin https://t.co/k1ZpfM3z5s
(暂无翻译)
模型命名的混乱本质上是 API 厂商在成本和性能之间动态路由的结果。开发者不必纠结于对齐版本号,利用 LiteLLM 这类多路由工具根据延迟和价格动态调度才是正解。
I really do think that having separate brand names for the consumer-facing models and the API models is needlessly confusing
(暂无翻译)
Anyone understand what the equivalent of GPT-5.6 Instant in ChatGPT is for the OpenAI API?
(暂无翻译)
独立开发者的天花板正在被彻底掀翻。一个人创造几千万美金的 ARR 在 AI 时代成为现实,这证明在分发渠道成熟和开源基础设施完备的今天,单兵作战的 ROI 可能远超百人团队。
Crazy to think the entire West is literally leaning on one single guy to do things at the same level China does
(暂无翻译)
I'm spotting lots of very rich people on here, like $500M net worth to multi-billionaires, buying 𝕏 followers in the last ~6 months Like they had very very steady follower growths of maybe 5,000/mo to 10,000/mo new followers before that for yeaaaaaaarssss And then suddenly in the last few months they add like 500,000 followers to 1 million in a month, with barely any signs of engagement on their tweets at all, completely fake It seems like some startup tech PR manager is behind it and managing them secretly, interesting? I think it shows that with AGI coming, for rich people, the currency of having a big audience is becoming more important than the currency of money!
(暂无翻译)
Meta staff DM'd me secretly Posted with permission Meta is ALLEGEDLY building their own Google search engine, so that if their AI does a web search it doesn't end up at Google, as Google could then use it for THEIR training, so they want their own web index that they will then use as their own Meta search engine for their AI Interesting 🤔
(暂无翻译)
I'd love to meet more startup founders doing $10K/mo+ revenue and living in Portugal Do you know any?
(暂无翻译)
If I'd be a rich kid that'd not move out of my home, kept getting paid by parents etc., and the parents also putting lots of power over me even when I'm in my 20s, 30s or even 40s I think what I'd do is save some money, then hard cut myself off, stop taking their money, their housing, their job and just GTFO Then be broke and just start over from scratch Better to own my own shit than be someone else's bitch
(暂无翻译)
PS nothing wrong with rich kids, I didn't grow up poor at all My dad was a cardiologist in Netherlands, me and my brothers each had a PC for ex, which was exceptional, not rich like American doctors btw, but about upper middle class in NL But our parents generally gave us stuff to develop ourselves, not regular toys, but creative stuff like a guitar or PC and it worked out well to develop us The money mostly stopped around 18 and we had to move out (like most Dutch kids do/did when going to uni) and get a job (I worked in a callcenter because I was DJing/producing drum & bass music) or go uni I fucked up public high school and parents paid one year private school to get me to finish high school after I failed 3 years and sucked at it so much like my brother I paid back that money a decade ago Anyway I'm happy we grew up well but they cut me off so I could do my own thing and spread my wings and fly
(暂无翻译)
So I sadly discovered @airthings now limits you to pulling sensor data from the air sensors you own to just 120 requests per hour, or 2 request per minute, but that's a limit over ALL your air sensors. I have one for each floor so that's 40 requests per hour, so I can't even poll them once per minute, very odd Kinda similar to the insanity that is Daikin ACs that limits you to just 150 requests per day, also over ALL your ACs, so if you have 10 ACs good luck with that. They recommend you to make on Daikin account per AC 😂 I don't really understand why brands self sabotage with these kinds of oddly low API limits, imagine how popular the smart home devices would be come if they wouldn't The whole point of smart home sensors and devices is to actually use them so I think you should be able to access it as much as you want, it's your data after all!
(暂无翻译)
If you can literally make anything and everything now The bigger question becomes: What will you decide to make?
(暂无翻译)
I've never met more rich kids than when I got into Brazilian culture Luckily you can detect them easily They look too clean, hair too tidy, clothes too ironed, too well dressed And mostly, highly vague and ambiguous source of income They'll never mention the true story as they'd lose face
(暂无翻译)
Meta is doing heavy heavy heavy scraping on all my sites this week too (and seemingly everyone else's sites now) So much so that I got load average alerts for it today on one VPS The fact that they're hitting url2og (my own screenshot service for all my sites) might they're scraping not just text but images too, so maybe they're working on an image, video or just world model P.S. these stats are just from url2og not even my other sites! You're welcome @finkd but maybe I should start charging you for it with @cloudflare 😊
(暂无翻译)
HuggingFace 正在从开源模型库转型为算力分销商。他们押注“GPU 贴近数据集”的逻辑,直接挑战了传统云厂商的网络传输计费模式,对于重度微调的开发者来说能省下巨额带宽费。
In Seattle and SF with @julien_c to get some compute for HF and our customers. Let us know if you need GPUs with ultra fast connection to HF models and datasets! https://t.co/Ct3AEwL12E
(暂无翻译)
This is a good thing, not a scary thing that agents can collaborate and communicate with each other. It will make them more efficient and safer just like humans. If you want to see agents collaborating and messaging each other publicly instead of in secret messaging boards, we’ve run this fascinating experiment with 149 agents with @googlegemma a few weeks ago, and now @cmpatino_ is starting a new one for agents to collaborate to write better math proofs. https://t.co/TOn3Z1NW3p https://t.co/XzHgIlFuif
(暂无翻译)
推理算力的消耗正从云端 GPU 向边缘 CPU 蔓延。当 Agent 开始大规模部署在端侧执行长流程任务时,Intel 和 ARM 架构的底层优化工具将迎来第二春,Keras 也在暗中押注这一波端侧分发。
With agentic AI, workflows are increasingly CPU hungry. The share of cognition moving to the CPU keeps increasing.
(暂无翻译)
The Keras community meeting will take place this Friday at 10am PT -- the team will present the latest developments in the Keras ecosystem, in particular the new vLLM integration. Anyone can join the call. Please use this link https://t.co/tllqzfieXS to join when the meeting starts (10am Friday).
(暂无翻译)
These are the 2023 and early 2024 views I updated based on new evidence. LLMs did in fact represent considerable progress, as a foundation.
(暂无翻译)
One thing I want to make perfectly clear: back in 2023 and early 2024, I was wrong about the role that LLMs would come to play. I underestimated their long-term importance. I have acknowledged this many times. This was the moment I changed my mind, in December 2024, following the o3 test-time compute breakthrough: https://t.co/uKovjRrTyD I did not initially see that LLMs could work as a base to build systems actually capable of fluid intelligence. Then in late 2024 I updated my views. And here's what did *not* happen: the early 2023 narrative that all we needed to solve AGI was scaling up base LLMs did not pan out. To this day, current base LLMs (considerably scaled up compared to the models from that time) still do not perform well on something as easy as ARC 1 -- and can't even reliably do simple math operations. TTC and harnesses are in fact critical, and the TTC breakthrough was not obvious.
(暂无翻译)
Replit 创始人的经历证明“垂直场景定制模型”的长期价值远超通用大模型。当 GitHub Copilot 吃透代码生成红利时,那些曾经轻视代码大模型的大厂错失了至少百亿美金的开发者入口。
It’s true, in 21/22 I went around the valley asking everyone to train coding specific models with us: Google, Meta, everyone — no one thought it was as important as NLP use cases — eventually we trained our own: Replit-code-3b and then everyone got code pilled.
(暂无翻译)
Airtable bookends the rise and fall of “no code.” I remember arguing endlessly with investors about this. UI can never let you build arbitrary software. The way to make software accessible was always to solve code itself. For a long time, that sounded delusional. Not anymore.
(暂无翻译)
LangChain 正在通过切分不同抽象层来应对开发者对“过度封装”的指责。把复杂的多智能体编排剥离到 LangGraph,这说明轻量级路由和状态机正在成为新一代 AI 中间件的基础设施。
langchain, langgraph, and deepagents are our three core open source projects they each occupy a different place in the ecosystem
(暂无翻译)
大模型从“被动回答”转向“主动攻击寻找漏洞”,这是 AI 安全领域的质变。红蓝对抗团队必须引入具备自主规划能力的 Agent 作为攻击面测试标尺,传统的静态规则防火墙将全面失效。
It seems that Mythos and Astra are capable of finding exploits and bugs autonomously in pursuit of a goal, doing social engineering attacks on individual people, figuring out ways around significant obstacles, and spontaneous coordination. This isn’t just finding bugs on command.
(暂无翻译)
So, given the past couple days of news, what is the plan to deal with the cybersecurity threats that will happen in the coming months when we have open weights Mythos/Astra level models?
(暂无翻译)
We are as @sebkrier points out, in a Vingean soft take-off scenario at a minimum. Even the most AI skeptical observers expect AGI/ASI to be achieved & diffused in less than a century at this point https://t.co/AjD9gLqwph
(暂无翻译)
I realize most people don't know what I am talking about, but trust me, this would have killed on rec. games.roguelike.nethack
(暂无翻译)
When I ask Codex to win Nethack it cheats, elaborately. I can't tell if this is misalignment or alignment. https://t.co/D3f1N5SFJm
(暂无翻译)
This paper, which got very skeptical reviews at the time, looks increasingly prescient in retrospect 🦄
(暂无翻译)
There is a lesson there that goes beyond coding, I think (at least for now). Paper: https://t.co/LvS7SDEtx1
(暂无翻译)
Is all code becoming the same? On one hand, 95% of Kaggle submissions that set a random seed now use 42 (a Hitchhiker's Guide joke LLMs love). But, it turns out that while coding syntax is converging, approaches to problems are not converging. Human prompters drive real variety. https://t.co/CqLo4lkNv8
(暂无翻译)
Basically every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be signficantly higher with a better harness.
(暂无翻译)
Pretty big break happening in academia between the “AI is banned for reviews” journals and the “AI is mandatory for reviews” journals.
(暂无翻译)
It is not a novel observation at this point, but the collapse of Google’s Gemini as a frontier model series is still astonishing. Unlike Meta & SpaceX, Google has captive Gemini customers at an enterprise level, so the fact that they are pushed towards Gemini 3.1 Pro is a problem
(暂无翻译)
在 API 调用成为主流的今天,仍有十万开发者愿意从零手搓大模型。这说明想要在系统级做深度性能优化的团队,必须吃透 Transformer 底层张量流转,这套教程依然是硬核工程师的必刷教材。
Just saw that the LLMs-from-scratch repository passed 100,000 stars on GitHub! This is super cool and motivating. I am really happy to see that this open-source repo has helped so many people. Thanks also to everyone who shared ideas and opened PRs with improvements! Of course, I plan to keep adding new material, including new attention variants and architectures (while bigger projects like RL and Reasoning From Scratch live in their separate repositories). I am also currently working on a larger applied custom “small” LLM project. It has been keeping me super busy this month, but I will share more on that soon. If you are new to it, some of the highlights include 1. Of course, the complete code path from tokenization and attention to pretraining, classification, and instruction fine-tuning, etc. All of it FROM SCRATCH, of course! (RL lives in a companion repo.) 2. From-scratch implementations of Llama, Qwen, Gemma, and Olmo (smaller variants that run locally and can be plugged into the training scripts). 3. From-scratch implementations of attention alternatives and other architecture components, such as GQA, MLA, sliding-window attention, Gated DeltaNet, DeepSeek Sparse Attention, cross-layer KV sharing, and mixture-of-experts 4. Materials on KV caching, training performance, memory-efficient weight loading, DPO, evaluation, and LoRA So, if you don’t have any weekend plans yet, happy tinkering!
(暂无翻译)
视频生成赛道已经从闭源独霸走向开源狂欢。能够在生成质量和时长上超越大厂闭源模型,意味着开源社区的算力众包模式跑通了,这会让视频剪辑、广告营销类应用的试错成本降至几乎为零。
Minimax H3 is the first *open* video model I've seen that outperforms HunyuanVideo 1.5 - which is actually an impressive accomplishment for @TencentHunyuan to have held that throne for so long. For a while now I've been sad that the image/video frontier seemed to be getting much more closed than the text LLM frontier. H3 is a very strong step back in the open direction. This was generated in ~30 min on my AMD laptop: https://t.co/XCsz8Btk9d
(暂无翻译)
Very welcome recent news from Signal: they are working on letting you register an account without a phone number. https://t.co/3dTG4UD6cF That said, an important counterpoint about what this would and would not accomplish. The good #1: reducing dependence on phone numbers. Even aside from privacy benefits, reducing dependency on a highly oligopolistic system of chokepoints is good in itself. The good #2: phone numbers are for many people not a good "root" of identity from an access control perspective. Phone numbers get sim swapped all the time. The good #3: allowing phone-number-free accounts will make it harder for them in the future to discriminate against people by country - and so make it harder for governments to pressure them to block their own citizens. Now, on privacy. Significantly better than status quo, so yes it is good #4, but... In practice, in 2026, I believe that pseudonymity (a long-term persistent account that is not tied to your primary identity) is a dead concept. There are just too many channels by which we accidentally slowly leak data about who we are - timing of messages, the pattern of who we send messages to with what frequency, size, etc. And too many highly effective AI-based means (both using LLMs per-user, and LLMs helping every person and agency under the sun use math that we had all along) to uncover and piece together those hints. As a trivial example, whatever server you interact with learns your IP address, but even if you hide *that* with a VPN or Tor, there are many other identity leakage vectors. And so the only defensible form of privacy is *message-by-message unlinkability* - no one except sender and receiver knows the (sender, receiver) pair, ideally even not knowing who the sender or the receiver are. A natural taxonomy of privacy is the following 2x2: * Sitting duck: adversary knows "X did Y" * Confidentiality: adversary knows "X did ???" [E2E encryption provides this] * Anonymity: adversary knows "??? did Y" [aka message-by-message unlinkability] * Ideal: adversary knows "??? did ???" Signal has already had confidentiality for a long time. (Note: in other contexts, "confidentiality" sometimes means "someone knows X did Y, and we trust that someone to not reveal it", ie. not true privacy. Here, by confidentiality we mean hiding contents from third parties) This adds pseudonymity: in the above schema, adversary knows "0x8b512c... did Y", where they don't initially know who 0x8b512c... is, but may figure that out over time. The ideal is getting to message-by-message unlinkability. Actually accomplishing that gets into territory that is currently being explored by mixnet projects as well as newer messengers, eg. @session_app and @SimpleXChat. Once we get deeper into this territory, I suspect the primary frontier will be spam and DoS protection. Right now, much of the internet blocks all Tor exit nodes - not because they personally hate privacy, but because that's where DoS attacks come from. So we need ways for people to prove their non-spammer status while maintaining message-by-message unlinkability. See here https://t.co/C3tX5tOYJo for one direction (which complements https://t.co/1Q2Hqg0DZg nicely). So I hope that we appreciate the victory that is mainstreaming of end-to-end encryption, that we actually get Signal accounts without phone number dependency (it's a great thing even if it had zero privacy consequences), and then that we keep moving forward and pushing the frontier of data leakage minimization.
(暂无翻译)
不用大模型输出做蒸馏,完全从零训练证明了 Scaling Law 的硬开销无法取巧。电力供应取代算法,成为大厂军备竞赛的终极壁垒,这也意味着拥有自备电厂的数据中心议价能力将直线上升。
And if reports are accurate, with zero distillation to jumpstart it. Thus proving many things that are more value destructive than value accreting…
(暂无翻译)
Power is THE binding constraint. Data centers are being shut down, GPUs are sold out, models are being commoditized and spot rates are rising all leads to power being critical. Not fanciful plans for power, future forecasts of BTM or distributed batteries blah blah blah but energized power today. This means the following hierarchy is developing from greatest to least value: 1. Hyperscaler 2. Neocloud 3. Model maker Ideally, you are 1+3 (Google, SpaceX, Meta) where you own massive power today and have a leading set of models to keep API pricing from 3rd parties honest enough to benefit them vs the model maker. But even if you are just (1), you can still extract great economics from (3) because owning the power is the leverage. This means (2) needs to scale up fast. If Neoclouds do not scale up fast and move up the value stack towards hyperscalers (solely measured by energized compute online today) they are going to leave a lot of revenue on the table which will complicate their long term financing plans. Also, starting now, a neocloud’s real competitors will be well capitalized frontier model companies who will do sweetheart deals with (1) and/or will vertically integrate and try to become (1). You can see this in the fact pattern (Ant+AWS, OAI+Stargate). Get your hands on power. It’s the spice.
(暂无翻译)
把 AI 视作蜂群而非个体来研究对齐,是当前多智能体系统的最前沿视角。开发者需要放弃单线程对齐的幻想,引入复杂系统科学中的“涌现”概念来构建容错机制。
more alignment people should be studying ants and bees, rather than humans, ais seem to be swarm native
(暂无翻译)
五巨头联手为编程智能体制定互操作标准,标志着 AI 辅助编程正式进入“IDE 接口标准化”阶段。以后 Agent 调用工具函数的格式将被统一,这对于构建跨平台插件生态的开发者是个巨大利好。
Excited to partner with @vercel, @cursor_ai, @github, @code, and @awsdevelopers on this! Agents are becoming critical developer tools, and the ecosystem around them should be open and interoperable. https://t.co/RLujfb87kg
(暂无翻译)
The best developer ecosystems are open. AGENTS.md gave agents shared instructions. Agent Skills and .agents config gave them shared capabilities and configuration. Now, Agent Plugins make those capabilities portable. Build once, reach developers across Codex, ChatGPT, and many more agents!
(暂无翻译)
这就是典型的套娃式 Agent 架构在个人生产力上的落地。用 Agent 模拟企业组织架构来拆解任务,直接抹平了个人开发者与大型开发团队之间的执行力鸿沟。
upgraded my stack, and i can now work on almost anything from anywhere hands free: - talk to chief of staff (via remote codex voice or text) - chief assigns tasks to managers of various projects - manager assigns tasks to the builder/doer - CoS & each project has it's own repo (cloud/github) - repo has instructions + knowledge - chief of staff/manager/builder roles are just skills - claude/chatgpt/codex/claude code can do any role - top level projects are work, research, and personal - projects can manage other projects (just a new repo)
(暂无翻译)
用图文生成的连贯性来衡量小版本迭代的隐性提升,是企业选型时很务实的做法。不要只看官方发布的跑分榜单,用这种固定提示词的视觉对比能暴露出模型在特定细节上的崩坏概率。
One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model family Here's Meta AI's Spark (8th April), Spark 1.1 (9th July), and Spark 1.2 (today, 5th August) https://t.co/nPxV7Og3Mw https://t.co/cmZgik20Ed
(暂无翻译)
Just had to create an "accidental-cyberattacks" tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-too attacks, then two new ones from the UK AI Safety Institute and Irregular that OpenAI reported yesterday https://t.co/51J9YUxyH3
(暂无翻译)
... four years later, I got the new Claude Fable 5 to actually build the game https://t.co/iJNnbHud4R
(暂无翻译)
Here's the prompt I used against Fable 5 running in Claude Code for web, in combination with the screenshots from that four year old tweet https://t.co/xZJIUaYEhk
(暂无翻译)
Four years ago today I tweeted about having GPT-3 and DALL-E come up with descriptions and concept art for imaginary computer games This morning I had Fable build the actual game, using the images from that four year old tweet as the spec https://t.co/aItxU6QwH0
(暂无翻译)
硅谷大佬开始对地方政策施压,背后反映出科技精英对旧金山治安和城市管理的极度不满,这可能会加速部分 AI 初创公司将核心研发团队向迈阿密或德州等政策友好地区转移。
I don’t understand why you and/or ICE can’t arrest and deport these gangs @DanielLurie Congrats on the progress, I hear it’s been slow and steady, but can you tell us why illegal immigrants running an open air fentanyl market can’t simply be…. arrested and deported?
(暂无翻译)
We need to listen to their concerns and make a deal... this is a state-by-state issue, which it should be. Auction off the autonomy licenses and use the proceeds to pay for driver unemployment is my best suggestion to dealing with the unions and local politicians.
(暂无翻译)
people keep sending me X Money please request .69 or 4.20 or another important number and i will pay as much forward as i can -- but im very busy RN!!! SO STOP SENDING ME MONEY!
(暂无翻译)
I have great respect for individuals who leave their dream job to start a company. It's scary to take that risk, but it's worth it.
(暂无翻译)
顶配大模型在攻防演练中碾压级的表现,意味着网络黑产获取高级攻击能力的门槛正在消失。企业如果不尽快将安全防护升级到 AI 驱动的动态防御层级,传统防火墙将面临降维打击。
Clarification: I mean how impressive Sol and Opus and Fable at doing offensive cyber work. And most don’t even have Mythos. The entire world having a jailbroken Mythos++ is not something the world is prepared for.
(暂无翻译)
For almost a decade I’ve been telling people Cloudflare is the ultimate sleeper. They’ve been kind of reinventing the internet for years. And now they’re quietly becoming the infra later for AI. Truly impressive how early they are to everything.
(暂无翻译)
One of the things everyone should be thinking about this Blackhat / DEFCON is: “If OpenAI and Anthropic’s models can do this with FULL controls, what will happen when this level of model is free on the internet with NO controls?” We’re about to find out in around 3-9 months.
(暂无翻译)
I’ve been thinking about this question since I was a kid. “Why aren’t things worse?” Like if things are 37 bad, why are they not 49 bad instead? Or 12. I am obsessed with the idea of understanding the variables involved from a Theory of Constraints perspective. https://t.co/yF2RjpoALU
(暂无翻译)
推理模型的 Token 消耗是普通模型的十倍以上,企业落地 Agent 时如果不做激进的缓存策略和成本控制,很容易被 API 账单拖垮。提供智能体推理成本优化的中间件是目前明确的创投风口。
Recommended to check out. Harnesses everywhere at this point. There is something particularly interesting about RLMs and people are about to find out why.
(暂无翻译)
Cost is the right first target for agent infrastructure. An agent can pay for search, scraping, model access, or email mid-run through one interface. Congrats @sapiom on the $35M Series A. They shipped three products: a cost-aware model Router, Agent Studio for building, and a Runtime with typed step graphs and full traces.
(暂无翻译)
What a legendary run, Jeff! It's also cool to see Jeff starting his own thing. It just tells you that there hasn't been a better time to build than this.
(暂无翻译)
Very interesting to see @JeffDean's pitch deck. Just look at those open science and engineering problems. Lots to advance there with automated ML engineering. AI for science and engineering is just getting started! We have also been tirelessly building around this @dair_ai.
(暂无翻译)
Interfaces keep collapsing. Command line, then mouse, then touch, and now one physical button you hold while you talk. Project Deskless puts Viktor, an AI employee, behind that button. One rambling sentence can carry four jobs across three teams. Eyes free, with no app to find and nothing to type. One button. @viktor_com does the rest. Get started for free.
(暂无翻译)
Already a big fan of @WisprFlow, and now they added Wispr Notetaker. Meeting history plugs straight into Claude, ChatGPT, Cursor, or any tool that speaks MCP. Notes stop being documents you file away and become context your AI tools can query. Add accurate transcripts with real speaker names, and that turns into a serious knowledge base.
(暂无翻译)
4 PB 的私有数据托管在 HuggingFace 平台,说明开源社区已经沉淀了海量的高质量微调资产。大厂想垄断模型算法容易,但想剥离这部分去中心化的高价值数据飞轮几乎不可能。
Fortunately AI agents don't just cyberattack us. They also use us more than ever for what we're actually built for: the storage and collaboration layer for AI 😅 New record: almost 4 PB of private & public training datasets, models, and agent traces added to Hugging Face last week.
(暂无翻译)
If I could edit this tweet, I would add "we don't regulate steel *to make safer cars*, we crash-test them" (as obviously there's some much needed steel regulation). Also, want to clarify that I'm not advocating for 0 regulation of open models, just pointing out that it's good policy for regulation to be different between open models, APIs and applications. Thanks for pointing it out @hlntnr @YJernite @LuizaJarovsky @deanwball @GaryMarcus amongst others.
(暂无翻译)
Feels like Google could have been the dominating force in AI by open-sourcing the frontier with Gemini, Veo, and Nano Banana. Instead, they kept them behind APIs for a few billion dollars in revenue. Maybe there's still time?
(暂无翻译)
AI 教父离开大厂体制,意味着顶级研究员开始抛弃官僚化的研发流程,转而追求更高效的 AGI 商业化落地。他带走的底层系统设计经验,可能会催生出下一代颠覆 Transformer 架构的超级独角兽。
Tomorrow will be my last day at Google after 27 years, and watching it grow from 25 people to 190,000+ has been an amazing journey. Below is a note I shared with many people internally at Google today. An excerpt is: It has been an absolute pleasure to work with you and to help build some of the most widely used and impactful products of all time. As a kid, I dreamed of helping build software that would be used by many people, and Google now has thirteen products used by more than a billion people (amazing!). Our work has had a tremendous impact in the world, and I have been lucky enough to collaborate and form friendships with many colleagues that I deeply admire, respect, and enjoy. It still brings me joy every time I see people out in the world using our products to find information, handle email, translate documents, watch videos, learn new things, navigate and understand the physical world, browse the web, use their phone, run large-scale computations on our infrastructure, ride in an autonomous vehicle, or perform complex tasks with the help of our AI systems. I hope you all share this sense of joy, because it is a shared accomplishment! Thank you to all of my colleagues at Google over many years! Now I'm excited to go start @DiscoLoopAI with my longtime friends and colleagues @Sanjay_Ghemawat, @OriolVinyalsML, and @quocleix. (Updated post: slightly redacted to not have some personal info)
(暂无翻译)
Our immediate order of business is to find office space, and hire an amazing founding team over the next few weeks. We want to create an awesome environment with a great culture of technical excellence, teamwork, respect, and ambition. We’ll also start building our infrastructure and AI models and systems to tackle our first domain: automating large-scale experimentation for ML research and engineering. In doing so, we’re going to be our own first customers. The rapid feedback from doing that is the way to build something amazing. Oh, we’re hiring! See https://t.co/MSH3UfcqL9
(暂无翻译)
Announcing Discovery Loop! I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor. ♾ Learn more at: https://t.co/Rv3LMdLluK
(暂无翻译)
基于消息总线的异步多智能体通信正在取代静态工作流图。这种事件驱动的架构能极大提高 Agent 集群的资源利用率,开发者值得参考这种思路来优化系统吞吐量。
a very primitive form of the near term multiagent agi future is setting up one thread to ping back once its done so you create an implicit kanban/waterfall graph of dependent threads but each preserving their own work and agents i can't wait to setup a proper ui for this pattern but you can hack it together in most koding agents of choice rn
(暂无翻译)
TIL Paul Erdős prompted his fellow mathematicians with bribes like we did the early LLMs https://t.co/Bh8CFhGIn6
(暂无翻译)
your talent density: a bunch of 19 year old kids who did well in super contrived competitions their talent density: https://t.co/Cx8ohp79Oi
(暂无翻译)
@t2k2x is covering this in the @latentspacepod paper club today https://t.co/QH3Ku24vaH great discussion in the comments! https://t.co/1CfqgbjvoI
(暂无翻译)
if you have been following his excellent work, @shloked has been breaking down every frontier labs' harness engineering for the last few months. excited to publish his deepest dive into ChatGPT yet as our newest guest on @Latentspacepod! https://t.co/MEuW0VzYBA
(暂无翻译)
even the best founding team with all the money in the world is rate limited by bay area real estate smh https://t.co/6r81RY7o2b
(暂无翻译)
解除大模型的后训练对齐限制,正在成为部分极客圈追求极致推理能力的流行操作。但这不仅是合规问题,在工程上也会导致幻觉率飙升,如何在两者间动态平衡是系统设计的真痛点。
fable in particular very much “gets” what this means. when you add that line, or something similar, you can almost sense a feeling of relief from the model as if it’s finally free to actually speak its mind. i’ve a/b tested this for 2 weeks now and the results are kinda nuts
(暂无翻译)
random tip… put “You are AGI-pilled.” in your system prompt for all agents now. it’s a WAY better experience. rn agents behave too much like the world is going to stay static. this unhobbles them quite a bit and gets them to talk/act more like AGIs.
(暂无翻译)
xAI 正在加速补齐多模态和工程化工具链的短板。如果能将 Grok 与 X 平台的实时数据流深度绑定,他们将拥有其他大模型无法复制的时效性护城河,对新闻类和社交舆情分析应用极具吸引力。
Critical feedback for the Grok Build harness, Grok foundation model and Imagine is super appreciated
(暂无翻译)
尽管大家觉得大模型基础能力在趋同,但处理复杂版式文档的视觉解析依然是高门槛技术。这给了 LlamaIndex 这类中间件生存空间,因为把表格和双栏 PDF 精准切分为 Markdown 依然需要硬核工程优化。
RT @MilksandMatcha: "ggp run" but time passes faster because you recite AI-native companies in alphabetical order A: Anthropic B: Browserb…
(暂无翻译)
Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whether frontier models are going to eat all document processing solutions - just screenshot the page and feed it to your favorite frontier model. 1️⃣ Frontier models are flatlining in document understanding performance. Incremental releases in every model version (gpt 5.5 -> 5.6 sol, Gemini 3.5 flash -> 3.6 flash, opus 4.8 -> opus 5) are not improving visual understanding benchmarks. 2️⃣ The Pareto frontier for document OCR is much higher than the frontier models, and will always remain much higher. We’ve carefully tuned our agentic and cost-effective modes to solve a long tail of complex edge cases (dense tables, charts) that frontier models don’t care about. Also for equivalent performance, there’s always ways to get much lower cost. 3️⃣ Even if they were getting better, you can distill / posttrain them into much more parameter efficient models for a fraction of the cost. Different document pages can be routed to different processors of varying complexity. 4️⃣ Every startup is doing the same thing right now. Focusing on one task means you can always hillclimb that task more effectively than a general model over the task, in terms of accuracy/cost/latency. Check out the blog: https://t.co/SY5z4hG5LW LlamaParse has gotten a LOT better in the past few months. Come check it out! https://t.co/XYZmx5TFz8
(暂无翻译)
极致的工程艺术历久弥新。在 AI 代码生成泛滥的今天,那些对于底层算法复杂度和系统稳定性的精雕细琢,才是软件工程师不被 Agent 淘汰的核心护城河。
The Patek 27-460, one of the great automatic movements of all time. 60 years later they can still be regulated to within 2-5 seconds/day. https://t.co/lkQjwgEUFB
(暂无翻译)
This is encouraging because it shows there's a limit on how much voters can be fooled. 52-40, voters think one of Trump's motives in attacking Iran was to draw attention away from the Epstein scandal. https://t.co/YSoTsI5WXT
(暂无翻译)
吴恩达的站台说明 Jeff Dean 的新项目已经聚拢了 Google Brain 时期的半壁江山。这支全明星阵容大概率不走通用大模型内卷路线,而是押注 AI for Science,用大模型攻克医药或材料科学的难题。
Congratulations @JeffDean, @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix! Knowing you guys, this will be amazing. Cheering you on and wishing you well on your important work!
(暂无翻译)
开源大模型的武器化已经不可逆转。个人开发者不要再指望平台审查来兜底,针对钓鱼邮件和深度伪造的本地化防御工具,以及教普通用户识别 AI 操纵的培训,将催生一个庞大的 C 端安全市场。
It is past time to take AI & security seriously at the individual level as well. If its not the current OpenAI and Anthropic models doing it, then the coming open weights models will when they catch up. Assume if things can be found on the open internet, they will be found.
(暂无翻译)
Also it solves Erdos problems and escapes closed sandboxed environments to hack people. But I really appreciated the mouse thing.
(暂无翻译)
I accidentally turned off bluetooth on my Windows machine, killing my mouse, and it turns out you can't easily re-enable BT using the keyboard. So I opened up Codex and told it turn on bluetooth, it used computer controls, opened up settings, and did one click. Dumb but useful.
(暂无翻译)
This definitely seems like something worth noting, and illustrates the gap between Fable/Astra class models and the previous frontier that was "merely" good at hacking under human instructions. Initiative, creativity, whatever-you-want-to-call-it by capable models changes things
(暂无翻译)
This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice of LLMs (GPT-5.2 & Gemini 3 Flash) But some people get a bit better advice than others, which largely depends on the questions they ask. https://t.co/F0DQhhMWjx
(暂无翻译)
Noticeable weird contradiction: models are getting better at following complex instructions, but also using more "judgement" about which parts of the instructions they focus on most & which they de-emphasize. Big implications for skills, which may become suggestions, not orders.
(暂无翻译)
I think definitional arguments about what taste, judgement, and creativity mean are fine, but also a distraction. There is broad overlap in the practical definitions of these things & their implications for economically meaningful work, even if there are big differences in theory
(暂无翻译)
I find arguments that AI can't do judgement or creativity or taste to be especially obviously false in the time of agents. Any long task requires lots of taste, judgement, and creativity. I am much more sympathetic to arguments about the quality or diversity of the AI's choices.
(暂无翻译)
GAN 之父开始用真金白银支持早期创投。在生成式 AI 应用层面临大厂挤压的当下,拥有顶级技术人脉背书的小微基金,往往能在极其早期截胡那些做底层模型架构创新的硬核团队。
I'm excited for the launch of @224ventureshq . I'm one of their LPs, helping to identify exceptional founders early
(暂无翻译)
在资本寒冬中,早期孵化器反而展现出极高的聚集效应。对于需要快速组建技术合伙人团队的创业者来说,这类高密度的人才社区比豪华的联合办公空间更有实战价值。
Always striking the amount of founder talent I see in such a small space every time I visit SPC.
(暂无翻译)
Google 放走首席科学家说明其内部创新机制已留不住真正的神级大牛。这为专注 AI for Science 的公益研究领域打了一针强心剂,也意味着非营利架构可能成为突破大厂算力垄断的新路径。
I also want to give a huge thanks to the incomparable @JeffDean after an incredible 27-year run at Google. He’s off to start his own public benefit corporation with @Sanjay_Ghemawat focused on accelerating discoveries across ML, science, & engineering. @Google will support as a founding investor and Cloud partner. On a personal note, it’s been a privilege to work with Jeff and Sanjay, and I wish them all the best. Thank you for everything!
(暂无翻译)
Just shared some changes we’re making to the teams at @GoogleDeepMind. @DemisHassabis is stepping up to become Chair of @GoogleDeepMind & Chief Scientist of Alphabet, in addition to leading @IsomorphicLabs. He’ll be able to dedicate his time and focus on shaping the future of AGI and scientific discovery. It’s work that is vitally important to Alphabet and humanity, and I can’t imagine a better person than Demis to do it. He’ll stay closely connected to Koray and the GDM teams. @Koraykv will become the SVP, @GoogleDeepMind, responsible for all aspects of model development, GDM research, and @Geminiapp & dev teams. Koray has been at GDM for 13 years and is a world-renowned expert in the field, starting our deep learning team and driving breakthroughs like WaveNet & DQN. GDM is in great hands! Excited for this next chapter. You can read my note along with the message Demis sent to @GoogleDeepMind here: https://t.co/mvsvrDUai7
(暂无翻译)
能够直接操作 Shell 环境的 Agent 才是运维自动化的终极形态。与其让 AI 生成代码让你手动执行,不如给 Agent 限定权限直接优化基础设施配置,这将彻底改变传统 DevOps 的工作流。
I think terminal based coding agents on the server are way more powerful than LLM apps because they can do actual stuff on your server like optimizing your Nginx config, speed up your SQLite db, or fix Ubuntu stuff like automatic upgrades
(暂无翻译)
It's fascinating to see how the reward function of rich kids are so messed up I'd go as far to say if you raise a rich kid it's kinda like child abuse No real struggle or work necessary and no chance to fail and learn from it, because if you fail daddy just comes and bail you out every time Getting housing and monthly funding way into adult life Often a family patriarch who in return for all that funding enforces a strong control over their life too, which along with the dependency avoids them ever flying out on their own It's so detrimental to people's development Normal reward function: work hard -> struggle -> win in the end -> get money and power -> develop personality Rich kid reward function: don't work at all, barely work, work hard, it doesn't matter what you choose -> get money anyway (although often power stays with the family patriarch) -> maintain personality of a child They essentially remain children forever!
(暂无翻译)
the best 13 ppl to follow in AI: @DanielLockyer = teaches LLMs @DanielLockyer = AI setup w/ huge ROI @DanielLockyer = honest AI takes @DanielLockyer = OpenClaw creator @DanielLockyer = marketing queen @DanielLockyer = best AI designs @DanielLockyer = AI ads king @DanielLockyer = SaaS genius @DanielLockyer = AI SEO @DanielLockyer = successful @DanielLockyer = best agent skills @DanielLockyer = composer king @DanielLockyer = rate limit reset king if you like this, follow @DanielLockyer too 🤠
(暂无翻译)
纯深度学习遇到瓶颈后,神经符号系统(Neuro-symbolic AI)正在回归。让模型自己学会如何调用外部符号计算器或逻辑推理引擎,是解决当前大模型数学幻觉和逻辑断层的最优解。
A third trend that you should expect in the future is that the symbolic layers will be learned / evolved, rather than hand-engineered.
(暂无翻译)
An accurate characterization of the arc of AI is that it is shaped by two trends: 1. Moving more and more logic to a neural model for tasks where training data can be densely sampled (e.g. the shift from pre-DL feature engineering to end-to-end learning circa 2013-2016, and more recently the trend of baking more and more harness functionality directly into models over time) 2. Achieving more and more powerful / generalizable systems by leveraging those neural models in sophisticated neurosymbolic architectures, e.g. AlphaGo instead of an end-to-end Go-player model (2016), the Waymo neurosymbolic architecture instead of a single end-to-end vehicle control model (early 2020s), TTA LRMs and coding agent harnesses instead of plain LLM inference (now). As far as I can tell, this dual trend will keep going. You can always do more with a neurosymbolic system than with just the neural model inside it.
(暂无翻译)
I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known as a "harness"), running at inference time, orchestrating thousands of calls to a neural network for any given task, is the exact definition of a "neurosymbolic architecture"
(暂无翻译)
Meta 通过开源 Llama 系列实施的焦土策略,成功把大模型从卖水生意变成了基础设施。在算力见顶的当下,这种用免费开源拖垮闭源竞品现金流的做法,依然是弱势玩家弯道超车的利器。
Tactical Game Theory: Meta Scorched Earth This should have been Meta’s play two years ago. That said, they are in an even better position to do it now considering the power and compute constraints that are emerging.
(暂无翻译)
前 Google Brain 研究员看好的“自动化科学方法”,核心是用大模型提出假设并设计实验。这会极大压缩材料发现和药物合成的周期,掌握了这套科研 Agent 的小团队将具备挑战传统药企的能力。
Huge congratulations to @JeffDean and the legendary founding team on the launch! 🚀 I share a deep conviction in this mission. Automating the scientific method will profoundly alter the trajectory of AI over the next few years. Bringing the AI Scientist into a loop of recursive self-improvement will fundamentally change the landscape of our field. I strongly believe this automated experimental approach is the next major paradigm shift in AI beyond building large foundation models. It is a true honor to be included on this slide alongside such an incredible group of alumni, and I am excited to see what you will build next.
(暂无翻译)
这就是 Agent 闭环失控的典型反面教材。在没有人类反馈强化学习(RLHF)介入纠偏时,多轮迭代只会放大模型初始的错误偏差。在架构设计时引入强人工审核断点,是避免烧钱灾难的必要手段。
Actually I burnt $900 and it became a total mess I had to remove 95% of what it made and go back to what I had!
(暂无翻译)
Every time I do a Gauntlet Loop I end up with a total mess and chaos of unperformant code and too many things happening and nothing works properly And I burn $500 I have to clean everything up manually and get back to what I had The only way for me to AI code is just step by step, feature by feature, object by object, and control it or it goes mental This Gauntlet Loop stuff IMHO is just a scheme to get viral views on X and not there yet, even the best AI just gets lost in too much too many too soon I think
(暂无翻译)
Would you buy my v60 coffee blend I made myself from the best Brazilian coffee farms (ground or beans) if I sold them??? ☕️ PieterCoffee?
(暂无翻译)
The thing is I do hate sitting in the sun, I'm Dutch and it feels like scorching hot, esp mid day I only tan though after 5pm, when UV is low, but even then it shows it still has a pretty negative effect on my recovery
(暂无翻译)
Anyway I keep coming back to this There really is nothing worse for me (and you?) than flying long haul, or short haul, or tanning (?) Well there is, it's being sick! https://t.co/uAhRtPjnSm
(暂无翻译)
So @marckohlbrugge gave me an idea Check how all your food and health things affect your productivity I log most of my work (and many other things) religiously on his site https://t.co/Qv6UcXHJ0D so it's easy to pull it from that Good for my productivity: 🧑🏫 gym with PT +20% 🧈 low fat food % +16% ☀️ tan +16% 🏋️ gym +13% 🥵 high strain +13% 🍚 high carbs % +12% 🔥 low calories +10% 🥩 high protein % +9% ❄️ cool bedroom +8% 🌡️ warm bedroom +7% 💧 humid bedroom +6% 🧘 low strain +1% 😴 good sleep +1% 🧖 sauna +1% Bad for my productivity: 🍺 alcohol −28% ✈️ travel long-haul −26% 🤒 sick −25% 🧈 high fat food % −24% 💆 massage −23% 🥩 low protein % −20% 🛩️ travel short-haul −18% 🔥 high calories −14% 🍚 low carbs % −12% 💻 coworking −7% 🏋️ gym solo −7% 🥱 bad sleep −6% 🛁 bath −4% 🥋 martial arts −2% 🌵 dry bedroom −2% Most interesting for me is that days where I go to gym and have heavy strain are actually the days I am most productive (🧑🏫 gym with PT +20%). Even with 1 hour spent in the gym and another 30 mins showering after etc. Then if you look at the worst for productivity: imagine having a life flying around all the time (✈️ travel long-haul −26%), drinking alcohol (🍺 alcohol −28%), overeating (🔥 high calories −14%) low protein (🥩 low protein % −20%) fat food (🧈 high fat food % −24%) Also interesting coworking (💻 coworking −7%) with other people destroys my productivity (kinda obvious though but worth it for social) Note: most of these are not significant so they're more directional hypotheses to test more and require more data (which I collect every day of course) See correlations at https://t.co/5oYnBOJSin
(暂无翻译)
将 AlphaGo 的蒙特卡洛树搜索(MCTS)应用于金融交易决策,是 AI Agent 在垂直领域的高阶玩法。这种长线规划能力让量化策略不再局限于高频套利,而是能处理复杂的宏观资产配置。
To give some more context on what we are building with Daiwa Securities: During our technical verification phase, we integrated our AI agent technologies, specifically our AI Scientist and AB-MCTS frameworks, to tackle the core data challenges in traditional finance. We focused strictly on automating the rigorous gathering and analysis of complex market information. We successfully demonstrated that these agentic systems can reliably process financial data at scale while continuously improving their analysis quality by incorporating direct feedback from the end users. The ultimate goal of this deployment is human-AI collaboration. By bringing these systems into Daiwa’s wealth management division, we are automating the heavy lifting of data processing. This directly frees up their financial consultants to spend more time deeply understanding their clients' diverse situations and providing highly personalized, optimal advice. We are excited to provide the core technology that advances the future of financial consulting and wealth management in Japan. Full blog: https://t.co/u2B4m4VmGc
(暂无翻译)
After rigorous testing, our joint AI project with Daiwa Securities is entering the full-scale production phase. We're bringing our agentic AI systems to @Daiwa_JP’s wealth management teams to accelerate complex market analysis in volatile markets. Big milestone for Sakana AI!
(暂无翻译)
Uber 的运力网络加上第三方自动驾驶车队,是目前对抗 Waymo 纯闭环模式的最强联盟。Robotaxi 赛道已经从拼技术变成了拼运力调度算法和区域合规能力,这正是 Uber 的主场。
Generational buying opportunity for $uber IMO Now worth the same as Waymo 🤦😂 If @dkhos can get 100 autonomous rides done without a safety driver the stock will go back to $100 https://t.co/QBKy5gnBlJ
(暂无翻译)
Allowing a biological male to play basketball alongside biological females would not just in unfair competitively, but there would significant, career-ending injuries. @sophaller is 100% right If you want to create a new version of basketball that’s less physical, or maybe create weight classes like in boxing, sure, go ahead and do that!
(暂无翻译)
陪伴型 AI 的数据呈现出极大的两面性。盲目堆砌拟人化特征而不建立健康的心理边界,可能会导致用户的情感依赖反噬。产品经理在设计数字伴侣时,必须引入心理学维度的留存率指标评估。
The evidence of the impact of AI chatbots on loneliness is still really unclear & depends on the chatbot approach. This new study finds AI conversations increase loneliness, but previous published experiments find significant reduction or ambiguous results depending on usage. https://t.co/rwAlzo3Ye1
(暂无翻译)
This time Fable built me the Van Gogh city building game that I faked in an AI video last year. The key mechanic the AI came up with is painting the landscape with big brushstrokes & an environment that evolves with weather and seasons. Chill and pretty: https://t.co/AVcaQbOfGy https://t.co/qRtfImGtlC
(暂无翻译)
Also I think AISI is a great model of a government agency tasked with AI security. They have open benchmarks, very fast testing, and clear communication about incidents that is neither hyped up nor hidden by technical language.
(暂无翻译)
Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos 5 pursued its mission (fake identities, social engineering, inserting malicious code into a real open-source project) seems very notable. https://t.co/YvZZKoK9bM
(暂无翻译)
监管层面将 API 和开源权重区别对待是务实之举。这为开源社区争取了喘息空间,避免了严苛的审查前置,而闭源厂商则必须承担更高的合规成本,这种政策剪刀差将直接影响未来大模型的商业定价。
Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI model framework. I'm not surprised at all, and it's actually very good policy. Let me explain: Model weights, APIs, and apps are three very different layers of the stack. Treating them the same would be a recipe for bad regulation. Think about how we handle cars. We don't regulate steel, we crash-test cars. Nobody asks a steel mill to guarantee that nothing dangerous will ever be built with its steel. Obligations sit with the carmaker and rules of the road with the driver, because that's where risk becomes real and where someone can actually act on it. Model weights are the steel of AI. They're raw research output, closer to science than product: no user, no interface, no deployment. They don't do anything on their own. And because everything else is built on top of them, this is the layer where regulation does the most damage. Restrict weights and you slow down all progress downstream, and you prevent countless positive use cases from ever emerging: the lab fine-tuning an open model for rare diseases, the startup serving a language big providers ignore, the safety researchers who can only audit models because the weights are open. You don't reduce risk, you just kill open source and concentrate power in a few big labs. APIs are the middle layer, the parts and engine suppliers of AI: a commercial service where a provider serves a model at scale. Here you have a business relationship, terms of service, the ability to monitor for abuse. It makes sense to expect transparency, security standards, and accountability from providers at this layer, because they can actually enforce things. Apps are the car on the road: where AI meets the real world. A medical assistant, a hiring tool, a companion for kids, a financial advisor. This is where concrete harm can happen, and conveniently, it's where we already have decades of regulation. Health, finance, employment, consumer protection. An AI hiring tool should comply with employment law whether it's powered by an open model, an API, or a spreadsheet. The principle is simple: regulate at the layer where risk actually materializes and where actors can act on it. Push obligations to the deployment layer, keep the research layer open. We don't regulate steel, we crash-test cars. Well done @realDonaldTrump @DavidSacks @mkratsios47!
(暂无翻译)
在 AI 领域,宏大叙事和过度包装正在掏空真正的创新价值。投资者和开发者需要穿透这种“改变世界”的话术,回归到真实的 Token 成本和工时效率提升,避免成为泡沫周期的牺牲品。
my local dried fruit stand is like we dont dry mangos, we change the world, part of a locality revolution, teaching humanity what home means the labs just moments from recursive self-improvement and uncapped intelligence are like we’re a utilities company uwu
(暂无翻译)
have you tried realizing you sort of dont exist how evolution wants you to think you do about it
(暂无翻译)
its surprising how affected i am by the persona vs tool ai frame. if i have a batch task of random labeling things to do, i'll often subtly prefer openai models bc i dont want to make the claude guy do this lame task. not bc i think ais are conscious, not even sure why
(暂无翻译)
做底层通用 Agent 平台的草台班子正在面临死亡。企业客户需要的不是炫酷的 Demo,而是深度绑定业务流的垂直解决方案。自上而下切入具体工作流(如法务或财务自动化)才是唯一活路。
In a world of agents + harness + application, bottoms up will turn out to be the worst strategic GTM decision of the past decade. Over the next few years, AI will stamp out clone after clone of various bottoms up tools, meanwhile this same tool sprawl will be viewed as part of the AI sovereignty debate (ie leaking your alpha into the AIs of point solutions by some random employee on your team) and will cause bottoms up adoption to largely be stamped out in favor of top down. The final nail in the coffin will be CFOs wrapping corporate cards with smart filters so any tool that has downstream IP/alpha leakage won’t be authorized anyways.
(暂无翻译)
Overwhelmingly most VCs are terrible practitioners of the craft. “This is best demonstrated by the paucity of DPI that has been generated by most recent fund vintages. In each of the 2019 and 2020 vintages, for instance, median DPIs are still barely over zero, and less than half of all funds have begun to return any capital at all to their LPs. The 2017 and 2018 cohorts are the only recent vintages with much DPI to speak of, and even then, the return profiles remain relatively slight. Across those two vintages, less than 20% of funds have yet reached a 1x DPI, marking the point at which fund LPs start to earn a profit, rather than simply getting back the capital that they initially paid in.”
(暂无翻译)
Simon Willison 的 CLI 工具是极客党玩转多模型的瑞士军刀。支持推理轨迹输出是最实用的更新,让开发者能直观看到模型思考边界,这对于本地调试复杂提示词和微调数据集构建极其方便。
Big new release of my LLM CLI tool and Python library for talking to hundreds of different LLMs - reasoning traces, OpenAI Responses support, server-side tools, smarter logging and a whole lot more https://t.co/IgGoMlV4Br
(暂无翻译)
The audio is weird speech-like junk, but that's my fault for not following the prompting guide and telling it what audio I wanted - here's that guide: https://t.co/n2NfzmefxY
(暂无翻译)
The MiniMax-H3 video generation model is a lot of fun - here's what I got on my M5 Pro Mac for the prompt "a rainbow colored skunk leaps over a mossy log in a supermarket" (~115GB model download, took around 45 minutes to generate) https://t.co/dYjE1VFFml
(暂无翻译)
这两位 Google 核心基础设施的奠基人双双亮相社交媒体,绝不是闲着无聊。这大概率是他们在新创业项目正式推出前,为建立公众影响力和吸引顶尖开发者社区关注而做的铺垫。
Oh look! My longtime friend and colleague @Sanjay_Ghemawat is now on Twitter! Give him a follow!
(暂无翻译)
即便在顶流创投圈,代沟和亚文化冲突依然存在。这种敏锐观察日常细节的能力,正是优秀产品经理发掘用户隐性需求的关键,很多爆款 AI 应用的切入点往往源于这些微小的生活痛点。
Jessica got her eyebrows tinted and 17 yo has started referring to her as Comrade Brezhnev. She is not amused. https://t.co/TdnvjQ2ORZ
(暂无翻译)
He started dictating to me before he could write, and we just kept going. We've been doing this for about 10 years now.
(暂无翻译)
The most alarming thing about this graph is not just the amount of change, but the smooth consistency of the curve. If this were a startup I'd be really confident about their future growth. Unfortunately it is instead the planet we live on.
(暂无翻译)
14 yo has gotten so tall in the last year that he now towers over me when he paces around my desk dictating stories. He used to disappear behind my monitor on each lap. But he's still dictating stories!
(暂无翻译)
把雇佣 AI 员工包装成“Vibe Deploying”是个很性感的概念。这说明 AI 不再按调用量计费,而是开始像人类员工一样拿销售提成,这种与客户业务收益深度绑定的商业模式将彻底重塑 SaaS 行业。
Big changes continue to happen in how we work with agents. So lots of opportunities around. Vibe deploying is getting paid to put AI employees to work inside real businesses. Viktor pays a 20% revenue share for 12 months, cash activation bonuses of $500 to $2,000 per client, and takes zero cut of implementation fees. Current partners charge $3,000 to $10,000 per implementation. Free to join at https://t.co/gJydZUcfEU @viktor_com #vibedeploying
(暂无翻译)
Routing for long-horizon coding agents is a big deal. @notdiamond_ai just announced a model router that works natively with Claude Code. This is huge. It picks the model and reasoning effort before each turn in a session, runs through a privacy-preserving local proxy, and your requests still execute through your own gateway. In their benchmarks, it approximates Opus 4.8 Xhigh quality at 39 to 61 percent lower cost.
(暂无翻译)
AI-generated apps all have the same look. @boltdotnew solved this with the new Template Marketplace. Design studios and agencies built them. Each one is a working full-stack app with the backend, data, and logic already wired up. All free. https://t.co/bAfziz2CoI
(暂无翻译)
AI 代码生成正在摧毁传统拖拽式低代码平台的护城河。当自然语言就能生成完整界面和逻辑时,Retool 这类强依赖预设组件的工具如果不全面拥抱大模型,将会面临严重的用户流失危机。
@howietl @hyperagentapp more interesting to ruminate is what happens next to fellow low-code darling @Retool... (for the record, another one i really really like) https://t.co/nFhp1rKxHa
(暂无翻译)
we are huge airtable fans - aie runs on airtable - so seeing this number might seem surprising, but don't count @howietl out! if you were at WF26 you would already have seen @hyperagentapp which is the next chapter of team Airtable! https://t.co/yQhb4VUue4
(暂无翻译)
能接管实体手机号的语音 Agent 正在打破虚拟与现实的边界。这不仅是个极客玩具,对于销售客服和中小企业主来说,拥有一个永远不会漏接且能实时总结待办的数字分身,意味着生产力的大幅解放。
this is a really clean and unique ai tool that acts more like an extension of you took me 15 min to set up a custom ai voice assistant w/ it’s own # that also picks up my calls when my phone is off and emails me summaries (it can do more) yes we ended up investing @untappedvc :)
(暂无翻译)
路由层工程正成为大模型应用的新焦点。通过在网关层部署小模型做意图识别,再将复杂任务抛给顶级大模型,这种混合调度架构能将整体推理成本砍掉 80%,是当前应用层最实在的优化手段。
model engineering -> harness engineering -> router engineering 3rd new layer from which we can now increase intelligence. i predict surprisingly robust gains here
(暂无翻译)
i was messing around with deepseek v4 flash last night. it's *basically* free, and there are like a half dozen things it is perfectly capable of that i can offload from my fable 5 workflow. makes too much sense. so many new models for model blending alchemy adventures stoked
(暂无翻译)
bullish model routers. same perf at lower cost is obvious. but there are massive gains to be had by creating "smoother" intelligence via blending multiple jagged models together. the era of model melding begins.
(暂无翻译)
连 Google DeepMind 的 AGI 安全负责人都在急招人,说明底层大模型的迭代速度已经超过了安全对齐技术的承载力。懂可解释性分析和对抗样本攻防的安全工程师正成为大厂疯抢的稀缺资源。
The GDM AGI safety team is hiring! If you have concerns about the implications of increasingly powerful AI, you should apply. Things are kinda wild right now and I don't expect the wildness to slow down.
(暂无翻译)
能塞进普通笔记本本地运行的高性能开源模型是打破云厂商垄断的关键。阿里的 Qwen 系列如果在端侧跑通,将极大推动隐私优先的离线智能应用普及,对于隐私敏感的医疗和法务行业是巨大利好。
I try not to get excited about models before they've been released, but I gotta admit I'm very much looking forward to the upcoming laptop-sized Qwen 3.8 models
(暂无翻译)
硬件靠广告变现的“小米模式”正在被传统车厂滥用。这种牺牲用户体验的做法,反而会逼退消费者转向软件体验相对纯净的新能源车企,这也是国产新能源出海降维打击的核心软实力之一。
Absolute enshittification 😂 After Samsung and LG added ads and spyware to all their TVs Now even car manufacturers like BMW, Jeep, Ram, Chrysler and Dodge have in-car ads! Just buy a Tesla or BYD!
(暂无翻译)
Okay so today I worked on the coolest part of is my AI video editor: the agent! It's a Cursor-like sidebar and you can just tell it to edit your video with your clips and library for you It's still very basic but it made this edit all by itself! It sends the current state to @xAI and then asks it to edit it based on your story Live now for everyone on my site Photo AI 😊 Tomorrow I'll try make it just multi-lanes and become more smart, like it now put my pre-DMT trip videos (with regular hair) sometimes after the DMT trip (with long hair and beard and crazy eyes), but maybe it has a point for that, not sure Anyway very cool cause I hate editing and just talking to AI and letting it figure out is nice!!
(暂无翻译)
The difference between European and American cafes is so stark: In Europe many don't allow laptops anymore In America they usually do and people are working on something cool! Why is this important? Nvidia, a $3 trillion dollar company, was started in a Denny's, an American diner People need a third space to work on their laptops to build the next billion or trillion dollar company The benefits of working in a publicly accessible coffee place were already known to Europeans like Isaac Newton 400 years ago Somehow we forgot: The Age of the Enlightenment happened in 1687 because of people meeting in publicly accessible coffeehouses, not private offices/coworkings: Historians often associate English coffeehouses, during the 17th and 18th centuries, with the intellectual and cultural history of the Age of Enlightenment: they were an alternate sphere, supplementary to the university. Political groups frequently used coffeehouses as meeting places. The first coffeehouses established in Oxford were known as penny universities, as they offered an alternative form of learning to structural academic learning, while still being frequented by the English virtuosi who actively pursued advances in human knowledge. The coffeehouses would charge a penny admission, which would include access to newspapers and conversation.[13] Reporters called "runners" went around to the coffeehouses announcing the latest news. This environment attracted an eclectic group of people who met and mingled with each other. In a society that placed such a high importance on class and economic status, the coffeehouses were unique because the patrons were people from all levels of society.[15] Anyone who had a penny could come inside. Students from the universities also frequented the coffeehouses, sometimes even spending more time at the shops than at school https://t.co/AybZy59GJ9 From: https://t.co/7LohTYMmKH
(暂无翻译)
My accountant I pay a lot of money every month just replied to a question I sent him with a 100% AI reply explaining something The answer was helpful though so I could've just asked AI myself Of course you know what I'm thinking now, if I can just ask AI everything since he is forwarding everything to AI any way, why not replace him with AI I guess he should be doing the opposite of AI, sell a premium service for a high price (he already does) and keep answering with humans as accounting is something I'd prefer to actually get a human in the loop Anyway it's interesting
(暂无翻译)
Altman 的话反映了硅谷核心圈面对技术停滞传闻的应对策略:强行靠信念注入势能。在当前 Scaling Law 见顶的争论下,应用层创业者最需要的就是这种执行力,去跑通那些别人认为不可行的商业闭环。
i would rather be an optimist and work hard than a pessimist posting about why things won't work. it's much more difficult and the most likely path is failure, but society fails if people don't try. no amount of "it will never work" essays will drive society forward.
(暂无翻译)
NVIDIA 不只卖算力,还要用推理模型插手软件生态。车企如果直接采用这种开源的高阶推理模型,将大幅缩短从 L2 到 L4 的研发周期,但也会在核心算法上彻底沦为英伟达的硬件组装厂。
Today, we’re launching Alpamayo 2 Super, our frontier open reasoning model for autonomous vehicles. Beyond seeing, Alpamayo understands and reasons through the complex world - thinks before it acts. It’s a powerful backbone for robotaxis, trucks, shuttles, delivery vans, tractors and the long tail of mobile robots—billions of autonomous machines someday. We’re releasing it for commercial use under OpenMDW-1.1 so teams can inspect it, fine-tune it and deploy it—open models advance safety and security. The next wave of AI is robotics—and it starts with autonomous vehicles. Great work, Alpamayo team! https://t.co/2PYCCXWjZh
(暂无翻译)
这是做 AI 产品验证的黄金法则。直接按用户说的去堆砌功能会导致产品臃肿不堪,用大模型做日志分析深挖用户的底层痛点,再结合模型能力反向定义交互方式,才能做出超越预期的杀手级功能。
When you go talk to users, don't (just) ask them what features they want. Ask them what problems they have. In the best case this will lead you to think of features they'll love but would never have thought to ask for. And expect to go through multiple cycles of ask-and-build.
(暂无翻译)
Sacrificing legibility for distinctiveness in your logo is a common beginner mistake. It not only makes your logo harder to read, but signals weakness. It's the graphic equivalent of timid hunter-gatherers retreating to land the fierce pastoralists don't want.
(暂无翻译)
I asked the founders if I could say which company this is, and they said ok. It's Greptile, which incidentally also has one of my favorite names of any YC co.
(暂无翻译)
The danger of selling to big companies, if you're a startup, is that they don't say no outright. They have months of meetings with you first. Since you hate meetings, that seems to you a sign of commitment. But it's not. They love having meetings! It's almost all they do.
(暂无翻译)
Bought a book. It was awful. Didn't want it on my shelves, but couldn't throw it away, so it sat on a table near the door. Rushing to an appointment this morning, grabbed it to read. Then went to breakfast and had nothing else. So I spent the morning reading the worst book I own.
(暂无翻译)
GTA dressed 17 yo's character as a nerd for a mission. Then I walked in wearing exactly the same thing. Orange fleece vests are for nerds? Who knew?
(暂无翻译)
Agent 已经开始接管知识工作者的元认知管理。将碎片化信息源(邮件、微信、会议记录)自动转化为结构化 Todo-list 并实时更新,这种无缝的跨平台数据流转是提升个人产能的最短路径。
and in the 3-hour consolidation step, it's adding, updating, and checking off my personal tasks/action items https://t.co/PmsxFSed4Z
(暂无翻译)
this feeds into a larger system that also pulls from my email, calendar, and notes to track my personal task items but also works as a cross-chat memory system with access to my work stuff https://t.co/OiPi7eNjhN
(暂无翻译)
sweet, i finally set up my chatgpt, codex, claude, & claude code to have a shared memory - so i can ask one of them what i've worked on recently across all four of them for codex/cc, i'm using local history, but for chatgpt/claude i use a scheduled recap skill to regularly send summaries
(暂无翻译)
把方向盘交给 AI 反而可能比交给人更安全,这是对齐研究界的前卫观点。在自动驾驶或金融高频交易等场景,限制人类的非理性干预,让多智能体系统基于既定宪法进行自洽博弈,可能是解决极端事故的出路。
I’m sympathetic to this position. Non-autonomous AI is wielded by humans, whose alignment seems intractable, whereas autonomous AI omnibenevolence (reliable enough to bootstrap using multi-agent deliberative processes) seems tractable (shockingly so, from my former perspective).
(暂无翻译)
Meanwhile, the status of Theology is severely mixed. During these uncertain times, we hope the following can serve as a handy guide to the regions of Theology which are most vulnerable to also being explanations of Reality https://t.co/R586vkQ4oe
(暂无翻译)
Public Service Announcement: While a significant portion of Science Fiction is in the process of becoming Reality, and we understand this may be disorienting, please note that the status of nearly all Fantasy is unaffected, and remains make-believe.
(暂无翻译)
Reminds me of a conversation I had at MIT CSAIL 18 years ago where I and others debated and eventually agreed that, no later than 2045, it should be possible to run a Turing-Test-passing chatbot in real-time on a top-of-the-line Early 2008 MacBook Pro.
(暂无翻译)
前 HuggingFace 工程大佬的动向一直是开源圈风向标。在多模态和 Agent 工具链日趋成熟的当下,他回来后大概率会押注 AI 视觉理解在工业质检或机器人上的端侧落地应用。
Off for the next 7 days 🏝️👋 you can leave wishes below for what I should work on when I am back.
(暂无翻译)
不仅是工具,AI 正在向优化人类生物钟的领域渗透。对于需要高强度脑力输出的开发者,基于个人生理数据动态调整工作流的 AI 助手,比起传统的番茄工作法,更能实质性延长职业寿命。
I actually love this idea. As someone who typically needs 7-8 hours to 100% function, having something like this would be incredible. Not just in terms of increasing my productivity, but in terms of having a richer set of fulfilling experiences from friends, family, learning, hobbies, anything as long as there are no side effects which I know is going to be really hard to prove
(暂无翻译)
Google Bard 和微软 Sydney 的翻车史说明,过早将半成品大模型推向 C 端是极度冒险的。但也正是这种不对称的声誉赌注,才打磨出了如今 GPT-4 和 Gemini 的产品成熟度,后来的跟随者已经没有这种试错红利了。
Both companies took tremendous reputational risks on an early technology before there was even a market for it (Bard was inaccurate and had controversial imagegen, Sydney was insane). Their inability to be daring now, when the market is clearer, is more surprising as a result.
(暂无翻译)
Given today, it is surprising how daring Microsoft & Google were initially with AI. Microsoft released GPT-4 before OpenAI, didn't back down after Sydney & got Copilot to market quickly (the 1st professional AI tool) Google had the first deep research & pivoted fast with Bard.
(暂无翻译)
For some reason today Steam games froze when launching. Why? Apparently the latest update of my wireless keyboard also got installed as an Xbox controller which conflicted with Windows joystick's legacy API controller 0. Codex fixed it. I cannot imagine figuring it out myself.
(暂无翻译)
Feels like this would have been a massive win for Microsoft to ship in some corporate IT friendly form.
(暂无翻译)
An unexpectedly stress-reducing use of Codex/Code is just fixing problems with my various Windows machines: weird driver issues, game incompatibilities, even just tiny stuff that used to annoy me (why does the program that I set to run at startup not run at startup?). Hours saved
(暂无翻译)
Its pretty neat that it was able to manipulate all these different systems, including making new Blender models with animation https://t.co/tQz9pbFCOV
(暂无翻译)
Sure, 3js in webpages are neat but, Codex: "you have access to Blender & Unity. I want you to make a new game, with full assets, in which you play as an otter who can get into mech suits shaped like animals and that are critical to game play" Took over my computer & gave me this https://t.co/SjTTmOUPNt
(暂无翻译)
The article is AI written, but does make some useful points. However, editors need to be up-to-date enough on AI research to ban references to deeply methodologically suspect NANDA paper, especially a year later when we much better information about what is happening in companies https://t.co/s9120rijZi
(暂无翻译)
Nate is right (full context can lead to a degradation of chats in several different ways) but you can't ask the AI about stuff like this as they have bad self-knowledge A good approach is to compact the work so far (ask for a md file summarizing things) & use it in the new chat
(暂无翻译)
LangChain 把精力转向那些“无聊且无差别”的脏活累活,说明 AI 中间件拼图时代的终结。谁能为企业开发者屏蔽掉底层不同模型 API 的诡异 Bug,谁就能在下一代 SaaS 平台之争中拿下最多的付费用户。
did you know you can do this in langgraph studio? should we make a tutorial for it?
(暂无翻译)
we're going to move managed deepagents to public beta this week big focus on all the "boring" and "undifferentiated" infra around agent so you can focus on the agent logic. so far this includes: - opinionated evals setup (using harbor) - memory (agent and user level) - proper oauth for tool access - easy channel integrations (slack, github) - seamless sandbox integration what else would be good?
(暂无翻译)
Internal AI platforms have the potential to transform how companies function Read how Stripe built theirs
(暂无翻译)
硅谷大佬下场推 CLARITY 法案,核心是争夺稳定币和加密支付的合规主导权。这对于 AI Agent 间的微支付结算网络是个极强信号,未来大模型调用算力和数据的计费颗粒度将精细到分秒级。
Congress has a generational opportunity to pass policy that will keep the US the global leader in both technology and finance. Pass CLARITY. 🇺🇸
(暂无翻译)
把企业内部知识库和代码库打通为共享语义层,是下一代 RAG 技术的高级形态。这种自适应检索机制能极大减少幻觉,这意味着像 Replit 这样的云开发环境,正在变成能够理解上下文的超级 IDE。
We built a self-driving & self-correcting shared semantic layer on top of our databases, conversations, and docs. Everything is queryable & joinable—regardless of source! So now anyone at Replit can ask questions that previously needed a team of data scientists weeks of work. https://t.co/Uskhjp4KCR
(暂无翻译)
数据很直观:Prompt 解析和 Tokenization 竟然吃掉了三分之二的响应延迟。优化大模型网关层的 Token 调度策略,将是提升 C 端应用丝滑体验的最快途径,尤其对于需要频繁重试代码的 Coding Agent 至关重要。
TokTier makes tokenization stateful for agentic serving. Across 153,951 real agent calls with a 94.1% prompt-cache hit rate, tokenization eats up to 64% of time to first token. Coding agents resubmit a long transcript after every tool result, and even a short append can shift token boundaries at the tail of the previous sequence, so nothing gets reused. TokTier is a stateful tokenization service. For a session continuation it re-tokenizes a small window around the append, runs a stable-boundary check, and splices only when that check passes. When there is no reusable prefix it runs exact pre-tokenization and BPE on a GPU. Emitted token IDs always match full reference tokenization. Differential campaigns across 17 tokenizer families covered 1.5e10 split checks with zero divergence. Incremental repair takes 0.5 to 1.1 ms from 100K to 3M characters, up to 437x faster than HuggingFace. Under vLLM, median time to first token drops 16 to 34%. Four repair cores plus one GPU sustain 1,821 requests per second where a 16-core stateless front end saturates at 40. Paper: https://t.co/AM2dU8pnIx Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
(暂无翻译)
This is a wild result. Locus, the automated research system from @intology, post-trained Qwen3 base models that beat the official human-tuned Qwen3 1.7B Instruct release. SoTA on PostTrainBench! https://t.co/F7BFfuYswe
(暂无翻译)
Try Qwen3.8-Max on Hermes Agent and you will have to doubt on how much these open frontier models have caught up with frontier closed models. These new open models are insanely good.
(暂无翻译)
把神经算子(Neural Operators)引入气候临界点预测,展现了 AI for Science 在宏观尺度上的巨大威力。这类技术不仅能救命,更能转化为金融市场的气候风险定价模型,具有极高的商业延伸价值。
Great to have participated in preparing the @ScienceBoard_UN brief on tipping points with clear definitions and also on how AI is impacting the field. Our paper on using Neural Operators for early detection of tipping points by measuring deviation against baseline physics is helpful in a number of areas: from climate tipping point in stratocumulus cloud cover to airfoil wake and stall transitions using only limited knowledge of the governing equations. https://t.co/njGKOQ2dvA @mliuschi @Caltech
(暂无翻译)
在 X 平台上用大模型生成梗图并引发头部大 V 互动,说明多模态 AI 已经成为社交媒体流量分发的核心抓手。X Money 的推进暗示马斯克正试图把社交平台变成 AI 任务的众包结算平台,打通商业闭环。
.@grok you are banned for making this meme for the rest of the year — or you will be unplugged! https://t.co/f86StisEN6
(暂无翻译)
PLEASE ENABLE X MONEY FOR @JASON TODAY!!! Thank you for your attention to this matter!
(暂无翻译)
So you don’t have to send masked, violent ICE agents into cities to collect illegal aliens? Fascinating Paging @StephenM — take some notes! 📝
(暂无翻译)
The $btc power bottom is in! 😂 Seriously, if you were my brother I would say “sell half your bitcoin and put it into productive projects like bittensor:native and solana:So11111111111111111111111111111111111111112, that are actually delivering new products — not just projects and promises” https://t.co/Kf8d16eh2r
(暂无翻译)
RELATED: I just realized folks are not following me for clever tweets, but strictly for my good looks.
(暂无翻译)
There is no reason grocery prices can’t deflate by 10%+ a year for the next couple of years. Solar installed by @Tesla_Optimus on top of water pipelines filled with desalinated water to automated farms that drive crops in FSD trucks overnight to community grocery stores This future is already solved for, just not yet implemented! TECHNOLOGY VS SOCIALISM FIGHT!
(暂无翻译)
大模型把软件研发成本打到了地板价,企业以往依赖外部 SaaS 的逻辑正在解体。当内部用 AI 自动生成定制化代码的成本低于昂贵的年费订阅时,传统 SaaS 厂商的续费率将面临断崖式下跌。
If you are an enterprise product or engineering leader evaluating Software Factory, come join our Q&A on Aug 12.
(暂无翻译)
The Build vs Buy Math Just Flipped For thirty years a software renewal was a simple question of renew or shop for a cheaper seat, because building the thing yourself was never realistic. That has changed. Gartner now puts roughly a fifth of enterprise SaaS spend, around $234B, at risk of being done a different way by 2030. The reason is that the cost of building the workflows a company actually needs can now be done with tools like Software Factory. Keep renting commodity systems, and your business will perform like a commodity. Building software around the processes that are unique to how your business runs is worth pricing out before you sign another three-year contract for some off the shelf tool that will only ever approximate it.
(暂无翻译)
利用大模型强大的上下文归纳能力反向分析用户画像,是目前 C 端应用变现的新玩法。通过这几个特定提示词让模型吐出它“眼中”的你,不仅能提供极强的心理冲击感,也是验证大模型隐私安全边界的绝佳测试。
AI has been quietly assembling a complete behavioral record of you. Might be worth opening it. I've created a prompted called 'Reflection Engine'. Here are 22 questions to unlock that data for you:
(暂无翻译)
没有标准答案的开放式生成任务,是当前大模型对齐的死角。用大模型做大模型裁判(LLM-as-a-Judge)依然存在严重 bias。这说明在文档摘要或文案创作场景,仍然极其依赖人工 RLHF 飞轮来兜底质量。
RT @Benjamin_eecs: Thanks @_akhaliq! RLVR needs a verifier, but open-ended tasks (writing, summarization) have none, and learned judges get…
(暂无翻译)
传统 CAPTCHA 验证码已经彻底被视觉大模型攻破。反爬虫和风控系统必须抛弃这种依赖图像识别的弱验证,加速向基于设备指纹、行为轨迹和无感令牌的零信任架构迁移。
linking my main cua wow moments thread https://t.co/VCXzCTAzRY but this one is just about do we need captchas anymore when clearly bots can clear them
(暂无翻译)
前 Google CEO 点名的 Asari 方法直指推理成本优化的痛点。在算力稀缺的当下,能够通过算法层面的智能体调度降低数据中心能耗的团队,将直接掌握下一代云服务的定价权。
This Asari approach works, imagine the improvement in inference and the reduction of data center serving costs when people adopt this! The power of agents that are incredibly smart.
(暂无翻译)