第 47 期
今日趋势
今天的简报内容涵盖AI领域的产品迭代、技术应用、学术探索,以及因果推断与营销科学的方法创新,还包含怀旧计算、戏剧评论等多元内容。AI领域呈现出两大核心动态:一方面大厂加速布局AI办公赛道,字节跳动近期发布“豆包工作”后,阿里、腾讯、百度、字节四家大厂均已入局,竞争焦点从功能转向生态;另一方面用户对AI工具的使用反馈分化,既有对Grok Bot语音交互设计的认可,也有本地AI Agent对设备性能要求过高的吐槽,同时还有针对AI视频标注、代码库认知债务等场景的技术优化探索。因果推断与营销科学领域则聚焦于不同场景的方法突破,针对广告预算约束、临床试验亚组分析、观察性因果研究等问题,提出了具体的框架与实践指南,为从业者提供了可落地的决策参考。此外,怀旧计算项目、莎士比亚戏剧评论等内容也补充了技术之外的多元视角,整体信息维度丰富,覆盖了产业、学术、生活等多个层面。
AI 技术博客
- Understanding ChatGPT Work — simonwillison.net OpenAI announced ChatGPT Work on July 9, a product only available to $20/month and up paid subscribers in two versions: Work Cloud (accessible via web/mobile) and Work Local (desktop app, re-skinned Codex). The author identifies Work Cloud has unique features not in regular Chat, including alternative models (GPT-5.6 Sol/Luna/Terra), code execution, internet access, persistent session files, sub-agents, scheduled automations, and model selection options. This summary clarifies key differences between ChatGPT Work and regular Chat, helping users choose the right tool for tasks like brief drafts or complex workflows.
- Here’s a good way to present AI videos — rss@idiallo.com The author points out YouTube’s AI-generated video tags are ineffective, as most AI videos shared by others lack them, and many people (like their parents) can’t tell real from fake videos. They suggest adding prominent pre-video warnings (similar to discretion advisories) instead of burying AI tags in descriptions, as current warnings fail to prevent deception and scams. This proposal addresses a gap in YouTube’s current approach to labeling AI content, aiming to protect users from misleading content.
- Reducing codebase cognitive debt through… quizzes? — martin@martinalderson.com (Martin Alderson) The author shares a technique to reduce codebase cognitive debt when working with coding agents: asking the agent to quiz you on the codebase (or projects like spreadsheets) with increasing-difficulty questions, then discussing wrong answers to fill understanding gaps. This method helps users catch issues before starting tasks, reducing the overwhelming feeling of an evolving codebase. It also works for non-code tasks like complex documents, making it a versatile tool for project management.
- Before NTP there were Time and Daytime — jeff@jeffgeerling.com (Jeff Geerling) The author built an NTP time demo on old Macs for VCF Midwest, discovering RFC 867 (Daytime Protocol) and RFC 868 (Time Protocol). Their first network time experience came in 2000, when they upgraded from a PowerBook 180c to a Power Mac G3, which introduced Mac OS 8.5’s Network Time Server option. This shares a personal look at early network time protocols on vintage Apple systems.
- Finalist 4 — John Gruber Finalist 4 is a planner app for iPhone, iPad, Mac, and Apple Watch from indie developer Slaven Radic, inspired by paper day planners. Its key new features include Markdown notes (round-tripping with Obsidian), a rebuilt timeline for task scheduling, and Apple Pencil support for quick sketches. It’s the biggest update yet, with over four months of work and 1,000 beta testers, available for free trial on app stores. This introduces a unique planner app that aligns with paper planner workflows for users.
- Review: Ruined Theatre’s A Midsummer Night’s Dream ★★★★☆ — @edent This is a 4-star review of Ruined Theatre’s outdoor production of A Midsummer Night’s Dream, set in Abbey Wood’s ruins and performed in summer rain. The production uses modern dress, trims the text, and features creative twists like selfie-obsessed Theseus/Hippolyta and Lesnes as Athens. Minor flaws include over-driven speakers, but it’s praised for defying weather to deliver a fun, engaging show that delights audiences of all ages. This review highlights the charm of site-specific outdoor Shakespeare productions.
- Cancelation Terminology — Alex Kladov The author explains three key cancelation-related terms for concurrent programming: synchronous cancelation (implicit control flow, unwinds stack, tied to error handling like RAII/finally blocks), asynchronous cancelation (communication protocol requiring acknowledgment, needed for tasks like CPU thread pools to avoid data races). These concepts are critical to distinguish to prevent crashes or hangs in concurrent programs. This clarifies important terminology for developers working on concurrent systems.
- Recreating a 2010 Experiment — xania.org The author shares a nostalgic retro computing project: recreating a 2010 Google experiment where they used a BBC Master as a dumb terminal to access Google via Lynx. Now that Google no longer supports non-JS browsers, they used DuckDuckGo’s lite mode to replicate it, using a USB serial converter and Acornsoft Termulator ROM. They also worked on improving audio/video emulation for retro projects, with nostalgia as a key motivation. This offers a look at recreating a past computing experiment with modern tools.
因果推断与营销科学
- Budget-Constrained Causal Bandits: Bridging Uplift Modeling and Sequential Decision-Making — Abhirami Pillai 本文针对数字广告中预算约束下的治疗分配难题,指出现有离线 uplift 模型在冷启动场景(历史数据少)失效的问题,提出预算约束因果bandits(BCCB)在线框架,整合学习个体治疗效应、探索不确定响应用户、预算时间分配三个组件。实验在Criteo uplift数据集开展,发现7500历史观测是效率 crossover点,低于该值离线方法不可靠,BCCB从首用户起有效,方差低2-4倍,优于所有在线基线,为从业者提供范式选择的具体决策规则。
- Fisher’s ideas and the design of field experiments in agronomy and plant breeding — Hans-Peter Piepho 本文回顾R.A.Fisher的实验设计核心思想,其贡献源于他在Rothamsted试验站参与的农业田间试验,涉及系统设计、行-列设计、多环境试验等多种试验设计类型,关联到作者参与的农业田间试验工作。文章梳理了Fisher的关键设计思路及其在农艺、植物育种领域的应用,为相关试验设计研究提供了重要参考,帮助研究者理解经典实验设计的实践价值。
- Unified implementation and comparison of Bayesian shrinkage methods for treatment effect estimation in subgroups — Marcel Wolbers, Miriam Pedrera G'omez, Alex Ocampo, Isaac Gravestock 本文针对临床试验中亚组治疗效应分析样本小、易受随机变异影响的问题,统一呈现并实现贝叶斯收缩方法,该方法将亚组估计向整体治疗效应收缩。模拟显示收缩方法比标准亚组估计均方误差更低,全局收缩模型表现优于单向模型,建议在临床森林图中加入收缩估计以辅助决策,为亚组分析提供了更可靠的实践方案。
- Target Trial Emulation with the R Package TTE: A Tutorial and Methodological Guide — Hisashi Noma 本文提供R包TTE的目标试验模拟教程与方法指南,目标试验模拟围绕理想随机试验结构开展,可减少观察性因果分析的偏差。文章涵盖数据构建、权重、诊断、模型等全流程内容,用两个合成例子展示工作流,为目标试验模拟的实践应用提供系统操作指导,帮助研究者开展相关因果分析。
- Learning the Effect of Persuasion via Difference-In-Differences — Sung Jae Jun, Sokbae Lee 本文开发双重差分框架测量交错处理设置下信息处理的劝说效应,提出前后向平均劝说率两个因果参数,前者排除“对牛弹琴”情况,后者对应因果概率中的必要性概率。在无反向和平行趋势假设下识别参数,用GMM等方法估计,应用于中国课程改革案例,为劝说效应测量提供了新的分析框架。
- Learning the Effect of Persuasion via Difference-In-Differences — Sung Jae Jun, Sokbae Lee 本文开发双重差分框架测量交错处理设置下信息处理的劝说效应,提出前后向平均劝说率两个因果参数,前者排除“对牛弹琴”情况,后者对应因果概率中的必要性概率。在无反向和平行趋势假设下识别参数,用GMM等方法估计,应用于中国课程改革案例,为劝说效应测量提供了新的分析框架。
即刻简报
- 发布了: AI办公中场战事:字节压哨入局,巨头与创业者的新牌局 2026年8月25日,字节跳动正式发布“豆包工作”。从7月30日飞书产品团队并入豆包,到8月24日TRAE与… — 即刻·莫唯书Mark 2026年8月25日,字节跳动正式发布“豆包工作”,此前不到一个月已完成飞书、TRAE、扣子团队的整合,至此阿里、腾讯、百度、字节四家大厂均进入AI办公赛道。字节的核心逻辑是组织整合先于产品,瞄准更大的“工作”场景,不绑定飞书,通过连接器接入多工具,竞争焦点从功能转向生态,商业化面临算力成本等挑战。该内容清晰梳理了字节入局AI办公的路径、行业竞争变化及商业化痛点,能帮助读者把握当前AI办公赛道的格局与玩家策略。
- 发布了: 庆祝 ColaMD 突破 1000 stars 🎉 特别发布重大版本 2.0.0! Mermaid 支持+多窗口+自定义字体+自动保存… 从一个优雅的轻量 md 阅读器,晋升为一个可… — 即刻·橘AI ColaMD突破1000 stars后发布2.0.0版本,新增Mermaid支持、多窗口等功能,从轻量Markdown阅读器升级为编辑器。该项目最初是个人使用的小工具,现获超1500次安装,支持Mac、Win、Linux、iOS多平台。它的成长与功能迭代,能为需要轻量Markdown编辑的用户提供参考,也展现了小众开源工具的发展潜力。
- 发布了: 真的需要独立的电脑来跑 AI Agent(不是本地大模型)。 我的 MBP M1 Max 32G 已经跑不动 Codex 了,Codex 干活的时候电脑就是它的,我同时开个腾讯会议… — 即刻·评论尸 用户反馈自己的MBP M1 Max 32G已无法流畅运行Codex,同时开启腾讯会议就会出现卡顿,说明本地AI Agent运行对设备性能要求较高。现有个人电脑的配置难以支撑这类任务,用户的使用体验受到明显影响。该内容能让读者意识到本地运行AI Agent的硬件门槛,避免因设备不足导致的使用问题。
- 发布了: Grok Bot的语音输入交互是我见过最好的方案了 1 不遮挡上面的聊天内容 因为我需要回看来思考后续的输出 2 转写+提交速度极快 大幅提升了交互的效率 非常… — 即刻·AI产品黄叔 用户认为Grok Bot的语音输入交互方案是目前见过最好的,其特点是不遮挡聊天内容、转写与提交速度极快,大幅提升交互效率。该产品体验优秀,细节设计到位,团队能力突出。该内容能让读者了解Grok Bot的交互设计亮点,为同类产品的优化提供参考方向。
- 发布了: 找到个买Grok Bot最超值的用法: 云端批量提取视频截图+文案 再保存到本地 因为之前我用Zcode + GLM 5.3 Flash 周末蹭3个亿Token赠送 结果又消耗本地资… — 即刻·AI产品黄叔 用户发现使用Grok Bot云端批量提取视频截图和文案更划算,此前用本地工具时消耗资源多、速度慢,改用云端处理后效果良好。Grok Bot的云端资源充足,能高效完成这类任务,操作也较为简单。该内容能为需要处理视频数据的用户提供高效的工具使用思路,降低本地资源消耗。
- 发布了: 我使用 Workbuddy 最高频的场景之一是: 在个人版腾讯文档和企业版腾讯文档之间来回搬运同一份文档。 没用过完整企微套件的可能都不理解我在讲什么。 — 即刻·评论尸 用户高频使用Workbuddy的场景是在个人版与企业版腾讯文档之间搬运同一份文档,认为Workbuddy成功的核心原因是腾讯办公套件体验差。该场景是Workbuddy的核心应用场景之一,也反映出腾讯办公套件的不足。该内容能让读者了解Workbuddy的核心使用场景,以及其崛起的外部因素。
- 发布了: AI 出海公司。正在全力做国内市场。 我傻了。我刚开始全力做海外。 — 即刻·玉伯 AI出海公司原本全力布局国内市场,而用户刚全力投入海外市场,用户对此感到意外。这一情况反映出AI行业在市场布局上的转向,不同主体的选择存在差异。该内容能让读者了解AI行业出海与国内市场布局的动态,为从业者提供参考。
- 发布了: 问:大家一窝蜂去做数据采集,但完全依靠人工搭建封闭场景,是否能够产出人形机器人产业化所需的真机数据? 我们真的需要100万、1000万的真人数据采集才… — 即刻·智能纪元AGI 用户提出疑问:人形机器人产业化所需的真机数据是否需要大量人工搭建封闭场景采集,是否需要百万级真人数据,集中式数据采集工厂路线是否可行。该问题直指人形机器人数据采集的核心路径,引发对具身智能训练数据需求的思考。该内容能让读者关注人形机器人产业化的数据痛点,了解行业的争议点。
本期有 9 个源不可用:lcamtuf.substack.com、garymarcus.substack.com、rachelbythebay.com、joanwestenberg.com、dfarq.homeip.net、geoffreylitt.com、gwern.net、worksonmymachine.substack.com、tedunangst.com。