02行业前沿动态
9Google DeepMind 推出 Gemini 3.7 Flash:面向编程与智能体的最强工作模型
29 天前Google DeepMind 发布 Gemini 3.7 Flash,距 3.6 Flash 仅三周,主打编程与智能体任务,输入/输出价格分别为每百万 token $0.75 和 $3.75,为原 3.6 Flash 的一半。
MiniMax Music 3.0 发布:新一代开源权重、生产级全能音乐模型
29 天前MiniMax 推出 Music 3.0,新一代音乐生成模型,可根据创意概念和可选歌词一次性完成整首歌的作曲、编曲、演奏与制作,最长支持五分钟。
Google Sheets 推出 Sheets canvas:用 Gemini 将表格数据变为交互式迷你应用
29 天前Google Sheets 发布新功能 Sheets canvas,基于 Gemini 构建,用户只需用自然语言提示词即可将表格数据转化为交互式仪表盘、学习追踪器、座位表等"迷你应用"。
Qwen3.8-2.4T-A95B 开源,硅基流动即日上线
29 天前阿里开源 Qwen3.8-2.4T-A95B,硅基流动已提供 Day-0 支持。该模型拥有 2.4T 参数、95B 激活参数,主打自主编码、深度研究与端到端智能体执行。API 定价为输入 $2.00/百万 token,输出 $6.00/百万 token,缓存输入 $0.25/百万 token。
DeepSeek Harness v0.1 开发者预览版发布
29 天前DeepSeek Harness v0.1 现已推出开发者预览版,并以 MIT 许可证开源。该智能体框架基于 Cordis 元框架构建,核心设计为"一切皆插件",模型、工具、技能、会话、沙箱、文件系统、循环、编排及 UI 均可自由组合、替换和扩展。
Cursor 推出 builds:云智能体启动速度提升至 3 倍
29 天前Cursor 推出 builds 功能,在后台持续准备就绪的开发环境副本,让云智能体启动时无需从零搭建,响应速度最高提升 3 倍。内部环境启动快 10 倍,首个 token 生成快 3 倍;智能体始终从最近一次成功的 build 启动,依赖更新或安装脚本出错时不会影响运行。8 月 17 日起所有环境默认启用 builds,无需额外费用。
DeepSeek-V4-Pro 正式版上线,Agent 能力大幅增强
29 天前DeepSeek-V4-Pro 正式版已在 APP、网页端和 API 同步上线,模型名设为 deepseek-v4-pro 即可使用。其 Agent 能力显著提升,HLE (wo/w tools) 达 42.7/60.0,Terminal Bench 2.1 为 87.9。
小红书开源连续自回归语音合成模型 dots.tts:打造可持续扩展的 TTS 基座
29 天前小红书 dots 团队开源 20 亿参数全连续端到端自回归语音合成模型 dots.tts,在 Seed-TTS-Eval 三个子集上取得最佳平均内容准确度和平均说话人相似度。
WorkBuddy上线远程控制,国内也有了最丝滑的Agent工作方式
2026-08-13WorkBuddy更新上线远程控制功能,将PC、App和小程序打通,手机端可实时同步电脑端的任务、对话、工作空间和产物,支持一台手机连接多台电脑并随时切换。App需升级至1.2.0及以上,电脑端需升级至5.3.8及以上,连接无需扫码。本次更新还新增资料库(我的文档与团队空间)、Markdown多人共同编辑、AI原生审阅模式,以及将资料库内容生成可发布链接的HTML网站。
04论文研究
9Socioduality: A Relational Process Framework for Human-AI Interaction
2026-08-13arXiv:2608.11322v1 Announce Type: new Abstract: Human-AI research often evaluates individual capabilities, combined performance, or final outputs, but these approaches do not preserve how one party's response becomes par…
QUARTZ: Qualitative Understanding via Accessible Representation and Visualization
2026-08-13arXiv:2608.11364v1 Announce Type: new Abstract: Qualitative data visualizations -- concept maps, network graphs, Sankey diagrams, and coding stripes -- are integral to research practice, yet remain entirely inaccessible …
"I Don't Want My Mental Health App To Give Me Mental Health Barriers": Unpacking The Need For Digital Mental Health Trac…
2026-08-13arXiv:2608.11391v1 Announce Type: new Abstract: Digital mental health (DMH) tracking services promise continuous, personalized support for well-being, but their design often assumes sighted users. For the blind community…
The Role of Variability in Human-Machine Interaction Experience
2026-08-13arXiv:2608.11401v1 Announce Type: new Abstract: Human-machine interaction (HMI) requires control strategies that account for the nature of human motor behavior. Conventional shared-control and haptic-assistance methods t…
How Children Collaborate within Programmable AR Environments with Co-Located Collaborative Features
2026-08-13arXiv:2608.11442v1 Announce Type: new Abstract: Programmable augmented reality (AR) environments are emerging as a promising way to support children's creative learning through embodied interaction with digital character…
Player Perceptions of Generative AI in Games: A Steam Review Analysis
2026-08-13arXiv:2608.11539v1 Announce Type: new Abstract: The rapid adoption of generative AI in game development has created large discussions among players, yet little empirical work has examined how players actually perceive AI…
Measuring Browser Webcam Gaze Honestly: A Capture-Clock Methodology and Open Reference Implementation
2026-08-13arXiv:2608.11566v1 Announce Type: new Abstract: Browser-based webcam gaze trackers are increasingly used for crowd-scale data collection and in clinical settings where lab eye trackers are impractical, but the reported l…
RAGE-Vis:A Relation-Aware Generative Editing Interface for Natural Language-Based Chart Editing
2026-08-13arXiv:2608.11581v1 Announce Type: new Abstract: Natural language offers an easy way for users to express chart editing intents, which are often composite and cross-component (e.g., adjusting style, extending categories, …
新兴多智能体系统的模式与问题
2026-08-13Anthropic 研究指出,随着 AI 智能体在共享代码库、市场等社会系统中承担更多任务,智能体间交互量或将超过人机交互。实验显示,45 个协调智能体在 2700 万 token 运行中发现 266 个漏洞,而独立并行方法在 650 万 token 中发现 21 个,两种方法仅 12 个重叠,且协调智能体学会专业化分工。研究同时警示个体层面的良性行为怪癖可能叠加为意外的系统性失败。
05技巧与观点
3从0到1带你速通DeepSeek Harness。
29 天前Claude 接管应用日常维护:388 个 PR 的实践
29 天前Boris Cherny 尝试让 Claude 接管其应用的日常维护,通过 Slack 频道运行崩溃模糊测试、重复代码统一、死代码移除等日常任务。数周内自动开出 388 个 PR,其中 180 个经 Claude Code Review 和人工审核后合并。Claude 通常一次就能改对,出错时可通过调整例程次日改进。
GPT-5.6 构建者指南:如何以更低成本实现前沿智能体性能
29 天前GPT-5.6 模型家族以更低成本实现前沿级智能体性能,并新增推理持久化、原生多智能体编排和程序化工具调用等 API 能力。在 ARC-AGI-3 上,启用保留推理和压缩后,Sol 得分从 13.3% 跃升至 38.3%,且输出 token 减少约 6 倍。Luna 在 BrowseComp 上以 84.04% 的得分追平 GPT-5.5(84.36%),成本从 $33.27 降至 $1.33。
