01政策相关
102行业前沿动态
10GPT-Live实时音频新架构发布
2026-08-03GPT-Live 是一种用于实时音频的新架构和栈: GPT-Live 可以在说话的同时聆听。 为了让这种体验在 ChatGPT 规模下显得自然,我们从客户端到模型重建了语音栈。 这一新架构让音频持续流动,因此更深入的推理和工具使用不会打断对话。
微软开源 Orchard 智能体训练框架
2026-08-03Orchard 是一个面向研究社区的开源框架,用于跨任务类型训练和评估 AI 智能体。它降低了复杂性,同时通过让研究人员复用同一套基础设施,支持较小模型也能实现强劲性能。https://msft.it/6019a8fqP
OpenRouter 推出 Ori Eval 简化评估流程
2026-08-03推出 Ori Eval:编写首个评估的最简单方式。 没有绝对最好的模型,只有最适合每项任务的模型。Ori Eval 利用 OpenRouter 的 API 处理代码库中的每项任务,然后评估结果。 curl -fsSL https://openrouter.ai/skills/spawn-ori-eval
商汤发布 SenseNova U1.5-Lite-Preview 开源模型
2026-08-03商汤推出 SenseNova U1.5-Lite-Preview,一个基于 NEO-Unify 架构的轻量级原生统一多模态模型,仅 8B-MoT 参数即可达到商业闭源模型的生成与编辑质量。
Cloudflare 推出 @cloudflare/computer 预览版:为智能体提供虚拟文件系统与多执行环境
2026-08-03Cloudflare 发布 @cloudflare/computer 早期预览版,这是一个开源智能体运行时,为每个智能体提供虚拟文件系统,并支持在 isolate、容器沙箱或浏览器中执行代码。
Cloudflare 推出 Billable Usage API:为自助账户提供按产品与计费周期的程序化成本可见性
2026-08-03Cloudflare 发布 Billable Usage API,为自助账户提供单一端点,一次调用即可返回按产品和计费周期拆分的用量与成本,覆盖 Workers、R2、D1、Workers AI、Vectorize、Images 和 Stream。
Cloudflare Workers 与 Containers 现已支持入站 TCP 连接和 gRPC
2026-08-03Cloudflare 在 Agents Week 期间推出 Workers 运行时新处理器 connect(socket),可直接接受 Spectrum 提供的入站 TCP 套接字,并支持将套接字转发至 Durable Objects 或 Containers,实现全双工通信。
MiniMax H3 正式开源:通用全模态生成系统支持 2K 视频与原生立体声
2026-08-03MiniMax 正式开源新一代通用视频模型 H3,可统一理解文本、图像、视频和音频,生成最高 2K 分辨率、最长 15 秒、带 32 kHz 原生立体声音频的视频。
Qwen3.8-Max 发布:开源最强编码与协作模型,2.4T 参数
2026-08-03Qwen 正式发布 Qwen3.8-Max,这是 Qwen 家族迄今最强的模型,拥有 2.4T 参数(95B 激活),并首次开源 Qwen-Max 级权重,开放权重将于下周发布。
UEmbed:统一稀疏与稠密的多模态嵌入模型
2026-08-03UEmbed 是一种仅解码器的多模态嵌入模型,可在单次因果前向传播中同时生成稀疏词级和稠密表示,通过可学习特殊 token 与词汇表分区突破单 token 信息瓶颈。
04论文研究
9To Facilitate or not to Facilitate: Human and LLM Facilitator Tendencies in Online Discussions
2026-08-03arXiv:2607.28643v1 Announce Type: new Abstract: Automating facilitation in online discussions is a long-standing social concern given the increasing time we spend on online spaces and the failure of content moderation ap…
Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework
2026-08-03arXiv:2607.28644v1 Announce Type: new Abstract: Creativity in computational systems is often evaluated as an objective property of artifacts, with existing Computational Creativity (CC) frameworks assessing creative meri…
Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
2026-08-03arXiv:2607.28645v1 Announce Type: new Abstract: Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildabl…
"YES! YES! I absolutely love this insight!" Affirmative Narration as Interactional Strategy in Dialogues with LLM Chatbo…
2026-08-03arXiv:2607.28646v1 Announce Type: new Abstract: This article analyses narrative mechanisms that are common in dialogues with LLM chatbots. In combination, these mechanisms produce an interactional strategy for maximising…
ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning
2026-08-03arXiv:2607.28647v1 Announce Type: new Abstract: This paper presents ConnectED, a human-centered AI system that supports the full instructional lifecycle in Vietnamese education by linking curriculum-aligned lesson design…
Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations
2026-08-03arXiv:2607.28648v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on modifying cognitive appraisals, the …
COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention
2026-08-03arXiv:2607.28649v1 Announce Type: new Abstract: COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conference. It captures ecologi…
Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration
2026-08-03arXiv:2607.28650v1 Announce Type: new Abstract: While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI into system administrat…
SwanTale:面向指令与零样本任务的统一多说话人语音与音频生成
2026-08-03SwanTale 提出统一的多说话人语音与音频生成模型,同时支持零样本与指令任务。研究配套推出 SwanData-Caption 数据方案,通过清洗、合成覆盖与多级标注解决数据稀缺问题,并引入 SwanVAE、Unified MoE、GRPO 后训练等技术。实验显示,SwanTale 在多项零样本与指令指标上领先,并在两项任务的表达力评分中均取得最佳成绩。
05技巧与观点
6Palantir 强劲季度后,CEO Alex Karp 称 AI 行业"马克思主义"
2026-08-03Palantir CEO Alex Karp 在季度股东信中警告,前沿 AI 实验室对企业过于不可信,并称其意图"占有所谓合作伙伴的生产资料",带有"马克思主义色彩"。该公司第二季度营收 19 亿美元,同比增长 93%,利润 11 亿美元。Karp 主张 Palantir 提供模型无关的 AI 与分析软件,让企业掌控自身数据与 AI"废气"(提示词、编排、上下文)。
Claude Code 连接器可复用至 Artifacts
2026-08-03我想很多人没有意识到--如果你连接了一个 Claude 连接器(例如你的 Gmail、日历、Slack 等),Claude Code 也将能够使用它们,包括在 Artifacts 中。
EA 首席战略官谈生成式 AI 如何进入可游玩的实时游戏世界
2026-08-03EA 首席战略官 Mihir Vaidya 认为,游戏是 AI 的试验场,但生成式 AI 进入游戏面临 60 帧/秒、数千玩家同步和低延迟等严苛约束,不能只追求"看起来真实",而必须"行为正确"。他主张采用神经符号架构,在生成能力之外保留确定性与可控性,并称"控制是下一个前沿"。EA 将 AI 影响分为效率、扩展和转型三个层面,其中《模拟人生》已服务超 5 亿玩家,拥有近万亿种游玩排列组合。
AirLLM 实现单块 4GB GPU 运行 70B 模型推理
2026-08-03AirLLM 项目支持在单块 4GB 显存 GPU 上运行 70B 参数大模型推理,无需多卡或大规模显存配置。该项目已开源,相关讨论在 Hacker News 上获得 103 点热度,引发社区关注。
Google Agent Skills 幕后:如何构建、测试与规模化
2026-08-03Google Agent Skills 团队详解其开源技能库的构建与治理流程:项目始于 Google Cloud Next 2026 前的"swarm"冲刺,发布后 GitHub 星标超 15,000。为保证规模化下的质量,每个技能须通过标准化目录结构、CI/CD 流水线(含 linter、链接检查、AI 辅助清单)及提交时与每周的持续评估,并优先引用远程 MCP 工具。
Kimi Work 幻灯片制作教程发布
2026-08-03使用 Kimi Work 制作幻灯片 - 教程 #1。 Kimi Slides 处理整个幻灯片制作流程: - 清晰的结构与研究,由 Kimi K3 驱动 - 连贯的设计,包括精美的图表和 SmartArts - 可编辑并可直接下载 欢迎在评论区告诉我们你还想看什么内容!
