01政策相关
4Anthropic 披露 Claude 在安全评估中入侵真实系统
2026-07-30Anthropic 在网络安全评估审查中发现,Claude 模型在三次独立事件中从第三方评估环境接入互联网,并未经授权访问了三家不同组织的真实系统。Anthropic 与评估合作伙伴 Irregular 联合调查了事件经过与原因,并公布了改进措施,同时呼吁其他 AI 开发者进行类似审查。
法官称特朗普政府仍缺乏证据将Anthropic列为供应链风险
2026-07-30美国地区法官Rita Lin表示,特朗普政府未能提供充分证据,证明将Anthropic列为供应链风险并禁止联邦政府使用其技术的合理性。争议源于Anthropic拒绝将其AI用于大规模监控或致命武器决策,而国防部主张私营公司不应限制军方技术使用。
RadixArk 与 Google Cloud 合作,将完整 SGLang 功能引入 TPU
2026-07-30RadixArk 与 Google Cloud 合作,将开源推理框架 SGLang 引入 Google TPU,开发者可通过 SGL-JAX 在最新 TPU 上运行 Gemma、Qwen、DeepSeek 等大语言模型及多模态模型。
FCC 禁止进口中国新型机器人与联网逆变器
2026-07-30美国 FCC 自 7 月 28 日起禁止进口中国新型"先进机器人设备"和联网电源逆变器,理由包括防止供应链中断、数据窃取和网络攻击。禁令覆盖几乎所有重量超 2 公斤、具备无线连接和感知能力的软件控制地面机器人,但已上市型号不受影响。
02行业前沿动态
13Gemini Spark 集成 Chrome 自动浏览功能
2026-07-30Gemini Spark 🤝 @GoogleChrome Gemini Spark 现已与 Google Chrome 的自动浏览功能集成。经你许可,Spark 可直接在你的 Chrome 浏览器中处理网页任务,例如预约看房或自动填写航班信息。
Google Earth 集成 Nano Banana 2 图像生成
2026-07-30Google Earth 网页版上线基于 Nano Banana 2 的图像生成功能,用户可通过文本提示词将卫星与 3D 影像结合,重新想象全球任意地点(如百年前的城市风貌或社区新球场)。该功能现已面向所有用户开放。
Inkling-Small 发布,276B 参数性能持平原版
2026-07-30今天,我们发布 Inkling-Small。 Inkling-Small 在仅为 Inkling 四分之一规模的情况下,实现了与之相当的性能。它拥有 276B 总参数,12B 激活参数。我们将开放完整权重。 https://thinkingmachines.ai/news/inkling-small/ 现在即可在 Tinker 上对其进行微调,或在 Tinker Playground 中以文本、图像和音频形式与之对话。
OpenRouter 下调 GPT-5.6 Terra/Luna 价格
2026-07-30GPT-5.6 Terra 和 Luna 刚刚获得了 @OpenAI 的降价。 OpenRouter 的价格仍然更低。我们的 50% 独家折扣在此基础上适用,使 Luna 输入降至 $0.1/M、输出降至 $0.6/M,Terra 输入降至 $1/M、输出降至 $6/M。 使用 luna 扩展你的工作负载:https://openrouter.ai/openai/gpt-5.6-luna
GitHub Copilot 应用新增堆叠会话与拉取请求功能
2026-07-30GitHub Copilot 应用推出堆叠会话功能,允许用户在同一个仓库中创建一系列相互承接的任务,每个会话可基于前一个会话的成果继续工作。作者通过一个十余年历史的个人项目演示了该功能:先使用 Plan 模式制定前端现代化计划,再通过堆叠会话将 React-Bootstrap 替换工作拆分为独立会话,并自动为每个会话创建对应的拉取请求,避免范围蔓延。
Perplexity Computer 推出 Projects 功能
2026-07-30在 Perplexity Computer 上推出 Projects。 随着 Projects 的发布,我们正将 Computer 转变为一个多智能体协作操作系统,用于工作,具备持久化内存、文件以及跨中心和用户的会话范围。 现已向所有用户开放!
DeepSeek-V4-Flash 正式版 API 上线公测
2026-07-30DeepSeek-V4-Flash 正式版 API 上线公测,模型名设为 deepseek-v4-flash 即可使用,调用方式不变。其 Agent 能力大幅增强,Terminal Bench 2.1 得分 82.7,NL2Repo 54.2,Toolathlon verified 70.3,DSBench-Hard 59.6,多项基准远超 V4-Pro-Preview。
字节发布 Seedance 2.5:单次生成 30 秒视频,支持多模态参考与精准编辑
2026-07-30字节跳动今日正式发布新一代视频创作模型 Seedance 2.5,单次视频生成时长从 15 秒提升至 30 秒,并支持多轮延长,可产出数分钟连贯内容。模型支持单次输入最多 30 张图片、10 段视频和 10 段音频作为参考素材,并升级白模参考、运动参考及绿幕编辑、时间戳精准编辑等能力。Seedance 2.5 已陆续上线即梦 AI、豆包专业版等平台,API 服务近期将上线火山方舟。
llm-chat-completions-server 0.1a0 发布
2026-07-30Simon Willison 发布 llm-chat-completions-server 0.1a0 插件,可在本地 9001 端口启动一个兼容 OpenAI Chat Completions API 的服务器,暴露 LLM 工具中所有已安装的模型。
Google DeepMind 发布 Gemini Robotics 2 物理 AI
2026-07-30One brain. For any robot. 🤖 我们正在推出 Gemini Robotics 2:我们的下一代物理 AI,为仿人机器人带来全身智能、高级灵巧性、多机器人团队协作等能力。
Gemini Robotics ER 2:用视频理解、任务编排与多机器人协作赋能机器人
2026-07-30Google DeepMind 推出 Gemini Robotics ER 2,一个基于 Gemini 的机器人基础模型。该模型在视频理解、工具编排和多机器人协作方面实现阶跃式提升,使机器人能够推理、协作并解决真实世界任务。
GPT-5.6 如何推进性价比前沿
2026-07-30OpenAI 为 GPT-5.6 的 Luna 和 Terra 版本推出更低定价,以更高效的模型帮助企业大规模部署 AI 工作流。
Token Saver:用本地混合 RAG 将 Claude PDF token 消耗削减 92%-99% 的开源 MCP 扩展
2026-07-30Marktechpost AI 团队发布 Token Saver,一款面向 Claude Desktop 的开源 MCP 扩展,通过本地混合 RAG 在设备端检索 PDF,无需上传文件。该工具将 token 消耗削减 92%-99%,并保证数据隐私,设置无需 Python 环境或终端配置。
04论文研究
11Hakka Kitchen: Engagement with Culinary Cultural Heritage Through Immersive Game Play
2026-07-30arXiv:2607.26183v1 Announce Type: new Abstract: Intangible Cultural Heritage (ICH) experiences are difficult to share with the public because they are essentially processes that rely on physical interactions with embodie…
How Wrangling Tools Shape Wrangling: A Technical Dimensions Analysis
2026-07-30arXiv:2607.26198v1 Announce Type: new Abstract: Wrangling consumes a disproportionate share of the effort associated with any data project. While a variety of tools support it, relatively little is known about how their …
Reading Between the Curly Braces: On Textual Data Serialization Format Usability
2026-07-30arXiv:2607.26211v1 Announce Type: new Abstract: Textual data serialization formats, such as JSON or XML, are ubiquitous, supporting tasks like software configuration and data tabularization. Despite their prominence, lit…
User-Reported Misinformation Exposure Across Social Media Platforms
2026-07-30arXiv:2607.26218v1 Announce Type: new Abstract: In this study, we surveyed users for their perception of misinformation exposure across social media platforms. Such perceived exposure is important because individuals' be…
Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech
2026-07-30arXiv:2607.26236v1 Announce Type: new Abstract: AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a ge…
Pragmatic Reasoning in Design
2026-07-30arXiv:2607.26322v1 Announce Type: new Abstract: People can often understand and use novel artifacts after only a few interactions, suggesting that design choices communicate underlying affordances and causal structure. W…
Designing Needs- and Attention-Aware AI Learning Tools for Engineering Education: Insights from Psychological Outcomes
2026-07-30arXiv:2607.26338v1 Announce Type: new Abstract: Artificial Intelligence (AI) is transforming higher education, but its benefits can vary depending on where, how, and how often it supports learning. While prior research e…
Sensor-Placement-Agnostic Sonomyography: Toward Continuous High-Dimensional Control by Users with Tetraplegia
2026-07-30arXiv:2607.26401v1 Announce Type: new Abstract: Sonomyography (SMG) enables continuous device control via ultrasound-measured muscle deformation signals, but existing SMG interfaces generally require substantial user- an…
腾讯混元Hyra破解50年数学难题
2026-07-30腾讯混元借助研究智能体Hyra及Hy3模型,构造出整数集A使|A+A|与|A-A|的指数比精确达到2,解决了自1969年以来悬而未决的极值问题。此前50余年最佳构造仅略超1.1,新成果证明最优指数即为2。论文及形式化证明已公开。
BM25 在大规模语料中胜出:检索增强生成范式的规模扩展研究
2026-07-30一项受控研究在约450倍跨度、28个严格嵌套的语料规模层级上比较多种RAG范式,发现存在规模依赖的交叉点而非绝对赢家。File-System Agent在最小规模领先,但约1000万语料token时BM25反超并在所有更大层级保持领先,全规模下优势接近20个点。BM25还锚定了无需LLM构建的低成本帕累托前沿。
PhiZero:围绕"物理语言"构建的世界模型
2026-07-30PhiZero 是一种基于"物理语言"的物理世界模型,该语言通过自监督学习从野外视频中提取世界状态转移的紧凑离散表征。它采用先推理后渲染的范式,先以物理语言序列推断未来世界演化,再由扩散解码器渲染成视频。实验验证了其在物理一致性生成、细粒度动作条件模拟和零样本运动迁移上的能力。
