01政策相关
3英伟达、微软和Meta联合警告:应避免对开放权重模型过度监管
2026-07-24英伟达、微软和Meta联合签署公开信,警告对开放权重AI模型的过度监管将削弱美国在AI领域的竞争力。信中指出,开放权重模型能促进创新、降低准入门槛,并支持学术研究。OpenAI和Anthropic未签署该信函。
微软阐述开源模型助力美国竞争力路径
2026-07-24开放权重模型对健康的 AI 生态系统至关重要。我们与行业同仁一道,正在规划一条路径,让开放权重模型在保护国家安全的同时,增强美国竞争力并扩大经济机会。
Kimi K3 在网络安全漏洞利用测试中大幅落后美国前沿模型,知识蒸馏或为原因
2026-07-24英国AI安全研究所与美国AI标准与创新中心联合评估显示,月之暗面的Kimi K3在ExploitBench基准上得分32.2%,远低于美国领先模型的76.2%,但优于智谱GLM-5.2的24.4%。
02行业前沿动态
8Midjourney V8.2 发布:专注美学提升与个性化理解
2026-07-24Midjourney 今日推出 V8.2 图像模型,重点提升美学质量、图像创意与个性化表现。低质量图像出现频率将显著降低,个性化功能能更精准理解用户审美偏好。V8.2 的个性化配置文件拥有更大、更优的图像选择池,建议用户尝试新旧配置文件体验最新版本。
Anthropic 发布 Claude Opus 5
2026-07-24Anthropic 发布 Claude Opus 5,其智能水平接近 Claude Fable 5,但价格减半。该模型在 Frontier-Bench v0.1 上性能超过 Opus 4.8 两倍以上,在 ARC-AGI 3 上得分是次优模型的三倍。Opus 5 即日起成为 Claude Max 的默认模型和 Claude Pro 的最强模型。
蚂蚁百灵发布Ling-3.0-flash原生混合推理模型
2026-07-24蚂蚁百灵发布新一代原生混合推理模型Ling-3.0-flash,总参数量124B,激活参数量仅5.1B,在传统推理、指令遵循与长文本等指标上对标甚至超越上一代旗舰Ring-2.6-1T。模型采用原生混合线性注意力架构与1/64稀疏MoE,并扩展至10,000+可交互训练环境,长输入下TTFT降低60%至80%以上。
Runway Agent 推出自然语言工作流功能
2026-07-24在 Runway Agent 中引入工作流。现在你可以通过自然语言构建、运行或编辑基于节点的工作流。工作流可大规模解锁高质量输出。 立即尝试,点击下方链接调用 / Workflow 技能。
百度搭子更新:电脑手机接力、桌面端内嵌浏览器上线,复杂任务可跨端连续执行
2026-07-24百度搭子在近期AI Day上推出多项升级,支持电脑与手机双端互联,同步任务上下文与执行进度,用户可跨设备接力完成复杂工作。桌面端内嵌浏览器正式上线,能自动打开多个网页执行调研、下载等操作,手机端支持云端远程操控。智能路由自动匹配任务模式,平均任务耗时降低20%,Token利用率提升25%;简单任务完成度达100%,复杂任务高交付率94%,积分消耗最高降低75%。
FLUX 3 x mimic:新一代视频动作模型
2026-07-24Black Forest Labs 发布多模态基础模型 FLUX 3,联合训练图像、视频和音频,其中视频预测占训练算力的 95% 以上。该模型与机器人公司 mimic 合作推出 FLUX-mimic,已在奥迪生产线上测试部署。加入动作预测后,视频生成质量最初下降最多 10%,但经 3500 步训练后恢复原有水平。
Black Forest Labs 发布 FLUX 3 多模态模型,支持单次生成 20 秒视频与原生音频
2026-07-24Black Forest Labs 以 Early Access 方式推出 FLUX 3 多模态基础模型,采用统一架构联合学习图像、视频和音频。该模型基于 Self-Flow 学习框架扩展,可在单次生成中输出最长 20 秒视频并附带原生音频,支持文生视频、图生视频、多镜头串联等任务。
OpenRouter 推出 Classifiers 测试版:自动标记 AI 请求的用途与成本归属
2026-07-24OpenRouter 上线 Classifiers 测试版,允许用户通过自定义分类法(最多 8 个维度)自动标记每次 AI 请求的任务类型、部门归属、合规类别等信息。分类异步运行,不增加推理延迟;支持采样率控制成本,推荐使用 Gemini 3.5 Flash Lite 作为分类模型。标记结果写入日志,并可在 Activity Explorer 中按维度聚合分析模型使用分布与成本流向。
04论文研究
9Anthropic 联合 Andon Labs 发布 Drone-Bench,评估 AI 模型自主操控无人机执行定位追踪任务的能力
2026-07-24Anthropic 与 Andon Labs 合作推出 Drone-Bench,用于测试 AI 模型自主操控四旋翼无人机在室内环境中定位并追踪指定人员的能力。该基准将任务分解为 3D 地图重建、定位、导航、目标检测与跟随五个子任务,并通过软件复现实现快速评估。实验表明,该任务链的难度足以区分不同智能水平的模型,并揭示 AI 在物理世界操控能力上的进步轨迹。
HARP: The Human--AI Research Platform
2026-07-24arXiv:2607.20773v1 Announce Type: new Abstract: Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys toward more conversational exchanges. Researchers studying HCI …
Flint: A Semantics-Driven Data Visualization Intermediate Language
2026-07-24arXiv:2607.20775v1 Announce Type: new Abstract: We present Flint, an intermediate language that enables authors to create high-quality visualizations from concise, semantics-driven specifications without explicitly confi…
Sonic Stage: Automatically Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Bli…
2026-07-24arXiv:2607.20835v1 Announce Type: new Abstract: Audio description (AD) makes film and television accessible to blind and low-vision (BLV) audiences by narrating characters' actions. However, in scenes with lots of dialog…
Exploring the Design Space of LLM-Based Programming Support in CS Education: A Scoping Review through the Lens of Assist…
2026-07-24arXiv:2607.21257v1 Announce Type: new Abstract: As large language models (LLMs) become integrated into programming education, learner-facing systems increasingly differ in how that assistance is bounded, enacted, and con…
Reimagining the Augmented Reality Accessibility Ecosystem for Deaf Students: Service Provider Perspectives in Experienti…
2026-07-24arXiv:2607.21289v1 Announce Type: new Abstract: In experiential learning environments, Deaf and hard of hearing (DHH) students often experience ``split attention,'' dividing their focus among tasks, instructors, and acce…
CRAFT: Exploring Wearable Creative AI on Smart Glasses for Fiction Writing in Real-World Contexts
2026-07-24arXiv:2607.21394v1 Announce Type: new Abstract: Creative writing increasingly integrates AI assistance, yet current tools miss in-situ moments when writers draw inspiration from real-world experiences. We envision Contex…
Thinkink: 2D Spatial Ink-native Interaction with LLMs
2026-07-24arXiv:2607.21468v1 Announce Type: new Abstract: People often use handwritten notes and sketches to externalize ideas for ideation. To integrate large language models (LLMs) into this practice, we propose Thinkink. Prompt…
A Needs Assessment for Measuring Geographic - Legislative Associations in the U.S. House of Representatives
2026-07-24arXiv:2607.21502v1 Announce Type: new Abstract: Political legislation affects the well-being and livelihoods of constituents. In the U.S. Congress a representative's voting record on bills and legislation is public. Thes…
05技巧与观点
2Claude 5 代模型上下文工程新规则:Claude Code 系统提示词精简超 80%
2026-07-24Anthropic 为 Claude Opus 5 和 Claude Fable 5 等新一代模型删除了 Claude Code 超过 80% 的系统提示词,且编码评测无显著损失。
Claude-thermos:保持 Claude 会话缓存热度,避免重新编码费用
2026-07-24Claude-thermos 通过本地反向代理监控 Claude Code 会话,在主智能体因等待子智能体而空闲超过 5 分钟时,自动发送预热请求刷新提示缓存。实测约 185 次本地会话中,缓存过期导致的重新编码占账单约 22%。工具以 uvx 运行,支持自定义空闲阈值和预热间隔。
