01政策相关
102行业前沿动态
8Claude in Chrome 侧边栏升级为 Claude Cowork 会话
2026-08-12Claude in Chrome 浏览器扩展的侧边栏现已升级为 Claude Cowork 会话,对话会保存至历史记录,技能和连接器可在浏览器中工作,且任务可在桌面、网页和移动端应用间无缝切换。
阿里开放 Qwen3.8-2.4T-A95B 模型权重:2.4T MoE、激活 95B、原生 256K 上下文
2026-08-12阿里 Qwen 团队正式开放 Qwen3.8-2.4T-A95B 模型权重,这是 Qwen-Max 级别模型首次开源。模型总参数 2.4T,每个 Token 激活 95B,原生支持 262,144 Token 上下文并可扩展至 1,010,000 Token。
微软首发自研推理模型MAI-Thinking-1
2026-08-12我们的首个推理模型 MAI-Thinking-1 从零开始构建,现已在 Microsoft Foundry 上线。为团队点赞!更多详情如下。
Meta 开源 Muse Glimmer 登陆 OpenRouter
2026-08-12Meta AI 超级智能实验室的首个开放权重模型 Muse Glimmer 已在 OpenRouter 上线! 这是一款 30B 密集文本+图像模型,采用 Apache 2.0 许可证,旨在成为可靠的本地智能体,MCP Atlas 得分 75.5,SWE-Bench Pro 得分 51.2。 https://openrouter.ai/meta/muse-glimmer-30b
LTX-2.5 模型登场:AI 生成 10 秒 720P 视频仅需 6.8 秒,原生集成 ComfyUI
2026-08-12LTX 推出 LTX-2.5 模型,原生集成 ComfyUI,在 2 张英伟达 GB200 配置下生成 10 秒 720P 视频仅需 6.8 秒。LTX-2.5 Fast 以每秒 0.09 美元生成带音频 720p 视频,10 秒片段成本 0.90 美元;年度经常性收入低于 1,000 万美元的组织可免费使用。
Cursor 与 SpaceXAI 联合发布 Grok 4.6
2026-08-12Cursor 与 SpaceXAI 今日发布 Grok 4.6,重点强化长时运行智能体与交互式视觉任务,在多项智能体编程与知识工作基准上达到前沿水平,并在 Artificial Analysis Intelligence Index 上追平 GPT-5.6 Sol。
xAI 发布 Grok 4.6,强化长时运行智能体能力
2026-08-12xAI 今日发布 Grok 4.6,在 Grok 4.5 基础上重点强化长时运行智能体及更复杂的交互式与视觉工作能力。该模型在多项智能体编码与知识工作基准上达到前沿水平,在 Artificial Analysis Intelligence Index(九项基准综合分)上追平 GPT-5.6 Sol。
OpenRouter 推出实时网页搜索基准测试:如何为智能体选择引擎、深度与模型
2026-08-12OpenRouter 发布实时排行榜,系统评测模型、搜索引擎、搜索方法与预算四类配置组合。数据显示,将搜索预算从 1 轮增至 25 轮可使 BrowseComp 得分近乎翻倍,成本仅增 2.5-7 倍;模型选择比引擎更重要,平均分差 15 分 vs 10 分。失败率高的任务应降低搜索深度以控制成本。
04论文研究
9空货架还是丢钥匙?Google 研究:Recall 是参数化事实性的瓶颈
2026-08-12Google Research 提出知识画像框架,发现前沿 LLM(如 Gemini3、GPT-5)的事实编码接近饱和,但回忆(recall)能力不足,多数事实错误源于"丢钥匙"而非"空货架"。该框架将事实分为编码失败、回忆失败等五类画像,并配套推出 WikiProfile 基准,含 2,150 条维基百科事实,每条配 10 个问题,用于分别探测编码、回忆与识别能力。
How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation
2026-08-12arXiv:2608.09939v1 Announce Type: new Abstract: Production teams deploying LLM chat agents face a specific quality assurance gap: existing evaluation tools test individual responses or simulate social interactions, but n…
The impact of design factors of virtual and augmented reality on tertiary students user experience in Metaverse
2026-08-12arXiv:2608.09940v1 Announce Type: new Abstract: The Metaverse is a convergent space integrating virtual reality (VR) and augmented reality (AR) technologies, with market projections rising from \$65.5 billion in 2022 to …
EweAcT: Ewe behaviour aligned to accelerometer data for activity monitoring in extensive grazing systems
2026-08-12arXiv:2608.09943v1 Announce Type: new Abstract: Monitoring livestock behaviour under extensive conditions would provide valuable insights to assess animal adaption to environmental perturbations in agroecological systems…
Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents
2026-08-12arXiv:2608.09944v1 Announce Type: new Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure behind visual layout. Re…
Co-Lecturing With the DED: Explaining Circuit Design via the Draw Encode Display Loop
2026-08-12arXiv:2608.09945v1 Announce Type: new Abstract: When representing digital circuits, 2 dimensional hand drawings free us from the linear structure of hardware description languages, enabling intuitive reasoning and making…
HoosierHelp: Benchmarking LLM Agents for Social Service Navigation
2026-08-12arXiv:2608.09946v1 Announce Type: new Abstract: Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising…
Mapping Multimodal Pilot Stress and Fatigue During Flight Sessions
2026-08-12arXiv:2608.09947v1 Announce Type: new Abstract: This study analyzes patterns of stress and exhaustion among student pilots throughout flight training using a combination of physiological and self-reported measurements. T…
Immersive Micromanipulation Integrating Pipette and Injector Operations with McKibben-Based Haptic Sensations for Worklo…
2026-08-12arXiv:2608.10033v1 Announce Type: new Abstract: Intracytoplasmic sperm injection (ICSI) requires advanced micromanipulation techniques but relies solely on visual feedback and involves frequent interface switching betwee…
05技巧与观点
4DeepSeek V4 Pro与Grok 4.6同日发布,双双逼近Claude Fable 5体验
2026-08-12DeepSeek V4 Pro正式版与Grok 4.6在2小时内先后发布,均为1.6T/1.5T参数模型,逼近Claude Fable 5体验。
AutoGPT 如何用 AGENTS.md 和技能门控管理 AI 生成的拉取请求
2026-08-12AutoGPT 维护者发现,AI 智能体不会主动阅读文档,因此将指令放在 AGENTS.md 和技能文件中,并置于代码目录旁。他们通过强制 PR 模板、测试计划、CI 覆盖率门槛和 CLA 签名等门控机制,将智能体提交的 PR 从"不可用"转变为"可用但不符合路线图"。其中 CLA 签名因需浏览器和 OAuth 流程,被用作区分人类与智能体的"人类探测器"。
我写了一本 AI 教科书--AI 还要多久才能写得更好?
2026-08-12作者在完成一本 RLHF 教科书后反思,LLM 在长文非虚构写作上进展停滞,GPT 4.5 和 Kimi K2 等写作强模型已显老旧,而编码、数学等任务已接近超人水平。模型能改错字、做编辑,但组织整章内容时仍混乱且易出错,作者认为这阻碍了模型自主解决开放科学问题。
零基础用户半天上手AI的12步实操流程
2026-08-12文章给出一套零基础用户半天上手AI的12步实操流程:准备内存不低于16G的电脑,订阅ChatGPT并安装Codex或使用WorkBuddy,用语音输入法以【背景、痛点、需求】框架向AI交代任务,经苏格拉底提问澄清需求后投喂文件让AI直接完成,最后沉淀为可复用Skill。文中建议Codex选GPT-5.6 Sol最高模型,WorkBuddy选Kimi K3。
