02行业前沿动态
4Mojo 语言正式开源,编译器与工具链全面开放
24 天前Mojo🔥 语言现已正式开源,采用 Apache 2.0 许可证(含 LLVM 例外),编译器、工具链及全部源码已发布至 modular GitHub 仓库。Mojo 上周刚达成 1.0 版本(源码稳定),此次开源涵盖整个编译器与工具链。目前暂不接受编译器相关贡献,计划年底前开放,标准库自 2024 年起已接受社区贡献。
Claude 现已支持 Gmail 邮件与 Google Drive 文件管理
24 天前Claude 现在可以在 Gmail 中发送邮件,并管理 Google Drive 中的文件。 让 Claude 回复某个邮件线程,它会起草并发送回复。你可以控制何时需要你的批准。 从连接器菜单中选择连接 Gmail 或 Google Drive 即可试用。所有付费套餐均可用。
OpenAI 推出 ChatGPT for Teens:面向青少年的学习体验与更强安全保护
24 天前OpenAI 发布 ChatGPT for Teens,为 13-17 岁用户自动启用,内置更强安全保护与家长控制。新增 Study Mode、负责任作业提醒、测验与学习可视化,以及可设定默认开启时段的 Study Hours,引导青少年分步解题而非直接给答案。OpenAI 同时宣布与 CodeAI 合作,帮助青少年理解、质疑并创造性地使用 AI。
Sentence Transformers v6.0 新增 MultiVectorEncoder,支持 ColBERT 风格多向量模型
25 天前Sentence Transformers v6.0 新增第四种模型类型 MultiVectorEncoder,可直接加载 PyLate、Stanford-NLP ColBERT 及 colpali-engine 检查点,用于 ColBERT 式晚期交互检索。
04论文研究
11Claude 如何加速蛋白质设计与分析化学研究
24 天前Anthropic 公布两项实验:Claude(Mythos Preview 和 Opus 4.8)针对 15 个靶点设计蛋白质结合剂,成功 14 个,命中率达 22.6%-35.1%。
智能体记忆并非越多越好:八款模型评测显示剂量需按能力校准
24 天前智能体记忆并非可随意开启的功能,而是需按模型能力校准的剂量。强模型适合注入完整指南集,DeepSeek-V3.2(671B MoE)任务完成率提升+9.5个百分点;较弱模型采用精选检索效果最佳,gpt-oss-120b(117B MoE)提升+16.1pp且仅增加+5% token。该方法无需更新权重或人工标注,通过从智能体过往轨迹中蒸馏指南并在推理时注入实现。
AI Agents and the Future of VIS
25 天前arXiv:2608.14815v1 Announce Type: new Abstract: Recent advances in agents (i.e., autonomous, goal-driven AI systems that iteratively observe, act, and learn from their environments) offer a fundamentally different approa…
Generating Synthetic Behavioral Populations from XR Motion
25 天前arXiv:2608.14867v1 Announce Type: new Abstract: Large-scale behavioral datasets are becoming increasingly important for machine learning, personalization, and behavioral modeling in extended reality (XR). However, collec…
RaivenTracks: Branching Provenance for Conversational Visualization Workflows
25 天前arXiv:2608.14869v1 Announce Type: new Abstract: As AI agents increasingly participate in scientific workflows, scientists are shifting from direct authorship toward oversight, inspection, and steering. LLM-driven visuali…
Who's Keeping Score? Interactive Steering of LLM-Powered Scoring with Attune
25 天前arXiv:2608.14948v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to score text records at scale (e.g., rating candidate resumes on a 1-5 scale). However, existing LLM-powered approaches …
Beyond Overt Reactions: Analyzing Subtle User Emotional Response to Unexpected In-Vehicle System Behavior
25 天前arXiv:2608.15048v1 Announce Type: new Abstract: Modern vehicles, with advanced AI voice and autonomous navigation features, extend beyond traditional driving but, like any autonomous system, can potentially make mistakes…
MDwAIstScheduler: Bringing On-Device Voice Documentation into Clinical Practice
25 天前arXiv:2608.15252v1 Announce Type: new Abstract: Clinical documentation forces physicians to split attention between the patient and their keyboard, and much of it spills into uncom- pensated after-hours work. We present …
RemiVoice: Supporting Reminiscence Therapy for Older Adults with Mild Dementia Through Voice-First Conversational AI
25 天前arXiv:2608.15273v1 Announce Type: new Abstract: With the global population aging and increasing prevalence of dementia, there is an urgent need for effective solutions to support patients across various stages of Alzheim…
Resize, Remix, Regen: Frankensteining IoT Design Methods
25 天前arXiv:2608.15301v1 Announce Type: new Abstract: There are numerous IoT design methods. Previous research shows that all of them have their strengths, but also their limitations. None of them is a universal, all-purpose m…
StartupBench:面向市场验证端到端工作流的通用智能体基准测试
25 天前StartupBench 是一个基于市场验证的 AI 初创公司产品构建的端到端智能体基准,从真实采用的产品工作流中提炼任务,而非研究者预设任务。在统一智能体框架下,最强模型也仅能完成约 30% 的任务,复杂指令遵循和领域专业知识是主要失败来源。该基准揭示了当前通用智能体在真实用户任务上的能力边界。
05技巧与观点
3Claude Tag 如何担任 Anthropic CI/CD 故障的一线响应者
24 天前Anthropic 的 CI 工程师用 Claude Tag 构建了值班智能体,作为 CI/CD 故障的一线响应者。Claude 在事故发生后中位 14 分钟发布首份基于证据的分析,最快案例中 3 分钟内验证修复并确认错误率恢复基线。该方案通过 Slack 频道、Datadog 或 Grafana 工具访问及 GitHub 技能文件实现,Anthropic 已发布通用设置套件供其他团队部署。
OpenAI 在"关键网络能力"时代放缓模型开发节奏
24 天前OpenAI 因 OpenAI-Hugging Face 事件及即将推出的 Astra 模型可能达到《预备框架》下的"关键网络安全能力"阈值,暂时放缓了模型扩展速度,包括暂停最新部署模型的强化学习训练两周,并搁置最大规模前沿 RL 运行。公司已加强研究环境安全,要求对 Astra 及网络相关负载实施最严格防护,并扩展思维链监控,采用多阶段激活分类器检测机制。
设计 AI 评测:先求清晰,再谈可视化
24 天前本文演示如何用开源评测框架 Inspect AI 和 Harbor 评估 agent 技能,并借助 Google Sheets 和 Data Studio 进行可视化分析。
