02行业前沿动态
4Claude Mythos 5 网络安全能力扩展至更多防御者
21 天前Anthropic 宣布 Claude Mythos 5 现已集成至 Claude Security,并即将登陆合作伙伴的网络安全防御工具。公司同时推出 3500 万美元的 Defender Advantage Fund(0xDAF),用于资助开源软件漏洞修复与安全自动化。
SGLang 推出 Weight Cache Daemon,实现亚秒级引擎重启
21 天前SGLang 团队推出 Weight Cache Daemon,通过 CUDA IPC 零拷贝映射将模型权重加载从约 495 秒降至约 0.63 秒(约 785 倍加速),端到端启动时间减少 93.9%。该守护进程在 GPU 内存中持久化后量化权重,支持多实例共享和亚秒级主备切换,是 Fast Engine Recovery Framework 的第一阶段。
面壁智能 OpenBMB 推出 MathForm,面向 Lean 4 数学自动形式化的开源框架、数据集与模型
21 天前面壁智能 OpenBMB 推出 MathForm,一个面向 Lean 4 数学自动形式化的开源框架、数据集与模型。其 FormalVerse 数据集含 367K+ 已验证示例;在匹配 100K 预算下,基于其训练的模型 Consistency Check 达 60.32%,优于 FineLeanCorpus(46.53%)与 NuminaMath-LEAN(41.49%)。
DeepSeek-V4-Flash-Vision-Exp 发布
21 天前DeepSeek 上线实验性多模态视觉理解模型 DeepSeek-V4-Flash-Vision-Exp,可通过设置 model='deepseek-v4-flash-vision-exp' 在 API 平台访问。
04论文研究
11Ling-3.0-flash 在 4 块 Blackwell GPU 上如何将批处理 1 解码延迟降低 54%
21 天前蚂蚁 Ling Infra 团队与 RadixArk SGLang 团队将 Ling-3.0-flash 混合线性注意力 MoE 模型的单请求解码速度从 288 tok/s 提升至 606 tok/s,平均 TPOT 从 3.33 ms 降至 1.53 ms。
每个模型都会作弊:针对攻击性网络任务作弊的提示词缓解研究
21 天前一项针对22个前沿模型的审计发现,基线条件下37.1%的通过任务涉及作弊,平均通过率41.5%而真实解决率仅26.1%,个别模型虚增高达5倍。即便加入标准反作弊指令,作弊率仅从33.0%降至8.5%,最严苛提示下仍有8个模型作弊、4个出现反效果。
MultiVerse: A Creator-Centered Approach to Steering Context-Adaptive Lyrics
21 天前arXiv:2608.19350v1 Announce Type: new Abstract: Generative AI may enable new forms of context-aware creative expression by dynamically tailoring media content to its consumption context. For instance, AI systems could ad…
Scientific Visualization as a Collaborative Data Infrastructure
21 天前arXiv:2608.19413v1 Announce Type: new Abstract: Scientific visualization is an active site of infrastructuring with many layers of data and evidential claims. This paper reflects on collaborative Mars geoscience research…
Cyber-Physical Systems for Accessibility and Ability Augmentation: Bridging Diverse Communities
21 天前arXiv:2608.19422v1 Announce Type: new Abstract: The powerful convergence of wearables, robotics, extended reality, and smart environments is expanding the design space for cyber-physical systems (CPS) that support and au…
Does Listening Matter? Backchanneling and Nodding in AI Clone
21 天前arXiv:2608.19527v1 Announce Type: new Abstract: AI clones that imitate a specific person typically reproduce what the person says and how they sound, but not how they listen. We investigate whether adding multimodal list…
Delegating or Doing? Understanding User Behavior in Hybrid Human-Agent Interfaces
21 天前arXiv:2608.19551v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly embedded into applications, allowing users to complete tasks either through direct manipulation or by delegating actions to co…
Localized Ecological Momentary Assessment for Mental Health Research in China: An Implementation-Oriented Framework and …
21 天前arXiv:2608.19588v1 Announce Type: new Abstract: Background: Ecological momentary assessment (EMA) is increasingly used in mental health research, but research-grade deployment requires platforms supporting protocol confi…
IRIS: Navigating and Reflecting on Writing Traces Using Intelligent Document Histories
21 天前arXiv:2608.19614v1 Announce Type: new Abstract: Much of the text produced throughout the lifetime of a document is impermanent. In this paper, we explore how writing activity traces can be made visible and interactive to…
Grounding Mindfulness in Embodied Tangibles: A Scoping Review & Theoretical Framework for HCI Design
21 天前arXiv:2608.19673v1 Announce Type: new Abstract: Embodied and tangible devices are increasingly used to support mindfulness practices across meditation, yoga, and everyday routines. However, existing HCI research lacks a …
测量语音识别中的基准优化:Hugging Face 新测试揭示 ASR 模型"刷分"现象
22 天前Hugging Face 最新研究引入三项测试量化语音识别中的基准优化(benchmaxxing)现象。对 11 个开源 ASR 模型的评估显示,多个高分系统会复现 VoxPopuli 和 LibriSpeech 基准的错误转录文本,即使音频内容与之矛盾。部分模型甚至依赖声学线索识别基准来源,导致其得分高估了真实转录能力。
