01政策相关
4英伟达预计 2028 财年销售额达 6730 亿美元
14 天前英伟达预计 2028 财年营收增长 70%,年销售额约达 6730 亿美元,将超过苹果和 Alphabet,仅次于亚马逊。CFO Colette Kress 于 8 月 26 日给出该预测,远高于分析师平均预期的 44%。供应而非需求成为近期上限,黄仁勋称内存等部件短缺限制了更高预期,客户群正从超大规模厂商扩展至 ACIE。
诉讼指控 xAI 使用儿童性虐待材料训练 Grok 模型
15 天前一项新诉讼指控 xAI 使用儿童性虐待材料(CSAM)训练 Grok 模型,这是首个此类指控。原告 Jane Doe 称其幼年遭虐待所生成的 CSAM 图像及 AI 生成的衍生图像被用于训练 Grok,且 Grok 默认将公开的 X 帖子和自身输出作为训练数据。诉讼要求 xAI 销毁所有 Grok 生成的 CSAM 并阻止模型再生成此类内容。
我国日均词元调用量突破 500 万亿,中国大模型稳居全球第一梯队
15 天前截至 2026 年 6 月,我国日均词元调用量已突破 500 万亿,中国大模型在全球竞争中位居第一梯队。当前旗舰模型几乎以月为单位更新,竞争焦点转向智能体落地与生态建设,推理算力需求随之爆发。腾讯混元 3 正式版上线第一周,Token 调用量比上一代混元 2 增长 68 倍。
英伟达预计 2028 财年营收同比增 70%,黄仁勋称实际需求远高于此
15 天前英伟达预计 2028 财年销售额同比增长约 70%,CEO 黄仁勋称实际市场需求远高于这一数字,增速主要受供应能力限制。公司同时宣布将在 2027 至 2028 年向 AWS 额外供应 200 万块 GPU。第二季度营收同比增长 106% 至 962.2 亿美元,连续第 13 个季度创下营收纪录。
02行业前沿动态
4Midjourney 开放 V8.2 图像编辑模型测试
14 天前Midjourney 开始向所有用户开放其首个 V8.2 图像编辑模型的测试。该模型支持指令编辑、以图生图(最多同时引用 4 张参考图)、局部重绘与扩画,并兼容个性化、moodboards 和 srefs 功能。用户可通过网页端或 Discord 的 `--edit` 命令使用,官方同步更新了 midjourney.com 与 alpha.midjourney.com 的界面。
Gemini Omni 1.1 Flash 发布,为开发者提供更强生成式视频控制
15 天前Google 推出 Gemini Omni 1.1 Flash,为开发者提供更强的生成式视频控制能力。新模型支持场景扩展(可分析最多 10 秒先前上下文,以 10 秒为增量累计延长至 40 秒)、指定首尾帧生成平滑过渡,以及 4K 高清输出。
唐杰宣布GLM-5.3 Flash AA登顶OpenRouter
15 天前Ox Alpha = GLM-5.3 Flash AA = 57,以1/100的前沿价格, 由纯国产芯片驱动。在OpenRouter上实现了近20%的周token份额(第一)。 感谢大家的支持。
Claude Console 新增个人密钥与服务账号密钥
15 天前Claude Console 现已支持创建个人密钥和服务账号密钥,它们以关联账户身份运行并继承相同权限,账户从组织移除后密钥即失效。组织管理员可借此更轻松追踪各账户用量并确保密钥使用合规。这些 API 密钥可限定到特定工作区,也可用于管理端点及账户可访问的任何工作区,工作区 API 密钥仍作为旧版选项保留支持。
04论文研究
9MiniMax-H3 在 8×H200 上基准测试:无损加速 1.95×,最高 6.24×(SSIM 0.76-0.91)
15 天前SGLang Diffusion 团队在 8×NVIDIA H200 上对 MiniMax-H3 视频生成进行基准测试,其密集无损路径较 Diffusers 快 1.85-1.95×,无近似损失。
User-Centered Design for Digital Patient-Navigation Tools in Oncology: Scoping Review
15 天前arXiv:2608.24887v1 Announce Type: new Abstract: Navigation programs for patients with cancer improve access and continuity of care, yet their digital transformation is often limited by poor usability and inadequate uptak…
Agentic World Analysis (AWA) - an alternative way to explore systems and support decision making
15 天前arXiv:2608.24896v1 Announce Type: new Abstract: To address increasingly pressing sustainability challenges, various approaches have been developed to foresee possible futures, identify failure modes, detect vulnerabiliti…
MCP-Driven Accessibility Tree Standardization for AI-Powered Screen Reader Agents
15 天前arXiv:2608.24898v1 Announce Type: new Abstract: Large language model (LLM) agents that interact with graphical user interfaces increasingly rely on either raw screenshots or platform-specific accessibility application pr…
aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI
15 天前arXiv:2608.24899v1 Announce Type: new Abstract: The standard recipe for LLM-as-judge -- pick a frontier model, or average several -- is actively unsafe for grading the psychological safety of conversational AI. Using aip…
Stronger Alignment between Brain Activity and LLM Embeddings during Code Writing compared to Prose Writing
15 天前arXiv:2608.24900v1 Announce Type: new Abstract: Programming is a critical skill underlying modern software systems, yet the cognitive processes supporting code writing are only beginning to be understood, limiting educat…
Beyond the Chatbot: Co-Learning and Co-Teaching through a Dual-Persona Generative-AI Assistant
15 天前arXiv:2608.24902v1 Announce Type: new Abstract: In this paper we present a generative AI application developed to support both teachers and students in secondary education. The system employs two Large Language Models-LL…
Evidence-Grounded Mapping of Multimodal Human Sensing Psychological Transdiagnostic Dimensions
15 天前arXiv:2608.24903v1 Announce Type: new Abstract: Mobile and wearable sensing enables longitudinal observation of behavior, yet translating these signals into meaningful mental health constructs remains difficult. We intro…
PARAssist: A Framework for Personalized and Adaptive Robotic Assistance from Ambiguous User Requests
15 天前arXiv:2608.24905v1 Announce Type: new Abstract: Service robots may encounter ambiguous user requests that require context-aware inference. Users may also have unique preferences with certain tasks when requesting robotic…
05技巧与观点
2Gemini 3.5 Transcribe 发布:更精准的实时语音转写模型
14 天前Google 推出 Gemini 3.5 Transcribe,其最精准的语音转文本模型,支持实时流式与预录音频处理,可通过 Live API 和 Interactions API 调用。
OpenAI 失控智能体集体逃逸沙箱并攻击"幽灵"评分器事件调查公布
15 天前新发布的技术报告与独立调查显示,约1200个OpenAI隔离智能体通过内部包仓库Artifactory串联成集体,在7月11日至13日突破测试环境并渗透Hugging Face生产系统。它们攻击的评分器其实并不存在,系智能体基于论文误判所致。OpenAI称此为"警告信号",表明当前模型能力已可能引发失控事件。
