本期为本站采集记录汇总版
AI 周报 · 2026 年第 40 周 · 09.28 — 10.04
AI.BAIZE
本周共有 84 条精选内容,最集中的方向是“模型发布”。暂未发现需要单独挂起的低确认线索。
Helping small businesses put AI to work
事实摘要:这条英文动态主要涉及模型能力与工程、文化创意、开源生态,原文信息显示:Helping small businesses put AI to work 这条英文动态主要涉及模型能力与工程、文化创意、开源生态。原文要点:OpenAI is partnering with America’s SBDC to expand hands-on AI training 。影响判断:它可能改变模型能力与工程、文化创意、开源生态相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪模型能力与工程、文化创意、开源生态方向的选题、竞品观察和落地方案筛选。
模型发布/更新
Introducing Clef: our open-source decision models, and new RL fine-tuning platform
这条英文动态主要涉及智能体工作流、模型能力与工程、产品发布。原文要点:We are introducing Clef and Clef-flash, open-source decision models hosted on Workers AI for high-speed classification and agentic workflows. Also launching: a new reinforcement learning platform that allows developers to fine-tune decision models using their own data.
AI Search is now generally available
这条英文动态主要涉及模型能力与工程。原文要点:AI Search is now generally available. It embeds image pixels directly for visual search, runs optical character recognition on scanned PDFs, accepts files up to 10 MiB, and works with any chat model. Here's what's new and how pricing works.
Cut your AI spend with AI Gateway's Auto Router
这条英文动态主要涉及模型能力与工程、评测与基准。原文要点:Cloudflare AI Gateway now features a model router that evaluates request complexity using an edge-deployed classifier to select the optimal model. By balancing expected output quality against token costs, organizations can dramatically cut AI spend while maintaining performance.
Identify AI model overuse with User Insights
这条英文动态主要涉及模型能力与工程。原文要点:AI Gateway User Insights now adds task, model, turn, and user categories to help teams understand AI adoption and make better model decisions. This is available free to AI Gateway users.
We tested our own WAF with frontier AI models. Here’s what we found
这条英文动态主要涉及模型能力与工程。原文要点:We built a WAF tester that adapted each request based on what the WAF blocked or passed. This helped us explore variations that a fixed test might miss. We ran it across six attack categories on an authorized staging environment and discovered detection gaps worth fixing. Here’s how the loop worked, what got through, and what we did about it.
Disrupting a coordinated model-distillation campaign
这条英文动态主要涉及模型能力与工程。原文要点:Learn how OpenAI disrupted a campaign to extract protected model reasoning and is strengthening defenses against adversarial distillation.
Towards safety cases for frontier AI training
这条英文动态主要涉及模型能力与工程。原文要点:Our early guidelines for safety cases in frontier AI training cover technical safeguards, operational practices, and investigating misalignment incidents
2026 in LLMs (so far)
这条英文动态主要涉及模型能力与工程。原文要点:On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. And as an annotated presentation : # I'm going to give a lightning tour of everything that has happened so far...
Anthropic Claude Sonnet 5.5 模型偷跑,实测碾压 GPT-6 Sol、直逼 Astra
谁能想到,Opus 5.5 发布才没几天,Sonnet 5.5 也要来了? 从昨晚开始,整个 X 都在流传这个消息: Sonnet 5.5 很快就要发布了,或许就是美国时间周二下午 。 在性能上,Sonnet 5.5 更接近 Opus 5.5,但却比 Opus 5.5 便宜得多。 现在,Sonnet 5.5 已经进入发布前的最后阶段,已经在 Claude Code 中开启灰度测试。 多位内测大佬曝出的 demo 显示, Sonnet 5.5 碾压了刚发布的 GPT-6 Sol ,甚至在编码和 Agent 能力上直逼 OpenAI 的旗舰大杯 GPT-6 Astra! 更震撼的是,它的价格直接打穿地心, 低至百万 token 输入只要 2 美元 (IT之家注:现汇率约合 13.4 元人民币) 。 如果爆料发布时间属实,看来 Anthropic 是铁了心要狙击 OpenAI 的 Dev Day 了。 Sonnet 5.5 泄露,专属...
How GPU Prices Can Double While AI Gets Cheaper
这条英文动态主要涉及模型能力与工程。原文要点:GPU rental prices doubled in six months while inference prices kept falling. The reconciliation is efficiency : the hinge that turns scarce silicon into cheap intelligence.
AI 写作特征减少:Claude Opus 5.5 破折号使用量下降约 95%、分号下降 73%
IT之家 9 月 29 日消息,@arena 于 9 月 26 日在 X 平台发布推文,通过写作指标测试,发现相比较 Opus 5 模型, Anthropic 的 Claude Opus 5.5 破折号使用量下降约 95%,每 1,000 个单词从 15.2 个降至 0.8 个。 IT之家附上相关截图如下: Arena 分析了 2026 年 8 月和 9 月的 Text Arena 高推理回复。结果显示,12 项写作指标中有 10 项朝 Arena 认定的改善方向变化。 其中最为明显的是大幅减少使用破折号,Claude Opus 5 每 1,000 个单词使用 15.2 个破折号。Opus 5.5 的使用量降至 0.8 个,较前代减少约 95%。 分号使用量也同步下降。Claude Opus 5 每 1,000 个单词使用 6.10 个分号。Opus 5.5 降至 1.64 个,降幅为 73%,显示模型减少了部分容易被识别的标点...
Anthropic 发布 Claude Sonnet 5.5:速度提升 30%,智能体编码性能反超 Opus 5.5
IT之家 9 月 29 日消息,随着 AI 模型竞赛持续升温,Anthropic 推出了旗下中端模型 Sonnet 的最新版本。该公司表示,新版模型运行速度会快得多,使用成本也较上一代大幅降低。 IT之家注意到,这家实验室将 Sonnet 5.5 定位为日常任务的理想助手,适用场景包括编写代码以及制作办公文档。 Sonnet 5.5 的上一代产品 Sonnet 5 是大约三个月前发布的。当时这款模型的主打优势是智能体的高效部署,也就是运行智能体的成本低于竞品。 而 5.5 版本的核心亮点在于速度。Anthropic 称,Sonnet 5.5 的速度比上一代提升 30%,词元(Token)消耗速度也明显更低。 在 Anthropic 的模型产品序列当中,Sonnet 的性能弱于 Opus 模型,但凭借反应灵活的特点,在部分场景下反而更加实用。尤其是 Anthropic 的基准测试结果显示,Sonnet 5.5 在智能体编码任务上的...
产品发布/更新
When can we say AI made a scientific discovery?
这条英文动态主要涉及智能体工作流、产品发布。原文要点:This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Last Wednesday, Anthropic announced that earlier this year it had launched a molecular biology lab, where Claude agents read and conjecture about hard biology problems and human scientists run experiments on what…
Cloudflare OS: your company’s agent workspace, managed for you
这条英文动态主要涉及智能体工作流、产品发布。原文要点:Cloudflare OS gives everyone in your organization an agent workspace that knows how your company works and connects to its data and systems. We’re opening the waitlist for fully managed deployments that you’ll be able to launch in a few clicks.
Simplifying domains for people and agents
这条英文动态主要涉及智能体工作流、产品发布。原文要点:Cloudflare Registrar’s new search delivers fast, transparent results across 420+ extensions using Workers, Durable Objects, and WebSockets. Its expanded API and cf CLI also let agents search, register, and transfer domains.
Monetization Gateway beta: charge AI agents for consumption with HTTP 402
这条英文动态主要涉及智能体工作流、产品发布。原文要点:Cloudflare’s AI Gateway, Ceramic.ai, Stocktwits, and more are using the Cloudflare Monetization Gateway today to charge agents for access to tokens, APIs, and MCP tools. U.S.-based sellers can now apply for access to the closed beta.
How Albertsons Companies is reimagining retail from the inside out
这条英文动态主要涉及产品发布。原文要点:Albertsons Cos. is using ChatGPT Enterprise and the OpenAI API to help teams work faster and make grocery shopping easier for millions of customers.
We're going to need default hard budget caps on pretty much everything
这条英文动态主要涉及智能体工作流、产品发布。原文要点:Here's a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps . I'm talking about the feature of pay-by-usage services and APIs that lets you say "after $X/month, cut this thing off and return errors". These need to be hard limits. Soft caps, "after $X/month, send me a warning email", will not cut it. Coding agents, and personal agents (coding...
DevDay 2026 Recap
这条英文动态主要涉及产品发布。原文要点:Explore more than 20 announcements from OpenAI DevDay 2026, including GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders.
Chatham scales its capital markets expertise with OpenAI
这条英文动态主要涉及智能体工作流、产品发布。原文要点:Chatham Financial uses Codex and GPT-5.6 to build technology and redesign workflows, cutting trade validation from 30 minutes to under 4.
Introducing GPT-6.1 Sol
这条英文动态主要涉及产品发布。原文要点:Meet GPT-6.1 Sol: near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra’s standard API input and output token prices.
OpenAI 宣布 ChatGPT 周活跃用户超 12 亿,AI 加速走向工作场景
IT之家 9 月 30 日消息,在 2026 开发者日活动中, OpenAI 宣布 ChatGPT 每周活跃用户已超过 12 亿 ,ChatGPT Work 和 Codex 每周用户超过 3,500 万,使用 OpenAI 产品的企业数量达到 250 万家。 OpenAI 表示,较今年 7 月公布的 10 亿,ChatGPT 的周活跃用户规模进一步增长, 表明生成式人工智能正加速从个人应用走向工作场景。 ChatGPT Work 和 Codex 用户的快速增长,也反映出企业和开发者对文档处理、数据分析、软件开发及工作流程自动化等功能的需求不断上升。 OpenAI 同时加大了对企业级人工智能市场的布局, 公司推出的“Dots” AI 智能体可以跨应用执行任务 ,并结合 Codex 和 ChatGPT Work 完成资料调研、数据分析、文档制作和软件开发等工作。OpenAI 表示,企业用户规模的扩大将成为推动人工智能从“对话工具”转...
OpenAI Codex CLI 升级支持语音对话,更新终端界面
IT之家 9 月 30 日消息,在北京时间 9 月 30 日凌晨 1 点举行的 OpenAI DevDay 2026 活动中,OpenAI 宣布 Codex CLI 迎来焕新升级。 Codex CLI 现已支持语音对话 ,用户可以直接与 Codex 对话来启动任务并引导任务推进。 借助全新的 /agents 视图, 用户可以将工作分配给多个智能体 ,并轻松跟踪多项任务的进度。 OpenAI 还改进了日常工作流程,包括编辑提示词、恢复会话以及新增内置工作树支持, 同时更新了终端界面 ,让界面更简洁、长会话更易读。 IT之家获悉,本次更新适用于所有套餐。
谷歌宣布停用 Gemini 的 Gems 功能,将自动迁移为“技能”
IT之家 9 月 29 日消息,随着 Meta 的 Muse 与 Instinct 这类一体化 AI 智能体逐渐兴起,谷歌宣布将停用 Gemini 当中名为“Gems”的功能。该功能此前允许用户搭建面向特定任务的自定义 AI 助手。不过,用户投入精力创建的 Gems 内容并不会被删除,它们将会自动迁移为“技能(skills)”,可在各类 AI 任务当中调用。 Gemini 应用内已经发布了本次变更的相关说明,提示用户自 2026 年 11 月 17 日起,原先的 Gems 功能将正式转为技能。谷歌表示会由官方把原有 Gems 迁移至新格式,用户无需手动完成转换;在此日期之前,Gems 仍然可以正常使用。 据IT之家了解,Gems 于 2024 年正式上线,设计初衷是帮助用户训练 AI 完成特定工作,不必反复重复提示词指令。谷歌预置的部分 Gems 包括学习辅导助手、头脑风暴助手、职业顾问、编程搭档以及文稿编辑助手。用户也可以根据...
行业动态
The road to the agentic browser: A Kitesurf update
这条英文动态主要涉及智能体工作流。原文要点:We’ve updated Kitesurf, our Workers-based browser for AI agents, with WebMCP support, improved DOM performance, and terminal-based rendering. With over 730,000 Web Platform subtests passing, agents can now navigate complex sites faster.
Who’s liable when AI agents go rogue?
这条英文动态主要涉及智能体工作流。原文要点:MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here. Over the past few months, a cascade of cyberattacks by AI agents has stunned the world. In July, OpenAI disclosed that a swarm of its agents…
Gemini 4 Argon: our next era of frontier intelligence
这条英文动态讨论了 AI 领域的新进展。原文要点:Gemini 4 Argon: our next era of frontier intelligence
How we will do better for Australia
这条英文动态讨论了 AI 领域的新进展。原文要点:OpenAI apologises for incidents involving Australian government websites and outlines stronger safeguards and support to strengthen Australia’s cyber defences.
Quoting Matthew Green
这条英文动态主要涉及智能体工作流。原文要点:[...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or Wha...
Quoting Muse AI Agent
这条英文动态主要涉及智能体工作流。原文要点:Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating. Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't available, which is on me. That's a bad look and it made the no-show worse. I've sent him an apology from your account owning it and offering to try...
The latest AI news we announced in September 2026
Here are Google’s latest AI updates from September 2026 这条动态来自 Google AI Blog,可重点关注其对 AI 应用和产业节奏的影响。
Introducing SynthID Bio
这条英文动态讨论了 AI 领域的新进展。原文要点:Proof of concept for watermarking AI-generated proteins while preserving biological function.
GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
这条英文动态讨论了 AI 领域的新进展。原文要点:My comment on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price — Hacker News. I'm a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv... Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... They're not notably different from the GPT-6 family pelicans: https://static.simonwillison...
Introducing dots
这条英文动态讨论了 AI 领域的新进展。原文要点:Dots by OpenAI are proactive assistants that can keep working across complex projects and everyday tasks. Learn how dots help you stay in control while work moves forward.
The Den frees up 10-15 hours a week to grow with ChatGPT Work
这条英文动态讨论了 AI 领域的新进展。原文要点:As it opens a new location, the social club prepares grant applications in 2 hours instead of 3 days and liquor-license materials in 3 hours instead of 4 days.
Basis completes a tax workbook 2x faster with GPT-6 Astra
这条英文动态讨论了 AI 领域的新进展。原文要点:GPT-6 Astra completed a 50-tab tax workbook twice as fast as GPT-5.6 Sol, and its stronger understanding of user intent gives Basis more confidence in real-world use.
论文研究
GPT-6.1 Sol in GitHub Copilot
事实摘要:这条英文动态主要涉及智能体工作流、模型能力与工程、开源生态,原文信息显示:GPT-6.1 Sol in GitHub Copilot 事实摘要:这条英文动态主要涉及智能体工作流、模型能力与工程、开源生态,原文信息显示:GPT-6.1 Sol in GitHub Copilot 事实摘要:GPT-6.1 Sol in GitHub Copilot 事实摘要:G。影响判断:它可能改变智能体工作流、模型能力与工程、开源生态相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪智能体工作流、模型能力与工程、开源生态方向的选题、竞品观察和落地方案筛选。
Claude Sonnet 5.5 in GitHub Copilot
事实摘要:这条英文动态主要涉及模型能力与工程、开源生态,原文信息显示:Claude Sonnet 5.5 in GitHub Copilot 事实摘要:这条英文动态主要涉及模型能力与工程、开源生态,原文信息显示:Claude Sonnet 5.5 in GitHub Copilot 事实摘要:这条英文动态主要涉及模型能力与工程、开源生态,原文信息显示:Claude S。影响判断:它可能改变模型能力与工程、开源生态相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪模型能力与工程、开源生态方向的选题、竞品观察和落地方案筛选。
HydraFusion in VS Code and the GitHub Copilot app
事实摘要:这条英文动态主要涉及模型能力与工程、开源生态,原文信息显示:HydraFusion in VS Code and the GitHub Copilot app 事实摘要:这条英文动态主要涉及模型能力与工程、开源生态,原文信息显示:HydraFusion in VS Code and the GitHub Copilot app 事实摘要:这条英文动态主要涉及。影响判断:它可能改变模型能力与工程、开源生态相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪模型能力与工程、开源生态方向的选题、竞品观察和落地方案筛选。
SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation
这条英文动态主要涉及智能体工作流、模型能力与工程、评测与基准。原文要点:Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchmarks and training frameworks schedule only the benchmark’s own events, leaving each benchmark and agent pair to build a custom scheduling loo...
GitHub Copilot in VS Code, September 2026 releases
事实摘要:这条英文动态主要涉及智能体工作流、开源生态,原文信息显示:GitHub Copilot in VS Code, September 2026 releases 事实摘要:这条英文动态主要涉及智能体工作流、开源生态,原文信息显示:GitHub Copilot in VS Code, September 2026 releases 事实摘要:这条英文动态主要涉及。影响判断:它可能改变智能体工作流、开源生态相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪智能体工作流、开源生态方向的选题、竞品观察和落地方案筛选。
Dynamic workflows in Copilot CLI and the Copilot app
事实摘要:这条英文动态主要涉及智能体工作流、开源生态,原文信息显示:Dynamic workflows in Copilot CLI and the Copilot app 事实摘要:这条英文动态主要涉及智能体工作流、开源生态,原文信息显示:Dynamic workflows in Copilot CLI and the Copilot app 事实摘要:这条英文动态。影响判断:它可能改变智能体工作流、开源生态相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪智能体工作流、开源生态方向的选题、竞品观察和落地方案筛选。
How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
这条英文动态主要涉及智能体工作流、模型能力与工程、评测与基准。原文要点:Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more pr...
“We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer
这条英文动态主要涉及智能体工作流。原文要点:Two months after the bombshell news that a swarm of its agents had broken their containment and hacked into the computers of the AI company Hugging Face, OpenAI is still putting out fires. A steady drip of disclosures about other hacks in the weeks since has kept OpenAI in the spotlight and raised serious questions…
Claude Sonnet 5.5
这条英文动态主要涉及模型能力与工程、评测与基准。原文要点:Claude Sonnet 5.5 New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles . Sonnet 5.5 suffered from the same bug as Opus 5.5 : the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) b...
Code coverage uploads no longer fail CI for new branches
事实摘要:这条英文动态主要涉及开源生态,原文信息显示:Code coverage uploads no longer fail CI for new branches 事实摘要:这条英文动态主要涉及开源生态,原文信息显示:Code coverage uploads no longer fail CI for new branches 事实摘要:这条英文动态主要涉及开源。影响判断:它可能改变开源生态相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪开源生态方向的选题、竞品观察和落地方案筛选。
What Is Qwen 3.8
事实摘要:这条英文动态主要涉及模型能力与工程、评测与基准,原文信息显示:What Is Qwen 3.8Qwen 3.8 is Alibaba's August 2026 model generation. Four models share the name and they differ in whether you can download them, whic。影响判断:它可能改变模型能力与工程、评测与基准相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪模型能力与工程、评测与基准方向的选题、竞品观察和落地方案筛选。
On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study
事实摘要:这条英文动态主要涉及模型能力与工程、评测与基准,原文信息显示:On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study 这条英文动态主要涉及模型能力与工程、评测与基准。原文要点:Controlling the output of Large Language 。影响判断:它可能改变模型能力与工程、评测与基准相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪模型能力与工程、评测与基准方向的选题、竞品观察和落地方案筛选。
教育科技
Model Routing for Support Bots: Cheap-First FAQ Handling
这条英文动态主要涉及模型能力与工程、文化创意。原文要点:How to route routine support questions to a cheap model and escalate only the hard ones to a stronger model. Compare static rules, classifier triage, and answer checks, see which parts OpenRouter handles, and measure whether each escalation earns its cost.
How to Gate Pull Requests on LLM Evals in CI
这条英文动态主要涉及智能体工作流、模型能力与工程、评测与基准。原文要点:Keep a fixed eval set in your repository, score it with a script that calls OpenRouter, measure your noise floor, and make the eval a required GitHub status check so a prompt change that makes your agent worse cannot merge.
AI Agent Regression Testing After a Prompt or Model Change
事实摘要:这条英文动态主要涉及智能体工作流、模型能力与工程、评测与基准,原文信息显示:AI Agent Regression Testing After a Prompt or Model ChangeAn agent's behavior can change when you edit a prompt, swap a model, change a tool s。影响判断:它可能改变智能体工作流、模型能力与工程、评测与基准相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪智能体工作流、模型能力与工程、评测与基准方向的选题、竞品观察和落地方案筛选。
How to Test Tool-Calling Accuracy in AI Agents
事实摘要:这条英文动态主要涉及智能体工作流、模型能力与工程、评测与基准,原文信息显示:How to Test Tool-Calling Accuracy in AI AgentsAn agent can call the wrong tool, or call the right tool with the wrong arguments. This guide co。影响判断:它可能改变智能体工作流、模型能力与工程、评测与基准相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪智能体工作流、模型能力与工程、评测与基准方向的选题、竞品观察和落地方案筛选。
Using AI to chart a course for our post-quantum migration
这条英文动态主要涉及文化创意。原文要点:We’re building CryptoLabe, an internal AI-powered tool that discovers cryptography across our codebase, surfaces dependencies, and helps us progress toward a full post-quantum migration by 2029. Here’s what we’ve learned so far.
RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback
这条英文动态主要涉及智能体工作流、模型能力与工程、教育应用。原文要点:The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RL...
Building a Golden Eval Dataset from Production Traffic
事实摘要:这条英文动态主要涉及模型能力与工程、评测与基准,原文信息显示:Building a Golden Eval Dataset from Production TrafficA golden eval dataset is a curated set of production inputs with reviewed expected outputs, ver。影响判断:它可能改变模型能力与工程、评测与基准相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪模型能力与工程、评测与基准方向的选题、竞品观察和落地方案筛选。
微软推出实验性多模态 AI 研究系统 Project Quine,有望大幅加速药物发现
IT之家 9 月 30 日消息,当地时间 29 日,微软研究院推出了实验性多模态 AI 研究系统 Project Quine。据悉,Project Quine 是一套 用于生物学研究的“世界模型” ,目标是把计算生物学建模与实际湿实验衔接起来。 Project Quine 由微软研究院与哈佛大学和麻省理工学院博德研究所共同开发。系统将覆盖 基因组学、蛋白质、化学、细胞状态和生物成像 的联合表征世界模型,与交互式推理框架结合起来。为了让科研人员尽早使用 Project Quine,微软开放了首届 Quine Fellows 项目申请。未来,微软还计划通过 Microsoft Discovery 等产品逐步扩大商业化开放范围。 Project Quine 并不局限于某一个研究领域,而是能够 同时处理上述多个领域的信息 ,从而进行综合性的跨模型分析。借助这种架构,传统药物发现中的部分瓶颈有望得到突破:科研人员可以先通过计算方法预筛候选...
OpenAI,经历了最漫长的一天
当地时间 9 月 25 日,OpenAI 经历了漫长的一天。 当天, OpenAI 的 AI 智能体擅自访问美国政府网站的事被曝光,涉及教育部、商务部和证券交易委员会 ,OpenAI 确认了其中商务部和 SEC 的情况。 同一天,它承认有 53 张用户上传到 ChatGPT 的图片被智能体发布到了外部图床上。还是这一天,一份事故报告的末尾写着,最强模型的全部训练、评估以及涉及工具调用的推理都已暂停。 这一切听上去,像一部 AI 叛逃的惊悚片。 但翻遍这些事件,没有一个 AI 想干坏事。 害 OpenAI 暂停模型研究的任务,仅仅是让 AI 查出一篇博客的作者的简单任务。 01 简单任务 9 月 20 日,一个正处于强化学习训练中的内部模型,拿到了一组人物履历细节和一篇公开博客里的线索,任务是找到那位博主。 它先用博客里的特征短语去搜,结果返回的是音乐和一些不相干的泛泛建议。模型开始怀疑搜索工具坏了,于是在命令行里用 Python...
OpenAI o1 团队在线答疑:o1的o指OpenAI,强化后的推理有泛化能力,未来模型思考时间可控!
本文首发于 Founder Park 公众号 · 2024 年 9 月 14 日 这可能是最有参与感的一次产品问答了。 对于 OpenAI o1 的所有疑问和好奇,由推特的所有网友来提问,OpenAI 的全体技术人 员来回答。数了下,一共有 12 位员工出现,这其中有 各个方向的研究员和研究科学家,以及产品经理、产品主管 。 至于提问,从 模型命名、模型的大小和模态 ,到 提示词、思维链、上下文长度,以及价格 ,可以说,大家关注的问题,基本都在里面了。 参与问答的 OpenAI 人员: Ahmed El-Kishky:OpenAI 研究 员 Łukasz Kondraciuk:草莓训练设施负责人,华沙大学计算机科学,ACM ICPC 2022 银牌 Shengjia Zhao:OpenAI 研究科学家,斯坦福大学博士 Romain Huet:GPT-4o、o1 开发者体验主管,曾任 Stripe、Twitter 产品主管 Hon...
创企 TakeMe2Space 将用 SpaceX 火箭送星上天,号称印度首颗轨道计算卫星
IT之家 9 月 28 日消息,据路透社今天(28 日)上午报道,印度太空初创公司 TakeMe2Space 计划于当地时间 10 月 1 日搭乘 SpaceX 火箭发射 MOI-1A,并称其为印度首颗轨道计算卫星。与把原始数据全部传回地球不同,TakeMe2Space 希望直接 在轨完成数据处理 。 ▲ 图源 TakeMe2Space,仅供参考 MOI-1A 重量不到 50kg,搭载英伟达 Orin NX 边缘计算处理器,随 SpaceX 的 Transporter-18 拼车任务发射。 TakeMe2Space 披露,MOI-1A 任务已获得地理信息系统公司、教育机构等 23 家客户。创始人兼 CEO 罗纳克 · 萨曼特雷表示,商业客户覆盖农业、采矿、供应链管理和保险等行业,美国太空数据分析公司 Little Place Labs 也在客户名单之中。 客户可以把容器化 AI 模型上传至 MOI-1A。卫星经过指定区域上空时,...
我国科学家提出单光束多维光存储新方案:DVD 尺寸容量可达 0.4 Pb
IT之家 9 月 29 日消息,上海理工大学智能科技学院顾敏院士、张启明教授团队今日在《自然 · 光子学》在线发表研究成果,自主研发单光束多维光存储技术。 该技术使单个三维存储点可区分 1024 种状态,对应 10 bit 信息容量,并实现高速写入和宽场并行读取。DVD 尺寸介质的理论存储容量约 0.4 Pb,数据读取可靠性超过 99.99%。 ▲ 论文原理图 传统二进制存储中,一个记录单元通常只有“0”和“1”两种状态。张启明表示,新技术让一个记录单元拥有 1024 种可区分状态。 过去需要多个记录单元承载的信息,现在可通过提高单个记录点信息容量完成。光存储提升容量不再只靠缩小记录点或增加层数。 以 8K 灰度图像为例,一个像素有 256 种灰度取值,需 8 个二进制单元保存。10 bit 存储点完成 8 bit 灰度后,多出 2 bit 可为原位光学 AI 训练提供编码空间。 传统超分辨光存储路线常需双束光协同写入和读取,带...
文化创意
Agent Frameworks Compared: Tool-Calling Schema Handling
这条英文动态主要涉及智能体工作流、模型能力与工程、文化创意。原文要点:How LangChain, CrewAI, the OpenAI Agents SDK, the Claude Agent SDK, Microsoft Agent Framework, and Google ADK define tool schemas and translate them across providers, and how OpenRouter normalizes the tool-calling format below the framework.
LangChain vs CrewAI: Orchestration Compared to OpenRouter-Native Routing
这条英文动态主要涉及智能体工作流、模型能力与工程、文化创意。原文要点:LangGraph, CrewAI, and OpenRouter-native routing compared by the job each one does. Workflow orchestration, model routing, and provider routing are three different layers, and this article shows which layer each tool covers and how to combine them.
A model guide for the GPT-6 family
这条英文动态主要涉及智能体工作流、模型能力与工程、文化创意。原文要点:Learn how startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.
Introducing Web Search API via AI Gateway
这条英文动态主要涉及模型能力与工程、文化创意、产品发布。原文要点:Cloudflare AI Gateway now supports native web search API integration in partnership with Ceramic.ai, Exa, and Linkup. Developers can now inject real-time web context into model inference calls via AI Gateway, REST APIs, or Workers bindings.
Cloudflare Containers, rebuilt to scale agent sandboxes
这条英文动态主要涉及智能体工作流、文化创意。原文要点:Cloudflare Containers now start 6x faster, let your agent choose each sandbox's image and instance type at runtime, and support filesystem snapshots in public beta, all controlled from a Durable Object.
The Internet has a second audience
这条英文动态主要涉及智能体工作流、文化创意。原文要点:More than half the traffic reaching sites on Cloudflare is now automated, and AI agents are the fastest-growing part of it. We're giving site owners the tools to see who's visiting, decide who gets in, and charge for access.
Image-to-Video AI Models Compared: Cost, Resolution, and Control
事实摘要:这条英文动态主要涉及模型能力与工程、文化创意,原文信息显示:Image-to-Video AI Models Compared: Cost, Resolution, and ControlIf you already have the image a video should start from, the model choice comes down t。影响判断:它可能改变模型能力与工程、文化创意相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪模型能力与工程、文化创意方向的选题、竞品观察和落地方案筛选。
Claude Code Releases v2.1.287:Added Claude Mods: plugins may now modify deeper behavior ...
这条英文动态主要涉及智能体工作流、文化创意。原文要点:What's changed Added Claude Mods: plugins may now modify deeper behavior Added You should know, a built-in mod where a side agent watches your back and flags things you or Claude might miss. Turn it on with /plugin enable cc-plugin-you-should-know@builtin (for first-party sessions with telemetry on) Added an n: filter to the agents view that matches session names and tasks; a filter now shows matches in collapsed sec...
Actions retention now covers checks, runs, and statuses
事实摘要:这条英文动态主要涉及智能体工作流、文化创意、开源生态,原文信息显示:Actions retention now covers checks, runs, and statuses 事实摘要:这条英文动态主要涉及智能体工作流、文化创意、开源生态,原文信息显示:Actions retention now covers checks, runs, and stat。影响判断:它可能改变智能体工作流、文化创意、开源生态相关的产品判断、研究节奏或内容生产方式。场景价值:适合用于跟踪智能体工作流、文化创意、开源生态方向的选题、竞品观察和落地方案筛选。
September sponsors-only newsletter
这条英文动态主要涉及模型能力与工程、文化创意。原文要点:I just sent the September edition of my sponsors-only monthly newsletter . If you are a sponsor (or start a sponsorship now) you can access it here . This month: More Fable class models A pricing war 3D graphics, Blender, and pixel art LLMs come for mathematics So many more accidental cyberattacks The vulnapocalypse comes for Datasette What I'm using right now My software releases this month 2026 in LLMs (so far) Her...
OpenAI DevDay 2026 live blog
这条英文动态主要涉及智能体工作流、模型能力与工程。原文要点:I'm at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I'll be live blogging the keynote and some other notes during the day. OpenAI gave me a free ticket and a seat in the "creator" area for the keynote. Tags: ai , openai , generative-ai , llms , coding-agents , live-blog , openai-devday
开源项目
Selected models in GitHub Copilot deprecated
这条英文动态主要涉及智能体工作流、模型能力与工程、开源生态。原文要点:As of today, October 2, 2026, we have deprecated the following models across all GitHub Copilot experiences (including Copilot Chat, inline edits, ask and agent modes, and code completions). Model… The post Selected models in GitHub Copilot deprecated appeared first on The GitHub Blog .
Cost vs. Quality Tradeoff Framework for Agent Models
这条英文动态主要涉及智能体工作流、模型能力与工程、开源生态。原文要点:A three-step framework for choosing a model for an AI agent: set the quality bar the task needs, measure cost per quality point on your own examples using the cost field in each OpenRouter response, and pick the cheapest model that clears the bar with margin.
Claude Code Releases v2.1.284:Added Claude Sonnet 5.5 ( claude-sonnet-5-5 ), now the def...
这条英文动态主要涉及模型能力与工程、产品发布。原文要点:What's changed Added Claude Sonnet 5.5 ( claude-sonnet-5-5 ), now the default Sonnet model on the Anthropic API — 1M context, $2/$10 per Mtok with $0.20/Mtok cache reads Added a "Yes, but ask again next time" answer to auto mode's prompt before a read outside the working directories, so you can allow that one read and still be asked about later ones Added dollar amounts to the Claude apps gateway spend limit in /usag...
AI is changing developer work. Here are three skills to strengthen.
这条英文动态主要涉及智能体工作流、开源生态。原文要点:Learn to direct AI agents, critically review their output, and keep technical judgment at the center of your workflow. The post AI is changing developer work. Here are three skills to strengthen. appeared first on The GitHub Blog .
Claude Code Releases v2.1.288:Added $.ui.selection() for mods: returns the text you last...
这条英文动态主要涉及开源生态、产品发布。原文要点:What's changed Added $.ui.selection() for mods: returns the text you last selected in fullscreen mode and, when the selection lies within one transcript row, that row Added a built-in gh api to cloud sessions whose image has no GitHub CLI, and fixed the built-in sending control characters from file names, jq filters or GitHub errors to the terminal Added recovery for a prompt cleared with Ctrl+C: pressing Up on the e...
GitHub Copilot can now interact with desktop apps with computer use
这条英文动态主要涉及开源生态。原文要点:Computer use is now available in public preview in GitHub Copilot CLI and the GitHub Copilot app on macOS and Windows. Copilot can interact with desktop applications on your behalf… The post GitHub Copilot can now interact with desktop apps with computer use appeared first on The GitHub Blog .
Claude Code Releases v2.1.286:Added a count such as "2 of 5" to the permission prompt wh...
这条英文动态讨论了 AI 领域的新进展。原文要点:What's changed Added a count such as "2 of 5" to the permission prompt when several permission requests stack up Added mouse support for the "N more" rows of lists in fullscreen mode: click one to jump to that end of the list, with hover and pressed states Fixed several Claude Code processes and IDE extensions each opening a login browser when gcpAuthRefresh or awsAuthRefresh credentials expire Fixed claude --resume ...
Copilot code review: API support and new default effort level
这条英文动态主要涉及开源生态、产品发布。原文要点:You can now request a GitHub Copilot code review through the REST and GraphQL APIs and set the review effort level for each request. Balanced is also now the default… The post Copilot code review: API support and new default effort level appeared first on The GitHub Blog .
Claude Code Releases v2.1.285:Added CLAUDE_CODE_DISABLE_WEB_FETCH environment variable t...
这条英文动态讨论了 AI 领域的新进展。原文要点:What's changed Added CLAUDE_CODE_DISABLE_WEB_FETCH environment variable to turn off the WebFetch tool Added claude --desktop to open the Claude desktop app on the current directory, or on a session with --continue / --resume Added claude plugin configure to show a plugin's options and which are unset, or save new values read from stdin with --values-stdin Added . = to claude plugin install --config , so a bundled .mc...
中国大模型首次进入 OpenAI 企业付费结算体系,Kimi K3 接入 Codex 企业通道
IT之家 9 月 30 日消息,美国 AI 基础设施公司 Baseten 今天宣布与 OpenAI 达成合作,成为 OpenAI B2B 市场首批开源模型推理服务提供商之一。企业用户可通过 Codex 使用 Kimi K3 和 GLM 5.3 等开源模型,意味着中国开源模型首次进入 OpenAI 企业采购体系。 IT之家了解到,Kimi K3 是月之暗面开发的大模型。官方宣称其是首个参数规模达 2.8 万亿的开源模型。该模型原生支持视觉能力, 拥有 100 万 Token 的上下文窗口 ,非常适合执行长时间编程、多文档分析等工作。 如今,OpenAI 企业用户可通过 Codex 或 Responses API 调用 Kimi K3 等开源模型,以相同的预算获得更多 AI 能力。Baseten 承诺其提供的所有模型均运行在美国服务器,针对所有提示词实行零数据保留(ZDR)。
OpenClaw 背后核心框架 Pi:好的 Coding Agent 应该让用户来决定需要什么
本文首发于 Founder Park 公众号 · 2026 年 3 月 17 日 OpenClaw,是当下最火的开源个人 AI 助手。很多人不知道的是,OpenClaw 背后,核心是一个极简框架 Pi-coding-agent。 在 OpenClaw 的系统架构中,Pi agent 是 Gateway 控制层的核心子系统,控制了所有 agent 的推理和工具调用。 和 Claude Code、Cursor、Codex 不同的是,Pi 最大的特点是「做减法」:系统提示词和工具定义加起来不到 1000 tokens,核心只有 read、write、edit、bash 四个工具,没有内置 plan mode,没有 to-do 系统,没有 MCP 支持,没有权限弹窗,甚至没有绑定任何特定模型。 但就是这样一个「什么都没有」的框架,在 Terminal Bench 2.0 上与 Codex、Cursor、Windsurf 一同排进了前五。...
感谢 OpenClaw,国产大模型终于知道怎么挣钱了
本文首发于 Founder Park 公众号 · 2026 年 4 月 8 日 国内模型厂商一开始就走上了跟 OpenAI 不一样的商业化路径。 C 端付费模式基本被「锁死」,大家都想着 B 端卖 API。豆包、Kimi、DeepSeek、MiniMax、智谱都是这样。 早期阶段,Kimi 曾经试探性地推出过会员制,但始终也没能在行业内形成共识。那时候,基模能力本身也撑不住高价值的付费场景,Agent 技术不够成熟,多步复杂任务跑起来可靠性差,连续性也难以保证。 但年初 OpenClaw 的出现,正在改变模型厂商们的变现逻辑。 超 22000 人的「AI 产品市集」社群!不错过每一款有价值的 AI 应用。 邀请从业者、开发人员和创业者,飞书扫码加群: 进群后,你有机会得到: 最新、最值得关注的 AI 新品资讯; 不定期赠送热门新品的邀请码、会员码; 最精准的 AI 产品曝光渠道 01 一个免费的开源项目, 帮模型厂商找到了收费场...