Coding
FrontierSWE v2
长时间工程能力:在 20 小时内完成高难度软件、科研与视觉任务
它测什么
FrontierSWE v2 由 Proximal Labs 推出,考察编程 Agent 能否长时间自主推进复杂项目。它包含 34 个任务,覆盖系统实现、性能优化、科学计算、视觉推理和 AI 研究,每项任务运行 5 次,每次最多 20 小时。
你可以把它理解成一场持续一整天的工程挑战:不是回答“怎么写”,而是实际写出程序、测试、发现问题并反复改进。例如,用 Zig 重新实现 Git,在 SQLite 上构建兼容 PostgreSQL 的服务,使用 Remotion 复刻一段视频,或训练仅凭摄像头画面驾驶赛车的机器人。
官方使用 Proximus 框架,为模型提供 Linux 环境,部分任务配有 GPU;框架支持查看图片、压缩上下文和保存候选提交,鼓励模型在剩余时间内继续优化。成绩由自动验证器根据最终成果计算,因此反映的是模型与这套 Agent 框架组合的表现,不等同于它在 Codex 或 Claude Code 中的表现。
怎么看分数
每项任务先得到 0–1 的分数,再将五次试验的平均成绩跨任务汇总为 mean@5,以百分比显示;65% 不代表完整解决了 65% 的任务。分数下方的 worst@5–best@5 是每项任务五次试验中最差与最好成绩汇总后的范围,不是 95% 置信区间。费用以美元计,耗时以小时计,均为单次试验的平均值,不是跑完整套评测的总开销。本页收录官网主榜 All 视图的全部 14 个模型,统一使用主榜的得分、费用和耗时口径;分析图与主榜的汇总值有细微差异,未混用。v2 新增 21 项任务、移除 4 项旧任务,且改变了评分方式,不能与 v1 的排名直接比较。官方仓库已公开任务,但独立运行器与公共预构建镜像仍标为待发布。
模型表现
GPT-6 AstraOpenAIAgent · Proximus | Proximus | 65.5%55.1–73.0%worst–best@5 | $1,029.65 | 12.1 小时 |
|---|---|---|---|---|
Claude Fable 5.1AnthropicAgent · Proximus | Proximus | 56.3%144.5–66.7%worst–best@5 | $138.55 | 11.6 小时 |
Claude Opus 5AnthropicAgent · Proximus | Proximus | 52.0%39.1–60.7%worst–best@5 | $196.71 | 16.1 小时 |
Claude Fable 5AnthropicAgent · Proximus | Proximus | 47.0%33.9–61.8%worst–best@5 | $301.28 | 15.8 小时 |
GPT-5.6OpenAIAgent · Proximus | Proximus | 32.2%22.2–44.2%worst–best@5 | $179.64 | 8.6 小时 |
GLM-5.3Z AIAgent · Proximus | Proximus | 30.2%17.7–40.6%worst–best@5 | $97.22 | 17.0 小时 |
Kimi K3Moonshot AIAgent · Proximus | Proximus | 25.9%13.9–37.6%worst–best@5 | $109.71 | 18.4 小时 |
Grok 4.6xAIAgent · Proximus | Proximus | 25.3%12.9–37.3%worst–best@5 | $243.43 | 13.8 小时 |
Gemini 3.7 FlashGoogleAgent · Proximus | Proximus | 20.3%10.3–31.2%worst–best@5 | $34.14 | 7.9 小时 |
Gemini 3.8 FlashGoogleAgent · Proximus | Proximus | 19.6%9.5–31.6%worst–best@5 | $38.75 | 7.2 小时 |
Qwen3.8-MaxAlibabaAgent · Proximus | Proximus | 15.8%8.7–24.3%worst–best@5 | $55.14 | 18.5 小时 |
DeepSeek V4 Flash Vision ExpDeepSeekAgent · Proximus | Proximus | 14.8%7.0–26.5%worst–best@5 | $8.57 | 14.8 小时 |
Muse Spark 1.2MetaAgent · Proximus | Proximus | 12.0%6.8–18.4%worst–best@5 | $27.81 | 4.6 小时 |
InklingThinking MachinesAgent · Proximus | Proximus | 4.1%0.1–10.8%worst–best@5 | $9.15 | 1.2 小时 |