核心信息
Perplexity 在 WANDR 基准上对 GPT-6 Astra 进行了评测,其得分达到 0.682,单任务成本为 11.98 美元,是目前所有已测模型中的最高分。
要点
- GPT-6 Astra 比 Fable 5.1 得分高 13.5%,而单任务成本低 6.1%。
- 相比 Opus 5,得分高出 27.0%,成本仅高 3.3%。
- 这一结果使 GPT-6 Astra 成为 Perplexity 在 WANDR 上目前测试过的领头模型。
Perplexity 在 WANDR 基准上对 GPT-6 Astra 进行了评测,其得分达到 0.682,单任务成本为 11.98 美元,是目前所有已测模型中的最高分。
We evaluated GPT-6 Astra on WANDR. It scored 0.682 at $11.98 per task, the highest score of any model we tested. GPT-6-Astra scored 13.5% higher than Fable 5.1 at 6.1% lower cost, and 27.0% higher than Opus 5 at 3.3% higher cost. https://t.co/SyYmD38qvq