Перейти к содержимому
Оригинал
Sebastian Raschka· @rasbt · X·· 21 день назадОценка ИИ36

Заголовок и краткое изложение на выбранном языке ожидают перевода.

Краткий обзор ИИ

Sebastian Raschka 用 Paint UI 复刻同一张图,对比 GPT-5.6 Astra 与 Qwen3.8 Max 的计算机使用(视觉)能力:Astra 默认用几何图形分层绘制,Qwen 则逐像素还原,后者结果天然更接近原图、得分会更高。但他认为不能据此断定 Qwen 的计算机使用或视觉理解更强,这恰恰说明只看最终结果的 benchmark 有多不可靠。

Полный текст

Полный текст на выбранном языке ожидает перевода. Пока показан оригинал.

Some food for thought when designing benchmarks...

So, here's a little computer-use (visual) comparison between GPT-5.6 Astra and Qwen3.8 Max.

The task here was to recreate the image in the center using the Paint UI.

Super interesting how the two different LLMs+Harnesses approached this totally differently by default.

I.e., Astra tried to approach this by drawing and layering geometric shapes. Qwen approached this pixel by pixel. (Of course, the pixel-by-pixel result looks closer to the original, it's essentially a low-res version of that by nature.)

So, the Qwen-generated image would surely score higher in the sense that it's closer to the original.

But I wouldn’t conclude from this example that one LLM generalizes better than the other on other tasks. Also, I wouldn't say Qwen has better compute-use capabilities or better visual understanding than Astra.

But it highlights an interesting point about how slippery benchmarks are when they only compare final results.

Источник: Sebastian Raschka · x.com