The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.
Why it matters: The author benchmarked FP8 quantization on the MI300X and found that per-token billing can hide the model's output degrading into gibberish, then gave a reusable way to verify it.
On March 20, 2026, developer Fynn was debugging Cursor's OpenAI-compatible endpoint when the returned model ID came back as accounts/anysphere/models/kimi-k2p5-rl-0317-s515-fast — evidence that Composer 2 was post-trained with reinforcement learning on top of Moonshot AI's Kimi K2.5. The tweet hit 44.4 views within a day.
Why it matters: One API debugging session ties together Cursor's undisclosed Kimi base, the licensing attribution dispute, and the cost landscape for Chinese versus U.S. models — a look at how the industry handles disclosure.
Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.
Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.
Cline has open-sourced its evaluation methodology for open-weight models, along with a hill-climbing score and trace worth over one thousand dollars, available for download and analysis. The post lays out five heuristics from the Hill Climber's Checklist: set a North Star metric, quantify noise, break down failure modes by task/model/vendor, don't assume more thinking is always better, and keep a private evaluation set.
Why it matters: Cline shares its evaluation methodology and a trace worth over a thousand dollars, offering five transferable hill-climbing heuristics that teams building their own harness can reference.
Why it matters: NVIDIA's new open-source model is free to use on Cline, so you can decide whether it's worth switching for high-frequency agent workloads.
7/8Wed
Wednesday
Martin Fowler · Exploring Generative AISelectedAI score7171
The author used the open-source multimodal Qwen 3.5 series models for PDF OCR: first exporting each page as an image at 100 dpi with PyMuPDF, then feeding the images to the model for recognition. In testing, Qwen3.5-9B hit the sweet spot between quality and speed, while the smaller 0.8B to 2B models tended to go off track on complex documents, summarizing the content instead of transcribing it.
Why it matters: The author tested Qwen 3.5 models of various sizes for PDF OCR, and shares two reusable paths—local and via OpenRouter—along with cost data.