Writing
Commentary
37 write-ups filed under commentary.
Inference-Time Compute Is a New Scaling Axis
o1 improves when it is allowed to spend longer on a problem. A benchmark score without a compute budget is an incomplete number.
The Final Pixel Won't Come From the Renderer
Geometry, camera and scene structure stay reliable ground truth. More of final appearance is moving into learned systems.
Open Models Are Becoming Research Infrastructure
Llama 3.1 gets discussed as a benchmark result. The licence terms change which experiments are possible at all.
The Model Is Becoming a Runtime
Function calling, code execution and structured output turn inference into a loop. The model stops being a text generator and starts being a control layer.
Multimodality Changes the Architecture, Not Just the Interface
GPT-4o is easy to read as a faster interface. Training one model end to end across text, vision and audio is an architectural change.
Benchmark Scores Are Not Model Capability
Claude 3 posts 86.8% on MMLU and 50.4% on GPQA Diamond. The chart is useful and it is not the same thing as capability.
Long Context Is Not Memory
Gemini 1.5 makes a million tokens usable. A larger working set is not a system that decides what should survive the session.