2026-09 - W3
Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
LLM benchmarks are mostly marketing to be honest at this point.
GitHub - mattpocock/dictionary-of-ai-coding: AI coding jargon, explained in plain English. · GitHub
Interesting glossary.
The contagion of fear | The Observation Deck
No one actually wants simplicity - lukeplant.me.uk
I think a good test of whether you truly love simplicity is whether you are able to remove things you have added, especially code you’ve written, even when it is still providing some value, because you realise it is not providing enough value.
Another test is what you are tempted to do when a problem arises with some of the complexity you’ve added. Is your first instinct to add even more stuff to fix it, or is it to remove and live with the loss?
The only path I can see through all this is to cultivate an almost obsessive suspicion of FOMO. I think that’s probably key to learning to say no.
Don’t Feed the Thought Leaders - Earthly Blog
The more universal a solution someone claims to have to whatever software engineering problem exists, and the more confident they are that it is a fully generalized solution, the more you should question them.
Introducing System One Models & Jev - TypeSafe AI Blog
Really interesting! It could unlock some new use cases by having a LLM that respond super fast!
Three +1s and a prayer - Minid.net
There was one rule that gave the whole thing away: if nobody left any comments on a pull request, you could merge it after a day. That was bonkers. Silence became a form of approval, which meant the most reliable way to ship anything was to write a pull request so boring that nobody answered it … The goal was never to produce code that other programmers enjoy reading, and it was certainly never to accumulate approvals on GitHub. The goal was to build correct systems. For a long time we used humans reviewing the work of other humans because that was the best approximation we had, and now we are starting to build tools that can turn it into something much closer to search, measurement and verification. The further we go in that direction, the less room is left for opinions about implementation, and the more weight falls on the part I cannot see any way to delegate, which is somebody sitting down and deciding what the machine is actually supposed to do. We have historically been quite bad at that part.
The AI Competence/Judgement Gap - Itamar Gilad

My Old Boss Handed Me a Playbook. I’m Giving It to You

AI Model & API Providers Analysis | Artificial Analysis
Yet another benchmark. Their pareto line graph is an interesting way of looking at the model’s overall performance + trade-offs.
GitHub - ag-libs/lathe · GitHub
An alternative java LSP from jdtls.
”Do You Still Read the Code?” • zanlib
source code is only one kind of product of the activity of programming, but because it’s more visible, it’s treated as more valuable than the understanding acquired while producing it. It’s as if we treated the steam coming out of a coal power plant cooling tower as the main output, rather than the electricity, merely because electricity is invisible.
Why I Think You Should Almost Never Use AI to Write Anything Substantive
I think so because (1) the writing process is an essential part of the thinking process, (2) AI writing is vague and wrong in hard-to-notice ways, and (3) writing with AI (and not labeling it as such) is rude and misleading. …
All the stuff I wrote about above, about subtle errors and vagueness, and all the stuff about how, when a text is AI-written, you have no idea whether the author put a lot of thought into it — all these things violate that contract. So when I read a text and notice that it is fully or partly AI-written, my trust in the text and in the author is immediately, and I think rationally, lowered.
And for all those reasons, when you promote AI-written text, or send a draft of AI-written text to someone, I think you are being rude. I think it’s sort of like sending a really sloppily written draft to someone and hiding the fact that it’s really sloppily written.