• Coupling models and harnesses

    This is very interesting. The more time I spend in Claude Code, the more I thought that coding agent harnesses were the perfect sort of software to be open sourced. The blossoming popularity of Pi seemed to reinforce this. But then…

    Armin reports on a weird problem he ran into while hacking on Pi: 

    The short version is that newer Claude models sometimes call Pi’s edit tool with extra, invented fields in the nested edits[] array. And not Haiku or some small model: Opus 4.8. The edit itself is usually correct but the arguments do not match the schema as the model invents made-up keys and Pi thus rejects the tool call and asks to try again.

    That alone is not too surprising as models emit malformed tool calls sometimes. Particularly small ones. What surprised me is that this is getting worse with newer Anthropic models as both Opus 4.8 and Sonnet 5 show it but none of the older models. In other words, the SOTA models of the family are worse at this specific tool schema than their older siblings.

    Armin theorizes that this is because more recent Anthropic models have been specifically trained (presumably via Reinforcement Learning) to better use the edit tools that are baked into Claude Code. This has the unfortunate effect that other coding harnesses, such as Pi, may find that their own custom edit tools are more likely to be used incorrectly.

    – Simon Willison, Better Models, Worse Tools

  • Understanding the amount of bullshit work

    AI is going to help humanity by revealing just much human time is spent on bullshit work.

    Consulting giant Accenture is trying to figure out how to stop non-technical workers from blowing through companies’ AI token budget on trivial tasks like converting PDFs to presentation slides, according to leaked audio obtained by 404 Media. Across the industry Accenture is seeing “soaring token spend,” according to the audio.

    The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI

  • Local journalism

    South Shore news: AI-generated newsletter has paying audience (Boston Globe)

    This feels like the future. AI agents, directed by assignment editors, gather local intelligence and context. Other agents draft and edit stories. Humans handle quality control and accountability. Meanwhile, separate software tracks traffic and prominence, feeding signals back to adjust agent behavior — optimizing for the business goals or cultural impact that publication leadership is after.

  • The new skill: task imagination

    With models that can run for days, NLW argues we’ll all have to up-level ambition — and become token-efficiency optimizers who match models to use cases. Nate B. Jones’s framing: most of us have nothing that’s ever taken even an hour on AI, so the scarce skill is imagining tasks worth handing to a model that works for days.

    — The AI Daily Brief · https://aidailybrief.ai/e/2026-06-10#task-imagination

  • Jevon’s Paradox and AI

    Jevons paradox is coming to knowledge work. By making it far cheaper to take on any type of task that we can possibly imagine, we’re ultimately going to be doing far more. The vast majority of AI tokens in the future will be used on things we don’t even do today as workers: they will be used on the software projects that wouldn’t have been started, the contracts that wouldn’t have been reviewed, the medical research that wouldn’t have been discovered, and the marketing campaign that wouldn’t have been launched otherwise.

    — Aaron Levie, Jevon’s Paradox for Knowledge Work

    Via Simon Willison’s Webblog.

  • Be YouTube, not Qwest

    Today, no amount of model training is too much. No price for that training is too high. The builders bet is that inference will be orders of magnitude cheaper in 5 years. Build for that. Be Netflix or YouTube, not Qwest.

    This year, American tech companies will spend $300 billion to $400 billion on artificial intelligence, which is in nominal dollars more than any group of companies have ever spent to do anything. Notably, these companies are not remotely close to earning $400 billion on artificial intelligence.

    That’s why you’re starting to hear some people wonder whether the AI build-out is turning into the mother of all economic bubbles.

    The prospect of an AI bubble should scare us. Roughly half of last quarter’s GDP growth came from infrastructure spending on AI, and more than half of stock market appreciation in the last few years has come from companies associated with AI. If the AI spending project blows up in the next few years, as our next guest says it might, the implications for technology, the economy, and politics would be immense.

    This Is How the AI Bubble Could Burst – Plain English with Derek Thompson

  • On learning, open-mindedness and objectivity

    One regret: not being better at learning from people who you may not like – at first, or at all. It is so tempting to close one’s mind off in response to a visceral dislike of a person. Especially if they appear successful. Jealousy, and indulging one’s sense of self-righteousness, are enemies of learning and wisdom.

  • Friction, difficulty and independence

    I think about this a lot with my own kids: making sure they have the right amount of friction in there lives, having time and space to be bored, and to even get into low-grade trouble.

    School hasn’t been a problem so far (acknowledging my privilege); they get challenge from other places. This piece is a good reminder of how difficulty – and potential boredom – are important parts of growing up.