Reverse-engineered token Value of coding agent Plans
The article reverse-engineers token pricing of AI coding agent plans (Cursor, Windsurf, Copilot, Codeium). Pro plans are often 1.5–3× cheaper per token than buying the model directly, while Ultra plans can be 10× more expensive. It offers a framework for developers to compare plans objectively.
Background
- AI coding agents (Cursor, Copilot, Devin) break coding tasks into step-by-step "plans." Each step consumes LLM tokens — the basic unit AI models use to process text, which cost real money on APIs like GPT-4.
- The author reverse-engineered popular agents to measure each step's "token value": how many tokens burned vs. how much actual coding work got done. Key finding: many agents waste tokens on redundant file reads, verbose internal chatter, and low-value self-checks.
- This matters because tokens = cost and speed. An agent that burns 50,000 tokens on planning before writing one line of code feels slow and expensive. The article offers a framework for evaluating whether an agent is efficient or just burning compute.
- For developers building or choosing coding agents, this reveals the hidden economics behind AI tooling — beyond the demo videos and hype.
Max Weinbach says he had early access to OpenAI's new model GPT-5.6 Sol, calling it his favorite model by far. He highlights that it never gives up and will keep reasoning until it's done. OpenAI announced that GPT-5.6 Sol, along with Terra and Luna, will launch publicly on Thursday, with preview access expanding globally now.
The US government ordered Anthropic to suspend access to its Fable 5 and Mythos 5 models for all customers, citing a potential jailbreak technique that involved asking the model to review a codebase for vulnerabilities—a capability Anthropic says is available in other public models. Access was abruptly cut off on June 12.
Andrej Karpathy announces the release of Claude Fable 5, the same underlying model as Mythos but with added safeguards. He calls it a major step forward, particularly for long problem-solving sessions on difficult tasks, and describes it as state-of-the-art on nearly all benchmarks with exceptional performance in software engineering, research, and vision.
Roman Storm warns that the legal theory in his case could set a precedent making open-source developers liable for how others use their code, potentially criminalizing the mere publication of privacy, messaging, or crypto tools. He notes that developer Michael Lewellen cannot publish lawful code due to prosecution fears, and argues this chilling effect extends beyond any single case.
Meta's engineering culture is deteriorating under Mark Zuckerberg and Scale AI CEO Alexandr Wang, who have introduced keyboard tracking, reassignments to data labeling, and AI-centric performance metrics. Critics argue this incentivizes performative AI use, drives away experienced engineers, and contributed to a major Instagram hijacking incident caused by AI-written and AI-reviewed code.