Ask HN: Does anybody still FEEL improvements between latest LLMs for coding?
A user on Hacker News asks whether people still feel meaningful improvements between the latest LLM generations for coding, suggesting that recent models seem roughly equal in usefulness. The post invites anecdotes from users who have experienced the opposite.
Background
- This is a discussion on Hacker News (HN), a forum popular with software developers and tech professionals. "Ask HN" is a standard post format where a user poses a question to the community.
- The post asks whether people still notice a real, practical difference (an "improvement you can feel") between newer and older large language models (LLMs) — specifically for the task of writing or debugging code.
- The underlying context: Over the past ~2 years, models like GPT-3.5, GPT-4, GPT-4o, Claude 3.5 Sonnet, Claude 3 Opus, and various open models (Llama, Qwen, DeepSeek) have rapidly succeeded each other. Early on, each generation brought dramatic leaps in coding ability. The questioner is suggesting that for many developers, the improvements have plateaued or become marginal — newer models feel about as useful as the previous ones for everyday programming tasks.
- Why it matters: If coding performance has plateaued, it affects whether developers pay for premium subscriptions, which model they use, and the broader narrative about AI progress. A "feeling" of stagnation among practitioners is itself noteworthy data about the state of the art.