「一晩で解決」と謳われたClaudeの不安定なテスト対策、実用化までに2週間を要した理由
thoughtbotのエンジニアが、不安定なテスト(flaky tests)を自動修正するClaude AIソリューションを評価・実用化した経験を報告。当初「一晩で解決」と謳われたこの手法は、実際にはラベル抽出やCIログのパースなど様々な改良を経て2週間かけて実用レベルに達した。テスト結果の差分検出やログ解析の精度向上など、実運用に向けた地道な調整の重要性が語られている。
thoughtbotのエンジニアが、不安定なテスト(flaky tests)を自動修正するClaude AIソリューションを評価・実用化した経験を報告。当初「一晩で解決」と謳われたこの手法は、実際にはラベル抽出やCIログのパースなど様々な改良を経て2週間かけて実用レベルに達した。テスト結果の差分検出やログ解析の精度向上など、実運用に向けた地道な調整の重要性が語られている。
Newer Claude models sometimes invent extra keys in tool call arguments, breaking validation in Pi's edit tool. The author suspects post-training for Claude Code's forgiving harness makes alternative schemas fail. This suggests closed RL training can degrade general tool-use reliability.
The article argues that Claude for Mac is an Electron app because its lead developer, Felix Rieseberg, is a former Microsoft engineer known for porting Slack to Electron and has a professional history strongly tied to the Electron framework.
Simon Willison released llm-coding-agent 0.1a0, an experimental coding agent built on his LLM library. The agent provides tools for reading, editing, searching files and executing commands, and ships with a Python API and CLI. It was developed using Claude Code (Fable 5) via TDD with a spec-first approach, and is available as a slop-alpha on PyPI.
Truth Social remains an outlet primarily for Trump's own posts, while other administration officials continue using X. The platform functions as a blog-like channel for Trump's messages rather than a genuine social network competitor.