Skip to content
TopicTracker
From HackerNewsView original
TranslationTranslation

Ask HN: What are you go to LLM models for the following

A Hacker News user asks the community for recommended LLM models across three categories: coding, text-to-speech (TTS) and speech-to-text (STT), and image creation/understanding. The author shares their own setup using Qwen3-Coder-Next for coding, Fish Audio S2 Pro and Whisper for audio, and Gemma and Flux for image tasks, all running locally on an M5 Max MacBook Pro with 128GB of memory.

Background

- This is a discussion thread on Hacker News (HN), a tech-focused social news site where users ask and answer practical questions. The "Ask HN" format means a community member is polling other readers for advice. - The user lists models they run locally on a high-end MacBook Pro (Apple's latest M5 Max chip, 128GB RAM). This matters because most people use LLMs via cloud APIs (like ChatGPT's website), but running them locally requires expensive hardware and technical know-how. - Qwen3-Coder-Next: A specialized code-generation model from Alibaba's Qwen family, designed for programming tasks. Alternative: Claude's Sonnet, GPT-4, or Cursor. - Fish Audio S2 Pro: A text-to-speech model. Whisper: OpenAI's speech-to-text model, widely used for transcription. - Gemma: Google's family of lightweight open-weight models (used here for image analysis). Flux: A high-quality open image generation model from Black Forest Labs (rivals Midjourney/DALL-E). - The thread is a practical reference for HN readers who self-host AI models (not via cloud services), comparing toolchains across different tasks.

Related stories