Skip to content
TopicTracker
From HackerNewsView original
TranslationTranslation

Ask HN: Is anyone running local LLMs in their organization?

A Hacker News user asks the community about experiences running local LLMs within an organization, including hardware, model choice, resource allocation, and access management, seeking advice on key considerations and interesting use cases.

Background

Hacker News (HN) is the social news site run by startup incubator Y Combinator, deeply influential in tech and startup culture. This is an "Ask HN" post — a community Q&A thread where users pose questions directly to the audience. - The post asks about running large language models (LLMs) *locally* inside an organization rather than using cloud APIs like OpenAI's ChatGPT or Anthropic's Claude. This matters because of growing concerns around data privacy, API costs, vendor lock-in, and regulatory compliance. - Running local LLMs requires significant hardware (usually high-end GPUs like NVIDIA A100/H100 or consumer RTX cards), and the models themselves (e.g., LLaMA, Mistral, DeepSeek) have become dramatically more capable in smaller sizes over the past two years, making on-premise deployment more feasible. - The question reflects a broader industry shift: as LLMs become cheaper and smaller, many companies are experimenting with self-hosting to keep sensitive data off third-party servers. The comments thread typically surfaces real-world setups, pain points around infrastructure and access control, and comparisons of model quality versus cloud alternatives.

Related stories