Skip to content
TopicTracker
From HackerNewsView original
TranslationTranslation

GitHub Handles Git LFS

The article explains how GitHub manages Git Large File Storage (LFS), detailing the architecture and processes used to efficiently store and deliver large files in repositories. It covers the separation of LFS metadata from binary data, the use of content-addressable storage, and the integration with GitHub's existing infrastructure to ensure performance and reliability.

Background

- Git LFS (Large File Storage) is an open-source extension that replaces large files (binaries, images, archives) in a Git repository with text pointers, storing the actual file content on a remote server. This keeps repository clones fast and history manageable. - GitHub has offered Git LFS since 2015; every LFS-enabled repo uses a certain amount of bandwidth and storage that counts toward a paid plan (the free tier is quite limited). - The article's author, Scott Berrevoets, worked on the Git LFS team at GitHub. He is describing a significant internal redesign of how GitHub stores and serves LFS data. - The key prior context: GitHub historically stored LFS objects on S3 (Amazon's cloud storage) behind a custom Ruby-on-Rails proxy called the "LFS API." As GitHub grew, this setup hit performance ceilings (latency, throughput) and operational complexity issues. - This matters because millions of developers rely on Git LFS daily for game development, machine learning datasets, design assets, and other large files. How GitHub handles LFS scaling directly affects their workflows.

Related stories