Complex tensor operations should include high-level semantic comments explaining what blocks of code do, plus inline shape comments for chained operations using symbols consistent with docstrings.
All utility functions specific to a model must be in the same module file as the model itself, not in separate utility files, to maintain self-contained modules.
All forward and public methods must validate tensor shapes at the beginning, wrapped in torch.compiler.is_compiling() guard, with standardized error messages.
All tensor arguments in model methods must have jaxtyping type annotations with shape specifications (e.g. Float[torch.Tensor, "b c h w"]) for runtime-checkable shape information.
Cannot add new required parameters to production model __init__ or public methods without default values, as this breaks backward compatibility with existing code and checkpoints.
Cannot remove or rename parameters in production models without implementing _backward_compat_arg_mapper and incrementing __model_checkpoint_version__ to maintain compatibility.
Every model must have CI tests verifying constructor instantiation and all public attributes (excluding buffers/parameters) using pytest parameterization.
Every model must have non-regression tests comparing outputs against reference data saved in .pth files, using realistic tensor shapes and pytest parameterization.
Every model must have tests that load from checkpoint files (.mdlus), verify attributes, and compare outputs against reference data to ensure serialization works correctly.
Avoid string-based class selection with many options (>3 choices) in model constructors; prefer dependency injection with instances for better type safety and clearer APIs.
Use check_min_version() to check optional dependencies without importing, and @require_version decorator to protect version-specific features; pyproject.toml is the single source of truth for dependencies.
Core development principles and workflow for the Nanotron project. Apply when discussing project architecture, planning new features, or reviewing code structure.
A PipelineBlock wraps modules for efficient pipeline-parallel execution. Each block operates on a specific device and communicates with other pipeline stages through P2P. When passing data between pipeline stages: Design modules with clean interfaces for pipeline parallelism: Use TensorPointer to reference tensors across pipeline stages without copying: In pipeline parallelism, different tensors are only available on specific pipeline ranks (see src/nanotron/data/dataloader.py):
Overview of Nanotron project structure and key components, useful for onboarding and orientation. Apply when discussing project architecture or organization.