upskill
huggingface/upskill/AGENTS.md
Before finishing a change, run: CI enforces the same flow in .github/workflows/ci.yml. Formatting and linting are enforced with Ruff.
How real projects brief Codex, Cursor and every other agent that reads AGENTS.md.
huggingface/upskill/AGENTS.md
Before finishing a change, run: CI enforces the same flow in .github/workflows/ci.yml. Formatting and linting are enforced with Ruff.
huggingface/Repo2RLEnv/docs/reference/AGENTS.md
A reference for what's possible when consumers run Repo2RLEnv-emitted Harbor tasks: which agent harnesses can run them, what inputs each agent accepts (LLM endpoint, model, etc.), and the data path that makes RL training possible — how token IDs and logprobs flow out of the sandbox into a trainer. This is Harbor's responsibility, not Repo2RLEnv's — but understanding it matters because it determines what kinds of consumers our datasets work with out of the box. From src/harbor/models/agent/name.py — the canonical…
huggingface/hub-docs/docs/hub/agents.md
Hugging Face Agents connect AI agents to the Hub. Using MCP (Model Context Protocol), Skills, or open-source tooling, agents can search models, explore datasets, run Spaces, and use community tools. You can connect agents via the HF MCP Server, install pre-built Skills for coding agents, or build agents programmatically with the huggingface_hub SDK. Agents work with any MCP-compatible client, including ChatGPT, Claude Desktop, Cursor, VS Code, and more.
huggingface/funes/AGENTS.md
Read this before changing the code or reporting an issue. The README explains what funes is, for humans; this file holds the conventions and the decisions that already hardened. A verb is described in four places. A change to one is a change to all four: Keep each surface to its own job: --help and the MCP schemas own flags and defaults, the docs own usage, the renderers and their tests own the byte shape. A fact repeated across two…
huggingface/hf-hub/AGENTS.md
Async Rust client library for the Hugging Face Hub API. This is the Rust equivalent of the Python huggingface_hub library. The primary entry point is the HFClient struct, which wraps an Arc<HFClientInner> for cheap cloning. All methods are async and use reqwest as the HTTP client. Paginated endpoints return impl Stream<Item = Result<T>> via futures::stream::try_unfold. Methods take parameters via bon per-method builders finished with .send().await (mirroring reqwest / aws-sdk / octocrab style). Key capabilities: These rules apply to ALL code…
huggingface/hf-mcp-server/AGENTS.md
No summary in the file. Open it to read it.
huggingface/optimum-neuron/AGENTS.md
Optimum Neuron bridges Hugging Face libraries (Transformers, Diffusers, PEFT) with AWS Trainium/Inferentia accelerators. Use this file for project-wide guidance and the model-specific guides below: - Inference models guide: optimum/neuron/models/inference/AGENTS.md - vLLM guide: optimum/neuron/vllm/AGENTS.md When tasked with work in a specific subdirectory, read the relevant AGENTS.md before starting — these are automatically loaded when Claude Code is invoked inside those directories, but must be read manually when working from the project root: When adding a new model, create a CLAUDE.md containing…
huggingface/optimum-neuron/optimum/neuron/cache/AGENTS.md
CLI commands live in optimum/commands/neuron/cache.py (not in this directory). Failed entries permanently block recompilation of the same graph. This is the primary reason cleanuplocalcache() exists. The Hub cache integration works by replacing libneuronxla.createcompilecache at runtime: CompileCacheHfProxy wraps a local CompileCacheFs (or CompileCacheS3) and adds Hub lookup/download as a fallback. It delegates all writes to the local cache — Hub upload only happens via explicit synchronize(). The patching utility is in optimum/neuron/utils/patching.py. synchronizehubcache() uploads the local cache to a HF Hub…
huggingface/optimum-neuron/optimum/neuron/models/inference/AGENTS.md
This guide focuses on NxD inference models, decoder graphs, attention, and porting practices. For project-wide workflows see AGENTS.md. Architecture documents with class hierarchies, graph variants, forward dispatch, and generation loop diagrams: - Text generation (CausalLM): ARCHITECTURECAUSALLM.md - Image-text-to-text (VLM): ARCHITECTUREIMAGETEXTTOTEXT.md Implementation details live in: - optimum/neuron/models/inference/backend/modules/decoder/modelingdecoder.py - optimum/neuron/models/inference/backend/modules/decoder/decoderbuilders.py - optimum/neuron/models/inference/backend/modules/decoder/decoder_wrappers.py See reference implementation: - optimum/neuron/models/inference/llama/modeling_llama.py Use NxDI for neuron-specific graph changes and HF Transformers for base architecture. The Optimum Neuron implementation prioritizes stability, maintainability, and HF ecosystem compatibility over cutting-edge…
huggingface/optimum-neuron/optimum/neuron/models/inference/backend/modules/attention/AGENTS.md
This directory contains the attention implementation for NxD inference models. For broader context see optimum/neuron/models/inference/AGENTS.md. - attentionbase.py — NeuronAttentionBase: base class for all decoder attention layers. Handles TP-aware QKV/O projections, RoPE, flash attention dispatch, and KV cache management. - flashattentionnki.py — Custom NKI kernel (flashfwdlarged) for models with headdim > 128 (e.g. Gemma3-27B with headdim=256). - gqa.py — Grouped Query Attention sharding strategies (GroupQueryAttentionQKV, GroupQueryAttentionO). - utils.py — Shared attention helpers (RoPE application, repeat_kv, manual softmax for token generation). NeuronAttentionBase.getflashattentionstrategy()…
huggingface/optimum-neuron/optimum/neuron/models/inference/gemma3/AGENTS.md
This directory contains the Neuron-optimized Gemma3 inference implementation. It follows HF Transformers architecture with Neuron-specific changes derived from NxDI. Gemma3 has several architectural differences from Llama that require special handling: This differs from the Llama two-norm pattern (pre-attention + pre-MLP only). `Gemma3NxDModelForCausalLM.converthftoneurons
huggingface/optimum-neuron/optimum/neuron/models/inference/granite/AGENTS.md
This directory contains the Neuron-optimized Granite inference implementation. It follows HF Transformers structure with only Neuron-specific graph changes from NxDI. Reference: NxDI Granite implementation Granite shares Llama's architecture and inherits similar removals: Why removed: Requires unstable NKI private APIs. Optimum Neuron uses compiler auto-optimization. Why removed: Production serving features delegated to vLLM integration. Why removed: Optimum uses HF Transformers config base extended via neuron_config.json. Optimum Neuron prioritizes stability and HF compatibility over experimental optimizations.
huggingface/optimum-neuron/optimum/neuron/models/inference/llama4/AGENTS.md
This directory contains the Neuron-optimized Llama4 inference implementation, aligned to HF Transformers with Neuron-specific changes from NxDI. Reference: NxDI Llama4 implementation Llama4 follows Llama architecture with shared removals: Why removed: Requires neuronxcc.nki private APIs not stabilized for external use. Optimum relies on compiler auto-optimization. Why removed: Advanced serving features for production deployments. vLLM integration handles production needs. Why removed: Optimum uses HF config base with neuron_config.json extensions. Optimum Neuron prioritizes stability, maintainability, and HF ecosystem compatibility.
huggingface/optimum-neuron/optimum/neuron/models/inference/llama/AGENTS.md
This directory contains the Neuron-optimized Llama inference implementation. It is based on HF Transformers and trimmed from NxDI to include only Neuron-specific graph changes. Reference: NxDI Llama modeling_llama.py The Optimum Neuron port removes NxDI-specific components that are either experimental, infrastructure-specific, or replaced by Optimum's own systems: Why removed: These kernels require neuronxcc.nki private/pre-prod APIs not stabilized for external use. Optimum Neuron relies on stable NeuronX compiler auto-optimization rather than hand-tuned NKI kernels. Users needing these optimizations should use NxDI directly.
huggingface/optimum-neuron/optimum/neuron/models/inference/mixtral/AGENTS.md
This directory contains the Neuron-optimized Mixtral (MoE) inference implementation. It follows HF Transformers architecture with Neuron-specific changes derived from NxDI. Reference: NxDI Mixtral implementation Mixtral is a Mixture-of-Experts (MoE) model with additional MoE-specific removals: Why removed: Requires unstable NKI APIs. Compiler handles expert optimization automatically. Why removed: Complex serving optimizations for production MoE deployments. Focus on foundational MoE export/inference. Why removed: Optimum uses standardized MoE configuration via neuron_config.json. Optimum Neuron provides stable foundational MoE support while delegating advanced serving to…
huggingface/optimum-neuron/optimum/neuron/models/inference/phi3/AGENTS.md
This directory contains the Neuron-optimized Phi-3 inference implementation, aligned with HF Transformers and limited to Neuron-specific graph changes. Reference: NxDI Phi-3 implementation Phi-3 uses Llama-like architecture with shared removals: Why removed: Requires unstable neuronxcc.nki private APIs. Optimum uses compiler optimization. Why removed: Production serving features handled by vLLM integration. Why removed: Optimum uses HF Transformers config with neuron_config.json extensions. Optimum Neuron prioritizes stability and HF ecosystem compatibility for Phi-3 models.
huggingface/optimum-neuron/optimum/neuron/models/inference/qwen2/AGENTS.md
This directory contains the Neuron-optimized Qwen2 inference implementation. It follows HF Transformers architecture with only Neuron-specific changes from NxDI. Reference: NxDI Qwen2 implementation Qwen2 follows Llama-like architecture with shared removals: Why removed: Requires neuronxcc.nki private APIs. Optimum relies on compiler auto-optimization. Why removed: Production serving delegated to vLLM integration. Why removed: Optimum uses HF config base with neuron_config.json. Optimum Neuron prioritizes stability and HF compatibility for Qwen2 models.
huggingface/optimum-neuron/optimum/neuron/models/inference/qwen3/AGENTS.md
This directory contains the Neuron-optimized Qwen3 inference implementation. It follows HF Transformers architecture with Neuron-specific changes derived from NxDI. Reference: NxDI Qwen3 implementation Qwen3 follows Llama architecture with shared removals: Why removed: Requires unstable neuronxcc.nki APIs. Optimum uses compiler optimization. Why removed: Production serving features handled by vLLM integration. Why removed: Optimum uses HF Transformers config with neuron_config.json. Optimum Neuron prioritizes stability and HF ecosystem compatibility.
huggingface/optimum-neuron/optimum/neuron/models/inference/qwen3_moe/AGENTS.md
This directory contains the Neuron-optimized Qwen3 MoE inference implementation. It follows HF Transformers architecture with Neuron-specific changes from NxDI. Reference: NxDI Qwen3 MoE implementation Qwen3 MoE is a Mixture-of-Experts model with MoE-specific removals: Why removed: Requires unstable NKI APIs. Compiler handles MoE optimization. Why removed: Complex MoE serving optimizations for production. Focus on foundational support. Why removed: Optimum uses standardized config via neuron_config.json. Optimum Neuron provides stable foundational MoE support with vLLM for advanced serving.
huggingface/optimum-neuron/optimum/neuron/models/inference/smollm3/AGENTS.md
This directory contains the Neuron-optimized SmolLM3 inference implementation. It follows HF Transformers architecture with Neuron-specific changes from NxDI. Reference: NxDI SmolLM3 implementation SmolLM3 follows Llama architecture with shared removals: Why removed: Requires neuronxcc.nki private APIs. Optimum relies on compiler auto-optimization. Why removed: Production serving features handled by vLLM integration. Why removed: Optimum uses HF Transformers config with neuron_config.json extensions. Optimum Neuron prioritizes stability and HF compatibility for SmolLM3 models.
An open format for instructions to coding agents, read by Codex, Cursor and others. Think of it as a README written for agents.
At the repository root, with more specific files in subdirectories. Agents read the one closest to the file they're editing.
Setup and test commands, code style, and the rules a new contributor would need to know.
Claude Code reads CLAUDE.md. A one-line CLAUDE.md that points at AGENTS.md covers both.