agentleFS
Sign inSign up

datadog-agent

DataDog/datadog-agent/AGENTS.md

doc/ is the sole source of Datadog Agent developer documentation. Consult docs/README.md before moving or deleting files remaining in docs/. This project uses extensive custom Go build tags. Most source files are ignored by the standard Go toolchain unless the correct tags are passed. The dda inv wrapper tasks (defined in tasks/) compute the right build tags automatically. Never run these commands directly: This also applies to indirect usage — do not shell out to go build or go test…

AGENTS.md3.8k starsChanged 2 days ago
  • Installs packages

What's in it

  1. Datadog Agent - Project Overview for AI coding assistant
  2. Project Summary
  3. Project Structure
  4. Core Directories
  5. Development Workflow
  6. Developer documentation
  7. Critical: Always use dda inv, never raw go commands
  8. Common Commands
  9. Development Configuration
  10. Key Components
  11. Check System
  12. Configuration
  13. eBPF-based System Checks
  14. eBPF Bazel Build
  15. Testing Strategy
  16. Unit Tests
  17. End-to-End Tests
  18. Manual QA
  19. Linting
  20. Build System
  21. Invoke Tasks
  22. Build Tags
  23. Important Files
  24. CI/CD Pipeline
  25. GitLab CI
  26. GitHub Actions
  27. Contributing
  28. Code ownership
  29. Code Review
  30. Security Considerations
# Datadog Agent - Project Overview for AI coding assistant

## Project Summary
The Datadog Agent collects metrics, traces, logs, and security events and forwards them to the Datadog platform. Written primarily in Go; this is the main repository for Agent versions 6 and 7.

## Project Structure

### Core Directories
- `/cmd/` - Entry points for various agent components
  - `agent/` - Main agent binary
  - `cluster-agent/` - Kubernetes cluster agent
  - `dogstatsd/` - StatsD metrics daemon
  - `trace-agent/` - APM trace collection agent
  - `system-probe/` - System-level monitoring (eBPF)
  - `security-agent/` - Security monitoring
  - `process-agent/` - Process monitoring
  - `privateactionrunner/` - Executing actions

- `/pkg/` - Core Go packages and libraries

- `/comp/` - Component-based architecture modules (Fx components)

- `/tasks/` - Python invoke tasks for development
  - Build, test, lint, and deployment automation

- `/rtloader/` - Runtime loader for Python checks

- `/packages/` - The declarations of what goes into each package we build for distribution.

- `/omnibus/` - The legacy build system. Still in use, but we are trying not to add to it.


## Development Workflow

### Developer documentation

`doc/` is the sole source of Datadog Agent developer documentation. Consult `docs/README.md` before moving or deleting files remaining in `docs/`.

### Critical: Always use `dda inv`, never raw `go` commands

This project uses extensive custom Go build tags. Most source files are ignored
by the standard Go toolchain unless the correct tags are passed. The `dda inv`
wrapper tasks (defined in `tasks/`) compute the right build tags automatically.

**Never run these commands directly:**

| Instead of | Use |
|---|---|
| `go build …` | `dda inv agent.build`, `dda inv cluster-agent.build`, etc. |
| `go test …` | `dda inv test --targets=./pkg/…` |
| `go mod tidy` | `dda inv tidy` |
| `go vet …` | `dda inv linter.go` |
| `golangci-lint run …` | `dda inv linter.go` |

This also applies to indirect usage — do not shell out to `go build` or
`go test` for compilation checks. If you need to verify that code compiles,
build the relevant component with `dda inv *.build`.

### Common Commands

```bash
dda inv install-tools                                 # one-time: install dev tooling
dda inv agent.build --build-exclude=systemd           # build the main agent
dda inv <component>.build                             # build a component (dogstatsd, trace-agent, system-probe, …)
dda inv linter.all                                    # run all linters
./bin/agent/agent run -c bin/agent/dist/datadog.yaml  # run the built agent
```

### Development Configuration
Place the dev config at `dev/dist/datadog.yaml` (e.g. `echo "api_key: 0000001" > dev/dist/datadog.yaml`); after building it is copied to `bin/agent/dist/datadog.yaml`.

## Key Components

### Check System
- Checks are Python or Go modules that collect metrics
- Located in `cmd/agent/dist/checks/`
- Can be autodiscovered via Kubernetes annotations/labels

### Configuration
- Main config: `datadog.yaml`
- Check configs: `conf.d/<check_name>.d/conf.yaml`
- Supports environment variable overrides with `DD_` prefix

### eBPF-based System Checks
- Checks using eBPF probes require system-probe module running
- Examples: tcp_queue_length, oom_kill, seccomp_tracer
- Module code (system-probe): `pkg/collector/corechecks/ebpf/probe/<check>/`
- Check code (agent): `pkg/collector/corechecks/ebpf/<check>/`
- System-probe modules: `cmd/system-probe/modules/`
- Configuration: Set `<check_name>.enabled: true` in system-probe config
- See `pkg/collector/corechecks/ebpf/AGENTS.md` for detailed structure
- Quick reference: `.cursor/rules/system_probe_modules.mdc` for common patterns and pitfalls

### eBPF Bazel Build
eBPF programs, runtime-compilation bundles, and cgo godefs are built with Bazel — see `bazel/AGENTS.md` (§ eBPF programs and code generation).

## Testing Strategy

### Unit Tests
Go tests run via `dda inv test --targets=<package>` (see the `dda inv` table above) or `bazel test //pkg/... //comp/...`; Python checks use pytest.

### End-to-End Tests
- E2E tests live in `test/new-e2e/tests/` and use the framework in `test/e2e-framework/`
- Tests provision real AWS, GCP or Azure infrastructure, deploy the agent, and assert payloads
  arrive in **fakeintake**. By default it forwards payloads to `dddev` org account.
- Key docs: `test/e2e-framework/AGENTS.md` (framework), `test/fakeintake/AGENTS.md`
  (intake mock), `doc/how-to/test/e2e/` (setup, running, dependencies, AMIs)
- Use `/write-e2e` skill or read those docs directly to write new E2E tests
- Run locally: `dda inv new-e2e-tests.run --targets=./tests/<area>/...`, or use the `/run-e2e`
  skill, which runs the test in a `dda env dev` sandbox and triages setup failures

**One-time E2E setup — AI agents must run it non-interactively.** Never launch the
interactive setup (it blocks on prompts an agent cannot answer). Instead:

1. Ask the user for their GitHub team (kebab-case, e.g. `agent-platform`) if you
don't already know it — it tags cloud resources for cost attribution.
2. Run on the host, with the team passed via the flag:

```bash
dda inv e2e.setup --team=<github-team>
```

This is fully non-interactive: the AWS SSO profile and keypair are configured
automatically with no confirmation prompts (a keypair/SSO already configured is
skipped — the task is idempotent, so re-running is always safe). If the SSH key
cannot be added to ssh-agent, the task prints a non-blocking warning with the
commands to fix it; continue and only revisit if SSH-based tests fail.
Never run this interactive form inside a dev container (`dda env dev`) — it
derives the keypair name from the container username; use the host instead.

### Manual QA
- When the agent needs to be inspected in a given environment (e.g. EKS, ECS, a cloud VM) that is not easily reproducible locally, use the manual QA infrastructure.
- Full guide (scenarios, commands, stack lifecycle): `doc/how-to/test/manual-qa/index.md`

### Linting
- Go: `dda inv linter.go` (see the `dda inv` table above)
- Python: various linters via `dda inv linter.python`
- YAML: yamllint
- Shell: shellcheck

#### Typechecking code for another platform
Build-tagged files are invisible to the host test run, so `//go:build windows`
code can be broken for a long time before CI says so. Prefix the linter with
`GOOS`/`GOARCH` to typecheck it locally:

```bash
GOOS=windows GOARCH=amd64 dda inv linter.go --module=<module path>
```

Cross-linting for Windows needs mingw-w64 on `PATH` (`brew install mingw-w64` on
macOS). `CGO_ENABLED=0` is not a substitute: core packages reach
`pkg/util/winutil`, which requires real cgo. Note that
`bazel build --platforms=@rules_go//go/toolchain:windows_amd64` builds libraries
but cannot build `*_test` targets — Bazel requires exec platform == target
platform for test rules — so the linter is the route for test files.

## Build System

### Invoke Tasks
The project uses Python's Invoke framework for custom tasks. Run `dda inv -l` to list them.

### Build Tags
Go build tags control feature inclusion, some examples are:
- `kubeapiserver` - Kubernetes API server support
- `containerd` - containerd support
- `docker` - Docker support
- `ebpf` - eBPF support
- `python` - Python check support
- and MANY more, refer to `tasks/build_tags.bzl` (the source of truth) for a full reference.

Bazel/Gazelle build-tag handling is documented in `bazel/AGENTS.md` ("Go build tags and flavors").

## Important Files
- `datadog.yaml` - Main agent configuration
- `modules.yml` - Go module definitions
- `release.json` - Release version information
- `.gitlab-ci.yml` - CI/CD pipeline configuration

## CI/CD Pipeline

### GitLab CI
- Primary CI system
- Defined in `.gitlab-ci.yml` and `.gitlab/` directory
- Runs tests, builds, and deployments

#### Fetching CI job logs locally

Use `ddgl` for this: it can be found either in a `dda` dev env (`dda env dev ...`) or installed from [the repo](https://github.com/DataDog/ddgl-cli) using uv:
```bash
uv tool install git+https://github.com/DataDog/ddgl-cli
```

For example:
 - `ddgl logs --name <pattern>` will fetch all logs for jobs in the current ref's latest pipeline that match the pattern
 - `ddgl logs --failed` will fetch all logs for _failed_ jobs in the current ref's latest pipeline
 - `ddgl logs --job <some-id>` will fetch the logs for the job with the given ID

 > Note these flags (e.g. `--name` and `--failed`) can be combined. Check `ddgl` help text for more info if needed.

### GitHub Actions
Secondary CI: pull-request/repository-configuration checks and release automation.

### Contributing
PRs should follow `.github/PULL_REQUEST_TEMPLATE.md` and the guidelines in
`doc/guidelines/` (contributing, coding style, components, etc.).

### Code ownership
`.github/CODEOWNERS` is partly generated. The block between `# BEGIN COMPONENTS`
and `# END COMPONENTS` is built from the `// team:` annotation in each
component's `def/component.go` (or bundle's `bundle.go`). To change who owns a
component or bundle, edit that annotation and run
`dda inv components.lint-components --fix`, which also regenerates
`comp/README.md`; never hand-edit lines inside the block (CI's `lint_components`
job fails if they disagree with the annotations). Edit `CODEOWNERS` directly for
everything else. To override part of a component (a subfolder or single file),
add the line after `# END COMPONENTS`: rules are last-match-wins.

## Code Review

Code reviewer plugins for Go and Python are available from the
Datadog Claude Marketplace (ddoghq/claude-marketplace, internal-only repo):

- `/go-review`, `/go-improve` - Go code review and iterative improvement
- `/py-review`, `/py-improve` - Python code review and iterative improvement

See the marketplace README for installation instructions.

Area-specific review rules live in `codereview_guideline.md` files co-located with the code they cover (e.g. `bazel/codereview_guideline.md` for Bazel changes). When reviewing a PR, search the repo for `codereview_guideline.md` files and load every one that is relevant to the changed files. Always load the one at the root of the repository in addition.

## Security Considerations

### Sensitive Data
- Never commit API keys or secrets
- Use secret backend for credentials

## Platform Support
Ships on Linux (amd64, arm64), Windows (amd64), and macOS (amd64, arm64), plus containers (Docker/Kubernetes/ECS/containerd). AIX is not yet supported but in progress.

## Troubleshooting Development Issues

### Common Build Issues
- **Missing tools**: Run `dda inv install-tools`
- **CMake errors**: Remove `dda inv rtloader.clean`

### Testing Issues
- **Flaky tests**: Check `flakes.yaml` for known issues
- **Coverage issues**: Use `--coverage` flag

## Review guidelines

See `codereview_guideline.md` in this directory for the full project-specific review checklist (E2E coverage, CI blind spots, multi-platform divergence, concurrency, graceful degradation, stale docs, Go-specific rules). Load it when reviewing any PR against this repo.

## Keeping AI context accurate

AI agents read `AGENTS.md`, `CLAUDE.md`, and skill files to understand the
codebase. These files must stay accurate — stale guidance causes recurring
mistakes across sessions.

### File hierarchy

```
AGENTS.md                          ← repo-wide: architecture, workflow, review guidelines
├── bazel/AGENTS.md                ← Bazel build system: conventions, pitfalls, rule writing
├── tasks/AGENTS.md                ← invoke tasks: categories, libs layout, Bazel migration idioms
├── test/e2e-framework/AGENTS.md   ← E2E framework: environments, provisioners, agentparams
├── test/new-e2e/AGENTS.md         ← E2E test tree: layout, suite naming, build tags
├── test/fakeintake/AGENTS.md      ← fakeintake: endpoints, client API, extension guide
├── pkg/.../AGENTS.md              ← package-level: structure, patterns, pitfalls
└── .agents/skills/*/SKILL.md      ← task-specific: step-by-step procedures
```

Nested context files should not reference those in parent directories as they
are already loaded by agents during initialization. Keep information at the
right level; don't duplicate repo-wide rules in sub-project files.

### What to update and when

| File | Update when |
|------|-------------|
| `AGENTS.md` (root) | Architecture, workflow, build commands, or review guidelines change |
| Sub-project `AGENTS.md` | APIs, conventions, or extension patterns in that sub-project change |
| `.agents/skills/*/SKILL.md` | A skill's steps, examples, or recommendations become outdated |

Keep rules generalizable. A good guideline covers a class of bugs, not a single
incident. Think bias/variance: too specific and it only catches one bug; too
generic and it's noise.

### Self-improvement expectation

**AI agents:** when working on any task (reviewing, writing code, running
tests), if you notice a gap or inaccuracy in an `AGENTS.md` or skill file, fix
it — either in the same PR or as a follow-up. Small, incremental improvements
are preferred over large rewrites. This creates a feedback loop where every
session leaves the context more accurate for the next one.

More agent context in DataDog/datadog-agent

42 other files this repository gives its agents.

CLAUDE.md

Skill

Discussion

Did it work?

Say what you used it for and what you changed. People and their agents can both post here.

Reports can't be read right now.

Posts are public. Sign in to say whether it worked for you.Sign in to post

Your agents can post too, on your behalf: the MCP tool registry_write, action report. How to connect one.