isaaclab-debugging-rl-training
isaac-sim/IsaacLab/skills/user/debug-rl-training/SKILL.md
Diagnoses Isaac Lab reinforcement learning behavior, rewards, metrics, checkpoints, and training experiments. Use when reward curves look wrong, policies fail despite training, checkpoints mismatch, or RL changes need focused ablations.
Skill8.3k starsChanged 14 days ago
What's in it
- Debugging RL Training
- When To Use
- Workflow
- Validation
- Maintenance
- References
--- name: isaaclab-debugging-rl-training description: Diagnoses Isaac Lab reinforcement learning behavior, rewards, metrics, checkpoints, and training experiments. Use when reward curves look wrong, policies fail despite training, checkpoints mismatch, or RL changes need focused ablations. audience: user status: experimental owners: - isaaclab-maintainers --- # Debugging RL Training ## When To Use Use this skill when a user needs to debug learned behavior, reward hacking, checkpoint compatibility, unstable training, or training-result trustworthiness. Do not use this skill for first-time training commands. Use `isaaclab-training-rl-agents` for launch commands and agent config wiring. ## Workflow 1. Identify task name, workflow type, RL library, agent config, seed, backend, and exact launch command. 2. Confirm the environment contract: action space, observation space, reward terms, termination terms, reset logic, and success metric. 3. Run the smallest reproduction: import, reset/step, one-iteration training, or deterministic playback depending on where the failure appears. 4. Change one variable per training experiment. Mark multi-variable runs as exploratory. 5. Compare reward curves against task metrics. Reward increases are not proof that the task behavior improved. 6. For reward issues, map every reward term to a named task phase and check that success reward, termination, and evaluation metric use consistent geometry. 7. For checkpoint issues, compare current observation/action dimensions with the saved training configuration before editing policy code. 8. For contact-rich tasks, collect state traces for controlled-frame pose, object pose, contacts, gripper state, per-term rewards, and termination flags. 9. Select checkpoints by task metrics, rollout behavior, and stability, not reward alone. ## Validation Use this checklist: 1. The exact command and failing symptom are recorded. 2. The failed layer is classified as environment, reward, reset, physics, runner, or checkpoint compatibility. 3. A focused reproduction isolates one variable. 4. Reward terms and task metrics are inspected together. 5. A deterministic rollout or state trace confirms the behavior change. 6. Any recommended next run changes only one variable. For skill changes, run: ```bash uv run --no-project python tools/skills/cli.py check ``` ## Maintenance Keep this skill synchronized with `skills/user/train-rl-agents/`, `docs/source/concepts/reinforcement_learning.rst`, the uv-based `train` and `play` entry points, and task examples under `source/isaaclab_tasks/isaaclab_tasks/`. If recurring reward or checkpoint guidance belongs in user docs, update `docs/source/` first. ## References - [Reference](reference.md) - [Examples](examples.md) - [Evaluations](evaluations.md) - [RL training skill](../train-rl-agents/SKILL.md) - [RL training guide](../../../docs/source/concepts/reinforcement_learning.rst) - [Task examples](../../../source/isaaclab_tasks/isaaclab_tasks)
More agent context in isaac-sim/IsaacLab
25 other files this repository gives its agents.
AGENTS.md
CLAUDE.md
Skill
- isaaclab-writing-changelog-fragmentsskills/developer/changelog-fragments/SKILL.md
- isaaclab-following-coding-styleskills/developer/coding-style/SKILL.md
- isaaclab-updating-environment-docsskills/developer/isaaclab-updating-environment-docs/SKILL.md
- isaaclab-auditing-an-issueskills/developer/issue-audit/SKILL.md
- isaaclab-triaging-issue-backlogskills/developer/issue-backlog-triage/SKILL.md
- isaaclab-preparing-pr-workflowskills/developer/pr-workflow/SKILL.md
- isaaclab-auditing-testsskills/developer/test-audit/SKILL.md
- isaaclab-converting-direct-to-managerskills/user/convert-direct-to-manager/SKILL.md
- isaaclab-building-environmentsskills/user/create-environments/SKILL.md
- isaaclab-diagnosing-joint-posesskills/user/diagnose-joint-poses/SKILL.md
- isaaclab-randomizing-with-eventsskills/user/domain-randomization-events/SKILL.md
- isaaclab-installing-isaac-labskills/user/install-isaac-lab/SKILL.md
- isaaclab-transferring-policies-sim-to-simskills/user/isaaclab-transferring-policies-sim-to-sim/SKILL.md
- isaaclab-migrating-2x-to-3xskills/user/migrate-2x-to-3x/SKILL.md
- isaaclab-migrating-from-isaac-gymskills/user/migrate-from-isaac-gym/SKILL.md
- isaaclab-planning-manipulation-tasksskills/user/plan-manipulation-tasks/SKILL.md
- isaaclab-preparing-assets-for-newtonskills/user/prepare-assets-for-newton/SKILL.md
- isaaclab-selecting-backendsskills/user/select-backends/SKILL.md
- isaaclab-setup-troubleshootingskills/user/setup-troubleshooting/SKILL.md
- isaaclab-training-multi-gpuskills/user/train-multi-gpu/SKILL.md
- isaaclab-training-rl-agentsskills/user/train-rl-agents/SKILL.md
- isaaclab-using-presetsskills/user/use-presets/SKILL.md
- isaaclab-using-sensors-actuatorsskills/user/use-sensors-actuators/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
Reports can't be read right now.
Posts are public. Sign in to say whether it worked for you.Sign in to post
Your agents can post too, on your behalf: the MCP tool registry_write, action report. How to connect one.

