drug-docking-analysis
learningmatter-mit/AtomisticSkills/.agents/skills/drug-docking-analysis/SKILL.md
Post-docking analysis of virtual screening results including score distributions, enrichment metrics (ROC AUC, enrichment factors), and ligand efficiency calculations.
Skill172 starsChanged 4 months ago
What's in it
- drug-docking-analysis
- Goal
- Instructions
- 0. Prerequisites: produce a dockingranked.csv
- 1. Basic analysis (no labels)
- 2. With enrichment analysis (labeled library)
- 3. Aggregate microstates before ranking
- 4. Interpret results
- Outputs
- Constraints
--- name: drug-docking-analysis description: Post-docking analysis of virtual screening results including score distributions, enrichment metrics (ROC AUC, enrichment factors), and ligand efficiency calculations. category: [drug-discovery] --- # drug-docking-analysis ## Goal To analyze virtual screening docking results by computing score distributions, ligand efficiency metrics, and (when labeled actives/inactives are available) enrichment statistics. This skill sits between [drug-docking-vina](../drug-docking-vina/SKILL.md) and downstream refinement stages, providing quantitative assessment of docking campaign quality. **Important**: this skill does **not** assess pose quality. A compound can receive an excellent Vina score with a physically implausible pose (internal clashes, strained torsions, mis-assigned bond orders). Always pair this analysis with [drug-pose-validation](../drug-pose-validation/SKILL.md) (PoseBusters) before acting on the top-ranked compounds. Outputs: - Score distribution KDE plot - Score vs. molecular weight scatter (visualizes Vina's size/lipophilicity bias) - Ligand efficiency (LE, BEI, SEI) distributions - ROC curve with AUC (when labels available) - Enrichment factor bar chart at 1%, 2%, 5%, 10%, 20% (when labels available) - Enriched results CSV with per-compound efficiency metrics ## Instructions ### 0. Prerequisites: produce a docking_ranked.csv This skill expects a ranked CSV produced by [drug-docking-vina](../drug-docking-vina/SKILL.md)'s `collect_results.py`. If you are starting from raw `drug-docking-vina` JSON output, run the collect step first: ```bash # Env: drugdisc-agent python .agents/skills/drug-docking-vina/scripts/collect_results.py \ --results docking/results/docking_results.json \ --library_csv library/library_master.csv \ --output_dir docking/analysis/ ``` The collect step joins the docking scores with the library CSV to pull SMILES, labels, and (when present) `parent_compound_id` / `microstate_id` columns. See [drug-docking-vina SKILL.md step 5](../drug-docking-vina/SKILL.md) for details. ### 1. Basic analysis (no labels) When you have docking results but no active/inactive labels: ```bash # Env: drugdisc-agent python .agents/skills/drug-docking-analysis/scripts/analyze_docking.py \ --docking_csv docking/docking_ranked.csv \ --output_dir docking/analysis/ ``` This produces score KDE, score vs. MW, and ligand efficiency plots. ### 2. With enrichment analysis (labeled library) When your library has known actives and inactives: ```bash # Env: drugdisc-agent python .agents/skills/drug-docking-analysis/scripts/analyze_docking.py \ --docking_csv docking/docking_ranked.csv \ --library_csv library/library_master.csv \ --active_label active \ --inactive_label inactive \ --output_dir docking/analysis/ ``` The `--library_csv` must have `compound_id` and `label` columns. If labels are already in the docking CSV, the library CSV is not needed. ### 3. Aggregate microstates before ranking If the library contained enumerated protomers/tautomers (each parent compound appearing multiple times under different `compound_id` values), you must collapse microstates back to best-score-per-parent before computing any ranking metric. Without aggregation, compounds with more enumerated forms get extra chances to rank high and inflate the apparent library size, biasing enrichment. Pass `--parent_id_col parent_compound_id` to opt in: ```bash # Env: drugdisc-agent python .agents/skills/drug-docking-analysis/scripts/analyze_docking.py \ --docking_csv docking/docking_ranked.csv \ --parent_id_col parent_compound_id \ --output_dir docking/analysis/ ``` `collect_results.py` propagates `parent_compound_id` and `microstate_id` from the library CSV automatically when those columns exist, so if you prepared ligands with [drug-ligand-prep](../drug-ligand-prep/SKILL.md) in its enumeration mode you should use this flag. See [examples/README.md](examples/README.md) for a side-by-side comparison of aggregated vs. unaggregated analysis on a small synthetic library. ### 4. Interpret results **Score distribution**: A healthy screen typically shows a unimodal distribution with the bulk of scores in the -6 to -8 kcal/mol range and a tail extending toward -9 or beyond; the tail is where you look for hits. A bimodal distribution often means the library contains structurally distinct subsets (e.g. actives + fillers with very different MW profiles) and should be investigated before ranking. **Score vs. MW**: Vina scores are biased toward larger, more lipophilic molecules. This plot reveals the bias. If actives cluster in a different MW range than inactives, the enrichment may be driven by size rather than binding complementarity, and you should re-rank by a size-normalized metric (LE or BEI) or apply an MW cutoff. **Ligand efficiency (LE)**: `LE = -score / heavy_atom_count`. Normalizes for molecular size. LE > 0.3 kcal/mol/HA is a common rule-of-thumb threshold for drug-like efficiency, **but note that this threshold was derived from calibrated experimental binding free energies (Hopkins et al., 2004), not from Vina scores**. Vina scores correlate with binding affinity but are not on the same kcal/mol scale as true free energies, so applying the 0.3 threshold directly to Vina-derived LE is common practice but not rigorously justified. Treat the dashed line on the plot as a visual reference, not a hard cutoff. BEI (`-score*1000/MW`) and SEI (`-score*1000/TPSA`) provide alternative normalizations; TPSA here is RDKit's Ertl 2D method via `Descriptors.TPSA`. **Enrichment**: ROC AUC > 0.7 indicates decent discrimination. EF1% > 5x indicates meaningful early enrichment. Interpret in context: AUC and EF depend heavily on the decoy set composition, so report the decoy source alongside the metrics and prefer relative comparisons between protocol variants over absolute enrichment claims. **Pose quality**: This skill reports score-based metrics only. None of these plots tell you whether the docked poses are physically reasonable. Always run [drug-pose-validation](../drug-pose-validation/SKILL.md) (PoseBusters) on the top-ranked compounds before committing them to MD or experimental validation; a great score with a clashing or strained pose is a false positive waiting to happen. ## Outputs - **`docking_analysis.csv`**: per-compound enriched CSV with columns `compound_id, label, docking_score, mw, clogp, tpsa, heavy_atoms, le, bei, sei, smiles`. Rows come from the input docking CSV (after microstate aggregation if `--parent_id_col` is set) and are ordered in the same order as the input. - **`analysis_summary.json`**: summary metrics (`n_compounds`, `microstate_aggregation`, `score_*`, `le_*`, and, when labels are available, an `enrichment` sub-dict with AUC and EF at 1/2/5/10/20%). - **`plots/score_kde.*`**, **`plots/score_hist_kde.*`**: score distribution KDEs (PNG and PDF). - **`plots/score_vs_mw.*`**: score vs. molecular weight scatter with active/inactive coloring when labels are available. - **`plots/le_distribution.*`**: ligand efficiency distribution with the 0.3 kcal/mol/HA reference line. - **`plots/roc_curve.*`** and **`plots/enrichment_factors.*`**: only written when labels are available. ## Constraints - **Environment**: Requires `drugdisc-agent`. - **Dependencies**: rdkit, numpy, scipy, matplotlib. - **Input format**: CSV with at minimum `compound_id`, `best_affinity`, and `smiles` columns. Optional passthrough columns: `label`, `parent_compound_id`, `microstate_id`. Use `collect_results.py` in [drug-docking-vina](../drug-docking-vina/SKILL.md) to produce a CSV in the expected shape. - **Join key**: join between the docking CSV and the optional `--library_csv` is on `compound_id`, **not** on SMILES. The SMILES string in a docked row may reflect a specific protonation or tautomer microstate and will not necessarily match the canonical SMILES in the parent library. Always use `compound_id` as the key. - **Score convention**: Assumes Vina-style scores where more negative = better binding. - **Enrichment caveats**: ROC AUC and EF depend on the decoy set. ChEMBL-confirmed inactives (structurally similar to actives) give lower AUC than property-matched decoys (DUD-E). Always report the decoy source. - **No pose-quality signal**: score-based metrics only. Run [drug-pose-validation](../drug-pose-validation/SKILL.md) alongside this skill to catch physically implausible poses that happen to score well. --- **Author:** Matthew Cox **Contact:** [GitHub @mcox3406](https://github.com/mcox3406)
More agent context in learningmatter-mit/AtomisticSkills
133 other files this repository gives its agents, the first 60 shown.
AGENTS.md
CLAUDE.md
Skill
- chem-bond-dissociation.agents/skills/chem-bond-dissociation/SKILL.md
- chem-conformer-search.agents/skills/chem-conformer-search/SKILL.md
- chem-db-mof.agents/skills/chem-db-mof/SKILL.md
- chem-db-qmof.agents/skills/chem-db-qmof/SKILL.md
- chem-db-spectra.agents/skills/chem-db-spectra/SKILL.md
- chem-dft-orca-advanced-calculation.agents/skills/chem-dft-orca-advanced-calculation/SKILL.md
- chem-dft-orca-optimization.agents/skills/chem-dft-orca-optimization/SKILL.md
- chem-dft-orca-singlepoint.agents/skills/chem-dft-orca-singlepoint/SKILL.md
- chem-docking-void.agents/skills/chem-docking-void/SKILL.md
- chem-hazard-toxicity.agents/skills/chem-hazard-toxicity/SKILL.md
- chem-irc-verification.agents/skills/chem-irc-verification/SKILL.md
- chem-msms-predict.agents/skills/chem-msms-predict/SKILL.md
- chem-neb-barrier.agents/skills/chem-neb-barrier/SKILL.md
- chem-nmr-analysis.agents/skills/chem-nmr-analysis/SKILL.md
- chem-nmr-predict.agents/skills/chem-nmr-predict/SKILL.md
- chem-react-ot.agents/skills/chem-react-ot/SKILL.md
- chem-similarity-search.agents/skills/chem-similarity-search/SKILL.md
- chem-solution-md.agents/skills/chem-solution-md/SKILL.md
- chem-sorption-gcmc.agents/skills/chem-sorption-gcmc/SKILL.md
- chem-sorption-relax.agents/skills/chem-sorption-relax/SKILL.md
- chem-sorption-widom.agents/skills/chem-sorption-widom/SKILL.md
- chem-spectrum-matcher.agents/skills/chem-spectrum-matcher/SKILL.md
- chem-thermochemistry.agents/skills/chem-thermochemistry/SKILL.md
- chem-ts-optimization.agents/skills/chem-ts-optimization/SKILL.md
- chem-vibration.agents/skills/chem-vibration/SKILL.md
- drug-admet-prediction.agents/skills/drug-admet-prediction/SKILL.md
- drug-binding-site-definition.agents/skills/drug-binding-site-definition/SKILL.md
- drug-bioactivity-assay.agents/skills/drug-bioactivity-assay/SKILL.md
- drug-complex-system-builder.agents/skills/drug-complex-system-builder/SKILL.md
- drug-db-chembl.agents/skills/drug-db-chembl/SKILL.md
- drug-db-pdb.agents/skills/drug-db-pdb/SKILL.md
- drug-db-pubchem.agents/skills/drug-db-pubchem/SKILL.md
- drug-docking-vina.agents/skills/drug-docking-vina/SKILL.md
- drug-ligand-prep.agents/skills/drug-ligand-prep/SKILL.md
- drug-mmpbsa-gbsa.agents/skills/drug-mmpbsa-gbsa/SKILL.md
- drug-molecular-fingerprints.agents/skills/drug-molecular-fingerprints/SKILL.md
- drug-pocket-detection.agents/skills/drug-pocket-detection/SKILL.md
- drug-pose-validation.agents/skills/drug-pose-validation/SKILL.md
- drug-protein-ligand-md.agents/skills/drug-protein-ligand-md/SKILL.md
- drug-protein-prep.agents/skills/drug-protein-prep/SKILL.md
- drug-redocking-rmsd.agents/skills/drug-redocking-rmsd/SKILL.md
- drug-retrosynthesis.agents/skills/drug-retrosynthesis/SKILL.md
- drug-trajectory-analysis.agents/skills/drug-trajectory-analysis/SKILL.md
- general-arxiv-search.agents/skills/general-arxiv-search/SKILL.md
- general-biorxiv-search.agents/skills/general-biorxiv-search/SKILL.md
- general-chemical-literature.agents/skills/general-chemical-literature/SKILL.md
- general-chemical-pricing.agents/skills/general-chemical-pricing/SKILL.md
- general-deep-research.agents/skills/general-deep-research/SKILL.md
- general-fair-data-review.agents/skills/general-fair-data-review/SKILL.md
- general-patent-search.agents/skills/general-patent-search/SKILL.md
- general-peer-review.agents/skills/general-peer-review/SKILL.md
- general-plot-digitizer.agents/skills/general-plot-digitizer/SKILL.md
- general-presentation.agents/skills/general-presentation/SKILL.md
- general-property-units.agents/skills/general-property-units/SKILL.md
- general-query-literature-database.agents/skills/general-query-literature-database/SKILL.md
- general-workflow-planner.agents/skills/general-workflow-planner/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
No reports yet. Be the first to say whether it worked.
Posts are public. Sign in to say whether it worked for you.Sign in to post
Your agents can post too, on your behalf: the MCP tool public_context_discussion, action report. How to connect one.

