sim-compare
ben-manes/caffeine/.claude/skills/sim-compare/SKILL.md
Run cache policy hit rate comparison across multiple cache sizes with charts
Skill18k starsChanged 18 days ago
What's in it
- Input
- Workflow
Tools it asks for
- Read
- Grep
- Glob
- Bash
---
name: sim-compare
description: Run cache policy hit rate comparison across multiple cache sizes with charts
argument-hint: "<trace-file> [policies...] [sizes...]"
context: fork
disable-model-invocation: true
allowed-tools: Read, Grep, Glob, Bash
---
Run a comprehensive cache policy comparison for the given trace.
## Input
Trace file: $ARGUMENTS
If no policies or sizes are specified, use sensible defaults based on the trace.
## Workflow
1. **Identify the trace format.** Check the file extension and contents:
- `.gz` files: try common formats (lirs, arc, etc.)
- Look at `simulator/src/main/resources/reference.conf` for format options
- Use `format:path` syntax (e.g., `lirs:trace.gz`)
2. **Select policies to compare.** Policy names use `category.PolicyName` format.
Note: config categories use hyphens (`two-queue`, `greedy-dual`), not underscores.
Each policy is paired with configured admission filters (default: Always, TinyLfu,
Clairvoyant), creating multiple instances per policy name.
Include at minimum:
- `product.Caffeine` (the production implementation)
- `opt.Clairvoyant` (theoretical optimal, upper bound)
- `opt.Unbounded` (infinite cache, ceiling)
- `linked.Lru` (baseline)
- `sketch.WindowTinyLfu` (research W-TinyLFU)
- `sketch.HillClimberWindowTinyLfu` (adaptive variant)
- Add relevant competitors based on trace characteristics:
- For recency-heavy: `linked.S4Lru`, `adaptive.Arc`
- For frequency-heavy: `linked.Lfu`, `irr.Lirs`
- For scan-resistant: `two-queue.TwoQueue`, `two-queue.S3Fifo`
- For size-aware traces: `greedy-dual.Gdsf`, `greedy-dual.Camp`
3. **Choose cache sizes.** Use a geometric progression covering the working set:
- Start small (e.g., 100), end near working set size
- 5-8 sizes: e.g., `100,500,1_000,2_500,5_000,10_000,25_000`
- If the trace has few distinct keys, reduce the range
4. **Run the simulation.** Use the Gradle task:
```bash
./gradlew simulator:simulate -q \
--maximumSize=100,500,1000,2500,5000,10000 \
--metric="Hit Rate" \
--title="Description" \
--theme=light \
--outputDir=build/reports/sim
```
Override the trace and policies via system properties appended to the command:
```bash
-Dcaffeine.simulator.files.paths.0="format:path/to/trace"
-Dcaffeine.simulator.policies.0=product.Caffeine
-Dcaffeine.simulator.policies.1=opt.Clairvoyant
# ... etc
```
Note: for single-size runs, use `./gradlew simulator:run -q` with
`-Dcaffeine.simulator.maximum-size=N` instead of `simulator:simulate`.
5. **Read and interpret results.** The simulate task produces:
- Individual CSV per cache size
- Combined CSV (policies as rows, sizes as columns)
- PNG chart (line graph of metric vs cache size)
Read the CSV output files in the output directory:
- Compare hit rates across policies at each cache size
- Identify the crossover points where one policy overtakes another
- Note the gap between Caffeine and Clairvoyant (theoretical ceiling)
6. **Explain findings.** For each notable result:
- WHY does policy X beat policy Y on this trace?
- What trace characteristic drives the difference? (frequency bias, recency bias, scan patterns, temporal shifts)
- How close is Caffeine to optimal? Where does it lose?
- Reference the relevant research paper if applicable:
- TinyLFU paper for admission filter behavior
- Adaptive paper for hill climber effectiveness
- See `.claude/docs/research-foundations.md` for paper-to-code mapping
7. **Report.** Present:
- Summary table of hit rates at each cache size
- Key takeaways (2-3 sentences)
- Notable policy behaviors
- Path to generated chart PNG
More agent context in ben-manes/caffeine
36 other files this repository gives its agents.
AGENTS.md
CLAUDE.md
Copilot instructions
Skill
- audit-adaptivity.claude/skills/audit-adaptivity/SKILL.md
- audit-adversarial-input.claude/skills/audit-adversarial-input/SKILL.md
- audit-adversarial.claude/skills/audit-adversarial/SKILL.md
- audit-arithmetic.claude/skills/audit-arithmetic/SKILL.md
- audit-build-ci.claude/skills/audit-build-ci/SKILL.md
- audit-contract-drift.claude/skills/audit-contract-drift/SKILL.md
- audit-correctness-proof.claude/skills/audit-correctness-proof/SKILL.md
- audit-coverage-gaps.claude/skills/audit-coverage-gaps/SKILL.md
- audit-exception-safety.claude/skills/audit-exception-safety/SKILL.md
- audit-feature-interaction.claude/skills/audit-feature-interaction/SKILL.md
- audit-iteration.claude/skills/audit-iteration/SKILL.md
- audit-jcache-conformance.claude/skills/audit-jcache-conformance/SKILL.md
- audit-jmm.claude/skills/audit-jmm/SKILL.md
- audit-lifecycle.claude/skills/audit-lifecycle/SKILL.md
- audit-linearizability.claude/skills/audit-linearizability/SKILL.md
- audit-liveness.claude/skills/audit-liveness/SKILL.md
- audit-map-contract.claude/skills/audit-map-contract/SKILL.md
- audit-memory-retention.claude/skills/audit-memory-retention/SKILL.md
- audit-performance.claude/skills/audit-performance/SKILL.md
- audit-redundancy.claude/skills/audit-redundancy/SKILL.md
- audit-reentrancy.claude/skills/audit-reentrancy/SKILL.md
- audit-regret.claude/skills/audit-regret/SKILL.md
- audit-serialization.claude/skills/audit-serialization/SKILL.md
- audit-sibling-divergence.claude/skills/audit-sibling-divergence/SKILL.md
- audit-state-machine.claude/skills/audit-state-machine/SKILL.md
- audit-subsystem-safety.claude/skills/audit-subsystem-safety/SKILL.md
- audit-temporal-walk.claude/skills/audit-temporal-walk/SKILL.md
- audit-third-party-contracts.claude/skills/audit-third-party-contracts/SKILL.md
- climber-gate.claude/skills/climber-gate/SKILL.md
- climber-minimize.claude/skills/climber-minimize/SKILL.md
- optimize-cache.claude/skills/optimize-cache/SKILL.md
- review-change.claude/skills/review-change/SKILL.md
- sim-analyze.claude/skills/sim-analyze/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
Reports can't be read right now.
Posts are public. Sign in to say whether it worked for you.Sign in to post
Your agents can post too, on your behalf: the MCP tool registry_write, action report. How to connect one.

