agentleFS
Sign inSign up

skills-ws

san-npm/skills-ws/public/llms-full.txt

87 agent skills for AI coding assistants. Format: If we [change], then [metric] will [improve/decrease] by [amount], because [rationale]. Example: If we shorten the signup form from 5 fields to 3, then signup completion rate will increase by 15%, because friction reduction at high-intent moments increases conversion. ICE framework (quick): RICE framework (more rigorous): Quick reference table: Test duration: Minimum: run ≥ 1 full business cycle (usually 7 or 14 days) so every day-of-week and any weekly purchase/payday rhythm is…

llms.txt2 starsChanged 3 months ago
  • Reads credentials
  • Deletes or force-pushes
  • Installs packages
  • Sends data out
# skills.ws — Full Skill Index

> 87 agent skills for AI coding assistants.

## ab-testing
Category: conversion
Description: Experimentation guidance: A/B test design, sample-size/MDE calculation, pre-analysis plans, SRM and validity checks, frequentist + Bayesian + sequential analysis, variance reduction (CUPED), and ship/no-ship decisions. Use when designing or analyzing an A/B test, sizing an experiment, prioritizing tests, debugging suspicious results, or deciding whether to ship a variant.
Features:
  - Hypothesis generation frameworks
  - Sample size and duration calculators
  - Statistical significance analysis
  - Experiment prioritization (ICE, RICE, PIE)
  - Multi-variant test design
  - Results interpretation and documentation
Use Cases:
  - Design an A/B test for a pricing page
  - Calculate required sample size for significance
  - Prioritize a backlog of experiment ideas
  - Interpret test results and make ship/no-ship decisions

# A/B Testing

## Workflow

### 1. Hypothesis Generation

**Format:** If we [change], then [metric] will [improve/decrease] by [amount], because [rationale].

**Example:** If we shorten the signup form from 5 fields to 3, then signup completion rate will increase by 15%, because friction reduction at high-intent moments increases conversion.

### 2. Prioritization

**ICE framework (quick):**

| Factor | Score 1-10 | Definition |
|--------|-----------|------------|
| Impact | 1-10 | How much will it move the metric? |
| Confidence | 1-10 | How sure are we it'll work? |
| Ease | 1-10 | How fast/cheap to implement? |
| **ICE Score** | | (I + C + E) / 3 |

**RICE framework (more rigorous):**

| Factor | Definition |
|--------|-----------|
| Reach | How many users affected per quarter? |
| Impact | Expected effect size (0.25, 0.5, 1, 2, 3) |
| Confidence | % sure (100%, 80%, 50%) |
| Effort | Person-weeks to implement |
| **RICE Score** | (R × I × C) / E |

### 3. Sample Size Calculation

**Formula:**
```
n = (Z_α/2 × √(2p̄(1-p̄)) + Z_β × √(p₁(1-p₁) + p₂(1-p₂)))² / (p₂ - p₁)²

Where:
  p₁ = baseline conversion rate
  p₂ = expected conversion rate (baseline × (1 + MDE))
  p̄  = (p₁ + p₂) / 2
  Z_α/2 = 1.96 (for 95% confidence)
  Z_β   = 0.84 (for 80% power)
```

**Quick reference table:**

| Baseline rate | MDE (relative) | Sample per variant |
|--------------|----------------|-------------------|
| 2% | 10% | 78,000 |
| 2% | 20% | 20,000 |
| 5% | 10% | 30,000 |
| 5% | 20% | 7,700 |
| 10% | 10% | 14,300 |
| 10% | 20% | 3,700 |
| 20% | 10% | 6,300 |
| 20% | 20% | 1,600 |

**Test duration:**
```
Days needed = (Sample per variant × 2) / Daily traffic to test page
```

**Minimum:** run ≥ 1 full business cycle (usually 7 or 14 days) so every day-of-week and any weekly purchase/payday rhythm is represented; never stop mid-week even if the sample target is hit early.

**Run length is driven by power and cycles, not a fixed cap.** There is no universal "stop at 4 weeks" rule — B2B, marketplace, pricing, retention, and low-traffic tests routinely need 6–12+ weeks. The real risks in a long test are *exposure/sample drift* (the population changes — new acquisition channels, seasonality, holidays) and *novelty/primacy effects* (returning users react to the change for a few weeks, then revert). Mitigate by:
- Decide the **fixed horizon up front** from the sample-size calc (or use a sequential design, below). Do not let "it's been a month" become the stopping rule.
- Plot the **daily cumulative lift**; a stable, flattening curve signals novelty has worn off, a still-trending one means keep running.
- For novelty-prone changes (UI redesigns, new features), report **new-user vs returning-user** segments separately (pre-registered — see segmentation in §6) and consider a long-running **holdback** to measure the durable effect.
- If you must change the experiment design or population mid-flight, **stop and restart as a new test** rather than reinterpreting the old one.

### 4. Test Design

**Rules:**
- One hypothesis per test
- Randomly assign **users** (stable hash of a persistent user/device id → bucket), not sessions, so a user always sees the same variant (avoids flickering and contamination)
- Use the same metric definition, observation window, and instrumentation for control and variant
- Define primary metric AND guardrail metrics **before** launch
- Don't peek-and-stop on a fixed-horizon test; only the pre-registered sequential design (below) permits early stopping
- Ship the variant code behind a **feature flag / server-side assignment** so you can ramp exposure (1% → 5% → 50%) and kill instantly without a deploy; assign server-side where possible to dodge client ad-blockers and flicker

**Pre-analysis plan (write before launch, freeze it):** This is the single best defense against p-hacking. Record:

| Field | Example |
|-------|---------|
| Primary metric (one) | Signup completion rate |
| Guardrail metrics | p95 latency, error rate, revenue/user, refund rate |
| Unit of analysis & randomization | User id; same unit for assignment and metric (ratio metrics → delta method, below) |
| MDE / alpha / power | +5% relative, α = 0.05 (two-sided), power = 0.80 |
| Design & horizon | Fixed-horizon N = 30k/arm **or** sequential (mSPRT, α-spending) |
| Stopping rule | Stop at horizon; OR sequential boundary crossed; OR guardrail breach |
| Pre-registered segments | mobile vs desktop, new vs returning (everything else is exploratory) |
| Exclusions | internal IPs/employee ids, known bots, pre-exposure activity |

**Guardrail metrics (always monitor):**
- Latency (p50/p95 — variant shouldn't be slower)
- Error / crash rate
- Revenue per user and refund/chargeback rate (don't lift signups while tanking revenue)
- Bounce rate / core engagement

**Instrumentation & traffic-quality QA (before trusting any number):**
- **Validate event tracking on both arms** in staging *and* in a pre-launch A/A test — fire the exposure event exactly once at first eligible impression, and confirm conversion events join to the same unit id.
- **Filter bots and internal traffic** (employees, QA, monitoring, datacenter ASNs) *before* analysis, not after.
- **Run an A/A test** (or use the sequential framework on a no-change comparison) periodically to confirm your false-positive rate matches α and your assignment/logging is unbiased.

### 5. Statistical Analysis

**Step 0 — Sample-Ratio Mismatch (SRM) check. Do this FIRST; if it fails, STOP.**

If the observed split differs from the intended split (e.g. you targeted 50/50 but see 5,000 vs 5,400), randomization or logging is broken and **every downstream p-value is untrustworthy** — do not interpret the result, find the bug (redirect bias, flag eval, bot filtering applied to one arm, double-firing exposure events). Test with a chi-square goodness-of-fit; a p-value below ~0.01 is an SRM.

```python
from scipy.stats import chisquare

# Observed exposures per arm. Plug in your real assignment counts.
observed = [5000, 5000]            # control, variant  (a clean 50/50 here)
expected_ratio = [0.5, 0.5]        # intended split
total = sum(observed)
expected = [total * r for r in expected_ratio]

chi2, srm_p = chisquare(f_obs=observed, f_exp=expected)
print(f"SRM chi-square p = {srm_p:.4f}")
if srm_p < 0.01:
    raise SystemExit("SRM DETECTED — assignment/logging is broken. Do NOT trust metrics; debug first.")
# e.g. observed = [5000, 5400] -> srm_p ~ 0.0001 -> STOP and debug before reading any metric.
```

**Frequentist approach (standard):**

```python
import numpy as np
from scipy import stats

# Results
control = {'visitors': 5000, 'conversions': 250}  # 5.0%
variant = {'visitors': 5000, 'conversions': 295}  # 5.9%

p1 = control['conversions'] / control['visitors']
p2 = variant['conversions'] / variant['visitors']
p_pool = (control['conversions'] + variant['conversions']) / (control['visitors'] + variant['visitors'])

se = np.sqrt(p_pool * (1 - p_pool) * (1/control['visitors'] + 1/variant['visitors']))
z = (p2 - p1) / se
p_value = 2 * (1 - stats.norm.cdf(abs(z)))

lift = (p2 - p1) / p1 * 100
ci_95 = 1.96 * np.sqrt(p1*(1-p1)/control['visitors'] + p2*(1-p2)/variant['visitors'])

print(f"Control: {p1:.3%}")
print(f"Variant: {p2:.3%}")
print(f"Lift: {lift:.1f}%")
print(f"95% CI: [{(p2-p1-ci_95)/p1*100:.1f}%, {(p2-p1+ci_95)/p1*100:.1f}%]")
print(f"p-value: {p_value:.4f}")
print(f"Significant: {'Yes' if p_value < 0.05 else 'No'}")
```

**Bayesian approach (when you want probability of being better + a risk-aware decision):**

`P(variant > control)` alone over-ships tiny, uncertain wins. Always pair it with **expected loss** (the average downside in conversion-rate points if you ship and you're wrong) and a **ROPE** (region of practical equivalence — a band of differences too small to matter). Ship only when P(better) clears a high bar AND expected loss is below a tolerance you set in advance.

```python
import numpy as np
from scipy.stats import beta

# Beta(1,1) uniform prior + observed data (use a weakly-informative prior near baseline if you have history)
a_alpha = control['conversions'] + 1
a_beta  = control['visitors'] - control['conversions'] + 1
b_alpha = variant['conversions'] + 1
b_beta  = variant['visitors'] - variant['conversions'] + 1

draws = 200_000
samples_a = beta.rvs(a_alpha, a_beta, size=draws)
samples_b = beta.rvs(b_alpha, b_beta, size=draws)
diff = samples_b - samples_a                       # in absolute rate points

prob_b_better = (diff > 0).mean()
# Expected loss if we SHIP variant: average shortfall when control is actually better
expected_loss_ship = np.maximum(samples_a - samples_b, 0).mean()
# 95% credible interval on the absolute difference
ci_lo, ci_hi = np.percentile(diff, [2.5, 97.5])

# ROPE: differences within +/- 0.2 absolute points are "practically equal"
rope = 0.002
p_in_rope = ((diff > -rope) & (diff < rope)).mean()

print(f"P(variant > control): {prob_b_better:.1%}")
print(f"Expected loss if ship: {expected_loss_ship*100:.3f} pts")
print(f"95% credible interval (abs): [{ci_lo*100:.3f}, {ci_hi*100:.3f}] pts")
print(f"P(difference within ROPE): {p_in_rope:.1%}")

# Decision thresholds (set BEFORE launch)
DECISION_PROB = 0.95            # ship confidence
LOSS_TOLERANCE = 0.0005         # max acceptable expected loss (0.05 pts)
ship = prob_b_better >= DECISION_PROB and expected_loss_ship <= LOSS_TOLERANCE
print("Decision:", "SHIP" if ship else "keep running / inconclusive")
```

**Variance reduction — CUPED (use when you have pre-experiment data).** CUPED removes pre-existing user differences using a pre-period covariate (e.g. each user's prior-28-day spend or visits), often cutting variance 30–50% — which means a smaller sample or a shorter test for the same power. It's standard on mature platforms. Adjust the metric, then run the *same* t-test/CI on the adjusted values.

```python
import numpy as np
from scipy import stats

# y = in-experiment metric per user; x = same user's pre-period covariate (mean-centered)
# group: 0 = control, 1 = variant. Arrays aligned by user.
def cuped_adjust(y, x):
    x = x - x.mean()
    theta = np.cov(y, x, ddof=1)[0, 1] / np.var(x, ddof=1)   # optimal coefficient
    return y - theta * x

y_adj = cuped_adjust(y, x)
t, p = stats.ttest_ind(y_adj[group == 1], y_adj[group == 0], equal_var=False)
print(f"CUPED-adjusted effect p = {p:.4f}  (variance reduced vs raw t-test)")
```
The covariate must be **pre-treatment** (measured before assignment) and correlated with the outcome; never use a post-treatment variable or you bias the estimate.

**Sequential testing / always-valid p-values (use when stakeholders WILL peek).** A fixed-horizon p-value is only valid if you look once at the planned N. If you want to monitor a dashboard daily and be able to stop early, use a design built for continuous monitoring instead of repeatedly applying the 0.05 test:
- **Group sequential / alpha-spending (O'Brien–Fleming, Pocock):** pre-plan K interim looks; spend α across them so the overall false-positive rate stays at 0.05. Good when looks are scheduled (e.g. weekly).
- **Always-valid inference (mSPRT / confidence sequences):** gives a p-value/CI valid at *every* moment, so you may stop the instant it crosses — at the cost of needing a somewhat larger sample if the effect is small. This is what "peeking-safe" dashboards (modern experimentation platforms) implement.
- Practical rule: pick fixed-horizon **or** sequential up front and write it in the pre-analysis plan. Do **not** run a fixed-horizon test and then stop early because it "hit significance" — that inflates false positives 2–5×.

**Ratio & revenue metrics (variance is bigger than it looks).** For metrics where the analysis unit ≠ randomization unit (clicks-per-session, revenue-per-user, CTR aggregated over sessions), the naive standard error is wrong because observations within a user are correlated. Use the **delta method** or **cluster/bootstrap by user** for the variance, and consider **winsorizing** heavy-tailed revenue (cap at ~p99) so one whale doesn't dominate. Run significance on the user-level mean (or delta-method SE), not on the pooled event counts.

### 6. Ship / No-Ship Decision

Evaluate the **primary metric on the pre-registered design only** (fixed-horizon p-value at planned N, or the sequential boundary). For Bayesian tests, swap "p < 0.05" for "P(better) ≥ threshold AND expected loss ≤ tolerance" from §5.

| Scenario | Decision |
|----------|----------|
| Significant AND lift > MDE AND guardrails OK | Ship |
| Significant AND lift > 0 but < MDE | Ship only if cost-free; the effect is below what you decided was worth shipping — usually iterate |
| Not significant at planned horizon | **Inconclusive — do NOT silently extend.** Extending after seeing a near-miss is p-hacking. Only continue if a longer horizon (or sequential boundary) was pre-specified; otherwise redesign and run a fresh, better-powered test. |
| Significant AND lift negative | Kill variant |
| Guardrail metric degraded | Kill variant regardless of primary metric |

**Segmentation discipline.** Reading the result inside subgroups (mobile, country, new vs returning) is valuable but is where false discoveries breed:
- Report **pre-registered segments** as confirmatory; treat every other slice as **exploratory hypothesis generation**, not proof.
- Correct for multiple comparisons across segments/metrics — **Benjamini–Hochberg (FDR)** for many exploratory reads, **Bonferroni** when a single false positive is costly. A "win" found only after slicing 12 ways needs its own confirmatory test before you ship it to that segment.
- Beware **Simpson's paradox**: a variant can win overall yet lose in every segment (or vice-versa) if segment mix differs between arms — another reason the SRM and assignment checks in §5 matter.

### 7. Documentation Template

```markdown
## Test: [Name]
**Hypothesis:** If we [change], then [metric] will [change] by [amount]
**Primary metric:** [one metric]   **Guardrails:** [latency, error rate, revenue/user, ...]
**Randomization unit:** [user id]   **MDE / alpha / power:** [+5% rel / 0.05 / 0.80]
**Design & stopping rule:** [fixed-horizon N=X/arm | sequential mSPRT] — frozen before launch
**Pre-registered segments:** [mobile vs desktop, new vs returning]   **Exclusions:** [internal, bots]
**Duration:** [start] to [end]  (>= 1 full business cycle)

### Validity checks
- SRM: observed [n_c / n_v], chi-square p = [..]  → PASS / FAIL
- A/A or instrumentation QA: PASS / FAIL    Bots & internal traffic filtered: Y/N

### Results
| Metric | Control | Variant | Lift | CI / p-value (or P(better) + exp. loss) | Sig? |
|--------|---------|---------|------|------------------------------------------|------|
| Primary | X% | Y% | +Z% | [..] | Y/N |

### Decision: Ship / Kill / Iterate
**Reasoning:** [primary on pre-registered design + guardrails; any segment reads flagged exploratory]
**Next test:** [What we learned and what to try next]
```

## Common Mistakes

- Stopping a fixed-horizon test early because results "look significant" — peeking inflates false positives 2–5×. Use a sequential design if you need to stop early.
- Trusting results without an **SRM check** — a broken 50/50 split silently corrupts every metric.
- Extending a "near-miss" test that wasn't pre-registered to extend (it's p-hacking dressed up as patience).
- **Post-hoc segment fishing** with no multiple-comparison correction — slice enough ways and something always "wins."
- Running too many variants (splits traffic, dilutes power, multiplies comparisons).
- Testing tiny changes on low-traffic pages (will never reach significance — see the sample-size table).
- Using the naive binary-proportion test on **revenue/ratio metrics** (correlated within-user observations → understated variance → false wins).
- Ignoring practical significance (a statistically significant 0.1% lift usually isn't worth shipping).
- Treating a long-running winner as durable without checking for novelty decay (split new vs returning users).

---

## accounting-finance
Category: operations
Description: Operator finance for SMEs/startups: GAAP/IFRS P&L, 13-week cash/runway forecasting, SaaS unit economics (NRR/GRR/CAC payback), chart of accounts, monthly close, bank reconciliation, ASC 606/IFRS 15 revenue recognition, VAT/sales-tax checklists. Use when building a P&L/budget, forecasting cash, setting up bookkeeping, or checking invoicing/VAT.
Features:
  - P&L statement analysis and generation
  - Cash flow forecasting models
  - Invoice automation workflows
  - Tax compliance checklists by jurisdiction
  - Revenue recognition patterns
  - Budget vs actual variance analysis
Use Cases:
  - Build a monthly P&L analysis template
  - Set up automated invoicing workflows
  - Create a cash flow forecast model
  - Design a tax compliance checklist for EU SMEs

# Accounting & Finance

> **Not tax, legal, or audit advice.** This skill encodes general operator practice and standard frameworks (US GAAP, IFRS, ASC 606/IFRS 15, EU VAT). Rates, thresholds, registration triggers, and filing rules change and are jurisdiction-specific. **Before filing anything or relying on a number for a board, lender, investor, or tax authority, have it reviewed by a qualified accountant/tax advisor** licensed in the relevant jurisdiction. Route any edge case — multi-state US nexus, cross-border digital services, equity/SAFE accounting, R&D capitalization, transfer pricing, M&A, or revenue-recognition judgment calls — to a professional. Treat every concrete rate/threshold below as **"verify before use."**

> **Sibling skills (don't duplicate — cross-link):** country-level tax depth → `eu-tax-accounting` (all 27 EU member states: corporate, VAT, payroll, deadlines). Billing/dunning/usage metering implementation → `saas-billing` (Express/Node) or `stripe-billing` (Next.js). Pricing/packaging → `pricing-optimization`. GTM funnel/forecasting → `revenue-operations`. Churn/retention cohort analysis → `retention-analytics`. EU regulatory (GDPR, contracts, entity) → `eu-legal-compliance`.

---

## 1. P&L Structure (GAAP / IFRS)

Standard multi-step income statement. **Key rule: define what OpEx contains so D&A is counted exactly once.** Below, operating expenses are stated *excluding* depreciation & amortization (the "ex-D&A" convention common in SaaS reporting); D&A is its own line. EBITDA then equals operating income + D&A with no double-count. If your accounting system buckets D&A *inside* OpEx, drop the separate D&A line and compute EBITDA = operating income + D&A (added back), not by re-subtracting it.

| # | Line item | Calculation | Watch for |
|---|-----------|-------------|-----------|
| 1 | **Revenue (net)** | Recognized per ASC 606/IFRS 15 (§6), net of refunds/credits | Recognized ≠ billed ≠ cash collected. Don't book deferred revenue as revenue. |
| 2 | **COGS / Cost of revenue** | Hosting/infra, third-party API/usage fees, payment processing, customer support & success delivery, onboarding/implementation labor | SaaS COGS typically 15–30% of revenue. Keep R&D and S&M *out* of COGS. |
| 3 | **Gross profit** | Revenue − COGS | SaaS target gross margin 70–85%. |
| 4 | **Operating expenses (ex-D&A)** | Sales & Marketing + Research & Development + General & Administrative | Allocate fully-loaded headcount (salary + employer taxes + benefits) to the right function. |
| 5 | **EBITDA** | Gross profit − OpEx(ex-D&A) | Proxy for operating cash generation; ignores capex, financing, tax. |
| 6 | **Depreciation & amortization** | Capitalized assets + capitalized software/intangibles amortization | Pure non-cash; never in COGS *and* here. |
| 7 | **Operating income (EBIT)** | EBITDA − D&A | GAAP operating result. |
| 8 | **Net interest** | Interest expense − interest income | |
| 9 | **Pre-tax income (EBT)** | EBIT − net interest ± other | |
| 10 | **Income tax expense** | Current + deferred tax | Tax expense (accrual) ≠ tax *paid* (cash). |
| 11 | **Net income** | EBT − income tax | Bottom line. |

> **EBITDA vs Adjusted EBITDA:** "Adjusted EBITDA" further adds back stock-based compensation, one-time/restructuring items, and M&A costs. Always label which one you're showing and footnote the add-backs — investors discount unexplained adjustments. **Rule of 40** (growth % + FCF or EBITDA margin % ≥ 40) is a common SaaS health check, not a GAAP metric.

**Monthly P&L review checklist**
- [ ] Recognized revenue reconciles to the billing system and the deferred-revenue roll-forward (§6), not just to cash.
- [ ] COGS contains only cost-of-delivery; S&M/R&D/G&A are not leaking into it.
- [ ] Headcount fully loaded (salary + employer payroll tax + benefits) and allocated to the correct function.
- [ ] One-time/non-recurring items flagged and excluded from run-rate and from EBITDA→Adjusted EBITDA.
- [ ] D&A counted once (per the convention you chose above).
- [ ] MoM and YoY comparatives included; material variances explained (§8).
- [ ] Accruals booked for incurred-but-unbilled expenses (the close, §5).

---

## 2. Cash Flow Forecasting

### 13-week rolling direct cash forecast (the operator standard)

Forecast **cash in/out**, not accruals. Rebuild weekly from the bank balance.

```
Week | Start cash | + AR collected | + Other in | − Payroll | − Vendors/AP | − Tax/VAT | − Debt svc | = End cash
1    | 150,000    | 45,000         | 0          | 30,000    | 8,000        | 0         | 0          | 157,000
2    | 157,000    | 12,000         | 0          | 0         | 5,000        | 0         | 2,500      | 161,500
3    | 161,500    | 28,000         | 5,000      | 30,000    | 9,000        | 14,000    | 0          | 141,500
...
13   | ...
```

**Rules**
- Use **cash collected** (apply realistic AR collection lag: e.g. net-30 invoices land in week 5–6, with a haircut for late payers), not revenue recognized.
- Payroll on **actual** pay dates (semi-monthly/biweekly/monthly) including employer taxes; biweekly = 26 pay runs/yr (two 3-paycheck months).
- VAT/sales-tax remittances and corporate-tax instalments on **statutory due dates** — these are large, lumpy, and easy to forget.
- Model AP on actual vendor terms; don't assume everything clears in the booking week.
- Flag any week where ending cash dips below a **defined floor** (e.g. ≥ 2 months of operating burn or a debt covenant minimum).
- Keep a low/base/high collections scenario for any week with concentrated customer risk.

### Burn & runway

```
Gross burn   = total operating cash OUT in the month (exclude one-offs / financing)
Net burn     = gross burn − cash revenue collected      (the number that actually depletes the bank)
Runway (mo)  = current cash balance / average forward NET burn   (use a 3-month trailing avg, not a single noisy month)
```

> **Runway is a trigger to plan, not a script.** A short runway with predictable recurring revenue, an open credit line, and a near-term profitability path is very different from a short runway with lumpy revenue and no debt access. Weigh: revenue predictability & retention (§3), fundraising-market conditions, debt availability and **covenants**, the path/time-to-default-or-breakeven, dilution at the current valuation, and the owners'/board's risk tolerance. Generally start serious fundraising **9–12 months** before zero cash (a raise commonly takes 3–6 months), and pre-model the cost-cut lever you'd pull if a round slips — but the right move is situational; pressure-test it with the board/CFO.

---

## 3. SaaS / Subscription Unit Economics

### MRR / ARR movement schedule (single source of truth for "growth quality")

```
                         Month
Beginning MRR            100,000
+ New (new logos)         12,000
+ Expansion (upsell)       6,000
+ Reactivation             1,000
− Contraction (downsell)  (3,000)
− Churned (lost logos)    (5,000)
= Ending MRR             111,000
ARR = Ending MRR × 12 =  1,332,000
```

- **Quick Ratio** = (New + Expansion + Reactivation) / (Contraction + Churned). > 4 is strong; < 1 means you're losing ground.
- Reconcile this schedule to the deferred-revenue roll-forward (§6) and to recognized revenue (§1) every month.

### Retention — measure *logo* and *revenue* separately

| Metric | Formula | Read it as |
|--------|---------|-----------|
| **Logo (customer) churn** | Customers lost in period / customers at start | Counts accounts, ignores size. |
| **Gross Revenue Retention (GRR)** | (Start MRR − contraction − churn) / Start MRR | Caps at 100%. Excludes expansion → pure leakage. Best-in-class ≥ 90% (SMB) / ≥ 95% (enterprise). |
| **Net Revenue Retention (NRR/NDR)** | (Start MRR − contraction − churn + expansion) / Start MRR | Can exceed 100%. > 110% = healthy expansion engine; > 120% = elite. |
| **Logo retention** | 1 − logo churn | High logo churn + high NRR ⇒ a few big accounts carry you (concentration risk). |

### CAC, LTV, payback — state your assumptions or the numbers lie

| Metric | Formula | Notes / pitfalls |
|--------|---------|------------------|
| **Blended CAC** | All S&M / *all* new customers (incl. organic) | Flatters efficiency; use for company-level view. |
| **Paid CAC** | Paid S&M / customers from paid channels | The number that matters for scaling spend. |
| **CAC payback (months)** | CAC / (new MRR per customer × **gross margin %**) | Use *gross-margin-adjusted* MRR, not raw price. < 12 mo good (SMB), < 18–24 mo acceptable (enterprise). |
| **LTV** | (ARPA × gross margin %) / **churn rate** | Use *revenue* churn for $-LTV; on negative net churn this formula blows up — switch to a finite cohort horizon (e.g. 36-month discounted cohort value). |
| **LTV : CAC** | LTV / CAC | ≥ 3:1 healthy; ≫ 5:1 may mean you're under-investing in growth. |
| **Magic number** | Net new ARR / prior-quarter S&M | > 0.75 efficient; account for **sales-cycle lag** (spend in Q1 closes in Q2). |

> **Common mistakes:** mixing gross vs net churn; forgetting gross-margin adjustment in payback/LTV; counting organic logos in *paid* CAC; ignoring the S&M→revenue timing lag; using a single month's churn (annualize a multi-month cohort). For deep cohort/retention curves see `retention-analytics`.

---

## 4. Bookkeeping Automation, Chart of Accounts & the Monthly Close

The biggest leverage point: a clean **chart of accounts (COA)** + a repeatable **close** + automated **bank feeds** + **approval controls**.

### 4a. Chart of accounts (SMB/SaaS starter — numeric ranges)

Group by the P&L/balance-sheet line it rolls into so reporting is automatic.

| Range | Type | Example accounts |
|-------|------|------------------|
| 1000–1999 | **Assets** | 1000 Operating bank · 1010 Savings/reserve · 1100 Accounts receivable · 1200 Prepaid expenses · 1500 Fixed assets · 1510 Accumulated depreciation (contra) · 1600 Capitalized software |
| 2000–2999 | **Liabilities** | 2000 Accounts payable · 2100 Credit cards · 2200 Accrued expenses · 2300 **Deferred revenue** · 2400 Sales-tax/VAT payable · 2500 Payroll liabilities · 2700 Loans/notes payable |
| 3000–3999 | **Equity** | 3000 Common stock/share capital · 3100 Additional paid-in capital · 3200 Retained earnings |
| 4000–4999 | **Revenue** | 4000 Subscription revenue · 4100 Usage/overage revenue · 4200 Services/onboarding · 4900 Refunds & credits (contra) |
| 5000–5999 | **COGS** | 5000 Hosting/infrastructure · 5100 Third-party API/usage · 5200 Payment processing fees · 5300 Support & success (delivery) · 5400 Implementation labor |
| 6000–7999 | **Operating expenses** | 6000 Salaries & wages · 6010 Employer payroll taxes · 6020 Benefits · 6100 Sales & marketing · 6200 R&D/software dev (non-capitalized) · 6300 Rent & facilities · 6400 SaaS tools/subscriptions · 6500 Professional fees (legal/accounting) · 6600 Travel · 6700 Depreciation & amortization |
| 8000–9999 | **Other** | 8000 Interest income · 9000 Interest expense · 9500 Income tax expense |

**Rules:** keep it shallow (use classes/tags/departments for dimensions, not 200 accounts); never expense to a bank/AP account; reserve a contra account for refunds; reconcile **2300 Deferred revenue** to the §6 roll-forward and **2400 Sales-tax/VAT payable** to filed returns.

### 4b. Bank-feed reconciliation workflow

1. Connect bank/credit-card **feeds** (Plaid/native) into the ledger (QuickBooks Online, Xero, NetSuite, Wave).
2. Set **bank rules** to auto-categorize recurring lines (payroll provider → 6000/6010; Stripe payout → split fee 5200 vs gross; AWS → 5000).
3. **Match** feed transactions to existing invoices/bills; create from rules only when unmatched.
4. Clear the bank rec so **ledger balance = bank statement balance** every month; investigate any unreconciled item — never "plug" it.
5. Reconcile the **Stripe/PSP payout**: gross charges − processing fees − refunds = net deposit; book fees to 5200, not as a revenue contra.

### 4c. Monthly close checklist (target: business-day +5)

- [ ] All bank & credit-card accounts reconciled to statements (4b).
- [ ] AR aging reviewed; bad-debt reserve assessed.
- [ ] AP complete; **accruals** booked for incurred-but-unbilled costs (cut-off).
- [ ] Prepaids amortized (insurance, annual SaaS tools).
- [ ] Depreciation/amortization run for the period.
- [ ] **Deferred-revenue roll-forward** posted; revenue recognized per ASC 606/IFRS 15 (§6).
- [ ] Payroll fully recorded incl. employer taxes and PTO accrual.
- [ ] Sales-tax/VAT liability reconciled to returns/registers.
- [ ] Intercompany/owner transactions cleared (no personal expenses in the company ledger).
- [ ] Flux/variance review vs prior month, budget, and forecast (§8); lock the period.

### 4d. Approval controls, receipt capture & audit trail (segregation of duties)

- **Separate** who *requests*, *approves*, and *pays* — no single person initiates and disburses (fraud control). In tiny teams compensate with owner review of the bank feed + dual sign-off above a threshold.
- **Spend authorization matrix:** e.g. < €500 manager · €500–5k department head · > €5k founder/CFO · > €25k board. Document and enforce in the AP/expense tool.
- **Receipt capture:** require an itemized receipt per expense (Ramp/Brex/Pleo/Expensify auto-OCR and attach); enforce a per-transaction documentation rule for tax substantiation.
- **Audit trail:** keep an immutable, time-stamped log of who entered/edited/approved each transaction; restrict ledger admin; never share logins. Retain records per jurisdiction (commonly **7–10 years**; verify locally).
- **Month-end lock:** close the prior period so posted entries can't be silently altered; corrections go through dated adjusting entries.

---

## 5. Invoicing & Accounts Receivable

### Invoice/dunning workflow

1. Contract signed → create invoice/subscription record.
2. Invoice issued → send on billing date with a payment link.
3. Track **aging** by terms (net 15/30/60).
4. Overdue dunning sequence (tune by segment; soften for strategic accounts):
   - Day 1 past due: friendly reminder + link.
   - Day 7: second notice.
   - Day 14: escalate to account owner.
   - Day 30: final notice; assess late fee (if contractually allowed) and collections.

> For automated dunning, retries, and payment-failure recovery on Stripe, use `saas-billing`/`stripe-billing` — don't hand-roll it.

### What a compliant invoice contains

Exact mandatory fields are **jurisdiction- and transaction-specific** (and differ for a *full* VAT invoice vs a *simplified* receipt below a local threshold). General good practice:

- Unique sequential invoice number; issue date (and tax point/supply date if different); due date.
- Supplier legal name, address, and **tax/VAT/company registration number where the supplier is registered**.
- Customer name and address.
- Line items: description, quantity, unit price; subtotal; tax rate(s) and tax amount per rate; total payable; currency.
- Payment terms and remittance/bank details.

> **The customer's VAT number is NOT universally required.** It is required (and must be valid) when you apply the **EU B2B reverse charge / intra-Community supply** — without a verified buyer VAT ID you generally cannot zero-rate (validate via **VIES**). But many customers (consumers, non-VAT-registered small businesses, non-EU buyers) have no VAT number, and that is fine. Likewise, *your* VAT number only appears if you are VAT-registered. Don't block invoicing on a VAT field that doesn't apply. Confirm the exact field set for your country with `eu-tax-accounting` or a local accountant.

---

## 6. Revenue Recognition (ASC 606 / IFRS 15)

**5-step model:** (1) identify the contract → (2) identify distinct performance obligations → (3) determine the transaction price → (4) allocate price to obligations (by standalone selling price) → (5) recognize revenue as/when each obligation is satisfied.

**SaaS patterns**

| Arrangement | Recognition |
|-------------|-------------|
| Monthly subscription | Ratably as service is delivered (each month). |
| Annual prepaid (e.g. €12,000 upfront) | €1,000/mo recognized; remainder sits in **deferred revenue (2300)**. |
| Multi-element (license + implementation + support) | Allocate price across distinct obligations by standalone selling price; recognize each on its own pattern. |
| Setup/onboarding fee | Defer and recognize over the period it relates to (often the contract/expected-life) unless it's a distinct obligation delivered upfront. |
| Usage/consumption | Recognize as usage occurs. |

### Deferred-revenue roll-forward (must tie to the balance sheet and §3)

```
Beginning deferred revenue        80,000
+ Billings (new + renewals)       30,000
− Revenue recognized this period  (26,000)
= Ending deferred revenue         84,000
```

> Multi-element allocation, contract modifications, variable consideration, and capitalized contract costs (ASC 340-40) involve **judgment** — get auditor/accountant sign-off before relying on the policy externally.

---

## 7. Tax Compliance Checklists

> **All rates/thresholds: "verify before use."** Jurisdiction-specific; they change. This is not tax advice — confirm with a qualified advisor and file via the official authority. For 27-country EU depth, use `eu-tax-accounting`.

### 7a. EU VAT

| Scenario | Treatment |
|----------|-----------|
| B2B, same EU country | Charge local VAT. |
| B2B, cross-border EU (valid buyer VAT ID) | **Reverse charge** — 0% on the invoice; buyer self-accounts. Validate the ID via **VIES**; note "reverse charge" on the invoice. |
| B2C goods/services within EU | Charge the **destination** country's rate; report via **OSS** (or IOSS for ≤ €150 imported goods). |
| **B2C digital services** (TBE: telecom, broadcasting, e-services) | Taxed where the **customer** is; charge that country's rate; report via OSS. (No small intra-EU threshold for cross-border digital B2C beyond the €10k pan-EU micro-threshold for goods+TBE.) |
| Sale to non-EU customer | Often outside EU VAT — **but the destination country's VAT/GST/sales tax may apply and may require local registration** (see 7d). Don't assume "no tax." |

> **OSS / IOSS:** register in one EU country to report all EU B2C (OSS) / low-value imports (IOSS), avoiding many separate registrations. Standard rates change — verify each at the national tax authority (links in `eu-tax-accounting`).

| Country (standard VAT, **as of Jun 2026 — verify**) | Rate | Official source |
|------|------|-----------------|
| Luxembourg | 17% | guichet.public.lu / AED |
| Germany | 19% | bzst.de |
| France | 20% | impots.gouv.fr |
| Netherlands | 21% | belastingdienst.nl |
| Spain | 21% | agenciatributaria.es |
| Italy | 22% | agenziaentrate.gov.it |
| Ireland | 23% | revenue.ie |

EU-wide consolidated/standard-rate list: European Commission *Taxes in Europe Database* / VAT rates page. **Re-verify before invoicing.**

### 7b. EU e-invoicing & digital reporting (ViDA) — 2026 readiness

Structured (machine-readable) e-invoicing and near-real-time digital reporting are being mandated unevenly across the EU and under **VAT in the Digital Age (ViDA)**, which also extends **deemed-supplier** VAT rules to platforms/marketplaces. Indicative B2B timeline (**verify per country**): **Germany** Jan 2025 (must *receive*), issuance phasing in ~2027–2028; **Belgium** Jan 2026 (B2B); **France** Sep 2026 (must *receive*), Sep 2027 (must *issue*, large/mid first). Use formats like EN 16931 / Peppol BIS / Factur-X (ZUGFeRD) / FatturaPA (Italy, already live). Action: confirm dates and formats per country in `eu-tax-accounting`, and ensure your billing tool can emit a compliant structured invoice.

### 7c. UK VAT (post-Brexit, separate from EU)

- Standard rate **20%** (verify at gov.uk); registration threshold historically **£90k** taxable turnover (verify current figure).
- **Making Tax Digital (MTD):** VAT returns must be filed via MTD-compatible software with digital record-keeping.
- B2B sales to UK from abroad and low-value imports have specific rules; UK is **not** in EU OSS — UK VAT is handled separately.

### 7d. US sales tax / SaaS nexus

- The US has **no VAT**; sales tax is **state (and local)**, ~45 states + DC, each with its own rules and rates.
- **Economic nexus** (post-*Wayfair*): you can owe collection in a state with no physical presence once you cross its sales/transaction threshold — a common one is **$100,000 in sales or 200 transactions/yr**, but **thresholds, the "200 transactions" prong, and whether SaaS is taxable all vary by state — verify each**. (Many states dropped the transaction count; some never tax SaaS; a few always do.)
- Steps: track sales by state → monitor each threshold → register where you have nexus → collect at the correct combined (state+county+city+district) rate → file/remit on each state's cadence. Tools: Stripe Tax, Avalara, TaxJar, Anrok automate calc/registration/filing.
- **Marketplace facilitator laws:** if you sell *through* a marketplace (Amazon/App Store/etc.), the **platform** often collects and remits sales tax on your behalf — but you may still have filing obligations; confirm per state.

### 7e. Payroll tax

- Withhold and remit employee income tax + employee/employer **social/payroll contributions** on each pay run; deposit on the statutory schedule (penalties for late deposits are steep).
- US: federal income tax withholding + FICA (Social Security/Medicare, employee & employer) + FUTA + state withholding/SUTA; file (e.g.) Form 941 quarterly and W-2s annually — **verify current forms/limits**.
- EU/UK: employer social security + PAYE-type withholding vary by country; see `eu-tax-accounting` per state.
- **Worker classification (employee vs contractor) is a high-risk audit area** — misclassification creates back-tax and penalty exposure; get advice.

### 7f. Corporate income tax

- Tax **expense** (accrual, on the P&L) differs from tax **paid** (cash); most jurisdictions require **estimated/instalment** payments during the year — model these in the cash forecast (§2).
- Track **deferred tax** (timing differences) and any usable **loss carryforwards**; R&D credits/incentives may apply.
- EU corporate rates range widely (e.g. Ireland 12.5%, Germany ~30% effective) — see `eu-tax-accounting`; **verify** before relying.

### 7g. Vendor / contractor information reporting

- **US 1099:** generally issue **1099-NEC** for ≥ **$2,000/yr** (raised from $600 for tax years beginning after 2025; may be inflation-adjusted from 2027) paid to US non-corporate contractors (collect a **W-9** before paying); **1099-K** is issued by payment processors/marketplaces; **verify the current dollar threshold** (it has been changing). Foreign contractors: collect **W-8BEN/W-8BEN-E** instead.
- **EU/UK:** equivalents include **DAC7** (platform reporting of seller income) and local contractor-reporting/withholding rules — confirm per country.

---

## 8. Budget vs Actual (variance analysis)

Both the **Variance** and **% Var** columns below use a single **favorable(+) / unfavorable(−)** convention — the *raw* arithmetic delta is shown separately so the two columns never disagree on sign:

| Category | Budget | Actual | Raw Δ (Actual−Budget) | Variance (F/U) | % Var (F/U) | Flag |
|----------|--------|--------|-----------------------|----------------|-------------|------|
| Revenue | 100,000 | 95,000 | −5,000 | −5,000 (unfavorable) | −5% | Review |
| COGS | 25,000 | 23,000 | −2,000 | +2,000 (favorable) | +8% | OK |
| Marketing | 30,000 | 38,000 | +8,000 | −8,000 (unfavorable) | −27% | Alert |
| R&D | 40,000 | 41,000 | +1,000 | −1,000 (slightly unfavorable) | −2.5% | OK |

> **Sign convention (the one rule that prevents misreads):** for a **revenue/income** line, actual *above* budget is favorable; for a **cost/expense** line, actual *below* budget is favorable. So COGS coming in €2,000 under budget is **+€2,000 favorable** even though the raw `Actual−Budget` is −€2,000 — that is why the "Raw Δ" and "Variance (F/U)" columns can have opposite signs for cost lines. Pick the F/U convention for *both* the amount and the % and apply it to every row so the table can't be misread. `% Var (F/U)` = |Actual−Budget| / Budget, signed favorable/unfavorable.

**Rules**
- Flag variances > **10%** for review, > **20%** for action.
- Always explain **WHY** (price vs volume vs timing), not just the delta.
- Distinguish **timing** variances (will reverse) from **permanent** ones (reforecast).
- Reforecast at least quarterly off actuals; feed the result back into the cash forecast (§2) and the MRR schedule (§3).

---

## affiliate-marketing
Category: growth
Description: Affiliate/partner program design, commission economics, fraud-resistant tracking & attribution, compliance (FTC/GDPR/CPRA/tax/KYC), recruitment, and payout ops. Use when launching an affiliate program, designing commission terms, building click/conversion tracking, reviewing affiliate fraud, vetting payouts, or recruiting partners.
Features:
  - Affiliate program structure design
  - Commission model optimization (CPA, CPS, tiered)
  - Partner recruitment and onboarding
  - Tracking pixel and attribution setup
  - Affiliate content and creative guidelines
  - Performance reporting and payout automation
Use Cases:
  - Launch an affiliate program from scratch
  - Design a tiered commission structure
  - Set up affiliate tracking with proper attribution
  - Recruit and onboard the first 50 affiliates

# Affiliate Marketing

## Workflow

### 1. Program Structure

**In-house vs network:**

| Factor | In-house | Network (Awin, Impact, CJ, etc.) |
|--------|----------|-----------------------------------|
| Setup cost | Higher (build/integrate tracking) | Lower (platform onboarding fee) |
| Ongoing cost | SaaS tracker fee + payment ops + fraud ops + tax/1099 ops + eng maintenance | Network override (commonly ~20-30% on top of commission) + per-payout fees |
| Control | Full | Limited by platform rules/TOS |
| Recruitment | You do it all | Access to affiliate marketplace |
| Tracking | Custom or SaaS (Rewardful, PartnerStack, FirstPromoter, Tolt) | Built-in |
| Best for | SaaS, high-value products, brand control | E-commerce, consumer products, fast volume recruitment |

**"In-house = free" is a myth.** Even self-hosting, you pay: SaaS tracker subscription (or build/maintain a click table), payment rails (PayPal/Wise/Tipalti fees, FX, reversals), fraud review labor, tax compliance (W-9/W-8BEN collection, 1099-NEC/1042-S filing), sanctions screening, and ongoing engineering. Budget 3-8% of affiliate GMV for ops on top of commissions; networks bundle most of this into their override.

**Recommendation:** Start in-house with a SaaS tracker (Rewardful, PartnerStack, FirstPromoter, Tolt) so you keep first-party data and brand control. Add a network only when you need volume recruitment in a marketplace and can absorb the override. Prices/overrides change: verify current network rates at awin.com / impact.com and tracker pricing on each vendor's site (as of Jul 2026); note ShareASale was folded into Awin and its platform closed in late 2025.

### 2. Commission Models

| Model | Structure | Best for | Example |
|-------|-----------|----------|---------|
| CPA (Cost Per Acquisition) | Flat fee per signup/sale | SaaS free trials, lead gen | $50 per paid signup |
| CPS (Cost Per Sale) | % of sale value | E-commerce, variable pricing | 20% of first purchase |
| Recurring | % of subscription revenue | SaaS with monthly billing | 20% for a defined window (see below) |
| Tiered | Increasing % at volume thresholds | Motivating top performers | 20% (1-10), 25% (11-50), 30% (50+) |
| Hybrid | Base CPA + recurring bonus | Balanced motivation | $25 CPA + 10% recurring |

**Setting commission rates (margin/LTV-anchored, not a fixed rule):**
- Compute blended CAC and gross-margin LTV per segment. Cap total affiliate payout (CPA + lifetime recurring) so it stays a fraction of contribution margin — a common target is ≤25-40% of gross-margin LTV, but the right number depends entirely on your margins and competitive landscape.
- Recurring duration is a business choice, not a universal "12 months." Pick the window from margin and partner type:

| Partner / product | Typical recurring term | Why |
|------|------|------|
| SMB SaaS, thin margin | First 12 months | Caps liability where churn + payback risk is high |
| High-margin SaaS, sticky product | 24 months or lifetime | High LTV/long retention justifies sharing more; lifetime is a recruiting magnet (e.g., many infra/dev tools) |
| Ecommerce | One-time % of first order (sometimes 30-day repeat) | No subscription to share |
| High-ticket / enterprise | Flat CPA or % of first contract, sometimes Y1 only | Long sales cycle, large deal size, finance prefers a fixed liability |
| Agency / reseller | Margin share or revenue share for life of the account | They own the relationship and support |
| Influencer / large creator | Higher % or flat fee + bonus, often negotiated per-deal | Reach commands a premium; negotiate per partner |

- Trade-off to state explicitly: lifetime/long terms maximize recruitment and partner loyalty but create perpetual liability and harder unit-economics forecasting; short windows protect margin but recruit fewer top affiliates. Model both against gross-margin LTV before committing.
- Review rates quarterly using affiliate-sourced cohort LTV, refund/chargeback rate, and payback period vs other channels.

### 3. Tracking Implementation

Track in your database, not just a cookie. The cookie is a pointer to a server-side **click record** that carries the data you need to attribute, de-fraud, and reverse. Never derive a payout directly from a raw cookie value.

**Schema (Postgres):**
```sql
CREATE TABLE affiliates (
  id            BIGSERIAL PRIMARY KEY,
  status        TEXT NOT NULL DEFAULT 'pending', -- pending | active | paused | banned
  payout_hash   BYTEA,            -- hash of payout destination (detect linked accounts)
  created_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE TABLE affiliate_clicks (
  click_id      UUID PRIMARY KEY DEFAULT gen_random_uuid(),
  affiliate_id  BIGINT NOT NULL REFERENCES affiliates(id),
  landing_path  TEXT,
  utm_source    TEXT, utm_medium TEXT, utm_campaign TEXT,
  ip_hash       BYTEA,            -- hash, not raw IP (privacy)
  ua_hash       BYTEA,            -- coarse device/UA fingerprint hash
  consent       BOOLEAN NOT NULL DEFAULT FALSE,  -- analytics/marketing consent at click time
  created_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX ON affiliate_clicks (affiliate_id, created_at);

CREATE TABLE affiliate_conversions (
  id            BIGSERIAL PRIMARY KEY,
  click_id      UUID REFERENCES affiliate_clicks(click_id),
  affiliate_id  BIGINT NOT NULL REFERENCES affiliates(id),
  customer_id   BIGINT NOT NULL,
  event_type    TEXT NOT NULL,          -- 'signup' | 'sale' | 'rebill'
  amount_cents  BIGINT NOT NULL,
  commission_cents BIGINT NOT NULL,
  idempotency_key  TEXT UNIQUE NOT NULL, -- e.g. order_id + event_type
  status        TEXT NOT NULL DEFAULT 'pending', -- pending | locked | approved | paid | reversed | rejected
  locked_until  TIMESTAMPTZ,            -- payout hold (refund/chargeback window)
  created_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);

-- Session fingerprints for fraud joins in §5 (self-referral / IP & device overlap)
CREATE TABLE customer_sessions (
  customer_id   BIGINT NOT NULL,
  ip_hash       BYTEA,
  ua_hash       BYTEA,
  created_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX ON customer_sessions (customer_id);
```

**On click — validate, then store a click record (signed click_id in the cookie):**
```javascript
const crypto = require('crypto');
const SECRET = process.env.AFFILIATE_COOKIE_SECRET; // 32+ random bytes

const sign = (v) => crypto.createHmac('sha256', SECRET).update(v).digest('base64url');

app.get('/ref/:affiliateId', async (req, res) => {
  // 1. Validate the affiliate exists AND is approved/active (not pending, banned, or paused)
  const aff = await getAffiliate(req.params.affiliateId);
  if (!aff || aff.status !== 'active') return res.redirect('/'); // silently drop, no cookie

  // 2. Record the click server-side with fraud + consent signals
  const click = await createClick({
    affiliateId: aff.id,
    landingPath: req.query.lp || '/',
    utm: { source: req.query.utm_source, medium: req.query.utm_medium, campaign: req.query.utm_campaign },
    ipHash: hashIp(req.ip),          // hash; do not store raw IP
    uaHash: hashUa(req.get('user-agent')),
    consent: req.cookies.consent === 'granted', // see Compliance: set strictly-necessary only pre-consent
  });

  // 3. Cookie holds an HMAC-signed click_id, not a guessable affiliate id
  const value = `${click.click_id}.${sign(click.click_id)}`;
  res.cookie('aff_click', value, {
    maxAge: cookieWindowMs(aff),  // per-program window, see table below
    httpOnly: true, secure: true, sameSite: 'lax', path: '/',
  });
  res.redirect(click.landingPath);
});
```

**On conversion — verify signature, enforce window + attribution rules, idempotent, reversible:**
```javascript
app.post('/api/checkout/complete', async (req, res) => {
  const raw = req.cookies.aff_click;
  if (!raw) return res.json({ ok: true }); // organic / direct — no attribution, do not invent one

  const [clickId, sig] = raw.split('.');
  if (!clickId || sig !== sign(clickId)) return res.json({ ok: true }); // tampered cookie

  const click = await getClick(clickId);
  if (!click) return res.json({ ok: true });

  // Attribution-window check (the click, not the cookie, is the source of truth)
  if (Date.now() - click.created_at.getTime() > cookieWindowMs({ id: click.affiliate_id })) {
    return res.json({ ok: true }); // expired
  }

  // Exclusions: existing customers, self-referral, paused affiliate
  const aff = await getAffiliate(click.affiliate_id);
  if (!aff || aff.status !== 'active') return res.json({ ok: true });
  if (await isExistingCustomerBeforeClick(req.user.id, click.created_at)) return res.json({ ok: true });
  if (await isSelfReferral(aff, req.user, click)) return res.json({ ok: true });

  // Idempotent write — survives retries / duplicate webhooks; hold for refund/chargeback window
  await recordConversion({
    clickId,
    affiliateId: aff.id,
    customerId: req.user.id,
    eventType: 'sale',
    amountCents: req.body.amount_cents,
    commissionCents: computeCommission(aff, req.body.amount_cents),
    idempotencyKey: `${req.body.order_id}:sale`, // UNIQUE — duplicate => no-op
    status: 'locked',
    lockedUntil: addDays(new Date(), 30), // do not pay until refund window passes
  });
  res.json({ ok: true });
});
```

**Reversal on refund/chargeback (clawback before payout):**
```javascript
// Stripe webhook: charge.refunded / charge.dispute.created
await db.query(
  `UPDATE affiliate_conversions SET status='reversed'
     WHERE idempotency_key=$1 AND status IN ('locked','approved')`,
  [`${orderId}:sale`]
);
// If already paid, record a negative adjustment against the affiliate's next payout.
```

**Cookie window standards:**

| Product type | Cookie window | Rationale |
|-------------|--------------|-----------|
| SaaS | 30-90 days | Longer consideration cycle |
| E-commerce | 7-30 days | Shorter purchase cycle |
| High-ticket | 90-180 days | Enterprise sales cycle |

**Coupon-code attribution (cookieless fallback):** Map a unique discount code → affiliate. At order time, if a tracked coupon is used, attribute to that affiliate. Resolve conflicts explicitly with the cookie/click (define which wins — usually the **explicitly-entered coupon** since it is a stronger intent signal than a stale cookie). Codes survive cross-device and consent-blocked tracking, so most programs offer both a link and a code.

**Attribution rules:**
- **Last click wins** — standard and simplest. The most recent *valid* affiliate click within the window gets credit.
- **First click wins** — rewards discovery (Amazon historically used variants of this). Keep the earliest valid click; later clicks don't overwrite.
- **Linear / multi-touch split** — complex and rarely worth it for affiliate; avoid unless you have multiple affiliates per journey and a reason to split.
- **Direct traffic does NOT erase a valid affiliate click.** A user clicking an affiliate link and later returning via direct/bookmark/branded-search should still convert to the affiliate while the click is within the window — that is the normal, fair behavior. Wiping it under-credits affiliates and is *not* an anti-fraud measure. (Self-referral fraud is handled by the `isSelfReferral` check and the fraud rules in §5, not by nuking direct traffic.) If you intentionally run *paid-search-last-touch* rules to stop affiliates poaching your own branded-search traffic, document that policy in the affiliate agreement — don't bury it in code.

### 4. Partner Recruitment

**Ideal affiliate profiles:**

| Type | Characteristics | Approach |
|------|----------------|----------|
| Content creators | Blog/YouTube in your niche | Outreach with free product + custom commission |
| Review sites | G2, Capterra, niche review blogs | Ensure listing, offer affiliate tracking |
| Influencers | Social following in target audience | Custom landing page + higher commission |
| Existing customers | Happy users with audience | In-app referral prompt + affiliate upgrade option |
| Agencies | Serve your target market | Reseller/referral hybrid program |

**Recruitment outreach template:**
```
Subject: Partner with [Product] — [X]% commission

Hi [Name],

I've been following your content on [specific topic] — [genuine compliment].

We're building [Product], which helps [audience] with [value prop].
I think it'd be a natural fit for your audience.

Our affiliate program:
- [X]% recurring commission (or flat $X per signup)
- [X]-day cookie window
- Dedicated affiliate dashboard
- Custom landing pages and creatives

Interested in trying it out? Happy to set you up with a free account
and walk through the program.

[Name]
```

### 5. Compliance

> This is operational guidance, not legal advice. Affiliate programs move money and touch privacy, tax, and advertising law across jurisdictions — have counsel review your agreement, disclosures, and data flows. Rules change; verify against the linked primary sources.

**FTC disclosure (US) — Endorsement Guides, updated 2023 and actively enforced through 2026:**
- Affiliates MUST disclose a material connection clearly and conspicuously, *before/near* the link, in the same medium (in-video for video, in-stream for audio), not only in a description or "link in bio."
- A bare `#ad`/`#sponsored` can be sufficient if unavoidable; vague terms like `#collab`, `#sp`, `#ambassador`, or `#partner` are not. Platform "paid partnership" toggles do not replace a clear disclosure.
- The brand can be liable for affiliates' deceptive claims. The FTC's 2024 Consumer Reviews and Testimonials Rule (effective October 2024) allows civil penalties for fake/incentivized reviews and undisclosed insider endorsements; your contract must require truthful, substantiated claims and ban fake reviews and "review gating." Primary source: FTC Endorsement Guides + the Rule on Consumer Reviews and Testimonials (ftc.gov).
- Put disclosure obligations, an approved-claims list, and audit rights in the affiliate agreement, and monitor (don't just promise to monitor — the FTC expects active monitoring).

**Privacy & consent (table stakes by 2026 — do not ship cookie tracking without this):**
- **EU/UK (GDPR + ePrivacy):** affiliate/analytics cookies are not "strictly necessary," so you need prior opt-in consent before setting them. Pre-consent, set only a strictly-necessary cookie; gate the `/ref` click cookie and any device hashing on `consent === granted`. Honor IAB TCF signals if you use a CMP. Use coupon-code attribution as the cookieless fallback for non-consenting EU users.
- **US (CPRA/CCPA + state laws):** offer a "Do Not Sell or Share My Personal Information" / opt-out, honor Global Privacy Control (GPC) browser signals, and disclose affiliate tracking in your privacy policy. Sharing click data with networks can count as a "sale/share."
- Store IP/UA as salted hashes, set a data-retention limit on click logs, and write affiliate data sharing into your DPA / privacy policy. Verify current obligations at gdpr.eu, ico.org.uk, and oag.ca.gov/privacy (as of Jun 2026).

**Advertising-channel & trademark rules (protect brand + avoid account bans):**
- Email: affiliates emailing on your behalf must comply with **CAN-SPAM** (US), **CASL** (Canada — express consent + identification), and GDPR/PECR (EU/UK). Ban purchased lists and require working unsubscribe + sender identification in the agreement.
- Paid search: forbid bidding on your brand/trademark terms and trademark + "coupon/discount/promo" combos, and ban direct-linking / typosquatting / fake domains. Forbid running ads that impersonate you.
- Platform policies: Google/Meta/TikTok/Amazon Associates each restrict how affiliate links and incentivized content run — affiliates violating these can get your assets flagged. Reference them in the agreement.

**Tax, KYC & sanctions (before you pay anyone):**
- Collect tax forms before first payout: **W-9** (US persons) / **W-8BEN(-E)** (non-US). File **1099-NEC** for US payees over the IRS threshold (verify the current-year threshold at irs.gov) and **1042-S** for applicable foreign payees; withhold where required.
- KYC/identity verification on affiliates (especially high-payout) to prevent payout fraud and money laundering; many payout providers (Tipalti, Trolley/PayPal, Wise) bundle this.
- **Sanctions screening:** screen affiliates and payout destinations against OFAC SDN and equivalent EU/UK lists; block payouts to sanctioned persons/countries. Bake "we may withhold payment for legal/sanctions/fraud reasons" into the agreement.
- VAT/GST: in some jurisdictions affiliate commission is a taxable supply — clarify who issues invoices and whether commissions are inclusive/exclusive of VAT.

**Affiliate agreement must-have clauses:** disclosure & truthful-claims obligations + audit rights; prohibited methods (brand bidding, spam, cookie stuffing, self-referral, incentivized/fake reviews, adware); payout terms, hold/lock period, clawback on refund/chargeback/fraud; consent/data-handling obligations; right to withhold for legal/sanctions/fraud; termination + survival of clawback.

**Fraud detection — run these as scheduled reviews, hold suspicious conversions before payout** (uses the §3 schema, including `customer_sessions` and the hashed `affiliates.payout_hash` of each affiliate's payout destination):

```sql
-- 1. Self-referral / IP & device overlap (click and conversion share fingerprint)
SELECT cv.affiliate_id, cv.customer_id
FROM affiliate_conversions cv
JOIN affiliate_clicks ck ON ck.click_id = cv.click_id
JOIN customer_sessions cs ON cs.customer_id = cv.customer_id
WHERE ck.ip_hash = cs.ip_hash OR ck.ua_hash = cs.ua_hash;

-- 2. Cookie stuffing / forced clicks: huge click volume, near-zero conversion, sub-second dwell
SELECT affiliate_id,
       COUNT(*) AS clicks,
       AVG(EXTRACT(EPOCH FROM (first_conv.created_at - ck.created_at))) AS avg_dwell_s
FROM affiliate_clicks ck
LEFT JOIN LATERAL (
  SELECT created_at FROM affiliate_conversions c
  WHERE c.click_id = ck.click_id ORDER BY created_at LIMIT 1
) first_conv ON true
WHERE ck.created_at > now() - interval '7 days'
GROUP BY affiliate_id
HAVING COUNT(*) > 5000
   AND COUNT(*) FILTER (WHERE first_conv.created_at IS NOT NULL)::float / COUNT(*) < 0.001;

-- 3. Abnormal conversion rate (suspiciously high CVR vs program median)
WITH per_aff AS (
  SELECT a.id AS affiliate_id,
         COUNT(DISTINCT ck.click_id) AS clicks,
         COUNT(DISTINCT cv.id)       AS conversions
  FROM affiliates a
  LEFT JOIN affiliate_clicks ck      ON ck.affiliate_id = a.id
  LEFT JOIN affiliate_conversions cv ON cv.affiliate_id = a.id
  WHERE ck.created_at > now() - interval '30 days'
  GROUP BY a.id
)
SELECT affiliate_id, clicks, conversions,
       ROUND(conversions::numeric / NULLIF(clicks,0), 4) AS cvr
FROM per_aff
WHERE clicks > 100 AND conversions::numeric / NULLIF(clicks,0) > 0.20  -- tune to your niche
ORDER BY cvr DESC;

-- 4. High refund/chargeback rate (low-quality or fraudulent traffic)
SELECT affiliate_id,
       COUNT(*) AS conversions,
       COUNT(*) FILTER (WHERE status = 'reversed') AS reversed,
       ROUND(COUNT(*) FILTER (WHERE status='reversed')::numeric / COUNT(*), 3) AS reversal_rate
FROM affiliate_conversions
WHERE created_at > now() - interval '90 days'
GROUP BY affiliate_id
HAVING COUNT(*) >= 10
   AND COUNT(*) FILTER (WHERE status='reversed')::numeric / COUNT(*) > 0.10
ORDER BY reversal_rate DESC;

-- 5. Duplicate / linked accounts (same payout destination across "different" affiliates)
SELECT payout_hash, array_agg(id) AS affiliate_ids, COUNT(*)
FROM affiliates
GROUP BY payout_hash
HAVING COUNT(*) > 1;
```

Additional non-SQL checks: brand-bidding violations (monitor paid-search SERPs for your trademark via a rank/ad monitor), minimum click→conversion dwell (reject sub-second), and a manual quarterly review of the top affiliates by revenue. Default-hold (`status='locked'`) all conversions through the refund window so fraudulent ones can be reversed before any payout.

### 6. Performance Optimization

**Monthly affiliate dashboard:**

| Metric | Calculate | Benchmark |
|--------|-----------|-----------|
| Active affiliates | Affiliates with ≥1 conversion/month | 10-20% of total |
| Revenue per affiliate | Total affiliate revenue / Active affiliates | Track trend |
| Conversion rate | Conversions / Clicks | 2-5% (depends on niche) |
| EPC (Earnings Per Click) | Total commissions / Total clicks | $0.50-2.00 |
| Average commission | Total paid / Total conversions | Track vs CAC |
| Affiliate-sourced % | Affiliate revenue / Total revenue | 10-30% target |

**Top performer strategy:**
- Identify top 10% of affiliates by revenue
- Offer exclusive commission rates (+5-10%)
- Provide early access to new features for content
- Quarterly check-in call with affiliate manager
- Custom creatives and co-branded landing pages

---

## ai-agent-building
Category: dev
Description: Build production AI agents — LangGraph state machines, CrewAI teams, tool design, memory, RAG, MCP, multi-agent orchestration, evals, cost control, and safety. Use when building LangGraph/CrewAI agents, designing or validating tools, wiring RAG or MCP, adding human-in-the-loop, or running agent evals and safety reviews.
Features:
  - CrewAI agent and task configuration
  - LangGraph stateful workflow patterns
  - Tool use and function calling patterns
  - Memory systems: short-term, long-term, episodic
  - Multi-agent orchestration and delegation
  - Production deployment with observability
Use Cases:
  - Build a multi-agent research pipeline
  - Create an agent with persistent memory
  - Orchestrate agents with LangGraph workflows
  - Deploy agents to production with monitoring

# AI Agent Building

## Reference guide

Read only the references needed for the current request:

- **Agent Architecture Fundamentals**: [references/agent-architecture-fundamentals.md](references/agent-architecture-fundamentals.md)
- **LangGraph: State Machine Agents**: [references/langgraph-state-machine-agents.md](references/langgraph-state-machine-agents.md)
- **CrewAI: Multi-Agent Teams**: [references/crewai-multi-agent-teams.md](references/crewai-multi-agent-teams.md)
- **Tool Design: Best Practices**: [references/tool-design-best-practices.md](references/tool-design-best-practices.md)
- **Memory Patterns**: [references/memory-patterns.md](references/memory-patterns.md)
- **RAG Pipeline: Production Patterns**: [references/rag-pipeline-production-patterns.md](references/rag-pipeline-production-patterns.md)
- **Multi-Agent Patterns**: [references/multi-agent-patterns.md](references/multi-agent-patterns.md)
- **Production Concerns**: [references/production-concerns.md](references/production-concerns.md)
- **Modern Agent Surfaces (2025-2026)**: [references/modern-agent-surfaces-2025-2026.md](references/modern-agent-surfaces-2025-2026.md)
- **Safety: Prompt Injection Defense**: [references/safety-prompt-injection-defense.md](references/safety-prompt-injection-defense.md)
- **Evaluation**: [references/evaluation.md](references/evaluation.md)
- **Checklist: Production Agent**: [references/checklist-production-agent.md](references/checklist-production-agent.md)
- **MCP (Model Context Protocol) Integration**: [references/mcp-model-context-protocol-integration.md](references/mcp-model-context-protocol-integration.md)
- **Deployment: Containerized Agent**: [references/deployment-containerized-agent.md](references/deployment-containerized-agent.md)
- **Cost Control**: [references/cost-control.md](references/cost-control.md)

### Resource: references/agent-architecture-fundamentals.md

## Agent Architecture Fundamentals

An AI agent is an LLM that can take actions. That's it. Everything else is engineering around that core loop:

```
Observe → Think → Act → Observe → Think → Act → ...
```

The complexity comes from: which actions? how to recover from failures? how to know when to stop? how to not bankrupt you on API calls?

---

### Resource: references/checklist-production-agent.md

## Checklist: Production Agent

- [ ] Tools have clear descriptions, input validation, and error handling
- [ ] Timeouts on all tool calls and LLM invocations
- [ ] Cost tracking per conversation/user
- [ ] Fallback models configured
- [ ] Streaming for user-facing responses
- [ ] Conversation memory with size limits
- [ ] Prompt injection defense (input sanitization)
- [ ] Output validation (no system prompt leaks)
- [ ] Human-in-the-loop for high-stakes actions
- [ ] Checkpointing for long-running workflows
- [ ] Evaluation suite with regression tests
- [ ] Token usage monitoring and alerts
- [ ] Rate limiting per user
- [ ] Logging of all tool calls and responses
- [ ] Graceful degradation when tools fail

---

### Resource: references/cost-control.md

## Cost Control

```python
# Cost-aware model routing — use cheap models when possible
from datetime import datetime, timezone
from langchain_openai import ChatOpenAI

class BudgetExceededError(Exception):
    pass

# Prices in comments are USD/1M input tokens, list as of Jul 2026; verify before relying on them.
# gpt-5-family models reject temperature; steer with reasoning effort instead.
MODELS = {
    "fast": ChatOpenAI(model="gpt-5.4-nano"),                        # cheapest tier: classification, routing
    "smart": ChatOpenAI(model="gpt-5.5"),                            # ~$5/1M in, general work
    "reasoning": ChatOpenAI(model="gpt-5.5", reasoning_effort="high"),  # multi-step logic/math
}

def select_model(task_type: str, input_length: int) -> str:
    """Route to cheapest model that can handle the task."""
    if task_type == "classification" or input_length < 500:
        return "fast"
    if task_type in ("code_generation", "complex_reasoning"):
        return "reasoning"
    return "smart"

# Budget enforcement
class BudgetTracker:
    def __init__(self, daily_limit_usd: float = 10.0):
        self.daily_limit = daily_limit_usd
        self.spent_today = 0.0
        self.last_reset = datetime.now(timezone.utc).date()

    def check_budget(self, estimated_cost: float) -> bool:
        if datetime.now(timezone.utc).date() > self.last_reset:
            self.spent_today = 0.0
            self.last_reset = datetime.now(timezone.utc).date()
        if self.spent_today + estimated_cost > self.daily_limit:
            raise BudgetExceededError(f"Daily budget ${self.daily_limit} exceeded")
        return True

    def record_spend(self, cost: float):
        self.spent_today += cost
```

### Resource: references/crewai-multi-agent-teams.md

## CrewAI: Multi-Agent Teams

```python
# pip install crewai crewai-tools
from crewai import Agent, Task, Crew, Process
from crewai_tools import SerperDevTool, ScrapeWebsiteTool

# Define specialized agents
researcher = Agent(
    role="Senior Research Analyst",
    goal="Find comprehensive, accurate information about the given topic",
    backstory="You're a seasoned researcher with 15 years of experience in market analysis.",
    tools=[SerperDevTool(), ScrapeWebsiteTool()],
    verbose=True,
    allow_delegation=False,
    llm="gpt-5.5",
)

writer = Agent(
    role="Technical Writer",
    goal="Create clear, engaging content based on research findings",
    backstory="You're a technical writer who excels at making complex topics accessible.",
    verbose=True,
    llm="gpt-5.5",
)

editor = Agent(
    role="Editor",
    goal="Review and polish the content for accuracy, clarity, and engagement",
    backstory="You're a meticulous editor with an eye for detail and factual accuracy.",
    verbose=True,
    llm="gpt-5.5",
)

# Define tasks
research_task = Task(
    description="Research the current state of {topic}. Find key trends, statistics, and expert opinions.",
    expected_output="A comprehensive research brief with key findings, statistics, and sources.",
    agent=researcher,
)

writing_task = Task(
    description="Write a 1500-word article based on the research brief.",
    expected_output="A well-structured article with introduction, key sections, and conclusion.",
    agent=writer,
    context=[research_task],  # Uses output from research
)

editing_task = Task(
    description="Edit the article for clarity, accuracy, and engagement. Fix any factual errors.",
    expected_output="A polished, publication-ready article.",
    agent=editor,
    context=[writing_task],
)

# Assemble crew
crew = Crew(
    agents=[researcher, writer, editor],
    tasks=[research_task, writing_task, editing_task],
    process=Process.sequential,  # or Process.hierarchical with a manager
    verbose=True,
)

result = crew.kickoff(inputs={"topic": "AI agents in production"})
```

---

### Resource: references/deployment-containerized-agent.md

## Deployment: Containerized Agent

```dockerfile
# Dockerfile — production agent with health checks
FROM python:3.12-slim AS base

RUN pip install --no-cache-dir langgraph langchain-openai redis uvicorn fastapi

WORKDIR /app
COPY . .

# Non-root user
RUN useradd -m agent && chown -R agent:agent /app
USER agent

# python:3.12-slim has no curl — use a stdlib check (no extra packages, no shell deps)
HEALTHCHECK --interval=30s --timeout=5s --retries=3 \
  CMD ["python", "-c", "import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://localhost:8000/health', timeout=4).status==200 else 1)"]

EXPOSE 8000
CMD ["uvicorn", "server:app", "--host", "0.0.0.0", "--port", "8000"]
```

```python
# server.py — FastAPI wrapper with streaming, cost tracking, rate limiting
import json, time, tiktoken
from collections import defaultdict
from fastapi import FastAPI, Request, HTTPException
from fastapi.responses import StreamingResponse
from langchain_core.messages import HumanMessage

from my_agent import agent  # your compiled LangGraph app (see "Basic Agent" above)

MODEL = "gpt-5.5"
PRICE_IN, PRICE_OUT = 5.00, 30.00  # USD/1M tokens for MODEL; keep in sync with CostTracker.PRICES

app = FastAPI()
start_time = time.time()
try:
    enc = tiktoken.encoding_for_model(MODEL)
except KeyError:
    enc = tiktoken.get_encoding("o200k_base")  # fallback for models tiktoken doesn't know yet

# In-memory rate limiter (use Redis in production)
request_counts: dict[str, list[float]] = defaultdict(list)
RATE_LIMIT = 20  # requests per minute

@app.middleware("http")
async def rate_limit(request: Request, call_next):
    api_key = request.headers.get("x-api-key", "anonymous")
    now = time.time()
    request_counts[api_key] = [t for t in request_counts[api_key] if now - t < 60]
    if len(request_counts[api_key]) >= RATE_LIMIT:
        raise HTTPException(429, "Rate limit exceeded")
    request_counts[api_key].append(now)
    return await call_next(request)

@app.post("/chat")
async def chat(request: Request):
    body = await request.json()
    user_msg = body["message"]
    api_key = request.headers.get("x-api-key")

    # Token counting for cost tracking
    input_tokens = len(enc.encode(user_msg))

    async def stream():
        total_output_tokens = 0
        async for event in agent.astream_events(
            {"messages": [HumanMessage(content=user_msg)]},
            version="v2",
        ):
            if event["event"] == "on_chat_model_stream":
                chunk = event["data"]["chunk"].content
                if chunk:
                    total_output_tokens += len(enc.encode(chunk))
                    yield f"data: {json.dumps({'text': chunk})}\n\n"

        # Log cost using the model's own price (see PRICE_IN/PRICE_OUT above).
        # Note: tiktoken counts only the raw text; it does NOT include tool-call
        # args, system prompt, or reasoning tokens — for exact billing read
        # usage_metadata off the final message instead of estimating here.
        cost = (input_tokens * PRICE_IN + total_output_tokens * PRICE_OUT) / 1_000_000
        yield f"data: {json.dumps({'done': True, 'tokens': {'in': input_tokens, 'out': total_output_tokens}, 'cost_usd': round(cost, 6)})}\n\n"

    return StreamingResponse(stream(), media_type="text/event-stream")

@app.get("/health")
async def health():
    return {"status": "ok", "model": MODEL, "uptime": time.time() - start_time}
```

---

### Resource: references/evaluation.md

## Contents

- Evaluation
- LLM-as-Judge
- Regression Testing

## Evaluation

### LLM-as-Judge

```python
from langchain_core.messages import HumanMessage
from langchain_openai import ChatOpenAI
from pydantic import BaseModel, Field

class Judgement(BaseModel):
    accuracy: int = Field(ge=1, le=5, description="Does it match the reference?")
    completeness: int = Field(ge=1, le=5, description="Does it cover all key points?")
    clarity: int = Field(ge=1, le=5, description="Is it well-written and clear?")
    reasoning: str

# Use a strong, separate judge model; structured output removes brittle json.loads parsing.
eval_model = ChatOpenAI(model="gpt-5.5").with_structured_output(Judgement)

EVAL_PROMPT = """Rate the AI response on a 1-5 scale for accuracy, completeness, and clarity.

Question: {question}
Response: {response}
Reference Answer: {reference}"""

async def evaluate_response(question: str, response: str, reference: str) -> Judgement:
    return await eval_model.ainvoke(
        EVAL_PROMPT.format(question=question, response=response, reference=reference)
    )

# Run evaluation suite
async def run_eval_suite(agent, test_cases: list[dict]) -> dict:
    results = []
    for case in test_cases:
        out = await agent.ainvoke({"messages": [HumanMessage(content=case["question"])]})
        answer = out["messages"][-1].content
        score = await evaluate_response(case["question"], answer, case["expected"])
        results.append({"case": case["question"], "score": score})

    n = len(results)
    avg_accuracy = sum(r["score"].accuracy for r in results) / n
    avg_completeness = sum(r["score"].completeness for r in results) / n
    return {"results": results, "avg_accuracy": avg_accuracy, "avg_completeness": avg_completeness}
```

> **Bias note:** an LLM judge favors verbose, confident, same-family answers and is itself promptable. Calibrate against a human-labeled gold set, randomize answer order for pairwise comparisons, and never let a model grade its own output unchecked in CI.

### Regression Testing

```python
# tests/test_agent.py  (pytest-asyncio; `agent` is your compiled app from above)
import pytest
from langchain_core.messages import HumanMessage
from my_agent import agent

REGRESSION_CASES = [
    {
        "input": "What's the refund policy?",
        "must_contain": ["30 days", "full refund"],
        "must_not_contain": ["no refunds"],
    },
    {
        "input": "How do I cancel my subscription?",
        "must_contain": ["settings", "billing"],
        "must_use_tools": ["search_knowledge_base"],
    },
]

@pytest.mark.parametrize("case", REGRESSION_CASES)
async def test_agent_regression(case):
    result = await agent.ainvoke({"messages": [HumanMessage(content=case["input"])]})
    answer = result["messages"][-1].content.lower()

    for phrase in case.get("must_contain", []):
        assert phrase.lower() in answer, f"Missing: {phrase}"

    for phrase in case.get("must_not_contain", []):
        assert phrase.lower() not in answer, f"Should not contain: {phrase}"
```

---

### Resource: references/langgraph-state-machine-agents.md

## Contents

- LangGraph: State Machine Agents
- Basic Agent with Tool Calling
- Human-in-the-Loop with interrupt() and Checkpointing
- TypeScript LangGraph

## LangGraph: State Machine Agents

LangGraph is the production-grade choice for complex agents. It gives you explicit control flow, checkpointing, and human-in-the-loop — things you need in production but that simple chains don't offer.

### Basic Agent with Tool Calling

```python
# pip install langgraph langchain-openai langgraph-checkpoint-sqlite
from typing import Annotated, TypedDict
from langgraph.graph import StateGraph, START, END
from langgraph.graph.message import add_messages
from langgraph.prebuilt import ToolNode
from langchain_openai import ChatOpenAI
from langchain_core.tools import tool

# Define state
class AgentState(TypedDict):
    messages: Annotated[list, add_messages]

# Define tools
@tool
def search_database(query: str) -> str:
    """Search the product database for items matching the query."""
    # Real implementation here
    return f"Found 3 products matching '{query}': Widget A ($10), Widget B ($20), Widget C ($30)"

@tool
def create_order(product_name: str, quantity: int) -> str:
    """Create an order for a product."""
    order_id = f"ORD-{hash(product_name) % 10000:04d}"
    return f"Order {order_id} created: {quantity}x {product_name}"

tools = [search_database, create_order]
model = ChatOpenAI(model="gpt-5.5").bind_tools(tools)  # gpt-5-family models reject temperature; use reasoning effort to steer

# Define nodes
def agent(state: AgentState) -> AgentState:
    response = model.invoke(state["messages"])
    return {"messages": [response]}

def should_continue(state: AgentState) -> str:
    last_message = state["messages"][-1]
    if last_message.tool_calls:
        return "tools"
    return END

# Build graph
graph = StateGraph(AgentState)
graph.add_node("agent", agent)
graph.add_node("tools", ToolNode(tools))

graph.add_edge(START, "agent")
graph.add_conditional_edges("agent", should_continue, {"tools": "tools", END: END})
graph.add_edge("tools", "agent")

app = graph.compile()

# Run
result = app.invoke({
    "messages": [{"role": "user", "content": "Find me a widget under $15 and order 2 of them"}]
})
```

### Human-in-the-Loop with `interrupt()` and Checkpointing

The modern pattern (LangGraph 0.2.x+) uses the `interrupt()` function to pause *inside* a node and `Command(resume=...)` to feed a decision back. The value passed to `Command(resume=...)` becomes the return value of `interrupt()`, so you must **actually check it** before executing the side-effecting tool — never blindly continue into the tool node. Requires a checkpointer and a stable `thread_id`.

```python
from typing import Annotated, TypedDict
from langgraph.graph import StateGraph, START, END
from langgraph.graph.message import add_messages
from langgraph.prebuilt import ToolNode
from langgraph.types import interrupt, Command
from langgraph.checkpoint.sqlite import SqliteSaver  # pip install langgraph-checkpoint-sqlite
# For pure in-memory dev use: from langgraph.checkpoint.memory import InMemorySaver

class AgentState(TypedDict):
    messages: Annotated[list, add_messages]

def agent(state: AgentState) -> AgentState:
    return {"messages": [model.invoke(state["messages"])]}

def route_after_agent(state: AgentState) -> str:
    last = state["messages"][-1]
    if not getattr(last, "tool_calls", None):
        return END
    # High-stakes tools go through approval; everything else runs directly.
    if any(tc["name"] == "create_order" for tc in last.tool_calls):
        return "approval"
    return "tools"

def approval(state: AgentState) -> Command:
    """Pause and surface the pending order to a human. The resumed value is the decision."""
    last = state["messages"][-1]
    order_calls = [tc for tc in last.tool_calls if tc["name"] == "create_order"]

    # interrupt() returns whatever the human passes via Command(resume=...)
    decision = interrupt({
        "action": "approve_order",
        "orders": [tc["args"] for tc in order_calls],
        "prompt": "Approve these orders? Reply {'approved': bool, 'reason': str}",
    })

    if not decision.get("approved"):
        # Reject: feed a tool message back so the agent can apologize / replan.
        # Do NOT fall through to the tools node.
        from langchain_core.messages import ToolMessage
        return Command(
            goto="agent",
            update={"messages": [
                ToolMessage(
                    content=f"Order rejected by human: {decision.get('reason', 'no reason given')}",
                    tool_call_id=tc["id"],
                ) for tc in order_calls
            ]},
        )
    # Approved: now (and only now) proceed to execute the tool.
    return Command(goto="tools")

graph = StateGraph(AgentState)
graph.add_node("agent", agent)
graph.add_node("tools", ToolNode(tools))
graph.add_node("approval", approval)  # returns Command, so its targets are dynamic

graph.add_edge(START, "agent")
graph.add_conditional_edges("agent", route_after_agent,
                            {"tools": "tools", "approval": "approval", END: END})
graph.add_edge("tools", "agent")

# Compile with a checkpointer — required for interrupt/resume.
with SqliteSaver.from_conn_string(":memory:") as checkpointer:
    app = graph.compile(checkpointer=checkpointer)
    config = {"configurable": {"thread_id": "order-123"}}

    # First run stops at interrupt(); the payload appears under "__interrupt__".
    result = app.invoke(
        {"messages": [{"role": "user", "content": "Order 5 Widget As"}]},
        config=config,
    )
    print(result["__interrupt__"])  # show the orders to the human / UI

    # Human decides. Resume by passing the decision into interrupt() via Command(resume=...).
    final = app.invoke(Command(resume={"approved": True}), config=config)
    # To deny instead:  app.invoke(Command(resume={"approved": False, "reason": "over budget"}), config=config)
```

> `interrupt()` replaces the old `interrupt_before=[...]` / `app.invoke(None, config)` resume idiom, which paused *before* a node but did not let you pass or inspect an approval value. Note `SqliteSaver.from_conn_string` is now a context manager; for persistence on disk use a file path instead of `":memory:"`.

### TypeScript LangGraph

```typescript
import { StateGraph, START, END, Annotation } from "@langchain/langgraph";
import { ChatOpenAI } from "@langchain/openai";
import { ToolNode } from "@langchain/langgraph/prebuilt";
import { tool } from "@langchain/core/tools";
import { z } from "zod";
import { BaseMessage, HumanMessage } from "@langchain/core/messages";

// State definition
const AgentState = Annotation.Root({
  messages: Annotation<BaseMessage[]>({
    reducer: (prev, next) => [...prev, ...next],
  }),
});

// Tools
const searchTool = tool(
  async ({ query }) => {
    return `Results for "${query}": Product A, Product B`;
  },
  {
    name: "search",
    description: "Search the product database",
    schema: z.object({ query: z.string() }),
  }
);

const model = new ChatOpenAI({ model: "gpt-5.5" }).bindTools([searchTool]);

// Nodes
async function agent(state: typeof AgentState.State) {
  const response = await model.invoke(state.messages);
  return { messages: [response] };
}

function shouldContinue(state: typeof AgentState.State) {
  const lastMsg = state.messages[state.messages.length - 1];
  if ("tool_calls" in lastMsg && lastMsg.tool_calls?.length) {
    return "tools";
  }
  return END;
}

// Graph
const graph = new StateGraph(AgentState)
  .addNode("agent", agent)
  .addNode("tools", new ToolNode([searchTool]))
  .addEdge(START, "agent")
  .addConditionalEdges("agent", shouldContinue, { tools: "tools", [END]: END })
  .addEdge("tools", "agent");

const app = graph.compile();

const result = await app.invoke({
  messages: [new HumanMessage("Find products related to widgets")],
});
```

---

### Resource: references/mcp-model-context-protocol-integration.md

## Contents

- MCP (Model Context Protocol) Integration
- Building an MCP Server
- Connecting LangGraph to MCP Tools

## MCP (Model Context Protocol) Integration

MCP is the standard for connecting agents to external tools. Instead of hardcoding tool implementations, agents connect to MCP servers that expose tools over a standardized protocol.

### Building an MCP Server

```typescript
// mcp-server.ts — expose tools for any MCP-compatible agent
import { McpServer } from '@modelcontextprotocol/sdk/server/mcp.js';
import { StreamableHTTPServerTransport } from '@modelcontextprotocol/sdk/server/streamableHttp.js';
import { z } from 'zod';
import express from 'express';

const server = new McpServer({ name: 'my-tools', version: '1.0.0' });

// Register tools with Zod-typed parameters (registerTool replaces the deprecated server.tool)
server.registerTool('search_docs', {
  description: 'Search internal documentation by query',
  inputSchema: {
    query: z.string().describe('Search query'),
    limit: z.number().optional().describe('Max results (default 10)'),
  },
}, async ({ query, limit = 10 }) => {
  const results = await searchIndex(query, limit);
  return {
    content: [{ type: 'text', text: JSON.stringify(results, null, 2) }],
  };
});

server.registerTool('create_ticket', {
  description: 'Create a support ticket in Jira',
  inputSchema: {
    title: z.string().describe('Ticket title'),
    priority: z.string().describe('low | medium | high | critical'),
    description: z.string().describe('Detailed description'),
  },
}, async ({ title, priority, description }) => {
  // Validate before acting — agents will pass garbage sometimes
  if (!['low', 'medium', 'high', 'critical'].includes(priority)) {
    throw new Error(`Invalid priority "${priority}". Must be: low, medium, high, critical`);
  }
  const ticket = await jira.createIssue({ summary: title, priority, description });
  return {
    content: [{ type: 'text', text: `Created ticket ${ticket.key}: ${ticket.self}` }],
  };
});

// Streamable HTTP transport (replaces deprecated SSE transport)
const app = express();
app.use(express.json());

app.post('/mcp', async (req, res) => {
  const transport = new StreamableHTTPServerTransport({
    sessionIdGenerator: undefined, // stateless
  });
  await server.connect(transport);
  await transport.handleRequest(req, res);
});

app.listen(3100, () => console.log('MCP server on :3100'));
```

### Connecting LangGraph to MCP Tools

Don't hand-roll an MCP client. Use the official `langchain-mcp-adapters`, which speaks **Streamable HTTP** (the transport the server above exposes at `/mcp`) and returns ready-to-use LangChain tools — handling schema conversion, sessions, and reconnects for you. The deprecated `sse_client` transport will not talk to a `StreamableHTTPServerTransport` server.

```python
# pip install langchain-mcp-adapters langchain langgraph langchain-openai
import asyncio
import os
from langchain_mcp_adapters.client import MultiServerMCPClient
from langchain.agents import create_agent  # replaces deprecated langgraph.prebuilt.create_react_agent
from langchain_openai import ChatOpenAI

async def main():
    client = MultiServerMCPClient({
        "my-tools": {
            "transport": "streamable_http",         # matches the server's /mcp endpoint
            "url": "http://localhost:3100/mcp",
            "headers": {"Authorization": f"Bearer {os.environ['MCP_TOKEN']}"},  # optional auth
        },
        # add more servers here; tools are merged into one list
    })

    tools = await client.get_tools()  # list[BaseTool], names/schemas come from the server
    agent = create_agent(ChatOpenAI(model="gpt-5.5"), tools)

    result = await agent.ainvoke(
        {"messages": [{"role": "user", "content": "Search the docs for CORS config and open a ticket."}]}
    )
    print(result["messages"][-1].content)

asyncio.run(main())
```

`MultiServerMCPClient` is **stateless by default** — each tool call opens a fresh session and tears it down. For tools that need a persistent session (e.g. sampling, server-side state), wrap calls in `async with client.session("my-tools") as session:`. To call a remote MCP server directly from a frontier model without an adapter, use the provider's native MCP tool type (see the OpenAI Responses example above, and the `mcp-client` / `mcp-server-builder` sibling skills).

---

### Resource: references/memory-patterns.md

## Contents

- Memory Patterns
- Conversation Buffer with Sliding Window
- Summary Memory
- Vector Store Memory (Long-term)

## Memory Patterns

### Conversation Buffer with Sliding Window

```python
from langchain_core.messages import trim_messages

# Keep last N messages, but always keep the system message
trimmer = trim_messages(
    max_tokens=4000,
    strategy="last",
    token_counter=model,
    include_system=True,
    allow_partial=False,
)

# In your agent node
def agent(state: AgentState) -> AgentState:
    trimmed = trimmer.invoke(state["messages"])
    response = model.invoke(trimmed)
    return {"messages": [response]}
```

### Summary Memory

```python
from langchain_core.messages import SystemMessage

async def maybe_summarize(state: AgentState) -> AgentState:
    messages = state["messages"]
    if len(messages) < 20:
        return state

    # Summarize older messages, keep recent ones
    old_messages = messages[1:-10]  # Skip system, keep last 10
    recent = messages[-10:]

    summary = await model.ainvoke([
        SystemMessage(content="Summarize this conversation concisely, preserving key facts and decisions:"),
        *old_messages,
    ])

    return {
        "messages": [
            messages[0],  # System message
            SystemMessage(content=f"Previous conversation summary: {summary.content}"),
            *recent,
        ]
    }
```

### Vector Store Memory (Long-term)

```python
# pip install langchain-chroma langchain-openai
from datetime import datetime, timezone
from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma

embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
memory_store = Chroma(
    collection_name="agent_memory",
    embedding_function=embeddings,
    persist_directory="./memory_db",
)

@tool
def recall_memory(query: str) -> str:
    """Search past conversations and learned facts for relevant information."""
    docs = memory_store.similarity_search(query, k=5)
    if not docs:
        return "No relevant memories found."
    return "\n\n".join([
        f"[{doc.metadata.get('timestamp', 'unknown')}] {doc.page_content}"
        for doc in docs
    ])

@tool
def store_memory(fact: str, category: str = "general") -> str:
    """Store an important fact or learning for future reference."""
    memory_store.add_texts(
        texts=[fact],
        metadatas=[{
            "category": category,
            "timestamp": datetime.now(timezone.utc).isoformat(),
        }],
    )
    return f"Stored: {fact}"
```

---

### Resource: references/modern-agent-surfaces-2025-2026.md

## Contents

- Modern Agent Surfaces (2025-2026)
- Anthropic Memory Tool (public beta)
- OpenAI Responses API (March 2025)

## Modern Agent Surfaces (2025-2026)

### Anthropic Memory Tool (public beta)

Lets Claude store and retrieve files across turns so long-running agents don't blow context. Operations: `view`, `create`, `str_replace`, `insert`, `delete`, `rename`. You implement the storage backend (a per-conversation `/memories/` directory on disk or object store) by handling `tool_use` blocks named `"memory"` and returning `tool_result` blocks.

```python
# Still public beta as of Jun 2026 — pass the memory tool + beta header.
# Verify the current tool-type version string and header at:
# https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool

response = client.beta.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=4096,
    betas=["context-management-2025-06-27"],          # current beta flag as of Jun 2026
    tools=[{"type": "memory_20250818", "name": "memory"}],  # confirm latest memory_* version in docs
    messages=conversation,
)
```

Pair with **prompt caching** on a long system prompt so the agent's "personality + memory index" is cached across turns: cached input is billed at ~10% of the base input price (a ~90% discount). Combine with **tool-use context clearing** (same beta header) to drop stale tool results from the window automatically.

### OpenAI Responses API (March 2025)

Stateful successor to Chat Completions: tools, file/web/MCP, reasoning models, and conversation `store: true` for server-held state.

```python
# pip install openai
from openai import OpenAI
client = OpenAI()

resp = client.responses.create(
    model="gpt-5.5",
    input="Summarize the latest issues in repo X and open one for the worst.",
    store=True,
    reasoning={"effort": "medium"},
    tools=[
        {
            "type": "mcp",
            "server_label": "github",
            "server_url": "https://mcp.github.com",  # remote MCP server
            # Reserve "never" for trusted, read-only servers; write actions stay behind approval.
            "require_approval": "always",
        },
    ],
)
print(resp.output_text)
```

The `mcp` tool type lets the model call any remote MCP server (Streamable HTTP) without you proxying every call. See the `mcp-client` skill for client patterns and `mcp-server-builder` for shipping your own.

---

### Resource: references/multi-agent-patterns.md

## Contents

- Multi-Agent Patterns
- Supervisor Pattern

## Multi-Agent Patterns

### Supervisor Pattern

```python
import json
from typing import Annotated, TypedDict
from langgraph.graph import StateGraph, START, END
from langgraph.graph.message import add_messages
from langchain_core.messages import SystemMessage

class SupervisorState(TypedDict):
    messages: Annotated[list, add_messages]
    next_agent: str

from typing import Literal
from pydantic import BaseModel

class Route(BaseModel):
    next: Literal["researcher", "coder", "writer", "FINISH"]

# with_structured_output guarantees a parsed Route — don't json.loads(content),
# which breaks the moment the model wraps JSON in prose or a code fence.
router_model = supervisor_model.with_structured_output(Route)

def supervisor(state: SupervisorState) -> SupervisorState:
    """Route to the appropriate specialist agent."""
    decision = router_model.invoke([
        SystemMessage(content="""You are a supervisor routing tasks to specialists:
- researcher: for finding information
- coder: for writing or reviewing code
- writer: for creating content
Pick the next worker, or FINISH when the task is complete."""),
        *state["messages"],
    ])
    return {"next_agent": decision.next}

def route(state: SupervisorState) -> str:
    return state["next_agent"]

graph = StateGraph(SupervisorState)
graph.add_node("supervisor", supervisor)
graph.add_node("researcher", researcher_agent)
graph.add_node("coder", coder_agent)
graph.add_node("writer", writer_agent)

graph.add_edge(START, "supervisor")
graph.add_conditional_edges("supervisor", route, {
    "researcher": "researcher",
    "coder": "coder",
    "writer": "writer",
    "FINISH": END,
})
# All agents report back to supervisor
for agent in ["researcher", "coder", "writer"]:
    graph.add_edge(agent, "supervisor")

app = graph.compile()
```

---

### Resource: references/production-concerns.md

## Contents

- Production Concerns
- Cost Tracking
- Streaming Responses
- Fallback Models

## Production Concerns

### Cost Tracking

```python
import tiktoken
from contextlib import contextmanager

class CostTracker:
    # USD per 1M tokens (input/output). List prices as of Jul 2026 (these move often);
    # treat as a starting point and re-check the official pricing pages, ideally generating
    # this dict from a dated constants file in CI:
    #   OpenAI:    https://openai.com/api/pricing
    #   Anthropic: https://platform.claude.com/docs/en/about-claude/pricing
    PRICES = {
        "gpt-5.6-sol":      {"input": 5.00, "output": 30.00},  # flagship
        "gpt-5.6-terra":    {"input": 2.50, "output": 15.00},  # balanced
        "gpt-5.6-luna":     {"input": 1.00, "output": 6.00},   # cost-optimized
        "gpt-5.5":          {"input": 5.00, "output": 30.00},
        "gpt-5.4":          {"input": 2.50, "output": 15.00},  # production workhorse
        "gpt-5.1":          {"input": 1.25, "output": 10.00},
        "claude-opus-4-8":   {"input": 5.00, "output": 25.00},
        "claude-sonnet-4-6": {"input": 3.00, "output": 15.00},
        "claude-haiku-4-5":  {"input": 1.00, "output": 5.00},
    }

    def __init__(self):
        self.total_input_tokens = 0
        self.total_output_tokens = 0
        self.total_cost = 0.0
        self.calls = []

    def track(self, model: str, input_tokens: int, output_tokens: int):
        prices = self.PRICES.get(model, {"input": 0, "output": 0})
        cost = (input_tokens * prices["input"] + output_tokens * prices["output"]) / 1_000_000
        self.total_input_tokens += input_tokens
        self.total_output_tokens += output_tokens
        self.total_cost += cost
        self.calls.append({"model": model, "input": input_tokens, "output": output_tokens, "cost": cost})

    def report(self) -> str:
        return (
            f"Total: {len(self.calls)} calls, "
            f"{self.total_input_tokens} input + {self.total_output_tokens} output tokens, "
            f"${self.total_cost:.4f}"
        )
```

### Streaming Responses

```python
# LangGraph streaming (assumes `app` and HumanMessage from the Basic Agent setup above)
from langchain_core.messages import HumanMessage

async for event in app.astream_events(
    {"messages": [HumanMessage(content="Hello")]},
    version="v2",
):
    if event["event"] == "on_chat_model_stream":
        chunk = event["data"]["chunk"]
        print(chunk.content, end="", flush=True)
    elif event["event"] == "on_tool_start":
        print(f"\n[Using tool: {event['name']}]")
```

### Fallback Models

```python
from langchain_openai import ChatOpenAI
from langchain_anthropic import ChatAnthropic

primary = ChatOpenAI(model="gpt-5.5", timeout=30)
fallback = ChatAnthropic(model="claude-sonnet-4-6", timeout=30)

model = primary.with_fallbacks([fallback])
# Automatically tries fallback if primary fails (cross-provider is the point —
# survives a single vendor's outage or rate-limit spike)
```

---

### Resource: references/rag-pipeline-production-patterns.md

## Contents

- RAG Pipeline: Production Patterns
- Chunking Strategies
- Hybrid Search (Vector + Keyword)
- Reranking
- Citation Pattern

## RAG Pipeline: Production Patterns

### Chunking Strategies

```python
from langchain_text_splitters import RecursiveCharacterTextSplitter, Language

# For general documents
splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=200,
    separators=["\n\n", "\n", ". ", " ", ""],
    length_function=len,
)

# For code
code_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.PYTHON,
    chunk_size=1500,
    chunk_overlap=200,
)

# For markdown with structure preservation
markdown_splitter = RecursiveCharacterTextSplitter.from_language(
    language=Language.MARKDOWN,
    chunk_size=1000,
    chunk_overlap=100,
)
```

### Hybrid Search (Vector + Keyword)

```python
from langchain_community.retrievers import BM25Retriever
from langchain.retrievers import EnsembleRetriever

# Vector search (semantic)
vector_retriever = vector_store.as_retriever(search_kwargs={"k": 5})

# Keyword search (BM25)
bm25_retriever = BM25Retriever.from_documents(documents, k=5)

# Combine with weights
hybrid_retriever = EnsembleRetriever(
    retrievers=[vector_retriever, bm25_retriever],
    weights=[0.6, 0.4],  # Favor semantic, but keyword catches exact matches
)
```

### Reranking

```python
from langchain.retrievers import ContextualCompressionRetriever
from langchain_cohere import CohereRerank

# Retrieve broadly, then rerank for precision
reranker = CohereRerank(model="rerank-english-v3.0", top_n=3)
retriever = ContextualCompressionRetriever(
    base_compressor=reranker,
    base_retriever=hybrid_retriever,  # Gets 20 candidates
)

# Usage: retriever.invoke("How do I configure CORS?")
# Returns top 3 most relevant chunks from the initial 20
```

### Citation Pattern

```python
from langchain_core.prompts import ChatPromptTemplate

RAG_PROMPT = ChatPromptTemplate.from_messages([
    ("system", """Answer the question based on the provided context.
Include citations using [1], [2] etc. referencing the source documents.
If the context doesn't contain the answer, say so — don't make things up.

Context:
{context}"""),
    ("human", "{question}"),
])

def format_docs_with_citations(docs):
    formatted = []
    for i, doc in enumerate(docs, 1):
        source = doc.metadata.get("source", "unknown")
        formatted.append(f"[{i}] (Source: {source})\n{doc.page_content}")
    return "\n\n".join(formatted)
```

---

### Resource: references/safety-prompt-injection-defense.md

## Contents

- Safety: Prompt Injection Defense
- Input Validation
- Output Validation

## Safety: Prompt Injection Defense

### Input Validation

```python
import re

def sanitize_user_input(text: str) -> str:
    """Basic prompt injection defense."""
    # Remove common injection patterns
    suspicious_patterns = [
        r"ignore (?:all )?(?:previous |prior |above )?instructions",
        r"you are now",
        r"new instructions:",
        r"system prompt:",
        r"</s>|<\|im_end\|>|<\|endoftext\|>",
    ]
    for pattern in suspicious_patterns:
        if re.search(pattern, text, re.IGNORECASE):
            return "[Input contained suspicious patterns and was filtered]"
    return text
```

### Output Validation

```python
from pydantic import BaseModel, field_validator

class AgentResponse(BaseModel):
    answer: str
    sources: list[str]
    confidence: float

    @field_validator("answer")
    @classmethod
    def no_system_leaks(cls, v: str) -> str:
        forbidden = ["system prompt", "you are an AI", "as an AI language model"]
        for phrase in forbidden:
            if phrase.lower() in v.lower():
                raise ValueError("Response contained forbidden content")
        return v

    @field_validator("confidence")
    @classmethod
    def valid_range(cls, v: float) -> float:
        if not 0 <= v <= 1:
            raise ValueError("Confidence must be between 0 and 1")
        return v
```

---

### Resource: references/tool-design-best-practices.md

## Contents

- Tool Design: Best Practices
- Error Recovery and Timeout Handling
- Tool Design Rules

## Tool Design: Best Practices

### Error Recovery and Timeout Handling

```python
import asyncio
from functools import wraps
from langchain_core.tools import tool

def with_timeout(seconds: int = 30):
    def decorator(func):
        @wraps(func)
        async def wrapper(*args, **kwargs):
            try:
                return await asyncio.wait_for(func(*args, **kwargs), timeout=seconds)
            except asyncio.TimeoutError:
                return f"Error: Tool timed out after {seconds}s. Try a simpler query."
        return wrapper
    return decorator

def with_retry(max_retries: int = 3):
    def decorator(func):
        @wraps(func)
        async def wrapper(*args, **kwargs):
            last_error = None
            for attempt in range(max_retries):
                try:
                    return await func(*args, **kwargs)
                except Exception as e:
                    last_error = e
                    if attempt < max_retries - 1:
                        await asyncio.sleep(2 ** attempt)
            return f"Error after {max_retries} retries: {str(last_error)}"
        return wrapper
    return decorator

@tool
@with_retry(3)
@with_timeout(30)
async def query_database(sql: str) -> str:
    """Run a read-only SELECT against the analytics warehouse and return rows.

    Args:
        sql: A single SELECT statement. No DML/DDL, no multiple statements.
    """
    try:
        validated = validate_readonly_sql(sql, allowed_tables={"orders", "products", "customers"})
    except ValueError as e:
        return f"Error: {e}"

    # Defense in depth: the LLM-facing connection uses a DB role that only has
    # SELECT on the allowed schema (see note below) AND a per-statement timeout.
    rows = await ro_db.execute(validated, timeout_s=10)  # ro_db = read-only-role pool
    if len(rows) > 50:
        return f"Query returned {len(rows)} rows (showing first 20):\n{format_rows(rows[:20])}"
    return format_rows(rows)
```

**Why the old `"DROP" in sql.upper()` blocklist is not production-safe:** substring checks are trivially bypassed (`/*DROP*/`, `dr"||"op`, a column literally named `update_ts`), they still allow stacked statements (`SELECT 1; DELETE ...`), CTE-wrapped writes, `pg_sleep()`-style DoS, schema enumeration via `information_schema`/`pg_catalog`, and cross-tenant reads. **Allowlist with a real SQL parser instead of blocklisting.** Use `sqlglot` to parse to an AST, reject anything that isn't exactly one `SELECT`, and enforce table allowlist + tenant scoping:

```python
# pip install sqlglot
import sqlglot
from sqlglot import exp

def validate_readonly_sql(sql: str, allowed_tables: set[str], tenant_id: str | None = None) -> str:
    statements = sqlglot.parse(sql, read="postgres")
    if len(statements) != 1:
        raise ValueError("Exactly one statement is allowed (no stacked queries).")
    tree = statements[0]

    # 1. Top level must be a pure SELECT (this also rejects INSERT/UPDATE/DELETE/DDL,
    #    and SELECT ... INTO / data-modifying CTEs at the root).
    if not isinstance(tree, exp.Select):
        raise ValueError("Only SELECT statements are allowed.")

    # 2. No write expressions or unsafe constructs anywhere in the tree.
    banned = (exp.Insert, exp.Update, exp.Delete, exp.Drop, exp.Alter,
              exp.Create, exp.Command, exp.Merge, exp.Into, exp.Set)
    if any(node for node in tree.walk() if isinstance(node, banned)):
        raise ValueError("Query contains a forbidden write/DDL operation.")

    # 3. Allowlist every referenced table; block catalog/schema probing.
    for tbl in tree.find_all(exp.Table):
        name = tbl.name.lower()
        if tbl.db and tbl.db.lower() in ("information_schema", "pg_catalog"):
            raise ValueError("System catalog access is not allowed.")
        if name not in allowed_tables:
            raise ValueError(f"Table '{name}' is not allowed.")

    # 4. Force a hard row cap (LLMs forget LIMIT; large scans cost money / leak data).
    if not tree.args.get("limit"):
        tree = tree.limit(1000)

    # 5. (Multi-tenant) inject a tenant filter so the agent can never read other tenants.
    if tenant_id is not None:
        tree = tree.where(exp.condition(f"tenant_id = {sqlglot.exp.Literal.string(tenant_id)}"))

    return tree.sql(dialect="postgres")
```

Layer this with infrastructure controls — the validator is the inner ring, not the only ring:

- **Dedicated read-only role.** Run agent queries on a connection whose Postgres role has `SELECT` only, on a restricted schema/view: `GRANT SELECT ON orders, products, customers TO agent_ro;` and nothing else. Even a parser bypass then cannot write.
- **Statement timeout.** `SET statement_timeout = '10s'` on that role/session to kill `pg_sleep`-style or runaway scans.
- **Prefer views.** Expose curated, pre-joined, already tenant-scoped views (e.g. `agent_orders_v`) and allowlist only those — never base tables.
- **Parameterize the tenant id**; never string-format untrusted values into SQL elsewhere in your app.

### Tool Design Rules

1. **Clear descriptions** — the LLM reads them to decide when to use the tool
2. **Validate inputs** — never trust LLM-generated parameters
3. **Return errors as strings** — don't throw exceptions, let the agent recover
4. **Limit output size** — truncate large results, the context window is precious
5. **Make tools idempotent** where possible — agents retry
6. **Include examples in docstrings** — helps the LLM use tools correctly

---

---

## aleph-cloud-self-deployment
Category: web3
Description: Deploy and operate VMs on Aleph Cloud with the aleph-client CLI — single node or multi-node fleets, Tailscale mesh, HAProxy distribution, backup/recovery, cost control, and security hardening. Use when deploying confidential/persistent VMs (e.g. an OpenClaw agent runtime) on Aleph Cloud, or building an Aleph node fleet.
Features:
  - Aleph Cloud account setup with credits or $ALEPH tokens
  - Programmatic CRN discovery and scoring
  - Non-interactive VM instance creation (bypasses TUI)
  - OpenClaw agent framework installation and configuration
  - Gateway auth with auth-profiles.json (not auth.json)
  - SSH keypair generation and secure key transfer
  - Recursive multi-instance deployment enablement
  - All-in-one deploy.sh and setup-agent.sh scripts
  - Cost planning across all 6 compute tiers
  - Comprehensive troubleshooting for real-world issues

# Aleph Cloud Self-Deployment: VM & Multi-Node Fleet Management

Framework for deploying and managing persistent/confidential VMs on Aleph Cloud's decentralized compute network using the official `aleph-client` CLI, with patterns for running an OpenClaw agent runtime across one or many nodes (Tailscale mesh, HAProxy distribution, pull-based backup, and hardening).

> **Verify before you ship.** Aleph CLI flags, OpenClaw install commands, and pricing change over time. This skill is current as of **Jun 2026**. Note: docs.aleph.cloud's CLI reference now documents a rewritten `aleph-cli` (installed via Homebrew, apt, or cargo) whose syntax differs from the Python `aleph-client` used throughout this skill; this skill targets the Python `aleph-client` (PyPI, v1.9.x). Authoritative sources, used throughout: the Aleph CLI command reference at https://docs.aleph.cloud/devhub/sdks-and-tools/aleph-cli/ (instance subcommands: https://docs.aleph.cloud/devhub/sdks-and-tools/aleph-cli/commands/instance.html), and OpenClaw docs at https://docs.openclaw.ai/. Run `aleph instance create --help` and `aleph pricing instance` to confirm current flags and prices on your machine.
>
> If you installed the rewritten `aleph-cli` from the docs instead of the Python `aleph-client`, translate commands: `aleph pricing instance` -> `aleph instance price` (`--size 2vcpu-4gb`, `--json`); `aleph account address` -> `aleph account show`; `aleph account create --private-key ...` -> `aleph account import <name> --private-key`; `instance create --name X --compute-units N --rootfs-size MIB` -> `instance create X --vcpus/--memory/--disk-size`; `--crn-url`/`--crn-auto-tac` -> `--crn-hash`. Run `aleph instance create --help` to see which client you have.

> **Sibling skills.** This skill focuses on Aleph-specific provisioning and fleet orchestration. For deep, vendor-neutral coverage prefer: `security-hardening` (SSH/firewall/CIS), `monitoring-observability` (metrics, alerting, log pipelines), and `docker-production` (Compose v2, image hygiene). Use those alongside this one rather than duplicating their depth here.

## Safety gate

Before executing commands or changing external systems, confirm scope, credentials, target environment, rollback, and required approval. Pin and verify third-party artifacts; never expose secrets to client code or logs.

## Reference guide

Read only the references needed for the current request:

- **Table of Contents**: [references/table-of-contents.md](references/table-of-contents.md)
- **Infrastructure Planning & Architecture**: [references/infrastructure-planning-architecture.md](references/infrastructure-planning-architecture.md)
- **Quick Start — tested single-VM happy path**: [references/quick-start-tested-single-vm-happy-path.md](references/quick-start-tested-single-vm-happy-path.md)
- **Single Node Deployment Foundation**: [references/single-node-deployment-foundation.md](references/single-node-deployment-foundation.md)
- **Multi-Node Fleet Management**: [references/multi-node-fleet-management.md](references/multi-node-fleet-management.md)
- **Auto-Provisioning Protocol (SRP)**: [references/auto-provisioning-protocol-srp.md](references/auto-provisioning-protocol-srp.md)
- **Inter-VM Communication Networks**: [references/inter-vm-communication-networks.md](references/inter-vm-communication-networks.md)
- **Load Distribution & Orchestration**: [references/load-distribution-orchestration.md](references/load-distribution-orchestration.md)
- **Disaster Recovery & Auto-Recreation**: [references/disaster-recovery-auto-recreation.md](references/disaster-recovery-auto-recreation.md)
- **Emergency Response Procedures**: [references/emergency-response-procedures.md](references/emergency-response-procedures.md)
- **Backup Verification**: [references/backup-verification.md](references/backup-verification.md)
- **Contact Information**: [references/contact-information.md](references/contact-information.md)
- **Post-Incident Procedures**: [references/post-incident-procedures.md](references/post-incident-procedures.md)
- **Cost Optimization Strategies**: [references/cost-optimization-strategies.md](references/cost-optimization-strategies.md)
- **Security Hardening Framework**: [references/security-hardening-framework.md](references/security-hardening-framework.md)
- **Monitoring & Maintenance**: [references/monitoring-maintenance.md](references/monitoring-maintenance.md)

### Resource: references/auto-provisioning-protocol-srp.md

## Contents

- Auto-Provisioning Protocol (SRP)
- Agent Continuity System

## Auto-Provisioning Protocol (SRP)

### Agent Continuity System

**Auto-Provisioning Framework:**
> **What this is.** An OPTIONAL OpenClaw-specific "agent continuity" layer that
> replicates an agent's workspace (`SOUL.md`/`AGENTS.md`/`MEMORY.md`/skills) from the
> primary to workers, so a worker can take over the agent's state. It is independent
> of OpenClaw's own config and only meaningful if you run OpenClaw with such a
> workspace. Skip this whole section if you just need plain VMs. This script runs
> **on the primary node** (it is installed there by `setup_continuous_replication`).

```bash
#!/bin/bash
# auto-provisioning-protocol.sh  — runs ON the primary node.
set -euo pipefail

# SRP Configuration
SRP_VERSION="2.0.0"
REPLICATION_DIR="/opt/openclaw/replication"
FLEET_CONFIG="${FLEET_CONFIG:-/opt/fleet-manager/fleet.json}"   # node-local copy if present
BACKUP_RETENTION_DAYS=30

# Control-plane access for replicate_to_fleet(): the fleet manager listens on this
# node's Tailscale IP and needs the shared key. Both come from the root-owned
# EnvironmentFile that the fleet manager itself uses.
[[ -f /etc/fleet-manager.env ]] && { set -a; . /etc/fleet-manager.env; set +a; }
FLEET_MGR_HOST="${BIND_HOST:-127.0.0.1}"

echo "Auto-Provisioning Protocol v$SRP_VERSION"

initialize_srp() {
    echo "🔬 Initializing Auto-Provisioning Protocol..."

    # Create replication directory structure
    mkdir -p "$REPLICATION_DIR"/{soul,agents,memory,skills,config,logs}

    # Initialize replication manifest
    cat > "$REPLICATION_DIR/manifest.json" << 'MANIFEST'
{
  "srp_version": "2.0.0",
  "initialized": null,
  "last_replication": null,
  "replication_count": 0,
  "source_node": null,
  "target_nodes": [],
  "integrity_hash": null,
  "components": {
    "soul": {
      "path": "SOUL.md",
      "required": true,
      "last_modified": null,
      "hash": null
    },
    "agents": {
      "path": "AGENTS.md",
      "required": true,
      "last_modified": null,
      "hash": null
    },
    "memory": {
      "path": "MEMORY.md",
      "required": false,
      "last_modified": null,
      "hash": null
    },
    "skills": {
      "path": "skills/",
      "required": false,
      "last_modified": null,
      "hash": null
    },
    "user_data": {
      "path": "USER.md",
      "required": false,
      "last_modified": null,
      "hash": null
    }
  }
}
MANIFEST

    local tmpfile=$(mktemp)
    jq '.initialized = now | .source_node = env.HOSTNAME' "$REPLICATION_DIR/manifest.json" > "$tmpfile"
    mv "$tmpfile" "$REPLICATION_DIR/manifest.json"

    echo "✅ SRP initialized"
}

collect_replication_data() {
    echo "📦 Collecting replication data..."

    local openclaw_root="/opt/openclaw"
    local workspace_root="$openclaw_root/workspace"

    # Core agent files
    if [[ -f "$workspace_root/SOUL.md" ]]; then
        cp "$workspace_root/SOUL.md" "$REPLICATION_DIR/soul/"
        echo "✅ SOUL.md collected"
    fi

    if [[ -f "$workspace_root/AGENTS.md" ]]; then
        cp "$workspace_root/AGENTS.md" "$REPLICATION_DIR/agents/"
        echo "✅ AGENTS.md collected"
    fi

    if [[ -f "$workspace_root/MEMORY.md" ]]; then
        cp "$workspace_root/MEMORY.md" "$REPLICATION_DIR/memory/"
        echo "✅ MEMORY.md collected"
    fi

    # User configuration
    if [[ -f "$workspace_root/USER.md" ]]; then
        cp "$workspace_root/USER.md" "$REPLICATION_DIR/"
        echo "✅ USER.md collected"
    fi

    # Skills directory
    if [[ -d "$workspace_root/skills" ]]; then
        rsync -av "$workspace_root/skills/" "$REPLICATION_DIR/skills/"
        echo "✅ Skills directory synchronized"
    fi

    # Memory files (daily logs) — last 30 days. -print0/xargs -0 is space-safe.
    if [[ -d "$workspace_root/memory" ]]; then
        find "$workspace_root/memory" -type f -name "*.md" -mtime -30 -print0 \
            | xargs -0 -I{} cp {} "$REPLICATION_DIR/memory/"
        echo "Recent memory files collected"
    fi

    # Configuration backups
    cp -r "$openclaw_root/config" "$REPLICATION_DIR/" 2>/dev/null || true

    # Calculate integrity hashes
    update_integrity_hashes
}

# Stable content hash of a directory: hashes per-file (filename + bytes), sorted,
# then hashes that list. NUL-delimited so spaces/newlines in names are safe.
# Always EXCLUDES manifest.json so verification is repeatable (the manifest itself
# is mutated by this very function and must not feed back into the hash).
hash_tree() {
    local dir="$1"
    [[ -d "$dir" ]] || { echo "MISSING"; return; }
    find "$dir" -type f ! -name 'manifest.json' -print0 \
        | sort -z \
        | xargs -0 -r sha256sum \
        | sha256sum | cut -d' ' -f1
}

update_integrity_hashes() {
    echo "Calculating integrity hashes..."
    local manifest_file="$REPLICATION_DIR/manifest.json" tmpfile

    # Per-component hashes
    for component in soul agents memory skills; do
        local path="$REPLICATION_DIR/$component"
        if [[ -d "$path" ]]; then
            local hash; hash="$(hash_tree "$path")"
            tmpfile="$(mktemp)"
            jq --arg comp "$component" --arg hash "$hash" \
                '.components[$comp].hash = $hash' "$manifest_file" > "$tmpfile"
            mv "$tmpfile" "$manifest_file"
        fi
    done

    # Overall hash over the whole replication set, EXCLUDING the mutable manifest.
    local overall_hash; overall_hash="$(hash_tree "$REPLICATION_DIR")"
    tmpfile="$(mktemp)"
    jq --arg hash "$overall_hash" '.integrity_hash = $hash | .last_replication = now' \
        "$manifest_file" > "$tmpfile"
    mv "$tmpfile" "$manifest_file"
    echo "Integrity hashes updated (overall: ${overall_hash:0:12}...)"
}

# Verify a replicated set on the receiving node: recompute the overall hash
# (excluding manifest.json) and compare to manifest.integrity_hash.
verify_integrity() {
    local dir="${1:-$REPLICATION_DIR}"
    local expected actual
    expected="$(jq -r '.integrity_hash' "$dir/manifest.json")"
    actual="$(hash_tree "$dir")"
    if [[ "$expected" == "$actual" ]]; then
        echo "Integrity OK ($actual)"; return 0
    else
        echo "Integrity MISMATCH: expected $expected, got $actual"; return 1
    fi
}

replicate_to_node() {
    local target_node=$1
    local target_ip=$2

    echo "🔄 Replicating to node: $target_node ($target_ip)"

    # Create replication package
    local package_name="replication-$(date +%Y%m%d-%H%M%S).tar.gz"
    local package_path="/tmp/$package_name"

    cd "$REPLICATION_DIR"
    tar -czf "$package_path" .

    local ssh_user; ssh_user="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"
    # SSH key on the primary: provisioning copies it here (see "key distribution" note).
    local ssh_key="${ALEPH_SSH_KEY:-/root/.ssh/aleph_ed25519}"
    # Transfer package to target node
    scp -i "$ssh_key" -o StrictHostKeyChecking=accept-new \
        "$package_path" "$ssh_user@$target_ip:/tmp/"

    # Execute replication on target node. We pass the package name + the SAME
    # hash_tree() implementation so the receiver can VERIFY BEFORE INSTALLING.
    ssh -i "$ssh_key" -o StrictHostKeyChecking=accept-new \
        "$ssh_user@$target_ip" "PKG='$package_name' bash -s" << 'REMOTE_SCRIPT'
#!/bin/bash
set -euo pipefail

echo "Receiving replication package..."
WORK="$(mktemp -d /tmp/repl.XXXXXX)"   # unique dir — no collisions between runs
trap 'rm -rf "$WORK"' EXIT
tar -xzf "/tmp/$PKG" -C "$WORK"
cd "$WORK"

# Same stable, manifest-excluding hash used on the sender.
hash_tree() {
    find "$1" -type f ! -name 'manifest.json' -print0 | sort -z \
        | xargs -0 -r sha256sum | sha256sum | cut -d' ' -f1
}

# VERIFY BEFORE INSTALLING — abort if the package is corrupt/tampered.
if [[ -f manifest.json ]]; then
    expected="$(jq -r '.integrity_hash' manifest.json)"
    actual="$(hash_tree "$WORK")"
    if [[ "$expected" != "$actual" ]]; then
        echo "Integrity MISMATCH (expected $expected, got $actual) — NOT installing."
        exit 1
    fi
    echo "Integrity OK ($actual)"
fi

# Install atomically-ish into the workspace, owned by the login user.
LOGIN_USER="$(logname 2>/dev/null || echo "${SUDO_USER:-$USER}")"
sudo mkdir -p /opt/openclaw/workspace/{memory,skills}
sudo chown -R "$LOGIN_USER":"$LOGIN_USER" /opt/openclaw/workspace
[[ -f soul/SOUL.md ]]     && cp soul/SOUL.md     /opt/openclaw/workspace/
[[ -f agents/AGENTS.md ]] && cp agents/AGENTS.md /opt/openclaw/workspace/
[[ -f memory/MEMORY.md ]] && cp memory/MEMORY.md /opt/openclaw/workspace/
[[ -f USER.md ]]          && cp USER.md          /opt/openclaw/workspace/
[[ -d skills ]] && rsync -a skills/ /opt/openclaw/workspace/skills/
[[ -d memory ]] && cp memory/*.md /opt/openclaw/workspace/memory/ 2>/dev/null || true

# Reload OpenClaw to pick up new workspace state (daemon-managed).
sudo systemctl restart openclaw || true
rm -f "/tmp/$PKG"
echo "Replication complete on $(hostname)"
REMOTE_SCRIPT

    rm -f "$package_path"
    echo "Replication to $target_node completed"
}

replicate_to_fleet() {
    echo "Initiating fleet-wide replication..."
    collect_replication_data

    # Ask the local fleet manager (Tailscale) for the worker list, authenticated.
    : "${FLEET_API_KEY:?FLEET_API_KEY not found in /etc/fleet-manager.env}"
    local nodes
    nodes="$(curl -fsS -H "x-api-key: $FLEET_API_KEY" "http://$FLEET_MGR_HOST:8080/fleet/status" \
        | jq -r --arg me "$(hostname)" '.nodes[] | select(.node_id != $me) | .ip_address')"

    for node_ip in $nodes; do
        replicate_to_node "worker" "$node_ip" &
    done
    wait
    echo "Fleet replication complete."

    local tmpfile; tmpfile="$(mktemp)"
    jq '.replication_count += 1' "$REPLICATION_DIR/manifest.json" > "$tmpfile"
    mv "$tmpfile" "$REPLICATION_DIR/manifest.json"
}

setup_continuous_replication() {
    echo "Setting up continuous replication..."

    # Install THIS script at a stable path so the cron job can call it. We copy the
    # currently-running file rather than assuming it already exists there.
    install -D -m 755 "$(readlink -f "$0")" /opt/openclaw/replication/auto-provisioning-protocol.sh

    # Cron wrapper INVOKES the script's subcommand (does NOT `source` it — sourcing
    # would run the command dispatcher at the bottom with no args and execute
    # `initialize_srp`, clobbering the manifest as a side effect).
    cat > /opt/openclaw/replication-cron.sh << 'CRON_SCRIPT'
#!/bin/bash
export PATH="/usr/local/bin:/usr/bin:/bin"
SRP=/opt/openclaw/replication/auto-provisioning-protocol.sh
# Only the primary (the node running fleet-manager) drives fleet replication.
if [[ -f /opt/fleet-manager/fleet-manager.js ]]; then
    echo "$(date -Iseconds): scheduled replication from primary"
    "$SRP" replicate
else
    echo "$(date -Iseconds): worker node — skipping"
fi
CRON_SCRIPT
    chmod +x /opt/openclaw/replication-cron.sh

    (crontab -l 2>/dev/null; echo "*/5 * * * * /opt/openclaw/replication-cron.sh >> /var/log/replication.log 2>&1") | crontab -
    echo "Continuous replication configured (every 5 min, primary only)"
}

# Emergency replication trigger
emergency_replicate() {
    local reason="${1:-manual_trigger}"

    echo "🚨 Emergency replication triggered: $reason"

    # Force immediate collection and replication
    collect_replication_data
    replicate_to_fleet

    # Log emergency replication
    echo "$(date -Iseconds): Emergency replication completed - $reason" >> "$REPLICATION_DIR/logs/emergency.log"
}

# Command dispatcher
case "${1:-init}" in
    "init")
        initialize_srp
        ;;
    "collect")
        collect_replication_data
        ;;
    "replicate")
        replicate_to_fleet
        ;;
    "continuous")
        setup_continuous_replication
        ;;
    "emergency")
        emergency_replicate "$2"
        ;;
    *)
        echo "Usage: $0 {init|collect|replicate|continuous|emergency}"
        exit 1
        ;;
esac
```

---

### Resource: references/backup-verification.md

## Backup Verification

**Daily Checks:**
- [ ] Backup completion status: `tail /var/log/backup.log`
- [ ] Backup size consistency
- [ ] Recovery snapshot validity

**Weekly Checks:**
- [ ] Test restore procedure on staging
- [ ] Verify backup accessibility
- [ ] Check backup retention policy

### Resource: references/contact-information.md

## Contact Information

**Emergency Contacts:**
- Primary Admin: [Your contact info]
- Backup Admin: [Backup contact info]
- Aleph Cloud support / community: https://docs.aleph.cloud and the Aleph Cloud Telegram/Discord (see the docs site footer)

**Service URLs:**
- Fleet Manager (Tailscale only): http://<PRIMARY_TAILSCALE_IP>:8080
- Load Balancer (public): http://<PRIMARY_PUBLIC_IP>
- HAProxy stats (Tailscale only): http://<PRIMARY_TAILSCALE_IP>:9090/haproxy-stats

### Resource: references/cost-optimization-strategies.md

## Contents

- Cost Optimization Strategies
- Dynamic Resource Management

## Cost Optimization Strategies

### Dynamic Resource Management

**Cost Optimization Framework:**
```bash
#!/bin/bash
# cost-optimization.sh

set -e

FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"

echo "💰 Setting up cost optimization strategies..."

analyze_costs() {
    echo "Analyzing current fleet costs from LIVE pricing..."
    local worker_count; worker_count="$(jq '.worker_nodes | length' "$FLEET_CONFIG")"

    # Pull real per-hour USD pricing from the CLI rather than hardcoding ALEPH/mo.
    # Tier 3 ~= the 4 CU primary; Tier 2 ~= the 2 CU workers (adjust to your tiers).
    local primary_hr worker_hr
    primary_hr="$(aleph pricing instance --tier 3 --payment-type credit --json 2>/dev/null \
        | jq -r '.price_per_hour // .usd_per_hour // empty' 2>/dev/null || true)"
    worker_hr="$(aleph pricing instance --tier 2 --payment-type credit --json 2>/dev/null \
        | jq -r '.price_per_hour // .usd_per_hour // empty' 2>/dev/null || true)"
    : "${primary_hr:=0.0132}"   # dated fallback (~Jun 2026); verify with `aleph pricing instance`
    : "${worker_hr:=0.0066}"

    local hours=730  # ~1 month
    local monthly; monthly="$(echo "($primary_hr + $worker_count * $worker_hr) * $hours" | bc -l)"

    cat > ~/.aleph-deploy/cost-analysis.json << COST_ANALYSIS
{
  "analysis_date": "$(date -Iseconds)",
  "pricing_source": "aleph pricing instance (USD/hour, PAYG)",
  "rates_usd_per_hour": { "primary": $primary_hr, "worker": $worker_hr },
  "node_breakdown": [
    { "type": "primary", "count": 1, "usd_per_hour": $primary_hr, "specs": "4 vCPU / 8 GiB / 80 GiB" },
    { "type": "worker",  "count": $worker_count, "usd_per_hour": $worker_hr, "specs": "2 vCPU / 4 GiB / 40 GiB" }
  ],
  "estimated_total_monthly_usd": $(printf '%.2f' "$monthly")
}
COST_ANALYSIS

    printf 'Estimated monthly cost: $%.2f USD (1 primary + %s workers, PAYG)\n' "$monthly" "$worker_count"
    echo "Source rates from 'aleph pricing instance'. Saved to cost-analysis.json."
    echo "NOTE: 'hold' payment locks ALEPH instead of streaming USD — see the pricing note at the top."
}

setup_cost_tiers() {
    echo "Setting up cost optimization tiers..."

    # estimated_monthly_usd uses the dated Jun-2026 PAYG example rates
    # (primary ~$10/mo, worker ~$5/mo). These are ESTIMATES — confirm with
    # `aleph pricing instance`. They are NOT ALEPH-token amounts.
    cat > ~/.aleph-deploy/cost-tiers.json << 'COST_TIERS'
{
  "_note": "estimated_monthly_usd are dated (~Jun 2026) PAYG examples; verify with 'aleph pricing instance'.",
  "tiers": {
    "minimal": {
      "description": "Single node for development/testing",
      "nodes": {
        "primary": 1,
        "workers": 0
      },
      "estimated_monthly_usd": 10,
      "use_cases": ["Development", "Testing", "Personal projects"]
    },
    "balanced": {
      "description": "Cost-effective production setup",
      "nodes": {
        "primary": 1,
        "workers": 2
      },
      "estimated_monthly_usd": 20,
      "use_cases": ["Small production", "Side projects", "Limited budget"]
    },
    "standard": {
      "description": "Recommended production configuration",
      "nodes": {
        "primary": 1,
        "workers": 4
      },
      "estimated_monthly_usd": 30,
      "use_cases": ["Production workloads", "Medium traffic", "Business use"]
    },
    "high_availability": {
      "description": "Enterprise-grade reliability",
      "nodes": {
        "primary": 1,
        "workers": 6,
        "backup": 1
      },
      "estimated_monthly_usd": 45,
      "use_cases": ["Critical applications", "High traffic", "Enterprise"]
    }
  },
  "optimization_strategies": {
    "spot_instances": {
      "description": "Use lower-cost CRNs for worker nodes",
      "savings_potential": "15-30%",
      "risk_level": "medium"
    },
    "auto_scaling": {
      "description": "Scale workers based on demand",
      "savings_potential": "20-40%",
      "risk_level": "low"
    },
    "mixed_crn": {
      "description": "Distribute across different CRN pricing",
      "savings_potential": "10-25%",
      "risk_level": "low"
    },
    "scheduled_scaling": {
      "description": "Reduce capacity during off-hours",
      "savings_potential": "25-50%",
      "risk_level": "low"
    }
  }
}
COST_TIERS

    echo "✅ Cost tiers configuration created"
}

setup_auto_scaling() {
    echo "📈 Setting up auto-scaling for cost optimization..."

    local primary_ip=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")

    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$primary_ip" << 'AUTOSCALE_SETUP'
#!/bin/bash

# Create auto-scaling service
cat > /opt/auto-scaler.sh << 'AUTOSCALER'
#!/bin/bash

FLEET_CONFIG="/opt/fleet-manager/nodes.json"
MIN_WORKERS=2
MAX_WORKERS=8
CPU_THRESHOLD_UP=75
CPU_THRESHOLD_DOWN=25
SCALE_COOLDOWN=300  # 5 minutes

log_message() {
    echo "$(date -Iseconds): $1" | tee -a "/var/log/auto-scaler.log"
}

get_average_cpu_usage() {
    local total_cpu=0
    local node_count=0

    # Use process substitution (< <(...)) instead of pipe (|).
    # A pipe runs `while` in a subshell, so variable updates to
    # total_cpu and node_count are lost when the subshell exits.
    while read -r ip; do
        local cpu_usage=$(ssh -i /root/.ssh/aleph_ed25519 \
                             -o ConnectTimeout=5 "${REMOTE_USER:-root}@$ip" \
                             "top -bn1 | grep 'Cpu(s)' | awk '{print \$2}' | cut -d'%' -f1" 2>/dev/null || echo "0")

        if [[ "$cpu_usage" =~ ^[0-9.]+$ ]]; then
            total_cpu=$(echo "$total_cpu + $cpu_usage" | bc -l)
            node_count=$((node_count + 1))
        fi
    done < <(jq -r '.nodes[] | select(.status == "active" and .node_id != "primary") | .ip_address' "$FLEET_CONFIG")

    if (( node_count > 0 )); then
        echo "scale=2; $total_cpu / $node_count" | bc -l
    else
        echo "0"
    fi
}

# Real scale-up: create + provision a new worker via the aleph CLI, then let it
# register. Requires the aleph CLI + funded account + key + env on the primary.
scale_up() {
    local current_workers; current_workers="$(jq '[.nodes[] | select(.status=="active" and .node_id!="primary")] | length' "$FLEET_CONFIG")"
    (( current_workers >= MAX_WORKERS )) && { log_message "At MAX_WORKERS ($MAX_WORKERS)"; return 1; }
    command -v aleph >/dev/null || { log_message "aleph CLI absent on primary — cannot scale up."; return 1; }
    : "${FLEET_API_KEY:?}"; : "${PRIMARY_TS_IP:?}"

    local name="auto-worker-$(date +%s)" out hash ip
    log_message "Scaling up: creating $name"
    out="$(aleph instance create --name "$name" --compute-units 2 --rootfs-size 40960 \
            --ssh-pubkey-file /root/.ssh/aleph_ed25519.pub --payment-type credit --payment-chain BASE 2>&1)"
    hash="$(printf '%s\n' "$out" | grep -oE '[0-9a-f]{64}' | head -1)"
    for _ in $(seq 1 30); do
        ip="$(aleph instance list --json | jq -r --arg n "$name" '.[]|select(.name==$n)|(.ipv4//.ipv6//empty)' | head -1)"
        [[ -n "$ip" ]] && break; sleep 10
    done
    [[ -z "$ip" ]] && { log_message "Scale-up: $name got no IP"; return 1; }
    # ITEM_HASH is passed through so the registry records this instance's hash;
    # scale_down() reads .item_hash to delete the right instance.
    ssh -i /root/.ssh/aleph_ed25519 -o StrictHostKeyChecking=accept-new "root@$ip" \
        "NODE_ID='$name' PRIMARY_TS_IP='$PRIMARY_TS_IP' FLEET_API_KEY='$FLEET_API_KEY' TAILSCALE_AUTH_KEY='${TAILSCALE_AUTH_KEY:-}' ITEM_HASH='$hash' bash -s" <<'REPROV'
set -euo pipefail; export DEBIAN_FRONTEND=noninteractive
apt-get update && apt-get install -y curl jq iproute2
installer_1="$(mktemp)"
curl -fsSL https://get.docker.com -o "$installer_1"
less "$installer_1"  # Review before execution; verify the release checksum/signature when published.
sh "$installer_1"
rm -f "$installer_1"
installer_2="$(mktemp)"
curl -fsSL https://deb.nodesource.com/setup_22.x -o "$installer_2"
less "$installer_2"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_2" - && apt-get install -y nodejs
rm -f "$installer_2"
installer_3="$(mktemp)"
curl -fsSL https://tailscale.com/install.sh -o "$installer_3"
less "$installer_3"  # Review before execution; verify the release checksum/signature when published.
sh "$installer_3"
# file: pattern keeps the auth key out of the process list (see Tailscale section)
rm -f "$installer_3"
[[ -n "${TAILSCALE_AUTH_KEY:-}" ]] && { printf '%s' "$TAILSCALE_AUTH_KEY" > /tmp/ts && chmod 600 /tmp/ts && tailscale up --auth-key="file:/tmp/ts" --hostname="$NODE_ID"; rm -f /tmp/ts; }
installer_4="$(mktemp)"
curl -fsSL https://openclaw.ai/install.sh -o "$installer_4"
less "$installer_4"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_4"
rm -f "$installer_4"
TS_IP="$(tailscale ip -4 2>/dev/null || hostname -I | awk '{print $1}')"
curl -fsS -X POST "http://$PRIMARY_TS_IP:8080/fleet/register" -H "x-api-key: $FLEET_API_KEY" \
  -H 'Content-Type: application/json' -d "{\"node_id\":\"$NODE_ID\",\"ip_address\":\"$TS_IP\",\"item_hash\":\"${ITEM_HASH:-}\",\"capabilities\":[\"compute\",\"openclaw\"]}"
REPROV
    log_message "Scale-up complete: $name ($ip). haproxy-fleet-sync will add it within 60s."
    echo "$(date +%s)" > /tmp/last-scale-action
}

# Real scale-down: drain in HAProxy, deregister, then DELETE the Aleph instance.
scale_down() {
    local current_workers; current_workers="$(jq '[.nodes[]|select(.status=="active" and .node_id!="primary")]|length' "$FLEET_CONFIG")"
    (( current_workers <= MIN_WORKERS )) && { log_message "At MIN_WORKERS ($MIN_WORKERS)"; return 1; }
    command -v aleph >/dev/null || { log_message "aleph CLI absent on primary — cannot scale down."; return 1; }

    local victim; victim="$(jq -r '[.nodes[]|select(.status=="active" and .node_id!="primary")]|sort_by(.cpu_usage // 0)|first|.node_id' "$FLEET_CONFIG")"
    [[ -z "$victim" || "$victim" == "null" ]] && return 0
    local hash; hash="$(jq -r --arg n "$victim" '.nodes[]|select(.node_id==$n)|.item_hash // empty' "$FLEET_CONFIG")"
    log_message "Scaling down: draining $victim"
    # 1. Mark draining; 2. remove from HAProxy; 3. delete instance; 4. drop from state.
    local tmpfile; tmpfile="$(mktemp)"
    jq --arg n "$victim" '.nodes = (.nodes | map(if .node_id==$n then .status="draining" else . end))' "$FLEET_CONFIG" > "$tmpfile" && mv "$tmpfile" "$FLEET_CONFIG"
    /opt/manage-haproxy-backends.sh remove "$victim" 2>/dev/null || true
    sleep 10   # let in-flight requests finish
    if [[ -n "$hash" ]]; then
        aleph instance delete "$hash" && log_message "Deleted instance $hash ($victim)"
    fi
    tmpfile="$(mktemp)"
    jq --arg n "$victim" '.nodes |= map(select(.node_id != $n))' "$FLEET_CONFIG" > "$tmpfile" && mv "$tmpfile" "$FLEET_CONFIG"
    log_message "Scale-down complete: removed $victim"
    echo "$(date +%s)" > /tmp/last-scale-action
}

check_scaling_needed() {
    log_message "🔍 Checking if scaling is needed..."

    # Check cooldown period
    if [[ -f /tmp/last-scale-action ]]; then
        local last_action=$(cat /tmp/last-scale-action)
        local current_time=$(date +%s)
        local time_diff=$((current_time - last_action))

        if (( time_diff < SCALE_COOLDOWN )); then
            log_message "⏳ Still in cooldown period ($((SCALE_COOLDOWN - time_diff))s remaining)"
            return 0
        fi
    fi

    local avg_cpu=$(get_average_cpu_usage)
    log_message "📊 Current average CPU usage: $avg_cpu%"

    if (( $(echo "$avg_cpu > $CPU_THRESHOLD_UP" | bc -l) )); then
        log_message "🔺 CPU usage above threshold ($CPU_THRESHOLD_UP%), scaling up..."
        scale_up
    elif (( $(echo "$avg_cpu < $CPU_THRESHOLD_DOWN" | bc -l) )); then
        log_message "🔻 CPU usage below threshold ($CPU_THRESHOLD_DOWN%), scaling down..."
        scale_down
    else
        log_message "✅ CPU usage within acceptable range"
    fi
}

# Dispatcher: `daemon` runs the loop (used by systemd); the others let the
# scheduled-scaler (and operators) invoke a single action.
case "${1:-daemon}" in
    daemon)     while true; do check_scaling_needed; sleep 60; done ;;
    once)       check_scaling_needed ;;
    scale-up)   scale_up ;;
    scale-down) scale_down ;;
    *) echo "Usage: $0 {daemon|once|scale-up|scale-down}"; exit 1 ;;
esac
AUTOSCALER

chmod +x /opt/auto-scaler.sh

# Create systemd service (disabled by default)
cat > /etc/systemd/system/auto-scaler.service << 'SCALER_SERVICE'
[Unit]
Description=Fleet Auto Scaler
After=network.target fleet-manager.service

[Service]
Type=simple
User=root
EnvironmentFile=/etc/fleet-manager.env
ExecStart=/opt/auto-scaler.sh
Restart=always
RestartSec=30
Environment=AUTO_SCALING_ENABLED=false

[Install]
WantedBy=multi-user.target
SCALER_SERVICE

# Disabled by default. Auto-scaling CREATES and DELETES paid instances, so enable
# it only after confirming the aleph CLI, a funded account, the fleet SSH key, and
# FLEET_API_KEY/PRIMARY_TS_IP/TAILSCALE_AUTH_KEY are present in /etc/fleet-manager.env.
echo "Auto-scaler configured (disabled by default)"
echo "To enable: systemctl enable --now auto-scaler"
AUTOSCALE_SETUP

echo "✅ Auto-scaling configured on primary node"
}

setup_scheduled_scaling() {
    echo "⏰ Setting up scheduled scaling for off-hours cost savings..."

    local primary_ip=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")

    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$primary_ip" << 'SCHEDULED_SETUP'
#!/bin/bash

# Create scheduled scaling script
cat > /opt/scheduled-scaler.sh << 'SCHEDULER'
#!/bin/bash
set -euo pipefail

# cron has a bare environment — load the shared key/host so the delegated
# auto-scaler actions (which call the aleph CLI over the mesh) have what they need.
[[ -f /etc/fleet-manager.env ]] && { set -a; . /etc/fleet-manager.env; set +a; }

FLEET_CONFIG="/opt/fleet-manager/nodes.json"

log_message() {
    echo "$(date -Iseconds): $1" | tee -a "/var/log/scheduled-scaler.log"
}

# Drive worker count to a target by invoking the auto-scaler's single-step actions
# (which perform real aleph create/delete). One step per loop, with a short pause.
scale_to_count() {
    local target_count="$1" reason="$2"
    log_message "Scaling to $target_count workers: $reason"
    local current_count; current_count="$(jq '[.nodes[]|select(.status=="active" and .node_id!="primary")]|length' "$FLEET_CONFIG")"

    if (( target_count == current_count )); then
        log_message "Already at target capacity ($target_count)"; return 0
    fi
    if (( target_count > current_count )); then
        local n=$((target_count - current_count))
        log_message "Adding $n worker(s) via auto-scaler"
        for ((i=0; i<n; i++)); do /opt/auto-scaler.sh scale-up || break; sleep 15; done
    else
        local n=$((current_count - target_count))
        log_message "Removing $n worker(s) via auto-scaler"
        for ((i=0; i<n; i++)); do /opt/auto-scaler.sh scale-down || break; sleep 5; done
    fi
}

# Scaling schedules based on time
current_hour=$(date +%H)
current_day=$(date +%u)  # 1=Monday, 7=Sunday

# Business hours scaling (9 AM - 6 PM weekdays)
if (( current_day <= 5 && current_hour >= 9 && current_hour <= 18 )); then
    scale_to_count 4 "Business hours scaling"
# Evening hours (6 PM - 11 PM)
elif (( current_day <= 5 && current_hour >= 19 && current_hour <= 23 )); then
    scale_to_count 2 "Evening hours scaling"
# Night/weekend minimal capacity
else
    scale_to_count 1 "Off-hours minimal scaling"
fi
SCHEDULER

chmod +x /opt/scheduled-scaler.sh

# Setup cron jobs for scheduled scaling
(crontab -l 2>/dev/null; echo "0 9 * * 1-5 /opt/scheduled-scaler.sh >> /var/log/scheduled-scaler.log 2>&1") | crontab -
(crontab -l 2>/dev/null; echo "0 18 * * 1-5 /opt/scheduled-scaler.sh >> /var/log/scheduled-scaler.log 2>&1") | crontab -
(crontab -l 2>/dev/null; echo "0 23 * * * /opt/scheduled-scaler.sh >> /var/log/scheduled-scaler.log 2>&1") | crontab -

echo "✅ Scheduled scaling configured"
echo "Schedules:"
echo "- Business hours (9 AM): Scale to 4 workers"
echo "- Evening hours (6 PM): Scale to 2 workers"
echo "- Night/weekends (11 PM): Scale to 1 worker"
SCHEDULED_SETUP

echo "✅ Scheduled scaling configured"
}

create_cost_monitoring() {
    echo "📈 Setting up cost monitoring dashboard..."

    cat > ~/.aleph-deploy/scripts/cost-monitor.sh << 'COST_MONITOR'
#!/bin/bash
# cost-monitor.sh — fleet cost report from LIVE `aleph pricing` (USD, PAYG).
# Run from a machine on the tailnet (queries the fleet manager over Tailscale).
set -euo pipefail

FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"
: "${FLEET_API_KEY:?Set FLEET_API_KEY (see fleet.env)}"
MGR_HOST="$(jq -r '.primary_node.tailscale_ip // .primary_node.ip' "$FLEET_CONFIG")"
mkdir -p ~/.aleph-deploy/reports

# Live USD/hour rates (tier 3 ~ primary, tier 2 ~ worker). Dated fallbacks if the
# CLI is unavailable; ALWAYS verify with `aleph pricing instance`.
rate() { aleph pricing instance --tier "$1" --payment-type credit --json 2>/dev/null \
    | jq -r '.price_per_hour // .usd_per_hour // empty' 2>/dev/null || true; }

generate_cost_report() {
    local report_date; report_date="$(date +%Y-%m-%d)"
    local fleet_status active_workers
    fleet_status="$(curl -fsS -H "x-api-key: $FLEET_API_KEY" "http://$MGR_HOST:8080/fleet/status" 2>/dev/null || echo '{"nodes":[]}')"
    active_workers="$(echo "$fleet_status" | jq '[.nodes[]|select(.status=="active" and .node_id!="primary")]|length')"

    local p_hr w_hr; p_hr="$(rate 3)"; w_hr="$(rate 2)"
    : "${p_hr:=0.0132}"; : "${w_hr:=0.0066}"     # ~Jun 2026 fallback — verify!
    local monthly daily
    monthly="$(echo "($p_hr + $active_workers * $w_hr) * 730" | bc -l)"
    daily="$(echo "$monthly / 30" | bc -l)"

    cat > ~/.aleph-deploy/reports/cost-report-$report_date.json << REPORT
{
  "report_date": "$report_date",
  "pricing_source": "aleph pricing instance (USD/hour, PAYG)",
  "rates_usd_per_hour": { "primary": $p_hr, "worker": $w_hr },
  "fleet": { "primary_nodes": 1, "worker_nodes": $active_workers, "total_nodes": $((active_workers + 1)) },
  "estimated_monthly_usd": $(printf '%.2f' "$monthly"),
  "estimated_daily_usd": $(printf '%.2f' "$daily"),
  "recommendations": ["Enable scheduled scaling for off-hours", "Right-size worker count to real load"]
}
REPORT

    echo "COST SUMMARY ($report_date)"
    echo "Active nodes: $((active_workers + 1)) (1 primary + $active_workers workers)"
    printf 'Estimated monthly: $%.2f USD   daily: $%.2f USD (PAYG)\n' "$monthly" "$daily"
    echo "Rates from 'aleph pricing instance'. Report: ~/.aleph-deploy/reports/cost-report-$report_date.json"
}

generate_cost_report
(crontab -l 2>/dev/null; echo "0 8 * * * $HOME/.aleph-deploy/scripts/cost-monitor.sh >> /var/log/cost-monitor.log 2>&1") | crontab -
COST_MONITOR

chmod +x ~/.aleph-deploy/scripts/cost-monitor.sh

echo "✅ Cost monitoring configured"
}

# Execute cost optimization setup
analyze_costs
setup_cost_tiers
setup_auto_scaling
setup_scheduled_scaling
create_cost_monitoring

echo "💰 Cost optimization setup complete!"
echo ""
echo "Available cost optimization features:"
echo "- Auto-scaling based on CPU usage (disabled by default)"
echo "- Scheduled scaling for off-hours savings"
echo "- Daily cost reporting and monitoring"
echo "- Multiple deployment tiers (minimal to high-availability)"
echo ""
echo "Enable auto-scaling: ssh root@PRIMARY_IP 'sudo systemctl enable auto-scaler && sudo systemctl start auto-scaler'"
echo "View cost reports: ls ~/.aleph-deploy/reports/"
echo "Monitor costs: ~/.aleph-deploy/scripts/cost-monitor.sh"
```

---

### Resource: references/disaster-recovery-auto-recreation.md

## Contents

- Disaster Recovery & Auto-Recreation
- Automated Backup System

## Disaster Recovery & Auto-Recreation

### Automated Backup System

**Comprehensive Backup Framework:**
```bash
#!/bin/bash
# disaster-recovery-system.sh

set -e

FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"
BACKUP_RETENTION_DAYS=30
BACKUP_STORAGE_PATH="/opt/openclaw/backups"

echo "🛡️ Setting up Disaster Recovery System..."

setup_backup_infrastructure() {
    local primary_ip=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")

    echo "📦 Setting up backup infrastructure..."

    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$primary_ip" << 'BACKUP_SETUP'
#!/bin/bash
set -e

# Create backup directories. Own them by the ACTUAL login user (root on Aleph
# base images, ubuntu on some) — never hardcode "ubuntu", which does not exist on
# root-only images and would abort this script under `set -e`.
LOGIN_USER="$(logname 2>/dev/null || echo "${SUDO_USER:-root}")"
sudo mkdir -p /opt/openclaw/backups/{fleet,nodes,data,logs}
sudo chown -R "$LOGIN_USER":"$LOGIN_USER" /opt/openclaw/backups

# Install backup tools
sudo apt-get update
sudo apt-get install -y rsync rclone jq awscli

# Create comprehensive backup script
cat > /opt/openclaw/backup-system.sh << 'BACKUP_SCRIPT'
#!/bin/bash
set -uo pipefail

BACKUP_BASE="/opt/openclaw/backups"
TIMESTAMP=$(date +%Y%m%d-%H%M%S)
RETENTION_DAYS=30
# SSH login user for reaching workers (image-dependent; Aleph base images use root).
REMOTE_USER="${REMOTE_USER:-root}"
SSH_KEY="/root/.ssh/aleph_ed25519"

log_message() {
    echo "$(date -Iseconds): $1" | tee -a "$BACKUP_BASE/backup.log"
}

backup_fleet_config() {
    log_message "📋 Backing up fleet configuration..."

    local backup_dir="$BACKUP_BASE/fleet/$TIMESTAMP"
    mkdir -p "$backup_dir"

    # Fleet registry
    cp /opt/fleet-manager/nodes.json "$backup_dir/" 2>/dev/null || true

    # HAProxy configuration
    cp /etc/haproxy/haproxy.cfg "$backup_dir/" 2>/dev/null || true

    # Service configurations
    cp /etc/systemd/system/fleet-manager.service "$backup_dir/" 2>/dev/null || true
    cp /etc/systemd/system/haproxy-fleet-sync.service "$backup_dir/" 2>/dev/null || true

    # Network configurations
    cp /opt/tailscale-info.json "$backup_dir/" 2>/dev/null || true

    log_message "✅ Fleet configuration backed up to $backup_dir"
}

backup_node_data() {
    local node_ip=$1
    local node_name=$2

    log_message "💾 Backing up data from $node_name ($node_ip)..."

    local backup_dir="$BACKUP_BASE/nodes/$TIMESTAMP/$node_name"
    mkdir -p "$backup_dir"

    # Backup OpenClaw workspace
    rsync -av --compress --delete \
        -e "ssh -i $SSH_KEY -o StrictHostKeyChecking=accept-new" \
        "$REMOTE_USER@$node_ip:/opt/openclaw/workspace/" \
        "$backup_dir/workspace/" 2>/dev/null || true

    # Backup configurations
    rsync -av --compress \
        -e "ssh -i $SSH_KEY -o StrictHostKeyChecking=accept-new" \
        "$REMOTE_USER@$node_ip:/opt/openclaw/config/" \
        "$backup_dir/config/" 2>/dev/null || true

    # Backup logs (last 7 days only)
    ssh -i "$SSH_KEY" "$REMOTE_USER@$node_ip" \
        "find /var/log -name '*.log' -mtime -7 -exec tar -czf /tmp/logs-$node_name.tar.gz {} +" 2>/dev/null || true

    scp -i "$SSH_KEY" \
        "$REMOTE_USER@$node_ip":/tmp/logs-$node_name.tar.gz \
        "$backup_dir/" 2>/dev/null || true

    log_message "✅ Node data backed up for $node_name"
}

backup_all_nodes() {
    log_message "🌐 Starting full fleet backup..."

    # Backup fleet configuration
    backup_fleet_config

    # Get fleet nodes
    if [[ -f /opt/fleet-manager/nodes.json ]]; then
        local nodes=$(jq -r '.nodes[] | select(.status == "active") | .node_id + "," + .ip_address' /opt/fleet-manager/nodes.json)

        # Backup each node in parallel
        while IFS=',' read -r node_id ip_address; do
            backup_node_data "$ip_address" "$node_id" &
        done <<< "$nodes"

        # Wait for all backups to complete
        wait
    fi

    log_message "✅ Full fleet backup completed"
}

cleanup_old_backups() {
    log_message "🧹 Cleaning up old backups..."

    # Remove backups older than retention period
    find "$BACKUP_BASE" -type d -name "20*" -mtime +$RETENTION_DAYS -exec rm -rf {} + 2>/dev/null || true

    log_message "✅ Old backups cleaned up"
}

create_recovery_snapshot() {
    log_message "📸 Creating recovery snapshot..."

    local snapshot_file="$BACKUP_BASE/recovery-snapshot-$TIMESTAMP.json"

    # Create comprehensive recovery information
    cat > "$snapshot_file" << SNAPSHOT
{
  "timestamp": "$TIMESTAMP",
  "fleet_config": $(cat /opt/fleet-manager/nodes.json 2>/dev/null || echo '{"nodes":[]}'),
  "system_info": {
    "hostname": "$(hostname)",
    "uptime": "$(uptime)",
    "disk_usage": $(df -h / | awk 'NR==2{print "{\\"used\\": \\""$5"\\", \\"available\\": \\""$4"\\"}"}'),
    "memory_usage": $(free -h | awk 'NR==2{print "{\\"total\\": \\""$2"\\", \\"used\\": \\""$3"\\", \\"free\\": \\""$7"\\"}"}')
  },
  "services_status": {
    "fleet_manager": "$(systemctl is-active fleet-manager 2>/dev/null || echo 'inactive')",
    "haproxy": "$(systemctl is-active haproxy 2>/dev/null || echo 'inactive')",
    "openclaw": "$(systemctl is-active openclaw 2>/dev/null || echo 'inactive')"
  },
  "network_info": {
    "tailscale_status": $(tailscale status --json 2>/dev/null || echo '{}'),
    "public_ip": "$(curl -s http://checkip.amazonaws.com 2>/dev/null || echo 'unknown')"
  }
}
SNAPSHOT

    log_message "✅ Recovery snapshot created: $snapshot_file"
}

# Main backup execution
case "${1:-full}" in
    "full")
        backup_all_nodes
        create_recovery_snapshot
        cleanup_old_backups
        ;;
    "config")
        backup_fleet_config
        ;;
    "snapshot")
        create_recovery_snapshot
        ;;
    "cleanup")
        cleanup_old_backups
        ;;
    *)
        echo "Usage: $0 {full|config|snapshot|cleanup}"
        exit 1
        ;;
esac
BACKUP_SCRIPT

chmod +x /opt/openclaw/backup-system.sh

# Setup automated backups via cron
(crontab -l 2>/dev/null; echo "0 2 * * * /opt/openclaw/backup-system.sh full >> /var/log/backup.log 2>&1") | crontab -
(crontab -l 2>/dev/null; echo "0 */6 * * * /opt/openclaw/backup-system.sh snapshot >> /var/log/backup.log 2>&1") | crontab -

echo "✅ Backup infrastructure setup complete"
BACKUP_SETUP

echo "✅ Backup infrastructure configured on primary node"
}

setup_node_monitoring() {
    echo "👁️ Setting up node monitoring and auto-recreation..."

    local primary_ip=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")

    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$primary_ip" << 'MONITORING_SETUP'
#!/bin/bash

# Create node monitoring service
cat > /opt/node-monitor.sh << 'MONITOR_SCRIPT'
#!/bin/bash

FLEET_CONFIG="/opt/fleet-manager/nodes.json"
CHECK_INTERVAL=60
FAILURE_THRESHOLD=3

log_message() {
    echo "$(date -Iseconds): $1" | tee -a "/var/log/node-monitor.log"
}

check_node_health() {
    local node_id=$1
    local node_ip=$2

    # SSH login user is image-dependent (root on Aleph base images).
    local ru="${REMOTE_USER:-root}"
    # Check SSH connectivity
    if ! ssh -i /root/.ssh/aleph_ed25519 \
            -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new \
            "$ru@$node_ip" "echo 'alive'" &>/dev/null; then
        return 1
    fi

    # Check OpenClaw service
    if ! ssh -i /root/.ssh/aleph_ed25519 \
            "$ru@$node_ip" "systemctl is-active openclaw" &>/dev/null; then
        return 2
    fi

    # Check OpenClaw gateway health via the CLI over SSH (no HTTP /health
    # endpoint is documented; the gateway binds loopback by default anyway).
    if ! ssh -i /root/.ssh/aleph_ed25519 \
            "$ru@$node_ip" "openclaw gateway status || openclaw health" &>/dev/null; then
        return 3
    fi

    return 0
}

mark_node_unhealthy() {
    local node_id=$1
    local failure_reason=$2

    log_message "❌ Node $node_id marked as unhealthy: $failure_reason"

    # Update node status in fleet registry
    local tmpfile=$(mktemp)
    jq --arg node "$node_id" --arg status "unhealthy" \
        '.nodes = (.nodes | map(if .node_id == $node then .status = $status else . end))' \
        "$FLEET_CONFIG" > "$tmpfile"
    mv "$tmpfile" "$FLEET_CONFIG"
}

# Recreate a dead worker. REQUIREMENTS on the primary: the aleph-client CLI must be
# installed and a funded account configured (so `aleph instance create` can run
# unattended), plus the fleet SSH private key at /root/.ssh/aleph_ed25519 and the
# fleet API key in /etc/fleet-manager.env. Without these, recreation is skipped
# with a clear log line rather than silently "succeeding".
auto_recreate_node() {
    local node_id="$1"
    log_message "Auto-recreating failed node: $node_id"

    local node_config; node_config="$(jq -c --arg n "$node_id" '.nodes[] | select(.node_id==$n)' "$FLEET_CONFIG")"
    [[ -z "$node_config" || "$node_config" == "null" ]] && { log_message "No config for $node_id"; return 1; }

    command -v aleph >/dev/null || { log_message "aleph CLI not on primary — cannot recreate; alerting operator."; return 1; }
    [[ -f /root/.ssh/aleph_ed25519 ]] || { log_message "Fleet SSH key missing on primary — cannot provision replacement."; return 1; }
    : "${FLEET_API_KEY:?}"; : "${PRIMARY_TS_IP:?PRIMARY_TS_IP must be set in the unit env}"

    # 1. Delete the dead instance if we have its item-hash (frees PAYG billing / held tokens).
    local old_hash; old_hash="$(jq -r '.item_hash // empty' <<< "$node_config")"
    if [[ -n "$old_hash" ]]; then
        log_message "Deleting dead instance $old_hash"
        aleph instance delete "$old_hash" || log_message "WARN: delete failed (already gone?)"
    fi

    # 2. Create a like-for-like replacement (2 CU / 40 GiB worker).
    local out new_hash new_ip
    out="$(aleph instance create --name "$node_id" --compute-units 2 --rootfs-size 40960 \
            --ssh-pubkey-file /root/.ssh/aleph_ed25519.pub \
            --payment-type credit --payment-chain BASE 2>&1)"
    log_message "create: $out"
    new_hash="$(printf '%s\n' "$out" | grep -oE '[0-9a-f]{64}' | head -1)"

    # 3. Wait for an IP via the REAL `aleph instance list`.
    for _ in $(seq 1 30); do
        new_ip="$(aleph instance list --json | jq -r --arg n "$node_id" '.[] | select(.name==$n) | (.ipv4 // .ipv6 // empty)' | head -1)"
        [[ -n "$new_ip" ]] && break; sleep 10
    done
    [[ -z "$new_ip" ]] && { log_message "Replacement $node_id got no IP"; return 1; }

    # 4. Re-provision over SSH: install OpenClaw + Tailscale, re-register with the primary.
    #    ITEM_HASH carries the NEW instance hash so the registry stays able to
    #    delete/recreate this node on the next failure.
    ssh -i /root/.ssh/aleph_ed25519 -o StrictHostKeyChecking=accept-new "root@$new_ip" \
        "NODE_ID='$node_id' PRIMARY_TS_IP='$PRIMARY_TS_IP' FLEET_API_KEY='$FLEET_API_KEY' \
         TAILSCALE_AUTH_KEY='${TAILSCALE_AUTH_KEY:-}' ITEM_HASH='$new_hash' bash -s" <<'REPROV'
set -euo pipefail
export DEBIAN_FRONTEND=noninteractive
apt-get update && apt-get install -y curl jq iproute2
installer_1="$(mktemp)"
curl -fsSL https://get.docker.com -o "$installer_1"
less "$installer_1"  # Review before execution; verify the release checksum/signature when published.
sh "$installer_1"
rm -f "$installer_1"
installer_2="$(mktemp)"
curl -fsSL https://deb.nodesource.com/setup_22.x -o "$installer_2"
less "$installer_2"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_2" - && apt-get install -y nodejs
rm -f "$installer_2"
installer_3="$(mktemp)"
curl -fsSL https://tailscale.com/install.sh -o "$installer_3"
less "$installer_3"  # Review before execution; verify the release checksum/signature when published.
sh "$installer_3"
# file: pattern keeps the auth key out of the process list (see Tailscale section)
rm -f "$installer_3"
[[ -n "${TAILSCALE_AUTH_KEY:-}" ]] && { printf '%s' "$TAILSCALE_AUTH_KEY" > /tmp/ts && chmod 600 /tmp/ts && tailscale up --auth-key="file:/tmp/ts" --hostname="$NODE_ID"; rm -f /tmp/ts; }
installer_4="$(mktemp)"
curl -fsSL https://openclaw.ai/install.sh -o "$installer_4"
less "$installer_4"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_4"
rm -f "$installer_4"
TS_IP="$(tailscale ip -4 2>/dev/null || hostname -I | awk '{print $1}')"
curl -fsS -X POST "http://$PRIMARY_TS_IP:8080/fleet/register" -H "x-api-key: $FLEET_API_KEY" \
  -H 'Content-Type: application/json' \
  -d "{\"node_id\":\"$NODE_ID\",\"ip_address\":\"$TS_IP\",\"item_hash\":\"${ITEM_HASH:-}\",\"capabilities\":[\"compute\",\"openclaw\"]}"
REPROV

    # 5. Update fleet state atomically: new hash/ip, status active, reset failures.
    local tmp; tmp="$(mktemp)"
    jq --arg n "$node_id" --arg h "$new_hash" --arg ip "$new_ip" \
       '.nodes = (.nodes | map(if .node_id==$n then (.item_hash=$h | .ip_address=$ip | .status="active" | .failure_count=0) else . end))' \
       "$FLEET_CONFIG" > "$tmp" && mv "$tmp" "$FLEET_CONFIG"
    log_message "Node $node_id recreated: $new_ip ($new_hash)"
}

monitor_fleet() {
    log_message "🔍 Starting fleet monitoring cycle..."

    if [[ ! -f "$FLEET_CONFIG" ]]; then
        log_message "⚠️ Fleet configuration not found"
        return 1
    fi

    local nodes=$(jq -r '.nodes[] | select(.status != "unhealthy") | .node_id + "," + .ip_address' "$FLEET_CONFIG")

    while IFS=',' read -r node_id ip_address; do
        [[ -z "$node_id" ]] && continue

        log_message "Checking health of $node_id ($ip_address)..."

        if ! check_node_health "$node_id" "$ip_address"; then
            local failure_count=$(jq -r --arg node "$node_id" '.nodes[] | select(.node_id == $node) | .failure_count // 0' "$FLEET_CONFIG")
            failure_count=$((failure_count + 1))

            # Update failure count
            local tmpfile=$(mktemp)
            jq --arg node "$node_id" --argjson count "$failure_count" \
                '.nodes = (.nodes | map(if .node_id == $node then .failure_count = $count else . end))' \
                "$FLEET_CONFIG" > "$tmpfile"
            mv "$tmpfile" "$FLEET_CONFIG"

            if (( failure_count >= FAILURE_THRESHOLD )); then
                mark_node_unhealthy "$node_id" "Health check failed $failure_count times"

                # Auto-recreate if enabled
                if [[ "$AUTO_RECREATE" == "true" ]]; then
                    auto_recreate_node "$node_id"
                fi
            else
                log_message "⚠️ Node $node_id health check failed ($failure_count/$FAILURE_THRESHOLD)"
            fi
        else
            # Reset failure count on successful check
            local tmpfile=$(mktemp)
            jq --arg node "$node_id" '.nodes = (.nodes | map(if .node_id == $node then .failure_count = 0 else . end))' \
                "$FLEET_CONFIG" > "$tmpfile"
            mv "$tmpfile" "$FLEET_CONFIG"

            log_message "✅ Node $node_id healthy"
        fi
    done <<< "$nodes"
}

# Continuous monitoring loop
while true; do
    monitor_fleet
    sleep $CHECK_INTERVAL
done
MONITOR_SCRIPT

chmod +x /opt/node-monitor.sh

# Create systemd service for monitoring. AUTO_RECREATE defaults to FALSE — it
# deletes+recreates paid instances and needs the aleph CLI, a funded account,
# TAILSCALE_AUTH_KEY, and PRIMARY_TS_IP. Turn it on deliberately once those are
# in /etc/fleet-manager.env. With it off, the monitor only marks nodes unhealthy
# and logs, so an operator can decide.
cat > /etc/systemd/system/node-monitor.service << 'MONITOR_SERVICE'
[Unit]
Description=Fleet Node Monitor
After=network.target fleet-manager.service

[Service]
Type=simple
User=root
EnvironmentFile=/etc/fleet-manager.env
ExecStart=/opt/node-monitor.sh
Restart=always
RestartSec=30
# Set AUTO_RECREATE=true in /etc/fleet-manager.env to enable destructive recreation.
Environment=AUTO_RECREATE=false

[Install]
WantedBy=multi-user.target
MONITOR_SERVICE

sudo systemctl daemon-reload
sudo systemctl enable node-monitor
sudo systemctl start node-monitor

echo "Node monitoring service configured (AUTO_RECREATE off by default)"
MONITORING_SETUP

echo "Node monitoring and auto-recreation configured"
}

create_disaster_recovery_runbook() {
    echo "📖 Creating disaster recovery runbook..."

    cat > ~/.aleph-deploy/DISASTER_RECOVERY_RUNBOOK.md << 'RUNBOOK'
# Disaster Recovery Runbook

### Resource: references/emergency-response-procedures.md

## Contents

- Emergency Response Procedures
- 1. Primary Node Failure
- 2. Multiple Worker Node Failures
- 3. Complete Fleet Failure
- 4. Data Loss Recovery

## Emergency Response Procedures

### 1. Primary Node Failure

**Symptoms:**
- Fleet manager unreachable
- Load balancer not responding
- Cannot access fleet status API

**Recovery Steps:**
1. Check instance status: `aleph instance list` (find the primary by name; note its item-hash/IP).
2. If the instance is gone, recreate the primary and restore its state from your
   off-node backups (the backup target on a different CRN, or local pulls):
   ```bash
   cd ~/.aleph-deploy
   ./deploy-fleet.sh openclaw-fleet 1     # deploy a fresh primary
   # Restore /opt/fleet-manager and /opt/openclaw/config from the latest backup
   # under ~/.aleph-deploy/backups (or the backup node), e.g.:
   rsync -a ~/.aleph-deploy/backups/fleet/<latest>/  "$SSH_USER@<new_primary_ip>:/tmp/restore/"
   ```
3. Update DNS/routing to the new primary IP.
4. Workers re-register automatically once the fleet manager is back on the mesh.

### 2. Multiple Worker Node Failures

**Symptoms:**
- Reduced capacity
- Load balancer showing failed backends
- High response times

**Recovery Steps:**
1. Check fleet status: `./fleet-control.sh status` (uses the authenticated mgr helper).
2. Identify failed nodes.
3. If AUTO_RECREATE is enabled it triggers automatically; otherwise restore capacity:
   ```bash
   ./fleet-control.sh scale 5   # recreate workers up to the target (confirms deletes)
   ```
4. Monitor recovery progress.

### 3. Complete Fleet Failure

**Symptoms:**
- All nodes unreachable
- Complete service outage

**Recovery Steps:**
1. Confirm what still exists: `aleph instance list`.
2. Deploy a fresh primary, then workers:
   ```bash
   ./deploy-single-vm.sh openclaw-recovery-primary
   ./deploy-fleet.sh openclaw-recovery 5
   ```
3. Restore fleet/config/workspace from your latest off-node backup (see case 1).
4. Update external DNS/routing.

### 4. Data Loss Recovery

**Symptoms:**
- Missing user data
- Corrupted configurations
- Lost agent workspace state

**Recovery Steps:**
1. List available backups: `ls -la ~/.aleph-deploy/backups/  /opt/openclaw/backups/`
2. Verify a backup's integrity, then restore the needed components (rsync the relevant
   `nodes/<ts>/<node>/workspace` or `fleet/<ts>` directory back to the node).
3. If using the OpenClaw replication layer, re-run a verified replication:
   ```bash
   ssh "$SSH_USER@<primary_ts_ip>" '/opt/openclaw/replication/auto-provisioning-protocol.sh replicate'
   ```
4. Verify data integrity and restart affected services.

### Resource: references/infrastructure-planning-architecture.md

## Contents

- Infrastructure Planning & Architecture
- Aleph Cloud Architecture Overview
- CRN Selection Strategy

## Infrastructure Planning & Architecture

### Aleph Cloud Architecture Overview

**Network Topology:**
```
┌─────────────────────────────────────────────────────────┐
│                   Aleph Cloud Network                   │
├─────────────────┬─────────────────┬─────────────────────┤
│   Primary Node  │  Worker Node 1  │   Worker Node 2     │
│   (Orchestrator)│   (Compute)     │    (Compute)        │
│                 │                 │                     │
│ • Fleet Manager │ • OpenClaw      │  • OpenClaw         │
│ • Load Balancer │ • Tailscale     │  • Tailscale        │
│ • Backup Coord  │ • Health Mon    │  • Health Mon       │
│ • SSH Gateway   │ • Auto-Restart  │  • Auto-Restart     │
└─────────────────┴─────────────────┴─────────────────────┘
         │                 │                 │
         └─────────────────┼─────────────────┘
                  Tailscale Mesh Network
                     SSH Tunnels
```

**Resource Planning Matrix.** Aleph instances are sized in **compute units** (1 CU ≈ 1 vCPU + 2 GiB RAM); you can override with explicit `--vcpus`/`--memory`/`--rootfs-size`. Persistent/confidential VMs run on a specific **CRN** (Compute Resource Node) that you choose by URL or hash.

```yaml
Node Types:
  Orchestrator (Primary):
    Tier: 4 vCPU / 8 GiB RAM / 80–100 GiB rootfs  (≈ 4 compute units)
    CRN: a high-uptime CRN you have verified (see "CRN selection" below)
    Role: fleet manager, HAProxy, backup coordinator, SSH gateway

  Compute Nodes (Workers):
    Tier: 2 vCPU / 4 GiB RAM / 40–50 GiB rootfs  (≈ 2 compute units)
    CRN: spread across 2–3 distinct CRNs for fault isolation
    Role: OpenClaw agent runtime, task execution

  Backup Node (Optional):
    Tier: 1 vCPU / 2 GiB RAM / 20 GiB rootfs  (≈ 1 compute unit)
    CRN: a *different* CRN/region than the primary, for redundancy
    Role: off-node backup target, emergency recovery
```

**Cost model (read this — it changed).** Aleph supports two payment modes, selected with `--payment-type`:

- **`hold`** — lock (don't spend) a quantity of $ALEPH tokens for as long as the VM runs; tokens are released on `delete`. No ongoing burn.
- **`superfluid` / `credit`** — pay-as-you-go streaming (per second) priced in **USD**, settled in $ALEPH or credits. This is the model most users want for fleets.

Do **not** hardcode "ALEPH/month" figures — the token price floats and tiers change. Always read live pricing with the CLI:

```bash
aleph pricing instance                 # all tiers, all payment types
aleph pricing instance --tier 1 --json # one tier, machine-readable
aleph pricing instance --payment-type credit
```

As of **Jun 2026**, pay-as-you-go instance pricing is roughly (confirm with `aleph pricing instance`, do not quote these as fixed):

| Tier | vCPU / RAM / rootfs | Approx. PAYG (USD/hr) | Approx. (USD/mo, 730h) |
|------|---------------------|-----------------------|------------------------|
| 1    | 1 / 2 GiB / 20 GiB  | ~$0.0036              | ~$2.6                  |
| 2    | 2 / 4 GiB / 40 GiB  | ~$0.0066              | ~$4.8                  |
| 3    | 4 / 8 GiB / 80 GiB  | ~$0.0132              | ~$9.6                  |

> These are dated examples for planning only. Confirm current numbers at the Aleph console (https://app.aleph.cloud) or via `aleph pricing instance` before budgeting. A 1 primary + 4 worker fleet on these tiers lands around $30–40/mo PAYG as of Jun 2026 — but verify.

### CRN Selection Strategy

A CRN is the physical node that hosts your persistent/confidential VM. Pick CRNs by **compute availability, payment-mode support, terms acceptance, region, and (for confidential VMs) SEV support** — not by hitting an Aleph API messages endpoint. Discover and inspect CRNs with the CLI rather than guessing URLs:

```bash
#!/bin/bash
# crn-discovery.sh — list and shortlist real CRNs for instance deployment.
set -euo pipefail

echo "=== Available Compute Resource Nodes ==="
# `aleph instance` deployments resolve CRNs from the network; the node index
# is also browsable at https://app.aleph.cloud (Console > Compute) and
# https://docs.aleph.cloud/nodes/compute/ . Prefer the console for capacity,
# version, and reward/uptime score; use --crn-url / --crn-hash from there.

# When creating an instance you may omit --crn-url to let the CLI auto-select
# a CRN, or pin one explicitly. For confidential or Pay-As-You-Go instances a
# CRN is REQUIRED, and you must accept its Terms & Conditions:
#   aleph instance create ... --crn-url "https://<crn-host>" --crn-auto-tac

# Sanity-check a candidate CRN's compute API (this is the CRN's own
# /about endpoint — NOT the Aleph message API):
check_crn() {
    local crn_url="$1" crn_name="$2"
    echo "=== $crn_name ($crn_url) ==="
    echo -n "  Reachable: "
    if curl -fsS --max-time 8 "$crn_url/about/usage/system" >/dev/null 2>&1; then
        echo "yes"
        echo "  Capacity/usage:"
        curl -fsS --max-time 8 "$crn_url/about/usage/system" \
            | jq '{cpu: .cpu, mem: .mem, period}' 2>/dev/null || true
    else
        echo "NO — skip this CRN"
        return 1
    fi
    # Confidential support advertised under /about/capability on SEV-capable CRNs
    echo -n "  Confidential (SEV) capable: "
    curl -fsS --max-time 8 "$crn_url/about/capability" 2>/dev/null \
        | jq -r '.confidential // "unknown"' 2>/dev/null || echo "unknown"
    echo "------------------------"
}

# Replace these with real CRN hosts from https://app.aleph.cloud (Console).
# Do NOT use unrelated services (e.g. storage gateways) as CRNs — they cannot
# host an Aleph instance and `aleph instance create` will fail against them.
# check_crn "https://<crn-1-host>" "CRN 1"
# check_crn "https://<crn-2-host>" "CRN 2"

echo "=== SELECTION GUIDANCE ==="
echo "Primary : highest-uptime CRN with spare capacity and recent node version"
echo "Workers : 2-3 DISTINCT CRNs/regions for fault isolation"
echo "Backup  : a CRN on a different operator/region than the primary"
```

---

### Resource: references/inter-vm-communication-networks.md

## Contents

- Inter-VM Communication Networks
- Tailscale Mesh Network Setup

## Inter-VM Communication Networks

### Tailscale Mesh Network Setup

**Tailscale Integration Script:**
```bash
#!/bin/bash
# setup-tailscale-mesh.sh

set -e

TAILSCALE_AUTH_KEY="${1:-}"
FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"

if [[ -z "$TAILSCALE_AUTH_KEY" ]]; then
    echo "❌ Error: Tailscale auth key required"
    echo "Get your key from: https://login.tailscale.com/admin/settings/keys"
    echo "Usage: $0 <tailscale-auth-key>"
    exit 1
fi

setup_tailscale_node() {
    local node_ip=$1
    local node_name=$2
    local ssh_user="${SSH_USER:-$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG")}"

    echo "Setting up Tailscale on $node_name ($node_ip)..."

    ssh -i ~/.aleph-deploy/keys/aleph_ed25519 -o StrictHostKeyChecking=accept-new \
        "$ssh_user@$node_ip" << TAILSCALE_SETUP
#!/bin/bash
set -euo pipefail

echo "Installing Tailscale..."

# Use the official OS-detecting installer instead of pinning the Ubuntu 22.04
# ("jammy") apt repo — this works on Ubuntu 24.04 and other distros without edits.
installer_1="$(mktemp)"
curl -fsSL https://tailscale.com/install.sh -o "$installer_1"
less "$installer_1"  # Review before execution; verify the release checksum/signature when published.
sh "$installer_1"

# Connect to Tailscale network
rm -f "$installer_1"
# WARNING: Passing --auth-key on the command line exposes it in the process list.
# For production, write the key to a file and use --auth-key=file:/path/to/key
echo "$TAILSCALE_AUTH_KEY" > /tmp/ts-authkey && chmod 600 /tmp/ts-authkey
sudo tailscale up --auth-key="file:/tmp/ts-authkey" --hostname="$node_name"
rm -f /tmp/ts-authkey

# Enable IP forwarding for subnet routing
echo 'net.ipv4.ip_forward = 1' | sudo tee -a /etc/sysctl.conf
echo 'net.ipv6.conf.all.forwarding = 1' | sudo tee -a /etc/sysctl.conf
sudo sysctl -p

# Get Tailscale IP
TAILSCALE_IP=\$(tailscale ip -4)
echo "✅ Tailscale configured. IP: \$TAILSCALE_IP"

# Update local network configuration
cat > /opt/tailscale-info.json << INFO
{
  "tailscale_ip": "\$TAILSCALE_IP",
  "node_name": "$node_name",
  "connected": true,
  "setup_date": "\$(date -Iseconds)"
}
INFO

# Configure Tailscale service for auto-start
sudo systemctl enable tailscaled
sudo systemctl start tailscaled

echo "🎉 Tailscale setup complete on $node_name"
TAILSCALE_SETUP

    echo "✅ Tailscale configured on $node_name"
}

configure_mesh_network() {
    echo "🕸️ Configuring Tailscale mesh network..."

    # Get all fleet nodes
    local primary_ip=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")
    local primary_name=$(jq -r '.primary_node.name' "$FLEET_CONFIG")

    # Setup Tailscale on primary node
    setup_tailscale_node "$primary_ip" "$primary_name"

    # Setup Tailscale on worker nodes
    local workers=$(jq -r '.worker_nodes[] | .name + " " + (.ip // "unknown")' "$FLEET_CONFIG")

    while IFS=' ' read -r worker_name worker_ip; do
        if [[ "$worker_ip" != "unknown" ]]; then
            setup_tailscale_node "$worker_ip" "$worker_name"
        fi
    done <<< "$workers"

    echo "⏳ Waiting for mesh network to stabilize..."
    sleep 30

    # Verify mesh connectivity
    echo "🔍 Verifying mesh connectivity..."
    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$primary_ip" << 'VERIFY'
#!/bin/bash
echo "Testing Tailscale mesh connectivity..."

tailscale status --json | jq -r '.Peer[] | .HostName + " -> " + .TailscaleIPs[0]' | while IFS=' -> ' read -r hostname tailscale_ip; do
    echo -n "Ping $hostname ($tailscale_ip): "
    if ping -c 1 -W 2 "$tailscale_ip" >/dev/null 2>&1; then
        echo "✅ Connected"
    else
        echo "❌ Failed"
    fi
done
VERIFY

    echo "✅ Tailscale mesh network configured"
}

setup_ssh_tunnels() {
    echo "🚇 Setting up SSH tunnels as backup communication..."

    local primary_ip=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")

    # Create SSH tunnel configuration
    cat > ~/.aleph-deploy/configs/ssh-tunnels.conf << 'TUNNEL_CONFIG'
# SSH Tunnel Configuration for Fleet Communication
# Format: LocalPort:RemoteHost:RemotePort

# Fleet Manager Access (Primary -> Workers)
8080:localhost:8080

# OpenClaw gateway access (default gateway port 18789)
18789:localhost:18789

# Health Monitoring
9090:localhost:9090

# Log Aggregation
5514:localhost:514
TUNNEL_CONFIG

    # Setup tunnel management script
    cat > ~/.aleph-deploy/scripts/manage-tunnels.sh << 'TUNNEL_SCRIPT'
#!/bin/bash
# manage-tunnels.sh — SSH tunnels as a BACKUP path when Tailscale is unavailable.
# Prefer the Tailscale mesh; use this only as fallback. Tracks its own PIDs so
# `stop` never kills unrelated SSH sessions belonging to the same user.
set -euo pipefail

TUNNEL_CONFIG="$HOME/.aleph-deploy/configs/ssh-tunnels.conf"
FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"
PID_DIR="$HOME/.aleph-deploy/run/tunnels"
mkdir -p "$PID_DIR"

start_tunnels() {
    local target_ip="$1" target_name="$2"
    echo "Starting SSH tunnels to $target_name ($target_ip)..."
    local last_octet; last_octet="$(echo "$target_ip" | awk -F. '{print $NF+0}')"
    while IFS=':' read -r local_port remote_host remote_port; do
        [[ "$local_port" =~ ^#.*$ || -z "$local_port" ]] && continue
        local unique_port=$((local_port + last_octet))
        # -f backgrounds AFTER auth; capture the resulting PID via a control socket
        # so we can stop exactly this tunnel later (no broad pkill).
        local ctl="$PID_DIR/${target_name}-${unique_port}.ctl"
        ssh -i "$SSH_KEY" -f -N -L "$unique_port:$remote_host:$remote_port" \
            -o StrictHostKeyChecking=accept-new -o ServerAliveInterval=60 \
            -o ControlMaster=yes -o ControlPath="$ctl" \
            "$SSH_USER@$target_ip"
        echo "  Tunnel: localhost:$unique_port -> $target_name:$remote_port (ctl: $ctl)"
    done < "$TUNNEL_CONFIG"
}

stop_tunnels() {
    echo "Stopping SSH tunnels started by this tool..."
    shopt -s nullglob
    for ctl in "$PID_DIR"/*.ctl; do
        # Address the exact control socket; -O exit cleanly closes only that tunnel.
        local host; host="$(basename "$ctl")"
        ssh -O exit -o ControlPath="$ctl" placeholder 2>/dev/null || true
        rm -f "$ctl"
        echo "  closed $host"
    done
}

list_tunnels() {
    echo "Active tunnels (control sockets in $PID_DIR):"
    shopt -s nullglob
    for ctl in "$PID_DIR"/*.ctl; do
        echo -n "  $(basename "$ctl"): "
        ssh -O check -o ControlPath="$ctl" placeholder 2>&1 || echo "stale"
    done
}

case "${1:-start}" in
    "start")
        jq -r '.worker_nodes[] | .name + " " + (.ip // "unknown")' "$FLEET_CONFIG" \
        | while IFS=' ' read -r name ip; do
            [[ "$ip" != "unknown" ]] && start_tunnels "$ip" "$name"
        done
        ;;
    "stop")    stop_tunnels ;;
    "list")    list_tunnels ;;
    "restart") stop_tunnels; sleep 2; "$0" start ;;
    *)
        echo "Usage: $0 {start|stop|list|restart}"
        exit 1
        ;;
esac
TUNNEL_SCRIPT

    chmod +x ~/.aleph-deploy/scripts/manage-tunnels.sh

    echo "✅ SSH tunnel management configured"
}

# Command dispatcher
case "${1:-configure}" in
    "configure")
        configure_mesh_network
        ;;
    "tunnels")
        setup_ssh_tunnels
        ;;
    *)
        echo "Usage: $0 <tailscale-auth-key> [configure|tunnels]"
        echo ""
        echo "Steps:"
        echo "1. Get Tailscale auth key from https://login.tailscale.com/admin/settings/keys"
        echo "2. Run: $0 <auth-key> configure"
        echo "3. Run: $0 <auth-key> tunnels"
        exit 1
        ;;
esac
```

---

### Resource: references/load-distribution-orchestration.md

## Contents

- Load Distribution & Orchestration
- Load Balancer Configuration
- Request Distribution Strategies

## Load Distribution & Orchestration

### Load Balancer Configuration

**HAProxy Load Balancer Setup:**
```bash
#!/bin/bash
# setup-load-balancer.sh
set -euo pipefail

FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"
PRIMARY_IP=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")          # public IP (SSH hop only)
PRIMARY_TS_IP=$(jq -r '.primary_node.tailscale_ip' "$FLEET_CONFIG")

# Stats credentials: generate a random password (never a static one) and store it
# locally so you can look it up. The stats page is bound to Tailscale only.
STATS_USER="${STATS_USER:-admin}"
STATS_PASS="${STATS_PASS:-$(openssl rand -hex 16)}"
echo "HAProxy stats login: $STATS_USER / $STATS_PASS"
echo "STATS_USER=$STATS_USER"$'\n'"STATS_PASS=$STATS_PASS" > ~/.aleph-deploy/configs/haproxy-stats.env
chmod 600 ~/.aleph-deploy/configs/haproxy-stats.env

echo "Setting up HAProxy load balancer..."

# Install HAProxy on primary node. Unquoted heredoc so STATS_* and PRIMARY_TS_IP
# expand HERE into the remote script.
ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER@$PRIMARY_IP" << HAPROXY_SETUP
#!/bin/bash
set -euo pipefail

echo "Installing HAProxy..."
sudo apt-get update
sudo apt-get install -y haproxy

sudo cp /etc/haproxy/haproxy.cfg /etc/haproxy/haproxy.cfg.backup

# Resolve this node's Tailscale IP for the (private) stats listener.
TS_IP="\$(tailscale ip -4 2>/dev/null || echo '${PRIMARY_TS_IP}')"

# TLS: HAProxy terminates HTTPS with a single COMBINED PEM (fullchain + private
# key concatenated, key last) at this path. Drop a real cert here for production:
#   sudo cat fullchain.pem privkey.pem > \$TLS_PEM   # order matters: cert(s) then key
# (Let's Encrypt: \`cat \$LE/fullchain.pem \$LE/privkey.pem\`.) If no cert is present
# we DO NOT open a bogus plaintext :443 — the 443 listener is added only when the
# PEM exists, so the advertised URL matches what actually serves TLS.
TLS_PEM="/etc/haproxy/certs/site.pem"
sudo mkdir -p /etc/haproxy/certs && sudo chmod 700 /etc/haproxy/certs
if [[ -s "\$TLS_PEM" ]]; then
    sudo chmod 600 "\$TLS_PEM"
    TLS_BIND="bind *:443 ssl crt \$TLS_PEM alpn h2,http/1.1"
    echo "TLS cert found at \$TLS_PEM — enabling HTTPS on :443"
else
    TLS_BIND="# bind *:443 ssl crt \$TLS_PEM   # no cert present — HTTPS disabled (drop a combined PEM here to enable)"
    echo "No TLS cert at \$TLS_PEM — serving HTTP only on :80 (HTTPS not advertised)."
fi

# Create HAProxy configuration. Single-quoted inner heredoc keeps HAProxy's own
# \$-free syntax literal; we inject TS_IP / creds via sed right after.
cat > /tmp/haproxy.cfg << 'HAPROXY_CONFIG'
global
    daemon
    user haproxy
    group haproxy
    log stdout local0 info
    chroot /var/lib/haproxy
    stats socket /run/haproxy/admin.sock mode 660 level admin
    stats timeout 30s

defaults
    mode http
    timeout connect 5000ms
    timeout client 50000ms
    timeout server 50000ms
    option httplog
    option dontlognull
    option redispatch
    retries 3

# Statistics interface — bound to the TAILSCALE IP only (never *:9090), with a
# randomly generated password. Reachable only over the private mesh.
listen stats
    bind __TS_IP__:9090
    stats enable
    stats uri /haproxy-stats
    stats realm HAProxy\ Statistics
    stats auth __STATS_USER__:__STATS_PASS__

# Frontend - public entry point. Always listens on :80. The :443 line below is
# injected by sed: a real `bind *:443 ssl crt <combined.pem>` when a cert exists,
# otherwise a commented-out placeholder (so we never expose a plaintext :443 that
# masquerades as HTTPS). See the cert-provisioning note above.
frontend openclaw_frontend
    bind *:80
    __TLS_BIND__

    # Health check endpoint (matches the fleet manager's UNAUTHENTICATED /health)
    monitor-uri /health

    default_backend openclaw_nodes

# Backend - OpenClaw nodes
backend openclaw_nodes
    balance roundrobin
    option httpchk GET /health

    # Health check configuration
    default-server check maxconn 50 rise 2 fall 3 inter 2s

    # Primary node (higher weight)
    # NOTE: the port MUST match the service actually exposed on the node: the
    # OpenClaw gateway's configured port (default 18789, loopback-bound until
    # you bind it to a reachable interface) or your own app's port.
    server primary-node localhost:3000 weight 150 check

    # Worker nodes will be added dynamically
HAPROXY_CONFIG

# Inject the Tailscale IP, stats credentials, and the TLS bind line (use | as the
# sed delimiter since values contain no pipes; credentials were generated, not
# hardcoded). __TLS_BIND__ becomes a real ssl bind only when a cert exists.
sed -i "s|__TS_IP__|\${TS_IP}|; s|__STATS_USER__|${STATS_USER}|; s|__STATS_PASS__|${STATS_PASS}|; s|__TLS_BIND__|\${TLS_BIND}|" /tmp/haproxy.cfg

# Validate the config BEFORE replacing the live one (avoids a broken restart).
if sudo haproxy -c -f /tmp/haproxy.cfg; then
    sudo mv /tmp/haproxy.cfg /etc/haproxy/haproxy.cfg
    sudo systemctl enable haproxy
    sudo systemctl restart haproxy
    if [[ -s "\$TLS_PEM" ]]; then
        echo "HAProxy installed: HTTP on :80, HTTPS on :443 (cert \$TLS_PEM); stats on \${TS_IP}:9090 (Tailscale only)"
    else
        echo "HAProxy installed: HTTP on :80 only (no TLS cert); stats on \${TS_IP}:9090 (Tailscale only)"
    fi
else
    echo "HAProxy config invalid — not applying."; exit 1
fi
HAPROXY_SETUP

echo "Configuring dynamic backend management..."

# Create backend management script
ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER@$PRIMARY_IP" << 'BACKEND_SCRIPT'
#!/bin/bash

cat > /opt/manage-haproxy-backends.sh << 'MANAGE_BACKENDS'
#!/bin/bash

HAPROXY_STATS_SOCKET="/run/haproxy/admin.sock"

# Control-plane access: the fleet manager listens on the node's Tailscale IP and
# requires the shared API key. Both come from the root-owned EnvironmentFile that
# the fleet manager also uses (FLEET_API_KEY, BIND_HOST).
[[ -f /etc/fleet-manager.env ]] && { set -a; . /etc/fleet-manager.env; set +a; }
FLEET_MGR_HOST="${BIND_HOST:-127.0.0.1}"
: "${FLEET_API_KEY:?FLEET_API_KEY not found in /etc/fleet-manager.env}"

add_backend_server() {
    local server_name=$1
    local server_ip=$2
    # Default port must match the service actually exposed on the worker (the
    # OpenClaw gateway's configured port, default 18789, or your own app).
    local server_port=${3:-3000}
    local weight=${4:-100}

    echo "Adding backend server: $server_name ($server_ip:$server_port)"

    # Add server to HAProxy backend
    echo "add server openclaw_nodes/$server_name $server_ip:$server_port weight $weight check" | \
        sudo socat stdio "$HAPROXY_STATS_SOCKET"

    echo "✅ Server $server_name added to load balancer"
}

remove_backend_server() {
    local server_name=$1

    echo "Removing backend server: $server_name"

    # Disable server first
    echo "disable server openclaw_nodes/$server_name" | sudo socat stdio "$HAPROXY_STATS_SOCKET"

    # Remove server from backend
    echo "del server openclaw_nodes/$server_name" | sudo socat stdio "$HAPROXY_STATS_SOCKET"

    echo "✅ Server $server_name removed from load balancer"
}

list_backend_servers() {
    echo "📋 Current backend servers:"
    echo "show servers state openclaw_nodes" | sudo socat stdio "$HAPROXY_STATS_SOCKET"
}

update_server_weight() {
    local server_name=$1
    local new_weight=$2

    echo "Updating weight for $server_name to $new_weight"
    echo "set weight openclaw_nodes/$server_name $new_weight" | sudo socat stdio "$HAPROXY_STATS_SOCKET"
}

sync_with_fleet() {
    echo "🔄 Syncing backends with fleet registry..."

    # Get current fleet status (over Tailscale, authenticated)
    local fleet_nodes=$(curl -fsS -H "x-api-key: $FLEET_API_KEY" "http://$FLEET_MGR_HOST:8080/fleet/status" | jq -r '.nodes[] | .node_id + "," + .ip_address + "," + .status')

    # Get current HAProxy backends
    local current_backends=$(echo "show servers state openclaw_nodes" | sudo socat stdio "$HAPROXY_STATS_SOCKET" | awk '{print $4}' | grep -v "#" | sort)

    # Add new nodes to HAProxy
    while IFS=',' read -r node_id ip_address status; do
        if [[ "$status" == "active" && "$node_id" != "primary" ]]; then
            # Check if server already exists in HAProxy
            if ! echo "$current_backends" | grep -q "$node_id"; then
                add_backend_server "$node_id" "$ip_address" 3000 100
            fi
        fi
    done <<< "$fleet_nodes"

    # Remove offline nodes from HAProxy
    echo "$current_backends" | while read -r backend_name; do
        [[ -z "$backend_name" ]] && continue

        # Check if this backend still exists in fleet
        if ! echo "$fleet_nodes" | grep -q "$backend_name,"; then
            echo "⚠️  Backend $backend_name not found in fleet, removing..."
            remove_backend_server "$backend_name"
        fi
    done

    echo "✅ Backend synchronization complete"
}

# Auto-sync with fleet every 60 seconds
auto_sync() {
    while true; do
        sync_with_fleet
        sleep 60
    done
}

case "${1:-sync}" in
    "add")
        add_backend_server "$2" "$3" "$4" "$5"
        ;;
    "remove")
        remove_backend_server "$2"
        ;;
    "list")
        list_backend_servers
        ;;
    "weight")
        update_server_weight "$2" "$3"
        ;;
    "sync")
        sync_with_fleet
        ;;
    "auto")
        auto_sync
        ;;
    *)
        echo "Usage: $0 {add|remove|list|weight|sync|auto}"
        echo ""
        echo "Commands:"
        echo "  add <name> <ip> [port] [weight] - Add backend server"
        echo "  remove <name>                   - Remove backend server"
        echo "  list                            - List all backend servers"
        echo "  weight <name> <weight>          - Update server weight"
        echo "  sync                            - Sync with fleet registry"
        echo "  auto                            - Auto-sync daemon"
        exit 1
        ;;
esac
MANAGE_BACKENDS

chmod +x /opt/manage-haproxy-backends.sh

# Install socat for HAProxy socket communication
sudo apt-get install -y socat

# Create systemd service for auto-sync
cat > /etc/systemd/system/haproxy-fleet-sync.service << 'SYNC_SERVICE'
[Unit]
Description=HAProxy Fleet Synchronization
After=haproxy.service fleet-manager.service

[Service]
Type=simple
User=root
EnvironmentFile=/etc/fleet-manager.env
ExecStart=/opt/manage-haproxy-backends.sh auto
Restart=always
RestartSec=30

[Install]
WantedBy=multi-user.target
SYNC_SERVICE

sudo systemctl daemon-reload
sudo systemctl enable haproxy-fleet-sync
sudo systemctl start haproxy-fleet-sync

echo "HAProxy backend management configured"
BACKEND_SCRIPT

echo "Load balancer setup complete."
# Only advertise HTTPS if the combined PEM is actually present on the primary
# (the same condition the HAProxy config uses to add the :443 ssl bind).
if ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER@$PRIMARY_IP" \
       "test -s /etc/haproxy/certs/site.pem" 2>/dev/null; then
    echo "Public load balancer: https://$PRIMARY_IP  (HTTP on http://$PRIMARY_IP)"
else
    echo "Public load balancer: http://$PRIMARY_IP"
    echo "  (HTTPS not enabled — add a combined fullchain+key PEM at"
    echo "   /etc/haproxy/certs/site.pem on the primary and re-run to serve TLS on :443.)"
fi
echo "HAProxy stats (Tailscale only): http://$PRIMARY_TS_IP:9090/haproxy-stats"
echo "Stats login is in ~/.aleph-deploy/configs/haproxy-stats.env"
```

### Request Distribution Strategies

**Load Distribution Algorithm:**
```bash
#!/bin/bash
# intelligent-load-distribution.sh

FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"
PRIMARY_IP=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")

setup_intelligent_distribution() {
    echo "🧠 Setting up intelligent load distribution..."

    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER@$PRIMARY_IP" << 'DISTRIBUTION_SETUP'
#!/bin/bash
set -euo pipefail

# Node.js 22.x (OpenClaw and our tooling require Node >= 22.19)
curl -fsSL https://deb.nodesource.com/setup_22.x | sudo -E bash -
sudo apt-get install -y nodejs

# Create intelligent distribution service
mkdir -p /opt/load-distributor
cd /opt/load-distributor

cat > intelligent-distributor.js << 'DISTRIBUTOR_JS'
const express = require('express');
const axios = require('axios');

const app = express();
app.use(express.json());

// Fleet manager URL + API key come from the environment (set by the systemd unit
// via /etc/fleet-manager.env). All control-plane calls MUST send x-api-key.
const FLEET_API_KEY = process.env.FLEET_API_KEY;
const FLEET_MGR_URL = `http://${process.env.BIND_HOST || '127.0.0.1'}:8080`;
if (!FLEET_API_KEY) { console.error('FATAL: FLEET_API_KEY missing'); process.exit(1); }
const fleet = axios.create({ baseURL: FLEET_MGR_URL, headers: { 'x-api-key': FLEET_API_KEY }, timeout: 5000 });

class IntelligentDistributor {
    constructor() {
        this.nodes = new Map();
        this.requestHistory = [];
        this.loadMetrics = new Map();

        // Load balancing strategies
        this.strategies = {
            'round_robin': this.roundRobin.bind(this),
            'least_connections': this.leastConnections.bind(this),
            'weighted_response_time': this.weightedResponseTime.bind(this),
            'resource_aware': this.resourceAware.bind(this),
            'session_affinity': this.sessionAffinity.bind(this)
        };

        this.currentStrategy = 'resource_aware';
        this.updateMetrics();
    }

    async updateMetrics() {
        try {
            // Get fleet status (authenticated, over Tailscale)
            const fleetResponse = await fleet.get('/fleet/status');
            const nodes = fleetResponse.data.nodes || [];

            // Update node metrics
            for (const node of nodes) {
                if (node.status === 'active') {
                    const metrics = await this.collectNodeMetrics(node);
                    this.loadMetrics.set(node.node_id, metrics);
                }
            }
        } catch (error) {
            console.error('Error updating metrics:', error.message);
        }

        // Schedule next update
        setTimeout(() => this.updateMetrics(), 30000); // 30 seconds
    }

    async collectNodeMetrics(node) {
        try {
            // Mock metrics collection - replace with actual implementation
            return {
                cpu_usage: Math.random() * 100,
                memory_usage: Math.random() * 100,
                active_connections: Math.floor(Math.random() * 50),
                avg_response_time: Math.random() * 1000,
                error_rate: Math.random() * 0.1,
                last_updated: new Date().toISOString()
            };
        } catch (error) {
            console.error(`Error collecting metrics for ${node.node_id}:`, error.message);
            return null;
        }
    }

    // Round Robin Strategy
    roundRobin(availableNodes) {
        if (!this.roundRobinIndex || this.roundRobinIndex >= availableNodes.length) {
            this.roundRobinIndex = 0;
        }
        return availableNodes[this.roundRobinIndex++];
    }

    // Least Connections Strategy
    leastConnections(availableNodes) {
        let selectedNode = availableNodes[0];
        let minConnections = Infinity;

        for (const node of availableNodes) {
            const metrics = this.loadMetrics.get(node.node_id);
            if (metrics && metrics.active_connections < minConnections) {
                minConnections = metrics.active_connections;
                selectedNode = node;
            }
        }

        return selectedNode;
    }

    // Weighted Response Time Strategy
    weightedResponseTime(availableNodes) {
        let selectedNode = availableNodes[0];
        let minResponseTime = Infinity;

        for (const node of availableNodes) {
            const metrics = this.loadMetrics.get(node.node_id);
            if (metrics && metrics.avg_response_time < minResponseTime) {
                minResponseTime = metrics.avg_response_time;
                selectedNode = node;
            }
        }

        return selectedNode;
    }

    // Resource Aware Strategy (CPU + Memory + Response Time)
    resourceAware(availableNodes) {
        let selectedNode = availableNodes[0];
        let bestScore = Infinity;

        for (const node of availableNodes) {
            const metrics = this.loadMetrics.get(node.node_id);
            if (metrics) {
                // Calculate composite score (lower is better)
                const score = (
                    metrics.cpu_usage * 0.4 +
                    metrics.memory_usage * 0.3 +
                    (metrics.avg_response_time / 10) * 0.2 +
                    metrics.error_rate * 100 * 0.1
                );

                if (score < bestScore) {
                    bestScore = score;
                    selectedNode = node;
                }
            }
        }

        return selectedNode;
    }

    // Session Affinity Strategy
    sessionAffinity(availableNodes, sessionId) {
        if (!sessionId) return this.resourceAware(availableNodes);

        // Simple hash-based affinity
        const hash = this.simpleHash(sessionId);
        const nodeIndex = hash % availableNodes.length;
        return availableNodes[nodeIndex];
    }

    simpleHash(str) {
        let hash = 0;
        for (let i = 0; i < str.length; i++) {
            const char = str.charCodeAt(i);
            hash = ((hash << 5) - hash) + char;
            hash = hash & hash; // Convert to 32-bit integer
        }
        return Math.abs(hash);
    }

    async selectNode(requestInfo = {}) {
        try {
            // Get available nodes (authenticated)
            const fleetResponse = await fleet.get('/fleet/status');
            const availableNodes = fleetResponse.data.nodes.filter(n => n.status === 'active');

            if (availableNodes.length === 0) {
                throw new Error('No available nodes');
            }

            // Apply distribution strategy
            const strategy = this.strategies[this.currentStrategy];
            const selectedNode = strategy(availableNodes, requestInfo.sessionId);

            // Log request for analysis
            this.requestHistory.push({
                timestamp: new Date().toISOString(),
                selected_node: selectedNode.node_id,
                strategy: this.currentStrategy,
                request_info: requestInfo
            });

            // Keep only last 1000 requests
            if (this.requestHistory.length > 1000) {
                this.requestHistory = this.requestHistory.slice(-1000);
            }

            return selectedNode;

        } catch (error) {
            console.error('Error selecting node:', error.message);
            throw error;
        }
    }
}

const distributor = new IntelligentDistributor();

// API Endpoints
app.get('/distribute/node', async (req, res) => {
    try {
        const requestInfo = {
            sessionId: req.headers['x-session-id'],
            requestType: req.query.type,
            clientIp: req.ip
        };

        const selectedNode = await distributor.selectNode(requestInfo);
        res.json({
            node_id: selectedNode.node_id,
            ip_address: selectedNode.ip_address,
            strategy: distributor.currentStrategy
        });

    } catch (error) {
        res.status(500).json({ error: error.message });
    }
});

app.get('/distribute/metrics', (req, res) => {
    const metrics = {};
    distributor.loadMetrics.forEach((value, key) => {
        metrics[key] = value;
    });
    res.json(metrics);
});

app.get('/distribute/history', (req, res) => {
    res.json(distributor.requestHistory.slice(-100)); // Last 100 requests
});

app.post('/distribute/strategy', (req, res) => {
    const { strategy } = req.body;
    if (distributor.strategies[strategy]) {
        distributor.currentStrategy = strategy;
        res.json({ success: true, strategy });
    } else {
        res.status(400).json({ error: 'Invalid strategy' });
    }
});

const PORT = 8081;
// Bind to localhost only — this is an internal control API consumed by the
// primary's own routing logic, not a public endpoint.
app.listen(PORT, '127.0.0.1', () => {
    console.log(`Intelligent Load Distributor on 127.0.0.1:${PORT}`);
});
DISTRIBUTOR_JS

# Install dependencies
npm init -y
npm install express axios

# Create systemd service
cat > /etc/systemd/system/load-distributor.service << 'DISTRIBUTOR_SERVICE'
[Unit]
Description=Intelligent Load Distributor
After=network.target fleet-manager.service

[Service]
Type=simple
User=root
WorkingDirectory=/opt/load-distributor
EnvironmentFile=/etc/fleet-manager.env
ExecStart=/usr/bin/node intelligent-distributor.js
Restart=always
RestartSec=10
Environment=NODE_ENV=production

[Install]
WantedBy=multi-user.target
DISTRIBUTOR_SERVICE

sudo systemctl daemon-reload
sudo systemctl enable load-distributor
sudo systemctl start load-distributor

echo "Intelligent load distributor configured (localhost:8081, internal only)"
DISTRIBUTION_SETUP

echo "Intelligent load distribution setup complete."
echo "Distribution API is internal (localhost:8081 on the primary)."
echo "From the primary: curl http://127.0.0.1:8081/distribute/node"
}

# NOTE: collectNodeMetrics() returns randomized placeholders; see the Metrics note below.

# Execute setup
setup_intelligent_distribution
```

> **Metrics note.** `collectNodeMetrics()` above returns **randomized placeholder values** so the strategy code is runnable out of the box. For real distribution, replace it with actual per-node metrics, e.g. scrape `node_exporter`/cAdvisor over the Tailscale mesh, or have each worker POST CPU/mem/conn counts to the fleet manager. See the sibling `monitoring-observability` skill for a production metrics pipeline.

---

### Resource: references/monitoring-maintenance.md

## Contents

- Monitoring & Maintenance
- Routine Maintenance Checklist
- Quick Reference Commands
- Troubleshooting

## Monitoring & Maintenance

### Routine Maintenance Checklist

**Daily:**
- Check fleet status: `./fleet-control.sh status`
- Review backup logs: `tail /var/log/backup.log`
- Check security events: `tail /var/log/security-events.log`

**Weekly:**
- Review cost reports: `ls ~/.aleph-deploy/reports/`
- Check node health: `./fleet-control.sh health`
- Verify backup integrity: run a test restore on staging

**Monthly / as needed:**
- Update system packages: `./fleet-control.sh deploy update-packages.sh`
- Re-check CRN pricing and availability: `aleph pricing instance`
- Rotate `FLEET_API_KEY` if a node/operator may be compromised (regenerate, update `/etc/fleet-manager.env` on the primary, restart fleet-manager/sync/distributor)

**On a security event (not on a fixed schedule):**
- Rotate SSH keys: `~/.aleph-deploy/scripts/rotate-ssh-keys.sh rotate` (verify-before-activate; old key kept until new one is proven)

### Quick Reference Commands

```bash
# Fleet operations
./fleet-control.sh status        # View fleet status
./fleet-control.sh health        # Health check all nodes
./fleet-control.sh restart openclaw  # Restart service on all nodes
./fleet-control.sh logs openclaw 100 # Collect last 100 log lines

# Backup & Recovery
ssh root@PRIMARY_IP '/opt/openclaw/backup-system.sh full'
ssh root@PRIMARY_IP '/opt/openclaw/backup-system.sh snapshot'

# Security
~/.aleph-deploy/scripts/security-status.sh
~/.aleph-deploy/scripts/rotate-ssh-keys.sh rotate

# Cost monitoring
~/.aleph-deploy/scripts/cost-monitor.sh

# Auto-scaling (enable/disable)
ssh root@PRIMARY_IP 'sudo systemctl enable auto-scaler && sudo systemctl start auto-scaler'
ssh root@PRIMARY_IP 'sudo systemctl stop auto-scaler && sudo systemctl disable auto-scaler'

# Replication
ssh root@PRIMARY_IP '/opt/openclaw/replication/auto-provisioning-protocol.sh replicate'
ssh root@PRIMARY_IP '/opt/openclaw/replication/auto-provisioning-protocol.sh emergency manual'

# Tailscale mesh
ssh root@PRIMARY_IP 'tailscale status'
```

### Troubleshooting

| Problem | Cause | Fix |
|---------|-------|-----|
| Fleet manager 401 | Missing x-api-key header | Add `-H "x-api-key: $FLEET_API_KEY"` to curl calls |
| Worker can't register | Fleet manager not reachable | Check Tailscale connectivity and UFW rules |
| nodes.json ENOENT | File not created before service start | Create `echo '{"nodes":[]}' > /opt/fleet-manager/nodes.json` and restart |
| HAProxy backend stale | Fleet sync not running | Check `systemctl status haproxy-fleet-sync` |
| SSH key rotation fails | New key not propagated | Old key still works (rotation is verify-before-activate); re-run `rotate-ssh-keys.sh rotate`, or manually append: `ssh-copy-id -i KEY "$SSH_USER@NODE"` |
| Auto-scaler variables lost | Pipe subshell scoping | Use `while read ... done < <(cmd)` process substitution |
| Replication files missing | Wrong extract paths | Files are under `soul/`, `agents/`, `memory/` subdirectories |
| High CPU but no scale-up | Cooldown period active | Wait 5 minutes or reset `/tmp/last-scale-action` |

### Resource: references/multi-node-fleet-management.md

## Contents

- Multi-Node Fleet Management
- Fleet Deployment Orchestrator
- Fleet Management Commands

## Multi-Node Fleet Management

### Fleet Deployment Orchestrator

**Master Deployment Script:**
**Before you run this:** generate ONE persistent `FLEET_API_KEY` locally and export it. Both the deploy script and the fleet manager must use the *same* key, and it must survive restarts (the manager must not invent a new random key each boot).

```bash
# Generate once and store it safely (NOT in git, NOT in shell history files):
export FLEET_API_KEY="$(openssl rand -hex 32)"
echo "FLEET_API_KEY=$FLEET_API_KEY" >> ~/.aleph-deploy/configs/fleet.env   # chmod 600 this file
chmod 600 ~/.aleph-deploy/configs/fleet.env
```

```bash
#!/bin/bash
# deploy-fleet.sh
set -euo pipefail

# Fleet Configuration
FLEET_NAME="${1:-openclaw-fleet}"
NODE_COUNT="${2:-5}"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="${ALEPH_SSH_USER:-root}"
: "${FLEET_API_KEY:?Set FLEET_API_KEY (see fleet.env) before deploying}"

# Pin CRNs you have ACTUALLY verified with crn-discovery.sh. Leave empty to let
# the CLI auto-select. Never list non-compute services here (a storage gateway
# or NFT pinning API is NOT a CRN and cannot host an instance).
PRIMARY_CRN="${PRIMARY_CRN:-}"                 # e.g. https://<verified-crn-host>
WORKER_CRNS=(${WORKER_CRNS:-})                 # e.g. ("https://<crn-a>" "https://<crn-b>")

echo "Deploying fleet: $FLEET_NAME with $NODE_COUNT nodes"

# Fleet configuration. worker_nodes entries WILL record ip + item_hash (added at
# create time) so networking/backup/security scripts can find every worker.
cat > ~/.aleph-deploy/configs/fleet.json << EOF
{
  "fleet_name": "$FLEET_NAME",
  "deployment_date": "$(date -Iseconds)",
  "node_count": $NODE_COUNT,
  "ssh_user": "$SSH_USER",
  "primary_node": null,
  "worker_nodes": [],
  "network": {
    "ssh_tunnel_port": 2222,
    "load_balancer_port": 8080
  },
  "replication": {
    "enabled": true,
    "sync_interval": 300,
    "backup_retention": 7
  }
}
EOF

deploy_primary_node() {
    echo "📊 Deploying Primary Node (Orchestrator)..."

    local node_name="${FLEET_NAME}-primary"
    # The primary setup script is parameterized with the (persistent) fleet key so
    # the manager and workers share ONE key. We export it into the heredoc env.
    local setup_script
    setup_script=$(FLEET_API_KEY="$FLEET_API_KEY" envsubst '$FLEET_API_KEY' << 'PRIMARY_SETUP'
#!/bin/bash
set -euo pipefail
export DEBIAN_FRONTEND=noninteractive

# Standard VM setup. Modern package set: Docker Engine + Compose v2 plugin
# (installed via get.docker.com below), Node 22 via NodeSource, iproute2 for `ss`.
apt-get update && apt-get -y upgrade
apt-get install -y curl wget git htop jq fail2ban ufw ca-certificates iproute2 gettext-base

curl -fsSL https://get.docker.com -o /tmp/get-docker.sh && sh /tmp/get-docker.sh
installer_1="$(mktemp)"
curl -fsSL https://deb.nodesource.com/setup_22.x -o "$installer_1"
less "$installer_1"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_1" - && apt-get install -y nodejs

# Create a dedicated non-root user for fleet services. Running all services as
rm -f "$installer_1"
# root is a security risk — a compromise in any service gives full system access.
useradd -r -s /usr/sbin/nologin -d /opt/fleet-manager fleetmgr || true

# Install fleet management tools
mkdir -p /opt/fleet-manager
cd /opt/fleet-manager

# Fleet Manager Application
cat > fleet-manager.js << 'FLEET_MANAGER'
const express = require('express');
const fs = require('fs');
const app = express();

app.use(express.json());

// API key auth. The key MUST be provided via the environment (root-owned
// EnvironmentFile, see below) so it is stable across restarts and is never
// generated/logged. Fail fast if it is missing rather than minting a random one.
const FLEET_API_KEY = process.env.FLEET_API_KEY;
if (!FLEET_API_KEY || FLEET_API_KEY.length < 32) {
    console.error('FATAL: FLEET_API_KEY env var missing or too short. Refusing to start.');
    process.exit(1);
}
// Constant-time comparison; header-only (never accept keys in the query string —
// URLs are logged and cached, leaking the secret).
const crypto = require('crypto');
function keyMatches(provided) {
    if (typeof provided !== 'string') return false;
    const a = Buffer.from(provided);
    const b = Buffer.from(FLEET_API_KEY);
    return a.length === b.length && crypto.timingSafeEqual(a, b);
}
function requireAuth(req, res, next) {
    if (!keyMatches(req.headers['x-api-key'])) {
        return res.status(401).json({ error: 'Unauthorized' });
    }
    next();
}

// Health check FIRST and UNAUTHENTICATED — HAProxy/`option httpchk` calls this
// without an API key. Keep it non-sensitive (no node data).
app.get('/health', (req, res) => {
    res.json({ status: 'healthy', timestamp: new Date().toISOString() });
});

// Everything below requires the API key.
app.use(requireAuth);

// Fleet status endpoint
app.get('/fleet/status', (req, res) => {
    try {
        const data = fs.readFileSync('/opt/fleet-manager/nodes.json', 'utf8');
        res.json(JSON.parse(data));
    } catch (err) {
        if (err.code === 'ENOENT') {
            res.json({ nodes: [] });
        } else {
            res.status(500).json({ error: err.message });
        }
    }
});

// (Health check is defined above, before requireAuth, so HAProxy/httpchk can
// reach it without an API key. Do not re-add an authenticated /health here.)

// Node registration endpoint
app.post('/fleet/register', (req, res) => {
    const { node_id, ip_address, capabilities, item_hash } = req.body;

    let fleet;
    try {
        fleet = JSON.parse(fs.readFileSync('/opt/fleet-manager/nodes.json', 'utf8'));
    } catch {
        fleet = { nodes: [] };
    }

    // Update or add node. We persist item_hash (the Aleph instance hash captured
    // at create time) so the autoscale/auto-recreate paths can delete/recreate
    // this exact instance later. Heartbeats omit item_hash, so on re-register we
    // preserve whatever hash we already stored for this node.
    const existingIndex = fleet.nodes.findIndex(n => n.node_id === node_id);
    const prior = existingIndex >= 0 ? fleet.nodes[existingIndex] : {};
    const nodeData = {
        node_id,
        ip_address,
        capabilities,
        item_hash: item_hash || prior.item_hash || null,
        last_seen: new Date().toISOString(),
        status: 'active'
    };

    if (existingIndex >= 0) {
        fleet.nodes[existingIndex] = nodeData;
    } else {
        fleet.nodes.push(nodeData);
    }

    fs.writeFileSync('/opt/fleet-manager/nodes.json', JSON.stringify(fleet, null, 2));
    res.json({ success: true });
});

// Load distribution endpoint
app.get('/fleet/distribute/:task', (req, res) => {
    const task = req.params.task;
    let nodes;
    try {
        nodes = JSON.parse(fs.readFileSync('/opt/fleet-manager/nodes.json', 'utf8'));
    } catch {
        nodes = { nodes: [] };
    }

    // Simple round-robin distribution
    const activeNodes = nodes.nodes.filter(n => n.status === 'active');
    if (activeNodes.length === 0) {
        return res.status(503).json({ error: 'No active nodes available' });
    }

    const assignedNode = activeNodes[Math.floor(Math.random() * activeNodes.length)];
    res.json({
        task,
        assigned_node: assignedNode.node_id,
        node_ip: assignedNode.ip_address
    });
});

const PORT = process.env.PORT || 8080;
// Bind to the Tailscale interface (or localhost) — NEVER 0.0.0.0. The systemd
// unit sets BIND_HOST to the node's Tailscale IP so workers on the mesh can
// register, while the public internet cannot reach the control plane.
const BIND_HOST = process.env.BIND_HOST || '127.0.0.1';
app.listen(PORT, BIND_HOST, () => {
    console.log(`Fleet Manager listening on ${BIND_HOST}:${PORT}`);
});
FLEET_MANAGER

# Install dependencies and start fleet manager
npm init -y
npm install express
chmod +x fleet-manager.js

# Provision the SHARED, PERSISTENT FLEET_API_KEY via a root-owned EnvironmentFile.
# The key was injected into this setup script by deploy-fleet.sh (envsubst) and is
# never logged. BIND_HOST is resolved to the Tailscale IP after the mesh is up
# (a drop-in updates it; until then it stays on localhost).
install -o root -g root -m 600 /dev/null /etc/fleet-manager.env
{
  echo "FLEET_API_KEY=${FLEET_API_KEY}"
  echo "PORT=8080"
  echo "BIND_HOST=127.0.0.1"
} > /etc/fleet-manager.env

# Create systemd service
cat > /etc/systemd/system/fleet-manager.service << 'SERVICE'
[Unit]
Description=OpenClaw Fleet Manager
After=network.target

[Service]
Type=simple
User=fleetmgr
WorkingDirectory=/opt/fleet-manager
EnvironmentFile=/etc/fleet-manager.env
ExecStart=/usr/bin/node fleet-manager.js
Restart=always
RestartSec=10
# Harden: no new privileges, read-only system except its own dir.
NoNewPrivileges=true
ProtectSystem=strict
ReadWritePaths=/opt/fleet-manager

[Install]
WantedBy=multi-user.target
SERVICE

# Set ownership so fleetmgr user can read/write
chown -R fleetmgr:fleetmgr /opt/fleet-manager

# Initialize nodes registry BEFORE starting fleet-manager.
# fleet-manager.js reads this file on startup — if it doesn't exist,
# the readFileSync call will throw ENOENT and crash the service.
echo '{"nodes": []}' > /opt/fleet-manager/nodes.json
chown fleetmgr:fleetmgr /opt/fleet-manager/nodes.json

systemctl daemon-reload
systemctl enable fleet-manager
systemctl start fleet-manager

# Install OpenClaw on the primary (official installer + onboarding daemon).
# Docs: https://docs.openclaw.ai/install . Requires Node >= 22.19 (installed above).
installer_2="$(mktemp)"
curl -fsSL https://openclaw.ai/install.sh -o "$installer_2"
less "$installer_2"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_2"
# `openclaw onboard --install-daemon` is interactive; run it manually (or with a
rm -f "$installer_2"
# pre-seeded config/secret store) to create the systemd daemon. Do NOT hand-write
# /opt/openclaw/config/*.json — OpenClaw manages its own config via onboard.

echo "Primary node base setup complete (fleet-manager active on Tailscale)."
PRIMARY_SETUP
    )

    # Create the instance with CURRENT flags, then provision over SSH (the Python
    # aleph-client has no --setup-script / --image-ref / --disk-size / --crn; the
    # rewritten aleph-cli does have --disk-size). See `... create --help`.
    local create_args=(
        --name "$node_name"
        --compute-units 4 --memory 8192
        --rootfs-size 81920
        --ssh-pubkey-file "$SSH_KEY.pub"
        --payment-type credit --payment-chain BASE
        --persistent-volume "name=fleet,mount=/opt/fleet-manager,size_mib=10240"
    )
    [[ -n "$PRIMARY_CRN" ]] && create_args+=(--crn-url "$PRIMARY_CRN" --crn-auto-tac)

    local out item_hash primary_ip
    out="$(aleph instance create "${create_args[@]}")"
    echo "$out"
    item_hash="$(printf '%s\n' "$out" | grep -oE '[0-9a-f]{64}' | head -1)"
    primary_ip="$(wait_for_ip "$node_name")" || { echo "Primary got no IP"; return 1; }

    # Provision over SSH using the injected, persistent FLEET_API_KEY.
    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER@$primary_ip" \
        "FLEET_API_KEY='$FLEET_API_KEY' bash -s" <<< "$setup_script"

    # Record name, IP, and item_hash so every later script can reach/destroy it.
    local tmpfile; tmpfile="$(mktemp)"
    jq --arg n "$node_name" --arg ip "$primary_ip" --arg h "$item_hash" \
       '.primary_node = {name:$n, ip:$ip, item_hash:$h}' \
       ~/.aleph-deploy/configs/fleet.json > "$tmpfile"
    mv "$tmpfile" ~/.aleph-deploy/configs/fleet.json

    echo "Primary node deployed: $primary_ip ($item_hash)"
    return 0
}

# Poll the REAL `aleph instance list` for a named instance's IP (no fake commands).
wait_for_ip() {
    local name="$1" ip=""
    for _ in $(seq 1 30); do
        ip="$(aleph instance list --json \
            | jq -r --arg n "$name" '.[] | select(.name==$n) | (.ipv4 // .ipv6 // empty)' \
            | head -1)"
        [[ -n "$ip" ]] && { echo "$ip"; return 0; }
        sleep 10
    done
    return 1
}

deploy_worker_node() {
    local node_id="$1" crn_url="$2" primary_ip="$3"
    local node_name="${FLEET_NAME}-worker-${node_id}"
    echo "Deploying worker node $node_id ($node_name)..."

    # Worker provisioning script. The worker JOINS the Tailscale mesh FIRST, then
    # registers with the primary over that mesh (primary_tailscale_ip), so the
    # address it registers is always its reachable Tailscale IP — never a
    # firewalled public/private address. We pass primary's Tailscale IP, the
    # shared key, the Tailscale auth key, and the instance ITEM_HASH in as env
    # vars at SSH time.
    local setup_script
    setup_script=$(cat <<'WORKER_SETUP'
#!/bin/bash
set -euo pipefail
export DEBIAN_FRONTEND=noninteractive

apt-get update && apt-get -y upgrade
apt-get install -y curl wget git htop jq ca-certificates iproute2
curl -fsSL https://get.docker.com -o /tmp/get-docker.sh && sh /tmp/get-docker.sh
installer_3="$(mktemp)"
curl -fsSL https://deb.nodesource.com/setup_22.x -o "$installer_3"
less "$installer_3"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_3" - && apt-get install -y nodejs

# Join the Tailscale mesh BEFORE registering, so `tailscale ip -4` returns a
rm -f "$installer_3"
# reachable mesh address (the control plane is only reachable over the mesh).
# Official Tailscale installer; review first via:
#   curl -fsSL https://tailscale.com/install.sh -o /tmp/ts-install.sh && less /tmp/ts-install.sh
installer_4="$(mktemp)"
curl -fsSL https://tailscale.com/install.sh -o "$installer_4"
less "$installer_4"  # Review before execution; verify the release checksum/signature when published.
sh "$installer_4"
rm -f "$installer_4"
: "${TAILSCALE_AUTH_KEY:?TAILSCALE_AUTH_KEY required to join the mesh before registering}"
printf '%s' "$TAILSCALE_AUTH_KEY" > /tmp/ts && chmod 600 /tmp/ts
tailscale up --auth-key="file:/tmp/ts" --hostname="$NODE_ID"
rm -f /tmp/ts
# Confirm we actually have a mesh IP before going any further.
for _ in $(seq 1 12); do
  TS_IP="$(tailscale ip -4 2>/dev/null || true)"
  [[ -n "$TS_IP" ]] && break
  sleep 5
done
[[ -n "${TS_IP:-}" ]] || { echo "Worker never obtained a Tailscale IP — aborting"; exit 1; }

# Install OpenClaw (official installer; Node already present).
installer_5="$(mktemp)"
curl -fsSL https://openclaw.ai/install.sh -o "$installer_5"
less "$installer_5"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_5"
# Run `openclaw onboard --install-daemon` to set up the daemon (see docs).
rm -f "$installer_5"

# Registration: POST to the primary over Tailscale, key from EnvironmentFile.
# NODE_ID / PRIMARY_TS_IP / FLEET_API_KEY / ITEM_HASH are provided via /etc/worker.env.
install -o root -g root -m 600 /dev/null /etc/worker.env
cat > /etc/worker.env <<ENVF
NODE_ID=${NODE_ID}
PRIMARY_TS_IP=${PRIMARY_TS_IP}
FLEET_API_KEY=${FLEET_API_KEY}
ITEM_HASH=${ITEM_HASH}
ENVF

cat > /opt/register-worker.sh <<'REGISTER'
#!/bin/bash
set -euo pipefail
set -a; . /etc/worker.env; set +a
# Use the Tailscale IP as our reachable address. We already joined the mesh in
# the setup phase, so this must succeed; bail rather than register an
# unreachable public/private fallback address.
LOCAL_IP="$(tailscale ip -4 2>/dev/null || true)"
[[ -n "$LOCAL_IP" ]] || { echo "No Tailscale IP yet — not registering an unreachable address"; exit 1; }
curl -fsS -X POST "http://${PRIMARY_TS_IP}:8080/fleet/register" \
  -H "Content-Type: application/json" \
  -H "x-api-key: ${FLEET_API_KEY}" \
  -d "{\"node_id\":\"${NODE_ID}\",\"ip_address\":\"${LOCAL_IP}\",\"item_hash\":\"${ITEM_HASH}\",\"capabilities\":[\"compute\",\"openclaw\"]}"
REGISTER
chmod +x /opt/register-worker.sh

# Register once now (Tailscale is already up), then keep a heartbeat going.
/opt/register-worker.sh

# Heartbeat as a supervised systemd timer (re-registers every 30s; updates last_seen).
cat > /etc/systemd/system/heartbeat.service <<'HB_SVC'
[Unit]
Description=Worker node heartbeat
After=network-online.target tailscaled.service
[Service]
Type=oneshot
EnvironmentFile=/etc/worker.env
ExecStart=/opt/register-worker.sh
HB_SVC
cat > /etc/systemd/system/heartbeat.timer <<'HB_TIMER'
[Unit]
Description=Run worker heartbeat every 30s
[Timer]
OnBootSec=30
OnUnitActiveSec=30
[Install]
WantedBy=timers.target
HB_TIMER
systemctl daemon-reload
systemctl enable --now heartbeat.timer

echo "Worker node setup complete (joined mesh, registered over Tailscale)."
WORKER_SETUP
    )

    local create_args=(
        --name "$node_name"
        --compute-units 2 --memory 4096
        --rootfs-size 40960
        --ssh-pubkey-file "$SSH_KEY.pub"
        --payment-type credit --payment-chain BASE
    )
    [[ -n "$crn_url" ]] && create_args+=(--crn-url "$crn_url" --crn-auto-tac)

    local out item_hash worker_ip
    out="$(aleph instance create "${create_args[@]}")"
    echo "$out"
    item_hash="$(printf '%s\n' "$out" | grep -oE '[0-9a-f]{64}' | head -1)"
    worker_ip="$(wait_for_ip "$node_name")" || { echo "Worker $node_id got no IP"; return 1; }

    # primary_ip here is the primary's TAILSCALE IP (resolved by the caller after
    # the mesh is up). Provision over SSH with NODE_ID/PRIMARY_TS_IP/FLEET_API_KEY/
    # TAILSCALE_AUTH_KEY (so the worker joins the mesh first) and ITEM_HASH (so the
    # primary's registry records the instance hash for later delete/recreate).
    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER@$worker_ip" \
        "NODE_ID='$node_name' PRIMARY_TS_IP='$primary_ip' FLEET_API_KEY='$FLEET_API_KEY' \
         TAILSCALE_AUTH_KEY='$TAILSCALE_AUTH_KEY' ITEM_HASH='$item_hash' bash -s" \
        <<< "$setup_script"

    # Record name, id, crn, IP, and item_hash. IP is REQUIRED by Tailscale/backup/
    # security scripts — never omit it.
    local worker_info tmpfile
    worker_info="$(jq -n --arg n "$node_name" --argjson id "$node_id" \
        --arg crn "$crn_url" --arg ip "$worker_ip" --arg h "$item_hash" \
        '{name:$n, id:$id, crn:$crn, ip:$ip, item_hash:$h}')"
    tmpfile="$(mktemp)"
    jq --argjson w "$worker_info" '.worker_nodes += [$w]' \
        ~/.aleph-deploy/configs/fleet.json > "$tmpfile"
    mv "$tmpfile" ~/.aleph-deploy/configs/fleet.json

    echo "Worker node $node_id deployed on ${crn_url:-auto-selected CRN}: $worker_ip"
}

# Main deployment sequence
echo "Starting fleet deployment sequence..."

# 1. Deploy + provision the primary (installs the fleet manager on its Tailscale IP).
deploy_primary_node
primary_public_ip="$(jq -r '.primary_node.ip' ~/.aleph-deploy/configs/fleet.json)"

# 2. Bring the primary onto Tailscale and capture its mesh IP. Workers register
#    against THIS address (the control plane is never reachable on the public IP).
#    Requires TAILSCALE_AUTH_KEY in the environment (see "Tailscale Mesh" section).
: "${TAILSCALE_AUTH_KEY:?Set TAILSCALE_AUTH_KEY before deploying the fleet}"
ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER@$primary_public_ip" \
    "TAILSCALE_AUTH_KEY='$TAILSCALE_AUTH_KEY' bash -s" <<'TS_BOOT'
set -euo pipefail
installer_6="$(mktemp)"
curl -fsSL https://tailscale.com/install.sh -o "$installer_6"
less "$installer_6"  # Review before execution; verify the release checksum/signature when published.
sh "$installer_6"        # official, OS-detecting installer
rm -f "$installer_6"
printf '%s' "$TAILSCALE_AUTH_KEY" > /tmp/ts && chmod 600 /tmp/ts
tailscale up --auth-key="file:/tmp/ts" --hostname="$(hostname)"
rm -f /tmp/ts
# Re-point the fleet manager at the Tailscale interface and restart it.
TS_IP="$(tailscale ip -4)"
sed -i "s/^BIND_HOST=.*/BIND_HOST=${TS_IP}/" /etc/fleet-manager.env
systemctl restart fleet-manager
echo "PRIMARY_TS_IP=${TS_IP}"
TS_BOOT

primary_ts_ip="$(ssh -i "$SSH_KEY" "$SSH_USER@$primary_public_ip" "tailscale ip -4")"
jq --arg ip "$primary_ts_ip" '.primary_node.tailscale_ip=$ip' \
    ~/.aleph-deploy/configs/fleet.json > /tmp/fleet.$$ && \
    mv /tmp/fleet.$$ ~/.aleph-deploy/configs/fleet.json
echo "Primary Tailscale IP: $primary_ts_ip"

# 3. Deploy workers. Each worker's setup script JOINS the Tailscale mesh FIRST
#    and only THEN registers — so it registers against the primary's Tailscale IP
#    using its own reachable Tailscale IP (never a firewalled public address).
worker_total=$((NODE_COUNT - 1))
for i in $(seq 1 "$worker_total"); do
    if (( ${#WORKER_CRNS[@]} > 0 )); then
        crn_url="${WORKER_CRNS[$(((i - 1) % ${#WORKER_CRNS[@]}))]}"
    else
        crn_url=""   # let the CLI auto-select a CRN
    fi
    deploy_worker_node "$i" "$crn_url" "$primary_ts_ip" &
    sleep 30   # stagger to avoid overwhelming CRNs
done
wait

echo "Fleet deployment complete."
echo "Fleet manager (PRIVATE, Tailscale only): http://$primary_ts_ip:8080"
echo "Status: curl -H \"x-api-key: \$FLEET_API_KEY\" http://$primary_ts_ip:8080/fleet/status"
echo "Next: run setup-tailscale-mesh.sh to verify the mesh, then setup-load-balancer.sh."
jq . ~/.aleph-deploy/configs/fleet.json
```

> **Ordering note.** Workers reach the fleet manager over Tailscale, so the primary joins the mesh *before* workers are provisioned (step 2). Each worker's setup script then joins Tailscale **first** and only **then** registers, so the address it registers is always its reachable Tailscale IP. This requires `TAILSCALE_AUTH_KEY` in the environment (passed through to each worker at SSH time). You can still run `setup-tailscale-mesh.sh` (next section) afterward to verify mesh connectivity. The public IPs are used only for the initial SSH provisioning hop.

> **Primary needs the fleet SSH key (one-time).** Several primary-resident services (replication, backups, node monitor, key rotation) SSH from the primary to workers, so the primary must hold the **private** key. Copy it once, locked down, after the primary is up — prefer Tailscale for the hop:
>
> ```bash
> PRIMARY_TS_IP="$(jq -r '.primary_node.tailscale_ip' ~/.aleph-deploy/configs/fleet.json)"
> scp -i "$SSH_KEY" "$SSH_KEY" "$SSH_USER@$PRIMARY_TS_IP:/root/.ssh/aleph_ed25519"
> ssh -i "$SSH_KEY" "$SSH_USER@$PRIMARY_TS_IP" "chmod 600 /root/.ssh/aleph_ed25519"
> ```
>
> Primary-side scripts read `ALEPH_SSH_KEY` (default `/root/.ssh/aleph_ed25519`). Treat this key as sensitive: it grants root on every worker. Rotate it (see the rotation tool) if the primary is ever compromised, and never bake the private key into an instance setup message.

### Fleet Management Commands

**Fleet Control Script.** Run this from a machine that is **on the tailnet** (the control plane lives on the primary's Tailscale IP). It reads SSH user/key and the manager host from config/env.

```bash
#!/bin/bash
# fleet-control.sh
set -euo pipefail

FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"
SSH_OPTS=(-i "$SSH_KEY" -o StrictHostKeyChecking=accept-new)
# All fleet manager endpoints require x-api-key auth. Keep the key in fleet.env
# (chmod 600), source it before running, or export it; never hardcode it.
FLEET_API_KEY="${FLEET_API_KEY:?FLEET_API_KEY env var is required (see fleet.env)}"

# Control-plane base URL = primary's TAILSCALE IP:8080 (NOT the public IP).
MGR_HOST="$(jq -r '.primary_node.tailscale_ip // .primary_node.ip' "$FLEET_CONFIG")"
mgr() {  # mgr <path> [curl args...]
    curl -fsS -H "x-api-key: $FLEET_API_KEY" "http://$MGR_HOST:8080$1" "${@:2}"
}

fleet_status() {
    echo "Fleet status:"
    mgr /fleet/status | jq '.' || { echo "Unable to reach fleet manager (on tailnet?)"; return 1; }
}

fleet_health() {
    echo "Fleet health check:"
    local nodes; nodes="$(mgr /fleet/status | jq -r '.nodes[].ip_address')"
    for node_ip in $nodes; do
        echo "Checking node: $node_ip"
        if ssh "${SSH_OPTS[@]}" -o ConnectTimeout=5 "$SSH_USER@$node_ip" \
               "systemctl is-active openclaw" &>/dev/null; then
            echo "  OK   $node_ip - OpenClaw running"
        else
            echo "  DOWN $node_ip - OpenClaw not responding"
        fi
    done
}

fleet_restart() {
    local service_name=$1
    [[ "$service_name" =~ ^[a-zA-Z0-9_.-]+$ ]] || { echo "Invalid service: $service_name"; return 1; }
    echo "Restarting $service_name on all nodes..."
    for node_ip in $(mgr /fleet/status | jq -r '.nodes[].ip_address'); do
        echo "  $node_ip"
        ssh "${SSH_OPTS[@]}" "$SSH_USER@$node_ip" "sudo systemctl restart $service_name"
    done
}

fleet_deploy() {
    local script_path=$1
    echo "Deploying script to all nodes: $script_path"
    [[ -f "$script_path" ]] || { echo "Script not found: $script_path"; return 1; }
    for node_ip in $(mgr /fleet/status | jq -r '.nodes[].ip_address'); do
        echo "  $node_ip"
        scp "${SSH_OPTS[@]}" "$script_path" "$SSH_USER@$node_ip":/tmp/deploy-script.sh
        ssh "${SSH_OPTS[@]}" "$SSH_USER@$node_ip" "chmod +x /tmp/deploy-script.sh && sudo /tmp/deploy-script.sh"
    done
}

# Real scale operation. Up: create+provision new workers, register them, let
# haproxy-fleet-sync pick them up. Down: drain HAProxy backend, deregister, then
# DELETE the Aleph instance (irreversible for non-persistent volumes — confirmed).
# Requires FLEET_API_KEY, TAILSCALE_AUTH_KEY, and the deploy/worker helpers; the
# simplest robust approach is to re-invoke deploy-fleet.sh's worker function. Here
# we implement it inline so fleet-control.sh is self-contained.
fleet_scale() {
    local target=$1
    [[ "$target" =~ ^[0-9]+$ ]] || { echo "Target must be an integer"; return 1; }
    local cur; cur="$(jq '.worker_nodes | length' "$FLEET_CONFIG")"   # worker count
    local want=$((target - 1))                                        # minus the primary
    (( want < 0 )) && { echo "Target must be >= 1 (includes primary)"; return 1; }
    echo "Scaling workers from $cur to $want (fleet total $((cur+1)) -> $target)..."

    local primary_ts; primary_ts="$(jq -r '.primary_node.tailscale_ip' "$FLEET_CONFIG")"
    local fleet_name; fleet_name="$(jq -r '.fleet_name' "$FLEET_CONFIG")"

    if (( want > cur )); then
        : "${TAILSCALE_AUTH_KEY:?TAILSCALE_AUTH_KEY required to add workers}"
        command -v aleph >/dev/null || { echo "aleph CLI required on this host to add workers"; return 1; }
        for ((i=cur+1; i<=want; i++)); do
            local wname="${fleet_name}-worker-${i}" wout whash wip
            echo "Adding worker $i ($wname)..."
            # Real create (current flags), then poll the REAL `aleph instance list`.
            wout="$(aleph instance create --name "$wname" --compute-units 2 --rootfs-size 40960 \
                    --ssh-pubkey-file "$SSH_KEY.pub" --payment-type credit --payment-chain BASE 2>&1)"
            echo "$wout"
            whash="$(printf '%s\n' "$wout" | grep -oE '[0-9a-f]{64}' | head -1)"
            wip=""
            for _ in $(seq 1 30); do
                wip="$(aleph instance list --json \
                    | jq -r --arg n "$wname" '.[]|select(.name==$n)|(.ipv4//.ipv6//empty)' | head -1)"
                [[ -n "$wip" ]] && break; sleep 10
            done
            [[ -z "$wip" ]] && { echo "  $wname got no IP — skipping"; continue; }
            # Provision over SSH: Tailscale join + register with the primary over the mesh.
            # ITEM_HASH is passed through so the primary's registry records this
            # instance's hash for later delete/recreate.
            ssh "${SSH_OPTS[@]}" "$SSH_USER@$wip" \
                "NODE_ID='$wname' PRIMARY_TS_IP='$primary_ts' FLEET_API_KEY='$FLEET_API_KEY' \
                 TAILSCALE_AUTH_KEY='$TAILSCALE_AUTH_KEY' ITEM_HASH='$whash' bash -s" <<'REPROV'
set -euo pipefail; export DEBIAN_FRONTEND=noninteractive
apt-get update && apt-get install -y curl jq iproute2
installer_7="$(mktemp)"
curl -fsSL https://get.docker.com -o "$installer_7"
less "$installer_7"  # Review before execution; verify the release checksum/signature when published.
sh "$installer_7"
rm -f "$installer_7"
installer_8="$(mktemp)"
curl -fsSL https://deb.nodesource.com/setup_22.x -o "$installer_8"
less "$installer_8"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_8" - && apt-get install -y nodejs
rm -f "$installer_8"
installer_9="$(mktemp)"
curl -fsSL https://tailscale.com/install.sh -o "$installer_9"
less "$installer_9"  # Review before execution; verify the release checksum/signature when published.
sh "$installer_9"
# file: pattern keeps the auth key out of the process list (see Tailscale section)
rm -f "$installer_9"
[[ -n "${TAILSCALE_AUTH_KEY:-}" ]] && { printf '%s' "$TAILSCALE_AUTH_KEY" > /tmp/ts && chmod 600 /tmp/ts && tailscale up --auth-key="file:/tmp/ts" --hostname="$NODE_ID"; rm -f /tmp/ts; }
installer_10="$(mktemp)"
curl -fsSL https://openclaw.ai/install.sh -o "$installer_10"
less "$installer_10"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_10"
rm -f "$installer_10"
TS_IP="$(tailscale ip -4 2>/dev/null || hostname -I | awk '{print $1}')"
curl -fsS -X POST "http://$PRIMARY_TS_IP:8080/fleet/register" -H "x-api-key: $FLEET_API_KEY" \
  -H 'Content-Type: application/json' \
  -d "{\"node_id\":\"$NODE_ID\",\"ip_address\":\"$TS_IP\",\"item_hash\":\"${ITEM_HASH:-}\",\"capabilities\":[\"compute\",\"openclaw\"]}"
REPROV
            # Append {name,id,ip,item_hash} to fleet.json atomically.
            local wtmp; wtmp="$(mktemp)"
            jq --arg n "$wname" --argjson id "$i" --arg ip "$wip" --arg h "$whash" \
               '.worker_nodes += [{name:$n, id:$id, crn:"", ip:$ip, item_hash:$h}]' \
               "$FLEET_CONFIG" > "$wtmp" && mv "$wtmp" "$FLEET_CONFIG"
            echo "  Added $wname ($wip)."
        done
        echo "New workers register automatically; haproxy-fleet-sync adds them within 60s."
    elif (( want < cur )); then
        local n=$((cur - want))
        echo "Removing $n least-recently-active worker(s)..."
        # Pick workers to remove (last in the list = most recently added).
        local victims; victims="$(jq -r '.worker_nodes[-'"$n"':][] | .name + " " + .ip + " " + .item_hash' "$FLEET_CONFIG")"
        while read -r name ip hash; do
            [[ -z "$name" ]] && continue
            echo "Draining $name ($ip)..."
            # 1. Drain in HAProxy so no new requests go to it, then deregister.
            ssh "${SSH_OPTS[@]}" "$SSH_USER@$primary_ts" \
                "sudo /opt/manage-haproxy-backends.sh remove '$name' || true"
            # 2. Confirm before the irreversible delete.
            echo "About to DELETE Aleph instance $name ($hash). Non-persistent data is lost."
            read -r -p "Type the item-hash to confirm: " ans
            if [[ "$ans" == "$hash" ]]; then
                aleph instance delete "$hash"
                # 3. Atomically drop it from fleet.json.
                local tmp; tmp="$(mktemp)"
                jq --arg n "$name" '.worker_nodes |= map(select(.name != $n))' \
                    "$FLEET_CONFIG" > "$tmp" && mv "$tmp" "$FLEET_CONFIG"
                echo "Removed $name."
            else
                echo "Skipped $name (hash mismatch)."
            fi
        done <<< "$victims"
    else
        echo "Fleet already at target size."
    fi
    # Keep node_count in sync with reality.
    local tmp; tmp="$(mktemp)"
    jq --argjson c "$target" '.node_count=$c' "$FLEET_CONFIG" > "$tmp" && mv "$tmp" "$FLEET_CONFIG"
}

fleet_logs() {
    local service_name="${1:-openclaw}" lines="${2:-50}"
    [[ "$service_name" =~ ^[a-zA-Z0-9_.-]+$ ]] || { echo "Invalid service: $service_name"; return 1; }
    [[ "$lines" =~ ^[0-9]+$ ]] || { echo "Invalid line count: $lines"; return 1; }
    echo "Collecting logs from all nodes..."
    for node_ip in $(mgr /fleet/status | jq -r '.nodes[].ip_address'); do
        echo "=== $node_ip ==="
        ssh "${SSH_OPTS[@]}" "$SSH_USER@$node_ip" "sudo journalctl -u $service_name -n $lines --no-pager"
        echo ""
    done
}

# Command dispatcher
case "${1:-status}" in
    "status")
        fleet_status
        ;;
    "health")
        fleet_health
        ;;
    "restart")
        fleet_restart "${2:-openclaw}"
        ;;
    "deploy")
        fleet_deploy "$2"
        ;;
    "scale")
        fleet_scale "$2"
        ;;
    "logs")
        fleet_logs "$2" "$3"
        ;;
    *)
        echo "Usage: $0 {status|health|restart|deploy|scale|logs}"
        echo ""
        echo "Commands:"
        echo "  status          - Show fleet status"
        echo "  health          - Check health of all nodes"
        echo "  restart [svc]   - Restart service on all nodes"
        echo "  deploy <script> - Deploy script to all nodes"
        echo "  scale <count>   - Scale fleet to N nodes"
        echo "  logs [svc] [n]  - Collect logs from all nodes"
        exit 1
        ;;
esac
```

---

### Resource: references/post-incident-procedures.md

## Post-Incident Procedures

1. Document incident in `/opt/openclaw/incidents/`
2. Review and update recovery procedures
3. Test improvements on staging environment
4. Update team on lessons learned
RUNBOOK

echo "✅ Disaster recovery runbook created at ~/.aleph-deploy/DISASTER_RECOVERY_RUNBOOK.md"
}

# Execute all disaster recovery setup
setup_backup_infrastructure
setup_node_monitoring
create_disaster_recovery_runbook

echo "🛡️ Disaster Recovery System setup complete!"
echo ""
echo "Key Components:"
echo "- Automated daily backups at 2 AM"
echo "- Node health monitoring every 60 seconds"
echo "- Auto-recreation of failed nodes (configurable)"
echo "- Comprehensive recovery runbook"
echo ""
echo "View backup logs: ssh root@PRIMARY_IP tail -f /var/log/backup.log"
echo "View monitoring logs: ssh root@PRIMARY_IP tail -f /var/log/node-monitor.log"
```

---

### Resource: references/quick-start-tested-single-vm-happy-path.md

## Quick Start — tested single-VM happy path

Do this end-to-end first; the fleet machinery below builds on it. Tear-down is included so you never leak a paid VM.

```bash
# 0. Prereqs (mid-2026): Python 3.10+, jq, an SSH client, and a funded Aleph
#    account. macOS: `brew install libsecp256k1`; Debian/Ubuntu:
#    `sudo apt-get install -y libsecp256k1-dev`.

# 1. Install the Python aleph-client CLI (PyPI, v1.9.x): the flag set below
#    targets this client, not the newer aleph-cli documented on docs.aleph.cloud
python3 -m pip install --user pipx && python3 -m pipx ensurepath   # if needed
pipx install aleph-client
aleph --version

# 2. Create or import an account, then check balance
aleph account create                      # interactive; or import an existing key:
# aleph account create --private-key "0xYOUR_PRIVATE_KEY"   # or --private-key-file PATH
aleph account address                      # your public address (fund it on the right chain)
aleph account balance                      # ALEPH balance / available credits

# 3. Generate an SSH key dedicated to Aleph VMs (ed25519 — modern, small, fast)
mkdir -p ~/.aleph-deploy/keys
ssh-keygen -t ed25519 -f ~/.aleph-deploy/keys/aleph_ed25519 -N "" -C "aleph-fleet-$(date +%Y%m%d)"

# 4. See live pricing, then create ONE pay-as-you-go instance (2 vCPU / 4 GiB / 40 GiB)
aleph pricing instance --payment-type credit
aleph instance create \
  --name openclaw-primary \
  --compute-units 2 \
  --rootfs-size 40960 \
  --ssh-pubkey-file ~/.aleph-deploy/keys/aleph_ed25519.pub \
  --payment-type credit \
  --payment-chain BASE \
  --crn-auto-tac
# The CLI prints the instance item-hash and (for PAYG/confidential) allocates it
# on a CRN, returning the assigned IPv6/IPv4. Save the item-hash it prints:
#   ITEM_HASH=<hash from CLI output>

# 5. Find the instance + its IP from authoritative CLI output (no fake `instance get`)
aleph instance list --json | jq '.[] | {name, item_hash, ipv4: .ipv4, ipv6: .ipv6}'
VM_IP=$(aleph instance list --json | jq -r '.[] | select(.name=="openclaw-primary") | (.ipv4 // .ipv6)' | head -1)

# 6. Verify SSH (accept-new = trust first key, reject changed keys = MITM protection)
ssh -i ~/.aleph-deploy/keys/aleph_ed25519 -o StrictHostKeyChecking=accept-new \
    root@"$VM_IP" "echo 'SSH OK'; cat /etc/os-release | grep PRETTY_NAME"
# Note: the default cloud user is image-dependent (often `root` on Aleph base
# images; `ubuntu` on some). Check with the command above and adjust below.

# 7. Tear down when done so you stop paying / release held tokens
aleph instance delete "$ITEM_HASH"     # irreversible for non-persistent volumes — see guardrails
```

> **Destructive-operation guardrail.** `aleph instance delete` is irreversible and **destroys non-persistent (rootfs/ephemeral) data**. Before any delete/scale-down/recreate: (1) `aleph instance list` and confirm the exact item-hash, (2) back up persistent data first, (3) require an explicit confirmation in scripts. A reusable confirm helper:
>
> ```bash
> confirm_destructive() {  # usage: confirm_destructive "<action>" "<target>"
>     echo "About to: $1 -> $2"
>     read -r -p "Type the target hash to confirm: " ans
>     [[ "$ans" == "$2" ]] || { echo "Aborted."; return 1; }
> }
> ```

### Resource: references/security-hardening-framework.md

## Contents

- Security Hardening Framework
- Comprehensive Security Configuration

## Security Hardening Framework

### Comprehensive Security Configuration

**Security Hardening Script:**
```bash
#!/bin/bash
# security-hardening.sh

set -e

FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"

echo "🔒 Implementing comprehensive security hardening..."

setup_firewall_rules() {
    local node_ip=$1
    local node_type=$2
    local ssh_user="${SSH_USER:-$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG")}"
    local ssh_key="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"

    echo "Configuring UFW firewall on $node_type ($node_ip)..."

    # Unquoted heredoc so $node_type expands HERE (operator side) into the remote
    # script. Tailscale's CGNAT range is 100.64.0.0/10; the mesh interface is
    # tailscale0. The OpenClaw agent runtime port is NEVER opened to the internet.
    ssh -i "$ssh_key" -o StrictHostKeyChecking=accept-new "$ssh_user@$node_ip" << FIREWALL_SETUP
#!/bin/bash
set -euo pipefail

echo "Configuring UFW firewall rules..."
sudo ufw --force reset
sudo ufw default deny incoming
sudo ufw default allow outgoing
sudo ufw allow ssh
sudo ufw limit ssh   # rate-limit SSH brute force

# Allow all traffic on the Tailscale mesh interface (private, authenticated mesh).
sudo ufw allow in on tailscale0
sudo ufw allow 41641/udp   # Tailscale direct connections

if [[ "$node_type" == "primary" ]]; then
    # Public edge: only the load balancer's HTTP/HTTPS.
    sudo ufw allow 80
    sudo ufw allow 443
    # Fleet Manager (8080), Load Distributor (8081), HAProxy stats (9090) bind to
    # the Tailscale IP and are reachable ONLY over the mesh (allowed above by the
    # 'in on tailscale0' rule). Do NOT open them publicly.
    echo "Primary node firewall rules applied"
else
    # Worker: NO public OpenClaw port. The gateway (default 18789, loopback-bound)
    # is reachable ONLY over Tailscale (handled by 'allow in on tailscale0');
    # HAProxy on the primary also reaches workers over the mesh. A public gateway
    # port on an agent that can execute actions is a critical exposure; never do it.
    echo "Worker node firewall rules applied (OpenClaw private to mesh)"
fi

# Security hardening rules
sudo ufw deny 23    # Telnet
sudo ufw deny 135   # RPC
sudo ufw deny 139   # NetBIOS
sudo ufw deny 445   # SMB

# Enable firewall
sudo ufw --force enable

# Display status
sudo ufw status verbose

echo "🛡️ Firewall configuration complete"
FIREWALL_SETUP

    echo "✅ Firewall configured on $node_type node"
}

setup_ssh_hardening() {
    local node_ip=$1

    echo "🔑 Hardening SSH configuration on $node_ip..."

    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$node_ip" << 'SSH_HARDENING'
#!/bin/bash
set -e

echo "🔧 Hardening SSH configuration..."

# Backup original SSH config
sudo cp /etc/ssh/sshd_config /etc/ssh/sshd_config.backup

# Create hardened SSH configuration
sudo tee /etc/ssh/sshd_config << 'SSHD_CONFIG'
# SSH Hardened Configuration for Aleph Cloud Fleet

# Basic settings
Port 22
Protocol 2
HostKey /etc/ssh/ssh_host_rsa_key
HostKey /etc/ssh/ssh_host_ecdsa_key
HostKey /etc/ssh/ssh_host_ed25519_key

# Authentication
PubkeyAuthentication yes
AuthorizedKeysFile .ssh/authorized_keys
PasswordAuthentication no
PermitEmptyPasswords no
ChallengeResponseAuthentication no
UsePAM yes

# Security restrictions. NOTE: PermitRootLogin is set dynamically below based on
# the actual login user — on Aleph base images the only user is root, so we use
# `prohibit-password` (key-only root) there rather than `no`, which would lock you out.
MaxAuthTries 3
MaxSessions 2
MaxStartups 2:30:10
LoginGraceTime 30

# Disable dangerous features by default
X11Forwarding no
AllowTcpForwarding no
GatewayPorts no
PermitTunnel no
AllowAgentForwarding no

# Network settings
AddressFamily inet
ListenAddress 0.0.0.0
TCPKeepAlive yes
ClientAliveInterval 300
ClientAliveCountMax 2

# Logging
SyslogFacility AUTH
LogLevel VERBOSE

# Miscellaneous
PrintMotd no
PrintLastLog yes
Compression no
UseDNS no

# Subsystem
Subsystem sftp /usr/lib/openssh/sftp-server -l INFO
SSHD_CONFIG

# Restrict logins and re-enable TCP forwarding for OUR login user only (the user
# is image-dependent — root on Aleph base images, ubuntu on some — so derive it
# at runtime rather than hardcoding "ubuntu". TCP forwarding is needed for SSH
# tunnels (Section 5) and is harmless for Tailscale, which doesn't use sshd.)
LOGIN_USER="$(logname 2>/dev/null || echo "${SUDO_USER:-$USER}")"
if [[ "$LOGIN_USER" == "root" ]]; then ROOT_POLICY="prohibit-password"; else ROOT_POLICY="no"; fi
{
  echo ""
  echo "PermitRootLogin ${ROOT_POLICY}"
  echo "AllowUsers ${LOGIN_USER}"
  echo "Match User ${LOGIN_USER}"
  echo "    AllowTcpForwarding yes"
} | sudo tee -a /etc/ssh/sshd_config >/dev/null

# Validate BEFORE reloading; if invalid, restore the backup so we keep access.
if sudo sshd -t; then
    sudo systemctl reload ssh
    echo "SSH hardening complete (login user: ${LOGIN_USER})"
else
    echo "sshd config invalid — restoring backup, NOT reloading."
    sudo cp /etc/ssh/sshd_config.backup /etc/ssh/sshd_config
    exit 1
fi
SSH_HARDENING

    echo "SSH hardened on node: $node_ip"
}

setup_key_rotation() {
    echo "Installing SSH key rotation tool..."

    # ── scripts/rotate-ssh-keys.sh ───────────────────────────────────────────
    # Correct, verify-before-activate rotation. Key invariants:
    #  - generates an ed25519 key into aleph_ed25519-new (matches the key TYPE);
    #  - tests the NEW key against EVERY node BEFORE activating it (rollback-safe);
    #  - the OLD key stays authorized until the new key is proven, so you can never
    #    lock yourself out; cleanup removes the old key by EXACT LINE (grep -Fvx).
    cat > ~/.aleph-deploy/scripts/rotate-ssh-keys.sh << 'KEY_ROTATION'
#!/bin/bash
set -euo pipefail

FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"
KEY_DIR="$HOME/.aleph-deploy/keys"
BACKUP_DIR="$HOME/.aleph-deploy/key-backups"
mkdir -p "$BACKUP_DIR" "$HOME/.aleph-deploy/logs"

log() { echo "$(date -Iseconds): $1" | tee -a "$HOME/.aleph-deploy/logs/key-rotation.log"; }

# All node IPs (primary + every worker). Workers without an IP are skipped with a warning.
all_node_ips() {
    jq -r '[.primary_node.ip] + [.worker_nodes[].ip] | .[] | select(. != null and . != "")' "$FLEET_CONFIG"
}

generate_new_keys() {
    local d; d="$(date +%Y%m%d-%H%M%S)"
    log "Backing up current key and generating a new ed25519 pair..."
    [[ -f "$KEY_DIR/aleph_ed25519" ]] && {
        cp "$KEY_DIR/aleph_ed25519"     "$BACKUP_DIR/aleph_ed25519-$d"
        cp "$KEY_DIR/aleph_ed25519.pub" "$BACKUP_DIR/aleph_ed25519.pub-$d"
    }
    # ed25519 (not RSA) — matches the active key type and the file name.
    ssh-keygen -t ed25519 -f "$KEY_DIR/aleph_ed25519-new" -N "" -C "aleph-fleet-$d"
}

deploy_new_keys() {
    log "Appending NEW public key to authorized_keys on all nodes (old key stays valid)..."
    local newpub; newpub="$(cat "$KEY_DIR/aleph_ed25519-new.pub")"
    while read -r ip; do
        log "  -> $ip"
        # Still authenticate with the CURRENT (old) key; just append the new one.
        ssh -i "$KEY_DIR/aleph_ed25519" -o StrictHostKeyChecking=accept-new "$SSH_USER@$ip" \
            "mkdir -p ~/.ssh && touch ~/.ssh/authorized_keys && \
             grep -qxF '$newpub' ~/.ssh/authorized_keys || echo '$newpub' >> ~/.ssh/authorized_keys && \
             chmod 600 ~/.ssh/authorized_keys"
    done < <(all_node_ips)
}

# Verify the NEW key works against EVERY node BEFORE we activate it.
test_new_keys() {
    log "Verifying NEW key connectivity on all nodes..."
    local ok=0 fail=0
    while read -r ip; do
        if ssh -i "$KEY_DIR/aleph_ed25519-new" -o ConnectTimeout=10 \
               -o StrictHostKeyChecking=accept-new "$SSH_USER@$ip" "true" &>/dev/null; then
            ok=$((ok+1))
        else
            log "  FAILED on $ip"; fail=$((fail+1))
        fi
    done < <(all_node_ips)
    log "New-key check: $ok ok, $fail failed"
    (( fail == 0 ))
}

activate_new_keys() {
    log "Promoting NEW key to active (old key archived for rollback)..."
    mv "$KEY_DIR/aleph_ed25519"     "$KEY_DIR/aleph_ed25519-old"
    mv "$KEY_DIR/aleph_ed25519.pub" "$KEY_DIR/aleph_ed25519.pub-old"
    mv "$KEY_DIR/aleph_ed25519-new"     "$KEY_DIR/aleph_ed25519"
    mv "$KEY_DIR/aleph_ed25519-new.pub" "$KEY_DIR/aleph_ed25519.pub"
    chmod 600 "$KEY_DIR/aleph_ed25519"; chmod 644 "$KEY_DIR/aleph_ed25519.pub"
}

cleanup_old_keys() {
    local oldpub; oldpub="$(cat "$KEY_DIR/aleph_ed25519.pub-old" 2>/dev/null || true)"
    [[ -z "$oldpub" ]] && return 0
    log "Removing OLD key from all nodes (exact-line match)..."
    while read -r ip; do
        # grep -Fvx: fixed-string, whole-LINE, inverted — removes ONLY the exact old
        # key line, never a substring or an unrelated key. Connect with the new key.
        ssh -i "$KEY_DIR/aleph_ed25519" "$SSH_USER@$ip" \
            "grep -Fvx '$oldpub' ~/.ssh/authorized_keys > ~/.ssh/ak.tmp && \
             mv ~/.ssh/ak.tmp ~/.ssh/authorized_keys && chmod 600 ~/.ssh/authorized_keys" \
            || log "  WARN: could not clean $ip (left old key in place)"
    done < <(all_node_ips)
    rm -f "$KEY_DIR/aleph_ed25519-old" "$KEY_DIR/aleph_ed25519.pub-old"
}

rotate_keys() {
    log "Starting SSH key rotation..."
    generate_new_keys
    deploy_new_keys
    sleep 5
    if test_new_keys; then          # MUST pass on every node before we switch over
        activate_new_keys
        cleanup_old_keys            # old key only removed AFTER new key is active+proven
        log "SSH key rotation completed successfully."
    else
        log "New key failed on at least one node — NOT activating. Old key still works."
        rm -f "$KEY_DIR/aleph_ed25519-new" "$KEY_DIR/aleph_ed25519-new.pub"
        return 1
    fi
}

case "${1:-rotate}" in
    rotate) rotate_keys ;;
    test)   test_new_keys ;;
    *)      echo "Usage: $0 {rotate|test}"; exit 1 ;;
esac
KEY_ROTATION

    chmod +x ~/.aleph-deploy/scripts/rotate-ssh-keys.sh

    # Rotate ON DEMAND, not on a forced monthly schedule. Automatic forced rotation
    # of SSH keys provides little security benefit and risks lock-out if a node is
    # unreachable when the cron fires. Rotate when a key may be compromised or when
    # an operator leaves. To opt into scheduled rotation, uncomment:
    # (crontab -l 2>/dev/null; echo "0 3 1 * * $HOME/.aleph-deploy/scripts/rotate-ssh-keys.sh rotate >> $HOME/.aleph-deploy/logs/key-rotation.log 2>&1") | crontab -
    echo "SSH key rotation tool installed: ~/.aleph-deploy/scripts/rotate-ssh-keys.sh rotate"
}

setup_intrusion_detection() {
    echo "👁️ Setting up intrusion detection system..."

    local primary_ip=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")

    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$primary_ip" << 'IDS_SETUP'
#!/bin/bash
set -e

echo "🔍 Installing and configuring intrusion detection..."

# Install fail2ban
sudo apt-get update
sudo apt-get install -y fail2ban

# Create custom jail configuration
sudo tee /etc/fail2ban/jail.local << 'JAIL_CONFIG'
[DEFAULT]
# Ban time: 1 hour
bantime = 3600
# Find time: 10 minutes
findtime = 600
# Max retry: 3 attempts
maxretry = 3
# Ignore local IPs
ignoreip = 127.0.0.1/8 10.0.0.0/8 172.16.0.0/12 192.168.0.0/16

[sshd]
enabled = true
port = ssh
filter = sshd
# fail2ban >= 0.10 merged the old standalone sshd-ddos filter into the sshd
# filter; "mode = aggressive" covers those patterns. A separate [sshd-ddos]
# jail would reference a missing filter and abort fail2ban startup entirely.
mode = aggressive
logpath = /var/log/auth.log
maxretry = 3
bantime = 3600

# OpenClaw service protection. DISABLED by default: fail2ban refuses to start
# a jail whose logpath is missing, and OpenClaw does not create
# /var/log/openclaw/access.log out of the box. Enable only after the log file
# exists (touch it with correct ownership, or point logpath at a real log).
[openclaw]
enabled = false
port = 18789
filter = openclaw
logpath = /var/log/openclaw/access.log
maxretry = 10
bantime = 1800

# Fleet manager protection. DISABLED by default for the same reason: the fleet
# manager logs to journald via systemd, not to /var/log/fleet-manager.log.
# Enable after either forwarding its journal to that file (syslog rule) or
# switching this jail to "backend = systemd".
[fleet-manager]
enabled = false
port = 8080
filter = fleet-manager
logpath = /var/log/fleet-manager.log
maxretry = 5
bantime = 1800
JAIL_CONFIG

# Create custom filters
sudo mkdir -p /etc/fail2ban/filter.d

# OpenClaw filter
sudo tee /etc/fail2ban/filter.d/openclaw.conf << 'OPENCLAW_FILTER'
[Definition]
failregex = .*Failed authentication from <HOST>.*
            .*Invalid request from <HOST>.*
            .*Rate limit exceeded from <HOST>.*
ignoreregex =
OPENCLAW_FILTER

# Fleet manager filter
sudo tee /etc/fail2ban/filter.d/fleet-manager.conf << 'FLEET_FILTER'
[Definition]
failregex = .*Unauthorized access attempt from <HOST>.*
            .*Invalid API key from <HOST>.*
ignoreregex =
FLEET_FILTER

# Enable and start fail2ban
sudo systemctl enable fail2ban
sudo systemctl start fail2ban

# Create monitoring script
cat > /opt/security-monitor.sh << 'SEC_MONITOR'
#!/bin/bash

log_security_event() {
    local event_type=$1
    local details=$2
    echo "$(date -Iseconds): [$event_type] $details" | tee -a /var/log/security-events.log
}

check_failed_logins() {
    local failed_logins=$(grep "Failed password" /var/log/auth.log | grep "$(date +%b\ %d)" | wc -l)

    if (( failed_logins > 10 )); then
        log_security_event "HIGH_FAILED_LOGINS" "Detected $failed_logins failed login attempts today"
    fi
}

check_banned_ips() {
    local banned_count=$(sudo fail2ban-client status sshd | grep "Currently banned:" | awk '{print $3}')

    if (( banned_count > 0 )); then
        local banned_ips=$(sudo fail2ban-client status sshd | grep "Banned IP list:" | cut -d: -f2)
        log_security_event "IPS_BANNED" "Currently banned IPs: $banned_ips"
    fi
}

check_unusual_processes() {
    # Check for processes consuming high CPU
    local high_cpu_procs=$(ps aux --sort=-%cpu | head -6 | tail -5 | awk '$3 > 80')

    if [[ -n "$high_cpu_procs" ]]; then
        log_security_event "HIGH_CPU_USAGE" "Processes consuming high CPU detected"
    fi
}

check_network_connections() {
    # Check for unusual network connections
    # `ss` is the default on modern Ubuntu (netstat needs the net-tools package).
    local external_connections=$(ss -tn state established | tail -n +2 | grep -v "127.0.0.1\|10.\|172.16\|192.168" | wc -l)

    if (( external_connections > 50 )); then
        log_security_event "HIGH_EXTERNAL_CONNECTIONS" "Detected $external_connections external connections"
    fi
}

# Run security checks
check_failed_logins
check_banned_ips
check_unusual_processes
check_network_connections

# Generate daily security summary
if [[ "$(date +%H:%M)" == "23:59" ]]; then
    log_security_event "DAILY_SUMMARY" "Security monitoring completed for $(date +%Y-%m-%d)"
fi
SEC_MONITOR

chmod +x /opt/security-monitor.sh

# Setup security monitoring cron
(crontab -l 2>/dev/null; echo "*/15 * * * * /opt/security-monitor.sh") | crontab -

echo "✅ Intrusion detection system configured"
IDS_SETUP

echo "✅ Intrusion detection configured on primary node"
}

setup_log_monitoring() {
    echo "📋 Setting up centralized log monitoring..."

    local primary_ip=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")

    # Setup log aggregation on primary node
    ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$primary_ip" << 'LOG_SETUP'
#!/bin/bash
set -e

echo "📊 Setting up centralized logging..."

# Install rsyslog for log aggregation
sudo apt-get update
sudo apt-get install -y rsyslog

# Configure rsyslog as log server
sudo tee /etc/rsyslog.conf << 'RSYSLOG_CONFIG'
# Provides TCP syslog reception
$ModLoad imtcp
$InputTCPServerRun 514

# Provides UDP syslog reception
$ModLoad imudp
$InputUDPServerRun 514

# Log templates
$template RemoteLogs,"/var/log/remote/%HOSTNAME%/%PROGRAMNAME%.log"
*.* ?RemoteLogs
& ~

# Local logging
$ActionFileDefaultTemplate RSYSLOG_TraditionalFileFormat
auth,authpriv.*                 /var/log/auth.log
*.*;auth,authpriv.none         -/var/log/syslog
daemon.*                       -/var/log/daemon.log
kern.*                         -/var/log/kern.log
mail.*                         -/var/log/mail.log
user.*                         -/var/log/user.log

# Emergency messages to all logged in users
*.emerg                         :omusrmsg:*
RSYSLOG_CONFIG

# Create log directories
sudo mkdir -p /var/log/remote
sudo chown -R syslog:syslog /var/log/remote

# Restart rsyslog
sudo systemctl restart rsyslog

# Create log analysis script
cat > /opt/log-analyzer.sh << 'LOG_ANALYZER'
#!/bin/bash

LOG_DIR="/var/log"
REPORT_DIR="/opt/log-reports"
REPORT_DATE=$(date +%Y-%m-%d)

mkdir -p "$REPORT_DIR"

generate_security_report() {
    echo "🔍 Generating security log analysis..."

    local report_file="$REPORT_DIR/security-report-$REPORT_DATE.txt"

    {
        echo "SECURITY LOG ANALYSIS - $REPORT_DATE"
        echo "=================================="
        echo ""

        echo "SSH Login Attempts:"
        grep "sshd" "$LOG_DIR/auth.log" | grep "$(date +%b\ %d)" | grep "Failed password" | wc -l
        echo ""

        echo "Successful SSH Logins:"
        grep "sshd" "$LOG_DIR/auth.log" | grep "$(date +%b\ %d)" | grep "Accepted password" | wc -l
        echo ""

        echo "Fail2ban Actions:"
        grep "fail2ban" "$LOG_DIR/fail2ban.log" | grep "$(date +%Y-%m-%d)" | tail -10
        echo ""

        echo "Top Source IPs (Failed Logins):"
        grep "Failed password" "$LOG_DIR/auth.log" | grep "$(date +%b\ %d)" | awk '{print $(NF-3)}' | sort | uniq -c | sort -nr | head -5
        echo ""

        echo "OpenClaw Service Status:"
        systemctl status openclaw --no-pager || echo "Service not found"
        echo ""

        echo "Fleet Manager Status:"
        systemctl status fleet-manager --no-pager || echo "Service not found"

    } > "$report_file"

    echo "✅ Security report generated: $report_file"
}

generate_performance_report() {
    echo "📈 Generating performance log analysis..."

    local report_file="$REPORT_DIR/performance-report-$REPORT_DATE.txt"

    {
        echo "PERFORMANCE LOG ANALYSIS - $REPORT_DATE"
        echo "====================================="
        echo ""

        echo "System Load Average:"
        uptime
        echo ""

        echo "Memory Usage:"
        free -h
        echo ""

        echo "Disk Usage:"
        df -h
        echo ""

        echo "Top Processes by CPU:"
        ps aux --sort=-%cpu | head -6
        echo ""

        echo "Top Processes by Memory:"
        ps aux --sort=-%mem | head -6
        echo ""

        echo "Network Connections (established):"
        ss -tn state established | tail -n +2 | wc -l

    } > "$report_file"

    echo "✅ Performance report generated: $report_file"
}

# Generate reports
generate_security_report
generate_performance_report

# Cleanup old reports (keep 30 days)
find "$REPORT_DIR" -name "*.txt" -mtime +30 -delete
LOG_ANALYZER

chmod +x /opt/log-analyzer.sh

# Setup daily log analysis
(crontab -l 2>/dev/null; echo "0 1 * * * /opt/log-analyzer.sh") | crontab -

echo "✅ Centralized logging configured"
LOG_SETUP

echo "✅ Log monitoring configured on primary node"
}

# Execute security hardening for all nodes
harden_all_nodes() {
    echo "🔒 Hardening security on all fleet nodes..."

    local primary_ip=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")
    local worker_ips=($(jq -r '.worker_nodes[] | .ip // empty' "$FLEET_CONFIG"))

    # Harden primary node
    echo "🛡️ Hardening primary node..."
    setup_firewall_rules "$primary_ip" "primary"
    setup_ssh_hardening "$primary_ip"

    # Harden worker nodes
    for worker_ip in "${worker_ips[@]}"; do
        [[ -z "$worker_ip" || "$worker_ip" == "null" ]] && continue

        echo "🛡️ Hardening worker node: $worker_ip..."
        setup_firewall_rules "$worker_ip" "worker"
        setup_ssh_hardening "$worker_ip"
    done
}

# Create security status checker
create_security_checker() {
    echo "🔍 Creating security status checker..."

    cat > ~/.aleph-deploy/scripts/security-status.sh << 'SEC_STATUS'
#!/bin/bash

FLEET_CONFIG="$HOME/.aleph-deploy/configs/fleet.json"
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="$(jq -r '.ssh_user // "root"' "$FLEET_CONFIG" 2>/dev/null || echo root)"

check_node_security() {
    local node_ip=$1
    local node_type=$2

    echo "🔍 Checking security status of $node_type node ($node_ip)..."

    # Check UFW status
    echo -n "  Firewall: "
    if ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$node_ip" "sudo ufw status" | grep -q "Status: active"; then
        echo "✅ Active"
    else
        echo "❌ Inactive"
    fi

    # Check SSH configuration
    echo -n "  SSH Security: "
    local ssh_score=0
    if ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$node_ip" "grep -q 'PasswordAuthentication no' /etc/ssh/sshd_config"; then
        ssh_score=$((ssh_score + 1))
    fi
    # Accept either 'no' or 'prohibit-password' (the latter is used on root-only images).
    if ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$node_ip" "grep -Eq 'PermitRootLogin (no|prohibit-password)' /etc/ssh/sshd_config"; then
        ssh_score=$((ssh_score + 1))
    fi
    if ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$node_ip" "grep -q 'MaxAuthTries 3' /etc/ssh/sshd_config"; then
        ssh_score=$((ssh_score + 1))
    fi

    if (( ssh_score >= 2 )); then
        echo "✅ Hardened ($ssh_score/3)"
    else
        echo "⚠️ Needs attention ($ssh_score/3)"
    fi

    # Check fail2ban (primary node only)
    if [[ "$node_type" == "primary" ]]; then
        echo -n "  Intrusion Detection: "
        if ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$node_ip" "systemctl is-active fail2ban" &>/dev/null; then
            echo "✅ Active"
        else
            echo "❌ Inactive"
        fi
    fi

    # Check system updates
    echo -n "  System Updates: "
    local updates=$(ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER"@"$node_ip" "apt list --upgradable 2>/dev/null | grep -c upgradable || echo 0")
    if (( updates == 0 )); then
        echo "✅ Up to date"
    else
        echo "⚠️ $updates updates available"
    fi

    echo ""
}

# Check all fleet nodes
check_fleet_security() {
    local primary_ip=$(jq -r '.primary_node.ip' "$FLEET_CONFIG")

    echo "🔒 FLEET SECURITY STATUS"
    echo "========================"
    echo ""

    check_node_security "$primary_ip" "primary"

    local worker_ips=($(jq -r '.worker_nodes[] | .ip // empty' "$FLEET_CONFIG"))
    for worker_ip in "${worker_ips[@]}"; do
        [[ -z "$worker_ip" || "$worker_ip" == "null" ]] && continue
        check_node_security "$worker_ip" "worker"
    done
}

check_fleet_security
SEC_STATUS

chmod +x ~/.aleph-deploy/scripts/security-status.sh

echo "✅ Security status checker created"
}

# Execute all security hardening
harden_all_nodes
setup_key_rotation
setup_intrusion_detection
setup_log_monitoring
create_security_checker

echo "🔒 Security hardening complete!"
echo ""
echo "Security components:"
echo "- UFW firewall configured on all nodes"
echo "- SSH hardened (key-only; root login set to prohibit-password on root-only images)"
echo "- On-demand SSH key rotation (verify-before-activate; not a forced schedule)"
echo "- Fail2ban intrusion detection"
echo "- Centralized logging"
echo ""
echo "Check security status: ~/.aleph-deploy/scripts/security-status.sh"
```

---

### Resource: references/single-node-deployment-foundation.md

## Contents

- Single Node Deployment Foundation
- Prerequisites & Setup
- Single VM Deployment

## Single Node Deployment Foundation

### Prerequisites & Setup

**Local Environment Setup:**
```bash
#!/bin/bash
# setup-aleph-environment.sh
set -euo pipefail

echo "Setting up Aleph Cloud deployment environment..."

# Install the Python aleph-client CLI (PyPI, v1.9.x) in an isolated environment.
# The flags used throughout this skill target this client, not the newer
# aleph-cli documented at https://docs.aleph.cloud/devhub/sdks-and-tools/aleph-cli/
if ! command -v aleph &>/dev/null; then
    echo "Installing aleph-client via pipx..."
    if ! command -v pipx &>/dev/null; then
        python3 -m pip install --user pipx
        python3 -m pipx ensurepath
        echo "Restart your shell, then re-run this script."; exit 1
    fi
    # System dep: libsecp256k1 (macOS: brew install libsecp256k1;
    # Debian/Ubuntu: sudo apt-get install -y libsecp256k1-dev)
    pipx install aleph-client
fi

aleph --version

# Create deployment directory structure
mkdir -p ~/.aleph-deploy/{keys,configs,scripts,backups,logs,reports}

# Generate an ed25519 SSH key pair for VMs (preferred over RSA in 2026)
if [[ ! -f ~/.aleph-deploy/keys/aleph_ed25519 ]]; then
    echo "Generating SSH key pair..."
    ssh-keygen -t ed25519 -f ~/.aleph-deploy/keys/aleph_ed25519 -N "" \
        -C "aleph-fleet-$(date +%Y%m%d)"
fi

echo "Environment setup complete."
echo "Next steps:"
echo "  1. aleph account create        # or import: --private-key / --private-key-file"
echo "  2. Fund the address shown by 'aleph account address' on your payment chain"
echo "  3. aleph pricing instance      # check current tiers/prices"
```

> **Key type note.** This skill standardizes on **ed25519** keys at `~/.aleph-deploy/keys/aleph_ed25519`. If you are upgrading an older deployment that used a different key (for example `~/.aleph-deploy/keys/aleph_rsa`), either re-run the generator above or `export ALEPH_SSH_KEY=<path-to-your-existing-key>` and keep it consistent; every script below reads the same variable.

**Account Creation & Funding:**
```bash
#!/bin/bash
# account-setup.sh
set -euo pipefail

echo "Setting up Aleph account..."

read -rp "Do you want to (c)reate new account or (i)mport existing? " choice
case "$choice" in
    c|C)
        echo "Creating new account..."
        aleph account create               # add --replace only to overwrite an existing default
        ;;
    i|I)
        echo "Importing an existing key..."
        # Documented import path is `account create` with a key source — there is
        # no `aleph account import-private-key` command.
        read -rsp "Paste private key (hidden), or leave blank to use a file: " pk; echo
        if [[ -n "$pk" ]]; then
            aleph account create --private-key "$pk" --replace
        else
            read -rp "Path to private key file: " pkfile
            aleph account create --private-key-file "$pkfile" --replace
        fi
        ;;
    *)
        echo "Invalid choice"; exit 1 ;;
esac

echo "Active account:"
aleph account show
ADDR=$(aleph account address)
echo "Address: $ADDR"

# Check balance (correct command is `aleph account balance`, not `aleph balance`)
echo "Balance / credits:"
aleph account balance

echo
echo "Funding: send ALEPH (or buy credits) to the address above on your chosen"
echo "payment chain (ETH / BASE / AVAX / SOL). Manage funds in the console:"
echo "  https://app.aleph.cloud"
echo "Budget guidance: run 'aleph pricing instance' for current per-tier USD pricing."
echo "Account setup complete."
```

### Single VM Deployment

**Why provision over SSH, not `--setup-script`.** The Aleph CLI does **not** take a `--setup-script` flag, and there is no `aleph instance status --wait` / `aleph instance get`. The reliable pattern is: create the instance, poll `aleph instance list` for its IP, then run a provisioning script over SSH. This also keeps the (large) setup logic out of the on-chain message.

**Basic VM Deployment Script:**
```bash
#!/bin/bash
# deploy-single-vm.sh — create one instance, then provision it over SSH.
set -euo pipefail

# Configuration
VM_NAME="${1:-openclaw-primary}"
COMPUTE_UNITS="${2:-2}"          # 1 CU ~= 1 vCPU + 2 GiB RAM
ROOTFS_MIB="${3:-40960}"          # 40 GiB in MiB (rootfs-size is in MiB)
PAYMENT_TYPE="${4:-credit}"       # hold | superfluid | credit | nft
PAYMENT_CHAIN="${5:-BASE}"        # ETH | BASE | AVAX | SOL
CRN_URL="${6:-}"                  # optional: pin a CRN you verified; else auto-select
SSH_KEY="${ALEPH_SSH_KEY:-$HOME/.aleph-deploy/keys/aleph_ed25519}"
SSH_USER="${ALEPH_SSH_USER:-root}"   # image-dependent (root on Aleph base images)

echo "Deploying single VM: $VM_NAME ($COMPUTE_UNITS CU, $((ROOTFS_MIB/1024)) GiB)"

# 1. Create the instance with CURRENT flags. (Run `aleph instance create --help`
#    to confirm flags for your CLI version.)
create_args=(
    --name "$VM_NAME"
    --compute-units "$COMPUTE_UNITS"
    --rootfs-size "$ROOTFS_MIB"
    --ssh-pubkey-file "$SSH_KEY.pub"
    --payment-type "$PAYMENT_TYPE"
    --payment-chain "$PAYMENT_CHAIN"
    # Keep agent state on a persistent volume so a VM stop/rebuild doesn't wipe it.
    # Syntax: name=...,mount=...,size_mib=... (see `aleph instance create --help`).
    --persistent-volume "name=data,mount=/data,size_mib=20480"
)
if [[ -n "$CRN_URL" ]]; then
    create_args+=(--crn-url "$CRN_URL" --crn-auto-tac)   # auto-accept that CRN's T&C
fi

# Capture output so we can extract the item-hash the CLI prints.
CREATE_OUT="$(aleph instance create "${create_args[@]}")"
echo "$CREATE_OUT"
ITEM_HASH="$(printf '%s\n' "$CREATE_OUT" | grep -oE '[0-9a-f]{64}' | head -1)"
echo "Instance item-hash: ${ITEM_HASH:-<parse from output above>}"

# 2. Poll `aleph instance list` (the real command) for the assigned IP.
echo "Waiting for an IP to be assigned..."
VM_IP=""
for _ in $(seq 1 30); do
    VM_IP="$(aleph instance list --json \
        | jq -r --arg n "$VM_NAME" '.[] | select(.name==$n) | (.ipv4 // .ipv6 // empty)' \
        | head -1)"
    [[ -n "$VM_IP" ]] && break
    sleep 10
done
[[ -z "$VM_IP" ]] && { echo "No IP yet; check 'aleph instance list' and CRN allocation."; exit 1; }
echo "VM IP: $VM_IP"

# 3. Verify SSH (accept-new: trust first host key, reject changed keys = MITM defense).
echo "Testing SSH connection..."
ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new -o ConnectTimeout=15 \
    "$SSH_USER@$VM_IP" "echo 'SSH connection successful'"

# 4. Provision over SSH (kept out of the on-chain message; rerun-safe).
echo "Provisioning $VM_NAME..."
ssh -i "$SSH_KEY" -o StrictHostKeyChecking=accept-new "$SSH_USER@$VM_IP" \
    "OPENCLAW_GATEWAY_PORT=18789 bash -s" <<'PROVISION'
#!/bin/bash
set -euo pipefail
export DEBIAN_FRONTEND=noninteractive

# --- Base packages (modern names; ss replaces netstat; net-tools optional) ---
apt-get update && apt-get -y upgrade
apt-get install -y curl wget git htop unzip jq fail2ban ufw ca-certificates \
                   iproute2   # provides `ss`

# --- Docker Engine + Compose v2 plugin (NOT the deprecated docker-compose binary) ---
# Verify checksums in high-security environments; get-docker.sh is Docker's official script.
curl -fsSL https://get.docker.com -o /tmp/get-docker.sh
sh /tmp/get-docker.sh
SUDO_USER_NAME="${SUDO_USER:-$(logname 2>/dev/null || echo root)}"
usermod -aG docker "$SUDO_USER_NAME" || true
docker compose version    # Compose v2 ships as a Docker plugin: `docker compose ...`

# --- Node.js 22.x (OpenClaw requires Node >= 22.19; 24 is recommended) ---
# Official NodeSource installer; to review first:
#   curl -fsSL https://deb.nodesource.com/setup_22.x -o /tmp/ns.sh && less /tmp/ns.sh
installer_1="$(mktemp)"
curl -fsSL https://deb.nodesource.com/setup_22.x -o "$installer_1"
less "$installer_1"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_1" -
rm -f "$installer_1"
apt-get install -y nodejs
node --version

# --- Firewall: deny inbound by default; HTTP/HTTPS only on the public edge.
#     The OpenClaw port is NOT opened publicly (see Security section) — reach it
#     via the Tailscale mesh or behind HAProxy. ---
ufw default deny incoming
ufw default allow outgoing
ufw allow ssh
ufw limit ssh
ufw allow 80
ufw allow 443
ufw --force enable

# --- Install OpenClaw (official installer; handles Node if missing) and onboard ---
# Docs: https://docs.openclaw.ai/install . The installer sets up a systemd daemon
# via `--install-daemon`; do not hand-roll an ExecStart=/usr/bin/node server.js unit.
# Official OpenClaw installer (docs.openclaw.ai/install); download and review the
# script first in high-security environments.
installer_2="$(mktemp)"
curl -fsSL https://openclaw.ai/install.sh -o "$installer_2"
less "$installer_2"  # Review before execution; verify the release checksum/signature when published.
bash "$installer_2"
# Onboarding is interactive by design; for headless provisioning configure tokens
rm -f "$installer_2"
# via env/secret store first, then:
#   openclaw onboard --install-daemon
# Verify once installed:
#   openclaw --version && openclaw doctor && openclaw gateway status
# The gateway listens on 18789 by default (OPENCLAW_GATEWAY_PORT or gateway.port
# override it) and binds loopback by default: configure it to listen on the
# Tailscale interface before HAProxy or mesh peers can reach it.

echo "VM base provisioning complete."
PROVISION

echo "Deployment complete."
echo "SSH:     ssh -i $SSH_KEY $SSH_USER@$VM_IP"
echo "OpenClaw: reach it over Tailscale or via HAProxy (gateway defaults to 18789, loopback-bound; NEVER a public port)."
echo "Tear down: aleph instance delete ${ITEM_HASH:-<item-hash>}"
```

> **OpenClaw config note.** OpenClaw is configured through `openclaw onboard` (writing to its own workspace under `~/.openclaw`/the daemon home), **not** by hand-authored `/opt/openclaw/config/production.json` files with keys like `server.cluster` or `aleph.node_id` — those were invented in earlier drafts and OpenClaw does not read them. Treat any per-node "role" we track (primary/worker) as *our* fleet metadata in `fleet.json`, separate from OpenClaw's own config.

---

### Resource: references/table-of-contents.md

## Table of Contents

1. [Infrastructure Planning & Architecture](#infrastructure-planning--architecture)
2. [Single Node Deployment Foundation](#single-node-deployment-foundation)
3. [Multi-Node Fleet Management](#multi-node-fleet-management)
4. [Auto-Provisioning Protocol (SRP)](#auto-provisioning-protocol-srp)
5. [Inter-VM Communication Networks](#inter-vm-communication-networks)
6. [Load Distribution & Orchestration](#load-distribution--orchestration)
7. [Disaster Recovery & Auto-Recreation](#disaster-recovery--auto-recreation)
8. [Cost Optimization Strategies](#cost-optimization-strategies)
9. [Security Hardening Framework](#security-hardening-framework)
10. [Monitoring & Maintenance](#monitoring--maintenance)

---

---

## api-design
Category: dev
Description: Production HTTP API design — REST conventions, pagination, error models, versioning, rate limiting, auth, and idempotency. Use when designing or reviewing public/internal HTTP APIs, OpenAPI contracts, pagination, error models, rate limits, auth, or idempotent write endpoints.
Features:
  - REST best practices (naming, methods, status codes, pagination)
  - OpenAPI 3.1 specification generation
  - Authentication patterns (JWT, OAuth2, API keys)
  - Rate limiting and error handling (RFC 7807)
  - GraphQL schema design patterns
  - Webhook design with signature verification
Use Cases:
  - Design a RESTful API from scratch
  - Generate OpenAPI specs for documentation
  - Implement rate limiting and auth
  - Design webhook delivery with retry logic

# API Design

## Reference guide

Read only the references needed for the current request:

- **REST Conventions That Actually Matter**: [references/rest-conventions-that-actually-matter.md](references/rest-conventions-that-actually-matter.md)
- **Pagination: Cursor vs Offset**: [references/pagination-cursor-vs-offset.md](references/pagination-cursor-vs-offset.md)
- **Error Handling: RFC 9457 Problem Details**: [references/error-handling-rfc-9457-problem-details.md](references/error-handling-rfc-9457-problem-details.md)
- **API Versioning**: [references/api-versioning.md](references/api-versioning.md)
- **Rate Limiting**: [references/rate-limiting.md](references/rate-limiting.md)
- **Authentication Patterns**: [references/authentication-patterns.md](references/authentication-patterns.md)
- **Idempotency**: [references/idempotency.md](references/idempotency.md)
- **OpenAPI 3.1 Specification**: [references/openapi-3-1-specification.md](references/openapi-3-1-specification.md)
- **GraphQL vs REST: Decision Matrix**: [references/graphql-vs-rest-decision-matrix.md](references/graphql-vs-rest-decision-matrix.md)
- **Response Envelope**: [references/response-envelope.md](references/response-envelope.md)
- **Checklist: Production-Ready API**: [references/checklist-production-ready-api.md](references/checklist-production-ready-api.md)

### Resource: references/api-versioning.md

## Contents

- API Versioning
- URL Versioning (Preferred for Public APIs)
- Header Versioning (Alternative)
- Deprecation Strategy
- Versioning Timeline

## API Versioning

### URL Versioning (Preferred for Public APIs)

```
/api/v1/users
/api/v2/users
```

Simple, explicit, easy to route. The pragmatic choice.

### Header Versioning (Alternative)

```
Accept: application/vnd.myapi.v2+json
```

More "RESTful" but harder to test (can't just paste a URL).

### Deprecation Strategy

```typescript
// middleware/deprecation.ts
function deprecationWarning(sunset: string, alternative: string) {
  return (req: Request, res: Response, next: NextFunction) => {
    // RFC 9745: the value is a structured-field Date (@unix-timestamp), not a
    // boolean. Using the sunset date satisfies RFC 9745's rule that Sunset must
    // not be earlier than Deprecation; pass a separate deprecation date if the
    // API was deprecated before the sunset.
    res.setHeader('Deprecation', `@${Math.floor(new Date(sunset).getTime() / 1000)}`);
    res.setHeader('Sunset', sunset);  // RFC 8594
    res.setHeader('Link', `<${alternative}>; rel="successor-version"`);
    next();
  };
}

// Usage: Sunset must be an HTTP-date (RFC 8594 / RFC 9110), in the future
app.get('/api/v1/users',
  deprecationWarning('Wed, 01 Jul 2026 00:00:00 GMT', '/api/v2/users'),
  v1UserHandler,
);
```

### Versioning Timeline

```
v1 released → v2 released → v1 deprecated (6 month warning) → v1 sunset (returns 410 Gone)
```

---

### Resource: references/authentication-patterns.md

## Contents

- Authentication Patterns
- JWT Access + Refresh Token (Fastify)
- API Keys (Service-to-Service)

## Authentication Patterns

### JWT Access + Refresh Token (Fastify)

```typescript
import Fastify, { FastifyRequest, FastifyReply } from 'fastify';
import jwt from '@fastify/jwt';

const app = Fastify();

await app.register(jwt, {
  secret: process.env.JWT_SECRET!,
  sign: { expiresIn: '15m' },  // Short-lived access tokens
});

// Decorate the `authenticate` preHandler used by protected routes below.
// Without this decorator the `preHandler: [app.authenticate]` example throws.
app.decorate('authenticate', async (request, reply) => {
  try {
    await request.jwtVerify();  // populates request.user from the Bearer token
  } catch {
    throw new AppError(401, 'UNAUTHENTICATED', 'Missing or invalid access token');
  }
});

// TypeScript: augment Fastify so `app.authenticate` and `request.user` type-check.
declare module 'fastify' {
  interface FastifyInstance {
    authenticate: (request: FastifyRequest, reply: FastifyReply) => Promise<void>;
  }
}
declare module '@fastify/jwt' {
  interface FastifyJWT {
    payload: { sub: string; role: string };  // sign() input
    user: { sub: string; role: string };     // request.user shape
  }
}

// Refresh tokens use a SELECTOR.SECRET design so lookup is a single indexed
// query, never a scan over every active hash:
//   - selector: random id, stored in plaintext, UNIQUE-indexed — used to find the row
//   - secret:   random, stored only as an argon2 hash — verified in constant time
//   - familyId: groups every token descended from one login, so reuse of a
//               rotated token can revoke the whole family (theft detection)
// Wire format handed to the client is `${selector}.${secret}`.
function issueRefreshToken(userId: string, familyId: string) {
  const selector = crypto.randomBytes(16).toString('base64url');
  const secret = crypto.randomBytes(32).toString('base64url');
  return { token: `${selector}.${secret}`, selector, secret, familyId, userId };
}

const REFRESH_TTL_MS = 30 * 24 * 60 * 60 * 1000; // 30 days

// Login
app.post('/api/v1/auth/login', async (req, reply) => {
  const { email, password } = req.body as { email: string; password: string };

  const user = await db.findUserByEmail(email);
  if (!user || !await argon2.verify(user.passwordHash, password)) {
    throw new AppError(401, 'INVALID_CREDENTIALS', 'Invalid email or password');
  }

  const accessToken = app.jwt.sign({ sub: user.id, role: user.role });
  const familyId = crypto.randomUUID();
  const rt = issueRefreshToken(user.id, familyId);

  await db.storeRefreshToken({
    selector: rt.selector,
    secretHash: await argon2.hash(rt.secret),  // never store the raw secret
    userId: rt.userId,
    familyId: rt.familyId,
    expiresAt: new Date(Date.now() + REFRESH_TTL_MS),
  });

  reply.send({ accessToken, refreshToken: rt.token, expiresIn: 900 });
});

// Refresh — rotate, and detect reuse of an already-rotated token
app.post('/api/v1/auth/refresh', async (req, reply) => {
  const { refreshToken } = req.body as { refreshToken: string };
  const [selector, secret] = (refreshToken ?? '').split('.');
  if (!selector || !secret) {
    throw new AppError(401, 'INVALID_TOKEN', 'Malformed refresh token');
  }

  // Single indexed lookup by selector — O(1), no hash scan.
  const row = await db.findRefreshTokenBySelector(selector);
  if (!row || !await argon2.verify(row.secretHash, secret)) {
    throw new AppError(401, 'INVALID_TOKEN', 'Invalid refresh token');
  }

  // Reuse detection: a token that's already been consumed/revoked but is
  // presented again means it was likely stolen → kill the whole family.
  if (row.consumedAt || row.revokedAt || row.expiresAt < new Date()) {
    await db.revokeRefreshTokenFamily(row.familyId);
    throw new AppError(401, 'TOKEN_REUSE_DETECTED', 'Refresh token reuse detected; session revoked');
  }

  // Rotate atomically: mark this token consumed and insert its successor in
  // one transaction so a crash can't leave the user with zero valid tokens.
  const rt = issueRefreshToken(row.userId, row.familyId);
  await db.rotateRefreshToken({
    consumeSelector: selector,
    next: {
      selector: rt.selector,
      secretHash: await argon2.hash(rt.secret),
      userId: rt.userId,
      familyId: rt.familyId,
      expiresAt: new Date(Date.now() + REFRESH_TTL_MS),
    },
  });

  const user = await db.findUser(row.userId);
  const accessToken = app.jwt.sign({ sub: user.id, role: user.role });

  reply.send({ accessToken, refreshToken: rt.token, expiresIn: 900 });
});

// Protected route
app.get('/api/v1/me', {
  preHandler: [app.authenticate],
}, async (req, reply) => {
  const user = await db.findUser(req.user.sub);
  reply.send({ data: user });
});
```

### API Keys (Service-to-Service)

```typescript
// Generate API keys
function generateApiKey(): { key: string; hash: string; prefix: string } {
  // Pick a prefix unique to your product; do not imitate another vendor's
  // format (sk_live_ is Stripe's), it confuses secret scanners.
  const key = `myapp_live_${crypto.randomBytes(32).toString('base64url')}`;
  const prefix = key.slice(0, 15);  // For identification without exposing key
  const hash = crypto.createHash('sha256').update(key).digest('hex');
  return { key, hash, prefix };
}

// Validate — always compare hashes, never raw keys
async function validateApiKey(key: string): Promise<ApiKeyRecord | null> {
  const hash = crypto.createHash('sha256').update(key).digest('hex');
  return db.findApiKeyByHash(hash);
}

// Middleware
async function apiKeyAuth(req: Request, res: Response, next: NextFunction) {
  const key = req.headers['x-api-key'] as string
    || req.headers.authorization?.replace('Bearer ', '');

  if (!key) throw new AppError(401, 'MISSING_API_KEY', 'API key required');

  const record = await validateApiKey(key);
  if (!record) throw new AppError(401, 'INVALID_API_KEY', 'Invalid API key');
  if (record.revokedAt) throw new AppError(401, 'REVOKED_API_KEY', 'API key has been revoked');

  req.apiKey = record;
  next();
}
```

---

### Resource: references/checklist-production-ready-api.md

## Checklist: Production-Ready API

- [ ] Consistent URL patterns (plural nouns, max 2 levels nesting)
- [ ] Cursor pagination for list endpoints
- [ ] RFC 9457 Problem Details error responses (`application/problem+json`) with field-level errors
- [ ] Rate limiting with `RateLimit-*` headers (IETF draft, draft-ietf-httpapi-ratelimit-headers), optionally legacy `X-RateLimit-*`
- [ ] Idempotency keys for POST endpoints (required for money-moving writes; bound to method+route+body+principal)
- [ ] Request validation from OpenAPI spec
- [ ] API versioning with deprecation/sunset headers
- [ ] Authentication (JWT for users, API keys for services)
- [ ] CORS configured correctly
- [ ] Request/response logging with correlation IDs
- [ ] Compression (gzip/brotli)
- [ ] Health check endpoint (/healthz)
- [ ] OpenAPI spec as source of truth
- [ ] Generated client SDKs from OpenAPI spec

### Resource: references/error-handling-rfc-9457-problem-details.md

## Contents

- Error Handling: RFC 9457 Problem Details
- Standard Error Response
- Error Handler Middleware
- Usage
- Error Response Examples

## Error Handling: RFC 9457 Problem Details

RFC 9457 (2023) obsoletes RFC 7807 — the wire format is unchanged, so the
object is still universally called "Problem Details" and uses the same
`application/problem+json` media type. Set that content type on error
responses so generic clients and gateways can parse them:

```typescript
res.type('application/problem+json');
```

### Standard Error Response

```typescript
// types/error.ts
interface ProblemDetail {
  type: string;          // URI reference identifying the error type
  title: string;         // Human-readable summary
  status: number;        // HTTP status code
  detail?: string;       // Human-readable explanation specific to this occurrence
  instance?: string;     // URI reference identifying this specific occurrence
  // Extensions
  errors?: FieldError[]; // Field-level validation errors
  code?: string;         // Machine-readable error code
  traceId?: string;      // For debugging
}

interface FieldError {
  field: string;
  message: string;
  code: string;
}
```

### Error Handler Middleware

```typescript
// middleware/error-handler.ts
import { Request, Response, NextFunction } from 'express';

class AppError extends Error {
  constructor(
    public statusCode: number,
    public code: string,
    message: string,
    public errors?: FieldError[],
  ) {
    super(message);
    this.name = 'AppError';
  }
}

// Specific error classes
class NotFoundError extends AppError {
  constructor(resource: string, id: string) {
    super(404, 'RESOURCE_NOT_FOUND', `${resource} with id '${id}' not found`);
  }
}

class ValidationError extends AppError {
  constructor(errors: FieldError[]) {
    super(422, 'VALIDATION_ERROR', 'Request validation failed', errors);
  }
}

class ConflictError extends AppError {
  constructor(message: string) {
    super(409, 'CONFLICT', message);
  }
}

class RateLimitError extends AppError {
  constructor(retryAfter: number) {
    super(429, 'RATE_LIMITED', `Rate limit exceeded. Retry after ${retryAfter}s`);
  }
}

// The error handler
function errorHandler(err: Error, req: Request, res: Response, _next: NextFunction) {
  const requestId = req.headers['x-request-id'] as string;

  if (err instanceof AppError) {
    return res.status(err.statusCode).json({
      type: `https://api.example.com/errors/${err.code.toLowerCase()}`,
      title: err.code.replace(/_/g, ' ').toLowerCase(),
      status: err.statusCode,
      detail: err.message,
      instance: req.originalUrl,
      code: err.code,
      errors: err.errors,
      traceId: requestId,
    });
  }

  // Unexpected errors — log full details, return generic message
  req.log?.error({ err }, 'Unhandled error');

  res.status(500).json({
    type: 'https://api.example.com/errors/internal',
    title: 'Internal Server Error',
    status: 500,
    detail: 'An unexpected error occurred',
    instance: req.originalUrl,
    code: 'INTERNAL_ERROR',
    traceId: requestId,
  });
}

app.use(errorHandler);
```

### Usage

```typescript
app.get('/api/v1/users/:id', async (req, res) => {
  const user = await db.findUser(req.params.id);
  if (!user) throw new NotFoundError('User', req.params.id);
  res.json({ data: user });
});

app.post('/api/v1/users', async (req, res) => {
  const errors: FieldError[] = [];
  if (!req.body.email) errors.push({ field: 'email', message: 'Email is required', code: 'REQUIRED' });
  if (!req.body.name) errors.push({ field: 'name', message: 'Name is required', code: 'REQUIRED' });
  if (errors.length) throw new ValidationError(errors);

  const existing = await db.findUserByEmail(req.body.email);
  if (existing) throw new ConflictError('A user with this email already exists');

  const user = await db.createUser(req.body);
  res.status(201).json({ data: user });
});
```

### Error Response Examples

`404 Not Found`:

```json
{
  "type": "https://api.example.com/errors/resource_not_found",
  "title": "resource not found",
  "status": 404,
  "detail": "User with id 'abc-123' not found",
  "instance": "/api/v1/users/abc-123",
  "code": "RESOURCE_NOT_FOUND",
  "traceId": "req-xyz-789"
}
```

`422 Unprocessable Entity` with field-level errors:

```json
{
  "type": "https://api.example.com/errors/validation_error",
  "title": "validation error",
  "status": 422,
  "detail": "Request validation failed",
  "code": "VALIDATION_ERROR",
  "errors": [
    { "field": "email", "message": "Must be a valid email address", "code": "INVALID_FORMAT" },
    { "field": "age", "message": "Must be at least 18", "code": "MIN_VALUE" }
  ]
}
```

---

### Resource: references/graphql-vs-rest-decision-matrix.md

## GraphQL vs REST: Decision Matrix

| Factor | REST | GraphQL |
|--------|------|---------|
| **Use when** | CRUD-heavy, well-defined resources | Complex relationships, varying client needs |
| **Caching** | HTTP caching works perfectly | Requires custom caching (Apollo, Relay) |
| **Versioning** | URL versioning, straightforward | Schema evolution, deprecation directives |
| **File uploads** | Multipart form, straightforward | Requires separate upload endpoint or multipart spec |
| **Real-time** | SSE, WebSocket (separate) | Subscriptions (built-in) |
| **Tooling** | Mature (Postman, curl) | Specialized (GraphiQL, Apollo DevTools) |
| **N+1 problem** | Solved by design (one endpoint = one response) | Requires DataLoader |
| **Mobile** | Over-fetching without field selection | Precise data fetching |
| **Team size** | Any | Better with dedicated frontend/backend teams |

**Strong REST signals:** Public API, simple CRUD, caching matters, small team.
**Strong GraphQL signals:** Multiple clients (web, mobile, partners) with different data needs, deeply nested relationships, rapid frontend iteration.

**Don't use GraphQL because it's trendy.** Use it when you genuinely have the data-fetching complexity that justifies it.

---

### Resource: references/idempotency.md

## Contents

- Idempotency
- Idempotency Keys for Safe Retries

## Idempotency

### Idempotency Keys for Safe Retries

Design rules this middleware enforces:

- **Bind the cache to the full request, not just the key.** Cache under a
  hash of `method + route + authenticated principal + request-body`. Reusing
  one key across two different POSTs (or with a changed body) must NOT replay
  the first response — return `422` on a key/body mismatch instead.
- **Release the lock on every exit path** (`finish`, `close`, and errors),
  not only inside a `res.json` patch — otherwise thrown errors, non-JSON or
  streaming responses, and crashes strand the lock until its short TTL.
- **Cache outcomes intentionally.** Persist deterministic results — `2xx`
  and client errors (`4xx`, e.g. validation) — so retries are stable. Do NOT
  cache `5xx`/timeouts: those are transient and the client should be able to
  retry into a fresh attempt.
- **Two TTLs.** A short *lock* TTL (seconds, in case the process dies mid-flight)
  and a longer *result* TTL (hours/days) for the cached response.

```typescript
// middleware/idempotency.ts
import { createHash } from 'crypto';

const LOCK_TTL = 60;          // seconds — bounds a crash that strands the lock
const RESULT_TTL = 24 * 3600; // seconds — replay window for a completed request

// Endpoints where a missing key is a hard error (money-moving / side-effectful).
const REQUIRE_KEY = [/^\/api\/v1\/payments/, /^\/api\/v1\/transfers/];

function fingerprint(req: Request, key: string): string {
  const principal = (req as any).user?.id ?? (req as any).apiKey?.id ?? 'anon';
  const body = createHash('sha256').update(JSON.stringify(req.body ?? {})).digest('hex');
  // route (not originalUrl) so query strings don't fragment the key
  const route = (req as any).route?.path ?? req.path;
  return createHash('sha256')
    .update([req.method, route, principal, key, body].join('\n'))
    .digest('hex');
}

async function idempotency(req: Request, res: Response, next: NextFunction) {
  if (req.method !== 'POST') return next();

  const idempotencyKey = req.headers['idempotency-key'] as string | undefined;
  if (!idempotencyKey) {
    if (REQUIRE_KEY.some((re) => re.test(req.path))) {
      throw new AppError(400, 'IDEMPOTENCY_KEY_REQUIRED',
        'Idempotency-Key header is required for this endpoint');
    }
    return next();  // optional elsewhere
  }

  const fp = fingerprint(req, idempotencyKey);
  const resultKey = `idem:res:${fp}`;
  const lockKey = `idem:lock:${fp}`;
  // Detects "same key, different request" → reject rather than replay.
  const keyGuard = `idem:key:${idempotencyKey}`;

  const cached = await redis.get(resultKey);
  if (cached) {
    const { statusCode, body } = JSON.parse(cached);
    res.setHeader('Idempotent-Replayed', 'true');
    return res.status(statusCode).json(body);
  }

  // Reject reuse of the same key with a different method/route/body.
  const priorFp = await redis.set(keyGuard, fp, 'EX', RESULT_TTL, 'NX', 'GET') as string | null;
  if (priorFp && priorFp !== fp) {
    throw new AppError(422, 'IDEMPOTENCY_KEY_REUSED',
      'This Idempotency-Key was already used with a different request');
  }

  const locked = await redis.set(lockKey, '1', 'EX', LOCK_TTL, 'NX');
  if (!locked) {
    throw new AppError(409, 'REQUEST_IN_PROGRESS',
      'A request with this idempotency key is already being processed');
  }

  // Capture the final payload, then persist + unlock on ANY terminal event.
  let captured: { statusCode: number; body: unknown } | undefined;
  const originalJson = res.json.bind(res);
  res.json = (body: unknown) => {
    captured = { statusCode: res.statusCode, body };
    return originalJson(body);
  };

  let settled = false;
  const settle = async () => {
    if (settled) return;
    settled = true;
    // Cache deterministic outcomes (2xx + client errors); never cache 5xx.
    if (captured && captured.statusCode < 500) {
      await redis.set(resultKey, JSON.stringify(captured), 'EX', RESULT_TTL);
    } else {
      await redis.del(keyGuard);  // let the client retry a failed attempt cleanly
    }
    await redis.del(lockKey);     // always release, even on error/stream/abort
  };
  res.on('finish', settle);  // response fully sent
  res.on('close', settle);   // client aborted before finish

  next();
}

app.use('/api/v1', idempotency);
```

> Note: the `SET ... GET` option requires Redis ≥ 7.0. On older servers,
> replace the `keyGuard` step with a `GET` then a `SET ... NX`.

Client usage:
```typescript
// Client retries safely
const response = await fetch('/api/v1/payments', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'Idempotency-Key': crypto.randomUUID(),  // Generate once, retry with same key
  },
  body: JSON.stringify({ amount: 5000, currency: 'usd' }),
});
```

---

### Resource: references/openapi-3-1-specification.md

## Contents

- OpenAPI 3.1 Specification
- Complete Example
- Validation Middleware from OpenAPI Spec

## OpenAPI 3.1 Specification

### Complete Example

```yaml
openapi: 3.1.0
info:
  title: Users API
  version: 1.0.0
  description: User management API
  contact:
    email: api@example.com
  license:
    name: MIT

servers:
  - url: https://api.example.com/v1
    description: Production
  - url: https://staging-api.example.com/v1
    description: Staging

security:
  - bearerAuth: []

paths:
  /users:
    get:
      operationId: listUsers
      summary: List users
      tags: [Users]
      parameters:
        - name: cursor
          in: query
          schema:
            type: string
        - name: limit
          in: query
          schema:
            type: integer
            minimum: 1
            maximum: 100
            default: 20
        - name: status
          in: query
          schema:
            type: string
            enum: [active, inactive, suspended]
        - name: sort
          in: query
          schema:
            type: string
            default: -created_at
      responses:
        '200':
          description: Users list
          content:
            application/json:
              schema:
                type: object
                required: [data, pagination]
                properties:
                  data:
                    type: array
                    items:
                      $ref: '#/components/schemas/User'
                  pagination:
                    $ref: '#/components/schemas/CursorPagination'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '429':
          $ref: '#/components/responses/RateLimited'

    post:
      operationId: createUser
      summary: Create user
      tags: [Users]
      parameters:
        - name: Idempotency-Key
          in: header
          schema:
            type: string
            format: uuid
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/CreateUserRequest'
      responses:
        '201':
          description: User created
          content:
            application/json:
              schema:
                type: object
                properties:
                  data:
                    $ref: '#/components/schemas/User'
        '422':
          $ref: '#/components/responses/ValidationError'

components:
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      bearerFormat: JWT
    apiKey:
      type: apiKey
      in: header
      name: X-API-Key

  schemas:
    User:
      type: object
      required: [id, email, name, status, created_at]
      properties:
        id:
          type: string
          format: uuid
        email:
          type: string
          format: email
        name:
          type: string
        status:
          type: string
          enum: [active, inactive, suspended]
        created_at:
          type: string
          format: date-time
        updated_at:
          type: string
          format: date-time

    CreateUserRequest:
      type: object
      required: [email, name]
      properties:
        email:
          type: string
          format: email
        name:
          type: string
          minLength: 1
          maxLength: 100
        role:
          type: string
          enum: [user, admin]
          default: user

    CursorPagination:
      type: object
      properties:
        next_cursor:
          type: [string, "null"]
        has_more:
          type: boolean

    ProblemDetail:
      type: object
      required: [type, title, status]
      properties:
        type:
          type: string
          format: uri
        title:
          type: string
        status:
          type: integer
        detail:
          type: string
        code:
          type: string
        errors:
          type: array
          items:
            type: object
            properties:
              field:
                type: string
              message:
                type: string
              code:
                type: string

  responses:
    Unauthorized:
      description: Authentication required
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ProblemDetail'

    ValidationError:
      description: Validation failed
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ProblemDetail'

    RateLimited:
      description: Rate limit exceeded
      headers:
        Retry-After:
          schema:
            type: integer
        RateLimit-Limit:        # IETF draft rate-limit headers (draft-ietf-httpapi-ratelimit-headers)
          schema:
            type: integer
        RateLimit-Remaining:
          schema:
            type: integer
        RateLimit-Reset:        # seconds until reset (delta), not epoch
          schema:
            type: integer
        X-RateLimit-Limit:      # legacy, optional
          schema:
            type: integer
        X-RateLimit-Remaining:
          schema:
            type: integer
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ProblemDetail'
```

### Validation Middleware from OpenAPI Spec

```typescript
import { OpenApiValidator } from 'express-openapi-validator';

app.use(
  OpenApiValidator.middleware({
    apiSpec: './openapi.yaml',
    validateRequests: true,
    validateResponses: process.env.NODE_ENV !== 'production',  // Dev only
    validateSecurity: false,  // Handle auth separately
  }),
);
```

---

### Resource: references/pagination-cursor-vs-offset.md

## Contents

- Pagination: Cursor vs Offset
- Offset Pagination (Simple, Flawed)
- Cursor Pagination (Production-Grade)
- Keyset Pagination for Large Datasets

## Pagination: Cursor vs Offset

### Offset Pagination (Simple, Flawed)

```typescript
// Simple but problematic for large datasets
app.get('/api/v1/users', async (req, res) => {
  const page = parseInt(req.query.page as string) || 1;
  const limit = Math.min(parseInt(req.query.limit as string) || 20, 100);
  const offset = (page - 1) * limit;

  const [users, total] = await Promise.all([
    db.query('SELECT * FROM users ORDER BY id LIMIT $1 OFFSET $2', [limit, offset]),
    db.query('SELECT COUNT(*) FROM users'),
  ]);

  res.json({
    data: users.rows,
    pagination: {
      page,
      limit,
      total: parseInt(total.rows[0].count),
      totalPages: Math.ceil(parseInt(total.rows[0].count) / limit),
    },
  });
});
```

**Problems with offset pagination:**
- `OFFSET 100000` scans and discards 100k rows — O(n)
- Inserting/deleting rows between pages causes duplicates/gaps
- COUNT(*) on large tables is slow

### Cursor Pagination (Production-Grade)

```typescript
// Cursor-based — consistent, performant, no skipping
app.get('/api/v1/users', async (req, res) => {
  const limit = Math.min(parseInt(req.query.limit as string) || 20, 100);
  const cursor = req.query.cursor as string | undefined;

  let query = 'SELECT * FROM users';
  const params: any[] = [limit + 1]; // Fetch one extra to detect hasMore

  if (cursor) {
    const decoded = decodeCursor(cursor); // { id: 123, created_at: '2024-01-01' }
    query += ' WHERE (created_at, id) < ($2, $3)';
    params.push(decoded.created_at, decoded.id);
  }

  query += ' ORDER BY created_at DESC, id DESC LIMIT $1';

  const result = await db.query(query, params);
  const hasMore = result.rows.length > limit;
  const items = hasMore ? result.rows.slice(0, -1) : result.rows;

  const nextCursor = hasMore
    ? encodeCursor({
        id: items[items.length - 1].id,
        created_at: items[items.length - 1].created_at,
      })
    : null;

  res.json({
    data: items,
    pagination: {
      next_cursor: nextCursor,
      has_more: hasMore,
    },
  });
});

// Cursor encoding — base64 JSON (not security, just obfuscation)
function encodeCursor(data: Record<string, any>): string {
  return Buffer.from(JSON.stringify(data)).toString('base64url');
}

function decodeCursor(cursor: string): Record<string, any> {
  return JSON.parse(Buffer.from(cursor, 'base64url').toString());
}
```

### Keyset Pagination for Large Datasets

For tables with 10M+ rows, keyset pagination on an indexed column:

```sql
-- Requires composite index: CREATE INDEX idx_users_created_id ON users(created_at DESC, id DESC);
SELECT * FROM users
WHERE (created_at, id) < ('2024-06-15 10:30:00', 12345)
ORDER BY created_at DESC, id DESC
LIMIT 20;
-- Avoids O(offset) scans entirely: one indexed seek to the cursor position,
-- then reads O(limit) rows — cost stays flat no matter how deep you paginate.
-- The trailing id is the tie-breaker: include a unique column in BOTH the
-- WHERE comparison and ORDER BY, or rows sharing a created_at can be skipped
-- or duplicated across pages.
```

---

### Resource: references/rate-limiting.md

## Contents

- Rate Limiting
- Sliding Window Log with Redis (Production)
- Atomic Token Bucket (Lua) — constant memory, allows bursts

## Rate Limiting

### Sliding Window Log with Redis (Production)

A sorted set stores one member per request, scored by timestamp. Each call
trims entries older than the window, adds the current request, and counts
what remains — giving an exact rolling count with no fixed-window burst
seam. Cost is O(log N) per request and memory is O(requests-in-window) per
key, so for very high-volume limits prefer a token-bucket / GCRA counter
(constant memory) — see the atomic Lua variant below.

```typescript
import Redis from 'ioredis';

const redis = new Redis(process.env.REDIS_URL);

interface RateLimitResult {
  allowed: boolean;
  remaining: number;
  resetAt: number;
  retryAfter?: number;
}

async function checkRateLimit(
  key: string,
  maxRequests: number,
  windowSeconds: number,
): Promise<RateLimitResult> {
  const now = Math.floor(Date.now() / 1000);
  const windowStart = now - windowSeconds;

  // Sliding window log using sorted set
  const pipeline = redis.pipeline();
  pipeline.zremrangebyscore(key, 0, windowStart);      // Remove old entries
  pipeline.zadd(key, now.toString(), `${now}:${Math.random()}`);  // Add current
  pipeline.zcard(key);                                   // Count in window
  pipeline.expire(key, windowSeconds);                   // TTL cleanup

  const results = await pipeline.exec();
  const count = results![2][1] as number;

  if (count > maxRequests) {
    const oldestInWindow = await redis.zrange(key, 0, 0, 'WITHSCORES');
    const retryAfter = oldestInWindow.length >= 2
      ? parseInt(oldestInWindow[1]) + windowSeconds - now
      : windowSeconds;

    return {
      allowed: false,
      remaining: 0,
      resetAt: now + retryAfter,
      retryAfter,
    };
  }

  return {
    allowed: true,
    remaining: maxRequests - count,
    resetAt: now + windowSeconds,
  };
}

// Middleware
function rateLimit(maxRequests: number, windowSeconds: number) {
  return async (req: Request, res: Response, next: NextFunction) => {
    // Per-user if authenticated, per-IP otherwise
    const key = req.user
      ? `ratelimit:user:${req.user.id}`
      : `ratelimit:ip:${req.ip}`;

    const result = await checkRateLimit(key, maxRequests, windowSeconds);

    // IETF draft headers (draft-ietf-httpapi-ratelimit-headers, still an
    // Internet-Draft, not an RFC). The latest draft consolidates these into
    // RateLimit and RateLimit-Policy structured fields; the Limit/Remaining/Reset
    // trio below matches earlier drafts and stays the most widely deployed form.
    // `RateLimit-Reset` is seconds-until-reset
    // (a delta), not an epoch timestamp — that's the key difference from the
    // legacy `X-RateLimit-Reset` convention below.
    const resetDelta = Math.max(0, result.resetAt - Math.floor(Date.now() / 1000));
    res.setHeader('RateLimit-Limit', maxRequests);
    res.setHeader('RateLimit-Remaining', result.remaining);
    res.setHeader('RateLimit-Reset', resetDelta);

    // Legacy headers — keep for older clients; `X-RateLimit-Reset` is an epoch.
    res.setHeader('X-RateLimit-Limit', maxRequests);
    res.setHeader('X-RateLimit-Remaining', result.remaining);
    res.setHeader('X-RateLimit-Reset', result.resetAt);

    if (!result.allowed) {
      res.setHeader('Retry-After', result.retryAfter!);  // seconds (RFC 9110)
      throw new RateLimitError(result.retryAfter!);
    }

    next();
  };
}

// Different limits for different endpoints
app.use('/api/v1/auth', rateLimit(10, 60));       // 10/min for auth
app.use('/api/v1/', rateLimit(100, 60));           // 100/min general
app.use('/api/v1/search', rateLimit(30, 60));      // 30/min for search
```

### Atomic Token Bucket (Lua) — constant memory, allows bursts

The sliding-window pipeline above is two round-trips and stores one key per
request. A token bucket runs as a single atomic Lua script (no race between
read and write under concurrency), uses O(1) memory per key, and naturally
permits short bursts up to `capacity` while enforcing a steady refill rate.

```typescript
// Refills `refillRate` tokens/sec up to `capacity`; each request costs 1 token.
// KEYS[1] = bucket key. ARGV: capacity, refillRate, now (sec, fractional), cost.
const TOKEN_BUCKET = `
local key        = KEYS[1]
local capacity   = tonumber(ARGV[1])
local refillRate = tonumber(ARGV[2])
local now        = tonumber(ARGV[3])
local cost       = tonumber(ARGV[4])

local state   = redis.call('HMGET', key, 'tokens', 'ts')
local tokens  = tonumber(state[1])
local ts      = tonumber(state[2])
if tokens == nil then tokens = capacity; ts = now end

-- Refill based on elapsed time, cap at capacity
tokens = math.min(capacity, tokens + (now - ts) * refillRate)

local allowed = 0
if tokens >= cost then
  allowed = 1
  tokens = tokens - cost
end

redis.call('HSET', key, 'tokens', tokens, 'ts', now)
-- Expire when the bucket would be full again (idle reclaim)
redis.call('EXPIRE', key, math.ceil(capacity / refillRate) + 1)

-- Seconds until enough tokens for one request (0 if allowed now)
local retry = 0
if allowed == 0 then retry = (cost - tokens) / refillRate end
return { allowed, tostring(tokens), tostring(retry) }
`;

const sha = await redis.script('LOAD', TOKEN_BUCKET);

async function checkTokenBucket(
  key: string, capacity: number, refillRate: number, cost = 1,
): Promise<RateLimitResult> {
  const now = Date.now() / 1000;
  const [allowed, tokensStr, retryStr] = (await redis.evalsha(
    sha, 1, key, capacity, refillRate, now, cost,
  )) as [number, string, string];
  const remaining = Math.floor(parseFloat(tokensStr));
  const retryAfter = Math.ceil(parseFloat(retryStr));
  return {
    allowed: allowed === 1,
    remaining,
    resetAt: Math.floor(now) + Math.ceil((capacity - remaining) / refillRate),
    ...(allowed === 1 ? {} : { retryAfter }),
  };
}
// e.g. checkTokenBucket('ratelimit:user:42', 100, 100 / 60) → 100 burst, refills to 100/min
```

---

### Resource: references/response-envelope.md

## Response Envelope

```typescript
// Consistent response format
interface ApiResponse<T> {
  data: T;
  meta?: Record<string, any>;
  pagination?: CursorPagination;
}

// Always wrap in { data: ... }
// Single item:  { "data": { "id": "123", "name": "John" } }
// List:         { "data": [...], "pagination": { "next_cursor": "...", "has_more": true } }
// Error:        RFC 9457 Problem Details (no data wrapper)

// Why? Consistent parsing, easy to add metadata, forward-compatible
```

---

### Resource: references/rest-conventions-that-actually-matter.md

## Contents

- REST Conventions That Actually Matter
- URL Design
- Filtering, Sorting, Pagination

## REST Conventions That Actually Matter

Forget the academic debates about REST maturity levels. Here's what matters in practice:

### URL Design

```
# Resources are nouns, plural
GET    /api/v1/users              # List users
POST   /api/v1/users              # Create user
GET    /api/v1/users/:id          # Get user
PATCH  /api/v1/users/:id          # Partial update
PUT    /api/v1/users/:id          # Full replace (rare)
DELETE /api/v1/users/:id          # Delete user

# Nesting: max 2 levels deep
GET    /api/v1/users/:id/orders           # User's orders
GET    /api/v1/users/:id/orders/:orderId  # Specific order

# Don't nest deeper — use query params instead
# BAD:  /api/v1/users/:id/orders/:orderId/items/:itemId
# GOOD: /api/v1/order-items/:itemId
# GOOD: /api/v1/orders/:orderId/items?expand=product

# Actions that don't map to CRUD — use verb sub-resources
POST   /api/v1/users/:id/verify-email
POST   /api/v1/orders/:id/cancel
POST   /api/v1/reports/generate
```

### Filtering, Sorting, Pagination

```
# Filtering — use query params with field names
GET /api/v1/users?status=active&role=admin&created_after=2024-01-01

# Sorting — comma-separated, prefix with - for descending
GET /api/v1/users?sort=-created_at,name

# Field selection — reduce payload
GET /api/v1/users?fields=id,name,email

# Search — use q for full-text
GET /api/v1/users?q=john&status=active

# Combining
GET /api/v1/orders?status=pending&sort=-created_at&limit=20&cursor=eyJ...
```

---

---

## ascii-banner
Category: design
Description: Build animated ASCII/Unicode banners for CLI tools and web UIs — frame animation, ANSI color, terminal capability detection, flicker-free rendering, accessibility, and canvas/WebGL ASCII shaders. Use when adding a CLI startup banner, building terminal-aesthetic web UIs, converting images/3D scenes to ASCII, or making banner animation TTY-safe.
Features:
  - Frame-based CLI animation with flicker-free rendering
  - ANSI color role system (4-bit, 8-bit, 24-bit with detection)
  - Terminal capability detection and graceful degradation
  - Accessibility: reduced motion, screen reader safe, opt-in animation
  - Web canvas ASCII renderer (image/video/3D to ASCII)
  - Three.js ASCII post-processing for web UIs
  - figlet text banner generation
  - Static image-to-ASCII conversion (Python)
Use Cases:
  - Create an animated splash screen for a CLI tool
  - Build a web hero section with ASCII shader effect
  - Convert a logo to ASCII art for terminal display
  - Add a branded animation to a dev tool startup

# Animated ASCII Banners

## Overview

Animated ASCII banners create personality in CLI tools and terminal-aesthetic web UIs. This skill covers both terminal-native (Node.js/Python CLI) and web-based (canvas/WebGL) implementations.

**Key challenges:** Terminal inconsistency, ANSI color fragmentation, screen reader accessibility, flicker prevention, and cross-platform rendering.

## Part 1: Terminal ASCII Animation (CLI)

### 1. Frame-Based Animation Architecture

```
project/
  frames/           # Each .txt file is one animation frame
    frame-001.txt
    frame-002.txt
    ...
  colors/           # Color map per frame (optional)
    frame-001.json
  src/
    renderer.ts     # Animation engine
    palette.ts      # ANSI color role mapping
    detect.ts       # Terminal capability detection
```

### 2. Basic Animation Loop (Node.js)

```javascript
import fs from "fs";
import readline from "readline";

const frames = fs
  .readdirSync("./frames")
  .filter(f => f.endsWith(".txt"))
  .sort()
  .map(f => fs.readFileSync(`./frames/${f}`, "utf8"));

let current = 0;
let running = true;

function render() {
  if (!running) return;
  readline.cursorTo(process.stdout, 0, 0);
  readline.clearScreenDown(process.stdout);
  process.stdout.write(frames[current]);
  current = (current + 1) % frames.length;
}

// 75ms = ~13fps — safe for most terminals
const interval = setInterval(render, 75);

// Graceful cleanup
process.on("SIGINT", () => {
  running = false;
  clearInterval(interval);
  readline.cursorTo(process.stdout, 0, 0);
  readline.clearScreenDown(process.stdout);
  process.exit(0);
});

// Auto-stop after one loop
setTimeout(() => {
  clearInterval(interval);
  running = false;
}, frames.length * 75);
```

### 3. ANSI Color System

**Use semantic color roles, not hardcoded values.** Terminals remap colors based on user themes.

```javascript
// Color role mapping — degrade gracefully across terminals
const ANSI_ROLES = {
  primary:   "\x1b[32m",   // Green (accent)
  secondary: "\x1b[36m",   // Cyan
  highlight: "\x1b[97m",   // Bright white
  shadow:    "\x1b[90m",   // Dark gray
  dim:       "\x1b[2m",    // Dim modifier
  reset:     "\x1b[0m",
};

function colorize(char, role) {
  if (!role || role === "none") return char;
  return `${ANSI_ROLES[role] || ""}${char}${ANSI_ROLES.reset}`;
}
```

**ANSI color modes:**

| Mode | Colors | Support | Use |
|------|--------|---------|-----|
| 4-bit | 16 colors | Universal | Safe default — use this |
| 8-bit | 256 colors | Most modern terminals | Extended palette |
| 24-bit (truecolor) | 16M colors | iTerm2, Kitty, modern terminals | Brand-exact colors |

**Terminal detection:**
```javascript
function getColorSupport() {
  const env = process.env;
  if (env.NO_COLOR) return "none";
  if (env.COLORTERM === "truecolor" || env.COLORTERM === "24bit") return "24bit";
  if (env.TERM_PROGRAM === "iTerm.app") return "24bit";
  if (env.TERM?.includes("256color")) return "8bit";
  if (process.stdout.isTTY) return "4bit";
  return "none";
}
```

### 4. Flicker Prevention

**Problem:** `clearScreen` + full repaint causes visible flicker.

**Solution:** Differential rendering — only repaint changed characters:

```javascript
let previousFrame = "";

function renderDiff(frame) {
  const lines = frame.split("\n");
  const prevLines = previousFrame.split("\n");

  for (let y = 0; y < lines.length; y++) {
    if (lines[y] !== prevLines[y]) {
      readline.cursorTo(process.stdout, 0, y);
      process.stdout.write(lines[y] + "\x1b[K"); // Clear to end of line
    }
  }
  previousFrame = frame;
}
```

**Additional techniques:**
- Use alternate screen buffer (`\x1b[?1049h` to enter, `\x1b[?1049l` to exit)
- Hide cursor during animation (`\x1b[?25l`, restore with `\x1b[?25h`)
- Batch writes using a string buffer, write once per frame

### 5. Accessibility

Terminals have **no `aria-live` equivalent** — a screen reader reads stdout linearly, so every cursor-repositioned frame you write risks being announced as new text (a "chatterbox" of garbage). The accessible pattern is the opposite of the web one: **emit a single static text line, then do all animation with raw escape codes that assistive tech ignores or that you suppress entirely when not on an interactive TTY.**

**Mandatory requirements:**

| Requirement | Concrete CLI implementation |
|-------------|------------------------------|
| Opt-in / opt-out | Animate only behind intent (`--banner`/`--animate`), and always honor `--no-banner`. Never auto-play in CI or pipes. |
| Static alternative | Before animating, print one plain line (e.g. `skills.ws — CLI v1.2.0`). This is what a screen reader and a logfile actually read; the animation is decoration layered on top. |
| Don't re-announce | Never rewrite the *text content* every frame. Either animate inside the alternate screen buffer (which most SR setups skip) or only repaint changed cells (see §4) so unchanged glyphs aren't re-emitted. |
| Reduced motion | There is **no standardized terminal "reduce motion" signal.** Disable animation when stdout is not a TTY, when `NO_COLOR` is set (users who set it generally want quiet output), when `TERM=dumb`, and via your own opt-out flag/config. On the **web** use the real standard: `@media (prefers-reduced-motion: reduce)`. |
| Graceful degradation | Static ASCII art fallback (no escapes) whenever animation is disabled by any check above. |
| Color-independent | Art must read by shape, not color — verify it's recognizable in `NO_COLOR=1` mode. |

```javascript
// CLI gate: only animate when it's safe and wanted.
// Note: NO_ANIMATION / REDUCE_MOTION are conventions, not standards —
// support them defensively but rely on the TTY/flag checks as the real signal.
function shouldAnimate({ flags = {} } = {}) {
  if (flags.noBanner) return false;            // explicit opt-out wins
  if (!process.stdout.isTTY) return false;     // piped, redirected, or CI
  if (process.env.CI) return false;            // build logs
  if (process.env.TERM === "dumb") return false;
  if (process.env.NO_COLOR) return false;      // user wants quiet output
  if (process.env.NO_ANIMATION) return false;  // de-facto convention
  if (process.env.REDUCE_MOTION) return false; // some apps export this; non-standard
  return true;
}

// Always emit the accessible line; animate only as decoration on top.
function showBanner(flags) {
  console.log("skills.ws — CLI v1.2.0");   // read by SR / captured in logs
  if (shouldAnimate({ flags })) startAnimation();
}
```

> Web equivalent: gate animation with the standardized media query, not an env var:
> ```javascript
> const reduce = window.matchMedia("(prefers-reduced-motion: reduce)").matches;
> if (!reduce) startWebAnimation();
> ```

### 6. ASCII Art Design

**Pick a character set deliberately — "ASCII" and "terminal art" are not the same.** True ASCII is bytes 0x20–0x7E and renders identically everywhere (logs, dumb terminals, Windows `cmd`, CI). The box/block/arrow glyphs below are **Unicode** — gorgeous in iTerm2/Kitty/Windows Terminal but they mojibake on legacy code pages or under the wrong locale. Treat them as two distinct modes and pick based on `process.stdout.isTTY` + a UTF-8 locale check (`/utf-?8/i.test(process.env.LC_ALL || process.env.LC_CTYPE || process.env.LANG || "")`).

**Mode A — ASCII-safe (universal, 0x20–0x7E only):**
```
Shading (light → dense):  . : - = + * # %  @
Borders:                  + - | =   (corners: + , edges: - and |)
Fills:                    . , ; * # @   (no solid blocks exist in ASCII)
Geometry / arrows:        / \ < > ^ v  (use v for down-arrow)
```

**Mode B — Unicode-enhanced (UTF-8 TTYs only; falls back to Mode A):**
```
Box-drawing:  ┌ ─ ┐ │ └ ┘ ╔ ═ ╗ ║ ╚ ╝ ├ ┤ ┬ ┴ ┼
Block fills:  ░ ▒ ▓ █ ▄ ▀ ▐ ▌   (▁▂▃▄▅▆▇█ for vertical bars/sparklines)
Geometry:     ╱ ╲ △ ▽ ◇ ○ ●
Arrows:       → ← ↑ ↓ ⟶ ⟵
```

```javascript
const supportsUnicode =
  process.stdout.isTTY &&
  /utf-?8/i.test(process.env.LC_ALL || process.env.LC_CTYPE || process.env.LANG || "");
const charset = supportsUnicode ? UNICODE_SET : ASCII_SET;
```

**figlet for text banners.** The figlet npm package ships its own `figlet` CLI (since v1.6.0), so no separate wrapper package is needed. Use `npx` (downloads/runs the CLI without a global install), or install it globally, or call the library from code.

```bash
# No-install, one-off (recommended): npx fetches the CLI on demand
npx figlet -f Slant "SKILLS"

# OR install globally so `figlet` is on PATH
npm i -g figlet
figlet -f Slant "SKILLS"

# Python equivalent (pyfiglet ships a console script):
pip install pyfiglet
pyfiglet -f slant "SKILLS"
```

```javascript
// Or use the figlet npm package as a library (after `npm install figlet`):
import figlet from "figlet";
console.log(figlet.textSync("SKILLS", { font: "Slant" }));
```

**Popular figlet fonts:** `Slant`, `Banner3`, `Big`, `Doom`, `Standard`, `Small` (npm font names are capitalized; the system `figlet`/`pyfiglet` binaries accept lowercase like `slant`). List installed fonts with `figlet -l` or `pyfiglet -l`.

## Part 2: Web ASCII Animation (Canvas/WebGL)

### 7. Canvas-Based ASCII Renderer

Convert any visual (3D scene, video, image) to ASCII in the browser:

```javascript
const CHARS = " .:-=+*#%@";

function renderAscii(ctx, canvas, source, cellW, cellH) {
  // Draw source to small offscreen canvas
  const cols = Math.floor(canvas.width / cellW);
  const rows = Math.floor(canvas.height / cellH);
  // Create the offscreen canvas once outside the render loop and reuse it
  // (resize only when cols/rows change) instead of allocating per frame.
  const offscreen = new OffscreenCanvas(cols, rows);
  const offCtx = offscreen.getContext("2d", { willReadFrequently: true });
  offCtx.drawImage(source, 0, 0, cols, rows);
  const pixels = offCtx.getImageData(0, 0, cols, rows).data;

  ctx.fillStyle = "#0a0a0a";
  ctx.fillRect(0, 0, canvas.width, canvas.height);
  ctx.font = `${cellH - 2}px monospace`;

  for (let y = 0; y < rows; y++) {
    for (let x = 0; x < cols; x++) {
      const i = (y * cols + x) * 4;
      const brightness = (pixels[i] * 0.299 + pixels[i+1] * 0.587 + pixels[i+2] * 0.114) / 255;
      if (brightness < 0.02) continue;

      const char = CHARS[Math.floor(brightness * (CHARS.length - 1))];
      const green = Math.floor(40 + brightness * 215);
      ctx.fillStyle = `rgba(0,${green},${Math.floor(green*0.55)},${0.3 + brightness * 0.7})`;
      ctx.fillText(char, x * cellW, y * cellH + cellH - 2);
    }
  }
}
```

### 8. Three.js + ASCII Post-Processing

Render an animated 3D scene to an offscreen WebGL buffer, then feed that buffer to the `renderAscii` function from §7. Complete, runnable example (Three.js r150+; verify the import path and API against your installed version — `THREE.WebGLRenderer` and ES-module imports are stable through mid-2026, but minor APIs drift):

```javascript
import * as THREE from "three";

// --- Sizing (drives BOTH the WebGL buffer and the visible ASCII canvas) ---
const WIDTH = 480, HEIGHT = 320;
const CELL_W = 8, CELL_H = 14; // monospace cell size in px

// --- Visible ASCII canvas (what the user sees) ---
const asciiCanvas = document.getElementById("ascii");
asciiCanvas.width = WIDTH;
asciiCanvas.height = HEIGHT;
const asciiCtx = asciiCanvas.getContext("2d", { willReadFrequently: false });

// --- Scene ---
const scene = new THREE.Scene();
const geometry = new THREE.TorusKnotGeometry(1, 0.35, 128, 32);
const material = new THREE.MeshStandardMaterial({ color: 0x00ff88 });
const mesh = new THREE.Mesh(geometry, material);
scene.add(mesh);

// --- Camera (REQUIRED — was missing) ---
const camera = new THREE.PerspectiveCamera(50, WIDTH / HEIGHT, 0.1, 100);
camera.position.z = 4;

// --- Lighting (MeshStandardMaterial renders black without lights) ---
scene.add(new THREE.AmbientLight(0xffffff, 0.4));
const key = new THREE.DirectionalLight(0xffffff, 1.2);
key.position.set(3, 4, 5);
scene.add(key);

// --- Offscreen WebGL renderer (its canvas is the SOURCE for renderAscii) ---
const renderer = new THREE.WebGLRenderer({ antialias: true });
renderer.setSize(WIDTH, HEIGHT);
// renderer.domElement is NOT added to the DOM — it's our pixel source.

const reduceMotion = window.matchMedia("(prefers-reduced-motion: reduce)").matches;

function animate() {
  mesh.rotation.x += 0.01;
  mesh.rotation.y += 0.007;
  renderer.render(scene, camera);
  // renderAscii is defined in §7; it reads pixels and draws characters.
  renderAscii(asciiCtx, asciiCanvas, renderer.domElement, CELL_W, CELL_H);
  requestAnimationFrame(animate);
}

if (reduceMotion) {
  // Honor reduced motion: render a single static frame instead of looping.
  renderer.render(scene, camera);
  renderAscii(asciiCtx, asciiCanvas, renderer.domElement, CELL_W, CELL_H);
} else {
  animate();
}
```

### 9. Performance Optimization

| Technique | Impact | Implementation |
|-----------|--------|---------------|
| Skip black pixels | 30-50% fewer draw calls | `if (brightness < threshold) continue` |
| Throttle FPS | Reduce CPU usage | `requestAnimationFrame` with timestamp check |
| Reduce resolution | Fewer cells to render | Smaller offscreen canvas |
| Cache character metrics | Avoid repeated `measureText` | Pre-compute once |
| Use `willReadFrequently` | Faster `getImageData` | Pass to canvas context options |
| Gradient fade | Visual polish | CSS gradient overlay at edges |

### 10. Static ASCII Art Generation

**From image to ASCII (Python):**
```python
from PIL import Image

CHARS = " .:-=+*#%@"

def image_to_ascii(path, width=80):
    img = Image.open(path).convert("L")
    aspect = img.height / img.width
    height = int(width * aspect * 0.5)  # Terminal chars are ~2:1
    img = img.resize((width, height))

    ascii_art = ""
    for y in range(height):
        for x in range(width):
            brightness = img.getpixel((x, y)) / 255
            ascii_art += CHARS[int(brightness * (len(CHARS) - 1))]
        ascii_art += "\n"
    return ascii_art
```

**From text to ASCII banner** (uses `npx figlet`; see §6):
```bash
# Quick branded banner, indented two spaces
npx figlet -f Slant "skills.ws" | sed 's/^/  /'

# With green color (bash) — wrap in ANSI SGR codes
echo -e "\033[32m$(npx figlet -f Slant 'skills.ws')\033[0m"
```

## Checklist

- [ ] Terminal capability detection (TTY + UTF-8 locale) before rendering
- [ ] Print a static text alternative first; fall back to static art when animation disabled
- [ ] Disable animation when not a TTY, in CI, with `NO_COLOR`, or `TERM=dumb`; web uses `prefers-reduced-motion`
- [ ] Choose ASCII-safe vs Unicode charset by capability (don't assume UTF-8)
- [ ] Hide cursor during animation, restore after
- [ ] Use alternate screen buffer for full-screen animations
- [ ] Differential rendering to prevent flicker
- [ ] Test on: iTerm2, Terminal.app, Windows Terminal, Alacritty, VS Code terminal
- [ ] Cleanup on SIGINT (restore cursor, clear buffer)
- [ ] Keep animation under 3 seconds (respect user's time)
- [ ] Web: add gradient fade, throttle to 30fps max

---

## auth-implementation
Category: dev
Description: Secure authentication & authorization — OAuth 2.1/OIDC with PKCE & state, JWT/JWKS verification, hashed-rotating refresh tokens, sessions/BFF, passkeys/WebAuthn, MFA/TOTP, RBAC/ABAC, password hashing, and CSRF. Use when implementing or reviewing auth, authz, MFA, passkeys, OAuth/OIDC, sessions, tokens, or access control.
Features:
  - OAuth 2.0 flows: authorization code, PKCE, client credentials
  - JWT structure, signing, validation, and refresh token rotation
  - Session management: cookie-based, token-based, Redis sessions
  - NextAuth.js / Auth.js setup and provider configuration
  - Passport.js strategies for Express applications
  - Passkeys and WebAuthn implementation
  - RBAC and ABAC authorization patterns
  - Password hashing with bcrypt and argon2
  - MFA/2FA with TOTP (Google Authenticator)
  - CSRF protection, secure cookies, and rate limiting
Use Cases:
  - Implement OAuth 2.0 PKCE flow for a single-page application
  - Set up JWT auth with refresh token rotation
  - Add Google, GitHub, and Apple social login
  - Implement role-based access control for an API
  - Add passkey/WebAuthn authentication to a web app
  - Set up NextAuth.js with multiple providers
  - Implement TOTP-based two-factor authentication
  - Configure secure session management with Redis

# Authentication & Authorization

Security-critical patterns for AuthN/AuthZ in 2026. Code here is meant to be copied, so it is written to be correct and safe by default: every secret stored hashed, every token rotation atomic, every redirect-based flow CSRF-protected via `state`/PKCE. Vendor endpoints and library APIs drift — when a value here is dated, the inline note tells you where to re-verify.

**Threat-model defaults**: assume the browser is hostile (XSS can read anything JS can), assume tokens leak, assume requests are replayed and races happen. Prefer short-lived access tokens + server-held session/refresh state. For SPAs, prefer a **BFF (Backend-for-Frontend)** holding tokens server-side over putting access tokens in `localStorage`.

---

## Safety gate

Before executing commands or changing external systems, confirm scope, credentials, target environment, rollback, and required approval. Pin and verify third-party artifacts; never expose secrets to client code or logs.

## Reference guide

Read only the references needed for the current request:

- **1. OAuth 2.1 / OIDC Flows**: [references/1-oauth-2-1-oidc-flows.md](references/1-oauth-2-1-oidc-flows.md)
- **2. JWT (JSON Web Tokens)**: [references/2-jwt-json-web-tokens.md](references/2-jwt-json-web-tokens.md)
- **3. Session Management**: [references/3-session-management.md](references/3-session-management.md)
- **4. Auth.js v5 (NextAuth) Setup — App Router**: [references/4-auth-js-v5-nextauth-setup-app-router.md](references/4-auth-js-v5-nextauth-setup-app-router.md)
- **5. Passport.js Strategies**: [references/5-passport-js-strategies.md](references/5-passport-js-strategies.md)
- **6. Passkeys / WebAuthn (`@simplewebauthn/server` v13)**: [references/6-passkeys-webauthn-simplewebauthn-server-v13.md](references/6-passkeys-webauthn-simplewebauthn-server-v13.md)
- **7. RBAC & ABAC**: [references/7-rbac-abac.md](references/7-rbac-abac.md)
- **8. Password Hashing**: [references/8-password-hashing.md](references/8-password-hashing.md)
- **9. MFA / 2FA with TOTP**: [references/9-mfa-2fa-with-totp.md](references/9-mfa-2fa-with-totp.md)
- **10. Security Best Practices**: [references/10-security-best-practices.md](references/10-security-best-practices.md)

### Resource: references/1-oauth-2-1-oidc-flows.md

## Contents

- 1. OAuth 2.1 / OIDC Flows
- Discover endpoints (don't hardcode)
- Authorization Code + PKCE — full server-side flow with state
- Authorization Code + PKCE — public client (SPA/mobile) caveat
- Client Credentials Flow (Machine-to-Machine)

## 1. OAuth 2.1 / OIDC Flows

OAuth 2.1 (the consolidation of 2.0 + best-practice RFCs) makes **PKCE mandatory for all clients**, forbids the implicit and password grants, and requires exact redirect-URI matching. Use the **Authorization Code flow + PKCE everywhere** (yes, even confidential server-side clients).

### Discover endpoints (don't hardcode)

Prefer the provider's OIDC discovery document over hardcoded URLs so endpoints and the JWKS URI stay correct:

```javascript
// Fetch once at boot, cache in memory (respect Cache-Control)
const discovery = await fetch(
  'https://accounts.google.com/.well-known/openid-configuration'
).then(r => r.json());
// => { authorization_endpoint, token_endpoint, jwks_uri, issuer, ... }
// Google (as of Jun 2026): authorization_endpoint = https://accounts.google.com/o/oauth2/v2/auth
//                          token_endpoint         = https://oauth2.googleapis.com/token
// Verify: https://accounts.google.com/.well-known/openid-configuration
```

### Authorization Code + PKCE — full server-side flow with `state`

This is the canonical flow. **`state` and the PKCE `code_verifier` are both persisted server-side before redirect and verified on callback** — skipping either reopens CSRF / login-injection / code-injection. Below, secrets live on the server and only `code_challenge` + `state` ever hit the browser.

```javascript
import crypto from 'node:crypto';

const b64url = (buf) => buf.toString('base64url'); // Node >=16 supports 'base64url'

function pkcePair() {
  const verifier = b64url(crypto.randomBytes(32));               // 43-128 chars, high entropy
  const challenge = b64url(crypto.createHash('sha256').update(verifier).digest());
  return { verifier, challenge };
}

// --- Step 1: begin login (server route) ---
app.get('/auth/login', async (req, res) => {
  const state = b64url(crypto.randomBytes(32));
  const nonce = b64url(crypto.randomBytes(32));   // OIDC: binds id_token to this session
  const { verifier, challenge } = pkcePair();

  // PERSIST state + verifier + nonce server-side, keyed to THIS session, BEFORE redirecting.
  // Short TTL; single use. (Express-session shown; a signed httpOnly cookie also works.)
  req.session.oauth = { state, nonce, verifier, createdAt: Date.now() };

  const url = new URL(discovery.authorization_endpoint);
  url.searchParams.set('client_id', process.env.OAUTH_CLIENT_ID);
  url.searchParams.set('redirect_uri', process.env.OAUTH_REDIRECT_URI); // must EXACTLY match registered URI
  url.searchParams.set('response_type', 'code');
  url.searchParams.set('scope', 'openid email profile');
  url.searchParams.set('code_challenge', challenge);
  url.searchParams.set('code_challenge_method', 'S256');
  url.searchParams.set('state', state);
  url.searchParams.set('nonce', nonce);
  res.redirect(url.toString());
});

// --- Step 2: callback (server route) — VALIDATE everything ---
app.get('/auth/callback', async (req, res) => {
  const { code, state } = req.query;
  const saved = req.session.oauth;
  delete req.session.oauth; // consume immediately so it can't be replayed

  // (a) provider returned an error?
  if (req.query.error) return res.status(400).send(`OAuth error: ${req.query.error}`);
  // (b) we actually started a flow, and it hasn't expired
  if (!saved || Date.now() - saved.createdAt > 10 * 60 * 1000) {
    return res.status(400).send('No pending OAuth flow / expired');
  }
  // (c) STATE MUST MATCH — constant-time compare to avoid timing oracles
  const ok = typeof state === 'string'
    && state.length === saved.state.length
    && crypto.timingSafeEqual(Buffer.from(state), Buffer.from(saved.state));
  if (!ok) return res.status(403).send('Invalid OAuth state'); // CSRF / login-injection blocked here
  // (d) need a code
  if (typeof code !== 'string' || !code) return res.status(400).send('Missing authorization code');

  // (e) exchange code — token request is form-encoded; include the PKCE verifier we persisted
  const body = new URLSearchParams({
    grant_type: 'authorization_code',
    code,
    redirect_uri: process.env.OAUTH_REDIRECT_URI,
    client_id: process.env.OAUTH_CLIENT_ID,
    code_verifier: saved.verifier,            // proves we started this exact flow
  });
  // Confidential clients add their secret (Basic auth header preferred over body params):
  const headers = { 'Content-Type': 'application/x-www-form-urlencoded' };
  if (process.env.OAUTH_CLIENT_SECRET) {
    const basic = Buffer.from(
      `${process.env.OAUTH_CLIENT_ID}:${process.env.OAUTH_CLIENT_SECRET}`
    ).toString('base64');
    headers.Authorization = `Basic ${basic}`;
  }

  const tokenRes = await fetch(discovery.token_endpoint, { method: 'POST', headers, body });
  if (!tokenRes.ok) return res.status(401).send('Token exchange failed');
  const tokens = await tokenRes.json(); // { access_token, refresh_token, id_token, expires_in }

  // (f) Validate the OIDC id_token signature/iss/aud/exp AND that nonce matches saved.nonce
  //     (see §2 verifyToken; pass audience = OAUTH_CLIENT_ID and check decoded.nonce === saved.nonce)
  // (g) Establish your OWN session here (don't hand provider tokens to the browser).
  res.redirect('/');
});
```

### Authorization Code + PKCE — public client (SPA/mobile) caveat

A SPA cannot keep `state`/`verifier` truly secret from XSS. `sessionStorage` survives a redirect but is JS-readable. **Preferred 2026 pattern: run the code exchange in a BFF** so the browser never holds tokens. If you must do it browser-side, still generate and check `state`, and still send the `code_verifier`:

```javascript
// Browser: begin
const { verifier, challenge } = await pkcePairWebCrypto(); // Web Crypto version below
const state = crypto.randomUUID();
sessionStorage.setItem('pkce_verifier', verifier);
sessionStorage.setItem('oauth_state', state);
const url = new URL(discovery.authorization_endpoint);
url.searchParams.set('client_id', CLIENT_ID);
url.searchParams.set('redirect_uri', REDIRECT_URI);
url.searchParams.set('response_type', 'code');
url.searchParams.set('scope', 'openid email profile');
url.searchParams.set('code_challenge', challenge);
url.searchParams.set('code_challenge_method', 'S256');
url.searchParams.set('state', state);
location.href = url.toString();

// Browser: callback — VALIDATE state before exchanging
const params = new URLSearchParams(location.search);
const code = params.get('code');
const returnedState = params.get('state');
const savedState = sessionStorage.getItem('oauth_state');
const verifier = sessionStorage.getItem('pkce_verifier');
sessionStorage.removeItem('oauth_state');
sessionStorage.removeItem('pkce_verifier');
if (!code || !returnedState || returnedState !== savedState) {
  throw new Error('Invalid OAuth state or missing code'); // stop — do not exchange
}
const res = await fetch(discovery.token_endpoint, {
  method: 'POST',
  headers: { 'Content-Type': 'application/x-www-form-urlencoded' },
  body: new URLSearchParams({
    grant_type: 'authorization_code', code,
    client_id: CLIENT_ID, redirect_uri: REDIRECT_URI, code_verifier: verifier,
  }),
});
const tokens = await res.json();

// Web Crypto PKCE (browser)
async function pkcePairWebCrypto() {
  const bytes = crypto.getRandomValues(new Uint8Array(32));
  const verifier = btoa(String.fromCharCode(...bytes))
    .replace(/\+/g, '-').replace(/\//g, '_').replace(/=+$/, '');
  const digest = await crypto.subtle.digest('SHA-256', new TextEncoder().encode(verifier));
  const challenge = btoa(String.fromCharCode(...new Uint8Array(digest)))
    .replace(/\+/g, '-').replace(/\//g, '_').replace(/=+$/, '');
  return { verifier, challenge };
}
```

### Client Credentials Flow (Machine-to-Machine)

For backend services and API-to-API calls. No user. Token requests are **form-encoded** per RFC 6749, the secret stays server-side, and you should **cache the token** until shortly before `expires_in` rather than minting one per call.

```javascript
let cached = { token: null, exp: 0 };

async function getServiceToken() {
  if (cached.token && Date.now() < cached.exp - 60_000) return cached.token; // 60s safety margin
  const res = await fetch(process.env.OAUTH_TOKEN_URL, {
    method: 'POST',
    headers: { 'Content-Type': 'application/x-www-form-urlencoded' },
    body: new URLSearchParams({
      grant_type: 'client_credentials',
      client_id: process.env.CLIENT_ID,
      client_secret: process.env.CLIENT_SECRET, // server-only; never ship to a browser
      audience: 'https://api.example.com',       // provider-specific (Auth0 uses `audience`)
      scope: 'read:things write:things',
    }),
  });
  if (!res.ok) throw new Error(`Token request failed: ${res.status}`);
  const { access_token, expires_in } = await res.json();
  cached = { token: access_token, exp: Date.now() + expires_in * 1000 };
  return access_token;
}
```

---

### Resource: references/10-security-best-practices.md

## Contents

- 10. Security Best Practices
- CSRF Protection (do NOT use csurf)
- Secure Cookie Configuration
- Rate Limiting Login Attempts
- Social Login Setup (Google, GitHub, Apple)
- Secure logout / session invalidation

## 10. Security Best Practices

### CSRF Protection (do NOT use `csurf`)

`csurf` has been **deprecated/unmaintained since 2022** — don't add it to new code. Use SameSite cookies as a baseline plus an explicit token defense, and validate `Origin`/`Referer` on state-changing requests.

```javascript
// Option A (recommended for sessions): double-submit signed token via `csrf-csrf`
import { doubleCsrf } from 'csrf-csrf';

const { generateCsrfToken, doubleCsrfProtection } = doubleCsrf({
  getSecret: () => process.env.CSRF_SECRET,
  getSessionIdentifier: (req) => req.session?.id ?? '', // REQUIRED: binds the token to this session
  cookieName: '__Host-csrf',
  cookieOptions: { sameSite: 'strict', secure: true, path: '/', httpOnly: true },
  getCsrfTokenFromRequest: (req) => req.headers['x-csrf-token'],
});

app.get('/csrf-token', (req, res) => res.json({ csrfToken: generateCsrfToken(req, res) }));
app.use(doubleCsrfProtection); // rejects unsafe methods without a matching token

// Option B (defense in depth): reject state-changing requests from foreign origins
app.use((req, res, next) => {
  if (['POST', 'PUT', 'PATCH', 'DELETE'].includes(req.method)) {
    const origin = req.get('origin') || req.get('referer') || '';
    const allowed = ['https://app.example.com'];
    if (!allowed.some(a => origin.startsWith(a))) {
      return res.status(403).json({ error: 'Cross-origin request blocked' });
    }
  }
  next();
});
```

For SPAs/mobile using JWT in the `Authorization` header (not cookies), CSRF tokens aren't required — the browser won't auto-attach a header cross-site. **But** if you store any auth state in cookies (including a BFF session), you DO need CSRF defense. `SameSite=Lax/Strict` reduces risk but is not complete: it doesn't cover same-site subdomain attacks, and `Lax` still allows top-level cross-site GETs — so never perform state changes on GET.

### Secure Cookie Configuration

```javascript
res.cookie('__Host-session', token, {
  httpOnly: true,     // JS can't read it (XSS mitigation)
  secure: true,       // HTTPS only
  sameSite: 'strict', // 'lax' only if you need top-level cross-site navigation to stay logged in
  maxAge: 86_400_000, // 24h
  path: '/',
  // __Host- prefix => browser enforces Secure + path=/ + NO Domain (locks cookie to exact host)
});
```

### Rate Limiting Login Attempts

Throttle on **IP and account, normalized**, not just one. The default IP key generator mishandles IPv6 (a whole /64 shares an address) — use `express-rate-limit`'s `ipKeyGenerator` helper. Combine a coarse per-IP limit with a stricter per-account limit, and add account lockout/backoff for repeated failures.

```javascript
import rateLimit, { ipKeyGenerator } from 'express-rate-limit';

const loginLimiter = rateLimit({
  windowMs: 15 * 60 * 1000, // 15 minutes
  max: 10,                  // 10 attempts per key per window
  standardHeaders: 'draft-7',
  legacyHeaders: false,
  message: { error: 'Too many login attempts. Try again later.' },
  // Key by account + IP. Normalize email; ipKeyGenerator() handles IPv6 correctly.
  keyGenerator: (req) => {
    const email = String(req.body?.email ?? '').trim().toLowerCase();
    return `${email}|${ipKeyGenerator(req.ip)}`;
  },
});

app.post('/auth/login', loginLimiter, loginHandler);
```

For distributed deployments back the limiter with a shared store (e.g. `rate-limit-redis`) so limits are global, not per-instance. Track consecutive failures per account and apply exponential backoff or temporary lockout, with an audit log entry per failure.

### Social Login Setup (Google, GitHub, Apple)

**Required env vars per provider** (verify scopes/console layout at each link; layouts change):

| Provider | Vars | Console |
|----------|------|---------|
| Google | `GOOGLE_CLIENT_ID`, `GOOGLE_CLIENT_SECRET` | console.cloud.google.com → APIs & Services → Credentials |
| GitHub | `GITHUB_CLIENT_ID`, `GITHUB_CLIENT_SECRET` | github.com/settings/developers → OAuth Apps |
| Apple | `APPLE_CLIENT_ID`, `APPLE_TEAM_ID`, `APPLE_KEY_ID`, `APPLE_PRIVATE_KEY` | developer.apple.com → Certificates, IDs & Profiles |

(With Auth.js v5, prefer the auto-inferred `AUTH_GOOGLE_ID`/`AUTH_GOOGLE_SECRET` form.)

**Callback URLs:** register the exact callback URL(s); no wildcards in production. Keep all client secrets and Apple private keys server-side only — never expose them to the browser bundle.

### Secure logout / session invalidation
- Server sessions: destroy the session record (`req.session.destroy`) and clear the cookie. Don't rely on the client to "forget" the cookie.
- JWTs: keep them short-lived; on logout, delete/revoke the refresh-token family (§2) and, for high-value apps, add the access token's `jti` to a short-TTL blocklist until it expires.
- On password change or detected compromise, revoke **all** of the user's refresh-token families and active sessions.

### Resource: references/2-jwt-json-web-tokens.md

## Contents

- 2. JWT (JSON Web Tokens)
- Structure
- JWT Validation (Node.js) — pin algorithm, issuer, audience
- Refresh Token Rotation — hashed, atomic, reuse-detecting

## 2. JWT (JSON Web Tokens)

### Structure

```
header.payload.signature

Header:  { "alg": "RS256", "typ": "JWT", "kid": "key-id-1" }
Payload: { "iss": "https://auth.example.com", "aud": "https://api.example.com",
           "sub": "user123", "role": "admin", "iat": 1706000000, "exp": 1706003600 }
Signature: RS256(base64url(header) + "." + base64url(payload), privateKey)
```

### JWT Validation (Node.js) — pin algorithm, issuer, audience

The most common JWT bugs: accepting `alg: none`, accepting whatever `alg` the token claims (HS256/RS256 confusion), or not checking `iss`/`aud`. **Always pin `algorithms` explicitly, and always validate `issuer` and `audience`.** Cache JWKS keys (jwks-rsa caches + rate-limits by default).

```javascript
import jwt from 'jsonwebtoken';
import jwksClient from 'jwks-rsa';

const client = jwksClient({
  jwksUri: 'https://auth.example.com/.well-known/jwks.json',
  cache: true, cacheMaxEntries: 5, cacheMaxAge: 10 * 60 * 1000, // 10 min
  rateLimit: true, jwksRequestsPerMinute: 10,
});

function getKey(header, callback) {
  client.getSigningKey(header.kid, (err, key) => callback(err, key?.getPublicKey()));
}

function verifyToken(token) {
  return new Promise((resolve, reject) => {
    jwt.verify(token, getKey, {
      algorithms: ['RS256'],                  // pin; never allow 'none' or HS*/RS* mixing
      issuer: 'https://auth.example.com',     // must match token iss
      audience: 'https://api.example.com',    // must match token aud (your API/client id)
      clockTolerance: 5,                      // seconds, for minor clock skew
    }, (err, decoded) => (err ? reject(err) : resolve(decoded)));
  });
}

// Express middleware
async function authMiddleware(req, res, next) {
  const token = req.headers.authorization?.replace(/^Bearer /, '');
  if (!token) return res.status(401).json({ error: 'No token provided' });
  try {
    req.user = await verifyToken(token);
    next();
  } catch {
    return res.status(401).json({ error: 'Invalid token' });
  }
}
```

`jsonwebtoken` is the workhorse; for ESM/edge runtimes consider **`jose`** (`jwtVerify` + `createRemoteJWKSet`), which is promise-native and runs on Web Crypto.

### Refresh Token Rotation — hashed, atomic, reuse-detecting

Refresh tokens are long-lived bearer credentials, so treat them like passwords:

1. **Store only a hash** (SHA-256 is fine for a high-entropy random token; you don't need bcrypt/argon2 for 256-bit randomness). The raw token exists only in the client's secure cookie.
2. **Look up by hash**, never by raw value.
3. **Rotate atomically** in a DB transaction with a compare-and-set so two concurrent refreshes can't both succeed.
4. **Detect reuse**: a refresh token is single-use. If a *used/rotated* token is presented again, treat it as theft and **revoke the whole token family**.

```javascript
import crypto from 'node:crypto';

const sha256 = (s) => crypto.createHash('sha256').update(s).digest('hex');
const newOpaqueToken = () => crypto.randomBytes(32).toString('base64url'); // 256-bit, unguessable

// Issue at login: create a family, store HASH, return RAW token to client (httpOnly cookie)
async function issueRefreshToken(userId, meta) {
  const raw = newOpaqueToken();
  const familyId = crypto.randomUUID();
  await db.refreshToken.create({ data: {
    tokenHash: sha256(raw), userId, familyId,
    used: false, revoked: false,
    expiresAt: new Date(Date.now() + 30 * 24 * 60 * 60 * 1000), // 30 days
    userAgent: meta?.userAgent, ip: meta?.ip,                   // device/session metadata
  }});
  return raw; // caller sets it as a Secure; HttpOnly; SameSite cookie (path=/auth/refresh)
}

app.post('/auth/refresh', async (req, res) => {
  const raw = req.cookies?.refresh_token;            // delivered via httpOnly cookie, not body
  if (!raw) return res.status(401).json({ error: 'No refresh token' });
  const tokenHash = sha256(raw);

  try {
    const result = await db.$transaction(async (tx) => {
      // Lock the row (Postgres) so concurrent refreshes serialize on it.
      const [stored] = await tx.$queryRaw`
        SELECT * FROM "RefreshToken" WHERE "tokenHash" = ${tokenHash} FOR UPDATE`;

      if (!stored) throw { code: 'INVALID' };

      // REUSE DETECTION: a used or revoked token presented again => credential theft.
      if (stored.used || stored.revoked) {
        await tx.refreshToken.updateMany({
          where: { familyId: stored.familyId },
          data: { revoked: true },                  // nuke the entire family
        });
        throw { code: 'REUSE' };
      }
      if (stored.expiresAt < new Date()) throw { code: 'EXPIRED' };

      // Atomic compare-and-set: only the first concurrent caller flips used:false -> true.
      const claim = await tx.refreshToken.updateMany({
        where: { id: stored.id, used: false },
        data: { used: true },
      });
      if (claim.count !== 1) throw { code: 'RACE' };  // someone else won; reject this one

      // Mint the next token in the SAME family and store its hash.
      const nextRaw = newOpaqueToken();
      await tx.refreshToken.create({ data: {
        tokenHash: sha256(nextRaw), userId: stored.userId, familyId: stored.familyId,
        used: false, revoked: false,
        expiresAt: new Date(Date.now() + 30 * 24 * 60 * 60 * 1000),
        userAgent: req.get('user-agent'), ip: req.ip,
      }});

      const user = await tx.user.findUnique({ where: { id: stored.userId } });
      const accessToken = jwt.sign(
        { sub: stored.userId, role: user.role },
        process.env.JWT_PRIVATE_KEY,
        { algorithm: 'RS256', issuer: 'https://auth.example.com',
          audience: 'https://api.example.com', expiresIn: '15m' }
      );
      return { accessToken, nextRaw };
    });

    res.cookie('refresh_token', result.nextRaw, {
      httpOnly: true, secure: true, sameSite: 'strict', path: '/auth/refresh',
      maxAge: 30 * 24 * 60 * 60 * 1000,
    });
    res.json({ accessToken: result.accessToken });
  } catch (e) {
    if (e?.code === 'REUSE') {
      // Optional: log/audit + force user re-auth on all devices.
      return res.status(401).json({ error: 'Token reuse detected; family revoked' });
    }
    return res.status(401).json({ error: 'Invalid refresh token' });
  }
});
```

> The `FOR UPDATE` row lock + `updateMany(where: { used: false })` compare-and-set is what makes this safe under concurrency. On databases without `SELECT ... FOR UPDATE`, rely solely on the conditional update's affected-row count (`claim.count === 1`) as the gate — never on a read-then-write without it.

**Token lifetimes:**
- Access token: 15 minutes (short-lived, stateless, RS256)
- Refresh token: 30 days max, rotated on every use, hashed at rest
- ID token: ~1 hour (OIDC user info; validate `nonce` on login)

---

### Resource: references/3-session-management.md

## Contents

- 3. Session Management
- Cookie-Based Sessions (Traditional / BFF)
- Cookie vs Token Comparison

## 3. Session Management

### Cookie-Based Sessions (Traditional / BFF)

```javascript
import session from 'express-session';
import { RedisStore } from 'connect-redis';
import { createClient } from 'redis';

const redisClient = createClient({ url: process.env.REDIS_URL });
await redisClient.connect();

app.set('trust proxy', 1); // required behind a TLS-terminating proxy so secure cookies work
app.use(session({
  store: new RedisStore({ client: redisClient }),
  secret: process.env.SESSION_SECRET,   // rotate via an array: [newSecret, oldSecret]
  resave: false,
  saveUninitialized: false,
  name: '__Host-session',               // __Host- prefix: requires Secure, path=/, no Domain
  cookie: {
    secure: true,       // HTTPS only
    httpOnly: true,     // no JS access (XSS can't read it)
    sameSite: 'lax',    // mitigates cross-site POST CSRF (not a complete defense — see §10)
    maxAge: 24 * 60 * 60 * 1000,
    path: '/',
    // NOTE: __Host- forbids `domain`. Drop the prefix if you need a shared parent domain.
  },
}));

// Regenerate the session ID on privilege change to prevent session fixation:
app.post('/auth/login', loginLimiter, async (req, res) => {
  // ... verify credentials ...
  req.session.regenerate((err) => {
    if (err) return res.status(500).end();
    req.session.userId = user.id;
    res.json({ ok: true });
  });
});
```

### Cookie vs Token Comparison

| Aspect | Cookie Sessions | JWT Access Tokens |
|--------|----------------|-------------------|
| Storage | Server (Redis/DB) | Client — see XSS row |
| Stateless | No (server lookup) | Yes (self-contained) |
| Revocation | Easy (delete from store) | Hard (need blocklist or short TTL) |
| Scalability | Need shared store | No shared state needed |
| XSS exposure | `httpOnly` cookie is unreadable by JS | **Avoid `localStorage`** (JS-readable → XSS steals it). Keep in memory, or use a BFF that holds the token server-side. |
| CSRF exposure | Needs CSRF defense (§10) | Safe **only** if sent in `Authorization` header, not a cookie |
| Mobile | Needs cookie support | Works everywhere |
| Best for | Server-rendered apps, **BFF for SPAs** | Native/mobile, service-to-service |

**2026 browser recommendation:** do not store access tokens in `localStorage`/`sessionStorage`. Either (a) use a **BFF** where the browser holds only an `httpOnly` session cookie and the server attaches tokens to upstream calls, or (b) keep the access token in a JS variable in memory and refresh via an `httpOnly` cookie hitting `/auth/refresh`.

---

### Resource: references/4-auth-js-v5-nextauth-setup-app-router.md

## 4. Auth.js v5 (NextAuth) Setup — App Router

Auth.js v5 splits config into a root `auth.ts` (exports `handlers`, `auth`, `signIn`, `signOut`) and a thin route handler. Provider imports are now `next-auth/providers/*` named like `Google`, `GitHub`, `Credentials` (the old `GoogleProvider` names are v4). Env vars are auto-inferred from `AUTH_<PROVIDER>_ID/SECRET`; set `AUTH_SECRET` (generate with `npx auth secret`) and `AUTH_TRUST_HOST=true` when self-hosting behind a proxy.

```typescript
// auth.ts (project root)
import NextAuth from 'next-auth';
import Google from 'next-auth/providers/google';
import GitHub from 'next-auth/providers/github';
import Credentials from 'next-auth/providers/credentials';
import { PrismaAdapter } from '@auth/prisma-adapter';
import { prisma } from '@/lib/prisma';
import argon2 from 'argon2';
import { z } from 'zod';

const signInSchema = z.object({
  email: z.string().email(),
  password: z.string().min(8),
});

export const { handlers, auth, signIn, signOut } = NextAuth({
  adapter: PrismaAdapter(prisma),
  session: { strategy: 'jwt' },           // 'jwt' for Credentials; 'database' for OAuth-only is also fine
  providers: [
    Google,                               // reads AUTH_GOOGLE_ID / AUTH_GOOGLE_SECRET
    GitHub,                               // reads AUTH_GITHUB_ID / AUTH_GITHUB_SECRET
    Credentials({
      credentials: { email: {}, password: {} },
      authorize: async (credentials) => {
        const parsed = signInSchema.safeParse(credentials);
        if (!parsed.success) return null;
        const user = await prisma.user.findUnique({ where: { email: parsed.data.email } });
        // Same generic outcome whether the user is missing or the password is wrong (no enumeration):
        if (!user?.hashedPassword) return null;
        const ok = await argon2.verify(user.hashedPassword, parsed.data.password);
        return ok ? user : null;
      },
    }),
  ],
  callbacks: {
    async jwt({ token, user }) {
      if (user) { token.role = user.role; token.id = user.id; }
      return token;
    },
    async session({ session, token }) {
      session.user.role = token.role as string;
      session.user.id = token.id as string;
      return session;
    },
  },
  pages: { signIn: '/login', error: '/auth/error' },
});
```

```typescript
// app/api/auth/[...nextauth]/route.ts
import { handlers } from '@/auth';
export const { GET, POST } = handlers;
```

```typescript
// types: augment the session/JWT so token.role/id type-check
// types/next-auth.d.ts
import 'next-auth';
declare module 'next-auth' {
  interface User { role?: string }
  interface Session { user: { id: string; role?: string } & DefaultSession['user'] }
}
```

Read the session in a Server Component/route via `const session = await auth();`. Protect routes in `middleware.ts` by exporting `auth` as the middleware.

---

### Resource: references/5-passport-js-strategies.md

## 5. Passport.js Strategies

`Invalid email` vs `Invalid password` are different responses, which lets an attacker enumerate accounts. **Return one generic message** and run the password compare even when the user is missing (constant-ish work) so timing doesn't leak existence either.

```javascript
import passport from 'passport';
import { Strategy as GoogleStrategy } from 'passport-google-oauth20';
import { Strategy as LocalStrategy } from 'passport-local';
import argon2 from 'argon2';

// A precomputed dummy argon2 hash so we do equivalent work when the user doesn't exist.
const DUMMY_HASH = process.env.DUMMY_ARGON2_HASH; // generate once: argon2.hash('not-a-real-password')

passport.use(new LocalStrategy(
  { usernameField: 'email' },
  async (email, password, done) => {
    const user = await db.findUserByEmail(String(email).trim().toLowerCase());
    const hash = user?.hashedPassword ?? DUMMY_HASH;
    const valid = await argon2.verify(hash, password).catch(() => false);
    if (!user || !valid) {
      return done(null, false, { message: 'Invalid credentials' }); // generic; no enumeration
    }
    return done(null, user);
  }
));

passport.use(new GoogleStrategy({
  clientID: process.env.GOOGLE_CLIENT_ID,
  clientSecret: process.env.GOOGLE_CLIENT_SECRET,
  callbackURL: '/auth/google/callback',
}, async (accessToken, refreshToken, profile, done) => {
  const email = profile.emails?.[0]?.value;
  let user = await db.findUserByGoogleId(profile.id);
  if (!user) {
    user = await db.createUser({ googleId: profile.id, email, name: profile.displayName });
  }
  return done(null, user);
}));

passport.serializeUser((user, done) => done(null, user.id));
passport.deserializeUser(async (id, done) => done(null, await db.findUserById(id)));
```

Also add **account lockout / progressive delay** after repeated failures (see §10 rate limiting) and **audit-log** auth events (login success/failure, password change, MFA enrollment) with user id, IP, and timestamp.

---

### Resource: references/6-passkeys-webauthn-simplewebauthn-server-v13.md

## 6. Passkeys / WebAuthn (`@simplewebauthn/server` v13)

The v13 API differs from older snippets you'll find online:
- `generateRegistrationOptions` takes **`userName`/`userDisplayName`** — there is no `userID` option you pass a string to anymore.
- After registration, persist **`registrationInfo.credential`** = `{ id, publicKey, counter, transports }` (plus `credentialDeviceType`/`credentialBackedUp`). `id` is a Base64URL string; `publicKey` is a `Uint8Array` (store as bytes/`bytea`).
- `verifyAuthenticationResponse` takes a single **`credential`** object (the old `authenticator:` key is gone) and a **`requireUserVerification`** boolean.
- Always store and re-send `transports`, and **write back `newCounter`** to detect cloned authenticators.

```javascript
import {
  generateRegistrationOptions, verifyRegistrationResponse,
  generateAuthenticationOptions, verifyAuthenticationResponse,
} from '@simplewebauthn/server';

const rpName = 'My App';
const rpID = 'example.com';                 // domain only, no scheme/port
const origin = 'https://example.com';       // full origin the browser sends

// --- Registration: options ---
app.post('/auth/passkey/register/options', async (req, res) => {
  const user = req.user;
  const existing = await db.getCredentialsByUserId(user.id);
  const options = await generateRegistrationOptions({
    rpName, rpID,
    userName: user.email,
    userDisplayName: user.name ?? user.email,
    attestationType: 'none',
    excludeCredentials: existing.map(c => ({ id: c.credentialId, transports: c.transports })),
    authenticatorSelection: {
      residentKey: 'preferred',       // 'required' for usernameless/discoverable login
      userVerification: 'preferred',  // 'required' to force biometric/PIN
    },
    supportedAlgorithmIDs: [-7, -257], // ES256, RS256
  });
  await db.saveChallenge(user.id, options.challenge); // store server-side, short TTL
  res.json(options);
});

// --- Registration: verify ---
app.post('/auth/passkey/register/verify', async (req, res) => {
  const user = req.user;
  const expectedChallenge = await db.getChallenge(user.id);
  let verification;
  try {
    verification = await verifyRegistrationResponse({
      response: req.body,
      expectedChallenge,
      expectedOrigin: origin,
      expectedRPID: rpID,
      requireUserVerification: true,
    });
  } catch (err) {
    return res.status(400).json({ error: err.message });
  }
  if (verification.verified && verification.registrationInfo) {
    const { credential, credentialDeviceType, credentialBackedUp } = verification.registrationInfo;
    await db.saveCredential(user.id, {
      credentialId: credential.id,        // Base64URLString
      publicKey: credential.publicKey,    // Uint8Array -> store as bytes
      counter: credential.counter,
      transports: credential.transports,  // e.g. ['internal','hybrid']
      deviceType: credentialDeviceType,
      backedUp: credentialBackedUp,
    });
  }
  await db.clearChallenge(user.id);
  res.json({ verified: verification.verified });
});

// --- Authentication: options ---
app.post('/auth/passkey/login/options', async (req, res) => {
  // Optionally scope to a known user's credentials; omit allowCredentials for usernameless flow.
  const options = await generateAuthenticationOptions({
    rpID,
    userVerification: 'preferred',
    // allowCredentials: creds.map(c => ({ id: c.credentialId, transports: c.transports })),
  });
  await db.saveSessionChallenge(req.sessionID, options.challenge);
  res.json(options);
});

// --- Authentication: verify ---
app.post('/auth/passkey/login/verify', async (req, res) => {
  const expectedChallenge = await db.getSessionChallenge(req.sessionID);
  const stored = await db.getCredentialById(req.body.id); // req.body.id is Base64URLString
  if (!stored) return res.status(400).json({ error: 'Unknown credential' });

  let verification;
  try {
    verification = await verifyAuthenticationResponse({
      response: req.body,
      expectedChallenge,
      expectedOrigin: origin,
      expectedRPID: rpID,
      credential: {                       // v13 shape (was `authenticator`)
        id: stored.credentialId,
        publicKey: stored.publicKey,      // Uint8Array
        counter: stored.counter,
        transports: stored.transports,
      },
      requireUserVerification: true,
    });
  } catch (err) {
    return res.status(400).json({ error: err.message });
  }
  if (verification.verified) {
    // CRITICAL: persist newCounter to detect cloned authenticators / replay.
    await db.updateCounter(stored.id, verification.authenticationInfo.newCounter);
    await db.clearSessionChallenge(req.sessionID);
    req.login(stored.user, () => res.json({ verified: true }));
  } else {
    res.status(401).json({ verified: false });
  }
});
```

Pair with `@simplewebauthn/browser` (`startRegistration`/`startAuthentication`) on the client. Verify the current API at https://simplewebauthn.dev (this skill targets v13).

---

### Resource: references/7-rbac-abac.md

## Contents

- 7. RBAC & ABAC
- Role-Based Access Control (RBAC)
- Attribute-Based Access Control (ABAC)

## 7. RBAC & ABAC

### Role-Based Access Control (RBAC)

```typescript
const PERMISSIONS = {
  admin: ['read', 'write', 'delete', 'manage_users', 'manage_billing'],
  editor: ['read', 'write'],
  viewer: ['read'],
} as const;

type Role = keyof typeof PERMISSIONS;
type Permission = (typeof PERMISSIONS)[Role][number];

function requirePermission(permission: Permission) {
  return (req, res, next) => {
    const userRole = req.user?.role as Role | undefined;
    const perms = (userRole && PERMISSIONS[userRole]) || [];
    if (!perms.includes(permission)) return res.status(403).json({ error: 'Forbidden' });
    next();
  };
}

app.delete('/api/posts/:id', requirePermission('delete'), deletePost);
app.get('/api/posts', requirePermission('read'), listPosts);
```

### Attribute-Based Access Control (ABAC)

```typescript
interface PolicyContext {
  user: { id: string; role: string; department: string };
  resource: { ownerId: string; type: string; status: string; department?: string };
  action: string;
}

function evaluatePolicy(ctx: PolicyContext): boolean {
  if (ctx.user.role === 'admin') return true;
  // Owners can edit their own resources
  if (ctx.action === 'edit' && ctx.resource.ownerId === ctx.user.id) return true;
  // Editors can edit published resources in their own department
  if (ctx.action === 'edit' && ctx.user.role === 'editor'
      && ctx.resource.status === 'published'
      && ctx.resource.department === ctx.user.department) return true;
  return false; // default-deny
}
```

**Default-deny** is the rule: if no policy explicitly allows the action, reject. Always enforce authorization **server-side per request** — never trust a client-sent role/permission, and re-check ownership on every mutating route (most IDOR bugs are a missing per-object check).

---

### Resource: references/8-password-hashing.md

## 8. Password Hashing

Use a memory-hard algorithm. **`argon2id` is the OWASP first choice**; `bcrypt` is an acceptable, widely-supported fallback (note bcrypt silently truncates inputs beyond 72 bytes — pre-hash with SHA-256 if you must allow long passphrases). Never roll your own.

```javascript
import argon2 from 'argon2';
import bcrypt from 'bcryptjs';

// argon2id (recommended) — OWASP 2026 baseline params
const hash = await argon2.hash(password, {
  type: argon2.argon2id,
  memoryCost: 19456,   // 19 MiB (OWASP min); raise to 46–64 MiB on capable servers
  timeCost: 2,         // iterations
  parallelism: 1,
});
const valid = await argon2.verify(hash, password);

// bcrypt fallback (cost >=12 in 2026)
const bhash = await bcrypt.hash(password, 12);
const bvalid = await bcrypt.compare(password, bhash);
```

**Never:** MD5, SHA-1, plain SHA-256/512 (fast → brute-forceable), or any unsalted/un-stretched hash for passwords. Check new passwords against a breached-password list (e.g., HIBP k-anonymity range API) and enforce a sane minimum length over arbitrary complexity rules (NIST SP 800-63B).

---

### Resource: references/9-mfa-2fa-with-totp.md

## 9. MFA / 2FA with TOTP

Three correctness rules this implements:
1. **Never return the TOTP secret to the client as a "backup code."** The secret is the authenticator seed; if it's also a "backup code" then anyone who saw setup can mint valid TOTP codes forever. Backup codes are **separate, random, single-use, stored hashed**.
2. **Verify enrollment before activating**, and tolerate one time-step of clock drift (`epochTolerance: 30`, one 30-second step; otplib v12 called this `window: 1`).
3. **Bind the login MFA step to a pending first-factor session** — never trust a `userId` from the request body, or anyone can complete MFA "as" any user id.

```javascript
import { generateSecret, generateURI, verify } from 'otplib';
import QRCode from 'qrcode';
import crypto from 'node:crypto';

const sha256 = (s) => crypto.createHash('sha256').update(s).digest('hex');

// --- Setup: generate secret + QR. Do NOT return the secret as a backup code. ---
app.post('/auth/mfa/setup', requireAuth, async (req, res) => {
  const secret = generateSecret();
  const otpauth = generateURI({ issuer: 'MyApp', label: req.user.email, secret });
  const qrCode = await QRCode.toDataURL(otpauth);
  await db.saveTempMfaSecret(req.user.id, secret); // pending; not yet active
  // Return ONLY the QR/otpauth so the user can scan it. The raw secret is shown once for
  // manual entry only if you choose to; it is NOT a backup code.
  res.json({ qrCode, otpauth });
});

// --- Verify enrollment + issue SEPARATE single-use backup codes (stored hashed) ---
app.post('/auth/mfa/verify', requireAuth, async (req, res) => {
  const { token } = req.body;
  const secret = await db.getTempMfaSecret(req.user.id);
  if (!secret || !(await verify({ token, secret, epochTolerance: 30 })).valid) {
    return res.status(400).json({ error: 'Invalid code' });
  }
  await db.activateMfa(req.user.id, secret);

  // Backup codes: random, single-use, displayed ONCE, stored only as hashes.
  const plain = Array.from({ length: 10 }, () => crypto.randomBytes(5).toString('hex')); // 10×10-hex
  await db.replaceBackupCodes(req.user.id, plain.map(c => sha256(c))); // store HASHES only
  res.json({ success: true, backupCodes: plain }); // show once; never retrievable again
});

// --- Login MFA challenge: bind to a pending first-factor session, NOT a body userId ---
app.post('/auth/mfa/challenge', async (req, res) => {
  const { token } = req.body;
  // mfaPending was set by the password/first-factor step after it verified credentials.
  const userId = req.session?.mfaPending?.userId;
  if (!userId) return res.status(401).json({ error: 'No pending login' });

  const user = await db.findUserById(userId);
  let ok = (await verify({ token, secret: user.mfaSecret, epochTolerance: 30 })).valid;
  if (!ok) {
    // Backup code path: look up by HASH and consume single-use.
    ok = await db.consumeBackupCode(userId, sha256(String(token)));
  }
  if (!ok) return res.status(401).json({ error: 'Invalid MFA code' });

  // First + second factor both satisfied → now establish the authenticated session.
  delete req.session.mfaPending;
  req.session.regenerate((err) => {
    if (err) return res.status(500).end();
    req.session.userId = userId;
    res.json({ accessToken: issueAccessToken(user) });
  });
});

// The first-factor handler that sets mfaPending (sketch):
// after verifying email+password, if user.mfaEnabled:
//   req.session.mfaPending = { userId: user.id, at: Date.now() };  // short TTL
//   return res.json({ mfaRequired: true });
```

This targets otplib v13 (async, returns `{ valid }`); code written for the v12 `authenticator` API should pin `otplib@^12`.

Rate-limit `/auth/mfa/challenge` per pending session (e.g., 5 attempts) and expire `mfaPending` after a few minutes. For recovery, require re-verification (email link + cooldown) before disabling MFA, and re-issue fresh backup codes when the user regenerates them.

---

---

## aws-production-deploy
Category: operations
Description: Production AWS infra-as-code in Terraform & CDK: 3-tier VPC, ECS Fargate, Aurora, CloudFront/S3/WAF, OIDC CI/CD, monitoring, security hardening. Use when deploying a web app to AWS for production, writing/reviewing Terraform or CDK, setting up GitHub Actions OIDC deploys, or hardening an AWS account (remote state, GuardDuty, KMS, IAM).
Features:
  - ECS Fargate service deployment with auto-scaling
  - RDS PostgreSQL with read replicas and automated backups
  - CloudFront CDN with custom domain and SSL
  - Route53 DNS with health checks and failover
  - CloudWatch alarms, dashboards, and log aggregation
  - Infrastructure as Code with CDK and Terraform examples
Use Cases:
  - Deploy a production Next.js app on AWS
  - Set up a highly available database layer
  - Configure CDN with cache invalidation
  - Build monitoring dashboards for production services

# AWS Production Deploy

Production-grade AWS infrastructure patterns. Not hello-world — real modules you'd ship to production with VPC isolation, ECS Fargate, RDS, CloudFront, and full CI/CD.

## Safety gate

Before executing commands or changing external systems, confirm scope, credentials, target environment, rollback, and required approval. Pin and verify third-party artifacts; never expose secrets to client code or logs.

## Reference guide

Read only the references needed for the current request:

- **Architecture Overview**: [references/architecture-overview.md](references/architecture-overview.md)
- **1. VPC with Proper Network Isolation — Terraform**: [references/1-vpc-with-proper-network-isolation-terraform.md](references/1-vpc-with-proper-network-isolation-terraform.md)
- **2. ECS Fargate with Auto-Scaling**: [references/2-ecs-fargate-with-auto-scaling.md](references/2-ecs-fargate-with-auto-scaling.md)
- **3. RDS Aurora with Read Replicas**: [references/3-rds-aurora-with-read-replicas.md](references/3-rds-aurora-with-read-replicas.md)
- **4. CloudFront + S3 + WAF**: [references/4-cloudfront-s3-waf.md](references/4-cloudfront-s3-waf.md)
- **5. CI/CD — GitHub Actions to ECS**: [references/5-ci-cd-github-actions-to-ecs.md](references/5-ci-cd-github-actions-to-ecs.md)
- **6. Monitoring & Cost Alerts**: [references/6-monitoring-cost-alerts.md](references/6-monitoring-cost-alerts.md)
- **7. Database Migration Strategy**: [references/7-database-migration-strategy.md](references/7-database-migration-strategy.md)
- **8. CDK Alternative**: [references/8-cdk-alternative.md](references/8-cdk-alternative.md)
- **9. Cost Optimization**: [references/9-cost-optimization.md](references/9-cost-optimization.md)
- **10. Debugging ECS in Production**: [references/10-debugging-ecs-in-production.md](references/10-debugging-ecs-in-production.md)
- **11. Terraform Remote State (do this first)**: [references/11-terraform-remote-state-do-this-first.md](references/11-terraform-remote-state-do-this-first.md)
- **12. Production Guardrails (don't skip these)**: [references/12-production-guardrails-don-t-skip-these.md](references/12-production-guardrails-don-t-skip-these.md)

### Resource: references/1-vpc-with-proper-network-isolation-terraform.md

## 1. VPC with Proper Network Isolation — Terraform

Most tutorials give you a flat VPC. Production needs three tiers: public (ALB only), private (compute), isolated (database). NAT Gateway per AZ for HA.

```hcl
# modules/vpc/main.tf

variable "project" { type = string }
variable "environment" { type = string }
variable "vpc_cidr" { default = "10.0.0.0/16" }
variable "az_count" { default = 3 }

data "aws_availability_zones" "available" {
  state = "available"
}

locals {
  azs = slice(data.aws_availability_zones.available.names, 0, var.az_count)
  public_cidrs   = [for i in range(var.az_count) : cidrsubnet(var.vpc_cidr, 4, i)]
  private_cidrs  = [for i in range(var.az_count) : cidrsubnet(var.vpc_cidr, 4, i + var.az_count)]
  isolated_cidrs = [for i in range(var.az_count) : cidrsubnet(var.vpc_cidr, 4, i + var.az_count * 2)]
}

resource "aws_vpc" "main" {
  cidr_block           = var.vpc_cidr
  enable_dns_hostnames = true
  enable_dns_support   = true
  tags = { Name = "${var.project}-${var.environment}", Environment = var.environment }
}

# VPC Flow Logs — mandatory for debugging and compliance
resource "aws_flow_log" "main" {
  vpc_id               = aws_vpc.main.id
  traffic_type         = "ALL"
  log_destination_type = "cloud-watch-logs"
  log_destination      = aws_cloudwatch_log_group.flow_logs.arn
  iam_role_arn         = aws_iam_role.flow_logs.arn
}

resource "aws_cloudwatch_log_group" "flow_logs" {
  name              = "/vpc/flow-logs/${var.project}-${var.environment}"
  retention_in_days = 30
}

resource "aws_iam_role" "flow_logs" {
  name = "${var.project}-${var.environment}-flow-logs"
  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Action = "sts:AssumeRole", Effect = "Allow"
      Principal = { Service = "vpc-flow-logs.amazonaws.com" }
    }]
  })
}

resource "aws_iam_role_policy" "flow_logs" {
  role = aws_iam_role.flow_logs.id
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect = "Allow"
      Action = ["logs:CreateLogGroup","logs:CreateLogStream","logs:PutLogEvents","logs:DescribeLogGroups","logs:DescribeLogStreams"]
      Resource = "*"
    }]
  })
}

# Public subnets — ALB lives here
resource "aws_subnet" "public" {
  count                   = var.az_count
  vpc_id                  = aws_vpc.main.id
  cidr_block              = local.public_cidrs[count.index]
  availability_zone       = local.azs[count.index]
  map_public_ip_on_launch = true
  tags = { Name = "${var.project}-${var.environment}-public-${local.azs[count.index]}" }
}

# Private subnets — ECS tasks, NAT for outbound
resource "aws_subnet" "private" {
  count             = var.az_count
  vpc_id            = aws_vpc.main.id
  cidr_block        = local.private_cidrs[count.index]
  availability_zone = local.azs[count.index]
  tags = { Name = "${var.project}-${var.environment}-private-${local.azs[count.index]}" }
}

# Isolated subnets — RDS, ElastiCache. NO internet access.
resource "aws_subnet" "isolated" {
  count             = var.az_count
  vpc_id            = aws_vpc.main.id
  cidr_block        = local.isolated_cidrs[count.index]
  availability_zone = local.azs[count.index]
  tags = { Name = "${var.project}-${var.environment}-isolated-${local.azs[count.index]}" }
}

resource "aws_internet_gateway" "main" {
  vpc_id = aws_vpc.main.id
}

# One NAT per AZ for production HA (cross-AZ NAT is a single point of failure
# AND incurs cross-AZ data charges). Single NAT for dev cuts the per-NAT hourly
# fee — roughly one gateway's hourly + data cost; verify current NAT Gateway
# pricing for your region at https://aws.amazon.com/vpc/pricing/.
resource "aws_eip" "nat" {
  count  = var.environment == "production" ? var.az_count : 1
  domain = "vpc"
}

resource "aws_nat_gateway" "main" {
  count         = var.environment == "production" ? var.az_count : 1
  allocation_id = aws_eip.nat[count.index].id
  subnet_id     = aws_subnet.public[count.index].id
}

resource "aws_route_table" "public" {
  vpc_id = aws_vpc.main.id
  route {
    cidr_block = "0.0.0.0/0"
    gateway_id = aws_internet_gateway.main.id
  }
}

resource "aws_route_table_association" "public" {
  count          = var.az_count
  subnet_id      = aws_subnet.public[count.index].id
  route_table_id = aws_route_table.public.id
}

resource "aws_route_table" "private" {
  count  = var.environment == "production" ? var.az_count : 1
  vpc_id = aws_vpc.main.id
  route {
    cidr_block     = "0.0.0.0/0"
    nat_gateway_id = aws_nat_gateway.main[count.index].id
  }
}

resource "aws_route_table_association" "private" {
  count          = var.az_count
  subnet_id      = aws_subnet.private[count.index].id
  route_table_id = aws_route_table.private[var.environment == "production" ? count.index : 0].id
}

# Isolated — no internet route at all
resource "aws_route_table" "isolated" {
  vpc_id = aws_vpc.main.id
}

resource "aws_route_table_association" "isolated" {
  count          = var.az_count
  subnet_id      = aws_subnet.isolated[count.index].id
  route_table_id = aws_route_table.isolated.id
}

output "vpc_id" { value = aws_vpc.main.id }
output "public_subnet_ids" { value = aws_subnet.public[*].id }
output "private_subnet_ids" { value = aws_subnet.private[*].id }
output "isolated_subnet_ids" { value = aws_subnet.isolated[*].id }
```

---

### Resource: references/10-debugging-ecs-in-production.md

## 10. Debugging ECS in Production

```bash
# Open an interactive shell via ECS Exec (Session Manager, NOT SSH).
# Requires: enable_execute_command on the service, the four ssmmessages:* perms
# on the TASK role (section 2), and a shell in the image. Distroless/no-shell
# images have no /bin/sh — bake in a debug shell or use an ephemeral sidecar.
aws ecs execute-command --cluster myapp-prod --task TASK_ID \
  --container app --interactive --command /bin/sh

# Verify Exec is actually enabled on a running task (look for enableExecuteCommand):
aws ecs describe-tasks --cluster myapp-prod --tasks TASK_ARN \
  --query 'tasks[0].enableExecuteCommand'

# Tail logs
aws logs tail /ecs/myapp-production/app --since 30m --follow

# Check why tasks are failing
aws ecs describe-tasks --cluster myapp-prod --tasks TASK_ARN \
  --query 'tasks[0].stoppedReason'

# Force redeploy (ECS rolling controller only — a CODE_DEPLOY service rejects
# this; trigger a CodeDeploy deployment instead, see section 5 variant A).
aws ecs update-service --cluster myapp-prod --service myapp-prod --force-new-deployment
```

---

### Resource: references/11-terraform-remote-state-do-this-first.md

## 11. Terraform Remote State (do this first)

Local state is unacceptable for a team or for production. With S3-native state locking (Terraform 1.10+/1.11+) you no longer need a DynamoDB lock table — set `use_lockfile = true`. The state bucket must be encrypted and versioned.

```hcl
# backend.tf — bootstrap the bucket ONCE with local state, then migrate.
terraform {
  backend "s3" {
    bucket       = "myorg-tfstate-prod"
    key          = "app/production/terraform.tfstate"
    region       = "us-east-1"
    encrypt      = true
    use_lockfile = true            # native S3 lock; Terraform >= 1.10
    kms_key_id   = "alias/tfstate" # CMK, not the default aws/s3 key
  }
}

# state bucket resources (apply with a temporary local backend first)
resource "aws_s3_bucket" "tfstate" { bucket = "myorg-tfstate-prod" }
resource "aws_s3_bucket_versioning" "tfstate" {
  bucket = aws_s3_bucket.tfstate.id
  versioning_configuration { status = "Enabled" }
}
resource "aws_s3_bucket_server_side_encryption_configuration" "tfstate" {
  bucket = aws_s3_bucket.tfstate.id
  rule {
    apply_server_side_encryption_by_default {
      sse_algorithm     = "aws:kms"
      kms_master_key_id = aws_kms_key.tfstate.arn
    }
  }
}
resource "aws_s3_bucket_public_access_block" "tfstate" {
  bucket                  = aws_s3_bucket.tfstate.id
  block_public_acls       = true
  block_public_policy     = true
  ignore_public_acls      = true
  restrict_public_buckets = true
}
resource "aws_kms_key" "tfstate" {
  description         = "tfstate"
  enable_key_rotation = true
}
resource "aws_kms_alias" "tfstate" {
  name          = "alias/tfstate"
  target_key_id = aws_kms_key.tfstate.key_id
}
```

If you are on Terraform < 1.10, keep a DynamoDB lock table and set `dynamodb_table` in the backend block instead of `use_lockfile`.

---

### Resource: references/12-production-guardrails-don-t-skip-these.md

## Contents

- 12. Production Guardrails (don't skip these)
- CI role: scope it and bound it
- Detection: turn it on account-wide
- ECR: scan on push + expire old images
- RDS: parameter group + KMS + tested restores
- Canary / synthetic alarm
- Tagging & least-privilege defaults

## 12. Production Guardrails (don't skip these)

The modules above ship a working stack; these turn it into something you can defend in an audit and operate at 3am.

### CI role: scope it and bound it
The `github-actions-deploy` role assumed in section 5 must be locked to your repo via the OIDC `sub` claim and capped with a permissions boundary so a compromised workflow can't escalate.

```hcl
data "aws_iam_openid_connect_provider" "github" { url = "https://token.actions.githubusercontent.com" }

data "aws_iam_policy_document" "gha_assume" {
  statement {
    actions = ["sts:AssumeRoleWithWebIdentity"]
    principals {
      type        = "Federated"
      identifiers = [data.aws_iam_openid_connect_provider.github.arn]
    }
    condition {
      test     = "StringEquals"
      variable = "token.actions.githubusercontent.com:aud"
      values   = ["sts.amazonaws.com"]
    }
    # Lock to one repo + ref. NEVER use repo:org/*:* — that lets any repo assume it.
    condition {
      test     = "StringLike"
      variable = "token.actions.githubusercontent.com:sub"
      values   = ["repo:myorg/myapp:ref:refs/heads/main"]
    }
  }
}

resource "aws_iam_role" "gha_deploy" {
  name                 = "github-actions-deploy"
  assume_role_policy   = data.aws_iam_policy_document.gha_assume.json
  permissions_boundary = aws_iam_policy.gha_boundary.arn # caps max privilege
}
```

Pair this with GitHub Environments: the `environment: production` in section 5 should have **required reviewers** so a human approves each prod deploy (an approval gate, not just a label).

### Detection: turn it on account-wide
```hcl
resource "aws_guardduty_detector" "main" { enable = true }
resource "aws_securityhub_account" "main" {}
resource "aws_config_configuration_recorder" "main" {
  name     = "default"
  role_arn = aws_iam_role.config.arn
  recording_group {
    all_supported                 = true
    include_global_resource_types = true
  }
}
```
GuardDuty (threat detection), Security Hub (CIS/AWS Foundational Security Best Practices scoring), and AWS Config (resource compliance + drift) are the baseline three. Add Access Analyzer to catch public/cross-account exposure.

### ECR: scan on push + expire old images
```hcl
resource "aws_ecr_repository" "app" {
  name                 = "myapp"
  image_tag_mutability = "IMMUTABLE"          # tags can't be overwritten
  image_scanning_configuration { scan_on_push = true }
  encryption_configuration { encryption_type = "KMS" }
}
resource "aws_ecr_lifecycle_policy" "app" {
  repository = aws_ecr_repository.app.name
  policy = jsonencode({ rules = [{
    rulePriority = 1, description = "keep last 20 images"
    selection    = { tagStatus = "any", countType = "imageCountMoreThan", countNumber = 20 }
    action       = { type = "expire" }
  }] })
}
```
Note: `IMMUTABLE` tags mean the `:latest` retag in section 5's build step will fail — push only the immutable `:$IMAGE_TAG` and reference that, or use a mutable repo for `:latest`.

### RDS: parameter group + KMS + tested restores
- Attach an `aws_rds_cluster_parameter_group` to enforce `rds.force_ssl = 1`, sane `log_min_duration_statement`, and `log_statement = 'ddl'`.
- Encrypt with a customer-managed KMS key (`kms_key_id` on the cluster), not the default `aws/rds` key, so you control rotation and cross-account sharing.
- `backup_retention_period` (35 in section 3) is worthless if you've never restored. Periodically `aws rds restore-db-cluster-to-point-in-time` into a scratch cluster and smoke-test it. Consider `aws_backup` with cross-region copy for DR.

### Canary / synthetic alarm
The section 6 alarms are reactive. Add a CloudWatch Synthetics canary hitting a real user path and alarm on its `SuccessPercent`, so you detect "site is down" before customers do. Wire canary failure into the CodeDeploy `auto_rollback_configuration` alarms (section 2a) so a bad deploy rolls back automatically.

### Tagging & least-privilege defaults
Set a provider-level `default_tags` block (`Environment`, `Project`, `Owner`, `CostCenter`) so every resource is attributable in Cost Explorer and the budget alarm in section 6 is actionable. Run `tfsec`/`checkov`/`trivy config` in the `test` job (section 5) to catch insecure Terraform before apply.

### Resource: references/2-ecs-fargate-with-auto-scaling.md

## Contents

- 2. ECS Fargate with Auto-Scaling
- 2a. CodeDeploy blue/green resources
- 2b. Simpler alternative: ECS rolling deploy with circuit breaker

## 2. ECS Fargate with Auto-Scaling

```hcl
# modules/ecs/main.tf

variable "project" { type = string }
variable "environment" { type = string }
variable "vpc_id" { type = string }
variable "private_subnet_ids" { type = list(string) }
variable "public_subnet_ids" { type = list(string) }
variable "container_image" { type = string }
variable "container_port" { default = 3000 }
variable "cpu" { default = 512 }
variable "memory" { default = 1024 }
variable "desired_count" { default = 2 }
variable "min_count" { default = 2 }
variable "max_count" { default = 10 }
variable "health_check_path" { default = "/health" }
variable "secrets_arn" { type = string }
variable "certificate_arn" { type = string }
variable "admin_cidr" { type = string } # trusted CIDR for the blue/green test listener

resource "aws_ecs_cluster" "main" {
  name = "${var.project}-${var.environment}"
  setting {
    name  = "containerInsights"
    value = "enabled"
  }
}

resource "aws_cloudwatch_log_group" "app" {
  name              = "/ecs/${var.project}-${var.environment}/app"
  retention_in_days = 30
}

# Task execution role — pulls images, writes logs, reads secrets
resource "aws_iam_role" "task_execution" {
  name = "${var.project}-${var.environment}-task-exec"
  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{ Action = "sts:AssumeRole", Effect = "Allow", Principal = { Service = "ecs-tasks.amazonaws.com" } }]
  })
}

resource "aws_iam_role_policy_attachment" "task_execution" {
  role       = aws_iam_role.task_execution.name
  policy_arn = "arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy"
}

resource "aws_iam_role_policy" "task_execution_secrets" {
  role = aws_iam_role.task_execution.id
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{ Effect = "Allow", Action = ["secretsmanager:GetSecretValue"], Resource = [var.secrets_arn] }]
  })
}

# Task role — what YOUR CODE runs as. Least privilege.
resource "aws_iam_role" "task" {
  name = "${var.project}-${var.environment}-task"
  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{ Action = "sts:AssumeRole", Effect = "Allow", Principal = { Service = "ecs-tasks.amazonaws.com" } }]
  })
}

resource "aws_iam_role_policy" "task" {
  role = aws_iam_role.task.id
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      { Effect = "Allow", Action = ["s3:GetObject","s3:PutObject"], Resource = ["arn:aws:s3:::${var.project}-${var.environment}-uploads/*"] },
      # Permissions for the ADOT collector sidecar below: forward traces to
      # X-Ray and pull centralized sampling rules.
      { Effect = "Allow",
        Action = ["xray:PutTraceSegments","xray:PutTelemetryRecords","xray:GetSamplingRules","xray:GetSamplingTargets","xray:GetSamplingStatisticSummaries"],
        Resource = ["*"] },
      # Required for ECS Exec (enable_execute_command below). Without these four
      # SSM Messages actions on the TASK role, `aws ecs execute-command` fails with
      # "execute command failed because execute command was not enabled".
      { Effect = "Allow",
        Action = ["ssmmessages:CreateControlChannel","ssmmessages:CreateDataChannel","ssmmessages:OpenControlChannel","ssmmessages:OpenDataChannel"],
        Resource = ["*"] }
    ]
  })
}

data "aws_region" "current" {}

resource "aws_ecs_task_definition" "app" {
  family                   = "${var.project}-${var.environment}"
  network_mode             = "awsvpc"
  requires_compatibilities = ["FARGATE"]
  cpu                      = var.cpu
  memory                   = var.memory
  execution_role_arn       = aws_iam_role.task_execution.arn
  task_role_arn            = aws_iam_role.task.arn

  container_definitions = jsonencode([
    {
      name  = "app"
      image = var.container_image
      portMappings = [{ containerPort = var.container_port, protocol = "tcp" }]
      secrets = [
        { name = "DATABASE_URL", valueFrom = "${var.secrets_arn}:DATABASE_URL::" },
        { name = "REDIS_URL", valueFrom = "${var.secrets_arn}:REDIS_URL::" }
      ]
      environment = [
        { name = "NODE_ENV", value = var.environment },
        { name = "PORT", value = tostring(var.container_port) }
      ]
      logConfiguration = {
        logDriver = "awslogs"
        options = { "awslogs-group" = aws_cloudwatch_log_group.app.name, "awslogs-region" = data.aws_region.current.region, "awslogs-stream-prefix" = "app" }
      }
      healthCheck = {
        command = ["CMD-SHELL", "curl -f http://localhost:${var.container_port}/health || exit 1"]
        interval = 30, timeout = 5, retries = 3, startPeriod = 60
      }
    },
    # Tracing sidecar: ADOT collector (OpenTelemetry), forwards traces to X-Ray.
    # The X-Ray daemon and SDKs entered maintenance mode on February 25, 2026
    # (security fixes only); AWS recommends OpenTelemetry instrumentation. If you
    # must keep the daemon for an existing app, pin amazon/aws-xray-daemon:3.x,
    # never :latest.
    {
      name = "aws-otel-collector", image = "public.ecr.aws/aws-observability/aws-otel-collector:latest"
      cpu = 32, memory = 256, essential = false
      command = ["--config=/etc/ecs/ecs-default-config.yaml"]
      portMappings = [{ containerPort = 4317, protocol = "tcp" }, { containerPort = 4318, protocol = "tcp" }]
      logConfiguration = { logDriver = "awslogs", options = { "awslogs-group" = aws_cloudwatch_log_group.app.name, "awslogs-region" = data.aws_region.current.region, "awslogs-stream-prefix" = "otel" } }
    }
  ])
}

# Security groups
# Only CloudFront's origin-facing ranges may reach the ALB. Opening 80/443 to
# 0.0.0.0/0 would let clients hit the ALB DNS name directly and bypass the WAF
# and rate limits attached to CloudFront in section 4. For defense in depth,
# also verify a secret custom origin header at the ALB.
data "aws_ec2_managed_prefix_list" "cloudfront" {
  name = "com.amazonaws.global.cloudfront.origin-facing"
}

resource "aws_security_group" "alb" {
  name_prefix = "${var.project}-${var.environment}-alb-"
  vpc_id      = var.vpc_id
  ingress {
    from_port       = 443
    to_port         = 443
    protocol        = "tcp"
    prefix_list_ids = [data.aws_ec2_managed_prefix_list.cloudfront.id]
  }
  ingress {
    from_port       = 80
    to_port         = 80
    protocol        = "tcp"
    prefix_list_ids = [data.aws_ec2_managed_prefix_list.cloudfront.id]
  }
  egress {
    from_port   = 0
    to_port     = 0
    protocol    = "-1"
    cidr_blocks = ["0.0.0.0/0"]
  }
  lifecycle { create_before_destroy = true }
}

resource "aws_security_group" "ecs" {
  name_prefix = "${var.project}-${var.environment}-ecs-"
  vpc_id      = var.vpc_id
  ingress {
    from_port       = var.container_port
    to_port         = var.container_port
    protocol        = "tcp"
    security_groups = [aws_security_group.alb.id]
  }
  egress {
    from_port   = 0
    to_port     = 0
    protocol    = "-1"
    cidr_blocks = ["0.0.0.0/0"]
  }
  lifecycle { create_before_destroy = true }
}

# ALB
resource "aws_lb" "main" {
  name               = "${var.project}-${var.environment}"
  internal           = false
  load_balancer_type = "application"
  security_groups    = [aws_security_group.alb.id]
  subnets            = var.public_subnet_ids
  enable_deletion_protection = var.environment == "production"
  drop_invalid_header_fields = true
}

resource "aws_lb_listener" "https" {
  load_balancer_arn = aws_lb.main.arn
  port              = 443
  protocol          = "HTTPS"
  ssl_policy        = "ELBSecurityPolicy-TLS13-1-2-2021-06"
  certificate_arn   = var.certificate_arn
  default_action {
    type             = "forward"
    target_group_arn = aws_lb_target_group.blue.arn
  }
  lifecycle { ignore_changes = [default_action] }
}

resource "aws_lb_listener" "http_redirect" {
  load_balancer_arn = aws_lb.main.arn
  port              = 80
  protocol          = "HTTP"
  default_action {
    type = "redirect"
    redirect {
      port        = "443"
      protocol    = "HTTPS"
      status_code = "HTTP_301"
    }
  }
}

# --- Two target groups for CodeDeploy blue/green ---
# CodeDeploy swaps the production listener between these two groups. Both must
# exist up front; the running service is attached to exactly one at a time and
# CodeDeploy shifts traffic to the other on each deploy.
resource "aws_lb_target_group" "blue" {
  name_prefix          = "blue-"
  port                 = var.container_port
  protocol             = "HTTP"
  vpc_id               = var.vpc_id
  target_type          = "ip"
  deregistration_delay = 30
  health_check {
    path                = var.health_check_path
    healthy_threshold   = 2
    unhealthy_threshold = 3
    timeout             = 5
    interval            = 15
    matcher             = "200"
  }
  lifecycle { create_before_destroy = true }
}

resource "aws_lb_target_group" "green" {
  name_prefix          = "green-"
  port                 = var.container_port
  protocol             = "HTTP"
  vpc_id               = var.vpc_id
  target_type          = "ip"
  deregistration_delay = 30
  health_check {
    path                = var.health_check_path
    healthy_threshold   = 2
    unhealthy_threshold = 3
    timeout             = 5
    interval            = 15
    matcher             = "200"
  }
  lifecycle { create_before_destroy = true }
}

# Test listener on :8443 — lets CodeDeploy validate the green stack before it
# receives production traffic. Reuse the prod cert or a separate test cert.
resource "aws_lb_listener" "test" {
  load_balancer_arn = aws_lb.main.arn
  port              = 8443
  protocol          = "HTTPS"
  ssl_policy        = "ELBSecurityPolicy-TLS13-1-2-2021-06"
  certificate_arn   = var.certificate_arn
  default_action {
    type             = "forward"
    target_group_arn = aws_lb_target_group.green.arn
  }
  lifecycle { ignore_changes = [default_action] }
}

# Allow the test-listener port through the ALB and into the tasks. The green
# task set on :8443 is not yet validated, so never expose it to 0.0.0.0/0:
# scope it to a trusted admin CIDR (or the CloudFront prefix list above).
resource "aws_security_group_rule" "alb_test_ingress" {
  type              = "ingress"
  security_group_id = aws_security_group.alb.id
  from_port         = 8443
  to_port           = 8443
  protocol          = "tcp"
  cidr_blocks       = [var.admin_cidr] # e.g. your office/VPN CIDR
}

# ECS Service — CodeDeploy-controlled blue/green with auto-rollback.
# NOTE: deployment_controller = CODE_DEPLOY is INCOMPATIBLE with the ECS
# deployment_circuit_breaker / deployment_configuration blocks; rollback is
# configured on the CodeDeploy deployment group instead (see section 2a). If you
# prefer plain ECS rolling deploys, swap to the variant in section 2b — do NOT
# mix the two.
resource "aws_ecs_service" "app" {
  name            = "${var.project}-${var.environment}"
  cluster         = aws_ecs_cluster.main.id
  task_definition = aws_ecs_task_definition.app.arn
  desired_count   = var.desired_count
  launch_type     = "FARGATE"
  enable_execute_command = true

  deployment_controller { type = "CODE_DEPLOY" }

  network_configuration {
    subnets          = var.private_subnet_ids
    security_groups  = [aws_security_group.ecs.id]
    assign_public_ip = false
  }

  load_balancer {
    target_group_arn = aws_lb_target_group.blue.arn
    container_name   = "app"
    container_port   = var.container_port
  }

  # CodeDeploy mutates task_definition and load_balancer on each deploy; ignore
  # them so Terraform does not fight CodeDeploy.
  lifecycle { ignore_changes = [task_definition, load_balancer] }
}

# Auto-scaling on CPU and request count
resource "aws_appautoscaling_target" "ecs" {
  max_capacity       = var.max_count
  min_capacity       = var.min_count
  resource_id        = "service/${aws_ecs_cluster.main.name}/${aws_ecs_service.app.name}"
  scalable_dimension = "ecs:service:DesiredCount"
  service_namespace  = "ecs"
}

resource "aws_appautoscaling_policy" "cpu" {
  name               = "${var.project}-${var.environment}-cpu"
  policy_type        = "TargetTrackingScaling"
  resource_id        = aws_appautoscaling_target.ecs.resource_id
  scalable_dimension = aws_appautoscaling_target.ecs.scalable_dimension
  service_namespace  = aws_appautoscaling_target.ecs.service_namespace

  target_tracking_scaling_policy_configuration {
    predefined_metric_specification { predefined_metric_type = "ECSServiceAverageCPUUtilization" }
    target_value       = 65
    scale_in_cooldown  = 300
    scale_out_cooldown = 60
  }
}

resource "aws_appautoscaling_policy" "requests" {
  name               = "${var.project}-${var.environment}-requests"
  policy_type        = "TargetTrackingScaling"
  resource_id        = aws_appautoscaling_target.ecs.resource_id
  scalable_dimension = aws_appautoscaling_target.ecs.scalable_dimension
  service_namespace  = aws_appautoscaling_target.ecs.service_namespace

  target_tracking_scaling_policy_configuration {
    predefined_metric_specification {
      predefined_metric_type = "ALBRequestCountPerTarget"
      # With CodeDeploy blue/green the live target group alternates blue<->green,
      # so per-target request scaling is approximate right after a deploy. If you
      # need exact request-based scaling under blue/green, prefer a CPU/memory
      # target (above) or a custom CloudWatch metric on the ALB request count.
      resource_label = "${aws_lb.main.arn_suffix}/${aws_lb_target_group.blue.arn_suffix}"
    }
    target_value       = 1000
    scale_in_cooldown  = 300
    scale_out_cooldown = 60
  }
}
```

### 2a. CodeDeploy blue/green resources

These complete the blue/green deploy the service above declares. CodeDeploy needs an app, a deployment group bound to the ECS service + ALB listeners + both target groups, and an IAM role. The `AppSpec` and the GitHub Actions invocation are in section 5.

```hcl
# modules/ecs/codedeploy.tf

resource "aws_codedeploy_app" "app" {
  name             = "${var.project}-${var.environment}"
  compute_platform = "ECS"
}

resource "aws_iam_role" "codedeploy" {
  name = "${var.project}-${var.environment}-codedeploy"
  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{ Action = "sts:AssumeRole", Effect = "Allow", Principal = { Service = "codedeploy.amazonaws.com" } }]
  })
}

resource "aws_iam_role_policy_attachment" "codedeploy" {
  role       = aws_iam_role.codedeploy.name
  policy_arn = "arn:aws:iam::aws:policy/AWSCodeDeployRoleForECS"
}

resource "aws_codedeploy_deployment_group" "app" {
  app_name               = aws_codedeploy_app.app.name
  deployment_group_name  = "${var.project}-${var.environment}"
  service_role_arn       = aws_iam_role.codedeploy.arn
  deployment_config_name = "CodeDeployDefault.ECSCanary10Percent5Minutes"

  deployment_style {
    deployment_type   = "BLUE_GREEN"
    deployment_option = "WITH_TRAFFIC_CONTROL"
  }

  blue_green_deployment_config {
    # Spin up the green task set, run validation, then shift traffic.
    deployment_ready_option { action_on_timeout = "CONTINUE_DEPLOYMENT" }
    # Keep old (blue) task set for 15 min so you can roll back instantly.
    terminate_blue_instances_on_deployment_success {
      action                           = "TERMINATE"
      termination_wait_time_in_minutes = 15
    }
  }

  auto_rollback_configuration {
    enabled = true
    events  = ["DEPLOYMENT_FAILURE", "DEPLOYMENT_STOP_ON_ALARM"]
  }

  ecs_service {
    cluster_name = aws_ecs_cluster.main.name
    service_name = aws_ecs_service.app.name
  }

  load_balancer_info {
    target_group_pair_info {
      prod_traffic_route { listener_arns = [aws_lb_listener.https.arn] }
      test_traffic_route { listener_arns = [aws_lb_listener.test.arn] }
      target_group { name = aws_lb_target_group.blue.name }
      target_group { name = aws_lb_target_group.green.name }
    }
  }
}

output "ecs_cluster_name" { value = aws_ecs_cluster.main.name }
output "ecs_service_name" { value = aws_ecs_service.app.name }
output "codedeploy_app_name" { value = aws_codedeploy_app.app.name }
output "codedeploy_deployment_group" { value = aws_codedeploy_deployment_group.app.deployment_group_name }
output "alb_arn_suffix" { value = aws_lb.main.arn_suffix }
output "alb_dns_name" { value = aws_lb.main.dns_name }
output "ecs_security_group_id" { value = aws_security_group.ecs.id }
```

### 2b. Simpler alternative: ECS rolling deploy with circuit breaker

If you do NOT need blue/green (no per-deploy test traffic, faster rollouts are fine), drop section 2a, drop the test listener, and use the standard ECS rolling controller. Pick exactly one of 2a or 2b — `CODE_DEPLOY` and the circuit-breaker block are mutually exclusive.

```hcl
# Replacement for the aws_ecs_service.app body in section 2.
# deployment_controller defaults to ECS, so just omit it.
  enable_execute_command             = true
  deployment_minimum_healthy_percent = 100
  deployment_maximum_percent         = 200
  deployment_circuit_breaker {
    enable   = true
    rollback = true
  }
  # ECS-native rolling deploys mutate the task definition, so do NOT ignore it:
  lifecycle { ignore_changes = [] }
```

With 2b, the GitHub Actions "Deploy" step in section 5 (`aws ecs update-service --force-new-deployment`) is the correct deploy mechanism. With 2a, use the CodeDeploy step shown there instead.

---

### Resource: references/3-rds-aurora-with-read-replicas.md

## 3. RDS Aurora with Read Replicas

```hcl
# modules/rds/main.tf

variable "project" { type = string }
variable "environment" { type = string }
variable "vpc_id" { type = string }
variable "isolated_subnet_ids" { type = list(string) }
variable "ecs_security_group_id" { type = string }

resource "aws_db_subnet_group" "main" {
  name       = "${var.project}-${var.environment}"
  subnet_ids = var.isolated_subnet_ids
}

resource "aws_security_group" "rds" {
  name_prefix = "${var.project}-${var.environment}-rds-"
  vpc_id      = var.vpc_id
  ingress {
    from_port       = 5432
    to_port         = 5432
    protocol        = "tcp"
    security_groups = [var.ecs_security_group_id]
  }
  egress {
    from_port   = 0
    to_port     = 0
    protocol    = "-1"
    cidr_blocks = ["0.0.0.0/0"]
  }
}

# Pin a specific supported minor and pick your upgrade policy. As of mid-2026
# Aurora PostgreSQL supports the 14 / 15 / 16 / 17 / 18 major lines (13.x left
# standard support Feb 2026; 18.3 arrived Jun 2026); 16.x and 17.x carry LTS
# minors (16.8 and 17.7). Use a recent minor
# (e.g. 16.x LTS for stability, 17.x for newest features) and let AWS apply
# patch upgrades in the maintenance window. Verify the current minor list at
# https://docs.aws.amazon.com/AmazonRDS/latest/AuroraPostgreSQLReleaseNotes/AuroraPostgreSQL.Updates.html
variable "engine_version" { default = "16.8" } # LTS line; bump deliberately

resource "aws_rds_cluster" "main" {
  cluster_identifier                  = "${var.project}-${var.environment}"
  engine                              = "aurora-postgresql"
  engine_version                      = var.engine_version
  allow_major_version_upgrade         = false # set true only for a planned major upgrade
  apply_immediately                   = false # batch changes into the maintenance window
  preferred_maintenance_window        = "sun:05:00-sun:06:00"
  database_name                       = replace(var.project, "-", "_")
  master_username                     = "dbadmin"
  manage_master_user_password         = true
  iam_database_authentication_enabled = true
  db_subnet_group_name                = aws_db_subnet_group.main.name
  vpc_security_group_ids              = [aws_security_group.rds.id]
  backup_retention_period             = 35
  preferred_backup_window             = "03:00-04:00"
  copy_tags_to_snapshot               = true
  deletion_protection                 = var.environment == "production"
  storage_encrypted                   = true
  enabled_cloudwatch_logs_exports     = ["postgresql"]

  serverlessv2_scaling_configuration {
    min_capacity = var.environment == "production" ? 2 : 0.5
    max_capacity = var.environment == "production" ? 16 : 4
  }
}

resource "aws_rds_cluster_instance" "writer" {
  identifier                   = "${var.project}-${var.environment}-writer"
  cluster_identifier           = aws_rds_cluster.main.id
  instance_class               = "db.serverless"
  engine                       = aws_rds_cluster.main.engine
  engine_version               = aws_rds_cluster.main.engine_version
  performance_insights_enabled = true
  monitoring_interval          = 30
  monitoring_role_arn          = aws_iam_role.rds_monitoring.arn
}

resource "aws_rds_cluster_instance" "reader" {
  count                        = var.environment == "production" ? 2 : 1
  identifier                   = "${var.project}-${var.environment}-reader-${count.index}"
  cluster_identifier           = aws_rds_cluster.main.id
  instance_class               = "db.serverless"
  engine                       = aws_rds_cluster.main.engine
  engine_version               = aws_rds_cluster.main.engine_version
  performance_insights_enabled = true
  monitoring_interval          = 30
  monitoring_role_arn          = aws_iam_role.rds_monitoring.arn
}

resource "aws_iam_role" "rds_monitoring" {
  name = "${var.project}-${var.environment}-rds-mon"
  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{ Action = "sts:AssumeRole", Effect = "Allow", Principal = { Service = "monitoring.rds.amazonaws.com" } }]
  })
}

resource "aws_iam_role_policy_attachment" "rds_monitoring" {
  role       = aws_iam_role.rds_monitoring.name
  policy_arn = "arn:aws:iam::aws:policy/service-role/AmazonRDSEnhancedMonitoringRole"
}

output "cluster_endpoint" { value = aws_rds_cluster.main.endpoint }
output "reader_endpoint" { value = aws_rds_cluster.main.reader_endpoint }
```

---

### Resource: references/4-cloudfront-s3-waf.md

## 4. CloudFront + S3 + WAF

```hcl
# modules/cdn/main.tf

variable "project" { type = string }
variable "environment" { type = string }
variable "domain_name" { type = string }
variable "alb_dns_name" { type = string }
# CloudFront + CLOUDFRONT-scoped WAF certs MUST live in us-east-1. Pass an ACM
# cert ARN from us-east-1 here (see the provider alias note below).
variable "certificate_arn" { type = string }

# CloudFront and a CLOUDFRONT-scoped WAFv2 ACL can only be created in us-east-1.
# Declare a us-east-1 provider alias in the ROOT module and pass it to this
# module via `providers = { aws = aws, aws.us_east_1 = aws.us_east_1 }`:
#
#   # root main.tf
#   provider "aws" { region = "eu-west-1" }            # your primary region
#   provider "aws" {
#     alias  = "us_east_1"
#     region = "us-east-1"
#   }
#   module "cdn" {
#     source    = "./modules/cdn"
#     providers = { aws = aws, aws.us_east_1 = aws.us_east_1 }
#     ...
#   }
#
# and require both in the module:
terraform {
  required_providers {
    aws = { source = "hashicorp/aws", configuration_aliases = [aws.us_east_1] }
  }
}

resource "aws_s3_bucket" "assets" {
  bucket = "${var.project}-${var.environment}-assets"
}

resource "aws_s3_bucket_public_access_block" "assets" {
  bucket                  = aws_s3_bucket.assets.id
  block_public_acls       = true
  block_public_policy     = true
  ignore_public_acls      = true
  restrict_public_buckets = true
}

resource "aws_cloudfront_origin_access_control" "s3" {
  name                              = "${var.project}-${var.environment}-s3"
  origin_access_control_origin_type = "s3"
  signing_behavior                  = "always"
  signing_protocol                  = "sigv4"
}

# OAC requires a bucket policy granting the CloudFront SERVICE principal
# s3:GetObject, scoped to THIS distribution via AWS:SourceArn. Without it,
# every object 403s because public access is blocked above.
resource "aws_s3_bucket_policy" "assets" {
  bucket = aws_s3_bucket.assets.id
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Sid       = "AllowCloudFrontOAC"
      Effect    = "Allow"
      Principal = { Service = "cloudfront.amazonaws.com" }
      Action    = "s3:GetObject"
      Resource  = "${aws_s3_bucket.assets.arn}/*"
      Condition = { StringEquals = { "AWS:SourceArn" = aws_cloudfront_distribution.main.arn } }
    }]
  })
}

# Managed cache/origin-request/response-header policies (replace legacy
# forwarded_values). These IDs are AWS-managed and stable across accounts.
data "aws_cloudfront_cache_policy" "caching_optimized" { name = "Managed-CachingOptimized" }
data "aws_cloudfront_cache_policy" "caching_disabled" { name = "Managed-CachingDisabled" }
data "aws_cloudfront_origin_request_policy" "all_viewer_except_host" { name = "Managed-AllViewerExceptHostHeader" }
data "aws_cloudfront_response_headers_policy" "security" { name = "Managed-SecurityHeadersPolicy" }

resource "aws_cloudfront_distribution" "main" {
  enabled         = true
  is_ipv6_enabled = true
  aliases         = [var.domain_name]
  price_class     = "PriceClass_100"
  web_acl_id      = aws_wafv2_web_acl.main.arn

  origin {
    domain_name = var.alb_dns_name
    origin_id   = "alb"
    custom_origin_config {
      http_port              = 80
      https_port             = 443
      origin_protocol_policy = "https-only"
      origin_ssl_protocols   = ["TLSv1.2"]
    }
  }

  origin {
    domain_name              = aws_s3_bucket.assets.bucket_regional_domain_name
    origin_id                = "s3-assets"
    origin_access_control_id = aws_cloudfront_origin_access_control.s3.id
  }

  # Static assets — immutable, long cache. CachingOptimized strips cookies,
  # compresses, and respects Cache-Control from the origin.
  ordered_cache_behavior {
    path_pattern               = "/_next/static/*"
    allowed_methods            = ["GET", "HEAD"]
    cached_methods             = ["GET", "HEAD"]
    target_origin_id           = "s3-assets"
    compress                   = true
    viewer_protocol_policy     = "redirect-to-https"
    cache_policy_id            = data.aws_cloudfront_cache_policy.caching_optimized.id
    response_headers_policy_id = data.aws_cloudfront_response_headers_policy.security.id
  }

  # Default — dynamic, forward to ALB. CachingDisabled = no caching;
  # AllViewerExceptHostHeader forwards query strings, cookies, and headers
  # (minus Host, which must resolve to the ALB origin).
  default_cache_behavior {
    allowed_methods          = ["DELETE","GET","HEAD","OPTIONS","PATCH","POST","PUT"]
    cached_methods           = ["GET","HEAD"]
    target_origin_id         = "alb"
    viewer_protocol_policy    = "redirect-to-https"
    compress                  = true
    cache_policy_id           = data.aws_cloudfront_cache_policy.caching_disabled.id
    origin_request_policy_id  = data.aws_cloudfront_origin_request_policy.all_viewer_except_host.id
  }

  viewer_certificate {
    acm_certificate_arn      = var.certificate_arn
    ssl_support_method       = "sni-only"
    minimum_protocol_version = "TLSv1.2_2021"
  }

  restrictions { geo_restriction { restriction_type = "none" } }
}

# WAF — rate limiting + OWASP managed rules.
# A CLOUDFRONT-scoped WAFv2 ACL MUST be created in us-east-1, hence the aliased
# provider declared in the module header above.
resource "aws_wafv2_web_acl" "main" {
  provider = aws.us_east_1
  name     = "${var.project}-${var.environment}"
  scope    = "CLOUDFRONT"

  default_action { allow {} }

  rule {
    name     = "rate-limit"
    priority = 1
    action { block {} }
    statement {
      rate_based_statement {
        limit              = 2000
        aggregate_key_type = "IP"
      }
    }
    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "rate-limit"
      sampled_requests_enabled   = true
    }
  }

  rule {
    name     = "aws-managed-common"
    priority = 2
    override_action { none {} }
    statement {
      managed_rule_group_statement {
        name        = "AWSManagedRulesCommonRuleSet"
        vendor_name = "AWS"
      }
    }
    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "common"
      sampled_requests_enabled   = true
    }
  }

  rule {
    name     = "aws-managed-sqli"
    priority = 3
    override_action { none {} }
    statement {
      managed_rule_group_statement {
        name        = "AWSManagedRulesSQLiRuleSet"
        vendor_name = "AWS"
      }
    }
    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "sqli"
      sampled_requests_enabled   = true
    }
  }

  visibility_config {
    cloudwatch_metrics_enabled = true
    metric_name                = "${var.project}-waf"
    sampled_requests_enabled   = true
  }
}
```

---

### Resource: references/5-ci-cd-github-actions-to-ecs.md

## 5. CI/CD — GitHub Actions to ECS

```yaml
# .github/workflows/deploy.yml
name: Deploy to ECS
on:
  push:
    branches: [main]

concurrency:
  group: deploy-${{ github.ref }}
  cancel-in-progress: false

permissions:
  id-token: write
  contents: read

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: actions/setup-node@v6
        # Use a current LTS. Node 24 is Active LTS (through Oct 2026); Node 22
        # moved to Maintenance LTS in Oct 2025 (security fixes to Apr 2027);
        # Node 20 is end of life (Apr 2026), do not use. Match this to the
        # runtime in your Dockerfile.
        with: { node-version: 24, cache: npm }
      - run: npm ci && npm test && npm run lint && npm run typecheck

  deploy:
    needs: test
    runs-on: ubuntu-latest
    environment: production
    steps:
      - uses: actions/checkout@v7

      - uses: aws-actions/configure-aws-credentials@v6
        with:
          role-to-assume: arn:aws:iam::${{ secrets.AWS_ACCOUNT_ID }}:role/github-actions-deploy
          aws-region: us-east-1

      - uses: aws-actions/amazon-ecr-login@v2
        id: ecr

      - name: Build and push
        id: build
        env:
          ECR_REGISTRY: ${{ steps.ecr.outputs.registry }}
          IMAGE_TAG: ${{ github.sha }}
        run: |
          docker build --cache-from $ECR_REGISTRY/myapp:latest \
            -t $ECR_REGISTRY/myapp:$IMAGE_TAG -t $ECR_REGISTRY/myapp:latest .
          docker push $ECR_REGISTRY/myapp:$IMAGE_TAG
          docker push $ECR_REGISTRY/myapp:latest
          echo "image=$ECR_REGISTRY/myapp:$IMAGE_TAG" >> $GITHUB_OUTPUT

      # Requires a `myapp-production-migrate` task definition with a `migrate`
      # container that has DATABASE_URL injected as an ECS secret (valueFrom the
      # same Secrets Manager secret as the app) and the SAME task/execution roles
      # as the app. Register it in Terraform (a second aws_ecs_task_definition with
      # the migrate command), or reuse the app task def and only override the
      # command as below. SUBNETS/SG secrets must each be a JSON-array-safe,
      # COMMA-separated list with NO spaces, e.g. subnet-aaa,subnet-bbb — the CLI
      # parses subnets=[a,b]. Quote them if a single value to avoid shell globbing.
      - name: Run migrations
        env:
          # comma-separated, no spaces: "subnet-aaa,subnet-bbb,subnet-ccc"
          SUBNETS: ${{ secrets.PRIVATE_SUBNET_IDS }}
          SG: ${{ secrets.ECS_SECURITY_GROUP_ID }}
        run: |
          set -euo pipefail
          NETCFG="awsvpcConfiguration={subnets=[${SUBNETS}],securityGroups=[${SG}],assignPublicIp=DISABLED}"
          TASK_ARN=$(aws ecs run-task --cluster myapp-production \
            --task-definition myapp-production-migrate --launch-type FARGATE \
            --network-configuration "$NETCFG" \
            --overrides '{"containerOverrides":[{"name":"migrate","command":["npx","prisma","migrate","deploy"]}]}' \
            --query 'tasks[0].taskArn' --output text)
          aws ecs wait tasks-stopped --cluster myapp-production --tasks "$TASK_ARN"
          EXIT=$(aws ecs describe-tasks --cluster myapp-production --tasks "$TASK_ARN" \
            --query 'tasks[0].containers[?name==`migrate`].exitCode | [0]' --output text)
          [ "$EXIT" = "0" ] || { echo "migration exited $EXIT"; exit 1; }

      # Register the new task-definition revision (shared by both deploy styles).
      - name: Register task definition
        id: taskdef
        run: |
          set -euo pipefail
          TASK_DEF=$(aws ecs describe-task-definition --task-definition myapp-production --query 'taskDefinition')
          NEW_DEF=$(echo "$TASK_DEF" | jq --arg IMG "${{ steps.build.outputs.image }}" \
            '.containerDefinitions[0].image = $IMG | del(.taskDefinitionArn,.revision,.status,.requiresAttributes,.compatibilities,.registeredAt,.registeredBy)')
          NEW_ARN=$(aws ecs register-task-definition --cli-input-json "$NEW_DEF" --query 'taskDefinition.taskDefinitionArn' --output text)
          echo "arn=$NEW_ARN" >> "$GITHUB_OUTPUT"

      # --- Deploy variant A: CodeDeploy blue/green (matches section 2a) ---
      # update-service is REJECTED on a CODE_DEPLOY-controlled service, so drive
      # the deploy through CodeDeploy with an AppSpec that names the new revision.
      - name: Deploy (CodeDeploy blue/green)
        run: |
          set -euo pipefail
          APPSPEC=$(jq -n --arg TD "${{ steps.taskdef.outputs.arn }}" '{
            version: "0.0",
            Resources: [{ TargetService: { Type: "AWS::ECS::Service", Properties: {
              TaskDefinition: $TD,
              LoadBalancerInfo: { ContainerName: "app", ContainerPort: 3000 }
            }}}]
          }')
          DEP_ID=$(aws deploy create-deployment \
            --application-name myapp-production \
            --deployment-group-name myapp-production \
            --revision "revisionType=AppSpecContent,appSpecContent={content='$APPSPEC'}" \
            --query 'deploymentId' --output text)
          aws deploy wait deployment-successful --deployment-id "$DEP_ID"

      # --- Deploy variant B: ECS rolling (use INSTEAD of A if you chose 2b) ---
      # - name: Deploy (ECS rolling)
      #   run: |
      #     aws ecs update-service --cluster myapp-production --service myapp-production \
      #       --task-definition ${{ steps.taskdef.outputs.arn }} --force-new-deployment
      #     aws ecs wait services-stable --cluster myapp-production --services myapp-production

      - name: Verify
        run: |
          set -euo pipefail
          for i in {1..5}; do
            [ "$(curl -so /dev/null -w '%{http_code}' https://api.example.com/health)" = "200" ] && break
            [ "$i" = "5" ] && { echo "health check never returned 200"; exit 1; }
            sleep 2
          done
```

---

### Resource: references/6-monitoring-cost-alerts.md

## 6. Monitoring & Cost Alerts

```hcl
resource "aws_sns_topic" "alerts" {
  name = "${var.project}-${var.environment}-alerts"
}

resource "aws_cloudwatch_metric_alarm" "alb_5xx" {
  alarm_name          = "${var.project}-high-5xx"
  namespace           = "AWS/ApplicationELB"
  metric_name         = "HTTPCode_Target_5XX_Count"
  statistic           = "Sum"
  period              = 300
  evaluation_periods  = 2
  threshold           = 50
  comparison_operator = "GreaterThanThreshold"
  alarm_actions       = [aws_sns_topic.alerts.arn]
  dimensions          = { LoadBalancer = var.alb_arn_suffix }
}

resource "aws_cloudwatch_metric_alarm" "latency_p99" {
  alarm_name          = "${var.project}-high-latency"
  namespace           = "AWS/ApplicationELB"
  metric_name         = "TargetResponseTime"
  extended_statistic  = "p99"
  period              = 300
  evaluation_periods  = 3
  threshold           = 2
  comparison_operator = "GreaterThanThreshold"
  alarm_actions       = [aws_sns_topic.alerts.arn]
  dimensions          = { LoadBalancer = var.alb_arn_suffix }
}

resource "aws_cloudwatch_metric_alarm" "ecs_cpu" {
  alarm_name          = "${var.project}-ecs-cpu"
  namespace           = "AWS/ECS"
  metric_name         = "CPUUtilization"
  statistic           = "Average"
  period              = 300
  evaluation_periods  = 3
  threshold           = 80
  comparison_operator = "GreaterThanThreshold"
  alarm_actions       = [aws_sns_topic.alerts.arn]
  dimensions          = { ClusterName = var.ecs_cluster_name, ServiceName = var.ecs_service_name }
}

resource "aws_budgets_budget" "monthly" {
  name         = "${var.project}-monthly"
  budget_type  = "COST"
  limit_amount = "500"
  limit_unit   = "USD"
  time_unit    = "MONTHLY"

  notification {
    comparison_operator        = "GREATER_THAN"
    threshold                  = 80
    threshold_type             = "PERCENTAGE"
    notification_type          = "ACTUAL"
    subscriber_email_addresses = [var.alert_email]
  }
}
```

---

### Resource: references/7-database-migration-strategy.md

## Contents

- 7. Database Migration Strategy
- Safe migration pattern:
- Dangerous vs safe:
- Rollback:

## 7. Database Migration Strategy

**Golden rule: migrations must be backward-compatible.** Old and new code run simultaneously during deployment.

### Safe migration pattern:
```
Deploy 1: ADD new column (nullable)
Deploy 2: Write to BOTH columns
Deploy 3: Backfill old rows in batches
Deploy 4: Read from new column only
Deploy 5: DROP old column
```

### Dangerous vs safe:
```sql
-- NEVER (locks table):
ALTER TABLE users ADD COLUMN verified boolean NOT NULL DEFAULT false;

-- SAFE (two steps):
ALTER TABLE users ADD COLUMN verified boolean;
-- Backfill in batches:
UPDATE users SET verified = false WHERE verified IS NULL AND id BETWEEN $1 AND $2;
-- Then:
ALTER TABLE users ALTER COLUMN verified SET DEFAULT false;
ALTER TABLE users ALTER COLUMN verified SET NOT NULL;
```

### Rollback:
```bash
aws ecs describe-services --cluster myapp-prod --services myapp-prod \
  --query 'services[0].taskDefinition' --output text > /tmp/last-good
# If things break:
aws ecs update-service --cluster myapp-prod --service myapp-prod \
  --task-definition $(cat /tmp/last-good) --force-new-deployment
```

---

### Resource: references/8-cdk-alternative.md

## 8. CDK Alternative

```typescript
import * as cdk from 'aws-cdk-lib';
import * as ec2 from 'aws-cdk-lib/aws-ec2';
import * as ecs from 'aws-cdk-lib/aws-ecs';
import * as ecs_patterns from 'aws-cdk-lib/aws-ecs-patterns';
import * as rds from 'aws-cdk-lib/aws-rds';

export class ProductionStack extends cdk.Stack {
  constructor(scope: cdk.App, id: string, props?: cdk.StackProps) {
    super(scope, id, props);

    const vpc = new ec2.Vpc(this, 'Vpc', {
      maxAzs: 3, natGateways: 3,
      subnetConfiguration: [
        { name: 'Public', subnetType: ec2.SubnetType.PUBLIC, cidrMask: 20 },
        { name: 'Private', subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS, cidrMask: 20 },
        { name: 'Isolated', subnetType: ec2.SubnetType.PRIVATE_ISOLATED, cidrMask: 20 },
      ],
    });

    const db = new rds.DatabaseCluster(this, 'Database', {
      // CDK v2. The AuroraPostgresEngineVersion enum often lags AWS's released
      // minors, so prefer `.of()` with an explicit supported version (see the
      // RDS module note in section 3). Use `VER_16_x`/`VER_17_x` if present in
      // your aws-cdk-lib version.
      engine: rds.DatabaseClusterEngine.auroraPostgres({
        version: rds.AuroraPostgresEngineVersion.of('16.8', '16'),
      }),
      serverlessV2MinCapacity: 2, serverlessV2MaxCapacity: 16,
      writer: rds.ClusterInstance.serverlessV2('writer'),
      readers: [rds.ClusterInstance.serverlessV2('reader1', { scaleWithWriter: true })],
      vpc, vpcSubnets: { subnetType: ec2.SubnetType.PRIVATE_ISOLATED },
      backup: { retention: cdk.Duration.days(35) },
      deletionProtection: true, storageEncrypted: true,
    });

    const service = new ecs_patterns.ApplicationLoadBalancedFargateService(this, 'Service', {
      vpc, taskSubnets: { subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS },
      cpu: 512, memoryLimitMiB: 1024, desiredCount: 2,
      taskImageOptions: {
        image: ecs.ContainerImage.fromAsset('.'),
        containerPort: 3000,
        secrets: { DATABASE_URL: ecs.Secret.fromSecretsManager(db.secret!, 'url') },
        environment: { NODE_ENV: 'production' },
      },
      circuitBreaker: { rollback: true },
    });

    const scaling = service.service.autoScaleTaskCount({ minCapacity: 2, maxCapacity: 10 });
    scaling.scaleOnCpuUtilization('Cpu', { targetUtilizationPercent: 65 });
    scaling.scaleOnRequestCount('Req', { requestsPerTarget: 1000, targetGroup: service.targetGroup });
    db.connections.allowDefaultPortFrom(service.service);
  }
}
```

---

### Resource: references/9-cost-optimization.md

## 9. Cost Optimization

| Resource | Dev | Production |
|----------|-----|------------|
| NAT Gateway | 1 | 1 per AZ |
| RDS | Serverless min 0.5 | Serverless min 2 |
| ECS | 256/512 | 512/1024+ |
| Logs retention | 7 days | 30-90 days |

**Biggest cost trap: NAT Gateway data charges.** Route ECR pulls and log shipping through VPC endpoints so they bypass NAT. Pulling images needs ALL of: `ecr.dkr` + `ecr.api` (interface) + `s3` (gateway — ECR layers live in S3). Interface endpoints also need a security group that allows 443 from the ECS tasks.

```hcl
# Interface endpoints need 443 ingress from the workloads using them.
resource "aws_security_group" "vpce" {
  name_prefix = "${var.project}-${var.environment}-vpce-"
  vpc_id      = aws_vpc.main.id
  ingress {
    from_port   = 443
    to_port     = 443
    protocol    = "tcp"
    cidr_blocks = [aws_vpc.main.cidr_block]
  }
  lifecycle { create_before_destroy = true }
}

# Gateway endpoints (S3 + DynamoDB) are FREE — no hourly or data charge.
resource "aws_vpc_endpoint" "s3" {
  vpc_id            = aws_vpc.main.id
  service_name      = "com.amazonaws.${data.aws_region.current.region}.s3"
  vpc_endpoint_type = "Gateway"
  route_table_ids   = aws_route_table.private[*].id
}

# Interface endpoints (ECR + logs) bill per-AZ-hour + per-GB; still far cheaper
# than NAT data transfer for steady image pulls and log volume.
locals {
  interface_endpoints = toset(["ecr.dkr", "ecr.api", "logs", "secretsmanager"])
}
resource "aws_vpc_endpoint" "interface" {
  for_each            = local.interface_endpoints
  vpc_id              = aws_vpc.main.id
  service_name        = "com.amazonaws.${data.aws_region.current.region}.${each.key}"
  vpc_endpoint_type   = "Interface"
  subnet_ids          = aws_subnet.private[*].id
  security_group_ids  = [aws_security_group.vpce.id]
  private_dns_enabled = true
}
```

Endpoints trade NAT data-transfer cost for per-endpoint hourly + per-GB fees, so the net saving depends on traffic and region — measure with Cost Explorer and verify current rates at https://aws.amazon.com/privatelink/pricing/ and https://aws.amazon.com/vpc/pricing/. Gateway endpoints (S3/DynamoDB) are free, so add them unconditionally.

---

### Resource: references/architecture-overview.md

## Architecture Overview

```
                    ┌─────────────┐
                    │  Route 53   │
                    └──────┬──────┘
                           │
                    ┌──────▼──────┐
                    │ CloudFront  │──── S3 (static assets)
                    └──────┬──────┘
                           │
                    ┌──────▼──────┐
                    │     ALB     │  (public subnets)
                    └──────┬──────┘
                           │
              ┌────────────┼────────────┐
              │            │            │
         ┌────▼───┐  ┌────▼───┐  ┌────▼───┐
         │ECS Task│  │ECS Task│  │ECS Task│  (private subnets)
         └────┬───┘  └────┬───┘  └────┬───┘
              │            │            │
              └────────────┼────────────┘
                           │
                    ┌──────▼──────┐
                    │   RDS       │  (isolated subnets)
                    │  Primary +  │
                    │  Read Replica│
                    └─────────────┘
```

---

---

## bing-webmaster
Category: analytics
Description: Bing Webmaster Tools setup, IndexNow protocol, URL submission, backlink analysis, AI Performance report (Feb 2026), Generative Engine Optimization (GEO) per Bing's official guidelines, Copilot citation tracking, meta-directive controls for AI surfaces, and Bing-specific SEO. Use when setting up Bing Webmaster Tools, implementing IndexNow, tracking Copilot citations, optimizing for Bing search or AI answers.
Features:
  - Bing Webmaster Tools setup and verification
  - AI Performance report — first-party Copilot citation analytics (Feb 2026)
  - Generative Engine Optimization per Bing's official guidelines
  - Meta-directive controls for AI surfaces (NOARCHIVE, NOCACHE, DATA-NOSNIPPET)
  - IndexNow protocol for AI freshness
  - URL submission and crawl control
  - Backlink profile analysis
  - 2026 abuse policy compliance (prompt injection, artificially engineered language)
Use Cases:
  - Set up Bing Webmaster Tools for a new site
  - Track Copilot citations via AI Performance report
  - Implement GEO best practices Bing officially endorses
  - Configure meta directives to control AI vs classic indexing
  - Implement IndexNow for instant AI-eligible indexing

# Bing Webmaster Tools & GEO

## Workflow

### 1. Setup & Verification

**Verification methods (pick one):**
- XML file upload (`BingSiteAuth.xml` to root)
- Meta tag (`<meta name="msvalidate.01" content="XXXX" />`)
- CNAME DNS record
- Auto-verify if already in Google Search Console (import)

**Import from GSC:** Bing offers one-click import of all your GSC properties — fastest path.

### 2. IndexNow

IndexNow tells search engines about URL changes instantly. Bing explicitly recommends it for AI freshness: *"IndexNow helps ensure that AI systems reference the most current version of a page when generating answers."*

**Key file (one-time setup):** generate a key (8–128 hex chars), serve it as a static text file whose body is exactly the key.
```bash
KEY=$(openssl rand -hex 16)          # e.g. a1b2c3...; treat as a public token, not a secret
echo -n "$KEY" > "public/$KEY.txt"   # served verbatim at https://example.com/$KEY.txt
```
On Next.js/Vercel anything in `public/` is served at the domain root automatically; on other hosts drop the file in the web root. Verify with `curl https://example.com/$KEY.txt` before submitting.

**Single URL** — always URL-encode the submitted URL so query strings/anchors don't break the request:
```bash
curl -G "https://api.indexnow.org/indexnow" \
  --data-urlencode "url=https://example.com/updated-page?ref=launch" \
  --data-urlencode "key=$KEY"
```

**Batch (up to 10,000 URLs per request):** all URLs must share the same host as `host`/`keyLocation`.
```bash
curl -X POST "https://api.indexnow.org/indexnow" \
  -H "Content-Type: application/json" \
  -d '{
    "host": "example.com",
    "key": "'"$KEY"'",
    "keyLocation": "https://example.com/'"$KEY"'.txt",
    "urlList": [
      "https://example.com/page1",
      "https://example.com/page2"
    ]
  }'
```

**Response codes to handle:** `200` accepted · `202` accepted, key validation pending · `400` invalid format · `403` key not found/invalid at `keyLocation` · `422` URL doesn't match host or key mismatch · `429` too many requests (back off and retry with exponential delay). Submitting to any one participating engine (Bing, Yandex, etc.) propagates to the others, so one POST to `api.indexnow.org` is enough. Log the URLs submitted plus the returned status so you can audit coverage; only re-submit a URL when its content actually changes — repeated pings of unchanged URLs add no value.

**Auto-trigger on deploy (Next.js example):**
```javascript
const changedUrls = getChangedPages();
if (changedUrls.length > 0) {
  await fetch('https://api.indexnow.org/indexnow', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({
      host: 'example.com',
      key: process.env.INDEXNOW_KEY,
      keyLocation: `https://example.com/${process.env.INDEXNOW_KEY}.txt`,
      urlList: changedUrls
    })
  });
}
```

**Deriving `getChangedPages()`** — pick the source that matches your stack; the goal is "files that changed in this deploy → public URLs":

- **Git diff (CI / GitHub Actions, most reliable):** diff the deployed commit against the previous one and map page files to routes.
  ```javascript
  import { execSync } from 'node:child_process';

  function getChangedPages() {
    // GitHub Actions exposes the previous SHA as github.event.before; locally fall back to HEAD~1
    const base = process.env.GITHUB_EVENT_BEFORE || 'HEAD~1';
    const files = execSync(`git diff --name-only ${base} HEAD`, { encoding: 'utf8' })
      .split('\n')
      .filter(Boolean);

    return files
      .filter(f => /^src\/app\/.*\/page\.(tsx|mdx)$/.test(f) || /^content\/.*\.mdx$/.test(f))
      .map(fileToUrl)
      .filter(Boolean);
  }

  function fileToUrl(file) {
    // src/app/blog/my-post/page.tsx -> https://example.com/blog/my-post
    const route = file
      .replace(/^src\/app/, '')
      .replace(/\/page\.(tsx|mdx)$/, '')
      .replace(/^content/, '')          // content/blog/x.mdx -> /blog/x
      .replace(/\.mdx$/, '')
      .replace(/\/\(.*?\)/g, '')         // strip Next.js route groups: /(marketing)/about -> /about
      .replace(/\/index$/, '');          // /blog/index -> /blog
    return `https://example.com${route || '/'}`;
  }
  ```
- **Sitemap diff (CMS/SSG without clean Git→route mapping):** fetch the freshly built `sitemap.xml`, compare each `<loc>`'s `<lastmod>` against the previous build's sitemap (cache it as a build artifact), and submit only URLs whose `lastmod` advanced.
- **CMS webhook:** trigger from the CMS publish event and read the changed slug(s) straight out of the webhook payload (e.g. Sanity/Contentful/WordPress `post.permalink`) — most precise for editorial sites.
- **Build manifest:** for incremental static regeneration, map regenerated paths from the build output (e.g. Next.js `.next/server/app` route manifest) to URLs.

Whichever source you use, dedupe and cap each POST at 10,000 URLs.

### 3. AI Performance report (public preview, Feb 2026)

Bing Webmaster Tools → **AI Performance**. Microsoft positions this as the first first-party AI-citation analytics shipped by a major search engine; it's the most direct citation telemetry available today regardless of who got there first.

What it shows:
- How often your content is cited in Copilot, Bing AI summaries, and partner integrations
- Which URLs are referenced
- Citation activity over time
- Since June 2026 (preview rollout): Intents (query-intent categories behind citations), Topics (thematic query clusters), Citation Share (your share of all citations for a grounding query), and Compare (overlay prior time periods)

How to use it:
- **Audit gaps:** URLs ranking well in classic Search but missing from AI citations are GEO opportunities.
- **Track wins:** Watch citation count after structural rewrites (answer-first, schema, IndexNow ping).
- **Compare with referrer logs:** Cross-check `bing.com/chat` and `copilot.microsoft.com` referrers against the report.

### 4. Generative Engine Optimization (GEO) — Bing's official guidance

Bing's updated webmaster guidelines now define **GEO** as *"focused on content eligibility for grounding and reference in AI responses."* Bing states GEO doesn't guarantee citation — same as SEO doesn't guarantee ranking.

**Best practices Bing lists for becoming a grounding source:**

| Practice | What Bing says |
|---|---|
| Clear facts | Present facts directly, no vague claims |
| Entity references | Avoid ambiguous references; use full names |
| Naming consistency | Same entity names across text, images, video |
| Topic focus | One topic per URL |
| Information placement | Key info near the top of the page |
| Structure | Clear headings, tables, FAQ sections |
| Evidence | Examples, data, cited sources |
| Freshness | Regular updates + IndexNow notifications |
| Authority | Deepen coverage in related areas |

**Schema markup:** Not officially mandated for citation — clear, well-structured prose can be grounded without it — but JSON-LD helps engines disambiguate entities, prices, authorship, and dates, which is exactly what grounding relies on. Treat it as the de facto structured-data standard across Bing, Google, Perplexity, and ChatGPT; add it where it removes ambiguity rather than chasing a specific citation-rate multiplier (no public, methodologically sound study quantifies the lift as of Jun 2026).

### 5. Meta directives — effect on Copilot

Bing's 2026 guidelines spell out per-directive AI behavior:

| Directive | Effect on Copilot / AI answers |
|---|---|
| `NOARCHIVE` | Prevents Copilot use entirely |
| `NOCACHE` | Limits Copilot to URLs, titles, and snippets |
| `DATA-NOSNIPPET` | May reduce citation quality |
| `NOINDEX` | Removes from both classic Search and AI surfaces |

Example — allow classic indexing but limit AI use:
```html
<meta name="robots" content="index, follow, noarchive">
```

### 6. Updated abuse policies (2026)

Bing expanded its abuse guidelines alongside GEO:

- **"Prompt Injection and AI Manipulation"** — new dedicated section. Attempts to interfere with Bing's or Copilot's language models will be demoted.
- **"Keyword Stuffing and Artificially Engineered Language"** — renamed and expanded. Covers content designed to trigger AI citations, not just classic rankings.
- **Scaled content** — language softened from "malicious" to: *"Large-scale content generated without oversight, quality control, or editorial review often lacks usefulness, accuracy, and originality, and may be excluded from indexing."* — aligns with Google's stance: targets intent, not automation itself.

### 7. URL Submission API

```bash
curl -X POST "https://ssl.bing.com/webmaster/api.svc/json/SubmitUrl?apikey=$BING_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"siteUrl":"https://example.com","url":"https://example.com/new-page"}'
```

**Daily quota:** 10,000 URLs/day for verified sites. Use for bulk migrations; prefer IndexNow for incremental updates.

### 8. Bing vs Google — what still differs

| Factor | Google | Bing |
|---|---|---|
| Social signals | Minimal | Moderate ranking factor |
| Exact-match domains | Discounted | Mildly rewarded |
| Multimedia weight | Moderate | Higher (images, video, transcripts) |
| Keyword in URL | Minor | Moderate |
| AI citation telemetry | Limited (Search Console partial) | **First-party AI Performance report** |
| GEO in guidelines | Implicit | **Explicitly named** |
| `llms.txt` stance | Not adopted for Search; says standard robots/meta govern Search surfaces | No official position |

> `llms.txt` adoption is moving fast in 2026 — recheck both engines' current stance before relying on this row. Google has stated it does not use `llms.txt` for Search crawling and that standard `robots.txt`/meta-robots controls govern Search surfaces (this is "not adopted," not a formal ban). Bing has published no official position. Treat `llms.txt` as low-cost optional housekeeping, not a ranking or citation lever.

### 9. Backlink analysis

Bing Webmaster provides free backlink data:
- Inbound links by domain
- Anchor text distribution
- Top linked pages
- New + lost links

**Audit checklist:**
- [ ] Check anchor text diversity (over-optimized exact-match anchors are a spam signal)
- [ ] Monitor new + lost links weekly
- [ ] Compare profile vs top 3 competitors

**Disavow — guarded, not routine.** Disavowing is a high-risk action that can suppress legitimate links if misapplied; modern engines already ignore most low-quality links automatically. Only disavow when you have a **manual action / link spam notice** in Bing Webmaster Tools or Google Search Console, or a clear, audited pattern of paid/spam/negative-SEO links you cannot get removed at the source. Before submitting: export your full link profile and the disavow list as a backup, prefer disavowing at the `domain:` level over individual URLs, and keep the file under version control so changes are reversible. Do not disavow as a precautionary or recurring task.

### 10. Reporting cadence

**Monthly Bing audit:**
- [ ] Crawl errors → fix
- [ ] Search performance (impressions, clicks, CTR)
- [ ] **AI Performance** — citation count + cited URLs
- [ ] Bing vs Google rankings for top 20 keywords
- [ ] IndexNow submission success rate
- [ ] Sitemap freshness

## Sources

URLs and policy wording shift fast in this area — confirm against the live page before citing. Retrieval dates below are when wording in this skill was last checked.

- Bing AI Performance announcement (public preview, Feb 2026): https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview (retrieved Jul 2026)
- AI Performance expansion, Intents/Topics/Citation Share/Compare (Jun 16, 2026): https://blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare (retrieved Jul 2026)
- Bing Webmaster Guidelines — GEO definition, per-directive AI behavior, and the updated abuse/scaled-content policies (the primary source for §4–§6, not third-party coverage): https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a (retrieved Jun 2026)
- IndexNow protocol + endpoints, response codes, and key-file rules: https://www.indexnow.org/documentation (retrieved Jun 2026)
- Bing URL Submission API reference: https://learn.microsoft.com/bingwebmaster/ (retrieved Jun 2026)

---

## blog-engine
Category: marketing
Description: End-to-end pipeline for one long-form, answer-engine-ready blog post — brief, primary-source research, intent mapping, outline, draft, on-page + AI-search optimization, JSON-LD, internal links, QA, and refresh, with 50+ headline formulas and 8 post-type templates. Use when researching, outlining, drafting, optimizing, or refreshing a blog post or SEO article.
Features:
  - Research and outline generation
  - 50+ headline formulas
  - Featured snippet optimization
  - SEO checklist and meta tag generation
  - Internal linking strategy
  - Blog post templates by type
Use Cases:
  - Write a complete blog post from topic to publish-ready
  - Optimize existing posts for featured snippets
  - Generate headline variants for A/B testing
  - Build a content production pipeline

# Blog Engine

Production pipeline for **one excellent long-form post**, from blank page to publish-ready and through its first refresh. This skill owns *execution of a single article*. It deliberately does **not** re-derive program strategy or engine internals — cross-link instead:

- **Topic clusters, editorial calendar, pillar/cluster model** → `content-strategy`
- **Template/directory/comparison pages generated at scale** → `programmatic-seo`
- **Engine-by-engine GEO stance, schema catalog, E-E-A-T, Core Web Vitals, hreflang** → `seo-geo`
- **Headline/CTA frameworks (PAS/AIDA/4U/BAB), voice calibration, before/after rewrites** → `copywriting`
- **Local intent (city/service pages, NAP, GBP)** → `local-seo`
- **Post-publish distribution sequences** → `email-sequence`, `social-media-kit`, `social-media-growth`

## 2026 ground rules (read first)

Search in mid-2026 is split between classic blue-link ranking and **answer engines** (Google AI Overviews & AI Mode, Bing Copilot, ChatGPT Search, Perplexity, Gemini, Claude). A post must work for both. Non-negotiables:

1. **Information gain over imitation.** Copying the SERP's structure/word count produces derivative pages that Google's helpful-content systems and AI engines both ignore. Every post must add something the top results don't have: original data, a first-hand test, a named expert quote, a calculator, a decision table, or a clearer synthesis. Aim to be *the* source an answer engine quotes, not the tenth paraphrase.
2. **First-party experience (the extra "E").** Show you actually did the thing: screenshots you took, numbers you measured, a methodology paragraph, a dated byline with a real author bio and credentials. This is what separates a quotable post from spun content.
3. **Length follows intent, not a target.** There is no minimum word count. A definition query deserves 600 focused words; a "best X for Y" comparison may need 3,000 with a table. Match the depth a satisfied reader needs and stop.
4. **Disclose AI assistance + keep an editorial gate.** AI-drafted copy must be fact-checked, edited, and reviewed by a named human before publish. Add a transparency note in your content policy (e.g., "Drafted with AI assistance, reviewed and edited by [author]") where your jurisdiction or audience expects it. Mass-produced, unreviewed AI pages are the exact pattern Google's scaled-content-abuse policy (March 2024, enforced through 2026) demotes — see `programmatic-seo` for the scale-safe variant.
5. **Citation hygiene.** Cite primary sources (the study, the docs, the filing — not a blog citing a blog). Record the publish/last-updated date of every source; drop or re-verify anything older than ~18 months for fast-moving topics. Never fabricate a statistic, quote, or study; if you can't verify it, cut it.

---

## Pipeline

### 0. Brief (define before you research)

Write these seven lines before touching a draft. They prevent scope creep and a post that ranks for nothing.

```
Primary keyword     : best crm for solo consultants
Search intent       : commercial-investigation (wants a shortlist + how to choose)
Audience + stage    : solo consultant, evaluating tools, low technical depth
Target query / JTBD : "which CRM should a one-person shop actually pay for?"
Information gain     : our own 30-day test of 6 tools + pricing table they can't find collated elsewhere
Primary CTA + path  : free trial of [product] (mid-article soft, end hard)
Author + credibility : [Name], ran a consulting practice 6 yrs (real bio + photo)
```

**Map the intent → format** (this replaces "look at the top 5 and copy them"):

| Intent | Query signals | Winning format | Primary CTA |
|---|---|---|---|
| Informational / definition | "what is", "meaning", "how does X work" | Concise answer-first explainer | Subscribe / related deep-dive |
| How-to / procedural | "how to", "steps", "tutorial" | Numbered steps + screenshots + pitfalls | Tool/template download |
| Commercial investigation | "best", "top", "vs", "alternatives", "review" | Comparison table + criteria + verdict | Trial / demo |
| Transactional | "buy", "pricing", "coupon", "near me" | Short, decision-oriented, fast path | Buy / contact |
| Navigational | brand + feature | Direct, brand-led | Login / product page |

### 1. Research (primary sources, not the SERP)

Goal: collect material that lets you *out-cover* the field, and verify every fact.

- **Define the entity set.** List the people, products, specs, standards, and sub-questions a complete answer must cover (the topic's "entities"). Coverage gaps here are why thin posts lose to thorough ones. For cluster-level entity planning see `content-strategy`.
- **Harvest real questions.** People Also Ask, `AlsoAsked`/`AnswerThePublic`-style tools, Reddit/forum threads, support tickets, sales-call objections, and YouTube comments. These become H2s and FAQ candidates.
- **Pull primary sources.** Original studies, official docs, regulatory filings, manufacturer specs, first-party analytics. Capture: source name, URL, **publish/updated date**, and the exact figure. Prefer the source over any blog summarizing it.
- **Generate first-party information gain.** Pick at least one: run a hands-on test, survey your list, pull anonymized data from your product, screenshot a real workflow, or interview an expert. This is the single highest-leverage step and the one competitors skip.
- **Read the SERP for *gaps*, not a template.** Skim the top results to find what's missing, outdated, or wrong — then fill that hole. Do **not** target their word count or mirror their headings; that's how you produce a forgettable near-duplicate.
- **Note the answer-engine angle.** Check whether an AI Overview/Copilot answer already appears for the query and what it cites. Aim to become a more citable source (clear claims, a stat with attribution, a definition block). Engine-specific tactics live in `seo-geo`.

**Fact-check gate before drafting:** every statistic has a primary source + date; every quote is attributed and real; nothing is older than your freshness threshold without re-verification; AI-suggested "facts" are independently confirmed.

### 2. Outline

Structure follows the intent format from step 0; this skeleton is the common case for an informational/how-to post. Use the per-type templates below for comparison, listicle, alternatives, case study, thought leadership, and product-led pages.

```
# {Headline — primary keyword + a specificity hook (number, year, outcome)}

> Author byline + publish/updated date + 1-line credibility ("ran X for Y years")

## Intro (80–150 words)
- Open with the reader's problem or a concrete promise (NOT a generic "in today's world")
- State the payoff: what they'll be able to do/decide by the end
- For definition/how-to intent only: include a tight 40–55 word answer block near the top
- Optional ToC for posts with 5+ H2s

## {H2 — first sub-question / step}        ← mirror real PAA / entity gaps
### {H3 if a step has detail}

## {H2 — second}
## {H2 — third}
## {H2 — Our test / data / methodology}     ← the information-gain section

## {H2 — Comparison / decision table}        ← if commercial intent

## FAQ  (optional; see schema note in §4 — schema ≠ visible rich result)
- 3–6 genuine residual questions not answered in the body

## Conclusion / Next step
- One-line synthesis (the takeaway, not a recap of every H2)
- Single primary CTA matched to intent
```

### 3. Draft

Write for a satisfied reader first, the crawler second. Rules that hold up in 2026:

- **Lead by intent, not by reflex.** Answer-first is right for *definition/how-to* queries (it earns the snippet and the AI citation). For comparison, investigative, narrative, or transactional posts a hard answer-first sentence reads robotically — open with the stakes, the surprising finding, or the scenario, and place the crisp answer where it belongs.
- **One idea per paragraph; vary length.** Mostly 2–4 sentences, but don't mechanize it — a one-line paragraph for emphasis and an occasional longer one for nuance read more human and dodge the spun-content feel.
- **Skip the formulaic filler.** Avoid manufactured "bucket brigades" ("Here's the thing:", "But wait — it gets better:") and AI throat-clearing ("In today's fast-paced world", "It's important to note that"). They pad word count, signal low-effort content, and don't help the reader. Earn engagement with specifics and a real point of view instead. (Genuine transitions are fine; canned hype lines are not.)
- **Be concrete and original.** Replace "studies show" with the named study + number; replace "many businesses" with a real example or your own data. Specificity is the whole game for both readers and answer engines.
- **Format for scannability and extraction.** Descriptive H2/H3s phrased like real questions; bulleted criteria; HTML `<table>` for any comparison; a 40–55 word definition block under a "What is X?" heading. Clean, self-contained chunks are what get lifted into AI Overviews and snippets.
- **Insert images where they explain, not on a timer.** Add a diagram/screenshot/chart wherever it carries information (every ~300–500 words is typical, not a rule). Each needs a real alt text and ideally is original (your screenshot > stock).
- **Readability ~grade 7–9** (Flesch-Kincaid ~60–70) for general audiences — but *don't* dumb down technical posts for technical readers. Match the audience in the brief.
- **Voice & deeper headline/CTA craft** → `copywriting`. **Repurposing one post into many formats** → `content-strategy`.

### 4. On-page & answer-engine optimization

Checklist — apply naturally, never stuff:

- [ ] Primary keyword in: title tag, H1, first ~100 words, URL slug, meta description (once each, no forcing)
- [ ] Secondary keywords + entities covered in H2s and body where they read naturally
- [ ] **Title tag** ≤ ~60 chars, front-loaded keyword + a hook (number/year/benefit)
- [ ] **Meta description** ~150–160 chars, includes keyword + a reason to click (it's a CTR lever, not a ranking factor)
- [ ] **URL slug**: short, lowercase, hyphenated, keyword, no dates/stop-words (`/best-crm-solo-consultants`)
- [ ] **Alt text** on every image: describes the image; keyword only if genuinely accurate
- [ ] **Internal links**: 3–6 to relevant posts/pillars with descriptive anchors (see §5)
- [ ] **External links**: 2–4 to primary/authoritative sources (studies, docs, official pages)
- [ ] **One H1 only**; logical H2→H3 nesting; no skipped levels
- [ ] **Author byline + bio + visible publish/updated date** (E-E-A-T signal)
- [ ] **Article/BlogPosting JSON-LD** present (see below)
- [ ] Quotable assets in place: a stat-with-source, a definition block, a decision table — the things answer engines cite

**Structured data — what actually does something in 2026:**

| Schema type | Use it for | Visible rich result today? |
|---|---|---|
| `Article` / `BlogPosting` | Every post — author, dates, publisher, image | Helps machine understanding; powers Top Stories/Discover eligibility |
| `BreadcrumbList` (items are `ListItem`) | Site hierarchy | Yes — breadcrumb trail in SERP |
| `HowTo` | Genuine step-by-step procedures | Largely **removed** from Google rich results (2023); still aids comprehension/AI — don't expect the visual widget |
| `FAQPage` | Real FAQ blocks | **Do not expect a rich result.** Google restricted FAQ rich results to authoritative gov/health sites in 2023 and has retired general FAQ visibility since. Keep `FAQPage` only as machine-readable context / possible AI-citation help when you have *genuine* Q&A — never add fake questions to chase a snippet. |
| `Product` / `Review` / `AggregateRating` | Product or review posts (only with real, verifiable ratings) | Yes — review stars (policy-gated; self-serving ratings are penalized) |

Full engine-by-engine schema catalog and validation workflow: `seo-geo`. Validate any JSON-LD in Google's Rich Results Test and `schema.org` validator before publish.

**Minimal `BlogPosting` JSON-LD** (fill the placeholders; dates in ISO 8601):

```json
{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "headline": "Best CRM for Solo Consultants (2026): 6 Tested",
  "description": "We tested six CRMs for 30 days. Here's the shortlist, pricing, and how to choose.",
  "image": "https://example.com/img/best-crm-solo-cover.png",
  "datePublished": "2026-06-07T09:00:00+00:00",
  "dateModified": "2026-06-07T09:00:00+00:00",
  "author": {
    "@type": "Person",
    "name": "Author Name",
    "url": "https://example.com/about/author-name",
    "jobTitle": "Independent consultant"
  },
  "publisher": {
    "@type": "Organization",
    "name": "Your Brand",
    "logo": { "@type": "ImageObject", "url": "https://example.com/logo.png" }
  },
  "mainEntityOfPage": { "@type": "WebPage", "@id": "https://example.com/best-crm-solo-consultants" }
}
```

**Optional `FAQPage` JSON-LD** — add *only* for real Q&A; treat as machine context, not a guaranteed widget:

```json
{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [{
    "@type": "Question",
    "name": "Do solo consultants need a CRM?",
    "acceptedAnswer": {
      "@type": "Answer",
      "text": "If you track more than ~20 active relationships or any repeat pipeline, yes — a lightweight CRM beats a spreadsheet for follow-up reminders and history."
    }
  }]
}
```

### 5. Internal-link plan

Internal links spread authority and define your topic structure — plan them, don't sprinkle randomly.

- **Up to the pillar:** link this post to the cluster's pillar page (and the pillar back down to it). Cluster architecture is owned by `content-strategy`.
- **Sideways to siblings:** 2–4 links to closely related posts using **descriptive anchor text** (the target's topic, e.g. "lightweight CRM pricing"), never "click here".
- **Down to conversion:** 1–2 links to the relevant product/landing/pricing page on the natural path to your CTA. Landing-page craft: `landing-page-builder`.
- **Backfill inbound links:** add links *from* existing high-authority posts *to* this new one — new posts have no internal equity until you do.
- **Audit anchors:** vary anchor text; avoid 10 posts all linking with the identical exact-match phrase (looks manipulative).

Quick plan template:

```
This post → pillar:        /guides/crm  (anchor: "complete CRM guide")
This post → sibling:       /crm-vs-spreadsheet  (anchor: "CRM vs a spreadsheet")
This post → conversion:    /pricing  (anchor: "[product] pricing")
Inbound (existing → this): /freelancer-tools, /consulting-ops  add contextual links
```

### 6. Image brief

Give a designer/generator everything in one block; originality beats stock for trust and reuse:

```
Cover (1200×630, also OG image):
  - Concept: 6 CRM logos on a comparison grid, brand colors, post title overlaid
  - Format: WebP, < 150 KB, alt: "Comparison grid of six CRM tools tested for solo consultants"
In-body:
  1. Screenshot — your actual pipeline view in [tool] (original, not stock)
  2. Chart — 30-day response-time results from our test (label axes + source)
  3. Decision diagram — "Which CRM should you pick?" flow
Specs: WebP/AVIF, lazy-load below the fold, explicit width/height (CLS), descriptive alt on each.
```

### 7. QA & publish checklist

- [ ] **Facts re-verified** against primary sources; dates current; no fabricated stats/quotes
- [ ] **Human editorial review** done by a named editor; AI assistance disclosed per policy
- [ ] Reads well aloud; no filler/hype lines; intro delivers on the title (no clickbait gap)
- [ ] Grammar/spelling clean; consistent terminology and capitalization
- [ ] One H1; heading hierarchy valid; ToC anchors resolve
- [ ] All links work (no 404s); external links open authoritative, live sources
- [ ] Images compressed (WebP/AVIF), sized, lazy-loaded, all with alt text; cover doubles as OG
- [ ] JSON-LD (`BlogPosting` + breadcrumbs; `FAQPage` only if real) validates in Rich Results Test
- [ ] **Open Graph + Twitter Card**: `og:title`, `og:description`, `og:image` (1200×630), `og:type=article`, `twitter:card=summary_large_image`
- [ ] Title tag ≤ ~60 chars, meta description ~150–160 chars, slug clean
- [ ] Canonical tag set; not accidentally `noindex`; in the sitemap (verify in `search-console`)
- [ ] Mobile render + Core Web Vitals sane (LCP/INP/CLS) — see `web-performance`, `seo-geo`
- [ ] Internal-link plan executed both directions
- [ ] Distribution queued: email (`email-sequence`), social (`social-media-kit`)

### 8. Refresh (the post-publish loop most teams skip)

A blog post is an asset to maintain, not ship-and-forget. Schedule reviews; refreshed content often outperforms net-new for the same effort.

- **Cadence:** evergreen posts every 6–12 months; fast-moving topics (pricing, tools, regulations, "best of 2026") quarterly.
- **Triggers to refresh now:** ranking/clicks sliding (check `search-console`), facts/stats/screenshots gone stale, a year in the title rolling over, a competitor now out-covering you, or a new AI Overview answering the query (rework to become the cited source).
- **What to do:** update stats + dates to current primaries, add new sub-questions/entities, re-shoot stale screenshots, prune dead/outdated sections, strengthen the information-gain asset, fix the internal-link graph, then bump `dateModified` and re-request indexing.
- **Decide:** update in place when intent is unchanged (keeps the URL's history); split into a new post when you're really targeting a different intent; consolidate/redirect thin overlapping posts into one strong page.

---

## Headline formulas (50+)

Pick a frame that matches intent, then make it **specific** — add a number, a year, a named outcome, or a constraint. Promise only what the post delivers; a clickbait gap between title and intro tanks dwell time. Deeper headline frameworks (PAS/AIDA/4U/BAB) and split-testing live in `copywriting`. `{}` = fill in.

**How-to / instructional**
1. How to {achieve outcome} in {timeframe}
2. How to {achieve outcome} (Even If {common obstacle})
3. How to {do task} the Right Way: {N} Steps
4. The Complete Guide to {topic} for {audience}
5. {Task}: A Step-by-Step Guide for {year}
6. How I {achieved specific result} — and How You Can Too
7. The Beginner's Guide to {topic}
8. How to {outcome} Without {pain/cost/tool}
9. {N} Steps to {outcome} (With Examples)
10. The Lazy Person's Guide to {outcome}

**Listicle / number**
11. {N} {tools/tips/ways} to {achieve outcome}
12. {N} {category} Every {audience} Should {action} in {year}
13. {N} {mistakes} That Are {negative consequence}
14. {N} Surprising {facts/stats} About {topic}
15. {N} Best {products} for {use case}, Tested
16. Top {N} {category}: Ranked for {audience}
17. {N} {things} You're Doing Wrong (and How to Fix Them)
18. {N} Underrated {tools/tactics} for {outcome}
19. {N} Examples of {thing} Done Right
20. {N}-Minute {task}: {N} Quick Wins

**Comparison / alternatives**
21. {Product A} vs {Product B}: Which Is Better for {use case}?
22. {Product A} vs {Product B} vs {Product C}: An Honest Comparison
23. The {N} Best {Product} Alternatives in {year}
24. {Expensive option} Too Pricey? {N} Cheaper Alternatives
25. Is {product/approach} Worth It? An Honest {year} Review
26. {Approach A} or {Approach B}: How to Choose
27. We Tested {N} {products} — Here's the Winner
28. {Product}: Pros, Cons, and Who It's Actually For

**Question**
29. What Is {term}? (And Why It Matters for {audience})
30. Why Does {phenomenon} Happen — and What to Do About It
31. Should You {action}? Here's How to Decide
32. Can You Really {desirable outcome}? We Checked
33. What's the Best Way to {achieve outcome}?
34. Is {common belief} Actually True?

**Negative / mistake / warning**
35. Stop {doing common thing} — Do This Instead
36. The {N} {topic} Mistakes Costing You {money/time}
37. Why Your {effort} Isn't Working (and the Fix)
38. {N} Myths About {topic}, Debunked
39. The Hidden Cost of {common choice}
40. Avoid These {N} {topic} Pitfalls

**Curiosity / contrarian / data**
41. The Surprising Truth About {topic}
42. What Nobody Tells You About {topic}
43. We Analyzed {N} {things}. Here's What We Found
44. {Counterintuitive claim}: The Data Says {finding}
45. I {did unusual thing} for {timeframe}. Here's What Happened
46. The {topic} Trend Everyone's Ignoring in {year}

**Outcome / benefit-led**
47. {Achieve outcome} in {timeframe} — Without {sacrifice}
48. The {adjective} Way to {achieve outcome}
49. Double Your {metric} With This {tactic/framework}
50. From {bad state} to {good state}: A {timeframe} Playbook

**Thought-leadership / opinion**
51. Why {prediction} Will Change {industry} by {year}
52. The Case for (and Against) {approach}
53. {Industry} Has a {problem} Problem. Here's the Fix
54. Unpopular Opinion: {contrarian take}

---

## Post-type templates

Eight reusable skeletons. Each maps to an intent from §0; swap in your brief's keyword, entities, and information-gain asset. Keep one H1, descriptive H2s, and the schema noted.

### A. How-to / tutorial — *(intent: procedural)*
```
# How to {outcome} in {timeframe}: A Step-by-Step Guide
Byline + date · 40–55 word answer block (what they'll achieve + rough time/cost)
## What you'll need / Prerequisites
## Step 1: {action}        ← screenshot + the "why"
## Step 2: {action}
## Step N: {action}
## Common mistakes / Troubleshooting     ← your first-hand pitfalls = info gain
## FAQ (real residual questions)
## Next step  → tool/template CTA
Schema: BlogPosting (+ HowTo for comprehension; expect no visual widget)
```

### B. Listicle — *(intent: informational/commercial)*
```
# {N} {items} to {outcome} in {year}
Intro: who this list is for + how you chose (selection criteria = trust signal)
## 1. {Item}  — what it is · best for · 1 pro · 1 con · price
## 2. {Item}
## … N
## How to choose the right one for you      ← decision guidance, not just a list
## FAQ
Schema: BlogPosting · use ordered list markup
```

### C. Comparison ("X vs Y") — *(intent: commercial-investigation)*
```
# {Product A} vs {Product B}: Which Is Better for {use case}? ({year})
Verdict up top: 2–3 sentences naming the winner *for each use case*
## Comparison at a glance       ← HTML table: price, key features, best-for, free tier
## {Product A}: strengths & weaknesses
## {Product B}: strengths & weaknesses
## Head-to-head: {dimension 1}, {dimension 2}, {pricing}
## Our test / methodology       ← information gain: how you evaluated
## Which should you choose?     ← map persona → pick
## FAQ
Schema: BlogPosting (+ Review/AggregateRating ONLY with real, verifiable ratings)
```

### D. Alternatives ("best X alternatives") — *(intent: commercial-investigation)*
```
# The {N} Best {Product} Alternatives in {year}
Intro: why someone leaves {Product} (price, missing feature, lock-in) + how you picked
## Quick comparison table       ← alternative · best for · price · key differentiator
## 1. {Alternative} — who it's for · vs {Product} · pricing · catch
## … N
## How to migrate from {Product}        ← practical info gain
## FAQ
Schema: BlogPosting · ordered list
```

### E. Ultimate guide / pillar — *(intent: informational, broad)*
```
# The Complete Guide to {topic} ({year})
ToC (this is long) · who it's for · what you'll learn
## What is {topic}?             ← 40–55 word definition block
## Why {topic} matters
## {Core subtopic 1}  → links to a dedicated cluster post
## {Core subtopic 2}  → cluster post
## {Core subtopic 3}  → cluster post
## Common mistakes / best practices
## Tools & resources
## FAQ
Schema: BlogPosting · this is the hub — link out to and back from cluster posts (content-strategy)
```

### F. Case study / results — *(intent: informational + proof, BOFU)*
```
# How {subject} {achieved result} in {timeframe}
Result up front: the headline number + context (this IS the information gain)
## Background / starting point  ← baseline metrics
## The challenge
## What we did                  ← specific, replicable steps
## Results                      ← before/after table or chart, real numbers
## What we'd do differently
## How to apply this to your {situation}
## CTA → relevant product/service
Schema: BlogPosting
```

### G. Thought leadership / opinion — *(intent: informational + brand authority)*
```
# {Contrarian or forward-looking thesis}
Stake your claim in the first paragraph — say something a base model wouldn't
## The conventional wisdom (and why it's incomplete)
## My/our argument             ← backed by data, experience, or first-hand examples
## Counterpoints & honest limits   ← engaging with objections builds credibility
## What this means for {audience}
## Conclusion: the takeaway
Schema: BlogPosting · author bio + credentials are essential here
```

### H. Product-led / "jobs-to-be-done" SEO — *(intent: informational → product)*
```
# How to {solve problem the product solves} (with or without {product category})
Teach the solution genuinely first — earn trust before the pitch
## Understanding {the problem}
## Method 1: {manual / free approach}      ← real, usable — don't gatekeep
## Method 2: {using a tool like ours}      ← natural, honest product fit
## Comparison: when each method makes sense
## FAQ
## Get started  → product CTA
Schema: BlogPosting · keep teaching:selling ratio high; for scaled JTBD pages see programmatic-seo
```

---

## Anti-patterns (do not ship)

- Mirroring the SERP's word count/headings → derivative, demoted, never cited.
- Fabricated or unverifiable stats/quotes; citing a blog that cites the real source instead of the source.
- Fake FAQ blocks added only to chase a (now mostly gone) rich result.
- Manufactured bucket brigades, hype lines, and AI throat-clearing as filler.
- A rigid answer-first sentence forced onto narrative/comparison/transactional posts.
- Unreviewed, undisclosed mass-produced AI pages (scaled-content-abuse risk — `programmatic-seo`).
- Exact-match anchor text on every internal link; "click here" anchors.
- Keyword stuffing in title/meta/alt; a clickbait title the intro doesn't pay off.
- Publishing once and never refreshing.

---

## brand-strategy
Category: marketing
Description: Frameworks and templates for building and managing a cohesive brand system. Use when defining or auditing brand positioning, messaging hierarchy, voice/tone, visual identity, brand architecture, naming, or brand guidelines for a product or company.
Features:
  - Brand positioning framework (fill-in-the-blank)
  - Messaging hierarchy (tagline to proof points)
  - Voice and tone spectrum guide
  - Visual identity system design
  - Competitive positioning map
  - Brand guidelines document structure
Use Cases:
  - Define brand positioning for a new product
  - Create a brand voice and tone guide
  - Design a visual identity system
  - Build a complete brand guidelines document

# Brand Strategy

## Brand Positioning Framework

Complete this statement — if you can't, your positioning isn't clear enough:

```
For [TARGET AUDIENCE] who [NEED/SITUATION],
[BRAND] is the [CATEGORY]
that [KEY DIFFERENTIATOR]
because [REASON TO BELIEVE].
```

**Example:**
> For growth-stage SaaS teams who need to ship marketing pages fast,
> Webflow is the visual development platform
> that gives designers production-level control without engineering dependencies
> because it generates clean, production-ready code with built-in CMS and hosting.

### Positioning Inputs Checklist

- [ ] Target audience defined with specificity (not "everyone")
- [ ] Category clearly named (or intentionally created)
- [ ] 1-2 differentiators that are true, relevant, AND defensible
- [ ] Proof points for each differentiator (data, patents, methodology)
- [ ] Competitive alternatives identified (including "do nothing")

### Positioning Worksheet (run in order)

**Step 1 — ICP segmentation.** Don't position for "the market." Pick one beachhead segment and describe it precisely. Score candidate segments and start with the highest total.

| Segment | Urgency of pain (1-5) | Budget/willingness to pay (1-5) | Reachability (1-5) | Strategic value / expansion (1-5) | Total |
|---------|----|----|----|----|----|
| e.g. Seed-stage B2B SaaS founders | 5 | 3 | 4 | 4 | 16 |
| e.g. Enterprise marketing ops | 3 | 5 | 2 | 5 | 15 |

For the winning segment, write a one-paragraph ICP: firmographics (size, industry, stage), the buyer's role and the user's role (often different), the trigger event that starts the search, and the budget they already spend on the problem.

**Step 2 — Competitive alternatives.** Frame against what the customer would actually do instead, not just direct competitors. Four buckets:
- **Direct** — does the same job a similar way (e.g., another design tool).
- **Indirect** — does the job a different way (e.g., hiring a contractor).
- **Status quo / "do nothing"** — spreadsheet, manual process, or living with the pain. This is usually the #1 competitor — name it explicitly.
- **In-house build** — for technical buyers, "we'll build it ourselves."

**Step 3 — Category strategy.** Decide one of three plays and commit:
- **Join an existing category** — fastest; you compete on differentiation inside a known frame. Use when buyers already budget for the category.
- **Subdivide a category** — claim a niche ("the X for Y", e.g., "the CRM for solo realtors"). Use when the broad category is crowded but a segment is underserved.
- **Create a new category** — most expensive (you fund the education), highest ceiling. Only attempt with strong funding and a genuinely new mechanism; most "new categories" should have been a subdivision.

**Step 4 — Differentiator stress test.** List each claimed differentiator and test it. Cut any that fail the first three columns; the survivors are your positioning.

| Differentiator | True? (provable today) | Relevant? (buyer cares) | Defensible? (hard to copy in 12 mo) | Verdict |
|----------------|------------------------|-------------------------|--------------------------------------|---------|
| Generates clean production code | Yes | Yes (devs inherit it) | Partly (architectural moat) | Keep |
| "Easy to use" | Subjective | Yes | No (everyone claims it) | Cut — too generic |
| 50+ integrations | Yes | For some segments | No (table stakes soon) | Demote to proof point |

Defensibility sources to look for: proprietary data/network effects, switching costs, a patented or genuinely novel mechanism, brand/community, or unit-economics advantage. "We work harder" and "better UX" are not durable moats on their own.

**Step 5 — Reasons to believe (RTB) evidence table.** Every differentiator needs ranked proof. Lead with the strongest verifiable evidence.

| Differentiator | RTB (proof) | Type | Strength |
|----------------|-------------|------|----------|
| Ship 10x faster | Independent benchmark: median deploy 14 min → 90 sec across 40 teams | 3rd-party data | Strong |
| Designer control | Patent US-XXXXXXX on the visual-binding compiler | IP | Strong |
| Reliability | 99.99% uptime SLA, public status page | Operational | Medium |
| Trust | "Used by Fortune-500 brand X (logo, with permission)" | Social proof | Medium |

**Step 6 — Final positioning variants.** Write 2-3 versions of the positioning statement (one safe/category-joining, one bolder/subdividing), then pressure-test each with 5-10 real prospects: read it aloud, ask "what does this company do, and who is it for?" Keep the version they can play back accurately and that makes the target lean in. Revisit positioning when the segment, competitive set, or product fundamentally changes — not on a calendar.

## Messaging Hierarchy

```
Tagline (5-8 words)
├── Value Proposition 1
│   ├── Proof Point 1a
│   └── Proof Point 1b
├── Value Proposition 2
│   ├── Proof Point 2a
│   └── Proof Point 2b
└── Value Proposition 3
    ├── Proof Point 3a
    └── Proof Point 3b
```

| Level | Purpose | Example |
|-----------------|-------------------------------|--------------------------------------|
| Tagline | Memorable, emotional hook | "Think Different" |
| Value props | Rational benefits (3 max) | "Ship 10x faster" |
| Proof points | Evidence for each value prop | "Used by 200K+ teams at Fortune 500" |
| RTBs | Why you can deliver | Patent, methodology, team expertise |

**Rules:**
- Taglines usually lead with emotion; value props usually lead with rational benefit. The best taglines fuse both — "Just Do It" is emotional but implies a functional promise, "The Ultimate Driving Machine" is rational framed emotionally. Pick a primary register per element so the message stays sharp, but don't force a wall between them.
- 3 value propositions maximum — more dilutes the message
- Every proof point must be verifiable
- Test messaging with real prospects, not your team

### Messaging Matrix (audience × message mapping)

Build one row per priority segment. This is the bridge from positioning to actual copy — it forces a different proof point and CTA per audience instead of one generic pitch.

| Segment | Job-to-be-done | Top objection | Proof to counter it | Primary CTA | Channel | Sample headline |
|---------|----------------|---------------|---------------------|-------------|---------|-----------------|
| Seed-stage founder | "Launch a credible site without hiring a dev" | "I'll outgrow a no-code tool" | Exports clean code; 200K+ teams scaled on it | Start free | Search/PLG | "Ship your launch site this weekend — no engineer required" |
| Marketing lead (Series B) | "Update pages without filing a Jira ticket" | "Will IT/eng approve it?" | SOC 2; role-based publishing; clean code review | Book a demo | Paid + outbound | "Your team ships pages. Engineering keeps its sprint." |
| Agency / freelancer | "Deliver client sites faster, bill more" | "Client lock-in / handoff" | White-label, client billing, code export | See partner program | Community/referral | "Build, bill, and hand off in one platform" |
| Enterprise procurement | "De-risk the buy" | "Security, SLA, compliance" | 99.99% SLA, SSO/SAML, DPA, EU data residency | Contact sales | ABM/sales | "Enterprise-grade governance for the web team" |

How to use it:
- **Job-to-be-done** is the customer's words, not your feature name. Pull these verbatim from interviews and support tickets.
- **Objection → proof** is the most valuable column: it pre-empts the reason each segment says no. If you have no proof for a real objection, that's a product/roadmap gap, not a copy gap.
- One **primary CTA** per segment per touchpoint — competing CTAs reduce conversion.
- Map message **stage** too if you sell over time: awareness (lead with the problem/JTBD), consideration (lead with differentiator + objection-handling proof), decision (lead with risk reduction — SLA, guarantee, references).

## Brand Voice & Tone Guide

**Voice** = personality (constant). **Tone** = mood (varies by context).

### Voice Definition Template

Define your voice on 4 spectrums:

| Spectrum | Our Position | Example |
|----------------------|--------------------------|-------------------------------|
| Formal ↔ Casual | Casual but competent | "Here's the deal" not "Hereby" |
| Serious ↔ Playful | Mostly serious, wit OK | Humor in social, not in legal |
| Technical ↔ Simple | Simple with depth option | Lead simple, link to deep dives |
| Bold ↔ Humble | Confident, not arrogant | "We built X" not "We're the best" |

### Tone by Context

| Context | Tone Shift | Example |
|------------------|----------------------------|---------------------------------|
| Marketing site | Confident, aspirational | "Build something remarkable" |
| Error messages | Helpful, calm | "Something went wrong. Here's what to try." |
| Social media | Conversational, human | "Okay this feature is *chef's kiss*" |
| Legal/compliance | Clear, neutral | "Your data is stored in the EU" |
| Crisis comms | Direct, empathetic | "We messed up. Here's what happened." |

### Full Voice & Tone Framework

**1. Three voice pillars.** Distill the brand into 3 adjectives, each made operational with a "this, not that" pair. Adjectives alone are useless to a writer; the contrast is what makes them usable.

| Pillar | We are… | We are not… | In practice |
|--------|---------|-------------|-------------|
| Direct | Plain-spoken, gets to the point | Curt or cold | "This costs $20/mo." not "Pricing varies based on a number of factors." |
| Encouraging | Optimistic, on the reader's side | Hype-y or condescending | "You've got this — here's step one." not "It's super easy!!" |
| Expert | Precise, evidence-led | Jargon-heavy or arrogant | "Latency dropped 60% in our tests." not "Blazing-fast, period." |

**2. Lexicon — words we use / words we avoid.** A shared word list keeps a 20-person team sounding like one brand.

| Use | Avoid | Why |
|-----|-------|-----|
| "people", "teams", "you" | "users", "consumers" | Humanize; speak to the reader |
| "help", "free up" | "leverage", "utilize", "synergy" | Plain English, no corporate filler |
| "we got this wrong" | "mistakes were made" | Own it; passive voice dodges accountability |
| product names exactly as styled | ad-hoc capitalization | Consistency builds recognition |

**3. Mechanics & house style.** Lock the small decisions so they aren't relitigated: capitalization (sentence case vs. title case for headings), Oxford comma (yes/no), contractions (yes for warmth), em dash vs. parenthesis, numerals ("spell out one-nine" vs. "always digits"), emoji policy by channel, and how you write dates/times/currency. Adopt a base reference (e.g., a major brand or AP/Chicago) and document only your deviations.

**4. Readability targets.** Set a reading-grade ceiling per surface (marketing/help ~grade 7-9; legal as required). Prefer short sentences, active voice, second person. Test copy against the targets before publishing.

**5. Accessibility & inclusivity in language.** Use plain language; expand acronyms on first use; write descriptive link text ("read the pricing guide", never "click here"); default to gender-neutral and people-first phrasing; avoid idioms that don't translate for a global/ESL audience. This makes copy work for screen readers and non-native speakers alike.

**6. Worked before/after example.**
- ❌ "Our best-in-class, enterprise-grade solution leverages cutting-edge AI to synergistically optimize your workflows."
- ✅ "Our AI drafts your first reply in seconds, so your team spends time on the hard tickets — not the easy ones."

Ship the voice guide with 5-8 such real before/after rewrites; writers copy patterns far faster than they internalize adjectives.

## Visual Identity System

| Element | Specification | Deliverable |
|---------------|--------------------------------------|-------------------------------|
| Logo | Primary, secondary, icon, monochrome | SVG + PNG at standard sizes |
| Color palette | Primary, secondary, neutral, semantic | Hex, RGB, HSL, CMYK values |
| Typography | Headings, body, mono, display | Font files + usage rules |
| Imagery | Photography style, illustration style | Mood board + do/don't examples |
| Iconography | Style, stroke weight, grid | Icon library + creation rules |
| Spacing/grid | Base unit, layout grid | Design tokens or spec sheet |

**Color palette structure:**
- Primary: 1-2 brand colors (used for CTAs, key elements)
- Secondary: 2-3 supporting colors
- Neutrals: 4-5 grays from near-white to near-black
- Semantic: Success, warning, error, info

### Visual Identity Audit & Accessibility Checklist

Run this when building a new system or auditing an existing one. Accessibility is not optional polish in 2026 — WCAG 2.2 AA is the de facto baseline and is referenced by the EU Accessibility Act (in force June 28, 2025) and US ADA/Section 508 expectations.

**Color & contrast (WCAG 2.2 AA):**
- [ ] Body text contrast ≥ **4.5:1** against its background.
- [ ] Large text (≥ 24px, or ≥ 18.7px bold) and UI/graphical components/focus indicators contrast ≥ **3:1**.
- [ ] Information is never conveyed by color alone (error states also use icon/text; chart series use labels/patterns).
- [ ] Brand color usable for CTAs at AA against white **and** the dark surface — if not, define an accessible "action" tint distinct from the marketing brand color.
- [ ] Verify combinations with a contrast checker (e.g., WebAIM, or the contrast lint in your design tool); document pass/fail per pairing.

**Dark mode:**
- [ ] Dedicated dark palette (don't just invert — pure-black #000 + pure-white #fff causes halation; use near-black ~#0E0F12 and off-white text).
- [ ] Elevation shown via lighter surfaces, not just shadows (shadows are weak on dark).
- [ ] Brand and semantic colors re-tuned for dark backgrounds and re-checked for AA contrast.

**Motion & animation:**
- [ ] Honor `prefers-reduced-motion`; provide a non-animated path for essential content.
- [ ] No content flashes more than **3 times per second** (seizure safety, WCAG 2.3.1).
- [ ] Animation is purposeful (feedback/continuity), short, and never blocks interaction.

**Logo system:**
- [ ] Minimum sizes specified: digital ~24px height for the icon/favicon, ~120px width for the full logo; print ~10mm height (set real numbers per logo and test legibility).
- [ ] Clear space defined as a ratio of a logo element (e.g., "= height of the wordmark's cap height").
- [ ] Variants for light bg, dark bg, monochrome, and a single-color knockout.
- [ ] Misuse examples documented (don't stretch, recolor, add effects, place on busy imagery, or rotate).

**Typography & layout:**
- [ ] Type scale defined (e.g., modular scale 1.250) with min body size ~16px on web.
- [ ] Line length ~45-75 characters; line-height ~1.5 for body.
- [ ] Webfont loading strategy set (`font-display: swap`, subset, preload) to avoid layout shift.

**Design tokens (naming conventions):**
Use a 3-tier token architecture so brand changes propagate without touching components:
- **Primitive / global** — raw values: `color.blue.500 = #2563EB`, `space.4 = 16px`.
- **Semantic / alias** — intent: `color.text.primary`, `color.bg.surface`, `color.action.default`, `color.feedback.error`.
- **Component** — scoped: `button.primary.bg`, `card.border.color`.
Name by role, never by appearance (`color.action.default`, not `color.green`) so re-skinning doesn't create lies. Define both `light` and `dark` themes at the semantic tier. Export to platforms via a tool like Style Dictionary or the W3C Design Tokens format so design and code stay in sync.

**Asset deliverables:**
- [ ] Logo: SVG (primary) + PNG @1x/@2x/@3x, favicon set, social avatars and OG/share images at correct dimensions.
- [ ] Color: hex, RGB, HSL, and CMYK + Pantone for print.
- [ ] Type: licensed webfont + desktop files, and named fallback stacks.
- [ ] Tokens published as JSON; icon set as an SVG sprite/library.

## Brand Audit Methodology

**Run annually or before major repositioning.**

1. **Internal audit:** Survey employees on brand perception, review all touchpoints
2. **External audit:** Customer interviews (10-15), prospect surveys, social listening
3. **Competitive audit:** Map competitors on key perception dimensions
4. **Touchpoint inventory:** List every place the brand appears, score consistency
5. **Gap analysis:** Internal perception vs external perception vs desired perception

### Touchpoint Inventory + Consistency Scoring

List every place the brand appears and score each 1-5 on visual, verbal, and experience consistency against the guidelines. Sort by `priority × gap` to find the highest-leverage fixes.

| Touchpoint | Owner | Reach/priority (1-5) | Visual (1-5) | Verbal/voice (1-5) | Experience (1-5) | Avg | Notes / gap |
|------------|-------|---------------------|--------------|--------------------|--------------------|-----|-------------|
| Homepage | Marketing | 5 | 4 | 3 | 4 | 3.7 | Voice drifts formal vs. guide |
| Onboarding emails | Lifecycle | 4 | 2 | 3 | 3 | 2.7 | Old logo, off-palette |
| Sales deck | Sales | 4 | 2 | 2 | — | 2.0 | Rebuilt ad-hoc per rep |
| Support macros | Support | 3 | — | 2 | 4 | 3.0 | Tone too robotic |
| App empty states | Product | 3 | 4 | 2 | 3 | 3.0 | No voice applied |
| Social profiles | Marketing | 4 | 4 | 4 | — | 4.0 | On-brand |

**Brand health score** = average of all touchpoint averages, weighted by priority. Track it over time; a single number makes drift visible to leadership.

### Customer Interview Script (30 min, 8-12 people)

Mix current customers, churned customers, and prospects who chose a competitor. Record verbatims — exact words become messaging copy.
1. Walk me through the last time you needed [the job]. What did you do? *(uncovers real JTBD and the status-quo alternative)*
2. What were you using before us, and why did you switch — or why haven't you? *(switching triggers and friction)*
3. In your own words, what do we do? Who is it for? *(positioning clarity check)*
4. If we disappeared tomorrow, what would you use instead, and what would you miss? *(differentiation and stickiness)*
5. Describe our brand as a person — how would you introduce them at a party? *(personality / voice perception)*
6. When did we frustrate or surprise you? *(experience gaps)*
7. Who else should hear about us, and how would you describe us to them? *(referral language)*

### Competitive Perception Axes

Score yourself and 3-5 competitors on the attributes your audience actually buys on (pull these from interview frequency, not internal opinion), 1-5 each. Then pick the two highest-variance, highest-importance attributes as the axes for the positioning map below.

| Attribute | Us | Comp A | Comp B | Comp C | Importance to buyer (1-5) |
|-----------|----|--------|--------|--------|----------------------------|
| Easy to adopt | 4 | 2 | 3 | 5 | 5 |
| Powerful/extensible | 3 | 5 | 4 | 2 | 4 |
| Trustworthy/secure | 4 | 4 | 5 | 3 | 5 |
| Value for money | 5 | 3 | 2 | 4 | 4 |

Look for a high-importance attribute where you outscore everyone and rivals are clustered — that whitespace is your positioning wedge.

### Competitive Positioning Map

Plot brands on a 2×2 matrix using the two dimensions that matter most to your audience:

```
        High Price
            │
  Premium   │   Luxury
  Niche     │   Established
            │
Low ────────┼──────── High
Innovation  │         Trust
            │
  Disruptor │   Value
  Challenger│   Incumbent
            │
        Low Price
```

Pick axes that reveal whitespace. Common pairs: price/quality, innovation/trust, simple/powerful.

## Brand Architecture

| Model | Structure | Example | Best When |
|------------------|-----------------------------|-----------------|-------------------------------|
| Branded house | Master brand drives all | Google, Virgin | Strong parent, related offerings |
| House of brands | Independent brands | P&G, Unilever | Diverse categories, M&A strategy |
| Endorsed | Sub-brands + parent endorsement | Marriott Bonvoy, Courtyard by Marriott | Credibility transfer needed |
| Hybrid | Mix of above | Amazon (AWS, Alexa, Whole Foods) | Large portfolio, some overlap |

**Decision criteria:**
- How related are the offerings? → Related = branded house
- Does the parent brand help or hurt? → Helps = endorsement
- Different audiences entirely? → House of brands
- Need to acquire and keep separate? → House of brands

## Naming Strategy

**Name types:**

| Type | Example | Pros | Cons |
|--------------|-------------|---------------------|--------------------------|
| Descriptive | General Motors | Instant clarity | Hard to trademark, boring |
| Invented | Spotify | Highly ownable | Requires education spend |
| Metaphor | Amazon | Evocative, memorable | Can feel random |
| Acronym | IBM | Short, professional | Meaningless until established |
| Founder | Goldman Sachs | Heritage, trust | Succession risk |

**Naming checklist (2026 realities):**
- [ ] **Domain:** exact-match `.com` is ideal but increasingly scarce; a strong brandable `.com` with a modifier (`getX.com`, `Xhq.com`, `tryX.com`) or a well-established alt TLD (`.ai`, `.io`, `.co`, `.app`) is acceptable. Buy obvious **defensive domains** (common misspellings, the `.com` if you launch on an alt TLD, and your country TLD).
- [ ] **Trademark:** clearance search in every target jurisdiction (USPTO, EUIPO, UKIPO, WIPO Global Brand DB) within the **Nice classification classes** you'll actually operate in — a mark can be free in your class but taken in an adjacent one that matters. Distinctive/invented marks register far more easily than descriptive ones. Budget for a trademark attorney before you commit spend; a clear search is not legal clearance. *(This is not legal advice — verify with a qualified IP attorney.)*
- [ ] **Marketplace / app-store conflicts:** check the Apple App Store, Google Play, GitHub, npm/PyPI, and Chrome/extension stores — these have their own naming uniqueness rules and reject confusingly similar names regardless of trademark status.
- [ ] **AI / search disambiguation:** Google the exact name and ask a couple of LLMs "what is [name]?" — if it collides with a famous brand, a common word, or another startup, you'll fight for SERP and AI-answer real estate forever. Prefer names that return *you* on the first page within months.
- [ ] **Social handles:** exact handle available (or acquirable) on the platforms you'll use; reserve them immediately even before launch.
- [ ] No negative or unintended meanings/slang in your key markets' languages.
- [ ] Pronounceable and spellable by the target audience — passes the "phone test" (say it aloud; can they spell it back?).
- [ ] Not tied to a single feature or geography you may outgrow.

## Brand Story Framework

```
1. ORIGIN:    Why we started (the problem we couldn't ignore)
2. MISSION:   What we do and for whom (present tense)
3. VISION:    The world we're building toward (future tense)
4. VALUES:    How we operate (3-5, actionable not generic)
5. PROOF:     Evidence we're living this (metrics, stories, milestones)
```

**Values anti-patterns:** "Innovation," "Integrity," "Excellence" — if every company claims it, it's not a differentiator. Make values specific and behavioral: "Ship before it's comfortable" > "Innovation."

## Brand Guidelines Document Structure

```
1. Brand Overview (positioning, story, values)
2. Logo Usage (versions, spacing, minimum size, misuse examples)
3. Color System (palettes, accessibility ratios, usage rules)
4. Typography (typefaces, hierarchy, sizing scale)
5. Imagery & Illustration (style, dos and don'ts)
6. Voice & Tone (guide + examples by context)
7. Layout & Grid (spacing system, templates)
8. Digital Applications (web, email, social templates)
9. Print Applications (business cards, signage, swag)
10. Co-branding Rules (partner lockups, minimum requirements)
```

### Starter Brand Guidelines Template

Copy this skeleton and fill the bracketed values. Aim for *prescriptive* (a writer/designer can execute without asking) over *descriptive*. Ship it as a living doc (web page or Figma) with a version number and changelog, not a frozen PDF.

```markdown
# [Brand] Brand Guidelines — v[1.0] · last updated [YYYY-MM-DD]

## 1. Foundation
- Positioning statement: For [audience] who [need], [brand] is the [category] that [differentiator] because [RTB].
- Mission (present): [...]   Vision (future): [...]
- Values (behavioral): [e.g., "Ship before it's comfortable"; "Default to transparency"]
- One-liner / boilerplate (50 words) for press & footers: [...]

## 2. Logo
- Files: /logo (svg, png @1x/2x/3x, favicon, social avatar, OG image)
- Clear space: ≥ [cap-height] on all sides.   Min size: icon [24px], full [120px] / print [10mm].
- Variants: full-color-light-bg, full-color-dark-bg, monochrome-black, monochrome-white (knockout).
- Misuse: do not stretch, recolor, add shadows/outlines, rotate, or place on low-contrast imagery.

## 3. Color
| Token (semantic)        | Light    | Dark     | Use                       |
|-------------------------|----------|----------|---------------------------|
| color.brand.primary     | #______  | #______  | Logo, key brand moments   |
| color.action.default    | #______  | #______  | Buttons, links (AA-safe)  |
| color.text.primary      | #______  | #______  | Body copy                 |
| color.bg.surface        | #______  | #______  | Cards, panels             |
| color.feedback.error    | #______  | #______  | Errors (never color-only) |
- All text/background pairs documented as passing WCAG AA (≥4.5:1 body, ≥3:1 large/UI).

## 4. Typography
- Headings: [Typeface], scale [modular 1.250], weights [600/700].
- Body: [Typeface], 16px base, line-height 1.5, measure 45-75ch.
- Mono (code/data): [Typeface].   Fallback stacks + webfont loading: font-display: swap.

## 5. Imagery & Illustration
- Photography: [style — e.g., natural light, candid, no stock clichés]. Do / Don't examples linked.
- Illustration: [style, stro

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.