server-ops
wlkerwong-boop/china-web-deployment/skills/server-ops/SKILL.md
Alibaba Cloud / Linux server operations: swap, SSH hardening, acme.sh SSL, PM2 proxy, domain checks.
Skill0 starsChanged 17 days ago
- Reads credentials
- Deletes or force-pushes
---
name: server-ops
description: "Alibaba Cloud / Linux server operations: swap, SSH hardening, acme.sh SSL, PM2 proxy, domain checks."
version: 1.0.0
author: Hermes Agent
license: MIT
metadata:
hermes:
tags: [server, alibaba-cloud, ssl, ssh, pm2, swap, acme-sh, domain, proxy]
---
# Server Operations
Covers recurring server-management patterns for the Alibaba Cloud lightweight server (203.0.113.10) and Mac mini local services. Use when the user asks about server setup, SSL certificates, swap, SSH, or PM2 process management.
## Trigger conditions
- User mentions: server, 阿里云, swap, SSL, 证书, SSH, PM2, proxy, 部署, 域名查询, launchd
- Any task involving acme.sh cert issuance or renewal
- Setting up or troubleshooting reverse proxies (Node http-proxy)
- Domain availability checking
---
## 1. Domain Availability Checking
When the user needs to check if domains are available:
1. **Alibaba WHOIS** (`https://wanwang.aliyun.com/whois/<domain>`) — first stop for .cn/.com.cn
2. **System `whois`** — for .com/.net/.org; if no record found, domain is available
3. **`dig +short <domain> A`** — if no DNS resolution, domain is likely unregistered
4. **Registry-specific WHOIS** — `whois -h whois.verisign-grs.com <domain>` for .com; `whois -h whois.cnnic.cn <domain>` for .cn
5. **Web search** — search for the domain name to see if any site actively uses it
A domain with no DNS resolution + no WHOIS match + no search results is almost certainly available.
### Pitfalls
- `.academy` and other new gTLDs may not be queryable via standard WHOIS — use dig + web search
- CNNIC WHOIS rate-limits; space requests by at least 2 seconds
- `whois` default behavior queries IANA first — always follow referrals or use `-h` with the specific registry server
---
## 2. Swap Setup (Alibaba Cloud)
The Alibaba Cloud lightweight server often loses swap after reboot. Persistent setup:
```bash
fallocate -l 2G /swapfile || dd if=/dev/zero of=/swapfile bs=1M count=2048
chmod 600 /swapfile
mkswap /swapfile
swapon /swapfile
echo "/swapfile swap swap defaults 0 0" >> /etc/fstab
```
Verify: `free -h` should show the swap; `swapon --show` confirms persistence.
### Pitfalls
- `fallocate` may not be available on older kernels — fall back to `dd`
- The server (759MB RAM) needs swap just to survive; without it, new PM2 processes will OOM-kill existing ones
- Swap was previously configured but lost after reboot — always verify `/etc/fstab` entry survived
---
## 3. SSH Hardening
After SSH key auth is confirmed working:
```bash
# Install public key
cat ~/.ssh/id_ed25519.pub | ssh root@<server> 'cat >> ~/.ssh/authorized_keys && chmod 600 ~/.ssh/authorized_keys'
# Test key auth BEFORE disabling password
ssh -o PasswordAuthentication=no root@<server> 'echo OK'
# Disable password
cp /etc/ssh/sshd_config /etc/ssh/sshd_config.bak-$(date +%Y%m%d)
sed -i 's/^PasswordAuthentication yes/PasswordAuthentication no/' /etc/ssh/sshd_config
sshd -t && systemctl reload sshd
```
### Pitfalls
- Never disable password until key auth is confirmed working in a SEPARATE ssh call
- If `authorized_keys` is created with wrong permissions, key auth silently fails
- SSH key must be ed25519 (RSA may be rejected by newer OpenSSH)
---
## 4. acme.sh SSL Certificate Management
### ALPN mode (port 443 — MUST stop proxy first)
The proxy occupies port 443. acme.sh ALPN mode also needs port 443. **Always stop the proxy first**, then issue, then restart:
```bash
pm2 stop proxy
sleep 2
~/.acme.sh/acme.sh --issue --alpn --standalone \
-d domain.com -d www.domain.com \
--keylength ec-256 --force
# Verify SAN
openssl x509 -in ~/.acme.sh/domain_ecc/fullchain.cer -text -noout | grep "DNS:"
pm2 start proxy # or pm2 restart proxy if already running
```
### Extending an existing cert with new domains
Use `--issue --force` (NOT `--renew` — renew keeps existing domain list):
```bash
pm2 stop proxy
~/.acme.sh/acme.sh --issue --alpn --standalone \
-d existing.com -d new.subdomain.com \
--keylength ec-256 --force
pm2 start proxy
```
### HTTP-01 pitfalls
HTTP-01 standalone mode on port 80 frequently fails with 403 even when the port is confirmed free. The Python fallback server (used when `socat` is absent) gets terminated mid-validation. **Always prefer ALPN mode.**
If HTTP-01 is the only option:
- Install `socat`: `yum install -y socat` (removes the Python fallback issue)
- Or use DNS validation with Alibaba Cloud API (`dns_ali.sh` in acme.sh dnsapi/)
### Installing cert with auto-reload
```bash
~/.acme.sh/acme.sh --install-cert -d domain.com \
--key-file /root/.acme.sh/domain_ecc/domain.key \
--fullchain-file /root/.acme.sh/domain_ecc/fullchain.cer \
--reloadcmd "pm2 restart proxy"
```
### Auto-renewal
acme.sh cron runs daily: `crontab -l` should show the acme.sh `--cron` entry. The `--reloadcmd` from `--install-cert` handles proxy restart on renewal.
### Pitfalls
- Server has **no socat** — the Python fallback is unreliable for HTTP-01
- Renewal reload cmd must use `pm2 restart proxy`, NOT `nginx -s reload` (no nginx installed)
- The server runs a Node.js reverse proxy, not nginx; acme.sh's `--nginx` mode will fail
- **`--renew` does NOT extend domain list** — use `--issue --force` with all domains (old + new) to add SAN entries
- **ALPN needs port 443 free** — always `pm2 stop proxy` first, or acme.sh will fail with "tcp port 443 is already used"
- **HTTP-01 `--webroot` may fail** if port 80 redirects from proxy interfere — prefer ALPN
---
## 5. PM2 Reverse Proxy with SNI (Multi-Cert)
## 5. PM2 Reverse Proxy with SNI (Multi-Cert)
### ✅ SNICallback: Working and Reliable (Node v22, 2026-07-26 verified)
SNICallback + `tls.createSecureContext` works correctly on Node.js v22.23.1. The multi-cert SNI pattern is now deployed in production proxy-v4.js, serving 5 domains across 3 certificates with 0 crashes.
```javascript
// SNICallback with correct syntax
https.createServer({
SNICallback: (servername, cb) => {
const cert = getCert(servername);
// 🔥 MUST use explicit {key, cert} — direct cert object SILENTLY crashes
cb(null, tls.createSecureContext({ key: cert.key, cert: cert.cert }));
},
}, routeRequest).listen(443);
```
### 🔥 Critical Syntax Pitfall: tls.createSecureContext Object vs Props
**WRONG** — silently crashes TLS handshake with no error log:
```javascript
tls.createSecureContext(cert) // cert = {key: Buffer, cert: Buffer}
```
**RIGHT** — must explicitly pass properties:
```javascript
tls.createSecureContext({ key: cert.key, cert: cert.cert })
```
This pitfall was the root cause of repeated proxy crash loops on 2026-07-26. The proxy starts normally, prints "HTTPS proxy on :443", but every TLS connection is reset. PM2 restarts climb with no useful error. `node --check` passes (syntax is valid). Only `openssl s_client` on the server reveals the cert is never served.
### 🏆 No-SNI Alternative: Single Multi-SAN Cert
When all domains fit in one certificate's SAN, use the simpler single-cert pattern:
```javascript
https.createServer(certB, routeRequest).listen(443);
```
Only add SNI when domains genuinely need separate certs (different CA, different expiry, or SAN overflow). The multi-cert approach adds complexity but is reliable when the syntax above is used correctly.
### getCert Domain Ordering
Check **subdomains before parent domains** to avoid unintended matches:
```javascript
function getCert(servername) {
if (servername && servername.includes("sub.site-a.com")) return certC; // FIRST
if (servername && servername.includes("site-a.com")) return certA; // catch bare + www
return certB; // default: main site
}
```
If the parent-domain check comes before the subdomain check, `sub.site-a.com` matches `site-a.com` first and gets the wrong certificate.
### headersSent Guards (Mandatory)
ALL `proxyReq.on("error", ...)` handlers MUST check `res.headersSent`:
```javascript
proxyReq.on("error", () => {
if (!res.headersSent) { res.writeHead(502); res.end("Bad Gateway"); }
});
```
Without this guard, partial responses followed by upstream failure → `ERR_HTTP_HEADERS_SENT` crash.
---
## 6. PM2 Process Management Quick Reference
```bash
pm2 list # Status overview
pm2 logs <name> --lines 50 # Recent logs (add --nostream for non-following)
pm2 stop/restart <name> # Control individual processes
pm2 start script.js --name X # New process
pm2 delete <name> # Remove from list (then pm2 save)
pm2 save # Persist process list for reboot
```
### Pitfalls
- `pm2 restart` preserves the script path from the initial `pm2 start` — to change the script, `delete` then `start`
- After `delete` + `start`, PM2 warns "not in sync with saved list" — just `pm2 save` again
- **`pm2 restart` may not pick up file changes**: if restart count climbs with the same error, PM2 caches the old script. Fix: `pm2 delete` + `pm2 start` fresh.
---
## 7. Proxy Edit Protocol (Alibaba Server)
Any edit to the live proxy (`/root/proxy-v4.js`):
1. `cp /root/proxy-v4.js /root/proxy-v4.js.bak-$(date +%Y%m%d-%H%M)`
2. Make changes — **prefer Python inline script over sed for multi-line or special-character edits**
3. **Syntax check BEFORE restart**: `node -c /root/proxy-v4.js` — skip this and the proxy crashes in a loop
4. `pm2 restart proxy`
5. Verify: `curl -sI https://<domain>/ --max-time 5 | head -3`
### 🔥 sed pitfall: special characters silently corrupt the file
`sed` with `&&`, `\n`, or other special chars in the replacement string often produces broken JS:
```bash
# BROKEN: && gets interpreted, \n produces literal backslash-n
sed -i "s/old/if (x && y) {\n return z;\n}/" proxy-v4.js
# RIGHT: use Python for any non-trivial edit
ssh root@host 'python3 -c "
with open(\"/root/proxy-v4.js\", \"r\") as f:
content = f.read()
content = content.replace(\"old_string\", \"\"\"new
multi-line
string\"\"\")
with open(\"/root/proxy-v4.js\", \"w\") as f:
f.write(content)
print(\"OK\")
"'
```
After sed corruption: `node -c` catches the syntax error, but the proxy may already be in a crash loop. Restore from `.bak` immediately, then redo with Python.
### Zero-downtime constraint
The proxy restart causes ~2 seconds of downtime. Avoid during peak usage. The backend services are NOT affected — only the proxy itself is briefly down.
---
## 8. DeepSeek API Diagnostics
When a report/API on the server returns empty content or appears broken:
### Quick Check Pipeline
```bash
# 1. Check balance
curl -s https://api.deepseek.com/user/balance -H "Authorization: Bearer $KEY"
# 2. Verify API key works (Python — most reliable)
python3 -c "
import json, urllib.request
# read key from api-service/index.js
req = urllib.request.Request('https://api.deepseek.com/v1/chat/completions', ...)
resp = urllib.request.urlopen(req, timeout=15)
print(json.loads(resp.read())['choices'][0]['message']['content'])
"
# 3. Test from Node.js (what api-service actually uses)
node -e "
const https=require('https');
// ... same API call
"
```
### Pitfall: Node.js + deepseek-v4-pro = empty content
**Symptoms:** API returns HTTP 200 with `finish_reason: length` and empty content (`""`). No error in response body. Same API call works perfectly from Python `urllib` or `curl`. At the application level, this manifests as `{success: true, report: ""}` — calculations succeed but the LLM-generated text is silently empty.
**Root cause:** Node.js 22 `fetch()` / `https.request()` on Alibaba Cloud CentOS returns empty content from `deepseek-v4-pro` model. `deepseek-chat` works fine in both Node.js and Python. This is a TLS/client-library compatibility issue — not a balance or key problem.
**Fix:** Use `model:'deepseek-chat'` instead of `model:'deepseek-v4-pro'` in Node.js code.
**Diagnostic flow:**
1. Check balance first: `curl -s https://api.deepseek.com/user/balance -H "Authorization: Bearer $KEY"`
2. Test from Python (most reliable): same API call with `urllib` — if this works, the key/balance are fine
3. Test from Node.js: same API call with `https.request` — if this returns empty, it's this bug
4. The telltale signal: `finish_reason: length` + `content: ''` from Node.js, but `finish_reason: stop` + valid content from Python
**How to confirm:** Test the same API call from both Python and Node.js on the server. If Python works and Node.js returns empty with `finish_reason: length`, it's this bug. See `references/deepseek-nodejs-bug.md` for a full reproduction recipe and the api-service specific diagnostic flow.
### Report-API specific pitfalls
- **calcBazi can silently break**: The function was previously modified to use non-existent method names (`lunar['getYEARInGanZhi']()`). Always verify against backups (`index.js.bak`, `index.js.bak3`). Correct methods: `getYearInGanZhiExact()`, `getMonthInGanZhi()`, `getDayInGanZhi()`, `getTimeInGanZhi()`.
- **PM2 restart is mandatory after code change**: PM2 caches the original file. Changes to `index.js` won't take effect until `pm2 restart api-service`.
- **Port 3003 may have zombie instances**: After repeated crashes, use `lsof -iTCP:3003 -sTCP:LISTEN` to verify only one instance is running.
### Prevention: Don't assume "402 = balance"
Always verify the actual API response before concluding balance issues. The DeepSeek `user/balance` endpoint gives definitive balance info. A 402 in logs may be historical; the current issue could be a model compatibility bug. Use the diagnostic flow above (Python first, then Node.js) to isolate the real cause before topping up balance.
## 10. API & Deployment Debugging
### Streaming API Testing (SSE/EventStream)
When testing streaming endpoints (e.g. `/api/master-report/stream`) with curl, `-s` hides headers and curl hangs waiting for stream data. Use `-sv` to see response headers immediately:
```bash
# WRONG: hangs with no output — looks broken
curl -s --max-time 10 -X POST https://example.com/api/stream -d '{}'
# RIGHT: shows HTTP 200 + content-type immediately
curl -sv --max-time 10 -X POST https://localhost/api/stream -H "Host: dom" -k -d '{}' 2>&1 | head -15
# Look for: HTTP/1.1 200, content-type: text/event-stream
```
A 200 with `text/event-stream` means the API IS working — the problem is upstream (proxy timeout, browser fetch handling, missing SSE event listener), not the endpoint itself. Always test from BOTH internal (port 3005) AND external (through proxy port 443) to isolate proxy vs app issues.
### `.next/BUILD_ID` Deployment Validation
After rsync of `.next/` to the server, **always verify BUILD_ID exists** before restarting the app:
```bash
ssh root@host 'ls /root/app/.next/BUILD_ID || echo "MISSING — rsync failed, rebuild and re-rsync"'
```
Missing BUILD_ID → Next.js throws `Could not find a production build` → PM2 restart loops with 50+ crashes/minute. Root cause: rsync race condition where `.next/` directory exists but files are incomplete. **Fix: `rm -rf /root/app/.next` on server, rebuild locally, re-rsync with `--delete`.**
When the user's task document says "调用 Codex 完成":
1. **First, quick-test Codex**: `CODEX=$(find ~/.hermes/node/lib/node_modules/@openai/codex -name codex -type f | head -1) && $CODEX exec "hello"` — if it fails with config errors (wire_api, provider, etc.), skip to step 5.
2. Find the real Codex binary: `CODEX=$(find ~/.hermes/node/lib/node_modules/@openai/codex -name codex -type f | head -1)`
3. Use `terminal(pty=true, background=true, notify_on_complete=true)` — Codex is interactive, needs PTY
4. Prompt structure: numbered steps, explicit boundaries, verification commands at each step
5. **FALLBACK: delegate_task** — When Codex fails to start (e.g., wire_api removed in 0.128+), use `delegate_task` with `toolsets: ["terminal","file","web"]`. Subagents are MORE reliable — they use the same model as Hermes, have full tool access, and are unaffected by Codex version changes. This is now the recommended pattern for all deployment/ops tasks.
6. Include the "boundary" section: "不得修改仓库代码;不得重启 X/Y/Z 服务;遇到错误完整记录并报告"
7. Include the "security" section: "密钥不进代码、不进输出"
See `references/codex-task-template.md` for a reusable prompt template.
## Support Files
- `references/codex-task-template.md` — reusable prompt template for dispatching tasks to Codex
- `references/deploy-verify-checklist.md` — mandatory browser verification after deploying frontend changes
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

