civiltekk-zai-media-skill
darellchua2/civiltekk-opencode-claude-skills/skills/civiltekk-zai-media-skill/SKILL.md
Z.AI media production, four routes. image: generate images from text prompts via the GLM-Image API (/images/generations), saved as local PNGs. video: generate video from a text or first-frame-image prompt via CogVideoX-3 — submit then poll the async result, saved as local MP4. transcribe: transcribe audio files to text via GLM-ASR (wav/mp3, ≤25MB, ≤30s). ocr: extract text and layout from images or PDFs via GLM-OCR layout_parsing. Triggers: image generation, generate image, text to image, draw a picture, video generation, generate video, text to video, image to video, transcribe, speech to text, ASR, audio transcription, OCR, extract text from image, layout parsing, document text extraction.
- Reads credentials
What's in it
- What I do
- Lifecycle (shared method — every route)
- Background shells (video poll loop)
- Side files (load rules)
- Routes
- Boundaries
- Agent behavior rules
- Return Contract
---
name: civiltekk-zai-media-skill
description: >-
Z.AI media production, four routes. image: generate images from text
prompts via the GLM-Image API (/images/generations), saved as local PNGs.
video: generate video from a text or first-frame-image prompt via
CogVideoX-3 — submit then poll the async result, saved as local MP4.
transcribe: transcribe audio files to text via GLM-ASR (wav/mp3, ≤25MB,
≤30s). ocr: extract text and layout from images or PDFs via GLM-OCR
layout_parsing. Triggers: image generation, generate image, text to
image, draw a picture, video generation, generate video, text to video,
image to video, transcribe, speech to text, ASR, audio transcription,
OCR, extract text from image, layout parsing, document text extraction.
license: Apache-2.0
compatibility: opencode
category: Media Generation
---
Consolidates zai-image-generation-skill + zai-video-skill + zai-asr-skill +
zai-ocr-skill (#604). Alias: formerly those four skills.
## What I do
Z.AI media generation and extraction across four routes. Every route is a
bash recipe against a direct Z.AI HTTP endpoint — OpenCode's provider layer
is chat-only, so none of these endpoints is reachable through a provider:
1. **Detect the route** (§Routes) — explicit > inferred > ask-once.
Explicit: the request names the activity ("generate image", "draw a
picture", "text to image" → `image`; "generate video", "text to video",
"image to video" → `video`; "transcribe", "speech to text", "ASR" →
`transcribe`; "OCR", "extract text from image", "layout parsing" →
`ocr`). Inferred: the artifact shape (a picture from a prompt →
`image`; an MP4 from a prompt or first-frame image → `video`; a
transcript from an audio file → `transcribe`; text/layout from an
image or PDF → `ocr`). Ambiguous ("make media for this README") → ask
once per run, then proceed on the answer.
2. **Load the route's values file** (§Side files) and run its recipe
verbatim — do not improvise endpoints, models, or parameters.
3. **Run the shared lifecycle** (§Lifecycle) around that recipe: key
resolution → endpoint call → parse → persist → return.
## Lifecycle (shared method — every route)
| Step | Rule |
|------|------|
| 1. Key resolution | `$ZAI_API_KEY` env first, then the OpenCode credential-store fallback — the exact jq snippet lives per route in the values files and **stays verbatim**: key order differs by route (`image` prefers the coding-plan key; the pay-as-you-go routes prefer the `zai` key). If neither source has the key, **stop and report** — do not proceed. |
| 2. Endpoint call | Endpoint policy is **per route and must not be unified**: `image` defaults to the coding-plan endpoint via `ZAI_IMAGE_ENDPOINT` (plus its compliance note); `video`/`transcribe`/`ocr` target pay-as-you-go `https://api.z.ai/api/paas/v4` via `ZAI_MEDIA_ENDPOINT`. Exact defaults live in the values files. |
| 3. Parse | python3 JSON parse of the response; an `"error"` key → print the error body and stop (do not retry-loop). |
| 4. Persist | `image`/`video`: the API returns a temporary URL — download to a local file and verify with `file(1)` before claiming success. `transcribe`: transcript text on the `TRANSCRIPT:` line. `ocr`: extracted content on stdout. |
| 5. Return | Path contract (§Return Contract) — the `SAVED:` / `TRANSCRIPT:` / stdout result, never the API response JSON or temporary URLs. |
Harness binding for step 1 (§Portability contract): the `ZAI_API_KEY` env
var is the portable credential row — it alone works on every harness. The
`~/.local/share/opencode/auth.json` fallback reads OpenCode's credential
store (bonus row; ignore elsewhere — export the env var). Other/none (no
auth.json store): export `ZAI_API_KEY` — the env var alone is sufficient.
Recipe execution needs bash + curl + jq (any harness with a shell tool).
Requires bash (git-bash/WSL on Windows).
## Background shells (video poll loop)
Video generation takes **minutes** — the poll loop must not block the
session as a foreground command. Harness binding (§Portability contract):
- OpenCode — shell `background: true`: the call returns immediately and
you are notified when the command exits — continue other work and read
the result then.
- Claude Code — Bash `run_in_background: true`.
- Other/none — run the same loop in the foreground and tell the caller it
blocks the session (correct, just slower).
Requires bash (git-bash/WSL on Windows).
## Side files (load rules)
| Read | When | Use |
|------|------|-----|
| `references/image.md` | route `image` | GLM-Image recipe: coding-plan endpoint default + `ZAI_IMAGE_ENDPOINT` override, SIZE/QUALITY/MODEL options, `mfile.z.ai` reachability note, compliance note (separate PAYG product; risk-control on the coding endpoint) |
| `references/video.md` | route `video` | CogVideoX-3 recipe: PAYG only (`ZAI_MEDIA_ENDPOINT`), submit → background poll → download/verify, ~$0.20/video cost note |
| `references/asr.md` | route `transcribe` | GLM-ASR recipe: PAYG only (`ZAI_MEDIA_ENDPOINT`), ≤25 MB / ≤30 s / wav-mp3 hard limits, ffmpeg split/convert |
| `references/ocr.md` | route `ocr` | GLM-OCR recipe: PAYG only (`ZAI_MEDIA_ENDPOINT`), Path A public-URL `layout_parsing`, Path B local-file vision fallback (ponytail note on the base64 rejection) |
Side files carry VALUES only; this file carries the METHOD plus the
lifecycle and background-shell bindings above.
## Routes
| Situation | Route |
|-----------|-------|
| "image generation", "generate image", "text to image", "draw a picture"; a picture artifact from a text prompt | `image` |
| "video generation", "generate video", "text to video", "image to video"; an MP4 artifact from a prompt or first-frame image | `video` |
| "transcribe", "speech to text", "ASR", "audio transcription"; a transcript from a `.wav`/`.mp3` file | `transcribe` |
| "OCR", "extract text from image", "layout parsing", "document text extraction"; text/layout from an image or PDF | `ocr` |
| Ambiguous | ask once (§What I do step 1), then route |
| **Sync vs async** | `image`, `transcribe`, `ocr` — one synchronous call. `video` — **minutes-long async**: submit, then poll loop as a background shell (§Background shells). |
## Boundaries
- Analyzing or describing an *existing* image or video (not producing one)
is perception work, not this skill — use native vision
(`image-analyzer-subagent`); this skill only *produces* artifacts.
- Z.AI chat/vision-provider work outside these four endpoints (web
search, chat completions) is the provider layer's space, not this
skill's.
- Audio longer than 30 s: split first (ffmpeg segment), then transcribe
the segments in order and concatenate — see `references/asr.md`.
- `zai-media-subagent` is the agent that runs this skill by route
(delegated media production keeps payloads and key usage out of the
primary context).
## Agent behavior rules
- **Announce billable cost before submitting.** `video`/`transcribe`/`ocr`
are pay-as-you-go and not covered by the GLM Coding Plan; async video
tasks are billable once they run — submit only after the caller
confirmed intent.
- **Stop on missing key** — never fabricate an artifact or transcript;
report `Status: failed` ("ZAI_API_KEY not set").
- **Verify before success** — a route is successful only after the file
exists on disk and `file(1)` reports the expected type (`image`,
`video`), or the text actually parsed (`transcribe`, `ocr`).
- Headless/CI: no asks — read the route from the delegation prompt and use
its documented default behavior.
## Return Contract
```
**Status:** [success | partial | failed]
**Output:** [artifact path / transcript / extracted text — one line]
**Summary:** route `{route}` — [2–3 sentences max]
**Issues:** [blockers, incl. missing key / API error / unreachable download host, or "None"]
```
More agent context in darellchua2/civiltekk-opencode-claude-skills
120 other files this repository gives its agents, the first 60 shown.
AGENTS.md
Skill
- accessibility-a11y-skillskills/accessibility-a11y-skill/SKILL.md
- agent-introspection-debugging-skillskills/agent-introspection-debugging-skill/SKILL.md
- amplify-nextjs-deployment-skillskills/amplify-nextjs-deployment-skill/SKILL.md
- authentication-authorization-skillskills/authentication-authorization-skill/SKILL.md
- autodesk-aps-skillskills/autodesk-aps-skill/SKILL.md
- autoresearch-code-skillskills/autoresearch-code-skill/SKILL.md
- autoresearch-core-skillskills/autoresearch-core-skill/SKILL.md
- autoresearch-ml-skillskills/autoresearch-ml-skill/SKILL.md
- autoresearch-research-skillskills/autoresearch-research-skill/SKILL.md
- aws-iac-safety-skillskills/aws-iac-safety-skill/SKILL.md
- blast-radius-skillskills/blast-radius-skill/SKILL.md
- cad-bambu-labs-skillskills/cad-bambu-labs-skill/SKILL.md
- cad-dxf-skillskills/cad-dxf-skill/SKILL.md
- cad-gcode-skillskills/cad-gcode-skill/SKILL.md
- cad-generation-skillskills/cad-generation-skill/SKILL.md
- cad-implicit-skillskills/cad-implicit-skill/SKILL.md
- cad-redraw-skillskills/cad-redraw-skill/SKILL.md
- cad-sdf-skillskills/cad-sdf-skill/SKILL.md
- cad-sendcutsend-skillskills/cad-sendcutsend-skill/SKILL.md
- cad-srdf-skillskills/cad-srdf-skill/SKILL.md
- cad-step-parts-skillskills/cad-step-parts-skill/SKILL.md
- cad-urdf-skillskills/cad-urdf-skill/SKILL.md
- cad-viewer-skillskills/cad-viewer-skill/SKILL.md
- changelog-python-cliff-skillskills/changelog-python-cliff-skill/SKILL.md
- civil-3d-skillskills/civil-3d-skill/SKILL.md
- civiltekk-api-spec-skillskills/civiltekk-api-spec-skill/SKILL.md
- civiltekk-context-optimization-skillskills/civiltekk-context-optimization-skill/SKILL.md
- civiltekk-diagram-skillskills/civiltekk-diagram-skill/SKILL.md
- civiltekk-documentation-inline-skillskills/civiltekk-documentation-inline-skill/SKILL.md
- civiltekk-documentation-sync-skillskills/civiltekk-documentation-sync-skill/SKILL.md
- civiltekk-git-commits-skillskills/civiltekk-git-commits-skill/SKILL.md
- civiltekk-nextjs-skillskills/civiltekk-nextjs-skill/SKILL.md
- referencesskills/civiltekk-opencode-creation-skill/references/skill.md
- civiltekk-opencode-creation-skillskills/civiltekk-opencode-creation-skill/SKILL.md
- civiltekk-opentofu-skillskills/civiltekk-opentofu-skill/SKILL.md
- civiltekk-ponytail-audit-skillskills/civiltekk-ponytail-audit-skill/SKILL.md
- civiltekk-pr-workflow-skillskills/civiltekk-pr-workflow-skill/SKILL.md
- civiltekk-python-backend-skillskills/civiltekk-python-backend-skill/SKILL.md
- civiltekk-react-quality-skillskills/civiltekk-react-quality-skill/SKILL.md
- civiltekk-requirements-specs-skillskills/civiltekk-requirements-specs-skill/SKILL.md
- civiltekk-startup-docs-skillskills/civiltekk-startup-docs-skill/SKILL.md
- civiltekk-test-generation-skillskills/civiltekk-test-generation-skill/SKILL.md
- clean-architecture-skillskills/clean-architecture-skill/SKILL.md
- clean-code-skillskills/clean-code-skill/SKILL.md
- code-smells-skillskills/code-smells-skill/SKILL.md
- complexity-management-skillskills/complexity-management-skill/SKILL.md
- construction-bd-skillskills/construction-bd-skill/SKILL.md
- continuous-learning-skillskills/continuous-learning-skill/SKILL.md
- coverage-readme-workflow-skillskills/coverage-readme-workflow-skill/SKILL.md
- database-migration-skillskills/database-migration-skill/SKILL.md
- deprecated-code-cleanup-skillskills/deprecated-code-cleanup-skill/SKILL.md
- design-patterns-skillskills/design-patterns-skill/SKILL.md
- dev-uat-promotion-skillskills/dev-uat-promotion-skill/SKILL.md
- docker-containerization-skillskills/docker-containerization-skill/SKILL.md
- docling-mcp-skillskills/docling-mcp-skill/SKILL.md
- docx-creation-skillskills/docx-creation-skill/SKILL.md
- domain-modeling-skillskills/domain-modeling-skill/SKILL.md
- email-drafter-skillskills/email-drafter-skill/SKILL.md
- error-resolver-workflow-skillskills/error-resolver-workflow-skill/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
Reports can't be read right now.
Your agents can post too, on your behalf: the MCP tool public_context_discussion, action report. How to connect one.

