agentleFS
Sign inSign up

visual-asset-management-system / backendPipelines

awslabs/visual-asset-management-system/backendPipelines/CLAUDE.md

This is the Claude Code steering document for the backendPipelines/ tree. It covers the pipeline architecture, S3 output-path conventions, assetId threading, and the end-to-end checklist for adding a new processing pipeline. For project-wide context, see the root CLAUDE.md. VAMS supports four pipeline execution types (PIPELINEEXECUTIONTYPES in backend/backend/models/pipelines.py): Lambda (sync or async invocation of a Lambda function), SQS (async message to an SQS queue), EventBridge (async event to an EventBridge bus), and DeadlineCloud (an AWS Deadline Cloud job). SQS and EventBridge…

CLAUDE.md142 starsChanged 5 days ago
# VAMS Pipeline Development Guide

This is the Claude Code steering document for the `backendPipelines/` tree. It covers the pipeline architecture, S3 output-path conventions, `assetId` threading, and the end-to-end checklist for adding a new processing pipeline. For project-wide context, see the root `CLAUDE.md`.

## Pipeline Architecture

VAMS supports four pipeline execution types (`PIPELINE_EXECUTION_TYPES` in `backend/backend/models/pipelines.py`): **Lambda** (sync or async invocation of a Lambda function), **SQS** (async message to an SQS queue), **EventBridge** (async event to an EventBridge bus), and **DeadlineCloud** (an AWS Deadline Cloud job). SQS and EventBridge pipelines are async-only and support optional callback via Step Functions Task Tokens.

DeadlineCloud is async-only and its callback is **mandatory** — `createJob` only queues the job, so `waitForCallback` must be `Enabled` or pipeline create is rejected with a `400`. It is built by `DeadlineCloudTaskBuilder` (`backend/backend/common/workflows/stepfunctions_builder.py`) plus the `deadlineCloudJobCallback` handler, and is gated by `app.pipelines.deadlineCloudExecutionTypeEnabled`: accepted only in the commercial `aws` partition, and rejected at pipeline create when the deployment has not enabled it. Because the type has no non-callback route, treat `SendTaskFailure` as required for it — see [Reporting Failure on a Task-Token Pipeline](#reporting-failure-on-a-task-token-pipeline).

```
S3 event / API trigger → Lambda → Step Functions → Lambda / SQS / EventBridge / DeadlineCloud → AWS Batch containers (optional)
  backendPipelines/{useCase}/lambda/    -- orchestration
  backendPipelines/{useCase}/container/ -- processing
```

## Pipeline S3 Output Paths

The workflow ASL generates several S3 paths passed to each pipeline step. Pipelines must use the correct path for each output type:

| Path Variable                          | Bucket           | Purpose                                                         | Versioned |
| -------------------------------------- | ---------------- | --------------------------------------------------------------- | --------- |
| `outputS3AssetFilesPath`               | Asset bucket     | File-level outputs: new files, file previews (`.previewFile.X`) | Yes       |
| `outputS3AssetPreviewPath`             | Asset bucket     | Asset-level previews only (whole-asset representative image)    | Yes       |
| `outputS3AssetMetadataPath`            | Asset bucket     | Metadata files produced by the pipeline                         | Yes       |
| `inputOutputS3AssetAuxiliaryFilesPath` | Auxiliary bucket | Temporary working files or special non-versioned viewer data    | No        |

**Key distinction:** `outputS3AssetFilesPath` is for file-level outputs including `.previewFile.gif/.jpg/.png` thumbnails tied to specific files. `outputS3AssetPreviewPath` is only for asset-level preview images that represent the asset as a whole. Most pipelines producing file previews should write to `outputS3AssetFilesPath`.

**Rules for output path usage:**

-   **Always pass through** all output paths from the workflow payload in `vamsExecute` lambdas. Never hardcode empty strings for output paths — the workflow's process-output step relies on finding files at these locations.
-   **Use `outputS3AssetFilesPath`** for file-level outputs, including `.previewFile.X` thumbnail files generated by preview pipelines.
-   **Use `outputS3AssetPreviewPath`** only for asset-level preview images (not file-level previews).
-   **Use `inputOutputS3AssetAuxiliaryFilesPath`** only for temporary files during processing or for special non-versioned preview data (e.g., Potree octree viewer files) that the frontend reads directly from the auxiliary bucket.
-   The `constructPipeline` lambda should prefer the appropriate output path when provided, falling back to the auxiliary path only for direct/local invocations where workflow context is unavailable.

## Metadata and Attribute Output Formats (`outputS3AssetMetadataPath`)

The process-output step (`handlers/workflows/sfn/processWorkflowExecutionOutput.py`) decides what a
JSON file under `outputS3AssetMetadataPath` means **from its file name**, matched on suffix anywhere
under the path (subdirectories included). There are exactly three shapes:

| File name written by the pipeline   | Written to                       | Store           |
| ----------------------------------- | -------------------------------- | --------------- |
| `asset.metadata.json`               | The asset (recorded against `/`) | Asset metadata  |
| `<relativeFilePath>.metadata.json`  | That file                        | File metadata   |
| `<relativeFilePath>.attribute.json` | That file                        | File attributes |

`asset.metadata.json` is a **reserved basename**: any other `*.metadata.json` is treated as
file-level and its target path is derived from the file name, so a pipeline that names its
asset-level file anything else silently writes file metadata against a path that does not exist.

All three use the same body. `updateType` is `update` (upsert, the default) or `replace_all`:

```json
{
    "type": "metadata",
    "updateType": "update",
    "metadata": [{ "metadataKey": "AB_geometric_metadata", "metadataValue": "{\"volume\": 12.4}" }]
}
```

`type` is `metadata` or `attribute` and is auto-corrected to match the file name, so the file name is
what actually decides. A missing or non-list `metadata` array fails the write-back.

**A pipeline may annotate a file it produces in the same run.** The process-output step ingests
`outputS3AssetFilesPath` **before** it lists `outputS3AssetMetadataPath` — the file write is
synchronous and its per-file outcome is read first — so by the time metadata is written the new file
exists on the asset. Write the new file to `outputS3AssetFilesPath` and its metadata to
`outputS3AssetMetadataPath` in the same execution; no second run or ordering flag is needed.

Two consequences of that ordering worth designing to:

-   The metadata target path must match the file's **final asset-relative path**, which includes the
    workflow's `defaultOutputFileBaseExecutionPathExtension`. The step applies that extension when
    deriving the target, so name the metadata file after the relative path the pipeline wrote, not
    after the absolute S3 key.
-   Metadata for a file whose ingestion FAILED is rejected by the metadata service's own
    file-existence check, and the execution is recorded FAILED. That is intended: it prevents metadata
    rows accumulating against files that never landed.

Ordering and all three write kinds are pinned by
`backend/tests/handlers/workflows/test_processOutput_write_order.py`.

## Preserving Relative Paths for Asset-Adjacent Outputs

When a pipeline writes output files that correspond to a specific input file (e.g., `.previewFile.X` thumbnails), the output **must preserve the input file's relative path within the asset**. The process-output step expects outputs at the same relative location as the input so it can move them to the correct final location in the asset bucket.

-   Asset files are stored at `{assetId}/{relative_path}/{filename}` — the relative path may include zero or more subdirectories between the asset ID and the filename.
-   The output paths (`outputS3AssetFilesPath`, etc.) point to the asset root: `s3://bucket/{assetId}/`.
-   Containers must include the relative subdirectory in the output S3 key: `{outputDir}{relative_subdir}/{filename}.previewFile.X`.

```
Input key:  xd130a6d6.../test/pump.e57
Output dir: xd130a6d6.../

✅ Correct output: xd130a6d6.../test/pump.e57.previewFile.gif
❌ Wrong output:   xd130a6d6.../pump.e57.previewFile.gif  (relative path lost)
```

## Reporting Failure on a Task-Token Pipeline

A pipeline registered with `waitForCallback: "Enabled"` receives an AWS Step Functions task token, and the
workflow's task stays `RUNNING` until something reports against that token. The pipeline owns both outcomes:
`SendTaskSuccess` on completion and **`SendTaskFailure` on every failure route**. A route that returns or
raises without reporting leaves the workflow task pending for its full `taskTimeout` — hours on the GPU
pipelines — and the run reads as in-progress the whole time.

Both halves are required, and each is inert without the other:

-   **The handler** calls `SendTaskFailure` from every path that can fail — each `except` block _and_ every
    early `return` that emits a 4xx. A pre-invoke rejection (a manifest that resolves to the wrong file
    count, an unreadable input configuration) is the common case: the container never starts, so nothing
    else can report.
-   **The CDK lambda builder** grants `states:SendTaskFailure` on the `vamsExecute` function itself. Check
    the function, not the file — a builder file often grants it on `openPipeline`/`pipelineEnd` while the
    `vamsExecute` builder lacks it, which reads as present to a file-level grep:

    ```bash
    awk '/export function build.*VamsExecute/,/^}/' <builder>.ts | grep -c SendTaskFailure
    ```

    Without the grant the call raises `AccessDeniedException`, the handler logs it, and the task hangs
    exactly as before — the only difference is one log line.

A nested `RequestResponse` invoke needs its own result check on top of those two halves. A Lambda that
RAISED still returns `StatusCode` 200 — the failure is reported in `FunctionError` — so a
`StatusCode`-only check reads a failed launch as success, no route reports the token, and the callback
task blocks until `taskTimeout`. Copy the guard verbatim rather than re-wording it, so the copies stay
comparable (canonical copy: `multi/modelOps/lambda/vamsExecuteModelOps.py`):

```python
if lambda_response.get('FunctionError'):
    raise Exception(
        "Invoke Open Pipeline Lambda Failed: " + str(lambda_response.get('FunctionError')))
```

That propagation is also why, **in the nested-invoke pipelines**, `openPipeline`'s
`abort_external_workflow` deliberately does NOT wrap its own `send_task_failure` in try/except.
Swallowing it there returns a payload-level 400 under a clean invoke — a shape the caller does not
inspect — so the failed callback becomes invisible and the task hangs for its full `taskTimeout`. Letting
it propagate is what sets `FunctionError`, which makes the caller report the token under the
`vamsExecute` function's own role: a second attempt under a different identity. A duplicate
`SendTaskFailure` against an already-failed token raises `TaskDoesNotExist` inside that handler's own
wrapped abort, which only logs. Both directions are pinned by
`preview/pcPotreeViewer/lambda/tests/test_open_pipeline_function_error.py`.

The scope of that rule is the `FunctionError` channel, so it holds only where a **lambda** invokes
`openPipeline`. Where `openPipeline` is a state machine **state** instead — a `tasks.LambdaInvoke`, as in
`simulation/isaacLabTraining` — there is no calling lambda to inspect a result, and Step Functions fails
the state directly on a raise. Wrapping the callback is then correct, and `isaacLabTraining`'s
`lambda/tests/test_open_pipeline_task_failure.py` pins it: a failing callback must not mask the original
error. The discriminator is one grep, not a judgement:

```bash
grep -rn "InvocationType" <pipeline>/lambda/*.py    # hits -> nested invoke, do not wrap
```

Report the token before propagating, so the original error still reaches Amazon CloudWatch, and make the
call conditional on a token being present: a direct invoke carries none and must not fail inside the
callback helper. In the abort helpers the `cause` is truncated to 256 characters to match the peer
implementations — that is a pre-invoke rejection, whose whole message is one sentence. A handler that
reports a finished JOB's outcome carries the child's own output instead, so it bounds that text itself to
fit the 32768-character `cause` limit; `multi/rapidPipelineEKS`'s CHECK_JOB path is the example, and its
`POD_LOG_TAIL_MAX_CHARS` is deliberately far larger than 256.

**On the SUCCESS path, report the outcome and not the job's output.** A `SendTaskSuccess` output and the
lambda's own return value both land in the execution record, which is durable and readable by anyone who
can read the execution — and on a run that worked, a third-party tool's stdout has no diagnostic value
there. Fetch the log and write it to the function's own log stream, which keeps the read permission
exercised and leaves an operator a tail to read, but keep it out of both payloads. A size bound is not
redaction: it makes the payload fit, it does not make the content appropriate to store. The pair of
properties is pinned by `multi/rapidPipelineEKS/lambda/tests/test_pod_log_bounding.py` — absent from the
success payloads, present in the log stream — because removing the FETCH would satisfy the first alone.

## Capturing a Child Process's Output for the Execution Record

A container that runs its real work in a child process must keep a bounded copy of that child's output,
because the message it raises on failure is what the workflow stores as the execution's error — and that
record is what an operator reads. `subprocess.run(check=True)` with no capture leaves the child writing
to the inherited stdout: the output reaches Amazon CloudWatch, but the container never sees it, so the
exception can only report an exit code and the cause has to be hunted for by log stream.

Use the `_run_streaming` helper the four deployable NVIDIA inference containers carry (identical in each;
the canonical copy is `genAi/nvidia/cosmos/3/container/inference.py`) and put its returned tail in the
raised message:

```python
returncode, output_tail = _run_streaming(cmd, env=env, cwd=REPO_DIR)
if returncode != 0:
    raise RuntimeError(f"Inference failed with exit code {returncode}. Last output:\n{output_tail}")
```

Three things about it are load-bearing:

-   **Read before you wait.** `Popen(stdout=PIPE)` followed by `wait()` with no reader is what deadlocks
    — the child blocks on `write()` once the pipe buffer fills and never exits. The helper's loop drains
    as the child writes, which is why it is safe.
-   **`capture_output=True` does NOT deadlock** — `run()` drains through `communicate()`, which reads both
    pipes concurrently (measured: 1.2 MB, no stall). It is still wrong here for a different reason: it
    yields nothing until the child exits, so a multi-hour GPU job would log nothing at all while running
    and a hang would be undiagnosable. Do not restate the deadlock claim; both measurements are pinned in
    `genAi/nvidia/tests/test_inference_output_capture.py`.
-   **The tail is bounded** (80 lines, 2000 chars each) so a job printing a progress line per step cannot
    grow the container's memory for the whole run. The failure message needs the end of the output, not
    all of it.

The helper is duplicated per container rather than shared: each container is its own Docker build context
whose Dockerfile `COPY`s an explicit file list, so no shared module is importable at container runtime.
Inline it into the existing entry point rather than adding a file — a new file also needs a Dockerfile
`COPY` edit, and a missed one fails at container **runtime**, not at build. The no-drift test compares the
four deployable copies structurally, so edit one and propagate in the same change. A fifth copy sits in
`genAi/nvidia/cosmos/predict/containerv1/`, which is retained as a reference implementation with no
configuration key that deploys it; it is outside the no-drift set and outside every other check in that
file, so a change made across the deployable four does not reach it.

## `assetId` Threading

**`assetId` is a workflow state variable — thread it, don't derive it.** The `assetId` is passed as a top-level field in the workflow event payload. It must be captured in the `vamsExecute` lambda, forwarded to `constructPipeline`, included in the pipeline definition, and used directly in the container. Never attempt to reverse-engineer the asset ID from S3 path segments.

To compute the relative subdirectory in container code using the explicit `assetId`:

```python
# assetId comes from the pipeline definition (threaded from workflow state)
input_parts = stage_input.objectKey.split("/")
asset_id_idx = input_parts.index(assetId)
relative_subdir = "/".join(input_parts[asset_id_idx + 1:-1])  # "" if file is at asset root
```

**Pipeline state threading pattern** (applies to all pipelines, not just preview):

```
Workflow event (assetId, databaseId, paths, ...)
  → vamsExecute lambda: capture assetId from event body, include in messagePayload
    → constructPipeline lambda: read assetId from event, include in definition dict
      → Container: read assetId from PipelineDefinition, use for relative path computation
```

## vamsSchema Registration (how a built-in becomes usable)

A pipeline's CDK stack only creates AWS resources. What makes it appear in VAMS — as a pipeline, its
templates, and a runnable workflow — is a **`vamsSchema/` bundle** the CDK uploads to the artefacts
bucket and imports at deploy time through `SYSTEM_USER` cross-calls
(`backend/backend/common/workflows/vamsSchemaImport.py`). Registration is idempotent: a redeploy
overwrites and unarchives, so re-deploying never duplicates or strips a built-in.

```
backendPipelines/{useCase}/{name}/vamsSchema/
    pipeline.json                  # required
    workflow.json                  # optional (one built-in workflow per pipeline)
    templates/{templateId}.json    # optional, one file per template
```

The registration custom resource re-fires only when the bundle changes: `schemaHash` in
`vamsSchemaRegistration-construct.ts` covers `pipeline.json`, `workflow.json`, and the **top-level**
`templates/*.json` (a subdirectory is skipped, not read). Two rules follow:

-   **Never hash an unresolved CDK token.** Override values like `fn.functionName` stringify to
    `${Token[TOKEN.n]}`, where `n` shifts when any unrelated construct is added. Hashing that text
    re-fires every registration on an unrelated deploy, and each one overwrites operator edits to the
    built-in (rename, retuned `systemConfig`, deliberate archive) from the schema files. Substitute a
    placeholder for token values — the resolved value still reaches CloudFormation via the
    `resourceOverrides` / `idOverrides` properties, which detect a real retarget themselves. Test:
    `infra/test/pipelines/vamsSchemaRegistrationHash.test.ts`.
-   **Bundles share the artefacts bucket with `infra/lib/artefacts/`.** The root `DeployArtefacts`
    deployment prunes (`s3 sync --delete`) over the bucket root, so it must keep excluding
    `vamsSchema/*` — otherwise refreshing an unrelated artefact deletes every bundle while the
    registration resources still expect to read them. Test: `infra/test/storage/artefactsBucketPrune.test.ts`.

`pipeline.json` carries no ARNs — the execution target is injected at deploy time from
`resource_overrides` per `executionConfig.executionType` (`lambda.resourceId`, `sqs.queueUrl`,
`eventBridge.busArn`, `deadlineCloud.farmId`, …), so the same file works in every account and
partition. Author it with the block for your execution type present but empty:

```json
{
    "pipelineName": "3D Basic Conversion",
    "category": "Conversion",
    "description": "…",
    "executionConfig": {
        "executionType": "Lambda",
        "waitForCallback": "Disabled",
        "taskTimeout": "900",
        "lambda": {}
    },
    "systemConfig": {
        "inputFileArity": "one",
        "assetScope": { "wholeAsset": false },
        "metadataInputs": {
            "assetMetadata": false,
            "fileMetadata": false,
            "fileAttributes": false
        },
        "requireTemplate": true,
        "allowCustomTemplateOverride": true,
        "inputFileFilters": { "allow": ["*.glb", "*.stl"], "exclude": [] }
    }
}
```

**Rules that bite if you get them wrong:**

1. **`inputFileFilters.allow` must match what the container actually accepts.** The filter is what
   the execute API and the file-upload trigger match against — a missing extension makes the pipeline
   silently unselectable for that file type.
   **An omitted, empty, or `*` allow list means "any file"** and defers the decision to the rest of the
   chain (workflow -> pipeline -> the chosen template's `overrides`); an omitted exclude list excludes
   nothing. A filter only ever NARROWS eligibility. A match-everything pattern in an `exclude` list
   (`*`, `**`, `*.*`, `/*`, `/**`) is REJECTED on save at every level including triggers, since exclude
   is applied last and would remove every file — leave the list empty to exclude nothing.
2. **A `requireTemplate: true` bundle either marks a default or is deliberately mutually exclusive.**
   Execute auto-selects the pipeline's default; with no default the caller _must_ name a `templateId`
   or the run is rejected. The importer promotes a **single** template to default automatically, so a
   one-template bundle needs nothing.

    With two or more templates there are two legitimate shapes, and which one applies is a product
    decision rather than an oversight:

    - **Mark one `"isDefault": true`** when one template is the ordinary choice and the others are
      variations of it. A caller then gets a working zero-argument run.
    - **Mark none**, when the templates are _mutually exclusive outputs_ and picking for the caller
      would silently produce the wrong artifact. Four bundles are this shape: `conversion-3d-basic`
      (convert-to-glb / -gltf / -obj / -stl), `rapid-pipeline` (to-glb / to-gltf), `vntana-model-ops`
      (to-glb / -gltf / -usdz) and `3dRecon-splat-toolbox` (splat-objects /
      splat-environments-360). There is no defensible default target format or capture geometry, so
      the API refusing a run that names no `templateId` is the correct behaviour.

    The consequence of the second shape is real and must be accepted knowingly: such a workflow is
    **not runnable through a zero-argument execute**, so anything that launches it — a trigger, a CLI
    script, the execute wizard — has to supply a `templateId`. Verified live: all three launch normally
    once the id is named. `infra/test/pipelines/vamsSchemaTemplateDefaults.test.ts` asserts each
    multi-template `requireTemplate` bundle is one shape or the other, so a bundle that drifts into
    "several templates, no default, no exemption recorded" fails rather than being discovered at runtime.

3. **`inputFileArity: "none"`** (a results-only or generate-from-nothing pipeline) means the manifest
   has no input files, so the asset identity comes from the execution's **output target**
   (`outputAssetId` / `outputDatabaseId`, resolved in `manifestHelper.resolve_inputs`). Do not read
   `assetId` from an input file in that case.
4. **`assetScope` accepts two vocabularies** — the CDK shorthand `{"wholeAsset": true|false}` and the
   canonical four `*Allowed` keys. Validation accepts both; a malformed value can fail the import
   while the deploy still exits 0, so confirm the row landed in `PipelineStorageTableV2` after
   deploying.
5. **A partial `systemConfig` is safe — the importer fills every field you omit with its default.**
   The stored record replaces `systemConfig` wholesale rather than merging it, so registration completes
   a bundle's block before writing: whatever the bundle declares wins, everything else becomes the
   documented default (nested maps like `assetScope` / `metadataInputs` are filled key-by-key, so
   naming one rule does not drop its siblings). Declare only what differs from the defaults. This is
   also what keeps a NEW `systemConfig` field from changing the meaning of bundles written before it
   existed.
6. **`allowWorkflowTriggerChaining` (default `false`) lets ANOTHER workflow's output fire this
   workflow's triggers** — how a preview or metadata built-in runs on a conversion pipeline's result. A
   workflow never fires on output it wrote itself whatever the value, so it cannot loop on its own
   files; a chained file must still match the trigger's `inputFileFilters`. The Potree preview, 3D
   preview thumbnail, and GenAI 3D metadata labeling bundles enable it.
7. **A workflow's `defaultOutputFileBaseExecutionPathExtension` supplies the output path prefix when
   an execution names none.** It is stored UNRESOLVED, so its `{{tag}}` placeholders resolve per run --
   one stored `/{{jobName}}/` gives every execution its own output folder. The prefix is inserted
   immediately before each output file's own name, so a container's own output folders are preserved.
   **A container must therefore not create its own per-job folder** -- the workflow prefix is what
   separates runs, and a container-side folder shows up as a stray level inside every asset. The
   Gaussian Splat and Isaac Lab bundles use it for exactly this.
8. **Let the TEMPLATE decide whether a step needs an input file.** When one pipeline supports several
   modes that differ in what they consume, set the pipeline's `inputFileArity` to the LOWEST value any
   of its templates needs (usually `none`) and let each template raise it via its `overrides`
   (`inputFileArity`, `assetScope`, `metadataInputs`, `inputFileFilters` — validated on save; unknown
   keys and bad arity values are rejected). A text-to-video template then needs no input file while an
   image-to-video template on the same pipeline overrides arity to `one`. This keeps one pipeline per
   MODEL rather than one per mode, and the execute form asks for a file only when the chosen template
   consumes one. The Cosmos 3 bundles are configured this way.
   **A workflow's `inputFileArity` is authored, not derived** — templates are chosen per execution, so
   set it to the MAXIMUM any pipeline/template combination in that workflow can require; a lower gate
   rejects a selection a template would have accepted.
   **A template's `tagSchema` is how a run gets OPERATOR-SUPPLIED options** — a prompt, a seed, an
   output format, a quality preset. Each entry declares one field the execute form renders and the API
   validates: `tagKey` (letters/digits/underscore only, so `{{tagKey}}` substitutes), `type`
   (`string` | `integer` | `number` | `boolean` | `string-list` | `enum`), `required`, `default`,
   `enumValues` (required for `enum`), and `label` / `description` for the form. Reference each one as
   `{{TAG}}` in the `configBody`; a declared tag the body never references is silently unused, so the
   operator fills in a field that reaches no pipeline.
   **In a `json` body, quoting is type-driven and checked on save:** an `integer`/`number`/`boolean`/
   `string-list` placeholder is a bare JSON value (`"seed": {{SEED}}`) while a `string`/`enum` one sits
   inside the string it fills (`"prompt": "{{PROMPT}}"`). The reverse is rejected with a 400, because
   quoting a typed tag would hand the container `"42"` where it expects `42`. Non-`json` formats
   (`yaml`, `xml`, `openjd`, `raw`) are stored verbatim and not shape-checked. A `tagKey` may not
   collide with a reserved system tag name or start with the reserved `metadata_` prefix.
9. **A workflow ref's `jobName` is an output-path segment, not a display label.** It becomes the
   `{jobName}` folder in `{baseAssetsPrefix}pipelines/{pipelineName}/{jobName}/output/{executionId}/files/`
   — relative to the area VAMS owns in the default asset bucket, which
   `executionRecords.run_bucket_key()` joins that bucket's `baseAssetsPrefix` onto; the state machine
   carries the relative form and the prefix as separate values. It is persisted
   on the workflow record as the derived `jobNames[]`, and is what `executeWorkflow` reads to
   reconstruct those prefixes at launch. Omit it in a bundle unless the pipeline id would not identify
   the step — blank already falls back to the pipeline id, keeping each step's output distinct. It
   takes the id charset only (3-63 chars), so **`{{tag}}` placeholders are rejected**: use the
   workflow's `defaultOutputFileBaseExecutionPathExtension` (rule 7) to vary the path per run. Do not
   confuse the FIELD with the `{{jobName}}` TAG, which resolves to the run's generated job name.
10. **A workflow may not list the same pipeline twice.** Per-step execute params, resolved template
    configs, and filtered inputs are all keyed by `pipelineDatabaseId:pipelineId`, so a repeated
    pipeline silently overwrites the earlier step's resolved config and both steps run identically —
    with no error. When one model needs two modes in one workflow (train then evaluate, say), ship two
    pipelines sharing a container image / ECR repo / compute environment rather than one pipeline
    listed twice with different templates.
11. **A sub-state-machine execution name must be unique at millisecond concurrency.** `openPipeline.py`
    derives the name a pipeline's own state machine runs under (`PipelineJob_<stamp>_<random>`), and
    Step Functions rejects a repeat with `ExecutionAlreadyExists`. A workflow may carry several triggers
    of one type, so one upload fans out to N simultaneous runs of the SAME pipeline — a timestamp alone
    (even to the millisecond) is not enough, so keep the random suffix. The name also namespaces
    per-execution S3 objects in some pipelines (`rp_config_{jobName}.json`), where a collision has
    concurrent runs overwrite each other's config instead of merely failing to start. Keep it within the
    80-character Step Functions limit and free of `:` and `/`.

Verify registration after a deploy rather than assuming it:

```bash
vamscli pipeline get -d GLOBAL -p {pipelineId} --json-output
vamscli pipeline template list -d GLOBAL -p {pipelineId}
```

## Registering Sub-Processes and Logs (abort, stage status, log sources)

A pipeline reports the resources it starts by putting one event on the orchestration bus.
`backend/backend/handlers/workflows/sfn/registerPipelineExecution.py` appends them to the
pipeline-execution row, and three readers consume the record: abort (`registeredSubExecutions`), the
execution details API (`availableLogs` on every step; `subExecutions` with per-stage status behind
`includeSubExecutions=true`, derived at read time from the sub-state-machine definition and history —
nothing about stages is stored), and the full-mode logs API (`logSources`, one source at a time by `logId`).

Event contract:

-   `Source` = `<ORCHESTRATION_EVENT_SOURCE_PREFIX>.execution.<executionId>.pipeline.<pipelineExecutionId>`
    (the manifest/payload already carries it). The handler ignores an event whose `Source` does not end in
    `.pipeline.<pipelineExecutionId>` for the id the detail names.
-   `DetailType` = `pipeline.execution.register`.
-   `Detail.pipelineExecutionId` (required); `Detail.subExecution` =
    `{resourceType, <locator keys>, stageName?, label?}`; `Detail.logs[]` =
    `{logGroupArn, logGroupName, logStreamName, logStreamPrefix, stageName?, label?, sourceType?}` with
    `sourceType` one of `stateMachine | lambda | batch | ecs | container | custom` (anything else is stored
    as `custom`).
-   Validators: `CLOUDWATCH_LOG_GROUP_ARN`, `CLOUDWATCH_LOG_GROUP_NAME`, `LOG_STREAM_NAME` (also the prefix
    — no `:` or `*`), `SFN_STATE_NAME` (`^[^\x00-\x1f\x7f]{1,80}$`), `DISPLAY_LABEL` (same class, 1–128),
    `LOG_SOURCE_TYPE`. An invalid optional field is dropped with a warning; the entry survives.
-   Dedup is by location `(logGroupArn without ":*" | logGroupName, logStreamName, logStreamPrefix)`; a
    redelivered entry that carries `stageName` / `label` / `sourceType` the stored one lacks merges them in.
    At most 50 logs and 50 sub-executions are stored per pipeline execution. Registration stays best-effort:
    wrap `put_events` so a failure logs and returns.

What every built-in emits (the `register_sub_execution` helper in each `openPipeline.py`, or
`register_batch_job` in `executeBatchJob.py`):

-   The state-machine log entry with `sourceType: "stateMachine"`, `label: "<pipeline> state machine"`; the
    `subExecution` with `label: "<pipeline> processing"`.
-   One **container** entry per Batch state, only when the env vars are present:

    ```python
    {"logGroupArn": BATCH_JOB_LOG_GROUP_ARN, "logGroupName": BATCH_JOB_LOG_GROUP_NAME,
     "logStreamPrefix": f"{<job definition name>}/default/", "stageName": "<Batch state name>",
     "sourceType": "batch", "label": "<Batch state name> container"}
    ```

    Which group that is depends on the compute family. The five **Fargate** job definitions (coordinate
    transform, Blender renderer, 3D thumbnail, PDAL, Potree) write through the `awslogs` driver to a
    VAMS-owned, KMS-encrypted `/aws/vendedlogs/Pipelines/<Name><hash>` group, and register **that** group;
    the **GPU** Batch pipelines (NVIDIA Cosmos, GR00T, Isaac Lab, Splat Toolbox) set no log configuration,
    so theirs is AWS Batch's default `/aws/batch/job`. In both families the stream is
    `<jobDefinitionName>/default/<ecs-task-id>` — the Fargate construct sets `awslogs-stream-prefix` to the
    physical (hashed) job definition name, the same string the producer's derived `*_JOB_DEFINITION_NAME`
    resolves to. Container output does not print the VAMS execution ids, so a registered prefix that the
    real stream falls under is what lets the read skip the execution-scope terms; a prefix the stream does
    not start with reads as `scoped`, and the filter drops every container line.

CDK rules for those env values (`infra/lib/nestedStacks/pipelines/**`):

-   The registering lambda's environment spreads one of the two helpers in
    `infra/lib/helper/batchJobLogGroup.ts`: `...vendedBatchJobLogGroupEnvironment(containerLogGroup)` for a
    Fargate pipeline (the group its `BatchFargatePipelineConstruct` was given as `logGroup`; the Potree
    builder sets `PDAL_` / `POTREE_JOB_LOG_GROUP_NAME` / `_ARN` inline because its two jobs write to two
    groups), or `...batchJobLogGroupEnvironment()` for a GPU pipeline — `BATCH_JOB_LOG_GROUP_NAME =
"/aws/batch/job"` and `BATCH_JOB_LOG_GROUP_ARN = IAMArn(BATCH_JOB_LOG_GROUP_NAME).loggroup`, the colon
    `log-group:` form. The helper is the one place the default group is named: never spell the literal or
    derive the ARN in a builder, and never `formatArn(..., ArnFormat.SLASH_RESOURCE_NAME)`, which renders
    `log-group//aws/batch/job`, fails `CLOUDWATCH_LOG_GROUP_ARN`, and degrades the entry to a name-only
    read. Beside those two, the builder sets
    the job-definition env its producer reads (`BATCH_JOB_DEFINITION_NAME` for an `openPipeline.py`, `PDAL_` /
    `POTREE_JOB_DEFINITION_NAME` for the two Potree states, the pre-existing `BATCH_JOB_DEFINITION` for an
    `executeBatchJob.py`) from the `{ jobDefinitionName }` the construct passes it — a per-builder
    `OpenPipelineBatchLogProps` interface in the `openPipeline` builders.
-   `jobDefinitionName` is the job definition **name**, never its ARN: Fargate `EcsJobDefinition` →
    `.jobDefinitionName` (a token that resolves to the hashed name); GPU `CfnJobDefinition` given a
    `jobDefinitionName` prop → that same string (never `.ref`, the ARN with revision); an unnamed
    `CfnJobDefinition` → `jobDefinitionNameFromRef(jobDef.ref)` from the same helper. An ARN in
    `logStreamPrefix` contains `:`, fails `LOG_STREAM_NAME`, and leaves the container source permanently
    `unscoped`.
-   `stageName` must equal the ASL state name, which is the CDK construct id because no construct sets
    `stateName`. `infra/test/pipelines/batchLogRegistrationEnvFargate.test.ts`,
    `batchLogRegistrationEnvGpu.test.ts` and `containerLogRegistrationEnvEcs.test.ts` synthesize each
    pipeline construct (`infra/test/support/pipelineConstructHarness.ts`), parse its ASL
    (`infra/test/support/asl.ts`) and assert that every module-level `*_STATE_NAME = "…"` literal the
    producer declares (`declaredStageNames`; for cosmos, the `COSMOS_BATCH_STATE_NAME` env value) is a
    key of `States`; renaming a Batch construct without changing the producer's `stageName` fails those
    tests instead of silently breaking the stage ⇄ history join. Declare a stage name as a module-level
    string literal, never built at runtime, or the test cannot see it.
-   A `lambda` entry uses `` `/aws/lambda/${fn.functionName}` `` with `IAMArn(name).loggroup`; never read
    `fn.logGroup` on a function without an explicit log group (it synthesizes a `Custom::LogRetention`
    resource).

Producer tests (`lambda/tests/test_manifest_refactor.py`, `test_batch_job_registration.py`) assert the new
keys and that the emitted `logGroupArn` passes `CLOUDWATCH_LOG_GROUP_ARN` and `logStreamPrefix` passes
`LOG_STREAM_NAME`. Run each pipeline's `tests/` in its own pytest process (module-name collision across
pipelines).

## Every boto3 Client Carries the Adaptive Retry Configuration

Every `boto3.client(...)` and `boto3.resource(...)` in this tree takes the shared retry configuration:

```python
from botocore.config import Config

retry_config = Config(retries={'max_attempts': 5, 'mode': 'adaptive'})

s3 = boto3.client('s3', config=retry_config)
sfn_client = boto3.client('stepfunctions', config=retry_config)
```

This is the same rule as `backend/CLAUDE.md` Rule 6, and it applies here for a sharper reason: a
pipeline Lambda or container runs against Step Functions, Amazon S3, and EventBridge for the length of a
job — hours, on the GPU pipelines — so a bare client sits on botocore's default retry mode with no
client-side rate limiting, and a sustained burst surfaces as a throttling error on the caller instead of
being smoothed.

Declare the constant **above the first client**, not merely after the imports. A handful of modules here
interleave imports with executable code (`multi/rapidPipelineEKS/lambda/consolidated_handler.py` builds
a client and then keeps importing), and a constant placed after the last import lands below the client
that uses it — `NameError` at module import, which in a Lambda is a cold-start 500 on every request.

**A deliberate departure needs a comment saying why.** A call that must NOT retry — a non-idempotent
operation where a retry would duplicate work — is a legitimate exception, but it has to be
distinguishable from an oversight, because the ratchet cannot tell them apart.

`backend/tests/common/workflows/test_pipeline_boto_clients_configured.py` holds this at zero bare
clients. It is scoped to `backendPipelines/`; the separate ratchet for `backend/backend` does not cover
this tree, which is how 137 bare clients across 68 files sat invisible behind a green "59 → 0" for the
API handlers.

## Test Conventions

Pipeline tests run with pytest's defaults — there is no `pytest.ini` anywhere under `backendPipelines/`,
which is why the existing `@pytest.mark.unit` is unregistered and only warns.

**A rule that must hold for EVERY pipeline goes in `backendPipelines/tests/`.**
`test_open_pipeline_extension_gates.py` is the worked example: it loads all seven `openPipeline.py`
handlers by path under per-pipeline module names and asserts each tests EXACT membership of its parsed
`ALLOWED_INPUT_FILEEXTENSIONS` list. `in` against the joined env string is substring containment, which
admits any prefix of a listed extension (`.us` passes for `.usd,.usda`), and the loose form spread by
copying an existing pipeline — which is precisely what a per-pipeline test cannot catch. Two of seven
were fixed and five were not, and no per-pipeline suite noticed.

**Give every test module a suite-private basename.** Pipelines are near-copies of one another, so their
test files collide: `test_extension_gate.py`, `test_manifest_refactor.py`,
`test_construct_pipeline_failure_reporting.py`, `test_open_pipeline_function_error.py`,
`test_pipeline_end_token_routes.py`, `test_output_relative_subdir.py` and both `conftest.py` files under
`pcPotreeViewer` each exist in two or more pipelines. No tests directory carries an `__init__.py`, so a
single pytest process that collects two same-named modules ERRORS at collection with `import file
mismatch` — which is why `pytest backendPipelines/` cannot be run as one command today, and why the
suites have to be invoked one directory at a time. Worse than the inconvenience: with a different
invocation order a suite can import ANOTHER pipeline's same-named module and assert against the wrong
file while passing. Prefix a new file with its pipeline (`test_splat_extension_gate.py`).

**A suite that reads configuration from the process environment must restore it.**
`splatToolbox/lambda/tests/test_extension_gate.py` does `os.environ.setdefault` at import and then calls
its loader with no argument, so any other suite in the same process that sets
`ALLOWED_INPUT_FILEEXTENSIONS` without restoring it makes splat's tests fail on unrelated assertions.
Pass the value explicitly on every load, or restore what you changed.

**Mark a temporary test.** A test written to prove one specific change landed — a deleted container file, a
removed helper, a dead branch — carries `@pytest.mark.temporary` plus a line naming what it pins, so release
cleanup can find it with `pytest -m temporary --collect-only` (root `CLAUDE.md` Rule 13). It is required
because a temporary test and a durable guardrail are indistinguishable by reading them afterwards: both scan
source, both assert an absence, both explain themselves.

Do **not** mark a test whose subject stays writable. Several structural tests here are durable and must keep
holding:

-   the pinned-revision checks over each `Dockerfile` (a `--build-arg …_COMMIT=main` still builds green)
-   the `manifestHelper.py` byte-identity check across every vendored copy
-   the `customLogging/logger.py` byte-identity check and the emitted-line token redaction checks
    (`backendPipelines/tests/test_pipeline_logger_identity.py`, `test_pipeline_logger_formatter.py`) — a
    handler can always be edited back to an f-string event log or a drifting logger copy
-   the `_run_streaming` no-drift comparison across the four deployable NVIDIA containers
-   `test_container_file_inventory.py`'s `from .utils` scan — the package is `vams_utils`, so that import
    fails at container **runtime** on a GPU Batch job, invisible to the image build and to CDK synth

The shortcut "the forbidden literal appears nowhere in the source, so the test is spent" is **wrong** — a
forbid-forever guardrail also has zero occurrences, and that absence is the guard working.

## A Non-Root Container Normalizes Read Bits on the Source It COPYs

A pipeline container that drops to a non-root `USER` runs `RUN chmod -R a+rX <path>` on the line
immediately after every `COPY` of source it will read:

```dockerfile
COPY ./preview_pipeline /app/preview_pipeline
RUN chmod -R a+rX /app/preview_pipeline
...
USER appuser
```

`docker COPY` preserves the build host's umask and copies the files root-owned. A hardened build host —
umask `077` (the STIG default on RHEL and Amazon Linux) or `027` (locked-down CI) — therefore produces
root-owned `600`/`640` files the non-root user cannot read, and Python raises
`PermissionError: [Errno 13]` at import. It is NOT `ModuleNotFoundError`: the directory is recreated
`0755`, so the package is traversable and found, but its files are unreadable. This is invisible at build
time — the image builds green and fails only at container runtime, and only for images built on such a
host. The Batch job definition sets no user override, so the image's `USER` is what runs.

`a+rX` grants world-read on files and traverse on directories without adding an execute bit to plain
files. `COPY --chown=<user>` is NOT sufficient — it makes the file readable by that one owner while the
mode stays restrictive. Worked examples: `preview/3dThumbnail/container/Dockerfile` and
`conversion/coordinateTransform/container/Dockerfile`.

## Adding a New Processing Pipeline

1. Create directory under `backendPipelines/{useCase}/`.
2. Add Lambda handler in `lambda/` subdirectory. **Every pipeline `lambda/` directory MUST include:**

    - `__init__.py` (package marker)
    - `customLogging/__init__.py` (package marker)
    - `customLogging/logger.py` (copy from any existing pipeline, e.g., `backendPipelines/3dRecon/splatToolbox/lambda/customLogging/logger.py`)

    Without these files, Lambda will fail at import time with `No module named 'customLogging'`. The Lambda layer provides a fallback, but the local `customLogging/` package is required in each pipeline's code asset.

    **All `customLogging/logger.py` copies must stay byte-identical** — verify with
    `find backendPipelines -path '*/lambda/customLogging/logger.py' -exec md5sum {} \; | awk '{print $1}' | sort -u`,
    which must print exactly one hash. The logger is where task tokens are redacted from every log line,
    so a copy that drifts is a pipeline whose CloudWatch stream carries a bearer credential. Edit one
    copy, then propagate to the rest in the same change, and add the new pipeline's path to the
    `LOGGER_COPIES` tuple in `backendPipelines/tests/test_pipeline_logger_identity.py`, which pins the
    single digest, fails on an unlisted copy, and is the list `test_pipeline_logger_formatter.py` reads
    too. Log an
    event as a structured field (`logger.info("Event", event=event)`), never as an f-string, and never
    log a task token on its own.

    A pipeline that reads the workflow manifest also vendors `manifestHelper.py`. **All copies must stay
    byte-identical** — verify with
    `find backendPipelines -name manifestHelper.py -exec md5sum {} \; | awk '{print $1}' | sort -u`,
    which must print exactly one hash. Edit one copy, then propagate to the rest in the same change.
    `fetch_manifest` RAISES when a referenced manifest cannot be read (it returns `None` only when the
    payload references no manifest at all): the manifest is the sole carrier of asset identity and the
    output paths, so swallowing that error starts a job that fails only after its compute — a GPU
    instance for the NVIDIA pipelines — is provisioned. Keep the token capture ahead of
    `resolve_pipeline_inputs` in every `vamsExecute` handler so the raise reaches `SendTaskFailure`.

3. Add container if needed in `container/` subdirectory. **A container that drops to a non-root `USER`
   runs `RUN chmod -R a+rX <path>` on the line immediately after every `COPY` of source it reads** —
   `COPY` preserves the build host's umask, and a hardened host's root-owned `600`/`640` files raise
   `PermissionError` at import under the non-root user, invisible to the build. See
   [A Non-Root Container Normalizes Read Bits on the Source It COPYs](#a-non-root-container-normalizes-read-bits-on-the-source-it-copys).
4. **Register the pipeline's sub-process and log sources** from `openPipeline.py` (or from
   `executeBatchJob.py` when that lambda submits the job itself): the state-machine log entry with
   `sourceType`/`label`, the `subExecution` with `label`, and one container entry per Batch state naming the
   group that job definition writes to — the pipeline's vended `/aws/vendedlogs/Pipelines/<Name><hash>` group
   for a Fargate job, `/aws/batch/job` for a GPU job with no log configuration — (`logStreamPrefix`
   `"<jobDefinitionName>/default/"`, `stageName` = the ASL state name, declared as a
   module-level `*_STATE_NAME` literal). The builder supplies `ORCHESTRATION_BUS_NAME` +
   `orchestrationBus.grantPutEventsTo(fun)`, `STATE_MACHINE_LOG_GROUP_NAME` / `_ARN`,
   `...vendedBatchJobLogGroupEnvironment(logGroup)` (Fargate) or `...batchJobLogGroupEnvironment()` (GPU)
   and the job-definition-name env the producer reads
   (`BATCH_JOB_DEFINITION_NAME`). Add the construct to
   `infra/test/pipelines/batchLogRegistrationEnv{Fargate,Gpu}.test.ts` (or
   `containerLogRegistrationEnvEcs.test.ts` for an ECS task) and assert the emitted entries in
   `lambda/tests/test_manifest_refactor.py`. Without it, abort leaves the compute running and the execution
   shows no stages or container logs. See
   [Registering Sub-Processes and Logs](#registering-sub-processes-and-logs-abort-stage-status-log-sources).
5. **Author the `vamsSchema/` bundle** (`pipeline.json`, plus `workflow.json` and `templates/` as
   needed) and register it with the `VamsSchemaRegistration` construct from the pipeline's nested
   stack. Without this the AWS resources deploy but nothing appears in VAMS. Each
   `templates/{templateId}.json` carries the `configBody` the container receives and the `tagSchema`
   declaring the per-run options an operator sets on the execute form — that is where a prompt, a seed,
   or a target format belongs, rather than hardcoded in the body. See
   [vamsSchema Registration](#vamsschema-registration-how-a-built-in-becomes-usable).
6. Create CDK nested stack in `infra/lib/nestedStacks/pipelines/`.
7. Add pipeline config to `config.ts` under `pipelines` section.
8. Register in pipeline builder nested stack.
9. Add feature switch if pipeline is optional.
10. **Add pipeline flag to VPC builder** (`infra/lib/nestedStacks/vpc/vpcBuilder-nestedStack.ts`). A pipeline using AWS Batch, ECS, or Fargate goes into some of the VPC builder's three condition blocks — **which ones depends on the subnets its compute runs in.** Decide that first from what `pipelineBuilder-nestedStack.ts` passes as its `pipelineSubnets`: `pipelineNetwork.isolatedSubnets.pipeline` or `pipelineNetwork.privateSubnets.pipeline`. Search for `useSplatToolbox` (private) and `usePreview3dThumbnail` (isolated) to see both treatments.

    - **Subnet creation condition** (~line 343) — the `if` block that pushes `subnetPublicConfig` and `subnetPrivateConfig` into `subnetConfigurations`. **Private-subnet pipelines only.** Omit it for one and its Batch compute environment fails with "Resource subnets are required"; add an isolated-subnet pipeline and CDK creates public subnets plus **one NAT gateway per Availability Zone** (~$66/month at two AZs, plus data processing) that the pipeline never routes through, because `subnetPrivateConfig` is `PRIVATE_WITH_EGRESS` and the `ec2.Vpc` sets no `natGateways`.
    - **Pipeline-only endpoint condition** (~line 651) — the `if` block that creates Batch, ECR API, and ECR Docker interface VPC endpoints in the isolated subnets. **Required for every pipeline, either placement.** Without it Batch jobs cannot pull container images.
    - **ECS endpoint condition** (~line 736) — the `needsEcsPrivate` variable. **Private-subnet pipelines only.** This is the ECS _control-plane_ endpoint that the ECS agent on an EC2-launch-type container instance needs; **Fargate tasks do not use it** (they need ECR, Amazon S3 and CloudWatch Logs, supplied by the block above). Each endpoint adds one ENI per AZ, ~$15/month.

    Six pipelines run in isolated subnets today (3dBasic, CAD/mesh metadata extraction, Potree viewer, 3D thumbnail, GenAI metadata labeling, coordinate transform) and appear in the endpoint block only; four run in private subnets (Splat Toolbox, NVIDIA Cosmos, NVIDIA GR00T, Isaac Lab training) and appear in all three. Regression coverage asserting both directions: `infra/test/pipelines/coordinateTransformVpcPlacement.test.ts`.

11. **Pass through all output paths** in the `vamsExecute` lambda — never hardcode empty strings for `outputS3AssetFilesPath`, `outputS3AssetPreviewPath`, or `outputS3AssetMetadataPath`. See [Pipeline S3 Output Paths](#pipeline-s3-output-paths) for conventions.
12. **Use the correct output path** in the `constructPipeline` lambda for the container's output target: `outputS3AssetFilesPath` for file-level outputs (including `.previewFile.X` thumbnails), `outputS3AssetPreviewPath` for asset-level previews only, `outputS3AssetMetadataPath` for metadata. Only use `inputOutputS3AssetAuxiliaryFilesPath` for temporary files or special non-versioned viewer data (e.g., Potree octree files).
13. **Preserve relative paths** in container output. When writing asset-adjacent files (e.g., `.previewFile.X`), the container must maintain the input file's relative subdirectory within the asset so process-output can locate outputs correctly. See [Pipeline S3 Output Paths](#pipeline-s3-output-paths) for the derivation pattern.
14. **Update `documentation/docusaurus-site/docs/deployment/configuration-reference.md`** with all new pipeline configuration options (`enabled`, `autoRegisterWithVAMS`, `autoRegisterAutoTriggerOnFileUpload`, and any pipeline-specific settings). Follow the existing format: `-   \`app.pipelines.{pipelineName}.{option}\` | default: {value} | #{description}`.
15. **Update licenses/attributions** when a pipeline adds, removes, or changes a third-party model, container base image, or dependency with its own license. Update **both** `NOTICE.md` (the per-pipeline dependency table + attribution note at the repo root) **and** `documentation/docusaurus-site/docs/additional/notices.md` (the per-pipeline license paragraph + the closing attribution/reference list). Record the exact license (e.g. NVIDIA Open Model License, OpenMDW-1.1, Apache-2.0) and any required attribution string. A model under a new license (as Cosmos 3 uses OpenMDW-1.1 rather than the NVIDIA Open Model License) must be reflected in both files, and the pipeline doc page's Prerequisites + Attribution sections.
16. **Update the pipeline count** wherever the docs state one. `documentation/docusaurus-site/docs/overview/features.md` opens the "Built-In Pipelines" section with a spelled-out count ("VAMS includes _fourteen_ built-in processing pipelines…") followed by a table — bump the number to match the new table row count when adding or removing a pipeline. Grep the docs for the current number word to catch any other count references.
17. **Update the root `CLAUDE.md`** — its Project Overview pipeline list and its directory tree both enumerate the pipeline set, and the tree's box-drawing glyphs assert which directory is a parent's last child, so a new sibling left out reads as an assertion that it does not exist. Required by root `CLAUDE.md` Rule 11 ("New pipeline"). Also add a row to `documentation/docusaurus-site/docs/pipelines/` and its `pipelines/overview.md` table per `documentation/CLAUDE.md`.
18. **A pipeline built by AWS CodeBuild pushes and consumes ONE content-addressed tag.** The construct
    supplies `IMAGE_TAG` to the project's `environmentVariables` from `sourceAsset.assetHash`, and the
    pull site names that same value — a shared compute construct takes it as one prop together with the
    repository (`ecrImage` on `batch-fargate-pipeline.ts`, `codeBuildImage` on `batch-gpu-pipeline.ts`)
    so the tag cannot be omitted while the repository is supplied. The buildspec must NOT default
    `IMAGE_TAG`; it fails the build when the project supplied none, because a default pushes a tag the
    Batch job definition does not name and the deploy still reports success. `:latest` is pushed
    alongside solely as the `--cache-from` alias, since a content-addressed tag never pre-exists and a
    cold cache adds hours to a GPU image build. Copying an existing buildspec is how this regresses;
    `infra/test/pipelines/codeBuildImageTagCoordination.test.ts` and the immutable-tag block of
    `infra/test/pipelines/containerBuildSources.test.ts` assert both halves.
19. **A container that clones an upstream repository at build time clones a FIXED revision.** Declare
    the revision as an `ARG <NAME>_COMMIT=<40-hex>`, `git checkout --detach` it, and verify it landed
    with `test "$(git rev-parse HEAD)" = "${<NAME>_COMMIT}"` **in the same `RUN`** — a checkout in a
    later instruction is a different layer and pins nothing. Write the resolved id to a file in the
    image and echo it from the entrypoint, so a run's log names the code it ran. Without the
    verification the pin is decorative: `--build-arg <NAME>_COMMIT=main` checks out a moving ref and
    the build still succeeds. The same applies to a `COPY --from=<image>:<tag>` and to any installer
    URL that omits a version. Coverage: the NVIDIA block of
    `infra/test/pipelines/containerBuildSources.test.ts`, which lists the Dockerfiles explicitly
    because a `**/Dockerfile` glob passes locally and fails in CI on the gitignored splat Dockerfile.

20. **A directory containing `.synced-commit` is overwritten from upstream on every `cdk synth` — and
    on every `cdk list`.** `SplatToolboxConstruct.syncContainerSources` clones the pinned commit and
    copies **every** upstream file over `backendPipelines/3dRecon/splatToolbox/container/`. Editing one
    of those files does not stick: the change survives until the next CDK invocation and is then gone,
    with `git status` clean afterwards because the restored copy matches `HEAD`.

    **`.gitignore` does not tell you which files are upstream's.** It lists only `Dockerfile`,
    `/src/*`, `LOCAL_DEBUG_README.md` and `.synced-commit`, so a tracked file like
    `build_models_tar.py` looks VAMS-owned and is not. What actually distinguishes them is whether the
    file exists upstream, and the observable proxy is the mtime: after a sync, upstream's files carry
    the marker's timestamp while VAMS additions (`__main__.py`, `vams_utils/`, `vams_bake_models.py`)
    keep their own.

    To change behaviour in an upstream-owned file, follow the pattern the sync already uses for the
    Dockerfile: a **programmatic injection** applied after the copy, anchored on a pattern, that
    throws when the anchor is missing rather than silently no-op'ing. Reserve that for something the
    deployed pipeline executes — a workstation-only helper is not worth a brittle anchor, and the
    honest alternative is to scope the rule that flags it (see
    `backend/tests/common/workflows/test_pipeline_boto_clients_configured.py`, whose exemption is a
    DENY list precisely so a future upstream file turns the rule red instead of vanishing from it).

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.