secedgar-mcp-server
cyanheads/secedgar-mcp-server/AGENTS.md
Server: secedgar-mcp-server Version: 0.15.9 Framework: @cyanheads/mcp-ts-core ^0.13.6 Engines: Bun ≥1.4.0, Node ≥24.0.0 Query SEC EDGAR filings, XBRL financials, and company data through MCP. Read-only, no API keys required. Full design: docs/sec-edgar-mcp-design.md. Read the framework docs first: node_modules/@cyanheads/mcp-ts-core/CLAUDE.md contains the full API reference — builders, Context, error codes, exports, patterns. This file covers server-specific conventions only. When the user asks what's next or needs direction, suggest options based on the current project state. Common next steps: Tailor suggestions to what's actually…
- Reads credentials
# Agent Protocol
**Server:** secedgar-mcp-server
**Version:** 0.15.9
**Framework:** [@cyanheads/mcp-ts-core](https://www.npmjs.com/package/@cyanheads/mcp-ts-core) `^0.13.6`
**Engines:** Bun ≥1.4.0, Node ≥24.0.0
Query SEC EDGAR filings, XBRL financials, and company data through MCP. Read-only, no API keys required. Full design: `docs/sec-edgar-mcp-design.md`.
> **Read the framework docs first:** `node_modules/@cyanheads/mcp-ts-core/CLAUDE.md` contains the full API reference — builders, Context, error codes, exports, patterns. This file covers server-specific conventions only.
---
## Core Rules
- **Logic throws, framework catches.** Tool/resource handlers are pure — throw on failure, no `try/catch`. Plain `Error` is fine; the framework catches, classifies, and formats. Use error factories (`notFound()`, `validationError()`, etc.) when the error code matters.
- **Use `ctx.log`** for request-scoped logging. No `console` calls. Attach operation-specific fields with `withExtra(ctx, { … })` — `RequestContext` is closed, so a bare `{ ...ctx, field }` spread no longer type-checks.
- **Use `ctx.state`** for tenant-scoped storage. Never access persistence directly.
- **Ask for missing input** with `return ctx.requestInput(...)`, reading the answers back through `ctx.inputs` when the handler is re-entered. Never await a caller mid-handler.
- **Secrets in env vars only** — never hardcoded.
- **Close the loop on issues.** When implementing work tracked by a GitHub issue, comment on the issue with what landed and close it. Do both — a comment without a close leaves stale issues open; a close without a comment leaves no record of what shipped. The comment is for future readers — state the concrete changes, not the conversation that produced them.
---
## What's Next?
When the user asks what's next or needs direction, suggest options based on the current project state. Common next steps:
1. **Add tools/resources/prompts** — scaffold new definitions using the `add-tool`, `add-app-tool`, `add-resource`, `add-prompt` skills
2. **Add services** — scaffold domain service integrations using the `add-service` skill
3. **Add tests** — scaffold tests for existing definitions using the `add-test` skill
4. **Field-test definitions** — exercise tools/resources/prompts with real inputs using the `field-test` skill, get a report of issues and pain points
5. **Run the `security-pass` skill** — audit handlers for MCP-specific security gaps: output injection, scope blast radius, input sinks, tenant isolation
6. **Run the `polish-docs-meta` skill** — finalize README, CHANGELOG, metadata, and agent protocol for shipping
7. **Run the `maintenance` skill** — investigate changelogs, adopt upstream changes, and sync skills after `bun update --latest`
Tailor suggestions to what's actually missing or stale — don't recite the full list every time.
---
## MCP Surface
### Tools
| Name | Description | Key Inputs |
|:-----|:------------|:-----------|
| `secedgar_company_search` | Find companies and retrieve entity info with optional recent filings | `query`, `include_filings?`, `forms?`, `filing_limit?`, `filed_after?`, `filed_before?` |
| `secedgar_search_filings` | Search EDGAR filings since 1993 — full-text (2001+) plus archive-backed browse for pre-2001 ranges | `query?`, `forms?`, `filed_after?`, `filed_before?`, `limit?` |
| `secedgar_get_filing` | Fetch a specific filing's metadata and document content | `accession_number`, `cik?`, `content_limit?`, `document?` |
| `secedgar_get_financials` | Get historical XBRL financial data for a company | `company`, `concept`, `taxonomy?`, `period_type?`, `limit?` |
| `secedgar_get_snapshot` | One-call financial profile — latest value of every supported concept from a single companyfacts read | `company`, `taxonomy?`, `period_type?` |
| `secedgar_get_material_events` | 8-K filings with item codes decoded, filterable by item | `company`, `items?`, `filed_after?`, `filed_before?`, `limit?` |
| `secedgar_get_insider_transactions` | Form 4 / 4-A insider transactions parsed from ownership XML | `company`, `transaction_type?`, `limit?`, `filed_after?`, `filed_before?` |
| `secedgar_get_institutional_holdings` | 13F-HR quarterly institutional holdings parsed from the information table | `company`, `quarter?`, `limit?`, `offset?`, `consolidate?` |
| `secedgar_find_holders` | Reverse 13F lookup — which institutional managers reported holding an issuer | `issuer`, `cusip?`, `quarter?`, `limit?` |
| `secedgar_get_beneficial_owners` | 5%+ blockholders of an issuer from structured SCHEDULE 13D/13G XML | `issuer`, `form_kind?`, `include_amendments?`, `limit?` |
| `secedgar_get_fund_holdings` | Fund portfolio holdings from the quarterly NPORT-P report | `fund`, `series_id?`, `report_date?`, `limit?`, `offset?` |
| `secedgar_fetch_frames` | Fetch SEC XBRL frames for one concept × one period across all reporting companies | `concept`, `period`, `taxonomy?`, `unit?`, `limit?`, `offset?`, `sort?` |
| `secedgar_compare_companies` | Period-aligned comparison of 2-10 named companies across 1-8 concepts | `companies`, `concepts`, `taxonomy?`, `period_type?`, `periods?` |
| `secedgar_search_concepts` | Discover supported XBRL concept names or reverse-lookup a raw tag | `search?`, `group?`, `taxonomy?` |
| `secedgar_dataframe_describe` | List canvas dataframes with provenance, TTL, and schema | `name?` |
| `secedgar_dataframe_query` | Run a single-statement SELECT across dataframes (DuckDB SQL) | `sql`, `preview?`, `register_as?` |
| `secedgar_dataframe_drop` | Drop a canvas dataframe by name. Opt-in via `EDGAR_DATAFRAME_DROP_ENABLED=true` | `name` |
### Resources
| URI | Description |
|:----|:------------|
| `secedgar://concepts` | Common XBRL financial concepts grouped by statement, mapping friendly names to XBRL tags |
| `secedgar://filing-types` | Common SEC filing types with descriptions, cadence, and use cases (NPORT-P included), plus the 8-K item-code decode tables for both numbering regimes |
### Prompts
| Name | Description |
|:-----|:------------|
| `secedgar_company_analysis` | Guides structured analysis of a company's SEC filings |
---
## Domain Notes
- **Rate limit:** 10 req/s per IP. Every SEC request goes through one process-wide `createPacer` queue in `EdgarApiService` that spaces request *starts* `1000 / EDGAR_RATE_LIMIT_RPS` ms apart, with no concurrency cap and no wait budget — an ordinary burst drains at the configured rate and is never shed, and this server's own traffic can never exceed it, so `rateLimitRps` stays at 10 and headroom for a co-tenant on the outbound IP is the operator's call via `EDGAR_RATE_LIMIT_RPS`. A 429 is **not** retried: SEC holds the block until the request rate has stayed under the threshold for ten minutes and every request sent meanwhile restarts that clock, so the ~3s retry budget (`MAX_RETRIES` 3 × `BASE_BACKOFF_MS` 1000, exponential) cannot outlast it (#112). Instead a 429 closes the service's **block gate** for `EDGAR_RATE_LIMIT_COOLDOWN_SECONDS` (default 600). While it is closed nothing is sent: a call fails locally as `RateLimited` with `retryAfter` set to the whole seconds remaining (never below 1), and the gate is re-read at dispatch, so a call already queued in the pacer when the 429 landed — or a 5xx retry sleeping through its backoff — is refused too. The first call after the cool-down is sent alone as a probe while every other call waits on it: any HTTP answer but a 429 reopens normal pacing, a probe 429 closes the gate for the full cool-down again, and a probe that gets no HTTP answer (a network error) proves nothing, so the next call probes in its place. **The gate lives in the service, not in the pacer's `cooldown` option** — that gate holds queued calls until it reopens and then releases them all at the pacing rate, so it can neither refuse locally nor keep the reopening to one request. The upstream 429 and the local refusal carry one `data.reason: 'rate_limited'`, declared (`thrownBy: 'service'`, `retryable: true`) on every tool that reaches SEC, plus `data.retryable: true` — `ctx.fail` copies a contract entry's `retryable` onto `data`, a service throw does not, so both throw sites set it themselves (#122) — and one recovery hint naming SEC's ten-minute window and the seconds to wait; SEC's HTML block page stays out of the payload (`captureBody: false`). Mirror-served reads (the ticker base behind `resolveCik`, and `tryGetCompanyFacts` / `tryGetCompanyConcept` / `tryGetFrames` through `mirrorOrLive`) never reach `rawFetch`, so they keep answering during a block. 500/502/503/504 still retry with the backoff. `createApp({ teardown })` disposes the pacer through `disposeEdgarApiService()` (#116)
- **Parameter names:** one name per concept across tools — `company` (the entity a tool reads; on `get_institutional_holdings` that is the 13F filer), `filed_after` / `filed_before` (inclusive filing-date bounds), `forms`. Every tool carrying one of those keys declares `inputAliases` for the other spellings in use — `ticker`, `cik`, `ticker_or_cik` → `company`; `start_date`, `date_from` → `filed_after`; `end_date`, `date_to` → `filed_before`; `form_types` → `forms` — so a caller carrying a name from one tool to the next is not rejected (#115). An alias is declared only where it maps to exactly one key of that tool and its value shape fits: a singular `form_type` cannot fill the `forms` array, so it is left out rather than widening the schema. `issuer`, `fund`, `companies`, and `query` name different roles and are not aliased. `search_filings` still filters by date only with both bounds; `company_search`, `get_material_events`, and `get_insider_transactions` accept either alone
- **User-Agent required:** `"AppName contact@email.com"` on every SEC request or IP gets blocked
- **CIK zero-padding:** URLs require 10-digit zero-padded CIK (`String(cik).padStart(10, '0')`)
- **ETF/mutual-fund tickers:** `loadTickerCache` fetches `company_tickers_mf.json` alongside `company_tickers.json` and merges fund symbols (VOO, SCHD, JEPI…) into `byTicker`. Fund CIKs are registrant trusts (1:many with series), so fund rows go into `byTicker` only — never `byCik`, whichever source they come from. Resolved fund matches carry `seriesId`/`classId` from the live MF file, the only source of either. On a mirror-served ticker base the live MF fetch is merged in when `EDGAR_MIRROR_FALLBACK_LIVE` is on, so fund tickers resolve even if the mirror predates MF ingestion (#43); strict mirror-only (`FALLBACK_LIVE=false`) uses the mirror's own MF rows (ticker→CIK, no series/class). The mirror stores a fund symbol with an empty name — its schema has no series column — and `mirrorIndexRow` reads that empty name as the fund marker; a symbol claims `byTicker` in the order equity → live fund → mirror fund, so a live entry supersedes the mirror's bare row and keeps its series (#135). **The live fund slice has its own failure window:** a failed MF load (any HTTP failure, a malformed body, the block gate's refusal) is not cached for the ticker-cache TTL — `loadMfTickers` returns no rows and a `retryAt` `FUND_SLICE_RETRY_MS` (60 s) out, or the error's `retryAfter` when longer, so nothing is retried while the gate is closed. The first resolution after it refetches the file once, single-flight, onto the cached base (`reloadFundSlice`), which keeps its own load time: the equity index is not refetched and its TTL does not move. A per-resolution retry was rejected — against a persistent 5xx it adds three requests and ~3 s of backoff to every lookup (#119). Each failed load logs one warning through the framework's global `logger` (no request context to attach to) carrying the URL, the HTTP status or `malformed_body`, and the retry delay (#43).
- **Name and ticker resolution:** `resolveCik` walks numeric CIK → the `/^[A-Z]{1,5}$/` ticker fast path → the name passes (exact → prefix → substring over `allEntries`) → a catch-all `byTicker` lookup for symbols whose shape skipped the fast path. Two normalizations hang off that chain. The name passes' **exact** tier also compares suffix-normalized forms (`normalizeCompanySuffix`): the terminal whitespace-delimited token only, a trailing `.`/`,` stripped as part of recognizing it, expanded to one of four canonical long forms — `corp`→`corporation`, `inc`→`incorporated`, `co`→`company`, `ltd`→`limited`. **The four buckets are never merged.** The registry carries distinct registrants differing only in which suffix they use (`TORO CO` 0000737758 vs `TORO CORP.` 0001941131, and the same shape on Blue Owl and Fluent), so a single "has a corporate suffix" marker would compare two companies equal. Only the terminal token is rewritten — the same words sit mid-name in titles ending in a different suffix (`American Water Works Company, Inc.`) — and a terminal token that merely contains one (`Ltd./ADR`) is not a suffix. Mid-name punctuation is out of scope and still blocks a match (`Fluent Inc` does not reach `Fluent, Inc.`); so are `llc`/`plc` and the `lp`/limited-partnership family, which must never fold into the `ltd` bucket (#107). The **catch-all ticker lookup** retries a dotted symbol with `.`→`-` on a miss — brokers and market data write `BRK.B` where `company_tickers.json` lists `BRK-B`, and that file carries exactly one dotted ticker (a placeholder), so the substitution cannot land on the wrong registrant. The early gate stays untouched: it is a fast path, not the only ticker lookup, and `search_filings`' `ticker:` token inherits the fix through `resolveCik` with no change of its own (#110)
- **Former-name resolution:** `src/services/edgar/data/former-names.json` is a committed asset (generated offline via `bun run gen:former-names`) containing `[lowercasedFormerName, zeroPaddedCIK]` tuples. `buildTickerCache` folds them into `allEntries` for name search and the trigram pass. Regenerate on demand after major M&A waves — former names are immutable, so the file rarely needs updating.
- **Near-match suggestions:** `suggestCompanies(query, allEntries)` runs a Dice-coefficient trigram pass on the zero-hit search path, returning up to 3 candidates above a threshold. Each entry is scored on its **name and its ticker**, keeping the higher of the two — one ranked, deduped, top-N list rather than two disjoint ones, so a strong ticker match can outrank a weak name match and vice versa, and a ticker-shaped miss gets help instead of a bare `no_match` (#111). An entry missing either field is simply not scored on it: former-name entries carry a name and no ticker, and MF rows carry neither a name nor a place in `allEntries`, so fund tickers stay structurally unreachable here regardless of what the scorer does. Suggestions appear in the `no_match` error `data.suggestions` and in the error message. Never auto-resolves.
- **XBRL friendly names:** `concept-map.ts` maps `"revenue"` → real XBRL tags; handles historical tag changes (ASC 606), including both assessed-tax variants — a filer presenting revenue gross of excise/sales tax reports only `RevenueFromContractWithCustomerIncludingAssessedTax`, and without it the walk fell through to tags SEC retired in 2018 (#98). The Including variant sits BEHIND `Revenues`, not ahead of it — the two disagree for filers that tag gross sales and net-of-excise sales under separate elements, so promoting it over `Revenues` swaps the definition partway back through the history and prints a step change that is a tag switch, not a business fact; behind `Revenues` it only fills frames no current tag covers. The same precedent places successor tags: `capex` walks `PaymentsToAcquireProductiveAssets` behind the PP&E-only tag (it also counts software and intangibles) and `interest_expense` walks `InterestExpenseNonoperating` last, so NVIDIA, Amazon, Visa, and Microsoft reach current values and no frame the older tags cover changes (#125). A different-definition tag still goes in `relatedTags`, never `tags` — `ppe_net`'s finance-lease-inclusive total is the primary line for Alphabet, Meta, and Tesla (#36, #130). `ifrsTags` covers every concept with an IFRS element confirmed present in a live 20-F filer's companyfacts; `notes_payable` is deliberately unmapped (IFRS presents borrowings as one caption, with no notes/debt split) and `shares_outstanding` is `dei`. Never add an `ifrsTags` entry from the taxonomy alone — an element existing is not evidence filers tag it (#99). `ifrsTags` order is semantic the same way `tags` is: index 0 is the counterpart of the us-gaap index-0 tag, so the same concept answers the same question under either taxonomy (`accounts_receivable` leads with the trade-only caption, not the combined trade-and-other one; `share_repurchases` with the financing outflow, not the treasury-only line a share-cancelling filer tags zero). Verify a candidate order against filers that report BOTH tags — presence alone does not settle which one leads. One concept has no right order: IFRS filers split between the employee-scheme share-based-payment element and the IFRS 2.51(a) expense total, and the two invert between filers (SAP's total is its real line with the employee element a two-period fringe; Sanofi is the exact opposite), so `stock_based_compensation` carries `ifrsTagSelection: 'coverage'` and picks per filer by which tag covers more standard periods (#101). It is the only concept that does, and the flag is scoped to `ifrsTags` on purpose — its us-gaap pair IS a ladder, since `AllocatedShareBasedCompensationExpense` is the disaggregated line and out-covers the total for Alphabet, IBM, Tesla, and P&G.
- **Unknown concept names:** `findUnknownConcept` in `concept-map.ts` runs in `get_financials`, `compare_companies`, and `fetch_frames` before company resolution or any SEC request. After trimming, a catalog name in any case or separator form passes, and so does anything matching `/^[A-Z][A-Za-z0-9]*$/` — every item element in us-gaap, ifrs-full, dei, and srt has that shape and SEC matches tags case-sensitively, so no input that could return data is rejected, and a well-formed tag that 404s stays `no_data` / `no_concept_data` (#45). Everything else fails as `unknown_concept` (`NotFound`) and never reaches a request URL, where the tag is interpolated as a path segment. The hint is a formula from the `DERIVATIONS` table (`free_cash_flow`, `ebitda`, `working_capital`; read before the tag-shape test, since no taxonomy element is named `EBITDA`) or up to three catalog names by Dice trigram score ≥ 0.3 against name and label, ties by name — the label is the only route from `capital_expenditures` to `capex`. `total_debt` deliberately has no formula: `debt` is long-term only. `compare_companies` fails only when every concept is unknown; otherwise it answers the rest and lists each unknown once in `unknown_concepts`, never as a per-company gap (#128). Known concepts are deduped the same way, by what they resolve to — the catalog entry, or the trimmed raw tag — so `revenue`, `Revenue`, and ` revenue` are compared once under the first spelling and a caveat names the merged inputs; a catalog name and a raw tag it walks (`revenue` and `Revenues`) are different concepts and stay separate (#145). `trigramSimilarity` lives in the `trigram-similarity.ts` leaf so the catalog scores names without importing the HTTP service.
- **Series staleness:** two causes, one symptom, one caveat — `seriesStalenessCaveats(tag, label, newest, reference)` in `concept-series.ts` owns both, so a series is never flagged twice in different words. A friendly name's tag array is walked in order, so a tag SEC retired can win when every current tag reports nothing for the filer; SEC ships that tell in the taxonomy `label` (`"… (Deprecated 2018-01-31)"`), already in hand — no extra upstream call. Generalized rather than pinned to one pair: `revenue` and `cogs` both carry retired fallbacks (#98). A *current* tag whose series simply ends carries no tell at all, so the fallback signal is the gap itself: **730 days**, two full fiscal years behind the reference. The floor has to clear one year plus a filing window — an annual-only concept sits a fiscal year back until the next report lands, and a 20-F is due four months after year end — which puts a healthy filer around 500 days. Each tool passes the freshest reference it has without an extra call: `get_snapshot` and `compare_companies` pass `newestReportedPeriod(facts, taxonomy)`, the newest standard period the filer reports across the whole concept catalog — it isolates a line lagging its siblings and collapses to silence for a filer that stopped filing, whose every line is equally old, and it deliberately reads the whole catalog rather than the requested concepts, or a single-concept comparison would measure a series against itself. The reference skips the `dei` namespace — cover-page facts track the filing, not the statements, so they stay current for a registrant that stopped reporting under this taxonomy; counting them put Toyota's us-gaap profile (last reported fiscal 2020) six years behind a cover-page date and flagged all 28 of its lines at once. `get_financials` passes today, because one companyconcept payload holds no filer-wide period. All three measure the *full deduped set*, not the period-filtered slice — an annual view trailing a live quarterly series is a property of the filter. The retirement stamp wins when both apply; it names the concrete cause. `get_financials` therefore reads the clock, so its tests pin one (`vi.useFakeTimers({ toFake: ['Date'] })`) (#102).
- **XBRL deduplication:** Filter to entries with `frame` field to get one value per standard calendar period. SEC frames the latest-filed fact for a period, and since the pay-versus-performance rule a DEF 14A re-tags five years of annual `NetIncomeLoss`, so the proxy's figure — often rounded, sometimes mis-scaled or sign-flipped (TheRealReal's CY2024 at +134.2B against the 10-K's −134.2M) — holds the frame. A frame held by any Schedule 14A/14C form (`isProxyForm`) is answered with the latest-filed fact from another form with the same tag, unit key, `start`, and `end` (`reportingFormTwin`), keeping the frame; with no such fact the proxy value stands. Only proxy holders are replaced — a periodic-form allow-list would revert Bank of America's CY2013, framed on a 2016 8-K recast (10,539M) that no later 10-K repeats. The swap runs per tag and per unit key BEFORE cross-tag priority, so tag 0's proxy-held frame takes its twin from tag 0 and is never handed to a lower tag, and it costs two regex tests per framed value when no holder needs correcting. SEC also frames any year-long duration as `CY####`, including the trailing-twelve-month figure a 10-Q discloses, so a later 10-Q can hold an annual frame for a year the filer has not closed (Amazon's CY2026 `NetIncomeLoss` is 135.3B for July 2025–June 2026). An annual frame held by a 10-Q/10-QT takes the latest same-period fact from a form that is neither a quarterly report nor a proxy (the 10-K, in practice), else stays only when it covers a closed fiscal year a later 10-Q happens to disclose — ended before filing, and ending within a week of a year-long fact's end elsewhere in the series (or the series shows none to test against) — else drops out of the series (#142). Leaving every twinless 10-Q holder out would delete real years: across seven filers' companyfacts, 9 of 74 such holders are closed fiscal years reported only in a later 10-Q (Merck's `AccountsReceivableSale` CY2020). `frameHolderFact` owns both holder rules. The mirror's `getFrames` applies them and marks its response `holderFormsResolved`; the live frames API carries no form, so `fetch_frames` caveats annual `NetIncomeLoss` frames (#123) and annual frames for a calendar year still inside its 10-K window, through April of the next year (#142). When a friendly name maps to multiple tags, same-frame collisions resolve by **tag priority** — the lower-index tag wins (index 0 = preferred total, e.g. IFRS `Revenue` over the `RevenueFromContractsWithCustomers` sub-line), ties within one tag by latest `filed` (restatement). Tag-array order in `concept-map.ts` is therefore semantic (#44). The rules live once in `concept-series.ts` (`resolveFrameSeries` over `CompanyConceptUnit`s carrying a `tagIndex`, their source `tag`, and their unit key; `seriesFromCompanyFacts` for the whole-company reads) — every tool reading XBRL facts goes through it so their numbers agree. `resolveConceptTarget` in `concept-map.ts` is the matching taxonomy+tag resolution (mapping taxonomy wins over the `us-gaap` default; `ifrs-full` uses `ifrsTags` when present) and carries the `tagSelection` for the array it picked. `fetch_frames` resolves through it too, so its `taxonomy` input — `us-gaap` or `dei`, the only namespaces SEC publishes frames for — picks the namespace a raw tag is read from, a catalog name keeps its own mapped taxonomy under the default, and an explicit `dei` reads a catalog name's tags from `dei` exactly as `get_financials` does (#143). **`priority` is the default and coverage ranking is never the general rule** — it answers "which element does this filer maintain", which is the question only when the elements mean the same thing. Where they mean different things the array order is the answer and period count is noise: Molson Coors keeps eleven calendar years under a 2018-retired revenue element and ten under the assessed-tax-inclusive one against five under `Revenues`, so ranking `revenue` by coverage hands it either the dead tag or its gross-of-excise line in place of the continuous net series; Smucker and Spotify survive it only on a coincidental tie in period count (#101). `preferredTagIndex` picks the winning tag under either rule. Every resolved value names its own source tag — `get_financials` rows, `get_snapshot` points, and `compare_companies` cells each carry one — and the line-level `tag`/`label`/`unit` follow the tag behind the newest value (`describingTag`): under `coverage` that is the winner, and under `priority` a successor appended behind an older leader answers the recent frames, so describing the line by the leader would name a tag that produced none of its current values (#125). A companyconcept or companyfacts unit served as anything but an array is dropped at the service edge (SEC serves Visa's and Coca-Cola's `NetIncomeLoss` companyconcept as `"units":{"USD":{}}`); a companyconcept payload left with no units reads as served-empty, and `get_financials` answers every such tag from one companyfacts read, which it reuses for the no-data probe (#141). The two rules also differ in what the losers contribute: under `priority` the array is a ladder, so a lower tag fills frames the leader does not report, while under `coverage` the tags are alternates and only the winner contributes — gap-filling from an alternate splices two definitions into one series, and the filers reporting both tags disagree where they overlap (Ferrari tags CY2022 at EUR 16.2M under the employee element and EUR 20.9M under the IFRS 2.51(a) total), so the filled series prints a step that is a tag switch rather than a business fact. Seven percent of the 224 IFRS filers reporting either `stock_based_compensation` tag have that shape; a winner-only series that then ends years back is what the staleness caveat is for.
- **Fiscal-period caveats:** SEC reports a filer's fiscal Q4 as the 10-K residual, never as a discrete quarterly fact, so the calendar quarter that a filer's fiscal Q4 *spans* carries no frame-tagged quarterly value. That is not always the quarter holding the fiscal year-end date — a January year-end closes a fiscal Q4 running Nov–Jan, which SEC frames as calendar Q4 — and calendar-year filers are affected too (no discrete Q4). `fiscal-periods.ts` owns both directions: `fiscalQ4Caveats(period)` for the cross-company frame view (`fetch_frames`), `missingQuarterCaveats(frames)` for the per-filer series view (`get_financials`, `get_snapshot`, `compare_companies`) (#95). The per-filer detection is data-driven — no maintained filer list — and reads only the newest 4 years carrying quarterly frames, applying the per-year quorum inside that window rather than before it, because SEC frame-tagged all four calendar quarters before ~CY2021 and drops the fiscal-Q4 quarter after; a whole-history read would let the older years mask the live gap, and filtering first would let a filer whose recent years are too sparse silently describe the older regime. **Keep that slice-then-filter order.** The detector reports a *set* of absent quarters, not one: a filer whose remaining fiscal quarters also span non-calendar durations loses two (Costco tags only calendar Q1 and Q4), and the per-year quorum is two of four — low enough to admit that shape, high enough to bound the absent set at two and keep a one-quarter stub year from dragging a third in (#100). Message wording and the derivation advice differ between the one- and two-quarter cases; the three consuming tools all iterate `caveats`, so nothing downstream assumes at most one entry.
- **`get_financials` period default:** with `period_type` unset the series defaults to `annual` (clean FY series); if the annual filter empties a non-empty series whose frames are all instant (`CY####Q#I` — balance-sheet, shares-outstanding, raw instant tags), it falls back to the full series so the first call returns data (#48). An explicit `period_type` is honored as-is and still errors when it excludes everything.
- **8-K item codes:** SEC replaced the original single-integer item numbering with the current `x.xx` scheme in Release 33-8400, effective **2004-08-23**, so EDGAR's `items` field carries two disjoint vocabularies. `eight-k-items.ts` decodes by the code's **shape**, never the filing date — dotted is current, a bare integer is legacy — because filers reported pre-changeover events under the old numbers for months after the effective date, and a date-keyed table would mislabel exactly those. The two vocabularies reuse integers for unrelated events (legacy `5` is Other Events, current `5.02` is officer departures; legacy `12` is the ancestor of `2.02`), so a dotted `items` filter never matches a pre-2004 filing — `get_material_events`' zero-hit notice says so and echoes the codes actually present. The tables back three surfaces: the tool's `items` enum, the decoded per-filing rows, and the `secedgar://filing-types` resource. A code matching a shape but absent from that regime's table comes back with no label rather than a guess.
- **EFTS quirks:** `dateRange=custom` must be set when using date params; the singular `entity` param is ignored — `cik:`/`ticker:` targeting passes the resolved CIK via the plural `ciks` param (server-side scope, independent of the filing's name text, so former-name filings on the same CIK are matched); `size` is ignored entirely and every request answers with 100 hits, so `find_holders` pages with `from` alone and expresses its budget in pages (5 → 500 rows)
- **Submissions archive walk:** the submissions feed's `filings.recent` holds the last year or 1,000 filings, whichever is more — so 26,000+ rows for JPMorgan, and a filing day is never split (Microsoft's window runs to 1,002 rows to finish its oldest day); everything older lives on `filings.files[]` archive pages of 200–360 KB each. Every tool that reads past `recent` goes through `SubmissionsArchiveWalk` (`submissions-archive.ts`): one `ARCHIVE_PAGE_SCAN_CAP` (10) for all, pages read one at a time through the service's pacer and page cache, and no page selected when the lower bound falls on or after `recent`'s oldest date — so a call answered by `recent` sends no archive request. The caller picks the order — `newest-first` (`company_search`, `search_filings`' pre-2001 arm, `get_material_events`, `get_insider_transactions`' window), or `oldest-first` forward from a date (`get_institutional_holdings`' quarter lookup, since newest-first cannot reach JPMorgan's 2020-Q4 13F-HR on page 41 of 70 under the cap) — and can stop after any page; `scannedThrough` (the oldest page start read) and `truncated` (selected pages left unread, whether the cap, an early stop, or a caller that never started the walk left them) are reported the same way everywhere. Worst case per call is the submissions document plus 10 pages plus the tool's own document reads (`get_insider_transactions`: 111). `get_material_events` without a date window walks only to fill `limit`, stopping on the page that fills it with rows passing the `items` filter, and its `dataset.truncated` is true whenever selected pages went unread — including when `recent` alone fills `limit` and no page is read, since the dataframe then holds only the recent window (#134); `get_institutional_holdings` stops on the first page holding a 13F-HR or 13F-NT for the quarter, and runs the #86 operating-company routing only when neither `recent` nor the pages read hold any 13F-HR or 13F-NT — a notice-only manager asked for a quarter it filed nothing for gets the plain quarter miss (#117, #133); without a `quarter`, a manager whose `recent` holds 13F-NT notices and no 13F-HR gets the notice message for its newest notice rather than the #86 classification (#133). **Page selection widens both date bounds by a day:** a manifest page's `filingFrom` / `filingTo` can sit a day off its own rows in either direction (JPMorgan's page 041 is listed to 2021-02-17 but holds 42 filings dated 2021-02-18; Apple's page 001 is listed a day past its newest row), so an exact overlap test skipped the page holding a window's edge day. The slack costs at most one extra page per bound, and every caller filters rows by their own filing dates (#137)
- **Filing metadata outside `recent`:** `get_filing` takes `form`, `filing_date`, and `period_ending` from `recent` when the accession is there, else from the filing's own SEC header — the raw `<TYPE>` / `<FILING-DATE>` / `<PERIOD>` in `index-headers.html`, which the tool already fetches to type the documents, and only when that page 404s (common before 2014) the sub-1 KB `<accession>.hdr.sgml`. `FILING-DATE` is the filing date, not the acceptance timestamp, and matches the submissions feed; an S-8 or Form 4 header has no `PERIOD`. `company_name` stays the current registrant name (#126)
- **Filing content:** HTML → text via `html-to-text` library; pre-2005 filings produce noisier output. The library walks the DOM recursively, so `CONVERT_OPTIONS` caps it at `limits.maxDepth: 512` with a `[…]` ellipsis where the cut falls — without it deep markup overflowed the stack and one such candidate failed `search_filings`' whole pre-2001 text scan (#118). The bound must hold on both runtimes the server ships under: a fresh Node process overflows near 1,360 levels of nested lists (its most stack-hungry shape), Bun near 13,000; real HTML filings measure under 25 levels. Legacy SGML `<PAGE>` markers are replaced with a line break before parsing (`neutralizePageMarkers`) — the parser never closes them, so a text document otherwise nests one level per page and a long one would be cut at the cap; with them gone the deepest measured `.txt` sits under 100 (an EX-27 schedule's unclosed tags). Both changes leave every sampled filing's text byte-identical. `limits` deep-merges with the library defaults, so its 16,777,216-character `maxInputLength` (silent truncation, no ellipsis) still applies
- **Heading detection:** `detectHeadings` runs three line-anchored patterns — the all-caps forms, the mixed-case same-line `Item N. Title` / `Part II` forms (#71), and a bare `Item N` marker whose title lands on a later line. Every whitespace match is horizontal-only: `\s` matches `\n`, which fused runs of consecutive all-caps lines into composite entries (`"TABLE OF CONTENTS\n\nPART I"`) (#105). Horizontal-only is not space-or-tab — the all-caps arm alternates `[^\S\n]` in rather than listing literal members, because filings put NBSP inside all-caps headings too (`MICROSOFT<NBSP>CORPORATION`). The Item suffix spans `16A`–`16K` and the marker's period is optional — 20-F filers write the lettered markers unpunctuated. **A bare marker is admitted on the shape of what follows, not just on a title being there:** a blank line between marker and title admits the pair; without one, the pair is admitted only when no bare page-number line sits under the title, since marker/title/page-number on three tight lines is the TOC-cell artifact #71's guard exists to reject. That test is one-way — it rejects a TOC row, it does not certify a heading — so a marker line anywhere else (a page footer, an unwrapped paragraph) is admitted too, taking the next line as its title; the entry carries the marker's own offset, so `section` still lands in the right place. A heading text repeated `RUNNING_HEADER_MIN_OCCURRENCES` (5) times at distinct offsets is page furniture — C3is' FY2025 20-F repeats `TABLE OF CONTENTS` 129 times and a 10-Q stamps `PART I` on all 47 of its pages, while a heading printed once in the TOC and once in the body occurs twice. **Furniture is then handled by shape, two ways:** a bare marker (`Item 1`, `PART I`) carries no navigable text, so no occurrence of it earns an entry and every one is dropped; anything else keeps its FIRST occurrence and drops the rest, inverting the keep-later rule — a filing that runs a section's own heading across every page of that section (`NOTES TO THE CONSOLIDATED FINANCIAL STATEMENTS` through an S-1's financial statements) prints it first at the section start, and keep-later would land it on the section's last page instead. So the outline carries no fused composite and at most one entry per running header, at its first occurrence. **A bare marker is counted on its own line, never on the joined heading:** a page footer's marker takes its title from whatever prose the next page opens with, so every composite it produces is a different string and only the marker line repeats — count composites instead and a 10-Q's per-page `Item 1` footer yields one bogus entry per page, crowding the real Items past `maxEntries`. That count is per rendered form (`Item 1` and `Item 1.` are separate), since a filing stamping the bare form on every page still writes the punctuated one for the real heading. Dedup (keep the later, i.e. body, occurrence) keys on `foldForHeadingMatch` plus a terminal-punctuation strip, never raw lowercase: the same filing prints one marker with different space characters and a trailing period in its TOC and its body (`Item 3.` with U+0020 vs U+2009; `…Policies` vs `…Policies.`), and a raw key leaves both in the outline. Trailing bare page numbers are stripped from Item headings for the same reason. Lettered sub-item markers (a bare `A.`/`D.` inside an Item) are deliberately not detected — single letters recur throughout filing prose.
- **Section matching:** `get_filing`'s `section` compares through `foldForHeadingMatch` (`filing-to-text.ts`) on **both** operands — Unicode whitespace runs collapse to one space, U+2018/U+2019 → `'`, U+201C/U+201D → `"`, then lowercase. Comparison-time only: `FilingHeading.heading`, the `outline` output, `content[]`, and every character offset keep the document's own bytes. EDGAR headings carry NBSP runs and typographic quotes, so a caller re-sending an outline heading through a UI that normalizes either one never byte-matched the heading it came from — which broke the `section_not_found` recovery loop the error's own hint points at (#106). Still a plain substring test, not fuzzy matching — `"item 8"` resolves only to Item 8 — but run on folded operands, so a single-spaced needle now also reaches a heading whose marker carries an NBSP run or a double space (`Item 9.` in a 2004 10-K). Folding only one operand does not fix the round trip.
- **Binary filing documents:** a filing index routinely carries scanned pages, PDF exhibits, `Financial_Report.xlsx`, and `-xbrl.zip` — `filingToExtract` runs `html-to-text` over whatever bytes arrive and never inspects a content type, so requesting one used to return decoded binary as `content` with no error (#96). `get-filing.tool.ts` detects them by filename extension (`binaryTypeFromName`), because that is the only signal available before the body is fetched: the canonical `GRAPHIC` type lives in the submission header, which the archive path fetches *after* the document, and requiring it would trade today's parallel header/submissions fetch for a sequential one. Two halves, both needed — `documents[]` entries carry `binary: true` (`isBinaryDocument` also honors a header `GRAPHIC` type), and `resolveFilingArchive` throws `binary_document` before `tryGetFilingDocument` runs, since the `document` input is an unconstrained string and a catalog marker alone cannot stop a caller. Extension-driven and bucket-independent on purpose: a PDF exhibit is typed `EX-99.*` and files under `exhibits`, not among the graphics.
- **Dataframes:** `secedgar_search_filings`, `secedgar_get_financials`, `secedgar_fetch_frames`, `secedgar_compare_companies`, `secedgar_get_insider_transactions`, `secedgar_get_institutional_holdings`, `secedgar_get_material_events`, `secedgar_find_holders`, `secedgar_get_beneficial_owners`, and `secedgar_get_fund_holdings` materialize their full result set as `df_<id>` on a shared DuckDB-backed canvas (one per tenant). The two ownership tools register the full parsed set (inline response is a preview capped at `limit`); `get_insider_transactions` scans at least `INSIDER_CANVAS_FILING_SCAN` recent filings when canvas is available (so small inline limits don't under-scan the window), capped at 100, and flags `dataset.truncated` when more Form 4 filings exist beyond the scanned window (#63); with a `filed_after` / `filed_before` window it reads the in-window Form 4 / 4-A rows with an XML primary document — `recent`, then the archive pages overlapping the window, newest first, stopping once the scan cap plus a sentinel row are in hand — and with a canvas parses every one up to 100 instead of the 40-filing floor, since the window already bounds the set; `dataset.truncated` is then the sentinel or unread overlapping pages, and `history_scanned_through` the oldest filing parsed (#127); its `shares_traded` is an unsigned magnitude paired with a `direction` (acquire/dispose) column, so net activity is `SUM(CASE WHEN direction='dispose' THEN -shares_traded ELSE shares_traded END)` (#46); `get_financials`' dataframe materializes the source-filing fiscal keys as `source_filing_fy`/`source_filing_fp` (renamed from bare `fiscal_year`/`fiscal_period` — every comparative period restated in one filing carries that filing's fy/fp, so `period_end` is the time key for ORDER BY/GROUP BY/windowing) (#72); `compare_companies`' dataframe holds the FULL aligned series across every period while the inline matrix is capped at `periods`, and each row carries both the aligned `period` key and the underlying XBRL `frame` (they differ for point-in-time concepts) (#85); the inline window is the newest periods of the union across companies, so a company-concept pair whose values all predate it is named in `caveats` — one line per concept, each company with its newest period and the dataframe pointer — and never in `gaps`, which keeps meaning no value in any period (#144). Each row set carries an all-nullable schema (framework's `inferSchemaFromRows` always emits `nullable: true` since mcp-ts-core 0.10.4). Per-table TTL is delegated to the framework via `RegisterTableOptions.ttlMs` (resolved in mcp-ts-core#140); bridge still tracks metadata expiry in `ctx.state` for provenance. Every producer emits the describe-then-query pointer on the branch that registers the dataframe, composed once by `dataframeGuidance(dataset)` in the bridge so the wording cannot drift; where a truncation call already fires there it rides as that call's `guidance`, since `ctx.enrich.notice` is last-wins across `notice`/`truncated` and a second call would clobber the first (#104). `secedgar_dataframe_query` detects a `row_limit`-bound result from `QueryResult.truncated` rather than `rowCount > rows.length` — `row_limit` is pushed into the query as the provider's cap, so a capped result comes back with `rowCount === rows.length` and the arithmetic cannot see it; `rowCount` is then the cap and not a total, so that branch names the cap instead of printing "of N rows", `row_count_capped` carries the same disclosure into `format()`, and the fallback comparison stays for the `registerAs` path, where `truncated` is never set. A SQL `LIMIT` exactly equal to the cap is indistinguishable from a result that holds that many rows and is reported as exact (#109). `secedgar_dataframe_query` runs framework's SQL gate with `denySystemCatalogs: true` to block `information_schema`, `pg_catalog`, `sqlite_master`, and `duckdb_*` catalog access. Raw DuckDB execution errors (missing table, syntax) are caught in `bridge.query` and re-thrown as structured `missing_table`/`invalid_sql` reasons, with DuckDB's "Did you mean …?" hint stripped so internal catalog names don't leak (#47). The framework's own classifications pass through untouched and are declared on the tool's contract (`thrownBy: 'service'`, since the handler never names them): a SELECT that prepares and then fails on the data — a `Conversion` / `Invalid Input` / `Out of Range` error, e.g. `CAST('abc' AS INTEGER)` — is `ValidationError` with `sql_execution_error` (mcp-ts-core 0.13.4; `DatabaseError` with no reason before), and the gate's `multi_statement`, `denied_function`, and `plan_operator_not_allowed` rejections arrive with their own recovery hints. `get_beneficial_owners` registers one row per reporting person rather than one per filing, since powers and percent of class are per person; `get_fund_holdings` registers every position in the report while the inline list is one `limit`-sized page from `offset`, and its rows carry the fund keys (`series_id`, `registrant_cik`, `report_period_date`, `accession_number`) so a portfolio joins the 13F and insider rows on `cusip`.
- **Beneficial ownership (13D/13G):** SEC replaced the legacy `SC 13D` / `SC 13G` text filings with structured XML under the `SCHEDULE 13D` / `SCHEDULE 13G` names on **2024-12-18**; the changeover is total, so a legacy-name filter matches nothing recent and a current-name filter matches nothing older. The two schedules are **separate schemas, not one shape with optional fields** — 13D repeats `reportingPersons/reportingPersonInfo` with the power fields flat on each person and the percentage in `percentOfClass`; 13G repeats `coverPageHeaderReportingPersonDetails` with the powers nested inside `reportingPersonBeneficiallyOwnedNumberOfShares`, the percentage in `classPercent`, no reporting-person CIK anywhere, and no purpose-of-transaction field at all (its item 4 is structured ownership data, which is what makes it the passive form). `beneficial-ownership-parser.ts` therefore branches per schema and dispatches on the submission type, falling back to whichever repeated reporting-person element the document carries. **Ownership powers and percent of class are per reporting person, including on joint filings** — a fund, its adviser, and its controlling principal each report the same underlying shares, so the parsed owner list is never collapsed to a group total and the percentages must not be summed. Discovery runs off the issuer's own `getSubmissions()`: SEC cross-lists a schedule under every CIK party to it, so no entity-scoped EFTS step is needed.
- **NPORT-P series routing:** a fund's public portfolio report covers **exactly one series**, a registrant trust holds many (iShares Trust: 390 listed series, 200+ reports for a single period), and the series identity appears **only inside the document** — neither the submissions feed nor EFTS carries `seriesId` on an NPORT-P row. Routing is therefore series-first: EDGAR's company browse accepts a series ID in place of a CIK (`getFundSeriesFilings`), which is the one SEC surface mapping a series to its accessions and which also names the registrant, so a bare series ID resolves in one request. The browse feed's registrant outranks any registrant the `fund` input implied, since `series_id` can name a series of an entirely different trust and the archive path is keyed on the registrant. The bounded header scan (`tryGetFilingDocumentHead` + `parseNportHeader`, cap `SERIES_SCAN_LIMIT`) is the fallback and the disambiguator for a trust whose series carry no listed ticker; SEC's archive host **ignores `Range`** (a ranged request answers 200 with the whole body), so the saving is client-side stream cancellation only — and the cutoff cannot cut below one read chunk, which Bun sets at 262,144 bytes, so `HEADER_MAX_BYTES` binds only on a document arriving in several chunks. **A registrant's own NPORT-P list is never the candidate set for a series-structured trust**: series fiscal quarters are staggered, so the list interleaves funds and one report at the newest period does not mean one fund — it means the fund whose quarter ends latest. The scan therefore runs on the newest batch even when that batch holds one report, and a series found there re-routes through the browse feed; only a report naming no series at all (a closed-end fund, which also carries no `classId`) is answered from the registrant's list. Candidates are ordered by **reported period, not filing date**: an `NPORT-P/A` restating an old period is filed after every report of the periods that followed it. Reports publish on a ~60-day lag, so `report_period_date` and `publication_lag_days` are first-class output and the `as_of` enrichment states both.
- **Local mirror (opt-in, `EDGAR_MIRROR_ENABLED`):** routes `resolveCik`, `tryGetCompanyConcept`, `tryGetCompanyFacts`, and `tryGetFrames` to a local SQLite mirror (framework `MirrorService`) of `company_tickers.json` + the `companyfacts.zip` bulk archive; the live API is the fallback on a miss (`EDGAR_MIRROR_FALLBACK_LIVE`). Bootstrap out-of-band with `bun run mirror:init`; refresh nightly via cron (HTTP) or `bun run mirror:refresh`. The three `mirror:*` commands also ship in the production Docker image — run them against a deployed container with `docker exec <container> bun run mirror:<init|refresh|verify>`. Node/Bun only — skipped on Workers. Frames are assembled from the company-facts store, so `loc` (business location) is absent. The ingester stores a tag's own name when SEC serves it with no label (`InterestExpenseNonoperating`), so every read path treats a stored label equal to a camel-case compound tag as no label (`storedLabel`) — callers fall back to the concept's label, and `getFrames` answers an empty `label` as the live frames API does; a one-word element's label is the word itself (`Revenues`) and is kept. Read-side, so stored rows need no migration. No FTS5 — every routed lookup is exact/indexed (cik+taxonomy+tag point, taxonomy+tag scan, cik scan for the whole-company read, ticker/CIK). Mirror ingests MF fund symbols from `company_tickers_mf.json`; on the mirror path a live MF merge supplements them when `FALLBACK_LIVE` is on (superseding the mirror's bare fund rows with series/class, and covering a mirror synced before MF ingestion), so funds resolve there too (#43, #135).
---
## Config
| Env Var | Required | Default | Description |
|:--------|:---------|:--------|:------------|
| `EDGAR_USER_AGENT` | **Yes** | — | User-Agent for SEC compliance. Format: `"AppName contact@email.com"` |
| `EDGAR_RATE_LIMIT_RPS` | No | `10` | Max requests/second to SEC APIs. Do not exceed 10. |
| `EDGAR_RATE_LIMIT_COOLDOWN_SECONDS` | No | `600` | Seconds to stop sending to SEC after a 429. Calls are refused locally for this long, then the first one goes out alone as a probe. SEC lifts its block only after ten quiet minutes, so a shorter value just probes into it. |
| `EDGAR_TICKER_CACHE_TTL` | No | `3600` | Seconds to cache the ticker index (company_tickers.json + company_tickers_mf.json). A failed fund-file load is retried after 60 s (or the rate-limit cool-down) instead of standing for the TTL. |
| `EDGAR_DATASET_TTL_SECONDS` | No | `86400` | Per-table TTL for canvas-registered dataframes. Sliding window touched on every dataframe op. |
| `EDGAR_DATAFRAME_DROP_ENABLED` | No | `false` | Set to `true` to expose `secedgar_dataframe_drop`. TTL handles cleanup otherwise. Off, the tool is registered through `disabledTool()`: absent from `tools/list` and uncallable, but rendered on the HTTP landing page in a `disabled` group carrying the reason and `EDGAR_DATAFRAME_DROP_ENABLED=true`, so an operator can see the capability exists. The list `createApp()` receives — and `buildServerManifest()`'s `definitionCounts.tools` — is 17 either way; `/.well-known/mcp.json` on mcp-ts-core 0.13.6 carries no per-tool definitions at all, so nothing changes there (#103). |
| `EDGAR_MIRROR_ENABLED` | No | `false` | Enable the local SQLite mirror of company_tickers + XBRL company-facts. Node/Bun only (skipped on Workers). Bootstrap once with `bun run mirror:init`. |
| `EDGAR_MIRROR_PATH` | No | `./data/edgar-mirror` | Directory holding the mirror SQLite databases (tickers + companyfacts). |
| `EDGAR_MIRROR_REFRESH_CRON` | No | — | In-process nightly refresh cron (HTTP transport only). Recommended `0 9 * * *`. Omit to refresh out-of-band via `bun run mirror:refresh`. |
| `EDGAR_MIRROR_FALLBACK_LIVE` | No | `true` | Fall back to the live SEC API on a mirror miss (unsynced, or a filing newer than the last refresh). Set `false` for strict mirror-only reads. |
| `CANVAS_PROVIDER_TYPE` | No | `duckdb` | Canvas engine. Set to `none` to disable the canvas (e.g. on Cloudflare Workers). |
---
## Services
| Module | Path | Purpose |
|:-------|:-----|:--------|
| `EdgarApiService` | `src/services/edgar/edgar-api-service.ts` | Paced HTTP client with the rate-limit block gate, CIK resolution, all SEC API calls |
| `submissions-archive` | `src/services/edgar/submissions-archive.ts` | `SubmissionsArchiveWalk` — the one walk over `filings.files[]` archive pages: shared page cap, caller-chosen order, early stop, scan depth and truncation |
| `filing-headers` | `src/services/edgar/filing-headers.ts` | SEC submission-header parsers: per-document types and the submission's form / filing date / period |
| `concept-map` | `src/services/edgar/concept-map.ts` | Static friendly name → XBRL tag mapping; `resolveConceptTarget` picks the taxonomy + tag list; `findUnknownConcept` rejects names that are neither |
| `trigram-similarity` | `src/services/edgar/trigram-similarity.ts` | Dice trigram scorer behind company and concept near-match suggestions |
| `concept-series` | `src/services/edgar/concept-series.ts` | Shared frame dedup + tag selection (priority or per-filer coverage); one value per standard calendar period; series-staleness caveats |
| `fiscal-periods` | `src/services/edgar/fiscal-periods.ts` | Fiscal-Q4 frame caveat and per-filer missing-quarter detection |
| `eight-k-items` | `src/services/edgar/eight-k-items.ts` | 8-K item-code decode tables for both numbering regimes; shape-driven `decodeEightKItem` |
| `xml-nodes` | `src/services/edgar/xml-nodes.ts` | Prefix-insensitive tag lookup + value coercion shared by the SEC XML parsers |
| `ownership-parser` | `src/services/edgar/ownership-parser.ts` | Form 4 ownership XML, 13F information tables, and the 13F cover page (filer name, period) |
| `beneficial-ownership-parser` | `src/services/edgar/beneficial-ownership-parser.ts` | Sibling SCHEDULE 13D / 13G parsers with a submission-type dispatcher |
| `nport-parser` | `src/services/edgar/nport-parser.ts` | NPORT-P identity block (cheap, for series routing) and full portfolio report |
| `filing-to-text` | `src/services/edgar/filing-to-text.ts` | HTML → readable plain text conversion; heading detection and the comparison fold `section` matches through |
| `CanvasBridge` | `src/services/canvas-bridge/canvas-bridge.ts` | Adapter over framework `DataCanvas`: `df_<id>` minting, all-nullable schema derivation, per-table TTL, shared-canvas acquire |
| `sql-gate-extras` | `src/services/canvas-bridge/sql-gate-extras.ts` | System-catalog SQL deny on top of the framework's read-only gate |
| `EdgarMirror` | `src/services/edgar/mirror/` | Opt-in local SQLite mirror (framework `MirrorService`) of company_tickers + XBRL company-facts; ready-gated read helpers back `resolveCik`/`tryGetCompanyConcept`/`tryGetCompanyFacts`/`tryGetFrames` |
---
## Context
Handlers receive a unified `ctx` object. Key properties:
| Property | Description |
|:---------|:------------|
| `ctx.log` | Request-scoped logger — `.debug()`, `.info()`, `.notice()`, `.warning()`, `.error()`. Auto-correlates requestId, traceId, tenantId. Dual-sink: Pino **and** `notifications/message` to the client, so treat it as client-visible. |
| `ctx.state` | Tenant-scoped KV — `.get(key)`, `.set(key, value, { ttl? })`, `.delete(key)`, `.getMany(keys)`, `.list(prefix, { cursor, limit })`. Accepts any serializable value. |
| `ctx.requestInput` | Suspend and ask the caller for more input — `return ctx.requestInput({ inputRequests: { key: inputRequired.elicit({ message, requestedSchema }) } })`. Never returns; the handler is re-entered with the answers. Always present. |
| `ctx.inputs` | Reader over a retried request's responses — `.accepted(key, schema)`, `.view(key)`, `.state()`, `.dropped`. Empty on the first round. |
| `ctx.enrich` | Success-path agent context (empty-result notices, query echo, pagination totals) — `ctx.enrich(...)` or `.notice()` / `.total()` / `.echo()` / `.truncated()`. Reaches `structuredContent` and `content[]`; lands only when the definition declares an `enrichment` block (no-op otherwise). |
| `ctx.content` | Non-text content blocks — `.image(data, mimeType)`, `.audio(data, mimeType)`, or `ctx.content(block)` for a raw block. Prepended to `content[]` after `format()`; never enters `structuredContent`. |
| `ctx.signal` | `AbortSignal` for cancellation. |
| `ctx.requestId` | Unique request ID. |
| `ctx.tenantId` | Tenant ID from JWT; `'default'` for stdio or HTTP with auth off. |
---
## Errors
Handlers throw — the framework catches, classifies, and formats.
**Recommended: typed error contract.** Declare `errors: [{ reason, code, when, recovery, retryable?, severity?, thrownBy? }]` on `tool()`. The handler then receives `ctx.fail(reason, msg?, data?)` typed against the reason union, and `data.reason` is auto-populated for observability. The `recovery` field is required (≥5 words, lint-validated). Use `ctx.recoveryFor('reason')` to spread the contract recovery onto the wire (mirrored into `content[]` unless the message already contains it verbatim); pass an explicit `{ recovery: { hint } }` when runtime context matters. Forwarding it is lint-enforced per throw site (`error-contract-recovery-unforwarded`). Mark an entry raised below the handler — a module-level helper, a service, the canvas bridge — with `thrownBy: 'service'` so `error-contract-unthrown` skips it; lint-only metadata, nothing at runtime reads it. Baseline codes (`InternalError`, `ServiceUnavailable`, `Timeout`, `ValidationError`, `SerializationError`, `RequestCancelled`) bubble freely without declaration. **All EDGAR tools that have known failure modes use this pattern.**
```ts
errors: [
{ reason: 'company_not_found', code: JsonRpcErrorCode.NotFound,
when: 'The company input does not resolve to a CIK',
recovery: 'Use a ticker symbol or 10-digit CIK number for an exact match.' },
],
async handler(input, ctx) {
if (!match) throw ctx.fail('company_not_found', `Company '${input.company}' not found.`, {
...ctx.recoveryFor('company_not_found'),
});
}
```
**Service-layer pattern (no `ctx`).** Throw an `McpError` with `data: { reason, recovery: { hint } }`, and mark the tool's matching `errors[]` entry `thrownBy: 'service'`. The auto-classifier preserves `data` on the wire so clients see the same `error.data.reason` they'd see from `ctx.fail`. The tool error text closes with a term line rendering only what `data` carries — `(reason <reason>)` for a reason alone, `· retryable` / `· not retryable` appended when `data.retryable` is a boolean — so tests assert that `content[0].text` contains the diagnostic rather than pinning it exactly. A contract entry's `retryable` reaches `data` only through `ctx.fail`; a service throw sets `data.retryable` itself (#122).
**Fallback for ad-hoc throws** (no contract entry fits, prototype code): use error factories or plain `Error`.
```ts
import { notFound, validationError } from '@cyanheads/mcp-ts-core/errors';
throw notFound('Item not found', { itemId }); // explicit code
throw new Error('Item not found'); // auto-classified → NotFound
```
**HTTP responses.** `httpErrorFromResponse(response, { service, data })` from `/utils` maps the full status table (401/403/408/422/429/5xx) and captures body + `Retry-After`. Used in `EdgarApiService`.
See framework CLAUDE.md and the `api-errors` skill for the full reference.
---
## Structure
```text
src/
index.ts # createApp() entry point — stateless sessionMode, teardown stops the SEC pacer and closes the mirror
config/
server-config.ts # EDGAR env vars (Zod schema)
services/
edgar/
edgar-api-service.ts # Paced HTTP client + rate-limit block gate, CIK resolution
submissions-archive.ts # Shared archive-page walk (page cap, order, early stop)
filing-headers.ts # SEC submission-header parsers
concept-map.ts # Friendly name → XBRL tag mapping + taxonomy/tag resolution
concept-series.ts # Shared frame dedup, tag selection, staleness caveats
trigram-similarity.ts # Dice trigram scorer for near-match suggestions
fiscal-periods.ts # Fiscal Q4 / missing-quarter caveats
eight-k-items.ts # 8-K item-code decode tables (current + legacy regimes)
xml-nodes.ts # Shared XML tag lookup + value coercion helpers
ownership-parser.ts # Form 4, 13F information table + cover page parsers
beneficial-ownership-parser.ts # SCHEDULE 13D / 13G parsers + dispatcher
nport-parser.ts # NPORT-P header (series routing) + full report
filing-to-text.ts # HTML → plain text
types.ts # Domain types
mirror/ # Opt-in local SQLite mirror (framework MirrorService)
edgar-mirror.ts # two stores + ready-gated read helpers
tickers-sync.ts # company_tickers.json ingester
companyfacts-sync.ts # companyfacts.zip streaming ingester (fflate)
index.ts # barrel + server-side singleton
types.ts # row shapes + constants
canvas-bridge/
canvas-bridge.ts # Framework DataCanvas adapter, df_<id> minting, per-table TTL
sql-gate-extras.ts # Bridge-layer system-catalog deny on top of framework SQL gate
mcp-server/
tools/definitions/
index.ts # buildToolDefinitions() — the createApp() tool list, drop tool gated via disabledTool()
company-search.tool.ts
search-filings.tool.ts
get-filing.tool.ts
get-financials.tool.ts
get-snapshot.tool.ts
get-material-events.tool.ts
get-insider-transactions.tool.ts
get-institutional-holdings.tool.ts
find-holders.tool.ts
get-beneficial-owners.tool.ts
get-fund-holdings.tool.ts
fetch-frames.tool.ts
compare-companies.tool.ts
search-concepts.tool.ts
dataframe-describe.tool.ts
dataframe-query.tool.ts
dataframe-drop.tool.ts # Opt-in via EDGAR_DATAFRAME_DROP_ENABLED
resources/definitions/
concepts.resource.ts
filing-types.resource.ts
prompts/definitions/
company-analysis.prompt.ts
```
---
## Naming
| What | Convention | Example |
|:-----|:-----------|:--------|
| Files | kebab-case with suffix | `search-docs.tool.ts` |
| Tool/resource/prompt names | snake_case | `search_docs` |
| Directories | kebab-case | `src/services/doc-search/` |
| Descriptions | Single string or template literal, no `+` concatenation | `'Search items by query and filter.'` |
---
## Skills
Skills are modular instructions in `framework-skills/` at the project root. Read them directly when a task matches — e.g., `framework-skills/add-tool/SKILL.md` when adding a tool. `bun run list-skills` prints the full registry. The directory is deliberately not `skills/`: Claude Code and Codex auto-load a plugin's root `skills/`, so a server that ships `.claude-plugin/` or `.codex-plugin/` would hand these development skills to every agent that installs it. Keep `skills/` free for skills meant for those agents.
**Agent skill directory:** Copy skills into the directory your agent discovers (Claude Code: `.claude/skills/`, others: equivalent). Skills then load as context without referencing `framework-skills/` paths. After framework updates, run the `maintenance` skill — Phase B re-syncs the agent directory.
Available skills:
| Skill | Purpose |
|:------|:--------|
| `setup` | Post-init project orientation |
| `design-mcp-server` | Design tool surface, resources, and services for a new server |
| `add-tool` | Scaffold a new tool definition |
| `add-app-tool` | Scaffold an MCP App tool + paired UI resource |
| `add-resource` | Scaffold a new resource definition |
| `add-prompt` | Scaffold a new prompt definition |
| `add-service` | Scaffold a new service integration |
| `add-test` | Scaffold test file for a tool, resource, or service |
| `field-test` | Exercise tools/resources/prompts with real inputs, verify behavior, report issues |
| `tool-defs-analysis` | Read-only audit of MCP definition language across the surface — voice, leaks, defaults, recovery hints, output descriptions |
| `security-pass` | Audit server for MCP-flavored security gaps: output injection, scope blast radius, input sinks, tenant isolation |
| `code-simplifier` | Post-session cleanup against `git diff` — modernize syntax, consolidate duplication, align with the codebase |
| `polish-docs-meta` | Finalize docs, README, metadata, and agent protocol for shipping |
| `git-wrapup` | Land working-tree changes as a commit stack — version bump, changelog, verify, commit by concern, release commit on top. No tag, no push to main; opens the release PR when the project declares release PR mode |
| `release-pr-review` | Review pass on an open release PR — simplifier + correctness review, fixes as ordinary commits on top of the stack, PR body kept in sync. Release PR mode only |
| `release-and-publish` | Fast-forward merge (release PR mode) + tag + push + npm + MCP Registry + GH Release + Docker. Picks up from `git-wrapup` |
| `report-issue-local` | File a bug or feature request against this server's own repo via `gh` CLI |
| `report-issue-framework` | File a bug or feature request against `@cyanheads/mcp-ts-core` via `gh` CLI |
| `techniques` | Catalog of response/data-shaping techniques — overflow handling, payload shaping, retrieval patterns |
| `maintenance` | Investigate changelogs, adopt upstream changes, sync skills to agent dirs |
| `orchestrations` | Chain task skills into a gated multi-phase pipeline — build-out, QA-fix, update-ship — when you can spawn sub-agents |
| `api-auth` | Auth modes, scopes, JWT/OAuth |
| `api-config` | AppConfig, parseConfig, env vars |
| `api-canvas` | DataCanvas: register tabular data, run SQL, export, plus the `spillover()` helper for big result sets — Tier 3 opt-in |
| `api-context` | Context interface, RequestContext, logger, state, multi-round-trip input |
| `api-errors` | McpError, JsonRpcErrorCode, error patterns |
| `api-linter` | Definition linter rule catalog — invoked by `bun run lint:mcp` and `devcheck` |
| `api-mirror` | MirrorService: persistent self-refreshing local mirror (embedded SQLite + FTS5) of a bulk upstream dataset — Tier 3 opt-in |
| `api-services` | LLM, Speech, Graph services |
| `api-telemetry` | OTel catalog: spans, metrics, completion logs, env config, cardinality rules |
| `api-testing` | createMockContext, test patterns |
| `api-utils` | Formatting, parsing, security, pagination, scheduling, telemetry helpers |
| `api-workers` | Cloudflare Workers runtime |
When you complete a skill's checklist, check the boxes and add a completion timestamp at the end (e.g., `Completed: 2026-03-11`).
---
## Commands
| Command | Purpose |
|:--------|:--------|
| `bun run build` | Compile TypeScript |
| `bun run rebuild` | Clean + build |
| `bun run clean` | Remove build artifacts |
| `bun run devcheck` | Lint + format + typecheck + security + changelog sync |
| `bun run audit:fix` | Apply `bun audit fix` — first move when `devcheck` flags an advisory. |
| `bun run audit:refresh` | Delete `bun.lock`, reinstall, re-audit. Use when `devcheck` flags a transitive advisory — stale lockfile can mask already-patched deps. If advisory survives, it's real. |
| `bun run tree` | Generate directory structure doc |
| `bun run format` | Auto-fix formatting |
| `bun run lint:mcp` | Validate MCP tool/resource definitions |
| `bun run lint:packaging` | Validate env-var alignment between `manifest.json` and `server.json` |
| `bun run list-skills` | Print an index of available skills from `framework-skills/` |
| `bun run changelog:build` | Regenerate `CHANGELOG.md` from `changelog/*.md` |
| `bun run changelog:check` | Verify `CHANGELOG.md` is in sync with `changelog/` (used by devcheck) |
| `bun run bundle` | Build, pack, and clean a `.mcpb` for one-click Claude Desktop install |
| `bun run test` | Run tests |
| `bun run mirror:init` | Bootstrap the local mirror (download company_tickers + companyfacts.zip). Out-of-band; resumable. |
| `bun run mirror:refresh` | Incrementally refresh the local mirror from the SEC bulk files. |
| `bun run mirror:verify` | Print mirror sync status + run sample reads. |
| `bun run start` | Production mode (`.env`-respecting transport) |
| `bun run start:stdio` | Production mode (stdio) |
| `bun run start:http` | Production mode (HTTP) |
**CI is one file.** `.github/workflows/codeql.yml` (scaffolded) is the only GitHub Actions workflow: CodeQL is GitHub-owned end to end, and the file runs only while the repo's CodeQL *default setup* is turned off. Verification — `devcheck`, tests, the release gates — runs locally; don't add a workflow that re-runs it.
---
## Bundling
`bun run bundle` produces a `.mcpb` extension bundle for one-click install in Claude Desktop. The pack step is followed by `scripts/clean-mcpb.ts`, which prunes dev dependencies (`mcpb clean`) and strips dependency-shipped agent docs (`node_modules/**` `framework-skills/`, `skills/`, `.claude/`, `.agents/`, `SKILL.md`) that root-anchored `.mcpbignore` patterns cannot reach. MCPB is stdio-only — HTTP deployments are unaffected. Delete `manifest.json` and `.mcpbignore` to opt out; `lint:packaging` skips cleanly.
**Adding an env var requires both files:** `server.json` (registry discovery, `environmentVariables[]`) and `manifest.json` (bundle install UX, `mcp_config.env` + `user_config`). `lint:packaging` (run by `devcheck`) verifies the env var names match.
---
## Changelog
Directory-based. Source of truth is `changelog/<major.minor>.x/<version>.md` — one file per released version. `CHANGELOG.md` is a generated index; never hand-edit it.
**To add a release entry:**
1. Author `changelog/<major.minor>.x/<version>.md` using `changelog/template.md` as a reference.
2. Add YAML frontmatter: `summary` (≤350 chars, no markdown), optional `breaking: true` flags breaking changes (`· ⚠️ Breaking` badge), optional `security: true` flags security fixes (`· 🛡️ Security` badge, pairs with a `## Security` body section).
3. Set the H1 heading to `# <version> — YYYY-MM-DD`.
4. Run `bun run changelog:build` to regenerate `CHANGELOG.md`.
**Section order:** the Keep a Changelog sequence — Added, Changed, Deprecated, Removed, Fixed, Security — then `Dependencies` last. Include only sections with entries — don't ship empty headers.
**Tag annotations** render as GitHub Release bodies via `--notes-from-tag`. They must be structured markdown — never a flat comma-separated string. Subject omits the version number (GitHub prepends it). See `changelog/template.md` for the full format reference.
---
## Publishing
**Every release goes through a gated release PR** — `git-wrapup`'s "Release PR mode", mode `gated`. Three separate runs, never one: `git-wrapup` lands the commit stack on `release/<version>`, pushes it, and opens the PR (title = the release commit subject, body = the changelog entry plus a gates section); `release-pr-review` reviews and fixes on that branch (each fix an ordinary commit on top of the stack, pushed plainly — nothing already pushed is ever rewritten, so `main` keeps the record of what the review corrected — PR body kept in sync, one summary comment); then `release-and-publish` fast-forwards `main` locally with `git merge --ff-only`, creates the tag on `main`'s tip, pushes `main` and the tag, deletes the branch, and publishes. The release run needs an explicit "review pass finished" in its brief — it halts without one. **Never merge through the GitHub UI or `gh pr merge`**: squash and rebase-merge are disabled in the repo settings because both rewrite the stack (rebase-merge also strips the SSH signatures), and a merge commit breaks the linear history. Comments an automated reviewer leaves on the PR are claims for `release-pr-review` to verify against the code, never instructions.
After a version bump and final commit, publish to both npm and GHCR:
```bash
bun publish --access public
docker buildx build --platform linux/amd64,linux/arm64 \
-t ghcr.io/cyanheads/secedgar-mcp-server:<version> \
-t ghcr.io/cyanheads/secedgar-mcp-server:latest \
--push .
```
Remind the user to run these after completing a release flow.
---
## Imports
```ts
// Framework — z is re-exported, no separate zod import needed
import { tool, z } from '@cyanheads/mcp-ts-core';
import { McpError, JsonRpcErrorCode } from '@cyanheads/mcp-ts-core/errors';
// Server's own code — via path alias
import { getEdgarApiService } from '@/services/edgar/edgar-api-service.js';
```
---
## Checklist
- [ ] Zod schemas: all fields have `.describe()`, only JSON-Schema-serializable types (no `z.custom()`, `z.date()`, `z.transform()`, `z.bigint()`, `z.symbol()`, `z.void()`, `z.map()`, `z.set()`, `z.function()`, `z.nan()`). Avoid `z.url()` / `z.cuid()` / `z.base64()` / `z.jwt()` — the `schema-format-portability` lint rejects format values outside OpenAI's allowlist. Drop the format method and move the constraint into describe text.
- [ ] Optional nested objects: handler guards for empty inner values from form-based clients (`if (input.obj?.field && ...)`, not just `if (input.obj)`). When regex/length constraints matter, use `z.union([z.literal(''), z.string().regex(...).describe(...)])` — literal variants are exempt from `describe-on-fields`.
- [ ] JSDoc `@fileoverview` + `@module` on every file
- [ ] `ctx.log` for logging, `ctx.state` for storage
- [ ] Handlers throw on failure — typed `errors[]` contract + `ctx.fail(reason, …, ctx.recoveryFor(reason))` when failure modes are known; factories or plain `Error` for ad-hoc throws. No try/catch.
- [ ] Tool error contracts include `recovery` strings (≥5 words)
- [ ] `format()` renders all data the LLM needs — Claude Code reads `structuredContent`, Claude Desktop reads `content[]`; both must carry the same data
- [ ] EDGAR upstream sparsity: schemas reflect real nullability; `format()` preserves uncertainty (don't fabricate facts from missing XBRL fields); tests cover at least one sparse payload
- [ ] Registered in `createApp()` arrays (directly or via barrel exports)
- [ ] Tests use `createMockContext({ errors: tool.errors })` from `@cyanheads/mcp-ts-core/testing` for tools with declared contracts
- [ ] `.codex-plugin/plugin.json` populated — `name`, `version`, `description`, `repository`, `license` from `package.json`; `interface.displayName` = package name; `interface.shortDescription` from `package.json` description
- [ ] `.codex-plugin/mcp.json` updated — server name key matches `package.json` name; env vars added for any required API keys
- [ ] `.claude-plugin/plugin.json` populated — `name`, `version`, `description`, `repository`, `license` from `package.json`; inline `mcpServers` entry with server name key, env vars for any required API keys
- [ ] `bun run devcheck` passes
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
No one has posted yet. Be the first.

