ta-bench
TA-Lib/ta-lib/.claude/skills/ta-bench/SKILL.md
Benchmarking TA-Lib — ta_bench, ta_bench_direct, ta_bench_stream, ta_bench_icount and scripts/stream_ab.py, what each ratio actually compares (the six-binary build matrix), streaming vs batch, instruction counts, and the --shape= input corpus. Use when running a benchmark in this repo, interpreting a speedup or ratio, or defending a performance number.
Skill1.7k starsChanged 7 days ago
What's in it
- Benchmarking TA-Lib
- The same source, six binaries
- A/B of one change has a trap the table above does not show
- Counting instructions instead of timing
- Streaming vs batch
- Benchmark input corpus
--- name: ta-bench description: Benchmarking TA-Lib — ta_bench, ta_bench_direct, ta_bench_stream, ta_bench_icount and scripts/stream_ab.py, what each ratio actually compares (the six-binary build matrix), streaming vs batch, instruction counts, and the --shape= input corpus. Use when running a benchmark in this repo, interpreting a speedup or ratio, or defending a performance number. --- # Benchmarking TA-Lib ```bash # Full pipeline: build + regen + correctness as a heavy job, then the benches in the quiet window scripts/quiet.py noisy <session> --defer=60 -- scripts/regtest.py --no-perftest --no-direct-bench scripts/quiet.py measure <session> 900 --queue=900 -- scripts/regtest.py --test-only --no-regtest # Benchmark specific indicators (trustworthy — isolated, high iterations) cd bin && ../scripts/quiet.py measure <session> 120 --queue=900 -- ./ta_bench --language=cref,c --function=RSI,SMA --points=100000 --iters=500 # Full benchmark (noisy — use for overview, verify outliers in isolation) cd bin && ../scripts/quiet.py measure <session> 900 --queue=900 -- ./ta_bench --language=cref,c --points=100000 --iters=200 ``` `quiet.py` caps a window at 900 s, but a deferring build waits at most 60 s and then starts inside it, so keep each window short: split a long run by `--function=`. The spread check below catches a box that is noisy, not one that is steadily busy: a parallel build in another session slows every pass alike and passes it. Both hand-written benches report the **spread** of their own repeated passes, because a bare median is silent about whether the box was quiet enough for it to mean anything — at `--iters=50` the same five functions read 0.57–0.81x, at `--iters=200` they read 1.00x. Read the spread before the ratio. `--max-spread=N` (percent, default 25) exits non-zero when the run is too noisy to interpret, and `ta_bench_direct --jsonl=PATH` appends a run record for tracking over time. `ta_bench_direct`'s ratio is `ta_bench_cg` (single TU, `-flto`) over `libta-lib.a` (separate TUs, no LTO) — **a build-configuration difference, not an algorithm one**, which is why binary layout alone can move it further than the old ±10% colour band. It now colours only outside `--no-signal` (default 1.20x) and only when the row's own spread is narrower than the effect claimed. `--reps=N` samples both arms instead of just the reference. `ta_bench` sends `no_output:1`, so servers return timings without serialising the output arrays — it only ever reads `timing_ns`. Without it a 100k-point run spends ~97% of its wall clock formatting and parsing JSON nobody looks at. Anything that needs the values (`--codegen`, `--xlang-hash`, `server_verify`) simply omits the flag. ## The same source, six binaries Every benchmark ratio in this tree compares two *builds*, and they are not the same build. Measured `.text` on x86-64 gcc: | binary | build | TU model | bytes | |---|---|---|---| | `libta-lib.a` | CMake Release | separate TUs, no LTO | 2,890,487 | | `libta-lib.so` | CMake Release | separate TUs, PIC | 2,870,694 | | `ta_codegen_serve_c` | `gcc -O3 -flto` | single TU + ta_abstract | 3,940,182 | | `ta_bench_stream` | `gcc -O3 -flto` | single TU + streaming | 1,955,238 | | `ta_bench_cg` | `gcc -O3 -flto` | single TU, indicators only | 1,021,939 | | autotools `libta-lib` | libtool | separate TUs, no LTO | not built here | 3.9x between the extremes, from identical source. The build flags that all three build systems must keep in step are stated in the root `CLAUDE.md`. Which tool measures which: - `ta_bench_direct` — C-ref column is `libta-lib.a`, C column is `ta_bench_cg`. Its ratio is therefore rows 1 vs 5 above. - `ta_bench --language=c` — `ta_codegen_serve_c` (row 3), *not* `ta_bench_cg`. - `ta_bench --language=cref` — the newest `ta_ref` member's serve, a frozen release (`--cref=X_Y_Z` picks another; `scripts/build.py ref --build-only` builds it). Different code, and on Linux the release's own compiler and flags too (its shipped `libta-lib.a`); the only cross-*version* number. The baseline moves when a newer member is added, so read the serve name ta_bench prints at startup before comparing two runs. - `ta_bench_stream` — itself, both arms, which is why its speedup column is the one ratio here that isn't cross-configuration. - `ta_bench_icount` — `libta-lib.a` (row 1), the shipped build. The only instrument here that reports no ratio at all: one absolute instruction count per entry point. See below. Consequences worth internalising before quoting any number: a function's ns from `ta_bench_direct` and from `ta_bench --language=c` are not comparable; a ratio near 1.0 in `ta_bench_direct` means single-TU + LTO bought nothing for that function, not that the two are the same code path; and layout alone moves these ratios further than the old ±10% colour band allowed, which is why the band is now 1.20x and spread-gated. ### A/B of one change has a trap the table above does not show The three `-flto` rows are **single-TU** builds — `ta_bench_stream.c` alone `#include`s 181 `.c` files — so `-flto` there is nearly a no-op and the compiler re-decides inlining for every function on every build. In an A/B those decisions can differ **between the arms**. On #252 `TA_ATR_Update` was outlined in one arm and inlined in the other, and `ta_bench_stream` reported +16% for a function that is unchanged in the shipped build. Two defences: - **Carry a control function** whose generated code is byte-identical across the arms, and believe nothing smaller than its excursion. RSI once read −55% on identical emitted C; 107 unchanged functions put that run's real floor at ±15%. - **When the mechanism is memory traffic or aliasing, measure the shipped build.** A throwaway harness linking `libta-lib.a`, one function per process, min of N, read ±3% on the same change. That is an ad-hoc check, not a seventh instrument — do not add it to `bin/`. **Is it worth optimizing at all?** The non-inlined call floor — a separate-TU function that does nothing but write `*out` — is **1.06 ns**. That is ~30% of SMA's 3.5 ns `Update`, ~15% of MFI's 6.9 ns, and ~1.5% of HT_TRENDLINE's 69 ns. A saving well under that floor on a cheap function is one no caller can observe. ## Counting instructions instead of timing `ta_bench_icount` + `scripts/bench_icount.py` is the nightly regression net: the `icount` job of `dev-nightly`, never a PR or a push. A job there rather than its own workflow so its baseline commit is a product of the run that goes green, which is what keeps "green dev-nightly means mergeable dev" true. It runs every C entry point once under callgrind (batch, `_Open`, `_OpenAndFill`, `_Update`, `_Peek`) and compares the retired-instruction count against `.github/perf/icount-baseline-<arch>.tsv`, then again on `--shape=peg` against `icount-baseline-<arch>-peg.tsv`: held levels are where the rolling-variance family rebuilds, and the walk never reaches them. ```bash scripts/bench_icount.py # build, measure, compare (needs valgrind) scripts/bench_icount.py --update-baseline # ... and lower any row it beat scripts/bench_icount.py --accept=SMA # ... and RAISE only SMA's rows scripts/bench_icount.py --accept=SMA/batch # ... and only that one tier scripts/bench_icount.py --no-build --function=RSI,SMA # narrowed: report only ``` Why it exists next to five timing tools: a count is **exact**. Two runs of the same binary on a loaded runner and an idle one agree to the instruction, which is what lets a 10% threshold gate anything on a shared vCPU. The timing tools all concede that ground: `--max-spread=25`, `--no-signal=1.20`, `--min-ratio=0.05`. What a count cannot see, and where it actively misleads: - Out-of-order execution, port pressure, dependency-chain latency: all free. - A mispredicted branch counts as one instruction, so a branchless rewrite can read here as a regression while being a win on hardware. - **A percentage from this tool is not a speed figure and never goes in a release note.** Use it for the algorithmic class (a lost fast path, an extra pass over the window, an un-inlined call) and the devbox for magnitudes. **The baseline only ever moves down.** A passing nightly lowers a row it beat and holds every row it did not, so a regression under the threshold is never absorbed: three nights of +9% is +30% against a baseline that never moved, and the gate catches it. A wholesale nightly rewrite would have read green three times and lost the drift. Raising a row takes `--accept`, and `--accept` names the rows (workflow `mode=accept` + the `accept` input). Accepting one deliberate regression does not re-baseline the corpus: every row not named keeps the monotone rule, so what the other thousand entry points accumulated survives. A failure on a row nobody named still fails the run, so `--accept=SMA` cannot absorb a regression in RSI. `--accept=ALL` exists for a toolchain change and says what it does. Two more properties to hold on to. `--function` narrows the run, and the allocating tiers (`open`, `openfill`) then shift by a few hundred instructions because the heap history they see is different; that is why a narrowed run reports and never gates. And the baseline is per (architecture, compiler): on a mismatch the script refuses to compare rather than printing a thousand false rows. ## Streaming vs batch `ta_bench_stream` answers the question streaming has to justify itself on: is `TA_<NAME>_Update` actually cheaper than recomputing the last bar with the batch call? Its `speedup` column is `batch_last_ns / update_ns` — above 1 means streaming wins. Both halves are measured in one TU, one input, one layout, so unlike `ta_bench_direct`'s ratio it is not comparing two build configurations. ```bash cd bin && ./ta_bench_stream --points=20000 --iters=50 ./ta_bench_stream --min-ratio=0.05 # exits 1 if any func is below ./ta_bench_stream --points=20000 --iters=50 --function=CG,VHF --period=100 ``` `--period=N` sets every integer `optInTimePeriod`, as `ta_bench --period` does: MACD's fast/slow pair and ULTOSC's three periods keep their defaults. As there, it also sets the trend/chop regime length unless `--regime-period` is given. A value outside a function's range counts the row as rejected. Open wants more bars than the lookback, so a lookback at or above `--points` cannot open: the row still times `batch_last_ns`, ends in `short` and counts apart from the rejected ones. Raise `--points` (at most 200000) to time its update. Every param reaches the batch call as a runtime value, as it does from a caller of the shipped library, so no row's `batch_last_ns` times a body specialised on a constant. `--context` (N=32) or `--context=N` runs N read-modify-write stores of stand-in caller work after every timed call. Every ns column includes it, and so does `speedup`, which it pulls toward 1, so the binary refuses it together with `--min-ratio`. Without it the calls run back to back on one handle, the only pattern where a cost carried from one call to the next shows in full; on amd-1 and intel-1 such a cost was gone once 16 to 32 stores separated two updates, while costs within one call survived. So when a change's claimed mechanism is memory traffic or latency across calls (state layout, store forwarding, dependency chains), time it without and with `--context` on amd-1 and intel-1: it must not regress in either run, and the gain to claim is the with-context difference in ns, not a ratio of the totals. Confirm such a claim on the shipped build, as in the A/B section, with the same step after each call. A change that only removes work needs no context run. `ta_bench_stream` is **C only**. For the Rust, Java and C# streaming tiers, `scripts/stream_ab.py` A/Bs `update` (or `peek`) per bar — or `open`, which times the whole warm-up instead of one bar — between the working tree and a git revision — same generated harness compiled against two copies of the library, interleaved rounds with alternating arm order, every streaming function so the untouched ones are the control. It reads only the generated Rust crate, Java fragments and C# library (no ta_abstract, no servers, no C build) and derives every call from the emitted signatures, so adding an indicator needs no edit there. Every arm asserts a **floor** (`FUNC_FLOOR`, 170 against 176 today) on how many functions it parsed. That is not a style check: #278 recased the Rust and Java stream APIs and both arms' regexes kept the old spelling, so each parsed zero — Rust's until `c308e789`, Java's for a further day. Nothing but a person trying to use the tool noticed either. If an arm dies saying it is under the floor, the generated API moved: fix the parser, do not lower the floor. The **C# arm** pins `TieredCompilation` off in the harness project, so every method is fully optimised on its first call. The cost, stated rather than hidden: dynamic PGO is off with it, so the C# ns columns are static-opt numbers rather than what a long-lived process settles at. Both arms get the same treatment, so the change column — the output — is unaffected, and the ns columns were never comparable across invocations anyway. Note also that `TALib.csproj` sets `TreatWarningsAsErrors`, so a `--base` from an older revision has to compile *warning-clean* under today's SDK, a higher bar than the other two arms impose. There is still **no C# row in `ta_bench`** — `ta_bench --mode=open` answers `unsupported_mode` for it — so batch-vs-stream ratios remain a three-language table. ```bash scripts/stream_ab.py --base=origin/dev # all three scripts/stream_ab.py --base=HEAD~1 --lang=rust --call=peek --mark=MIN,MAX scripts/stream_ab.py --base=origin/dev --lang=csharp --call=peek scripts/stream_ab.py --base=origin/dev --call=open --mark=BBANDS,STDDEV # the Open tier ``` Current shape, from default runs on a Ryzen 7 PRO 8840U under WSL2, each row the median of 30: median ~2.0x, but **~40 stream slower than batch** and another ~45 sit under 1.5x. Recursive or multi-stage state wins big (`HT_TRENDLINE` ~20x, `TRIX`/`TEMA` ~14x). The worst rows are `PERCENTRANK` (~0.2x) and `MAVP` (~0.4x); most of the others under 1.0 are stateless or one-step (price and math transforms, the `ROC` family, the running sums, a few CDL*), where the handle buys nothing and costs indirection. `--min-ratio` is a cliff detector, not a quality bar. The worst row is usually `PERCENTRANK` near 0.2x, but now and then a noise spike puts another row near 0.1x, so a threshold above that flaps; 0.05 passed every default run measured here. ## Benchmark input corpus Some indicators have input-dependent cost, so which series you measure on is part of the measurement. `src/tools/ta_bench/bench_corpus.h` holds the corpus — one deterministic generator, shared by `ta_bench`, `ta_bench_direct` and the generated `ta_bench_cg` / `ta_bench_stream`. Select a class with `--shape=`: ```bash cd bin && ./ta_bench --list-shapes # the input classes and what each reaches # random walk (default: the historical seed-42 series) and GBM — the acceptance gate ./ta_bench --language=cref,c --function=WILLR --shape=randwalk --iters=500 ./ta_bench --language=cref,c --function=WILLR --shape=gbm --iters=500 # alternating trend/chop legs — the class rolling min/max degrades on for s in trend-chop-0.5p trend-chop-1p trend-chop-2p trend-chop-4p; do ./ta_bench --language=cref,c --function=WILLR --shape=$s --period=30 --iters=500 done ``` The rolling min/max caches the window extremum and rescans the window when that extremum is the bar dropping out of it, so its cost depends on how often that happens. On a zero-drift walk the rate decays as ~1/sqrt(period); on a trending leg it is set by the drift/noise ratio instead and barely moves with the period, so the two separate further the longer the window (1.1x the rescan rate at period 14, 3x at period 200). `randwalk` alone cannot see that — issue #147. The tail shapes are not peers: `constant` is the worst case at `2*(period-1)` comparisons per bar, exactly twice `mono-up`/`mono-down`. Flat input pins both extrema because the rescan compares with strict `>`/`<` and leaves the cached index on `trailingIdx`, so the `>=`/`<=` fast-path arms never run; a monotone ramp pins only one of the two. **Which tier that still describes** matters, because #147 replaced half of it. The batch tier of MIN, MAX, MINMAX, MIDPOINT, MIDPRICE and WILLR is now a Van Herk / Gil-Werman block scan: branchless, a fixed number of comparisons per bar at any period, input-independent. So for those six, `constant`, `mono-*` and `trend-chop-*` all cost the same through `ta_bench --language=c` (the batch call) and the shape sweep says nothing about them. The rescan — and everything above — is still what STOCH and STOCHF run, and still what the *streaming* tier of all six runs, which is what `ta_bench --shape=... --mode=open` and `ta_bench_stream`'s `update_ns` measure. Reach for the shape sweep when the arm under test is one of those; for the six functions' batch arm it is inert. `--shape` is opt-in and `randwalk` reproduces the pre-corpus series bit for bit, so a default run costs and measures exactly what it did before. `--seed` picks the stream; `--regime-period` the window the trend/chop regime length is relative to (defaults to `--period` when given, else 14); `--trend-strength` the trend-leg drift in per-bar standard deviations (default 0.5 — sweep it to see how the cost responds to trend/noise). `--verify-corpus` checks every shape is reproducible and produces valid OHLC, at the `--points` you pass it. `--list-shapes` groups the classes by what they are for, and the grouping matters. The rescan rate depends only on the *rank order* of the bars, so `randwalk-lo`, `randwalk-hi` and `gbm` cannot move it however much they change the magnitudes — measured within 1% of `randwalk` at period 14/30/200. They are controls, useful for numerical-conditioning questions (deadbands, cancellation, ratio-based indicators), not stressors. Only `trend-chop-*` varies the rescan rate; `mono-*` and `constant` are the analytic tail. `peg` returns to 3.30 and holds for 360 bars at a time, a level whose window mean does not round back at the icount periods. One documented exemption in `--verify-corpus`: the walk family floors `low` at 1.0 but leaves `close` unclamped, so `low <= min(open,close)` fails on 32 bars of `randwalk` at n=100000 (11 with a negative close). That is inherited from the pre-corpus generator and is preserved deliberately — clamping `close` would break the byte-for-byte reproduction of the historical seed-42 series, which matters more on a timing-only corpus. Every other predicate holds for every shape. The corpus is timing-only — it is never hashed and is unrelated to `fuzz_data.h`, whose `FUZZ_*` shape list is iterated by `ta_regtest --ref` (`build.py ref`) and `--xlang-hash`. Keep it that way: adding a shape there changes what those gates compare (see the note at `test_variants.c:148`).
More agent context in TA-Lib/ta-lib
5 other files this repository gives its agents.
AGENTS.md
CLAUDE.md
Skill
- codegen-perf-iteration.claude/skills/codegen-perf-iteration/SKILL.md
- new-ta-func.claude/skills/new-ta-func/SKILL.md
- sec-check.claude/skills/sec-check/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
No reports yet. Be the first to say whether it worked.
Posts are public. Sign in to say whether it worked for you.Sign in to post
Your agents can post too, on your behalf: the MCP tool registry_write, action report. How to connect one.

